Tech Meridian ← ENTITY INDEX
RU

TOPIC · ENTITY #176

dense models

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

MODELS · 1 SOURCE · NVIDIA Developer

Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose Each

An NVIDIA Developer post compares dense and Mixture-of-Experts (MoE) architectures, showing how a 30B-parameter model can activate only about 3B parameters per token and discussing the resulting capacity and throughput trade-offs; Nemotron 3.5 Lightning is used as an illustrative example.

6.0

MODELS · 1 SOURCE · Mistral AI

Mistral AI introduces Forge to let enterprises train models on proprietary data

Mistral AI launched Forge, a system for enterprises to train frontier-grade AI models grounded in their proprietary documentation, codebases, structured records and policies. Forge supports pre-training, post-training, reinforcement learning, dense and MoE architectures, multimodal inputs, and is already being used in partnerships with organizations such as ASML, DSO National Laboratories Singapore, Ericsson, ESA, HTX Singapore and Reply.

7.0