NEWS · MODELS · #428
Mistral releases Small 4: 119B MoE multimodal model with 256k context (Apache 2.0)
Mistral announced Mistral Small 4, a 119B-parameter hybrid Mixture-of-Experts model (128 experts, 4 active) that accepts text and image inputs, offers a 256k context window, and includes a configurable reasoning_effort parameter; it is released under the Apache 2.0 license. The company says Small 4 unifies capabilities from its Magistral, Pixtral, and Devstral lines, targets chat, coding/agentic, and complex-reasoning use cases, claims substantial latency and throughput gains versus Mistral Small 3, and is available across vLLM, llama.cpp, SGLang, Transformers and other runtimes.
KEY POINTS
- Mistral announced Mistral Small 4, a 119B-parameter hybrid Mixture-of-Experts model (128 experts, 4 active) that accepts text and image inputs, offers a 256k context window, and includes a configurable reasoning_effort parameter; it is released under the Apache 2.0 license.
- The company says Small 4 unifies capabilities from its Magistral, Pixtral, and Devstral lines, targets chat, coding/agentic, and complex-reasoning use cases, claims substantial latency and throughput gains versus Mistral Small 3, and is available across vLLM, llama.cpp, SGLang, Transformers and other runtimes.
- This matters because an open-source, unified MoE model with very long context and configurable reasoning could simplify deployments, lower inference costs, and broaden access to powerful multimodal and reasoning-capable models.
WHY IT MATTERS
This matters because an open-source, unified MoE model with very long context and configurable reasoning could simplify deployments, lower inference costs, and broaden access to powerful multimodal and reasoning-capable models.