Tech Meridian ← ENTITY INDEX
RU

MODEL · ENTITY #2662

llama.cpp

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

MODELS · 1 SOURCE · NVIDIA Developer

TensorRT Edge-LLM runs Qwen3.6-27B on Jetson AGX Thor, completes MLPerf Edge Agentic 6.4× faster

NVIDIA's TensorRT Edge-LLM ran Qwen3.6-27B on a single Jetson AGX Thor Developer Kit and achieved 52.33 tokens/sec in the MLPerf Inference v6.1 Edge Agentic performance workload, completing all 1,007 turns in 24 minutes 36 seconds — 6.4× faster than the llama.cpp Jetson reference (2h37m). The submission used NVFP4 quantization for weights/activations, FP8 for the KV cache, tree-based multi-token prediction, and KV-cache/recurrent-state reuse to accelerate long-context agent decoding.

7.0

MODELS · 1 SOURCE · Mistral AI

Mistral releases Small 4: 119B MoE multimodal model with 256k context (Apache 2.0)

Mistral announced Mistral Small 4, a 119B-parameter hybrid Mixture-of-Experts model (128 experts, 4 active) that accepts text and image inputs, offers a 256k context window, and includes a configurable reasoning_effort parameter; it is released under the Apache 2.0 license. The company says Small 4 unifies capabilities from its Magistral, Pixtral, and Devstral lines, targets chat, coding/agentic, and complex-reasoning use cases, claims substantial latency and throughput gains versus Mistral Small 3, and is available across vLLM, llama.cpp, SGLang, Transformers and other runtimes.

9.0