Tech Meridian ← ENTITY INDEX
RU

TOPIC · ENTITY #835

vLLM

Related event timeline, sources and context from the news index.

EVENT TIMELINE

5

COMPANIES · 1 SOURCE · NVIDIA Developer

NVIDIA releases AIPerf, a multiprocess LLM benchmarking tool

NVIDIA AIPerf replaces GenAI-Perf with a ground-up multiprocess architecture designed to avoid client-side bottlenecks for high-concurrency LLM inference benchmarking. It supports 15+ endpoint types, public datasets and trace replay formats (ShareGPT, Mooncake, Baseten, WEKA/AgentX), configurable arrival patterns, and reports TTFT, ITL, latency percentiles and GPU telemetry when DCGM or pynvml are available.

6.0

CODING · 1 SOURCE · AWS Machine Learning

Hugging Face publishes six open-source Skills to deploy models on Amazon SageMaker AI via coding agents

Hugging Face published six open-source 'Skills' (GitHub) that let coding agents orchestrate end-to-end deployment of Hugging Face models to Amazon SageMaker AI. The skills automate selecting the correct serving container from AWS Deep Learning Containers, creating real-time or serverless endpoints with autoscaling and CloudWatch alarms, and provide verified teardown paths; they run using Python and the AWS CLI and are designed to prevent fragile or costly agent-made deployment mistakes.

6.0

CODING · 1 SOURCE · Cohere

Cohere publishes megakernel serving engine for North Mini Code with faster H100 decoding

Cohere describes a megakernel-based serving engine for its North Mini Code 30B model that runs BF16 on a single NVIDIA H100 and claims 1.25×–1.41× end-to-end speedup over vLLM, with a reported 292 tok/s (62% of SoL) at batch size 1 — about 1.58× faster than vLLM. The system supports production features (continuous batching, paged attention, ragged sequences), an OpenAI-compatible endpoint with tool calling, is implemented as a single CUDA file, and the code is available on GitHub.

7.0

MODELS · 1 SOURCE · AWS Machine Learning

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

AWS Machine Learning provides a walkthrough for deploying Qwen3.8-2.4T-A95B, a 2.4‑trillion‑parameter open‑weight model, on Amazon SageMaker HyperPod using vLLM. The guide covers cluster provisioning, NVFP4 quantization, and creating an OpenAI‑compatible endpoint with built‑in reasoning, tool calling, and native MTP speculative decoding.

7.0

MODELS · 1 SOURCE · Mistral AI

Mistral releases Small 4: 119B MoE multimodal model with 256k context (Apache 2.0)

Mistral announced Mistral Small 4, a 119B-parameter hybrid Mixture-of-Experts model (128 experts, 4 active) that accepts text and image inputs, offers a 256k context window, and includes a configurable reasoning_effort parameter; it is released under the Apache 2.0 license. The company says Small 4 unifies capabilities from its Magistral, Pixtral, and Devstral lines, targets chat, coding/agentic, and complex-reasoning use cases, claims substantial latency and throughput gains versus Mistral Small 3, and is available across vLLM, llama.cpp, SGLang, Transformers and other runtimes.

9.0