Tech Meridian ← ENTITY INDEX
RU

COMPANY · ENTITY #714

KV cache

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

COMPANIES · 1 SOURCE · AWS Machine Learning

Amazon SageMaker Inference adds prefix-aware routing to reduce LLM latency

Amazon SageMaker Inference now offers prefix-aware routing, a strategy that sends requests sharing the same prompt prefix to the same instance so the KV cache stays warm. In benchmarks on Llama 3.1 70B, this reduced P50 time-to-first-token by up to 77% and increased KV cache hit rates from about 25% to over 80%.

7.0

MODELS · 1 SOURCE · DeepSeek

DeepSeek launches DeepSeek‑V4.1‑Flash with new causal encoder–decoder and lower costs

DeepSeek announced DeepSeek‑V4.1‑Flash, a smaller-native-vision model using a new Causal Encoder–Decoder design (8B active params for input, 16B for output), new pretraining and larger-scale RL post-training, plus KV‑cache compression. The company is retiring V4‑Flash variants, routing deepseek-v4-pro traffic to V4.1‑Flash from 04:00 UTC on Sept 14, 2026, and changing pricing (new rates take effect 04:00 UTC on Sept 10, 2026) while partners WorkBuddy, CodeBuddy and OpenCode already support V4.1‑Flash.

8.0