RELEASE · MODELS · #505
DeepSeek launches DeepSeek‑V4.1‑Flash with new causal encoder–decoder and lower costs
DeepSeek announced DeepSeek‑V4.1‑Flash, a smaller-native-vision model using a new Causal Encoder–Decoder design (8B active params for input, 16B for output), new pretraining and larger-scale RL post-training, plus KV‑cache compression. The company is retiring V4‑Flash variants, routing deepseek-v4-pro traffic to V4.1‑Flash from 04:00 UTC on Sept 14, 2026, and changing pricing (new rates take effect 04:00 UTC on Sept 10, 2026) while partners WorkBuddy, CodeBuddy and OpenCode already support V4.1‑Flash.
KEY POINTS
- DeepSeek announced DeepSeek‑V4.1‑Flash, a smaller-native-vision model using a new Causal Encoder–Decoder design (8B active params for input, 16B for output), new pretraining and larger-scale RL post-training, plus KV‑cache compression.
- The company is retiring V4‑Flash variants, routing deepseek-v4-pro traffic to V4.1‑Flash from 04:00 UTC on Sept 14, 2026, and changing pricing (new rates take effect 04:00 UTC on Sept 10, 2026) while partners WorkBuddy, CodeBuddy and OpenCode already support V4.1‑Flash.
- This release claims better performance and much lower inference costs while changing routing and pricing, so it affects deployment plans, costs, and compatibility for DeepSeek users and partners.
WHY IT MATTERS
This release claims better performance and much lower inference costs while changing routing and pricing, so it affects deployment plans, costs, and compatibility for DeepSeek users and partners.