Tech Meridian ← ENTITY INDEX
RU

MODEL · ENTITY #607

DeepSeek

Related event timeline, sources and context from the news index.

EVENT TIMELINE

8

RESEARCH · 1 SOURCE · arXiv cs.AI

Characterizing web search behavior of conversational LLM agents across four platforms

arXiv:2609.19244v1 reports the first study of agentic Web search across four conversational platforms (ChatGPT, Claude, Grok, DeepSeek), combining real-world user interactions (in vivo) with controlled API experiments (in vitro). The paper analyzes when agents choose to invoke Web search, their query strategies, domain preferences in returned results, and how they transform results into grounded responses, finding substantial variability across platforms, platform-specific result biases, and some reliance on uncited search results.

7.0

MODELS · 1 SOURCE · The Decoder

OpenRouter shows 25,000% spike in weekly token consumption since Jan 2025

The Decoder reports that OpenRouter’s weekly token consumption rose from about 0.5 trillion tokens in January 2025 to roughly 126.2 trillion tokens, a >25,000% increase. The piece cautions this mostly reflects inflated token metrics — e.g., reasoning/agentic models generating large 'thinking' token volumes and GPT-5.6 Luna producing many tokens per prompt — rather than a commensurate jump in end-user adoption or business value; OpenAI’s Astra leads revenue and some Chinese models (Kimi, GLM, DeepSeek) show fast growth from a small base.

6.0

MODELS · 1 SOURCE · NVIDIA Developer

Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine

NVIDIA Developer describes how to use the NVIDIA Transformer Engine to accelerate dropless Mixture‑of‑Experts (MoE) training workloads in JAX, outlining implementation details and considerations for integrating the engine with MoE models. The article situates this work amid recent MoE models such as DeepSeek, Qwen, and Mixtral and discusses practical steps to improve training efficiency.

6.0

MODELS · 1 SOURCE · DeepSeek

DeepSeek launches DeepSeek‑V4.1‑Flash with new causal encoder–decoder and lower costs

DeepSeek announced DeepSeek‑V4.1‑Flash, a smaller-native-vision model using a new Causal Encoder–Decoder design (8B active params for input, 16B for output), new pretraining and larger-scale RL post-training, plus KV‑cache compression. The company is retiring V4‑Flash variants, routing deepseek-v4-pro traffic to V4.1‑Flash from 04:00 UTC on Sept 14, 2026, and changing pricing (new rates take effect 04:00 UTC on Sept 10, 2026) while partners WorkBuddy, CodeBuddy and OpenCode already support V4.1‑Flash.

8.0

MODELS · 1 SOURCE · DeepSeek

DeepSeek publishes open-source DeepSeek‑V4 preview with 1M-token context

DeepSeek has open-sourced a preview of DeepSeek‑V4, offering two variants: DeepSeek‑V4‑Pro (1.6T total / 49B active params) and DeepSeek‑V4‑Flash (284B total / 13B active params). The company says both models support a default 1M-token context, introduce novel token-wise compression and DSA (DeepSeek Sparse Attention), claim open-source state‑of‑the‑art agentic and reasoning performance (trailing only Gemini‑3.1‑Pro among models they compare to), and are available now via chat.deepseek.com and updated APIs with compatibility for OpenAI ChatCompletions & Anthropic endpoints; older models deepseek-chat and deepseek-reasoner will be retired on Jul 24, 2026 at 15:59 UTC.

8.0

MODELS · 1 SOURCE · DeepSeek

DeepSeek launches experimental model V3.2-Exp with sparse attention and big API price cuts

DeepSeek introduced DeepSeek-V3.2-Exp, an experimental model built on V3.1-Terminus that debuts DeepSeek Sparse Attention (DSA) to improve long-context efficiency with minimal quality loss; published benchmarks show parity with V3.1-Terminus. DeepSeek also cut DeepSeek API prices by over 50% effective immediately, provides TileLang and CUDA GPU kernels, and keeps V3.1-Terminus available via a temporary API until Oct 15, 2025, 15:59 UTC for comparison testing.

7.0

MODELS · 1 SOURCE · DeepSeek

DeepSeek-V3.1 renamed and released as DeepSeek-V3.1-Terminus

The DeepSeek team has released DeepSeek‑V3.1‑Terminus, an update that builds on V3.1 while addressing user feedback. Changes include improved language consistency (fewer CN/EN mix-ups and elimination of stray characters), stronger Code Agent and Search Agent performance, and reportedly more stable, more reliable outputs across benchmarks; open-source weights are available on Hugging Face.

6.0

MODELS · 1 SOURCE · DeepSeek

DeepSeek releases DeepSeek‑V3.1 with hybrid 'Think & Non‑Think' agent mode and open-source weights

DeepSeek released DeepSeek‑V3.1, introducing a hybrid inference mode called Think & Non‑Think (toggleable via the “DeepThink” button at chat.deepseek.com), faster 'Think' responses vs. DeepSeek‑R1‑0528, improved tool use and multi-step agent skills, Anthropic API format support, and beta Strict Function Calling support. V3.1 Base received 840B tokens of continued pretraining for long-context extension; the tokenizer and chat template were updated (tokenizer_config.json); open-source weights for DeepSeek‑V3.1 and DeepSeek‑V3.1‑Base are published on Hugging Face, and new pricing/off-peak discount changes take effect Sep 5, 2025, 16:00 UTC.

7.0