RELEASE · MODELS · #503
DeepSeek launches experimental model V3.2-Exp with sparse attention and big API price cuts
DeepSeek introduced DeepSeek-V3.2-Exp, an experimental model built on V3.1-Terminus that debuts DeepSeek Sparse Attention (DSA) to improve long-context efficiency with minimal quality loss; published benchmarks show parity with V3.1-Terminus. DeepSeek also cut DeepSeek API prices by over 50% effective immediately, provides TileLang and CUDA GPU kernels, and keeps V3.1-Terminus available via a temporary API until Oct 15, 2025, 15:59 UTC for comparison testing.
KEY POINTS
- DeepSeek introduced DeepSeek-V3.2-Exp, an experimental model built on V3.1-Terminus that debuts DeepSeek Sparse Attention (DSA) to improve long-context efficiency with minimal quality loss; published benchmarks show parity with V3.1-Terminus.
- DeepSeek also cut DeepSeek API prices by over 50% effective immediately, provides TileLang and CUDA GPU kernels, and keeps V3.1-Terminus available via a temporary API until Oct 15, 2025, 15:59 UTC for comparison testing.
- DSA and accompanying kernel/tools aim to reduce compute and improve long-context performance while an immediate >50% API price cut changes user cost dynamics.
WHY IT MATTERS
DSA and accompanying kernel/tools aim to reduce compute and improve long-context performance while an immediate >50% API price cut changes user cost dynamics.