
DeepSeek's V4.1-Flash beats its own V4-Pro flagship while shrinking KV cache 4x
DeepSeek released V4.1-Flash, a 552B-parameter MoE under an MIT license that activates 8B parameters for input and 16B for generation, and scores 74.2 on DeepSWE v1.1 against V4-Flash's 54.4. The real signal is economics — its KV cache runs 4x smaller than V4-Flash and 437x smaller than V1, cutting HBM to a quarter — and DeepSeek is retiring V4-Pro by routing that traffic to the cheaper Flash model on Sept 14. It still trails V4-Pro on knowledge benchmarks like SimpleQA-Verified, 42.3 to 55.2, so the trade is cost and speed over deep recall.
Source: deepseek.com ↗
V4.1-Flash is now live on the DeepSeek API with native multimodal support.
DeepSeek
Why this matters
- → DeepSeek replaces flagship with cheaper model, signals aggressive cost-performance optimization
- → 4x smaller KV cache cuts inference costs and HBM requirements dramatically
- → Architectural breakthrough (causal encoder-decoder) enables smaller models to outperform larger ones
The $0.01 flagship