415.tech
AI & tech, from the frontlines of Silicon Valley
DeepSeek's V4.1-Flash beats its own V4-Pro flagship while shrinking KV cache 4x

DeepSeek's V4.1-Flash beats its own V4-Pro flagship while shrinking KV cache 4x

DeepSeek released V4.1-Flash, a 552B-parameter MoE under an MIT license that activates 8B parameters for input and 16B for generation, and scores 74.2 on DeepSWE v1.1 against V4-Flash's 54.4. The real signal is economics — its KV cache runs 4x smaller than V4-Flash and 437x smaller than V1, cutting HBM to a quarter — and DeepSeek is retiring V4-Pro by routing that traffic to the cheaper Flash model on Sept 14. It still trails V4-Pro on knowledge benchmarks like SimpleQA-Verified, 42.3 to 55.2, so the trade is cost and speed over deep recall.

Source: deepseek.com

Post on XEmail

V4.1-Flash is now live on the DeepSeek API with native multimodal support.

DeepSeek

Why this matters

  • → DeepSeek replaces flagship with cheaper model, signals aggressive cost-performance optimization
  • → 4x smaller KV cache cuts inference costs and HBM requirements dramatically
  • → Architectural breakthrough (causal encoder-decoder) enables smaller models to outperform larger ones
The $0.01 flagship
Also in this edition