
DeepSeek's MIT-licensed V4-Flash-0731 beats its larger V4-Pro preview on agentic benchmarks
DeepSeek released V4-Flash-0731 under an MIT license — a sparse mixture-of-experts model with 13B active parameters and a 1M-token context that scores 82.7 on Terminal-Bench 2.1 and 76.7 on Cybergym, ahead of the bigger V4-Pro preview. The lift came from post-training alone, and at $0.14 per million input tokens and $0.28 per million output, weights competitive with the strongest proprietary models are now downloadable, speculative decoding module included.
Source: huggingface.co ↗
DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.
DeepSeek
Why this matters
- → Smaller model beats larger preview on agent benchmarks — efficiency gain from training alone
- → MIT-licensed weights downloadable at $0.14/$0.28 per million tokens — competitive with proprietary leaders
- → 1M-token context + speculative decoding enables local inference at scale
Efficiency beats scale