415.tech
AI & tech, from the frontlines of Silicon Valley
DeepSeek's MIT-licensed V4-Flash-0731 beats its larger V4-Pro preview on agentic benchmarks

DeepSeek's MIT-licensed V4-Flash-0731 beats its larger V4-Pro preview on agentic benchmarks

DeepSeek released V4-Flash-0731 under an MIT license — a sparse mixture-of-experts model with 13B active parameters and a 1M-token context that scores 82.7 on Terminal-Bench 2.1 and 76.7 on Cybergym, ahead of the bigger V4-Pro preview. The lift came from post-training alone, and at $0.14 per million input tokens and $0.28 per million output, weights competitive with the strongest proprietary models are now downloadable, speculative decoding module included.

Source: huggingface.co

Post on XEmail

DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.

DeepSeek

Why this matters

  • → Smaller model beats larger preview on agent benchmarks — efficiency gain from training alone
  • → MIT-licensed weights downloadable at $0.14/$0.28 per million tokens — competitive with proprietary leaders
  • → 1M-token context + speculative decoding enables local inference at scale
Efficiency beats scale