
DeepSeek's V4-Flash API enters public beta with DeepSWE up from 7.3 to 54.4
DeepSeek put its V4-Flash-0731 build into public beta with DeepSWE jumping from 7.3 to 54.4 and Terminal Bench 2.1 from 61.8 to 82.7 — same architecture, same size, only re-post-trained. That makes the agentic gain a post-training result rather than a scale result, and the build adds Responses API support and Codex adaptation for agent tooling. The V4-Pro API and the models behind DeepSeek's app and website are unchanged.
Source: technode.com ↗
The company says the model scored 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE, among other benchmarks.
DeepSeek
Why this matters
- → DeepSeek's post-training boost (agentic score 7.3→54.4) proves gains don't require scale.
- → V4-Flash beta enables agent tooling via Responses API + Codex adaptation.
- → Same model size, higher capability — efficiency gains matter for cost-conscious deployments.
Post-training beats scale