
Inception ships Mercury 2.5, a diffusion LLM at 1,107 tokens/sec with a 260K context
Inception released Mercury 2.5, a diffusion language model hitting 1,107 tokens per second on widely available Nvidia GPUs with a 260K-token context, listing at $0.20 and $0.75 per million input and output tokens, discounted 80% at launch. The production numbers carry more weight than the claimed 40% intelligence gain: Augment Code cut context-compaction latency 82% — roughly 150 seconds to 27 — and cost 90%, while OpenCall's P99 voice response fell from several minutes to about one second. Diffusion models are now fast and cheap enough to absorb the high-frequency supporting calls inside search, voice, and coding agents, where per-call latency compounds.
Source: inceptionlabs.ai ↗
After we switched to Mercury, our P99 response time dropped from several minutes to just one second, and our P50 dropped from 0.4 seconds to under 0.2.
Why this matters
- → Sub-200ms latency in voice agents eliminates multi-second pauses.
- → 82% latency cut in context compaction unlocks long-context coding workflows.
- → 80% cheaper than frontier models at launch; production-proven across search, voice, coding.