415.tech
AI & tech, from the frontlines of Silicon Valley
Inception ships Mercury 2.5, a diffusion LLM at 1,107 tokens/sec with a 260K context

Inception ships Mercury 2.5, a diffusion LLM at 1,107 tokens/sec with a 260K context

Inception released Mercury 2.5, a diffusion language model hitting 1,107 tokens per second on widely available Nvidia GPUs with a 260K-token context, listing at $0.20 and $0.75 per million input and output tokens, discounted 80% at launch. The production numbers carry more weight than the claimed 40% intelligence gain: Augment Code cut context-compaction latency 82% — roughly 150 seconds to 27 — and cost 90%, while OpenCall's P99 voice response fell from several minutes to about one second. Diffusion models are now fast and cheap enough to absorb the high-frequency supporting calls inside search, voice, and coding agents, where per-call latency compounds.

Source: inceptionlabs.ai

Post on XEmail

After we switched to Mercury, our P99 response time dropped from several minutes to just one second, and our P50 dropped from 0.4 seconds to under 0.2.

Oliver Silverstein, Co-founder and CEO, OpenCall

Why this matters

  • → Sub-200ms latency in voice agents eliminates multi-second pauses.
  • → 82% latency cut in context compaction unlocks long-context coding workflows.
  • → 80% cheaper than frontier models at launch; production-proven across search, voice, coding.
Speed meets scale
Also in this edition