415.tech
AI & tech, from the frontlines of Silicon Valley
OpenAI previews Ultrafast, running GPT-5.6 Sol at 14x speed on Cerebras chips

OpenAI previews Ultrafast, running GPT-5.6 Sol at 14x speed on Cerebras chips

OpenAI is previewing Ultrafast, a tier that pushes GPT-5.6 Sol to roughly 750 output tokens per second — 14x standard — on Cerebras wafer-scale hardware, breaking the usual trade of frontier quality for a smaller, faster model. Real-time latency at full model strength opens workflows that previously could not wait for one: incident response, live customer support, financial market analysis, and e-commerce. Access is limited to a small set of customers until capacity grows, so the capability is a signal about where inference silicon is heading more than something buyable today.

Source: techcrunch.com

Post on XEmail

Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.

OpenAI

Why this matters

  • → Frontier model speed at full capability enables real-time workflows previously impossible
  • → 14x speedup signals where inference silicon and model deployment are heading
  • → Breaks the speed-vs-quality tradeoff that has defined model tiers
Speed meets scale