415.tech
AI & tech, from the frontlines of Silicon Valley
OpenAI's Jalapeño posts 1.5-1.9x more inference work per watt than Nvidia Blackwell

OpenAI's Jalapeño posts 1.5-1.9x more inference work per watt than Nvidia Blackwell

OpenAI published Jalapeño's first benchmarks at Hot Chips: on SemiAnalysis' InferenceX across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, the Broadcom-built chip returned 1.5-1.9x more work per watt and 1.7-3.6x lower end-to-end latency than Nvidia Blackwell systems, with hardware lead Richard Ho calling it a significant advance over state of the art. The gains come from a full-stack design that keeps model state including the KV cache local to cut prefill and communication delays — but the Blackwell comparison ages fast, since Jalapeño ships in very small volumes at the end of 2026 and only reaches broad deployment in 2027.

Source: techcrunch.com

Post on XEmail

Jalapeño can serve more AI work per unit of power, while also returning responses more quickly.

Richard Ho, OpenAI head of hardware

Why this matters

  • → OpenAI's custom chip outperforms Nvidia's latest by 1.5–1.9× on inference efficiency.
  • → Full-stack design eliminates KV-cache bottlenecks that slow response times.
  • → Narrow 2026 deployment window means Nvidia has 12 months to iterate.
Custom silicon strikes back
Also in this edition