
OpenAI's Jalapeño posts 1.5-1.9x more inference work per watt than Nvidia Blackwell
OpenAI published Jalapeño's first benchmarks at Hot Chips: on SemiAnalysis' InferenceX across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, the Broadcom-built chip returned 1.5-1.9x more work per watt and 1.7-3.6x lower end-to-end latency than Nvidia Blackwell systems, with hardware lead Richard Ho calling it a significant advance over state of the art. The gains come from a full-stack design that keeps model state including the KV cache local to cut prefill and communication delays — but the Blackwell comparison ages fast, since Jalapeño ships in very small volumes at the end of 2026 and only reaches broad deployment in 2027.
Source: techcrunch.com ↗
Jalapeño can serve more AI work per unit of power, while also returning responses more quickly.
Why this matters
- → OpenAI's custom chip outperforms Nvidia's latest by 1.5–1.9× on inference efficiency.
- → Full-stack design eliminates KV-cache bottlenecks that slow response times.
- → Narrow 2026 deployment window means Nvidia has 12 months to iterate.