
OpenAI's Jalapeño beats Nvidia GB200/GB300 on performance per watt in first benchmarks
OpenAI published first benchmark numbers for Jalapeño, its custom inference chip, claiming 1.5-1.9x more work per watt, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x higher performance on highly interactive workloads against Nvidia GB200/GB300 systems, drawing 550W or less on a 700W rating. Deployment into OpenAI's own infrastructure begins by year-end with a second generation underway, which moves inference economics for at least one frontier lab off Nvidia's price list — packaging and foundry capacity remain the hard bottleneck.
Source: latent.space ↗
Jalapeño delivered 1.5–1.9× more work per watt at peak throughput and 1.7–3.6× lower end-to-end latency, with 2.1–4.1× higher performance for highly interactive workloads
OpenAI, Jalapeño benchmark results
Why this matters
- → OpenAI's custom chip outperforms Nvidia's fastest inference systems on efficiency and latency
- → Inference economics may shift away from Nvidia's price list for frontier labs
- → Model-assisted kernel optimization is merging compiler work into the model improvement loop
Nvidia's grip loosens