415.tech
AI & tech, from the frontlines of Silicon Valley
OpenAI's Jalapeño beats Nvidia GB200/GB300 on performance per watt in first benchmarks

OpenAI's Jalapeño beats Nvidia GB200/GB300 on performance per watt in first benchmarks

OpenAI published first benchmark numbers for Jalapeño, its custom inference chip, claiming 1.5-1.9x more work per watt, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x higher performance on highly interactive workloads against Nvidia GB200/GB300 systems, drawing 550W or less on a 700W rating. Deployment into OpenAI's own infrastructure begins by year-end with a second generation underway, which moves inference economics for at least one frontier lab off Nvidia's price list — packaging and foundry capacity remain the hard bottleneck.

Source: latent.space

Post on XEmail

Jalapeño delivered 1.5–1.9× more work per watt at peak throughput and 1.7–3.6× lower end-to-end latency, with 2.1–4.1× higher performance for highly interactive workloads

OpenAI, Jalapeño benchmark results

Why this matters

  • → OpenAI's custom chip outperforms Nvidia's fastest inference systems on efficiency and latency
  • → Inference economics may shift away from Nvidia's price list for frontier labs
  • → Model-assisted kernel optimization is merging compiler work into the model improvement loop
Nvidia's grip loosens
Also in this edition