
OpenAI's Jalapeño rack hits 1.7 exaFLOPS, beating GB300 on throughput per kilowatt
OpenAI detailed Jalapeño at Hot Chips: a Broadcom-co-designed inference accelerator whose 128-chip rack delivers 1.7 exaFLOPS of 4-bit compute, 27.5TB of HBM4, and roughly 2PB/s of memory bandwidth, with claimed gains of 1.5x-1.9x throughput per kilowatt and 1.7x-3.6x lower end-to-end latency versus Nvidia's GB200 and GB300 racks. The numbers are vendor-reported and the silicon is inference-only, but a frontier lab shipping its own accelerator in volume in 2027 marks the point where OpenAI's largest cost line stops being fully controlled by Nvidia.
Source: theregister.com ↗