
OpenAI and Broadcom unveil Jalapeño — first custom inference ASIC reporting ~50% lower cost than GPU clusters
OpenAI and Broadcom unveiled Jalapeño, OpenAI's first custom inference ASIC — TSMC 3nm, 8 HBM stacks, built in nine months — with early tests showing inference costs roughly 50% below GPU clusters. If the cost figure holds at production scale, it marks a structural shift in OpenAI's unit economics and the start of a real reduction in NVIDIA GPU dependence for its production inference workloads.
Source: openai.com ↗