OpenAI and Broadcom reveal Jalapeño chip
OpenAI's first custom inference chip, an ASIC built with Broadcom, targets cheaper LLM serving, with large-scale deployment set for late 2026.
OpenAI and Broadcom unveiled Jalapeño on June 24, OpenAI’s first custom chip and its first real move into hardware after years spent on models and products. It is an application-specific chip built only for large language model inference, the work of running a model once it is trained, rather than the general-purpose silicon a GPU offers.
The headline is speed and efficiency. OpenAI says the design went from schematics to tape-out in about nine months, which it calls the fastest high-performance ASIC cycle it knows of, and credits using its own models to accelerate the engineering. It claims performance per watt “substantially better” than current state-of-the-art hardware. Worth noting honestly: those are self-reported figures OpenAI says are not yet final.
Broadcom supplies the silicon and networking, while Celestica builds the boards and racks. Large-scale deployment is planned for late 2026 at gigawatt scale, and Microsoft is expected to buy about 40 percent of the first batch.
What it means for you
Nothing here is a product you can buy. The signal is strategic. OpenAI joining Google and Amazon in designing its own inference chips is a push to control cost and supply, and a direct shot at Nvidia. Inference economics are the whole game now that OpenAI’s serving costs run into the billions. If Jalapeño works, the payoff reaches you as cheaper or more available model serving down the line, not a chip on your desk. It is one more reason to choose AI tools you can actually swap out rather than betting your stack on today’s prices.