OpenAI and Broadcom reveal Jalapeño chip
Archive item — written before sources were shown.
OpenAI's first custom inference chip, an ASIC built with Broadcom, targets cheaper LLM serving, with large-scale deployment set for late 2026.
OpenAI and Broadcom unveiled Jalapeño on June 24, OpenAI’s first custom chip and its first real move into hardware after years spent on models and products. It is an application-specific chip built only for large language model inference, the work of running a model once it is trained, rather than the general-purpose silicon a GPU offers.
The headline is speed and efficiency. OpenAI says the design went from schematics to tape-out in about nine months, which it calls the fastest high-performance ASIC cycle it knows of, and credits using its own models to accelerate the engineering. It claims performance per watt “substantially better” than current state-of-the-art hardware. Worth noting honestly: those are self-reported figures OpenAI says are not yet final.
Broadcom supplies the silicon and networking, while Celestica builds the boards and racks. Large-scale deployment is planned for late 2026 at gigawatt scale, and Microsoft is expected to buy about 40 percent of the first batch.
What it means for you
Nothing here is a product you can buy. The signal is strategic. OpenAI joining Google and Amazon in designing its own inference chips is a push to control cost and supply, and a direct shot at Nvidia. Inference economics are the whole game now that OpenAI’s serving costs run into the billions. If Jalapeño works, the payoff reaches you as cheaper or more available model serving down the line, not a chip on your desk. It is one more reason to choose AI tools you can actually swap out rather than betting your stack on today’s prices.
Update, August 25: the first real numbers
Two months after the reveal, OpenAI published its first measured results, and the headline claim moved from “should be faster” to actual benchmark numbers across three different models: its own GPT-OSS 120B, DeepSeek’s R1 670B, and Moonshot’s Kimi K2.5 1T. Testing the same chip on models it did not build is the more interesting part, since it argues the design is a general inference platform rather than a one-model showcase.
Across all three models, OpenAI says Jalapeño delivered 1.5 to 3.6 times better results than the comparison hardware depending on the metric: at peak throughput it squeezes 1.5 to 1.9 times more useful work out of every watt, and cuts end-to-end response time by a factor of 1.7 to 3.6. For latency-sensitive, highly interactive workloads, the gap widens further, to 2.1 to 4.1 times higher performance. OpenAI frames the significance as architectural, not incremental: existing inference hardware usually forces a tradeoff between throughput and latency, and it says Jalapeño delivers gains on both at once. The company also disclosed that its own models accelerated the chip’s development twice over, helping design and bring up the hardware, then optimizing how it gets programmed.
That tradeoff is exactly what matters for anyone running or evaluating agentic workloads: an agent that has to wait longer per step to get more total throughput is a worse product experience even if the aggregate numbers look fine. If OpenAI’s claimed numbers hold up under third-party testing once Jalapeño reaches production scale later this year, the practical effect lands the same way the original announcement predicted: not a chip you buy, but a real chance at cheaper, more responsive model serving from OpenAI as this hardware comes online.
- 01OpenAI and Broadcom unveil LLM-optimized inference chipopenai.com · primary
- 02OpenAI and Broadcom Unveil LLM-Optimized Intelligence Processorinvestors.broadcom.com · primary
- 03OpenAI and Broadcom reveal Jalapeno, first AI chip in partnershipcnbc.com · reporting
- 04OpenAI and Broadcom unveil Jalapeño, a custom chip built for LLM inferencethe-decoder.com · reporting
- 05Jalapeño's first results show industry-leading speed and efficiency in AI inferenceopenai.com · primary, August 25 update
- 06OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks showtechcrunch.com · independent reporting, August 25 update
