OpenAI and Broadcom unveil Jalapeno inference chip
OpenAI and Broadcom announced Jalapeno on June 24, OpenAI's first custom Intelligence Processor for running large language models after they have been trained. The companies say the accelerator was designed around inference workloads from ChatGPT, Codex, the API and future agentic products, rather than adapted from a general-purpose chip. OpenAI says engineering samples are already running machine-learning workloads in its lab at target frequency and power, including GPT-5.3-Codex-Spark.
The announcement matters because inference is where AI products meet users. Every ChatGPT reply, Codex task or API call consumes compute after a model has been trained. If OpenAI can lower the cost and energy needed for those responses, it can make its services faster, more reliable and potentially cheaper to operate. OpenAI says early testing shows substantially better performance per watt than current state-of-the-art systems, while a detailed technical report will follow later.
Broadcom is handling silicon implementation and networking technologies, with Celestica contributing board, rack and system expertise. OpenAI frames the chip as the first step in a multi-generation platform, with initial deployment planned by the end of 2026 and expansion with data center partners after that. Broadcom CEO Hock Tan said the roadmap includes gigawatt-scale data centers with Microsoft and other partners beginning in 2026.
For the AI market, the signal is clear: frontier labs are moving deeper into infrastructure. OpenAI has depended heavily on Nvidia GPUs, especially for training, but Jalapeno gives it more control over a critical part of serving models at scale. It also follows a broader pattern in which Google, Amazon, Microsoft and Meta use custom chips alongside Nvidia hardware to manage cost, supply and performance.
For developers and companies, the short-term effect will not be a new product to buy. The bigger implication is capacity. If OpenAI's own inference hardware works as promised, users may see fewer bottlenecks, faster agents and more stable access during demand spikes. It also raises the competitive bar: the next AI race is not only about better models, but about who can run them efficiently enough for everyday use.