OpenAI and Broadcom Unveil Jalapeño, a Custom Inference Chip Targeting Late 2026 Deployment
- OpenAI's first custom chip, Jalapeño, was designed in nine months — described as the fastest ASIC development cycle for high-performance semiconductors [1]
- Engineering samples are running GPT-5.3-Codex-Spark at production target frequency and power, with 'substantially better' performance per watt than current state-of-the-art [2]
- The chip anchors a 10-gigawatt strategic collaboration between OpenAI and Broadcom, with deployment starting late 2026 and completing by end of 2029 [3]
- Microsoft is expected to purchase roughly 40% of initial Jalapeño production for its data centers [4]
- Nvidia shares fell roughly 1% on the day of the announcement, while Broadcom was flat [5]
OpenAI and Broadcom on Tuesday unveiled Jalapeño, a custom accelerator chip designed exclusively for large language model inference — marking OpenAI's first foray into silicon and a direct challenge to Nvidia's dominance in AI computing hardware. The chip went from initial design to manufacturing tape-out in nine months, a timeline the companies described as the fastest ASIC development cycle ever achieved in high-performance semiconductors [1].
Engineering samples of Jalapeño are already running machine learning workloads in the lab at production target frequency and power, including OpenAI's GPT-5.3-Codex-Spark model [2]. The companies are targeting initial deployment by late 2026, with the chip serving as the first in a multi-generation compute platform they are building together [3].
The announcement comes alongside a broader strategic collaboration to deploy 10 gigawatts of OpenAI-designed AI accelerators. Racks of accelerator and network systems are targeted to begin shipping in the second half of 2026, with the full buildout completing by the end of 2029 [3]. Microsoft, OpenAI's largest investor and cloud partner, is expected to purchase approximately 40% of the initial production run [4].
The Chip
Jalapeño is a purpose-built inference accelerator — not a general-purpose GPU — architected from a blank slate around the specific kernels, memory movement, networking, and serving patterns that matter most for frontier AI models [2]. OpenAI said its own AI models helped accelerate the design process, contributing to the compressed nine-month development timeline [1].
Early testing shows the chip delivers performance per watt 'substantially better' than current state-of-the-art hardware, though detailed independent benchmarks have not yet been published [2]. OpenAI's hardware lead Richard Ho said the architecture is designed to push utilization 'much closer to theoretical peak performance' than existing solutions [2]. Reports indicate the chip targets roughly 50% cost savings compared with typical AI GPUs for inference workloads [5].
Broadcom handled the silicon implementation and is providing networking technology, including its Tomahawk chips. Celestica is responsible for boards, racks, and full system integration [4].
The Partnership
The chip announcement is the visible centerpiece of a larger infrastructure play. OpenAI and Broadcom disclosed a strategic collaboration to deploy 10 gigawatts of OpenAI-designed AI accelerators — an enormous commitment that would represent a meaningful share of total U.S. data center capacity [3].
OpenAI president Greg Brockman framed the effort as a full-stack infrastructure strategy. 'By designing more of the stack ourselves, we can serve more intelligence with greater efficiency,' Brockman said [6]. Broadcom CEO Hock Tan called it 'just the beginning of a multi-generation roadmap' that will enable gigawatt-scale data centers with Microsoft and other partners [2].
Why It Matters
OpenAI joins Google, Amazon, Microsoft, and Meta in the growing club of hyperscale AI companies designing custom silicon to reduce dependence on Nvidia's GPU ecosystem. Nvidia remains the dominant supplier of AI training and inference hardware, with a market capitalization of $4.8 trillion, but each major custom chip announcement chips away at its inference monopoly [5].
For Broadcom, the deal validates its position as the go-to ASIC design partner for hyperscalers building custom AI processors — a business that has already fueled the company's stock to a 44% gain over the past year. Broadcom shares were roughly flat on the day at $381, while Nvidia fell about 1% to $198 [5].
The inference market is the key battleground. As AI models move from training to mass deployment, inference accounts for a growing share of total compute spending. Custom ASICs optimized for specific model architectures can deliver significant power and cost advantages over general-purpose GPUs for inference workloads, making this a natural entry point for OpenAI's hardware ambitions.
What's Next
OpenAI said a detailed technical report on Jalapeño's architecture and benchmarks will be published in the coming months [2]. The initial deployment in late 2026 will test whether the chip can deliver on its performance-per-watt promises at production scale.
The 10-gigawatt buildout with Broadcom is scheduled to complete by the end of 2029, suggesting a cadence of successive chip generations over the next three years [3]. Whether Jalapeño meaningfully reduces OpenAI's reliance on Nvidia — or merely supplements it — will depend on how quickly production ramps and how the chip performs against Nvidia's own inference-optimized hardware, including the Blackwell and Rubin architectures.
Further sources
The stories that matter, in one email. Free — unsubscribe anytime.