OpenAI's custom Jalapeno AI inference ASIC is built for internal use but the company leaves the door open to eventual broader commercial rollout.
Following OpenAI's reveal of the Jalapeño ASIC, a fundamental question emerged: who is this for? OpenAI's messaging has been mixed. The company suggested it was building ASICs for internal use when announcing its partnership with Broadcom last year, yet it subsequently benchmarked Jalapeño against Nvidia's Blackwell accelerators and detailed a multi-generational roadmap at Hot Chips 2026.
Jalapeño is designed for OpenAI's compute needs, according to Richard Ho, Head of Hardware at OpenAI. However, Ho told Tom's Hardware Premium that "you could use it for anybody, honestly," signaling openness to wider adoption. "We have such a strong demand for compute within the company. It's going to take us a good long time to even fill our own demand, which is growing all the time," Ho said. "I think that we're going to have our hands full just providing compute for OpenAI for a good long time. That's not to say that it can't be used elsewhere. I believe it could be, but I think our priority is to make sure that OpenAI's compute needs are met first and foremost."
The chip's competitive positioning rests on benchmarks OpenAI shared at Hot Chips, using SemiAnalysis' InferenceX benchmark to compare Jalapeño against Nvidia's GB200 and GB300. ASICs are common, but competitive performance among them is rare—they're optimized for specific workloads. Jalapeño proved notable by accelerating not just OpenAI's own open-weight GPT-OSS model, but also DeepSeek R1 and Kimi K2.5.
Originally, OpenAI hadn't planned to present benchmarks at Hot Chips or even attend the event. Ho recounted how a team of engineers got Kimi and DeepSeek running on Jalapeño in just two months between the A0 sample and the Hot Chips presentation—a feat that impressed even OpenAI's leadership.
Ho was explicit that Jalapeño remains deployed internally for now, yet he left the door deliberately ajar. "What we really wanted to demonstrate, to put to rest, the misperception in the industry that our custom inference chip was only for OpenAI models," Ho said. "It's programmable, and it's general purpose, and it's not hard-coded for OpenAI models."
Supply constraints may explain OpenAI's reluctance to enter the external hardware market. Ho noted that there's "a new baseline for supply," reflecting two years of efforts by himself and CEO Sam Altman touring fabrication plants to secure additional capacity. While OpenAI is "in good shape" on supply internally, sourcing for external customers presents a different challenge.
Efficiency, not raw performance, was the primary design driver. Ho stressed efficiency as a cousin of compute, pointing to power-constrained modern AI data centers where more efficient inference translates directly to more effective compute. Codesign with OpenAI's internal models also played a crucial role. "Codesign is something that you can't do with a third-party silicon merchant really well because there's a lot of research IP in the models, and so you just can't share that widely because it will leak. It will get out there no matter how many NDAs you put in place," Ho explained.
If Jalapeño were destined for external release, it would face competition not from Nvidia's Grace Blackwell platform but from the newer Vera Rubin. Ho said Blackwell comparisons were made "because those were the best published results that we could find," though OpenAI has run internal benchmarks against Vera Rubin and tested larger context windows. The published InferenceX benchmarks focused on 8k1k scenarios—8,000 input tokens and 1,000 output tokens.
OpenAI's internal testing suggests the ASIC "seems to perform even better than the existing benchmarks from some of the other devices that are available." The Vera Rubin comparison looks particularly strong. "Obviously, by the time we deploy, it'll be Vera Rubin, maybe even VR Ultra in some parts of the deployment schedule. We've done our internal ones, but obviously we don't publish those. Those have to come from Nvidia and other people who are able to do that," Ho said. "Yeah, we're doing really well on those."