OpenAI's Hardware chief details the Jalapeño inference chip design in an exclusive interview transcript.
OpenAI revealed its Jalapeño inference ASIC at Hot Chips in August 2026, a chip that leaned heavily on AI to deliver an remarkably compressed design window. Following the reveal, Tom's Hardware sat down with OpenAI's VP of Hardware, Richard Ho, to discuss the chip's genesis, future ambitions, and the role of AI in silicon development. The following transcript has been edited for flow and clarity.
**On the core motivation**
Jake Roach, Senior CPU Analyst at Tom's Hardware, opened by asking what drove the decision to develop a proprietary ASIC. Richard Ho identified efficiency as the primary factor. "The main thing we're aiming for is efficiency," he explained, "because obviously, as Sam has been saying, we are going to be compute-limited, and a compute limitation is really how much power we can get into data centers."
Ho emphasized that the team wanted to maximize intelligence delivery to users within constrained power budgets. He highlighted inference as the critical bottleneck: while training consumes substantial compute during pre-training, the user-facing cost—and their perception of intelligence—hinges on inference performance. "Their user experience in terms of how fast ChatGPT responds, or how fast Codex responds, or how fast the agents respond—the latency really matters," Ho said. He noted that Jalapeño was designed to offer both strong low-latency performance for latency-sensitive applications and the ability to easily shift toward higher throughput to reduce inference costs.
**Why build in-house**
When asked why OpenAI developed its own ASIC rather than rely on existing commercial offerings, Ho stressed the value of full-stack co-design. "We wanted to take advantage of the co-design opportunity," he said. "It's something that you can't do with a third-party silicon merchant really well, because there's a lot of research IP in the models. You just can't share that widely because it will leak."
Having an internal team working directly with researchers gave OpenAI complete visibility across the stack, enabling intelligent trade-offs about where optimizations belong—in software, the model itself, or hardware. "We can make those trade-offs intelligently because we have that full visibility," Ho explained. He credited this full-stack visibility for Jalapeño's benefits, allowing the team to understand exactly what hardware was needed and optimize the compiler and broader stack accordingly.
**Dispelling misconceptions about optimization scope**
Roach sought to clarify a point of industry confusion: some had suggested Jalapeño was optimized specifically for OpenAI's proprietary models. Ho corrected this mischaracterization. OpenAI deliberately used SemiAnalysis's open-source InferenceX benchmark featuring different model architectures and sizes to demonstrate broad applicability. "The whole point was that our custom inference chip was not only for OpenAI models—we've shown with the Hot Chips results that it flies on open-source models, flies on any LLM, in a sense. All transformer-based LLM models will be very performant," he said.
To underscore the chip's generality, Ho highlighted that the team didn't examine the open-source models until after receiving silicon. Yet within roughly two months of optimization work, they achieved strong performance results and were able to present them publicly. "It shows it's programmable, it's general-purpose, and it's not hard-coded for OpenAI models," he stated.
**Current and future deployment**
When asked whether Jalapeño would remain exclusively internal or serve external customers, Ho acknowledged the chip could theoretically serve others but emphasized OpenAI's overwhelming internal demand. "We have such a strong demand for compute within the company. It's going to take us a good long time to even fill our own demand, which is growing all the time," he said. With expanding daily and weekly active user bases, new capabilities in Codex, reasoning features, and unannounced products in the pipeline, OpenAI expects to be fully occupied meeting its own needs for the foreseeable future.
**The nine-month design timeline**
Roach noted the remarkable turnaround from initial design to tape-out—roughly nine months—accomplished with AI assistance, and asked whether timelines could compress further as roadmaps scale. Ho framed this as establishing a new baseline. "In the old baseline, you're talking 18 months to two years, roughly. Often that's even with some existing IP or some more legacy architecture design. We're starting from scratch here. There's not a line of code here to refer to," he explained.
The nine-month result represents what a talented team can achieve with AI support, though timeline improvements for future projects depend on design complexity. Jalapeño involved pragmatic architectural trade-offs to hit an aggressive schedule driven by urgent compute needs. More complex designs incorporating advanced packaging and co-packaged optics may take longer than nine months, Ho acknowledged, but should still improve significantly with AI assistance. Derivative designs based on Jalapeño should proceed much faster.
Ho emphasized that the nine-month achievement establishes a proof point for the industry. The AI models used in design—primarily Codex, its predecessor Sol, and the newer Astra—proved exceptionally capable. Models improved substantially between November 2025, when work began, and tape-out, and continued improving through May when kernel optimization commenced. "Even from when we started doing the kernel optimization in May, when the chips were first coming online, we ourselves were shocked at how much better Codex was and what it could do," Ho said. The team was pleasantly surprised by the performance gains achieved during the two-month sprint phase, driven largely by how effectively the models handled kernel optimization.
Ho stressed that AI augmented rather than replaced engineers. "We didn't replace our engineers; they just became super productive. With a smaller team of really good engineers with a lot of this AI stuff, you could do things faster and better than you could otherwise," he said.
**Tools and methodology**
Regarding design tools, Roach asked whether OpenAI used standard EDA software from vendors like Cadence and Synopsys. Ho confirmed the team leaned open-source philosophically but necessarily relied on standard industry flows for sign-off. "For sign-off, you need to use the standard EDA flows, and we did, because you want to make sure those results are good and correct. There's no real alternative today," he explained. The approach combined standard flows optimized by AI and expert engineers.
**Industry response and future direction**
Ho reported substantial industry interest following Jalapeño's public reveal, with conversations beginning even before the announcement once results were confirmed. While declining to preempt any specific announcements, he indicated the industry reaction was positive and OpenAI intended to broadly share its learnings. "This is not something that, 'Hey, we have this, and we're going to keep it to ourselves.' We want to make the industry more productive in general because better compute from everybody helps us as well," Ho stated.
When Roach clarified that industry interest focused on the design methodology rather than chip distribution, Ho confirmed: "Right, exactly. What we did to make those timelines, what we did to get the performance boost at the end. How did we do it? What did we use? I think those are learnings that we want to bring out to the industry as well."
Ho noted an active startup ecosystem around AI-assisted chip design and positioned OpenAI as wanting to share its approach concretely. "There is a pretty active startup scene around AI for chip design, and it's good. There's a lot of smart people thinking about it," he said. He emphasized that OpenAI's contribution would not be theoretical but grounded in tangible results: Codex, Astra, and the Jalapeño achievement itself provide concrete proof points. "We can point to it and say, 'Here's what we did, here's how we did it, and here's what we got.' It's going to be very concrete."