OpenAI claims its proprietary Jalapeño chip outperforms Nvidia Blackwell in AI inference benchmarks.
OpenAI’s custom Jalapeño AI chip has outperformed Nvidia’s Blackwell-based GB300 in key inference benchmarks, according to results released by the company. The findings underscore OpenAI’s ongoing efforts to reduce its reliance on Nvidia for running AI models.
Presented at the Hot Chips conference at Stanford University on Tuesday, the test results demonstrated that Jalapeño delivers superior performance per watt and faster response times. The evaluations were conducted using SemiAnalysis’ InferenceX benchmark. Unlike training-focused silicon, Jalapeño is specifically engineered for AI inference—the stage where trained models process user requests.
Richard Ho, OpenAI’s head of hardware, stated that the chip was designed to handle both high-volume workloads and applications demanding rapid responses. Testing incorporated OpenAI’s smaller open-source model alongside models from DeepSeek and Moonshot AI. Operating at a lower power draw of 700 watts, the chip aims to significantly reduce data centre expenses. “It's a really good chip — it should drop it by a lot,” Ho told Bloomberg regarding potential cost savings.
SemiAnalysis, which conducted the tests at OpenAI’s laboratories, confirmed that Jalapeño surpassed Blackwell in performance per watt across most scenarios. However, the research firm noted the comparison was not entirely like-for-like, as Jalapeño utilizes newer HBM4 memory. Nvidia’s upcoming Vera Rubin platform will also feature HBM4, offering a more direct competitive benchmark. “Jalapeño is really competing against chips like Rubin that also use HBM4,” SemiAnalysis analysts told CNBC. They added, “Vera Rubin systems are starting to ship to customers right now, while it will still be some time before OpenAI has anything beyond engineering samples of Jalapeño.”
OpenAI first unveiled Jalapeño in June in partnership with semiconductor manufacturer Broadcom. Early samples suggested potential cost savings of approximately 50% compared to conventional AI GPUs. The chip was developed in roughly nine months before being handed off to Taiwan Semiconductor Manufacturing Company (TSMC) for production. OpenAI plans a limited deployment toward the end of 2026, with a broader rollout anticipated in 2027. Development of a second-generation Jalapeño is already underway.
This initiative aligns with a broader industry trend as major technology firms invest heavily in custom semiconductors for AI workloads. Google, Meta, and Amazon are similarly advancing their own AI chip programs. Despite this shift, OpenAI maintains that it will continue relying on Nvidia, particularly for computationally intensive AI training tasks. “Nvidia is a really good partner, and we continue to need a lot of Nvidia,” Ho said.
Industry analysts expect Nvidia’s GPUs to remain critical due to their proven performance, flexibility, and entrenched CUDA software ecosystem. Meanwhile, Jalapeño provides OpenAI with greater direct control over the cost and efficiency of scaling AI model inference.