Friday, August 28, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

OpenAI unveiled its Jalapeño custom accelerator, claiming faster AI inference speeds and lower power consumption than flagship GPUs.

Marks a strategic shift by top-tier labs toward proprietary silicon to reduce dependency on third-party foundries and optimize inference economics.
Trade pressSlicast · August 27, 2026 · US · Source: Google News
importance 85

OpenAI has published the first detailed performance results for Jalapeño, its custom AI inference chip, positioning it directly against NVIDIA’s latest high-end systems. According to the company, Jalapeño delivers greater AI work per watt while significantly reducing response latency—two critical metrics as AI workloads shift from simple prompts toward always-on agents. OpenAI clarifies, however, that Jalapeño is not intended to replace NVIDIA across the entire AI infrastructure stack. Instead, it functions as an inference-focused accelerator, providing OpenAI with an additional hardware option alongside existing NVIDIA and third-party systems.

To evaluate performance, OpenAI utilized InferenceX, a public benchmark developed by SemiAnalysis designed to measure AI serving under realistic inference conditions. The company tested workloads involving GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. Across these operating points, Jalapeño delivered 1.5–1.9 times more peak throughput per kilowatt and achieved 1.7–3.6 times lower end-to-end latency compared to the commercial reference systems. For AI agents—which often require multiple model calls per task—these latency reductions can compound into noticeably faster interactions. These figures reflect OpenAI’s internal measurements rather than independent industry-wide benchmarks.

Beyond raw performance, the architectural approach behind Jalapeño warrants attention. OpenAI designed the chip, memory subsystem, networking fabric, and serving software as an integrated stack, allowing precise tuning around modern language models. This architecture minimizes unnecessary data movement and keeps critical model states closer to compute resources, specifically optimized for LLM inference rather than repurposed from general-purpose hardware. The accelerator was developed in partnership with Broadcom and Celestica as part of a broader multi-generation infrastructure platform.

In direct comparisons, OpenAI pitted Jalapeño against commercial systems built on NVIDIA’s GB200 and GB300 accelerators. On the Kimi K2.5 workload, Jalapeño recorded 18,195 mixed tokens per second per kilowatt, outperforming the GB300-based system’s 11,862, while also cutting end-to-end latency from 5.31 seconds to 1.56 seconds. Despite these gains, OpenAI emphasizes that Jalapeño does not signal a departure from NVIDIA. The company plans to maintain a diversified hardware strategy, selecting accelerators based on specific workload requirements rather than relying on a single vendor.

Jalapeño is exclusively for OpenAI’s internal infrastructure and will not be sold commercially. The first-generation accelerator represents the initial phase of a multi-generation compute platform, with deployment targeted for the end of 2026 and scaling to gigawatt-level capacity over time. Notably, OpenAI completed the chip’s journey from initial design to manufacturing tape-out in approximately nine months, leveraging its own AI models to accelerate parts of the design and optimization pipeline. The company reports that AI-generated implementations for attention and mixture-of-experts workloads ran 1.5–1.8 times faster than previous human-authored versions, though these improvements apply to specific components rather than complete models.

By developing its own accelerator, OpenAI gains tighter control over latency, power efficiency, memory bandwidth, and inference economics, aligning hardware directly with the models, serving software, and products it operates daily. For users, this could translate to faster ChatGPT responses, accelerated Codex workflows, and reduced API inference costs. Strategically, the move secures infrastructure flexibility as demand for AI agents and real-time inference continues to expand. While Jalapeño’s initial benchmark results establish a strong foundation in custom silicon—particularly regarding inference efficiency and latency—the true validation will occur when these gains transition from controlled testing environments into large-scale production.

Read the original
OpenAI unveiled its Jalapeño custom… · Slicast