OpenAI의 커스텀 Jalapeño ASIC 배포는 추론 워크로드를 위해 Nvidia의 Vera Rubin 가속기 대신 AMD의 EPYC Turin 프로세서를 선택했다.
OpenAI's Jalapeño ASIC deployment pairs its new inference chips with AMD EPYC Turin hosts, despite the company's deep infrastructure relationship with Nvidia. Each host carries two Turin-class processors and 1.5TB of DRAM. Nvidia's new Vera CPU did not make the first production design.
That choice is more than a component swap. OpenAI designed Jalapeño as an application-specific integrated circuit optimized for language-model inference, yet built the surrounding host layer with a mature x86 server platform instead of Nvidia's purpose-built Arm CPU. Richard Ho, OpenAI's vice president and head of hardware, described the Turin decision as "pragmatic," telling Tom's Hardware that Vera remained "a little bit behind" at the required maturity level. The comment places deployment certainty ahead of tighter ownership of the computing stack. OpenAI still depends heavily on Nvidia accelerators elsewhere, and Jalapeño remains an internal platform whose early performance claims require broader validation. Even so, the host decision shows how hyperscalers can challenge Nvidia selectively without abandoning its hardware entirely.
OpenAI has turned Jalapeño from a chip announcement into a rack design built around separate AMD host and custom accelerator layers. The architecture uses one CPU host rack beside one Jalapeño ASIC rack. Each Katsu CPU tray contains two AMD EPYC Turin-class CPUs and 1.5TB of DRAM, along with local storage and 400-gigabit frontend networking. Eight external PCIe cables connect each CPU tray to its corresponding Vindaloo accelerator tray. The neighboring rack holds 128 Jalapeño chips across 16 accelerator trays, with eight Chana switch trays connecting those accelerators inside the rack and across larger installations. The published rack architecture can extend its scale-up network across 16 racks, connecting as many as 2,048 Jalapeño accelerators through copper and optical links.
The host processors do not replace the inference ASICs; they handle the supporting CPU work required to feed, coordinate, schedule, and manage accelerator workloads. Production inference needs tokenization, request handling, storage access, networking, model orchestration, safety services, and other CPU-bound operations. Agentic applications increase those demands—an agent can alternate between model inference, Python execution, retrieval, database access, and external tools. Poor host performance can leave expensive accelerators waiting while those stages complete. OpenAI therefore needed more than a fast inference chip: a host platform with sufficient memory capacity, network support, software compatibility, and operational history.
Power illustrates the system-level challenge. The host rack uses approximately 31 kilowatts in production, while the accelerator rack draws roughly 130 kilowatts. Together, the paired system consumes about 160 kilowatts. These figures show why Jalapeño cannot be assessed through chip specifications alone—rack networking, host utilization, cooling, software, and workload placement all influence the useful work delivered by the installation. The same principle applies to OpenAI's benchmark claims: a favorable accelerator result matters only if the complete system can repeat it under production traffic. Choosing an established host platform reduces one source of uncertainty during that transition.
AMD EPYC Turin hosts gave OpenAI a known platform while the company attempted an unusually compressed accelerator development schedule. OpenAI and Broadcom say Jalapeño moved from initial design to manufacturing tape-out in nine months. That schedule describes a program with little room for avoidable integration problems. OpenAI designed the accelerator architecture, while Broadcom contributed silicon implementation, networking, and connectivity expertise. Celestica worked on the board, rack, and complete system design. Ho said the team wanted aggressive performance and cost goals without accepting unnecessary risks. Turin met the program's requirements, and OpenAI's partners already had relevant experience with the platform.
A new CPU architecture would expand the validation surface. Differences in instruction sets, compiler behavior, management tooling, and application compatibility can create delays even when the underlying processor performs well. Turin belongs to AMD's fifth-generation EPYC server family, uses the established x86 instruction set, and supports twelve memory channels per socket, giving system designers substantial memory bandwidth and capacity. OpenAI's rack design places 1.5TB of DRAM beside each pair of processors to retain application state, prepare requests, manage data, and support CPU-side services surrounding inference.
OpenAI also had to prepare software for an accelerator without an existing developer ecosystem. Nvidia benefits from years of CUDA adoption, optimized libraries, deployment tools, and operator familiarity. Jalapeño begins without that installed base. OpenAI can control its internal software environment, but its engineers must still build compilers, kernels, monitoring systems, and scheduling logic for the new platform. Using familiar host hardware keeps that work concentrated on the custom accelerator, letting the team separate Jalapeño-specific defects from problems caused by an additional CPU transition.
The decision reflects schedule discipline rather than a sweeping judgment about x86 and Arm. OpenAI selected the component that reduced integration risk for this generation. Ho did not argue that Turin will remain the best host for every future OpenAI system; he said it satisfied the immediate program's needs and maturity requirements. The first Jalapeño generation is consequently a hybrid strategy: OpenAI takes architectural risk where specialization promises meaningful inference gains, while retaining commodity server technology where maturity provides greater value.
The immediate pressure falls on Nvidia's attempt to make Vera the default CPU around next-generation AI infrastructure. Nvidia presents Vera as a processor designed for the CPU work surrounding agents—Python runtimes, sandboxed code, orchestration, analytics, and other tasks that occur between accelerator calls. The processor uses 88 custom Olympus cores and an LPDDR5X memory subsystem that provides as much as 1.2TB per second of bandwidth. Vera also connects to Rubin GPUs through NVLink-C2C, supplying up to 1.8TB per second of coherent bandwidth between the CPU and GPU. That tight coupling supports Nvidia's larger sales argument: customers can buy CPUs, GPUs, networking, interconnects, libraries, and rack systems as one coordinated platform. Nvidia says Vera systems will become available through system builders and cloud partners during fall 2026. OpenAI's choice exposes a timing problem inside that strategy: Vera may offer attractive specifications, but Jalapeño needed a host platform that partners could deliver on the accelerator's schedule.