ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Jalapeño Enters Internal Deployment: OpenAI Picks AMD EPYC Hosts, Not Nvidia Vera

Jalapeño Enters Internal Deployment: OpenAI Picks AMD EPYC Hosts, Not Nvidia Vera

AI information • Admin • • 8 views

The Jalapeño chip has entered internal deployment at OpenAI, and the host configuration beside it has just been confirmed. On October 2, 2026, Tom's Hardware reported, citing SemiAnalysis' description of the rack-scale deployment and an interview with Richard Ho, OpenAI's vice president and head of hardware, that the Jalapeño ASIC is being deployed alongside AMD EPYC Turin host CPUs rather than the Vera CPUs used in Nvidia's own platform, with 1.5TB of memory per host.

Two racks side by side: hosts on one side, 128 chips on the other

SemiAnalysis describes a two-rack pairing. The host rack, called Katsu, has 16 trays, each with two Turin-class EPYC processors, 1.5TB of standard DRAM, local NVMe storage and 400G front-end networking. The accelerator rack, called Vindaloo, also has 16 trays, each holding eight Jalapeño chips — 128 chips per rack, delivering up to 1.7 ExaFLOPs of 4-bit compute with 27.5TB of HBM4 memory. Eight PCIe direct-attach copper cables link corresponding trays one to one across the two racks. Each Jalapeño package draws about 700W and uses six 12-high HBM4 stacks for 15.4 TB/s of memory bandwidth.

Why the host vote went to AMD

In an AI rack, the host CPU handles data preprocessing, scheduling and feeding requests to the accelerators. It is not the most visible component, but it decides the procurement structure of the whole rack. Nvidia's approach bundles its GPUs with its own Vera CPUs: the deeper a customer buys in, the tighter the lock-in. With the accelerator replaced by its in-house Jalapeño and the host going to AMD, OpenAI's chain contains no Nvidia part at all. That fits its long-running multi-vendor strategy — Nvidia and AMD GPUs keep handling training, while inference gradually shifts toward custom silicon and multiple suppliers, so that its fastest-growing cost is not priced by a single vendor.

The pace is unchanged: small volumes this year, and the scores are still vendor-run

Some cold water is in order. Jalapeño was unveiled by OpenAI and Broadcom on June 24, 2026, and the first results published at Hot Chips on August 25 — 1.5 to 1.9 times the throughput per kilowatt of Nvidia's GB200 and GB300 and 1.7 to 3.6 times lower end-to-end latency on SemiAnalysis' InferenceX benchmark — come from OpenAI's own testing. Ho's timeline has deployment at the end of 2026 in very small volumes, with significant deployment in 2027. Confirmation of the host configuration shows the rollout is on track, but large-scale replacement of GPUs is still far off. For observers, what matters next is not the benchmark chart but whether these two-rack systems, once live, hold a real cost-per-inference advantage over GPU setups.

Recommended Tools

More