Summary

Jalapeño is OpenAI's first custom inference chip, designed with Broadcom (unveiled 2026-06-24, TSMC 3nm, roughly a nine-month design cycle with AI assistance). At Hot Chips on 2026-08-25 OpenAI presented first results on the public SemiAnalysis InferenceX benchmark: against Nvidia GB200 and GB300 systems, Jalapeño delivered 1.5-1.9x more AI work per watt at peak and 1.7-3.6x lower end-to-end latency, with 2.1-4.1x higher performance for interactive workloads (GPT-OSS on GB200; DeepSeek R1 and Kimi K2.5 on GB300 as comparison points). The chip is rated at 700W and sustains at most 550W in production profiles. AI-generated kernels beat human experts by 1.5-1.8x on selected blocks of the stack. First deployment is planned for the end of 2026 in very small volumes, with second- and third-generation parts on the roadmap; Nvidia GPUs remain the bulk of OpenAI's fleet.

Why it matters
These are vendor-presented numbers, but on a public, independently defined benchmark rather than an in-house suite — a step more verifiable than most chip launches. If the per-watt lead holds at deployment, OpenAI's own inference cost trajectory decouples from Nvidia's pricing, and first-generation custom ASICs for LLM serving will have beaten current flagship GPUs on efficiency. Infra teams should watch end-of-2026 deployment and second-gen scaling before drawing architecture conclusions.
Technical details
Benchmark SemiAnalysis InferenceX (public methodology); comparison systems Nvidia GB200 (GPT-OSS) and GB300 (DeepSeek R1, Kimi K2.5)
Results 1.5-1.9x more AI work per watt at peak; 1.7-3.6x lower end-to-end latency; 2.1-4.1x higher performance on interactive workloads
Power rated 700W; sustained <=550W in production profiles
Design with Broadcom; unveiled 2026-06-24; TSMC 3nm; ~9-month design cycle with AI assistance; AI-generated kernels 1.5-1.8x faster than human experts on selected blocks
Deployment end of 2026 in very small volumes; Gen2/Gen3 roadmap; Nvidia GPUs remain the bulk of OpenAI capacity
Venue Hot Chips 2026, 2026-08-25
Tags
inference-chipcustom-siliconbroadcombenchmarkinference-infrastructure