Hardware

OpenAI Jalapeño Chip Beats Nvidia on Inference

OpenAI has released performance data for Jalapeño, its custom Broadcom-built inference chip, which beats Nvidia's top hardware and signals a shift toward proprietary AI silicon.

AlphaSignal2 days agoHardware
Image: AlphaSignal

OpenAI has published the first performance metrics for Jalapeño, its custom inference processor co-designed with Broadcom. Benchmarked on InferenceX using GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, the 700-watt chip—which sustained power at or below 550 watts—significantly outperformed Nvidia's 1,200-watt GB200 and 1,400-watt GB300 systems. Across these models, Jalapeño delivered 1.5 to 1.9 times more work per watt at peak throughput, 1.7 to 3.6 times lower end-to-end latency, and 2.1 to 4.1 times higher performance on highly interactive workloads. On Kimi K2.5 1T, it achieved roughly 1.5 times higher peak performance per watt and 3.4 times lower latency than the GB300.

The hardware went from early schematics to fabrication readiness in just nine months, a speed OpenAI attributes to using its own models for software-hardware co-development. Earlier GPT models optimized arithmetic circuits, while Codex and an internal model named GPT-Astra helped write kernels. This AI-assisted workflow allowed engineers to optimize three unplanned open-weight models in two months. For certain GPT-OSS attention and mixture-of-experts blocks, the AI-generated code ran 1.5 to 1.8 times faster than human-written alternatives.

Deployment of the inference-only chip will begin inside OpenAI's infrastructure by the end of this year, marking the start of a multi-generational platform that includes a 10-gigawatt Broadcom buildout slated for completion by late 2029. While OpenAI will continue purchasing Nvidia hardware for training and general use, Jalapeño targets the operational costs of serving models. Early estimates suggest a roughly 50 percent reduction in cost per inference token compared to Nvidia GPU clusters. For developers, this efficiency translates directly to faster ChatGPT and Codex responses, lower API pricing, and greater capacity for complex, multi-step agent workloads without throttling.

This is our own summary of reporting by AlphaSignal

More in Hardware