OpenAI Jalapeño Beats NVIDIA Blackwell by Up to 1.9x in Inference Efficiency
OpenAI has published the first detailed performance results for Jalapeño, its first custom AI inference ASIC, showing up to 1.9x higher peak throughput per kilowatt than NVIDIA Blackwell systems across a public inference benchmark. According to OpenAI's official Jalapeño results, the Broadcom assisted accelerator delivered between 1.5x and 1.9x more AI work per unit of power while also producing between 1.7x and 3.6x lower end to end latency across GPT OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.
The tests used SemiAnalysis InferenceX, a public benchmark designed to measure the complete process of serving AI requests rather than theoretical accelerator compute. Jalapeño achieved approximately 85,448 mixed tokens per second per kilowatt with GPT OSS 120B compared with 44,960 for NVIDIA GB200, giving OpenAI roughly a 1.9x advantage. On DeepSeek R1, Jalapeño reached 19,641 compared with 11,781 for GB300, or around 1.7x higher throughput per kilowatt, while Kimi K2.5 reached 18,195 compared with 11,862 for GB300, giving Jalapeño approximately a 1.5x advantage.
Power efficiency is a major part of that result. Jalapeño carries a published 700W package rating, compared with 1,200W for GB200 and 1,400W for GB300 in the tested configurations. OpenAI says actual sustained Jalapeño power remained at or below 550W during these workloads. The company normalized its benchmark comparisons using each accelerator's published chip power rating, making the results particularly relevant for AI data centers where available electricity increasingly limits how much inference capacity can be deployed.
Latency may be even more important for agentic AI. Jalapeño delivered approximately 3.6x lower end to end latency than GB300 on DeepSeek R1 and approximately 3.4x lower latency with Kimi K2.5. OpenAI says highly interactive workloads showed between 2.1x and 4.1x higher performance, which matters because AI agents can execute many sequential model calls before completing a task. A small latency reduction repeated across dozens or hundreds of reasoning steps can substantially change the responsiveness of the complete agent workflow.
The results provide the performance data that was still missing when Jalapeño was first unveiled in June. OpenAI and Broadcom moved the accelerator from initial design to tapeout in only 9 months, with OpenAI models assisting parts of chip design, verification and optimization. The company says Jalapeño was designed specifically around large language model inference, including memory movement, networking and serving behavior observed across its own production workloads rather than adapting a more general purpose GPU architecture.
OpenAI is also attacking another part of NVIDIA's advantage through software. Jalapeño was designed as a predictable programming target that allows AI systems to optimize workload placement, scheduling and communication. Using Codex with GPT Astra, OpenAI says it brought 3 open weight models that were not originally planned for Jalapeño to high performance within 2 months. For selected GPT OSS attention and mixture of experts blocks, AI generated implementations were between 1.5x and 1.8x faster than existing implementations written by human experts. Those results apply only to selected kernels, not complete model performance, but they demonstrate how AI assisted programming could reduce some of the development friction traditionally associated with moving away from mature accelerator ecosystems.
1/n OPENAI JALAPENO - Scale Up Domain:
— GDP (@bookwormengr) August 25, 2026
The most striking detail of Jalapeno is that its scale up domain can have 2048 XPUs compared to only 72 of NVL72.
They divide scale up in local and global domains - use copper for local (higher bandwidth) and use optical transreceivers… https://t.co/buWNwcgJGA pic.twitter.com/8qe8ZGsHtv
That is where the comparison with CUDA becomes strategically important. NVIDIA CEO Jensen Huang has described CUDA as NVIDIA's biggest competitive moat because decades of libraries, tools, developer experience and optimized software make replacing NVIDIA hardware considerably harder than simply producing a faster chip. Jalapeño does not eliminate that advantage. OpenAI controls its own models, serving software, networking and hardware, giving it an unusually strong environment for custom silicon that most companies cannot reproduce. OpenAI also explicitly says it will continue deploying NVIDIA accelerators for both training and inference.
OpenAI's new Jalapeño chip announced at Hot Chips beats Vera Rubin's July results on Output Throughput per MW. (1/7)🧵 pic.twitter.com/mVHQ5P41BY
— SemiAnalysis (@SemiAnalysis_) August 25, 2026
There is another important limitation to the headline numbers. Jalapeño was compared against commercially available GB200 and GB300 Blackwell systems, not NVIDIA's newer Vera Rubin platform. By the time Jalapeño reaches broader deployment during 2027, NVIDIA will be ramping a newer accelerator generation with major improvements to compute, networking and memory. The InferenceX benchmark itself is public, but the Jalapeño measurements were published by OpenAI because the custom chip is not yet available for independent external testing. The results are therefore significant measured silicon data, but they should not be interpreted as proof that OpenAI has permanently overtaken NVIDIA across AI infrastructure.
Jalapeño is more threatening to NVIDIA than another startup claiming a faster AI chip because OpenAI already owns enormous inference demand. It does not need to convince the wider market to abandon CUDA. It only needs to move enough ChatGPT, Codex and API inference onto its own silicon to reduce cost, power consumption and dependence on external accelerators.
The 1.5x to 1.9x efficiency advantage is impressive, but CUDA remains a much larger platform than a single inference benchmark. NVIDIA supports training, inference, scientific computing, networking, libraries and thousands of production workloads across an enormous developer ecosystem. Jalapeño instead demonstrates how vertical integration can attack that moat from another direction. If AI can increasingly optimize the software required to program specialized ASICs, custom silicon becomes easier to deploy, and that may ultimately be a bigger strategic challenge for NVIDIA than the raw benchmark numbers themselves.
Could AI programmed custom ASICs like Jalapeño eventually weaken NVIDIA's CUDA advantage, or is the broader CUDA ecosystem still too large for specialized chips to seriously challenge?
