OpenAI reported that Jalapeño delivered 1.5–1.9× more AI work per watt and 1.7–3.6× lower end to end latency than the tested Nvidia GB200 and GB300 systems in August 2026. The tests used GPT OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T on SemiAnalysis’s InferenceX benchmark, with Jalapeño compared against GB200 for...
Research answer

Create a landscape editorial hero image for this Studio Global article: What did OpenAI’s custom Jalapeño inference chip achieve in its August 2026 benchmarks against Nvidia’s GB200 and GB300 systems, which model. Article summary: OpenAI’s August 2026 results position Jalapeño as a potentially much more efficient, lower-latency option for high-volume LLM serving than the specific Nvidia GB200 and GB300 configurations tested—not as a general replac. Topic tags: general, news, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts w
OpenAI’s first detailed Jalapeño results make a focused claim: its custom silicon may serve large language models with less power and lower response latency than the specific Nvidia GB200 and GB300 systems used for comparison. On SemiAnalysis’s public InferenceX benchmark, OpenAI reported 1.5–1.9× more AI work per watt and 1.7–3.6× lower end-to-end latency across three open-weight models.
That is significant for a company serving AI requests at enormous scale, but it is not the same as proving that Jalapeño is a faster or cheaper replacement for GPUs everywhere. The tests were narrow, the chip is designed only for inference, and the results were primarily generated by OpenAI on engineering silicon. 14
OpenAI compared Jalapeño with Nvidia’s GB200 and GB300 systems on InferenceX, a benchmark intended to measure the full process of serving an AI request. The company normalized efficiency figures using the published power ratings of the accelerators.
Across the reported tests, Jalapeño achieved:
The largest latency gap reported in the supplied results was on DeepSeek R1 670B: Jalapeño completed a request in 1.65 seconds, compared with 5.99 seconds on the GB300-based system. 1112
These are relative benchmark results, not a promise of a matching reduction in ChatGPT response times or API prices. Real-world performance depends on the model, serving software, batching, networking, context length, output length, and the target latency or throughput level.
The benchmark covered three publicly available open-weight models:
| Model | Nvidia comparison | Reported Jalapeño result |
|---|---|---|
| GPT-OSS 120B | GB200 | About 1.9× more work per watt; 1.03 seconds versus 1.80 seconds end-to-end latency |
| DeepSeek R1 670B | GB300 | About 1.7× more work per watt; 1.65 seconds versus 5.99 seconds latency |
| Kimi K2.5 1T | GB300 | About 1.5× more work per watt; 1.56 seconds versus 5.31 seconds latency |
The figures above come from reported benchmark summaries rather than an independently reproduced, full-suite test. 61112
The workload was nominally single-turn, with 8,000 input tokens and 1,000 output tokens. That setup creates a controlled comparison, but it does not represent every type of AI serving traffic. It leaves important questions about multi-turn conversations, tool calls, agent workflows, variable-length responses, batching behavior, and very long-context requests. 813
The benchmark also measured particular software and hardware configurations. Changes to compilers, kernels, quantization, scheduling, model architecture, or Nvidia’s runtime stack could change the size of the gap.
Jalapeño is an application-specific integrated circuit, or ASIC, co-developed by OpenAI and Broadcom for large-language-model inference. Unlike a general-purpose GPU, it is optimized for the operation of running trained models to generate responses. 15
Reported system specifications include:
At the chip level, the reported B0 version uses a reticle-scale compute die manufactured on TSMC’s N3P process and is rated at 13.4 PFLOPS of MXFP4 compute. 18
The architecture reflects a full-system approach rather than a standalone accelerator strategy. Memory, interconnect, networking, and rack-scale communication are central to serving very large models, where requests may need to be distributed across many chips. 1819
OpenAI and Broadcom reportedly moved from concept to tape-out in roughly nine months, an unusually short timeline for a high-performance custom chip. 720
AI-assisted engineering was part of the design process, but that does not mean an AI model independently designed and manufactured the processor. The reported development still involved human architecture and engineering teams, Broadcom’s semiconductor implementation work, manufacturing partners, and system integration. 713
The more defensible takeaway is that OpenAI used its knowledge of LLM workloads—and AI tools within the engineering workflow—to co-design hardware and software around a narrow serving target. That specialization can improve efficiency, but it also reduces flexibility compared with a programmable GPU.
Jalapeño is built to serve trained models, not to train new frontier models. OpenAI will therefore continue to need programmable accelerators, especially GPUs, for model training, research, experimentation, and workloads that do not map cleanly onto a fixed-function inference design. 813
This distinction limits the meaning of “beating Nvidia.” Jalapeño may outperform the tested Nvidia configurations on selected inference metrics while Nvidia hardware remains essential elsewhere in OpenAI’s compute stack. A specialized ASIC can be excellent at a predictable, high-volume job without replacing the broader platform that supports training and rapid model changes.
The results should be read as “better in these reported tests,” not “the best AI system overall.” Several caveats matter:
There is also no basis in these results for claiming a universal 50% cost reduction, a guaranteed lower API price, or immediate Nvidia displacement. The supplied evidence supports performance and efficiency claims under specified conditions, not those broader commercial conclusions.
OpenAI plans to begin deploying Jalapeño in its own infrastructure by the end of 2026, with volume production and broader rollout expected in 2027. The longer-term plan involves gigawatt-scale data centers with Microsoft and other infrastructure partners across multiple chip generations. 16
The chip is intended primarily for OpenAI’s internal demand rather than for sale as a merchant processor. That means outside companies should not expect to buy Jalapeño hardware or rent a public Jalapeño cloud instance based on these announcements. 16
Keeping the hardware internal lets OpenAI optimize it for its own models, kernels, and serving systems. It also avoids the software support, customer competition, sales, and ecosystem obligations that come with becoming a conventional semiconductor supplier. Those strategic benefits are an interpretation of the deployment model; the confirmed plan is internal use rather than external chip sales. 16
Jalapeño is best understood as one part of OpenAI’s broader full-stack compute portfolio. The company’s strategy combines Microsoft and other cloud capacity, Nvidia-class GPUs for flexible and training-heavy workloads, and custom silicon for stable inference workloads where specialization can pay off. OpenAI’s own description of its compute portfolio includes Microsoft, Nvidia, AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, and SoftBank.
That approach mirrors a wider industry shift. Google has developed TPUs, Amazon offers Trainium and Inferentia, Microsoft has Maia, AMD competes with general-purpose AI GPUs, and Nvidia continues to supply broadly programmable systems. Anthropic is also pursuing increasingly customized infrastructure through its cloud relationships, although the supplied benchmark evidence does not provide a direct performance comparison with Anthropic’s systems.
The near-term strategic impact is therefore more nuanced than a GPU replacement story. Jalapeño could give OpenAI more control over inference economics, latency, supply, and its hardware roadmap. Nvidia remains important for training and flexibility, while OpenAI’s custom chip gives it a specialized option for the predictable, high-volume serving workloads that dominate its own infrastructure needs. 813
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
OpenAI reported that Jalapeño delivered 1.5–1.9× more AI work per watt and 1.7–3.6× lower end to end latency than the tested Nvidia GB200 and GB300 systems in August 2026.
OpenAI reported that Jalapeño delivered 1.5–1.9× more AI work per watt and 1.7–3.6× lower end to end latency than the tested Nvidia GB200 and GB300 systems in August 2026. The tests used GPT OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T on SemiAnalysis’s InferenceX benchmark, with Jalapeño compared against GB200 for GPT OSS and GB300 for the two larger models.
Jalapeño is a 700 watt, inference only ASIC developed with Broadcom. OpenAI plans internal deployment from late 2026, with broader volume rollout expected in 2027.