OpenAI says Jalapeño, its first Broadcom co developed custom accelerator, beats the compared Nvidia GB200/GB300 rack systems on the tested inference workloads in both energy efficiency and response speed. These are early, vendor reported benchmark results—not an independent, apples to apples measure of all productio...
Research answer

Create a landscape editorial hero image for this Studio Global article: What did OpenAI’s custom Jalapeño inference chip, co developed with Broadcom, demonstrate in its latest benchmarks against Nvidia’s GB200 an. Article summary: OpenAI says Jalapeño, its first Broadcom co developed custom accelerator, beats the compared Nvidia GB200/GB300 rack systems on the tested inference workloads in both energy efficiency and response speed.. Topic tags: general web, openai, llm, agents, ai. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
OpenAI says Jalapeño, its first Broadcom co-developed custom accelerator, beats the compared Nvidia GB200/GB300 rack systems on the tested inference workloads in both energy efficiency and response speed. These are early, vendor-reported benchmark results—not an independent, apples-to-apples measure of all production AI workloads. 47
Across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, OpenAI reported 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency than its comparison systems. 4
On GPT‑OSS 120B versus GB200, OpenAI reported 1.9× higher peak mixed throughput per kW (85,448 versus 44,960), 1.7× lower end-to-end latency (1.03 s versus 1.80 s), and a minimum time-between-tokens of 0.69 ms—or 1,459 tokens/s/user—versus 1.87 ms, or 535 tokens/s/user. 4
On DeepSeek R1 MXFP4 versus GB300, OpenAI reported 1.7× higher peak mixed throughput per kW (19,641 versus 11,781), 3.6× lower end-to-end latency (1.65 s versus 5.99 s), and 1.43 ms minimum time-between-tokens: about 700 tokens/s/user versus 169 on GB300. 4
On Kimi K2.5 MXFP4 versus GB300, it reported 1.5× higher peak throughput per kW, 3.4× lower end-to-end latency, and about 694 versus 182 tokens/s/user at minimum time-between-tokens. 4
Jalapeño is an LLM-inference-specific accelerator rather than a general-purpose GPU adapted for the task. OpenAI, Broadcom, and Celestica respectively contributed the accelerator design, silicon implementation/networking, and board/rack/system integration. 1
The disclosed rack-scale design uses a large connected network domain so a workload can remain within one system; it explicitly places and retains model state, including KV cache, locally to reduce inter-chip movement and communication stalls. 4
OpenAI says the design does not separate prefill and decode into distinct resource pools: the same fungible compute, memory, and network resources are intended to handle both phases. That matters because prompt processing is compute-heavy, while autoregressive token generation is memory-bandwidth-heavy. 4
Reported system specifications beyond that are limited in the available public evidence. One technical report describes a 128-chip system with roughly 1.7 exaFLOPS and 27 TB of HBM, but OpenAI’s benchmark post does not itself provide a complete rack bill of materials. 4
OpenAI says it moved from initial design to manufacturing tape-out in nine months, calling it an exceptionally fast advanced-ASIC cycle. 1
Its models assisted design, implementation exploration, measurement, verification, arithmetic-circuit optimization, and programming. OpenAI says Codex with GPT‑Astra brought three previously unplanned open-weight models to high performance within two months; on selected GPT‑OSS attention and mixture-of-experts blocks, AI-generated implementations were 1.5–1.8× faster than existing human-written ones. Those block-level figures are not whole-model results. 4
Jalapeño is intended for inference, not training. OpenAI will therefore continue to rely on Nvidia, AMD, and other accelerators for training as well as some inference capacity. 6
The benchmark was SemiAnalysis’s public InferenceX suite, using nominal single-turn 8k-token input/1k-token output workloads and measuring full request serving. It compared the systems at matched user experience and normalized by published chip power ratings. 47
OpenAI lists Jalapeño at 700 W package power, while saying sustained measured power was at or below 550 W on these workloads; GB200 and GB300 were modeled at 1,200 W and 1,400 W respectively. This makes the reported per-watt advantage sensitive to workload and power-accounting choices. 4
SemiAnalysis says it witnessed runs in OpenAI’s lab but did not run the full benchmark suite itself. It also cautions that the single-turn 8k/1k InferenceX setup does not test longer-context, multi-turn agent workloads, routing, prefix caching, cache management, or offload behavior; such factors can materially affect production performance. 4
The comparison is also time-bound: GB200/GB300 are commercial Blackwell-generation systems, while newer Nvidia platforms may be more appropriate future comparators. Nvidia separately reports GB300 DeepSeek-R1 throughput under MLPerf’s different offline methodology, so those results should not be directly compared with Jalapeño’s interactive tokens-per-user figures. 24
OpenAI plans a limited deployment in its own compute infrastructure by the end of 2026, followed by greater capacity in 2027. Gen 2 is “deep in development,” and Gen 3 is beginning to take shape. 46
It does not plan to sell Jalapeño externally: OpenAI hardware VP Richard Ho said internal demand is too high to envisage outside sales. 6
This is principally a strategy to supplement—not displace—Nvidia and AMD. OpenAI says it will continue broad deployment of Nvidia and other partners’ accelerators, particularly because Jalapeño does not address training. 46
In the broader competitive landscape, Google already has TPUs for training and inference; Amazon and Microsoft have in-house AI chips; and Anthropic has said it is building an in-house chip-design effort. Jalapeño puts OpenAI in that same vertical-integration race, but its current advantage claim is narrowly about LLM inference on the disclosed benchmark configurations. 6
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
OpenAI says Jalapeño, its first Broadcom co developed custom accelerator, beats the compared Nvidia GB200/GB300 rack systems on the tested inference workloads in both energy efficiency and response speed.
OpenAI says Jalapeño, its first Broadcom co developed custom accelerator, beats the compared Nvidia GB200/GB300 rack systems on the tested inference workloads in both energy efficiency and response speed. These are early, vendor reported benchmark results—not an independent, apples to apples measure of all production AI workloads.
[4][7] Benchmark results Across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, OpenAI reported 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end to end latency than its comparison systems.