OpenAI reports that Jalapeño delivered 1.5–1.9× more AI work per watt and 1.7–3.6× lower end to end latency than Nvidia GB200 and GB300 systems across three tested language models. Jalapeño is an inference focused ASIC co developed with Broadcom, rated at 700 watts and measured at 550 watts or less in the reported t...
Research answer

Create a landscape editorial hero image for this Studio Global article: What are the reported performance results, power and cost-efficiency benefits, testing details and caveats, deployment timeline, development. Article summary: OpenAI reports that Jalapeño—a Broadcom co-developed, inference-focused ASIC—beats Nvidia’s Blackwell GB200/GB300 systems on the tested LLM-serving workloads for work per watt and latency. It is a meaningful proof point . Topic tags: general, general web, user generated, news. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts w
OpenAI’s first published results for Jalapeño make a focused claim, not a universal one: its custom inference ASIC outperformed Nvidia’s GB200 and GB300 systems on work per watt and response latency across selected large-language-model serving tests. That is an important signal for AI infrastructure, where electricity and cooling increasingly shape the cost of serving models. But Jalapeño is designed for inference—not as a general-purpose replacement for Nvidia GPUs—and the comparison has meaningful limitations. 57
Jalapeño is an application-specific integrated circuit (ASIC) co-developed by OpenAI and Broadcom for large-language-model inference. Inference is the stage where a trained model generates responses, rather than the training process used to create the model in the first place. Purpose-built silicon can trade some flexibility for better efficiency on predictable, high-volume workloads. 619
That specialization is central to interpreting the results. Nvidia’s GPU systems are intended to support a much wider range of workloads, while Jalapeño is optimized around OpenAI’s serving requirements and its broader hardware-software stack. 56
OpenAI says it tested Jalapeño against Nvidia GB200- and GB300-based systems using three publicly available models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI’s Kimi K2.5 1T. Across those comparisons, Jalapeño delivered 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency, according to the company’s results. 5713
The reported model-level figures illustrate the range:
OpenAI also reported larger advantages for highly interactive operating points, where the priority is rapid response rather than maximum aggregate throughput. The company put that advantage at 2.1–4.1× higher performance in those workloads. 413
The power figures help explain the appeal. Jalapeño is rated at 700 watts, while the compared GB200 and GB300 accelerators are rated at 1,200 and 1,400 watts, respectively. OpenAI says Jalapeño’s measured sustained power stayed at or below 550 watts on the tested workloads. 61213
Lower accelerator power does not automatically translate into an equivalent reduction in total data-center cost. However, serving more model work with less power can reduce pressure on electricity, cooling, and power-distribution systems. At OpenAI’s anticipated scale, those operational effects are the core economic rationale for custom inference silicon. 5
The tests used SemiAnalysis’s public InferenceX benchmark and measured end-to-end serving across several models and operating points, rather than relying only on a theoretical chip specification. That makes the results more relevant to real inference operations. 7
They should still be read as OpenAI-reported results, not as an independently replicated, vendor-neutral verdict. 45
There is also a hardware-generation caveat. Jalapeño uses HBM4 memory, while the tested GB200 and GB300 systems do not. SemiAnalysis characterized the comparison as incomplete and not entirely fair for that reason. 5
This matters because memory bandwidth and capacity are important in serving large models. Nvidia’s Rubin platform also uses HBM4, making it a more relevant generational comparator than Blackwell systems that lack the same memory technology. Jalapeño’s reported lead over GB200 and GB300 therefore does not establish that it will outperform Rubin. 517
Another limitation is that an inference-only chip is being compared with systems built for broader use. Nvidia GPUs can be used for training, fine-tuning, inference, and changing model architectures, while a custom ASIC is most valuable when the target workload is stable enough to justify specialization. 520
OpenAI plans to begin deploying Jalapeño in its own compute infrastructure in very small volumes by the end of 2026, followed by a more substantial ramp in 2027. 1133
The program is intended to be multigenerational. OpenAI has said that second- and third-generation chips are already in development, while Broadcom describes a roadmap aimed at gigawatt-scale deployments with Microsoft and other data-center partners. 31130
The broader strategy is vertical co-design: coordinating models, inference software, memory, networking, systems, and silicon rather than treating the accelerator as an isolated component. The potential benefit is tighter control over latency and serving costs; the trade-off is that the resulting hardware is less broadly reusable than a general-purpose GPU. 3633
The fast development schedule is notable, but a rapid tape-out is not the same as proven volume production. Manufacturing yield, system integration, software maturity, supply, and the ability to support future models will determine whether the benchmark advantage becomes a durable infrastructure advantage. The available reporting confirms the planned rollout, but not yet its eventual production scale. 3033
OpenAI has said it will continue relying on Nvidia for training and as a compute partner. 411
That division of labor is logical. Training frontier models requires highly flexible systems that can accommodate changing architectures, algorithms, and experimentation. Nvidia’s GPUs also benefit from the CUDA software ecosystem and a full stack that includes networking and complete data-center systems. 5
Jalapeño is therefore best understood as a targeted substitution strategy. OpenAI may use it for recurring, predictable inference workloads where it controls the models, serving software, and infrastructure. Nvidia remains important for training, research, model iteration, and workloads where flexibility matters more than specialization. 45
The strongest case for Nvidia is not that a custom ASIC cannot win a narrowly defined inference benchmark. It is that customers often need one platform to handle many jobs as models and software change. GPUs can train, fine-tune, and serve models, whereas a purpose-built ASIC concentrates its advantage on a narrower workload. Nvidia’s software, networking, and systems ecosystem also make the comparison broader than chip-level throughput alone. 5
Deployment scale further limits the immediate competitive impact. Jalapeño will initially serve OpenAI’s own infrastructure rather than becoming a generally available alternative that customers can buy and deploy across the market. 14
In that sense, the announcement changes the bargaining position around inference more immediately than it changes the overall AI-compute market. OpenAI has demonstrated a possible alternative for part of its workload, but it has not ended its dependence on Nvidia.
Analysts view Jalapeño as evidence that a hyperscaler-designed ASIC can match or exceed Blackwell-class systems on inference efficiency. That could pressure Nvidia in inference, especially as large AI operators seek to lower serving costs and reduce dependence on a single supplier. 1732
OpenAI is not alone in pursuing custom silicon. Google, AWS, and Meta are also cited as major technology companies developing their own AI chips, reinforcing the shift toward workload-specific infrastructure. 5
The most plausible outcome is not an immediate GPU collapse but a more segmented market:
The key question is whether Jalapeño’s reported gains survive fairer comparisons, changing models, and production-scale deployment. For now, the evidence supports a narrower conclusion: OpenAI has shown that custom silicon can be highly competitive with Blackwell for selected inference workloads, while Nvidia remains deeply embedded in the rest of the AI-compute stack.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
OpenAI reports that Jalapeño delivered 1.5–1.9× more AI work per watt and 1.7–3.6× lower end to end latency than Nvidia GB200 and GB300 systems across three tested language models.
OpenAI reports that Jalapeño delivered 1.5–1.9× more AI work per watt and 1.7–3.6× lower end to end latency than Nvidia GB200 and GB300 systems across three tested language models. Jalapeño is an inference focused ASIC co developed with Broadcom, rated at 700 watts and measured at 550 watts or less in the reported tests; OpenAI plans small volume deployment by the end of 2026 and a broader ramp...
The comparison is not fully like for like because Jalapeño uses HBM4 while the tested Blackwell systems do not.