OpenAI says its Broadcom-built Jalapeño chip delivered 1.5x-1.9x more throughput per kilowatt and 1.7x-3.6x lower latency than Nvidia's GB200/GB300 on SemiAnalysis's InferenceX suite, though the lead narrows to about 1.5x under utility-power and multi-token-prediction comparisons, and Nvidia's upcoming Vera Rubin platform was not tested.
How much faster is Jalapeño than Nvidia's flagship GPUs?
OpenAI says its Jalapeño chip delivered 1.5 times to 1.9 times more throughput per kilowatt and 1.7 times to 3.6 times lower end-to-end latency than Nvidia's GB200 and GB300 rack systemsCITE:E1. The figures come from SemiAnalysis's public InferenceX benchmark suiteCITE:E1. The comparison pits a 700W Jalapeño part against accelerators rated at 1,200W (GB200) and 1,400W (GB300)CITE:E3. Testing covered three open models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's 1-trillion-parameter Kimi K2.5CITE:E4. OpenAI reports its widest leads at low-latency operating points, where it claims 8.6 times to 104.3 times more throughput per kilowatt versus the GB300's fastest previous time-between-tokens settingsCITE:E5.
| Metric | Jalapeño | GB200 | GB300 |
|---|
| Rated power | 700W | 1,200W | 1,400W |
| Measured sustained power | ≤550W | — | — |
| Throughput per kW vs GB200/GB300 | 1.5x–1.9x higher | baseline | baseline |
| End-to-end latency | 1.7x–3.6x lower | baseline | baseline |
| Throughput per kW at low-latency settings | 8.6x–104.3x higher | — | baseline |
| All-in utility power per accelerator | 1.18kW | — | 2.55kW |
| HBM memory | 216 GiB HBM4 @ 15.4 TB/s | — | 288GB HBM3E |
What power did Jalapeño actually draw during testing?
OpenAI said Jalapeño's measured sustained power stayed at or below 550W in testing, below its 700W ratingCITE:E6. That figure applies across the same three-model test set — GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5CITE:E4 — where OpenAI's steepest per-kilowatt margins over the GB300 appeared at low-latency settings, reaching up to 104.3 timesCITE:E5.
Do the efficiency gains hold under different test conditions?
OpenAI's own appendix data shows the efficiency gap narrows once power accounting changesCITE:E7. A comparison using all-in utility power per accelerator — 1.18kW for Jalapeño against 2.55kW for the GB300 — produces narrower gaps than the headline rated-power comparisonCITE:E7. Pitting Jalapeño against a GB300 running multi-token prediction shrinks the peak efficiency lead to roughly 1.5 timesCITE:E8.
What does an independent evaluator say about Jalapeño?
SemiAnalysis, which ran InferenceX together with OpenAI engineers in the company's lab, described the chip as "beating every Nvidia, AMD, and Google chip we have been able to test"CITE:E9.
How does Jalapeño's memory compare to the GB300's?
Each Jalapeño package pairs its compute die with six HBM4 stacks, totaling 216 GiB at 15.4 TB/s, versus the GB300's 288GB of HBM3E at a 1,400W ratingCITE:E10CITE:E11. Per watt of rated power, OpenAI's chip packs roughly 50% more memory than the GB300CITE:E11. Memory capacity carries a manufacturing cost: Micron told the Hot Chips conference on August 23 that HBM consumes roughly three times the wafer area of DDR5 for equivalent capacity, a penalty that widens with each generationCITE:E17.
What is OpenAI's stance on Nvidia given its Broadcom partnership?
OpenAI signed a 10GW custom AI chip deployment agreement with Broadcom last OctoberCITE:E16, and separately, on August 17, Nvidia agreed to provide up to $105 billion in financing for an OpenAI-leased data center campus in OhioCITE:E13. Despite fielding its own chip, Richard Ho, OpenAI's vice president of hardware, told Bloomberg following the announcement: "Nvidia is a really good partner, and we continue to need a lot of Nvidia"CITE:E14.
What's next for Jalapeño's development roadmap?
Jalapeño was unveiled in June after a nine-month RTL-to-tapeout cycleCITE:E10. According to Bloomberg, a second-generation chip is approaching tapeout, expected within months, and concept work on a third generation is already underwayCITE:E15.
What deployment risks and supply pressure could affect the rollout?
OpenAI plans to begin deploying Jalapeño in its own data centers later this year, but it has not yet been benchmarked against Nvidia's next platformCITE:E12. Jalapeño wasn't tested against Vera Rubin, the Nvidia platform slated to power the first gigawatt of Nvidia systems OpenAI agreed to deploy in the second half of 2026CITE:E20. Scaling Jalapeño across the 10GW Broadcom deployment agreement would make OpenAI a substantial new claimant to HBM4 supplyCITE:E16, at a time when Samsung, SK Hynix, and Micron have sold their HBM capacity through 2027CITE:E18. SK Hynix (SK海力士) CEO Kwak Noh-jung has warned that 2027 will be the worst year of that shortageCITE:E19.
What this means
The size of Jalapeño's advantage depends entirely on which numbers are compared: the headline 1.5x–1.9x throughput-per-kilowatt and 1.7x–3.6x latency gainsCITE:E1 shrink to about 1.5x once all-in utility powerCITE:E7 or Nvidia's multi-token predictionCITE:E8 are factored in, and none of it was measured against Nvidia's upcoming Vera Rubin platformCITE:E20. At the same time, OpenAI is deepening rather than replacing its Nvidia relationship — accepting up to $105 billion in Nvidia financing for its Ohio campusCITE:E13 even as Richard Ho affirms continued Nvidia dependenceCITE:E14 — while its 10GW Broadcom commitmentCITE:E16 arrives just as HBM capacity is sold out industry-wide through 2027CITE:E18CITE:E19.