SemiconductorsBRIEF

OpenAI's 700W Jalapeño Chip Claims Up to 1.9x Efficiency Over Nvidia's GB300 in First Published Benchmarks

E
EffectStory 編輯部Editorial Team
Published · Updated
OpenAI says its Broadcom-built Jalapeño chip delivered 1.5x-1.9x more throughput per kilowatt and 1.7x-3.6x lower latency than Nvidia's GB200/GB300 on SemiAnalysis's InferenceX suite, though the lead narrows to about 1.5x under utility-power and multi-token-prediction comparisons, and Nvidia's upcoming Vera Rubin platform was not tested.

How much faster is Jalapeño than Nvidia's flagship GPUs?

OpenAI says its Jalapeño chip delivered 1.5 times to 1.9 times more throughput per kilowatt and 1.7 times to 3.6 times lower end-to-end latency than Nvidia's GB200 and GB300 rack systemsCITE:E1. The figures come from SemiAnalysis's public InferenceX benchmark suiteCITE:E1. The comparison pits a 700W Jalapeño part against accelerators rated at 1,200W (GB200) and 1,400W (GB300)CITE:E3. Testing covered three open models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's 1-trillion-parameter Kimi K2.5CITE:E4. OpenAI reports its widest leads at low-latency operating points, where it claims 8.6 times to 104.3 times more throughput per kilowatt versus the GB300's fastest previous time-between-tokens settingsCITE:E5.

MetricJalapeñoGB200GB300
Rated power700W1,200W1,400W
Measured sustained power≤550W
Throughput per kW vs GB200/GB3001.5x–1.9x higherbaselinebaseline
End-to-end latency1.7x–3.6x lowerbaselinebaseline
Throughput per kW at low-latency settings8.6x–104.3x higherbaseline
All-in utility power per accelerator1.18kW2.55kW
HBM memory216 GiB HBM4 @ 15.4 TB/s288GB HBM3E

What power did Jalapeño actually draw during testing?

OpenAI said Jalapeño's measured sustained power stayed at or below 550W in testing, below its 700W ratingCITE:E6. That figure applies across the same three-model test set — GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5CITE:E4 — where OpenAI's steepest per-kilowatt margins over the GB300 appeared at low-latency settings, reaching up to 104.3 timesCITE:E5.

Do the efficiency gains hold under different test conditions?

OpenAI's own appendix data shows the efficiency gap narrows once power accounting changesCITE:E7. A comparison using all-in utility power per accelerator — 1.18kW for Jalapeño against 2.55kW for the GB300 — produces narrower gaps than the headline rated-power comparisonCITE:E7. Pitting Jalapeño against a GB300 running multi-token prediction shrinks the peak efficiency lead to roughly 1.5 timesCITE:E8.

What does an independent evaluator say about Jalapeño?

SemiAnalysis, which ran InferenceX together with OpenAI engineers in the company's lab, described the chip as "beating every Nvidia, AMD, and Google chip we have been able to test"CITE:E9.

How does Jalapeño's memory compare to the GB300's?

Each Jalapeño package pairs its compute die with six HBM4 stacks, totaling 216 GiB at 15.4 TB/s, versus the GB300's 288GB of HBM3E at a 1,400W ratingCITE:E10CITE:E11. Per watt of rated power, OpenAI's chip packs roughly 50% more memory than the GB300CITE:E11. Memory capacity carries a manufacturing cost: Micron told the Hot Chips conference on August 23 that HBM consumes roughly three times the wafer area of DDR5 for equivalent capacity, a penalty that widens with each generationCITE:E17.

What is OpenAI's stance on Nvidia given its Broadcom partnership?

OpenAI signed a 10GW custom AI chip deployment agreement with Broadcom last OctoberCITE:E16, and separately, on August 17, Nvidia agreed to provide up to $105 billion in financing for an OpenAI-leased data center campus in OhioCITE:E13. Despite fielding its own chip, Richard Ho, OpenAI's vice president of hardware, told Bloomberg following the announcement: "Nvidia is a really good partner, and we continue to need a lot of Nvidia"CITE:E14.

What's next for Jalapeño's development roadmap?

Jalapeño was unveiled in June after a nine-month RTL-to-tapeout cycleCITE:E10. According to Bloomberg, a second-generation chip is approaching tapeout, expected within months, and concept work on a third generation is already underwayCITE:E15.

What deployment risks and supply pressure could affect the rollout?

OpenAI plans to begin deploying Jalapeño in its own data centers later this year, but it has not yet been benchmarked against Nvidia's next platformCITE:E12. Jalapeño wasn't tested against Vera Rubin, the Nvidia platform slated to power the first gigawatt of Nvidia systems OpenAI agreed to deploy in the second half of 2026CITE:E20. Scaling Jalapeño across the 10GW Broadcom deployment agreement would make OpenAI a substantial new claimant to HBM4 supplyCITE:E16, at a time when Samsung, SK Hynix, and Micron have sold their HBM capacity through 2027CITE:E18. SK Hynix (SK海力士) CEO Kwak Noh-jung has warned that 2027 will be the worst year of that shortageCITE:E19.

What this means

The size of Jalapeño's advantage depends entirely on which numbers are compared: the headline 1.5x–1.9x throughput-per-kilowatt and 1.7x–3.6x latency gainsCITE:E1 shrink to about 1.5x once all-in utility powerCITE:E7 or Nvidia's multi-token predictionCITE:E8 are factored in, and none of it was measured against Nvidia's upcoming Vera Rubin platformCITE:E20. At the same time, OpenAI is deepening rather than replacing its Nvidia relationship — accepting up to $105 billion in Nvidia financing for its Ohio campusCITE:E13 even as Richard Ho affirms continued Nvidia dependenceCITE:E14 — while its 10GW Broadcom commitmentCITE:E16 arrives just as HBM capacity is sold out industry-wide through 2027CITE:E18CITE:E19.

📊 Evidence

FAQ

How much faster is Jalapeño than Nvidia's flagship GPUs?

OpenAI says its Jalapeño chip delivered 1.5 times to 1.9 times more throughput per kilowatt and 1.7 times to 3.

What power did Jalapeño actually draw during testing?

OpenAI said Jalapeño's measured sustained power stayed at or below 550W in testing, below its 700W ratingCITE:E6.

Do the efficiency gains hold under different test conditions?

OpenAI's own appendix data shows the efficiency gap narrows once power accounting changesCITE:E7. A comparison using all-in utility power per accelerator — 1.

What does an independent evaluator say about Jalapeño?

SemiAnalysis, which ran InferenceX together with OpenAI engineers in the company's lab, described the chip as "beating every Nvidia, AMD, and Google chip we hav…

📎 Sources

  1. tomshardware.com

Related data

Author's TakeEffectStory 編輯部

The headline 1.9x figure is a nameplate-power comparison; once all-in utility power and Nvidia's multi-token prediction are included, Jalapeño's edge compresses to roughly 1.5x, and OpenAI still excluded Vera Rubin from the test entirely. That gap between the marketing number and the appendix data matters more than the benchmark itself, especially since Richard Ho says OpenAI still needs Nvidia at scale even as it accepts $105 billion in Nvidia financing for its Ohio campus. The indicator worth watching is what happens when Jalapeño's second generation approaches tapeout within months — if OpenAI publishes a Vera Rubin comparison then, the 10GW Broadcom ramp becomes a real substitution story rather than a footnote inside a deepening GPU partnership.

E
EffectStory 編輯部Editorial Team

Related

BRIEF

Android's Motion Assist Adds Moving Dots to Fight Car Sickness

Google is rolling out Motion Assist, a Play Services feature for Android 17 that places moving dots along the screen edges to track a vehicle's motion using the phone's accelerometer and gyroscope, aiming to reduce or eliminate motion sickness — two years after Apple shipped a similar Motion Cues feature in iOS 18.

EffectStory 編輯部 ·
BRIEF

OpenAI Restores 5-Hour Usage Limit for ChatGPT Plus's Codex and Work, Effective August 25

OpenAI reinstated the 5-hour rolling usage limit for Codex and ChatGPT Work on Plus accounts starting August 25, 2026, restoring a dual-limit system that pairs the 5-hour cap with the existing weekly quota. OpenAI's Codex and ChatGPT lead Thibault Sottiaux confirmed the change on X, linking it to a usage surge after the GPT-5.6 Sol launch. Pro, Enterprise, and Edu plans remain unaffected for now.

EffectStory 編輯部 ·