Hero image for OpenAI's Jalapeño Chip Just Beat Nvidia Blackwell
By AI Tool Briefing Team

OpenAI's Jalapeño Chip Just Beat Nvidia Blackwell


At the Hot Chips conference on August 25, OpenAI published the first independent benchmark results for Jalapeño, its in-house inference chip — and for the first time, this isn’t a rendering or a roadmap slide. Tested on SemiAnalysis’s public InferenceX benchmark against Nvidia’s Blackwell-generation GB200 and GB300 systems, Jalapeño delivered 1.5 to 1.9 times more work per watt and up to 4.1 times faster response times on the kind of chatty, back-and-forth workload ChatGPT actually runs. Six days earlier, Anthropic hired the founder of Google’s TPU program to build its own competing silicon. The “every AI lab wants off Nvidia” story just stopped being a rumor and started being a spec sheet.

Quick Summary: What Happened

DetailInfo
AnnouncedAugust 25, 2026, at Hot Chips, Stanford University
ChipJalapeño, OpenAI’s first custom “Intelligence Processor,” co-developed with Broadcom
Benchmark usedSemiAnalysis InferenceX, a public, independently run suite
Baseline comparedNvidia GB200/GB300 (Blackwell-generation) rack systems
Power efficiency1.5x–1.9x more work per watt
End-to-end latency1.7x–3.6x lower
Interactive/chatty workloads2.1x–4.1x faster
Deployment timelineSmall-volume by end of 2026; larger-scale in 2027
Related newsAnthropic hired ex-Google TPU chief Amir Salek on Aug 21 to lead its own chip push

Bottom line: OpenAI has real, third-party-verified numbers showing its own chip beating Nvidia’s current generation on the metric that actually matters for a chatbot — and it landed the same week a second major lab confirmed it’s building silicon of its own.


What Actually Happened

OpenAI first announced the Jalapeño partnership with Broadcom back in October 2025, and up to now that’s all it was: a partnership announcement, a name, and a promise. That changed at Hot Chips. OpenAI put actual numbers next to actual competitors on a benchmark it doesn’t control.

The results, run on SemiAnalysis’s InferenceX suite across three models — GPT-OSS-120B, DeepSeek R1 (670B), and Kimi K2.5 — showed Jalapeño beating Nvidia’s GB200 and GB300 systems on both axes that matter for inference: how much work the chip does per watt of power, and how fast it returns an answer. On GPT-OSS-120B specifically, Jalapeño hit roughly 85,448 mixed tokens-per-second per kilowatt against Blackwell’s 44,960 — call it 1.9x. On DeepSeek R1, the gap was about 1.7x on throughput and 3.6x on latency (1.65 seconds versus 5.99 seconds to first meaningful output). The advantage widened further, to 2.1x–4.1x, on the highly interactive, low-latency scenarios that map most closely onto how people actually use ChatGPT: short prompts, fast back-and-forth, lots of concurrent users.

That’s the headline. The context matters too: a single Jalapeño rack packs 128 accelerators, 1.7 exaFLOPS of 4-bit compute, and 27.5 TB of HBM4 memory at just under 2 petabytes per second of bandwidth — and it does it at a 700W-per-chip power budget against Blackwell parts rated at 1,200W to 1,400W. Richard Ho, OpenAI’s VP of hardware, put it plainly: “The bottom line is that the results show a very, very significant performance advance,” adding that Jalapeño can “serve more AI work per unit of power, while also returning responses more quickly” — the two things chip vendors usually trade off against each other, not stack.

OpenAI also said Jalapeño went from initial design to manufacturing tape-out in nine months, which it credits partly to using its own models to accelerate parts of the design process. Deployment is modest to start: very small volumes by the end of 2026, scaling up through 2027. This isn’t replacing Nvidia hardware in OpenAI’s data centers next quarter. It’s the first real evidence that it eventually could, for at least part of the workload.

What Did Jalapeño’s Benchmark Results Actually Show?

Stripped down to the numbers that got repeated across every outlet covering Hot Chips, OpenAI’s own claims for Jalapeño versus Nvidia’s Blackwell-generation GB200/GB300 systems break into three figures:

  1. 1.5x to 1.9x more AI work per watt — the core efficiency metric, measured across GPT-OSS-120B, DeepSeek R1, and Kimi K2.5.
  2. 1.7x to 3.6x lower end-to-end latency — how long a request takes from prompt to completed response.
  3. 2.1x to 4.1x faster on highly interactive workloads — the ChatGPT-shaped use case, where speed and concurrency matter more than raw throughput.

All three numbers come from OpenAI’s own results run on SemiAnalysis’s independently operated InferenceX benchmark, not an internal-only test.

Why This Matters

If you use ChatGPT or build on the OpenAI API, this is the closest thing to a leading indicator you’ll get for where speed and pricing head next. Inference cost is the single biggest lever behind what a chatbot response costs to serve and how fast it comes back to you. A chip that does meaningfully more work per watt, at meaningfully lower latency, is a chip that lets OpenAI either drop prices, raise usage limits, or push more reasoning into every response without the bill exploding — assuming the small-volume 2026 deployment actually scales the way OpenAI is projecting for 2027.

None of that shows up in your ChatGPT plan today. Small-volume deployment by the end of this year means Jalapeño is not the chip answering your prompts right now, and won’t be for a while. But the direction is the story: OpenAI has spent two years as one of Nvidia’s largest customers. It just published proof, from a benchmark it doesn’t operate, that it can build something competitive with Nvidia’s own current-generation hardware for the specific job of running its own models.

It’s also worth being precise about what this comparison is and isn’t. Analysts who’ve looked closely, including SemiAnalysis’s Dylan Patel, called it unusual for a first-generation custom chip to beat a merchant-silicon leader like Nvidia at all — most hyperscaler first attempts land somewhere behind. But the baseline here is GB200/GB300, Nvidia’s 2024-2025 Blackwell generation. Nvidia’s newer Vera Rubin platform, which Nvidia detailed at GTC 2026 promising 3.3x to 5x inference gains over Blackwell Ultra, had already started shipping by the time these numbers came out and wasn’t part of the test. Comparing a brand-new chip to a shipping generation that’s about to be a generation behind is standard practice — SemiAnalysis and OpenAI both flagged it — but it’s the caveat to hold onto before assuming Jalapeño “beats Nvidia” in any permanent sense.

The Bigger Picture

Jalapeño didn’t land in isolation. Six days before the Hot Chips results, Bloomberg reported that Anthropic had hired Amir Salek, who ran Google’s TPU program and helped ship its first seven generations of chips, to join its compute team reporting to James Bradbury. Salek isn’t a random hire, either — he spent eight years at Nvidia before Google, which means Anthropic just recruited someone who has now built chips at both the company everyone buys from and the company that proved a hyperscaler could build its own credible alternative.

Anthropic had already confirmed on August 5 that it’s standing up an in-house silicon team, with the explicit stated goal of co-designing chips and Claude models together to cut inference costs by roughly half. It also hired Clive Chan, an early member of OpenAI’s own custom-chip effort, back in June. Put those together and the picture is less “Anthropic might build a chip someday” and more “Anthropic is actively assembling the same kind of team OpenAI used to ship Jalapeño.”

That’s the pattern now across the industry, not a one-off. Google has run its own TPUs for years and supplies them to Anthropic alongside Nvidia and Amazon chips. Amazon has Trainium. OpenAI has Jalapeño. Anthropic is now visibly building toward something of its own, on top of existing deals with UK chip startup Fractile and compute capacity agreements with Riot Platforms and Volta Infra Holdings. Every major AI lab either already has custom silicon or just publicly staffed up to build it. Nvidia’s own customers are becoming its competitors, one hire and one benchmark at a time.

Reaction from chip analysts has split, which is itself useful signal. Yole Group’s Adrien Sanchez called the results a “threat to Nvidia’s inference margins,” per CNBC’s coverage — inference being the fastest-growing, and increasingly most contestable, part of Nvidia’s business. Atreides Management’s Gavin Baker called Jalapeño the “first good ASIC outside of TPU/Trainium,” a real compliment given how many hyperscaler chip projects have quietly underdelivered, while noting it could still lose to more specialized disaggregated systems on some workloads. Not everyone’s convinced: Nvidia dominates AI training outright, and Jalapeño isn’t for sale or rent outside OpenAI — it only ever serves OpenAI’s own workloads, which caps how much market share it can actually take even if the performance numbers hold up at scale.

Our Take

We think the benchmark is the real story here and the “beats Nvidia” framing in a lot of headlines oversells it. Jalapeño winning on an independent benchmark, against a shipping Nvidia generation, on the specific metrics that determine chatbot economics, is genuinely notable — first-generation custom silicon usually doesn’t get there. That’s not spin; SemiAnalysis runs InferenceX independently and multiple outlets not just running an OpenAI press release confirmed the same numbers.

What we’d push back on is the leap from “beat Blackwell on a benchmark” to “Nvidia has a problem.” Jalapeño isn’t a product Nvidia loses a sale to — OpenAI isn’t buying fewer GPUs next quarter because of it, and small-volume 2026 deployment means the actual fleet impact is close to zero this year. What it does change is the multi-year story: every lab that used to be a pure Nvidia customer is now also a potential competitor building the exact class of chip that eats into Nvidia’s highest-margin, fastest-growing business. Anthropic hiring Amir Salek six days before this benchmark landed isn’t a coincidence of timing so much as confirmation that the rest of the industry read the writing on the wall well before OpenAI published the proof.

For anyone building on the OpenAI API or paying for ChatGPT: nothing changes in your bill or your latency today. What’s worth tracking is 2027, when Jalapeño’s larger-scale deployment is supposed to start. That’s the point where “OpenAI published a good benchmark” either turns into “your API calls got cheaper and faster” or it doesn’t, and you’ll know which within a year.

Frequently Asked Questions

Q: What is OpenAI’s Jalapeño chip? A: Jalapeño is OpenAI’s first custom inference chip, co-developed with Broadcom and first announced in October 2025. OpenAI calls it an “Intelligence Processor” — hardware purpose-built to run inference for its own models rather than general-purpose GPU work.

Q: How much faster is Jalapeño than Nvidia’s Blackwell chips? A: On SemiAnalysis’s InferenceX benchmark, OpenAI reported 1.5x–1.9x more work per watt, 1.7x–3.6x lower end-to-end latency, and 2.1x–4.1x faster performance on highly interactive, ChatGPT-style workloads, compared to Nvidia’s GB200/GB300 Blackwell-generation systems.

Q: Was this benchmark run by OpenAI or an independent party? A: OpenAI ran the tests, but on SemiAnalysis’s InferenceX, a benchmark suite SemiAnalysis operates and publishes independently — not an internal-only OpenAI metric.

Q: When will Jalapeño actually power ChatGPT or the OpenAI API? A: OpenAI says small-volume deployment starts by the end of 2026, with larger-scale rollout in 2027. It’s not running production traffic at meaningful scale yet.

Q: Can other companies buy or rent Jalapeño chips? A: No. OpenAI has no announced plans to sell or rent Jalapeño externally — it’s built to run OpenAI’s own workloads, similar to how Google’s TPUs primarily serve Google’s own infrastructure.

Q: Why did Anthropic hire Amir Salek? A: Salek founded and ran Google’s TPU program through its first seven generations before spending time at Nvidia and Cerberus Capital Management. Anthropic hired him to its compute team, reporting to James Bradbury, as part of an in-house silicon effort it confirmed on August 5, aimed at cutting inference costs by roughly half through co-designed chips and models.

Q: Does this mean Nvidia is losing its lead in AI chips? A: Not in training, where Nvidia’s dominance is largely unchallenged. In inference specifically, analysts are split — some call Jalapeño a genuine “threat to Nvidia’s inference margins,” others note Nvidia’s newer Vera Rubin platform wasn’t part of this comparison and that Jalapeño only serves OpenAI’s own workloads rather than competing for external sales.

Q: Will this make ChatGPT cheaper? A: Not immediately. Any pricing or speed impact depends on Jalapeño’s 2027 larger-scale deployment actually delivering the efficiency gains shown in these early benchmarks across OpenAI’s full production traffic.


Last updated: August 27, 2026. Sources: OpenAI — Jalapeño’s first results · SemiAnalysis — OpenAI Jalapeño: Better Than Nvidia Blackwell · CNBC — OpenAI’s Jalapeño AI chip brings new ‘threat’ to Nvidia margins · TechCrunch — OpenAI’s Jalapeño chip is built for fast inference at scale · The Register — OpenAI’s upcoming Jalapeño chip looks like it’ll be an inference beast · Bloomberg — Anthropic Taps Google Chip Veteran as Part of Push Into Hardware · Broadcom — OpenAI and Broadcom Unveil LLM-Optimized Intelligence Processor.

Related reading: NVIDIA GTC 2026: Vera Rubin, NVIDIA Agent Toolkit, and Agentic AI · Anthropic’s IPO Could Top SpaceX. Here’s the Math · OpenAI Files for IPO · Amazon Bets $200B on AI: What Changes for Users · Google’s $40B Anthropic Bet: What Changes for Claude