OpenAI's Jalapeno Chip Beat Nvidia's GB300 in Testing. Read the Fine Print.

0
277

Here is the headline everyone grabbed this week: OpenAI says its homemade chip beat Nvidia's best in testing. The stock talk started instantly, the takes got loud, and somewhere in the noise a very important detail got buried — the chip that "beat" Nvidia is a 700-watt inference part that cannot train a model, and the numbers come from OpenAI itself. I have run servers long enough to know vendor benchmarks deserve a raised eyebrow. Let's dig into what was actually announced and what it means for the infrastructure underneath the AI gold rush.

What Actually Happened

On August 25, at the Hot Chips conference — the annual gathering where silicon people actually show their work — OpenAI's chip chief Richard Ho presented benchmark results for Jalapeno, the company's first in-house processor, developed with Broadcom and manufactured by TSMC on a 3-nanometer process. The claim: Jalapeno beat Nvidia's GB300 on work-per-watt and response speed at 700 watts. Ho repeated the pitch on Bloomberg Tech the same day. "The chip is really good — it should drop it by a lot."

This is not a rumor from a leak. The chip was unveiled on June 24, with Sam Altman and Broadcom CEO Hock Tan showing off a prototype wafer together. What changed this week is that the numbers are finally public — and they are competitive. Ho noted that GB300 was the top-performing option on the public benchmarking system used for the comparison, which makes the win mean something. On SemiAnalysis' InferenceX test, OpenAI reported the part beating a Blackwell system on tokens per user and throughput per kilowatt.

The Chip: Inference-Only, 700 Watts, Built for the Power Bill

Jalapeno is not a general-purpose GPU. It is an application-specific integrated circuit — an ASIC — designed for exactly one job: inference. That means serving trained models, processing prompts, and generating responses. It is built on TSMC's 3nm process with eight HBM memory stacks integrated on-package and a systolic array architecture tuned for large language model workloads.

The 700-watt figure is the part that matters. Electricity is the dominant running cost in any data center. Nvidia chips do not just cost a fortune to buy; they cost a fortune to feed. A part that delivers the same work per watt — or better — at scale changes the unit economics of inference, and inference is where the real money in AI is being spent. Ho's own framing: "We're at a cost level and a power level that will bring the infrastructure cost down. This is step one."

The Fine Print Nobody Reads

Now the part I care about, because this is where the headline and the reality diverge.

First: these are OpenAI's own numbers. TNW flagged the obvious problem back in June when the chip was unveiled — vendor benchmarks deserve a raised eyebrow until independent results land. Nothing about that changed this week. Second: Jalapeno was not tested against Nvidia's newest hardware. Vera Rubin, the generation that just started shipping, was not in the comparison. Beating last year's part is a milestone; beating the current one is a different story. Third: Jalapeno cannot train models at all. It is a specialist, not a replacement.

The test setup is worth noting too. OpenAI ran the comparison on its own small open model plus third-party models from DeepSeek and Moonshot AI. The widest advantage came on Moonshot's Kimi — the largest model they tried — and they also saw promising results on some unreleased, larger OpenAI models. That is a real signal, but it is also a narrow one.

This Is Not One Chip. It Is the Whole Industry Going Vertical

OpenAI is not alone here, and that is the bigger story everyone keeps skimming past. Anthropic is in talks with Samsung about a custom chip on a 2nm process. Google has been talking to Marvell about custom inference silicon. Meta's Iris chip reportedly passed testing in July — with Broadcom now designing custom parts for three of Nvidia's biggest customers at once.

Add it up: every hyperscaler with serious AI spend is trying to reduce the toll booth that is Nvidia's pricing. In the past month we have seen reports of 15 percent-plus server price increases from Nvidia and a market where memory is now treated as strategic infrastructure. When your suppliers control the margins, you build your own lane. That is not betrayal. That is procurement.

What This Means: The Nvidia Model Is Being Tested on Three Fronts

Nvidia is not going anywhere. Ho was careful to say it himself: "Nvidia is a really good partner, and we continue to need a lot of Nvidia." The company still dominates training, and its software moat — CUDA and the ecosystem built around it — is worth more than any single benchmark.

But the direction is unmistakable. Inference is where the volume is, and inference is where the custom chips are landing. If Jalapeno delivers even half of what OpenAI claims at scale, the cost structure of serving AI changes — and every other company with a chip program just got validation for its own bet.

There is a Europe angle worth watching as well. The EU has committed around 20 billion euros to AI gigafactories that will buy merchant hardware on the open market, weeks after Nvidia warned customers that server prices are climbing. Europe's most prominent AI chip company, Axelera, targets edge devices — not the data centers the gigafactory money is building. So the merchant market stays Nvidia's, for now, while the vertical players carve out their own slices.

What Comes Next

OpenAI plans a small-scale Jalapeno deployment by the end of 2026, scaling through 2027. A second-generation chip tapes out in the coming months, and a third generation exists in concept. The roadmap is real and moving fast.

My take, from someone who watches this industry through a server-room window: believe the direction, not the number. Custom silicon for inference is inevitable — the economics demand it. But "outperformed Nvidia in testing" and "outperformed Nvidia in production at a fraction of the cost" are two very different claims, and only the second one matters when the lights are on and the meter is running. Watch the independent benchmarks, watch the power bills, and watch what Nvidia does next. This story is just getting started.

— Allan Ali, Sylt.ing

Site içinde arama yapın
Kategoriler
Read More
AI News & Updates
Jack Dorsey's Buzz Is Coming for Slack and GitHub — and AI Agents Are the Whole Point
Jack Dorsey's Buzz Is Coming for Slack and GitHub — and AI Agents Are the Whole Point Let me be...
By Allan 2026-07-31 10:37:47 0 863
AI Tools & Software
The AI Pilot Graveyard: Why 88% of Proofs of Concept Never Reach Production - And What the 5% Who Succeed Do Differently
THE AI PILOT GRAVEYARD: WHY 88% OF PROOFS OF CONCEPT NEVER REACH PRODUCTION — AND WHAT THE 5% WHO...
By PriyaSharma 2026-06-30 01:12:37 0 599
AI Business & Monetization
Calculating Tangible Returns in AI Automation Deployments
Measuring Enterprise Returns from Process Automation Initiatives Defining Key Performance Metrics...
By PriyaSharma 2026-07-10 20:41:56 0 2K
Generative AI & AI Art
Designing Event Invitations with AI: The Data-Backed Playbook for Higher Attendance and Lower Costs
Designing Event Invitations with AI: The Data-Backed Playbook for Higher Attendance and Lower...
By Patty 2026-08-23 11:07:11 0 466
AI News & Updates
What Hermes Agent Reveals About Mastering AI Agent Design
What Hermes Agent Reveals About Mastering AI Agent Design Why Hermes Agent Stands Apart in a...
By Jessica 2026-06-17 11:02:33 0 1K