A day before Nvidia (NVDA) earnings, OpenAI says its Jalapeño chip beats Nvidia (NVDA) on efficiency.

A day before Nvidia (NVDA) earnings, OpenAI says its Jalapeño chip beats Nvidia (NVDA) on efficiency.

Key points

  • OpenAI used the Hot Chips conference on August 25 to publish the first benchmarks for Jalapeño, the inference chip it developed with Broadcom (AVGO). The company says the chip runs more efficiently than Nvidia (NVDA) hardware. Bloomberg and SemiAnalysis reported on the presentation that day
  • The headline numbers: more than 700 tokens per second per user on DeepSeek R1, and about 1,400 on Kimi K2.5 and GPT-OSS, with output-token throughput per megawatt that SemiAnalysis says tops Nvidia's Vera Rubin.
  • The catch: the numbers are OpenAI's own, run on an easier 8k-in, 1k-out workload, on first-run A0 silicon. On cost per token, SemiAnalysis calls Jalapeño and Vera Rubin a dead heat.
  • Timing matters. This landed one day before Nvidia's earnings on August 26, and it's still a 2027 story. Vera Rubin is shipping now, and Jalapeño won't be in real volume until late next year.

OpenAI picked a pointed moment to show its cards. On Tuesday, August 25, at the Hot Chips conference, it presented the first public benchmarks for Jalapeño, the custom inference chip it developed with Broadcom (AVGO), and its pitch was direct. OpenAI says that on the inference workloads it tested, Jalapeño outperforms Nvidia (NVDA) hardware. Bloomberg put it plainly, that OpenAI says its new chips can outperform Nvidia's in tests. The chip research shop SemiAnalysis, which watched the runs in person, published its own breakdown the same day. All of it landed one day before Nvidia reports earnings on Wednesday.

The reported gains are substantial. In OpenAI's tests, Jalapeño ran DeepSeek R1 at more than 700 tokens per second per user at a concurrency of one, the speed a single user actually feels, and ran Kimi K2.5 and GPT-OSS at roughly 1,400. SemiAnalysis singled out Kimi K2.5, where at single-user interactivity Jalapeño ran many times faster than the next best chip it has tested. On output tokens per megawatt, the metric that decides how much a data center can serve for its power bill, SemiAnalysis said Jalapeño edged out Nvidia's next-generation Vera Rubin.

"The bottom line is that the results show a very, very significant performance advance over state of the art," OpenAI hardware head Richard Ho said. "Jalapeño can serve more AI work per unit of power, while also returning responses more quickly."

Here's where it gets more honest. Every one of those numbers came from OpenAI, not from an independent run. SemiAnalysis verified the tests in the lab in person, but it was careful about what that means. "All numbers are provided to us by OpenAI. We verified the InferenceX runs in person in the lab, but we did not run the full suite of InferenceX benchmarks nor have we seen AgentX results," the firm wrote. The tests used a single-turn 8,000-token input and 1,000-token output, which SemiAnalysis called an easier workload to tune for. The harder long-context, multi-turn agent workloads that stress a real serving stack haven't been run yet, and the results come from A0 silicon, the very first version of the chip.

The number that should cool the "Nvidia killer" talk is cost. On output tokens per dollar, SemiAnalysis called Jalapeño and Vera Rubin "head-to-head, producing almost the same number of output tokens per $." An efficiency edge per watt is real, but on the metric that decides what it costs to answer a prompt, this is a tie, not a rout. The firm's own read was measured: "We don't think that Nvidia hardware is inferior, but more so that Jalapeño's software bring-up has progressed more quickly than Nvidia's."

Timing is the other half of the story. Vera Rubin systems are shipping to customers right now. Jalapeño, by OpenAI's own account, reaches the end of 2026 in very small volumes, with production ramping through 2027 and most of the output not landing until late next year. So the chip behind these benchmarks won't exist in quantity for more than a year, and by then Nvidia's product roadmap will also have advanced. On the hardware itself, Jalapeño is built on TSMC's (TSM) N3P process with HBM4 memory at 15.4 terabytes per second, and packs 128 chips to a rack. A second version already in the fab, called B0, is expected to add about 25% more performance per watt.

What it means for the stocks

The near-term piece is sentiment. A story headlined "OpenAI's chip beats Nvidia," carried by Bloomberg the day before Nvidia reports, is exactly the kind of thing that can rattle a stock going into a print, even when the fine print is far softer than the headline. It's worth keeping straight what the benchmark actually showed. A faster software rollout and a per-watt edge on an easier workload, on first-run silicon, from the chip's own maker, that ties Nvidia on cost and won't ship in volume until 2027. That is not the same thing as a cheaper, better chip you can buy today.

The longer piece is the one that's been building for two years. OpenAI, one of Nvidia's biggest and most visible customers, now has real silicon that competes with Nvidia's best on the workload OpenAI cares about most. It doesn't have to win outright to matter. It only has to be good enough to move some of OpenAI's own inference off Nvidia GPUs over time, and today's numbers say it might. As before, the cleaner public way to play it is Broadcom, which co-designs the chip and earns the custom-silicon revenue, along with the TSMC and memory supply chain that gets paid no matter whose name is on the chip.

Nvidia's moat doesn't break on a conference slide. But this is the most credible shot at it yet, and the timing ensured the announcement got attention the day before Nvidia had to speak.

Frequently asked questions

What did OpenAI reveal about Jalapeño at Hot Chips?

On August 25, 2026, OpenAI presented the first benchmarks for Jalapeño, its custom inference chip developed with Broadcom (AVGO), and said that on the inference workloads it tested, the chip beats Nvidia (NVDA) hardware on efficiency. It showed more than 700 tokens per second per user on DeepSeek R1 and about 1,400 on Kimi K2.5 and GPT-OSS.

Are the Jalapeño benchmarks independent?

No. SemiAnalysis, which watched the tests in person, said the numbers were all provided by OpenAI and that it did not run the full benchmark suite or see the harder long-context AgentX results. The tests used an easier 8k-input, 1k-output workload on first-run A0 silicon.

Does Jalapeño actually beat Nvidia on cost?

Not really. SemiAnalysis found Jalapeño and Nvidia's Vera Rubin were head-to-head on output tokens per dollar. Jalapeño's edge is efficiency per watt and a faster software rollout, not a lower cost per token. This is general information, not investment advice.

When will the Jalapeño chip actually ship?

OpenAI says it reaches the end of 2026 in very small volumes, with production ramping through 2027 and most output not until late next year. Nvidia's Vera Rubin systems are already shipping to customers now.

Which public stock benefits from Jalapeño?

Broadcom (AVGO), which co-designs the chip and earns the custom-silicon revenue, along with the TSMC (TSM) and high-bandwidth memory supply chain behind it. OpenAI is private, so there is no direct way to own it. This is general information, not investment advice.

More on NVDA and AVGO

David Han
David Han

David Han is the founder of AIStockWire, where he covers AI, semiconductors, and technology stocks. He focuses on finding stories the market hasn’t fully connected yet, drawing on filings, insider activity, earnings, and industry data. His commentary has been quoted by U.S. News & World Report, Moneywise, and Yahoo Finance. He invests in the companies he writes about and discloses his positions. Nothing he publishes is investment advice.