Key points
- OpenAI used the Hot Chips conference on August 25 to publish the first benchmarks for Jalapeño, the inference chip it developed with Broadcom (AVGO). The company says the chip runs more efficiently than Nvidia (NVDA) hardware. Bloomberg and SemiAnalysis reported on the presentation that day
- The headline numbers: more than 700 tokens per second per user on DeepSeek R1, and about 1,400 on Kimi K2.5 and GPT-OSS, with output-token throughput per megawatt that SemiAnalysis says tops Nvidia's Vera Rubin.
- The catch: the numbers are OpenAI's own, run on an easier 8k-in, 1k-out workload, on first-run A0 silicon. On cost per token, SemiAnalysis calls Jalapeño and Vera Rubin a dead heat.
- Timing matters. This landed one day before Nvidia's earnings on August 26, and it's still a 2027 story. Vera Rubin is shipping now, and Jalapeño won't be in real volume until late next year.
OpenAI picked a pointed moment to show its cards. On Tuesday, August 25, at the Hot Chips conference, it presented the first public benchmarks for Jalapeño, the custom inference chip it developed with Broadcom (AVGO), and its pitch was direct. OpenAI says that on the inference workloads it tested, Jalapeño outperforms Nvidia (NVDA) hardware. Bloomberg put it plainly, that OpenAI says its new chips can outperform Nvidia's in tests. The chip research shop SemiAnalysis, which watched the runs in person, published its own breakdown the same day. All of it landed one day before Nvidia reports earnings on Wednesday.
The reported gains are substantial. In OpenAI's tests, Jalapeño ran DeepSeek R1 at more than 700 tokens per second per user at a concurrency of one, the speed a single user actually feels, and ran Kimi K2.5 and GPT-OSS at roughly 1,400. SemiAnalysis singled out Kimi K2.5, where at single-user interactivity Jalapeño ran many times faster than the next best chip it has tested. On output tokens per megawatt, the metric that decides how much a data center can serve for its power bill, SemiAnalysis said Jalapeño edged out Nvidia's next-generation Vera Rubin.
"The bottom line is that the results show a very, very significant performance advance over state of the art," OpenAI hardware head Richard Ho said. "Jalapeño can serve more AI work per unit of power, while also returning responses more quickly."
Here's where it gets more honest. Every one of those numbers came from OpenAI, not from an independent run. SemiAnalysis verified the tests in the lab in person, but it was careful about what that means. "All numbers are provided to us by OpenAI. We verified the InferenceX runs in person in the lab, but we did not run the full suite of InferenceX benchmarks nor have we seen AgentX results," the firm wrote. The tests used a single-turn 8,000-token input and 1,000-token output, which SemiAnalysis called an easier workload to tune for. The harder long-context, multi-turn agent workloads that stress a real serving stack haven't been run yet, and the results come from A0 silicon, the very first version of the chip.
The number that should cool the "Nvidia killer" talk is cost. On output tokens per dollar, SemiAnalysis called Jalapeño and Vera Rubin "head-to-head, producing almost the same number of output tokens per $." An efficiency edge per watt is real, but on the metric that decides what it costs to answer a prompt, this is a tie, not a rout. The firm's own read was measured: "We don't think that Nvidia hardware is inferior, but more so that Jalapeño's software bring-up has progressed more quickly than Nvidia's."
Timing is the other half of the story. Vera Rubin systems are shipping to customers right now. Jalapeño, by OpenAI's own account, reaches the end of 2026 in very small volumes, with production ramping through 2027 and most of the output not landing until late next year. So the chip behind these benchmarks won't exist in quantity for more than a year, and by then Nvidia's product roadmap will also have advanced. On the hardware itself, Jalapeño is built on TSMC's (TSM) N3P process with HBM4 memory at 15.4 terabytes per second, and packs 128 chips to a rack. A second version already in the fab, called B0, is expected to add about 25% more performance per watt.
What it means for the stocks
The near-term piece is sentiment. A story headlined "OpenAI's chip beats Nvidia," carried by Bloomberg the day before Nvidia reports, is exactly the kind of thing that can rattle a stock going into a print, even when the fine print is far softer than the headline. It's worth keeping straight what the benchmark actually showed. A faster software rollout and a per-watt edge on an easier workload, on first-run silicon, from the chip's own maker, that ties Nvidia on cost and won't ship in volume until 2027. That is not the same thing as a cheaper, better chip you can buy today.
The longer piece is the one that's been building for two years. OpenAI, one of Nvidia's biggest and most visible customers, now has real silicon that competes with Nvidia's best on the workload OpenAI cares about most. It doesn't have to win outright to matter. It only has to be good enough to move some of OpenAI's own inference off Nvidia GPUs over time, and today's numbers say it might. As before, the cleaner public way to play it is Broadcom, which co-designs the chip and earns the custom-silicon revenue, along with the TSMC and memory supply chain that gets paid no matter whose name is on the chip.
Nvidia's moat doesn't break on a conference slide. But this is the most credible shot at it yet, and the timing ensured the announcement got attention the day before Nvidia had to speak.



