Cerebras (CBRS) bets its giant AI chip can outrun Nvidia (NVDA) GPUs

Cerebras Wafer Scale Engine 3 shown beside the two chips of an Nvidia B200

Image: Cerebras

Key points

  • Cerebras builds one giant chip across a silicon wafer
  • Memory beside the computing cores helps speed up AI responses
  • The design has limits, and manufacturing still depends on TSMC

Cerebras Systems (CBRS) builds an AI chip so large it describes it as "the size of a dinner plate." The point is to fit more computing cores and memory together, reducing the time spent moving data between them. Cerebras says that design lets its systems generate AI responses faster than competing GPU systems.

Its Wafer-Scale Engine 3 measures 46,225 square millimeters, about 58 times the area of one of the two silicon dies inside Nvidia's (NVDA) B200, according to Cerebras's IPO prospectus.

The technology hasn't kept the stock above its IPO price. Cerebras closed at $166.43 on Oct. 2, about 10% below its $185 offering price. More insider shares become eligible for sale on Oct. 14.

How Cerebras makes one giant chip

Chipmakers usually make many separate chips on a round silicon wafer, then cut them apart. Conventional designs stay within the area that lithography equipment can expose at once. Nvidia's Blackwell GPUs, including the B200, join two maximum-size silicon dies with a high-speed link. Together, they contain 208 billion transistors.

Cerebras connects computing cores across those boundaries to make one giant processor. "We used the entire wafer for one chip: a technique called wafer-scale integration," the company said in its prospectus. Manufactured by TSMC (TSM) on its 5-nanometer process, the WSE-3 holds 4 trillion transistors, 900,000 computing cores, and 44 gigabytes of on-chip memory.

Why keeping memory close matters

Generating an AI answer requires repeatedly reading the model's weights, the numerical values learned during training. GPUs typically store those weights in high-bandwidth memory, or HBM, beside the processor. Moving them to the computing cores can become a bottleneck as the model produces tokens, the pieces of text that make up an answer. In Cerebras's example, a 70-billion-parameter model using 16-bit weights needs about 140 gigabytes just to store them.

Cerebras places memory beside the cores on the wafer itself. The company says the WSE-3 has 2,625 times the memory bandwidth of a B200 and that its systems deliver inference up to 15 times faster than leading GPU systems on leading open-source models. Those are different measures: memory bandwidth describes how quickly data can move, while inference speed describes how quickly the system generates an answer.

OpenAI is a major customer for that speed. In January, Cerebras announced a multiyear agreement for OpenAI to deploy 750 megawatts of its computing power. Cerebras valued the deal at more than $20 billion in its prospectus.

Cerebras avoids HBM packaging, but still depends on TSMC

Nvidia's GPUs need an advanced packaging step that mounts the processor and its HBM stacks together, and that capacity has been one of the tightest bottlenecks in AI. A Cerebras wafer has no HBM stacks to attach, so it doesn't need that step.

It doesn't escape the supply chain, though. Cerebras relies on TSMC alone to make its wafers, and it buys almost all of its manufacturing services and parts on purchase orders "without capacity or volume commitments," it said in the prospectus. Bigger, better-financed customers, or ones with long-term agreements, could get its suppliers' capacity first, the company warned. Its own packaging is hard, too. Cerebras said that when it started, "nobody knew how to package such a big chip without cracking it," or how to power and cool one.

The giant chip still has limits

Every wafer has defects, and a flaw would normally ruin a chip this size. Cerebras built in spare capacity instead. The wafer has 970,000 physical cores, with 900,000 turned on, and it routes around the bad ones.

Memory is the other limit. Each wafer has 44 gigabytes of fast on-chip memory. The weights of a 70-billion-parameter model take about 140 gigabytes at 16-bit precision, so keeping them entirely on-chip requires multiple wafers. Lower-precision weights need less space. Nvidia's DGX B200 server, by comparison, holds 1,440 gigabytes of HBM across its eight GPUs.

Cerebras is trying to scale up. Its newest system, the CS-4, combines three upgraded WSE-3 Turbo chips. Cerebras says it delivers inference up to 30 times faster than GPU systems, with results varying by model and configuration. Whether that's enough to win more customers like OpenAI is the question behind the stock.

Frequently asked questions

How big is the Cerebras chip?

The Cerebras Wafer-Scale Engine 3 measures 46,225 square millimeters, 58 times the area of one of the two silicon dies in Nvidia's B200, according to Cerebras's IPO prospectus. It has 4 trillion transistors, 900,000 cores and 44 gigabytes of on-chip memory, and it's made by TSMC on a 5-nanometer process.

Why does Cerebras make such a large chip?

Keeping memory and computing on one piece of silicon cuts down on moving data between memory and processors, which can slow AI models as they generate answers. Cerebras says its chip has 2,625 times more memory bandwidth than an Nvidia B200 and delivers inference up to 15 times faster than leading GPU systems on leading open-source models.

What are the downsides of the Cerebras wafer-scale chip?

Each wafer has 44 gigabytes of on-chip memory. The weights of a 70-billion-parameter model take about 140 gigabytes at 16-bit precision, so keeping them entirely on-chip requires multiple wafers. Cerebras also relies on TSMC alone to make its wafers, without long-term capacity commitments, and it has had to engineer its own packaging, power and cooling for a chip that size.

More on CBRS and NVDA

Dennis Singleton
Dennis Singleton

Dennis Singleton was born in Australia and later moved to the United States. He has spent years following the markets, but what keeps his attention is how AI is built. He writes about the companies behind the technology, from semiconductor designers and advanced packaging to photonics, memory, networking, and the hardware powering modern AI. His approach starts with filings, earnings, and industry research, then translates the important details into clear, straightforward analysis without unnecessary hype.