Key points
- Cerebras builds one giant chip across a silicon wafer
- Memory beside the computing cores helps speed up AI responses
- The design has limits, and manufacturing still depends on TSMC
Cerebras Systems (CBRS) builds an AI chip so large it describes it as "the size of a dinner plate." The point is to fit more computing cores and memory together, reducing the time spent moving data between them. Cerebras says that design lets its systems generate AI responses faster than competing GPU systems.
Its Wafer-Scale Engine 3 measures 46,225 square millimeters, about 58 times the area of one of the two silicon dies inside Nvidia's (NVDA) B200, according to Cerebras's IPO prospectus.
The technology hasn't kept the stock above its IPO price. Cerebras closed at $166.43 on Oct. 2, about 10% below its $185 offering price. More insider shares become eligible for sale on Oct. 14.
How Cerebras makes one giant chip
Chipmakers usually make many separate chips on a round silicon wafer, then cut them apart. Conventional designs stay within the area that lithography equipment can expose at once. Nvidia's Blackwell GPUs, including the B200, join two maximum-size silicon dies with a high-speed link. Together, they contain 208 billion transistors.
Cerebras connects computing cores across those boundaries to make one giant processor. "We used the entire wafer for one chip: a technique called wafer-scale integration," the company said in its prospectus. Manufactured by TSMC (TSM) on its 5-nanometer process, the WSE-3 holds 4 trillion transistors, 900,000 computing cores, and 44 gigabytes of on-chip memory.
Why keeping memory close matters
Generating an AI answer requires repeatedly reading the model's weights, the numerical values learned during training. GPUs typically store those weights in high-bandwidth memory, or HBM, beside the processor. Moving them to the computing cores can become a bottleneck as the model produces tokens, the pieces of text that make up an answer. In Cerebras's example, a 70-billion-parameter model using 16-bit weights needs about 140 gigabytes just to store them.
Cerebras places memory beside the cores on the wafer itself. The company says the WSE-3 has 2,625 times the memory bandwidth of a B200 and that its systems deliver inference up to 15 times faster than leading GPU systems on leading open-source models. Those are different measures: memory bandwidth describes how quickly data can move, while inference speed describes how quickly the system generates an answer.
OpenAI is a major customer for that speed. In January, Cerebras announced a multiyear agreement for OpenAI to deploy 750 megawatts of its computing power. Cerebras valued the deal at more than $20 billion in its prospectus.
Cerebras avoids HBM packaging, but still depends on TSMC
Nvidia's GPUs need an advanced packaging step that mounts the processor and its HBM stacks together, and that capacity has been one of the tightest bottlenecks in AI. A Cerebras wafer has no HBM stacks to attach, so it doesn't need that step.
It doesn't escape the supply chain, though. Cerebras relies on TSMC alone to make its wafers, and it buys almost all of its manufacturing services and parts on purchase orders "without capacity or volume commitments," it said in the prospectus. Bigger, better-financed customers, or ones with long-term agreements, could get its suppliers' capacity first, the company warned. Its own packaging is hard, too. Cerebras said that when it started, "nobody knew how to package such a big chip without cracking it," or how to power and cool one.
The giant chip still has limits
Every wafer has defects, and a flaw would normally ruin a chip this size. Cerebras built in spare capacity instead. The wafer has 970,000 physical cores, with 900,000 turned on, and it routes around the bad ones.
Memory is the other limit. Each wafer has 44 gigabytes of fast on-chip memory. The weights of a 70-billion-parameter model take about 140 gigabytes at 16-bit precision, so keeping them entirely on-chip requires multiple wafers. Lower-precision weights need less space. Nvidia's DGX B200 server, by comparison, holds 1,440 gigabytes of HBM across its eight GPUs.
Cerebras is trying to scale up. Its newest system, the CS-4, combines three upgraded WSE-3 Turbo chips. Cerebras says it delivers inference up to 30 times faster than GPU systems, with results varying by model and configuration. Whether that's enough to win more customers like OpenAI is the question behind the stock.



