Cerebras Sidesteps AI Chip Supply Bottlenecks With 5nm Wafer-Scale Design
Key Takeaways
- •Cerebras' WSE-3 wafer-scale chip avoids HBM, CoWoS packaging, and 3nm manufacturing by using TSMC's 5nm process with all memory on-chip as SRAM, sidestepping three major semiconductor supply bottlenecks.
- •The company reported a $25.4 billion backlog as of June 2026, with more than $20 billion tied to a multi-year agreement with OpenAI.
- •Cerebras has over 600 MW of data-center capacity either live or contracted, and its Q2 revenue nearly doubled year-over-year.
- •In September 2026, Cerebras agreed to supply CS-4 systems for approximately 100 MW to Gimlet Labs, with a separate long-term General Compute deployment planned for the first quarter of 2027.
- •CEO Andrew Feldman claims the architecture delivers the fastest inference in the world by an order of magnitude, though the 44 GB on-chip SRAM limit and scaling for the largest models are expected to face scrutiny from real-world workloads.

The AI hardware race has largely been fought over scarce parts, and Cerebras Systems is trying to win it by needing fewer of them. The company's wafer-scale chips dispense with high-bandwidth memory (HBM), advanced CoWoS packaging, and 3nm manufacturing—three of the tightest chokepoints in the semiconductor industry right now. Those are the inputs most rival AI accelerator designs depend on, and the reason queues for them shape who can build cutting-edge hardware, and when.
A chip built around what it leaves out
Cerebras took a different route from most of the industry. Its WSE-3 and the derivative WSE-3T are built on TSMC's 5nm process node, and all of their memory lives on the chip itself as SRAM rather than in separate HBM stacks. Where a conventional accelerator is a small die cut from one part of a silicon wafer, the wafer-scale approach uses nearly the whole wafer as a single chip—the source of the design's name.
The WSE-3 was announced in March 2024, and its specifications read more like a city plan than a chip sheet: 900,000 AI cores, 44 GB of on-chip SRAM, and approximately 4 trillion transistors, all sitting within a 46,225 mm² area.
Chief Executive Andrew Feldman has argued that the design pays off where it counts most for customers: speed when AI models generate responses. Speaking in June 2026, he described the architecture as delivering "the fastest inference in the world by an order of magnitude."
According to the research findings, Feldman's claim centers on boosting token-generation speed while removing the slowdowns typical of GPU-based designs.
The backlog behind the bet
As of June 2026, the company reported a backlog of $25.4 billion. More than $20 billion of that total comes a multi-year deal with OpenAI.
Cerebras also says it has over 600 MW of data-center capacity either live or contracted, and its Q2 revenue nearly doubled year-over-year, according to the research findings.
In September 2026, Cerebras agreed to supply CS-4 systems for approximately 100 MW of capacity to Gimlet Labs. The company also struck a long-term agreement with General Compute, with deployment planned for the first quarter of 2027.
All of this follows Cerebras' Nasdaq IPO in May 2026. Since listing, the company has begun expanding US manufacturing and building out partnerships to increase its operational capacity.
Why skipping the queue matters
The research findings note that AI infrastructure demand increasingly hinges on data-center space, power, and construction, not just on whether chips are available. That helps explain why Cerebras now talks about capacity in megawatts rather than chip counts: the 600 MW figure and the roughly 100 MW Gimlet Labs commitment are measures of power and floor space. By sidestepping HBM, CoWoS, and 3nm allocation, the company swaps dependence on chip-industry supply lines for the industry-wide constraints the findings identify—power, floor space, and construction capacity.
What to watch
For Cerebras investors, the core question is execution. A $25.4 billion backlog is a promise, not revenue, and converting it depends on getting data centers built, powered, and running on schedule. The research findings point to data-center ramp costs as a challenge the company is working to overcome. The roughly 100 MW Gimlet Labs rollout and the General Compute deployment targeted for the first quarter of 2027 give that execution question near-term checkpoints.
Customer concentration is the second factor to track. With more than $20 billion of the backlog tied to OpenAI, the health of that single relationship carries outsized weight. Deals like the Gimlet Labs and General Compute agreements help diversify the mix, though they are much smaller by comparison.
On performance, Feldman's order-of-magnitude inference claim will face scrutiny from customers running real workloads. On-chip SRAM is extremely fast, but 44 GB per chip is a fixed budget, and how systems scale for the largest models will shape how broadly the design gets adopted.