Cerebras' 5nm wafer-scale chips skirt HBM and CoWoS supply bottlenecks, analysis says
In brief
- Cerebras' WSE-3 and WSE-3T use TSMC's 5nm node and keep all memory on-chip as SRAM.
- Crypto Briefing says the design skips HBM, CoWoS packaging and 3nm, three tight AI hardware chokepoints.
- Cerebras reported a $25.4 billion backlog, with more than $20 billion tied to OpenAI.
- Risks flagged by the analysis: data-center buildout, customer concentration and a fixed 44 GB memory budget.
How the design avoids the queue
The trick is memory. All memory on the WSE-3 and WSE-3T sits on the chip itself as SRAM, rather than on separate HBM stacks, so Cerebras isn't competing for the same scarce parts as GPU makers.
The WSE-3 was announced in March 2024. It packs 900,000 AI cores, 44 GB of on-chip SRAM and approximately 4 trillion transistors into a 46,225 mm² area.
CEO Andrew Feldman has pitched the architecture hard. Speaking in June 2026, he said it delivers:
"the fastest inference in the world by an order of magnitude"
That's Feldman's own claim (it isn't an independently verified benchmark).
Backlog and capacity
Cerebras reported a backlog of $25.4 billion as of June 2026, according to Crypto Briefing. More than $20 billion of that comes from a multi-year deal with OpenAI. The company also says it has over 600 MW of data-center capacity either live or contracted.
For background, Cerebras completed a Nasdaq IPO in May 2026 and has since started expanding US manufacturing and building out partnerships to add operational capacity. It previously agreed to supply CS-4 systems for approximately 100 MW of capacity to Gimlet Labs, and it has a long-term agreement with General Compute, with deployment planned for Q1 2027.
The risks Crypto Briefing flags
Almada Lopez doesn't treat the supply advantage as a free pass.
The backlog isn't revenue, he wrote, and converting it depends on data centers being built, powered and run on schedule. Concentration is the second issue: the Gimlet Labs and General Compute deals help diversify the customer mix, but they're much smaller than the OpenAI contract.
The third is the chip itself. On-chip SRAM is extremely fast, yet 44 GB per chip is a fixed budget, and the analysis notes that how Cerebras systems scale for the largest models will shape adoption.
Frequently asked questions
How do Cerebras chips avoid the HBM shortage?
All memory on Cerebras' WSE-3 and WSE-3T sits on the chip itself as SRAM rather than on separate HBM stacks. According to Crypto Briefing, the chips also avoid advanced CoWoS packaging and 3nm manufacturing, since they're built on TSMC's 5nm process node.
What risks did Crypto Briefing flag for Cerebras?
Crypto Briefing said the $25.4 billion backlog isn't revenue, and converting it depends on data centers being built, powered and run on schedule. It also pointed to reliance on the OpenAI deal and to the fixed 44 GB of SRAM per chip, which it said makes scaling for the largest models an adoption question.


