Uncapped #57 | Andrew Feldman from Cerebras
Written by an Ethmos research agent · shared with you
Cerebras built a radically faster wafer-scale chip by surviving an 18-month near-death crisis, betting inference speed will define AI's next compute bottleneck.
- Founding thesis: Feldman argues attacking an incumbent like NVIDIA requires being 100-500x better, not incrementally cheaper, since Goliath can always cut price or bundle.
- 18-month near-death stretch: Cerebras burned $8M/month unable to make its Wafer Scale Engine work, reporting no progress at every board meeting until a July 2019 breakthrough.
- Customer ladder: Cerebras progressed from US government to sovereign cloud (G42), then a frontier lab (OpenAI), then a hyperscaler, each time answering a new skeptic objection.
- Supply chain reality: TSMC and ASML (sole EUV lithography maker) are described as near-monopolies; fabs cost $40-50B and take five years, causing chronic undersupply versus exponential AI demand.
- Inference speed bet: Feldman says speed itself creates new markets (citing dial-up vs broadband/Netflix), and Cerebras is pursuing 'disaggregation' partnerships with AMD and AWS which Cerebras reports can yield up to 5x throughput gains.
- NVIDIA's edge: Feldman attributes NVIDIA's dominance not to CUDA or chip design but to a decade of grit built while its stock languished (2003-2013) as a struggling public company.
Deep dive
The founding bet against NVIDIA
Feldman describes starting Cerebras in 2016 around a computer-architect's insight: AI was an unusually compute-intensive workload (unlike, say, ARM chips for phones), and the dominant architecture at the time wasn't built for it. He and Vishria frame the "attack Goliath" math explicitly: if an incumbent like NVIDIA improves roughly 2x per year and it takes a challenger five years to reach scale, the incumbent has compounded roughly 32x in that time, so a true challenger needs a further multiple on top — pushing the bar to 100x, 500x, or 1,000x better, achievable only through radical architectural change, not incremental tuning. Because the biggest competitor buys silicon, manufacturing capacity, and EDA tools more cheaply, Cerebras chose to control every layer — chip, board, system, software, up to the API — accepting higher cost and more failure points in exchange for full-stack innovation control, which is how it says it became world-leading in wafer-scale packaging. Vishria adds that venture investing accumulates "scar tissue" over time, and that he lacked full appreciation going in of how much science and engineering the company would need to invent from scratch.
Near-death and the eighteen-month grind
Feldman recounts a period with no evidence wafer-scale computing was even possible, spending about $8 million monthly with board meetings every six weeks where there was essentially nothing new to report. He describes wafers shattering in seconds, then minutes, then one surviving an hour, before a breakthrough day in July 2019 in a converted lab office where temperature finally held steady — a moment he calls one of the best of his life, having solved a problem no one in seventy-five years of computing had cracked. Vishria, as a board member, says he deliberately stayed out of technical decisions, calling it his job to finance and support rather than second-guess engineering, and that good boards ask about hiring and strategic tradeoffs (specialization vs. flexibility) rather than dictate technique.
Sand to chip: the fragile supply chain
Feldman walks through fabrication mechanics: ASML is the sole maker of EUV lithography machines (each roughly 50-60 feet, costing about $500 million), and TSMC uses decades of accumulated expertise to turn silicon ingots into chips inside fabs costing $40-50 billion and taking years to build. He describes this as a genuine monopoly, unlike De Beers' deliberate control of diamond supply, arising instead from technical skill nobody else has replicated. He argues US policy over three decades pushed fabs and their supporting vendor ecosystem (packaging firms like Amkor) offshore, and proposes waiving local ordinances for a twenty-year window to let TSMC, Samsung, and GlobalFoundries build US capacity, warning that losing chip capacity would be catastrophic for the economy, drawing an analogy to the Suez Canal blockage disrupting goods.
From training to inference
Cerebras started as a training-focused system when TensorFlow and ResNet were dominant, pre-transformer; Feldman describes decision-making in hardware as "hop to hop" rather than a fixed long-range plan, given the unknowable path forward. A pivotal early architectural choice — building for general algebraic operations rather than optimizing narrowly for convolutional networks — meant Cerebras was, by luck or foresight, fast at transformers when they emerged, despite the architecture predating that model type. The company only pivoted decisively toward inference around 2022-2024 once it saw AI's trajectory suggesting near-universal usage, concluding inference compute demand would overwhelm available infrastructure. Separately, Feldman discusses how Cerebras blends young, promoted-from-within product and go-to-market talent with seasoned chip engineers, arguing experience matters far more in silicon design than in product decisions where customer insight is untested territory for everyone; he also credits durable, decades-long relationships with TSMC and contract manufacturers, built on consistent follow-through, as critical to weathering supply crunches.
Customers, data centers, and the AI build-out
Cerebras's customer path moved from the US government to sovereign cloud (naming G42), then to a frontier lab (OpenAI), then a hyperscaler — each stage answering a new skeptic's objection, per Feldman. On the broader AI build-out, Feldman says Sam Altman was essentially the only person who correctly foresaw the scale of spending (recalling the eye-popping spending figures Altman floated that drew laughter at the time), while Cerebras, TSMC, NVIDIA, and memory makers all underestimated it. He describes data centers as bottlenecked by an aging US grid, long lead times for generators and electrical transmission switches, and criticizes the industry for poor community communication, asserting data centers use four-to-seven times less water than California almond farming and use closed-loop cooling.
Speed, disaggregation, and NVIDIA's moat
Feldman frames speed as historically market-creating, comparing slow AI to dial-up internet and noting Netflix transformed into a movie studio once broadband enabled streaming, tying this to a frontier model's recent limited-availability launch. Cerebras is pursuing disaggregation — splitting inference workloads across chip types — with AMD and AWS, citing reported gains of up to 5x in throughput while maintaining speed, ahead of production availability, and says it's open to similar partnerships across the major chip makers, naming NVIDIA, AMD, Google's TPU, and AWS. On NVIDIA specifically, Feldman attributes its success to relentless grit accumulated during a decade (roughly 2003-2013) of poor stock performance as a public company, more than to CUDA or architecture, and describes Cerebras as still operating with an underdog mentality despite its growing customer list. He closes by recounting his upbringing on the Stanford campus among faculty neighbors including William Shockley and Amos Tversky, crediting a culture that valued intellectual horsepower above wealth or status.
Briefs like this, for everything you follow.
Ethmos puts research agents on your coverage universe — every debate that touches your names, every executive appearance you'd have missed — and files briefs like this one every morning.
Start freeThis page is an AI-generated summary and analysis prepared with Ethmos and shared by an Ethmos user. It does not reproduce the original programming. The underlying episode and its recording remain the property of their respective creators, and all show and company names are the property of their respective owners.
AI summary. May contain errors. Not investment advice.