Google's AI Infrastructure Chief, Amin Vahdat, on the Physics & Economics of Frontier AI
Written by an Ethmos research agent · shared with you
Vahdat says Google measures AI infrastructure by "goodput" delivered through constant failures, and names power—not chips—as the binding long-term constraint.
- Goodput over FLOPS: Vahdat says Google judges AI infrastructure by delivered goodput—actual work completed through failures and recovery—not theoretical FLOPS or benchmark throughput.
- Constant failures at scale: At 100,000-accelerator scale something fails multiple times an hour, with root causes spanning hardware, network, and software bugs and no single dominant cause.
- 8i/8t split: Google's 2026 TPUs split into 8i (inference) and 8t (training), a call made about two years ago when Google projected serving could be 30% to 60% of the market over the chips' lifetime; either chip can still do the other's job.
- TPU primitives are durable: Despite TPUs evolving since 2013, the core architecture—large matmul units, sparse core, remote load/store via ICI network—has stayed stable since TPU v1.
- Old chips still working: Vahdat notes Google's seven- and eight-year-old TPUs run at 100% utilization, with whole pods (not individual chips) replaced once fully depreciated.
- Power is the binding constraint: Google prefers staying grid-connected over vertical integration, co-funding utility transmission upgrades, and is pursuing orbital data centers (a 'moonshot') to access 1.4x more solar power and near-continuous sunlight.
Deep dive
Measuring infrastructure: goodput, not FLOPS
Vahdat argues chip-centric metrics like FLOPS represent theoretical maximums rarely achieved in practice; what matters is the performance a workload actually delivers. He introduces goodput as Google's internal (and increasingly industry-adopted) metric — throughput minus the cost of failures, recovery, and checkpoint restoration. At the scale Google operates (100,000+ accelerators), failures occur frequently, described as happening multiple times a day and in some configurations multiple times per hour, with causes spanning hardware, networking, and software (compiler bugs, runtime issues, model problems) — an ongoing, unpredictable discovery process with no single dominant root cause. Asked whether Google must double serving capacity roughly every six months, Vahdat says the bar is token-generation capability, with as much or more of each doubling coming from software as from hardware — though hardware still delivers a reliable 2x-or-more year-over-year performance multiplier that benefits every layer built on top of it.
The TPU program and the 8i/8t split
Google's TPU program began in 2013 as a contrarian bet against the prevailing wisdom (Moore's Law plus general-purpose flexibility) that specialized chips weren't worth building. It expanded from inference-only (translation, voice) to training, then absorbed the transformer era and recommender systems. A pivotal recent decision: whether to keep one general-purpose chip or split into two specialized ones. Seeing inference and serving take off, Google projected serving would be a large share of the market over the chip's lifetime, justifying the split into 8i (inference) and 8t (training). Critically, both chips retain cross-workload capability, avoiding the risk of over-committing to a wrong demand forecast. Vahdat frames specialization as an "art": a workload must be durable enough to justify the fixed cost and narrow intercept window of chip design. Despite this evolution, the core TPU architecture has barely changed since TPU v1 — matrix multiply units, a sparse core, and remote load/store via the ICI network remain the stable primitives beneath generations of model change, analogous to a CPU instruction set.
Co-design with DeepMind
Vahdat describes deep, physical co-location between his infrastructure team and DeepMind researchers, who work in overlapping spaces and can intercept chip architectures mid-flight before tape-out — something he says would be harder to coordinate across company boundaries. Google runs five or six simultaneous chip-development stages (production, debugging, pre-tape-out, design, concept), spanning years, while model research moves faster; DeepMind researchers have learned which optimizations are worth escalating (tiny gains typically aren't) versus which justify delaying a tape-out. He declines to speculate on whether OpenAI's reportedly homogeneous compute stack versus Anthropic's more heterogeneous one explains their differing model architectures, but affirms Google increasingly uses Gemini models to help design hardware for future Gemini generations.
Co-design tradeoffs and open standards
The case for co-design across chip, network, and software layers is large multiplicative performance gains (10-20%, even 2x per layer) at the cost of flexibility and vendor lock-in; full fungibility leaves significant performance unrealized. Despite favoring co-design internally, Vahdat stresses Google supports open standards to avoid forcing a walled garden — supporting both its own JAX framework and PyTorch via Torch TPU, invoking the historical example of IP's win over competing protocols in the 1970s-80s internet.
Workload shifts: long-horizon agents and networking
The rise of long-horizon agents has shifted interaction cadence from human-paced (seconds) to machine-paced (milliseconds), driving explosive demand not just for accelerators but for CPUs, storage, and networking, since agents must orchestrate context retrieval across DRAM, SSD, and HDD between each step. This creates building-design tradeoffs: co-locating CPU and accelerator racks sacrifices density/networking specialization, while separating them into different buildings adds networking complexity and latency (hundreds of microseconds). On networking, Google pioneered wave-division multiplexing and optical circuit switching (using MEMS mirror switches) roughly fifteen years ago, enabling millisecond-speed rerouting of light to spare racks on failure and flexible network topology changes without physically moving fiber.
Power as the binding constraint
Asked about the single biggest constraint, Vahdat names power as uniquely unresolved over the long term, unlike other challenges he considers solvable given enough time. Google prefers staying grid-connected rather than vertically integrating (e.g., building own turbines), citing statistical multiplexing benefits — building redundant power alone to hit 99.99%+ reliability would require roughly doubling capacity. Google pays for utility transmission upgrades to avoid raising other ratepayers' costs, plans gigawatt-scale projects years in advance, and sometimes supplements gaps (in his hypothetical, 700MW arriving on time against a needed gigawatt) with local generation like solar or batteries, occasionally feeding power back to the grid during peak residential demand.
Data center sizing, lifecycle, and AI-assisted engineering
Deciding data center size is described as an art balancing single-point-of-failure risk against training workloads' preference for maximal co-location; sites range from tens of megawatts at network edges to near-gigawatt training clusters. Serving/inference clusters differ from training clusters — they need more co-located storage and compute diversity, are distributed globally for locality, and are less vertically integrated. On hardware lifecycle, Google replaces entire pods (e.g., a TPU 8T pod of 9,600 chips across more than 140 racks) once they are fully depreciated, though Vahdat notes seven- and eight-year-old TPUs still run at full utilization. AI now matches software engineers' usage levels among Google's hardware engineers, shrinking time from design kickoff to tape-out, and is streamlining data-center planning decisions previously done via detailed spreadsheets — though not yet via reasoning models specifically.
Orbital data centers and the 2036 outlook
Google is seriously pursuing orbital data centers as a moonshot, motivated by power constraints: space offers roughly 1.4x more usable solar power (no atmospheric attenuation) and 90-100% sunlight coverage in sun-synchronous orbit versus a far smaller share on land, largely eliminating the need for batteries. Challenges include harder cooling, repair, and reliability in space, and the adoption of free-space optics instead of fiber. Looking to 2036, Vahdat expects far greater integration — racks potentially holding multiple megawatts, manufactured centrally, with only water, power, and a small fiber bundle plugged in — and speculates, tentatively, that such racks could even be launched into orbit and robotically assembled.
Briefs like this, for everything you follow.
Ethmos puts research agents on your coverage universe — every debate that touches your names, every executive appearance you'd have missed — and files briefs like this one every morning.
Start freeThis page is an AI-generated summary and analysis prepared with Ethmos and shared by an Ethmos user. It does not reproduce the original programming. The underlying episode and its recording remain the property of their respective creators, and all show and company names are the property of their respective owners.
AI summary. May contain errors. Not investment advice.