EP 192: Talking Apple Silicon with Apple's Tom Boger, Kaiann Drance, and Sri Santhanam
Written by an Ethmos research agent · shared with you
Apple's 2nm chips pair new MCM packaging and dual neural engines to push sustained on-device AI performance while keeping security compute costs hidden but real.
- 2nm transition: A20 Pro and M6 move to 2nm; on A20 Pro that delivers a 20% faster CPU and 40% faster GPU versus its predecessor, plus a dual Neural Engine (new to iPhone) and extra GPU/memory-channel bandwidth.
- Dual Apple Neural Engine comes to iPhone: A20 Pro packs two discrete 16-core ANEs (already in M6) that can run parallel workloads or split one workload, powering Siri AI and DeepFusion low-light photography.
- New packaging cuts heat: Placing memory beside the processor die (an MCM approach inspired by M1) frees the vapor chamber to boost sustained performance, aided by nano-twin copper and graphite materials.
- Unified memory strategy for edge AI: Apple frames on-device AI plus Private Cloud Compute as a scalable hybrid architecture to cut token costs and preserve privacy across iPhone, Mac Mini, Mac Studio, and clustered M5 Ultra systems.
- Security has a silicon cost: Secure Enclave and the newer Secure Exclave (isolated hardware compartments for biometric and sensor data) require dedicated engineering effort, including memory-tagging techniques introduced to avoid CPU performance hits.
- Vertical integration as moat: Apple designs silicon and products 'in concert,' letting it commit chip real estate years ahead for future camera/AI needs rather than picking from a merchant-silicon menu.
Deep dive
The 2nm jump and its yields
Sri Santhanam (Apple's Silicon Engineering Group head) says the SoC vertical stack — architecture, design, implementation, packaging — matters as much as the process node, since Apple has shipped meaningful gains in years without a node bump. This year's 2nm transition gave Apple headroom to double the neural engine, add a GPU core, and add memory channels, yielding a 20% faster CPU and 40% faster GPU on A20 Pro versus its predecessor. Kaiann Drance (iPhone/Watch marketing) notes the architecture is scalable across the whole line, from the Watch's S11 chip up through M-series, sharing packaging and design learnings.
Thermals as a hidden lever
Executives repeatedly frame thermal design as the real bottleneck behind sustained (not just peak) performance for AI and gaming workloads. Tom Boger (Mac/iPad marketing) cites the Mac Mini's power efficiency as newly critical for agentic AI workloads that run continuously, and describes users clustering multiple Mac Studios to run frontier models on a standard wall outlet. On iPhone, Drance says the team must re-optimize thermal design each year within a tighter thermal envelope than Mac, balancing frame-rate stability and hand comfort against AI and gaming loads. Santhanam adds that A20 Pro's new packaging — placing memory beside the processor die — was designed explicitly to extract heat more efficiently via the vapor chamber, unlocking both burst and sustained performance.
Packaging innovation and its ripple effects
A20 Pro brings an MCM packaging approach to iPhone. Santhanam says placing memory beside the die increases the chip's XY footprint but improves bandwidth and thermal contact with the vapor chamber, which Drance says is now physically larger and uses new nano-twin copper and graphite materials, contributing to the 40% sustained-performance gain. Drance describes this as a "domino effect" requiring redesign of adjacent components (including fitting a larger battery) — work she says a consumer never sees, only feels as better sustained performance and camera capability. Boger notes the same packaging philosophy scales up to M5 Ultra's quad-die architecture, which he calls unmatched in the industry for interdie bandwidth.
Vertical integration and silicon-product co-design
A recurring theme, echoed by all three guests, is that Apple designs silicon and end products "in concert" rather than picking chips from a vendor menu — Drance references Apple's stated position as "not a merchant silicon vendor." This lets Apple commit chip real estate (e.g., extra data lanes) years ahead of features it knows are coming in camera or AI, while building in headroom so older iPhones keep running new software well, supporting device longevity.
Dual Neural Engine and on-device AI architecture
A20 Pro's two discrete 16-core Apple Neural Engines can run separate workloads in parallel or split one workload, with Apple's Core AI framework — not app developers — deciding placement for performance or efficiency. Santhanam cites Siri AI, camera processing, and the DeepFusion low-light photography network as dual-ANE workloads; Drance adds that this underpins new variable-aperture and improved low-light/video capture. Boger frames the ANE as one piece of a "balanced approach" alongside neural accelerators embedded in the GPU and CPU, all fed by high memory bandwidth, calling Apple's CPU cores the fastest at AI orchestration. Drance also cites a Netflix app leveraging dual displays on the iPhone Duo as an early example of developers exploiting the new performance and thermal headroom, made possible without developers making explicit ANE placement decisions.
Edge AI, unified memory and security architecture
Executives connect unified memory and Private Cloud Compute into a single "scalable" strategy: run as much AI as possible on-device for cost and privacy reasons, and offload only more complex models to Apple's private cloud while preserving on-device-level privacy protections. Boger says customer demand is driven by rising token costs and a desire to keep data local, plus a shift toward specialized models suited to different tasks. On security, Drance and Santhanam describe hardware investments — the long-standing Secure Enclave (originally for Touch ID/Face ID biometrics) and a newer Secure Exclave isolating sensor data such as audio — as compartments the application processor cannot access, plus a memory-tagging technique introduced last year that required significant engineering to avoid hurting CPU performance. Boger frames this as part of a holistic security stack spanning FileVault, locked-down boot ROMs, and the removal of kernel extensions.
Briefs like this, for everything you follow.
Ethmos puts research agents on your coverage universe — every debate that touches your names, every executive appearance you'd have missed — and files briefs like this one every morning.
Start freeThis page is an AI-generated summary and analysis prepared with Ethmos and shared by an Ethmos user. It does not reproduce the original programming. The underlying episode and its recording remain the property of their respective creators, and all show and company names are the property of their respective owners.
AI summary. May contain errors. Not investment advice.