TMT Breakout

TMT Breakout

TMTB Podcast Intel: AWS (AMZN) CEO Matt Garman on Rationing GPUs, Bernstein’s Rasgon, Skydance’s (SKYD) Ellison and Kreiz on Synergies and Delevering, OpenAI’s Sottiaux, Cloudflare (NET) CEO on Agents

Ethmos's avatar
TMT Breakout's avatar
Ethmos and TMT Breakout
Oct 08, 2026
∙ Paid
TMTB Podcast Intel — Powered by Ethmos

TMTB Podcast Intel is produced by Ethmos, research agents that listen to every episode in TMTB’s coverage universe and write the briefs below.

Try Ethmos — 14 days free, then 25% off your first month for TMTB readers.


Today’s Essential Episodes:

  1. AWS CEO Matt Garman on Rationing Scarce GPUs for Startups, Trainium as an Inference Chip and Amazon’s $220B Capex
    The a16z Show, with a16z’s Raghu Raghuram and Matt Garman, CEO of Amazon Web Services

  2. Bernstein’s Stacy Rasgon on Why the Semis Cycle Hasn’t Peaked, Power and Clean Rooms as the Real Constraints and NVIDIA as a Laggard
    Monetary Matters with Jack Farley, with Stacy Rasgon, senior analyst covering U.S. semiconductors and semicap at Bernstein

  3. Skydance CEOs David Ellison and Ynon Kreiz on $6B of Synergies, the Path to 3x Leverage and Streaming Unification
    CNBC’s Squawk on the Street, with David Faber, David Ellison (chairman and CEO of Skydance) and Ynon Kreiz (co-CEO of Skydance) (the Skydance segment)

  4. OpenAI’s Tibo Sottiaux on Holding Back Astra 6.1 Over Safety, Early Recursive Self-Improvement and the Dots Agents
    Pioneers of AI, with Rana el Kaliouby and Thibault “Tibo” Sottiaux, head of Product & Platform at OpenAI

  5. Cloudflare CEO Matthew Prince on Agent Traffic Heading to 1,000x Human, a Micropayment Rail Beyond Visa’s Scale and the Post-Ad Internet
    CoinDesk Podcast Network, with CoinDesk’s Jennifer Sanasie and Matthew Prince, co-founder and CEO of Cloudflare


Produced by Ethmos: AI-assisted summaries of public podcast episodes. For information only. May miss context or contain errors. Verify before relying on them. Summaries are not substitutes for the episodes, which belong to their creators. Not investment advice, offer, or solicitation. Not affiliated with or endorsed by the shows, hosts, guests or companies discussed. Corrections and publisher opt-outs: support@ethmos.ai.
AWS CEO Matt Garman on Rationing Scarce GPUs for Startups, Trainium as an Inference Chip and Amazon’s $220B Capex, on The a16z Show

Listen · Watch

THE SIGNAL

  • Garman put AWS revenue at about $169-170B, growing roughly 37%, with on-prem migration still a tailwind alongside AI; holding that growth rate at this scale backs up the $220B capex number.

  • AWS estimates 30-40% of its revenue comes from companies that started on AWS as startups, which explains why it keeps GPU capacity aside for startups instead of selling everything to frontier labs.

  • Garman says AWS used to build ahead of demand, but GPU elasticity is largely gone while core compute and storage still have headroom; AI capacity is effectively rationed and fully used, supporting near-term revenue visibility (our read).

  • Garman admits agent deploys from users with no cloud account often go to partners with simpler layers on top of AWS; the no-card, 30-second signup aims to win that entry back, flagging middleman risk (our read).

01 The capex and allocation calculus

Garman confirms Amazon’s 2026 capital expenditure is roughly $220 billion, which he calls possibly the largest single-year expense of any company, driven by AI demand he says is massive and shows no signs of slowing. AWS could sell all its GPU capacity to frontier labs alone, but Garman says it intentionally withholds capacity to support the broader ecosystem — large frontier labs (Anthropic, OpenAI, Meta), large enterprises (Salesforce, JPMC), and startups. Despite this intent, AWS still fulfills only around 60% of startup GPU requests, sometimes with delays, different regions, or different configurations. The company recently committed to buying 2 million NVIDIA GPUs over the next few years, with no certainty that will be sufficient.

On bubble concerns, Garman contends AWS’s revenue concentration per customer is in the single digits, unlike some neoclouds at 30-60%. He says the bulk of AWS usage is in core compute, storage, and inference tied to production workloads where enterprises report positive ROI today, implying that spend won’t simply evaporate. He draws an explicit analogy to the dot-com era: many individual bets will fail, but durable businesses (he cites Google and Amazon as historical analogs) survive and justify the infrastructure buildout.

02 Shifting bottlenecks and the invisible benefits of data centers

Garman frames infrastructure constraints as a moving target, citing the business book The Goal and its idea that there’s never one bottleneck — power, memory, TSMC capacity, HBM, networking components, or construction labor each become binding at different times and in different geographies (e.g., abundant power in Indonesia but not Germany). AWS now plans power and memory needs years in advance (out to 2028), including direct investment in renewable and nuclear power projects, both behind-the-meter and via the grid — planning horizons he describes as entirely new for the company at this scale.

On public perception, Garman acknowledges the industry has done a poor job explaining data centers’ community benefits, citing an example where a county’s residents pay $5,000 less in taxes annually because of AWS’s presence, a fact largely invisible to residents. He attributes negative sentiment partly to a subset of “not great” operators who ignore regulations, and says AWS needs to be more vocal about its environmental and community practices, including water-positive cooling and renewable power purchasing.

03 Custom silicon: Graviton and Trainium

Garman traces AWS’s chip strategy back roughly 13-14 years to network virtualization offload cards addressing customer demand for “bare metal performance,” which led to the Annapurna Labs acquisition and eventually Graviton, AWS’s ARM-based server chip. Graviton now ships in greater volume than any other chip type, offers roughly 20% lower cost and 20% better performance, and is used by over 90% of AWS’s top 100 customers — Garman calls it customers’ easiest way to cut their bill, in some cases halving server counts.

Trainium, now in its third generation (with a fourth announced), was launched about five-six years ago anticipating AI compute growth. Despite its training-oriented name, Garman says it has proven to be a strong inference chip on cost and performance, driving most Bedrock inference traffic, with deals with both Anthropic and OpenAI plus roughly six to a dozen smaller startups building on it. Capacity is sold out roughly through the end of next year, he says.

04 Designing infrastructure for agents

Garman describes several infrastructure shifts driven by agentic workloads: agents care disproportionately about tail latency (not just average throughput), prompting AWS to optimize underlying services like S3; a new preview service called AWS Context lets agents find data across Aurora, S3, and other stores more easily than human-oriented access patterns allow. AWS has also simplified account creation — new accounts can now be created in under 30 seconds with just a Gmail login, with no credit card or manual VPC/IAM setup required, while still allowing full scaling later without migration.

New agent-specific building blocks are emerging, Garman says: databases that agents can create quickly and throw away for transient tasks without giving up durability if they need to persist, fine-grained and time-boxed agent permissions distinct from human/service-role permissions, and lightweight compute sandboxes built on AWS’s Firecracker micro-VM technology (invented roughly a decade ago, now widely used by sandbox startups) because of its fast spin-up and strong security isolation.

05 Enterprise adoption, trust, and security

Garman says most enterprise agents today are simple, non-autonomous, and still have humans in the loop; the biggest opportunity is pushing them toward safe autonomy. Two barriers dominate: enterprises tend to replicate existing human workflows step-by-step rather than redesigning them for parallelized, agent-native approaches, and they lack trust that autonomous agents won’t make catastrophic mistakes (e.g., deleting production databases), lacking mature eval systems, data labeling, and drift-testing practices. AWS’s professional-services teams run roughly 45-day engagements designed to train customers to own these capabilities themselves rather than depend on long-term consulting.

On data strategy, Garman describes Bedrock’s core guarantee, that customer data stays inside the customer’s own trusted environment and model providers never see their prompts, as a key reason enterprises are moving proofs of concept into production there, citing growing OpenAI and Anthropic workloads inside Bedrock. He also flags growing enterprise interest in open-weights models fine-tuned or distilled on proprietary data, noting most such work currently happens in SageMaker, which he describes as finding renewed relevance as an open-model customization platform. On security, Garman says AWS recently launched Continuum, a service using powerful models to scan customer environments for vulnerabilities and prioritize them using environmental context, framing AI as both an attack-surface risk and a security opportunity requiring machine-speed defense.

06 Agents inside AWS itself

Garman describes AWS’s internal tool Quick, rolled out to every Amazon employee, enabling HR, finance, and other non-engineering teams to build their own agents — for example, automating team-planning tasks that used to take teams weeks down to hours, or pulling tax compliance rules automatically. The clearest measured gains, he says, are in software and product development speed, where frontier teams now practice agent-first development: agents write the code while engineers manage fleets of agents, producing a measurable acceleration in how fast AWS ships new features. He says AWS is actively experimenting with organizational structure — potentially shrinking teams from ten people to three or four — but has not yet settled on how organizations should permanently restructure around agent-driven work.


Bernstein’s Stacy Rasgon on Why the Semis Cycle Hasn’t Peaked, Power and Clean Rooms as the Real Constraints and NVIDIA as a Laggard, on Monetary Matters

Listen

THE SIGNAL

  • Nvidia’s 70% guide is capped by visible land, power and shell builds, not demand (enough for 100%); Rasgon models EPS of $15-16 at 70% and ~$20 at 100%, versus a ~$240 stock.

  • Clean rooms gate tool shipments: Micron’s facilities spend will outgrow its equipment spend next fiscal year, and Rasgon thinks WFE could top $300B next year if unconstrained, versus his ~$204B model, so tool upside lags (our read).

  • Lam’s NAND risk is smaller than its reputation: it still takes ~25-30% of NAND tool spend, yet when post-COVID NAND WFE fell from $20B+ to ~$5B, it guided ~$20 annualized pre-split EPS with almost no NAND.

  • Beyond Nvidia and Broadcom, Rasgon puts next-year AI silicon at ~$15B+ for MediaTek, ~$10B for Marvell and ~$5B for Qualcomm; separately, Intel gained ~200bp of gross margin selling server parts it had written down to zero.

01 Why multiples fell as earnings doubled

Rasgon frames the SOX index’s ~80% YTD gain against earnings growth exceeding 100% as classic late-cycle investor behavior: as a sector approaches a perceived peak, beats get rewarded less and multiples compress even as forward estimates keep rising. He notes forward estimates have actually risen since the June peak even as stocks fell further, meaning multiple compression has worsened. His own view: fundamentals look better, not worse, and he doesn’t see a peak this year or next, though he says he doesn’t know beyond that.

02 Roadshow findings and the demand debate

After meeting 11 companies over two days in his annual Silicon Valley roadshow, Rasgon says the universal theme is demand so strong that order visibility is exceptionally high — though he cautions semiconductor company forecasting generally is an unsolved problem, since companies only see orders in front of them, not true end demand. He references Jensen Huang’s prior forecast of $3-4 trillion in annual AI infrastructure spending by decade’s end as having once sounded implausible but now plausible given top-five hyperscalers approaching $1.5 trillion in spend next year. On returns, he cites neoclouds renting capacity at roughly $30 billion per gigawatt (implying a one-to-1.5-year payback), foundational model companies’ vertical revenue growth, and Meta’s consumer AI agent Muse as early evidence that real-world monetization is emerging, concluding every AI bear case ultimately reduces to whether there’s a return — and he believes there is.

03 Double ordering, stockpiling, and power constraints

Rasgon explains the mechanics of double ordering (customers over-order when lead times stretch, inflating apparent demand) as a classic driver of semiconductor cycles, but argues this cycle looks different: there’s no sign of GPUs stockpiled in warehouses the way COVID-era supply chains saw hoarding, since built AI capacity gets used immediately. He dismisses narratives of idle chips awaiting undelivered data centers as largely false, and downplays NIMBY data-center backlash as a political talking point unlikely to stop build-outs given available land and the option to build overseas or, per Elon Musk’s ambitions, in orbit. He identifies power — not land or capital — as the binding constraint on reaching Huang’s $3-4 trillion figure, forcing “behind the meter” on-site generation.

04 Nvidia and Broadcom as laggards

Rasgon details how investors rotated through sequential bottleneck trades in 2026 — accelerators, memory, semicap, networking, optical, power semis, and most recently CPUs — each rallying as capacity constraints hit. Oddly, Nvidia and Broadcom underperformed despite comparable earnings torque, partly because investors used them as stable funding sources (selling safe winners to buy higher-torque bottleneck names) and partly due to technical portfolio limits given Nvidia’s roughly $5 trillion market cap (8-9% of the S&P 500). He frames guided growth numbers (Nvidia 70%+ next year, potentially 100% with more supply; Broadcom 100% growth for two straight years reaching $30+ EPS by 2028) as implying roughly 11-12x earnings if those numbers are realized (by his math, Nvidia is in the mid-teens on 70% growth, $15-16 of EPS against a ~$240 stock, and 11-12x only if revenue doubles to about $20 of EPS) — cheap if this isn’t a peak.

05 Memory mechanics: DRAM, NAND, and HBM

Rasgon walks through why memory is especially tight: the current cycle follows the worst memory downturn since the tech bubble, leaving producers under-invested, while high bandwidth memory (HBM) requires 3-4x the wafers of standard DRAM due to stacking yield losses, sharply cutting effective capacity. He notes memory and storage comprise roughly 80% of silicon wafer area inside an Nvidia NVL72 rack. NAND faces separate dynamics — more competitors (Samsung, Hynix, Micron, Solidigm, Kioxia, YMTC) and new demand from AI-related KV cache storage — with Micron saying supply will get even tighter in coming years than it is now. Material new memory capacity isn’t expected until 2027-2028.

06 Semicap equipment and clean room bottleneck

Keep reading with a 7-day free trial

Subscribe to TMT Breakout to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 TMT Breakout · Publisher Terms
Substack · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture