Long-form 23 items
Thu, 24 Sept 2026
Long-form 05:25 ET
Alpha Exchange — Amanda Lynam: GS credit on AI capex financing
GS Chief Credit Strategist: hyperscaler IG $250bn/'26→$400bn/'27; $6tn capex '26–'30; hyperscalers 40% of AI issuance; little crowding-out; IG absorb.
asr Alpha Exchange · Goldman Sachs · Meta · Bloomberg
Long-form 05:20 ET
SemiAnalysis — ClusterMAX 3.0: Nebius platinum, rankings, financing
Ep.033: 77 providers ranked; Nebius joins CoreWeave platinum; Google gold; Azure/AWS down; GPU-hour backwardation; NVDA backstop ~$588bn→$2tn; SLAs.
asr SemiAnalysis Weekly · Nebius · CoreWeave · Oracle
Wed, 23 Sept 2026
Long-form 05:15 ET
Latent Space — Diogo Almeida / Jev: System One for prod, not God
InstructGPT coauthor: frontier chat/RLHF APIs wrong for software; Jev as code-consumed System One; >1T tokens/day machine traffic; dark data + agents.
asr Latent Space · TypeSafe · Jev · OpenAI
Tue, 22 Sept 2026
Long-form 05:40 ET
ILTB — Gabe Stengel / Rogo: investing superintelligence, harness, last mile
Rogo CEO: o1 Pro→Opus 4.5 unlocked junior-analyst work; next 2–5y is firm reinvention; harness/compliance/last-mile beat raw models for buy-side.
asr Invest Like the Best · Rogo · Jane Street · Goldman Sachs
Long-form 05:30 ET
All-In — Naveen Rao: AI energy wall, 4D computing, 1000x efficiency bet
Unconventional AI CEO: Google-scale token energy already ~12GW; ~50% of token cost is power; aims 1000x efficiency in ~3.5y via dynamical chips.
asr All-In Summit · Unconventional AI · Nervana · Intel
Mon, 21 Sept 2026
Long-form 05:30 ET
Excess Returns — Jason Hsu: China AI gap, capex arms race, S&P seven
Rayliant CIO: China models on-par/open-source; energy grid edge; hardware rents until overcapacity; Mag7 CapEx arms race; S&P is one-tree, not diversifier.
asr Excess Returns · Rayliant Global Advisors · Research Affiliates · DeepSeek
Long-form 05:25 ET
Odds on Open — Lihong Wang: ex-IMC semis quant, AI stack portfolio
Ex-IMC semis options MM on flow/V, NVDA–AMD relative vol, DeepSeek corr blowups; 50-name AI stack book at ~2×; models-beat-S&P claim needs harness.
asr Odds on Open · IMC · NVIDIA · AMD
Sun, 20 Sept 2026
Long-form 05:20 ET
MiB — Glen Kacher: AI boom is catch-up, not overbuild
Light Street CIO on AI5 semis concentration, NVDA ~85% share, demand ahead of supply, 10–20y stack cycle, agents→~5× tokens; DC politics as education risk.
asr Masters in Business · Light Street Capital · NVIDIA · AMD
Sat, 19 Sept 2026
Long-form 05:20 ET
a16z — Ali Ghodsi: enterprise stall is context, not IQ
Databricks CEO on pacing PR vs cyber risk, four-test RSI bar, ontology/Genie as the adoption bind, Uni Gateway cost control + GLM shift.
asr The a16z Show · Databricks · OpenAI · Hugging Face
Long-form 05:20 ET
No Priors — Ermon/Inception: diffusion wins inference parallelism
Stefano Ermon on Mercury ≈ Haiku/Flash/mini speed tier, ~10× decode vs AR at GPT-2 scale, OpenCall leaving Cerebras for NVDA GPUs, 20–30% latency wedge.
asr No Priors · Inception · Mercury · OpenAI
Fri, 18 Sept 2026
Long-form 05:20 ET
Dwarkesh — Noam Brown: agent swarms, RSI speedup, alignment bind
OpenAI's Noam Brown on 10k-agent Navier-Stokes solve (130B tokens/88h), Ultra Mode multi-agent, Codex $7–8k/day internal, RSI ≠ 100x overnight.
transcript Dwarkesh Podcast · OpenAI · Hugging Face · Astra
Thu, 17 Sept 2026
Long-form 05:20 ET
Latent Space — AIUC: trust/liability as the agent adoption bind
Rune Kvist (ex-Anthropic) on $40M Series A: AIUC-1 quarterly agent standard, Lloyd's-backed policies, Waymo/Air Canada liability — eval+insurance stack.
transcript Latent Space · AIUC · Anthropic · Cursor
Long-form 05:15 ET
All-In — Gerstner: no AI bubble; semis = ~70% of Nasdaq return
Altimeter's Brad Gerstner on All-In: earnings-driven tape, offtake must fund Mag5 capex, Dylan 43GW too hot (~25GW), lab RR as takeoff switch. [asr]
asr All-In Podcast · NVIDIA · Anthropic · OpenAI
Wed, 16 Sept 2026
Long-form 05:20 ET
SemiAnalysis Ep.031 — pacing may eat more compute, not less
Emergency ep on Amodei pacing: OpenAI CoT monitoring ~20% of rollup compute; safety spend likely raises, not cuts, infra demand; HF as shot across bow. [asr]
asr SemiAnalysis Weekly · Anthropic · OpenAI · Hugging Face
Long-form 05:15 ET
All-In — Satya: pace with common sense; MSFT builds, leases, rents
Nadella on All-In: broad diffusion over mystical slowdown; ~30m enterprise Copilot users of ~250–300m TAM; kit ~60% of cost; Quincy DC ~400–500 MW. [asr]
asr All-In Podcast · Microsoft · OpenAI · Anthropic
Tue, 15 Sept 2026
Long-form 07:30 ET
Elon Musk & Gwynne Shotwell — AI Peer Review, Starship, Terafab, SpaceX/Tesla Merger (All-In)
SpaceX President Gwynne Shotwell says SpaceX is as much an AI business as a space business by revenue, with compute rental 'a heck of a business' and Starlink at ~1.5–2% penetration; Elon Musk joins from Memphis to push cross-lab model peer review, handicap Starship ship-catch at ~50–60%, frame Terafab as build-or-fail-to-scale, and non-deny a Tesla–SpaceX combination.
asr All-In Podcast · SpaceX · Tesla · xAI
Long-form 07:30 ET
All-In — Jensen: doomer math fails; open models carry apps
Huang on All-In: extinction %s unscientific; ~$400bn AI VC ~80% open-model; NVIDIA goes "as deep as needed." Trump brands DC opposition a hoax. [asr]
asr All-In Podcast · NVIDIA · Anthropic · OpenAI
Mon, 14 Sept 2026
Long-form 20:50 ET
Jensen Huang — Nvidia's Future, Physical AI, Rise of the Agent, Inference Explosion (All-In)
NVIDIA CEO Jensen Huang tells the All-In hosts that agentic workloads drove a ~10,000x compute step in two years, that a higher-capex Vera Rubin factory can still deliver the lowest token cost via ~10x throughput, and that Physical AI is already a near-$10bn NVIDIA line while open-weight agents redefine the desktop OS — with China licenses restarting and consensus growth paths rejected as undersized.
asr All-In Podcast · NVIDIA · Groq · Anthropic
Long-form 20:30 ET
Gavin Baker — Why AI Demand Is Outrunning Compute Supply (a16z Show)
Atreides CIO Gavin Baker tells David George that AI fundamentals accelerated through July–August while related equities drew down; argues sub-one-year compute paybacks and thin heavy-user penetration make undersupply through 2028 the base case, with NVIDIA’s financeable stack and hybrid open-source routers as the durable structure.
asr The a16z Show · NVIDIA · OpenAI · Anthropic
Long-form 17:29 ET
TBPN: The AI Slowdown Debate
Metadata-only: TBPN's Sep 14 episode (full + Diet cut) titled The AI Slowdown Debate, with guests including Nico Wittenborn, Scott Keogh, Mitchell Green, David Rosenthal, Ben Gilbert, and Faraj Aalaei.
metadata-only TBPN
Long-form 13:00 ET
SemiAnalysis Weekly: why 4-Hi HBM may win on inference economics
Metadata-only capture of SemiAnalysis Weekly Ep. 030: Myron Xie and Jordan Nanos on Rubin Ultra shipping 192GB HBM versus a 1TB preview, supply-driven decontenting, and why less memory per chip can still be the right call.
metadata-only SemiAnalysis Weekly · NVIDIA · SemiAnalysis
Long-form 12:06 ET
Eisman Playbook: Big Short partners on rates, AI, gold
Metadata-only: Steve Eisman with Vincent Daniel and Porter Collins on bonds, Treasury buybacks, OpenAI risk, gold, shorting mechanics, and two live short ideas.
metadata-only The Real Eisman Playbook · OpenAI
Long-form 12:04 ET
Latent Space: Richard Socher on recursive self-improvement
Metadata-only: Latent Space interviews Richard Socher (Recursive / You.com) on recursive self-improvement as the next major AI step; show notes flag AI×finance conference promo.
metadata-only Latent Space · You.com · Recursive
Long-form · Thu, 24 Sept 2026 · 05:20 ET

SemiAnalysis — ClusterMAX 3.0: Nebius platinum, rankings, financing

Ep.033: 77 providers ranked; Nebius joins CoreWeave platinum; Google gold; Azure/AWS down; GPU-hour backwardation; NVDA backstop ~$588bn→$2tn; SLAs.

asr Jordan NanosSam HarshePratt Bhatt NebiusCoreWeaveOracleGoogle CloudAzureAWSLambdaFermusGMI CloudTensorWaveCrusoeNvidiaOpenAI Source ↗
Venue: SemiAnalysis WeeklyHost: Jordan NanosDuration: 69mPublished: Wed, 23 Sept 2026 · 18:50 ET

Opening

ClusterMAX 3.0 — SemiAnalysis's hands-on ranking of managed GPU clusters — puts Nebius into platinum beside CoreWeave, elevates Google Cloud into gold with Oracle, and marks Azure/AWS/Crusoe moving down, while the conversation pivots from NCCL/health-check bugs to why lenders and parametric insurers need actuarial cluster quality for the AI buildout. Jordan Nanos with Sam Harshe and Pratt Bhatt (~69m) reviewing the Wed Sep 23 article: methodology after months of testing, tier moves, reliability/networking/storage war stories, agentic-coding stress on ops, and financing (SLAs, Nvidia backstop universe, old-GPU terminal value). YouTube auto-captions asr — Watch. Figures marked (asr).

Key takes

Headline ranks: Nebius joins CoreWeave in platinum; gold is Oracle + Google Cloud (Google newly up); silver has Azure moving down, Fermus/Lambda stay, GMI up, TensorWave stays; bronze includes AWS and Crusoe moving down. Scope claim: 77 providers ranked, market view expanded to 323 in the subscriber Neocloud model, >200 end-user interviews. ClusterMAX 4.0 teased after Vera Rubin lands in DCs; TPUs/Trainium deliberately deferred ("too many secrets"). [asr]

Buyers' top-of-mind is extreme near-term GPU-hour backwardation — pay almost anything for chips in weeks; big discount if willing to wait months — so even broken Kubernetes/storage still clears at good margins if you have GPUs. Providers that failed to sell early (quality signal historically) are now closing at better margins. Neoclouds expanding up-stack into managed inference / hosted training / endpoints because cluster margins are rich and labs capture cake — Pratt's frame: everything points to centralization of compute. [asr]

Reliability differentiation is health checks that actually work under failure injection — active (preemptable dummy workloads) vs passive (low/zero GPU overhead) — plus timed detect→drain→hot-spare (HGX) or reboot/repair (NVL72) workflows. Amazon HyperPod Slurm cited as health-check logic that required a healthy node to run remediation that was itself needed to restore the node — "worse than no health checks." Sample SLAs sold as industry standard for downtime definitions / credits. [asr]

Networking next after reliability: providers should ship a golden NCCL recipe with smooth message-size throughput; Google called out for baking custom networking into NGC containers so NCCL works out of box without Nvidia switch silicon. EFA/custom scale-out framed as a false economy — saves a few % on fabric but blocks DeepSeek-style collectives / modern recipes; Perplexity blog posts cited as evidence of engineer-hours spent making EFA work. Fermus GB300 example: out-of-box TorchTitan MoE job ~445–455 tok/s/GPU vs ~3500 on better-configured peers (~50% gap) from exposed all-to-all latency — fixed after joint debugging. [asr]

Financing is why ClusterMAX matters for the ~$7tn buildout narrative: lenders and neoclouds insure against late delivery / failed acceptance and against SLA violations that trigger cancellation rights; parametric insurers (e.g. Parametrics named) need actuarial inputs (redundant power/cooling/uplinks, parts depot, staff, flood/hurricane). SemiAnalysis circulating example SLAs/contracts; seed startups spending "literally 100% of seed" on neocloud capacity makes provider pick existential. Credit/markets model tease: Nvidia off-balance-sheet backstop universe >$588bn by end FY27, forecast >$2tn by end 2031; separate supply-chain-only backstop figure ~$827bn FY31 for HBM etc. (ex LPS/DC build). Amazon/Google also putting balance sheets to work. Terminal-value debate: 4y H100 contracts still signing; OpenAI never gave up A100s (2020 → 6+ years); GB300 racks ~140–200 kW vs H100 ~30 kW facilities that cannot host Vera Rubin. [asr]

Hosted training / RL infra is harder than pre-training clusters — inference + training sync + environments/sandboxes; market "not yet" for pure hosted training; Anthropic paid heavily for environments. Endpoint teaser (not in ClusterMAX proper): famous unnamed provider ~75% cache-hit vs good providers ~99% on ~800-token replay traces — 75% roughly doubled total cost vs 99%. Framework wars: SGLang/vLLM → Miles/PrimeRL etc.; trainer–inference mismatch compounding over long trajectories collapses training. Agentic coding helps debug but blows feet off on orchestration; CMax CLI open-sourced as front-end. [asr]

Key math

77 providers ranked; 323 in Neocloud market model; >200 end-user interviews (asr) — coverage vs prior ClusterMAX. [asr]

Platinum: Nebius + CoreWeave; Gold: Oracle + Google Cloud (asr) — tier headline. [asr]

Near-term GPU price backwardation: pay up for weeks-out chips; large discount for months wait (asr — qualitative) — why mediocre clusters still clear. [asr]

Fermus GB300 MoE out-of-box ~445–455 tok/s/GPU vs ~3500 on better peers (~50% gap) (asr) — networking config as realized performance. [asr]

Endpoint cache-hit ~75% vs ~99%; ~75% roughly doubles replay cost on ~800-token traces (asr — preliminary, unnamed famous provider) — inference margin trap. [asr]

Nvidia off-BS backstops >$588bn end FY27; >$2tn end 2031 forecast (asr — SemiAnalysis credit model / Dan article) — financing backdrop. [asr]

Supply-chain-only Nvidia backstop ~$827bn FY31 HBM etc., ex LPS/DC (asr — speaker hedge "may get numbers wrong") — narrower universe. [asr]

GB300 racks ~140 kW → ~200 kW peak vs H100 racks ~30 kW / facilities designed 30–40 kW (asr) — why old H100 halls can't take Vera Rubin without rebuild. [asr]

Quotes

"There's a new addition to the platinum tier. We've got Nebius up there joining CoreWeave." — Jordan Nanos [asr]

"The ridiculous backwardation of the price curve at the moment… people are willing to pay just about anything for any chips that they can get in the next few weeks." — Sam Harshe [asr]

"Everything points to centralization of compute." — Pratt Bhatt [asr]

"Nvidia's backstop universe off-balance sheet has grown… to be over $588 billion… forecasts… over 2 trillion by the end of 2031." — Jordan Nanos [asr]

"OpenAI has never given up an A100 GPU that they've gotten." — Jordan Nanos [asr]

"I'd rather not be fighting to get back your… 5% because it was down for three weeks out of the year." — Jordan Nanos [asr]

Variant perception

Priced in — Neocloud quality is heterogeneous; CoreWeave as reference premium; Google/Oracle serious GPU cloud competitors; financing/SLA risk on AI buildout; Nvidia ecosystem backstops and offtake circularity already a macro debate; HBM/power as physical constraints.

What's new — Nebius into platinum on measured ClusterMAX 3.0; Google into gold; Azure/AWS/Crusoe marked down on hands-on reliability/networking; quantified Fermus out-of-box MoE gap; endpoint cache-hit cost doubling teaser; explicit link from cluster quality → parametric insurance / lender underwriting; SemiAnalysis Nvidia backstop stock/flow figures ($588bn / $2tn / $827bn supply-chain); A100 never-retired + multi-year H100 contracts as terminal-value evidence against steep depreciation narratives; RL/hosted-training harder than pretrain-cluster product.

Bear case — Rankings are SemiAnalysis's product (sample SLAs, credit model, consulting) talking book; "77/323" and backstop dollars are model outputs not audited disclosures; tier moves may lag commercial reality (AWS/Azure still default for many enterprises); GPU backwardation could flip fast if supply catches; hosted-training "no market yet" could be wrong within a year; unnamed endpoint shame-test is preliminary.

Discount — SemiAnalysis sells ClusterMAX access, sample SLAs, site visits, and a new credit/markets model — the episode is both research and distribution for those products. Nebius/CoreWeave platinum elevates names SemiAnalysis already covers heavily. Financing segment steers listeners to cluster@semianalysis.com and the Nvidia Backstop Universe article.

Positioning

AI capex durability — STRENGTHENS. Near-term GPU-hour backwardation, multi-year H100 contracts still clearing, A100s never retired at OpenAI, and expanding Nvidia/hyperscaler backstops all read as demand still outrunning usable supply — with quality/SLA as the binding financing filter rather than a demand cliff.

HBM supply binds — NEUTRAL. Supply-chain backstop figure (~$827bn FY31 HBM etc.) and Vera Rubin / GB300 power-density migration reinforce physical-world bottlenecks, but the episode does not add a new HBM wafer/capacity datapoint beyond financing exposure.

Inference margin inversion — WEAKENS (narrow). Endpoint cache-hit miss roughly doubling cost and "80% margins on inference" as a joke undercuts the idea that list/serving math alone captures realized inference P&L when routing/cache is broken.

The Open/Close  ·  Research commentary, not investment advice. Positions may be held in securities mentioned.