Long-form 23 items
Thu, 24 Sept 2026
Long-form 05:25 ET
Alpha Exchange — Amanda Lynam: GS credit on AI capex financing
GS Chief Credit Strategist: hyperscaler IG $250bn/'26→$400bn/'27; $6tn capex '26–'30; hyperscalers 40% of AI issuance; little crowding-out; IG absorb.
asr Alpha Exchange · Goldman Sachs · Meta · Bloomberg
Long-form 05:20 ET
SemiAnalysis — ClusterMAX 3.0: Nebius platinum, rankings, financing
Ep.033: 77 providers ranked; Nebius joins CoreWeave platinum; Google gold; Azure/AWS down; GPU-hour backwardation; NVDA backstop ~$588bn→$2tn; SLAs.
asr SemiAnalysis Weekly · Nebius · CoreWeave · Oracle
Wed, 23 Sept 2026
Long-form 05:15 ET
Latent Space — Diogo Almeida / Jev: System One for prod, not God
InstructGPT coauthor: frontier chat/RLHF APIs wrong for software; Jev as code-consumed System One; >1T tokens/day machine traffic; dark data + agents.
asr Latent Space · TypeSafe · Jev · OpenAI
Tue, 22 Sept 2026
Long-form 05:40 ET
ILTB — Gabe Stengel / Rogo: investing superintelligence, harness, last mile
Rogo CEO: o1 Pro→Opus 4.5 unlocked junior-analyst work; next 2–5y is firm reinvention; harness/compliance/last-mile beat raw models for buy-side.
asr Invest Like the Best · Rogo · Jane Street · Goldman Sachs
Long-form 05:30 ET
All-In — Naveen Rao: AI energy wall, 4D computing, 1000x efficiency bet
Unconventional AI CEO: Google-scale token energy already ~12GW; ~50% of token cost is power; aims 1000x efficiency in ~3.5y via dynamical chips.
asr All-In Summit · Unconventional AI · Nervana · Intel
Mon, 21 Sept 2026
Long-form 05:30 ET
Excess Returns — Jason Hsu: China AI gap, capex arms race, S&P seven
Rayliant CIO: China models on-par/open-source; energy grid edge; hardware rents until overcapacity; Mag7 CapEx arms race; S&P is one-tree, not diversifier.
asr Excess Returns · Rayliant Global Advisors · Research Affiliates · DeepSeek
Long-form 05:25 ET
Odds on Open — Lihong Wang: ex-IMC semis quant, AI stack portfolio
Ex-IMC semis options MM on flow/V, NVDA–AMD relative vol, DeepSeek corr blowups; 50-name AI stack book at ~2×; models-beat-S&P claim needs harness.
asr Odds on Open · IMC · NVIDIA · AMD
Sun, 20 Sept 2026
Long-form 05:20 ET
MiB — Glen Kacher: AI boom is catch-up, not overbuild
Light Street CIO on AI5 semis concentration, NVDA ~85% share, demand ahead of supply, 10–20y stack cycle, agents→~5× tokens; DC politics as education risk.
asr Masters in Business · Light Street Capital · NVIDIA · AMD
Sat, 19 Sept 2026
Long-form 05:20 ET
a16z — Ali Ghodsi: enterprise stall is context, not IQ
Databricks CEO on pacing PR vs cyber risk, four-test RSI bar, ontology/Genie as the adoption bind, Uni Gateway cost control + GLM shift.
asr The a16z Show · Databricks · OpenAI · Hugging Face
Long-form 05:20 ET
No Priors — Ermon/Inception: diffusion wins inference parallelism
Stefano Ermon on Mercury ≈ Haiku/Flash/mini speed tier, ~10× decode vs AR at GPT-2 scale, OpenCall leaving Cerebras for NVDA GPUs, 20–30% latency wedge.
asr No Priors · Inception · Mercury · OpenAI
Fri, 18 Sept 2026
Long-form 05:20 ET
Dwarkesh — Noam Brown: agent swarms, RSI speedup, alignment bind
OpenAI's Noam Brown on 10k-agent Navier-Stokes solve (130B tokens/88h), Ultra Mode multi-agent, Codex $7–8k/day internal, RSI ≠ 100x overnight.
transcript Dwarkesh Podcast · OpenAI · Hugging Face · Astra
Thu, 17 Sept 2026
Long-form 05:20 ET
Latent Space — AIUC: trust/liability as the agent adoption bind
Rune Kvist (ex-Anthropic) on $40M Series A: AIUC-1 quarterly agent standard, Lloyd's-backed policies, Waymo/Air Canada liability — eval+insurance stack.
transcript Latent Space · AIUC · Anthropic · Cursor
Long-form 05:15 ET
All-In — Gerstner: no AI bubble; semis = ~70% of Nasdaq return
Altimeter's Brad Gerstner on All-In: earnings-driven tape, offtake must fund Mag5 capex, Dylan 43GW too hot (~25GW), lab RR as takeoff switch. [asr]
asr All-In Podcast · NVIDIA · Anthropic · OpenAI
Wed, 16 Sept 2026
Long-form 05:20 ET
SemiAnalysis Ep.031 — pacing may eat more compute, not less
Emergency ep on Amodei pacing: OpenAI CoT monitoring ~20% of rollup compute; safety spend likely raises, not cuts, infra demand; HF as shot across bow. [asr]
asr SemiAnalysis Weekly · Anthropic · OpenAI · Hugging Face
Long-form 05:15 ET
All-In — Satya: pace with common sense; MSFT builds, leases, rents
Nadella on All-In: broad diffusion over mystical slowdown; ~30m enterprise Copilot users of ~250–300m TAM; kit ~60% of cost; Quincy DC ~400–500 MW. [asr]
asr All-In Podcast · Microsoft · OpenAI · Anthropic
Tue, 15 Sept 2026
Long-form 07:30 ET
Elon Musk & Gwynne Shotwell — AI Peer Review, Starship, Terafab, SpaceX/Tesla Merger (All-In)
SpaceX President Gwynne Shotwell says SpaceX is as much an AI business as a space business by revenue, with compute rental 'a heck of a business' and Starlink at ~1.5–2% penetration; Elon Musk joins from Memphis to push cross-lab model peer review, handicap Starship ship-catch at ~50–60%, frame Terafab as build-or-fail-to-scale, and non-deny a Tesla–SpaceX combination.
asr All-In Podcast · SpaceX · Tesla · xAI
Long-form 07:30 ET
All-In — Jensen: doomer math fails; open models carry apps
Huang on All-In: extinction %s unscientific; ~$400bn AI VC ~80% open-model; NVIDIA goes "as deep as needed." Trump brands DC opposition a hoax. [asr]
asr All-In Podcast · NVIDIA · Anthropic · OpenAI
Mon, 14 Sept 2026
Long-form 20:50 ET
Jensen Huang — Nvidia's Future, Physical AI, Rise of the Agent, Inference Explosion (All-In)
NVIDIA CEO Jensen Huang tells the All-In hosts that agentic workloads drove a ~10,000x compute step in two years, that a higher-capex Vera Rubin factory can still deliver the lowest token cost via ~10x throughput, and that Physical AI is already a near-$10bn NVIDIA line while open-weight agents redefine the desktop OS — with China licenses restarting and consensus growth paths rejected as undersized.
asr All-In Podcast · NVIDIA · Groq · Anthropic
Long-form 20:30 ET
Gavin Baker — Why AI Demand Is Outrunning Compute Supply (a16z Show)
Atreides CIO Gavin Baker tells David George that AI fundamentals accelerated through July–August while related equities drew down; argues sub-one-year compute paybacks and thin heavy-user penetration make undersupply through 2028 the base case, with NVIDIA’s financeable stack and hybrid open-source routers as the durable structure.
asr The a16z Show · NVIDIA · OpenAI · Anthropic
Long-form 17:29 ET
TBPN: The AI Slowdown Debate
Metadata-only: TBPN's Sep 14 episode (full + Diet cut) titled The AI Slowdown Debate, with guests including Nico Wittenborn, Scott Keogh, Mitchell Green, David Rosenthal, Ben Gilbert, and Faraj Aalaei.
metadata-only TBPN
Long-form 13:00 ET
SemiAnalysis Weekly: why 4-Hi HBM may win on inference economics
Metadata-only capture of SemiAnalysis Weekly Ep. 030: Myron Xie and Jordan Nanos on Rubin Ultra shipping 192GB HBM versus a 1TB preview, supply-driven decontenting, and why less memory per chip can still be the right call.
metadata-only SemiAnalysis Weekly · NVIDIA · SemiAnalysis
Long-form 12:06 ET
Eisman Playbook: Big Short partners on rates, AI, gold
Metadata-only: Steve Eisman with Vincent Daniel and Porter Collins on bonds, Treasury buybacks, OpenAI risk, gold, shorting mechanics, and two live short ideas.
metadata-only The Real Eisman Playbook · OpenAI
Long-form 12:04 ET
Latent Space: Richard Socher on recursive self-improvement
Metadata-only: Latent Space interviews Richard Socher (Recursive / You.com) on recursive self-improvement as the next major AI step; show notes flag AI×finance conference promo.
metadata-only Latent Space · You.com · Recursive
Long-form · Sat, 19 Sept 2026 · 05:20 ET

No Priors — Ermon/Inception: diffusion wins inference parallelism

Stefano Ermon on Mercury ≈ Haiku/Flash/mini speed tier, ~10× decode vs AR at GPT-2 scale, OpenCall leaving Cerebras for NVDA GPUs, 20–30% latency wedge.

asr Stefano ErmonSarah Guo InceptionMercuryOpenAICerebrasNVIDIAOpenCallOpenRouterStanford Source ↗
Venue: No PriorsHost: Sarah GuoDuration: 38mPublished: Fri, 18 Sept 2026 · 06:00 ET

Opening

Diffusion language models are a bet that inference — not training — is the scarce axis: parallel token generation maps to GPUs the way transformers mapped training off RNNs. Stanford professor / diffusion pioneer Stefano Ermon (Inception co-founder & CEO) joins Sarah Guo on No Priors (~38m). Ground covered: 2019 score-based → modern diffusion; 2024 discrete-text breakthrough (~GPT-2 scale, matched AR perplexity, ~10× faster generation); Mercury models claimed on-par with Haiku/Flash/mini-nano while faster; custom serving stack (not vLLM/SGLang); ~50-person / ~2-year company; OpenCall voice agents switching from Cerebras custom silicon to Mercury on NVIDIA GPUs; 20–30% of OpenRouter-style workloads as latency-bound addressable wedge; RSI still human-bottlenecked. YouTube auto-captions (asr); Watch.

Key takes

Architecture claim: autoregressive decode is still sequential (token N needs 1…N−1) → memory-bound on GPUs; diffusion generates many tokens in parallel → training-like matmul-heavy workload at inference. Historical rhyme: 2017 RNN→transformer unlocked training parallelism; Ermon argues diffusion is the inference-time analogue. Bitter lesson: more parallel wins. Economics = intelligence per watt / per dollar; test-time compute and RL rollouts are both inference-bound, so better inference scaling compounds post-training too. [asr]

2024 academic proof point: discrete diffusion transformer matched AR perplexity at <1B / GPT-2 scale on same data — and generated text ~10× faster. That result justified spinning out Inception to scale commercially. Continuous modalities (image/video/music/protein) already diffusion-dominated; discrete text/code was the open bet vs lab AR consensus. [asr]

Product claim today: Mercury models benchmark-comparable to frontier labs' speed-optimized small models (Haiku / Flash / mini-nano class) while significantly faster; served in production with an in-house engine. Cannot drop onto vLLM/SGLang — serving IP is part of the moat. Still transformer+attention underneath; scrapes public data/evals; differentiation is speed now, possible intelligence/data-efficiency later (denoising as augmentation). [asr]

Company snapshot: ~2 years old, ~50 people; bottleneck more compute than headcount; product team serves current Mercury + research trains next (training/RL/inference). Attracts green-field researchers; uses frontier-lab models to accelerate own R&D but says human idea-pruning still binds — "not there yet" on RSI. [asr]

Where speed wins now: latency-sensitive apps, especially voice agents. Named customer OpenCall — previously on Cerebras for required speed; switched to diffusion LLMs to match custom-chip latency on commodity NVIDIA GPUs (availability + cost). Software acceleration multiplicative with future hardware. Controllability: coarse-to-fine generation steerable earlier than AR (academic literature); API kept OpenAI-compatible text-in/text-out so harnesses transfer. [asr]

Workload split (Ermon estimate): ~20–30% of tasks (OpenRouter taxonomy) are latency-critical enough that highest-quality-within-latency-budget models address them — floor on diffusion TAM near-term. Frontier intelligence still required for much of the rest; diffusion not yet at that frontier. Adoption friction = immature ecosystem (kernels, serving) built in-house. [asr]

Key math

2024 paper: GPT-2-scale (<1B params) discrete diffusion matched AR perplexity; ~10× faster text generation (asr — Ermon) — founding technical claim. [book]

Mercury ≈ Haiku / Flash / OpenAI mini-nano quality tier on benchmarks, significantly faster (asr — Ermon; vendor claim) — commercial positioning. [book]

Inception ~2 years old; ~50 people (asr) — scale. [book]

OpenCall: Cerebras custom silicon → Mercury on NVIDIA GPUs at comparable voice latency (asr — Ermon) — hardware-substitution datapoint. [book]

~20–30% of OpenRouter-classified workloads latency-bound / addressable by speed-optimized models (asr — Ermon estimate) — TAM floor. [book]

Software × hardware gains claimed multiplicative (asr) — stacking argument vs custom ASIC only.

Quotes

"Autoregressive models are still sequential… You cannot generate the 10th token until you've generated everything that comes before it. That kind of workload… is extremely memory bound." — Stefano Ermon [asr]

"The equivalent at inference time is a diffusion model because a diffusion model is built to… process many tokens at the same time." — Stefano Ermon [asr]

"We were able to generate text like 10x faster compared to the autoregressive model." — Stefano Ermon [asr] (GPT-2-scale 2024 result)

"Our Mercury models are on par with the Haiku models, Flash models, mini nano models from OpenAI… while being significantly faster." — Stefano Ermon [asr]

"They switched over to our diffusion based LLMs because they can essentially get the same speed as… an autoregressive model on custom hardware… on Nvidia GPUs." — Stefano Ermon [asr] (OpenCall)

"Between 20 and 30% where latency is really really important… as a lower bound I think it could be addressable." — Stefano Ermon [asr]

"The bottleneck is more compute than people." — Stefano Ermon [asr]

Variant perception

Priced in — inference efficiency matters; small/fast model tiers (Haiku/Flash/mini); voice latency as a product constraint; custom inference chips (Cerebras/Groq/etc.) as a speed path. Markets already trade "inference is the next bottleneck."

What's new — a principal (diffusion's academic father) shipping production discrete diffusion LLMs with a named custom-silicon→GPU substitution (OpenCall), a quantified ~10× academic speed claim, and an explicit 20–30% latency-wedge TAM. Reframes the hardware race: software parallelism on NVIDIA may substitute for scarce specialty inference silicon in some voice/agent stacks. Serving-engine moat is underappreciated vs model weights alone.

Bear case — Mercury quality claims are vendor-benchmark; not independently audited here. 10× is GPT-2-scale research, not frontier Mercury vs frontier AR. Diffusion may lose on hard reasoning even if it wins latency. Labs can add speculative decoding / other AR speedups and close the gap without architectural switch. 20–30% is Ermon's OpenRouter eyeball, not revenue mix.

Discount — Ermon is CEO of Inception raising/selling Mercury; Sarah Guo (Conviction) is a friendly AI-infra interviewer. Episode is a category-creation pitch. Treat "will win inference" as thesis, not settled fact — but the GPU-mapping argument and customer silicon-substitution are concrete enough to track.

Positioning

Inference margin inversion — STRENGTHENS (mechanism). Parallel decode + software speed on commodity GPUs is a direct path for cost/latency per token to fall without waiting for a new chip cycle; voice agents paying for speed today validate willingness-to-pay for the fast tier. If diffusion (or AR speedups it forces) works, serving gross margins can widen even as list prices drop.

AI capex durability — MIXED / watched. Better tokens-per-watt could reduce accelerator intensity per query — but also unlocks more real-time agent/voice usage (volume up). Ermon's own bottleneck ("more compute than people") and RL-rollout inference hunger still pull demand. Net: supports continued GPU scarcity even if architecture mix shifts.

Enterprise agent stall — NEUTRAL / mild help. Faster, more controllable generation helps voice/tool agents feel production-ready, but does not solve eval/liability. OpenAI-compatible API eases harness swap — adoption friction is stack maturity, not just model IQ.

HBM supply binds — NEUTRAL / slight weaken if true. Memory-bound AR decode is partly an HBM/bandwidth story; compute-heavy parallel diffusion reframes the binding constraint toward FLOPs vs bandwidth for some workloads — watch, do not reprice yet.

The Open/Close  ·  Research commentary, not investment advice. Positions may be held in securities mentioned.