Diogo Almeida 1 item
‹ All posts People /Diogo Almeida /Long-form ARCHIVE
Long-form · Wed, 23 Sept 2026 · 05:15 ET

Latent Space — Diogo Almeida / Jev: System One for prod, not God

InstructGPT coauthor: frontier chat/RLHF APIs wrong for software; Jev as code-consumed System One; >1T tokens/day machine traffic; dark data + agents.

asr Diogo AlmeidaswyxAlessio Fanelli TypeSafeJevOpenAIAnthropicChatGPT Source ↗
Venue: Latent SpaceHost: swyxDuration: 141mPublished: Mon, 21 Sept 2026 · 18:13 ET

Opening

Frontier chat/RLHF APIs are the wrong product for production software — TypeSafe's Jev is pitched as a "System One" / machine-native model class optimized for code-as-consumer, intelligence-per-dollar, calibration, and reliability nines that behave like a database query, not a coworker chatbot. Latent Space (~141m) with Diogo Almeida (TypeSafe CEO; InstructGPT coauthor; ex-OpenAI) days after the Jev launch that dominated the AI timeline. Ground covered: System One vs pretrained LM / RLHF chat / RLVR; refusals as type errors in dependencies; launch metrics (machine token traffic, Discord); use-case map (dark data, coding agents, real-time); KV-cache tyranny for agents; roadmap toward an "AWS of intelligence." YouTube auto-captions asr — Watch. Figures marked (asr).

Key takes

Thesis: code should consume the model — "machine native System One / large programmable" — vs pretrained LMs (internet autocomplete), RLHF chatbots (reply-to-text), and RLVR (gray area). Almeida: TypeSafe designs internals for software integration; Jev is the first of that class, branded for the intelligence-per-dollar frontier (name → Jevons). Tradeoff stack he wants debated: reliability, cost, calibration, speed — not chat Elo. [asr]

RLHF mode-dropping collapses calibration and is poison for programmable use. Soapbox: RLHF drops minority modes (GAN analogy) so long strings look fine while subtle wrong answers are hard to see; Yan LeCun's "LLMs doomed as sequence length grows" slide is "mathematically obvious but empirically wrong" once mode coverage is understood. Calibration / confidence is the under-discussed failure mode relative to chat polish. [asr]

API refusals are a type error — tolerable in ChatGPT, insane in a background dependency. Product chat can refuse NSFW; an API that stochastically breaks because a user message tripped safety is "anti-user" and comes from "horseless carriage / AI coworker" obsession rather than making intelligence as boring as a database query. He wants many more reliability nines before intelligence is "just there." [asr]

Launch claim: machine-driven traffic already surpassed a trillion tokens/day milestone — night churn, not only demo signups. Signups/waitlist offboarding called a mistake for a developer platform (many non-devs); platform uptime claimed "more up nines than Anthropic" through an "unprecedented launch." Discord cited at ~100k. If TypeSafe disappeared, catch-up framed as maybe ~1–2 years if model quality matters. [asr]

Volume thesis: "dark data" (corp piles too expensive for frontier LMs) + coding agents are the big money; real-time / e-commerce / assistants love every ~10ms shaved. Structured JSON for state / instructions / criteria — "thinking in templates/system messages is the old way." Brand constraint: stay pre-frontier on intelligence-per-dollar (and intelligence-per-second as a separate metric); may temporarily LTS Jev 1.13.0 rather than fracture the fleet while shipping fast. Aspiration: layers toward an "AWS of intelligence" / TCP of System One. [asr]

Exit narrative: left OpenAI worried about AI winter from RLHF overpromise-vs-underdeliver; launch week framed as proof automation utility is back. TFP growth "3% in 5 years" line as economic-revolution north star vs lab charter drift to "$100 billion in profit." "Stuff in the tank" beyond this "low-key research preview." [asr]

Key math

>1 trillion tokens/day milestone passed; night-continuous machine traffic (asr — vendor claim) — production-vs-demo tell. [asr]

Discord ~100,000 (asr — host cite) — community scale at interview. [asr]

Catch-up if TypeSafe vanishes: ~1–2 years (asr — Almeida guess, model-quality contingent) — moat framing. [asr]

User-facing latency budgets ~100ms–1s "magical"; halving latency ≈ double sequential intelligence calls (asr) — intelligence-per-second logic. [asr]

~10ms shave known to CEOs/CTOs in real-time products (asr — conjecture) — e-commerce/assistant demand. [asr]

Jev 1.13.0 as possible temporary LTS (asr) — versioning/fleet constraint. [asr]

TFP growth ~3% in 5 years as stated aspiration (asr) — macro north star, not forecast model. [asr]

Quotes

"We need a new class of models… system one models… the goal is for code to be the consumer." — Diogo Almeida [asr]

"Jev is meant to be optimized for intelligence per dollar." — Diogo Almeida [asr]

"Refusal is just like obviously a type error." — Diogo Almeida [asr]

"A trillion tokens a day is a lot… surpassing that is awesome." — Diogo Almeida [asr]

"We want to be deep in the guts of programs because that's how you make software powerful." — Diogo Almeida [asr]

"I don't care how much smarter it is. It needs to be in the pre-frontier." — Diogo Almeida [asr]

Variant perception

Priced in — ChatGPT-shaped APIs dominate spend; agent/coding demos are hot; enterprise wants cheaper/faster/more reliable inference; RLHF refusal friction is a known developer complaint; Jev launch was already timeline-saturating.

What's new — Explicit "System One / machine-native / code-as-consumer" product class vs chat; RLHF mode-drop as calibration poison with programmable consequences; refusal-as-type-error in dependencies; claimed >1T tokens/day machine traffic within days of launch; dark-data + coding-agent volume map; pre-frontier intelligence-per-dollar brand discipline; structured JSON I/O as first-class vs prompt templates.

Bear case — Token/day and uptime claims are launch-week vendor metrics; "System One" may be rebranded classification/small-model stack; chat frontier labs can ship cheap fast calibrated endpoints and erase the wedge; coding-agent volume may still prefer KV-cache-native frontier models (his own "tyranny of the KV cache" essay admits the gap); 1–2y catch-up admits thin proprietary moat if quality is the scarce input.

Discount — CEO mid-launch on a friendly technical podcast; maximum book-talking and hiring funnel. Auto-caption asr garbles proper nouns (Diogo/Diego, Jevons/Jevad, RLHF/RHF). Show notes/launch video hype (~40M views cited on site, not verified here) surrounds the interview.

Positioning

Enterprise agent stall — WEAKENS (soft) on mechanism, NEUTRAL on Fortune-500 $ proof. Argues production failure mode is wrong model class (chat/refusal/calibration), not missing capability — and claims machine traffic already in real work. Still no named F500 production deployment with headcount or dollars attached.

Inference margin inversion — soft STRENGTHENS (demand side). Intelligence-per-dollar / pre-frontier brand and dark-data volume imply serving demand that does not require frontier list prices; no lab GM arithmetic.

AI capex durability — NEUTRAL / soft STRENGTHENS. Machine token churn and "AWS of intelligence" aspiration assume continued serving scale; no hyperscaler build evidence.

HBM supply binds — NEUTRAL. No semiconductor supply content.

The Open/Close  ·  Research commentary, not investment advice. Positions may be held in securities mentioned.