Person
Stefano Ermon
1 document in the archive.
Appearances newest first
Stefano Ermon on Mercury ≈ Haiku/Flash/mini speed tier, ~10× decode vs AR at GPT-2 scale, OpenCall leaving Cerebras for NVDA GPUs, 20–30% latency wedge.
1 document in the archive.