The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Sep 2, 2026 · 44:03
A comedic episode parodying AI hardware marketing, featuring Sean Lie discussing fictional Cerebras CS4 chips with absurd specs like 30 billion parameter models running at 4,400 TPS with wafer-scale integration and jalapeño-flavored SRAM designs.