Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 · 1:42:54
Philip Kiely and Ali Taha of Baseten join Swyx to explain what actually happens when a 200,000-token request hits a production inference system, arguing that stacking quantization, speculative decoding, and disaggregated prefill/decode can make open models like GLM-5.2 up to 10x faster. They detail cache-aware routing, training traffic-specific speculators, and why quantization errors can cancel out so a more-quantized model beats a less-quantized one. They also reveal how Baseten grafted Kimi's vision encoder onto GLM-5.2, discuss NVIDIA Dynamo as a toolkit rather than a turnkey speedup, and explain why they are bearish on mega kernels. The conversation covers video generation's compute barriers, the trend toward ASIC-like GPUs, and the emerging loop where GLM-5.2 writes the GPU kernels that serve itself.