A product discussed on Latent Space.

Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 · 1:42:54
Philip Kiely and Ali Taha of Baseten join Swyx to explain what actually happens when a 200,000-token request hits a production inference system, arguing that stacking quantization, speculative decoding, and disaggregated prefill/decode can make open models like GLM-5.2 up to 10x faster. They detail cache-aware routing, training traffic-specific speculators, and why quantization errors can cancel out so a more-quantized model beats a less-quantized one. They also reveal how Baseten grafted Kimi's vision encoder onto GLM-5.2, discuss NVIDIA Dynamo as a toolkit rather than a turnkey speedup, and explain why they are bearish on mega kernels. The conversation covers video generation's compute barriers, the trend toward ASIC-like GPUs, and the emerging loop where GLM-5.2 writes the GPU kernels that serve itself.

Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 8, 2026 · 1:26:00
NVIDIA's Nader Khalil and Kyle Kranen join Swyx and Vibhu to explain how the company moves like a $4.3 trillion startup through speed-of-light (SOL) first-principles thinking, agent security boundaries, and the Dynamo inference engine. They argue agents should only do two of three things (files, internet, code) to prevent vulnerabilities, and detail Brev's acquisition to improve developer UX with one-click GPU access and DGX Spark integration. Kyle describes Dynamo as a data center scale inference engine that optimizes serving by scaling out, using prefill/decode disaggregation, Kubernetes-based scheduling, and model-hardware co-design to improve cost, latency, and quality. The episode covers SOL's role in creating urgency, long-context limits and potential 'unhobblers' like multi-head latent attention, and the shift toward CLI-first agent workflows for enterprise tools.
Powered by PodHood