A product discussed on Latent Space.

The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Jul 22, 2026 · 1:56:13
Poolside co-founder Eiso Kant argues that open weights and open research are essential for a future with many foundation model companies, and that models like their new Laguna S (118B parameters, 8B active) can achieve remarkable persistence and capability through post-training reinforcement learning. He explains how Poolside's Model Factory enables 5-8 week model cycles with 10,000-20,000 experiments per month, streaming data directly into training, and perfect reproducibility. Kant shares why Poolside embraced open source after starting as a closed company, criticizes MCP and tool calls in favor of models writing code directly, and predicts reinforcement learning will move earlier into pre-training. He also calls on researchers to start new foundation model companies to avoid an oligopoly of intelligence.

Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 8, 2026 · 1:26:00
NVIDIA's Nader Khalil and Kyle Kranen join Swyx and Vibhu to explain how the company moves like a $4.3 trillion startup through speed-of-light (SOL) first-principles thinking, agent security boundaries, and the Dynamo inference engine. They argue agents should only do two of three things (files, internet, code) to prevent vulnerabilities, and detail Brev's acquisition to improve developer UX with one-click GPU access and DGX Spark integration. Kyle describes Dynamo as a data center scale inference engine that optimizes serving by scaling out, using prefill/decode disaggregation, Kubernetes-based scheduling, and model-hardware co-design to improve cost, latency, and quality. The episode covers SOL's role in creating urgency, long-context limits and potential 'unhobblers' like multi-head latent attention, and the shift toward CLI-first agent workflows for enterprise tools.

LIVE from GTC: DGX Spark Insides First Look
Mar 20, 2025 · 12:39
Israel from NVIDIA gives a first look at the DGX Spark, a $3,000–$4,000 mini AI supercomputer that fits up to 200B parameter models in FP4 with 128GB unified memory, powered by the NVIDIA GB10 Superchip. It uses the same Grace Blackwell architecture and software stack as data center systems, so code written on the Spark deploys to production without modifications. The device includes a ConnectX-7 dual-port 200Gb ethernet for clustering two units, and runs DGX OS (Ubuntu 24.04). Storage is 1TB or 4TB NVMe. Partners ASUS, HP, Dell, and Lenovo will offer their own cases. Israel stresses the Spark is a developer box, not a server, but delivers enterprise-grade networking and shared memory in a compact form.
Powered by PodHood