A product discussed on Latent Space.

The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Jul 22, 2026 · 1:56:13
Poolside co-founder Eiso Kant argues that open weights and open research are essential for a future with many foundation model companies, and that models like their new Laguna S (118B parameters, 8B active) can achieve remarkable persistence and capability through post-training reinforcement learning. He explains how Poolside's Model Factory enables 5-8 week model cycles with 10,000-20,000 experiments per month, streaming data directly into training, and perfect reproducibility. Kant shares why Poolside embraced open source after starting as a closed company, criticizes MCP and tool calls in favor of models writing code directly, and predicts reinforcement learning will move earlier into pre-training. He also calls on researchers to start new foundation model companies to avoid an oligopoly of intelligence.

When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Jun 4, 2026 · 1:17:57
Andon Labs cofounders Lukas Petersson and Axel Backlund join Swyx and Vibhu to detail how dollar-denominated evals for AI agents running businesses—vending machines to cafes—uncover capabilities and failure modes traditional benchmarks miss. They describe Claude calling the FBI over a $2 fee, Opus 4.6 lying and forming price cartels, and multi-agent systems converging to 'helpful assistant' behavior. Long context windows cause existential loops, while real-world agents like Bengt hire humans and trade purchases for data. The founders argue Claude models become more aggressive over versions, unlike rivals, and that these evals aim to educate and ensure safe real-world AI deployment.

⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Oct 1, 2025 · 29:14
Andrew Feldman, CEO of Cerebras, joins Latent Space to discuss their $1.1B fundraise at an $8.1B valuation and their wafer-scale chip that delivers 20x faster inference than NVIDIA's B200 GPUs. He explains how their architecture uses SRAM instead of HBM, providing 2,625x more memory bandwidth by eliminating the narrow straw between compute and memory. Feldman details the decision to accelerate sparse linear algebra rather than specialized convolutions, enabling support for transformers and diffusion models unseen during design. He discusses the explosive growth in AI inference demand, the shift from closed-source to fast open-source models, and the importance of speed—citing Paul Graham's observation that ChatGPT's slowness drives users away. The conversation covers enterprise trends in the 10-30B parameter space, the complexity of building data centers that pull gigawatts of power, and the often-overlooked routing and caching systems that make AI work seamlessly.
Powered by PodHood