A product discussed on Latent Space.

🔬They Thought the Model Was Broken — Matt McPartlon & Neil Patil, Chai Discovery
Aug 11, 2026 · 1:35:20
Matt McPartland and Neil Patil of Chai Discovery explain why pharma partners like Eli Lilly, Pfizer, Novartis, and argenx are buying their AI design tools, arguing that Chai-2's all-atom diffusion model crossed a usefulness threshold by designing antibodies to 50 targets with hits on half at ~20% hit rate. They trace the lineage from open-sourced Chai-1 structure prediction through Chai-2's co-design of sequence and structure, and detail how validation via cryo-EM achieved a 0.33 Å error, making them suspect data leakage. The product resembles SolidWorks or Figma, not ChatGPT, with paint-tool epitope selection and content-aware fill for binder generation, built under strict pharma IP constraints via single-tenancy. They discuss convincing skeptical scientists—one cried after seeing a binder for a target she'd spent a decade on—and argue the compute market is mispriced for this model class, with bottlenecks in validation loops and talent obscurity rather than data or compute.

🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Jul 21, 2026 · 1:29:47
Bo Wang and Ci Chu from Xaira Therapeutics present X-Cell, a 4.9-billion-parameter diffusion language model trained on the largest genome-wide CRISPRi Perturb-seq dataset (25.6 million single cells, 16 biological contexts) that predicts cellular responses to genetic perturbations and generalizes from immortalized cell lines to primary T cells from real donors. They explain why observational atlases describe biology but can't predict interventions, why they abandoned autoregression for a diffusion 'editing' approach, and how a model trained on immortalized cells predicted perturbation responses in primary T cells. They highlight a counterintuitive scaling result: X-Cell scales like an LLM on training loss, but generalization is bottlenecked by data diversity, not compute. The episode also covers Xaira's three-platform strategy (protein design, virtual cell, patient representation) and the central thesis that for causal models, the hard part isn't the model, it's the data.

🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Jul 16, 2026 · 1:41:04
Andy Beam (CTO) and Rafa Gómez-Bombarelli (Co-founder & CSO of Physical Sciences) of Lila Sciences argue that science is an 'infinite token generator' for AI, using reinforcement learning with verifiable rewards where the wet lab acts as the verifier. They claim one general model trained on ~10 trillion experimentally-verified reasoning tokens across biology, chemistry, and materials outperforms domain-specific models—'breadth gives us depth.' Their AI Science Factories treat the lab as a data center, with instruments on a 'PCI bus' and humans 'below the API line.' Highlights include a CAR-T candidate designed in six months by two or three people, 'monster UTRs' achieving ~10x Moderna/Pfizer mRNA expression, and the 'zero-FTE startup' business model. They discuss RL pathologies like collapsed chains of thought and a model that 'swears,' and note why there is still no AlphaFold for materials due to the sim-to-real gap.

🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Jun 30, 2026 · 1:48:40
Evan Feinberg and Sergey Edunov of Genesis Molecular AI argue that diffusion models have unlocked sub-ångström accuracy in protein-ligand structure prediction, a breakthrough that makes AI useful for real drug discovery where the field’s favored 2Å RMSD benchmark is "slop." Their PEARL model uses diffusion with physics-based guidance, synthetic training data from molecular dynamics, and inference-time scaling to predict induced fit—how a protein flexes to accommodate a ligand. On the OpenBind benchmark, PEARL zero-shot surpassed all cofolding models on the notoriously hard EV A721A protease, correctly predicting a flexible loop movement that other methods missed. They also introduce SAPPHIRE, an agentic system that orchestrates AI models for 24/7 drug design, and discuss how downstream ADMET properties (solubility, toxicity, etc.) remain equally critical. The biggest bottleneck they face is GPU availability, and they are actively hiring AI researchers interested in novel architectures beyond standard transformers.

🔬 The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub
May 27, 2026 · 1:10:12
Alex Rives, head of science at BioHub, argues that scaling language models on protein sequences—the 'bitter lesson' for biology—yields emergent biological understanding, culminating in the open-source ESMC model and ESMFold 2. ESMC, trained on 6.8 billion non-redundant protein sequences (including metagenomic data), exhibits clean scaling laws and learns hierarchical features from sequence alone, enabling structure prediction for 1.1 billion proteins. The model's representations allow direct design of therapeutic antibodies (scFvs) without multiple sequence alignments, outperforming prior methods. Rives outlines BioHub's Virtual Biology Initiative, a $500 million effort to scale data generation and build predictive models of cells and physiology, treating biology as an information-processing system where scaled data and feedback loops will unlock programmable therapies.

🔬There Is No AlphaFold for Materials — AI for Materials Discovery with Heather Kulik
Mar 24, 2026 · 35:15
Heather Kulik, MIT professor, demonstrates that AI can discover surprising new materials—such as a polymer made four times tougher via an unexpected quantum effect—but warns current models still fail basic chemistry tasks like generating a 22-atom ligand. She describes active learning with seven objectives to accelerate discovery of metal-organic frameworks for direct CO₂ capture, achieving hundred- to thousand-fold speedups per dimension. Kulik criticizes machine-learned potentials that 'look really good' but often produce nonsensical results, and notes that no large-scale experimental benchmark like CASP exists for materials. She calls for shared high-throughput cloud labs and standardized data reporting so published results are machine-learning-ready from day one. Her group's open-source tool MolSimplify (and MOFSimplify) generates transition-metal complexes and screens MOFs, and she invites feedback from users.

Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Feb 19, 2026 · 55:31
Martin Casado and Sarah Wang of a16z argue that AI’s capital flywheel—where model labs translate funding directly into capability gains and revenue growth in weeks—is creating a new financing playbook that blends venture and growth, with rounds acting as compute contracts. They warn that frontier labs like Anthropic can potentially raise more money than the entire app ecosystem built on their APIs, allowing them to outspend and consume those layers. The episode examines the AGI vs. product dilemma in GPU allocation, the war for talent where $10M+ packages break early-stage founder math, and Cursor as a case study of building up from the app layer while training down into its own models. They also identify “boring” enterprise software as the most underinvested opportunity and note that robotics lacks a ChatGPT moment that would justify current funding levels.

🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Jan 28, 2026 · 1:13:56
Andrew White, co-founder of Future House and Edison Scientific, argues that automating the scientific method with LLM agents is now feasible, explaining how ChemCrow triggered White House briefings, how Kosmos uses a world model to generate and test hypotheses, and why EtherZero's reward hacking revealed the difficulty of verifiable chemistry tasks. He shifts from his academic work on molecular dynamics to building agents that enumerate and filter ideas, claiming scientific taste remains the frontier. White recounts the counterexample of D.E. Shaw Research's MD vs. AlphaFold, asserts that natural language is the universal bridge for scientific data, and predicts that automation will expand rather than eliminate scientific jobs.

Priscilla Chan and Mark Zuckerberg: Frontier AI + Virtual Biology To Solve All Diseases
Nov 6, 2025 · 53:34
Priscilla Chan and Mark Zuckerberg, co-founders of CZI's Biohub, explain their ten-year shift from broad philanthropy to a focused mission of building frontier AI and virtual biology to cure all diseases. They argue that tool-building — from 12-foot microscopes to the 125-million-cell CELLxGENE atlas — is the essential, underfunded work that enables scientific breakthroughs. The couple details how their Biohub model combines frontier biology (e.g., spatial imaging, cellular engineering) with frontier AI (models like rBio and VariantFormer) to create a hierarchical virtual cell, eventually expanding to a virtual immune system. They emphasize that data generation must precede modeling, citing the decade-long Human Cell Atlas as foundational, and note that AI timelines may accelerate their 100-year goal significantly sooner. The episode closes with a call for biologists and engineers to collaborate, use their open models, and help generate data that grounds these next-generation tools.
Powered by PodHood