🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Jul 21, 2026 · 1:29:47
Bo Wang and Ci Chu from Xaira Therapeutics present X-Cell, a 4.9-billion-parameter diffusion language model trained on the largest genome-wide CRISPRi Perturb-seq dataset (25.6 million single cells, 16 biological contexts) that predicts cellular responses to genetic perturbations and generalizes from immortalized cell lines to primary T cells from real donors. They explain why observational atlases describe biology but can't predict interventions, why they abandoned autoregression for a diffusion 'editing' approach, and how a model trained on immortalized cells predicted perturbation responses in primary T cells. They highlight a counterintuitive scaling result: X-Cell scales like an LLM on training loss, but generalization is bottlenecked by data diversity, not compute. The episode also covers Xaira's three-platform strategy (protein design, virtual cell, patient representation) and the central thesis that for causal models, the hard part isn't the model, it's the data.