Guest on Latent Space.

Trillion Token Context. No, Really — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding
Sep 4, 2026 · 27:02
Accelerated Understanding co-founders Anima Anandkumar and Benedikt Jenik argue the universality and scale of language models can extend to physics via a single foundation model trained across fluid dynamics, semiconductors, and energy. They report that models trained across multiple physics domains outperform equally sized single-domain models, evidencing genuine transfer learning. Jenik says they train at up to trillion context and infer at five trillion, with 22-terabyte outputs, requiring reinvented sharding infrastructure. Anandkumar says neural operators give resolution invariance where transformers' quadratic complexity fails. Simulator-generated curricula and physics-based self-improvement provide dense training signals, with first customers in semiconductor design and geothermal.

🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Aug 26, 2026 · 1:23:32
Anima Anandkumar, Caltech professor and ex-NVIDIA lead, argues physical-world modeling needs neural operators that bake in physics, not language-style scaling; it already rivals supercomputers. Physics-informed nets fail; Fourier neural operators capture multi-scale non-local phenomena, and FourCastNet, trained on 50,000 reanalysis samples, forecasts weather tens of thousands of times faster on a consumer GPU. FourCastNet 3’s spherical harmonics keep month-long rollouts stable for ensemble climate prediction; the same operators build a fusion digital twin a million times faster than simulation and inverse design. TorchLean ports PyTorch to Lean for formal robustness proofs, and on the UN Science Advisory Board she urges AI for science not be regulated like chatbots. Bottleneck: more compute.
Powered by PodHood