
Latent Space
Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs
June 4, 20261h 15m
Show notes
The new AIEWF website is live! Get your tickets booked ASAP as they -will- sell out. Take the AI Engineering Survey and get >$2k in credits and free AIE WF tickets! Most industry benchmarks compress intelligence and reasoning ability into scores. SWE-Bench Pro, MMLU, Humanity’s Last Exam, etc. These metrics are useful, but don’t always represent the full extent of how a model performs in the real world.
Transcript
Transcript not available for this episode yet.
More from Latent Space

The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Aug 3, 20261h 41m

Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Jul 28, 20261h 9m

Inside the Model Factory — Eiso Kant, Poolside AI
Jul 23, 20261h 54m

🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
Jul 21, 20261h 29m

🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences
Jul 16, 20261h 41m