No Priors
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown
June 26, 202636 min
Show notes
When a new AI model drops, it’s judged based on a static benchmark grid that doesn’t account for how long the model is allowed to think. How then should we measure a model’s true capability? OpenAI research scientist Noam Brown returns to talk with Sarah Guo about his latest essay on why the AI industry’s traditional benchmark grids are broken, and how large-scale test-time compute is fundamentally changing how models are evaluated.
Transcript
Transcript not available for this episode yet.
More from No Priors
Coinbase’s Everything Exchange: Agentic Finance, Stablecoins, and Tokenization with CEO Brian Armstrong
Sep 10, 202645 min
Redefining Chip Architecture with Arm CEO Rene Haas
Sep 3, 202637 min
Rethinking Legacy Data Infrastructure with Eon Co-Founders Ofir Ehrlich and Gonen Stein
Aug 27, 202634 min
From Restoring Sight to Reimagining the Brain, with Max Hodak
Aug 20, 202631 min
What Chess.com Teaches US About Superhuman Capabilities, with CEO Erik Allebest
Aug 13, 202646 min