Steadcast
No Priors cover art
No Priors

Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown

June 26, 202636 min

Show notes

When a new AI model drops, it’s judged based on a static benchmark grid that doesn’t account for how long the model is allowed to think. How then should we measure a model’s true capability? OpenAI research scientist Noam Brown returns to talk with Sarah Guo about his latest essay on why the AI industry’s traditional benchmark grids are broken, and how large-scale test-time compute is fundamentally changing how models are evaluated.

Transcript

Transcript not available for this episode yet.

More from No Priors

Coinbase’s Everything Exchange: Agentic Finance, Stablecoins, and Tokenization with CEO Brian Armstrong

Sep 10, 202645 min

Redefining Chip Architecture with Arm CEO Rene Haas

Sep 3, 202637 min

Rethinking Legacy Data Infrastructure with Eon Co-Founders Ofir Ehrlich and Gonen Stein

Aug 27, 202634 min

From Restoring Sight to Reimagining the Brain, with Max Hodak

Aug 20, 202631 min

What Chess.com Teaches US About Superhuman Capabilities, with CEO Erik Allebest

Aug 13, 202646 min