

Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown
When a new AI model drops, it’s judged based on a static benchmark grid that doesn’t account for how long the model is allowed to think. How then should we measure a model’s true capability? OpenAI research scientist Noam Brown returns to talk with Sarah Guo about his latest essay on why the AI industry’s traditional…
We score an episode from what we can actually measure — its reach and engagement, what listeners say about it, and what it covers. We don't have those signals for this one yet, so it doesn't get a number.
Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.
Rate this episode
Add your vote to the Hive — listeners rate every episode after they finish.
- No reviews yet — be the first.
More from No Priors: Artificial Intelligence | Technology | Startups
See all episodes →
From Restoring Sight to Reimagining the Brain, with Max Hodak

Redefining Chip Architecture with Arm CEO Rene Haas

Rethinking Legacy Data Infrastructure with Eon Co-Founders Ofir Ehrlich and Gonen Stein
Similar episodes from other shows
More like No Priors: Artificial Intelligence | Technology | Startups →
Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI

AI won't plateau — if we give it time to think | Noam Brown

Why AI Needs Better Benchmarks

From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki
One great episode in your inbox, daily — free.
One email a day, unsubscribe anytime. No spam, ever.