

Benchmarks 201: Why Leaderboards > Arenas >> LLM-as-Judge
The first AI Engineer World’s Fair talks from OpenAI and Cognition are up! In our Benchmarks 101 episode back in April 2023 we covered the history of AI benchmarks, their shortcomings, and our hopes for better ones. Fast forward 1.5 years, the pace of model development has far exceeded the speed at which benchmarks…
We score an episode from what we can actually measure — its reach and engagement, what listeners say about it, and what it covers. We don't have those signals for this one yet, so it doesn't get a number.
Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.
Rate this episode
Add your vote to the Hive — listeners rate every episode after they finish.
- No reviews yet — be the first.
More from Latent Space: The AI Engineer Podcast
See all episodes →
🔬Scaling Past Informal AI - Carina Hong, Axiom Math

🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

Runway’s WorldPrompt and the Engineering of Real-Time Worlds
Similar episodes from other shows
More like Latent Space: The AI Engineer Podcast →
Why AI Needs Better Benchmarks

Optimizing for efficiency with IBM’s Granite

Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown

20VC: 70% of Neolabs Will Die | There Will be a $100BN US Open-Source Model | Data is a Trillion $ Market | Governments Cannot Regulate Models: It is Too Late | The Cyber Attacks to Come Will be Insane with Anastasios Angelopoulos @ Arena
One great episode in your inbox, daily — free.
One email a day, unsubscribe anytime. No spam, ever.