
All episodes
BUZZING· 73%58m·2024-07-12·Technology

Benchmarks 201: Why Leaderboards > Arenas >> LLM-as-Judge
Latent Space: The AI Engineer Podcast
The first AI Engineer World’s Fair talks from OpenAI and Cognition are up! In our Benchmarks 101 episode back in April 2023 we covered the history of AI benchmarks, their shortcomings, and our hopes for better ones. Fast forward 1.5 years, the pace of model development has far exceeded the speed at which benchmarks…
Buzzmeter · cross-platform

73%
BUZZING
Active swarm energy. Strong reception, real community buzz.
Hive Vote · Skep listeners
New on SkepAwaiting the hive's first votes
Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.
Rate this episode
Add your vote to the Hive — listeners rate every episode after they finish.
Recent reviews
- No reviews yet — be the first.
More from Latent Space: The AI Engineer Podcast
See all episodes →
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
Latent Space: The AI Engineer Podcast · 57m

🔬Doing Vibe Physics — Alex Lupsasca, OpenAI
Latent Space: The AI Engineer Podcast · 1h 31m

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
Latent Space: The AI Engineer Podcast · 1h 25m
One great episode in your inbox, daily — free.
One email a day, unsubscribe anytime. No spam, ever.