
All episodes
SCOUT· 62%26m·2026-02-23·Technology

⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data
Latent Space: The AI Engineer Podcast
Olivia Watkins (Frontier Evals team) and Mia Glaese (VP of Research at OpenAI, leading the Codex, human data, and alignment teams) discuss a new blog post (https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/) arguing that SWE-Bench Verified—long treated as a key “North Star” coding benchmark—has…
Buzzmeter · cross-platform

62%
SCOUT
Scout bees are checking it out — divided opinions, the swarm hasn't committed.
Hive Vote · Skep listeners
New on SkepAwaiting the hive's first votes
Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.
Rate this episode
Add your vote to the Hive — listeners rate every episode after they finish.
Recent reviews
- No reviews yet — be the first.
More from Latent Space: The AI Engineer Podcast
See all episodes →
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
Latent Space: The AI Engineer Podcast · 57m

🔬Doing Vibe Physics — Alex Lupsasca, OpenAI
Latent Space: The AI Engineer Podcast · 1h 31m

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
Latent Space: The AI Engineer Podcast · 1h 25m
One great episode in your inbox, daily — free.
One email a day, unsubscribe anytime. No spam, ever.