

⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data
Olivia Watkins (Frontier Evals team) and Mia Glaese (VP of Research at OpenAI, leading the Codex, human data, and alignment teams) discuss a new blog post (https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/) arguing that SWE-Bench Verified—long treated as a key “North Star” coding benchmark—has…
We score an episode from what we can actually measure — its reach and engagement, what listeners say about it, and what it covers. We don't have those signals for this one yet, so it doesn't get a number.
Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.
Rate this episode
Add your vote to the Hive — listeners rate every episode after they finish.
- No reviews yet — be the first.
More from Latent Space: The AI Engineer Podcast
See all episodes →
🔬Scaling Past Informal AI - Carina Hong, Axiom Math

🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

Runway’s WorldPrompt and the Engineering of Real-Time Worlds
Similar episodes from other shows
More like Latent Space: The AI Engineer Podcast →
How OpenAI Built Its Coding Agent

Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AI

20VC: Codex vs Claude Code vs Cursor: Who Wins, Who Loses | Will All Coding Be Automated - Do We Need PMs | The Real Bottleneck to AGI | The Three Phases of Agents and What You Need to Know with Alex Embiricos, Head of Codex at OpenAI

Why Everyone Is Obsessed With Claude Code
One great episode in your inbox, daily — free.
One email a day, unsubscribe anytime. No spam, ever.