
![Latent Space: The AI Engineer Podcast — [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang](/_next/image?url=https%3A%2F%2Fsubstackcdn.com%2Ffeed%2Fpodcast%2F1084089%2Fpost%2F186610569%2F183dd75aed4203e2c58adcc0da042dcd.jpg&w=3840&q=75)
[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
From creating SWE-bench in a Princeton basement to shipping CodeClash, SWE-bench Multimodal, and SWE-bench Multilingual, John Yang has spent the last year and a half watching his benchmark become the de facto standard for evaluating AI coding agents—trusted by Cognition (Devin), OpenAI, Anthropic, and every major lab…
We score an episode from what we can actually measure — its reach and engagement, what listeners say about it, and what it covers. We don't have those signals for this one yet, so it doesn't get a number.
Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.
Rate this episode
Add your vote to the Hive — listeners rate every episode after they finish.
- No reviews yet — be the first.
More from Latent Space: The AI Engineer Podcast
See all episodes →
🔬Scaling Past Informal AI - Carina Hong, Axiom Math

🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

Runway’s WorldPrompt and the Engineering of Real-Time Worlds
Similar episodes from other shows
More like Latent Space: The AI Engineer Podcast →
How OpenAI Built Its Coding Agent

Why Everyone Is Obsessed With Claude Code

Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AI

20VC: Codex vs Claude Code vs Cursor: Who Wins, Who Loses | Will All Coding Be Automated - Do We Need PMs | The Real Bottleneck to AGI | The Three Phases of Agents and What You Need to Know with Alex Embiricos, Head of Codex at OpenAI
One great episode in your inbox, daily — free.
One email a day, unsubscribe anytime. No spam, ever.