
All episodes![Latent Space: The AI Engineer Podcast — [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang](/_next/image?url=https%3A%2F%2Fsubstackcdn.com%2Ffeed%2Fpodcast%2F1084089%2Fpost%2F186610569%2F183dd75aed4203e2c58adcc0da042dcd.jpg&w=3840&q=75)
BUZZING· 75%17m·2025-12-31·Technology
![Latent Space: The AI Engineer Podcast — [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang](/_next/image?url=https%3A%2F%2Fsubstackcdn.com%2Ffeed%2Fpodcast%2F1084089%2Fpost%2F186610569%2F183dd75aed4203e2c58adcc0da042dcd.jpg&w=3840&q=75)
[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Latent Space: The AI Engineer Podcast
From creating SWE-bench in a Princeton basement to shipping CodeClash, SWE-bench Multimodal, and SWE-bench Multilingual, John Yang has spent the last year and a half watching his benchmark become the de facto standard for evaluating AI coding agents—trusted by Cognition (Devin), OpenAI, Anthropic, and every major lab…
Buzzmeter · cross-platform

75%
BUZZING
Active swarm energy. Strong reception, real community buzz.
Hive Vote · Skep listeners
New on SkepAwaiting the hive's first votes
Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.
Rate this episode
Add your vote to the Hive — listeners rate every episode after they finish.
Recent reviews
- No reviews yet — be the first.
More from Latent Space: The AI Engineer Podcast
See all episodes →
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
Latent Space: The AI Engineer Podcast · 57m

🔬Doing Vibe Physics — Alex Lupsasca, OpenAI
Latent Space: The AI Engineer Podcast · 1h 31m

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
Latent Space: The AI Engineer Podcast · 1h 25m
One great episode in your inbox, daily — free.
One email a day, unsubscribe anytime. No spam, ever.