SkepSkep
All episodes
Latent Space: The AI Engineer Podcast — [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Not scored yet17m·2025-12-31·Technology

[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang

Latent Space: The AI Engineer Podcast

From creating SWE-bench in a Princeton basement to shipping CodeClash, SWE-bench Multimodal, and SWE-bench Multilingual, John Yang has spent the last year and a half watching his benchmark become the de facto standard for evaluating AI coding agents—trusted by Cognition (Devin), OpenAI, Anthropic, and every major lab…

Listen in your app
Buzzmeter · cross-platform
Not scored yet

We score an episode from what we can actually measure — its reach and engagement, what listeners say about it, and what it covers. We don't have those signals for this one yet, so it doesn't get a number.

Hive Vote · Skep listeners
New on SkepAwaiting the hive's first votes

Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.

Rate this episode

Add your vote to the Hive — listeners rate every episode after they finish.

Rate this episode
Recent reviews
  • No reviews yet — be the first.

One great episode in your inbox, daily — free.

One email a day, unsubscribe anytime. No spam, ever.