SkepSkep
All episodes
Latent Space: The AI Engineer Podcast — [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Buzzing ratingBUZZING· 75%17m·2025-12-31·Technology

[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang

Latent Space: The AI Engineer Podcast

From creating SWE-bench in a Princeton basement to shipping CodeClash, SWE-bench Multimodal, and SWE-bench Multilingual, John Yang has spent the last year and a half watching his benchmark become the de facto standard for evaluating AI coding agents—trusted by Cognition (Devin), OpenAI, Anthropic, and every major lab…

Listen now
Buzzmeter · cross-platform
Buzzing rating
75%
BUZZING

Active swarm energy. Strong reception, real community buzz.

Hive Vote · Skep listeners
New on SkepAwaiting the hive's first votes

Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.

Rate this episode

Add your vote to the Hive — listeners rate every episode after they finish.

Rate this episode
Recent reviews
  • No reviews yet — be the first.

One great episode in your inbox, daily — free.

One email a day, unsubscribe anytime. No spam, ever.