
![Latent Space: The AI Engineer Podcast — [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI](/_next/image?url=https%3A%2F%2Fsubstackcdn.com%2Ffeed%2Fpodcast%2F1084089%2Fpost%2F186610564%2Fe1bc34b9ec6a2ebd16157f848fb57b2d.jpg&w=3840&q=75)
[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
From pre-training data curation to shipping GPT-4o, o1, o3, and now GPT-5 thinking and the shopping model, Josh McGrath has lived through the full arc of OpenAI’s post-training evolution—from the PPO vs DPO debates of 2023 to today’s RLVR era, where the real innovation isn’t optimization methods but data quality,…
We score an episode from what we can actually measure — its reach and engagement, what listeners say about it, and what it covers. We don't have those signals for this one yet, so it doesn't get a number.
Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.
Rate this episode
Add your vote to the Hive — listeners rate every episode after they finish.
- No reviews yet — be the first.
More from Latent Space: The AI Engineer Podcast
See all episodes →
🔬Scaling Past Informal AI - Carina Hong, Axiom Math

🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

Runway’s WorldPrompt and the Engineering of Real-Time Worlds
Similar episodes from other shows
More like Latent Space: The AI Engineer Podcast →
From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki

Richard Sutton – Father of RL thinks LLMs are a dead end

EP 41: The Reward Signal: The Missing Ingredient in Every AI System You’ve Built

Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AI
One great episode in your inbox, daily — free.
One email a day, unsubscribe anytime. No spam, ever.