

FlashAttention 2: making Transformers 800% faster w/o approximation - with Tri Dao of Together AI
FlashAttention was first published by Tri Dao in May 2022 and it had a deep impact in the large language models space. Most open models you’ve heard of (RedPajama, MPT, LLaMA, Falcon, etc) all leverage it for faster inference. Tri came on the podcast to chat about FlashAttention, the newly released FlashAttention-2,…
We score an episode from what we can actually measure — its reach and engagement, what listeners say about it, and what it covers. We don't have those signals for this one yet, so it doesn't get a number.
Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.
Rate this episode
Add your vote to the Hive — listeners rate every episode after they finish.
- No reviews yet — be the first.
More from Latent Space: The AI Engineer Podcast
See all episodes →
🔬Scaling Past Informal AI - Carina Hong, Axiom Math

🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

Runway’s WorldPrompt and the Engineering of Real-Time Worlds
Similar episodes from other shows
More like Latent Space: The AI Engineer Podcast →
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic

Reiner Pope – The math behind how LLMs are trained and served

EP 58: Every Millisecond Matters: Diffusion LLMs and the Future of Voice AI | Aditya Grover, Inception

Large models on CPUs
One great episode in your inbox, daily — free.
One email a day, unsubscribe anytime. No spam, ever.