
All episodes
SCOUT· 65%54m·2023-07-26·Technology

FlashAttention 2: making Transformers 800% faster w/o approximation - with Tri Dao of Together AI
Latent Space: The AI Engineer Podcast
FlashAttention was first published by Tri Dao in May 2022 and it had a deep impact in the large language models space. Most open models you’ve heard of (RedPajama, MPT, LLaMA, Falcon, etc) all leverage it for faster inference. Tri came on the podcast to chat about FlashAttention, the newly released FlashAttention-2,…
Buzzmeter · cross-platform

65%
SCOUT
Scout bees are checking it out — divided opinions, the swarm hasn't committed.
Hive Vote · Skep listeners
New on SkepAwaiting the hive's first votes
Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.
Rate this episode
Add your vote to the Hive — listeners rate every episode after they finish.
Recent reviews
- No reviews yet — be the first.
More from Latent Space: The AI Engineer Podcast
See all episodes →
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
Latent Space: The AI Engineer Podcast · 57m

🔬Doing Vibe Physics — Alex Lupsasca, OpenAI
Latent Space: The AI Engineer Podcast · 1h 31m

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
Latent Space: The AI Engineer Podcast · 1h 25m
One great episode in your inbox, daily — free.
One email a day, unsubscribe anytime. No spam, ever.