

Collaboration & evaluation for LLM apps
Small changes in prompts can create large changes in the output behavior of generative AI models. Add to that the confusion around proper evaluation of LLM applications, and you have a recipe for confusion and frustration. Raza and the Humanloop team have been diving into these problems, and, in this episode, Raza…
We score an episode from what we can actually measure — its reach and engagement, what listeners say about it, and what it covers. We don't have those signals for this one yet, so it doesn't get a number.
Buzzmeter says it's buzzing across platforms — the Hive Vote says whether the people who actually listened liked it.
Rate this episode
Add your vote to the Hive — listeners rate every episode after they finish.
- No reviews yet — be the first.
Similar episodes from other shows
More like Practical AI →
Building the Foundation Model Ops Platform — with Raza Habib of Humanloop

Stop being skeptical about AI for development with Charity Majors

Evals, error analysis, and better prompts: A systematic approach to improving your AI products | Hamel Husain (ML engineer)

How to Prompt GPT-5
One great episode in your inbox, daily — free.
One email a day, unsubscribe anytime. No spam, ever.


