# Hamel Husain

قسمت های برتر پادکست
به روز شده:

Building an AI Sleep Coach: How Rest is Making CBTI Principles Accessible to DIY Sleep Hackers

Guests Martin Siniawski, CEO and co-founder, Rest Ignacio, CTO, Rest You'll hear how they: Discovered the sleep use case from podcast app user behavior (10% of users, but high willingness to pay) Used jobs-to-be-done research to identify "DIY sleep hackers" as an underserved segment Chos…

Communities of Practice

Key Topics: What a “community of practice” really means in modern product work The difference between learning from people vs. learning with people How to find like-minded peers for collaborative learning Building your Personal Learning Network (PLN) Personal knowledge management as a produc…

How to Do AI Evals Step-by-Step with Real Production Data | Tutorial by Hamel Husain and Shreya Shankar

Today’s Episode Everyone’s demoing AI features. Few are shipping them to production reliably. The gap? Evals. Not the theoretical kind. The real-world kind that catches bugs before users do. Hamel Husain and Shreya Shankar train people at OpenAI, Anthropic, Google, and Meta on how to build AI produ…

Evals are the new PRD. Here is the playbook with the CEO of the leader in the space (Ankur Goyal, Founder and CEO, Braintrust)

Today’s episode Most PMs treat evals like a quality gate. Something you run right before shipping, just to check the box. That is backwards. The best AI product teams treat evals as the starting point. They write the eval before the prompt. They iterate on the scoring function before the model. The…

The AI Testing Framework Every Business Needs (But Few Use)

Keith Richman sits down with Hamel Husain, machine learning engineer and founder of Parlance Labs, to demystify AI evaluations (evals). Hamel breaks down why generic AI testing metrics fall short and how businesses can actually measure, debug, and improve their AI applications in the real world. Th…

AI Engineering

Topics Covered Teresa's accidental path into AI engineering Her partnership with Vistaly Building "Teresa Bot" How she learned — and how you can too Discovery skills transfer directly to AI engineering On her engineering background (and why it's not what you think) The real takeaway Key …

How to Run Evals in Claude Code with Aparna Dhinakaran, Founder and CPO of Arize

Today’s episode Many of the smartest AI teams I know are running their evals on Arize. Teams at Uber, Booking.com, Pepsi, and others. It’s become one of the most important skills for PMs. I already had on the CEO of Braintrust, Hamel Husain and Shreya Shankar, and Ankit Shukla. Today I’m adding to …

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

Today’s Episode A developer posted this workflow in March, and it is the clearest picture of where PM is heading that I’ve seen all year. Rasty Turek spent the past year building with coding agents, and he mapped how his process changed over that time. He reckons he now spends around 90% of his tim…

The Rise of the AI Scientist

“If you feel that a product is hard to eval, or you feel like, ‘I don’t even know how to eval this,’ it’s a strong smell that your product isn’t good.” — Hamel Husain, on AI product design An AI data agent tells you last quarter’s net revenue. It doesn’t show the metric definition, source tables, f…

Episode 6: Half the Tests Failed and the Build Went Green

`PROMPTFOO_PASS_RATE_THRESHOLD` takes a **percentage, not a fraction** — so writing `0.9` when you meant 90% gives you a **0.9%** threshold, and CI goes green with half your tests failing. Nothing warns you. This episode builds the eval gate that makes swapping a model a check instead of a vibe: on…