# Shreya Shankar

Najpopularniejsze odcinki podcastów
Zaktualizowano:

Episode 72: Why Agents Solve the Wrong Problem (and What Data Scientists Do Instead)

I often see what I would consider to be b******t evals, especially in data, like write this dumb SQL. Almost every one of these dumb SQL questions that I’ve seen for benchmarks are just so either obviously easy or overwhelmingly adversarial. They just, they don’t feel valuable as a data scientist, …

Episode 13: More questions to improve the return on your AI investments

In this episode, Collin and I continue building our list of questions to help increase the chances of your AI investments delivering a return. We focus in particular on the role non-technical C-level leaders should play in this effort. We use Chip Huyen’s AI Engineering as a practical framework, e…

Data Agents with Shreya Shankar - Weaviate Podcast #135!

Shreya Shankar from UC Berkeley joins the Weaviate Podcast to discuss data agents, the Data Agent Benchmark, and DocETL. The conversation opens with defining what a data agent actually is, not just text-to-SQL over a single table, but an AI system that can reason across dozens of heterogeneous data…

AI Engineering

Topics Covered Teresa's accidental path into AI engineering Her partnership with Vistaly Building "Teresa Bot" How she learned — and how you can too Discovery skills transfer directly to AI engineering On her engineering background (and why it's not what you think) The real takeaway Key …

How to Run Evals in Claude Code with Aparna Dhinakaran, Founder and CPO of Arize

Today’s episode Many of the smartest AI teams I know are running their evals on Arize. Teams at Uber, Booking.com, Pepsi, and others. It’s become one of the most important skills for PMs. I already had on the CEO of Braintrust, Hamel Husain and Shreya Shankar, and Ankit Shukla. Today I’m adding to …

Episode 38 - Meet the Interns

Send us Fan Mail On this week’s episode we’re joined by a remarkable group of college interns who have spent the semester working with us. Throughout their time here, they’ve contributed their talents, learned new skills, and become a valued part of the DECAL team. Let's meet Gabrielle Banks- Commu…

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

Today’s Episode A developer posted this workflow in March, and it is the clearest picture of where PM is heading that I’ve seen all year. Rasty Turek spent the past year building with coding agents, and he mapped how his process changed over that time. He reckons he now spends around 90% of his tim…

The Rise of the AI Scientist

“If you feel that a product is hard to eval, or you feel like, ‘I don’t even know how to eval this,’ it’s a strong smell that your product isn’t good.” — Hamel Husain, on AI product design An AI data agent tells you last quarter’s net revenue. It doesn’t show the metric definition, source tables, f…

Discovery In The AI Era

Christian Idiodi sits down with Teresa Torres, author of "Continuous Discovery Habits" and host of the "Just Now Possible" podcast, for a candid conversation about what 15 months of building AI products has changed in her thinking, and what it hasn't. Teresa explains why she started out refusing to…

Episode 6: Half the Tests Failed and the Build Went Green

`PROMPTFOO_PASS_RATE_THRESHOLD` takes a **percentage, not a fraction** — so writing `0.9` when you meant 90% gives you a **0.9%** threshold, and CI goes green with half your tests failing. Nothing warns you. This episode builds the eval gate that makes swapping a model a check instead of a vibe: on…