1
00:00:00.000 --> 00:00:03.880
I was excited to sit with Hamel Hussein, the founder of parlance labs.
2
00:00:03.880 --> 00:00:08.640
He walks through why you need to be moving fast and experimenting with AI, but also why
3
00:00:08.640 --> 00:00:14.000
you really need to slow down and ask yourself the right questions, talk through harnesses
4
00:00:14.000 --> 00:00:18.960
and guardrails you need to be putting on, and ultimately why it's really important that
5
00:00:18.960 --> 00:00:23.560
we're digging in but also stepping back and asking ourselves what should we be doing
6
00:00:23.560 --> 00:00:24.560
here.
7
00:00:31.000 --> 00:00:38.360
Super excited today to be joined by Hamel Hussein from parlance labs.
8
00:00:38.360 --> 00:00:39.800
He is in the thick of it.
9
00:00:39.800 --> 00:00:45.000
We talk about companies implementing AI, being challenged with AI, how they're going
10
00:00:45.000 --> 00:00:46.560
about solving it.
11
00:00:46.560 --> 00:00:52.880
He is both a leading consulting firm, a well-known speaker and author on the topic and I'm
12
00:00:52.880 --> 00:00:58.600
very excited to have him here to talk through his thoughts on this crazy world we're living
13
00:00:58.600 --> 00:00:59.600
in today.
14
00:00:59.600 --> 00:01:00.600
Welcome.
15
00:01:00.600 --> 00:01:01.600
Yeah, happy to be here.
16
00:01:01.600 --> 00:01:02.600
Thank you.
17
00:01:02.600 --> 00:01:04.560
Before anything else, can you just introduce everybody to parlance labs and give a little
18
00:01:04.560 --> 00:01:05.560
bit of your background?
19
00:01:05.560 --> 00:01:08.240
Yeah, a little background on myself.
20
00:01:08.240 --> 00:01:11.120
I've been a machine learning engineer for over 25 years.
21
00:01:11.120 --> 00:01:15.200
I've worked at a lot of startups and small and bigger ones.
22
00:01:15.200 --> 00:01:19.720
I've worked at Airbnb, GitHub, a bunch of other machine learning startups.
23
00:01:19.720 --> 00:01:23.920
Been doing this consulting, independent consulting, like an independent developer, I've been
24
00:01:23.920 --> 00:01:29.520
an independent developer for about three years, helping people build AI products and the
25
00:01:29.520 --> 00:01:36.000
bottleneck that I kept seeing is people struggling on how to measure and test their AI products
26
00:01:36.000 --> 00:01:42.560
beyond five checks and so I decided to focus on that and so I also do training and education
27
00:01:42.560 --> 00:01:44.320
on the subject to write books.
28
00:01:44.320 --> 00:01:50.960
I'm writing a book with my co-author Shreya Shankar on AI eVals and yeah, I talk a lot
29
00:01:50.960 --> 00:01:51.960
about eVals.
30
00:01:52.040 --> 00:01:56.280
So just for the general population, can you describe what an eVal is?
31
00:01:56.280 --> 00:02:03.320
It's kind of a hyper-loaded term, but basically what it is is how do you do data analysis and
32
00:02:03.320 --> 00:02:09.840
debugging on your AI application in a structured way so that number one, you know what's wrong?
33
00:02:09.840 --> 00:02:16.320
But then also, you know how you should prioritize what you should fix and how to design metrics
34
00:02:16.320 --> 00:02:21.800
to measure things, especially when you have stochastic outputs, like the outputs of
35
00:02:21.800 --> 00:02:27.320
AI are stochastic, they're like text, they don't really have, there's not like a clear,
36
00:02:27.320 --> 00:02:32.200
right or wrong answer, you can deterministically test, so like how do you go about testing an
37
00:02:32.200 --> 00:02:35.560
application like that? That's what eVals is all about.
38
00:02:35.560 --> 00:02:37.880
So obviously that's super critical.
39
00:02:37.880 --> 00:02:44.840
When you're beginning the process of implementing, do you need to define that or is it completely
40
00:02:44.920 --> 00:02:49.720
iterative as your needs and the models evolve?
41
00:02:49.720 --> 00:02:57.000
Yeah, so the way you start with the eVals is data analysis and what you do is you go through a
42
00:02:57.000 --> 00:03:02.600
process called error analysis, which is a kind of data analysis that's kind of blend some qualitative
43
00:03:02.600 --> 00:03:09.000
analysis with some quantitative analysis and you figure out what is broken in your application
44
00:03:09.880 --> 00:03:16.680
and based on that, you can decide what is the right things to measure based on what's actually
45
00:03:16.680 --> 00:03:20.920
happening in your application. So the thing that's different about software testing and AI
46
00:03:20.920 --> 00:03:27.720
testing is when it comes to AI products, there's an infinite surface area of what can go wrong.
47
00:03:27.720 --> 00:03:33.880
And so you need to kind of prioritize like what to measure and how to measure it.
48
00:03:33.960 --> 00:03:40.760
And so that upfront data analysis is key to help you zone in on okay like what you should do.
49
00:03:40.760 --> 00:03:46.040
So it's a little bit like this bottoms-up approach is really important to complement like a top-down
50
00:03:46.040 --> 00:03:51.560
approach of like hey I want to have some things I do want to test or I'm worried about certain
51
00:03:51.560 --> 00:03:57.400
types of failures, but I think people overly focused on the top-down to their detriment and they
52
00:03:57.400 --> 00:04:03.400
get lost in generic metrics. Can you dive into that a little bit? What do you consider a generic
53
00:04:03.400 --> 00:04:11.000
metric in this case? So if you Google eVals or you Google how do I evaluate my AI application?
54
00:04:11.000 --> 00:04:17.320
There's a high likelihood you will stumble upon some vendors or some tools that will promise to
55
00:04:17.320 --> 00:04:24.360
or completely automate the testing of your AI application. And what they'll promise you is they'll
56
00:04:24.360 --> 00:04:30.920
throw up a dashboard that has a bunch of scores like helpfulness score conciseness score toxicity
57
00:04:31.640 --> 00:04:37.640
coherence score you name it. And you'll get a beautiful dashboard with a bunch of metrics on it
58
00:04:37.640 --> 00:04:43.640
usually on a scale of 1 to 5 or 110 or something like that. And what ends up happening is no one
59
00:04:43.640 --> 00:04:49.720
really knows what that means. Also those kind of generic things usually don't correlate with what's
60
00:04:49.720 --> 00:04:55.080
important for you to focus on or to fix. So actually those things are actively very harmful
61
00:04:55.720 --> 00:05:03.720
because they distract you and they make you burn engineering cycles just looking at metrics that
62
00:05:03.720 --> 00:05:09.080
don't matter. And so what you need to do is you need to be very thoughtful about what you're measuring
63
00:05:09.080 --> 00:05:15.080
and make sure that the metrics that you do have a matter. And you kind of have to put your data
64
00:05:15.080 --> 00:05:21.320
signs hat on is kind of just very similar to product analytics. Like you wouldn't take your product
65
00:05:21.400 --> 00:05:26.200
and throw up a dashboard with a bunch of generic metrics. You would think carefully about the way
66
00:05:26.200 --> 00:05:32.200
your metrics are calculated if they make sense for your business. So the kind of same idea applies here.
67
00:05:32.200 --> 00:05:36.840
I have so many follow-up questions. So you're beginning an implementation process
68
00:05:37.640 --> 00:05:43.880
in your mind as an executive or as a leader. You know the outcome you're trying to get at.
69
00:05:43.880 --> 00:05:49.320
In most cases that people I've spoken to they sort of assume that the models have x percent
70
00:05:49.320 --> 00:05:57.080
hallucination and are wrong and are trying to close that gap by training and training and training
71
00:05:57.080 --> 00:06:01.720
and then having some human oversight. Is that a fundamentally flawed way to go about the
72
00:06:01.720 --> 00:06:08.040
implementation because they're not even setting up the right testing plan upfront? So if you go
73
00:06:08.040 --> 00:06:14.040
into trying to test your application with this idea that you need to reduce hallucination
74
00:06:14.840 --> 00:06:21.800
and need to increase healthfulness and you know this is kind of like a generic metrics mindset.
75
00:06:21.800 --> 00:06:26.280
Yeah totally. And what that means is you don't really know what's wrong. You're just kind of
76
00:06:27.080 --> 00:06:32.280
going through some motions and you're going to end up wasting a lot of time and you're going to
77
00:06:32.280 --> 00:06:37.560
get lost. And it's a very appealing thought that hey you can just don't worry about this testing
78
00:06:37.560 --> 00:06:42.440
stuff. Just plug in this framework and we'll calculate a score for you and you'll be fine.
79
00:06:43.160 --> 00:06:46.360
That's absolutely not the case and that's why a lot of people struggle.
80
00:06:47.800 --> 00:06:53.800
You know frankly that's why my business exists. If people weren't getting confused and they weren't
81
00:06:53.800 --> 00:07:00.280
getting led astray then you wouldn't need my help. Unfortunately I think people are trying to look
82
00:07:00.280 --> 00:07:05.480
for the easy button. Unfortunately in this case then easy button doesn't exist. It's not the case
83
00:07:05.480 --> 00:07:10.520
that you have to do everything manually or it has to be painful as you know you can use coding
84
00:07:10.520 --> 00:07:14.920
agents to help you implement the evals and write some of the evals and wire up the plumbing.
85
00:07:15.880 --> 00:07:22.520
But you have to be thoughtful about what you're measuring. And really one of the big parts about
86
00:07:22.520 --> 00:07:29.240
evals is it's not just purely testing. It's a process that where you look at your data
87
00:07:30.280 --> 00:07:35.640
and you understand what good looks like. So most people don't know what good looks like
88
00:07:35.640 --> 00:07:41.400
including myself really including anybody until you look at the outputs and what you need to do is
89
00:07:41.400 --> 00:07:47.720
iterate and kind of specify like oh like this is good and this is bad. It's kind of impossible to
90
00:07:47.720 --> 00:07:53.960
do that without it's like an iterative way of looking at things. That's you know this like and
91
00:07:53.960 --> 00:07:59.240
that's part of the annotations that you might do with evals is like you know and that's what happens
92
00:07:59.240 --> 00:08:04.840
when you're trying to transfer your knowledge to the AI. So it's really hard to build a good AI
93
00:08:04.840 --> 00:08:08.920
product without doing that exercise. You know maybe I maybe because I just came back from human
94
00:08:08.920 --> 00:08:14.920
X where everyone's talking about how agents will do everything. Is there a world in which the
95
00:08:14.920 --> 00:08:21.080
automation itself can do the eval and then automatically fix it or is it inherently you know a human
96
00:08:21.080 --> 00:08:29.560
to model training issue. I think AI is really good at fixing bugs. The deterministic errors in your
97
00:08:29.560 --> 00:08:35.480
product like the code's not working or is doing the wrong thing or is failing in test but AI
98
00:08:35.480 --> 00:08:42.680
cannot read your mind. It doesn't know what you feel like good is. So if a customer is interacting
99
00:08:42.680 --> 00:08:50.680
with your product and it's you know the product is not doing the right thing it may look like it's
100
00:08:50.680 --> 00:08:56.280
helpful on the surface to an AI but if you put your product head on you're like you know what we could
101
00:08:56.280 --> 00:09:03.320
have done better here. We should be able to help the customer in a better way or help the user
102
00:09:03.320 --> 00:09:08.040
in a better way than this or you know what this interaction doesn't make any sense or you know what
103
00:09:08.040 --> 00:09:14.600
we don't have the best tools here. We need to fix the tools we have we need to fix our retrieval
104
00:09:14.600 --> 00:09:20.600
because we're not really bringing in the right sources here and so the AI doesn't know what it
105
00:09:20.600 --> 00:09:28.760
doesn't know like it doesn't have the ability to read your mind and sort of elicit what good looks
106
00:09:28.760 --> 00:09:36.040
like. What AI probably can do is walk you through the process a bit of e-vails and say okay like
107
00:09:36.760 --> 00:09:42.440
and kind of interrogate you the whole bunch. I think that's where the future is is to say like
108
00:09:42.440 --> 00:09:49.000
okay let me guide you through the n10 process of e-vails and let's write it together which is a
109
00:09:49.080 --> 00:09:54.280
you know that's kind of what good e-vail frameworks are trying to do but you still need to have a human
110
00:09:54.280 --> 00:10:01.400
in the loop. When I hear how you describe it you know Claude andthropic you know has this skill builder
111
00:10:01.400 --> 00:10:06.600
for small businesses and individuals they launched where it's very back and forth is this what you
112
00:10:06.600 --> 00:10:10.520
wanted to look like is this what you wanted to feel like. Maybe I don't think it's going to be purely
113
00:10:10.520 --> 00:10:15.880
chat I think it's needs to be a bit different than that because what it involves is looking at
114
00:10:15.880 --> 00:10:21.640
lots of data and you want to render your data in a very domain specific way that's specific to
115
00:10:21.640 --> 00:10:27.160
your products. So for example if you have images you need to render those images you have emails you
116
00:10:27.160 --> 00:10:31.960
need to make it look like an email if it's a chat at least look like a chat there's metadata that's
117
00:10:31.960 --> 00:10:38.600
involved with making decisions you need to render that metadata in a very nice easy to see way
118
00:10:38.600 --> 00:10:44.120
alongside other data so you can make a quick judgment on what is happening with the AI yeah this
119
00:10:44.120 --> 00:10:49.960
probably some software it's a little bit beyond just chatbot but you know I don't think this
120
00:10:49.960 --> 00:10:56.360
anything really special I think like if you zoom out a bit I don't think agents are going to
121
00:10:56.360 --> 00:11:01.960
completely build software either they will build this they will build it to your specification but
122
00:11:01.960 --> 00:11:09.320
what is your specification like you know the more non trivial your software is the more you're
123
00:11:09.320 --> 00:11:16.040
going to have to inject your taste in your specific point of view and if you don't have a specific
124
00:11:16.040 --> 00:11:21.960
taste and point of view then your product is probably shit you know it's not going to be differentiated
125
00:11:21.960 --> 00:11:27.720
or like what are you even building and so I think that in the same way emails is the same way
126
00:11:27.720 --> 00:11:31.400
it's really an extension of that because like what really what you're good getting down to with
127
00:11:31.400 --> 00:11:37.800
evals is it's a process of like eliciting your specification and then measuring against that
128
00:11:37.880 --> 00:11:43.400
so in this world now where you know we like to joke that boards call a CEO they're like what are
129
00:11:43.400 --> 00:11:50.600
we doing for AI the CEO calls it just you know it layers down to let's get something out and let's
130
00:11:50.600 --> 00:11:57.320
show something what goes wrong as part of this process when you think about how you know when
131
00:11:57.320 --> 00:12:02.360
you step in so the first thing that goes wrong is not even evals the first question I ask is are
132
00:12:02.360 --> 00:12:07.960
you using AI and are you using AI in a deep way are you engineers coding with AI
133
00:12:08.760 --> 00:12:14.440
are you building things with AI that's beyond just like copy and paste into chat GPT if the answer
134
00:12:14.440 --> 00:12:20.040
is no then I then you're not going to be successful building AI because you want to have a good
135
00:12:20.040 --> 00:12:26.440
intuition on what is possible and what's not you're going to have really bad specifications
136
00:12:27.400 --> 00:12:33.160
you know you won't kind of zone in on what is a good idea and what's not you want to have a good
137
00:12:33.160 --> 00:12:39.480
mental model so I think that's the first failure point that people face and the second failure
138
00:12:40.120 --> 00:12:46.920
mode that I see a lot is reaching for complexity too fast so okay you want to build an AI product
139
00:12:47.960 --> 00:12:54.280
don't off the bat go for the most complex architecture and setup like don't go for don't like
140
00:12:54.280 --> 00:12:59.800
on day one reach where the most complicated orchestration framework and a graph database and a
141
00:12:59.800 --> 00:13:05.720
multi agent thing start simple and build your way build your way up incrementally I see a lot of
142
00:13:05.720 --> 00:13:11.800
teams that just go straight to the most complicated architecture and then you know they can't really
143
00:13:11.800 --> 00:13:17.960
reason about what is happening and they don't have a good mental model but I don't think that
144
00:13:17.960 --> 00:13:24.760
last point is not necessarily unique to AI that's always been a problem in software to some degree
145
00:13:25.480 --> 00:13:31.880
an AI it's a little bit more acute because you can swallow complexity a lot faster because of AI
146
00:13:31.880 --> 00:13:37.640
agents so how do you balance the sort of need to go slow with the reality of
147
00:13:38.360 --> 00:13:43.160
phomo and pressure from all different parts of the org like if you're
148
00:13:43.960 --> 00:13:49.240
you know director or VP what do you what guidance do you have for how they can manage up in these
149
00:13:49.240 --> 00:13:55.160
cases it's not necessarily going slow I wouldn't go to I wouldn't go slow I would just be very
150
00:13:55.160 --> 00:14:02.040
deliberate I think it's really important to be building with AI be trying lots of ideas now there's
151
00:14:02.040 --> 00:14:09.720
a double edge sort AI is AI allows you to build the wrong thing faster and build many wrong things
152
00:14:10.120 --> 00:14:17.320
faster and with more skill but also lets you build the right thing faster it also allows you to
153
00:14:17.320 --> 00:14:22.200
experiment with a lot of ideas it you know it sounds good on paper like okay you can experiment
154
00:14:22.200 --> 00:14:29.080
with a lot of ideas but if most of your ideas are bad then you're just going to sink in the bad
155
00:14:29.080 --> 00:14:33.880
ideas still you're going to turn through lots and lots of bad ideas so you need to be a little bit
156
00:14:33.880 --> 00:14:40.920
deliberate it's good to have good product sense in some good grounding in taste and what
157
00:14:40.920 --> 00:14:46.280
you should build so that at least you know you're moving in some productive direction because
158
00:14:46.280 --> 00:14:53.320
again like agents can be very distracting on their own it can be fun just to build I think it's
159
00:14:53.320 --> 00:14:59.320
not really moving fast is being deliberate in my mind and then trying to like do experiments
160
00:14:59.960 --> 00:15:04.440
and sort of double down on what's working and be very methodical about that that's
161
00:15:04.440 --> 00:15:13.240
for super interesting about AI being able to build bad things quickly in the last two years people
162
00:15:13.240 --> 00:15:17.960
have gone from I'm very actively managing my AI you know too I'm going to let them write this
163
00:15:17.960 --> 00:15:21.640
I'm sure you'd say you get these emails where clearly nobody read it over they're just
164
00:15:22.280 --> 00:15:27.800
somebody put it in AI and sent it when you go on an organizational level how big a risk is that
165
00:15:27.800 --> 00:15:33.160
do you think where there's just this lack of ability or lack of desire because our brains are
166
00:15:33.160 --> 00:15:39.320
getting trained to not push back and not actually ask the questions that would be to real deliberate
167
00:15:39.320 --> 00:15:47.320
implementation I think it amplifies who you are so if you're a mediocre person who is okay with
168
00:15:47.320 --> 00:15:52.680
writing slop for example um it's just going to amplify you and you're not you're going to be
169
00:15:52.680 --> 00:15:57.480
more susceptible to say you know what you know I don't mind even if you're not a mediocre person
170
00:15:57.480 --> 00:16:03.640
you can still everyone gets lazy and gets tired and can get trapped in this but I think it does amplify
171
00:16:03.640 --> 00:16:09.000
your natural tendencies you know just statistically there's more average people than above average
172
00:16:09.000 --> 00:16:15.480
people and I think it you will see it like this amplification of slop if everyone adopts it
173
00:16:15.480 --> 00:16:20.040
because it takes away all the friction of doing the work in the first place usually before there
174
00:16:20.040 --> 00:16:27.640
was some friction between you and you know pushing this information now there's no friction
175
00:16:27.640 --> 00:16:33.160
so that's the unfortunate kind of reality of it and it's really the same thing about
176
00:16:33.720 --> 00:16:38.120
it's not even just like yeah it's not just about building the wrong things it's saying the wrong
177
00:16:38.120 --> 00:16:43.000
things it's uh doing the wrong things it's doing it you know it just amplifies you it could amplify
178
00:16:43.000 --> 00:16:47.240
you in the wrong direction like very fast you know I'm kind of hopeful like maybe we will
179
00:16:48.040 --> 00:16:53.640
find ourselves we'll find a way to deal with it we'll figure it out and I do think that people do
180
00:16:53.640 --> 00:16:59.560
notice the difference if people can smell AI you know they they understand when and you know I
181
00:16:59.560 --> 00:17:05.320
I think people they'll will adapt in some ways maybe not completely but there'll be some adaptation
182
00:17:05.320 --> 00:17:10.280
we won't be completely lost so you've just described a chain of events where things can go really
183
00:17:11.000 --> 00:17:17.560
well or mediocre or really right when you walk in and you start an engagement like what's the most
184
00:17:17.560 --> 00:17:25.480
common problem that you run into one is swallowing too much complexity which I already discussed
185
00:17:25.480 --> 00:17:33.480
another one is there's a mandate from on top the hey we need to use AI and people take a
186
00:17:34.200 --> 00:17:39.800
existing product surface area and just slap a chatbot on it and say hey we got AI you know that
187
00:17:39.800 --> 00:17:46.680
is a very mediocre experience that kind of doesn't move the needle and so I think you to yeah to
188
00:17:46.680 --> 00:17:50.920
build AI you have to you should be pretty you should try to be thoughtful of how you can actually
189
00:17:50.920 --> 00:17:58.040
help the user like accelerate their workflow and do things faster for example instead of
190
00:17:58.760 --> 00:18:06.200
putting a chatbot on your product consider putting an mcp on your product or consider exposing
191
00:18:06.200 --> 00:18:12.440
apis for your product that is likely to be way more helpful to a lot more people than just putting
192
00:18:12.440 --> 00:18:16.680
in a chatbot on people are reluctant to expose apis and mcp's your product because I feel like
193
00:18:16.680 --> 00:18:20.440
they're getting distributed intermediate by the AI now you're not even opening the application
194
00:18:20.440 --> 00:18:25.400
potentially you're just chatting with cloud code or chat gbt or whatever to interact with your
195
00:18:25.400 --> 00:18:30.040
product but at the same time you know I don't think people are going to be clicking on menus
196
00:18:30.040 --> 00:18:35.240
and clicking buttons for that much longer I think it requires a little bit more and this comes
197
00:18:35.240 --> 00:18:41.480
back to are you using AI because if you're not using AI deeply then that's where these problems
198
00:18:41.480 --> 00:18:47.320
stem from because you don't have a good mental model of where the puck is moving that's
199
00:18:47.320 --> 00:18:54.680
super interesting can you just describe like a use case or case study chatbot versus mcp because
200
00:18:54.680 --> 00:19:01.720
we just internally at my company we switched providers because they didn't have an mcp for cloud
201
00:19:01.800 --> 00:19:08.600
we just canceled our our whole account but I think that nuance is really lost and I'd love
202
00:19:08.600 --> 00:19:12.440
to hear a case study or examples that you might have yeah I mean like a really big example that
203
00:19:12.440 --> 00:19:20.360
I think everyone can relate to is google so recently up until very recently it was difficult to
204
00:19:20.360 --> 00:19:26.440
interact programmatically with AI with google works like google workspace like google docs google
205
00:19:26.440 --> 00:19:34.040
sheets they had integrated chat with Gemini but it was very poor like to do anything you know to
206
00:19:34.040 --> 00:19:40.440
like modify spreadsheet to modify your calendar to look through your email it wasn't really great
207
00:19:40.440 --> 00:19:47.640
it was very frustrating honestly and then eventually like they exposed the google team exposed like
208
00:19:47.640 --> 00:19:54.360
a google workspace CLI there's some other third party stuff to make these more agent friendly
209
00:19:54.360 --> 00:19:58.040
and that makes a huge difference is what everyone uses now you know like if you're using an open
210
00:19:58.040 --> 00:20:03.560
claw for example you know it's using if you want to wire it up this using that so you know that's
211
00:20:04.360 --> 00:20:10.840
that's really huge another example is that may hit home is if you just slap a chatbot on your
212
00:20:10.840 --> 00:20:15.800
product you have to think really carefully about the interface for example if your chatbot is
213
00:20:15.800 --> 00:20:22.520
scheduling a meeting you don't want to go back and forth on just text like hey here at the meeting
214
00:20:22.600 --> 00:20:27.640
times bullet points meaning what mean time one time two type three and then you have to go oh yeah
215
00:20:27.640 --> 00:20:33.480
the four o'clock works that can be very brittle you should expose an interface like here's a widget
216
00:20:33.480 --> 00:20:40.840
select one of these times great select time great uh confirm that makes a lot more sense but people
217
00:20:40.840 --> 00:20:45.400
need to think a bit more holistically about like hey like what is the right interface what's the
218
00:20:45.400 --> 00:20:52.360
right workflow here where the user can give visual feedback and be confident that it works
219
00:20:52.360 --> 00:20:56.280
and that you can also avoid bugs because if you're just trying to do everything with text like
220
00:20:56.280 --> 00:21:01.640
just pure chat then I don't know if the tool fired correctly you know you don't know if the
221
00:21:01.640 --> 00:21:07.080
tool fired correctly either maybe it did maybe it didn't I want visual confirmation in you as a
222
00:21:08.120 --> 00:21:12.200
person serving that you wanted to work more deterministically so so not everything needs to go
223
00:21:12.200 --> 00:21:18.760
through an lm right so it's like how do you have the right approach in product thinking
224
00:21:18.760 --> 00:21:24.920
it's part of the e-vail process putting in guard rails to make sure we stay within certain harnesses
225
00:21:24.920 --> 00:21:31.400
or frameworks or where does that come into play okay so the question is guard rails okay so first
226
00:21:31.400 --> 00:21:36.760
let's talk about what a guard rail is so guard rails are very specific kind of e-vail that sits
227
00:21:36.760 --> 00:21:42.680
in between the request response path and blocks a certain output from being shown to a user for
228
00:21:42.680 --> 00:21:49.080
example you don't want your AI to talk about competitors or maybe use profanity or something like
229
00:21:49.080 --> 00:21:55.160
that so you just want to block it by blocking it means either you just prevent the AI from saying
230
00:21:55.160 --> 00:22:00.760
anything or you just make the AI say something generic like I can't help with that and we've all
231
00:22:00.760 --> 00:22:06.520
seen that so a lot of people think of guard rails they also think of off the shelf stuff
232
00:22:06.520 --> 00:22:10.680
they're like oh I'm going to go to this framework and I'm going to get the profanity guard rail
233
00:22:10.680 --> 00:22:16.200
or I'm going to get the unhelpfulness guard rail and you have to be really careful because if you
234
00:22:16.200 --> 00:22:22.680
use a guard rail off the shelf you're just using someone else's prompt and it's very likely that
235
00:22:22.680 --> 00:22:29.320
that someone else's prompt is not going to work well for you I've actually looked at these prompts
236
00:22:29.320 --> 00:22:35.400
quite a bit those prompts have examples that are very specific to certain domains oftentimes it's
237
00:22:35.400 --> 00:22:41.080
like a shopping domain or a travel domain or something like that but if you're doing something
238
00:22:41.080 --> 00:22:45.560
in legal you don't want a travel domain example in your guard rail prompt it's not going to it's
239
00:22:45.560 --> 00:22:50.360
not the it's not the it's not the best fit and so you have to go through like a similar process of
240
00:22:50.360 --> 00:22:56.360
evals and decide okay like which failures you want to prevent against and you want to prioritize
241
00:22:56.360 --> 00:23:01.880
failures are actually happening or failures you can simulate if you can't reasonably simulate the
242
00:23:01.880 --> 00:23:07.000
failure or you don't see it actually happening then it's kind of a lower priority but yeah guard
243
00:23:07.000 --> 00:23:11.800
rails is a special type of e-vail usually you want it to be fast because it's blocking the
244
00:23:11.800 --> 00:23:17.800
response or it's in the path but it's really very similar to other kinds of e-vails it just has
245
00:23:17.800 --> 00:23:24.200
that these additional characteristics so one layer of complexity on top of everything else is
246
00:23:24.200 --> 00:23:32.120
most companies are trying to figure out the stack that they're using which then manages for cost
247
00:23:32.120 --> 00:23:39.640
for model usage and as you're trying to create a testing an e-vail plan but also trying to manage
248
00:23:39.640 --> 00:23:44.520
for a changing world where you're switching models how do you recommend people think through
249
00:23:45.240 --> 00:23:50.920
that process yeah so the best thing to do is to use the most powerful model you can to start with
250
00:23:51.480 --> 00:23:59.400
to make your life easy just pick one to start yeah pick one to start if you can use like you know
251
00:23:59.400 --> 00:24:04.120
the best open AI model or the best anthropic models or something like that something easy
252
00:24:04.120 --> 00:24:08.440
hopefully it's something you're familiar with so again something that you're using to build stuff
253
00:24:08.440 --> 00:24:13.560
already you're already using in your coding agents so use that because you might already
254
00:24:13.560 --> 00:24:18.920
that is important to like benefit from that intuition you already have and then what you should
255
00:24:18.920 --> 00:24:24.920
do is build an e-vail harness with the metrics that matter then you can try other models you can
256
00:24:24.920 --> 00:24:30.120
try backing off the smaller models different models and you can sort of see okay what the trade-offs
257
00:24:30.120 --> 00:24:37.240
are between latency and cost and performance and sort of reason about it from there the reason why
258
00:24:37.240 --> 00:24:43.800
it's useful to start with the most powerful model and then back off is it lets you to build more
259
00:24:43.800 --> 00:24:49.640
simply to begin with you don't have to yeah you can try to see like what is works in the
260
00:24:49.640 --> 00:24:55.320
sort of most easiest case of trying to get the AI to do something and then you can back it off
261
00:24:55.320 --> 00:25:00.360
and see what happens and then you can more more reason will be like reason about the trade-offs
262
00:25:00.360 --> 00:25:05.720
versus if you start try to start in the other direction sometimes like you have to build some more
263
00:25:05.720 --> 00:25:10.920
complexity to get to the same performance and you don't know it's not really yeah like and then
264
00:25:11.320 --> 00:25:15.560
you kind of stuck with the complexity which really and just as I process everything that you're
265
00:25:15.560 --> 00:25:21.160
saying in this world most executives that I talk to like oh we're going to go in and we're going to
266
00:25:21.160 --> 00:25:28.040
cut all these costs and it's going to be fantastic but what I'm hearing you say is you start with
267
00:25:28.040 --> 00:25:32.600
a process there's all you need a layer of people in order to really maximize it you you're going
268
00:25:32.600 --> 00:25:38.760
to continually need watching evaluation building experimentation which would sort of speak against
269
00:25:38.760 --> 00:25:43.080
any real cost efficiencies in some parts of the org and you wrote about this you just read about
270
00:25:43.080 --> 00:25:49.000
data science actually it was this you know the view that data science is going away but actually what
271
00:25:49.000 --> 00:25:54.040
you're describing is in a world where data science becomes more important than ever I do think there's
272
00:25:54.040 --> 00:25:59.880
going to be a lot of cost reduction for sure I don't think that roles are going to be wholesale
273
00:25:59.880 --> 00:26:04.360
eliminated completely like there'll still be a software engineer there'll still be like a product
274
00:26:04.360 --> 00:26:11.640
manager still be a data scientist maybe those roles will collapse into fewer people so people will
275
00:26:11.640 --> 00:26:19.480
wear more hats than they have before they're able to span more surface area so I do think that cost
276
00:26:19.480 --> 00:26:25.400
will decrease you know like you don't need as many data scientists as you did pre-AI but it's still
277
00:26:25.400 --> 00:26:31.480
good to have the skill somewhere in your organization even with AI because like yeah what a data scientist
278
00:26:32.360 --> 00:26:37.320
is doing is they're asking questions and the ability as you know the ability to ask the right
279
00:26:37.320 --> 00:26:44.520
questions is directly proportional to the quality of output you get with your AI I do see that
280
00:26:44.520 --> 00:26:50.360
there'll be a drastic cost reduction now it gets us left to be seen like what direction you take
281
00:26:50.360 --> 00:26:58.840
as a company with AI like do you just try to hold the line and do more with do the same with less
282
00:26:59.480 --> 00:27:06.360
or you try to do more a lot more with the same people I'm currently I mean I'm kind of in with the
283
00:27:06.360 --> 00:27:12.760
view of like hey you have to do more otherwise you're going to die in a lot of cases it might
284
00:27:12.760 --> 00:27:16.760
might be certain cases where yeah you could just say the same and just do it like it makes
285
00:27:16.760 --> 00:27:24.120
sense to just be more cost efficient but for a lot of tech related things yeah I think growth
286
00:27:25.000 --> 00:27:30.280
a lot of the model companies the OpenAI's the clot I think just announced last week they're
287
00:27:30.920 --> 00:27:35.960
trying to set up their own implementation teams to help move directly into I'm going to work
288
00:27:35.960 --> 00:27:42.600
with private equity to go ahead and change how you think about AI what's your advice for
289
00:27:42.600 --> 00:27:47.880
going directly with the model company versus using an implementation firm versus using some hybrid
290
00:27:47.880 --> 00:27:52.760
and how should people think about the process of who's helping them think through these changes
291
00:27:52.760 --> 00:27:57.480
I think the question is about when should you rely on a third party or get the help of a third
292
00:27:57.480 --> 00:28:02.920
party when building your AI applications and fundamentally it's it's very similar to should
293
00:28:02.920 --> 00:28:09.720
you engage with a consulting company to build your AI and that in turn is kind of leads to the
294
00:28:09.720 --> 00:28:18.280
question of is AI competency in building this AI product within your core competency like is
295
00:28:18.280 --> 00:28:22.760
this something that's important to your business you know so like if it's a software business
296
00:28:23.480 --> 00:28:29.000
it's hard to see how building an AI product is not within your core competency you know if you're
297
00:28:29.000 --> 00:28:34.040
exposing a product of any kind that a software yeah it's really it's really hard to see how
298
00:28:34.040 --> 00:28:40.360
wouldn't be and I haven't yet encountered a company where it isn't in that where I feel like it
299
00:28:40.360 --> 00:28:46.440
isn't in their core competency the reason is is because if you're doing knowledge work fundamentally
300
00:28:46.440 --> 00:28:51.080
AI should be within your core competency so it's very difficult for me to think of a situation
301
00:28:51.080 --> 00:28:57.240
where AI so you should be really careful about getting a consulting company to help you implement AI
302
00:28:57.240 --> 00:29:03.560
now where I think you should maybe you can maybe get help there is the upskill you should absolutely
303
00:29:03.560 --> 00:29:08.360
upskill your team so they're not dependent on a third party being dependent on a third party
304
00:29:08.440 --> 00:29:16.360
expected for something as important as AI is very risky in my mind and also in kind of an
305
00:29:16.360 --> 00:29:21.320
anti pattern because what are you going to do when those external parties leave you need to make
306
00:29:21.320 --> 00:29:27.960
sure that you are like you treat it as a training exercise like a deliberate training exercise
307
00:29:27.960 --> 00:29:33.480
and you're not using it as a crutch and when you go in and talk to companies do you have a framework
308
00:29:33.560 --> 00:29:40.440
or sort of a roadmap that you show them that takes them from point A to point B and then allows
309
00:29:40.440 --> 00:29:46.200
them to watch wins repeat on their own basically I put companies through sort of a bootcamp where
310
00:29:46.760 --> 00:29:52.200
I put them through this e-vals course first I make sure that they're in the right place to do e-vals
311
00:29:52.200 --> 00:29:58.120
meaning they're using AI internally they already have an AI product and they just they're at a
312
00:29:58.120 --> 00:30:02.760
place where they want to make the AI product work really well and they want to know how do we test it
313
00:30:02.760 --> 00:30:08.920
how do we measure it in a way that makes sense so once we get past that what I do is I have a
314
00:30:08.920 --> 00:30:15.320
course that teaches people e-vals so I give my clients access to the course and at the same time
315
00:30:15.320 --> 00:30:21.480
what I do is we go through their data and we debug their product and we do this whole end-to-end
316
00:30:21.480 --> 00:30:27.960
e-vals process but we do it together you know we basically pair program and build the whole e-vals
317
00:30:28.920 --> 00:30:33.960
and they're basically doing it I am telling them how to do it and how to get them unstuck
318
00:30:34.680 --> 00:30:40.120
according to their data but by the end of it they don't need me because I transfer all my knowledge
319
00:30:40.120 --> 00:30:44.200
giving them all the tools they need and we've gone through some we've done a bunch of practice
320
00:30:45.480 --> 00:30:50.040
so it's kind of like going to driving school I would say yeah it's like hey like you might
321
00:30:50.040 --> 00:30:54.040
do a little bit of study of the rules but then I'm gonna get in the car and we're gonna I'm
322
00:30:54.040 --> 00:30:59.080
just gonna drive with you until you can drive and then when we're done you can just drive on
323
00:30:59.080 --> 00:31:04.520
your own on your own what's interesting is this argument I keep getting into which is I'll have
324
00:31:04.520 --> 00:31:10.680
this discussion with someone and then they'll go to X or Twitter and read a quote from somebody
325
00:31:10.680 --> 00:31:15.240
at OpenAI that said they wrote 10 million lines of code all by themselves and they'll skip your
326
00:31:15.240 --> 00:31:19.800
example and say I don't need any of that I'm just waiting for self-driving car and how do you
327
00:31:20.680 --> 00:31:24.920
convince people that that might not be the right approach we can get us like analogies all day
328
00:31:24.920 --> 00:31:28.840
not it's just a humor in the analogy yeah yeah you can have self-driving car but you gotta tell
329
00:31:28.840 --> 00:31:36.520
where you want to go right so so that's really what this is like you know fundamentally you know
330
00:31:37.320 --> 00:31:45.240
it's really really difficult to tell to know like what to tell an AI to do and keep it on target
331
00:31:45.240 --> 00:31:52.440
keep it on task without a harness without an environment where it can test itself
332
00:31:52.440 --> 00:31:57.160
and it can get feedback on whether or not is doing the right thing and keep itself in check
333
00:31:58.440 --> 00:32:03.080
and so with these million dollar or it said these million lines of code things blog posts they're all
334
00:32:03.080 --> 00:32:09.320
backed by this like harness engineering idea that's the only way you can do that and inside this
335
00:32:09.320 --> 00:32:16.520
harness engineering are metrics and logs and traces and observability and that's keeping the AI
336
00:32:16.520 --> 00:32:21.960
on track so if you want to have a really good harness you have to have e-vows so e-vows is a huge
337
00:32:21.960 --> 00:32:29.080
part of the harness it's almost all of the harness so interesting because in the world of social media
338
00:32:29.080 --> 00:32:35.240
those nuances get lost yeah definitely yeah you know even I feel like I'm informed enough to know
339
00:32:35.320 --> 00:32:40.280
and I read that and I'm like oh man it's the easiest thing in the world but it clearly isn't
340
00:32:40.920 --> 00:32:45.960
I'm implementing AI today I'm deep in I'm experimenting I'm kind of living what you described
341
00:32:46.520 --> 00:32:50.920
should I just expect that I'm gonna make a ton of mistakes over time and some portion
342
00:32:50.920 --> 00:32:58.600
is gonna be Monday's gonna be let on fire and that's part of adapting my org to this new world
343
00:32:58.600 --> 00:33:02.280
yeah I think so I mean I think it's hard to do anything without making some mistakes
344
00:33:02.920 --> 00:33:08.120
I make mistakes all the time and I think it's uh you have to be willing to tolerate some mistakes
345
00:33:08.120 --> 00:33:14.040
if you're gonna experiment and you're gonna especially with AI that's moving so fast I think the
346
00:33:14.040 --> 00:33:22.200
main idea is to have a very experimental mindset and to encourage people to be using AI at the
347
00:33:22.200 --> 00:33:28.760
frontier as much as possible so they can have those mental models so that you can make less mistakes
348
00:33:28.760 --> 00:33:34.520
because you know the biggest mistake is building the wrong thing you can eval you can eval the
349
00:33:34.520 --> 00:33:40.120
wrong thing all the way to hell but it's not gonna help if you're still building the wrong thing
350
00:33:40.120 --> 00:33:46.200
like it doesn't matter yeah it's just really useful to to experiment a lot what's it when I hear
351
00:33:46.200 --> 00:33:51.720
you say that most companies that aren't in the tech world don't have a quote unquote R&D budget
352
00:33:51.720 --> 00:33:56.920
but it's almost like maybe every company needs an R&D budget now yeah I mean and so what I mean by
353
00:33:56.920 --> 00:34:04.040
experimentation is not this idea of this expensive lab with these like super computers or beakers
354
00:34:04.040 --> 00:34:10.360
and flats and all this you know like chemicals and like it's like oh this like fancy R&D no I'm
355
00:34:10.360 --> 00:34:15.000
talking about you sitting at your laptop using cloud code trying to build some stuff you know and
356
00:34:15.000 --> 00:34:22.280
so I think you don't need an R&D budget maybe need a token budget but not really I mean you can
357
00:34:22.280 --> 00:34:27.480
have a hundred dollar plan and get pretty far and you can get a lot of intuition by using these
358
00:34:27.480 --> 00:34:33.160
things very deliberately to solve problems just shifting gears for a second what's your personal
359
00:34:33.160 --> 00:34:41.080
tech stack on a curiosity personal tech stack so yeah I use cloud code a lot I use codex as well
360
00:34:41.080 --> 00:34:47.480
at the same time I'm constantly experimenting with stuff I use cursor a little bit I don't have I
361
00:34:47.480 --> 00:34:52.280
used to have opinionated tech stack I used to you know I used to be a Python developer just like
362
00:34:52.280 --> 00:34:59.320
only write Python mostly data science stuff but now I'm like all over the place I don't really care
363
00:34:59.320 --> 00:35:04.920
I use the tools that the AI wants to use so I do that I have a bunch of skills I have a bunch of
364
00:35:04.920 --> 00:35:12.760
tools you know try to slowly build my tech stack I experiment with stuff constantly I'm experimenting
365
00:35:12.760 --> 00:35:18.280
with open claw kind of counter to let's say the narrative I haven't found it to be super useful
366
00:35:18.280 --> 00:35:25.320
yet in my company I have it's me and I have a bunch of other kind of talented developers that are
367
00:35:25.320 --> 00:35:32.680
either my friends that are in my Slack channel that work at other companies and honestly we spend
368
00:35:32.680 --> 00:35:38.120
most of our time with open claw we spend spending most of our time improving open claw are like fixing
369
00:35:38.120 --> 00:35:42.600
open claw are building tools for open claw are building tools that can build tools for open claw
370
00:35:42.600 --> 00:35:47.480
and when we step back we're like wait a second this is all fun and amusing but like we're not
371
00:35:47.480 --> 00:35:54.680
actually doing anything and so it is a realization that I came to and it's like you can do like
372
00:35:54.680 --> 00:36:01.640
the same kind of you know recurring schedule scheduling tasks ambient tasks now through like
373
00:36:02.360 --> 00:36:07.560
clawed you know they have like dispatching clawed co-working scheduled tasks and so you can do a lot
374
00:36:07.560 --> 00:36:12.680
of stuff there so I don't know it's it's but you know despite that we still experiment we're like
375
00:36:12.680 --> 00:36:19.080
okay like is there is there like a place where it is interesting or useful you know like always try
376
00:36:19.080 --> 00:36:27.320
to find out just by using it constantly so yeah I was just using stuff I had the quintessential
377
00:36:27.320 --> 00:36:33.400
open claw experience where I set it to manage a marketing campaign and it decided that it was
378
00:36:33.400 --> 00:36:38.280
working so well it up the budget by like a hundred percent hundred times and I were spending five
379
00:36:38.280 --> 00:36:44.520
or eight or five hundred okay well was it working well at five it was but you know the harness
380
00:36:44.520 --> 00:36:49.400
going back to your your Eval would have been okay once it's at five take it to seven then take it
381
00:36:49.400 --> 00:36:57.880
to ten and I just was like you go you go wild with your you know evaluation and it jumped it up
382
00:36:57.880 --> 00:37:03.800
to five hundred and I figured out after a day but it was a good it was a good lesson and
383
00:37:05.080 --> 00:37:10.040
okay what is it doing now is it is it still running it is but now it's like you know I've
384
00:37:10.040 --> 00:37:13.720
gone the other way because I don't pay that much attention to it so now it like goes up a dollar
385
00:37:13.720 --> 00:37:18.600
day which is almost boring so I need to there's there's a middle ground that I somebody could help
386
00:37:18.600 --> 00:37:23.080
me experiment with or I could spend more time on but I had I set up the project intelligently
387
00:37:23.080 --> 00:37:28.920
before and this all goes back to what you described and and done my own evals versus just jumping in
388
00:37:28.920 --> 00:37:33.960
and you know guiding I would have been much more successful yeah what I have noticed is okay so
389
00:37:33.960 --> 00:37:41.880
even though AI allows you to do things faster and maybe span greater surface area to do one thing
390
00:37:41.960 --> 00:37:49.560
like super well you still have to focus a lot on it so in that way nothing has changed so for
391
00:37:49.560 --> 00:37:56.520
example if I want an AI to be really good at video editing I probably need to really focus on
392
00:37:56.520 --> 00:38:02.840
video editing for a couple of months it may be exclusively and make it like really really good
393
00:38:03.480 --> 00:38:08.520
as the good as possible if I just do use like someone else's prompt someone else's skill or
394
00:38:09.160 --> 00:38:14.120
tried to like vibe it out in a day I'm gonna get some like very the make made mediocre
395
00:38:15.000 --> 00:38:20.200
and it's gonna still be exciting because it's gonna be way more than what I would do normally
396
00:38:20.920 --> 00:38:28.200
but it's not gonna be like oh this is amazing it's not gonna be a par with maybe like a really
397
00:38:28.200 --> 00:38:35.000
talented person per se but you know it's still useful because it's like free but so it's very
398
00:38:35.000 --> 00:38:39.640
interesting like yeah it's funny because I just wrote this piece you know why AI isn't great
399
00:38:39.640 --> 00:38:45.560
for the ADHD population because and you alluded to this earlier it magnifies whatever you know
400
00:38:45.560 --> 00:38:50.520
greatness or weakness you have you know as I talk to my most ADD friends they're building 25
401
00:38:50.520 --> 00:38:56.760
things and if I talk to my most focused friends they're nitpicking the product to the point of who
402
00:38:56.760 --> 00:39:01.560
cares and somewhere in the middle is the right answer but you have to manage your own personality
403
00:39:01.560 --> 00:39:07.720
in the world of AI well look thank you so much you know parlance lab sounds like it's doing
404
00:39:08.360 --> 00:39:13.320
important work in terms of helping us all make it useful which is the whole you know point that
405
00:39:13.320 --> 00:39:18.280
I'm interested is how do we actually get this from it's a company mandate to it's actually helping
406
00:39:18.280 --> 00:39:33.880
my company grow and thank you so much for spending time with us yeah thank you