00:00:00.000 --> 00:00:03.180
AI demos are easy. That is kind of the
problem.
00:00:03.379 --> 00:00:07.120
You can wire an LLM into a tool, give it a
few
00:00:07.120 --> 00:00:11.119
prompts, let it call an API, and suddenly
it
00:00:11.119 --> 00:00:15.019
looks like magic. It writes queries, it
explains
00:00:15.019 --> 00:00:19.019
dashboards, it investigates issues. It
might
00:00:19.019 --> 00:00:23.339
even suggest a fix. And in a demo, that
feels
00:00:23.339 --> 00:00:27.440
great. But production does not care about
demos.
00:00:27.800 --> 00:00:31.440
Production cares if the answer was right.
If the
00:00:31.440 --> 00:00:35.140
agent used the right tool. If it leaked
data. If
00:00:35.140 --> 00:00:38.899
it got confused by logs. If it made things
worse
00:00:38.899 --> 00:00:42.780
after a deploy. If the cost graph quietly
turned
00:00:42.780 --> 00:00:45.659
into a crime scene. And that is where this
gets
00:00:45.659 --> 00:00:48.939
interesting because the hard part is not
proving
00:00:48.939 --> 00:00:51.960
AI can do something impressive in a
controlled
00:00:51.960 --> 00:00:54.979
demo. The hard part is figuring out what
happens
00:00:54.979 --> 00:00:58.170
after people start depending on it. When
it is
00:00:58.170 --> 00:01:02.270
connected to real tools, real telemetry,
real
00:01:02.270 --> 00:01:05.950
production systems, real costs, real
failure
00:01:05.950 --> 00:01:09.790
modes, at that point, AI is not just a
feature
00:01:09.790 --> 00:01:13.510
anymore. It is another thing you have to
operate.
00:01:13.909 --> 00:01:16.790
I'm Brian Teller from Teller's Tech, and
this
00:01:16.790 --> 00:01:37.170
is Ship It Weekly. Welcome back to Ship It
Weekly,
00:01:37.329 --> 00:01:40.090
where I filter the noise and focus on what
actually
00:01:40.090 --> 00:01:42.530
matters when you are the one running
infrastructure
00:01:42.530 --> 00:01:46.390
and owning reliability. Most weeks, it's a
quick
00:01:46.390 --> 00:01:50.430
DevOps, SRE, platform, cloud, and security
news
00:01:50.430 --> 00:01:54.069
recap. In between those, I do conversation
episodes
00:01:54.069 --> 00:01:57.700
with people building and operating the
systems we
00:01:57.700 --> 00:02:00.340
all end up depending on. Today, I'm joined
by
00:02:00.340 --> 00:02:03.500
Mat Ryer from Grafana Labs. Mat is Senior
00:02:03.500 --> 00:02:07.260
Director of AI at Grafana, a longtime Go
developer,
00:02:07.579 --> 00:02:10.879
author, open source contributor, and
podcast
00:02:10.879 --> 00:02:14.439
host. And in this episode, we talk about
what
00:02:14.439 --> 00:02:18.120
happens when AI moves out of the demo
booth and
00:02:18.120 --> 00:02:21.219
into production. We get into Grafana
Assistant,
00:02:21.580 --> 00:02:26.199
AI observability, evals, LLM-as-judge
patterns,
00:02:26.669 --> 00:02:30.389
telemetry cost, OpenTelemetry, agent
guardrails
00:02:30.389 --> 00:02:34.610
and why a chat interface is not enough if
the
00:02:34.610 --> 00:02:37.210
system underneath is making decisions
people
00:02:37.210 --> 00:02:40.509
need to trust one of the threads i liked
most
00:02:40.509 --> 00:02:43.750
in this conversation is Mat's point that
AI
00:02:43.750 --> 00:02:47.129
gives us a new primitive that sounds big
and
00:02:47.129 --> 00:02:51.110
abstract but his example is pretty simple
a conversation
00:02:51.110 --> 00:02:56.090
plus a loop suddenly becomes an agent Or
as Mat
00:02:56.090 --> 00:02:59.949
put it, for loops are back. But once
people rely
00:02:59.949 --> 00:03:03.449
on that agent, the problem changes. Now
you need
00:03:03.449 --> 00:03:06.430
to know whether it helped, whether it
answered
00:03:06.430 --> 00:03:08.889
the question, whether it picked the right
tools,
00:03:09.150 --> 00:03:12.550
whether a prompt change improved one
workflow
00:03:12.550 --> 00:03:16.050
while breaking another, whether your AI
feature
00:03:16.050 --> 00:03:20.069
is actually observable enough to operate
like
00:03:20.069 --> 00:03:23.550
the rest of your production stack. We also
talk
00:03:23.550 --> 00:03:27.250
about why AI observability is not just
latency,
00:03:27.449 --> 00:03:32.270
logs, and HTTP 200s. Those still matter,
obviously.
00:03:32.629 --> 00:03:37.469
But now you also care about behavior,
cost, tool
00:03:37.469 --> 00:03:42.110
selection, user feedback, eval results,
model
00:03:42.110 --> 00:03:45.729
versions, prompt changes, and whether the
agent
00:03:45.729 --> 00:03:48.840
is producing something useful. or just
confidently
00:03:48.840 --> 00:03:52.360
wandering around your telemetry matt also
gets
00:03:52.360 --> 00:03:55.780
into UX which i think is more important
here
00:03:55.780 --> 00:03:58.580
than people give it credit for because if
the
00:03:58.580 --> 00:04:01.939
AI gives you a wall of text you still have
to
00:04:01.939 --> 00:04:04.919
trust that text but if it can show you the
graph
00:04:04.919 --> 00:04:08.620
deep link you into the right Grafana view
apply
00:04:08.620 --> 00:04:12.280
the filters and let you inspect the source
data
00:04:12.280 --> 00:04:15.580
yourself that is a very different
experience
00:04:15.580 --> 00:04:20.120
and near the end We talk about where AI
can actually
00:04:20.120 --> 00:04:23.680
help operations teams today, not some
giant six
00:04:23.680 --> 00:04:27.259
-month AI transformation project. More
like pick
00:04:27.259 --> 00:04:30.220
something small. Use it to enhance the
workflows
00:04:30.220 --> 00:04:33.699
you already have. Let it write the query.
Let
00:04:33.699 --> 00:04:36.680
it help with the first pass. Let it
automate
00:04:36.680 --> 00:04:40.019
the boring parts, but build the guardrails
and
00:04:40.019 --> 00:04:43.319
feedback loops around it, which honestly
fits
00:04:43.319 --> 00:04:46.420
the name of the show pretty well. Ship
something
00:04:46.420 --> 00:04:50.319
small. learn from it, then ship again. All
right,
00:04:50.399 --> 00:05:00.079
let's jump in. Today, I'm joined by Mat
Ryer
00:05:00.079 --> 00:05:03.060
from Grafana Labs. Mat is a senior
director
00:05:03.060 --> 00:05:06.740
of AI at Grafana, where he's focused on
how AI
00:05:06.740 --> 00:05:09.699
fits into observability and production
systems.
00:05:10.120 --> 00:05:13.819
He's also a longtime Go developer, author,
open
00:05:13.819 --> 00:05:17.290
source contributor, and podcast host. And
we're
00:05:17.290 --> 00:05:19.870
talking about what happens when AI moves
from
00:05:19.870 --> 00:05:23.230
demos into production, why observability
for
00:05:23.230 --> 00:05:26.810
AI systems is not the same as basic
service monitoring,
00:05:27.009 --> 00:05:30.610
and what teams should be thinking about as
telemetry
00:05:30.610 --> 00:05:33.889
volume, cost, and operational complexity
keep
00:05:33.889 --> 00:05:36.709
climbing. Mat, thank you for joining me.
Thank
00:05:36.709 --> 00:05:39.949
you, Brian. Pleasure to be here. So
starting
00:05:39.949 --> 00:05:43.750
out, I'm curious, as AI systems move from
experiments
00:05:43.750 --> 00:05:47.220
into production, What are teams
underestimating?
00:05:48.439 --> 00:05:52.199
Yes. Well, first of all, I think it's very
exciting,
00:05:52.339 --> 00:05:55.339
this whole space. And I think that's where
I
00:05:55.339 --> 00:05:59.500
always start with this. We have a new
primitive
00:05:59.500 --> 00:06:02.079
now. We have a new capability, something
that
00:06:02.079 --> 00:06:05.720
we couldn't do before. And that, I think,
is
00:06:05.720 --> 00:06:08.199
just a very exciting thing. There are
concerns
00:06:08.199 --> 00:06:13.230
with AI and valid concerns. alongside them
are
00:06:13.230 --> 00:06:16.350
i think is the fact that we are now in a
new
00:06:16.350 --> 00:06:19.470
world and this is a very exciting thing
you see
00:06:19.470 --> 00:06:23.029
people making little agents to solve all
kinds
00:06:23.029 --> 00:06:26.610
of little problems we see people building
big
00:06:26.610 --> 00:06:28.829
agents and big complex ones as we are
doing as
00:06:28.829 --> 00:06:31.529
well at Grafana Labs to solve more complex
and
00:06:31.529 --> 00:06:35.490
deeper problems and just now like having a
a
00:06:35.490 --> 00:06:38.560
new primitive like a conversation Being
able
00:06:38.560 --> 00:06:41.000
to now, even just putting that in a loop,
by
00:06:41.000 --> 00:06:43.720
the way, which was when things became
agentic.
00:06:43.800 --> 00:06:45.959
That's just a for loop. That's just like,
oh,
00:06:46.019 --> 00:06:49.279
for loops are back. They're so back.
They're
00:06:49.279 --> 00:06:54.139
cool again. And I think agents are just
amazing.
00:06:54.319 --> 00:06:56.480
Just a simple primitive put together with
something
00:06:56.480 --> 00:06:58.939
that we're already very familiar with can
produce
00:06:58.939 --> 00:07:01.439
something new. So we're still in that
phase.
00:07:01.819 --> 00:07:04.459
And I love to see all the little things
that
00:07:04.459 --> 00:07:08.620
people are building. At some point you do
end
00:07:08.620 --> 00:07:11.379
up with, like we have a product in
production,
00:07:11.639 --> 00:07:15.100
Grafana Assistant, and people are relying
on
00:07:15.100 --> 00:07:16.660
it. They're using it as part of their
day-to
00:07:16.660 --> 00:07:19.560
-day. Some people have leaned very much so
into
00:07:19.560 --> 00:07:22.560
it and it's really unlocked them for all
the
00:07:22.560 --> 00:07:24.800
observability stuff they're trying to do.
So
00:07:24.800 --> 00:07:27.220
now suddenly we can't just vibe it. We
can't
00:07:27.220 --> 00:07:29.620
just feel it out. We've got to grow up and
make
00:07:29.620 --> 00:07:32.589
sure that this is going to work for them.
And
00:07:32.589 --> 00:07:35.209
I think people underestimate how hard that
is.
00:07:35.310 --> 00:07:39.430
It is quite easy to wire up an LLM to
anything,
00:07:39.709 --> 00:07:45.149
do a few simple tools, and suddenly you
unlock
00:07:45.149 --> 00:07:48.410
this new capability. But how do you really
know
00:07:48.410 --> 00:07:50.149
that's good? How do you know when you
change
00:07:50.149 --> 00:07:52.170
it, you aren't making it worse in some
other
00:07:52.170 --> 00:07:54.649
place? You know, we used to have unit
tests,
00:07:54.889 --> 00:07:58.329
test suites protecting us from all this.
We didn't
00:07:58.329 --> 00:08:01.310
really have that now. So how do we make
sure?
00:08:01.790 --> 00:08:04.670
when we make changes to it it is getting
better
00:08:04.670 --> 00:08:06.829
and how do we make sure it's even like
correct
00:08:06.829 --> 00:08:10.209
in the first place you know it demos very
well
00:08:10.209 --> 00:08:13.230
like you say we've kind of i think checked
off
00:08:13.230 --> 00:08:16.470
that from the mission like has to demo
well for
00:08:16.470 --> 00:08:21.449
people to care to it demos pretty well uh
but
00:08:21.449 --> 00:08:24.230
beyond the demo yeah how do you how do you
make
00:08:24.230 --> 00:08:26.889
sure it does that and we we have we had to
solve
00:08:26.889 --> 00:08:28.850
this problem in assistant which we built
our
00:08:28.850 --> 00:08:31.810
Grafana Assistant we launched that I think
it
00:08:31.810 --> 00:08:35.990
was last year now. And yeah, we've had to
solve
00:08:35.990 --> 00:08:39.330
that problem. So at a high level, how do
you
00:08:39.330 --> 00:08:41.889
start with that problem statement? How do
you,
00:08:41.909 --> 00:08:45.029
do you build special testing around that?
How
00:08:45.029 --> 00:08:48.289
do you actually solve that? Yeah, so I'd
say,
00:08:48.470 --> 00:08:51.789
right in the beginning, I think people
still
00:08:51.789 --> 00:08:55.049
vibe testing, manual testing, I still
think is
00:08:55.049 --> 00:08:57.230
the thing to do in the beginning. Because
it's
00:08:57.230 --> 00:08:58.929
the really when you're just feeling it
out, you
00:08:58.929 --> 00:09:01.820
need to know. Is this even possible? Is
this
00:09:01.820 --> 00:09:05.139
feasible? Does this make sense? And
there's so
00:09:05.139 --> 00:09:08.240
much innovation and there's so much
opportunity
00:09:08.240 --> 00:09:11.620
to learn that actually you still have to
have
00:09:11.620 --> 00:09:14.340
that same attitude of build things and
ship it.
00:09:14.419 --> 00:09:17.100
That's why I came on the Ship It Weekly
podcast
00:09:17.100 --> 00:09:20.360
because it's all about shipping it. So I
would
00:09:20.360 --> 00:09:23.340
say don't wait for there to be good evals.
Don't
00:09:23.340 --> 00:09:25.360
wait for you to have this problem solved
before
00:09:25.360 --> 00:09:28.940
you progress. But it is something to think
about
00:09:28.940 --> 00:09:31.370
in the back. in the background and yeah
when
00:09:31.370 --> 00:09:33.730
when we did assistant we essentially had
to build
00:09:33.730 --> 00:09:36.490
we had lots of different tools of course
there's
00:09:36.490 --> 00:09:39.210
like the normal telemetry that we're very
good
00:09:39.210 --> 00:09:42.629
at Grafana Labs metrics for measuring
latencies
00:09:42.629 --> 00:09:45.350
and things like this and we even have
metrics
00:09:45.350 --> 00:09:48.049
for costs measuring the the cost of
different
00:09:48.049 --> 00:09:52.210
operations and things uh logs of course
kind
00:09:52.210 --> 00:09:55.250
of makes sense i think tracing becomes
more interesting
00:09:55.250 --> 00:09:57.750
now because you could think of a
conversation
00:09:57.750 --> 00:10:00.370
really as a as a sort of trace you could
imagine
00:10:00.370 --> 00:10:05.269
it in that you know in that model um and
and
00:10:05.269 --> 00:10:08.350
i think so so you get a lot already just
of the
00:10:08.350 --> 00:10:10.710
basics but there are new things you care
about
00:10:10.710 --> 00:10:13.330
suddenly uh you know it's not just getting
a
00:10:13.330 --> 00:10:17.610
200 back from an LLM but now how do we how
do
00:10:17.610 --> 00:10:20.970
we make sure that the content the the way
it
00:10:20.970 --> 00:10:23.389
behaved the tools it chose to use how do
we make
00:10:23.389 --> 00:10:27.590
sure that stuff is right and we did it um
partly
00:10:28.620 --> 00:10:32.679
By vibing tools internally, literally we'd
have
00:10:32.679 --> 00:10:36.299
like, we've got a handful of them and
various
00:10:36.299 --> 00:10:40.139
dashboards and things and different
databases
00:10:40.139 --> 00:10:42.320
and different ways of storing things, you
know,
00:10:42.320 --> 00:10:44.440
just trying to sort of cobble together the
solution
00:10:44.440 --> 00:10:46.899
of this. And again, I think that was the
right
00:10:46.899 --> 00:10:49.820
approach in the beginning because it's a
playground
00:10:49.820 --> 00:10:52.379
where we can experiment and we can try
things.
00:10:52.500 --> 00:10:55.480
We can solve problems without having to
think
00:10:55.480 --> 00:10:58.299
about solving the entire problem. of this
we
00:10:58.299 --> 00:11:00.179
can just pick one piece of it and just
solve
00:11:00.179 --> 00:11:02.500
it with a standalone thing so there's a
lot of
00:11:02.500 --> 00:11:04.700
freedom in that and i think the culture
that
00:11:04.700 --> 00:11:06.980
we had in the engineering team when we're
building
00:11:06.980 --> 00:11:09.860
this played a lot into that because yeah
we had
00:11:09.860 --> 00:11:12.840
to make sure that people could try things
just
00:11:12.840 --> 00:11:15.620
solve that particular problem not go for
months
00:11:15.620 --> 00:11:18.460
and have big design sessions and just do
this
00:11:18.460 --> 00:11:20.960
enormous project to try and solve the
whole thing
00:11:22.269 --> 00:11:24.610
After we'd done all that, of course, you
then
00:11:24.610 --> 00:11:26.529
end up in a situation where you do have
quite
00:11:26.529 --> 00:11:29.070
a complex little set of internal tools
that are
00:11:29.070 --> 00:11:32.470
all slightly overlapping and doing things
in
00:11:32.470 --> 00:11:34.230
different ways. They don't all agree with
each
00:11:34.230 --> 00:11:37.830
other. Some of them are agentic
themselves. So
00:11:37.830 --> 00:11:39.809
that's even more getting like Inception,
where
00:11:39.809 --> 00:11:43.090
the agents are monitoring the agents to
make
00:11:43.090 --> 00:11:46.639
sure that they're behaving properly. And
so then
00:11:46.639 --> 00:11:49.159
we were able to take all that learning and
put
00:11:49.159 --> 00:11:51.980
that together into our AI observability
product,
00:11:52.100 --> 00:11:55.139
which is a new feature that we have in
Grafana
00:11:55.139 --> 00:12:00.120
Cloud. And that's our answer to this. And
we
00:12:00.120 --> 00:12:03.059
think we've picked the most important
things
00:12:03.059 --> 00:12:04.720
that people should care about and built
that
00:12:04.720 --> 00:12:07.539
into the product. And you can see we're
different
00:12:07.539 --> 00:12:10.159
to some kind of competitive products.
We're different
00:12:10.159 --> 00:12:12.600
in some interesting ways because I think
some
00:12:12.600 --> 00:12:16.279
of them were imagined. Whereas we built
ours
00:12:16.279 --> 00:12:18.919
from the real day-to-day of actually
operating
00:12:18.919 --> 00:12:23.279
the assistant at scale. Yeah, very cool.
What
00:12:23.279 --> 00:12:26.779
does production level observability for AI
actually
00:12:26.779 --> 00:12:31.220
look like? Well, you do need evals. You
need
00:12:31.220 --> 00:12:35.620
to be able to test whether what's actually
being
00:12:35.620 --> 00:12:38.539
asked for is being answered. And we have
a...
00:12:38.809 --> 00:12:41.409
We have a technique inside where we
actually
00:12:41.409 --> 00:12:44.350
have an LLM as a judge looking at a
conversation.
00:12:44.590 --> 00:12:46.070
And you can just do this. You can try
this. It
00:12:46.070 --> 00:12:48.549
works. Just ask it, did you answer the
user's
00:12:48.549 --> 00:12:50.590
request? Because you can see their first
message.
00:12:50.850 --> 00:12:53.129
You can see the path you went on and where
you
00:12:53.129 --> 00:12:55.450
ended up. Did you answer it? Did the user
have
00:12:55.450 --> 00:12:59.389
to intervene in any way? Were you
ambiguous?
00:12:59.470 --> 00:13:02.549
You can kind of ask it these things in,
and you're
00:13:02.549 --> 00:13:04.289
dealing with natural language now. So it's
like
00:13:04.289 --> 00:13:06.870
you write these tests also in natural
language.
00:13:08.639 --> 00:13:11.460
And then we can measure that. We can start
to
00:13:11.460 --> 00:13:13.320
measure and start to put some data and
some science
00:13:13.320 --> 00:13:18.399
behind the behavior of agents. So evals, I
think,
00:13:18.399 --> 00:13:21.580
are going to be increasingly more and more
and
00:13:21.580 --> 00:13:25.480
more important. We want these AIs, we want
these
00:13:25.480 --> 00:13:28.580
technologies to be trusted and to be used.
Some
00:13:28.580 --> 00:13:31.559
people write them off just because they
know
00:13:31.559 --> 00:13:33.840
that they can be wrong. So they can't
trust it,
00:13:33.919 --> 00:13:35.139
and so therefore they're not going to use
it
00:13:35.139 --> 00:13:37.610
at all. I don't think that's the right
attitude,
00:13:37.750 --> 00:13:40.669
personally. You know, the internet, the
network
00:13:40.669 --> 00:13:43.129
connections aren't reliable. They error
all the
00:13:43.129 --> 00:13:45.830
time. And we kind of paper over those
errors
00:13:45.830 --> 00:13:48.350
so you don't really notice it. Same kind
of thing
00:13:48.350 --> 00:13:52.289
with AI. We have these non-deterministic
systems.
00:13:52.370 --> 00:13:55.009
In some ways, when it comes to doing
investigations
00:13:55.009 --> 00:13:58.129
and things, for like digging around your
telemetry,
00:13:58.149 --> 00:14:01.629
that can actually be turned into an asset
where...
00:14:01.919 --> 00:14:04.559
The fact that multiple agents take random
routes
00:14:04.559 --> 00:14:08.360
through things can actually be good.
That's what
00:14:08.360 --> 00:14:10.440
humans might do if they all take a
different
00:14:10.440 --> 00:14:12.580
approach. But then you've got an even
bigger
00:14:12.580 --> 00:14:16.019
observability problem because now this one
thing,
00:14:16.159 --> 00:14:18.860
suddenly it's multiple conversations with
models,
00:14:19.019 --> 00:14:22.100
not even the same model, multiple models
often,
00:14:22.279 --> 00:14:24.879
certainly different system prompts,
different
00:14:24.879 --> 00:14:28.820
sets of tools. All this complexity now is
kind
00:14:28.820 --> 00:14:32.309
of a big... mess that you have to deal
with and
00:14:32.309 --> 00:14:35.929
we think yeah simple observability simple
uh
00:14:35.929 --> 00:14:38.990
and new signals as well as part of this is
the
00:14:38.990 --> 00:14:40.789
solution to so that you can just answer
those
00:14:40.789 --> 00:14:43.049
questions and make sure that your agents
are
00:14:43.049 --> 00:14:45.429
on the right track so kind of following up
with
00:14:45.429 --> 00:14:48.450
that a lot of teams instrumented
everything then
00:14:48.450 --> 00:14:51.169
they realized they were drowning in
telemetry
00:14:51.169 --> 00:14:56.370
costs uh what changed yes I think we're
going
00:14:56.370 --> 00:14:58.870
to go through that as well for AI. So I
think
00:14:58.870 --> 00:15:01.350
you're right. People would just monitor
everything.
00:15:01.629 --> 00:15:04.809
It feels really cheap. Like when you're
writing
00:15:04.809 --> 00:15:07.490
something, you just add a counter in code.
It
00:15:07.490 --> 00:15:08.889
feels like the cheapest thing you can do.
As
00:15:08.889 --> 00:15:10.690
soon as you then hit that scale, suddenly
you're
00:15:10.690 --> 00:15:14.049
counting. You're just counting at scale is
hard.
00:15:14.610 --> 00:15:17.830
So, yeah, you kind of create these new
kind of
00:15:17.830 --> 00:15:20.889
problems when you hit scale. The same
thing is
00:15:20.889 --> 00:15:23.799
definitely true for AI. I think a lot of
people
00:15:23.799 --> 00:15:26.539
aren't really monitoring much at the
moment.
00:15:27.299 --> 00:15:31.200
But when they do, and as this increases,
we'll
00:15:31.200 --> 00:15:32.740
have the same kind of problem. And, you
know,
00:15:32.759 --> 00:15:34.960
Grafana Labs has this technology called
adaptive
00:15:34.960 --> 00:15:37.860
telemetry, which is basically what they do
is
00:15:37.860 --> 00:15:41.440
they look at all the telemetry that you've
used.
00:15:41.500 --> 00:15:43.950
They look at all the queries you're
making. And
00:15:43.950 --> 00:15:47.570
then they identify bits where you've never
queried
00:15:47.570 --> 00:15:50.549
this metric. So we're going to
automatically
00:15:50.549 --> 00:15:52.809
sample it down. So we'll just take a
sample of
00:15:52.809 --> 00:15:56.429
it. We won't capture all of it. And, of
course,
00:15:56.429 --> 00:15:59.370
you can tune it and tweak it and things.
And
00:15:59.370 --> 00:16:01.429
this was really an answer to that problem,
the
00:16:01.429 --> 00:16:04.370
fact that people's telemetry was just
growing
00:16:04.370 --> 00:16:08.230
and growing and growing, and not just in
the
00:16:08.230 --> 00:16:10.389
history of it, but even just the now,
because
00:16:10.389 --> 00:16:12.409
people's systems and complexity is going
up.
00:16:13.019 --> 00:16:15.639
So I think that we'll end up with another
kind
00:16:15.639 --> 00:16:17.899
of adaptive telemetry. We have already in
our
00:16:17.899 --> 00:16:21.200
solution sampling, so you can choose how
often
00:16:21.200 --> 00:16:27.120
the evaluators will run on conversations.
So
00:16:27.120 --> 00:16:28.659
we're kind of thinking about that already.
And
00:16:28.659 --> 00:16:31.379
I think this is the advantage of building
this
00:16:31.379 --> 00:16:34.980
solution at Grafana Labs, because we had
to solve
00:16:34.980 --> 00:16:36.940
a lot of the observability problems
already.
00:16:37.179 --> 00:16:39.659
And a lot of the thinking applies
directly, which
00:16:39.659 --> 00:16:41.740
is very cool. We have these drill down
apps,
00:16:41.879 --> 00:16:45.659
which... basically present the data and
you can
00:16:45.659 --> 00:16:48.080
just click and drill into it. That turns
out
00:16:48.080 --> 00:16:49.679
to be good for humans, but actually quite
good
00:16:49.679 --> 00:16:51.620
for agents as well, because it's the same
problem.
00:16:51.940 --> 00:16:55.799
There's just too much stuff there and to
look
00:16:55.799 --> 00:16:58.899
through it all and find it is hard. So how
do
00:16:58.899 --> 00:17:01.360
we, yeah, the fact that we solve that for
humans
00:17:01.360 --> 00:17:03.820
turns out it's quite good to solve for
agents.
00:17:04.579 --> 00:17:08.539
Where does OpenTelemetry fit into this AI
observability
00:17:08.539 --> 00:17:12.299
story? I think it's going to have a... big
part
00:17:12.299 --> 00:17:15.400
of it we actually are talking already with
um
00:17:15.400 --> 00:17:17.759
ted young who was one of the co-founders
of
00:17:17.759 --> 00:17:22.500
OpenTelemetry yeah and we're kind of um
i've
00:17:22.500 --> 00:17:24.259
seen some proposals and i've seen some
things
00:17:24.259 --> 00:17:26.920
where the conversation data gets put into
trace
00:17:26.920 --> 00:17:29.440
data so that they they are stitched
together
00:17:29.440 --> 00:17:32.200
in that way we didn't take that approach
because
00:17:32.200 --> 00:17:36.619
of scale issues um that we saw so we
haven't
00:17:36.619 --> 00:17:39.019
got that approach in our solution but It
could
00:17:39.019 --> 00:17:41.740
well end up being some new standard like
that.
00:17:41.819 --> 00:17:44.140
But I just think with OpenTelemetry, it's
going
00:17:44.140 --> 00:17:46.519
to just take some time. You know, we like
to
00:17:46.519 --> 00:17:49.680
move very fast in the AI department in
Grafana
00:17:49.680 --> 00:17:52.680
Labs and in Grafana Labs in general. In
some
00:17:52.680 --> 00:17:55.380
places, you do have to go slowly. Where
you're
00:17:55.380 --> 00:17:58.140
dealing with standards or foundational
things
00:17:58.140 --> 00:18:02.140
that you really need to get them right,
you're
00:18:02.140 --> 00:18:04.200
going to go a bit slower. So I think it'll
take
00:18:04.200 --> 00:18:07.859
some time. But I do think OpenTelemetry
should...
00:18:08.240 --> 00:18:11.819
embrace this new world and and see which
new
00:18:11.819 --> 00:18:14.839
signals are common enough that we are
going to
00:18:14.839 --> 00:18:17.420
then capture them we got to make those
decisions
00:18:17.420 --> 00:18:20.460
ourselves when we built our solution so of
course
00:18:20.460 --> 00:18:23.660
we capture you know the the basic things
you
00:18:23.660 --> 00:18:26.640
expect metrics for latency and things we
add
00:18:26.640 --> 00:18:29.700
cost into there using metrics and stuff
that
00:18:29.700 --> 00:18:31.579
that's also turns out to be quite
complicated
00:18:31.579 --> 00:18:34.220
to do but very powerful if you have
accurate
00:18:34.220 --> 00:18:37.529
cost data um And with the way that we have
labels
00:18:37.529 --> 00:18:39.410
and things, kind of like the Prometheus
model,
00:18:39.589 --> 00:18:42.829
you can use labels for the different
versions
00:18:42.829 --> 00:18:45.289
of your app or different agent names or
something.
00:18:45.470 --> 00:18:48.829
So now you can compare performance, not
just
00:18:48.829 --> 00:18:51.630
cost and metrics, but the other dimensions
as
00:18:51.630 --> 00:18:53.970
well. And if you've got evals running, you
can
00:18:53.970 --> 00:18:56.470
compare behavior between different models
as
00:18:56.470 --> 00:18:59.349
well. So you can say, version one of our
model.
00:18:59.690 --> 00:19:02.089
Took a long time to answer this kind of
question.
00:19:02.329 --> 00:19:04.910
And it's because we hadn't prompted it
much or,
00:19:04.950 --> 00:19:06.750
you know, we just gave it one tool that it
could
00:19:06.750 --> 00:19:09.549
use. We did some work. We added a couple
of tools.
00:19:09.670 --> 00:19:13.470
We tweaked some prompts and our new
version is
00:19:13.470 --> 00:19:16.349
now much faster. It solves this problem.
It's
00:19:16.349 --> 00:19:18.829
much more accurate. The user's happy. We
capture
00:19:18.829 --> 00:19:21.470
ratings as well of users, you know, thumbs
up,
00:19:21.470 --> 00:19:24.799
thumbs down, as well as comments. All the
sort
00:19:24.799 --> 00:19:27.220
of practical things you need to actually
have
00:19:27.220 --> 00:19:29.599
some kind of feedback loop where you can
improve
00:19:29.599 --> 00:19:34.119
your agents. Is that what you have found
to be,
00:19:34.140 --> 00:19:36.660
I guess, the best feedback loop is like
providing
00:19:36.660 --> 00:19:41.940
like right in the chat, the ability to
give feedback
00:19:41.940 --> 00:19:44.099
immediately like, hey, this was not the
right
00:19:44.099 --> 00:19:46.819
response or hey, this was not detailed
enough
00:19:46.819 --> 00:19:49.519
or it went off on a tangent that was not
really
00:19:49.519 --> 00:19:53.500
related to the question. Or are you
feeding that
00:19:53.500 --> 00:19:56.619
in with other information as well? And
then how
00:19:56.619 --> 00:20:00.259
does that work? Yeah, I think user
feedback is
00:20:00.259 --> 00:20:03.019
the most valuable. And it is nice that you
can,
00:20:03.119 --> 00:20:05.019
even just a thumbs up, thumbs down is
enough
00:20:05.019 --> 00:20:10.420
to give you something to go on. I think
actually
00:20:10.420 --> 00:20:14.480
typing in a comment is even better, right?
Because
00:20:14.480 --> 00:20:17.980
we can genuinely have LLMs consider that
commentary
00:20:17.980 --> 00:20:22.309
and even suggest PRs. into our repo to
improve
00:20:22.309 --> 00:20:25.710
prompts and things like this um so mostly
though
00:20:25.710 --> 00:20:28.470
we don't get uh we don't get as much of
that
00:20:28.470 --> 00:20:31.230
as we would like so we can't rely on that
and
00:20:31.230 --> 00:20:32.789
the other thing is we're very sensitive
about
00:20:32.789 --> 00:20:37.109
learning from people's data without them
realizing
00:20:37.109 --> 00:20:39.690
that's what we're doing so yeah we are we
are
00:20:39.690 --> 00:20:42.829
very um yeah we want to make sure that we
are
00:20:42.829 --> 00:20:46.910
uh not yeah just being very trustworthy
basically
00:20:46.910 --> 00:20:49.390
if we're going to build AI that we expect
people
00:20:49.390 --> 00:20:51.880
to use I think the standards have to be
very,
00:20:51.900 --> 00:20:54.920
very high for things like that. Make sure
that
00:20:54.920 --> 00:20:57.559
your data is protected. It's easier than
ever
00:20:57.559 --> 00:21:02.440
to leak data. We had a case where, luckily
it
00:21:02.440 --> 00:21:06.660
was test data, but we set up a GitHub
permission
00:21:06.660 --> 00:21:09.660
for an MCP server. We didn't realize that
you
00:21:09.660 --> 00:21:12.140
don't need write permission to open an
issue
00:21:12.140 --> 00:21:16.019
in a public repo. So we didn't give it
write
00:21:16.019 --> 00:21:18.079
permission, but still it started to open
an issue.
00:21:20.349 --> 00:21:23.190
it was a test box, so it was no problem.
It was putting data that
00:21:23.190 --> 00:21:26.150
it saw in there into the issues very cool
feature
00:21:26.150 --> 00:21:29.329
of course and extremely useful for like
IRM and
00:21:29.329 --> 00:21:31.950
incident management and things especially
useful
00:21:31.950 --> 00:21:35.490
when it opens a PR for you very nice but
we've
00:21:35.490 --> 00:21:37.609
got to do that carefully and safely and i
think
00:21:37.609 --> 00:21:40.109
knowing that your systems are safe and
aren't
00:21:40.109 --> 00:21:44.589
behaving and that they will react uh
properly
00:21:44.589 --> 00:21:48.160
and appropriately to prompts that are kind
of
00:21:48.160 --> 00:21:51.279
probing and trying to do malicious things.
I
00:21:51.279 --> 00:21:54.059
think that all is a very important part of
this
00:21:54.059 --> 00:21:58.440
that we're still scratching the surface
on. Yeah.
00:21:58.759 --> 00:22:01.940
I think it also kind of highlights the
need for
00:22:01.940 --> 00:22:06.740
guardrails because LLMs are over-eager,
right,
00:22:06.819 --> 00:22:10.220
to solve the problem, get to the root
cause and
00:22:10.220 --> 00:22:13.599
provide a solution. And if it has the
means to
00:22:13.599 --> 00:22:17.650
do that, it will. Yeah. It's true for
tools as
00:22:17.650 --> 00:22:20.450
well. It's true for like, if you give it
lots
00:22:20.450 --> 00:22:22.710
of tools, it's just going to be like a kid
in
00:22:22.710 --> 00:22:24.730
a candy shop. It's going to be so excited
to
00:22:24.730 --> 00:22:27.170
just go and use all the tools. And given
that
00:22:27.170 --> 00:22:29.269
it's non-deterministic, you can easily end
up
00:22:29.269 --> 00:22:32.089
in a situation where it's using quite
random
00:22:32.089 --> 00:22:35.430
tools to try and solve problems. And it's
trying
00:22:35.430 --> 00:22:38.170
to use a hammer to open a door, for
example,
00:22:38.349 --> 00:22:40.829
which I'm not a DIY expert, but I don't
think
00:22:40.829 --> 00:22:45.420
that's how you do that. But yeah, so. So
how
00:22:45.420 --> 00:22:47.839
do you sort of tackle that? I think that's
kind
00:22:47.839 --> 00:22:51.920
of another interesting angle on this.
Yeah, so
00:22:51.920 --> 00:22:56.339
from a guardrails perspective, I'm sure at
the
00:22:56.339 --> 00:22:58.859
IAM level, like you're providing
read-write
00:22:58.859 --> 00:23:01.259
permissions or scoped permissions based on
what's
00:23:01.259 --> 00:23:03.579
needed. But are you doing like
pre-prompting
00:23:03.579 --> 00:23:06.180
as well? Or are you not relying on
pre-prompting
00:23:06.180 --> 00:23:08.279
at all? Because that's just too close to
the
00:23:08.279 --> 00:23:13.259
LLM and, you know, it could be influenced.
uh
00:23:13.259 --> 00:23:15.339
well yeah so we have how do you build that
system
00:23:15.339 --> 00:23:19.259
yeah yeah we've got a few different things
one
00:23:19.259 --> 00:23:22.079
of the so our AI observability solution
lets
00:23:22.079 --> 00:23:26.359
us uh write kind of security type tests
and assert
00:23:26.359 --> 00:23:28.740
them basically on the model on the model
so that
00:23:28.740 --> 00:23:31.240
kind of gets us we can do some kind of
testing
00:23:31.240 --> 00:23:33.819
there just to set test the end point to
see what
00:23:33.819 --> 00:23:35.740
how are you going to behave if a user asks
for
00:23:35.740 --> 00:23:40.019
this um you know it's also dealing with
customer
00:23:40.019 --> 00:23:43.259
data so it's also dealing with data in
logs which
00:23:43.259 --> 00:23:45.740
can just be any text so you have to be
careful
00:23:45.740 --> 00:23:47.680
to make sure that it doesn't just load
some logs
00:23:47.680 --> 00:23:50.140
and think oh i'm just going to execute
something
00:23:50.140 --> 00:23:53.940
in in here um and the real answer in in
assistant
00:23:53.940 --> 00:23:56.480
is we just have lots of different checks
and
00:23:56.480 --> 00:23:59.319
different gates and things happening
because
00:23:59.319 --> 00:24:02.559
there's so much it's a very complex um
solution
00:24:02.559 --> 00:24:05.240
i think it looks simple when you use it
and the
00:24:05.240 --> 00:24:07.440
UX there's another thing i will get to
talk about
00:24:07.440 --> 00:24:09.619
hopefully is the the user experience of
this
00:24:09.619 --> 00:24:12.319
is paramount here more than ever and it's
always
00:24:12.319 --> 00:24:16.420
been important i think in tech but the um
you
00:24:16.420 --> 00:24:20.619
know the fact that you you the fact that
you
00:24:20.619 --> 00:24:25.220
need to uh just just have well the fact
that
00:24:25.220 --> 00:24:28.119
we have this complex system we use
multiple models
00:24:28.119 --> 00:24:30.759
we've got like specialist tools inside
assistant
00:24:30.759 --> 00:24:33.660
all running in Grafana Cloud there's
preemptive
00:24:33.660 --> 00:24:36.000
work that happens around discovering the
infrastructure
00:24:36.539 --> 00:24:40.140
And so we put that into a semantic vector
database.
00:24:40.240 --> 00:24:43.680
So it's the agents have access to kind of
learning
00:24:43.680 --> 00:24:47.339
that it's had before. When incidents and
investigations
00:24:47.339 --> 00:24:50.220
are happening, that also is a great source
of
00:24:50.220 --> 00:24:53.019
information for agents. But again, how do
you
00:24:53.019 --> 00:24:55.420
make sure that's all safe? So I think at
any
00:24:55.420 --> 00:24:57.660
point when you're integrating and bringing
some
00:24:57.660 --> 00:25:00.660
data in, you have that security question
of in
00:25:00.660 --> 00:25:04.930
what ways could this be malicious? What
protections
00:25:04.930 --> 00:25:09.190
can we do here to help and make sure it's
not
00:25:09.190 --> 00:25:12.410
going to be a risky system for people to
use?
00:25:13.230 --> 00:25:17.049
And how have you thought about the UX side
of
00:25:17.049 --> 00:25:22.890
the user experience side of the AI agents?
How
00:25:22.890 --> 00:25:25.170
is that architected or how do you think
about
00:25:25.170 --> 00:25:29.130
providing that experience in a quality
environment
00:25:29.130 --> 00:25:33.240
that is actually useful? Yeah, this is
somewhere
00:25:33.240 --> 00:25:35.759
where I think the team have excelled. And
we
00:25:35.759 --> 00:25:38.700
have heard this feedback. One of the big
labs
00:25:38.700 --> 00:25:42.240
told us that at the time, this was the
best looking
00:25:42.240 --> 00:25:45.220
version of something that they'd seen
using their
00:25:45.220 --> 00:25:48.180
models, which was a big, big compliment,
of course.
00:25:48.319 --> 00:25:52.839
I think the key is we on the team just got
really
00:25:52.839 --> 00:25:56.930
annoyed by walls of text. It was a kind of
a
00:25:56.930 --> 00:25:59.509
byproduct of this is that, you know, the
fact
00:25:59.509 --> 00:26:02.549
that you are dealing in conversation means
it's
00:26:02.549 --> 00:26:04.450
generating text as what kind of how the
whole
00:26:04.450 --> 00:26:08.170
thing works. And it's just so much text.
Now,
00:26:08.230 --> 00:26:13.509
when it comes to trusting LLMs, you know,
in
00:26:13.509 --> 00:26:15.829
that text generation, that's where they
can hallucinate
00:26:15.829 --> 00:26:19.990
things. So we, instead of using text
wherever
00:26:19.990 --> 00:26:22.740
we can in the assistant. We'll use a
Grafana
00:26:22.740 --> 00:26:25.680
visualization. So, you know, luckily we
have
00:26:25.680 --> 00:26:27.819
this library of these beautiful tools.
Grafana
00:26:27.819 --> 00:26:30.240
has been around a while. So anything we
needed
00:26:30.240 --> 00:26:33.980
to represent from logs or from traces or
metrics
00:26:33.980 --> 00:26:39.420
or profiles or other ways with graph views
and
00:26:39.420 --> 00:26:42.519
Mermaid diagrams, things like this, we
would
00:26:42.519 --> 00:26:44.920
ask the assistant to do that and to prefer
that.
00:26:45.099 --> 00:26:48.829
And so it's... It's really about, I think,
having
00:26:48.829 --> 00:26:51.789
people with excellent taste, kind of
designer
00:26:51.789 --> 00:26:56.970
engineers in one, really, having them
obsessed
00:26:56.970 --> 00:26:59.069
with that mission of we're going to fight
this
00:26:59.069 --> 00:27:03.069
wall of text. Because if, you know, a
picture
00:27:03.069 --> 00:27:04.950
tells a thousand words. So if we can show
you,
00:27:05.049 --> 00:27:07.849
here's what the data is doing, not only
can you
00:27:07.849 --> 00:27:10.750
see it as a user, but that is still better
as
00:27:10.750 --> 00:27:14.190
an experience for users than reading what
someone
00:27:14.190 --> 00:27:16.900
tells you about that graph. Someone might
say,
00:27:17.000 --> 00:27:20.920
oh, it's a flat graph. Is it flat but
high? Is
00:27:20.920 --> 00:27:24.220
it flat but low? You can describe it. It's
actually
00:27:24.220 --> 00:27:26.299
easier. Just show me the graph. This is
why I
00:27:26.299 --> 00:27:28.339
don't think dashboards are going to go
away in
00:27:28.339 --> 00:27:30.579
the AI world because we're always going to
care
00:27:30.579 --> 00:27:33.920
about what's true. And if the agents are
correct,
00:27:34.200 --> 00:27:37.519
we're going to want to see that data. So
the
00:27:37.519 --> 00:27:40.240
Grafana Assistant is very visually
beautiful
00:27:40.240 --> 00:27:42.880
for that reason because it piggybacks a
lot on
00:27:42.880 --> 00:27:48.440
Grafana. Yeah, it's important to, it's
going
00:27:48.440 --> 00:27:50.400
to make a big difference. And I think you
can
00:27:50.400 --> 00:27:52.720
separate yourself if you put a bit more
effort
00:27:52.720 --> 00:27:56.380
into making sure that that experience is
excellent.
00:27:56.859 --> 00:27:59.299
I think you can separate yourselves from
similar
00:27:59.299 --> 00:28:02.460
capabilities. It turns out for some
reason, and
00:28:02.460 --> 00:28:04.220
we have a few theories on this, but it
turns
00:28:04.220 --> 00:28:08.880
out that LLMs are quite good at Grafana by
default.
00:28:09.839 --> 00:28:11.839
And we think it's because of so much of
its open
00:28:11.839 --> 00:28:14.740
source. Grafana. The whole code base is
open
00:28:14.740 --> 00:28:17.079
source, but also all the community stuff,
all
00:28:17.079 --> 00:28:19.640
the discussions around there, like
community
00:28:19.640 --> 00:28:22.500
dashboards out there. And a lot of the
connected
00:28:22.500 --> 00:28:25.859
ecosystem, Prometheus and OpenTelemetry,
all
00:28:25.859 --> 00:28:28.299
these other things are also kind of very
open
00:28:28.299 --> 00:28:31.380
source. And I think the LLMs have hoovered
all
00:28:31.380 --> 00:28:33.259
of that up, so they naturally have quite
good
00:28:33.259 --> 00:28:36.859
context. So when we first stitched
together the
00:28:36.859 --> 00:28:40.009
LLM... with Grafana and gave it a few
tools and
00:28:40.009 --> 00:28:45.549
had the UI in the Grafana front end. It
was surprisingly
00:28:45.549 --> 00:28:47.869
already quite good. And this is what
people are
00:28:47.869 --> 00:28:49.769
finding when they wire up Claude Code with
our
00:28:49.769 --> 00:28:53.509
gcx, which is a CLI tool that we released.
That
00:28:53.509 --> 00:28:56.349
thing, you know, for certain tasks, it's
very
00:28:56.349 --> 00:28:59.190
good. And if you've got your code right
there
00:28:59.190 --> 00:29:01.589
as well, like it's very handy to be
working alongside
00:29:01.589 --> 00:29:04.759
with production code. with telemetry you
know
00:29:04.759 --> 00:29:06.740
and having the having them considered at
the
00:29:06.740 --> 00:29:10.200
same time but all the work we then did in
assistant
00:29:10.200 --> 00:29:12.559
with the prompting the fine tuning of
certain
00:29:12.559 --> 00:29:16.259
little pieces of it the specialist tools
and
00:29:16.259 --> 00:29:18.920
lots of other kind of secret sauce that's
gone
00:29:18.920 --> 00:29:23.490
in created this uh complexity that we have
to
00:29:23.490 --> 00:29:26.049
deliver in a simple way so i think if you
use
00:29:26.049 --> 00:29:28.609
assistant it'll look hopefully just looks
and
00:29:28.609 --> 00:29:31.309
feels very intuitive you know it does cool
things
00:29:31.309 --> 00:29:34.609
like it can build deep links so it knows
how
00:29:34.609 --> 00:29:36.970
to jump you to the right page in Grafana
but
00:29:36.970 --> 00:29:39.369
not just the right page because all the
your
00:29:39.369 --> 00:29:41.250
because all the filters are in the URL
parameters
00:29:41.250 --> 00:29:43.750
it will apply those filters as well so it
takes
00:29:43.750 --> 00:29:47.289
you straight to the right view um and so
that
00:29:47.289 --> 00:29:50.380
was another good thing i think we did was
We
00:29:50.380 --> 00:29:52.059
aren't going to reinvent everything and
give
00:29:52.059 --> 00:29:55.039
you an AI-only experience here. We are
going
00:29:55.039 --> 00:29:57.759
to use all the capabilities that you're
already
00:29:57.759 --> 00:30:00.099
familiar with and some of the amazing
tools that
00:30:00.099 --> 00:30:03.799
already exist in Grafana as part of this
too.
00:30:04.160 --> 00:30:08.259
And the fact that the assistant knows what
page
00:30:08.259 --> 00:30:10.319
you're looking at. Little things like
that. If
00:30:10.319 --> 00:30:13.099
you go to a page and ask and say, tell me
about
00:30:13.099 --> 00:30:15.750
this, it knows what you're talking about.
So
00:30:15.750 --> 00:30:18.109
that really matters. Those kinds of little
details,
00:30:18.150 --> 00:30:20.849
that experience that the user's having, I
think
00:30:20.849 --> 00:30:23.089
is what makes it feel like excellent,
helps you
00:30:23.089 --> 00:30:25.970
trust it. You see the data itself as well.
So
00:30:25.970 --> 00:30:30.069
you can see if it's made a mistake. And I
think
00:30:30.069 --> 00:30:34.069
UX is more important than ever. Very true.
Okay,
00:30:34.150 --> 00:30:38.670
so for a team that is starting out wanting
to
00:30:38.670 --> 00:30:42.890
introduce AI, where... Where do you think
AI
00:30:42.890 --> 00:30:45.450
can actually help operations teams today?
Like
00:30:45.450 --> 00:30:47.529
where should they start? What's the most
important
00:30:47.529 --> 00:30:53.309
telemetry or piece of data to start using
to
00:30:53.309 --> 00:30:56.650
analyze? Well, it will depend probably on
each
00:30:56.650 --> 00:30:59.269
of their cases. The cool thing is they
will know,
00:30:59.450 --> 00:31:02.349
they'll have an instinct already for
things that
00:31:02.349 --> 00:31:04.650
they already care about. So the thing is,
the
00:31:04.650 --> 00:31:08.150
nice thing is probably people already are
using
00:31:08.150 --> 00:31:10.490
some kind of telemetry. They're using some
kind
00:31:10.490 --> 00:31:13.599
of observability. And really, AI is there
to
00:31:13.599 --> 00:31:17.220
enhance that experience. It's there to
make that
00:31:17.220 --> 00:31:19.759
easier to do. Since it writes the queries
for
00:31:19.759 --> 00:31:22.200
you, you don't have to learn PromQL and
LogQL.
00:31:22.799 --> 00:31:26.799
And I've used those languages before, of
course,
00:31:26.920 --> 00:31:29.140
but I would always have to go and look up
again
00:31:29.140 --> 00:31:32.200
how to do something or prefer the query
builder,
00:31:32.259 --> 00:31:35.460
honestly. Now you don't even need to do it
at
00:31:35.460 --> 00:31:37.559
all because genuinely, through natural
language,
00:31:37.660 --> 00:31:39.819
it can generate those things. And it
generates
00:31:39.819 --> 00:31:42.700
very complex ones as well. It's far out,
you
00:31:42.700 --> 00:31:45.400
know, surpassed me, my abilities. But I
would
00:31:45.400 --> 00:31:48.440
say like, use it initially to enhance what
you're
00:31:48.440 --> 00:31:53.559
already doing and see ways that you can
automate.
00:31:53.740 --> 00:31:56.119
In some ways, that's really what we're
doing.
00:31:56.160 --> 00:31:58.099
It's just automating it, but automating it
in
00:31:58.099 --> 00:32:00.359
a trustworthy way. And I would say from
those
00:32:00.359 --> 00:32:03.140
few little early things and do stuff, ship
it,
00:32:03.200 --> 00:32:05.519
Ship It Weekly. I like that actually as a
motto,
00:32:05.660 --> 00:32:09.079
Ship It Weekly. Not just the fact that
this podcast
00:32:09.079 --> 00:32:12.680
runs every week. But I actually ship
things every
00:32:12.680 --> 00:32:17.099
week. I like that because you're kind of
forced
00:32:17.099 --> 00:32:20.220
to focus on scope. Well, the scope that's
most
00:32:20.220 --> 00:32:22.619
important. You know, if you have a short
window
00:32:22.619 --> 00:32:25.700
for shipping things, then you've got to
really
00:32:25.700 --> 00:32:28.039
focus on what's important. And I think
that applies
00:32:28.039 --> 00:32:30.180
to people picking up AI things. Don't do
some
00:32:30.180 --> 00:32:33.099
enormous AI project that's going to take
months
00:32:33.099 --> 00:32:35.000
and months. Because honestly, everything's
different
00:32:35.000 --> 00:32:38.000
by then anyway. Pick something small, do
it,
00:32:38.059 --> 00:32:41.000
action it, and talk about it and share it.
and
00:32:41.000 --> 00:32:43.440
see what see if there's excitement you can
drum
00:32:43.440 --> 00:32:45.700
up it's what happened at Grafana Labs and
then
00:32:45.700 --> 00:32:48.099
now you see people all over the company
building
00:32:48.099 --> 00:32:50.480
all sorts of things and doing amazing
things
00:32:50.480 --> 00:32:53.900
with agents some of it will eventually be
a product
00:32:53.900 --> 00:32:57.279
i'm sure but it's it comes from play and
that's
00:32:57.279 --> 00:32:59.200
something that we do all right at Grafana
Labs
00:32:59.200 --> 00:33:03.299
we we have a lot of space for people to
play
00:33:03.299 --> 00:33:05.960
and get creative and try things and fail
for
00:33:05.960 --> 00:33:10.160
things to fail um there's a great analogy
that
00:33:10.160 --> 00:33:13.160
i love where they they did a study where
they
00:33:13.160 --> 00:33:15.059
got these two groups and they told they
were
00:33:15.059 --> 00:33:17.140
making pottery you know on the pottery
wheel
00:33:17.140 --> 00:33:19.500
but they were all amateurs they told one
group
00:33:19.500 --> 00:33:22.299
you've got to make the best pot you can
and they
00:33:22.299 --> 00:33:24.660
told the other group just just make as
many pots
00:33:24.660 --> 00:33:26.859
as you can doesn't matter what they're
like just
00:33:26.859 --> 00:33:29.259
make lots of them and then by the end of
it the
00:33:29.259 --> 00:33:32.859
the group that were making as many pots as
they
00:33:32.859 --> 00:33:35.240
could were making better pots than the
group
00:33:35.240 --> 00:33:38.180
that was trying to make the best pot and i
think
00:33:38.730 --> 00:33:40.950
That tells you a lot, I think. It's the
trial
00:33:40.950 --> 00:33:42.809
and error that you learn. You know, there
was
00:33:42.809 --> 00:33:45.710
a lot more waste on the floor of this
other team
00:33:45.710 --> 00:33:48.210
where pots hadn't worked and they had to
throw
00:33:48.210 --> 00:33:51.789
it away or smash it up and tear it down
and turn
00:33:51.789 --> 00:33:55.410
it into something else. So it feels
wasteful,
00:33:55.410 --> 00:33:58.289
I think, to some. And they might be
tempted to
00:33:58.289 --> 00:34:01.329
kind of optimize that away. But honestly,
that's
00:34:01.329 --> 00:34:04.029
often where the learning is. So I think
people
00:34:04.029 --> 00:34:07.819
should pick up something, tackle it.
something
00:34:07.819 --> 00:34:11.500
small have a play see what you can do and
from
00:34:11.500 --> 00:34:14.500
there yeah you create new problems for
yourself
00:34:14.500 --> 00:34:17.800
but they're good problems to have where
where
00:34:17.800 --> 00:34:20.619
do you think AI is going with
observability this
00:34:20.619 --> 00:34:22.699
year i mean obviously it's it's the talk
at every
00:34:22.699 --> 00:34:29.039
conference but is there overarching areas
in
00:34:29.039 --> 00:34:32.960
observability that you think that AI um
that
00:34:32.960 --> 00:34:35.000
we're just at the forefront like we're
able to
00:34:35.000 --> 00:34:39.690
do now with larger token uh more tokens
more
00:34:39.690 --> 00:34:44.730
uh better models um faster models is there
anything
00:34:44.730 --> 00:34:47.150
that like lends itself now that maybe we
couldn't
00:34:47.150 --> 00:34:49.869
have done six months ago or we could do in
six
00:34:49.869 --> 00:34:53.130
months um based on these these rapid
changes
00:34:53.130 --> 00:34:58.510
to AI i think seeing seeing the models get
better
00:34:58.510 --> 00:35:01.849
and get smarter is nice we are finding
that we
00:35:01.849 --> 00:35:04.630
have to tell it less which is interesting
we're
00:35:04.630 --> 00:35:08.570
also finding that we We are basically
tuning
00:35:08.570 --> 00:35:12.670
our prompts to a specific model. We
noticed with
00:35:12.670 --> 00:35:16.989
Claude 3.5 to 3.7 even, 3.7 didn't perform
00:35:16.989 --> 00:35:19.670
as well as 3.5 by default. We had to make
it
00:35:19.670 --> 00:35:21.429
work by changing it and tweaking it
because it
00:35:21.429 --> 00:35:25.210
was a kind of different thing. And so I
think
00:35:25.210 --> 00:35:30.210
we'll get to the point where agents are,
yeah,
00:35:30.269 --> 00:35:32.889
you have to tell them a lot less. I think
their
00:35:32.889 --> 00:35:36.130
instincts would be good. We already have
plenty
00:35:36.130 --> 00:35:40.050
of tools to help with skills and things.
In Assistant,
00:35:40.250 --> 00:35:42.449
you can write skills, which lets you
basically
00:35:42.449 --> 00:35:45.070
control how the agents behave. This is
nice if
00:35:45.070 --> 00:35:47.590
you've got expert SREs where they've spent
years
00:35:47.590 --> 00:35:50.489
building these patterns and practices and
they've
00:35:50.489 --> 00:35:53.070
got these high standards. You can then
give those
00:35:53.070 --> 00:35:54.710
same high standards to the agents and
they'll
00:35:54.710 --> 00:35:57.469
follow your behaviors. So you can already
customize
00:35:57.469 --> 00:36:03.130
and tune it like that. But I think...
that's
00:36:03.130 --> 00:36:05.130
going to just improve. And I sort of
expect that
00:36:05.130 --> 00:36:08.969
to happen. Bigger context windows, not
necessarily
00:36:08.969 --> 00:36:13.989
seeing that improve things too much. We
already
00:36:13.989 --> 00:36:17.670
have kind of too much context. A big
challenge
00:36:17.670 --> 00:36:20.230
that we're doing at Grafana Labs is, and
with
00:36:20.230 --> 00:36:22.989
Assistant when it does an investigation,
it's
00:36:22.989 --> 00:36:25.349
about going and finding the right context
and
00:36:25.349 --> 00:36:27.269
keeping the right context and discarding
the
00:36:27.269 --> 00:36:30.360
things that's not right. So big context
windows,
00:36:30.500 --> 00:36:32.639
kind of they're tempting. You just think,
oh,
00:36:32.679 --> 00:36:34.079
we could just fill it all up with
everything.
00:36:34.280 --> 00:36:35.840
I don't think that's going to help it at
all.
00:36:36.199 --> 00:36:40.460
It's just more for it to think about. But
I do
00:36:40.460 --> 00:36:43.900
think we are going to see more things
automated.
00:36:44.239 --> 00:36:49.500
And one example I think is an easy one is
rolling
00:36:49.500 --> 00:36:53.440
back a deploy. Let's say you kick off a
deploy
00:36:53.440 --> 00:36:56.000
through GitHub, whatever your process is,
at
00:36:56.000 --> 00:36:58.929
the end, once it's out, Maybe, you know,
there's
00:36:58.929 --> 00:37:01.630
some, you do this in staging and things,
whatever
00:37:01.630 --> 00:37:04.550
you already have, all the great stuff. In
production,
00:37:04.710 --> 00:37:07.550
have agents actually go and do a series of
kind
00:37:07.550 --> 00:37:10.530
of critical tests, check things, check all
the
00:37:10.530 --> 00:37:13.489
dashboards, you know, they know how to use
Grafana,
00:37:13.590 --> 00:37:17.269
so they know how to use all your
telemetry. So
00:37:17.269 --> 00:37:22.170
how can we get all of that and make good
decisions
00:37:22.170 --> 00:37:27.150
from it? And then a decision might be
that. that
00:37:27.150 --> 00:37:28.969
deploys not i'm not happy with it for
whatever
00:37:28.969 --> 00:37:31.130
reason and i'm i as the agent i'm just
going
00:37:31.130 --> 00:37:34.230
to decide to roll it roll it back so i
think
00:37:34.230 --> 00:37:36.250
that's a safe quite a safe operation
because
00:37:36.250 --> 00:37:38.610
it's you should always be able to roll
back a
00:37:38.610 --> 00:37:41.329
release that should be a safe operation um
it's
00:37:41.329 --> 00:37:43.130
not always so it's you know it's not
without
00:37:43.130 --> 00:37:47.809
risk but um yeah that feels like something
like
00:37:47.809 --> 00:37:48.969
that i think we're going to get we're
going to
00:37:48.969 --> 00:37:50.469
see more and more of that and then i want
to
00:37:50.469 --> 00:37:52.510
go further what else are we happy to let
it do
00:37:52.510 --> 00:37:55.190
can it can it deal with scale issues based
on
00:37:56.750 --> 00:38:00.230
unexpected events can it um what else can
it
00:38:00.230 --> 00:38:02.630
do where else can it take action and i
think
00:38:02.630 --> 00:38:06.269
uh we need to get a bit braver sometimes
uh on
00:38:06.269 --> 00:38:08.409
some of these things it's on us to build
the
00:38:08.409 --> 00:38:11.289
tools and make sure that they are behaving
well
00:38:11.289 --> 00:38:14.550
and you can see what they're doing and you
know
00:38:14.550 --> 00:38:16.730
people will some people will always have
the
00:38:16.730 --> 00:38:19.190
the human in the loop to give the approval
and
00:38:19.190 --> 00:38:21.349
hit the button and i actually think that's
also
00:38:21.349 --> 00:38:25.230
fine but but all the work that leads up to
that
00:38:25.230 --> 00:38:27.829
point Because there's still a big space
there
00:38:27.829 --> 00:38:30.250
to explore. And we are exploring it. It is
very
00:38:30.250 --> 00:38:32.590
exciting, some of the things that can
happen.
00:38:33.230 --> 00:38:35.510
But I do think we're going to have agents
doing
00:38:35.510 --> 00:38:39.150
more of this operating things for us. At
least
00:38:39.150 --> 00:38:41.389
the simple stuff. And that frees us up to
then
00:38:41.389 --> 00:38:43.489
focus on the cases that aren't simple and
that
00:38:43.489 --> 00:38:47.690
are more complicated. Or frees us up to
build
00:38:47.690 --> 00:38:50.989
more things and do more other things. Very
true.
00:38:51.639 --> 00:38:54.199
Okay, so wrapping up, is there any belief
about
00:38:54.199 --> 00:38:56.760
AI or observability that you think is just
wrong?
00:38:57.000 --> 00:39:02.340
I think one thing that's wrong is because
it's
00:39:02.340 --> 00:39:05.559
not 100%, you therefore can't trust it. As
we
00:39:05.559 --> 00:39:10.760
mentioned earlier, yes, it's not perfect.
It
00:39:10.760 --> 00:39:13.460
gets things wrong. So do people, but you
still
00:39:13.460 --> 00:39:19.750
hire people, hopefully, for now. So I
think that,
00:39:19.829 --> 00:39:23.369
I think, but that's not just AI. I think
people
00:39:23.369 --> 00:39:27.250
often struggle with that. I remember, I
remember
00:39:27.250 --> 00:39:30.349
like in COVID times, I was wearing a mask
and
00:39:30.349 --> 00:39:33.150
a guy in the elevators like thought I was
silly.
00:39:33.250 --> 00:39:35.690
Like I thought I just believed some hype
or whatever.
00:39:36.369 --> 00:39:38.429
So he's like, why are you wearing that?
They're
00:39:38.429 --> 00:39:40.690
not even effective. They only like, only
prevents
00:39:40.690 --> 00:39:44.230
20% of spread or something. So it's like,
oh,
00:39:44.329 --> 00:39:47.429
is it? So it's only 20% fewer people going
to
00:39:47.429 --> 00:39:51.849
be. But because it wasn't 100%, he
dismissed
00:39:51.849 --> 00:39:54.710
it. I see people doing the same thing with
AI.
00:39:56.550 --> 00:40:03.489
Or they think with coding, it's not as
good as
00:40:03.489 --> 00:40:05.150
me because I prompted it, asked it to do
something,
00:40:05.210 --> 00:40:06.750
it didn't do it as well as I would have
done
00:40:06.750 --> 00:40:09.550
it. And so I did that last year and I'm
not going
00:40:09.550 --> 00:40:13.030
to touch it. AI is not for me. That's also
a
00:40:13.030 --> 00:40:15.789
mistake. Keep trying them. It's usually
you haven't
00:40:15.789 --> 00:40:17.630
prompted it well enough if it's not
working.
00:40:17.980 --> 00:40:19.699
You can prompt it differently and there's
different
00:40:19.699 --> 00:40:22.920
things you can try. And you can also be
much
00:40:22.920 --> 00:40:25.159
more specific. In fact, the more context
you
00:40:25.159 --> 00:40:28.400
give it, the better. So, yeah, it's not
100%.
00:40:28.400 --> 00:40:32.659
It's not perfect, but it's good. It's
changing
00:40:32.659 --> 00:40:35.940
the world. Yeah, for sure. Mat, any
closing
00:40:35.940 --> 00:40:38.960
thoughts for our audience? No, but I just
want
00:40:38.960 --> 00:40:42.739
to say keep shipping it weekly because,
you know,
00:40:42.880 --> 00:40:44.840
that's the only real way to know you're
doing
00:40:44.840 --> 00:40:48.420
anything useful. Iterate, yeah. And keep
listening
00:40:48.420 --> 00:40:51.400
to this great podcast as well with Brian.
Appreciate
00:40:51.400 --> 00:40:53.320
it. Thanks, Mat. Thank you so much for
your
00:40:53.320 --> 00:40:54.699
time. Thank you for coming on. Really,
really
00:40:54.699 --> 00:40:57.719
appreciate it. Pleasure. All right. That
was
00:40:57.719 --> 00:41:00.639
my conversation with Mat Ryer from Grafana
00:41:00.639 --> 00:41:03.599
Labs. The thing that stuck with me the
most is
00:41:03.599 --> 00:41:06.940
that AI in production has to grow up
pretty fast.
00:41:07.199 --> 00:41:10.920
It is fun when it is a prototype. It is
fun when
00:41:10.920 --> 00:41:13.949
someone wires up a model. gives it a few
tools
00:41:13.949 --> 00:41:17.250
and suddenly it can answer questions that
used
00:41:17.250 --> 00:41:21.269
to require bouncing between dashboards
logs traces
00:41:21.269 --> 00:41:25.389
docs and half a dozen slack threads that
part
00:41:25.389 --> 00:41:29.010
is genuinely cool but once people start
depending
00:41:29.010 --> 00:41:32.829
on it the bar changes now it needs
observability
00:41:32.829 --> 00:41:37.650
it needs evals it needs feedback loops it
needs
00:41:37.650 --> 00:41:41.780
cost awareness It needs guardrails. It
needs
00:41:41.780 --> 00:41:45.179
some way to tell whether the agent
actually helped
00:41:45.179 --> 00:41:48.480
or just produced a convincing answer. And
that
00:41:48.480 --> 00:41:51.340
is where I think this conversation gets
useful
00:41:51.340 --> 00:41:55.179
for platform and SRE teams. Because the
trap
00:41:55.179 --> 00:41:58.579
is thinking AI observability just means
watching
00:41:58.579 --> 00:42:02.039
the AI service like any other service. Did
it
00:42:02.039 --> 00:42:05.420
return a 200? How long did it take? How
much
00:42:05.420 --> 00:42:08.550
did it cost? Those are good signals. but
they
00:42:08.550 --> 00:42:11.969
are not enough. With agents, you care
about the
00:42:11.969 --> 00:42:15.570
path it took, the tools it used, the
context
00:42:15.570 --> 00:42:19.010
it kept, the context it threw away,
whether it
00:42:19.010 --> 00:42:21.809
answered the actual user request, whether
the
00:42:21.809 --> 00:42:24.449
user had to correct it, whether the model
changed
00:42:24.449 --> 00:42:28.449
behavior after a prompt update or model
upgrade.
00:42:28.670 --> 00:42:31.690
That is a much messier kind of production
system.
00:42:31.969 --> 00:42:35.449
I also liked Mat's take on trust. A lot of
people
00:42:35.449 --> 00:42:39.550
treat AI like it has to be perfect or it
is useless.
00:42:39.789 --> 00:42:43.010
And I get the instinct, especially when we
are
00:42:43.010 --> 00:42:46.170
talking about production systems. But we
already
00:42:46.170 --> 00:42:50.289
operate plenty of imperfect systems.
Networks
00:42:50.289 --> 00:42:55.150
fail. APIs time out. Humans miss things.
Dashboards
00:42:55.150 --> 00:42:58.989
lie by omission. Runbooks rot. The answer
is
00:42:58.989 --> 00:43:02.869
not blind trust, but it is also not
refusing
00:43:02.869 --> 00:43:07.360
to use the thing because it is not 100%.
The
00:43:07.360 --> 00:43:10.880
answer is instrumentation, feedback,
constraints,
00:43:11.360 --> 00:43:14.599
and judgment. It's the same as most other
production
00:43:14.599 --> 00:43:17.380
problems, honestly. The other big takeaway
for
00:43:17.380 --> 00:43:21.860
me was the UX side. AI should not just
bury operators
00:43:21.860 --> 00:43:25.500
in a wall of generated text. If the system
can
00:43:25.500 --> 00:43:29.360
show the graph, link to the right view,
apply the
00:43:29.360 --> 00:43:32.500
right filters, and expose the evidence
behind
00:43:32.500 --> 00:43:36.190
the answer, That is way more useful than a
paragraph
00:43:36.190 --> 00:43:38.630
that sounds right. Because at the end of
the
00:43:38.630 --> 00:43:41.869
day, operators still want to see what is
true,
00:43:42.070 --> 00:43:45.269
not just what the model said. So my
takeaway
00:43:45.269 --> 00:43:49.550
is pretty simple. Use AI to make the work
easier.
00:43:49.750 --> 00:43:52.750
Let it help search through telemetry. Let
it
00:43:52.750 --> 00:43:56.289
summarize patterns. Let it assist with the
first
00:43:56.289 --> 00:44:00.090
pass of an investigation. But do not skip
the
00:44:00.090 --> 00:44:04.420
production discipline. Measure it. Test
it. Watch
00:44:04.420 --> 00:44:08.659
the cost. Build evals. Keep humans in the
loop
00:44:08.659 --> 00:44:12.440
where the blast radius is real. And when
you
00:44:12.440 --> 00:44:15.659
give an agent tools, remember that it may
actually
00:44:15.659 --> 00:44:18.980
use them. That sounds obvious, but it is
going
00:44:18.980 --> 00:44:22.559
to matter a lot. I'll have links to Mat,
Grafana
00:44:22.559 --> 00:44:25.679
Labs, and anything else we mentioned in
the show
00:44:25.679 --> 00:44:29.099
notes. If you enjoyed this conversation,
Follow
00:44:29.099 --> 00:44:32.059
or subscribe to Ship It Weekly wherever
you listen
00:44:32.059 --> 00:44:35.380
to podcasts. It helps the show and it
makes sure
00:44:35.380 --> 00:44:38.380
you get both these conversation episodes
and
00:44:38.380 --> 00:44:42.280
the weekly DevOps, SRE, platform, cloud,
and
00:44:42.280 --> 00:44:45.159
security news recaps. You can also find
everything
00:44:45.159 --> 00:44:48.440
over at shipitweekly.fm. Thanks for
listening
00:44:48.440 --> 00:44:50.380
and I'll see you later this week.