WEBVTT
00:00:00.000 --> 00:00:02.580
Welcome to this deep dive. You know, our mission
00:00:02.580 --> 00:00:04.740
today is really about cutting through the massive
00:00:04.740 --> 00:00:08.679
amount of noise in AI evaluation. Yeah, and there
00:00:08.679 --> 00:00:11.199
is a lot of noise out there right now. It really
00:00:11.199 --> 00:00:13.900
is. We're doing that by unpacking this fascinating
00:00:13.900 --> 00:00:17.300
October 2026 archive paper. It was by Jingji
00:00:17.300 --> 00:00:20.899
Ning, Xueqi Li, and Yibo Kang. Right, the one
00:00:20.899 --> 00:00:24.160
titled, Agentic RAG Evaluation, Budget Allocation
00:00:24.160 --> 00:00:27.100
Across Questions, Trajectories, and Reads. Exactly.
00:00:27.280 --> 00:00:29.519
And the core tension they're exploring is something,
00:00:29.640 --> 00:00:32.520
well, pretty much every AI developer is agonizing
00:00:32.520 --> 00:00:35.859
over right now. Companies are burning staggering
00:00:35.859 --> 00:00:38.600
amounts of money testing their AI agents. Oh,
00:00:38.640 --> 00:00:40.820
absolutely. I mean, we're talking millions of
00:00:40.820 --> 00:00:44.039
tokens, which translates directly to hard dollars.
00:00:44.159 --> 00:00:46.640
And there's this very real possibility that they're
00:00:46.640 --> 00:00:49.579
completely wasting those budgets by measuring
00:00:49.579 --> 00:00:52.539
the exact wrong things. Yeah, that's the scary
00:00:52.539 --> 00:00:55.280
part for a lot of teams. So, OK, let's unpack
00:00:55.280 --> 00:00:57.590
this. Think about this business problem like
00:00:57.590 --> 00:01:00.390
trying to raid a brand new restaurant in your
00:01:00.390 --> 00:01:02.509
neighborhood. Okay, I like that analogy. Right,
00:01:02.590 --> 00:01:04.769
because you assume there's a certain level of
00:01:04.769 --> 00:01:07.709
objective truth to be found. You go in, you eat
00:01:07.709 --> 00:01:09.829
the food, and you decide if it's good or bad.
00:01:10.069 --> 00:01:12.409
Right, simple enough. But imagine if that experience
00:01:12.409 --> 00:01:15.329
was just totally chaotic, like you order the
00:01:15.329 --> 00:01:17.269
spaghetti. But the chef cooks it differently
00:01:17.269 --> 00:01:19.090
every single time you order it. That would be
00:01:19.090 --> 00:01:20.969
frustrating, yeah. And to make matters worse,
00:01:21.170 --> 00:01:23.590
your own taste buds randomly change depending
00:01:23.590 --> 00:01:26.049
on what day of the week it is. Suddenly, trying
00:01:26.049 --> 00:01:27.950
to give that restaurant a definitive rating becomes
00:01:27.950 --> 00:01:31.909
this massive, expensive headache. That evaluative
00:01:31.909 --> 00:01:34.489
muddy water is exactly the landscape development
00:01:34.489 --> 00:01:37.489
teams face when testing agentic rag systems today.
00:01:37.689 --> 00:01:40.319
It really is. Because, you know, we aren't dealing
00:01:40.319 --> 00:01:42.680
with static data sets anymore. When a system
00:01:42.680 --> 00:01:44.760
actively goes out and searches for information
00:01:44.760 --> 00:01:47.599
to build its answer, evaluating its capability
00:01:47.599 --> 00:01:50.700
introduces three distinct layers of variance.
00:01:51.060 --> 00:01:53.900
Okay, break those down for us. So first you have
00:01:53.900 --> 00:01:56.739
the question variance, just the differences in
00:01:56.739 --> 00:01:59.750
the prompts themselves. Right. Then, if you ask
00:01:59.750 --> 00:02:02.510
the agent the exact same question twice, it might
00:02:02.510 --> 00:02:04.709
follow a completely different search trajectory,
00:02:04.810 --> 00:02:07.650
like it pulls different sources or maybe pulls
00:02:07.650 --> 00:02:10.030
them in a different order. So that's the chef
00:02:10.030 --> 00:02:11.810
cooking the spaghetti differently every time.
00:02:12.169 --> 00:02:14.830
Exactly. And finally, you have the evaluation
00:02:14.830 --> 00:02:18.229
itself. This is usually performed by an AI judge
00:02:18.229 --> 00:02:20.789
grading the final answer and that search trajectory.
00:02:21.050 --> 00:02:23.310
And that would be the unpredictable taste buds.
00:02:23.509 --> 00:02:26.909
Spot on. So the industry dilemma is that development
00:02:26.909 --> 00:02:29.710
teams have a fixed token budget, but they don't
00:02:29.710 --> 00:02:31.810
know how to spend it. Do they ask the AI new
00:02:31.810 --> 00:02:34.110
questions? Right. Or do they repeat the same
00:02:34.110 --> 00:02:36.389
question to map out different search trajectories?
00:02:36.530 --> 00:02:38.909
Or do they have the AI judge read the final answer
00:02:38.909 --> 00:02:41.189
multiple times to ensure the grade is consistent?
00:02:41.330 --> 00:02:43.469
They just don't know. It's a massive resource
00:02:43.469 --> 00:02:46.069
allocation problem. Like if you have a budget
00:02:46.069 --> 00:02:48.430
of $500 to figure out if that restaurant is good,
00:02:48.569 --> 00:02:51.569
how do you spend it? Yeah, exactly. Do you order
00:02:51.569 --> 00:02:54.509
50 different dishes? Or order the exact same
00:02:54.509 --> 00:02:57.810
spaghetti 10 times? Or do you buy one dish and
00:02:57.810 --> 00:03:02.370
have five friends taste it? And without any mathematical
00:03:02.370 --> 00:03:05.129
guidance, most teams just default to repeated
00:03:05.129 --> 00:03:07.590
runs. They just brute force it. Right. They test
00:03:07.590 --> 00:03:09.689
the same questions and have the judge do multiple
00:03:09.689 --> 00:03:12.330
reads, basically hoping an average will give
00:03:12.330 --> 00:03:15.069
them a reliable baseline. But that doesn't actually
00:03:15.069 --> 00:03:17.469
work, does it? No. Doing that simply ends up
00:03:17.469 --> 00:03:20.509
remeasuring noise. They're burning token budgets
00:03:20.509 --> 00:03:22.469
without getting a clearer picture of how capable
00:03:22.469 --> 00:03:25.030
the agent truly is. So how do we actually fix
00:03:25.030 --> 00:03:27.110
this? How do we find the mathematical truth here?
00:03:27.509 --> 00:03:29.229
Well, to solve this, the researchers applied
00:03:29.229 --> 00:03:32.530
a robust statistical framework called generalizability
00:03:32.530 --> 00:03:35.330
theory. Wait, isn't that from like the 1970s?
00:03:35.370 --> 00:03:38.009
It is, yeah. It was originally developed for
00:03:38.009 --> 00:03:40.449
psychological and educational testing. You'll
00:03:40.449 --> 00:03:43.250
often hear it associated with Kronbach. Oh, interesting.
00:03:43.490 --> 00:03:46.370
So how does a 1970s psychology framework help
00:03:46.370 --> 00:03:49.759
with AI tokens? So... Generalizability theory
00:03:49.759 --> 00:03:52.120
is designed to split measurement variance into
00:03:52.120 --> 00:03:55.039
specific facets. Instead of treating an inconsistent
00:03:55.039 --> 00:03:58.979
AI score as just one big lump of unexplainable
00:03:58.979 --> 00:04:01.520
error. Which is what everyone does now. Right.
00:04:01.860 --> 00:04:04.639
Instead of that, this framework isolates the
00:04:04.639 --> 00:04:07.379
noise. Okay, so it's essentially taking a scalpel
00:04:07.379 --> 00:04:11.099
to the variance. You can mathematically say this
00:04:11.099 --> 00:04:13.099
percentage of the fluctuation is caused by our
00:04:13.099 --> 00:04:16.180
question set being too diverse. Exactly. And...
00:04:16.709 --> 00:04:19.230
This percentage is from the agent taking weird
00:04:19.230 --> 00:04:22.689
search paths. And this percentage is just our
00:04:22.689 --> 00:04:25.970
AI judge being moody. Exactly. And that breakdown
00:04:25.970 --> 00:04:27.910
allows you to model and predict the uncertainty
00:04:27.910 --> 00:04:30.910
of your evaluation, which in this paper is quantified
00:04:30.910 --> 00:04:33.329
as the standard error. And just to clarify, a
00:04:33.329 --> 00:04:35.290
smaller standard error means a tighter, more
00:04:35.290 --> 00:04:38.459
reliable evaluation, right? Correct. So to figure
00:04:38.459 --> 00:04:40.699
out how to minimize that standard error, the
00:04:40.699 --> 00:04:43.519
researchers used repeated sampling on two standard
00:04:43.519 --> 00:04:46.980
multi -hop QA benchmarks, Hot Pot QA and Music
00:04:46.980 --> 00:04:49.139
Q. Okay, those are pretty standard in the industry?
00:04:49.519 --> 00:04:52.779
Yeah, very standard. So they systematically varied
00:04:52.779 --> 00:04:55.079
the number of questions, the trajectories, and
00:04:55.079 --> 00:04:57.519
the reads, and they recorded the variance at
00:04:57.519 --> 00:04:59.540
every step. So they just mapped everything out?
00:04:59.759 --> 00:05:02.459
They did. And from all that data, they built
00:05:02.459 --> 00:05:05.089
predictive models for budget allocation. Get
00:05:05.089 --> 00:05:08.230
this, it achieved an incredibly impressive 3
00:05:08.230 --> 00:05:11.709
.5 % to 4 .0 % accuracy in predicting the standard
00:05:11.709 --> 00:05:14.069
error based on how tokens were spent. Wow, that's
00:05:14.069 --> 00:05:17.410
really tight. But I have to push back on one
00:05:17.410 --> 00:05:19.949
thing here. Sure, go for it. I'm struggling to
00:05:19.949 --> 00:05:22.110
buy that the AI judge is actually a significant
00:05:22.110 --> 00:05:24.829
source of that noise. I mean, we work with these
00:05:24.829 --> 00:05:27.430
frontier models daily. They don't just wildly
00:05:27.430 --> 00:05:30.529
hallucinate their grading criteria on the exact
00:05:30.529 --> 00:05:33.290
same text upon a second read, do they? You would
00:05:33.290 --> 00:05:35.629
think so, but the researchers actually documented
00:05:35.629 --> 00:05:39.410
a massive 14 .3 % disagreement rate. Wait, really?
00:05:39.610 --> 00:05:44.160
14 %? Yeah. When the AI judge evaluated the exact
00:05:44.160 --> 00:05:47.100
same answer twice using standard default parameters,
00:05:47.379 --> 00:05:50.959
it contradicted its own previous grade over 14
00:05:50.959 --> 00:05:54.319
% of the time. That is insane. So doing multiple
00:05:54.319 --> 00:05:57.680
reads really would eat up a lot of the budget
00:05:57.680 --> 00:06:01.040
trying to fix that. But here is the kicker. No.
00:06:01.569 --> 00:06:04.430
They uncovered a highly practical fix for it.
00:06:04.550 --> 00:06:07.550
Okay, I'm listening. By simply dropping the judge's
00:06:07.550 --> 00:06:10.009
temperature setting to zero, basically forcing
00:06:10.009 --> 00:06:13.129
determinism, that disagreement rate plummeted
00:06:13.129 --> 00:06:17.089
down to just 3 .4%. Oh, wow. Yeah. Okay, so just
00:06:17.089 --> 00:06:19.110
by neutralizing the model's creativity dial,
00:06:19.350 --> 00:06:21.870
they almost entirely eliminated the read variance.
00:06:22.110 --> 00:06:25.329
Exactly. The judge stops changing its mind. Which
00:06:25.329 --> 00:06:27.370
means you solve that specific facet of noise
00:06:27.370 --> 00:06:30.149
with a basic settings tweak rather than throwing
00:06:30.149 --> 00:06:32.790
more of your actual token budget at it. And that
00:06:32.790 --> 00:06:35.009
fundamentally alters the math for your budget
00:06:35.009 --> 00:06:37.410
allocation. You immediately know that spending
00:06:37.410 --> 00:06:40.170
tokens, having a judge reread the same output,
00:06:40.250 --> 00:06:42.910
is economically inefficient. Right, because a
00:06:42.910 --> 00:06:45.089
simple parameter change gives you the stability
00:06:45.089 --> 00:06:48.470
you need. Exactly. Okay, so armed with that knowledge
00:06:48.470 --> 00:06:51.589
of how the variance splits, what does this actually
00:06:51.589 --> 00:06:54.649
change for a team building AI today? Let's talk
00:06:54.649 --> 00:06:56.670
about the hard numbers. Let's do it. If I'm managing
00:06:56.670 --> 00:06:58.750
a development team today and I have a budget
00:06:58.750 --> 00:07:01.769
of, say, 34 million model tokens sitting in my
00:07:01.769 --> 00:07:04.610
cloud account, how does this paper tell me to
00:07:04.610 --> 00:07:07.550
actually deploy them? Well, expanding the number
00:07:07.550 --> 00:07:09.949
of questions is the ultimate winner. The paper
00:07:09.949 --> 00:07:12.529
provides concrete results for that exact 34 million
00:07:12.529 --> 00:07:14.670
token scenario. Okay, so what's the breakdown?
00:07:15.370 --> 00:07:18.029
Going wide with your question coverage lowers
00:07:18.029 --> 00:07:21.069
the standard error by a massive 33 % compared
00:07:21.069 --> 00:07:23.879
to a strategy of doing five repeated reads. A
00:07:23.879 --> 00:07:26.860
33 % drop in uncertainty just by expanding the
00:07:26.860 --> 00:07:29.920
prompt pool. That's definitive. It really is.
00:07:30.279 --> 00:07:32.720
And expanding the question set also lowers the
00:07:32.720 --> 00:07:35.660
standard error by 12 .6 % compared to asking
00:07:35.660 --> 00:07:37.480
the same question and mapping three different
00:07:37.480 --> 00:07:40.240
search trajectories. So the mathematical truth
00:07:40.240 --> 00:07:43.620
just points in one direction. Breath beats depth.
00:07:43.879 --> 00:07:46.040
Yep. When you're trying to find the true capability
00:07:46.040 --> 00:07:48.639
of the system, you go wide. But what about the
00:07:48.639 --> 00:07:51.509
cost analysis across different models? Because
00:07:51.509 --> 00:07:55.110
we know API price tiers vary wildly, right? Between
00:07:55.110 --> 00:07:57.610
a cheaper, faster model and like a heavy state
00:07:57.610 --> 00:07:59.189
-of -the -art reasoning model. Oh, for sure.
00:07:59.310 --> 00:08:02.810
Prices are all over the place. So does this breadth
00:08:02.810 --> 00:08:05.370
-first strategy hold up when the cost of input
00:08:05.370 --> 00:08:08.110
versus output tokens completely changes the weight
00:08:08.110 --> 00:08:11.250
of the budget? It does. The researchers actually
00:08:11.250 --> 00:08:13.310
built that directly into their predictive function.
00:08:13.670 --> 00:08:15.529
Oh, they accounted for the pricing tiers. They
00:08:15.529 --> 00:08:18.009
did. They found that the single -read variance
00:08:18.009 --> 00:08:21.269
penalties, so the... tiny bit of accuracy you
00:08:21.269 --> 00:08:23.189
lose by only having the judge read the answer
00:08:23.189 --> 00:08:27.769
once, only ranged from 0 to 9 .9 % relative to
00:08:27.769 --> 00:08:30.850
the absolute optimal allocations. Wait, so even
00:08:30.850 --> 00:08:32.429
at different price points, it's still better
00:08:32.429 --> 00:08:35.330
to just read it once? Yeah. Having an AI judge
00:08:35.330 --> 00:08:38.250
reread a massive context window of an agent's
00:08:38.250 --> 00:08:41.610
search trajectory is wildly expensive compared
00:08:41.610 --> 00:08:44.169
to the marginal informational gain you get. So
00:08:44.169 --> 00:08:47.429
the cost analysis universally favors adding more
00:08:47.429 --> 00:08:50.730
questions over adding repeats. regardless of
00:08:50.730 --> 00:08:53.950
your API tier. That's the takeaway. So if you're
00:08:53.950 --> 00:08:56.870
running an AI project, the actionable takeaway
00:08:56.870 --> 00:08:59.750
is clear. Stop wasting your token budget asking
00:08:59.750 --> 00:09:02.590
your AI the same question over and over. Stop
00:09:02.590 --> 00:09:04.490
trying to see if the chef cooks the spaghetti
00:09:04.490 --> 00:09:07.970
identically every time. Exactly. Go wide. Ask
00:09:07.970 --> 00:09:10.529
the agent 50 different questions to find out
00:09:10.529 --> 00:09:13.409
if the restaurant is actually good. It is mathematically
00:09:13.409 --> 00:09:16.210
the most economically efficient way to buy certainty
00:09:16.210 --> 00:09:19.000
in a noisy environment. Here's where it gets
00:09:19.000 --> 00:09:21.500
really interesting for me, though. When you say
00:09:21.500 --> 00:09:24.539
it out loud, going wide with more questions is
00:09:24.539 --> 00:09:26.220
better than asking the same question repeatedly.
00:09:27.179 --> 00:09:29.879
It honestly sounds like common sense, you know?
00:09:29.960 --> 00:09:32.940
It is highly intuitive, yeah. So how genuinely
00:09:32.940 --> 00:09:36.620
new is this paper? Are they just using sophisticated
00:09:36.620 --> 00:09:39.139
math to prove something developers already suspected?
00:09:39.960 --> 00:09:42.159
Rating the novelty here, I'd give it a solid
00:09:42.159 --> 00:09:45.200
four out of five. Really? A four out of five
00:09:45.200 --> 00:09:48.429
for proving we need more questions. Well, because
00:09:48.429 --> 00:09:51.470
the true novelty isn't the headline that breadth
00:09:51.470 --> 00:09:55.269
is good. The novelty is the quantification. Okay,
00:09:55.330 --> 00:09:58.049
I see. This paper is essentially a statistician's
00:09:58.049 --> 00:10:01.649
answer to an LLM evaluation problem. What's fascinating
00:10:01.649 --> 00:10:04.269
here is that they took the exact same logic used
00:10:04.269 --> 00:10:07.850
in clustered AB power analysis and retrofitted
00:10:07.850 --> 00:10:11.509
it for AI agent evaluations. Wait, hold on. How
00:10:11.509 --> 00:10:14.450
does clustered AB power analysis translate to
00:10:14.450 --> 00:10:17.740
AI tokens? So in a standard marketing A -B test,
00:10:17.940 --> 00:10:20.799
you need to calculate statistical power to know
00:10:20.799 --> 00:10:22.700
if your sample size is large enough to prove
00:10:22.700 --> 00:10:25.940
an ad actually worked, right? Rather than just
00:10:25.940 --> 00:10:28.120
getting lucky. Right. You need a big enough sample
00:10:28.120 --> 00:10:31.440
so it's statistically significant. Exactly. Here,
00:10:31.620 --> 00:10:34.820
developers need statistical power to know if
00:10:34.820 --> 00:10:38.259
an updated agentic RAG system is actually better
00:10:38.259 --> 00:10:41.259
than their previous baseline. Tokens are your
00:10:41.259 --> 00:10:43.960
sample size budget. Oh, that makes total sense.
00:10:44.419 --> 00:10:47.220
So instead of a team manager relying on a gut
00:10:47.220 --> 00:10:49.759
feeling that they should ask more questions,
00:10:49.960 --> 00:10:52.500
this predictive mathematical model tells you
00:10:52.500 --> 00:10:55.059
exactly what your margin of error will be. With
00:10:55.059 --> 00:10:58.580
that 3 .5 % to 4 % accuracy. Exactly. Based entirely
00:10:58.580 --> 00:11:01.480
on how you slice your token budget. So it's the
00:11:01.480 --> 00:11:03.399
difference between knowing eating vegetables
00:11:03.399 --> 00:11:05.759
is good for you and having a system calculate
00:11:05.759 --> 00:11:08.779
the exact optimal milligrams of vitamin C you
00:11:08.779 --> 00:11:11.059
need today based on your specific metabolism.
00:11:11.539 --> 00:11:14.019
That's a great way to put it. It moves us from
00:11:14.019 --> 00:11:17.580
a heuristic to a highly actionable formula. It
00:11:17.580 --> 00:11:20.000
separates the true advance from just established
00:11:20.000 --> 00:11:22.659
practice by providing the exact mathematical
00:11:22.659 --> 00:11:26.039
mechanism to execute the evaluation flawlessly.
00:11:26.360 --> 00:11:29.659
Okay, but let's step back a bit. That 33 % drop
00:11:29.659 --> 00:11:32.639
in uncertainty is fantastic for standard QA datasets
00:11:32.639 --> 00:11:36.240
like Hotpot QA. Sure. But those are sterile academic
00:11:36.240 --> 00:11:39.110
benchmarks. If an agent is writing production
00:11:39.110 --> 00:11:40.909
code instead of answering multi -hop trivia,
00:11:41.679 --> 00:11:43.820
A single trajectory error, like pulling the wrong
00:11:43.820 --> 00:11:46.179
API documentation, crashes the entire application.
00:11:46.559 --> 00:11:49.519
That's very true. Does this math hold up in the
00:11:49.519 --> 00:11:52.179
messy real world? I mean, to really understand
00:11:52.179 --> 00:11:54.200
this, we have to look at what else is happening
00:11:54.200 --> 00:11:56.399
in the evaluation space right now, specifically
00:11:56.399 --> 00:11:59.299
projects like GrabBench. Yeah, GrabBench is a
00:11:59.299 --> 00:12:02.080
great counterweight to discuss here because GrabBench
00:12:02.080 --> 00:12:04.580
focuses intensely on the practical reality of
00:12:04.580 --> 00:12:06.399
production environments. Right. It's not just
00:12:06.399 --> 00:12:09.509
trivia. Exactly. It evaluates whether an AI works
00:12:09.509 --> 00:12:12.429
when integrated into actual company software,
00:12:12.629 --> 00:12:15.590
dealing with real -world friction and complex
00:12:15.590 --> 00:12:19.169
tool use. It takes the chef out of the test kitchen
00:12:19.169 --> 00:12:21.190
and puts him on the line during a Friday night
00:12:21.190 --> 00:12:23.669
dinner rush. Great analogy. And there's also
00:12:23.669 --> 00:12:26.850
the Beyond Semantic Similarity paper. That actually
00:12:26.850 --> 00:12:28.889
released on the exact same day as this budget
00:12:28.889 --> 00:12:32.049
allocation one. Really? Same day? Yeah. And that
00:12:32.049 --> 00:12:34.129
one looks at the performance and costs of the
00:12:34.129 --> 00:12:37.110
agentic retrieval mechanism itself. So how do
00:12:37.110 --> 00:12:39.929
all these fit together? Well, those papers tackle
00:12:39.929 --> 00:12:42.769
the same overarching problem of trusting AI agents,
00:12:42.830 --> 00:12:45.429
but from completely different angles. Beyond
00:12:45.429 --> 00:12:48.409
Semantic Similarity evaluates the retrieval algorithm's
00:12:48.409 --> 00:12:52.169
efficiency. Okay. GrabBench evaluates the agent's
00:12:52.169 --> 00:12:55.110
environmental success. But this budget allocation
00:12:55.110 --> 00:12:58.250
paper, it is uniquely laser -focused on the underlying
00:12:58.250 --> 00:13:00.970
measurement framework itself. So it's telling
00:13:00.970 --> 00:13:03.190
you how to set up the test in the first place.
00:13:03.509 --> 00:13:07.269
Exactly. It uses generalizability theory to tell
00:13:07.269 --> 00:13:09.470
you exactly how to divide your pie chart of token
00:13:09.470 --> 00:13:12.049
spend before you even begin running those production
00:13:12.049 --> 00:13:14.450
workloads. It's foundational, whereas the others
00:13:14.450 --> 00:13:16.590
are more situational. It's the blueprint for
00:13:16.590 --> 00:13:20.460
the test. Yes, exactly. But I want to offer some
00:13:20.460 --> 00:13:22.820
expert skepticism here before a listener takes
00:13:22.820 --> 00:13:25.519
this blueprint and treats it as an absolute law
00:13:25.519 --> 00:13:27.879
of physics for their own projects. And that's
00:13:27.879 --> 00:13:30.629
totally fair. Because like we discussed, the
00:13:30.629 --> 00:13:32.970
nature of a coding task is fundamentally different
00:13:32.970 --> 00:13:35.879
than a QA task. Yeah, we definitely have to acknowledge
00:13:35.879 --> 00:13:39.259
a major caveat here. And just as my own opinion,
00:13:39.279 --> 00:13:42.039
looking at the space, these precise results are
00:13:42.039 --> 00:13:45.240
based entirely on a hotpot QA and music. Two
00:13:45.240 --> 00:13:48.299
specific benchmarks. Right. So if we connect
00:13:48.299 --> 00:13:50.700
this to the bigger picture, the optimal budget
00:13:50.700 --> 00:13:52.879
allocation might shift dramatically when applied
00:13:52.879 --> 00:13:55.639
to completely different tasks. Like creative
00:13:55.639 --> 00:13:59.500
writing AIs or coding agents. Exactly. If you're
00:13:59.500 --> 00:14:02.179
testing a coding agent. trajectory variance might
00:14:02.179 --> 00:14:04.360
suddenly become massively more important than
00:14:04.360 --> 00:14:06.940
question variance because the sequential logic
00:14:06.940 --> 00:14:09.379
of the search path dictates the compilation success.
00:14:09.820 --> 00:14:12.240
So while eval teams across the industry will
00:14:12.240 --> 00:14:15.259
likely copy this framework because mathematically
00:14:15.259 --> 00:14:17.299
predicting standard error is just too useful
00:14:17.299 --> 00:14:20.460
to ignore, they need to treat these exact allocation
00:14:20.460 --> 00:14:24.259
ratios as indicative guardrails, not an absolute
00:14:24.259 --> 00:14:27.269
law. Right. The framework provides the tools
00:14:27.269 --> 00:14:29.850
to map your own variants. It helps you find the
00:14:29.850 --> 00:14:32.169
optimal percentages for your unique use case.
00:14:32.389 --> 00:14:34.850
But the numbers won't be exactly the same everywhere.
00:14:35.070 --> 00:14:37.470
That makes total sense. So let's briefly recap
00:14:37.470 --> 00:14:40.570
this journey. Sure. Evaluating AI agents is inherently
00:14:40.570 --> 00:14:43.710
noisy. But by applying the old school statistical
00:14:43.710 --> 00:14:47.049
rigor of generalizability theory, we now have
00:14:47.049 --> 00:14:49.470
a mathematical roadmap. A very precise one, yeah.
00:14:49.690 --> 00:14:52.029
And that roadmap proves that when cutting through
00:14:52.029 --> 00:14:54.909
the noise with a fixed token budget, breadth...
00:14:55.049 --> 00:14:58.179
beats repetition. Funding a wider variety of
00:14:58.179 --> 00:15:00.980
questions is vastly superior to retesting the
00:15:00.980 --> 00:15:03.600
same ones. Absolutely. And you know, this raises
00:15:03.600 --> 00:15:05.740
an important question to mull over as we wrap
00:15:05.740 --> 00:15:07.879
up today. What's that? Well, if our understanding
00:15:07.879 --> 00:15:11.059
of an AI agent's core capability changes so drastically
00:15:11.059 --> 00:15:13.740
just based on how we mathematically allocate
00:15:13.740 --> 00:15:16.139
our testing budget, are we actually building
00:15:16.139 --> 00:15:18.580
AI that is good for the real world, or are we
00:15:18.580 --> 00:15:21.120
just building AI that is hyper -optimized to
00:15:21.120 --> 00:15:24.059
pass our highly specific, heavily budgeted tests?
00:15:24.559 --> 00:15:27.419
Wow. Are we just training chefs who only know
00:15:27.419 --> 00:15:29.519
how to cook for the health inspector? Basically,
00:15:29.679 --> 00:15:32.419
yeah. It's a serious friction the entire industry
00:15:32.419 --> 00:15:35.200
has to grapple with as these systems deploy into
00:15:35.200 --> 00:15:38.259
the real world. Well, thanks for joining us on
00:15:38.259 --> 00:15:40.879
this deep dive. Keep questioning the data, keep
00:15:40.879 --> 00:15:43.279
looking at how the tests are designed, and we'll
00:15:43.279 --> 00:15:43.879
catch you next time.
00:00:00.000 --> 00:00:02.580
Welcome to this deep dive. You know, our mission
00:00:02.580 --> 00:00:04.740
today is really about cutting through the massive
00:00:04.740 --> 00:00:08.679
amount of noise in AI evaluation. Yeah, and there
00:00:08.679 --> 00:00:11.199
is a lot of noise out there right now. It really
00:00:11.199 --> 00:00:13.900
is. We're doing that by unpacking this fascinating
00:00:13.900 --> 00:00:17.300
October 2026 archive paper. It was by Jingji
00:00:17.300 --> 00:00:20.899
Ning, Xueqi Li, and Yibo Kang. Right, the one
00:00:20.899 --> 00:00:24.160
titled, Agentic RAG Evaluation, Budget Allocation
00:00:24.160 --> 00:00:27.100
Across Questions, Trajectories, and Reads. Exactly.
00:00:27.280 --> 00:00:29.519
And the core tension they're exploring is something,
00:00:29.640 --> 00:00:32.520
well, pretty much every AI developer is agonizing
00:00:32.520 --> 00:00:35.859
over right now. Companies are burning staggering
00:00:35.859 --> 00:00:38.600
amounts of money testing their AI agents. Oh,
00:00:38.640 --> 00:00:40.820
absolutely. I mean, we're talking millions of
00:00:40.820 --> 00:00:44.039
tokens, which translates directly to hard dollars.
00:00:44.159 --> 00:00:46.640
And there's this very real possibility that they're
00:00:46.640 --> 00:00:49.579
completely wasting those budgets by measuring
00:00:49.579 --> 00:00:52.539
the exact wrong things. Yeah, that's the scary
00:00:52.539 --> 00:00:55.280
part for a lot of teams. So, OK, let's unpack
00:00:55.280 --> 00:00:57.590
this. Think about this business problem like
00:00:57.590 --> 00:01:00.390
trying to raid a brand new restaurant in your
00:01:00.390 --> 00:01:02.509
neighborhood. Okay, I like that analogy. Right,
00:01:02.590 --> 00:01:04.769
because you assume there's a certain level of
00:01:04.769 --> 00:01:07.709
objective truth to be found. You go in, you eat
00:01:07.709 --> 00:01:09.829
the food, and you decide if it's good or bad.
00:01:10.069 --> 00:01:12.409
Right, simple enough. But imagine if that experience
00:01:12.409 --> 00:01:15.329
was just totally chaotic, like you order the
00:01:15.329 --> 00:01:17.269
spaghetti. But the chef cooks it differently
00:01:17.269 --> 00:01:19.090
every single time you order it. That would be
00:01:19.090 --> 00:01:20.969
frustrating, yeah. And to make matters worse,
00:01:21.170 --> 00:01:23.590
your own taste buds randomly change depending
00:01:23.590 --> 00:01:26.049
on what day of the week it is. Suddenly, trying
00:01:26.049 --> 00:01:27.950
to give that restaurant a definitive rating becomes
00:01:27.950 --> 00:01:31.909
this massive, expensive headache. That evaluative
00:01:31.909 --> 00:01:34.489
muddy water is exactly the landscape development
00:01:34.489 --> 00:01:37.489
teams face when testing agentic rag systems today.
00:01:37.689 --> 00:01:40.319
It really is. Because, you know, we aren't dealing
00:01:40.319 --> 00:01:42.680
with static data sets anymore. When a system
00:01:42.680 --> 00:01:44.760
actively goes out and searches for information
00:01:44.760 --> 00:01:47.599
to build its answer, evaluating its capability
00:01:47.599 --> 00:01:50.700
introduces three distinct layers of variance.
00:01:51.060 --> 00:01:53.900
Okay, break those down for us. So first you have
00:01:53.900 --> 00:01:56.739
the question variance, just the differences in
00:01:56.739 --> 00:01:59.750
the prompts themselves. Right. Then, if you ask
00:01:59.750 --> 00:02:02.510
the agent the exact same question twice, it might
00:02:02.510 --> 00:02:04.709
follow a completely different search trajectory,
00:02:04.810 --> 00:02:07.650
like it pulls different sources or maybe pulls
00:02:07.650 --> 00:02:10.030
them in a different order. So that's the chef
00:02:10.030 --> 00:02:11.810
cooking the spaghetti differently every time.
00:02:12.169 --> 00:02:14.830
Exactly. And finally, you have the evaluation
00:02:14.830 --> 00:02:18.229
itself. This is usually performed by an AI judge
00:02:18.229 --> 00:02:20.789
grading the final answer and that search trajectory.
00:02:21.050 --> 00:02:23.310
And that would be the unpredictable taste buds.
00:02:23.509 --> 00:02:26.909
Spot on. So the industry dilemma is that development
00:02:26.909 --> 00:02:29.710
teams have a fixed token budget, but they don't
00:02:29.710 --> 00:02:31.810
know how to spend it. Do they ask the AI new
00:02:31.810 --> 00:02:34.110
questions? Right. Or do they repeat the same
00:02:34.110 --> 00:02:36.389
question to map out different search trajectories?
00:02:36.530 --> 00:02:38.909
Or do they have the AI judge read the final answer
00:02:38.909 --> 00:02:41.189
multiple times to ensure the grade is consistent?
00:02:41.330 --> 00:02:43.469
They just don't know. It's a massive resource
00:02:43.469 --> 00:02:46.069
allocation problem. Like if you have a budget
00:02:46.069 --> 00:02:48.430
of $500 to figure out if that restaurant is good,
00:02:48.569 --> 00:02:51.569
how do you spend it? Yeah, exactly. Do you order
00:02:51.569 --> 00:02:54.509
50 different dishes? Or order the exact same
00:02:54.509 --> 00:02:57.810
spaghetti 10 times? Or do you buy one dish and
00:02:57.810 --> 00:03:02.370
have five friends taste it? And without any mathematical
00:03:02.370 --> 00:03:05.129
guidance, most teams just default to repeated
00:03:05.129 --> 00:03:07.590
runs. They just brute force it. Right. They test
00:03:07.590 --> 00:03:09.689
the same questions and have the judge do multiple
00:03:09.689 --> 00:03:12.330
reads, basically hoping an average will give
00:03:12.330 --> 00:03:15.069
them a reliable baseline. But that doesn't actually
00:03:15.069 --> 00:03:17.469
work, does it? No. Doing that simply ends up
00:03:17.469 --> 00:03:20.509
remeasuring noise. They're burning token budgets
00:03:20.509 --> 00:03:22.469
without getting a clearer picture of how capable
00:03:22.469 --> 00:03:25.030
the agent truly is. So how do we actually fix
00:03:25.030 --> 00:03:27.110
this? How do we find the mathematical truth here?
00:03:27.509 --> 00:03:29.229
Well, to solve this, the researchers applied
00:03:29.229 --> 00:03:32.530
a robust statistical framework called generalizability
00:03:32.530 --> 00:03:35.330
theory. Wait, isn't that from like the 1970s?
00:03:35.370 --> 00:03:38.009
It is, yeah. It was originally developed for
00:03:38.009 --> 00:03:40.449
psychological and educational testing. You'll
00:03:40.449 --> 00:03:43.250
often hear it associated with Kronbach. Oh, interesting.
00:03:43.490 --> 00:03:46.370
So how does a 1970s psychology framework help
00:03:46.370 --> 00:03:49.759
with AI tokens? So... Generalizability theory
00:03:49.759 --> 00:03:52.120
is designed to split measurement variance into
00:03:52.120 --> 00:03:55.039
specific facets. Instead of treating an inconsistent
00:03:55.039 --> 00:03:58.979
AI score as just one big lump of unexplainable
00:03:58.979 --> 00:04:01.520
error. Which is what everyone does now. Right.
00:04:01.860 --> 00:04:04.639
Instead of that, this framework isolates the
00:04:04.639 --> 00:04:07.379
noise. Okay, so it's essentially taking a scalpel
00:04:07.379 --> 00:04:11.099
to the variance. You can mathematically say this
00:04:11.099 --> 00:04:13.099
percentage of the fluctuation is caused by our
00:04:13.099 --> 00:04:16.180
question set being too diverse. Exactly. And...
00:04:16.709 --> 00:04:19.230
This percentage is from the agent taking weird
00:04:19.230 --> 00:04:22.689
search paths. And this percentage is just our
00:04:22.689 --> 00:04:25.970
AI judge being moody. Exactly. And that breakdown
00:04:25.970 --> 00:04:27.910
allows you to model and predict the uncertainty
00:04:27.910 --> 00:04:30.910
of your evaluation, which in this paper is quantified
00:04:30.910 --> 00:04:33.329
as the standard error. And just to clarify, a
00:04:33.329 --> 00:04:35.290
smaller standard error means a tighter, more
00:04:35.290 --> 00:04:38.459
reliable evaluation, right? Correct. So to figure
00:04:38.459 --> 00:04:40.699
out how to minimize that standard error, the
00:04:40.699 --> 00:04:43.519
researchers used repeated sampling on two standard
00:04:43.519 --> 00:04:46.980
multi -hop QA benchmarks, Hot Pot QA and Music
00:04:46.980 --> 00:04:49.139
Q. Okay, those are pretty standard in the industry?
00:04:49.519 --> 00:04:52.779
Yeah, very standard. So they systematically varied
00:04:52.779 --> 00:04:55.079
the number of questions, the trajectories, and
00:04:55.079 --> 00:04:57.519
the reads, and they recorded the variance at
00:04:57.519 --> 00:04:59.540
every step. So they just mapped everything out?
00:04:59.759 --> 00:05:02.459
They did. And from all that data, they built
00:05:02.459 --> 00:05:05.089
predictive models for budget allocation. Get
00:05:05.089 --> 00:05:08.230
this, it achieved an incredibly impressive 3
00:05:08.230 --> 00:05:11.709
.5 % to 4 .0 % accuracy in predicting the standard
00:05:11.709 --> 00:05:14.069
error based on how tokens were spent. Wow, that's
00:05:14.069 --> 00:05:17.410
really tight. But I have to push back on one
00:05:17.410 --> 00:05:19.949
thing here. Sure, go for it. I'm struggling to
00:05:19.949 --> 00:05:22.110
buy that the AI judge is actually a significant
00:05:22.110 --> 00:05:24.829
source of that noise. I mean, we work with these
00:05:24.829 --> 00:05:27.430
frontier models daily. They don't just wildly
00:05:27.430 --> 00:05:30.529
hallucinate their grading criteria on the exact
00:05:30.529 --> 00:05:33.290
same text upon a second read, do they? You would
00:05:33.290 --> 00:05:35.629
think so, but the researchers actually documented
00:05:35.629 --> 00:05:39.410
a massive 14 .3 % disagreement rate. Wait, really?
00:05:39.610 --> 00:05:44.160
14 %? Yeah. When the AI judge evaluated the exact
00:05:44.160 --> 00:05:47.100
same answer twice using standard default parameters,
00:05:47.379 --> 00:05:50.959
it contradicted its own previous grade over 14
00:05:50.959 --> 00:05:54.319
% of the time. That is insane. So doing multiple
00:05:54.319 --> 00:05:57.680
reads really would eat up a lot of the budget
00:05:57.680 --> 00:06:01.040
trying to fix that. But here is the kicker. No.
00:06:01.569 --> 00:06:04.430
They uncovered a highly practical fix for it.
00:06:04.550 --> 00:06:07.550
Okay, I'm listening. By simply dropping the judge's
00:06:07.550 --> 00:06:10.009
temperature setting to zero, basically forcing
00:06:10.009 --> 00:06:13.129
determinism, that disagreement rate plummeted
00:06:13.129 --> 00:06:17.089
down to just 3 .4%. Oh, wow. Yeah. Okay, so just
00:06:17.089 --> 00:06:19.110
by neutralizing the model's creativity dial,
00:06:19.350 --> 00:06:21.870
they almost entirely eliminated the read variance.
00:06:22.110 --> 00:06:25.329
Exactly. The judge stops changing its mind. Which
00:06:25.329 --> 00:06:27.370
means you solve that specific facet of noise
00:06:27.370 --> 00:06:30.149
with a basic settings tweak rather than throwing
00:06:30.149 --> 00:06:32.790
more of your actual token budget at it. And that
00:06:32.790 --> 00:06:35.009
fundamentally alters the math for your budget
00:06:35.009 --> 00:06:37.410
allocation. You immediately know that spending
00:06:37.410 --> 00:06:40.170
tokens, having a judge reread the same output,
00:06:40.250 --> 00:06:42.910
is economically inefficient. Right, because a
00:06:42.910 --> 00:06:45.089
simple parameter change gives you the stability
00:06:45.089 --> 00:06:48.470
you need. Exactly. Okay, so armed with that knowledge
00:06:48.470 --> 00:06:51.589
of how the variance splits, what does this actually
00:06:51.589 --> 00:06:54.649
change for a team building AI today? Let's talk
00:06:54.649 --> 00:06:56.670
about the hard numbers. Let's do it. If I'm managing
00:06:56.670 --> 00:06:58.750
a development team today and I have a budget
00:06:58.750 --> 00:07:01.769
of, say, 34 million model tokens sitting in my
00:07:01.769 --> 00:07:04.610
cloud account, how does this paper tell me to
00:07:04.610 --> 00:07:07.550
actually deploy them? Well, expanding the number
00:07:07.550 --> 00:07:09.949
of questions is the ultimate winner. The paper
00:07:09.949 --> 00:07:12.529
provides concrete results for that exact 34 million
00:07:12.529 --> 00:07:14.670
token scenario. Okay, so what's the breakdown?
00:07:15.370 --> 00:07:18.029
Going wide with your question coverage lowers
00:07:18.029 --> 00:07:21.069
the standard error by a massive 33 % compared
00:07:21.069 --> 00:07:23.879
to a strategy of doing five repeated reads. A
00:07:23.879 --> 00:07:26.860
33 % drop in uncertainty just by expanding the
00:07:26.860 --> 00:07:29.920
prompt pool. That's definitive. It really is.
00:07:30.279 --> 00:07:32.720
And expanding the question set also lowers the
00:07:32.720 --> 00:07:35.660
standard error by 12 .6 % compared to asking
00:07:35.660 --> 00:07:37.480
the same question and mapping three different
00:07:37.480 --> 00:07:40.240
search trajectories. So the mathematical truth
00:07:40.240 --> 00:07:43.620
just points in one direction. Breath beats depth.
00:07:43.879 --> 00:07:46.040
Yep. When you're trying to find the true capability
00:07:46.040 --> 00:07:48.639
of the system, you go wide. But what about the
00:07:48.639 --> 00:07:51.509
cost analysis across different models? Because
00:07:51.509 --> 00:07:55.110
we know API price tiers vary wildly, right? Between
00:07:55.110 --> 00:07:57.610
a cheaper, faster model and like a heavy state
00:07:57.610 --> 00:07:59.189
-of -the -art reasoning model. Oh, for sure.
00:07:59.310 --> 00:08:02.810
Prices are all over the place. So does this breadth
00:08:02.810 --> 00:08:05.370
-first strategy hold up when the cost of input
00:08:05.370 --> 00:08:08.110
versus output tokens completely changes the weight
00:08:08.110 --> 00:08:11.250
of the budget? It does. The researchers actually
00:08:11.250 --> 00:08:13.310
built that directly into their predictive function.
00:08:13.670 --> 00:08:15.529
Oh, they accounted for the pricing tiers. They
00:08:15.529 --> 00:08:18.009
did. They found that the single -read variance
00:08:18.009 --> 00:08:21.269
penalties, so the... tiny bit of accuracy you
00:08:21.269 --> 00:08:23.189
lose by only having the judge read the answer
00:08:23.189 --> 00:08:27.769
once, only ranged from 0 to 9 .9 % relative to
00:08:27.769 --> 00:08:30.850
the absolute optimal allocations. Wait, so even
00:08:30.850 --> 00:08:32.429
at different price points, it's still better
00:08:32.429 --> 00:08:35.330
to just read it once? Yeah. Having an AI judge
00:08:35.330 --> 00:08:38.250
reread a massive context window of an agent's
00:08:38.250 --> 00:08:41.610
search trajectory is wildly expensive compared
00:08:41.610 --> 00:08:44.169
to the marginal informational gain you get. So
00:08:44.169 --> 00:08:47.429
the cost analysis universally favors adding more
00:08:47.429 --> 00:08:50.730
questions over adding repeats. regardless of
00:08:50.730 --> 00:08:53.950
your API tier. That's the takeaway. So if you're
00:08:53.950 --> 00:08:56.870
running an AI project, the actionable takeaway
00:08:56.870 --> 00:08:59.750
is clear. Stop wasting your token budget asking
00:08:59.750 --> 00:09:02.590
your AI the same question over and over. Stop
00:09:02.590 --> 00:09:04.490
trying to see if the chef cooks the spaghetti
00:09:04.490 --> 00:09:07.970
identically every time. Exactly. Go wide. Ask
00:09:07.970 --> 00:09:10.529
the agent 50 different questions to find out
00:09:10.529 --> 00:09:13.409
if the restaurant is actually good. It is mathematically
00:09:13.409 --> 00:09:16.210
the most economically efficient way to buy certainty
00:09:16.210 --> 00:09:19.000
in a noisy environment. Here's where it gets
00:09:19.000 --> 00:09:21.500
really interesting for me, though. When you say
00:09:21.500 --> 00:09:24.539
it out loud, going wide with more questions is
00:09:24.539 --> 00:09:26.220
better than asking the same question repeatedly.
00:09:27.179 --> 00:09:29.879
It honestly sounds like common sense, you know?
00:09:29.960 --> 00:09:32.940
It is highly intuitive, yeah. So how genuinely
00:09:32.940 --> 00:09:36.620
new is this paper? Are they just using sophisticated
00:09:36.620 --> 00:09:39.139
math to prove something developers already suspected?
00:09:39.960 --> 00:09:42.159
Rating the novelty here, I'd give it a solid
00:09:42.159 --> 00:09:45.200
four out of five. Really? A four out of five
00:09:45.200 --> 00:09:48.429
for proving we need more questions. Well, because
00:09:48.429 --> 00:09:51.470
the true novelty isn't the headline that breadth
00:09:51.470 --> 00:09:55.269
is good. The novelty is the quantification. Okay,
00:09:55.330 --> 00:09:58.049
I see. This paper is essentially a statistician's
00:09:58.049 --> 00:10:01.649
answer to an LLM evaluation problem. What's fascinating
00:10:01.649 --> 00:10:04.269
here is that they took the exact same logic used
00:10:04.269 --> 00:10:07.850
in clustered AB power analysis and retrofitted
00:10:07.850 --> 00:10:11.509
it for AI agent evaluations. Wait, hold on. How
00:10:11.509 --> 00:10:14.450
does clustered AB power analysis translate to
00:10:14.450 --> 00:10:17.740
AI tokens? So in a standard marketing A -B test,
00:10:17.940 --> 00:10:20.799
you need to calculate statistical power to know
00:10:20.799 --> 00:10:22.700
if your sample size is large enough to prove
00:10:22.700 --> 00:10:25.940
an ad actually worked, right? Rather than just
00:10:25.940 --> 00:10:28.120
getting lucky. Right. You need a big enough sample
00:10:28.120 --> 00:10:31.440
so it's statistically significant. Exactly. Here,
00:10:31.620 --> 00:10:34.820
developers need statistical power to know if
00:10:34.820 --> 00:10:38.259
an updated agentic RAG system is actually better
00:10:38.259 --> 00:10:41.259
than their previous baseline. Tokens are your
00:10:41.259 --> 00:10:43.960
sample size budget. Oh, that makes total sense.
00:10:44.419 --> 00:10:47.220
So instead of a team manager relying on a gut
00:10:47.220 --> 00:10:49.759
feeling that they should ask more questions,
00:10:49.960 --> 00:10:52.500
this predictive mathematical model tells you
00:10:52.500 --> 00:10:55.059
exactly what your margin of error will be. With
00:10:55.059 --> 00:10:58.580
that 3 .5 % to 4 % accuracy. Exactly. Based entirely
00:10:58.580 --> 00:11:01.480
on how you slice your token budget. So it's the
00:11:01.480 --> 00:11:03.399
difference between knowing eating vegetables
00:11:03.399 --> 00:11:05.759
is good for you and having a system calculate
00:11:05.759 --> 00:11:08.779
the exact optimal milligrams of vitamin C you
00:11:08.779 --> 00:11:11.059
need today based on your specific metabolism.
00:11:11.539 --> 00:11:14.019
That's a great way to put it. It moves us from
00:11:14.019 --> 00:11:17.580
a heuristic to a highly actionable formula. It
00:11:17.580 --> 00:11:20.000
separates the true advance from just established
00:11:20.000 --> 00:11:22.659
practice by providing the exact mathematical
00:11:22.659 --> 00:11:26.039
mechanism to execute the evaluation flawlessly.
00:11:26.360 --> 00:11:29.659
Okay, but let's step back a bit. That 33 % drop
00:11:29.659 --> 00:11:32.639
in uncertainty is fantastic for standard QA datasets
00:11:32.639 --> 00:11:36.240
like Hotpot QA. Sure. But those are sterile academic
00:11:36.240 --> 00:11:39.110
benchmarks. If an agent is writing production
00:11:39.110 --> 00:11:40.909
code instead of answering multi -hop trivia,
00:11:41.679 --> 00:11:43.820
A single trajectory error, like pulling the wrong
00:11:43.820 --> 00:11:46.179
API documentation, crashes the entire application.
00:11:46.559 --> 00:11:49.519
That's very true. Does this math hold up in the
00:11:49.519 --> 00:11:52.179
messy real world? I mean, to really understand
00:11:52.179 --> 00:11:54.200
this, we have to look at what else is happening
00:11:54.200 --> 00:11:56.399
in the evaluation space right now, specifically
00:11:56.399 --> 00:11:59.299
projects like GrabBench. Yeah, GrabBench is a
00:11:59.299 --> 00:12:02.080
great counterweight to discuss here because GrabBench
00:12:02.080 --> 00:12:04.580
focuses intensely on the practical reality of
00:12:04.580 --> 00:12:06.399
production environments. Right. It's not just
00:12:06.399 --> 00:12:09.509
trivia. Exactly. It evaluates whether an AI works
00:12:09.509 --> 00:12:12.429
when integrated into actual company software,
00:12:12.629 --> 00:12:15.590
dealing with real -world friction and complex
00:12:15.590 --> 00:12:19.169
tool use. It takes the chef out of the test kitchen
00:12:19.169 --> 00:12:21.190
and puts him on the line during a Friday night
00:12:21.190 --> 00:12:23.669
dinner rush. Great analogy. And there's also
00:12:23.669 --> 00:12:26.850
the Beyond Semantic Similarity paper. That actually
00:12:26.850 --> 00:12:28.889
released on the exact same day as this budget
00:12:28.889 --> 00:12:32.049
allocation one. Really? Same day? Yeah. And that
00:12:32.049 --> 00:12:34.129
one looks at the performance and costs of the
00:12:34.129 --> 00:12:37.110
agentic retrieval mechanism itself. So how do
00:12:37.110 --> 00:12:39.929
all these fit together? Well, those papers tackle
00:12:39.929 --> 00:12:42.769
the same overarching problem of trusting AI agents,
00:12:42.830 --> 00:12:45.429
but from completely different angles. Beyond
00:12:45.429 --> 00:12:48.409
Semantic Similarity evaluates the retrieval algorithm's
00:12:48.409 --> 00:12:52.169
efficiency. Okay. GrabBench evaluates the agent's
00:12:52.169 --> 00:12:55.110
environmental success. But this budget allocation
00:12:55.110 --> 00:12:58.250
paper, it is uniquely laser -focused on the underlying
00:12:58.250 --> 00:13:00.970
measurement framework itself. So it's telling
00:13:00.970 --> 00:13:03.190
you how to set up the test in the first place.
00:13:03.509 --> 00:13:07.269
Exactly. It uses generalizability theory to tell
00:13:07.269 --> 00:13:09.470
you exactly how to divide your pie chart of token
00:13:09.470 --> 00:13:12.049
spend before you even begin running those production
00:13:12.049 --> 00:13:14.450
workloads. It's foundational, whereas the others
00:13:14.450 --> 00:13:16.590
are more situational. It's the blueprint for
00:13:16.590 --> 00:13:20.460
the test. Yes, exactly. But I want to offer some
00:13:20.460 --> 00:13:22.820
expert skepticism here before a listener takes
00:13:22.820 --> 00:13:25.519
this blueprint and treats it as an absolute law
00:13:25.519 --> 00:13:27.879
of physics for their own projects. And that's
00:13:27.879 --> 00:13:30.629
totally fair. Because like we discussed, the
00:13:30.629 --> 00:13:32.970
nature of a coding task is fundamentally different
00:13:32.970 --> 00:13:35.879
than a QA task. Yeah, we definitely have to acknowledge
00:13:35.879 --> 00:13:39.259
a major caveat here. And just as my own opinion,
00:13:39.279 --> 00:13:42.039
looking at the space, these precise results are
00:13:42.039 --> 00:13:45.240
based entirely on a hotpot QA and music. Two
00:13:45.240 --> 00:13:48.299
specific benchmarks. Right. So if we connect
00:13:48.299 --> 00:13:50.700
this to the bigger picture, the optimal budget
00:13:50.700 --> 00:13:52.879
allocation might shift dramatically when applied
00:13:52.879 --> 00:13:55.639
to completely different tasks. Like creative
00:13:55.639 --> 00:13:59.500
writing AIs or coding agents. Exactly. If you're
00:13:59.500 --> 00:14:02.179
testing a coding agent. trajectory variance might
00:14:02.179 --> 00:14:04.360
suddenly become massively more important than
00:14:04.360 --> 00:14:06.940
question variance because the sequential logic
00:14:06.940 --> 00:14:09.379
of the search path dictates the compilation success.
00:14:09.820 --> 00:14:12.240
So while eval teams across the industry will
00:14:12.240 --> 00:14:15.259
likely copy this framework because mathematically
00:14:15.259 --> 00:14:17.299
predicting standard error is just too useful
00:14:17.299 --> 00:14:20.460
to ignore, they need to treat these exact allocation
00:14:20.460 --> 00:14:24.259
ratios as indicative guardrails, not an absolute
00:14:24.259 --> 00:14:27.269
law. Right. The framework provides the tools
00:14:27.269 --> 00:14:29.850
to map your own variants. It helps you find the
00:14:29.850 --> 00:14:32.169
optimal percentages for your unique use case.
00:14:32.389 --> 00:14:34.850
But the numbers won't be exactly the same everywhere.
00:14:35.070 --> 00:14:37.470
That makes total sense. So let's briefly recap
00:14:37.470 --> 00:14:40.570
this journey. Sure. Evaluating AI agents is inherently
00:14:40.570 --> 00:14:43.710
noisy. But by applying the old school statistical
00:14:43.710 --> 00:14:47.049
rigor of generalizability theory, we now have
00:14:47.049 --> 00:14:49.470
a mathematical roadmap. A very precise one, yeah.
00:14:49.690 --> 00:14:52.029
And that roadmap proves that when cutting through
00:14:52.029 --> 00:14:54.909
the noise with a fixed token budget, breadth...
00:14:55.049 --> 00:14:58.179
beats repetition. Funding a wider variety of
00:14:58.179 --> 00:15:00.980
questions is vastly superior to retesting the
00:15:00.980 --> 00:15:03.600
same ones. Absolutely. And you know, this raises
00:15:03.600 --> 00:15:05.740
an important question to mull over as we wrap
00:15:05.740 --> 00:15:07.879
up today. What's that? Well, if our understanding
00:15:07.879 --> 00:15:11.059
of an AI agent's core capability changes so drastically
00:15:11.059 --> 00:15:13.740
just based on how we mathematically allocate
00:15:13.740 --> 00:15:16.139
our testing budget, are we actually building
00:15:16.139 --> 00:15:18.580
AI that is good for the real world, or are we
00:15:18.580 --> 00:15:21.120
just building AI that is hyper -optimized to
00:15:21.120 --> 00:15:24.059
pass our highly specific, heavily budgeted tests?
00:15:24.559 --> 00:15:27.419
Wow. Are we just training chefs who only know
00:15:27.419 --> 00:15:29.519
how to cook for the health inspector? Basically,
00:15:29.679 --> 00:15:32.419
yeah. It's a serious friction the entire industry
00:15:32.419 --> 00:15:35.200
has to grapple with as these systems deploy into
00:15:35.200 --> 00:15:38.259
the real world. Well, thanks for joining us on
00:15:38.259 --> 00:15:40.879
this deep dive. Keep questioning the data, keep
00:15:40.879 --> 00:15:43.279
looking at how the tests are designed, and we'll
00:15:43.279 --> 00:15:43.879
catch you next time.