WEBVTT
00:00:00.000 --> 00:00:03.399
Imagine this scenario for a second. You are overseeing
00:00:03.399 --> 00:00:06.139
a massive A -B test for a new product feature.
00:00:06.299 --> 00:00:09.820
Oh, yeah. The classic tech setup. Right. So the
00:00:09.820 --> 00:00:11.820
engineering team runs the numbers. You check
00:00:11.820 --> 00:00:14.859
the dashboard. And the data is just flawless.
00:00:15.080 --> 00:00:18.000
I mean, the p -value is practically zero. So
00:00:18.000 --> 00:00:20.980
you pop the champagne. Exactly. You deploy the
00:00:20.980 --> 00:00:25.100
new feature. You celebrate. And then, well, next
00:00:25.100 --> 00:00:27.000
quarter, the company loses millions of dollars.
00:00:27.059 --> 00:00:29.730
And literally nobody can explain why. It's the
00:00:29.730 --> 00:00:32.250
ultimate nightmare for any data -driven team,
00:00:32.329 --> 00:00:34.829
really. You followed all the rules, the math
00:00:34.829 --> 00:00:37.530
said you had a clear winner, but the real -world
00:00:37.530 --> 00:00:40.649
deployment was just a complete disaster. And
00:00:40.649 --> 00:00:43.350
that silent, expensive disaster is exactly what
00:00:43.350 --> 00:00:45.770
we are tearing apart today. Welcome to the Deep
00:00:45.770 --> 00:00:48.600
Dive. Glad to be here. We are looking at a brand
00:00:48.600 --> 00:00:51.640
new paper from September 2026 by Max Farrell,
00:00:51.740 --> 00:00:55.420
Malika Korgan -Bekova, and Sanjog Misra. It's
00:00:55.420 --> 00:00:58.299
called Robust AB Decisions. Yeah, and it is a
00:00:58.299 --> 00:01:01.479
fascinating read. It really is. We're going to
00:01:01.479 --> 00:01:03.560
unpack why the most common battle cry in tech
00:01:03.560 --> 00:01:06.079
and business, you know, P is less than 0 .05,
00:01:06.239 --> 00:01:08.719
ship it, might actually be answering the completely
00:01:08.719 --> 00:01:10.980
wrong question. And that's a big deal because,
00:01:11.040 --> 00:01:14.780
well, entire companies run on that rule. Exactly.
00:01:15.450 --> 00:01:18.329
So if you have ever relied on an A -B test to
00:01:18.329 --> 00:01:20.950
make a decision, whether you are prepping for
00:01:20.950 --> 00:01:23.909
a big board meeting, launching a new user interface,
00:01:24.170 --> 00:01:26.829
or you're just insanely curious about how these
00:01:26.829 --> 00:01:29.849
massive digital platforms operate, this research
00:01:29.849 --> 00:01:32.849
reveals a massive blind spot in that entire process.
00:01:33.209 --> 00:01:35.489
What makes this paper such a breakthrough for
00:01:35.489 --> 00:01:37.689
the industry isn't just that it points out a
00:01:37.689 --> 00:01:40.310
flaw in standard hypothesis testing, you know.
00:01:41.019 --> 00:01:42.959
I mean, people have been complaining about statistical
00:01:42.959 --> 00:01:45.219
significance for decades. Right. It's not a new
00:01:45.219 --> 00:01:48.540
complaint. Exactly. What we are going to explore
00:01:48.540 --> 00:01:52.180
today is a genuinely new practical solution rooted
00:01:52.180 --> 00:01:55.400
in something called robust decision theory. Robust
00:01:55.400 --> 00:01:58.120
decision theory. Yeah. Okay. We are shifting
00:01:58.120 --> 00:02:00.359
the goalpost for merely detecting a mathematical
00:02:00.359 --> 00:02:03.379
difference in a vacuum to actually optimizing
00:02:03.379 --> 00:02:06.340
for economic survival in a highly unpredictable,
00:02:06.739 --> 00:02:09.629
messy world. Okay, let's unpack this standard
00:02:09.629 --> 00:02:11.889
industry workflow first, because you see it everywhere,
00:02:12.030 --> 00:02:14.210
from tiny startups to Fortune 50 companies. You
00:02:14.210 --> 00:02:16.930
really do. Step one, you run your A -B test.
00:02:17.150 --> 00:02:18.930
You have your control group, you have your treatment
00:02:18.930 --> 00:02:21.229
group. Step two, you check the math. Usually
00:02:21.229 --> 00:02:23.569
some form of a T -test, right? Hunting for that
00:02:23.569 --> 00:02:25.830
magical threshold of statistical significance.
00:02:26.169 --> 00:02:29.370
The holy grail. Right. And step three, if the
00:02:29.370 --> 00:02:32.069
new feature wins, you deploy it to all your users,
00:02:32.229 --> 00:02:35.849
permanently. But reading through Farrell, Korganbekova,
00:02:35.990 --> 00:02:38.569
and Mishra's work, I realized how flawed this
00:02:38.569 --> 00:02:41.860
is. Oh, it is fundamentally broken. Using a standard
00:02:41.860 --> 00:02:43.879
t -test to make a final business deployment decision
00:02:43.879 --> 00:02:47.340
feels like using a metal detector to buy a plot
00:02:47.340 --> 00:02:49.780
of land on an active fault line. Wow. Right?
00:02:49.860 --> 00:02:51.539
That's a great way to put it. Sure. It tells
00:02:51.539 --> 00:02:53.960
you there's some valuable ore right beneath your
00:02:53.960 --> 00:02:56.280
feet today, but it completely ignores the fact
00:02:56.280 --> 00:02:58.039
that the ground is mathematically guaranteed
00:02:58.039 --> 00:03:00.199
to shift tomorrow. That is a much more accurate
00:03:00.199 --> 00:03:02.979
way to visualize the danger. To understand why
00:03:02.979 --> 00:03:04.960
we are building on fault lines, you really have
00:03:04.960 --> 00:03:07.180
to look at what the t -test was actually designed
00:03:07.180 --> 00:03:09.060
to do. It wasn't built for business, was it?
00:03:09.219 --> 00:03:12.849
No, not at all. A t -test's null hypothesis framing
00:03:12.849 --> 00:03:15.370
was never built to serve as a business deployment
00:03:15.370 --> 00:03:19.569
policy. It is a scientific tool. Its original
00:03:19.569 --> 00:03:21.949
purpose was strictly to control the rate of false
00:03:21.949 --> 00:03:24.530
claims in academic literature or like clinical
00:03:24.530 --> 00:03:28.449
trials. It answers a very narrow historical question.
00:03:28.650 --> 00:03:31.750
Which is what exactly? It basically asks, did
00:03:31.750 --> 00:03:34.250
the treatment differ from the control in this
00:03:34.250 --> 00:03:37.189
exact specific sample of data that I just happened
00:03:37.189 --> 00:03:39.889
to observe over the last two weeks? In this specific
00:03:39.889 --> 00:03:43.169
sample, which is looking backward. Exactly. It's
00:03:43.169 --> 00:03:45.770
not asking, will this specific feature maximize
00:03:45.770 --> 00:03:48.009
our revenue over the next five years? Precisely.
00:03:48.009 --> 00:03:51.229
It is fundamentally not evaluating the real economic
00:03:51.229 --> 00:03:54.789
payoff. A standard t -test doesn't weigh the
00:03:54.789 --> 00:03:57.370
magnitude of your potential win against the future
00:03:57.370 --> 00:03:59.610
risks of deploying it. Yeah, that makes sense.
00:04:00.000 --> 00:04:02.020
Yet a massive share of real -world corporate
00:04:02.020 --> 00:04:05.060
decisions are gated exclusively by this metric.
00:04:05.340 --> 00:04:09.199
A product team sees a p -value below 0 .05, and
00:04:09.199 --> 00:04:11.479
they automatically ship the feature. Assuming
00:04:11.479 --> 00:04:13.439
the future will look exactly like their two -week
00:04:13.439 --> 00:04:16.860
test window. Which it... Never does. Never. Which
00:04:16.860 --> 00:04:19.180
brings us to the core danger that authors highlight,
00:04:19.379 --> 00:04:22.120
the concept of fragile winners. I want to ground
00:04:22.120 --> 00:04:24.220
this in a real example. Let's do it. Okay. So
00:04:24.220 --> 00:04:26.980
let's say you run a massive A -B test on a new
00:04:26.980 --> 00:04:29.540
checkout button design for an e -commerce platform.
00:04:29.779 --> 00:04:32.480
You run it over a long holiday weekend. Okay.
00:04:32.519 --> 00:04:35.399
Very specific time. Right. The users interacting
00:04:35.399 --> 00:04:37.680
with your site have a very specific mindset.
00:04:37.920 --> 00:04:40.060
They're browsing casually, maybe looking for
00:04:40.060 --> 00:04:42.829
sales. In that specific environment, the new
00:04:42.829 --> 00:04:45.589
button crushes the control. It is a clear winner.
00:04:45.870 --> 00:04:48.310
But what happens when you deploy that button
00:04:48.310 --> 00:04:51.189
globally and suddenly it's a Tuesday in November?
00:04:51.490 --> 00:04:54.490
Right. The users are now, I don't know, stressed
00:04:54.490 --> 00:04:57.290
corporate buyers making bulk purchases on tight
00:04:57.290 --> 00:05:00.310
deadlines. The demographics have shifted slightly.
00:05:00.589 --> 00:05:03.649
The user intent has changed. The broader economic
00:05:03.649 --> 00:05:06.569
environment might even be different. So the environment
00:05:06.569 --> 00:05:09.819
drifted. Yes. That is what the research calls
00:05:09.819 --> 00:05:13.180
distributional drift. The real world is not a
00:05:13.180 --> 00:05:16.819
sealed laboratory vacuum. It is volatile. So
00:05:16.819 --> 00:05:20.180
it's messy. Very. A fragile winter is a feature
00:05:20.180 --> 00:05:22.519
that looked absolutely fantastic in the highly
00:05:22.519 --> 00:05:25.019
specific rigid conditions of your test sample.
00:05:25.439 --> 00:05:28.000
But it falls apart the moment the environment
00:05:28.000 --> 00:05:30.759
drifts even slightly from those exact conditions.
00:05:31.040 --> 00:05:33.220
Because the t -test never checked for drift.
00:05:33.339 --> 00:05:35.300
It just confirmed that the metal detector beeped
00:05:35.300 --> 00:05:37.399
on that one specific holiday weekend. Exactly.
00:05:38.430 --> 00:05:41.509
totally blind to the future. So if fragile winners
00:05:41.509 --> 00:05:43.949
are silently bleeding corporate margins, which
00:05:43.949 --> 00:05:46.670
they are because teams deploy these features
00:05:46.670 --> 00:05:48.850
and then act shocked when the quarterly revenue
00:05:48.850 --> 00:05:52.250
mysteriously dips, how do we fix it? Right. That's
00:05:52.250 --> 00:05:54.170
the million dollar question. Literally, how do
00:05:54.170 --> 00:05:57.629
you mathematically test for a chaotic, unpredictable
00:05:57.629 --> 00:06:00.389
future? This is where the paper introduces its
00:06:00.389 --> 00:06:03.069
core paradigm shift. We have to stop treating
00:06:03.069 --> 00:06:05.910
A -B testing as a traditional hypothesis test.
00:06:06.379 --> 00:06:09.259
Okay. And start treating it as what? As a robust
00:06:09.259 --> 00:06:12.399
decision theory problem. Instead of just looking
00:06:12.399 --> 00:06:14.759
at this single point estimate, the exact numerical
00:06:14.759 --> 00:06:17.839
average you observed in your test, this new framework
00:06:17.839 --> 00:06:21.160
forces you to calculate an ambiguity penalized
00:06:21.160 --> 00:06:24.639
value. Whoa. Okay. That is a heavy bit of jargon.
00:06:24.740 --> 00:06:27.920
I know. Ambiguity penalized value, or as the
00:06:27.920 --> 00:06:29.379
authors refer to it throughout the research,
00:06:29.600 --> 00:06:32.540
ambiguity aversion. We really need to break that
00:06:32.540 --> 00:06:35.019
down for the listener. Yeah. Let's do that. What's
00:06:35.019 --> 00:06:37.540
fascinating here is that ambiguity aversion is
00:06:37.540 --> 00:06:40.839
actually a very intuitive stance on risk. Imagine
00:06:40.839 --> 00:06:42.740
you were packing for a trip to a city you've
00:06:42.740 --> 00:06:45.339
never visited before. Okay, I'm with you. A standard
00:06:45.339 --> 00:06:47.060
A -B test approach would look at the weather
00:06:47.060 --> 00:06:49.839
forecast for today, see that it says 70 degrees
00:06:49.839 --> 00:06:52.660
and sunny, take that as absolute truth, and pack
00:06:52.660 --> 00:06:55.439
only t -shirts. Relying entirely on that single
00:06:55.439 --> 00:06:58.180
point estimate. Right, but an ambiguity averse
00:06:58.180 --> 00:07:00.899
traveler does not fully trust that single forecast.
00:07:01.360 --> 00:07:04.269
They ask, what if a cold front moves in? What
00:07:04.269 --> 00:07:06.410
if the forecast is slightly wrong and it rains?
00:07:06.689 --> 00:07:10.129
So they plan for the drift. Exactly. They inherently
00:07:10.129 --> 00:07:13.050
penalize the t -shirt only option because its
00:07:13.050 --> 00:07:16.269
success relies entirely on one rigid, perfectly
00:07:16.269 --> 00:07:18.810
assumed model of the weather. Right. Instead,
00:07:18.970 --> 00:07:21.569
they prefer the option of packing layers. It
00:07:21.569 --> 00:07:23.949
might be slightly less optimal to carry a heavier
00:07:23.949 --> 00:07:26.629
suitcase if it ends up being exactly 70 degrees,
00:07:26.910 --> 00:07:29.769
but that option remains robust and functional
00:07:29.769 --> 00:07:33.029
across a wide range of plausible nearby weather
00:07:33.029 --> 00:07:35.860
scenarios. Okay, let me translate this back to
00:07:35.860 --> 00:07:38.279
the product dashboard. Under this new framework,
00:07:38.420 --> 00:07:40.899
instead of a data scientist asking, did this
00:07:40.899 --> 00:07:43.720
feature win in our exact historical test? Yeah.
00:07:43.759 --> 00:07:46.879
The deployment rule explicitly asks, would this
00:07:46.879 --> 00:07:49.779
feature still win across a neighborhood of slightly
00:07:49.779 --> 00:07:53.339
different plausible realities? You got it. That's
00:07:53.339 --> 00:07:56.000
exactly it. It evaluates the candidate feature
00:07:56.000 --> 00:07:58.879
under the exact data you collected, but also
00:07:58.879 --> 00:08:01.240
under distributions that are mathematically close
00:08:01.240 --> 00:08:04.610
to it. Oh, wow. If the feature's advantage is
00:08:04.610 --> 00:08:07.149
incredibly fragile, like if it only beats the
00:08:07.149 --> 00:08:10.149
control under highly rigid assumptions, the ambiguity
00:08:10.149 --> 00:08:13.329
-averse rule penalizes it. It deliberately selects
00:08:13.329 --> 00:08:16.050
options that remain strong across those plausible
00:08:16.050 --> 00:08:20.509
nearby realities. I love the philosophy of penalizing
00:08:20.509 --> 00:08:23.230
options that just got lucky in one specific scenario,
00:08:23.430 --> 00:08:27.290
but I have a massive glaring structural problem
00:08:27.290 --> 00:08:30.730
with this. Oh, let's hear it. Calculating a winner
00:08:30.730 --> 00:08:34.120
across infinite... plausible alternative realities
00:08:34.120 --> 00:08:36.419
sounds like an absolute computational nightmare.
00:08:36.659 --> 00:08:39.460
Oh, you are not wrong. I mean, if I am running
00:08:39.460 --> 00:08:42.220
a data engineering pipeline, you are asking me
00:08:42.220 --> 00:08:44.919
to not just parse my test data, but to somehow
00:08:44.919 --> 00:08:47.759
simulate millions of different ways the future
00:08:47.759 --> 00:08:50.620
user base could drift. And average out the worst
00:08:50.620 --> 00:08:53.519
case scenarios? Yeah, it sounds impossible. That
00:08:53.519 --> 00:08:56.480
would require immense server power. The compute
00:08:56.480 --> 00:08:59.000
cost alone would mean you'd spend weeks just
00:08:59.000 --> 00:09:01.039
deciding what shade of blue to make a hyperlink.
00:09:01.549 --> 00:09:03.289
No engineering team is going to approve that.
00:09:03.470 --> 00:09:06.870
And your skepticism is entirely justified. That
00:09:06.870 --> 00:09:09.289
exact computational bottleneck is the primary
00:09:09.289 --> 00:09:11.950
reason this kind of robust decision -making has
00:09:11.950 --> 00:09:14.029
been kept out of mainstream A -B testing for
00:09:14.029 --> 00:09:17.389
so long. It's just too heavy. Right. Finding
00:09:17.389 --> 00:09:20.169
the absolute worst -case scenario across an infinite
00:09:20.169 --> 00:09:22.129
neighborhood of possible future distributions
00:09:22.129 --> 00:09:25.950
is an intractable optimization problem. It would
00:09:25.950 --> 00:09:28.950
completely crash a standard server. So how do
00:09:28.950 --> 00:09:31.190
they get around it? Well, this brings us to the
00:09:31.190 --> 00:09:33.629
most significant breakthrough in the paper. To
00:09:33.629 --> 00:09:36.509
bypass that infinite simulation, the authors
00:09:36.509 --> 00:09:38.889
apply a concept from large deviations theory
00:09:38.889 --> 00:09:41.990
called the Donsker -Varden variational representation.
00:09:43.169 --> 00:09:46.049
Donsker -Varden? That sounds like something you'd
00:09:46.049 --> 00:09:48.509
need a PhD in advanced physics to even pronounce,
00:09:48.649 --> 00:09:51.370
let alone code. It is definitely a mouthful.
00:09:51.669 --> 00:09:53.629
Please tell me we don't have to simulate the
00:09:53.629 --> 00:09:55.850
multiverse. You don't have to simulate anything.
00:09:56.269 --> 00:09:59.250
The Donsker -Varden representation is essentially
00:09:59.250 --> 00:10:02.220
a profound mathematical shortcut. A shortcut.
00:10:02.919 --> 00:10:05.840
How does it work? Here is how it works under
00:10:05.840 --> 00:10:09.440
the hood. Instead of forcing you to painstakingly
00:10:09.440 --> 00:10:11.899
simulate every single possible path the future
00:10:11.899 --> 00:10:14.559
might take, this theorem finds a mathematical
00:10:14.559 --> 00:10:18.120
upper bound on your worst case risk. Okay. It
00:10:18.120 --> 00:10:20.419
recognizes that the probability of extreme drift
00:10:20.419 --> 00:10:24.159
decays in a very predictable way. So it takes
00:10:24.159 --> 00:10:26.799
that infinite, impossible search over every possible
00:10:26.799 --> 00:10:29.850
future. and it completely collapses it into a
00:10:29.850 --> 00:10:33.049
single, closed -form algebraic expression. Wait,
00:10:33.190 --> 00:10:35.649
wait. Here's where it gets really interesting.
00:10:36.049 --> 00:10:38.889
You're saying it collapses infinite possibilities
00:10:38.889 --> 00:10:43.110
into just a standard equation? Yes. It expresses
00:10:43.110 --> 00:10:45.809
the penalty for that future ambiguity purely
00:10:45.809 --> 00:10:48.049
as a function of the variance in your existing
00:10:48.049 --> 00:10:51.269
data. Oh, wow. If your test data has massive
00:10:51.269 --> 00:10:53.789
variance, the formula automatically applies a
00:10:53.789 --> 00:10:56.460
heavy penalty. recognizing that the feature is
00:10:56.460 --> 00:10:59.299
highly vulnerable to drift. If the variance is
00:10:59.299 --> 00:11:03.100
tight, the penalty is small. That is wild. The
00:11:03.100 --> 00:11:05.740
brilliant part is that the result is a deployment
00:11:05.740 --> 00:11:08.940
rule that needs literally no new instrumentation.
00:11:09.179 --> 00:11:11.279
You do not need to rewrite your data collection
00:11:11.279 --> 00:11:14.460
pipeline. You just need the exact same arm -level
00:11:14.460 --> 00:11:16.919
outcome data that your standard t -test is already
00:11:16.919 --> 00:11:20.320
using. Wait, really? So I can just take the exact
00:11:20.320 --> 00:11:22.740
same dashboard data my team is already collecting,
00:11:22.940 --> 00:11:26.080
throw away the t -test, and run those same numbers
00:11:26.080 --> 00:11:29.240
through this Donsker -Vardon formula with one
00:11:29.240 --> 00:11:32.679
single critical new input. The formula requires
00:11:32.679 --> 00:11:35.220
the decision maker to set what is called a trust
00:11:35.220 --> 00:11:38.960
parameter. A trust parameter, like a volume dial
00:11:38.960 --> 00:11:42.639
for risk? That is a perfect analogy. It is a
00:11:42.639 --> 00:11:45.440
literal dial that controls how ambiguity -averse
00:11:45.440 --> 00:11:48.100
the business wants to be. Okay, so how does the
00:11:48.100 --> 00:11:51.100
dial work? Well, if you turn the dial all the
00:11:51.100 --> 00:11:54.159
way down to zero, you are stating that you completely
00:11:54.159 --> 00:11:57.240
trust the test data and expect zero environmental
00:11:57.240 --> 00:11:59.759
drift. At that setting, the formula basically
00:11:59.759 --> 00:12:02.600
reverts to standard, naive, best arm selection.
00:12:02.919 --> 00:12:05.379
You just pick the winner from the test. Right,
00:12:05.460 --> 00:12:08.220
you're back to the old way. Exactly. But as you
00:12:08.220 --> 00:12:11.019
turn that dial up, as you acknowledge that your
00:12:11.019 --> 00:12:13.580
real -world environment is volatile and you want
00:12:13.580 --> 00:12:16.220
to be protected against drift, the formula becomes
00:12:16.220 --> 00:12:19.620
increasingly rigorous. So it filters out the
00:12:19.620 --> 00:12:22.039
weak ones. Yes. It will block the deployment
00:12:22.039 --> 00:12:25.059
of a feature unless it possesses a robust, uncertainty
00:12:25.059 --> 00:12:27.919
-adjusted edge that can survive that volatility.
00:12:28.379 --> 00:12:31.360
That is incredibly elegant. It gives the executive
00:12:31.360 --> 00:12:33.840
team a literal knob to turn based on their financial
00:12:33.840 --> 00:12:35.860
risk tolerance without having to build a supercomputer
00:12:35.860 --> 00:12:38.919
to run the simulations. Precisely. It bridges
00:12:38.919 --> 00:12:42.600
that gap perfectly. But an elegant theory on
00:12:42.600 --> 00:12:45.159
paper is still just theory. I want to know if
00:12:45.159 --> 00:12:47.480
this actually survives in the wild. Because the
00:12:47.480 --> 00:12:49.279
authors didn't just publish the math, did they?
00:12:49.720 --> 00:12:52.259
No, they took it much further. And this is where
00:12:52.259 --> 00:12:55.379
the paper moves from interesting theory to industry
00:12:55.379 --> 00:12:58.539
-altering practice. They validated their robust
00:12:58.539 --> 00:13:02.700
rule against 552 real digital advertising experiments.
00:13:04.039 --> 00:13:07.059
552. And just to emphasize this for you listening,
00:13:07.240 --> 00:13:10.039
we are talking about clean, simulated laboratory
00:13:10.039 --> 00:13:13.179
data. These experiments were run on a major U
00:13:13.179 --> 00:13:16.200
.S. online platform. Real -world data. Yeah.
00:13:16.259 --> 00:13:19.919
We are talking about messy, noisy, chaotic human
00:13:19.919 --> 00:13:22.440
behavior. Ad placements, click -through rates,
00:13:22.580 --> 00:13:25.080
revenue per impression. The kind of data that
00:13:25.080 --> 00:13:28.110
is notoriously prone to drift. And the results
00:13:28.110 --> 00:13:30.590
of applying this ambiguity -averse rule to those
00:13:30.590 --> 00:13:34.649
552 real -world tests were definitive. What did
00:13:34.649 --> 00:13:36.990
they find? The research shows that deploying
00:13:36.990 --> 00:13:39.669
features via this new rule substantially reduces
00:13:39.669 --> 00:13:42.929
what decision theorists call regret, especially
00:13:42.929 --> 00:13:45.070
when compared to conventional hypothesis testing.
00:13:45.470 --> 00:13:47.649
Okay, let's do a quick jargon buster on regret.
00:13:47.850 --> 00:13:50.429
Because in normal life, regret is buying a gym
00:13:50.429 --> 00:13:53.149
membership you never use. What does regret mean
00:13:53.149 --> 00:13:56.289
mathematically for these ad platforms? In decision
00:13:56.289 --> 00:13:59.129
theory, regret is a highly specific measurable
00:13:59.129 --> 00:14:02.590
metric. It is the gap between the economic payoff
00:14:02.590 --> 00:14:05.049
you actually achieved by following your deployment
00:14:05.049 --> 00:14:07.750
rule versus the payoff you would have achieved
00:14:07.750 --> 00:14:10.090
if you had made the absolute perfect choice in
00:14:10.090 --> 00:14:13.250
hindsight. Oh, I see. It essentially measures
00:14:13.250 --> 00:14:16.159
the financial cost of your mistakes. So to say
00:14:16.159 --> 00:14:19.899
this formula reduces regret means that the Donsko
00:14:19.899 --> 00:14:22.259
-Vorodan equation made choices that were significantly
00:14:22.259 --> 00:14:25.379
closer to the perfect hindsight choices than
00:14:25.379 --> 00:14:27.980
the standard t -test did. Let's look at the mechanics
00:14:27.980 --> 00:14:31.059
of those ad tests to make it clear. In a typical
00:14:31.059 --> 00:14:33.639
batch of experiments, a standard t -test might
00:14:33.639 --> 00:14:37.799
flag 100 new ad features as statistically significant
00:14:37.799 --> 00:14:41.299
winners. Right. P is less than 0 .05 for all
00:14:41.299 --> 00:14:45.000
of them. Exactly. A naive pipeline deploys all
00:14:45.000 --> 00:14:48.620
100. But the robust rule, applying its penalty
00:14:48.620 --> 00:14:51.659
for variance and ambiguity, might look at those
00:14:51.659 --> 00:14:55.580
same 100 and say, actually, 40 of these are incredibly
00:14:55.580 --> 00:14:58.440
fragile. If the user base shifts even slightly,
00:14:58.740 --> 00:15:01.580
these will lose money. So it spots the traps.
00:15:01.879 --> 00:15:05.419
Yes. So it only deploys the 60 robust winners.
00:15:05.700 --> 00:15:09.580
It acts as a much smarter filter. If we connect
00:15:09.580 --> 00:15:12.159
this to the bigger picture. Over the lifespan
00:15:12.159 --> 00:15:16.080
of those 552 experiments, the robust rule simply
00:15:16.080 --> 00:15:19.019
left less money on the table. It prevented the
00:15:19.019 --> 00:15:22.139
platform from deploying those 40 fragile features
00:15:22.139 --> 00:15:24.039
that would have eventually crashed and burned
00:15:24.039 --> 00:15:26.679
when the environment inevitably drifted. Wow.
00:15:26.799 --> 00:15:29.899
It minimized the downside risk without sacrificing
00:15:29.899 --> 00:15:33.299
the genuine robust wins. And because it is a
00:15:33.299 --> 00:15:36.220
closed -form algebraic formula, the transferability
00:15:36.220 --> 00:15:39.759
of this is staggering. It really is. Any data
00:15:39.759 --> 00:15:41.600
science team currently gating their deployments
00:15:41.600 --> 00:15:44.159
on p -values can take Ferrell, Korgambikova,
00:15:44.179 --> 00:15:46.460
and Mishra's paper, extract the formula, and
00:15:46.460 --> 00:15:49.240
swap it into their pipeline this afternoon. Literally
00:15:49.240 --> 00:15:51.740
today. The impact is immediate. You don't have
00:15:51.740 --> 00:15:53.940
to re -architect your data warehouse. You just
00:15:53.940 --> 00:15:55.860
change the math at the very final checkpoint
00:15:55.860 --> 00:15:58.679
of the pipeline. It slots directly into the existing
00:15:58.679 --> 00:16:01.620
infrastructure. It is simply a better calibrated,
00:16:01.720 --> 00:16:04.200
reality -adjusted lens applied to the data you
00:16:04.200 --> 00:16:06.659
are already spending millions of dollars to collect.
00:16:07.080 --> 00:16:09.659
Given how clean the implementation is and how
00:16:09.659 --> 00:16:11.860
clear the financial benefit is in reducing regret,
00:16:12.139 --> 00:16:15.620
I have to ask, how genuinely novel is this? Are
00:16:15.620 --> 00:16:18.639
we just rebranding an old idea or is this a fundamental
00:16:18.639 --> 00:16:21.460
leap forward? That is a crucial distinction to
00:16:21.460 --> 00:16:25.019
make. If we evaluate the novelty here, we have
00:16:25.019 --> 00:16:27.179
to separate the complaint. From the solution.
00:16:27.460 --> 00:16:30.139
Okay, let's start with the complaint. The complaint
00:16:30.139 --> 00:16:33.200
itself, the idea that p -values are poorly suited
00:16:33.200 --> 00:16:36.600
for business deployment, is very old. It is practically
00:16:36.600 --> 00:16:39.720
a genre of blog posts in the tech industry. Right,
00:16:39.759 --> 00:16:42.620
everyone knows the t -test has limitations. But
00:16:42.620 --> 00:16:44.980
if this paper merely reiterated that t -tests
00:16:44.980 --> 00:16:47.559
are flawed, it wouldn't be worth our time. The
00:16:47.559 --> 00:16:50.159
massive advance here, what makes this a 4 out
00:16:50.159 --> 00:16:52.919
of 5 on the novelty scale, is that they engineered
00:16:52.919 --> 00:16:56.029
a highly specific drop -in replacement. They
00:16:56.029 --> 00:16:58.389
actually built the fix. They breached a massive
00:16:58.389 --> 00:17:01.529
gap between incredibly abstract, dense, large
00:17:01.529 --> 00:17:04.049
deviations mathematics, the Donsker -Varadin
00:17:04.049 --> 00:17:07.490
representation, and an immediate practical engineering
00:17:07.490 --> 00:17:10.809
solution. Proving it on a massive scale with
00:17:10.809 --> 00:17:13.950
live ad platform data is exceptionally rare for
00:17:13.950 --> 00:17:16.829
a paper rooted in theoretical math. It actually
00:17:16.829 --> 00:17:19.250
reminds me of a broader shift we're seeing across
00:17:19.250 --> 00:17:21.349
the tech sector. We've talked in previous deep
00:17:21.349 --> 00:17:24.359
dives about major companies like Spotify. overhauling
00:17:24.359 --> 00:17:26.359
their experimentation platforms. Yeah, that's
00:17:26.359 --> 00:17:28.579
been a huge trend. Sometimes they're moving entirely
00:17:28.579 --> 00:17:31.539
away from Bayesian A -B testing because it just
00:17:31.539 --> 00:17:35.119
gets too complex. There is this constant, exhausting
00:17:35.119 --> 00:17:38.019
industry churn of trying to find the right statistical
00:17:38.019 --> 00:17:40.180
framework. And that highlights another reason
00:17:40.180 --> 00:17:43.119
this research is so refreshing. It brilliantly
00:17:43.119 --> 00:17:45.980
sidesteps the entire decades -old religious war
00:17:45.980 --> 00:17:47.960
between frequentist and Bayesian statistics.
00:17:48.480 --> 00:17:50.759
Wait, really? How does it manage to avoid that
00:17:50.759 --> 00:17:53.809
debate entirely? Well... Because both frequentists
00:17:53.809 --> 00:17:56.089
and Bayesians are fundamentally arguing about
00:17:56.089 --> 00:17:58.329
the best way to estimate the truth of the experiment
00:17:58.329 --> 00:18:01.430
itself. They are fighting over the science of
00:18:01.430 --> 00:18:04.650
inference. Okay. This paper steps back and says,
00:18:04.730 --> 00:18:07.630
we actually don't care how you estimate the initial
00:18:07.630 --> 00:18:10.430
test. We are entirely reframing the nature of
00:18:10.430 --> 00:18:13.490
the decision that happens after the test. Oh,
00:18:13.529 --> 00:18:17.710
that's brilliant. If focus is... purely on deployment
00:18:17.710 --> 00:18:20.589
as a robust decision problem. Treating your initial
00:18:20.589 --> 00:18:23.450
data merely as an input, regardless of how you
00:18:23.450 --> 00:18:27.109
sourced it. That makes so much sense. It's like
00:18:27.109 --> 00:18:30.029
two meteorologists arguing over what brand of
00:18:30.029 --> 00:18:32.910
thermometer is the most accurate, and this paper
00:18:32.910 --> 00:18:35.130
walks in and says, it honestly doesn't matter
00:18:35.130 --> 00:18:37.210
which thermometer you use, just pack a sweater
00:18:37.210 --> 00:18:39.930
because the weather is going to change. I couldn't
00:18:39.930 --> 00:18:42.250
have phrased it better myself. And intellectually,
00:18:42.609 --> 00:18:45.089
while the application is new, the philosophical
00:18:45.089 --> 00:18:48.750
ancestry of this idea is deeply respected. Where
00:18:48.750 --> 00:18:51.589
does it come from? It traces back to classical
00:18:51.589 --> 00:18:54.329
statistics, specifically to Wald's statistical
00:18:54.329 --> 00:18:57.029
decision theory and the concept of minimax regret
00:18:57.029 --> 00:18:59.970
from the mid -20th century. Oh, minimax regret.
00:19:00.329 --> 00:19:03.410
Right. The foundational idea is that you should
00:19:03.410 --> 00:19:06.900
evaluate a decision rule by its worst case. ambiguity
00:19:06.900 --> 00:19:10.059
-penalized performance rather than blindly assuming
00:19:10.059 --> 00:19:13.079
average case behavior under one perfect utopian
00:19:13.079 --> 00:19:16.359
scenario. So they took an old battle -tested
00:19:16.359 --> 00:19:19.140
philosophical stance on risk management, updated
00:19:19.140 --> 00:19:21.740
it with modern optimization math, and applied
00:19:21.740 --> 00:19:24.160
it to the hyper -modern problem of digital A
00:19:24.160 --> 00:19:27.619
-B testing. Exactly. But let's be clear on the
00:19:27.619 --> 00:19:29.819
limitations here, too. This is not a magic bullet
00:19:29.819 --> 00:19:32.700
for bad experimental design, is it? Definitely
00:19:32.700 --> 00:19:36.470
not. That is a vital caveat. This framework only
00:19:36.470 --> 00:19:39.369
dictates what to do after you have your results.
00:19:39.650 --> 00:19:43.410
Right. Garbage in, garbage out. Always. If you
00:19:43.410 --> 00:19:46.329
are tracking the wrong metrics or testing on
00:19:46.329 --> 00:19:49.450
a heavily biased audience segment, this formula
00:19:49.450 --> 00:19:52.190
won't save you. There is a whole separate discipline
00:19:52.190 --> 00:19:55.529
dedicated to experiment design. So you still
00:19:55.529 --> 00:19:57.250
have to do the hard work up front. Absolutely.
00:19:57.690 --> 00:20:00.329
This paper assumes you already know how to run
00:20:00.329 --> 00:20:03.430
a clean test. It simply fixes the final crucial
00:20:03.430 --> 00:20:06.930
step. the deployment trigger. So what does this
00:20:06.930 --> 00:20:08.450
all mean for the people listening right now?
00:20:09.049 --> 00:20:11.789
Obviously, a solo developer testing button colors
00:20:11.789 --> 00:20:13.690
on a personal blog probably doesn't need to stress
00:20:13.690 --> 00:20:16.750
over the Danske Veraden representation. True,
00:20:16.869 --> 00:20:19.329
the stakes there are pretty low. But anyone making
00:20:19.329 --> 00:20:21.490
high -stakes, capital -intensive choices based
00:20:21.490 --> 00:20:24.329
on data absolutely needs to pay attention. So
00:20:24.329 --> 00:20:27.029
product managers, data science leads? Exactly.
00:20:27.150 --> 00:20:29.849
If you are a product manager, an executive at
00:20:29.849 --> 00:20:32.079
an e -commerce giant, or a marketing director
00:20:32.079 --> 00:20:34.220
allocating millions of dollars based on A -B
00:20:34.220 --> 00:20:36.839
test readouts, you need to recognize that your
00:20:36.839 --> 00:20:40.119
current tools are likely ignoring future volatility.
00:20:40.619 --> 00:20:43.000
You could be shipping those fragile winners.
00:20:43.400 --> 00:20:45.619
You might be systematically deploying fragile
00:20:45.619 --> 00:20:47.619
winners that are slowly bleeding your margins
00:20:47.619 --> 00:20:50.240
because they simply cannot survive in the wild.
00:20:50.579 --> 00:20:53.819
It is a complete paradigm shift. We are looking
00:20:53.819 --> 00:20:55.900
at a clear path to move away from the flawed,
00:20:55.960 --> 00:20:58.819
fragile world where we pretend P less than 0
00:20:58.819 --> 00:21:02.240
.05 equals guaranteed future success. Yeah, that
00:21:02.240 --> 00:21:06.079
world is ending. We can move into a robust, ambiguity
00:21:06.079 --> 00:21:09.259
-averse future using the Donsker -Varadin formula
00:21:09.259 --> 00:21:11.960
to essentially immunize our deployment decisions
00:21:11.960 --> 00:21:14.640
against the chaos of the real world, fundamentally
00:21:14.640 --> 00:21:18.480
minimizing our regret. And looking ahead, this
00:21:18.480 --> 00:21:20.460
raises an important question that goes entirely
00:21:20.460 --> 00:21:22.559
beyond what the authors covered in this paper.
00:21:23.039 --> 00:21:26.700
Oh, lay it on us. Well, we've established that
00:21:26.700 --> 00:21:28.720
the real world drifts away from test environments
00:21:28.720 --> 00:21:31.700
naturally, right? But think about the era we
00:21:31.700 --> 00:21:35.019
are entering right now with AI -generated hyper
00:21:35.019 --> 00:21:38.140
-personalized content. The digital environment
00:21:38.140 --> 00:21:41.160
isn't just drifting organically anymore. It is
00:21:41.160 --> 00:21:44.019
being actively algorithmically mutated by the
00:21:44.019 --> 00:21:48.319
second. Oh, wow. That's a scary thought. If the
00:21:48.319 --> 00:21:50.400
environment shifts so rapidly that the future
00:21:50.400 --> 00:21:53.519
looks entirely alien to the past, can even a
00:21:53.519 --> 00:21:56.859
robust ambiguity -averse formula keep up? Or
00:21:56.859 --> 00:21:58.980
will the speed of AI -driven drift eventually
00:21:58.980 --> 00:22:01.900
outpace our ability to mathematically penalize
00:22:01.900 --> 00:22:05.799
it? That is a staggering thought. If the ground
00:22:05.799 --> 00:22:08.140
is shifting that fast, even packing layers might
00:22:08.140 --> 00:22:11.980
not be enough. We'll have to see. Thank you for
00:22:11.980 --> 00:22:13.980
joining us on this deep dive. Keep questioning
00:22:13.980 --> 00:22:16.079
the metrics you rely on, and we will catch you
00:22:16.079 --> 00:22:16.519
next time.
00:00:00.000 --> 00:00:03.399
Imagine this scenario for a second. You are overseeing
00:00:03.399 --> 00:00:06.139
a massive A -B test for a new product feature.
00:00:06.299 --> 00:00:09.820
Oh, yeah. The classic tech setup. Right. So the
00:00:09.820 --> 00:00:11.820
engineering team runs the numbers. You check
00:00:11.820 --> 00:00:14.859
the dashboard. And the data is just flawless.
00:00:15.080 --> 00:00:18.000
I mean, the p -value is practically zero. So
00:00:18.000 --> 00:00:20.980
you pop the champagne. Exactly. You deploy the
00:00:20.980 --> 00:00:25.100
new feature. You celebrate. And then, well, next
00:00:25.100 --> 00:00:27.000
quarter, the company loses millions of dollars.
00:00:27.059 --> 00:00:29.730
And literally nobody can explain why. It's the
00:00:29.730 --> 00:00:32.250
ultimate nightmare for any data -driven team,
00:00:32.329 --> 00:00:34.829
really. You followed all the rules, the math
00:00:34.829 --> 00:00:37.530
said you had a clear winner, but the real -world
00:00:37.530 --> 00:00:40.649
deployment was just a complete disaster. And
00:00:40.649 --> 00:00:43.350
that silent, expensive disaster is exactly what
00:00:43.350 --> 00:00:45.770
we are tearing apart today. Welcome to the Deep
00:00:45.770 --> 00:00:48.600
Dive. Glad to be here. We are looking at a brand
00:00:48.600 --> 00:00:51.640
new paper from September 2026 by Max Farrell,
00:00:51.740 --> 00:00:55.420
Malika Korgan -Bekova, and Sanjog Misra. It's
00:00:55.420 --> 00:00:58.299
called Robust AB Decisions. Yeah, and it is a
00:00:58.299 --> 00:01:01.479
fascinating read. It really is. We're going to
00:01:01.479 --> 00:01:03.560
unpack why the most common battle cry in tech
00:01:03.560 --> 00:01:06.079
and business, you know, P is less than 0 .05,
00:01:06.239 --> 00:01:08.719
ship it, might actually be answering the completely
00:01:08.719 --> 00:01:10.980
wrong question. And that's a big deal because,
00:01:11.040 --> 00:01:14.780
well, entire companies run on that rule. Exactly.
00:01:15.450 --> 00:01:18.329
So if you have ever relied on an A -B test to
00:01:18.329 --> 00:01:20.950
make a decision, whether you are prepping for
00:01:20.950 --> 00:01:23.909
a big board meeting, launching a new user interface,
00:01:24.170 --> 00:01:26.829
or you're just insanely curious about how these
00:01:26.829 --> 00:01:29.849
massive digital platforms operate, this research
00:01:29.849 --> 00:01:32.849
reveals a massive blind spot in that entire process.
00:01:33.209 --> 00:01:35.489
What makes this paper such a breakthrough for
00:01:35.489 --> 00:01:37.689
the industry isn't just that it points out a
00:01:37.689 --> 00:01:40.310
flaw in standard hypothesis testing, you know.
00:01:41.019 --> 00:01:42.959
I mean, people have been complaining about statistical
00:01:42.959 --> 00:01:45.219
significance for decades. Right. It's not a new
00:01:45.219 --> 00:01:48.540
complaint. Exactly. What we are going to explore
00:01:48.540 --> 00:01:52.180
today is a genuinely new practical solution rooted
00:01:52.180 --> 00:01:55.400
in something called robust decision theory. Robust
00:01:55.400 --> 00:01:58.120
decision theory. Yeah. Okay. We are shifting
00:01:58.120 --> 00:02:00.359
the goalpost for merely detecting a mathematical
00:02:00.359 --> 00:02:03.379
difference in a vacuum to actually optimizing
00:02:03.379 --> 00:02:06.340
for economic survival in a highly unpredictable,
00:02:06.739 --> 00:02:09.629
messy world. Okay, let's unpack this standard
00:02:09.629 --> 00:02:11.889
industry workflow first, because you see it everywhere,
00:02:12.030 --> 00:02:14.210
from tiny startups to Fortune 50 companies. You
00:02:14.210 --> 00:02:16.930
really do. Step one, you run your A -B test.
00:02:17.150 --> 00:02:18.930
You have your control group, you have your treatment
00:02:18.930 --> 00:02:21.229
group. Step two, you check the math. Usually
00:02:21.229 --> 00:02:23.569
some form of a T -test, right? Hunting for that
00:02:23.569 --> 00:02:25.830
magical threshold of statistical significance.
00:02:26.169 --> 00:02:29.370
The holy grail. Right. And step three, if the
00:02:29.370 --> 00:02:32.069
new feature wins, you deploy it to all your users,
00:02:32.229 --> 00:02:35.849
permanently. But reading through Farrell, Korganbekova,
00:02:35.990 --> 00:02:38.569
and Mishra's work, I realized how flawed this
00:02:38.569 --> 00:02:41.860
is. Oh, it is fundamentally broken. Using a standard
00:02:41.860 --> 00:02:43.879
t -test to make a final business deployment decision
00:02:43.879 --> 00:02:47.340
feels like using a metal detector to buy a plot
00:02:47.340 --> 00:02:49.780
of land on an active fault line. Wow. Right?
00:02:49.860 --> 00:02:51.539
That's a great way to put it. Sure. It tells
00:02:51.539 --> 00:02:53.960
you there's some valuable ore right beneath your
00:02:53.960 --> 00:02:56.280
feet today, but it completely ignores the fact
00:02:56.280 --> 00:02:58.039
that the ground is mathematically guaranteed
00:02:58.039 --> 00:03:00.199
to shift tomorrow. That is a much more accurate
00:03:00.199 --> 00:03:02.979
way to visualize the danger. To understand why
00:03:02.979 --> 00:03:04.960
we are building on fault lines, you really have
00:03:04.960 --> 00:03:07.180
to look at what the t -test was actually designed
00:03:07.180 --> 00:03:09.060
to do. It wasn't built for business, was it?
00:03:09.219 --> 00:03:12.849
No, not at all. A t -test's null hypothesis framing
00:03:12.849 --> 00:03:15.370
was never built to serve as a business deployment
00:03:15.370 --> 00:03:19.569
policy. It is a scientific tool. Its original
00:03:19.569 --> 00:03:21.949
purpose was strictly to control the rate of false
00:03:21.949 --> 00:03:24.530
claims in academic literature or like clinical
00:03:24.530 --> 00:03:28.449
trials. It answers a very narrow historical question.
00:03:28.650 --> 00:03:31.750
Which is what exactly? It basically asks, did
00:03:31.750 --> 00:03:34.250
the treatment differ from the control in this
00:03:34.250 --> 00:03:37.189
exact specific sample of data that I just happened
00:03:37.189 --> 00:03:39.889
to observe over the last two weeks? In this specific
00:03:39.889 --> 00:03:43.169
sample, which is looking backward. Exactly. It's
00:03:43.169 --> 00:03:45.770
not asking, will this specific feature maximize
00:03:45.770 --> 00:03:48.009
our revenue over the next five years? Precisely.
00:03:48.009 --> 00:03:51.229
It is fundamentally not evaluating the real economic
00:03:51.229 --> 00:03:54.789
payoff. A standard t -test doesn't weigh the
00:03:54.789 --> 00:03:57.370
magnitude of your potential win against the future
00:03:57.370 --> 00:03:59.610
risks of deploying it. Yeah, that makes sense.
00:04:00.000 --> 00:04:02.020
Yet a massive share of real -world corporate
00:04:02.020 --> 00:04:05.060
decisions are gated exclusively by this metric.
00:04:05.340 --> 00:04:09.199
A product team sees a p -value below 0 .05, and
00:04:09.199 --> 00:04:11.479
they automatically ship the feature. Assuming
00:04:11.479 --> 00:04:13.439
the future will look exactly like their two -week
00:04:13.439 --> 00:04:16.860
test window. Which it... Never does. Never. Which
00:04:16.860 --> 00:04:19.180
brings us to the core danger that authors highlight,
00:04:19.379 --> 00:04:22.120
the concept of fragile winners. I want to ground
00:04:22.120 --> 00:04:24.220
this in a real example. Let's do it. Okay. So
00:04:24.220 --> 00:04:26.980
let's say you run a massive A -B test on a new
00:04:26.980 --> 00:04:29.540
checkout button design for an e -commerce platform.
00:04:29.779 --> 00:04:32.480
You run it over a long holiday weekend. Okay.
00:04:32.519 --> 00:04:35.399
Very specific time. Right. The users interacting
00:04:35.399 --> 00:04:37.680
with your site have a very specific mindset.
00:04:37.920 --> 00:04:40.060
They're browsing casually, maybe looking for
00:04:40.060 --> 00:04:42.829
sales. In that specific environment, the new
00:04:42.829 --> 00:04:45.589
button crushes the control. It is a clear winner.
00:04:45.870 --> 00:04:48.310
But what happens when you deploy that button
00:04:48.310 --> 00:04:51.189
globally and suddenly it's a Tuesday in November?
00:04:51.490 --> 00:04:54.490
Right. The users are now, I don't know, stressed
00:04:54.490 --> 00:04:57.290
corporate buyers making bulk purchases on tight
00:04:57.290 --> 00:05:00.310
deadlines. The demographics have shifted slightly.
00:05:00.589 --> 00:05:03.649
The user intent has changed. The broader economic
00:05:03.649 --> 00:05:06.569
environment might even be different. So the environment
00:05:06.569 --> 00:05:09.819
drifted. Yes. That is what the research calls
00:05:09.819 --> 00:05:13.180
distributional drift. The real world is not a
00:05:13.180 --> 00:05:16.819
sealed laboratory vacuum. It is volatile. So
00:05:16.819 --> 00:05:20.180
it's messy. Very. A fragile winter is a feature
00:05:20.180 --> 00:05:22.519
that looked absolutely fantastic in the highly
00:05:22.519 --> 00:05:25.019
specific rigid conditions of your test sample.
00:05:25.439 --> 00:05:28.000
But it falls apart the moment the environment
00:05:28.000 --> 00:05:30.759
drifts even slightly from those exact conditions.
00:05:31.040 --> 00:05:33.220
Because the t -test never checked for drift.
00:05:33.339 --> 00:05:35.300
It just confirmed that the metal detector beeped
00:05:35.300 --> 00:05:37.399
on that one specific holiday weekend. Exactly.
00:05:38.430 --> 00:05:41.509
totally blind to the future. So if fragile winners
00:05:41.509 --> 00:05:43.949
are silently bleeding corporate margins, which
00:05:43.949 --> 00:05:46.670
they are because teams deploy these features
00:05:46.670 --> 00:05:48.850
and then act shocked when the quarterly revenue
00:05:48.850 --> 00:05:52.250
mysteriously dips, how do we fix it? Right. That's
00:05:52.250 --> 00:05:54.170
the million dollar question. Literally, how do
00:05:54.170 --> 00:05:57.629
you mathematically test for a chaotic, unpredictable
00:05:57.629 --> 00:06:00.389
future? This is where the paper introduces its
00:06:00.389 --> 00:06:03.069
core paradigm shift. We have to stop treating
00:06:03.069 --> 00:06:05.910
A -B testing as a traditional hypothesis test.
00:06:06.379 --> 00:06:09.259
Okay. And start treating it as what? As a robust
00:06:09.259 --> 00:06:12.399
decision theory problem. Instead of just looking
00:06:12.399 --> 00:06:14.759
at this single point estimate, the exact numerical
00:06:14.759 --> 00:06:17.839
average you observed in your test, this new framework
00:06:17.839 --> 00:06:21.160
forces you to calculate an ambiguity penalized
00:06:21.160 --> 00:06:24.639
value. Whoa. Okay. That is a heavy bit of jargon.
00:06:24.740 --> 00:06:27.920
I know. Ambiguity penalized value, or as the
00:06:27.920 --> 00:06:29.379
authors refer to it throughout the research,
00:06:29.600 --> 00:06:32.540
ambiguity aversion. We really need to break that
00:06:32.540 --> 00:06:35.019
down for the listener. Yeah. Let's do that. What's
00:06:35.019 --> 00:06:37.540
fascinating here is that ambiguity aversion is
00:06:37.540 --> 00:06:40.839
actually a very intuitive stance on risk. Imagine
00:06:40.839 --> 00:06:42.740
you were packing for a trip to a city you've
00:06:42.740 --> 00:06:45.339
never visited before. Okay, I'm with you. A standard
00:06:45.339 --> 00:06:47.060
A -B test approach would look at the weather
00:06:47.060 --> 00:06:49.839
forecast for today, see that it says 70 degrees
00:06:49.839 --> 00:06:52.660
and sunny, take that as absolute truth, and pack
00:06:52.660 --> 00:06:55.439
only t -shirts. Relying entirely on that single
00:06:55.439 --> 00:06:58.180
point estimate. Right, but an ambiguity averse
00:06:58.180 --> 00:07:00.899
traveler does not fully trust that single forecast.
00:07:01.360 --> 00:07:04.269
They ask, what if a cold front moves in? What
00:07:04.269 --> 00:07:06.410
if the forecast is slightly wrong and it rains?
00:07:06.689 --> 00:07:10.129
So they plan for the drift. Exactly. They inherently
00:07:10.129 --> 00:07:13.050
penalize the t -shirt only option because its
00:07:13.050 --> 00:07:16.269
success relies entirely on one rigid, perfectly
00:07:16.269 --> 00:07:18.810
assumed model of the weather. Right. Instead,
00:07:18.970 --> 00:07:21.569
they prefer the option of packing layers. It
00:07:21.569 --> 00:07:23.949
might be slightly less optimal to carry a heavier
00:07:23.949 --> 00:07:26.629
suitcase if it ends up being exactly 70 degrees,
00:07:26.910 --> 00:07:29.769
but that option remains robust and functional
00:07:29.769 --> 00:07:33.029
across a wide range of plausible nearby weather
00:07:33.029 --> 00:07:35.860
scenarios. Okay, let me translate this back to
00:07:35.860 --> 00:07:38.279
the product dashboard. Under this new framework,
00:07:38.420 --> 00:07:40.899
instead of a data scientist asking, did this
00:07:40.899 --> 00:07:43.720
feature win in our exact historical test? Yeah.
00:07:43.759 --> 00:07:46.879
The deployment rule explicitly asks, would this
00:07:46.879 --> 00:07:49.779
feature still win across a neighborhood of slightly
00:07:49.779 --> 00:07:53.339
different plausible realities? You got it. That's
00:07:53.339 --> 00:07:56.000
exactly it. It evaluates the candidate feature
00:07:56.000 --> 00:07:58.879
under the exact data you collected, but also
00:07:58.879 --> 00:08:01.240
under distributions that are mathematically close
00:08:01.240 --> 00:08:04.610
to it. Oh, wow. If the feature's advantage is
00:08:04.610 --> 00:08:07.149
incredibly fragile, like if it only beats the
00:08:07.149 --> 00:08:10.149
control under highly rigid assumptions, the ambiguity
00:08:10.149 --> 00:08:13.329
-averse rule penalizes it. It deliberately selects
00:08:13.329 --> 00:08:16.050
options that remain strong across those plausible
00:08:16.050 --> 00:08:20.509
nearby realities. I love the philosophy of penalizing
00:08:20.509 --> 00:08:23.230
options that just got lucky in one specific scenario,
00:08:23.430 --> 00:08:27.290
but I have a massive glaring structural problem
00:08:27.290 --> 00:08:30.730
with this. Oh, let's hear it. Calculating a winner
00:08:30.730 --> 00:08:34.120
across infinite... plausible alternative realities
00:08:34.120 --> 00:08:36.419
sounds like an absolute computational nightmare.
00:08:36.659 --> 00:08:39.460
Oh, you are not wrong. I mean, if I am running
00:08:39.460 --> 00:08:42.220
a data engineering pipeline, you are asking me
00:08:42.220 --> 00:08:44.919
to not just parse my test data, but to somehow
00:08:44.919 --> 00:08:47.759
simulate millions of different ways the future
00:08:47.759 --> 00:08:50.620
user base could drift. And average out the worst
00:08:50.620 --> 00:08:53.519
case scenarios? Yeah, it sounds impossible. That
00:08:53.519 --> 00:08:56.480
would require immense server power. The compute
00:08:56.480 --> 00:08:59.000
cost alone would mean you'd spend weeks just
00:08:59.000 --> 00:09:01.039
deciding what shade of blue to make a hyperlink.
00:09:01.549 --> 00:09:03.289
No engineering team is going to approve that.
00:09:03.470 --> 00:09:06.870
And your skepticism is entirely justified. That
00:09:06.870 --> 00:09:09.289
exact computational bottleneck is the primary
00:09:09.289 --> 00:09:11.950
reason this kind of robust decision -making has
00:09:11.950 --> 00:09:14.029
been kept out of mainstream A -B testing for
00:09:14.029 --> 00:09:17.389
so long. It's just too heavy. Right. Finding
00:09:17.389 --> 00:09:20.169
the absolute worst -case scenario across an infinite
00:09:20.169 --> 00:09:22.129
neighborhood of possible future distributions
00:09:22.129 --> 00:09:25.950
is an intractable optimization problem. It would
00:09:25.950 --> 00:09:28.950
completely crash a standard server. So how do
00:09:28.950 --> 00:09:31.190
they get around it? Well, this brings us to the
00:09:31.190 --> 00:09:33.629
most significant breakthrough in the paper. To
00:09:33.629 --> 00:09:36.509
bypass that infinite simulation, the authors
00:09:36.509 --> 00:09:38.889
apply a concept from large deviations theory
00:09:38.889 --> 00:09:41.990
called the Donsker -Varden variational representation.
00:09:43.169 --> 00:09:46.049
Donsker -Varden? That sounds like something you'd
00:09:46.049 --> 00:09:48.509
need a PhD in advanced physics to even pronounce,
00:09:48.649 --> 00:09:51.370
let alone code. It is definitely a mouthful.
00:09:51.669 --> 00:09:53.629
Please tell me we don't have to simulate the
00:09:53.629 --> 00:09:55.850
multiverse. You don't have to simulate anything.
00:09:56.269 --> 00:09:59.250
The Donsker -Varden representation is essentially
00:09:59.250 --> 00:10:02.220
a profound mathematical shortcut. A shortcut.
00:10:02.919 --> 00:10:05.840
How does it work? Here is how it works under
00:10:05.840 --> 00:10:09.440
the hood. Instead of forcing you to painstakingly
00:10:09.440 --> 00:10:11.899
simulate every single possible path the future
00:10:11.899 --> 00:10:14.559
might take, this theorem finds a mathematical
00:10:14.559 --> 00:10:18.120
upper bound on your worst case risk. Okay. It
00:10:18.120 --> 00:10:20.419
recognizes that the probability of extreme drift
00:10:20.419 --> 00:10:24.159
decays in a very predictable way. So it takes
00:10:24.159 --> 00:10:26.799
that infinite, impossible search over every possible
00:10:26.799 --> 00:10:29.850
future. and it completely collapses it into a
00:10:29.850 --> 00:10:33.049
single, closed -form algebraic expression. Wait,
00:10:33.190 --> 00:10:35.649
wait. Here's where it gets really interesting.
00:10:36.049 --> 00:10:38.889
You're saying it collapses infinite possibilities
00:10:38.889 --> 00:10:43.110
into just a standard equation? Yes. It expresses
00:10:43.110 --> 00:10:45.809
the penalty for that future ambiguity purely
00:10:45.809 --> 00:10:48.049
as a function of the variance in your existing
00:10:48.049 --> 00:10:51.269
data. Oh, wow. If your test data has massive
00:10:51.269 --> 00:10:53.789
variance, the formula automatically applies a
00:10:53.789 --> 00:10:56.460
heavy penalty. recognizing that the feature is
00:10:56.460 --> 00:10:59.299
highly vulnerable to drift. If the variance is
00:10:59.299 --> 00:11:03.100
tight, the penalty is small. That is wild. The
00:11:03.100 --> 00:11:05.740
brilliant part is that the result is a deployment
00:11:05.740 --> 00:11:08.940
rule that needs literally no new instrumentation.
00:11:09.179 --> 00:11:11.279
You do not need to rewrite your data collection
00:11:11.279 --> 00:11:14.460
pipeline. You just need the exact same arm -level
00:11:14.460 --> 00:11:16.919
outcome data that your standard t -test is already
00:11:16.919 --> 00:11:20.320
using. Wait, really? So I can just take the exact
00:11:20.320 --> 00:11:22.740
same dashboard data my team is already collecting,
00:11:22.940 --> 00:11:26.080
throw away the t -test, and run those same numbers
00:11:26.080 --> 00:11:29.240
through this Donsker -Vardon formula with one
00:11:29.240 --> 00:11:32.679
single critical new input. The formula requires
00:11:32.679 --> 00:11:35.220
the decision maker to set what is called a trust
00:11:35.220 --> 00:11:38.960
parameter. A trust parameter, like a volume dial
00:11:38.960 --> 00:11:42.639
for risk? That is a perfect analogy. It is a
00:11:42.639 --> 00:11:45.440
literal dial that controls how ambiguity -averse
00:11:45.440 --> 00:11:48.100
the business wants to be. Okay, so how does the
00:11:48.100 --> 00:11:51.100
dial work? Well, if you turn the dial all the
00:11:51.100 --> 00:11:54.159
way down to zero, you are stating that you completely
00:11:54.159 --> 00:11:57.240
trust the test data and expect zero environmental
00:11:57.240 --> 00:11:59.759
drift. At that setting, the formula basically
00:11:59.759 --> 00:12:02.600
reverts to standard, naive, best arm selection.
00:12:02.919 --> 00:12:05.379
You just pick the winner from the test. Right,
00:12:05.460 --> 00:12:08.220
you're back to the old way. Exactly. But as you
00:12:08.220 --> 00:12:11.019
turn that dial up, as you acknowledge that your
00:12:11.019 --> 00:12:13.580
real -world environment is volatile and you want
00:12:13.580 --> 00:12:16.220
to be protected against drift, the formula becomes
00:12:16.220 --> 00:12:19.620
increasingly rigorous. So it filters out the
00:12:19.620 --> 00:12:22.039
weak ones. Yes. It will block the deployment
00:12:22.039 --> 00:12:25.059
of a feature unless it possesses a robust, uncertainty
00:12:25.059 --> 00:12:27.919
-adjusted edge that can survive that volatility.
00:12:28.379 --> 00:12:31.360
That is incredibly elegant. It gives the executive
00:12:31.360 --> 00:12:33.840
team a literal knob to turn based on their financial
00:12:33.840 --> 00:12:35.860
risk tolerance without having to build a supercomputer
00:12:35.860 --> 00:12:38.919
to run the simulations. Precisely. It bridges
00:12:38.919 --> 00:12:42.600
that gap perfectly. But an elegant theory on
00:12:42.600 --> 00:12:45.159
paper is still just theory. I want to know if
00:12:45.159 --> 00:12:47.480
this actually survives in the wild. Because the
00:12:47.480 --> 00:12:49.279
authors didn't just publish the math, did they?
00:12:49.720 --> 00:12:52.259
No, they took it much further. And this is where
00:12:52.259 --> 00:12:55.379
the paper moves from interesting theory to industry
00:12:55.379 --> 00:12:58.539
-altering practice. They validated their robust
00:12:58.539 --> 00:13:02.700
rule against 552 real digital advertising experiments.
00:13:04.039 --> 00:13:07.059
552. And just to emphasize this for you listening,
00:13:07.240 --> 00:13:10.039
we are talking about clean, simulated laboratory
00:13:10.039 --> 00:13:13.179
data. These experiments were run on a major U
00:13:13.179 --> 00:13:16.200
.S. online platform. Real -world data. Yeah.
00:13:16.259 --> 00:13:19.919
We are talking about messy, noisy, chaotic human
00:13:19.919 --> 00:13:22.440
behavior. Ad placements, click -through rates,
00:13:22.580 --> 00:13:25.080
revenue per impression. The kind of data that
00:13:25.080 --> 00:13:28.110
is notoriously prone to drift. And the results
00:13:28.110 --> 00:13:30.590
of applying this ambiguity -averse rule to those
00:13:30.590 --> 00:13:34.649
552 real -world tests were definitive. What did
00:13:34.649 --> 00:13:36.990
they find? The research shows that deploying
00:13:36.990 --> 00:13:39.669
features via this new rule substantially reduces
00:13:39.669 --> 00:13:42.929
what decision theorists call regret, especially
00:13:42.929 --> 00:13:45.070
when compared to conventional hypothesis testing.
00:13:45.470 --> 00:13:47.649
Okay, let's do a quick jargon buster on regret.
00:13:47.850 --> 00:13:50.429
Because in normal life, regret is buying a gym
00:13:50.429 --> 00:13:53.149
membership you never use. What does regret mean
00:13:53.149 --> 00:13:56.289
mathematically for these ad platforms? In decision
00:13:56.289 --> 00:13:59.129
theory, regret is a highly specific measurable
00:13:59.129 --> 00:14:02.590
metric. It is the gap between the economic payoff
00:14:02.590 --> 00:14:05.049
you actually achieved by following your deployment
00:14:05.049 --> 00:14:07.750
rule versus the payoff you would have achieved
00:14:07.750 --> 00:14:10.090
if you had made the absolute perfect choice in
00:14:10.090 --> 00:14:13.250
hindsight. Oh, I see. It essentially measures
00:14:13.250 --> 00:14:16.159
the financial cost of your mistakes. So to say
00:14:16.159 --> 00:14:19.899
this formula reduces regret means that the Donsko
00:14:19.899 --> 00:14:22.259
-Vorodan equation made choices that were significantly
00:14:22.259 --> 00:14:25.379
closer to the perfect hindsight choices than
00:14:25.379 --> 00:14:27.980
the standard t -test did. Let's look at the mechanics
00:14:27.980 --> 00:14:31.059
of those ad tests to make it clear. In a typical
00:14:31.059 --> 00:14:33.639
batch of experiments, a standard t -test might
00:14:33.639 --> 00:14:37.799
flag 100 new ad features as statistically significant
00:14:37.799 --> 00:14:41.299
winners. Right. P is less than 0 .05 for all
00:14:41.299 --> 00:14:45.000
of them. Exactly. A naive pipeline deploys all
00:14:45.000 --> 00:14:48.620
100. But the robust rule, applying its penalty
00:14:48.620 --> 00:14:51.659
for variance and ambiguity, might look at those
00:14:51.659 --> 00:14:55.580
same 100 and say, actually, 40 of these are incredibly
00:14:55.580 --> 00:14:58.440
fragile. If the user base shifts even slightly,
00:14:58.740 --> 00:15:01.580
these will lose money. So it spots the traps.
00:15:01.879 --> 00:15:05.419
Yes. So it only deploys the 60 robust winners.
00:15:05.700 --> 00:15:09.580
It acts as a much smarter filter. If we connect
00:15:09.580 --> 00:15:12.159
this to the bigger picture. Over the lifespan
00:15:12.159 --> 00:15:16.080
of those 552 experiments, the robust rule simply
00:15:16.080 --> 00:15:19.019
left less money on the table. It prevented the
00:15:19.019 --> 00:15:22.139
platform from deploying those 40 fragile features
00:15:22.139 --> 00:15:24.039
that would have eventually crashed and burned
00:15:24.039 --> 00:15:26.679
when the environment inevitably drifted. Wow.
00:15:26.799 --> 00:15:29.899
It minimized the downside risk without sacrificing
00:15:29.899 --> 00:15:33.299
the genuine robust wins. And because it is a
00:15:33.299 --> 00:15:36.220
closed -form algebraic formula, the transferability
00:15:36.220 --> 00:15:39.759
of this is staggering. It really is. Any data
00:15:39.759 --> 00:15:41.600
science team currently gating their deployments
00:15:41.600 --> 00:15:44.159
on p -values can take Ferrell, Korgambikova,
00:15:44.179 --> 00:15:46.460
and Mishra's paper, extract the formula, and
00:15:46.460 --> 00:15:49.240
swap it into their pipeline this afternoon. Literally
00:15:49.240 --> 00:15:51.740
today. The impact is immediate. You don't have
00:15:51.740 --> 00:15:53.940
to re -architect your data warehouse. You just
00:15:53.940 --> 00:15:55.860
change the math at the very final checkpoint
00:15:55.860 --> 00:15:58.679
of the pipeline. It slots directly into the existing
00:15:58.679 --> 00:16:01.620
infrastructure. It is simply a better calibrated,
00:16:01.720 --> 00:16:04.200
reality -adjusted lens applied to the data you
00:16:04.200 --> 00:16:06.659
are already spending millions of dollars to collect.
00:16:07.080 --> 00:16:09.659
Given how clean the implementation is and how
00:16:09.659 --> 00:16:11.860
clear the financial benefit is in reducing regret,
00:16:12.139 --> 00:16:15.620
I have to ask, how genuinely novel is this? Are
00:16:15.620 --> 00:16:18.639
we just rebranding an old idea or is this a fundamental
00:16:18.639 --> 00:16:21.460
leap forward? That is a crucial distinction to
00:16:21.460 --> 00:16:25.019
make. If we evaluate the novelty here, we have
00:16:25.019 --> 00:16:27.179
to separate the complaint. From the solution.
00:16:27.460 --> 00:16:30.139
Okay, let's start with the complaint. The complaint
00:16:30.139 --> 00:16:33.200
itself, the idea that p -values are poorly suited
00:16:33.200 --> 00:16:36.600
for business deployment, is very old. It is practically
00:16:36.600 --> 00:16:39.720
a genre of blog posts in the tech industry. Right,
00:16:39.759 --> 00:16:42.620
everyone knows the t -test has limitations. But
00:16:42.620 --> 00:16:44.980
if this paper merely reiterated that t -tests
00:16:44.980 --> 00:16:47.559
are flawed, it wouldn't be worth our time. The
00:16:47.559 --> 00:16:50.159
massive advance here, what makes this a 4 out
00:16:50.159 --> 00:16:52.919
of 5 on the novelty scale, is that they engineered
00:16:52.919 --> 00:16:56.029
a highly specific drop -in replacement. They
00:16:56.029 --> 00:16:58.389
actually built the fix. They breached a massive
00:16:58.389 --> 00:17:01.529
gap between incredibly abstract, dense, large
00:17:01.529 --> 00:17:04.049
deviations mathematics, the Donsker -Varadin
00:17:04.049 --> 00:17:07.490
representation, and an immediate practical engineering
00:17:07.490 --> 00:17:10.809
solution. Proving it on a massive scale with
00:17:10.809 --> 00:17:13.950
live ad platform data is exceptionally rare for
00:17:13.950 --> 00:17:16.829
a paper rooted in theoretical math. It actually
00:17:16.829 --> 00:17:19.250
reminds me of a broader shift we're seeing across
00:17:19.250 --> 00:17:21.349
the tech sector. We've talked in previous deep
00:17:21.349 --> 00:17:24.359
dives about major companies like Spotify. overhauling
00:17:24.359 --> 00:17:26.359
their experimentation platforms. Yeah, that's
00:17:26.359 --> 00:17:28.579
been a huge trend. Sometimes they're moving entirely
00:17:28.579 --> 00:17:31.539
away from Bayesian A -B testing because it just
00:17:31.539 --> 00:17:35.119
gets too complex. There is this constant, exhausting
00:17:35.119 --> 00:17:38.019
industry churn of trying to find the right statistical
00:17:38.019 --> 00:17:40.180
framework. And that highlights another reason
00:17:40.180 --> 00:17:43.119
this research is so refreshing. It brilliantly
00:17:43.119 --> 00:17:45.980
sidesteps the entire decades -old religious war
00:17:45.980 --> 00:17:47.960
between frequentist and Bayesian statistics.
00:17:48.480 --> 00:17:50.759
Wait, really? How does it manage to avoid that
00:17:50.759 --> 00:17:53.809
debate entirely? Well... Because both frequentists
00:17:53.809 --> 00:17:56.089
and Bayesians are fundamentally arguing about
00:17:56.089 --> 00:17:58.329
the best way to estimate the truth of the experiment
00:17:58.329 --> 00:18:01.430
itself. They are fighting over the science of
00:18:01.430 --> 00:18:04.650
inference. Okay. This paper steps back and says,
00:18:04.730 --> 00:18:07.630
we actually don't care how you estimate the initial
00:18:07.630 --> 00:18:10.430
test. We are entirely reframing the nature of
00:18:10.430 --> 00:18:13.490
the decision that happens after the test. Oh,
00:18:13.529 --> 00:18:17.710
that's brilliant. If focus is... purely on deployment
00:18:17.710 --> 00:18:20.589
as a robust decision problem. Treating your initial
00:18:20.589 --> 00:18:23.450
data merely as an input, regardless of how you
00:18:23.450 --> 00:18:27.109
sourced it. That makes so much sense. It's like
00:18:27.109 --> 00:18:30.029
two meteorologists arguing over what brand of
00:18:30.029 --> 00:18:32.910
thermometer is the most accurate, and this paper
00:18:32.910 --> 00:18:35.130
walks in and says, it honestly doesn't matter
00:18:35.130 --> 00:18:37.210
which thermometer you use, just pack a sweater
00:18:37.210 --> 00:18:39.930
because the weather is going to change. I couldn't
00:18:39.930 --> 00:18:42.250
have phrased it better myself. And intellectually,
00:18:42.609 --> 00:18:45.089
while the application is new, the philosophical
00:18:45.089 --> 00:18:48.750
ancestry of this idea is deeply respected. Where
00:18:48.750 --> 00:18:51.589
does it come from? It traces back to classical
00:18:51.589 --> 00:18:54.329
statistics, specifically to Wald's statistical
00:18:54.329 --> 00:18:57.029
decision theory and the concept of minimax regret
00:18:57.029 --> 00:18:59.970
from the mid -20th century. Oh, minimax regret.
00:19:00.329 --> 00:19:03.410
Right. The foundational idea is that you should
00:19:03.410 --> 00:19:06.900
evaluate a decision rule by its worst case. ambiguity
00:19:06.900 --> 00:19:10.059
-penalized performance rather than blindly assuming
00:19:10.059 --> 00:19:13.079
average case behavior under one perfect utopian
00:19:13.079 --> 00:19:16.359
scenario. So they took an old battle -tested
00:19:16.359 --> 00:19:19.140
philosophical stance on risk management, updated
00:19:19.140 --> 00:19:21.740
it with modern optimization math, and applied
00:19:21.740 --> 00:19:24.160
it to the hyper -modern problem of digital A
00:19:24.160 --> 00:19:27.619
-B testing. Exactly. But let's be clear on the
00:19:27.619 --> 00:19:29.819
limitations here, too. This is not a magic bullet
00:19:29.819 --> 00:19:32.700
for bad experimental design, is it? Definitely
00:19:32.700 --> 00:19:36.470
not. That is a vital caveat. This framework only
00:19:36.470 --> 00:19:39.369
dictates what to do after you have your results.
00:19:39.650 --> 00:19:43.410
Right. Garbage in, garbage out. Always. If you
00:19:43.410 --> 00:19:46.329
are tracking the wrong metrics or testing on
00:19:46.329 --> 00:19:49.450
a heavily biased audience segment, this formula
00:19:49.450 --> 00:19:52.190
won't save you. There is a whole separate discipline
00:19:52.190 --> 00:19:55.529
dedicated to experiment design. So you still
00:19:55.529 --> 00:19:57.250
have to do the hard work up front. Absolutely.
00:19:57.690 --> 00:20:00.329
This paper assumes you already know how to run
00:20:00.329 --> 00:20:03.430
a clean test. It simply fixes the final crucial
00:20:03.430 --> 00:20:06.930
step. the deployment trigger. So what does this
00:20:06.930 --> 00:20:08.450
all mean for the people listening right now?
00:20:09.049 --> 00:20:11.789
Obviously, a solo developer testing button colors
00:20:11.789 --> 00:20:13.690
on a personal blog probably doesn't need to stress
00:20:13.690 --> 00:20:16.750
over the Danske Veraden representation. True,
00:20:16.869 --> 00:20:19.329
the stakes there are pretty low. But anyone making
00:20:19.329 --> 00:20:21.490
high -stakes, capital -intensive choices based
00:20:21.490 --> 00:20:24.329
on data absolutely needs to pay attention. So
00:20:24.329 --> 00:20:27.029
product managers, data science leads? Exactly.
00:20:27.150 --> 00:20:29.849
If you are a product manager, an executive at
00:20:29.849 --> 00:20:32.079
an e -commerce giant, or a marketing director
00:20:32.079 --> 00:20:34.220
allocating millions of dollars based on A -B
00:20:34.220 --> 00:20:36.839
test readouts, you need to recognize that your
00:20:36.839 --> 00:20:40.119
current tools are likely ignoring future volatility.
00:20:40.619 --> 00:20:43.000
You could be shipping those fragile winners.
00:20:43.400 --> 00:20:45.619
You might be systematically deploying fragile
00:20:45.619 --> 00:20:47.619
winners that are slowly bleeding your margins
00:20:47.619 --> 00:20:50.240
because they simply cannot survive in the wild.
00:20:50.579 --> 00:20:53.819
It is a complete paradigm shift. We are looking
00:20:53.819 --> 00:20:55.900
at a clear path to move away from the flawed,
00:20:55.960 --> 00:20:58.819
fragile world where we pretend P less than 0
00:20:58.819 --> 00:21:02.240
.05 equals guaranteed future success. Yeah, that
00:21:02.240 --> 00:21:06.079
world is ending. We can move into a robust, ambiguity
00:21:06.079 --> 00:21:09.259
-averse future using the Donsker -Varadin formula
00:21:09.259 --> 00:21:11.960
to essentially immunize our deployment decisions
00:21:11.960 --> 00:21:14.640
against the chaos of the real world, fundamentally
00:21:14.640 --> 00:21:18.480
minimizing our regret. And looking ahead, this
00:21:18.480 --> 00:21:20.460
raises an important question that goes entirely
00:21:20.460 --> 00:21:22.559
beyond what the authors covered in this paper.
00:21:23.039 --> 00:21:26.700
Oh, lay it on us. Well, we've established that
00:21:26.700 --> 00:21:28.720
the real world drifts away from test environments
00:21:28.720 --> 00:21:31.700
naturally, right? But think about the era we
00:21:31.700 --> 00:21:35.019
are entering right now with AI -generated hyper
00:21:35.019 --> 00:21:38.140
-personalized content. The digital environment
00:21:38.140 --> 00:21:41.160
isn't just drifting organically anymore. It is
00:21:41.160 --> 00:21:44.019
being actively algorithmically mutated by the
00:21:44.019 --> 00:21:48.319
second. Oh, wow. That's a scary thought. If the
00:21:48.319 --> 00:21:50.400
environment shifts so rapidly that the future
00:21:50.400 --> 00:21:53.519
looks entirely alien to the past, can even a
00:21:53.519 --> 00:21:56.859
robust ambiguity -averse formula keep up? Or
00:21:56.859 --> 00:21:58.980
will the speed of AI -driven drift eventually
00:21:58.980 --> 00:22:01.900
outpace our ability to mathematically penalize
00:22:01.900 --> 00:22:05.799
it? That is a staggering thought. If the ground
00:22:05.799 --> 00:22:08.140
is shifting that fast, even packing layers might
00:22:08.140 --> 00:22:11.980
not be enough. We'll have to see. Thank you for
00:22:11.980 --> 00:22:13.980
joining us on this deep dive. Keep questioning
00:22:13.980 --> 00:22:16.079
the metrics you rely on, and we will catch you
00:22:16.079 --> 00:22:16.519
next time.