00:00:00.000 --> 00:00:04.000
This week, containerd disclosed a stack of CRI
00:00:04.000 --> 00:00:06.820
plugin vulnerabilities in the runtime layer,
00:00:07.080 --> 00:00:10.800
a huge number of Kubernetes nodes trust to start
00:00:10.800 --> 00:00:14.759
your containers. Datadog ran a PostgreSQL
00:00:14.759 --> 00:00:17.660
gameday and learned their database could fail over
00:00:17.660 --> 00:00:21.300
just fine. It just couldn't do it safely. AWS
00:00:21.300 --> 00:00:25.039
DevOps Agent and Datadog's MCP Server are both
00:00:25.039 --> 00:00:28.660
now generally available. And the new AWS integration
00:00:28.660 --> 00:00:33.179
means AI incident response just graduated from
00:00:33.179 --> 00:00:37.500
demo to on-call rotation. And EKS will now route
00:00:37.500 --> 00:00:40.060
your Kubernetes control plane's outbound traffic
00:00:40.060 --> 00:00:44.020
through your own VPC, which is great, right up
00:00:44.020 --> 00:00:47.439
until a stale route table quietly kills your
00:00:47.439 --> 00:00:50.280
admission webhooks. Put those together and the
00:00:50.280 --> 00:00:53.259
shape of the episode is pretty clear. The control
00:00:53.259 --> 00:00:56.859
plane keeps getting wider. Runtimes. Databases.
00:00:57.079 --> 00:01:00.899
Incident agents. API-server egress. credentials,
00:01:01.039 --> 00:01:04.219
even the cloud console. One by one, they are
00:01:04.219 --> 00:01:07.140
all sliding into your production blast radius.
00:01:07.439 --> 00:01:10.159
And here's the part that matters. Your users
00:01:10.159 --> 00:01:13.700
don't care which control plane failed. They just
00:01:13.700 --> 00:01:16.799
feel the wait. I'm Brian Teller from Teller's
00:01:16.799 --> 00:01:36.859
Tech, and this is Ship It Weekly. Welcome back
00:01:36.859 --> 00:01:39.959
to Ship It Weekly, the show about the DevOps,
00:01:40.280 --> 00:01:45.000
SRE, cloud, platform, and security stories that
00:01:45.000 --> 00:01:47.980
actually matter when you are the person who has
00:01:47.980 --> 00:01:51.219
to keep the thing running at 3 a.m. If you are
00:01:51.219 --> 00:01:54.540
new here, follow or subscribe wherever you are
00:01:54.540 --> 00:01:57.519
watching or listening. And if you want the weekly
00:01:57.519 --> 00:02:01.219
story list and source links, check out OnCallBrief.com
00:02:01.219 --> 00:02:05.069
For past episodes, full show notes, and
00:02:05.069 --> 00:02:08.069
more from the show, head over to ShipItWeekly.fm
00:02:08.069 --> 00:02:12.469
We open with the containerd CRI plugin vulnerabilities,
00:02:13.009 --> 00:02:16.650
because your node runtime is the trust boundary
00:02:16.650 --> 00:02:20.530
underneath the trust boundary. Then, Datadog's
00:02:20.530 --> 00:02:24.430
PostgreSQL HA gameday, where the scary discovery
00:02:24.430 --> 00:02:28.530
wasn't that failover was hard, it was that failover
00:02:28.530 --> 00:02:32.800
was unsafe. After that, AWS DevOps Agent and
00:02:32.800 --> 00:02:36.639
Datadog MCP Server going GA. And what it means
00:02:36.639 --> 00:02:40.080
when an AI agent gets a seat near your control
00:02:40.080 --> 00:02:43.680
plane. Then, EKS customer-routed control-plane
00:02:43.680 --> 00:02:47.599
egress. Because your API server is now part of
00:02:47.599 --> 00:02:50.180
your network perimeter, whether you plan for
00:02:50.180 --> 00:02:53.680
it or not. In the lightning round, GitHub Credential
00:02:53.680 --> 00:02:57.840
Revocation. AWS Console Private Access. Vercel
00:02:57.840 --> 00:03:01.780
Connect, and S3 annotations. And we close with
00:03:01.780 --> 00:03:05.620
Marc Brooker on waiting, on why your customers
00:03:05.620 --> 00:03:08.860
live in the tail of your latency distribution,
00:03:09.240 --> 00:03:12.360
even when your dashboards swear everything's
00:03:12.360 --> 00:03:20.039
fine. Let's get into it. First up, containerd
00:03:20.039 --> 00:03:23.780
has a batch of CRI plugin vulnerabilities. And
00:03:23.780 --> 00:03:26.949
if you run Kubernetes, this one's yours. AWS
00:03:26.949 --> 00:03:30.030
published a security bulletin spanning
00:03:30.030 --> 00:03:35.030
containerd branches 1.7 through 2.3. And the list is
00:03:35.030 --> 00:03:38.490
not a fun read. Image cache poisoning through
00:03:38.490 --> 00:03:41.770
checkpoint image references. Host command execution.
00:03:42.490 --> 00:03:45.949
through unsanitized image labels, CDI annotation
00:03:45.949 --> 00:03:49.710
handling that can inject devices and host mounts,
00:03:50.030 --> 00:03:53.050
host file reads through symlinked container
00:03:53.050 --> 00:03:57.150
log paths during checkpoint restore, and a denial
00:03:57.150 --> 00:04:01.110
of service from crafted images that exhaust memory.
00:04:01.370 --> 00:04:05.319
So not exactly a relaxing Patch Tuesday. Here's
00:04:05.319 --> 00:04:08.840
why it matters. containerd sits underneath an
00:04:08.840 --> 00:04:12.099
enormous number of clusters, and we spend almost
00:04:12.099 --> 00:04:15.219
all of our security attention on the layers above
00:04:15.219 --> 00:04:18.800
it. Pod specs, admission control, image scanning,
00:04:19.040 --> 00:04:23.220
RBAC, network policy, runtime classes, all the
00:04:23.220 --> 00:04:26.180
familiar Kubernetes machinery. But eventually,
00:04:26.519 --> 00:04:29.680
something has to actually pull the image, unpack
00:04:29.680 --> 00:04:34.100
it, restore it, wire up devices. handle the logs,
00:04:34.180 --> 00:04:37.699
and start the container. That layer is a trust
00:04:37.699 --> 00:04:41.040
boundary too. And in some ways, it's the more
00:04:41.040 --> 00:04:44.019
dangerous one. Because by the time a workload
00:04:44.019 --> 00:04:47.399
reaches the runtime, the rest of the system has
00:04:47.399 --> 00:04:50.800
already decided this thing is allowed to exist.
00:04:51.160 --> 00:04:54.240
That's why the boring fields turn out to matter.
00:04:54.459 --> 00:04:57.939
Labels, annotations, checkpoint and restore paths,
00:04:58.319 --> 00:05:02.459
CDI, log paths, every field. that feels like
00:05:02.459 --> 00:05:05.819
plumbing can become an input to privileged behavior
00:05:05.819 --> 00:05:09.920
on the node. A malicious image isn't just application
00:05:09.920 --> 00:05:14.079
code. It's metadata, build time weirdness, and
00:05:14.079 --> 00:05:17.240
a set of assumptions the runtime makes about
00:05:17.240 --> 00:05:20.800
what it can trust. The takeaway is direct. Patch
00:05:20.800 --> 00:05:23.980
containerd. Check your managed node groups,
00:05:24.240 --> 00:05:28.519
your self-managed nodes, your AMIs, your Bottlerocket
00:05:28.519 --> 00:05:31.899
versions, your distro packages, anything
00:05:31.899 --> 00:05:35.319
that controls the runtime. If you lean on checkpoint
00:05:35.319 --> 00:05:39.899
restore, CDI devices, or GPU workloads, look
00:05:39.899 --> 00:05:43.079
harder. And if you don't use any of that, don't
00:05:43.079 --> 00:05:46.699
relax. At least one of these issues doesn't need
00:05:46.699 --> 00:05:49.620
checkpoint and restore turned on at all. Your
00:05:49.620 --> 00:05:53.379
node runtime is the trust boundary under the
00:05:53.379 --> 00:05:56.139
trust boundary. Stop treating it like invisible
00:05:56.139 --> 00:06:03.699
plumbing. Second story. Datadog published a genuinely
00:06:03.699 --> 00:06:07.160
good engineering write-up on running high availability
00:06:07.160 --> 00:06:11.500
PostgreSQL on Kubernetes. And it's one of those
00:06:11.500 --> 00:06:14.839
pieces that sounds boring until the real problem
00:06:14.839 --> 00:06:17.980
comes into focus. The problem wasn't that the
00:06:17.980 --> 00:06:20.779
database couldn't fail over. It was that it couldn't
00:06:20.779 --> 00:06:24.189
fail over safely. During a gameday, Datadog
00:06:24.189 --> 00:06:27.269
simulated a zonal failure. That added network
00:06:27.269 --> 00:06:30.930
latency, replication lag grew, and when the cluster
00:06:30.930 --> 00:06:34.209
needed a new primary, Patroni couldn't safely
00:06:34.209 --> 00:06:37.410
promote a standby without risking data loss.
00:06:37.730 --> 00:06:40.930
So the system got stuck in the worst possible
00:06:40.930 --> 00:06:44.889
spot. The old primary was unhealthy. The standbys
00:06:44.889 --> 00:06:47.550
weren't safe to promote, and the only correct
00:06:47.550 --> 00:06:50.199
move was to wait. That's the kind of failure
00:06:50.199 --> 00:06:53.680
mode that ages every SRE in the room about three
00:06:53.680 --> 00:06:56.680
years. Because on paper, you have everything.
00:06:56.939 --> 00:07:00.519
Multiple nodes, standbys, Kubernetes, automation,
00:07:01.040 --> 00:07:04.079
failover machinery. And then the actual failure
00:07:04.079 --> 00:07:07.699
arrives and the system says, yes, but not safely.
00:07:07.920 --> 00:07:10.839
Which, honestly, is the right answer. Promoting
00:07:10.839 --> 00:07:14.120
a stale standby might hand you a writable primary
00:07:14.120 --> 00:07:17.569
faster. But if it costs you data loss, split
00:07:17.569 --> 00:07:21.110
brain, or a broken consistency guarantee, you
00:07:21.110 --> 00:07:23.829
haven't fixed the outage. You've traded it for
00:07:23.829 --> 00:07:26.470
a corruption event. That's not an improvement.
00:07:26.769 --> 00:07:29.910
It's just a different postmortem. The real lesson
00:07:29.910 --> 00:07:33.009
is that HA isn't only about whether the service
00:07:33.009 --> 00:07:35.829
comes back. It's about whether the recovery path
00:07:35.829 --> 00:07:39.329
itself is safe. Can you fail over without losing
00:07:39.329 --> 00:07:42.509
writes? Can you prove which standby is safe to
00:07:42.509 --> 00:07:45.060
promote? Can your automation tell the difference
00:07:45.060 --> 00:07:47.959
between available and correct? And does your
00:07:47.959 --> 00:07:51.139
whole team agree on which one it should prefer
00:07:51.139 --> 00:07:54.899
before the incident call is on fire? Datadog's
00:07:54.899 --> 00:07:57.360
answer was to move toward synchronous replication
00:07:57.360 --> 00:08:01.300
and stronger Patroni guardrails. So a promoted
00:08:01.300 --> 00:08:05.019
standby is guaranteed to have the writes it needs.
00:08:05.300 --> 00:08:07.899
And that's the part that's worth copying. They
00:08:07.899 --> 00:08:11.040
didn't just ask how to recover faster. They asked
00:08:11.040 --> 00:08:14.480
how to recover safely. So test your database
00:08:14.480 --> 00:08:18.079
HA against real constraints, not the easy ones.
00:08:18.360 --> 00:08:22.019
Ask what happens under replication lag. Ask what
00:08:22.019 --> 00:08:25.259
happens during a zone failure. Ask what happens
00:08:25.259 --> 00:08:28.579
when the network is slow instead of cleanly dead.
00:08:28.860 --> 00:08:32.039
Ask what happens when every standby is behind.
00:08:32.399 --> 00:08:35.419
And ask whether your automation prefers safety
00:08:35.419 --> 00:08:38.600
or availability. And whether everyone actually
00:08:38.600 --> 00:08:41.779
agrees with that choice. Because failover is
00:08:41.779 --> 00:08:44.860
useless, if the only safe option is waiting.
00:08:45.039 --> 00:08:52.919
But unsafe failover can be a lot worse. Third
00:08:52.919 --> 00:08:56.679
story, AWS DevOps Agent is now generally available
00:08:56.679 --> 00:09:00.820
and Datadog's MCP Server is GA as a standard
00:09:00.820 --> 00:09:04.379
way for AI agents to reach Datadog monitoring
00:09:04.379 --> 00:09:07.379
data. This is one of those announcements. where
00:09:07.379 --> 00:09:10.019
the slide says autonomous incident resolution
00:09:10.019 --> 00:09:13.919
and the operator says, cool, but what exactly
00:09:13.919 --> 00:09:17.720
is it allowed to touch? The idea is solid. AWS
00:09:17.720 --> 00:09:21.820
DevOps Agent can work through Datadog MCP Server
00:09:21.820 --> 00:09:26.179
to investigate an incident across logs, metrics,
00:09:26.419 --> 00:09:30.340
traces, deployment events, and AWS infrastructure
00:09:30.340 --> 00:09:33.980
context. Instead of one engineer bouncing between
00:09:33.980 --> 00:09:38.679
CloudWatch, Datadog, deploy history, traces, dashboards,
00:09:38.679 --> 00:09:42.700
and Slack, the agent correlates the signals and
00:09:42.700 --> 00:09:45.779
helps push the incident forward and nobody wants
00:09:45.779 --> 00:09:48.620
to spend the first 30 minutes of an outage doing
00:09:48.620 --> 00:09:51.879
browser-tab archaeology if an agent can gather
00:09:51.879 --> 00:09:55.460
context, summarize what changed, flag a suspicious
00:09:55.460 --> 00:09:59.259
deploy and propose likely causes that's real
00:09:59.259 --> 00:10:02.820
time saved but this is also the moment AI incident
00:10:02.820 --> 00:10:06.549
response stops being a chatbot and becomes a
00:10:06.549 --> 00:10:10.289
production workflow. It's an agent reading operational
00:10:10.289 --> 00:10:13.549
telemetry, interpreting signals, recommending
00:10:13.549 --> 00:10:17.909
fixes, and potentially wired into Slack, PagerDuty,
00:10:18.049 --> 00:10:21.669
ServiceNow, your code, your deploys, and your
00:10:21.669 --> 00:10:24.549
runbooks. That puts it right next to the control
00:10:24.549 --> 00:10:27.870
plane. And once something sits next to the control
00:10:27.870 --> 00:10:31.029
plane, the question stops being, is it smart?
00:10:31.210 --> 00:10:34.840
And becomes, what authority does it have? Can
00:10:34.840 --> 00:10:38.240
it only read? Can it write? Can it open tickets?
00:10:38.539 --> 00:10:42.200
Trigger automation? Roll back a deploy? Restart
00:10:42.200 --> 00:10:46.679
a service? Change config? Page a human at 4 a.m.?
00:10:46.940 --> 00:10:51.120
Can it make things worse quickly and very confidently?
00:10:51.460 --> 00:10:54.919
That last one is the whole game. Incident response
00:10:54.919 --> 00:10:58.840
isn't about speed. It's about safe speed. So
00:10:58.840 --> 00:11:02.659
treat AI incident tooling like any other production
00:11:02.659 --> 00:11:05.759
automation. Give it the least privilege that
00:11:05.759 --> 00:11:09.059
still leaves it useful. Log what it sees and
00:11:09.059 --> 00:11:11.860
what it does. Make the human approval boundary
00:11:11.860 --> 00:11:15.879
impossible to miss. And draw a hard line between
00:11:15.879 --> 00:11:19.299
what it can recommend and what it can execute.
00:11:19.600 --> 00:11:23.279
Have rollback rules. Know what happens when it's
00:11:23.279 --> 00:11:26.419
wrong. And don't grade it only on time to answer.
00:11:26.580 --> 00:11:29.480
Grade it on whether the answer was safe, auditable,
00:11:29.620 --> 00:11:33.100
and actually useful under pressure. AI incident
00:11:33.100 --> 00:11:36.679
response is moving from demo to production. That's
00:11:36.679 --> 00:11:44.080
exciting. Production just needs guardrails. Fourth
00:11:44.080 --> 00:11:48.120
story. Amazon EKS now supports customer-routed
00:11:48.120 --> 00:11:51.940
control-plane egress. That's a very AWS phrase.
00:11:52.240 --> 00:11:55.000
So here's the human version. The Kubernetes API
00:11:55.000 --> 00:11:58.940
server sometimes needs to call outward to admission
00:11:58.940 --> 00:12:02.940
webhooks, OIDC providers, aggregated API servers,
00:12:03.159 --> 00:12:05.940
other endpoints that you control. Historically,
00:12:06.019 --> 00:12:09.639
that outbound traffic took AWS managed egress
00:12:09.639 --> 00:12:12.679
paths. Now you can route it through your own
00:12:12.679 --> 00:12:16.299
VPC, which hands platform teams control over
00:12:16.299 --> 00:12:19.700
routing, inspection, firewalls, NAT, private
00:12:19.700 --> 00:12:22.419
connectivity, and compliance boundaries. For
00:12:22.419 --> 00:12:25.519
regulated environments, that's a real win. It
00:12:25.519 --> 00:12:28.320
also makes the control plane feel a lot more
00:12:28.320 --> 00:12:30.899
like part of your network. which of course it
00:12:30.899 --> 00:12:34.000
always was. The difference is that now you own
00:12:34.000 --> 00:12:37.559
the outbound path and AWS is blunt about what
00:12:37.559 --> 00:12:40.379
that ownership means. In customer routed mode,
00:12:40.639 --> 00:12:43.639
you are responsible for making sure the control
00:12:43.639 --> 00:12:46.700
plane can reach the endpoints it needs. Wrong
00:12:46.700 --> 00:12:50.100
route table, too-tight security group, a NACL
00:12:50.100 --> 00:12:52.940
that blocks the wrong thing, a broken firewall
00:12:52.940 --> 00:12:55.980
hop, and control plane operations start failing.
00:12:56.120 --> 00:13:00.139
That includes admission webhook calls, and OIDC
00:13:00.139 --> 00:13:03.700
authentication. So yes, great feature. But it
00:13:03.700 --> 00:13:06.580
isn't a checkbox. It's a failure mode change.
00:13:06.919 --> 00:13:10.480
If your API server can't reach an admission webhook,
00:13:10.539 --> 00:13:14.460
do pod creates fail? Do deploys hang? Does authentication
00:13:14.460 --> 00:13:18.039
break? Does your incident response now depend
00:13:18.039 --> 00:13:21.779
on a firewall path some other team owns? And
00:13:21.779 --> 00:13:25.019
do you have a metric? a test, and a name on the
00:13:25.019 --> 00:13:28.120
pager for when it breaks? This is a feature you
00:13:28.120 --> 00:13:31.100
bring to a design review. Not because it's risky,
00:13:31.259 --> 00:13:34.580
but because it's powerful. Map the traffic. Map
00:13:34.580 --> 00:13:38.440
the dependencies. Test the webhooks. Test OIDC.
00:13:38.580 --> 00:13:41.679
Test the failure modes. Make the routing visible.
00:13:41.899 --> 00:13:44.679
And write the runbook before the control plane
00:13:44.679 --> 00:13:47.820
starts failing in creative ways. The Kubernetes
00:13:47.820 --> 00:13:50.659
control plane is becoming part of your network
00:13:50.659 --> 00:14:00.740
perimeter. Treat it like one. Quick lightning
00:14:00.740 --> 00:14:03.980
round. First, GitHub added self-service credential
00:14:03.980 --> 00:14:07.059
revocation for incident response. Enterprise
00:14:07.059 --> 00:14:11.159
owners now get a break-glass capability to revoke
00:14:11.159 --> 00:14:14.179
a compromised user's credentials in one move.
00:14:14.399 --> 00:14:17.200
This matters because credential cleanup should
00:14:17.200 --> 00:14:20.200
never be a scavenger hunt. You do not want to
00:14:20.200 --> 00:14:22.639
be hand-hunting through SSO authorizations,
00:14:22.899 --> 00:14:27.149
personal access tokens, SSH keys, and OAuth grants
00:14:27.149 --> 00:14:30.490
while everyone argues in Slack. Revocation is
00:14:30.490 --> 00:14:33.169
incident response infrastructure. Know who can
00:14:33.169 --> 00:14:35.889
trigger it, know what it kills, know what it
00:14:35.889 --> 00:14:39.169
logs, and put it in the compromised-account runbook.
00:14:39.370 --> 00:14:42.429
Second, AWS Management Console private access
00:14:42.429 --> 00:14:45.529
now works without internet connectivity. Console
00:14:45.529 --> 00:14:48.490
traffic for supported services can flow over
00:14:48.490 --> 00:14:51.549
VPC endpoints instead of the public internet.
00:14:51.899 --> 00:14:54.460
It's a strong story for regulated environments.
00:14:54.899 --> 00:14:57.399
Even the console is getting pulled behind private
00:14:57.399 --> 00:15:00.259
network boundaries. The lesson? Console access
00:15:00.259 --> 00:15:03.179
is part of your control plane too. And private
00:15:03.179 --> 00:15:06.779
link, endpoint policies, and known-account restrictions
00:15:06.779 --> 00:15:10.360
are becoming cloud operations, not just app networking.
00:15:10.879 --> 00:15:13.960
Third, Vercel shipped Vercel Connect. And the
00:15:13.960 --> 00:15:17.179
idea worth catching is runtime credential exchange.
00:15:17.659 --> 00:15:20.419
Instead of stashing a long-lived provider token
00:15:20.419 --> 00:15:23.799
for an agent, The app proves its identity and
00:15:23.799 --> 00:15:26.500
gets a short-lived task-scoped credential.
00:15:26.799 --> 00:15:28.980
That's the pattern that we've been tracking for
00:15:28.980 --> 00:15:31.720
weeks. Agent credentials moving from store this
00:15:31.720 --> 00:15:35.679
token forever to prove who you are and get scoped
00:15:35.679 --> 00:15:38.600
access when you need it. Short-lived credentials
00:15:38.600 --> 00:15:41.720
don't solve every agent security problem, but
00:15:41.720 --> 00:15:44.360
they beat long-lived secrets sitting around
00:15:44.360 --> 00:15:47.879
waiting to become next quarter's incident. Fourth,
00:15:48.039 --> 00:15:52.009
Amazon S3 annotations are here. mutable, queryable
00:15:52.009 --> 00:15:56.149
context attached directly to S3 objects. Sounds
00:15:56.149 --> 00:15:59.909
dull, but object metadata has driven a lot of
00:15:59.909 --> 00:16:02.850
awkward platform design over the years. Side
00:16:02.850 --> 00:16:06.470
tables, DynamoDB metadata stores, Lambda sync
00:16:06.470 --> 00:16:10.490
jobs, custom catalogs, and constant drift between
00:16:10.490 --> 00:16:14.029
the object and whatever's describing it. If annotations
00:16:14.029 --> 00:16:16.720
shrink that glue layer, That's worth watching.
00:16:16.820 --> 00:16:20.259
Object metadata is quietly becoming a first-class
00:16:20.259 --> 00:16:24.000
platform layer, especially for data, AI, search,
00:16:24.200 --> 00:16:27.799
and agent workflows that need to know what an
00:16:27.799 --> 00:16:38.779
object is, not just where it lives. The human
00:16:38.779 --> 00:16:42.340
closer this week comes from a Marc Brooker post
00:16:42.340 --> 00:16:47.370
about waiting, latency, MTTR, and why averages
00:16:47.370 --> 00:16:50.990
can lie. The point is that your users don't experience
00:16:50.990 --> 00:16:54.070
your averages the way your dashboards report
00:16:54.070 --> 00:16:57.570
them. You measure mean latency, mean time to
00:16:57.570 --> 00:17:00.529
recovery, average outage duration. But people
00:17:00.529 --> 00:17:03.590
are far more likely to land in the long waits
00:17:03.590 --> 00:17:07.289
simply because long waits take up more of the
00:17:07.289 --> 00:17:10.970
time. That's the inspection paradox. A 10-minute
00:17:10.970 --> 00:17:15.019
outage catches a few users. A 10-hour outage
00:17:15.019 --> 00:17:18.000
catches a lot of them. Your incident tracker
00:17:18.000 --> 00:17:22.000
counts both as one outage. Your dashboard says
00:17:22.000 --> 00:17:26.279
MTTR looks fine. Your users say they spent all
00:17:26.279 --> 00:17:29.140
morning waiting. Both are true. And that's the
00:17:29.140 --> 00:17:31.460
whole episode, really. When the system breaks,
00:17:31.740 --> 00:17:34.460
nobody experiences your architecture diagram.
00:17:34.759 --> 00:17:37.859
They experience waiting. Waiting for a request.
00:17:38.160 --> 00:17:40.890
Waiting for recovery. Waiting for a credential
00:17:40.890 --> 00:17:44.150
to get revoked. Waiting for a deploy to stop
00:17:44.150 --> 00:17:46.950
failing. Waiting for the control plane to come
00:17:46.950 --> 00:17:50.190
back. Waiting for someone to find the right context.
00:17:50.549 --> 00:17:53.930
So here's the takeaway. Don't only measure the
00:17:53.930 --> 00:17:56.990
system from the server side. Measure it from
00:17:56.990 --> 00:18:00.650
the waiting side. Because your users don't live
00:18:00.650 --> 00:18:04.309
in your average. They live in the tail. And the
00:18:04.309 --> 00:18:07.789
tail is usually where the real reliability story
00:18:07.789 --> 00:18:10.799
is hiding. That's it for this week of Ship It
00:18:10.799 --> 00:18:13.519
Weekly. We covered containerd runtime risk,
00:18:13.799 --> 00:18:17.000
Postgres failover safety, AI incident response,
00:18:17.460 --> 00:18:20.839
EKS control-plane egress, and why your users
00:18:20.839 --> 00:18:23.819
feel the wait more than your dashboards show.
00:18:24.039 --> 00:18:27.019
If this episode was useful, follow or subscribe
00:18:27.019 --> 00:18:29.960
wherever you are watching or listening. If you're
00:18:29.960 --> 00:18:32.940
on YouTube, hit subscribe. If you're in a podcast
00:18:32.940 --> 00:18:36.250
app, follow the show there. And if you know someone
00:18:36.250 --> 00:18:39.109
wrestling with Kubernetes runtime security, database
00:18:39.109 --> 00:18:42.750
failover, AI incident response, or platform control
00:18:42.750 --> 00:18:45.930
planes, send them this one. It genuinely helps
00:18:45.930 --> 00:18:48.549
the show grow, and it helps me keep making this
00:18:48.549 --> 00:18:51.549
for people who actually live with these systems.
00:18:51.849 --> 00:18:54.430
You can find the weekly brief at OnCallBrief.com
00:18:54.430 --> 00:18:57.549
and the full show notes, links, and past
00:18:57.549 --> 00:19:01.230
episodes at ShipItWeekly.fm. I'm Brian Teller
00:19:01.230 --> 00:19:03.670
from Teller's Tech. Thanks for listening. And
00:19:03.670 --> 00:19:06.140
remember, your dashboards measure the average.
00:19:06.339 --> 00:19:08.220
Your users feel the wait.