00:00:00.339 --> 00:00:09.820
Thank you. Hey, I'm Brian Teller. I work in DevOps
00:00:09.820 --> 00:00:13.179
and SRE and I run Teller's Tech. Ship It Weekly
00:00:13.179 --> 00:00:15.900
is where I filter the noise and pull out what
00:00:15.900 --> 00:00:18.379
actually matters when you're the one running
00:00:18.379 --> 00:00:22.260
infrastructure and owning reliability. If something's
00:00:22.260 --> 00:00:25.079
hype, I'll call it hype. If it changes how you
00:00:25.079 --> 00:00:27.800
operate, I'll break it down in plain English.
00:00:28.059 --> 00:00:31.839
Most weeks, this is a quick news recap. In between
00:00:31.839 --> 00:00:34.579
those, I drop interview episodes with folks across
00:00:34.579 --> 00:00:37.880
the DevOps world. Happy holidays. Merry Christmas,
00:00:38.079 --> 00:00:40.899
all that good stuff. It's the day after Christmas,
00:00:41.020 --> 00:00:43.280
so if you're listening while hiding from your
00:00:43.280 --> 00:00:46.280
family, checking one thing real quick, I respect
00:00:46.280 --> 00:00:49.000
it. Quick piece of housekeeping, the new site
00:00:49.000 --> 00:00:53.280
is live at shipitweekly .fm. That's where I'm
00:00:53.280 --> 00:00:55.939
putting links and show notes. Also, I'm looking
00:00:55.939 --> 00:00:58.179
for interview guests. If you're building real
00:00:58.179 --> 00:01:01.000
infra or platform stuff and you want to come
00:01:01.000 --> 00:01:04.140
on for a chill conversation episode, hit the
00:01:04.140 --> 00:01:07.659
email on shipitweekly .fm. If you've got war
00:01:07.659 --> 00:01:10.439
stories, even better. And one quick ask before
00:01:10.439 --> 00:01:13.219
we start. If the show has been useful, hit follow
00:01:13.219 --> 00:01:15.900
or subscribe wherever you are listening. And
00:01:15.900 --> 00:01:18.459
if you've got 10 seconds, a rating or review
00:01:18.459 --> 00:01:21.060
really helps way more than it should. All right,
00:01:21.180 --> 00:01:24.500
three main stories for today. First, Cloudflare
00:01:24.500 --> 00:01:27.280
wrote up how they built an internal maintenance
00:01:27.280 --> 00:01:31.060
scheduler on workers. This is real platform engineering.
00:01:31.439 --> 00:01:34.680
Memory limits. data modeling, query optimization,
00:01:35.379 --> 00:01:39.120
caching, Parquet for historical analysis. Second,
00:01:39.340 --> 00:01:43.140
AWS databases are now available directly in Vercel
00:01:43.140 --> 00:01:46.200
marketplace. It's a quiet shift, but it's a big
00:01:46.200 --> 00:01:50.040
one. Devs can click button real AWS databases
00:01:50.040 --> 00:01:53.079
from inside Vercel, but you still have to own
00:01:53.079 --> 00:01:55.900
governance, billing, and the blast radius. Third,
00:01:56.120 --> 00:01:59.060
there's an open source project from AWS called
00:01:59.060 --> 00:02:03.629
Team. Temporary Elevated Access Management, and
00:02:03.629 --> 00:02:06.530
it's built around IAM Identity Center. It's approval
00:02:06.530 --> 00:02:09.710
-based, time -bound access. This is one of those
00:02:09.710 --> 00:02:12.370
everybody wants it, few implement it cleanly
00:02:12.370 --> 00:02:14.990
problems. Then we'll do a lightning round, and
00:02:14.990 --> 00:02:17.930
we'll close with Mark Brooker's What Now? Handling
00:02:17.930 --> 00:02:21.449
Errors in Large Systems. Let's get into it. Cloudflare
00:02:21.449 --> 00:02:24.530
has a really good post on how they built an internal
00:02:24.530 --> 00:02:27.349
maintenance brain on workers. The core problem
00:02:27.349 --> 00:02:29.939
is kind of obvious once you hear it. when you
00:02:29.939 --> 00:02:33.400
run infra at their scale you cannot rely on humans
00:02:33.400 --> 00:02:36.340
to remember every dependency and every weird
00:02:36.340 --> 00:02:39.780
routing rule and every if these two things go
00:02:39.780 --> 00:02:42.979
down at the same time a customer special setup
00:02:42.979 --> 00:02:46.560
gets wrecked scenario so they built a centralized
00:02:46.560 --> 00:02:49.879
scheduler that treats maintenance like a set
00:02:49.879 --> 00:02:53.560
of constraints like we must always have at least
00:02:53.560 --> 00:02:56.979
one of these routers active or this customer
00:02:56.979 --> 00:03:00.419
pins traffic through these data centers, so don't
00:03:00.419 --> 00:03:03.560
take all of them out at once. The fun part is
00:03:03.560 --> 00:03:06.039
how they got it working within worker's limits.
00:03:06.379 --> 00:03:09.000
Their first naive approach was basically load
00:03:09.000 --> 00:03:12.419
everything into one worker. All the relationships,
00:03:12.939 --> 00:03:16.379
all the product config, all the metrics, then
00:03:16.379 --> 00:03:19.639
compute constraints. And even in proof of concept,
00:03:19.860 --> 00:03:22.680
they hit out of memory errors. So they took a
00:03:22.680 --> 00:03:25.979
step back and said, okay. Workers have limits.
00:03:26.240 --> 00:03:29.180
We can't treat this like a giant in -memory analytics
00:03:29.180 --> 00:03:32.520
job. We need to only load the data that matters
00:03:32.520 --> 00:03:35.539
for the specific maintenance request. If you
00:03:35.539 --> 00:03:37.819
get a maintenance request for your router in
00:03:37.819 --> 00:03:40.740
Frankfurt, you probably do not need to load Australia.
00:03:41.099 --> 00:03:43.840
You need the dependency neighborhood around that
00:03:43.840 --> 00:03:46.810
router. that pushed them into graph modeling
00:03:46.810 --> 00:03:50.330
they describe constraints as objects and associations
00:03:50.330 --> 00:03:54.270
basically vertices and edges routers are objects
00:03:54.270 --> 00:03:57.870
pulls are objects and the dependencies are associations
00:03:57.870 --> 00:04:01.150
and then they built a typed association interface
00:04:01.150 --> 00:04:04.250
so the constraint logic stays simple but the
00:04:04.250 --> 00:04:07.469
backing implementation can get smarter over time
00:04:07.469 --> 00:04:10.469
then they flip their data fetching style instead
00:04:10.469 --> 00:04:13.330
of pulling down huge responses and filtering
00:04:13.330 --> 00:04:16.069
locally They started doing targeted requests
00:04:16.069 --> 00:04:19.089
through the graph interface. They claim response
00:04:19.089 --> 00:04:23.189
sizes dropped by 100 times in one spot. That's
00:04:23.189 --> 00:04:26.670
huge. Also, it's the exact kind of win you get
00:04:26.670 --> 00:04:30.370
when you stop shipping your entire dataset into
00:04:30.370 --> 00:04:32.829
your app layer just to throw most of it away.
00:04:33.089 --> 00:04:36.279
Of course, that created the next problem. sub
00:04:36.279 --> 00:04:39.279
-request limits. They traded a few massive requests
00:04:39.279 --> 00:04:43.579
for a ton of tiny requests. And then they started
00:04:43.579 --> 00:04:46.439
breaching sub -request limits. So they built
00:04:46.439 --> 00:04:50.600
a fetch pipeline with request deduping, a small
00:04:50.600 --> 00:04:54.800
LRU cache, edge caching via caches .default,
00:04:54.860 --> 00:04:59.019
and sane retry and back off. After tuning, they
00:04:59.019 --> 00:05:02.899
were seeing about a 99 % cache hit rate on those
00:05:02.899 --> 00:05:05.819
fetches. That's wild. And it's It's also how
00:05:05.819 --> 00:05:08.259
you make something like this survive at scale
00:05:08.259 --> 00:05:11.959
without tuning your internal APIs into a creator.
00:05:12.160 --> 00:05:14.600
Then there's a super relatable metric story.
00:05:14.980 --> 00:05:18.720
They use Thanos for Prometheus queries, and they
00:05:18.720 --> 00:05:21.519
call out what a lot of teams do by accident.
00:05:21.800 --> 00:05:25.500
Ask for everything. Get megabytes back, parse
00:05:25.500 --> 00:05:29.160
JSON into single -threaded runtime, then filter
00:05:29.160 --> 00:05:32.220
most of it out. That's basically self -inflicted
00:05:32.220 --> 00:05:35.839
pain. instead they use the graph to find the
00:05:35.839 --> 00:05:39.040
specific relationships first then issue much
00:05:39.040 --> 00:05:41.879
more targeted thanos queries they say average
00:05:41.879 --> 00:05:44.939
response size went from multiple megabytes to
00:05:44.939 --> 00:05:48.079
about one kilobyte in one case so the theme of
00:05:48.079 --> 00:05:51.079
this story is stop dragging huge blobs of data
00:05:51.079 --> 00:05:54.660
into your application just so you can toss 99
00:05:54.660 --> 00:05:57.699
of it and then they bring it home with historical
00:05:57.699 --> 00:06:01.459
analysis real time is one thing historical is
00:06:01.459 --> 00:06:04.500
a different because now you're scanning months
00:06:04.500 --> 00:06:07.620
of data to see if your logic is actually safe
00:06:07.620 --> 00:06:10.399
and accurate. They talk about how Prometheus
00:06:10.399 --> 00:06:14.019
TSDB blocks are not really designed for object
00:06:14.019 --> 00:06:16.899
storage access patterns and how that turns into
00:06:16.899 --> 00:06:19.920
a lot of random reads. So they adopt a parquet
00:06:19.920 --> 00:06:23.199
conversation layer for historical data. Columnar
00:06:23.199 --> 00:06:26.069
format. better stats, and you can fetch what
00:06:26.069 --> 00:06:28.670
you need without slamming object storage with
00:06:28.670 --> 00:06:31.750
random IO. Takeaways you can steal even if you're
00:06:31.750 --> 00:06:34.689
not Cloudflare. If you're building platform brains
00:06:34.689 --> 00:06:38.089
like schedulers, deploy orchestrators, policy
00:06:38.089 --> 00:06:42.149
evaluators, you will hit limits. Memory, CPU,
00:06:42.750 --> 00:06:47.329
API quotas, request fanout. You win by changing
00:06:47.329 --> 00:06:50.470
the shape of the problem, not brute forcing it.
00:06:50.550 --> 00:06:53.509
Graph interfaces are a cheat code for dependency
00:06:53.509 --> 00:06:57.110
-heavy domains. And targeted queries plus caching
00:06:57.110 --> 00:07:00.649
plus backoff is still undefeated. All right,
00:07:00.790 --> 00:07:03.949
let's move from Cloudflare did adult engineering
00:07:03.949 --> 00:07:07.750
to developers can click -button databases. AWS
00:07:07.750 --> 00:07:10.889
announced that AWS databases are now available
00:07:10.889 --> 00:07:14.110
on the Vercel marketplace. The headline is simple.
00:07:14.560 --> 00:07:17.319
From Vercel, you can provision and connect to
00:07:17.319 --> 00:07:22.240
Aurora Postgres, Aurora DSQL, and DynamoDB in
00:07:22.240 --> 00:07:25.160
seconds. And here's the part platform folks should
00:07:25.160 --> 00:07:28.019
not ignore. The onboarding path is basically
00:07:28.019 --> 00:07:32.199
create a new AWS account from Vercel with some
00:07:32.199 --> 00:07:35.480
starter credits. So the dev experience is you're
00:07:35.480 --> 00:07:38.439
already in Vercel. You click a thing, and now
00:07:38.439 --> 00:07:40.980
you have a database, and the app is wired up.
00:07:41.259 --> 00:07:44.480
That is great for Velocity. It is also the kind
00:07:44.480 --> 00:07:47.439
of thing that bypasses governance if you don't
00:07:47.439 --> 00:07:50.180
get in front of it. Even the AWS side is legit.
00:07:50.420 --> 00:07:53.600
You still have to deal with who owns that AWS
00:07:53.600 --> 00:07:57.519
account long term. Is it inside your AWS organization
00:07:57.519 --> 00:08:00.699
or is it an orphan account that exists because
00:08:00.699 --> 00:08:05.120
Vercel made it easy? How do SCPs apply? How do
00:08:05.120 --> 00:08:07.959
guardrails apply? How do you do tagging and cost
00:08:07.959 --> 00:08:11.279
allocation so finance doesn't show up later asking
00:08:11.279 --> 00:08:14.420
why there are mystery accounts? What does networking
00:08:14.420 --> 00:08:17.720
look like? Are public endpoints acceptable? Do
00:08:17.720 --> 00:08:20.660
you need private connectivity? Do you have a
00:08:20.660 --> 00:08:24.480
VPC strategy that fits Vercel -first teams? What's
00:08:24.480 --> 00:08:28.699
your audit baseline? Cloud trail? Config? Detective
00:08:28.699 --> 00:08:32.240
controls? All the boring stuff. Also, Region
00:08:32.240 --> 00:08:35.980
selection matters for data, residency, and latency.
00:08:36.340 --> 00:08:39.759
It's not just pick whatever. Vercel also hinted
00:08:39.759 --> 00:08:42.240
this is evolving. They're talking about coming
00:08:42.240 --> 00:08:45.519
soon support for provisioning into an existing
00:08:45.519 --> 00:08:48.879
AWS account, not just a new one. If that lands
00:08:48.879 --> 00:08:51.320
cleanly, that's the version I'd actually want
00:08:51.320 --> 00:08:54.460
as a platform team because now you can meet developers
00:08:54.460 --> 00:08:57.399
where they are without losing governance. So
00:08:57.399 --> 00:09:00.379
the takeaway, if your dev platform can create
00:09:00.379 --> 00:09:03.480
AWS resources, your governance has to meet it
00:09:03.480 --> 00:09:06.919
there. The database is easy. The ownership model
00:09:06.919 --> 00:09:10.240
is the hard part. All right. Story 3 is about
00:09:10.240 --> 00:09:13.259
access, which matters even more once you have
00:09:13.259 --> 00:09:15.919
more accounts and more surfaces. TEAM stands
00:09:15.919 --> 00:09:19.620
for Temporary Elevated Access Management. It's
00:09:19.620 --> 00:09:22.820
an open -source solution built around AWS IAM
00:09:22.820 --> 00:09:26.139
Identity Center. The pitch is basically approval
00:09:26.139 --> 00:09:30.139
-based, time -bound elevated access to AWS accounts.
00:09:30.480 --> 00:09:32.919
Users request elevated access for a specific
00:09:32.919 --> 00:09:36.059
period of time, with a reason. Approvers approve
00:09:36.059 --> 00:09:39.720
or deny it. If approved, access is granted. When
00:09:39.720 --> 00:09:42.860
time expires, it is automatically removed. That
00:09:42.860 --> 00:09:45.600
automatic removal is the whole point. Because
00:09:45.600 --> 00:09:48.840
most orgs fail here. Someone gets admin just
00:09:48.840 --> 00:09:51.759
for this incident, and then it stays for months.
00:09:52.000 --> 00:09:55.000
Privilege creep becomes the default. Team also
00:09:55.000 --> 00:09:58.279
leans into auditing and visibility. Who requested
00:09:58.279 --> 00:10:01.879
what, who approved it, when it expired, plus
00:10:01.879 --> 00:10:05.129
session logging. How I'd frame it. This is not
00:10:05.129 --> 00:10:08.470
your break glass story. Break glass is the world
00:10:08.470 --> 00:10:11.970
is on fire and we need access right now. It should
00:10:11.970 --> 00:10:16.190
be rare, noisy, and heavily monitored. Team is
00:10:16.190 --> 00:10:19.269
the daily, I need admin for 45 minutes to do
00:10:19.269 --> 00:10:21.809
this legit change workflow. If you want this
00:10:21.809 --> 00:10:24.769
to actually work culturally, approvals need to
00:10:24.769 --> 00:10:27.309
be fast enough that people don't route around
00:10:27.309 --> 00:10:30.429
it. And the default permission sets need to be
00:10:30.429 --> 00:10:33.789
sane so elevation is actually meaningful, not
00:10:33.789 --> 00:10:37.169
just ceremony. If you are already on IAM Identity
00:10:37.169 --> 00:10:39.629
Center and you've been hand -waving we should
00:10:39.629 --> 00:10:43.190
do JIT access, team is at least worth a look.
00:10:43.309 --> 00:10:45.789
Even if you don't adopt it, it's a good reference
00:10:45.789 --> 00:10:48.710
for what time -bound elevation can look like
00:10:48.710 --> 00:10:51.389
without building everything from scratch. Alright,
00:10:51.590 --> 00:10:54.139
time for the lightning round. GitHub Actions
00:10:54.139 --> 00:10:57.000
improved performance on the workflows page. Small
00:10:57.000 --> 00:10:59.700
change, but if you live in Actions, you'll feel
00:10:59.700 --> 00:11:02.960
it. Big workflows render better now, lazy loading,
00:11:03.220 --> 00:11:06.159
and you can filter jobs by status so you can
00:11:06.159 --> 00:11:09.179
just see failures or in -progress stuff. During
00:11:09.179 --> 00:11:12.279
an incident, this is a real quality of life upgrade.
00:11:12.799 --> 00:11:15.860
Next, the weird one, Lambda managed instances.
00:11:16.299 --> 00:11:19.340
This is basically run Lambda functions on EC2,
00:11:19.519 --> 00:11:22.519
but AWS manages the lifecycle of those instances
00:11:22.519 --> 00:11:25.419
for you. It's meant for steady state workloads
00:11:25.419 --> 00:11:28.679
and specialized compute needs. It also changes
00:11:28.679 --> 00:11:31.480
concurrency and execution assumptions a bit.
00:11:31.580 --> 00:11:34.240
So you actually need to care about thread safety
00:11:34.240 --> 00:11:37.639
and shared state in ways you might not with regular
00:11:37.639 --> 00:11:40.960
Lambda. Interesting, slightly cursed. But I get
00:11:40.960 --> 00:11:43.879
why it exists. Atmos quick hit. There's a Cloud
00:11:43.879 --> 00:11:47.080
Posse Atmos issue where vendoring a component
00:11:47.080 --> 00:11:50.960
with an invalid URL triggers a weird GitHub username
00:11:50.960 --> 00:11:53.919
prompt. I'm mentioning it less because of the
00:11:53.919 --> 00:11:57.000
bug and more because Atmos is clearly turning
00:11:57.000 --> 00:12:01.279
into a bigger workflow ecosystem now. CLI, dev
00:12:01.279 --> 00:12:04.580
containers, IDE integrations, the whole thing.
00:12:04.720 --> 00:12:09.919
And last, k8sdiagram .fun. It's a free Kubernetes
00:12:09.919 --> 00:12:13.340
diagram builder that can also generate YAML for
00:12:13.340 --> 00:12:17.159
common resources. I would not blindly apply auto
00:12:17.159 --> 00:12:20.960
-generated YAML to prod, but for teaching, prototyping,
00:12:21.080 --> 00:12:23.940
or explaining architecture to humans, it's actually
00:12:23.940 --> 00:12:27.200
super handy. All right, let's close with the
00:12:27.200 --> 00:12:29.980
human story, because this ties into everything
00:12:29.980 --> 00:12:32.620
we just talked about. Mark Brooker wrote a post
00:12:32.620 --> 00:12:36.070
called, What Now? Handling errors in large systems.
00:12:36.350 --> 00:12:39.029
It's basically an interactive error handling
00:12:39.029 --> 00:12:42.169
game. You decide whether a system should crash
00:12:42.169 --> 00:12:45.269
or keep going when something goes wrong. And
00:12:45.269 --> 00:12:48.470
then he explains his take. The key idea is simple.
00:12:48.629 --> 00:12:51.250
And it's something we forget all the time. Error
00:12:51.250 --> 00:12:54.190
handling isn't a local decision. It's a global
00:12:54.190 --> 00:12:57.350
property of the system. We love to argue about
00:12:57.350 --> 00:13:00.610
one line of code. Should this crash? Should this
00:13:00.610 --> 00:13:03.960
retry? Should this be best effort? Mark's point
00:13:03.960 --> 00:13:07.139
is, that decision only makes sense if you understand
00:13:07.139 --> 00:13:10.059
the architecture around it. He asks questions
00:13:10.059 --> 00:13:13.620
like, are failures correlated? If the same bad
00:13:13.620 --> 00:13:17.059
input can hit every node, crashing can amplify
00:13:17.059 --> 00:13:20.299
the blast radius. Can a higher layer handle the
00:13:20.299 --> 00:13:23.340
air? Some architectures are designed to tolerate
00:13:23.340 --> 00:13:26.139
a few crashes. None are designed to tolerate
00:13:26.139 --> 00:13:29.169
a ton of crashes continuously. Is it actually
00:13:29.169 --> 00:13:32.169
safe to keep crashing? Is it actually safe to
00:13:32.169 --> 00:13:35.690
keep running? Sometimes continuing means silent
00:13:35.690 --> 00:13:38.750
corruption, which is worse than a crash. Then
00:13:38.750 --> 00:13:42.230
he ties it to blast radius reduction. Cell -based
00:13:42.230 --> 00:13:45.230
architectures, independent regions, isolating
00:13:45.230 --> 00:13:48.549
failures so you don't have a single mistake become
00:13:48.549 --> 00:13:51.789
a global outage. That connects to today's stories
00:13:51.789 --> 00:13:55.190
perfectly. Cloudflare's scheduler exists because
00:13:55.190 --> 00:13:58.389
humans will guess wrong sometimes, and the systems
00:13:58.389 --> 00:14:01.309
need to prevent correlated failure. The Vercel
00:14:01.309 --> 00:14:05.149
Marketplace DB story is a new surface area. The
00:14:05.149 --> 00:14:08.090
failure modes aren't just technical, they're
00:14:08.090 --> 00:14:11.110
governance and ownership failures. And team is
00:14:11.110 --> 00:14:13.970
literally about reducing the risk of one person
00:14:13.970 --> 00:14:17.350
with standing admin turning a mistake into a
00:14:17.350 --> 00:14:20.379
disaster. So yeah. If you want a good mindset
00:14:20.379 --> 00:14:23.340
going into next year, stop treating air handling
00:14:23.340 --> 00:14:27.000
like a code style preference. Treat it like architecture.
00:14:27.610 --> 00:14:30.169
All right, that's it for this episode of Ship
00:14:30.169 --> 00:14:32.409
It Weekly. We covered Cloudflare's maintenance
00:14:32.409 --> 00:14:35.190
scheduler on workers and the platform limits
00:14:35.190 --> 00:14:38.950
force better design lessons. AWS databases inside
00:14:38.950 --> 00:14:41.830
the Vercel marketplace and what that means for
00:14:41.830 --> 00:14:45.649
governance and blast radius. And team as a practical
00:14:45.649 --> 00:14:49.769
path to time -bound elevated access with IAM
00:14:49.769 --> 00:14:52.529
Identity Center. If you got something out of
00:14:52.529 --> 00:14:55.409
this, hit follow or subscribe wherever you are
00:14:55.409 --> 00:14:58.429
listening. And if you can, leave a quick rating
00:14:58.429 --> 00:15:01.090
or review. It's annoying how much that helps
00:15:01.090 --> 00:15:04.570
the show. Links and show notes are on shipitweekly
00:15:04.570 --> 00:15:08.309
.fm. And one last reminder, I'm looking for interview
00:15:08.309 --> 00:15:11.070
guests. If you want to come on and talk through
00:15:11.070 --> 00:15:14.250
real DevOps or platform work you're doing, hit
00:15:14.250 --> 00:15:16.950
the email on the site. I'm Brian. Thanks for
00:15:16.950 --> 00:15:19.330
listening. Happy holidays. And I'll see you next
00:15:19.330 --> 00:15:21.330
week, which is technically next year.