WEBVTT
NOTE
Transcription provided by Podhome.fm
Created: 09/24/2026
23:18:26Duration: 2980.310
Channels: 1
1
00:00:11.360 -->
00:00:15.440Hello, and welcome to the Data Engineering Podcast, the show about modern data management.
2
00:00:16.075 -->
00:00:20.074Today's episode is sponsored by Parallel, where agents find answers.
3
00:00:20.475 -->
00:00:27.195Most engineers today closely follow new model releases, but don't pay attention to their agent's most important tool, web search.
4
00:00:27.675 -->
00:00:40.140Parallel develops enterprise grade infrastructure for agents to retrieve high quality context from the web. Their core products are a suite of APIs for retrieving high quality information from the web with Pareto optimal quality, cost, and speed.
5
00:00:40.780 -->
00:00:53.245Whether you work on voice agents that need two hundred millisecond latency, chatbots that balance speed, depth, and quality, or long horizon agents to do thorough overnight research for you, Parallel is a single platform for all of your agentic research.
6
00:00:53.645 -->
00:00:57.805Get started for free today at dataengineeringpodcast.com/parallel.
7
00:00:58.220 -->
00:01:09.180Your host is Tobias Maci. And today, I'm interviewing Christopher Doidge about a data warehousing approach called agile ledger architecture that is designed to drive down data debt. So Christopher, can you start by introducing yourself?
8
00:01:09.740 -->
00:01:18.195Absolutely. Thank you so much, Tobias, and so happy to be here. My name is Christopher Doidge, founder and principal data architect at UNI Consulting LLC
9
00:01:18.354 -->
00:01:21.954and author of the book that we're gonna talk about today. So where are the insights?
10
00:01:22.274 -->
00:01:24.994It's the blueprint for agile ledger architecture.
11
00:01:25.555 -->
00:01:36.970And, core focus here is bridging the gap between technical data engineering and executive decision making. Essentially, we're looking to eliminate data debt and reduce overall executive time to insight.
12
00:01:37.369 -->
00:01:41.530And do you remember how you first got started working in the data space and what kept you there?
13
00:01:42.170 -->
00:01:50.185Yeah. This goes back fifteen years ago. I think starting off as a intern to be a sales rep at a medical
14
00:01:50.505 -->
00:01:55.065device company. Right? And I always thought I wanted to be an entrepreneur,
15
00:01:55.065 -->
00:02:03.870and that's really where it started. But I found very quickly that that was not my true forte, And it was very much more in organization
16
00:02:04.110 -->
00:02:08.270and problem solving, which really led me into the world of data itself.
17
00:02:08.670 -->
00:02:10.990It grew from there. I mean, my experience
18
00:02:11.150 -->
00:02:15.630ranges from working in strawberry plant distribution with Driscolls here in California,
19
00:02:16.145 -->
00:02:18.625all the way over to the leading classifieds
20
00:02:18.625 -->
00:02:22.945in Europe with, eBay Klein and Seigen, which is headquartered in Berlin.
21
00:02:23.105 -->
00:02:23.665So
22
00:02:24.065 -->
00:02:29.505a lot of it went from being in the actual trenches of analytics and just witnessing the frustration of
23
00:02:30.130 -->
00:02:35.170executives argue what is the true count of our new customers for an hour and a half
24
00:02:35.650 -->
00:02:41.010into then deciding, you know, it's not a tooling problem that we're having here.
25
00:02:41.410 -->
00:02:43.730It's really a question of architecture,
26
00:02:44.275 -->
00:02:48.435And that's really what led me to agile ledger architecture and how to build this.
27
00:02:48.915 -->
00:02:52.035Digging now into the agile ledger architecture,
28
00:02:52.035 -->
00:02:59.555I'm wondering if you can just give a little bit more color around what it is and some of the story about how you came to that design.
29
00:03:00.330 -->
00:03:08.250Yeah. Absolutely. It's a really great question. So the easiest way to explain this is we're missing a layer when it comes to our data architecture.
30
00:03:08.730 -->
00:03:13.290So, generally, when engineering and analytics works together to build a platform,
31
00:03:13.875 -->
00:03:17.475we're stopping at the analyst as our end deliverable.
32
00:03:17.954 -->
00:03:22.355So what I've witnessed in my career is we have these wonderful data tables. They're fantastic.
33
00:03:22.754 -->
00:03:35.810But in order to get it consumable to reach the actual insight that we need, we need another step beyond that. And so what agile ledger architecture does is it enforces strict ledger accounting
34
00:03:35.890 -->
00:03:37.010in the intake.
35
00:03:37.489 -->
00:03:41.569That way, we're essentially pushing business definitions all the way to the very front.
36
00:03:42.575 -->
00:03:46.015And our ultimate goal is the creation of a gold layer table.
37
00:03:46.254 -->
00:03:58.990So imagine this is a ready to consume table that anybody at your enterprise can connect to and pull the answers exactly how they need it. What I've typically seen is that when we think of a source of truth,
38
00:03:59.230 -->
00:04:04.590we imagine one giant Excel table that just has every possible thing that you can think of.
39
00:04:04.990 -->
00:04:06.670And so my whole
40
00:04:07.070 -->
00:04:14.725thought around this is let's deconstruct that. The true source of truth is actually understanding where does this data point come from,
41
00:04:15.125 -->
00:04:18.325what is the actual source around it, and when is it refreshed,
42
00:04:18.325 -->
00:04:20.005is generally what I need to know.
43
00:04:20.645 -->
00:04:24.645Using those definitions and then a series of taxonomies
44
00:04:24.860 -->
00:04:28.140to link our data together, we can essentially create
45
00:04:28.300 -->
00:04:29.740end state tables
46
00:04:29.900 -->
00:04:34.940that allow any stakeholder in our business to get to the same answer repeatedly
47
00:04:34.940 -->
00:04:53.105over and over again. And this is what the heart of Agile Ledger is. It's serving up our data not just so that your analysts have a nice clean playground to play in, but it's also so that you have a consumable layer so that your stakeholders can get to the answers they need without having to write a nice big complex SQL query.
48
00:04:53.660 -->
00:04:54.620One of the,
49
00:04:55.180 -->
00:05:02.380I guess, temptations whenever you're talking about data warehousing is to say, okay. Well, I'm running into troubles with
50
00:05:02.540 -->
00:05:10.885whatever approach. I'm going to develop a new one. Obviously, there is forty years plus or minus of data warehouse,
51
00:05:11.044 -->
00:05:13.445best practices and principles and
52
00:05:13.605 -->
00:05:15.285designs and architectures,
53
00:05:15.365 -->
00:05:21.680and everybody seems to have developed their own opinion. Some of them are an addendum to preexisting
54
00:05:21.680 -->
00:05:33.474approaches such as the Kimball star style or the Inman third normal form, or there's DataVault, which is sort of going in its own direction. I've seen anchor modeling. There's the whole medallion architecture,
55
00:05:33.474 -->
00:05:40.035ETL versus ELT. So there's a lot to try and figure out. And I'm wondering if you can give your
56
00:05:40.435 -->
00:05:48.230summary of how you see the Agile Ledger architecture fitting in that overall ecosystem ecosystem of everybody having their own opinion and there being
57
00:05:48.789 -->
00:05:52.389I don't even know how many different architectures available to choose from.
58
00:05:52.949 -->
00:05:54.470Yeah. Absolutely.
59
00:05:54.470 -->
00:06:23.980And you're a 100% right. I mean, I'm remembering back to the days when we just had an Excel file, and it was this is the column is our source of truth, so to say. And now there's these wonderful tools where we can pull data in from different disconnected systems and take them through a series of training, so to say, in order to get them ready to consume and analyze on. So, essentially, ALA doesn't try to replace any of these. It's actually built in coexistence with it. Right? It's just the way to amplify it.
60
00:06:24.380 -->
00:06:34.965Essentially, what we're talking about at the core is that you do not release a gold layer table that's ready to consume for the business without having clear definitions of how this is being put together.
61
00:06:35.365 -->
00:06:51.420So this is where I said earlier where we sort of stop at the analyst when it comes to delivering up data. We think, okay. Let's get the data as clean as we can, and we'll place it here. But then the analyst still has to take it through particular steps in order to get it sort of business ready.
62
00:06:51.900 -->
00:07:10.405And I've even seen this working at, as I mentioned, in eBay in Germany. It's we had tables where I still needed to run deduplication. I still needed to filter out specific statuses, and this is never gonna end. You know, business logic constantly evolves over time. And so it's impossible
63
00:07:10.405 -->
00:07:11.445to pour
64
00:07:11.604 -->
00:07:18.199sort of like a concrete layer of data and assume that it's always gonna solve every need that you have in the future.
65
00:07:18.680 -->
00:07:23.240So this is where ALA comes into play because what you can do is
66
00:07:23.479 -->
00:07:31.635lock in your source exactly how you have it. And then as the future changes and the business starts asking different questions of the data,
67
00:07:31.875 -->
00:07:35.315you can evolve your gold layer tables to answer those questions.
68
00:07:35.875 -->
00:07:39.794And this is where I introduced the idea or the notion of a clay model.
69
00:07:40.410 -->
00:07:47.290So you're working with clay. That structure isn't always necessarily solidified until your absolute end.
70
00:07:47.530 -->
00:08:01.495But the idea around the clay model is that you can build it up to whatever sort of architecture that you need for the current sprint or for the current need and then tear it down and rebuild it as you need it. So what ALA does is it really focuses
71
00:08:01.735 -->
00:08:02.695hyperfocuses
72
00:08:02.695 -->
00:08:31.525actually on specific data debt that is needed for the business right now. That way I can take to attach some context to this. I could take my orders out of, like, a system like Shopify. I can break it down to particular products. I can go down to counties that I'm shipping it to, go down to whatever level of granularity that we need to. But we build those taxonomies and those bridges as our business needs evolve instead of just trying to tackle everything at the same time, which is what I've commonly seen.
73
00:08:31.925 -->
00:08:45.880So it's a way to organize this and prioritize it with the business. That way you see what's actually working and you enable analytics to be your true compass. You know, leveraging your data to tell you you're going in the wrong direction. Let's retry and test.
74
00:08:46.200 -->
00:09:02.695So it really complements those systems well because ALA works especially well with the medallion architecture as well too. So it's not about tearing any of those down. It's just about attaching a philosophy or an ideology into the way that you architect your end state data.
75
00:09:03.495 -->
00:09:06.055And one of the interesting things is
76
00:09:06.390 -->
00:09:14.950that medallion architecture in particular isn't actually its own style of data warehousing. It's just a way of sequencing the transformations.
77
00:09:15.670 -->
00:09:17.590Yeah. Absolutely. Yeah. Great point.
78
00:09:19.045 -->
00:09:20.245And so
79
00:09:20.404 -->
00:09:42.950digging now into the overall question of data debt where sometimes you want to actually take on some of that debt because it's necessary to get to some learning or discovery to actually help cement what is the thing that you're actually trying to do with it. Sometimes it's just something that happens incidentally, and I'm curious if you can talk to some of the ways that that debt manifests
80
00:09:42.950 -->
00:09:46.205and some of the decision structure about
81
00:09:46.445 -->
00:09:51.325whether and how to be intentional about it and some of the ways that it happens accidentally.
82
00:09:52.365 -->
00:10:00.620Yeah. No. That's excellent point. So I I think back to my experience with Driscolls. And so, Driscolls were dealing with plant operations
83
00:10:00.620 -->
00:10:02.460all over the world. And
84
00:10:02.860 -->
00:10:08.780large companies like this have a lot of partners that they work with along the entire supply chain.
85
00:10:09.180 -->
00:10:17.845So you could have the most amazing data architecture in the world, but your partners aren't going to. And a lot of times you cannot control this.
86
00:10:18.165 -->
00:10:35.970And so this is a great example of data debt. If I receive an input from a partner telling me what they have in their three p l or whatever storage concern they're giving me or data they're giving me, I can't necessarily always receive it in a format that works for me and then have to convert it over.
87
00:10:36.290 -->
00:10:43.569So this is a classic form of data that you're gonna have your tech savvy analyst jump in there and start, you know, adding helper formulas to
88
00:10:43.925 -->
00:10:49.045convert this data to something that fits into your system. That way you can correct the intake.
89
00:10:49.204 -->
00:10:56.084You know, not everybody has an API. Not everyone can serve up the information the way that you need it. And so
90
00:10:56.350 -->
00:10:58.430this just snowballs over time.
91
00:10:58.830 -->
00:10:59.390And,
92
00:10:59.950 -->
00:11:27.019ideally, what really helps you to drive the decision is is this going to impact the business and move us forward in a way, or is this just some easy, simple, busy work that we can continue to do because it's not really worth the struggle to actually cover this or not? So I would say the business needs and what's really gonna drive your needle should really dictate which data that you actually cover or not because all of them are required different levels of resources.
93
00:11:28.380 -->
00:11:30.540One of the other interesting aspects of
94
00:11:30.779 -->
00:11:37.660debt, whether it's technical debt, data debt, financial debt, is some people unfortunately figure out is that it's not always,
95
00:11:38.145 -->
00:12:07.810I guess, financial debt's pretty straightforward. But technical debt and data debt, it's not always obvious that you even have that until you get to a point where you're actually trying to achieve some outcome and there are things standing in your way. The way that I have usually seen technical debt kind of identified is that it is the thing that adds friction to what you're trying to do. And I'm wondering if you could just talk to some of the ways that people can even identify whether and how much debt they have and then more specifically at a pinpoint where it's coming from.
96
00:12:09.305 -->
00:12:20.105Yeah. Absolutely. So for me, I found that the easiest way to pinpoint this is if I look at two different datasets and I know that I need to combine these to get to the next level of my analysis,
97
00:12:20.505 -->
00:12:35.990my first initial question is how do I join this? Right? And to me that right there is your data debt. If you cannot figure that out within just a couple seconds of looking at it because there's no columns that seem to be the same or there's no primary key, surrogate key,
98
00:12:36.150 -->
00:12:45.475there's your data debt right there. So it's looking you right in the face. And so with this, it comes with, alright, are we able to do a simple proration
99
00:12:45.555 -->
00:12:46.915and move it forward?
100
00:12:47.795 -->
00:12:49.555Is it something that's gonna take
101
00:12:49.875 -->
00:12:53.330a lot of resources? Cause now I need to build a taxonomy
102
00:12:53.330 -->
00:13:03.250to then say, okay. When you see this record in system a, convert it to this record in system b as a high level example. Right? There's some way that we gotta build up that join,
103
00:13:03.490 -->
00:13:09.115and that's exactly what Datadet is here. So the way that I like to do it is I'll identify
104
00:13:09.115 -->
00:13:11.835what it is that needs to be done for this analysis
105
00:13:12.555 -->
00:13:19.835and, what are the pieces that we'll need to put resources to in order to tackle this. How much time do we estimate this will take?
106
00:13:20.699 -->
00:13:36.315That will also help us understand whether we should even tackle this data debt or leave it on our never ending backlog because it's not that big of a deal. What I've typically found, what happens is that you tend to have these hero analysts in the team that are fantastic,
107
00:13:36.555 -->
00:13:41.355but they tend to just do this in the background with scripts on their local machines or
108
00:13:41.835 -->
00:13:57.579some business context that they've heard in a meeting at some point. And the idea around Datadet is let's get this documented and actually bring life to it. A lot of times, they're just living in somebody's head, and then no one's even aware of the debt, just as you mentioned.
109
00:13:57.899 -->
00:14:05.535But at some point, we all have to pay it forward. So we need to find a way to, you know, bring this to light, bring visibility to it.
110
00:14:05.935 -->
00:14:16.015And I found that the best way to do that is with the construction of these gold tables. Because as I try to make data more and more consumable to the nontechnical
111
00:14:16.015 -->
00:14:16.335user,
112
00:14:17.060 -->
00:14:22.580that to me quickly identifies which pieces I just cannot join without something simple.
113
00:14:23.300 -->
00:14:34.795And so digging now into the architecture itself, I'm wondering if you can give a bit of an overview about what are some of the core principles that you are using as the foundation of the architecture
114
00:14:34.795 -->
00:14:38.795and what are some of the ways that people should be identifying
115
00:14:39.115 -->
00:14:44.555whether and how it might fit into their overall data modeling practice and data warehouse design.
116
00:14:45.830 -->
00:14:53.430Yeah. Absolutely. So for me, it really boils down to a very simple test. I just call it the 15 litmus test.
117
00:14:53.670 -->
00:15:03.415The idea around this is your total time to insight. So if you imagine a stakeholder coming in and asking a very simple question, how many customers
118
00:15:04.135 -->
00:15:09.255did we have yesterday? As simple as that. If this takes you more than fifteen minutes
119
00:15:09.495 -->
00:15:13.575to answer, then you are in need of agile ledger architecture.
120
00:15:15.040 -->
00:15:20.959If it's less than that, you're fine. Just move on. Why why improve something that's already working quite well?
121
00:15:21.360 -->
00:15:36.845But the idea is that whole test is if it's taking longer than fifteen minutes to answer basic business questions and each time you wanna understand, okay, do I have the data to back the answer for this? And that takes longer than fifteen minutes.
122
00:15:37.245 -->
00:15:40.685Then that's when you know that you need some work in your architecture.
123
00:15:41.404 -->
00:15:48.370We're built off of some particular principles as you mentioned. So obviously, the core principles are that fifteen minute litmus test.
124
00:15:48.610 -->
00:15:53.330Medallion pipeline works very, very well for this. It's built off of the foundations.
125
00:15:53.649 -->
00:16:01.355So, to reiterate, that's using bronze as your loading dock to just ingest the basic data. I know the listeners probably already know this.
126
00:16:01.675 -->
00:16:09.835Silver for cleaned ledgers and then ultimately gold, which would be our fluid clay model. Another core principle is the upstream integrity.
127
00:16:09.995 -->
00:16:17.870So this is fixing data anomalies right at the entry instead of waiting towards downstream or trying to correct them in complex queries
128
00:16:18.029 -->
00:16:20.910that live outside of our medallion pipeline.
129
00:16:21.470 -->
00:16:41.555For me, the easiest way to identify this is another core principle, the unsung hero, which is my favorite report since the first days of analytics. It's the data error report. It's the one thing no one ever wants to see. It's the continuous scan that looks through our structures and looks for all the gaps. It's essentially in layman's terms, select star
130
00:16:41.555 -->
00:16:43.875where everything is wrong and shouldn't be that way.
131
00:16:44.699 -->
00:16:51.740And then I could go through and scrub the list and actually correct this. That way our goal layer remains intact and accurate.
132
00:16:52.379 -->
00:16:54.540Given that the medallion
133
00:16:54.540 -->
00:16:58.255approach is the presumed basis,
134
00:16:58.335 -->
00:17:02.495what are some of the ways that that informs or influences
135
00:17:02.495 -->
00:17:03.855the technological
136
00:17:03.855 -->
00:17:05.535substrates that you are
137
00:17:06.255 -->
00:17:26.659supposing for this architectural approach? So are you presuming that people have any sort of data streaming? Are you presuming that they're using a traditional data warehouse appliance or a lake house architecture and just some of the ways that the underlying technology has an impact on the way that you actually adopt and implement the principles in in the architecture that you're proposing?
138
00:17:27.914 -->
00:17:29.754Yeah. Absolutely. So, actually,
139
00:17:30.315 -->
00:17:37.274I developed this architecture with the thoughts when I was working in the trenches with very sophisticated tooling.
140
00:17:37.595 -->
00:17:40.075Right? So this is where it actually came to life.
141
00:17:40.720 -->
00:17:47.600However, I didn't even realize it until I was supporting a client who strictly worked off of Google Sheets.
142
00:17:47.920 -->
00:17:57.935So we're talking the lowest code that you can go. There was no API here. They were downloading data from a Shopify store, exporting into a CSV,
143
00:17:58.255 -->
00:18:02.015opening it up in Excel, and then starting their full on analysis.
144
00:18:02.415 -->
00:18:05.775And that's where it really dawned on me because what happened was
145
00:18:06.210 -->
00:18:09.009every single time they wanted to repeat this analysis,
146
00:18:09.010 -->
00:18:12.369they had to redo that entire process from beginning to end
147
00:18:13.169 -->
00:18:23.075every single time. And so that's when it clicked where I thought, well, why are you doing it that way? We can take the principles that we learn from more sophisticated tools
148
00:18:23.155 -->
00:18:26.034that allow something like a medallion pipeline
149
00:18:26.115 -->
00:18:27.794innate with its UI
150
00:18:28.035 -->
00:18:32.675and actually push it all the way back to the very basics of an Excel or a Google Sheet.
151
00:18:33.160 -->
00:18:38.600So you can use this similar philosophy to your architecture regardless of what tooling
152
00:18:38.600 -->
00:18:45.400and software that you actually have. The whole idea is around, like, loading the actual data,
153
00:18:46.294 -->
00:18:47.095transforming
154
00:18:47.095 -->
00:18:47.734it,
155
00:18:47.975 -->
00:18:54.614and then presenting it in a way that can be consumed by everyone at the business. And that's down to its basic, basic core.
156
00:18:54.855 -->
00:19:11.629You can do that with absolutely any system that you have. Of course, it gets easier when you have better tools because you don't have to manually export CSVs from files. You could just let an API do it for you. So it works better when you're have more sophisticated toolings.
157
00:19:11.950 -->
00:19:12.749However,
158
00:19:12.750 -->
00:19:16.429it still functions fine when you're talking a low code environment.
159
00:19:17.505 -->
00:19:29.825So for people who already have a substantial amount of investment in a given data warehouse architecture, obviously, there's a lot of inertia built in to how they're doing things, way that they have their system designed.
160
00:19:29.825 -->
00:19:34.460It's not easy to immediately pivot to a new way of doing things.
161
00:19:34.780 -->
00:19:40.780And I'm wondering if you can talk through what the adoption process looks like for people who do have
162
00:19:41.100 -->
00:19:54.334even just a medium scale data warehouse, and they've already built out a lot of pipelines around those assumptions and just some of the upfront work that's necessary to start planning out that adoption and how you can do that incrementally.
163
00:19:55.455 -->
00:19:56.255Yeah. Absolutely.
164
00:19:56.770 -->
00:20:02.690So what I like to do when faced with this scenario is to take the quarterly sprint approach.
165
00:20:02.929 -->
00:20:03.490So,
166
00:20:04.049 -->
00:20:07.570obviously, when you're talking a new architecture, a new system,
167
00:20:07.890 -->
00:20:21.835I think every company that I've worked at always has a new system incoming. It seems to be the buzzword that corporate likes to use all the time is, oh, there's a new system coming that'll solve all our problems, and then it never comes. So what I found is
168
00:20:22.475 -->
00:20:25.675this type of architecture can be built alongside
169
00:20:25.675 -->
00:20:26.960what you already have.
170
00:20:27.280 -->
00:20:33.679So you're not needing to burn the entire warehouse down and waste everything that you've already built. What you already have is fine.
171
00:20:34.000 -->
00:20:42.934What I recommend that you do is you look at your upcoming next quarter and you prioritize what is the most important questions that I know stakeholders are going to have.
172
00:20:43.495 -->
00:20:51.414What is sort of a KPI shield that I can erect that can let them pull the data that they need at a self-service volume,
173
00:20:51.495 -->
00:21:03.639answer the questions that they have, and then my team can actually focus on the things that are gonna actually help us instead of trying to give you the kind of customers to fill for your Google slide deck. Alright?
174
00:21:04.200 -->
00:21:07.480So it's bringing to light the data that we have now
175
00:21:07.774 -->
00:21:11.854and then logging any data debt for anything that we cannot do.
176
00:21:12.095 -->
00:21:18.254So as a crystal example or crystal clear example to give here, if the quarterly objective is to
177
00:21:18.654 -->
00:21:21.455improve picking velocity by 15%
178
00:21:21.455 -->
00:21:22.894in the warehouse, let's say.
179
00:21:23.630 -->
00:21:27.470And we understand, okay. One of the metrics that they're gonna wanna constantly
180
00:21:27.470 -->
00:21:31.229review is what is our overall packing time on orders.
181
00:21:31.630 -->
00:21:35.229So for every order, how many minutes is it taking as an example?
182
00:21:35.710 -->
00:21:48.625Perhaps I don't have the data to do that yet. So my data debt is I know where the items are located in the warehouse, but for some reason, I cannot tie the items to the actual order itself yet. Right? Unfortunately,
183
00:21:48.785 -->
00:21:51.585that would be a huge problem, but let's pretend that's our problem.
184
00:21:52.180 -->
00:21:58.660If that's the case, that is a data debt ticket. That's exactly what I would log, and that would be one of the first things that I would tackle
185
00:21:58.820 -->
00:22:03.940in order to give the business the data that they need for this upcoming sprint.
186
00:22:04.500 -->
00:22:06.820So this is the way that I would prioritize it
187
00:22:07.345 -->
00:22:26.460because, again, the beauty of it being a clay model is that at the end of our quarter or even throughout it, throughout the various feedback loops that you have, where stakeholders are gonna be very clear about what's not working with the data model that you presented to them. You can iterate on it. So you can see what is it that's missing,
188
00:22:27.020 -->
00:22:33.900what is it that we need to add to it, and this can obviously illuminate further data that so that you can continue down this.
189
00:22:34.395 -->
00:22:37.995It's not a rabbit hole. It's the exact opposite of a rabbit hole.
190
00:22:38.235 -->
00:22:48.155This value driver, I guess, is what I could call it. You can keep going down this path of obtaining the critical data that business is going to need to make their best decisions.
191
00:22:48.820 -->
00:22:52.020And a lot of times in business that doesn't happen
192
00:22:52.180 -->
00:22:56.500until you have everybody looking at the same thing and speaking the same language.
193
00:22:56.900 -->
00:23:04.660So this helps align everybody completely across the board. I can see too many times where I've been in these meetings and everybody is
194
00:23:05.435 -->
00:23:12.715presenting the exact same metric with totally different numbers. And it's because they all obtained it a completely different way.
195
00:23:13.115 -->
00:23:16.875So this aligns the business to the same metrics and definitions.
196
00:23:17.880 -->
00:23:20.360The other challenge for any
197
00:23:20.600 -->
00:23:24.360sort of correctness campaign in an engineering organization
198
00:23:24.360 -->
00:23:26.760is that there's always the
199
00:23:27.640 -->
00:23:29.640issue of backsliding
200
00:23:29.640 -->
00:23:32.680or
201
00:23:30.635 -->
00:23:39.595blocking people getting real work done. So one of the analogies that comes to mind is something like the introduction of a type checker for something like a Python code base or
202
00:23:39.835 -->
00:23:50.570the introduction of a new linting rule that has 10,000 violations at the start, but that you wanna start driving down. And so there's the overall approach of, okay. Well, I'm going to accept this baseline,
203
00:23:50.570 -->
00:24:01.445but I'm going to ratchet it so that every time it gets better, I'd never allow it to regress and nobody can introduce a new violation. And so I'm curious what are some of those technical controls that are available
204
00:24:01.605 -->
00:24:15.179as you are adopting this Agile Ledger architecture for being able to say, okay. This is how we want it to be at the end state. This is where we are right now, and this is how we're going to make sure that we don't have any of that regression as we make these changes.
205
00:24:16.460 -->
00:24:18.540Yeah. Exactly. With
206
00:24:18.540 -->
00:24:20.460the ALA architecture,
207
00:24:20.620 -->
00:24:24.860for me, the easiest way around this is to have a strict gatekeeper
208
00:24:24.860 -->
00:24:28.700on the data dictionary. So it's something we haven't spoken too much about.
209
00:24:29.100 -->
00:24:30.865But with ALA,
210
00:24:30.945 -->
00:24:32.465one of the deliverables
211
00:24:32.465 -->
00:24:35.745is every time you present a gold layer table,
212
00:24:36.065 -->
00:24:39.024you must have an update in your data dictionary.
213
00:24:39.185 -->
00:24:40.145So historically,
214
00:24:40.145 -->
00:24:42.225this is also known as a business lexicon.
215
00:24:42.799 -->
00:24:48.159It's essentially a warehouse of knowledge for your stakeholders to be able to go and understand
216
00:24:48.240 -->
00:24:50.159how are we defining this metric,
217
00:24:50.720 -->
00:24:55.440what is the business definition, and what is the data definition on it. Ultimately,
218
00:24:55.520 -->
00:24:59.524they wanna know that if I think of a metric as new customers,
219
00:24:59.525 -->
00:25:06.004that we're all speaking the exact same thing. What does new customers actually mean to the business and having this crystal clear?
220
00:25:06.245 -->
00:25:15.149So for me, enforcing a strict rule where no gold table hits production unless its sources, transformations, and business definitions are a 100% documented
221
00:25:15.309 -->
00:25:28.724helps us to stay on track with this because it will keep you from backstepping and taking shortcuts to let's just throw it into the gold table now and then figure it out later because it's gonna come back to bite you very quickly.
222
00:25:29.125 -->
00:25:35.765The whole point with ALA is that since it has a logical setup and you're cleanly labeling definitions
223
00:25:35.765 -->
00:25:41.659all the way through from inception, you're bringing a lot of trust to your data, which is the most important thing.
224
00:25:42.060 -->
00:25:50.700It's easy to have data and to perform an analysis. The hard thing is having trust on it, especially as you scale up an enterprise.
225
00:25:51.020 -->
00:25:52.780When you're moving to larger companies,
226
00:25:53.195 -->
00:25:59.914such as eBay or Driscoll's, and you're dealing with billions and billions of rows of record, there's a lot of noise.
227
00:26:00.235 -->
00:26:03.355And so to really strip it away and get to the true insight,
228
00:26:03.675 -->
00:26:08.155it's not about scaling and performing massive rules on absolutely everything.
229
00:26:08.740 -->
00:26:20.340It's focusing on what's truly important to the business now so that we can learn our best practices along the way. And then those are what we enforce as we go through. But to me, it's about prioritization
230
00:26:20.340 -->
00:26:25.904is what I would say. So stick to one path and that can easily be your quarterly objectives
231
00:26:26.065 -->
00:26:32.465and then build it out from there. But don't try to sprint to everything at the same time. The other challenge,
232
00:26:32.784 -->
00:26:35.825particularly in warehouse architecture,
233
00:26:35.905 -->
00:26:46.279is that every time you introduce a new principle, you have to educate everyone about how it's supposed to work, how to do it. And, I mean, the star schema approach has been around for,
234
00:26:46.759 -->
00:27:00.394I think, more than thirty years at this point, and it's still shocking how many people who work in the data space who have never actually come to grips with its formal definitions or even the a casual approach of how to do it correctly. And so
235
00:27:00.635 -->
00:27:01.755for people who
236
00:27:02.220 -->
00:27:07.340do have an existing warehouse or even people who are saying, I'm gonna go and build a warehouse right now,
237
00:27:07.820 -->
00:27:12.779what are some of the ways that they need to be thinking about sharing some of the
238
00:27:13.179 -->
00:27:21.735core tenets and principles and then also layering in some of those technical constraints about how to do this properly so that you don't have one person
239
00:27:22.215 -->
00:27:35.929having one interpretation and doing their own thing over on this set of tables and somebody else thinks about it differently and has a different way of managing that in a different area of the warehouse and just some of the ways of building cohesion and shared understanding?
240
00:27:36.810 -->
00:27:46.810Yeah. Absolutely. No. That's a fantastic question. And I think that's where the simplicity of the medallion architecture comes into play because you enforce strict rules at each layer.
241
00:27:47.455 -->
00:27:50.575So for example, bronze, there's no rules. It's
242
00:27:50.735 -->
00:27:57.135all about speed. All I'm trying to do is get the source original source data loaded into my warehouse
243
00:27:57.215 -->
00:28:01.535as fast as humanly and machinely possible. That's my ultimate goal. No rules.
244
00:28:02.299 -->
00:28:06.619When it comes to silver, this is where I'm applying all of our business context.
245
00:28:06.860 -->
00:28:08.379So for example,
246
00:28:08.620 -->
00:28:21.184every order ID is always duplicated, so make sure you deduplicate it. Make sure you run a window sum across this. You know? Any other rules that you can think of on it. Strip out the dollar sign that seems to be coming into the revenue field,
247
00:28:21.505 -->
00:28:22.945convert it into decimal,
248
00:28:23.825 -->
00:28:29.825whatever the case may be that you wanna apply to that. But those rules are documented and loaded into Silver itself.
249
00:28:30.169 -->
00:28:37.849So as you're constructing that ledger and you're making a very clean, immutable ledger with the strict accounting philosophy,
250
00:28:38.490 -->
00:28:41.450you're setting a core list of rules.
251
00:28:41.690 -->
00:28:50.325And so it cannot change from those. These are how it is because the if you make a change into here, you're gonna obviously break everything downstream
252
00:28:50.485 -->
00:28:52.565that happens after that as well too.
253
00:28:52.965 -->
00:28:53.524So
254
00:28:53.845 -->
00:29:02.359that is how I found strict enforcement onto it and also by isolating them. So silver should be specific to the source that it comes into.
255
00:29:02.679 -->
00:29:19.934And then as you mentioned with the it it actually fills straight into a star schema because if you think of the visual layout of it, each of those independent tables are your clean silver ledgers is what you're doing here. So it's taking the exact same principles that exist. We're just saying
256
00:29:20.335 -->
00:29:34.049on top of the architecture that you're doing, you're focusing on the wrong things to start off with. So continue what you're doing. Just your focus should be aligned closer to what the business needs to move forward
257
00:29:34.450 -->
00:29:39.010and bring that information visible so that everybody's on the same page.
258
00:29:39.575 -->
00:29:45.254There are so many times that I've seen us create data tables and construct a full star schema
259
00:29:45.255 -->
00:29:47.975for then the business rule to completely change.
260
00:29:48.535 -->
00:29:50.055And then when that happens,
261
00:29:50.055 -->
00:29:53.335your immediate instinct is to go back and rebuild the ledgers,
262
00:29:53.830 -->
00:30:00.870But that doesn't necessarily have to be the case. You can simply adjust the logic that goes into one of the ledgers to now exclude
263
00:30:01.030 -->
00:30:07.750order IDs that come from one channel as an example. And that is how you can update and fix these without having to
264
00:30:08.625 -->
00:30:10.145break anything downstream.
265
00:30:10.625 -->
00:30:11.985With the idea of
266
00:30:12.465 -->
00:30:15.745ledgers, it brings up a number of different
267
00:30:15.985 -->
00:30:17.184associations.
268
00:30:17.185 -->
00:30:29.990So there's the ledger for an accounting ledger where you've got double entry bookkeeping. There's also the ledger where in the sort of crypto coin area of having the blockchain of an ordered sequence of auditable events.
269
00:30:30.230 -->
00:30:40.354There's the ledger of sort of a decision ledger of this is why we did it this way, and I'm wondering if you can just talk through some of the ways that you want to think about the
270
00:30:40.674 -->
00:30:46.674durability of that ledger. You said you don't necessarily wanna go and wipe it and rebuild it. You just wanna modify it and just some of the
271
00:30:47.970 -->
00:30:54.690technical aspects of that ledger and making sure that it doesn't have too much of a
272
00:30:54.770 -->
00:31:01.010oh, we'll just rebuild it whenever we want to sort of from an event stream principle, but the ledger itself actually
273
00:31:01.010 -->
00:31:03.414being a durable system of record.
274
00:31:04.294 -->
00:31:08.455Yeah. And it it stems from the movement from bronze to silver.
275
00:31:08.695 -->
00:31:15.735Exactly. Because if you think about it, when your bronze stumps in and it creates a specific now set of data that's available
276
00:31:16.140 -->
00:31:19.419for you to now transform into your Silver Ledger.
277
00:31:19.820 -->
00:31:22.299If you think about it as a
278
00:31:22.460 -->
00:31:28.860hard coded list is the way I try to think of a ledger is this is the truth. This is what it is that we want.
279
00:31:29.500 -->
00:31:31.580And so
280
00:31:31.205 -->
00:31:36.085if you're adding new filters, new business context that comes down into play,
281
00:31:36.405 -->
00:31:42.565it's about recording that logic into here so that that can evolve over time because that will happen.
282
00:31:43.780 -->
00:31:53.220Businesses will make critical decisions and decide to change the way that we count customers, change the way that we actually want to classify our orders or categorize them.
283
00:31:53.540 -->
00:32:00.934And those things can happen. But what I've commonly found is that this context doesn't live in our systems and it's not documented.
284
00:32:01.414 -->
00:32:04.774What tends to happen is that it lives in private scripts.
285
00:32:05.095 -->
00:32:05.894And so
286
00:32:06.135 -->
00:32:21.539the core philosophy and ideology of ALA is pushing this documentation through to the inception. It's so making sure that we do not lose it over time because there's been too many times that you have that one data person leave your company,
287
00:32:21.700 -->
00:32:27.614and then there goes the entire business context and logic of how this is constructed and going on.
288
00:32:28.014 -->
00:32:40.760The idea around this is constructing a system that's going to outlive you and is going to be there. Right? And in order for that to happen, you have to have the rules listed into there.
289
00:32:41.080 -->
00:32:52.455It helps when you're in those conversations with the stakeholders because then it's not so much of, I don't know the source and I'm not quite sure the rules that we're applying to it to get you this list of customers.
290
00:32:52.775 -->
00:32:54.295It's clearly documented.
291
00:32:54.295 -->
00:32:56.855We can go back and review it together and see it.
292
00:32:57.175 -->
00:33:00.295So for me, it's really about the documentation
293
00:33:00.295 -->
00:33:02.055that enables the governance for this.
294
00:33:03.030 -->
00:33:04.070One of the
295
00:33:04.470 -->
00:33:14.630other aspects of the time in which we are right now is that you can't really get through any conversation about anything without AI coming up. And in particular, in the data engineering space,
296
00:33:15.575 -->
00:33:22.054Text to SQL is one of the earlier examples of, oh, this can actually be really useful to help me answer questions.
297
00:33:22.135 -->
00:33:26.615But as we invest more and more into agentic approaches,
298
00:33:27.255 -->
00:33:36.270being able to make sure that the warehouse is navigable and scrutable by an LLM or an agent is increasingly valuable.
299
00:33:36.590 -->
00:33:43.550And I'm wondering what are the benefits that that ledger and governance as part of the warehouse adds to
300
00:33:43.745 -->
00:33:49.505some of those workloads where you don't necessarily have a human with all of that organizational knowledge
301
00:33:49.745 -->
00:33:55.585ready at hand. And instead, you have an agent that needs to build that context
302
00:33:55.665 -->
00:33:56.785fresh every time,
303
00:33:57.430 -->
00:34:07.750you know, where maybe that context lives as a durable system of record somewhere, but it's still every time the LLM launches, it's coming up fresh. It needs to make sure that it's pulling in all the right detail.
304
00:34:08.550 -->
00:34:15.375Yeah. Absolutely. And it's a wonderful point. I mean, we can think back to the age old analogy of garbage in garbage out.
305
00:34:15.615 -->
00:34:20.335And one thing that we've seen with AI is that it's phenomenal at accelerating
306
00:34:20.335 -->
00:34:22.895very bad information when it can.
307
00:34:23.535 -->
00:34:30.820So in order to prevent that and ensure that you're actually, you're having the AI build off of useful business context
308
00:34:30.900 -->
00:34:34.820exactly to what you're trying to accomplish here is through this system.
309
00:34:35.380 -->
00:34:37.700What I have seen so many times is
310
00:34:37.940 -->
00:34:40.420users will take a massive dataset,
311
00:34:40.734 -->
00:34:44.175hand it to an AI agent and say, go find me the insights.
312
00:34:44.655 -->
00:34:47.615And that can be incredibly useless because
313
00:34:47.615 -->
00:34:52.015it's gonna dive through and perform multiple regression analysis.
314
00:34:52.015 -->
00:35:00.849It's gonna try to find different links that it can. But the point is you've been in business for x years. You've already done that for a period of time.
315
00:35:01.250 -->
00:35:04.770So why would you have the AI start from scratch every single time?
316
00:35:05.089 -->
00:35:09.730It translates directly into the context windows because there's nothing more frustrating
317
00:35:10.265 -->
00:35:13.705with AI systems than the retention windows of its context.
318
00:35:13.945 -->
00:35:16.185So you'll talk through and spend hours
319
00:35:16.505 -->
00:35:20.585getting down to the core of an analysis with an AI system.
320
00:35:20.825 -->
00:35:25.625You'll upload the dataset to it, you'll figure something out, and you'll get a wonderful moment.
321
00:35:25.760 -->
00:35:32.400Oh, wow. All of the marketing campaigns that we've been doing have been completely wrong. We've been leaving out this little URL
322
00:35:32.720 -->
00:35:34.880or this little ID in the URL.
323
00:35:35.040 -->
00:35:39.520So when it comes into our warehouse, we have absolutely no way to tie it to the true orders.
324
00:35:39.920 -->
00:35:40.480And
325
00:35:41.015 -->
00:35:45.575you finally discovered this through digging through with the AI and figuring that all out.
326
00:35:46.215 -->
00:35:48.375Why would you want your next conversation
327
00:35:48.775 -->
00:35:59.410to completely start from the beginning and now go into a different route when you've already uncovered from that? You wanna start with that as the context base and for it to load that up and then go.
328
00:35:59.890 -->
00:36:00.530So
329
00:36:00.770 -->
00:36:05.890what you're doing is you're enabling these tools to work off a very clean, refined data.
330
00:36:06.385 -->
00:36:12.945And so now instead of co connecting AI directly to the bronze layer, which is what most people are doing,
331
00:36:13.185 -->
00:36:23.490you're now letting it play with clean data that has already gone through this. To put it also into simple terms is you can make it cheaper for you. It costs a lot of tokens
332
00:36:23.570 -->
00:36:25.970to load a gigantic dataset
333
00:36:26.370 -->
00:36:29.250and have it find insights for you.
334
00:36:29.730 -->
00:36:32.850Instead, you could give it a much more refined list
335
00:36:33.345 -->
00:36:39.265that is quite a bit smaller, already has gone through a rigor of everything that you know is just flat out wrong,
336
00:36:39.585 -->
00:36:52.190eliminated that, and then it can help you to take it to the next level for that. It's an exciting time with AI where I know businesses wanna build a lot of tools. I've helped companies with building AI chatbots.
337
00:36:52.190 -->
00:37:08.175You know, how do we guide our stakeholders to the right datasets into what lives out there? Well, it doesn't just know the AI agent didn't just wake up or get programmed, and now it knows where everything resides. You need to teach it and train it over time just like a standard human being.
338
00:37:08.655 -->
00:37:12.815So the purpose of this is by building it throughout these
339
00:37:13.055 -->
00:37:19.695levels with your data and architecting in a certain way, you can give it clean context. You can give it the business rules.
340
00:37:20.550 -->
00:37:23.350By forcing yourself to have clean documentation
341
00:37:23.590 -->
00:37:25.990at your gold layer and understanding
342
00:37:25.990 -->
00:37:31.990this is how we define this metric. This is how we count it. It's not gonna send you down a random rabbit hole.
343
00:37:32.310 -->
00:37:39.655It won't give you an insight that you think is gonna be really exciting to share with your VP just to realize that it overinflated
344
00:37:39.655 -->
00:37:43.175leads because it's not using the definition that your business uses.
345
00:37:44.135 -->
00:37:46.615And so as people are
346
00:37:47.015 -->
00:37:51.255going through the process of adopting this architecture,
347
00:37:51.335 -->
00:37:54.990they're either designing a new warehouse from scratch or
348
00:37:55.150 -->
00:38:04.589migrating an existing warehouse. What are some of the most interesting or innovative or unexpected ways that you've seen people use some of the principles from this design in actually
349
00:38:04.835 -->
00:38:05.714facilitating
350
00:38:05.714 -->
00:38:07.075their overall
351
00:38:07.075 -->
00:38:10.115objective of the organizational warehouse?
352
00:38:11.154 -->
00:38:15.954Yeah. Absolutely. So I've seen this apply in something as simple as
353
00:38:16.194 -->
00:38:18.035smaller consulting operations.
354
00:38:18.035 -->
00:38:19.875You know, that would to me was really interesting.
355
00:38:20.360 -->
00:38:22.920It went from somebody having to
356
00:38:23.400 -->
00:38:25.160here here's the business scenario.
357
00:38:25.240 -->
00:38:34.040They had about fifteen, twenty contractors on their team, all with independent time sheets, and and they needed to do a
358
00:38:33.485 -->
00:38:35.885payroll at the end of the week and pay everybody.
359
00:38:36.205 -->
00:38:39.565So you had a massive problem of data ingestion,
360
00:38:39.565 -->
00:38:40.925data unification,
361
00:38:41.325 -->
00:38:48.740cleansing, and then being able to take an operational step on it from there, which is typically what we're doing with our data anyway.
362
00:38:49.140 -->
00:38:51.220So using an ALA approach,
363
00:38:51.380 -->
00:38:54.020they were able to take something that repeatedly
364
00:38:54.020 -->
00:38:59.140was taking them hours every single week into simply running in a matter of minutes.
365
00:38:59.905 -->
00:39:06.945By taking this approach and breaking it out, they could load independent bronze layers for each of the independent time sheets,
366
00:39:07.185 -->
00:39:11.985unify it based on their master data that they wanted to view it as in their business,
367
00:39:12.225 -->
00:39:13.985and then be able to spit out an output.
368
00:39:14.589 -->
00:39:22.750I've seen this work incredibly well in very fascinating ways because it's able to take the data that exists,
369
00:39:23.150 -->
00:39:30.965apply a logical base to what your business needs are, and then help you get to that deliverable much quicker than you can imagine.
370
00:39:31.925 -->
00:39:37.045The great thing is it'll help you to even use more sophisticated tooling,
371
00:39:37.125 -->
00:39:39.125such as Python scripts or,
372
00:39:39.525 -->
00:39:47.750other reporting software even. They can help you elevate this and take it to the next level because you're feeding it already a cleaned object.
373
00:39:48.150 -->
00:39:50.710So by taking it through the rigor of ALA
374
00:39:50.710 -->
00:39:52.390where it surpasses
375
00:39:52.390 -->
00:39:54.070three different levels,
376
00:39:54.470 -->
00:40:00.665and then it's coming out to a consumable layer that is easily ingestible by any stakeholder in the business,
377
00:40:00.905 -->
00:40:03.865you can then operate on that and get to your completion.
378
00:40:03.945 -->
00:40:06.665So I've seen it used in multiple different
379
00:40:07.225 -->
00:40:08.105scenarios.
380
00:40:08.185 -->
00:40:09.785Mentioned the payroll.
381
00:40:10.250 -->
00:40:13.210I've also seen this applied to operations
382
00:40:13.450 -->
00:40:14.170where
383
00:40:14.250 -->
00:40:18.010this was relating back to a hyper growth company in Berlin.
384
00:40:18.410 -->
00:40:21.130They were using this to actually plan operations
385
00:40:21.210 -->
00:40:22.730for their solar panel installation.
386
00:40:23.335 -->
00:40:25.255So they were able to unify
387
00:40:25.335 -->
00:40:32.135their data across very many different partners and systems that they had to then plan out the operation
388
00:40:32.135 -->
00:40:36.455of which locations were they gonna go to in the upcoming week and,
389
00:40:36.775 -->
00:40:40.650actually install solar panels based on team availability,
390
00:40:40.650 -->
00:40:41.530proximity,
391
00:40:41.770 -->
00:40:42.730and scheduling.
392
00:40:43.130 -->
00:40:47.210So a lot of these you can think are very manual processes that take a long time.
393
00:40:48.010 -->
00:40:56.305But by tackling very specific prioritized data debt, for example, where are these teams located? Where are our customers located?
394
00:40:56.464 -->
00:41:00.945What does their schedule look like? If you can bring to light this data debt,
395
00:41:01.185 -->
00:41:03.425clean it up and ingest it properly,
396
00:41:03.585 -->
00:41:08.080then you can enable your business to fly very quickly with the data that you actually have.
397
00:41:08.720 -->
00:41:13.360And in your own experience of working in this space and developing
398
00:41:13.360 -->
00:41:18.320this overall design system, what are some of the most interesting
399
00:41:17.375 -->
00:41:22.415or challenging lessons that you've learned to the process of data warehouse design and implementation?
400
00:41:23.775 -->
00:41:30.895Yeah. So, first one is if the time to insight is already fifteen minutes, don't try to improve it beyond that.
401
00:41:31.214 -->
00:41:33.910That was a hard lesson that I learned.
402
00:41:34.150 -->
00:41:41.830It's if you already have a system that's working great for you, then leave it as is. The whole philosophy here is to help you to
403
00:41:42.230 -->
00:41:44.790get to the true insights that you need quickly.
404
00:41:45.605 -->
00:41:49.205And to easily define insight here, it's
405
00:41:49.525 -->
00:42:06.920you need to also trust that data. So it's not just the fact of I can give you the answer very quickly. If you don't believe it, it's it's not gonna be an insight. It's like that simple joke that I've heard where someone says, I'm very fast at math. And they say, okay, what's 14 plus 35? And they just say a 120.
406
00:42:07.240 -->
00:42:18.095It's like, that wasn't right. Yeah. But it was fast. So that's a great example of what an insight is not. And so the whole point here is to deliver delivering credible
407
00:42:18.095 -->
00:42:23.695data. So it has to also be understood at the same time. But, yeah, definitely,
408
00:42:23.695 -->
00:42:30.180I would say if time to insight is already under the fifteen minutes, it's just not needed in this sort of way.
409
00:42:30.660 -->
00:42:34.180Other particular challenges that I've seen along the way too
410
00:42:34.580 -->
00:42:35.300is
411
00:42:35.780 -->
00:42:41.700it's easy to get excited when there's a new system and wanting to tackle multiple data debt at the same time.
412
00:42:42.425 -->
00:42:45.625So my recommendation is always to stay on track,
413
00:42:45.945 -->
00:42:51.545and this can easily be prioritized based on what your objectives are for the next upcoming period.
414
00:42:52.105 -->
00:42:57.865And so you touched on this a little bit, but what are the cases where the agile ledger architecture is the wrong choice?
415
00:42:59.590 -->
00:43:04.950Yeah. Absolutely. So if you're in the early stages of still trying to figure out what you're trying to do as a business,
416
00:43:05.430 -->
00:43:11.510this is not for you. If you have a simple system where everything is done in just Shopify,
417
00:43:11.910 -->
00:43:13.430you're strictly running campaigns,
418
00:43:14.045 -->
00:43:27.165and you can get to the answers that you need to very quickly. This is a 100% not for you. Right? So to me, it's ultimately down to that litmus test that we talked about. It's the time to insight. If you're
419
00:43:27.590 -->
00:43:31.910can already get to what you need and it's quick enough for you, then this is not either.
420
00:43:32.310 -->
00:43:34.310Also, if we're talking about unstructured
421
00:43:34.310 -->
00:43:35.830exploratory dumps.
422
00:43:35.830 -->
00:43:38.950So if for purely exploratory unindexed
423
00:43:38.950 -->
00:43:39.430datasets,
424
00:43:39.815 -->
00:43:42.775like raw computer vision streams or server logs,
425
00:43:43.175 -->
00:43:47.415where business definitions don't really matter, then this is also not for you.
426
00:43:47.735 -->
00:43:49.495That takes a different approach,
427
00:43:49.655 -->
00:43:53.255and usually that data is not trying to be consumed
428
00:43:53.750 -->
00:43:56.390in order to generate business insights.
429
00:43:56.710 -->
00:44:07.750So if that's not the ultimate goal, then, agile ledger architecture is also not for you. And as you continue to invest time and effort into popularizing
430
00:44:07.750 -->
00:44:12.285this and flushing out the ideas? What are some of the ideas
431
00:44:12.285 -->
00:44:14.205that you have planned for
432
00:44:14.925 -->
00:44:16.925new and exploratory aspects
433
00:44:16.925 -->
00:44:17.565of
434
00:44:17.725 -->
00:44:21.405what the architecture can enable, how it can be extended,
435
00:44:21.485 -->
00:44:27.850some of the ways to maybe build tooling around its adoption or just some of the other plans you have going forward?
436
00:44:28.570 -->
00:44:42.835Yeah, absolutely. It's it's a very exciting time for it. So I really see it, lined into three pillars for the future of this. One is really the mentality shift. So it's helping data teams to transition from reactive ticket takers
437
00:44:43.075 -->
00:44:45.395into strategic business partners.
438
00:44:45.795 -->
00:44:51.555So helping them to step out of that daily query grind. You know, one of the most, demoralizing
439
00:44:51.555 -->
00:44:52.435tasks is
440
00:44:52.830 -->
00:44:58.590if somebody comes asking for leads and every time you need to run a twenty, thirty minute SQL query,
441
00:44:58.750 -->
00:44:59.630you know, that
442
00:44:59.870 -->
00:45:04.030can be very weighing on you as an analyst or a data person.
443
00:45:04.590 -->
00:45:05.310And so
444
00:45:05.585 -->
00:45:14.464enabling this sort of a structure to where you're already taking that logic and putting it into your consumption layer, as we mentioned, that gold layer tables,
445
00:45:14.944 -->
00:45:23.340you're essentially moving that process to being done one time. You're making it incredibly repeatable. You're making it visible, and you're ensuring adoption.
446
00:45:23.340 -->
00:45:25.100Everybody's doing it the same way.
447
00:45:25.420 -->
00:45:29.580Along with this mentality shift, it's really viewing data as a compass,
448
00:45:30.220 -->
00:45:34.785and it's reminding teams that data is a true compass. It's not ball.
449
00:45:35.025 -->
00:45:38.465You know, a good data compass shows you which direction not to go,
450
00:45:38.944 -->
00:45:46.065and it empowers leaders to execute high velocity decisions by telling us, well, that path definitely did not work. Let's not do that again.
451
00:45:47.420 -->
00:46:01.740Introduction of feedback loops to say this is where the data is pointing us to and then seeing if it's the right track. Let's go that way and see if it's actually helping our business in the right way or not. And if it's not, then we adjust our compass and adjust again.
452
00:46:02.140 -->
00:46:10.275So it's a bit more around mentality shifts and bringing to light that you don't always have a problem with your data.
453
00:46:10.595 -->
00:46:16.835It's more about the way that you've stopped short. You've made it strictly for an analyst, the environment,
454
00:46:17.270 -->
00:46:18.390and essentially
455
00:46:18.390 -->
00:46:19.750created a bottleneck.
456
00:46:19.910 -->
00:46:23.990We should take it a step further and enable the data for all stakeholders
457
00:46:23.990 -->
00:46:35.075within the company. Last pillar for this, for the future of the architecture is, open source and low code templates. So I'd love to get to the stage where I have prepackaged modeling templates already built,
458
00:46:35.555 -->
00:46:40.275maybe even DBT libraries that can help a team implement this ideology quickly,
459
00:46:40.675 -->
00:46:50.369set what their bronze would be, have it suggest some silver rules to get you a clean layer, and then spin up your very first gold tables to see how your nontechnical
460
00:46:50.369 -->
00:46:52.530stakeholders can self serve data.
461
00:46:53.089 -->
00:47:08.885Well, for anybody who wants to get in touch with you and follow along with the work that you're doing, I'll have you add your preferred contact information to the show notes. And as the final question, I'd like to get your perspective on what you see as being the biggest gap in the tooling or technology that's available for data and AI systems today.
462
00:47:10.210 -->
00:47:11.570Yeah. Absolutely.
463
00:47:11.570 -->
00:47:16.850So the gap between technical data engineering speed and operational business understanding,
464
00:47:17.250 -->
00:47:19.570we have incredible tools to stream
465
00:47:19.890 -->
00:47:20.850gigabytes,
466
00:47:20.850 -->
00:47:22.770petabytes of data in seconds.
467
00:47:23.170 -->
00:47:28.635But if data analysts are are not embedded in operational business meetings
468
00:47:28.714 -->
00:47:31.195and understanding how decisions are being made,
469
00:47:31.515 -->
00:47:33.435we just end up building these
470
00:47:33.595 -->
00:47:35.115fast monuments
471
00:47:35.194 -->
00:47:36.635to past priorities.
472
00:47:37.115 -->
00:47:44.010And I always had this joke because it seemed like every data team that I worked in, we always built this monument to past priorities.
473
00:47:44.250 -->
00:47:47.450By the time we finally got it built, the business had already moved on.
474
00:47:47.770 -->
00:47:50.570So all hail the monument to past priorities.
475
00:47:50.570 -->
00:47:51.290So,
476
00:47:51.850 -->
00:47:56.095it's exactly the biggest gap that I see. We spend so much of our time focused on
477
00:47:56.335 -->
00:48:08.750tooling and ensuring that we have the flashiest software to accomplish this, but you don't need it. I mean, you can get to the true insights using a Google Sheet. It just architecting your data in the proper way.
478
00:48:09.310 -->
00:48:11.870It's about convincing your stakeholders
479
00:48:11.950 -->
00:48:17.870by showing them how you're applying logic to the data in order to get it ready for them to consume.
480
00:48:18.350 -->
00:48:23.445And then trusting that the number that you're giving them is coming from the sources that is agreed upon.
481
00:48:23.765 -->
00:48:26.965If you can cross that bridge, I feel like you're definitely
482
00:48:26.965 -->
00:48:29.685reducing that gap between tooling and technology.
483
00:48:31.365 -->
00:48:53.095Alright. Well, thank you very much for taking the time today to join me and share the work that you're doing on the Agile Ledger architecture. It's definitely a very interesting approach. It's great to see that it is an addition to, not necessarily replacement of the warehouse that I've been building all along. So I appreciate the effort that you're putting into that, and I hope you enjoy the rest of your day. Absolutely. Thank you so much for your time. Appreciate it.
484
00:49:00.454 -->
00:49:04.670Thank you for listening, and don't forget to check out our other shows. Podcast.net
485
00:49:04.670 -->
00:49:13.869covers the Python language, its community, and the innovative ways it is being used. And the AI engineering podcast is your guide to the fast moving world of building AI systems.
486
00:49:14.430 -->
00:49:24.385Visit the site to subscribe to the show, sign up the mailing list, and read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email hosts@dataengineeringpodcast.com
487
00:49:24.385 -->
00:49:30.625with your story. Just to help other people find the show, please leave a review on Apple Podcasts and tell your friends and coworkers.