WEBVTT
NOTE
Transcription provided by Podhome.fm
Created: 7/6/2024
12:33:24 PM
Duration: 2779.506
Channels: 1
1
00:00:11.264 -->
00:00:15.445 Hello, and welcome to the Data Engineering podcast, the show about modern data management.
2
00:00:15.985 -->
00:00:25.100Your host is Tobias Macy, and today I'm interviewing Satish Jayanti about the practice and promise of building a column aware data architecture through intentional modeling. So, Satish,
3
00:00:25.480 -->
00:00:29.995for folks who haven't listened to your previous appearance on the show, if you can just start by introducing yourself.
4
00:00:30.455 -->
00:00:31.755 Yes. Thanks, Tobias,
5
00:00:32.295 -->
00:00:33.275for having me.
6
00:00:33.735 -->
00:00:38.390My name is, Satish Jayanti. I'm a a cofounder and CTO at Callas.
7
00:00:39.170 -->
00:00:41.670 And do you remember how you first got started working in data?
8
00:00:42.370 -->
00:00:43.910 Oh, yeah. Absolutely. So,
9
00:00:45.035 -->
00:00:46.095I've started
10
00:00:46.555 -->
00:00:47.295my career,
11
00:00:48.715 -->
00:00:50.575you know, in in software engineering,
12
00:00:51.515 -->
00:00:52.575application development.
13
00:00:53.260 -->
00:00:53.980And I was,
14
00:00:55.019 -->
00:00:58.320dabbled in that for a few years and was working for a
15
00:00:58.940 -->
00:01:01.825a start up in year 2000 in Los Angeles.
16
00:01:03.105 -->
00:01:03.605And,
17
00:01:04.865 -->
00:01:09.205this we we were selling online courses for a large financial firms.
18
00:01:10.600 -->
00:01:12.140You know, there was a lot of,
19
00:01:12.440 -->
00:01:14.200lot of boom at that time as you know.
20
00:01:15.080 -->
00:01:17.740You may have heard how, you know, the online,
21
00:01:18.725 -->
00:01:20.745Internet was taking off and everything was
22
00:01:21.125 -->
00:01:21.865very, very,
23
00:01:22.565 -->
00:01:25.545getting very hot. So at that time, I was
24
00:01:26.740 -->
00:01:28.120working there. And
25
00:01:29.140 -->
00:01:33.545a lot of times what happened was people would come to me and say, hey. Can you give me this report?
26
00:01:34.585 -->
00:01:35.244You know,
27
00:01:35.545 -->
00:01:39.805can can I have that report? Can I have this report? And I I I started kinda
28
00:01:40.185 -->
00:01:42.445taking those requests as they came in
29
00:01:42.760 -->
00:01:44.939and started building stuff and and
30
00:01:45.240 -->
00:01:48.220giving them. But at some point, I realized that
31
00:01:48.920 -->
00:01:50.060this was just overwhelming.
32
00:01:50.825 -->
00:01:56.525It was just not I was not able to keep up. And then that's when I started looking more and more into
33
00:01:57.225 -->
00:01:59.705how this is done usually. And,
34
00:02:00.200 -->
00:02:02.619you know, that that's where I kinda hit the,
35
00:02:03.159 -->
00:02:09.420the the topic of data warehousing, how to build it right, how the dimensional modeling and all of that. So I started
36
00:02:09.805 -->
00:02:13.505like that, accidentally getting into it, and, never looked back.
37
00:02:15.005 -->
00:02:15.825 And so
38
00:02:16.125 -->
00:02:18.160in terms of the overall space
39
00:02:18.560 -->
00:02:20.580of data warehousing, data platforms,
40
00:02:20.880 -->
00:02:22.820obviously, the cloud has been
41
00:02:23.360 -->
00:02:24.420the most impactful
42
00:02:24.800 -->
00:02:26.100change in the ecosystem
43
00:02:26.400 -->
00:02:30.015for probably the last decade, if not longer. And
44
00:02:31.195 -->
00:02:32.015in the process
45
00:02:32.315 -->
00:02:35.935of moving some of those data warehousing and data platform workloads,
46
00:02:37.050 -->
00:02:38.190large organizations
47
00:02:38.570 -->
00:02:39.230in particular
48
00:02:39.610 -->
00:02:43.150have taken a bit of a longer time to do that because they have
49
00:02:43.450 -->
00:02:45.370more existing infrastructure. They have,
50
00:02:45.885 -->
00:02:51.424bigger risk involved than actually making that migration. There are more moving pieces to coordinate around that.
51
00:02:51.885 -->
00:02:53.984And so a lot of the
52
00:02:54.560 -->
00:03:03.380patterns and practices around how to do data platforms and data warehousing in the cloud have been developed by smaller and faster moving organizations.
53
00:03:03.945 -->
00:03:21.130And I'm curious what you have seen as the overall impact of the cloud and some of that early adoption on the overall practice of data modeling and the data architectures that manifest in cloud environments as compared to some of the on prem environments that these larger organizations are still managing?
54
00:03:21.605 -->
00:03:22.985 Yeah. Absolutely. So
55
00:03:23.285 -->
00:03:26.345and and and, you know, I mean, you you said it right. So
56
00:03:27.045 -->
00:03:27.945smaller companies
57
00:03:28.325 -->
00:03:30.345moved to the cloud much quicker,
58
00:03:30.790 -->
00:03:32.010obviously, because,
59
00:03:32.550 -->
00:03:42.925they didn't perceive the the the risk or anything. I mean, they didn't they didn't have that much to begin with. They're getting started. They don't have large volumes of data. So they just quickly
60
00:03:43.705 -->
00:03:45.565took advantage of the cloud platforms
61
00:03:45.945 -->
00:03:49.005compared to larger firms. There, you know, there's so many
62
00:03:49.430 -->
00:04:10.360legacy systems running and security was a big thing. When cloud started, everybody every CTO, every executive was questioning the security. Like, how secure is it going to be once you put it on the on the cloud? So that was a big thing. So that took a while for a lot of companies to even kind of think about how, you know, how to address that. As far as the data modeling goes,
63
00:04:10.819 -->
00:04:13.480I think when these small companies or or smaller,
64
00:04:14.005 -->
00:04:14.825you know, smaller,
65
00:04:15.925 -->
00:04:17.545environments move to cloud,
66
00:04:18.165 -->
00:04:20.585things or most companies like startups,
67
00:04:21.610 -->
00:04:23.069which move pretty quickly.
68
00:04:23.690 -->
00:04:26.990And data modeling was seen as
69
00:04:27.690 -->
00:04:30.030a somewhat of a hurdle.
70
00:04:30.815 -->
00:04:31.715And and,
71
00:04:32.335 -->
00:04:35.075or in some cases, they may not even aware
72
00:04:35.535 -->
00:04:36.595of data modeling
73
00:04:37.055 -->
00:04:40.889because they haven't done that in the past. So they may not know that was
74
00:04:41.270 -->
00:04:42.569an important piece.
75
00:04:42.949 -->
00:04:49.955So for that reason, it just went on and people started building this really fast, and the data modeling took the back seat.
76
00:04:50.335 -->
00:04:52.995And it's just, for the last few years,
77
00:04:53.295 -->
00:04:57.770it's just disappeared from the analytics space. It's coming back now, but that's
78
00:04:58.229 -->
00:05:00.490that that's what, that's what happened
79
00:05:00.870 -->
00:05:01.530in in,
80
00:05:02.229 -->
00:05:04.009as these things move to cloud.
81
00:05:04.345 -->
00:05:05.565 And in particular,
82
00:05:05.945 -->
00:05:07.884with the data warehousing,
83
00:05:08.345 -->
00:05:09.884data modeling approach
84
00:05:10.185 -->
00:05:10.685of,
85
00:05:11.225 -->
00:05:13.884in particular, dimensional modeling that
86
00:05:14.300 -->
00:05:17.360was developed in the nineties in the era of
87
00:05:17.740 -->
00:05:23.425these large and expensive data warehouse appliances that were typically resource constrained. There were a lot of,
88
00:05:24.224 -->
00:05:24.724conflicting
89
00:05:25.185 -->
00:05:28.724priorities around what you needed to be able to answer and when.
90
00:05:29.025 -->
00:05:38.410I'm wondering how you have seen the ongoing conversation around the relative merits of dimensional modeling versus some of the newer paradigms of wide tables, etcetera,
91
00:05:39.104 -->
00:05:44.485and how that manifests and what the actual requirements are around modeling
92
00:05:44.785 -->
00:05:45.285priorities
93
00:05:45.664 -->
00:05:47.685in cloud data warehouse environments?
94
00:05:48.680 -->
00:05:49.419 Yeah. So
95
00:05:50.280 -->
00:05:52.539if you take the data modeling practice,
96
00:05:53.000 -->
00:06:00.794I mean, we should probably look at why people used to model in the first place. Like what were the benefits that people were getting out of modeling. Right?
97
00:06:01.175 -->
00:06:09.930I mean, there were 3 stages of modeling, generally speaking. A lot of people agree with this. The conceptual modeling, the logical modeling, and the physical modeling. Right?
98
00:06:10.550 -->
00:06:14.010So the conceptual modeling is primarily where you're saying,
99
00:06:14.354 -->
00:06:15.815hey. I need to understand
100
00:06:16.755 -->
00:06:17.735how the business,
101
00:06:18.595 -->
00:06:20.695or what the business needs. So
102
00:06:21.155 -->
00:06:23.975what what is what is that they want? What are the requirements?
103
00:06:24.820 -->
00:06:29.880So in a conceptual model, you start putting all the the entities, the business entities,
104
00:06:30.500 -->
00:06:31.480and you start
105
00:06:32.085 -->
00:06:36.105drawing lines between these entities to show relationships between them.
106
00:06:36.725 -->
00:06:42.650And then you use this conceptual modeling as a way to gather requirements as a communication tool
107
00:06:43.189 -->
00:06:44.650to work with the business.
108
00:06:45.509 -->
00:06:53.835It's much easier when you, you know, show something visually. You can sit in a room, talk to the business people and say, hey, is this what you're thinking? Is this how the relationships are?
109
00:06:54.375 -->
00:06:55.835That's a thinking exercise.
110
00:06:56.535 -->
00:07:11.435That's very important for the business and the people who are building the solutions to get on the same page. And then as you move towards to the logical model, it's more like the solution. You're talking about, okay, this is how I'm going to do it. It's still abstracted away from the physical implementation,
111
00:07:12.055 -->
00:07:12.955but it's still
112
00:07:13.414 -->
00:07:14.715a way to,
113
00:07:15.414 -->
00:07:17.275kind of draw out the solution.
114
00:07:18.080 -->
00:07:18.980And then finally,
115
00:07:20.160 -->
00:07:20.900the physical,
116
00:07:21.520 -->
00:07:25.220model. Physical models are where now they're closer,
117
00:07:25.685 -->
00:07:28.025and they're very they could be different
118
00:07:28.405 -->
00:07:31.705depending on the target platform. So if it's a cloud platform,
119
00:07:32.325 -->
00:07:35.305you may have some decisions that you
120
00:07:35.750 -->
00:07:38.330would not have done when you were doing an on prem
121
00:07:38.710 -->
00:07:41.930database platform. So as as you get to the physical aspect
122
00:07:42.230 -->
00:07:50.115of this, it's coming closer to the target platforms. So because Cloud platforms do perform well well, you can make some
123
00:07:51.169 -->
00:07:53.430decisions like, hey. I don't care about the storage.
124
00:07:53.889 -->
00:08:00.384You know? I I think the the system can handle the compute, so this is fine. So you're basically kinda
125
00:08:00.685 -->
00:08:01.504doing those,
126
00:08:02.125 -->
00:08:06.305kind of decisions as you move towards this. Now coming to the dimensional models.
127
00:08:06.820 -->
00:08:07.480Now dimensional
128
00:08:08.340 -->
00:08:11.720models are are the same. I mean, dimensional models are for analytics mostly,
129
00:08:12.180 -->
00:08:13.720and they are kind of
130
00:08:14.205 -->
00:08:14.705report
131
00:08:15.085 -->
00:08:17.105or analytics oriented models.
132
00:08:17.565 -->
00:08:21.585Makes it much much easier for the business to understand a model
133
00:08:21.910 -->
00:08:22.410without
134
00:08:22.870 -->
00:08:33.095getting overwhelmed with something like an operational model. Like, if you have an operational data model, there's a lot of tables, a lot of relationships, and just become very, very hard to even write queries.
135
00:08:33.475 -->
00:08:45.389But whereas if you show a domestic model and if you do it right, it becomes much much easier for for not only kinda start with requirements and understand, but also once you do it,
136
00:08:45.855 -->
00:08:48.435then business consumption, when they look at this,
137
00:08:49.214 -->
00:08:50.995they can kind of relatively
138
00:08:51.615 -->
00:08:54.115easily understand this these type of models.
139
00:08:55.200 -->
00:09:05.025 And as far as that dimensional modeling practice, just to expand that a bit more for folks who aren't familiar, this largely refers to the star and snowflake schema
140
00:09:05.405 -->
00:09:06.925versus the data vaults,
141
00:09:07.325 -->
00:09:12.449schema approach of the way that you structure the different tables and relationships within the warehouse
142
00:09:12.750 -->
00:09:17.329as opposed to what's typically 3rd normal form in transactional databases.
143
00:09:17.790 -->
00:09:19.810And as far as that actual
144
00:09:20.175 -->
00:09:20.675exercise
145
00:09:21.215 -->
00:09:21.715of
146
00:09:22.335 -->
00:09:27.635defining these dimensional models and figuring out what are the actual attributes and dimensions,
147
00:09:28.180 -->
00:09:34.760That can be a very complex and time consuming process. I'm wondering what are some of the ways that teams have
148
00:09:35.595 -->
00:09:40.815worked through that to be able to manage the, very rapid pace of
149
00:09:41.195 -->
00:09:41.695requirements
150
00:09:42.315 -->
00:09:43.695and the the changing requirements
151
00:09:44.839 -->
00:09:55.884with the upfront investment needed to be able to actually build out those models and cement them and, some of the ways the teams should be thinking about that balance of, I need to get some answers right now versus
152
00:09:56.264 -->
00:10:01.550this is going to help me in the long run. Yeah. That that's a very good point, and
153
00:10:01.850 -->
00:10:08.805 that's the challenge. Right? So people fall in these traps. And 1 trap is, hey, don't do modeling because I can do it fast.
154
00:10:09.425 -->
00:10:11.925Or do modeling and perfect the model
155
00:10:12.384 -->
00:10:13.685so I'll get it right.
156
00:10:14.330 -->
00:10:16.110Both are wrong, in my opinion,
157
00:10:16.970 -->
00:10:21.470because you got to get the right balance. You if you skip the modeling exercise,
158
00:10:21.985 -->
00:10:24.725regardless of whether we aren't doing our cloud or not,
159
00:10:25.105 -->
00:10:29.285then you're not taking the time to think. You're not taking the time to
160
00:10:29.825 -->
00:10:36.120build that foundation correctly. As opposed to saying, okay, I'm gonna spend all my time in modeling, then you're ignoring the business.
161
00:10:36.580 -->
00:10:40.920You know, you need to do your end goal is to provide high quality data,
162
00:10:41.855 -->
00:10:43.795that can be, you know, expanded,
163
00:10:44.495 -->
00:10:48.274upon more requirements. As you get more requirements, you need to have a way to scale this.
164
00:10:49.320 -->
00:10:50.220To do that,
165
00:10:50.760 -->
00:10:51.740you need to balance.
166
00:10:52.120 -->
00:10:55.660So so the I mean, what I have done in the past is
167
00:10:56.135 -->
00:10:57.115my goal was
168
00:10:57.575 -->
00:10:58.075to
169
00:10:58.615 -->
00:11:00.315always kinda take the data
170
00:11:00.775 -->
00:11:05.435and regardless of I mean, you don't have to model everything. You don't understand. You do that first iteration
171
00:11:06.009 -->
00:11:21.735of just get a feel for it. Do that concept to model or logical model, whatever that is. Don't try to perfect that. Just get the first iteration of it. Build something. It doesn't have to be purely dimensional model yet. It could be just views. And then put it out there as soon as possible.
172
00:11:22.275 -->
00:11:23.095And then
173
00:11:23.550 -->
00:11:25.970let the business kinda and set the expectation
174
00:11:26.510 -->
00:11:34.405that, hey, this is the first iteration. I want you to take a look at the data first and tell me, you know, I have these questions and do you see anything wrong?
175
00:11:34.945 -->
00:11:42.370So and then they'll look at the data. It could be just dashboards or whatever way that you want to, you know, give them the data. It could be sort of Excel.
176
00:11:42.829 -->
00:11:47.970They can download into that. But they can look at the data quickly and then give that feedback.
177
00:11:48.695 -->
00:11:51.115Then you start thinking about, oh, now I understand
178
00:11:51.495 -->
00:11:55.115the relationships. Now I can kinda improve my model
179
00:11:55.709 -->
00:12:22.875and say, hey. This is how this needs to be. And then kinda follow those iterations as you go through. It will never be perfect, and that's okay because everything is changing. Now even though it may seem perfect today, it's not going to be as soon as the company goes and buys another company. You know, everything is like the rules are changed. Now everything has changed. So that's okay. But but understanding that that balance is important and that as long as you're trying to improve it,
180
00:12:23.495 -->
00:12:23.995continuously,
181
00:12:24.774 -->
00:12:25.675you'll be successful.
182
00:12:26.375 -->
00:12:26.875 Another
183
00:12:27.334 -->
00:12:27.834outgrowth
184
00:12:28.350 -->
00:12:30.930of the cloud and cloud warehouses
185
00:12:31.310 -->
00:12:37.810and the decreasing storage costs and the decoupling of compute and storage is the shift in the
186
00:12:38.415 -->
00:12:41.635integration path from ETL, which was very,
187
00:12:42.735 -->
00:13:08.480upfront investment. You had to make sure that you got things right because, otherwise, you're going to lose information or you're going to have the wrong assumptions and do the transformations wrong, and you had no way of going back and doing it right afterwards. Or if you did, it required a lot of extra engineering to give you that fallback versus the ELT paradigm that has become prevalent of just load all of the data and then transform it so so that you don't have to think about modeling up front, which I think has had this,
188
00:13:09.740 -->
00:13:24.259downstream impact of saying, well, if I don't have to do all the modeling on load, well, I don't really have to do any modeling at all. Because if I do do it wrong, I can just rewrite it and rebuild everything, and it's fine. And I'm curious what you have seen as the,
189
00:13:24.880 -->
00:13:34.555relative merits of the upfront investment requirements of ETL versus the very fast moving iterative capability of ELT and some of the ways that that influences
190
00:13:34.935 -->
00:13:43.819the overall thinking and patterns within a team or an organization as far as that data modeling, data integration, data analytics workflow?
191
00:13:44.279 -->
00:13:46.540 Yeah. Absolutely. So, I mean, e ELT,
192
00:13:47.240 -->
00:13:51.755definitely is has changed the game in in terms of how you build these things.
193
00:13:52.295 -->
00:13:55.834In in the in the old days, you know, when you had ETL systems,
194
00:13:56.590 -->
00:13:59.410they became the bottleneck. They are very rigid tools.
195
00:13:59.870 -->
00:14:01.890The GUI tools mostly, but,
196
00:14:02.270 -->
00:14:09.495they brought the efficiencies, but also rigid in terms of, you know, you didn't when you hit a corner case, you didn't have much control on it.
197
00:14:10.355 -->
00:14:12.134And and also you were not leveraging
198
00:14:12.760 -->
00:14:21.500the investment that you made in the target platforms. Right? The database platforms. But ELT paradigm has changed that. Now you have this full power,
199
00:14:22.095 -->
00:14:23.235that you can utilize,
200
00:14:23.935 -->
00:14:26.995on the the using the database platform to do the transformations.
201
00:14:27.615 -->
00:14:30.755So what that means is, yeah, you can you can iterate
202
00:14:31.209 -->
00:14:32.190through much much,
203
00:14:32.890 -->
00:14:33.470you know,
204
00:14:34.250 -->
00:14:34.750rapidly,
205
00:14:35.370 -->
00:14:36.670as opposed to the ETL,
206
00:14:37.529 -->
00:14:39.390ETL paradigm. So,
207
00:14:39.745 -->
00:14:43.205yeah, as I said before, I think it's it's much more
208
00:14:43.585 -->
00:14:44.085easier
209
00:14:44.465 -->
00:14:46.005to do those kind of iterations,
210
00:14:47.800 -->
00:14:48.699doing ELT
211
00:14:49.320 -->
00:14:53.420on a platform like for example, Snowflake or some kind of cloud platforms.
212
00:14:53.880 -->
00:15:00.545 Another aspect of this question of data modeling, particularly with dimensional models where it does make the
213
00:15:01.085 -->
00:15:02.545understandability and discoverability
214
00:15:03.085 -->
00:15:03.520of
215
00:15:03.920 -->
00:15:06.020the core domain objects clearer
216
00:15:06.640 -->
00:15:17.745than if you have these 1 off models that answer a specific question. But if you wanna do any further exploration, then you have to go back to from the source and build back up to it. Is this question of
217
00:15:18.285 -->
00:15:19.185data self-service
218
00:15:19.649 -->
00:15:24.870for people in the organization, people in the business suite who don't necessarily
219
00:15:25.329 -->
00:15:30.905have the capacity or time to dig deep into the data and understand it semantically.
220
00:15:31.765 -->
00:15:35.145And I'm curious what you have seen for the teams who
221
00:15:35.685 -->
00:15:44.360fall into that pit of, I just need to build this 1 model to answer this 1 specific question versus investing in that upfront data modeling or dimensional modeling,
222
00:15:44.845 -->
00:15:46.545who is driving the actual
223
00:15:47.005 -->
00:15:49.185effort of doing the analysis,
224
00:15:49.565 -->
00:16:05.875and how does that how how does that manifest in terms of the ways that the teams think about what their priorities are of doing the upfront modeling versus just answer questions really quick? Classic challenge. Right? I mean, you know, the business units are saying, I know my data. I understand my data.
225
00:16:06.575 -->
00:16:10.595 Even if you give me a model, I'm gonna look at this model for this particular
226
00:16:11.135 -->
00:16:20.790process. And I'm gonna just build it for this 1 and and I'm happy. But the problem with that is there's not just 1 business process in a company. There's multiple business processes,
227
00:16:21.165 -->
00:16:24.065And these business processes share a lot of data.
228
00:16:24.525 -->
00:16:25.025So
229
00:16:25.565 -->
00:16:27.505you have to have a way to
230
00:16:27.965 -->
00:16:29.265kind of plan this
231
00:16:29.750 -->
00:16:32.490from that higher level. Like, to understand,
232
00:16:32.870 -->
00:16:34.490hey, what are my business processes?
233
00:16:35.030 -->
00:16:36.810What is the shared data?
234
00:16:37.334 -->
00:16:52.500What become dimensions? What become facts? And how are these dimensions linked? This is all very, very important to make sure that you're not building silos. Because just because you build a dimensional model doesn't mean that you're doing it right. You know, you need to make sure that
235
00:16:52.880 -->
00:16:53.699you are,
236
00:16:54.639 -->
00:16:55.139building
237
00:16:55.774 -->
00:17:05.315with the intention to share data across these business processes and understand those. And there's a term Kimbal mentions, you know, called conformed dimensions.
238
00:17:05.760 -->
00:17:06.740That was the idea.
239
00:17:07.360 -->
00:17:27.660Now we don't wanna get into how we do that and what are the pros and cons and all of that. But generally speaking, the idea is, hey, there's a shared dataset. You build that dataset for 1 process, for 1 unit, and then make sure that you leverage and use that same thing to connect other, you know, other datasets. So that way you kind of build in this integration.
240
00:17:28.680 -->
00:17:29.180 And
241
00:17:29.480 -->
00:17:31.340as far as the overall
242
00:17:31.800 -->
00:17:32.300ecosystem
243
00:17:32.760 -->
00:17:34.460of data modeling practices,
244
00:17:35.305 -->
00:17:35.805Obviously,
245
00:17:36.345 -->
00:17:44.365star and snowflake schema and data vault have been around possibly the longest, or they've been the longest lived. But I'm wondering if there are any other
246
00:17:45.000 -->
00:18:02.655patterns or data modeling approaches that have been developed in recent years that you have found to be useful or beneficial or things that are maybe built off of some of the star and snowflake or Datavault approaches that make it easier to iterate into those dimensional approaches?
247
00:18:03.330 -->
00:18:04.070 Yeah. So,
248
00:18:04.850 -->
00:18:06.950I mean, to be honest, I haven't seen anything
249
00:18:07.650 -->
00:18:13.235big like that that, you know, when dimension models came out and data wall came out.
250
00:18:13.855 -->
00:18:22.740Of course, there was the enterprise data model from Inman. So these were the major kind of data modeling, you know, data modeling methodologies.
251
00:18:23.120 -->
00:18:29.225Yeah. As far as I mean, there's there's time series modeling. There's like all these subsets of modeling that that you can talk about, but
252
00:18:29.605 -->
00:18:44.270nothing major in terms of revamping the whole thing. So so if you're talking about data warehouses, you have a few options. You know, you're saying, okay, I want to build my, you know, use Data Vault. If you're if you're seeing that there's a need to build a Data Vault for your business,
253
00:18:44.705 -->
00:18:51.205then that would be a good approach to build that raw data, organize that. And then on top of that, you build your star schemas.
254
00:18:51.640 -->
00:19:02.815You're still building the dimensional models, but you're building on top of the the raw data vault models. And there are situations where you may say, I I we don't do Data Vault link in Data Vault.
255
00:19:03.435 -->
00:19:05.135We can just go straight into,
256
00:19:05.515 -->
00:19:12.120dimension models. There's pros and cons in each 1 of those. But on top of this, there are things that you can do custom
257
00:19:12.420 -->
00:19:16.840kind of modeling, I would say. You know, you can create aggregate tables if you think there's,
258
00:19:17.315 -->
00:19:27.270you know, there is a performance thing. And your data is, like, huge and you want to aggregate it. But make sure you aggregate it after you bring the granular data into your,
259
00:19:27.810 -->
00:19:42.365into your air warehouse. Right? Because if you if you if you capture the data from the operational systems and you aggregate it right away, then if somebody's asking a deeper question to, hey. Show me how you got this aggregated number, then you don't have a way
260
00:19:42.799 -->
00:19:52.205to drill into it because now you have to go to the operational system, which is probably already update. So yeah. I think there's not not I have not seen, like, a big big,
261
00:19:52.605 -->
00:19:54.945model methodology since these 2.
262
00:19:57.725 -->
00:20:11.384 RudderStack helps you build a customer data platform on your warehouse or data lake. Instead of trapping data in a black box, they enable you to easily collect customer data from the entire stack and build an identity graph on your warehouse, giving you full visibility and control.
263
00:20:11.764 -->
00:20:22.960Their SDKs make event streaming from any app or website easy, and their extensive library of integrations enable you to automatically send data to hundreds of downstream tools. Sign up for free today at dataengineeringpodcast.com/rudderstack.
264
00:20:27.425 -->
00:20:38.300Another element of the conversation today is the question of column awareness in the space of data modeling and data transformation. And I'm curious if you can just start by describing
265
00:20:38.760 -->
00:20:46.755what you mean by that term and some of the ways that that will impact the approach of data modeling or the ways that data transformations
266
00:20:47.295 -->
00:20:51.010and table structures are built up? Yeah. So
267
00:20:51.470 -->
00:20:52.690 if you look at the
268
00:20:52.990 -->
00:20:54.690data, the I think columns
269
00:20:54.990 -->
00:21:04.695columns are the the building blocks pretty much. That's the lowest grain that you can go to in a in a dataset. Right? I mean, if you if you think look at an Excel sheet, each cell is a,
270
00:21:05.200 -->
00:21:15.605you know, is is the data value in a particular column. When we say column aware, it's basically column metadata. That's what we're talking about here. But how we can leverage that,
271
00:21:15.905 -->
00:21:20.725that's the distinction here. You know, how we can use it to its maximum potential
272
00:21:21.540 -->
00:21:22.840to do automation
273
00:21:23.540 -->
00:21:26.520is what we are talking about. By storing the column
274
00:21:26.980 -->
00:21:29.560metadata or being column aware
275
00:21:30.025 -->
00:21:33.165and using that leveraging that from the ground up
276
00:21:33.945 -->
00:21:35.805to do whether it's modeling,
277
00:21:36.265 -->
00:21:37.805whether you're generating code,
278
00:21:38.270 -->
00:21:40.290whether you're doing column lineage,
279
00:21:41.150 -->
00:21:42.930whether you're doing impact analysis,
280
00:21:43.390 -->
00:21:51.975whether you're determining the state of your target database platform. You get all of those benefits just by leveraging that column awareness.
281
00:21:53.075 -->
00:21:57.015And and and again, that's what we we did at at Colas.
282
00:21:57.539 -->
00:21:58.039So
283
00:21:58.419 -->
00:22:01.240if you if you look at anything that you're sharing with the business,
284
00:22:01.780 -->
00:22:04.200you know, whether it's a KPI or metric,
285
00:22:04.945 -->
00:22:07.045A KPI is made out of metrics.
286
00:22:07.585 -->
00:22:11.205Like KPI is a key performance indicator, you know, it's made out of metrics.
287
00:22:11.505 -->
00:22:15.010And if a metric is usually made out of 1 or more columns,
288
00:22:15.710 -->
00:22:19.330and these columns are coming from tables. So it's always comes down to columns.
289
00:22:19.870 -->
00:22:22.290You know, you're applying transformations on columns.
290
00:22:23.065 -->
00:22:32.340You're adding new columns. You're deleting new columns. You're working with columns. Data quality is about columns. So it just becomes so crucial
291
00:22:33.039 -->
00:22:34.740to understand the relationships
292
00:22:36.000 -->
00:22:43.815and the dependencies between these columns. And, of course, columns in tables, in tables search collection of columns. Right? So what what
293
00:22:44.115 -->
00:22:46.455I we think is that that's the best way
294
00:22:46.870 -->
00:22:47.370to
295
00:22:47.830 -->
00:22:48.330kinda
296
00:22:48.870 -->
00:22:51.370build an analytics system so you can scale.
297
00:22:51.670 -->
00:22:53.290Because when you're small,
298
00:22:53.725 -->
00:23:29.040it doesn't matter as much. When you're meaning when you're building a small, system or a small database, small model, it doesn't matter much. Because you were thinking, I can write a query. I can get this done pretty fast. Fine. But usually, that's not the intention. Like, whenever you're working for a business you're building, it it may be small at that time, but the whole intention is to grow. You know, they'll go and get more datasets. You're selling more. You're acquiring other companies. Whatever that's but you have to have a way to scale. So that's where this comes into play. The automation comes into play. And to do proper automation, you need to be column aware. So many different things to dig into there.
299
00:23:29.884 -->
00:23:54.195 So from the perspective of a table is just a collection of columns, and the columns are the thing that you actually care about, I'm wondering if you can talk to some of the more nuanced aspects of that where, yes, a table is a collection of columns, but the columns in within a table typically all relate to each other in some fashion. So there is some contextual awareness that you need to have to be able to operate properly on that column. And I'm wondering
300
00:23:54.780 -->
00:24:16.200what are some of the pitfalls that somebody might fall into if they say, I only really care about the columns, so this is the thing I'm going to track. What what are the pieces of metadata that they might drop on the floor by the fact that they are maybe ignoring some of the other, contextual and semantic aspects of the table that that column is contained within and the ways that that column is propagated to other differently named tables or collections of columns, etcetera?
301
00:24:17.305 -->
00:24:18.905 Very good question. So so,
302
00:24:19.545 -->
00:24:20.845you know, when I say columns,
303
00:24:21.305 -->
00:24:24.685tables, a collection of columns, you know, it's true. However,
304
00:24:25.250 -->
00:24:38.135there are dependencies. There are relationships between those columns. There's a primary key. There's a foreign key. There's there's all these other relationships that are very, very important. And that's where I guess the modeling comes into play. But there is additional metadata
305
00:24:38.515 -->
00:24:46.380we store on these columns. You know, it's about not just the name of the data type, but there is, like, you can attach now. You can attach
306
00:24:46.760 -->
00:24:48.060or you can tag things
307
00:24:48.575 -->
00:24:51.235as foreign keys, as business keys,
308
00:24:51.615 -->
00:24:54.755as change type 2 columns in in a dimension.
309
00:24:55.055 -->
00:24:56.035By doing those,
310
00:24:56.370 -->
00:25:11.815now we can kinda leverage those additional metadata tags to do even more automation. We can generate scripts that would say, hey, if I know the business key, if I know the change tracking column, if I know the foreign key, I can just, you know, write this code or generate this code automatically.
311
00:25:12.440 -->
00:25:13.340Same thing with,
312
00:25:13.960 -->
00:25:16.700data quality rules. I mean, I can take a domain,
313
00:25:17.160 -->
00:25:21.260like a financial domain, and I can tag these columns as
314
00:25:21.684 -->
00:25:34.970certain attributes. Like for example, a cussip in a in a financial world which stands for for stock. You know, you you can do those things and apply automation on data quality rules. So it's not just the columns, but also taking into consideration
315
00:25:35.590 -->
00:25:37.690those relationships with the columns.
316
00:25:38.150 -->
00:25:39.610What is this column
317
00:25:40.335 -->
00:25:46.755do in a particular table, is it the business key, is it a composite key, more than 1 column. Right? So all those are very, very,
318
00:25:47.375 -->
00:25:52.460important. I wanna make sure that people when they do modeling, that those are some of the key,
319
00:25:52.760 -->
00:25:54.940things that you need to start considering.
320
00:25:55.545 -->
00:26:00.605 And for people who are building their transformations, they're doing their modeling,
321
00:26:00.985 -->
00:26:04.685what are some of the elements of a given tool or workflow
322
00:26:05.690 -->
00:26:06.430that are necessary
323
00:26:06.890 -->
00:26:09.870for column awareness to be an effective component
324
00:26:10.250 -->
00:26:13.150of that overall effort? So what are the capabilities
325
00:26:13.450 -->
00:26:18.965in a tool that need be present for it to be column aware? What are some of the ways that that column awareness
326
00:26:19.424 -->
00:26:22.325will be incorporated into the
327
00:26:22.880 -->
00:26:39.475transformation and workflow of saying, okay. I've got this information in this column. I need to combine it with this column from this other table to merge it into this new attribute or this new metric that is then being propagated into this new table, which is a collection of columns and some of the ways that lineage,
328
00:26:40.090 -->
00:26:46.430at the column level needs to be present to inform the transformations and the modeling efforts, etcetera.
329
00:26:47.025 -->
00:26:49.605 Yeah. So so being being column aware,
330
00:26:49.985 -->
00:26:51.125as I said before,
331
00:26:51.585 -->
00:26:53.605it's not just storing column metadata
332
00:26:54.210 -->
00:26:56.710or not just creating a lineage diagram.
333
00:26:57.090 -->
00:27:00.390Those are just are the byproducts of being column aware.
334
00:27:00.930 -->
00:27:07.715We have to leverage it from the ground up on everything that you do in the tool. So whether you're generating
335
00:27:08.414 -->
00:27:08.914code
336
00:27:09.215 -->
00:27:11.235or you're generating impact analysis
337
00:27:11.740 -->
00:27:12.480or you're
338
00:27:12.940 -->
00:27:14.080generating, lineage,
339
00:27:14.380 -->
00:27:20.320or even when you're modeling, you're selecting business keys. You you're drawing foreign key relationships.
340
00:27:20.845 -->
00:27:25.825You're applying data quality rules. So all of these should be part of your
341
00:27:26.445 -->
00:27:31.570of the tool in a way that you can easily interact with the with the metadata
342
00:27:32.429 -->
00:27:35.250and kinda enhance and enrich the metadata.
343
00:27:35.710 -->
00:27:42.875So so you have these columns. And now as a user, you understand what these columns are and what they are supposed to do,
344
00:27:43.335 -->
00:27:49.570and what their function is and what do they mean. So you can enrich that that repository, that metadata repository
345
00:27:50.110 -->
00:27:51.890by adding additional metadata.
346
00:27:52.510 -->
00:27:58.955And as you do that and how the tool would take that column metadata plus the enriched metadata
347
00:27:59.414 -->
00:28:00.955that that the user input
348
00:28:01.255 -->
00:28:01.995and generate,
349
00:28:03.015 -->
00:28:06.460these other artifacts that I was talking about, like scripts
350
00:28:07.240 -->
00:28:08.460or models or
351
00:28:08.840 -->
00:28:15.404 linear diagrams or and so on. And so from that, another question is, how much of the column awareness
352
00:28:16.505 -->
00:28:22.764is intended for the actual consumer of the tool of somebody who's building the the transformations
353
00:28:23.889 -->
00:28:31.269versus the column awareness being a function of the tool itself and is something that is used to drive that automation
354
00:28:32.054 -->
00:28:35.274without the need for as much direct human involvement?
355
00:28:35.654 -->
00:28:37.274 Yeah. Very good question. So
356
00:28:39.080 -->
00:28:40.860we don't wanna burden the user
357
00:28:41.320 -->
00:28:42.300by making them,
358
00:28:43.000 -->
00:28:45.020understand or or or having them
359
00:28:45.320 -->
00:28:54.165forcing them to use the tool in a certain way because we are calling away. You know, they it should be an easy to use tool, you know, whether you're a GUI person or you're writing
360
00:28:54.945 -->
00:28:56.325some code, some template.
361
00:28:56.630 -->
00:29:01.290It should be very easy to use and you don't have to think about the column awareness. You're just given.
362
00:29:01.590 -->
00:29:09.915But behind the scenes, because we are column aware, we can do a lot of automation for you. So it's mostly how the tool functions,
363
00:29:10.535 -->
00:29:11.595but there is some
364
00:29:12.055 -->
00:29:16.235some things that kinda manifest into the UI. For example,
365
00:29:16.680 -->
00:29:27.375because we know every column, we can show you a list of columns that you can select which columns are business keys. So that that just becomes easy for you. Because because we are column aware, we can show, hey, here are the columns.
366
00:29:27.675 -->
00:29:35.150Pick the business key, 1 or 2 is the composite key, and also pick what columns are type 2 columns, for example.
367
00:29:35.610 -->
00:29:48.845And the user can just kind of just go and pick rather than remembering and typing and putting it in a YAML file or whatever. So it's a it's a combination, but, mostly, it's a function of the tool behind the scenes. In terms of the overall tooling ecosystem,
368
00:29:49.880 -->
00:30:05.245 the default mode for a while, at least, has been to operate on the table level or even just on the task level where there is potentially no awareness at all of whether or not there are tables or columns or files being manipulated.
369
00:30:06.105 -->
00:30:06.605And
370
00:30:07.200 -->
00:30:14.020as we have gone through successive generations of transformation tools, underlying compute platforms,
371
00:30:14.615 -->
00:30:27.380We have gotten to a more granular and nuanced level of the tools having awareness of what are the underlying objects or resources that are being operated on. And I'm wondering what are some of the complexities
372
00:30:27.680 -->
00:30:29.860of managing this column awareness
373
00:30:30.480 -->
00:30:31.380that have,
374
00:30:32.025 -->
00:30:36.685taken so long for us to get to where maybe the table was the previous level of
375
00:30:37.145 -->
00:30:41.860understanding within the tooling and, some of the ways that that poses complexity
376
00:30:42.240 -->
00:30:50.045in the tool development to be able to expose these capabilities and this semantic understanding of what is being done? So we we,
377
00:30:50.845 -->
00:30:51.345 we
378
00:30:51.805 -->
00:30:58.545extract this metadata as upon discovery. Right? I mean, as soon as you connect to something, we try to kinda extract this metadata
379
00:30:59.090 -->
00:31:07.990and and save that metadata. And then you're basing off of that metadata. You're you're you're you're gonna use that as a starting point and build additional objects
380
00:31:08.404 -->
00:31:09.544and additional workflows.
381
00:31:09.845 -->
00:31:15.784As far as, you know, there there could be down the road, there could be things where we may not be able to extract this metadata
382
00:31:16.164 -->
00:31:16.825out of
383
00:31:17.320 -->
00:31:19.419out of some I don't know, like a video
384
00:31:20.600 -->
00:31:25.820message. Maybe maybe we don't need to. Or or maybe it's possible because the AI is,
385
00:31:26.205 -->
00:31:32.545you know, coming on and I don't want to go take this discussion in that direction, but I'm just saying things are changing. We understand.
386
00:31:32.845 -->
00:31:33.585But regardless,
387
00:31:35.070 -->
00:31:35.890the importance
388
00:31:36.590 -->
00:31:38.049of having this column
389
00:31:38.350 -->
00:31:39.250column awareness,
390
00:31:41.549 -->
00:31:43.090it's not gonna go away.
391
00:31:44.135 -->
00:31:44.955But, you know,
392
00:31:46.054 -->
00:31:46.554whatever
393
00:31:47.414 -->
00:31:53.035in the near future, I can't see where we say, hey. It's too complex, so let's not do column level.
394
00:31:53.510 -->
00:32:02.010Just go back to the table level or or or it's just like summarizing the data. You know, once you summarize, you lose the grand narrative. That's that's exactly what what would happen.
395
00:32:02.855 -->
00:32:05.755Is it difficult for the tools to build this?
396
00:32:06.295 -->
00:32:08.955I think we should do as as much as we can
397
00:32:09.335 -->
00:32:11.275as vendors to do this.
398
00:32:11.639 -->
00:32:17.5801 key thing though, what I've seen in the, how this was done in the past versus how we're doing now is
399
00:32:17.960 -->
00:32:20.299column metadata was a afterthought.
400
00:32:20.715 -->
00:32:22.735And lineage diagrams were an afterthought.
401
00:32:23.435 -->
00:32:29.455They were like, okay. After you do everything, whatever you're doing, then we're gonna run something and scan it and do this.
402
00:32:29.899 -->
00:32:36.640Well, that's not how we think it should be done. We think the calm awareness should be the building block
403
00:32:37.055 -->
00:32:40.275for for everything. It should be the starting point for everything,
404
00:32:40.895 -->
00:32:43.715and that's the fundamental difference. Yeah. It's challenging
405
00:32:44.175 -->
00:32:47.040problem, but, you know, we have taken upon ourselves
406
00:32:47.820 -->
00:32:53.200 to solve this. Yeah. That was another aspect that I wanted to discuss is this question of
407
00:32:53.605 -->
00:33:00.664column awareness and lineage tracking within the transformation tool and within the workflow of actually manipulating the data
408
00:33:01.200 -->
00:33:06.020versus bolting it on or discovering it after the fact with some of these
409
00:33:06.400 -->
00:33:07.620metadata platforms
410
00:33:08.215 -->
00:33:12.795and the lineage analysis from the SQL logs in the data warehouse
411
00:33:13.495 -->
00:33:17.840and some of the ways that pushing that information into the transformation and data manipulation
412
00:33:18.700 -->
00:33:21.040layer impacts the effectiveness
413
00:33:21.420 -->
00:33:28.725and utility of these metadata platforms, whether it serves as a means of improving them or just serves as a means of obviating them entirely.
414
00:33:29.105 -->
00:33:31.345 So so you're you're you're basically asking if,
415
00:33:32.310 -->
00:33:35.050the way that the current tools, like, after
416
00:33:35.430 -->
00:33:39.130you build something, the way they extract those, is that useful?
417
00:33:39.845 -->
00:33:40.825 Yeah. So, basically,
418
00:33:41.605 -->
00:33:48.590given the fact that metadata platforms have become popular because of the need for being able to expose this lineage information, etcetera.
419
00:33:48.970 -->
00:33:50.510Does having that lineage
420
00:33:50.810 -->
00:33:59.115information and column level awareness within the transformation tool serve to improve those metadata platforms, or does it serve to help make them obsolete?
421
00:33:59.415 -->
00:34:03.755 Yeah. I I yeah. I think it's the former. So it's gonna help improve.
422
00:34:04.215 -->
00:34:07.750Or the the metadata yeah. Yeah. Definitely improve. Because
423
00:34:08.050 -->
00:34:11.750the reason for that is I think when you do it as an
424
00:34:12.175 -->
00:34:16.515exercise at the end of, you know, using a metadata tool to extract everything,
425
00:34:17.135 -->
00:34:20.435they do a pretty solid job. But however, they won't be 100%.
426
00:34:21.030 -->
00:34:31.505You know, they they're going to have some, you know, accuracy issues. But when you do it from the ground up like us, because we are using this as a basis to generate the code.
427
00:34:31.885 -->
00:34:41.930So our lineage is going to be, for example, just since we're talking about lineage. There's more than just the lineage. I wanna make sure that, you know, the common awareness is not lineage is just a byproduct.
428
00:34:42.390 -->
00:34:43.610But to answer your question,
429
00:34:44.070 -->
00:34:52.375we would actually improve the metadata extraction that happens with these tools. So it's gonna be more accurate. And another point here is
430
00:34:52.755 -->
00:34:56.710it also of course, the whole point of this thing is it it improves the productivity
431
00:34:57.490 -->
00:35:01.430of building this in the 1st place. And and the metadata tools
432
00:35:02.130 -->
00:35:04.309would would would get more accurate metadata.
433
00:35:04.744 -->
00:35:13.869 Yeah. It's interesting too because of the recent efforts of some of the metadata platforms to use their contextual awareness to feedback into the
434
00:35:14.170 -->
00:35:26.405transformation process, some of the data platform automation. So figuring out when do you want to schedule your different transformations to happen so that more of them can happen within this given time frame for, you know, cost management, etcetera.
435
00:35:26.730 -->
00:35:34.510So definitely interesting to see how that plays out. Another avenue that I like to explore in this space of column awareness in the transformation
436
00:35:34.855 -->
00:35:58.715layer is some of the ways that that manifests outside of the context of the data warehouse where within the warehouse, it's very clear what the benefits are because everything has happened within that bounded context of everything is a column, and it gets moved into another column. And so there there is this similarity in terms of the operations being performed. But as you start to leave the context of the warehouse into something like a machine learning workflow
437
00:35:59.390 -->
00:36:00.770or a streaming system,
438
00:36:01.150 -->
00:36:07.329or as the data leaves the warehouse and feeds back into maybe a SaaS platform through things such as reverse ETL,
439
00:36:07.765 -->
00:36:09.945How does the column awareness
440
00:36:10.484 -->
00:36:11.785help in those
441
00:36:12.085 -->
00:36:13.065external contexts?
442
00:36:13.845 -->
00:36:18.810 Yeah. So I don't have a particular use case in mind, but I'm I'm pretty sure that
443
00:36:19.270 -->
00:36:23.290just like how we are saying we are going to improve the metadata
444
00:36:23.695 -->
00:36:25.555extraction process to improve the accuracy,
445
00:36:26.335 -->
00:36:29.635the same thing can be applied to other outside processes.
446
00:36:30.335 -->
00:36:32.994Because we have this level of information
447
00:36:33.780 -->
00:36:41.480and the quality of output that we are giving, that basically helps improve any external process that depends on on this.
448
00:36:42.055 -->
00:36:44.155Whether it's reverse ETL or
449
00:36:44.615 -->
00:36:45.115some
450
00:36:45.494 -->
00:36:49.760visualization tools pulling data, you know, it's all connected,
451
00:36:50.060 -->
00:36:50.560interconnected.
452
00:36:50.940 -->
00:36:53.360I just don't have a use case on top of my mind.
453
00:36:53.980 -->
00:36:58.000But but definitely, I see that it's it's it can only benefit these systems.
454
00:36:58.475 -->
00:37:03.535 And in terms of your work at Coalesce and helping to
455
00:37:04.075 -->
00:37:06.400think through the ways that column awareness
456
00:37:06.799 -->
00:37:28.080manifest within your tool and just as a general practice of column awareness being this core capability of the transformation layer. What are some of the most interesting or innovative or unexpected ways that you have seen that used either in the context of data modeling or the overall impact that it has had on the engineering workflow to be able to build out these different analytical datasets?
457
00:37:28.860 -->
00:37:29.520 Yeah. So,
458
00:37:30.015 -->
00:37:32.275I mean, when when we came up with this and,
459
00:37:33.055 -->
00:37:36.355you know, we you know, our clients started using our product.
460
00:37:36.815 -->
00:37:56.555What I have seen is they started building these, you know, what we call nodes in in the in the platform. So a node is just a, you know, a pattern. Basically, a node type is a pattern. And behind the pattern, you can write a template. And the template is actually leveraging the metadata behind the scenes, this column level metadata. So what I'm super excited is
461
00:37:56.860 -->
00:37:57.360about,
462
00:37:57.660 -->
00:38:00.640you know, these customers coming up with innovative
463
00:38:00.940 -->
00:38:03.680templates. They are coming up with this their own ideas
464
00:38:04.140 -->
00:38:06.080of how they can leverage this
465
00:38:06.565 -->
00:38:07.705column level metadata
466
00:38:08.005 -->
00:38:16.690and start creating templates to meet their needs. So every time you write a new template, you're basically creating a new note type that can be leveraged by anybody.
467
00:38:16.990 -->
00:38:20.369If they share it in the in the future with other clients,
468
00:38:20.990 -->
00:38:24.369I don't have to rebuild anything. I just use the note type.
469
00:38:25.185 -->
00:38:27.605But but the way that they're doing is is,
470
00:38:28.145 -->
00:38:29.045is very exciting.
471
00:38:29.345 -->
00:38:34.050How they are able to come up with innovative solutions to meet their needs riding
472
00:38:34.430 -->
00:38:35.330these no pips.
473
00:38:35.869 -->
00:38:38.690 And in your own work of building
474
00:38:39.070 -->
00:38:46.025Coalesce and building this column aware tooling, what are some of the most interesting or unexpected or challenging lessons you've learned in the process?
475
00:38:46.645 -->
00:38:54.960 It was it was definitely a big it was a journey for us. You know, we have we have done went back and forth in a lot of deep technical
476
00:38:55.420 -->
00:38:56.640architectural things.
477
00:38:56.975 -->
00:38:59.875What what I have learned, let's see.
478
00:39:00.175 -->
00:39:01.955I think the number 1 thing is
479
00:39:02.255 -->
00:39:06.690the potential. I did not realize how much we can do with this, but the
480
00:39:06.990 -->
00:39:25.900more we do it, the more we think we can do with this. For example, when we when we came out with it, we had some things in mind. Okay. Yeah. We'll get lineage. Yeah. We'll get lineage. Lineage is gonna be a great thing. But then we're saying, okay. Not only that, we can start building this, what we call custom node types.
481
00:39:26.280 -->
00:39:28.780That's another thing. But now we're saying
482
00:39:29.080 -->
00:39:32.140how we can, you know, apply data quality, for example,
483
00:39:32.495 -->
00:39:34.595with with the same column level metadata.
484
00:39:35.375 -->
00:39:40.435And and now we're entering this AI world, and I can already see how we can
485
00:39:41.340 -->
00:39:42.720even do some of those
486
00:39:43.340 -->
00:39:44.800automation using AI
487
00:39:45.420 -->
00:39:47.040with the column level metadata.
488
00:39:47.500 -->
00:39:54.755So so so my point is as we do more, we see that this architecture that we chose is
489
00:39:55.615 -->
00:39:57.875is still has a lot of
490
00:39:58.200 -->
00:39:58.700potential,
491
00:39:59.560 -->
00:40:01.260to be used in the future use cases.
492
00:40:01.800 -->
00:40:03.660 And for people who are
493
00:40:04.119 -->
00:40:05.500building out their
494
00:40:05.875 -->
00:40:19.000data transformation workflows, they're deciding on how best to approach their modeling, what are some of the cases where a column aware approach is the wrong choice? I I think there's no such thing in my opinion as
495
00:40:19.300 -->
00:40:21.400 column awareness is incorrect
496
00:40:21.700 -->
00:40:22.600because especially
497
00:40:23.454 -->
00:40:25.635we're not asking them to
498
00:40:26.494 -->
00:40:27.795spend a lot of time
499
00:40:28.095 -->
00:40:31.714to do this. Now, of course, if you have a tool that is column aware,
500
00:40:32.119 -->
00:40:36.220they don't have to do it. But if if they have to build something,
501
00:40:36.520 -->
00:40:39.740then from the ground up in their in their environment,
502
00:40:40.200 -->
00:40:40.700and
503
00:40:41.095 -->
00:40:44.635they have to make it column aware, then it's a different story.
504
00:40:45.015 -->
00:40:54.019Because then there might be situations where you might you might say, I don't need the column awareness for this. Yeah. It's possible that I can see some use cases where they,
505
00:40:54.880 -->
00:41:00.259maybe they might it might be an overkill for them. But but if you adapt a tool that already has it,
506
00:41:00.705 -->
00:41:09.765then, you know, then it's not a big deal because you're not you're not you're only leveraging it. You're not trying to build everything from the ground up. And for teams who are
507
00:41:10.220 -->
00:41:14.160 looking to improve their data modeling approaches, they're looking to
508
00:41:14.540 -->
00:41:17.260simplify their overall workflow, what are some of the,
509
00:41:17.580 -->
00:41:31.430resources that you recommend that they look to for being able to learn more about some of the data modeling principles, some of the ways that column awareness can be added to or integrated into their, overall workflow or tool chains, etcetera.
510
00:41:31.810 -->
00:41:35.910 Yeah. You know, I've I've been in this space for a long time. So my recommendations
511
00:41:36.210 -->
00:41:44.695would be, you know, the the classic books, the data warehouse toolkit from Ralph Kimball and and, you know, things like that.
512
00:41:45.210 -->
00:41:48.670I don't expect everybody to just go through the whole Ralph Kimball
513
00:41:49.210 -->
00:41:52.109book, which might be, first of all, a very
514
00:41:52.555 -->
00:41:55.694lengthy book and some people might feel like
515
00:41:55.994 -->
00:41:57.775some of the concepts are outdated.
516
00:41:58.155 -->
00:42:00.015But I think, as I said before,
517
00:42:00.340 -->
00:42:08.280just focus on, like, the 1 on 1 of modeling for us. Start there. There's I can't name 1 resource. You can find so much on the on the Internet.
518
00:42:08.795 -->
00:42:12.175But make sure that you understand why what's the benefit of modeling?
519
00:42:12.475 -->
00:42:18.590Why do we need to model? And don't fall into the trap of perfecting the model. That's another thing I my advice.
520
00:42:18.970 -->
00:42:23.695Always keep iterating because it's going to change regardless. There's no perfect
521
00:42:24.075 -->
00:42:24.635model. And,
522
00:42:25.195 -->
00:42:39.280 I I would say, you know, there's a there's a lot of resources, but but start small and and kinda play with them. Alright. Are there any other aspects of the overall space of data modeling and data architecture in the context of cloud warehousing in particular
523
00:42:39.580 -->
00:42:40.320and the
524
00:42:40.805 -->
00:42:41.305capabilities
525
00:42:41.685 -->
00:42:47.420and impact of column aware tooling that we didn't discuss yet that you'd like to cover before we close out the show?
526
00:42:48.059 -->
00:42:55.520 1 thing that that I didn't talk about with the Calm awareness is the the documentation that you get out of this. I mean,
527
00:42:55.994 -->
00:42:57.934the governance aspect. I mean,
528
00:42:58.714 -->
00:43:03.055you you we have seen a lot of times where people take pride in coding,
529
00:43:03.515 -->
00:43:09.890which is great. And you quote all these things, and as you get your environments get bigger and bigger, you have a lot more code,
530
00:43:10.270 -->
00:43:23.550and then the parsing leaves. And then there you go. It's becomes very, very difficult for anybody to kinda go and understand the code. And this is a classic thing. We all know this. But column awareness again, another byproduct of that is the
531
00:43:23.850 -->
00:43:26.430the documentation that you can generate out of this.
532
00:43:26.810 -->
00:43:28.350Now how this was built.
533
00:43:28.895 -->
00:43:48.435And you can also add context to also show why this was built the way it was built. And by just pushing a button, it can read this column metadata and generate this documentation for you. And it's always going to be accurate whether you agree or not with that, but that's the truth. This is how it was built,
534
00:43:48.815 -->
00:43:50.275and this is what we're showing.
535
00:43:50.895 -->
00:43:52.995So that's another big value
536
00:43:53.375 -->
00:43:54.835that you get from that.
537
00:43:55.339 -->
00:44:11.819 Alright. Well, for anybody who wants to get in touch with you and follow along with the work that you and your team are doing, I'll have you add your preferred contact information to the show notes. And as the final question, I'd like to get your perspective on what you see as being the biggest gap in the tooling or technology that's available for data management today. Yeah. I think,
538
00:44:12.440 -->
00:44:13.180 a tool
539
00:44:13.480 -->
00:44:19.660or a platform, I would say. We we like to call ourselves a platform because of our we think big. And
540
00:44:19.995 -->
00:44:21.855where business and IT
541
00:44:22.475 -->
00:44:23.615power business users
542
00:44:23.995 -->
00:44:25.535and IT architects,
543
00:44:26.395 -->
00:44:30.410data architects, data engineers, where they can all work together
544
00:44:30.710 -->
00:44:33.450on 1 single platform to build their data foundation
545
00:44:33.830 -->
00:44:34.810and data assets.
546
00:44:35.295 -->
00:44:45.549And that's the that's the gap right now. Either the tools are too extreme to to the business needs, so the business users are just using it, using those tools,
547
00:44:46.250 -->
00:44:49.150or it's the other way around. Like, the IT
548
00:44:49.450 -->
00:44:50.990oriented tools that
549
00:44:51.369 -->
00:45:02.255business doesn't have any visibility or an or or have the training to understand how these tools work. What Coalesce, for example, does is to bring these 2 types of personas,
550
00:45:03.460 -->
00:45:07.960or many other personas into the same platform. So everything is being built there,
551
00:45:08.500 -->
00:45:15.995 and it's all governed in 1 place. Alright. Well, thank you very much for taking the time today to join me and share your thoughts on
552
00:45:16.375 -->
00:45:31.135the benefits and challenges of data modeling and some of the ways that column aware tooling can help with that overall effort. Definitely appreciate the time and energy that you and your team are putting into making that available. So, thank you for that. I hope you enjoy the rest of your day. Thanks, Tobias. Thanks
553
00:45:31.435 -->
00:45:32.415 for having me.
554
00:45:38.430 -->
00:45:39.330 Thank you for listening.
555
00:45:39.630 -->
00:45:48.635Don't forget to check out our other shows, podcast dot in it, which covers the Python language, its community, and the innovative ways it is being used, and the Machine Learning podcast,
556
00:45:49.095 -->
00:45:53.595which helps you go from idea to production with machine learning. Visit the site at dataengineeringpodcast.com
557
00:45:55.015 -->
00:46:11.925to subscribe to the show, sign up for the mailing list, and read the show notes. And if you've learned something or tried out a product from the show, then tell us about it. Email hosts at dataengineeringpodcast.com with your story. And to help other people find the show, please leave a review on Apple Podcasts and tell your friends and coworkers.