WEBVTT
NOTE
Transcription provided by Podhome.fm
Created: 7/6/2024
12:32:03 PM
Duration: 3461.910
Channels: 1
1
00:00:12.594 -->
00:00:16.695 Hello, and welcome to the Data Engineering Podcast, the show about modern data management.
2
00:00:17.710 -->
00:00:24.849When you're ready to build your next pipeline and want to test out the projects you hear about on the show, you'll need somewhere to deploy it. So check out our friends over at Linode.
3
00:00:25.685 -->
00:00:34.265With our managed Kubernetes platform, it's now even easier to deploy and scale your workflows or try out the latest Helm charts from tools like Pulsar, Pacaderm, and Dagster.
4
00:00:35.190 -->
00:00:42.410With simple pricing, fast networking, object storage, and worldwide data centers, you've got everything you need to run a bulletproof data platform.
5
00:00:43.335 -->
00:00:44.955Go to data engineering podcast.com/linode
6
00:00:46.455 -->
00:00:56.360today. That's l I n o d e, and get a $100 credit to try out a Kubernetes cluster of your own. And don't forget to thank them for their continued support of this show.
7
00:00:57.140 -->
00:00:58.680Struggling with broken pipelines,
8
00:00:59.155 -->
00:00:59.895stale dashboards,
9
00:01:00.355 -->
00:01:01.255missing data?
10
00:01:01.875 -->
00:01:04.055If this resonates with you, you're not alone.
11
00:01:04.515 -->
00:01:12.280Data engineers struggling with unreliable data need look no further than Monte Carlo, the world's 1st end to end fully automated data observability
12
00:01:12.580 -->
00:01:13.080platform.
13
00:01:14.340 -->
00:01:21.075In the same way that application performance monitoring Monte Carlo monitors and alerts for data issues across your
14
00:01:22.815 -->
00:01:23.315data
15
00:01:24.549 -->
00:01:35.235Monte Carlo monitors and alerts for data issues across your data warehouses, data lakes, ETL, and business intelligence, reducing time to detection and resolution from weeks or days to just minutes.
16
00:01:37.215 -->
00:01:41.395Start trusting your data with Monte Carlo today. Go to data engineering podcast.com/
17
00:01:42.390 -->
00:01:44.090monte carlo to learn more.
18
00:01:44.390 -->
00:01:48.170The first 10 people to request a personalized product tour will
19
00:01:49.270 -->
00:01:55.205the first 10 people to request a personalized product tour will receive an exclusive Monte Carlo swag box.
20
00:01:56.305 -->
00:02:09.705Your host is Tobias Macy. And today, I'm welcoming back Maxime Beauchmann to talk about the impacts that the evolution of the modern data stack has had on the role and responsibilities of data engineers. So Max, for anybody who isn't familiar with you, can you give a brief introduction?
21
00:02:10.084 -->
00:02:14.345 For sure. Yeah. And thank you for having me on the show again. Excited to be here. So
22
00:02:14.805 -->
00:02:20.880how to best introduce myself? I think at this point in my career, I'm probably best known for the work that I've done around
23
00:02:21.340 -->
00:02:27.265Apache Airflow and Apache Superset. So I started both these open source projects when I was at Airbnb
24
00:02:28.045 -->
00:02:29.105back in 2014
25
00:02:29.485 -->
00:02:30.145and 15.
26
00:02:31.005 -->
00:02:31.745Since then,
27
00:02:32.045 -->
00:02:34.625I went on to start a company called Preset
28
00:02:35.150 -->
00:02:44.865where we essentially offer Apache Superset as a service. For those not familiar with Apache Superset, it is a very much a data visualization exploration platform
29
00:02:45.405 -->
00:02:49.985that caters to like all of the business intelligence type use cases and beyond.
30
00:02:50.510 -->
00:02:55.570And then talking a tiny bit more about my career. So over the past 20 years, I've been,
31
00:02:55.870 -->
00:02:59.330you know, a business intelligence engineer, data warehouse architect.
32
00:02:59.775 -->
00:03:01.954By the time I joined Facebook in 2012,
33
00:03:02.814 -->
00:03:08.355I believe, is when we started calling ourselves data engineers internally at Facebook, and that's the title
34
00:03:08.800 -->
00:03:11.380that followed me for much of the decade
35
00:03:12.160 -->
00:03:17.620to come. And then I've just been building a lot of data tools. Like, I really enjoyed building
36
00:03:18.245 -->
00:03:21.065tooling around data, so that's really my my passion.
37
00:03:21.685 -->
00:03:22.185 And
38
00:03:22.565 -->
00:03:24.665in terms of your introduction
39
00:03:24.965 -->
00:03:36.835to data, we've gone over that a couple of times in past episodes you've been on, so we'll make you rehash that. I'll just I'll let a link in the show notes for people who wanna go back and hear about that. But as far as the topic at hand today,
40
00:03:37.215 -->
00:03:43.155you recently had a post that was talking about how the modern data stack is reshaping data engineering.
41
00:03:44.080 -->
00:03:59.525And before we dig too much into that, I'll also call out the previous articles you had done almost 5 years ago on the rise and the downfall of the data engineer, and then we also did an episode all the way back in episode 3 talking about defining data engineering
42
00:04:00.460 -->
00:04:07.680because it was still very early in the journey of data engineering being a dedicated role and being something that people would actually go out and get as their job title.
43
00:04:08.175 -->
00:04:21.130So now almost 5 years later, it's almost incomprehensible that that was even the state of the world at the time. So I'm wondering if you can just give your current working definition of what a data engineer is and does
44
00:04:21.510 -->
00:04:25.324given how much it has shifted since the last time we covered that?
45
00:04:25.625 -->
00:04:33.965 Things have have changed quite a bit, right, over the 4 or 5 years, and that's really what I was interested in and that latest blog post. So I highly encourage people to
46
00:04:34.330 -->
00:04:42.750even, like, pause this podcast and read the post because I think today we're gonna talk about a lot of these trends and how they're shaping and changing
47
00:04:43.384 -->
00:04:45.725the data the modern data team and
48
00:04:46.104 -->
00:04:51.485the modern data engineer. But I think it's the definition of what a data engineer
49
00:04:52.180 -->
00:04:56.840is at its core hasn't really changed, right? It says the practice of designing and building
50
00:04:57.540 -->
00:05:07.555systems and processes for collecting, sorting, analyzing data at scale. Like at a at a high level, you know, a data engineer is someone who built kind of systems processes
51
00:05:08.495 -->
00:05:10.275around data and metadata.
52
00:05:10.930 -->
00:05:14.789It's a super broad field, right, like that touches just about every
53
00:05:15.250 -->
00:05:17.349industry, every team, every department.
54
00:05:17.945 -->
00:05:23.0051 thing that has changed quite a bit, I think, over the past decade is just how
55
00:05:23.385 -->
00:05:27.165mainstream data has become. So, you know, 20 years ago
56
00:05:27.650 -->
00:05:38.495when I was a data warehouse architect, it was the craft of a very small group of people in the company to be kind of the librarians of the data in the company. It was very kinda
57
00:05:39.035 -->
00:05:55.710focused and targeted to a small group of people to take care of that. And now it's like every company wants to be data driven. Every team's a data team. People are investing a lot in their data team and that kind of skills. So things have become much more mainstream over the past decade, and then there's been,
58
00:05:56.074 -->
00:05:59.615like, new roles that have been kinda shaping up. So there's some specializations,
59
00:06:00.074 -->
00:06:01.055some new tooling.
60
00:06:01.595 -->
00:06:11.880There's some tooling that makes some of the things that used to take a lot of time now, just kind of something that doesn't take time anymore. So it will be interesting to talk about all of these things and all these changes.
61
00:06:12.895 -->
00:06:25.539 And to your point, at the time that we first visited the idea of what is data engineering, it was still very much a low level operation where you had to be, as you said, well versed in the
62
00:06:26.000 -->
00:06:33.685mechanical aspects of how data was laid out on disk, how the processing engines were going to work with it, whereas now a lot of that has been
63
00:06:33.985 -->
00:06:41.840pushed into the software layer. You don't need to think about it. You just say, I wanna take the data and put it from point a to point b. There are services to do that. You don't need
64
00:06:42.220 -->
00:06:54.835to think about all of the retry logic and error conditions that go into all of those things that wasted a lot of time and caused a lot of headaches. So I'm wondering how the growing availability of these data infrastructure services
65
00:06:55.215 -->
00:06:55.715and
66
00:06:56.060 -->
00:06:56.720the utilities
67
00:06:57.099 -->
00:06:59.919that are built on top of and around them have
68
00:07:00.220 -->
00:07:05.360shifted the foundational skills and knowledge that are necessary to be effective as a data engineer
69
00:07:06.045 -->
00:07:15.260and some of the ways that new and aspiring data engineers should think about spending their time and energy to actually break into that role? There is a lot in this question
70
00:07:15.560 -->
00:07:18.300 on PAC. 1 thing is, like, the rise of
71
00:07:18.760 -->
00:07:19.500the services
72
00:07:20.440 -->
00:07:25.695that automate a lot of what a data engineer used to do. So on 1 front, there's these cloud data warehouses,
73
00:07:26.395 -->
00:07:31.935commoditizing kind of the database administrator type workloads. Right? Even like the infrastructure
74
00:07:33.340 -->
00:07:37.840load of the data engineer. I think in the rise of the data engineer, I talked about, like, a, some people
75
00:07:38.300 -->
00:07:49.025include the infrastructure work of, like, setting up your data structure as part of the data engineering role. And I think like that as definitely that with the cloud services that exist today,
76
00:07:49.460 -->
00:08:02.705you don't need to go and kind of set up your own data warehouse. Right? Like, what all you need to do is kinda create a a Snowflake account, a BigQuery account, and you're up and running fairly quickly. You pay as you go, or you don't even have to necessarily
77
00:08:03.965 -->
00:08:08.030size and kinda grow your cluster based on needs. Like, all that stuff is done
78
00:08:08.510 -->
00:08:09.250for you.
79
00:08:09.710 -->
00:08:21.545That does mean, though, that there's still a burden around provisioning, like choosing the technology that you're gonna use and giving access to people and then procurement. Right? Like, making sure you're choosing the right thing for the right reason
80
00:08:22.085 -->
00:08:23.465and perhaps containing
81
00:08:24.085 -->
00:08:28.910costs in some ways or monitoring costs is becoming maybe more of a concern
82
00:08:29.690 -->
00:08:31.150over time. There's
83
00:08:31.775 -->
00:08:43.589another part of, like, the squish in some way of, like, on 1 end. Right? Like, the data engineer doesn't have to to do some infrastructure work. Maybe it doesn't have to do as much of the scripting to hoard data.
84
00:08:44.130 -->
00:08:49.990We're we're in a phase where, like, data warehousing is, like, a lot of it is about hoarding the data from all of the different
85
00:08:50.505 -->
00:08:51.645systems and subsystems
86
00:08:52.185 -->
00:08:56.845in your company into a central place. Nowadays, that means getting a bunch of
87
00:08:57.360 -->
00:09:01.380data from your SaaS services. Right? A modern company uses 100
88
00:09:02.080 -->
00:09:03.940of SaaS services to operate,
89
00:09:04.320 -->
00:09:08.855whether it's around, you know, recruiting or customer CRM type things.
90
00:09:09.315 -->
00:09:14.055In all areas of business nowadays, we use, you know, targeted specialized SaaS
91
00:09:14.580 -->
00:09:22.440systems. Right? Whether it's payroll or pretty much in all areas. So bringing all this data, hoarding the data to the data warehouse
92
00:09:22.895 -->
00:09:26.595is also something that's getting commoditized by tooling with things like
93
00:09:26.975 -->
00:09:29.075Fivetran and Meltano and Airbyte.
94
00:09:29.695 -->
00:09:30.435It's becoming
95
00:09:31.700 -->
00:09:45.634easier to bring all of the data into a central place, so then there's a little bit of some workload is, like, setting that stuff up, making sure it works, you know, monitoring these things and getting all, you know, the procurement and the operation
96
00:09:46.415 -->
00:09:48.834and the selection of that tooling is still big.
97
00:09:49.380 -->
00:09:50.279Another area
98
00:09:51.139 -->
00:09:59.334that we've seen changes is, like, the rise of the analytics engineer. That means now we have a data analyst who speaks SQL and
99
00:09:59.714 -->
00:10:01.894knows Git pretty well, so that means they can
100
00:10:02.195 -->
00:10:11.340start automating some of the the t and EL t. Right? So these people are able now to kind of solve their own problem and automate their pipelines.
101
00:10:12.095 -->
00:10:16.194So that pushes the data engineer a little bit further away from this.
102
00:10:16.654 -->
00:10:22.490So back to your question now, what does that mean in terms, like, what skills do you need as a data engineer
103
00:10:23.270 -->
00:10:29.665nowadays? I think it's an interesting question. Right? I think, like, there's a question around specialization,
104
00:10:30.205 -->
00:10:38.560like, how broad do you wanna go? Do you wanna be more full stack, right, and be able to cover some of the data analyst to data infrastructure
105
00:10:39.420 -->
00:10:56.750spectrum, or do you wanna really focus and be the person who manages core datasets? There's also an area that seems to be just as complex and as kind of very still well attributed to data engineers, all the streaming the streaming pipelines. This area, clearly, if your company needs to have streaming data pipelines,
106
00:10:57.370 -->
00:10:58.430stream data processing,
107
00:10:58.810 -->
00:11:02.190actually, you know, I think that stays under the the realm of
108
00:11:02.524 -->
00:11:05.105specialty of the the modern data engineer as well.
109
00:11:05.485 -->
00:11:15.120 With the introduction of these managed services where a lot of the work to set up the foundational data platform is just sign up for the service,
110
00:11:15.500 -->
00:11:19.040put in the, you know, credentials, do some of the integration work,
111
00:11:19.384 -->
00:11:24.444where that starts to sound a lot more like an infrastructure engineer than a data engineer responsibility.
112
00:11:25.144 -->
00:11:38.394And I'm wondering what you see as the potential for at least some of the more service and infrastructure level work to be pushed into the domain of the DevOps engineer or the platform engineer
113
00:11:39.014 -->
00:11:49.520and less so in the realm of the data engineer where the data engineer is maybe just the 1 making the selection of which tools to actually purchase and integrate and less of doing that actual integration work.
114
00:11:49.980 -->
00:11:58.665 I think that's clear. Right? Like, we we can kinda hand that over to the infra cloud infrastructure team. They can handle it just as they handle other
115
00:11:59.365 -->
00:12:08.640systems and piece of infrastructure that they do handle. So that means procurement process and, you know, even doing, like, security review. Right? Is that tool matching our security
116
00:12:09.100 -->
00:12:11.360type requirements? It's SOC 2 compliant.
117
00:12:12.045 -->
00:12:13.425Something can be done
118
00:12:13.964 -->
00:12:16.944in tandem or even led by your normal
119
00:12:17.245 -->
00:12:19.024infra cloud infrastructure team.
120
00:12:19.340 -->
00:12:24.080I think that in terms, like, wiring all these things together, so we buy all of these,
121
00:12:24.620 -->
00:12:28.320you know, services and tooling, and there's there's still, like, a responsibility
122
00:12:28.620 -->
00:12:29.120of
123
00:12:29.805 -->
00:12:35.905making these things work and then tying them together. Right? Like so I think that's true of infrastructure
124
00:12:36.540 -->
00:12:41.520in general. It certainly is true with data infrastructure too. Right? It's not just like buying
125
00:12:41.900 -->
00:12:52.5555 Tran and dbt cloud or whatever, you know, or astronomer cloud and then getting these things to I should be like, okay. We've bought them. Now we're done. And I say you still need to go and make these things work
126
00:12:53.270 -->
00:12:56.570very well together. So there's, like, you know, metadata integration.
127
00:12:57.510 -->
00:13:04.274There is essentially, like, just the duct tape and the chicken wire that's required to get all these things to work together.
128
00:13:04.654 -->
00:13:07.475And I think the reality of the modern software engineer,
129
00:13:08.095 -->
00:13:12.980just as much as, like, the modern, you know, data engineer, is to get all these services
130
00:13:13.920 -->
00:13:24.585to work well together and then to take, you know, all the business rules and the the things that are specific to your company and and, you know, put those and make sure that those things are reflected in those systems.
131
00:13:25.365 -->
00:13:32.090 1 of the big themes in the idea of the modern data stack and all the services that are becoming available
132
00:13:32.550 -->
00:13:33.450and the
133
00:13:33.990 -->
00:13:55.985sort of areas of focus for data engineers and data product managers is the idea of democratization of data where you want to make data access more universal throughout the organization. You want to lower the barrier of entry, lower the level of sophistication that's necessary to be able to actually explore these different datasets that are powering the business.
134
00:13:56.605 -->
00:14:06.950And in your post about the downfall of the data engineer, you called out the pressure on data engineers to maintain control with so many different contributors with varying levels of skill and understanding.
135
00:14:07.810 -->
00:14:17.565And I'm wondering how you see the modern data stack balancing some of those concerns of giving everybody access to the data, you know, empowering them to actually ask and answer questions,
136
00:14:18.024 -->
00:14:22.285but at the same time, not overwhelming the data engineer or not introducing
137
00:14:22.904 -->
00:14:23.530sort of
138
00:14:23.930 -->
00:14:28.750uncontrolled manipulation of the data in a way that actually causes there to be
139
00:14:29.450 -->
00:14:39.435invalid assumptions based on the, you know, unknown quality of the datasets or people who are creating new datasets without necessarily understanding what the original context was?
140
00:14:39.920 -->
00:14:46.899 Yeah. So I think, like, if we think about this problem of, you know, if we give more access to more things to more people just in the abstract,
141
00:14:47.355 -->
00:14:51.935know, there's this danger of, like, people getting lost or, you know, shooting themself in the foot.
142
00:14:52.235 -->
00:14:57.055And and I think that's a general problem. Once something becomes more
143
00:14:57.420 -->
00:15:01.440accessible to more people, there's a risk that it might be misused or misunderstood.
144
00:15:02.140 -->
00:15:11.415So a big thing is the education gap. Right? So do we make sure that we provide resources for these people to use the tooling right? And is the tool
145
00:15:12.030 -->
00:15:19.090well structured and organized to provide all the context that the people need to succeed in what they're trying to achieve
146
00:15:19.630 -->
00:15:20.450with the tool.
147
00:15:20.765 -->
00:15:24.145So that probably mean 1 big thing is like metadata accessibility.
148
00:15:24.525 -->
00:15:31.850Right? So if you're building a dashboard from a dataset, like how do you know all the context on this dataset? Like, is it fresh?
149
00:15:32.150 -->
00:15:33.130Who's the owner?
150
00:15:33.670 -->
00:15:35.770Is it reliable? Is it certified?
151
00:15:36.545 -->
00:15:40.645So I think some of these questions can be answered through the use of, like, good
152
00:15:41.025 -->
00:15:48.430metadata and metadata management and maybe, like, data dictionary, that kind of that space of call it data graph or the metadata graph,
153
00:15:48.890 -->
00:15:51.310understanding where is it coming from, who owns it,
154
00:15:51.770 -->
00:16:00.425how reliable is this, who are other users of this dataset. Right? It's very popular and used every day by many of your colleagues. Your colleagues is probably,
155
00:16:01.009 -->
00:16:01.829you know, reliable.
156
00:16:02.690 -->
00:16:13.995So that's, like, somewhat, like, beyond the tribal knowledge of going to the Slack channel called data questions and asking people, like, hey. Does anyone know about this dataset and whether I should use this?
157
00:16:14.375 -->
00:16:16.315Metadata accessibility is important.
158
00:16:16.860 -->
00:16:19.279Educating your workforce. So at
159
00:16:19.660 -->
00:16:25.839previous companies, we had programs to make sure that we push data literacy forward
160
00:16:26.635 -->
00:16:35.295internally and make that accessible to make that accessible to just about everyone within a company. So at Airbnb, we had data university
161
00:16:35.899 -->
00:16:46.915where we we taught, and I'm sure the program is still going on, maybe has changed over time, but my memory of it is we trained people on the tooling that we had, and there was, like, a progression
162
00:16:47.215 -->
00:16:51.235of, like, you know, learning airflow 101 and airflow 201 and airflow 301.
163
00:16:51.935 -->
00:17:06.304There was also classes around data structures and the tables and the datasets that are most popular, and then just also, like, orientation of, like, how do you ask a question? How do you find out what the dataset you might wanna use or might not wanna use is.
164
00:17:06.765 -->
00:17:11.565Similarly, at Facebook, there was something called DataCamp, and that was a little bit more. Instead of being,
165
00:17:12.520 -->
00:17:22.645you know, a series of classes, maybe with a commitment a few hours a week, so that was the approach at Airbnb. At Facebook, I believe it was a full 1 week called data camp, and you would just
166
00:17:22.945 -->
00:17:25.765almost kinda check out of your team for a whole week
167
00:17:26.065 -->
00:17:28.245and then go sit through a bunch of
168
00:17:28.705 -->
00:17:29.205classes.
169
00:17:29.585 -->
00:17:30.085And
170
00:17:30.420 -->
00:17:36.600it was, like, classes and exercises too. Right? So they would ask some questions, get little projects,
171
00:17:36.980 -->
00:17:37.640and play,
172
00:17:38.035 -->
00:17:49.760you know, data analysts for for a whole week, which was pretty exciting. And they made it pretty fun too where you would learn about the datasets. You would go and answer some really kinda key intricate questions of, like, hey. How does
173
00:17:50.460 -->
00:17:50.960engagement
174
00:17:51.260 -->
00:17:56.545work for different age groups? And are teens as engaged on Facebook as,
175
00:17:56.845 -->
00:18:04.350you know, your different groups of people? And then go and run your your own analysis and learn about all the tooling that we had available internally.
176
00:18:04.730 -->
00:18:14.815So there's this education gap. I think that's a big 1. There's for the tooling to show more context. Think, is 1 way to help with that. I'll open up on, like, the topic of,
177
00:18:15.115 -->
00:18:22.140call it data literacy or call it democratization of access to data. There's some bigger topics there, like, do we wanna democratize
178
00:18:22.920 -->
00:18:32.865the entire analytics process? Right? Do we want to make it possible for more people to write pipelines, to for more people to go and instrument more events and application,
179
00:18:33.565 -->
00:18:43.080or more people to go and define, you know, business rules and things like that too. So I think the I think the answer is yes. And then the question is like, what are the right set of guardrails
180
00:18:43.465 -->
00:18:47.085in education we need to enable more people to do more of this?
181
00:18:47.545 -->
00:18:49.485 Yeah. The democratization
182
00:18:49.865 -->
00:19:04.735is definitely something that's worth kind of enumerating where it could just mean giving people access to read it. But as you said, maybe you wanna be able to give everybody access to write their own pipelines, to be able to build their own datasets that power their specific segment of the business where,
183
00:19:05.035 -->
00:19:23.044you know, 1 of the areas that's most recent that's seeing a lot of attention is the idea of the metrics layer where you wanna bring the business users in to understand the definitions of what that metric is supposed to mean semantically and some of the ways that the data that we have can be used to actually formulate that metric
184
00:19:23.820 -->
00:19:27.760because the sales manager is more likely to know what the actual
185
00:19:28.220 -->
00:19:29.760semantics around a conversion
186
00:19:30.140 -->
00:19:41.184should be versus a data analyst or a data engineer because, you know, we're working at the layer of the data so we can see, okay, these are all the numbers. These are the different events that tie together. But from the business perspective,
187
00:19:41.645 -->
00:20:04.090what does it actually mean to be a conversion, and how is that being used in the data? So we wanna make sure that everybody's working together on that. From the pipeline perspective. You know? And we have the core set of data pipelines that are kind of protected, and you don't just grant access to everyone. But from that base set of datasets that we're pulling in from the, you know, application databases,
188
00:20:04.470 -->
00:20:10.905from Salesforce, whatever it might be, we then wanna be able to give people access to build their own downstream
189
00:20:11.285 -->
00:20:12.825pipelines, downstream datasets.
190
00:20:13.125 -->
00:20:23.289But to the point of guardrails, maybe we say here is the kind of cookie cutter template of your DBT model to say, you know, these are the core datasets you're able to pull from.
191
00:20:23.669 -->
00:20:47.445Here is the initial set of operations to build a new transformation or build a new table so that it's using maybe the agreed upon vocabulary as far as column names. But you can now go and build some other view on this dataset that you can consume in this dashboard. So you have kind of the templated out set of steps in the pipeline so that all ties together with your kind of paved path. And then if they go
192
00:20:47.825 -->
00:20:54.565a field of that, then they're kind of on their own, and you make no guarantees about the validity of their datasets that they're building.
193
00:20:55.000 -->
00:21:09.825 Yeah. I mean, I'm interested to talk about, like, the metrics layer and, like, what it's after and what's novel about it and what's maybe so not so novel around it too. Because, like, you know, if you go back to the artificial data warehousing books that are, you know, 25 years old now,
194
00:21:10.125 -->
00:21:15.950so the Ralph Kimball books, the Bill Inman books, it was always about metrics and dimensions.
195
00:21:16.730 -->
00:21:18.830It was about, you know, conformed dimensions,
196
00:21:19.529 -->
00:21:21.549conformed metrics, conformed facts,
197
00:21:21.985 -->
00:21:24.805getting consensus, defining these things very, very well,
198
00:21:25.345 -->
00:21:27.605having these things be a reflection of the business.
199
00:21:28.145 -->
00:21:39.230So I think that these ideas are not new. Like, there's even, like, metric centric data modeling. Like, to say, like, oh, the metric is the most important thing, you know, at the heart of data modeling
200
00:21:39.530 -->
00:21:42.684or or even from a data governance standpoint.
201
00:21:43.385 -->
00:21:52.140You know? I think it does kinda make sense, but it really I think what is screaming to me, you know, looking at these metrics layer and, you know, different
202
00:21:52.680 -->
00:21:59.755entities and people, companies gonna emerge in the space, like, talk about different things when they say metrics layer 2.
203
00:22:00.215 -->
00:22:09.130But I think some common things and themes that we see, 1 is, like, beyond the dbt world of templated SQL, like, we need higher level abstractions
204
00:22:10.070 -->
00:22:19.415that maybe come with more constraints and guarantees than just like your raw templated SQL. So templated SQL is too free form. You know, anybody can do anything.
205
00:22:19.875 -->
00:22:22.695It's kind of the far west, so maybe the metrics layer
206
00:22:23.100 -->
00:22:30.559is a little bit more prescriptive in what you can and cannot do and how you have to define, say, ownership of things or how
207
00:22:30.985 -->
00:22:33.085things are derived or how
208
00:22:33.785 -->
00:22:40.880things like time window, you know, are expressed more semantically instead of, like, writing, like, these more complex, you know, unreadable
209
00:22:41.420 -->
00:22:42.559mountains of SQL.
210
00:22:42.940 -->
00:22:45.520So I think there's, like, this higher level abstraction
211
00:22:46.540 -->
00:22:48.320with more constraints and guarantees.
212
00:22:49.215 -->
00:23:00.9801 thing that makes a lot of sense to me that people are not necessarily talking about too much is this idea of, like, more entity centric data modeling. So when you think about, like, metric centric data modeling, that means, like, hey, we're gonna make the metric
213
00:23:01.440 -->
00:23:03.539to the kind of unit, that really
214
00:23:03.840 -->
00:23:04.740strong entity
215
00:23:05.485 -->
00:23:06.145in information
216
00:23:06.605 -->
00:23:07.745architecture around
217
00:23:08.125 -->
00:23:09.905how we manage data and metadata.
218
00:23:10.285 -->
00:23:16.280Right? I think that makes sense. Like if you think at Airbnb or like bookings, you know, bookings is really important thing.
219
00:23:16.660 -->
00:23:20.360Let's define, like, who owns certain, like, subsets of dimensions
220
00:23:21.220 -->
00:23:21.720around
221
00:23:22.100 -->
00:23:23.640how bookings are defined.
222
00:23:24.245 -->
00:23:26.985And then we all need to align on a definition of this stuff.
223
00:23:27.445 -->
00:23:37.660I like entity centric data modeling, which is, like, when you think about it, like, the Kimbell book is all about, like, dimensional modeling, which a dimension is an entity. Right? So it is very much entity
224
00:23:38.120 -->
00:23:41.179centric data modeling. I I like to push this idea
225
00:23:41.605 -->
00:23:43.225more forward and say
226
00:23:43.684 -->
00:23:45.384beyond dimensional modeling,
227
00:23:45.764 -->
00:23:52.620if you look at, like, feature stores nowadays that are more emerging in the field of, like, ML type area and feature engineering.
228
00:23:53.080 -->
00:23:57.020I think it's really interesting to bring a lot more metrics to inside
229
00:23:57.400 -->
00:24:01.635dimensional modeling or call it entity centric data modeling to bring things like,
230
00:24:02.095 -->
00:24:09.809you know, 7 day visits and 7 day clicks and 7 day page views and 28 day to pivot these metrics inside these
231
00:24:10.190 -->
00:24:11.809entity centric datasets.
232
00:24:12.510 -->
00:24:13.409So that's happening
233
00:24:13.789 -->
00:24:17.865quite a bit in the fields of ML and, you know, historically in dimensional modeling.
234
00:24:18.805 -->
00:24:31.710Going back into like the metrics layer, I think to me it's a little bit of a misnomer because it's still like metrics are not useful without dimensions. Right? So it's like it's it's we still live in a world of, like, metrics and dimensions.
235
00:24:32.325 -->
00:24:34.904I guess now we're just looking to add, like,
236
00:24:35.284 -->
00:24:36.184more governance,
237
00:24:37.205 -->
00:24:40.825kinda construct and ideas, like, around the metric.
238
00:24:41.640 -->
00:24:52.505And then these higher level, like, less SQL and more YAML, there's, like, kind of this tweak of, like, oh, let's be more configuration driven and less, like, in a code or declarative, like, transformation, low level transformation.
239
00:24:53.365 -->
00:24:55.225Like, let's operate a higher level
240
00:24:55.605 -->
00:25:00.240a little bit, which I think is a great idea. Something we could we could talk a lot more,
241
00:25:00.560 -->
00:25:01.060about.
242
00:25:03.680 -->
00:25:13.765 Are you bored with writing scripts to move data into SaaS tools like Salesforce, Marketo or Facebook ads? Hightouch is the easiest way to sync data into the platforms that your business teams rely on.
243
00:25:14.385 -->
00:25:18.005The data you're looking for is already in your data warehouse and BI tools.
244
00:25:18.390 -->
00:25:25.049Connect your warehouse to Hightouch, paste a SQL query, and use their visual mapper to specify how data should appear in your SaaS systems.
245
00:25:25.725 -->
00:25:34.890No more scripts, just SQL. Supercharge your business teams with customer data using Hightouch for reverse ETL today. Get started for free at data engineering podcast.com/hitouch.
246
00:25:38.070 -->
00:25:44.475Yeah. To your point about operating at the higher level, 1 of the other interesting things that's been happening lately is the reemergence
247
00:25:44.855 -->
00:25:51.195of these visual pipeline builders and low code slash no code solutions where, you know, maybe
248
00:25:51.950 -->
00:26:01.09010, 15 years ago, it was the world of, you know, SQL Server Integration Studios and Pentaho, and everything was a drag and drop builder for defining your pipelines.
249
00:26:01.765 -->
00:26:07.625And then with the advent of airflow and the series of tools that followed it, they went back to
250
00:26:07.925 -->
00:26:18.029everything is software. So it was software defined pipeline, so you needed to be able to write code and reason about the flows and with the, you know, map reduce world of Hadoop. And
251
00:26:18.394 -->
00:26:23.135now we're starting to build back up to this higher level of you can, you know, take these
252
00:26:23.514 -->
00:26:40.895visual elements, drag and drop them together, but then you have a way to actually drop down into the code layer. So I think it was prophecy IO is 1 of the interesting entrances in that space where you have this visual mapper. But then when you want to actually dig in and maybe tweak things specifically, it actually generates
253
00:26:41.755 -->
00:26:49.600the spark code so that you can modify it yourself if you have sufficient knowledge. And so it's an interesting world where we have this kind of hybrid of
254
00:26:49.900 -->
00:26:54.560low code visual builders, but also the ability to drop down into the software level.
255
00:26:54.975 -->
00:26:57.715 Yeah. It's really interesting to see these cycles too.
256
00:26:58.175 -->
00:27:08.300I think both use cases are valid. I think, like, 1 realization, you know, as the person originally created Airflow is, like, the pipeline world is, like, too complicated
257
00:27:08.680 -->
00:27:10.620to kinda express inversion
258
00:27:11.080 -->
00:27:12.700and diff and review.
259
00:27:13.345 -->
00:27:16.085It's so complex that it has to be
260
00:27:16.465 -->
00:27:18.805represented as code at a certain level.
261
00:27:19.265 -->
00:27:25.559When you start getting into, like, those GUIs and you try to do, like, source control type things that
262
00:27:25.860 -->
00:27:32.664now are kind of a given. Right? Like, reviewing a pipeline, seeing what it looked like before and after and forking and testing
263
00:27:33.284 -->
00:27:37.144and CICD type things. Like, that stuff to me feels like
264
00:27:37.684 -->
00:27:43.770as you go up the level of complexity, the there's more need to be in that very kind of version control and
265
00:27:44.150 -->
00:27:44.970as code
266
00:27:45.430 -->
00:27:57.610environment. We also see, like, infrastructure as as code. Right? Like, it is also a movement and seems to be pretty well settled. There's tension there between, like, declarative and templated too. Right? Like, you expressed it as YAML.
267
00:27:57.910 -->
00:27:59.610And if so, like, I'll parameterize
268
00:27:59.990 -->
00:28:11.665it is and, you know, I've seen, like, a lot of YAML with a lot of logic in it, you know, to a point where it doesn't feel like a static declaration of anything. It's very much more like code.
269
00:28:12.270 -->
00:28:21.405So, yeah, there's that tension. You know? So to me, I like the idea of, like, being able to have it both ways. So if you could have, you know, the drag and dropiness of, say, Informatica
270
00:28:21.865 -->
00:28:23.085and code orientation
271
00:28:23.465 -->
00:28:26.525of something like Airflow and have bidirectional
272
00:28:26.825 -->
00:28:31.400workflow and being able to, you know, pivot from 1 to the other and vice versa.
273
00:28:32.100 -->
00:28:34.040Maybe that's the best of
274
00:28:34.500 -->
00:28:45.164all worlds, but if you can just add it 1 way, it probably should be code. Right? Like, I don't know. At a certain level of complexity, like these GUIs did just seem to break down pretty intricately.
275
00:28:45.800 -->
00:28:51.740 Absolutely. I think that, as you said, if you need to go 1 direction or the other, it should be software because
276
00:28:52.280 -->
00:29:10.250at a certain point, you can't express the necessary logic in these constrained environments without having a very long iteration cycle of needing to say, okay. Well, now I need to, you know, go in and define a completely new visual block with some different input types that will map to the specific use case that I have.
277
00:29:10.554 -->
00:29:25.840And then, you know, you end up with a proliferation of blocks that are very similar to each other with slight tweaks. And so then it's just a a different explosion of complexity where you'd be better off, you know, having just defined a function that it takes a few parameters and, you know, does these different conditional steps.
278
00:29:26.140 -->
00:29:34.485 I think the best 1 that I've seen in terms of the best GUI that I prefer is the Abenisho. It's not well known because it was very kind of special purpose and
279
00:29:34.785 -->
00:29:42.390I believe very, very expensive, but it was very good in the visual drag and drop realm. And you could go pivot from
280
00:29:42.850 -->
00:29:59.290code to visual to visual to code bidirectionally pretty well. And then the parallelism specification, like the way that you could monitor the pipeline as it executed was pretty great. It could see the flow visual flow of rows through intricate, you know,
281
00:29:59.750 -->
00:30:00.250transformation
282
00:30:00.550 -->
00:30:34.830phases. So it felt a little bit like when you look at a query execution plan, you know, from a complex, like, you know, parallel database, you can kinda see all the blocks and how the different phases of your query. They kinda just expose that as an API. So it could be I'm gonna have, like, a a group by, you know, and I'm gonna have a, you know, a parallelization phase with a round robin, so you could define all these things very, very well and visually. And for me, it helped me early in my career to think in parallel. It's just the fact of seeing it and seeing the rows flow through and seeing the declaring the parallel phases and the computation,
283
00:30:35.529 -->
00:30:38.029like, really helped me understand, like, data processing
284
00:30:38.835 -->
00:30:40.455on, like, distributed architecture
285
00:30:40.755 -->
00:30:43.255early on because the visuals were so great.
286
00:30:43.715 -->
00:30:46.054 Another element of the kind of guardrails,
287
00:30:46.470 -->
00:30:57.715and you hinted at it earlier, is the idea of data governance. And that has also gone through a few different shifts where, you know, earlier on, it was a very sort of process oriented manual
288
00:30:58.415 -->
00:31:07.010endeavor where you had to have the data dictionaries, and you said, you know, these are the different data owners, and, you know, maybe you had very coarse grained
289
00:31:07.549 -->
00:31:13.650access layers to say, you know, you can only access this dataset if you have this role in the LDAP system.
290
00:31:14.085 -->
00:31:15.065And now with
291
00:31:15.525 -->
00:31:16.025more
292
00:31:16.645 -->
00:31:18.585code driven and more flexible
293
00:31:19.285 -->
00:31:23.385data systems and layers on top of that, thinking in terms of things like
294
00:31:23.820 -->
00:31:27.040the introduction of tools like Immuta, which have more
295
00:31:27.900 -->
00:31:32.560data sort of attribute oriented access controls versus just role based access controls
296
00:31:33.034 -->
00:31:37.695and some more of these flexible metadata layers to be able to understand
297
00:31:37.995 -->
00:32:11.855as the data flows through the systems, these permissions need to flow with it and being able to do sort of just in time access control where somebody wants to query a given table, but it has maybe somebody's address in it so that you need to request access to it, and then that propagates to somebody else to say yes or no rather than having to, you know, go through a very manual process of trying to, you know, submit a request to the IT department, waiting for them to turn it around in a week or so before you could run the query, and at that point, you've forgotten what you were trying to figure out, you know, we can have these more flexible
298
00:32:12.635 -->
00:32:13.135processes
299
00:32:13.595 -->
00:32:16.500to manage data access so that people
300
00:32:17.279 -->
00:32:29.265are constrained and that they're not just gonna query everything if they don't have the necessary context or they don't have the necessary access. But in the cases where they do need to be able to run a query across a set of data,
301
00:32:29.565 -->
00:32:33.425they can do so. Another interesting element of that is some of the more
302
00:32:33.929 -->
00:32:34.429sophisticated
303
00:32:34.809 -->
00:32:39.309sort of data privacy algorithms and cryptographic algorithms to be able to actually
304
00:32:39.929 -->
00:32:48.785run queries on encrypted data without ever having to actually decrypt it in flight to be able to do aggregations, but you don't ever actually see, you know, the individual values.
305
00:32:49.165 -->
00:32:55.450I'm curious to get your thoughts on some of the more modern sort of data governance aspects of how
306
00:32:55.909 -->
00:33:03.544that plays into the data engineering role and some of the ways that that also manifests in this, sort of data democratization play?
307
00:33:03.845 -->
00:33:06.345 There's, like, many things to unpack here.
308
00:33:06.804 -->
00:33:08.1841 topic is
309
00:33:08.690 -->
00:33:09.510kinda inheritance
310
00:33:10.450 -->
00:33:15.429in the data schemas or data access policy. Right? So you you mentioned data governance, and to me there's
311
00:33:15.975 -->
00:33:20.715subfields there. 1 is data access policy, like who can access what,
312
00:33:21.095 -->
00:33:25.890and then there's data governance more like who created what, who owns what, who's responsible
313
00:33:26.350 -->
00:33:26.850for
314
00:33:27.150 -->
00:33:31.570say the change management, the SLAs around a certain dataset. So
315
00:33:31.950 -->
00:33:34.884let's get in to the more data access policy.
316
00:33:35.264 -->
00:33:37.365I think it's pointing in the direction that
317
00:33:37.745 -->
00:33:40.965if the database is aware of the dataset
318
00:33:41.370 -->
00:33:44.110graph, right, like, which column is coming from where,
319
00:33:44.490 -->
00:33:50.395then you can apply good inheritance kind of scheme into, like, data access policy, and there seem to be
320
00:33:50.875 -->
00:33:55.055value in that. It's interesting to see, like, with ELT kind of winning
321
00:33:55.435 -->
00:34:07.350pretty significantly. Right? Like, a lot of, like, the bulk of, like, the batch processing nowadays written in SQL and it's done by the database engine. So that means the database engine should be able to track the provenance of any given
322
00:34:08.015 -->
00:34:09.555column and dataset and
323
00:34:10.015 -->
00:34:15.795have some inheritance rules around that. Right? And then databases like Dremio, for instance, are baked. The transformations
324
00:34:16.890 -->
00:34:17.790and the derivatives
325
00:34:18.090 -->
00:34:19.470see the database is aware
326
00:34:19.770 -->
00:34:21.550of where things are coming from,
327
00:34:21.850 -->
00:34:24.350and I think there's a need for that. So that means
328
00:34:25.465 -->
00:34:29.805maybe, you know, over time, what we see is, like, the database engine
329
00:34:30.665 -->
00:34:32.765being very aware of the dataset
330
00:34:33.390 -->
00:34:34.210kinda semantic.
331
00:34:34.510 -->
00:34:40.289Right? Like, so whatever is done in something like dbt inside the database, the database knows about and
332
00:34:40.589 -->
00:34:41.809surfaces that information.
333
00:34:42.454 -->
00:35:01.420There's a lot you can do with that beyond just data access policy. Right? There's, like, aggregate awareness. Right? I could ask a question to the database and it would be like, hey. I know I have an aggregate that's fresh here that will better serve your query. So so they're they said that I know about that you don't know about that I'm gonna use to answer your query more efficiently.
334
00:35:02.015 -->
00:35:05.075So I think, like, we will see the rise of database
335
00:35:05.455 -->
00:35:12.900engines that are aware of the ELT or the transformation semantic and leverage that for all sorts of consideration.
336
00:35:13.760 -->
00:35:14.480And probably
337
00:35:15.119 -->
00:35:21.905is Dremio the only example I can think of that? I mean, Vertica kinda does that with projections, you know, Dremio with reflections,
338
00:35:22.625 -->
00:35:28.005but it's aware of different maybe perhaps, like, projections or different show of the same dataset.
339
00:35:28.465 -->
00:35:34.869That's 1 thought. I don't know if I wanna get again, there was 2 aspects of your question. I'm kinda tempted to go in the data governance.
340
00:35:35.569 -->
00:35:38.630That's more kinda ownership, validity, SLA,
341
00:35:38.930 -->
00:35:39.430SLO
342
00:35:40.130 -->
00:35:49.095type thing, but that's a very large unsolved problem right now that I think, like, a lot of the interest in data mesh currently
343
00:35:49.850 -->
00:35:50.590are around
344
00:35:51.290 -->
00:35:53.870the fact that the data mesh is talking about
345
00:35:54.570 -->
00:36:03.725data governance. Like, who owns what and what is private, and what is public in terms of datasets, and what's the API to the data warehouse, and what are the guarantees, like, you know,
346
00:36:04.105 -->
00:36:07.470treating data assets as and they call it data products. Right? Like,
347
00:36:07.930 -->
00:36:14.590treating your datasets like they're little products with little kinda API with dual binding contracts around them. So I think that's an interesting
348
00:36:15.205 -->
00:36:23.385area where we, you know, haven't figured out as an industry, you know, the the answers there. Interesting parallels in this area with, like, microservices
349
00:36:23.685 -->
00:36:28.910and the DevOps world of, like, you know, microservice is, like, a really kinda clear contract and service.
350
00:36:29.450 -->
00:36:34.430This question of, like, can could we have, like, datamarts or datasets, you know, some similar kinda
351
00:36:34.755 -->
00:36:38.455contracts and guarantees in the data world around sets of datasets.
352
00:36:39.235 -->
00:36:40.535 In terms of the actual
353
00:36:41.315 -->
00:36:42.935data engineering position,
354
00:36:43.820 -->
00:36:48.160as the usage of data has become more widespread, more data sources have become available,
355
00:36:48.540 -->
00:36:52.640it's easier to actually get a data infrastructure set up with all these different services.
356
00:36:53.195 -->
00:36:53.935It has
357
00:36:54.555 -->
00:37:07.840caused an increase in demand for data engineers, which has made it difficult for companies to be able to actually hire for it because there are so many opportunities out there. And I'm curious what your sense of the
358
00:37:08.380 -->
00:37:11.600sort of order of dependency has been in the
359
00:37:12.235 -->
00:37:19.055sort of rise of the modern data stack and the demand for data engineers as to which has driven the other more
360
00:37:19.835 -->
00:37:20.335prominently?
361
00:37:21.090 -->
00:37:22.390 I think on 1 front
362
00:37:23.010 -->
00:37:30.390with the analytics engineer, you know, as this materialized and if we can get enough people that have those skills of, like, the
363
00:37:30.715 -->
00:37:40.175analytical mindset and then kind of the curiosity of the data analysts and the for, like, someone who's, like, vertically aligned, right, that sits in a product team and wants to answer
364
00:37:40.619 -->
00:37:42.640to solve problems with data
365
00:37:43.180 -->
00:37:55.805while being able to write pipelines and check down source control and have decent kind of data engineering IG. And I think maybe that creates a new need that removes some of the load and the pressure to have so many data engineers.
366
00:37:56.345 -->
00:38:01.230And the fact that they're vertically aligned, I think their odds of succeeding is probably better
367
00:38:01.610 -->
00:38:05.470than a data engineer maybe trying to do that across verticals. So
368
00:38:05.770 -->
00:38:08.110that removes some of the pressure there.
369
00:38:08.535 -->
00:38:10.555I think, like, as any discipline
370
00:38:11.335 -->
00:38:11.835matures,
371
00:38:12.615 -->
00:38:17.035it becomes more the essence of itself. Right? So everything that is automatable
372
00:38:18.160 -->
00:38:18.980in the role
373
00:38:19.840 -->
00:38:21.060becomes, you know, served
374
00:38:21.760 -->
00:38:31.595by tooling and by practices. And what is left is the things that cannot, you know, be solved, like, with a single solution or, like, the kind of 1 size fits all type of solution.
375
00:38:32.055 -->
00:38:33.835Like, what does that mean in the world
376
00:38:34.535 -->
00:38:39.660of the data engineer? What's the essence of it once everything that is automatable
377
00:38:40.200 -->
00:38:40.780is automated.
378
00:38:41.480 -->
00:38:50.214I think there's, like, less and less left. Like, there's probably a page to read from the DevOps movement there too. Like, you know, there's still, like, very much a need for
379
00:38:50.595 -->
00:38:51.575DevOps engineers
380
00:38:52.674 -->
00:38:56.330even, like, kinda 10 years in to the DevOps move too.
381
00:38:56.710 -->
00:39:21.3191 question is, like, what are some of the things that every data engineer does that are gonna go away maybe in the next, you know, 5 years, like, what services are gonna pop up. So there's, like, these common patterns and data pipelines. Right? And then in the past, I've been talking about I call it, like, parametric pipelines, which is this idea, like, these higher level abstractions we were talking about a little bit earlier. So everyone does, like,
382
00:39:21.655 -->
00:39:22.155sessionization,
383
00:39:22.535 -->
00:39:27.835for instance, to provide answers around, you know, click stream analysis and segmentation.
384
00:39:28.135 -->
00:39:29.900Right? So we all do this stuff.
385
00:39:30.220 -->
00:39:33.440And then companies, as they mature up, they build their own
386
00:39:33.819 -->
00:39:41.884AB testing framework that computes, you know, p values, confidence intervals, and does all sorts of magic and complex computation
387
00:39:42.265 -->
00:39:42.765around,
388
00:39:43.144 -->
00:39:54.300you know, subjects and experiments and metric sets and all this stuff. So that's another 1. You know, I've seen people build and rebuild, you know, cohort analysis frameworks.
389
00:39:54.600 -->
00:39:57.020And I think all of these, we're gonna see
390
00:39:57.725 -->
00:40:02.065a company maybe or or people or open source projects or
391
00:40:02.525 -->
00:40:03.025abstractions
392
00:40:03.325 -->
00:40:14.609that help people solve these problems without having to reinvent the wheel so that every single company is, like, kind of building essentially a variation on the same theme. Like, I would love to see these abstractions
393
00:40:15.484 -->
00:40:18.625coming into existence so that next time I need to do sessionization,
394
00:40:18.925 -->
00:40:23.185I can just, like, you know, download the package and and solve that problem.
395
00:40:23.560 -->
00:40:37.225We're not quite there yet. Right? Like, I think, we might see, like, you know, airflow tags or airflow, like, libraries or, like, DBT projects as reference implementation. But we're in the phase where it's even hard to find
396
00:40:37.685 -->
00:40:39.599some good reference implementation
397
00:40:39.900 -->
00:40:46.559for the things I talked about. Right? If I go today, I'm like, I wanna write my AB testing framework or I wanna do some sessionization,
398
00:40:46.964 -->
00:40:48.345like, what are the resources?
399
00:40:48.805 -->
00:40:51.305You'll probably find some reference implementation
400
00:40:51.684 -->
00:40:56.505that if you're lucky, you might be able to reuse tiny portions of it
401
00:40:56.950 -->
00:41:11.365 and kind of bend into submission to get to where you need to be. Right? Yeah. And to your point of the selection of these different, you know, prebuilt packages, but also at the level of the different services that are being built and composed together,
402
00:41:11.905 -->
00:41:21.060you know, that has definitely become 1 of the responsibilities of the data engineer to say, okay. You know, do I wanna use Fivetran? Do I wanna use Stitch? Do I wanna use Meltano?
403
00:41:21.520 -->
00:41:33.225You know, which data warehouse do I need? There are multiple offerings for each of these different layers of the stack, and so a big part of it is just tool selection and integration of those systems. And I'm wondering
404
00:41:33.799 -->
00:41:37.019what you have found to be some of the useful strategies
405
00:41:37.400 -->
00:41:53.430for approaching that selection process and being able to understand how well each of those different layers integrates together, some of the potential edge cases that might come about where, you know, maybe I want to use Fivetran, but it doesn't work with Firebolt yet kind of a thing and being able
406
00:41:53.990 -->
00:42:03.875to discover some of those edge cases before you get too far down the road of trying to get it integrated and find out that it actually doesn't work yet. The part of beauty of the modern data stack,
407
00:42:04.175 -->
00:42:09.955 I think, you know, as we try to to define it, like 1 of the properties that we've seen is the pay as you go
408
00:42:10.570 -->
00:42:18.990and try at will or at least, like, try for cheap. So if it's pay as you go and you wanna do a proof of concept, you're able to self serve into that.
409
00:42:19.345 -->
00:42:21.445Where in the past, you might have to,
410
00:42:21.905 -->
00:42:23.445like, spin up some infrastructure
411
00:42:23.745 -->
00:42:29.410to do a POC or to pay or deal with a vendor process and have an official POC
412
00:42:30.110 -->
00:42:32.450approach. And then the POC become an institution
413
00:42:32.750 -->
00:42:35.090where, like, now you have to involve 3 or 4 vendors.
414
00:42:35.790 -->
00:42:44.165If you wanna do a horizontal or vertical kinda integration through it too, you would have to to involve multiple vendors for each layer and then align
415
00:42:44.465 -->
00:42:50.539them, and then the combinations just becomes really heavy. So at least, like, now I think you can go pretty easily
416
00:42:50.839 -->
00:42:52.059and try, you know,
417
00:42:52.359 -->
00:42:59.155if you could go from having nothing to having a pretty decent proof of concept with a kind of full stack integration pretty quickly.
418
00:42:59.695 -->
00:43:03.395So if you wanna try Fivetran today, I think it's pretty easy to get started
419
00:43:03.970 -->
00:43:06.150and to get some data starting to flow.
420
00:43:06.529 -->
00:43:09.589And similarly, I think with like reverse CTL or
421
00:43:09.890 -->
00:43:12.789some of these things that used to be very like non trivial.
422
00:43:13.405 -->
00:43:36.464Then you probably want to, you know, talk to peers, similar companies, like, you know, tap into the collective wisdom in terms of, like, for people that that are kinda like you, and then make sure it works for you. And, hopefully, you can get going. I think our our story with reverse CTL, that preset is we're like, hey. You know, do these tools like, we need to send data back to HubSpot, some product analytics back to HubSpot.
423
00:43:37.085 -->
00:43:42.570You know, how are we gonna do this? I was like, oh, let me just try 1. I'm just gonna try HiteTouch, and within
424
00:43:42.950 -->
00:43:47.610you know, it's, like, 25 minutes. I was connected to my database and sending data over,
425
00:43:47.985 -->
00:43:56.165and everything was working pretty well. And we only needed 1 integration, which fits under the pre the freemium plan, and we're like, okay. Well, problem solved. You know?
426
00:43:56.579 -->
00:43:57.880So build confidence
427
00:43:58.339 -->
00:44:06.200very quickly, and I think that's where the more old school vendors need to worry a little bit. It's like for this generation of people,
428
00:44:06.885 -->
00:44:14.425you know, we wanna self serve. We wanna run our own POC, and we wanna get, like, time to value down to, like, sub 1 hour,
429
00:44:15.049 -->
00:44:29.255and that's just not compatible with the more traditional sales cycle. Like, you gotta talk to someone, and they're gonna ask you a bunch of question. They're gonna qualify you, and they're not gonna be interested in selling you anything unless, like, your contract value is gonna be above, you know, 20 or $50, 000.
430
00:44:29.875 -->
00:44:35.280So I think they're gonna miss out on the more traditional vendors are gonna miss out on these, these opportunities.
431
00:44:35.980 -->
00:44:37.760Kinda so you're gonna sneak up on them.
432
00:44:38.140 -->
00:44:39.040 Another interesting
433
00:44:39.420 -->
00:44:45.265element of wordplay is the idea of the modern data stack has gained a lot of popularity
434
00:44:45.565 -->
00:44:50.145as well as the idea of building a data platform. And I'm wondering if you see those as
435
00:44:50.740 -->
00:44:51.880disjoint concerns
436
00:44:52.180 -->
00:45:00.315or something where you start with the modern data stack, and then you have to build the platform on top of it and some of the sort of skills and responsibilities
437
00:45:00.775 -->
00:45:01.755that are
438
00:45:02.455 -->
00:45:04.875implied in each of those phrases.
439
00:45:05.619 -->
00:45:13.319 I don't know what is the data platform and what is the modern data stack. They're, like, both a little bit unanswered. But, like, 1 way I would paint a picture
440
00:45:13.745 -->
00:45:15.605for me is, like, my data platform
441
00:45:16.385 -->
00:45:18.805at the start up that I'm part of is
442
00:45:19.105 -->
00:45:20.405the collection of
443
00:45:20.705 -->
00:45:28.440building blocks that we selected from the modern data stack and made work together with our business logic. Right? So we pick a a certain number of things,
444
00:45:28.820 -->
00:45:32.245invested in making them work together. They're all modern data stack.
445
00:45:32.645 -->
00:45:34.665I would say, like, building blocks.
446
00:45:35.045 -->
00:45:39.625And then our data platform is, like, the fabric or the mesh of services
447
00:45:40.210 -->
00:45:42.470and this logic that we built on top of it.
448
00:45:42.849 -->
00:45:48.295 Going back to what you're saying earlier about the role of templated SQL
449
00:45:48.835 -->
00:45:54.215and the current prominence that it has in the form of DBT and, a few other systems.
450
00:45:55.060 -->
00:46:01.240But as you were saying, we need some higher level constructs to be able to have appropriate guardrails and appropriate
451
00:46:01.994 -->
00:46:24.305kind of proofs around the validity of the workflows that we're trying to build where SQL is a little bit too free form because it's just text. You know, it's parsable. You can make some assumptions about it, but it's very easy to kinda shoot yourself in the foot without necessarily having some advanced warning of that fact. And I'm curious what you see as the long term viability
452
00:46:24.685 -->
00:46:28.170of tools like DBT and the idea of templated SQL
453
00:46:28.710 -->
00:46:41.950 as a core workflow and maybe some of the ideas that might succeed that is a more Is it a more fitted abstraction, maybe? Right? And so and there's a question as to whether, you know, dbt
454
00:46:42.490 -->
00:46:44.910or airflow or template SQL
455
00:46:45.370 -->
00:46:47.710can be the building block of
456
00:46:48.010 -->
00:46:54.885these higher level construct, and then I'd like to shoot that down. I think it's not. So I think, like, dbt or template SQL
457
00:46:55.345 -->
00:46:58.160is a great way, I would say, to express
458
00:46:58.700 -->
00:47:06.400ETL primitives. And by ETL primitives, I mean, like, you wanna source from a dataset, you wanna apply filters, you wanna group by, you wanna
459
00:47:06.795 -->
00:47:13.855join so that ETL primitives or data transformation primitives are these simple things that are very, very well expressed
460
00:47:14.395 -->
00:47:15.135in SQL.
461
00:47:15.790 -->
00:47:18.690And with a little bit of YAML in there or templating,
462
00:47:19.150 -->
00:47:21.089a little bit of Jinja and YAML
463
00:47:21.550 -->
00:47:22.130and prioritization,
464
00:47:22.430 -->
00:47:26.365you can achieve a lot, and it's great. I think that the rise of DBT
465
00:47:26.825 -->
00:47:30.765and by the way, like, I would say, like, airflow as templated, like, Jinja
466
00:47:31.224 -->
00:47:35.020baked into it very deeply too. So you're gonna achieve, like, very similar things
467
00:47:35.400 -->
00:47:45.765with Airflow. Right? So Airflow, I would say, is a superset of what you can do with dbt in many ways. Right? So you can also have, like, all these other operators and your SQL operators and
468
00:47:46.065 -->
00:47:49.200the Jinja templating. But I would say dbt
469
00:47:49.500 -->
00:47:51.760does a better job at, you know,
470
00:47:52.140 -->
00:48:00.285showing you exactly at just the subset of what you need if all you care about is you have a single data warehouse or using just templated SQL.
471
00:48:00.665 -->
00:48:04.685I think, like, dbt is just very elegant in terms of, like, coordinating
472
00:48:04.985 -->
00:48:05.965a lot of SQL
473
00:48:06.430 -->
00:48:09.010very, very well. It solves that in a very good way.
474
00:48:09.310 -->
00:48:14.030So now if you wanna build these higher level constructs, so let's just take 1 and we'll take
475
00:48:14.955 -->
00:48:21.855I don't know which 1 is the best 1. We could take, like, the AB testing framework. Right? So you can go and write an AB testing framework
476
00:48:22.490 -->
00:48:29.070in DBT today. Right? Like with YAML, you could say, like, go and define your your metrics, your metrics group,
477
00:48:29.575 -->
00:48:33.195where you have, like, your subjects, right, your user IDs
478
00:48:33.495 -->
00:48:37.915and all these metrics, and then what are your experiments and your exposure tables.
479
00:48:38.700 -->
00:48:42.160And you can go and build all of that. But then
480
00:48:42.619 -->
00:48:48.445what you're building is really hard to reuse for a variety of reasons. Like, 1 is that
481
00:48:49.005 -->
00:48:58.430as you become kinda logic heavy, you have a lot more Jinja than SQL, and then that just not as very expressive. Like, SQL with a lot of Jinja in it,
482
00:48:58.910 -->
00:49:08.985where every field list is a 4 loop on a collection of fields stored somewhere else. It's just, like, very hard to read and reason about, and it's not expressive enough
483
00:49:09.605 -->
00:49:21.420to do that well to have, like, these very dynamic pipelines. And then there's the other core issue, which is, like, d v 2 doesn't really solve you know, you're writing a certain dialect, so you're only solving the problem
484
00:49:22.040 -->
00:49:26.805for people who use the same dialect as you. Or if you're trying to say, like,
485
00:49:27.265 -->
00:49:34.680oh, you know, I'm gonna write something that works kinda cross SQL languages, then your template is gonna become even more overloaded
486
00:49:35.140 -->
00:49:38.840with Jinja. Right? So you would not use something like limit or.
487
00:49:40.075 -->
00:49:42.655You would use some sort of, like, more intricately
488
00:49:43.755 -->
00:49:46.850complex abstraction on top of it. So I think, like,
489
00:49:47.490 -->
00:49:52.630DBT doesn't seem like the right place to build these, like, higher level constructs.
490
00:49:53.170 -->
00:49:56.685Right? You know, maybe it's a great place to do a reference implementation
491
00:49:56.985 -->
00:49:58.525and say, like, I have this simple
492
00:49:59.065 -->
00:50:01.005dbt project where I do obsessionization.
493
00:50:01.625 -->
00:50:23.330I'm gonna share this in a GitHub repo and you can take it and reuse what you want and alter it to kinda fit your need, which might be the first phase. Like, we I think we need people doing that today so we can identify the patterns and share and talk about these things and have all these reference like, a a good library of reference implementation so people can compare and try things. So It's a good place to start.
494
00:50:23.710 -->
00:50:24.210Spark
495
00:50:24.670 -->
00:50:28.770maybe? It seemed like a better place to do some of these things, the way you can write these
496
00:50:29.445 -->
00:50:34.665more dynamic pipelines, it can be more dry. It's, like, more expressed as code.
497
00:50:35.205 -->
00:50:42.170It seems more like a natural place for some of these frameworks to be in a higher level construct to be expressed.
498
00:50:43.030 -->
00:50:46.010I don't know. There's a real question there of, like,
499
00:50:46.704 -->
00:50:48.805if you're trying to build these abstractions
500
00:50:49.105 -->
00:50:50.724today, right, reusable
501
00:50:51.105 -->
00:50:52.645kinda high level transformations,
502
00:50:53.809 -->
00:50:56.150And I called them parametric pipelines
503
00:50:57.010 -->
00:51:03.665or competition frameworks in the past. It's like kinda this idea of these higher level construct that solves certain, like, data engineering
504
00:51:04.204 -->
00:51:05.025high level
505
00:51:05.964 -->
00:51:13.359challenges. Like, what's the right tool set if you're trying to build, like, a 1 size fits all solution or reusable component that all companies can use
506
00:51:13.980 -->
00:51:15.039to solve these problems?
507
00:51:15.500 -->
00:51:20.704I don't know. I think I'd use Spark as probably what I would try to use if I was to work on that.
508
00:51:21.085 -->
00:51:26.464Does that work for everyone? Like, does everyone has a Spark cluster or is able to run a Spark workload?
509
00:51:26.990 -->
00:51:33.250Does it make sense for people to get data out their warehouse to compute it somewhere else and send it back in the ELT heavy world?
510
00:51:33.870 -->
00:51:35.970Maybe. I don't know. It's unclear.
511
00:51:36.585 -->
00:51:40.045 Just put everything into, Delta Lake, and then problem solved.
512
00:51:41.705 -->
00:51:45.005 Put into Lake and, let people write, MapReduce
513
00:51:45.305 -->
00:51:51.300to solve it, and and we're done. And that's the way we used to do it a long time ago. We're 20 years ago. Yeah. Yeah.
514
00:51:52.240 -->
00:51:57.375But, yeah, I mean, I would love to see a lot more of these, like, reference implementation, people sharing, like, hey.
515
00:51:57.675 -->
00:51:58.494This is how
516
00:51:58.955 -->
00:52:06.930this team at this company solved you know, build a core analysis framework on top of airflow. And here's, you know, things you
517
00:52:07.550 -->
00:52:10.849might wanna try to reuse, right, and alter and and
518
00:52:11.244 -->
00:52:12.384make sense of.
519
00:52:12.765 -->
00:52:15.505And you could have, like, more people sharing more
520
00:52:15.964 -->
00:52:16.944of these things.
521
00:52:17.565 -->
00:52:27.559I think it would become more clear what the different variation of on that topic are and make make it easier for someone eventually to solve that problem once and for all for for everyone.
522
00:52:28.155 -->
00:52:37.455 There are a number of other sort of hot take topics that it would be fun to dig into. Maybe we'll have to slate those for another interview to go a little deeper on them. So
523
00:52:37.800 -->
00:52:39.180I guess just quickly,
524
00:52:39.640 -->
00:52:41.900in your work participating
525
00:52:42.200 -->
00:52:49.045and contributing to the data ecosystem, what have been some of the interesting or unexpected or challenging lessons that you've learned in the process?
526
00:52:49.505 -->
00:52:53.445 So 1 thing that's interesting is to see how there's these cycles.
527
00:52:54.210 -->
00:52:56.390And if you've been around long enough
528
00:52:56.690 -->
00:53:02.950in any given discipline, you'll see getting new people come in and have a fresh take on these old problems
529
00:53:03.635 -->
00:53:16.940without having necessarily the context of some of the failures in the past. I think that there's both, like, a beauty in that. Right? That kind of the innocence of giving an old problem a completely new shot with a new environment
530
00:53:17.240 -->
00:53:18.700and then you set of
531
00:53:19.385 -->
00:53:24.365maybe tools and and solution. Right? The world has changed, so you don't think the right way.
532
00:53:24.665 -->
00:53:29.790And I'm sure you think the same way about the problem and can get really creative and fresh ideas.
533
00:53:30.250 -->
00:53:36.910There's also on the other side, the stupidity of kind of missing out and kind of this teenage, like, not being able to
534
00:53:37.285 -->
00:53:41.865leverage previous experiences of this innocence of, like, not missing out on
535
00:53:42.245 -->
00:53:43.545learning from previous
536
00:53:44.005 -->
00:53:46.025achievements and learning and struggles.
537
00:53:46.610 -->
00:53:50.790So it's interesting to be that person, something that points out to, you know,
538
00:53:51.330 -->
00:53:58.454technologies or mythologies that, you know, were born or existed, you know, 10, 15, 20 years ago that
539
00:53:58.994 -->
00:54:02.375went pretty far. Like, in some cases, like, there's some of these efforts
540
00:54:02.914 -->
00:54:11.160are very notable and solve not necessarily the same set of problems in the same way, but sometimes they would have take optimized for a different
541
00:54:11.540 -->
00:54:21.385kinda outcome or a different facet of the problem and, like, much better on that facet than what we're doing now. So it's been really interesting to see, could everyone
542
00:54:21.790 -->
00:54:23.730rebuild everything on new premises?
543
00:54:24.030 -->
00:54:31.010Like, you know, everything's gotta be on the cloud and everything is as a service and everything is as pay as you go and everything is distributed first.
544
00:54:31.345 -->
00:54:35.444But in terms of, like, user experience and some of the expressivity
545
00:54:35.825 -->
00:54:38.724of how we solve the problem, there's, like, shortcomings
546
00:54:39.025 -->
00:54:42.100on that side as we optimize for new kind of constraints.
547
00:54:42.640 -->
00:54:44.980 To close out the show as the final question,
548
00:54:45.680 -->
00:54:47.860in episode 3, when we
549
00:54:48.195 -->
00:54:50.455took a crack at defining data engineering,
550
00:54:51.075 -->
00:54:57.815we closed out with some predictions for the following years of what would come for the data engineering role.
551
00:54:58.260 -->
00:55:02.200And many of those have actually been proven out pretty well. So you're very prescient in that.
552
00:55:02.740 -->
00:55:06.775So now that we're kind of recapping some of those ideas and the
553
00:55:07.795 -->
00:55:13.575definition of data engineering, I'm interested in what your next set of predictions are for the upcoming years. I think there's a there's a question around, like, how are data engineers gonna
554
00:55:16.260 -->
00:55:30.805 work with analytics engineers. And that's a similar question, I think, to, like, what does a DevOps specialist like, how do they work with developers or engineers elsewhere in the company? And, you know, it's kinda transfer of, like, the vertically aligned versus horizontally aligned.
555
00:55:31.105 -->
00:55:34.100But I think on the short term, we're gonna see a little bit of a struggle
556
00:55:34.480 -->
00:55:39.380and tension and and kinda identifying, like, the border between the 2 roles.
557
00:55:39.760 -->
00:55:48.265And maybe the data engineer is gonna feel like they're hurting kind of a little bit more reckless analytics engineers that, you know, they wanna solve business problems first.
558
00:55:48.610 -->
00:55:53.990They're maybe oriented a little bit more short term, and they don't care about performance and costs and,
559
00:55:54.290 -->
00:55:59.654like, naming conventions and best practices and hygiene. Right? So we're gonna see some tension there form
560
00:56:00.194 -->
00:56:02.615until we can create, like, all of
561
00:56:03.075 -->
00:56:15.770the tooling and the rules and the guidelines of the best practices that are required for to make sure to keep these people in check and make sure that they're not, you know, accumulating depth as they solve, problems in their respective particles.
562
00:56:16.605 -->
00:56:27.800 Alright. Well, thank you very much for taking the time again today to talk through the sort of current definition of data engineering. So appreciate all the time and energy that you have put into
563
00:56:28.180 -->
00:56:30.920contributing to the data ecosystem and your continued
564
00:56:31.305 -->
00:56:49.710sort of thought leadership, if you will. So always a pleasure to have you on the show. Definitely have to have you back again sometime. So thank you again for all of that, and I hope you enjoy the rest of your day. It's been a pleasure, and I know there was a lot more questions on your list that we did not cover. So happy to come back on the show at some point and then push the conversation forward.
565
00:56:55.365 -->
00:56:58.265Listening. Don't forget to check out our other show, podcast.init
566
00:56:59.010 -->
00:56:59.510atpythonpodcast.com
567
00:57:00.849 -->
00:57:05.190to learn about the Python language, its community, and the innovative ways it is being used.
568
00:57:05.569 -->
00:57:16.750And visit the site of data engineering podcast dotcom to subscribe to the show, sign up for the mailing list, and read the show notes. If you've learned something or tried out a project from the show, then tell us about it. Email host atdataengineeringpodcast.com
569
00:57:18.250 -->
00:57:23.549with your story. And to help other people find the show, please leave review on Itunes and tell your friends and coworkers.