WEBVTT
NOTE
Transcription provided by Podhome.fm
Created: 7/6/2024
12:35:46 PM
Duration: 2565.710
Channels: 1
1
00:00:13.955 -->
00:00:18.055 Hello, and welcome to the Data Engineering Podcast, the show about modern data management.
2
00:00:18.710 -->
00:00:29.085Have you ever woken up to a crisis because a number on a dashboard is broken and no 1 knows why? Or sent out frustrating Slack messages trying to find the right dataset? Or tried to understand what a column name means?
3
00:00:29.865 -->
00:00:34.685Our friends at Outland started out as a data team themselves and faced all this collaboration chaos.
4
00:00:35.110 -->
00:00:37.850They started building Outland as an internal tool for themselves.
5
00:00:38.870 -->
00:00:44.795Outland is a collaborative workspace for data driven teams, like GitHub for engineering or Figma for design teams.
6
00:00:45.675 -->
00:00:55.220By acting as a virtual hub for data assets ranging from tables and dashboards to SQL snippets and code, Atlan enables teams to create single source of truth for all of their data assets
7
00:00:55.680 -->
00:01:01.780and collaborate across the modern data stack through deep integrations with tools like Snowflake, Slack, Looker, and more.
8
00:01:02.535 -->
00:01:03.434Go to dataengineeringpodcast.com/outland
9
00:01:05.814 -->
00:01:09.994today. That's a t l a n, and sign up for a free trial.
10
00:01:10.600 -->
00:01:13.820If you're a data engineering podcast listener, you get credits worth
11
00:01:14.200 -->
00:01:15.659$3, 000 on an annual subscription.
12
00:01:17.615 -->
00:01:24.755When you're ready to build your next pipeline and want to test out the projects you hear about on the show, you'll need somewhere to deploy it. So check out our friends over at Linode.
13
00:01:25.650 -->
00:01:34.229With our managed Kubernetes platform, it's now even easier to deploy and scale your workflows or try out the latest Helm charts from tools like Pulsar, Pacaderm, and Dagster.
14
00:01:35.065 -->
00:01:42.285With simple pricing, fast networking, object storage, and worldwide data centers, you've got everything you need to run a bulletproof data platform.
15
00:01:43.240 -->
00:01:44.859Go to data engineering podcast.com/linode
16
00:01:46.359 -->
00:02:14.284today. That's l I n o d e, and get a $100 credit to try out a Kubernetes cluster of your own. And don't forget to thank them for their continued support of this show. Your host is Tobias Macy. And today, I'm interviewing Satish Jayanti about how organizations can use data architectural patterns to stay competitive in today's data rich environment. So, Satish, can you start by introducing yourself? Thank you for having me on this. My name is, Satish Jayanti. I'm 1 of the cofounders of Coalesce,
17
00:02:15.064 -->
00:02:18.989 and I currently play the chief technology officer role, in the company.
18
00:02:19.450 -->
00:02:23.069 And do you remember how you first got started working in the area of data?
19
00:02:23.450 -->
00:02:24.349 Yes. Absolutely.
20
00:02:24.890 -->
00:02:31.155I have started my career as an, you know, application programmer, dabbled with that, and soon became a DBA,
21
00:02:31.615 -->
00:02:33.395database administrator by accident.
22
00:02:33.855 -->
00:02:34.755I was responsible
23
00:02:35.920 -->
00:02:41.860to run the, you know, database servers on a regular basis, to make sure the business is running smoothly.
24
00:02:42.400 -->
00:02:43.460This is for an
25
00:02:43.845 -->
00:02:44.345online
26
00:02:44.805 -->
00:02:45.944e learning platform
27
00:02:46.325 -->
00:02:46.825startup
28
00:02:47.204 -->
00:02:48.105in Los Angeles.
29
00:02:48.405 -->
00:02:55.420And because it's a startup, I was kind of playing many, many roles as you can expect in a startup.
30
00:02:56.040 -->
00:02:58.460And 1 of the things that I was doing was
31
00:02:59.325 -->
00:03:00.065I was writing
32
00:03:00.365 -->
00:03:00.865and
33
00:03:01.165 -->
00:03:01.665providing
34
00:03:02.125 -->
00:03:02.625insights,
35
00:03:02.925 -->
00:03:08.380like writing queries and generating reports for the business as part of my b b a role as well.
36
00:03:09.340 -->
00:03:15.520And it got to a point where it was just not sustainable. The amount of request that I was getting
37
00:03:15.875 -->
00:03:18.055and the amount of work that I have to do
38
00:03:18.355 -->
00:03:28.930to put something together and give it to business. These were, like, some basic questions like, hey, how many people are, you know, using this particular course? Or what are the top 10 courses? Things like that.
39
00:03:29.390 -->
00:03:29.890And
40
00:03:30.430 -->
00:03:33.490I questioned myself. Like, there must be a better way to do this.
41
00:03:33.925 -->
00:03:34.584And that's
42
00:03:34.965 -->
00:03:35.465when
43
00:03:35.845 -->
00:03:39.545I had my first encounter with the concept of data warehousing.
44
00:03:40.325 -->
00:03:42.185So I picked up Ralph Kimball's
45
00:03:42.960 -->
00:03:44.180data warehouse toolkit,
46
00:03:44.880 -->
00:03:46.340read it many, many times.
47
00:03:46.720 -->
00:03:47.780It was very interesting.
48
00:03:48.320 -->
00:03:50.955And then implemented my first data mark.
49
00:03:51.355 -->
00:04:05.080Then that was, like, a big light bulb for me at that time. And that's how I got into do that and continue to build a lot of data warehouses, data marts, and eventually also manage some groups of, you know, data professionals
50
00:04:05.700 -->
00:04:07.400and so on for several companies.
51
00:04:08.114 -->
00:04:23.389 And so that brings us now to where you are today at Coalesce. I'm wondering if you can share a bit about what it is that you're building there and some of the story behind how it came to be and why this is the problem space that you wanted to spend your time and energy on. Yeah. Absolutely. So
52
00:04:23.765 -->
00:04:28.345 in my, you know, several years of data warehousing and data mart and data analytics
53
00:04:28.645 -->
00:04:29.145experience,
54
00:04:30.085 -->
00:04:32.985the main challenge, it was always data transformations
55
00:04:33.320 -->
00:04:35.340for me. It was pretty clear that
56
00:04:35.800 -->
00:04:36.700we were spending
57
00:04:37.320 -->
00:04:38.300a lot of time
58
00:04:38.760 -->
00:04:41.660to take the raw data and change it to
59
00:04:42.235 -->
00:04:47.215a form that is useful and that can be consumed for decision making.
60
00:04:48.074 -->
00:04:48.574So
61
00:04:49.010 -->
00:04:51.670when I was leading a group in a financial firm,
62
00:04:52.450 -->
00:04:57.190we were building data warehouses. We have all the tools because people were like, we were
63
00:04:57.705 -->
00:04:58.685acquiring companies,
64
00:04:59.065 -->
00:05:00.685so it was growing really fast.
65
00:05:01.225 -->
00:05:05.550And we had a big ETL team and pretty much any tool that you can think of.
66
00:05:06.030 -->
00:05:08.770But we were still unable to keep up with the demand.
67
00:05:09.389 -->
00:05:14.610And then at that time, I came across this concept from a company, especially it's it's called Warescape.
68
00:05:15.085 -->
00:05:16.865That was my first encounter of
69
00:05:17.325 -->
00:05:18.705data warehouse automation.
70
00:05:19.245 -->
00:05:22.065The whole idea is there's so many patterns
71
00:05:22.599 -->
00:05:23.259in data
72
00:05:23.800 -->
00:05:26.139warehousing and so many mundane tasks.
73
00:05:26.759 -->
00:05:28.220You know, how can you
74
00:05:28.680 -->
00:05:31.900automate those things in a way that you make the
75
00:05:32.205 -->
00:05:33.665engineers very productive.
76
00:05:34.285 -->
00:05:36.785And it's just not 1 thing, but it's those,
77
00:05:37.405 -->
00:05:47.110you know, opportunities wherever you can automate from, you know, 1 end to the other, the entire data data warehouse life cycle. And the aggregation of those automations collectively
78
00:05:47.490 -->
00:05:47.990will
79
00:05:48.325 -->
00:05:49.465make you more productive.
80
00:05:50.005 -->
00:05:51.145So that was the concept.
81
00:05:51.605 -->
00:05:58.650And I was hooked on to that and, you know, I implemented it and saw really, like, a lot of benefit from that.
82
00:05:59.430 -->
00:06:07.215You know, when that company got acquired and I moved on, and that concept was what stuck in my mind in that was a legacy product.
83
00:06:07.675 -->
00:06:09.135So we took that concept
84
00:06:09.915 -->
00:06:12.095and build it for the modern data stack.
85
00:06:12.395 -->
00:06:17.120That's how we got here. And my cofounder and I, we've worked for that company
86
00:06:17.660 -->
00:06:18.160implementing
87
00:06:18.539 -->
00:06:20.639large data warehouses for large companies
88
00:06:21.099 -->
00:06:22.160with great results.
89
00:06:22.620 -->
00:06:28.155There are a lot of drawbacks in their solution, and we saw that. And we found an opportunity to
90
00:06:28.695 -->
00:06:33.340modernize and build it for the modern stack. But the core idea was still automation
91
00:06:33.640 -->
00:06:35.260and automating data transformation,
92
00:06:35.800 -->
00:06:37.820which we think is still
93
00:06:38.505 -->
00:06:41.645not automated. There's a lot of other areas like the database
94
00:06:42.105 -->
00:06:45.805platforms now. Snowflake has automated that data acquisition.
95
00:06:46.185 -->
00:06:51.490Like, if you look at Fivetran, it's doing a great job there. However, when it comes to data transformation,
96
00:06:52.190 -->
00:06:55.715it's still right for automation, I would say. And
97
00:06:56.254 -->
00:06:57.235 the most direct
98
00:06:57.535 -->
00:06:58.035comparison
99
00:06:58.495 -->
00:07:05.370that comes to mind right now is obviously the work that the folks at DBT are doing. And I'm curious if you can speak to some of the
100
00:07:05.910 -->
00:07:12.729overlap and maybe potential coexistence of what you're building at Coalesce as compared to where the focus of DBT is.
101
00:07:13.105 -->
00:07:16.325 So what we have seen in the last, you know, few years,
102
00:07:16.785 -->
00:07:17.765if you go back
103
00:07:18.225 -->
00:07:22.325several years, you'll see that people were at the beginning. They were just hand coding.
104
00:07:22.990 -->
00:07:24.290And that's a lot of work.
105
00:07:24.670 -->
00:07:29.650And then they said, okay, let's do something graphical. Then the ETL tools were born.
106
00:07:30.255 -->
00:07:34.995The ETL tools are graphical tools that you know, graphic the GUI based tools.
107
00:07:35.375 -->
00:07:38.035They would give you a lot of efficiency.
108
00:07:39.370 -->
00:07:44.110Pretty much with some training could use that. So it's all like widget based drag and drop
109
00:07:44.650 -->
00:07:46.030data pipeline development.
110
00:07:46.665 -->
00:07:50.445However, the problem was, you know, when you go out of
111
00:07:50.985 -->
00:07:52.045its boundaries
112
00:07:52.425 -->
00:07:53.645and you have this
113
00:07:54.080 -->
00:07:55.300special use case,
114
00:07:55.840 -->
00:08:03.699then you have to resort to leading the tool, go out, and kind of do something like a short procedure or something in the database itself.
115
00:08:04.305 -->
00:08:07.205So that was a limitation. They were pretty inflexible.
116
00:08:08.065 -->
00:08:09.445So what happened is,
117
00:08:10.065 -->
00:08:11.845you know, the whole industry kinda,
118
00:08:12.530 -->
00:08:15.190you know, took a 1 80 degree turn and went
119
00:08:15.570 -->
00:08:16.630everything as code.
120
00:08:17.169 -->
00:08:21.190And that's what DBT is. You know, everything is core. Now
121
00:08:21.685 -->
00:08:27.065it gives you a lot of flexibility. Of course, core is the most flexible thing. You can write anything you want.
122
00:08:27.365 -->
00:08:27.865However,
123
00:08:28.325 -->
00:08:30.900the cost of that is you lose the efficiency.
124
00:08:31.520 -->
00:08:33.380It's how we see it. Now
125
00:08:33.920 -->
00:08:35.620with everything as core paradigm,
126
00:08:36.080 -->
00:08:40.755you need, you know, highly skilled people in the organization, especially large organizations.
127
00:08:41.295 -->
00:08:43.075It's gonna be hard to
128
00:08:43.695 -->
00:08:46.275have that many, you know, highly skilled
129
00:08:46.680 -->
00:08:51.580data engineers given that it's so hard to find it engineers these days. And on top of that,
130
00:08:51.960 -->
00:08:53.340because you don't have efficiency,
131
00:08:53.915 -->
00:08:56.735you're gonna be coding a lot and still
132
00:08:57.355 -->
00:09:01.215not able to accomplish in a certain and keep up the demands.
133
00:09:01.839 -->
00:09:03.700So what we think is
134
00:09:04.160 -->
00:09:07.380a solution that has best of the both worlds
135
00:09:07.839 -->
00:09:12.545is what is needed, and that's what we are. We are the solution that,
136
00:09:13.005 -->
00:09:15.665you know, you can do 80% of the work
137
00:09:16.205 -->
00:09:19.080GUI because it does give you a lot of productivity.
138
00:09:19.459 -->
00:09:25.080There's a lot of patterns that can be automated. There is no reason why I should be coding the same thing over and over.
139
00:09:25.495 -->
00:09:52.535 And when it comes to core cases, that's when I'll focus on the coding aspects of it. So that's how you get the the results that you need on time. To your point about the 1st generation of ETL tools and the drag and drop workflow builders is that when you do hit the edges, you're kind of left to your own devices, and you have to figure out how do I build some additional component that I can somehow jam into this GUI builder and get them to work together.
140
00:09:53.140 -->
00:10:06.165And I'm curious if you can talk to some of the escape hatches that you've built into Coalesce for being able to move from that initial process of here's the rough workflow. This is 80% of what I need, but now I actually need to dig in
141
00:10:06.545 -->
00:10:10.165and customize this to fit my specific use case and being able to
142
00:10:10.520 -->
00:10:16.220have that be an affordance in the system rather than something that you have to fight against the system to achieve?
143
00:10:16.680 -->
00:10:20.685 1 of the things, again, you know, we wanted to build it in a way that
144
00:10:21.465 -->
00:10:22.445we can provide
145
00:10:23.305 -->
00:10:27.16580% of the solution out of the box and easy to use by anybody.
146
00:10:27.580 -->
00:10:32.560What that means is we kinda guide, you know, the user in a certain direction
147
00:10:33.100 -->
00:10:36.320in building a pipeline. There is a certain flow to it.
148
00:10:36.865 -->
00:10:44.005Basically, we call it, you know, we call it the graph. Everybody calls it it a graph, whatever you're building as a pipeline. Each 1 is a node.
149
00:10:44.529 -->
00:10:46.790And these nodes have certain configurations
150
00:10:47.330 -->
00:10:49.510and certain behavior. Right?
151
00:10:50.130 -->
00:10:54.425How to create or how to materialize a particular object on Snowflake,
152
00:10:54.725 -->
00:10:59.865or how do you load that object? If it's a table, how do you load Versa logic to load DML, basically?
153
00:11:00.405 -->
00:11:04.490What we have done is we have built these components as Lego
154
00:11:04.870 -->
00:11:05.530blocks, just
155
00:11:05.990 -->
00:11:10.570at a very granular level. So you can assemble these things on your own
156
00:11:10.875 -->
00:11:14.654to build a different kind of note pad that fits a certain pattern.
157
00:11:15.514 -->
00:11:19.135And as an architect, you can build these user defined nodes
158
00:11:19.730 -->
00:11:23.510and kind of, you know, meet that or address those edge cases.
159
00:11:24.770 -->
00:11:28.790So when people start off, they start off with a whole bunch of nodes that are available.
160
00:11:29.385 -->
00:11:30.125For example,
161
00:11:30.425 -->
00:11:31.485type 2 dimensions.
162
00:11:31.945 -->
00:11:35.405And out of the box, you don't have to think about it. You just go use it.
163
00:11:35.785 -->
00:11:43.420But if you say, I don't want this to behave this way. I wanna do some changes. Like, maybe I don't want to use surrogate key. I wanna use hash keys.
164
00:11:43.800 -->
00:11:48.495Then you go behind the scenes. You go into that note type. You make some minor adjustments.
165
00:11:48.955 -->
00:11:54.575You have a new note type. Everything is like a Lego block that you can control and configure
166
00:11:54.955 -->
00:11:57.210and work with. So that's the idea
167
00:11:57.830 -->
00:12:00.330 here. 1 of the challenges with warehouses
168
00:12:01.030 -->
00:12:03.210has often been SQL and
169
00:12:03.605 -->
00:12:07.865the fact that it is very declarative and flexible, but not always very composable.
170
00:12:08.325 -->
00:12:12.825And I'm curious how you have approached that challenge of being able to
171
00:12:13.550 -->
00:12:15.650encapsulate these nodes so that
172
00:12:16.270 -->
00:12:29.805the handoff between them is as pluggable as you want it to be so that you can combine them into these workflows without having to worry about how the underlying SQL is actually going to mesh together and what the sort of contract is between these different stages of the workflows.
173
00:12:30.560 -->
00:12:48.860 The way that it works right now is there are several I mean, you can build as many stages as you want in the pipeline, and you can have a raw layer, which is basically the raw data that's coming in. And then you can build a CDC layer, for example, to capture the deltas of what is being, you know, loaded by
174
00:12:49.400 -->
00:12:50.860a data ingestion system.
175
00:12:51.320 -->
00:12:56.425You can have a staging layer, which is basically now we can materialize them as views or tables.
176
00:12:57.225 -->
00:13:06.520But it's all happening in Snowflake. The data is moving from 1 layer to the other in Snowflake. So whether it's views or a set of tables.
177
00:13:07.060 -->
00:13:11.800But today, it's all SQL. Now you can change that because we are we are giving you templates.
178
00:13:12.420 -->
00:13:14.985There is no reason why you can't
179
00:13:15.445 -->
00:13:18.425generate a different type of code other than SQL
180
00:13:18.805 -->
00:13:22.745in the tool. You know, it's you have full control to override the template
181
00:13:23.070 -->
00:13:25.410and generate, for example, some other language
182
00:13:25.790 -->
00:13:29.795as long as Snowflake has the native capabilities to do so.
183
00:13:30.115 -->
00:13:35.575And we're seeing more and more of that where Snowflake is supporting all these other paradigms in the platform.
184
00:13:36.035 -->
00:13:48.160So the handshake or the flow can be pretty much customizable in in the way that you want. Today, we have only SQL support. So it has to go through table to table to view or view, you know,
185
00:13:48.714 -->
00:13:54.735however, you know, SQL functions. But I can see that the handshake could change down the road depending on what's not their problem.
186
00:13:57.720 -->
00:14:01.660 Are you looking for a structured and battle tested approach for learning data engineering?
187
00:14:01.960 -->
00:14:05.980Would you like to know how you can build proper data infrastructures that are built to last?
188
00:14:06.305 -->
00:14:11.764Would you like to have a seasoned industry expert guide you and answer all of your questions? Join Pipeline Academy,
189
00:14:12.065 -->
00:14:23.330the world's first data engineering boot camp. Learn in small groups with like minded professionals for 9 weeks part time to level up in your career. The course covers the most relevant and essential data and software
190
00:14:23.725 -->
00:14:28.385topics that enable you to start your journey as a professional data engineer or analytics engineer.
191
00:14:28.765 -->
00:14:35.260Plus, they have ask me anythings with world class guest speakers every week. The next cohort starts in April of 2022.
192
00:14:36.120 -->
00:14:36.620Visitdataengineeringpodcast.com/academytodayandapplynow.
193
00:14:42.435 -->
00:14:57.690In terms of the overall workflow, looking at the site and through the documentation, it seems to be fairly opinionated. And I'm curious if you can talk to the design principles and philosophies that you have embedded into the user experience and how you make decisions about
194
00:14:58.025 -->
00:15:01.005where to prioritize features and how to
195
00:15:01.705 -->
00:15:05.485present the different capabilities of the system in a manner that's
196
00:15:05.945 -->
00:15:07.005internally cohesive?
197
00:15:07.639 -->
00:15:13.899 Again, it goes back to our philosophy of, hey, 80% of this can be automated. And the design principle is
198
00:15:14.575 -->
00:15:17.075you got to be very easy to use.
199
00:15:17.775 -->
00:15:20.355And that's number 1, you know, since
200
00:15:20.895 -->
00:15:27.810we are bringing best of the both worlds here. We are saying, hey, you have to have the flexibility. But at the same time, you want to have the efficiency.
201
00:15:28.350 -->
00:15:30.130In order for it to be efficient,
202
00:15:30.875 -->
00:15:33.615you need to kinda interact with the tool pretty easily.
203
00:15:34.155 -->
00:15:35.695And also the personas
204
00:15:36.395 -->
00:15:39.295that are going to be working with this tool, it also varies
205
00:15:40.010 -->
00:15:42.350depending on their experience. Right? If I'm an architect,
206
00:15:42.970 -->
00:15:46.910my experience should be that I should be able to go and set standards
207
00:15:47.505 -->
00:15:59.660so that junior engineers can just consume those standards without even thinking about them. If there's, you know, extensibility that needs to happen, I can do that as well. I can set like, extend the product behavior
208
00:16:00.120 -->
00:16:10.005because there is a new feature or new like, something that came out in Snowflake that we want to support. Now you go create a new node, and then you make that available. That can be consumed by data engineers.
209
00:16:10.460 -->
00:16:12.880Now as a data engineer, the experiences could
210
00:16:13.260 -->
00:16:16.960be different. If you're a junior engineer, you may just want to go build pipelines
211
00:16:17.420 -->
00:16:19.840based on the standards set by my architect.
212
00:16:20.154 -->
00:16:21.855So we wanna make sure that that is
213
00:16:22.315 -->
00:16:26.095possible and that they are getting the productivity that they are expecting out of this tool.
214
00:16:26.475 -->
00:16:29.774But on the other hand, if you're a data analyst, like a business analyst
215
00:16:30.220 -->
00:16:31.680building dashboards reports,
216
00:16:32.140 -->
00:16:33.520for them, it's all about
217
00:16:33.900 -->
00:16:35.440understanding what was built
218
00:16:35.820 -->
00:16:40.725and why it was built the way it was built. Like, what is this dimension mean? What does this column mean?
219
00:16:41.264 -->
00:16:42.245How do I
220
00:16:42.865 -->
00:16:50.960understand or how do I know that this data is correct and where is it coming from? So for them, the experience is all about
221
00:16:51.420 -->
00:16:51.920documentation,
222
00:16:52.860 -->
00:16:53.360lineage,
223
00:16:54.105 -->
00:16:58.365understanding what was built. Because there is no data project where
224
00:16:58.825 -->
00:17:04.940you can just kind of remove these data persona like, professionals or personas from at the end of the day, in the real
225
00:17:05.320 -->
00:17:09.500world, all of these people have to come together to make a data project successful.
226
00:17:10.075 -->
00:17:11.695So we are making sure that
227
00:17:11.995 -->
00:17:13.615these people have the right experience
228
00:17:14.235 -->
00:17:26.200 for what they're doing in the in the tool. And so in terms of the actual Coalesce platform, I'm wondering if you can speak to the technical architecture and how you've approached the implementation of the system.
229
00:17:26.985 -->
00:17:35.405 Yeah. So this is, again, we have implemented on the cloud. We are on Google Cloud. You know, we have a Kubernetes cluster that would serve as an application
230
00:17:36.110 -->
00:17:39.409to our, you know, clients. It's a it's a multi tenant environment.
231
00:17:40.029 -->
00:17:42.210And we have a metadata database
232
00:17:42.669 -->
00:17:44.690that is in Google
233
00:17:45.165 -->
00:17:45.665Firebase.
234
00:17:46.525 -->
00:17:50.065So, you know, we get all that scalability from the Google's system.
235
00:17:50.765 -->
00:17:56.620And as far as scalability of processing goes, the data processing, of course, we rely on Snowflake.
236
00:17:57.080 -->
00:17:59.180We have a template render that
237
00:17:59.560 -->
00:18:05.515takes all the metadata as input and generates the code according to whatever template logic that was written.
238
00:18:05.975 -->
00:18:09.995And those are submitted to Snowflake, and Snowflake is doing the heavy lifting
239
00:18:10.509 -->
00:18:20.405and returns the result sets, you know, whatever it has done and how many rows affected or or is there an error. Those things get back to this
240
00:18:20.785 -->
00:18:22.005to the system from Snowflake.
241
00:18:22.385 -->
00:18:27.925Yeah. So essentially, it is a cloud based system, which where we have a cluster that is serving,
242
00:18:28.260 -->
00:18:30.200working as a multi tenant platform.
243
00:18:30.660 -->
00:18:34.680 In terms of the data architectural patterns, you mentioned that
244
00:18:35.060 -->
00:18:37.160Coalesce is designed to enable
245
00:18:37.764 -->
00:18:47.225a junior or intermediate level data engineer to be productive while staying within the guardrails that are set by a more senior engineer or a data architect.
246
00:18:47.590 -->
00:18:52.170And I'm curious if you can talk to some of the patterns that you have seen organizations
247
00:18:52.470 -->
00:19:07.610fall prey to where they end up spending wasted cycles or they start to design themselves into a situation where they're gradually losing productivity rather than gaining it? We have seen some amazing things that are happening with our
248
00:19:08.230 -->
00:19:12.570 this whole user defined node concept that we have provided to our customers.
249
00:19:13.294 -->
00:19:14.515But at the same time,
250
00:19:14.895 -->
00:19:18.275you know, people people can make mistakes with that. You know?
251
00:19:18.735 -->
00:19:24.919So far, I would say it's been more positive than negative. I can talk about just some pitfalls if that's what you're looking for
252
00:19:25.380 -->
00:19:26.440that people can,
253
00:19:26.820 -->
00:19:33.135you know, do and get into trouble. I myself had had those kind of pitfalls in the past.
254
00:19:33.435 -->
00:19:36.575You know, sometimes under pressure, I would take some
255
00:19:37.250 -->
00:19:38.230band aid approach
256
00:19:38.769 -->
00:19:40.149and do something that
257
00:19:40.929 -->
00:19:43.909is, you know it doesn't address the foundational
258
00:19:44.835 -->
00:19:46.695aspect of the data analytics
259
00:19:47.075 -->
00:19:54.750solution, but it just like a band aid. And then you end up with that band aid forever. You think you can get rid of it, but you don't. That's a pitfall.
260
00:19:55.450 -->
00:19:56.190And and, also,
261
00:19:56.889 -->
00:19:58.350you know, when you plan
262
00:19:58.889 -->
00:20:00.429these data projects
263
00:20:01.225 -->
00:20:04.764and if you quickly create a standard and you give it to the business,
264
00:20:05.385 -->
00:20:21.524the business might take that as a solution. You already built the solution, so you're done. You know? And I got what I want. So it's over. Right? But in your mind, you're thinking, hey. That was just a Band Aid. I still haven't built it the right way. I need more budget. I need more people that there is a part 2 for this project.
265
00:20:21.904 -->
00:20:25.445So that is another pitfall that I myself encountered in the past
266
00:20:25.950 -->
00:20:30.049where it's nothing to do with the technology itself. It's just more about
267
00:20:30.590 -->
00:20:33.809how you approach this whole thing. You know, if I'm building a foundation,
268
00:20:34.424 -->
00:20:57.155you gotta say part 1, part 2, part 3 or phase 1, phase 2, phase 3. Phase 1 is probably a quick and dirty solution. Phase 2 is the improvement on that. Phase 3 is the real output. And you gotta plan for that 3 years or whatever number of years and get the budget for the whole thing, not just for 1 thing. So that's the lesson that I learned myself when I was doing it. So, again, I know you're looking for more technical
269
00:20:57.535 -->
00:21:08.250side of these things, but I think sometimes it's the nontechnical is more important than the technical, I would say. As far as the technical aspects of this, it's pretty straightforward, you know, because
270
00:21:08.915 -->
00:21:10.535you know what Snowflake does.
271
00:21:10.915 -->
00:21:13.175If you have somebody who has built data warehouses,
272
00:21:13.715 -->
00:21:15.255we are providing you a platform
273
00:21:15.875 -->
00:21:16.695that can
274
00:21:17.580 -->
00:21:18.080automate
275
00:21:18.460 -->
00:21:19.600those patterns.
276
00:21:20.299 -->
00:21:20.799So
277
00:21:21.419 -->
00:21:24.480if you do that, you're gonna be pretty good, pretty satisfied.
278
00:21:24.985 -->
00:21:34.365 In terms of those prebuilt templates, you mentioned that it comes out of the box with a certain set of them. End users are able to add and customize their own templates.
279
00:21:34.830 -->
00:21:39.170I'm curious what your approach has been to figuring out what is the
280
00:21:39.950 -->
00:21:42.450minimum base set of templates that you want to provide,
281
00:21:42.965 -->
00:21:43.705the specific
282
00:21:44.085 -->
00:21:52.025data modeling styles that you want to work with, maybe providing templates to be able to work with specific data sources and use cases,
283
00:21:52.870 -->
00:22:05.735and how you think about what you wanted to have available at start, whether it was, like, the Snowflake approach to data modeling in terms of the star schemas and slowly changing dimensions or Data Vault and just how you think about that overall process of
284
00:22:06.035 -->
00:22:07.415providing data modeling
285
00:22:07.955 -->
00:22:17.000out of the box to get people started moving faster and helping them to discover what are the actual problems that they care about as a business that are the rest of the 20%.
286
00:22:18.035 -->
00:22:19.575 So what we're seeing is,
287
00:22:19.955 -->
00:22:28.850as far as the data warehousing solutions go, you know, people have certain methodologies that they want to adopt. Right? I mean, Kimbell has been the standard 1 for a long time.
288
00:22:29.230 -->
00:22:31.410There is a lot of momentum around Datawalt,
289
00:22:32.030 -->
00:22:36.605and there is everything in between. Right? I mean, you know, variations of Datawalt,
290
00:22:37.145 -->
00:22:38.925variations of other methodologies.
291
00:22:39.625 -->
00:22:43.805So what we are doing is we are giving a set of out of the box,
292
00:22:44.340 -->
00:22:45.480again, these nodes,
293
00:22:45.780 -->
00:22:49.160you know, for dimensions, for facts, for precision staging,
294
00:22:49.620 -->
00:22:50.440stage nodes,
295
00:22:50.900 -->
00:22:51.400hubs,
296
00:22:51.875 -->
00:22:55.255links, satellites, you name it. So we provide that
297
00:22:55.795 -->
00:22:57.175those things out of the box.
298
00:22:57.555 -->
00:23:01.850For the most part, that will satisfy a lot of use cases right out of the
299
00:23:02.150 -->
00:23:06.090box. And anything that is beyond that, they can change it. They can
300
00:23:06.470 -->
00:23:10.570branch off of an existing 1, and they can enhance it to meet their needs.
301
00:23:10.904 -->
00:23:12.924But what we are seeing also is, like,
302
00:23:13.465 -->
00:23:15.804people coming up with stuff that we didn't expect.
303
00:23:16.345 -->
00:23:21.140For example, you know, Snowflake has the streaming and tasks as a functionality.
304
00:23:21.760 -->
00:23:24.260They just identify deltas and things like that.
305
00:23:24.720 -->
00:23:25.7801 of our customers,
306
00:23:26.400 -->
00:23:27.220they just
307
00:23:27.600 -->
00:23:33.695went ahead and built a c and c node. And now that node is available in the graph, and you can just take a bunch
308
00:23:34.154 -->
00:23:36.654of raw tables and say, add a stream,
309
00:23:37.090 -->
00:23:37.830add a
310
00:23:38.210 -->
00:23:39.910task, and run every whatever
311
00:23:40.370 -->
00:23:47.68410 minutes or whatever and dump the delta into another table. It has become such an easy task for everybody else to
312
00:23:48.144 -->
00:23:51.445consume that type of functionality and build that functionality into their pipelines.
313
00:23:51.904 -->
00:23:55.060That's what we're seeing. You know? And we also see that
314
00:23:55.680 -->
00:24:03.220the nodes that we are getting out of the box, that is going to grow because we wanna create a marketplace where people can actually share these things,
315
00:24:03.605 -->
00:24:07.065notes or packages, make up a bunch of notes together
316
00:24:07.524 -->
00:24:09.225that perform a certain function.
317
00:24:09.684 -->
00:24:11.300So that's where we're going with that.
318
00:24:11.700 -->
00:24:14.120 And as far as the workflow for
319
00:24:14.660 -->
00:24:15.640adopting Coalesce
320
00:24:15.940 -->
00:24:20.200and starting to integrate it into the usage and the
321
00:24:20.565 -->
00:24:21.304data platform
322
00:24:21.764 -->
00:24:23.465and the sort of organizational
323
00:24:23.924 -->
00:24:24.904analytics capabilities.
324
00:24:25.284 -->
00:24:37.465Wondering if you can just talk through that process and some of the background knowledge that's useful to have as you're figuring out what the overall workflow and the node structures are going to look like. Again, there's several personas
325
00:24:38.165 -->
00:24:44.185 that are going to be using the tool. It's not just built for 1 type because we want to address the entire problem
326
00:24:44.630 -->
00:24:48.170as much as we can, not just 1 piece. Although we're focused on transformations,
327
00:24:48.870 -->
00:24:52.490there's there's other things that are on the edge that also are important.
328
00:24:52.845 -->
00:24:58.625So transformation seems like, how do I be changing the data from, you know, 1 form to another as quickly as possible?
329
00:24:59.085 -->
00:25:01.105But what if I don't have column lineage?
330
00:25:01.570 -->
00:25:02.950Right? Then you don't have
331
00:25:03.410 -->
00:25:04.870a way to really
332
00:25:05.330 -->
00:25:10.790kinda see, you know, what's going on. So to answer your question, it depends on
333
00:25:11.485 -->
00:25:12.065the organizational
334
00:25:12.524 -->
00:25:14.065structure. They they have
335
00:25:14.605 -->
00:25:20.465data engineers that would be ideal kind of persona to deal with the tool because they understand SQL.
336
00:25:20.840 -->
00:25:22.300They understand the methodology
337
00:25:23.080 -->
00:25:25.420to some degree. And, you know, if you have architects,
338
00:25:26.200 -->
00:25:30.220you know, for them, they're gonna work with the customization aspect of the tool.
339
00:25:30.665 -->
00:25:43.120So I think that's what is expected. To work with the tool, we're seeing you have to be an architect or a engineer or a power user who would get some help from IT, but they also can build the pipelines
340
00:25:43.500 -->
00:25:45.715whether they are proficient in SQL or not.
341
00:25:46.115 -->
00:25:54.135 And as I was looking at the Coalesce product, I noticed that it's very closely tied to Snowflake as the underlying
342
00:25:54.780 -->
00:26:06.455storage and warehouse layer, and I'm curious if you can speak to the thinking that went into that decision and some of the ways that you have implemented Coalesce to potentially allow for
343
00:26:06.755 -->
00:26:14.780additional storage and query engines in the future and some of the other directions that you see as potential expansions to coalesce?
344
00:26:15.480 -->
00:26:27.774 When we started this, you know, Snowflake was an obvious choice. Every other prospect that we talked to is moving to Snowflake pretty much. It was pretty clear for us to focus on Snowflake.
345
00:26:28.530 -->
00:26:31.350However, the tool is built in a way that
346
00:26:31.970 -->
00:26:35.830the communication with the query engine has been abstracted.
347
00:26:36.355 -->
00:26:42.295It's just a best practice. Right? I mean, to build a software in a way that the front end is agnostic to what it's talking to.
348
00:26:42.675 -->
00:26:44.375So that's how the tool is built.
349
00:26:44.950 -->
00:26:47.050The middle layer is the template.
350
00:26:47.510 -->
00:26:52.330And we are not thinking about this at this time because we are hyper focused on Snowflake,
351
00:26:53.165 -->
00:26:56.065And we wanna expand the platform even more and be
352
00:26:56.445 -->
00:27:02.225in lockstep with Snowflake's features and things like that. But, however, because we built it in a way that
353
00:27:02.669 -->
00:27:03.970that part is abstracted,
354
00:27:04.429 -->
00:27:06.530if we really want to support another platform,
355
00:27:07.150 -->
00:27:09.410all it is that we have to do is
356
00:27:09.789 -->
00:27:10.770build those templates
357
00:27:11.345 -->
00:27:13.684that would generate the flavor of the SQL
358
00:27:14.225 -->
00:27:16.645that can run on a particular target platform.
359
00:27:17.184 -->
00:27:18.565 In terms of the
360
00:27:19.390 -->
00:27:28.530ways that you have been working with some of your early design partners, I'm wondering what are some of the most interesting or innovative or unexpected ways that they're using Coalesce?
361
00:27:29.184 -->
00:27:30.804 1 of the things I was
362
00:27:31.184 -->
00:27:32.325surprised is
363
00:27:32.945 -->
00:27:33.764people are
364
00:27:34.225 -->
00:27:34.725building
365
00:27:35.505 -->
00:27:38.780these nodes that I was talking about that we never thought of.
366
00:27:39.560 -->
00:27:40.060And,
367
00:27:40.600 -->
00:27:42.380you know, 1 example I gave you,
368
00:27:42.760 -->
00:27:46.140which is, you know, streaming and tasks, CDC type of functionality,
369
00:27:46.924 -->
00:27:49.265But there's also people building, like,
370
00:27:49.645 -->
00:27:50.625a data profiling
371
00:27:51.325 -->
00:27:52.865functionality into this.
372
00:27:53.245 -->
00:27:56.065So people can create a profiling node
373
00:27:56.590 -->
00:27:58.850and capture their profiling metrics
374
00:27:59.549 -->
00:28:01.010to monitor data quality.
375
00:28:01.390 -->
00:28:05.650And that is something that we haven't, you know, built or given to anybody.
376
00:28:06.054 -->
00:28:08.715But because the platform enables them to do that,
377
00:28:09.335 -->
00:28:11.515so they are rapidly building these
378
00:28:11.895 -->
00:28:13.674nodes that we never thought of.
379
00:28:14.215 -->
00:28:18.400And the other aspect, you know, is very interesting from my standpoint is
380
00:28:18.780 -->
00:28:19.600how quickly
381
00:28:19.900 -->
00:28:22.720people are able to build, you know, complex
382
00:28:23.304 -->
00:28:23.804solutions.
383
00:28:24.504 -->
00:28:27.164I'll tell you a recent thing that happened. So
384
00:28:27.705 -->
00:28:31.310we have an alliance director who was, you know, responsible to work with
385
00:28:31.710 -->
00:28:35.410large firms and their, you know, system integrators and partners.
386
00:28:36.030 -->
00:28:38.930You know, he's been handing out trial accounts and things like that.
387
00:28:39.465 -->
00:28:40.285And there was
388
00:28:40.825 -->
00:28:45.085this SI, you know, a a big firm. You know, they got some trial accounts.
389
00:28:45.785 -->
00:28:52.050And, you know, we got in a call. After they got the trial accounts, 2 days later, we were on a call. And before the call started,
390
00:28:52.670 -->
00:29:09.500they were saying that, hey. They want to talk to another firm to talk and and show them this tool. And we were like, we just sent you a recording. You know, we just need to talk and not make make sure that you understand the tool. And they said, but I think we figured out. Let me show you. And then then they started sharing screen.
391
00:29:09.800 -->
00:29:11.420They built this gigantic
392
00:29:12.040 -->
00:29:12.540graph
393
00:29:12.975 -->
00:29:13.795that is basically
394
00:29:14.815 -->
00:29:16.595a implementation of an SAP,
395
00:29:17.135 -->
00:29:18.195you know, SAP
396
00:29:18.575 -->
00:29:21.950module or whatever, SAP thing that they built. I'm not an expert in SAP,
397
00:29:22.650 -->
00:29:25.390but whatever they build is called a calculation view or something.
398
00:29:25.770 -->
00:29:26.270And
399
00:29:26.570 -->
00:29:28.990they built that in a matter of hours
400
00:29:29.804 -->
00:29:31.985to see how it performs on Snowflake.
401
00:29:32.524 -->
00:29:40.019On SAP, it takes, like, a long time because of whatever the SAP architecture is and how it works. But when they moved that to Snowflake
402
00:29:40.480 -->
00:29:42.019and and they built using coalesce,
403
00:29:42.559 -->
00:29:51.285you know, they got the performance boost. But what I was surprised was how quickly they picked up the tool based on a recording. Yeah. It's definitely very cool.
404
00:29:51.665 -->
00:29:54.245 And so in terms of your own experience
405
00:29:54.625 -->
00:29:57.125of building Coalesce and
406
00:29:57.610 -->
00:30:03.230iterating on the technical aspects, working with your design partners to figure out the product direction
407
00:30:03.850 -->
00:30:09.934and grow the business, I'm curious, what are some of the most interesting or unexpected or challenging lessons that you've learned in the process?
408
00:30:10.395 -->
00:30:15.375 First of all, there's a lot of learning. As soon as I started, I was constantly looking at
409
00:30:15.700 -->
00:30:17.000a lot of other tools
410
00:30:17.940 -->
00:30:22.440and what they're doing and what are the gaps that they're trying to fill, you know, starting with,
411
00:30:22.740 -->
00:30:23.400you know,
412
00:30:23.975 -->
00:30:24.475schedulers,
413
00:30:25.095 -->
00:30:26.235orchestrating tools,
414
00:30:27.175 -->
00:30:27.675you
415
00:30:28.055 -->
00:30:28.795know, data
416
00:30:29.335 -->
00:30:37.420observability tools and whatnot. So there's a lot of learning in that regard. It was enjoyable for me. I enjoyed that part. I wouldn't call that as a challenge. But
417
00:30:37.880 -->
00:30:43.665I think 1 of the challenging aspects for us right now that I see is, you know, we show the tool to people and
418
00:30:44.285 -->
00:30:45.425they get very excited,
419
00:30:45.965 -->
00:30:50.100but they always have something to add to it in terms of what they need.
420
00:30:51.460 -->
00:31:01.885And it's almost like, hey. I love your tool. I wanna use it, you know, but can you add this functionality? Can you make sure that you have this by this time? Or when are you planning to have it?
421
00:31:02.424 -->
00:31:04.044And we get that from all directions.
422
00:31:04.745 -->
00:31:18.154So to manage all of that and to prioritize, which is an obvious thing for and this is a common problem for any vendor, I guess. That is very challenging, in my opinion, to be able to focus on what we're doing, but also prioritize
423
00:31:19.174 -->
00:31:20.635and pivot if necessary
424
00:31:21.335 -->
00:31:23.835and address those in a timely manner
425
00:31:24.429 -->
00:31:26.450is definitely challenging. And
426
00:31:26.750 -->
00:31:33.695I knew that, but I'm experiencing it now. So that's different just from knowing and then actually experiencing it. Yeah. Absolutely.
427
00:31:34.555 -->
00:31:36.175 In terms of the
428
00:31:36.955 -->
00:31:39.135sort of management of these
429
00:31:39.675 -->
00:31:44.390graphs and the execution plans that you have. I'm wondering if you can talk to the,
430
00:31:44.690 -->
00:31:49.030I guess, change management process there and how you're able to maybe
431
00:31:49.554 -->
00:31:52.294automate construction of some of these graphs
432
00:31:53.075 -->
00:31:57.254or being able to say, I've built this graph. I'm going to test it in
433
00:31:57.620 -->
00:32:07.000either a test account for Snowflake or, you know, on on a test subset of the data and then being able to manage that rollout to the full production environment.
434
00:32:07.684 -->
00:32:25.495 Absolutely. And change management is very, very important thing that we have kinda made sure we focus on that right from the beginning. So, you know, first of all, we have of course, we integrate with Git. We save the state of what was built and what was deployed in our metadata and in Git. So whatever you build,
435
00:32:25.875 -->
00:32:34.700it goes to Git. From there, it goes to different environments that you want to deploy to. So we have this concept of, you know, creating an environment that has
436
00:32:35.320 -->
00:32:37.980credentials, that has a Snowflake account details.
437
00:32:38.440 -->
00:32:39.820It also has something
438
00:32:40.325 -->
00:32:48.745called storage mappings. So it's basically saying, hey. What database key ones that I need to work with when you push this code to this particular environment?
439
00:32:49.210 -->
00:32:53.870So an environment kinda encapsulates all of those things. So when we take this
440
00:32:54.490 -->
00:32:57.310git state and we push that to that environment,
441
00:32:57.930 -->
00:32:58.590you know,
442
00:32:58.894 -->
00:33:03.875you can do this with command line or you can do this via front end. But it goes through certain
443
00:33:04.254 -->
00:33:07.154process where it compares what's on the target,
444
00:33:07.550 -->
00:33:09.890and it will show you the differences
445
00:33:10.270 -->
00:33:14.450in the delta that it's going to execute. This is what we call plan and deploy,
446
00:33:14.830 -->
00:33:16.770where you always have a
447
00:33:17.295 -->
00:33:20.115plan that you can see before you actually deploy.
448
00:33:20.495 -->
00:33:32.230So you get to approve that. So that's how the change management is done. That's how you promote from 1 environment to the other. It's going from dev to get get to any number of environments that you want to push push to.
449
00:33:35.045 -->
00:33:39.865 Modern data teams are dealing with a lot of complexity in their data pipelines and analytical code.
450
00:33:40.245 -->
00:33:47.269Monitoring data quality, tracing incidents, and testing changes can be daunting and often takes hours to days or even weeks.
451
00:33:47.649 -->
00:33:58.794By the time errors have made their way into production, it's often too late and the damage is done. DataFold built automated regression testing to help data and analytics engineers deal with data quality in their pull requests.
452
00:33:59.414 -->
00:34:08.030DataFold shows how a change in SQL code affects your data, both on a statistical level and down to individual rows and values before it gets merged to production.
453
00:34:08.890 -->
00:34:13.285No more shipping and praying. You can now know exactly what will change in your database.
454
00:34:13.745 -->
00:34:21.510DataFold integrates with all major data warehouses as well as frameworks such as airflow and DBT and seamlessly plugs into CI workflows.
455
00:34:22.130 -->
00:34:22.630Visitdataengineeringpodcast.com/datafold
456
00:34:25.250 -->
00:34:27.110today to book a demo with DataFold.
457
00:34:29.295 -->
00:34:31.875For people who are looking to
458
00:34:32.175 -->
00:34:35.155accelerate their rate of development
459
00:34:35.535 -->
00:34:39.075and the speed at which they're able to go from idea
460
00:34:39.539 -->
00:34:42.839to analysis, what are the cases where Coalesce is the wrong choice?
461
00:34:43.380 -->
00:34:45.480 It depends on the use case, for sure.
462
00:34:46.019 -->
00:35:01.120And Coalesce is built definitely more for the preparing data for analytics. Let's say that. And there are certain proven methodologies that people adopt. You know, for example, you know, Kimball or Datavault or things like that that, you know, if you're building something central,
463
00:35:01.500 -->
00:35:02.320that's core,
464
00:35:02.780 -->
00:35:11.695you probably follow 1 of these methodologies to build that. Now people also build something in between. As I said, you know, they can just build flat tables. That's fine too.
465
00:35:12.155 -->
00:35:13.375But where it doesn't
466
00:35:13.860 -->
00:35:20.760fit is if you're just moving data from 1 point a to point b for application to application integration, for example,
467
00:35:21.155 -->
00:35:30.535that would not be call us, I would say. I mean, you can bend the tool to do that, but that's not the purpose of the tool. It's definitely in the data analytics domain.
468
00:35:31.000 -->
00:35:53.099So, you know, if you want to do application to application integration, that would be something else that you should look at. As you continue to iterate on the product and keeping in mind these competing priorities and feature requests, I'm wondering if you can speak to some of the things you have planned for the near to medium term. Yeah. Definitely. I mean, you know, this space is vast. There's a lot of need out there for clients.
469
00:35:53.640 -->
00:35:55.9801 of the things that we are focusing is,
470
00:35:56.359 -->
00:35:59.339you know, you're going to call us and start building the pipelines.
471
00:35:59.915 -->
00:36:03.915You know, build these nodes and you just build the pipeline, right, from from,
472
00:36:04.315 -->
00:36:05.615from the beginning. But
473
00:36:05.915 -->
00:36:10.510we wanna add some kind of modeling to this down the road, a way to kinda
474
00:36:10.970 -->
00:36:12.990look at the source data and see
475
00:36:13.450 -->
00:36:21.815if you can, you know, kinda connect the dots and say, hey, this field in here is related to this field in this table. You know, and the tables could be coming from different sources.
476
00:36:22.195 -->
00:36:33.000But with that kind of information and input from the user, now we can take the automation to the next level because, you know, you don't even have to specify the joins anymore because we already did from that interface
477
00:36:33.460 -->
00:36:35.720by looking at the data at the beginning.
478
00:36:36.055 -->
00:36:38.875Therefore, we can automate those things. We call it the discovery
479
00:36:39.415 -->
00:36:40.155of datasets.
480
00:36:40.695 -->
00:36:42.875So that's another piece that we're
481
00:36:43.255 -->
00:36:45.035very focused on. But,
482
00:36:45.460 -->
00:36:45.960also,
483
00:36:46.260 -->
00:36:48.040Snowflake is adding so much functionality
484
00:36:48.820 -->
00:37:07.050as we speak. I mean, we want to be in lockstep with that, you know, whether it's a data science use case, you know, like, for example, the Snowpark, I think that's what it's called. We wanna make sure what we can do there. That's definitely on our minds as well to support those data science use cases and other languages that you can generate code for,
485
00:37:07.430 -->
00:37:07.930Snowflake.
486
00:37:08.470 -->
00:37:12.555 Are there any other aspects of the work that you're doing at Coalesce or the
487
00:37:13.015 -->
00:37:15.355overall problems of how to approach
488
00:37:15.655 -->
00:37:17.994data architectural patterns and stay
489
00:37:18.454 -->
00:37:25.660sort of ahead of the game in terms of being able to build out analysis and drive the business that we didn't discuss yet that you'd like to cover before we close out the show?
490
00:37:25.960 -->
00:37:27.579 1 thing I am, you know,
491
00:37:27.895 -->
00:37:31.435passionate about, and this is something that I got introduced to recently,
492
00:37:32.295 -->
00:37:33.195is the decentralization
493
00:37:33.975 -->
00:37:37.035paradigm, which is, you know, in other words, they call it data mesh.
494
00:37:37.700 -->
00:37:40.980I'm really very passionate about that idea because it just seems
495
00:37:41.380 -->
00:37:49.415going back in all my, you know, years of experience with this, if I look at this particular paradigm, it makes a lot of sense.
496
00:37:49.795 -->
00:37:55.415Just to give you a very high level overview of what that means, you know, basically, there's underlying 4 principles.
497
00:37:56.130 -->
00:37:57.8291 is make the domain
498
00:37:58.289 -->
00:38:04.069responsible for building the pipelines and producing high quality data. So in other words, rather than having
499
00:38:04.369 -->
00:38:05.510a central team
500
00:38:05.895 -->
00:38:06.635that does
501
00:38:07.015 -->
00:38:13.275build this big, gigantic data warehouse and has a team of data engineers building this large pipelines,
502
00:38:13.815 -->
00:38:19.619you know, instead of doing that, you know, how about we kinda decentralize this? Like, take
503
00:38:20.079 -->
00:38:26.045that same idea, but do it at a domain level. Do it at a line of business level. That's the first principle.
504
00:38:26.505 -->
00:38:51.109But, obviously, once you say that, now aren't you creating silos is the next question. Right? If you do that, now you're creating silos. But then the answer to that is, you know, they have to create in a way that it is a product based mentality. It's like a data as a product. That's what you call. When you go to the supermarket, you buy a product. You expect certain quality. You expect certain documentation. You expect it to be safe. That's the same idea. So if my domain,
505
00:38:51.410 -->
00:39:07.125you know, produces some data and publishes some data, the other domains, other people who want to consume that data expect certain quality. That's the second principle. And the third principle is the self serving aspect to it. Like, people don't have to rely on IT or some kind of specialist
506
00:39:07.700 -->
00:39:08.359that they
507
00:39:08.740 -->
00:39:13.400need to talk to to use this dataset. Instead, they can just kinda do self-service.
508
00:39:14.099 -->
00:39:14.839And finally,
509
00:39:15.380 -->
00:39:16.599some kind of governance
510
00:39:17.235 -->
00:39:27.070on all of this. Governance at the local level, it is at the domain level. At the same time, governance at the broader level, especially from IT, to make sure that there's no duplication
511
00:39:27.690 -->
00:39:32.350and things are being shared correctly and things are have some consistency, some standards.
512
00:39:32.895 -->
00:39:45.190So all of this whole data mesh paradigm, I'm very excited about. And the good news that we have from Cholera side is I think Cholera is just right out out of the box. It checks all these boxes pretty much right away.
513
00:39:45.490 -->
00:39:45.990And
514
00:39:46.290 -->
00:39:50.070I'm very, very curious to see, you know, if an organization
515
00:39:51.525 -->
00:39:56.345is going in that direction, I want to see callouts play an important role in their organization.
516
00:39:56.805 -->
00:39:58.345 Yeah. Data mesh is definitely
517
00:39:58.965 -->
00:39:59.715an interesting
518
00:40:00.359 -->
00:40:11.395approach that has been gaining a lot of attention, so definitely appreciate your enthusiasm for it. Spoken to Zhamak a couple of times, and it's definitely a subject that comes up repeatedly on this show.
519
00:40:11.935 -->
00:40:15.955 Cool. I'm glad it is. Because in my opinion, I think that is, a way to
520
00:40:16.510 -->
00:40:17.890scale for an organization
521
00:40:18.350 -->
00:40:24.050moving forward. However, there's a lot more in there to learn and and make sure you do it right. Absolutely.
522
00:40:24.605 -->
00:40:40.815 Well, for anybody who wants to get in touch with you and follow along with the work that you're doing, I'll have you add your preferred contact information to the show notes. And as a final question, I'd like to get your perspective on what you see as being the biggest gap in the tooling or technology that's available for data management today. I talked about the the modeling
523
00:40:41.115 -->
00:40:46.095 piece that, you know, a lot of people are asking about it because, you know, it's 1 thing to understand
524
00:40:46.910 -->
00:40:50.930the raw data and how that data is linked, especially coming from different sources.
525
00:40:51.470 -->
00:40:55.569On the other hand, once you build something, you also want to see what you built.
526
00:40:56.135 -->
00:40:59.835For example, if you built a data vault, you want to be able to visualize that
527
00:41:00.135 -->
00:41:00.795and see
528
00:41:01.095 -->
00:41:06.180as you're building it, hey, is this what I want? You you cannot comprehend everything just by looking at code
529
00:41:06.480 -->
00:41:07.619or even a pipeline.
530
00:41:08.079 -->
00:41:14.660There is another perspective to this, which is the the final kind of model view that people will see and kinda
531
00:41:14.984 -->
00:41:16.845can use that as a communication tool
532
00:41:17.305 -->
00:41:20.525for them to understand and also communicate with with other
533
00:41:20.905 -->
00:41:22.365business, you know, users.
534
00:41:22.670 -->
00:41:25.490We think that is very, very critical and important. And
535
00:41:25.790 -->
00:41:36.435I know we tied our road map, but it's gonna be coming very soon. That's 1 thing I can say. I mean, there's a whole lot of other things as well. But I would I would just say 1 since that is the near term. Absolutely.
536
00:41:37.135 -->
00:41:53.545 Well, thank you very much for taking the time today to join me and share the work that you're doing at Coalesce. It's definitely a very interesting product and tackling a real problem that people are experiencing. So I appreciate all of the time and energy that you and your team are putting into that, and I hope you enjoy the rest of your day. Thank you so much. It was a pleasure.
537
00:41:58.964 -->
00:42:02.140For listening. Don't forget to to check out our other show, podcast.init@pythonpodcast.com
538
00:42:04.680 -->
00:42:09.020to learn about the Python language, its community, and the innovative ways that is being used.
539
00:42:09.434 -->
00:42:10.734And visit the site at dataengineeringpodcast.com
540
00:42:12.155 -->
00:42:21.570to subscribe to the show, sign up for the mailing list, and read the show notes. If you've learned something or tried out a project from the show, then tell us about it. Email hosts at data engineering podcast.com
541
00:42:22.110 -->
00:42:27.410with your story. And to help other people find the show, please leave your view on Itunes and tell your friends and coworkers.