WEBVTT
NOTE
Transcription provided by Podhome.fm
Created: 04/07/2026
23:53:31Duration: 3563.887
Channels: 1
1
00:00:11.360 -->
00:00:15.440Hello, and welcome to the data engineering podcast, the show about modern data management.
2
00:00:16.075 -->
00:00:29.195If you lead a data team, you know this pain. Every department needs dashboards, reports, custom views, and they all come to you. So you're either the bottleneck slowing everyone down, or you're spending all your time building one off tools instead of doing actual data work.
3
00:00:29.970 -->
00:00:37.010Retool gives you a way to break that cycle. Their platform lets people build custom apps on your company data while keeping it all secure.
4
00:00:37.489 -->
00:00:46.155Type a prompt like build me a self-service reporting tool that lets teams query customer metrics from Databricks, and they get a production ready app with the permissions and governance built in.
5
00:00:46.555 -->
00:00:50.395They can self serve, and you get your time back. It's data democratization
6
00:00:50.395 -->
00:00:51.595without the chaos.
7
00:00:51.915 -->
00:00:54.875Check out Retool at dataengineeringpodcast.com
8
00:00:54.875 -->
00:01:01.879slash Retool today, that's r e t o o l, and see how other data teams are scaling self-service.
9
00:01:01.960 -->
00:01:05.560Because let's be honest, we all need to retool how we handle data requests.
10
00:01:06.200 -->
00:01:16.085Your host is Tobias Macy, and today I'm welcoming back Gleb Majanski to talk about predictions for the impact of AI on data engineering through the remainder of 2026.
11
00:01:16.165 -->
00:01:20.885And so, Gleb, for folks who haven't heard any of your past appearances, if you could just give a quick introduction.
12
00:01:21.924 -->
00:01:29.070Yeah. Thanks for hosting me again, Tobias. Always great to be back. Yeah, I'm Gleb. I am CEO and co founder of DataFold.
13
00:01:29.070 -->
00:01:30.190I spent
14
00:01:30.270 -->
00:01:32.030pretty much my entire career
15
00:01:32.350 -->
00:01:34.990being a data engineer and building data platforms.
16
00:01:35.070 -->
00:01:39.230I got a chance to build data platform for Autodesk
17
00:01:39.395 -->
00:01:51.314using the new data cloud tools that were new at the time. That was over ten years ago. And spent a few years scaling data platforms at Lyft, which was a really great and insightful time because
18
00:01:51.475 -->
00:01:52.594Lyft is an incredibly
19
00:01:53.090 -->
00:01:54.850data driven business.
20
00:01:55.010 -->
00:01:58.530And then I built the third data platform at a
21
00:01:58.690 -->
00:02:05.010startup building teleoperation for autonomous vehicles and had to deal with a lot of IoT data and telemetry,
22
00:02:05.010 -->
00:02:08.755which was also really cool. And finally started DataFold,
23
00:02:08.755 -->
00:02:09.555which actually
24
00:02:10.035 -->
00:02:24.930I've been working on for now I realized six years already this March. And at DataFold, we have always worked on automating data engineering. And pre AI, we focused a lot on automating different data quality workflows. And now with AI,
25
00:02:25.170 -->
00:02:28.930we focus on AI automation, which includes providing
26
00:02:28.930 -->
00:02:30.050specialized agents
27
00:02:30.305 -->
00:02:34.705for certain things like automating data platform migrations and optimization
28
00:02:34.705 -->
00:02:42.225of data platform costs. And we also provide tools for anyone's agents to be more efficient with context and various specialized tooling.
29
00:02:42.989 -->
00:02:44.590And so you recently
30
00:02:44.670 -->
00:02:50.190put up a blog post on your company blog talking about some of the predictions that you have for
31
00:02:50.430 -->
00:02:52.909the data engineering ecosystem
32
00:02:52.909 -->
00:02:58.375now that we do have AI that is capable of performing a lot of operational
33
00:02:58.375 -->
00:03:04.375and development tasks. And so that definitely changes the scope and skills
34
00:03:04.375 -->
00:03:23.590of what people who are actually working in these spaces need to be thinking about. And so I'm just wondering if you can just quickly run through some of the key points, and then we'll explore some of the impact that that has on how people should be thinking about structuring their day to day work and their systems to be able to actually take advantage of some of these shifts in the industry.
35
00:03:24.575 -->
00:03:26.175Yeah, absolutely. And
36
00:03:26.895 -->
00:03:38.175title of the post saying predictions is not particularly modest, but I can talk about how it came together. So AI has been kind of in our life for what, let's say three years at least.
37
00:03:38.829 -->
00:03:41.790And we've all been using it in different forms,
38
00:03:41.950 -->
00:03:47.870starting from Charge GPT obviously, and kind of trying to build things and automations and using coding tools.
39
00:03:48.189 -->
00:03:52.349And I think it's all has been kind of coming and evolving.
40
00:03:52.855 -->
00:03:53.735And then
41
00:03:53.975 -->
00:03:56.295probably around November,
42
00:03:56.295 -->
00:04:00.295remember I had a kind of like an awakening myself where
43
00:04:00.615 -->
00:04:10.170after the release of Opus 4.5 Biontropic and then OpenAI releasing their own version of the similarly capable model, if I remember that order correctly.
44
00:04:10.170 -->
00:04:14.730I think there's been a very sharp increase that I felt personally in my workflows
45
00:04:14.810 -->
00:04:21.050in terms of coding, both software and data engineering coding. And that was super profound
46
00:04:21.325 -->
00:04:22.925on me over Christmas
47
00:04:23.165 -->
00:04:23.965holidays.
48
00:04:24.045 -->
00:04:31.085I was able to automatically code. I wouldn't say wipe code because I think that's a little bit understating the impact,
49
00:04:31.405 -->
00:04:35.005two new products for DataFold. And then I also tried
50
00:04:35.170 -->
00:05:00.395what's it like to do data engineering with a full agentic experience and I was completely blown away. I always thought of myself as someone who can write really good SQL and in my data, data engineers, I think that was a superpower and that's how you advanced in your career and that's how you got things done. And with the current capabilities of agentic coding, I just thought that this completely changes the job and the experience.
51
00:05:00.715 -->
00:05:18.130And I thought it's really important to reflect on what it means for data engineering. And at that time I didn't really have very concrete thoughts. So I thought, well, would just try to put them on paper and make them a little bit more structured. And yeah, I can kind of go through the biggest things that at least for me were important takeaways.
52
00:05:18.130 -->
00:05:25.215So I think one is that right now it feels like the world is divided in Harry Potter into wizards and muggles.
53
00:05:25.775 -->
00:05:52.295And I was myself a muggle until I went full into agentic coding. And I think it's important to define what actually agentic coding is because there's so many different terms and buzzwords thrown around. So if you ask anyone, Hey, how are using AI in the work? Everyone would say that they use it at some capacity. But if you actually drill in and say, Well, how are you actually using it in your workflow? Let's say data engineer, Linux engineer. A lot of people would say, well, I use
54
00:05:52.535 -->
00:05:56.135ChatGPT or I use Claude or another LLM
55
00:05:56.535 -->
00:06:13.730to help me write code. And then if you ask them, walk me through what you actually are doing. They would say, well, say I need to do analysis. And I would prompt the chat to write me a code and then I would run this code against my data platform, my database, and then get results and then maybe ask the chat again for suggestions.
56
00:06:14.050 -->
00:06:16.770And this is actually helpful because writing
57
00:06:16.930 -->
00:06:20.585SQL is tedious, but it's not agentic coding.
58
00:06:20.585 -->
00:06:21.865And the way that
59
00:06:22.264 -->
00:06:30.585I would define agentic is that Magent is a AI system that is capable of achieving a goal by
60
00:06:31.050 -->
00:06:43.370choosing the path for how to get there with the tools that you give it and the context that you give it. And the difference is that if you use an agent to accomplish the similar task, let's say building a new model in DBT Wireflow,
61
00:06:43.675 -->
00:06:51.995is that an agent would actually be able to not only write the code for you, but execute that code against the database, get the results, evaluate the results,
62
00:06:52.235 -->
00:07:00.479put the code into, let's say, DBT model, run DBT, debug it, write tests, debug it, and then present you with the complete outcome.
63
00:07:00.879 -->
00:07:02.960And the difference between
64
00:07:03.039 -->
00:07:09.039that kind of workflow where an agent actually not only writes something for you, but executes actions
65
00:07:09.199 -->
00:07:14.405in the context of data engineering, that means executing something in the database, is enormous.
66
00:07:14.485 -->
00:07:22.805I would say 10 to 50 x relative to not manual work, but relative to, like, using just a chat experience or tab
67
00:07:22.965 -->
00:07:24.405autocomplete experience.
68
00:07:24.405 -->
00:07:26.245And the reason is because
69
00:07:26.870 -->
00:07:33.990it's a loop that just gets not just single task automated and everything automated. And in that environment,
70
00:07:33.990 -->
00:08:22.215the human stops being the bottleneck of having to execute things and evaluate things. We're just now becoming the driver of a really autonomous process. And this is something that has been going on for at least over a year. So Cloud Code was released, I think about a year ago, but it took some time both in terms of evolving the agents and evolving the models that power them to get to the point where this process can be truly autonomous. And so I think that that's the first prediction that I preface by saying it's obvious, because I think to many that tried it, it's very obvious. Agenting and data engineering will boom in 2026 and will just be the default mode of how most data teams are operating. But the reason why I think it's still important to talk about it is because back to the muggles and wizards, so many people haven't tried it. And even at DataFold, we had some of our
71
00:08:22.534 -->
00:08:24.455most impactful engineers
72
00:08:24.610 -->
00:08:25.490adopting
73
00:08:25.650 -->
00:08:26.770this workflow
74
00:08:27.170 -->
00:08:39.145months after, like for example, I had opened it as a CEO. And it was really hard to convince some folks because they would be like, well, really like my workflow. I need to see all the code. I need to write all the code. And they had perfectly reasonable
75
00:08:39.305 -->
00:08:46.025explanations for why that happens. But once you get into trying agent decoding, I think there's no going back. You're
76
00:08:46.665 -->
00:08:50.505never the same person again. And so I think that's why it's important to talk about it is because
77
00:08:50.880 -->
00:08:54.240I still think that the majority of data practitioners out there,
78
00:08:54.640 -->
00:09:04.640especially at big companies, especially in enterprise that's like a bit more regulated, a bit more conservative are just not yet in this mode of working. And I think that we need to change that very fast.
79
00:09:05.335 -->
00:09:38.904Yeah. That's definitely one that's worth spending some time on because even just in the software engineering space, putting aside the differences between just writing the software and working in the data engineering space, which is a topic that I have covered on this podcast, I don't even know how many times, but just even talking to software engineers about the use of agentic coding tools, there's a huge degree of variance in terms of their adoption and their willingness to seed that much control to the agents. And that I'm speaking as somebody who is leading a team who is going through some of that
80
00:09:39.225 -->
00:09:55.540growing pains, and I'm definitely on the side of being very enthusiastic about it and trying to push people into it. So you still have those people who are saying, no. You're going to pry my IDE out of my cold dead hands. I'm not going to let an AI write the code for me. And I think that there's definitely a
81
00:09:55.700 -->
00:10:05.525a spectrum, and I think that a lot of people are trying to fight what seems to be inevitable. I mean, obviously, time will tell. I'm interested too in maybe digging into some of the
82
00:10:06.085 -->
00:10:07.525distinctions of
83
00:10:07.765 -->
00:10:17.470using AgenTex software engineering and getting the AI to write the code and run the validation suite and execute the tests and create that feedback loop and some of the additional
84
00:10:17.630 -->
00:10:42.325ways that we need to think about extending that workflow or integration points or various tools or MCPs or context layers that we need to incorporate to be able to bring that level of functionality into the data engineering space where it's the code working on the actual data that determines its effectiveness and correctness and just some of the ways that that is maybe underserved by the current suite of tools and focus from
85
00:10:43.380 -->
00:10:48.820the very active and constantly shifting landscape that we're currently trying to navigate.
86
00:10:49.220 -->
00:10:51.620Yeah. Tobias, you actually raised a very good point about
87
00:10:51.940 -->
00:10:56.735the resistance being kind of giving up the control. And I
88
00:10:57.055 -->
00:10:57.695someone
89
00:10:59.055 -->
00:11:06.975that has always been trying to get to the solution as fast as possible. I don't necessarily have that problem, but I do see folks that are way more
90
00:11:07.389 -->
00:11:31.455technically capable and have much deeper knowledge than I am having that friction more because I think they just hone their craft way more than I have when it comes to writing code. But I think that what's important to understand is and remember is that at least right now as a human, you're still in control of the process, right? So however you want to do review, whether you want to review every single line of code that AI writes or define the tests or define the QA process,
91
00:11:31.615 -->
00:11:43.589you can do that, right? It doesn't mean that you don't have the control of the outcome. It's just the whole process can be that much faster if you leverage agentic coding. And then the second point that you brought up is how does
92
00:11:44.870 -->
00:11:49.029agentic data engineering differ from the same pattern in software?
93
00:11:49.190 -->
00:11:54.149And you're right, I think the access to data is incredibly important because if I'm a software engineer,
94
00:11:54.625 -->
00:12:18.180I typically develop on a sandbox environment with synthetic data or with data that I come up with just for executing tests. I don't really code live on the system. Whereas if I'm a data engineer, it's usually the opposite. I am writing code and all the preliminary exploratory queries I'm doing, I'm doing this on production data because that's how I get the insight in terms of what the data contains.
95
00:12:18.420 -->
00:12:23.540And that obviously presents challenges for AI adoption because a lot of enterprises are
96
00:12:23.925 -->
00:12:24.725rightfully
97
00:12:24.885 -->
00:12:25.765hesitant
98
00:12:25.765 -->
00:12:28.245and anxious about letting
99
00:12:28.325 -->
00:13:02.155AI, which means large language models that are hosted can be hosted in different providers accessing their proprietary data. Because if you look at the terms of service of different LLM providers, there's actually quite a bit of a range of their guarantees in terms of whether data and prompts are used for post training and evaluation, whether it's not used for post training and evaluation. And so if we have a coding agent that leverages a third party LLM and it has access to your database, that actually does present quite a bit of security surface area and privacy surface area that you need to be aware about. Now,
100
00:13:02.395 -->
00:13:06.870I don't think this is a good excuse not to use AI because
101
00:13:07.030 -->
00:13:10.870by now, as of today, there are multiple ways
102
00:13:11.030 -->
00:13:24.155to ensure that your agentic data engineering coding is perfectly in line with even the strictest possible compliance. For example, I think that all data teams, even in the most regulated industries,
103
00:13:24.315 -->
00:13:29.115use a cloud data platform as their core sender of operation,
104
00:13:29.275 -->
00:13:32.155whether it's Databricks or Snowflake or GCP.
105
00:13:32.395 -->
00:13:47.410And each one of those platforms offers their own LLM endpoints that are governed by the same terms of service as the rest of the platform. And you can use those LLM endpoints for agentic coding. You can use your even favorite agents like
106
00:13:47.730 -->
00:13:52.925Code with the LLM endpoints that are hosted within Databricks or Snowflake
107
00:13:53.005 -->
00:13:53.725for
108
00:13:54.125 -->
00:14:18.725coding. And that means that none of the data leaves your security perimeter. So LLM and data are all within the same security perimeter from the data flow perspective, but also from the legal perspective. And furthermore, we've seen data platforms like Snowflake and Databricks very aggressively roll out their own agents that are even more, I would say, out of the box ready and security compliant because
109
00:14:18.805 -->
00:14:27.444they just by definition work within the same environment and not using data for kind of training something that is completely outside the environment
110
00:14:27.770 -->
00:14:34.090as far as I know, but obviously ask your lawyer. So I think I think that the maturity of those solutions
111
00:14:34.410 -->
00:14:43.450has evolved for enterprise to be able to adopt those tools pretty aggressively. And I think that the adoption is lagging way behind the capabilities right now.
112
00:14:44.195 -->
00:14:49.955Now digging into some of the second and third order impacts of using
113
00:14:49.955 -->
00:14:53.555these agentic workflows for data engineering,
114
00:14:53.555 -->
00:14:55.635there's also the differentiation
115
00:14:55.635 -->
00:15:23.435that we need to think about of not just am I using the AI to do the work of data engineering as far as writing the code and validating it, but what is the role once I move that agent off of my laptop and turn it into an always on mode and give it that goal oriented execution to say, you're actually going to live within the execution path of my data engineering, whether it is operational monitoring to do validation of data as it lands or
116
00:15:23.820 -->
00:15:43.795doing in flight data transformation to do things like adding structure to unstructured data, doing things like entity extraction, and some of the ways that that shifts even just the nature of the work beyond just I'm gonna write a bunch of SQL, write a bunch of transforms, and then verify everything is correct at the end of the day. Yeah, I think in terms of the implications
117
00:15:44.035 -->
00:15:46.115on the data engineer's
118
00:15:46.115 -->
00:15:46.675job,
119
00:15:47.075 -->
00:15:49.075they are quite profound.
120
00:15:49.714 -->
00:15:51.315And I think that the
121
00:15:51.395 -->
00:15:52.915value of a data engineer
122
00:15:53.139 -->
00:16:02.100as it has been for the past ten, fifteen years, as we've seen the rise of cloud data warehouses and big data in terms of writing the code, maintaining the code for data pipelines,
123
00:16:02.260 -->
00:16:04.660is now shifting towards
124
00:16:04.820 -->
00:16:05.700operating,
125
00:16:05.860 -->
00:16:14.025like you said, agents or teams of agents that are performing different tasks that previously would be completely owned by human engineers.
126
00:16:14.185 -->
00:16:18.345And I think that that's a very important shift that data engineers
127
00:16:18.505 -->
00:16:26.070and data analysts and analytics engineers need to recognize and be aware of is because if your current role is
128
00:16:26.550 -->
00:16:35.575writing code and that's a current value prop, that will probably no longer be relevant over the course of this year. And it's just a matter of time. I And don't think the timeline is very long until
129
00:16:35.655 -->
00:16:44.695that type of skill will be completely eliminated by automation. But I don't think that means that we don't need data engineers or we don't need that many data professionals because
130
00:16:44.855 -->
00:16:45.575the
131
00:16:45.655 -->
00:16:46.375agentic
132
00:16:46.490 -->
00:16:48.410data engineering patterns
133
00:16:48.490 -->
00:16:55.449drop the cost of creating data pipelines, managing data pipelines, operationalizing them. And that means
134
00:16:55.610 -->
00:16:56.410that
135
00:16:56.649 -->
00:16:57.370the
136
00:16:57.529 -->
00:17:10.085business can do much, much more with their data. So it doesn't mean that like, Oh, okay, we will just do the same but with fewer people. I think that when I talk to, I would say more forward looking
137
00:17:10.245 -->
00:17:20.809data leaders and CDOs at enterprise, I hear them being very excited about the new capabilities that previously they just weren't able to unlock because their
138
00:17:21.210 -->
00:17:32.405teams were completely bogged down doing the basics. For years, over a decade, all we talked about was how to deliver dashboards, machine learning model and maintain SLAs and data quality for stakeholders.
139
00:17:32.405 -->
00:17:35.445And now because those things can be automated,
140
00:17:35.685 -->
00:17:40.885the types of things that I see data leaders wanting to tackle is
141
00:17:41.130 -->
00:17:50.730incredibly exciting. For example, I've been chatting with one of our customers who runs data platform for a very large parcel delivery service. And they've been talking about how after
142
00:17:50.890 -->
00:18:06.794adopting a modern cloud warehouse and also embracing AI, they're now thinking, okay, we can go way beyond dashboards. We can actually create a simulation for a business so we can simulate every single parcel, we can simulate the bottlenecks. And then instead of being reactive with operational dashboards, being proactive.
143
00:18:06.875 -->
00:18:21.470So having solvers that just run our business based on the data. And so I think this means that the demand for data engineering as a way to deliver high quality data to power data driven decisions is going to actually grow.
144
00:18:21.710 -->
00:18:23.310And in economics,
145
00:18:23.310 -->
00:18:29.154is this famous Jevan's paradox that essentially says that if the price for a given
146
00:18:29.475 -->
00:18:59.445resource or capability drops, we'll actually see more of that being consumed. And you see a lot of talk about Jevan's paradox in the context of GPUs and AI and how the cost AI dropping and people will be using more AI. I think the same is true for the output of data engineering. Because it's going to be cheaper to create data pipelines, we'll see more data pipelines being created, more data products being created. Because I think historically data has been underutilized by businesses in terms of what's possible to do to run businesses more efficiently, and the economics will just create really strong motivation to to do more.
147
00:19:00.005 -->
00:19:10.920Yeah. You're seeing the references to Jevan's paradox a lot in the software space as well of people having that debate over our software engineers going to be completely
148
00:19:11.080 -->
00:19:13.640obviated and removed as a result of AI,
149
00:19:13.720 -->
00:19:14.920but instead,
150
00:19:14.920 -->
00:19:52.875you're just seeing software engineers doing more. I I know in my own work, there are dozens of different little scripts or tweaks or improvements or side projects that I've done in my day job that I wouldn't have otherwise bothered with. I actually just recently had a project that's been waiting on the shelf for about two years that I did a full rate up on. This is exactly how I would do it if I had the time. It's probably gonna take a full time engineer about two months to actually do the whole thing and validate it. And then just last week, for whatever reason, that project came back to my mind, and I said, oh, well, I'm just gonna go ahead and throw the document at my agentic engineering tool. And within two days, I had it complete and validated, and now it's in production.
151
00:19:53.595 -->
00:19:58.235Yep. And so so I think one of the interesting side effects as well of the
152
00:19:59.035 -->
00:20:00.075acceleration
153
00:20:00.075 -->
00:20:02.555of capability and productivity,
154
00:20:02.715 -->
00:20:07.115but also the broadening of who can do which parts of the workflow
155
00:20:07.490 -->
00:20:09.010brings an interesting question
156
00:20:09.330 -->
00:20:10.530about the
157
00:20:10.770 -->
00:20:12.130role proliferation
158
00:20:12.130 -->
00:20:14.850that the data space in particular has been seeing,
159
00:20:15.250 -->
00:20:28.964I think, even just since 2020 where the idea of analytics engineers and machine learning engineers and ML ops and LLM ops data engineers and pipeline engineers and SQL engineers, you've been seeing this fragmentation of specialization because
160
00:20:29.044 -->
00:20:35.445the work is complicated. It does require a lot of domain and technical expertise to be able to do effectively.
161
00:20:35.910 -->
00:20:49.110I'm curious what impact you are either seeing or predicting on just the ways that we think about what the roles and responsibilities are for data oriented professionals and whether we will maybe see a coalescing
162
00:20:49.110 -->
00:20:53.554of role definitions because every person can have a broader
163
00:20:53.795 -->
00:20:58.355scope because of the fact that they're able to get the AI to take on a lot of the heavy lifting.
164
00:20:58.835 -->
00:21:11.110Yeah. Tawai, I think this is an excellent insight and I do think that we will see something similar to what we've seen we have been seeing in the software world where there has been consolidation
165
00:21:11.110 -->
00:21:23.045and at the same time, all of a sudden, the kind of product engineer and product manager, so roles that have been more focused on tying the business
166
00:21:23.045 -->
00:21:26.165problem solving to the actual technical solution,
167
00:21:26.485 -->
00:21:51.304has been elevated massively because now if you're a product manager, if you're product engineer, you don't have to rely on a team of more specialized engineers to get what you need to get done. And I think that just like for software engineers, having more product mindset, wearing more hats, being able to think more strategically, interfacing with people. The same will be true for data space as well. I think that there will be less value
168
00:21:51.385 -->
00:22:27.540in being hyper specialized. For example, in my day of data engineering, we've had streaming data engineering experts who would just work on streaming ingestion pipelines. And then we would have analytics engineers just work on turning the data already ingested into then data products that are consumed by data analysts and data scientists. I think that we will see far more demand for cross functional specialists who can take a business problem and then solve it end to end from the very, very beginning, which could actually start in, Oh, we need new instrumentation and we need to bring in new data streams. All the way to, okay, how this now powers
169
00:22:27.540 -->
00:22:46.265the business through either humans making decisions or increasingly so probably machines making decisions about the business. And I think that has a really important implication on, again, how data professionals need to think about their career evolution. I think that the soft skills, the business acumen, the domain expertise
170
00:22:46.425 -->
00:22:54.230will start to matter way more than highly specialized technology skills and product thinking as well. Because ultimately
171
00:22:54.310 -->
00:23:17.515the word beta product has been kind of en vogue in the past few years, but it never really truly picked up. I think now it's actually worth revisiting because everything we do as data practitioners, every single streaming pipeline or, you know, machine learning model, it all is in the service of solving a business problem. So it all is some sort of an internal external product that we're we're building.
172
00:23:18.350 -->
00:23:20.190And the natural
173
00:23:20.190 -->
00:23:21.870next question is
174
00:23:22.190 -->
00:23:29.549if the nature of the work and the people who are doing the work collapses down to a smaller number who are doing more,
175
00:23:29.790 -->
00:23:37.164how does that also impact the way that we think about the underlying platforms and infrastructure that we need where
176
00:23:37.325 -->
00:23:38.3642020,
177
00:23:38.365 -->
00:23:40.125maybe starting in 2019,
178
00:23:40.125 -->
00:23:46.219really saw the growth of the whole modern data stack that caused huge proliferation
179
00:23:46.220 -->
00:23:54.139also because of the fact that we had zero interest rates. So VCs were throwing money at everybody with an idea, and now we're in another
180
00:23:54.220 -->
00:24:08.325another phase where a lot of those early movers are getting acquired or put out of business because their adjacencies are being consumed by other systems. I think maybe one of the best examples is
181
00:24:08.485 -->
00:24:11.684the work that's happening with Fivetran and DBT
182
00:24:11.684 -->
00:24:18.019and SQL Mesh where they were all separate tools, they were all separate companies, and now they have all been,
183
00:24:18.500 -->
00:24:23.380aggregated into one company that is trying to own more of the process.
184
00:24:23.780 -->
00:24:27.540And as the actual entities
185
00:24:26.905 -->
00:24:30.024are that are interacting with all these technical layers
186
00:24:30.184 -->
00:24:35.304cease to be human increasingly and are instead AI and agentic workloads,
187
00:24:35.385 -->
00:24:37.465how does that change the requirements
188
00:24:37.465 -->
00:24:52.150of what the systems need to be able to do, what the integration patterns are, what the surface area of that technology stack needs to look like to enable these agentic workloads to be able to execute more effectively and with the appropriate context.
189
00:24:53.425 -->
00:25:03.105Yeah. Well, there's so so much to unpack here, Tobias. But maybe, like, yeah, let let's start with the consolidation. I think the the consolidation of the data stack is part of the
190
00:25:03.345 -->
00:25:22.005more global phenomena. I think that the fragmentation that historically, like you said, been caused by a lot of the available funding I think the available funding is one of the causes, but I think the other is that historically writing software has been very expensive across the board, not just engineering, but also
191
00:25:22.965 -->
00:25:42.780good product managers have always been very rare and expensive and hard to find and hard to nurture and grow. And then for each product team, need to be staffed with great software teams to be able to ship pictures. And so for any single vendor to be able to, let's say, go very deep in a problem or expand into adjacent product
192
00:25:43.100 -->
00:26:10.769area, it always has been quite expensive and a risky bet at many times because you have to invest a lot, this hasn't been your focused area. Do you go there? How fruitful it's going to be? And so that's why many software renderers have been very focused on just their domains and core competencies. And that's why also we've seen a lot of startups pop up that have been solving things that were falling through the cracks among the larger vendors or focusing on niche problems. And now,
193
00:26:11.409 -->
00:26:14.850because the cost of writing software is drastically, drastically
194
00:26:15.345 -->
00:26:26.864cheaper, you can experiment way faster, you can ship MVPs and test them way faster, and you can expand into adjacent product areas that create more value for your target user persona
195
00:26:27.424 -->
00:26:31.920way quicker. And that's far less risky because of the whole compression of the shipping
196
00:26:32.080 -->
00:26:39.760cycle and costs. Just for example, with DataFold, we were able to expand into areas such as data platform costs optimization
197
00:26:39.920 -->
00:26:45.534very quickly despite being a very small team and into the platform migrations here earlier
198
00:26:45.615 -->
00:26:54.735that would not be possible at all without AI. And so I think that's kind of the expansion of the platforms and consolidation of the more fragmented
199
00:26:54.735 -->
00:27:01.820market. That's definitely one force. I think the other forces that we touched on is do we need that many people building
200
00:27:01.820 -->
00:27:19.605data products, building pipelines? We're seeing a lot of layoffs happening at companies that are actually seemingly doing quite well. And so there's definitely a lot of anxiety around, well, do we need the main tech professionals? Do we need that many people in data space? And I don't think we know for sure. I don't think anyone knows for sure, at least
201
00:27:20.085 -->
00:27:22.804not until It's hard to say that yes,
202
00:27:23.525 -->
00:27:35.679will need far, far fewer humans to just run everything Because at the same time, we see that companies that are doing really well and growing really fast, like the big AI labs and a lot of players in the AI space, they are hiring very aggressively.
203
00:27:36.400 -->
00:27:52.695What I think even though they have a lot of AI automation and arguably they're best in class in being able to ship and automate with AI. So I think that maybe a few things are true at once. Think that if what you possess as a data professional is what I would call a commodity skill, like writing SQL, it's just
204
00:27:53.335 -->
00:28:05.800no longer defensible nor remarkable, you are at risk. But at the same time, if you're able to stretch and combine multiple roles and you bring product thinking and you're great at working with people and navigating and understanding business environments,
205
00:28:05.800 -->
00:28:15.794business context and business goals, then I think you will be in demand as much as ever because you can do so much with your impact can be 50x
206
00:28:15.794 -->
00:28:20.595and companies would value that. And I think that means that just
207
00:28:21.395 -->
00:28:25.155the market becomes far less even or kind of uniformly
208
00:28:25.155 -->
00:28:33.820distributed. Now we're probably gonna see a bimodal distribution of data professionals who are adopting very quickly and are very marketable
209
00:28:33.900 -->
00:28:50.975because they possess skills that are valued in the AI world. Then a lot of folks who are kind of lagging behind because their skills are no longer in demand. And I think that's what makes this whole labor market very turbulent for data professionals. And that's why I think it's very important to be in the first camp and not the second camp.
210
00:28:51.830 -->
00:28:58.870Expanding a little bit on what you were referring to earlier as well of because we can move faster, because we can experiment faster,
211
00:28:59.110 -->
00:29:22.059I am no longer willing to settle for, hey, give me a static dashboard that I can look at when I remember to and instead moving to these more proactive use cases. And I think this is the overall dream of what we wanted business intelligence to be of, okay. Great. You've told me something. Now what? And the now what, I think, is more in the loop and more automatic and more exploratory.
212
00:29:22.059 -->
00:29:30.620And I'm wondering what are some of the ways that you are seeing some of the potential for these agentic use cases and asynchronous
213
00:29:30.620 -->
00:29:43.515discovery that can happen, even looking at things like the Orion project from Gravity or the Compass project from Dagster where you just have this agent that's churning through your data asynchronously
214
00:29:43.515 -->
00:30:02.600and finding these little nuggets of insight to say, hey. Did you know this? Or, hey. I just found out this interesting fact. And then being able to take that and turn it into, okay. Now that I know this, what is the next step? And actually having some recommendation of here are the five things that you should try and then being able to actually have the capacity to do more of that experimentation
215
00:30:02.600 -->
00:30:05.864of whether it's AB testing on a website or
216
00:30:06.025 -->
00:30:09.544changing some of the features of a given product to
217
00:30:09.705 -->
00:30:19.999and I'm just wondering what are some of the ways that you're seeing folks leverage this unlocked capability and this the the fact that we're not spending so much time on toil and we can instead focus on these higher leverage elements.
218
00:30:20.480 -->
00:30:36.955Yeah. I think the high leverage elements is really important here, Tobias, because there is also a lot of, I would say, kind of like shallow use cases where just having an agent come through your data and come up with things that look interesting doesn't necessarily mean that it is impactful
219
00:30:37.115 -->
00:30:39.595for the business. I think ultimately,
220
00:30:39.675 -->
00:30:42.955the value of the data comes from the
221
00:30:43.115 -->
00:30:58.419value of decisions that we're making based on that data. And that value could be also quantified through risk. Like If the business is thinking about whether to expand or kill certain products or to invest more money in this particular acquisition channel for its customers,
222
00:30:58.580 -->
00:31:00.980then if the data helps you reduce that risk,
223
00:31:01.140 -->
00:31:03.380that is very quantifiable value.
224
00:31:03.975 -->
00:31:16.215If we're being a little bit more academic from information theory, and I think that it all has to come down to what is the business problem we're trying to solve and what's the best way to solve it. And I think that if we identify
225
00:31:16.215 -->
00:31:19.480these bottlenecks that are currently very manual or
226
00:31:19.480 -->
00:31:20.280decisions
227
00:31:20.280 -->
00:31:41.075made not based on data and better data can help us reduce the risk of those decisions and ultimately get the business more efficient and grow faster, then I think the automation possibilities are completely limitless. Because like you said, we can go from the world where a human will look at a dashboard, make a suggestion to another human who would then make a product decision, who would then
228
00:31:41.555 -->
00:31:44.995write the code and change something about your product or
229
00:31:45.200 -->
00:31:46.160propagate
230
00:31:46.160 -->
00:31:47.600the decision through organization.
231
00:31:47.760 -->
00:31:55.120Now you can have an agent that evaluates the data and makes the decision immediately. Those decisions can be quite diverse.
232
00:31:55.440 -->
00:32:13.095So that kind of automation is not new. For example, ride sharing businesses like Uber and Lyft pioneered automatic decision making and balancing decided markets. And they have been incredibly, incredibly data driven even without AI. But the types of decisions that could be automated were limited to just high frequency
233
00:32:13.175 -->
00:32:18.669use cases like driver passenger matching and algorithmic pricing. But now we can automate way more
234
00:32:19.149 -->
00:32:27.389people heavy and diverse processes ranging from support to getting your marketing resources and in general optimization
235
00:32:27.389 -->
00:32:31.865of the entire business. But again, I think that it all has to come down from
236
00:32:32.265 -->
00:32:35.865the business use case rather than from kind of, oh, let's agent
237
00:32:36.105 -->
00:32:48.179figure out things on its own. I do think that occasionally we can stumble on a treasure chest in the data if we let agents lose, but I would be surprised if that is a systematically winning paradigm than coming from,
238
00:32:48.420 -->
00:32:55.700okay, actually what business needs? I think the other important aspect here is that there's a lot of obviously, agenda coding is important,
239
00:32:55.700 -->
00:32:59.614but it's still a problem to supply these agents with the context
240
00:32:59.774 -->
00:33:12.620because a lot of the context exists outside of database and outside of immediate code base that the agent has access to. It exists in people's heads. It exists in email and Slack and Teams and documentation
241
00:33:12.620 -->
00:33:21.019tools and Excel spreadsheets. And so I think we're still figuring out how to collect all this context and feed that in the agent so that the agent can actually work
242
00:33:21.340 -->
00:33:28.495with the data efficiently. And it is a real bottleneck that I think that we will see being solved over the next year.
243
00:33:29.295 -->
00:33:30.335And so
244
00:33:30.495 -->
00:33:32.815we've discussed all of these
245
00:33:33.054 -->
00:33:33.934wonderful
246
00:33:34.255 -->
00:33:50.860exciting futures that we're looking forward to. And so as an individual contributor or as an engineering leader, what are the things that I need to be thinking about and concrete steps that I should be taking today to make sure that I'm able to actually realize some of this promise
247
00:33:50.940 -->
00:34:28.205and not just get bogged down with all kinds of bugs or errors or problems that are introduced because somebody else with AI is spewing all kinds of problematic code and data into my systems or just ways that we should be thinking about as professionals? What are the skills? What are the day to day practices that we need to be investing in to be able to build that flywheel of letting AI do more of the drudgery and move beyond just, can I do something locally on my machine to what are the confidence building steps that I need to take to be able to actually let the AI run-in that inner loop of my data system?
248
00:34:28.765 -->
00:34:56.085Yeah. What a great question. Well, maybe we should start a little bit with the basics foundation of infrastructure. So in the startup world, everyone runs on really exciting tools and there is lots of great tools in the modern data stack today that are just very easy. And then also they are AI first and very friendly to agents, have MCPs and everyone can move fast and be heavy. But the larger world, the most of the world, data world still runs on legacy data infrastructure.
249
00:34:56.085 -->
00:35:01.045So much so that if you talk to the leaders in the data space,
250
00:35:01.205 -->
00:35:02.725like the leading data platforms,
251
00:35:02.725 -->
00:35:07.830they still estimate that there is fifty to one hundred and fifty billion in
252
00:35:07.910 -->
00:35:16.870data platform spend going into legacy tooling. So those billion dollar companies themselves, they think that there's like 10x to 30x more
253
00:35:16.950 -->
00:35:19.670workloads that are currently locked in legacy platforms.
254
00:35:20.085 -->
00:35:24.805And the risk is that if you are a data team operating on a legacy platform,
255
00:35:25.045 -->
00:35:33.685you can't really take full advantage of AI because those platforms are not AI native. It's really hard to get data out of them. It's really hard to ensure interoperability
256
00:35:34.405 -->
00:35:49.420with the modern tools. Your data stack is fragmented and that just slows you down. So I think that's just a very important foundation. Make sure that you're running on the modern data platform because otherwise you're going to be fundamentally slowed down. Now, the good news is that with AI, migrations
257
00:35:50.005 -->
00:36:01.925to modern platforms are far, far easier and DataFold has kind of pioneered the software first approach to data platform migrations, but we are obviously not the only ones doing this. And it's
258
00:36:02.210 -->
00:36:04.609obvious that with agent decoding,
259
00:36:04.769 -->
00:36:07.810moving code, which is the primary cost of migration,
260
00:36:07.890 -->
00:36:13.329has become much, much easier. And with the cost of data platform migrations
261
00:36:13.329 -->
00:36:14.609and the timelines
262
00:36:14.609 -->
00:36:31.235really plummeting, I think there is just no good reason to be stuck on the legacy infrastructure because I think it will present a very substantial long term risk for your business if you do so. The side effect of this is I think that legacy data platforms are completely cooked like ETL
263
00:36:31.235 -->
00:36:33.075tools and
264
00:36:32.550 -->
00:36:33.910on prem installations
265
00:36:33.910 -->
00:36:34.630because
266
00:36:34.870 -->
00:37:04.420the only reason why they're still in business and still have all these enterprise customers running those legacy software is because of the migration friction. And if that's going away, then I don't think they have any chance to stand against the modern players. So that's one. So first, make sure that your foundations are solid as a data leader. I think the second thing is even, like I said, even at a startup, ensuring that your team takes full advantage of AI is challenge. It's hard. It's hard because not only you're fighting some inertia and
267
00:37:04.900 -->
00:37:27.545people having their own ways about the workflows, but everything is changing so fast. And today you think like, Oh, this coding agent is the state of the art. Tomorrow, a new model comes out and everyone says, This new thing is state of the art. And so there's a lot of noise, there's lot of confusion. And then like you said, there's obviously risks in terms of you don't want to blow things up and there are real security and privacy risks. And so I don't think that
268
00:37:27.730 -->
00:37:29.250there is any perfect answer,
269
00:37:30.210 -->
00:37:46.994really embracing this new world and investing and learning about it and trying things out. I don't think there's a perfect recipe of do this or use this agent or use this model. I think everyone needs to try for themselves and learn what works, what doesn't, and also invest in ways that allow them to do AI data engineering,
270
00:37:47.075 -->
00:37:49.555AI data product building safely
271
00:37:49.555 -->
00:37:50.355and
272
00:37:50.595 -->
00:37:51.395securely.
273
00:37:53.555 -->
00:38:15.545And basically, you need to bring agency to your team, no pun intended. There is no one that will magically provide you a solution, I think, this year at least that will just work and solve all of your problems. You'll have to figure it out. But that process of figuring it out is important. I've seen that teams that invested in it that encouraged their teams to try things, figured out a way to securely deploy
274
00:38:15.705 -->
00:38:30.310AI and let people try it, enable engineering coding. These teams are moving way faster and the gap between teams that are proactive versus just waiting and not investing in education and trying things and piloting new
275
00:38:30.390 -->
00:38:38.375ways is growing rapidly. So I think that's really, really important. And then if you're a leader, we've talked about what are the career
276
00:38:38.695 -->
00:39:15.615implications for, let's say, individual contributors. But for data leaders, think it's also quite important to recognize that if you're not riding the AI wave right now, and I don't want to sound buzzword, if you're not embracing AI with your team, your job is at risk because at some point, the leadership of the organization will recognize that you're slowing everyone down. That's kind of a negative way to say it. But the positive way of saying it is as a leader, you can multiply your impact by 10 to 50x if you invest in making your team fully AI enabled. So I think that's just a very kind of black and white world that we're embracing just because of how disruptive this technology is.
277
00:39:16.095 -->
00:39:35.915And I think too, some of the concrete steps and ways that we should be thinking about how to keep that flywheel moving and get it moving faster is as you're doing your day to day work, if you are interacting with an AI, whether it's Claude code or GitHub Copilot or what have you, anytime you find yourself repeating something,
278
00:39:36.235 -->
00:39:43.915that is a signal that maybe you should start to codify that either into an agents. M d or a Claude dot m d that lives in the repository
279
00:39:44.235 -->
00:39:54.110or codifying it in a skills dot m d so that the agent can incorporate that context for that particular style of workflow and just documenting
280
00:39:54.110 -->
00:39:54.750those
281
00:39:54.990 -->
00:40:06.775workflows and the ways that you work and the types of work that need to be done so that you're not repeating yourself every time and so that everybody on the team is able to take advantage of those context cues collectively
282
00:40:06.775 -->
00:40:20.310rather than it being a single player mode where you as an individual contributor have figured out all of the tricks, and so you're able to move at 50 times your regular pace, but everybody else on your team is left behind. And just even just getting the agent to generate
283
00:40:20.710 -->
00:40:34.255those skills files or agents files to say, hey. Make sure that this gets added to the team context or wherever that might live and just thinking through what are some of the ways that you can accelerate that workflow and reduce the need to do that constant rediscovery
284
00:40:34.255 -->
00:40:45.030of best practices and then even letting the agent loose on your code base to say, hey. Tell me what are some of the patterns that we have established, codify that so that it's easier to understand,
285
00:40:45.110 -->
00:41:17.950and then you can determine is that something that you want to invest in going forward or not. But you can use these tools for more than just writing the code. You can use it for understanding it, doing some analysis of maybe what are some of the areas of duplicative effort that we have across these different tool chains and how and then the other key piece is use those agent decoding tools to help you write more utilities to give you better and faster feedback. So for instance, I'm using a combination of DBT and Superset for business intelligence.
286
00:41:18.190 -->
00:41:33.425And one of the challenges that we've been going through recently is I wrote a utility that lets me actually build a better work workflow of being able to go from QA to production with superset instead of just being point and click. So there are a bunch of YAML files that are fairly opaque and inscrutable to a human,
287
00:41:33.825 -->
00:41:41.500and it's hard to tell is everything lined up. So I wrote I I got Copilot to radio utility to say, hey. Here are all the YAML files.
288
00:41:41.660 -->
00:41:45.180Write a validation script that will look at the dashboards,
289
00:41:45.180 -->
00:41:49.260make sure that all the charts that it references exist, all the charts that are referenced,
290
00:41:49.420 -->
00:41:58.195all the columns that they're looking for actually exist in the datasets and that those datasets are actually appropriately pinned to DBT models so that I can have an end to end confidence building
291
00:41:58.355 -->
00:42:14.329exercise before I ever ship it and just being able to build some of those tools where the agent could actually execute that as it's making changes to say, hey, am I going in the right direction or did I just break everything and I need to back up? Yeah, well, I think Tobias, you are here illustrating
292
00:42:14.329 -->
00:42:17.290a very important point which is that
293
00:42:17.450 -->
00:42:32.215there is effort and there is skill and there is craft in mastering the AI first workflow. And I think this is one of the arguments that I hear thrown a lot is that folks who are maybe more resistant to using AI
294
00:42:32.375 -->
00:42:43.940are labeling people who are very AI first as lazy because like, Oh, sure, you'll just tell AI to do stuff and then there's no craft. I actually think it's the opposite. Coming back from the Harry Potter analogy,
295
00:42:44.180 -->
00:42:47.540yes, you can become a wizard but you have to go to Hogwarts first.
296
00:42:48.420 -->
00:42:50.900It's not something that you can magically
297
00:42:51.060 -->
00:42:56.605wave your magic wand and then all of a sudden great things are happening.
298
00:42:57.085 -->
00:42:57.805Even though
299
00:42:58.045 -->
00:43:08.125the bar for making great things is definitely way lower because of thanks to AI, I do think that there is a skill and you have to invest
300
00:43:08.619 -->
00:43:13.660time and energy and learn how to leverage AI most effectively.
301
00:43:13.660 -->
00:43:17.340And back to our question about what differentiates the
302
00:43:17.420 -->
00:43:18.860kind of most successful
303
00:43:19.020 -->
00:43:20.300data practitioners
304
00:43:20.380 -->
00:43:26.905who are gonna be leveraging AI and will be very relevant in this new economy versus those that could struggle,
305
00:43:27.145 -->
00:43:28.025I think that
306
00:43:28.345 -->
00:43:29.785the AI mastery
307
00:43:29.785 -->
00:43:32.185is a craft and skill that
308
00:43:32.425 -->
00:43:37.785actually is quite defensible and important in the market. For example, I can't imagine any
309
00:43:37.310 -->
00:43:42.590effective data team at a fast moving organization today who wouldn't evaluate
310
00:43:42.670 -->
00:43:47.630new hires in terms of their AI skills. I personally would ask questions,
311
00:43:48.030 -->
00:43:54.665how have you been using AI in your workflow? What have you built to improve your workflow? What have you done to enable your coworkers
312
00:43:54.985 -->
00:43:58.185to improve the workflow? Just to your point, right? Because
313
00:43:58.425 -->
00:44:01.625even though it is magic, it doesn't come necessarily
314
00:44:01.865 -->
00:44:03.960easy. You have to also
315
00:44:04.040 -->
00:44:12.520invest in understanding how it works, invest in education, invest in building tools and scripts for yourself. And I do think that AI improving will
316
00:44:12.680 -->
00:44:29.705probably make some things easier. I think some of the patterns that existed a year ago right now are not necessarily relevant because AI can figure things out more on its own. But I still think there's always gonna be something to learn and to master that can differentiate you from everyone else who hasn't invested in learning and mastering.
317
00:44:30.359 -->
00:44:55.464Absolutely true. And one of the other at least perceived barriers to entry that can happen, particularly if you are in a company that doesn't have free access to unlimited compute or an expansive budget is that there is a nonzero amount of cost involved in using most of these leading edge AI tools. So Cloud co Cloud Code Max, you're talking about $200
318
00:44:55.464 -->
00:44:59.470per user per month, which is in the grand scheme of things very affordable.
319
00:44:59.710 -->
00:45:07.869But if you have a large team, can be quite substantial. And I'm just wondering what are some of the ways that you're seeing teams address some of that consideration
320
00:45:07.869 -->
00:45:08.510of
321
00:45:08.750 -->
00:45:11.630how do I justify this initial expense
322
00:45:11.765 -->
00:45:23.045before I'm actually getting all of the benefits and just some of that balancing act of the catch 22 that you're in where I want to move faster but I can't afford to but I can't afford not to because I have to move faster.
323
00:45:23.445 -->
00:45:29.240Yeah. I think we're starting to see the cost of LLM inference be quite substantial.
324
00:45:29.320 -->
00:45:44.365So for example, a year ago when our customers asked us, well, how worried should they be about LLM inference for data fold features? I would say, well, if you're a multi tenant environment, don't worry about it because we pay for it. If it's a single tenant environment, we use LLM and points, That's negligible.
325
00:45:44.765 -->
00:45:57.885Now it's a very different story because of how much we were able to automate and how much we actually are consuming in terms of LLM costs so that we had to establish kind of LLM FinOps at DataFold because of how significant
326
00:45:58.110 -->
00:46:04.190those costs have become. And I think pretty recently they surpassed the costs of overall infrastructure.
327
00:46:04.350 -->
00:46:09.710All of the infrastructure combined versus the LLM costs, the LLM costs have become more substantial,
328
00:46:09.950 -->
00:46:12.990which I think tells about just how much work is
329
00:46:13.494 -->
00:46:27.020actually being done. The other thing I'll say is that it shouldn't really stop you from automating your work. It certainly doesn't stop us. And I think that if you're running a team of engineers or data engineers and
330
00:46:27.420 -->
00:46:33.900team, all team members are on, let's say, ClotCode Max Plan at $200 a month
331
00:46:34.220 -->
00:46:37.740and they're reaching limits on those plan, well, congratulations,
332
00:46:38.220 -->
00:46:55.530you're running an extremely efficient and impactful team. That's the world we want to be at. I would be far more worried about a team that uses a couple of $15 a month subscriptions for the tools that don't help much, just because again, how much leverage it gives us. Think relative to the cost of engineers,
333
00:46:55.530 -->
00:47:00.010relative to the cost of our attention and our time as humans,
334
00:47:00.090 -->
00:47:25.760this is still low relative to how much we can actually do and how much leverage we're getting. And then again, that's not to say that there could not be completely wasteful AI spend. So if you task an agent with some poorly defined task without guard rails and using the most expensive model, because why not, then well, I wouldn't be surprised if you recap a very large bill. That can happen, happen to us, and will continue to be happening. But I don't
335
00:47:26.720 -->
00:47:29.120see it as being prohibitively expensive for the industry.
336
00:47:29.360 -->
00:47:30.080Furthermore,
337
00:47:30.240 -->
00:47:54.730the advancements in the models and in the efficiency are also quite rapid. So we're seeing open source models like Quen that are maybe lagging behind the frontier, most intelligent models like Opus being quite on par in terms of baseline coding. And those models you can run on, let's say 32 gigabyte Mac these days, not even the top line Mac. And then larger models can fit into more upgraded
338
00:47:54.730 -->
00:47:58.010machines. And I think these are just one
339
00:47:58.010 -->
00:48:00.250off data points, but they are signals that
340
00:48:01.130 -->
00:48:02.970at least for the individual
341
00:48:03.130 -->
00:48:04.810productivity standpoint,
342
00:48:05.369 -->
00:48:15.895LLM costs are far from being prohibited for modern organizations. And I think there's way way way lots of things that we're spending on that are providing far less value than than that.
343
00:48:16.375 -->
00:48:17.335Absolutely.
344
00:48:17.415 -->
00:48:29.500So as you have been going through this journey yourself where you have moved from being a data engineering company as these AI models became more capable,
345
00:48:29.580 -->
00:48:31.020you reoriented
346
00:48:31.020 -->
00:48:39.345your product vision to be AI native and brought the LLMs into the inner loop of what you're offering, and now you're also using it
347
00:48:39.505 -->
00:48:48.945extensively for doing the coding and data engineering within your own company. What are some of the key takeaways that you want to make sure that folks are
348
00:48:49.569 -->
00:48:55.730listening to and aware of and maybe some of the other companies or individuals that you're looking to for inspiration
349
00:48:55.730 -->
00:48:58.210who are maybe a couple of steps ahead of where you are?
350
00:48:58.770 -->
00:49:05.010Yeah. Well, I can speak for the transformation on DataFold. So in terms of how we use
351
00:49:05.705 -->
00:49:11.145AI internally to ship our own product, over the course of past couple of months, we shifted from
352
00:49:11.305 -->
00:49:13.145most of the code was
353
00:49:13.385 -->
00:49:24.960written by software engineers and then Wassamia Automation to none of the engineers actually writing code anymore. They're writing prompts and having conversations with the AI agents, different kinds of agents for
354
00:49:25.040 -->
00:49:28.160writing code and then QA ing the code and then
355
00:49:28.400 -->
00:49:29.200operationalizing
356
00:49:29.200 -->
00:49:30.880that and managing infrastructure.
357
00:49:30.880 -->
00:49:46.425But that has been a very, very important shift. And then we have completely autonomous agents that bring up the whole software stack in the cloud, build a new feature, bring up a preview with a URL that anyone on the team can look at, test everything. And that's pretty much the entire software engineering
358
00:49:46.425 -->
00:49:50.745loop being automated end to end from a, let's say, a linear
359
00:49:50.825 -->
00:49:52.960task all the way to production.
360
00:49:53.040 -->
00:49:59.440And I think that that's the future for sure. I think there's more that we can automate, but it has been profoundly
361
00:49:59.440 -->
00:50:10.205impactful on how much we were able to ship as a company. Externally in terms of what we were able to do for customers because of that is, I think one of the most interesting
362
00:50:10.445 -->
00:50:32.700changes for us as a company has been that pre AI we were selling tools from a SaaS model to data teams. And they would use those tools to be more productive. Let's say more productive at validating their data, more productive at discovering their data, more productive in terms of communicating with the stakeholders about their data. And AI enabled us not just to make those tools better,
363
00:50:32.860 -->
00:50:35.375but to also start offering
364
00:50:35.455 -->
00:50:37.215solutions for customers
365
00:50:37.215 -->
00:50:40.335that instead of offering a productivity gain,
366
00:50:40.415 -->
00:50:46.015offer a complete business outcome. And an example of that is a data platform migration. So historically,
367
00:50:46.015 -->
00:51:08.415to execute data platform migration, you would hire a team of consultants and then they may bring some internal tools that would help them translate some of their code and then there were anything they would do by hand, or they would use some of the native tools from the data platforms. But all in all, it was kind of like an extremely manual effort and the market was kind of separated into software companies building so called accelerators
368
00:51:08.415 -->
00:51:14.655that would then sell that software to service companies who would then use accelerators to perform services.
369
00:51:14.655 -->
00:51:21.060And AI enabled us to combine both in one offering where we have our own software,
370
00:51:21.619 -->
00:51:23.060which essentially is
371
00:51:23.380 -->
00:51:28.820a team of AI agents that have different roles that we deploy to provide migration as an outcome. So in that model,
372
00:51:29.060 -->
00:51:44.289there is no human billable hours that we need to sell, yet we're able to provide a full service where the customer gets migration done, completed as an outcome. So we're kind of replacing both the service plus the tool with full AI automation.
373
00:51:44.289 -->
00:52:05.465And there are other use cases that we are going to launch that are following the similar model where instead of buying tools that make your team more productive to accomplishing certain tasks, you can just buy an outcome of the task being done. And when I say task, I mean a very complete large business outcome. For example, to execute a large data platform migration historically
374
00:52:05.465 -->
00:52:06.745has taken
375
00:52:07.305 -->
00:52:11.545one to three years at the cost of millions of dollars at enterprise.
376
00:52:12.069 -->
00:52:13.990Now it can be done in weeks
377
00:52:14.150 -->
00:52:22.470at a fraction of the price. And there are other problems like that that exist in the enterprise that can be solved in a similar model. So I do feel that
378
00:52:22.630 -->
00:52:27.510that's just one example of one company operating in the data domain. But I do think that this model where
379
00:52:27.775 -->
00:52:29.295you can replace
380
00:52:29.295 -->
00:52:30.095services
381
00:52:30.095 -->
00:52:32.895plus scattered tools with a single offering
382
00:52:33.055 -->
00:52:37.615and sell an outcome that's way faster, way better, way economically efficient
383
00:52:37.855 -->
00:52:40.015with full end to end AI automation,
384
00:52:40.470 -->
00:52:41.830I think this model has
385
00:52:42.310 -->
00:52:44.630a lot of promise in this world and
386
00:52:45.109 -->
00:52:48.070I would expect to see more problems being solved like that.
387
00:52:48.550 -->
00:52:51.350And are there any other aspects of
388
00:52:51.430 -->
00:52:53.830this shift in capabilities
389
00:52:53.830 -->
00:52:54.869and workflow
390
00:52:55.235 -->
00:52:56.755in outcomes
391
00:52:56.755 -->
00:53:06.915and just the overall predictions that you have as we continue to accelerate into this uncertain future that we didn't discuss yet that you'd like to cover before we close out the show?
392
00:53:07.859 -->
00:53:23.494Yeah. I would say one thing, it may be worth a deeper dive, but I think it's still important for folks to kind of dwell on because it certainly was a big realization for me. I I do think that the way we were thinking about data quality in a pre AI world and
393
00:53:23.654 -->
00:53:28.935post AI world is very different because if you think about the last five years,
394
00:53:29.095 -->
00:53:30.295there's been a huge
395
00:53:30.375 -->
00:53:40.770conversation in the data community about how do we ensure data quality? How do we make sure that the products we're delivering to the business are correct? There has been multiple frameworks that were started,
396
00:53:40.930 -->
00:53:46.610open source and features and existing frameworks to tackle data quality. Multiple companies got funded.
397
00:53:47.625 -->
00:53:51.225Hundreds of millions of VC dollars went into the space. DataFold,
398
00:53:51.225 -->
00:54:16.755no exception to this. Data quality has been entire focus area for us for the first three years of the company. And I think all of this is now irrelevant completely because the way we thought about data quality before was, okay, how do we instrument all these different rules and checks and tests and assertions about the data so that we can signal to data consumers that this data is good? But it's always been a moving target because data is changing, the business requirements are changing and
399
00:54:16.995 -->
00:54:40.190I don't think we ever got to the point where a data team would be like, Okay, great. We've invested so much in data quality and we are perfectly happy. I think what happened actually is that we learned how to work with imperfect data over time. And I think that in the AI world, this is irrelevant because ultimately all the solutions we've been building were built for humans in mind, and humans have very limited context and very limited capacity for processing.
400
00:54:40.270 -->
00:54:50.295So how can we curate one dataset that can describe this thing that we know is gold? It's good. Humans, please use it. Don't use anything else. Just this dataset.
401
00:54:50.935 -->
00:54:51.415And
402
00:54:51.735 -->
00:54:54.215make sure the task coverage is great. I think with AI,
403
00:54:54.789 -->
00:55:16.735the whole thing flips upside down. You want AI to have access to all of your data, even the data that's imperfect, because none of the data is perfect to start with. And then have AI actually figure out how to provide the better quality answer to your question, considering everything and considering the various degrees of reliability or accuracy or precision of your data points.
404
00:55:16.975 -->
00:55:20.975Now, that is not to say that AI will do a marvelous job
405
00:55:21.295 -->
00:55:51.630if you just throw all the data that you have at it, right? I think this presents its own challenge. But I do think that the way we're thinking about data quality is different in the AI world. What matters in the AI world is that A, AI has ability to act. So we talked about agentic loop, execute queries, evaluate queries, run tools, run tools like dbt. And two, AI has access to all of your data because again, the more data points you have, the more complete picture you can construct. But three, you also need to make sure AI has the right context on that data. So
406
00:55:51.869 -->
00:55:58.030what does each dataset mean? Where it comes from? How does it relates to the business entities and the nature of the business?
407
00:55:58.510 -->
00:56:14.585What about those data sources and how they all interlink? And that third component is actually currently missing. And I think that this year is going be a big year for figuring out how do we feed the right context into the agents so they can actually figure out how to solve end to end problems with us, considering that all data is
408
00:56:14.905 -->
00:56:24.650imperfect. And it's also a very big investment area for us at DataFold, because we had to build what we call data knowledge graph in order to power migrations. And we're going to be offering that to
409
00:56:25.050 -->
00:56:26.330all customers to
410
00:56:26.570 -->
00:56:46.300help their agents provide more context. But in general, I think that's the really big shift from let's write a bunch of tests for humans and curate data for humans and kind of focus everyone on very small pieces to let's embrace the full complexity of data, but figure out how to make AI work reliably with it. Yeah. Definitely a valuable
411
00:56:46.619 -->
00:57:18.540thing to be thinking about where we should be spending our time most effectively and what are some of the pieces that the AI doesn't care about and we should maybe just seed control of to hark back to our earlier points. So for anybody who wants to get in touch with you and follow along with the work that you're doing and the rest of the team, I'll have you add your preferred contact information to the show notes. And, as the final question, I'm sure that your answer has changed a few times since we last spoke, but what do you see as being the biggest gap in the tooling or technologies that's available for data? And I'll add AI management today.
412
00:57:19.260 -->
00:57:21.740I would say that right now,
413
00:57:22.060 -->
00:57:24.780I firmly believe that tooling capability
414
00:57:24.780 -->
00:57:33.095exceeds our ability to adopt those tools. And I think that instead of trying to find the perfect tool, now is the time to experiment
415
00:57:33.415 -->
00:58:03.345with what's available and build your own workflow. Just like you gave an example, Tobias, how you've been using coding agents to write scripts, to automate the workflow and kind of build skills to make those agents more effective at accomplishing more and more tasks. Now is the world where you can fill your own gaps, you can build your own tools. And I think that's a very important mind shift that is still very, very few people in the data community fully embrace, and I would love for everyone lean on this as hard as possible. Go and build your tools,
416
00:58:03.505 -->
00:58:05.184make your workflows better,
417
00:58:05.600 -->
00:58:10.640increase your quality of life at job. There has never been a time like this.
418
00:58:10.960 -->
00:58:14.560And that's that's what I think is missing. Go build your own tools.
419
00:58:14.960 -->
00:58:15.840Absolutely.
420
00:58:16.080 -->
00:58:22.235Well, thank you as always for taking the time today to join me and share your own thoughts and experiences
421
00:58:22.235 -->
00:58:24.635of living out this weird
422
00:58:24.635 -->
00:58:35.515and exciting future that we're all barreling into. So as always, I appreciate you taking the time, and I hope enjoy the rest of your day. Thank you, Tobias. I will let you go back to your agentic coding.
423
00:58:43.980 -->
00:58:48.220Thank you for listening, and don't forget to check out our other shows. Podcast.net
424
00:58:48.220 -->
00:58:57.445covers the Python language, its community, and the innovative ways it is being used. And the AI engineering podcast is your guide to the fast moving world of building AI systems.
425
00:58:58.005 -->
00:59:05.205Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes. And if you've learned something or tried out a project from the show, then tell us about
426
00:59:05.529 -->
00:59:14.250Email hosts at data engineering podcast dot com with your story. Just to help other people find the show, please leave a review on Apple Podcasts and tell your friends and coworkers.