1
00:00:00.100 --> 00:00:10.070
Welcome to the Data Strategy Gurus podcast. In this show, we bring together the brightest minds in the world of data strategy, data management, artificial intelligence, and disruptive technologies.
2
00:00:10.290 --> 00:00:16.750
Thought leaders and experts share their insights, knowledge, and experience on how to stay ahead of the game in an ever-evolving data landscape.
3
00:00:17.180 --> 00:00:23.740
Whether you're a data professional, a business leader, or simply someone who is passionate about the power of data, this podcast is for you.
4
00:00:23.800 --> 00:00:34.780
So sit back, relax, and join us on a journey to explore the world of data, analytics, artificial intelligence, tech, and beyond. Hi, and welcome to the Data Strategy Gurus podcast.
5
00:00:35.120 --> 00:00:42.620
Today, we have Tanya Bragin, uh, the VP of product, uh, at ClickHouse. Tanya, uh, welcome to the show.
6
00:00:42.800 --> 00:00:56.080
Uh, maybe you can help us out and tell us a bit more about your journey in the data and analytics space, and how you came, uh, to be the VP of product at ClickHouse. Thank you for having me on the show.
7
00:00:56.120 --> 00:00:57.520
Yeah, happy to share.
8
00:00:57.620 --> 00:01:09.320
Um, as probably many professionals, I didn't know growing up or even when I was, uh, at the university that I would, first of all, end up in a data space, uh, and certainly I didn't even know what a product manager was.
9
00:01:09.420 --> 00:01:29.440
Um, [lips smack] so what happened was I was in grad school actually, and before that, I spent a few years in consulting, and I was trying to figure out how to combine my deeply technical background, I was studying, um, distributed systems at University of Washington, with, uh, kind of business orientation and practical application of what we were kind of researching at that time.
10
00:01:29.480 --> 00:01:35.840
And my professor recommended looking into product management. And so I was lucky enough to join a startup.
11
00:01:35.900 --> 00:01:45.400
I considered bigger companies, but I was lucky enough to join a Seattle-based startup called ExtraHop, and that's really where I learned what it means to be a product manager in a deeply technical space.
12
00:01:45.940 --> 00:01:49.680
At that point, I didn't consider myself to be in a data space, but of course I dealt with data.
13
00:01:49.740 --> 00:02:00.500
This was a company focused on network analytics, and we actually built our own database to store network telemetry, because at that point, this was embedded devices and it had to be hyper-optimized, uh, to run in that device.
14
00:02:01.210 --> 00:02:12.380
And, um, eventually, I went, uh, to Elastic, a company behind Elasticsearch after, you know, spending quite a bit of time with ExtraHop, and that was really when I realized data is just a huge space.
15
00:02:12.620 --> 00:02:23.200
Um, machines produce data, applications produce data. It needs to be searched, it needs to get, uh, aggregated. There's real time aspects to it, there's historical aspects to it, and frankly, I just fell in love.
16
00:02:23.750 --> 00:02:35.260
Uh, and so my, my time at Elastic, I spent, uh, seven years there. I learned a lot, both about data and open source. And now, now I'm with ClickHouse. This is another company that's also focused on real time analytics.
17
00:02:35.300 --> 00:02:48.060
Happy to talk about it, but, um, you know, the idea is that I really enjoy working, um, on products that help data practitioners and developers build effective applications. And, you know, happy to share my experience.
18
00:02:48.180 --> 00:02:58.290
And again, thank you for having me on the show. Yeah. You, you just said, uh, you, you did some distributed, um, systems, very technical, low level, if I understand. Yeah.
19
00:02:58.290 --> 00:03:05.700
You even helped build the, the, the base layer out of that. But what made you fell in love with, with the data management aspect?
20
00:03:05.800 --> 00:03:17.940
Data is still the level a bit up, but still pretty technical and tangible for, for a lot of people. Yes. Yes. So, um, I guess what I realized, uh, you know, when I was at ExtraHop, we, we took the hard road, right?
21
00:03:18.000 --> 00:03:28.920
We built our own database for a very specific use case, but most, you know, teams don't need that. They don't need to do that. Um, and so they need to adopt an off-the-shelf, uh, data store.
22
00:03:29.000 --> 00:03:38.590
And again, at that time, you know, of course, there were, there were, there were commercial databases. Oracle, uh, is a great example, right? A great database that has helped a lot of teams scale their analytics.
23
00:03:39.060 --> 00:03:47.920
But over time, there's been just this explosion in different types of data stores, including open source options, and it's just, um, such a vibrant landscape.
24
00:03:48.100 --> 00:04:04.440
And so data management becomes really a set of choices of which technology do you adopt for which use case, uh, what are the trade-offs in having kind of a sing- single focus on, on one system that tries to do everything versus best of breed, and helping teams make those decisions.
25
00:04:04.500 --> 00:04:10.300
It's non-trivial, uh, to, to actually make these trade-offs. I talk to a lot of customers that, you know, kind of struggle with that.
26
00:04:10.360 --> 00:04:19.220
You know, they say, "Okay, like, do I, you know, look at the new shiny thing all the time, or do I, you know, really take care about which technologies I adopt when?"
27
00:04:19.260 --> 00:04:32.080
Because it's, it's a non-zero investment for teams, uh, to, to kind of go with the new technology. Yeah. And, and where you say, okay, uh, where you see the market is going, you, you have...
28
00:04:32.140 --> 00:04:43.160
I see, I see the two paths, uh, going. Uh, so the best of breed, so one, one tool that fits all, but you have the hyper-specialized, uh, databases as well, or data platforms what I see developing.
29
00:04:43.560 --> 00:04:53.700
What is your feeling on, on where it's going, or is it really, do we really need to make a choice, and it depends a bit on, on the use case and the industry you're in? Yeah.
30
00:04:54.440 --> 00:05:06.000
So again, my perspective comes a little bit from working with analytical workloads. Uh, ExtraHop, Elastic, now ClickHouse, have all focused on helping customers analyze events primarily. And the...
31
00:05:06.030 --> 00:05:15.780
kind of the way I see the world is, uh, there's source of truth, which are typically objects, and you have to write those into transactional databases. It's very important that those are kept consistent.
32
00:05:16.140 --> 00:05:34.880
But then when it comes to, um, doing something with that data, uh, to either analyze it internally or to build applications for your customers with the focus on analytics, um, event data, which is typically time indexed, is a very different workload from a transactional, um, workload that might be, you know, present in another database.
33
00:05:34.969 --> 00:05:46.820
And so, um, the way I see the world is increasingly, I think a pattern, uh, needs to emerge where teams realize from day one that they need both. Um, I know there's, uh, this concept of HTAP, trying to combine the two.
34
00:05:47.480 --> 00:05:50.100
My perspective on that is it's hard. It's hard to do both.
35
00:05:50.160 --> 00:05:58.870
I'm not saying it's impossible, but it's hard, and in the end, it kind of locks you into a single vendor with their choice of a transactional database and an analytical database in one.
36
00:05:59.640 --> 00:06:07.220
And for such critical choices, I actually do believe a best of breed approach makes sense. Pick a transactional database or databases that make sense for your workloads.
37
00:06:07.720 --> 00:06:20.230
Pick analytical database or databases, um, that are a fit for your workload, and really recognize that you need both-Um, and both are, again, I call them databases, but these are really data platforms because they have to perform at scale.
38
00:06:20.460 --> 00:06:32.580
Usually, you have to deal with clustering beyond a single node. And so this is something that is fundamental, I think, to the way that we, um, build applications, write data, analyze data in the business.
39
00:06:32.700 --> 00:06:47.200
And again, I come at it more from the analytical side, given my background, but what I am seeing is just increasingly the team- teams recognize that they need both, and their choices, uh, in each space are kind of a n-- it's a non-trivial choice, like what technology to bet on and why.
40
00:06:48.760 --> 00:07:00.520
Yeah. You see, um, really when you're pushing the edge, not to the edge, but pushing the edge of the technology, then you're better off with, with the hyper-specialized, uh, solutions, if I understand what you're saying.
41
00:07:01.100 --> 00:07:04.440
Yes. C-correct. Correct. Yeah. Especially for teams dealing with a lot of data.
42
00:07:04.580 --> 00:07:13.280
Like, I can understand if you're a startup and you're just starting out and it, you know, time to market it is everything, and you just really want, like, to pick a simple path, and you're not yet at scale.
43
00:07:13.320 --> 00:07:22.580
You know, like just... Yeah. At that point, we even see people just pick one database. Maybe it's a transactional database. They don't have that much analytical data yet, and they just go with it. And that's okay.
44
00:07:22.660 --> 00:07:28.420
I'm not saying that's the wrong approach. But the-- even a small startup can generate a lot of analytical data very quickly.
45
00:07:28.460 --> 00:07:47.680
So we see a lot of teams quickly recognizing that now they have to migrate this workload, and it becomes, uh, kind of a, a fire drill because they may already have, again, as a smaller team, paying customers, and they're having to go through this, you know, fi- like fire drill migration to something like ClickHouse, you know, say from Postgres, for, for the analytical portion of their workload.
46
00:07:48.240 --> 00:07:54.400
So I, I get the choices, you know, like sometimes you just go with what you know because you're just starting to build. It makes sense.
47
00:07:54.500 --> 00:08:05.030
But for the teams to recognize that you need both, you know, and kind of prepare for that eventual migration, even at an early stage of, say, a startup or a, an, a mid-sized comp-- like, like, company.
48
00:08:05.080 --> 00:08:12.740
Again, machines can generate a lot of data. Applications can generate a, a lot of data. There's lots of public data sets out there. Data is ever present, right?
49
00:08:12.760 --> 00:08:24.439
And so, um, our capacity to generate more just keeps growing every day. Yeah, exactly. In your recent article a-as well, you talked about the unbundling of cloud data warehouses.
50
00:08:24.999 --> 00:08:30.720
Can you elaborate a bit on, on what that means, uh, for businesses, uh, today? Absolutely.
51
00:08:30.860 --> 00:08:41.200
So now we're going a little bit more of from like developers building applications to internal teams, um, you know, dealing with data, oftentimes historical data that ha-- that their businesses have amassed over the years.
52
00:08:41.240 --> 00:08:50.009
Of course, this data keeps growing because new data sets come in. Um, one thing that, um, I've seen in the past, you know, like five or so years, right?
53
00:08:50.040 --> 00:09:01.539
I mean, so, um, data warehouse, just kind of like rewinding, of course. Uh, these are the systems that were established, you know, decades ago to help businesses centralize important data for decision making.
54
00:09:01.800 --> 00:09:11.300
This was a time when, um, applications each kind of had their own embedded database, so how would a business, you know, kind of keep track of that data centrally, uh, for effective decision making?
55
00:09:11.400 --> 00:09:18.540
This is why a data warehouse was born. It was a really critical development for businesses that wanted to be and needed to be data-driven.
56
00:09:18.600 --> 00:09:29.140
And vendors like IBM, Oracle, others, eventually Hadoop as like the open source alternative, really helped businesses, uh, become data-driven and start valuing da-data as part of their business.
57
00:09:29.800 --> 00:09:35.360
Uh, but for a long time, these systems remained, uh, on premise, and I would say they were a bit of a walled garden, right?
58
00:09:35.400 --> 00:09:48.240
So unless Oracle or IBM had an integration with some other part of your ecosystem, you kind of had that, um, one place where you did everything, and you couldn't necessarily leverage it for, you know, many use cases outside of whatever was enabled by that vendor.
59
00:09:49.240 --> 00:10:06.880
And so to me, like, when, uh, the shift to the cloud started happening and, uh, cloud data warehouses like Redshift, Snowflake, BigQuery, um, enabled teams to move these workloads to the cloud environment, this was a fundamental shift in kind of ability of teams to leverage this data in multiple use cases.
60
00:10:07.380 --> 00:10:16.680
Because cloud is just natively more enabled for integrations, and, uh, the team started saying, "Okay, this data set is now in the cloud. What should we actually be using it for?"
61
00:10:16.780 --> 00:10:25.080
And sometimes they said, "You know, let's start building new types of applications with this data. Let's build analytical, analytical applications maybe that we haven't built in the past, in the past. Why not?"
62
00:10:25.800 --> 00:10:34.720
Um, but this is where, you know, um, if you think about the kinds of applications that we all use today, they're very different from, you know, ten years ago, twenty years ago.
63
00:10:34.859 --> 00:10:45.300
Um, back when, uh, traditional data warehouses were built, they were optimized for more batch-oriented, ad hoc analytics workload. You know, you configure a report, you run it.
64
00:10:45.400 --> 00:10:55.700
Um, and today, uh, if you're talking about building applications both for external and internal users, it's increasingly expected to have the kind of UI that is data-driven but very interactive.
65
00:10:55.880 --> 00:11:04.800
You have charts, and you're filtering them, or you have a search bar, and you're searching for something. It's not a report that you schedule, or it's not a query that you run and that you come back to kind of see it.
66
00:11:05.360 --> 00:11:15.900
Those workloads still exist, but when, you know, uh, users in the cloud using cloud data warehouses, uh, now try to build these interactive applications, they run into challenges.
67
00:11:16.210 --> 00:11:29.920
And so what we're seeing now, and that was-- that's what we called an unbundling of the cloud data warehouse, is that, um, teams are looking at their cloud data warehouse, and they're basically saying, "Okay, which of these workloads still make sense to keep in a data warehouse or a data lake?"
68
00:11:30.000 --> 00:11:34.610
There's kind of this question of like, should I be using a data lake, data lakehouse, data warehouse?
69
00:11:35.080 --> 00:11:45.300
And which ones maybe make sense to run in a real-time database or a real-time data warehouse, as I call it, uh, for these interactive analytics. And I'm happy to talk about the choices.
70
00:11:45.340 --> 00:12:01.760
I mean, there are some technical constraints on the current, um, cloud data warehouse architecture that makes it difficult to adopt for these interactive data-driven applications, and why the teams are having to make the choice to kind of move some of these workloads out for a more optimized data store for real-time analytics.
71
00:12:02.920 --> 00:12:12.700
Yeah. Interesting that you bring that on. I mean, traditionally, the data warehouses, batches, we were already happy to see some, some static information and see how last week went.
72
00:12:12.880 --> 00:12:26.130
But this-- these days with online, a lot of data, e-commerce, you, you have the tendency, and to my feeling as well in newer architecture, you should look at, uh, having that real-time aspect or near real-time aspect- Mm-hmm...
73
00:12:26.160 --> 00:12:33.060
and not with the batch load. I think it comes as well in how you use the compute and, and the storage of the cloud data warehouses.
74
00:12:33.100 --> 00:12:39.830
And if you do that in exactly the same architectural approach as the on-premYeah, your costs are skyrocketing. Yes.
75
00:12:39.860 --> 00:12:53.980
So can you dive into how these real-time data warehouses differ, uh, are different from the traditional ones, and why businesses really need to focus on that real-time aspect? Sure, yeah. So there's a couple of things.
76
00:12:54.150 --> 00:13:04.390
W- you mentioned one of them, which is the freshness of data. Um, you know, h- how, how often is a, is a new dataset loaded? Is it every week? Is it every day? Is it every hour, or is it continuous?
77
00:13:04.680 --> 00:13:11.519
Uh, like is it real-time or near real-time, you know, where it's seconds or minutes f- like fresh in terms of new events coming in.
78
00:13:11.560 --> 00:13:21.660
And again, frankly, for some use cases, it's not important to have continuously loading data. You know, for some use cases, it's okay to have a, a data loading strategy that is a little bit more batch oriented.
79
00:13:21.700 --> 00:13:25.960
But for instance, for, uh, these data-driven applications... Like, I'll, I'll give you an example, right?
80
00:13:26.040 --> 00:13:38.790
So I'm a product professional, and I use SaaS services to help me understand what users are coming to my website, and, you know, when it comes to, like, my cloud business that I'm building, like what is the experience, uh, that the users are having right now.
81
00:13:38.880 --> 00:13:46.120
I need to break down multiple datasets, uh, to kind of understand up to the minute performance of the product that I'm responsible for building inside ClickHouse.
82
00:13:46.650 --> 00:13:54.240
And, you know, many professionals today rely on tools like this. For these types of tools, it's very important that the dataset is continuously loading. So that's number one.
83
00:13:54.420 --> 00:14:02.960
Uh, real-time data warehouse combines both continuously loading datasets up to the minute at least kind of data, and then also historical look back.
84
00:14:03.020 --> 00:14:14.660
Because as much as I wanna see the experience the users are having right now, I need to go back, you know, weeks, months, sometimes years if I'm, if I need to answer a question that's a little bit more business intelligence oriented in nature.
85
00:14:15.440 --> 00:14:24.900
So that's number one, um, kind of how, how fresh is your data. The second one is how concurrent, um, is, is, is like the number of users that your, uh, data warehouse is meant to support.
86
00:14:25.700 --> 00:14:37.360
In traditional data warehouses, including, uh, the cloud data warehouses, the assumption has been that these are internally facing workloads and, you know, maybe there's a dozen users, maybe there's even up to 100, but not like thousands.
87
00:14:37.500 --> 00:14:41.120
You know, there's not like necessarily thousands of clients coming in.
88
00:14:41.200 --> 00:14:50.690
Whereas if you're building a SaaS offering with this kind of externally driven workload, this-- these could be thousands of users, you know, coming to, to use your SaaS application.
89
00:14:51.000 --> 00:14:59.760
Of just a very different, um, set of concurrent user sessions to handle, and, you know, some cloud data warehouses just even just hard limit you.
90
00:14:59.860 --> 00:15:11.200
Others, you know, you could try to s- scale it out to that, but this is where you run into a challenge of cost, which I'll talk about in a minute. Um, and sort of the last part is, uh, latency.
91
00:15:11.260 --> 00:15:20.350
Again, how interactive is the query? 'Cause in the end, if I'm building an application and the user is meant to be interacting, uh, there's something about, like, you know, our...
92
00:15:20.780 --> 00:15:28.460
Like we detect latency at about like a second. Like, ideally, you want like a sub-second latency in terms of, like, clicking a button and something happening.
93
00:15:28.500 --> 00:15:38.280
We can tolerate, like maybe up to like a second and a half, two seconds. It'll feel sluggish, but it'll still be interactive enough. Once you go into like five, 10 seconds, it's slow. People are not go- going to...
94
00:15:38.480 --> 00:15:45.840
Like, think about yourself. You come to a website and you search. If it takes 10 seconds for a reaction, you're not going to use that- Mm-hmm... in an interactive manner.
95
00:15:45.900 --> 00:15:55.960
So if you're building an interactive app, that latency of sometimes very complex queries on a lot of data has to be ideally a second or below, millisecond as we call it.
96
00:15:56.400 --> 00:16:09.380
So those are the three things, fresh data, concurrent sessions, you know, up to potentially thousands, and sub-second latency of complex queries on, you know, sometimes terabytes of data, uh, if, if not more.
97
00:16:09.440 --> 00:16:30.080
So those are the, the requirements, and again, you can see how this workload, um, you know, is, I, you know, it can be thought of as specialized, but actually a lot of applications today have this exact characteristic, and a lot of businesses, uh, to differentiate in the marketplace are building these types of applications both for external and internal users to kind of, uh, make decisions faster.
98
00:16:31.980 --> 00:16:38.010
Yeah, I think these are very good guidelines, what you give on the freshness, concurrency, and, and the scalability.
99
00:16:38.040 --> 00:16:49.900
And as well, where people previously were only building systems for internal usage, where they now expose it to, uh, external users, which comes into play where you have more concurrent users at the same time.
100
00:16:50.200 --> 00:16:58.380
So your business, uh, usage of the platforms become differently- Yeah... uh, in such a way. How do you see as well, uh, generative AI?
101
00:16:58.460 --> 00:17:06.060
It's, it's all over the place, uh, right now, and large language models are tapping into the data management space and, and the strategies.
102
00:17:06.080 --> 00:17:17.560
Do you think they will help us become faster data-driven, and how do you see that happening? So generative AI is an interesting, um, space because it's new to a lot of us.
103
00:17:17.780 --> 00:17:22.880
Uh, we're still trying to figure out, first of all, what kind of applications we will-- will we be building using this technology.
104
00:17:23.160 --> 00:17:33.780
The obvious ones businesses are exploring right now is, you know, simply better chatbots or, you know, kind of better help systems as- assistance, and those are non-trivial new applications to build.
105
00:17:33.800 --> 00:17:38.720
But I think frankly, we just don't even know what new applications, uh, will emerge from the space.
106
00:17:39.360 --> 00:17:48.020
So a couple of things, I guess, for data teams to keep in mind and what we're seeing, like customers asking us about where to apply our technology there, 'cause there's a few stages.
107
00:17:48.540 --> 00:18:02.380
Um, all teams have data that could, could potentially be used for training, so there is a question of ad hoc analyzing existing data and determining which datasets you might be using for training your own models or maybe tuning existing large, large lang- language models.
108
00:18:02.780 --> 00:18:09.360
This is ad hoc analytics. Sometimes it's even local, right? And so for that, do you need like a real-time data warehouse? Probably not.
109
00:18:09.400 --> 00:18:14.680
You probably have existing ad hoc analytics tools that you use to kind of reach out to your data lake.
110
00:18:14.700 --> 00:18:22.080
You know, like this is again, like whatever as an analyst you're using right now for local or semi-local analysis is probably what you need to continue using.
111
00:18:22.860 --> 00:18:32.910
Um, but then once you decide on this dataset, uh, you may need to build a data pipeline for kind of transforming these datasets on the way to actually training models. This is platforming work.
112
00:18:33.280 --> 00:18:43.220
This is again, where potentially a real-time data warehouse helps because this needs to happen, um, you know, kind of in a continuous manner. And again, you may need visibility into these pipelines.
113
00:18:43.300 --> 00:18:52.580
You may need to, again, have reporting. And so some of our users started using us kind of as an event platform to transform this data and analyzing this data, right?
114
00:18:52.620 --> 00:19:00.680
Because again, it's just-- it's yet another data problem. Um-So that's on the m- like, model producing side, which data sets you use, and how do you get them to the models?
115
00:19:00.820 --> 00:19:11.060
On the model consuming side, there's a question of, um, do you need to do something... Like, you have, you know, a model that, that you're able to leverage as part of your application, but how do you leverage it?
116
00:19:11.160 --> 00:19:14.580
And sometimes you have to-- There's this concept of vector search. What is that?
117
00:19:14.700 --> 00:19:30.620
This is basically a way to take an existing data set, transform it into vectors, so that you can do a certain, um, uh, type of search that, that hasn't been done before, a certain type of lookup, basically, that is, um, you know, e-e-exactly leveraging these large language models.
118
00:19:31.120 --> 00:19:38.320
This process of taking existing data, transforming it into vectors, storing vectors, and then searching those fast, that's the, the, the consumption side.
119
00:19:39.020 --> 00:19:59.700
And right now, the big questions for the data team is, again, talking about specialization or not, do we need to adopt a very specialized vector database, or will our, our existing databases, you know, Postgres, Elastic, ClickHouse, others, develop sufficient sort of vector search capabilities that we don't have to, because we have data gravity, uh, in these systems?
120
00:19:59.840 --> 00:20:05.640
And I think the jury is still out. Again, uh, you know, I kind of see it both ways. If you have a super large- Mm-hmm...
121
00:20:05.660 --> 00:20:10.980
scale, super specialized vector search use case, probably you should look at one of these specialized databases.
122
00:20:11.020 --> 00:20:32.300
But if you don't, um, you know, at least like where we see our customers right now, if they're using ClickHouse for fraud analytics at scale, and they have heuristic-based fraud analytics, and they just need to add, you know, some maybe even lower scale vector search, like, way to do fraud analytics, they may just leverage the same database, assuming that performance characteristics of vector search that we added is sufficient.
123
00:20:32.440 --> 00:20:40.380
You know, we're not gonna be leveraging GPUs, we're not gonna be doing it at gigantic scale, but it may be sufficient for your use case. So I think the teams will have to make that choice.
124
00:20:40.600 --> 00:20:50.100
Comes back to, like, where does it make sense to adopt a specialized database, and where does it not? And the teams will have to, I think, evaluate it workload by workload. Yeah.
125
00:20:50.160 --> 00:20:59.200
It's, it's kind of choosing starting up with your Excel data set, and if it's not enough, then you scale up to a, a real- Yeah... data platform, a real database. Uh, but it's...
126
00:20:59.560 --> 00:21:14.170
I see as well the, the tendency of a lot of the vendors for, for the databases, uh, they're embedding as, as well some, some vector, uh, capability in, in their platform because they see a lot of people are integrating the large language models in the platform- Exactly...
127
00:21:14.170 --> 00:21:16.760
to build on top of the data what they already have.
128
00:21:16.780 --> 00:21:32.340
But I see a different, uh, approach as well in helping, uh, the vector databases and the generative AI models as well in doing data management, so spotting, uh, data quality issues, uh, optimizing your, your, your data pipelines and so on.
129
00:21:32.420 --> 00:21:38.900
So that's another, uh, direction what I see, and not purely on, on the technical level. Absolutely right. I had forgotten about this one. Exactly.
130
00:21:38.960 --> 00:21:42.780
Like, now, like this concept of observability, you have to now bring it to this use case.
131
00:21:42.820 --> 00:21:52.720
I men- I mentioned the, the data pipelines, but the consumption side as well needs to be monitored and observed, and that's a really hard problem. So at ClickHouse, again, we built an AI assistant.
132
00:21:52.760 --> 00:21:59.260
And, again, as a product person responsible for it, my biggest question is not even just, like, at which latency are requests being served.
133
00:21:59.269 --> 00:22:08.460
That, of course, like, I wanna make sure, again, like, it's interactive, this AI assistant loads fast enough. But my biggest question is actually, are the suggestions being served to users actually valid?
134
00:22:08.600 --> 00:22:16.580
Like, are they ga- gaining value from my suggestions, or are they just wrong? Is the model hallucinating? And that is actually a still a really hard problem for teams to solve.
135
00:22:16.640 --> 00:22:20.280
Like, we can collect data, but unfortunately, actually, there's data privacy concerns there.
136
00:22:20.320 --> 00:22:28.160
Like, I don't wanna be collecting all of the queries that our users run because, like, again, for me, like I, I may- maybe I shouldn't be looking at that.
137
00:22:28.240 --> 00:22:33.880
And so that question is actually very hard to answer for teams, especially building externally facing products.
138
00:22:33.980 --> 00:22:39.260
Um, and I think that's gonna be really interesting to see what kind of innovation happens here because I need answers to those questions.
139
00:22:39.320 --> 00:22:52.820
If I'm gonna introduce, you know, something that, for instance, helps ClickHouse cloud users, uh, you know, who don't know, like some of our specialized sen- sings to come up with queries faster, I wanna make sure these queries are actually valid and performant.
140
00:22:53.100 --> 00:22:59.600
And, like, answering that question is non-trivial right now, even if you collect the data. [chuckles] Yeah. So, [laughs]
141
00:22:59.800 --> 00:23:11.080
so you're u- you're using that as well to, to analyze what users are doing on the platform to help you build a better product, but then we turn... well, we, we stumble again into data privacy and- Yes...
142
00:23:11.180 --> 00:23:20.020
all those regulations. Yeah. Uh, it's, it's a hard choice. Uh, it's very helpful. You can build better products, but it's building that trusted source. Yes.
143
00:23:20.630 --> 00:23:31.420
And well, giving that information to you as a, as a product vendor to help me serve better- Yeah... um, yeah, I think it's gonna be a, a fun discussion in the future. Exactly. Exactly right.
144
00:23:31.440 --> 00:23:38.970
In many ways, like, when I build these features right now, I put a lot of trust. Like, right now we're using, uh, off-the-shelf models. I'm putting a lot of trust in those models today. Mm-hmm.
145
00:23:38.980 --> 00:23:45.000
Like, we've tested it, of course, internally, but again, like, we don't know how, how well they're working at scale for our customers.
146
00:23:45.540 --> 00:23:54.450
So far, this is a beta feature, as you mentioned, you know, and, and, like, there's always caveats, and you... of course, you always opt in to any sort of data collection. Um, but yeah. Mm-hmm.
147
00:23:54.450 --> 00:24:03.100
I, I think that's g- it's gonna be interesting. Like, the observability piece of it is gonna be very interesting in the coming years. I agree. Yeah. It's, it's where we draw the line for privacy.
148
00:24:03.240 --> 00:24:11.580
Am I willing to share, uh- Mm-hmm... my, my query, what I'm sharing? I have that same aspect when I'm playing around with ChatGPT. You think, "Am I training the model, or am I really- Yeah...
149
00:24:11.600 --> 00:24:21.080
benefiting from what I'm getting?" So it's both ways on where you say, "Okay, I'm kind of beta testing this, this, uh, this product, uh, but it's returning me value as, as well."
150
00:24:21.920 --> 00:24:34.248
Uh, Sonja, you're also an advocate, uh, for open source. How do you see this, uh, approach influencing the, the future of data analytics?Yeah. So this is my second open source company, but I probably should rewind again.
151
00:24:34.288 --> 00:24:42.528
When I joined Elastic, I didn't know much about open source. I came... Again, ExtraHop was, um, we used open source, but our solution was a B2B commercial solution.
152
00:24:43.038 --> 00:24:49.208
And so with Elastic, it was really eye-opening for me, how powerful the open source community is in terms of innovation.
153
00:24:49.308 --> 00:25:04.028
What I would say, like, the biggest benefit of adopting open source and being, like, open inside your organization and data m- like, management strategies to open source, uh, being leveraged, is that you can do experimentation at scale without paying licensing fees up front.
154
00:25:04.068 --> 00:25:23.588
Of course, if you're hosting open source, you know, on, on your machines, like, you're paying for the infrastructure, but the point is, before you've proved- proven out a use case, you can take an open source data system, deploy it, and just experiment and just sort of prove the value to the business before you, again, go, uh, all in on some commercial platform and start paying these licensing fees.
155
00:25:24.188 --> 00:25:30.628
Um, what I've seen is that it allows teams and individuals within larger teams sometimes move faster on ideas.
156
00:25:30.768 --> 00:25:40.768
So again, generative AI, I think we'll see a lot of open source leveraged to experiment with it because, again, there are less or no strings attached to kind of an open source way to experiment with something.
157
00:25:41.458 --> 00:25:50.608
Then I think for the data teams, the question is, do we keep the open source experiment or do we move it into an existing data platform? And again, each team will have to kind of make the determination themselves.
158
00:25:51.058 --> 00:26:01.548
Sometimes, especially if, like, the open source-based solution already got productized, like, why fix what's not broken, then you may just say, "Okay," like, "we leave this one-off and, you know, we get, like, some commercial support.
159
00:26:01.648 --> 00:26:09.318
Maybe we move to a hosted model." What I've seen at Elastic is usually if that happens, this platform may eventually start getting leveraged for other use cases.
160
00:26:09.628 --> 00:26:18.588
Because the reason you've kept it around is probably it has something unique about it in the way it serves its workloads, and usually data platform teams find other users of those workloads.
161
00:26:18.628 --> 00:26:23.828
So what we've seen actually work very well is if you go that route, do a lunch and learn within your organization.
162
00:26:23.868 --> 00:26:31.968
You might find out there's other usage of this open source, and you can just onboard them all onto this platform over time. So you can consolidate and...
163
00:26:32.008 --> 00:26:52.288
But it's just a natural channel for innovation, and I just would encourage teams to be open to, like, open source, uh, uh, solutions kind of being leveraged, experiment with, but then sort of h- follow up on how you productize it within your organization, h-how you make sure it's, you know, managed in, in a compliant way, in a way that's actually supportable by the data platform team.
164
00:26:53.908 --> 00:26:59.308
Yeah. It's, it's a, it's a, it's a, it's a platform. It's a tool for enabling innovation.
165
00:26:59.408 --> 00:27:08.278
Uh, and, and you brought up a-an interesting point as well, where you say lunch and learn and, uh, involve other people to make benefit of it, or they have- Yeah...
166
00:27:08.308 --> 00:27:11.258
different ideas on the use cases and what they can bring in.
167
00:27:11.368 --> 00:27:24.588
I saw in a lot of companies as well, uh, especially public s-services, they are kind of more afraid of open source because they don't have anybody, uh, which they can hold responsible- Yeah... for the software in there.
168
00:27:25.048 --> 00:27:37.148
Uh, who do we hit on the head if something goes wrong, uh, and, and where can we go to, to get that real support? Yeah. They solve it by intermediate agencies that take up that, that, that responsibility.
169
00:27:37.728 --> 00:27:46.848
Uh, do you think that that could be a blocker, uh, for the future? Well, so it depends on the open source. Um, a lot of open source do have a commercial backing company.
170
00:27:46.888 --> 00:27:55.248
That was the case with Elasticsearch, and eventually Elastic became the, the backing entity for it. Um, and then with ClickHouse, again, there's ClickHouse Inc. behind the open source.
171
00:27:55.288 --> 00:28:05.698
So w-when that's the case, you always have the ability to go to the vendor that's, um, you know, um... Most open source projects ha-ha-have, like, a, a company that's the major contributor to it.
172
00:28:06.128 --> 00:28:15.208
Even if it's part of foundation, there's usually s-somebody, so-some entity driving the innovation in that open source project, even though there's a, if there's a major set of contributors coming from the outside.
173
00:28:15.248 --> 00:28:17.208
So find that main contributor.
174
00:28:17.268 --> 00:28:29.068
If there's a commercial entity behind it, they may offer both training services, professional services, support, um, you know, we do, Elastic did, uh, obviously, you know, like with any technology, that is, like, table stakes.
175
00:28:29.668 --> 00:28:36.598
Um, increasingly with open source, there is also a hosted model that these commercial vendors provide. And again, it's not for everyone.
176
00:28:36.688 --> 00:28:43.928
Um, we recognize that many of the people that a-adopt ClickHouse will continue to self-manage it, and that's great. Like, we want them to be successful in that model.
177
00:28:44.348 --> 00:28:52.777
But for some teams, especially if they're already self-managing in a cloud provider, right? So say you're self-managing in, uh, AWS in a- Mm-hmm... on an EC2 instance.
178
00:28:53.308 --> 00:29:02.548
Why not just buy that as a hosted offering through, say, even AWS Marketplace from the vendor? Now you don't have to, you know, deal with upgrades. You don't have to deal with, um, ongoing maintenance.
179
00:29:02.708 --> 00:29:11.008
Often these cloud, hosted cloud offerings actually have an evolved architecture that's easier for both maintenance as well as kind of running at scale.
180
00:29:11.088 --> 00:29:19.478
So as an example, ClickHouse went, uh, with a separated storage and compute architecture in the cloud. There's automatic scaling of compute, um, both up and down.
181
00:29:19.608 --> 00:29:28.568
So from a cost perspective, actually, a cloud could be cheaper because, again, on a self-managed version, you don't have some of these capabilities, and you're kind of overprovisioning for peak use.
182
00:29:28.948 --> 00:29:36.928
So carefully examine what makes sense for you to stay self-managed, to go with a hosted model. Again, not everybody can, uh, due to, again, data privacy.
183
00:29:37.188 --> 00:29:48.298
Um, but definitely take a look at, like, what the vendor behind the open source provides and evaluate if it makes sense for your team. Yeah. So really taking benefit of the knowledge that has been put into the- Yeah...
184
00:29:48.308 --> 00:29:51.728
the managed service. Yeah. It also leverages community- And what image... Mm-hmm... I just...
185
00:29:52.068 --> 00:29:58.138
Sorry, I forgot, like, one, one other subject because I didn't know much about what an open source community is before Elastic and before ClickHouse.
186
00:29:58.188 --> 00:30:07.698
I mean, this is basically everybody that's adopted it and that's using it at scale and, um, like with any technology, but especially in open source, there's a really strong culture towards sharing.
187
00:30:07.898 --> 00:30:13.668
I think that that's one of the way users give back, is they publish technical blogs. They say, "This is how we use it. This is how we solve these problems."
188
00:30:14.288 --> 00:30:22.688
Uh, there's always going to be a community chat, uh, that is free and supported both by kind of the creators of the project and the community. So definitely leverage that.
189
00:30:22.768 --> 00:30:27.128
Like, that is one aspect that's really vibrant about open source, and if you can participate.
190
00:30:27.648 --> 00:30:34.658
Another thing that's come, come back again after, you know-COVID, like we're all now a little bit more open to meeting in person, is in-person meetups.
191
00:30:34.938 --> 00:30:43.958
Often they run in almost every city globally, and so you actually have ability to meet with some of these experts locally that are maybe leveraging it and talk about your experience, right?
192
00:30:43.998 --> 00:30:48.788
So for instance, like how do you operationalize open source? Like, how do you make some of the choices I just made?
193
00:30:48.818 --> 00:30:57.938
Leverage others in your region, um, to kind of meet with experts and, and see what their experience has been like. Sorry. Yeah, just wanted to mention that part. [laughs] No, no, no.
194
00:30:57.978 --> 00:31:05.518
I'm really intrigued and, and I'm happy that, that you share this, this community working that is so valuable to a lot of people. What you...
195
00:31:05.558 --> 00:31:14.758
Even sometimes in organizations, what, what you are missing out, people are afraid of sharing because they're afraid of losing their position. Mm-hmm. And that's why they keep everything to themselves.
196
00:31:14.798 --> 00:31:23.578
But that's something I like as well by, by sharing it. Uh, you see sh- startups are, are more open to that and hel- uh, willing to, to help the community.
197
00:31:23.588 --> 00:31:35.668
Uh, so definitely there is still a big future for open source and backed up by, by commercial organizations. Yes, this is the, the way to go and, and share more knowledge instead of, uh, putting a patent on- Mm-hmm...
198
00:31:35.698 --> 00:31:40.578
and everything and getting big money out [laughs] of that. So that's a bit my vision I, I have on that.
199
00:31:41.358 --> 00:31:52.598
Based upon what we've been discussing, what are the emerging trends that you see in the data and the analytic space that people working in the data and analytic space, uh, should be aware of? Yeah.
200
00:31:52.678 --> 00:32:02.838
So we already talked about the, like, im- the rising importance of analytical systems. And, um, again, maybe I'm biased, but again, I think that HTAP, you know, it's, it's hard to make it work right.
201
00:32:02.918 --> 00:32:11.038
Like, even, you know, like talking with users that leverage Snowflake, you know, a wonderful vendor, right, that, that now has Unistore. It's hard to build...
202
00:32:11.078 --> 00:32:19.678
Like, and it's, it's, it's an amazing analytical system, but like building a transactional store right next to it, like, it's just hard. It's a hard problem. Why not specialize? You know?
203
00:32:19.818 --> 00:32:26.518
So again, my perspective is use a, ch- choose, you know, choose a transactional platform, choose an analytical platform, combine the best of the breed.
204
00:32:26.958 --> 00:32:32.278
But d- definitely, like rising importance of analytical, like systems is something to keep in mind.
205
00:32:32.378 --> 00:32:40.378
Um, on the analytical side, we already talked about unbundling, and one thing that I forgot to mention then is the aspect of cost because I talked about the three challenges there.
206
00:32:40.778 --> 00:32:44.998
Uh, the freshness of data, the, uh, concurrency of queries, and the latency.
207
00:32:45.658 --> 00:32:53.548
You know, to be completely transparent, you can solve that in a data, a cloud data warehouse with just throwing more compute at the problem, but then your costs rise- Yeah...
208
00:32:53.548 --> 00:33:00.438
you know, like not exponentially, but, uh, what we've seen is like anywhere three to five to 10X, which is really significant.
209
00:33:00.478 --> 00:33:09.538
And if you can solve a problem, you know, i- in a much more efficient way, you know, like usually kind of where I see switching costs start to make sense is anywhere over 2X, basically.
210
00:33:10.078 --> 00:33:14.837
Uh, you know, because you, you still, you have to move the data set out. You may need to organize it differently. You have to learn a new technology.
211
00:33:14.878 --> 00:33:21.898
But if your costs at scale are more than 2X kind of covered by this move, um, the team should carefully look at it.
212
00:33:22.038 --> 00:33:31.638
And again, what we see is, again, nothing bad to say about Snowflake, Redshift, and BigQuery because these systems, again, enable teams to take critical workloads, move them to the cloud.
213
00:33:31.958 --> 00:33:56.087
But when you're looking at real-time analytics and kind of the, the, the cost, uh, benefit of moving a workload out, just keep in mind that like the teams I s- see today, the way they are solving, uh, their real-time needs inside a cl- cloud data warehouse is providing more and more compute or provisioning more and more compute, but the costs are often three to five to 10X of what you could get, uh, if you took a real-time data warehouse instead.
214
00:33:56.807 --> 00:34:06.138
Um, so anyway, so that's, that's on the analytical side. The last thing I wanted to mention, um, there's this concept of like X as a data problem, right? Like, so there's...
215
00:34:06.158 --> 00:34:16.158
But in the end, like if you're using some point solution and it's driven by data, you know, there's a question of like, do you leverage an off-the-shelf solution, or do you build an internal data platform for it?
216
00:34:16.218 --> 00:34:17.778
Especially for bigger organizations.
217
00:34:18.378 --> 00:34:27.138
And the trend that we're seeing right now in observability is kind of like it started off as something you kind of leveraged, um, something like Splunk, you know, say for log analytics.
218
00:34:27.778 --> 00:34:37.118
Um, then you ha- had a trend towards something like Elastic maybe, you know, uh, to adopt. Then there was Datadog, again, kind of into SaaS. Right now we're seeing a little bit of a backlash.
219
00:34:37.158 --> 00:34:43.928
Pe- teams again asking themselves for observability, does it really make sense for us to leverage SaaS solutions, or is it another data problem?
220
00:34:44.158 --> 00:34:49.468
Does it make sense for us, at least for a subset of the workload, again, to go with a best-of-breed solution?
221
00:34:49.598 --> 00:34:59.348
At least at ClickHouse, we're seeing quite a few users, especially with bigger teams that have data platform teams, explore building, um, their own observability stack.
222
00:34:59.658 --> 00:35:07.718
With ClickHouse being part of it, we don't, we're not a full solution, but ClickHouse could be, for instance, your log store. You may leverage something like Grafana for UI, maybe another UI.
223
00:35:08.238 --> 00:35:31.638
But teams that have, um, the skill set already to manage event data platforms are carefully looking at their SaaS bills and basically are saying, "Okay, like if we have something like Datadog, again, a wonderful solution, but if it's from a cost perspective, not the right thing for our business, maybe we should treat observability as a data problem and kind of again, unbundle that basically and move that into our data platform."
224
00:35:31.718 --> 00:35:40.868
So just wanted to mention it. Uh, again, the way I see it is these cycles of bundling and unbundling, they kind of, the pendulum swings back and forth. There's no right answer.
225
00:35:41.098 --> 00:35:50.718
In the end, every organization has to figure out what's right for them. Yeah. Be- being in the data space for more than 30 years as, as well, [laughs] I've seen that happening.
226
00:35:50.898 --> 00:36:01.498
Uh, and I was just thinking about that, where you say it's really a cycle, and depending on the new technology that, that becomes available, you have the tendency to go left or right, uh, in this respect.
227
00:36:01.938 --> 00:36:13.538
And, and how do you see the evolution of data management, uh, in the coming years? A, a lot of companies are still shouting, "We want to become data-driven," but don't really- Yeah... really get it right away.
228
00:36:14.038 --> 00:36:27.730
A lot of challenges on data management in itself, the, the, the frameworks, the, the management, but on the techna- technology level, what, what do you see is, is happening, and will it become easier?That's a good question.
229
00:36:27.790 --> 00:36:39.710
I think the biggest challenge, um, internal users face to adopting, uh, like, systems like... I, I, I can tell like internally for us, we leverage ClickHouse as a, a data warehouse. Of course we do, right?
230
00:36:39.790 --> 00:36:45.290
We, we drink our own champagne. [laughs] One of our biggest challenges, of course, is, well, what is this data actually?
231
00:36:45.330 --> 00:36:57.710
Like, having, like, an omnipresent data catalog that would tell anybody in the human terms how to actually even, like, think about what this data point represents is still the hardest part between, like, the computer and human interface.
232
00:36:58.050 --> 00:37:10.430
The catalog technology is still not, not standardized. I wish it was. I think we-- I wish we all had just, like, this one understanding of, like, how a data catalog is built, and then every vendor integrates with it.
233
00:37:10.470 --> 00:37:21.390
The closest I've seen there, um, in the data lake side with the Iceberg, um, kind of like, f- community- Yep... they're, they're advocating for this standardized open c- data catalog. This is what I miss.
234
00:37:21.570 --> 00:37:25.879
This is what I wish, you know, the, the, the data community really kind of aligned on.
235
00:37:25.990 --> 00:37:34.810
I think from a human perspective, like leveraging data, this would really help for internal teams to have this as a standard everywhere, and again, all vendors sort of aligning with that standard.
236
00:37:35.550 --> 00:37:37.060
Um, I guess you asked me about trends.
237
00:37:37.110 --> 00:37:56.730
My prediction, I would say, is that this technology does get developed, uh, because internally, again, like, it's gonna be just business critical for all of us to leverage data, for instance, to, uh, build better applications that are enabled by machine learning and AI, so it's gonna become business critical to enable more and more people to work with the data to discover what data sets to train.
238
00:37:57.250 --> 00:38:03.170
So I think that is going to change and evolve, and there's gonna be industry-level conversations about it.
239
00:38:03.230 --> 00:38:11.590
The other prediction I have is that, uh, we're all gonna be thinking very deeply about where to leverage some of these large language models to help external users, right?
240
00:38:11.630 --> 00:38:21.330
Again, my AI assistant that I've built is just a beta product. I'm thinking very hard next year is what else to build for my users to f- to help them, right? Adopt our cloud, to learn ClickHouse.
241
00:38:21.870 --> 00:38:31.150
Um, that's always the biggest challenge in adopting new products is learning. I think AI is gonna help us learn faster and just kind of achieve proficiency faster.
242
00:38:31.250 --> 00:38:43.610
Um, but again, it's an unsolved problem, how to leverage those technologies. Yeah, yeah. Challenging times, but, uh, I think the, the, the- Exciting times, I would say... the speed will just in- [laughs] Yeah. Yeah.
243
00:38:43.650 --> 00:38:52.570
Exciting. More exciting than challenging. [laughs] That's how I experience themself, but, uh, the speed is, is just going up. But we're very well supported by the tools. That's, that's what I see.
244
00:38:52.690 --> 00:38:58.310
Still a lot of problems to solve as well, so we're not yet there. Uh, but it's, it's really exciting.
245
00:38:58.990 --> 00:39:10.950
We had a lot of topics we've been discussing, so I'd like to ask you a few questions, and you just have to say the one or the other. Okay. So what do you think? Cloud data warehouses or on-prem data solutions?
246
00:39:12.670 --> 00:39:23.170
Cloud, but in the middle. I, um, I'm gonna explain my, my answer. So again, our, our customers, when it comes to analytical data, I think they prefer the ease of use of cloud, uh, systems.
247
00:39:23.250 --> 00:39:30.510
But one of the challenges with the hosted model is that it goes against data gravity if you already have a lot of data in your VPC.
248
00:39:30.590 --> 00:39:39.830
Um, and also again, from a privacy perspective, if you're building applications for external users, you don't wanna give that data to third parties. So a new model is evolving. That's kind of a combination of both.
249
00:39:39.970 --> 00:39:45.970
Bring your own cloud, as is often called. Uh, it's not new. Databricks has been, in many ways, kind of the pioneer there.
250
00:39:45.990 --> 00:39:54.809
But the new architecture there for analytical system vendors like us is to deploy the data plane in your VPC and deploy the control plane in ours.
251
00:39:54.850 --> 00:40:02.410
So you have the combination of, like, the ease of use for SaaS, uh, with the benefits of data staying in your VPC.
252
00:40:02.670 --> 00:40:17.990
So I think cloud, but bring your own cloud type of model is really where the analytical space is gonna go. Yeah. And then the next one. Early bird or a night owl? Early bird, and I'll explain that one.
253
00:40:18.030 --> 00:40:24.390
For me, so I've been working... I, I... So both ClickHouse and Elastic are distributed teams, fully distributed. Europe, Asia.
254
00:40:24.570 --> 00:40:30.750
I'm based in, in, um, uh, the West Coast of the US, uh, so, so in, in the San Francisco Bay Area.
255
00:40:31.370 --> 00:40:43.790
And, um, w- I've noticed early on that my European colleagues, in order to work with me and teams, you know, based on in the US, had to stay up late to, to, um, you know, like even during dinner time to kind of meet with us, and I don't like that.
256
00:40:43.850 --> 00:40:50.190
I think, you know, it's very important for teams have to have balance, work-life balance, and so I was thinking about how to really correct for that.
257
00:40:50.730 --> 00:40:58.570
So when I manage teams, first of all, I tell, uh, my teams based in Europe, "Block your dinner time. Do not give that up. Like, this is your family time. This is your time to go to the gym.
258
00:40:58.650 --> 00:41:06.510
This is your time to take a walk and be sane." On the flip side, how do you meet with me? Like, it kind of leaves only, like, deep evenings. Like, so I get up at 5 a.m.
259
00:41:06.690 --> 00:41:16.780
actually, and I meet with my teams if I need to at 5 or 6 a.m. I've completely adjusted to this lifestyle. I go to sleep very early. Oftentimes, my daughter is like, "I'm not tired." I'm like, "I'm tired."
260
00:41:16.890 --> 00:41:26.870
[laughs] It's like, "We're gonna go to sleep." [laughs] Um, so definitely early bird, but because... Not because of preference, it's because just distributed teams, it takes work to, you know...
261
00:41:26.950 --> 00:41:38.130
And it has to, it has to be on both sides. It takes two to tango, and I think it's very important that everybody recognizes they have a hand, like, to play, uh, in making distributed teams work well in a sustainable way.
262
00:41:40.050 --> 00:41:53.690
Yeah, fully understand. I'm in the same kind of position, uh, working globally as well. Real-time analytics or batch processing? Real-time analytics. [laughs] Batch processing only when- Yes, Maria... yeah, of course.
263
00:41:53.770 --> 00:42:02.330
I mean, look, again, coming back to, like, the, the reasons, uh, I think more and more applications require it. It's not... We don't do it because it's fun. It's actually harder to build.
264
00:42:02.350 --> 00:42:07.450
It's actually more expensive, right? But more and more applications simply require it, and so that's kinda how I see the world.
265
00:42:07.490 --> 00:42:19.290
I like to say that real-time apps are eating the world, you know, both in, in, in kind of external apps and internal app side. Yeah, definitely that's, that's the, the way where we are going.
266
00:42:19.330 --> 00:42:30.030
That's, that's what I see as well. Open source software or pr- proprietaria solutions? Open source. I mean, again, that's my bias.
267
00:42:30.250 --> 00:42:38.978
Uh, but we talked a- about the challenges of sort of like adopting open source unless there's a commercial vendorSomehow affiliated with it, right? So again, it's a balance.
268
00:42:39.138 --> 00:42:46.618
I think open source is great for innovation experimentation. In the end though, teams will need commercial entities behind them in some model that works.
269
00:42:46.788 --> 00:42:56.458
And sort of like I think that we are still in the evolution where we're trying to find which models work best with an open source distribution model from a software perspective. But yeah, I'm betting on open source.
270
00:42:57.657 --> 00:43:06.798
[laughs] Next one. What do you want to do? Invent a new technology or redis- well, rediscover a lost art?
271
00:43:08.838 --> 00:43:19.638
Ooh, that's a tough one, uh, because I feel like every time you invent a new technology, almost every time you're rediscovering a lost art. Uh, that's happened multiple times. Um- [laughs] So yeah.
272
00:43:19.918 --> 00:43:29.338
I actually think, uh... I, I'm gonna say invent a new technology though. I think there's still enough to invent that's new. At some point- Yeah, but your tweak is interesting where you say they're pretty close related.
273
00:43:29.638 --> 00:43:46.378
Yeah. It's, it's like the, yeah, old wine in new bags, what they say. Yeah. Mm-hmm. Yeah. [laughs] Next one. The time travel to the past or to the future? Ooh. Hmm. I think the future. Yeah.
274
00:43:46.418 --> 00:43:57.278
Yeah, I think the past, I mean, there's a lot of roman- like romanticism about the past, but I think, like, I can't imagine right now a life without a smartphone. Like, I can't. So future. For me, it's future.
275
00:43:57.458 --> 00:44:12.998
[laughs] So you're quite happy with the, with the new developments and the new technology which, uh, should enlighten our lives. Yeah. And a final one, AI-driven insights or the human-centric analysis? Um,
276
00:44:14.458 --> 00:44:26.258
AI, I would say. I, I'm, I'm pretty bullish. I'm pretty bullish on AI. I don't think it's a fad. I don't think it's hype. Um, I think what, what... the breakthrough that we have seen in the past year is just incredible.
277
00:44:26.278 --> 00:44:31.478
It's just incredible. How it gets leveraged, I mean, that's gonna come back to humans leveraging it. I don't think AI is gonna take over.
278
00:44:31.638 --> 00:44:35.578
Like, I'm not worried about that at all, frankly, but I'm excited about what it brings.
279
00:44:35.838 --> 00:44:50.298
I think that, I think we will figure out how to use it responsibly, um, and I think that, you know, uh, we're all going to benefit from this leap. Definitely bullish on AI. [laughs] Thanks. So what, what...
280
00:44:50.358 --> 00:44:57.448
finally, what drives your passion for, for data and analytics outside of your passion, uh, of your professional life?
281
00:44:58.178 --> 00:45:08.318
So what I really like is discovering new things that are driven by data, and sometimes it has to do with these inflection points. My husband and I, we have this tradition. He cooks me omelet on Saturday mornings.
282
00:45:08.378 --> 00:45:14.718
We have a coffee. That's when we often catch up. We're both in tech, right? So we're both kind of deep in our jobs, and then we kind of come up for air and we talk.
283
00:45:15.238 --> 00:45:20.838
And one of the things that we like to look at together sometimes is there's a Reddit called Data is Beautiful. I don't know if you've ever seen it.
284
00:45:21.218 --> 00:45:30.078
But it's basically a Reddit where members post interesting, just amazing charts driven by data, and sometimes these data visualizations are just mind-blowing, right?
285
00:45:30.138 --> 00:45:36.168
Like, and so I don't know, I just looked at the top one, uh, today and just randomly. It says, "How couples meet in the US."
286
00:45:36.658 --> 00:45:46.098
And you see all these channels kind of through family, through friends, work, bar, grade school, neighbors, college, go way down, and then online going straight up, right?
287
00:45:46.118 --> 00:45:56.618
And again, I've been out of the dating scene forever. Like, I, I'm happily married now for 15 years plus, you know? [laughs] But it's just amazing to still look at data because the trend is so clear. And so I love data.
288
00:45:56.758 --> 00:46:02.118
I think data is beautiful. Um, and that's what I enjoy about it outside of my professional life.
289
00:46:03.718 --> 00:46:15.758
Yeah, it's really the visualization of the data and then seeing the trends and understanding what, what is happening really. But it's, it's a good takeaway. So Data is Beautiful on one of the Reddit, uh, knots. Okay.
290
00:46:15.858 --> 00:46:30.998
Let's, uh, have a look in, into depth. Uh, good takeaway. So Tanya, finally, to wrap up, data connects us all, but music connects us as well. So what is your favorite band or type of music? Okay. So this one's easy.
291
00:46:31.678 --> 00:46:35.718
I am, uh, totally just a huge fan of Taylor Swift. She is incredible.
292
00:46:35.958 --> 00:46:45.948
So how I, I, I sort of lear- learned about her, I was at Elastic and I still remember Shay, our, our CTO, CEO of Elastic, talking about Taylor Swift, and I was like, "Isn't that, like, a pop singer?
293
00:46:46.018 --> 00:46:50.358
Like, why are you talking about her?" And he was like, "No, no, no. You have to understand that, like, this woman is incredible.
294
00:46:50.418 --> 00:46:58.798
Like, you have to just go listen to, like, the breadth of, like, talent that she exhibits through, through her work." And I remember at that point sort of, like, I didn't quite get it. I s- you know?
295
00:46:58.938 --> 00:47:09.298
And so over the years I started listening more and more, and I'm just transformed by her music. My daughter and I are both, um, you know, really into, like, just the music that she has. My daughter is 10.
296
00:47:09.678 --> 00:47:20.138
We'll listen to it together. It's really empowering for her as, you know, a, a, a girl growing up right now to see this amazing female artist coming up from humble beginnings, like really making in this.
297
00:47:20.208 --> 00:47:30.398
And, and she worked hard for it. Like, I mean, like, there's stories of her, like, playing the guitar until her fingers were raw and like, like, this is, like, what it takes, I think, to succeed. So it's very inspiring.
298
00:47:30.738 --> 00:47:39.058
Um, it's really amazing that she is Person of the Year, uh, this year. I think it's very appropriate. Yeah. So to me, this one's easy. Taylor Swift. Yeah.
299
00:47:39.548 --> 00:47:49.558
[laughs] Not, not only the music, but, but the story behind and, uh, the passion and, uh, the drive to get where she wants to- Exactly... uh, where she is right now. Exactly. Yeah. Yeah. Great one.
300
00:47:50.038 --> 00:47:59.358
Tanya, it was so great having you on the show. Uh, maybe you can share as a last note where people can find you online, uh, on Twitter, LinkedIn, your website.
301
00:47:59.738 --> 00:48:05.558
Absolutely, and maybe we can add it to the show notes, but yes, I'm very active on Twitter, um, as well as m- Medium.
302
00:48:05.598 --> 00:48:12.578
I have a personal Medium where I read some articles, more on the meta, what is it like to be a product manager? What is it to, like, to work with open source pricing models?
303
00:48:12.718 --> 00:48:21.827
So if you're interested in those topics, you can find me on Medium. Uh, and finally, um, I also, you know, I'm on LinkedIn, so happy to add it, but I'm kind of less active there.
304
00:48:21.858 --> 00:48:30.538
It's really more for professional networking. But m- would be very happy for folks to reach out to me if they want to chat or learn more. [outro music] Okay. Thank you very much, Tanya. Thank you.
305
00:48:30.738 --> 00:48:35.378
Thank you for having me on the show. Thank you so much. Thank you for joining us on this awesome podcast.
306
00:48:35.678 --> 00:48:43.338
As senior executives, data and analytics architects, and AI professionals, your time is valuable and we appreciate you choosing to hang out with us.
307
00:48:43.378 --> 00:48:49.658
If you liked what you heard, please give us a thumbs up, hit the subscribe button, and leave a comment. We love hearing from our audience.
308
00:48:49.898 --> 00:49:05.418
Don't forget to spread the word on social media, and let's continue to drive innovation in the industry together. Thanks for listening, and we'll catch you on the next episode. [outro music]