WEBVTT
NOTE
Transcription provided by Podhome.fm
Created: 7/6/2024
12:45:38 PM
Duration: 2726.446
Channels: 1
1
00:00:11.264 -->
00:00:15.525 Hello, and welcome to the Data Engineering podcast, the show about modern data management.
2
00:00:16.384 -->
00:00:20.260Legacy CDPs charge you a premium to keep your data in a black box.
3
00:00:20.720 -->
00:00:26.180RudderStack builds your CDP on top of your data warehouse, giving you a more secure and cost effective solution.
4
00:00:26.560 -->
00:00:31.115Plus, it gives you more technical controls so you can fully unlock the power of your customer data.
5
00:00:31.574 -->
00:00:32.074Visitdataengineeringpodcast.com/rudderstack
6
00:00:34.614 -->
00:00:36.875today to take control of your customer data.
7
00:00:37.240 -->
00:00:52.254Your host is Tobias Macy. And today, I'm interviewing DeVerus Brown about the impact of real time data on business opportunities and risk profiles. So, DeVerus, welcome back to the show. For people who haven't listened to your previous episode where you introduced Maroxa, can you just give a bit of an introduction?
8
00:00:52.875 -->
00:00:58.960 Yeah. Thank you again for having me on. I know it's been I think we talked at what at the beginning of the COVID, and it's been,
9
00:01:00.220 -->
00:01:00.720wow.
10
00:01:01.100 -->
00:01:02.000Time flies.
11
00:01:03.155 -->
00:01:06.455Yeah. We are, we've evolved a a little bit.
12
00:01:06.835 -->
00:01:09.735Still the same mission. We want to make real time
13
00:01:10.170 -->
00:01:12.110data the default for any organization,
14
00:01:12.410 -->
00:01:13.710but, we have
15
00:01:14.170 -->
00:01:15.770changed our focus to,
16
00:01:16.170 -->
00:01:16.670empowering,
17
00:01:17.450 -->
00:01:19.070soft regular software engineers
18
00:01:19.875 -->
00:01:22.775to utilize real time data with their,
19
00:01:23.395 -->
00:01:25.495favorite programming language and existing
20
00:01:25.955 -->
00:01:26.935developer workflows
21
00:01:27.315 -->
00:01:29.495to saw innovate and solve problems.
22
00:01:29.860 -->
00:01:31.720I think that that's really what we do.
23
00:01:32.340 -->
00:01:35.800We built a a platform that allows you to ingest
24
00:01:36.500 -->
00:01:37.000process,
25
00:01:38.135 -->
00:01:40.475orchestrate, and stream real time data,
26
00:01:41.015 -->
00:01:42.395with just regular code.
27
00:01:43.175 -->
00:01:53.960You know, kind of a big middle finger out to the the rest of the drag and drop low code, no code world where, you know, you you you kinda automate copy paste, and then you have to add another point solution
28
00:01:54.615 -->
00:01:59.115on at the end of that to actually get usage out of that data. Like, you can do everything,
29
00:01:59.494 -->
00:02:00.475with the code,
30
00:02:00.935 -->
00:02:12.790you know, in in place and not have to rip and replace anything from a from a platform or a process standpoint. So that's Maroxa in a nutshell. And how about DaVaerys in a nutshell?
31
00:02:13.475 -->
00:02:15.495 How about just you? Yeah. Me?
32
00:02:16.435 -->
00:02:18.775 CEO, cofounder, you know, Maraksa.
33
00:02:19.875 -->
00:02:24.170Yeah. That's that's pretty much it. That's it's all encompassing. And I'm originally from Chicago.
34
00:02:24.550 -->
00:02:25.930Went to University of Illinois.
35
00:02:26.230 -->
00:02:28.410Been out in the in the valley,
36
00:02:29.270 -->
00:02:33.145in total for probably 13 years now. I was here too, moved away,
37
00:02:33.605 -->
00:02:34.005then,
38
00:02:34.565 -->
00:02:43.140you know, I kinda came back, when I had a company that that got acquired by another company out here. So yeah. I mean, I've been a product manager.
39
00:02:44.240 -->
00:02:45.300So if you use
40
00:02:45.840 -->
00:02:48.820Windows I mean, not Windows, but Microsoft Azure,
41
00:02:49.360 -->
00:02:49.860Zendesk,
42
00:02:50.395 -->
00:02:50.895Heroku.
43
00:02:51.275 -->
00:02:59.775You know, you use the Devaris stack in some way, shape, or form. My whole life has been should say all life, but professional life has been mostly focused on
44
00:03:00.100 -->
00:03:02.040empowering developers to be more productive.
45
00:03:02.580 -->
00:03:06.885And so a lot of places I've worked at, that's that's what I've done and
46
00:03:07.364 -->
00:03:09.545not stopping that while I'm doing my rocks.
47
00:03:11.685 -->
00:03:20.569 You've mentioned already a bit about kind of what it is that you're building there and some of the ways that the focus and goals of the platform and company have evolved. And
48
00:03:20.870 -->
00:03:32.105in terms of the target customers, you mentioned that you're now focused on software engineers, developers. I'm wondering if you can give a bit of nuance as to are there any particular kind of industries
49
00:03:32.405 -->
00:03:34.550or problem areas that you
50
00:03:34.850 -->
00:03:42.944are seeing a lot of adoption or that you have been kind of focusing on kind of developing towards? Or if there's a particular
51
00:03:43.325 -->
00:03:44.845avenue that you have found,
52
00:03:45.165 -->
00:04:01.375works well for your entree into a given company, whether it's bottom up from developers who are tasked with just make this thing work, and then they find Meroxa and say, oh, hey. This does everything for me, or is it more of a kind of top down, or do they meet in the middle? Yeah. Right now is is completely top down. Like, the
53
00:04:01.675 -->
00:04:06.130 basically, what we found is is that, you know, we do really, really great work for defense,
54
00:04:06.530 -->
00:04:10.790in the highly regulated industries. So because of, you know, their requirements
55
00:04:11.090 -->
00:04:13.830to be secure in the moment all the time,
56
00:04:14.370 -->
00:04:14.770and so,
57
00:04:15.855 -->
00:04:18.355you know, banks, insurance, health, defense,
58
00:04:19.055 -->
00:04:19.555intelligence,
59
00:04:20.095 -->
00:04:21.075that type of stuff,
60
00:04:21.695 -->
00:04:32.909that's really where we found a a a general nexus into the, you know, center of gravity around, like, the people who utilize us and gravitate towards us. I mean, really, any large organization
61
00:04:33.210 -->
00:04:46.919now, we basically go in and say, like, look. Your data team is siloed off, you know, inundated with with requests from all parts of the business and they're frankly overwhelmed. So what does it look like to turn your entire engineering organization
62
00:04:47.699 -->
00:04:51.879into a data team using the Maroxa tools? And you don't have to rip and replace anything.
63
00:04:52.485 -->
00:04:56.585And, you know, and they can do it with the with the existing code. And it's like
64
00:04:57.765 -->
00:05:01.385right? Like, everybody's just like, tell me more. So,
65
00:05:01.730 -->
00:05:10.205you know, especially in this time when when people have been, you know, laying folks off and all of that. Right? Like, every single company is a data company,
66
00:05:10.604 -->
00:05:12.865whether you like it or not. Right? And so
67
00:05:13.245 -->
00:05:16.384that really is a is a gift and a curse for us because,
68
00:05:16.925 -->
00:05:25.470you know, everybody uses different data in different ways. Everybody stores it in different ways. Right? Like, you know, I go into some spots, and it's just like everything's in Excel.
69
00:05:25.930 -->
00:05:36.715Or, you know, you go to some place and there's, like, we have a microservice for everything. Right? And that's just I don't know, man. It's it's it's wild to see, but really all they were doing is just trying to reduce that complexity
70
00:05:37.270 -->
00:05:47.675and give people the tools to to innovate with real time data. And, you know, that's, by and large has been the reason why people have been able to to, you know, build on a motive relationship with us and our platform and the things that we're doing. And, data
71
00:05:55.175 -->
00:05:58.150 that is flowing data that is flowing at us in real time, continuously,
72
00:05:58.610 -->
00:05:59.110unbounded,
73
00:05:59.490 -->
00:06:01.590whatever nomenclature you want to use,
74
00:06:02.050 -->
00:06:04.550generally, that requires a lot of upfront
75
00:06:04.985 -->
00:06:17.320technical and infrastructure investment. You know, usually, that means, okay, we're going to use Kafka, or lately, there are a few more entrants into the market for, you know, whether you wanna use Pulsar or Red Panda or what have you. And, you know, the overall
76
00:06:17.620 -->
00:06:21.080ideal of I want to be able to build real time applications
77
00:06:21.975 -->
00:06:25.115When it first came into the realm of possibility, it was usually
78
00:06:25.575 -->
00:06:29.355only for the big tech companies of the world, you know, Netflix, Google,
79
00:06:29.780 -->
00:06:37.640Facebook, etcetera, because they had all of the engineers who built all these systems from the first place. They had the investment money to be able to put into actually building these systems,
80
00:06:38.165 -->
00:06:38.985and so
81
00:06:39.365 -->
00:06:44.585everybody else was saying, oh, I want real time too, and would start down the path, realize that it was prohibitively
82
00:06:45.130 -->
00:06:48.910complicated and expensive, and then settle for some kind of half measure.
83
00:06:49.210 -->
00:06:57.655And I'm wondering what you have seen as the overall shifts in the industry, in the technology, in the kind of level of understanding and sophistication
84
00:06:58.435 -->
00:07:03.590across kind of your standard engineers that have made this a tractable problem
85
00:07:03.970 -->
00:07:14.675 and have brought real time into the realm of possibility for a possibility for a larger audience. Yeah. I mean, everybody everybody, if you look at, all of I mean, that's really, like, the reason why we exist. Right? This is that we looked at,
86
00:07:16.509 -->
00:07:25.169you know, the writing was on the wall and saw that all these big companies are using real time data as their competitive advantage. Right? You mentioned Netflix. Like, could you imagine,
87
00:07:25.555 -->
00:07:30.194you know, you it's late at night or it's the weekend, you bring somebody over for some Netflix until
88
00:07:30.754 -->
00:07:33.895and you gotta wait 10 minutes to get a movie recommendation because
89
00:07:34.360 -->
00:07:39.020somebody set up a a a data app that is, you know, doing full table scans
90
00:07:39.400 -->
00:07:40.220and compiling,
91
00:07:40.760 -->
00:07:54.470you know, a pre you know, doing models and recommendations and all that type of stuff, right, on your data warehouse. Like, that that wouldn't fly. Or, you know, you're coming home from from a bar late at night and you have to wait 10 minutes, you know, 5 to 10 minutes
92
00:07:54.770 -->
00:08:01.270to get a recommended driver. Right? Like, real time data is customer experience today. Like, it is a prerequisite
93
00:08:02.045 -->
00:08:08.705for that. And so you're right in that, you know, you look at these these these companies, Netflix got 2, 000 data engineers.
94
00:08:09.300 -->
00:08:16.440Right? Like like, I mean, that's an that's an arm they have more data engineers than some people have regular engineers. Right? Like, that's insane.
95
00:08:17.145 -->
00:08:27.060So who who's gonna compete with that? Yeah. I I and that's really what why, you know, I'm not trying to turn this into a more actual commercial, but that's really why we exist is that that that infrastructure
96
00:08:27.920 -->
00:08:58.460should be democratized so that people can start bringing that business value. And then the innovation and, you know, like like, the the guardrails that we provide to for people to innovate at that level, that's something that is gonna usher in a whole new set of experiences and apps that weren't possible to people because they just simply didn't have the money and resources and expertise to to make those things happen. And, like, that's the the the benefit that that that we have. Right? It's like, you can go from, you know, single person startup, you know, or somebody that's tinkering on the idea
97
00:08:58.920 -->
00:09:21.029all the way up to, you know, old world legacy, ecommerce, banking, defense, you know, like, all that type of stuff. I think that that's really the, the mindset that people have to start thinking. And what you know, I'm a big proponent of is that real time data should just be the default. Alright? Real time data doesn't necessarily like having data flowing in real time. Doesn't mean event driven. Alright? But it just means that as,
98
00:09:21.495 -->
00:09:45.485you know, actions are happening across your platform, your app, whatever it is, that that that activity is getting tracked somewhere in real time. And it's very granular. It's very actionable, and and you can do things with it if you want to in the moment that you need to. And that's really the you know, if if if we can get all of these companies at a at a base level to start doing that, I think I think you're gonna start to see
99
00:09:46.020 -->
00:09:49.720better apps and and and and better customer experiences.
100
00:09:50.260 -->
00:10:27.600I mean, like, that that's really the goal. Right? It's it's it's just to to to do that. And, you know, for me, it's like there's a ton of these what I call intelligent insights and and analytic companies where it's like, yo, you usually gotta get the data to us and then, like, we can give you all of this. But the hard part is actually, like, getting the data structured in a format that they're, you know, kind of black box and go do the thing and provide the value at that point. Right? And it's like, we handle that first part, but also we give you the guardrails to make that second part easier, scalable, and, you know, that that type of thing. So that's that's really the the the value prop and sell that we have when we go talk to an organization.
101
00:10:28.275 -->
00:10:31.175 And beyond the baseline infrastructure
102
00:10:31.555 -->
00:10:39.500of I have the pipes, I'm able to get data in, and it's able to get processed and spit back out the other end in real time. What are some of the other
103
00:10:39.880 -->
00:10:50.675barriers or aspects of the kind of complexity of working with real time data that you see teams run into once they say, okay. I've got my rocks in it. It solves all my plumbing problems. Now what? Yeah. Exactly.
104
00:10:51.154 -->
00:10:55.270 It's a different access pattern. Right? Like, people are used to writing
105
00:10:55.650 -->
00:11:00.310a select star from orders, you know, kinda query. Right? They aren't really, like,
106
00:11:00.775 -->
00:11:03.835select star from, you know, you know,
107
00:11:05.095 -->
00:11:11.870cart update or something like that. Right? Like, just the granularity of the data. Like, people are just just used to dealing with snapshots,
108
00:11:12.250 -->
00:11:15.709right, in the CSV world. Right? And it's like
109
00:11:16.170 -->
00:11:24.125I I I think part of the problem is is that that folks really aren't oriented for that type of world where it's super granular
110
00:11:24.580 -->
00:11:31.480and it's micro versus, like, we get in a CSV dump full of just data upon data upon data upon data. Right? And I think, like,
111
00:11:31.855 -->
00:11:38.355framing your mindset around the different types of data that can come in, whether and and have it be consistent.
112
00:11:38.790 -->
00:12:04.885I think that's the thing. The others the the other challenge around working with real time data is, like, order guarantees and delivery semantics and, like, you know, like, all those types of things. Right? And so we handle a lot of that, but, I mean, you know, deduping is a problem for everybody. Right? Like, you know, scaling is a problem for for most people, not us, but, you know, a lot of people. Right? And so just understanding how to deal with that those types of of patterns and architectural decisions and things, That's a,
113
00:12:05.385 -->
00:12:20.279you know, it's just a new new world that we're trying to usher in. Right? Like, everybody's talking about data apps and, you know, data management, all this stuff, but they talk about it from, like, the the outskirts of it. Right? Like, you know, this data lineage. This data governance. This data observability.
114
00:12:20.660 -->
00:12:25.195This metadata management. It's like, all of that doesn't necessarily
115
00:12:25.655 -->
00:12:26.955talk about the actual
116
00:12:27.495 -->
00:12:29.755experience of how you're using that data,
117
00:12:30.215 -->
00:12:51.390to to to provide customer experiences. Alright? Like, all that stuff is just just pure overhead. And, yeah, it can it can add things, but it's like, I just need to go from point a to point b. And you're just like, you know, let me throw a spoiler on the car. And it's like, I don't really need that right now. Like, I just need to go from here to here. Right? And so I feel like there needs to be a a a simplification
118
00:12:51.850 -->
00:12:57.385of the the data landscape so people don't see this as as as too much of a herculean task.
119
00:12:57.705 -->
00:13:01.325 Now that real time data is at the point of democratization,
120
00:13:01.785 -->
00:13:03.805more people are starting to
121
00:13:04.185 -->
00:13:10.910embark on their own journeys of actually incorporating that into their product offerings or building whole applications around this paradigm?
122
00:13:11.210 -->
00:13:26.030What are some of the new categories of the types of products and applications and experiences that you have seen being unlocked by this broader availability of the underlying technology? Yeah. I mean, I I think, 1 of the things that this unlocks is that it can
123
00:13:26.730 -->
00:13:41.060 move decision making and processing closer to the to the edge. Right? And because now I can if I can get a granular kind of bite of of activity coming in, I can action you know, do some action on that versus,
124
00:13:41.360 -->
00:14:19.410alright. Well, I've got I, you know, I gotta wait every hour to get a dump of the last hour of activity, then I gotta run some analytics, and then I get some to some some tables and blah blah blah. And maybe 3, 4 hours down the line. Assume that I have everything already set up and automated. Right? Like which most people don't, but, like, you know, 3 to 4 hours down the line, then I can actually, you know, give that result or provide some value back. And it's like, yeah. No. I can actually push that back a little bit further. Right? Or I don't have to wait. I don't have to choose either or. That's the other part too. Right? Where it's like, oh, am I doing this for analysis or I'm actually being Right? And, like, the analysis part, you're looking backwards, but, you know, being proactive, like, you know, with the real time data, you're actually, you know, doing the stuff in time. So I think
125
00:14:22.645 -->
00:15:13.820pushing the decision making and processing closer to the edge is really, like, 1 of the biggest benefits that that that we've seen that's kind of unlocked the next generation of, you know, kind of customer experiences and value. And that's that's really where people wanna be at. And it's it's not, yo, let's throw more infrastructure or more point products at the at the problem is, no. We already have a great foundation. Let's build upon that and start thinking about more about how we could be more applicable and relevant to our customers. Because, I mean, that's what it is. It's the name of the game. Right? It's like, you know, whether your customer's external or internal, you're trying to use the data to tell a story or provide value to whoever whoever needs it. Right? And so, you know, that that that's really the the the the benefit of of of using us and question the decision making and the ability to to to do all those different types of things closer to the to the edge and then not having to sacrifice, like, historical
126
00:15:14.464 -->
00:15:16.165analysis and then proactive,
127
00:15:16.704 -->
00:15:18.084you know, customer engagement.
128
00:15:18.464 -->
00:15:23.285 And with that experience of pushing more of that decision making to the edge, pushing the experience
129
00:15:23.610 -->
00:15:25.470to being much lower latency
130
00:15:25.850 -->
00:15:39.485that also brings the possibility of increasing the risk profile because you have a much shorter window to be able to react to any bad data, any, you know, bugs that get introduced, any errors in kind of the source systems. And I'm wondering
131
00:15:40.105 -->
00:15:40.925how that impacts
132
00:15:41.320 -->
00:16:03.360the kind of the appetite for risk and the overall risk profile of the applications that are being built and how that also influences the types of applications that companies are comfortable building as they first start to explore the space. I mean, that's not gonna change regardless if it's a data app or it's just a regular web app. Right? Like, you know, that that that risk profile is gonna be there. I think, actually,
133
00:16:03.740 -->
00:16:04.400 for us,
134
00:16:05.115 -->
00:16:21.600because we are just playing regular code and we're end to end, you know, you can just write a unit test or a functional test. You don't have to set up a whole bunch of infrastructure to, like, do regular testing. Right? Like, it can be done locally on your machine before you actually deploy it and all that type of stuff. Right? Like, if you are
135
00:16:22.425 -->
00:17:08.960have good testing practices in general for your web app and, like, all that type of stuff, like, porting that over to Morax is super easy because it's just regular code, man. You know, you can use your Datadog to to to see the output of the functions and, you know, monitor, do all those, Datadog or Splunk or whatever it is. Right? Like, it it just literally just fits into your existing workflows. So it hasn't changed. I wouldn't say, like, yeah, it hasn't changed by and large the types of apps that people are are building. I think because we're we're giving them the guardrails to experiment, they're more willing to to do some of the riskier things because, you know, they can have more repeated chances at bat. They don't have to wait 2 to 3 weeks for the data team to give give an update or blah blah blah. Like, it fits within their existing software development life cycle. Right? And, like, I think that confidence
136
00:17:09.260 -->
00:17:10.160and that foundation
137
00:17:10.700 -->
00:17:16.560literally gives them the ability to be more risky and to do things that that they didn't think were were possible with their existing
138
00:17:17.075 -->
00:17:21.255 kinda architecture and stacks. And for any type of
139
00:17:21.715 -->
00:17:23.735more kind of potentially sensitive
140
00:17:24.230 -->
00:17:26.010data applications or organizations
141
00:17:26.470 -->
00:17:27.130that are
142
00:17:27.590 -->
00:17:29.290in a more kind of rigorous
143
00:17:29.670 -->
00:17:30.890regulatory environment,
144
00:17:31.270 -->
00:17:33.465what are some of the types of technical
145
00:17:34.165 -->
00:17:45.850controls that they should be thinking about either in terms of validating the source data as it's coming into the pipeline or, you know, as it's traversing the pipeline before it gets delivered, just kind of what what are the available points
146
00:17:46.630 -->
00:17:55.885of mitigation or overall strategies for mitigating some of those risk profiles? Yeah. I mean, you know, we run, you know, on premise, off premise, edge,
147
00:17:56.424 -->
00:18:14.404 hybrid, like, every combination that you could think of and especially for some of these highly regulated environments. Right? Like, 1, you know, we're encrypted end to end. Alright? So, you know, as soon as soon as we ingest something, we basically do, like, PKI at scale, right, which is gets embedded in the key,
148
00:18:14.945 -->
00:18:16.164you know, on the ingest.
149
00:18:16.520 -->
00:18:19.980And to access that data downstream, you need to basically have the key,
150
00:18:20.600 -->
00:18:34.710in order to do that. So and then it's encrypted in the end. And that's whether we're regardless of where we're deployed and how we're deployed. Right? And so the other piece of that too is is that, you know, right now, because a lot of the work that we do is in the department of defense is multiclass.
151
00:18:35.010 -->
00:18:41.909It's, you know, kinda all over the place as far as, like, connectivity and things like that. So we've really had to build in a lot of order guaranteeing,
152
00:18:42.370 -->
00:18:48.245a lot of, like, you know, resiliency into the platform, a lot of, you know, all of that so that way we can
153
00:18:48.545 -->
00:18:50.325make sure that your data gets delivered.
154
00:18:50.705 -->
00:19:11.270Right? Like like like that. And and, know, just as a general ethos as a company. Right? You know, we got 2 jobs. Never lose data. Never expose data. So I think, you know, 1 of the things, like, as we were going through complaint and I'm not, like, saying all this stuff to, like, hackers come come and mess with us. Right? Like, you know, it's not like a open challenge because I'm sure that there's something we miss. But, you know,
155
00:19:12.870 -->
00:19:25.220we get regular pen tests. We you know, all all the compliance stuff that we have to deal with. Like, we're we go overboard just to make sure that we never lose data and never expose data. I mean, so much of the so to the point where
156
00:19:25.760 -->
00:19:27.140even when we're troubleshooting,
157
00:19:27.520 -->
00:19:59.600we don't actually get to see the details of the record. We just see that, like, hey. This source system sent x y z over over and it's, you know, these fields, these types. We don't actually see the values. Right? Even, that's kinda 1 of the things. And, like, we give our users the the ability to to to tune their their caching or retention because, you know, underneath we use Kafka. Right? And so, you know, you have to even though it's processed and all of that, there's a log record that goes through that. Right? And so, like, we're just making sure that that we are, you know,
158
00:19:59.980 -->
00:20:01.759we are doing right by our customers
159
00:20:02.220 -->
00:20:05.519to make sure that, you know, those risks are mitigated as much as possible.
160
00:20:05.980 -->
00:20:07.200Because, I mean, you know,
161
00:20:07.755 -->
00:20:21.030data comes from everywhere. And, you know, we do a lot of work to make sure that when you connect to something, it's gonna always say connected. And then when you're, you know, moving that data and orchestrating it, it's always gonna be in the 4th bet that's unique.
162
00:20:21.490 -->
00:20:24.309 And for developers who are
163
00:20:25.164 -->
00:20:30.225building these real time applications, I'm wondering what are some of the technical capabilities,
164
00:20:30.765 -->
00:20:31.585the architectural
165
00:20:31.965 -->
00:20:32.310design
166
00:20:32.950 -->
00:20:36.650kind of background that they need, some of the ways that this real time data
167
00:20:36.950 -->
00:20:39.770introduces new architectural paradigms or new
168
00:20:40.195 -->
00:20:48.294strains at the different integration points between systems and just some of the overall application and system design process that needs to be brought into
169
00:20:48.755 -->
00:20:56.090 the the process of being able to actually build these applications? Yeah. I mean, I I would say the the biggest thing is we,
170
00:20:57.634 -->
00:20:59.414all we care about is,
171
00:21:00.835 -->
00:21:16.285or, you know, the the skill set. I would say it could be like a junior engineer because all you really need to know is where your data's coming from, where it's going, and what format it needs to be when it gets there. Like, that's literally the access pattern for how you do that. It's like, you know, dot connect, dot process,
172
00:21:16.825 -->
00:21:18.605and dot write. You're like,
173
00:21:19.065 -->
00:21:21.325done. You know what I'm saying? And it's like,
174
00:21:21.700 -->
00:21:31.240oh, that that that that makes a lot of sense. And and I think that, like like, we tried to map that inside of our SDK. So if you if you know how to declare a a method
175
00:21:31.755 -->
00:21:35.375off of a, you know, object instant, you know, like, that type of stuff,
176
00:21:35.835 -->
00:21:40.400you too now can have be a data engineer. Right? Like, that's pretty much it.
177
00:21:41.020 -->
00:21:47.920So that's really the thing. And, like, just learning the architecture underneath, like, that's our whole value prop. Like, you don't really need to know the architecture underneath,
178
00:21:49.144 -->
00:21:51.804Because why? Right? Like, if if I'm in, you know,
179
00:21:52.345 -->
00:21:56.205let's just say I'm in the New York Times. Right? Like, am I a distributed data company?
180
00:21:56.640 -->
00:22:08.795I mean, not really. I'm a news company. So all I really want is, like like, I wanna be able to get people to create content faster and, like, those sorts of things. I don't care about setting up a, you know, ingest service and
181
00:22:09.095 -->
00:22:24.595airflow orchestration and blah blah blah. Like, that's a means to an end, and it shouldn't be the thing that you're you're spending the most time and resources on. And, like, that's really the the value. So, yeah, I mean, junior engineer can as long as they know how to declare methods, they can use Maroxa,
182
00:22:25.294 -->
00:22:28.034 and nobody should really have to worry about the underlying infrastructure.
183
00:22:28.414 -->
00:22:41.615And as a service provider and a platform operator focused on this real time space, I know, as you've mentioned, that you're built on top of Kafka as your kind of core backbone of the system. I'm wondering
184
00:22:41.995 -->
00:22:45.455over the past couple of years as you have matured the platform,
185
00:22:45.914 -->
00:22:47.294kind of adapted the capabilities,
186
00:22:47.760 -->
00:22:48.160added,
187
00:22:48.800 -->
00:22:51.220this interface on top to be able to simplify,
188
00:22:51.680 -->
00:22:58.555or in kind of abstract way all of the internals for people. What are some of the overall kind of system design evolutions that have gone
189
00:22:59.255 -->
00:23:00.315internally into Meroxa
190
00:23:00.695 -->
00:23:02.235and your perspective
191
00:23:02.615 -->
00:23:05.200on the level of kind of maturity and capability
192
00:23:05.500 -->
00:23:07.360of the kind of streaming technologies
193
00:23:08.059 -->
00:23:08.720more broadly.
194
00:23:09.259 -->
00:23:19.880 Yeah. I mean, our architecture has changed quite a bit. Right? Like, we've these are changes that we knew to expect. Right? But 1 of the biggest things that we ended up doing was that we rewrote,
195
00:23:20.500 -->
00:23:26.679our our Kafka Connect, and and we rewrote that in Go into a open source project called Conduit.
196
00:23:27.075 -->
00:23:27.475And,
197
00:23:27.875 -->
00:23:33.414you know, there there there there there were platform reasons and, you know, just kinda functionality reasons,
198
00:23:33.715 -->
00:23:43.830why we ended up doing that. So 1, Kafka Connect runs on the the the JVM, and that's just a huge resource hog. 2, the connectors are just kinda all over the place as far as quality
199
00:23:44.130 -->
00:23:45.350and and and efficiency.
200
00:23:45.955 -->
00:24:08.150We found out that, like, oh, snap. If you use the red chip connector, that's like a a gigabyte of RAM. Now, you know, magnify that across, multiply that across all of our customer base. Right? And it's like, you get all of these provisioned resources and not a lot of consistent, you know, throughput. And so I was just like, why are we why are we doing that? Right? Like, why do we have all these beefy boxes just for the Kafka Connect connector,
201
00:24:08.530 -->
00:24:12.290and it's just it's just not super performant. Other thing too is is that,
202
00:24:12.770 -->
00:24:16.790you know, we wanted people to be able to write connectors in their favorite languages.
203
00:24:17.575 -->
00:24:33.400And so, you know, to start building this this ecosystem like Kafka, and I get why they had to, I mean, Confluent, why they had to to, you know, kinda close down the ecosystem because of business reasons. But, look, at the end of the day, right, like, if I wanna build a connector, like, I shouldn't have to go through some weird esoteric
204
00:24:33.745 -->
00:24:36.885kinda, you know, license approval process. Or if I wanna
205
00:24:37.185 -->
00:24:43.680update a connector based off of my use case, I shouldn't have to, you know, be, subjected to to, you know, my vendors,
206
00:24:44.540 -->
00:24:58.775you know, product road map like thing. Like, it just didn't make sense. So we built our own for for that reason. And then just on the operational side, like, we started to use that on our for ourselves. So, you know, every all of our service run services run on Kubernetes,
207
00:24:59.590 -->
00:25:03.530all containerized, all that type of stuff. We have custom Kubernetes controllers
208
00:25:03.830 -->
00:25:16.900that that, you know, help talk to our control plane to to, you know, do all the automation and things like that. But also 1 of the interesting things that we found out is, like, on Kafka Connect just randomly creates new topics
209
00:25:24.340 -->
00:25:51.320get for some of the workflows that we needed to do, right, and lights, you know, be able to manage streams and some of the work work stuff that we have down the line with, like, stateful stream processing, we need to have a little bit more control over the life cycle of an event. Alright? And, like, Kafka Connect just didn't give that to you. And so, you know, it just turns out, like, oh, yeah. This this actually works. This is better for us. And now we can scale this better. We can, you know, customize it better. And it's just a better customer experience,
210
00:25:51.700 -->
00:25:55.960down the line for for data integration. And that thing can sit alone,
211
00:25:56.340 -->
00:26:08.465you know, be be be a standalone thing, or you're gonna put it inside of this giant infrastructure and, like, do amazing things. So, you know, we've got it, you know, copy the net, reza, go sing you know, standalone go binary,
212
00:26:08.810 -->
00:26:12.910single binary. But, like, we can it's, like, 30 megabytes. We can deploy that anywhere.
213
00:26:13.370 -->
00:26:18.270Right? Do do all types of things. Right? And so, I mean, like, it's super cool.
214
00:26:18.875 -->
00:26:26.095And that's probably the biggest architectural change that we've had. I mean, pretty soon, I think, you know, either mid this year towards end of towards
215
00:26:26.510 -->
00:26:44.485end of this year, we're gonna 1 0 it. You know? It's a bunch of connectors that we got for it. It actually we included a run time so you can run your existing Confid Connect connectors. But, I mean, you know, for us, it's that's really been our biggest competitive advantage. Because the other part of it is, like, we can generate high super high quality connectors
216
00:26:44.865 -->
00:26:55.825very, very, very fast. Right? And, like, it's already Kafka compatible and, you know, like, all that type of stuff. So, yeah, I mean, it's, it's pretty cool. And if you were to start over today,
217
00:26:56.205 -->
00:27:00.225 brand new with the infrastructure with the ecosystem as it is now, wondering
218
00:27:00.560 -->
00:27:22.780 what are some of the other architectural primitives that you would orient around or ways that you would rethink the overall implementation of your stack. Yeah. I mean, I think I think 1 of the things that we were looking at is, like, doing the stateful stream processing. You can probably doing that a little sooner. Alright? So right now, we're stateless. Right? And it's just like, you know, you can pass this stuff through, but, like, a lot of the use cases around
219
00:27:23.400 -->
00:27:29.125analytics, which is where the you know, a lot of people find the value in, need stateful student
220
00:27:29.505 -->
00:27:30.245processing. So,
221
00:27:30.545 -->
00:27:31.105you know,
222
00:27:31.745 -->
00:27:48.385building in Beam and Flink into the into the platform. Right? Like, that's something that that we're doing now, which also to your last question, like, how's the architecture changed? But, like, yeah, we're introducing, like, you know, Flink as a as a, you know, our stateful stream process of runtime. Right? And, like,
223
00:27:48.765 -->
00:27:53.025that's the you know, because of the the customer demand and use cases,
224
00:27:53.470 -->
00:27:58.850all this stuff that have evolved. You know, now, Maroxa or will be, you know, once we release out to the public,
225
00:27:59.150 -->
00:28:01.970again, you won't have to worry about the infrastructure. You can just do dotwindow.aggregate.join.
226
00:28:03.534 -->
00:28:04.034Right?
227
00:28:04.495 -->
00:28:25.505Like, it's pretty pretty cool to do. Right? Like, you know, instead of like, oh, man. I gotta get this Flink job. I gotta do this. I gotta manage the schema. I gotta, like, all this things. I was like, no. No. No. No. No. This is just this. Alright? Need to make sure I'm checkpointing it properly. Exactly. Like, you ain't gotta worry about that. Like, we'll handle all of that for you. Right? And I think that that's something that is is super useful.
228
00:28:26.365 -->
00:28:31.240Other thing I forgot to mention, because, like, we do so much cool stuff underneath the hood, It's like, we'll have a,
229
00:28:32.100 -->
00:28:34.580a dot analytics thing where it's basically, like,
230
00:28:35.380 -->
00:28:36.120our own,
231
00:28:36.900 -->
00:28:38.040real time store.
232
00:28:38.345 -->
00:28:51.530So for people that, you know, probably are used to using, like, ClickHouse or or Druid or Pinot or something like that. Right? Like, again, you don't have to worry about it. Like, we'll pre compile the schemas for you. We'll do all of that and,
233
00:28:51.830 -->
00:29:13.585make it easy for people to to to do these queries in real time as well. So I think, like, you know, and just put it behind an endpoint. Right? Let's just go to, you know, whatever your data app name is and slash query and, like, you'll be able to, you know, write SQL and things like that. Right? And so, you know, for us, it's just really making that experience easier end to end so that people can, you know, interact with the data,
234
00:29:14.125 -->
00:29:23.650as it is in in real time. And, like, you know, again, it's it's for me, I always say it's kinda like the the the the phone system. Right? Like, you know, we've had, 7 digit, 10 digit phone numbers forever. The technology underneath it has changed, but the the user interface part of that
235
00:29:29.284 -->
00:29:32.904is always the same. I just dial a number, and I can pick up on the other end. Alright?
236
00:29:33.205 -->
00:29:37.890And that's really kinda what we're doing underneath the hood. It's just like, oh, we'll give you additional functionality,
237
00:29:38.190 -->
00:29:47.935but you don't necessarily need to know how the sausage is made. Alright? We'll give you the ability to tune it and, you know, kinda choose your own flavors. But at the end of the day, right, like, for 90% of the world,
238
00:29:48.475 -->
00:29:59.049 you know, the base offering is gonna be okay. Absolutely. And you mentioned kind of statefulness and kind of being able to do windowing functions is something that you're investing in now.
239
00:29:59.485 -->
00:30:00.385Another aspect
240
00:30:01.005 -->
00:30:21.304of streaming data that people will typically run into is wanting to be able to actually join data across streams as they're traversing the different pipes, as well as being able to do transactional workflows where I only want to actually commit this record to the stream if this other, you know, record is also committed at the same time. So being able to do kind of atomic operations
241
00:30:21.765 -->
00:30:32.730 on top of the streams, and I'm wondering what your level of investment has been on those types of capabilities as well. So for a stateful stream processing, like, you just named a whole bunch of stuff that were abstract in a way. Right? Like
242
00:30:33.110 -->
00:30:40.644like like, at the end of the day, right, like like, the should be able to help you do a lot of that stuff. But, again, when you wanna join the stream, what what do you wanna do? Like, just just talk out that algorithm.
243
00:30:41.105 -->
00:30:42.085Right? Just like,
244
00:30:45.260 -->
00:30:55.155oh, I got data source a, data source b. I wanna join on a specific field or a specific key. Right? Like, oh, that sounds like a method I could do. Give
245
00:30:55.775 -->
00:30:56.595me give me
246
00:30:57.054 -->
00:31:10.385whatever it is, data source dot, you know, or or whatever it is, like, app dot join. And I you know, the first argument is stream 1, stream 2, and then the 3rd you know, second argument, stream 2. The third argument is the key. Right? Like
247
00:31:11.085 -->
00:31:13.585and then I could just store that in a variable or a collection.
248
00:31:14.285 -->
00:31:25.370You know what I'm saying? Like like like, which 1 would you rather have? Right? And, like, that's the experience that that that we're bringing out to to to the world. And then over time as as more paradigms get introduced to
249
00:31:25.865 -->
00:31:33.005underlying platforms, we'll add more functionality as we see fit. But, like, for most people, all you need is dot join or dot aggregate
250
00:31:33.305 -->
00:31:46.795or, you know, something like that. Like, the maintenance and operation of that, just leave that to us so we can handle it. But, like yeah. I mean, that that that's really the the the user interface that we're thinking about, bringing out via code. You know?
251
00:31:47.175 -->
00:31:54.169 And in terms of the kind of API design, the interface, the ways that you kind of document and communicate about the
252
00:31:54.549 -->
00:32:27.195problems that you're solving and kind of trying to avoid having to get too deep into the weeds about the ways that those problems are being solved, I'm curious what your overall philosophy has been, your approach for actually building out those APIs and validating them and ensuring that developers are able to intuit what they are actually doing without maybe saying, like, oh, I thought it was doing this thing, but it actually did this completely different random thing that I didn't want. Want. Yeah. Yeah. Yeah. Yeah. We don't want any, like, weirdo side effects from from that stuff. But we start I mean, this is just us being being product focused. Right? It's like
253
00:32:27.559 -->
00:32:29.340 we don't fall in love with the technology.
254
00:32:29.799 -->
00:32:37.664We start with the experience and the user journey first and then work backwards, right, to and figure out, like, okay. Well, what pieces of technology
255
00:32:38.125 -->
00:32:39.745or or combinations of technology
256
00:32:40.205 -->
00:32:44.945can we come you know, leverage to to to drive the ideal user experience?
257
00:32:45.250 -->
00:32:47.750So when I say, like, dot join or dot aggregate,
258
00:32:48.130 -->
00:32:49.830right, like, that's where we start.
259
00:32:50.210 -->
00:32:58.735Right? And it's like, okay. Well, when we do that, what is what you know, we do customer interviews and things like that, and we just start thinking about, like, okay. Well, what are
260
00:32:59.115 -->
00:33:02.095the the expectations of this output? Or what are the expectations
261
00:33:02.395 -->
00:33:35.745of this, you know, kind of the high level API. And then once we get get that in, and then that's when we kinda, you know, go down the the go through the paces of, okay. Well, this is, you know, how we how we how we surface that view. And I think that that's really the the approach that, you know, if you're focused on the product and the experience, it's it's easier because people aren't necessarily, like, you know, you build the thing first, then you go back and refactor and make it more performant, more secure, blah blah blah blah. Right? Like, that's the thing that, you know, with everything that we do, we just try to be mindful of that.
262
00:33:36.285 -->
00:33:47.395 And in terms of the ways that people are employing maroxa, the types of applications that they're building, what are some of the most interesting or innovative or unexpected ways that you've seen it used? Oh, man.
263
00:33:48.174 -->
00:33:52.495 A lot of stuff I can't talk about because we do a lot of lot of stuff in defense, but,
264
00:33:53.135 -->
00:33:55.154I mean, look, man. We've we've
265
00:33:56.030 -->
00:34:00.210I wouldn't say a lot of it is super interesting because but it's it's
266
00:34:00.510 -->
00:34:03.330interesting because of, like, from a use case perspective,
267
00:34:03.915 -->
00:34:20.165there's a lot of data migration. There's a lot of, like, transformations and, you know, going from legacy to to to kinda modern things. But I would say it's interesting because of the people who use it. Right? And, like, the things that they're trying to solve is interesting to them. That's all I care about. Right? Long as it's interesting to you,
268
00:34:20.865 -->
00:34:30.860you know, I real time search indexing, is that gonna get people like, oh, man. I I I man, my my inventory on my ecommerce website is always up to date.
269
00:34:31.560 -->
00:34:32.780On people's Like,
270
00:34:33.640 -->
00:34:45.850you know what I'm saying? Like like, that type of thing. But 1 of the fun things I was you know, I had we had a a a some people, you know, kinda building apps. Right? And, like, no. You just start to see, like, people are, you know, doing interesting
271
00:34:47.210 -->
00:34:54.490stock watching and, you know, building models there in the financial services. Right? And, like, the ability for people to to to kinda,
272
00:34:56.075 -->
00:34:58.655do active learning instead of, like, using,
273
00:34:58.955 -->
00:34:59.355you know,
274
00:34:59.994 -->
00:35:01.275you know, some of these, like,
275
00:35:01.915 -->
00:35:09.570GPT things or generative AI things. Right? Like, some of them having to use billions of data points, they can, like, build their own,
276
00:35:10.830 -->
00:35:13.490you know, kinda data driven or data focused,
277
00:35:13.950 -->
00:35:14.770data centric
278
00:35:15.205 -->
00:35:20.025models and stuff like that. Right? And, like, the applicability on that and for, like, you know,
279
00:35:20.405 -->
00:35:21.225risk mitigation
280
00:35:21.765 -->
00:35:22.985and and and
281
00:35:23.490 -->
00:35:40.255simulations and all that type of stuff. It's actually kinda cool to see, but it's like, you know, that's in the Fintech world. That's what they do. Right? Like, not really interesting to me, but long as you like it and you pay me every month to be able to do that, hey. I love it, man. So that's just kinda 1 of the things.
282
00:35:40.770 -->
00:35:50.585 And in your experience of building MiraXa, building out this product, exploring this problem space, what are some of the most interesting or unexpected or challenging lessons that you've learned personally?
283
00:35:50.965 -->
00:35:51.625 Oh, man.
284
00:35:52.485 -->
00:35:58.050You know, it's it's 1 of those things where where we're inter inter you know, we stand on the shoulders of giants.
285
00:35:58.850 -->
00:36:00.390Right? And what I mean by that is
286
00:36:00.850 -->
00:36:03.350we're system engineer first, software engineer second,
287
00:36:04.050 -->
00:36:07.430which means that, you know, we'll we'll curate existing components
288
00:36:07.805 -->
00:36:13.105and try to put them together. And if it doesn't necessarily work, then we'll build something. And so a lot of these, like,
289
00:36:13.485 -->
00:36:19.130very, very, very, very popular products in open source world, they're just really bad developer experience.
290
00:36:19.430 -->
00:36:21.050Right? Like like, this
291
00:36:21.590 -->
00:36:24.250you know? Or or or, like, people don't
292
00:36:24.865 -->
00:36:26.245necessarily think about
293
00:36:26.545 -->
00:36:27.045scaling
294
00:36:27.585 -->
00:36:30.085and, you know, kind of those things. Right? And it's like,
295
00:36:30.465 -->
00:36:31.275we've had to,
296
00:36:33.000 -->
00:36:51.520you know, make some changes and upstream changes to to to repos and things like that because of, you know, people not understanding how these things can be used in the use cases. Yeah. Which is, again, this is a generic platform. It's an open source thing. Like, people have their their windows and lenses and into all these things. And so it's, it's just like,
297
00:36:51.820 -->
00:37:26.000you know, sometimes we'll we'll get into things. It's like, yo, there has to be somebody that's complaining about this stuff because this legit does not make sense. Right? Like like, those types of things. So we find ourselves running into that, like, quite a bit. And and, you know, so, you know, for better or worse, like, you know, you'll read a a project and be like, oh, this sounds great. I can use it. And then once you start digging into the nitty gritty, it's like, oh, man. Like, yeah, this is there's some some weirdo kinda side effects or, you know, negative negative interactions that you might have when you start integrating those things. So, yeah, I I think I think from a technical aspect, that's 1 of the things that that we've learned.
298
00:37:26.300 -->
00:37:31.385I mean, you know, just just in general. Right? Like, talking to people and understanding
299
00:37:31.765 -->
00:37:51.595how how important data is to them on a day to day basis and how shitty their experiences have been. Like, that's super surprising. Like, you're okay waiting 3 months to get a result back for, like, mission critical stuff. Other than the super surprises, like, the of legacy stuff that's out there, I'm like, yo. Whoever sold IBM DB 2,
300
00:37:51.895 -->
00:37:56.635like, that person is probably in, like, the salesperson all hall of fame at some point. Like,
301
00:37:56.940 -->
00:38:00.960there's so much IBM DB 2 out there. And I was just like, yo.
302
00:38:01.500 -->
00:38:09.865Why is this a thing? How was this a thing? Like, you know, it's just like, wow, man. This is, you know, the eighties called and they want their database back. So,
303
00:38:10.645 -->
00:38:45.825you know, it's just it's just some of that stuff where it's just like it just surprises you. Like, y'all y'all still are using it. Like, all I hear in my day to day is, like, Snowflake is everywhere. Databridge is everywhere. Then when you go start talking to these enterprises, they're like, we literally just got on Hadoop. So now you're telling me there's this new thing that I gotta worry about? I hope, you know, pump your brakes kinda thing. So I think I think that's, that's 1 of the the things that that is the most surprising for us. Some of the things that are the most surprising for us. Absolutely. Especially on the legacy point. As somebody who, in my career, has been in in the process of actually installing a brand new AS 400.
304
00:38:47.040 -->
00:38:48.900That's a thing? That's a thing.
305
00:38:49.440 -->
00:38:54.395Wait. Wait. Wait. You said brand new install of it. Like, on what machine?
306
00:38:55.015 -->
00:39:02.440 Right? Like No. An act a s 400. Like, that is the machine. They they carted it on a pallet, and they say, here you go. Like, they still manufacture
307
00:39:03.300 -->
00:39:07.800those? Or was it I mean, as of about 10, 15 years ago. Yeah. Wild.
308
00:39:08.260 -->
00:39:11.525 That's why I'm like, you said brand new at AS 400. They're like, that is
309
00:39:12.085 -->
00:39:18.505 that's tripping me out right now, man. Okay. Yep. Yeah. Yeah. I had somebody on my team who was the dedicated RPG developer.
310
00:39:19.680 -->
00:39:21.859 I will say this, I've ridden more COBOL
311
00:39:22.160 -->
00:39:27.655and and and Fortran in the, you know, probably the last, like, year than I've ever had to.
312
00:39:29.655 -->
00:39:30.155 Yep.
313
00:39:30.695 -->
00:39:41.190These things never die. They just go dormant. Ever. Die. And so for people who are exploring the space of real time, what are the cases where maroxa is the wrong choice?
314
00:39:41.490 -->
00:39:43.250 There is no. No. No.
315
00:39:43.890 -->
00:39:51.735I would say, like, until we get the, like, the stream process and stuff, like, fully baked, you know, there there there's some, like, you know, weird ergonomics
316
00:39:52.035 -->
00:39:57.175around, like, that type of stuff, and you just you know, you wouldn't use it for stateful stream processing,
317
00:39:57.730 -->
00:40:02.950where you need to do joins and aggregations and, you know, kind of window things and blah blah blah. Like,
318
00:40:03.250 -->
00:40:07.695that's not something that that that's the best experience for us right now.
319
00:40:08.315 -->
00:40:10.815Everything else, fair game. So, you know,
320
00:40:11.115 -->
00:40:14.575there's a lot of cool things that people are doing and and and,
321
00:40:15.560 -->
00:40:19.740you know, find a utility on our platform, right, in a in a myriad of ways. So
322
00:40:20.200 -->
00:40:33.734 And as you continue to build and iterate on the product, what are some of the things you have planned for the near to medium term of maroxa or any particular problem areas that you're excited to dig into? Yeah. Definitely. So, obviously, you keep hearing me say stateful state processing,
323
00:40:34.580 -->
00:40:38.840 adding support for for, you know, some additional languages like,
324
00:40:39.460 -->
00:40:42.885c sharp and Kotlin and maybe even Rust. You know?
325
00:40:43.685 -->
00:40:45.705You know, Atom's support for more connectors,
326
00:40:46.485 -->
00:40:47.545you know, interoperability
327
00:40:47.845 -->
00:40:54.190with or streaming systems like, you know, Pulsar and Red Panda and, you know, some of those things. Right?
328
00:40:55.130 -->
00:40:59.230I also think that, we're getting into the generative space.
329
00:40:59.905 -->
00:41:03.205So much like everybody else in their mama right now is
330
00:41:03.665 -->
00:41:05.205is, you know, has some, like,
331
00:41:05.585 -->
00:41:07.445dope kinda coding experiences
332
00:41:07.905 -->
00:41:12.200driven by generative AI. We will play in that space on multiple levels as well,
333
00:41:12.900 -->
00:41:24.154including, like, in our dashboard when you'll be able to ask questions about the data, you know, as it moves in real time, build the apps in real time. And I think, like, you know, we'll have a a pretty cool,
334
00:41:24.934 -->
00:41:30.400experience there. And then, you know, eventually, like, start, you know, moving away from just the developer
335
00:41:30.940 -->
00:41:36.645persona and moving into to different personas. And and, you know, do we build our own REPL
336
00:41:37.185 -->
00:41:39.345kinda workbook based flow for,
337
00:41:39.744 -->
00:42:00.615for data scientists and analysts? You know, what does that look like for streaming data? You know? That's not something that we people have had. Alright? Like, what does data cataloging and data governance look like for streams? Right? Like, you got it for static data, but what about data in motion? Right? Like, what does that look like? You know, some of those things are near term that that we're thinking about. And,
338
00:42:01.520 -->
00:42:08.500you know, I think I think that, again, it's it's really about building the best experience for for for engineers
339
00:42:08.965 -->
00:42:12.185and and people that are gonna use data to to solve problems,
340
00:42:12.725 -->
00:42:15.145answer questions, and and innovate for customers.
341
00:42:15.605 -->
00:42:31.045 Are there any other aspects of the overall space of building real time data applications and the work that you're doing at Meroxa to support it that we didn't discuss yet that you'd like to cover before we close out the show? No, man. I think I think we we talked about a lot of it. Everybody needs to start thinking real time first.
342
00:42:31.425 -->
00:42:48.015Real time first. That's it. Well, for anybody who wants to get in touch with you and follow along with the work that you're doing, I'll have you add your preferred contact information to the show notes. And as the final question, I'd like to get your perspective on what you see as being the biggest gap in the tool on our technology that's available for data management today. Oh,
343
00:42:48.875 -->
00:42:51.375 it's just so it's it's it's just
344
00:42:51.789 -->
00:42:52.289so
345
00:42:52.589 -->
00:42:57.010fragmented, man. Like like, there's literally a point solution for everything.
346
00:42:57.470 -->
00:43:01.425And I I'm I'm hoping that at some point that there's gonna be consolidation
347
00:43:02.125 -->
00:43:03.025and then consolidation
348
00:43:03.325 -->
00:43:09.745around, like like, jobs, you know, jobs to be done. Right? But, like, man, there's just so much so much just
349
00:43:10.330 -->
00:43:17.790I was talking to somebody this other day. It's like, I don't know if these companies are actually doing well or they're just posting content every day. Right? Like,
350
00:43:18.105 -->
00:43:25.325but, you know, they're posting content every day and it's making people feel like, oh, I need this thing. But, like, when you get into, like, a large organization,
351
00:43:25.705 -->
00:43:31.310you don't really need those things. Right? And it's just like that's the thing where I'm just kinda like, we've just created,
352
00:43:32.250 -->
00:43:33.470you know, this demand,
353
00:43:34.490 -->
00:43:46.650and it's just kinda like you know? I don't know, man. It's it's just that fragmentation is just killing us right now. So I don't think that there is a gap. I think there's honestly too much. Right? Like like, maybe the gap is
354
00:43:47.110 -->
00:43:48.330understanding, like,
355
00:43:49.030 -->
00:43:49.530interoperability
356
00:43:49.990 -->
00:43:52.650between those systems. Right? And how do we make that better?
357
00:43:53.135 -->
00:44:11.425How do we reduce the the the kinda, operational and decision making overhead that people have to make when traversing each of those systems. Right? Because the more things you add on to it, it gets, you know, infinitely more complex. So, yeah, man. That those are the kind of things that that I I just think, like, if we all thought about this in a much better fashion,
358
00:44:11.965 -->
00:44:25.570 and there's some consolidation, like, the state of things for customer experience and providing that value can get a lot better. Well, thank you very much for taking the time today to join me and share the work that you've been doing at MiraXa to help support real time application developers.
359
00:44:26.349 -->
00:44:38.320Definitely a very interesting problem space, and it's great to see that evolve and mature. So thank you for all the time and work that you and your team are putting into that, and I hope you enjoy the rest of your day. I appreciate that, man. Thank you very much for having me
360
00:44:38.780 -->
00:44:39.280 on.
361
00:44:45.345 -->
00:44:46.244 Thank you for listening.
362
00:44:46.545 -->
00:44:49.125Don't forget to check out our other shows, podcast.init,
363
00:44:49.585 -->
00:44:55.530which covers the Python language, its community, and the innovative ways it is being used, and the Machine Learning podcast,
364
00:44:55.990 -->
00:45:00.490which helps you go from idea to production with machine learning. Visit the site at dataengineeringpodcast.com
365
00:45:01.910 -->
00:45:13.390to subscribe to the show, sign up for the mailing list, and read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email hosts at data engineering podcast dotcom with your story.
366
00:45:13.690 -->
00:45:18.910And to help other people find the show, please leave a review on Apple Podcasts and tell your friends and coworkers.