WEBVTT
NOTE
Transcription provided by Podhome.fm
Created: 06/18/2026
23:18:51Duration: 2991.405
Channels: 1
1
00:00:11.440 -->
00:00:15.440Hello, and welcome to the Data Engineering Podcast, the show about modern data management.
2
00:00:16.155 -->
00:00:24.314This episode is sponsored by Data Driven dot I o, the free data engineering interview prep platform built by data engineers for data engineers.
3
00:00:24.795 -->
00:00:33.560Have you ever walked into a data engineering interview and gotten a question that has nothing to do with real data engineering work? Interviewing is its own skill separate from the job.
4
00:00:33.960 -->
00:00:40.360Watch your code execute live, inspect Spark internals, and whiteboard your data models and pipelines and defend your decisions.
5
00:00:41.160 -->
00:00:44.840Unlike SQL only or Python only practice, datagerman.io
6
00:00:44.840 -->
00:00:50.784covers the full interview loop. Star schemas, slowly changing dimensions, grain and fact table design,
7
00:00:51.105 -->
00:00:55.984item potency, watermarks, dead letter queues, change data capture, and back pressure.
8
00:00:56.385 -->
00:01:03.830Every question comes from real data engineer interview loops at Google, Amazon, Meta, Stripe, Databricks, Netflix, and Airbnb.
9
00:01:04.390 -->
00:01:18.195Go to data engineering podcast dot com slash data driven today to start practicing. Your host is Tobias Maci, and today I'm interviewing Jevan Meltay about the challenges of building a reliable streaming system with Kafka. So, Jevin, can you start by introducing yourself?
10
00:01:18.435 -->
00:01:26.275Sure. Hey, everyone. Yeah. Thanks for having me, Tobias. My name is Jevin Malte. I'm a here in Ottawa, Canada. I'm a fractional CTO
11
00:01:26.275 -->
00:01:27.555staff engineer
12
00:01:27.635 -->
00:01:43.030type of person, usually, like, early stage companies. So lots of experience like Zapier or Clio, but also, like, early stuff. Yep. That's that's me. Don't don't necessarily come from, like, a data engineering background, pharma product development, but I've learned a lot from the product the data folks. So excited to be here.
13
00:01:43.954 -->
00:01:50.994Yeah. And it's been increasingly the case where no matter where you're working in the stack, data becomes
14
00:01:50.994 -->
00:02:19.205a bigger and bigger portion of what you have to do even more so with all of these AI powered features that we're jamming into products willy nilly. Yeah. For sure. I mean, data at DataJudaires, you know, they're they're in the background and and not necessarily appreciated, but it's it's gonna be going to become increasingly like the the moat, right, of, having access to maybe customers' data that you can run AI on because it's so easy just to I mean, it's far easier, I should say, for building those AI products,
15
00:02:19.365 -->
00:02:26.805like building products with AI instead. But the data, of course, that's gonna be really the proprietary stuff that is gonna be hard for other competitors to build.
16
00:02:27.569 -->
00:02:32.610And so digging more into that, how did you first get started working in this space of data?
17
00:02:33.170 -->
00:02:45.745Yeah. So I I started I got recruited into Zapier as an engineering manager on the data team. And I think liked my profile, like moving fast, moving scrappy. And so getting embedded there, learned a lot about just really the ETL,
18
00:02:45.825 -->
00:02:49.345like just building traditional ETL pipelines with like Airflow
19
00:02:49.345 -->
00:02:56.980and just where the purpose is really for BI. But of course at Zapier, their whole thing is just moving data around. And so I ended
20
00:02:57.780 -->
00:03:05.220up moving to the growth team and we had a lot of problems with just trying to enrich our data for doing leads and managing sales.
21
00:03:05.300 -->
00:03:21.985And so that's really how I got started into thinking about data pipelines and how we do this and why do we do this only at midnight every day? Why can't we do it in real time? Is there a better way? So that was kind of my first touch with it. And then I moved to a company called Humi, which is a payroll company
22
00:03:22.360 -->
00:03:23.640up here in Canada.
23
00:03:23.800 -->
00:03:28.280I was leading the payroll engineering team. We were moving about $20,000,000
24
00:03:28.280 -->
00:03:29.880a month just in payroll,
25
00:03:30.040 -->
00:03:35.400and we had acquired a bunch of these different systems. And so trying to connect up our payroll,
26
00:03:35.885 -->
00:03:40.284like our front end with our backend, where we were tracking all of our finances.
27
00:03:40.284 -->
00:03:43.485Then we acquired like a time tracking company for like
28
00:03:44.125 -->
00:04:03.230restaurants. And so trying to sync all these things up was like really challenging when you have two different databases at a front end that wants async stuff. So how I started thinking about like, yeah, how do we work with data coming from the product side? And then kind of gig more recently over at Clio where I was managing a staff engineer on this documents team,
29
00:04:03.470 -->
00:04:04.750where we had like 1,800,000,000
30
00:04:04.750 -->
00:04:05.230documents.
31
00:04:05.625 -->
00:04:24.110We had to kind of manage the metadata or move these documents around between the different, a lot of different systems. And so, yeah, so I had to start looking into how to do that. So I didn't go into it to be like, I want to, yeah, just move this data around in the backend, but really coming from product requirements. And so that's where I started coming to the school of learning all these data people.
32
00:04:24.830 -->
00:04:29.150And so now you have started building TypeStream,
33
00:04:29.230 -->
00:04:33.310which is focused on addressing some of these streaming use cases,
34
00:04:33.745 -->
00:04:38.305particularly for people who are more product oriented and don't necessarily
35
00:04:38.384 -->
00:04:53.020have or want to develop a lot of deep knowledge of things like Kafka. And I'm wondering if you can just talk through some of what it is that you're building there and why and how it is that you decided that this is where you wanted to spend at least some portion of your time and energy? Sure.
36
00:04:53.180 -->
00:05:04.915So most places where I've worked, they they had Kafka already. So I mean, three places that I in two of the three places I mentioned, we already had Kafka, which I thought was cool. You know, I'm learning about this technology, but I I started looking into it and understanding
37
00:05:04.995 -->
00:05:08.995all of the functionality that it has that we can be leveraging.
38
00:05:09.155 -->
00:05:15.395I realized we weren't really using any of the cool stuff at all. In fact, I think the reason why we use Kafka is people were just
39
00:05:15.715 -->
00:05:18.740thought it was the right thing to do because it sounds like
40
00:05:19.139 -->
00:05:36.645it's a smart queuing system. It's a queuing system for people who have to move a lot of data when in fact, it's not that at all. In a lot of the cases it was moving. We were using it because we wanted to move data from one database to a different database. We would do either synchronous requests
41
00:05:36.805 -->
00:05:40.245from one to the other at a certain time, or maybe you do it asynchronously.
42
00:05:40.245 -->
00:05:47.180And I was like, Man, why aren't we just using a simple queuing system for that? But I started looking into it. Was like, Wow, why don't Kafka's
43
00:05:47.180 -->
00:06:00.060got this really cool thing called K tables or interactive queries where it can actually be the source of truth that could As you pass these events through, we don't need a database at all. We could just query these little endpoints with the particular
44
00:06:00.595 -->
00:06:12.435optimized and indexed query that I want. Like, wow, that's great. Why are we moving this data at all? Or why don't we just send a notification of a new user signing up into Kafka and then all of our systems
45
00:06:12.650 -->
00:06:24.410react to that instead of having to asynchronously copy it to a different database and then check these things. So I'm blabbering a little bit, but as I was just exploring Kafka more to understand the technology from first principles,
46
00:06:24.905 -->
00:06:29.225I saw that we just weren't using it properly. And there was a lot of functionality on top of it that could
47
00:06:29.625 -->
00:06:34.345that I thought would be really applicable to our team, despite, you know, it being a very heavyweight
48
00:06:34.505 -->
00:06:41.970product to be able to go and deploy and maintain. So we can get to all of those, but that's kind of what led me down this path of wanting to build TypeStream
49
00:06:41.970 -->
00:06:57.375is enabling both data teams, but also product teams to leverage the best that Kafka has to offer without having to be a super heavy data engineer, or even if you are a data engineer, to build this for your product teams who are requesting it, you know, trying to enable those teams.
50
00:06:58.735 -->
00:07:00.575Talking through some of the
51
00:07:00.975 -->
00:07:02.975product engineer challenges
52
00:07:03.055 -->
00:07:11.110around some of these streaming use cases, as you pointed out, you had Kafka already, so there were certainly somebody who at least had some familiarity
53
00:07:11.110 -->
00:07:49.840with what it can provide. But what are some of the situations that you experienced that led to the problem of you're holding it wrong? Yeah. Sure. So I can talk about two examples kind of at Clio. So the first one was we had kind of a central model of database that was huge, and we acquired a different company that was really a different product that was very marketing focused. So in both cases, we had like these user entities. So we had users users tables, organization tables, you know, the first two things that you probably built. And two these are two very separate products. And so we wanted to make sure it's seamless with SSO and and everything. And so so what we do is you'd have, you know, a sync that would happen periodically or
54
00:07:50.294 -->
00:08:00.055worse yet, a user would have to log in and press the sync button from one platform, and it would start just to dump all of the changes from the previous timeframe to the other database.
55
00:08:00.215 -->
00:08:27.045So I think the way I was looking at this is why don't we just instead, from both systems, have a singular event that's schema'd called create user. And we just have this that goes on to this Kafka topic or a channel or, you know, if if you're not as familiar with the Kafka space, and both systems just listen to that. They just create it in their own database if they want. So that would be like the simple place. Simple thing is regardless of where you sign up, you could just have the single place that would just submit it, and you can robustly
56
00:08:27.125 -->
00:08:42.450go and know where to go to to get this. That would trigger an event. And then we could have we could take it further. We didn't. So that was just first thing is let's try to admit these events in a singular place and have a single source of truth. And the other is like, why don't we just build like a central customer database where it would be updated in real time?
57
00:08:42.769 -->
00:08:43.410And so
58
00:08:43.970 -->
00:08:50.085trying to build out like just what's called a K table, where as you have these events come through, it will like materialize
59
00:08:50.165 -->
00:09:30.300these tables that you can go and query using, you know, gRPC or HTTP or whatever to get the most recent up to date information. So you can get a picture of users and you can combine stuff from different topics. So, you know, if they're joining an organization or whatever. So I think this is a far more robust way to try to pull information from different places. The other major project that I think was pretty interesting for this use case was what we called Clio desktop. And this is our own version of a Dropbox type of competitor where, if you're a law firm, you've got your easily 5,000,000 documents that you have stored with us and you want to sync those to your computer, but you don't want to always have them all on your computer. You want to have them maybe cloud synced. Anyways, free product that we were giving away kind of as a lead gen.
60
00:09:30.940 -->
00:09:49.205And we were having a hard time just trying to keep both sides up to date. You have to make a change maybe on the web UI, or you make a change on the client UI, or heaven forbid, you try to do two things within the same few seconds and you get some sort of conflict. And so that's where it would be it was a great use case. Was like, hey, let's
61
00:09:49.525 -->
00:10:06.520be able to whenever you add a file, whether it's local or on server, you have a centralized place for a source of truth saying like create file and both both systems downstream would consume that both on the server and your own desktop client would have, like, an endpoint, and it would just be able to stream in all of those changes.
62
00:10:06.840 -->
00:10:11.160Far better than trying to have two big data stores that you have to do reconciliation
63
00:10:11.160 -->
00:10:14.365between. So pretty interesting for, like, a very large scale product.
64
00:10:14.845 -->
00:10:17.245With the use of Kafka,
65
00:10:17.325 -->
00:10:23.405just the Kafka infrastructure itself can be very heavyweight. It requires a lot of operational
66
00:10:23.565 -->
00:10:24.445overhead,
67
00:10:24.524 -->
00:10:25.325particularly
68
00:10:25.404 -->
00:10:33.180if you aren't using some of the newer builds that have given you a way to drop at least some dependency on Zookeeper.
69
00:10:33.180 -->
00:10:43.095There's the challenge of how much data you have to keep resident on disk for Kafka if you're not using the automatic offload to things like object storage.
70
00:10:43.335 -->
00:10:51.015And so for teams who say, yeah, this is great. I want all of these streaming use cases. I wanna be able to query things automatically,
71
00:10:51.175 -->
00:10:54.210but they don't necessarily have that operational
72
00:10:54.370 -->
00:11:00.610capacity or necessarily the budget to go to a Confluent or an AWS managed Kafka.
73
00:11:00.770 -->
00:11:06.210What are some of the ways that they should be thinking about the trade offs of Kafka
74
00:11:06.475 -->
00:11:07.515operationally
75
00:11:07.755 -->
00:11:08.955and architecturally
76
00:11:08.955 -->
00:11:25.480versus some of these other approaches for building streaming systems, whether that's something like an AWS Kinesis or there have been a number of competitors who are coming and nipping up the heels of Kafka through effectively ground up rewrites that are protocol compatible. So I'm thinking RedPanda,
77
00:11:25.480 -->
00:11:32.920AutoMQ, and Pulsar being the main contenders there. Just how much of it is the interface and how much of it is Kafka specifically
78
00:11:33.000 -->
00:11:34.680that is truly beneficial?
79
00:11:35.295 -->
00:11:47.615Yeah. That's a great question. So I actually don't know the difference, the underlying difference between too many of these, like Kafka compatible ones. I mean, for me, it's like, I don't care. Just as long as it's Kafka compatible, it should be swap like swappable.
80
00:11:48.175 -->
00:12:16.085My my key that I you know, I'm interested as data as a product engineer is like, what are those things that I can build on top of it? Does it still work with all of those great libraries? Kafka Streams, Kafka Connect has all the stuff out of the box that I can connect with. So I actually don't know the difference between a lot of these. I'm maybe a pleb in that respect. But I think the interesting, again, for product engineers is the competitors with like temporal and ingest, which are doing some of the similar things that are far more like orchestration
81
00:12:16.085 -->
00:12:33.985and workflow oriented. So the case that we talked about, say with, you you have a user sign up, right? You could have a temporal workflow where you have user sign up event that's submitted and then you have the workflow that will maybe it'll do the dual right for you to your two different data stores. But Tempora will also manage all of the retries and all of the
82
00:12:34.305 -->
00:12:37.425fallbacks and all of the logic around that, which
83
00:12:37.825 -->
00:13:03.775I think goes a long way to be able to help you build out these very event driven workflows. And so that's where I see like the biggest competitor to something like a Kafka adoption, which I think is actually a very positive thing, right? So they're giving a lot more tooling out of the box to enable people to build these very durable workflows instead of, you know, the traditional thing we'd go and build, or like, I'll go build a cron job or I'll use like a Redis store to manage a cache and
84
00:13:04.495 -->
00:13:08.815just pray it doesn't go down because I didn't really think through about, you know, item potency
85
00:13:09.055 -->
00:13:23.160and things. And so, yeah, that's kind of where I'm seeing things going from from my side. I know it's very I mean, I think I love the idea that I think AutoMQ came up with of, let's just use s three for, like, the backing instead. It's like, I don't know what the the folks at Confluent
86
00:13:23.160 -->
00:13:26.840thought of when that first came out. It was like, oh, shoot. That's a great idea, you know, probably.
87
00:13:27.605 -->
00:13:38.565Yeah. There's definitely a lot of second mover advantage in terms of all of these competitors where they saw Kafka and said, that's great, but here are all the things that make it harder than necessary
88
00:13:38.565 -->
00:13:41.060because of the fact that it was engineered
89
00:13:41.060 -->
00:13:51.220quite a while ago and very much in the, I guess, middle point of Hadoop where everything was Java, everything was focused on the Hadoop ecosystem
90
00:13:51.535 -->
00:13:54.255and how you can work with that. And before
91
00:13:54.255 -->
00:13:58.655s three really came and ate the lunch of the entire HDFS
92
00:13:58.655 -->
00:14:02.815and then led to things like Spark and MapReduce being
93
00:14:03.055 -->
00:14:03.855largely
94
00:14:04.095 -->
00:14:04.655obviated.
95
00:14:05.570 -->
00:14:22.075Well and, I mean, I don't I know those words, but I don't never used any of this stuff. Okay. Just to just to show, like, my bias of where I'm coming from here. Yeah. And and so I struggle, like, at at actually all three places to try to advocate for getting Kafka set up was a burden. Right? It's like, okay, I think the
96
00:14:22.475 -->
00:14:25.915two places we already had it, they had Kafka, but they didn't have like the schema registry,
97
00:14:26.075 -->
00:14:43.600which I would probably argue is like the most important thing that you need. When we were at Zapier, we had a data governance team and it was made of like, I think three people and one person full time, she was amazing. She went and wrangled all of the different teams of all different things that they were sending as just raw JSON
98
00:14:43.760 -->
00:14:49.200onto the Kafka topics, like just unmarshaled. And so like, okay, these are the 123
99
00:14:48.795 -->
00:15:01.195different events that we have. I made them in the schema registry. Could you all please go and migrate to that? And of course they're like, no, there's no benefit to me doing it. I've already done it. But yet like, but yet there was downtime
100
00:15:01.355 -->
00:15:19.435all the time because we'd have an event that would you know, that would be changed upstream and downstream. They were expecting things a certain way. And so that's the other piece is, so we got the Kafka. Kafka itself is like a big thing you gotta deploy. The schema registry is another thing that you have to kind of understand and then want to have some, this idea of type safety
101
00:15:20.075 -->
00:15:39.170and then deploy that. And then you probably wanna have some config as code around that. So, you know, if something goes down, you wanna have all your topics rebuilt in a cert with a certain configuration that they want. So, yeah, it's a lot, it's a lot to manage. You know, we, with TypeStream, the idea there's this lovely thing called the Bayesium, if you're familiar with this and probably a lot of people are, but
102
00:15:39.650 -->
00:15:40.690in a bunch of
103
00:15:41.010 -->
00:15:41.570places,
104
00:15:41.970 -->
00:16:07.589I've talked to, they're like, yeah, Kafka sounds cool. They're like, I just don't see the need for it. I'm like, listen, let me just expose all of your events for you. So within TypeStream, we have it bundled where you could just connect up to your Postgres or your MongoDB from within the project itself, and we'll stand up a durable Debissium for you, which will basically just listen to all of your changes in your database, all of your tables, and put them onto their own respective topics. So you're you're immediately getting all of your you're baking your database asynchronously.
105
00:16:07.910 -->
00:16:18.615You can respond to your database asynchronously out of the box. So that's a pretty common way that I've seen people do it is just throw Devisium on there with your schema registry and put that all start reacting to your database.
106
00:16:19.335 -->
00:16:22.535Talking through a bit more that product focus
107
00:16:22.535 -->
00:16:23.255where
108
00:16:23.350 -->
00:16:25.830Kafka is this powerful substrate,
109
00:16:25.910 -->
00:16:26.870but is
110
00:16:27.030 -->
00:16:41.145very easy to shoot yourself in the foot if you don't know exactly what you're doing that provides a challenge of how do you build an appropriate abstraction layer that makes it more approachable and easier to incorporate in these product focused features
111
00:16:41.225 -->
00:16:42.825without inadvertently
112
00:16:42.904 -->
00:16:54.680going down a path where maybe you have too many topics and not enough consumers or you haven't charted it appropriately ahead of time and how you can deal with some of those early decisions that can have long term consequences.
113
00:16:54.839 -->
00:17:09.695Yeah. The one that that I was it was hard to get my head around is this whole idea of partitions, which probably is like a very newbie data engineering thing to to get their head around. But like, it's no, I just want to have one partition. Like I can understand that. I just want to take all But of course you don't want to do that. You want to be able to support simultaneous
114
00:17:09.695 -->
00:17:20.809consumption of your topics. So, yeah. So this idea with TypeStream is even before that is like, I discovered this thing called Kafka Streams, which is a very powerful Java Kotlin library for Kafka,
115
00:17:20.970 -->
00:17:24.489where you can write kind of these minuscule little Java snippets,
116
00:17:24.570 -->
00:17:25.450and it will
117
00:17:26.090 -->
00:17:28.009allow you just to pass data.
118
00:17:28.090 -->
00:17:40.745You can react to whatever Kafka topic, but you could do transformations in line and you can connect up all these different Kafka streams together and then drop them into say a topic. And then you can have that topic being monitored
119
00:17:40.745 -->
00:17:42.104by a Kafka
120
00:17:42.184 -->
00:17:42.585connector,
121
00:17:43.730 -->
00:17:55.570which is effectively dropping the data into whatever type of sync that you want, S3 bucket, database, elastic search. You know, there's lots of these built in. And so when building I these these things out of these companies that we talked about,
122
00:17:55.890 -->
00:18:02.684I was trying to deploy this and they're like, Java? Jevan, are you serious? Why are you trying to deploy Java? We're a Rails shop. We're
123
00:18:02.684 -->
00:18:05.724Rails people. We are literally Java refugees.
124
00:18:05.804 -->
00:18:18.470That's why we started with Rails because we didn't want to compile things or do this massive configuration. So that was a big barrier. And so the path that I took with TypeStream is like, Hey, let's just do config as code. You can write your JSON.
125
00:18:18.710 -->
00:18:33.705Everything is strictly typed within kind of the, within all the topics that we built, have TypeScreen built on top of. So you can be really confident that that when you build your little JSON defined definition for your whole pipeline, that everything's fully typed inside of it. So it's very composable
126
00:18:33.785 -->
00:18:56.985and you can have all the niceties of tooling around building your, doing your code generators for your clients on top of your topics and stuff. And so I think that's far more digestible for people somehow, if even though it gets compiled down into Java, it's less offensive than if, you know, we just have this JSON file that defines the entire pipeline. And so, yeah, that's kind of the idea is like, let's take the best of Kafka.
127
00:18:57.065 -->
00:19:06.960Let's use Kafka streams. Let's use K tables. Let's have these gRPC endpoints. So I don't have to build a database for every service that I want to build instead of define in
128
00:19:07.840 -->
00:19:26.335JSON and just have this deploy for me. And so, yeah, so the vision is like, yeah, if I'm a data engineer and people are asking me to do this, or if I'm observing that the product teams are like, man, they just want to put up another database for this little service. Like, Hey, here's a tool that we can give you. Use Kafka tables.
129
00:19:26.335 -->
00:19:36.210We've got TypeStream set up for you. And you could just instead query this endpoint and just define the query that you want to have index for. I think as far, will just, will allow us to scale far further and
130
00:19:37.170 -->
00:19:45.330we could, but we can keep that nice scent like source of truth of Kafka in one single place. And we don't have to manage 50 databases across, you know, 70 services.
131
00:19:46.495 -->
00:19:48.975One of the challenges that inevitably
132
00:19:48.975 -->
00:19:52.014happens whenever you are trying to build this
133
00:19:52.174 -->
00:19:56.015more approachable facade on top of some underlying
134
00:19:56.015 -->
00:19:56.895infrastructure
135
00:19:56.895 -->
00:19:57.615is that
136
00:19:58.030 -->
00:20:19.495every abstraction is going to leak. And so you have to determine, okay, where do I put the escape hatches so that if somebody does need to reach underneath and twiddle some knobs to get exactly what they want out of it without having to disrupt everything that's built on top. How do I do that in a way that doesn't feel terrible and actually feels as though it's part of the overall
137
00:20:19.655 -->
00:20:32.490intended experience and becomes intentionally designed rather than an accidental hack that you have to retrofit in in there afterwards. I'm sure you could talk through some of the process that you went through to figure out what were these appropriate
138
00:20:32.570 -->
00:20:33.530interfaces
139
00:20:33.690 -->
00:20:38.570and escape hatches to build into this more product focused layer on top of Kafka.
140
00:20:38.905 -->
00:21:01.110Yeah. I think so because the heart of it is really, let's just expose all of the greatness that has that Kafka has to offer. I'm gushing a little bit. I am a Kafka fanboy. There's no doubt. But let's let's expose what it has to offer to people to make it as approachable as possible. And so with that, the layer that we have on top is quite thin. So really what we have is the compiler that will go
141
00:21:01.590 -->
00:21:02.789and take your JSON,
142
00:21:02.950 -->
00:21:09.724build a graph from that, calculate the types, and then configure your topics in the appropriate way, and then deploy
143
00:21:09.965 -->
00:21:22.845Kafka streams at in a single process there. So that's that's really that's really it. You can do it through a CLI or you can do it through JSON. And so it's it's quite lightweight. The nice advantage of that is you can use your favorite Kafka observability
144
00:21:22.845 -->
00:21:29.860tools on top of that. With using Kafka Kaf Bat, which is like the the successor to Kafka
145
00:21:29.860 -->
00:21:52.809UI, you can go and observe the topics right away. And so you can just the data is not hidden and the abstraction's pretty thin. So and all of that's kind of linked within the TypeStream UI there. But I so I totally agree. And that's what we've really been thoughtful about is like, what is what's the right level of abstraction here? Where if you're like a data engineer who know what you're doing, you know, you can build this JSON, you can go copilot,
146
00:21:53.130 -->
00:22:11.635deploy it, go observe how the topics work in your favorite tool. Or if you're just a product engineer who wants to go and deploy this thing, you don't care, you know, exactly what the underlying thing is. But if something breaks, yeah, you can go take a look and observe the topics, look at the consumer groups, see what's being consumed, and all of that kind of maps
147
00:22:11.795 -->
00:22:14.275to the the TypeStream interface to see,
148
00:22:14.675 -->
00:22:21.989you know, so you can see it all there. So anyway, it's something I was struggling with. That's why why I'm babbling a bit is I totally hear where you're coming from. And
149
00:22:22.870 -->
00:22:23.509so
150
00:22:23.990 -->
00:22:26.309for teams who do have
151
00:22:26.710 -->
00:22:32.214the opportunity to take advantage of some of the niceties of TypeStream where they don't have to
152
00:22:32.615 -->
00:22:43.815do all of the schema registry and manage all of the partition mappings and determine the appropriate topics, etcetera. They say, I just want to be able to get data from over here to over there as fast as possible.
153
00:22:44.299 -->
00:22:45.820What are some of those
154
00:22:46.059 -->
00:22:48.619typical use cases that you're seeing people
155
00:22:49.020 -->
00:22:50.460use TypeStream
156
00:22:50.460 -->
00:22:55.019or even Kafka Sans TypeStream for in a typical
157
00:22:55.260 -->
00:22:57.740application led
158
00:22:56.794 -->
00:22:58.315architecture and design?
159
00:22:58.635 -->
00:23:11.515Yeah. So common ways for talking about moving data between databases and there's other use cases, but I think that's that's a very common one that I've seen and a popular one. So say you have like Postgres and you wanna transform something into Elasticsearch.
160
00:23:11.920 -->
00:23:19.760You want to maybe do some full text search for documents that people have uploaded, for example. And so in Postgres, you've got your row that has the file,
161
00:23:20.080 -->
00:23:21.600it's like the file name,
162
00:23:22.000 -->
00:23:23.920and it's got the reference to S3, for example.
163
00:23:24.745 -->
00:23:34.984And you don't have, let's say you don't have Kafka set up at all. So the first way would be like, let's connect to Bezium, which is a, you know, does the change data capture that can sit on top of Postgres, listen for all the changes.
164
00:23:35.225 -->
00:23:37.705And anytime that what that particular table changes,
165
00:23:38.200 -->
00:23:44.440it will emit an event into Kafka topic. So great. Now we've got an asynchronous
166
00:23:44.440 -->
00:23:54.914list of events that come through. So we've got But it's still got all of the It's just the raw metadata around the file. So the next thing to do is you want to be able to go and extract. Let's get
167
00:23:55.155 -->
00:23:59.394that content from the file itself. So typically what you'd have to do is
168
00:23:59.794 -->
00:24:29.645make a little Java type Java Kafka stream thing that would connect it to the topic and it would have to have the Google or whatever it is, the library to go and download that file. You'd have to parse it. If it's PDF, you got to go and read that. And so great, now you've got the content. And so you'd emit that into a So you take that Let's say simple case, you take that raw content, emit it into a new event or a new topic. Here's the extracted stuff. And then what you could do is connect up what's called Kafka Connect, which I didn't talk about it much, but these are just libraries
169
00:24:30.365 -->
00:24:34.710that exist, most of them free, that connect to all of your favorite data sources
170
00:24:35.029 -->
00:24:56.005that you like. You can pull from data sources or you can drop them. So in this case, we're talking about Elasticsearch. There's a great Kafka Connect for Elasticsearch there. So you could stand up a process that would do that, that would listen to the topic, and you do a little bit of configuration to say, okay, here's from this metadata, you can go and drop it into this particular type of row or whatever it is, the document in Elasticsearch.
171
00:24:56.005 -->
00:24:58.005So it's a very common path
172
00:24:58.440 -->
00:25:07.719of trying to do that whole thing. So with TypeStream, it's a bit simpler. You you could just have open up your TypeStream interface. You can type in your Postgres
173
00:25:07.799 -->
00:25:08.679database
174
00:25:08.840 -->
00:25:09.559and that
175
00:25:10.279 -->
00:25:13.855will launch Debezium. It'll start pulling the tables that you've selected.
176
00:25:14.175 -->
00:25:21.054And we have a bunch of kind of pre configured internal nodes that will like say download file or extract text from file, do OCR,
177
00:25:21.215 -->
00:25:26.810and it'll expose that into another topic. And then you just drag and drop the Elasticsearch
178
00:25:26.810 -->
00:25:40.810sync on there, and it will go and launch the Kafka Connect for you. So that's like one very, very common example, I think, of how people are using it today or should be using it instead of doing like a cron job or a Redis or something to to manage that pipeline for you.
179
00:25:42.275 -->
00:25:45.554When you're talking about Debizium and change data capture,
180
00:25:46.035 -->
00:25:48.835it's definitely great when it works,
181
00:25:48.995 -->
00:26:07.400but there are also lots of edge cases and failure modes that can be very difficult to deal with, particularly when you have a large initial sync that you have to do to be able to populate the current state and then get to the point where you can catch up and just replay with the write ahead logs.
182
00:26:07.880 -->
00:26:10.680The biggest problem is usually that all of a sudden,
183
00:26:11.365 -->
00:26:28.929Kafka or Debesium goes down and fails to consume the logs for a substantial period of time, and so you have to reset the replication state and do the whole sync all over again. And so how do you try to mitigate some of those challenges for people who are
184
00:26:29.169 -->
00:26:43.855dealing with having to run these complex CDC jobs? Or are there cases where you're saying, we recommend it for maybe databases of this scale beyond that, kind of good luck, figure it out. Here are some useful pointers.
185
00:26:44.175 -->
00:26:44.735Right.
186
00:26:45.375 -->
00:26:50.014This comes back to the question of how much abstraction are we building on top of it?
187
00:26:50.495 -->
00:26:55.200And there's people that are just way smarter than we are in the doing the Debizium stuff.
188
00:26:55.440 -->
00:26:58.720And so we just we're just deploying Debizium vanilla
189
00:26:58.800 -->
00:27:02.000with basically the customization that they can compatible with ours.
190
00:27:02.240 -->
00:27:02.960And so
191
00:27:03.760 -->
00:27:06.720when we're getting into like the complexities or edge cases with Debizium,
192
00:27:07.304 -->
00:27:12.664we're pretty like, I'm pretty clear with the product. Like, listen, you gotta, like, you just gotta make this work with Debizium.
193
00:27:12.825 -->
00:27:42.924Most of the time, this is not the issue, especially people who are just getting started with it. Like they're not going to be depending on these events immediately, right? Like they're just starting up for the first time. Certainly downstream. I'm hoping that if people get excited about just using Kafka and they're actually really like leveraging the events that they're emitting downstream from Debesium, then yeah, this is like, you need to really make it production ready or, you know, think about the different edge cases of what can happen, but I'm pretty upfront. It's like, Hey, like Debesium is a whole of the beast to itself, probably very specific expertise
194
00:27:43.245 -->
00:27:52.210that I don't know too much about, but I can stand it up multiple times. I've done that and we could set that up for you super, super easily. But I hear you that there's yeah. There's definitely
195
00:27:52.210 -->
00:27:57.570edge cases around things going down and having delays in consumption, but I
196
00:27:57.890 -->
00:27:59.090I don't touch that.
197
00:28:00.690 -->
00:28:01.250Fair enough.
198
00:28:02.145 -->
00:28:08.065To that point of streaming being very beneficial, but also potentially very complex,
199
00:28:08.385 -->
00:28:13.825one of the ways to make it more approachable is something akin to what you're building with TypeStream of
200
00:28:14.260 -->
00:28:16.100make it friendlier APIs,
201
00:28:16.100 -->
00:28:20.100things interfaces that are more familiar to application developers.
202
00:28:20.740 -->
00:28:27.299But there is also that risk of once you adopt something and you integrate it tightly enough, then
203
00:28:27.455 -->
00:28:33.774you're committed and you have to be able to operationalize it and manage it for all those edge cases.
204
00:28:33.935 -->
00:28:34.894And so
205
00:28:35.055 -->
00:28:39.934for people who do come to TypeStream and say, oh, hey, great. I don't have to figure out Kafka anymore.
206
00:28:40.095 -->
00:28:42.175What are some of the ways that you
207
00:28:42.600 -->
00:28:56.119manage some of that expectation setting of this is the easy on ramp, but the, the on ramp goes much farther than maybe as far as you're willing to go. And so some of those ways of mitigating
208
00:28:56.120 -->
00:28:56.679the
209
00:28:56.995 -->
00:28:57.955expectations
210
00:28:57.955 -->
00:29:04.274and, making sure that people who do go down that path, whether it's TypeStream or some other interface,
211
00:29:04.914 -->
00:29:07.394are appropriately committed to
212
00:29:08.115 -->
00:29:09.715the progressive
213
00:29:09.715 -->
00:29:12.034level of complexity that they're going to have to
214
00:29:12.620 -->
00:29:13.420deal with?
215
00:29:13.820 -->
00:29:16.299Yeah. So I think the
216
00:29:17.740 -->
00:29:20.940phrase that the way that we're framing the
217
00:29:21.340 -->
00:29:23.260products, we're calling it Terraform
218
00:29:23.260 -->
00:29:24.780for Kafka streams.
219
00:29:25.020 -->
00:29:48.579And so it's pretty narrow in terms of what we want to enable. We want to enable you to not have to build a microservice for every single thing. We want you to leverage the data that you already have in Kafka or get the data into Kafka. So you have that network effect of being able to leverage a single source of truth and not having to move data around all the time. And so we really help you on ramp into that. And we have some of the niceties
220
00:29:48.580 -->
00:29:51.139for, you know, beginners
221
00:29:51.139 -->
00:29:55.779ways on ramp, including dBzium and stuff. But I think when you get to a certain scale,
222
00:29:56.595 -->
00:30:04.674some of these things are just gonna have to manage yourself. Debesium, you should have your own instance of that's connected to the database. So, I mean, Shopify got started.
223
00:30:05.235 -->
00:30:20.779They just connected Debesium to everything is what I understand when I was interviewing there. And they just like exposed everything. It's like, cool. Now we have everything there. And then from there they had to learn how to be able to work through all the edge cases and just have a team that's dedicated on just the data, data emission.
224
00:30:21.605 -->
00:30:24.645So I think as people get as teams,
225
00:30:24.885 -->
00:30:29.924as they get more experience or they get larger and they start to have some certain bottlenecks,
226
00:30:30.085 -->
00:30:36.165there's certain things that you should be able to take over yourself. TypeStream, again, because the abstraction is pretty straightforward,
227
00:30:36.540 -->
00:30:43.580you can go and build your own Kafka streams that can connect two different topics, say within the TypeStream pipeline,
228
00:30:43.900 -->
00:30:48.700if you want. Again, because it's all perfectly visible and observable, see how that works.
229
00:30:49.725 -->
00:30:52.764But that's probably with most things, whereas you can start easy.
230
00:30:53.165 -->
00:30:58.205And then once you My son, he's trying to learn to code. I asked a guy who's
231
00:30:58.605 -->
00:31:05.940a coworking space. I'm a heavy AI guy. He was like, How would you recommend it? He's like, you know, Do you start with Claude code and just like build these things?
232
00:31:06.260 -->
00:31:10.259Or do you just like learn the old HTML CSS stuff ground up?
233
00:31:10.659 -->
00:31:15.299And you can go either way. So my son's like, Yeah, I wanna do AI to be able to go and rebuild ArcRaters.
234
00:31:15.575 -->
00:31:20.215I'm like, cool, you can try that. And my hope is that if, as he looks at this abstraction
235
00:31:20.375 -->
00:31:24.375and he's like, okay, now I have to figure out how to integrate like an AI library for these
236
00:31:24.615 -->
00:31:34.029little guys moving, he'll start to go deeper and get some experience into these kind of deeper nuanced things. So that's what comes to mind here is like start with this abstraction.
237
00:31:34.030 -->
00:31:43.390And then as you need to get more experience or you're hitting edge cases, gain more experience with the lower layers down so you can see all the different nuance there.
238
00:31:44.455 -->
00:31:45.575One of the other
239
00:31:46.135 -->
00:31:46.935details
240
00:31:46.935 -->
00:31:48.055that is
241
00:31:48.455 -->
00:31:49.495very useful
242
00:31:49.655 -->
00:31:53.975when you are building when any of these data focused
243
00:31:55.255 -->
00:31:55.895layers
244
00:31:56.055 -->
00:32:01.049is being able to understand how do I appropriately model these systems, what are the
245
00:32:01.770 -->
00:32:11.530proper structures for the objects, how much information do I need to pack into this particular event to make sure that I have it available, how do I manage things like schema evolution,
246
00:32:12.575 -->
00:32:13.215And
247
00:32:13.455 -->
00:32:23.774what are the cases where maybe I'm putting too much information in and I should actually split that into two separate topic streams? I'm just curious how you're seeing people tackle some of those questions,
248
00:32:23.775 -->
00:32:25.535particularly from an evolutionary
249
00:32:25.535 -->
00:32:37.090context of, I just wanna get started. I'll just jam everything in there. Oh, shoot. That blew up, and now I need to figure out how to decompose it a little bit and just some of those data modeling questions about how to appropriately
250
00:32:37.330 -->
00:32:40.610manage the semantics and expectations about those
251
00:32:41.475 -->
00:32:42.754discrete events?
252
00:32:42.995 -->
00:32:43.634Sure.
253
00:32:44.195 -->
00:32:44.914First,
254
00:32:45.155 -->
00:32:50.275I know I mentioned before, but like get that schema registry because you talked about an example with Zapier
255
00:32:50.275 -->
00:32:52.835is like, no one will ever try to
256
00:32:53.395 -->
00:32:55.635be excited about trying to
257
00:32:56.220 -->
00:33:03.980add types to their events later on. Implement that right away and do your marshaling if you can still convince people to use it after
258
00:33:04.300 -->
00:33:04.860that
259
00:33:05.340 -->
00:33:06.220heavy lift.
260
00:33:07.020 -->
00:33:12.154Thankfully, the schema registry does have a bunch of niceties in there, including forward and reverse compatible
261
00:33:12.155 -->
00:33:14.475for your evolution of your events.
262
00:33:14.715 -->
00:33:18.394And so that helps quite a bit where as soon as you're
263
00:33:19.195 -->
00:33:20.955emitting your events upstream,
264
00:33:21.620 -->
00:33:25.620you don't have to necessarily worry about having your downstream consumers immediately
265
00:33:25.940 -->
00:33:26.980able to
266
00:33:27.620 -->
00:33:29.460adapt and consume that if
267
00:33:29.780 -->
00:33:40.395you've thought through the evolution piece there. So that's like a major, major advantage by having the schema registry is you can do some of this evolution and have your two systems decoupled
268
00:33:40.395 -->
00:33:42.635from each other. Hopefully they're
269
00:33:42.795 -->
00:33:49.595not. Or the other thing too with Kafka is you can have it set up where as soon as there's an error consuming,
270
00:33:49.595 -->
00:33:53.209just stop consuming from the pipeline. That buys you time,
271
00:33:53.610 -->
00:34:00.570whether by design or not, to be able to go and do your upgrade and then set it off again and it'll be able to reconsume again.
272
00:34:01.370 -->
00:34:04.729That's where I think there's just so much power in this decoupling,
273
00:34:04.809 -->
00:34:06.250especially when you have this type safety.
274
00:34:06.784 -->
00:34:08.865But as as as yours, you know,
275
00:34:09.425 -->
00:34:23.680yours you mentioned, like, how do we think about this as we go? I'm a pretty scrappy guy. So I'm kinda like, just have to do do your do your day design or your half day design, trying to think through this data model, especially when we talk about
276
00:34:24.160 -->
00:34:30.400windowing and having to combine data across topics. So you have to join a user and an organization,
277
00:34:30.560 -->
00:34:33.280maybe with a file upload, you have to sync through what
278
00:34:34.320 -->
00:34:41.105are the race conditions there? How do I join all that data up? Which you can still do all in Kafka streams and stuff. But it's
279
00:34:41.105 -->
00:34:50.865like most things, sometimes we just have to put it out there and try. But before we put it out there, is this a one way we have to think to ourselves, is this a one way door, two way door to be able to
280
00:34:51.425 -->
00:34:52.545see how much risk
281
00:34:53.550 -->
00:34:56.830we we have when we actually go deploy this thing. And
282
00:34:57.390 -->
00:34:58.030so
283
00:34:58.350 -->
00:35:00.190for people who are
284
00:35:00.670 -->
00:35:01.630adopting
285
00:35:02.110 -->
00:35:03.070TypeStream,
286
00:35:03.310 -->
00:35:17.835what does a typical workflow look like and some of who is the typical persona who actually goes and says, hey. That thing solves the problem that I'm trying to deal with right now and brings it in, and then just the overall process of getting it integrated and deployed?
287
00:35:18.570 -->
00:35:21.370Yeah. So the primary group
288
00:35:21.370 -->
00:35:27.450who really get us and it resonates with are people who are using Kafka and already have an idea of what
289
00:35:28.010 -->
00:35:29.210Kafka streams are.
290
00:35:29.770 -->
00:35:36.685Two major benefits for them is they can just have it config as code, makes it a lot easier for them to deploy and test their workflows.
291
00:35:37.325 -->
00:35:42.205And the second thing is when those who are really kind of in this enabling role,
292
00:35:42.845 -->
00:35:53.540where they're trying to help their team move faster, where the alternative is going to have to build like a bunch of infrastructure, like build a microservice here and a database to support it, or a Redis instance,
293
00:35:53.620 -->
00:35:54.500you know, whatever.
294
00:35:54.980 -->
00:35:57.540Or it's like a high risk thing where they want to explore
295
00:35:57.540 -->
00:36:02.525a new database because that's maybe a lot to stand up. Those are the people who really resonates with.
296
00:36:02.845 -->
00:36:03.565Folks
297
00:36:03.885 -->
00:36:08.605who are more product engineers like myself, unless they've really encountered the same problems that we've described,
298
00:36:09.165 -->
00:36:12.685they just are generally fine with having
299
00:36:13.085 -->
00:36:35.815a SQS queue or using Redis or having a number of different databases around. But I think the primary people are those who are like, yeah, I've used Kafka, but maybe, you know, or they're hearing today, like I had no idea. It can do all those crazy things of like these materialized tables and replays so I can build something two months down the road and it can save all, have all my data there. Those are really the people I think are benefit
300
00:36:35.815 -->
00:36:45.190the most from it. We have the Debussy and thing built in, which I think is pretty exciting for a lot of folks, but to stand it up, it's really straightforward. We can deploy with or without the schema registry.
301
00:36:45.670 -->
00:36:46.470We
302
00:36:46.710 -->
00:36:57.545need it, it needs to be there, but we can deploy it with you. If you have a schema registry in Kafka, you literally just connect TypeStream to that and that's really it. We can manage the processes within our own
303
00:36:58.265 -->
00:37:01.865We use Kafka to manage all the processes. We use a compacted
304
00:37:02.185 -->
00:37:05.225topic if you're really geeking out on the Kafka stuff.
305
00:37:05.465 -->
00:37:07.945And so, yeah, it's really straightforward to be able to deploy.
306
00:37:08.690 -->
00:37:15.329And for teams who you are working with and who bring in TypeStream to solve their
307
00:37:15.490 -->
00:37:16.770data replication
308
00:37:16.930 -->
00:37:17.890challenges,
309
00:37:18.049 -->
00:37:26.905particularly if they don't necessarily have a dedicated data team to work with, what what are some of the most interesting or innovative or unexpected ways that you've seen it used?
310
00:37:27.865 -->
00:37:34.905It's I think people want to try out new new data stores. So folks who are wanting to do now analytics
311
00:37:34.905 -->
00:37:35.385or
312
00:37:35.760 -->
00:37:39.040or really look at some stuff at scale, and maybe they're sitting in
313
00:37:39.920 -->
00:37:44.960a Postgres database and they're like, I really want to do ClickHouse. It's like, great. You can stand up TypeStream.
314
00:37:45.120 -->
00:37:48.240You can just connect to Bezium up right to your Postgres,
315
00:37:48.800 -->
00:37:50.000stand up your ClickHouse,
316
00:37:50.615 -->
00:37:57.175connect, type in the credentials for each and type stream and press play, and it'll just stream it all for you.
317
00:37:58.055 -->
00:38:01.095To be able to do that, by hand, you're deploying
318
00:38:01.734 -->
00:38:02.934There's
319
00:38:02.934 -->
00:38:15.470lots that you're deploying to be able to do that. And so that those are some of the that's probably the most exciting things. People are like, I wanna explore this interesting data store that I could really leverage and speed things up by 10 x to make that really seamless,
320
00:38:15.710 -->
00:38:17.870really easy to do. It's pretty exciting.
321
00:38:19.105 -->
00:38:22.945And as you have been building this product and
322
00:38:23.185 -->
00:38:25.745working on Kafka and
323
00:38:25.905 -->
00:38:28.545getting to know its internals more closely
324
00:38:29.185 -->
00:38:29.905and
325
00:38:30.145 -->
00:38:36.630building a product on top of it, what are some of the most interesting or unexpected challenging lessons that you've learned in the process?
326
00:38:37.590 -->
00:38:42.390I think the question that keeps coming to mind is why is it so hard to deploy this thing?
327
00:38:45.750 -->
00:38:52.035That's that's just such a It's a shame really. And maybe it's by design so that we all go and buy the Confluence mega
328
00:38:52.115 -->
00:38:52.835price,
329
00:38:53.075 -->
00:38:55.234buy the Confluence package for very expensive.
330
00:38:55.555 -->
00:38:58.835But it is a shame that it's so hard to be able to go and deploy
331
00:38:59.140 -->
00:38:59.940Kafka,
332
00:38:59.940 -->
00:39:01.300the schema registry,
333
00:39:01.540 -->
00:39:03.460all of this with high availability
334
00:39:04.500 -->
00:39:06.420in your cloud provider of choice
335
00:39:06.740 -->
00:39:10.580is the thing that's really the thing that keeps coming back to is like, man,
336
00:39:11.060 -->
00:39:13.859I've Or talking to people and just the
337
00:39:15.355 -->
00:39:17.195I'm trying to be always look
338
00:39:17.195 -->
00:39:31.530at first principles for products and be able to go and figure out the best way to use the right tools you have available. But when I'm talking to people, just like, Man, that's a lot to be able to go and deploy. I already have Redis. I've already got Postgres. Why don't I just write all my events
339
00:39:31.930 -->
00:39:32.570as
340
00:39:32.809 -->
00:39:35.850Postgres rows? You're like, yeah, I get it. And I totally do.
341
00:39:36.890 -->
00:39:43.450It's a heavy lift to get to a place, but I think So that's one. And the second is like, yeah, just how underutilized I think Kafka is
342
00:39:43.885 -->
00:39:46.765or all the functionality on top of Kafka
343
00:39:47.165 -->
00:39:48.685that people are just not using,
344
00:39:49.005 -->
00:39:51.565even though they have Kafkas already fully available.
345
00:39:51.965 -->
00:39:53.885And that's also disappointing
346
00:39:54.925 -->
00:40:02.580because it's like, you have all of this that you can build on top of, but yet we just haven't pushed through to find out what those things are because we're,
347
00:40:02.740 -->
00:40:03.220especially
348
00:40:03.940 -->
00:40:05.460like us who've
349
00:40:05.780 -->
00:40:10.820been doing this for more than fifteen, twenty years is like this whole asynchronous event driven
350
00:40:11.140 -->
00:40:19.855architecture is not something that we're very, you know, that grew up with necessarily, right? We have our data store and if you wanna do some sort of asynchronous events, like,
351
00:40:20.415 -->
00:40:21.615you know, it's still
352
00:40:22.575 -->
00:40:24.415new in the past fifteen years for us.
353
00:40:25.380 -->
00:40:30.260So these are it's just a different way of thinking, being event driven versus just looking at synchronous table.
354
00:40:31.460 -->
00:40:32.579One of the
355
00:40:32.900 -->
00:40:35.460common themes that I've been seeing,
356
00:40:36.019 -->
00:40:37.700particularly with this whole
357
00:40:38.145 -->
00:40:39.505boom of
358
00:40:39.505 -->
00:40:41.825foundation model AI providers
359
00:40:41.825 -->
00:40:46.465is that anytime you're building on top of something that you don't fully control,
360
00:40:46.785 -->
00:40:53.720there is that element of platform risk where if everything changes out from under you, you don't really have a lot of recourse.
361
00:40:53.720 -->
00:40:57.080And so given the fact that we do have so many different
362
00:40:57.400 -->
00:40:59.720API compatible implementations
363
00:40:59.800 -->
00:41:01.000that are
364
00:41:01.240 -->
00:41:07.915at least intended to be drop in replacements for Kafka, I'm curious how you're thinking about that from a
365
00:41:08.075 -->
00:41:08.795sort of
366
00:41:09.755 -->
00:41:12.635operational security standpoint for your own work.
367
00:41:13.275 -->
00:41:17.275Do you are we talking about, like, the AI the AI world?
368
00:41:17.515 -->
00:41:21.500I'm thinking of I'm thinking in terms of what you're building where you're currently
369
00:41:21.660 -->
00:41:23.980deploying on top of Kafka specifically.
370
00:41:24.220 -->
00:41:36.065But there are other alternatives that you could incorporate such as the Pulsar with their Kafka protocol layer or Red Panda that is intended to be a drop in replacement or AutoMQ
371
00:41:36.065 -->
00:41:42.065where Right. You're not vendor locked as it were despite the fact that they're open sourced to Kafka
372
00:41:42.065 -->
00:41:42.865specifically.
373
00:41:43.185 -->
00:41:47.745Yeah. Yeah. The the major advantage that we have within this realm, I think, is that
374
00:41:48.180 -->
00:41:48.900we,
375
00:41:49.140 -->
00:42:00.260there's so many good like Kafka connectors that we can go and drop backups into S3 of like all, everything that we have, if we really want to, right? Or I want to now move my Kafka,
376
00:42:00.900 -->
00:42:02.819like all my Kafka topics
377
00:42:04.085 -->
00:42:14.244a different distribution of it or into a different kinesis or something. Well, there's compatibility layers strictly on by connecting topics together where I can now progressively
378
00:42:14.404 -->
00:42:15.925have a new consumer,
379
00:42:17.045 -->
00:42:18.885having whatever it is, like we have
380
00:42:19.500 -->
00:42:28.620Pulsar, I've never tried it, but presumably Pulsar has a way where you can subscribe to Kafka topics or you use a Kafka connector and you can start progressively
381
00:42:28.620 -->
00:42:30.060dropping all of your data,
382
00:42:30.460 -->
00:42:36.055like having it as a second consumer effectively for it, where now you have like the second
383
00:42:36.135 -->
00:42:36.935cluster
384
00:42:37.015 -->
00:42:39.255that's a full replica
385
00:42:39.255 -->
00:42:39.895of
386
00:42:40.535 -->
00:42:42.375the initial one, which I think is really,
387
00:42:42.775 -->
00:42:47.015which is good. So I think because there's already so many compatibility
388
00:42:48.750 -->
00:42:49.630layers,
389
00:42:49.630 -->
00:42:50.830not just like magic
390
00:42:51.310 -->
00:43:02.990connection of these, like of a replication of a cluster, we can just move the data and have another consumer that will translate it to a different place. Or, you know, you can write a very thin layer that will just, you know, read one message,
391
00:43:03.515 -->
00:43:07.675copy it to your new place that you wanna store it yourself.
392
00:43:07.675 -->
00:43:08.235So
393
00:43:09.915 -->
00:43:15.035I'm not I'm not super concerned about vendor lock in that way because it's so easy for me to get my data and move it around.
394
00:43:16.029 -->
00:43:18.750And so for people who are
395
00:43:19.150 -->
00:43:21.069dealing with these challenges
396
00:43:21.069 -->
00:43:22.750of data replication
397
00:43:22.750 -->
00:43:23.550or
398
00:43:23.710 -->
00:43:32.195data freshness or latencies, what are the cases where you would advocate against using TypeStream and maybe just go whole hog into Kafka directly?
399
00:43:33.075 -->
00:43:33.635Right.
400
00:43:34.195 -->
00:43:36.035We wouldn't solve the problem.
401
00:43:36.835 -->
00:43:43.800So usually, you know, Kafka streams will be one of the fastest ways latency wise because it's going to be Kafka native, it's compiled
402
00:43:44.200 -->
00:43:48.359and it's designed to be able to work in between to do the transformations
403
00:43:48.359 -->
00:43:50.520natively, like in line in the pipeline.
404
00:43:50.920 -->
00:43:56.680So moving from, we're not losing anything with Kafka because we're just compiling down into that same library.
405
00:43:57.275 -->
00:43:58.475There's not a
406
00:43:59.115 -->
00:44:00.715ton in terms of
407
00:44:00.955 -->
00:44:13.275like latency or performance that you would gain. It's more if you have a very particular type of transformation that I think you want to be able to write and you really need to have kind of that Java
408
00:44:12.420 -->
00:44:24.099accessing attributes and doing some sort of weird tangly thing that maybe we don't support in native TypeStream node that maybe you want to go and write your own Kafka stream process that will go and consume directly off a topic.
409
00:44:24.795 -->
00:44:29.195Performance wise though, where TypeStream doesn't you'll be fully
410
00:44:29.595 -->
00:44:30.235the
411
00:44:30.635 -->
00:44:36.714same latency you'd get by writing this natively in Java. But you might have to write something in Java if
412
00:44:36.840 -->
00:44:41.480if TypeStream doesn't provide, like, a node or the flexibility that you need in terms of the transformation.
413
00:44:41.640 -->
00:44:43.000Likely, that'd be the reason.
414
00:44:43.880 -->
00:44:45.960And as you continue to
415
00:44:46.120 -->
00:44:48.680build and invest in TypeStream
416
00:44:48.680 -->
00:45:00.295and continue to evolve along with Kafka and the broader ecosystem around that? What are some of the things you have planned for the near to medium term or any particular projects or capabilities you're excited to explore?
417
00:45:00.535 -->
00:45:05.175Yeah. Wanna try it with more with more of these Kafka compatible distributions
418
00:45:05.175 -->
00:45:05.975to see
419
00:45:06.400 -->
00:45:07.040how
420
00:45:07.600 -->
00:45:17.200they perform, I think. Because if one of the advantages, if you have with, I think this thin abstraction that we have on top of it is whether you deploy it with
421
00:45:18.400 -->
00:45:21.280vanilla Kafka or you go deploy with AutoMQ
422
00:45:22.265 -->
00:45:28.345or Memphis or whatever, it should be entirely portable. And so I want to be able to ensure people that
423
00:45:28.745 -->
00:45:33.705regardless of where you have as your Kafka deployment, that TypeStream should be
424
00:45:34.825 -->
00:46:03.675good to work for you. And the other thing is, yeah, we're looking for more people to try it out. It doesn't have a ton of users right now, but if you think there's a, if someone listening thinks they have a particular use case where they're like, I want to move data between these two data stores, it's a real pain. Or like, I want to be able to get my data and use Kafka as my source of truth. And I'd love just to have like a REST endpoint to offer to our engineers to be able to query for the latest materialized version. I'd love to talk to them and see if this
425
00:46:04.235 -->
00:46:06.320would be a fit for you. It's
426
00:46:06.560 -->
00:46:22.965it's BSL. So it's like you can go and take a look at the source code yourself. You can deploy it, you know, for fun, try it out. But we think it's so cool. We just don't want Confluent going, just dragging all that code into their into their deployment, not giving us anything. So
427
00:46:24.484 -->
00:46:29.205Are there any other aspects of the work that you're doing on TypeStream specifically
428
00:46:29.205 -->
00:46:30.885or applications
429
00:46:30.885 -->
00:46:33.640for Kafka and the streaming systems,
430
00:46:33.640 -->
00:46:38.760particularly with a product oriented focus that we didn't discuss yet that you'd like to cover before we close out the show?
431
00:46:39.640 -->
00:46:43.160No. I don't think so. I think those we saw some that's some pretty common
432
00:46:43.800 -->
00:46:49.635yeah. Probably common use cases, maybe some unique ones. Yeah. No. Not really. I think that's
433
00:46:50.115 -->
00:46:51.075kind of it.
434
00:46:51.315 -->
00:47:07.140For anybody who wants to get in touch with you and follow along with the work that you're doing, I'll have you add your preferred contact information to the show notes. And as the last question, I'd like to get your perspective on what you see as being the biggest gap in the tooling technology or training that's available for data and AI management today.
435
00:47:07.460 -->
00:47:09.700Oh, great. Okay. Yeah. You can reach me,
436
00:47:10.100 -->
00:47:12.500jevin @typestream.io,
437
00:47:12.740 -->
00:47:14.580or you can check out typestream.io
438
00:47:14.580 -->
00:47:15.460if you wanna learn more.
439
00:47:16.415 -->
00:47:17.055And
440
00:47:17.295 -->
00:47:18.895the biggest gap,
441
00:47:19.535 -->
00:47:27.775I'm most concerned about kind of data security with the agentic stuff. So I do lots of work with legal tech in my fractional CTO work
442
00:47:28.095 -->
00:47:28.895and
443
00:47:30.080 -->
00:47:33.920passing your data, even though you have like zero data retention to these
444
00:47:34.560 -->
00:47:41.600AI providers is still very concerning for our clients who are working on multimillion or billion dollar deals.
445
00:47:42.240 -->
00:47:43.120And so
446
00:47:43.265 -->
00:47:45.425how we're going to be able to manage that
447
00:47:45.745 -->
00:48:01.780where, you know, just like that whole self hosted idea, we're not going to help self host a model, but like, how do we have that same assurance of like, our data is not going somewhere else, but still running frontier models is like one of the top things that I think we're going to really need to be thinking about. And then the other thing I think about a lot is what
448
00:48:02.340 -->
00:48:04.420we opened with is how
449
00:48:04.660 -->
00:48:07.380where is going to be the competitive edge for
450
00:48:07.540 -->
00:48:08.180companies
451
00:48:10.020 -->
00:48:10.900with their data
452
00:48:11.300 -->
00:48:12.740in the age of AI.
453
00:48:13.395 -->
00:48:14.035A
454
00:48:14.275 -->
00:48:16.835lot of people are just using wrappers around
455
00:48:17.635 -->
00:48:19.075those frontier models,
456
00:48:19.315 -->
00:48:28.250but I think the real advantage is going to be when we can provide a lot more context and offer up our context of the data that we manage as data and product engineers,
457
00:48:28.810 -->
00:48:37.210to the AI. And I think there's still quite a bit of work on tooling around that and and kind of the security and the assurances of how we're managing that data,
458
00:48:37.610 -->
00:48:39.770for privacy. That's we still need to work through.
459
00:48:40.484 -->
00:48:50.165Alright. Well, thank you very much for taking the time today to join me and share the work that you're doing on TypeStream and just your overall efforts of helping product teams
460
00:48:50.244 -->
00:48:52.085gain more capability
461
00:48:52.085 -->
00:49:03.790out of things like Kafka and these complex data focused tools to solve real world problems. I appreciate all the time and effort that you're putting into that, and I hope you enjoy the rest of your day. Yeah. Thanks a lot, Tobias. Same to you.
462
00:49:11.415 -->
00:49:15.735Thank you for listening, and don't forget to check out our other shows. Podcast.net
463
00:49:15.735 -->
00:49:24.615covers the Python language, its community, and the innovative ways it is being used. And the AI Engineering podcast is your guide to the fast moving world of building AI systems.
464
00:49:25.510 -->
00:49:35.510Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email hosts@dataengineeringpodcast.com
465
00:49:35.510 -->
00:49:41.755with your story. Just to help other people find the show, please leave a review on Apple podcasts and tell your friends and coworkers.