WEBVTT
NOTE
Transcription provided by Podhome.fm
Created: 06/01/2026
01:24:29Duration: 3260.302
Channels: 1
1
00:00:11.440 -->
00:00:15.440Hello, and welcome to the Data Engineering Podcast, the show about modern data management.
2
00:00:16.155 -->
00:00:24.314This episode is sponsored by Data Driven dot I o, the free data engineering interview prep platform built by data engineers for data engineers.
3
00:00:24.795 -->
00:00:33.560Have you ever walked into a data engineering interview and gotten a question that has nothing to do with real data engineering work? Interviewing is its own skill separate from the job.
4
00:00:33.960 -->
00:00:40.360Watch your code execute live, inspect Spark internals, and whiteboard your data models and pipelines and defend your decisions.
5
00:00:41.160 -->
00:00:44.840Unlike SQL only or Python only practice, datagerman.io
6
00:00:44.840 -->
00:00:50.815covers the full interview loop. Star schemas, slowly changing dimensions, grain and fact table design,
7
00:00:51.135 -->
00:00:55.855item potency, watermarks, dead letter cues, change data capture, and back pressure.
8
00:00:56.415 -->
00:01:03.830Every question comes from real data engineer interview loops at Google, Amazon, Meta, Stripe, Databricks, Netflix, and Airbnb.
9
00:01:04.310 -->
00:01:08.870Go to data engineering podcast dot com slash data driven today to start practicing.
10
00:01:09.350 -->
00:01:20.125Your host is Tobias Macey, and today, I'm interviewing Weimo Liu about the engineering behind Puppy Graph's zero copy ETL for querying your lakehouse as a graph. So Weimo, can you start by introducing yourself?
11
00:01:20.765 -->
00:01:24.445Hello, everyone. This is Weimo, co founder of PuppyGraph.
12
00:01:24.445 -->
00:01:26.845The name sounds like self driving. And
13
00:01:27.325 -->
00:01:40.329before that, I worked at a graph database staff called Tiger Graph and also Google F1 team. F1 is a unified SQL query engine inside Google. It can query all the data across Google without ETL,
14
00:01:40.330 -->
00:01:43.530and it's serving billions query per day. Yeah. So that's me.
15
00:01:44.695 -->
00:01:47.895And do you remember how you first got started working in the data space?
16
00:01:48.615 -->
00:01:52.615Oh, it's a long story. When I was in college, I'm working on some
17
00:01:53.015 -->
00:01:56.055research projects on some open source spatial database.
18
00:01:56.550 -->
00:02:02.710And after that, I went to GW, the George Washington University, for my PhD degree
19
00:02:02.950 -->
00:02:05.270about database sampling technique.
20
00:02:05.430 -->
00:02:13.685After that, I joined TIGR graph with a closing series A, because the CTO and the co founder is a good friend of my PhD advisor,
21
00:02:13.925 -->
00:02:21.285because he was a professor in database area as well. And laterally, I joined Google working on the SQL query engine. And
22
00:02:21.900 -->
00:02:23.420finally, I
23
00:02:23.420 -->
00:02:25.420convinced my friend to find
24
00:02:25.500 -->
00:02:26.540PuppyGraph.
25
00:02:27.500 -->
00:02:30.299And so digging now into PuppyGraph,
26
00:02:30.299 -->
00:02:39.395can you give a bit of an overview about what it is that you're building and some of the story behind how it got started and why you decided that this is where you want to spend your time and energy.
27
00:02:40.595 -->
00:02:52.275Yeah. So PuppyGraph is a federated graph engine on your tables or even Mongo and other data source. You don't need to load your data to somewhere else, but just connect the PuppyGraph
28
00:02:52.550 -->
00:02:58.470and run graph query, graph pattern, and the graph algorithm on top of it. The story is that
29
00:02:59.030 -->
00:03:01.110since 2022,
30
00:03:01.190 -->
00:03:02.470ChatBet was
31
00:03:02.790 -->
00:03:03.830becoming
32
00:03:03.830 -->
00:03:06.390popular, and some of my friends is
33
00:03:06.710 -->
00:03:08.230some founders of
34
00:03:08.515 -->
00:03:10.355BigLaran model project.
35
00:03:10.355 -->
00:03:12.035And they share with me that
36
00:03:12.674 -->
00:03:21.715no one will write a SQL or any other query in the future, and the agent will do everything. And we feel that, oh, this is a big opportunity.
37
00:03:21.715 -->
00:03:24.710So we're trying to build an engine for agent.
38
00:03:25.030 -->
00:03:26.550And then we think about
39
00:03:27.110 -->
00:03:31.270what's the agent need. And we try to follow the first principle,
40
00:03:31.590 -->
00:03:32.390and then
41
00:03:32.550 -->
00:03:37.750we start to build the Puppy graph. Yeah. Since recently, there's a lot of buzzword
42
00:03:37.365 -->
00:03:46.965like agent harness, but at that time we don't know it yet. But we're trying to, we believe agent need something like that, so we're trying to follow the principle.
43
00:03:47.365 -->
00:03:48.565And currently,
44
00:03:48.725 -->
00:03:53.930we collect a lot of customer feedback as well. And we believe there are three
45
00:03:54.170 -->
00:03:56.650hard requirements to have a successful
46
00:03:56.650 -->
00:03:57.690data agent.
47
00:03:58.170 -->
00:03:59.210One is the
48
00:03:59.769 -->
00:04:03.769process unlimited data. The second is the sub second real
49
00:04:03.769 -->
00:04:04.569time performance.
50
00:04:05.065 -->
00:04:06.505The third is that
51
00:04:06.745 -->
00:04:08.425so called agent harness.
52
00:04:08.665 -->
00:04:12.905And ourselves, we pick a graph query language to meet this,
53
00:04:13.145 -->
00:04:13.785to
54
00:04:14.105 -->
00:04:15.305enable the
55
00:04:15.625 -->
00:04:17.145ontology enforced,
56
00:04:17.225 -->
00:04:22.080which is how we believe we are pretty unique in this area. As you mentioned,
57
00:04:22.480 -->
00:04:23.680the introduction
58
00:04:23.680 -->
00:04:26.400of these agentic capabilities
59
00:04:26.720 -->
00:04:41.165and the graph based knowledge retrieval that a lot of them benefit from has caused a pretty substantial resurgence in the overall interest in graph databases and graph query engines overall.
60
00:04:41.485 -->
00:04:44.125And I know that fundamentally
61
00:04:44.444 -->
00:04:49.780graph data, particularly if you're using a graph native storage layer,
62
00:04:50.100 -->
00:05:01.220is very challenging to scale horizontally because of the constraints of graph topologies, difficulties of figuring out where and how to shard, the impact of super nodes within the overall graph structure.
63
00:05:01.925 -->
00:05:07.125And I'm wondering if you can talk to some of the ways that the architecture of PuppyGraph
64
00:05:07.365 -->
00:05:17.525helps to address some of that or just some of the overall challenges in dealing with graph data, particularly as you move to larger volumes and more complex queries.
65
00:05:19.100 -->
00:05:29.180Yes, yes. So this is a long time challenge in the graph space. I think the first project that solved the graph scalability issue is Prego from Google,
66
00:05:29.340 -->
00:05:32.220and Google published a paper about that. It's called
67
00:05:33.745 -->
00:05:36.065large scale graph processing.
68
00:05:36.384 -->
00:05:36.945And
69
00:05:37.185 -->
00:05:43.425TagGraph is kind of a base that paper to develop it. And after that But internally,
70
00:05:43.425 -->
00:05:55.250we have something involved because it's kind of an iteration run by run, and we feel it's too static and too consuming like MapReduce. And then we make it more flexible.
71
00:05:55.490 -->
00:06:07.505And after I joined Google, I realized, since I was at Telegraph and people like it, but they spent, like, big banks spent eighteen months to load data into it. And our
72
00:06:07.505 -->
00:06:08.785customer complained
73
00:06:09.025 -->
00:06:09.665about
74
00:06:10.065 -->
00:06:16.305this a lot. And after I joined Google, I figured out what's wrong. Since when I was at Google, we query everything.
75
00:06:16.840 -->
00:06:18.920And like if you have logs,
76
00:06:19.000 -->
00:06:19.960because logs
77
00:06:20.440 -->
00:06:31.240follow the certain format, so we just define that table. We don't need to load it, rewrite it, and then query it as a table. And then I think about whether Graph can do it as well. And
78
00:06:31.765 -->
00:06:36.325in this case, if we can just query, for example, your data lake like iceberg
79
00:06:36.325 -->
00:06:43.365and the scalability for storage is no longer a problem, then the problem will be how we can scale the computation.
80
00:06:43.845 -->
00:06:44.325And
81
00:06:45.400 -->
00:06:50.120we do something different since most graph database is sharded by the
82
00:06:50.920 -->
00:06:56.040nodes, and then they have a partition of different nodes. But then it's highly dependent
83
00:06:56.280 -->
00:07:00.440on what's the distribution of a graph. But we are more like
84
00:07:00.775 -->
00:07:02.855shard the data by edges.
85
00:07:02.935 -->
00:07:06.455In this case, even you have a super node like Justin Bieber.
86
00:07:07.014 -->
00:07:09.014He may have too many followers.
87
00:07:09.815 -->
00:07:15.014But if we can just shard it on our edges, for example, if we have 3,000,000
88
00:07:15.014 -->
00:07:18.160followers, we can make it a three partition.
89
00:07:18.320 -->
00:07:22.320And for some others like me, I don't have a lot of followers.
90
00:07:22.480 -->
00:07:23.040But
91
00:07:23.840 -->
00:07:24.640we can,
92
00:07:25.600 -->
00:07:27.360for example, all my teammates,
93
00:07:27.919 -->
00:07:29.360we can be in
94
00:07:29.440 -->
00:07:30.639all the edge between
95
00:07:31.145 -->
00:07:34.505our teammates and the followers can be in the same partition.
96
00:07:34.825 -->
00:07:42.105In this case, the partition side can be similar to each other, and then we can just shred by the edges to solve
97
00:07:42.985 -->
00:07:48.070the hub node situation. And also, we make the graph traverse much easier by
98
00:07:48.470 -->
00:07:52.710doing this. Before that, for example, if you do 10 hop neighbor traverse,
99
00:07:52.790 -->
00:07:55.910it really one more hops and the
100
00:07:56.310 -->
00:07:58.870complexity is increased
101
00:07:58.870 -->
00:07:59.350especially.
102
00:07:59.914 -->
00:08:02.795But now it's kind of a linear increase.
103
00:08:02.955 -->
00:08:04.795And we can solve some
104
00:08:05.035 -->
00:08:07.03510 hop network queries in
105
00:08:07.275 -->
00:08:10.635one or two seconds with cluster because
106
00:08:10.715 -->
00:08:14.010we can shut the edge and also shard the computation.
107
00:08:14.410 -->
00:08:16.330During the hops,
108
00:08:16.890 -->
00:08:23.450during the graph traverse, we can do the shuffling between different nodes. So this is a kind of way to solve the problem.
109
00:08:24.985 -->
00:08:30.425You mentioned too that one of the features or one of the capabilities
110
00:08:30.425 -->
00:08:34.025that you're aiming for with PuppyGraph is to have low latency,
111
00:08:34.345 -->
00:08:41.339and lakehouse architectures in particular struggle with that broadly just because of the architectural
112
00:08:41.339 -->
00:08:46.380fundamentals of it, and there's a lot of work going on at all the different layers to help mitigate that.
113
00:08:46.699 -->
00:08:47.740But I'm wondering,
114
00:08:47.980 -->
00:08:50.699particularly when you're in a data exploration
115
00:08:50.699 -->
00:08:53.500phase or if you are using
116
00:08:54.035 -->
00:08:55.475agent generated
117
00:08:55.555 -->
00:09:01.635graph traversal queries, which might be very complex or require multiple hops, how you
118
00:09:01.955 -->
00:09:06.035mitigate some of those challenges of being able to cut down on latency
119
00:09:06.310 -->
00:09:10.470as well still being able to allow for exploratory
120
00:09:10.470 -->
00:09:12.230and discoverable
121
00:09:12.550 -->
00:09:13.270connections?
122
00:09:14.230 -->
00:09:25.225Yeah. First, there are certain overhead if we read data from a lake house. Like, we don't optimize for ten milliseconds or twenty milliseconds query at all. So
123
00:09:26.905 -->
00:09:29.945in that case, there are a lot of all in memory solution,
124
00:09:30.105 -->
00:09:31.145and we just
125
00:09:31.530 -->
00:09:33.130kind of give up it.
126
00:09:33.370 -->
00:09:34.410What
127
00:09:34.810 -->
00:09:42.490we are good at is like a sub second or a single digit second one. In this case, we still have the overhead, but the overhead usually
128
00:09:42.810 -->
00:09:48.250may be fifteen milliseconds or one hundred milliseconds to fetch from
129
00:09:47.215 -->
00:09:48.575S3, for example.
130
00:09:49.055 -->
00:09:52.255But at the same time, we optimize for the computation.
131
00:09:52.495 -->
00:09:56.015And we also have a vectorized evaluation
132
00:09:58.495 -->
00:10:03.250and also MPP. In this case, we can handle more nodes and edge at the same time.
133
00:10:03.810 -->
00:10:04.690And also,
134
00:10:05.009 -->
00:10:17.089we can scale out more machine, better performance. In this case, we can try to optimize for sub cache second query. And at the same time, because Iceberg has metadata, so we can do active cache or
135
00:10:17.465 -->
00:10:20.185adaptive cache. And then we can just
136
00:10:20.585 -->
00:10:22.105gather data and
137
00:10:22.505 -->
00:10:25.145store the cache in memory and also
138
00:10:25.145 -->
00:10:29.065local disk. And the next time after we
139
00:10:29.705 -->
00:10:30.425read the
140
00:10:30.879 -->
00:10:31.680metadata
141
00:10:31.680 -->
00:10:34.959and we see that, oh, this parquet file was not updated
142
00:10:35.279 -->
00:10:37.360in the last several minutes or
143
00:10:38.160 -->
00:10:44.240is a kind of a cache hit, and then we can just load from local disk or even memory, and then the
144
00:10:44.480 -->
00:11:00.899performance will be much better. There have been other attempts at being able to add a graph traversal layer on top of other storage. I think the graphics package from Apache Spark is probably the most notable one. Some of the recent
145
00:11:01.060 -->
00:11:09.779attempts at that are the KoozieDB project, I think, was aiming for that, and that I think has been taken over by the LadybugDB
146
00:11:09.779 -->
00:11:20.695fork. And then maybe the most notable recent entry is the LanceGraph package for being able to do graph traversals on top of the underlying Lance table format.
147
00:11:20.695 -->
00:11:30.135And I'm wondering what are some of the areas of inspiration or comparison that you would like to highlight between what you're doing with PuppyGraph and some of those other technologies?
148
00:11:31.240 -->
00:11:36.760Yeah. I think our projects are kind of inspired by the GraphX and the GraphFrame.
149
00:11:36.920 -->
00:11:41.639But the tricky part is that Spark is not optimized for graph at all.
150
00:11:42.935 -->
00:11:45.495And also, we saw some friends
151
00:11:45.575 -->
00:11:57.575build on top of other SQL query engine like Treno. And in this case, you highly depend on the compute framework itself. And usually, it's not designed for Graph. For example, if Spark optimized
152
00:11:57.575 -->
00:11:57.975for
153
00:11:58.830 -->
00:12:07.710Spark jobs and also Spark SQL and also like TreeNote optimized for SQL. In this case, the engine to optimize for something
154
00:12:08.430 -->
00:12:09.310else first,
155
00:12:09.630 -->
00:12:13.355and then on top of it, it optimized for Graph. The
156
00:12:13.355 -->
00:12:15.275direction is not the
157
00:12:15.915 -->
00:12:19.835There are no alignment, so the performance is a big bottleneck.
158
00:12:20.155 -->
00:12:26.715And we also see that Lens, Graph, and Kudu. I think they are a very great product,
159
00:12:26.715 -->
00:12:28.315but it's more like
160
00:12:29.089 -->
00:12:35.490small data. Maybe the storage can be scalable because they can just try to raise the iceberg.
161
00:12:35.889 -->
00:12:40.050But I think I will propose this first. And laterally, both of them support
162
00:12:40.050 -->
00:12:41.570a read from object store.
163
00:12:42.275 -->
00:12:48.755But the issue is that if the data is big, and also even the data is not very big, like 200
164
00:12:48.755 -->
00:12:51.075gigabytes, but because computation
165
00:12:51.075 -->
00:12:53.875of graph is very heavy, the data is highly connected.
166
00:12:54.350 -->
00:13:06.430And in this case, shuffle is a necessary feature for this kind of workload. And we are very good at this kind of stuff. And since we saw a lot of all in memory solution like MemGraph,
167
00:13:06.785 -->
00:13:09.985and it's very good at small data because
168
00:13:10.225 -->
00:13:13.345all data load to memory first and then do the computation.
169
00:13:13.585 -->
00:13:16.545And I think CoolGraph and the LessGraph
170
00:13:17.185 -->
00:13:22.010are also kind of a single machine one. And this is a long term problem
171
00:13:22.649 -->
00:13:27.930in the graph world. Since the graph data is highly connected, so no one wants to do the shuffling.
172
00:13:27.930 -->
00:13:28.329And
173
00:13:29.290 -->
00:13:36.135then the bottleneck is that if it's a small data, everyone is publishing benchmark and is very fast.
174
00:13:36.454 -->
00:13:39.895But when it really scale in the industry
175
00:13:40.375 -->
00:13:52.510data size, it will be a potential problem. And we also see most of our customers coming to us for this because their data is too big, because they're already leveraging Databricks,
176
00:13:52.510 -->
00:14:00.430Iceberg, or train or hire things. Data is already there, and it's super big. And when they're trying to have a graph solution,
177
00:14:00.885 -->
00:14:02.005and I
178
00:14:02.725 -->
00:14:10.005think we are very unique in this position. That's why we are a small company, but we have a lot of big logos
179
00:14:10.405 -->
00:14:11.685as our customers.
180
00:14:12.085 -->
00:14:15.845Now digging into some of the data modeling question,
181
00:14:16.620 -->
00:14:22.540your focus is on that zero ETL aspect of you don't have to move your data into a different layer,
182
00:14:22.779 -->
00:14:25.900but a lot of data maybe doesn't necessarily
183
00:14:25.900 -->
00:14:28.700have that natural graph topology
184
00:14:28.700 -->
00:14:34.955or you need to do some explicit modeling of it. And, also, there are a few different flavors of graph
185
00:14:35.435 -->
00:14:36.395definitions,
186
00:14:36.395 -->
00:14:40.475whether it's a labeled property graph or the RDF triples.
187
00:14:40.555 -->
00:14:49.320And I'm wondering if you can talk to some of the ways that you approach some of that data discovery and data modeling aspect of being able to take the existing data
188
00:14:49.639 -->
00:14:53.399as it is in its natural, probably tabular structure,
189
00:14:53.720 -->
00:15:01.515and be able to represent that as a graph and manage the evolution, particularly as the underlying schemas evolve?
190
00:15:01.995 -->
00:15:04.235Yeah. Yeah. You definitely have
191
00:15:05.755 -->
00:15:07.995a deep expertise
192
00:15:07.995 -->
00:15:21.820in this area. Yeah. So this is a typical question. And first, let me talk about the ideal case. The ideal case is that all the table are normalized, and then it's nature to be a graph. Like you have a customer table,
193
00:15:21.980 -->
00:15:24.300you have a product table, you have order history.
194
00:15:25.055 -->
00:15:28.654Actually, history is an edge between customer and
195
00:15:28.975 -->
00:15:32.015product, means customer A, product B.
196
00:15:32.255 -->
00:15:36.975Those are perfect. And at the same time, since in database 101
197
00:15:36.975 -->
00:15:38.575and there are some principle,
198
00:15:38.990 -->
00:15:40.190like everybody
199
00:15:40.589 -->
00:15:48.510do the data modeling. Please create the data tables in a normalized way, and it saves a lot of cost for storage. And
200
00:15:48.990 -->
00:15:49.709also,
201
00:15:50.029 -->
00:15:50.910we provide
202
00:15:51.149 -->
00:16:01.165a very easy and straightforward data modeling on this situation. But of course, this is the ideal case, and some customer are okay if they already normalize
203
00:16:01.165 -->
00:16:02.285some tables,
204
00:16:02.525 -->
00:16:05.885but now they feel, But denormalize is actually for
205
00:16:06.045 -->
00:16:06.765predrawing
206
00:16:07.010 -->
00:16:20.765something and to have better performance. But after they reach out to us, they realize, oh, maybe we don't need to do predrawing, we just run a graph pattern. And it's also very faster, either very fast and in real time. And
207
00:16:21.085 -->
00:16:22.685so this is the ideal case.
208
00:16:23.485 -->
00:16:36.440And some customer either have normalized table or they are okay with normalizing their current table, because before that, the denormalized table is for the single table, wide table performance.
209
00:16:36.600 -->
00:16:40.279And if they can normalize it and still have a
210
00:16:40.839 -->
00:17:06.200very fast performance on graph pattern, which means they can tile drawings but without the slow performance of drawings. So they are pretty happy. Another is that because table is already in production for some other use case, they don't want to change it at all. There is some tricky part for us. And one way that we have a logical view and then define graph on a logical view. Another possibility
211
00:17:06.200 -->
00:17:13.480that we have a very flexible mapping. For example, if you already draw in the customer profile and
212
00:17:13.880 -->
00:17:15.080product profile
213
00:17:15.080 -->
00:17:17.720as the wide table for order history.
214
00:17:18.184 -->
00:17:29.864In this case, we can define one column like a ZIP code as a node. So it's not a node table, it's just an attribute in a wide column. But we will dedupe it for you logically,
215
00:17:29.865 -->
00:17:35.520and then you can have a flexible mapping from the graph schema to your tables.
216
00:17:35.920 -->
00:17:44.320So this is another way we do it. And of course, because for the same 100 table, for example, we can create a lot of
217
00:17:44.644 -->
00:17:46.164different graph schema.
218
00:17:46.404 -->
00:17:52.324And different graph schema have a benefit and have an advantage and a disadvantage.
219
00:17:53.044 -->
00:17:58.644And usually, they highly depend on the use case and the query they are running. And then we can
220
00:17:59.890 -->
00:18:04.130suggest the best GRASS schema. But usually,
221
00:18:04.210 -->
00:18:05.169our customer,
222
00:18:05.410 -->
00:18:07.490because they want to build some
223
00:18:07.890 -->
00:18:09.730customer facing agent system,
224
00:18:09.810 -->
00:18:15.015so they are pretty familiar with their tables. In this case, we will discuss together and
225
00:18:15.175 -->
00:18:22.934build some graph schema best for their use case. Yeah, but of course, sometimes it's not that good, but
226
00:18:23.655 -->
00:18:24.295we just,
227
00:18:25.630 -->
00:18:32.509and if it works, it's fine, but definitely there are always space to optimize it.
228
00:18:34.030 -->
00:18:36.669Beyond the relational structures,
229
00:18:36.990 -->
00:18:38.705there are also potentially
230
00:18:38.705 -->
00:18:43.185document models you mentioned that you're able to execute across MongoDB
231
00:18:43.185 -->
00:18:44.785as a storage layer.
232
00:18:45.025 -->
00:18:48.465And also, increasingly, we're looking to unstructured
233
00:18:48.465 -->
00:18:52.065data sources and doing some transformation of that into
234
00:18:52.830 -->
00:18:53.789semantically
235
00:18:54.029 -->
00:19:12.924enriched deep data or extracting structured data from free text. And I'm wondering how you're addressing some of those as well and being able to map that into a graph structure and maybe some of the interesting use cases that you unlock because of the fact that you're able to work across these different storage layers.
236
00:19:13.485 -->
00:19:14.845Yeah. So for
237
00:19:15.085 -->
00:19:15.565Mongo,
238
00:19:17.630 -->
00:19:24.190so it really depends on how unstructured the data will be. And for Mongo, it's actually pretty
239
00:19:24.430 -->
00:19:28.350structured already since most of the collection follow the same pattern and
240
00:19:29.310 -->
00:19:36.385something like a JSON file as well. And even though they are not flattened at the table, but logically,
241
00:19:37.185 -->
00:19:38.625you can just flatten,
242
00:19:38.625 -->
00:19:40.865you can have a flattened table. Like,
243
00:19:41.345 -->
00:19:46.865when I was at Google, we also do a lot of this work, like if it's a nested field, like aws.c,
244
00:19:47.240 -->
00:19:54.040we can just use SQL to select adobe. C from like a collection A, something like that. And
245
00:19:54.280 -->
00:19:57.640in this case, we can just connect MongoDB
246
00:19:57.640 -->
00:20:06.144with a JDBC interface, or maybe it's a surprise to some of our audience, like MongoDB can access by JDBC.
247
00:20:06.144 -->
00:20:24.639And in this case, we can just run the query similar to the table ones. And also for Mongo collection, there are still something like a foreign key. And then we can use the key to link to each other to form a graph. And for even more structured data, like, for example, PDFs,
248
00:20:25.120 -->
00:20:25.919we have
249
00:20:26.085 -->
00:20:32.565two partners. One is from Treno team, one is my old friend at Google. They are building the
250
00:20:32.885 -->
00:20:34.085index of
251
00:20:34.565 -->
00:20:40.325Google Search. So what they are doing is that they already have a bunch of documents like PDF,
252
00:20:40.640 -->
00:20:42.880and they do the entity expression
253
00:20:42.880 -->
00:20:50.480and store it in, for example, expert table. This is an interesting part because for a lot of abstract data like PDF,
254
00:20:50.720 -->
00:20:55.195assuming the Bank of America bank statement, even they are unstructured.
255
00:20:55.275 -->
00:20:58.475But the unstructured data itself has
256
00:20:58.475 -->
00:20:59.515some potential
257
00:20:59.515 -->
00:21:05.915structure inside. Like you have all the PDFs follow the same pattern, and they have a bank
258
00:21:06.760 -->
00:21:07.960account number,
259
00:21:08.440 -->
00:21:10.519they have the home address,
260
00:21:10.760 -->
00:21:21.785they have a account holder name, and they also have transaction tables there. So there are certain rules, and our partner will help us to do the extraction
261
00:21:21.785 -->
00:21:23.465from the PDFs,
262
00:21:23.625 -->
00:21:25.785and then to markdown, and then to tables.
263
00:21:26.025 -->
00:21:26.585And
264
00:21:26.905 -->
00:21:30.025since both of our partners work
265
00:21:30.185 -->
00:21:31.625in this area
266
00:21:31.625 -->
00:21:38.230for many years, like the Google Search team and also the Trindle team. So we just
267
00:21:38.470 -->
00:21:40.150partner with them closely.
268
00:21:40.150 -->
00:21:43.990And what we are doing is just you already have tables,
269
00:21:43.990 -->
00:21:45.590and then you can have a
270
00:21:45.830 -->
00:21:48.390graph query, and also you can have
271
00:21:48.784 -->
00:21:50.864agent system to query it.
272
00:21:51.105 -->
00:21:51.664And
273
00:21:52.225 -->
00:21:55.744we already have a lot of drawing customer, and
274
00:21:56.065 -->
00:21:59.345it's pretty smooth. And it's even better than,
275
00:21:59.505 -->
00:22:03.265for example, just to make chunks and then do the embedding.
276
00:22:03.780 -->
00:22:04.500Since
277
00:22:04.660 -->
00:22:09.220when do the chunks and embedding, actually you didn't leverage your
278
00:22:09.460 -->
00:22:11.860potential structure
279
00:22:12.100 -->
00:22:17.380inside of your documents. For example, if in the case that every PDF have a
280
00:22:17.940 -->
00:22:18.660different pattern,
281
00:22:19.045 -->
00:22:20.485then maybe the
282
00:22:20.725 -->
00:22:27.044embedding is better. But if you have like 1,000,000 PDFs and all of them follow the
283
00:22:27.525 -->
00:22:28.644exact pattern,
284
00:22:28.645 -->
00:22:30.965it's actually a structured data,
285
00:22:31.365 -->
00:22:32.565a structure
286
00:22:32.960 -->
00:22:34.080representation.
287
00:22:34.560 -->
00:22:40.240So in our experience, if we can flatten those tables, it'll have a lot of benefits.
288
00:22:41.120 -->
00:22:42.320Digging into
289
00:22:42.400 -->
00:22:48.254PuppyGraph itself, I'm wondering if you can give a bit more detail on the architecture,
290
00:22:48.254 -->
00:23:03.620some of the technology choices that you're investing in to enable this use case and some of the core ecosystem primitives that you're leaning on to be able to manage the complexity of the space that you're working in?
291
00:23:05.220 -->
00:23:12.180Yeah. So since the beginning, because we want to build a system for agents, so we consider a different ecosystem
292
00:23:12.180 -->
00:23:13.300and which
293
00:23:13.545 -->
00:23:17.225part of the system and which component to pick.
294
00:23:17.465 -->
00:23:23.785Like the first is that I think three years ago, we believe that the even before that, after
295
00:23:24.345 -->
00:23:26.345the expert team left Netflix,
296
00:23:26.750 -->
00:23:33.710we believe they will be the standard for OLAP in the very near future. So we pinged the founders and
297
00:23:34.110 -->
00:23:49.205showed them how we can run a three hops graph query on Iceberg without any change. And it's much faster than most graph database on the market, and they also feel surprised because Iceberg is not optimized for graphs. And
298
00:23:49.765 -->
00:23:51.605we work this closely.
299
00:23:51.605 -->
00:23:57.125And at the beginning, we only support Iceberg, but my teammates question, what if Iceberg
300
00:23:57.610 -->
00:24:00.969won't be popular or is not become
301
00:24:01.529 -->
00:24:23.885popular fast enough? And then we will just bankrupt. And then we support other different data source, but our favorite is Iceberg. And another thing we are waiting for is that we're waiting for the agent is capable enough to generate all the query automatically without human in the loop. And in this case, we're to pick up an interface.
302
00:24:24.285 -->
00:24:29.005And we consider SQL at the beginning, but we feel that
303
00:24:28.190 -->
00:24:43.684because when human writes the SQL, there are lots of context in their mind. And when they write something wrong and they know it because they know the building logic, like a student won't join with a teacher's salary table, otherwise the student will have 100
304
00:24:43.684 -->
00:24:49.124income last year or something like that, but without a throw out an error. And
305
00:24:49.605 -->
00:24:57.360then since I work at TechRaf, I think the graph is the best because the graph itself not just contains data
306
00:24:57.680 -->
00:25:02.240but also contains the ontology. And if people are doing ontology,
307
00:25:02.400 -->
00:25:11.144the ontology results in the graph, why we just query the graph rather than use the graph to generate better SQL? It's against the first principle.
308
00:25:11.785 -->
00:25:23.384And then we support the Cypher and the Gremlin at the same time because they are popular and there are enough public data on GitHub, like the large model and stand, and generate the query.
309
00:25:23.465 -->
00:25:23.945And
310
00:25:24.289 -->
00:25:25.009also,
311
00:25:26.049 -->
00:25:32.929we only provide a Docker to our customer because then they can just use Kubernetes to deploy it easier.
312
00:25:33.010 -->
00:25:49.835And so this is basically like the interface layer, the graph query, like Cypher and Gramming. And for the computation part is what we are doing by ourselves. And for the storage layer, we just leverage the table format. And the best one, of course, for us is
313
00:25:50.155 -->
00:25:57.210the iceberg. Then we design this system, and also it will be easy to leverage by a different community.
314
00:25:57.210 -->
00:26:02.330Like we connect the community of graph work and the community of common
315
00:26:02.330 -->
00:26:04.169data engineer work. In
316
00:26:04.649 -->
00:26:05.129the
317
00:26:05.690 -->
00:26:08.570last ten years, I think the graph community
318
00:26:08.570 -->
00:26:09.769is separate
319
00:26:10.255 -->
00:26:22.575with data engineer community. Since the data engineer community is involved a lot, but the graph community has still not changed a lot since the last ten years. And I think if we can bring the capabilities together,
320
00:26:22.990 -->
00:26:29.789and, it's not just benefits the agent as we design, but also benefits the human user as well.
321
00:26:30.269 -->
00:26:34.029One of the other use cases for graphs,
322
00:26:34.269 -->
00:26:44.275particularly in the data engineering community that I've come across a few times is for master data management and being able to do things like named entity reconciliation
323
00:26:44.275 -->
00:26:50.835to be able to say these two documents that are talking about slightly differently worded
324
00:26:50.929 -->
00:26:55.250entities are actually the same thing and being able to do some of that resolution
325
00:26:55.250 -->
00:27:03.169there. And I'm wondering what you're seeing as far as applications of PuppyGraph in that more I'm gonna use air quotes and say traditional
326
00:27:03.575 -->
00:27:17.255data warehousing use cases in addition to these more agentic workloads and maybe some of the cases where those two coincide of being able to use agents to do some of that master data management and entity linking.
327
00:27:18.390 -->
00:27:19.990Yeah. So we saw some
328
00:27:20.230 -->
00:27:23.990user are using Graph database to do entity resolution,
329
00:27:24.150 -->
00:27:26.790and also there are some other ways in SQL.
330
00:27:26.870 -->
00:27:27.429And
331
00:27:27.750 -->
00:27:30.310we're trying to see, since now
332
00:27:30.625 -->
00:27:31.744you know, our
333
00:27:32.385 -->
00:27:33.024theory
334
00:27:33.985 -->
00:27:34.624and
335
00:27:35.025 -->
00:27:40.624every data warehouse and the data lake now support graph. So we're trying to apply the data
336
00:27:40.625 -->
00:27:41.585resolution
337
00:27:41.585 -->
00:27:42.625solution
338
00:27:42.625 -->
00:27:44.625of a graph database to the
339
00:27:45.260 -->
00:27:47.420SQL tables ecosystems.
340
00:27:47.580 -->
00:27:52.060And then the user can just apply the existing solution
341
00:27:52.300 -->
00:28:02.655to the SQL one. And also, what we're doing is that we can write back the result to tables. In this case, and for example, Iceberg will be the bus of the
342
00:28:03.055 -->
00:28:06.095data pipeline. And like a bus,
343
00:28:06.495 -->
00:28:29.274write the result back, and we read the result from Iceberg and other engine like Treno and the Spark SQL, read tables and write tables. In this case, we don't need to talk with each other and do the data loading. Everybody just read from Iceberg and writes to Iceberg. And then your output can be our input, our output can be other input. And we're trying to make the data pipeline easier.
344
00:28:30.315 -->
00:28:34.874Now circling back around to what you were commenting with some of these other
345
00:28:35.115 -->
00:28:42.539more point solution graph engines that are very efficient on smaller scales and volumes of data.
346
00:28:42.860 -->
00:28:46.059What are some of the ways that you're seeing people maybe
347
00:28:46.220 -->
00:28:51.580use both in concert where they've got a dedicated graph engine for
348
00:28:51.865 -->
00:28:57.864their data that needs to be low latency and in the hot path of a certain workload,
349
00:28:58.105 -->
00:29:11.160but then using PuppyGraph for more of that scale out across larger volumes of data and maybe being able to transfer data to and from the hot path and into the more warm path at the iceberg layer.
350
00:29:11.880 -->
00:29:18.440Yes. Exactly. So this is what we're expecting. Like, in SQL world, it's pretty common. Like, you have PostgreSQL,
351
00:29:18.555 -->
00:29:22.395and also you have Snowflake Databricks Trino.
352
00:29:22.395 -->
00:29:27.355And then you have Postgres to handle the transactional CRUD with ACID.
353
00:29:27.435 -->
00:29:30.875And for the large streaming data or batch
354
00:29:31.035 -->
00:29:35.820data, and then you use Trino or Snowflake or Databricks.
355
00:29:36.220 -->
00:29:41.260But before that, in graph word, it seems all the stuff similar to PostgreSQL
356
00:29:41.260 -->
00:29:42.140position.
357
00:29:42.540 -->
00:29:45.820And no one cares about the OLAP
358
00:29:45.820 -->
00:29:56.914one. And we're trying to be the OLAP one. And at the same time, for graph database, because people won't store all the data in graph database. Usually, they handle the hot data or transactional data.
359
00:29:57.155 -->
00:30:20.375And in this case, in SQL world, it's just made like a CDC or some data loading things, like you wear the AirBiz jacket, right? So you download the data from Postgres SQL to Iceberg, for example, and then to do the OLAP. I think, hopefully, in graph world, this can be the common practice as well. Like, you don't want to store all the historical
360
00:30:20.775 -->
00:30:26.535data in graph database. That's too expensive. And also, it affects your
361
00:30:26.535 -->
00:30:28.215transactional QPS
362
00:30:28.215 -->
00:30:29.335when you run
363
00:30:29.495 -->
00:30:31.495some heavy analytical query.
364
00:30:32.000 -->
00:30:37.040And also it's very slow since we see a lot of time out and auto memory
365
00:30:37.040 -->
00:30:38.160to run
366
00:30:38.240 -->
00:30:44.720that kind of query. But if you can still use a graph database as transactional updates and then
367
00:30:45.200 -->
00:30:46.320load the data
368
00:30:46.565 -->
00:30:51.684to the iceberg and then use, for example, PuppyGraph to run the OLAP query.
369
00:30:52.005 -->
00:31:06.539And then it will have a lot of benefits, which proved very well in the SQL world. So hopefully this can be the common practice. And also we help some customer keep in the graph world. Before that, they're trying to migrate away from
370
00:31:06.780 -->
00:31:08.299Neo4j to
371
00:31:08.380 -->
00:31:09.340PostgreSQL
372
00:31:09.340 -->
00:31:21.995because of ecosystem problem. And we said that you don't have to do it now. Graph have OLAP as well. And then they just do the data loading on the CDC from Neo4j to Iceberg, and we
373
00:31:22.075 -->
00:31:24.554run a query on top of Iceberg.
374
00:31:25.275 -->
00:31:32.300I'm interested in digging a little bit more into some of that translation layer and the data modeling and representation
375
00:31:32.380 -->
00:31:34.300of graph structures
376
00:31:34.300 -->
00:31:36.300in the Iceberg ecosystem
377
00:31:36.300 -->
00:31:38.540because Iceberg was designed
378
00:31:38.940 -->
00:31:39.820primarily
379
00:31:40.060 -->
00:31:54.424with tabular structures in mind. And I know that, for instance, Kuzu DB actually uses a columnar representation under the hood, but I'm just curious what are some of the points of impedance mismatch between a graph native representation
380
00:31:54.424 -->
00:31:58.559on disk, for instance, from something like a MemGraph or a Neo four j
381
00:31:58.720 -->
00:32:01.840and how to actually do that translation
382
00:32:01.840 -->
00:32:05.759into Iceberg and mapping to and from the structural
383
00:32:05.840 -->
00:32:06.480semantics?
384
00:32:07.615 -->
00:32:08.575Yeah. So
385
00:32:08.975 -->
00:32:09.695in
386
00:32:10.895 -->
00:32:13.054our design, we decoupled the computation
387
00:32:13.455 -->
00:32:17.854and the storage at all. In this case, the computation is still on
388
00:32:18.174 -->
00:32:20.734graph mode. But the storage,
389
00:32:20.815 -->
00:32:27.440just since what we're doing then, we define the node operator and the edge operator. And the operator's
390
00:32:27.440 -->
00:32:32.399input and output are collection of nodes and edges. And in this case, we're assuming
391
00:32:32.720 -->
00:33:07.084all the graph query, graph pattern, or graph algorithm can be a combination of node operator and edge operator. Then we can do the cost based optimization. And for a single operator, because the input and output are collection, so we can do the MPP and also vectorize the evaluation. And of course, in this case, because it's a collection, so the column based storage is really, really important. And then final stage, we still need to fetch the data. In this case, we just run the Parquet file reader. And then read the Parquet file and translate it into collection,
392
00:33:07.325 -->
00:33:09.084and then to the computation.
393
00:33:09.085 -->
00:33:16.204And I think this is an interesting part. Before that, all the graph database, they're trying to speed up to support
394
00:33:16.605 -->
00:33:28.210a complex query, but it's still close to row based. So because it's hard to support the high QPS transactional updates. But in SQL world, everybody know that the OLAP and the OLTP
395
00:33:28.290 -->
00:33:31.245need to have a different storage, and OLTP
396
00:33:31.245 -->
00:33:34.685need a row based one and the OLAP need a column based. And
397
00:33:35.085 -->
00:33:36.205with the
398
00:33:36.685 -->
00:33:40.125column based one, it's much more memory efficient.
399
00:33:40.285 -->
00:33:40.845And
400
00:33:41.245 -->
00:33:44.045in this case, we can handle much larger data.
401
00:33:44.480 -->
00:33:49.440And at the same time, the query complexity is no longer a problem since, for example,
402
00:33:49.840 -->
00:34:06.635one CPU instruction can handle a vector of nodes and edge. And at the same time, because we only access the necessary attribute, like even one node or add have 100 attributes. But maybe for single query, only three or four is related. If column based, we can just leave
403
00:34:06.955 -->
00:34:08.075all the other
404
00:34:08.235 -->
00:34:10.07597 or 96
405
00:34:10.075 -->
00:34:13.515attributes on disk. In this case, it's much more memory efficient.
406
00:34:14.650 -->
00:34:16.490One of the other challenges
407
00:34:16.570 -->
00:34:18.890for a product like PuppyGraph
408
00:34:18.890 -->
00:34:21.050is that graph engines,
409
00:34:21.130 -->
00:34:26.970as we mentioned before, have been somewhat niche for a while. They're not as broadly adopted
410
00:34:26.970 -->
00:34:29.530as a Postgres or a Snowflake.
411
00:34:29.745 -->
00:34:36.865And I'm wondering what are some of the areas of education that you've had to invest in to help people understand
412
00:34:36.865 -->
00:34:42.785the power and benefits of having that native graph structure and graph traversal capability
413
00:34:43.250 -->
00:34:47.010available to the underlying data that they're already investing in?
414
00:34:47.730 -->
00:34:56.770Well, I think this is a kind of a chicken egg problem since before that, the investment before you run the first graph queries are too heavy.
415
00:34:57.010 -->
00:35:01.685So even some perfect graph use case, people still want to, for example,
416
00:35:02.165 -->
00:35:07.525write a SQL or write a Spark job to do it. Because even it's slow and complicated,
417
00:35:07.605 -->
00:35:12.485but you don't need to do a lot of have another copy of data and have another pipeline.
418
00:35:13.170 -->
00:35:18.450And in this case, I think it is hard for the user to adopt the native
419
00:35:18.530 -->
00:35:19.490graph engine.
420
00:35:19.730 -->
00:35:24.210And at the same time, we feel that actually there are certain requirements,
421
00:35:24.210 -->
00:35:29.755and a lot of users just give up the use case after they try different ways. Like
422
00:35:29.755 -->
00:35:31.755when they have very complex things,
423
00:35:32.235 -->
00:35:34.715they try the complex SQL, but
424
00:35:35.275 -->
00:35:36.795it's very soon
425
00:35:36.955 -->
00:35:37.675become
426
00:35:37.755 -->
00:35:40.155no longer, it's no longer human readable.
427
00:35:40.155 -->
00:35:42.635And there are hundreds of lines SQL
428
00:35:42.900 -->
00:35:43.540or
429
00:35:43.700 -->
00:35:45.380either too slow
430
00:35:45.619 -->
00:35:59.015or like if a lot of customers have 1,000 tables, but in daily work, there are only 20 or 30 tables are used. All the others are just left there and no one accesses it at all. But now
431
00:35:59.255 -->
00:36:03.095because we show the possibility to a lot of our customer
432
00:36:03.335 -->
00:36:07.735and they see that, for example, they can write very complex query
433
00:36:07.815 -->
00:36:09.735short way, like ten
434
00:36:10.295 -->
00:36:17.050ten lines of query. The expression capability is the, more than 100 lines SQL.
435
00:36:17.130 -->
00:36:24.010And then they feel that and the the interesting part is that when they feel that, oh, this works and can just return the result,
436
00:36:24.250 -->
00:36:25.370they will keep
437
00:36:25.755 -->
00:36:27.035trying our
438
00:36:27.035 -->
00:36:27.835capability
439
00:36:27.835 -->
00:36:30.075and write more complex query.
440
00:36:30.394 -->
00:36:32.234And this is more
441
00:36:32.555 -->
00:36:38.635common when they are using an agent since agent don't care how complexity the query will be because
442
00:36:38.875 -->
00:36:48.060people just ask us some question and assign a task to agent, and the agent will decouple into subtasks.
443
00:36:48.220 -->
00:36:50.300And each subtask can be more
444
00:36:51.099 -->
00:36:52.140complexity,
445
00:36:52.220 -->
00:36:53.339more and more complex.
446
00:36:53.895 -->
00:36:55.095And in this case,
447
00:36:55.815 -->
00:36:59.495some sometimes they send the logs to us to help let us debug.
448
00:36:59.735 -->
00:37:14.190We feel that even the graph query, the one hundredth line, this no longer can be readable. But the agent can just write the correct one, which is a surprise for us. And we feel that if we show the stronger capability
449
00:37:14.589 -->
00:37:18.430and, like, a more complex query can be handled in a short time
450
00:37:19.310 -->
00:37:22.510and the response is in real time, and then
451
00:37:22.805 -->
00:37:27.365people don't care, like, they want to issue more complex query
452
00:37:27.605 -->
00:37:28.245and
453
00:37:28.485 -->
00:37:31.125assign the agent more complex tasks.
454
00:37:31.365 -->
00:37:39.849So we believe that the usage will be larger and larger. Before that, maybe it's limited because the limitation
455
00:37:39.930 -->
00:37:45.849of the tools. So people have to give up some wonderful idea. But now they can just try, and
456
00:37:46.490 -->
00:37:47.450also the
457
00:37:47.775 -->
00:37:50.095agent that can help them to try. So
458
00:37:50.255 -->
00:38:00.895is the cost is pretty low now. So they can try some fancy idea without heavy invest. Another element of the overall graph ecosystem
459
00:38:00.974 -->
00:38:01.934that has
460
00:38:02.440 -->
00:38:09.320varying levels of support depending on the underlying engine or the language that you're working within is the
461
00:38:09.560 -->
00:38:12.120core graph query and traversal,
462
00:38:12.280 -->
00:38:15.880and then a lot of engines will add another layer of
463
00:38:16.315 -->
00:38:22.395out of the box graph algorithms or graph machine learning or data science capabilities
464
00:38:22.395 -->
00:38:26.475such as between the scores and centrality scores, etcetera.
465
00:38:26.715 -->
00:38:28.155And I'm wondering what the
466
00:38:28.640 -->
00:38:36.560capabilities are around PuppyGraph for being able to do some of those more native graph feature extraction and discovery.
467
00:38:37.280 -->
00:38:46.714Yeah. I think this is something different from PuppyGraph to other graph solutions. Since most of the graph solutions, their query engine and the graph algorithm are implemented
468
00:38:47.275 -->
00:38:48.555separately. Like,
469
00:38:49.115 -->
00:38:55.435have a query engine, and also they have an independent implementation of a graph algorithm one by one.
470
00:38:55.835 -->
00:39:06.420But for us, we try to all leverage our engine, as I mentioned, the node operator and the add operator stuff. And in this case, when we implement a new graph algorithm,
471
00:39:06.500 -->
00:39:08.900we don't need to start from zero,
472
00:39:08.900 -->
00:39:14.745like how to parallel process the data or how to share the data. We just need to
473
00:39:15.225 -->
00:39:17.225implement an algorithm like decode
474
00:39:17.225 -->
00:39:18.025pretext.
475
00:39:18.025 -->
00:39:24.505And after that, we can deliver a new algorithm within one week, something like that. Some of the customers even
476
00:39:24.990 -->
00:39:25.790implement
477
00:39:25.790 -->
00:39:28.910their own graph algorithm by the query language.
478
00:39:29.070 -->
00:39:30.110Since, you know,
479
00:39:30.430 -->
00:39:32.430Grammarly is Turing complete.
480
00:39:32.430 -->
00:39:32.910So
481
00:39:33.790 -->
00:39:38.270before that, people don't do it just because if we implement through
482
00:39:38.510 -->
00:39:41.655query engine, it will be too slow. And like
483
00:39:41.655 -->
00:39:44.455GNS Graph, it's a single thread engine.
484
00:39:44.535 -->
00:39:54.855So if you implement on this, the algorithm will only run on single thread. So it's a potential issue for the performance. But for us, some of our customers just customize their algorithm
485
00:39:55.350 -->
00:40:00.310based on the query language, and that is another option. So we feel that. And also,
486
00:40:00.550 -->
00:40:11.474we won't charge additional for that. The people just charge our usage for the engine, whether it's algorithm or Cypher query or grammar query, we charge the same. We don't have additional,
487
00:40:11.474 -->
00:40:13.955like, enterprise feature charging for that.
488
00:40:14.675 -->
00:40:17.555And so for somebody who is interested
489
00:40:17.555 -->
00:40:18.195in
490
00:40:18.595 -->
00:40:19.315using
491
00:40:19.395 -->
00:40:20.115the
492
00:40:20.195 -->
00:40:21.155capabilities
493
00:40:21.155 -->
00:40:26.490of graph engines for doing some of that graph traversal discovery,
494
00:40:27.130 -->
00:40:28.650semantic capture,
495
00:40:28.650 -->
00:40:29.690and ontological
496
00:40:29.690 -->
00:40:30.730representation.
497
00:40:31.130 -->
00:40:38.415What are some of the guiding questions that you would ask them to help them determine whether PuppyGraph
498
00:40:38.415 -->
00:40:42.095or another solution is the appropriate
499
00:40:42.095 -->
00:40:42.815solution
500
00:40:43.055 -->
00:40:48.815or maybe even just say just use NetworkX and Python because it's a one off type of use case?
501
00:40:49.295 -->
00:40:50.335Usually,
502
00:40:50.495 -->
00:40:54.630if our customer already have data in data warehouse,
503
00:40:54.710 -->
00:40:56.710data lake, or even database,
504
00:40:56.869 -->
00:41:04.470we recommend you use a polygraph since it makes the pipeline much shorter and the system complexity
505
00:41:04.470 -->
00:41:05.350is much
506
00:41:05.589 -->
00:41:08.475lower. So in this case, they can just have,
507
00:41:08.954 -->
00:41:11.355for example, all the XBERG tables,
508
00:41:11.435 -->
00:41:15.755define graph on top of it, and then run the graph query and the graph algorithm.
509
00:41:16.155 -->
00:41:19.515Make the and also the results can write back to XBERG
510
00:41:19.515 -->
00:41:21.355and then leverage by Spark
511
00:41:21.570 -->
00:41:24.770or some other tools, and also even PyTorch,
512
00:41:24.770 -->
00:41:25.650this kind of stuff.
513
00:41:26.930 -->
00:41:33.170But some of them, for example, some of the data scientists, they don't care the data warehousing.
514
00:41:34.145 -->
00:41:35.745All the stuff already
515
00:41:35.984 -->
00:41:39.985They use Python all the way and all the stuff already in CSV file.
516
00:41:40.305 -->
00:41:40.865And if
517
00:41:41.585 -->
00:41:51.540We also support it, but it seems it is easier to use embedded one, like they can just use Python to read the CSV file like DuckDB
518
00:41:51.859 -->
00:41:56.100and then just run some network X on top of it. And
519
00:41:56.580 -->
00:42:00.980then we feel that the people are using Data Lake and Data Warehouse
520
00:42:01.565 -->
00:42:03.005like us because
521
00:42:03.565 -->
00:42:05.805the reason they use Data Lake and
522
00:42:06.205 -->
00:42:10.765Data Warehouse is because their data size is big, and they don't want to handle the
523
00:42:10.925 -->
00:42:11.725distribution
524
00:42:11.725 -->
00:42:13.645things. And so
525
00:42:13.965 -->
00:42:18.760we're in nature to be a good fit. But if people just use DuckDB
526
00:42:18.920 -->
00:42:33.695and have small data set, they can use DuckDB to handle the things, and it's embedded on Python, and all the stuff can be laptop at all. So we're trying to recommend our product to some of these users, but I
527
00:42:33.775 -->
00:42:36.095don't think we can convince them because,
528
00:42:36.095 -->
00:42:37.295frankly speaking,
529
00:42:37.615 -->
00:42:41.855just the Python ecosystem with DuckDB and also NetworkX
530
00:42:41.855 -->
00:42:45.290is better and more convenient than PuppyGraph.
531
00:42:45.850 -->
00:42:48.250One of the other interesting
532
00:42:48.250 -->
00:42:52.170aspects of where we are right now is the proliferation
533
00:42:52.170 -->
00:42:52.890of
534
00:42:53.130 -->
00:42:55.210vector embeddings for
535
00:42:55.290 -->
00:42:56.650some of these unstructured
536
00:42:57.035 -->
00:42:57.995sources.
537
00:42:58.155 -->
00:43:19.960And one of the patterns that I'm seeing is using a vector query to determine the starting point into a graph and then doing traversal from there. And I'm wondering how you're seeing people deal with that, particularly if they're using Iceberg as the underlying storage given that Iceberg doesn't really have native vector indexing capabilities.
538
00:43:20.760 -->
00:43:27.535Yeah. For ourself, we can query the Iceberg array as a vector. And also,
539
00:43:27.694 -->
00:43:30.734I hear some news from the Iceberg community.
540
00:43:30.974 -->
00:43:33.855They will support it very soon. And I think,
541
00:43:34.174 -->
00:43:40.520like, lessDB guys are also actively working with them. And hopefully, Iceberg can
542
00:43:40.760 -->
00:43:43.480have a vector type very soon.
543
00:43:43.560 -->
00:43:44.440But currently,
544
00:43:44.520 -->
00:43:47.720we just query the array type in Iceberg
545
00:43:47.800 -->
00:43:49.240as a vector.
546
00:43:49.400 -->
00:43:51.640And I think it works fine because
547
00:43:51.720 -->
00:43:59.725it's more like an index on top of a read type, and then we can run vector search on top of it.
548
00:44:00.285 -->
00:44:03.565And as you have been building PuppyGraph
549
00:44:03.565 -->
00:44:12.460and helping your customers get up to speed with it and understand its applications? What are some of the most interesting or innovative or unexpected ways that you've seen it used?
550
00:44:13.500 -->
00:44:16.060One for customer is Palo Alto Network.
551
00:44:16.220 -->
00:44:19.660Have several teams using our products. Some teams are using the
552
00:44:19.820 -->
00:44:24.125as a posture management, is a customer facing project. But the
553
00:44:24.205 -->
00:44:28.605interesting part is that the security research team, while they're doing that, they just
554
00:44:28.845 -->
00:44:30.765have all the logs in Iceberg
555
00:44:30.765 -->
00:44:35.645and use polygraph content to all their logs and look back, just to visualize all the data.
556
00:44:37.620 -->
00:44:38.420And then
557
00:44:38.740 -->
00:44:39.940they found some
558
00:44:40.740 -->
00:44:42.180botnet work. And
559
00:44:42.420 -->
00:44:44.820the author was arrested already.
560
00:44:44.900 -->
00:44:45.380And
561
00:44:46.260 -->
00:44:48.180people believe that the
562
00:44:48.180 -->
00:44:49.540malware was gone
563
00:44:49.955 -->
00:44:53.395and no one do the detection anymore. But
564
00:44:53.715 -->
00:44:56.755after they use public graph to look back at the logs,
565
00:44:57.155 -->
00:44:59.955there are still a lot of bot bot
566
00:44:59.955 -->
00:45:02.035network is attack
567
00:45:02.035 -->
00:45:05.235all the stuff. And so the bot network is still active,
568
00:45:06.090 -->
00:45:07.210even the
569
00:45:07.450 -->
00:45:08.890attacker was arrested.
570
00:45:08.890 -->
00:45:13.210So they feel surprised, but they also feel that this is very helpful.
571
00:45:13.690 -->
00:45:15.050And we also
572
00:45:15.370 -->
00:45:16.570didn't expect it.
573
00:45:17.895 -->
00:45:22.935Some the usage is more like a Splunk and is a more complex Splunk.
574
00:45:22.935 -->
00:45:27.175They can just use the public graph to to be the log reader
575
00:45:27.335 -->
00:45:29.495and then see what happened in the past.
576
00:45:30.630 -->
00:45:41.990And in your experience of building this product and platform and investing in this zero ETL capability for graph traversals and graph exploration,
577
00:45:41.990 -->
00:45:44.869what are some of the most interesting
578
00:45:43.545 -->
00:45:46.345unexpected or challenging lessons that you've learned in the process?
579
00:45:47.704 -->
00:45:53.945So we have some, like, one case is that at the beginning, we only support the Iceberg and later
580
00:45:54.025 -->
00:46:02.500Delta Lake and Hudi and Hive. But then some customers want us to support the database. We're trying to use the similar way, but
581
00:46:02.900 -->
00:46:11.315it's not a packet file reader, but projection with filters, like select attribute one from table A with some filter. But literally,
582
00:46:11.315 -->
00:46:18.035feedback from customer is that if we read too much from the transactional database, the QPS will be affected.
583
00:46:18.035 -->
00:46:27.380And then what we're doing is that we just do a cache layer and cache all the data from the database. But the lucky thing is that the database,
584
00:46:27.540 -->
00:46:41.395usually the data in database is not very big. So we just have a snapshot of the database data and then do the CDC, which is very different from our initial design. But I think with design partners and early customers,
585
00:46:41.555 -->
00:46:47.750their feedback is very, very important since what they care is not like if it's pure
586
00:46:47.750 -->
00:46:48.710in
587
00:46:48.869 -->
00:46:56.150this stage, it's just how we can fit into their production and how they can leverage the public graph technology.
588
00:46:56.230 -->
00:47:00.070So I think it's more like not just like we designed
589
00:47:00.245 -->
00:47:05.605and then everybody follow our pattern, but also after our early
590
00:47:05.925 -->
00:47:08.965adopter trial product, they provide feedback and
591
00:47:09.205 -->
00:47:10.405we're trying to
592
00:47:10.725 -->
00:47:12.085follow their request
593
00:47:12.549 -->
00:47:20.070and then have a different design for different use case. And we feel this is very valuable. And also like
594
00:47:20.390 -->
00:47:22.150the cybersecurity guys,
595
00:47:22.150 -->
00:47:26.070one and a half year ago, we are totally outsider of cybersecurity.
596
00:47:26.150 -->
00:47:28.645But after the different
597
00:47:28.885 -->
00:47:30.885leading cybersecurity company
598
00:47:30.964 -->
00:47:34.565reach out to us and they teach us how cybersecurity
599
00:47:34.565 -->
00:47:36.244industry can leverage
600
00:47:36.325 -->
00:47:37.285polygraph,
601
00:47:37.444 -->
00:47:40.805and they also let us to find like some other
602
00:47:41.460 -->
00:48:01.485company may use our product as well, and they gave us the names and let us to reach out to them. And their insight and the terminology, they're very helpful for us. Since just the engine is not that useful, but we can talk with the user a lot. We know what's their pinpoint and, how we can address their pinpoint.
603
00:48:02.525 -->
00:48:04.925And what are the cases where PuppyGraph
604
00:48:04.925 -->
00:48:16.160is the wrong choice and either the problem is just not a good fit for graph data generally or you'd be better served with a different graph engine.
605
00:48:17.280 -->
00:48:19.440Yeah. So really one
606
00:48:19.680 -->
00:48:23.725typical case is that, for example, you want to have some
607
00:48:24.445 -->
00:48:26.845personal AI memory storage.
608
00:48:27.085 -->
00:48:30.445And in this case, polygraph is not a good one since, really,
609
00:48:30.605 -->
00:48:33.645for personal AI memory, it's not big. And
610
00:48:34.045 -->
00:48:34.845embedded
611
00:48:34.845 -->
00:48:35.645solution
612
00:48:35.645 -->
00:48:46.480like maybe Kuso or some others is better. Like, you can just run on top of your run-in your laptop. And at the same time, you can support the transactional updates.
613
00:48:46.640 -->
00:48:47.120And
614
00:48:47.520 -->
00:48:52.015single data stack is good enough. And all staff can embed it in
615
00:48:52.415 -->
00:48:59.055Python program, for example. And in this case, I think it's a better solution. And also, we have a lot of similar case.
616
00:48:59.455 -->
00:49:00.015And
617
00:49:00.815 -->
00:49:26.765I think one good, when we do the judgment, I think whether the data size is very important. If the data size is small, I think the graph database is much better. Because you are writing data into it and you are reading data from it, And you don't need a data pipeline at all. In this case, I think, especially for the embedded graph database like Kudu, it's very good. And then you can have a you don't need to have a service right now. You just have,
618
00:49:27.085 -->
00:49:35.610for example, Python program and embedded Kudu in it, and then you can have all the functionality you need. So this is a one typical case.
619
00:49:36.010 -->
00:49:41.850And as you continue to build and iterate on PuppyGraph and its capabilities,
620
00:49:41.930 -->
00:49:45.305what are some of the areas of improvement
621
00:49:45.305 -->
00:49:50.984or new features or projects or problem areas that you're looking to dig into in the near to medium term?
622
00:49:51.785 -->
00:49:55.704Why is that? Definitely the enterprise features. Like, because
623
00:49:56.025 -->
00:50:02.710most of our customers are very big, and either our customers are big or our customers' customers are big. So
624
00:50:03.110 -->
00:50:09.670we are supporting the enterprise features, single sign on, rule based access, and all is in preview already.
625
00:50:09.830 -->
00:50:11.270And another thing is that
626
00:50:11.765 -->
00:50:19.205we want to have a better support of data warehouse. Same for data lake, because it's an open format and we can have access
627
00:50:19.285 -->
00:50:20.885to all the metadata
628
00:50:20.885 -->
00:50:22.085and table
629
00:50:22.165 -->
00:50:22.885stats
630
00:50:22.965 -->
00:50:32.430and all the related information, and then do the cost based optimization or some others. But for the warehouse, because some of them are pretty closed,
631
00:50:32.750 -->
00:50:39.710and so we need to collect all the information by ourselves and then have a better push down on the cost based optimization.
632
00:50:40.244 -->
00:50:43.605But another good information for us is that currently
633
00:50:43.605 -->
00:50:45.525the founders of
634
00:50:45.765 -->
00:50:49.204Parquet files and Apache Arrow are working on a
635
00:50:50.244 -->
00:50:51.765project called Columnar,
636
00:50:52.005 -->
00:50:52.565and
637
00:50:52.805 -->
00:50:54.325their project is ADPC.
638
00:50:54.650 -->
00:51:01.450And a lot of the warehouses are supporting ADPC now. And then we can read data through Apache Arrow.
639
00:51:01.770 -->
00:51:03.370You can see that is
640
00:51:03.610 -->
00:51:06.250much faster than just pure JDBC.
641
00:51:06.490 -->
00:51:08.010So I think because
642
00:51:08.010 -->
00:51:08.890ecosystem
643
00:51:08.890 -->
00:51:11.555evolve a lot and sometimes
644
00:51:11.635 -->
00:51:14.515we do need to implement what we need, we just wait.
645
00:51:14.595 -->
00:51:15.395The
646
00:51:15.395 -->
00:51:16.915feature we need is coming.
647
00:51:17.234 -->
00:51:17.715And
648
00:51:18.194 -->
00:51:18.915also,
649
00:51:19.154 -->
00:51:21.474culinary is our good partner.
650
00:51:21.635 -->
00:51:22.355And after
651
00:51:22.680 -->
00:51:31.400they share the project with us, we feel, oh, it's it's amazing. It's yeah. Chinese word is something like, when you're trying to sleep, you have a pillow.
652
00:51:33.160 -->
00:51:37.385Are there any other aspects of the work that you're doing on PuppyGraph
653
00:51:37.385 -->
00:51:40.745or this overall zero copy ETL
654
00:51:40.825 -->
00:51:45.945graph traversal capability that we didn't discuss yet that you'd like to cover before we close out the show?
655
00:51:46.985 -->
00:51:48.585I think these
656
00:51:48.309 -->
00:51:53.349covered all. You are super expert and you ask a lot of problem,
657
00:51:53.910 -->
00:52:00.630even we don't know before. After we engage with our customer users, they propose that. But definitely you have super
658
00:52:03.015 -->
00:52:07.415long users in this area. So I think it will cover all questions.
659
00:52:07.815 -->
00:52:25.280Thank you. And so for anybody who wants to get in touch with you and follow along with the work that you and your team are doing, I'll have you add your preferred contact information to the show notes. And as the final question, I'd like to get your perspective on what you see as being the biggest gap in the tooling or technology that's available for data and AI management today.
660
00:52:26.320 -->
00:52:30.320Yeah. I think we want to have a better agentical framework,
661
00:52:30.320 -->
00:52:37.255and currently, it's working, and we have a confidence to put it in production with a lot of customers already.
662
00:52:37.494 -->
00:52:38.055And
663
00:52:38.535 -->
00:52:40.855I think the improvement is more
664
00:52:40.855 -->
00:52:43.335easier to use. And also,
665
00:52:44.214 -->
00:52:45.255we want to
666
00:52:45.895 -->
00:52:50.510make it easier to connect with the fine tuning and reinforcement
667
00:52:50.670 -->
00:52:57.550learning things and the tools. And then we can have a better framework to embrace the ecosystem.
668
00:52:57.870 -->
00:53:14.405All right. Well, thank you very much for taking the time today to join me and share all the work that you're doing on PuppyGraph and the different use cases that it enables and some of the technological and architectural challenges of being able to act as that zero copy representation
669
00:53:14.405 -->
00:53:32.745on top of customers' underlying data. It's definitely a very interesting project and problem space, and I appreciate all of the work that you're doing to make graphs more available and accessible to a broader variety of use cases. So thank you again for that, and I hope you enjoy the rest of your day. Yeah. Thank you so much for the opportunity. Yeah. Have a good one.
670
00:53:43.830 -->
00:53:44.710Podcast.net
671
00:53:44.710 -->
00:53:53.830covers the Python language, its community, and the innovative ways it is being used. And the AI Engineering Podcast is your guide to the fast moving world of building AI systems.
672
00:53:54.310 -->
00:54:04.415Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email hosts@dataengineeringpodcast.com
673
00:54:04.415 -->
00:54:10.575with your story. Just to help other people find the show, please leave a review on Apple Podcasts and tell your friends and coworkers.