WEBVTT
NOTE
Transcription provided by Podhome.fm
Created: 08/02/2026
22:26:43Duration: 3734.519
Channels: 1
1
00:00:11.440 -->
00:00:15.440Hello, and welcome to the Data Engineering Podcast, the show about modern data management.
2
00:00:16.155 -->
00:00:25.994Your host is Tobias Macey. And today I'm interviewing Ragnor Comerford about OmniGraph, a lakehouse native graph storage layer with Git semantics. So Ragnar, can you start by introducing yourself?
3
00:00:26.314 -->
00:00:30.075Yeah, sure. Thanks, Tobias. Glad to be here. Yes, so I'm Ragnar.
4
00:00:30.450 -->
00:01:17.485I'm Dutch, but grew up mostly in France and Germany. And spent, I guess, most of my career, especially started my career really in kind of machine learning, computer science, because I guess the common thread was I was always obsessed with the idea of knowledge or information, or information theory as a whole. And then kind of kind of after sort of really working quite a bit in machine learning and computer science, I kind of dove quite a bit into the life sciences or biotechs. I spent quite a bit of time working in San Francisco at Edison's a company called Q Bio that built like a medical digital twin, and built lot of their internal data infrastructure and graphs, and then also spent quite a bit of time in protein design, essentially using a large language model trained on amino acids to design sort of new enzymes and proteins. So yeah, kind of the intersection of really computer science, ML and life sciences.
5
00:01:18.685 -->
00:01:21.885And do you remember how you first got started working in the data space?
6
00:01:22.525 -->
00:01:45.565Yeah, mean, as I mentioned quite early on, was I was really kind of interested in in sort of information theory and generally the power of like information and knowledge. So to me, you know, every kind of major outcome with, you know, both in research or in anything for like life science research to like hardware design was kind of downstream of how to use and operationalize information in a better way. And to me kind of that was also fundamentally an infrastructure problem.
7
00:01:46.205 -->
00:01:50.445And so quite early on when I was working in the life science, also realized
8
00:01:51.005 -->
00:02:06.870essentially how fragmented all the sort of knowledge information was when if you went into a lab, you had sort of a bunch of data and data set like on local computers, and you have essentially these huge biological databases like UniProt and, and PubChem and things like that, that contain
9
00:02:06.870 -->
00:02:08.550loads amounts of biological data.
10
00:02:09.675 -->
00:02:14.635And yeah, really kind of across this kind of fragmentation information and the kind of the subpar infrastructure.
11
00:02:14.795 -->
00:02:24.955And so I kind of really saw that as kind of the fundamental upstream problem, and got really interested in sort of building better data infra and systems to really make use of all this information.
12
00:02:26.580 -->
00:02:28.820And that brings us now to
13
00:02:28.980 -->
00:02:35.380Omnigraph. So I'm wondering if you can just start by giving a bit of an overview about what it is and how did it get started?
14
00:02:35.940 -->
00:02:43.745Yes. As you mentioned, so Omnigraph is a lakehouse native graph engine built on top of like the open data format lens.
15
00:02:44.225 -->
00:02:54.305And so they are kind of two lenses to look at it. One is actually working sort of backwards from essentially the new operator we were kind of optimizing for, and that is really the agent kind of workload.
16
00:02:54.939 -->
00:03:05.740And so like, obviously the upsets you're making now essentially is that going forward, most of the knowledge work is going to be performed by agents, or most likely also aged working in the background, not just kind of chat assistants.
17
00:03:06.140 -->
00:03:07.500And then working backwards
18
00:03:07.500 -->
00:03:14.115from essentially what the constraints that imposes on that kind of system. And I think fundamentally kind of split into kind of three different components.
19
00:03:14.595 -->
00:03:16.035One is like context,
20
00:03:16.035 -->
00:03:18.835the second one is coordination, third one is really governance.
21
00:03:19.235 -->
00:03:27.030So context is actually when an agent is supposed to perform a piece of knowledge work, know, how to sure that the agent actually has access to the right information with the right knowledge.
22
00:03:27.350 -->
00:03:40.755And then also for multiple agents means you'd actually have access to us like a shared world model, shared understanding essentially of the world. And the second one is coordination, okay, once you go from one, two to actually multiple agents, you know, running, how do you make sure that they all perform essentially
23
00:03:40.755 -->
00:03:54.570the correct global behavior rather than at the individual agent level. And then the third one is like, okay, when we have agents automating large parts of the knowledge work, you know, how to make sure that humans are still there where judgment their judgment matters,
24
00:03:54.650 -->
00:04:03.370and where they can provide essentially a signal to where things are going right or wrong. And this is kind of where also this kind of git start semantics that I missed earlier, and description
25
00:04:03.769 -->
00:04:04.569of Omnigraph
26
00:04:04.885 -->
00:04:13.845that became really important is like, okay, can you apply this sort of governance aspect of branching and merging also to the to the graph layer, to the structured data layer basically.
27
00:04:14.005 -->
00:04:22.840And so yeah, that was kind of really the the story behind it. How do you build like the optimal substrate for a multi agent system? But then at the same time, think we talked about this earlier,
28
00:04:23.560 -->
00:04:30.760like also looking at the existing ecosystem, obviously a lot of the technology or existing database that were designed for the human operator,
29
00:04:30.920 -->
00:04:40.245but also because of kind of historical path dependence of how we kind of build things. And so like the, the other sort of more bottom up reasoning or story behind Omnigrab is actually how would you build,
30
00:04:40.645 -->
00:04:44.405you know, a database essentially from first principles, so assembling
31
00:04:44.405 -->
00:05:04.425essentially all the building blocks of the amazing technology that has evolved, you know, fast forward to 2026, you know, and examples of these are like, you know, building a database on top of object storage, decoupling, you know, storage and compute, these open data formats, and then obviously the amazing query engines like DocDB or Data Fusion and things like that. So really sort of the
32
00:05:04.745 -->
00:05:07.865top down and bottom up sort of reasoning behind OmniGraph.
33
00:05:10.985 -->
00:05:15.145And so you mentioned that one of the core challenges
34
00:05:15.145 -->
00:05:19.500that you're focused on is this idea of multi agent coordination,
35
00:05:19.580 -->
00:05:24.220which is definitely the major concern of the day as the
36
00:05:25.100 -->
00:05:28.380number and capability of agents continues to expand
37
00:05:28.620 -->
00:05:37.235and the patterns are definitely still very much in flux. I'm actually just starting to read the agentic mesh book from O'Reilly, so good timing here.
38
00:05:37.555 -->
00:05:38.195And
39
00:05:38.995 -->
00:05:39.795graphs
40
00:05:39.795 -->
00:05:46.900have gained a lot of attention over the past year or two because of the fact that connectivity
41
00:05:46.900 -->
00:05:48.900is one of the key considerations
42
00:05:48.900 -->
00:05:49.540for
43
00:05:50.020 -->
00:05:57.060any intelligent system because the relationship between things is one of the key things that influences
44
00:05:57.060 -->
00:05:57.699the
45
00:05:58.020 -->
00:06:02.555types of decisions you might make. It adds nuance to the raw information,
46
00:06:03.115 -->
00:06:13.595and there are a number of graph engines that have been available for a number of years. I think d o four j is probably the, at least, most widely known one, if not necessarily the most widely deployed.
47
00:06:13.835 -->
00:06:17.410And I'm wondering if you can just talk to the core
48
00:06:17.730 -->
00:06:19.330foundational aspects
49
00:06:19.330 -->
00:06:19.970of
50
00:06:20.210 -->
00:06:29.410graphs and some of the graph engines that are in the market and what was lacking that led you to decide that you needed to create a another new one.
51
00:06:30.135 -->
00:06:37.255Yeah, no, sure. I think maybe taking a step back, think also I see in general, there's a problem with essentially the taxonomy
52
00:06:37.255 -->
00:06:52.400of databases, because obviously they were often formed for essentially forming like product categories or positioning basically, But I think we're from first principles, if we think about a database, I think we have to get a separate into like, along three axes, and one is essentially the
53
00:06:52.720 -->
00:07:10.395logical data model or essentially the semantics, right. And it means like the query language, like how you represent the data and things like that. Then there's the workload that is designed for, right, for example, let's say in OLTP OLAB, or in the case of time series for like heavy append only data. And the third is actually the physical manifestation or the physical data model that essentially
54
00:07:11.020 -->
00:07:30.245aims at the kind of solving the constraints imposed by essentially the logical data model and the workload. So, and then independent of that, while one can essentially design around these things, and then the problems obviously that graphs or graph engines were kind of put into this bucket of like, okay, it comes with, you know, oh, it's optimized just for large scale sort of, you know, edge traversal.
55
00:07:30.325 -->
00:07:52.290And it's also meant to be kind of slow and not scalable, essentially, right? But to me, essentially, graphs was always more about the semantics that it describes is like, how do you represent your data model? And also what query language and then essentially how you build it in a physical layer, right? It can evolve or kind of be optimized for different purposes. I think also because in the past, the database were built sort of in a very monolithic way, know, owning both the storage format, the indexes,
56
00:07:52.715 -->
00:08:03.995you know, the semantics, query language, but now we're seeing actually, you know, decompose ability being sort of the much better paradigm, right, and decoupling search and compute and learning from mistakes, you know, that we've made sort of in the past.
57
00:08:04.810 -->
00:08:09.130And so and the graph comes in as essentially at the semantics side,
58
00:08:09.130 -->
00:08:38.089essentially, right, in that the information in the world essentially exists also as, you know, objects that relate to each other, right? And this is kind of fundamentally how the world exists, and this is where sort of the ontology comes in, and this is why essentially the graph is kind of the correct representation of that world, because it's actually also in some sense the most generalized representation, because you can sort of represent any other things from a two dimensional array to, you know, like tables as a graph itself basically, right? So then it's more about, okay, you know, what workload
59
00:08:38.089 -->
00:08:41.690or query patterns is it kind of optimized for the kind of underlying storage layer.
60
00:08:42.345 -->
00:08:48.584But yeah, but to be more specific, if we had to kind of compare to existing graph engines, obviously there is, you know, Neo4j,
61
00:08:48.584 -->
00:08:52.265there's TigerGraph and so forth. But essentially the
62
00:08:52.745 -->
00:09:01.040biggest differentiator is one on the lakehouse side that we really built this on object storage and decoupling like storage compute and building also on open data formats like
63
00:09:01.360 -->
00:09:05.040our case, Lance, but you might know your Iceberg and sort of Delta Lake.
64
00:09:05.360 -->
00:09:12.640That means you can actually eventually also plug in any other query engine because in the end it's just files on object storage, So there's essentially not this kind of lock in.
65
00:09:13.295 -->
00:09:19.454And at the same time, also we optimize also for the Git style workflows, because if agents become the new operators,
66
00:09:19.535 -->
00:09:20.654they're fundamentally,
67
00:09:20.654 -->
00:09:32.770you know, probabilistic writers, whereas in the past, you know, databases were made for trusted writers, but we have a software that, you know, is programmed to kind of write data and it has all the essentially validation logic, you know, beforehand.
68
00:09:33.170 -->
00:09:37.170And agents are really probabilistic writers, so this kind of Gitstar paradigm
69
00:09:37.170 -->
00:09:59.150becomes super important. But I would have said even for humans, it would have been because since collaborating on shared data requires some kind of governance mechanism, coordination mechanism, and this is the exact one that, you know, Git or GitHub solved, essentially how do you allow, you know, hundreds of developers to collaborate on a shared, you know, artifact that needs to remain consistent at all times. So yeah.
70
00:10:00.270 -->
00:10:08.350I think the two engines that I'm drawing the most parallels with right now for what you're building with OmniGraph are PuppyGraph
71
00:10:08.350 -->
00:10:25.190on the lake house graph traversal side, and then the Git semantic side is the folks that are building the DOLT engine, which is a MySQL and now a Postgres compatible database engine that has native Git semantics all the way down to the storage layer. And I'm wondering,
72
00:10:25.670 -->
00:10:26.870what are some of the
73
00:10:27.110 -->
00:10:34.070engines or systems that you drew inspiration from as you were coming up with the initial ideas and architectural
74
00:10:34.070 -->
00:10:35.590layout for Omnigraph?
75
00:10:36.245 -->
00:10:40.485Yeah, no, I mean, we are we took a lot of inspiration from many kind of different systems,
76
00:10:40.644 -->
00:10:48.084including also around the kind of deployment model. Yeah, I can actually dive into that straight away. But like thinking again about the agent as kind of an operator,
77
00:10:48.165 -->
00:11:01.620I think became also really clear that essentially having a declarative system became would be super important, because essentially, in some sense, you want the agents to be able to inspect the whole state of the system, without too much like searching around.
78
00:11:02.019 -->
00:11:04.899And at the same time, also you want to kind of reduce
79
00:11:04.775 -->
00:11:23.550decision surface area for the agents of getting things wrong, right? And so this comes in those like the kind of query layer, right, where essentially you want your query language to be optimized for agents in the sense that you would post kind of guardrails on them. So one of these is essentially this schema type sort of query language, so you can actually validate the schema against sort of the query language.
80
00:11:23.710 -->
00:11:40.415And for, you know, humans have been annoyed because every time they get an error and they kind of construct their dataset, for an age they can just recover from it in the next sort of iteration of the loop essentially. But it's associated with declarative paradigm also kind of extends also to the actual whole operation of the system in that anything from the ontology
81
00:11:40.415 -->
00:11:42.495to the queries we define,
82
00:11:42.735 -->
00:11:45.695all kind of defining declarative race in order to kind of terraform,
83
00:11:46.070 -->
00:11:48.150to to be kind of all inspectable.
84
00:11:48.230 -->
00:11:50.390Digging into the query piece,
85
00:11:50.950 -->
00:11:55.270many of the graph engines that are out there have consolidated around
86
00:11:55.270 -->
00:11:58.230either one or both of Cypher,
87
00:11:58.310 -->
00:12:03.955which originated with Neo four j and has become effectively the stand in for the
88
00:12:04.275 -->
00:12:07.715I think it just got ratified, but the, GQL
89
00:12:08.035 -->
00:12:09.395ISO standard,
90
00:12:09.395 -->
00:12:16.035and then Gremlin is the other query language, and then, of course, Sparkle and the semantic web space. And I'm wondering
91
00:12:16.280 -->
00:12:17.880as you were developing
92
00:12:17.880 -->
00:12:23.080Omnigraph and thinking about the use cases and the useful constraints,
93
00:12:23.560 -->
00:12:30.040what are the pieces that you found to be most beneficial by having that more constrained typed precompiled
94
00:12:30.040 -->
00:12:31.080query interface?
95
00:12:31.904 -->
00:12:49.640Yeah, I guess when deciding whether to go with sort of an existing language or not, one had to kind of think of essentially what would be the advantage of what's actually advantage if a language is actually used by large amounts of people. And I think in the past, it was fundamentally about around the expertise or the understanding of people to write these queries,
96
00:12:49.959 -->
00:12:54.040But then if we thought about essentially Asians as a fundamental new operator, this
97
00:12:54.200 -->
00:12:56.520problem kind of, you know, vanished really.
98
00:12:57.000 -->
00:13:04.895And so and one could argue essentially that obviously the most have been kind of pre trained on like Cyphon this language, they should actually technically be better at writing these queries.
99
00:13:05.215 -->
00:13:16.570But actually, it turns out that essentially, most of the errors in actually constructing the queries are actually because of misunderstanding of the schema or types that exist rather than the underlying sort of, sort of syntax of the query language.
100
00:13:16.730 -->
00:13:33.955The LLMs are really good at sort of the in context learning, passing them essentially the syntax gives them the readability to construct like the query very easily. And so like the actually schema typing became the most important one, because if agents start writing knowledge back, essentially having this kind of strong validation
101
00:13:34.035 -->
00:13:35.555makes you the system really
102
00:13:36.035 -->
00:13:50.690makes it much easier to keep the system that consistent over time. And it's actually the beauty then also of the ontology, because the ontology in some sense, it actually imposes a kind of a reasoning process on the agent, Because if essentially you are forced to, when you insert knowledge to kind of
103
00:13:51.264 -->
00:13:54.144idea to the ontology, then you have to kind of follow the reasoning process,
104
00:13:54.625 -->
00:14:13.340process that the anthology kind of enforces, right? So for example, let's say you are analyzing transcripts of customer discovery calls, And now you want to essentially use a specific framework to extract information from it, like the jobs to be done framework, for example, right? Then you can include that directly in the ontology and make it typed so that it says like required
105
00:14:13.420 -->
00:14:22.954fields and required edges and things like that. And so essentially when the agents want to write knowledge and essentially gets an error because it's missing a piece of information, it forces them to go back, fetch that information,
106
00:14:23.195 -->
00:14:34.020and so really kind of imposes like the framework on the agent itself. So I think this was kind of one of the major sort of benefits of doing that. And then the second one is also kind of optimizing for
107
00:14:34.100 -->
00:14:35.220essentially the new
108
00:14:35.540 -->
00:14:41.540workload that, that Asia optimized for, and ultimately on the database side is about assembling the right piece of context.
109
00:14:42.180 -->
00:14:58.035And so that, that means that now actually writing queries is more of a sort of information retrieval task and including like search and multiple, multiple sort of modalities, like combining graph traversal, where you might look at an entity, follow its relationships, and combine that with vector or similarity
110
00:14:58.195 -->
00:15:05.160search and then full text search. So really kind of optimizing the query semantics and language for like context assembly,
111
00:15:05.480 -->
00:15:11.240rather than the past kind of the reason the kind of design constraints behind like the different languages itself.
112
00:15:11.480 -->
00:15:19.915Yeah, and then maybe the last bit is obviously decreasing essentially the decision surface area for agents when they actually write the query. Right? So for example, in Gremlin,
113
00:15:19.915 -->
00:15:27.195have sort of imperative paradigm. So it is kind of reason about sort of, you know, how to construct the query and execution itself.
114
00:15:27.435 -->
00:15:49.935Whereas like in a declarative manner, you kind of leave it to the engine that you decide how to kind of optimize that queries. So essentially, agents are also subject to that dilution of their attention, right? So the more context you pass it, the more decision they have to make at the same time that the worse they perform. So can you actually reduce the decision service area as much as possible for that sick task, like for example, retrieving a piece of information, and I think really optimizing for that essentially.
115
00:15:52.815 -->
00:15:53.295Now,
116
00:15:54.334 -->
00:15:59.900because of the fact that you have optimized for agents to be the predominant
117
00:16:00.380 -->
00:16:04.780consumer and producer for the data within Omnigraph,
118
00:16:04.780 -->
00:16:13.355how does that change some of the other potential applications that somebody might reach to it for, particularly in the case of graph
119
00:16:13.835 -->
00:16:19.035applications. So some of the bigger ones are things like knowledge graphs, fraud detection,
120
00:16:19.275 -->
00:16:26.260things like security information and event management for being able to correlate events. Obviously, many of those applications
121
00:16:26.660 -->
00:16:43.925can and do have agentic corollaries and agentic input as well as people are developing and evolving these systems. But how does it change the calculus of somebody who is looking gender generically for a graph engine? And what are the cases where they might turn away from OmniGraph because of that more
122
00:16:44.245 -->
00:16:45.764strict and constrained
123
00:16:45.764 -->
00:16:47.365query layer?
124
00:16:48.165 -->
00:17:11.235Yeah. I mean, that that's a good question. Because, obviously, whenever you kind of design for the operator, you kind of have also trade offs, right? So you can sort of make a system work for every type of, know, usage or workload. So in this case, yeah, we were not necessarily optimizing for the kind of traditional uses of like, you know, graphs, like as you as you mentioned, kind of in fraud detection, or, or these kind of type of queries, but really kind of an optimal substrate
125
00:17:11.395 -->
00:17:18.195for agents. So I think there's, there's already kind of really good solutions out there as well for these type of kind of workloads.
126
00:17:18.275 -->
00:17:26.289But in the end, yeah, we're really focused on, on kind of, especially the semantics of the kind of shared world model for agents.
127
00:17:26.530 -->
00:17:31.409And but yeah, at the same time, still like we are one has to kind of rethink about sort of the
128
00:17:31.730 -->
00:17:52.299constraints that one designs for it, because existing database engines like Postgres or LTP were designed for high concurrency, you know, rights that are all trusted rights. And now what does that, for example, that concurrency pattern look like now for agents? So for example, we think that the concurrency patterns are going be more at the branch level, because agents are going be performing like long running work independently,
129
00:17:52.300 -->
00:17:56.860and then it's going to be reviewed, you know, synthetically and semantically and then merged into the truth.
130
00:17:57.020 -->
00:17:59.500There's going be less of a need for high concurrency
131
00:17:59.500 -->
00:18:05.404rights, you know, that they might conflict with each other. But at the same time, you still need kind of transactional semantics,
132
00:18:05.404 -->
00:18:06.684because ultimately
133
00:18:06.684 -->
00:18:09.245it's also operational state that other agents
134
00:18:09.245 -->
00:18:11.325rely on for coordination.
135
00:18:11.644 -->
00:18:21.409And then similarly also like, it kind of merges some of the kind of analytical works in there as well, because partially some of the context assembly will involve like doing sort of more analytical queries.
136
00:18:21.410 -->
00:18:41.735And also we look at sort of the evolution of the markets like Databricks, also have this kind of convergence now of you know, OLTP and OLAP. But yeah, I mean, to answer your question, so we're not necessarily optimizing for the classical uses of graphs, essentially, because for us, it was mostly also about the semantics and not just essentially these kind of traversal workloads, although obviously,
137
00:18:42.230 -->
00:18:48.710we have kind of, you know, optimized indexes for graph traversal and things like that, but it's not kind of the primary use case.
138
00:18:49.190 -->
00:18:53.430So digging now into the architecture of OmniGraph,
139
00:18:53.430 -->
00:18:55.750we've touched a little bit on some of the pieces,
140
00:18:55.830 -->
00:19:02.065predominantly the fact that it is using Lance as the storage layer. I'm wondering if you can describe some more detail
141
00:19:02.065 -->
00:19:05.184of how it's built, how it's deployed,
142
00:19:05.345 -->
00:19:08.065and what the adoption and setup looks like.
143
00:19:08.970 -->
00:19:14.809Yeah, we kind of built on top of Lens, because Lens already kind of handled a lot of the lower level primitives
144
00:19:14.809 -->
00:19:16.570and really well engineered.
145
00:19:16.650 -->
00:19:30.945So those kind of don't know Lens. So that's a new new ish table format that's primarily optimized for three things. One was the fast random access, so they would kind of, you know, retrieve essentially low latency like individual
146
00:19:31.105 -->
00:19:38.470rows. And then second also kind of the multimodal modality of it to be able to actually blobs directly in the file format.
147
00:19:39.030 -->
00:19:50.575And then third was like cheap schema evolution, so that actually adding a column, right, doesn't require like rewriting this sort of whole data set. But apart, I mean, these are kind of the kind of major components that they kind of advertise,
148
00:19:50.575 -->
00:19:58.334but I mean, apart from that, I think they've done just a lot of amazing engineering on the format itself, right, handling things like obviously asset transactions,
149
00:19:58.575 -->
00:19:59.774also indexes,
150
00:20:00.174 -->
00:20:06.040and yeah, a lot of really good sort of primitives under the hood, and also really designed kind of like in a composable
151
00:20:06.040 -->
00:20:23.845way, and so we kind of treat that as kind of the L1 or the kind of the layer one, and we kind of build the layer two on top of that, and so if you think about, you know, as a team you have kind of a fixed engineering budget, and then where do you spend it on, and we kind of really spend on that layer rather than reinventing, let's say, a new table format,
152
00:20:23.925 -->
00:20:32.780and really we're focusing on essentially all the economics around like the graph, like the typing, the graph semantics, you know, the branching and merging mechanism,
153
00:20:32.780 -->
00:20:50.634obviously all the kind of graph traversal and the kind of deployment around it. But yeah, Lansen and also in sort of including also Data Fusion, who have as a kind of query layer or query execution engine, gave us a lot of like really good sort of building blocks to kind of build a really reliable sort of graph engine on top of that.
154
00:20:51.034 -->
00:20:52.554One of the interesting
155
00:20:52.875 -->
00:21:00.809pieces that I've seen from a number of the more recent graph engines is this adoption of columnar storage
156
00:21:00.890 -->
00:21:01.529and
157
00:21:01.850 -->
00:21:03.690tabular representation
158
00:21:03.769 -->
00:21:05.609at the file level.
159
00:21:05.850 -->
00:21:28.220And I'm curious particularly with the Data Fusion since that's largely a compiler for SQL in the the general case that I've seen it used. What are some of the aspects of impedance mismatch that you've had to work around between this graph as the core data model for what Omnigraph is aiming for versus this tabular representation
160
00:21:28.220 -->
00:21:29.660of the data on disk?
161
00:21:30.220 -->
00:21:40.614I mean, yeah, one of it obviously that we had to kind of design some of our like, custom like indexes as well, because I think a large part of kind of optimizing the queries actually was around the, you know, having the right indexes.
162
00:21:40.615 -->
00:21:41.815And so for example,
163
00:21:42.215 -->
00:21:46.534you would kind of build our own sort of adjacency matrix, basically, or CSR,
164
00:21:46.930 -->
00:21:52.210essentially for the query execution layer. But yeah, lot of kind of the sort of
165
00:21:52.770 -->
00:21:55.250query or traversal mechanism actually were
166
00:21:55.570 -->
00:22:05.495handled pretty well, by being able to push down essentially predicted down to Data Fusion or actually Lens itself, because Lens also itself handles a lot of the kind of the query logic behind the scenes.
167
00:22:05.815 -->
00:22:06.295But
168
00:22:06.615 -->
00:22:24.250yeah, so for now, but basically for us, most of the constraints are around like not having essentially primitives for graph traversal where we had to kind of build the index there. And also we're still making like lots of improvements on that side of things. But so far, yeah, the machinery kind of exposed by Lens already provide us like a lot of help essentially there.
169
00:22:24.650 -->
00:22:26.250One of the other
170
00:22:26.730 -->
00:22:33.225ways that using that columnar tabular storage instead of a graph native representation
171
00:22:33.225 -->
00:22:42.344on disk can impact the overall capabilities of a graph engine. And one of the reasons that graph engines are so difficult to scale horizontally
172
00:22:42.664 -->
00:22:43.384is the
173
00:22:44.290 -->
00:22:47.730number of degrees of removal that you're able to traverse
174
00:22:47.730 -->
00:23:00.835in a single query. And the fact that the impact of supernodes on a graph is one of the reasons that horizontal scaling of traditional graph engines has been so problematic over the years. And I'm curious what are some of the
175
00:23:01.155 -->
00:23:04.835design choices and intentional limitations that you're
176
00:23:04.995 -->
00:23:08.595accepting by virtue of not having a a graph native
177
00:23:08.835 -->
00:23:10.195storage representation
178
00:23:10.195 -->
00:23:12.995and instead relying on the ecosystem
179
00:23:12.995 -->
00:23:26.680advantages and the architectural advantages of Lance as that foundational layer, or some of the potential scaling problems that you hit if you're trying to do something like a 30 degree hop across multiple different interconnected nodes.
180
00:23:27.325 -->
00:23:38.125Yeah, so so far, we haven't really kind of optimized as much for these sort of very graph analytical queries, right? Because, as I mentioned, we were especially optimizing for the graph semantics or the representation
181
00:23:38.605 -->
00:23:45.110of the data in that sense. I And think a lot of the queries also are, can be essentially represented then as like, you know, really
182
00:23:45.590 -->
00:23:51.030well in terms of like SQL queries, or kind of these predicates sort of push downs, and also doing similarly
183
00:23:51.030 -->
00:23:52.950for vector search or full text search.
184
00:23:53.605 -->
00:23:57.364So yeah, so far we haven't really optimized for these sort of graph analytical
185
00:23:57.445 -->
00:24:14.860queries, but you know, we might essentially at some point, you know, eventually even create our own format or, or do like, like further optimization. But yeah, this wasn't currently like much of a design constraint at this stage, because again, we were kind of not optimizing for these sort of really graph analytical queries basically.
186
00:24:15.820 -->
00:24:17.820And as you have been
187
00:24:17.980 -->
00:24:21.660working through the early stages of building OmniGraph,
188
00:24:22.045 -->
00:24:38.719discovering some of the design space around it, understanding how people are using it, how it all fits together, what are some of the ways that your early assumptions have had to shift and some of the evolutions of the design and goals of the project have grown as you've gone through this process?
189
00:24:39.120 -->
00:24:39.919I mean,
190
00:24:40.240 -->
00:24:57.945less so in terms of kind of the design, but obviously, we were kind of aware that, for example, object storage give you sort of very, like, quite high latency, basically, right? And so initially, a lot of these the system kind of worked really well on fast on like a local file system, but then essentially when deploying it on object storage, so that essentially
191
00:24:57.945 -->
00:25:16.330using it essentially over like multiple weeks and accumulating low version history and things like that, we kind of obviously hit a lot of the end of the limits around both the latency imposed by object storage reads. So this is kind of on rise as well, and so this is kind of where we had to do a lot of optimizations of like minimizing the number of reads and writes and do some of the redesigns,
192
00:25:16.410 -->
00:25:20.650you know, around also like caching and the over kind of write and read patterns.
193
00:25:21.145 -->
00:25:26.105Of course, it's kind of obvious beforehand that, you know, object storage has this kind of latency constraints.
194
00:25:26.345 -->
00:25:31.304But of course, like as you kind of, you know, as the design kind of evolves and the workloads
195
00:25:31.625 -->
00:25:41.250increase, I think this is kind of where it was really hitting the limits and had to kind of redesign some of the mechanisms around that. And so yeah, that was probably more of a kind of technical
196
00:25:41.490 -->
00:25:42.290requirement,
197
00:25:42.290 -->
00:25:47.490basically. But so far in terms of kind of the, the design choice around the operator,
198
00:25:48.245 -->
00:26:06.799we we haven't kind of had we haven't had to do like too many changes yet, because I think a lot of the insights, you know, that went into Omnigravus who came from like, you know, many years working, you know, across like different data stacks and those having worked like with with agents and what they need and things like that. So, yeah, we'll see going forward if if things sort of change.
199
00:26:07.200 -->
00:26:07.840And then
200
00:26:08.240 -->
00:26:12.480because you are building on these other established projects
201
00:26:12.480 -->
00:26:15.600of Lance and Data Fusion and Arrow,
202
00:26:15.965 -->
00:26:22.924what are some of the ecosystem advantages that you gain from it as far as other integration points or
203
00:26:23.005 -->
00:26:24.284various other
204
00:26:24.445 -->
00:26:27.965interaction patterns that come for free because of that selection?
205
00:26:29.340 -->
00:26:55.375Yeah, I mean, I mean, first of all, so because it was sort of built in a in a composable way, you kind of have really this ecosystem advantage of be able to kind of plug in different parts in and out. For example, obviously, having open data formats means you use it with other query engines, also leverage different types of storage medium from S3 and, and SEF to like, be able to run it in an embedded format as well. But then you also have the whole kind of error ecosystem,
206
00:26:55.375 -->
00:27:09.110because you know, Lenses are based on error and then all the interoperability there with other libraries around zero copy with your pandas, polars, that DB and things like that. And then last as well, so we we kind of have made embeddings
207
00:27:09.110 -->
00:27:15.144also pluggable in terms of providers, it can sort of interact with, or you can use like any model on top of it.
208
00:27:15.625 -->
00:27:21.065And then lastly, we kind of provide also this sort of the data layer or the data substrate,
209
00:27:21.145 -->
00:27:30.120but we don't kind of impose essentially what it looks like on the agent orchestration layer. So we want to kind of design in a monolithic way, of defining also how agents are supposed to be running.
210
00:27:30.520 -->
00:27:35.160So in that sense also it kind of also as we see like the actual agent orchestration
211
00:27:35.160 -->
00:27:37.800libraries are kind of changing all the time and evolving,
212
00:27:37.960 -->
00:27:49.035and so be able to to kind of work with them in a sort of like composable way, think is really important, because you know, lot of companies or products out there kind of give you everything from the data layer to the agents, the models.
213
00:27:49.515 -->
00:27:55.435And so what it also means is that it pairs really well with, you know, I think companies or organizations that want
214
00:27:55.660 -->
00:28:05.259essentially a sovereign AI stack, be able to actually own essentially the AI stack from like using local and open source models to maybe even their custom hardware.
215
00:28:05.580 -->
00:28:10.780And so this kind of really plays well with sort of the sovereignty or local first kind of community there.
216
00:28:11.905 -->
00:28:19.424One of the other interesting pieces is, as you mentioned, Iceberg has a huge footprint in the data ecosystem.
217
00:28:19.425 -->
00:28:33.179There's a lot of information in Iceberg tables. And I know that Lance has at least a one way door of being able to migrate iceberg tables into Lance tables. And I'm curious if there are any thoughts that you have on being able to
218
00:28:33.660 -->
00:28:36.940introspect and generate a graph on preexisting
219
00:28:36.940 -->
00:28:37.580data
220
00:28:38.075 -->
00:28:41.914versus only being able to start from a net new
221
00:28:42.155 -->
00:28:45.355state storage and write data into the graph?
222
00:28:46.395 -->
00:28:55.780So yeah, in terms of lens, for example, we're definitely able to do that to kind of use like existing lens data as well, because also the nodes and edges are actually represented
223
00:28:55.780 -->
00:28:56.820as individual
224
00:28:56.980 -->
00:28:58.980lens files at several lens tables.
225
00:28:59.460 -->
00:29:04.340And yeah, Iceberg is definitely sort of the most, you know, widely used format out there. So,
226
00:29:05.085 -->
00:29:53.570so yeah, be able to kind of work with these, these existing data sets is quite important. But and LENS provides kind of like an interoperability that to be able to kind of migrate from Iceberg to LENS. But as the currency stands, we don't necessarily support like be able to query existing sort of iceberg data, because essentially, we designed really OmniGraph, not just as a kind of query layer or analytical engine, but but really use as an operational transaction database that for writing as well. And so and for the right side, we rely on lot of the lens sort of machinery right around, you know, as a transaction, things like that. So to the right side wouldn't necessarily work, but we have got quite a few requests as well to be able to kind of query other, you know, formats. So there's definitely something that we kind of have on the roadmap, but haven't optimized for that yet. Because really, we're focusing on really building sort of an operational
227
00:29:53.650 -->
00:30:11.815state layer rather than just kind of query layer, because I think there's also a lot, there's a lot of players already sort of under, you know, adding maybe a semantic layer on top of the existing data sets, maybe to query those, whereas our focus is really kind of on that, you know, shared substrate for for for like long horizon, like multi agent systems.
228
00:30:39.520 -->
00:30:40.320Yeah,
229
00:30:40.400 -->
00:31:15.200I mean, one of the kind of primary use cases emerging right now as companies are kind of understanding that, you know, the fragmentation of their context and also the lack of essentially an explicit, you know, ontology or world model is kind of the bottleneck to really be able to automate work, has meant essentially that this is the idea of this company group brain or the context graph has become like really popular. And so this is kind of one of the primary kind of use cases we're seeing where they would use like, they use Omnigrap essentially to represent their sort of shared ontology well off the organizations, and then using that essentially as a shared substrate for agents to to get you perform, you know, background work.
230
00:31:15.519 -->
00:31:53.115And at the same time, there's essentially different instantiations of that as well of like what the company brand means. And then, you know, one of kind of the two primary domains are on one of the kind of rev rev ops, like GTM side, essentially of building sort of, you know, more elaborate than a CRM, but something actually more custom in terms of like how they do kind of the whole revenue process around like prospecting and enriching, essentially building sort of that shared sort of world model of their sort of sales organization, and then running agents to enrich prospects to, you know, running all the process around that. And the second one also being now is like agents are now performing most of the coding work, also building
231
00:31:53.515 -->
00:31:54.955essentially graphs
232
00:31:54.955 -->
00:32:07.059or ontologies around the SDLC or the software development life cycle, Anything from, as we were talking about, like indexing a code base to kind of high level actually representing some of the things like projects, tasks, dependencies,
233
00:32:07.059 -->
00:32:12.820assumptions, but also, you know, different code bases and their dependencies between each other, and be able to see essentially
234
00:32:13.144 -->
00:32:44.844how changes propagate, might affect other systems, and the blast radius and things like that. Because I strongly believe that actually going forward, a lot of the energetic engineering is going to be performed essentially by having this kind of shared world model of essentially the whole engineering organization and all entities around it. And, and both like Simpson in production and existing code bases, and then having agents kind of work in loops around that sort of shared world brand. But yeah, I think there's also other interesting use cases you come across around also like, you know, in my background, like around like research,
235
00:32:45.164 -->
00:33:02.380R and D, or in life sciences and things like that, where you have also a large amount of entities that you know, agents need to be tracking and enriching information for, know, from proteins, molecules, etc. And there was some people using it even for essentially like a training like ML models, and they're of using the graph, you can have more as a kind of indexing
236
00:33:02.380 -->
00:33:14.065layer, then you can actually train directly on the lens files itself, because, you know, without essentially having to obviously, query the data, but essentially kind of be able to work directly on top of sort of the the lens files.
237
00:33:14.385 -->
00:33:14.945One
238
00:33:15.185 -->
00:33:16.065of the other
239
00:33:16.385 -->
00:33:17.265interesting
240
00:33:17.585 -->
00:33:18.385potential
241
00:33:18.865 -->
00:33:26.429free benefits that you might get as well is because you have the storage layer and object storage, which is generally
242
00:33:26.430 -->
00:33:30.669fairly inert. You don't have to have a lot of compute running for that to operate.
243
00:33:31.070 -->
00:33:46.205And the fact that Omnigraph itself is written in Rust, which has excellent support for WebAssembly. I'm curious if you've done any experimentation for being able to move the actual query interface into a browser and be able to ship an effectively serverless
244
00:33:46.549 -->
00:33:47.909graph introspection
245
00:33:47.909 -->
00:33:52.710to a web browser using open data in a lance store somewhere.
246
00:33:54.390 -->
00:33:59.190Yeah, no, that's a super interesting idea. And definitely kind of has been, you know, somewhere,
247
00:33:59.190 -->
00:34:01.030you know, on our roadmap,
248
00:34:01.030 -->
00:34:05.044you know, but haven't prioritized this yet. But I think that's definitely like a,
249
00:34:05.285 -->
00:34:08.485I think it's an interesting idea. And definitely, we'll explore that.
250
00:34:09.525 -->
00:34:10.085And
251
00:34:10.485 -->
00:34:11.525digging into
252
00:34:11.605 -->
00:34:18.820what you were discussing, as far as modeling some of the software development cycle and some of the agentic collaboration,
253
00:34:18.900 -->
00:34:22.580one of the projects that I'm actually building using OmniGraph
254
00:34:22.580 -->
00:34:27.140is this tool called Witten that I'm building internally for my day job
255
00:34:27.380 -->
00:34:27.940and
256
00:34:28.464 -->
00:34:37.665using it to track things like the agentic memory so that rather than it just being a markdown file that sits on my disk somewhere and nobody else ever sees,
257
00:34:37.825 -->
00:34:45.160it is stored as a node in the graph. It has the ability to be superseded by newer memories. So you have some evolution
258
00:34:45.160 -->
00:34:45.800of
259
00:34:46.520 -->
00:34:48.600some of that organizational knowledge.
260
00:34:48.600 -->
00:35:00.040It has the ability to track projects that have tasks and their dependencies between tasks, the dependencies between projects, etcetera. So you can have some of that richer detail beyond just a
261
00:35:00.755 -->
00:35:09.635linear view of GitHub issues somewhere of there's this GitHub issue, and maybe there's some linking between them, but it's still a fairly flat relationship.
262
00:35:09.635 -->
00:35:10.595You can actually
263
00:35:10.995 -->
00:35:19.960have remote URIs from the node in the graph to some of those GitHub issues so you can actually build some of that hierarchical state within the graph.
264
00:35:20.440 -->
00:35:26.600Because of the fact that it's a shared state space, you can have multiple different coding agents coordinating
265
00:35:26.600 -->
00:35:41.495on those projects and tasks and memories so that as I'm working in one piece of the overall system, I'm providing input about some of the broader organizational details such as this repository is actually slated for deprecation,
266
00:35:41.495 -->
00:36:08.515so don't bother doing a bunch of maintenance work on it. So that might change the way that somebody else's agent is thinking about the priorities of the tasks that they're picking off because it can see the fact that, oh, this is actually going to be deprecated. I'm not going to put in a bunch of effort. I'm gonna focus over here instead and just the ability to turn some of this agentic software engineering from a very siloed single player mode into a more collaborative multiplayer mode and have some of that
267
00:36:08.835 -->
00:36:09.795evolution
268
00:36:09.795 -->
00:36:11.155of capability
269
00:36:11.155 -->
00:36:39.305beyond what is reasonable within a single machine. And even just as I'm developing it and I haven't quite gotten it deployed to that multi user state, but even just having multiple agents running on my machine, I've seen it pick up some of the fact of, oh, well, this task over here touches on this thing I'm doing over here in this other session. And this memory was stored three sessions ago, but it gets resurfaced because it's relevant to the work that I'm doing now and just some of the ways that that self improving system comes about because you do have that
270
00:36:39.705 -->
00:36:41.065rich, detailed, useful
271
00:36:41.990 -->
00:36:45.990history and detail about what's being done and what's been done
272
00:36:46.230 -->
00:37:00.455changes the behaviors of the agents themselves. And I can only imagine some of the ways that that will change some of the things, like, as I evolve this to support something like a operations agent that lives in a Kubernetes cluster that's triaging errors
273
00:37:00.455 -->
00:37:11.220from the logs or incidents because of the fact that it's able to tap into some of the organizational memories and tap into some of the code graph, it can get some of those richer semantics and more accurate
274
00:37:11.540 -->
00:37:12.340debugging
275
00:37:12.340 -->
00:37:18.900than if it were just coming in fresh and only has some of the context that it's able to retrieve as part of the incident.
276
00:37:19.540 -->
00:37:30.235Yeah, no, for sure. I think like, I think also what it eventually boils down to is actually that, that agents really need more kind of explicitness of the world they operate in, right? Because I think with human collaboration,
277
00:37:30.235 -->
00:37:31.195human organizations,
278
00:37:31.195 -->
00:37:40.850lot of that kind of knowledge or information is kind of implicit, you know, of course, you know that the production database is, is there, or it depends on x, or like that person is involved there. Essentially,
279
00:37:40.930 -->
00:37:44.450age is supposed to be automated off the work, they don't have that kind of explicitness,
280
00:37:44.450 -->
00:37:50.050you know, essentially, that's where essentially they hit these blind spots that can have like severe kind of consequences.
281
00:37:50.355 -->
00:37:58.835Also on the other side, it can also like improve, improve as a massively also their ability to actually perform useful work, because they know exactly what are the dependencies between
282
00:37:58.915 -->
00:38:01.715the different things. So yeah, completely agree there.
283
00:38:01.955 -->
00:38:07.900I think what's kind of also interesting is that if we think about agents as kind of also kind of constructors
284
00:38:07.900 -->
00:38:13.100of new knowledge or enriches of knowledge, and we can also think of the, of the graph as kind of,
285
00:38:13.500 -->
00:38:17.820you know, an index or like indexing information, able to sort of like cumulatively
286
00:38:17.820 -->
00:38:26.425index more and more information, because the point is that from different kind of low level facts, you can essentially construct more information that you also store in the graph. You can essentially,
287
00:38:26.664 -->
00:38:27.305I think
288
00:38:27.625 -->
00:38:38.130this kind of, excuse me, maybe against like some of the system that say like, okay, we can just, we're just kind of retrieved information from the source system essentially at like query time, right? And essentially that
289
00:38:38.610 -->
00:38:45.890sort of means that at query time it has to kind of construct maybe more knowledge and one can only for example here is like entity resolution,
290
00:38:45.890 -->
00:39:16.150that if two different pieces of data essentially are the same, but they are not linked essentially by a foreign queue or something like that, then essentially has to kind of duplicate those entities at query time. But if you do that essentially beforehand, essentially having an agents enrich and for example link up entities that are know, duplicate or do this kind of semantic resolution. That means that essentially now you're kind of performing proactive work and constructing knowledge that then you can use that instantly at query time. And I think I kind of view that also as like a major paradigm that's going to emerge since you have agents just
291
00:39:16.655 -->
00:39:21.935cumulatively constructing knowledge that can then be built on by essentially other agents as well.
292
00:39:43.805 -->
00:39:51.005Yeah, no, for sure. I mean, because not just the open source, but also because we are kind of optimizing also for sovereignty,
293
00:39:51.005 -->
00:39:55.965because I think we believe that the old models of like cloud and SaaS of kind of locking in your data into an ecosystem
294
00:39:56.980 -->
00:40:27.230doesn't work anymore. There's just too much kind of, you know, market pressure right now. And it makes sense because essentially, if everybody has access to the same level of intelligence, essentially all your alpha is basically in the knowledge data you have, and then you don't want it to be either to be locked into either some kind of software as a service or essentially a cloud sort of database data provider. So essentially sovereignty is, I mean, something that we believe in like philosophically, but I think also what the market now demands, given that also, you can just build so much internally.
295
00:40:27.390 -->
00:40:42.635And so that also changes kind of like what then the business model, you know, looks like. And so the approach we took is kind of maybe similar to what we're taking is similar to kind of TerraForm in that, essentially, we kind of offer essentially the control pane right around essentially
296
00:40:42.795 -->
00:40:44.715managing this whole kind of state,
297
00:40:44.955 -->
00:40:50.235because you know, there's all kind of things like, you know, managing of different graphs and different clusters, and then also
298
00:40:50.475 -->
00:40:51.675schema migrations,
299
00:40:51.755 -->
00:41:06.819version migrations, and then the whole idea of kind of, you know, SLAs and ensuring like your system is actually operational and the observability around it. So this is kind of what we're optimizing for in terms of like, you know, bigger companies. And we're kind of also working on a
300
00:41:07.555 -->
00:41:20.995self serve offering to just get sort of like in a graph, instant graph to kind of work with, but in that sense, you know, the open source plays in well there, because I think bottom up adoptions was seen also with, you know, things on the success of Clickhouse is kind of a major,
301
00:41:22.370 -->
00:41:25.650you know, driver of like, you know, adoption of new technology.
302
00:41:26.210 -->
00:41:27.570And so, yeah.
303
00:41:30.050 -->
00:41:30.930And
304
00:41:30.930 -->
00:41:33.970as you have been building Omnigraph,
305
00:41:33.970 -->
00:41:41.255popularizing it, working with some of the early adopters, what are some of the most interesting or innovative or unexpected ways that you've seen it used?
306
00:41:41.575 -->
00:41:57.340Yes, one team of you using it also as a kind of for automating like, essentially, like an automated basically trading system. It's not a high frequency trading system, basically like a more one based on like fundamental research and uses kind of the graph as a shared artifact
307
00:41:57.420 -->
00:41:59.660to perform a kind of like market research,
308
00:41:59.740 -->
00:42:03.420and have, you know, multiple agents like enrich information producing like signals.
309
00:42:03.885 -->
00:42:30.990Think that was kind of an interesting one. And the other one that I mentioned earlier, also one that we hadn't necessarily optimized for is actually was actually using the graph as kind of an indexing mechanism for kind of like lens data sets, and then using that essentially for essentially training like ML models, basically. So these were kind of two interesting ones, because I mean, ones that we were expecting were the ones I mentioned earlier around the company brain, the engineering or code graphs, and then also kind of a multi agent research.
310
00:42:31.310 -->
00:42:32.990I think, going forward, think there's gonna be,
311
00:42:34.005 -->
00:42:36.965I think a lot of interesting work on how to kind of engineer
312
00:42:37.285 -->
00:42:38.405essentially the
313
00:42:38.724 -->
00:42:45.765system as a whole. So basically the combination of, you know, the ontology, the graph and the state in there, but also the agent
314
00:42:45.890 -->
00:42:46.770configuration,
315
00:42:46.770 -->
00:42:50.290so which agents run with what prompts, you know, at what cadence,
316
00:42:50.530 -->
00:42:58.050and in that sense we're also, you know, working on the event streaming layers, you'll be able to emit events from changes in the graphs of the CDC,
317
00:42:58.050 -->
00:43:23.110so that agents can react to specific changes in graph and then perform some piece of work. I think what's going be really interesting is basically how do you engineer this whole system as a unit, know, rather than, for example, so far, like in terms of the loops engineering, the focus has been on, you know, engineering the harness itself with specific agent, but now how do you kind of optimize the system as the old that is configured by, you know, the multiple agents with the default prompts, and then also the ontology
318
00:43:23.590 -->
00:43:38.205that defines basically the state space that evolves basically, right? I I think I kind of view this loop more as a sort of like, almost like a dynamic or system, know, that, you know, has a, you know, a timestamp, and then each timestamp you have like, you know, some evolution in the, in the, in the data.
319
00:43:38.445 -->
00:43:58.385And then essentially ontology or the schema forms basically the the state space that is kind of more constrained, right? Because you can imagine if you have, let's say, no constraints, or you just have a, let's say, one markdown file that multiple agents read and write to, of course, like the first problem is you will get conflicts, right? Because in terms of like concurrency, because there's no sort of decoupling modularity.
320
00:43:58.465 -->
00:44:04.785But I think the second part is actually just going to very quickly drift into like a high entropy sort of noise basically.
321
00:44:05.025 -->
00:44:17.480And so, but in the case of the graph and the agent configurations, like the question would be, okay, how do you engineer this kind of configuration of the system as a whole, that you kind of like reduce entropy kind of over time and kind of get essentially,
322
00:44:17.880 -->
00:44:31.415you know, a lower entropy state of that graph, but also with essentially not just like, you know, a lower like entropy or, but also that the information there is actually semantically also interesting or useful. And that will depend a lot, think actually on the
323
00:44:31.735 -->
00:44:49.470sort of the configuration of system. And maybe it also mean that we might be able to use some of that as a kind of reward signal for like post training or reinforcement learning, essentially when you have multiple agents acting on that system, and then optimizing the models to operate as a kind of joint system rather than the individual harness itself.
324
00:44:49.790 -->
00:45:00.325I think like the thinking there is still like quite early, but I think this is kind of my opinion, ultimately, what what's gonna matter is, like, really engineering the the that kind of outer loop as a whole, basically.
325
00:45:01.525 -->
00:45:08.090Yeah. And because of the fact that you do have the Git semantics as part of that core capability
326
00:45:08.330 -->
00:45:09.770and the versioning,
327
00:45:09.770 -->
00:45:17.450I can imagine that it would have a substantial benefit for things like experiment tracking or prompt evolution
328
00:45:17.610 -->
00:45:19.690through something like a DSpy
329
00:45:20.010 -->
00:45:22.330and just some of the ways that that can
330
00:45:22.914 -->
00:45:29.075simplify some of the MLOps space around the agent lifecycle and agent evolution.
331
00:45:29.555 -->
00:45:43.330Yeah, exactly. Yeah, because one way to view the branches is definitely not just as a sort of governance mechanism of like verifying outputs and things like that, but also as a way for agents to explore like parallel solutions in the solution spaces basically,
332
00:45:43.650 -->
00:45:46.450and then kind of discarding like some of the solutions.
333
00:45:46.610 -->
00:45:49.090And so you can imagine that kind of multi agent system in that
334
00:45:49.595 -->
00:45:50.714you might have the
335
00:45:51.595 -->
00:45:55.675graph itself that emits events based on changes in the data,
336
00:45:55.914 -->
00:46:09.080but essentially that agent subscribed to changes maybe only on the main branch, but then essentially on the other branches, and the other agents react to it, so that essentially only when it gets merged, does it trigger sort of the downstream, you know, effects basically.
337
00:46:09.320 -->
00:46:15.400And so yeah, I think it's really interesting to kind of maybe view these branches as kind of a way to kind of explore like parallel solution spaces.
338
00:46:15.924 -->
00:46:36.650I think similarly also to have loops itself on the branch so that you can have verification going on similar to how we do, you know, coding agent reviews on pull requests. Like right now, we're of doing loops on these pull requests, I just have a review, then we correct it, and we have another review. So So can we kind of, but mostly it's us kind of, they say like, oh, look at the PR comments, you know, and now you know, correct them, but we're doing that manually, but can you automate
339
00:46:36.650 -->
00:46:38.089that kind of process itself,
340
00:46:38.250 -->
00:46:41.930and then provide some kind of, you know, more compressed observability
341
00:46:41.930 -->
00:46:45.685to the human so that essentially they can add that judgment in a sort of,
342
00:46:46.085 -->
00:46:55.205you know, without being exposed to the high dimensional changes, but to kind of maybe what are the downstream impacts of, you know, merging that that truth, basically, and things like that.
343
00:46:55.765 -->
00:46:58.964In your experience of building this
344
00:46:59.440 -->
00:47:06.000system and technology and business, what are some of the most interesting or unexpected or challenging lessons that you've learned in the process?
345
00:47:06.560 -->
00:47:28.940So yeah, mean, definitely building a graph engine is like is like not the easiest thing to to build. And also like, there we were kind of also relying essentially on coding agents, and essentially you will see them now performing like essentially super well when it comes to like designing, like building like a CRUD app or some kind of like UI and software and things like that. But you know, in the case of like building GraphQL, there's a lot of kind of invariants
346
00:47:28.940 -->
00:47:30.619that you need to maintain,
347
00:47:30.700 -->
00:47:32.859For example, like, okay, what's the consistency
348
00:47:32.859 -->
00:47:40.539level you ensuring? Like what are some kind of latency constraints and things like that. And so like, a lot of changes will have a lot of kind of downstream effects,
349
00:47:40.855 -->
00:47:41.815and also
350
00:47:42.295 -->
00:47:48.695around also essentially the complexity of the task itself. And then that's where we kind of saw a lot of kind of the limitations of
351
00:47:49.015 -->
00:48:28.020the way kind of coding agents work. And so this kind of graph like, you know, loops of like finding the context that they need. And I think this is kind of informed as well a lot like building this kind of graph on top of a code base to be able to kind of have essentially a better management of the knowledge around the code base. Because really the problem with the with the with the agents that they will see like a function, and then assume it works a certain way, or like for example, the case of lens, where we, you have this kind of function that, you know, performs the operation on the underlying fragments, then you kind of expect they will behave in a certain way, and sort of, well, mean the agent expect that it will behave in a certain way and or assumes it because it's kind of obviously
352
00:48:28.100 -->
00:48:39.185optimal to not like read every piece of code file and like look at the actual source code, right? But it actually makes a lot of assumptions or hallucination in that case. I think this is kind of where it became like
353
00:48:39.505 -->
00:48:43.025actual botanite became is actually how do you represent that notion in a way that
354
00:48:43.345 -->
00:49:12.185essentially when agents are coding, they can actually retrieve the right piece of information without bloating the whole context size by saying, you know, read every kind of code file. And I think this is kind of where a lot of the intuitions emerge, because of the complexity of that engineering task around, you know, how to best use like, coding agents and actually building essentially a knowledge layer for them, they're able to kind of incrementally build up knowledge about like source dependencies or its own code base, and then rely on these facts when it's designing different pieces of the architecture.
355
00:49:12.185 -->
00:49:21.480And I think this is kind of one of the most interesting ones, because we're not necessarily hitting these limits when you're kind of designing more simple systems, or especially when it comes to green field projects,
356
00:49:21.640 -->
00:49:34.335but anything like when you have, you know, bigger upstream, like bigger library dependencies where we used to depend on the low level implementations of them, or your codebase itself is like brownfield and it's like a huge, you know,
357
00:49:34.735 -->
00:49:40.895huge code base, then then really the knowledge around the code base, how do you optimize the representation, the indexing,
358
00:49:41.215 -->
00:49:48.735such that essentially at a query time when it is performing coded, it can very quickly get the right information to make the design like optimal and things like that.
359
00:49:50.480 -->
00:49:53.760Yeah. And one of the things too that I was building in my project
360
00:49:54.160 -->
00:49:58.079that relies on OmniGraph is that graph representation
361
00:49:58.079 -->
00:50:06.815of the code based on tree sitter parsing of the code that exists on disk and then being able to generate heuristic
362
00:50:06.815 -->
00:50:08.974connections across repositories
363
00:50:08.974 -->
00:50:11.775so that for a service oriented architecture,
364
00:50:12.095 -->
00:50:19.934as the agent is making changes in one place, it can have some visibility into the potential impact across process boundaries
365
00:50:20.170 -->
00:50:34.090so that it can account for that either as designed to say, hey. I need to go do something in this other project now as well, or I need to do this in more of a two phase process of this evolution to make sure that you don't introduce breaking changes naively.
366
00:50:35.145 -->
00:50:44.985Yeah, exactly. 100%. I think it was the boils down to like our previous I think a lot of the the problems with Alex's kind of the the blind spots, right? And essentially,
367
00:50:45.065 -->
00:50:45.625like,
368
00:50:45.865 -->
00:50:56.690it doesn't know what it does know, and essentially that's where the the edge or the link in a graph comes in, because if you link something that when it looks at that node or entity that it sees
369
00:50:56.930 -->
00:51:05.825all the things that depend on it, that's how you kind of reduce the blind spots or in some sense the in kind of retrieval language, you kind of optimize them for the for recall
370
00:51:05.905 -->
00:51:06.945of the information,
371
00:51:07.105 -->
00:51:20.530because you kind of prepared it ahead of time in some sense. But yeah, I think like if I had brought down like the limitations of the agents, mostly boils down to really the like, blind spots or the or the recall side, maybe less so, for example, the the precision.
372
00:51:21.410 -->
00:51:23.650And so for people who are
373
00:51:23.970 -->
00:51:26.050either looking for some
374
00:51:26.290 -->
00:51:34.465substrate for a multi agent system or they're looking for some graph engine, what are the cases where Omnigraph is the wrong choice?
375
00:51:34.944 -->
00:51:39.105So yeah, if you say specifically for kind of agents and multi agent system,
376
00:51:39.345 -->
00:51:45.820then I would say it really depends on sort of the the complexity complexity level level of of the the system, basically, right? Because I think
377
00:51:46.460 -->
00:52:08.545I think once you kind of start simple, and that's where a lot of people get started, right, maybe a markdown file as a kind of shared state layer, right? And I think this works well. So if you have a kind of a single agent and not multiple agents reading or writing to it, but also that the actual knowledge you need to represent kind of, you know, fits into like a context size, because that's going to outperform like any other system, because you're able to kind of retrieve everything within the context.
378
00:52:08.704 -->
00:52:12.890But I think as kind of the amount of data increases,
379
00:52:13.130 -->
00:52:18.329both in terms of volume, but also in terms of the semantics that you're actually tracking multiple kinds of entities,
380
00:52:18.890 -->
00:52:22.250and at the same time you're also increasing the number of automations
381
00:52:22.250 -->
00:52:30.725you're running, so in that case, like number of agents you want to run-in the background, I think this is kind of where really Omnigraph starts kind of earning its,
382
00:52:31.125 -->
00:52:32.725its really kind of potential.
383
00:52:32.885 -->
00:52:33.845And I would say,
384
00:52:34.325 -->
00:52:40.245yeah, Omnigraph should be used kind of for very simple use case, because it's still quite a complex system to operate to understand.
385
00:52:40.325 -->
00:52:53.990I think obviously now things are becoming easier to operate with, with agents, right? I think most of the users of Omnigraph actually just been using the Omnigraph skill or pointing at essentially all our documentation
386
00:52:53.990 -->
00:53:22.505was optimized for agents and is able to actually guide them through it. And we have these cookbooks and things like that, right. So the onboarding is actually much easier now. And I think this is why in the past graphs were like less popular. I used to also be kind of quite obsessed about Sparkle and RDF and things like that. And like the amount of complexity that you need to understand around us all the semantics and that that's why then there was never kind of a wide adoption. But now essentially, agents are able to kind of really bridge that gap and actually make more complex things adoptable.
387
00:53:22.505 -->
00:53:36.265But still, would say kind of to answer the question, I think Omniglass shouldn't be used for like very kind of simple workloads. And so that it also mean maybe like, I think the two dimensions, like as I mentioned, the complexity at the volume of the data, then also single player versus multiplayer.
388
00:53:36.265 -->
00:53:39.570So Omnigraph makes much more sense when essentially multiplayer,
389
00:53:39.570 -->
00:53:43.090like multiple agents or multiple agents and humans essentially.
390
00:53:43.650 -->
00:53:48.530And as you continue to build and iterate on the Omnigraph project
391
00:53:48.945 -->
00:53:57.585and build the Modern Relay company, what are some of the things you have planned for the near to medium term, either for Omnigraph specifically
392
00:53:57.585 -->
00:53:59.505or any associated
393
00:53:59.505 -->
00:53:59.985projects?
394
00:54:01.040 -->
00:54:02.560Yeah, so for
395
00:54:02.560 -->
00:54:13.520Omnigraph, obviously, like we are kind of also quite focused on the actual deployment and essentially making deployment much easier with the control pane, right? And, and everything in Omnigraph is like declarative, so you know, the context,
396
00:54:14.995 -->
00:54:15.875ontology,
397
00:54:15.875 -->
00:54:16.755the policies,
398
00:54:16.995 -->
00:54:23.955and so making that kind of state management much easier is something that we're working on, and also a sort of like self self deployment.
399
00:54:24.035 -->
00:54:31.500But at the same time, I think the last gap as well is around also the kind of CDC or events sort of streaming,
400
00:54:31.820 -->
00:54:32.860because I think
401
00:54:33.740 -->
00:54:34.780the automation
402
00:54:34.780 -->
00:54:46.775eventually should be really event driven so that essentially agents can react to changes and kind of perform work and then entering that layer is going to be kind of a huge unlock to build this sort of this multi agent system because currently it relies on like crueling jobs, like polling,
403
00:54:46.855 -->
00:54:58.210essentially the state, basically be able to react very quickly to events going to be sort of really like an unlock to build these, multi agent systems. At the same time, kind of also the philosophy is like obviously again, like composability,
404
00:54:58.210 -->
00:55:14.905and also where does actually UI play a role now that essentially most of the work is performed by, you know, agents and where do kind of humans come in. I think sort of the UI is just to kind of evolve to be sort of more focused on like a job to be done, right? Maybe it's like, you know, you are you where such a human judgment observability
405
00:55:15.385 -->
00:55:26.450matters. And so for that, we kind of build these, we're building these notebooks similar to maybe the concept of like a Jupyter notebook or a dashboard, and that you can essentially generate this kind of UI blocks
406
00:55:26.450 -->
00:55:29.490that essentially look at a single lens or slice of the graph.
407
00:55:29.970 -->
00:55:35.490And so and also perform mutations that you know, might make sense for human records. Sometimes moving like your tasks
408
00:55:35.650 -->
00:55:48.735across a Kanban board is like more efficient in prompting the agent 20 times you know, to move like the different things around. So essentially, can you generate essentially individual UI blocks that are optimized for a very strict job to be done that maybe you have to do every Monday,
409
00:55:48.975 -->
00:56:15.945every day, or maybe just for an event that you're organizing or whatever, you kind of look at the slides and share the UI for it. And so we're kind of also burning out that layer. But again, it's supposed to be interoperable and sovereign, so we nothing is kind of monolithic in that, you you use everything, but essentially composable, you might, you know, vibe code your own UI on top of it, or plug into a via CLI SDK or sort of API essentially to any other sort of downstream system. So really kind of optimizing there also for
410
00:56:16.265 -->
00:56:16.905interoperability.
411
00:56:17.980 -->
00:56:24.220Yeah, and I haven't dug into it much. But I've also seen references to things like MCP UI,
412
00:56:24.300 -->
00:56:29.740where the agent can call into some UI generation capability that has some scaffolding,
413
00:56:29.740 -->
00:56:37.005but is largely driven by the agents. That'd be interesting to see as well how that could be factored into some of the data that comes out of Omnigraph.
414
00:56:37.085 -->
00:56:43.085Oh, yeah. For sure. Yeah. So you mean that basically the MCP will generate the UI within sort of the chat conversation itself?
415
00:56:43.980 -->
00:56:52.380I'm not entirely sure how it is intended to function. I just know that the idea is through tool calling, the agent can dynamically generate the UI
416
00:56:52.619 -->
00:57:04.615at the time of it interacting with whatever the objective is. Yeah, exactly. So I think yeah, there's definitely a major use case. Like our process into the notebook is also that it's kind of like in between generating,
417
00:57:04.615 -->
00:57:09.414let's say an AI generating the actual source code, in like React to build like the full application.
418
00:57:10.220 -->
00:57:14.940And it's kind of like one layer above, and it's based on like high level abstraction,
419
00:57:14.940 -->
00:57:19.180that's also declarative that that exposes like, for example, the queries and mutations,
420
00:57:19.339 -->
00:57:53.165so that you can just assemble, so the agent can actually write a whole dashboard in a few seconds on these like building blocks, rather than like, you know, building 20 different React apps that don't talk to each other, but that's generating these kind of blocks that you can then move around. I think, and it is again, like reducing the service area or decision service area for an agent on like, on the only corp file it needs to get right, right? In this case of generating UI, what should it bother with? Or the React code and nuances that you should focus on is like, okay, what data should be represented in what way, and essentially, and what mutations are allowed. And then obviously the,
421
00:57:53.325 -->
00:57:58.445like permissions and all these things are kind of handled by the data layer itself or the business logic.
422
00:57:58.525 -->
00:57:59.165But yeah.
423
00:58:00.180 -->
00:58:15.140Are there any other aspects of the work that you're doing on OmniGraph or any of the related projects or just this overall space of state management for agentic and multi agent systems that we didn't discuss yet that you'd like to cover before we close out the show?
424
00:58:15.795 -->
00:58:20.595No, I think we pretty much discussed everything. Yeah, that was super exciting conversation.
425
00:58:21.315 -->
00:58:38.850Alright, well, for anybody who wants to get in touch with you and follow along with the work that you and your team are doing and try out OmniGraph, we'll have you add your preferred contact information to the show notes. And as the final question, I'd like to get your perspective on what you see as being the biggest gap in the tooling or technology that's available for data or AI systems today.
426
00:58:40.369 -->
00:58:43.570Well, mean, mean, just like the obvious ones, for example, like the
427
00:58:44.255 -->
00:58:46.415latency of the
428
00:58:46.415 -->
00:58:48.175agents or the inference.
429
00:58:48.815 -->
00:58:52.175And then also like the constraints you're hitting when you're actually
430
00:58:52.575 -->
00:58:58.975running so many agents on your computer and the workload and things like that. But I guess they're not as interesting because they are kind of commonplace.
431
00:58:59.420 -->
00:59:04.380But yeah, I'm curious, like, what was your what was your perspective on that one? Or what's your answer on this?
432
00:59:04.940 -->
00:59:09.740I think really the biggest issue right now is just having any
433
00:59:10.300 -->
00:59:11.260coherent
434
00:59:11.260 -->
00:59:13.260way to even understand
435
00:59:13.660 -->
00:59:23.805what is out there and how it relates and being able to make informed decisions when you're making a technology choice because there's so much hype and
436
00:59:24.045 -->
00:59:29.900there's so much overlap between the different pieces. So being able to do that selection of understanding
437
00:59:30.059 -->
00:59:39.660what are the pieces that this actually solves? What are the pieces that it only sort of solves? And how do I need to think about composing solutions together to get to a complete platform capability?
438
00:59:40.555 -->
01:00:12.555I see. Okay. Maybe because there's like so many, so much technology out there, and it's difficult to make this sort of, like, right trade offs and knowing actually what the trade offs are beforehand and things like that. Exactly. Yeah. Yeah. Yeah, I think another one is also as sort of, you know, codegen becomes so, you know, widespread, essentially what happens to kind of the open source and sort the flooding, essentially of solutions, and they also difficult to assess was is that is actually serious effort behind it? Or is that going to be abundant in, you know, in a few weeks? Or, you know, was that kind of vibe coded, was the actual design
439
01:00:13.195 -->
01:00:19.115decision being made, you know, or like RFCs and things like that? I think this is something that maybe was less,
440
01:00:19.595 -->
01:00:34.340was definitely less of a problem in the past, because you know, if code was written, you knew, okay, humans had to think about it, you know, otherwise it wouldn't compile or it wouldn't have been written, or a test wouldn't pass. But now it's actually difficult to evaluate. Of course, you can tell from some of the
441
01:00:34.994 -->
01:00:44.755looking at some of the commits and also the way like maybe it's written, how much effort went into that and things like that. But yeah, I think this is gonna be like a major problem. I think now also in terms of the
442
01:00:45.394 -->
01:00:46.275reviewer
443
01:00:46.355 -->
01:00:48.515load, basically, when you get bombarded
444
01:00:48.515 -->
01:00:54.820with pull requests, it's actually much more expensive to review it than actually just for them to generate the code in the first place.
445
01:00:55.380 -->
01:00:56.340Absolutely.
446
01:00:56.500 -->
01:00:57.860Yeah. Well,
447
01:00:58.020 -->
01:01:03.860thank you very much for taking the time today to join me and share the work that you're doing on OmniGraph
448
01:01:03.860 -->
01:01:26.849and thinking through some of these challenges of being able to own and manage the state space across multi agent systems. It's definitely a very interesting and challenging problem domain. So I appreciate the work that you're putting into Omnigraph, and I've definitely have been taking advantage of it in my day job. So thank you again for your time on that, and I hope you enjoy the rest of your day. Yeah. Thanks so much. It was really fun.
449
01:01:34.415 -->
01:01:38.895Thank you for listening, and don't forget to check out our other shows. Podcast.net
450
01:01:38.895 -->
01:01:47.775covers the Python language, its community, and the innovative ways it is being used. And the AI engineering podcast is your guide to the fast moving world of building AI systems.
451
01:01:48.650 -->
01:01:58.650Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email hosts@dataengineeringpodcast.com
452
01:01:58.650 -->
01:02:04.855with your story. Just to help other people find the show, please leave a review on Apple Podcasts and tell your friends and coworkers.