WEBVTT
NOTE
Transcription provided by Podhome.fm
Created: 03/16/2026
01:50:35Duration: 3709.679
Channels: 1
1
00:00:11.360 -->
00:00:15.440Hello, and welcome to the data engineering podcast, the show about modern data management.
2
00:00:16.075 -->
00:00:29.195If you lead a data team, you know this pain. Every department needs dashboards, reports, custom views, and they all come to you. So you're either the bottleneck slowing everyone down, or you're spending all your time building one off tools instead of doing actual data work.
3
00:00:29.970 -->
00:00:37.010Retool gives you a way to break that cycle. Their platform lets people build custom apps on your company data while keeping it all secure.
4
00:00:37.489 -->
00:00:46.155Type a prompt like build me a self-service reporting tool that lets teams query customer metrics from Databricks, and they get a production ready app with the permissions and governance built in.
5
00:00:46.555 -->
00:00:50.395They can self serve, and you get your time back. It's data democratization
6
00:00:50.395 -->
00:00:51.595without the chaos.
7
00:00:51.915 -->
00:00:54.875Check out Retool at dataengineeringpodcast.com
8
00:00:54.875 -->
00:01:01.879slash Retool today, that's r e t o o l, and see how other data teams are scaling self-service.
9
00:01:01.879 -->
00:01:05.560Because let's be honest, we all need to retool how we handle data requests.
10
00:01:05.960 -->
00:01:16.445Your host is Tobias Maci, and today I'm interviewing Raj Shukla about building self improving AI systems and how they enable AI scalability in real production environments. Raj, can you start by introducing yourself?
11
00:01:16.765 -->
00:01:34.120Hi, Tobias. Very nice to be here. Thanks for having me. My name is Raj Shukla. I am the CTO at Symphony AI. Symphony AI is a vertical AI company, which means we take one domain or one vertical at a time, and we go very deep with
12
00:01:34.200 -->
00:01:42.600AI agents, AI models to achieve end to end business process automations, and just generally helping build the autonomous enterprise.
13
00:01:42.955 -->
00:01:57.820That's what we sell to our customers. So it'll be agents as a service, models as a service, and some applications that bring it all together. I've been at Symphony AI for for the last three and a half years. It's been a great journey working with true
14
00:01:57.900 -->
00:02:05.500real AI in real industries and factory floors and grocery stores and, you know, financial institutions fighting crime. So that's
15
00:02:05.500 -->
00:02:09.580my journey here. And before this, I used to be at Microsoft,
16
00:02:10.060 -->
00:02:13.635kind of had roles across the company, mostly in applied AI,
17
00:02:13.795 -->
00:02:17.155machine learning, engineering leadership positions at Microsoft.
18
00:02:17.715 -->
00:02:21.315And do you remember how you first got started working in the ML and AI space?
19
00:02:22.114 -->
00:02:25.970Yes, I do. I I I just happen to be
20
00:02:26.290 -->
00:02:27.410at the very
21
00:02:27.569 -->
00:02:29.489you know, right around my
22
00:02:29.730 -->
00:02:31.570I I was a computer science major.
23
00:02:31.730 -->
00:02:36.049And right about when I was finishing my bachelor's, coming into my master's,
24
00:02:36.745 -->
00:02:46.905just machine learning systems were becoming popular. So I've, you know, got introduced to machine learning as a true, you know, with its mathematical backgrounds in optimization
25
00:02:46.905 -->
00:03:00.160problems and all related areas. So it started with a bit of theory around machine learning and then got into applied machine learning. And as a part of my master's, I was doing anomaly detection systems and very large scale
26
00:03:00.719 -->
00:03:03.680network systems. So like, you know, telecom
27
00:03:04.035 -->
00:03:25.940networks and how you detect anomalies and how do you, how can you build systems that predict and react fast to anomalies and so on. So, you know, as a part of my kind of growing up in computer science itself, I was trained a bit in machine learning. And when I came into the industry, my first real jobs were in these areas of search and advertising.
28
00:03:25.940 -->
00:03:34.075So my first job was in click prediction models that allowed, you know, advertisers to figure out where to place their ads on pages, etcetera.
29
00:03:34.315 -->
00:03:50.690And once I got into Microsoft, my first work was on search ranking problems, and which was the kind of the early at scale machine learning models which used to be applied in the industry. Industry. So always been in this space, always enjoyed it. There is a element of magic
30
00:03:50.690 -->
00:03:54.450to how it all works. So that's what keeps me going.
31
00:03:54.849 -->
00:03:59.090And now digging into this concept of self improving AI,
32
00:03:59.645 -->
00:04:19.750obviously, system is a very necessary element of it because the AI models themselves are largely static until you retrain them or post train them. And I'm just wondering if you can just start by giving an outline of what constitutes a self improving AI system and what are the signals that you look to for whether the system is capable of that type of evolution?
33
00:04:20.310 -->
00:04:34.095Yeah, absolutely. This is very critical once you apply the modern foundation models and agentic system based on these foundation models out in the industry. So, you know, you have the concept of the environment.
34
00:04:34.415 -->
00:04:35.935An environment is
35
00:04:36.735 -->
00:04:42.975really a real world system which has where an agent is operating or one
36
00:04:42.975 -->
00:04:48.789of your models is operating and it's trying to do a task And it's operating under
37
00:04:48.789 -->
00:04:54.550some influence of certain external variables. So there could be triggers. For example,
38
00:04:54.949 -->
00:04:59.349people are buying things off of grocery shelves and the shelves are going empty.
39
00:05:00.025 -->
00:05:25.920And there is a monitoring system sitting there and which is observing it. And so that can create a trigger. And once the system takes actions on that, there is a reaction to it from the environment, right? And the reaction could be on the positive side that that action was good, or it could be on the negative side that that action wasn't good. So, for example, in financial crime fighting, the models and agents can detect whether something is fraudulent or looks like a money laundering activity.
40
00:05:26.000 -->
00:05:31.965And then at the end of it, there is a human who is investigating it, and the model makes a recommendation and
41
00:05:32.445 -->
00:05:34.845the human could say, No, that is not right because
42
00:05:35.005 -->
00:05:42.285of something that they know or they found out, or it's right and it's largely right. And so that's kind of, that's, so
43
00:05:43.730 -->
00:05:47.010there is a definition of an environment which kind of creates
44
00:05:47.330 -->
00:06:04.705the conditions under which the, all the input variables under which the agent is operating. And once the agent takes an action, there is some feedback coming whether that is right or not. And the idea is that whether it is with a self learning model inside the agent or whether it is through
45
00:06:04.865 -->
00:06:08.145intelligent kind of memory updates or other techniques,
46
00:06:08.385 -->
00:06:13.320that once you get the feedback, there's something should change, right? So the
47
00:06:13.480 -->
00:06:27.935the system should say, okay, I know I made a mistake and I got feedback for it. And I'm updating something, whether it's a memory database or whether it's it's core LLM configuration or the model itself
48
00:06:28.175 -->
00:06:32.495that I'm trying to improve. Next time, this will not going to happen, probabilistically,
49
00:06:32.495 -->
00:06:43.570of course. And that's what we try to achieve. I think there are practical considerations of it, and then there are theoretical ways of treating the environment as a reinforcement learning or RL environment.
50
00:06:43.650 -->
00:06:58.205And there is practical aspects of, you know, treating it as a self learning system with triggers and hooks and so on. So the field is very exciting, but I think it's primed for, you know, I would not say disruption, it's primed to go big
51
00:06:58.205 -->
00:06:59.565this year and next.
52
00:07:00.365 -->
00:07:02.605And in that element of
53
00:07:02.764 -->
00:07:05.485interacting with the environment
54
00:07:05.485 -->
00:07:08.125of the system that the model is operating
55
00:07:08.125 -->
00:07:13.430within, what are some of the components that comprise that environment,
56
00:07:13.430 -->
00:07:15.270whether that's digital infrastructure,
57
00:07:15.750 -->
00:07:17.510physical manipulation?
58
00:07:17.510 -->
00:07:30.335So particularly if you're starting to think towards things like robotics, I'm just wondering if you can talk through some of the pieces that are necessary in order for an AI system to have some of those reinforcement and learning loops.
59
00:07:30.974 -->
00:07:33.375Yeah. Yeah. That's a good question.
60
00:07:33.935 -->
00:07:44.290So I think, the the core aspect of the environment first is how are you digitizing the information that is coming in, whether it's the core data that is coming in. And sometimes that's obvious,
61
00:07:44.450 -->
00:07:47.890sometimes that's hard, right? When you're working with physical environments
62
00:07:47.970 -->
00:07:54.805like factory floors we operate in or grocery stores we operate in, you have to have some vision based, image based sensors
63
00:07:54.805 -->
00:08:06.965out there that is feeding in an input. And that is going to be digitized. That forms your core data layer on which these triggers are created, etcetera. And on the flip side, on the action side,
64
00:08:07.480 -->
00:08:10.600sometimes the actions are notifications
65
00:08:10.600 -->
00:08:11.720to humans
66
00:08:11.880 -->
00:08:49.170to change something. So for example, to go restock the items on the factory floor or on the grocery floor. And sometimes it's about ordering your next batch of goods to come in based on what the agent sees. So on both sides, it's you have to, I think fundamentally you have to look at what is the digital form of the information and then what is the physical translation of it. And it can be as complex as sort of edge computing or edge ML models, and it could be as simple as an API connection that actually triggers some action on
67
00:08:49.570 -->
00:08:50.210the action side.
68
00:08:51.545 -->
00:08:55.945One of the pieces that comes to mind as far as a way for
69
00:08:56.425 -->
00:09:03.625guiding a particular model and improving its capabilities is the what's broadly being termed context engineering,
70
00:09:03.625 -->
00:09:15.149but making sure that the model has access to particular information at a particular point in time. Another component is this idea of agentic memory where the model, as part of its execution,
71
00:09:15.149 -->
00:09:31.815decides that there's a particular piece of information that is relevant for future use, and so it will decide to push that into some form of context store, whether that's a dedicated memory layer or just a text file somewhere. And I'm just wondering if you can maybe give some differentiation
72
00:09:31.815 -->
00:09:36.510between the ways that you think about this overall concept of self improving systems
73
00:09:36.510 -->
00:09:37.390versus
74
00:09:37.550 -->
00:09:44.510the or agentic system terminology that's also broadly used and how you can decide which one you have.
75
00:09:44.910 -->
00:09:47.630Yeah, absolutely. I think there has been many evolutions
76
00:09:47.630 -->
00:09:57.785of it. So if I were to start from the very early evolution that existed even two years ago in the early agentic systems, there is this area of in context learning,
77
00:09:57.865 -->
00:10:13.610which is basically saying if you if an agent at a particular task or a step is making a decision, it is guided by a prompt. Now, what is the prompt guided by? The prompt is guided by some examples usually, which is a few shot kind of exercise.
78
00:10:13.769 -->
00:10:14.410Now,
79
00:10:14.570 -->
00:10:19.290even two years ago, people knew that you could adjust those few shots
80
00:10:20.024 -->
00:10:25.385to be dynamic to the problem. So depending on what type of input you're
81
00:10:25.385 -->
00:10:26.024getting,
82
00:10:26.425 -->
00:10:29.545you could select a different set of examples
83
00:10:29.785 -->
00:10:34.024to make the result a little better. And that itself gave huge
84
00:10:34.750 -->
00:10:35.709improvements
85
00:10:35.709 -->
00:10:37.389in what
86
00:10:37.389 -->
00:10:39.950the models could accomplish. So in this case,
87
00:10:40.269 -->
00:10:59.265the learning loop would be quite simple. The agent is making a decision or the step advocate makes a decision. You get feedback back in terms of examples of where it got wrong or examples of where it got right. And just in context of the next input coming in, you choose the right examples. Right? And that that's a very simplistic
88
00:11:01.890 -->
00:11:06.610extreme of it on the other side is to really set up the true RL environment
89
00:11:06.930 -->
00:11:08.690and to have
90
00:11:08.850 -->
00:11:11.650a true language model
91
00:11:12.050 -->
00:11:13.570be trained
92
00:11:13.570 -->
00:11:15.490with the right kind of techniques
93
00:11:15.745 -->
00:11:23.985through reinforcement learning, whether it's, in some cases, we have like verifiable feedback. So you can do these RLVR kind of loops.
94
00:11:24.225 -->
00:11:35.780In some cases, you do GRPO kind of policies where you are creating reward systems out of feedback you're getting. But in a sense, you are truly learning in the sense of learning
95
00:11:36.020 -->
00:11:42.820and not just updating artifacts anyway. So that's the true form of learning. And then this always learning
96
00:11:42.820 -->
00:11:44.100model,
97
00:11:44.100 -->
00:11:47.415it's what's guiding the true decision step that the
98
00:11:48.695 -->
00:11:53.655agent is operating on. I think what is getting very interesting lately with
99
00:11:54.295 -->
00:11:57.175agents like Claude and agentic
100
00:11:57.175 -->
00:11:59.575systems like OpenClaw recently
101
00:11:59.910 -->
00:12:02.870is that there is a middle ground somewhere where
102
00:12:03.110 -->
00:12:14.865the agent, well, the feedback loop is coming in, but the feedback loop is not going directly into an RL learned system. It's not going as raw prompt. It's actually being updated
103
00:12:14.865 -->
00:12:16.225as memory,
104
00:12:16.305 -->
00:12:25.585as an intelligent memory that gets pulled in at the right context at the right time. And the update on the of that memory is also not just a pure
105
00:12:25.930 -->
00:12:27.210append of that
106
00:12:27.530 -->
00:12:32.570feedback loop. It's an intelligent append in the sense there is an LLM call that says,
107
00:12:32.810 -->
00:12:44.465you can update taste and preferences of this user based on this interaction. And you can update, you know, there's episodic sort of memories, there is kind of stochastic or time based kind of memories and all. And
108
00:12:45.185 -->
00:12:50.065but with these, and you can keep all these memories in rather simple file system
109
00:12:50.705 -->
00:12:52.785based, you know, architectures.
110
00:12:53.105 -->
00:12:56.225And the agent harness is intelligent
111
00:12:56.385 -->
00:13:04.280enough to pick the right memory and update the memory at the right time as a background process. So it's really an engineering feat
112
00:13:04.280 -->
00:13:05.880that we are accomplishing
113
00:13:06.280 -->
00:13:23.135more than a science feat. I think it's definitely turning out to be better than prompt based in learning. And on the other side, the RL based true learning models are harder to implement. So from a practicality of it, it feels like the right middle ground,
114
00:13:23.375 -->
00:13:27.454which is gaining traction, and we are adopting it in our products as well.
115
00:13:28.759 -->
00:13:30.839And another piece of
116
00:13:31.000 -->
00:13:35.000the context engineering beyond the specific
117
00:13:35.079 -->
00:13:56.005memory system where the agent is involved in the creation and retrieval and curation of those pieces of information is also the idea of the various tool calls or MCP servers or even agent to agent interactions that are involved. And I'm wondering if you can give your sense of whether and how you differentiate
118
00:13:56.070 -->
00:13:59.910between those tool use capabilities and the evolution of information
119
00:14:00.150 -->
00:14:01.830available from those tools,
120
00:14:02.390 -->
00:14:08.070versus this idea of the memory system or reinforcement learning or model fine tuning.
121
00:14:09.154 -->
00:14:31.430Yeah. I think the tool use area overall has also seen a very rapid evolution in the last, I would say, one year or so. Right? I remember us, like, we at Symphony AI, we operate in very regulated industries like financial crime. And so our need for our agents to get things right is very high, in the high 90s, 99 percentages,
122
00:14:31.430 -->
00:14:31.990etc.
123
00:14:32.550 -->
00:14:34.310So we tended
124
00:14:34.310 -->
00:14:37.509to not leave things to LLMs.
125
00:14:38.045 -->
00:14:39.644And we had made
126
00:14:40.444 -->
00:14:43.964hundreds of tools, which are very specific deterministic
127
00:14:43.964 -->
00:14:44.685tools.
128
00:14:45.005 -->
00:14:52.285And we used to rely on our agents to do the right tool calling. Within the tools, you would do deterministic calculations so that
129
00:14:52.520 -->
00:14:57.160we don't leave a chance for LLMs to hallucinate while doing those sorts of calculations,
130
00:14:57.160 -->
00:15:03.400right? So it was a very complex system, but it used to get things right, etc. And then, you know, models kept improving
131
00:15:03.560 -->
00:15:07.320on just like, you know, better tool usage, etc. But particularly,
132
00:15:07.655 -->
00:15:09.655they kept improving in
133
00:15:09.975 -->
00:15:18.055certain kinds of tools and their usage a lot. Like, so search as a tool became very popular. Code execution or code writing,
134
00:15:18.295 -->
00:15:28.870obviously, became really good at it. So in today's day and age, you can actually assume that the code writing aspect of the model will get 95%
135
00:15:28.870 -->
00:15:33.910of those tools right. And so you don't have to pre write those tools. You can actually
136
00:15:35.045 -->
00:15:44.325explore that space at runtime and cache those tools later on if you have to be. But I think for me, the mind blowing moment was when these
137
00:15:44.644 -->
00:15:48.885kind of Unix tools and file system tools start getting really good.
138
00:15:49.510 -->
00:15:52.070And I think credit goes to Claude a lot,
139
00:15:52.390 -->
00:15:54.230the Claude team for exploring
140
00:15:54.230 -->
00:16:00.950that path. And just using the file system or Unix based tools as base tools, which can come together
141
00:16:01.435 -->
00:16:08.715to, you know, produce a lot of derived tools, really simplify the stack, right? And I think by the end of last year,
142
00:16:09.195 -->
00:16:44.084we had thrown away maybe 80% of our tools and gotten the same kinds of results with just relying on core basic tools. Now in certain cases for us, we cannot do that because I think you are effectively trading test time compute or long running tool usage and thinking models for latency and speed at which you're executing. And so certain domains we cannot. And we've we've kept the older architectures. But wherever we can trade off time to let the model do its kind of magic with the base tools that it brings and the business context
143
00:16:44.350 -->
00:16:44.990and,
144
00:16:45.230 -->
00:16:48.670the right knowledge context that we bring from our verticals,
145
00:16:48.910 -->
00:16:51.710that has been, like, really simplified our architecture.
146
00:16:52.750 -->
00:16:55.630And as you're talking about this idea of
147
00:16:55.870 -->
00:16:58.190creating the tools dynamically
148
00:16:58.190 -->
00:17:05.054at runtime, Obviously, as you pointed out, people who are using these models for software engineering use cases,
149
00:17:05.295 -->
00:17:17.429it's doing that all the time, and it's encouraged to do so. And I think that also opens up the overall question of what does it mean for an AI system to improve, improve along what axes and for what use cases.
150
00:17:17.830 -->
00:17:20.950Yeah, absolutely. I think if you look at how
151
00:17:21.430 -->
00:17:25.590the leading foundation model companies are thinking about it now, they don't think about
152
00:17:26.615 -->
00:17:32.934AGI as a pure LLM concept. Right? They are saying, oh, you know what? We realized in the middle something
153
00:17:33.175 -->
00:17:52.230interesting happened. Models are improving, but models got really good at code code writing. And that gave an ability for a other kinds of inner loops where when the model cannot figure it out, it can write code and it could create environments where it tests the code. And, you know, so it can create systems
154
00:17:52.310 -->
00:17:57.975not entirely as a model output, but as inner loop sub agents where it can keep improving.
155
00:17:58.215 -->
00:18:01.015And so that is seen as a
156
00:18:01.095 -->
00:18:03.895much clearer path towards AGI
157
00:18:03.895 -->
00:18:10.135than just seeing it just a model improve overall itself. And I think we take a similar approach in our
158
00:18:10.669 -->
00:18:17.950vertical AI systems and agents that we do. I think we gotta make sure that we build these sub agentic loops where
159
00:18:18.590 -->
00:18:21.389the which are vertical specific for us,
160
00:18:21.789 -->
00:18:23.789and we have to be the best at it where
161
00:18:24.085 -->
00:18:28.165the models write code, we give it an environment where it gets the feedback,
162
00:18:28.405 -->
00:18:33.445and and many times, based on the feedback, it rewrites the code and and updates
163
00:18:33.445 -->
00:18:44.789these systems to be more effective. Like, a very simple thing is, like, if you think of deep research agents and in enterprises you can think of deep research agents working on enterprise
164
00:18:44.790 -->
00:18:46.470APIs as the
165
00:18:47.430 -->
00:18:50.630main knowledge context and not search as an API. And,
166
00:18:51.605 -->
00:18:53.925you know, we realized internally
167
00:18:53.925 -->
00:19:04.005that if you're working with internal enterprise APIs and are trying to build a deep research agent on it, you could let the model write code in the middle. And as you give it feedback
168
00:19:04.005 -->
00:19:21.529that whether it was able to find the right kind of information or not, it will write more code to say in, you know, in the future when it does one API call, it gets an output. It does not pass that output to the next LLM call. It actually has written code where it is doing analytics
169
00:19:21.715 -->
00:19:27.555on that output and passing only a good summary of it to the next tool. So in simple terms,
170
00:19:27.795 -->
00:19:41.090our own sub agents evolved from Rag based systems to these really highly evolved agentic search, and then now into this agentic search plus code writing in the middle kind of sub agents. And that's that's really been
171
00:19:41.730 -->
00:19:47.330groundbreaking in some sense to help us achieve the level of autonomy we were expecting for these subsystems to do.
172
00:19:47.995 -->
00:19:52.154One of the other challenges that comes up when you do have
173
00:19:52.235 -->
00:19:53.754this dynamic
174
00:19:53.914 -->
00:19:54.715probabilistic
175
00:19:54.715 -->
00:19:55.514system
176
00:19:55.515 -->
00:19:59.274that is capable of creating its own tools,
177
00:19:59.595 -->
00:20:05.220exploring particular areas of focus is the potential for
178
00:20:05.460 -->
00:20:06.340misalignment
179
00:20:06.340 -->
00:20:09.460with the stated goals of that system.
180
00:20:09.860 -->
00:20:12.820And so that brings a lot of questions around security,
181
00:20:12.980 -->
00:20:14.900identity management, access controls,
182
00:20:15.325 -->
00:20:23.325guardrails, and I'm wondering how you're seeing some of those capabilities evolve in the ecosystem in terms of how people are starting to think about constraining
183
00:20:23.325 -->
00:20:24.045and
184
00:20:24.445 -->
00:20:32.669guiding these systems to ensure that they don't become misaligned, misaligned or if they do that they're quickly redirected to guiding the stated purpose?
185
00:20:33.470 -->
00:20:37.070Yeah. I think that is probably the biggest difference
186
00:20:37.070 -->
00:20:39.470between a consumer application
187
00:20:39.470 -->
00:20:42.029of agents and enterprise application.
188
00:20:42.030 -->
00:20:49.804And so we pay a lot of attention to it. We build a lot of kind of agent lifecycle management layers in our platforms
189
00:20:49.885 -->
00:20:55.725to deal with it. So just to give you an example in financial services of financial crime fighting,
190
00:20:55.885 -->
00:20:58.684there is a before an agent can make any decision.
191
00:20:59.170 -->
00:21:15.664There is an agentic process first which is on policy alignment and so the first step the agent has to say is it has gone through the policy and standard operating procedures and it proves that it got it right it actually goes one step further and highlights the policy gaps it sees
192
00:21:15.985 -->
00:21:20.945where the human should come in and clarify what to do in what what scenario.
193
00:21:21.025 -->
00:21:41.304And in a practical application when we go to our customers we'd say let's make sure we get this right. And the fact that it creates this big to do list and it maps how it created that to do list to what snippet in the standard operating procedure in the bank, that itself gives these, you know, every bank and every big
194
00:21:41.625 -->
00:21:47.624financial institution has these governance committees and model governance kind of teams. And just that is actually
195
00:21:47.785 -->
00:21:49.465far more transparent
196
00:21:50.184 -->
00:22:16.065than what predictive ML models used to do. I mean, these systems, these agents are actually telling, okay, I took this document, I created this to do list. I think when I execute this tool, I will write this kind of code. And I'm doing this because this mapping exists. They are actually being very receptive of it because they never got this level of transparency and explainability from ML models before. So that's like a conscious step we take when something has to be policy
197
00:22:16.145 -->
00:22:16.945guided.
198
00:22:16.945 -->
00:22:32.940That policy alignment itself is the first goal of an agent or a sub agent. Then comes the process of executing on it and while it's executing on it, it go off trail or whatnot? And yes, you just have to keep the right evals, you have to keep the right guardrails
199
00:22:33.100 -->
00:22:34.059around it.
200
00:22:34.540 -->
00:22:46.445Many of these guardrails are, many of these are while the agent is running, but many of these are also at an aggregate level at an agent monitoring step. While these agents are running in production,
201
00:22:46.605 -->
00:22:48.524you are seeing how they are performing,
202
00:22:48.924 -->
00:22:50.445why they are making certain decisions.
203
00:22:51.059 -->
00:22:54.339We have metrics in place which we are observing
204
00:22:54.659 -->
00:23:00.739as these agents are running and seeing how the performance is going. And I'll just go one step further. We are prepared
205
00:23:00.740 -->
00:23:13.945that these agentic systems will not go live right away. So the way we prep for it is we say, let these agents run-in the background and see how they are performing, right? And then over time you build
206
00:23:14.185 -->
00:23:17.065with these data driven, with these metrics driven
207
00:23:17.225 -->
00:23:23.570approaches, you build confidence that look, over the last three months, the agent is performing as good as your
208
00:23:24.210 -->
00:23:31.970human investigator in the case of financial crime, for example. And there also we take extra steps because there is three levels of investigators,
209
00:23:31.970 -->
00:23:37.984level one, level two, level three. And we'll first say, look, it is doing as good as a level one investigator
210
00:23:37.985 -->
00:23:38.624and
211
00:23:39.024 -->
00:23:58.460that gets adoption. And then you prove it doing as good as a level two investigator, kind of like software engineering, you would say these models are doing as good as a junior software engineer and as a senior software engineer and then as a staff or architect or whatnot. But, I mean, these stages are very clearly defined in the verticals and industries we operate in. And we are taking
212
00:23:58.620 -->
00:24:02.220gradual steps, keeping this transparency
213
00:24:02.380 -->
00:24:02.700and
214
00:24:03.184 -->
00:24:13.345strict guardrails in mind, proving it out while running in the background in production systems, not going live, and then stage wise going live. And I think it
215
00:24:13.345 -->
00:24:23.990takes all of that to actually develop confidence in a system like this. And we are also learning along the way as we are doing that. But I'm proud to say we feel like we are the furthest along amongst
216
00:24:24.070 -->
00:24:28.710every company we see in these domains to, you know, and how we are operating on this.
217
00:24:29.430 -->
00:24:38.315The other element of security comes in when you're talking about the agent's ability to dynamically create tools, and therefore, it's executing
218
00:24:38.554 -->
00:24:42.394unreviewed code in potentially a production environment,
219
00:24:42.475 -->
00:24:54.929which necessitates a certain level of sandboxing or constraints on what code or what features or functions can be executed. And I'm curious how you're seeing people manage that aspect of that dynamic
220
00:24:54.929 -->
00:24:55.970runtime compute.
221
00:24:56.945 -->
00:24:58.785Yeah. Yeah. 100%.
222
00:24:58.785 -->
00:25:07.904We we actually provide code execution sandboxes as a part of our platform, and they are being configured to run Python or TypeScript, but more importantly,
223
00:25:08.145 -->
00:25:12.869the how much of the network it can access, which APIs it has access to, whatnot.
224
00:25:13.990 -->
00:25:15.990Yeah, so as far as sandboxing
225
00:25:15.990 -->
00:25:17.029and securing
226
00:25:17.190 -->
00:25:25.429what these agents are doing, it's a very critical aspect. And we don't rely on any third party for that. We ship it as a part of our
227
00:25:25.909 -->
00:25:30.445platforms. What is interesting is file systems are getting
228
00:25:31.005 -->
00:25:43.039an interesting twist here because the agents are operating more and more as files in the file system and between sandbox local file systems and some, you know, more persistent
229
00:25:43.360 -->
00:25:44.559cloud storage,
230
00:25:44.640 -->
00:25:52.320there has to be a lot of system engineering done to keep those two things in sync as the agents are operating on it. So
231
00:25:52.560 -->
00:25:55.920more than the LLM magic, there is a lot of
232
00:25:56.245 -->
00:26:04.884rigorous system engineering that has to be done to to make sure that that happens. But, yeah, sandboxes come come by default as a part of all our deployments.
233
00:26:05.125 -->
00:26:17.990And the choice of, like, Python versus TypeScript versus what they are writing is based on evals in the use cases that we are doing. And then the final thing I would say is we try to make sure the agents
234
00:26:18.070 -->
00:26:23.350operate as real humans. So every agent gets its own auth and
235
00:26:23.670 -->
00:26:25.350every agent has to
236
00:26:25.855 -->
00:26:30.734you know, operate as an identity which can be reviewed and verified in a
237
00:26:31.455 -->
00:26:36.894customer that we deploy it in. And it follows the right authorization as well. So
238
00:26:37.375 -->
00:26:40.654whether it's the MCP servers it accesses
239
00:26:39.950 -->
00:26:41.629or whether it's the
240
00:26:41.950 -->
00:26:43.470API calls or
241
00:26:44.190 -->
00:26:45.710data that it accesses,
242
00:26:45.870 -->
00:26:58.825all of that is governed by the same RBAC and other access control policies that they have. And that's actually the hardest part to get right As like once going live, think POCs and proof
243
00:26:58.825 -->
00:26:59.864points are easy.
244
00:27:00.105 -->
00:27:01.624But when you're going live,
245
00:27:01.945 -->
00:27:06.104the concepts of these governed like for us, every agent,
246
00:27:06.265 -->
00:27:28.394every agentic system is a project with a set of agents operating in it, a master agent of sorts, and a set of resources that are being governed in the project. And the project comes with its own RBAC and access control policies and so on. And while onboarding, we have a step of onboarding an agent. It goes through kind of getting it the right authentication and the right
247
00:27:28.955 -->
00:27:34.394ID in the company that it operates in. So all that has to play in some sense
248
00:27:34.635 -->
00:27:36.154in getting that
249
00:27:36.315 -->
00:27:41.110right kind of security, secure setting in an enterprise for it to work efficiently.
250
00:27:41.270 -->
00:27:46.710Another aspect of this idea of self improving systems is that
251
00:27:46.870 -->
00:27:50.390the models themselves are no longer the differentiator
252
00:27:50.390 -->
00:28:09.659where before generative AI became so widely adopted, there were a number of businesses that were investing a lot of resources into building their custom models that were specific to the problem domain that they were focused on. Deep learning maybe expanded the bounds of what those models could do, but they were still purpose built for a particular use case.
253
00:28:09.900 -->
00:28:24.555Now everyone has access broadly to all of the same set of models. And so in order for it to be something that is useful and a differentiating capability of that business, it needs to have all of these other system level capabilities
254
00:28:24.555 -->
00:28:29.274around it. And I'm wondering how you're seeing organizations think about that level of
255
00:28:29.435 -->
00:28:40.990competition and ways that they think about the purpose of these machine learning and AI models in the broader context of their organization beyond just being an API call to a foundation model provider?
256
00:28:41.710 -->
00:28:44.909Yeah, that's a great question. And we see it in when
257
00:28:45.070 -->
00:28:58.625we talk to our real customers, there is an unknown in how much of their data and IP is leaking into the models. There is an unknown into what is what are they creating as dependencies on these models. I think we are very clear on our stance.
258
00:28:58.705 -->
00:29:13.210We know our industries, we know our domains, we bring that domain knowledge or the domain knowledge graph into every problem that we bring in. We know that there is a lot of context that is customer specific. We provide the right
259
00:29:13.450 -->
00:29:17.769semantics around how to bring that context in that is customer specific.
260
00:29:17.770 -->
00:29:19.129And we provide
261
00:29:19.290 -->
00:29:24.945sort of as the agent works on it, the context it uses, the context is updates,
262
00:29:25.105 -->
00:29:40.090as we discussed earlier in terms of memories. We are very clear to our customers on the IP that lies in there and much of the IP sits with them. So it is very interesting now that as we if we can create these self learning systems truly in production,
263
00:29:41.530 -->
00:29:43.450where the model remains the same,
264
00:29:43.930 -->
00:29:46.090but in updating memory
265
00:29:46.330 -->
00:29:46.810as
266
00:29:47.195 -->
00:29:49.355sitting in file systems and markdown
267
00:29:49.435 -->
00:29:54.875files and so on is the real magic. That's a very attractive proposition to most enterprises,
268
00:29:55.035 -->
00:29:59.274right? It feels like that they are, with that, they are creating that
269
00:29:59.435 -->
00:30:00.394layer of
270
00:30:00.715 -->
00:30:07.240kind of problem context and then process context. And we also use a term called action context.
271
00:30:07.480 -->
00:30:09.160And in some sense, that
272
00:30:09.720 -->
00:30:12.440is not captured by any CRMs and ERPs.
273
00:30:12.440 -->
00:30:14.120And as these systems go live,
274
00:30:14.625 -->
00:30:15.265these
275
00:30:15.425 -->
00:30:30.700folders with these markdown files that get created with the agent start becoming these sources of your processes and your actions and your escalations and all of that. I think it's we are very early in this journey right now, but it's playing out
276
00:30:30.780 -->
00:30:37.980is what I can say. And if we are successful with this, I think every enterprise can own their own
277
00:30:38.380 -->
00:30:40.460knowledge layer per se that is updating.
278
00:30:40.875 -->
00:30:42.955And we as vendors,
279
00:30:42.955 -->
00:30:46.635we are experts in our industry knowledge. We bring that context.
280
00:30:46.795 -->
00:30:51.034Every enter and we help every enterprise have their kind of sovereign
281
00:30:51.035 -->
00:30:53.355knowledge and and and process context.
282
00:30:53.700 -->
00:31:02.899That'll be a great place to land in. But I think we are very early. I think it'll take a whole year, maybe 2026, for that to evolve and see how how it is truly successful.
283
00:31:03.700 -->
00:31:04.739And another
284
00:31:04.820 -->
00:31:05.860piece of these
285
00:31:06.735 -->
00:31:33.669self improving systems is that beyond all of the context and data knowledge capture that these agents are capable of, there is also the underlying evolution of the models by these foundation providers as each generation adds new capabilities or focuses on particular use cases. But it also brings with it a certain level of platform risk as a particular model that you build a lot of your operational
286
00:31:33.830 -->
00:31:48.670capacity around gets deprecated in favor of a different model that maybe behaves slightly differently. And I'm wondering how you think about that aspect of owning your own destiny and how enterprises are thinking about their level of reliance on a particular
287
00:31:48.750 -->
00:31:51.950model governed via API access versus
288
00:31:51.950 -->
00:32:08.605self hosting one of these open weights models or even building some of their own large language models now that that has become, I'm gonna say, commoditized even though there is still a lot of knowledge and capability and in particular hardware required, but it's at least a a known quantity that people can do if they so choose.
289
00:32:09.005 -->
00:32:26.190Yeah. Yeah. No. That's a great question. I think, people don't realize practically what havocates creates to move from one version to another version or an update to like Claude move from 3.5 to 3.7 and then 4.5 and then 3.7 or 3.5 is deprecating.
290
00:32:27.325 -->
00:32:31.725The reality is enterprise systems need a lot of reliability,
291
00:32:31.885 -->
00:32:34.365right? And one taking it to production,
292
00:32:34.365 -->
00:32:46.500there are these strict, you know, we have an evaluation in our platform, which the first thing it checks is if I run this a 100 times, how many times it gets the same result. And every time we've updated the model version,
293
00:32:46.740 -->
00:32:47.299that
294
00:32:48.180 -->
00:32:50.580reliability metric has broken
295
00:32:51.220 -->
00:32:56.100in practice. So we've never been able to do just model upgrades
296
00:32:56.335 -->
00:33:03.775just, you know, by switching from one API to the other. There is always some prompt changes and there is always some,
297
00:33:04.575 -->
00:33:10.495you know, investigation and updates we have to do. So it's not a, I think your point is very right. Just
298
00:33:11.450 -->
00:33:23.450at the end of the day, the foundation model companies have a very wide set of benchmarks that they are operating against. And they can see that overall the model is improving, but in some local
299
00:33:24.034 -->
00:33:25.154benchmarks,
300
00:33:25.154 -->
00:33:31.394it does go down from one model to the other. There are some very popular examples of this when GPT-five
301
00:33:31.475 -->
00:33:37.714launched and 5.2 came along and all that. And in some of the coding benchmarks, it was actually going
302
00:33:37.875 -->
00:33:38.195worse.
303
00:33:38.700 -->
00:33:48.619So I think that's a real problem. Right now, what happens is what's happening is vendors like us, we take the hit of that for our customers. And
304
00:33:48.940 -->
00:33:58.415we are at least build systems where we catch this first and don't let it go live. And then we iterate and help our customers on it. But the whole thing is a little brittle.
305
00:33:59.615 -->
00:34:12.660The challenge is how do you make it, like how do you improve it and not let it be brittle? So of course, one way is to say, you know, have your small language model, small reasoning model that is very good at that task, host it yourself
306
00:34:12.820 -->
00:34:19.060and, you know, let it live out. The, of course, it comes with more investment and more,
307
00:34:20.145 -->
00:34:23.345it's more capital intensive. You have to run your own GPUs,
308
00:34:23.345 -->
00:34:24.065etcetera.
309
00:34:24.625 -->
00:34:31.585And, but it does come with more reliability and more bet on the future. So far, I haven't played that one,
310
00:34:31.985 -->
00:34:38.540like that thing play out primarily because even though benchmarks are getting saturated, models are getting commoditized.
311
00:34:38.540 -->
00:34:39.420Even now,
312
00:34:39.819 -->
00:34:42.700every new big release that comes out in models
313
00:34:42.780 -->
00:34:45.260does improve overall performance
314
00:34:45.579 -->
00:34:50.915in many things a lot, right? So having your own small language model is
315
00:34:50.994 -->
00:34:59.155seen as a bit of a maintenance project by enterprises and they are a little afraid to do it. I see it as an opportunity for
316
00:34:59.750 -->
00:35:04.390vendors like us, like Symphony AI to make it easy for them, where
317
00:35:04.869 -->
00:35:09.270they don't have to worry about it. And this is where I was saying this, if
318
00:35:09.510 -->
00:35:10.630we go from
319
00:35:10.790 -->
00:35:18.935models to the learning capabilities sitting in agentic systems as memory, memory upgrades and intelligent
320
00:35:19.015 -->
00:35:19.815invocation
321
00:35:19.815 -->
00:35:26.775of the right memory updates to the right memory. If that agentic harness can play a lot of that magic,
322
00:35:26.855 -->
00:35:28.935then everything will get simplified.
323
00:35:29.319 -->
00:35:34.839And so I think that is a more practical approach in my mind. Of course, if RL
324
00:35:34.839 -->
00:35:36.839learning systems become
325
00:35:36.839 -->
00:35:45.865very easy for every enterprise to operate with, I think, you know, what you said can play out. But right now, it it feels feels a little far out.
326
00:35:46.505 -->
00:35:47.705Regarding the
327
00:35:47.944 -->
00:35:51.065ways in which you're seeing organizations
328
00:35:51.065 -->
00:35:54.505actually invest in these self improving AI
329
00:35:54.505 -->
00:35:55.464systems,
330
00:35:55.464 -->
00:36:11.030what are some of the common patterns that you're seeing play out where it sounds like reinforcement learning and fine tuning are maybe further out or not as widely adopted, but just wondering how you're seeing people think through the levels of investment and sophistication
331
00:36:11.030 -->
00:36:13.430that they're willing to operationalize.
332
00:36:14.325 -->
00:36:15.605Yeah. I
333
00:36:15.605 -->
00:36:22.885think one very clear pattern is everyone's realizing is that I need to form any self learning
334
00:36:23.045 -->
00:36:23.925system,
335
00:36:24.244 -->
00:36:34.900the environment setup has to be right. And so it takes a lot to set up the right environment. It's that you have to start at all the way from the input data ingestion,
336
00:36:35.060 -->
00:36:47.595all the way to the action layer, and then getting the human feedback right. And and I think I see a lot of enterprises putting efforts in at least getting that right. Like getting the data right for this environment
337
00:36:47.595 -->
00:36:48.955to be ready.
338
00:36:49.115 -->
00:36:52.075And I see a lot of vendors asking questions,
339
00:36:52.155 -->
00:36:55.950rightly so. I see a lot of enterprises asking questions
340
00:36:56.110 -->
00:37:03.630of like, you know, everybody has agents. How is your agent better than the other? How is it learning? And when we go and we say,
341
00:37:03.950 -->
00:37:12.375you know, we show them how it learns, but we ask for it. How are we going to get this feedback back? Right? Is this feedback getting captured somewhere? We would like to
342
00:37:12.615 -->
00:37:20.535implement the right hooks in your systems for that feedback to flow in. And I think customers are receptive to that. They realize that
343
00:37:20.930 -->
00:37:24.370just digitizing that whole process from input
344
00:37:25.170 -->
00:37:27.330data capture, trigger capture,
345
00:37:27.490 -->
00:37:28.850to actions
346
00:37:29.010 -->
00:37:30.370being digitized,
347
00:37:30.370 -->
00:37:32.210to feedback loops
348
00:37:32.290 -->
00:37:33.090being
349
00:37:33.170 -->
00:37:35.730very clear kind of digital entities,
350
00:37:36.265 -->
00:37:37.145if you will,
351
00:37:37.545 -->
00:37:38.025as
352
00:37:38.665 -->
00:37:53.310sort of a task action being taken and a human feedback being written, etc. They are IC enterprises even being ready to put a human judge in the loop, or at least an LLM judge in the loop. And that is something as a property they want to own to convert
353
00:37:53.390 -->
00:38:08.965whether the task was right or not into an input format that the LLMs can use for feedback. So I think the readiness of, you know, kind of readying the enterprise for this environment is one pattern I see. Second pattern is where we play in, is that
354
00:38:09.205 -->
00:38:11.205building the right implementation
355
00:38:11.205 -->
00:38:25.290of the context layer, building the right implementation of the memory layer, and the fact that it sits in your file systems or as knowledge graphs or as something else. And I think that is becoming a differentiator
356
00:38:25.450 -->
00:38:26.090and
357
00:38:26.330 -->
00:38:43.825that's where the right architecture questions are getting asked when we go to sell inside enterprises and so on. So overall, I think the enterprises are getting ready for agents operating, a lot of agents operating. They know that they have to get their data and APIs and and just the whole feedback
358
00:38:44.385 -->
00:38:49.900loop ready. And I feel like that's practically speaking, that's that's where the efforts are right now.
359
00:38:50.460 -->
00:38:53.020As you're talking to these organizations
360
00:38:53.340 -->
00:38:54.780and enterprises
361
00:38:54.780 -->
00:39:05.545about investing in these self improving AI systems and really pushing forward in a focused and concerted manner on agentic capabilities,
362
00:39:05.545 -->
00:39:10.905what are some of the pitfalls or hidden costs that you
363
00:39:11.145 -->
00:39:14.025need to explore before they are able to actually
364
00:39:14.480 -->
00:39:18.560put these things into production in a repeatable and reliable fashion?
365
00:39:19.200 -->
00:39:22.720Yeah. I think it's full of landmines and not just pitfalls
366
00:39:22.960 -->
00:39:28.880once you go into putting these systems in place. I think the first thing you realize is, yes, there's databases,
367
00:39:29.335 -->
00:39:31.415there is ERPs and CRMs,
368
00:39:31.495 -->
00:39:34.935then there is policies and standard operating procedures.
369
00:39:35.015 -->
00:39:37.655Essentially, you're trying to see how do humans
370
00:39:37.895 -->
00:39:43.299run these complex business processes today. The first thing you realize is enterprises
371
00:39:43.299 -->
00:40:04.525think that they are running as per that policy or that as per that standard operating procedure, but they are not. So there's lots of policy gaps. So the reality is humans over time have found a way to work around those policy gaps and have developed a tribal knowledge around that. But agents fail with those policy gaps. And so how do you fill
372
00:40:04.685 -->
00:40:06.685that domain knowledge of
373
00:40:07.005 -->
00:40:08.045for the
374
00:40:08.125 -->
00:40:12.230LLM inside or the brain inside to say,
375
00:40:12.470 -->
00:40:18.870know, I did everything right as per what the policy said and yet my outcome was wrong. Well, was wrong because
376
00:40:18.950 -->
00:40:34.425there are gaps in the policy and there are hidden processes in every enterprise where they found ways to fill that gaps. Those processes could be people oriented like escalation paths, etc. Those those gaps could be just tribal knowledge built over
377
00:40:34.585 -->
00:40:39.705the type of scenarios that that you operate in, etc. But that data is just
378
00:40:40.025 -->
00:40:48.320isn't captured anywhere. And so but the there is no way to fill that also like that data just doesn't exist historically.
379
00:40:48.320 -->
00:40:52.800So you have to kind of start at day zero and build that build that
380
00:40:53.040 -->
00:40:56.485knowledge graph or build that sort of tribal knowledge
381
00:40:56.805 -->
00:40:58.965layer, if you will, in a
382
00:40:59.685 -->
00:41:10.180way that the LLMs can operate on it and build on it. And I think that's the hardest one to get right. The second one is, for these things to be truly truly actionable.
383
00:41:10.260 -->
00:41:14.340There's a lot of gaps and action steps like not everything
384
00:41:14.420 -->
00:41:20.260exists. Not every action is in an executable like API like format or you
385
00:41:20.260 -->
00:41:25.515know, or in a digital format. It is in a set of steps which goes through its own processes
386
00:41:25.595 -->
00:41:29.035of, you know, email channels or, you
387
00:41:29.355 -->
00:41:36.715know, some other formats of executing on it. So it is hard to capture the end to end loop because it is very fragmented.
388
00:41:36.970 -->
00:41:45.130And so in some sense the right way to attempt it is to first automate the sub processes
389
00:41:45.210 -->
00:41:49.130and to get that right while you figure out how to
390
00:41:50.184 -->
00:41:55.305digitize or get right the integration gaps across the set processes.
391
00:41:56.265 -->
00:41:59.065And as you have been working with these businesses
392
00:41:59.065 -->
00:42:03.880through that overall adoption curve and as they start to operationalize
393
00:42:03.880 -->
00:42:17.560and rely on these systems, what are some of the most interesting or innovative or unexpected ways that you've seen them either apply self improving AI or ways that you've seen them approach the creation of the systems that support that capability?
394
00:42:18.805 -->
00:42:20.085Yeah, I think we
395
00:42:20.565 -->
00:42:24.965we have a lot of success stories. I think one thing we pride ourselves
396
00:42:24.965 -->
00:42:27.125in is to go
397
00:42:27.285 -->
00:42:44.160from zero to production in the least amount of time, right, to And the way we are even in a POC to go from having signed a POC to going live on some kind of results at the fastest sort of time. So I think what
398
00:42:44.160 -->
00:42:45.280has been really
399
00:42:45.920 -->
00:42:47.359great to see is
400
00:42:47.575 -->
00:42:58.695what we bring once we go into a use case. We, as I said, we are a vertical AI company. We operate in few verticals where we know the end to end use cases perfect. So we have a process graph for it already.
401
00:42:58.935 -->
00:43:09.650We have the underlying entity graphs for what data is needed, what are the relationships between them and so on. And we now have agentic capabilities.
402
00:43:09.890 -->
00:43:13.010Once we map it to the customer's data,
403
00:43:13.170 -->
00:43:31.520we have agents that can map it, like do this mapping from their entity graph to our entity graph or rather an industry entity graph. And we can map it from their data to a process graph very fast. So there is an agent that can do that. That really speeds up, that kind of standardization
404
00:43:31.840 -->
00:43:44.560really speeds up the path to value that these enterprises are getting. Otherwise there is a whole discovery phase and coming up with what now you can only do it when you're operating in this vertical slices
405
00:43:44.355 -->
00:43:54.915in the use case you're aware of. If you start with like a general problem in an enterprise with which can go in n number of paths, it's a little hard to do it. So for us, that's been
406
00:43:55.315 -->
00:43:56.595practically speaking,
407
00:43:56.994 -->
00:44:01.510very fruitful exercise to develop these standard industry
408
00:44:01.510 -->
00:44:02.550or vertical
409
00:44:02.630 -->
00:44:10.070knowledge graphs and process graphs we land in day one when the enterprises that we go into. And on the flip side, on the
410
00:44:10.914 -->
00:44:14.595outcome side, it is actually quite remarkable what,
411
00:44:14.914 -->
00:44:22.275you know, some of these automations help you get to. Like I was saying, there are customers we have in financial services where
412
00:44:22.434 -->
00:44:31.130their L1 and L2 agents and an investigation step have been completely automated by AI. And so L1 agent is primarily,
413
00:44:31.130 -->
00:44:33.770it follows a policy and standard operating procedures.
414
00:44:33.770 -->
00:44:43.565There's no human involved. It takes in that document. It creates an agent for us that would follow it, and it would follow it exactly to the T. And it will get it like 100%
415
00:44:43.565 -->
00:44:51.885right. And so, and then you get to L2 investigator stage where there is some human intuition involved, and we are getting at very, very high accuracies
416
00:44:51.885 -->
00:44:56.960there as well. The ultimate surprising example I'll give you is that
417
00:44:57.360 -->
00:44:58.560by following
418
00:44:58.880 -->
00:45:07.600policies and by coming up with its own judgment of some of the ways these investigations can be done, we are actually finding
419
00:45:07.920 -->
00:45:09.215new detections
420
00:45:09.215 -->
00:45:17.615like banks policies were not getting updated with the latest in what the regulations say. But because these agents are always improving,
421
00:45:17.615 -->
00:45:18.735always discovering
422
00:45:18.735 -->
00:45:27.510new stuff, it started detecting new crime, which the policies were not ready to capture yet. So the self improving loops actually
423
00:45:27.589 -->
00:45:32.630fix the refresh cycle of what how quickly from once a
424
00:45:32.950 -->
00:45:38.869a US regulator passes a new regulation to how quickly it gets to it being caught in real
425
00:45:39.190 -->
00:45:39.829scenarios,
426
00:45:40.315 -->
00:45:57.830that time really shortened. And it was surprising to us, it's surprising to the end customer. But I mean, if you think about it, it does make sense. This is the kind of like you're cutting many layers of human processes in the middle, and that's why you kind of expect it. We just didn't expect it to come so fast.
427
00:45:58.310 -->
00:46:04.630And in your own experience of working in this space and exploring this overall capability
428
00:46:04.630 -->
00:46:05.190of
429
00:46:05.430 -->
00:46:15.945allowing these AI systems to learn and improve from their execution and interactions, interactions? What What are are some some of of the the most interesting or unexpected or challenging lessons that you've learned in the process?
430
00:46:16.745 -->
00:46:19.465I think for me and my teams,
431
00:46:19.625 -->
00:46:23.545we are following kind of two parallel paths. One is our own
432
00:46:23.945 -->
00:46:28.070kind of product building and system building, platform building is
433
00:46:28.549 -->
00:46:31.750a process. It's a software development lifecycle process,
434
00:46:31.910 -->
00:46:33.510and agents are
435
00:46:33.750 -->
00:46:42.565getting more and more autonomous there. And so we see that as a parallel to what our products and our agents and our products are doing. And it's kind of
436
00:46:42.965 -->
00:46:48.245how much autonomy can we get in our software development processes
437
00:46:48.325 -->
00:46:52.310kind of parallels how much autonomy we are able to build in our
438
00:46:52.550 -->
00:46:53.590end customer
439
00:46:53.670 -->
00:47:03.350process automation with agents. And there's a lot of lessons we learn in our internal code writing and PR and testing and DevOps processes
440
00:47:03.705 -->
00:47:23.430that we learn from, that we are able to apply to the end scenarios. And a lot of problems that we face with implementation here are similar to the ones that we face in practice there as well. Creating that right environment that self improves in a coding environment is just as complex as creating that environment in
441
00:47:23.990 -->
00:47:28.230a practical example. So I think without answering your question specifically,
442
00:47:28.230 -->
00:47:36.255I think that's been the most interesting journey for us. So in some sense, we are improving our own selves as developers
443
00:47:36.255 -->
00:47:38.575while we are learning how to improve
444
00:47:38.975 -->
00:47:46.095autonomy for the industries we operate in. When you're thinking about the application
445
00:47:46.095 -->
00:47:48.920of AI for a particular problem,
446
00:47:48.920 -->
00:47:50.359what are situations
447
00:47:50.359 -->
00:48:01.559where you would advise against the investment in those self reinforcing loops where you're fine just using an out of the box LLM or out of the box predictive model?
448
00:48:02.515 -->
00:48:06.115Yeah, that's a question we ask almost
449
00:48:06.115 -->
00:48:20.030every day, every use case we go into. I think there are levels of this question as well, right? Should you use an LLM? Yes. Should you use a large language model, small language model? You you should probably use a very small language model
450
00:48:20.190 -->
00:48:27.870in many cases and doesn't have to be large for the kind of task you are on. So I think a awareness of the task complexity
451
00:48:28.110 -->
00:48:39.045and whether an LLM can do it, a smaller language model can do it, or does it need a LLM in a loop to do it? That is the key kind of categorization.
452
00:48:39.365 -->
00:48:45.045Agents are effectively LLM in a loop with certain context and clever ways of
453
00:48:45.605 -->
00:48:51.290changing that context, etc. Because we operate in these specific use cases in the industries,
454
00:48:51.450 -->
00:48:53.930the tasks are very, very clear to us.
455
00:48:54.250 -->
00:48:59.130So for example, if you were to do a web research summary,
456
00:48:59.130 -->
00:49:07.715we actually have a small model for it. And it's like, you don't need to burn tokens on a very large language model there. When you are operating
457
00:49:07.795 -->
00:49:10.115in a more undefined domain
458
00:49:10.275 -->
00:49:11.155where,
459
00:49:11.395 -->
00:49:13.795you know, the agent is making decisions
460
00:49:14.319 -->
00:49:17.359with a lot of different contexts and variables,
461
00:49:17.599 -->
00:49:24.480it should probably be thinking or it should be reasoning. And so for us, every step of the process automation
462
00:49:24.640 -->
00:49:44.849is fairly well understood how deep it has to go in reasoning or thinking or how what kind of a task of kind of natural language understanding kind of task is it. And some of those tasks are not that hard. Some of those tasks actually don't even need an LLM. And so we have that fairly well laid out. The recommendation
463
00:49:44.849 -->
00:50:18.740I would give it's it's a bit of an art to figure out which task requires what level of thing. But but in general, you would say if it's the same task over and over again, and if it can be done by deterministic code, you can use an LLM to write that code, but just then save it, right? Just don't use an LLM again and again to do it. Just write an LLM once, save that code, execute that code again and again. If the input changes and if the decisions have to change again and again, that's when you have to decide whether you truly need an agent. So in our cases,
464
00:50:19.140 -->
00:50:29.700the truly hard, like you look at some of the truly hard scenarios of root cause analysis when some systems are going down in a plant and it has to go into
465
00:50:30.165 -->
00:50:48.090figuring out previous context like this, exploring the path of whether similar things have happened and, you know, making judgments based on that. When it's a complex task like that, something which takes humans also hours and days today, quite obvious you have to look into kind of agentic systems.
466
00:50:48.650 -->
00:50:55.050And when you were talking about that aspect of trying to avoid burning tokens unnecessarily,
467
00:50:55.210 -->
00:50:58.810that also brings up another axis of self improvement
468
00:50:58.810 -->
00:51:00.625of cost optimization
469
00:51:00.704 -->
00:51:04.545where maybe one of your objectives is to minimize
470
00:51:04.545 -->
00:51:08.464the expense of a particular agent use case or,
471
00:51:08.625 -->
00:51:17.810obviously, cost optimization on the infrastructure side is a separate question. But in terms of the agent itself or the AI system itself optimizing its own efficiency,
472
00:51:17.810 -->
00:51:34.775I'm wondering what you're seeing as far as capabilities on that horizon as well. Yeah, that's a great question. I mean, are facing that every day in with our cloud code and cursor costs. I wish that's my favorite feature from Cursor team if they could build a self improving
473
00:51:35.095 -->
00:51:36.615cost optimizer.
474
00:51:37.255 -->
00:51:53.020But I think it probably goes against their business model to try to do that. I think really the to be fair we haven't invested much in the area but you could see it in the horizon. As I said we do use many small language models in many of our tasks. I think the
475
00:51:53.180 -->
00:52:02.935the real challenge is how repeatable of a task is it Like how exactly the same pattern you see the next time you do it. And the
476
00:52:03.095 -->
00:52:06.454great thing about LLMs and bigger models is their
477
00:52:06.535 -->
00:52:07.255how
478
00:52:07.815 -->
00:52:16.180across task they are great at, right? So if your task just changes parameters a little bit, they will adapt to it, or rather they have already shown
479
00:52:16.260 -->
00:52:18.020good results in
480
00:52:18.020 -->
00:52:21.860many other similar tasks and so on. So the task specificity
481
00:52:21.860 -->
00:52:23.540versus
482
00:52:23.025 -->
00:52:30.224how general you want to go typically defines it. Wherever we can achieve the levels of task specificity
483
00:52:30.385 -->
00:52:32.385where, you know, it's
484
00:52:32.385 -->
00:52:49.080hard for it to deviate too much. I think that's where we should just optimize the hell out of it and and go to smaller models. You know, I'll tell you, practically speaking in the industry, we haven't built systems which try a whole lot of models. Like, I am very surprised by how good the Gemma
485
00:52:49.080 -->
00:52:51.845series of models are from Google.
486
00:52:52.085 -->
00:53:08.260And like, I see hardly anyone in the industry trying Gemma for many of their tasks. Right? And even within Gemini versus Gemma, like people don't even try Gemma. And so it's a little bit of a, I think we should build something that
487
00:53:08.420 -->
00:53:14.100once you've defined a task and an eval, you give it to an optimizer and it tries all these cheap models
488
00:53:14.260 -->
00:53:34.065and it says where it gets the most accuracy and it's able to achieve it. I think the question always comes, what happens when there is a drift in the input a little bit? Are you ready to absorb that? And I think with smaller models that has been a risk, but it's coming. I think the amount people are spending on coding agents,
489
00:53:34.420 -->
00:53:38.740it's getting quite crazy out there. Just, I think coding agents
490
00:53:38.740 -->
00:54:06.510will drive that cost optimization first. I've already seen a lot of our developers trying to use OLAMA and local models and trying to use cloud only for thinking and using a local model for writing the actual code. And so I think people have begun to develop these systems for wherever they are facing these cost pressures. I think the economy and the cost pressures and product margins will drive many of these things. I think right now people are just trying to get their
491
00:54:06.830 -->
00:54:07.950use cases
492
00:54:08.030 -->
00:54:10.670and automation right with high reliability.
493
00:54:10.830 -->
00:54:18.030Right after this will come the cost aspects. I think everyone feels they can reduce the cost, so they're kind of delaying it to
494
00:54:18.270 -->
00:54:23.145one year onwards, but it's coming. It's coming for sure. And
495
00:54:23.305 -->
00:54:29.545as you continue to work in this space and monitor the evolution of the ecosystem,
496
00:54:29.545 -->
00:54:32.025what are some of the ways that you anticipate
497
00:54:32.025 -->
00:54:34.700the tooling and substrates
498
00:54:34.700 -->
00:54:36.700and agentic frameworks
499
00:54:36.700 -->
00:54:45.660adapting to these concepts of self improvement and making the actual execution and integration of the supporting systems easier to do?
500
00:54:46.555 -->
00:54:52.475I think it's very, very hard to predict. That's the true honest answer. I think what we can see is
501
00:54:52.715 -->
00:55:08.060there'll be companies like us at Symphony AI who will build very good performance, reliable, agentic systems in some industries we are in. And our hope is that that results in a wide adoption in the industries we are in. But the enterprises will realize that the stacks are converging.
502
00:55:08.220 -->
00:55:08.619There is
503
00:55:09.259 -->
00:55:22.915we get a lot of requests from our customers that can you help us besides the use cases you are in, can you help us in our company in standardizing ways in which we should see our agentic systems? I think the
504
00:55:23.155 -->
00:55:24.835layers of data,
505
00:55:24.835 -->
00:55:34.109I think we already see a fair bit of standardization in, you know, MCP servers and the MCP protocols standardizing interfaces to agents.
506
00:55:34.109 -->
00:55:39.790I haven't seen that much pickup in A2A like protocols between multi agent
507
00:55:39.950 -->
00:55:40.589interoperability,
508
00:55:41.184 -->
00:55:41.825but,
509
00:55:42.785 -->
00:55:53.184you know, with systems like OpenClaw and all getting popular, there will be some standardization in the agent control planes as well that enterprises see, I think like with many agents running,
510
00:55:53.720 -->
00:55:59.320who's governing and controlling policies across those, Clearly that is emerging as an area. I
511
00:55:59.559 -->
00:56:02.440think the tooling is getting very, very standardized
512
00:56:02.680 -->
00:56:03.720going forward.
513
00:56:04.760 -->
00:56:13.875What is hard to predict is whether it'll be Postgres databases as the agent preferred layer or whether it'll be file systems and like how
514
00:56:14.835 -->
00:56:16.115do these things change?
515
00:56:18.115 -->
00:56:23.950That's quite evolving and that's hard to predict. But the concept of treating agents
516
00:56:23.950 -->
00:56:24.590as
517
00:56:24.910 -->
00:56:27.710like the agent lifecycle management
518
00:56:27.869 -->
00:56:30.510has become fairly standardized,
519
00:56:30.510 -->
00:56:31.950like just like model
520
00:56:32.109 -->
00:56:34.109ops and model lifecycle management.
521
00:56:35.105 -->
00:56:38.065And so those things are getting standardized. I think it's
522
00:56:39.265 -->
00:56:41.984a question of agents going into production,
523
00:56:42.305 -->
00:56:43.345creating value,
524
00:56:43.585 -->
00:56:47.105and then everything under them will start getting standardized in its layers.
525
00:56:47.880 -->
00:56:54.839All right. Are there any other aspects of this aspect of self improving AI systems or the ecosystem
526
00:56:54.839 -->
00:56:55.560of
527
00:56:55.640 -->
00:56:59.799AI applications that we didn't discuss yet that you'd like to cover before we close out the show?
528
00:57:00.625 -->
00:57:03.265I think for me, it's very interesting.
529
00:57:04.305 -->
00:57:05.265Mean, ultimately,
530
00:57:05.425 -->
00:57:10.625it's known to everybody that the foundation model companies are that reinforcement
531
00:57:10.625 -->
00:57:11.425learning
532
00:57:11.745 -->
00:57:13.185with real
533
00:57:13.185 -->
00:57:14.385world environments
534
00:57:14.670 -->
00:57:15.310and
535
00:57:15.710 -->
00:57:17.790especially with verifiable systems
536
00:57:17.790 -->
00:57:22.510is a big area of investments and models are improving at different domains
537
00:57:22.670 -->
00:57:31.865with more and more RL environments being created and so on. What's very interesting to me is if that same subset of capability
538
00:57:31.945 -->
00:57:39.465can come very quickly to enterprises in a way they can harness it. And so for just my business process,
539
00:57:39.625 -->
00:57:42.985can I get kind of that same level of capability?
540
00:57:43.225 -->
00:57:44.505Start with a smaller
541
00:57:45.065 -->
00:57:46.520model and
542
00:57:46.520 -->
00:57:49.720how do I apply the same reinforcement
543
00:57:49.720 -->
00:57:50.920learning steps
544
00:57:51.000 -->
00:57:54.120without needing to have a research scientist employed
545
00:57:54.280 -->
00:57:56.600in my company, right? But
546
00:57:56.920 -->
00:57:59.160knowing that the process is
547
00:57:59.640 -->
00:58:00.520dynamic,
548
00:58:00.599 -->
00:58:02.055knowing that if
549
00:58:02.535 -->
00:58:25.820I follow the process in one way, I know it is suboptimal. If I follow it another way, is better. How can I turn it into a reward for the model without knowing what rewards mean and RL means and all that? If we can simplify that, I think a lot of businesses, and that's an area we are trying to go deep into as well. I think that will enable companies to own sort of reasoning layers
550
00:58:26.460 -->
00:58:31.260of their own without relying on the big bottle companies. So I'm excited how this area
551
00:58:32.060 -->
00:58:32.780emerges forward.
552
00:58:33.635 -->
00:58:49.580All right. Well, for anybody who wants to get in touch with you and follow along with the work that you're doing, I'll have you add your preferred contact information to the show notes. And as the final question, I'd like to get your perspective on what you see as being the biggest gaps in the tooling technology or human training that's available for AI systems today.
553
00:58:50.140 -->
00:58:53.100Biggest gap in tooling and technology.
554
00:58:53.500 -->
00:58:57.740I think it's, I would probably say the other way. I think it's with
555
00:58:58.380 -->
00:59:08.415the way Cloud Code and some of these other agents are evolving. It is very easy to fill whatever gap exists, right? I think the gap I'm more worried about is
556
00:59:09.775 -->
00:59:10.895the integration
557
00:59:10.895 -->
00:59:11.454steps
558
00:59:12.140 -->
00:59:15.500when it needed to apply this to real practical
559
00:59:15.900 -->
00:59:20.140industries, real practical example. I think there is no clear
560
00:59:20.540 -->
00:59:23.980integration outline. Like everything has to be done
561
00:59:24.300 -->
00:59:29.785differently for every implementation that you go to. And so if there is a
562
00:59:30.425 -->
00:59:32.825way to use agents
563
00:59:32.825 -->
00:59:37.785to go and discover these processes and create like templates
564
00:59:37.785 -->
00:59:38.505of,
565
00:59:39.065 -->
00:59:44.490I think effectively we've been talking about digital twins for a while in the industry,
566
00:59:44.650 -->
00:59:57.845and there have been several attempts at it. We have our own attempt at it at an industry perspective in Symphony AI, but a true recognition of a digital twin in a company as a tooling, as an integration layer is
567
00:59:58.005 -->
01:00:04.405not there. And if it was there, then agents would onboard onto it very. But I think on the flip side,
568
01:00:04.565 -->
01:00:11.590you put Cloud Code in an environment in a company, you give it access to various live infra and resources,
569
01:00:11.670 -->
01:00:13.350and it starts to figure out.
570
01:00:13.670 -->
01:00:17.270I think we've got the best tooling ever in the history of
571
01:00:18.470 -->
01:00:27.535mankind and what developers had, especially with these AI tools. So I'm actually seeing the picture as very, very optimistic on
572
01:00:27.855 -->
01:00:31.295what we can do to fix the gaps and tooling where it exists.
573
01:00:32.095 -->
01:00:43.400All right. Well, thank you very much for taking the time today to join me and share your experiences and insights into how to build these AI systems in a way that they can continually evolve and improve
574
01:00:43.480 -->
01:01:02.315and some of the safety considerations around how to make sure that they stay well aligned with the organization's objectives. It's a fascinating and fast moving space. So I appreciate you taking the time to help share some of the expertise that you've developed through working through the hard bits and hope you enjoy the rest of your day. Thank you. Thank you for having me.
575
01:01:15.295 -->
01:01:23.215Its community, and the innovative ways it is being used. And the AI Engineering Podcast is your guide to the fast moving world of building AI systems.
576
01:01:23.695 -->
01:01:33.790Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email hosts@dataengineeringpodcast.com
577
01:01:33.790 -->
01:01:39.950with your story. And to help other people find the show, please leave a review on Apple Podcasts and tell your friends and coworkers.