OVER DEZE AFLEVERING
Coding agents are becoming more capable, but getting better results increasingly depends on the systems, interfaces, and feedback loops we build around them.
In this episode of the Superlinear podcast, we unpack four emerging practices in agentic engineering:
→ Shifting left in your agent harness
→ Using precise expert language to communicate intent
→ Choosing deliberately between MCP, CLI tools, Bash, and Code Mode
→ Using test oracles to help agents explore, verify, and reproduce complex behavior
We discuss ideas from AI Native DevCon London, and explore how better guides and sensors can improve a harness, why expert vocabulary helps you access more of an LLM’s capabilities, when code is a better interface than individual tool calls, and how working examples can communicate intent more precisely than a written specification.
The broader takeaway: building effectively with AI is not only about choosing a better model or writing a better prompt. It is also about designing a better environment for the agent to work within.
Superlinear is a podcast about emerging practices for building with AI.
Hosted by Brandon Kase and Christine Yip.
Homepage: https://superlinear.fm
Twitter: https://x.com/superlinear_fm
LinkedIn: https://www.linkedin.com/company/superlinearfm
Brandon Kase: https://x.com/bkase_
Christine Yip: https://x.com/christinetyip
Music licensed through Soundstripe. Code: O2UT6CCEQNH4KT0D
LAAT NOTITIES ZIEN 🔗
TRANSCRIPTIE 🔗
00:00:00.160 --> 00:00:07.759
Essentially, LLMs are experts in everything, but by default, you get the average.
00:00:08.160 --> 00:00:13.199
People building things in a space and they're like, I just built this in two days using AI.
00:00:13.359 --> 00:00:16.800
Did they build it from scratch or did they use an Oracle for it?
00:00:16.960 --> 00:00:18.320
Using an Oracle could be a shortcut.
00:00:18.719 --> 00:00:24.399
If something is triggering feedback, a self-correction, then it means that the model did something wrong.
00:00:24.559 --> 00:00:26.879
It's better if the model just doesn't do it wrong in the first place.
00:00:27.199 --> 00:00:32.799
If you want to have good taste, you definitely also want to have deeper understanding.
00:00:33.280 --> 00:00:39.759
Welcome to Super Linear, a podcast about emerging practices for building with AI agents.
00:00:40.000 --> 00:00:44.560
Each episode, we cut through the noise to find one new angle worth trying.
00:00:44.640 --> 00:00:52.159
Then we unpack why it matters, where it works, and how you can apply it in your own workflows so you can build more powerful things.
00:00:52.640 --> 00:00:57.280
We're your hosts, Christine Yet and Brandon Case.
00:00:57.679 --> 00:00:58.640
Hey everyone.
00:00:58.880 --> 00:01:04.799
Today we're breaking down four emerging practices that Brandon and I have noticed across the AI space.
00:01:05.040 --> 00:01:13.280
By the end of this episode, you'll start thinking of your harness more as an evolving system, one that is constantly shifting left.
00:01:13.519 --> 00:01:15.920
We'll explain what that means later in this episode.
00:01:16.640 --> 00:01:20.879
You'll also have a greater appreciation for using precise expert language.
00:01:21.519 --> 00:01:30.239
You'll think carefully before defaulting to MCP, and you will understand when you should reach for an Oracle as a tool for building more effectively.
00:01:31.200 --> 00:01:37.599
We re-recorded these takeaways right after the most recent AI native DEF CON in London.
00:01:37.840 --> 00:01:42.239
And we hope they give you a few new ideas to try in your own workflows.
00:01:42.640 --> 00:01:52.560
So the first emerging practice that we're going to talk about today was mentioned by Ryan Lopopolo from OpenAI at the AI Native DEF CON where you were last week in London.
00:01:52.799 --> 00:01:56.000
So, and that was leverage shifting left for agents.
00:01:56.079 --> 00:01:56.239
Yeah.
00:01:56.400 --> 00:01:56.959
Is that right?
00:01:57.200 --> 00:01:57.519
Okay.
00:01:57.840 --> 00:02:03.200
Can you tell us a little bit more about this emerging practice related to shifting left for agents?
00:02:03.439 --> 00:02:03.680
Yeah.
00:02:03.840 --> 00:02:05.599
So I I think this is interesting.
00:02:05.760 --> 00:02:13.840
This was one line that Ryan said sort of offhand during his presentation, which surprised me, and I had to ask a question about it at the end.
00:02:14.000 --> 00:02:20.080
So I want to take you all through the story of this point because I think it's interesting.
00:02:20.400 --> 00:02:22.240
What is the sentence that he mentioned?
00:02:22.639 --> 00:02:28.719
Well, let me share my screen because I I tweeted it so I can uh just show you.
00:02:29.120 --> 00:02:32.240
This wasn't his exact words, and maybe I misquoted him a little bit.
00:02:32.319 --> 00:02:46.719
But but the the sentence was something like you don't need to shift left as much, which was a good practice in software engineering, because harness engineering, which is like this new thing that a few people have talked about, which we can define a little better.
00:02:47.039 --> 00:02:50.560
Did that caught your attention because it's very controversial?
00:02:50.800 --> 00:02:54.319
Was shifting left like super important in software engineering?
00:02:54.560 --> 00:02:54.800
Yeah.
00:02:54.960 --> 00:02:59.680
Shifting left is um generally seen as a good practice.
00:02:59.759 --> 00:03:12.080
And in particular, it's a practice that I thought that Ryan really took advantage of in the stuff that he talks about in his presentations and and his blog posts.
00:03:12.240 --> 00:03:16.000
Um so that's that's in particular why this was very surprising to me.
00:03:16.319 --> 00:03:23.280
Um and also I I didn't believe it uh until I until I dug in and and sort of understood what what he was saying.
00:03:23.840 --> 00:03:24.240
Interesting.
00:03:24.639 --> 00:03:29.840
Before we dive into it, could you like explain maybe like in a few sentences what is shift left?
00:03:30.000 --> 00:03:31.280
So we're all on the same page.
00:03:31.599 --> 00:03:34.000
Yeah, so I'm gonna pull up this nice image.
00:03:34.240 --> 00:03:36.479
Shifting left, uh, let's see.
00:03:36.639 --> 00:03:45.199
Well, think about time on the x-axis, and uh what we're measuring is the time across the software or product development lifecycle.
00:03:45.360 --> 00:03:56.560
So, you know, you start with understanding the user's needs, requirements, gathering, architecture, design, coding, and then there's all sorts of testing, and then the product is released.
00:03:56.800 --> 00:04:05.120
The idea of shifting left is taking tests or errors that happen on the right hand side and push them as left as possible.
00:04:05.280 --> 00:04:28.000
And in the context of modern software engineering practices that I guess I believe, and and I think Ryan exemplifies in his some of the content that he's created, it's maybe focused on the testing side and it's moving tests that are slower and that run less often to tests that can run really quickly and at very close to the point that you're writing the software.
00:04:28.160 --> 00:04:35.759
You know, this intuitively is important with coding agents because coding agents are writing a lot more code and they're doing it fast.
00:04:35.920 --> 00:04:44.800
And so the, you know, you'd you'd think, you'd expect, and it's generally good practice, to get feedback to those coding agents as quickly as possible.
00:04:45.199 --> 00:04:49.759
Because then they don't go down the wrong route for too long and waste tokens.
00:04:50.000 --> 00:04:50.240
Exactly.
00:04:50.480 --> 00:04:50.720
Okay.
00:04:50.879 --> 00:04:54.720
But why does Ryan say like that's not as important anymore?
00:04:54.879 --> 00:04:56.319
It still makes sense to me to do it.
00:04:56.560 --> 00:04:56.800
Yeah.
00:04:56.959 --> 00:05:11.120
So so so Ryan is talking about this in the context of of harness engineering, which is a term that he wrote a long blog post about in February, and this went kind of viral, which kind of described how his team is creating software with coding agents.
00:05:11.279 --> 00:05:22.079
And I think this was again further described and elaborated upon Bergitta Burkler, who also gave a keynote presentation at the conference last week on the same topic.
00:05:22.319 --> 00:05:32.879
And so I guess I want to just describe harness engineering a little bit to give some context to help understand why Ryan claims that shifting left is less important.
00:05:32.959 --> 00:05:33.120
Okay.
00:05:33.680 --> 00:05:33.839
Yeah.
00:05:34.079 --> 00:05:37.920
So harness engineering, um, well, it's about engineering the harness.
00:05:38.160 --> 00:05:44.079
And the harness refers to everything that's not the model in the context of the coding agent.
00:05:44.240 --> 00:05:47.120
So the model is the thing that we're doing inference on.
00:05:47.279 --> 00:05:53.680
Then part of the harness is the actual software that you're using to control the coding agent, and maybe it's plugins.
00:05:53.759 --> 00:05:57.439
So for example, Cloud Code or Codex or Pi or Open Code.
00:05:57.519 --> 00:06:03.040
Uh, and it's also, you know, which plugins you have installed, what your system prompt is, these kinds of things.
00:06:03.199 --> 00:06:15.759
And then it also encompasses what uh Brigitte calls the user harness, which are made up of, here's another chart, are made up of feed forward and feedback controls, which she calls guides and sensors.
00:06:16.399 --> 00:06:25.920
So the guides are things that you put in markdown files that feed into the context of the agent ahead of time, or or maybe the the code of your project.
00:06:26.079 --> 00:06:30.480
So it's anything that goes in the context before the coding agent starts writing code.
00:06:31.600 --> 00:06:34.800
So like uh agents.md, for example.
00:06:35.040 --> 00:06:35.279
Yes.
00:06:35.600 --> 00:06:41.600
Agents.md, or if your agents.md tells the coding agent to read other markdown files, those would also be included.
00:06:41.759 --> 00:06:49.439
And then sensors are the things that give feedback to the coding agent as it's working and trying to perform its task.
00:06:49.600 --> 00:06:58.480
So these could be tests that run automatically every time the code is edited by a hook or because the agent decides to run it.
00:06:58.720 --> 00:07:19.040
And these these sensors, when when a test fails, for example, the the output of that test failure uh gets included in the agent's inference the next time um the next time the agent runs its loop.
00:07:19.199 --> 00:07:28.079
And so so it's a way to just in time inject the context into the coding agent, which which helps it correct itself if it's going down the wrong path.
00:07:28.319 --> 00:07:28.480
Yeah.
00:07:28.639 --> 00:07:35.040
So the agent is basically a loop and it just goes on until it it thinks it fulfilled its task.
00:07:35.199 --> 00:07:39.600
And in this feedback step, it gives itself more context.
00:07:39.759 --> 00:07:42.639
So in the next loop, it will uh give a better output.
00:07:42.800 --> 00:07:42.959
Yeah.
00:07:44.639 --> 00:08:00.240
So this is this is the harness, and this is uh harness engineering is the idea of thinking carefully and engineering the ways in which feedforward guides and feedback sensors are supplied for whatever project you're building or whatever task you're doing.
00:08:00.959 --> 00:08:03.920
That's I think that's like a high-level, high-level explanation.
00:08:04.079 --> 00:08:07.519
So I think we can go back to this point now.
00:08:07.839 --> 00:08:10.240
So I guess, okay, so why is it surprising?
00:08:10.319 --> 00:08:10.959
Let's start with that.
00:08:11.120 --> 00:08:40.879
It's surprising to me because harness engineering is about carefully engineering your sensors, for example, um, and and shifting left uh is one way shifting a test left from a slow test that runs once in a while or a pull request review by a human and shifting that into um a lint rule that runs every time the code is changed, it's much more efficient for coding agents to get that feedback and correct themselves.
00:08:41.039 --> 00:08:42.320
And so you are more productive.
00:08:42.559 --> 00:09:00.480
So if we're saying like shifting left is not as important anymore since we have hardness engineering, does that mean like, well, because we have these feedback steps, the self-correcting feedback steps, the coding agent will self-correct it anyway, so shifting left becomes less important?
00:09:00.720 --> 00:09:14.480
Um well, shifting left is the is the practice of introducing more of these feedback sensors um uh into your into your harness.
00:09:14.639 --> 00:09:23.360
So uh or at least that's one interpretation, and I think that's the interpretation that that Ryan was using when he was saying that he didn't like it as much.
00:09:24.080 --> 00:09:33.840
Uh and yeah, and and I think intuitively for me that was surprising because it seems it seems valuable, right?
00:09:34.080 --> 00:09:44.320
But I guess his his point when I asked for clarification is well, if something is triggering feedback, a self-correction, then it means that the model did something wrong.
00:09:44.480 --> 00:09:47.440
It's better if the model just doesn't do it wrong in the first place.
00:09:48.240 --> 00:10:06.240
And um and so rather than relying so much on lots and lots of sensors that get triggered all the time, it's better to put more information in the guides, more information in the markdown files that get fed into the the coding agent before it starts its work so that your sensors are triggered less often.
00:10:06.320 --> 00:10:09.440
Because whenever your sensor is triggered, you're spending a little bit of tokens.
00:10:09.600 --> 00:10:13.279
And then there's less tokens for focusing on whatever the task is at hand.
00:10:13.440 --> 00:10:14.320
Does that does that make sense?
00:10:14.559 --> 00:10:16.320
Yeah, yeah, that makes sense.
00:10:16.559 --> 00:10:29.519
So uh maybe sometimes when during when sensors pick up things, maybe some users are like, yes, it found something, good that it it caught it, but Ryan is like, no, if it caught something, then actually it's a failure because your guide wasn't good enough.
00:10:30.000 --> 00:10:33.519
You know, it would it would have been better, of course, if if you one-shotted it.
00:10:33.679 --> 00:10:46.320
But it almost, I don't know, maybe I'm just confused, but um it almost sounds like if you want the guide to be good from the start, that's even more left in the whole process, right?
00:10:46.480 --> 00:10:49.360
So that to me, it also sounds like shifting left.
00:10:49.600 --> 00:10:50.720
Yeah, yeah, yeah.
00:10:50.879 --> 00:10:52.320
I think I think you're right.
00:10:52.480 --> 00:11:03.200
I think, I mean, if we go back to this picture in this particular picture of shifting left, certainly um, I guess you would say the agents.md is kind of like your requirements.
00:11:03.440 --> 00:11:05.360
So in that sense, it is further left.
00:11:05.440 --> 00:11:08.960
But I think so so in some in some sense, yeah, it's actually shifting further left.
00:11:09.120 --> 00:11:20.240
But in another sense, I think like metaphorically, putting information in your guides is like having a conversation with your coworker or like telling them a rule, but it's telling them in a way that is fuzzy.
00:11:20.399 --> 00:11:21.919
It's like they have to remember that rule.
00:11:22.080 --> 00:11:29.679
And so it's important to still have sensors because a human or an LLM can still make mistakes from that information.
00:11:29.840 --> 00:11:32.639
So it's it's sort of, but you know, you are right though.
00:11:32.720 --> 00:11:34.720
It is, it is kind of another form of shifting left.
00:11:34.799 --> 00:11:42.399
And I wish that I could ask Ryan if what what he thinks about this or something or if we're just misunderstanding what he was saying.
00:11:42.559 --> 00:11:43.200
Yeah.
00:11:43.519 --> 00:11:58.399
Um, but but I I I think like the whole, even if this is whether it's whether it is precisely shifting left or not, I I think it's just really interesting to as a point around, you want to make sure that your your system is learning.
00:11:58.559 --> 00:12:09.519
So whenever the coding agent makes a mistake, even if the mistake is self-corrected, it should still you should try and sort of teach the model or sorry, teach your harness by putting that information in the guide.
00:12:09.600 --> 00:12:12.080
I think I think that's that's the the main insight here.
00:12:12.159 --> 00:12:15.200
So this is the this is what the emerging practice is, I suppose.
00:12:15.440 --> 00:12:15.919
Awesome.
00:12:16.080 --> 00:12:16.399
Cool.
00:12:16.559 --> 00:12:17.840
Yeah, thanks for sharing this.
00:12:18.000 --> 00:12:21.519
Also, very curious to listen to Ryan's talk.
00:12:21.679 --> 00:12:26.320
And if we interpreted what Ryan said wrongly, feel free to let us know.
00:12:26.480 --> 00:12:30.399
At least it did set us thinking and gave us a lot of insights.
00:12:30.559 --> 00:12:35.919
And so, yeah, the I believe the talks from ai native.com will be online eventually, right?
00:12:36.000 --> 00:12:36.960
Do you know Brendan?
00:12:37.360 --> 00:12:42.080
Yeah, some of some of the talks, the talks on the main stage are are going to be online, and this one was.
00:12:42.240 --> 00:12:45.279
Both both Ryan's and and Brigitte's were on the main stage.
00:12:45.360 --> 00:12:46.639
So both of those will be online.
00:12:47.279 --> 00:12:50.320
We can also share the links as we upload this recording.
00:12:50.559 --> 00:12:51.840
We'll share it in a description.
00:12:52.080 --> 00:12:52.480
All right.
00:12:52.639 --> 00:12:53.600
Thanks, Brandon.
00:12:53.679 --> 00:12:57.200
Then anything else you want to share before we go to emerging practice number two?
00:12:57.440 --> 00:12:58.159
No, let's move on.
00:12:58.320 --> 00:12:58.720
All right.
00:12:58.960 --> 00:12:59.519
Moving on.
00:12:59.759 --> 00:13:01.200
Emerging practice number two.
00:13:01.600 --> 00:13:04.639
Compress intent with expert language.
00:13:04.799 --> 00:13:05.600
What does this mean?
00:13:05.679 --> 00:13:08.080
And how does what how did this come up for you?
00:13:08.320 --> 00:13:10.879
Yeah, so this came from a tweet.
00:13:11.039 --> 00:13:12.879
So uh I'll pull up that tweet.
00:13:13.039 --> 00:13:18.159
Um, so so there's a a tweet, I guess it was a retweet, a quote retweet.
00:13:18.320 --> 00:13:22.320
So the original tweet was from Emil Kowalski from from Linear.
00:13:22.480 --> 00:13:24.240
And Emil, well, I'll just click on that.
00:13:24.399 --> 00:13:29.519
Emil said, to get good animations from an AI, you need to get good at telling it what you want.
00:13:29.679 --> 00:13:35.919
Stagger this list of items, make this animation direction aware, spatial consistency, crossfade, layout animation.
00:13:36.080 --> 00:13:38.080
I made a motion vocabulary for this.
00:13:38.240 --> 00:13:39.840
And then there's a link to a page.
00:13:40.000 --> 00:13:41.759
So so that was that was Emil's post.
00:13:41.840 --> 00:13:48.480
And then and then uh Guillermo Rausch from Vercell retweeted and said, This is what education in the age of AI looks like.
00:13:48.639 --> 00:13:49.600
Start with the language.
00:13:49.759 --> 00:13:51.600
The linguistic surface is your roadmap.
00:13:51.759 --> 00:13:56.879
Much like an SDK has a set of function definitions, human language is the new API to the world.
00:13:57.039 --> 00:13:58.320
It also always has been.
00:13:58.480 --> 00:14:02.000
There's never been an expert in any field without mastery over its language.
00:14:02.159 --> 00:14:05.759
The distinction is that English alone could not produce tangible things.
00:14:05.919 --> 00:14:10.000
You had to translate it into machine instruction through learning or delegation to others.
00:14:10.080 --> 00:14:11.120
You can now go direct.
00:14:11.279 --> 00:14:12.159
What a beautiful tweet.
00:14:12.480 --> 00:14:14.320
Yeah, and it it makes sense too.
00:14:14.639 --> 00:14:15.440
Yeah, so exactly.
00:14:15.519 --> 00:14:16.559
It does it does make sense.
00:14:16.720 --> 00:14:30.799
It seems it seems sort of obvious, but I think this really resonated with me in particular because I sort of studied this for like 10 years before LLMs came out in a slightly different context.
00:14:30.960 --> 00:14:46.240
Um so so when I when I whenever I read something about learning how to communicate about a topic with with precise language by learning carefully the language of experts in that in that field, it really gets my attention.
00:14:46.639 --> 00:15:00.960
Yeah, because I guess um also the the really good thing here is if you can convey your intents in a very succinct way to an LLM, then the output from the LLM will also be better while you use less tokens, right?
00:15:01.039 --> 00:15:04.080
So that's I guess what uh Emil is saying here.
00:15:04.240 --> 00:15:10.320
Like uh if you're going to build animations with AI, make sure you know the terminology related to animations.
00:15:10.480 --> 00:15:16.080
Um and I guess this this is for every other domain of any use case that people are building.
00:15:16.320 --> 00:15:26.399
For example, if you would be working on front end or building a website, you could either say, I want a pop-up window or a panel that appears on top of the current page.
00:15:26.559 --> 00:15:35.200
It should temporarily take focus, the rest of the page should be darkened or blurred, and visitors should not interact with the page behind it, blah, blah, blah, blah.
00:15:35.360 --> 00:15:39.200
Or you could just say to the LLM, build a modal for me.
00:15:39.360 --> 00:15:43.279
So you just have to know what the term is for the thing that you're trying to build.
00:15:43.440 --> 00:15:48.240
Or if you're building something related to legal or law, then you need to know the language of lawyers.
00:15:48.320 --> 00:15:51.360
And I guess the beautiful thing is LLMs know everything, right?
00:15:51.519 --> 00:15:53.440
They are an expert in many different domains.
00:15:53.679 --> 00:15:54.399
Yeah, exactly.
00:15:54.559 --> 00:15:56.240
So you don't really need to teach it.
00:15:56.399 --> 00:16:04.240
I mean, okay, I suppose if if you're doing something really, really on the cutting edge, maybe you need to give it context with some research papers or something.
00:16:04.320 --> 00:16:12.000
But but essentially, LLMs are experts in everything, but by default, you get the average.
00:16:12.080 --> 00:16:13.360
You get the average of everything.
00:16:13.519 --> 00:16:17.360
You get the average person who doesn't know much about all these things in its in its output.
00:16:17.519 --> 00:16:31.200
So not only are you saving tokens by being precise and using proper language, but you're also coercing the latent space of the model towards uh the expertise in whatever field you're you're talking to it with that correct language.
00:16:31.519 --> 00:16:32.000
Interesting.
00:16:32.159 --> 00:16:35.360
So basically, it's so interesting you're saying the latent space.
00:16:35.440 --> 00:16:44.320
So basically, it almost appears as an everyday coworker, but actually it knows so much more, but it's not really showing it as apparently.
00:16:44.559 --> 00:16:45.200
Yes.
00:16:46.080 --> 00:16:52.960
One other thing I thought was super interesting is that in our previous discussion, you went even a step further.
00:16:53.120 --> 00:16:58.960
In the examples that we just talked about, um, about ML's animation and and building a website.
00:16:59.120 --> 00:17:05.359
Basically, it was just about using the domain language related to the thing that you're trying to build.
00:17:05.599 --> 00:17:11.680
But when we talked previously, you pulled in also another domain when you were trying to build something.
00:17:11.839 --> 00:17:13.279
Could we dive a little bit into that?
00:17:13.519 --> 00:17:14.079
Yeah, yeah.
00:17:14.240 --> 00:17:20.000
So this is this is the thing that I was uh obsessed about for a long time, for many years.
00:17:20.240 --> 00:17:23.680
Um so and it's I guess I wrote about it.
00:17:23.839 --> 00:17:24.720
I'll just pull up.
00:17:24.880 --> 00:17:30.559
I have a a blog post from from 2020, and and there's just a a a paragraph here.
00:17:30.720 --> 00:17:38.079
Well, okay, so just quickly, the idea is using algebraic structures, uh, thinking about algebra when designing software.
00:17:38.240 --> 00:17:40.079
And um, I think it's interesting.
00:17:40.160 --> 00:17:43.839
In in 2020, in my intro, I I I I'll just read this out.
00:17:44.000 --> 00:17:51.359
I wrote, these structures give names to incredibly abstract notions, notions that we otherwise as humans would have a hard time discussing.
00:17:51.519 --> 00:17:54.160
When something has a name, our brains can reason about them.
00:17:54.400 --> 00:17:57.119
Shared vocabulary means more productivity for teams.
00:17:57.279 --> 00:18:04.400
Moreover, using these proper names introduces about a hundred years of mathematics and computer science content for further study.
00:18:04.640 --> 00:18:06.799
That's such a great way to phrase it.
00:18:07.039 --> 00:18:14.720
Basically, with LLMs, because they're so smart, you're trying, there's more opportunity to find shared vocabulary that you and LLM have.
00:18:14.799 --> 00:18:21.200
And because LLMs have such a they're an expert in everything, the shared vocabulary is huge, could be huge.
00:18:21.440 --> 00:18:23.279
So we could exploit more interesting.
00:18:23.680 --> 00:18:34.559
So in like something that I explored is hey, if we use algebra when building software, we can solve hard problems, build good libraries, communicate with our teams effectively, and all these things.
00:18:34.640 --> 00:18:37.279
And this is like doubly or triple E tree with LMs.
00:18:37.519 --> 00:18:40.000
Um, so just to like, I think it's interesting.
00:18:40.240 --> 00:18:44.160
One example that I used like a long time ago is also animations.
00:18:44.319 --> 00:18:47.519
So I thought it'd be interesting for me to just show that quickly.
00:18:47.759 --> 00:18:54.480
So the shared vocabulary that you're using here is not only about animation, but also related to algebra.
00:18:54.559 --> 00:18:58.079
So you're combining terms from both domains in this example.
00:18:58.319 --> 00:18:58.559
Yes.
00:18:58.640 --> 00:19:09.839
Although I will say I I'm definitely not an expert in the animation domain, and I did not study that too well in particular, but certainly from an algebra sense, this example I think does a good job.
00:19:10.559 --> 00:19:21.119
So I'm very curious how we could use algebra for building something else that initially when I hear it, it doesn't like ring a bell for me that it's related to algebra.
00:19:21.200 --> 00:19:22.720
So very curious about this example.
00:19:22.960 --> 00:19:23.200
Yeah.
00:19:23.279 --> 00:19:25.920
So just really quickly, I don't want to spend too much time on it.
00:19:26.079 --> 00:19:49.440
Um I guess anytime you can um think about a problem in terms of uh a small list of primitives that can be combined together in uh in a few ways, um, but you can just keep recombining them to build something really complex, you can take advantage of algebra.
00:19:49.680 --> 00:20:00.799
I guess the metaphor first is like Legos, you know, there's a few different Lego blocks, but even with just a few blocks, if you build them in different ways, you can build cities and towns and castles and all these things.
00:20:01.039 --> 00:20:03.920
So so in this case, let's do the same thing with animations.
00:20:04.079 --> 00:20:10.160
So let's say an animation is our core object, and we can just say an animation is something that happens over time.
00:20:10.240 --> 00:20:12.960
And and uh, you know, but it can be anything that happens over time.
00:20:13.119 --> 00:20:15.200
And then you can combine animations together.
00:20:15.359 --> 00:20:18.160
So this is an exploration of choreographed animations.
00:20:18.240 --> 00:20:22.319
So it's animations that you configure ahead of time in kind of like a movie.
00:20:22.559 --> 00:20:32.799
Um, and and so for example, if you have an animation for growing a circle and you have a different animation for fading out a circle, well, you can run them at the same time, for example.
00:20:32.960 --> 00:20:34.799
And we can call that addition.
00:20:34.960 --> 00:20:43.519
And if you run those two animations at the same time, you get what's on the video here, if you're watching the video, which is a circle that grows and fades out at the same time.
00:20:43.680 --> 00:20:44.559
Okay, makes sense.
00:20:44.720 --> 00:20:52.400
And you could also run animations one after another rather than at the same time, you could have one complete and then start the other one and then complete that.
00:20:52.559 --> 00:20:56.400
And that's sequencing one animation after another, we could call that multiplication.
00:20:56.559 --> 00:21:17.279
And it turns out if you choose multiplication for sequencing and addition for parallel, a lot of the algebraic laws that are true on integers, like addition commutes, one plus two is the same as two plus one, that multiplication distributes over addition, uh, which is like something you might have learned in in school, these things are true for animations as well.
00:21:17.519 --> 00:21:28.799
And what that does at the end of the day is in just a few operations, you can build really complex things and gives give a lot of power to users of software libraries.
00:21:28.960 --> 00:21:33.200
And and this is was true for humans, but it's true for LLMs as well.
00:21:33.359 --> 00:21:41.599
If you can describe a software surface area in less tokens, then you have more context to spend on you know doing whatever you need to do.
00:21:41.759 --> 00:21:47.920
So it's it's I think it's interesting as an example of something that you can apply to, like you were saying, something else.
00:21:48.079 --> 00:21:51.200
So it's animations, and I'm writing software with algebra.
00:21:51.359 --> 00:21:57.839
You could, for example, write software using logic and apply that to law, you know?
00:21:58.000 --> 00:21:59.680
So you can you can mix and match and pull.
00:22:00.160 --> 00:22:14.000
In language from different fields and use them at the same time, and you get uh you get to tap into the expertise of the LLM, uh you know, wearing different sort of expert hats at the same time, which is cool.
00:22:14.319 --> 00:22:22.000
It sounds like what you're doing in this example is like what you're asking is basically what are the algebraic properties that also work in animation?
00:22:22.160 --> 00:22:25.519
And here you've shown like, okay, addition is commutative.
00:22:25.680 --> 00:22:28.400
So no matter what sequence, you get the same results.
00:22:28.559 --> 00:22:33.119
So basically, what here on screen I see grow circle plus fade out.
00:22:33.279 --> 00:22:39.279
And then the LLM knows like, okay, these are the algebraic properties that also should be applied here for the animation.
00:22:39.440 --> 00:22:43.519
That's that's actually such a smart way to communicate um with an LLM.
00:22:43.680 --> 00:22:44.559
And yeah.
00:22:44.880 --> 00:22:59.279
You know, experimentally I have haven't done much with animations personally with with LLMs in my projects, but I have I have explored using the language of algebra to communicate with the LLMs to design libraries that are easy to use.
00:22:59.440 --> 00:23:07.039
And ultimately, it in my experience, it does pull the LLMs to build nice libraries that I think are nice to use.
00:23:07.200 --> 00:23:08.079
So that's just my opinion.
00:23:08.319 --> 00:23:09.039
That's so cool.
00:23:09.200 --> 00:23:13.759
So basically, what we're describing here is we're even going beyond what Guillermo has said here.
00:23:13.920 --> 00:23:29.279
So it's not only like, okay, try to communicate super succinctly but in a powerful way what your intent is by using the domain language of the thing that you're trying to build in in a specific domain, but you could even go beyond and also use terms from another domain.
00:23:29.519 --> 00:23:30.400
This is so interesting.
00:23:30.640 --> 00:23:35.759
Do any examples come to mind where you've seen other buildings doing this too?
00:23:36.000 --> 00:23:42.640
Where they go beyond one domain and combine it with terminology or libraries that are related to other domains?
00:23:42.880 --> 00:23:43.119
Yeah.
00:23:43.359 --> 00:23:52.799
Well, the uh it turns out using this like algebraic approach, it works really well when you're doing things that are very abstract.
00:23:52.960 --> 00:23:57.119
It's called abstract algebra, so it makes sense that when you're doing abstract things with software, it's useful.
00:23:57.279 --> 00:24:04.720
So like handling errors, just like the notion of handling errors, all the different ways you might want to handle errors, dealing with things running at the same time concurrently.
00:24:04.880 --> 00:24:09.119
That's another like abstract thing that works well with this kind of algebraic approach.
00:24:09.279 --> 00:24:10.480
Dependency management.
00:24:10.640 --> 00:24:16.480
So if you depend on certain resources and they need to load in certain orders and all of that.
00:24:16.640 --> 00:24:28.079
And so there's actually a library in TypeScript that has become very popular recently since the since the coding agent boom, which is called Effect TS.
00:24:28.240 --> 00:24:37.759
And Effect literally is just the combination of error handling, concurrency management, and dependency management treated very algebraically.
00:24:37.920 --> 00:24:43.599
And it's something that it's a little bit scary to developers at first, like this kind of thinking.
00:24:43.759 --> 00:24:54.000
I don't know people who've worked with Effect directly, but with this kind of approach, I've found resistance in you know different, different teams where I've tried to introduce this kind of programming before.
00:24:54.160 --> 00:24:59.039
But you get so much benefit out of it that it's kind of, in my opinion, it's worth getting through that resistance.
00:24:59.200 --> 00:25:02.240
Well, the the coding agents, they are experts in everything.
00:25:02.319 --> 00:25:04.000
So there's no resistance anymore.
00:25:04.160 --> 00:25:08.000
And and turns out, you know, doing things algebraically is a good idea.
00:25:08.160 --> 00:25:22.799
And so Effect TS suddenly is this super powerful library that people are finding empirically that when you tell your coding agent to use effect when it's building your TypeScript software, the coding agent just does a better job.
00:25:22.960 --> 00:25:25.519
There's less bugs, it can do more complex things more easily.
00:25:25.759 --> 00:25:34.480
This is the power of algebra, I guess, applied to, in this case, like this abstract idea of error handling, dependency management, and concurrency.
00:25:34.720 --> 00:25:35.119
Awesome.
00:25:35.279 --> 00:25:35.680
Very cool.
00:25:35.839 --> 00:25:43.519
So it's this powerful TypeScript library that also handles errors and dependencies and all those stuff in algebraic way.
00:25:43.680 --> 00:25:46.000
But this really sets people thinking like, what can we do?
00:25:46.160 --> 00:25:50.559
What else can we introduce to explore the latent space from the LLM?
00:25:50.720 --> 00:25:51.359
Yeah, cool.
00:25:51.440 --> 00:25:52.240
Thanks for sharing.
00:25:52.400 --> 00:25:54.079
Should we go dive into the next one?
00:25:54.240 --> 00:25:54.480
All right.
00:25:54.640 --> 00:25:58.960
Um, emerging practice number three, match the tool interface to the task.
00:25:59.119 --> 00:26:04.960
I think when we talked about this in our previous conversation, you you said be thoughtful with your agent tools.
00:26:05.119 --> 00:26:07.920
And we mentioned different tools in our conversation.
00:26:08.079 --> 00:26:13.039
There was this like hugely popular MCP that became super popular last year.
00:26:13.200 --> 00:26:18.720
And then the second half of last year, people started to shift more to CLI.
00:26:18.799 --> 00:26:22.160
And now it sounds like uh what's becoming more popular is code mode.
00:26:22.319 --> 00:26:24.000
Could you tell us a bit more about these?
00:26:24.240 --> 00:26:24.559
Yeah.
00:26:24.799 --> 00:26:31.599
Um, well, I don't think code mode is becoming popular yet, but uh I think it should become popular because it's a good idea.
00:26:31.680 --> 00:26:32.319
Let's put it that way.
00:26:32.559 --> 00:26:38.400
Before we start, maybe just to refresh everyone's um memory about it, why was everyone so excited about MCP?
00:26:38.480 --> 00:26:44.480
And what why was CLI perceived as better by some afterwards compared to MCP?
00:26:44.640 --> 00:26:47.200
And then why do you think Coke mode should be more popular now?
00:26:47.680 --> 00:26:47.839
Okay.
00:26:48.000 --> 00:26:55.759
Um so I I think in order to have this discussion, it's important to understand the agent loop carefully.
00:26:55.920 --> 00:27:06.160
So the the agent loop, which is what our coding agents do or any any agent, uh, is something like, you know, send prompt, get result, decide, call a tool or done.
00:27:06.240 --> 00:27:17.680
Um if it's call a tool, call the tool, get a result, put the result back into context, and then call inference again, get another result, and maybe you'll call a tool, and then you just keep looping until you're done.
00:27:17.759 --> 00:27:22.640
Um the early, uh in the early days, uh, you had to hard code your tools.
00:27:22.799 --> 00:27:36.079
So you in the in the system prompt for the agent, you'd say, okay, you have the ability to get the weather or use a calculator, and then you specified how the agent in its response can say, Hey, I want to call the weather and get the tool.
00:27:36.240 --> 00:27:38.079
Oh yeah, I still remember those days.
00:27:38.319 --> 00:27:38.960
Yeah.
00:27:39.680 --> 00:27:43.440
Um, and this was this was a little strict, I guess.
00:27:43.519 --> 00:27:47.680
And so, so enter MCP model context protocol.
00:27:47.839 --> 00:27:55.680
What MCP is, is it's literally a layer of indirection so that your tool call can be defined outside of the agent loop.
00:27:55.839 --> 00:28:00.720
So, so the tool that we put in the agent loop is call mcp, basically.
00:28:00.880 --> 00:28:06.160
And one of them is, you know, get get information from MCP server to tell me what the tool does.
00:28:06.319 --> 00:28:10.240
And the other one is call mcp tool with the information I got from the server.
00:28:10.400 --> 00:28:19.440
And and what this what this lets you do is it lets you get, it sort of introduces a plug-in ecosystem where lots of services can provide MCP interfaces.
00:28:19.599 --> 00:28:28.559
And you can, as a user of an agent, connect to any of these providers without the agent builder directly connecting them ahead of time.
00:28:28.720 --> 00:28:30.160
So that's like a cool thing.
00:28:30.319 --> 00:28:36.240
And in the context of coding agents, this is a way for you know cloud code to talk to GitHub.
00:28:36.319 --> 00:28:38.880
It's one way for Cloud Code to talk to GitHub, for example.
00:28:38.960 --> 00:28:39.200
Okay.
00:28:39.759 --> 00:28:40.000
Okay.
00:28:40.240 --> 00:28:44.000
So now there's a CLI, a CLI interface.
00:28:44.160 --> 00:28:48.079
So let's let's let's use this example of getting information from GitHub.
00:28:48.240 --> 00:28:52.640
You could use a GitHub MCP server, or you can use a GitHub CLI tool.
00:28:52.720 --> 00:28:59.200
And so rather than giving in the agent loop an MCP tool to the coding agent, you would give a bash tool.
00:28:59.359 --> 00:29:04.799
So the bash tool lets the coding agent run any bash command or a small bash script.
00:29:04.960 --> 00:29:12.079
So in this bash script, they can call one or more CLI tools and combine and the results together in interesting ways.
00:29:12.319 --> 00:29:25.200
Um, so for example, if we wanted to get the issues associated with a certain GitHub repository, I could make an MCP call to GitHub to get the issues, I get the result back, and then I give it to the user.
00:29:25.440 --> 00:29:28.880
Or with the CLI tool, I can make a bash call.
00:29:29.039 --> 00:29:36.160
The bash script can call, use the GitHub tool to get the information about the issues, and then I give it back to the user.
00:29:36.319 --> 00:29:37.359
Okay, it's about the same.
00:29:37.519 --> 00:29:41.680
Now, the interesting thing comes when you want to do something more complex.
00:29:41.839 --> 00:29:46.960
I'm gonna use another simple example, which is count the number of issues on a GitHub project.
00:29:47.039 --> 00:29:52.400
And let's just say there's no MCP call for that directly, or there's no CLI call for that directly either.
00:29:52.559 --> 00:29:54.480
You only have the ability to get all the issues.
00:29:54.640 --> 00:30:10.000
If you think about what the agent loop is doing, the first the loop goes and the agent decides, I want to call GitHub to get the issues, gets back the issues, those issues enter context, and then inference runs again on whatever the context was before plus the results of those issues.
00:30:10.160 --> 00:30:16.240
And then maybe the agent needs to use another tool to count, or maybe it counts by looking at the context and it gives you the answer.
00:30:16.400 --> 00:30:21.759
Now, with the CLI tool approach, the agent decides, okay, I need to use the GitHub tool.
00:30:21.920 --> 00:30:38.319
But rather than just getting the results of the issues back into context, the agent decides, okay, get the issues with the GH tool and pass the results of those issues to another CLI tool called WC, for example, which which can count the number of issues.
00:30:38.480 --> 00:30:40.559
And then what it gets back is just the number.
00:30:40.720 --> 00:30:44.480
And so the the context that's used is much smaller.
00:30:44.559 --> 00:30:47.279
There's only one tool call and and then and then you get the result.
00:30:47.440 --> 00:31:00.799
So so using the CLI, using bash, lets the agent put more information or more logic, can do more and compose together different tools in one logical agent tool call, which is just more fit.
00:31:00.960 --> 00:31:12.079
So because it was just one agent tool call, when you do it with the CLI, and it in your example, it did getting the issues and counting the issues right after one another.
00:31:12.240 --> 00:31:21.279
Um so basically the LLM doesn't feed all the intermediate results back into the loop and redo the loop basically before it does the counting.
00:31:21.359 --> 00:31:24.240
It just gets past the intermediate results from the previous step.
00:31:24.319 --> 00:31:26.480
Uh basically your context window stays cleaner.
00:31:26.559 --> 00:31:26.720
Yes.
00:31:26.880 --> 00:31:27.519
Well, interesting.
00:31:27.599 --> 00:31:28.480
That's that makes sense.
00:31:28.640 --> 00:31:37.039
And this is why uh the world has shifted a lot of momentum towards CLI interfaces rather than MCP.
00:31:37.839 --> 00:31:39.680
Um, because it's just more efficient.
00:31:41.119 --> 00:31:46.960
Um now, now, the interesting thing is you can take another step, and this is this is code mode.
00:31:47.119 --> 00:31:51.279
So this is something that people have been discussing, I think since this year.
00:31:51.440 --> 00:31:52.559
So it's been a few months.
00:31:53.039 --> 00:31:54.960
I'm gonna motivate this with another example.
00:31:55.119 --> 00:31:57.200
So I'm gonna use real example.
00:31:57.440 --> 00:32:02.799
So so uh Anthropic just this week released a let me share my screen.
00:32:02.880 --> 00:32:03.359
Let's see.
00:32:03.519 --> 00:32:03.920
Here it is.
00:32:04.160 --> 00:32:07.680
Anthropic released dynamic workflows this week or last week.
00:32:08.000 --> 00:32:09.599
Anyway, one of these weeks recently.
00:32:09.759 --> 00:32:10.960
What is a dynamic workflow?
00:32:11.119 --> 00:32:11.839
It's code.
00:32:13.359 --> 00:32:13.680
Okay.
00:32:14.079 --> 00:32:22.799
So um so let's let's use this as our example to explain what code mode is and why it's valuable when when you when you can use it.
00:32:24.559 --> 00:32:28.160
Is this uh their form of code mode or is this something different?
00:32:28.319 --> 00:32:29.519
So dynamic workflow.
00:32:30.000 --> 00:32:40.240
Code mode is a um you can think of code mode as the interface to um a set of tools or like a tool that you'd plug into your agent.
00:32:40.400 --> 00:32:52.000
So so in this case, dynamic workflows is uh a code mode style way of allowing cloud code to interact with subagents, to orchestrate subagents, basically.
00:32:52.240 --> 00:32:52.480
Cool.
00:32:52.640 --> 00:32:59.759
So maybe the analogy is like dynamic workflows and allows cloud code to interact with subagents using code.
00:32:59.839 --> 00:33:03.680
And in code mode, it allows the agent to interact with tools using code.
00:33:03.839 --> 00:33:04.640
That's the analogy.
00:33:04.799 --> 00:33:05.200
Is that right?
00:33:05.440 --> 00:33:11.839
Yeah, dynamic workflows allows cloud code to interact with subagents using code, but I haven't motivated why that's good yet.
00:33:12.079 --> 00:33:12.799
Okay, tell us.
00:33:12.960 --> 00:33:16.799
So, okay, so so so the task is to work with subagents.
00:33:16.880 --> 00:33:21.680
So that's spawn subagents, wait for subagents, and sort of organize subagents in different ways.
00:33:21.839 --> 00:33:25.839
So you could start with just a simple tool call or an MCP.
00:33:26.000 --> 00:33:34.240
And let's say I want to create five agents, and each of those creates two agents, and then I wait for all the results of those agents to come back.
00:33:34.400 --> 00:33:36.480
So that's that's the orchestration I want to do.
00:33:36.559 --> 00:33:40.880
So if I was using tool calls, I would make a tool call to spawn each agent.
00:33:41.039 --> 00:33:42.640
Maybe I could spawn five at a time.
00:33:42.720 --> 00:33:45.119
And then for each one, I would have to spawn more agents.
00:33:45.279 --> 00:33:55.759
So we're at, I don't know, 10 tool calls, maybe at the minimum, maybe like 20, and then wait for the results of all of these, which is at least another one.
00:33:56.000 --> 00:33:59.039
We're at like 25 tool calls, lots of intermediate context.
00:33:59.119 --> 00:34:02.960
It's kind of the same as the last example, a lot of context wasted, bad.
00:34:03.039 --> 00:34:03.200
Okay.
00:34:03.359 --> 00:34:05.839
So then you could think, okay, well, what if it was a CLI tool?
00:34:05.920 --> 00:34:06.400
That's better.
00:34:06.480 --> 00:34:15.199
So I could use bash and I could uh if there was a CLI tool that could spawn a subagent and wait for the results, well, I could spawn all the subagents.
00:34:15.440 --> 00:34:20.000
I could represent that structure in a bash script.
00:34:20.159 --> 00:34:23.840
Um, but it turns out bash is a bad programming language.
00:34:24.239 --> 00:34:30.400
And so when when when bash when a bash script gets really complicated, errors start happening.
00:34:30.639 --> 00:34:35.280
The like both humans and LLMs just can't like kind of struggle to make it work.
00:34:35.519 --> 00:34:40.639
Um and uh and you know, JavaScript is a better programming language than Bash.
00:34:40.719 --> 00:34:51.199
And so what code mode does is instead of saying bash is my tool and I can call some CLIs and connect them together, code mode says, here's a JavaScript interface and just write JavaScript.
00:34:51.280 --> 00:34:54.000
And and and that's exactly what dynamic workflows is.
00:34:54.159 --> 00:35:05.119
It's a it's a mechanism for for Claude when it's spawning subagents or describing a subagent orchestration workflow to use JavaScript or TypeScript for coding.
00:35:05.360 --> 00:35:17.679
Uh and so in the agent loop, the the agent, instead of making a tool call to MCP or instead of making a tool call to bash with a bash script, it makes a tool call to do dynamic workflows and gives it a bunch of JavaScript.
00:35:17.760 --> 00:35:27.840
And then that JavaScript can efficiently and effectively, and in a way that the LMs don't make many mistakes, describe very complex structures of workflow orchestration.
00:35:28.000 --> 00:35:29.360
Um interesting.
00:35:29.519 --> 00:35:35.840
And how's that where where how does it relate to using code mode for using tools?
00:35:36.239 --> 00:35:40.480
This is just an example of a tool that works really well with code mode.
00:35:40.559 --> 00:35:41.840
I can give more examples.
00:35:42.079 --> 00:35:51.679
So there's a Cloudflare in uh, I guess in February, presented a mechanism to do search with code mode.
00:35:51.920 --> 00:35:58.400
Um, I won't go through that whole blog post, but basically the the search API for Cloudflare is very powerful.
00:35:58.480 --> 00:36:06.159
And so rather than exposing all those belt, like little details through a CLI tool, maybe they are also exposed through the Cloudflare CLI tool.
00:36:06.239 --> 00:36:14.400
But in order to take advantage of all those features and being able to be flexible, the Cloudflare team found it helpful to provide a code mode interface.
00:36:14.559 --> 00:36:17.679
So the tool gives JavaScript to Cloudflare.
00:36:18.000 --> 00:36:22.960
Um, and that JavaScript describes how you want to do your search when the when the tool call comes back.
00:36:23.119 --> 00:36:24.880
So that's another example of code mode.
00:36:25.199 --> 00:36:25.519
Cool.
00:36:25.679 --> 00:36:38.719
So it sounds like code mode is really good if you want your agents to write arbitrary programs that are doing like more complex, or if you want it to be more flexible and do more complex logic.
00:36:38.880 --> 00:36:39.119
Yes.
00:36:39.280 --> 00:36:40.800
That's when code mode will shine.
00:36:40.960 --> 00:36:56.719
And bash is basically really good at combining like maybe CLI calls from different tools, but not really good for writing, again, writing arbitrary programs that can do very flexible things and puzzling them together, then then uh code mode is better in that situation.
00:36:56.880 --> 00:36:57.360
Is that right?
00:36:57.599 --> 00:36:58.400
Yes, exactly.
00:36:58.559 --> 00:36:58.800
Yeah.
00:36:58.960 --> 00:37:03.039
So so code mode is not something that can just replace using the bash tool.
00:37:03.119 --> 00:37:04.239
The bash tool is amazing.
00:37:04.400 --> 00:37:08.320
It's important for combining CLI tools together, like like you said, efficiently.
00:37:08.559 --> 00:37:21.599
But uh code mode is useful when um when their tool that you're using has enough complex behavior that you could design some kind of interface, some kind of API to control it with code.
00:37:21.760 --> 00:37:22.800
Let me give one more example.
00:37:22.960 --> 00:37:27.360
This is something that I used in in one of my projects a month or two ago.
00:37:27.760 --> 00:37:30.320
So let's see if I can I have it here.
00:37:31.119 --> 00:37:37.760
Um so so I one of my projects is uh building a Game Boy game that runs LLM inference.
00:37:37.840 --> 00:37:40.239
Um so it's just like a fun little process.
00:37:40.639 --> 00:37:47.119
When you're when you're building a Game Boy game, it's important to have ways to debug issues.
00:37:47.280 --> 00:37:58.480
And there's a tool called a debugger, which lets you step through instructions carefully uh and inspect hardware state of the Game Boy um when it's running in an emulator.
00:37:58.719 --> 00:38:06.480
So um you can build a debugger that has a bunch of buttons on it, and that's hard for an LLM to use.
00:38:06.639 --> 00:38:12.320
So then you can say, okay, well, I can build a CLI debugger, and then that's easier for the LLM to use, and it is.
00:38:12.480 --> 00:38:18.320
But I found debuggers are complex enough that actually it makes sense to provide a code mode interface.
00:38:18.559 --> 00:38:26.480
So so my debugger CLI tool has a code mode in it where you feed it JavaScript.
00:38:26.719 --> 00:38:32.159
Um so uh and and that lets you that sort of gives a lot more power to the to the system.
00:38:32.320 --> 00:38:44.639
So for example, one thing that's like very doesn't really make sense to do from a CLI tool context is like automatically running logic every time a certain part of the the Game Boy game executes.
00:38:44.800 --> 00:38:52.960
So for example, every time I hit this instruction, you know, look at the the Game Boy state, and if it's in this current state, set new breakpoints that do this over here.
00:38:53.199 --> 00:39:01.519
So these kinds of instructions, it's very hard to come up with a way for a CLI tool to express that logic, but it's very simple to express that in JavaScript.
00:39:01.599 --> 00:39:08.480
And when I tell my coding agent to use this debugger in code mode, it can come up with those complex scripts by itself.
00:39:08.639 --> 00:39:10.559
And I found it useful for debugging.
00:39:10.880 --> 00:39:11.280
Interesting.
00:39:11.440 --> 00:39:16.800
So basically you're telling like the LLM this is how you code the tool in order to use it.
00:39:17.119 --> 00:39:22.400
Yeah, so I have a skill for the Game Boy debugger that has a code mode interface.
00:39:22.559 --> 00:39:22.880
Yeah.
00:39:23.039 --> 00:39:31.199
And that skill teaches the LLM how to work with the complex interface and write software to sort of get the most out of it when it's debugging.
00:39:31.519 --> 00:39:32.079
Very cool.
00:39:32.239 --> 00:39:47.840
So basically, I guess uh the lesson here is understand why or the cost of using these different ways of giving your LLM tools so you can make good decisions, tasteful decisions to choose one over the other.
00:39:48.079 --> 00:39:48.559
Yeah, yeah.
00:39:48.719 --> 00:39:51.119
And by the way, sometimes you need to use MCP.
00:39:51.199 --> 00:39:59.599
If you're it's easier to deal with stateful services using MCP, it's still possible with CLI tools, but it's a little bit annoying.
00:39:59.840 --> 00:40:09.039
So if you have some kind of stateful service that you're interacting with, it can be nice to use MCP over CLI in that case.
00:40:09.280 --> 00:40:11.440
So yeah, it's be be thoughtful.
00:40:11.599 --> 00:40:17.679
And I think when you can use code mode, uh, you'll unlock a new superpower with your coding agent.
00:40:17.920 --> 00:40:18.480
That's very cool.
00:40:18.719 --> 00:40:28.320
It sounds like if you want to have good taste, you definitely also want to have deeper understanding so you can make those tasteful decisions.
00:40:28.480 --> 00:40:28.800
All right.
00:40:28.960 --> 00:40:30.880
Yeah, thanks for walking us through this.
00:40:31.039 --> 00:40:34.320
Anything you want to share before we go to our last emerging practice?
00:40:34.559 --> 00:40:36.000
No, let's let's do it.
00:40:36.159 --> 00:40:36.480
All right.
00:40:36.639 --> 00:40:38.000
I also really like this last one.
00:40:38.159 --> 00:40:39.599
Emerging practice number four.
00:40:39.760 --> 00:40:41.360
Use oracles when you can.
00:40:41.519 --> 00:40:42.559
Can you tell us more about it?
00:40:42.800 --> 00:40:43.119
Yeah.
00:40:43.280 --> 00:40:49.920
Um, so this idea was one of the points that a speaker at AI Native DevCon made.
00:40:50.000 --> 00:40:51.360
It was Justin Kormack.
00:40:51.440 --> 00:40:52.880
He's the ex-CTO of Docker.
00:40:53.119 --> 00:41:00.000
And Justin is rebuilding his own version of AWS S3 from scratch as a as a fun project.
00:41:00.159 --> 00:41:02.719
And I guess so Oracles.
00:41:03.039 --> 00:41:04.159
So what is an Oracle?
00:41:04.480 --> 00:41:05.199
Let's start with that.
00:41:06.159 --> 00:41:12.639
An Oracle is uh some system that you just say is correct.
00:41:12.880 --> 00:41:23.039
So in in Justin's example, since he's rebuilding AWS S3, the real using an Oracle, just quick question, so so you can all follow along.
00:41:23.440 --> 00:41:26.880
Well, so he's building a couple of AWS S3.
00:41:27.199 --> 00:41:36.559
Um there's uh the real Amazon S3 is uh an Oracle or can be an Oracle for his version.
00:41:36.880 --> 00:41:43.519
So is he making a case that using Oracle is so good and he it's a lesson from his own experience?
00:41:43.760 --> 00:41:53.840
Yes, it's one of one of the lessons from his presentation was using an Oracle uh during during exploration of behavior and testing was very useful.
00:41:54.000 --> 00:42:01.519
So he found that um the documentation for AWS S3 is actually not accurate when compared to the behavior.
00:42:01.679 --> 00:42:13.519
And so it's it's better, it was better for him to have his agent just explore what the real S3 does and learn that behavior so that he is able to implement something that works in the same way.
00:42:13.840 --> 00:42:18.400
And that's so that's why you meant he used AWS as the Oracle.
00:42:18.559 --> 00:42:19.679
Yeah, yeah.
00:42:19.920 --> 00:42:25.440
Um so there's a lot of ways to use Oracles in your development.
00:42:25.599 --> 00:42:31.440
Um and uh and you know, as Justin's saying, it's really useful.
00:42:31.679 --> 00:42:35.119
So I have did a project and and take took advantage of this.
00:42:35.280 --> 00:42:47.760
I guess just really quickly, I worked on another Game Boy emulator project where I was building a GPU emulator for Game Boy, and I used a known good CPU Game Boy emulator as an Oracle.
00:42:47.920 --> 00:43:04.000
And once I had infrastructure set up so that in my harness, so that the agent could just automatically check against the real working emulator and mine, the the agent was able to run autonomously for a few days and just get everything working perfectly, basically.
00:43:06.000 --> 00:43:12.960
So it sounds like it's good to use an Oracle when you want to convey the intent of the thing that you want to build.
00:43:13.039 --> 00:43:13.599
Is that right?
00:43:13.760 --> 00:43:15.119
So basically it's a shortcut.
00:43:15.360 --> 00:43:17.679
You're conveying it more specifically.
00:43:18.800 --> 00:43:24.800
So it's not, it's it's very useful in testing, but as you're saying, it's also useful as a communication mechanism.
00:43:24.960 --> 00:43:35.039
It's kind of like, you know, the saying like a picture is worth a thousand words, a working version of the thing is worth more words than you trying to describe how it works, just showing it to the agent.
00:43:35.199 --> 00:43:37.360
Something that Justin mentioned in his presentation.
00:43:37.440 --> 00:43:38.639
He just said a sentence on it.
00:43:38.719 --> 00:43:52.239
Again, something that was just one sentence that really stuck with me, which was he said, it's probably worth implementing a basic version of the service or application that you're trying to build first and get the behavior just perfect.
00:43:52.559 --> 00:43:53.920
But take shortcuts.
00:43:54.000 --> 00:43:55.360
It doesn't actually work fully.
00:43:55.440 --> 00:43:59.519
No, it's not ready to be shipped to people, but it but it works enough so that the behavior is exactly.
00:44:00.480 --> 00:44:04.880
And then use that as an Oracle to then go and implement the real thing that might be a lot more complicated.
00:44:05.039 --> 00:44:06.639
And I think that is a really cool idea.
00:44:06.960 --> 00:44:15.039
So, as an example, let's say I build a journaling app.
00:44:16.800 --> 00:44:19.280
And it only works locally on my machine.
00:44:19.440 --> 00:44:35.519
But if I want to scale it up and I want it to be able to serve thousands of users, it's okay if my own version on my machine works only well with me as the only user, because then we can use it as an Oracle and AI will know how it should behave even when there are thousands of people using it.
00:44:35.920 --> 00:44:36.239
Exactly.
00:44:36.480 --> 00:44:46.480
Shortcut to building the version that supports thousands of users, to have the version, to build the version that works with one user first and get that behavior exactly how you want it.
00:44:46.639 --> 00:44:57.199
And then you get a shortcut because you can say, okay, this is the Oracle, make Oracle tests, make it behave exactly like this, and then go implement this thing that's much more complicated that supports all these users.
00:44:57.440 --> 00:44:57.679
Right.
00:44:57.840 --> 00:45:03.599
Because if you even if you write a spec, sometimes there are still like great areas or maybe there are blind spots.
00:45:03.679 --> 00:45:12.559
But if you give your AI access to an Oracle whenever it has any doubts or doesn't know what decisions it should make, basically it will just query the Oracle.
00:45:12.719 --> 00:45:12.880
Yes.
00:45:13.119 --> 00:45:13.760
That's very smart.
00:45:13.920 --> 00:45:21.599
Actually, it makes me also wonder whenever I see things, uh, people building things in a space and they're like, oh, I just built this in two days using AI.
00:45:21.760 --> 00:45:26.159
I built this browser, you know, in in a day with an agent swarm.
00:45:26.320 --> 00:45:33.119
Did they build it from scratch, like in the traditional way, like how you would like start with a spec and just build it from the ground up?
00:45:33.280 --> 00:45:34.960
Or did they use an Oracle for it?
00:45:35.119 --> 00:45:37.519
Because it sounds like using an Oracle could be a shortcut.
00:45:37.760 --> 00:45:44.079
It's a real, it's it's a shortcut and it lets you do things that are otherwise impossible with this version of agents.
00:45:44.239 --> 00:45:47.599
Certainly impossible to do autonomously without human input.
00:45:47.679 --> 00:46:00.559
So all the examples that we've seen that are really impressive, building a browser, building an operating system, building a C compiler, these are things that have oracles that can be used to ensure that the thing is correct as it's being built.
00:46:00.800 --> 00:46:02.880
Yeah, and it's not, I don't think that's bad.
00:46:03.039 --> 00:46:03.599
I think it's good.
00:46:03.760 --> 00:46:05.599
I mean, it's something it's something to learn from.
00:46:05.760 --> 00:46:23.679
I think, like in the context of thinking about taste, um uh the best way to express little nuances in your taste is by having a prototype of the thing that behaves exactly how you want it to behave, and then telling the agent it needs to be exactly like this.
00:46:23.840 --> 00:46:24.000
Yeah.
00:46:24.159 --> 00:46:26.800
No, this makes me really excited to try it out.
00:46:27.039 --> 00:46:31.519
Not not only this emerging practice, but also the others that you just mentioned.
00:46:31.679 --> 00:46:33.360
This is so this is so good.
00:46:33.599 --> 00:46:40.079
So if we zoom out, uh, all four of these practices that we talked about pointed kind of in the same direction.
00:46:40.239 --> 00:46:52.079
Like good taste when building with AI is not just using better models or just writing better prompts, but well, among many things, you have to think about how to shape the environment around a model.
00:46:52.320 --> 00:46:58.960
You have to learn how to convey the intent more precisely, either through an Oracle or more precise expert language.
00:46:59.199 --> 00:47:05.039
And also you'd have to think about how to present agent tools to the agents in every situation.
00:47:05.199 --> 00:47:09.519
So, yeah, I guess these were the four emergent practices that we wanted to share.
00:47:09.679 --> 00:47:12.719
And this is also what we're going to keep exploring.
00:47:12.880 --> 00:47:19.440
So, how builders can develop better taste and how they can channel that taste into tools, agents, and workflows.
00:47:19.599 --> 00:47:24.000
Yeah, I got a lot out of it, and I'm very excited to try these things out.
00:47:24.239 --> 00:47:29.440
Yeah, and and these were just four things that happened to be mentioned in the AI space last week.
00:47:29.679 --> 00:47:34.320
And you know, a few of them were just one sentence that we turned into an hour-long discussion.
00:47:34.480 --> 00:47:37.039
So there's so many more of these practices.
00:47:37.199 --> 00:47:39.119
I mean, I have a hundred in my head right now.
00:47:39.280 --> 00:47:43.039
So Christine and I are gonna have a lot of conversations about this just for fun.
00:47:43.199 --> 00:47:47.840
And uh yeah, I guess let us know if you want us to keep sharing these, and we will.
00:47:48.079 --> 00:47:51.599
Thank you so much for listening, and thank you as well, Brendan.
00:47:51.679 --> 00:47:52.719
This was a lot of fun.