WEBVTT
00:00:00.149 -->
00:00:10.289putting a custom harness in every meaningful part of your workflow and not just at the root of your project, you can squeeze dramatically more intelligence out of your agents.
00:00:11.089 -->
00:00:16.629Basically, general-purpose harnesses are built to handle broad use cases and messy realities.
00:00:17.390 -->
00:00:22.250They are not optimal for any one task because they need to be decent at all of them.
00:00:22.885 -->
00:00:23.344The next step
00:00:23.344 -->
00:00:23.545Yeah. Yeah. Yeah.
00:00:24.024 -->
00:00:26.925is to treat harness engineering as something we do at every meaningful
00:00:26.925 -->
00:00:27.164Yeah. Yeah
00:00:27.304 -->
00:00:31.964section of your project and to do it even for every sufficiently hard task.
00:00:32.865 -->
00:00:37.564An agent should not just have its own job to do, it should have its own world.
00:01:02.148 -->
00:01:18.563In this episode we will briefly review what we mean when we say harness engineering. We will talk about what if you deliberately built a harness in every meaningful part of your workflow and even between sub-agents in the same prompt and not just on the top of your work tree.
00:01:19.582 -->
00:01:31.082We'll also talk about when narrow- when you have to narrow the capabilities of your agents and how to enforce behavior instead of just hoping that the model will follow your instructions.
00:01:31.677 -->
00:01:38.518So we've all probably experienced an agent making mistakes, like ignoring instructions or repeatedly producing poor results.
00:01:39.177 -->
00:01:43.058And often the way we respond is by adding more instructions.
00:01:43.878 -->
00:02:03.438For example, we clarify the prompt, we create a skill or we paste in more context or we write another role in agents dot MD[UM] but and sometimes this all works but it can also turn one general-purpose harness into an increasingly large collection of guidance that every task must carry around.
00:02:04.018 -->
00:02:11.677And Brandon, I know you and I like talked for hours the other day about this and you have a lot of thoughts about it.
00:02:08.957 -->
00:02:11.677Have you run into this as well
00:02:12.182 -->
00:02:17.117Yeah. We, we have to remember that the models have a certain amount of attention.
00:02:17.578 -->
00:02:24.573So, the more things that you put in context, the more distracted they get.
00:02:22.598 -->
00:02:24.573Right? So,
00:02:24.573 -->
00:02:25.612That makes sense.
00:02:25.953 -->
00:02:36.177if you're blindly shoving things in your agents dot MD whenever there's a mistake, if it's something that's rare, then you're distracting the agent when it's doing most of its other work, which is
00:02:36.177 -->
00:02:36.277Yeah.
00:02:36.818 -->
00:02:37.217not good,
00:02:37.217 -->
00:02:37.277Yeah,
00:02:37.538 -->
00:02:37.777not good.
00:02:37.777 -->
00:02:37.818yeah.
00:02:38.538 -->
00:02:38.918I think
00:02:38.918 -->
00:02:38.957Yeah,
00:02:39.217 -->
00:02:58.927a lot of people have started to do better. They're they're starting to do progressive disclosure of information. So if you're using skills, for example, then your agent will only pick up the skill if it's about to do something that is related to that skill, unless you, you know, manually include it[UH] and there's other ways to
00:02:58.927 -->
00:02:58.948yeah.
00:02:59.027 -->
00:02:59.288do this.
00:02:59.288 -->
00:02:59.328Yeah
00:02:59.348 -->
00:03:03.147You can have a directory of different markdown files and pull them in as needed.
00:03:03.707 -->
00:03:03.948So.
00:03:04.627 -->
00:03:06.247But this is only one half of the problem.
00:03:06.543 -->
00:03:06.582Mm-hmm.
00:03:06.582 -->
00:03:33.122Cuz as you as you know, listener [laughter] or maybe as as you might remember, agents get their context one from these feedforward guides, markdown that we're inserting ahead of time, but also from sensors, from things that are giving feedback, from tests that are running, from code reviews, things like this. So, these are giving feedback into the agent after they take certain steps. And
00:03:33.143 -->
00:03:33.144I
00:03:33.144 -->
00:03:33.323that's also
00:03:33.323 -->
00:03:33.582know.
00:03:33.582 -->
00:03:34.622adding to the context.
00:03:35.018 -->
00:03:53.603Yeah, I know we talked about this in our episode three, the latest episode, but basically guides are things and context that the agent gets before it starts its task, and the sensors are the things that the agent get after the task and they would give, for example, test lend rules. Right?
00:03:53.842 -->
00:03:53.943Yeah.
00:03:54.022 -->
00:03:54.263OK,
00:03:54.402 -->
00:03:54.483after
00:03:54.483 -->
00:03:54.682cool.
00:03:54.763 -->
00:03:56.997steps basically yeah. So.
00:03:56.997 -->
00:03:57.992OK. It it makes sense
00:03:57.992 -->
00:03:58.252So.
00:03:58.413 -->
00:03:58.652what you're
00:03:58.652 -->
00:03:58.712Mm-hmm.
00:03:58.712 -->
00:04:11.508saying because [UM][UM] also just [UH] [UH] to remind myself, because if you put for s- for I like how you said like some skills you could only invoke them when the time is right, because
00:04:11.728 -->
00:04:11.747Yeah
00:04:11.967 -->
00:04:21.028you know, a a bad example would be to put it in the agents or Claude dot MD because basically every agent or sub-agent that you spawn will also read that markdown file, right?
00:04:21.648 -->
00:04:22.028Exactly.
00:04:22.028 -->
00:04:31.348So, so I guess the problem is not, hey the model needs another instruction, but sometimes the model needs a better world, like you said, or an [UH] a better environment.
00:04:31.642 -->
00:04:31.783Right.
00:04:32.302 -->
00:04:32.562Exactly.
00:04:32.562 -->
00:04:32.923Awesome.
00:04:33.382 -->
00:04:37.723Instructions are just part of the environment. They're just part of the world.
00:04:35.742 -->
00:04:38.083But there's other parts in the world too
00:04:39.043 -->
00:04:46.038Gotcha. I would love to dive into that and and I think to talk about that we need to dive into harness engineering.
00:04:43.617 -->
00:04:47.057And I know you are jumping[laughter] to get
00:04:47.057 -->
00:04:47.117[laughter]
00:04:47.117 -->
00:05:02.718into harness engineering. Maybe before we like dive into this[UH] [UM] different way of thinking about harness engineering, can we first tell everyone what we mean when we say harness engineering? Like what what are the things exactly that are included in the harness
00:05:04.273 -->
00:05:16.127when we say harness engineering, we mean the coding agent harness, Claude Code, Codex, pi, maybe an SDK, how that's
00:05:16.127 -->
00:05:16.148Yeah.
00:05:16.168 -->
00:05:26.408customized globally and a layer on top of that, how it's customized for your particular project. This is how people typically talk about harness engineering.
00:05:28.067 -->
00:05:29.987plus other guides
00:05:29.987 -->
00:05:30.007Yeah.
00:05:30.048 -->
00:05:30.867and other sensors.
00:05:31.398 -->
00:05:34.858So the markdown that's part of Agents dot MD, any skills,
00:05:34.858 -->
00:05:35.197Mm-hmm.
00:05:36.077 -->
00:05:44.997any other reference files, even things that are automatically generated but inserted ahead of the agent doing its task.
00:05:45.677 -->
00:06:02.233And then the sensors that are relevant for your project or your system, like every time I run git commit, automatically run these tests and if there's a test failure, you know, exit loudly and it redirects to H.
00:06:02.233 -->
00:06:19.048Gotcha. So so the harness can be very broad, like it's not only the prompts, not only the project instructions and the skills and even the tools, but I think you're referring to like it's much broader, like tests, oracles, hooks probably as well, maybe profilers,
00:06:19.048 -->
00:06:19.668Plugins.
00:06:20.208 -->
00:06:25.067plugins, permissions, sandboxes, like everything in that environment.
00:06:25.528 -->
00:06:25.887Is that right
00:06:25.887 -->
00:06:25.927Yeah.
00:06:26.583 -->
00:06:57.642Yeah. So, I think this is important. When we talk about harness engineering, we can unambiguously talk about this whole thing if you just say harness online, different people will interpret it different ways. But the community has come together and there's a consensus definition for harness engineering, which is to mean that all that stuff as part of the harness that we're engineering, which means we're sort of taking care and setting all these things together holistically to build a system that helps us when we're solving our problems.
00:06:57.838 -->
00:07:05.858so it sounds like if you do harness engineering really well, you can push performance way higher than what a bare model can do.
00:07:06.677 -->
00:07:16.038And I remembered from our earlier conversation, you mentioned that that you listen to an interesting discussion as part of a podcast, another podcast that you listen to.
00:07:16.333 -->
00:07:16.572Yeah.
00:07:17.153 -->
00:07:36.088Yeah, so so you're going to listen to a clip from the Moonshots podcast, one of my favorite podcasts, where[UH] Dave Blundin is discussing that if you engineer your harness in a particular way, you can basically destroy this particular benchmark.
00:08:18.098 -->
00:08:38.018so so Dave used the term scaffold, but he means harness [laughter] in this in this context. So as as we heard, getting your harness right for a particular problem set or a particular focus can push your performance way higher than even an incrementally smarter model can. what we saw was going
00:08:38.018 -->
00:08:38.038Yeah.
00:08:38.357 -->
00:09:00.913from Fable to Opus-5 and even Opus-4 to Fable is dwarfed by having your harness set up properly with a with one of the dumber models. So. So what does this mean? I mean basically one, this is why harness engineering matters, and it's why lots of people have been talking about it, OK?
00:09:00.913 -->
00:09:03.048Mm-hmm. Yeah, indeed.
00:09:00.913 -->
00:09:11.442Because [UM] it was it wa- it was not only the folks on the Moonshot podcast, but we've also seen people [UM] discussing this on X.
00:09:11.982 -->
00:09:18.717I saw Dax from Open Code and Anthropic staff[UM] also touching on this topic. Maybe we can
00:09:18.717 -->
00:09:18.758Yeah.
00:09:18.758 -->
00:09:19.498pull up the tweets
00:09:19.793 -->
00:09:49.682Yeah, let's pull that up. So, so Dax is quoting something from the Anthropic stuff, so let's pull that up first so basically Jess is saying here this is the relevant statement. I think I'm quite biased, but I also think that it is impossible to get the maximum performance without tying together the harness and the model. Now, it's important to note that she's talking about harness in the context of cloud code and specifically in the context of a general-purpose harness.
00:09:50.363 -->
00:09:53.633So Dax is quoting this, OK?
00:09:54.153 -->
00:09:54.373Mm-hmm.
00:09:54.373 -->
00:10:06.778And he's saying people want to dunk on this by saying their weekend harness does better on some dumb one-shot benchmark. But at scale, this is true.
00:10:03.668 -->
00:10:09.832These tools are complicated and must factor in user behavior. And on and on and on.
00:10:10.493 -->
00:10:10.793Mm-hmm
00:10:10.832 -->
00:10:19.217So, so what is he saying? If you're trying to build a really good general-purpose harness, it's hard.
00:10:19.513 -->
00:10:19.812Mm-hmm.
00:10:20.352 -->
00:10:27.493It's hard and what you get is you get something that is pretty good at everything but it doesn't one-shot benchmarks.
00:10:27.788 -->
00:10:28.067Mm-hmm
00:10:28.067 -->
00:10:36.023Now, Dax is saying this saying, oh, you know, people are trying to build these these harnesses and they're thinking they're good because they can one-shot benchmarks.
00:10:36.643 -->
00:10:49.143I I think that's the wrong take. What the other way to look at this is to say it's possible to one-shot benchmarks by building a specific harness. This is the point of harness engineering.
00:10:47.062 -->
00:10:49.143Make your harness specific.
00:10:49.437 -->
00:10:49.817Mm-hmm.
00:10:49.918 -->
00:10:57.023Right[UM] there's a big consensus these days around harness engineering [UH] and you know we also talked about it last time.
00:10:57.482 -->
00:10:57.842Mm-hmm.
00:10:58.138 -->
00:10:58.238But
00:10:58.238 -->
00:10:58.258Mm.
00:10:58.857 -->
00:11:04.738people when they talk about harness engineering they're talking about the harness at the scope of their project. what matters
00:11:04.738 -->
00:11:04.758Mm-hmm
00:11:05.457 -->
00:11:06.618is specializing the harness.
00:11:07.118 -->
00:11:07.138Mm-hmm
00:11:07.432 -->
00:11:08.113To your work.
00:11:08.633 -->
00:11:09.332To your task.
00:11:09.623 -->
00:11:18.543Mm-hmm. Just the way how Dax was describing, if you specify it, if you adapt it to that specific benchmark, it will perform way better.
00:11:19.023 -->
00:11:21.962So, why don't we do that for our specific tasks?
00:11:22.503 -->
00:11:28.488And I think that definitely makes sense. I can imagine how this could apply in our own workflows.
00:11:28.988 -->
00:11:29.067Mm-hmm.
00:11:29.067 -->
00:11:37.383Like, a lot of[UH] projects that we probably a of people work on like a single project can contain very different kinds of work.
00:11:37.982 -->
00:11:48.383Maybe there's[throatclearing] a part that's[UM] front-end implementation, there's back-end architecture, there's performance optimization, maybe there's three-D modeling or data migrations.
00:11:49.182 -->
00:11:52.082And why should all of them use the same harness, right
00:11:52.298 -->
00:12:18.057Exactly. And I I know if you're like me maybe you've come across these kind of one-off instances through your work and I'll show a couple of them[UM] but it wasn't until this week that I realized you know actually what we should be doing is thinking about harness engineering holistically at these finer grained levels. so in order to maybe help you understand what we're saying, we
00:12:18.057 -->
00:12:18.077Yeah.
00:12:18.077 -->
00:12:19.378can pull up some examples. Yeah.
00:12:19.457 -->
00:12:20.378Yes, please
00:12:20.673 -->
00:12:39.467OK, so I'm gonna I'm gonna start with a solution to a problem that I've seen people talk about for months. I responded to a tweet even yesterday about this and it's something that you can solve in different ways depending on the kind of work that you're doing. So, what is the problem?
00:12:37.118 -->
00:12:41.182When I want to run multiple
00:12:41.182 -->
00:12:41.503Mm-hmm.
00:12:41.923 -->
00:12:45.143sub-agents or agents in parallel that do the same work.
00:12:45.437 -->
00:12:55.263What do I do if that same work is dependent on a shared resource so to be more specific, let's start with iOS I want
00:12:55.322 -->
00:12:55.423Yeah.
00:12:55.602 -->
00:13:22.732multiple agents to be testing in an iOS simulator and doing builds at the same time. Well, by default[laughter] that doesn't work because[UH] by default these things are they they conflict with each other. And so so this was a project I did in January in my in my harness I have something that that says that has a configuration[UM]
00:13:22.832 -->
00:13:23.113Sorry,
00:13:23.113 -->
00:13:23.312per agent.
00:13:23.312 -->
00:13:24.232f- just for the listeners
00:13:24.232 -->
00:13:24.413Mm-hmm.
00:13:24.653 -->
00:13:29.173who are not, who don't, who don't, who are listening on Spotify and don't have video, what are
00:13:29.192 -->
00:13:29.193And
00:13:29.193 -->
00:13:29.253you
00:13:29.253 -->
00:13:29.312good
00:13:29.312 -->
00:13:29.332showing
00:13:29.332 -->
00:13:29.592idea.
00:13:29.633 -->
00:13:30.592exactly on the screen?
00:13:30.888 -->
00:13:55.038Yeah, sorry about that so what I've pulled up is a terminal window and I've opened a an iOS project that I was working on. And within that iOS project I'm gonna open up in agents dot MD and just show you at the top of here it says Agent One simulator configuration and then there's a bunch of stuff. And there's a lot of these like hard-coded things in it[UM] but what this is what this is saying is
00:13:55.038 -->
00:13:55.077Yeah.
00:13:55.077 -->
00:13:56.337it's giving the agent
00:13:56.437 -->
00:13:56.457Yeah,
00:13:56.638 -->
00:13:58.118some specific commands that
00:13:58.118 -->
00:13:58.138yeah.
00:13:58.197 -->
00:13:59.498it can run to have
00:13:59.597 -->
00:13:59.618yeah.
00:13:59.638 -->
00:14:00.138a conflict-free
00:14:00.138 -->
00:14:00.158yeah.
00:14:00.873 -->
00:14:01.212access
00:14:01.212 -->
00:14:01.232Yeah.
00:14:01.273 -->
00:14:04.113to simulators. And the way that I set that up
00:14:04.293 -->
00:14:04.312Yeah.
00:14:04.552 -->
00:14:05.013is I have a
00:14:05.013 -->
00:14:05.033Yeah.
00:14:05.072 -->
00:14:05.633script which
00:14:05.633 -->
00:14:05.653Yeah.
00:14:05.673 -->
00:14:08.528I've stored in the scripts folder that is
00:14:08.707 -->
00:14:08.727Yeah.
00:14:08.788 -->
00:14:11.962run by every agent to get a unique
00:14:12.062 -->
00:14:12.102Yeah.
00:14:12.783 -->
00:14:20.773environment. And when you do that then you're able to run multiple iOS simulators at the same time. So this is an example of how I've set up the harness.
00:14:21.732 -->
00:14:32.495Yes, I've put some stuff in agents dot MD, but I've also created tools to help me isolate my environment so that I can have multiple agents running in parallel to build my iOS app.
00:14:30.812 -->
00:14:32.495[laughter] with my harness
00:14:32.495 -->
00:14:32.595Hey.
00:14:32.595 -->
00:14:33.534engineering hat on
00:14:33.894 -->
00:14:34.495[laughter]
00:14:34.495 -->
00:14:40.804[UH] I would and if skills existed at the time that I did this I would put this into a skill so
00:14:41.100 -->
00:14:41.419Right.
00:14:41.720 -->
00:14:47.934that it doesn't have to pollute the context and waste attention for agents that are not using the simulator
00:14:48.529 -->
00:14:48.870Gotcha.
00:14:49.004 -->
00:15:08.610Yeah, I think this example's interesting because a lot of examples are just about context narrowing. So it's like turning off the right tools, turning off certain guides. But the harness that you're building up through harness engineering, it's not just about what goes or doesn't go in in context, but it's about the whole environment.
00:15:09.129 -->
00:15:12.554And in this case, this is about configuring the execution environment.
00:15:13.154 -->
00:15:16.240It's setting up the execution environment to be able to run in parallel.
00:15:16.860 -->
00:15:25.919Gotcha. So [UM] I guess this example is more to show what can we set up in that environment to really help this agent to succeed in this specific task.
00:15:26.559 -->
00:15:27.940OK. Now I get it.
00:15:26.559 -->
00:15:27.940[laughter]
00:15:27.940 -->
00:15:43.799Exactly. Exactly. And I'm gonna just show I can quickly show an example of this same thing on the web. And I only this is exactly the same thing but for a web project Cuz I keep seeing people struggling with this[laughter] if you want to have concurrent sessions with let's
00:15:43.799 -->
00:15:44.019Mm-hmm.
00:15:44.019 -->
00:15:49.039say like the Chrome DevTools MCP. So what I've done is I've shared the docs from Chrome DevTools.
00:15:49.600 -->
00:16:02.110There's a section on concurrent sessions. There's a certain amount of flags that you pass in and there's some instructions here. You can feed this to your agent and ask your agent to set you up a harness to set up your environment so that it behaves in this way.
00:16:02.404 -->
00:16:02.985Oh, cool.
00:16:04.639 -->
00:16:04.640So
00:16:04.640 -->
00:16:11.039Yeah, we will add the link [UM] in the description too, so people who are only listening can find it too and watch later if they want.
00:16:11.335 -->
00:16:20.440Yeah So yeah, so that's, I think. This is an interesting example that came up in the past. For me, that doesn't have to do with context narrowing But context narrowing is a big part of it.
00:16:20.735 -->
00:16:20.754Mm-hmm.
00:16:20.754 -->
00:16:27.014So[UH] And we even have a whole section on context narrowing later [laughter]
00:16:27.565 -->
00:16:31.524Yeah, this is a very [UM] interesting example where you.
00:16:32.284 -->
00:16:40.664Yeah, I like it. I[UM] it's a good example that shows how you can give an agent [UM] the right tools to, specific to that task.
00:16:41.225 -->
00:16:43.184I'm wondering [UM] because we when
00:16:43.184 -->
00:16:43.245Yeah,
00:16:43.304 -->
00:16:48.504we had our conversation previously, we also had a few other like general examples.
00:16:49.164 -->
00:17:06.720Maybe we can show those now too and [laughter] for folks who are only listening, maybe we can just [UM] voice over them just to give people a little bit of [UM][UM] [UH] give give them more sense of how how this could be applied for different kind of use cases
00:17:06.720 -->
00:18:07.855So, what I've done is I've shared a diagram. And this diagram says, front-end implementer environment. What this agent needs in order to build UI well. So this is for tasks that involve the front-end The kinds of things that you're thinking about when you're designing your harness, when you're and you're engineering your harness and you're iterating on your harness, for this specific focus, you wanna understand what the goal is and then you wanna think, OK, what is the context here? What are the things that I need to specify in in files and pull them in In this case, design systems, reference screenshots for example[UM] what are the tools that I need? What are the tools that this harness needs to be able to control? Maybe if you're doing front-end web you need browser tools, you need a way to take screenshots you need a way to write code For and then there's automated checks. These are like your sensors that run in a hook or as a result of certain commands. So for example you might wanna do a visual diff for an accessibility check or you know integration tests.
00:18:08.615 -->
00:18:20.559And then and then you wanna think carefully about your definition of done. The criteria that the agent can that the agent can use to check that it's actually done with this work. So that's
00:18:20.559 -->
00:18:20.940Mm-hmm
00:18:21.460 -->
00:18:22.740you can der- derive that from the goal.
00:18:23.279 -->
00:18:34.250For example in this case it's like making sure that there aren't any accessibility issues that the the reference screenshot that you passed in is the same as the screenshot on the web.
00:18:34.789 -->
00:18:37.549And if you specify that, you can add feedback loops into your harness.
00:18:38.130 -->
00:18:42.589That for example, you can say, you know, keep going until the screenshot is in place.
00:18:43.269 -->
00:18:44.049Right? So.
00:18:44.049 -->
00:18:44.390Yeah?
00:18:44.684 -->
00:18:51.670So this is this is like a simple front-end example. I'll just really more quickly talk through a few more so
00:18:52.490 -->
00:18:56.549Yeah, how would the environment for other kinds of agents look like?
00:18:56.845 -->
00:19:57.029Yeah, so here's for maybe a performance agent. so I'm sharing a new diagram here you know, it's helpful to think in these in these categories. So what's the goal what context do we need, tools, feedback, exit criteria. I'm not gonna talk through all of these specifics, but just I'll pull a little handful context might be a baseline measurement of how fast something is, tools you might want a profiler the feedback might be the benchmark results and the exit criteria is that the benchmark result hits a certain point and then I'll just pull up one more so this is a debugger agent. The goal might be finding the root cause of a bug. Context might be things that you've seen, things that you've tried repro steps[UM] The tools, you might want a a debugger. This is like similar to the thing I talked about last time where I built a custom debugger for debugging my Game Boy projects[laughter]
00:19:57.190 -->
00:19:57.529Mm-hmm.
00:19:58.164 -->
00:19:59.414Yeah, and then and so on.
00:19:59.894 -->
00:20:00.194So.
00:20:00.194 -->
00:20:03.234This is very helpful.
00:20:00.194 -->
00:20:25.690This really [UM] [UH] helps me also see, like, when you mentioned, oh, build a world for your agent specific for to that task, how that could look like[UM] it also just just a a a quick side note, it also emphasize or shows me how a harness is much more than just a skill[UM] like skills are very useful, but a skill is usually just a additive.
00:20:25.690 -->
00:20:25.890Yeah.
00:20:26.230 -->
00:20:43.670It gives the agent instructions, knowledge, examples or maybe tools, but the harness environment that you just described and showed is much broader. It can add capabilities but[UM] basically it can also remove tools, restrict permissions[UM] or narrow available contexts.
00:20:44.589 -->
00:20:51.765So I guess maybe that brings us to is more capability for the agents always better?
00:20:52.809 -->
00:20:59.269so we we talked about what what should the agents have in their environments but also what should they not have.
00:20:59.769 -->
00:21:04.244And the reason why I ask is is because I got inspired by Tariq's post.
00:21:02.305 -->
00:21:12.759Maybe we can pull that up. Yes. So Tariq like I think it was a week ago when he sent out this [UM] post where he told
00:21:13.019 -->
00:21:13.180Yeah.
00:21:13.359 -->
00:21:33.085Claude Code users that he simplified the Claude MD file and I think he cut away like eighty percent of it but we this does not only need to be applied to context engineering you could also think about this from the perspective of harness engineering. Right?
00:21:34.125 -->
00:21:34.664What do you think
00:21:36.720 -->
00:21:48.474I think this is a really interesting an interesting thing to think about. So removing parts of your Markdown files is one thing. But [UM] let me pull up another tweet that was inspiring
00:21:49.414 -->
00:21:49.815Mm-hmm
00:21:50.015 -->
00:22:01.625that inspired interesting thoughts. So this is a tweet by Matt Pocock the creator of famous skills like grill me, et cetera[UM] [laughter]
00:22:01.744 -->
00:22:01.944Not
00:22:01.944 -->
00:22:01.984which
00:22:01.984 -->
00:22:02.164not real
00:22:02.164 -->
00:22:02.384not
00:22:02.484 -->
00:22:03.585me [laughter]
00:22:03.585 -->
00:22:30.220drill me, no [UM] so [UH] so Matt Matt posted this tweet the other day and[UH] I think the important part is here, the specs my skills create are intended to be deleted immediately, not kept around or treated as source code and he further clarified in a response and there was some discussion on this that this is because it poisons the context of your agent.
00:22:30.515 -->
00:22:30.634Mm-hmm
00:22:30.634 -->
00:23:00.819Because so the spec is used to derive the code the first time but as your agent starts iterating on the code as you make changes when you're reviewing as you add more features the code and the spec starts to drift. And you can try to keep them together but but it's tough and and it's not only about being careful not to put that context automatically in your agent's context[laughter] but even having it in your repo in your folder is enough to poison the mind of the agents.
00:23:01.440 -->
00:23:05.815So this is actually a real problem that I've been experiencing [laughter] and
00:23:05.815 -->
00:23:05.894Mm-hmm
00:23:05.894 -->
00:23:35.265and [UH] and didn't realize that I should do something about it until this week I have been checking in my my specs and frequently in the reasoning of my agents when I'm working with my agents I see something like, oh I thought the spec says this but when I went in the code and I tried to change it I noticed that actually the code says this other thing so I'll just forget what the spec says. That reasoning trace is polluting the mind of the agent.
00:23:32.710 -->
00:23:35.464It's distracting it from its work.
00:23:36.105 -->
00:23:41.025And so actually thinking carefully about how accessible is that information?
00:23:41.404 -->
00:23:54.025Even if it's not there, if it's accessible by the tools that the agent uses to search around, then you're still at risk of getting your contacts poisoned[UM] and I I think there's one other one other example of this. I mean,
00:23:54.085 -->
00:23:54.345Yeah.
00:23:54.345 -->
00:24:33.859the same sort of thing but so I use Beads Beads Rust which is a tool for task management. At some point we'll probably talk about this in more detail in an episode but the same thing. I had a conversation with a friend the other day and he was saying, oh how can I how can I use Beads but not check in the the Beads? And I'm like, hey that's the point. Like you wanna check it in because you you know that information's useful you can do a sort of all sorts of analysis on your data and and these kinds of things but but it's the same problem. If it's available and it's a spec basically it's like a description of a task not the actual result of the task then that can poison the content of your agent.
00:24:34.460 -->
00:24:48.289And that that then leads to this idea of OK well I should have two separate harnesses two separate environments. In one environment I have access to all the history beads, all the historical stacks.
00:24:45.434 -->
00:25:05.805Hopefully [UH] like ideally the reasoning traces of all my chat history. And I can do retros in that environment [UM] and in my environment where I'm actually implementing my work, I need to get rid of all that stuff. It needs to not even be available[UH] and accessible to the agent. This is something that I I only realized this week. So
00:25:05.805 -->
00:25:08.984[laughter] [UM] OK, no, that's super helpful.
00:25:09.565 -->
00:25:19.525I I could see I could see the context pollution [UM] point[UM] and also how we can how removing them can improve the work itself.
00:25:20.105 -->
00:25:33.859Cuz basically, agents have finite context, they have finite attention, they have finite tokens and finite compute. So every irrelevant spec or document loaded into the context will compete with something that is relevant.
00:25:34.440 -->
00:25:39.980So and also like every unnecessary tool that we introduce will also add another possible direction.
00:25:40.420 -->
00:25:50.640So this every generic instruction the agent must carry through the trajectory will consume attention that could have gone towards solving the actual problem.
00:25:51.200 -->
00:25:55.865Yeah. and you might think, OK OK, I'll I'll I'll keep my specs separate, you know?
00:25:55.924 -->
00:25:56.244Mm-hmm.
00:25:56.244 -->
00:26:07.609But it's not just specs, it's code too I wanna show another tweet that came up recently by by Victor Talin the creator of the BEND system, a prolific tweeter
00:26:08.529 -->
00:26:08.890[laughter]
00:26:09.009 -->
00:26:17.119[UH] so Talin noticed he has long tweets so let me highlight the the relevant part [UM] He he was begging
00:26:17.119 -->
00:26:17.200Mm-hmm.
00:26:17.200 -->
00:26:19.515the AI to remove and re-derive the work.
00:26:19.974 -->
00:26:31.904But notice that this does not work. It has a massive bias to keeping what is there [UM] so he he talks about how he spent so long trying to get Fable to just not [laughter] not keep what's there and and he couldn't.
00:26:32.404 -->
00:26:44.460And then he realized he can't keep something he doesn't see. So what he did is he he adjusted his harness to remove the parts of the code that he wanted Fable to re-derive.
00:26:45.259 -->
00:26:49.075And[UM] and only include a description of what it's meant to do.
00:26:49.694 -->
00:26:50.654And that and that fixed it.
00:26:50.950 -->
00:26:56.365Interesting. Do you have any examples, like when would you want to hide parts of the code? Like what parts of the code would you like to
00:26:56.365 -->
00:26:56.444Yeah.
00:26:56.444 -->
00:26:56.704hide?
00:26:57.224 -->
00:27:08.829in this particular case[UM] so Victor's working on a a core kernel for his new proof checker in programming language.
00:27:09.369 -->
00:27:16.545And and he's really conscious to keep the code and the complexity as small as possible, cuz the problem is already complex
00:27:17.065 -->
00:27:17.484Mm-hmm.
00:27:17.884 -->
00:27:52.480and also when the when the code is simpler it's easier to optimize. So[UM] so so he had he has a part of that code that he felt is too messy and he wants it to be cleaner and so so I guess like in general if you're refactoring if you have some part of your system where the code is messy and that messiness is making it hard to understand both for you and for your agents it's making it hard for new features to get added in it's making it hard to optimize this is a technique, a practice that you can do to to convince your agents to rewrite it. It's just get
00:27:52.480 -->
00:27:52.640[laughter]
00:27:53.819 -->
00:27:54.779rid of it [laughter]
00:27:54.779 -->
00:27:55.099[laughter]
00:27:55.259 -->
00:28:01.900and only include a s- like a placeholder rather than having it there and telling the agent that you need to to redo it more
00:28:01.900 -->
00:28:01.960mhm.
00:28:01.960 -->
00:28:02.660thing, one more
00:28:02.660 -->
00:28:02.740Yeah.
00:28:02.740 -->
00:28:16.825thing, but before we move off of this so we talked about, like, specs and plans, how that can pollute your, your environment if it's included in the wrong places. We talked about how code in the wrong places can ruin your environment for certain tasks
00:28:16.825 -->
00:28:17.105Mm-hmm.
00:28:17.609 -->
00:28:18.930but even sensors,
00:28:18.930 -->
00:28:18.950Mm-hmm.
00:28:19.529 -->
00:28:29.335even like tools that run automatically, tests that run automatically can pollute your environment[UM] and how is this?
00:28:25.960 -->
00:28:29.914Well, every unnecessary sensor slows
00:28:29.914 -->
00:28:30.115Mm-hmm.
00:28:30.115 -->
00:28:31.390down your agent. And
00:28:31.390 -->
00:28:31.609Yeah
00:28:31.904 -->
00:28:46.029as as our agents are getting faster, as we're willing to accept dumber models that run faster, we are controlled by the grips of of Amdahl's law.
00:28:46.650 -->
00:28:46.769OK?
00:28:46.769 -->
00:28:46.930[laughter]
00:28:47.220 -->
00:28:50.180So I found the article.
00:28:47.220 -->
00:28:55.315It was not a tweet, it was an article. so this was on the front page of Lobsters today[UH] it's a a post by Martin Alderson.
00:28:55.775 -->
00:29:01.855The title is I'm in parentheses mostly picking models on speed now, not intelligence.
00:29:02.595 -->
00:29:19.309So the point, I guess the whole point of his article is that he is reaching for dumber models that run faster. But here's the interesting part. There's a nice diagram that he built which is an illustration of, of Amdahl's law let's talk about this diagram first and then I'll explain Amdahl's law.
00:29:19.930 -->
00:29:19.950Mm-hmm.
00:29:19.950 -->
00:29:26.055OK? So So what is this diagram of on the bottom, on the X-axis, you see seconds per agent turn.
00:29:26.575 -->
00:29:26.934Mm-hmm.
00:29:26.934 -->
00:29:30.434And the Y is just a bar chart, so there's two bars
00:29:30.974 -->
00:29:31.295Mm-hmm
00:29:31.315 -->
00:29:33.634At the top is a bar that's fifty tokens per second.
00:29:34.095 -->
00:29:50.930And you see this this bar chart, it's broken up into sub-bars At the bottom is inference, and then tool tool calls, and then you[UM] So inference is the time that it takes to do inference, tool calls is the time it takes to wait for a tool to run and then get a result back.
00:29:51.490 -->
00:29:58.470And then U is the time it takes your brain to think about what happened and then decide what to do next.
00:29:59.069 -->
00:30:09.615And [UM] the the interesting thing here that's counterintuitive which is a consequence of what is called Amdahl's is if you if you compare the bar chart at the top, the fifty tokens per second one,
00:30:10.214 -->
00:30:10.255Mm-hmm.
00:30:10.255 -->
00:30:11.174and the bar chart at
00:30:11.174 -->
00:30:11.214Mm-hmm.
00:30:11.214 -->
00:30:11.494the bottom,
00:30:11.494 -->
00:30:11.734Mm. Mm-hmm
00:30:11.835 -->
00:30:12.835the two fifty tokens per
00:30:12.835 -->
00:30:12.855Mm-hmm.
00:30:12.894 -->
00:30:13.295second one
00:30:13.434 -->
00:30:13.454Mm-hmm.
00:30:13.835 -->
00:30:13.954[UM]
00:30:13.954 -->
00:30:13.994Mm-hmm
00:30:14.250 -->
00:30:16.829this is a rate of inference speed.
00:30:17.329 -->
00:30:24.990And so in this diagram the author shows inference take forty seconds at the top and eight seconds on the bottom. And that's a five X speedup
00:30:25.190 -->
00:30:25.230Mm-hmm
00:30:25.869 -->
00:30:26.549in the model speed.
00:30:27.230 -->
00:30:27.490But
00:30:27.920 -->
00:30:27.940Mm-hmm
00:30:28.000 -->
00:30:31.720the tool call speed and how fast your brain thinks does not change.
00:30:32.240 -->
00:30:36.079And so on net, your agent only speeds up by two X.
00:30:37.039 -->
00:30:44.424So a five X increase in speed in tokens per second is only a two X increase in your agency. And this
00:30:44.424 -->
00:30:44.904Mm-hmm
00:30:44.910 -->
00:31:04.484is Amdahl's law in action. So Amdahl's law I'm not gonna explain it in detail but essentially it's [UM] it relates the how fast your speed up of some section of your work is to how fast you can realize that speed up in the whole system. Because basically
00:31:04.484 -->
00:31:04.744Right.
00:31:04.785 -->
00:31:07.904if you have some fixed amount of time for other parts, that's
00:31:07.904 -->
00:31:07.944Mm-hmm
00:31:07.944 -->
00:31:10.964staying the same, but it's becoming a larger percentage of the total time.
00:31:11.644 -->
00:31:11.884Mm-hmm
00:31:12.845 -->
00:31:17.305So[laughter] so what's the point of this? How is it relevant to what we're we've been talking about?
00:31:17.424 -->
00:31:17.625Yeah,
00:31:17.625 -->
00:31:17.744Well,
00:31:17.744 -->
00:31:19.565because we were talking about having a lot
00:31:19.565 -->
00:31:19.625Mm-hmm.
00:31:19.625 -->
00:31:20.424of tests,
00:31:21.325 -->
00:31:41.400Yes. That's exactly how it's relevant. so so so yeah, as the models are getting faster, as you're willing to pick up faster models, if you actually want your agent to be faster, you have to make sure to shrink the tool call time as well. so and that's not just getting your your your tools to be faster, but it's choosing which tools you need at any specific point. That's the point.
00:31:42.059 -->
00:31:42.220Mm-hmm.
00:31:43.180 -->
00:32:00.309So that so you're so[UM] going back to how we started talking about this topic, about the [laughter] Amdahl's law. when there are so many tests in your own projects that have been accumulating over time while you were building it, you shouldn't keep all of them. Is that what you're or
00:32:00.490 -->
00:32:00.730Exactly.
00:32:00.730 -->
00:32:01.069you shouldn't
00:32:01.069 -->
00:32:01.230I mean
00:32:01.230 -->
00:32:01.250run
00:32:01.250 -->
00:32:01.269this
00:32:01.269 -->
00:32:01.369you
00:32:01.369 -->
00:32:01.589is.
00:32:01.670 -->
00:32:03.549shouldn't run them all the time for every
00:32:03.845 -->
00:32:03.944Yeah.
00:32:04.085 -->
00:32:04.345agent
00:32:04.845 -->
00:32:11.585So this is and this is just one kind of sensor, this is one kind of feedback that you have.
00:32:08.765 -->
00:32:12.505But let's use this as an example to ground ourselves
00:32:12.565 -->
00:32:13.005Mm-hmm.
00:32:13.565 -->
00:32:13.845yes.
00:32:14.569 -->
00:32:24.769what's happened to me several times in my projects is I have my agents write tests as they go and I have all the tests run automatically after each task so that there's no regressions
00:32:25.430 -->
00:32:25.890Mm-hmm.
00:32:25.990 -->
00:32:29.710but what happens is the the test time starts building up and up and up and it's slow
00:32:29.890 -->
00:32:30.349Yeah.
00:32:30.369 -->
00:32:57.440and and this is a problem that people might have run into in the olden days before agents as well it's just more acute now because the tests are running more often and so it's important to to take a step back and saying, OK, which test do I actually need for this part of my work? Which test can I run in parallel which tests are actually duplicates of each other [UH] so it's worth like going back and and that's part of the effort of like tending to your harness engineering garbage that you should be doing all the time.
00:32:57.940 -->
00:32:58.200And I
00:32:58.200 -->
00:32:58.460So,
00:32:58.720 -->
00:32:58.839I
00:32:58.900 -->
00:32:59.099mhm.
00:32:59.119 -->
00:33:07.480guess in the context of our conversation now, you should think about these environments distinctly even within a project.
00:33:07.775 -->
00:33:07.994Mm-hmm.
00:33:08.890 -->
00:33:18.529could you actually implement that in practice in your workflow? How can you [UM] for example in my current workflows I also have like almost four thousand tests.
00:33:18.990 -->
00:33:26.789How can I tell my agents what is the best way to tell my agents how to group them or filter them?
00:33:23.890 -->
00:33:28.769And how do I even like hide certain files from some agents
00:33:29.630 -->
00:33:52.430So for the for the first part, for for grouping tests[UM] that's that's just a prompt. So open a fresh prompt and in a fresh harness you know [UM] tell your tell your agent, OK look through my tests[UH] benchmark them, tell me if there's any that are duplicated, parallelized where you can.
00:33:52.950 -->
00:33:57.769And and you can just work with your agent and and fix that[UM] because that's applicable for everything.
00:33:58.630 -->
00:34:01.309So that's that's one step that's easy.
00:34:01.829 -->
00:34:02.150Mm-hmm.
00:34:03.029 -->
00:34:04.690The second part, actually removing them.
00:34:05.849 -->
00:34:12.690This is there's a few ways I can think of to do this. It's not something I've done yet myself since
00:34:13.250 -->
00:34:13.550Oh, wait.
00:34:13.550 -->
00:34:13.690we
00:34:13.690 -->
00:34:13.730Before
00:34:13.730 -->
00:34:13.731just
00:34:13.731 -->
00:34:13.789before
00:34:13.789 -->
00:34:14.989discovered it. Mm-hmm
00:34:15.090 -->
00:34:23.489removing them [UM] when[UM] when we ask the agent to help us categorize or filter the tests.
00:34:23.949 -->
00:34:26.110Probably it's not only the duplicates, right?
00:34:26.489 -->
00:34:32.949Because for example, some tests are more tests that touches front-end, some others tech- touches the back-end architecture.
00:34:33.809 -->
00:34:36.829So how would you how would you tell the agent to make a distinction there
00:34:37.684 -->
00:34:37.885Yeah.
00:34:38.784 -->
00:34:49.204It's funny, this is actually something that I spent a lot of my engineering career solving in different companies [UH] and there's lots of different ways to do it.
00:34:49.664 -->
00:34:51.844And there's different like levels of complexity you can go in.
00:34:52.344 -->
00:35:15.244So[UM] so there's so like we said this is a this is a problem that has existed for a while and so there's there's solutions[UM] so So the a simple thing you can do is for each test you can assign which part of the code base it's associated with[UM] and so in somewhere where
00:35:15.244 -->
00:35:15.304Yeah.
00:35:15.304 -->
00:35:24.144you're registering tests in in your strip that's running the test. You can work with your agent and say, OK, you know, match this prefix[UH] for these tests.
00:35:24.605 -->
00:35:30.885And if that prefix doesn't match of of like files or folders that were touched in this change, then you do don't have to run those tests.
00:35:31.885 -->
00:35:38.105So that's that's a simple thing that you can do that your agent can probably one-shot[UM] Does that make sense? So, let let me be specific.
00:35:38.605 -->
00:35:42.005You could say [UM] if you have a front-end directory and a back-end directory,
00:35:42.905 -->
00:35:43.184Mm-hmm.
00:35:43.425 -->
00:36:01.905in your in your test runner, your agent would change it for you to register[UH] tests that say like, OK, these are all the tests that run if a file is touched under the front-end folder. And these are all the tests that run if a file is touched under the back-end folder. And if you touch a file in both the front-end and back-end folder, then all those tests are gonna be run.
00:36:02.425 -->
00:36:02.625And that's exactly
00:36:02.625 -->
00:36:02.744Mm-hmm.
00:36:02.744 -->
00:36:03.224what you want
00:36:03.900 -->
00:36:04.239And you you
00:36:04.239 -->
00:36:04.420[UM]
00:36:04.940 -->
00:36:15.780ask the agents to make the decision which test needs to be run when[UH] code in a certain folder is touched.
00:36:15.780 -->
00:36:28.295Yeah. So you can you can work with your agent and and decide, OK, these are tests that are important for front-end and these are tests that are important for back-end. Then you tell your agent, OK, label them and put it into my test runner so that they run automatically in those situations.
00:36:28.855 -->
00:36:31.014So that's that's a really simple thing you can do.
00:36:31.655 -->
00:36:38.735There's a problem with that approach though, which is it's it's something that you assign that you make a judgment on
00:36:39.275 -->
00:36:39.355Mm-hmm.
00:36:39.355 -->
00:36:41.414and your judgment could be wrong [UH]
00:36:41.715 -->
00:36:42.054[laughter]
00:36:42.295 -->
00:36:42.414or
00:36:42.414 -->
00:36:42.574Mm-hmm.
00:36:42.574 -->
00:36:43.775it or it can go out of date.
00:36:44.414 -->
00:36:44.554And so
00:36:44.554 -->
00:36:44.914Mm-hmm.
00:36:45.494 -->
00:36:51.554[UH] so so that that's where this breaks down. So the other way to do it, which is much more effort, but
00:36:51.715 -->
00:36:51.735Mm-hmm
00:36:51.735 -->
00:37:05.155there's a bunch of test test runners that[UH] go the other way around where[UM] where the your whole build and test system is built around
00:37:05.355 -->
00:37:05.375Mm-hmm.
00:37:06.034 -->
00:37:06.375the idea
00:37:06.375 -->
00:37:06.394Mm-hmm.
00:37:06.534 -->
00:37:07.534of understanding
00:37:07.534 -->
00:37:07.554Mm-hmm
00:37:08.934 -->
00:37:12.355the the boundaries and the surface area of all of your code files.
00:37:12.835 -->
00:37:47.815And so when your build system understands that, it natively knows which which tests are associated with which modules. And this is something that[UH] there's tools like Bazel or its derivatives, Buck[UM] these things are sort of you have to buy buy into them opt into it [UM] if you use to set up your project that's another tool to look into but these are very heavy weight and you'll have to invest a lot of energy in or you'll have to change the structure of your project
00:37:48.315 -->
00:37:48.335Mm-hmm.
00:37:48.335 -->
00:37:55.414agents can do that for you so it's not that much energy[UM] but but[UH] it's also not always possible if you have if
00:37:55.414 -->
00:37:55.795Mm-hmm.
00:37:55.914 -->
00:38:08.215you have in your project code that has a lot of different programming languages and you're doing lots of physical custom things, it can be really hard to fit them into these frameworks.
00:38:08.849 -->
00:38:09.110So.
00:38:09.110 -->
00:38:19.170It sounds like it might be better if you used a second method, that you just build start building the project in that way instead of trying to[UM] implement those later
00:38:19.650 -->
00:38:27.630Yeah. Yeah. Yeah, that's something[UM] that's something you can do up front [UH] and with the agents it's not even that hard to change in the
00:38:27.630 -->
00:38:27.690[laughter]
00:38:27.690 -->
00:38:30.150middle [laughter] to be honest [UM] but it's
00:38:30.150 -->
00:38:30.309I'm just
00:38:30.309 -->
00:38:30.449[UH]
00:38:30.449 -->
00:38:32.329worried about mistakes or,
00:38:32.329 -->
00:38:32.349I
00:38:32.349 -->
00:38:32.351you
00:38:32.351 -->
00:38:32.489would.
00:38:32.489 -->
00:38:33.769know, misses and gaps
00:38:34.769 -->
00:38:58.789Yeah, I the thing that I would that I would discuss with your agent when you're trying to dig into this[UM] if you're listening and and you wanna play with it[UM] just talk to your agent about the trade-offs of of these different approaches[UM] because yeah, you'll you'll be limiting what is possible[UM] when you have your your tests set up that way, where it's
00:38:58.789 -->
00:38:58.869Mm-hmm.
00:38:58.869 -->
00:39:04.010like all tracked by the build system tracks it. But the benefit is it's perfect. Like,
00:39:04.050 -->
00:39:04.210[laughter]
00:39:04.210 -->
00:39:09.869if you change this line of code, it knows exactly what tests are associated with it and nothing else will run.
00:39:11.030 -->
00:39:11.250Awesome.
00:39:11.250 -->
00:39:15.829So, yeah, so so that's how to do th- that that side of it.
00:39:16.909 -->
00:39:16.969OK,
00:39:16.969 -->
00:39:17.309And then,
00:39:17.309 -->
00:39:18.150I I could [UM]
00:39:18.150 -->
00:39:18.349m-
00:39:18.670 -->
00:39:25.550also imagine if [UH] the second option, I don't know, is is not an option for some people and they go with option number one.
00:39:26.010 -->
00:39:34.849You could also add a retro step and run a loop to [UH] make sure that any test that you have missed the first time will be captured afterwards.
00:39:35.905 -->
00:39:36.144Yes.
00:39:36.684 -->
00:39:43.625Yeah. Yeah, yeah. Yeah, you can have a, you can have a review agent[UM] run periodically to make sure that things are still aligned.
00:39:44.425 -->
00:39:44.644Mm-hmm.
00:39:45.385 -->
00:39:45.925Awesome.
00:39:46.625 -->
00:39:56.364Yeah [UM] and I guess, and then there's one more piece, which is, you know, we we talked about actually making it impossible for the harness to see the code
00:39:56.364 -->
00:39:56.764Yeah.
00:39:57.204 -->
00:39:58.445or the files or the context.
00:39:59.085 -->
00:40:19.025How can we do that[UM][laughter] I don't have a great solution[UM] I I think[UH] I think to do this, so to to do it kind of in a in an ad-hoc way is easy, right? You just move the file[laughter] You
00:40:19.025 -->
00:40:19.244[UM]
00:40:19.485 -->
00:40:20.405know [UM] and
00:40:20.405 -->
00:40:20.425[UH]
00:40:20.425 -->
00:40:31.744and and that's that can be enough a lot of the time. I mean, the example I showed with with Victor with Victor Taelin, he he just deleted that part of the code and replaced it with a comment.
00:40:32.364 -->
00:40:32.585And and
00:40:32.585 -->
00:40:32.684Mm-hmm.
00:40:32.684 -->
00:40:33.344and that can be enough
00:40:33.344 -->
00:40:33.364[laughter]
00:40:33.985 -->
00:41:03.425[UM] I think there's an opportunity to build some kind of tooling around automatically[UM] showing and hiding or exposing and and removing from your s- execution environment, your sandbox[UM] the stacks and and plans and tasks, issue tracker, that that kind of thing[UH] I don't know of a tool that exists for that. So for now, what I've been doing so far is just manually moving files around
00:41:04.900 -->
00:41:11.360Mm-hmm. But if you're[UM] like in our examples that we mentioned during this episode we we for
00:41:11.460 -->
00:41:11.699Mm-hmm.
00:41:11.699 -->
00:41:17.039example we talked about front-end agents and agents that work on maybe the back-end architecture.
00:41:17.360 -->
00:41:17.579Yeah.
00:41:18.079 -->
00:41:23.840I could imagine that[UM] you want to give the front-end agents maybe the style guide
00:41:24.719 -->
00:41:24.840Mm-hmm
00:41:24.840 -->
00:41:41.320[UM] but when the back-end agent is working you might not want to remove the file, right? You you do want to like [UM] you don't want it to pollute the back-end agent's contacts, but you don't want to remove it and then hopefully you won't forget to re-add it again when you tell your front-end
00:41:41.320 -->
00:41:41.719Well, yeah.
00:41:41.719 -->
00:41:42.500agent to do work.
00:41:43.054 -->
00:42:20.054This is the, I mean, this is the thing to set up. It is like, we saw some examples where you, you do need to remove it. Cuz even it existing is gonna distract your agents because they'll, they'll pick it up when they're searching for files. So[UM] you know, more concretely, like a simple way to do that, if you're already using worktrees[UH] so if you use Claude or Codex, it's just a check box [laughter] when you, when you start a session[UM] if you're using the the app[UM] [UH] the you'll have a fresh work tree and you can just delete the files and [UH] and then you have to remember or you tell your agent like, hey don't commit these[UM]
00:42:20.054 -->
00:42:20.094Mm-hmm.
00:42:20.094 -->
00:42:35.175But yeah, even then they're they're sort of visible. So I think there's I think there's an opportunity to to to build tooling to automate this more[UM] But I, yeah.
00:42:35.889 -->
00:42:54.269So even if you spawn a sub-agent and you specify exactly what what files should be shared with the sub-agent, you think maybe since they have [UM] access to the files, like what Taelin said, Victor Taelin said[UM] they they might still access it in some way.
00:42:55.349 -->
00:43:02.429I guess[UH] in Victor's case it was in the same file in [UM] in the example with Matt Pocock.
00:43:03.050 -->
00:43:03.070The
00:43:03.070 -->
00:43:03.150Mm-hmm.
00:43:03.150 -->
00:43:15.670specs were just living in different directories and I've seen it. I've seen it happen. The agent, when you when you tell it change the start of the l- code, it looks and finds a spec that describes it and then it loads into context even if you didn't tell it to.
00:43:16.349 -->
00:43:16.469Mm-hmm.
00:43:16.469 -->
00:43:22.230And so, yeah, it existing is is a problem. I think [UM] I think you can probably.
00:43:23.909 -->
00:43:30.730You know, you can hack your way around it for now. You can just delete delete the files and maybe add something in your prompt that says like don't look at them.
00:43:31.269 -->
00:43:33.230But but that's that's not a good solution
00:43:34.125 -->
00:43:51.985Yeah, it's it's a good [UH] an interesting insight and practice to think about though [UM] because sometimes, you know, after building it maybe you don't need the spec anymore and and then the way how you described it, move it to a different directory that only the retro agent or some other agent who really needs it can access it that that makes sense.
00:43:52.704 -->
00:43:52.706Duh
00:43:52.706 -->
00:43:58.565Yeah, definitely [UH] worth thinking about even though, you know, there's not not a mature product or tool built yet.
00:44:13.974 -->
00:44:15.875Is limiting the blast radius.
00:44:16.434 -->
00:44:23.554That could also be another reason why to narrow the scope or capabilities of your agent, right? Could you maybe, yeah.
00:44:24.175 -->
00:44:27.284Yeah, yeah. [UM] when we say blast radius, we
00:44:27.284 -->
00:44:27.304[laughter]
00:44:27.324 -->
00:44:38.625mean like when we have an ask, how much of the world gets[laughter] gets destroyed by the blast of the bomb that's sort of the metaphor. So like how how many files are touched, how much stuff is changed.
00:44:39.184 -->
00:44:49.369And so yeah, it's good practice to limit the blast radius of your actions when you're using multiple agents in parallel, or just in general because you have less change to review and there's less chance of something getting away.
00:44:49.849 -->
00:44:59.969And so the best way to limit your blast radius is to remove the parts of your environment from the world and then the bomb cannot touch them because they're not even there.
00:45:00.625 -->
00:45:19.880Right, right. So for example, if [UM] if I was building an agent who would check my CRM to decide which clients I have to contact or reach out to, then give it only read access. Don't give it well[laughter][UM] Yeah, only give it read access. Don't let it edit the database.
00:45:20.289 -->
00:45:23.840yeah, this is I mean, that's an example of a of like the danger [laughter]
00:45:24.500 -->
00:45:24.880Mm-hmm.
00:45:25.139 -->
00:45:41.519the danger argument, which is also something, I mean, we haven't really touched on that it is something that a lot of people talk about I suppose, but it's another part of the harness so yeah, as you're saying, like we've seen people accidentally lose their database delete their hard drive.
00:45:42.159 -->
00:45:45.875These are examples of a harness that's not carefully guarded
00:45:47.014 -->
00:45:58.094Yeah. And maybe, you know, don't even give them the tools that they don't need, like don't give that CRM agent the tool to send out emails if you only want them to propose which people to email
00:45:58.389 -->
00:45:58.570Yeah.
00:45:59.130 -->
00:45:59.510Exactly
00:45:59.804 -->
00:46:10.565Cool. So yeah, I guess it's [UM] a narrowing, so I guess when designing a harness it's good to ask yourself what information does this agent genuinely need?
00:46:11.085 -->
00:46:14.545What tools will really help it complete its task?
00:46:15.164 -->
00:46:20.925What feedback or tests should it run automatically? What should it never be able to do?
00:46:21.405 -->
00:46:33.625And then [UM] the goal of this all is not to create the most powerful possible agent that can do everything, but it's to create the agent that most likely can succeed at a very specific task.
00:46:35.525 -->
00:46:35.945Awesome.
00:46:37.025 -->
00:46:37.264Perfect
00:46:37.264 -->
00:46:40.425OK. [laughter] Cool.
00:46:37.264 -->
00:46:47.485Then maybe this is a good point to also move to our[UM] one of the last things that we also want to touch on and it's also very important and interesting to touch on.
00:46:48.125 -->
00:46:58.164We talked about how narrowing the harness can help the agents make the agents work better but sometimes mistakes still happen.
00:46:58.905 -->
00:47:04.445And there are probably also different ways of enforcements that we can apply.
00:47:05.300 -->
00:47:18.780and that reminds me of something that I was trying to do with OpenClaw in January[UM] I was building this personal agent and I wanted to it to help me schedule my day based on my calendar and my priority list.
00:47:19.539 -->
00:47:30.159And [UM] but when I first tried to build it with a simple prompt I I learned that it sometimes does not follow the instructions or missed things.
00:47:30.719 -->
00:47:41.039And I remember that we talked back then as well and[UM] after [UM] chatting and discussing it, I realized there are actually different layers of reliability.
00:47:43.519 -->
00:47:45.460Yeah. Tell us, tell us about the layers.
00:47:45.655 -->
00:47:49.275So there are four different enforcement levels for agents.
00:47:50.275 -->
00:48:00.414That is, level one is prompting, level two is using templates, level three is using scripts, and level four is cron and heartbeat triggers.
00:48:02.114 -->
00:48:05.755So, level one is natural language instructions.
00:48:07.155 -->
00:48:10.355For example, wind down work by eight thirty PM.
00:48:10.835 -->
00:48:21.434And this is fine for tone and communication style, but it's not reliable for anything where accuracy matters because the agent will only follow it follow it most of the time.
00:48:22.014 -->
00:48:22.034And
00:48:22.034 -->
00:48:22.135Right,
00:48:22.135 -->
00:48:22.494that's not enough
00:48:22.494 -->
00:48:22.655right.
00:48:23.474 -->
00:48:23.614Yeah.
00:48:24.155 -->
00:48:25.295OK, and then what about level two
00:48:26.215 -->
00:48:28.355[UM] level two, using templates.
00:48:28.954 -->
00:48:34.815So, you could force your agent to use a structure that he n- that it needs to fill in before acting.
00:48:35.135 -->
00:48:44.094For example, I told my agents, check the calendar and fill out this section. And the section is calendar output and then blank section.
00:48:44.855 -->
00:48:54.295And if that section is empty, then the plan does not get sent to me because then it means the agent did not do a good job at checking the the calendar before sending it to me.
00:48:54.875 -->
00:48:55.195So that's a
00:48:55.195 -->
00:48:55.255Right.
00:48:55.255 -->
00:49:00.034forcing function to[UM] make sure the the the failure becomes visible.
00:49:01.094 -->
00:49:08.954Right. And then the and then OpenClaw would get that feedback loop. If that makes sense. And then And then what about [UH] what about the next layer?
00:49:26.054 -->
00:49:31.755So then the agent can read the output instead of interpreting raw API data.
00:49:32.275 -->
00:49:41.534And a script cannot misread a meeting time or forget to check[UM] so that is another enforcement level that's even more reliable.
00:49:43.594 -->
00:49:45.175Great. And then what's the last one
00:49:52.894 -->
00:50:11.175Fire on a schedule, then the agent does not need to remember at all to, for example, plan your day or ping you, send you a reminder at some[UH] [UH] at a certain hour of the day. The cron job will just fire and force it. And the heartbeat, you could also have a heartbeat that pulls every thirty minutes or something.
00:50:11.795 -->
00:50:25.815And I found that[UM] after I[UM] used these different layers for different things that I wanted my personal agent to do, finally it started [laughter] to be reliable enough[UM] for the task that I gave it.
00:50:27.574 -->
00:50:34.914And I could also imagine that [UM] these kind of enforcement levels could also be applied to harness engineering.
00:50:35.204 -->
00:50:47.684Yeah. I think so. The[UH] this like nomenclature I think is really interesting the there was a there was a post by John DeGoes this week. I'll just show it quickly.
00:50:48.340 -->
00:50:50.039Your old agent architecture is dead.
00:50:50.079 -->
00:50:50.119about
00:50:50.179 -->
00:50:50.380I mean,
00:50:50.519 -->
00:50:50.579that?
00:50:50.579 -->
00:51:22.135it's replacement John DeGoes is the CEO of Ziverge Tech he's the creator of ZIO, which is an inspiration to Effect and just in general like a functional programming guru and anyway this this post is interesting but in the very beginning he introduces this concept of of hope and and then how to fix it with enforcement. And I think that's like the best lens in which to look at what you were sharing regarding OpenClaw.
00:51:23.135 -->
00:51:23.195So.
00:51:23.195 -->
00:51:27.980Cool. How would you describe the different enforcement levels when people are building harnesses
00:51:28.269 -->
00:51:30.969So [UH] so there's you
00:51:31.050 -->
00:51:31.409and when
00:51:31.769 -->
00:51:32.170can think of it
00:51:32.170 -->
00:51:32.250people are
00:51:32.250 -->
00:51:33.010like a sliding scale
00:51:33.010 -->
00:51:33.030building
00:51:33.289 -->
00:51:33.469between
00:51:33.469 -->
00:51:33.489harnesses
00:51:33.530 -->
00:51:52.005hope and enforced. So so at all the way on the hope level, it's all in markdown it's all context driven, it's all non-deterministic, it's all the what will happen will be generally what you want but maybe not exact, and depending on how how complex your ask is.
00:51:49.304 -->
00:52:04.250And this is what you are seeing with layer one so and then as each of those layers that you described is moving us more and more towards enforcement, towards determinism, towards[UH] software one point O
00:52:04.250 -->
00:52:04.429a good
00:52:04.429 -->
00:52:04.690that does
00:52:04.690 -->
00:52:04.769point.
00:52:04.769 -->
00:52:28.914exactly what we want and ultimately the best designed harnesses are doing as much as possible in the determinism side, essentially, with speculative evaluation, with branch prediction I'm just using metaphors here, but with also guidance on the hope side. So this is exactly what we talked about in the last episode But just briefly to say it again
00:52:29.414 -->
00:52:29.715[laughter]
00:52:29.795 -->
00:52:50.175[UM] if you only have the determinism, if you only have your rules that yell at you if you do something wrong, then the agent is likely to do it wrong first and then fix itself [UM] if you only have the guidance in Agents dot MD, a lot of the time it'll do what you want the first time, and sometimes it'll fail. And so we
00:52:50.175 -->
00:52:50.394Yeah.
00:52:50.684 -->
00:53:03.289we want the best of both. We want to most of the time do the correct thing, and in the cases where we make a mistake, we correct it so that covers us to the layer two, I think. And then that
00:53:03.289 -->
00:53:03.309Yeah,
00:53:03.309 -->
00:53:03.349last
00:53:03.349 -->
00:53:03.469I think
00:53:03.750 -->
00:53:04.570layer that you talked about,
00:53:04.570 -->
00:53:04.590that's
00:53:05.010 -->
00:53:05.789the trigger,
00:53:05.869 -->
00:53:05.889a
00:53:06.150 -->
00:53:31.480this is just another part of the harness that we haven't talked about yet so how the how the agent actually starts doing its work is is another thing that you can control either non-deterministically by an orchestrator or if possible deterministically with a cron for example that runs on a timer and automatically deterministically always does that task at a certain time.
00:53:32.599 -->
00:53:41.079I definitely noticed when [UM] you go the deterministic route, your success rate is much higher[laughter] than staying on level one
00:53:42.159 -->
00:53:49.139And and this is this is exactly why context engineering is is not enough. And we need harness engineering.
00:53:49.719 -->
00:54:04.489Because if you're thinking in terms of context engineering, you're only in that hope world. you're[UM] you're changing your agent's dot MD. Maybe you're making lots of scales, you're making lots of different markdown files. But all of these are hopes. All of these are non-deterministic. All of these work sometimes but not all the time.
00:54:05.070 -->
00:54:11.090And you need that you know old school code.
00:54:08.250 -->
00:54:20.239You need that determinism. You need real software that's checking things and doing things and triggering things and that's a really important part of the harness and we can't leave that out
00:54:20.525 -->
00:54:20.905Yeah.
00:54:21.619 -->
00:54:54.855that hundred percent makes sense. I think I learned that when the cost of failures is higher[laughter] then the m- system should depend less on hoping that the model remembers a certain sentence[UM] in one of the[UH] markdown files [UM] I could also imagine that maybe sometimes not every smaller task needs a custom environment. Yeah, I I could imagine that maybe for a one-line copy change[UH] building a a a whole environment is not would be absurd
00:54:54.875 -->
00:54:56.405Yeah. Yeah. Yeah. And
00:54:56.405 -->
00:54:56.425Yeah,
00:54:56.425 -->
00:54:56.465it.
00:54:56.625 -->
00:54:58.960are there how would you describe like the type of
00:54:58.960 -->
00:54:59.159Mm-hmm.
00:54:59.559 -->
00:55:01.840situations to spend more
00:55:01.840 -->
00:55:01.980Yeah.
00:55:02.059 -->
00:55:03.800time on designing
00:55:03.800 -->
00:55:03.920Mm-hmm.
00:55:03.920 -->
00:55:04.280the harness
00:55:04.570 -->
00:55:06.815Yeah Yeah, I think this is also a scale.
00:55:07.269 -->
00:55:10.039so let's talk through the different endpoints and then the middle.
00:55:10.579 -->
00:55:10.739So
00:55:10.739 -->
00:55:11.099Mm-hmm.
00:55:11.295 -->
00:55:16.335like you mentioned, on one end if it's, you know, a one-line copy change, don't waste your time with this.
00:55:16.815 -->
00:55:27.375It's low stakes and the cost of the mistake is minimal as on the other side the complexity grows, safety becomes more paramount, more important,
00:55:27.695 -->
00:55:27.755Mm-hmm.
00:55:27.875 -->
00:55:30.114and And I guess in that world
00:55:30.894 -->
00:55:31.014Mm-hmm.
00:55:31.394 -->
00:55:32.034it's worth spending a
00:55:32.034 -->
00:55:32.074Mm-hmm
00:55:32.155 -->
00:55:59.465lot of energy and a lot of effort on your harness and actually taking the time building it, thinking carefully, going through loops, the human agent feedback loop and getting something really good before you start your test And then and then there's stuff in the middle. So, something that[UH] actually John mentioned in his article[UM] that [UH] I started playing with[UH] and and it seems to work. know, you don't have to design the harness yourself. You can just tell your agents to do it.
00:56:00.045 -->
00:56:09.414So, you can for example create a[UH] a nice sub-agent graph where the first step is to have a bootstrap agent build a harness for you for your task.
00:56:10.335 -->
00:56:11.695And and that that I found
00:56:11.695 -->
00:56:11.715Mm-hmm
00:56:11.735 -->
00:56:25.494works well for these in the middle cases. So you have a you have a reasonably complex problem you're trying to solve. Maybe the stakes of making a mistake are not too high, like it doesn't send an email to your boss [laughter] with something rude
00:56:25.494 -->
00:56:25.755Mm-hmm.
00:56:25.835 -->
00:56:28.019in it but you know, you're changing some code.
00:56:28.199 -->
00:56:34.190You're you're doing some, you know, large task and just asking your agent to do this for you
00:56:34.530 -->
00:56:34.610Mm-hmm.
00:56:35.909 -->
00:56:36.909is something that's helpful.
00:56:36.945 -->
00:56:48.844so it sounds like the amount of harness engineering should scale with factors like the difficulty of the task, the cost of failure[UM] maybe the number of different failure modes.
00:56:49.465 -->
00:56:58.844And probably I could imagine that if it's [UM] a recurring class of task, you might want to invest more time on building a reusable harness.
00:56:59.425 -->
00:57:00.380Yeah, that's a really good point
00:57:00.380 -->
00:57:19.119Yeah, so it sounds like that not every subtask needs the general-purpose harness and the elaborate infrastructure[UM] building the right harness that is better adapted performs really well on a specific task or bottleneck. So that is where the large productivity gains can come from.
00:57:19.699 -->
00:57:23.360So I guess the takeaway is match the harness investments to the task.
00:57:24.420 -->
00:57:24.800Is that right?
00:57:25.090 -->
00:57:25.550Exactly.
00:57:25.844 -->
00:57:26.264Alright.
00:57:26.824 -->
00:57:44.320OK. Well maybe to wrap things up, when setting up the environment for our workflows, what we often do is we set up the tools, the skills, the contacts, even hooks[UM] for our h- entire workflow. And we will have all our agents and sub-agents reuse all of them.
00:57:44.940 -->
00:57:51.280But hopefully, after listening to this episode, you're asking yourself also another question.
00:57:51.780 -->
00:57:58.440What harness could we build that is adaptive to this specific task and makes the outcome better?
00:57:58.920 -->
00:58:02.500Or what environment would make it hard for that agent to do wrong?
00:58:03.300 -->
00:58:11.380So, that question might change how you think about doing harness engineering from here on out and how you focus your harness on a problem you're working on.
00:58:12.880 -->
00:58:15.199Well, this was episode four of Superlinear.
00:58:15.820 -->
00:58:18.260We hope that you got something out of this episode.
00:58:18.820 -->
00:58:27.139Have a lot of fun with harness engineering. Please tell us in the comments about what you build and what you notice in your approach. Brandon and me are learning as well.
00:58:27.840 -->
00:58:36.300And we'd love to hear about what you experienced in your workflows and don't forget to subscribe so you don't miss out on our next episode.
00:58:37.019 -->
00:58:38.579Thank you all and see you next time.
00:58:39.139 -->
00:58:39.679Bye everyone.