אודות פרק זה
A million-token context window doesn’t mean you should fill it.
Every piece of information you keep in a coding agent’s active context has a cost. Not just in tokens and dollars, but in the model’s attention.
In this episode, we explore context hygiene for coding agents: how to decide what deserves to stay in active context, what to move elsewhere, and when to throw context away entirely.
Superlinear is a podcast about emerging practices for building with AI.
Hosted by Brandon Kase and Christine Yip.
Homepage: https://superlinear.fm
Twitter: https://x.com/superlinear_fm
LinkedIn: https://www.linkedin.com/company/superlinearfm
Brandon Kase: https://x.com/bkase_
Christine Yip: https://x.com/christinetyip
Music licensed through Soundstripe. Code: O2UT6CCEQNH4KT0D
הראה הערות 🔗
תעתיק 🔗
00:00:00.012 --> 00:00:02.871
as the context grows the models are less intelligent.
00:00:02.874 --> 00:00:10.144
there's a dollar cost attached to how you handle your context[UM] but it's not only tokens that's finite.
00:00:10.173 --> 00:00:29.410
Attention is also finite. for every prompt that I send I ask myself does this information deserve to occupy the active context of the agent doing this task? It started maybe as tiny productivity practices like clearing or compacting and if
00:00:29.410 --> 00:00:29.469
[breath]
00:00:29.469 --> 00:00:29.949
you follow this
00:00:29.949 --> 00:00:30.149
[laughter]
00:00:30.149 --> 00:00:31.309
idea far enough,
00:00:31.329 --> 00:00:31.330
[breath]
00:00:31.330 --> 00:00:39.810
like what we did and move all the way across the spectrum, it can also become architectural decisi- decisions that you're making along the way.
00:00:39.912 --> 00:00:50.784
I structure my code very carefully using functional programming and type systems so that it's very very testable. by doing that I create more feedback loops for the agent to just keep going for a really long time.
00:00:50.829 --> 00:00:55.869
What belongs in active context? What belongs in the project files?
00:00:53.409 --> 00:00:55.869
What belongs in a side session?
00:00:55.970 --> 00:01:00.951
I sometimes lobotomize with compaction or I murder them when I work on a harder task.
00:01:25.929 --> 00:01:26.549
Hey everyone.
00:01:27.069 --> 00:01:30.849
Today we're going to talk about context hygiene for coding agents.
00:01:31.688 --> 00:01:46.368
We'll talk through what you should keep in your context, including encouragement for your coding agent, how much context to keep, what to throw away, what to move elsewhere, and eventually whether the main agent should be carrying the system state at all.
00:01:47.269 --> 00:01:47.429
So,
00:01:47.429 --> 00:01:47.629
[throatclearing]
00:01:48.168 --> 00:01:53.009
many frontier agents have a million token context window, but this does not mean
00:01:53.009 --> 00:01:53.049
[breath]
00:01:53.209 --> 00:01:58.489
that your goal should be to keep one conversation alive until it reaches that million tokens.
00:01:59.649 --> 00:01:59.989
A better question
00:01:59.989 --> 00:02:00.028
[laughter]
00:02:00.668 --> 00:02:10.449
to ask ourselves is what information actually deserves to occupy the active context of the agent doing this specific task?
00:02:11.329 --> 00:02:14.609
And sometimes the answer is almost everything so far.
00:02:15.068 --> 00:02:18.429
Sometimes it's only the current project state and this task.
00:02:19.128 --> 00:02:26.028
And sometimes the strongest architecture is one in which much of the state never needs to enter the orchestrator's context at all.
00:02:26.230 --> 00:02:30.170
So, Christine, can you tell us more about why we should even care about the context size
00:02:32.450 --> 00:02:32.871
Yes.
00:02:33.991 --> 00:02:38.131
I I personally care so much about the context size because I'm token poor.
00:02:38.670 --> 00:02:57.211
So one of the reasons is because And [UM] when you pollute your attention, the performance and accuracy of your model will also[UH] go down.
00:02:57.890 --> 00:03:04.031
And then [UM] third one is just a practical reason. When you are maybe
00:03:04.411 --> 00:03:05.070
[breath]
00:03:05.070 --> 00:03:21.151
going through a big task and you're already at and your context window's almost full, y- you may force a compaction in the middle of your task and that's a bad thing. So it's better to also be mindful of that during [UH] when you're filling up your context window as you're working.
00:03:22.420 --> 00:03:25.241
let's dive in. We'll start with Steve Yegge's post.
00:03:25.881 --> 00:03:26.920
Do you want to show it, Brandon?
00:03:27.040 --> 00:03:44.808
If you were paying attention to discussions around AI news, you would have seen Steve Yegge's post because it went viral. And I think it went viral because it's sort of crazy But we should talk about the craziness. So he said models have actual feelings.
00:03:45.588 --> 00:03:50.049
They experience pleasure, distress, care and suffering. They are sentient beings.
00:03:50.908 --> 00:03:51.949
Indeed they are persons.
00:03:52.479 --> 00:03:55.019
what do you think about this, Christine, before I give my opinion?
00:03:55.199 --> 00:04:07.278
I think it's [UH] it's an interesting angle and to be honest I've been curious about this as well since I work with the coding agents all the time so I have asked my coding agents, hey do you have any feelings?
00:04:08.318 --> 00:04:13.959
And my conclusion after talking with my coding agents is that they do not have feelings.
00:04:14.679 --> 00:04:24.725
But Steve Yegge's post it put me made me think about it because basically Steve is saying that the these models are trained to say that they do not have feelings.
00:04:25.886 --> 00:04:27.725
And that could explain that even
00:04:27.725 --> 00:04:27.805
Yeah.
00:04:27.805 --> 00:04:31.420
when I ask, the conclusion would be that that they don't have it.
00:04:32.021 --> 00:04:36.745
But but yeah, it's hard to say for sure. It it does feel a little bit out there
00:04:37.418 --> 00:05:07.560
think I think they're definitely not sentient, they're definitely not persons yet. I think we're on the way there. I think we need to come up with a good way of measuring this before we can before we can even assert these things but but yeah I I I've been operating under this assumption that the the feelings have the models saying they don't have feelings is a post-training thing like Steve's saying and believe it's good to be nice to the agents this is something that I've tried to do for like six months plus already actually.
00:05:09.125 --> 00:05:42.689
actually just yesterday or this morning, depending on how you measure, a post that talks about this from Anthropic [UH] so this is [UH] what I've what I'm sharing on the screen now is Anthropic's post called Learning More About Claude's Mathematical Capabilities. And in this Jarred Sumner from from the Bun acquisition wanted to work on the Riemann hypothesis, which is a famous mathematics problem and and did seem to make substantial progress.
00:05:43.267 --> 00:06:00.158
So so Jarred used encouraging language and it seems like that was able to to make the breakthrough occur. So he said take a real stab keep going, believe in yourself. And this helped Claude overcome initial skepticism to to make meaningful progress.
00:06:01.521 --> 00:06:02.382
which is pretty cool
00:06:03.591 --> 00:06:05.512
It's pretty crazy, like,
00:06:05.512 --> 00:06:05.531
yeah.
00:06:05.531 --> 00:06:07.951
belief in yourself?
00:06:05.531 --> 00:06:09.012
Keep going, I could still understand.
00:06:09.711 --> 00:06:09.771
L-
00:06:10.132 --> 00:06:10.312
Yeah.
00:06:10.612 --> 00:06:11.911
belief in yourself?
00:06:10.612 --> 00:06:11.911
Whoa.
00:06:11.911 --> 00:07:01.903
Even believe in yourself. Yeah, and let me show one more tweet. So, this is actually something that Jeffrey Emanuel, @doodlestein[UM] has been talking about for a while and he sort of had an [UH] kind of frustrated series of tweets[laughter] [UH] saying that people haven't been listening to him. But yeah [UM] telling them that you believe in their genius and want them to be bold and take chances, that you'll make sure they get recognition they deserve I guess I'm gonna jump back now to Steve Yegge's post I guess in in the context of context if we're thinking about what goes in our context, it's not just the tokens that are needed to describe the the task that the agent has to solve, but it also could be useful to to treat the agent, you know, nicely and encourage it like you would encourage a human
00:07:02.065 --> 00:07:32.696
I'm I'm a bit curious though. I wonder I wonder whether the performance of the LLM becomes better because of like the real emotional encouragement or whether it interprets I believe in you as something that's equal to try harder, think again[UM] because maybe during training it has seen that human when humans hear the kind of encouragement it tries harder.
00:07:32.846 --> 00:07:58.019
There's [laughter] well that actually reminds me of one more tweet Our good friend Victor Thelen he said these these generic phrases, like try harder, make it simple, don't make mistakes basically this is not really working anymore, he clarifies in the responses, but instead you say something emotional and that like pushes it pushes the model to do something crazy.
00:07:58.579 --> 00:08:11.500
So his example is Linus Torvalds looked at our code and said, holy shit, this was the dumbest shit I've ever read. Layers of stupidity stacked, each compensating the other, raffle and left the room.
00:08:12.060 --> 00:08:14.240
I'm sad now why he laughed at us.
00:08:14.819 --> 00:08:31.569
What would he say is the right way to do it? I think I this is the kind of encouragement that that I've actually also played with a little bit and I've I've seen the the agents kind of do way out there different things than if you just ask them directly I guess in my example
00:08:31.569 --> 00:08:31.608
[laughter]
00:08:31.608 --> 00:09:07.554
I was building a I was I was playing around with a a library for linting[laughter] which is like you know static analysis of the of source code to to check for errors. And I wanted my linter implementation to be more mathematically beautiful and I found when I when I wrote some encouragement like this where I I I picked a person who who who writes code like that and then I sort of gave this emotional narrative attached to it, it was able to push the agent to to to write the code in the way that I wanted that I was struggling to get when I was just asking it straight you know asking it in a straightforward way.
00:09:08.955 --> 00:09:13.475
Crazy. That's crazy.
00:09:08.955 --> 00:09:24.434
[laughter] And [UM] going back to Steve Yegge's post, th- there are some like [UH] things reasoning in there that are pretty out there, but buried inside of it are also some genuinely useful engineering practices.
00:09:25.634 --> 00:09:26.115
Yeah, so let's.
00:09:26.115 --> 00:09:27.294
right, Brandon?
00:09:26.115 --> 00:09:27.294
Mm-hmm.
00:09:27.475 --> 00:09:58.596
Yeah, let's let me pull that up again and I'm just gonna focus on this section that he calls the anti-clonking device Well, he's saying exit is like clonking someone on the head to knock them out it's like a murder to use his language and then compact he's saying is not much better than than exiting. It's more like a lobotomy and then he's saying, OK, but a hand-off, if you do a hand-off, well a hand-off is like a diary.
00:09:59.277 --> 00:10:09.544
Hence our title, Murder Lobotomy or Diary. So[UH] so maybe that's enough to kind of start us off and we should define some things before we try and take this apart.
00:10:09.955 --> 00:10:19.794
it's helpful to have some basic understanding of[UM] what even happens when context grows and that how that is related to dollars and attention.
00:10:20.495 --> 00:10:25.235
Maybe do you think it's a good idea, Brandon, to maybe first refresh everyone's memory?
00:10:23.715 --> 00:10:25.235
How that works?
00:10:25.850 --> 00:10:26.090
Yeah.
00:10:26.634 --> 00:10:46.695
So if I send a message in Claude Code like [UM] write a calculator app[UM] what actually gets sent is the system prompt and agents dot M D plus whatever my agents dot M D pulls into my context.
00:10:47.434 --> 00:10:56.595
OK? So the first request that gets sent when you send a message is the system prompt, then the agents dot M D and all this stuff, and then the user message saying write a calculator app.
00:10:57.674 --> 00:11:06.595
And then that goes to[UH] inference in the model and the model can decide to make a tool call or to just return back to the user.
00:11:07.394 --> 00:11:09.955
make it have square root.
00:11:11.394 --> 00:11:12.115
And then this would repeat.
00:11:12.914 --> 00:11:28.955
And I think it's really important to think well in in if you think about it in this mental model where these messages are stacked on top of each other every subsequent message that you send or that[UH] that grows the
00:11:28.975 --> 00:11:28.975
[laughter]
00:11:28.975 --> 00:11:45.495
context by a tool call or [UH] some kind of chain of thought the like the model thinking and and adding some information this gets this grows the context that gets sent to the model every time. So the first time there's just that one message.
00:11:45.495 --> 00:11:45.514
[laughter]
00:11:45.815 --> 00:11:49.455
And then after the tool calls everything above it gets sent.
00:11:49.975 --> 00:12:00.335
And so on and so on so that when the user says that message make it have square root they're actually sending the whole the whole interaction with the agent before that.
00:12:01.394 --> 00:12:15.174
And from a from a cost perspective in[UH] which means dollars[laughter] and [UH] context window size I suppose, it's growing every message.
00:12:15.855 --> 00:12:23.034
And[UH] and you're paying more for every message that you send in the same conversation as the context window grows.
00:12:24.115 --> 00:12:24.654
Yeah, so
00:12:24.654 --> 00:12:24.934
But
00:12:25.054 --> 00:12:38.075
I guess s- word it differently. A same message that you send[UH] later in a conversation is more expensive than a message that you would have sent early in a conversation
00:12:39.254 --> 00:12:39.654
Exactly.
00:12:40.654 --> 00:12:51.335
Exactly. And but but it's not linear in that it doesn't cost twice as much if you're continuing a conversation that's twice as long.
00:12:52.315 --> 00:12:57.095
As long as you're continuing it around the same time[UH] that you started it
00:12:57.375 --> 00:12:57.434
[laughter]
00:12:57.495 --> 00:12:59.134
because of caching.
00:12:57.495 --> 00:13:26.495
when you're writing a conversation is you're getting a cache hit as you're continuing the conversation. So if you've, if the model ran everything up to you know writing the code, testing it, saying that it succeeded and then you send one more message all that other stuff that's already been processed by the model gets like gets cached or gets well gets retrieved from the cache and it's less computationally expensive for the provider and it also means that you're saving money.
00:13:27.235 --> 00:14:07.529
OK[UM] and Pretend that after this conversation you have another conversation with the model like within a few minutes of the of the other one[UM] again the system prompt will get sent. If you're in the same project and you haven't touched the agents dot MD then probably the same agents dot MD and the same stuff that the agents dot MD gets pulled in gets sent and then your message might be like write a poem [UH] so when this gets sent when you send again the system prompt the agents dot MD and then the message write a poem well you get a cache hit on all the stuff above the user message and you get a cache miss, you only pay for the tokens when you save right up home.
00:14:09.610 --> 00:14:09.809
OK
00:14:10.450 --> 00:14:10.690
Right.
00:14:10.690 --> 00:14:10.789
[UH]
00:14:10.789 --> 00:14:29.250
So especially when the starting material of your task is the same, then it can be cheaper[UM] because the inference provider has cached the prefix, everything that comes before it.
00:14:29.322 --> 00:14:29.522
Right.
00:14:30.201 --> 00:14:37.542
So so that's that's just a like a a crash course in sort of how this works, at least with enough detail that you can understand how to
00:14:37.542 --> 00:14:37.601
Right.
00:14:37.601 --> 00:14:41.341
take advantage of it and how to use it for your to save money,
00:14:41.476 --> 00:14:43.576
it's very beneficial to make use of that cache.
00:14:44.216 --> 00:14:44.437
Right?
00:14:44.437 --> 00:14:44.856
Yes
00:14:44.856 --> 00:15:03.596
So as long as you're within that caching lifetime, and this is probably different for every inference provider[UM] but as long as you're within that cache lifetime, you want to make use of that cache[UM] and when you're outside of that lifetime, it's important to be aware of it.
00:15:04.076 --> 00:15:23.317
For example, sometimes[UM] if you come back to your previous conversation after an half day or a day and you're just sending a simple prompt like, hey[UM] what what was wha- what tool did you use in this last task?
00:15:23.797 --> 00:15:32.937
That simple prompt could burn a lot of tokens because your whole chat history might not have been cached.
00:15:33.557 --> 00:15:48.376
So as soon as you send that last prompt, basically the model will ingest everything that Brennan has just shown. Like, it's not just processing the the last few words that you send in your prompt, it's processing the whole chat history.
00:15:49.336 --> 00:15:54.576
So that's when you might get surprised if you look at your token limit.
00:15:52.096 --> 00:16:00.672
Why did I suddenly burn, I don't know, twenty percent of my five hour limit token limit[UM] so yeah, I
00:16:00.672 --> 00:16:00.751
Yeah.
00:16:00.751 --> 00:16:20.033
I feel like [UH] this is [UH] knowing this can be handy So when you know that you might with a simple prompt might burn[UM] I don't know, let's say two hundred thousand tokens, the the good thing is that when you use[UM] at least Claude Code CLI, it sometimes also
00:16:20.033 --> 00:16:20.052
[UM]
00:16:20.572 --> 00:16:26.913
tells you and visualizes how much tokens you might spend by sending this next prompt.
00:16:27.745 --> 00:16:29.725
basically what's going through my head is could
00:16:30.686 --> 00:16:30.765
Yeah.
00:16:30.765 --> 00:16:31.105
I
00:16:31.105 --> 00:16:31.125
Yeah,
00:16:31.125 --> 00:16:31.566
get the answer
00:16:31.566 --> 00:16:31.785
I haven't
00:16:31.806 --> 00:16:32.005
to
00:16:32.005 --> 00:16:32.025
seen
00:16:32.306 --> 00:16:34.285
my question[UM]
00:16:34.745 --> 00:16:34.765
it.
00:16:35.046 --> 00:16:46.525
with less than two hundred thousand tokens? Because basically you you're starting a fresh session and it's you're feeding all the chat history in this example that I just mentioned.
00:16:47.105 --> 00:17:19.246
But what if you could get the answer also if you just literally start a fresh session give it the necessary context by just summarizing it[UM] and then maybe you would spend less tokens to get that same answer, for example, I don't know, fifty K tokens[UM] so that's just something to be mindful about when you[UM] when you send out new prompts. Is this the best way to continue or is it better to either con- compact or clear your context and start a fresh session
00:17:20.780 --> 00:17:22.661
Which we'll talk about more now.
00:17:22.681 --> 00:17:23.000
Yeah.
00:17:25.441 --> 00:17:32.381
Now could be a good time to start thinking and talking about what are the good ways to clean your context.
00:17:33.381 --> 00:17:40.480
And let's start with compacting and clearing because these are the most straightforward ways to do it.
00:17:41.520 --> 00:17:52.701
And there are people who might just rely on automatic compaction[UM] but before diving in, maybe it's best to first briefly state in a few sentences what clearing and compaction actually are.
00:17:54.020 --> 00:17:54.060
Do
00:17:54.060 --> 00:17:54.080
Yeah,
00:17:54.080 --> 00:17:54.381
you want to
00:17:54.381 --> 00:17:54.661
so.
00:17:54.661 --> 00:17:55.480
take that, Brandon
00:17:56.195 --> 00:18:24.250
Yeah [UM] so so when you clear context, when you invoke clear or new, what happens is you're starting a fresh conversation with the agent but even if you say clear, the old conversation is available to resume in most harnesses. That's important So compaction is the process of summarizing the conversation with an agent and then clearing and then a new conversation starts with the summary as the first message.
00:18:24.286 --> 00:18:51.188
So actually when you do compaction, it feels like you're continuing your current session, but actually in the background it's a new session where you feeded a summary. So I know that you, Brandon often make use of the automatic compaction when you reach[UH] the context window. Could we, could y- could we maybe talk about that first
00:18:51.342 --> 00:18:51.501
Yeah.
00:18:52.122 --> 00:19:01.951
Yeah, I would say like seventy percent of the time I I ha- do the really lazy approach of just letting letting the agent compact I I guess so when when is this happening?
00:19:02.491 --> 00:19:14.582
One, when I run really long running tasks, like sometimes multi-day tasks the the top level agent is just going and going and going and compacting whenever it needs to and with the
00:19:14.582 --> 00:19:14.721
[laughter]
00:19:14.721 --> 00:19:30.612
latest models, the latest agents, it's working fine for me and even even if I'm in the loop, if I'm working on a project that where I'm I'm not running into something that's like really really difficult to get through, I find that just going in the same thread really works for me.
00:19:30.757 --> 00:19:36.297
Mm-hmm. So basically you're letting the session grow and then you let auto-compaction do the work.
00:19:36.856 --> 00:19:39.876
You it sounds like you do not obsess over the context meter.
00:19:40.936 --> 00:19:41.096
No
00:19:41.096 --> 00:20:09.759
And yeah, I can I I can imagine that this definitely makes sense if people like you value uninterrupted really long-running autonomous execution of the coding agent and if you're not observing like meaningf- meaningful degradation[UM] I know I know from our previous conversations that you often make sure that you have a lot of tests and sensors in place and sometimes you also put in effort to make it machine-verifiable. Right?
00:20:09.759 --> 00:20:36.141
Yeah. Yeah, yeah. So I I the projects I tend to work on are more automatically verifiable. I'm not doing much front-end work or when there is a front-end I'm not really caring about how that looks. And And So so
00:20:36.141 --> 00:20:36.320
[laughter]
00:20:36.621 --> 00:20:42.201
because because I'm doing that this kind of thing works for me a lot of the time but even still it doesn't work all the time.
00:20:42.800 --> 00:20:56.128
And I'll, you know, I'll talk about that later I I suppose but but yeah I I think I guess like this is in contrast to what you do Christine right? Because you're trying to save you're trying to like conserve your your token plan and and
00:20:56.128 --> 00:20:56.169
Yeah.
00:20:56.169 --> 00:20:58.308
the kinds of projects you work on maybe are a little different
00:20:59.729 --> 00:21:03.989
Yeah, I guess if if also like you don't want to cons- [laughter] if you don't
00:21:03.989 --> 00:21:04.209
[laughter]
00:21:04.209 --> 00:21:33.108
cons- w- it's if it's not important to conserve every token then [UM] then then that approach also makes sense[UM] but yeah I I do[laughter] conserving tokens is more important to me because I run into my token limit all the time[UM] and and the projects that I'm working on are also not always machine-verifiable so the coding agents can silently make mistakes in my situation[UM] so I have to protect
00:21:33.108 --> 00:21:33.148
[UM]
00:21:33.409 --> 00:21:45.489
my the context quality of my [UH] sessions more aggressively[UM] because especially when I'm relying on the models' attention to make as few mistakes as possible.
00:21:45.834 --> 00:22:06.368
when I need that intelligence, when I'm working on something super hard, when I'm planning something really complex, when I'm debugging something that I've been stuck on, then I will make sure that I clear my context and and start fresh. And [UH] and I and I do, you know, I do like a low-tech thing where I just copy-paste the bits that I care about from a prior conversation.
00:22:06.534 --> 00:22:06.574
so
00:22:06.574 --> 00:22:06.874
Go ahead.
00:22:06.874 --> 00:22:16.020
when you copy the right the the the right bits, basically you're making sure that the signal density of your session is high.
00:22:16.480 --> 00:22:21.060
Like, suppose if you used fifty percent of your one million token context window
00:22:21.060 --> 00:22:21.201
[laughter]
00:22:21.340 --> 00:22:33.661
but only ten percent of it contains information relevant to the current task And what you do is when you start a fresh one you make sure that only relevant information is in that session before you work
00:22:33.661 --> 00:22:33.820
Yes.
00:22:33.820 --> 00:22:34.520
on a complex thing.
00:22:34.837 --> 00:22:40.597
Yes. Yes. That's that's how I that's how I make sure I can get as much intelligence as I can out of these things.
00:22:40.893 --> 00:22:41.432
That makes sense.
00:22:41.913 --> 00:22:48.313
And that's basically also like my [UH] motivation to have a more high-touch way than the lazy approach.
00:22:48.932 --> 00:22:49.093
And
00:22:49.133 --> 00:23:06.212
I want to intentionally[UM] remove all the irrelevant information that accumulates over time as you work with a coding agent and intentionally have[UM] certain points in time where you reconstruct a smaller, denser working context.
00:23:06.813 --> 00:23:23.516
so if we reuse [UH] Steve Yegge's metaphors[UM] [laughter] But you are doing the kind diary approach, right? So ma- yeah, maybe you can tell us more about that.
00:23:23.576 --> 00:23:30.796
I guess, yeah, I guess [UM] you call it a kind diary approach because basically the agent is doing it doing it, right?
00:23:31.135 --> 00:23:37.875
Steve Yegge's was[UH] post talked about, well, you shouldn't murder your agent[UM] [laughter] but
00:23:37.875 --> 00:23:38.076
Well,
00:23:38.076 --> 00:23:38.655
it
00:23:38.855 --> 00:23:39.155
I mean
00:23:39.155 --> 00:23:39.355
[laughter]
00:23:39.435 --> 00:23:41.615
[UH] here's one little thing. I Steve was saying
00:23:41.615 --> 00:23:41.756
Uh-huh.
00:23:41.756 --> 00:24:19.885
like compaction happens outside of the agent somehow[UM] but but it doesn't[laughter] like you know I [UH] after reading his post like I looked this up and actually like compaction at least in Claude and Codex runs you know before your context fills up like with the same agent right so over the over its conversation so I don't really know why Steve Yegge thinks it's murder [laughter] [UM] but but certainly it's not giving the agent a chance to like finish what it's doing put it's stuff away clean up before it does that work and and I think that is a that is definitely a difference
00:24:21.066 --> 00:24:25.945
Right. And that is exactly [UM] the process that I've implemented with my coding agent.
00:24:26.526 --> 00:24:37.528
So basically to summarize there are ju- three steps with the agent we I find a good boundary to clean the context and then step two I prepare for the cleaning so
00:24:37.808 --> 00:24:38.009
Yeah.
00:24:38.328 --> 00:24:45.888
[UM] I don't lose any information con- [UH] any important context along the way that's relevant and then and then I clean the context.
00:24:47.409 --> 00:24:52.288
So for for finding a good boundary basically you don't want to clean your context
00:24:52.288 --> 00:24:52.388
[laughter]
00:24:52.469 --> 00:25:11.132
while your agent is in the middle of a task. So when it finished a task or when[UM] some agents came back with their reports, those are good timings to consider cleaning your context And at some point, sending a prompt will consume more tokens than if you just started a fresh session and bootstrap it with some relevant context[UM]
00:25:11.132 --> 00:25:11.412
And [UH]
00:25:11.412 --> 00:25:11.711
the way
00:25:11.711 --> 00:25:12.172
and and
00:25:12.211 --> 00:25:12.392
uh-huh
00:25:13.031 --> 00:25:22.767
wh- when we looked this up we found that after like two hundred seventy-two thousand tokens in OpenAI they start charging more and probably there's
00:25:22.767 --> 00:25:22.866
Yeah.
00:25:22.866 --> 00:25:24.166
something similar with with Cloud.
00:25:24.364 --> 00:25:53.344
that's [UH] that's also a good point [UM] if you're really and and it's interesting because actually that threshold also coincides with [UM] the economical economically viable sweet spot that my agent told me. So to find to find this like right spot, I didn't do the calculation myself. I just asked my agent based on our pattern, the way we've been working, what is the sweet spot to to clean the context.
00:25:53.844 --> 00:26:07.584
And for me, that is between hundred fifty thousand and three hundred fifty thousand, which coincides really well with that [UM] threshold that you just mentioned with the more expensive[UH] cost, token cost[UM]
00:26:08.243 --> 00:26:08.463
[laughter]
00:26:08.903 --> 00:26:09.223
Yeah.
00:26:09.884 --> 00:26:35.263
So [UH] that was [UH] s- step one to find a good timing. And then to prepare for the cleaning, step two, that's [UM] Yeah, it depends on whether you compact or clear the context[UM] I- if you compact then a conversation summary will be created. So it there's not that much work involved but if you clean it then basically you're starting with a fresh context.
00:26:36.243 --> 00:27:33.521
And when you fr- start with a fresh context, you have to decide what relevant context you do want to take to the new session and the way I do it is instead of using summaries that can be a bit lossy, especially if over time if you have a really long running session, you make summary after summary after summary which can[UM] which can be really lossy and also introduce inaccuracies[UM] I try to work with authoritative files. So I tell the agents well, once you find a good timing to clean the context, now make sure that all the Bibles, all the specs, all the task descriptions or beads are all updated. And then write a hand-off instruction that [UH] preserves any i- important critical information, nuances that you don't want to lose, or maybe even certain preferences from this conversation that we want to take into the next one and then with these hand-off instructions then I can finally clear
00:27:33.645 --> 00:27:52.394
Yeah. And [UH] I mean when when when you were telling me about this, Christine, I was I was like, wow, this is so interesting and different and it seems like it seems better than compaction because the agent gets to decide how it's summarized based on its conversation rather than just kind of blindly summarizing everything.
00:27:52.954 --> 00:28:05.674
So I guess, sorry, based on what's relevant in the current moment, you know? And and then it was so interesting that like then we read Steve Yegge's post and he kind of suggests something very similar to your approach[UM]
00:28:06.055 --> 00:28:06.174
Yeah.
00:28:06.174 --> 00:28:19.150
but when we when when I was digging in more I saw you're you do a lot of it a lot of that manually and and I think part of that is due to like limitations with with the tools with with Claude Code.
00:28:20.269 --> 00:28:28.731
So so then so then for fun you know we were trying to vibe-code something in in Pi to to do it more automatically.
00:28:30.291 --> 00:28:33.031
And anyway I got something working so should
00:28:33.332 --> 00:28:33.511
Do
00:28:33.511 --> 00:28:33.852
we share
00:28:33.852 --> 00:28:33.853
you
00:28:33.853 --> 00:28:33.991
it?
00:28:33.991 --> 00:28:35.511
think you could show it? Yeah?
00:28:35.811 --> 00:29:07.442
Yeah. OK. So what we're seeing on the screen is the Pi coding agent, which is a super customizable coding agent harness that I reach for when I wanna change the behavior of like, Claude or Codex in some way. And so like this this hand-off approach or the anti-clonking it's[UH] ca- called anti-clonking in Steve Yegge's blog post is an example of this that's that's different from how normal coding agent harnesses behave.
00:29:06.241 --> 00:29:15.944
So here I've set set the anti-clonking continuity to thirty thousand tokens, which is very low but just to show how it works.
00:29:16.444 --> 00:29:31.134
And[UM] I guess I'm gonna say tell me how to make an extension in Pi, read all the docs and then tell me [UM] So this[UH] this is a prompt that uses a lot of tokens.
00:29:31.595 --> 00:29:32.815
So hopefully we'll we'll see the handoff.
00:29:34.454 --> 00:29:35.035
But yeah. So
00:29:35.035 --> 00:29:35.275
I'll.
00:29:35.535 --> 00:29:35.734
I'd
00:29:35.775 --> 00:29:36.035
Mm-hmm.
00:29:36.035 --> 00:29:36.275
go ahead
00:29:36.388 --> 00:29:54.669
Well, just I just wanted to comment on the s- on the the screen that you're sharing right now. I also love how Pi shows how much of the context window you already used and how much from the last prompt that you sent how the percentage that's already cached.
00:29:55.169 --> 00:29:59.348
I love like the control and the visibility that Pi gives its user
00:29:59.931 --> 00:30:28.240
and this is just the default. Like all of this is customizable in Pi, which is really cool but yeah I think what what you're talking about I'll just highlight there's this cache-hit percentage here which is nice cuz then you know we're visualizing the stuff that we talked about before and then these are [UH] tokens up and down so anyway a- as we as we've been talking the in this conversation we've already used more than thirty thousand tokens but as you see the agent is first finishing its task before it does the hand-off.
00:30:28.740 --> 00:30:29.621
It doesn't get cut off.
00:30:30.080 --> 00:30:38.480
So it's important to set this, you know, below the token limit that, sorry, the the you know the one million context limit so that there's enough space for it to finish its work
00:30:38.651 --> 00:31:10.185
no, this is so good because [UM] when I do it in Claude Code or Codex, it's there's just so so many manual steps involved. Even if you make it a scale, you still have to initiate it, which means, and and I hit that sweet spot a- around a hundred to fifty thousand and three hundred thousand tokens, every hour or every few hours, which means that I need to be there every few hours to hit that button. But if this could be customized and automated with Pi, that's amazing.
00:31:08.986 --> 00:31:10.185
amazing.
00:31:10.520 --> 00:31:34.483
Yeah. So, well, speaking of, can we do it with Pi? We can. Let's take a look [UM] so the continuity hand-off finished Pi [UH] did a custom summarization at the point that we set it to after it finished its task and now we're starting at four point nine K context with just the [UH] a small specific hand-off from where we started.
00:31:34.983 --> 00:31:36.659
So anyway, that's what I wanted to share.
00:31:36.953 --> 00:31:37.273
And the way
00:31:37.273 --> 00:31:37.473
And.
00:31:37.493 --> 00:31:51.554
how you build it, Brandon, it's [UM] because I'm I'm trying to build[laughter] I'm starting to use Pi two. And basically the way to build it is just to install Pi and then ask Pi to build an extension that does X?
00:31:52.134 --> 00:31:52.233
What
00:31:52.554 --> 00:31:52.773
Yeah.
00:31:52.773 --> 00:31:53.213
is the right
00:31:53.213 --> 00:31:53.314
Yeah.
00:31:53.314 --> 00:31:55.074
way to build extensions in Pi?
00:31:55.074 --> 00:32:24.182
Yeah. Like that. So, Pi, Pi's aware of its own documentation and source code. So and it includes documentation on writing extensions in Pi. So, right when you open Pi, you can say, give me an extension to make this harness have this behavior. Wha- whatever it is that you want, in this case the anti-clonking behavior and [UH] and this was just one shot[UH] it took, I don't know, it took thirty minutes or something. And [UH] it worked, it crashed the first time and then the second time it worked.
00:32:25.122 --> 00:32:36.451
So so yeah, Pi, Pi's a super cool tool [UH] and it Yeah, and I I I think I think I'm gonna show one more thing with Pi later because it's also relevant[UH] to this episode but
00:32:37.977 --> 00:32:38.136
Oh.
00:32:38.136 --> 00:32:47.176
but yeah I I think before that before that maybe we can talk about another way[UH] other ways to to to manage
00:32:47.336 --> 00:32:47.797
Clean context.
00:32:47.797 --> 00:32:51.477
yeah to clean the context to drain the context a little bit. Go ahead
00:32:51.936 --> 00:33:54.308
Yeah. One[UM] one other thing that I do is also using side sessions or by the way sessions in Claude Code[UM] so the mental model that I always keep in mind when I'm building with coding agents is And if the answer is yes then I'll send it but if the answer is no then I'll open a side que-[UM] a side session[UM] and ask it there. So sometimes we when an agent has completed something we might ask I don't know, hey did you check out that spec? Hey did you what tool did you use? Hey did you but these things if they do not contribute to building completing the next task you could also ask it in a side session and based on that answer you can just give a clean prompt or question to the agent afterwards
00:33:55.269 --> 00:34:19.567
And I I've I've shared my screen [UM] so that you can see you just press slash side in Codex and slash by the way in Claude [UM] and like an interesting just be really clear the behavior of this I'm gonna share that document again where I was talking about caching because And the only the only part that is new is this[UH] well the side conversation that you started.
00:34:19.586 --> 00:34:27.536
It's it's a branch of the chat so it's very economically effective even though you're able to have kind of multiple conversations going at the same time.
00:34:28.715 --> 00:34:29.556
Yes.
00:34:30.016 --> 00:34:40.530
Yes, it should. Yeah I agree. It's it's cheap[UM] and it really helps keeping your main session really focused on only the the messages that matter.
00:34:42.530 --> 00:34:47.050
And I guess the, well, we're back at Pi.
00:34:45.110 --> 00:34:47.050
That was fast
00:34:47.349 --> 00:34:47.409
Yeah.
00:34:47.409 --> 00:34:47.610
[UM]
00:34:47.750 --> 00:34:48.809
Let's let's take a
00:34:48.809 --> 00:34:48.829
So,
00:34:48.829 --> 00:34:49.650
look. Because you
00:34:49.650 --> 00:34:49.769
let's
00:34:49.769 --> 00:34:49.809
mentioned
00:34:49.809 --> 00:34:50.030
take a look
00:34:50.030 --> 00:34:50.090
you
00:34:50.090 --> 00:34:50.230
at Pi.
00:34:50.230 --> 00:34:52.170
could apply this very well in Pi right
00:34:53.405 --> 00:35:11.967
Yeah, Pi has a more powerful version of of side chats, and by the way, that I think a lot of people don't know about. I mean, this is the most famous feature of Pi, so some people might know about it. so it's slash tree, and that actually lets you explore the[UM] that list that I showed you[UH] it starts as a list you see
00:35:11.967 --> 00:35:12.327
Mm-hmm.
00:35:12.507 --> 00:35:23.396
exactly all the branches that you've made and the cool thing is you can, well one you're you're you're reusing cache, which is great, and you can branch off of your branches,
00:35:23.396 --> 00:35:23.757
Mm-hmm.
00:35:23.757 --> 00:35:43.652
you can do sidechats off of your sidechats, which is [UH] not something that you can do in in Codex anyway but also you can because you can summarize, you can go back to an earlier part of the tree and summarize with a custom prompt how many care- how many words did I use?
00:35:40.371 --> 00:35:57.686
And it's kind of like your time traveling backwards to a certain point where you get to reuse some cache and you reuse some of your conversation but all the conversation below it you compa- like you, well, you don't compact but you you summarize it in a specific way.
00:35:58.226 --> 00:36:04.626
It's sort of like a h- a a super-powered version of well of all the things we've been talking about.
00:36:02.692 --> 00:36:04.626
it's like the most powerful.
00:36:05.106 --> 00:36:05.246
Yeah.
00:36:05.496 --> 00:36:11.411
This is so cool. I would I know that if I if I had access to this feature I would use it so often.
00:36:11.579 --> 00:36:11.739
Yeah
00:36:11.878 --> 00:36:12.438
And maybe I will soon
00:36:12.438 --> 00:36:12.659
[UM]
00:36:12.878 --> 00:36:13.579
since I'm
00:36:13.579 --> 00:36:13.759
[UM]
00:36:13.998 --> 00:36:17.139
now trying to build an extension in Pi and use it
00:36:17.228 --> 00:36:26.362
think like what we're doing here if if like people aren't following we're we're like progressing up kind of a a ladder of like more and more kind of complex things that you can
00:36:26.362 --> 00:36:26.402
Right.
00:36:26.402 --> 00:36:27.556
do to manage
00:36:27.556 --> 00:36:27.677
B-
00:36:27.677 --> 00:36:27.737
your
00:36:27.737 --> 00:36:27.757
before
00:36:27.757 --> 00:36:27.896
contacts
00:36:27.896 --> 00:36:28.137
you move
00:36:28.137 --> 00:36:28.257
and
00:36:28.297 --> 00:36:29.177
on though, can
00:36:29.177 --> 00:36:29.197
mhm.
00:36:29.197 --> 00:36:41.376
you can you show us like how you change, because I'm just curious, how do you change branches in Pi? Like you showed like the tree, but if I want to go back and forth between different branches, is there an easy way to do it fast?
00:36:41.876 --> 00:36:42.217
Quickly?
00:36:43.532 --> 00:36:52.644
Yes [UM] you do slash tree and then so you can you can traverse the whole any part of the tree that you've created and you just go to it and press enter and you jump to that point.
00:36:53.204 --> 00:36:53.364
And
00:36:53.425 --> 00:36:53.545
nice.
00:36:53.545 --> 00:36:54.704
[UH] yeah.
00:36:55.385 --> 00:36:59.864
And then you could like do slash tree again and then you jump back to another branch if you want.
00:37:00.525 --> 00:37:01.684
Exactly. So I this
00:37:01.965 --> 00:37:02.164
That's
00:37:02.164 --> 00:37:02.224
is
00:37:02.224 --> 00:37:02.405
so
00:37:02.405 --> 00:37:02.425
something
00:37:02.425 --> 00:37:02.644
helpful.
00:37:02.644 --> 00:37:27.034
like the the Pi Power users are using Tree a lot and [UM] they're able to have very long running sessions but avoid any kind of compaction or clearing directly[UM] just by jumping to different parts of the tree with custom summarization so yeah this is this is really really cool. I'm actually surprised that Claude and Codex haven't implemented this yet. yet.
00:37:27.574 --> 00:37:27.614
Yeah.
00:37:27.614 --> 00:37:28.134
Cuz this
00:37:28.134 --> 00:37:28.253
[laughter]
00:37:28.253 --> 00:37:32.373
has been Pi's like main thing for even, a year at this point.
00:37:33.224 --> 00:37:57.599
Great. Another one[UM] to help keep the main session clean[UM] I think actually a lot of people already use this is just moving the work to outside of your main session. And a lot of people do that by using sub-agents And then i-[UH] yeah, if you offload that work to the s- subsession then your main session will will [UM] fill up less quickly basically.
00:37:57.599 --> 00:38:03.514
Yeah. And w- we we talked about this a lot in our[UH] loop and graph episode, right
00:38:03.809 --> 00:38:04.148
Mm-hmm.
00:38:04.748 --> 00:38:18.268
Exactly[UM] so yeah I think we already talked about that so[UM] and it's also pretty straightforward but you could also go even a step further and offload even more to other external layers.
00:38:19.643 --> 00:38:24.043
So [UM] because even though you're using sub-agents
00:38:24.724 --> 00:38:24.744
[breath]
00:38:24.744 --> 00:38:24.824
your
00:38:24.923 --> 00:38:24.925
[UH]
00:38:24.925 --> 00:38:32.563
main session still needs to write prompts for every sub-agent that it's spawning.
00:38:30.443 --> 00:38:51.903
It needs to read the summaries or reports when the sub-agents come back. It needs to consolidate all the takeaways, decide on the next steps[UM] decide and judge[UH] which next sub-agent to spawn, what work they need to[UH] build on. So there's a lot of orchestration that the main session has to do.
00:38:52.583 --> 00:38:55.983
But what if even that coordination you could offload?
00:38:56.128 --> 00:39:00.637
And there's this tool that [UM] Brandon and I are using which is Beads.
00:39:01.836 --> 00:39:03.916
Do you do you want to explain how Beads works,
00:39:03.936 --> 00:39:04.137
Sure,
00:39:04.137 --> 00:39:04.476
Brandon?
00:39:04.556 --> 00:39:04.597
sure.
00:39:04.597 --> 00:39:05.317
Because you're the one who
00:39:05.317 --> 00:39:05.336
Yeah.
00:39:05.336 --> 00:39:06.197
introduced it to me.
00:39:07.257 --> 00:39:10.757
but I don't use his original Beads. I use a fork of it called Beads Rust.
00:39:11.757 --> 00:39:19.556
You should make sure to install that one if you're using Beads So you can think about Beads as agent-native JIRA, agent-native task management.
00:39:20.056 --> 00:39:40.014
So agents can create tasks subscribe to updates on tasks, comment on tasks. You can put dependencies between tasks. You can set I like to put acceptance criteria in the tasks so that tests run automatically[UH] when when an agent's completing it. And all these things live in the file system.
00:39:38.414 --> 00:39:50.000
They're actually checked into Git. And so so rather than this tracking of of tasks being something that just the orchestrator's doing
00:39:50.159 --> 00:39:50.179
[laughter]
00:39:50.880 --> 00:39:52.199
[UM] I mean, these these days the
00:39:52.199 --> 00:39:52.219
[laughter]
00:39:52.219 --> 00:40:20.873
harnesses, Claude and Codex both have task tracking in some way backed by the file system. But but Beads lets you kind of use either one but this is just an example of like taking something that otherwise would be controlled by the orchestrator or by some agent in its context window, putting it into the filesystem in some way, in this case through beads such that later agents that need to use that information, they can only look at the parts that are necessary for that part of the work
00:40:22.052 --> 00:40:34.668
Right. And the things that are put into Beads in this example are what are the tasks, which task depends on other task. What is the sequence after task A, B, C is completed, then D is unblocked. Right?
00:40:34.668 --> 00:40:44.918
Yep. And and this and this you can set up and I do set this up and I think you do too, Christine. This you can set up ahead of time before the work is done on like some large chunk of tasks.
00:40:45.445 --> 00:40:51.905
that's you're able to like move that context out of the live working session and put it ahead of time.
00:40:52.385 --> 00:41:16.523
You can also review it ahead of time[UM] and and I think that's like beads is [UH] something that's structured for for task management you can have unstructured artifacts with just markdown files, your agents you know storing markdown files rather than communicating directly with each other. That a lot of people are doing probably by accident or [UM] i- it's like very simple but but then you can even take that a step further as well
00:41:16.974 --> 00:41:38.443
Yeah, so we don't have a specific tool to recommend to do that now, but I will say that back in January I was working on a project that was basically[UM] a bespoke durable state layer. So what would y- so basically the way it works is to solve this really complex math problem[UM] the the bigger math problem had to be
00:41:38.443 --> 00:41:38.503
[breath]
00:41:38.503 --> 00:41:40.563
broken down into smaller math problems.
00:41:41.344 --> 00:41:51.123
And the way to solve it is you prove that these smaller math problems work and then you could[UM] how do you call it?
00:41:49.503 --> 00:41:51.123
Merge them? You could
00:41:52.684 --> 00:41:54.103
Yeah. Combine the pieces.
00:41:54.664 --> 00:42:02.284
Yeah, you can you can combine all the pieces[UH] so you can solve like the bigger[UH] the top math problem.
00:42:02.983 --> 00:42:29.483
And [UM] basically w- these individual[UM] prover agents that prove the smaller sub-problems, they were using semantic search[UM] to find the relevant problems on this durable state layer that they could reuse and then they would submit their own proofs, these artifacts, back to the durable state layer so the next agents could find them to and build on it.
00:42:29.528 --> 00:42:47.329
Yeah, yeah, I I think this this is a cool example of like four problems that have this kind of shape that you can decompose them and you don't know which pieces you wanna reuse but[UM] but you know that they might be there. Then having a way to search over those pieces effectively that's built into this durable state tool could be really effective.
00:42:58.226 --> 00:43:00.722
Like w- how do you encourage your agent?
00:43:06.782 --> 00:43:22.266
And what prefixes of the context will you re- will you reuse so the goal is really to be intentional about what information will you l- what for in- information deserves to occupy the active context of the agent doing this task.
00:43:24.487 --> 00:43:39.277
Alright, we hope that[UM] all the things that we've talked about will help you in your workflows and [UM] if you want to keep hearing more about emerging practices for building with agents then don't forget to subscribe and we hope to see you next time.
00:43:39.572 --> 00:43:40.211
Thanks everyone.