ABOUT THIS EPISODE
Your project still carries decisions made before the latest model arrived: tests for removed features, slow checks, outdated code, and instructions that keep accumulating.
A new model is a useful prompt to revisit those decisions. Give your agent a concrete problem to investigate, then review what its proposed changes would improve or put at risk.
In this episode, we cover four maintenance tasks through our own projects: pruning low-value code and tests, turning recurring mistakes into lint rules, reassessing implementations, and updating agent instructions with official model guidance. Examples include an editor cleanup that removed about 6,000 net lines across 131 files, a custom Effect lint rule, and a solver investigation that cut round-trip time from 20 seconds to 1.5 seconds.
Superlinear is a podcast about emerging practices for building with AI.
Hosted by Brandon Kase and Christine Yip.
Homepage: https://superlinear.fm
Twitter: https://x.com/superlinear_fm
LinkedIn: https://www.linkedin.com/company/superlinearfm
Brandon Kase: https://x.com/bkase_
Christine Yip: https://x.com/christinetyip
Music licensed through Soundstripe. Code: O2UT6CCEQNH4KT0D
SHOW NOTES 🔗
TRANSCRIPT 🔗
00:00:00.141 --> 00:00:05.761
What in your project could actually be improved if you built this with the smarter model from the start?
00:00:06.102 --> 00:00:11.083
Ultimately, removed six thousand net lines across a hundred thirty-one files.
00:00:11.401 --> 00:00:18.161
Sometimes pause feature development and ask, knowing what we know now, what would we build differently?
00:00:18.554 --> 00:00:26.158
With Astra, you want less skills and typically with new models, you want less things that are steering.
00:00:27.719 --> 00:00:34.159
Welcome to Superlinear, a podcast about emerging practices for building with AI agents.
00:00:34.179 --> 00:00:46.020
Each episode we cut through the noise to find one new angle worth trying. Then we unpack why it matters, where it works, and how you can apply it in your own workflows, so you can build more powerful things.
00:00:47.119 --> 00:00:51.259
We're your hosts, Christine Yip and Brandon Kase.
00:00:51.823 --> 00:00:56.982
Hey everyone! We're living in an exciting time where new frontier models drop so often.
00:00:57.563 --> 00:01:01.883
So last week, GPT-6 Astra dropped and it's a lot smarter than Sol.
00:01:02.542 --> 00:01:04.563
So now we have access to Astra.
00:01:05.043 --> 00:01:10.683
What in your project could actually be improved if you built this with the smarter model from the start?
00:01:11.862 --> 00:01:13.802
So today we'll give you four prompts.
00:01:13.962 --> 00:01:17.643
you can ask your coding agents to deslop different parts of your project.
00:01:18.763 --> 00:01:27.742
prune code and tests, turn useful standards into fast checks, reassess an implementation, and revisit your instructions and skills.
00:01:28.623 --> 00:01:34.063
These are specific things not everyone might know to ask their agent to do once in a while.
00:01:34.783 --> 00:01:39.382
At the end of the episode, you should leave with a few ideas you can put straight to work in your own project.
00:01:40.703 --> 00:01:44.742
So, Brandon, what was the first thing you did when Astra just dropped?
00:01:45.503 --> 00:02:00.222
Well, the first thing I did was go on Twitter and then I saw a tweet, which I'll share from, well, well, I guess I was on Twitter to know that Astra dropped.
00:02:01.403 --> 00:02:07.962
Yes, And and an old tweet of Ryan Lopopolo's was shared.
00:02:09.483 --> 00:02:14.323
So I will now share it with you, with you all. I love Ryan's tweets too.
00:02:14.603 --> 00:02:38.423
So I saw that so what I've shared my screen and there's a tweet from Ryan and I'm gonna skim through it cuz it's too long but essentially the first tweet says basically look for tests that are doing stupid things and he has a more specific prompt than that tests must justify their presence.
00:02:39.522 --> 00:02:41.383
Based on what you learn, update some files.
00:02:41.763 --> 00:02:42.842
That's the gist of his tweet.
00:02:43.143 --> 00:03:08.902
And then the immediate first reply is another tweet or another prompt which he said, essentially, look through all of the code and make impossible states unrepresentable, propose simplifications, and then linked a few interesting resources that helps with that process.
00:03:09.643 --> 00:03:11.543
So why did this catch my eye?
00:03:11.562 --> 00:03:13.703
Cool. This sounds like an easy thing to start with.
00:03:14.223 --> 00:03:24.043
Yeah, Well, first of all, probably if you're working, building something, you always have tests and probably the number of tests also grows quite a bit while you're building your projects.
00:03:24.043 --> 00:03:38.122
Yeah. And this is something I'm already I've already faced and I think we even talked about it on an earlier episode, but that over time my test suite grows of control and I How does it grow out of control for you?
00:03:38.483 --> 00:03:49.103
It grows out of control because Well, I'm not paying close enough attention to what the agents are doing well, I've told my agents make sure you're testing everything that you're doing.
00:03:49.122 --> 00:03:53.462
I have some guidelines around the kinds of tests I want and expect when I have them do some work.
00:03:54.643 --> 00:04:20.483
But I'd rather have my agents make too many tests and stupid tests so that I can go later and clean it up than the agents make not enough tests because then they start writing code that's incorrect. And so because I on the side of too many tests, well, you end up in a situation where the where where things are messy. And so so this tweet is interesting.
00:04:20.742 --> 00:04:23.062
And but that's only one reason that it was interesting.
00:04:23.822 --> 00:04:25.723
I'll give you two more reasons that it's interesting to me.
00:04:27.382 --> 00:04:29.403
One, Ryan tweeted this.
00:04:29.622 --> 00:04:50.403
And anything Ryan tweets I have to try because he's just awesome and I I he's on the cutting edge of agentic engineering and two, his second prompt about cleaning up the cleaning up the code base is algebra.
00:04:51.062 --> 00:05:04.702
Or it it's sort of the kinds of like functional programming and algebraic thinking that I'm obsessed with specifically like making impossible states unrepresentable and Lexi Lambda's post is is something that I'm that I find interesting.
00:05:04.822 --> 00:05:10.622
I recommend you read it if you haven't read it yet, viewer, listener we'll link to it in the comments.
00:05:11.062 --> 00:05:18.343
So OK, enough already let me share a Codex window.
00:05:19.362 --> 00:05:22.062
So, we've been talking about this the last few episodes.
00:05:22.242 --> 00:05:37.802
One of the things that we're working on for this podcast is a podcast editor. So I decided to run these prompts in the podcast editor. I just put both of them in here I I tweaked it a little bit.
00:05:38.122 --> 00:05:53.862
My with my arrogance I said also review Brandon's work on algebra for software engineering and then linked my website which I have blog posts about this because why not yeah and what? This was really great.
00:05:53.882 --> 00:06:04.122
I'm I'll spare you I'm not gonna go through the results too carefully, but I think there's some interesting things here like it cleaned up a lot of stuff.
00:06:04.942 --> 00:06:07.202
It made it made code safer, made code better.
00:06:07.583 --> 00:06:16.663
Ultimately, you can see here, removed six thousand net lines across a hundred thirty-one files. Great. Deleting code is awesome.
00:06:17.223 --> 00:06:18.442
Deleting code is really good.
00:06:18.742 --> 00:06:20.273
It's not too interesting to go into the details here.
00:06:20.322 --> 00:06:24.502
I think that the point is this improved the code base.
00:06:24.903 --> 00:06:28.442
I followed up with some prompts to like really understand the exact things that it did.
00:06:28.882 --> 00:06:30.103
And I was happy with it.
00:06:30.783 --> 00:06:40.963
But the gist is there's a lot of stupid tests that are no longer there and there was some code that had edge cases that are no longer possible.
00:06:41.762 --> 00:06:41.963
So.
00:06:42.903 --> 00:06:43.103
When?
00:06:44.862 --> 00:06:45.322
Cool.
00:06:46.043 --> 00:06:46.702
Interesting.
00:06:47.663 --> 00:06:51.523
Yeah, and I. How many how many tests did it cut away for you?
00:06:52.202 --> 00:06:52.702
I don't even know.
00:06:53.283 --> 00:06:59.223
I don't know how many tests there are and I don't care about it. But I what I care about is how fast the test suite takes.
00:06:59.903 --> 00:07:03.822
And well, it's a little bit faster now.
00:07:04.663 --> 00:07:23.382
I Yeah yeah my I have a hard budget of thirty seconds for my for the test suite that runs whenever my agents make changes to this code base. And now it's like, now it's like twenty, twenty-five seconds instead of thirty.
00:07:24.023 --> 00:07:28.983
Wait, but what if it takes thirty seconds and and more tests are needed?
00:07:29.882 --> 00:07:37.403
Like, do you do you force your agent to remove tests that are lower priority, to keep it at thirty seconds?
00:07:39.062 --> 00:07:44.963
and like, why would you even, time bound it? Like, why does it need to be thirty seconds?
00:07:45.963 --> 00:07:47.362
Let me answer that question with a picture.
00:07:48.182 --> 00:08:25.762
Well, look, it's a tweet by me about one of our old episodes highlighting a very interesting picture. So, as model inference speeds up, or as our another way to think about this is as our models become smarter and they can do more with fewer reasoning tokens, which is what we see with GPT-6, the total time between asking a question or commanding my agent to do something and it having done so successfully, becomes less and less dominated, percentage-wise, by waiting for the LLM to run, which I can't control, and becomes more and more dominated by tool call time.
00:08:27.142 --> 00:08:36.423
And one tool call that runs all the time, that can be very slow if you aren't careful, is your test suite.
00:08:37.923 --> 00:08:43.462
So by making your test suite faster, your agents complete their tasks faster.
00:08:43.883 --> 00:08:48.222
It's even worse with flaky tests because then maybe the test suite fails and the agent has to run it again.
00:08:48.763 --> 00:08:49.383
Does that make sense?
00:08:49.383 --> 00:08:56.783
Mm-hmm. Yeah. So why did you choose to limit the duration of your testing to thirty seconds?
00:08:56.802 --> 00:08:59.543
Like why not five minutes or five seconds?
00:08:59.822 --> 00:09:12.342
Like why thirty? And also another follow-up question, how does your agent choose the tests if if there are too many to fit into thirty seconds?
00:09:13.023 --> 00:10:36.802
I, for me the the time budget on the test suite is sort of like a gut feeling as an engineer as I'm as I'm building my my project and as I'm seeing the test suite grow and shrink as I'm as I'm pruning it and cleaning it up in in this case the test suite ballooned to a few minutes and when I worked with my agent to figure out why, in earlier iteration than this one I once I cut the crap I was able to just get it down to like, twenty seconds or something, and then so I said, OK, just like yell at me when it hits thirty, so but it's but it's really a feeling I do think it's super important for I think it's super important for teams to really care about how long the test suite takes even more than before and to be really ruthless with the performance of your of your tests and even offloading them to other machines if you need to So that's the first part the for the second part of your question you asked you asked how does the agent know what to do or not do Wait before we, yeah, but before Mm-hmm we jump to the second question, Like, I'm sure that instincts come from somewhere, right?
00:10:36.822 --> 00:10:41.283
And maybe sometimes in certain situations you would allow perhaps a full minute.
00:10:42.163 --> 00:10:44.643
No? Or is it like always thirty seconds?
00:10:45.062 --> 00:10:47.202
I'm just trying to understand some reasoning.
00:10:47.222 --> 00:10:52.442
No No, it's a function it's a function of Yeah, it it's a function of the complexity of your application.
00:10:53.182 --> 00:11:23.342
OK, so is something that I think engineers build up over time as like an instinct because they've worked with so many different projects. And what I would say to those people is I urge you to decrease your your limits in this agentic world, cuz it's more important for it to be fast. So, for a given project's complexity, if as an engineer you feel like two minutes is fine, push it to sixty seconds but unfortunately I don't really have good guidance on how to come up with that.
00:11:23.363 --> 00:11:41.702
I the way it's really a feeling but it's a conversation you can have with your agents. You can say, OK, really, like like what I do is I say, oh, tell me tell me what are the different kinds of tests that we have?
00:11:41.722 --> 00:11:42.982
Let's organize them into buckets.
00:11:43.003 --> 00:11:44.062
What are the slowest ones?
00:11:44.423 --> 00:11:47.543
What are the ones what are the ones that don't matter?
00:11:47.842 --> 00:11:52.562
Like this is a new this one by Ryan is a different way to ask that question of what are the ones that don't matter.
00:11:53.082 --> 00:12:01.082
But but group the tests together, describe to me the different kinds of tests we have, and usually when I do that I can say, hey, these this kind of test is dumb.
00:12:01.403 --> 00:12:16.423
Delete it. And create a rule in AGENTS.md so you never write a test like this again but yeah, like I said, it's a feeling So you mentioned like it's so important in an agentic world that things happen fast.
00:12:16.903 --> 00:12:30.283
Like why? Is it because if is it in a like a in a setting where if you're building a product your competitor is also developing the same product you want to be faster than them?
00:12:31.243 --> 00:12:37.663
So it means like maybe for people who are just building things for themselves, the time is less relevant?
00:12:39.743 --> 00:12:50.663
Sorry, I'm just gonna pull up this again This blue here, if you see the image this so again I've pulled up the Amdahl's law diagram.
00:12:51.302 --> 00:13:11.182
I guess I didn't I didn't describe it for the listeners last so there's a there's a bar that says inference, there's a bar that says tool calls, and there's a that bar that says you. Right? And what this is saying is how long of a task are you waiting on inference, how long are you waiting on tool calls, and how long are you waiting for your brain to like understand what's going on?
00:13:11.903 --> 00:13:17.763
In the olden days, there was a big, bar of a human coding it.
00:13:18.062 --> 00:13:43.023
And at the same time, you're understanding what you're doing because you understand faster than you write code in the computer and that was so big that you can afford for a slower test suite or taking longer to to do the tool calls because percentage-wise it was not a big chunk of the work now it's such a big chunk.
00:13:43.523 --> 00:13:46.342
Like because we have agents that can write code so fast.
00:13:46.623 --> 00:13:47.602
And what's the point?
00:13:48.123 --> 00:14:03.942
I think like if you're if you're someone who's potentially interested in this podcast you probably have stopped reading the code at least a hundred percent of the time that your agents produce because you're just so excited that you can write software faster.
00:14:04.623 --> 00:14:16.003
But the important thing to realize is, like, delivering software is not just getting code on the computer. It's it's a lot of things.
00:14:16.023 --> 00:14:17.863
And one of those other things is testing that code.
00:14:18.163 --> 00:15:00.062
And so if testing the code takes such a long amount of time, then you don't realize the speed up in the development life cycle as well if you're not caring about that that time, that part of the of the development flow, Yeah. I guess I'm just very worried that, necessary tests might get scrapped from the process. And I'm trying to understand, like, what is the benefit of being super strict about choosing choosing a time duration, like thirty seconds.
00:15:00.562 --> 00:15:03.602
Cuz I would I would definitely so I did the same exercise.
00:15:03.623 --> 00:15:05.763
I also tried to remove redundant tests.
00:15:06.363 --> 00:15:18.023
But I did it mostly because I don't want to waste tokens, because I don't think it's good for the model's attention when you have a lot of redundants context.
00:15:18.822 --> 00:15:39.923
But time-wise I feel like if I had to trade it off like OK in this project where I want to make sure that everything works exactly as I want, I want to make sure that it's accurate, I think I would want to keep the tests instead of like making the process faster.
00:15:40.743 --> 00:15:44.802
So I'm very curious, what what your thought behind it is.
00:15:44.923 --> 00:15:49.342
Yeah this is one of those things that I actually still read.
00:15:50.222 --> 00:16:06.263
I think that's the answer like like I think it's actually important to understand what the tests are doing once in a while. Not every not necessarily every every minute that the agent's working.
00:16:06.623 --> 00:16:12.842
But once in a while it's important to go in and say, OK, what are these slow tests actually doing?
00:16:13.163 --> 00:16:30.523
And it and once you understand that, you can make a judgment call and say, OK, this test, even though it takes sixty seconds to run, it's so important that I'm gonna keep it. But I'm gonna run it in parallel with all my other stuff because I am not gonna add sixty seconds to the rest.
00:16:30.543 --> 00:16:43.702
I want it to be sixty seconds at the same time as something else and and this is one of those places where being a human in the loop is really important because the models are not good enough to make these judgment calls.
00:16:45.263 --> 00:17:13.462
And it's hard I think I think it's a case-by-case basis. But it but it is it is important to work with your agent and when I say the agent can't make the decision, the agent can't or I've found in practice while you're developing and while the agent is adding new features and adding tests, it doesn't do a good job at discerning which ones are important and and making these decisions.
00:17:13.742 --> 00:17:23.462
But if you later are investigating, you can work with your agents and your agent will help you make the decisions of like, OK, this is what this slow test is checking for.
00:17:23.682 --> 00:17:25.363
If we removed it, here would be the impact.
00:17:25.623 --> 00:17:29.942
If we removed it but added these other fast tests, this is what we would cover and this is what we would lose.
00:17:30.702 --> 00:17:40.502
And then you as a human read that and say, OK, yeah I think we should do it, or it's so important that this is tested that I'm gonna keep this tested. Mm-hmm. Yeah.
00:17:40.803 --> 00:17:47.022
OK. I guess it does make sense that at some point you have to timebound it OK.
00:17:47.903 --> 00:17:49.603
All right. Yeah, I don't I don't know.
00:17:49.623 --> 00:17:58.182
Yeah, I'm just thinking out loud while you talk and I'm like, I don't know, what how I would time-bound it for my projects, but it's definitely good to think about it.
00:17:59.042 --> 00:18:04.202
Yeah, what how long is your is your test suite taking for your game, roughly? Oh, it takes so long.
00:18:04.583 --> 00:18:17.403
But it's also taking long because I'm building this app for it on the phone and I connected two old phones to my coding to my computer so my coding agent can control it.
00:18:17.903 --> 00:18:34.803
But there's some waiting involved there for the for the coding agent to really take some actions in the app and then wait for the reactions but I think it's still worth it because it's not a simulation on my computer.
00:18:34.823 --> 00:18:59.202
It's like testing the real performance on the real end device so I would want to keep it I would I would what I would what I would question or what I would give to you as a question is do you you think it's important to run those every time you make any change to the code base?
00:19:00.563 --> 00:19:01.383
Or only sometimes?
00:19:01.542 --> 00:19:02.903
And when is that sometimes?
00:19:03.143 --> 00:19:05.702
Yeah. And that's and that's yeah.
00:19:06.462 --> 00:19:33.923
I think that's important to think about because it's slowing you down. I can imagine that would add like minutes to the to the cycle. And you know, at some point you just get slowed down so much that you never deliver the your your project, Right. So I guess what this exercise would do is cut down the total number of tests.
00:19:34.823 --> 00:19:46.762
But maybe you could even like build a process for you that only when you build certain features that touch certain parts of your project, then a certain part of your test will be executed.
00:19:47.462 --> 00:19:57.163
I did I did that I did this exercise too with Astra and it reviewed the test that Fable and Sol, GPT-5.6, Sol built.
00:19:57.643 --> 00:20:07.383
And it did find like redundant tests. Like maybe there were like, OK, I have like four to four thousand tests and twenty-three of them.
00:20:07.403 --> 00:21:02.563
Maybe it's not too much, but it's still like things that are not needed. Like twenty-three tests were Well, I mean, No no, but the twenty-three tests were protecting the behavior of something that I actually of a feature that I already removed but it's not just twenty-three in total, like this is, there were also other things like there were Astra also found a bunch of tests that where the title looked very impressive but when you check what they actually check what they actually test, it's actually it doesn't mean much, it doesn't add a lot of value so yeah there were definitely even though Fable and Sol were so smart just because of the building process or just how the test suite grew over time, Astra still found quite some improvements that make sense to change.
00:21:03.282 --> 00:21:37.123
Yeah. And this is something we I think we talked about it in the four emerging practices episode or maybe the harness engineering episode but yeah like there's a like as your project gets bigger there's more sophisticated things you can do, like you said, like running parts of your test suite when certain parts of your code is touched and but I think the point here is your tests are important to pay attention to as a human.
00:21:38.242 --> 00:21:57.143
And I think like even if even though you're not reading the code for the tests, understanding what kinds of tests are running and understanding where the bottlenecks are and being willing to accept those bottlenecks in certain situations or not as a human is something that is useful to still do when you're working with your agents.
00:21:57.782 --> 00:22:04.563
And so what we what we didn't show, I guess are prompts that can sort of do that.
00:22:04.682 --> 00:22:23.383
But I think the point of this section of the episode is A, yeah, you can just copy paste the stuff from Ryan always because it's awesome, But the the kind of broader thing that you should keep in the back of your head is, oh, my test suite matters.
00:22:23.903 --> 00:22:25.583
What are the tests that are running when?
00:22:25.843 --> 00:22:33.083
How often? On which changes are these slower tests running or not? And how long is the test suite taking in total?
00:22:33.403 --> 00:22:39.323
And and these kinds of investigations are great to do when new models come out, like Astra.
00:22:40.143 --> 00:22:53.643
So we talked about about removing tests, but for some recurring tests or patterns, can we get the feedback from those tests earlier while the agents are still writing the code?
00:22:55.303 --> 00:22:55.502
Yeah.
00:22:56.202 --> 00:23:09.383
Yeah. I think this is this is where lint linting comes in or types but basically some this is shifting left. Let's just say it again.
00:23:09.603 --> 00:23:11.202
This is a nice theme of our podcast.
00:23:11.222 --> 00:23:12.323
We've mentioned it in a lot of episodes.
00:23:12.643 --> 00:23:24.442
So, take something that you find at run time after your code is written when executing tests and try and move it into the into the shape of your code that's types.
00:23:24.462 --> 00:23:32.583
That's like the second part of Ryan's prompt that we talked about, the shape make impossible states unrepresentable.
00:23:33.722 --> 00:23:41.343
And then or. And the and the idea of shifting left is you do it right the first time instead of first doing it wrong and then let the agent test it and improve it.
00:23:41.722 --> 00:23:42.663
So that's shifting left.
00:23:43.282 --> 00:23:50.083
Just for some listeners who have not listened to our episode on the four emerging practices.
00:23:50.522 --> 00:24:04.863
That's the shifting left that we talked about in four emerging practices. You can think of shifting left as taking something that happens later in the software development life cycle, for example running a test, and moving it earlier, for example at when the code is compiling.
00:24:05.423 --> 00:24:07.123
That's like putting something into a type.
00:24:07.682 --> 00:24:18.202
Or another place to put something is at lint time during static analysis where we're not running code but we're running a tool over the code.
00:24:18.623 --> 00:24:19.782
That's like what a linter does.
00:24:20.482 --> 00:24:34.482
All this is an elaborate scheme to transition into linting and what came up what people were talking about when GPT-6 dropped and the kinds of things you can do to deslopify your project with lint rules. So, linting is super important.
00:24:35.583 --> 00:24:40.083
If you have a TypeScript codebase, you're probably already using a linter because it's sort of standard there.
00:24:40.303 --> 00:24:52.843
I strongly recommend you use Oxlint, it's the best one but if you're using other codebases, sometimes linting is not something that is popular in the ecosystem that you're working in.
00:24:53.242 --> 00:24:59.383
But usually there is a linter and you should absolutely make sure that you have a lint step in your project.
00:25:00.423 --> 00:25:07.563
Everything you do because it's so it's so such a useful place to give feedback to your agents, to correct them along the way.
00:25:08.282 --> 00:25:19.742
And yeah, and it's a place where as you're cleaning up your code base, as you're deslopifying you can notice patterns and you can make them enforced rules that are really fast.
00:25:20.383 --> 00:25:23.583
So you can they're sort of like a free win.
00:25:23.863 --> 00:25:24.383
It's a win-win.
00:25:25.522 --> 00:25:29.462
some things that you can only catch with a test, you have to pay for that with your test suite getting bigger.
00:25:30.083 --> 00:25:37.403
But Lint rules tend to run really fast so it's it's a good thing to do.
00:25:37.702 --> 00:25:51.222
So. What would you, what would should people ask the agents exactly? Yeah, work with your agents if you're unfamiliar with the with the tech stack that you're using to find the best linter and to install like the common set of lint rules.
00:25:51.542 --> 00:26:01.222
If you're using a popular programming language, there's probably some well-accepted, like, anti-slop lint collection of packages.
00:26:01.663 --> 00:26:18.282
Like, for example, in TypeScript, the common the commonly accepted package is Dylan Mulroy's anti-slop, which is an Oxlint plug-in where he's noticed specific things that the agent does when it writes TypeScript and made lint rules to catch it.
00:26:18.863 --> 00:26:20.542
So this is just this is just one example.
00:26:21.002 --> 00:26:50.722
But you should not just rely on plugins or collections of lint rules that other people have created. You have an agent, you can just ask your agent to make lint rules for you for your project to catch specific patterns, to catch specific things that it messes up. You can even ask your agent to detect patterns to detect specific parts of code that feels shaky to find it, and then you can ask your agent, hey, can you turn this into a lint rule for me in your project of choice?
00:26:51.323 --> 00:27:13.373
And this is something that I did in one of our projects so I'm gonna show that really quickly just as an example I made a custom Oxlint rule for the effect library which I'm using in my project I'm not gonna talk about effect again but it enforces a stylistic choice about how to use effect.
00:27:13.423 --> 00:27:16.803
In this case, making sure that all Effect.fn are named.
00:27:17.442 --> 00:27:18.923
So that's just one silly example.
00:27:19.482 --> 00:27:31.222
The point here is as you're cleaning up, as you're deslopifying your agent, it's good to install something to prevent that same mistake from happening again.
00:27:31.942 --> 00:27:40.843
And you can hope that a mistake doesn't happen again if you put something in your AGENTS.md, but you can enforce that it doesn't happen if you add a lint rule.
00:27:41.502 --> 00:27:42.512
Yeah, that's a good tip.
00:27:42.563 --> 00:27:46.163
I've not tried this I've only been pruning my tests.
00:27:46.182 --> 00:27:50.722
I've not been adding some tests during my maintenance sweep but this is a good thing to try.
00:27:52.583 --> 00:28:15.643
And they're better the I guess it's part of your test phase but like it's good to think about lint rules as super cheap high value things to add when you can and they're better they're like strictly better than tests if they can catch the kinds of errors that you need to catch but they're less powerful they aren't able to catch as many things as typical unit tests can catch, for example.
00:28:16.282 --> 00:28:30.863
Yeah, so one of the things that I'm always very excited to try whenever a new model drops is to ask the coding agents to use the smarter model to reassess the existing implementation.
00:28:31.303 --> 00:28:52.182
It's more like a periodic, well you could also do it periodically but it's interesting, especially interesting to do when a smarter model drops but basically it's a architectural reassessment you step back to check whether the implementation still fits what your project has become, especially if it evolved.
00:28:52.762 --> 00:28:57.682
Because when a project evolves, the code can keep reflecting decisions that no longer fit.
00:28:58.522 --> 00:29:09.083
And I think it's helpful to, like, sometimes pause feature development and ask, like, knowing what we know now, what would we build differently?
00:29:09.403 --> 00:29:21.282
And because it's a hard problem it's you get kind of the most out of new models when you're when you're using the latest and greatest to ask these questions.
00:29:21.863 --> 00:29:27.782
So Christine, can you can you tell us what what you did when GPT-6 came out?
00:29:27.803 --> 00:29:31.462
How you applied this technique to your code base?
00:29:32.583 --> 00:29:43.103
Yeah, so I'm developing a game and during the development process the game evolved a lot. Like I pivoted like three times or something.
00:29:43.643 --> 00:30:10.962
So I was hoping since the game has evolved so much I was hoping to get Astra to help me find a new structure, how the game could potentially be built and maybe there was a smarter or more straightforward or cleaner way to make the game work where you get the exactly same output but maybe it's it could be a more effective way to do it.
00:30:11.883 --> 00:30:19.222
So I asked the agents I want you to reassess the existing implementation.
00:30:21.682 --> 00:30:30.782
Yeah. I just asked a broad a broad question like that and then what I got back were actually more like bug reports?
00:30:31.803 --> 00:30:37.403
Like for example, the game is taking place in a in a in a building.
00:30:37.423 --> 00:30:39.843
There's a roof and you can place objects in the building.
00:30:40.522 --> 00:30:50.563
And during the development of the game, the roof changed from a flat slab to a sloped roof, but the placement validator still treated it as a flat slab.
00:30:50.563 --> 00:30:58.103
So players could aim at the visible roof but still have their placement rejected because the things changed.
00:30:58.222 --> 00:31:03.623
This caused some recurring bugs in a in a core mechanic but this feels.
00:31:05.063 --> 00:31:10.042
Yeah, perhaps I didn't ask it right because I felt I was actually looking for more than bug reports.
00:31:10.182 --> 00:31:25.903
I was looking more for, OK, this is a quite a significant architectural change you could do. Yeah. and I think I think that's that's a consequence of not being specific enough with the question, but it look, you still got value out of it, right?
00:31:25.923 --> 00:31:26.962
Cuz it surfaced bug reports.
00:31:27.462 --> 00:31:42.163
But I think, Oh yes, it surfaced like quite some things that were really helpful to change. Like I didn't even, well I guess maybe some test should have caught it, but there were all these things that test did not catch and those are really good findings.
00:31:43.282 --> 00:31:43.482
Yeah.
00:31:44.442 --> 00:31:50.583
But I think to I'll I'll give an example of what I did in this in this category when GPT-6 came out.
00:31:51.143 --> 00:32:11.303
So in the in the podcast video editor that I'm building there's an a core sort of component called the solver engine which figures out how to solve a request for a specific edit to turn it into a actual edit.
00:32:12.103 --> 00:32:13.002
And what does that mean?
00:32:13.002 --> 00:32:20.682
Well, you could request to cut a transcript, a part of a transcript for example, in the podcast that you're editing.
00:32:21.442 --> 00:32:31.163
But by default maybe that cut sort of is in between a word or there's a jump that doesn't feel right.
00:32:31.462 --> 00:32:36.682
And so there's there's a part of the system that checks for consistency around those things.
00:32:37.442 --> 00:32:54.423
And what I was finding is this was taking like after a few days of of vibe vibing that process was taking like twenty seconds, which is unconscionable, sible, unconscionable, I don't know how to say it.
00:32:54.823 --> 00:33:08.563
Anyway, it's unacceptable so I so anyway I asked my agents to fix it and and I specifically I was like, hey, this thing is slow.
00:33:09.682 --> 00:33:20.643
Can you go and investigate and figure out why it's slow, propose ways to make it faster, and then I got back a big list and I said, OK, do one, two, three, six, and seven on this list.
00:33:21.083 --> 00:33:28.962
And then in the end I took something that was taking twenty seconds, round-trip, to one and a half seconds, which is still slow but acceptable.
00:33:29.442 --> 00:33:29.643
Yeah.
00:33:30.383 --> 00:33:40.942
Yeah. But I the I guess the point is I was specific of like, hey, this part feels wrong and that this was like a performance thing but you could also do that with the architecture as well.
00:33:40.962 --> 00:33:51.923
You could say like, hey, my when we add features to this part of the project, it takes too long, it's buggy, like investigate the code in that section.
00:33:52.123 --> 00:33:53.522
Are we architecting it poorly?
00:33:54.022 --> 00:33:54.863
Can we do better?
00:33:55.522 --> 00:33:59.982
These kinds of things that are, like, directed at specific parts of the code base will get you.
00:34:02.063 --> 00:34:06.282
Anyway, it will help you deslop those specific parts of the code base. That's the point of the point of the episode.
00:34:06.923 --> 00:34:16.842
When I work with coding agents, I always fear that at some point when the project has grown to a certain size or complexity that it will become code spaghetti? So I'm trying to get ahead of it.
00:34:17.262 --> 00:34:23.182
And that's that was the intention of asking a question like this. If you if you ask I fear.
00:34:24.483 --> 00:34:32.543
Yeah. But if you literally ask the agent, I fear that the we're getting to complexity level where the code is becoming spaghetti.
00:34:32.882 --> 00:34:34.922
Can you investigate Is the code spaghetti?
00:34:35.202 --> 00:34:36.023
Where is it spaghetti?
00:34:36.182 --> 00:34:37.023
Where could we make it better?
00:34:37.643 --> 00:34:40.163
Then you'll probably get you'll get something out of that.
00:34:40.583 --> 00:34:57.603
And if you've never asked that question before or that kind of question before and you've been vibe coding for a while or like agentic engineering for a while, you definitely definitely will will find improvements if you ask Fable five point one or GPT-6 Astra, that question.
00:34:58.143 --> 00:35:18.983
Yeah. So. Yeah. I guess I try I should I should literally ask whether well, I should literally ask say I should change my prompts basically instead of saying like hey reassess the implementation is there a better way to implement it to get exactly the same result I should I should ask try asking.
00:35:21.123 --> 00:35:25.143
Are we on the road for a code base to become code spaghetti?
00:35:25.822 --> 00:35:43.023
And we're. Yeah, or just like saying what you feel and rambling. I find when we talk and we say a lot of stuff, the models are really good at picking up the parts that matter, and all that other extra stuff that we say, it gives it steers the model in the right direction.
00:35:43.563 --> 00:35:49.822
So so yeah, I think like in general, I guess, why are we talking about this at all?
00:35:50.123 --> 00:35:54.603
The point is, one part of your project that you should deslop is the code base itself.
00:35:55.123 --> 00:36:01.983
And I'm sure this isn't a surprise, but it's something that you should think about doing when whenever a new model drops.
00:36:02.583 --> 00:36:02.782
So.
00:36:04.302 --> 00:36:04.722
You should do it.
00:36:05.742 --> 00:36:09.782
So Before you wrapped up, I still wanted to Oh, I had something.
00:36:09.782 --> 00:36:40.003
also give another possibility which is maybe because I built it with Fable mostly, maybe Fable was smart enough and there was no way there was no there is nothing to discover for Astra the way that Fable built it was already like most efficient and when I pivoted it already like removed all the slop so the only thing that Astra gave back were the bug reports.
00:36:41.742 --> 00:36:44.702
Yeah, something that I'll have to test more.
00:36:44.722 --> 00:36:45.003
No. I don't believe that.
00:36:45.003 --> 00:37:09.583
No way. No way. No absolutely not. Cuz I mean if you if you had fully, like, designed the thing and one-shotted a specific a specific version of the thing with the with the best fable, it may produce something that has no way to improve like or no way to materially improve.
00:37:09.963 --> 00:37:12.103
But that's not how you've been building, right?
00:37:12.282 --> 00:37:38.342
You've you started with something, you've been iterating, iterating the best the best programmers, human or superintelligent will not the thing that you get after iterating, iterating is never as good as the thing that you get from understanding the system as a whole knowing the current state and sort of thinking about it holistically.
00:37:39.762 --> 00:37:53.583
So so I, yeah. So this is sort of obvious, right one thing you can deslop is the code base itself.
00:37:54.043 --> 00:38:09.702
The reason the reason why we're talking about this, the reason why we're bringing it up is we're reminding you that, hey, periodically, and especially when new models come out, you should be thinking about what parts of my code base can I clean up, and then go and try and do it.
00:38:10.722 --> 00:38:25.063
OK, we talked about different ways to maintain the code but there's probably in a case of working with coding agents, you can also think about maintaining the coding agent's instructions, like how you maintain the code.
00:38:25.422 --> 00:38:39.063
Like as a project grows, also the instructions accumulate, right? Past decisions, lessons from bugs or rules, tools and work-arounds, hand-off notes how would you, how do you think about that?
00:38:40.043 --> 00:38:44.842
Yeah, I think so your AGENTS.md will grow unchecked.
00:38:45.463 --> 00:39:05.842
Your you might have skills left over that are unnecessary, but most importantly, when a new model comes out, the model itself might behave differently when it interprets the same thing that was in your AGENTS.md in the in a prior instantiation of the model.
00:39:06.362 --> 00:39:11.842
So let me share a tweet. Let me share a tweet.
00:39:12.882 --> 00:39:20.182
So when GPT-6 came out, I found this nice post it was a reply to something that Victor Taelin said.
00:39:20.483 --> 00:39:43.063
This is 0xCritters and zero X Critters was quote tweeting an article from a an OpenAI member of technical staff that explained how GPT-6 interprets skills and handles prompts differently than prior models.
00:39:43.802 --> 00:39:53.262
And so this particular person, 0xCritters, said, hey, you should send something to your agent.
00:39:53.282 --> 00:39:54.023
Here's what you should send.
00:39:54.483 --> 00:39:59.822
Let's update our Codex, config, skills, custom instructions, et cetera for the new Astra model.
00:40:01.143 --> 00:40:02.253
And then it said some other stuff.
00:40:02.253 --> 00:40:10.643
But one of the things it said is find any official OpenAI docs on Astra and use this post.
00:40:11.523 --> 00:40:16.103
And then apply it. Anyway, do that.
00:40:16.922 --> 00:40:18.842
What can I say? It's helpful.
00:40:19.043 --> 00:40:26.603
Like usually when new models come out Anthropic or OpenAI will write about how it behaves differently from before.
00:40:27.043 --> 00:40:33.682
And you can just literally paste a link to those articles and give it to your agent and say, hey, read this and update my project.
00:40:34.382 --> 00:40:58.483
What are the things that people should expect for Astra with Astra you want less skills, and typically with new models you want less things that are steering because every new iteration the models become smarter at achieving goals.
00:40:59.123 --> 00:41:11.382
And if you steer them to the goal in a particular way, then it will hold back the model from the way that it would naturally steer, which is more effective and more efficient.
00:41:11.663 --> 00:41:28.902
And so basically every time a new model release comes out, you should be removing. Removing skills, removing parts of skills, removing lines from AGENTS.md, or rewriting parts of AGENTS.md to be less prescriptive about how the model should achieve a particular goal.
00:41:29.262 --> 00:41:34.422
These are these are the kinds of things that typically you wanna adjust every time there's a new model release.
00:41:34.722 --> 00:41:49.123
And these are the things that if you if you just sort of paste the link to the article from Anthropic or the article from OpenAI into your agent and say do all the things in here to my project it will it will apply those changes.
00:41:49.402 --> 00:42:04.543
And if you have if you have evals, if you have a ways to know if specific changes to your skills or AGENTS.md or harness in any way is better or worse, then you're in an even better shape and maybe that's something we'll talk about in a future episode.
00:42:04.983 --> 00:42:35.742
So it sounds like I think one of the key points here is when you rebuild your harness, when a new model drops, it's really important that you also give that context to the agent that tells the agent what the strengths are of that new model or how the changes were and one example where those what that resource is, the launch post or report of the of the lab that dropped the model.
00:42:36.222 --> 00:42:54.123
Because I tried to I tried to just broadly ask the agents, hey just check check my skills, check my agent's markdown file, but I didn't find like really specific recommendations. But again, just like the last time it's it's more the way of asking.
00:42:54.643 --> 00:42:57.882
And I see the difference.
00:42:58.063 --> 00:43:05.583
I could see like why giving that additional information about the model could really make the request more specific.
00:43:06.163 --> 00:43:08.382
Yeah. Yeah. I think I think.
00:43:08.922 --> 00:43:12.032
Yeah. Just something for myself when you were talking.
00:43:12.032 --> 00:43:18.643
Yeah. Being specific is important the way that you're specific is unimportant.
00:43:19.902 --> 00:43:20.862
It's in these latest models.
00:43:20.882 --> 00:43:35.663
You don't have to use specific magic words like, you are a scientist, do XYZ you just have to ramble about doing XYZ specifically enough so the agent knows how to, how to do whatever it needs to do.
00:43:36.623 --> 00:44:00.003
Right. And when you say about like specific enough good things to mention in this example for in this example when you want your coding agent to think about what how can it improve agent's dot MD and skills when using the new model, then you have to give specifics about the strengths of the model.
00:44:00.143 --> 00:44:00.443
Yes.
00:44:01.322 --> 00:44:09.302
Yeah, we're at wherever possible use official resources to aim the model in the right direction, aim the agent in the right direction. Cool.
00:44:10.003 --> 00:44:17.682
Well, we talked about different ways to improve the code base, the projects and the harness when a new model drops.
00:44:18.063 --> 00:44:20.682
But when when should we do this?
00:44:20.702 --> 00:44:26.182
Only when a new model drops or I could see a case where, sometimes you can do this periodically, right?
00:44:26.842 --> 00:45:19.362
But even if periodically, how periodic? Do you have thoughts on that yeah, I think it's definitely something you should be doing all the time where all the time is on a cadence that makes sense for your project. It's something it's something you can try and automate, but I find that these things typically still need the human in the loop to sort of pay attention to how things are being fixed and updated, and so and so for me because these new model releases are so frequent I just use whenever a model release happens as this alarm of oh I should go in and improve the improve the code, improve the tests, improve the lints rules and improve the way that I'm orchestrating you my agents with AGENTS.md and all of these things.
00:45:20.023 --> 00:45:20.963
Yeah, that makes sense.
00:45:21.922 --> 00:45:41.443
Huh. I was I was thinking of you probably you when you're working on your project you probably have a sense like how much how many changes have occurred after your latest maintenance round. So that probably also helps you decide when to do another maintenance round.
00:45:42.242 --> 00:45:48.682
And probably, whenever you've pivoted could also be a good point to ask some of these questions because we.
00:45:48.702 --> 00:45:50.702
After a big project finishes.
00:45:50.702 --> 00:45:54.682
Exactly. Maybe before a release or after a release.
00:45:55.302 --> 00:46:02.583
So when a stronger model arrives is dropped, then you can use it to revisit decisions that your project has accumulated.
00:46:03.222 --> 00:46:16.463
So you can ask your agent to deslopify your tests, your code turn bad patterns in your code that it sees into recurring tests.
00:46:17.222 --> 00:46:26.603
And you can also ask it to reassess the implementation and your harness. So we hope that this gave you some ideas for things to try.
00:46:26.623 --> 00:46:27.922
Now Astra has dropped.
00:46:28.402 --> 00:46:30.402
And we're very curious to hear from you.
00:46:30.842 --> 00:46:37.822
What's the one thing in your project that you've been putting up with that you now ask Astra to investigate?
00:46:39.983 --> 00:46:46.523
Alright! Drop that in the comments don't forget to subscribe to catch our next episode and thank you for watching!