درباره این اپیزود
GPT-6 Astra and Fable 5.1 are changing how we start new software projects:
Less harness engineering, less process, less prescribing. More deliberate steering on the few decisions that compound through the whole project.
We walk through the Codex conversation of an actual project that we built: where we applied engineering judgment, what we left to the model, and how we planned verification so the agent could detect mistakes and keep building autonomously.
From the first ramble to the stack and architecture decisions, this is our Fall 2026 workflow for starting a project with coding agents.
Resources mentioned
Matt Pocock — Grill with Docs The grilling skill Brandon uses to have the coding agent interview him, sharpen decisions, and build shared context before implementation. Grill with Docs
Kit Langton Effect educator and practitioner whose work Brandon references for Effect patterns and tooling. https://kitlangton.com/ https://x.com/kitlangton
Effect The TypeScript framework Brandon uses for typed errors, dependencies, concurrency, and application logic. Effect documentation
Effect Solutions Practical patterns and reference material for building applications with Effect. https://www.effect.solutions/
Effect Institute Interactive resources for learning Effect. https://effect.institute/
Alchemy TypeScript infrastructure-as-code built around Effect, which Brandon uses to make infrastructure part of the codebase the agent can inspect and modify. Alchemy
Lexi Lambda — “Parse, don’t validate” The type-driven design article behind the “make impossible states unrepresentable” principle discussed in the episode. Parse, don’t validate
Rich Sutton — “The Bitter Lesson” The essay behind the “Bitter Lesson” idea we refer to when discussing why more capable models can make some bespoke scaffolding less useful. The Bitter Lesson / Rich Sutton’s publications
Superlinear is a podcast about emerging practices for building with AI.
Hosted by Brandon Kase and Christine Yip.
Homepage: https://superlinear.fm
Twitter: https://x.com/superlinear_fm
LinkedIn: https://www.linkedin.com/company/superlinearfm
Brandon Kase: https://x.com/bkase_
Christine Yip: https://x.com/christinetyip
Music licensed through Soundstripe. Code: O2UT6CCEQNH4KT0D
در این اپیزود
یادداشت ها را نشان دهید 🔗
رونوشت 🔗
00:00:00.100 --> 00:00:14.140
I built an entire working, multiplayer port of this N64 game that works in the browser on phones in one day. A few hours of planning and then I let it run overnight. When I woke up, it was all working.
00:00:14.480 --> 00:00:16.019
It was playable and fun.
00:00:16.280 --> 00:00:18.760
And importantly, there was a robust foundation.
00:00:18.780 --> 00:00:20.640
It wasn't disgusting sloppy code.
00:00:20.960 --> 00:00:22.280
This is kind of crazy.
00:00:22.660 --> 00:00:34.460
I feel like a few weeks ago we did several episodes about harness engineering and we even like went deeper how to do harness engineering for every substantial task in the whole agent graph.
00:00:34.640 --> 00:00:39.280
As I removed things from my workflow, the agents just got better.
00:00:39.659 --> 00:00:41.659
Yeah, there's definitely some tension here.
00:00:41.780 --> 00:00:49.780
You don't want to dictate too much when you work with frontier models, but I still see that you are steering the agents.
00:00:50.119 --> 00:00:57.140
Yes, but I'm steering during the rambling and interviewing phase rather than steering ahead of that.
00:00:58.552 --> 00:01:04.992
Welcome to Superlinear, a podcast about emerging practices for building with AI agents.
00:01:05.013 --> 00:01:16.852
Each episode we cut through the noise to find one new angle worth trying. Then we unpack why it matters, where it works, and how you can apply it in your own workflows, so you can build more powerful things.
00:01:17.953 --> 00:01:22.093
We're your hosts, Christine Yip and Brandon Kase.
00:01:22.677 --> 00:01:26.037
Hey everyone, welcome to a new episode of Superlinear.
00:01:26.837 --> 00:01:33.218
Starting a software project from scratch looks very different in fall twenty twenty-six than it did a month ago.
00:01:34.058 --> 00:01:40.537
Coding agents can now take on huge chunks of implementation autonomously, which raises a new question.
00:01:41.177 --> 00:01:45.677
What should you still spend your own attention on before you let them run.
00:01:46.397 --> 00:01:58.837
So by the end of this episode, you should have a better sense of where your own judgment still matters when starting a project with coding agents, and you'll have a concrete way to approach starting a new software project.
00:01:59.578 --> 00:02:13.397
How to get the idea out of your head, which decisions are worth discussing and steering yourself, how far to plan before you build, and what feedback loops need to exist before you can safely give the agent more autonomy.
00:02:14.397 --> 00:02:24.277
We'll walk through one real project Brandon recently built with GPT Astra, but the workflow itself is the same basic structure he uses across all projects.
00:02:25.638 --> 00:02:29.038
So Brandon, what is the project that you built recently with Astra?
00:02:29.358 --> 00:02:30.057
OK.
00:02:30.758 --> 00:02:41.557
I built an entire working multiplayer port online multiplayer port of this N64 game that works in the browser, on GPUs, on phones.
00:02:42.298 --> 00:02:51.798
And it worked in one day, like a few hours of of planning of this workflow that we're gonna show, and then I let it run overnight. When I woke up, it was all working.
00:02:52.457 --> 00:02:55.657
So all features, everything was there.
00:02:56.598 --> 00:02:59.557
Everything worked exactly as I described it.
00:02:59.957 --> 00:03:14.677
There were no bugs I like there were there were little tiny things but like it was playable and fun and importantly it was it was there was a robust foundation.
00:03:14.698 --> 00:03:19.737
It wasn't disgusting sloppy code which some people have seen with their one-shot prompts with Astra.
00:03:20.258 --> 00:03:33.957
And I will just show what I woke up to so it really looks like the game, like it's the game, but it's in my browser.
00:03:34.318 --> 00:03:53.157
And I thought that Codex had cheated and just put a video of the emulator into my browser so I was like chatting with it and looking at the code and everything And I I included this here in this picture I took Yes, the agent confirmed.
00:03:53.157 --> 00:04:05.557
OK. Yes, the graphics and audio came from the N64 ROM, but the browser draws those original graphics using WebGPU while the Rust WASM engine recreates the gameplay.
00:04:06.598 --> 00:04:15.207
So cool. So this a variant of this is what we're gonna talk about today and you'll see more cuz we're gonna go into the details.
00:04:15.237 --> 00:04:19.838
Cool. So you so you let the agent run autonomously while you were asleep?
00:04:20.398 --> 00:04:42.017
Yes, So yeah. So I guess I guess the foreshadowing is my workflow, my current workflow is I do a bunch of upfront work, which is less structured than it used to be even a month ago and then I let the agents go for, overnight.
00:04:42.598 --> 00:04:45.538
And then come and check on it and do a little more, etcetera.
00:04:45.858 --> 00:04:48.718
Cool. So what is the high level shape of the workflow?
00:04:48.958 --> 00:04:55.478
So basically it's a series of rambling to the agent, then the agent interviewing you back.
00:04:56.197 --> 00:04:59.617
And you do this for different parts of the project.
00:04:59.637 --> 00:05:23.377
So specifically, ramble about the idea, get interviewed back, ramble about the tech stack, get interviewed back, ramble about the architecture, get interviewed back, ramble about the important parts of the project and more interviewing. you get the pattern and importantly, rambling about the feedback loops and the verification system.
00:05:23.718 --> 00:05:31.358
And then let it go. So in your experience, does this workflow really generalize across all different types of projects?
00:05:32.338 --> 00:05:51.617
Obviously, we don't have so much experience working with GPT-six because it just came out. But I have actually built a few things already web apps, this game thing little scripts and and also the parts of the podcast video editor.
00:05:51.978 --> 00:06:33.317
I I'm very confident that this approach generalizes and it sets you up to it sets you up with a nice foundation where you can actually build some bigger project on top of it after which you might have to revisit and, adjust your harness at some point. But it puts you in a much better position than just sort of like vibe one-shotting a prompt, which I guess is kind of one step below this approach. I'm surprised that I'm comfortable doing this little when starting a project, given how obsessed I've been about harnesses.
00:06:34.697 --> 00:06:43.197
And I wanna convince you all that this is the right thing to do with with the models that we have now.
00:06:44.817 --> 00:06:45.517
Oh.
00:06:46.177 --> 00:07:02.617
I want to dive into that because yeah, I've also been obsessed with harnesses and perhaps some listeners know we have been doing quite some episodes on harnesses. So I'm curious like how the workflow that we're going to talk about today changes things.
00:07:03.918 --> 00:07:26.278
So it sounds like with the different phases where you ramble and get interviewed by the agent, there's a lot of surface areas where you can inject your own taste and preferences. So I'm excited to dive into the workflow in a more detailed way and also maybe get some of your opinionated parts about software engineering the way how you ramble.
00:07:26.898 --> 00:07:32.478
So yeah, maybe we could dive into each step in more detail?
00:07:33.458 --> 00:07:51.517
Yeah. Yeah I'd love to and I am excited to push my agenda, my opinions about building software on this episode so let me just share one motivating tweet from the from the from the timeline.
00:07:51.877 --> 00:08:30.918
Just so that this just so that your all you listeners and viewers are motivated to pay attention and listen and not just write me off for being a heretic so this is a tweet that I'm showing from Uncle Bob Martin who actually has horrible takes on software engineering in general, but for the engineering stuff he seems to be doing a pretty good job of being on point and there's some general consensus in the the Twitterverse that people are saying the same thing, I'm saying the same thing as well. So here's his tweet.
00:08:30.937 --> 00:08:32.057
I'll read it in full, it's very short.
00:08:33.077 --> 00:08:38.577
Since I stopped using my harness, my token consumption has fallen by a huge factor.
00:08:39.217 --> 00:08:40.998
That harness was massively inefficient.
00:08:42.217 --> 00:09:02.738
And this is this is also sort of a subtweet of like a video that he posted maybe a day or two before this where in the video he said basically the same thing, that he built out this crazy harness, he's been iterating and getting all the guardrails just right and then he took a step back and and said, well, how much better is my system than just using vanilla a vanilla model?
00:09:03.337 --> 00:09:05.097
And the vanilla model was better.
00:09:05.898 --> 00:09:12.998
So these this kind of thing has been happening and it led me to experiment with it more as well.
00:09:13.317 --> 00:09:21.837
And I found as I removed things when I'm from my workflow the agents just got better.
00:09:22.357 --> 00:09:41.077
So so of crazy. I feel like a few weeks ago we did a whole episode about we did several episodes about harness engineering and we even like went deeper how to do harness engineering for every substantial task in the whole agent graph.
00:09:41.697 --> 00:09:45.337
So yeah very curious how this changes things.
00:09:46.158 --> 00:09:49.538
And I I wanna be like really specific here.
00:09:49.557 --> 00:10:07.898
We're talking about starting a project I think there's still space and it's still like valuable once your project starts growing to invest in that harness engineering stuff where it makes sense and you can make those decisions as you observe your agents working in your project, making mistakes, doing things wrong.
00:10:08.158 --> 00:10:18.317
Then you start adding some, harness things in there. And you have to reevaluate every time there's a model release, of course. But there's still time and place for that.
00:10:18.557 --> 00:10:31.937
happens, in my opinion, like further along in the development of a project. But when you're starting, it's very important to get out of the way, to get out of the way of the like super intelligent models that we have now.
00:10:32.378 --> 00:10:33.337
Gotcha. That makes sense.
00:10:33.357 --> 00:10:35.217
So this is more for the planning phase?
00:10:35.837 --> 00:10:43.477
And once you've started building, you when you have a better sense how your agent is building, then you can optimize or build the harness. Mm-hmm.
00:10:44.217 --> 00:11:02.138
I would say the like initial the initial seed of the project, I like to say, because there's some planning but also some building as well and quite a you can get quite far in this initial seed so, OK, without further ado, let's just jump into the thing OK, so I'm gonna share.
00:11:02.898 --> 00:11:09.258
OK, this is the real conversation that I had with Codex to make this game. Alright?
00:11:09.557 --> 00:11:17.317
So, you get it unfiltered, raw, this is the actual ramblings and interviewing that I did for this game.
00:11:17.557 --> 00:11:31.937
Alright? So, so for for the listeners who can't see, I'm gonna be sort of summarizing the ramblings that I did in these prompts I'm not gonna say them verbatim because it's too boring.
00:11:32.398 --> 00:11:35.217
But it's a ramble. So you should get the gist.
00:11:35.778 --> 00:11:37.837
So I guess let's just take a step back.
00:11:37.857 --> 00:11:38.778
What is a ramble?
00:11:38.798 --> 00:11:39.498
What am I talking about?
00:11:40.217 --> 00:11:41.977
Doing voice to text.
00:11:42.518 --> 00:11:45.158
OK? So use your favorite transcription software.
00:11:45.457 --> 00:11:52.977
There's WhisperFlow, Monolog, Hanzy, whatever. I don't even know all of them these days.
00:11:52.998 --> 00:11:58.357
SuperWhisperer maybe or you can use the one that's built into ChatGPT and Claude.
00:11:58.758 --> 00:12:07.717
But the important thing is it is actually really valuable to talk and to ramble because there's more context for the agent.
00:12:08.298 --> 00:12:14.398
And the agent the agent's smart enough to weave its way through your filler words and your mistakes.
00:12:15.738 --> 00:12:26.778
And it's just better if it can get more information out of you. Yeah, I definitely see a lot more people using voice to text when working with coding agents.
00:12:26.798 --> 00:12:30.038
It's so important. So let's dive in.
00:12:30.177 --> 00:12:31.798
So here is here is my first prompt.
00:12:32.077 --> 00:12:39.097
So this is me telling the agent like OK I wanna we have this ROM which is like a binary dump of the N64 game I put it in the folder.
00:12:39.477 --> 00:12:43.158
I said there's also a partial decomp for this game I pasted the link to that.
00:12:43.418 --> 00:12:44.878
It's fifty percent decompiled.
00:12:45.618 --> 00:13:00.898
And then I told Codex I wanna build a multiplayer mobile version of this game with a bit-for-bit identical clock and state machine for the core versus mode behavior so that someone who is a competitive level player of the original version has the same feel.
00:13:01.217 --> 00:13:06.798
That's important for me cuz I'm a competitive level player and I said it should be cross-platform.
00:13:07.278 --> 00:13:09.398
I wanna be able to do match play between devices.
00:13:09.937 --> 00:13:16.638
And so I said I think I think we want a web app and we only need the match play part anyway, et etcetera.
00:13:16.817 --> 00:13:21.317
I just rambled and then and then importantly I called the Grill with Docs skill.
00:13:22.878 --> 00:13:26.388
And then I said use Grill with Docs so that we can get the details right.
00:13:26.418 --> 00:13:27.977
And that triggers the interview.
00:13:28.618 --> 00:13:40.717
Cool. So you do specify a few things like you want the core wait, scroll up but you want it to be bit by bit the same?
00:13:41.878 --> 00:13:42.577
Yes.
00:13:43.618 --> 00:13:51.057
There's this particular game was remade a few times and the physics of it didn't feel quite like the original.
00:13:51.457 --> 00:14:09.077
And so by saying I want it to be bit for bit identical or I want a bit for bit identical state machine, that will ensure that that the game will feel the way that I want, the way that it did in that in that game.
00:14:09.097 --> 00:14:15.538
And I'll be diving deeper into that part in later in the conversation. So you're still strictly talking about what you want.
00:14:15.758 --> 00:14:16.457
OK, cool.
00:14:17.378 --> 00:14:21.577
Yes, this is we're at the idea level, so this is like fleshing out the idea.
00:14:23.977 --> 00:14:25.418
It's the beginning of fleshing out the idea.
00:14:26.138 --> 00:14:29.477
And this so this grilling with Docs is a skill.
00:14:31.258 --> 00:14:39.898
So you need to install the skill if you don't have it you can get it from like just Google Matt Pocock grilling skills.
00:14:40.677 --> 00:14:47.278
You can install grill with Docs and its dependencies there's like a It's like grilling and domain modeling.
00:14:47.557 --> 00:14:48.298
A couple of Matt Pocock skills.
00:14:48.298 --> 00:14:48.998
Why do you need dependencies?
00:14:49.118 --> 00:14:51.847
Isn't it just like an interviewing skill?
00:14:51.847 --> 00:14:52.977
Is it that big?
00:14:52.998 --> 00:14:56.878
I'm just surprised. Grill with Docs just calls two different skills.
00:14:56.898 --> 00:14:59.138
One of them's called grilling and one of them's called domain modeling.
00:14:59.337 --> 00:15:00.298
So you just need those two.
00:15:00.998 --> 00:15:09.018
But yeah. Basically it just it interviews you and during the interview it comes up with specific vocabulary and stores that as well.
00:15:09.477 --> 00:15:18.457
So that you can have a shared vocabulary with your agent so because I called this Grill with Docs skill, I start getting questions back.
00:15:19.217 --> 00:15:25.278
And actually in this version of the grilling skill I'm I get one question at a time.
00:15:25.498 --> 00:15:28.057
I actually like that I like that better.
00:15:28.658 --> 00:15:37.097
Matt, has updated his skills so that you get multiple questions at once but there's a way to change that so you can just go back to the one question at a time version.
00:15:38.097 --> 00:15:46.398
I, find it to be kind of more, more focused to just tackle them one at a time.
00:15:46.898 --> 00:16:06.238
And the Yeah. It gives you more space to dive deeper into one topic or one Yeah, and if you have questions about it, like there's no ambiguity about what you're talking about, if you, like if you have questions about the question, if you and your answer to one question might affect the next one.
00:16:06.258 --> 00:16:07.977
So I like to do them one at a time.
00:16:08.957 --> 00:16:15.197
Right. So it's saying like this first feedback is like, OK, what is this bit-for-bit identical?
00:16:15.518 --> 00:16:23.937
How do we make it testable do you like, reproducing obscure quirks might delay the first version?
00:16:24.758 --> 00:16:25.837
and I'm like, yes, it's important.
00:16:26.758 --> 00:16:47.817
And the models love saying that things take weeks and like we're gonna delay our development process, but it's so fast, you can just ignore that crap. So then there was what you get back, I'm not gonna go through each question, but it's like a series of questions to refine the context for the agent so that it knows exactly what you want.
00:16:48.357 --> 00:16:51.197
So it's saying like, are you targeting the original 2D?
00:16:51.197 --> 00:17:03.817
Yeah what about the controls and then, they asked a question about networking should we prioritize responsiveness?
00:17:03.837 --> 00:17:06.877
This I gave I had an idea.
00:17:08.377 --> 00:17:13.657
I was like, oh how does there's this online service for Mealy, how does that work?
00:17:13.857 --> 00:17:14.758
That seems to work well.
00:17:15.438 --> 00:17:49.518
And so I find that the interviewing is a mix of saying yes to the suggestions but also sometimes getting clarification to make sure you understand what it's asking and sort of steering in the direction that makes sense for you So it looks like this phase is not just a way for the agent to extract requirements or ideas that are already in your head, but you're actually also using this conversation to arrive at decisions that you didn't necessarily had at the beginning, like this networking question.
00:17:50.038 --> 00:17:51.478
Is that the right mental model?
00:17:51.817 --> 00:17:56.397
So more like thinking with the agent than specifying for the agent.
00:17:57.518 --> 00:18:07.778
Exactly. The initial this grilling process, the interview process helps you flesh out your idea.
00:18:07.958 --> 00:18:22.218
It you work out the edge cases that you weren't thinking about By being presented with questions, you get a different perspective than you otherwise were thinking about. It just makes the agent more aware of exactly what you wanna do.
00:18:23.238 --> 00:18:24.377
Cool. I like this approach.
00:18:24.637 --> 00:18:29.617
I I do a version of this in product and design work too with the coding agent.
00:18:30.458 --> 00:18:35.137
I'm I'm thinking with the agent rather than specifying for the agent.
00:18:35.998 --> 00:18:46.478
I might come in with a strong belief about the user or the product direction but I still ask the agent for alternatives, trade-offs, and what I might be missing.
00:18:46.897 --> 00:18:54.097
And then often the final decision ends up combining something that I knew about the product with something that the model surfaced.
00:18:54.637 --> 00:19:03.557
So it the point isn't to outsource the decision, it's it was it's to make the decision better together with the coding agent.
00:19:04.137 --> 00:19:14.077
Yeah. It's about working with the agent to get to get the best the best thing out of it it's like there's this term like the centaur.
00:19:14.778 --> 00:19:21.617
you wanna be the centaur, half human, half agent, because together a combination of human and agent gets the best result.
00:19:22.317 --> 00:19:25.018
And this yeah.
00:19:25.718 --> 00:19:43.758
This ideation phase is basically the product and design requirement gathering phase. But I've what I've done in this modern version of my workflow is I don't I don't tell the agent, OK, we're only focusing on product requirements.
00:19:43.778 --> 00:19:54.057
I just kind of let it ask me whatever makes sense to flesh out the idea so that's sort of one new, like, change in my workflow.
00:19:54.077 --> 00:20:01.897
I I've relaxed the structure a little bit so OK, I'm gonna scroll down through here.
00:20:02.077 --> 00:20:10.597
Oop. And because there's something interesting that happens around.
00:20:11.417 --> 00:20:14.698
Well, you rambled and talked a lot with the agents. There's a lot of yes.
00:20:14.978 --> 00:20:19.157
Yep. And look, there's a lot of me saying yes I think this first one.
00:20:19.178 --> 00:20:20.538
Is it like half hour, an hour?
00:20:21.597 --> 00:20:37.337
Yeah. This first series for the idea took about a half hour OK. So so at the end at the end of the grilling, the the agent asks like, OK, here's what I think, does this capture our shared understanding?
00:20:38.117 --> 00:20:42.778
And then in this case, the agent said I wanna begin the first engineering milestone.
00:20:43.637 --> 00:20:44.857
And I was like, no.
00:20:46.218 --> 00:20:58.458
I said yes, that's a shared understanding, but no we will agree on a tech stack before any meaningful work is done, So so this is me transitioning to grilling about the tech stack.
00:20:59.298 --> 00:21:03.897
OK and I was like, let's grill.
00:21:05.178 --> 00:21:18.678
So now we've transitioned from product idea I spent about thirty minutes going back and forth to tech stack. So, with this game, let me like give some context.
00:21:18.678 --> 00:21:40.218
What Codex is saying here is well, OK, so what I said is like, let's talk about the core of the of the engine, of the game engine first and so remember we're building this like bit-for-bit identical state machine and it needs to match the behavior of this N64 game that's partially decompiled to C.
00:21:40.917 --> 00:21:57.678
And so the agent recommended let's use C see, compile to WebAssembly using Emscripten if you don't know what any of this means, it's fine, but I know these terms if I didn't, I would ask and try and, understand what's going on here.
00:21:58.478 --> 00:22:03.218
And actually, this was a suggestion that I disagreed with.
00:22:04.218 --> 00:22:10.278
OK? we have an oracle in this in this project.
00:22:10.798 --> 00:22:27.577
So if you remember from if you're a follower of our podcast, we've talked about oracles once or twice It was in the episode about four emerging practices if you have an oracle, you can have a really powerful feedback loop.
00:22:28.238 --> 00:22:39.567
And because agents are really good at writing code and reproducing the same behavior against an oracle, you kind of have freedom to to change code and write it in any language.
00:22:39.617 --> 00:22:55.117
And so anyway this was like an important thing that happened and because I knew there was an oracle, I was like, OK, we are we don't have to use C I think using C is unethical.
00:22:56.298 --> 00:23:05.198
OK, I should pull up my tweet about this cuz it's funny lemme pull that up OK, listeners, this is the opinionated part.
00:23:07.057 --> 00:23:11.317
Here's a here's a banger from two thousand eighteen.
00:23:12.778 --> 00:23:18.057
I tweeted, if I've talked to you in person before, you've probably heard me rant about this.
00:23:18.778 --> 00:23:28.718
It is shortsighted, unethical, and immoral to write C or C plus in twenty If you need high performance or low level systems code, use Rust.
00:23:29.178 --> 00:23:55.417
I think that's even more true today than it was in twenty eighteen and the reason is because when you write C you get lots of undefined behavior, it's really hard to write code that's correct. This undefined behavior leads to it just really easy to have bugs and the the type system doesn't help you write better code this is even more important for agents today.
00:23:55.538 --> 00:23:57.657
So So what you can avoid.
00:23:57.657 --> 00:24:01.577
But if that's so important for agents, why did the agents suggest using C here?
00:24:04.317 --> 00:24:13.617
This is why we don't vibe one-shot prompts. I suppose this is as you mentioned, this is my opinion.
00:24:14.577 --> 00:24:26.298
There is no and it is not a it's a relatively common opinion, but it's not the average opinion.
00:24:26.557 --> 00:24:35.657
The average opinion is C is fine, you can use it sometimes and that is a bad opinion in my opinion.
00:24:36.637 --> 00:24:48.958
So and it really we're gonna talk we're about to talk more about this in this section but like agents can write any code.
00:24:49.617 --> 00:25:13.018
So you should not be picking a programming language that you think is like easier to understand if it's not safer for the agent. You may you may wanna pick a programming language that's easier for the agent to understand and you may wanna pick you do wanna pick languages that give guardrails to help the agents write correct software.
00:25:14.538 --> 00:25:28.958
But you should not there's no there's no reason to pick C this is just like a pet peeve of mine I hate C and it's unethical it's unethical Anyway, so so whenever I see C recommended, my alarms go off.
00:25:28.978 --> 00:25:37.337
Whenever I see the agent was suggesting like, oh we can we should just reuse the decomp code.
00:25:37.357 --> 00:25:38.478
I'm like, no, it's an oracle.
00:25:38.478 --> 00:25:40.557
Like, we can we can write it in whatever we want.
00:25:41.038 --> 00:25:52.367
And I said, the question is, should we use Rust Rust compiling to WebAssembly so it works in the browser or TypeScript. And I wasn't sure at this point and Mm-hmm.
00:25:52.417 --> 00:26:07.978
By the way, so part of your judgment that that the agent doesn't need to use C is because you knew it was because you were planning to use an oracle so there will be checks anyways.
00:26:08.758 --> 00:26:09.798
Did the agent know that?
00:26:10.038 --> 00:26:12.157
Like, did you give the agent that input?
00:26:13.958 --> 00:26:22.798
Cuz I'm wondering, maybe if the agent knew that you were going to use a test oracle, then it would have, like, put less weight on using C.
00:26:24.077 --> 00:27:02.597
Yeah, is this is an interesting it's an interesting discussion to have, I guess, like So, yeah, if you if you put some of your taste and your judgment into like an agents dot MD or something reachable from agents dot MD before you have this interview, then agent's more likely to be steered in the right direction so that you don't have to catch these errors so yeah, if I had said if there's an oracle, make sure you're using it then that might have steered the agent to not even make the suggestion.
00:27:03.538 --> 00:27:38.678
But I wonder especially since if you've seen the last episode, we talked about how GPT-6 Astra improves in performance as you delete things from agents dot MD and you delete skills it's all these all these things bind the bind the agent in in in a way that I'm curious, does that hurt its performance in a way that's unacceptable?
00:27:39.498 --> 00:27:58.097
Like, is it better to just expose the particular parts of your taste that matter for this project when they're relevant versus ahead of time kind of enumerating all the things that you care about I don't actually know.
00:27:58.678 --> 00:28:00.438
But my current workflow is.
00:28:00.458 --> 00:28:01.817
Yeah, there's definitely some tension here.
00:28:02.397 --> 00:28:11.538
Like you don't want to dictate too much when you work with frontier models, but I still see that you are steering the agents. Mm-hmm.
00:28:12.538 --> 00:28:37.971
Yes, but I'm steering I'm steering during the rambling and interviewing phase rather than steering ahead ahead of that and I haven't tried both enough to know empirically like which is better in the current landscape, but this approach is working for me Oh, you mean with ahead you mean like in agents MD?
00:28:38.877 --> 00:28:42.137
Right. In agents MD before starting this process. Right.
00:28:42.377 --> 00:28:53.377
Because in agents MDs basically a pre-written rulebook and you don't want that to be like super clouded and pre-define things prematurely.
00:28:54.178 --> 00:28:57.538
That's why I hesitate to do that in this in this age.
00:28:57.637 --> 00:29:07.038
And so but the flip side is like you should have these things in your head so you can steer the agent in real time when they come up. So.
00:29:08.117 --> 00:29:17.057
So that's Yeah. So anyway and well, OK, so why did I say should we use Rust to WASM versus TypeScript?
00:29:18.377 --> 00:29:21.238
I guess and.
00:29:22.417 --> 00:29:43.617
Yeah, it is so I will take this moment to pause the pause the sharing and tell you my taste my general judgment. I think by default every project should be TypeScript.
00:29:45.018 --> 00:29:54.178
And only if you have a reason should it not be TypeScript so so what are reasons why it might not be TypeScript?
00:29:54.337 --> 00:30:06.117
If you're doing some machine learning thing, Python probably if you're doing something that is systems code or needs to be really fast rust.
00:30:06.917 --> 00:30:15.438
And and then anything else, if you're doing something weird, if you're doing like a mobile app, you'll have to do something different.
00:30:15.438 --> 00:30:24.498
But I think this is the this is like a bitter lesson style opinion.
00:30:25.337 --> 00:30:37.137
So if you if we look at all of the all of the training data that these models have been using, which is all the code on the internet and all the people talking about code on the internet, number one is JavaScript and TypeScript.
00:30:37.698 --> 00:30:42.607
Well, number one's JavaScript but JavaScript doesn't have enough feedback loops built in, so TypeScript.
00:30:42.637 --> 00:31:24.298
So we need TypeScript and then Python is also really high up but Python doesn't have good feedback loops either so we only use it when we need to Rust also is very high in that list and it has the right kind of feedback loops to help the agents write good code I've tried I've experimented using like fancier programming languages and I think the benefits that you get from those fancy programming languages on the feedback side that I'm a fan of, like better type system, more functional programming, is not worth the degradation in intelligence that you see because the agents don't have enough information in their training data.
00:31:24.998 --> 00:31:36.238
So that's why TypeScript default to TypeScript and then, sometimes Rust, sometimes Python, basically nothing else. Mm-hmm.
00:31:36.978 --> 00:32:26.018
So these are the things you default to based on your own experience playing with coding agents and a mix of yeah my experience playing with coding agents and kind of like general philosophies that I have, which I guess I'll just quickly share cuz this is about hearing about my a particular rationale for things so like I said adding more feedback loops for the agents is really powerful. So there's type systems are a great way to give the agent a a machine-checkable tool to ensure that code way over here and code over here can connect to each other properly.
00:32:26.597 --> 00:33:17.038
So when you use a language that's typed, you can, some code over here defines its shape with a with a type, and then the code over here that's using it if the agent doesn't have this code in its context and it writes something that's wrong over here, the compiler, when it does its type checking, will surface that feedback to the agent and then the agent can steer and fix it and this is just like really really useful. And so and so I what I believe is you wanna turn that dial up to a thousand and whenever you can use more complicated types that typically might have been you might have been accused of overengineering, even though I tended to be on that scale anyway before agents in this day and age, like, overengineering is way more over because it's just so cheap to write code.
00:33:17.417 --> 00:33:22.978
And it's so cheap to write sophisticated code because the agents are really sophisticated.
00:33:23.397 --> 00:33:25.238
They're good at math and these things.
00:33:25.238 --> 00:33:48.317
So complicated advanced types, complicated advanced functional programming is really really useful as tools for these agents So. AI really changed the economics here, like previously it may have been expensive because humans had to learn and write all these boilerplate themselves.
00:33:49.238 --> 00:33:57.998
But now the agent can absorb all of that implementation costs while you can still get the benefit of these stronger machine-checkable structure.
00:33:59.238 --> 00:34:09.858
Exactly. Cool. So when you so when you discuss with the agents about the tech stack you also inject some of your preferences in it.
00:34:10.617 --> 00:34:22.177
I think the hard part is the hard part for me is that when you when the agent proposed C, you instantly had reasons to disagree.
00:34:23.038 --> 00:34:31.478
And I could imagine that someone less experienced might just think, oh, well whatever the agent's proposed sounds reasonable.
00:34:32.157 --> 00:34:44.177
So I'm wondering, like, how should someone who's still developing that engineering taste use this workflow without simply accepting whatever the model the agent is recommending?
00:34:44.938 --> 00:34:49.018
I think it's good to hold strong opinions loosely.
00:34:49.257 --> 00:34:56.878
And as you as you get more information, as you become more experienced, you should reshape your opinions.
00:34:57.898 --> 00:35:14.538
So like, the reason that I have this opinion is because I have so much experience working with C, trying C in different places, hearing stories about other people using it, and trying Rust and using Rust in different places.
00:35:15.338 --> 00:35:19.978
And so I'm just very confident in in this particular decision.
00:35:19.978 --> 00:35:24.818
And the way to build that confidence is to, yeah, to just try a lot of things.
00:35:25.057 --> 00:35:47.277
I think the best way the west, the best way to build this this discernment is to build a lot of different projects of different shapes as much as you can and think critically, like, oh, when I when I made these decisions, it fell apart in this way.
00:35:47.777 --> 00:35:52.177
When I made these decisions, this fell apart over here I got stuck here.
00:35:52.978 --> 00:35:56.077
When I work with, if I'm talking to people, I'm learning this.
00:35:57.257 --> 00:36:13.338
So it's kind of I don't have a great answer, but I think for beginners, first of all, I'm so sorry that you're a beginner when the AI is doing stuff for you for all of those people who are beginners cuz this is like a bad time to be to be learning stuff.
00:36:13.338 --> 00:36:19.137
But in some ways. In other ways, it's the best time because you have a tutor that you can ask questions to.
00:36:19.137 --> 00:36:44.038
So don't be lazy, ask questions, understand what's going on as much as you can you do a good job of that when you're working with the agents, Christine and then and then if you have if you have a big decision, if you're starting a project and you really don't wanna mess up, experienced engineers typically love this part of the process.
00:36:45.077 --> 00:36:47.097
Like, I love talking about this.
00:36:47.117 --> 00:36:50.878
Like, choosing the tech stack, making these initial architecture decisions is really fun for me.
00:36:51.297 --> 00:37:01.557
So hopefully in your network someone who is more experienced in this way and you can just ask them and hear their opinions.
00:37:02.389 --> 00:37:19.637
Yeah, or you could listen to podcasts like these right, Yeah, so this particular episode you're learning my opinions and I'm trying to share why I have them. So that's like another, that's a data point that you can use when you're building your next projects.
00:37:20.358 --> 00:37:29.378
And even if you, even if you don't follow all of my reasoning of the specific choices, and maybe we don't have time to even go into all of them for all the all the choices I'm gonna talk about here.
00:37:29.918 --> 00:37:37.217
Just take them and try it yourself and see what works and what doesn't work and what you agree with and what you don't Mm-hmm.
00:37:37.677 --> 00:37:46.018
Yeah. It definitely sounds that even with the smarter models, engineering judgment is still very, valuable.
00:37:47.757 --> 00:37:48.458
I think so.
00:37:49.217 --> 00:38:18.197
I I think it's important for it's important for setting the foundations right for your project so that it can it can scale up So but we have a super intelligent AI in our pockets that we can ask questions to when we don't understand something and yet it's helpful to be able to steer it so just copy what I did.
00:38:18.838 --> 00:38:26.157
Yeah and also a quick note to listeners if you have any like specific questions we would love to hear about them.
00:38:26.177 --> 00:38:42.358
We've been like talking a lot on the episode but I would love the conversation to be a bit more bidirectional too so if anyone has any specific questions feel free to leave them in the comments and maybe we can see if we can like talk about those topics too in a in the next episodes.
00:38:44.757 --> 00:38:48.257
OK. So OK. Sorry, that was a bit of a distraction.
00:38:48.498 --> 00:38:50.097
I was just wondering, like, OK.
00:38:50.117 --> 00:39:00.498
This looks the the process makes sense, but I was also when I looked at your the way how you were talking with the agent, would I also stop the agent right there?
00:39:00.697 --> 00:39:01.577
And I don't know about that.
00:39:02.518 --> 00:39:07.797
Yeah. No, it was a really it was a good question OK. So I'm going back to the chat now.
00:39:09.137 --> 00:39:15.958
So right. So now you should after this discussion you should understand why I said Rust versus TypeScript.
00:39:16.557 --> 00:39:41.378
I got some reasoning back and I decided, OK, Rust makes sense and then for the core and then I said then I was rambling and I was like OK I think TypeScript should go around the core on the server and in the web app and then and then I gave my TypeScript stack to the agent and what is my TypeScript stack?
00:39:42.217 --> 00:39:56.557
It's Effect TS V V4, Oxlint, and Bun on TypeScript. And it's set up in a particular way and alchemy. OK. So I'm gonna just quickly talk about these components.
00:39:57.498 --> 00:40:02.617
You can Google them to read more about them or talk to your agents Effect.
00:40:03.557 --> 00:40:04.257
Well, sorry.
00:40:04.958 --> 00:40:38.438
Before I go into the depths here I'm prescribing just enough and I'm prescribing particular frameworks but not too many frameworks, not too many libraries, not too many rules, enough that I can say it in a sentence but but I'm prescribing enough to increase the like the bitter lesson surface and the feedback loops for the agents. OK?
00:40:38.818 --> 00:40:49.117
So let me let me go through each of these carefully. So yeah we're talking about how we do not want to prescribe too much, right?
00:40:49.637 --> 00:40:51.818
Yeah, so yeah, that's a good point.
00:40:51.838 --> 00:41:11.978
When I say bitter lesson, I'm referring to like a famous essay by Rich Sutton where he says all these fancy machine learning algorithms are not gonna be as powerful, They're not gonna produce an AI that's as powerful as a really simple algorithm that you can just run on more compute.
00:41:13.117 --> 00:41:24.157
And this turned out to be true, right we're building these massive data centers and we're just running relatively simple algorithms on huge supercomputers and we're getting smarter and smarter models.
00:41:24.677 --> 00:41:25.378
So.
00:41:26.498 --> 00:41:46.938
And this like bitter lesson the kind of the theme is do the simplest thing that you can and all these fancy things that you try and add to make the system smarter, they actually end up getting in the way as the AI gets more intelligent.
00:41:47.657 --> 00:41:49.478
So that's the interpretation here.
00:41:49.478 --> 00:42:18.538
So so you want just enough to steer the agent in a direction that has your taste, which mine is more types, more functional programming, more feedback loops but not too much that you're relying on lots of buggy code and you're and not too much that you're sort of fixing the path of the agents so it can't use its intelligence to find the best route to the answer.
00:42:20.237 --> 00:42:20.938
Alright?
00:42:21.398 --> 00:42:22.097
Mm-hmm.
00:42:22.677 --> 00:42:23.717
OK. So.
00:42:24.757 --> 00:42:34.438
So I'm gonna quickly go through the the TypeScript stack I'm doing this because I want you all who are listening or watching to try this if you haven't yet.
00:42:34.858 --> 00:42:41.237
Because just default to TypeScript and try this stack and figure out what works or doesn't work for you.
00:42:41.637 --> 00:42:42.577
Alright. So.
00:42:43.518 --> 00:42:44.217
Effect.
00:42:44.797 --> 00:42:49.617
What is effect? It's what I consider like a missing standard library for TypeScript.
00:42:49.918 --> 00:43:38.277
It takes TypeScript and it makes it more type-safe, more functional programming it's very robust, it's mature, lots of people are using it it's fairly, it's starting to become really widely accepted as much as as much as React, which is also part of my stack so and what effect is like at its core is a data type that wraps concurrency, dependencies and errors into one type and it lets you express that the shape of those things in different parts of your code base so that your agents can not have concurrency bugs and not have bugs with dependencies and write tests better and lots of other nice things So, effect is really great.
00:43:38.998 --> 00:43:42.518
Oxlint is a really, fast linter.
00:43:43.217 --> 00:44:15.518
Linting is a static analysis step we've talked about on a few episodes including the last one but it it's a step that can run really fast over your code without actually building and running your code to give feedback to your agents and you can put in rules about rules that express your taste and style and prevent slop and all these things and then bun is a run time for TypeScript that is has batteries included. Again, a lot of people use it.
00:44:15.538 --> 00:44:19.898
It's very mature and when you have bun, you don't need lots of other libraries.
00:44:20.557 --> 00:44:39.757
So anyway, that's those and then alchemy, Alchemy's also really good Alchemy is a is a framework for expressing your infrastructure, the server, like your cloud in TypeScript with effect.
00:44:40.237 --> 00:45:17.177
So with using TypeScript in effect, you can your agents can describe all the different pieces of your like AWS or Cloudflare or Google Cloud components and you express it in code and then Alchemy magically makes that expression of what your cloud should look like appear in the cloud without you having to click on lots of stuff, which is nice. So it lets your agents build everything without you having to deal with it Cool and then there's some I think, well, I've talked about this on different episodes.
00:45:17.637 --> 00:45:32.898
Just briefly, there's like, there's some layering you can do with effect there's some tools by Kit Langton that are really nice and maybe we'll just link to that in the in the notes cuz I don't wanna waste too much time.
00:45:33.097 --> 00:45:39.458
So that's like that's my that's my TypeScript stack and then, sorry, one more thing, on the front-end React and Tailwind.
00:45:39.797 --> 00:45:46.797
And again maybe a lot of people know what these tools are but like if you think about why are they good, right?
00:45:47.157 --> 00:45:50.918
It's like a normal web app without React and Tailwind.
00:45:51.338 --> 00:45:54.737
You have HTML, CSS and JavaScript in three different places.
00:45:54.757 --> 00:46:13.657
You have descriptions of your project and when when you want code to manage layout or layout to think about styles, it's so messy and they're in different files Tailwind brings your CSS into your layout, your style into your layout, and React brings your layout into your code.
00:46:13.958 --> 00:46:21.918
And so if you use React and Tailwind, then your layout, your code and your style are all together and it's so easy to see what's going on.
00:46:22.038 --> 00:46:26.737
It's very powerful, programmable, your agents can understand it, there's nice functional programming things.
00:46:27.998 --> 00:46:29.518
So, that's that. Alright.
00:46:30.237 --> 00:46:50.208
Mm-hmm. So, if I listen to your list, the commonality in your TypeScript stack is that these are all choices that might affect large parts of the code base.
00:46:50.998 --> 00:47:04.077
Like TypeScript gives machine-checkable contracts, effects structures application logic, Oxlint's fast linter, Bun consolidates the tooling.
00:47:04.858 --> 00:47:05.737
What did you say more?
00:47:05.757 --> 00:47:08.338
React and Tailwind, they shape the front-end.
00:47:08.938 --> 00:47:24.358
So is the bar actually what deserves to be specified up front when building is the bar whether it it changes the shape of the whole project?
00:47:25.197 --> 00:47:31.358
And is it substantial enough that you don't want the agent to re-invent it?
00:47:32.157 --> 00:47:32.858
Yes.
00:47:34.018 --> 00:47:46.958
Yes. Because a lot of libraries you, well, it's better to have your agent just build the parts of the libraries that it needs from scratch that fits for your project.
00:47:47.277 --> 00:47:49.438
And so you don't wanna pull in lots and lots of libraries.
00:47:49.818 --> 00:48:24.717
But these big frameworks that change the way that all the code is written and expressed and managed, these are the kinds of things that are important to bring in and for my in my like opinionated stack it's everything to do with more feedback loops, more structure for the more more yeah machine-checkable guardrails for the for the agents so that they can make their own decisions but have consistency if they don't have all the code in their contacts at once.
00:48:26.197 --> 00:48:29.717
Mm-hmm. And I assume, this is your list.
00:48:30.077 --> 00:48:44.978
And I assume for people to come up with their own list, with their own preferences, is the recommendation again, well, just try a lot of different stuff, see what works, doesn't work Yeah. I guess, I guess now there are coding agents.
00:48:44.998 --> 00:48:49.557
It's also easier to try and faster to try a lot of different tools.
00:48:49.577 --> 00:48:52.338
You don't need to like learn new programming languages to try it.
00:48:53.077 --> 00:48:54.597
It's very cheap to try things, yeah.
00:48:54.978 --> 00:49:14.277
And I would really I would love to hear from people who disagree with this stack with your reasoning why and I would love to talk to you about it somewhere in the comments or on Twitter or whatever cuz I'm very I'm very happy with this stack and it's been it's been tested.
00:49:14.297 --> 00:49:30.077
It's been I've built now maybe like with at least with different pieces of this, I don't know, twenty or so projects with agents in the last few months and yeah it's just I'm very happy with it.
00:49:31.637 --> 00:49:46.177
Yeah. I feel I feel like I I'm starting to grasp where you do the steering where you do the steering but it's I but on the other hand I'm still trying to reconcile.
00:49:46.797 --> 00:50:18.038
Like when on side you're stripping away like you're not prescribing too much but on the other hand you're adding stronger types prescribing to use effects and linting and it's Yeah, it makes me wonder, like, why is one why is the one unnecessary and why is the other one valuable to prescribe?
00:50:19.057 --> 00:50:19.757
Yeah.
00:50:19.938 --> 00:50:39.257
this is this particular list I've arrived at from trying less and seeing things fall apart and trying more and seeing the agents be too slow and and like spend too many tokens.
00:50:39.657 --> 00:50:49.818
So this is just my like Goldilocks and I've and I don't know, I really recommend Effect.
00:50:50.237 --> 00:51:17.637
It's it's probably the least the least popular of this list, so maybe lots of people haven't seen it, but it's still fairly it's starting to become like more commonplace to see people who are like me that need effect and they don't want to write any code without effect and they hate the idea of their agents not having effect making their code safer so I really I really recommend trying it.
00:51:18.217 --> 00:51:44.057
So yeah, so I told it this and then I started and it's a grilling so I get more questions there's Yes, I told it yeah, Rust makes sense for the core, around it put TypeScript and use my TypeScript stack which is Effect, V4, Oxlint, Bun, Alchemy and React and React and Tailwind I my I think I say later.
00:51:44.217 --> 00:51:59.117
And then the agent says, that stack makes sense, they'll always say that and yeah and then and then they ask a question like, OK if we're using Cloudflare for the for the back-end how should we do multiplayer?
00:51:59.137 --> 00:52:00.757
It asks a specific question about that.
00:52:00.777 --> 00:52:05.737
I was like, yeah then it's asking a question I don't quite understand.
00:52:05.757 --> 00:52:09.117
I'm like, hey why are we why are we verifying this on the server?
00:52:10.958 --> 00:52:15.177
And then something interesting happens in the grilling and I wanna focus on this.
00:52:16.378 --> 00:52:20.338
I saw some words that made my alarms go off.
00:52:21.197 --> 00:52:28.057
I saw checksum and like submitted traces for inconsistencies.
00:52:28.478 --> 00:53:18.697
OK what I've observed is when you ask the agents to to build something and to like be responsible, write tests, have good feedback loops but you're generic and you're not paying attention they tend to become bureaucratic and like really over the top with there's something that Doodle Steen calls process porn they're just like they love They love they love checksumming things, hashing things, comparing hashes, signatures. They like ceremonies, signing ceremonies usually this is just slop that gets in the way.
00:53:19.057 --> 00:53:33.518
So I How do you, how do you recognize that? Because maybe maybe it could also be like, whoa, this looks very rigorous, but how do you distinguish that it's over the top?
00:53:34.838 --> 00:53:41.898
And. It's just one of the the first few times that the agents were doing this stuff, I was like, wow, that's cool.
00:53:41.918 --> 00:53:42.978
I never thought about that.
00:53:42.998 --> 00:53:58.057
And then I was like, oh, this is why this is why people don't do this. It's a waste of time and so this is something that you get a feel for as you just build a lot of things with agents and maybe one day the models will stop suggesting these dumb things, but right now they do.
00:53:58.277 --> 00:54:08.458
So If we want models to stop suggesting these things, can we shift left and tell the model just not to do process porn from the beginning?
00:54:09.398 --> 00:54:14.978
Yeah. I think that actually is the right kind of thing to put in your agents dot MD.
00:54:15.498 --> 00:54:30.217
These like principles or guidelines, the kinds of things that you would tell a competent human in an organization, these are the kinds of things that still make sense in your in your skills in your Agents dot MD because it's not prescriptive.
00:54:30.657 --> 00:54:35.637
It's just kind of giving a general rule that can help the agents make decisions that are not stupid.
00:54:36.297 --> 00:54:44.867
So yeah, so no process porn I would I would say I should have put that in my Agents dot MD here and I'm gonna start doing that.
00:54:44.898 --> 00:54:53.217
Right. So a general principle like, hey, prefer simple designs and avoid unnecessary bureaucracy.
00:54:53.498 --> 00:55:00.458
So something broad and not really prescriptive like never use checksums.
00:55:01.237 --> 00:55:18.858
Exactly. Cuz sometimes it makes sense so like I asked I said why do we even need checksums? And it's like oh we don't need them but they're good for debugging blah and then it said and then it said something else like the snapshot might omit random generator state and I was like why?
00:55:19.418 --> 00:55:28.697
Like our state is deterministic why do and then I said we don't need this process porn no unnecessary duplication of computation, no unnecessary signatures, no unnecessary checksums.
00:55:29.498 --> 00:55:39.757
And then and then the agent got what I meant so this is yeah this is just an interesting an interesting thing that happened during the grilling.
00:55:39.777 --> 00:56:00.637
And this is the kind of thing that I suggest you look for when you're doing this process, like impose your opinion, impose your taste and try different things out even if you don't hold that opinion strongly yourself yet, experiment, experiment with both sides and see what happens.
00:56:01.637 --> 00:56:02.338
Mm-hmm.
00:56:02.938 --> 00:56:25.438
Yeah, that makes sense cool. So then yep, React and V with Tailwind is like the front-end choice Vite is just nice for like it just makes React easier to use and it has some good testing things built in and then there's some details here that I don't I don't wanna go over. So I'm gonna just fast forward.
00:56:26.217 --> 00:56:29.998
Doo-doo-doo And OK.
00:56:30.898 --> 00:56:35.277
And then is this stack agreed?
00:56:35.297 --> 00:56:46.157
This is the this is the skill finishing up, like, is our shared understanding in place? And then I said yes, and then the agent tries to start building. And I said, no, don't start.
00:56:46.818 --> 00:56:49.717
Let's discuss technical architecture now that we pinned our stack.
00:56:50.697 --> 00:56:51.858
And then I said, grill me on that.
00:56:52.217 --> 00:57:14.788
OK, so now we're transitioning from tech stack to technical architecture and and this is where I like to kind of this is this is this is like the point at which you should be thinking to yourself, OK, how big am I gonna make this project?
00:57:14.788 --> 00:57:20.498
Like like is this something that I'm just doing for fun in an afternoon and I'm gonna throw it away?
00:57:20.737 --> 00:58:01.418
Then like maybe I shouldn't spend so much time getting all these details right or if it's the start of something that you want to, continue building for a while, then I'll spend more time and I will go over I'm actually not gonna just go in, I'm not gonna go into the details of the specific architecture decisions, cuz that that's just like engineering but what I am gonna say is after I did the general architecture I said let's see, scrolling scrolling That's a lot of back and forth.
00:58:01.538 --> 00:58:02.697
You kept. Yep, a lot of back and forth.
00:58:02.717 --> 00:58:04.498
You discussed a lot with the agents.
00:58:04.878 --> 00:58:05.577
Mm-hmm.
00:58:06.637 --> 00:58:24.157
And Yeah oh sorry. Let me just highlight this as I was scrolling I found something Another thing that the models like to do, which is really annoying and horrible and the code gets messy, is backwards compatibility.
00:58:25.358 --> 00:58:30.478
Like, usually, well, first of all, when you're starting a new project, you do not need backwards compatibility.
00:58:30.788 --> 00:58:38.918
So and even when you're even when you're changing existing code, a lot of times you don't need backwards compatibility.
00:58:39.438 --> 00:58:49.237
But by default, the agents will do that and it'll really make your code messy and it'll start to slow down development to try and, like, keep all these things working at the same time.
00:58:49.237 --> 00:59:03.958
So so whenever I feel like there's some kind of backwards compatibility conversation happening, I'm like no like don't like it's okay to make breaking changes, we don't need to save other formats until we release the thing.
00:59:04.637 --> 00:59:07.617
So anyway, that was just one little aside.
00:59:07.637 --> 00:59:08.637
I'm gonna keep scrolling.
00:59:09.318 --> 00:59:12.498
OK. So, again, the grilling ended.
00:59:13.318 --> 00:59:15.398
Does this capture a shared understanding of the architecture?
00:59:15.418 --> 00:59:21.717
I was like, yes, but I wanna talk more about this detail of the architecture, cuz I think it's really important.
00:59:22.197 --> 00:59:31.458
In this case I wanted to hear I wanted to talk through how do we keep the state machine really fast and I had some ideas.
00:59:31.898 --> 00:59:35.677
I said zero copies, pure math logic, SIMD where possible, et etcetera.
00:59:36.958 --> 00:59:44.518
But still composable and testable and make impossible states unrepresentable PILs a la Lexi Lambda.
00:59:45.197 --> 01:00:20.518
Grill me. OK? So so this first part is like some engineering stuff, so I'm not gonna talk about it, but I wanna share one more kind of general guideline or principle that I like to steer the agents towards at this maybe before we jump into that I have a question because at this point you've gone from the idea to the stack to the architecture and then you zoom in again on individual engineering decisions. Like, how do when to stop or how deep to go?
01:00:21.217 --> 01:00:26.637
what make you say like, OK, let's let this deserves another discussion before we start building.
01:00:27.257 --> 01:00:27.958
Mm-hmm.
01:00:28.657 --> 01:00:41.378
So this is the Yeah, so this, I would I would classify this not as, well, it's an engineering decision, but it's specifically another type of architecture decision.
01:00:42.338 --> 01:01:00.677
It's I got a feeling it was a mixture of knowing, OK, I care about this project enough that I wanna work on it for a little bit, maybe, and so I want to do a good job thinking through the architecture with the agent before we start.
01:01:02.318 --> 01:01:09.197
So there was some of that, and so that sort of pushes me to try and specify more parts of the architecture before starting.
01:01:10.518 --> 01:01:35.257
And then the other thing is I had this just idea nagging in my head of like, OK, we need to I think it's important to keep the state machine fast, but I also don't want really crappy code because this is this is a an important part of the system that is gonna be that's like it's the core for the whole game.
01:01:35.257 --> 01:01:39.338
Like the thing that I care about is that the mechanics the game feels the same as the original.
01:01:40.038 --> 01:01:56.018
So so that's why I chose this thing to go deep on Do you also have for like a general guideline or like some logic that people who want to follow this workflow can apply and know like, OK, this is how deep I should go?
01:01:56.617 --> 01:01:58.338
I would say it's it's hard.
01:01:58.358 --> 01:02:00.858
It it's a function of how big the project is.
01:02:00.898 --> 01:02:07.938
So if it's a if it's a small thing that you're gonna throw away in a day, then you might not even need to grill at all for the architecture.
01:02:08.557 --> 01:02:28.818
And if it's more, like, if you expect it to grow and be big and you're gonna be building on it for a long time, then you might wanna nail more of these details up front and go into more of the subsections of the architecture ahead of time I don't have any hard and fast rules, it's just kind of it's just kind of a, feeling, I guess.
01:02:28.958 --> 01:02:39.257
So it sounds like maybe if you're building simple workflows and applications, it might still work if you don't go super deep.
01:02:40.217 --> 01:02:45.737
Because if it even if it's not built in a most rigorous way, cracks might not really show yet.
01:02:45.858 --> 01:02:53.737
But if you're building something that needs to scale up, then if you don't do it in the right way, then cracks can start to show.
01:02:54.458 --> 01:02:59.697
And then it becomes really important that you when it's deep enough to steer the agents.
01:03:00.498 --> 01:03:07.557
Yeah. And in those cases, yeah, you may have to go back and rewrite parts of your project if you didn't do this ahead of time, which is doable but it's annoying.
01:03:08.077 --> 01:03:16.677
So I wanna just quickly say So so I decided to go deeper on this one part.
01:03:17.197 --> 01:03:20.137
Part of this was engineering details so I'm not gonna talk about it.
01:03:20.318 --> 01:03:46.398
But one part is a principle I have that I try to, like, put into my agent in certain places, like this one which is saying make impossible states unrepresentable this is like a phrase from a from Lexi Lambda.
01:03:46.418 --> 01:03:47.838
It's a popular blog post.
01:03:47.858 --> 01:03:51.838
We used it in the last episode it was from Ryan Lopopolo's prompt.
01:03:52.197 --> 01:04:20.858
I liked this as a way to give the agent guidance to do kind of exactly what I want, which is this make impossible states unrepresentable is it's informing the agent that you want to structure your code in a particular way so that more is caught at the type checking stage. So that the agents can get that feedback faster and your code is more likely to be correct by the time it's built.
01:04:21.378 --> 01:04:49.297
So because Now, sort of when you're writing TypeScript or when your agent is writing TypeScript in effect, it's gonna be doing a pretty good job of this, but Rust especially Rust code that's like a rewrite of this complicated N64 emulator, it's not going to had a feeling that the agent wouldn't do this unless I told it to. So I told it.
01:04:49.918 --> 01:04:56.797
And then and then and then again I get some nice grilling I'm not gonna go over those details.
01:04:57.197 --> 01:04:59.538
It's not important. But we did a little grilling back and forth.
01:04:59.858 --> 01:05:18.527
And actually I will say the I suggested a particular engineering choice that I thought would be helpful SIMD, that's what it's called and the agent told me why it's probably not worth doing this right now, maybe we can add it later if it becomes important.
01:05:18.557 --> 01:05:20.577
And I was like, yeah, that's a good point.
01:05:20.577 --> 01:05:21.458
We actually don't need this.
01:05:22.018 --> 01:05:35.918
So I think it's nice to it's nice to suggest things to the agent without commanding them when you're thinking through decisions like this because the agent will push back with reasons why it might not make sense and then you might change your change your opinion.
01:05:36.697 --> 01:05:44.637
So. Yeah, it's really like you're working together with an intelligent person developing like the plan together instead of specifying, dictating.
01:05:45.498 --> 01:05:46.197
Exactly.
01:05:46.677 --> 01:05:52.398
OK, so I'm gonna fast forward now to the point that we talked about. So here it is again.
01:05:52.777 --> 01:05:54.617
Does that capture our shared understanding?
01:05:55.277 --> 01:06:01.637
Yes. Now, now let's talk about testing approaches, the verification stage.
01:05:59.458 --> 01:06:01.637
This is very important.
01:06:02.277 --> 01:06:04.818
So this is all about the feedback loops.
01:06:04.838 --> 01:06:14.777
You wanna make sure that the feedback loops are or that the agent understands the kinds of feedback loops that work best for this project.
01:06:15.237 --> 01:06:27.018
And this is where in my current version of the workflow I try not to be too prescriptive. And I don't actually prescribe particular tests which I used to do.
01:06:27.398 --> 01:06:35.297
But now I just now, I start with my ramble was like, OK, what are the kinds of tests that we want?
01:06:35.318 --> 01:06:36.998
Here's ones that I think are important.
01:06:37.458 --> 01:06:40.898
We want some side effect-free things.
01:06:40.918 --> 01:06:43.498
This is nice to make testing more deterministic.
01:06:44.057 --> 01:06:54.038
And it's and it's nice with the effect framework I said we could have some, a small number of tests that are doing, that are actually like looking at the real game.
01:06:54.057 --> 01:07:11.458
But if you have too many, then the testing stage is too slow. And this is something we've talked about on a few episodes, including the last one, where if your testing phase is really long, then your development speed goes really slow because after every change you have to wait a long time for tests. So.
01:07:12.918 --> 01:07:38.838
So so basically I my rambling is trying to like tease out this kind of response from the agent where I wanna see the different parts of the of the code base, kind of broken up into pieces like this where there's different areas, and what is the main testing approach, the verification loop for that specific part and then I can audit that and see if it makes sense.
01:07:39.418 --> 01:07:45.498
And in this case, I don't remember if it made sense to me or not, but it did. I said it looks good to me.
01:07:46.117 --> 01:07:53.302
So but I I wanted to you Just a quick look.
01:07:53.322 --> 01:07:57.202
So it what did it? OK, I see.
01:07:58.362 --> 01:08:07.777
OK, cool. I'm sorry, I just wanted to see if there's any surprising things that it suggested, but it looks it looks fine. Remember I said Oracle already, so it's gonna.
01:08:08.117 --> 01:08:40.158
So there's the Oracle there's Fast Tests which I wanna see Effect Effect Logic so this is like part of the effect framework The multiplayer testing is gonna be complicated A few actual comparisons but a lot of precise state to sprite audio assertions that's like getting the graphics and audio to be the same.
01:08:41.658 --> 01:08:50.837
So anyway property-based testing is nice to see and effect giving control over side effects is what I wanted to see.
01:08:51.438 --> 01:09:04.747
So I guess what you're doing in this phase is when you as human are not longer in the loop, what feedback will the agent have to tell whether its work is correct?
01:09:05.698 --> 01:09:29.738
Yes. Yes and in this in this project I didn't I didn't go into a big back and forth with the agent, but usually I spend some time on this stage as well because yeah, like it like you said, it's it's one of the important things that makes sure when you wake up the next day the thing is working properly.
01:09:29.858 --> 01:09:41.757
Well, so if you get it wrong and there's too too many slow tests, then when you wake up the next day it won't be done. And if you get it wrong and there's not enough good tests, then you'll wake up the next day and it'll be done but it'll be incorrect.
01:09:42.778 --> 01:09:44.118
And both of those things are wrong.
01:09:44.698 --> 01:09:51.698
So So it sounds like successful autonomy is downstream of test and verification.
01:09:52.898 --> 01:10:06.537
like you could tell Astra now go build the entire thing, and that might sound very reckless in isolation, but if you by that point already defined idea and the target.
01:10:06.877 --> 01:10:20.097
You already made the architecture decisions very deliberately and you've given your agent ways to detect when it's wrong, then you can have more confidence in letting the coding agent build a really large chunk of the work autonomously.
01:10:21.858 --> 01:10:52.398
Exactly. And that's exactly that's exactly where we're at in the process there's the documents are in place and and at this point I was ready to start an agent run overnight and I literally just said read read all the documents build use use some sub-agents. Anyway, we can we can go into those details later. But Now I get how you built the Pokémon puzzle project.
01:10:52.398 --> 01:10:56.877
You didn't just say like, OK, build it now for me and wake up to a finished game.
01:10:57.257 --> 01:11:11.457
There was some really substantial work and back and forth before the implementation and because you steered ahead of time the agent worked autonomously throughout the night.
01:11:13.097 --> 01:11:43.677
Yeah. It's yeah. It also when I look at this workflow, I guess I can see how it generalizes across different projects too, like the substance, like the questions and the choices, like your list of preferred tools in stack, they might be different for every person or maybe even different for different kinds of projects, but the overall framework and the mechanism is still the same.
01:11:45.358 --> 01:11:51.757
Yes. And you'd be surprised I use the same TypeScript stack for every project now, almost every project.
01:11:52.057 --> 01:12:03.377
It's very rare that I that I deviate from that stack at this point and yeah and I like you said, it's a it's a function of my taste and my experience and my opinions.
01:12:03.757 --> 01:12:05.778
You but I just recommend that people try it.
01:12:05.898 --> 01:12:07.818
You should try it and see what works for you.
01:12:08.278 --> 01:12:13.757
And I wanna hear if you disagree with something, cuz I I'm open to trying other other tools also.
01:12:14.757 --> 01:12:15.457
Mm-hmm.
01:12:15.877 --> 01:12:20.738
Yeah, I think Yeah, I I'm excited to try this workflow too.
01:12:21.238 --> 01:12:28.358
I think the biggest tension that I predict when I try it out is how do I know what to specify and what to leave to the model?
01:12:28.938 --> 01:12:35.917
And initially when you were talking it felt very contradictory, like OK, models are getting smarter.
01:12:35.938 --> 01:12:40.457
Delete these instructions. Don't prescribe too much in agents dot MD.
01:12:40.478 --> 01:12:53.097
Delete the skills but in in our conversation I see that you keep piling up more decisions together with the agents like use TypeScript, use Effect, use Bun.
01:12:53.597 --> 01:13:00.137
And make sure don't apply process porn don't use checksums and stuff.
01:13:00.717 --> 01:13:22.097
It still feels like a lot of steering but what I hear is you try to preserve steering that has a compounding effect on the quality and correctness of the project while continually deleting steering that's mostly constraining how a capable model solves the problem.
01:13:23.358 --> 01:13:39.557
So there's not really a clean rule and that also means that AI's not removing engineering judgments and that engineering judgment is still important but I guess it's Yeah. But I guess it's moving the judgments more upstream, right?
01:13:40.078 --> 01:13:52.557
Like before coding agents, were you applying most of your engineering judgments, like where did you apply most of your engineering judgment before coding agents and compared to where you apply it now?
01:13:54.518 --> 01:14:51.497
I think a lot a lot I still did I still did upfront planning for bigger chunks of projects but a lot of it was discovered in real time And a lot of these choices around a tech stack are like you don't make them very often because before you weren't starting projects as often because it took longer to do work and and then it was it was it was still valuable to have these guardrails like good lints and tests and all these things for mature projects but when for smaller projects they were less important because there were less humans involved and it was easy to share the kind of I don't know get on the same page with people. So it I think what it's doing is it's taking things that used to only happen on big projects and bring them and make them valuable on everyday projects that even one person is making.
01:14:52.297 --> 01:15:01.177
And and yeah, like you said, there there's a little bit of some of the decisions that we used to make sort of in the middle of a project we're now doing more upfront.
01:15:01.818 --> 01:15:05.318
I'm doing more upfront and it's and it's working out.
01:15:07.217 --> 01:15:11.938
Yeah, I think maybe one more thing I wanna, well, just remind people.
01:15:11.957 --> 01:15:25.677
I think I said it earlier, but the it's good to have a set of principles and guidelines like as a way as an expression of your taste.
01:15:25.698 --> 01:15:52.698
And these are the things, like you said, that you can steer without being prescriptive so I I hope that people have were interested in hearing my collection of those things today. But yeah, I'm really interested in hearing if there's if there's well, if you all have specific ones that you use, Yeah, I think this was also the most interesting part for me in this conversation.
01:15:53.518 --> 01:16:05.778
LLMs have gotten really smart, but engineers building at the edge are still spending human attention and exposing technical decisions in order to build rigorous software with coding agents.
01:16:06.738 --> 01:16:18.018
So, yeah, if you are still developing this intuition what to specify for the model and what to leave to the model, we hope that this episode helped and gave you a useful starting point.
01:16:18.917 --> 01:16:29.658
Pay attention to decisions that shape the whole project, challenge complexity that cannot explain what it buys you and invest heavily in feedback loops that let the agent catch its own mistakes.
01:16:31.637 --> 01:16:36.658
So in the next episode we will again share more emerging practices so you can get more out of your coding agents.
01:16:37.278 --> 01:16:43.518
If you don't want to miss out then hit the subscribe button and thanks everyone for listening and see you next time.