OM DENNE EPISODE
Your coding agent does not need a frontier LLM call for every judgment. Deciding which tests to run, which skill to load, or whether a prompt needs more context can be a bounded question.
Jev and other decision models return typed choices, scores, or yes/no probabilities quickly enough to sit inside repeated workflows. A valid answer can still be wrong, so placement and fallback rules matter.
We show a live prompt checker and a small, six-commit test-selection experiment, then explore agent routing, local Laya, and how to choose between code, a decision model, a generative model, or a human for each step.
Superlinear is a podcast about emerging practices for building with AI.
Hosted by Brandon Kase and Christine Yip.
Homepage: https://superlinear.fm
Twitter: https://x.com/superlinear_fm
LinkedIn: https://www.linkedin.com/company/superlinearfm
Brandon Kase: https://x.com/bkase_
Christine Yip: https://x.com/christinetyip
Music licensed through Soundstripe. Code: O2UT6CCEQNH4KT0D
VIS NOTER 🔗
UDSKRIFT 🔗
00:00:00.239 --> 00:00:15.900
So now we can like put judgment in places where we previously maybe couldn't afford an LLM call and then the architecture might change of the things that we build. What a decision model does, like Jev, gives you a probability distribution over those choices.
00:00:16.820 --> 00:00:29.179
I tested six commits from our podcast video editor and there's two hundred fifty-three test files in our project and it only ran between zero and thirteen percent of the of the total tests.
00:00:29.500 --> 00:00:43.479
Definitely feels like a new primitive for for building with agents. Do you have any thoughts h- about how Jev or decision models in general change about software architecture or like, building with agents?
00:00:44.240 --> 00:00:50.460
Welcome to Superlinear, a podcast about emerging practices for building with AI agents.
00:00:50.679 --> 00:01:02.380
Each episode, we cut through the noise to find one new angle worth trying. Then we unpack why it matters, where it works, and how you can apply it in your own workflows, so you can build more powerful things.
00:01:03.619 --> 00:01:07.640
We're your hosts, Christine Yip and Brandon Kase.
00:01:08.480 --> 00:01:10.620
Hey everyone, welcome to Superlinear.
00:01:11.239 --> 00:01:18.620
About a week ago, TypeSafe launched Jev, which they call a system one model, and many people in the space are also calling it a decision model.
00:01:19.299 --> 00:01:25.799
And since then, we've already seen local decision models like Laya and others show up in the same emerging space too.
00:01:26.579 --> 00:01:32.359
So these models are fast and cheap, but that's not actually what we found most interesting.
00:01:33.379 --> 00:01:35.439
They don't just write code or prose.
00:01:35.959 --> 00:01:44.659
You give them some state, you define the kinds of decisions they're allowed to make, and then they return a judgment that your software can act on.
00:01:45.579 --> 00:01:59.799
So the question we'll explore in this episode is, if I already have a coding agent workflow or application built around frontier LLMs, where would I actually put a decision model?
00:01:56.799 --> 00:02:02.750
Would it make the workflow or application better?
00:01:59.799 --> 00:02:04.260
And how should I be thinking through how to use it?
00:02:05.819 --> 00:02:18.120
So, let's dive right in. Brandon, if I'm already happily using frontier models like Astra, Fable, or GP- GPT Sol, why should I care that decision models like Jev exist?
00:02:18.979 --> 00:02:34.159
Well, I I think I think we're we're reaching the point where the most frontier models were Astra, Fable they're so expensive to use. They're expensive in terms of like tokens, they use your s- your plan too fast, they cost too much money, and they're slow.
00:02:34.840 --> 00:02:47.599
And so I think all of us have been experimenting recently like What parts of our process can we delegate to, to dumber models?
00:02:48.719 --> 00:03:08.060
Like which, which sub-agents or whi- which parts of our, our workflows can we decompose into models that are less intelligent in order to complete the work faster, with lower latency, and cheaper so that we don't use our subscription too fast?
00:03:09.060 --> 00:03:31.199
And and Jev kind of just raised the stakes of of the potential of what that means so that's that that's sort of what initially got me excited. Like, I'm already playing around with delegating to to Luna for for certain things can I delegate to Jev instead?
00:03:32.120 --> 00:04:08.819
Because Jev is an order of magnitude cheaper and faster and lower latency. So that's that's like the first thing that got me excited. But then when I dug into it I realized that this the level of cheapness and speed and low latency that that Jev introduced totally opens a new class of use cases. Like both as both in the in the position of a like a coding workflow, like the the way in which I'm building, and and in applications. So Yeah. Go ahead.
00:04:09.159 --> 00:04:25.220
And it's and initially you were delegating tasks to Dumber models that are cheaper and faster, but in Jev's case, you're you don't even need to compromise on intelligence, right?
00:04:22.540 --> 00:04:25.579
I mean, Jev is also, is not that dumb.
00:04:27.060 --> 00:04:30.620
Right. That's the thing. It's it's smarter than Luna.
00:04:32.000 --> 00:04:33.199
And cheaper and faster.
00:04:34.420 --> 00:04:48.420
It's not it's not quite a Astra and Fable but it's like I think it's equivalent to like Sonnet five and Terra in the intelligence s- scale but then like way faster and cheaper than anything that's existed before.
00:04:49.220 --> 00:04:55.819
So yeah if we if we look this this is from the TypeSafe AI the the homepage for for Jev.
00:04:56.360 --> 00:05:06.279
The the the point of this graph is to show, look, Jev is really intelligent and super super cheap.
00:05:03.180 --> 00:05:06.279
Let me just show this as well.
00:05:06.560 --> 00:05:18.259
This is the price per million tokens and on the left is input and the right is output and we see it's about ten, well it's ten dollars for Astra and ten dollars for five point one.
00:05:19.160 --> 00:05:24.540
And then Jev is zero point zero four dollars per million.
00:05:25.480 --> 00:05:29.399
And then on output on output, LLMs are more expensive.
00:05:30.040 --> 00:05:39.339
So so Astra and Fable five point one are fifty dollars, for example and and Jev is free.
00:05:40.279 --> 00:05:42.620
It's too cheap to meter. That's their that's their phrase.
00:05:43.620 --> 00:05:46.920
So it's exciting. It's it's just it's way different.
00:05:47.600 --> 00:05:48.100
Right
00:05:48.240 --> 00:06:08.819
Yeah, I could already imagine that if so, it's so cheap, which is crazy, that it opens up a lot of possibilities for people to use it in their workflows or in their applications either when you maybe for steps that you want to use at scale or yeah where the cost really matters.
00:06:09.639 --> 00:06:22.779
So now we can like put judgment in places where we previously maybe couldn't afford an LLM call but now we put a deci- decision model there and then the architecture might change of the things that we build.
00:06:23.620 --> 00:06:38.589
But maybe maybe before we maybe but just to dive into that and for some listeners to really understand how Jev works, what does my workflow or application actually send to a decision model and what kind of answer do I get back?
00:06:36.579 --> 00:06:40.459
Like how is it different from an LLM? Yeah.
00:06:40.480 --> 00:06:46.740
So so with an LLM well, with with both, You you you send some context.
00:06:47.699 --> 00:06:58.699
And and actually I made a diagram. I'm gonna show the diagram. This is an example of sending a asking about an email.
00:06:56.139 --> 00:07:43.339
OK? So I'll I'll I'll explain it for the people who can't see the diagram, the the listeners so so the the question is which inbox should this email go in? And then you can imagine there's the subject and the body of the email. And then when you give that to an LLM, you get back a sequence of tokens one at a time, like this email looks like a pitch deck. The pitch deck should go in the category of preparing for presentations and then th- that same. So that's that's like what an LLM does and what a decision model does, like Jev or or a local one like Laya is an example of a local model.
00:07:44.199 --> 00:07:48.759
A decision model will you you give it what the options are.
00:07:49.540 --> 00:08:08.680
So is it so for example like I forget what I said but you know presentation decks bug reports, personal correspondence and then the decision model gives you a probability distribution over those choices or over that decision.
00:08:05.819 --> 00:08:20.399
It can be a choice or a yes-no question. There there's a couple different configurations and and so you can take that probability distribution and then decide on decide you make make your decision.
00:08:20.939 --> 00:08:28.959
Right. Right. So basically I guess decision model is a new name for classification models.
00:08:29.879 --> 00:08:50.379
Like when I when I was learning about machine learning I I learned that back then we called models that assign a class or a category to input data was called a classification model. So maybe this is the decision model is a new name for classification models.
00:08:51.700 --> 00:08:59.600
So it doesn't generate like text or prose but or even code. But it's really assigning a class label.
00:09:00.500 --> 00:09:02.259
And you you can define a class label.
00:09:04.299 --> 00:09:34.360
Right. Exactly but differently from historical classification models the or at least I didn't know of them but I'm pretty sure like in kind of classical machine learning classification models you need to s- specify the choices the classes when you're training the model But in in these new new-aged decision models, you provide the choices as part of the input.
00:09:36.620 --> 00:09:37.059
Mm-hmm.
00:09:37.419 --> 00:09:40.379
Yes, that is an, that is an interesting distinction.
00:09:40.860 --> 00:09:51.080
Before people start using these new decision models, is there s- is it, what is helpful to know about how it works before we apply it?
00:09:51.399 --> 00:10:09.360
Yeah, so, so let me go a little So in in JEV in particular there's there's three kinds of shapes of of questions so there's a choice and that's where you you give it the the classes.
00:10:10.559 --> 00:10:28.620
There's a score where you you kind of have a scale that you're scoring something on and then there's a yes no question which we would call like a bool but it's a non-deterministic bool so they call it a null spelled N O U L which is a probability distribution over yes or no.
00:10:29.340 --> 00:10:49.980
So these are these are the the shapes and this is important because as if you're thinking about using this thing you need to know kind of the type of the of the decision model. And that reminds me of a tangent that we that I wanna go on which is the company that created this is called TypeSafe AI. How cool is that? I love types.
00:10:50.100 --> 00:10:57.960
So that also caught my eye and I guess that leads me to another diagram.
00:10:59.860 --> 00:11:45.899
the by by type they mean you you a- you sort of ask for a shape and you'll always get that shape back. When you're when you're working with an LLM, if you ask the LLM to give you back a JSON, sometimes it doesn't which is so annoying or or it gives you the JSON in a slightly different shape. But with with Jev it's type-safe. And this is like an expectation that we have of decision models now But importantly, just because it's guaranteed to give you a valid answer of the valid shape, if you say the four options are bug, guess-pitch, feedback, and other, you'll get one of those four, but it might not be the correct answer. It still can be wrong. So it's smart.
00:11:46.179 --> 00:11:58.120
It's smart as SONNET five, but it's not necessarily, I don't know, it's not infinitely wise. It, it's important to still have that perspective, like these things still c- can make mistakes.
00:11:58.919 --> 00:12:19.039
I think that's it. It's like there's you, you, you give input, you, you give your context so and then you ask your question with the supplied answer type so one of these choices or yes, no, or a score and then you get an answer. That's like the core of of these decision models.
00:12:19.080 --> 00:12:19.399
Mm-hmm.
00:12:20.159 --> 00:12:26.320
I'm really excited to think about see some examples how how this is applied.
00:12:27.759 --> 00:12:38.379
What did you since it's been a week, like what have you seen like any favorite use cases or the way that people are applying Or maybe even like from your own experience?
00:12:39.039 --> 00:13:09.519
All of the above, all of the above so I guess let's let's show a few interesting use cases from Twitter to to start and then I can share about a couple things I created when I was playing with this. This is from Riley Brown and from from like like earl- late last week, Riley said he created a viral post analyzer.
00:13:10.580 --> 00:13:34.340
So as soon as you stop typing for point five seconds, it analyzes the viral potential and so the idea is, and then there's a little video of him using the tool, it kinda looks like Twitter, as he's typing his tweet he gets information about how viral the thing is so and as you know when he starts typing it's showing s- low score and then as he's explaining more detail and using better phrasing it, the score goes up.
00:13:36.240 --> 00:13:42.600
That's crazy. So Jev is so fast that it feels like, you know, live feedback.
00:13:44.360 --> 00:14:08.419
Exactly. Yeah. And so it's it's it's like a hu- between a hundred and two hundred milliseconds from the point of making the request and and getting the answer back. So that means you're paying for The network latency, then you have to run on on TypeSafe's hosting on their infrastructure and you're getting your response back. And that's happening in between a hundred and two hundred milliseconds, which is crazy.
00:14:09.220 --> 00:15:12.639
And I guess there's one maybe something in this in this example in in the little demo, Riley's not just getting a single number back, he's getting a score and a bunch of subscores and a bunch of basically he asked a bunch of questions and he gets all the answers to the questions at the same time. And this is an interesting feature of Jev that the the cost of asking one question versus asking five, for example in terms of time, it's almost zero time difference. And in terms of money, it's only twenty percent extra to do, you know, four times as much information. That makes sense. So so so this is kind of another another thing to keep in mind when you're when you're playing around with with Jev in particular. If you're using other decision models, they might not have this this feature. So it's important to to kind of be careful about that.
00:15:09.879 --> 00:15:39.440
Yeah, so that's one example I guess I was inspired by this kind of example to make a a live prompt rater so let me let me run it. So this is a this is a tool that I made when I was just playing with Jev. I thought, OK, how can how can this help me and my workflows to make me more productive?
00:15:40.039 --> 00:15:56.500
And I thought I create prompts to talk to my agents and you know what if there was a way to in real time get feedback about a prompt as I was writing it?
00:15:57.399 --> 00:16:22.080
And so you know so imagine I'm saying like build Tic Tac Toe. Yeah and I have so so in this case yeah the the idea is can we get real-time feedback as we're typing?
00:16:23.240 --> 00:16:45.799
Or if we can get real-time feedback as we're typing, what kinds of tools does this un- enable and unlock So getting I'm imagining obviously this is like a dumb prototype but I'm imagining I want to see fast feedback before I do an expensive action.
00:16:47.000 --> 00:16:57.240
That's like a something that's really interesting to me. You know, like sending sending a prompt to Astra or Fabel five point one is expensive.
00:16:58.559 --> 00:17:45.680
Like I wanna I have to wait my with my time and my tokens for an answer or for like a series of steps to happen even depending on the the prompt. And if there's a way to get that really fast feedback that doesn't slow me down, doesn't get in my way, but gives me that just little push of like, hey, you don't have enough context and so I I give a little more information and you know and now if you're looking at my screen the I got a little bit more of a percentage towards the prompt being clear and
00:17:45.700 --> 00:17:47.859
listeners who who do not see the screen.
00:17:48.259 --> 00:18:00.279
So in in this demo, Brandon put in his prompt, build Tic Tac Toe, and we saw in in we saw that this tool returns, OK.
00:18:01.200 --> 00:18:05.039
The cl- what the clearness of this prompt is at zero percent.
00:18:05.839 --> 00:18:17.119
And then Brandon just added build tic-tac-toe for the web. And the clearness jumped from zero percent to three percent and it almost felt like instant f- f- response.
00:18:19.940 --> 00:18:49.619
And if I said in React then then the clearness goes up a little bit more and the information or context missing goes down a little and these I mean so I I gave it three choices s- clear prompt for this task, information or context missing, duplicative or wasteful and and then I also gave like s- some other some other kind of annuals, non-deterministic yes-no's, like is the goal stated? Is there enough context included?
00:18:50.259 --> 00:18:53.380
What's the output format? And I get all of that back super fast as I'm typing.
00:18:54.299 --> 00:19:00.279
So that's this little little little silly example and I I. So so to
00:19:00.980 --> 00:19:23.019
just bef- before you continue, all those you decided on all those potential outcomes before you built this tool or when you were building this tool. So anyone who wants to build something with Jev, they need to make sure that they thoughtfully decide on the six.
00:19:24.539 --> 00:19:30.119
Well, in your example, those are six, but what is actually the classes that you want to get returned?
00:19:31.220 --> 00:19:40.200
Yes. And and either you need to decide up front or you can stick an LLM in the front.
00:19:40.579 --> 00:20:08.680
So now I've just shared another a diagram. So it's a a kind of a state machine. There's three blocks. One says propose, Then there's an arrow that says decide and then an arrow and then act and I'll just focus on that part. So the propose is this extra step where you can have you know, Jev can't determine its own choices but a normal LLM can generate choices based on some input.
00:20:09.019 --> 00:20:14.980
And then you can feed that into a decision model like Jev and then you can act on that.
00:20:15.519 --> 00:20:23.160
Yeah. And just to use your previous terms, you decide the choice and you decided that it should be scores per choice.
00:20:25.940 --> 00:20:48.799
Yes I think I think the in, in that example I had like newels, I had a I had a choice probability distribution which is probably the wrong type for for actually that that prompt now that I'm thinking about it. For my I wasn't thinking very carefully about the the the shape of the output, but when you're using this for something more real, you you wanna pay more attention to that.
00:20:49.279 --> 00:21:30.700
I actually can give an example of something where I paid a little more attention to the shape. So I asked Jev to pick which tests a commit needs to run. I think this was your idea actually, Christine, when we were first talking about Jev but I built I I I I put it into Claude as we were talking about it and built it and and kind of played with this a little bit. So Oh, cool. So so this is the this is a like disgusting slop website that a report website that Claude put together based on our conversation.
00:21:31.519 --> 00:21:36.819
So I I just wanna highlight the the outcome of this experiment. So so there was.
00:21:38.380 --> 00:21:42.960
Let me set the stage I tested six commits from our podcast video editor.
00:21:44.200 --> 00:21:49.960
And there's two hundred fifty-three test files in our project as of when I ran this.
00:21:51.119 --> 00:21:54.880
And and then there's a really simple question.
00:21:55.440 --> 00:22:00.279
Should this run based on the changes that we made in this commit?
00:22:01.119 --> 00:22:06.220
That that's the question t- given to Jev. And then Jev says yes or no with a probability.
00:22:07.180 --> 00:22:13.940
And and then if Jev says yes with probability greater than fifty percent, we run that test.
00:22:14.539 --> 00:22:49.539
And if Jev says yes with probability less than fifty percent, we would choose not to run it. And the question is, if we then go back and say, did we actually need to run that test for that commit did Jev s- successfully pick all the tests or not? And anyway, the answer is yes for all s- for all of them that I tried and it only ran between zero and thirteen percent of the of the total tests, which is cool. So and it costs a few cents, basically.
00:22:50.059 --> 00:23:05.619
Yeah. I think initially I thought it was interesting to try Jev for this was because you You know, this is a recurring step in your workflow. Like people building with agents, probably you're doing tests all the time.
00:23:06.160 --> 00:23:19.259
So it's interesting to use Jev because it's cheap and your test suite probably also grows over time. I in my project I have like more than four thousand tests.
00:23:20.160 --> 00:23:47.220
And probably you don't need four thousand tests you don't need to run all four thousand after building a feature. So yeah, so if every time an LLM would need to decide, go through all four thousand tests and decide which ones are relevant and it does it every time it builds something, it can get quite expensive and it will also probably slow down the process.
00:23:47.240 --> 00:23:54.740
But with Jev, now it suddenly, now it suddenly can be the test can be filtered in a fast and cheap way
00:23:56.700 --> 00:24:20.559
I mean, and and this is something we talked about in a prior episode, like, there's depending on how you're building your software and how your tests are structured, you might be able to build like a deterministic some deterministic code that chooses the tests because there's a way to tell like precisely exactly which test this file touches.
00:24:17.839 --> 00:24:47.700
But but that's not always true, especially if the the tests are really like are like playing a game in your case, for example, like playing through a game it's not it's so far away from the code that it's so the the the trace of like what files what changes in files could impact that end-to-end test run-through.
00:24:48.539 --> 00:24:57.339
It's like really hard and and yeah and and that's something you you need to use one of these tools that can that can reason to make that decision.
00:24:57.859 --> 00:25:06.240
So there is sometimes it is good to have deterministic code decide things.
00:25:07.380 --> 00:25:15.700
But sometimes it's also good to have still some intelligence involved And use a decision model like Jev instead.
00:25:16.279 --> 00:25:27.319
So you just mentioned deciding on tests as an example. Are there any o- other situations or use cases where that could also be relevant
00:25:29.000 --> 00:25:38.200
Yeah, I I think I think that's a really cool question what comes to mind for me is is like is like layered caching.
00:25:40.019 --> 00:25:46.140
OK? This is a real, real metaphor, so get ready for this so.
00:25:48.000 --> 00:25:52.359
Alright, so we're going on a, we're going on a journey here.
00:25:53.019 --> 00:25:54.059
Let's see how this works.
00:25:56.019 --> 00:26:01.599
back before we had LLMs, you had to draw things by hand or with Google slide shapes.
00:26:02.299 --> 00:26:11.980
So what I'm sharing on my screen is a slide from a presentation that I gave in two thousand and fourteen about or two thousand and sixteen? Anyway.
00:26:12.859 --> 00:26:13.599
Ten ten plus years
00:26:13.599 --> 00:26:13.920
ago.
00:26:14.200 --> 00:26:19.940
A while ago about caching and composable caches.
00:26:20.779 --> 00:27:06.779
But if we think about caching it's like it I think about it as this like abstract thing but I'll use a specific example here like like images on your phone when you're using like a social photo application how does your your the actual app sort of show you that picture as you go across different screens and if it's if it's built well then it will have abstractly a bunch of layered caching so the first time you open that app or you open that screen with that photo then you don't have it on your phone at all but it's in the cloud somewhere and so you download that image.
00:27:03.740 --> 00:27:20.900
And then let's say you you go to a different screen and come back that image might still be in memory and so you can load it from memory super fast. Now let's say you you like force quit the app and then or you open it like three days later.
00:27:21.880 --> 00:27:40.799
Well maybe you have it stored on disk so you don't have to go all the way to the cloud but you don't have it in memory and and so if you think about what the app is doing, it's saying, OK, I have this image I wanna load, like here's the name of it, here's a URL, something like that, and I first check RAM.
00:27:41.759 --> 00:27:58.700
And if I, if it's in RAM, then I'm like, yes, here's the image. And if it's not, then I, that's called a miss. And then we check disk. Is it in disk yes, then I get the image. Or no, it's not on disk, OK, now I need to go to the network. And so this is, this is a layering of, of caches.
00:27:59.180 --> 00:28:08.700
So now now that I did that aside the problem is, well for example, do I run this test?
00:28:11.240 --> 00:28:35.599
you could first say, OK do I have deterministic code that tells me with certainty that I can ru- that I should be running this or not? If the answer is yes, then great, you need to run it. Or or if the answer is you certainly don't need to run it, then that's also an answer and you can be done. Now, or that deterministic code that you wrote can't figure it out.
00:28:36.380 --> 00:28:50.619
And so then it's a miss. And then it goes to the next layer. And that next layer is now a decision model. This is our new thing that we have that we didn't have before. And now the decision model, it's a little more expensive, it's a little slower, and it costs a little bit of money. But it's still fast.
00:28:51.259 --> 00:28:57.460
And we can ask the question again, like, what test do I need to run for this change, for example.
00:28:58.519 --> 00:29:18.680
And you'll get back this one, or I couldn't determine and and if, you know, if you get back I can't determine, then maybe you fall back to a real, like, an old school LLM. I can't believe they're old school now. Then then then you go to an LLM and you pay more money and you wait more time to get the answer.
00:29:19.539 --> 00:29:29.720
And and this I'm using this this test ex- as an example but it's the same thing for for anything you could do the same thing for classifying an email.
00:29:30.539 --> 00:29:48.220
You could the email comes in and you can look at the subject and the body and you might have deterministic code like a regular expression that just extracts certain keywords and then you can classify it. Or if if you're missing those or there's ambiguity then you fall back to Jev or a decision model.
00:29:44.960 --> 00:30:14.000
And then again the decision model either gives you an answer or it can tell you if it gives you some uncertainty you might fall back again to a a smarter LLM. And so this kind of mental model of layered layered caching or you know a series of of of checks that are like cheap and fast at the front and expensive at the end and you kind of try them like dominoes until you get your answer. It's kind of like a really interesting thing
00:30:15.619 --> 00:30:18.819
that's that mental model makes sense.
00:30:19.779 --> 00:30:23.380
So, if you need exact computation, use code.
00:30:24.279 --> 00:30:32.180
If you need need judgments among known options, then you can use decision models.
00:30:33.160 --> 00:30:41.940
If you need a more open-ended answer, you use generative model. Or and then I guess if you need multi-step autonomous work then you use the coding agent.
00:30:43.180 --> 00:30:44.809
And and you can.
00:30:44.839 --> 00:30:52.819
Basically that's also like how it gets more expensive over time not over time but as you journey through this.
00:30:52.839 --> 00:30:56.799
As you journey down.
00:30:52.839 --> 00:31:04.099
Yeah. Yeah. And and and there's certain kinds of work that you can you can just hard-code which path you take.
00:31:04.880 --> 00:31:15.559
There's certain kinds of work where you can use Jev to decide which path you take This is like using decision models for routing.
00:31:15.980 --> 00:31:31.500
It's like a a a thing that people are playing with And then there's there's certain paths where you, yeah, ma- maybe you can do kind of the the caching example where you just try one at a time until you get something that works So it's it's I don't know. It like blows my mind.
00:31:31.900 --> 00:31:34.180
Cool. Yeah. So you applied that to.
00:31:35.720 --> 00:31:38.279
Well, we we could apply that to the test suite.
00:31:39.200 --> 00:31:54.539
Like is it easy to determine is this test relevant or not? Is it touching like the this file and and as and if you're not hunf- if you're not hundred percent sure you use Jev in the way that you just built it.
00:31:55.400 --> 00:32:11.140
I I'm just curious, can you also show like the prompts how you build your your first tool where you check the prompts live and also the one where you build the test I don't know how we call it, test classifier?
00:32:11.779 --> 00:32:14.599
Yes it's so easy.
00:32:11.779 --> 00:32:19.380
That's a good point though so you have to sign up for for Jev.
00:32:20.160 --> 00:32:29.799
So you go you go to TypeSafe TypeSafe dot I I and you sign up and there was a waitlist. I don't know if they're still having a waitlist Once you have access, OK, great.
00:32:30.400 --> 00:32:32.680
Yeah, so there's an agent skill reference.
00:32:33.119 --> 00:32:37.460
So if you're using Claude, you can install a plug-in. If you're using anything else, you can install s- install the skill.
00:32:38.279 --> 00:33:09.259
And then and then you literally just ask your agent. I'm sharing my my terminal now make me a one-page HTML web app with a text box and click 'Jev live' every time I write everything. I want a simple app where I can check if my prompts to the LLM are correct and I need more details. Use these following available choices and then I gave the you know clear clearer for this task information or context is missing and et cetera.
00:33:10.180 --> 00:33:12.319
And that's it. Anyway.
00:33:13.579 --> 00:33:34.099
It's just yeah. I mean you can you you you can build these simple prototypes in in one in a few sentences and that's why that's why we saw so much, like, innovation if you were paying attention to Twitter I guess I should show maybe let me just show one more thing. I mentioned like using Jev to to make a routing decision earlier.
00:33:34.680 --> 00:33:37.299
This is a cool example of a another routing decision.
00:33:37.660 --> 00:33:50.460
So I'm sharing a tweet by V E Chen it's at MIU twenty-one five ninety and this person said, people use Jev to pick a model before a task.
00:33:52.180 --> 00:33:57.099
I made it change GPT-six's reasoning effort inside Codex during the task.
00:33:58.359 --> 00:34:00.920
So this is a really interesting routing example
00:34:04.339 --> 00:34:11.800
I I feel like my first reaction is does does that not remove the cached, like the cache? It
00:34:12.019 --> 00:34:25.260
because you're not changing the model. If you if you change the model, then it would purge the cache. But changing the reasoning effort just changes, you know, how much the next prompt reasons, I guess.
00:34:26.099 --> 00:34:54.820
So so so it's cool. So you get you get faster. Anyway, he said fifty percent lower cost with Astra in in his example so I think this is like a really useful way that you can apply Jev to your workflow with your coding agents where it kind of like it shrinks into the background, you barely know it's there but the outcome is you can use Astra twice as much before you run out of tokens.
00:34:51.780 --> 00:34:54.820
So it's really valuable
00:34:56.219 --> 00:35:05.099
So Jev is so cheap that even if you do it every time you send a prompt, you still save money.
00:35:06.519 --> 00:35:06.960
Oh yeah.
00:35:06.960 --> 00:35:11.880
It's really cheap. I mean, think cuz, think about it.
00:35:09.300 --> 00:35:19.840
You're about to send a prompt to Astra which is a hundred to a thousand X more expensive.
00:35:20.539 --> 00:36:46.579
So, so you pay, you're paying a, a a zero point one to one percent overhead and that's like a really conservative upper bound cuz that's only the initial prompt, right? Not the whole back and forth with tool calls. Cuz that, it doesn't go to Jev every tool call in this example and what you get is a reasoning, a reasoning switch which makes Astra sort of do a lot less output tokens which are very expensive. The the other, I think the other just interesting routing example that people are, so many people have have tried this. I'm not gonna show a particular tweet but you can use Jev to choose your your tool or your skill so or you can do that sort of a- as a layer of indirection so for example the way that skills work right now before decision models is you load in the the description of the skill into your system prompt or into like sort of before you send your first user message. So if you have a hundred skills installed, you'll have like a hundred extra lines of of garbage that's in every prompt. And so your agent knows to call a skill if you're mentioning something about whatever whatever it is But you know, it's not super accurate.
00:36:47.099 --> 00:36:53.360
It's you're spending tokens every time And so what if you use Jev?
00:36:54.599 --> 00:37:08.480
Then re- Yeah, so the idea would be at every prompt if you're already using Jev to choose your reasoning, you can also ask another question about your prompt, which is are there any skills I should invoke?
00:37:10.500 --> 00:37:17.920
And then you you sort of run a Jev query against against your skills.
00:37:15.179 --> 00:37:33.239
And and you'll get back a probability distribution and you can then choose, kind of inject back into the harness what skills should run and that way, first of all, it's more accurate than relying on a- the attention of the model to decide and choose the skill, and then it's cheaper.
00:37:34.739 --> 00:38:09.260
And and same thing for tool calls, right? Imagine imagine you have like, there's a million ways to do this, but one way, you can have a single tool that you give to your coding agent, like you know, use GitHub and then you give it like a string or something and then the use GitHub tool actually kind of calls Jev indirectly and Jev decides which particular GitHub MCP call for example it's ma- is it makes rather than loading in the hundred GitHub MCP tools into your into your context if you're using MCP.
00:38:10.539 --> 00:38:11.119
Yeah so.
00:38:15.380 --> 00:38:32.139
Project and the step that you're in. And you you give it all the options which could be all the MCP tools or maybe skills or other tools and then basically that's that's a different model with with each with its own context.
00:38:32.820 --> 00:38:40.800
So you're not polluting your current main session with context about this specific choice.
00:38:41.639 --> 00:38:47.659
And once Jev has an answer, y- you just inject that answer into your current prompt.
00:38:48.719 --> 00:38:50.440
So you don't have that whole.
00:38:52.559 --> 00:38:58.880
Right. So these are like two parallel models working with its own context.
00:38:58.900 --> 00:39:10.260
you're chaining you're chaining together different model calls. You're using the Jev model as a tool for your main coding agent workflow. Yeah.
00:39:10.679 --> 00:39:34.780
Because I'm also wondering, like, how are any limitations? Like, I think if I remember correctly, Jev's context window is thirty-two K. So is there any limitation, like, OK, your if your s- number of your skills grows to a certain size or if the number of tools or if you're just searching the whole internet for possible tools then it will not be accurate with Jev
00:39:36.139 --> 00:40:03.760
may be important depending on what you're doing. And and that might be a reason you have to fall back to a cheap a cheap model rather than Jev. But I think a lot of for routing decisions, for example, you don't need that much context. You're not gonna you're not gonna decide to choose a skill based on ten prompts ago a message. You're really only gonna look at the most recent one, for example.
00:40:04.480 --> 00:40:19.900
But yeah, it is it is a limitation. And it's even more of a limitation when we look at other decision models that people are starting to create that run locally that have even smaller context windows like one K instead of thirty-two K or sixty-four K.
00:40:20.579 --> 00:40:30.920
Have have you run into any problems with Jev's context window? I'm just curious, like, whether you use it to a certain extent that w- where where Jev's context window was not sufficient anymore?
00:40:31.739 --> 00:40:33.599
And in and if so, in what situations?
00:40:35.239 --> 00:40:38.880
I haven't hit that limit yet in I mean I've only been experimenting a little bit.
00:40:39.860 --> 00:40:52.480
I have hit the context limit in local models that I've been playing with, local decision models, and that has caused them to not work so OK.
00:40:53.980 --> 00:40:57.159
I keep dancing. It's been less than a week so so no one That's
00:40:57.179 --> 00:41:06.440
played that much with Jev and so but it sounds like from your early experimenting sounds like TypeSafe chose a good context window for their products.
00:41:08.260 --> 00:41:16.699
It's it's fairly it's fairly it seems small compared to LLMs but it's big compared to other local decision models that are starting to pop up.
00:41:17.760 --> 00:41:37.340
I'm gonna just show one of the local ones that's popping up. May I? So Laya is all over my feed. The last day and a half. So this is by brain function collapse. I think that's a person's blog or something. I don't know. They claim twenty milliseconds per decision on a laptop GPU.
00:41:39.780 --> 00:41:40.300
Zero dollars.
00:41:41.119 --> 00:41:43.360
Cuz it runs locally.
00:41:41.119 --> 00:41:43.360
Cool. And it's small.
00:41:43.840 --> 00:41:59.539
Three hundred twenty-two parameters. Three hundred and twenty-two million parameters. And then there's this table that they, the Laya team claims that basically w- w- this is a comparison between Laya and Jev. Mm-hmm.
00:41:59.820 --> 00:42:05.940
Oh, perfect. I was just I just wanted to ask you for a comparison, like how big is Jev compared to Laya?
00:42:06.920 --> 00:42:13.500
Yes so we don't know how big Jev is because it's closed source and it's hidden.
00:42:10.280 --> 00:42:13.500
But we know how expensive
00:42:13.880 --> 00:42:14.159
it
00:42:14.159 --> 00:42:21.179
is. I don't think I haven't seen it but they we We know how expensive it is in terms of dollars and time.
00:42:22.659 --> 00:43:09.699
And and so we can compare a- accuracy, for example, and this table claims that Laya and Jev have similar accuracy on news, spam checking, choosing emotions. But then choosing a star rating, Jev is much smarter, for example and then you can also compare speed. So if you're seeing my screen there's a Flappy Bird game being played and Laya is making the decisions about whether, like what control should be sent to Flappy Bird for the bird to move through the pipes. And and Jev, Jev can play games too. Actually, let me just jump to this.
00:43:09.980 --> 00:43:30.039
Someone got Jev to play Mario. I think that's just a fun one to show for the the viewers so this was a tweet by Fadil Shaikh, if I'm saying that right and it's just a video of Super Mario Bros and Jev deciding on the controls to send to it.
00:43:30.360 --> 00:43:50.239
That's so interesting. So, I guess Jev is also very suitable for playing games because i- there's a a a bounded A number of choices that Jev can choose from, which is like, OK, you walk to the right or left or jump or shoot or whatever and then and then Jev can like choose the next action.
00:43:51.139 --> 00:43:54.820
And it's fast and it's cheap, so there it can take a lot of actions.
00:43:55.659 --> 00:43:59.000
it's not free and it's not in- it's not instant.
00:43:59.239 --> 00:44:21.280
And so the claim with these local models, such as Laya, is you can play games, Other games that you couldn't play with with Jev, like advanced Tetris because you need to make moves within like twenty milliseconds or, you know, fifty milliseconds instead of a hundred milliseconds or two hundred milliseconds.
00:44:22.079 --> 00:44:27.579
So, like, I saw this and I was like, oh my god, screw Jev. We can just use local models.
00:44:27.880 --> 00:45:02.940
It's faster, it's cheap, they're small anyway, and and if if the thing that matters is cost and latency for these super-fast models, then free is gonna be cheaper than any amount of money. It doesn't use too much of my GPU on my lo- on my computer, cuz the models are small, so it won't even won't even f- like phase me to be running it on my laptop. And latency is the thing that I care about. So if if I'm paying two hundred milliseconds and I could be paying twenty, then that will unlock a whole new set of use cases. And then and then Laya claims that it performs like as well as Jev.
00:45:00.739 --> 00:45:26.820
So, so then I tried it and I tried it on all of my examples and it just doesn't work for me. Well, I it didn't work on any of the things that I tried Jev on and I had to ask Claude like, hey, can you give me an example that Laya actually works on and and I was able to find something that worked.
00:45:27.719 --> 00:45:44.300
But but OK, so the same thing, it's easy to try Laya or to try any of these things. If you see another like local decision model, you just find the link to the thing and tell your agent, I wanna try this and then it'll download it and you can do whatever you want, right so and I I recommend everyone be doing this.
00:45:44.500 --> 00:45:54.760
Like you should just get in the habit of of knowing that there's no friction in trying things. You can try anything because your agent can f- can do the the hard work for you. So so.
00:45:55.619 --> 00:45:57.679
So I tried it. Yeah, everything didn't work.
00:45:58.260 --> 00:46:01.539
And why is that? Well, Laya has a one K context limit.
00:46:02.320 --> 00:46:21.639
So some of my examples actually needed more than one K of context. I didn't ever need thirty-two but I needed more than, more than one so that kind of broke a few it broke the the the diff, the example of picking the tests the other thing is, Laya does not have batching built in.
00:46:22.579 --> 00:46:28.460
So asking four questions is four times more expensive than asking one question. It's still free but it's four times slower.
00:46:29.420 --> 00:46:57.440
And so you can't actually ask fifty questions at the same time. Or or I I don't know what the limit is with Jev actually but let's say twenty or whatever if you ask twenty questions it's gonna be slower than Jev because Jev can answer all twenty at the same time. And so I was like OK what's up? What's going on here well other than those things I just said the the other thing is because Laya is a local model and it's small, you can fine-tune it.
00:46:58.199 --> 00:47:12.099
And if you fine-tune it, then the claim is it can get s- smarter and get close to Jev if you fine-tune it on a particular thing that you wanna learn so I'm gonna just show a diagram cuz I think this is really cool.
00:47:12.659 --> 00:47:16.159
So you can distill Jev into Laya.
00:47:17.059 --> 00:47:33.300
Cuz we have we have a decision model that's really smart and it's cheap enough to use to like generate a data set for fine-tuning. You know, it'll cost like a few cents probably to do like a few thousand pieces of synthetic data that you can generate with Jev.
00:47:30.860 --> 00:47:46.000
Yeah, and then you can you can fine-tune Laya and and the claim is that Laya will become smarter in that case. I haven't I haven't tried this yet. But but I think like these are the kinds of things you should be thinking about as decision models are popping up. Mm-hmm.
00:47:46.300 --> 00:47:49.099
How do you exactly s- distill Jev into Laya?
00:47:49.119 --> 00:48:08.699
Like, would you in that example where I have a test suite a four thousand tests and I want to have a local, cheap, fast model like Laya to be able to accurately decide which test is relevant for every feature that I build how can how how do I exactly
00:48:09.739 --> 00:49:07.820
I think, yeah I think for that example you would need a local decision model that has a bigger context window. So, but let me answer that question with an example that works with a smaller context window. Like the the prompt, the prompt checking example, the the first one I showed Yeah, so the the idea is you would you could use you could use an LLM to generate prompts for or you could use or you could mine your your like, chat history for real data of the prompts that you're sending to your coding agents then you give them all to Jev and get your answers from from Jev so for for wh- whatever you're you're trying to do for your application, like is this prompt, does it have enough information, is it clear, whatever, right?
00:49:08.460 --> 00:49:17.599
So you ask you and then you generate the these pairs of this prompt got this answer for all of your prompts from Jev.
00:49:18.739 --> 00:49:45.900
And then once you. And then o- once you have a labeled set then you then you you fine-tune Laya or whatever your your open-weight model is, which is, you know a bunch of matrix multiplies, I guess. I'm not a machine learning person, but but Claude is. So and then and then you get out you get out a model that will perform better on on that kind of of prompt.
00:49:46.599 --> 00:49:48.139
Or it'll perform closer to Jev.
00:49:48.679 --> 00:50:04.719
What comes to mind for me is it's very suitable when dur- w- when there's a Where you have a verification step where your agent can see or check or retrieve the results somewhere.
00:50:05.300 --> 00:50:38.840
And maybe we should actually build this for Superlinear the same like similar to the first example you showed from Riley where he built this viral tweet drafter we could we could build a tool where the agent would check the post that we sent out or maybe even the titles or the shorts that we sent out and then build a training data set based on how it performed.
00:50:39.699 --> 00:50:48.920
Like this is, you know, this short went viral, this short did not go viral. This was the description that was the title.
00:50:45.820 --> 00:50:55.420
And then based on that you could fine-tune a local model to help you draft viral posts or viral shorts.
00:50:56.159 --> 00:50:58.460
This must be how, like, machine learning people feel.
00:50:58.780 --> 00:51:38.719
It's like it's like, OK, like what like there's synthetic data we could create but how does the synthetic data understand virality and it it it doesn't, right? Like when we try and use Chatshubbety it doesn't it doesn't help us write tweets that go viral necessarily you can use like the real data from our channels but then you'd be You'd be getting things similar to the stuff we've already created. You could also pull in data from other people's posts and then you'd then you'd get things that are like, you know, good for whatever whichever posts we pulled applied to our stuff.
00:51:35.599 --> 00:51:54.820
But then then it's also like like things that are interesting to people change with time and so we'd we'd wanna like weigh things that are more recent, more heavy. There's all these kind of choices to make, which is interesting. But but I think like Yeah.
00:51:55.639 --> 00:52:09.659
I I think like the an interesting thing though that that now is unlocked is generating synthetic data has become an order of magnitude cheaper.
00:52:10.940 --> 00:52:17.780
Because Jev you can use Jev to make synthetic data where you used to only be able to use LLMs.
00:52:19.139 --> 00:52:21.380
And Jev is, you know, ten or a hundred times cheaper.
00:52:21.800 --> 00:52:23.599
And so now you can make your synthetic data. that
00:52:23.599 --> 00:52:37.099
work? Like, it would, you know, you would pull data and then Jev decides based on your based on how you framed it, is it how how suitable is it?
00:52:33.519 --> 00:52:37.099
And then that will be a a null?
00:52:37.619 --> 00:52:45.440
I'm not sure with with the v- A virality thing if if Jev will be how how good Jev would be at that. But
00:52:48.159 --> 00:52:53.019
Any real data, like when you're building a product and when you see how people interact with your products, like.
00:52:54.519 --> 00:53:27.480
Well, real data real data has a place and then synthetic data has a place. But but like if you were training if you were training something on on something where Jev is smart I can't think of I'll give you an example now playing playing a video game you can generate you can generate that data with Jev and then use that to train a new model or fine-tune an- an- another model. Whereas before you could only generate that data with you know, LLMs, I guess.
00:53:27.500 --> 00:53:28.739
Yes. No, I think it's
00:53:28.760 --> 00:53:31.679
interesting. I think it's yeah, it's a
00:53:31.679 --> 00:53:41.619
th- interesting thing to explore. Fine-tuning, you know, suddenly fine-tuning a local decision model is super-accessible now.
00:53:39.219 --> 00:53:43.300
It's it sounds like it's super-cheap and you get really high accuracy.
00:53:44.639 --> 00:53:58.420
Because at the beginning of the of this episode you said, l- whoa, you know, now with Jev I can build so many new things that wasn't possible or not really viable before what are other examples that make you feel that way?
00:53:59.300 --> 00:54:08.139
So we've seen like a live prompt checker, like almost like live test filtering What else?
00:54:08.940 --> 00:54:12.500
Oh and and of course the model routing, scale routing, and tool routing.
00:54:14.260 --> 00:54:21.300
I think the other kind of thing that really blew my mind which a a couple of people have built different things. Let me share my screen.
00:54:22.360 --> 00:54:22.940
Suspense.
00:54:24.260 --> 00:54:26.460
Something that blew my mind was browser use.
00:54:27.099 --> 00:54:59.480
This is an example I I pulled up a tweet by by Dax where he said he's the creator of Open Code he said, preview of how fast browser use can be powered by Jev and Open Code's browser use CLI and then there's a video of Open Code using browser use but if you can't watch the video it's like I mean it's browsing very fast, less than half a second on each page.
00:55:00.400 --> 00:55:08.420
And you know when when LLMs used to make these decisions, you had to wait for the LM and you had to pay for it.
00:55:08.940 --> 00:55:14.800
So so can you imagine, like, OK, I'm actually gonna share one more. I retweeted something.
00:55:15.559 --> 00:55:54.699
When you have browser use be fast and cheap, then you can build massively parallel browser-based adversarial testing. So there was a tweet by Rafal Volinsky where he shows a video of I don't know, what is this, like sixty-four sixty-four simultaneous browser use Windows? And and this is where, like, Jev earns its name. We haven't even talked about that. Anyway, this I'm showing a tweet where I say that I'll just read it.
00:55:56.000 --> 00:56:00.940
Now this is a Jev example that makes Jev earn its Jevons paradox name.
00:56:02.380 --> 00:56:45.420
Jev is named after Jevons paradox, which is the principle that as something becomes way more efficient, like cheap, fast, and smart, people end up using it more per unit cost of speed and intelligence, et cetera And so so this this is a real example where the fact that we can now do browser use really cheaply and by cheaply I mean time, we'll actually end up spending more money than when we had LLMs because before you might have ran like two or three of these kind of like browser runthroughs with an LLM trying to break your app if you were using any at all.
00:56:46.380 --> 00:56:52.920
But now that we have Jev it can run really fast, you're kind of incentivized to do like a thousand at the same time.
00:56:53.679 --> 00:58:34.960
And so, yeah, I I I think like Jev's paradox is quoted a lot when when people are talking about AI in this age where there's gonna be more jobs because AI makes people more productive even though it reduces the cost of completing a task. You'll just do more things and and like decision models kind of push us more in that direction, which I'm super excited about. Cuz cuz I think that's that's how I feel. Like I'm so excited about every time there's a new kind of tool that can like do more faster, cheaper because it makes us more productive as long as we're willing to jump in and try it and use it and experiment. And I think that's like the biggest that's the biggest thing that I want people to take away not like not any details about Jev or any of these local decision models but the fact that to try any of these things you just you just paste the URL into your coding agent and say I wanna try this and you give an idea you have or you say like help me come up with ideas and you can. And and that'll inspire you and you can create and do things and and I know, like, stuff changes so fast. Like, Christine and I were so upset that that Jev came out, like, the day after we recorded the last episode that we that we were then like edited and posted because, you know, it's so long now that Jev has been out and we're like, ah, we're behind we're I think all of us always feel behind but you can fight that by just acknowledging that there's no friction in trying things and lean into it.
00:58:35.599 --> 00:58:36.099
Yes.
00:58:36.639 --> 00:59:02.500
I think I think that examples are also starting to cluster for me. I think it was really interesting because now you can use like really cheap judgment broadly and then spend expensive reasoning only on a small number of cases that actually need it, like the caching layer that you talked about, Brandon so yeah, this is I could really I could see where people would apply or use Jev where scale matters.
00:59:03.139 --> 00:59:16.579
And then the last thing is then there's a combination of speed and scale where the decision model can basically sit inside the loop continuously and your browser use was a really good example.
00:59:16.880 --> 00:59:21.199
Like every step is another decision about what to click or n- to do next.
00:59:22.059 --> 00:59:27.099
And the ma- maybe the test selection could be another one.
00:59:27.440 --> 00:59:36.940
After every code change, quickly decide which tests are most likely relevant and run those tests.
00:59:32.679 --> 00:59:43.800
Like these workflows only really become practical if each decision is cheap and fast enough to repeat over and over. Mm-hmm.
00:59:44.539 --> 01:00:17.019
And these so so these buckets your like your coding something in your code agent workflow, speed, scale, or speed and scale hopefully these these like give you ideas of ways like if if you're running into problems or if you've thought of something in the past that that can help you or a friend or a colleague or or anyone if you if something you're thinking of kinda fits in these buckets then you should really try and and play with Jev and see what you can make
01:00:18.360 --> 01:00:46.940
Yeah. What what would be the first thing that you would advise people to try, like, right away, today, tomorrow? Because after we talked, I think the first thing that I would look for is think very critically about where in which places in my workflow or application do I not need s- really something generated, but I just need a good decision.
01:00:44.059 --> 01:00:48.019
And I need it cheaply, quickly, or many times.
01:00:51.039 --> 01:00:57.900
That's probably like how I would like go through my existing projects and and and think about it.
01:00:58.780 --> 01:01:18.639
Yeah, it's definitely easier to audit your existing workflow and kind of swap an LLM out for a decision model. That's easier conceptually than inventing a new place where a decision model should live when you You hadn't even thought of using any model in in that part of your workflow. So that that I think makes sense as a good starting point.
01:01:19.199 --> 01:01:35.659
And then I think, you know, get inspired. Hopefully you're inspired by some of the examples we talked about in this episode or you know, just search Jev on Twitter and you'll see lots of people playing with it. And and and then, you know, maybe that inspires you to to try something in your projects
01:01:37.219 --> 01:01:53.139
Mm-hmm. And after, you know, I think I now definitely see how Jev enables, like, new things to be built that weren't viable before. It definitely feels like a new primitive for for for building with agents.
01:01:54.119 --> 01:02:03.440
Do you have any thoughts, h- about how Jev or decision models in general change about software architecture or, like, building with agents?
01:02:06.019 --> 01:02:29.559
Every piece of, every workflow, every, it in, OK, in the same way that, like, when LLMs when LLMs came out, when coding agents came out, you had to rethink everything. You had to rethink everything that you were doing as a creator, as a builder. You had to rethink every part of the products that you were building for your users to see if, like, this primitive fits in.
01:02:29.860 --> 01:02:35.070
You have to do that same thing all over again with, with decision models now, which is so exciting.
01:02:35.099 --> 01:03:07.659
And hopefully this happens over and over and over again with new, new kinds of new kinds of primitives I, I, I think it's, there isn't a simple answer besides that. You just have to rethink everything. And, and as you said, make it clear, like anytime you're reaching for an LLM, are you making a decision or are you doing some kind of generation and even look at the deterministic parts of your workflows wherever you're making a decision that you didn't put an L M in cuz it had to be fast.
01:03:08.579 --> 01:03:14.639
Also question that, OK, now it can be fast to have this kind of artificial intelligence judgment.
01:03:15.679 --> 01:03:16.800
What can I do? What can I change?
01:03:17.579 --> 01:03:41.039
So so it feels like maybe similar to what we said before, like when you have a new workflow you could ask What kind of intelligence does this step actually need? And I remember you said, OK, if you need, sometimes you need code. For example, when you need exact logic, then you use code.
01:03:37.940 --> 01:03:50.360
When you need bounded fuzzy judgments then it could be a decision model. When you need novel output or reasoning, then it could be a generative model, an NLL an LLM.
01:03:51.119 --> 01:03:55.800
And when you need multi-step autonomy, then probably you want an agentic workflow.
01:03:57.179 --> 01:04:02.300
And maybe when you need accountability or maybe there's maybe there's also a place for humans.
01:04:03.719 --> 01:04:05.579
I don't know. Or don't you think so?
01:04:07.699 --> 01:04:23.599
Alright. Well I again I'm very excited to try out more things with decision models. I think the default move in the last few months have been whenever a software needs intelligence call an LLM.
01:04:24.260 --> 01:04:43.659
But now it seems like the default is changing and we have to be more thoughtful about OK which parts need deterministic code, which parts only need bounded judgment, which parts actually deserve expensive generation or reasoning and maybe which ones need only an agent or human.
01:04:44.460 --> 01:05:18.300
And if decision models make that middle layer good enough, fast enough, cheap enough, we may just start designing agent systems differently but yeah, we're still early. It's only been a week since Jev came out if you try Jev in your own harness and find a pattern that Brandon and I missed, please share it with us and we're super excited to hear about them. Every week we'll share emerging practices for building with coding agents and we hope to see you again next week.