00:00:00.239 --> 00:00:10.080
So imagine you're teaching a class on climate change, and instead of pulling up a stock diagram or a YouTube clip that you half remember from last semester, you just describe what you need.
00:00:10.320 --> 00:00:14.560
Show me a coastal city 50 years from now with the sea levels risen by a meter.
00:00:14.720 --> 00:00:16.160
Make it photorealistic.
00:00:16.399 --> 00:00:19.519
Now show me the same city with aggressive mitigation in place.
00:00:19.679 --> 00:00:21.359
Now let my students walk through it.
00:00:21.519 --> 00:00:27.440
Now that used to be a production budget, a specialist, a six-week timeline, and probably a specialist vendor.
00:00:27.920 --> 00:00:30.320
This week it became a text prompt.
00:00:30.559 --> 00:00:36.240
We're at one of those moments where the gap between what was possible and what is possible just moved.
00:00:36.560 --> 00:00:37.920
Fast and in public.
00:00:38.159 --> 00:00:42.079
And I think, we think, a lot of educators haven't quite clocked it yet.
00:00:42.240 --> 00:00:43.280
So that's the episode.
00:00:43.439 --> 00:00:48.960
What just shifted, why it matters, and what the people who are thinking hardest about this are actually betting on.
00:00:49.280 --> 00:00:50.320
And I want to be up front.
00:00:50.399 --> 00:01:00.479
I'm assuming this isn't one of these hype episodes that haven't been sponsored by Google or wherever we're going to be talking about it, but we want to tell you what's real, what's already in our hands, and what it actually means for the work.
00:01:00.799 --> 00:01:01.200
100%.
00:01:01.520 --> 00:01:04.560
So no paradigm shifts, no keynote voice.
00:01:05.040 --> 00:01:07.599
Maybe, maybe a little bit, maybe a little bit of a keynote voice.
00:01:07.760 --> 00:01:09.680
But first, let's roll the music.
00:01:18.640 --> 00:01:20.799
Welcome everyone to Adjunct Intelligence.
00:01:20.879 --> 00:01:22.400
My name is Dale Zinski.
00:01:23.040 --> 00:01:24.400
Head of AI Education.
00:01:24.480 --> 00:01:26.159
I'm joined by Nick McIntosh.
00:01:26.319 --> 00:01:27.760
Nick, how are you tracking?
00:01:27.920 --> 00:01:30.640
Is your mind just blown by this week?
00:01:30.719 --> 00:01:33.200
By all the exciting things that are around?
00:01:33.680 --> 00:01:36.799
I can't believe what's happened in this last week.
00:01:36.879 --> 00:01:41.599
I'm super excited and I'm especially happy to stray away from the dark path.
00:01:41.680 --> 00:01:43.359
Like, I mean, this is what gets me into this stuff.
00:01:43.439 --> 00:01:49.680
This is what gets me up in the morning, and I'm really excited just for a change of pace to get hyped, not hyped, appropriately excited with you.
00:01:49.760 --> 00:01:50.799
How are you doing with it?
00:01:51.200 --> 00:01:56.000
I think we've done enough of the like the good stuff the stuff where we can go to bed and rest at night.
00:01:56.079 --> 00:02:00.159
We warned everyone about the reasons why we should be cautious about AI, etc.
00:02:00.319 --> 00:02:00.640
etc.
00:02:01.040 --> 00:02:06.159
I just want to kind of geek out about it for a bit because there is some so much cool stuff happening in the AI world.
00:02:06.400 --> 00:02:07.760
But let's get into it.
00:02:08.560 --> 00:02:09.919
Let's start with some context.
00:02:10.080 --> 00:02:19.120
So for the last decade, and I mean that literally a decade, some of the most serious people in AI have been making a consistent argument that most of the world has ignored.
00:02:19.360 --> 00:02:25.680
Large language models, LLMs, they are remarkable, but they have a fundamental inbuilt limitation.
00:02:25.840 --> 00:02:27.120
They learn from text.
00:02:27.360 --> 00:02:29.520
Surprisingly, the world isn't text.
00:02:29.680 --> 00:02:32.159
The world is spatial, it's physical, it's dynamic.
00:02:32.319 --> 00:02:36.319
It has gravity and fluid dynamics and light and cause and effect.
00:02:36.560 --> 00:02:42.080
And no matter how much text you train a model on, you can't fully capture that through language alone.
00:02:42.240 --> 00:02:53.599
Jan Le Kun, so one of three people who won the Turing Award for essentially inventing modern deep learning, currently running his own AI lab, but XMeta, has been making this argument publicly since at least 2015.
00:02:53.680 --> 00:02:55.599
He's been almost evangelical about it.
00:02:55.759 --> 00:02:59.199
LLMs aren't the destination, they are a detour.
00:02:59.360 --> 00:03:10.960
Feife Li, aka the godmother of AI, the researcher who built ImageNet, one of the foundational data sets that made modern AI possible, published a landmark essay last year making the same case.
00:03:11.199 --> 00:03:20.080
She called LLM's quote, Wordsmiths in the Dark, eloquent but ungrounded, knowledgeable but disconnected from physical reality.
00:03:20.400 --> 00:03:25.120
Her argument was that the next real frontier is what she called spatial intelligence.
00:03:25.280 --> 00:03:35.919
Teaching machines not just to describe the world, but to understand it, to simulate it, to generate it, to let humans and agents move through it and interact with it.
00:03:36.159 --> 00:03:38.639
Now the systems that do this are called world models.
00:03:38.719 --> 00:03:42.159
And the vision was, is genuinely exciting.
00:03:42.240 --> 00:03:51.759
So these are machines that understand physics that can build consistent environments, that reason about what should happen next in a scene and not just what words should come next in a sentence.
00:03:52.000 --> 00:03:58.159
Now when Lee wrote that essay, it felt important, and it also felt, honestly, like a 2028 story.
00:03:58.400 --> 00:04:04.800
Feifei Lee's own company, World Labs, just made their Frontier World Model Marble available to everyone.
00:04:04.960 --> 00:04:07.120
So not a preview, not a wait list.
00:04:07.439 --> 00:04:17.199
Anyone can go to marble.worldlabs.ai right now and type a text description and get back a photorealistic 3D environment they can move through and edit.
00:04:17.360 --> 00:04:27.279
A Hobbit Kitchen, a castle courtyard, a space station, an ancient ruin, fully explorable, exportable, and editable through conversation.
00:04:27.759 --> 00:04:31.040
More recently, Google Deepmind released Genie 3.
00:04:31.120 --> 00:04:43.439
So we're talking a real-time interactive world model running at 20 to 24 frames per second, grounded in Google Street View data that generates photorealistic environments from text pumps that you can actually navigate.
00:04:43.600 --> 00:04:47.360
Now they named the education use cases themselves on the product page.
00:04:47.519 --> 00:04:51.839
So students exploring historical areas like ancient Rome, Greece.
00:04:52.079 --> 00:04:54.240
Google said that, not us.
00:04:54.720 --> 00:05:04.639
This week though, it got really exciting because Gemini Omni has rolled out to the Gemini app, to Google Flow, to YouTube Shorts for free, available now.
00:05:04.879 --> 00:05:12.800
So these are multi-turn conversational video editing where each instruction builds on the last, the scene stays consistent, and physics hold up.
00:05:12.959 --> 00:05:15.920
So protein folding as claymation, right?
00:05:16.079 --> 00:05:18.720
A violinist transported across environments.
00:05:19.120 --> 00:05:23.120
A coastal city, like I mentioned at the top, transformed through natural language.
00:05:23.360 --> 00:05:29.839
So these are three world model projects from two major organizations available to the public right now.
00:05:30.079 --> 00:05:31.279
So Dale, I've got to ask.
00:05:31.439 --> 00:05:35.199
So three products now available from all these different organizations.
00:05:35.360 --> 00:05:38.240
Is this, in your opinion, a coordinated moment?
00:05:38.399 --> 00:05:39.759
Is this coincidence?
00:05:40.000 --> 00:05:44.800
And for you, what's the specific thing, follow-up question, that that lands hardest for you?
00:05:44.879 --> 00:05:52.720
Is it the Marvel's 3D environments, the real-time interactivity of Genie, or just the cheer fact that Omni is free on YouTube Shorts?
00:05:53.040 --> 00:05:54.720
I'm blown away by Omni.
00:05:54.959 --> 00:05:58.639
I think Google Omni still has that kind of AI weirdness in certain bits.
00:05:58.800 --> 00:06:14.480
No doubt it kind of still puts it up, like the outputs are still climbing the uncanny valley, but actually on a technical level, so when you actually look at each individual frame or each individual like uh interaction within the video, it is incredible and I would argue on some occasions flawless.
00:06:14.560 --> 00:06:17.839
It's the way that fabric folds on like a jacket or a jumper.
00:06:17.920 --> 00:06:21.360
I was looking so closely for the weirdness, how water looks.
00:06:21.519 --> 00:06:27.680
I always wanted to be a visual effects artist kind of growing up, working on the likes of like your Transformers or like your Marvel movies.
00:06:27.920 --> 00:06:37.759
You would spend I would spend hours kind of learning how to simulate some of those really complex things where you can now just kind of say a few words, take a few snaps, and get very comparative results.
00:06:37.920 --> 00:06:40.399
I think it's a way to kind of milk the narrative a little bit.
00:06:40.480 --> 00:06:51.519
I personally subscribe to the fact that we've run out of training data, or so in terms of text-based training data, we've scanned all the books, we've looked at all of what's on the internet, and now it's just kind of in shittification.
00:06:51.839 --> 00:06:57.120
And so this next frontier of training data is simulating entire worlds to get over that hump.
00:06:57.360 --> 00:07:04.800
Think about the thousands of interactions at like a micron level when water collides compared to the words in a book or a post on Reddit.
00:07:04.959 --> 00:07:07.439
Then extrapolate that possibility on top of this.
00:07:07.600 --> 00:07:11.199
My mind actually can't keep up, it's just infinite spawn infinite spawn infinite.
00:07:11.759 --> 00:07:20.240
But I think the weirdest thing for this is similar to us eventually finding a cure for cancer with AI, that we first have to go through chatbots.
00:07:20.319 --> 00:07:24.959
For me, it's a bit odd that we have to go through what I would call snapchat filters.
00:07:25.120 --> 00:07:31.040
And I'll and I'll I'll I want to play some of the versions of this if you're watching on YouTube that I've been creating using Omni.
00:07:31.439 --> 00:07:36.879
But you have to go through this weird kind of Snapchat uh gimmicky version of this before we get to simulating entire worlds.
00:07:36.959 --> 00:07:40.240
So it's it's a weird path, but I guess that's the that consumer push.
00:07:40.399 --> 00:07:43.519
Um it's how you get the it's how you get the next money.
00:07:43.839 --> 00:07:44.639
A hundred percent.
00:07:44.800 --> 00:07:51.600
And uh, tying this back to to it the education side of things, because it's really easy to flatten this into AI makes cool visuals now.
00:07:51.759 --> 00:07:54.560
I think that's the least interesting version.
00:07:54.720 --> 00:08:01.680
The more interesting version is what happens to simulation as pedagogy when simulation stops requiring a budget.
00:08:01.920 --> 00:08:06.240
Now think of what think about what has historically been hard to teach visually, right?
00:08:06.399 --> 00:08:15.279
So anything inaccessible, ancient sites, cellular biology, dangerous industrial processes, anything where the complexity is in the dynamics, right?
00:08:15.360 --> 00:08:24.959
So how a disease spreads, how a structural load behaves, how a system responds under pressure, and then anything where personalization matters.
00:08:25.120 --> 00:08:29.600
So this concept applied to my context, my students, my region.
00:08:29.920 --> 00:08:32.320
Now, all of this is lived behind a wall.
00:08:32.559 --> 00:08:36.240
And the wall was expert production capability.
00:08:36.799 --> 00:08:41.519
And the thing that's so exciting about this, the great leveler, is that wall just came down.
00:08:41.840 --> 00:08:46.240
Now, not not fully, as you say, we're sort of in that Snapchat sort of filters era.
00:08:46.639 --> 00:08:50.879
Not perfectly, but the j the trajectory, the direction of travel is obvious.
00:08:51.120 --> 00:08:55.279
And the question for educators isn't whether this is impressive, because it clearly is.
00:08:55.600 --> 00:08:59.360
The question hopefully should be what do we build with it?
00:08:59.759 --> 00:09:03.039
And what does it do to us if we don't think about that deliberately?
00:09:03.759 --> 00:09:08.960
The promise of education, particularly in university context, is it's meant to be a safe environment to experiment.
00:09:09.120 --> 00:09:11.919
And so I think this is definitely a plus one in that column.
00:09:12.320 --> 00:09:15.679
But this is a this is a whole lot of safety.
00:09:16.000 --> 00:09:24.399
Like we're generating entire worlds here, and so I think even in the safety, there should be consequences, and that's the bit I'm I don't think we're fully expanded.
00:09:24.480 --> 00:09:28.559
Like, if we are able to create worlds, what does that mean to us who can now be to be gods?
00:09:28.639 --> 00:09:35.519
That is a different podcast episode, but it there's second and third aura effects to eventually contend with.
00:09:35.600 --> 00:09:37.919
But right now, let's focus on the Snapchat filters version of it.
00:09:38.000 --> 00:09:39.200
But what else is Google doing?
00:09:39.440 --> 00:09:43.519
No, but you're right, because Google is on an absolute tear at the moment.
00:09:43.600 --> 00:09:46.159
They had a really big week, and it's worth saying that out loud.
00:09:46.320 --> 00:09:53.840
So alongside Genie3, uh recently Gemini Omni, this last week, they also announced Gemini Spark.
00:09:53.919 --> 00:09:56.720
Now, this is their always on desktop agent.
00:09:56.879 --> 00:10:01.039
Think of it as an AI that doesn't wait for you to open it, it's running in the background.
00:10:01.120 --> 00:10:04.320
It watches context, it takes actions on your behalf.
00:10:04.480 --> 00:10:12.879
Rolling out in beta now, broader rollout this northern hemisphere summer, local file access and browser control by Q3.
00:10:13.360 --> 00:10:16.159
Okay, so we've spoken before about OpenClaw.
00:10:16.320 --> 00:10:22.399
Yeah, so the um the security nightmare keeping ITS directors awake and boarding since uh earlier this year.
00:10:22.559 --> 00:10:24.879
That was the open source version of the same idea.
00:10:25.039 --> 00:10:29.279
This is already out in the wild with many a fan and far more bots out there.
00:10:29.360 --> 00:10:31.279
We spoke about Maltbook and various other things.
00:10:31.440 --> 00:10:32.240
Uh the what was it?
00:10:32.320 --> 00:10:33.440
The Lethal Trifactor.
00:10:33.519 --> 00:10:35.279
There's a lot we can dig into there.
00:10:35.440 --> 00:10:40.240
But the thing is, people are running these things overnight, waking up to task completed.
00:10:40.960 --> 00:10:43.039
The agent layer is not theoretical.
00:10:43.200 --> 00:10:45.200
It's not next year, it's now.
00:10:45.440 --> 00:10:48.399
And Google's version is coming this summer.
00:10:48.639 --> 00:10:56.639
The gap between frontier AI capability and what's sitting in your students' pockets is basically close and it closed fast.
00:10:57.039 --> 00:11:11.600
It's that whole uh science experiment that like the guy and the the weird guy in the office or the weird guy at the playground would be like, I've spent a weekend building my my workflow in N8N and I've taped together these various APIs and it's on GitHub, it's so cool.
00:11:11.759 --> 00:11:14.000
And and you go, Yeah, sure, mate.
00:11:14.240 --> 00:11:17.759
But this is a consumer product available for like$9.99.
00:11:18.159 --> 00:11:22.080
And I'm gonna say it's based on technomas right now that look like they just work.
00:11:22.240 --> 00:11:30.399
And there's usually a bit of a chasm between the demo and the product, but I'd argue that chasm is maybe a year between actual functionality.
00:11:30.559 --> 00:11:32.960
Uh let's let's big note copilot for a little bit.
00:11:33.039 --> 00:11:38.399
It doesn't get much love in this space, but I think it's there now what they promised it would do a year ago.
00:11:38.639 --> 00:11:49.039
And so if we look at Google, who's got, I would say, in some spaces, a slightly better track record, the always-on agent, the personal assistant in your pocket on your phone, everyone has that now.
00:11:49.120 --> 00:11:51.519
So what does that what does that do to us as a consumer?
00:11:51.679 --> 00:11:52.960
People are gonna want this.
00:11:53.440 --> 00:11:54.159
No, 100%.
00:11:54.399 --> 00:12:01.840
I love that shout out, Copilot, because I mean, the other thing that we don't touch on enough in these kind of this is where we're airing away from the hype.
00:12:02.080 --> 00:12:17.200
The uh the the work system, the infrastructure built around that, so Microsoft Word plus Excel plus slides, or in the Google universe, docs, slidesheets, this is the thing that can talk to those things, and you have that entire ecosystem working together.
00:12:17.519 --> 00:12:19.360
My god, that's powerful.
00:12:19.600 --> 00:12:21.919
It's always the question around who's the best AI product.
00:12:22.000 --> 00:12:32.320
And I think if you've got Anthropic as the and Claude as the kind of single standalone products, and yes, you can tap into slides and tap into tools, and similar to um ChatGPT, it's very much standalone products.
00:12:32.399 --> 00:12:40.399
But some of these infrastructure plays, like your Geminis and your co-pilots, and maybe what Apple's gonna be doing in the next few weeks, that's probably where the future's gonna be.
00:12:40.559 --> 00:12:45.600
So they're two very different races, but one that the consumer probably wins at at this stage.
00:12:45.919 --> 00:12:46.960
Fingers crossed.
00:12:47.200 --> 00:12:50.080
But let's let's let's take a quick pivot.
00:12:50.159 --> 00:12:52.399
Let's go for let's go for a plot twist, okay?
00:12:52.720 --> 00:12:55.759
Because here is where it gets genuinely interesting.
00:12:55.919 --> 00:13:11.120
So in the middle of a week where the dominant story is about how fast and how capable AI is becoming, Mira Murati, so former CTO and briefly CEO of OpenAI, now running her own lab, made a very deliberate counter bit public.
00:13:11.440 --> 00:13:15.759
And to understand why it matters, you need to understand what she's actually building.
00:13:16.080 --> 00:13:23.039
I'm actually so glad you're bringing this because it's a story in the AI world that I haven't paid attention to because I'm just at information density.
00:13:23.120 --> 00:13:25.039
I've been too scared to look what Mira's doing.
00:13:25.120 --> 00:13:28.639
There's so many, only so many groundbreaking changes that one man can take.
00:13:30.159 --> 00:13:30.799
Yeah, no, 100%.
00:13:31.120 --> 00:13:34.799
Like, I mean, but you know who she is, and she has a voice worth listening to, right?
00:13:34.879 --> 00:13:40.639
Like, I mean, all this generation going AI, I feel like it was it was two days, if that.
00:13:40.960 --> 00:13:41.759
Yeah, yeah.
00:13:41.919 --> 00:13:44.080
But still, it's a line on the C V.
00:13:44.799 --> 00:13:45.919
Not that she needs it.
00:13:46.000 --> 00:13:47.840
Again, CTO for forever.
00:13:48.080 --> 00:13:53.759
But her lab, because she she walked away importantly from OpenAI, and she is a number, one of a number.
00:13:53.840 --> 00:13:59.039
Like, I mean, the Amade is, for example, Karpathi, who will come up uh probably later in this.
00:13:59.279 --> 00:14:05.200
The open AI alumni who have gone on to do their own things are interesting in themselves because of maybe what called them to do it.
00:14:05.279 --> 00:14:16.080
Okay, so Mira Murati though, her lab just released what they're calling interaction models, and the name is precise and it matters for what they're trying to do there.
00:14:16.320 --> 00:14:19.600
Because most AI voice interfaces work the same basic way.
00:14:19.679 --> 00:14:21.120
They're they're language-based, right?
00:14:21.200 --> 00:14:28.639
So you speak, the system transcribes, transcript goes to a language model, you get a response, cycle resets, turn-based, sequential.
00:14:28.720 --> 00:14:30.480
There's a very clear handoff point.
00:14:30.639 --> 00:14:34.639
Now, Marathi's interaction models don't work like that.
00:14:34.879 --> 00:14:40.879
They're built to natively understand the continuous, messy nature of human communication.
00:14:41.200 --> 00:14:57.440
Not just the words, but pauses, the parasocial timing of of of syntax and stress timing, the interruptions, the change in tone mid-sentence, the difference between a pause that means I'm thinking, and a pause that means I'm lost.
00:14:57.679 --> 00:15:01.279
Now the model is always perceiving, is always present.
00:15:01.360 --> 00:15:08.879
It's not waiting for its turn, it adapts in real time when you change direction, when you contradict yourself, when you're clearly not following.
00:15:09.360 --> 00:15:12.240
And that's fascinating technically.
00:15:12.480 --> 00:15:17.759
But what makes Marathi's position really striking is the explicit philosophy underneath it.
00:15:18.399 --> 00:15:26.399
So she said this week the best way to have many possible futures, so good futures, is to keep humans in the loop for as long as possible.
00:15:26.960 --> 00:15:32.559
And she's not saying that as a safety disclaimer, she is saying it as an architectural decision.
00:15:33.120 --> 00:15:41.679
So when most labs are racing towards AI that does increasingly complex work with less and less human involvement, she is explicitly building in the other direction.
00:15:41.919 --> 00:15:48.159
And her bet is that the most important variable isn't raw capability, is the quality of the collaboration.
00:15:48.399 --> 00:15:49.840
She used the word intent.
00:15:50.000 --> 00:15:52.240
So not capability, intent.
00:15:52.480 --> 00:15:56.320
Understanding what a person actually wants and not just what they typed.
00:15:56.559 --> 00:15:59.759
And that is a different vision of what AI is for.
00:15:59.919 --> 00:16:02.799
So, Dale, appreciate the information saturation point.
00:16:02.960 --> 00:16:04.720
I cannot imagine your to-do list.
00:16:04.879 --> 00:16:08.240
I'm gonna say not bullets, but measured in feet or flaws, perhaps.
00:16:08.399 --> 00:16:09.679
But does she have you convinced?
00:16:09.759 --> 00:16:11.519
Or at least do you see an argument here?
00:16:11.679 --> 00:16:13.759
Because I keep going back and forth, right?
00:16:13.919 --> 00:16:18.159
This whole keeping humans in the loop thing is an easy sentence.
00:16:18.480 --> 00:16:27.679
But on Marati's part, this is this a genuine philosophical commitment, or is it just smart positioning for an upstart lab that needs to differentiate?
00:16:27.840 --> 00:16:31.360
And what do you think it actually means for us, for for educators more broadly?
00:16:31.679 --> 00:16:34.399
Bear with me, because I think I've got a bit of an anecdote here.
00:16:34.559 --> 00:16:47.120
Centrini research, uh, they recently packed$15,000 of cash, a Xiaomi phone, and some nicotine patches into a pelican case and put a person on a speedboat into the Strait of Hamus.
00:16:47.679 --> 00:16:49.759
And they called him Analyst 3.
00:16:49.919 --> 00:16:51.279
Bear with me, there's a reason for this.
00:16:51.519 --> 00:16:52.320
Great, love it.
00:16:52.639 --> 00:17:02.720
Every major firm who's looking at what's happening in the Strait of Hamous, and to date this episode, currently it's closed or unclosed due to the tension between uh the Iranian government and the American government currently.
00:17:02.960 --> 00:17:09.759
So every major analyst firm is looking at satellite imagery regarding what's happening, you know, as how many boats are traveling through based on satellites.
00:17:09.839 --> 00:17:24.480
However, analyst 3, the guy with the the Pelican case, he was talking to smugglers, he was watching ships through a channel that the data said was barely open, half of them are running dark, there's no tracking seal, invisible to anyone who was sitting at a desk back in some analyst firm.
00:17:24.640 --> 00:17:28.960
And he came back with a very different picture of reality because he went there.
00:17:29.200 --> 00:17:33.680
That's what I keep coming back to when people are talking about humans being involved.
00:17:33.839 --> 00:17:45.839
There's a version of this that is about, and speaking about what uh Miriam Marietti's kind of her philosophical approach to this, it's about oversight and checkpoints, and this kind of human-dilute conversation and like what role does human play.
00:17:46.160 --> 00:17:48.880
It's it's not only exhausting, it's a little bit boring.
00:17:49.039 --> 00:17:51.359
I think the more interesting version is actually a lot simpler.
00:17:51.599 --> 00:18:03.680
Some things you can only know or be a part of because you were somewhere and you were involved, because someone hesitated, because they answered, because it smelled a certain way, although I'm not sure Marathi's uh model is looking at that.
00:18:03.839 --> 00:18:11.039
But I think Marathi's point is worth taking very seriously because intent isn't just what you want, it isn't just input, process, outcome.
00:18:11.279 --> 00:18:14.319
It's what you notice about somebody that you thought to uh to ask about.
00:18:14.480 --> 00:18:18.640
That's not really a design problem, that's just what it means to be a person in a place.
00:18:18.799 --> 00:18:29.119
So whilst education of humans remains a very human thing, humans with all their faculties and sense makings are not just in the loop here, but are all parts of it at all times.
00:18:29.359 --> 00:18:37.200
So I think plus one to the what Memorati's kind of working towards, I'm not sure I fully agree with her um her approach to this.
00:18:37.279 --> 00:18:40.400
I think there I think it's a slice of the pie versus the whole thing.
00:18:40.720 --> 00:18:43.359
And I think let's go back to kind of world models here.
00:18:43.519 --> 00:18:49.279
Perhaps they bridge some of the gap here as you as you don't need all the sensors and data to read a simulated world.
00:18:49.519 --> 00:18:53.519
So just a more mess to consider.
00:18:54.000 --> 00:18:54.960
No, no, I love it.
00:18:55.039 --> 00:19:10.799
And I mean the there's no there's no tension in that because it is I think what she offers is and the world models Yan Lakum, Fei Fei Li, and and a lot of them, is there are many different ways to approach this whole messy wicked problem, this open question.
00:19:11.200 --> 00:19:11.680
So yeah.
00:19:12.240 --> 00:19:13.519
It's the difference between humans.
00:19:13.920 --> 00:19:25.440
Humans aren't ones and zeros, and we're approaching one a ones, we're approaching a human world with a ones and zeros kind of lens through imports, outputs, and process, and I think that's where it kind of kind of starts to fall down.
00:19:25.519 --> 00:19:36.160
And I said, world simulation goes away to kind of bridge that, but um, I want to know how far we get before it is still just yes, no parts, it's all decision trees, it's all ones and zeros.
00:19:36.559 --> 00:19:40.160
I am trying so hard not to correct matrix jokes, you know what I mean?
00:19:40.240 --> 00:19:43.039
Like plank units, like I mean, just really nerding out.
00:19:43.119 --> 00:19:47.039
Like it's red pill, blue pill, like where we have to white rabbit, etc.
00:19:47.200 --> 00:19:47.359
etc.
00:19:47.519 --> 00:19:48.640
I haven't I might do that this weekend.
00:19:48.799 --> 00:19:50.880
I haven't done a matrix run in a while.
00:19:51.519 --> 00:19:52.640
Alright, let's bring this home.
00:19:52.880 --> 00:19:58.000
World models just shipped at consumer scale from multiple directions simultaneously.
00:19:58.160 --> 00:20:05.759
The agent layer is already running, and one of the most credible voices in AI is saying that humans need to stay genuinely in this.
00:20:06.000 --> 00:20:07.759
Those three things are in tension.
00:20:07.920 --> 00:20:11.519
And that tension is perhaps the most interesting thing of all.
00:20:11.680 --> 00:20:19.039
Because world models are about AI democratizing capability that used to require expertise and infrastructure, a wall, a gate, a moat.
00:20:19.200 --> 00:20:24.400
The agent story is about AI operating with less and less moment-to-moment human oversight.
00:20:24.799 --> 00:20:31.279
By contrast, Miradi's story is about deliberately designing for human present, human judgment, human intent at every step.
00:20:31.440 --> 00:20:35.519
Now, which of these trajectories shape education matters enormously.
00:20:35.759 --> 00:20:45.279
Because if it's world models, so AI is generative infrastructure, our job, the educated job, becomes curation, becomes context, the higher order thinking that AI can't simulate.
00:20:45.599 --> 00:20:54.799
If it's the agent layer, so AI is persistent autonomous worker, the question becomes what human oversight actually looks like when a lot of the scaffolding just runs.
00:20:55.200 --> 00:21:06.559
And if it's Murati's bit, so the AI is collaborative partner, then the educator's job is developing the human capacity to work with these systems well, to form good intent, to steer, to know when to push back.
00:21:07.279 --> 00:21:14.480
My honest read is that all three are happening simultaneously, and we're gonna spend the next few years figuring out how they coexist, right?
00:21:14.880 --> 00:21:18.480
A week ago I had a pretty clean map of where all of this was going.
00:21:18.640 --> 00:21:22.559
I think I arrogantly said, I am confident where AI is going, I can see the future.
00:21:22.880 --> 00:21:29.039
Agent layers on one side, machines are getting smarter, there's world models around, they're doing less, they're needing us less.
00:21:29.279 --> 00:21:33.839
I feel like it was a trajectory you could kind of draw, but this week the map doesn't fit anymore.
00:21:34.000 --> 00:21:35.839
My mind is once again blown.
00:21:36.000 --> 00:21:39.279
And the people who built this technology, they're not all walking in the same direction anymore.
00:21:39.359 --> 00:21:46.240
I think this tension across your FAFA le's, your Maratis, all kind of talking about different things, is really positive.
00:21:46.880 --> 00:21:49.039
Very different bets in the AI landscape.
00:21:49.200 --> 00:21:53.920
The one I keep getting pulled back to is probably the one that's a little more exciting, Maratis.
00:21:54.079 --> 00:21:59.839
And it's the same bet with that kind of centrini uh case I made when they put the Analyst III on the speedboat with the nicotine patches.
00:22:00.559 --> 00:22:04.000
For me, that's the most valuable position for the next decade.
00:22:04.160 --> 00:22:10.160
It's not about the machine doing the work, it's the person who's doing the work with the machine.
00:22:10.319 --> 00:22:15.599
They paid attention, they came back with what the system, uh, with information the system could do something with.
00:22:15.759 --> 00:22:18.720
So for anyone who teaches, that changes the brief a little bit.
00:22:18.880 --> 00:22:25.680
We've been preparing students to supervise the machine, to check it, to sit beside it, but maybe we should be preparing them for a different job.
00:22:25.839 --> 00:22:33.200
Maybe she would be preparing them for to be analyst three and and and collecting context and and working with the machine in that sense.
00:22:33.440 --> 00:22:35.839
But Nick, I heard there's one other little tidbit you've got for me.
00:22:35.920 --> 00:22:36.640
Yeah, yeah, yeah.
00:22:36.799 --> 00:22:43.440
So you you you're very much putting me in this sort of frame of mind with the analyst three and the speedboat and the nicotine patches, the zins.
00:22:43.599 --> 00:22:46.640
It feels like a heist movie, you son of a um I'm in.
00:22:46.880 --> 00:22:47.279
I'd do that.
00:22:47.359 --> 00:22:48.559
I'd do that bachelor.
00:22:48.880 --> 00:22:55.440
Undergrad of undergrad of awesome stories to be told after graduation.
00:22:55.680 --> 00:22:56.640
No, just really quick.
00:22:56.799 --> 00:23:01.440
This is a I don't I don't know quite how to team this without sounding like an absolute nerd, so I'm just gonna lean into it.
00:23:01.599 --> 00:23:03.920
Batman joining uh Superman, if you like.
00:23:04.000 --> 00:23:07.440
Or or you mentioned the the Avengers at the top, a very Avengers Assemble moment.
00:23:07.599 --> 00:23:09.599
Andre Carpathi, right?
00:23:09.759 --> 00:23:15.119
You know this name if you're not an Uber nerd, like I mean, uh my my colleague Taylor and myself.
00:23:15.440 --> 00:23:18.640
Carpathi was on the founding team at OpenAI.
00:23:18.799 --> 00:23:26.319
He was head of AI at Tesla, he did his PhD with Faye Fei Lee, who we mentioned earlier, the world models um expert godmother of AI.
00:23:27.039 --> 00:23:30.400
He has joined Anthropoc and we've been quiet on them today.
00:23:30.480 --> 00:23:37.359
And this is a get because he is going to be leading a team using Claude to accelerate Claude's own pre-training research.
00:23:37.519 --> 00:23:43.759
So, Claude getting better at training Claude with one of the most formidable people in the game in terms of this.
00:23:43.920 --> 00:23:49.200
I don't know what to make of this, but suffice it to say that this has been a hell of a week, right?
00:23:49.359 --> 00:23:54.079
And directions of travel hard to pick, and uh it's just an exciting time.
00:23:54.400 --> 00:23:55.680
It really is.
00:23:56.400 --> 00:23:57.839
Thank you everyone for listening.
00:23:57.920 --> 00:23:59.359
If you haven't already, leave us a review.
00:23:59.440 --> 00:24:02.559
It generally helps everyone who wants to find this podcast to find it.
00:24:02.799 --> 00:24:05.359
You can watch us on YouTube if you want to see our beautiful faces.
00:24:05.440 --> 00:24:11.279
We would love to see what's on what next week's episode, but it's a world that moves way too fast, as we've seen this very episode.
00:24:11.519 --> 00:24:13.440
Until then, stay curious, stay intelligent.
00:24:13.680 --> 00:24:15.680
Nick, I don't want to stay the human in the loop anymore.
00:24:15.759 --> 00:24:19.359
I think we just need to stay human in an AI world.