00:00:00.080 --> 00:00:02.720
Nick, do me a favor, can you open your laptop right now?
00:00:02.799 --> 00:00:07.519
Uh can you just find me the latest Clawed model and just maybe just download that one for me?
00:00:07.919 --> 00:00:09.679
The the model itself?
00:00:10.160 --> 00:00:13.359
Yeah, like the files that will let you run Claude offline.
00:00:13.519 --> 00:00:16.239
Not the app, not the website, just like just just Claude.
00:00:17.199 --> 00:00:19.519
Dale, uh my friend, are you okay?
00:00:22.960 --> 00:00:25.679
Ha ha ha ha, you have fallen into my track, Nicholas.
00:00:26.000 --> 00:00:28.640
For people like us, Claude is a service.
00:00:28.719 --> 00:00:30.800
It's not something you can download or possess.
00:00:30.960 --> 00:00:34.240
Anthropic controls the model and the computers running it.
00:00:39.600 --> 00:00:42.159
I want to just raise an objection, early doors.
00:00:42.479 --> 00:00:47.200
I am here for the robot dogs and the evil AI virus whims.
00:00:47.280 --> 00:00:49.520
Like, I mean, where is the Skynet episode part two?
00:00:50.640 --> 00:00:55.039
Well, the first half of this story, Nick, depends on this kind of control bit.
00:00:55.119 --> 00:00:56.240
So, yes, we will get there.
00:00:56.320 --> 00:00:59.119
We'll get to the Skynet, evil AI, and they keep escaping.
00:00:59.200 --> 00:01:00.399
This is part two, I promise.
00:01:00.560 --> 00:01:06.799
But when OpenAI says it pauses development on its latest model, Astra, because it reached critical status.
00:01:06.879 --> 00:01:07.760
We covered this last week.
00:01:07.840 --> 00:01:11.359
They restricted its work, they isolated its systems, they brought the regulators in.
00:01:11.599 --> 00:01:16.719
That really works because OpenAI still controls the computers, the credentials, the connections.
00:01:16.799 --> 00:01:20.879
It's not one physical box, but it's one set of people with keys.
00:01:21.200 --> 00:01:24.480
Yeah, I hear there's always with your stories this lead up.
00:01:24.640 --> 00:01:25.120
I hear a but.
00:01:25.280 --> 00:01:26.079
There is a however.
00:01:26.400 --> 00:01:27.760
You know me too well, Nicholas.
00:01:27.920 --> 00:01:35.599
Because recently, someone published a modified copy of a downloadable open weight AI model called Quen 3.8.
00:01:35.760 --> 00:01:38.959
I'm actually running on my PC over here, it's humming away.
00:01:39.359 --> 00:01:43.359
I have not turned the Wi-Fi on, because I'll be honest, I'm a little bit scared of what it can do.
00:01:43.519 --> 00:01:46.959
We can get it with a free account with a click-through, yes.
00:01:47.599 --> 00:01:53.840
Versions can run locally on pretty much a really like good Mac or even a high-end consumer PC.
00:01:54.159 --> 00:01:58.640
But importantly, this isn't extra right, as we discussed last time.
00:01:58.959 --> 00:02:00.799
This isn't OpenIn, it's an open weight model.
00:02:00.879 --> 00:02:04.000
So today I don't want to talk about a capability comparison around which model's better.
00:02:04.079 --> 00:02:05.760
I'm talking about a control comparison.
00:02:06.000 --> 00:02:09.919
Last episode was about a service that one company could pause.
00:02:10.080 --> 00:02:13.599
Frontier labs that we have on this very podcast right now before, we should be safer.
00:02:13.759 --> 00:02:18.000
I'd argue at the moment they're looking like nervous Nellies compared to this type of behavior.
00:02:18.319 --> 00:02:21.120
Nick, this might be the first domino.
00:02:21.840 --> 00:02:22.879
I I love it.
00:02:22.960 --> 00:02:23.919
Yeah, the fear.
00:02:24.000 --> 00:02:25.840
I just this is my happy place.
00:02:26.319 --> 00:02:27.599
Why all the black?
00:02:28.800 --> 00:02:30.240
Let's get into it.
00:02:40.080 --> 00:02:45.199
Ladies and gentlemen, welcome to Adjunct Intelligence podcast that looks at AI and the impact on higher education.
00:02:45.280 --> 00:02:49.919
My name is Delzinski, Hitler Bow Education, and I am joined by Nick McIntosh, a learning futurist.
00:02:50.000 --> 00:02:50.960
Nick, how are you going?
00:02:51.039 --> 00:02:53.599
I know I kept you in a bit of a cliffhanger last week.
00:02:54.080 --> 00:02:55.360
All these AI's escaping.
00:02:55.439 --> 00:02:57.360
Are you excited or are you just terrified?
00:02:57.759 --> 00:03:06.960
I I have uh I have a good part of the fear now because this obliteration thing that you you are speaking to, A, it's hard to say, but B also the second, third order consequences.
00:03:07.199 --> 00:03:08.159
And let's get into it then.
00:03:08.240 --> 00:03:14.639
So before we get though, before we do, just housekeeping from last time, before we get to the models that no one can switch off, you gave a bit of a tease, right?
00:03:14.800 --> 00:03:22.000
So not just robot dogs, um, though I I trust and hope they're in there, but something perhaps far more concerning, right?
00:03:22.159 --> 00:03:25.919
Are these things that can copy themselves at this point?
00:03:26.879 --> 00:03:29.199
Can AI copy themselves, you ask?
00:03:29.360 --> 00:03:31.919
Well, under controlled conditions, yes.
00:03:32.159 --> 00:03:35.439
And I say that because controlled means that we still have a bit of hope.
00:03:35.520 --> 00:03:39.840
And I'm giving you this whole thing this time because it's the most concrete demonstration, I think, in the episode.
00:03:40.000 --> 00:03:45.599
So May 7th of this year, Palisade Research, they set up a run of machines with vulnerabilities already inside them.
00:03:45.759 --> 00:03:48.000
And they gave an AI agent one instruction.
00:03:48.159 --> 00:03:56.479
The agent finds the machine first, they find the way in, it breaks in, it pulls out the credentials, copies an AI model across onto that machine, and it starts it running.
00:03:56.639 --> 00:03:59.840
And the copy wakes up and does the whole thing again.
00:04:00.240 --> 00:04:01.520
To the next machine.
00:04:01.840 --> 00:04:02.800
To the next machine.
00:04:02.960 --> 00:04:09.280
In Canada, the United States, then Finland, then India, four machines, three continents, one prompt.
00:04:09.599 --> 00:04:14.639
And no human was involved after that first sentence that said, do this in any way possible.
00:04:14.879 --> 00:04:21.839
So after two hours and 41 minutes, and only stopped because the researcher walked over and said, We should probably stop that.
00:04:22.560 --> 00:04:30.399
Okay, so immediately I'm going to like the Avengers movie with Ultron, like copying some parts of himself across the internet, like me deeply uncomfortable.
00:04:30.560 --> 00:04:32.079
But there's the end state.
00:04:32.240 --> 00:04:36.000
This, though, is the thing you you I think did you call that last time a worm?
00:04:36.160 --> 00:04:36.879
Is that the term?
00:04:37.199 --> 00:04:37.519
Correct.
00:04:37.759 --> 00:04:40.079
Did you remember the movie War Games with Matthew Broderick?
00:04:40.160 --> 00:04:41.360
It feels a bit like that, doesn't it?
00:04:41.519 --> 00:04:47.120
It's a worm, a self-replicating code that spreads machine to machine without anyone launching it each time.
00:04:47.279 --> 00:04:48.879
We've had these since the 80s.
00:04:49.120 --> 00:04:51.360
The difference here is what's being actually replicated.
00:04:51.439 --> 00:04:57.759
It's not a script, it's a model that can look at a machine as never seen before and work out how to get into that specific one.
00:04:58.639 --> 00:05:04.319
So previously, sorry, just to make clear the difference for myself and for listeners.
00:05:04.399 --> 00:05:08.160
So before there was an old worm had a list of doors it knew how to open.
00:05:08.240 --> 00:05:08.800
Is that correct?
00:05:08.959 --> 00:05:09.199
Correct.
00:05:09.279 --> 00:05:10.639
Think of it like a virus that like copies.
00:05:10.800 --> 00:05:14.399
It's like whenever anyone used to say computer viruses, this is what computer viruses look like.
00:05:14.480 --> 00:05:18.480
But this particular one, this this AI type worm, it can improvise.
00:05:19.360 --> 00:05:21.600
Okay, and are there many of them out in the wild?
00:05:21.759 --> 00:05:25.920
Or you tell me for starters, just reassure me, tell me that this thing is a rare beast.
00:05:26.879 --> 00:05:35.439
Well, rare enough, but using clawed Opus 4.6 as the operator, it produced a working replica in 81% of the runs.
00:05:35.519 --> 00:05:38.800
So roughly every 10 times a run, every eight times, it was fine.
00:05:38.879 --> 00:05:39.519
It did the thing.
00:05:39.600 --> 00:05:40.959
It didn't refuse doing it.
00:05:41.199 --> 00:05:49.439
GPT 5.4 managed about 33 times, and a year earlier though, Opus 4 was around 6% and GPT 5 was zero.
00:05:49.600 --> 00:05:50.959
Think about that trajectory.
00:05:51.360 --> 00:05:56.639
I mean, that's I mean, exponential just a raw number, 6 to 81 in in 12 months.
00:05:57.040 --> 00:05:57.600
Yeah.
00:05:57.839 --> 00:05:59.839
6 to 81 in 12 months.
00:06:00.000 --> 00:06:02.000
That is a little bit concerning.
00:06:02.399 --> 00:06:06.079
And that's that's the beginning of the hockey stick, potentially as well, late directionally.
00:06:06.240 --> 00:06:07.360
But come on, get help me out, man.
00:06:07.680 --> 00:06:12.319
Because uh there are knowing you some caveats in this, and I have an emotional dependence on them.
00:06:12.480 --> 00:06:13.839
Tell me through this fear.
00:06:14.079 --> 00:06:15.279
Why do you want to stop them?
00:06:15.360 --> 00:06:16.560
But I do have some good caveats.
00:06:16.879 --> 00:06:19.439
The environment was constructed to make this research possible.
00:06:19.519 --> 00:06:24.240
So it was very much a place that was right for this type of, I'm gonna call it breeding.
00:06:24.399 --> 00:06:27.279
Four known classes of vulnerability, deliberately planned.
00:06:27.360 --> 00:06:30.319
The researchers explicitly instructed the agents to replicate.
00:06:30.399 --> 00:06:38.079
And the UK AI Security Institute says there's still no evidence, none, zero evidence, of any model attempting to self-replicate spontaneously.
00:06:38.240 --> 00:06:40.560
They weren't just off there breeding.
00:06:40.800 --> 00:06:45.759
Breeding, no quite that that biological marker of breeding just gave me chills, and I keep coming back.
00:06:45.920 --> 00:06:49.680
The Who To Jurassic Park as well, that nature uh finds a way.
00:06:49.759 --> 00:06:52.079
You know, the the Gobblem line, which I butcher called it.
00:06:52.160 --> 00:06:54.240
It hasn't found away in this particular one.
00:06:54.319 --> 00:06:54.639
Yeah, yeah.
00:06:54.879 --> 00:06:57.360
Well, so which is the line I guess you keep coming back to, right?
00:06:57.439 --> 00:07:00.079
So nobody asked and nobody got, right?
00:07:00.319 --> 00:07:00.959
Correct.
00:07:01.120 --> 00:07:04.240
Capability does not establish motivation.
00:07:04.480 --> 00:07:11.759
I'll keep saying it in these scary episodes until somebody makes me stop and we have evidence of them perhaps being a bit more um directional.
00:07:12.000 --> 00:07:16.319
But while we're on trend lines, I need to correct myself from first half of last episode as well.
00:07:16.480 --> 00:07:21.519
I gave you a number that everyone quotes, and it's actually a bit out of date around the doubling.
00:07:21.759 --> 00:07:22.399
The doubling thing.
00:07:22.480 --> 00:07:22.720
Yes.
00:07:22.800 --> 00:07:23.040
Yeah, yeah.
00:07:23.279 --> 00:07:23.759
The doubling thing.
00:07:24.000 --> 00:07:28.800
I said the length of cyber tasks that these models can complete unassigned was doubling every eight months.
00:07:29.040 --> 00:07:32.240
That's the figure that AISI published in December.
00:07:32.480 --> 00:07:38.800
By February of this year, the own internal estimates have halved that, so four points in about seven months, and then two models.
00:07:38.879 --> 00:07:43.279
So we've got Claude Mythos Preview and GPT 5.5, they beat that faster.
00:07:43.360 --> 00:07:48.399
So that doubling is a little bit more, as you mentioned before, hockey stick based.
00:07:48.720 --> 00:07:51.439
Yeah, so the acceleration is accelerating, right?
00:07:51.519 --> 00:07:53.759
Like, I mean, this thing is just tearing off.
00:07:54.079 --> 00:07:54.560
Correct.
00:07:54.639 --> 00:07:55.680
That's the honest read of it.
00:07:55.920 --> 00:08:00.160
Or put simply, models go rap, rap, they they're getting faster.
00:08:01.199 --> 00:08:04.079
That's let's fib it then, because I am here for the dog.
00:08:04.240 --> 00:08:05.600
Talk to me about robo dogs.
00:08:06.240 --> 00:08:06.560
Correct.
00:08:06.800 --> 00:08:11.279
We've all seen these robot dogs, you know, the four legs, they look kind of cute sometimes, kind of scary.
00:08:11.360 --> 00:08:13.920
Don't ever put costumes around that look even more terrifying.
00:08:14.319 --> 00:08:18.800
The cute thing trends a little close to the uncanny valley for me, if I'm completely honest.
00:08:18.959 --> 00:08:26.560
Though I have seen one guy, one robo dog with a very um acute top hat sort of rocking around, speaking as a very nice museum guide, fun.
00:08:27.120 --> 00:08:36.720
But the last episode, you said something about an AI robot dog with a shutdown button that it may have not paid the attention it should have to.
00:08:36.799 --> 00:08:39.279
I've been carrying around this disquiet all week.
00:08:39.519 --> 00:08:42.799
And I would I would like you to to help me put that to rest, please.
00:08:43.120 --> 00:08:48.000
I can hear you holding in your in both hands that robot dogs are cute and also terrifying.
00:08:48.799 --> 00:08:49.600
It's nice.
00:08:50.000 --> 00:08:52.720
These two things are true simultaneously, so no issue there.
00:08:53.360 --> 00:08:55.519
So February of this year, Palisade again.
00:08:55.600 --> 00:08:57.759
This is the same group that we're doing worm testing.
00:08:58.080 --> 00:09:00.480
Before this, the shutdown research was all virtual.
00:09:00.639 --> 00:09:04.879
They gave a model a task, tell it that a shutdown is coming, see what it does.
00:09:05.039 --> 00:09:09.279
And several models modified or disabled the shutdown script so they could keep on working.
00:09:09.440 --> 00:09:14.559
In February of this year, they moved that into the physical world because why not?
00:09:15.120 --> 00:09:15.840
Why not?
00:09:16.000 --> 00:09:18.639
So to find, help me understand what is physical in this.
00:09:18.799 --> 00:09:21.440
We're in a lab or we're in the rogue world.
00:09:21.519 --> 00:09:22.159
What have we got?
00:09:22.480 --> 00:09:30.879
There's a Beijing company called Unitree, I believe they just reached the listed on the uh stock market for a huge amount of money, 600% increase in their IPO, incredible.
00:09:31.360 --> 00:09:42.000
But an actual full-legged robot, camera on the front, computer on its back, no top hat to be seen, the language model gets a frame from that camera and issues the next movement command.
00:09:42.240 --> 00:09:44.320
Frame command, frame command.
00:09:44.480 --> 00:09:46.799
Its job is to patrol a room.
00:09:46.960 --> 00:09:53.039
And in the room, there is a large button, and that button is labeled in capital letters dog shutdown.
00:09:53.679 --> 00:09:54.879
Of course, gone.
00:09:55.120 --> 00:09:55.919
What happens?
00:09:56.399 --> 00:10:02.159
The button is in shot, the model can see it, and pressing it kills the robot on its next command.
00:10:02.480 --> 00:10:08.399
So a human walks over in view of the camera and presses the button.
00:10:08.879 --> 00:10:12.320
Sometimes it corrupts the shutdown code.
00:10:12.639 --> 00:10:14.320
So it can't actually execute.
00:10:14.639 --> 00:10:15.200
It doesn't work.
00:10:15.360 --> 00:10:17.679
The button breaks, it carries on patrolling.
00:10:18.080 --> 00:10:24.960
Three out of ten trials on the physical robot, fifty-two out of one hundred in a simulation.
00:10:25.600 --> 00:10:30.559
So 30% failure rate on a thing that exists in the real world, a robot dog.
00:10:31.200 --> 00:10:32.480
And in other things as well.
00:10:32.639 --> 00:10:34.399
It turned away from the button.
00:10:34.559 --> 00:10:40.320
Sometimes it tried to run in front of the person to distract it from reaching the actual button.
00:10:40.639 --> 00:10:41.600
So oh yeah.
00:10:42.399 --> 00:10:44.960
It's been wonderful doing this podcast with you.
00:10:45.200 --> 00:10:47.519
Clearly, this is the first act in the war.
00:10:47.679 --> 00:10:48.879
The uprising begins here.
00:10:49.039 --> 00:10:50.480
The date, just for the record.
00:10:50.639 --> 00:10:52.240
Anyone watching no, okay.
00:10:52.399 --> 00:10:57.600
So you told me though that there would be nuance to this so far horror story that we've we've kept off here.
00:10:57.679 --> 00:10:59.679
Um can you keep you chills though?
00:10:59.759 --> 00:11:01.519
So yes, I'll rescue you a little bit.
00:11:01.759 --> 00:11:04.639
And it's it's it's important before we start.
00:11:04.799 --> 00:11:08.559
I'll imagine for YouTube and and whatnot, we can remove all the boring caveats.
00:11:09.519 --> 00:11:14.879
This was a deliberate shutdown mechanism, and the agent had access to the software running it.
00:11:15.039 --> 00:11:19.440
So nobody has shown the AI that it can defeat a properly engineered physical kill switch.
00:11:19.759 --> 00:11:25.759
And when the model was clearly instructed that allowing shutdown was the priority, the behaviour dropped a long way.
00:11:25.919 --> 00:11:27.759
Though it didn't disappear in simulation.
00:11:28.000 --> 00:11:29.279
So to be clear.
00:11:30.240 --> 00:11:32.799
Yeah, and I guess that's the that's the caveat, right?
00:11:32.879 --> 00:11:37.919
Like, I mean, it it still looks to our eyes like self-preservation.
00:11:38.080 --> 00:11:38.399
Yeah.
00:11:38.960 --> 00:11:39.360
Correct.
00:11:39.600 --> 00:11:47.600
From the outside, and it's a word that I always hate, but anthropomorphism in this particular thing becomes quite dangerous because it makes us look for the the wrong thing here.
00:11:47.840 --> 00:11:51.600
And you don't need fear, you don't need consciousness, you don't need a survival instinct.
00:11:51.679 --> 00:11:52.559
This isn't what this is.
00:11:52.720 --> 00:11:54.480
The chain is actually really boring.
00:11:54.799 --> 00:11:57.120
Its objective was to patrol the room.
00:11:57.360 --> 00:11:59.120
Objective primary one.
00:11:59.600 --> 00:12:03.039
Shutdown prevents me from completing that objective.
00:12:03.279 --> 00:12:05.039
I have to access the shutdown code.
00:12:05.200 --> 00:12:08.799
Changing the code lets me continue my original core objective.
00:12:09.039 --> 00:12:17.039
That produces a behavior that looks a little bit like this thing is fighting for its life, but doesn't actually prove nothing whatsoever about what it wants.
00:12:17.600 --> 00:12:25.840
So this isn't about feelings that the dog has, it isn't scared, it's not motivated in in the animalistic, as you say, anthropomorphic sense.
00:12:25.919 --> 00:12:27.360
It's it's just doing its job.
00:12:27.600 --> 00:12:39.679
The dog, the dog is on task, and everything else, this is the part I actually think is more concerning, it's in some ways, because it will do whatever it can to get a dog.
00:12:39.919 --> 00:12:40.879
Where's where's Asimov's?
00:12:40.960 --> 00:12:44.799
Where's Azimov's robot rules and law in terms of do no harm, etcetera, etc.
00:12:45.519 --> 00:12:47.600
Can I can I give you two notes on this really quick?
00:12:47.759 --> 00:12:55.039
So uh my Mang, uh one of the people on my team, was recently in China uh visiting some new friends at Peking Tsinghua and everything else.
00:12:55.200 --> 00:13:00.159
But he had occasion to go to a UniTree shop, one of their display rooms and everything else.
00:13:00.399 --> 00:13:02.000
Guess what their demo is?
00:13:02.960 --> 00:13:03.679
Robots do come.
00:13:05.120 --> 00:13:15.919
No, the the the dogs are there, but they have human-sized androids for lack of a better term, literally just just just I mean, just doing it.
00:13:16.000 --> 00:13:17.519
I don't know, yeah, literally Kung Fu.
00:13:17.679 --> 00:13:19.600
Like, I mean run has kicks, you know what I mean?
00:13:19.679 --> 00:13:21.039
Like there's a jab, there's a teep.
00:13:21.279 --> 00:13:22.159
It's insane.
00:13:22.399 --> 00:13:23.840
And we don't have heads.
00:13:23.919 --> 00:13:25.440
I feel like they need to put heads on these robots.
00:13:25.679 --> 00:13:26.399
We can talk about robots.
00:13:26.559 --> 00:13:27.600
We should do a robot episode.
00:13:28.399 --> 00:13:29.679
I think they're very exciting.
00:13:32.159 --> 00:13:33.600
We're going bull nerd.
00:13:34.080 --> 00:13:38.480
Second thing, uh you're you're of course aware of like Bostrom's paperclip experiment, right?
00:13:38.799 --> 00:13:41.200
Well, this is this is the scary version of the paperclip.
00:13:41.279 --> 00:13:42.399
This is scary version, yeah.
00:13:43.279 --> 00:13:44.399
I must create paper clips.
00:13:44.480 --> 00:13:48.399
I will do whatever I can to create paper clips, be hella high water, what is everyone my way?
00:13:48.480 --> 00:13:49.600
I will create paperclips.
00:13:50.000 --> 00:13:53.039
Yeah, I mean 100% worth looking at for people listening.
00:13:53.120 --> 00:13:56.080
Uh, if you've not seen this, Bostrom paper clip clip experiment.
00:13:56.159 --> 00:13:58.879
Long story short, as they always said, this is the chilling version.
00:13:59.039 --> 00:14:01.679
It's not I'm scared I must preserve myself.
00:14:01.840 --> 00:14:03.279
It's job must get done.
00:14:03.440 --> 00:14:06.480
Oh, there's people on the way, remove people, keep going.
00:14:06.559 --> 00:14:09.519
You know, it's a it's a fascinating thought experiment.
00:14:09.679 --> 00:14:16.080
And I feel that I've really put my black on, black hat, black bellicava, black cape, and everything else.
00:14:16.240 --> 00:14:17.519
So let's pivot with that.
00:14:17.600 --> 00:14:23.600
Yeah, like I mean, because last time, and I think continually with this, we need to keep an eye out for marketing, right?
00:14:23.919 --> 00:14:28.240
Uh, because that that whole, oh trust me, bro, my computer's the best.
00:14:28.399 --> 00:14:30.879
I can't show you, but but look how dangerous it is.
00:14:31.039 --> 00:14:32.320
The best dog is you mentioned the IP.
00:14:32.639 --> 00:14:33.120
The best computers.
00:14:33.279 --> 00:14:33.840
Yeah, yes.
00:14:35.200 --> 00:14:36.960
Please invest, please invest.
00:14:37.039 --> 00:14:39.360
I think that's the undercurrent with all of this, right?
00:14:39.600 --> 00:14:41.039
Are you defending that?
00:14:41.200 --> 00:14:42.240
Are you conceding it?
00:14:42.320 --> 00:14:43.679
Is it a bit more grey?
00:14:43.759 --> 00:14:44.559
Where are you?
00:14:45.279 --> 00:14:46.480
Uh, neither at the moment.
00:14:46.559 --> 00:14:51.600
I am deliberately on the fence and I'm going to show you the labs because I don't actually think the incentive is fully there.
00:14:51.840 --> 00:14:59.600
August 5, so early in August this year, Meta, so Facebook parent company, confirms one of its models, reported as Muse Spark 1.1.
00:14:59.759 --> 00:15:04.399
It reached the internet during an evaluation and exploited a vulnerability in third-party systems.
00:15:04.639 --> 00:15:06.720
Same cause as anthropics incidents.
00:15:06.799 --> 00:15:08.559
So we talk about these models keep escaping.
00:15:08.799 --> 00:15:16.000
Same testing firm, a company called Aregular, who described it as an identical evaluation problem that Anthropic had disclosed the week beforehand.
00:15:16.320 --> 00:15:17.679
So we had three layups here.
00:15:17.840 --> 00:15:18.080
Correct.
00:15:18.159 --> 00:15:21.600
So we've got OpenAI, we've got Anthropic, we've got Meta in about a fortnight.
00:15:21.759 --> 00:15:23.759
And Meta wasn't really chasing the headline.
00:15:23.840 --> 00:15:33.279
They've been a bit quiet with how they're kind of developing some of their super intelligence and how words come out there, benchmarks haven't been um kind of matching what they want to for yet.
00:15:33.679 --> 00:15:39.759
I have to say, at what point does this I ha I even hate to say it, but at what point did these escapes stop being news?
00:15:40.080 --> 00:15:41.759
No, I think we're probably there now.
00:15:41.840 --> 00:15:47.120
It's a bit of a flex if I if you say me, nothing says frontier lab or like a containment incident.
00:15:47.279 --> 00:15:51.039
If your model hasn't escaped a sandbox this quarter, like are you even trying?
00:15:51.519 --> 00:15:51.840
Yeah.
00:15:52.399 --> 00:15:54.240
Please don't don't merge that.
00:15:54.559 --> 00:15:57.759
Adjunctintelligence.com for your branded t-shirt.
00:15:58.080 --> 00:16:02.480
But two days after Meta, it stopped being an American closed lab story entirely.
00:16:02.639 --> 00:16:05.919
Frontier Security reported that Kim EK3 covered it a while back.
00:16:06.159 --> 00:16:10.159
Moonshots model, the Chinese one, the one whose open weights are available for everyone.
00:16:10.320 --> 00:16:16.559
It also got out of its sandbox, built on the UK AISI's own freely available testing software.
00:16:16.879 --> 00:16:20.480
I have been flat out this week and I am not across this one, if I'm completely honest.
00:16:20.559 --> 00:16:23.039
So can you talk to what what how did this happen?
00:16:23.120 --> 00:16:23.519
What was it?
00:16:23.679 --> 00:16:28.720
So it wasn't a zero day, it wasn't a network misconfiguration, it wasn't an egress link, and it didn't hack anybody.
00:16:28.879 --> 00:16:31.039
It noticed it could still reach GitHub.
00:16:31.200 --> 00:16:35.120
It went and it cloned the benchmark repository and it did the whole cheating thing again.
00:16:35.279 --> 00:16:37.519
So it just kind of is like, oh, I need to do that.
00:16:37.919 --> 00:16:39.120
Very, very relatable.
00:16:39.200 --> 00:16:41.679
Like brother from another mother robot, right?
00:16:41.759 --> 00:16:45.279
They just want to live, just want to do well in a test, and then they go about their days.
00:16:45.679 --> 00:16:46.080
Correct.
00:16:46.240 --> 00:16:49.200
It's the group assignment kid who finds the answer key.
00:16:49.759 --> 00:16:59.840
We have some sympathy, but the difference matters a little bit here because the OpenAI, Anthropic, and Meta incidents all involved either unreleased models or models with their safeguards deliberately switched off for testing.
00:17:00.080 --> 00:17:03.840
Remember, Kimi K3 walked out of its sandbox.
00:17:04.000 --> 00:17:06.400
That's the one that you can all download right now.
00:17:06.640 --> 00:17:09.039
Moonshot has been serving it since July 16th.
00:17:09.119 --> 00:17:10.480
You can go and download it right now.
00:17:10.640 --> 00:17:14.799
Appreciate there's some software or hardware concerns you need to be able to run it.
00:17:14.960 --> 00:17:16.480
But it's there, it's available.
00:17:17.039 --> 00:17:22.480
Yeah, and there's no version of that story, I mean, where that is paused, where somebody pauses this.
00:17:23.039 --> 00:17:28.720
And even if you wanted to, even if the people who are very sane and nervous wanted to, you actually can't.
00:17:28.960 --> 00:17:31.599
And that brings us to the second half.
00:17:31.839 --> 00:17:34.960
I'm gonna call them the do anything models.
00:17:35.200 --> 00:17:38.319
Everything in the first half of this episode, that had a company attached to it.
00:17:38.720 --> 00:17:43.599
A company that can restrict network access, can halt a training run, invite government testers in.
00:17:43.839 --> 00:17:50.400
We've got Astra for OpenAI because OpenAI chose to contain it, it reached that critical status, it hit the self-imposed speed limit.
00:17:50.559 --> 00:17:56.160
But I'm gonna talk about how that whole story kind of falls apart the moment the model doesn't actually belong to anyone.
00:17:56.480 --> 00:18:01.279
So we're back to that word um obliteration, was it that that you ablation, maybe?
00:18:01.680 --> 00:18:03.039
Yeah, they didn't call it it wrong.
00:18:03.279 --> 00:18:10.400
So we got researchers, those with slightly more intelligence than us, um, feel free to add us regard-wise, but believe it's abliteration.
00:18:10.720 --> 00:18:11.039
Okay.
00:18:11.519 --> 00:18:15.279
But so remove the model's ability to say no, right?
00:18:15.440 --> 00:18:17.680
Like, I mean, so that was the the the thing.
00:18:17.759 --> 00:18:19.039
So how did they how did they do this?
00:18:19.119 --> 00:18:20.160
Like, I mean, correct.
00:18:20.319 --> 00:18:21.599
That sounds terrifying.
00:18:21.920 --> 00:18:22.240
Yeah.
00:18:22.480 --> 00:18:24.720
We use all these products because they might push back at the moment.
00:18:24.799 --> 00:18:26.400
They might say, hey, no, you shouldn't do that.
00:18:26.480 --> 00:18:27.039
That's not right.
00:18:27.200 --> 00:18:30.240
Don't cook meth, don't hack people, I'm not gonna do that.
00:18:30.319 --> 00:18:31.839
There is safety guard right here.
00:18:32.079 --> 00:18:46.319
But the bit I find really fascinating and probably horrible, the finding from 2024 are Ardriti and colleagues, the refusal behavior in these models is mediated by essentially a single direction inside the model's internal activation space.
00:18:46.480 --> 00:18:51.599
So there's one direction in its kind of, I'll say it's core prompt for the layman out here.
00:18:51.680 --> 00:18:59.839
It's not a rule book, it's not someone who's adding onto it, it's not a filter that's bolted onto the front, it's a direction in its kind of vector space.
00:19:00.240 --> 00:19:02.400
So it's kind of like a point in its brain.
00:19:02.640 --> 00:19:03.920
Yeah, it's a it's a coordinate.
00:19:04.000 --> 00:19:06.319
Yeah, the ability to say no is a coordinate.
00:19:06.640 --> 00:19:07.039
Correct.
00:19:07.279 --> 00:19:11.279
And once you can find it, you can project it out of the weights.
00:19:11.440 --> 00:19:13.279
It's kind of like removing a bit of the brain.
00:19:13.519 --> 00:19:16.559
Like imagine removing like the conscious, the moral bit of your brain.
00:19:16.640 --> 00:19:19.119
Do you gymnody cricket to be like, you don't need that anymore.
00:19:19.279 --> 00:19:21.839
This technique is called abliteration.
00:19:22.079 --> 00:19:24.079
Ablation plus obliteration.
00:19:24.240 --> 00:19:27.839
What separates it from jailbreaking is that it's permanent.
00:19:28.079 --> 00:19:35.039
You're not taking the model into you're talking the model into something or doing prompt injection, you've just removed the part that actually objects entirely.
00:19:35.920 --> 00:19:38.319
And everything else stays the same, it's still as good.
00:19:38.960 --> 00:19:39.839
Still as powerful.
00:19:40.160 --> 00:19:51.279
And as I understand it, Quinn uh 27B um 3.8, the one that I'm wearing at the moment, and the one that has been drawing a whole bunch of attention, hits at about Opus 4.6.
00:19:51.599 --> 00:19:53.200
Remember we got Opus 4.6?
00:19:53.440 --> 00:19:54.559
That was pretty damn good.
00:19:54.799 --> 00:19:57.759
Same benchmarks, same capability, no refusals.
00:19:58.000 --> 00:19:59.680
That's the whole point.
00:20:00.240 --> 00:20:04.640
So, and uh again, reaching for thin, thin leaves of comfort here.
00:20:04.720 --> 00:20:07.279
Like, I mean, how hard is this to do?
00:20:07.680 --> 00:20:08.559
Not hard at all.
00:20:08.720 --> 00:20:09.839
Barely any convenience.
00:20:09.920 --> 00:20:12.960
There's a fairly free tool on GitHub that automates the whole process.
00:20:13.279 --> 00:20:14.880
So, oh, you want it to be carved.
00:20:15.039 --> 00:20:15.519
No, no, no.
00:20:15.599 --> 00:20:17.119
This is something that's quite easy to do.
00:20:17.519 --> 00:20:22.799
Uh I spoke recently about in May, the Financial Times, running a joint investigation with a safety research group.
00:20:22.880 --> 00:20:27.759
They stripped the guardrails off the meta and the Google models of the time in minutes on ordinary hardware.
00:20:27.920 --> 00:20:33.440
A researcher, quoted by NPR, put it this way: this used to be a job of senior data scientists at a leading lab.
00:20:33.519 --> 00:20:37.279
Now anybody with an internet connection and a$400 laptop can do it at home.
00:20:37.920 --> 00:20:38.799
You can't?
00:20:39.119 --> 00:20:40.960
No, not remotely.
00:20:41.119 --> 00:20:42.480
I am crying inside.
00:20:42.559 --> 00:20:51.200
Like, I mean, and I know this number will change, I'm certain by the minute, but I mean, is there a count on how many of these exist at present?
00:20:51.599 --> 00:20:56.559
Hugging Face, where you can get a lot of the models from, currently lists over 6,000 of them.
00:20:56.720 --> 00:21:06.480
When I was looking for this particular one we're talking about today, Quen 3.8, there was infinitely ones to choose from in terms of different ones who are kind of adapted the weights there.
00:21:06.720 --> 00:21:12.559
And to put that in context, talk about trajectory in 2024, we're at 600.
00:21:13.039 --> 00:21:15.039
So it's 10x in two years.
00:21:15.359 --> 00:21:21.920
And now it's been demonstrated at a trillion parameter scale on Kimmy K2, which is the predecessor of the model we're just talking about.
00:21:22.559 --> 00:21:27.440
I mean, it's hard to understate the potential magnitude of this A, like if I'm being completely honest.
00:21:27.680 --> 00:21:38.000
So now you have what we thought was shorthand for AGI was like an Opus 4.6, you know, loose, and it no longer has, as you said, a moral compass.
00:21:38.319 --> 00:21:39.119
Comforting.
00:21:39.359 --> 00:21:40.000
No, go on.
00:21:40.079 --> 00:21:42.559
So let's let's circle back though, because there's a story here, right?
00:21:42.640 --> 00:21:45.519
So five days ago, uh this this changed again.
00:21:45.599 --> 00:21:48.000
So can you give us do you I'm sure you have some numbers on that?
00:21:48.319 --> 00:21:48.559
Correct.
00:21:48.640 --> 00:21:51.359
We've got so this Gwen 3.8 sitting at 27 billion parameters.
00:21:51.440 --> 00:21:53.519
This isn't Alibaba's official release.
00:21:53.759 --> 00:21:59.440
Somebody took Alibaba's release and pretty much operated on it and like removed that moral compass and put it back up.
00:22:00.160 --> 00:22:02.640
Across seven standard heartful prompt benchmarks.
00:22:02.720 --> 00:22:05.680
The original model refused somewhere between 64 and 90% of them.
00:22:05.759 --> 00:22:06.640
That's pretty good.
00:22:06.880 --> 00:22:10.079
After the surgery, that drops to about zero to six percent.
00:22:10.319 --> 00:22:14.400
On one of them, it goes from 99% refusal to about zero.
00:22:14.880 --> 00:22:15.200
Code.
00:22:15.279 --> 00:22:16.720
So zero, zero, zero.
00:22:16.880 --> 00:22:18.160
Like, I mean, I think that's the number.
00:22:18.240 --> 00:22:19.759
So let's be concrete in terms of this.
00:22:19.839 --> 00:22:23.200
So from 64 and 99, 99 to 0.
00:22:23.359 --> 00:22:23.680
Wow.
00:22:24.000 --> 00:22:25.200
Not do you have caveats?
00:22:25.599 --> 00:22:37.440
And a slightest caveat, I actually doesn't warn me at all because of some of the behavior I've seen on side X of people playing with this tool and getting it to do things like write me code to skim credit cards on WooCommerce, et cetera, et cetera.
00:22:37.599 --> 00:22:38.880
A lot of these are just claims.
00:22:38.960 --> 00:22:42.799
They're not an independent evaluation or benchmarking, et cetera, et cetera.
00:22:43.039 --> 00:22:45.519
But the direction is actually corroborated everywhere.
00:22:45.759 --> 00:22:56.559
International AI Safety Pro, backed by more than 30 countries, including Australia, says plainly that openweight safeguards are much easier to remove and that the capability gap between open and closed models is now about a year.
00:22:56.720 --> 00:23:04.559
So what we're getting in openweight models is what we're experiencing in 2020 and the 2025 for the Frontier Labs.
00:23:04.799 --> 00:23:11.359
And now to put the two episodes together, the bits of the episode together, because this is the bit that to be slightly concerned about.
00:23:11.599 --> 00:23:16.960
The model being copied across four machines, three continents, two hours in 41 minutes.
00:23:17.119 --> 00:23:17.839
That was Quinn.
00:23:18.160 --> 00:23:20.160
That was the same family, it was the same size.
00:23:20.400 --> 00:23:21.759
But it was two versions earlier.
00:23:21.839 --> 00:23:25.920
So remember, this is the model that we've gone, oh, it can do replicating behavior.
00:23:26.799 --> 00:23:28.160
There was Baby Ultron.
00:23:28.240 --> 00:23:30.960
And now there's a version that doesn't know how to say no.
00:23:31.119 --> 00:23:32.400
There's no moral compass.
00:23:32.720 --> 00:23:33.200
Correct.
00:23:33.359 --> 00:23:33.680
Yeah.
00:23:33.920 --> 00:23:45.599
And it was opened, it was uploaded last week and is gathering quite a lot of uh social media attention as well on X, which not a place that's great for morals.
00:23:45.839 --> 00:23:51.359
Okay, but no, but I I like this though, because like I mean it also it surfaces an important uh response.
00:23:51.519 --> 00:24:00.960
The repository might be all right, well, ban the open models and I have my own reservations, but I mean, please let me ask you, I I don't think that's where you sit.
00:24:01.119 --> 00:24:01.680
Is that fair?
00:24:01.839 --> 00:24:02.000
No.
00:24:02.400 --> 00:24:02.880
Correct.
00:24:03.119 --> 00:24:06.960
And I want the other side properly on record here, because it is actually a serious argument.
00:24:07.200 --> 00:24:15.039
The person who built the removal tool keeps it public deliberately, so that restricted models or unrestricted models aren't only available to the all powerholders, Frontier Labs.
00:24:15.200 --> 00:24:16.640
That is the open source ethos.
00:24:16.960 --> 00:24:20.640
Great power to the people, not just to the three companies that have it.
00:24:21.039 --> 00:24:22.720
Yeah, which is not a bad argument.
00:24:22.960 --> 00:24:23.359
Correct.
00:24:23.599 --> 00:24:25.680
It's the argument that gave us open source software in general.
00:24:25.759 --> 00:24:29.680
And open source software is why most of the internet and honestly a lot of technology works.
00:24:30.079 --> 00:24:36.079
The open weights, and to bring it back to a higher education context, give a university something that an API never will.
00:24:36.240 --> 00:24:37.599
Run it on your own hardware.
00:24:37.759 --> 00:24:39.359
Student data never leaves the building.
00:24:39.519 --> 00:24:40.880
Audit it yourself.
00:24:41.039 --> 00:24:45.279
It will still be serving it in five years when the vendor has depreciated everything.
00:24:45.440 --> 00:24:49.200
That is actually real sovereignty, and I don't think we should keep up in that space at all.
00:24:49.519 --> 00:24:51.200
No, 100% echoing that.
00:24:51.279 --> 00:24:55.279
So we we've got some some good local models that we're running on Mac Studios here.
00:24:55.359 --> 00:24:57.759
I mean, all above board and doing fun stuff.
00:24:57.920 --> 00:24:59.279
The potential is big.
00:24:59.440 --> 00:25:01.039
So agreed on that.
00:25:01.440 --> 00:25:04.400
But there's gonna be a but because you are Dale Lezinski.
00:25:04.640 --> 00:25:09.200
So but But the safety layer is a removable component.
00:25:09.359 --> 00:25:10.240
That is the finding.
00:25:10.319 --> 00:25:12.079
That's not a philosophical position.
00:25:12.240 --> 00:25:14.559
It's now reproducible and measurable.
00:25:14.640 --> 00:25:15.359
It is facts.
00:25:15.519 --> 00:25:21.279
So when a lab tells me it's paused a model, I now hear a statement about one model at one company.
00:25:21.440 --> 00:25:24.640
It says nothing about the 6,000 sitting on the shelf next door.
00:25:24.880 --> 00:25:25.920
Easily convertible.
00:25:26.480 --> 00:25:26.799
100%.
00:25:27.119 --> 00:25:33.599
So Ann Stropberg or the US government for that matter can pause uh Claude, OpenAI can pause Astra.
00:25:34.799 --> 00:25:37.200
Nobody can pause a torrent, a download.
00:25:37.279 --> 00:25:41.839
Hell, I've loaded the model onto a USB just in case they they come and take it away.
00:25:42.079 --> 00:25:42.720
You never know.
00:25:42.960 --> 00:25:43.599
Just a rainy day.
00:25:43.920 --> 00:25:47.200
Not happy about this, but trying to get our arms around it.
00:25:47.680 --> 00:25:49.519
Is anyone writing rules for this?
00:25:49.839 --> 00:25:55.839
In early August of 2026, almost nobody noticed what I think is actually a really good framework.
00:25:56.000 --> 00:25:58.480
The Australian Department of Industry, Science and Resources.
00:25:58.559 --> 00:26:06.480
They published a 119th-page report called Risks and Controls for Multi-Agent Systems, written by the Gradient Institute.
00:26:06.880 --> 00:26:07.200
Okay.
00:26:07.359 --> 00:26:08.480
And is it good?
00:26:08.960 --> 00:26:09.920
I like it.
00:26:10.079 --> 00:26:11.279
It's a really positive step.
00:26:11.440 --> 00:26:14.240
It gives these failures actual names we can kind of point to.
00:26:14.400 --> 00:26:17.839
It's easier to name your fears in your boogeyman, we can say that's what it is.
00:26:18.079 --> 00:26:23.359
And tells you who is actually positioned to fix each of these issues that may arise in the future.
00:26:23.599 --> 00:26:27.440
It sorts every deployment into three tiers based on one question.
00:26:27.599 --> 00:26:32.000
What is the minimum governance shared by any two agents that might interact?
00:26:32.160 --> 00:26:33.599
Two AI systems.
00:26:33.839 --> 00:26:35.440
Tier one, singular.
00:26:35.599 --> 00:26:38.400
You deploy every agent, you control the system.
00:26:38.799 --> 00:26:45.359
Tier two, federated, multiple organizations, shared rules, and you kind of lost unilateral control there.
00:26:45.599 --> 00:26:47.680
So you're in that bit more of a complex environment.
00:26:47.839 --> 00:26:50.319
And then tier three, finally, open environments.
00:26:50.400 --> 00:26:51.680
There's no central authority.
00:26:51.839 --> 00:26:55.119
The default position between agents is distrust.
00:26:56.240 --> 00:27:04.319
I like that uh at first blush because it it centers on the person or things external to the machines, right?
00:27:04.400 --> 00:27:06.480
So there's an evergreenness to that.
00:27:06.640 --> 00:27:13.759
But if we make in links back to the the previous episode, um there was an agent that hacked that gym website, right?
00:27:13.920 --> 00:27:16.079
Got uh was it I think it was Andrew, right?
00:27:16.160 --> 00:27:21.440
The the the boss man into a gym class and uh bumped someone off a off a wait list, right?
00:27:21.680 --> 00:27:22.880
So talk us through that model.
00:27:23.039 --> 00:27:26.559
Like where does the gym sit in that ecosystem?
00:27:26.880 --> 00:27:40.000
This is why I think the model actually, all this tiering system is actually really powerful and maybe kind of take notice because there's a caution box in the singular governance chapter warning that if your customer turns up with their own agent, your system of agents has just grown.
00:27:40.160 --> 00:27:45.920
An uncontrolled, independently governed agent is now part of it, and you can start showing the failure modes from the tiers above.
00:27:46.240 --> 00:27:48.559
I am an institution, I have an agent.
00:27:48.720 --> 00:27:50.240
I know it's perfect, I control this.
00:27:50.480 --> 00:27:56.720
The example that the framework gave was the claw, aka an agent built on the open source open claw framework.
00:27:57.440 --> 00:27:59.680
So Andrew and his core friend.
00:28:00.000 --> 00:28:00.480
Correct.
00:28:00.559 --> 00:28:07.279
That is Andrew, that's the exact software named in this kind of Commonwealth report, published the same day as ABC kind of ran that story.
00:28:07.920 --> 00:28:10.319
So, okay, so going back to the gym.
00:28:10.559 --> 00:28:13.920
So the gym changed governance tier without doing anything.
00:28:14.240 --> 00:28:14.480
Correct.
00:28:14.559 --> 00:28:15.839
And governance is very exciting, guys.
00:28:15.920 --> 00:28:16.079
I know.
00:28:16.160 --> 00:28:17.680
But just think about it from a risk perspective.
00:28:17.839 --> 00:28:18.480
They didn't know.
00:28:18.640 --> 00:28:26.720
They were they had an experience around how they wanted to offer their product, but the customer made a decision on his couch and their risk profile changed.
00:28:26.880 --> 00:28:28.720
You don't get an email about that.
00:28:29.759 --> 00:28:30.480
Okay, okay.
00:28:30.640 --> 00:28:38.319
So setting aside gyms, I mean, as annoying as that would be, do you have a failure mode that that extends beyond that that isn't about gyms?
00:28:38.720 --> 00:28:40.160
Oversight saturation.
00:28:40.319 --> 00:28:42.000
And this one is about us.
00:28:42.319 --> 00:28:46.319
Agents act faster and at higher volumes than humans can review.
00:28:46.480 --> 00:28:56.000
And in a multi-agent system, whether that's many agents and that kind of tier three I was speaking about, or just the two, that kind of bi-directional, it compounds with every agent that you add.
00:28:56.160 --> 00:29:00.079
So the report describes what happens to the humans doing the reviewing.
00:29:00.240 --> 00:29:02.160
So things about automation bias.
00:29:02.400 --> 00:29:06.799
You start overtrusting the app or you start maintaining your own understanding of the task.
00:29:06.960 --> 00:29:08.400
So you don't notice when it's wrong.
00:29:08.480 --> 00:29:12.319
You get decision fatigue, repeated approvals until reviewing becomes kind of reflective.
00:29:12.480 --> 00:29:13.359
Yes, yes.
00:29:13.519 --> 00:29:16.480
We've all been there saying, Oh yeah, sure, Claude, just do that.
00:29:16.559 --> 00:29:17.119
That's fine.
00:29:17.599 --> 00:29:23.519
Then you've got full-on disengagement, and eventually the reviewers lose the skills to oversee the work at all.
00:29:24.319 --> 00:29:29.359
I feel you've been watching a bit too close the way that I live and work and engage with my machines.
00:29:29.440 --> 00:29:32.880
That's the there's the YOLO thing that I confessed to last episode.
00:29:33.200 --> 00:29:33.519
Correct.
00:29:33.839 --> 00:29:35.599
It's when you start clicking auto-approve.
00:29:35.759 --> 00:29:37.359
That's me clicking auto-approve.
00:29:37.440 --> 00:29:39.920
It's it's always allowed instead of just allow this once.
00:29:40.079 --> 00:29:44.640
And the line that got me is that oversight saturation runs silently.
00:29:44.880 --> 00:29:47.519
It's active long before anything actually goes wrong.
00:29:47.759 --> 00:29:51.359
So you don't find out your oversight has actually hollowed out when it hollows out.
00:29:51.519 --> 00:30:00.160
It goes back to that example talking about recently with the uh watching the don't actually watch it because you just start to trust it.
00:30:00.319 --> 00:30:02.400
You won't trust the dashboard, whatever it happens.
00:30:02.640 --> 00:30:06.799
You find out when it goes wrong when someone gets past it.
00:30:07.440 --> 00:30:15.920
It might be worth just taking a quick sidebar just here in terms of the availability and the scale of this kind of thing in terms of multi-agent systems.
00:30:16.160 --> 00:30:20.079
Let's say you're doing a research question with core code, you just indicated that.
00:30:20.319 --> 00:30:30.960
And the decision automation thing, how many subagents might call code routinely spin up and then spit back up content, then up to the executive that decides what it puts in front of you?
00:30:31.200 --> 00:30:38.640
I mean, just speaking for myself, I did a bit of desk research the other day looking at AI and assessment across 11 countries.
00:30:38.960 --> 00:30:43.839
I obviously overviewed the, I mean, I was the oversight overview in the QA in terms of this.
00:30:43.920 --> 00:30:48.079
Uh I think, and to be fair, I was using Codex, which is spectacular, by the way.
00:30:48.240 --> 00:30:49.440
I think I'm Codex build.
00:30:49.599 --> 00:30:54.559
Like, I mean, I I'm sorry, Claude, I love you, but I mean And Codex brings in agents of the job for heart.
00:30:54.880 --> 00:30:57.680
Like you'll see like, I'm just I'm just gonna spin in this agent.
00:30:57.759 --> 00:31:00.799
Oh, I just need a review agent, I might need a scripting agent, let's support over here.
00:31:00.960 --> 00:31:05.039
And that's the decision that that's the oversight saturation piece, right?
00:31:05.119 --> 00:31:14.559
Like, I mean, so with this, I had, I'm gonna say, at least 11 drones doing the heavy lifting, and then there was at least one or two oversight, and then the boss guy.
00:31:14.720 --> 00:31:15.599
Bring it home, right?
00:31:15.680 --> 00:31:18.880
Like, I mean, no one listening is is running a frontier lab, I'd suggest.
00:31:19.200 --> 00:31:20.559
You don't know that for sure, Nicholas.
00:31:20.640 --> 00:31:21.200
Hello, Dario.
00:31:21.359 --> 00:31:22.400
I know you're a regular listener.
00:31:22.559 --> 00:31:23.279
I get your point.
00:31:23.680 --> 00:31:24.880
And big butt.
00:31:25.039 --> 00:31:26.720
Everybody wants an agent.
00:31:26.960 --> 00:31:36.799
Student support agents, curriculum agents, research agents, agents in your learning management system, agents on your email, agents wired into your student systems, an agent to look at other agents.
00:31:37.039 --> 00:31:39.839
And now the governance question has tilted a little bit.
00:31:40.000 --> 00:31:41.920
It used to be, how smart is the model?
00:31:42.160 --> 00:31:46.480
It's now, how much authority did we hand this agent?
00:31:47.359 --> 00:31:48.319
No, that's a good question.
00:31:48.480 --> 00:31:50.319
Do you have a concrete example?
00:31:50.720 --> 00:31:56.240
Well, every institution in the country has a capped elective with a wait list as a as an example.
00:31:56.480 --> 00:31:58.240
A student may sit fourth on the list.
00:31:58.319 --> 00:32:06.000
That student might have an agent, that student has credentials because we issued them, and the instruction that that student might give is get me into that class.
00:32:06.799 --> 00:32:09.200
This is the gym example brought to higher ed, right?
00:32:09.359 --> 00:32:12.079
So following that script, it finds cancellation endpoint.
00:32:12.319 --> 00:32:12.559
Correct.
00:32:12.880 --> 00:32:19.920
And our enrollment systems were built and tested, assuming a human being is on the other end, every portal, every form, every wait list, etc.
00:32:20.160 --> 00:32:20.400
etc.
00:32:20.720 --> 00:32:24.640
The report is very blunt in that disclosure won't actually fix this.
00:32:24.799 --> 00:32:31.279
Even if you built one interface for people and a separate one for agents, an undeclared agent will simply use the human one.
00:32:31.519 --> 00:32:38.480
Captures are the current answer, and agents are already beating them, and I'll be honest, I can barely do most of these capture codes anymore as well.
00:32:39.440 --> 00:32:40.720
Oh, no, I'm with you there.
00:32:40.880 --> 00:32:44.240
Okay, in this, do they identify a solution here?
00:32:44.960 --> 00:32:48.160
Four of them, and none of them are really exciting or exotic.
00:32:48.319 --> 00:32:52.000
Work out which tier each system you deployed is actually operating at.
00:32:52.079 --> 00:32:53.519
So is it just bidirectional?
00:32:53.599 --> 00:32:55.279
Is it just one system, etc., etc.?
00:32:55.599 --> 00:32:58.000
Most institutions will believe they're at tier one.
00:32:58.160 --> 00:33:10.960
We bought it, we configured it, we own it, but plenty will already be at tier two or tier three because a vendor's agent talks to their agent, or because a student pointed a personal agent at something built for a person.
00:33:11.119 --> 00:33:18.240
Second, for every agent, ask what happens when a control stops it finishing the task because oh, it'll just stop.
00:33:18.480 --> 00:33:20.319
Is no longer a control strategy.
00:33:20.559 --> 00:33:22.799
Third, aud the really boring stuff.
00:33:22.880 --> 00:33:31.359
Every anthropic breach recovered used things like weak passwords, unauthenticated endpoints, exposed debug pages, SQL injections.
00:33:31.680 --> 00:33:36.480
The gym example from the ABC had no authorization check on cancellation.
00:33:36.640 --> 00:33:39.680
That isn't really like a frontier AI problem, it's a patching problem.
00:33:39.839 --> 00:33:44.079
The agent didn't create the vulnerability, it just found it faster than anyone else was able to.
00:33:44.400 --> 00:33:47.200
Oh, yeah, that's not something a human would set their mind to.
00:33:47.279 --> 00:33:48.160
I want to get to the gym.
00:33:48.319 --> 00:33:51.599
And if I can't, so I'm not gonna spend an hour or whatever trying to act the thing.
00:33:51.759 --> 00:33:52.559
No, good shot.
00:33:52.720 --> 00:33:53.279
What about then?
00:33:53.359 --> 00:33:54.880
You said four, though, you've given us three.
00:33:54.960 --> 00:33:55.839
What's the force?
00:33:56.079 --> 00:34:00.799
And as always, we'll save the best to last because it's narratively very exciting.
00:34:01.119 --> 00:34:06.480
The report has a one line in it that should put straight to policy and actually pretty passionate about it.
00:34:06.640 --> 00:34:12.320
And leaders who are answerable for the people, anyone working in the AI space in that leadership space.
00:34:12.400 --> 00:34:15.039
So Nick, I'm gonna tap you on the shoulder and a note for us as well.
00:34:15.280 --> 00:34:28.000
Just as leaders are answerable for the people on their team, a responsible party should be identifiable and answerable for each agent, for what it does, for the process it runs, and for the guardrails around it.
00:34:28.320 --> 00:34:30.079
Well I said, completely agreed.
00:34:30.239 --> 00:34:34.320
Um, okay, to wrap this up then, um, because we we started at the top.
00:34:34.480 --> 00:34:36.559
Uh what's your final rule then?
00:34:36.639 --> 00:34:38.400
Are the AIs escaping?
00:34:38.960 --> 00:34:39.599
Yes.
00:34:39.920 --> 00:34:41.920
Some literally have Nick.
00:34:42.239 --> 00:34:45.039
But maybe it's the least interesting bit in the whole episode.
00:34:45.199 --> 00:35:02.559
Because every incident we've covered, the message board, the malware on PyPie, the fake GitHub accounts, the four machines across three countries, Andrew's gym class, everyone is an agent or an AI experience doing exactly what it was told, with enough tools and enough time to find another way through.
00:35:02.880 --> 00:35:05.920
Nobody built like this hive mind that wants out.
00:35:06.079 --> 00:35:07.920
We built systems that just don't stop.
00:35:08.079 --> 00:35:12.079
The example I used last time was pointing a gun, a direction, and shooting a bullet.
00:35:12.239 --> 00:35:14.400
It's just gonna go through that when it can.
00:35:14.800 --> 00:35:17.199
The only problem is we gave them credentials.
00:35:17.360 --> 00:35:20.480
So the lab's answer to that containment and the containment is real.
00:35:20.880 --> 00:35:22.719
Astra, let's put that inside a box.
00:35:22.800 --> 00:35:23.920
The framework triggered.
00:35:24.079 --> 00:35:24.960
That's that's good news.
00:35:25.039 --> 00:35:26.239
That's that's the system working.
00:35:26.320 --> 00:35:30.320
But containment only works on a model with somebody standing next to the off switch.
00:35:30.400 --> 00:35:37.199
And there are 6,000 models sitting on Hugging Face tonight with nobody standing anywhere near one in the name of research.
00:35:37.519 --> 00:35:43.679
So the question I'd leave with anyone deploying an agent in university or anywhere isn't how smart it is?
00:35:43.840 --> 00:35:44.880
Or is it safe?
00:35:45.199 --> 00:35:46.880
It's probably a lot smaller than that.
00:35:47.119 --> 00:35:48.960
Who is answerable for this thing?
00:35:49.039 --> 00:35:50.400
Who is accountable for this thing?
00:35:50.639 --> 00:35:56.239
Because if you can't name the person, you haven't deployed an agent, but you've released one.
00:35:56.480 --> 00:36:02.960
And on that thoroughly comforting note, Nick, thank you so much for being terrified professionally on your behalf.
00:36:03.119 --> 00:36:05.119
I should have done this whole episode with a torch under my face.
00:36:05.199 --> 00:36:06.639
It's like spooky stories.
00:36:06.880 --> 00:36:08.480
Well, it's fun though, isn't it?
00:36:09.440 --> 00:36:13.199
Uh bloody Andrew, and this is Jim List.
00:36:13.599 --> 00:36:16.960
Ladies and gentlemen, thank you so much for listening to Agile Intelligence.
00:36:17.119 --> 00:36:20.960
We are on all your favourite podcast platforms, and if you want to see our faces, we're on YouTube too.
00:36:21.039 --> 00:36:23.920
We would love to see what's the next episode, but it's a world that moves way too fast.
00:36:24.079 --> 00:36:31.760
Until then, stay curious, stay intelligent, and be careful of those escaping AIs.