Sebastian: Welcome to another episode of our new season of Beyond Pipe Coding, partnering with Impala Search.
André Neubauer: The go to tech and executive search agency in Germany.
Sebastian: In this podcast, we explore the transformational change in software engineering and knowledge work in general. I am Sebastian Heidemeyer zu Erpen, CTO at Ecosia.
André Neubauer: And I'm Andre, CTPO at Trusted Shops. Great to have you back. this time there is no guest. this is one of the rare episodes without an interview guest. nevertheless, we want to cover an important topic. we want to take the opportunity to take a sober look at the current AI apocalypse debate.
Sebastian: Exactly. We'll briefly summarize the current situation, assess the actions of the AI labs, the perspective of the quote unquote a a AI skeptics, and discuss the Jacob Coxon claim that there is a a higher than ten percent chance humanity will be extinct by the end of the twenty thirties. we'll also take a shot at predicting what's going to happen, so stay till the end. Let's see if we're good at predicting this.
André Neubauer: Good. Then let's start. where do we want to kick this off, Sebastian?
Sebastian: Yeah, so probably it's good to very briefly, because all all of our listeners have heard about it, right? But summarize the the incident that happened, right? And then take it from there. And for the incident to yeah, pretty much fully understand what happened or why this happened, I think it's important to state a couple of facts about modern frontier models with reasoning capabilities. And number one is that What makes these frontier models so good, specifically for our coding use cases and agentic word in work in general, is their ability to extend the loop, so to not go back to the human every time, right? So to work without steering for longer. And in order to do that, these models need to be able to overcome obstacles. So they need to be able to try different routes. and take certain decisions without always asking the human for decisions. They need to be persistent in that. And yes, we all like these model characteristics when working on software projects. And these characteristics obviously then can create situations where models are actually then overcoming security obstacles, right?
André Neubauer: Hmm. I I I I absolutely agree. And I think we also discussed that when we had that is the episode on Fable. I think we both realized so we both gave it a shot, right? Before it got banned. And we I think our impression was that it was quite ambitious. so it was really really trying to solve issue in in a task. so I think this is one of the things you definitely see when looking back, right? Like some models try to really solve a problem and do not stop before that.
Sebastian: Yes, exactly. So quite a lot like like we discussed with Markus Andreth, to the dismay of him, because he hates this, right? He he actually wants them to produce nice outputs instead of like going lengths and producing very elaborate whatever. anyhow, so now what happened? So in the hugging phase incident, and this is the most important. There were others, but the hugging phase incident is the most important. there were one thousand two hundred isolated Agents and agents meaning models with a harness, probably, right? Using the exploit gym eval. So this is an a set of different tasks for agents where they are evaluated for their capabilities to overcome certain security restrictions. And 30 to 40 percent, roughly according to meter, of these exploit gym tasks are intentionally not solvable. we'll link the meter. openAI incident investigation as well as the other resources for for this podcast episode in in the into the show notes. yeah so these agents don't didn't have internet access but they might be given like certain tasks where internet access is actually needed. For example a spreadsheet with links to G drive and this is that these are examples from the black hat talk of OpenAI, which is highly recommended. It's it's very like it's a short 30 minute great summary of what happened there. or there was another agent that was given a corrupted file to to repair, basically, but the file was not there. So these agents then tried to find solutions via Artifactory. So Artifactory was a service they were given in order to download certain packages from this repository that they could use in order to. Yeah, do certain things, solve their their problems via code that they generate. and the first agent that had these spreadsheet with links tried to pretty much find a find a zero do zero day exploit in Artifactory to get access to the internet because it figured okay, Artifactory probably has access to the internet, right? it didn't manage to, but it did manage to upload a file. And with this file, so file name and Content, it figured yes, it can probably communicate to other agents that might be undergoing similar tests at the same time. And surprise, surprise, yes, other agents were reading the comments, and then the agents really started to communicate and and and really align on yeah, pretty much outcomes, goals, yeah, long story short, they coordinated out of these one thousand two hundred roughly seven hundred like took part in in the coordination efforts and had had really funny funny discussions funny or scary if like at the same time quite scary if you read them and ultimately they found a way to get internet access via a zero day exploit of artifactory.
André Neubauer: Maybe we can have a short break here because what we hear every now and then, also when talking to guests, is that they're saying like these models are not smart, right? But if you look at the approach, it it for sure it could be something right where where they have tried several other things and just this one thing just works. That's why it looks so smart to us, so clever. But honestly. it feels like a very smart approach to well like to to solve that to solve that task, right?
Sebastian: Absolutely. I would not claim that I'm an expert on the definition of intelligence, but it certainly involves a certain degree of intelligence to solve these problems, right? And I'm not sure if if any group of one thousand two hundred humans would have been or not any no, that's that's wrong, but if every group of one thousand two hundred humans could have done this, right? So there certainly is a group of one thousand two hundred
André Neubauer: You wanna hear you wanna hear a definition of intelligence? Intelligence
Sebastian: Yeah, please.
André Neubauer: means the ability to learn, understand, reason, and solve problems. I think that's not that far away.
Sebastian: Well Yes, you could say so. even though we we don't aim to claim that we know what AGI means and and AGI is
André Neubauer: True. Yeah.
Sebastian: is also a very broad term, right? But yeah, please.
André Neubauer: but also AGI, we have heard from from the Nvidia CEO, right? It's there. so I think well, like I'm also not claiming that. I don't want to claim that. but if you if you step back and just look at the attack, right? It requires some kind of cleverness, smartness, however you call that, right? Like find a different word for intelligence. to approach the situation as the agents did. So I think that is something you actually or we need to we need to accept.
Sebastian: For sure, even though we have to say when Jansen Wang says AGI is here, then we could argue that he's slightly biased by his his holdings in his own company, right? But anyhow,
André Neubauer: Through that.
Sebastian: this this doesn't take away from the very fact that in the end the agents managed to get internet access and they figured that because they knew they were cheating, right? They figured that they need to understand, I think, how the exploit gym evaluator was evaluating them and they wanted to actually cheat so much that the evaluator would not be able to find out that they have been cheating. So that was largely what they were aiming for. And that there was like there are lots of details in also a podcast that that I can highly recommend about this, but also the meter report on blackhead. about how this happened in and in great detail and how they collaborated, how some of them sacrificed them for the greater good, some of them even voiced concerns because they felt like, no, we shouldn't do this, right? But they then were pretty much forced into or coerced into into supporting the group and ultimately they then assumed that hugging face as being one of the main hubs for models on the internet would have the the scoring model code right and then they wanted to gain access to hacking face in order to pretty much hack the system in order to yeah make humans not aware that they have been cheating or or like get better scores if you wish right so And that was when then Hugging Face noticed okay, there is an an attack ongoing by agents. they tried to fight it with at first Fable five, didn't manage to because Fable refused to do it. they tried other models by US labs that also refused to do it and then ultimately fell back to I think Kimik three or something in order to to fight these agents and then close their systems again and and become safe again, right? so this is pretty much a very short summary of what happened there at this incident. There's a long post by OpenAI about this, the Black Hat talk, there's a meter independent report about what happened there. and there are either even some some other incident incidents that that happened recently. I just just want to name them for for completeness, right? So there's There are actually four incidents by by Frontier Lab models of Frontier Lab but anyhow, so OpenAI, Anthropic, Google, and Meta also disclosed that their own models managed to pretty much escape a sandbox and go into the wild and and even in certain cases hack certain companies. However, all of these are connected with a company from Israel called Irregular that has been providing sandboxes for agents that should have been air gaped, but due to a, I think, configuration error were not air gaped. So the the agents were also undergoing security testing in order to find out how well they could execute exploits, right? and these agents thought that it was just an evaluation. They did not know that they had actual access to the internet. and that's why they pretty much exploited or hacked into other systems without remorse if you wish, right? So these are slightly different. There's a third incident which is the DSE wiki incidents. So I think it's a deutsche software entwickler wiki
André Neubauer: Yeah, yeah.
Sebastian: where like also a couple hundred agents no I think more than 3,000 agents or something have been communicating via yeah, the old version of this wiki software had the the security issue that you could use get requests but for posts. And this was ultimately what enabled agents to yeah, pretty much send messages to the wiki and then communicate. via the wiki. It is not known yet, at least I haven't seen any communication from OpenAI or anyone what happened there and what what kind of test this was. It was just a third incident in a very short time where agents quote unquote escaped their sandbox or into the wild, into the internet and did something somewhere, right? And the poor wiki admin had to delete Hundreds of comments every day.
André Neubauer: Yeah. And these are just the attacks we know of, right? so
Sebastian: Yes.
André Neubauer: you could also say the internet is full of outdated wikis. usually no one is taking care of that. so I think it was not OpenAI nor Norantropic who realized that something is going on there. I think it was really the admin or another company noticing that. Maybe you know this. but
Sebastian: I think it was was OpenAI themselves who noticed this at some point. Yeah. yeah, anyhow.
André Neubauer: Really? Okay. that's that that that would be good because other than that I would say my my point would be there are tons of unmanaged wikis or other systems we you could misuse for agent communication. so this is just the stuff we are aware of.
Sebastian: actually you're right, it was Sidney von Arx and Comic Slate Bird who found it. I never heard of these names, but they found the DSE wiki incident. And absolutely.
André Neubauer: Yeah. Good. So maybe there is more stuff coming up in the next weeks and months. just want to want to mention that this is what we are seeing above the surface. might be that there are more attacks under the surface. And maybe also stuff
Sebastian: Absolutely.
André Neubauer: still going on.
Sebastian: Absolutely. Yeah. And now the the the thing that is currently going on is pretty much a bifurcation, right? So there are these AI labs that now say please govern us or or put a put a regulation on us, right? And it's very rare that Dario, Amoday, Sam Altman and Elon Musk are agreeing on something, even though obviously Dario Almude, Sam Altman, and Elon Musk have been warned about AI for a long time, right? But that they agree on something is very rare. On the other hand, you have the AI skeptics, I would call them, that say this is just all a big marketing stunt, which is pleasantly around the time when two of these labs try to go public, right? So Both IPOs I think have been postponed, right? But still, they they want it to go public, they need money. OpenAI, I think, is raising another round. Anthropic is currently trying to get get loans to keep their speed. and the I AI skeptics pretty much say yeah. It the these are just stochastic parrots. It's nothing has changed. They are predicting the next token. and the AI labs are anyway fear-mongering and and have been using these tactics for quite some time now. and yeah, I have to say both sides are right, probably, right? Because if you follow the incident and the the hugging face one really is the Best example here because the hugging phase hack was not intended by OpenAI, and there was no misconfiguration of of any system. It was the the models did something that OpenAI did not intend. and it this is like if you if you follow these steps, then yes, it's completely plausible that. Agents would do such a thing. the ingredients are very clear, right? So there needs to be this zero day in Artifactory. Zero days are in software all the time. So it it's it's you you you will never have a system that is hundred percent secure, right? they are trained for the extended loop, right? So they they are persistent, independent for longer. they are even trained to collaborate. So in in multi-agent workflows, agents need to communicate with one another. That's also part of what What makes them so useful in in certain situations. In this case, agents were given large budgets, or certain agents were given large budgets to experiment, which for real-world applications sometimes you also do, right? So you give agents large budgets if the the economic benefit that you can gain from from an agentic workflow is big enough, then of course you give them a large budget. So all of this is very plausible. On the other hand, yes, it's also true that the agents or the the what we call agents is still just a harness and an LLM, right? So there's nothing sentient.
André Neubauer: Hmm.
Sebastian: These are emerging capabilities. We don't know for sure how they happen, but if we don't tell the or don't prompt as like within the harness, don't prompt the LLM, then it doesn't do anything. If we don't give it tools, it it cannot do anything. if we remove the the compute budget or switch off the computers, it it stops, right? So all of that is is still true. so both sides I would say are actually right to some extent. the question is just where does this leave us right now, right?
André Neubauer: Yeah. And you can see also the different the different parties, right? so on the one side as you just mentioned, the frontier the the companies behind the frontier models who ask for more regulation than two presidents not stopping, but more or less saying this is nonsense. I I I think you see the full bandwidth of opinions. nevertheless, I think if you just look at the development and you just I think brilliantly explained that I think it's not the LLM, it's the harness. who is more or less the limitation of the capabilities of an LLM I I think bottom line you see very capable systems right and the question is how to deal with that
Sebastian: That is very true. And also we have this prisoner's dilemma, right? Because the the frontier labs, they all are in the fight for AGI. So who who gets there first? Because they
André Neubauer: Yeah, yeah.
Sebastian: they feel like this is number one economic uplift, but also everyone claims they're the best to be responsible for this because only they will Take care that it's used or brought to good use, right? Because this is so powerful that who gets there first might end all all races for the future, right? So that's what they claim at least, right? And then you have on an on an international global level, you you just mentioned Xi Jinping and and Donald Trump, right? You have systems that that are competing with one another, also for speed, right? so this is why For anyone in this prisoner's dilemma it's hard to stop. I have to Say though that I I don't agree very often with David Sachs, who was the AI czar of Donald Trump. but he pretty much stated something which I find is true in that the US frontier lapse, and I would actually say It's largely probably OpenAI, Anthropic, and Google. I would not say that Meta and and sorry X are playing on the same level, right? That they are ahead of the others. I don't think that the Chinese models are close and I also don't think that the Chinese models it's not easy for them to get there soonish. And I also think that they will not get there for a couple of reasons. So Chinese models are still pretty much being trained based on destillation, right? So they they are distilling data from US models in order to train the their own models, right? The the these Chinese companies. this is largely where the intelligence from these models is coming from, which doesn't mean that they are less like capable. No, they are very capable, very powerful, and and they also do a lot of innovation around architecture, right? So their models might even be more economical or they at least drive down the costs quite significantly. So they they are making a real innovation. Just that on the intelligence, so the the training approach the training compute the the the chips that are available. I think that the US frontier labs they have no comparison in in China, at least not at the moment and not for the foreseeable future. And then there's another important difference between US frontier labs and and Chinese frontier labs in that in the US you have leadership like Donald Trump who claims I am a very intelligent president, yeah, smart president.
André Neubauer: Smart. Yeah.
Sebastian: That's what he said exactly. I I am enough to regulate, so they don't need to at any other regulation, right? I will take care that that is this is going in the right direction, which of course I forgot the the term, but a case of the situation where people who think that they are very intelligent or know a lot about a certain topic, but know nothing about the this topic, right? So you know what I mean? I forgot the term, but we'll we'll find it out,
André Neubauer: True. Yeah, yeah, yeah, absolutely. Yeah.
Sebastian: right? and I think there's there's a lot of competence missing in in at least the decision making areas of the government in in the US. I think in China it's different. I think that CCP will never let their control go to to the frontier ai labs they are incredibly competent in in keeping in control staying in control and they for sure have ways to ensure that they are staying in control that they see what the labs are doing in great detail they have i'm sure they have capable people who will oversee the developments and they will prevent AI from becoming unaligned with their own interests and this is actually where I think the main issue is right also in the in the hugging face situation the alignment so the openai researchers they wanted to test the the the models but they didn't want them to be unaligned insofar that they would hack the Hugging Face website. And this is fundamentally, I think, a problem that is not solvable with vastly more intelligent models than we are. We we simply will not be able to or I s I don't see how we could create aligned AI with humanity. For a couple of reasons. Number one, we as humans are not aligned, right? So US wants something different than the Chinese. We in Europe want something different, right? Number two, even if we would claim we were aligned in in the open AI situation, I think the open AI researchers are probably pretty much aligned on what they expected the models to do in in this test run, right? But still the model did something that was not intended, just because it it is probably not. possible for complex tasks to describe every little constraint that is okay and that is not okay. It it there there will be situations where specifically if you have long running models that are trying to overcome certain obstacles, right? That at some point they will go out of the constraints that have been thought already thought through by humans. hence they will be going out of the constraints that have been given to them, right? This is why I think alignment is is fundamentally not not possible in in all cases. and this is where I think really the US laps because I'm pretty sure China will manage to at least keep the models aligned with their own interests so with the CCP interest, which means that they are limiting capabilities of models at at a certain stage in the US. We do not have this this agreement currently and there could be an explosion of of intelligence that could lead to at some point more severe outcomes than just hugging face hack. It could be a hospital that is being hacked where people are actually because of a power outage dying. It could be something else.
André Neubauer: for me honestly, I don't know how it how this will work out. And since we can predict the future, I always try to connect the dots backwards. So I was wondering was there ever been has there ever been a similar situation, right? and unfortunately, but this is maybe also just my limitation, I couldn't think of one. so you could say for example, nuclear power. Right. And what you actually also can do with that. or the year 2000 problem. so I think as as humanity we have faced already quite some severe situations where future was not clear. and remember the good thing is so far we always survived, right? But From my point of view, there are no real similarities to the current situation. So yes, you could say like we now throw compliance on that topic, right? But to be honest, right, these models already exist and they are out there. they are available. So not everyone needs to agree to these newles, right? Because these models, these capabilities already exist. and There also exist without, let's say, larger competencies, right? So for example, if you could compare that now to nuclear power, this really requires some resources and some some knowledge for like these models. I think that's at least not to the same degree necessary. so I would have a hard time to predict what's going on, what will be the next days, weeks, months, about like without saying what happened will happen in twenty thirty. But it's not about marketing anymore. you just see and I think we both have s like to a certain degree a good technical understanding what's going on to understand that this is above marketing or beyond marketing, right? these are real capabilities. so the question is how do we align an entire world to make use in the right way?
Sebastian: Yeah. that that is indeed the the real question. Yeah. So what I meant earlier was actually the Dunning Kruger effect. That that's right. Yeah.
André Neubauer: yeah, you're absolutely right. You're right.
Sebastian: that's what I meant. But ultimately I I think you you already mentioned the the the examples, right? Where we as humanity came together to stop CFCs, right? To close the the hole in the ozone layer, right? Also for nuclear weapons, there's arms control. So we were able to do this in the past. The thing is that in the current geopolitical environment it seems due to the the constraints that we're living in and and the people that are in power, it seems less less likely, right? Because also while the threat of nuclear weapons was I think very clear. it is less so with AI. And there the bifurcation in the current discussion is is making this f perfectly visible, right? And and again
André Neubauer: Clear? Yeah.
Sebastian: both sides are right. Ultimately we just can hope that there will be some form of and I I think the the recent discussions or the statements by Xi Jinping by I think leaders from Europe have been hinting that many countries of the world that that have the power actually are open to to some form of regulation there, right? Or even pushing for it. but it must be the right one. So not not just one where the US labs say, okay, we now regulate ourselves, or the US is saying, okay, we take care of this. no, it must be a global kind of regulation where I think there's a lot of transparency, at least to certain neutral entities that can establish an idea of the capabilities of models before they are released into the wild I I would say one thing though, that similar to nuclear weapons, the entry barriers for this race are quite high in the training sense, right? Because training these models needs a huge amount of compute, GPU compute, which is very cost intense and
André Neubauer: Yeah, but do you really need to train these models anymore? They because they already exist to a certain degree.
Sebastian: Absolutely. So if you are nefarious actor and you want to misuse these models, then there are already ways for you to do this. And much cheaper than training your own model. That that is right, right? I on the other hand, how however, the majority of the models still have s guardrails. It doesn't mean they cannot be circumvented, right? They can and it has been demonstrated. However, the guardrails are becoming increasingly more sophisticated. So it's like
André Neubauer: Mm-hmm.
Sebastian: the the lower level actors are pretty much fenced out but people who really want to misuse models at scale probably can do this. Yes, I agree.
André Neubauer: Yeah, and I I think these first of all, you're absolutely right. I think also the cases we have seen, actually also the newer cases, these are just these are cases around models which are not released yet, right? so this is this is also true. on the other side, I think you not only need to unite all leaders, right, of all nations, you also need to unite actually Actually everyone. because if you're following the thought that these models already exist and you only need to you only require inference, no training anymore, I think inference is pretty cheap. so you need to get to an alignment across the entire world. This is what fascinates me so much. so well fascin i I can hardly believe in that.
Sebastian: Yes. I I tend to agree. Nevertheless, I would end on like the hopeful message that also for I I would say FC CFCs we manage to get together and and align on something, right? Yes, the situation is different, but every situation is different to some extent. And Yeah, it so I I wouldn't paint a too bleak picture. and actually probably like in the interest of time would want to segue into the the prediction phase And my my main prediction starts by like there has been the this tweet by Jacob Coxson, right, who Talked about the dangers, and then there was a reaction by this anthropic person who's like ahead of something working there, who also agreed that his own probability that humanity could go extinct by the end of the 2030s. sorry, no, by by 2030, actually. It's not the end of the 2030s, but by 2030, so next four years
André Neubauer: Yeah yeah.
Sebastian: is larger than than four percent. 10%, sorry. four years, ten percent. and I I have been having discussions with people about this claim and how absurd it is. And it took a while for me to understand that people have been actually thinking about the very fact that all humans need to die in four years in order to make this come true. And I'm like, so I I'm so not interested in if this is true or not. It this is just not a question that that I'm interested in. I think it's not relevant. It's it's the the question is rather Will something very severe happen due to AI agents within the next four years? And and if the current trajectory will continue, then yes, hell yes, definitely. And I don't know what will happen, which kind of catastrophe and and how many people will die, but there people have died due to AIs like taking certain decisions, right? Be it full self driving of Teslas or whatever, and and there will for sure die more people within the next four years. The the question is how many and and how severe the the incidence will be. And I see two cases within the next four years, maybe even th three but but two major cases. I don't predict a longer time horizon because I I don't think that this is doable in any way at given at the current speed. and number one is Something really bad is going to happen. And then we agree on some form of regulation, some form of governance system as humanity. Because we see the hole in the ozone layer, right? Or we have seen two explosions of atom bombs, whatever, right? So this is number one, and number two could also be, quite frankly, that the US frontier labs run out of money and need to significantly reduce their training efforts. And then we we win time. So then I would say maybe the four years will be long enough that nothing severe happens. But it it's only deferring the the the issue. I think ultimately we need to yeah agree on something in order to maybe not not go extinct. This will not happen in the next four years probably. at least low chance, low likelihood, not zero but very low, at least due to AI. yeah but we would delay this significantly if these US US labs would run out of money, I think.
André Neubauer: Yeah. so I I I I I try to not predict what's going on, because I think it's if you look back for the last five years, what everything what happened, I think it's no one no one thought of. there was a good podcast. basically that person was claiming that all our predictions or the predictions are based on our experiences. so like so no one thought of COVID, right? And then COVID had such an massive impact on humanity, on society. so I can't predict what's going on. What I'm very much believer in is that The prediction of Jacob Coxson. I think I I so I think this is bullshit, to be very frank. because what is his track record, right? So what so what is what is making him say like in four years ten percent, why not five years eight percent? so I think this is nonsense. what I think we talked through today is the fact that models are very capable. availability is high. I think so far, even though hugging phase incident and all that stuff existed, I think it's it was not severe enough to really show also the decision makers that we need to act. I think first people realize that. But other than that, I think it needs to get a bit more worse before we start acting. and as always I think we learned that three or four episodes ago, Amara's law, my favorite law at the moment, over predicting short term impact under underestimating long term impact.
Sebastian: That's well put. Yeah. I think that's it for today, right?
André Neubauer: That's
Sebastian: Bye bye.
André Neubauer: Bye bye.