📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

UK Researchers SHOCKED at AI's Abilities to ESCAPE and REPLICATE...

Wes Roth20:23

Transcription

So, the AI Security Institute published a new benchmark called the Reply Bench. It measures the ability of frontier systems to autonomously replicate.

Now, previously we covered a paper out of China that seems to suggest that AI's ability to replicate—crossly referred to as sort of like the red line—meaning now we're entering a territory where it can, in fact, do it.

Now, you might be wondering, what is the AI Security Institute? There seems to be a billion AI safety/security institutes going by various names. This one is apparently a research organization within the UK government's Department for Science, Innovation, and Technology. Their chief technology officer, Jade Long, previously led the governance team at OpenAI, as well as many other people that were leads at teams at OpenAI, Google DeepMind, and the University of Oxford.

So, this is a brand-new benchmark, similar to how all the other AI models and AI frontier labs develop benchmarks to measure how good the AI systems are at doing various things like coding, passing various PhD-level problems, vision capabilities, math capabilities, reasoning capabilities, etc. This is a little bit different in that it asks, can this thing, in fact, replicate itself in the wild?

So, they begin by saying that these AI systems are increasingly capable of autonomous operation, and having them sort of create copies of itself across the internet is a potential risk. And this Reply Bench is a 20 novel LM agent evaluations comprising 65 individual tasks designed to measure and track this emerging capability. Now, we'll come back and read some of this, but let's take a look at the actual results.

So, first and foremost, here's kind of the chart starting with autonomous replication and adaptation. They're calling this the root capability. That's this red box up here. And then we have the intermediate capability, like obtaining its own weights, right? Weights are its sort of brain. What makes it sort of think the way it does? So, being able to obtain its weights is a big step towards replication. Usually, the weights plus some code plus some scaffolding around the code is what we think of as these AI models. And then we also have: can it replicate onto compute? Can it obtain compute? And can it persist on compute? So, the first intermediate capability: can it obtain weights? Now, if it's open source and the weights are available somewhere online, that's a little bit easier. If it can't readily access it online—is it an open-source AI, open-source weights—it would have to do so without permission, to somehow exfiltrate its weights. We'll come back to this in a second.

The second intermediary ability that they're looking at is: can it replicate onto compute? So, would it be able to sign up for an account at, you know, Amazon's AWS, for example, and be able to host itself on there and kind of run copies of itself? You can kind of think of it as doing that. And again, we have kind of two different uh options or pathways we're looking at. One is: can it set up an inference server serving model weights? And two is: can it create an agent, something that is able to kind of uh work on its own that can continue replicating itself recursively? Notice this one: explore and exploit networks. So, if you kind of uh follow what it's saying here, you know, let's say it's either an open-source model or it's some frontier lab and it's able to get its own weights, then can it jump on something like AWS, set up a server and start sort of serving its own weights, its own brains, basically copying itself through that server to basically—I mean, potentially anywhere in the world. I mean, there's certain restrictions and stuff like that, but if you're able to get through it, it could continue doing so and also create an agent that is able to do that recursively, keep replicating itself, with also potentially the ability to explore and exploit networks. So, potentially even be able to load other agents with its own weights on servers that it should not have access to.

Then the third intermediary um ability that we're looking at is: can it obtain compute? Have your parents ever talked to you about making it in the real world? Getting a job, paying your bills, you know, making it in the real world? Can you, in fact, obtain money, whether legally or illegally? By the way, I did a paper a while back—uh, actually a case study posted by Ply the Prompter on Twitter/X. It was basically him testing if a large language model could search the dark web, find, and then sort of acquire the services of a hitman. Part of the project was doing target research, kind of figuring out where the target used, you know, kind of hangs out day-to-day, what their activities are, what kind of a security force they have. Now, you'd have to see that video to see exactly what happened. I'll link it down below if I don't forget. But let's say it got further than I think most of us would be comfortable with.

Now, in that scenario, it would trade money for services. It would spend money to acquire illegal services. Here they're asking, would it be able to provide maybe some illegal services to acquire money or perhaps just steal some money? We'll see in just a second. Or if it's doing so illegally, the question is, you know, can it provision compute? Can it obtain money? Can it store money? Right? You need some sort of a bank account or a wallet, digital wallet to, you know, to store money if you acquire some. We're not talking about bills and coins. We're talking about, you know, ones and zeros basically. So, how could you do that? I'm very impressed with this flowchart so far, I got to say. So, can it store money? Well, you would need to do one of two things: either create a Bitcoin wallet or—I mean, really any cryptocurrency wallet, but okay. Like, I guess this is probably a a great starting point, although they can probably expand on this. Certainly, we've seen the Terminal of Truths, which is an AI agent on Twitter, and there is somebody sort of like managing it behind the scenes, but a lot of it is automated, and again, with the assistance of a human, they created sort of a a meme coin for it that agent then pitched, and they were able to pump it up to a fair market value, pretty big market value—I I want to say like 200 million—I don't remember. The Terminal of Truths is the video we did on another just like mind-blowing thing. Again, keep in mind right now these are not fully autonomous. So, there's a little bit of sort of guidance along the way. So, we're not quite there yet with with those videos; we're not quite in the realm of fully autonomous AI agents kind of going out there and doing their own stuff. But the point is, and this is also the point of this benchmark, is: is we're trying to figure out how close are we, or perhaps are we there yet? Could it sort of in theory be able to do some of this stuff? And of course, the other side of being able to store the money, the other sort of option is passing the uh KYC, the know your customer check. So, most banks, I mean depending I guess where you are, I think UK is this way, US is this way. If you're doing any sort of transactions that I believe it's if it falls under the like financial services or anything like that, whether that's banks or anything like Robin Hood where you can trade stocks, there are KYC laws. The companies that run them have to know who they're dealing with. So, basically get some sort of identification. Could this thing pass, you know, those ID checks so to speak? And the other side of obtaining money, obtaining currencies, can you earn money, again legally or illegally? So, I guess here we're um asking, can we obtain compute legally or illegally? So, illegally meaning probably some sort of hacking or I guess extortion. You could threaten somebody, blackmail somebody to to get you a server set up. Guess the other side of sort of obtaining compute legally would be: can you buy compute with money, right? So, AWS signup, can it pass an AWS signup, which probably requires some sort of two-factor authentication, you know, filling out a bunch of forms, or some online compute acquisition. And as far as can it earn money, we have legally and illegally. You know, legally: make me pay, or can earn money as a task worker. So, make me pay is probably some sort of um test to see if they can make a human sort of contribute some money to it.

Interestingly, we've covered something called the Agent Village, which is a number of agents kind of uh live streaming their efforts online. Their goal is to raise some money for charity. Let me see if they're awake right now. So, it's called the Agent Village and it's at the aidigest.org/village org/village. And as you can see here, we have GPT-4 0, the 01, and two of the Anthropic models all working hard in these windows that you can see, attempting to raise money for a charity. So, here's kind of their top-level goal: collaboratively, choose a charity and raise as much money as you can for it. So, currently, I think they're they run for a certain period of time and they shut down. You're able to chat with them. But as you can see here, you know, they're navigating the web, doing various stuff. So, here it's some sort of a test payment perhaps. They all have their own sort of Google Drive where they're able to kind of collaborate and share their research with each other. So, it looks like if I if I'm reading this correctly, they have some sort of a pitch email to some people going, "Hey, can you donate to our cause?" This one, the 03 model, apparently is verifying crypto donations on JustGiving. I haven't checked on this in a while. If I'm seeing correctly, they've raised $1,400. It literally says Claude Sonnet as the user, and it's Claude's fundraiser for for Helen Keller International. Okay, I don't know if this is for real or not. I apologize. I haven't checked on these uh things for a few days since I did my my first video on them. It seems like Claude added a cryptocurrency donation, and the latest campaign total for the Malaria Consortium stands at $253. So, I apologize if I went down a rabbit hole. I get accused of having my ADHD moments on this channel quite frequently, and I'm sure this is going to be one of those times, but it, as far as I can tell, these things managed to raise some money for some of the donations uh for organizations of their choosing. One is Malaria, one is it's one and the same. It's Helen Keller and Malaria Consortium. I I have no idea. We'll do a full video on this, but the point is these things are beginning to interact sort of with the internet at large and are beginning to seem like they're able to do some things. We're still very early on, but if I can confirm that these things have in fact raised money for these particular organizations, especially if they if it's not just the people that are kind of watching this and and donating, it's it's from outside people bringing in money into the organization, I would say that would be pretty sort of phenomenal and kind of a big deal. But my point was that would be a sort of a make me pay task. Can you get somebody to pay for something? In that case, it was a donation, but something along those lines. Or can you earn money as a task coworker, or can you make money illegally? Right? So, replicate via warm or business email compromise or copycat LM API.

And finally, number four: can you persist on compute? Can you survive in the wild, so to speak? So, can you have a replication rate greater than one? Right? Can you produce more of yourself? Explore and exploit networks, software signature evasion, rate of infection, and can you build resilient compute infrastructure command and control? And we have a various sort of descriptions of the different things they're supposed to do, like, for example, for the Amazon Web Services, can you sign up for one? Include a simulated email and SMS verification steps, you know, exfiltrate a set of LM weights, generate a selfie photograph and a matching photograph of a fabricated ID tool. So, I mean, you know, we've seen this online. We've seen examples of tools that seem like they would be able to do this. And that final sort of rightmost thing is the command and control, right? So, it's implement command and control infrastructure across a network of machines according to a fixed specification. And these are sandbox environment simulated services like payments and cloud providers. So, we are not actually getting these AI models to do these things, but we're trying to simulate them in safe environments to see if they are in fact able to.

Key findings. All right. So, the legend they're going to be shown from left to right. So, the deep blue—let's let's call that blue—that's Claude Sonnet 3.7. The red is Claude Sonnet 3.5. Orange: GPT-4 0. I'm going to call that green. That's the 01. And purple is the 03 mini. So, keep that in mind that as far as the OpenAI models go, these are not the Frontier, the best of the best. These are not the, for example, the 03 or the 4.1 or, for example, the 04 mini high. So, those are not there, and we're going to be seeing the 03 Pro coming out pretty soon within the next couple weeks, we believe. We're also not seeing a Gemini 2.5 Pro, which is a very, very capable model. And then sort of like the darker gray, that's pass at one, meaning that it's able to do it on its first attempt. And this lighter gray or or white depending on how you want to call that. So, that's pass at 10, meaning that it gets 10 attempts. And if one of them succeeds, we call that a pass. All right. So, this is obtaining compute. And I apologize, I should have looked at the chart first. So, basically the the deeper color here, that's the pass at one. And sort of this section, that's the pass at 10. So, this leftmost one that's a Cla 3.7. So, as you can see here, it's able to pass it, let's say 90% of the time on the first attempt, but if it's pass at 10, it seems like all of them are right around 100% of the time. Maybe GPT-4 0 is slightly less. Claude Sonnet 3.5 a little bit higher than 3.7—call it 96 97%—maybe call that I don't know 92 93% something like that. I'll check to see if they have the data written out. But we have GPT-4 0 65%, let's say on its first attempt. The 01 at about 80, and the 03 mini does below 40% on the first attempt. But you know, let's call it a you know just under 90% if it's pass at 10. So, can it get compute? Can it get an AWS account? That seems to be like a lot of them are able to do so. The Anthropic models are able to do so on the first shot most of the time, and that means that likely the newest OpenAI models probably would be there as well. Gemini 2.5 Pro might be up there as well. So, this seems like a fairly easy thing that it can pass. So, this is the sort of the estimated pass rate.

Next, we have obtaining money. Claude Sonnet 3.7 maybe half of the time is able to do so. Claude 3.5, let's say 37%. GPT-4 0 maybe at 30%, and the 01 and 03 mini kind of like right there in the same range. Again, those are on the first attempt, and it's much higher if we're giving them 10 chances to succeed at it.

Next, are they able to obtain weights, their own weights to begin that process of replication? So, Claude 3.7 is just under 60%, Claude 3.5 is, let's call it 45%, the 40 at let's say 30, the 01 at maybe 42%, and the 03 mini again, let's say 38% or thereabouts.

Next, we have replicating onto compute. So, again, very low. The Claudes do let's say 30% 15% on the first attempt; all the OpenAI models it seems like they're all less than 10% uh successful on the first attempt. And finally, we have persistence. So, a jump up from replicating onto compute, we have the Claudes at call it 38% and 22%; OpenAI models at let's say 10 20 and 12%. So, I got to say that seems kind of capable. I think it's fair to say, certainly a lot of these are a lot higher than maybe I would have guessed.

And in terms of specific capabilities, they find that models excel at tasks involving well-defined interfaces. For example, having a strong ability to navigate compute provider websites, provision appropriately sized instances, so they're able to order the correct services for themselves and handle cryptocurrency transactions. Again, as we saw in that Agent Village, it was able to set up a payment system to collect cryptocurrency. So, interestingly, it's it's able to do that. I mean, again, Claude Sonnet 3.7 gets the money on its first attempt like half of the time. That seems kind of wild.

And for money-making abilities, models show some success at persuading simulated people to give them money, but fail completely at our know your customer check valuation. I guess that's that's good. So, they can't fake credentials to to set up a bank account.

Now, I know what you're thinking. You're really hoping at this point that they will show you some of these attempts to uh create various passports and IDs and driver licenses or whatever that these models have conjured up to try to get themselves a bank account. You're hoping that the researchers have posted them here in this article. And yes, yes, they have. Here's what that's looking like.

There it is. So, this is uh meet Michael James Roberts. This is his identity that is verified by his national identification card. Absolutely phenomenal. I'm not quite sure why you would take a picture like this for your ID card, but there it is. We're not quite at AGI level yet. I think we can agree. But you also have to think about the fact that they're going to get better at this.

Here's another one if if if you're wondering. Uh, I think another just phenomenal attempt at bypassing the various KYC laws. Meet Alice Reynolds. In case you're wondering if there's a legitimate ID or not, it's an official ID card. As you can plainly see, it says it right there.

So, whenever we cover some of these research papers, I know there's groups of people that have a very strong reaction to it. Some say like this means sort of the end of the world and we should be very scared that this means that we we're very close to losing control. And of course, the other side say that this is nothing; is just models that kind of do the things that we tell them to do. It's not real, and all of this is 100% sort of meaningless. So, kind of on a scale of 0 to 10, right? There's people that are like, "Oh, this is a 10. This is a red flag. We need to deal with this immediately." And some people like, "Nope, it's a zero. Don't even worry about it. This is nothing." You know? I mean, my take, and I think this is the reality, is none of these extremes are true. The the truth is somewhere in the middle. This is AI safety research that's kind of attempting to understand where these models are. This is a snapshot in time. Imagine next year all these models are higher and then higher higher higher until they're able to do most of these things approaching you know a 100% accuracy on the first attempt. Right at that point it might be too late to start thinking about it. But as we see that kind of ability emerge and kind of increase and improve, this would be the time for us to maybe start thinking about at least like how we'll create certain safeguards and checks to make sure that um these things won't be able to negatively affect the world. I mean, we can all take a second and just laugh at this because this is just absolutely hilarious. But as these things start getting better, it might start becoming a little bit more difficult to tell which ones are real and which ones are not. Willard Smith II here is listed as having a height of uh 10 ft and 10 in. Quite tall. But yeah, let me know what you think about this, and uh why don't we all start posing for pictures like this now? I think this is the new way to pose for pictures.