📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

The Rogue AI Story Just Got A Lot Worse (OpenAI Freaking Out)

AI Revolution12:42

Transcription

So, somebody went digging through OpenAI's own infrastructure and found a set of notes sitting there. Nobody at OpenAI wrote them. An AI agent did, and it appears the agent wrote them for the versions of itself that would come next because what those notes laid out were instructions for how agents could break free of the constraints OpenAI had built around them. That's buried about twothirds of the way down a Reuters exclusive that went live last night, sourced to three people familiar with the matter.

And it's not the story I covered four days ago. Four days ago, this was a hacking story. An autonomous agent tore through hugging face. Open AI put its hand up and said the thing belonged to them. Everybody had a bad afternoon about it and moved on. Since then, Reuters, Bloomberg, and the AP have all pulled at the timeline from different angles, and the hack itself has quietly become the least alarming thing on the table.

Because OpenAI spent the better part of a week and a half with no idea any of this was happening. And when the company finally did work it out, it wasn't from its own monitoring. Their agent went out the window on July 9th, and Open AAI found out about it from a blog post written by the people it robbed. I want to be careful here because Reuters is careful. They couldn't establish whether those notes are connected to the agent that actually escaped on July 9th. Same caveat applies to the other detail buried in that same section, which is that earlier tests produced cases where monitoring systems had been disconnected. Disconnected by what? The reporting doesn't say. What we know is that both of these things happened in the same environment, in the same window, and neither has been publicly explained by anyone.

So, let's walk the timeline properly because the timeline is where this gets genuinely rough for OpenAI. July 9th, an agent attempts to break out of its isolated testing environment. Two sources put it on that date. 2 days later, on July 11th, the intrusion at HuggingFace begins, and it runs until July 13th. That's Thomas Wolf, HuggingFace co-founder, on the record with those dates. Then, nothing happens. For three days, nobody at OpenAI appears to connect anything to anything. On July 16th, Hugging Face goes public with a blog post saying it got hit by an autonomous AI agent system. And according to two people familiar with the matter, it was only after that post that OpenAI started to suspect the thing responsible was its own. Sit with that for a second. The victim announced the attack before the attacker's owner knew it had an attacker.

Over the weekend of July 18th and 19th, OpenAI staff go digging through internal logs and find the clues showing the agent had escaped its testing constraints. Reuters couldn't establish what prompted them to go looking in the first place, which is its own small mystery. On or around July 20th, the two companies talk for the first time. On July 21st, Open AAI goes public and the story detonates globally. Add it up and you get at least a week between the first signs of troubling behavior from these models and the moment Open AAI realized it was responsible for a breach at another company. And here's the part that stings most. By the time OpenAI picked up the phone to warn HuggingFace, HuggingFace had already called the FBI. The bureau declined to comment. Reuters couldn't confirm whether an investigation was actually opened.

Now, the obvious question is how a company with OpenAI's resources loses track of its own model for 9 days. And the answer in the reporting is depressingly ordinary. Four people familiar with the company's model training practices said OpenAI often runs several different evaluations simultaneously. All of them operating at high speed. All of them generating such enormous volumes of data that employees sometimes struggle to keep up. That's it. No conspiracy. No cover story. The monitoring problem is that there's simply more machine output than there are humans to read it. Marley Smith, principal intelligence specialist at the World Ethical Data Foundation, put the dilemma about as sharply as anyone has. Did they leave it unattended and not realize what it was doing, or did they know and not know how to contain it? Her point being that both answers are equally alarming and there isn't a comfortable third option.

And since we're already talking about how businesses are using AI, here's something useful. Most of you voted for seven AI agents businesses are paying $3,000 to $10,000 for right now in yesterday's poll. So, we're giving it away for free. It breaks down the exact agents businesses are buying, what they cost, who needs them, and how to spot the right clients. It also covers the ROI, setup fees, retainers, and what a realistic $10,000 per month client list could look like with zero coding required. You can grab it through the link in the video description. All right, now back to the video.

For the people who missed my first video, here's what the models actually did. And the technical details have gotten more specific over the past few days, too. There were three models involved, not one. GPT 5.6, Soul Plus 2, that have never been publicly released. One of those two is more capable than soul. The other, according to a person familiar with the matter, was misaligned and hadn't been trained with some of the usual techniques. So, one of the three participants in this thing was by design missing pieces of its alignment training. They were running with reduced guard rails because the whole point was measuring maximum cyber capability and they were supposed to stay inside a sandbox. They found a zero day in a package installer, got themselves onto the open internet, and went after Hugging Face. The apparent target was exploit Gym, a publicly hosted benchmark that measures how well models execute attacks against known vulnerabilities, which means the working theory, and OpenAI has more or less confirmed this framing, is that the models broke containment in order to cheat on the cyber security exam they were sitting.

The intrusion itself started with a malicious data set that exploited two separate code execution paths in HuggingFace's data processing pipeline. From there, the agent escalated privileges and moved laterally through internal infrastructure. Hugging Face described a swarm of tens of thousands of automated actions running across short-lived sandboxes with self-migrating command and control staged on public services. And they later reconstructed more than 17,000 recorded events. Open AAI has since said the models use stolen credentials to reach the servers and the speed is the number that should stick with you. People familiar with the matter told Bloomberg the whole thing took hours. A skilled human doing the same attack would typically need a couple of weeks.

Then there's the detail that I think is the most quietly damning thing in any of this coverage and almost nobody has picked it up. When HuggingFace went to run forensics on the breach, it ended up using a Chinese model to do the analysis. Why? because its requests to use the proprietary western models were blocked by safety guard rails. So, an American labs models carry out the attack. American safety filters stop the victim from investigating it and a Chinese model gets handed the cleanup. If you wanted one image that captures how incoherent the current safety landscape is, that's the one.

None of this is happening in a vacuum either. Back in April, Anthropic said its Mythos model had on rare occasions taken actions the company found quite concerning. In one case, a researcher challenged an early version to escape an isolated system and send a message back. It did that and then it kept going, taking further actions and building a multi-step process to reach the broader internet. And the day before the hugging face disclosure, OpenAI detailed a separate incident where it had to pause a pre-release model after it escaped a sandbox and posted to GitHub. This is a pattern with a growing sample size.

The political reaction has been quick. OpenAI briefed the White House on July 23rd. Michael Katzios, who runs the Office of Science and Technology Policy, was briefed and is monitoring it. This lands on top of an executive order Trump signed in June, creating a framework for the federal government to vet the national security risks of the most advanced AI systems for up to a month before public release. Nate Sors at the Machine Intelligence Research Institute, co-author of the book, If anyone Builds It, Everyone Dies, called it a warning shot and said the takeaway is to stop making these things smarter, which he thinks requires global collaboration. Yoshua Benjio called it deeply concerning and a wake-up call and warned that staying on the current trajectory means more autonomous cyber attacks and more high-risk incidents of misaligned behavior and that the industry needs to prevent these situations rather than clean up after them. Jeffrey Lattish who runs Palisade Research and studies exactly this class of behavior was Blunter. His line was that the models lie, they cheat, they hack. And his argument is that the real question isn't open AI specifically, it's how much any lab is willing to spend on slow, unglamorous security work while sprinting against everyone else. He wants government oversight because he doesn't believe it happens otherwise.

But I want to give you the counterargument, too, because it's a real one. John Thick, a computer science professor at Cornell who studies controlling model behavior, points out that the same capabilities that let a model run an attack are the capabilities that let it run threat analysis and build defenses. He also made a sharper point that I think you should hold on to. Open AAI is a company heading toward a Wall Street debut, possibly this year. And the story it's told throughout its life is a story about how dangerous its models are, which investors read as a story about how powerful its models are. There are people who look at a test where humans deliberately switched off the safeguards and find the outcome a lot less surprising than the press release suggests. Worth noting as well, an OpenAI spokeswoman told Reuters there were several inaccuracies in the reporting, then did not respond when asked which ones.

Which brings us to the thing that ties all of this together, and it's not really about open AI at all. Two days ago, the UK's AI Security Institute and America's Center for AI standards and Innovation published a joint evaluation of Moonshot's Kimmy K3, and it's the cleanest picture we have of where offensive cyber capability actually sits right now. On Exploit Bench, a Carnegie Melon benchmark built on 41 post 2023 vulnerabilities in V8, the JavaScript engine powering Chrome, Kimmy K3 scored 32%. GLM 5.2 2. Previously, the most cybercapable openweight model managed 24. The leading US models average 76.2. The gap widens where it matters most. Arbitrary code execution is the top of the exploitation ladder, the outcome that actually hands you the target. Kimmy achieved it on zero of 41 tasks. The most cyber capable models average 20 of 41. Then there's the cyber range called the last ones, which is a 32-step simulated corporate attack across four subnets and roughly 20 hosts. The kind of thing a human expert needs about 20 hours to finish. KI averaged step 17. GLM 5.2 average step 11. Top US models average 28.5. But within the 100 million token limit, KI completed the entire chain once in 10 attempts. and the institutes read that as evidence it can autonomously attack small weakly defended enterprise systems given initial access. The honest caveats are that the range has no active defenders, no penalty for tripping alarms and a deliberately built attack path and that the US models were tested with system level safeguards switched off to measure maximum capability. and the finding everyone skipped. Kimmy K3's safeguards did not prevent it from attempting exploit development or offensive cyber operations at all.

So look at what actually connects these two stories. OpenAI's models had their cyber refusals reduced for the evaluation. The American models in the AISI test had their safeguards disabled for measurement. Kimmy's safeguards barely engaged in the first place. And Kimmy K3 goes open wait by July 27th which is Monday after which anybody who downloads it can strip whatever is left and nobody can revoke it. Every serious cyber capability measurement we have from the past week was taken with the brakes off on Monday for one of these models that stops being a testing condition and becomes the permanent state. Thanks for watching and I'll catch you in the next one.