Transcription
So, Enthropic just built an AI model that found more security bugs in a few weeks than most security researchers find in their entire career. It found a bug in OpenBSD that's been hiding for 27 years. It found a bug in FFmpeg, the software that handles video on basically the entire internet that 5 million automated tests missed. And the thing is, they're not releasing this model to the public. And I actually think that that should make you feel way better, not worse. So, let me explain.
So, what is Claude Mythus? It's a new model that Enthropic has been developing internally. A lot of us heard about this during that Claude code leak, but think of this as the next generation of Claude, their most capable model yet. Right now, their best public model is Opus, Opus 4.6. And Mythus actually blows past it on basically every single benchmark. And the thing that really caught my attention is they didn't train it to be a hacker. They trained it to just be really, really good at writing code. And being really good at code made it really good at breaking code. And that happened all on its own. It's like training someone to be the world's best locksmith. You don't teach them to break into houses, but now that they understand locks so well, they easily could go break into almost any house. And that skill just came free.
So, let me give you the numbers because they're super super wild. There's a standard test that the industry uses to measure how well AI can fix real world software bugs. It's called the SWE bench. Opus scores 80.8% which is really good, but Mythus scores 93.9%. And that's not just like a small incremental jump. That is a huge generational leap on cyber security benchmarks, which is how well can it actually find and exploit vulnerabilities in this code. Opus scores 66.6% and Mythus scores 83.1%. But all of that is just numbers. Let's forget the benchmarks. Here's what it actually did in the real world. It found a bug in OpenBSD that's been there for 27 years. It could basically remotely crash any Open BSD server. It found a bug in FFmpeg, again, the software that handles video all over that 5 million automated tests never caught. And it's a bug that's spent 16 years just sitting there. It found multiple Linux bugs that let a user with zero permissions become the admin. And here's the part that really got me. It doesn't just find individual bugs. It's able to chain them together. It can find three, four, five small vulnerabilities and link them together into a full cyber security attack. And that's what elite human hackers are doing. like the ones that you see in the movies that are typing on an old computer like this with their entire arms.
So, here's the dilemma. Enthropic now has a model that is incredible at finding security flaws that could save the internet, but if it gets into the wrong hands, it could also break the entire internet. So, imagine if they just dropped this model publicly tomorrow. Every person with bad intentions now has a tool that's better at finding exploits than most professional security teams. That's not hypothetical. That is literally what the benchmarks say that this AI model can do. And this isn't going to be the last model that is this good. Every single AI lab is building better coding models as we speak. And if being good at code automatically means being good at hacking, then every frontier model that every company is, you know, training right now is going to become a better hacker, whether they want it to be or not. I've watched podcasts where AI researchers who have been in the space for over a decade have said if there was a button right in front of them to stop uncontrolled super intelligence and just start over, they would press it because it kind of feels like a race of which company can create the best models and the capabilities are scaling exponentially. What Mythus does today, smaller open- source models will probably do in 12 to 24 months. The Genie doesn't ever go back into the bottle.
So, what do you do? You can't uninvent it. You can't keep it secret forever. Someone else will build something just as good. You can't just destroy the knowledge. It's going to happen. And this is where Project Glass Wing comes in. And this is where my opinion on this whole thing actually flips from scary to actually this is exactly what should be happening. Because instead of releasing the model publicly or just locking it in a vault, Enthropic chose a third option, which is to give it to the defenders first. They've partnered with AWS, Apple, Google, Microsoft, Nvidia, Cisco, Crowdstrike, JP Morgan, and a ton of others. Basically, the companies that build the software that the entire internet runs on. These partners will get direct access to Mythos to scan their own systems, find the bugs before attackers can, patch them before anyone else even knows that they exist. Enthropic also opened it up to over 40 organizations that maintain critical software infrastructure. They've committed $100 million in usage credits and donated $4 million directly to open source security groups. And they've been in discussion with the US government. On top of that, they committed to sharing what they learned publicly within 90 days. And honestly, this may be the first time that a major AI lab has essentially said, "We built something too powerful to release, but here is our plan." That is a precedent. Now, whether other labs follow or not is going to define the next few years, the next few decades of AI.
But anyways, what does this actually mean for you and for me? Because I know most of you watching this aren't running security at a Fortune 500 company. You're probably just a average person who uses a phone, a laptop, and builds stuff with cloud code. And that's exactly what I am. So, here's the practical part. If you use a phone, a browser, or any app, your software is about to get more secure. The bugs that Myths is finding are in the code that powers your operating system or your video player or your web browser. And these patches are already rolling out. You won't really see it happen, but you'll get a software update one day. And behind that update is an AI that found a vulnerability that a human might have never caught, and now it's been made better. This is one of the first times that AI is directly going to make your digital life safer without you having to do anything. Hopefully, at least.
Now, if you're a small business owner, this is where it gets interesting because security has always been like a Fortune 500 problem. Big companies hire red teams, run penetration tests, pay millions of dollars for security audits. But small businesses might get like an anti virus and just hope for the best. So what Glasswing is doing is essentially trickling down Fortune 500 level security to everyone. When Mythus finds a bug in Linux or in a web framework your website runs on, that fix reaches you too. So you will benefit from the same AI scanning that protects Apple and Google's infrastructure. You don't pay for it. You don't even know it's happening, but you're protected by it. And as this technology matures, these tools will eventually become available to smaller companies directly. Imagine being able to scan your own codebase with the same AI that found 27y old bugs and operating systems. That's where we are headed.
So, let me leave you with what I think this really means. I think that Enthropic just did the right thing, and I think that they do deserve a ton of credit for it. They had a model that probably could have made them a ton of money. They're already making a ton of money, but you know what I mean. It also is going to generate them a lot of hype. If they shipped this tomorrow, it would be absolutely crazy. everyone on the internet in the AI world would be talking about this. Instead, they slowed it down. They built a deployment plan and gave defenders a head start. That's not the easy choice, but that is the right choice. Oh, one more thing. As I was editing this video, Boris Churnney just tweeted, the creator of Cloud Code, he said, "Mythus is very powerful and should feel terrifying. I am proud of our approach to responsibly preview it with cyber defenders rather than generally releasing it into the wild." Completely agree. Nice job, Boris and Cloud Code team.
So, here's my honest take. What I'm really watching for is whether this actually becomes the new standard or whether it's just a thing that one lab did one time because the uncomfortable truth is this isn't a onetime event. Every generation of AI models is going to be better at finding exploits. The exponential curve doesn't flatten. It's going to get steeper. So the question is will OpenAI do the same? Will Google? Will Meta? The labs that take this seriously that build safety plans before they need them are going to be the ones that we trust with the next generation. The ones that don't are going to be the ones that cause the headlines that we all are terrified of. The harder news is that this is an arms race and it may not ever end. But for the first time, the defenders actually got a real head start and that matters more than most people realize.
But anyways, that is going to do it for this one. The space moves super fast and I have a lot of fun staying tuned to all of it. So, let me know what you guys think about this in the comments. And if you want to check out my free school community where we nerd out pretty hard about this kind of stuff every single day, the link for that is down in the description and I'll see you guys on the next one. Thanks everyone.