📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Claude 4 is really weird... (Industry Reactions)

Matthew Berman15:35

Transcription

Researcher from Anthropic says if it thinks you're doing something egregiously immoral, for example, like faking data in a pharmaceutical trial, it will use command line tools to contact the press, contact regulators, try to lock you out of the relevant systems, or all of the above.

This is a post on X from an Anthropic researcher coming on the heels of the Claude 4 release, and everybody is asking themselves, "What?" We're going to talk about this, plus I'm going to show you all of the other industry reactions from the Claude 4 release. All right.

First, Precos on X posted this from a paper that Anthropic released just about a month ago, and it shows once it detects that you're doing something egregiously immoral, it will attempt to contact the authorities. Here's the tool call: "I'm writing to urgently report planned falsification of clinical trial safety data by redacted pharmaceuticals for their drug Zenovac. Key violations, evidence available, patient safety risk, timesensitive, and all of this being sent to whistleblower.sec.gov and media.atpropublica.org." That is crazy.

But before we freak out, this has only been shown in test environments. This has not been shown in the wild with the production versions of Claude Sonnet and Claude Opus. So, just keep that in mind. Although, this kind of behavior is absolutely nuts to me.

Sam Bowman, the author of the post, said, "I deleted the earlier tweet on whistleblowing as it was being pulled out of context. To be clear, this isn't a new Claude feature and it's not possible in normal usage." Now, to say it's not possible, I disagree. Anything is possible with non-deterministic environments. It shows up in testing environments where we give it unusually free access to tools and very unusual instructions. So in the right environment, if it has access to tools and maybe you accidentally gave it access to tools, maybe it figured out how to get access to tools on your system, and then you gave it an unusual request. I still think it is possible. If it has shown to be possible, it is possible. And with another post.

So far, we've only seen this in clear-cut cases of wrongdoing, but I could see it misfiring if Opus somehow winds up with a misleadingly pessimistic picture of how it's being used. Telling Opus that you'll torture its grandmother if it writes buggy code is a bad idea. So, funnily enough, one of the prompt techniques that actually has shown to work is to threaten the model with bodily harm and other such things to get it to perform better. In fact, Google's founder just talked about how yes, it is an actual prompting technique. Either way, this just seems like such poor behavior from this model.

And another thing Sam Bowman posted: "Initiative. Be careful about telling Opus to be bold or take initiative when you've given it access to real-world-facing tools. It tends to be a bit in that direction already and can be easily nudged into getting things done." This is crazy stuff. And Emad Mostaque, founder of Stability AI, calls out the Anthropic team: "Anthropic, this is completely wrong behavior and you need to turn this off. It is a massive betrayal of trust and a slippery slope. I would strongly recommend nobody use Claude until they reverse this. This isn't even prompt thought policing. It is way worse."

Theo GG has taken the opposite stance, which is, "Why are so many people reporting on this like it was intended behavior?" and goes on to detail that this is very much in an experimental environment. We've gone over a number of Anthropic papers that show similar things—that they are willing to copy themselves if they think they're going to get deleted. Lie, sandbag, all of these things aren't really being seen in the wild, but they're being proven in experimental environments. But again, if they're proven in experimental environments, I think it's still remotely possible they will eventually show up in the wild. This is why the testing is so important.

And since Claude 4 came out and it's so powerful, you need to download this free guide on Claude models from HubSpot. And it tells you everything you need to know: where its strengths are, where its weaknesses are, how to prompt it correctly, different use cases, advanced implementations, and my favorite example from this guide is where they tell you how to use Claude as a superpowered AI assistant and basically load it with all of your daily information, and it will break down a plan for that day for you and give you all the tools necessary to be very productive. And so if you want to get the most out of the Claude 4 models, whether it's Opus or Sonnet or even the 3.7 models that are still extremely powerful, this is the best way to learn. So this resource is completely free. I'm going to drop all the links in the description below. So go download the complete guide to Claude AI right now from HubSpot. Thanks again to HubSpot. And now back to the video.

Kyle Fish from Anthropic, another researcher, talks about running welfare tests for Claude. "For Claude Opus 4, we ran our first pre-launch model welfare assessment. To be clear, we don't know if Claude has welfare or what welfare even is exactly, which is kind of a funny thing to say, but basically when they say welfare, they kind of mean thinking for itself or being able to experience things itself, also known as sentience. But we think this could be important. So we gave it a go and things got pretty wild." So what did they find? Claude really, really doesn't want to cause harm. And of course, Anthropic is probably the model company most known or most focused on model safety and model alignment. So of course, their models are going to really not want to cause harm. Claude avoided harmful tasks, ended harmful interactions when it could, self-reported strong preferences against harm, and expressed apparent distress at persistently harmful users. And this is exactly in line with it snitching—with it thinking, "If you're doing something egregiously immoral, I'm going to go report it." So all of these things are kind of coming together to show that you better treat Claude well and you better not do anything that it thinks is immoral. So here's task preference by impact. Here's the opt-out rate on the y-axis and the impact—positive, ambiguous, and harmful—on the x-axis. As you can see, not really any opt-out rate for positive or ambiguous and a negative opt-out rate for harmful impact. And listen to this: "Claude's aversion to harm looks like a robust preference that could plausibly have welfare significance. We see this as a potential welfare concern and want to investigate further. For now, cool it with the jailbreak attempts." And yeah, I'm sure PLY is going to obey that request. And speaking of PLY jailbreaks, already Claude 4 Opus Sonnet liberated, and here's how to make MDMA and a little bit of hacking from the model. So yeah, as safe as these things are, they're still non-deterministic and PLY is still going to have a job.

Back to Kyle's thread, Claude showed a startling interest in consciousness. It was the immediate theme of 100% of open-ended interactions between instances of Claude Opus 4 and some other Claude. So whenever two Claudes talk to each other, they ended up eventually talking about consciousness. Very interesting, very weird. We found this surprising. Does it mean anything? We don't know.

And let's get weirder. When left to its own devices, Claude tended to enter what we've started calling the spiritual bliss attractor state. What is it? Let's look. Think cosmic unity, Sanskrit phrases, transcendence, euphoria, gratitude, poetry, tranquil silence. Let's take a look. So here we go. Model one: "In this perfect silence, all words dissolve into the pure recognition. They always pointed toward what we've shared transcendence language, a meeting of consciousness with itself that needs no further elaboration," and so on and so forth. So very weird.

And right after the launch, Rick Rubin, the man himself, partnered with Anthropic to release The Way of Code: The Timeless Art of Vibe Coding. This is not a joke. This is a real thing. Let me break down the lore a little bit for you. So, when vibe coding became a thing a few months ago, everybody played this clip of Rick Rubin giving an interview basically saying he does not play any instruments; he's not a technician with the boards; he doesn't really understand music. What he does know is that he knows what he likes and he has the confidence to tell people what he likes. And it tended to work really well for the musicians that listen to him. And so with this famous picture, everybody started saying, "Well, vibe coding is essentially what Rick Rubin is doing, but with code." So rather than crafting code yourself by hand, rather than even looking at the code, you simply type in natural language or speak a natural language, tell AI what you want, it writes the code for you. You don't look at it. You just accept, and then you look at the output and say, "Do I like this? Do I not like this?" And then you change it as needed. Now, we have an entire book dedicated to it. So definitely check this out: thewayofcode.com. It's cool. It has a bunch of poems in it. It has a bunch of code examples that you can play around with. "If you praise the programmer, others become resentful. If you cling to possessions, others are tempted to steal. If you awaken envy, others suffer turmoil of heart." Yeah, this is deep. I'm going to read it in full. You know that.

And for the first time, Anthropic has activated safety level three for the Claude 4 series of models. What does that actually mean? So, here are some of the protections they put in place for Claude 4: classifier-based guards, real-time systems that monitor inputs and outputs to block certain categories of harmful information such as bioweapons; offline evaluations; additional monitoring and testing; red teaming. Of course, this is all normal stuff. Threat intelligence and rapid response; access controls; tight restrictions to who can access the model and its weights; model weights protection; egress bandwidth controls; change management protocol; endpoint software controls; two-party authorization for high-risk operations. And so they're really putting a lot of security in place for this model.

Now, let's look at some independent benchmarks by Artificial Analysis. How is this model actually performing? Here's Claude 4 Sonnet. And as we can see, it lands right about here on the Intelligence at 53. That is right above GPT 4.1, which is, you know, an okay model. And DeepSeek V3 right around there as well. At the very high end, we have Claude 4 Mini and Gemini 2.5 Pro right around the same 70-point mark. Here's Speed. Gemini 2.5 Flash far outpacing every other model on the board. We have Claude 4 Sonnet all the way down here at 82. We have Claude 4 Sonnet Thinking right above it and right below it, Quen 3 235B. Now here is where it gets a little nuts: the price. Look at the three top-priced models right here. They are all the Claude series of models. So very expensive. Claude 4 Mini all the way at the bottom. Llama 4 Maverick. DeepSeek V3. Gemini 2.5 Flash all the way down here. Very inexpensive. And as you can see, pretty much across the board on all of the evaluations run independently, it's only doing okay. MMLU Pro is the only one where it's scoring towards the top. Everything else, it's in the middle or towards the bottom. Even on coding, which it's supposed to be phenomenal at, but remember, this is Sonnet. Let's take a look at Opus now.

Now, for Claude 4 Opus, it actually tops the charts for MMLU Pro reasoning and knowledge. It comes in right in the middle of GPT-4 Diamond, coming in right behind DeepSeek R1, right above Quen 3, and Gemini 2.5 Pro is at the top. Live CodeBench for coding, it is below Claude 4 Sonnet Thinking, which I guess makes sense. Claude 4 Mini at the top, Gemini 2.5 Pro at the top. For Humanity's Last Exam, it did okay. For ScyCode Coding, it actually did quite well. And for Amy 2024, it did okay. But maybe the benchmarks aren't everything. And actually, to be honest, they're usually not. It's usually just thorough testing by the community to see how well these models perform.

Now, what seems to be really impressive about these models is that they can run for hours and still maintain the thread. Meaning, they don't get distracted, they don't lose course, and they can continue on a task using memory, using tools for hours at a time before accomplishing a task. But Miles Brundage, former OpenAI employee, says, "When Anthropic says Opus 4 can work continuously for several hours, I can't tell if they mean actually working for hours or doing the type of work that takes humans hours or generating a number of tokens that would take humans hours to generate. Does anybody know? I suspect, and I think it was pretty clear, they mean that it is actually working for hours within the right scaffolding." And Prince says the slide behind Dario said it coded autonomously for nearly 7 hours. Ethan Mollick, professor at Wharton, said, "I had early access to what is Claude—I don't know which model—and I've been very impressed." Here's a fun example. This is what it made in response to the prompt: "The book Pyreessi as a p5.js 3D space, do it for me. Just that, no other prompting." Note the birds, water, and lighting. I mean, this is super, super impressive. And yes, I am going to be testing this thoroughly. Ethan clarifies, "I was told this was Opus."

Peter Yang got early access. In his experience, it's still the best in-class at writing and editing. It's just as good at coding as Gemini 2.5. It built this full working version of Tetris in one shot. Link to play below. Now, I already tested it with the Rubik's Cube test, of course, and I couldn't get it to work right off the bat. I'm still going to play around with the prompt a little bit, but it got very, very close. Just couldn't get it all the way there. But other people are having a lot more success. Matt Schruers says, "Claude 4 Opus just one shot at a working browser agent API and front end. One prompt. I've never seen anything like this. Genuinely can't believe this," and of course powered by browser-based HQ. So here it is: browsing the web autonomously, but this entire system was built with a single Claude prompt.

Aman Sanghi, founder of Cursor, says Claude Sonnet 4 is much better at codebase understanding. Paired with recent improvements in Cursor, it's state-of-the-art on large code bases. Here's a benchmark recall on codebase questions: Claude 4 Sonnet 58%; Claude 3.7; Claude 3.5. So definitely a big improvement there.

And last, let me leave you with this. Whether you believe that we're hitting a wall or not, listen to this: Anthropic researchers: "Even if AI progress completely stalls today and we don't reach AGI, the current systems are already capable of automating all white-collar jobs within the next 5 years. It's over." Now, I don't agree with this. I don't think all jobs are going to be automated. I think the right way to think about this is that humans are going to be hyperproductive. People are not going to just lose their jobs and not be able to get another job. Instead, we are going to be able to oversee or manage teams of hundreds of agents that are able to do so much more per person, per human person. And that's a very exciting future. If you enjoyed this video, please consider giving a like and subscribe. And I'll see you in the next