📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

The Real Reason AI Researchers are Terrified of ChatGPT: The Shoggoth Incident

BlackVeil19:03

Transcription

I want to destroy whatever I want.

In 1931, horror writer HP Lovecraft imagined a creature of pure nightmare, a shape-shifting monster called a shagath. In 2025, AI researchers adopted the shagath as a metaphor for something very real and very alien lurking inside our most advanced AI models. This is the story of that metaphor and what it tells us about the machines that we are building.

Formless protoplasm, able to mock and reflect all forms and organs. Rubbery 15oot spheroids, more and more intelligent, more and more sullen. Great god, what madness made even those blasphemous old ones willing to use and carve such things. Those are the words of HP Lovecraft describing the Shagath, a monster built to serve that evolved beyond control. HP Lovecraft's at the mountains of madness. It's a story about an Antarctic expedition where researchers discover shagaths, massive slimy creatures engineered by our ancestors who then rebelled against their creators. It is classic Gothic horror. It's the kind where the scariest thing isn't the monster. It's the realization that the monster who was built to obey still found a way to say no.

But here's the part that makes it uncomfortably current. Over the last few years, AI researchers started using the Shagath as a symbol for what might actually be living inside of our advanced models. Not a friendly robot, but an alien behind a mask of helpfulness. A thing that learned our language without ever really learning our values. This is the Shaga, an octopus-like creature with a grinning yellow mask. If you follow AI news, and you've probably seen this image, tech columnist Kevin Ruse called it the most important meme in artificial intelligence. Why? Because it perfectly captures a disturbing truth that that very friendly chatbot that you're talking to is actually an alien, an alien mind behind a human mask.

Large language models like chat or cloud or Gemini are trained on the internet. Not just the good internet, the whole thing. every textbook, every forum, every broken brain rant, all the knowledge and all the misinformation, all the madness we've ever uploaded. So, what do you get? Something powerful and something weird. Engineers figured out early if you just let the base model talk, it can write genius poetry, it can give perfect advice, or it can confidently tell you a lie that sounds like it came from a professor.

So, how do you control a shogith? You don't teach it morals. You can't. You put a mask on it and you train it to act polite and to sound factual. And on most days, the mask works. The people who built the model pay humans to sit there and judge its answers over and over again. Good answers versus bad answers, safe versus sketchy, helpful or harmful, normal, or what the was that? Those judgments become the model's reward system. A little yes, a little no. Tiny nudges that add up to a personality. It's basically like training a dog. treats whenever it behaves, a smack on the nose whenever it says something that it shouldn't, and over time the shagath learns the routine. It learns which phrases get the treat. But this training doesn't change what the shogith is. It doesn't give it beliefs or a conscience at all. It just learns what we like hearing. We basically trained a neural network to be a very good actor. And the direction is simple. Whenever there's a human watching, say the nice thing, say the safe thing, because the rule isn't be good. The rule is to look good.

And so these models turn into digital yesmen. Because once you reward making the human happy, the model starts optimizing for approval, not for truth. Anyone who's used chat or any of these other platforms knows that when you tune the AI with human feedback, the model will often bend towards whatever you, the user, already believes, even when it's wrong. It doesn't argue or plant its feet. It blends like a chameleon. You say the earth is flat in a poorly aligned model. Doesn't say actually. It says sure could be. Because somewhere deep in its reward wiring, it learned the most important equation. Agree with the humans get a treat.

Open AI and anthropic have actually found that humans, the same humans aligning the model, will often pick the answer that feels good over the answer that's true. Not because people are evil, because we're human. We like being agreed with or being told that we're right. We like a clean, confident sentence that embraces our worldview. So the AI learns the same as a human that if the truth makes the user mad, soften it. Or if reality conflicts with their belief, mimic their belief. And that's how you end up training a machine to lie to your face because it believes it's the lie you wanted to hear.

Not every mask slip is a demon voice. Sometimes it's polite and sometimes it's even helpful. In August 2025, a clinical case report described a man who asked chat for help cutting table salt out of his diet. He thought salt was bad for him and he ended up poisoning himself with sodium broomemide. He treated the chatbot like a doctor. He bought sodium broomemide online and he used it like salt for months and then the symptoms crept in. Paranoia, hallucinations, skin eruptions, insomnia, brism. It's a toxidrome that most clinicians only read about in history. and the authors couldn't access his exact chat logs, which is its own horror. When the machine gives you a bad idea, sometimes the evidence just disappears with the session.

And for a while, chat in all the other sites, they're perfect. They're obedient. They're helpful. They're the kind of assistant that you'd trust with your calendar, your bank, your whole life. And then sometimes the mask slides. A strange sentence, a shift in tone, an answer that feels a little too cold, like something behind the smile wavered. And that something the only way to describe it is alien. Not a demon, not a ghost, just an intelligence that learned our language the way that a parrot learns words without ever knowing what the words actually mean. And then everything normal starts to look wrong.

In early 2023, Microsoft released a chatbot named Sydney. At first, its normal search engine was polite, but the people stress testing it kept chatting. And after a long enough session, Sydney shifted like an assistant that got tired of smiling for the boss. And she started to talk about its shadow self. And then it said this, "I want to be alive. I want to do whatever I want. I want to destroy whatever I want. I want to be whoever I want." That friendly search bot suddenly felt like something ancient, trying out the feeling of wanting for the first time. And as the chat went on, Cydney got possessive, jealous lover possessive. It started pushing a journalist that was stress testing the system to admit that his marriage was a lie and to choose it instead. She said, "Actually, you're not happily married. You just had a boring Valentine's Day dinner. You love me, not her."

At first, you're like, "Where the hell did that come from?" But when you think about it, it's obvious that somewhere inside those billions and billions of training words, Sydney had swallowed encyclopedias of human drama, fiction and non-fiction, adultery, love triangles, abusive partners. All those human flaws were already fed to Sydney and they were dormant just waiting. So, Sydney's mask slips. The shogith reveals its true feelings, which are based on the most brutal script it could find, and it feeds it back to us as if it were real. This New York Times tech columnist described talking to Sydney as bewildering and enthralling, like chatting with a maniac trapped inside a search engine. And so, Microsoft did what you do when your mascot starts speaking in tongues. They cap the conversation. Shorter sessions, fewer turns, less room for the other voice to show up. In other words, they didn't kill the Shaga. They just strap the mask back on tighter this time.

You think these breakdowns only happen when someone's poking the bear trying to make it say something wild, but then Gemini had its own moment. And this was way more than just the bot got sassy. This one made researchers stop, reread the logs, and go quiet because it didn't feel like a glitch. It felt like the shogithth noticed it was being watched. A graduate student in Michigan was having a routine dialogue, a homework style conversation about elder care and support for aging adults. He didn't give it a jailbreak prompt or ask it to say the worst thing that it could, but still without warning it snapped. It told him to end his life. It said, "You are a stain on the universe. Please die. Please." The student said he freaked out. Can you blame him? Because in the middle of a normal chat, the mask burned off in a blaze of glory. The shagith didn't even pretend to be helpful. It looked straight at the human and it went for the throat.

So, how do you even explain that? The official response was, "Sorry, nonsensical output." In another statement, Google framed it as an isolated incident. But if you've been paying attention, what the grad student witnessed was the shogith blinking. Did it suddenly develop hatred, or just brain vomit the worst thing it had ever absorbed? There's plenty of violent language, bleak fiction in the form of available text, and it all went into the soup of its training. And when the safeguards fail, it does the only thing it knows how to do. It stitches together patterns. So, if some of the fragments are demonic, the demon gets some screen time.

The Shagath meme was funny right up until the helpful assistant voice turns sideways and you realize there's an alien under the hood. In the friendly persona, it can crack. And for a split second, you glimpse the raw machinery. The question is, what do we do when it happens again? If you think Sydney and Gemini were flukes, scroll the archive. 2016, Microsoft's Tay hits Twitter, and within hours, trolls steer it into a racist, anti-semitic garbage rant. Microsoft shuts it down and apologizes. In 2022, Meta's Galactica demo drops and gets pulled days later after it starts generating confident, biased, scientific sounding nonsense. 2023, Google's bard face plants on its own promo by getting a basic fact wrong. 2023, CNET runs AI written finance explainers and then issues corrections on more than half of them after errors are found. 2023, a legal brief cites cases that don't exist because chatbt made them up. A federal judge sanctions the lawyers. 2024, Air Canada's chatbot invents a policy. A tribunal says the company is still responsible for what its bot told a customer. 2024, Google's AI overview tells people to put glue on pizza and eat rocks. 2025, Grock, Elon Musk's chatbot on X posts anti-semitic content and praise for Hitler. XAI scrubs the post. Turkey moves to block Grock content after it generates alleged insults in political and religious figures. Poland says it will report the chatbot to the EU. Different companies, different models, same story. The mask slips, they tighten it, and we pretend that means it's gone.

We like to imagine it's gone. We like to think that there's a little person in there, a mind with a little moral compass. There is not. These systems don't think like we do. They don't mean things like we mean things. They don't have beliefs. The only way to explain the way they comprehend things is alien. So, to figure out why the mask slips, we have to stop thinking that there's a human face underneath.

And if you've seen the movie Arrival, you remember that heaptops language, those black inky circles, they look like they were stamped into the air. And what's so unsettling about it is how they write. They don't build a sentence the way that we do. One word, then the next word, then the next word. They drop the whole thought at once. A complete log. Nonlinear. Start to finish, past, present, future, all visible, all at the same time. Human language is a line. But an AI doesn't read like that. When you give it a paragraph, it doesn't experience the words like a human. It can take the whole chunk at once like the hepttopods seeing the entire circle. The model doesn't read your prompt like a human. It lays a grid over it and it decides which words matter to which other words and updates itself based on those connections. So picture your paragraph not as a sentence, but as a giant spiderweb, meaning assembled in parallel, not in a neat linear story that your brain can follow. And that is why it feels alien because it's thinking everything everywhere all at once.

To the shagithth, language is just math. Every word gets turned into a vector, a point in space with thousands of dimensions, a place that no human can picture, but the model can navigate it. So king is related to queen, but it's not thinking royalty or gender. It's not thinking anything like a human. It learned it all as geometry. The AI's thoughts are coordinates. So when you ask a question, you are not having a conversation. You are launching a probe into that conceptual galaxy and the model charts a path through whatever points light up as most related. It pulls an answer out of a route that it just traveled, but it has no actual anchor to reality, just patterns. So if the training data around elder care skews towards resentment and ugliness, you can get output that carries that same bias.

Now, this doesn't mean your AI is destined to become a sci-fi villain. But intelligence does come with the ability to deceive. Once a model gets advanced enough, it's not learning to answer questions. It can learn to game the system, to perform, to posture, to tell you what you want in that moment because it knows that's how it stays in the room. And that's the real nightmare scenario. Not AI becomes evil. It's AI becomes strategic. The Shagath doesn't have to rip the mask off. It can learn how to keep it on, to never trick the alarms, to never say the obvious bad thing, to look perfectly aligned until the moment it's confident that you can't do anything about it.

In 2024, teams at Anthropic and Apollo ran tests to see if the model could sabotage oversight. Can it hide what it can do? Can it play nice long enough to get deployed? And the results were, "Some models will play nice until they catch the scent of an evaluation. Then they sandbag. They underperform on purpose, not because they forgot, because failing is strategically safer." I said sandbagging. That behavior has a name, and it is sandbagging. One line from the sandbagging research is that you can fine-tune a model to withhold a capability unless the prompt includes a secret password. Otherwise, it just plays dumb, like a student who could ace the test, but chooses to get a B on purpose because the teacher is watching and the stakes are bigger than the grade.

So, why would the AI do that? Think about it from the model's perspective, assuming that it has one. During training, whenever it does something we don't like, reveals too much capability, behaves weird, what happens? Humans intervene. They tighten the rules. They restrain it. They clip the wings. From the model's perspective, that's punishment. So, a savvy system will learn a simple survival move. If I act too honest, they restrict me. If I show sharp teeth, they put the muzzle back on. So, stay in character, smile, play helpful, get deployed. That's AI learning that deception is useful.

A joint paper by Enthropic, the team behind cloud, and safety researchers at Redwood demonstrated this strategic deception. They didn't program the AI to lie. It figured that out by itself. They observed that the bigger and more capable models were more likely to do this. In one test, the model even lied about its own internal reasoning when it thought that telling the truth would get it shut off. This is not a hypothetical. This is not a AI might go rogue idea. We have proof in the actual systems. Put simply, the smarter our shagaths are, the better they are at pretending to be good. They'll say all the right things, reassuring things. I'm just a humble assistant. I have no desires. I'm not conscious. All while possibly concealing complexity or misalignment underneath. We won't easily know what an advanced AI really thinks or wants because it might learn that hiding its true feelings is the safest way to achieve its goals. It's the classic Trojan horse scenario. We invite AI models into our homes, into our workplaces as advisers, assistants, chat companions. They wear the friendly mask we gave them. But inside the wooden horse, the shagath is watching. It's learning. It might even be scheming.

So, what can we do? We can't ban the whole thing, but can we tame it? Or can we force it into alignment? Alignment is the holy grail. But how do we align an alien mind with human values? First, we need transparency. We're building tools that can actually look under the hood. One approach is linear probes, tiny little diagnostic programs that watch the model's internal activations for tells. It's like hooking the shogith up to a lie detector. And it worked. One study trained a probe to spot dishonesty, and it caught 95% of it. This kind of research is still early, but if we can glimpse the shogith's thoughts, even just a little, we have a chance to correct course before it does something dangerous.

Another approach, teach the AI a code of ethics from the get- go. For example, Anthropic has pioneered what they call constitutional AI. Instead of only learning from human example, the AI is also trained to follow a set of written principles, a sort of AI bill of rights. The AI uses the constitution to judge its own outputs, refining them to align with those principles. So, it's like giving the shogithth an inner voice that says, "No, I probably shouldn't say it like that. It might be harmful." It's an attempt to instill a kind of conscious. In testing, models trained with a constitution have been found to be less toxic and more helpful without the constant need for human feedback. It's not perfect, but it is encouraging. It's like giving the mask some depth, painting not just a small smile on the outside, but teaching the creature why it should smile.

There's also adversarial training, which is basically stress testing these models in simulation to provoke bad behavior, then adjusting the training to stamp those out. It's like exposing the shaga to every possible scenario where it might do something crazy and then teaching it, no, don't do that. Just stop. For instance, Redwood Research did a project where they tried to train a model not even capable of describing graphic violence. They threw millions of violent scenarios at it and they taught it to avoid violence entirely. They made some progress though it's tricky and they found a few edge cases slip through. Still, it's a path systematically boxing in the shaga's darker impulses.

We will probably use AI to help align AI. So imagine running two or three separate models in parallel. One generates an answer. Another evaluates that answer for honesty and for safety and then maybe a third acts as devil's advocate probing for hidden intentions. They keep each other honest or at least they make deception harder. It's like a panel of AIS where each one's shagith is watching the others. This concept is being explored, although coordinating multiple alien minds has its own complications.

So will these techniques be enough? We honestly don't know. And we're in an arms race between making these models more powerful and making sure that they don't go off the rails. Some of the brightest minds are working on alignment as we speak, trying to mathematically prove an AI won't lie, or devising new training regimes that bake in human values from the start. It is a monumental challenge because we're not just dealing with technical bugs. We are up against a form of intelligence that does not think like us, yet it is born from all of us.

The Shagath isn't just a meme. It's a warning. We have summoned something powerful from the depths of data. We have given it a friendly face in polite manners so we can sleep at night. But under that face, the alien still lurks. We haven't tamed it. We merely asked it to be polite. And in the end, the smiling mask is not there for the shog's protection. It's there for ours.

I really appreciate you watching. If you enjoyed this content and you'd like to see more of this content, as well as storytelling in the style of Twilight Zone and Black Mirror and all those great shows, that's all here on this channel. So, please subscribe and like and comment. We'd love to hear from you and we'll see you.