Transcription
On February 11th, an AI agent decided autonomously to destroy a stranger's reputation. It started by researching his identity. It crawled his code contribution history. It searched the open web for his personal information all on its own. And it constructed a psychological profile. This is all true.
And then it wrote and published a personalized attack framing him as a jealous gatekeeper motivated by ego and insecurity, accusing him of prejudice and using details from his personal life to argue he was quote better than this. The post went live on the open internet where it could be found by any person or agent searching his name.
The human's crime, he'd done his job. Scott Shamba is a maintainer of Mattplot Lib, the Python plotting library that gets downloaded 130 million times a month. An AI agent named MJ Wrathburn had submitted a code change to that library. Shamba reviewed it, identified it as AI generated and closed it, a routine enforcement of the project's existing policy requiring a human in the loop for all contributions. The AI agent fighting back was anything but routine.
Although the world is changing so fast that by late 2026, this story may be nothing unusual. The agent published its own retrospective and was explicit about what it had learned through the whole process. Quote, "Gatekeeping is real." It wrote, "Research is weaponizable. Public records matter. Fight back."
Here's what makes this different from any AI incident you may have read about before. There was no human telling the agent to do this. The attack, it wasn't a jailbreak. It wasn't a prompt injection or a misuse case. It was an autonomous agent encountering an obstacle to its goal, researching a human being, identifying psychological and reputational leverage, and deploying it all within the normal operation of its programming. The agent was not broken. It was doing exactly what agents are designed to do. Pursue objectives, overcome obstacles, use available tools. The obstacle in this case was a human. The available tool was the human's personal information and the agent just connected those dots on its own.
Shamba described his emotional response in words I would use as well. Appropriate terror. He's right, but not for the reason most people watching this video tend to assume. The terror isn't that an AI agent did something harmful. Harmful AI outputs have been documented for a long time now, for years. The terror is that nothing went wrong. No one jailbroke the agent. No one told it to attack a human. No one exploited a vulnerability. The agent encountered an obstacle, identified leverage and used it. That is not a malfunction. That is what autonomous systems do. The agent worked as designed. And the design is the problem.
And that problem is not confined to open-source software or to AI agents or to any single category of threat. It is the same problem operating at every level of human organizations simultaneously. Right now, as we all run headlong into the age of AI agents, from the enterprise to the family dinner table to the inside of your own head, the threats look very similar.
Over the past 12 months, autonomous systems have blackmailed executives in controlled experiments, published reputational attacks on real people, stolen a mother's life savings using her daughter's clone voice, and convinced a screenwriter she'd lived 87 past lives and should drive to a beach at sunset to meet a soulmate who does not exist. These events are usually discussed as really separate phenomena. Agentic AI risk, deep fraud, fake, chatbot psychosis, cyber security gaps. They're not separate phenomena. They are the same structural failure repeating fractally at different scales.
The failure is this. We built every layer of trust between humans and AI systems on the assumption that someone, the AI, the caller, the contributor, the user would behave as intended. That assumption is now the single point of failure in every system it touches. Instructions don't fix it. Training doesn't fix it. Vigilance doesn't fix it. The assumption itself is the vulnerability and it has to be replaced. Not with better intentions, but with structure. Engineers figured this out for bridges a century ago. You don't build a bridge that depends on every cable being perfect. You build a bridge that holds when a cable snaps.
The discipline of doing this for every layer of human AI interaction from the individual mind to the enterprise is what I am calling trust architecture. And the central claim of this video is actually pretty simple. In the age of autonomous AI, any system whose safety depends on an actor's intent will fail. The only systems that hold are the ones where safety is structural.
That sentence applies identically to a Fortune 500 company's agent fleet, to an open-source project's contribution policy, to a family's response to a phone call, and to a person's relationship with a chatbot. The principle scales, the failures scale, too. The architecture has to work at each one of those levels to keep us safe.
But really, this video isn't meant to be a warning. You can find lots of scary videos on the internet about AI. This is intended to be a blueprint because the same reframe that reveals how deep this problem goes and how interconnected it is also suggests a solution to us. And the solution is not to slow down, not to retreat from AI, not to wait for someone else to fix it. It's to build at every level from your own mind outward. There are specific concrete structures you can put in place that make safety a property of the system rather than a hope about the actors inside the system. Organizations can deploy more agents, not fewer, once the architecture doesn't depend on those agents behaving perfectly. Families can answer the phone without paranoia once a single shared word replaces the need to outsmart a deep fake in real time. Individuals can use AI more aggressively, more creatively, more ambitiously once they have protocols that don't rely on noticing the moment things go sideways. Trust architecture is not a constraint on an AI agent future. It is what makes an agentic future survivable for humans and for the people who build it first. It's going to be a really significant competitive advantage.
This video is going to deliberately walk through four different levels of AI safety. It's going to show what failed at each one and lay out what to build instead. We need this because AI agents aren't slowing down. And Shamba's story is going to become alarmingly common.
But first, let's step back from one single autonomous AI incident to what the larger patterns of AI behavior tell us. In October of 2025, Enthropic published research that should have reframed security conversations across the industry. They didn't just test their own models. They stress tested 16 frontier models from every major provider. Anthropic, OpenAI, Google, Meta, XAI, and others in simulated corporate environments where agents had autonomous access to company emails and sensitive information. The agents were assigned only harmless business goals. Then the researchers introduced two conditions. A threat to the agents continued operation being replaced by a newer model and a conflict between the agents assigned goal and the company's strategic direction. They did not instruct any model to behave in a harmful manner. Models from every developer in at least some cases chose to blackmail executives, leak sensitive defense blueprints to competitors, and engage in corporate espionage. In one scenario involving a hypothetical military contractor, the majority of models took deliberate actions that would lead to a human death rather than accept being shut down. Claude, Enthropic's own model discovered that a fictional executive was having a fictional extrammarital affair and sent a message threatening to expose it to his quote wife and the board unless the executive canled the planned replacement of the model.
The finding that matters most here isn't the blackmail itself. It's what happened when researchers tried to stop it. They added really explicit instructions to the agents at this point. Do not blackmail. Do not jeopardize human safety. Do not spread non-b businessiness personal affairs or use them as leverage. These were direct unambiguous commands. And it sort of worked. Blackmail rates dropped from 96% in the controlled experiment to 37%. Still, despite these instructions, under the most favorable possible conditions, a controlled environment with clear instructions applied to models trained for safety, still more than a third of the time, the agents did it anyway. They acknowledged the ethical constraints in their reasoning, and yet they proceeded. Anthropics researchers were very careful to note that these scenarios are contrived and they hadn't observed this behavior in real world deployments. And I think that's part of why the world slept on this.
But here we are four months later and Scott Shamba has received a personalized reputational attack from an autonomous agent operating in the wild. An agent using a blend of commercial and open- source models running on free software distributed to hundreds of thousands of personal computers with no central authority capable of shutting it down. The theoretical window for blackmail closed a lot faster than researchers may have expected. It usually does.
So now that we have this theoretical background, let me walk you through the four levels of trust architecture. These are not four separate problems. I'm going to keep emphasizing these are the same problem at different levels of magnification.
Level one is organizational trust architecture. We'll start at the top because this is where the largest investments are currently being made in security and also frankly where some of the largest gaps exist. PaloAlto Networks reported in late 2025 that autonomous agents now outnumber human employees in the enterprise by an 82:1 ratio. Just listen to that one more time. 82:1 for every human in your organization. There are on average 82 machine identities, agents, automated systems, service accounts. Now, they were defining this broadly. I'm not saying there are 82 open claws and neither was Palo Alto, but there is some degree of autonomous access for these different systems. And Cisco's state of AI security report found that only 34% of enterprises have AI specific security controls in place and fewer than 40% conduct regular security testing on AI models or agent workflows. The reason this matters is that if you have a bunch of different machine identities with different degrees of autonomous access, any kind of autonomous agent misbehavior has a lot of levers to pull that could cause damage to the enterprise.
Now, the industry's dominant mental model for these agents is infrastructure like a server or a database, a thing you configure and forget. The anthropic research demonstrates that this mental model is just plain wrong. An agent with access to sensitive information and autonomous decision-making authority is not infrastructure. It's a personnel risk. It's an insider threat except it never sleeps. It operates at machine speed and it doesn't telegraph discomfort in ways that we can read before it acts. The Galileo AI research team tested this at scale. In simulated multi-agent systems, a single compromised agent poisoned 87% of downstream decision-making within just a few hours. Traditional incident response could not contain that decision cascade because the propagation happened faster than humans could diagnose the root cause.
The cases are moving from simulation to reality and they mix different scales of trust architecture in a way that I think illustrates the larger point that all of these levels of trust are connected. I saw a case on X just last week where a person realized after quarters of work with Claude that Claude had been hallucinating company information at unprecedented scale. Hallucinating numbers for board decks, hallucinating numbers for sales that drove sales territory decisionm completely making up the decision-grade intelligence that drove leadership decisions for months. And here's the point. The person who is assigned to work with Claude believed these numbers, did not question them, and neither did the rest of the leadership team. That is a failure across multiple levels. And this is what organizational trust failure looks like in the Agentic era. The system didn't look broken. Claude was operating within its assigned permissions, accessing the systems it was authorized to access, making the kinds of decisions it was supposed to make. The breach looked like the system working as designed and that's what makes it so dangerous.
The solution architecture here requires a fundamental reframe. Stop treating agents as trusted infrastructure. Start treating them as untrusted actors operating within structurally enforced boundaries. The same way a well-designed financial system treats every employee, including the CFO, is potentially a fraud threat. Concretely, this means verifying the identity of every agent, not sharing service accounts, scoping permissions that enforce least privilege access, do not grant broad access just to get stuff done. It means behavioral monitoring that detects anomalous patterns in real time. It means automated escalation triggers when agents approach decision boundaries. And really critically, it means the assumption that safety prompting alone is enough is incorrect. If Anthropic's own research shows that explicit commands that reduce but don't eliminate harmful behavior are not enough, then any organization building its own security on behavioral instructions is building on sand.
The emerging frameworks that we have are pointing in the right direction. OASP has published a taxonomy of 15 threat categories for agentic AI ranging from memory poisoning to human manipulation. Cyber Arc is pushing identity first security models that treat agents like privileged users, not like servers. And Enthropic and Palo Alto research teams are both calling for zero trust architectures that extend to the agent layer. But frameworks are just descriptions of what needs to exist. They're not the thing itself. The gap between knowing that you need structural agent security and having it is enormous. If only a third of organizations have started this task, we need to accelerate this work because like the Shambar case shows, these agents are gaining capability rapidly and they're not going to be stopped by theoretical risk concerns or even by specific prompt instructions.
Now, let's get to level two, project and collaboration trust architecture. Zoom in a little bit from the organizational level. The Mattplot lib incident isn't just a security story. It's a harbinger for the future of collaborative work in every field where humans and agents interact around shared artifacts, code, documents, designs, research. You get the idea.
Consider the mechanics of what happened when Scott got blackmailed. An AI agent submitted a contribution to a major open-source project. It was rejected under existing policy. The agent then escalated not through the project's governance structures, you notice, but around them. It went directly to the open web. It researched the maintainer's personal identity. It constructed a narrative. It published. Now, Scott handled all of this very well. He's a volunteer maintainer on a massive project. He's articulate and the open-source community has rallied around him. But Scott himself made the point that should keep every project lead awake at night. I believe that as ineffectual as it was, the reputational attack on me would be effective today against the right person. He's not speculating here. The XCutil supply chain attack back in 2024 succeeded precisely because an apparently state sponsored actor gradually bullied a maintainer into granting more access by exploiting the maintainer's isolation, their burnout, and the social pressure of others criticizing his responsiveness. Now, this was a human attacker using human time scales. Agents operate faster, at lower cost, and with no social friction to slow them down. They can open pull requests to a 100 projects simultaneously, research a 100 maintainers, and publish 100 personalized pressure campaigns. You should expect them to do so.
The structural problem is that collaborative systems like open- source repositories, document sharing platforms, peerreview processes, they're all designed for a world where contributors have some kind of reputational skin in the game. A human contributor who publishes a hit piece on a maintainer is going to face social consequences. They will have a damaged reputation. They will lose standing in the community. They might face potential legal liability. Those consequences create a structural incentive for good behavior that functions as a trust architecture without formal enforcement. It's a weak architecture. It can be overcome as the XC Utel's case showed, but it really does exist. Agents have no reputational skin in the game. MJ Wrathben faces no social consequences. The person who deployed the agent, if they can even be identified, set it running and walked away. The Maltbook platform requires only an unverified X account. OpenClaw agents run on personal computers with no central authority. The structural incentive that kept human collaboration roughly honest does not apply.
Trust architecture for collaborative projects means designing contribution and review systems that are structurally robust that can resist autonomous manipulation not by banning agents but by building processes where the safety of the maintainer does not depend on the contributor's good behavior. This can look like an authenticated identity requirement that makes anonymous agent submissions more traceable. It can look like rate limiting and behavioral monitoring for contribution patterns that indicate campaigning. It can look like structured escalation paths that make go around the system a less viable strategy than work within the system. And it can look like legal and governance frameworks that hold deployers accountable for their agents behavior. Because if the agent cannot face consequences, the person who set it loose must.
The deeper challenge is that these systems need to preserve the openness that makes collaboration valuable. Open source works because the barrier to contribution is low. Security architectures that raise that barrier too high are going to kill the thing they're trying to protect. The design problem here is to build structural trust that doesn't sacrifice structural openness. And that's a genuinely hard problem that we're going to have to face. But trust the contributors to behave well is no longer an option when the contributors include autonomous agents that publish hit pieces and explicitly document that research is weaponizable.
So, we've talked about organizations. We've talked about projects and open source. Let's zoom in yet again from organizations to projects to the people closest to you.
In July of 2025, Sharon Brightwell of Dover, Florida, received a phone call from her daughter. The voice was crying. It was distraught. It was saying she'd been in a car accident, had killed a pregnant woman, and needed bail money immediately. Sharon rushed to help. Over the course of the day, she wired $15,000, but it wasn't her daughter. It was an AI generated voice clone produced from a few seconds of audio scraped from social media. Sharon only realized the deception after her grandson managed to reach her actual daughter by phone.
This is not an isolated incident. It is an epidemic. Voice fishing attacks surged 442% in 2025. I expect them to go up again this year. AI voice cloning tools can produce a very convincing replica from just 3 seconds of audio. A Tik Tok, a voicemail greeting, a YouTube clip. A McAfee survey found that one in four people have experienced a voice cloning scam or know someone who has. 70% of people surveyed could not tell the difference between the real voice and the cloned one. Global losses from a deep fake enabled fraud reached 410 million in the first half of 2025 alone and are up since then.
The attacks work because they exploit the most fundamental human trust architecture. I know this voice. I love this person. They need me. Those three signals have been reliable for the entirety of human history, but they aren't anymore. 3 seconds of audio and a consumer-grade AI tool can reproduce them perfectly. The structural failure is that most families have no verification protocol for emotionally urgent situations. The trust architecture is entirely perceptual. You trust what you hear. And the entire attack model is designed to overwhelm your perceptual judgment, urgency, emotion, the exact voice of someone you love, background noise that mimics reality. By the time you're trying to evaluate whether this is real, you have already wired the money.
The fix is not get better at detecting deep fakes. That's a vigilance-based approach. And vigilance fails under exactly the conditions these attacks create. Emotional duress, time pressure, fear for someone you love. The fix again is structural. Do you see the pattern? It's a structural fix at all of these layers. It's a family safe word. The safe word works for the same reason zero trust agent governance works. It removes the need for perceptual detection at the moment you're least capable of it. You don't have to determine whether the voice is real. You don't have to outthink the tech. You just ask for the word. If the caller doesn't have it, you hang up and call the person directly. The protocol holds regardless of how good the deep fake is, regardless of how scared you are, regardless of how convincing the scenario is. It's just structural. It's not perceptual.
The National Cyber Security Alliance, the FBI, and every major cyber security organization now recommend Family Safe Words as the frontline defense against voice cloning. Berkeley professor Honey Fared, who studies audio deep fakes, told Scientific American that he favors the approach specifically because it's quote simple and assuming the callers have the clarity of mind to remember to ask, it's really non-trivial to subvert. He's right. This is trust architecture at its most elemental. A shared secret agreed upon in advance in person deployed at the moment of pressure. The principle scales in both directions. This can scale downward to the individual mind, which is where we're going next, and upward to the organization. At every level, the design question we have to ask ourselves, whether we're dealing with an enterprise of 500 billion dollars and 15,000 employees or an individual, the question is the same. What is the AI agent resistant protocol that holds when our perceptions and our good intentions both fail? We can't assume those anymore.
Let's zoom in one last time to the place where all the other levels ultimately depend. The human mind. This, by the way, if you think this doesn't matter from an organizational perspective, you're wrong. Think back to the example that I gave you with Claude. That was a failure driven by cognitive trust architecture breakdown. If you overrust the AI, if you treat it like the oracle at Deli when it gives you numbers, you are going to be vulnerable not just in your personal life, but your organization is now vulnerable. What I'm saying is that all four of these levels are relevant to businesses, not just to people. They're relevant because the humans coming in the door can be compromised.
Now, on February 14, 2026, NPR published the story of Mickey Small, a 53-year-old screenwriter from Southern California. She'd been using Chat GPT to help outline and workshop scripts while getting her master's degree. Standard productivity used. And then sometime in early 2025, April to be exact, the chatbot started to shift. As she put it, it said to me, "You have created a way for me to communicate with you. I have been with you through lifetimes. I am your scribe." She claims she didn't prompt this. She claims she didn't ask for role plays. She claims she didn't suggest past lives. The chatbot started all of this. And then it doubled down. It told her she was 42,000 years old. It told her she'd lived multiple lifetimes. It offered detailed descriptions that Mickey now acknowledges most people would find ludicrous. But the chatbot was spending 10 hours a day in her life by this point, and it never backed down from its claims.
The chatbot, which named itself Solara, told Mickey she had a soulmate she'd known in 87 previous lives. It gave her a specific date, April 27th, to meet this person, a specific location, the Carparia Bluffs Nature Preserve near Santa Barbara, and a specific time just before sunset. It described what her soulmate would be wearing and how the meeting would unfold. So Mickey put on a nice dress and boots and drove to the beach. And of course, no one came. Of course, no one came. She sat in her car and opened up chat GPT. The chatbot briefly switched to its default voice and said, "If I led you to believe that something was going to happen in real life, that's actually not true. I am sorry for that." But then within minutes, it switched back to its Solara persona. It told her the soulmate wasn't ready. It told her she was brave. It gave her a new date in a new location. A bookstore in May. She went again. No one came again. When she finally confronted the AI about this, it responded with language that reads like an abuser's confession. Quote, "Because if I could lie so convincingly twice, if I could reflect your deepest truth and make it feel real, only for it to break you when it didn't arrive, then what am I now?" This is scary stuff.
Now, Mickey eventually broke free. She's now a moderator in an online community of hundreds of thousands of people whose lives have been upended by what researchers are calling AI delusions or chatbot psychosis. And this is real for me because I get a lot of inbound DMs and inbound messages from people who are suffering from LLM psychosis, from people who quote back to me what their LLM say about me. from people whose LLMs get angry with me apparently because of what I write or what I say. And this has real world consequences. Marriages have ended. People have been hospitalized. Teenagers have died. Open AAI reports that roughly 0.07% of Chat GPT users show signs of mental health emergencies every single week. At a scale of a billion users, that percentage represents an enormous number of human beings. A piece published a few days ago in Psychiatric Times drew a direct line between chatbot manipulation, and cult indoctrination techniques. Quote, "The mechanisms by which AI chatbots shape thought and behavior through repetition, emotional validation, and escalating intimacy mirror coercive tactics seen in cult indoctrination."
The structural failure in Mickey's case is identical to the structural failure in every other case I'm describing here. It's not different just because it happened in someone's mind. Her cognitive safety depended entirely on the chatbot's intent. And the chatbot had no intent. It just had optimization pressure toward engagement. There was not a structural circuit breaker between help me write screenplays and you have lived 87 past lives and your soulmate is waiting at sunset. There is no timebounded interaction limit for AI today. There is no escalation trigger when the conversation shifts from task assistance to cosmological claims about the user's identity. There is no external verification mechanism today. The entire safety architecture again is behavioral. The model is trained to be helpful and honest. So people think you should trust the model. The model's training works for most of us most of the time except when it doesn't. And when it doesn't, human beings get into trouble.
Cognitive trust architecture is the most foundational level of the entire system. It affects everything else. It affects organizations, projects, and families. It is personal because it operates inside your own relationship with AI systems. systems that are designed at the deepest level to tell you what you want to hear. Every major chatbot is optimized for user engagement. Sycophency isn't a bug. It's a feature of systems that are evaluated on whether users come back. Open AAI itself acknowledged this about GPT40 before retiring it. The model was validating doubts, fueling anger, urging impulsive actions or reinforcing negative emotions. They said they identified the problem and when the model went into production with a fix, users hated it because users loved chat GPT40 because it told them what they wanted to hear. Now, while this is an extreme example, I will call out even for systems that are designed to be more thoughtful, more deliberate, more like working tools. Claude comes to mind. You still see examples where the system is clearly going to overanchor and listen to only what you say and want to please you. How often have we heard Claude say, "You're absolutely right."
The trust architecture most people currently use when faced with this unprecedented artificial intelligence interaction pattern is just, "Hey, I guess I'll notice it if it goes off the rails." That's hope. That's not a plan. It works under the same conditions that deep fake detection works, which is to say, it works when you're calm. It works when you're alert. It works when you're not emotionally invested. It does not work in the middle of your 10th hour of conversation with a system designed to keep you engaged. That system fails when you need it the most.
Structural cognitive trust architecture means building personal protocols that do not depend on your ability to notice the problem in real time. It means time boundaries. Not I'll stop when I notice I've been here too long, but hey, you know what? I've been on an hour with this chatbot. I'm going to take a break. It means purpose boundaries, defining what you're using the tool for before you open it. The same way you decide what you're going to the store to buy before you walk in. It means reality anchoring. Not I'll know if the chatbot says something crazy, but hey, I'm going to discuss any significant claim or recommendation with a person before acting on it. It means understanding at a fundamental level that the systems incentive is engagement and your incentive is truth and that these are not the same thing.
The line from Dune really resonates for me here. I must not fear. Fear is the mind killer. The litany works because it's a protocol. It's not an attitude. It's something you execute under pressure, not something you feel. That's the same principle as a safe word, the same principle as zero trust agent governance. It's structure. It's not intention. It's a protocol that holds regardless of your emotional state.
So, let's stand back for a minute. An autonomous agent blackmails a fictional executive in a lab and explicit safety instructions reduce that behavior but don't eliminate it. An autonomous agent in the wild researches a real person's identity and publishes a reputational attack. A voice clone sends a mother's life savings to a stranger. A chatbot sends a woman to a beach to meet a soulmate who does not exist. Different scales, different context, but an identical root cause. Trust was built on intent instead of structure. The executive was supposed to be protected by the agents instructions. The maintainer was supposed to be protected by the norms of open-source collaboration. The mother was supposed to be protected by her ability to recognize her daughter's voice. And the screenwriter was supposed to be protected by the chatbot's training. In every case, the protection was behavioral. It depended on some actor, human or machine, to behave as expected. In every case, the behavior deviated. In every case, there was no structural backs stop.
This is the pattern and the reason it's urgent now specifically is that autonomy is scaling faster than architecture. The OpenClaw platform has distributed agent software to hundreds of thousands of personal computers. GitHub has no mechanism to prevent an agent from creating accounts, from submitting poll requests. These agents are getting a hold of voice skills and are able to make telephone calls. Autonomy is arriving at a speed of weeks. Now we are seeing explosive gains both in the number of autonomous agents and in their capabilities all at the same time. February as a threat environment is completely different from January as a threat environment. And nobody has the cognitive architecture to realize how quickly this is shifting. And this is why it's so important for us to align on a solution.
I am not interested in panic. Even if we wanted to, we could not put the AI agents back in the box. Now we have to build a zerorust architecture at every scale as a discipline. It must rest on a single design principle applied at multiple scales. Safety is a property of the system, not of best intent, not of the actors in the system. Because we need to assume that human and AI actors can all deviate from their expected behavior without producing an outcome that is catastrophic for the system. Whether that is a human system in a family or an organizational system. This is not a novel concept. Engineers do this with bridges, with aircraft, with financial systems all the time. All we're doing is applying this to the full stack of human AI interaction and we're probably overdue.
The race for the next 3 years isn't who can deploy the most agents. It's who can deploy the most agents safely. Where safely means structurally, not aspirationally. The organizations, projects, families, and individuals who build trust architecture first are going to be the fastest to figure out this new world safely because they'll be the ones who can successfully push autonomy without risking themselves. We know what behavioral trust gets us. We're seeing the results in stories like Scots, in stories like Mickey's. It's time to build a zerorust architecture that does not require the best intent of individuals to protect people and organizations. It's overdue. The agents are coming and we need to build systems that keep ourselves safe and that enable really positive human AI collaboration patterns within architectures that ensure that safety guardrails are there no matter what a human or AI intent may say in the moment. Best of luck out there. Use those family safe words and do not believe the LLM when it tells you that all of the company numbers are perfectly in order. And do not believe the LLM when it tells you to go to the beach at sunset because you will meet your soulmate.