📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Your Agent Produces at 100x. Your Org Reviews at 3x. That's the Problem.

AI News & Strategy Daily | Nate B Jones21:14

Transcription

The scary thing about the OpenClaw stories is that so many of them are real. You can actually build a complete $320,000 value SAS replacement suite by hooking your OpenClaw up to APIs. Yes, and that's true. You can also build a complete CRM replacement in just a few days with OpenClaw. Also true. Also a verified story. Yes, you can scale your ad created from 20 to 2,000. Also true. Also a verified story.

And the point of all of that is that you are taking a tool that is designed to be a general purpose agent and using it to cover over a lot of your own existing issues. And that's what I want to talk about today. Because when we talk about Open Claw, what we're talking about is the enthusiasm and the energy that comes from the world's first widely available general purpose agent. And I love all that hype. I love that energy. I love that you actually can do this stuff with OpenClaw. But what I don't love is that so many people are taking it as a blank slate permission slip and just saying, you know what, it doesn't matter. It's okay. I don't have to think about my data anymore. I don't have to think about my best practices anymore. I don't have to think about my software anymore. OpenClaw will make it better. No, it won't. Open Claw will not fix those things.

Look, I've spent the last couple of years arguing that the biggest risk that we face in software is not taking our entire stack seriously when we bring AI into it. If we're going to do a true reinvention for AI, we need to actually take seriously the idea that AI will need to be at the heart of a reinvented stack. We cannot just stick an open claw agent over the top, paper over all of the data issues, and pretend it's going to work. It won't. It just won't.

Now, I want to be precise about what OpenClaw actually is because the discourse has gotten very muddy with all of the hype and the hundreds of thousands of GitHub stars and everything else. OpenClaw is an open-source self-hosted model agnostic AI agent framework. It runs as a persistent Damon on your machine. It connects to your messaging app, Slack, WhatsApp, Telegram, Signal, you name it. And it acts on your behalf through shell access, browser automation, file operations, email, etc. It has a skill system that lets the agent extend itself into a bunch of communitybuilt capabilities and it has a set of memory that was originally stored as markdown files and is now undergoing some changes. The combination of all of those things as a modular architecture is new and was extremely explosive when it came out. We all know that. I've told that story.

The problem is we are now a few months into the open claw story and what we're finding is that the story gets more complicated as you try and use your open claw agent to paper over inefficiencies and weaknesses in your stack. And that's what I want to warn you about because you can absolutely do incredible things with openclaw 100%. But you have to make sure that you are building the layers of your software stack so that you can do that correctly. I think the simplest way to tell this story is to talk about real builds that we've had with OpenClaw and the risks that you have in that OpenClaw architecture if you don't take care of the rest of the stack.

So build number one, I talked about the CRM. It's a real story. It was a real non-coder who was able to build a real CRM using OpenClaw. That is both a tremendous accomplishment, a mark of how far agents have come and also something that is absolutely terrifying if you understand how CRM are built. Not because we want to go slow, we want to go really fast and there are real speed ups available. It's because CRM are fundamentally not pieces of software. We think of them as pieces of software. And the piece that came out, well, you can vibe code this in 12 days or whatever it was, it painted it as if OpenClaw built software that replaced a SAS. But but but that is not really what a CRM is. What a CRM is is encoded workflow logic that reflects the reality of your business in a way that makes sense for your sales process and for your customer care process. All of the things you know about your customers and how they buy and how they choose to make purchasing decisions and how they choose to keep their purchase and how they choose to expand and upgrade their purchase with you, all of that needs to have a place in your CRM.

And when we talk about people who are choosing not to go with Salesforce, who are vibe coding their own CRM, and this is one of N, right? There's so many. What we see when they're at their best is that people are moving quickly. They're encoding something and they make specific decisions about their CRM that reflect what they want out of the customer relationship, why their business model exists the way it is. They have that intent, that clarity of intent. I talk about clarity of intent a lot for a reason. They have that intent across their data structures and across their workflows. And all their agent does, all their open claw does is it helps them instantiate that intent. That's all. That's it. If you have any lack of clarity in your intent. If you just point your open claw at a CRM, say, "I want to build a CRM. Please just vibe code be one." You will get trash. You will get absolute trash. Not because it's not functioning software, not because your agent can't call it, but because what is built is going to reflect generic middle-of the road a workflow that works for everybody out of the box and therefore for nobody. And you're not going to get something that actually harnesses the power of custom software, which is the whole reason you use agents in the first place.

And so what I am begging you to do when you think about the promise of agentically developed software, and it is promise, and it's real, is take the high road here. There are two ways to build a Gentic software quickly. Both are fast. You're not trading speed. Number one, have clarity of intent around the workflows you are trying to encode and perhaps why they are different in the age of AI. Perhaps why your customer relationship is changing. Certainly, you should have a set of requirements that reflect what actually works on the ground in your business that's very clear, that's unique, that's different from other businesses. Otherwise, you would just buy the software, right? Then go ahead and build quickly with agents. The alternative is when you jump straight to building because that's the seductive part. That's the easy part and you don't take the time to get clear on your intent. Then you're in trouble because what you're going to get is generic average. It is whatever the average idea the LLM has of the software will be. And that's true whether you're using OpenClaw or agents uh of any kind. So that's lesson number one. It's about CRM, it's about software, it's about SAS.

Lesson number two in these openclaw agent deployments that we're starting to see is that you got to get your underlying data clean. This is part of the reason I put open brain together. I wanted people to have a clean data layer. It's I don't care if it's open brain or if it's something else for you. What I worry about is that you need to use your agent on day 30, day 300 as well as you use it on day one. And if you have a dirty memory system, a memory system that is not organized, you're going to be in trouble. And why are you going to be in trouble? You're going to be in trouble because Open Claw and other agents are not by default data organizers unless you tell them to be. Unless you give them guard rails, explicit constraints that require them to respect your data schemas and keep data clean, they won't. And so you need to start to think about the world as if it is a data-shaped problem that agents can help you accelerate, but the agents themselves are going to be messy, messy data engineers unless you structure them appropriately.

There is a story circulating about a team that spent $14,000 building a voice agent to handle lots of inbound calls. On paper, it seemed to work, but when you looked under the surface at the data the agent was handling, no one had specified how the schema was going to work. They found records all over the place. They had no good way to measure inbound and funnel, and it was just a complete mess, even though it was up and functioning and apparently took calls. So, take the time to get the schema right. Do not be satisfied with your open claw just because you can query it in a Slack message or a text message and it says something back. When I talk about legibility of surfaces as a key criterion for AI agents, this is what I mean. If if you just send a text message and it's like a void and you don't know what happens and then an answer comes back and you like the answer, you do not have an agent. You have a problem masquerading as a helpful answer. You really do because how do you know if it recorded it correctly? How do you know if it wrote it down anywhere? Unless you understand that stack, unless a company deliberately makes it clear and says, "We're going to be transparent. This is where your data lives. this is where it's going to be. This is how we update it. These are our guardrails. You're in trouble. And so, please, please, please, if you are building with agents, recognize that you don't want to be in the position of spending money, spending tokens on building something that looks good in a text message, but the data underneath is dirty.

So, we've talked about CRM, we've talked about data clarity. You know what the next thing we need to talk about is? It is mistaking a skill for a process. I love skills. Skills are great. I've talked about how important they are. I think they're critical. Do not mistake a skill or a tool call for a process. If you have a business workflow, it should be as much as you can hardwired in. You should not be trying to take your business workflow and tell yourself it's going to work good and stick it as a skill in OpenClaw and hope it works that way for production grade data every single time. Instead, your agent should be operating across hardwired process and speeding that along. And you should be able to evaluate the success of that agent across that process.

And what do I mean by that? What does that look like in practice? Let's say one of the skills your agent has is send an email. But that exists inside a larger process, right? It exists inside a larger process around triaging a ticket and talking to a customer and then recording an action. All of the stuff in between those actions like triage the ticket or write the email or contact the customer, whatever it is, when it's the in between glue, the stuff that passes the data, make that deterministic. In other words, make it as hardwired as you can because you want the agent to do the things it's really good at where it's actually composing the email in a tone that works for you and it's doing all of the wonderful things that also do. You don't want it to try to remember the whole process end to end and pretend it will follow that. You know what that's like? It's like ripping up your railroad and sticking your train on the ground and saying kind of go that way and you hope it's going to work. Don't take your rails out. Leave your rails in and let the agent do what it's good at. Agents are extremely good at text processing. Agents are actually really good at tool calling now. Agents can compose very very sharp solutions to difficult problems. But if you want an agent to follow a process the same way every time, you should be hardwiring those triggers. The agent should get the same exact trigger at the same exact time every time a ticket is open, for instance. That's how you build dependable working software. And I see a lot of people confusing this and they're like, "Well, it's a complicated process, so I'm just going to stick it all in a skill file and I'm going to I'm sure the agent will figure it out." No, it won't. Not predictably, not dependably, and I don't want to roll the dice with my customers. I hope you don't either. Please, please, please do not mistake the wonderful ability of an agent to call tools and skills with the ability to follow a workflow. Those are not the same things and we should stop thinking they are.

Look, I talk to a lot of leaders. The pattern I see is fairly consistent. The first month of an agentic deployment, if you're not careful, it all feels really good. It's the second and third month that things start to get scary because then you start to look underneath and you say, "Oh, okay. Wait, we didn't actually hardwire in this process." and a lot of things are slipping through the cracks that the agent is saying it's going well, it's going great and and and we just believed it initially. You you you got to stop letting agents tell you whether they're doing a good job or not. You got to actually evaluate them. You got to actually hardwire in that process wherever you can.

So, we've talked about CRM, we've talked about data, we've talked about tools and workflows and the difference between them. Now, we're going to talk about the org redesign that you're going to have to face if you take agents seriously. Because what people don't tell you, like the the story that went viral, and it's a real one about how OpenClaw scaled their number of creatives from 20 to 20,000. That sounds amazing if you're in the ad creative business. The problem is you just earned yourself a huge problem on the human side to review all that creative and figure out what you're going to do with it. So often we look at these tools as generative and we put like 10% of the thought into making these tools evaluative into making these tools critical thinkers and you can there are absolutely companies out there who are prioritizing I want my LLM to review the PRs that are sent to my GitHub repo. I want my LLM to think about how it can autofix bugs that are reported and autosubmit tickets that I can then review. and they're thinking as much about how the LLM is useful in evaluating quality as they are thinking about how the LLM could generate. We need to do more of that thinking. We cannot just sit there and hope that if we generate a bunch of stuff with our agent because it looks good on day one, it's not going to make ourselves really miserable on day 30 and day 60 when we have to evaluate it with the same size org. And the org redesign needs to reflect that reality. We need to be thinking in our org design about how people are going to move to be managers. I can't even say they're going to be moved to be like reviewers anymore. Like increasingly we need to think about getting ourselves abstracted out as much as we can of the agentic processes that are running because otherwise scale break points are going to pop up and you're going to see agents piling up work on a human's plate stressing out the human. The whole process slows down. you lose a lot of the benefits of all the tokens you're burning, etc.

As much as we can, we need to be thinking about it like there is a high-speed rail in the middle of a highway where the humans are driving on the highway and nobody touches the high-speed rail where the trains go back and forth and back and forth. That's kind of like the agent layer in enterprise right now. The more we architect it as if it needs its own dedicated high-spe speed rail and we are endto end thinking about the value it produces so that it's from inception to the end of the workflow it's agentic and laid out and clean and structured and evaluated and we know it works. That is what we're looking to do because if you start to mix if you want to move the train to run it on the highway you're going to have car accidents. It's going to be a problem. And increasingly what that means is the work that we do as humans is moving from we're going to be uh transporting goods in our analogy now well now the train does that right instead we're going to be focused on the beginning and the end the handoff points the way we design the system constructing that railroad all of that stuff that is an analogy in our transport networks that really works pretty well for tech right now like we're building the agentic pipeline we're clustering around the handoff points where data goes in and where data comes out this is what the future future of jobs looks like. And I think so often we have the mini me fallacy. We sit there and we think, well, these agents are good. My mini openclaw can do what my mini me does and it's going to be great. And we never think about the org design. Your open claw should not be a mini me. Your open claw, your agents that you take seriously should actually be something at the heart of the business that is configured for them first and you're just there to make sure that you are able to bring the judgment and the overall direction that you need to bring. And that is increasingly something you need to expect down the chain from individual contributors. And it's changing what we think individual contributors do. They're now managers of agents. And that is a skill set that we need to train into people.

Look, so much of the open claw story has been about security, right? I've talked about it. The fact that there are many vulnerabilities. In fact, it's such a big deal that folks like Jensen have spent a lot of time and unveiled entire tech stacks that are designed specifically to address that. safe. Open claw is now an entire category of software. The problem is deeper though. The problem is this. The reason why openclaw is unsafe is not necessarily a technical one. It's a people reason. It is because people are so hyped up on open claw and the promise of agents that they're moving really fast and they're skipping all of these foundations. They're just saying, "No, no, no, no. I'm good." And they're just moving on. Don't do that. Don't skip the foundational work. Eat your vegetables. Please, please, please take the idea that you need to have clean data seriously. That's the whole premise behind this open brain project. Take the idea that you need to have mapped out workflows seriously. That really matters. Take the idea that you need to have clarity of intent and evaluations and you need to think about these tools not just as generative but as evaluative seriously. If you don't, you are going to end up with something that is worth a cake and a party on day one and like you're going to be tearing your hair out on day 30 because it's not what you thought it was. And there are real stories. I told you some of those stories in this video. Real stories of people who spent real money on tokens, who celebrated success, who thought they launched something good, and then just like that telephone call story, the the LLM handled all the inbound calls. the open call seemed to be doing good and now they don't know where all the data was and like it's all mixed up and the data structures were never there in the first place. You got to take the rest of the stack seriously if you want agents to do good for you over the long term. Agents are not a magic wand that you can wave and fix everything. They are an incredibly powerful tool that you have to set up to use properly.

So what are my commandments for OpenClaw here? Number one, if you're going to put agents into your enterprise, if you're going to try and openclaw your enterprise, audit before you automate. Commandment number one, map the actual process. Not the idealized one, the one with all the edge cases, the one with the tribal knowledge, the one with the undocumented exception handling, all the things that are in your head. Map that. That's the first thing you do. Audit before you automate.

Commandment number two, fix the data before you give an agent access to it. Establish a source of truth. Define your schemas. Build your validation. Decide which system wins when there are two sources of truth that disagree. This is super boring work. It's also very essential if you want it to work well. So, audit before you automate. Fix the data.

Redesign your org for the throughput that you're about to get. If the agent is going to 10x your production capacity, plan your whole org around that. Do not assume the org will magically adjust. It won't. You have to think about job roles, where people sit, what they do, what tools they need access to. I can't tell you how many times people have been excited about OpenClaw and agents and then they find the IT department wasn't told about all the excitement and the hype and they have a bunch of restrictions on their computers and they can't even use the fancy new software they got. Please, please, please think about your or its provisioning and it's and what people do before you try and just strap on a rocket ship and go. You got to think it through. So, audit, fix the data, think about your org structure, redesign as you need to.

And number four, build observability from day one. Observability is not an afterthought. If you want to know if your agents are doing the scary stuff, you got to be observing them in production. You've got to actually look and see what did the agent do? What is the audit? What is the stack trace? How do you know that the agent was able to do task X or task Y and successfully get it done? Do not rely on agent self-reporting. Instead, have an independent perspective, preferably automated, that tells you if the agent got the job done correctly or not. If you don't have that, you are just rolling the dice with agents in production.

Number five, so we did build observability, we did redesign the org, we did fix the data and we did audit before you automate. All of those are great. Last one, last command. We only have five commandments here. Scope authority deliberately. Decide what the agent can do and cannot do. Make sure it's very clear. Make sure it's guardrailed. And do not give the agent free access to everything. That is one of the core sources of insecurity in a lot of open cloud deployments is that people just say, "Oh yeah, I can do anything. Dangerously skip permissions. Off you go." No, don't do that. That may make it faster on day one. It does not make your life better on day 30 or even possibly day two.

Look, I'm not against speed. If you're on here and you see my channel, you know I love how fast AI is making us go. I adore it. But do not get fooled by the possibility of going fast and think you can skip these thinking steps. You can't. The people who are going to go sustainably fast over a long period of time in the age of AI are people who take the formation of good intention and good structures that surround agents extremely seriously. They're the ones that are not just going to go fast on day one. They're going to go fast on day 60 and 90 and 120. They're the ones that have a system that enables sustained speed over time. And that's what I'm interested in building. I'm not interested in something where we can wave our hands and have a party on day one. I want something where we have sustained speed months and years later because we actually built our agents correctly. So listen to the five commandments for OpenClaw and good luck with your agent deployments and please please please do not dangerously skip permissions. Do not tell your agent to build something and not have clear intent. Please, please, please do not give your agent tasks to do that should properly be given to something with defined workflows, deterministic software, and excellent data underneath. Think about what each part of your stack does intentionally and you will be way way better off. I will see you on the other side and good luck with your deployments. I love OpenClaw. I love agents. I just want it done correctly because I want all of us to be able to speed up. Cheers.