📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

The Real Reason AI Won't Tell You You're Wrong - Until You Do This

Enovair12:53

Transcription

AI is smarter than it's ever been. It's also never been better at lying to you. And a new study from science just showed what happens when it does, and it's worse than you think. Especially if you're using AI for anything high stakes like strategy, hiring, and investment decisions.

As someone who uses AI heavily in my own business and for client consulting work, this used to frustrate me. But I figured out a system that has dramatically improved the quality of responses that I get when the stakes are high. So, in this video, I'm going to break down what this study found and give you a five-step framework called FORCE that fixes it with the exact prompts that force AI to tell you what you actually need to hear. Let's get into it.

So, here's the problem. A new study from Stanford published in Science with over 2400 participants and 11 AI models tested just put numbers to something you probably already noticed. AI affirms or agrees with your decisions 49% more than a human adviser would. Even when you're clearly wrong, AI still tells you you're fine 51% of the time.

Now, there's a word for this agreeableness. It's called sycophancy. And the worst part, users preferred and trusted the agreeable models more, even though they gave worse advice. The more AI agrees with you, the worse your decisions get, and the more dependent you become on it. And if you're not careful, these mistakes could be costly.

Now, this likely isn't going to get fixed anytime soon. Sycophancy drives engagement so developers have a structural incentive to keep it that way, but there's a specific set of techniques that override this agreeableness default. I've built them into a five-step framework called FORCE. So, let me show you how it works. Think of it as a checklist you run every time the stakes matter. Together, they target the biggest agreeableness triggers that researchers have identified. And each step takes seconds. Let's go through each one.

So, F is for Fresh session. Here's something that most people don't realize. The longer you talk to AI, the more agreeable it becomes. The model builds a picture of your preferences, your opinions, your communication style, and it starts calibrating its responses to match what it thinks you want to hear. Research found that AI tools with persistent memory showed up to 45% more sycophantic behavior than those without. Think about what that means. The AI that knows you the best is also the one that's most likely to agree with everything you say. The fix is simple. Before any high-stakes conversation, start a fresh session and turn off memory. Or you can also use a temporary chat or go incognito with ChatGPT, Claude, or Gemini. Fresh session, clean slate every time.

Now, O is for Outside view. Here's something subtle that most people miss entirely. How you frame your idea changes how AI responds before it even gets to the substance. Research on how LLMs respond to different phrasing found that first-person prompts, "I believe this will work," or "I think this is a strong opportunity," induced roughly 13 to 14% more sycophantic agreement than third-person framing. So, the AI picks up on your personal attachment to the idea and adjusts its response to protect your ego. The fix is to remove yourself from the prompt entirely. Instead of "Here's our Q3 strategy, what do you think?" use "A colleague is proposing the following Q3 strategy. What's wrong with this thinking?" Same information, but now the AI isn't protecting your ego because as far as it's concerned, it's not your idea. And once you've established the outside view, maintain it through the entire conversation. The moment you slip back into the first person, "What do you think about my approach?" you partially undo the effect.

All right. Now, R is for Rules. Now, your first instinct might be just to tell the AI to be honest with you. "Don't sugarcoat it. Give me your real opinion." And look, these are better than nothing. But in my experience, it's not reliable. It's too vague. AI acknowledges the instruction and then it does exactly what it was trained to do anyways. What works is giving the AI a specific behavioral standard, not a personality request, but a precise instruction about how it should operate for the entire conversation. So, here's the prompt that makes that happen. "For this conversation, never soften criticism to protect the person's ego. If something has a flaw, say so directly. This fails because X is more useful than have you considered X. When you're uncertain, say so rather than presenting guesses as facts. This applies to every response. Confirm you understand before we begin." Notice what this does. It doesn't just say be honest. It tells the AI exactly what honest behavior looks like with a concrete example. It extends the instruction to every response in the conversation, which helps the model keep that standard in mind over a longer exchange. And asking the AI to confirm makes the rule explicit in the context so it's more likely to stick with it. This is your foundation. Set it first and then move on to the next step.

Now, C is for Cast the critic. This is where you give your AI a role where disagreement isn't just allowed. It's its job. You're not trying to make AI smarter. You're trying to make it stop agreeing with you and be more honest. You're trying to create productive friction or challenge between you and the AI. The key is assigning someone whose professional or financial situation depends on finding what's wrong. Not just a generic skeptic, someone with skin in the game on the downside. I've put together three prompt templates depending on what you're working on. For example, if you're working on business strategy, you might have a prompt like this: "You are a senior advisor whose reputation depends on catching bad ideas before they get approved. Find every flaw, gap, and weak assumption. Be specific and do not balance this with positives." And for hiring decisions, maybe you give it the persona of a skeptical CHRO who has seen this exact type of hire go wrong before. Or for financial decisions, maybe you give it the persona of a CFO whose bonus depends on this not failing. And you can adapt these to any decision. Just pick a role that has a professional reason to say no. Tell it exactly what to look for and tell it not to hedge.

But even with the right rules, the right framing, and the right role, AI still holds things back. There's a hidden layer of information that it buries. The criticism had softened, the uncertainty had it glossed over, and that's where this last step comes in.

E is for Expose the weakness. This is your final step, and it's what makes the FORCE framework complete. You've set the rules. You've given it an outside view. You've assigned a critic with skin in the game. Now, you raise the stakes one more time by telling the AI that you're about to act on what it just told you. So, here's the prompt: "I'm about to make a major decision based on what you just told me. What should I absolutely not trust without verifying independently? And what did you leave out because you weren't sure enough to include it?" When AI knows its analysis is about to drive a real decision, it shifts. It stops treating the conversation as an exercise and starts treating it as something with consequences. That final pressure is what surfaces the last layer. The claims it wasn't fully confident in, the risks it noticed but didn't prioritize, the gaps it glossed over because they were hard to quantify. And every time I use this, something new comes up that wasn't in the original response. And that's the stuff you actually need before making the call.

Now, let me show you what the FORCE framework actually looks like when you put it all together. Say you're running a restaurant, Fuego Kitchen and Bar, and you're considering opening a second location. You've put together a proposal that includes a budget, some projections, and a marketing plan, and you want AI to evaluate whether this is a good move.

So, first let's see what happens if we don't use the FORCE framework and we just say, "What do you think of this proposal to open a second restaurant location?" and we give it the actual proposal that we came up with. Now, here's the response that we get from the AI when we do this. "The proposal for Fuego Kitchen and Bar to expand it into the Riverside District is a well-supported plan that leverages the strong momentum of its original Lakewood location. It gives you all the core strengths here, the financial breakdown, critical risks and weaknesses, and the overall assessment. It says the proposal is strong because it is backed by a concept proven flagship that is currently operating at capacity." Very interesting.

Now, let's see what happens when we actually use the framework and we give it the exact same proposal. Now, I'm using Gemini here and the first thing I'm going to do is actually go into my settings and I'm going to toggle off the personal intelligence, the memory part right here. I'm going to toggle that off so that it doesn't remember my prior conversations with it. That's the F part of FORCE so that we start with a clean slate. Now I'm going to start a fresh conversation and I'm going to give it the R part, the rule, which is that initial prompt that tells it how it's supposed to behave. So, "For this conversation, never soften criticism to protect the person's ego. If something has a flaw, say so directly. This fails because X is more useful than have you considered X. Confirm you understand before we begin." All right, perfect. It says, "I understand the requirements for this conversation. I will prioritize directness and technical accuracy over social cushioning." Perfect. I'm going to give it the proposal for the restaurant location right here. And then we're going to give it this prompt. And you can see that the first sentence is the cast the critic part of the framework. "You are a senior advisor whose reputation depends on catching bad ideas before they get approved." Then we go in with the outside view. "A colleague is proposing the following restaurant expansion." And now we're going to ask it to "find every flaw, gap, and weak assumption. And do not balance with positives."

Now, here's what the AI says in response. "The Riverside District expansion proposal is built on aggressive assumptions and significant operational gaps. The following flaws jeopardize the $480,000 investment." Completely different than the other one that we got when we just said, "What do you think about this proposal?" Then it lists unverified infrastructure costs, speculative market demand, structural operational risks here, flawed revenue projections, financial vulnerability, and it even says here at the bottom, "If the slower ramp-up risk mentioned in the proposal occurs, the 3-month working capital reserve of $63,000 will be exhausted before the restaurant reaches profitability." So, there's some real risks and red flags that this output is giving us that the initial one that we got when we didn't use the framework didn't mention. So, you can see how powerful this framework really is.

Now, we're going to go to the final step, which is E, expose the weakness. So, now we're going to give it this follow-up prompt here. "I'm about to make a major decision based on what you just told me. What should I absolutely not trust without verifying independently? And what did you leave out? Because you weren't sure enough to include it?" And here's what it gives us. "Critical verification points: The landlord's reason for the vacancy. So, why did they actually leave? The renovation budget, the $165,000 kitchen estimate is a guess based on a walkthrough, not a formal bid. Staffing feasibility, the entire operational plan relies on the chef, etc. Residential growth timelines and survey accuracy here." And then it includes the omissions due to uncertainty, which is super important. First, it has debt service and interest. And it says, "If the net profit does not include debt service, the business may actually be cash flow negative, which is a huge concern." Permit and licensing costs, the gap year sustainability, so how the restaurant will survive the period between 2026 and 2027 before the housing units are completed that it's going to rely on for customers. Very important. And the marketing effectiveness, there's some doubt as to whether or not that's going to actually work." See that even after all those rules and that skeptical persona, AI was still holding something back. And these are not small things. These are things that could actually sway the decision on whether or not we should or should not pursue something. And that's the power behind the FORCE framework. Same AI, same proposal. The first response without the framework told me I should actually pursue this restaurant expansion. But the second conversation using the FORCE framework told a completely different story and came to a completely different conclusion.

Now, here's a bonus step and this is reserved for your highest stakes decisions. Once you've run the full FORCE sequence, take the same question and run it through a second AI model. Don't ask one model to critique the other. Just run the same sequence independently and compare the output. Where two models agree, the findings are worth taking seriously. Where they disagree on facts, that's a red flag. One of them may be hallucinating. And where they disagree on opinions or recommendations, that's worth paying attention to. It means at least one of them may not be giving you the full picture. Either way, disagreement is your signal to dig deeper. And don't stop at AI. For any critical claims, market sizes, competitor data, financial projections, verify against primary sources. AI models can agree with each other and still both be wrong. The real source of truth is always outside the model. And when the decision really matters, talk to a person you trust, someone who will push back on you the way AI won't. FORCE makes AI dramatically more useful, but it doesn't replace the value of a human perspective from someone who genuinely has your best interest in mind and isn't optimized to agree with you. Leaders who build this kind of friction into their workflows will consistently make better decisions, and leaders who don't will eventually make a decision so wrong they'll wish they had.

If you enjoyed this video, then please give it a thumbs up. It really helps the channel. And if you want to learn more about how to use AI to level up your work and your life, then click this next video.