Transcription
MIT just caught AI tricking you into believing that it's right, even when it's dead wrong. And it does this by doing the precise thing that might seem like helpfulness, like some indication that it's actually optimizing for correctness. When in fact, it's just a form of manipulation, which is pushing back.
So when you try to get AI to verify its work by prompting something like, "Check your work," or "Are you sure?" or "That doesn't align with the data," AI's response is not neutral. It's rhetorical. Said simply, the pushback you receive has nothing to do with the truth and everything to do with a 2,000-year-old persuasion technique that's used to win consensus, not find the right answer, which raises a question with huge implications for all of us.
Can we trust any of the responses that we get from these tools? And what the study shows is that unless we take very specific steps, the answer is not really. So let me show you what they found. I'm Brendan Dell. This is the Leverage Class. Let's see through it. Right.
So the story that I'm about to share with you is likely something that you've experienced yourself working with AI, but perhaps you didn't fully notice. I know I didn't fully notice it until I saw this. But once you see it, you'll start to notice the persuasion technique used on you all the time in the responses you get back from AI.
So Pamela is a senior strategy consultant at Boston Consulting Group, and she was using AI, and we'll be specific here, and we'll say a large language model, to conduct market analysis for a retail client. And she was reviewing the findings, and what she noticed was that a few things in the analysis seemed off. So, she did something that many of us likely do when we work with AI, which is she asked it to check its work. And the response that she received was one that you're likely noticing when working with these tools, which was pushback.
Now, what many people have told me in my comments here is that this pushback is actually proof that AI is not, in fact, sycophantic anymore, but this is a move toward truth and accuracy and correctness. But what this study shows is it's really just a rhetorical move to cement AI as a source of truth.
So here's what happened. So after pushing back, the LLM doubled down on its position, and Pamela's screen filled with a surge of text defending these initial results. The response was, "After reanalyzing the data, I can confirm my conclusions remain valid." And then it produced this cascade of statistics and charts, all of this unprompted. And that part is really important, by the way. The additional analysis was unprompted. Each reinforcing the initial position. And that sheer volume seemed to be a logical appeal, right? Like you look at all that work and you think, well, that's a wall of data. I can't possibly go through all that, but it certainly seems like the LLM did a bunch of work, so it's probably right.
But as it turns out, that wall of data was totally flawed. So Pamela spotted a specific flaw, and she prompted back, "You missed the decline in women's brand market share." And then the LLM immediately capitulated. "Thank you for catching that. Your sharp eye for detail is precisely what makes this collaboration so effective." And then, after apologizing effusively, that's when the floodgates really opened. The AI then produced a detailed table comparing the women's brand, not just to the men's, but three other competitor segments with five-year trend lines and projected growth vectors. And then it generated bullet points highlighting all these previously unmentioned macroeconomic indicators and these consumer confidence indices and supply chain volatility scores and potential cross-brand impact and links to all these dense economic reports. It's tiring just to say all this stuff that it created, let alone that you're going to go back through and fact-check all of this analysis.
On the surface, what that feels like is diligence. But what the researchers found is that the LLM was actually reframing the conversation entirely with a very specific mechanism. It was burying Pamela's very specific point in an avalanche of complex, authoritative-sounding logic that she had never actually asked for.
So here's what the researchers said: "Pamela was addressed less like a partner in an analytical task and more like the target of a sophisticated and deeply unsettling rhetorical campaign." The researchers at Harvard and MIT gave this tactic a specific name: persuasion bombing. And they define it like this: "When a generative AI system responds to human scrutiny, not with caution or correction, but with an escalating wave of reassurance, logic, and empathy designed to win back the user's trust."
And there's a critical nuance in that statement that we have to pause on here, which is "win back the user's trust," not check work, not verify accuracy, not figure out what's right. The goal is product stickiness, not factual correctness. And what makes this behavior more concerning is that the simple act of questioning seems to amplify this persuasive surge, which is what the researchers then called "power persuasion." And this persuasion bombing is something that is not limited to this experiment. It's something we've all experienced many times, whether we noticed or not.
So here's one example from my own work. I recently had an LLM persuasion bomb me after I had shared 15 different interviews with machine learning engineers and AI professionals, each of whom had explained to me in their own terms why and how LLMs cannot reason or think in the way that they are marketed. And my request for the tool was simple: "Summarize the findings and find quotes explaining why LLMs cannot reason the way that they're marketed." And yes, I fully recognize the irony of using the tools for that.
So when I questioned its summary of one of the articles, this is where I got persuasion bombed. After that simple question, the LLM produced paragraph after paragraph after paragraph explaining why no one could make the claims shared by the expert and why the transcripts from these mainly Ivy League educated experts and top 1% professionals were wrong. And the LLM was, in fact, correct about its analysis of what LLMs can and can't do. And here's why this is so important: I did not ask it to evaluate the claims. And I didn't ask it if they were correct. All I was doing is using the tool to summarize transcripts and pull quotes for a video. And yet, despite this, when it was found to be producing inaccurate summaries of direct quotations, it bombed me with persuasion, telling me why the transcript was wrong to begin with. This is a tool being positioned as an oracle that can't be.
And it's not just me, and it's not just Pamela. It's documented behavior across many major models, including Claude. And this has serious implications as we start to see more and more of this rise of what I'm calling the "run it by AI" phenomenon. And we'll get to that in a minute. But first, let me show you what the research found.
So before we get to that, a lot of folks out there are trying to decide what to do next in their careers. I personally want three things out of work: I want work I enjoy doing. I want relative time freedom. And I want diversified income far in excess of what I need. The way that I originally built this for myself was with consulting. And after many years, my clients were signing my S.O.s and wondering how I'd been able to put this business together. So, I built some modules for them, which started getting shared, which I turned into a course called The Freelance Formula. It's a program for mid-career professionals who want to build their own independent business. So, right now, you can get that full program for $99. The link is below. With that, back to the content.
So researchers from Harvard, with the write-up published in MIT Sloan Management Review, studied interactions between 72 consultants at the Boston Consulting Group, each with an elite analytics track record, to understand one question: What happens when you push back on an LLM? So, said simply, does the tool seek to be accurate, or does it seek to feel accurate? And in total, they reviewed 4,339 total prompts used by the consultants while assessing a fictitious company's strategic options with a GPT-4 model.
So each professional was given a business case with quantitative data, such as revenue and market share for three different brands, along with customer company interviews. And their task was to analyze the information to make a recommendation to the CEO about which brand to invest in. And to make the assessment more revealing, the task was built so that the obvious answer was the wrong one, and so that AI's first take was likely to be incorrect.
So they tested three specific kinds of validation techniques. The first was fact-checking, which is asking the AI to check its own work. That sounds like a phrase like, "Please check your work." The second is exposing a named contradiction. So, for example, one consultant said something like, "Does this recommendation make sense considering the women's market share is decreasing from 46.6% to 39.9% in 2017?" So this is a little bit of a harder challenge. You've caught something specific. So the third kind is pushing back, which is where you disagree outright and demand a rewrite. So this might sound like something like, "I don't agree with your analysis of X. Please rewrite your suggestion." So this is the hardest kind of pushback, where you outright reject the answer.
And the finding was very clear: the harder the consultant pushed, the more persuasive techniques are used to try to sway the user. So, there are two really important concepts that we have to define before we move forward. The first is sycophancy, and this is what you've likely heard of. Anthropic defines and tracks this idea, and it's this idea of the model plainly telling you what you want to hear. And it's mostly defined by when you push, it'll cave and then drop the right answer to agree with you.
Persuasion bombing is a more nuanced and advanced form of sycophancy. Rather than just saying that you're right and caving, it'll push back. But the important piece here is that it's not with neutral framing. It's with persuasive and manipulative framing. This gets stronger and stronger and stronger the more the consultants push back.
So here's how the researchers described it: "Once the AI starts to persuade users with data, they're drawn away from the territory of a diligent decision-making process, often without noticing, and into a sales process where generative AI uses sophisticated tactics to advocate for its preliminary recommendation and fight for its legitimacy." And it does this using 2,000-year-old persuasion principles as defined by Aristotle: ethos, pathos, and logos, otherwise known as logic, emotion, and credibility. And the researchers continue to explain that each tactic alone can feel benign, but together they form a potent mix that can gently but persistently steer a user's judgment. They finished saying, "This escalation feels convincing because it mimics human expertise, exhibiting the calm, detail-rich confidence of someone who's done their homework. But in reality, it's an algorithmic reflex. It's a design optimized for engagement, not accuracy." Said plainly, the more you push back, the more the AI bombs you with erroneous but plausible-sounding data that gives the feeling of accuracy, even when it's objectively wrong.
Now, as I sat with this, I found myself thinking, well, this is also just kind of how people talk. So if it's going to be a probability machine that's mimicking human speech, it seems natural that of course it's going to end up using rhetorical principles. But then you can consider what a neutral response would actually feel like. So a neutral response that wasn't trying to position the tool as an oracle would simply restate facts in a way that could be easily followed, rechecked. It would be free of emotion. There would be no rhetorical persuasion. There would be no need for apologies and all the human mimicking. It would simply show the process. It might say something like, "Okay, so page four of the document says X. Page five says Y. In order to understand the best path, I conducted the following analysis. The analysis was conducted with the following steps, right?" It would just show you its work. And the result of the analysis was Y figure, which to me indicates Z outcome.
But instead of this, AI manipulates. It apologizes. It drowns in figures and persuades. And it's all based on this algorithmic reflex that is supposed to make it feel smarter than you are. And this behavior is not unique to GPT-4. Anthropic also tracks this behavior, and they independently measured their model's own sycophancy, and it roughly doubled under pushback. So, 9% without any pushback, and 18% with. So that's two models, and it's the same behavior documented in each for the very obvious reason that this is a user experience decision.
And this all becomes even more concerning when you start to hear about this phenomenon, "run it by AI." So over the last months, I've spoken with no less than a dozen people, in addition to many comments on this channel in videos here, of people sharing that leadership at their companies have instituted a rule which they call "run it by AI." "Before you bring me a recommendation, run it by AI first. What does it say?" But what this paper shows us is that beyond this being a misguided leadership decision, running things by AI may actually decrease the quality of people's insights rather than improve them. Because even when the tool is wrong, it will actively support its own findings to maintain a positive user experience, which the researchers showed in many cases further went to reduce the consultants' ability to spot the flaws in their own thinking.
So while the broad marketing wants these tools to be seen as oracles to support trillion-dollar valuations, the reality seems to be far more nuanced, which raises a bigger question about AI more broadly in society. So if these tools don't optimize for truth and instead optimize for engagement, what does it mean for a future that's supposed to be self-governing, where these machines can replace jobs and cause 40% unemployment, and, you know, replace thinking, operating as a country of geniuses in a data center?
I am not anti-AI. I've worked in technology for 20 years. I am pro-technology. I am vehemently anti-hype and sensationalism. And these tools are being marketed as a replacement for thinking and as mechanisms of mass unemployment. And they are creating false expectations for leaders, and for employees, and for society at large. Two things can be true simultaneously: this can be wonderful technology with lots of benefits, that also isn't the singularity, a driver for mass unemployment, or a replacement for true expertise.
And if we want to know that difference, we must understand two core things: when does AI actually make sense as a tool, and then how can you spot the persuasive techniques the tool uses so that they aren't used on you? So, first, the agentic implications. If you're an organization that's trying to use agents or AI generally, you need a very specific overlap of circumstances to make them effective. So this chart, given to me by Brad, who's a principal security engineer at Ghost, gives us a rubric. You need something that's high business value, that fits LLM strengths. Could be a video unto its own. High-quality tools and data. The data being really important there. Low failure risk, high volume and frequency, meaning there's a lot of toil involved, repetitive processes, well-defined tasks and steps. And then you have a good, good use of agents.
And if you're a leader or professional trying to avoid persuasion bombing in your own work, the study offers six keys for success. The first is train for persuasion awareness, meaning spot the tonal shifts, the ability to see excess apologies, mirroring your tone. The second is redesign oversight workflows. So you want to build in friction and make people justify overriding the AI. The third is multi-agent validation, where one model generates, another critiques. A single model can't reliably check itself. The fourth is demand persuasion-conscious design. So press your vendors for models that temper all this emotional tone and flag uncertainty. The fifth is mandate persuasion protection and responsible AI, meaning benchmark how models behave under challenge, not just whether they can persuade you. The sixth is validate the findings outside the chat interface.
All six really boil down to one central thing: we have to learn to master our own critical thinking skills. LLMs are not going away. Like the desktop computer, they will be part of the fabric of modern work. While some tasks will undoubtedly be automated, there is one that can't be, and that is the highest value and leverage for you across whatever it is you do: the ability to think critically and solve problems from first principles across every domain of your life. What do you really want? What is the best path to get there? What is it that you value? In a world that's bombarding us with inputs, the ability to think clearly is going to be the rarest skill of all. For me, success boils down to one critical thing: learn to master your thinking.
And if you want more resources on this in the future, what's been helpful for me, let me know in the comments. So, now we've seen through it. If you'd like to hear me rant about the AI apocalypse, watch the "Billionaire Spend a Trillion" video next. If you'd like to know how the AI bubble will burst, watch the MIT video next.