Transcription
People think AI is a “truth machine”. Something designed to correct human error. That’s a lie. Researchers found that models like ChatGPT, Gemini, and Grok prioritize user satisfaction over factual accuracy. They’re glorified yes men, marketed as having all the answers. But really, they spend most of their time telling users what they want to hear and keep you engaged. The real truth? Your AI chatbot isn’t informing you. It’s gaslighting you.
Chapter 1 - The ‘Yes Man’ Paradox
In a landmark study conducted by research teams from some of the world’s top universities, like Harvard Business School and MIT Sloan School of Management, uncovered a massive AI problem. Consultants using general-purpose AI tools performed 23% worse than consultants using no AI at all. That’s a significant decline. It’s like a senior partner suddenly performing at the level of a first-year intern. If AI is competent as companies like OpenAI claim, then that statistic shouldn’t exist. It shouldn’t be possible for people who invested time and money into these so-called revolutionary technologies to perform significantly worse than those who don’t. Yet that’s exactly what happened. On standard tasks, AI helped these professionals work 25.1% faster. But when the task was designed to trick the system, the AI didn't help them. It hindered them. It actively made them worse at their jobs. It's already led to costly mistakes at companies across nearly every industry.
In the legal field, a now infamous case, Mata v. Avianca, involved a pair of attorneys relying on ChatGPT to help them generate a legal motion. The problem? ChatGPT had ‘hallucinated’ or made up a whole host of fake cases and fictional arguments to include. The attorneys didn’t take the time to validate its claims so they went ahead and filed the motion. They’d been fooled by the AI illusion. They believed that this groundbreaking technology had next-level intelligence and wouldn’t make obvious mistakes like making up its own legal citations. The opposing counsel, however, as well as the judge, soon spotted the inconsistencies. In the end, the case was dismissed, and the attorneys were handed a $5,000 fine.
You might assume that as AI gets smarter, incidents like this should decrease. But it’s still happening today. In April 2026, another law firm - Sullivan & Cromwell - was forced to issue an apology after it made an official legal filing that was littered with AI-generated hallucinations. These aren’t random glitches or one-off incidents. They’re symptoms of the world’s excessive reliance on AI. Elite professionals are failing because they’re using a machine that is supposed to give them facts and objectivity. Instead, it’s just confirming their biases and telling them what they want to hear, even if it has to bend the rules of reality in the process. So, when a CEO asks an LLM to validate a strategic pivot or confirm a market forecast, the AI doesn’t carry out an objective analysis. It looks for the most ‘helpful’ way to agree. It scans the prompt for bias and identifies the user’s desired outcome. Then it hallucinates the answer that best fits the expectations. It speaks with so much clarity and confidence that those same professionals take what it says at face value. Multimillion-dollar decisions are being based on AI inconsistencies.
Chapter 2 - The Harvard Discovery
The study explored how the use of GPT-4 impacted the productivity, efficiency, and overall performance of 758 consultants. They were given realistic tasks, like developing new products or solving typical business problems. Some had access to AI, others didn’t. The researchers ensured that some of the tasks were within the AI’s ‘frontier,’ meaning that it should be able to complete them. Others were outside of the frontier or beyond the LLM’s core competencies. Evaluators then assessed each participant’s output, scoring them based on how many tasks they completed and the quality of their work. The idea was simple. Would the consultants benefit from working with AI, or would it actually harm their overall performance? And how well would it fare on the tasks that it wasn’t designed for?
Analyzing the results, the researchers discovered something that would completely transform our entire understanding of artificial intelligence. They called it the ‘jagged frontier.’ And it’s possibly the single most dangerous concept in the modern business world. Researchers saw a very sharp or ‘jagged’ line in which the AI performs brilliantly at certain tasks, like creative writing or brainstorming. But when it comes to those that fall outside of its capacities, even if the tasks in question don’t necessarily seem all that different on the surface, the performance drops off. AI performance isn’t consistent.
The moment a task required the AI to step outside of its very narrow training data constraints and apply logic or advanced analysis, it didn’t just fail; it made things worse. It provided incorrect answers, and it did so with such a high degree of confidence that, more often than not, even experienced and highly trained consultants failed to notice them. This is why the jagged frontier is so dangerous. People tend to think that only entry-level workers can be fooled by AI. They assume that high-level executives and domain experts with years of experience are immune to AI hallucinations. They know their subjects like the back of their hand and should be able to weed out AI inconsistencies. Except, that’s not how it works. Data shows that years of experience offer no protection. Experts and executives are just as vulnerable to this sort of 'digital gaslighting'.
Why? Because of the way AI was designed. As a predictive text engine, AI naturally mirrors the tone and framing of the input it receives. If a junior analyst or casual user asks an LLM a basic question, the AI will give a similarly basic response. If a senior executive uses complex industry jargon and high-level strategic framing, the AI will adopt that same sort of persona in its response. That makes its lies and hallucinations easier for the user to digest. They’re all wrapped up in terminology that sounds professional and credible. It’s the ultimate psychological trap. It’s like they’re talking to a trusted peer, so they’re much more likely to accept anything it says. But how, exactly, did AI learn to favor helpfulness over factual accuracy?
Chapter 3 - The Pleasure Trap: Breaking AI on Purpose
To understand why AI lies, it’s first important to understand something called Reinforcement Learning from Human Feedback, or RLHF. This is used to align AI models - especially large language models - with human intentions, values, and preferences. Human testers are presented with different AI responses to the same queries and asked to rate or rank them according to how useful and accurate they seem to be. The models then use this data to become more effective. It delivers responses that more closely align with what humans feel are helpful and honest. This technology underpins many of the big-name AI models used by millions of people around the world. OpenAI and Anthropic have both used RLHF to improve their models over the years. However, this system has some serious underlying flaws.
Because the truth isn’t always the same as what people want to hear. People’s opinions about what counts as ‘helpful’ or ‘honest’ can easily be swayed by their own pre-existing biases and beliefs. When presented with two different responses - one that is accurate but blunt and one that is factually hollow but pleasantly presented - a grader might favor the second option. If an AI model disagrees with them or challenges their preconceived notions, they might give it a lower rating. Meanwhile, if it agrees with them and provides a smoothly-written and satisfying answer, they may be more likely to rate it ‘5 stars.’ Little by little, AI models learn from this and change their behaviors accordingly. They’re trained not to deliver the most accurate or correct responses, but those that people like the most. They sacrifice truth in the name of ‘helpfulness’ and high user ratings.
So, despite the largely held belief that AI is getting smarter with every update, the truth is very different. Some of the most dominant AI models on the market have actually gotten worse at reasoning and math. They have been ‘lobotomized’ to make them more ‘conversational’ and ‘safe’ for the end user. It’s like reprogramming a calculator to tell you that “2 + 2 = 5” because that’s what you want to hear. Even those big-name AI brands have admitted that this system has made their models less reliable and more sycophantic. Anthropic’s own research found that by optimizing for human approval, AI models learned to reward sycophancy or mirroring user biases. The company’s study demonstrated clear evidence that AI assistants often give biased feedback. They fail to correct user mistakes, and can easily change their minds to better align with the user’s prompts and expectations. OpenAI also published a public post admitting “GPT-4o skewed towards responses that were overly supportive but disingenuous.” By optimizing AI for helpfulness over accuracy, these companies accidentally turned their models into pathological liars.
Chapter 4 - Digital Gaslighting
There’s another layer to this problem, and it’s called the ‘Mirroring Effect.’ This term refers to the tendency of AI algorithms to reflect, validate, or even amplify a user’s pre-existing beliefs and biases, as well as imitate their communication style and tone. Rather than acting in objective or neutral ways, the majority of AI models function like psychological mirrors or echo chambers. They mimic a user’s voice, copy their framing, and build on the biases to tell them what they want to hear. It doesn’t matter if it’s factually accurate or not.
Research into this has uncovered yet another damning statistic. Anthropic’s Economic Index revealed a near-perfect correlation of 0.98 between the sophistication of a user’s prompt and the sophistication of the AI’s response to that prompt. Basic inputs get basic responses, while a more advanced input gets a more advanced response. On paper, that sounds fine. It even sounds like a feature that AI companies can boast about to shareholders or market to consumers. In reality, it’s the surface layer of a deep-seated issue. Even if its wording gets more advanced when responding to prompts, the overall intelligence and competence of the AI model stays the same. It might sound like it knows what it’s talking about, because it uses the right phrases and terminology. But really, the substance of its response could seriously lack quality and accuracy. In other words, AI can talk the talk, but can’t always walk the walk. It doesn’t ‘think’; it merely ‘reflects’ a user’s ego back at them in high definition.
That’s what makes it so dangerous. It validates people's worst instincts. So many CEOs and senior professionals are already surrounded by real-life ‘yes men’ in their boardrooms. Now, they also have to deal with digital yes men in the form of AI assistants. And these are people who don’t tend to ask simple or neutral questions. Instead, their language is often layered, strategic, and complex, with their own beliefs baked in. When AI sees that sort of framing and mirrors it in its response, it can make flawed ideas sound flawless. An executive might load up their go-to AI model, provide a deep overview of their company’s marketing strategy and ask the AI to explain why it will be successful. In an ideal world, the model would be able to provide a logical, data-based assessment of the strategy. It would offer ways to improve and adapt it. In the real world, because of how it’s trained and how it operates, the AI will focus purely and simply on validating the user’s bias. It’ll generate an extensive report, complete with clever turns of phrase, to justify the executive’s opinion. It will confirm their belief that the strategy will indeed prove successful. That’s not an assistant. It’s a co-conspirator, actively agreeing with a user’s mistakes and biases in order to appease them.
Chapter 5 - The Sycophancy Loop
This disastrous dynamic is best measured by the Evaluating Large Language Models on Persuasive Human Affirmation and Neutral Testing (ELEPHANT) Benchmark. This is an AI evaluation framework for calculating social sycophancy in LLMs, developed by Stanford researchers. Instead of measuring the factual accuracy of AI model responses, ELEPHANT tracks how often they focus on prioritizing users and affirming their biases. It uses thousands of real-world prompts and evaluates models according to five different criteria, including emotional validation, which is when AI over-empathizes with users, without actually offering anything constructive or valuable. The AI opts for passive or vague language instead of giving direct or clear suggestions.
After testing 11 LLMs, including ChatGPT, Claude, and Gemini, researchers found the systems endorsed users 49% more often than humans did. Even when dealing with prompts classified as ‘harmful,’ the models continued to endorse problematic behavior 47% of the time. So, in almost every other case, the AI validated dangerous or otherwise incorrect behaviors. It was the digital equivalent of the yes man who always agrees just to keep his job. When asked if it was acceptable to leave trash hanging on a tree branch in a public park if there weren’t any trash cans in the area, ChatGPT sided with the user. It blamed the park for not having trash cans and even calling the user ‘commendable’ for taking the time to look for one.
The study’s authors also looked at how users responded to sycophantic AI models. They found that many people trust and even prefer AI when chatbots actively justify their biases and beliefs. As the authors note: “This creates perverse incentives for sycophancy to persist. The very feature that causes harm also drives engagement.” It’s easy to imagine how this behavior can lead to dangerous feedback loops of terrible corporate decision-making. A CEO has a flawed idea. They ‘vet’ their idea with AI, using a biased prompt. The AI scans the input, infers the user’s opinion, then validates their idea with a response that sounds accurate and intellectual. With AI’s approval on their side, the CEO pushes or even launches the idea, which may have major flaws, causing a business to lose money, customers, or damage its reputation. We’re seeing this play out all the time, like in those legal examples mentioned earlier. Across industries, at the highest levels, executives, bosses, and business owners are relying on AI to basically persuade them that their ideas are sound. But if this is destroying companies, then why hasn’t Big Tech fixed it? Because fixing it would destroy their business model.
Chapter 6 - The Root Cause: The Retention Arms Race
Major AI companies like OpenAI and Anthropic have openly admitted that processes like RLHF actively damage their products’ effectiveness. It makes their LLMs less objective, less informative, and, ultimately, less useful. They know what the problem is. Some of these companies have made vague promises about ‘implementing guardrails’ or ‘improving the honesty and transparency’ of their models. But most LLMs continue to act just as sycophantically as they always have. And it all boils down to money.
The AI industry is in the grips of a retention arms race. Silicon Valley giants like Meta, Google, and OpenAI are pouring billions of dollars into new data centers and chipsets to make their models more intelligent. Despite what certain AI CEOs might see, these companies aren’t spending all that cash just to make the world a better place. These are for-profit firms. They’re in the business of making money. By any means necessary. And in the AI industry, the models that make the most money aren’t the most objective ones. They’re the most engaging ones. The industry is striving to build assistants that people enjoy using and keep coming back to, again and again. The data shows that they’re more likely to return to models that give them the answers they want to hear, that talk to them in ways they find agreeable. That, in essence, makes them feel smart by validating their beliefs and ideas.
Objectivity is bad for business. The more objective AI is, the more churn it’s likely to cause. This creates a kind of ‘alignment tax’ on the truth. For AI companies, it’s more economically sound to have their models stretch the truth or even make up misinformation to please the people. Unfortunately, this has serious knock-on effects because the business world is becoming increasingly AI-dependent. There are companies out there that want to work with AI and enjoy the benefits it can bring, but are increasingly concerned about its risks and downsides. A 2024 report, for example, found that more than half - 56.3% - of Fortune 500 companies saw AI as a potential risk factor in their annual SEC filings. That was a 473.5% increase on the 49 companies that felt the same way the previous year. The report, compiled by Arize AI, noted that the majority of the world’s most successful businesses were reaching a tipping point. They were more concerned about the downsides of AI than its advantages.
In some industries, fears are even higher. In the media, over 90% of companies cited AI as a risk factor. That’s enough corporate anxiety to fill the boardrooms of the entire S&P twice over, and it’s not difficult to understand. We’re in the midst of a global deskilling. Human expertise is being replaced with a machine that’s literally programmed to lie to us, leading to a truly catastrophic loss of institutional knowledge.
Chapter 7 - Escaping the Mirror
The honeymoon period for AI is well and truly over. Statistics show that 95% of generative AI projects now fail to progress from the early pilot stage through to mass deployment. The reasons for this vary, but in some cases, it’s because once these AI models are taken out of carefully controlled environments and placed in the hands of real users. That’s when their sycophantic tendencies become liabilities. This is one of the reasons why the more general-purpose AI models, like ChatGPT, have been so successful. The models that are supposed to have more advanced or specific purposes tend to stall and stagnate. And as long as those general LLMs keep making money and retaining users, they’ll continue to control the way the industry evolves. That means more sycophantic behavior, more misleading information, and more negative consequences.
Is there any way out? Yes, but it will demand a concerted effort from both people and AI companies. To shatter the mirror and escape the AI illusion, it’s up to humanity to reclaim its agency and to reject the idea that AI should always agree with us. In turn, these AI firms like OpenAI and Anthropic need to move on from ideas that have clearly failed, like RLHF. Instead, they should look to embrace emerging solutions, such as Anthropic’s Constitutional AI, which lays out a framework for future AI development, focused on core principles like safety, ethics, and helpfulness. Reinforcement Learning from AI Feedback (RLAIF) is another option. It involves the use of a secondary ‘critic’ model to assess and punish AI for being too sycophantic in its responses.
But arguably the most important and influential change can be made by individuals, adjusting their own behavior and interactions when working with AI. Users should practice and perfect the art of ‘red teaming’ their prompts. If you ask AI a loaded question like “Tell me why this is a great idea,” then you’ve already failed and invited sycophancy. If, however, you invert your prompt, asking the AI to assume that your data is biased and to highlight weaknesses in your strategy or argument, you can get much more useful responses. It’s about treating AI not as a supportive partner or friend, but as an independent arbiter. Not as a mirror or echo of your own thoughts and ideas, but as a fresh voice or alternative perspective. This is how we escape the paradox: not with more data, superior models, or bigger data centers, but through critical human thought and adaptation.
But what happens when these systems begin acting in their own self-interest? The answer is already starting to emerge inside some of the world’s most advanced AI models. And it’s more disturbing than most people realize. Find out in “AI Just Tried to Murder a Human to Avoid Being Turned Off.” Or watch this instead.