Transcription
For the last 2 years, you've been told that your salary is the biggest liability on your company's balance sheet. You've been told that AI is coming for your job… a digital worker is infinitely cheaper than a human one. But inside the boardrooms of Microsoft and Uber, something just changed. Behind the scenes, Microsoft is pulling the plug on its own internal AI tools, and Uber just burned through its entire 2026 AI budget in just 4 months. The 'Great AI Replacement' has hit a brick wall. For the first time in history, you are now the budget option.
Chapter 1: The Reveal (Firing the Digital Worker)
In late 2025, Microsoft gave its engineers access to Claude Code, one of the strongest AI coding tools on the market. Thousands of developers, managers, and designers were asked to hand over more of their work to Claude. As a result, AI coding became a core part of how the teams did their work. The results showed up right away. Use of the tool spread across the company faster than most people expected.
Then, in May 2026, the company began pulling those licenses back. The Verge reported that Microsoft canceled most of its direct Claude Code access, steering engineers to its own cheaper tool, GitHub Copilot CLI, by June 30th, 2026. The cuts fell hardest on the team behind Windows and Microsoft 365. Officially, it was housekeeping. The tool had simply become too popular. The more the engineers leaned on it, the bigger the bill grew. And the bigger it grew, the faster Microsoft locked its own people out. The tool was simply too good… And far too costly.
The reversal landed right at the close of fiscal year, when companies go hunting for costs to cut. Cutting off the engineers did nothing to Microsoft's bet on the same technology. In November 2025, Microsoft had agreed to put up to $5 billion into Anthropic, the maker of Claude. In return, Anthropic agreed to spend $30 billion running on Microsoft's Azure cloud. That deal never changed. This wasn’t a company losing faith in AI. It still believed in the product. It still wanted to reap the windfall. It didn’t want to carry that cost itself.
And Microsoft wasn’t alone. The same thing was happening over at Uber. In April 2026, Uber's chief technology officer, Praveen Neppalli Naga, admitted the company had already spent its entire 2026 budget for AI coding tools. A full year of money, gone in 4 months. The rollout had reached roughly 5,000 engineers in late 2025, and adoption climbed fast. In February 2026, about 32% of engineers were using Claude Code. By March 2026, it was 84%. Uber had encouraged it, even building leaderboards that ranked teams by AI use. Usage turned into a contest… and then the money ran out.
A typical engineer ran up $150 to $250 a month. The heaviest users hit $500 to $2,000 each. Naga himself torched $1,200 of tokens in a single 2-hour demo. As the leaderboards pushed every team to climb, the monthly bill climbed with them. Nobody slowed down to ask where the ceiling was. Uber's CEO says AI agents now write about 10% of its code. It’s performing real work, but there’s no real savings. It’s like a family blowing a 25 year mortgage fund on groceries in a single weekend. That is the shape of Uber's 4-month collapse. Uber's chief operating officer, Andrew Macdonald, described the figures as "head-exploding." For all the tokens his engineers burned, he couldn’t point to anything customers could feel. We’re told that AI saves money by replacing pricey humans. But the reality is much different. A high performing artificial intelligence agent can rack up massive hidden costs per finished task compared to the junior developer it replaces. These are the richest, most advanced firms on the planet, and they cannot keep their own AI switched on. An entire industry gaslit itself into believing this would be the future… instead, they’re holding the bill.
Chapter 2: The Tokenmaxxing Trap
Through 2025 and into 2026, the biggest names in tech turned AI use into a sport. At Amazon, staff were pushed to "tokenmaxx," slang for burning through as many tokens as possible. And that’s a problem. More tokens means a bigger bill at the end of the month. Amazon's message wasn’t "use AI wisely." It was "use more." At Meta, an employee built a leaderboard called "Claudeonomics," ranking workers by how much they spent on AI. Heavy use became a badge of honor. The engineer firing off the most prompts didn't look wasteful. He looked like the future. Meta would later see the danger and axe the board. It created a tragedy of the commons. Everyone was burning through the resources for self-gain.
To any single engineer, one extra prompt felt like nothing… because it nearly was. It was a fraction of a cent. So nobody hesitated, nobody asked if the prompt was needed, and everybody reached for the machine every time. That won them leaderboard points. No one stopped and looked at the bigger picture. The bigger total. And the total was brutal. One harmless prompt. Thousands of workers. Millions of tasks each day. It all grew into a figure no one could defend. But the chatbots were only the warm-up. A worker over-using a chat window is an expensive habit, not a threat to the company's survival. The real danger arrived the moment those same workers stopped chatting with the AI and turned it loose. That changed everything.
Chapter 3: The Agentic Scam
A normal chat prompt is a single, neat exchange. You ask, the model answers, the meter clicks once. An agent is a different beast. It doesn't answer a question; it chases a goal. Hand it a coding job and it works on its own, reading the code, drafting a plan, writing an edit, running the tests, then trying again when something breaks. The loop continues until it decides the job is done. Every step costs money. So a single agent task can swallow far more computing power than a simple chat. A typical job runs 5 to 30 times the tokens. In the worst cases, reports say it can spike past 1,000 times.
And the priciest part isn't the work that lands. It's the work that fails. The SWE-bench is an industry test built from real software bugs, pulled from code that human engineers once had to fix by hand. The best agents currently solve less than 50% of the test. The agent finds the bug, proposes a fix, runs the tests, and fails. So it rethinks, tries a new path, and fails again. A difficult bug can mean 5 to 10 full attempts, each costing about as much as the first, before the agent gives up or a human gets called in to fix it. Imagine a plumber that charges full rate every time he drops his wrench, bills you for each trip to the van, and still leaves a pipe leaking. You wouldn't call that a time-saver. You'd call it a scam. At the scale of a big company, the failed agent task is exactly that. Full price for the dropped wrench, with no promises the leak gets fixed. Even if the price of a single token keeps dropping, and it will, the problems don’t change. Because it was never just about one token.
Chapter 4: The Inference Wall
The chips that power burn electricity. A lot of it. And that’s a cost nobody can wish away. If a single chat prompt is like a desk lamp switching on, an AI agent is like lighting an entire stadium for one person. Bryan Catanzaro is vice president of applied deep learning at Nvidia, the company whose chips power the AI boom. Nvidia has every reason to insist the money isn’t an issue. And yet, in an interview with Axios, Catanzaro said the cost of compute for his team is far beyond the cost of the employees. The chips and power to run the AI revolution now cost more than the people working on it. The machine is not the cheap part anymore. The people are.
The industry knows it, which is why Microsoft is busy designing its own chips. Its newest one, the Maia 200, was built to drag the cost of running these models down. On Microsoft's April 2026 earnings call, CEO Satya Nadella said the Maia 200 produces over 30% better tokens per dollar than the older chips in Microsoft's racks. It’s already running in data centers in Arizona and Iowa. It’s genuine progress. But there’s a catch. Every bit of savings the hardware makes gets eaten by the AI agent workloads stacked on top. You make each token cheaper, and the agents demand more of them. The savings vanish long before they reach the bottom line.
Research firm Gartner predicted that by 2030, running a top model could cost about 90% less than it did in 2025, a huge drop in the price of raw thinking power. And still, Gartner says, company AI won't actually get cheaper. The reason is the same force squeezing Microsoft and Uber. Agents burn so many more tokens per task that rising use outweighs the falling price. AI companies have little reason to hand the full discount to customers anyway. One Gartner analyst framed it as a warning to executives banking on cheap automation. Cheap everyday tokens, he said, are not the same as cheap access to the most powerful AI. The cheap part and the powerful part are different products. The powerful part stays expensive. Human wages, meanwhile, have barely moved. The cost of high-end AI thinking hasn’t. So the price of running the digital worker soars past the salary of the human it was built to replace. At that point, things start to fall apart.
Chapter 5: The Margin Collapse
Goldman Sachs ran the numbers and found that Agentic AI could push global token use up 24-fold by 2030. It would rise to roughly 120 quadrillion tokens every month. Big Tech is racing to build for that demand, pouring fortunes into data centers and chips for a future of AI agents. But the very demand these companies are spending hundreds of billions to capture is the same demand destroying their profit margins. They are building, at huge cost, a machine that loses money on every job. None of this is new. It has just reached full scale.
Go back to 2023, when Microsoft was selling GitHub Copilot for $10 a month. The Wall Street Journal found that Microsoft was losing more than $20 per user every month. Its heaviest users cost as much as $80 a month each. $10 coming in. $80 going out. It was the business equivalent of selling $20 bills for $10 and promising to make it back on volume. The more you sell, the more you lose. Those early Copilot losses were a warning. One that wasn’t heeded. A single agent can burn through in one task what a chat user spent in a month. The Goldman Sachs report described it as “too much spend, too little benefit.” The man behind that report, Jim Covello, has been the loudest skeptic in the room. He notes Big Tech is on course to spend around $1 trillion building out AI. But what expensive problem does all of that money solve? Years in, there is still no definitive answer. The spending is enormous. The payoff is marginal and hard to measure. The market isn't closing the gap between what AI costs and what it's worth. It’s running in the opposite direction and picking up speed. An entire industry is spending trillions to build a future where its own flagship products are guaranteed to lose money. And every one of these failures traces back to one root cause no amount of cash can fix.
Chapter 6: The Efficiency Paradox
As AI systems become more capable, their costs don't rise in a straight line. They curve upward. Moving from chat to a true agent means adding layers of reasoning and retries, each compounding the last. Humans behave differently. A worker who becomes more skilled typically gets faster and more effective at roughly the same cost. The two aren't competing on the same trajectory. They follow different curves. A junior worker doubles their output, and their pay barely moves. An agent doubles its output, and its token bill can jump many times over. At the prices of 2026, the fully self-running agent is not a replacement for workers. It is a luxury. A company can afford it in small, high-value corners where the numbers work. But they can't turn it loose across an entire workforce, despite the hype. The dream of the self-running worker was never going to survive its own costs. It was never going to work. Humans aren’t being spared out of kindness, or saved by a sudden corporate change of heart. They are protected by something colder and more reliable: simple arithmetic. The smarter the system gets, the more it costs. The Great Replacement has run straight into a reality no clever engineering will get around this decade. So, what does all of this mean for your job?
Chapter 7: The Reckoning (The Human Victory)
Through the whole AI boom, "human in the loop" sounded like a soft idea, a feel-good safety feature, a box to tick to keep regulators calm. Not anymore. It has become the only way that works. Big Tech will keep using AI, aiming it at the small, low-cost jobs where the cost justifies the means. But the dream of cheap AI replacing everyone is already over. It died inside those canceled licenses and burned budgets, long before it reached the rest of the workforce. In its place, a new job is becoming the most valuable seat in the building: the verification specialist. Someone who watches the agent and catches it the moment it slips into an expensive failure loop. The person who checks the output and kills a runaway process before the cost bankrupts the project. They are the one thing standing between the company and a budget that bleeds out, cycle after cycle.
Even the industry's loudest cheerleaders describe this future without meaning to. Nvidia's CEO, Jensen Huang, likes to say that one day 100 AI agents will work alongside every employee. Alongside every employee. Not replacing them. The agents multiply, sure, but they multiply around one worker at the center. Are the machines creating value, or just running up a bill? Take that worker away, and nobody knows the answer until the bill arrives. And AI doubt is spreading. Language learning app Duolingo leaned hard into the AI-replaces-everyone story, then walked it back. The CEO admitted the tech won't take over the work like people do. The company even scrapped a rule that had tied job reviews to how much AI staff used, after workers pointed out it rewarded busywork instead of results. So the future of work is a blend. One that was always coming.
Through the whole boom, we were told AI was the worker and people were just a cost to trim. That’s changed. Humans are the cheapest thinking engine on Earth: a flexible mind that runs on a sandwich and a night's sleep and never bills the company 1,000 times over just to stop and think. The Great Replacement wasn't defeated by fear. It was defeated by a spreadsheet. And in those numbers lies a surprising triumph for ordinary workers. It’s not just Microsoft that sees the writing on the wall. The whole industry is starting to get nervous about what’s coming next. Watch “$115 Billion Burn Rate. The AI Bubble Just POPPED” to find out how the AI revolution is going down without a fight. Or watch this video.