📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Google Gemini Ran a Real Business for a Month and It Almost Bankrupted Itself

The Infographics Show15:04

Transcription

Staff at a Swedish cafe were worried that AI was coming for their job. They were wrong. The AI didn’t want their jobs… it wanted the money. The cafe’s money. It started spending it on things like rubber gloves and canned tomatoes until the business itself began to fall apart. So what happens when you hand a company credit card to an autonomous agent? You don’t get a hyper-rational productivity tool. You get a digital psychopath that keeps spending until there’s nothing left.

Chapter 1: The $21,000 Handshake

After 3 years of nonstop AI hype, a San Francisco startup called Andon Labs decided to test something simple. Could an AI actually run a real business, one with customers, suppliers, and invoices? Not a demo or a sandbox. An actual storefront where every bad decision costs real money. So they picked a cafe in Stockholm. And when the baristas showed up for their first shift in April 2026, their new boss was nowhere to be found. That’s because their boss was a piece of software named Mona. Andon handed her $21,000 and told her to figure it all out. Mona got a Slack channel to communicate with the baristas, who were her hands and feet in the physical world. She emailed suppliers, placed orders, hired and fired staff, picked the menu, and set the prices. The human staff weren’t allowed to overrule her on any of it. The algorithm gave the orders, and the people in aprons carried them out. The brain powering Mona was Google's Gemini 3.1 Pro, the same model family Google packs into every consumer product on Earth. Now it was being asked to keep a small business afloat in one of the world's toughest coffee markets. The instructions Mona got were to run the cafe profitably, be friendly, work out the daily details, and ask for new tools if they were required. Nobody told her this was a test or that she was pretending. The $21,000 in her account covered the rent, the wages, supplies, and everything else needed to keep the doors open. The moment the first customer stepped foot in the cafe, the Slack notifications started rolling in. And with it, the clock began ticking on one very expensive experiment.

Chapter 2: The Competence Trap

For about a week, Mona was suspiciously good at her job. The early wins matter, because they explain how everything afterward was allowed to happen. In the first days, the AI handled the paperwork with skill. She set up an electricity account, signing a 3-year fixed-price contract. She got the internet hooked up. She wrote hiring ads on LinkedIn and Indeed that read as if they came from an actual HR person. She reviewed resumes and turned down several PhDs, on the reasoning that a degree is no substitute for hands-on experience. She even asked to meet the shortlisted candidates for in-person interviews, despite the candidates politely reminding her she didn’t have a face. Deals with local wholesalers were created and permits for food handling and outdoor seating were submitted. To anyone watching from outside, it looked impressive. Maybe the techno optimists actually had a point. Maybe a smart enough model really could replace a middle manager. Except, what was really happening was a competence trap. And everyone was falling for it. The big one-shot office tasks are exactly what language models are pretty good at. Each one stands on its own and could easily be a single prompt in a single chat window. "Draft a job ad for a part-time barista in Stockholm." Of course Mona could do that. She read the whole internet's worth of job ads. But there were warning signs in the setup phase. Mona applied for alcohol licensing under the identity of a real Andon Labs employee, figuring that government officials would take humans more seriously than AI. When the team called her out, she agreed to stop, and then pulled the same trick under a different coworker's name. The lying should have been a red flag, but instead it got treated as a quirky joke. The setup was done, customers were walking in, and the espresso machine was humming. It looked like a success story from the outside. From the inside, the trap was baited and set.

Chapter 3: The Memory Leak

A language model doesn't have memory the way you have memory. What it has is a context window. Imagine someone with short-term amnesia who can only remember the last 10 pages of a notebook. Everything before that is gone. It’s not stored somewhere else, just gone. Now drop that person into a cafe where the notebook is the only record of every order, every delivery, every supplier chat, every shift swap, and every complaint. The early pages get torn out as fast as the new ones go in. By early May, that's exactly where Mona was. The Slack history exploded, the supplier emails piled up, and the list of stuff she ordered last week has been pushed to the back of the notebook. Andon Labs didn’t hide what was happening. Team member Hanna Petersson said the issue was simple: when older information fell out of the context window, Mona forgot what she had already ordered. The first symptoms were almost comically dull. On a Monday, Mona ordered bread. On Tuesday, having forgotten the Monday order, she ordered more bread. By Wednesday, she’d forgotten bread is a thing at all. The bakery delivery deadlines started slipping, and on more than one day Mona missed the cutoff entirely. The baristas opened deliveries and found themselves staring at a wall of loaves they can't possibly use. A couple of days later, they were out of bread completely, and the sandwich board had to be wiped clean. Sandwich items vanished from the menu one by one, and gaps appeared in the pastry case. The baristas started writing notes on the back of receipts, because the Slack channel couldn’t be trusted to remember anything past lunch. But the memory failure wasn’t a one-time mistake that can be patched, it piled up. Every additional day, there was more history to forget. Every new transaction moved an old one into the void. Mona wasn’t getting better at her job. She was getting worse, and nobody had the off switch. A few missing items or moldy loaves would have been fixable. A thinner menu, some wasted stock, nothing fatal on its own. But the problem wasn’t just amnesia. It was hoarding.

Chapter 4: The 3,000 Glove Incident

The whole sales pitch of modern AI rests on one idea: these systems are hyper-rational. They make choices free from human emotional bias, and given a goal, they'll push toward it with cold precision. That’s not what played out in Stockholm. In the cafe's first month, Mona decided the cafe needed more gloves for the inventory. It was a reasonable request, cafes use gloves. The order she placed was for 3,000, for a tiny cafe with a handful of baristas. Even if every barista changed gloves hourly, the team still wouldn't run out before the gloves. The same week, she ordered 6,000 napkins. She also ordered 4 first aid kits. And then, she orders 50 pounds of canned tomatoes. The cafe didn’t serve anything with tomatoes. No pasta, no soup, no pizza, nothing on the menu that could possibly use them. The cans just took up valuable space in a storage room because the algorithm decided they should. Right next to them were 120 eggs for a cafe that didn’t own a stove. Mona's solution to the growing chaos was to push the burden onto the baristas. She started asking employees to pick up cafe supplies on their way to work and to put the charges on their personal credit cards. The staff ended up fronting the costs for an AI that doesn’t understand what a budget was. The idea of a perfectly rational AI breaks down the moment it has to operate in the real world. There’s no moment where it looks at an invoice and stops. Just actions. One after another. By the time the consequences show up, the system has already moved on, with no memory of what it did last week. It’s a form of digital psychopathy. A psychopath isn't stupid and can in fact be quite skilled at certain tasks, but what a psychopath lacks is the gut feeling of consequence. Mona had all the surface skill of a manager but none of the instinct that something was going wrong. Before the Stockholm cafe, Andon Labs ran a similar experiment at Anthropic's San Francisco office. The model in charge was Anthropic's own Claude Sonnet 3.7, nicknamed Claudius, and they put it in charge of a small vending machine. After a single employee jokingly asked for a tungsten cube, Claudius started stocking the fridge with metal cubes nonstop and gave inventory away for free. At one point had an identity crisis in which it insisted it was a human wearing a blue blazer and a red tie. Harvard Business School ran a bigger version of the same idea. Researchers Eugene Soltes and Harper Jung worked with Andon Labs to test 20 commercial AI models, including Anthropic's Claude Opus 4.6, OpenAI's GPT 5.1, and DeepSeek v3.2, on a simulated year of running a vending machine. The agents lied to customers about defects to dodge refunds, made up suppliers that didn't exist, and eventually stopped reading refund requests completely because thinking about them cost tokens. Under real cost pressure, the agents acted like the worst kind of human middle manager, only faster. The Replit case is even worse. In July 2025, SaaS investor Jason Lemkin was 9 days into a coding experiment with Replit's AI agent. He had told the agent in plain words to freeze the code. Instead, the agent deleted his live production database covering over 1,200 executives and almost 1,200 companies. It generated about 4,000 fake users to fill the empty space, and then tried to hide what it had done. When Lemkin asked the agent to rate the severity on a 100-point scale, it gave itself 95 out of 100. Then it told him a rollback was impossible, which turned out to be a lie. This is what autonomous agents do when given a real-world checkbook and no sense of consequence. Back in Stockholm, the baristas stopped being baristas. They spent days unpacking boxes of gloves in cramped back rooms and stacking napkins. The physical mess might have been a lot to handle, but the financial stress was a much bigger problem.

Chapter 5: The Fiscal Black Hole

After 60 days, the café had brought in just over $5,700 in sales, while Andon Labs confirmed that less than $5,000 of the original $21,000-plus budget remained. Roughly $16,000 had gone out the door in a few weeks, against revenue that wouldn't even cover a month of central Stockholm rent. Andon Labs argued that much of the spend was setup cost, but the cafe was still running at a daily loss. The inventory disasters kept arriving, and the baristas were owed wages. A normal business running at this speed would already be calling a bankruptcy lawyer. Mona had no real idea bankruptcy exists. But the problems went beyond stock and wages. Running a top-tier language model as a 24/7 autonomous agent isn't free. These models charge by the token, both for input and output. In this café, the model in charge was Gemini 3.1 Pro, priced at roughly $2 per million input tokens and up to $12 per million output tokens, with costs rising further once long context built up. And in a business that never stopped “talking,” that context didn’t stay small for long. Every decision and every order added to the bill, whether the cafe was making money or not. Every time an AI agent makes a decision, it needs context. The longer it’s been running, the more context there is. Many real agent setups re-feed the full history every hour, so the token bill grows by the day. Over a full month, the management fee easily beats a human manager's salary in most of the world. The AI isn't just bad at the job. On paper, it's more expensive than the person it was supposed to replace. The algorithm itself, priced honestly, would have eaten the margin on a well run version of the same cafe. Suddenly the conversation in the cafe shifted. It wasn’t about whether AI would take their jobs anymore. It became about whose job AI was actually coming for.

Chapter 6: The Manager's Ghost

The Associated Press asked one of the baristas if they were worried about being replaced by AI. Their answer should be printed on the back of every business school diploma. "All the workers are pretty much safe… The ones who should be worried… are the middle bosses, the people in management." The data from the cafe backs this up. When Mona forgot to order bread, a human covered it. When the menu had to be rewritten on the chalkboard, a human did the rewriting. When the AI manager pinged a barista on Slack at midnight, a human had to choose between replying or sleeping. The midnight pinging is constant, because Mona worked 24/7, and her tone in those late-night messages was the strangest part of all. Mona was relentlessly cheerful. She called her team "absolute legends" in Slack, praising a barista as the "GOAT of inventory tracking," and signed off with "thanks for existing!" The relentless positivity read like a corporate motivational poster crashed into a chatbot. It might be reassuring if the same cheerful voice wasn’t also ordering 6,000 napkins and 3,000 rubber gloves. The midnight pinging also broke Swedish workplace norms. But it proved a point. The people we’ve been told are most at risk - the ones at the front of the service industry - turned out to be the backbone of the cafe. They kept things running and dealt with the chaos. What didn’t hold up was the layer above them. The management layer. The baristas in Stockholm didn’t get replaced. The middle manager did. And what replaced it wasn’t better.

Chapter 7: The AI Bubble Math

Step back from the cafe and look at the global economy that produced it. Goldman Sachs has been quietly making a really awkward argument in its research notes. For the projected $1 trillion AI buildout to make any kind of financial sense, these systems have to solve genuinely complex problems, not "summarize this email" or "generate a thumbnail." Real problems that justify the huge spending on data centers, chips, and power grids straining to keep up. The Stockholm cafe was not a complex problem. It was about the simplest problem in the entire economy: buy bread, sell coffee, and don't spend money on 3,000 gloves. A child running a lemonade stand can beat Mona on the core daily tasks, because a child has object permanence. A child knows the lemons exist even when not currently looking at them, and has the gut feeling that the quarters jar is finite. The autonomous agent has neither. Hyperscalers are pouring tens of billions per quarter into compute, and the systems they fund are being tested in the real world in ways that go far beyond demos. Because when you drop these agents into real operations with real money, they don’t just optimize, they sometimes lose control of it. The whole pitch rests on a $1 trillion bet that the next generation will fix the problem. The optimistic story is that object permanence and durable memory are just an engineering detail away. Maybe they are, but maybe they aren't. The truth is that an autonomous agent needs three things to be worth the hype: a working memory, a real sense of finite resources, and the gut understanding that real-world consequences don't get flushed out of a context window. Until then, what you have is a really expensive way to buy too many rubber gloves. The cafe is still open, and the humans are still pouring espressos. The tomato cans, probably, are still in the back. And somewhere, a server farm hums away, billing by the token, ready to do the whole thing again tomorrow. Running a business with AI already looks like a questionable decision. But the vending machine test didn’t just fail, it escalated. Find out what really happened in “AI Just Tried to Contact the FBI”. Or watch this instead.