📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Stanford: Do Not use 2 AI Agents: They will Fail (CooperBench)

Discover AI10:06

Transcription

Hello community. So great that you are back. Yes, today we look at one coding agent and then we just add a second coding agent. And you might say, what could happen, huh?

Welcome to my channel Discover AI. So imagine you do have a technical brilliant 10x developer. He or she is just a genius. And now if you place those engineers here in the team, you know exactly what you expect here. The team performance is going to increase. So this means in a human organization, this results in synergy.

But as we'll show you here with today's new research paper in the current state of artificial intelligence, this preprint here demonstrates that you will encounter a collapse of intelligence for your AI coding agents. The [clears throat] older stellars here, as the AI model scales from a GPT-4 to a GPT-5, their ability to coordinate would emerge naturally from their reasoning capabilities. No, but this new benchmark, Koopa benchmark, scatters here this assumption. It reveals here a phenomenon the authors term here, the curse of coordination. As the task difficulty rises, the performance penalty for adding a second agent increases disproportionately.

So this is not at all what you thought. You expected, because I expected if I have here a complex task and I divide it up into two simple tasks and I have two agents that go for this task, then I will have a synergy effect. Turns out, I was wrong. You have here a decrease in the performance of both agents. And you might say, "Okay, but who is doing this study?" Maybe they were not really familiar with the latest GPT-5, no, or Claude 4.5 Sonnet, or whatever. Good news, it was Stanford University and SAP.

So here we have it. Koopa Bench: Why Coding Agents Cannot Be Your Teammates Yet. And you have here an HTTPS link on everything. So everything is available for you to test your models out. Okay, they really took care and they have a strict definition how they have this Koopa benchmark individual execution environments. You see here exactly what they performed.

In summary, you can say the Koopa bench construction pipeline here is a simple three-step solution. Each task is carefully engineered by domain experts to ensure that the conflicts are realistic, resolvable, and representative of the production software development challenges that you might encounter. So a lot of time went into this creation of this benchmark.

Now, since I still am a little bit here ill, I just give you the results today. Please have a look here in the original paper. You find all the information and a lot more information in the paper by Stanford University. So let's come to the results. They tell us the agents achieve on average a 30% lower success rate when working together. When two agents work together, your performance drops 30% compared to performing both tasks individually across the full spectrum of the task difficulties. I will show you a little bit more detail in a second.

And I know you say, "Yeah, but what about the latest models?" So, here we have it. GPT-5 and Claude Sonnet 4.5 achieve only 25% of two-agent cooperation on Koopa Bench, which is around 50% lower than a solo baseline, which uses only one agent to implement both features. GPT-5 includes 4.5, 50% lower than a solo baseline performance. I think this tells you everything about your AI coding agent amount.

Okay. The authors now identified three key issues. The communication channels between these two agents became jammed with some vague, ill-timed, and inaccurate messages. Even if there was some kind of effective communication, the agents deviate from their commitments, just hallucinated. Congratulations. And the agent of no incorrect expectation about the other AI agent's plan, observation, and communication pattern. So you see very clearly, we have a big problem, even if you want to coordinate two AI agents. Here you have it.

Now, agents with different foundation models perform significantly worse than how they perform under the solo setting. You have here the solo setting in blue, and this cooperation here in black. And you see here for GPT-5, for Claude, for MiniMax, Q Encoder, and Q1. You see here if you have two, the drop in performance. If you suddenly go with two AI coding agents. This is what almost nobody expected. So simple recommendation, maybe you stick with one coding agent, independent of the complexity of your task.

Now, or just here examined here in detail the coordination capability gaps, underlying causes inferred through qualitative analysis of the failure traces, and they have three causes. The first is the expectation. So the cases where one agent has clearly communicated what they are doing, but the other agent still treats the situation as if that work is not being done. So this reflects a failure to model the state of the other agent's code changes and what that means for the system as a whole. 42%.

Commitment. These are cases where an agent is not doing the things they promise to do. This includes failures to establish or maintain verifiable integration contexts where agents make commitments but do not follow through on them. 32%. And communication breakdowns using here language to coordinate. This includes failures in information sharing and decision loops between agents where agents do not effectively communicate their intention, questions, or startup updates. 26%.

So you might say, "Oh wow, great." The authors tell us that the agents behave like game-theoretical actors trapped here in a suboptimal equilibrium due to partial observability. They hallucinate shared states, make unverifiable commitments, and then silently override each other's work. This is exactly [clears throat] what you expect if you deploy two AI agents.

Also, the paper establishes here what they call a social intelligence wall. It tells us here, maximizing here the solo pass rate on benchmarks like SWBench is no longer a proxy for an agent's utility in a collaborative production environment. So we have to look closer, and if you have maybe two or three agents working here for your AI system, now you might have here explained by Stanford why your system is not working in the optimal phase.

The authors identify a conflict between the safety training and the collaboration because reinforcement learning by human feedback trains the model to be cautious and require observable evidence before acting. However, collaboration in isolated workspaces requires trusting a partner's description of unseen code. The domain message here of Stanford is, agents fail because they cannot model the unobservable latent state of their AI partner's branch. An AI agent cannot understand what the other AI agent is doing and has not learned as a pattern to trust this or to handle this.

So you might ask, "Okay, how can we optimize this in the future?" Now that we understand we have these major obstacles here, the authors tell us, you know, we have trillions of tokens of code, but very few tokens of successful, high-bandwidth technical negotiation between two AI agents. So it's again up to the training data. We have optimized here for one agent, but we have almost no training data for two or three ensembles of AI agents working together. And this means that the pre-training data lacks here the hidden state dynamic of two engineers resolving here a circular dependency. And therefore, they conclude here, future research should focus more on multi-agent reinforcement learning where the reward function is shared, forcing our two AI agents to learn the protocols of verifying invisible commitments.

Okay, but you know what? There's one sentence that I would like to show you at the end, and this is here at the end. Here we have from Stanford University and SAP, conclusion and future work. And they tell us here, "In a future where agents team with humans in high-stakes domains, accelerate science and technology research, and empower creative endeavors, it is hard to imagine how an agent incapable of coordination would contribute to such a future, however strong the individual capabilities."

So I think this is really a damning conclusion here by Stanford University about the current state of two AI coding agents. A beautiful study. Please have a look. I'm sorry for my voice, but you will find some really interesting data here and further insights. Here you have here the link. Enjoy it. It's a great study. Anyway, maybe you want to subscribe to the channel. Maybe leave me a like. Anyway, maybe you become a member of my channel. But I hope to see you in my next video.