📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Controlling Agent Swarms is your ONLY job...

Wes Roth24:05

Transcription

So, I got a feeling that this article is about to blow up. It's called “Age of the Agent Orchestrator.” It's written by Shyamal, who's applied AI at OpenAI. And every once in a while, we get a blog post or an article that just seems to kind of be a little bit prescient about the future. It's usually from people that are kind of on the inside of AI that can kind of see where it's going.

And more and more people are kind of hammering this point: What skills and abilities will be required in the future as AI intersects with more and more jobs? Now, I always kind of describe the future of jobs as looking like kind of kind of like playing the game of Factorio.

So, Factorio, if you haven't played it, it's this sort of factory-building 2D-ish simulation. You start by mining iron ore with a pickaxe, basically, eventually building up various drills and then smelting facilities and power plants, eventually making train stations, conveyor belts, and it just gets crazy from there. You get to design systems and slowly kind of over time see how those systems behave and optimize it here and there to improve outputs, and you're kind of trying to model how to improve that system and then keep adding complexity and optimizing it the whole way. If there's bottlenecks, you try to improve on the bottlenecks.

Now, more and more various AI researchers are kind of hinting that the people that are going to really succeed in the AI game are going to be people with some of these skills. So, here, for example, is Will DePue. So, he's Master of Slop at OpenAI, apparently. That's an interesting job title. I wonder what the job description is. How do you become the Master of Slop?

And this is what the Master of Slop has to say about kind of the future of work: “I do think the future of work is like Starcraft or Age of Empires. You have 200 micro-agents you're directing to fix problems, gather information, reach out to people, design new systems, etc.” And you know what? Yeah, it's it's about time. That's why I think this “Age of the Agent Orchestrator”—which is a great title, by the way—that's why it's so good because I think it kind of breaks down this idea, this concept, into more kind of easy-to-understand chunks. So, a lot of us kind of had this idea, and this person sat down and kind of wrote out why, kind of wrote the proof, if you will. So, let's take a look.

So, the question is, you know, a lot of people—his family are asking this person—what happens, what the world is going to look like when AI agents can do everything. Now, we're not close to that yet, but the abilities of these agents are rapidly improving. So, it's important to think about—and one great way to think about it—is what is scarce, what is in short supply. Markets tend to organize around whatever is in short supply.

So, for quite some time now, you know, the scarce thing that we've had now and leading up to this moment is knowing how to write software, for example, right? So, if you knew how to code something up that did a particular thing, you would create tons of value for everyone. When Excel showed up, the scarce thing was, for a while, was people that knew how to use it. If you knew how to use Excel, you could model a business. Talents like yours would be in short supply and therefore very in demand. You would get paid more. It's that idea of kind of supply and demand. Low supply, high demand, it means that the price goes up. The price that people are willing to pay for it goes up. And in in business, it's also a great competitive advantage.

And that now, as these AI models are getting better at, you know, doing the work, checking the work, improving the work, like doing the actual work, what's the scarce thing? What's the thing that's going to be missing? It's not going to be who knows how to do the task by hand. The scarce thing becomes who can orchestrate resources—well, compute, capital, access to data, and human/expert judgment.

To illustrate this point, just a few days ago, I wanted to scrape a certain website for—it had a bunch of data that I wanted to analyze. I wanted to see if there was some relationship between two sets of numbers. It doesn't really matter what they are. So, I went to ChatGPT, using the GPT-3 model, said, “Okay, what's the code to scrape this website,” allegedly. I think I'm supposed to say “allegedly,” but first of all, it comes back with and says, “Hey, this website actually has an API that has a free tier that allows you to get all that data without, you know, scraping because scraping is kind of a gray area.” It was willing to do it, but it said, “You know, why not do it the the right way with the API?” And I said, “Yes, that's terrific. Do that.” Five minutes later, I was able to get the data that I needed, put into an Excel spreadsheet. I uploaded the Excel spreadsheet to the model—to ChatGPT—and asked it to run a linear regression.

So, basically, figure out if there's some correlation between those two sets of data. So, in other words, like if one number goes up, does the other one kind of like go along with it? Are they kind of related? Are they correlated? And it just did all of that. It gave me a little chart, calculated the relationship. It said, “No, there's not really a trend there.” And then it said, “Maybe the trend isn't linear. Should we do something to see if there's a nonlinear trend?”

Again, keep in mind that I didn't necessarily have to know how to do all this stuff. Like, as long as I can kind of put into words what I wanted it to do, it would figure out the rest. So, we did the quadratic test to see if there's maybe like a quadratic—like a parabola—type relationship. Maybe the data fit on kind of like this U curve to see maybe there's some peak. And it did all the charts, all the plots, all the the math behind it.

My point is, not that long ago, if you wanted to reach that same point, you would have to know quite a bit. Number one, you would need to know how to program that thing in Python to scrape the data, use the API from the website to to get that data. You would need to know how to run that code. You would need to know how to get that code, you know, the data into an Excel spreadsheet, how to organize it, how to, you know, make the charts and the plots and how to do the various analysis and stuff like that. You would have to know what those things were to begin with, and then you would need to know how to do them in Excel. You would need to have expertise in those areas to be able to get those answers, right? The the answers to your questions. Now, more and more, you don't. You just need to know that you can ask ChatGPT for it and kind of be aware of the things that it could do.

So, as Shyamal is saying here, expertise gets democratized. And I think this is a a terrific way of saying it. So, if you spend 10,000 hours learning how to use Excel, I mean, that's great. You used to have expertise that a lot of other people lacked. Now, that kind of advantage is kind of flattened because most people, if they have ChatGPT or something similar, they can throw a spreadsheet in in there and have it do all the work.

So, that means that the premium shifts from “I know the tax code,” right? To “I can design the loop that gets the right answer and runs cheaply.” So, I did it because I had a subscription to ChatGPT, a monthly subscription. If I had to use the API and kind of pay per token or whatever, you know, kind of per use, I don't know, I estimate maybe that'll cost $3 to $5, something like that. But there's probably ways to set it up—like if you had to do it on a large scale—to where it would cost under a dollar, you know, 10, 20 cents, something like that. If you had to design, for example, software that just did that and you need to get the cost per run as cheap as possible, I mean, you can really get it much, much lower. The GPT-3 is a massive kind of general model. If you wanted to have kind of a single-serve model, kind of just focus on one thing, it it would be a lot cheaper to run.

So, right now you can hire an expert to do your taxes because the various possible mistakes that you can make, there's a lot of them, and the cost is is high if you're wrong. But if you can build or rent an agent that can do a good job of your taxes, they can ask for feedback, double-check edge cases and improve over time, then the—I like how he calls it here—“gatekeeping function of the expert” goes away.

Now, that's an interesting way of putting it because I mean, on one hand, you can think of experts as people that learn to do something and they provide value to you. Yeah, that's kind of like one take. The other take is, yeah, the gatekeeping function of the expert. If you want to have something done that you don't know how to do yourself, you're forced often to pay an expert for their services. And the more complicated it is and the less good people there are doing that service, the the more expensive it is. So, if you're trying to do something important and you need a very highly qualified person to do it, it's probably going to cost a lot of money. But for a lot of tasks, AI will kind of remove that moat. It will be easier to create images and video and do various math things and aggregate knowledge and do research and all sorts of things. Most things you can do in front of a computer will be affected by AI in one way or another.

Now, there still will be experts, but those experts—it'll be kind of a new category. This will be people to provide judgment, to set to set the high-level strategy, to handle the weird cases. So, for example, knowing the details of the tax code will be less valuable than knowing how to set up an autonomous workflow that produces a correct return, flags the ambiguous parts for a human to look at and costs 5 cents in compute instead of $5.

One thing that I really enjoy doing with ChatGPT is you can upload your, for example, blood work results and then have it explain to you like line by line what they mean. Like, you know those weird numbers and acronyms that they have on your blood work. I mean, if you get a full panel, like how many of those do you actually know what they mean? If you're in the medical profession, hopefully a lot, but for most of us, that that's not the case here. Not only can it walk you through everything and kind of also tell you kind of what it means and could even give you tips on how, you know, if it's a problem, how to improve it.

So, for example, if I wanted to, I can probably set up some sort of a loop that people just upload a PDF of their blood work and I have a certain loop that runs that kind of returns all the most sort of valuable answers to them, kind of gives them an overview, and it can be even targeted to specific conditions—like if you're feeling sluggish and low energy—like we can look at it specifically from that from that aspect, like what might be causing us not to be energetic, and I will go through and maybe highlight some things that are out of whack. Again, just an example. There's a million of other examples. These are just the things that I've have personally done that I I know that work and where these models can return excellent, excellent answers.

And so he continues: “It's also going to be uncomfortable for people whose status is tied up in being the only ones who can do something. Not only status, but also the wages and the salaries will also likely be impacted. And if you're training for a career because you believe the information is scarce, it's worth asking yourself what happens when it's not. If you're training to be the person who can take all the cheap information and turn it into a valuable outcome, you're on firmer ground. Resource optimization matters a lot more.”

So, in other words, kind of knowledge and ability to do stuff was scarce, and moving into the future, it won't be. Now, we have a massive abundance of abil—of an ability to do stuff. So, the question is, how do we make the most of it? How do we turn all those potential abilities into value for people using the least amount of resources? So, what does that actually look like?

So, he's saying we're going to have to get very good at assigning FLOPS/MP compute, liquidity, lab time, human review, etc., to autonomous workflows/agents. This is a new job. So, if you want to do some big marketing analysis, you send it to your analyst, and she tells you it'll be done in two weeks. So, she has other stuff to do. She has kind of an idea of how long it's going to take her. But in the world in which we're likely going to be living soon, let's say you can spin up 10,000 agents to do 10,000 analyses, your bottlenecks become very different. So, it's no longer kind of the idea of like you have between 9 to 5 to get the thing done. It's more about how to reduce, you know, the per-task costs as well as how to optimize the resources, right? Because your compute cluster might be finite, your budget is finite, the number of hours your expert reviewers is finite, right? So, you want to be doing more on the computer side. You're still going to need people to review it, but are there ways that you can reduce how much review they have to do? If your computer agents are producing code and you need somebody to do a final sort of look at it, is there automated checks that you can do that either kind of pass or fail that before the human has to review it, or maybe make it easier for the human to review it? Maybe some sort of a little printout or “Here are some potential red flags to look at,” etc. You got to think of things like how to queue up a million of tasks overnight when energy is cheap.

So, often when people say “liquidity,” they just mean like cash on hand. Do you have enough cash to do something? So, I'm actually not sure what he means by “low liquidity windows” in this context. Let me know in the comments if that makes sense to you. I mean, we're not talking about cash, right? We're talking about—I'm not getting that line. Let me know. Let me know if you do. What is a low liquidity window in in in this context?

So, in the past, various startups and businesses had to allocate capital carefully. In this new regime, you have to allocate compute carefully, too. It's not a “set it and forget it” type situation. Models will get better. Costs will go down. New tasks will appear. You will need people whose job it is to literally orchestrate fleets of AI agents and the scarce resources they need. Think of it like an air traffic controller or orchestrator for agents. The very best companies already do some of this. They optimize the cloud spend. They schedule jobs. They think about capital efficiency, but still, they're extremely wasteful of both human and machine time.

One thing that I found absolutely insane about the DeepMind, Google's DeepMind's AlphaDev, the recent paper that they published, is that they use this AI model to optimize Google's vast data centers that they call Borg. So, this new solution that the AI came up with has been in production for over a year, and it continuously recovers, on average, 7% of Google's worldwide compute resources. The sustained efficiency gain means at any given moment, more tasks can be completed on the same computational footprint. So, kind of exactly what he's talking about in the paper. Reducing waste means you can get more done. The people that are going to win are the ones that can optimize and make it better.

Now, I know what you're thinking. I can read your mind. You're probably thinking, “Yeah, but why would you need to be doing it if it this just says the AI did it and it did it probably better than human beings could? Does that mean that all these agents will also be orchestrated by an AI?” And interestingly, this is where I'm actually seeing one of the sort of obstacles for AI developments right now. Currently, a lot of the AI agents that are being built, they have a problem with long-term coherence. When you try having them run on their own and pursue long-horizon tasks, after a while, they tend to fall apart.

AlphaDev, interestingly, wasn't a fully autonomous AI agent. It was really—I think of it as an LLM, right?—that's kind of like the the pilot, the thing driving it. They had a certain structure around it, including automatic evaluations of the outputs. But a lot of the inputs, as you can see here, they came in from the the human scientist/engineer. They came up with the initial prompts and configurations. They they chose the LLMs that need to go in here. They wrote some evaluators for the outputs, programs that would judge how good the outputs were. And they also provide the initial sort of database—like, “Here's what we know so far, here's that. Why don't you try these things out, and here's how you judge if you're um doing well or not,” right? So, the humans set up all of that. And AlphaDev ran this insane sort of evolutionary search, weeded out the the bad things, and the good ideas that seemed prominent, they would kind of have offspring, if you will. They would propagate to future generations, and then they would keep searching through those branching chain of ideas, and at the end, they would come up with the best program, the best solution. But again, kind of what he's describing, you could almost think of these people up here as those agent orchestrators, right? So, they're they're sort of controlling it and checking things. They're monitoring this insane AI system and guiding it along. This thing isn't running autonomously, not fully.

So, for the time being, at least, there there seems to be a—let's call it an obstacle, a speed bump—towards fully autonomous agents just handling other agents or whatever. That thing is that sort of long-term coherence, and a lot of people are are struggling with it, and who knows, maybe tomorrow a paper will come out that's, “Oh, like we solved it,” who knows, maybe, but for the time being, that seems to be an open problem. So, literacy in managing agents becomes the new Excel. Knowing how to break down a task, set a reward, audit a run, is going to be the baseline skill. So, for companies and people trying to take advantage of this coming AI wave—so, setting up a lot of agents and doing kind of A/B tests—so, seeing what works, what doesn't, and staying flexible and just looking at the data. If you have that as your culture, you're going to compound faster. If you tried to bolt AI into old workflows, you're not going to have a good time, right? Just like when Excel showed up, if you knew how to use Excel for a time, you were very, very special. But after a while, it was just table stakes. You needed to know Excel, period.

So, if you're able to manage AI agents, do these A/B tests, product management for agents, as he describes it, that's really going to be the in-demand skill and a massive competitive advantage. So, here's uh his website.

So, he's saying, “One of the most satisfying frameworks I found for understanding AI progress is the MER curve, which measures AI performance based on length tasks AI agents can complete.” So, we started with just tasks that are a few seconds long—to GPT-3.5. Let's call that 40 seconds long. And over time, here we have, for example, the sonnet 3.7. So, let's say maybe that's like a just under an hour. So, it can complete a task that that takes an hour to do, and they're looking at a 50% success rate. There was this paper recently out of OpenAI called “Paper Bench,” and in it, they took human PhD-level participants, people that were working in machine learning, PhDs in machine learning, and they had them compete against, for example, in this case, the GPT-1 model, and they were tasked with replicating papers that were published in the sort of machine learning field in in AI. Usually, you have a paper that proposes some experiment, for example, and usually you have to write out the code that runs that experiment, and one one thing that's really helpful would be to be able to have some sort of a third-party verification of the code. So, they would look at that paper, reproduce the code, right? Not just use the code that the people wrote because they could have mistakes, but just look at their paper, the idea that they're proposing and kind of build that code from scratch and then run it to see if they can replicate the results. If they do, that means that the paper is probably good. If they can't replicate it, well, there's a problem. You probably remember that floating semiconductor thing from a few years ago that everybody thought was a massive breakthrough, but then when people tried to replicate it, it it didn't work, and they were like, “All right, so no.”

So, in order to replicate these things, you have to understand what the paper's talking about. You have to understand, read the paper and understand the concepts and then write the code to kind of like make those concepts come to life, to to run that experiment. And so the blue line here is the AI agents, right? As you can see here, they're much faster in the beginning than the human PhDs in machine learning. They're so fast out of the gate, but there's this kind of plateau effect, right? For the humans, well, it takes us a while to kind of warm up and and figure out what we're doing, right? To read the paper and kind of understand what it's talking about. Then, when we start working, we—after about 24 hours of actual work time, so not just a day pass, but act 24 hours of actually working on this problem—we pass the AI agents. So, this is from the mer.org, the research kind of behind the the chart that says that the length of task the AI just can do is increasing. They're saying that these AIs, they can serve as useful tools in many applications. Yet, even the best AI agents, they can't really carry out like long-term projects or substitute for human labor. They're unable to reliably handle even relatively low-skill computer-based work like remote executive assistance. We covered a paper here not that long ago called “Vending Bench,” where a lot of these AI models were tasked with running a vending machine business. They're supposed to handle ordering, inventory management, pricing, etc., like product research, etc., and they did well. Some of them did better than the human baseline, but that's on average largely because like when they do well, they they do well, but very often they go completely insane because they kind of lose the plot, if you will. In one of the cases, one of them couldn't figure out when the inventory is getting here, so it closed the business, realized it was still being charged for some of the the some of the fees for the services, thinks that it's being defrauded, and attempts to contact the FBI. We covered this in a different video. I'm not going to go into it, but it it was hilarious. But the the thing here is is that these things currently are superhuman at doing kind of short-horizon tasks. These discrete tasks that have a beginning and an end, they can like do it incredibly well. Stringing that together into some sort of a long-term project currently is not happening. And again, maybe tomorrow somebody will fix that. Somebody will figure out a workaround. We have no idea. But currently, and and for some time yet, we're we're probably not going to have these things be fully autonomous over long-horizon tasks.

So, the research from MIRI is saying, “Extrapolating this trend predicts that in under a decade, we will see AI agents that can independently complete a large fraction of software tasks that currently take humans days or weeks.” But none of this happens overnight. So, the point of the article is that your ability to turn cheap intelligence and expensive resources into valuable products is what will matter. I believe everything here is absolutely spot-on, and I think the skills required here, they'll resemble a little bit a video game—like what we talked about—Starcraft or Factorio or however you want to sort of…

Visualize that, and the name of the game will be kind of taking this cheap intelligence. By the way, if that doesn't make sense to you conceptually, what that means is: throughout history, if you had the money to hire smart people, you could build a lot of cool stuff. But you needed money because you needed to, you know, hire that intelligence. And those humans, those smart humans, can't work 24 hours a day.

With AI, you can have it work 24 hours a day. It's not very expensive. With this paper bench example, right? So, how much would it cost you to hire a PhD-level machine learning scientist to work on your project for 24 hours? Not one day, but 24 hours of work. So, here in this paper, they're saying it costs $400 in API credits to run that agent on a 12-hour rollout, and the 03 that is kind of judging how well it does, right?

So, think about this: like how they did this paper is they actually created kind of like what we've been just talking about. They have this thing doing the work, and they have this thing kind of judging the work output, right? So, it's $400 in API credits to run it for 12 hours, and then $66 to have the 03 mini judge the paper output. And there we're talking about PhD-level machine learning work. Obviously, I think the—I think for a lot of the other less demanding tasks, it would be a lot, a lot cheaper.

So, if there's some project or thing that the market, the people out there willing to, let's say, pay $10,000 for, and you're able to create a team of agents, right, some of them producing the things, other ones maybe troubleshooting it, grading it, doing quality assurance, etc.—let's say you're able to have that project completed for $2,000. Well, that's $8,000 cheaper than it would currently cost to do that project, right? And for a while there, you might be doing really well just at doing that.

Over time, as more and more companies catch up, they figure out how to build these AI agents. They also will figure out how to produce it for $2,000. Then the trick will be to kind of figure out how to use less resources, less compute, how to make it better, etc. Maybe you can cut the cost in half to where it costs $1,000. And I think kind of those two things, those are like the stages that we're going to be in over our near-term future.

Step one is: figure out how to replicate the stuff that humans are doing with AI agents. And then once that becomes more and more commonplace, when that alone won't be as exciting and as sort of as rare, then the idea will be: how do we—how do we optimize it? How do—how do we make it better, cheaper, faster, etc.

So, what do you think about that? Do you think he's spot on? Do you think something's missing? I think this is what, you know, software engineering was when computers start popping up, and it was very in demand. It's kind of the hot new profession. Not a lot of people that know how to do it, and the rewards for people that are able to do it well are huge. Whether you're trying to start your own startup, your own business, or doing this as a career perhaps. I personally think this stuff is going to be very much in demand, and to me personally, kind of interesting work to be doing.

So, I don't know about you, but I'm going to start practicing those skills as fast as I can right after just this one quick game of Factorio. If you made it this far, thank you so much for watching. My name is Wes Roth, and I'll see you in the next