Transcription
Do you think agents are promising? We have to talk about this. This was, uh, this is like the excitement of the year that agents are going to re—this is the generic hype term that a lot of business folks are using. AI agents are going to revolutionize everything.
Okay, so mostly the term "agent" is obviously overblown. We've talked a lot about reinforcement learning as a way to train for verifiable outcomes. Agents should mean something that is open-ended and is solving a task independently on its own and able to adapt to uncertainty.
There's a lot of term "agent" applied to things like Apple Intelligence, which we still don't have after the last WWDC, which is orchestrating between apps. That type of tool use is something that language models can do really well. Apple Intelligence, I suspect, will come eventually. It's a closed domain—it's your Messages app integrating with your photos with AI in the background. That will work.
That has been described as an agent by a lot of software companies to get into the narrative. The question is, what ways can we get language models to generalize to new domains and solve their own problems in real time? Maybe some tiny amount of training when they are doing this with fine-tuning themselves or in-context learning, which is the idea of storing information in a prompt. You can use learning algorithms to update that.
Whether or not you believe that is going to actually generalize to things like me saying, "Book my trip to go to Austin in two days. I have XYZ constraints," and actually trusting it, I think there's an HCI problem coming back for information. Well, what's your prediction there?
My gut says we're very far away from that. I think OpenAI's statement—you, I don't know if you've seen the five levels, right? Where chat is level one, reasoning is level two, and then agents is level three. I think there's a couple more levels, but it's important, right? We were in chat for a couple of years. We just theoretically got to reasoning. We'll be here for a year or two, right? And then agents.
But at the same time, people can try and approximate capabilities of the next level. But the agents are doing things autonomously, doing things for minutes at a time, hours at a time, etc. Right? Reasoning is doing things for tens of seconds at a time, right? And then coming back with an output that I still need to verify and use and try to check out, right?
And the biggest problem is, of course, like, um, it's the same thing with manufacturing, right? Like there's the whole Six Sigma thing, right? Like, you know how many nines do you get? And then you compound the nines onto each other. It's like if you multiply, you know, by the number of steps that are Six Sigma, you get to, uh, you know, a yield or something, right?
So like in semiconductor manufacturing, tens of thousands of steps, 9999999 is not enough, right? Because you multiply that by that many times, you actually end up with like 60% yield, right? Really low yield, yeah, or zero.
And this is the same thing with agents, right? Like chaining tasks together, each time LLMs—even the best LLMs in particularly pretty good benchmarks—don't get 100% right. They get a little bit below that because there's a lot of noise.
So how do you get to enough nines? Right? This is the same thing with self-driving. We can't have self-driving without it being like super geo-fenced, like Google, like Google's, right? And even then they have a bunch of teleoperators to make sure it doesn't get stuck, right? But you can't do that because it doesn't have enough.
And self-driving has quite a lot of structure because roads have rules. It's well-defined. There's regulation. When you're talking about computer use for the open web, for example, or the open operating system, like there's no—it's a mess.
So like the possibility—I'm always skeptical of any system that is tasked with interacting with the human world, with the open, messy human world. We can't get intelligence that's enough to solve the human world on its own. We can create infrastructure, like the human operators for WEO over many years, that enable certain workflows.
There is a company—I don't remember what it is—but that's literally their pitch. Yeah, we're just going to be the human operator when agents fail, and you just call us and we fix it. Yeah, it's an API call, and it's hilarious.
There's going to be teleoperation markets when we get human robots, which is—there's going to be somebody around the world that's happy to fix the fact that it can't finish loading my dishwasher when I'm unhappy with it. But that's just going to be part of the Tesla service package.
I'm just imagining like an AI agent talking to another AI agent. One company has an AI agent that specializes in helping other AI agents. But if you can make things that are good at one step, you can just stack them together.
So that's why I'm like, if it takes a long time, we're going to build infrastructure that enables it. You see the operator launch. They have partnerships with certain websites, with DoorDash, with OpenTable, with things like this. Those partnerships are going to let them climb really fast. Their model's going to get really good at those things.
It's going to proof of concept that might be a network effect where more companies want to make it easier for AI. Some companies will be like, "No, let's put blockers in place." Y—and this is the story of the internet. We've seen it now with training data for language models where companies are like, "No, you have to pay."
That said, I think like airlines have a very—and hotels have a high incentive to make their site work really well, and they usually don't. Like if you look at how many clicks it takes to order an airplane ticket, it's insane. I don't—you actually can't call an American Airlines agent anymore. They don't have a phone number.
I mean, it's horrible on many fronts. And to imagine that agents will be able to deal with that website when I, as a human, struggle—like I have an existential crisis every time I try to book an airplane ticket. I don't think it's going to be extremely difficult to build an AI agent that's robust.
But think about like United has accepted the Starlink term, which is they have to provide Starlink for free, and the users are going to love it. What if one airline is like, "We're going to take a year, and we're going to make our website have white text that works perfectly for the AIs every time anyone asks about an AI flight?"
They buy whatever airline it is, or they just like, "Here's an API, and it's only exposed to AI agents." And if anyone queries it, the price is 10% higher for any flight. But we'll let you see any of our flights, and you can just book any of them. Here you go, agent.
And then it's like, "Oh, and I made 10% higher price. Awesome." Yeah, and like am I willing to say that for like, "Hey, book me a flight to C Lex," right? And it's like, "Yeah, whatever."
I think, you know, computers and the real world and the open world are really, really messy. But if you start defining the problem in narrow regions, people are going to be able to create very, very productive things and ratchet down costs massively.
Right? Like now crazy things like, you know, robotics in the home, you know, those are going to be a lot harder to do, just like self-driving, right? Because there's just a billion different failure modes.
But agents that can navigate a certain set of websites and do certain sets of tasks or like look at your—take a photo of your grocery, uh, your fridge, and or like upload your recipes, and then like it figures out what to order from, you know, Amazon, SLH Foods, food delivery—like that's going to be pretty quick and easy to do, I think.
So it's going to be a whole range of business outcomes, and it's going to be tons of sort of optimism around people can just figure out ways to make money. To be clear, these sandboxes already exist in research. There are people who have built clones of all the most popular websites of Google, Amazon, blah, blah, blah, to make it so that there's—and I mean OpenAI probably has them internally to train these things.
It's the same as DeepMind's robotics team for years has had clusters for robotics where you like interact with robots fully remotely. They just have a lab in London, and you task to it. It arranges the blocks, and you do this research.
Obviously, there's texts there that fix stuff, but we've turned these cranks of automation before. You go from sandbox to progress, and then you add one more domain at a time and generalize.
I think in the history of NLP and language processing, instruction tuning in tasks per language model used to be like one language model did one task. And then in the instruction tuning literature, there's this point where you start adding more and more tasks together, where it just starts to generalize to every task.
And we don't know where on this curve we are. I think for reasoning with this RL and verifiable domains, we're very—we're early. But we don't know where the point is where you just start training on enough domains and poof, like more domains just start working, and you've crossed the generalization barrier.