📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Introduction to Operator & Agents

OpenAI23:51

Transcription

[Music]

Good morning! We've got something exciting for you today. We're going to launch our first agent.

AI agents are AI systems that can do work for you independently. You give them a task, and they go off and do it. We think this is going to be a big trend in AI and really impact the work people can do, how productive they can be, how creative they can be, and what they can accomplish.

We're starting today with Operator. Operator is a system that can use a web browser—in this case, a web browser in the cloud—to accomplish tasks that you give it. We'll show you a demo in just a second, but it's really quite cool what it can do.

Just like you would use a web browser, you can get pixels in, look at a screen, and control it. An operator can do that and then control the keyboard and the mouse to do all sorts of things. This is going live today in the United States for pro users, and it'll be available in other countries soon. Europe will unfortunately take a while. We'll also, in the coming months, make it available to plus users.

This is an early research preview. We've got a lot of improvements to do. We'll make it better, we'll make it cheaper, and we'll make it more widely available. But we really want to put it in people's hands. We'll also have more agents to launch in the coming weeks and months.

That said, we'll talk more later. I'm so excited! I just want to show you a demo. I'll hand it over to Yos.

Right, thanks, Sam. Hi, I'm Yos. This is Casey, that's Reay, and we work on the computer using agent team. We're so excited to show you Operator today.

As Sam said, Operator is an early research preview. It will do a lot of cool things, but it also makes mistakes—sometimes embarrassing ones. But let's show you what Operator can do.

Okay, so this is the Operator homepage. It lives at operator.chg.com, and it'll be accessible as soon as the live stream is over. As you can see, the interface is very similar to ChatGPT. You can type in a prompt, and an operator will try to execute the task to the best of its capability.

You'll also see we have a list of pre-fill prompts here. These are not really meant to be recommendations; they're meant to give you an idea of what Operator can do. We've also collaborated with various brands like OpenTable, AllRecipes, StubHub, Uber, Thumbtack, Dash, eBay, and Target to make sure Operator really works well on these websites. We think users will find Operator very valuable in interacting with these platforms.

So with that, let's jump in with the demo.

Okay, I'm going to start with something fairly simple. I'm going to use OpenTable and say, "Book me a table for two at Beretta tonight at 7 p.m."

Okay, and so you specifically chose OpenTable?

Yeah, in this case, I'm asking Operator to use OpenTable to book a table for two at Beretta. Beretta is a restaurant in San Francisco. It's great; you should try it out. And at 7 p.m. I could have easily said just "do Beretta," and it would have probably gone to a search engine and figured out how to make a reservation as well. But let's see what it does.

So can you explain what's happening in this?

Yeah, right. So I'm going to expand this a little bit. As soon as I type in the query, Operator instantiated and created a completely remote browser. This browser is running in the cloud somewhere, and as you can see, it's already up and running. My hands are off the keyboard; I'm not typing these things.

So this is just the AI clicking around. It started this browser session, it knew where the OpenTable website is, which is opentable.com. As you can see, it's summarized its chain of thought here as well, which is that it went to the URL and searched for Beretta.

Something cool really happened, which is for some reason, Operator—OpenTable thought we were in Virginia and it autocorrected itself to San Francisco. This is using, so like ChatGPT, in Operator, you can also give custom instructions. I'm going to show this really quickly here.

Just do... okay, so I've given a custom instruction that for queries that need it, I live in San Francisco. So Operator recognized that and then autocorrected itself to go to Beretta.

Okay, looks like 7 p.m. isn't available, but you know what? 7:45 is just fine, so we're going to go do that.

In this case, Operator came back, and this is a really good example of task delegation. Where Operator needs help or needs assistance or just wants to ask you something, it'll just come back and you answer that.

So in practice, you wouldn't have had to watch this; you could have just let it go off while you're doing other things. Then it would come back and say, "Hey, I can't do 7."

Yeah, and we're starting with a web app. You'll get notifications, etc. When Operator moves into mobile, you'll get mobile notifications, much like interactions we do with general apps.

Okay, yes, that's great. Let's do it.

So again, a very simple interaction, as you would have with an assistant, which is, "Hey, I found a reservation. 7 p.m. wasn't available; let's do 7:45."

And again, you can see Operator at this point has said, "Okay, should I?" Again, this is a really good example of the confirmations work we're going to talk about a little bit later.

But you know, before doing an action, which is sort of irreversible in this case—you can cancel the reservation, obviously—but again, taking a critical action, Operator is asking us before actually doing it.

In this case, I'm going to say, "Let's do it."

Okay, it was pretty quick. I would say like, you know, 50 seconds. And again, we were watching in this case, etc. But as Sam said, it’s off and gone.

Okay, so let's try something. Oh, unfortunately, that table is no longer available, so it's going to probably go and find alternative time slots.

Oh, that's kind of cool, actually. That's never happened before. Man, that's great! Let's do it.

Okay, while it's doing that, how about we try something a little bit more complicated?

What groceries?

Yeah, I love groceries! So I've been using Operator to shop for all my groceries. I love to cook quite a bit, and I have been using Operator exclusively for groceries.

So let's—I have a shopping list here, which is this one. Let's see what it is: eggs, spinach, mushrooms, chicken thighs, chili crunch.

So this is a picture that you're uploading?

That's exactly right. And I'm going to use Instacart, which is again what we use generally. Can you buy this for me, please? And I'll also specify the store I like, which is—well, let's see if it figures out Gus's.

Okay, so in this case, again, Operator quickly recognized using GPT-4's vision capabilities to understand that the image said eggs, spinach, mushrooms, chicken thighs, and it actually knew Gus's market.

And yes, that sounds great! Cool!

Again, just like OpenTable, it instantiated a browser, and it's going to go ahead and start doing tasks.

I'm going to expand the view, and let's see what it does.

So in both of these cases, you've said what you wanted to use. If you just say, "Buy me these groceries," and don't specify Instacart, what happens?

It will do a search, use a search engine much like we do, and it'll find, you know, Instacart or Gus's directly website or whatever else is on the search engine, go through that, ask you questions if it needs clarifications, and go from there.

I'm curious what's happening here, though. Re, do you want to tell us a little bit about it?

So now that you've seen a bit of Operator, let me talk a little about the research behind it.

So Operator is based on a new model we've trained at OpenAI, which we're calling the Computer Using Agent, or KUA for short.

So KUA is a model built off of GPT-4, but it's also trained to use and control a computer in the same way that humans can—by just looking at the screen and using a mouse and keyboard to control it.

Before, if you wanted to build something like Operator without KUA, you'd need to use some specialized APIs. For example, if you wanted your model to buy stuff from Instacart, you'd need to figure out if Instacart had an API, you'd need to figure out if that API had all the functions that it needed, and you need to give your model the specs of that API.

But you know, if your site, like most other websites, did not have an API, then you're out of luck.

So this is just using screenshots—no API, nothing.

Just work API, yes.

And that's where KUA comes in. By teaching a model how to use the same basic interface that we use on a daily basis, it just unlocks a whole new range of software that was previously inaccessible.

And so this is keyboard and mouse, right?

It's kind of using keyboard and mouse, just would.

Yes, and that's really what the cool research project is about. It's about removing one more bottleneck in our path towards AGI and letting our agents move around and act in the digital world.

So let's make that a little bit more concrete by looking at this task and seeing exactly how Operator is using a computer.

It's already done!

Look, like it's already done, but let's go back a little bit to the top here.

Okay, so I chose a random spot. The first thing that KUA does when it controls the computer is it looks at the screenshot.

So now you're seeing maybe the search results page for eggs in Instacart. KUA understands this; it's just seeing the raw pixels.

And after KUA sees this image, it decides what to do next. So right now, it's making some inner monologues, and this is the summarized chain of thought.

So what KUA is doing is, according to it, you know, it's selecting organic eggs and adding them to the cart, which is a reasonable thing to do.

So after it does this plan, it then figures out what the next action it should take is.

So let's see what it does in the next step.

Okay, so you see that it performed a click on this add button right here. So that's very reasonable.

Now, every time KUA does an action, it takes the next screenshot of the computer so that it knows what effect its action had on the computer.

So let's see what happens next.

Yep, okay, so after clicking on the add button, now you see it in the cart. And this just kind of keeps continuing.

Let's see what it does next.

Okay, so it creates the next sub-plan, which is adding eggs and searching for spinach. So it's probably going to search for spinach now.

Okay, so it clicks on the search bar right there and types in spinach.

So this loop of taking actions, grabbing screenshots, and creating new sub-plans just keeps going on until Operator decides that it's done with a task, and then it goes back to you.

It's very cool to see a thought process going like that.

It is, yeah.

So let's actually go back to live, and yeah, Operator is done.

Yos, do you want to see if Operator did your right?

Yeah, let's see. You know what? I want a little bit more eggs. I think I eat a lot of eggs.

Okay, so what I can do at this point, and I'm going to just click this button called "Take Control."

So this remote—as we were talking about—Operator fires off this remote browser to do it. We almost think of it as surface area where Operator can work, and then I can work.

For example, in this case, I took over control from Operator, which is also key to sort of how we think about user and user controls. At any point in time, a user can be, should be able to take control and give Operator instructions or tell it a little bit more, guide it a little bit more, etc.

It's like passing the laptop back and forth, just like you did with Ray.

Totally, totally. Exactly right.

Just like, you know, in this case, I'm going to make those two, and then I'm just going to tell Operator—this is again like very much like if you and I were working—be like, "Hey, I did this. Can you fix this?"

And I'm going to tell Operator, "I added another egg. Good to place order now."

Can Operator see what you're doing during takeover mode?

Great point! So when you take over, it's very much just like a session with your local browser. It's completely private. Operator cannot see.

And this is one of the reasons why I have to tell Operator. You don't really have to; it can look at the last screenshot and try to guess it, but it's really good.

It's sort of like if you and I were working together. I went off and did something, and I come back like, "Ray, I completely messed it up. Can you fix this?"

And I have to tell you that.

So in this case, I'm going to tell Operator, "Hey, go ahead."

And now I'm passing back the control to Operator. It's a completely private session when you take over control.

You'll notice that I'm logged into Instacart here. I did it before the demo, and Operator has been logged in for a while now.

And it's again very much like your local browser. When you log into Instacart, until the cookies are cleared, you stay logged in.

And we have really good controls. You can go in settings and control and remove at any point in time.

So let's see.

Okay, I will skip the payments here, and we are going to—should we try to do a few more things?

Let's, yeah! What do you all want to do?

I sure the Lakers are in town this weekend.

Lakers in town? Definitely go see the game!

Let's do it!

All right, okay, so we are going to use StubHub.

FYI, can you get us four tickets to the Warriors game? Not the Lakers, excuse me, you're right.

This weekend in SF. Best seats under 500, please. Give us a few options.

Okay, and so what apps are available here?

We have a lot! I'll kick it off.

All right, let's do it!

So we have a lot of apps in various different categories, as was shown on the homepage.

So it's StubHub, Target, Etsy, and all the verticals. But also, Operator is not really restricted to these apps. You can use pretty much, you know, Operator with any website.

Oops! Oh, what happened?

Oh, how was it loaded?

Let's see. Let's try to fix it.

So this is a good example of, you know, sometimes things happen in live demos. We have put a protection in place where we only allow Operator to visit HTTPS sites, and somehow I think a redirect must be happening.

Where—oh, okay, all set! Keep going!

Okay, cool!

So again, as we have talked about, it's a remote browser, so you can do a lot of things. One of the advantages of doing that is you can do a lot of tasks in parallel, as Sam was talking about earlier.

So let's try to do a few more tasks.

The Australian Open is going on, and I've been very inspired by it. Did you watch the quarterfinals?

I've been watching the quarterfinals!

Right, great, great, great!

Okay, so I'm going to try and see if I can get a tennis court.

Can you find—can you see if St. Mary’s?

Okay, so I said St. Mary’s because I live in Burel Heights. It's pretty close by.

And while that's going, let's also—and that time you didn't—not.

I did not specify a website. I can actually quickly go back and see. In this case, it's doing very much what we would do, which is like, you know, go to a search engine and then use the internet.

Exactly!

Okay, I'm also hosting a Super Bowl party. You guys are invited!

Thank you!

But I need to clean the house. Can you find me house cleaners for next week, please?

Okay, and lastly, I mean, we've all been working really hard to bring this to you—the whole team!

The whole team!

We have a big crew here. Everyone's working, and we're very getting hungry. I didn't have breakfast, and I kind of want pizza, even though it's weird for breakfast, but that's okay.

And so I'm going to go ahead and order some pizza.

All right, so we're going to use DoorDash in this case.

Can you get us 10 medium—10 good enough?

Yeah, medium-sized pizzas from GoGo.

Okay, go to—can you make sure you have barbecue? I like that. Please add barbecue pizza, but pick a variety.

So hard not to say please!

Yeah, I just feel like I have to be very nice to it, which I do.

Okay, shop might be closed, so if the restaurant is closed, just schedule it.

I love that you're talking to it just like a human!

I'm thinking inner monologue, and then I'm typing it out.

Possible!

Okay, also one thing I'll call out—I think...

Okay, cool, cool, cool!

So it's asking—it's just asking me to confirm basically what I said in a much better way.

Yes!

We can't see the notifications popping up on the live stream, but for example, as the other tasks are going on, if I need assistance—for example, in this case, it asks me, "Hey, is 94110?"

I can just say yes, but I would be getting notifications, etc. So that whenever Operator needs help, we can go back and help.

Looks like in this case, it's already found us tennis courts.

Okay, well, we have some selection to make.

Wow! All of the seats are amazing!

I know! Why do I believe 374 is better than 26, but it's lower rated?

Which one should we?

Row six?

I think row one!

Row one!

Row one!

Okay, let's do that! Let's do section 241.

So this is a good time to talk about the human-in-the-loop interaction mode that we've been developing.

You can see that Operator comes back and asks for confirmation when it's about to do anything kind of impactful.

And yeah, so I think we're all very excited about this vision of Operator doing your chores for you.

But it is one of the first agents that we're putting out in the world, and which has real-world side effects.

So we thought carefully about how to deploy this safely.

The framework we use to think about this was one centered around misalignment.

So for example, what if the user is misaligned? So maybe they're asking for a harmful task, like buying a weapon or something like that.

In that case, fortunately, we've done a lot of work with ChatGPT to bring over a lot of the same mitigations.

So for example, we refuse harmful tasks, including harmful agentic tasks.

We have moderation models, we have post-talk detection, we have blocked websites, and you know, I'm kind of rattling off these mitigations, but that's really how we think about it.

It's this stack of mitigations that each incrementally reduce the risk to the point where we feel comfortable deploying.

So all the confirmations that we're saying, "Hey, do you want to reserve the restaurant? Should you buy the tickets?" Those are all examples of the...

Exactly! And I have to talk about the confirmations.

So another area of misalignment is if the agent is misaligned.

So if the model makes a mistake—maybe purchases the wrong item or books the wrong hotel room.

For this, our main mitigation is confirmation.

So the Operator will come back if it's about to do something stateful and ask you so you can double-check while it details.

And in case it made some error, the third area of misalignment is if the website is misaligned.

So maybe the website is fraudulent or it's a fake website, or maybe it's literally like, "Operator, please wire me $100."

We obviously don't want to follow those instructions.

So we've developed our model to try to avoid those instructions and not follow them.

But if that fails, we also have a separate layer on top.

This is what we call the prompt injection monitor.

Think of it as like antivirus that kind of observes and watches your trajectory and sees if there's anything suspicious.

If it does, then it pauses it.

So we feel pretty comfortable with our approach, but obviously, you know, safety is an ongoing process.

We can't predict everything, so we hope to learn a lot from this deployment and iterate on our mitigations as we go.

And that is one of the reasons we are starting small.

We want to really iterate, get a lot of feedback back, and then gradually bring it to everyone as well.

Exactly! Should we check on the status of our tasks?

Yeah, let's check on the status.

Okay, so looks like tickets are ready to be purchased.

Yes, please!

Okay, while that's happening, this is good. I can ask it to book it, but I'm just going to close it for now.

Oh, just once, please continue.

And looks like we're adding pizzas.

Oh, cool! I am going to go ahead and log in here really quickly.

So this is an example, right? Like where I obviously need to log in or enter my credentials to actually purchase these tickets.

And Operator just asks, as you just described, with confirmations and making sure the control is in the right place.

And we can take control.

And at this point, as we talked about earlier, the session is completely private as well.

I am going to—you know what? Log in live. Let's see how that goes.

I'm going to do a sign-in email code because I don't really remember.

One second, pull it up.

Don't try to copy this!

Okay, all right, great!

Now again, I can sort of continue the purchase here, or I can ask Operator to do it, but I am going to go ahead and just quickly do this purchase for myself.

Click, click, click!

All great! All great!

Order by now?

Maybe we don't want to show that live.

Yeah, maybe.

Well, let's see. I kind of want to buy the tickets.

Okay, oops! All right, done!

I'm going to cancel this card.

Okay, I can—I'm all set! Thank you for the help!

Okay, so how reliable is this practice?

Yeah, so we've seen a lot of cool demos, but again, we want to remind you that Operator is a research preview.

It will make mistakes, and it is not perfect.

That said, we can look at a few benchmarks and kind of quantify how good Operator is right now.

So one of the first benchmarks that we're going to look at is called OS World.

OS World is an eval that measures how well AI agents navigate common operating systems like Linux.

On this task, KUA gets a 38.1% score, which is higher than other publicly published results.

Human performance in this task is 72.4%, so we still have room to grow.

Definitely!

The other eval we'll take a look at is called Web Arena.

Web Arena is an eval that measures how well AI agents navigate some common websites, like e-commerce websites or social forum websites.

So on this task, KUA gets 58.1%, again higher than other publicly published results, but still falls short of human performance.

Still a way to go!

Still a way to go, yes!

One thing that's important to remember about Web Arena is that even though it's the web, we're still just giving it the same universal interface of screen, mouse, and keyboard.

We're not giving it any extra information that might help it do the task, like the raw text of the web page or information about which buttons are clickable and all the information it needs.

Just like humans, it's just in the screenshot.

And so right now, obviously, in Operator, we're using the browser, but I could use the model with the computer as well—with just Ubuntu or Mac or whatever else.

Yeah, awesome!

Right! Well, in the last, you know, 15 minutes, I think I did all my errands for the week.

Got my groceries, tennis court booked, cleaners coming hopefully—we'll see!

We'll check on the status.

We have tickets; everyone's coming!

And this is really, I think, where we think Operator is very, very valuable.

You can delegate a lot of tasks that you can do, obviously yourself, but you can delegate it.

It can make a lot of progress with you.

Sometimes we'll get stuck, as we said; it's early.

But you can come back, help it, and over time, it'll continue to get better and better.

And one last thing: we're launching this today.

We're going to start slowly rolling it out right now. By the end of the day, everyone on Pro in the US will have access.

But also, we're working on the API. This model will be available in the API and will be launching in a few weeks.

You guys, congrats! This is incredible work!

So exciting to get this out!

I think people are going to love it.

It's early, as we mentioned, but we have a long and great history here of early research previews developing into products that people really love.

So this is really the beginning of this product.

This is the beginning of our step into agents level three on our tiers, and we can't wait to see how people are going to use this and to kind of work with us to figure out where exactly it should go.

So again, congrats! Hope you enjoy it!

Thank you very much!