📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Inside OpenAI's Sora: Surge to #1 App, Key Product Decisions & How Video Models Learn Physics

Unsupervised Learning: Redpoint's AI Podcast1:03:24

Transcription

I feel like I have the main characters of the AI ecosystem.

People are very creative. It's a social network. So, social networks are about people and relationships. Even if you change the ratio of creator to consumer by 1%, like that is a massive impact on the world. Celebrities and rights holders are going through like a very fast evolution of understanding this technology and like how to use it.

Jack is posting bangers. By the way, I think the first scientific breakthrough that comes as a result of simulating some phenomenon in video models would be an insane milestone, right?

You have a prediction of when that might be. I would be shocked if we are not making breakthroughs like this by Sora has taken the world by storm this past month. Uh it's number one in the app store has spawned endless hilarious videos and I had the privilege today of speaking with Bill Rohan and Thomas three of the main folks at Sora behind this magic at OpenAI. It was just a really fun conversation. I got to ask them everything that I've been thinking about that Sora spawned. We talked about their reactions to Ben Thompson's take on Sora as well as kind of the key product decisions that factored into how they thought about making it into a social experience. We talked about uh the state of video models, how they progressed over the past years and what's next future milestones including novel physics discoveries and their predictions for timelines on those. Uh we also hit on how video models in general will fit into the broader OpenAI ecosystem including chat GBT over time. Just a ton of fun to get to talk with these three brilliant folks about all these things in the video ecosystem. I think folks will really enjoy it. Without further ado, here's the Sor team.

Guys, I'm so excited to do this. Thank you so much for uh for for coming in here.

Thanks for having us.

I feel like I have the main characters of the AI ecosystem, [laughter] which is like as a podcast host, a fun thing to to get. So, thank you guys. I feel like a bunch of different things we'll want to hit on today. You know, maybe to start, I feel like there's been all these, you know, stories told of the chat GBT launch and how it was kind of not expected to be anything big and and out of nowhere uh turned into this just, you know, massively used consumer product. I'm curious whether like the same was true of Sora or I guess like did you guys kind of expect this reaction? Uh any kind of fun stories from that first day?

I mean I didn't expect to be number one in the app store for like a month. Um but it's you know one of those things you look at in retrospect like the research team absolutely cooked. We have a knack at opening eye of creating these viral moments really well and so you know h happy to have done it but I think it it surpassed my expectations.

I think Bill also has a hopeless optimism about him. It's not hopeless. Actually, [laughter]

not hopeless.

I think it was like there was a meme. I saw it the first uh it was a couple days after launch where it was like number three in the app store and then Bill was like demanding number one. I think Sam was putting you in the in the cage or something [laughter] like until you were number one.

So my most like 2.5K

and uh [laughter] willed it to life.

You know what I I I'm always I've worked a lot of different social media products in

Instagram for a while.

Yeah. and some successful, some failures within that ecosystem. So, I I tend to be very grounded and realistic and all this sort of stuff. Um, but I have to say like uh that optimism is actually very important. We were we ended up being number one on the app store which was wildly crazy for my expectations. Um, but it's delightful to see.

So, I imagine there's an aspect of this where you kind of have to reserve some sort of GPU capacity, right, to based on on what you expect. And so, I always wonder like how in the world do you figure out?

You know, it's it's not really a science. Uh, it's [laughter] a lot of, you know, I think everyone um feels computed across the entire industry, not just with an open AI. And certainly when we launch an entirely new product surface, right, that is as computensive as video, it requires some pain elsewhere in the company. I think one thing that's really nice about OpenAI and like to Sam's credit, uh like everyone kind of feels invested in everybody's success across the board, right? Certainly when chat GBPT uh you know releases like amazing new visual generation features we want to make sure those are great and like people are having an awesome experience in chat and likewise the company really feels invested in source success people see that you know we need to take these new bets and it's important for the company's longevity not just to have a super successful LM facing consumer product but also the winning one in video and everyone's willing to chip in you know do your part in uh in wartime

when did the idea to like create Sora as an independent app you know come about was that always obvious from the beginning it

it wasn't always obvious I think there was a clear vision vision and there has been like you know a great level and of ambition in terms of what this technology can be and how it can transform UGC and and media generally and so then if you pair that with what we were actually experiencing with the model with like hey there's something very special in multiplayer mode we're seeing this emergent behavior even before Sora 2 in terms of like what what is AI enabling people can participate in trends extremely quickly and create these trends extremely quickly like we've seen that with other video generation software in terms of like this is flooding social media more and more. Um and then with image genen when we launched image genen it's amazing it blew up in chatty but people were just putting themselves in these amazing scenes and so something about seeing yourself casting yourself in these videos and we just wanted to take that one step further like do it with your friends and we felt magic internally and I think once we felt that it was it was a very easy decision. The one thing to add also is it's it's not like inconceivable that maybe these the surfaces come closer together, but chat GBT does feel very single player and that's sacred in a way right now like it can feel a little jarring if you you you know you don't know when you do something in Chat GPT is this public or not. So we also we're sensitive to that as well. Yeah, there's actually there's a long journey. Yeah. Of uh

you know this year for me at least was like just trying to like think about what social means in this kind of bigger context and we had lots of wacky prototypes along the way. Um I'm very delighted it ended up in Sora but it was not clear to me like product journey is always extremely

what was the strangest thing you threw out?

I mean we we had a prototype of just a pure social media stream inside of JP um and we're trying it out internally. Um one of the things we saw, so we I was just like, it was like, think about social, okay, first thing that comes to mind, well, we have chatgbt, why don't we just put some social features in it, see what it looks like? And when we started using it, we used it internally. It happened to be right before the image gen. And what we were seeing, we had these like, it was like kind of like threaded sort of Reddit style. Um, and we would see these very long chains where what would happen is like somebody would put up an image and then somebody would be like, "Okay, it's smoking a cigar now." Like, say it's a duck. It's a duck with a cigar. Now it's upside down. Now it's red. And you're like, it was really cool to see it evolve. And it was like this these little meme chains. Um, and that seemed like something very special. And I'm thinking about more sort of what Rohan was saying is like that's actually something you can't really do without generative AI because it's too hard to create. It's too hard to remix to to to riff on something. Um, and so that was like the wackiest I think one that we had. Um and again through this long snaky product journey the Sora models were getting pretty good. Um we always wanted to do something really ambitious with them to really mass mass deploy them and fortunately it just kind of like fit like a glove in terms of you know image gen is very cool but videos are even even cooler in some ways especially in this context and

it sounds like from the image work you kind of knew cameos would be this like killer feature for uh for these powerful models. I don't know if we knew really [laughter] interesting.

It's just I mean,

yeah, it's one of those things in retrospect. Yeah, it's obvious, but I think this idea was in the ether, you know, from from when the research team was actually building the model. It wasn't until we actually like

had it like one of our engineers on the team, Bobo, was like, "Everyone send me a video of you saying, "Hey, Sora, it's Phil." Yeah. Hey, Sora, bring me to life. Um, we like in a Slack thread, literally. He uploaded all these in the back end. Then you could just tag the person. We had we had like talked about this idea. Yeah.

And then finally, like everything changed that day, I think. Like we all just started cameoing each other. We didn't have a verb for this at the time. We just started tagging each other. And like it happened before we even noticed. It wasn't like boom, this is it. Just like in a couple days I feel like we were like our feed is just [laughter] came like what's going on like and we were all using the app.

We actually had a moment where we're like is this bad? But yeah, wait a second. Like, did we just like we had interesting stuff before and now it's just people like it's like wait a second no that's actually amazing that's what we need.

Yeah.

One of the coolest parts of of building these products is you release it day one and I'm sure people start immediately using it in like a million different ways that you never intended. Like so it's been like the most surprising way you've seen people use or capabilities that you didn't know actually existed.

Some people were like putting themselves in motivational situations like like envisioning themselves in situations they wanted to be in. I think anytime I see a a style transform, it's maybe a too small of a word for this, but like you know, you take something, take a person, uh there's one review that I I didn't I couldn't believe this would work and ended up being one of the most remixed videos where I just had a kid opening a package for Christmas and it's like it's a Bill People's action [laughter] and it looks like Bill People. This is what's keeping me on the leaderboard video. I'm like 25th or something. [laughter]

Bill People's action. It's a it's I like it's it's remarkable that the model understands that and understands the context of just like from a couple of uh numbers can put you in a completely unusual scene. I was amazed that that emerged from the model. I see those type of things in my feed all the time. Not exactly the same, but it's it's always some some riff on that. That's very cool. Um and yeah, any anything that's massive transformed the claimation ones I love the video game ones. I think it's a little underexplored right now where you have like a Lucas Arts Adventure, but it's actually Rohan or Bill or something. Um, those I just eat that up. Immediate like see what the secret emoji is on those ones. Yeah. [laughter] Yeah.

Some people are just getting insane like the Ka Kaumu Matsumara like just like

things we we didn't even get close to that level of like stylistic output. So, just seeing the creativity of the world.

Yeah.

Yeah. I think after we launched storyboard, which lets you generate up to 25 seconds video, that's when like the quality bar for me like really spiked. Like I I'm actually surprised how in like one shot out of this model, you can just get unbelievably like cohesive stories. Like this is something that you cannot even like get with like a 100 attempts at like Sora 1. This is like really a new thing in Sora 2. It just really is reflective of the step function increase in intelligence.

You know, on on TPN, I think you shared that like 70% of Sora users are creators, which is mindboggling. I'm sure it's it's still super high. You know, one thing I thought was interesting around the discourse around Sora is when it was initially released, I think Ben Thompson wrote this piece and he was like, I'm skeptical of [laughter] Sora. Like, you know, most people don't want to create, they want to consume. Um, that's been the behavior of every other product. And then I think he like issued me a call. He was like, "Okay, I was wrong. [laughter] This is a cool product." Um, curious like what your reaction was to that and and you know, you think this level of creation kind of persists. Yeah, I mean we really designed the app from the ground up to be focused on the creation side and this was really our core hypothesis going into it. Like there's so many ways existing social media networks are like really amazing and people get like so much joy out of them and that can come from the consumption side. At the same time, I do think probably the worst like harms and like the worst case scenarios come from like the doom scrolling and like we really wanted to figure out how do we kind of avoid this very dystopian future of just like a slop feed that's like kind of optimized exactly to what you want to watch. It's very detached from your friends. It feels very lonely and isolating. And we kind of tackled this across the board. So like really I think the the single most important thing we did there was really with Cameo. Like Cameo is really what makes a gen feel personal to you and makes it feel very human in a way that just kind of like text to video or just like all these very simple ways of prompting these models maybe doesn't. Um and then we also did a lot of work Thomas specifically uh on the recommener side. Right. So like really the the way that these things really go off the rails if you're you're not careful or being kind of like a good faith actor is if you design this Rexus in a way which is sort of like maximally uh clickbaity/engaging and like kind of incentivizes the student scrolling behavior and Thomas did a lot of like very pioneering work to really rethink what like a modern Rexus stack looks like in this product to oriented towards creativity over consumption.

Yeah. And some of that is actually it's kind of self-fulfilling in a way or like at a circle. But the fact that you can participate in the network means that you can also encourage like you that's now an optimization objective that can be in there. And I think it's a very healthy one if you're actually making a decision to oh I'm going to remix this generation. Um that's like a very active behavior. It puts you in a very creative mode that's not that whole just consumptionoriented thing. And so that was the idea behind the Rexus is like because it's so easy to create, we can actually encourage people to create in a way that you don't normally optimize for because uh I mean having worked at Instagram, the idea that you saw this photo or this video and then you help you know opened your camera and then you took a photo and you have to attribute it back to the original photo you saw. It's impossible. It's or it's very close to impossible. You can get a little bit of signal out of that. Uh but in this world you can get signal out of everything. So, I think it's a it's a very uh nuanced but kind of interesting sort of behavior that you can see.

Yeah. And like on Ben Thompson, I think his arc, you know, I was crushed when he [laughter] when he first uh you know, I listen to Strateetri every morning and so I was crushed when his when he gave his first round of feedback. But his arc makes sense, right? Like he is a consumer. Coldstar Network is very hard problem. So the feed wasn't super dialed in. He's an early adopter too. So early adopter without like a feed that has had any time to bake. It wasn't that interesting. So, as a pure consumer and not a creator, I think he jumped in there. I think he had a quote about it just being like a Sam Alman advertisement [laughter] or something, which the feed did feel like that for a little bit to be fair. I mean, I thought it was funny, but um but then I think when he had his kind of like, you know, he brought it back to his mayulpa, he had a good point which is like even if you change the ratio of creator to consumer by 1%, like that is a massive impact on the world, right? And that's like such a different product. Um, and I think that was like a good summary of why we think Zora is special and what he might have missed the first time around because he is particularly like a pure consumer and not a creator

um, in his everyday life.

I do think it's also it's we play with this in some of our early prototypes um, off Sora. But the idea that there's a human creating it versus there's like a bot creating it is a very very foundational difference. It's like kind of hard to appreciate, but if you took the Sora feed today, may like maybe excluding cameos [clears throat] and you just stripped the person that posted it, it would be very uninteresting in my opinion. Like part of the essential thing is that somebody has looked at this and decided this is like kind of like they're putting their stamp of approval on it. They're also involved in the creative process. And I think that is it's like easy to think about. I've seen a few of these social networks that are completely AIdriven, but they last for like five minutes of interest. I get a like from a bot and I'm like, what does it like mean? I don't really understand. Whereas like if Ron likes my content like oh yeah you know I know what that means like hey [laughter] you saw my thing like did you like it like so

right something almost that's so visceral about the fact that a human created it because they're in the they're in it right other other tools you know one thing that's also been really interesting to see uh as you know sor has been out it's just like I feel like celebrities and rights holders are going through like a very fast evolution of of understanding this technology and like how to use it. Can you talk a little bit about like what that journey's been like over the past month and you know where where most folks are today?

Yeah, I mean we've chatted with kind of every facet of society since launching this and you know I I think if you rewind the clock like a month ago you know most people in the world really did not know video gen was a thing. Certainly they didn't know

you knew it was going to be number one but I guess higher [laughter] and Sora I was very confident if you go through my gen history. um uh over time, you know, we've really like just had the opportunity to sit down with these people and and hear out both where they're really excited about this platform, right? I mean, there's a huge value proposition here, especially for rights holders. We launched character cameos yesterday, right? You can imagine uh if you have some IP that people love and any kid now can generate a gen with them in that that character. I mean, that's going to be huge for these rights holders. Uh and we're also hearing them out on their concerns, right? like making sure that they have uh a lot of input on how their characters show up. They want restrictions. They don't want it to just kind of be a free-for-all.

Yeah, I think a pretty clever UI for uh for setting the basic rails around that. I'm sure that I'm sure the real celebrities have have more sophistications [laughter] flow and if there's like a celebrity who's particularly uh you know anxious about getting on the platform, we really step them through exactly like how they should think about setting these things. But when we kind of walk people through, you know, what we have on the docket and the features we've already integrated with character cameos, cameo restrictions, etc. There's honestly a lot of excitement around this. And, you know, just earlier today, we announced that we're starting to bring monetization to Sora. We imagine in the future this

very nice way to announce that, you know, 30 minutes before you come out, you got to break. This is like holding it for days right before this episode.

This exact moment. Um, and uh, you know, we are going to start piloting new ways for rights holders to monetize their content, and we're going to really prioritize folks who have invested in the platform from the beginning, have gotten on now, and we think amazing things are going to happen from that.

Have there been any specific like celebrities or rights holders that you're like, "Oh, they're they already get it." I mean, Mark Cuban immediately added cost plusd drugs.com to his cameo instructions. I feel like he,

you know, was first to figure out what could be like a cr, you know, an amazing industry in of itself in terms of how you think about, you know, brand advertisement and stuff like that. Um, I mean, there's been a lot of people like we we have a a moderation filter that prevents you from creating a public figure. So when you go through the cameo flow, uh, a lot of public figures actually get blocked, [laughter] um, because we don't want, you know, we're very sensitive to impersonation, that kind of stuff. So, uh, I think it's funny because we get a lot of inbound because they're like, we can't use your product. Um, but the amount of inbound we got has been amazing. So like tons of people are actually trying this out, using it, having an amazing time. Shaq incredibly creative. Like we talked to his team and they're just like, Shaq's just having a blast just using [laughter] like generating. Like I don't I don't even know if Shaq cares about like the name he's making for himself on this new network.

He is just genuinely having a creative fun [laughter] time.

Shaq is posting bangers by the way.

Go follow Shaq.

Um

you just got to open it up though.

Yeah. [laughter]

You've had both Google and Meta release, you know, similar kind of product, you know, products around the same time. You know, everyone's kind of racing forward on these video models. Wonder like what you make of those releases and kind of how you think this space evolves over time. Clearly there's a few folks with very good video models today.

I think the future is very bright for all these models. Like I think there's there's no world I think so will stay on top because of Bill's fine work. But uh you know like uh this is obviously a new medium and it's a new creative expression. I think the the industry will evolve completely and this is going to be a big part of that use case. I think Meta, Google etc. are right to be like I'm sure there's a code red or somebody joked yesterday there'd be a code purple or you know they have to invent a new color [laughter] for this one. Um, and so like I think that's that's actually correct because they're seeing something emerge here that has been created with this uh a magical technology. That said, uh I haven't been overly impressed with like some of the productization of the work so far. And I really think that goes back to our our idea of like people embracing people. People are very creative. It's a social network. So social networks are about people and relationships. And the thing that really enabled us to to move forward is this idea of Cameo and bringing bringing this like creative aspect to your life with your friends. Um, and I think that was kind of missed. I'm not surprised. Um, I'm not surprised both on like a model level, but I'm not totally surprised because it was just a it was actually a very again nonlinear snaky journey to get to this point. It wasn't like I think there were ideas of doing cameo, but there was also other ideas where like what does remix mean in this context? Like what does it We had one that was always weird. I don't think nobody else likes this one, but I always liked it, which was like you would you take a video of yourself reacting to the AI video kind of in the corner. There was a reaction. We should bring that one back.

I even generated like a dog reacting to myself or something. [laughter] It was cool. Uh but like we we're trying things. We tried a lot of different stuff. Even like the the UI, it was early on Joey had this idea for image gen where it was just a simple UI where you had your contact list and it was like you just like pick and choose people to be in the gen. That kind of makes sense, but it was not

it's not obvious. Of course, it's obvious in retrospect like that's why it's code purple. Uh but like it's it's not obvious when you go through this whole thing. There's a lot of micro decisions along the way that um can make or break your product with Instagram. I remember talking to Mikey, one of the founders there, and it like the square photos was like a limitation that they just imposed. I think they probably because they didn't want to deal with like aspect ratio nonsense and uh that was actually essential for its success. So, long story short, I think this space is going to get very very competitive very very fast. Um I'm reasonably confident we're going to stay on top. Hope so. And uh I do think you need to embrace people in this whole process and I I think empower them with creative creative tooling and that's what we're doing.

And it seems like on the product side you're continually lean into this. I know I think you guys have talked about like kind of focusing on communities and like other kind of ways in in the product. Talk a little bit about like how you kind of see yourselves leaning into this even further over time.

Yeah, I it's this is cold start social network impossible. Hardest problem ever. [laughter]

Fun one to work on.

Yeah, it's it's fun for sure. Especially with Oh, yeah. And by the way, it's a completely new medium. You're like, "Okay, [laughter] let me think about this." Uh, so we didn't really know exactly how things would evolve. And, uh, we gave it a very familiar form factor. You know, it looks very similar to other full screen media player apps. It defin has definitely has a different feel, but we did have this hypothesis and I think it's kind of proving correct that this is much more fun with your friends. And that aspect definitely is in the product. It's emphasized in the recommener system, but I don't know if it's completely emphasized in all the ways that it could be. Um, so I I see us lead leaning into that a lot more um over time. I think the public feed is also incredibly important and um it's a kind of a way of even giving you inspiration to be like, "Oh, actually, well, Shaq's here. It' be pretty cool to be with Shaq or like have a riff on Shaq's, you know, Marilyn Monroe meme or whatever he's doing today." [laughter] Uh but like um so I think there's some product work to be thought through and I'm also just excited about like you know when you think about like okay what does this technology enable that is more fun with your friends. There's probably things that we haven't thought of yet that we can actually lean into. Uh I can't say exactly what they are. Um certainly we're more emphasizing the DMs feature over time because we think that's like can be a very magical moment but you can imagine things that happen in small groups kind of like being very exciting. Even even large groups I think could be a very exciting angle. Um certainly our OpenAI group before we launched was very cool. Um you know it's a very dense connected uh large network that people were just having a lot of fun with. Um and I'd like to be able to support that in the product in the future as well.

That's interesting. One thing I'm struck by is just clearly like the way you put the feed algorithm like basically determines a lot of this behavior in the in the ecosystem, right? And like what you encourage there.

There's a famous quote. I'm trying to remember who did it, but it's like show me the incentives, I'll show you the outcome. And so this never been more true than in these large recommener systems. Uh and certainly when I worked on Instagram, there was there was a very explicit decision to prioritize your friends. We we chose a lot of things to not uh introduce unconnected content in the main feed. I worked on the explore feed, which was a secondary surface, but uh they made that a very explicit decision. Uh and it's kind of drifted over time for the reason that people are posting less and all that sort of stuff. But if you look at your Instagram feed, it's kind of kind of wacky. I'd say right now it's like it's pretty uh disconcerting. Uh same with X. And uh I think because we've brought this barrier to creation down so much, uh there's a good chance that we can actually emphasize that. It's very funny that AI videos are what bringing you more connected to your friends is a [laughter] little bit unusual, but like yeah, the hottest densely connected network is an AI gen. But I'm all here for it. Yeah.

Yeah.

I think it's very like like in some ways a very like heartening use case for AI, right? I think you know one uh one thing I was struck by is just like it's it's just fun in a way that like I feel like so many of the things that have been built on top of other open products are just very serious of like you know it's like every founder out of college wants to go build like a very specific vertical like enterprise app and it's like there haven't been there actually really haven't been that many consumer products that have been built I think a lot of people were playing around with this like how do you you know have AI generations and people in the same you know I feel like character AI was even riffing on some of this stuff uh in the early days and I feel like you guys have kind of nailed the first product around it and it's just cool. I mean, I don't know. It's like good cool for the world to see.

I mean, it wasn't like there's a whole crazy flow you have to go through to get your cameo, right, of like an audio challenge and moving your head around. There were many moments where Thomas and I were like, we're cooked. There's no way there's no way people are getting through this.

I mean, you're talking about funny launch stories. We had [laughter] cuz we were trying to dial in the cameos and we're like there I don't know, there's a bit of voodoo involved in that whole thing where you stand in this particular poll.

Oh, yeah.

The time of the day is 400 p.m. [laughter]

Literally, it was this is that we all went to this poll. We had recorded cameos of ourselves and somehow we ended up at a at a flow which actually is the optimal flow but is impossible impossible to do where you have to say the words while you move your head. [laughter] Um that was that was actually the opposite.

Yeah. It was like say a sentence move your head around [laughter] like this like

quick brown [snorts] It was insane. Uh but I mean it would actually produces slightly better better cameos. And so anyway that was that was a funny launch moment. We had to dial that back. Be like I don't think anyone's going to get through that one because we can't get through it.

Yeah. Is this a secret pro tip that my cameos would be better if I if I background looks very promising actually?

Yeah, good to know.

Lighting a very private place and uh and just roll my head in.

Exactly. [laughter]

One tension I imagine you get dragged into immediately because you're focused on creators is like

there's just such a range of sophistication of focus. We were talking about this before. There's like the kind of pure consumer type of creators that just want to like be able to remix something as easily as possible. Then you have these like proumer type use cases where they're like, you know, super sophisticated and you guys have obviously introduced some like basic editing functionality. How do you like think about this product service area over time?

I think one thing that's cool about this product is that at its best it's very democratizing in terms of who can create and who can kind of level up their game to become a traditionally procreator, right? If you see kind of any of these god tier gens from like, you know, people who have really mastered Sora, uh you can just remix it and have basically direct access, right, to all the ingredients that went into that. And you can sort of like learn the ropes gradually, right? About how you actually kind of master like prompting Sora, how you master like designing your cameo, etc. Um, I think an important part of this is continuing to let people who are at the frontier of creativity, these proconsumers, continue to be empowered with like even better tools, right? So they can keep pushing forward. And we are launching more features that are really aimed towards that group specifically. You know, storyboard was a big one. Yeah. Um, we're starting to introduce very basic kind of like editing features into the app like stitching which went out earlier this week. And so over time, you know, my hope is everyone kind of levels up, right? You know, we're going to really empower the top creators to do their thing, but just by virtue of seeing this stuff in the feed, right? And having all these amazing remixing and editing tools, everyone can kind of like gradually like turn into these people and it will really just like make the feed an amazing place to be. There's there's something beautiful about an an end state where more people can be creative and then also people it that can be a gateway drug for even going deeper. I think I had this experience growing up with like Garage Band. I think it's like a good a good analogy here where

um it was so accessible. It was crazy like you know at in the simplest you could just drag loops right. not even really sitting down and playing an instrument, but you're starting to get a sense of like, all right, what are the elements of creating? And you can actually create some interesting things, but then you can go deeper and you're like, oh, now now I'm actually curious enough to have the MIDI keyboard and like learn guitar and record and like it's a gateway drug, but you we got there by making the, you know, the barrier to entry much lower. And I'm curious, I feel like every app builder right now is is kind of thinking about like what do I build a scaffolding around these models versus like I just go to a beach for two years and like the models are going [snorts] to get way better. [laughter] How do you guys think about that as like a holistic team of like the the types of scaffolding you want to build around the shortcomings in the model versus ways it'll just get better and you know

Yeah. In general, I think a lot of OpenAI's magic comes from just having this incredibly ambitious AGI aligned research roadmap and sticking to it no matter what happens, no matter what the competitors release, no matter what the product pressure is. Um, and that's certainly our philosophy on Sora. You know, I think what's cool is as these models get more and more powerful, we kind of discover all these amazing capabilities that they have and that, you know, keeps like Roan and Thomas very busy. Um, so you know something like Cameo, something like remixing, right, is really uh not just uh a feature that we pioneer on research, but it's really a joint collaboration with the work the amazing guys on product are doing to figure out all the crazy things you can do with these models.

I do think it requires a very creative lens on what you do with it. Um, or a willingness to fail, you know, like let's try the AI, you know, completely AI feed with no cameos. It's actually not that cool. But I I can imagine a lot of things in gaming and other other surfaces that even today's LMS and video models could support in a very interesting way. And so I think it's just a little bit of thinking outside the box and not trying to mirror the exact road map of Open AI in any way. Uh but being like okay like what's a spin on this that actually makes sense. I'm always consumer build not talking about enterprise of which there are a thousand for these models of course but like the consumer things is just like what's creative? What's new? Let's embrace that. There's way more you can do with this. There's control. There's all kinds of exciting things. Yeah.

Yeah. This is a big part of the reason we started launching a sore API as well with this generation of models. There is so much stuff as Thomas is alluding to that you can build with these things. And we're a very small team on Sora.

Shockingly small. I mean,

shockingly um you were under 20 when you released it and now we're like 50 or so.

We're at like So there's like roughly nine or 10 researchers on Sora. I think product is at what like

under 20ish. Yeah. Um and then we have a systems team which is about like 13 or so folks. So it's it's like 40ish in total. It's pretty small. And so you know all of uh you know these creators who reach out or uh you know all these uh very entrepreneurial folks who want to build these new applications they now have the ability to do that with the sore API.

Have you seen anything cool already on the sore API or like what gets you guys excited about uh what people can do?

Mattel has done some cool stuff. So they've been prototyping

like Barbie Mattel.

Yeah they've been prototyping new toys with Sora which is just super cool to see. Um, you know, this is I think you know over time the applications will just get increasingly sophisticated. It's been I think 3 weeks since we launched the API but yeah already people are doing some pretty wacky stuff. I've seen some like CAD CAD to visualization pipelines where people are taking like a CAD file trans translating that to some caption that Soro could understand so they could visualize their parts which I'm like that doesn't seem very precise and they're like explained why this is actually critical and something missing and they're I'm like whoa that's [laughter] yeah amazing

you know maybe take a step back and talking about the video model side like

but let me just kind of contextualize for our listeners like how has the AI video space progressed over the past few years maybe some like the different milestones that matter to So I'd say for a while uh basically there was no progress in video and a lot of it was really on the image generation side. So one of the most important early papers in image generation um was Dolly 1 which came out of uh OpenAI several years back. Um that was like early work that Aditia did and that was really the first time we kind of saw this you know breakthrough general purpose capability that we were starting to see in the LMS come to the visual generation side. Before that point, you know, there were models that were very niche in terms of modeling very specific narrow distributions like human faces. Um, but there was really not a moment where it was clear these things could model kind of all the pixels that exist on the internet. Uh, I think from that point on it became clear where things were headed. It took a few years to really get image generation under full control. So, you know, Dolly 2, Dolly 3. Um, we started working on Sora in early 2023. It was co-developed at the time that Dolly 3 was being worked on. and we kind of had all of the necessary ingredients like make a big breakthrough at that point in time. We started to started to understand how scaling works. Uh diffusion models were starting to become much more principled uh both in terms of the actual formulation of diffusion as well as the architecture of them. And Sora 1 was really like the GBT1 moment for video. Um that was the first time you could even think about doing kind of high-res consistent generation uh for anything more than 1 second. Um we could do 60 seconds barely uh with that model. And since then, you know, we've really just been trying to push the frontier of intelligence and usability. So we for a long time we were actually trying to figure out what number of GPT is Sora 2. Yeah, I think it's we're very clear now it's GPT 3.5 in many ways both in terms of the breakthrough in capabilities. This is the only model that can you know do an Olympics gymnastic routine and like not have things just like go crazy,

right? I love that that's one of the eval uh everyone just like went off on Twitter with like all these Olympics proms. He saw limbs flying everywhere. Yeah, we got we got roasted. Um, so clearly, you know, it it's justified by the model intelligence boost, but also in terms of the usability side, right? GPT 3.5 ushered in chat GPT. That was the first time that these models became really valuable just a huge swath of society. And that's what we're seeing with Sora as well.

Yeah. I I guess like uh did that jump from GBT1 to 3.5 like surprise you? Uh or did it feel was it like people in the know were like yeah, if we keep scaling the way we're scaling this this is going to happen around. I mean we knew we were on to something you know when you again see uh just this like rapid improvement in how the model understands physics right it's very visceral to see uh you know these gymnastics routines start working or for you know a glass gets dropped it actually shatters right um you know you you know you're on to something we just see the videos right like seeing really is believing um so we we knew it was not going to be as slow of a ramp uh as it was for the LMS I mean it's funny to call that slow but you know we're really accelerating timelines here um uh so we knew knew that that it would eclipse the pace of progress in LM. We weren't sure by how much. Uh and so I you know I think we weren't sure will we end up okay like three 3.5 we ended up kind of being on the upper end of expectations.

Yeah. And what in your mind is like the jump from 3.5 to four. What are the the kind of like unsolved problems or things that you're you know that you think need to get better in these models?

I think the next like breakthrough capability in video is really going to be simulation of processes uh that need to last for hours or even longer on end. So we're really excited about how Sora will eventually be used for knowledge work or even for biology, physics research, etc. And if you think about what it takes

To simulate a wet lab, right, we really need a lot of improvements on the modeling side. You need to be able to run these models for days, weeks, years on end. And that requires solving a lot of very fundamental problems with generative modeling. I feel like this is one of the big questions.

I mean, maybe just take robotics as an example. I feel like there's this big question of like how much simulation data alone gets you. And like I think for a while it felt like maybe it could get you a bunch around locomotion but not like manipulation or some of the more sophisticated tasks. The video models have gotten a lot better since some of that early thinking. What what's your latest you know bet on you know how far just using these video models gets us in some of these domains where it felt like we'd have to build up massive real world data collection operations.

I think they're going to be instrumental for making progress here, right? And to your point, one of the really core issues with robotics, right? It's just hard to get a large data set of trajectories which is useful for pre-training and video models clearly right like deeply understand things like local motion like dexterity related tasks and so we're very bullish on the ability to repurpose models like Sora for these things. You know, I think I'm very bitter lesson as you know, many people are at an AAI. Uh, and I really believe that some of the older kind of lines of research which were built on very early video models but didn't necessarily bear fruit, uh, will start to be successful just with these increases in base model intelligence.

>> Are you going to basically need one of these like state-of-the-art video models to even do work in like bio, robotics, material sciences?

>> I think it will increasingly be the case. Yeah. So I guess like these these next frontiers it seems like for you are really around like longer running videos and the ability to kind of uh keep like true to physics in the world.

>> Yeah, it it's really in pursuit of this world simulation goal, right? Like how do we make it so Sora is not just kind of like a video generation system but really deeply understands, you know, every bit of reality and is useful for tasks that go well beyond just kind of entertainment. I think entertainment is an awesome thing um that people get tons of utility out of with these models as we're seeing with the Sor app right now. But I really think this is just phase one for video and this is this is kind of the simplest thing we can build right now with the current capability of models and their usefulness will just skyrocket over the coming years.

>> Yeah. Is there an eval that you have in mind that you're like oh video models can do this then that'll be

>> maybe GDP val eventually you know [laughter]

>> it's a good question. uh yeah you know I think the first scientific breakthrough that comes as a result of uh simulating some phenomenon video models will be an insane milestone right I mean that will just really open the floodgates and we've been debating on the team like what will that thing be uh high error of ours of course I think something related to classical physics is very likely video is just a good modality for understanding you know physical phenomenon that are well represented with like observational data right uh things like turbulence etc so I think that will be an insane moment. I guess it's not necessarily an eval, but certainly like a milestone that will really mark kind of like the beginning of a new era.

>> Yeah. And because I'm a shameless podcaster, I have to ask you, you have any a prediction on when that might be?

>> Oh, man. I'm really bullish on timelines. I I would be shocked if if we are not making breakthroughs like this by early 2028.

>> Yeah. Yeah.

>> And is it like uh I mean is some of the LM world where it feels like we have some scaling laws that we just need to follow or are there like other you know I guess on the the mix in the LLM world of like data algorithmic breakthroughs and and just more GPUs like is it a similar way to think about things in the video world or or how should people adjust their mental models?

>> I mean I think there's lots of axis along which you can make progress certainly scaling is a big one. You know, I think when you begin to think about how do we do these kinds of like yearslong simulation rollouts that maybe requires some new breakthroughs and you cannot just kind of port the existing techniques directly over that. Maybe you can and maybe it'll be successful. But, uh, when we really, you know, get to the point where you need to build this alternate reality, it remembers, you know, every, you know, character on your piece of paper right there, kind of the details of of the fabric of your jacket. Um, that seems like it might need something kind of new. Um uh so lots of uh green pastures to explore.

>> How do you do evals on on video models now? Like what what how do you know when they're getting better and uh uh you know how do you guys do on the product side? Do you have any go-to go-to things somewise? [laughter] But but we actually I mean this is something an area where we've matured a lot I think especially from Sora 1 to Sora 2 just learning how you know how high leverage good eval with with real use cases and before we launched having conviction in the in the use cases lets us build you know like basic product evals right like hey we're changing the model or some part of the stack in XYZ let's run you know the top Sora one prompt um through Sora 2 and get a sense of you know the diff and now that we actually have like production use cases just like like Cameo is a great one like anytime we make a change we want to understand how does it impact like one of these core use cases.

>> Yeah.

>> I'm curious with like with video models what obviously you could imagine if you like you know averaged out the preferences of everyone you may end up with like a personality list model that like is is perfect to no one. I wonder over time do you think there end up being like a bunch of different models with different aesthetics and different like use cases or does it kind of converge on a single model that you can kind of steer to your like way of liking things?

I think there's actually analogy here to recommener systems. It fits pretty nicely, which is like when people think of recommener systems, not that everyone necessarily thinks through their feed very often, but like um and let's just say it tasks you with like designing a feed, you your first thing you might think of is just like, oh, let me sort by popularity or like sort by something like that. And actually you find that that is exactly you're the problem this regression to the mean kind of thing where it's actually interesting to nobody because it's just just a trend that is only globally popular which is not suiting your own personal preference. And so um what the magic is you introduce personalization in different forms and that can be there's lots of different ways of doing this but the obvious one is look at your history and see what you're doing going to do next. Um and that really does change the feeling and the results of relevance all that sort of stuff in your feed. Um I can't imagine that we wouldn't have a similar phenomenon um basically everywhere. I mean we're seeing a lot in chatbt around personalization being a very important thing the way the model talks to you. I wouldn't necessarily I mean I'm not going to speak to video models directly but in recommener systems you learn very quickly that you don't want bespoke models for everybody right um it's an infrastructural nightmare. It's also like not really in the pursuit of AGI in many ways. Um you can kind of do it but it doesn't really scale. It also gets very hard to reason about. Um, and much better is to leverage the wisdom of the crowds where you find analogies between people. You do collaborative filtering and so it helps actually the scale benefits you in this beautiful way where it's like I know that this person has like videos that are similar in the past uh and I'm similar to that person and so now therefore I'm going to like very likely enjoy this video as well or be inspired to be creative from this video.

>> Yeah. I mean, I think models that perceive the world the way humans do, like if you measure intelligence in that way, diversity and understanding different stylistic behavior baked into the base model is is an important part of that intelligence. And I that was the craziest thing about Sora 2 to me, not the craziest thing, but one of the most amazing things. There's definitely mode collapse in I think almost all the other video models out there today. Even the ones that are great at like cinematic content and physics.

>> Zora feels like it's great at that and great at like a ring, you know, doorbell footage. [laughter] It's great at like an interview um like a podcast. that's great at all these different kind of things that feel very anime, very stylistically and different on sort of like just like I don't know the the range of the model is amazing and I I yeah I only anticipate that we we lean into that.

>> Yeah.

>> I mean you guys have obviously spoiled people on the LM side where it's just like the the default expectation is whatever you release is like 100 times cheaper in you know in in six 12 months like should we expect a similar thing on the video side?

>> 100%. And even if you rewind the clock to February 2024 when we showed off sore one to the world for the first time, that model cost him like $50 [laughter] in compute to sample like a 720p, you know, short video from. And if you look at like our API pricing for Sora 2, right, it's like cents on the dollar compared to that. So already we're seeing orders of magnitude decreases in cost with vastly higher model intelligence with this release and that trend will definitely continue.

>> You wanted to break big news on this podcast. So 30 minutes before we started, you tweeted that uh you were you were going to introduce some pricing uh which I guess entirely reasonable I think when you get 30 a day and then you have to start paying a little bit for uh for them. Um and I don't think the internet rioted from what at least I saw that seems like a natural first step in monetizing. I know you guys have talked before about like both, you know, monetization for the inference cost of running these models, but also monetization to figure out how to incentivize rights holders and all sorts of folks to get involved.

>> Yeah, I mean we really want to create an ecosystem where everyone is benefiting, right? Like we need to pay the GPU bills for Sora. uh we want you know even new creators who are coming up on this platform right they didn't necessarily have a following on Tik Tok or IG uh we want them to have a path to success and making money and then we want the rights holders who have you know this incredible library of characters who so many people will just love using to also benefit so when we were thinking about the initial ways to monetize you know we're going into this learning a lot every day we want to kind of take it slow and you know we want to make sure we're checking the boxes on the most important categories here which you know, giving a path to monetization to all of these different people within our ecosystem. And we really viewed credits as being kind of the primary way to do that um without necessarily overcommitting to a model that we don't know we can support long term or just like need to pull the plug on eventually. So, we're we don't think this will necessarily be the final way that we ultimately monetize Sora. We're really open-minded in this. Um, we're trying to be as transparent as possible and we kind of like make these decisions and just hear what everybody across the board has to say about it because we we want to land this in a spot, right, where, you know, certainly not OpenAI is like the only player like making money off of this, but everybody is benefiting. I think that's really important to the success of the platform long term, right? Like this should be somebody's full-time job. If they're a creator and they want to go viral on Sora, we should have a way to really make that possible for them.

Any other like types of pricing models you guys have ripped around or thought about?

nothing that's like super tangible in the short term. But again, going back to, you know, how you can just totally rethink the like branding and the ways you might want to bring some something that you sponsor um to the table now that you have generative video. Like, you know, right now if you're scrolling an Instagram feed, you get a new video for an advertisement. But you know what? If you're a creator and uh, you're okay with all the inanimate objects in your video being certain brands or something like that and you can auction off to certain brands. Like there's there's like some wacky ideas out there that I feel like are um yeah, we have like

>> I think other thing that is really interesting about this platform and different and I'm sure it'll grow over time. Um, like I'm use myself as an example. Like I made my cameo public of one of the first people, first early adopters that did because we were on the team. Uh, and I don't know, I think I have like 17,000 cameo appearances now. And like I if I sum the view count over those cameo appears I probably have never had more distribution in my entire life. [laughter]

>> And I I mean, I'm a nobody. Like >> my follower count is like creeping up to my one on Twitter. It's quite crazy actually.

>> That is crazy. [laughter] That is crazy. But and then you multiply that by the cameo count and the doing that like actually there's a crazy amount of reach for that one individual um that you don't it's like kind of an impossible level of reach on other platforms because uh they have to you know or you have to create it yourself necessarily but I actually like my like I like seeing what I do every day I change my cameo instructions I'm like oh I've got a cool shirt on now musical and I was like wonder what all the musical numbers [laughter] not just me looking at my own G like I'm also seeing what you know other people are doing so I think um it's it's a very interesting thing I don't I think we have a great analogy for this type of um

>> uh format. There's a lot of new medium things, but it it's just interesting to think that reach in a different way is like it's not just what you post necessarily.

>> And one thing I think that's cool about the product too is that it's it's like global and and I wonder if like you've seen any cool like different variations and ways people use it like across uh across geographies.

>> So we we just did uh the launch to some Southeast Asian countries yesterday night.

>> Literally last night hot off the press.

>> That was hot off the press. So that's that's out there. Um, we originally launched to uh US, Canada, and then opened it up to Korea and Japan. Uh, and there's definitely a huge different flavor of all kinds of creation. Comfortableness with cameos, comfortable with character cameos. Uh, it's been inspiring to see. Sometimes people complain about the cross-cultural content. I love it. Um, it's it's really everything just has a different flavor. The Japanese creators that we all love and follow.

>> There we go. Yeah, just like it it's like very aesthetic. It has a very different feel and and than anything I've seen from the US. So,

>> um yeah, I'm sure there's going to be more to more to come there as well. And um I like it when I can learn something from the Gens as well. One thing I saw uh was a somebody talking the Toronto accent and I was I'm from north of Toronto. I didn't know there was a Toronto accent. [laughter] As soon as I saw that, I was like that that is exactly how everybody talks. And I immediately opened Wikipedia and there's this huge article on the Toronto accent the dialect or whatever they call it and I was like I've just learned something [laughter] interesting. Um, which actually was my experience with Tik Tok too. Like I'm actually a big Tik Tok user and fan. Um, and I I so many things I've learned from that where somebody is like psychological things about myself or you know they're describing some behavior some attachment theory thing. I'm like wait a minute that's me. Like let me go look that up on Wikipedia. Uh so I think we'll see a lot more of that and I think crosscultural type of stuff is uh is also there. Also when you see a remix chain it's so fun to see all the different countries that are participating that always a different flavor. Totally the rap is a different you know slightly different

>> I need a translate button but it's like so close. Yeah

>> the latest feed I love. It's like my favorite um spotlight into what people are doing on the app. It's just like literally the stream of what's happening. One of the most to riff off of what Tom was talking about in terms of like learning things, just learning what what people want to what scenes they want to see themselves in and their friends in. Like it's it's like just so interesting the kind of things people are meing about and like I don't know like it's mundane to one person, but then if you if you're in that mode of like what do people care and laugh about? There's so many hilarious things people obsessed with. I was talking to a user today that's obsessed with cranes. She was like, "I just love visualizing myself on top of cranes." [laughter]

>> I mean, obviously you guys, I think, have been super thoughtful on the moderation side, too. And I don't know whether like I think both encouraging like a lot of fun content, but also I think have done a pretty good job of like navigating those waters. Like, was that an iterative process or do you just like nail it from step one? Like, you know, talk a bit about that.

>> Oh my god, we [snorts] long sleepless nights [laughter] um getting to the point where we are now and still, you know, still a lot of work. Shout out to amazing both people working on safety on the Sora team and like a ton of work has happened at OpenAI on on our moderation systems, moderation models. We have reasoning models that we've written about that are in our safety stacks which are amazing. Um, and so we really it's a great example of using our technology to make our products better. Um, and like we were talking about before, uh, we we are able to do a lot more with a smaller team now on this like moderation and safety front. But thank you for saying it's in a good place. not what people on Twitter tell us. But um

>> Twitter angry about something. I can't I can't [laughter]

>> there's you know the content moderation failures are are are understandably frustrating because you know we're trying to tow this line between user freedom but blocking all the bad stuff. And we're getting we're getting better every day. We're in new territory especially with cameos. We want to be really sensitive to um you feeling like you're in control and you know having you're comfortable on this network right. So with that comes uh a lot of thought in the guardrails. Yeah. Yeah.

You've kind of talked about this being a new social platform. I think you know Sam's spoken very publicly about like that being interesting surface area for open to explore. You talked about maybe this integrating with chatbt over time in some way. Like how do you see other like angles of things that are being worked on in OpenAI like one day you know co uh you know coalesing with us?

>> I mean Bill mentioned you know going beyond entertainment you can like chatt is is like sort of your super assistant. That's literally what we called it early days. um why why shouldn't it be able to respond with a with a really helpful video? Um I think that'll be really killer uh once we we get that kind of stuff in. Um but yeah, it's like I think the world's our oyster with ways that like all of our products could interact. I mean, we just released a browser. Um I don't know, you can imagine using a browser and having a little video assistant or something talking to you on the side, your agent help me book this flight. I don't know. There's crazy wacky ideas I've heard out there.

>> Yeah. And I think a lot of this stuff builds off each other like say the reasoning models which you know in the pursuit of AGI which very justified is uh you know I my first use case wouldn't be a moderation stack like wait a second that actually is a perfect use case and so a lot of this stuff builds on each other the ecosystem itself I do think like chatb does feel a little different a little bit sacred doesn't mean it can't change over time or like there's definitely things that we can do there but um you know it it is tends to be a more utility driven use case and you know mixing entertainment in a utility driven use case doesn't always work. Um, so I don't think it's just like jam it in there and it will just magically work. I think we went on to do that very thoughtfully. Um, but certainly as a as a utility use case, these models are

>> uh going to discover new modes of turbulence modeling. So there's clearly utility. Yeah, [laughter] I heard I heard 2028 the early not even not even by the second.

>> [laughter]

>> You should be able to ask Jad GBD about that. And you know, video as a modality is you see it on YouTube like half the videos on YouTube like how-to videos. That's something that we would want to see modeled. But in my case, you know, I here's my toilet fix figure out how I like adjust it and like fix this problem with it or whatever, whatever it is. Um so I I think that can be very naturally connected.

>> Yeah. Before we we move to our quick fire round, I I though I did want to ask you uh you know a shameless question which is maybe what's the simplest way to explain how these models learn physics?

>> It's a great question. Uh at a high level, right, these models are always doing prediction tasks. So in the case of for example diffusion models, you're getting a video which we've artificially added a lot of noise to and the goal of the neural network is to predict the underlying signal that we kind of obscured with all of that noise. LM's also a prediction test, right? you're predicting the next token condition on all the prior tokens. Uh we think actually both LM and Sora learn world models. They learn different flavors of world models. But fundamentally, right, if you're doing this prediction task, whether it's predicting signal or whether it's predicting the next token, you need to develop some understanding of kind of the dynamics of the world, right? So in the case of LMS, if I'm generating like a rhyme or something or maybe a poem, then it's very useful for me to have some kind of internal knowledge, right, about the structure of poems, about the linguistic elements of them that I just pick up from data, right? Because if I haven't gro these things, then it's very hard for me to predict the next token. And the analogy holds for video as well, right? So, if I'm a video model and the prompt is a guy playing basketball, I sure better have learned, right, how basketballs dribble, how light refracts through a scene and all these just little components of physics in order to paint this broader picture of like the entire scene. And so, it's basically an emergent property from lots of compute and lots of data, which is like a very, you know, benile thing to say these days. Everyone gets it. Um, better lesson works. Uh but fundamentally right if you don't have these kind of internal models of how the world functions then you will always be a worse model and thus have a higher loss than one that does and so that is kind of how we think about the optimization pressure uh which makes physics emergent from large scale video pre-training.

>> In the LLM world there's you know we went from just like the entire internet to now like there's you know incredibly valuable data that is like PhDs you know doing different things. Is there like an equivalent in the video world of like oh that is like a really complex you know piece of physics that like is great for the model to see or like

>> Right. Right. I mean, if you take any like video researcher and show them a bunch of videos and ask them which is their favorite, chances are, you know, if there's a gymnastics routine going on, they're going to be like, "That's the one gymnastics.

>> That's the one." But I think it's a really interesting question. And, you know, it's it's somewhat hard to answer because the way intelligence manifests in videos, I feel, is very different than the way it manifests in language. You could for example imagine there is a lecture video right where someone is giving uh you know a course on calculus and that clearly has an intelligence which is of the flavor of text right where it's teaching you about some like deep concept about the world mathematics physics what have you on the other hand though when you think about these gymnastics routines right there's no intellectual intelligence there beyond kind of the planning involved of doing like a gymnastics routine but there is so much interesting detail about the the small collisions that happen as part of that routine again about all of the people in the background shuffling amongst themselves, like building a model that can actually simulate not just the gymnast, right, but every single person in the broader scene. And so, you know, we're still trying to figure out what kind of data really makes for an amazing video model, but there's a lot to it um that's very multimodal in a way, right? Multimodal in the sense of kind of all these different pockets of intelligence which exist in video, but not necessarily in other modalities like text.

>> Um well, I always like to end our interviews with a with a standard quickfire round where I stuff some overly broad questions into the very end. Uh, and so maybe to start, um, I'm curious for each of you, one thing you've changed your mind on in the AI world in the last year.

>> I think my timelines for some things have accelerated and for other things have delayed. I think like we, this is a common one that people say, but I just think we we really overestimate estimate consumers and adoption and how people learn this technology. And we could be way ahead of like the general world in terms of the actual science, but way behind in terms of like the way we inter build a product interface that's acceptable uh uh accessible and how we teach the world about it. And then when you get into like actual enterprise use cases, there's like regulatory sort of uh just like yeah things you have to sort of battle through to get people to actually adopt this technology. So all of that, you know, working as like a a product person, even on the consumer product, just seeing like the crazy amount of sort of work that happens beyond behind the scenes to make things possible. Um, yeah, there's just a lot to do for people to actually adopt these things.

The aspect of uh AI I've updated most on is if you asked me a year ago, would I care if I watch, you know, some blockbuster movie uh versus say some movie generated by a model that's, you know, like Sora 3 caliber or something. I would say the Sora 3 movie would be just as good as the blockbuster one. I don't need any information about it. Like assuming it's, you know, grocked cinematic content and like interesting plots, like that's sufficient. Interestingly though, I feel like the more I've interacted with these models, I actually do feel a hollowess when there's not clear creative intent behind it. Um, which is something I would not have expected to have felt a priori in part because, you know, like, you know, I enjoy making the models. So, you know, it feels like weird for me to have that opinion. But I really have been struck by how much more interesting and compelling um, generations feel, even with Cameo, of course, right? Seeing the people you know and them feels amazing. seeing your pet Rocket uh my dog who's now as of yesterday available in Zora.

>> Um uh but even knowing someone uh I think Thomas was mentioning this earlier, even knowing someone just had the creative intent, right, to reject a sample, right? Or like really iterate on a prompt and the plot that they're actually using in the underlying generation has some kind of deeper connection to them emotionally, right? You actually do kind of feel that in a way that is very surprising to me. Um and I'm not sure exactly what triggered this change in me. Like I again I would have been like very convinced by these completely AI generated feeds that our bot made or something as long as the bot was good. But I I really now like do not actually think even I would find that compelling which is quite interesting to me. I'm not sure exactly actually why this is the case. But there's some you know there's still some like morsel of humanity which is like very important to communicate through these genens for for content to actually feel meaningful in some way. Um so I've really updated on that.

If you guys all, you know, left Open AI tomorrow and had to build things on top of the Sora API, uh, what would you go build?

>> Game

>> like a game where you put yourself in a video game where you

>> Yeah, I'm sorry. I could have provided more details than that, but like [laughter] are you leaving tomorrow?

>> Here's my very well thought out.

>> Why do you ask? [laughter]

>> Yeah, no, I I think there's um I am just I grew up in gaming. I learned to program in gaming. Uh, I can't imagine how much crazy opportunity there is in that space. Um, and the idea that a lot of the things that are very difficult at least for me when I was building games uh was like oh I have to go and uh source art or build the art myself and that was a labor labor intensive process. Learn the tools to do all that sort of stuff. Um, I don't think this is exactly what needs to manifest, but like what kind of gaming models are enabled by this generative technology in the same way that like we're talking about Cameo is that something uniquely enabled here is like what kind of things are out there. Um, so I'm excited about that and seeing what people can come up with in that space.

>> Totally. My answer is like I'll get a similar answer from a different angle which is just that I think I think the most I mean it probably comes through in the product we've built and how we're talking today but the most interesting thing about this technology is how it will transform and create new mediums rather than just you know uh like an AI film. I think it's like the least interesting uh concept to me. There's people doing cool things with film but like completely new things are most interesting. Um, and so I think interactive, which is like maybe gaming from Thomas' point of view, from my point of view, maybe it's interactive storytelling of some sort. Um, and no one's really nailed this, like you know, there's choose your own adventure books. I think those were pretty popular. Yeah.

>> And like some movies like Ber Snatch on Netflix where you kind of like, but I think

>> pushing that idea like a lot farther um of some like, you know, collaborative piece of piece of creative lore or something like that.

Why don't you think there have been more consumer product like fun consumer products built on top of these models? Is it just that like they weren't good enough until now or?

>> I think it's really one it's really freaking hard to build a consumer product. A good consumer product is like one of the hardest things

>> combined with the technology being so new. Um, and you know it's unclear what what actually sticks. I think I think that's actually still quite a hard

>> Yeah.

>> combination of things to nail. Yeah.

>> I mean you've just done it twice at the top so I guess [laughter]

>> still still difficult. Um, what about something you'd build on top of the API?

>> Oh, man.

>> I know we can't take you away from the models, but if we had to.

>> Yeah, I would just go train new models, you know. Um, I don't know. Something probably more on the science side or robotics facing. Um, I think there's all just a lot of cool new work to be done there.

>> Well, I want to leave the last word to you guys. Uh, I I usually say where can folks go to learn more, but I I'd be genuinely shocked if any of our listeners didn't use Sora. [laughter] Uh, so maybe I don't know if there's specific research you want to point people to or like cool new product features. Uh, the mic is yours. Anywhere you want to point our listeners. Uh, please go ahead.

>> We just launched character cameos yesterday. So, you know, if you use Sora, you know, you can upload yourself um and create videos with your friends. Now, you can upload anything, you know, pets, anatoms, you can create, you know, your own IP from Sora and create a character out of that. Um, I've got a Pickle that's going viral, Pickleton. So, [laughter] go cameo Pickleton. There you go. It's my piece.

>> Um, I have a wild wild one. Okay. [laughter] Just trying to think of something.

>> It's an open mic, please.

>> Open mic is uh

>> in Disneyland Hong Kong, there [laughter] is

>> Good start. Good start.

>> There is a It's the Haunted Mansion, but it's like I think it's called Magic Mansion. I don't know. They have a slight riff on it. And uh there's this lamp that it's a dark ride. So you just go through it and there's this lamp that pour like the whole concept is this monkey has this lamp and when it pours out the objects come to life and so it pours out and then there'll be like an armor on the wall and the arm will stop moving and talking to you and uh

>> we actually kind of just built that which [laughter] like you just record an object and then you say come make it alive

>> and so uh I think check out that ride is very cool but also back to character cameos it's like actually [laughter] it has that magical proper proper Yeah.

>> Yeah. I'm sold.

>> Um, yeah, we've temporarily made it so you don't need an invite code to get in. Come in with your friends. Sora is way more fun if you're in there with a group. And, uh, yeah, check it out. Give us feedback.

>> Yeah. Well, guys, thank you so much. This is a ton of fun. Um, I I want to make sure I get you guys back to actually building Sora. Uh, cuz the world will [laughter] will far more benefit from that. But, uh, seriously, this was a ton of fun.

>> Thank you so much for having us. See you in Disneyland, Hong Kong.

>> Yeah, [laughter] seriously.

>> [music] [music]