Transcription
With Genie 3, Google didn't just launch a prototype. It planted the seeds of entirely new realities nested within our own. A place where worlds generate themselves, machines learn inside them, and the lines between simulation and reality start to blur.
Now, I spent years working on real-world models at Google. So, I'm going to give you an insider perspective on what tech like Genie 3 and Cosmos means for the AI and spatial computing industry. Combine this with enterprise-grade simulation technology that's been evolving for years. Generative world models have real implications for 3D today. So let's get into it.
All right. So, quick primer on Genie. Essentially, there's a model where you provide a text prompt, a starting image, or even a video like Rui did over here, and you can suddenly take control just like you would in a video game and start playing this experience with the model generating the world you should be seeing in real time at 24 frames per second. The results are fantastic. There's also open-source variants of this tech such as the very appropriately named Matrix Game 2.0, uh, which is basically trained on a bunch of video game and GTA footage. But what makes Genie different is, of course, that this is trained on the entirety of YouTube, right? This is the motherlode of visual data sets. It almost makes it feel like playing a dream, right? Like this controlled hallucination of reality, and it makes you wonder the nature of our own existence.
But not to go too deep down the rabbit hole, take a look at this example of Google researchers playing with Genie 3 in a conference room where suddenly that video is put inside a Genie 3 world, and now you have a simulation within a simulation.
Now, a really cool byproduct of this model that's seen the visual complexity of the world and can generate worlds on demand is that it can basically do one-shot or single-image 3D reconstruction. You can take this old famous painting of Socrates and turn it into this explorable 3D world. And just consider for a second that all of this consistency is an emergent property of the model just getting bigger rather than the developers giving it some explicit underlying representation like a textured 3D mesh or a neural radiance field. To put this to the test, I basically took that Genie 3 world, painted out the UI, did an upscale in Topaz, and trained a Gaussian splat and a texture 3D mesh. And suddenly you can step inside this painting of Socrates from 1787. This is genuinely better than any image-to-3D model I've seen. Here's the texture 3D mesh version, which is also really, really good. You can drop this into any app that you love. Blender, Unreal Engine, the list goes on. And if you're curious, I've left links to both of these models on my Twitter post so you can check it out over there. I'll also leave it in the description below. I tried this same result with the famous Nighthawks painting, and the results are equally stellar.
But Google is far from the only one. Nvidia's got Cosmos. That's their version of a world foundation model, really focused on physical AI, which we'll explore in a second. Uh, Runway has been talking about general world models since late 2023. And the idea at least goes back to 2017. So, Crystalall says like, "If you want to understand why suddenly everyone's investing in video generation world models, re-watch this video from 2023." This is, of course, on the heels of the G3 launch. So, of course, the Genie 3 co-author dropped, uh, oh, and our ICML paper from 2021 on augmented world models. And of course, Crystal Ball clapback saying it's like, oh, and this 2018 paper, World Models by Schmidt Huber. Schmid Hub, one of the godfathers of AI and a, a rather controversial figure if you want to look into him. And of course, Elon, not one to be left out of the fun, says, "We will have real-time in 2 months." So, I guess Grock's going to have real-time video generation in a short while, too.
Now, if we even go back to the launch of Sora, they also talked about video generation as world simulators. Sort of this idea that if you want to create an AI that understands the physical world, a good way to test its understanding of the physical world is to have it generate the real world and have those be plausible and ideally factual reconstructions or renditions of, uh, what you would expect in reality. And really, if you take a step back, all of these companies are essentially trying to recreate the Holodeck from Star Trek, right? Like, sort of this idea of being able to generate worlds and experiences on demand that you can step inside and interact with just like you would with a video game, except this video game is indistinguishable from reality.
Now, look, there are different ways of going about creating these world models, right? With Genie 3, Google is basically taking all of YouTube and training this model on it, and suddenly it can predict what the next frame should be given a set of input events. That's kind of all it needs to do. And it turns out the bitter lesson is bittersweet for a reason because as you keep scaling up the data sets and the compute that you throw at it, it turns out you just get consistency as an emergent property. And it may well be that you keep scaling up this type of an approach and suddenly you would have something that would be like an Unreal Engine 6 with all of the physics that you've come to expect from it.
Now, Nvidia and other players in the space are taking more of a hybrid approach. This is how they're thinking about this problem of creating a simulated digital twin of reality. This implies this is a twin of something that exists in the physical world. So, we have a factual rendition of it. What do we mean when we're talking about robot simulations? And why do these world models need to play into it? Well, think about a self-driving car, for example. It's got some sensors, maybe has LiDAR, some RGB cameras. It can perceive the world, and that that produces these sensor tokens. And maybe you have some directions you're giving your car. Car is like, "Yo, take me home and take the highway." That produces these text tokens that goes into this policy model that basically produces action tokens, meaning like steering commands or maybe waypoints on the road in 3D space that it needs to take next to and essentially get you home. Um, so what happens is you want in a robotic simulation for the system to be able to take that action token, right, and then predict the next world state. That's where the simulation is so important because that prediction needs to be as close to the real world as possible. Like the margin of error is non-existent. Otherwise, you're going to have a dangerous robot in the real world. And we don't want that, right?
Nvidia is taking a hybrid approach where they're taking the best of classical 3D approaches and blending them with generative AI. Let me explain what that means. So, for example, you may have this amazing radiance field of a racetrack, right? Like that's super cool. This is very close to what the actual camera saw, but you do need the mesh reconstruction, right? So your physics engine knows how exactly to react and respond to the surfaces that exist in the world. So then when you take all of those pieces together of a classical approach, of a generative approach, you get something like this where you can basically reproduce the input and create infinite variations. So, for example, if you want to create a scenario where an animal walks in front of a car, boom, you've got a really good clip, and now you augment it with different types of animals. Maybe you're walking through San Francisco and you want to see it up in flames, which I guess happens on the regular these days. So, I don't know, maybe you'd be able to capture that in real life. But you can create all sorts of other variations, seasonality, weather conditions, the list goes on.
And this gets powerful, right? Because suddenly you can start using these world models as sort of like a jungle gym for robots, a generative training facility. So if you've got, you know, if you have a robot that's collected data in one environment, suddenly you can have that one skill replicated across all sorts of new environments. You could also perhaps have it do many skills in the same environment, and then of course many new skills in many new environments, right? And so you can imagine how a very small amount of really good real-world data can be augmented with synthetic training data to create a much larger corpus for your robot to train on. So who knew your robot would be docking every single night and dreaming just like we do, thinking about how it can do a better job the next day?
Robotics is very clearly a data problem. The reason we've had such an easy time with text and code generation, things like that, is because we've got the motherlode of data on the internet. We've spent decades essentially creating a digital twin of all human knowledge and creativity. And that's what everyone is training off the public internet right now. On the video side, you've got some advantages if you own something like YouTube, or maybe you're like ByteDance and you own TikTok, etc. But for this type of first-person view video that you need for like a robot, let's say doing a bunch of tasks, a humanoid robot specifically, it's super challenging. And same thing with driving, right? Despite all the dashcam footage, despite all the street view footage, you do not have enough observations of the physical world.
Now, there are a bunch of companies that are doing this type of synthetic training data generation for, you know, autonomous driving startups. In fact, Scale AI, which was recently purchased by Meta, got their start doing this type of labeling that you see over here. But these companies take all of that data and kind of create the final output that you'd need to train on for all the scenarios that you perhaps care about that you haven't captured in the real world. And as you can see, they're doing this with a mix of, you know, kind of photogrammetric and neural reconstruction approaches combined with traditional game engines and simulation combined with generative AI perhaps to give it that final pass to make it look more like that fisheye sensor in a rainy environment, you know, when you'd see it on a car itself.
So, you can think of something like Bifrost, which is another company doing this stuff, essentially as this like hybrid game engine where you can do all of these things, right? Like maybe there's a real-world reference that you're trying to replicate an interesting scenario on, and you essentially get this like Unreal Engine cloud stream builder environment to kind of build it all out. Let's say you're training an autonomous amphibian vessel. Boom. Suddenly you've got that type of data. And then somebody comes to you and says, "I need different weather and lighting conditions." And well, boom, you just go about doing that. And the beauty of doing this synthetically, especially with explicit 3D game engines or 3D content creation tools, is the fact that everything is already annotated, right? Like all your annotations are basically given to you for free, which you can then pass along to some ML model. So it's really good at detecting, let's say, if the seatbelt is on or off. Is the driver paying attention? Which direction were they looking in? And you could transfer that head pose data, not just bounding boxes, uh, to your model. And so that's sort of the beauty of generative AI. This love this guy taking a yawn. Um, you know, so for the uninitiated, some of these cars have this attention awareness feature where, you know, if you're not paying attention, uh, to the road, looking at your phone, etc., it'll start nudging you essentially.
Now, this doesn't just imply to robots in the traditional sense, like a car or a humanoid robot. It also applies to facilities or buildings or even entire cities. So this is what Nvidia means when they're talking about bringing physical AI to cities and industrial infrastructure. And if you think about it, this is exactly the kind of infrastructure where you need systems to be able to navigate, but you also already have a bunch of sensors in these spaces and places. Traffic is a really good example, right? So this is a company called Data from Sky. And basically what their tech lets you do is if you've got, you know, CCTV cameras or you're doing a drone survey, it can track every single car across camera view. So it'll detect a car, give it a stable ID, an identifier, and make sure that ID is persistent across every single camera view that you get. And this is a really good way of, for example, surveying how dense is traffic at different times of day. There's one way to do that, which is how Google does it with the location data on your phone, but that's so coarse-grained and it needs to be aggregated over very many users to be privacy-preserving. In this case, you can get a far more granular view of what's going on in an actual moment.
So, of course, extending our analogy, you can take something like Blender and the city generator plugin for it and create a synthetic rendition of that exact intersection and layer on those actual tracks. So, you've got very close to real-world traffic patterns inside your simulation in which you could, let's say, ask a car to drive around and do things. So, this gets very powerful, right? Suddenly both static and dynamic objects from the real world can be imported into the simulation where you create all sorts of permutations and combinations.
So to give you some examples, let's say I've got this real-world capture of the Lodi Garden in New Delhi. Suddenly I can imagine it in different lighting conditions using tools like Runway's Gen-2. Right here is me using Runway Gen-2 and Luma's video model for more of a media and entertainment use case. But you can imagine for like public safety if you want to train drones to be able to fly around after let's say a disaster has just happened, well, you want to train it on a destructed version of the city.
Now, if you saw my Google Earth documentary, then you know Google has amassed all this historical imagery. But think about it, if you can look back in time, you can also use AI to predict forwards in time too. And Google actually took a huge step towards this, like bringing together something like a ChatGPT for Earth. So think about it, right? Like you've got all of this optical data, this radar data, LiDAR data, climate data about every single part of the world. What if we just break up the world into these 10x10 meter squares and just create embeddings for them? Essentially distill down and create these compact summaries of all that data that you can literally ask questions to just like ChatGPT. We're going to have this on every possible scale that you can imagine, from the world to your country to your city to your home, where you can go and ask questions to your home.
Now, of course, not only can we train robots in these simulations, but we'll step inside them too with VR headsets. So, let's talk about the implications for media and entertainment, both immersive content and also just general flat content we distribute freely on the internet. Now, going back to the Holodeck idea, the fact that with some of these fully generative interactive world models, you're already embodied. You can step inside of puddles and recreate the physics that you can see. I think this is the killer use case for AR/VR, right? Like Google's got their XR headsets coming out. Of course, they're thinking about how generative media plays into that. And just think about it, since we've got an interactive world, creating that 65mm interpupillary distance to create stereoscopic depth is so possible. And you can do it without a game engine. Even if it's you just creating, let's say, a fisheye rendition of the world that you see, then another model for upscaling it, another model for doing 2D-to-3D conversion, like to create an even faster, lower-latency pipeline to go from like 24 fps to like 100 fps. I could see that being more of an engineering problem than a research one at this point. So, in my opinion, the Holodeck is closer than you think. It's certainly closer than I thought.
Now, of course, we also talked about how you can use these type of models to do 3D reconstructions of the world and extract 3D models or 3D radiance fields, right? So, you can extract 3D models out of these and use them in your traditional virtual production workflows, right? Like if you want to frame that shot of the talent walking by and record that with your phone camera just to see exactly what it'll look like, record those takes and then render them properly on the computer. You can totally do that. In fact, tools like Jetson even let you bring in that Gaussian splat and previsualize and use it for previs on the phone itself, which is really cool.
Now, I've talked about this in the past, but I'm really, really bullish on this hybrid approach, right? Clearly, it's the thing that works right now as we wait for fully neural or fully generative approaches to reach that level of simulation accuracy that you can get, um, you know, with explicit 3D approaches. My vision there is essentially you've got this 3D environment that you flesh out. You define, hey, these are the sets that I care about. These are the actors that I care about. These are the props. This is the time of day. This is the look that I'm going for. And then you essentially go into director mode, sort of like you can today with something like Convey Sim, which lets you create these sort of like AI agents that are inside Unreal Engine. And they give you a very lightweight GUI, by the way, to go create that without even needing to touch Unreal yourself. And then you can go and essentially talk to these characters and actually command them to do things too. Suddenly you're like the director on set, feeling like James Cameron in your freaking basement commanding your roster of talent to get exactly the shots that you want. So definitely check out Convey Sim as well if you want to play around with these technologies and think you could take something like this and run it through Runway Gen-2 or your favorite video-to-video model of choice like Luma.
In fact, if you want to see a workflow not so much optimized for talking characters, but everything else all in one UI, you got to check out this really cool tool called Intangible. And essentially, this is exactly what it does, right? And essentially, what you can do in this tool is use natural language to describe the scene that you want to create. It'll populate the world with actual explicit 3D assets that you can move around. You frame the cameras just the right way, and then you can use something like Flux and then V3 all in one interface to take it all the way. So having this combination of the editability of a 3D environment all in a browser along with the creativity of generative models feels like a really, really cool hybrid pipeline that is super underrated.
Now to wrap things up, let me talk to you about the future of 3D. You know, every year when Siggraph happens, it really puts it into focus for me just how much of an unsolved problem computer graphics and simulation really is. And this example might tell you why. Think about it. Let's say you wanted to create a photorealistic rendition of a bowl of strawberries down to the microstructures, the multi-layered materials, the freaking fur that you see on these strawberries. You could probably do it in a modern 3D tool. Absolutely could do it. You certainly do it with an AI generator. But then consider that that bowl of cherries actually sits on a breakfast table that itself sits in this apartment building that sits in this block inside of a massive city. To have simulations of the real world that scale seamlessly from strawberry scale down to city scale is a very open-ended problem. As Aaron Lefon said, who's the VP of Research at NVIDIA, that we're going to need completely new tools so that artists can even conceptualize, create, and iterate like order of magnitudes faster than you can today to make sense of this exact scale to have a simulation of the world inside your computer. And to be honest, we don't even know the 3D representations we'll need to get there. Is it going to be, you know, radiance field stuff like Gaussian splatting? We've got triangle splatting now. Is it going to be like this hybrid approach of bounding boxes plus fully generative? Is it going to be completely generative end-to-end? The bitter lesson reigns supreme. We just end up creating enough synthetic training data from all these simulations so that we can just get it all for free from one black box neural network. Maybe.
Not to mention, we're going to need entirely new ways of generating 3D content. Like the metaverse is a rather empty place without interesting 3D content, and we're just starting to scratch the surface of what we can do with those type of approaches. So, I don't think anyone knows which approach is going to reign supreme and win in the end. Is it going to be a fully neural approach like G3, or is it going to be some sort of a hybrid approach like we saw with Intangible, or what Nvidia is doing with Omniverse Plus Cosmos? But what I do know is that we are just scratching the surface of what we need to create both the ultimate digital twin of reality and the real-world Holodeck. And damn, it really makes you appreciate whoever wrote the rendering stack for reality.
That's it for this video. Bolav signing off, and I'll see y'all in the next one.