📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

ChatGPT Images 2.0 is the Best Visual AI in the World... By A Lot.

Limitless Podcast25:40

Transcription

Just yesterday, OpenAI released ChatGPT images 2.0. And the model blew my mind. I was up until 2:00 last morning playing around with it because of how powerful it is.

As I was watching Sam announce this model, he was talking about how image gen wasn't really that important to him. He felt like they already had a good image generation model. When he was presented with the outputs of this one, he had his holy moment. It's actually really phenomenal. And through trying it ourselves, we have uncovered that is actually true. I mean, we've really frequently used Nano Banana as the go-to default image generator, but now it's getting close to being indistinguishable from reality entirely. And we have a series of examples that we're going to show you that are probably useful for your actual applicable life. Things like interior design or generating comics or generating sales graphics. I don't think there's anyone who wouldn't find a beneficial use case of an image generator model that is as good as this one. So, let's get into the actual announcement. Let's walk through the examples. It's pretty amazing stuff.

Around uh midday yesterday, OpenAI tweeted this very mysterious post and it goes uh this is not a screenshot which is weird because it looks like a screenshot of someone's Mac desktop except this is completely AI generated and this was the precursor to their official announcement which is ChatGPT images 2.0. It's their new image model and it absolutely blows every other image model out of the water and I don't mean that as an understatement. It is number one across every single image benchmark. It's beaten Nano Banana 2 uh and any of the Chinese image gen models just completely don't weigh up.

So, what are some of the new things here? Well, the fidelity and quality of these images are incredibly high. You're seeing a demo video here where we have a chameleon in various different positions. The wording text is typically such a hard thing for image models to nail, especially within the AI world. It would like jumble up the letters or it wouldn't spell things correctly. Now we have that completely and utterly resolved. And so you can see some of these examples come to life here. Uh for example, look at the fidelity of this image of rice. Typically this would just look like a garbbled white mass. And now you can individually see each grain which is pretty uh nuts. And then you have examples which are a little scarier where this looks like a real screenshot of a handwritten note uh in someone's stylistic way but it is very much completely AI generated. So you can imagine this could be used for various different nefarious purposes which might get more malicious. But there's a ton of different examples and we want to get straight into it starting with ones that we've generated ourselves. There's one around furniture, right?

>> Yeah. But I actually want to start with the rice one because what you you mentioned with the rice is that it's precise enough to show the grains of rice, but it's also precise enough to write a single word on a grain of rice. And that fidelity is new. So what I did is I actually went to ChatGPT myself and tried to emulate this and I asked it to create a piece of rice with the word GPT image 2 generated on it and this was the output that I got. Um or actually this was the first output that I got and I spent maybe 5 minutes trying to find the grain of rice. I don't think it worked. Um so I asked it to draw a box around the grain of rice and it drew a box and then actually etched it in the middle. So there are some edge cases that don't quite work. I mean that was not in the original mentioned furniture.

I am currently living in an apartment that can use a little bit of extra what my apartment looks like. This is a much nicer variant of something that I would like to aspire to. So what I have prepared here is a reference image for ChatGPT along with a prompt of what I would like it to do. And that involves just doing things like um adding lamps and adding different furniture. basically swapping out the existing furniture that exists in this living room and moving it into a totally new vibe and style that I think I would more likely appreciate and resemble. So, while that's thinking, I guess we can kind of get into some more of the interesting parts of this model.

Well, I have a example that I actually have ready to go here. Um, I was kind of obsessed, don't tell anyone this, I was obsessed with manga as a kid. And so I was like, you know what would be cool? If we could turn our show, you and I, Josh, into a manga comic. So I created this detail prompt and I gave this beautiful photo of us. But my handsome guys, look at look at those handsome, very handsome guys. Um, and uh, I basically asked ChatGPT to generate the prompt for me. So, I gave it a rough idea of what I wanted to create, the scene, as it were, and it created a very detailed prompt with stylistic references, details, stuff that I wouldn't know because I'm not a storyboarder. I'm not a manga creator. But, funnily enough, I have an AI that can do it for me. So, I don't know if anyone is uh paying close attention to this story line here, but if you're not, that's great cuz I want to show you the end output. So, it's, as you can see, very long prompt, and this is the finished result. So, what you are looking at here is pretty good. Josh and I, let me explain this. Josh and I have been filming a podcast. As you can see, we've got our setup over here. But then we look out the window and there is a shadow and we notice that it is Sam Altman Godzilla size coming down upon us terrorizing uh New York City. This is I'd say the time estimate is roughly 5 years in the future, maybe even three. Don't know how how quickly AGI gets here. We grab our weapons. It is clawed. This is not a sponsored video, by the way. I just came up with this randomly. And it shoots out prompts that wrap around Sam Altman and eventually bring down GPT5 from taking over the world. Um, now you know what's going on in my head. But if you just notice this, like look at the fidelity of this. This like took 5 seconds to create the prompt and then another 2 minutes to create the actual image. Look at the fidelity of this. Like the writing is all accurate. the like this would cost like thousands and thousands of dollars and weeks maybe months of time to actually create from scratch and this did it in a in a bunch of seconds for a couple of cents. Like it's pretty impressive.

>> Yeah. Oh, it's so good. So, if if manga isn't your thing, we have the furniture example. It's ready to go. So, here I have the original that we're seeing on screen right now. This was the original living room. I fed it the prompt and here is the new one. It totally maintained the integrity of the room while swapping out just a few key pieces of furniture to change the vibe. And I think it's a testament to a practical use case that a lot of people might have is they want to design things. They want to make things look good and maintaining the personalized fidelity of whatever space it is. If you have a piece of clothing, I know this works for tryons. It's really good at maintaining continuity throughout these images. So, I thought that was a pretty interesting thing. If you have an apartment, if you have a closet worth full of clothes, you can just place those clothes out, take a picture of you, take a picture of their clothes, ask it to dress you, ask it to redo your living room, whatever it may be. Super powerful and works fairly quick. I mean, this output took maybe a minute to generate. And for those not sure, this is actually available to all users of ChatGPT, I believe. Very limited instances for the free users, but if you have the plus plan for $20 a month, you can just go off and start creating images and they will look just as good as this one.

Yeah, I mean if you're a professional um that has been toying around with using AI but it's just never been good enough. It's always got some form of error whether minor or big um now we have a tool that actually works for you. So if you're a designer, if you're a floor planner, there's a bunch of other examples I'll show here. This becomes a practical tool like GPT images one was very much a novelty and a toy. It was fun to see everyone in Studio Ghibli versions of ourselves. But now you can use this to create certain things. Now, not all use cases are good. If you're like me, I use social media to disseminate a lot of the breaking news that happens in the world of technology, AI, or whatever it might be. But you now have reached a point where we can't necessarily believe everything we see. And images too from ChatGPT doesn't make that any easier. Um, what you're seeing on the screen right now is not an official take or update on the Bloomberg terminal. Uh that is also not my desktop monitor. Um this is completely uh AI generated and you can probably tell parts of this like kind of gives it away. It's a little too zoomed in. Um unless of course you can like change the default settings on your Bloomberg terminal. But some of these things are are really good. Like this is exactly where this is on the Bloomberg terminal. The percentage mark isn't that large on the actual thing, but it's it's got all the sections pretty much nailed. So you know that the model looked up uh official Bloomberg terminal uh layouts and like recreated it. but it added a completely uh fake kind of like bit of news. So um you could change that bit of news to represent real news, but it would still be fake. So So there's a lot of like avenues here for misinformation or disinformation. So like not entirely accurate, but somewhat accurate. You can imagine the kind of social media frenzies that this would create if people were to believe and buy into these things. Like imagine if you read an announcement that wasn't actually real, bought a stock, and then it like realized that it was fake and then it crashed. You could end up losing money. You could fake data. There's a lot of avenues that this could go down.

>> Yeah, there there's two points on this. One is that like we're at the point now where even if you pixel peep, it is it is almost indistinguishable from real life. You can't really tell what is AI generated and what is not. And as that kind of gap converges, I imagine it will create a lot of chaos where there's just no way to tell what's real when these images are so good. The second thing that I'll mention is from this model in particular, anytime it's asked to generate a visual asset of a piece of software. For some reason, it's exceptionally good at understanding the nuances of every frame of every piece of software. Last night, I had it do Da Vinci Resolve, which is what I edit a lot of videos in. I had it emulate Photoshop and it got every single slider down to like the correct pixel, which leads me to believe that it appears as if there was some training customization around these software projects in particular. And you have to ask the second order question, why why is it so good at all this software? And I guess the answer for me is well, it's probably because they want their agents to understand how to navigate it and then eventually emulate it and then eventually replace it. And this nuanced understanding of how everything works is training for the image generation model, but also training for just I mean the future of what these agents are going to look like. So there might be some hidden stuff going on behind this image generation model as well.

So back to the demos. In addition to these capabilities, we have another one teed up right here, which is to create a premium infographic poster. Another strong suit of this model is text and how well it's able to render text that looks lifelike, looks accurate, and is able to kind of create a like a storyboard if you will, a poster. It can create multiple outputs. What I've asked it to do here is create an editorial infographic and this is the first time I'm actually seeing the output of this and it seems pretty cool. So this is for Limitless as you are familiar with and it kind of walks through our weekend review. So, the things that Limitless mentions, this is the poster that is serves as like the weekly roundup, the weekly review. It is pretty good. Is it Is it accurate? Let's Yeah, I'm I'm curious. You check the accuracy. I'll check the QR code. See if that works. Cuz word on the street is that QR codes work pretty well. Wow. I might need to replace the entire roundup newsletter, Josh, with something like this. Just a quick a quick glance, quick tip. You can imagine how this can kind of carry out to other applications, right? Is like if we want to juice the newsletter up a little bit. We could just create a graphic with one prompt by feeding it the context of everything we spoke about to give you this detailed infographic. This also applies to educators and people who are teaching things. Uh it's really easy to make graphics on particular lessons or mathematical equations or diagrams or anything you want visually represented. It's exceptionally good at that. So I thought this demo was kind of fun. It creates the Limitless we can review as a poster that's printable. The QR code does not work, but I've asked it to make the QR code scannable. So, while it finishes that up and we test that, maybe we can go on to another example.

>> Yeah, I was just going to say before we move on, um the educational point is uh a very precient one mainly because like if you're like me, >> uh you could like read as much text as you want, but sometimes a visual like that summarizes everything really helps. You can now plug like an entire book's worth of text into a single prompt. Like like a lot of these frontier models now have like a million context, which is like a couple of novels or like many many novels. And so if you can imagine if you're trying to learn about something and you want like the key points, you can not only ask the AI to summarize things and give you a bullet pointed list, you could get them to transform it into an illustrative poster that just you can look at in a glance before you go to bed and learn something brand new. So I can imagine this being used in science as well. um back when I had a bi biology degree or back when I was doing my degree um I remember we used to have these like research poster conferences and they used to be like like I don't know A1 size it was absolutely massive and you would have so much condensed information there and it took me weeks to make and the fact that I now have a tool here where you can just probably plug in a bunch of papers get extract the right information and then put it out in a very visual way just blows my mind like we are condensing a lot of frontier research and education tools like with this one simple update. It's it's very very cool.

But to move on to uh one more example that we generated, um one thing that's cool about images too is you can play around with one image and make it into several different aspect ratios. So what we have here is an individual, I don't know who this individual is, but it just generated it looking out onto the greatest city in the world in my opinion, New York City. And it's in like a nice little sunrise or sunset. I can't tell which one is which, but as you notice, um, it gives us different aspect ratios of the guy. Like over here we see him on the left. Over here we see him from a distance back. Over here we see a panoramic view where we can see him looking out onto what is this? This is Brooklyn Bridge. Um, so the details of it, um, you know, you can see some of it is like kind of like blurred aspect ratios as well is just very impressive. And you could start creating like storyboard sequences from this or just kind of like pitching visuals to whatever um you whatever kind of like idea or concept you want to make. You could use this in the product realm if you're like trying to figure out whether a model looks good advertising your product in like let's say the product was a coat in this particular way or it could just be something advertising completely different. It's it's very cool.

So how does this model perform so well I think is the question and one of the novel breakthroughs that this image gen model has that others don't is the detailed reasoning capabilities. This is an image generated model that will think before acting and will reason through the steps required to get the best image output. So generally it's just pure inference. You give it a context, you give it input and it just spits something out. This one actually reasons through the I guess the reasoning of why it's doing these things. And that's part of the reason why even though you're not giving it necessarily the best prompt, it's giving you a really powerful output. And I have another fun example here of um just like more comic books that you can make. This was a single prompt and that generated like an entire comic book with a really accurate character that's carried throughout. Another fun feature is the character continuity where you can generate a character. It will be prevailent throughout all the images. And then also one last example that we have here is of anyone who's involved in social media or just creating any sort of marketing material. I asked it to create an ad package for a masha shop in Williamsburg called Sage Bird. And Sage Bird now has a full kit of various aspect ratios to be posted on any platform that looks photoaccurate. If you'll notice, there's even a street sign that says Bedford Avenue, which is a street in Williamsburg, which is very funny. So, I think the the fidelity, the quality, the capabilities of this model are really endless. And again, the constraint is your imagination with how far you can push this thing because it it's just it's so powerful. I had so much fun using this. I must have generated at least 100 images so far just in the last 24 hours. And it is like it's so fun. I recommend everyone go and try it and figure out what use cases are best for you.

So, a question that came to mind immediately is okay, it's good, but how does it compare to some of its competitors? Primarily Nano Banana 2 from Google who has previously held the number one spot here. Um, now if you look at this image over here, it's not just number one, it's number one by a far mile. I think it has like 150 point increase on image arena. Um, if you don't know what this is, this is like the number one benchmark to test these image models. GPT images 2 isn't just number one overall. It is number one across every single category that is measured within this benchmark. So it has a long shot >> by a long shot. So it has a very distinctive lead. And if you're looking at this and you're saying, "Okay, well, whatever." People can like orient benchmarks around this. So like we don't know if it's real. I have a direct comparison for you. So the same prompt fed into GPT image 2 versus Nano Bonata Pro. And you can see that there is quite a lot of differences. You can see uh GPT Images 2 over here on the left. The lighting is much brighter. The fidelity is arguably a lot better. Um and as you can see, like, you know, there's more expression on her face. She's smiling. Um and there's a lot more things in the background. Like if you look at the plants in the back, it's it's way more hyperrealistic and harder to create for an image model. Now, if you look on the right, Nano Banana 2 is very good, but there's less complicated things going on behind them. The lighting is a little bit off. And you can kind of tell that I don't know, like maybe still on both sides, you can tell that they are kind of slightly AI generated. I would actually argue that images too, now that I'm looking at it for longer, uh looks like the glisten just seems too glistening. Um but Nanobon 2 can get away with it because the lighting is a little less. The point is these models are getting way way better and the examples keep coming but it's not just visual things like we're not like social media influencers don't have to be worried here. Um you can start using this for very practical purposes.

Now, there was this awesome example over here where a guy took an image of a book, right? And he said, "Could you generate me a barcode for this book?" And he generated the barcode and when you scan the barcode, it takes you it it's basically an embedded link. It takes you to a page where you can then purchase or buy the book. Now, this is very impressive for if you are like trying to sell a particular product, especially if it's physical. You now don't need to go through the complicated process of generating barcodes, getting it printed. You could feasibly create your own design book cover, print it out, and then wrap it around your actual product, and it actually works. It works with your internal system. Um, so I just thought this was pretty cool.

>> Yeah, it's amazing the clarity. And again, I think this is a testament to the reasoning where it can actually reason its way through and generate an accurate barcode in a world where it previously couldn't. So now, not only can it make infographics, but it could link these dynamic elements to real world artifacts, to a custom domain, to your book. They're actually usable without needing to take it into Photoshop and take it that final mile. And that's like a really cool unlock. We have another example here as well of uh the front page of the New York Times, which of course is entirely fabricated um or or at least partially. So like this isn't a real article. This isn't a real image of a paper, but all the information on it. So if you actually dig in here and read it, um all the information about open air unveiling GPT image 2 is accurate. They pulled it from the blog post. They didn't you didn't have to provide the blog post. They independently did it. It reasoned through it, pulled out the most important points and then wrote it in a stylistic manner of a New York Times writer. So you can start like imagining what this could do for press and media. If you are a reporter, you might be thinking, "Huh, so you're telling me I could just feed this the bullet points that I want it to make and it could write it in my voice, in my DNA." um that I like stylistically write an article for. That's amazing. You could also ask it to generate the image for you. So there's like this metro approach where like you're talking about the product, but then you use the product to generate a example image that you then put in. This is of course also generated by uh images too. So there's a lot of applications here. Again, I mentioned earlier disinformation is a very real thing. So you can imagine people sharing fake news articles about things that aren't real that might sway markets or inform people uh in the incorrect way. But cool nonetheless.

Yeah. And then there's more examples for people who are involved in architecture at all if you're doing floor plans. I mean, this one was cool where you you fed it an image of a house and then it generated a floor plan. But the the next example I think was even cooler because this was a digital rendering of a large building that had all the specs listed next to it. And using that spec sheet and using their 3D rendering, it created a fully rendered floor plan that you can actually use and send to an architecture to or send to an architect to to actually make blueprints and build the building. I'm not sure if this is up to code. I'm not an architect, but I imagine you can probably iterate your way through this with a proper architect to get it to be compliant, to get it up to spec if it's not already, and train it to do that. So this there's this unbelievable unlock that happens for pretty much any profession that's generating any sort of image. All you need to do is put a stamp on the bottom. It looks like it it already stamped it with some fake stamp of approval. >> Um, but I'm sure if you do this type of work, you can kind of you could put your own spin on and throw your own stamp on there. Um, if any of you are architects listen to this, I I encourage you to to try this out cuz I'm actually curious whether this is accurate or if not like how accurate is it? Because obviously like architects in training like train for seven years at school which is just insane. they have to um understand the physics behind the buildings that they're designing. And I'm wondering is this physically accurate? Are the estimates uh like do they make sense or is this completely made up and we still have a long way to go? It it looks legit to me but then I'm not an architect. So if you're if you're listening to this, let us know.

Um, there's another cool uh thing here where again I mentioned earlier if you are a visual learner, sometimes you just there's too much information. You can create these posters bracketed by uh a particular subject and it kind of like splits it up. So like with here we have like all the things going on in AI. You got AI models and agents, robotics, semiconductors and you just have images which explain the start to end process of creating these different things and what they actually do with a few words underneath it which I thought was cool. Um, and then there was this final example over here from Matt Schumer where um, I can relate to this cuz I formerly worked at a uh, big four consultancy and we had to create slide decks and it would take so long um, because you had to move things in a specific way or reformat the text. Um, and Matt Schumer oneshotted an entire slide deck um, by just providing it a bunch of information and it created it in the style of Spotify by the looks like it uh, by the looks of it. So very cool. Loads of different applications and I can't wait for more people to actually use this for professional purposes.

>> Yeah, the model's awesome. And I I guess the ask is to share whatever you're using it for because again like those prompts, those examples are the only limiting factors to really what this can do because it has the reasoning because it's so capable. It has the like pixel perfect fidelity. It's really just a matter of massaging it with prompts to get the output you want. Not really a limitation of the model anymore. And like to Sam's point early in the episode, it seemed like it was great before. Now this is just unbelievable. I can't imagine going back to Nano Banana Pro knowing that this exists. And it's just a testament again to how fast we're going and like what the downstream implications of this may be in the future when you can generate infinite images for cheap that are pixel perfect and indistinguishable from reality. What type of downstream effects does that have on every visual artifact that we interact with on a day-to-day basis? I mean, there's no way you could be sure. And this this has a lot of implications that I'm not sure we're fully aware of now, but will surely become known well as we kind of navigate through this. It creates a weird dynamic that seems a little uncomfortable. Like I And now I have to navigate the internet with such a strong filter to just try to parse through what's real and what's not. I'm curious um whether this tool can be used to generate visuals that humans hadn't thought of before necessarily as the AI becomes smarter and is trained on our prompts and largely our flaws like you know you can ask uh an AI to generate a detailed prompt to then prompt it itself because we don't know how to prompt it itself like it can do the same with images where it's like uh I I get that EAZ probably missed this point and so maybe if I create this visual in this particular way it's one that he hadn't thought of, but now it like breaks new ground for it. So I I wouldn't put it past this model to to like the model that we have to today to generate something a visual artifact that will soon be uh kind of like groundbreaking for humans to use. Like maybe it's not a poster, maybe it's not a slide deck, maybe it's something completely new that we haven't seen before. Pretty exciting stuff.

Yeah. So that's ChatGPT images 2.0, the newest and hottest image genen model in the world. I encourage anyone to try to displace it because uh that would be amazing if it gets better than this. But it's worth trying. It's worth sharing what prompts you use that give you some specific outputs that you may find helpful, interesting. The use cases are the currency. Please share yours in the comment section down below. If you enjoyed this video, don't forget to share it with a friend who may also want to generate some images. Perhaps they're involved in social media. Perhaps they just want to redesign their hypothetical apartment. Whatever it may be, it's fun. It's worth testing. It's worth trying to just like feel it and understand the intelligence. But yeah, I think that's pretty much it for today's episode. You guys any final part any thoughts here?

>> Nope. Um, if there's uh one request that I have, I want to see the images that you generate and try and surprise us. Try and do a use case that we haven't covered on this particular video cuz I'm curious of the creative purposes around this. Um, our social media profiles will be linked below. DM us there and yeah, I look forward to seeing what you have to to make.

Yeah. >> Awesome. Cool. All right. We'll see you guys in the next episode.