Transcription
Earlier today, OpenAI released a brand new image model called GPT image 2. And this new image model is best in the world by far, and it's not even close. Here is a simple prompt, an in-game screenshot of Rust. So, it can generate perfect images of video game. This image model is almost perfect at generating user interfaces. So, here is an image that was generated of YouTube. Here is an image that was generated of chat GPT. Here we can see GBT image 2 with the prompt, what if a hot sauce came out of a toothpaste tube? It turned this image into this image. This image model is going to change business and content creation forever. It's literally the best in every single category. Not only is it the best overall by far, it is the best at single image edits, multi-image edits, text rendering, product branding and commercial design, portraits, photorealistic and cinematic imagery, cartoon and fantasy art and 3D modeling. Today, we're going to dive into this brand new OpenAI image model. And in this video, I'm going to talk about my initial tests and one absurd use case that I was just so blown away by. We're going to talk about how to get started and we're going to do some live testing and then we're going to use it with AI agents in the codeex application. Let's not waste any more time. Let's dive into the video.
Okay, so today is April 21st, 2026 and OpenAI introduced images 2.0. The last version was 1.5. What was really cool about the last one is you could actually generate images with no background. It had like a transparent background. You actually can't do that yet with the new model. So before we dive into actually how to get started using this image model, I want to talk about my experiments this morning when I first tested this new GBT image 2 model. So the first thing that I did was I uploaded four images of me that are kind of in different angles and I just said, "These are four photos of me. Can you please create an image of me on the cover of a magazine?" And then this was the results. We have a pretty realistic image of me wearing the same shirt that I'm wearing in this one with basically perfect text on the image here. This is pretty cool, but it's not that crazy.
So, this isn't the most useful application of this technology, but this actually blew my mind. Check this out. I said, generate an image of the book, good to great. The barcode on the book should scan to the actual book. So, what we're going to do is we're going to full screen this and we're going to zoom in on this barcode. I have a barcode scanner right here. And let's go ahead and go like this. Bang. Good to great. The barcode scanner fully works. Like what? Do the same with the book, The Intelligent Investor. Okay, it's done. Now, what I'm going to do is I'm going to open this up full screen. We'll give this a scan. Boom. Look at that. The intelligent investor. Here are the videos. One thing I want to test here, I want to make sure that it's not picking up and just reading the ISBN number. So, I'm actually going to open up Canva. I'm going to paste it in here. Now, what I'm going to do is I'm going to I'm going to black out this part of this. I want to see if it's actually scanning the barcode. So, let's see. Bro, that's crazy. That is actually wild. What the hell?
In the next experiment, I uploaded four images of me and I said, "These are four photos of me. Can you please create an image of me as a cartoon with exaggerated physical traits? Almost make fun of me, but it should look like me." And this right here was the result. And obviously, as many of you have uh mentioned in the comments, I have big ears. So, it decided to exaggerate my ears. And it also analyzed the images very closely. Right? I had an AI powered journal in the background of this image and it put it here. It also um put not a morning person and it put uh a cup of coffee and in the background of this image there's a coffee menu. And overall I think this is pretty good and pretty funny. But this is where the first real experiment happened. So let me show you what happens next.
So then I asked for 11 different edits in one prompt. So in this one single prompt, I edited this image and I wanted it to make 11 different edits. And I'm going to show you exactly how many of those edits were made perfectly. So after the image edit was made, this was the result. I want to go through all 11 edits that I asked for and let's see how it did. So let's go through one by one. So, the first one is make his coffee say Riley Brown. Boom. Done. Then it said, "Get rid of the Red Bulls." As you can see here in the first one, there's Red Bulls and it got rid of the Red Bulls. Third one. Now, this one is a trick that I tried to do. I said, "Change his shirt to orange." And then later in the prompt, his shirt should become a brown turtleneck. And it it ended up being a brown turtleneck. So, three for three. Then it said, "Change the background screen to say vibecode.dev." You can see here it has some AI power journal. Here it says vibbecode.dev. Here I asked it to say today's plan should be for a content creator and it changed this. Here it says overink, overexlain, overcaffeinate, maybe work and it changed this to be for a content creator. I said change the sign at the top to GPT image 2. It did that. Uh make the bobblehead in the background a monkey. It did that perfectly in one shot. I said, "Make the microphone brand Palander." And then for the next one, I said, "Add an earring on his left ear." You see that this is the ear on the right, but it is his left ear. It did that correctly, and it should be a pink diamond. Boom. I said, "Make his hair a skin fade." You can see here, it's not really a skin fade. In this one, it is a skin fade. Um, and then, of course, it got the brown turtleneck correct. I said the sticky note should be orange and say, "Keep winning." You can see here it says, "Not a morning person down here at the bottom." And then it switched it to keep winning. And and then I I added at the very end I said make these changes and change nothing else. Keep all in the same position. So it executed this edit perfectly down to the pixel.
And then of course I had to test it on creating cartoons. So I said these are four photos of me just like the previous prompts. Could you please create a 2D comic of me in the8s? Should capture the politics of the time. Eight photos in one. So, here's the prompt and here's the result. You know, I don't really remember the 80s or even the 90s for that matter, but this does seem relatively accurate. You can pause the video to take a look because the real experiment of this one are in the next two prompts.
So, the next prompt that I did was an image edit. So, I used this image in the background here as a reference image and I said, "Do the same but for Europe. make me look European by changing my hair and clothing and surroundings, but I should be the same. And so here is the result. So here it made me more European. I'm wearing European European clothes, but you know, I am an oblivious American and I don't really understand the culture of Europe. So then I decided to try this prompt. This is a category of image generation that I call overlay explanation. So, I used this image as a reference or I added this to the prompt and I said, "I don't understand these references. Can you make an overlay in red which explains the references in simple text, please? Change nothing about the image except add an overlay which explains all of the references. Use a legible font that looks handwritten and arrows that look pendrawn." And then here's the result. it created this image and added these little uh red annotations over the top of every single image. And since this is a thinking image model, it's able to actually think about all of the different parts of the image and then add this annotation, which is really cool. We can see that this is a peace sign, which is the anti-nuclear movement. We have the labeling of the Berlin wall falling, the end of division in Europe. Here we can see that this is Cheranobyl disaster, a nuclear accident in the Soviet Union. And anyway, this is a pretty fun thing to do on top of images.
Okay, so how do you get started? Well, it's pretty easy actually. All you have to do is type in chat GPT into Google, then go to chat GPT, and then you can click create an image. And this will add this little image tag in here indicating to chat GPT that you are indeed generating an image. One of the coolest use cases that I saw was creating blueprint posters. And so I saw this prompt as a template and I can very easily upload an image of myself to this. And we can run this prompt. And when you enter your prompt, you will see the reference image up here. You'll see the prompt that you entered to create that image right here. And you will see your image being created. And I really like the UI. They made it a lot simpler. And then it takes around like 20 to 30 seconds to generate each image. But as I'll show you later, you can actually use agents to generate the images so that you can do them in parallel. We'll get to that in a little bit. All right. So, our image is done. And I'm just realizing now that this is meant for like interior design, but it created this pretty funny image of me. Uh, and it's measuring all of my measurements, which is actually kind of funny.
One cool thing that you can do is you can use this select tool. And this allows you to select a certain part of the image. And remember, you don't need to be exact by any means. You're basically just giving the AI a general idea of what you want. And so I'm just going to highlight the the hair here. And all I need to say, right, the AI is going to have context of this. I say make this solid uh white. Color it in. So I'm now giving context to the agent. And you can see here that instead of showing an image as the input reference, there's now a selection in the input reference. And I can't open the selection, but I can see that there is a selection that was made and it is now working. And then, as you can see here, it was perfectly colored in.
And now we're actually going to test iPhone mockups for mobile apps. So this is my app, Vibe Code app. As you can see here, we're going to upload one, two, three, four, five app screenshots into the prompt, as well as the logo. So, I'm just going to drag these six images in, and we're going to drag these in right here. Now, I'm going to say, here are uh five app screenshots of my iOS app and the logo, the orange logo. Can you please make a really cool horizontal image that shows five high quality iPhones with these exact screens on the iPhones hovering over a beautiful background with the logo uh on this and it should say welcome to the vibe code generation and the background should be insanely beautiful nature green and uh yeah very green uh background that kind of matches some of the green tones in the screenshots. But these the mock-ups of these phones that are floating should look very three-dimensional and beautiful. This should be the most beautiful image ever. So, as you can see here, all six images are entered as a reference. We will see how it does here. Okay, so it is done. Let's go ahead and open this up. I don't absolutely love the image overall, but I want to bring up the fact that if we go to Finder and we actually were to open this image right here, we can see that this phone screen is right here. This is very very similar. It did put the vibe code up here and there is five different uh images here. Let's go ahead and open up this one. As you can see here, this is this one. Look at this. This is very similar. So Nano Banana, which was the best image model previously, was not this good at getting them like pixel perfect on pixel perfect on these UI designs. This is basically exactly perfect with all of the icons, all of the text, the entire keyboard. This is actually mind-blowing.
Okay, so I actually got this idea from this guy's post. This one looks way better. So, I'm going to copy this and I'm going to go back to chat GPT here. I'm going to say the phone screens look really great. So, keep those exactly the same. I just want it to look a little bit more like this. And actually, don't have the text or logo. And I'm going to paste that new image. And now, let's try it. So, we're going to give one image basically a reference image. Okay, it's done. Let's go ahead and check this out. Wow, this is pretty damn good, right? This is looking good. So, it looks a lot better than this background. Something about the background is a little bit off. But here, after we gave one solid reference image, we have this like light coming down. Maybe what I can do, well, look at this. Now, I'm going to take a screenshot. Now, whenever you want to give visual context to an AI, I recommend using this tool called CleanShot Pro. I believe they have a free version. I don't know if this is paid. I forget. But we can actually um really quickly take a screenshot. Now I can just immediately draw on the image. And what I'm going to do is I'm going to go like this. And so I can very I can just easily copy this and be like, okay, this is fantastic. Please make them in the shape shown uh in this image here. Now, we're going to wait for the image to be done. Okay, and it's done. Check this out. Very similar. As you can see here, we have this original shape straight across, and now we have them in this nice arch. This is a pretty cool use case.
So, I have a theory that uh because I uploaded five screens, it didn't look quite as good as this. So, I'm going to please complete uh emulate the first image that has two mock-ups except with the other two screens. Uh, please do that and slightly change the background. Don't make it like the exact same thing, but the point is I want a really high quality 3D renders of phones showing off the screens and it should look absolutely beautiful. All right, so it's done. And I think my suspicion was correct. These are much higher quality screenshots. You know, five screenshots, there's a lot of de detail to pack into a single image. This is a much closer up image. I think this looks so much better. This is a very high quality image. This is really cool.
Okay, so before we get to using this in Codeex, their new AI agent super app, I want to talk about where else you can use this image model. And if you go to open AI playground and you open up their platform here on the left side here, you'll see an image tab. And in this image tab, you can select GPT image 2 and you can change the settings. And you can see here one of the coolest things about this new image model is you can generate images in 4K, you can generate images in 2K. Um, let's try the 4K one. I want you to generate an image of 175 people in the crowd. And there is one purple dinosaur in the crowd. And so this is a place where you can kind of generate uh images. It's a lot easier to uh rapid fire send images. Here you'll actually need to sign up for the OpenAI API and you'll actually have to go into billing and add some credits. It is a separate billing process, but this is a way where you can like rapid fire like I can type bunny and I can rapid fire generate images and I kind of have more granular control over the image model and all of the different settings. Okay, so I have no idea how many people are in this. Let's go ahead and open this up real quick. I do not want to have to count this manually. Please add a overlay of counting how many people are in this photo. The photo should stay exactly the same except from the top left down to the bottom right you should number each person and the bottom right person should be the last number which should tell us exactly how many people are in this image. So we're going to use AI image to count how many people are in this image. Let's see how it does. All right, we hit our first limitation. This is the first limitation that I've seen with this image model. It cannot count properly. It's like double labeling people. Um, and there isn't even one, right? We don't even see one. We see 1 2 3 4 5 6 nine. Um, it is not able to separate the different people and number them correctly. Here we have like all of these people have like half of their body showing and the people who have their full body showing this woman is 226, 227, 225 and 241. And you can see that pattern and here we have a total of 263 which is more than 175. So that's just one limitation of this image model.
Okay, my favorite part about this new announcement is that it is baked in to codeex. And you know, two days ago, I released a video on codeex. It's an hour and 40 minutes explaining why this is the number one tool to learn and everyone just automatically has access if you have a chat GBT account. Even if you have the free account, you do get access to codeex. And and so very simply, you could say that this uh agent platform codeex is a lot like cursor mixed with claw code mixed with lovable. that is a general agent tool kind of similar to clawed code or even open claw and mixed with cursor because now it is basically a full vibe coding platform. They have a browser that opens up here on the right side of the the app but you can also create documents. It's also like co-work because you can create powerpoints. And so that's the use case that I'm going to show you right now. So, you know how on chat GPT this whole time basically we've had to prompt the AI to create this like we actually had to type in our image prompt right here and say create this thing that does blah right we type in our prompt here and we have to do it every time we want an image well now image generation is just a tool for codecs codeex is the best coding model in the world it's one of the smartest general agents in the world and one of the tools it has access to now is image generation And so a simple version of the prompt I entered was check readwise which is my second brain tool where it's automatically hooked up to my Twitter bookmarks. So all of the tweets that I bookmark get saved to read wise. And so I just told Codex I said I said check readwise create a PowerPoint which is a skill right the same way you can create a claude skill you can also create a skill up here if you click on plugins. Um create a PowerPoint presentation of my recent saves with annotations. And I basically said each slide is a GPT- image-2 image that has the OG tweet and annotations over the top. And so then it did take like 10 minutes because it actually had to go out and research these annotations, right? It actually went out and learned about the new image model and then it realized it was a state-of-the-art image model and it add contextaware annotations over the slideshow. And you can go through here and we can see that there are 10 different slides and these are all images that I've saved. It even added the profile picture to this single image on each slide. And so we have context over all of the tweets that I save. Traditional SAS will become headless. UI becomes optional, right? You'll die unless you build for agents. Agent native product design. So, this is just a fun use case of a single prompt that generates a presentation of many images. And what's really cool about um codeex is you can just very easily export this to Canva, which I was already using at the beginning of this for my slides. And you can see here it's loading. And you can see here each slide has one GBT image. I did not prompt this specific image. The AI agent codeex used the GPT image 2 tool to generate these images on my behalf. And so if you give it enough good instructions, you could generate thousands of images for you based on your own liking. And so I believe that way more agents will be prompting GPT image 2 than humans because agents are getting so good at following patterns. Anyway, so this model came out only a few hours ago, I'll be covering it in great detail. I'll be covering image generation a lot going forward, but it'll probably be in the lens of AI agents. That's just my biggest passion right now. I think that's the highest leveraged thing to learn in AI. And so I'll be covering that a lot. But thank you guys for watching. This was a fun video. I'll see you here for the next.