📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Complete ChatGPT-5 Breakdown and First Impressions

Matt Wolfe25:36

Transcription

Earlier this week, OpenAI pushed out GPTO OSS, which was basically them putting out their most state-of-the-art model as an openweight model that anybody can use for free. And today, the reason they did that became a little more obvious because, well, we have a new state-of-the-art model. Today, August 7th, 2025, was the day they released GPT5.

The live stream that was promoting and launching GPT5 was nearly an hour and a half long. But in this video, I'm going to break down just the things that I think most people want to know about GPT5. And then after that, I'm going to share some of my thoughts on whether or not I was underwhelmed, overwhelmed, properly welmed, and what I wish I would have saw from this launch. So, let's dig right in.

First off, GPT5 got a lot smarter. So, they described GPT3 as like being as smart as chatting with a high schooler and GPT4 as smart as chatting with like a college student. Well, they described GPT5 as more like chatting with a helpful friend with PhD level intelligence. But not only is it like talking to somebody with a PhD, it's like talking to somebody with a PhD in pretty much every area of expertise. Now it's like talking to an expert, a legitimate PhD level expert in anything, any area you need on demand that can help you with whatever your goals are.

As with every new model release, it pretty much crushed every benchmark they put it up against. In the competition math benchmark here, when it had access to Python, it scored 100%. When it wasn't using tools like Python, it still got a 96.7%. It beats out all the other models at Frontier Math. Keep in mind, they're only comparing it to other OpenAI models, but I believe it's still the best benchmarks for any model. On the Harvard MIT mathematics tournament, when it had access to Python, scored 100%. In this Google proof exam here for PhD level science questions with Python, scored 89%. Humanity's Last Exam pretty much top of the pack does great on coding benchmarks again only comparing itself against the OpenAI benchmarks. Enthropic released their Claude Opus 4.1 which was the state-of-the-art coding model up until GPT5 came out. So to compare, we can see Opus 4.1 got a 74.5% on the SWE here where GPT5 got 74.9%. So slightly beating out even Anthropic's newest coding model. And as we scroll through the blog post here, we can see all of these different benchmarks where GPT5 is pretty much just crushing all of the older OpenAI models. On this college level visual problem solving, it got an 84.2%. Plot Opus, which came out just a day or two ago, got a 77.1%. So, even though they're not comparing their benchmarks with Anthropic's benchmarks, they're even beating out Anthropic's newest models benchmarks.

Saying all that and showing off all of those benchmarks. Even OpenAI is starting to admit that the benchmarks mean less and less to most people these days. Benchmarks, they're exciting numbers, but we're starting to saturate them. Like, when you're moving between 98 and 99% in some benchmark, it means you need something else to really capture how great the model is.

Another thing that they mentioned in the live stream is that they're going to go away with having to pick exactly which model to use. You know, do I use 40? Do I use 03 to get more thinking? 03 Pro to make it think even longer. They're going to just get rid of all of that. It's all going to be GPT5. And when you give it a prompt, GPT5 is going to basically decide how long it needs to think and the best way to get to the response you want it to get to.

They also mentioned that this new model is really great at picking up little nuances. So, if you give it a really long prompt with a whole bunch of subtle details, it actually does a good job at picking up those subtle details. You specify a long complicated task with lots of subtleties in the initial instructions. It's very good at picking up on those subtleties. It's also very good at if it's gone down a wrong path and actually goes and executes the code or hears back from you that it was incorrect. It's very good at backtracking, too.

During the live stream, they interviewed a handful of people that got to use GPT5 early, and pretty much every single one of them said that they were very impressed by the speed. So, apparently, it's a really, really fast model. And the part that probably most people care about who are watching this is, well, who's it available for and when will it be available? GPT5 is available to all users. So, whether you're on a free plan or a paid plan, you're getting it. Plus, subscribers get a little bit more usage and pro subscribers get access to GPT5 Pro, a version with extended reasoning for even more comprehensive and accurate answers. And they also mentioned on the live stream that pro users basically get unlimited access to it.

They actually touched on memory a little bit as well. And now they're giving it access to both your Gmail and your Google calendar so that it can be more of like an assistant where it can look at your calendar, look at your emails, and help sort of take actions and schedule things on your behalf. Next week, starting with pro users, followed by plus team and enterprise users. This is changing and we're giving chatbt access to Gmail and Google calendar. She then went on to show an example where it looked at her calendar and gave her a, you know, half hour by half hour day at a glance of what she had to do. Also looked at her emails to see if there was anything that she needed to take action on and even helped her pack for, you know, this flight that she had coming up that it noticed on her calendar. But again, as of the GPT5 launch date, August 7th, this isn't out yet. They said it's rolling out next week. So, this is not a feature that's going to be rolled into your chat GPT like immediately.

They also talked about some updates to the voice feature and now it's going to be available to everybody. Even free users can use the voice chat as much as they want. We are bringing our best voice experience to everyone. Free users can now chat for hours while paid subscribers can have nearly unlimited access and voice is also available in custom GPT. Plus, subscribers now can custom tailor the voice experience exactly to their need. It will follow your instruction closely. They went on to say that the voice feature is a lot more customizable now, so you can tell it to just answer in one word, and every single question you ask, it will just give you a one-word reply.

They also showed off that it was really, really good at translation. And they did a demo where they went back and forth, and it was actually helping him learn a new language. "Hey Chad, I'm learning Korean. Could you help me practicing it? Let's say, um, let's pretend I'm ordering at a cafe. Now, what should I say in Korean?" "Absolutely. I'd be happy to help you practice. So, if you're at a cafe and you want to keep it simple, you could start with something like, which means, 'Hello, I'd like one Americano, please.' And of course, you can adjust it based on what you want to order. Let me know if you want to try out more phrases."

They also mentioned that the personality is going to be more customizable. It's going to start in the text version, but eventually these personalities are going to roll out on the audio as well. We can see here from their blog post the personalities available initially for text chat and coming later to voice. So you can have concise and professional, thoughtful and supportive or a bit sarcastic. The four initial options, cynic, robot, listener, and nerd. And you do have to opt in and you can adjust them anytime in settings. And they also mention that they're all designed to reduce syphy. So, they're not necessarily going to just be super agreeable all of the time.

They also showed off a feature that personally I don't think was really worth showing in the keynote cuz like kind of who cares, but you can now change the colors inside of chat GPT and customize them how you want. We got to see a handful of demos during this keynote to show off what it's capable of. None of these demos honestly really blew me away. There was one that I thought was really cool, but most of them felt like stuff we've been able to do before. I guess what's impressive about these is that I think they were all done with just one shot, one prompt where we've been able to get very similar results using sort of multiple prompts. Now these models are capable of doing it with just like one shot.

One of the first examples they gave was getting a breakdown of the Berni effect and it explained it all. But then she asked, "Give me an SVG animation that better explains this concept," and this was the result they got from it. Created this interactive and engaging demo that I can actually play with. So I can actually change the air speed here to see how the lift and the pressure change accordingly. I can also tweak the angle of attack to see if my plane will actually fly or crash.

Another thing that they demoed was that when you're trying to code something, you can tell it to think harder or think quickly. Basically instruct it in your prompt how much you want it to think and the harder you tell it to think, probably the more detailed the output's going to be. But they also mentioned something pretty cool about the vibe coding capabilities was that you can go and open multiple tabs of chat GPT, give it the exact same prompt of the thing you want it to code in multiple windows and it will go and code three or four variations of it and you just pick the variation that came out the best for you. Like they vibe coded this French mouse game which looks a lot like snake but with a mouse head trying to chase cheese but other than that it looks exactly like snake but it also made another variation here in a separate window and they did have it code a third variation but they didn't actually show the third variation. They also showed it making this really impressive financial dashboard that was really well-designed had great colors when he hovers over the chart you can actually see that it's got the numbers in real time as he moves the mouse around, you can actually see the expenses in and revenue. A very beautiful looking dashboard all done with one prompt. And it took a couple minutes to generate, but the output was really, really good.

But probably the most impressive demo from the entire live stream was this castle game that they made that is a completely 3D game. You can see the little like guards walking around on the castle. They can rotate in 3D space and get any angle on it that you want. And I mean, it's got a lot of fine little details in there. There's a chat. You can see down in the bottom right, you've got the various characters and you can send chat messages and have conversations with the characters in the castle. And there's actually a little gameplay element to it where there's these little balloons that fly around and you can actually fire cannons and try to take down the balloons. For the game, they actually handed it over to Greg Brockman to try to play the game. And you can see that he's trying to pop the balloons as they fly around with the cannons from the castle. I mean, it's a pretty impressive looking game. I'm assuming they used 3JS to code it, but they didn't specifically specify in the demo.

They also talked a bit about safety in this live stream and how it's designed to minimize hallucinations. We can see this chart that they created. I'm guessing lower is better on this chart because pink is GPT5 with thinking. Open AAI is 03. And so yeah, that's the hallucination rate. So 0.7% versus 4.5%. So they're getting the amount of hallucinations down by quite a bit. They also mentioned that they're doing their best to minimize deception. And to me, it's just crazy that we even need to think about the fact that these AI models might try to deceive people. But they're working hard to minimize the deception rate on these models. You can see in coding deception, GPT5 16.5% versus 47.4% in missing image 9.9% versus 86.7%. I don't totally understand what these benchmarks mean other than they're trying to get the models to deceive you less which, yeah, that seems like a pretty important thing.

They also mentioned that the model is going to give you a response less often that sounds like, "Sorry, I can't help you with that." So, if you ask it to give you instructions on how to make fireworks, I believe that was one of their examples. Instead of saying, "I'm sorry, I can't help you with that," it might point you to ethical legal resources to get help with that. And it also seems to have a little bit better understanding of intent. So, it could kind of tell if your prompt was intended for malicious use versus an actual helpful use case, which seems interesting. I don't that seems like jailbreaking opportunities for me to tweak with intent to get what you want out of it, but I'm no expert on that so I don't totally know how it works.

When it comes to the API, GPT5 is available in OpenAI's API starting immediately. And it comes in three flavors. You've got GPT5, the, you know, highest-end model. You've got GPT5 mini, which is like, I guess the middle ground, and GPT5 Nano, which is the smaller, likely faster but not quite as smart model. And then the pricing to use the APIs reflects, you know, how fast and how smart the model is. And they seem pretty on par with most other APIs. So it'll probably be the new standard for most people using OpenAI APIs.

For those interested in the API, there is a new variable called reasoning effort. So you can tell it how long or how little to reason. So if you need a very, very fast response for whatever the app is you're building, you can have a very low reasoning effort. And if you want it to be smarter and give a more thought-out response, you can give it a higher reasoning effort. They're also adding a new verbosity parameter with a low, medium, high value. So you can actually tell it whether you want the output to have a lot of words or a little bit of words. The context window boosted to 400,000 tokens, which is roughly 300,000 words of input and output.

They also mentioned that they're giving free access in Cursor for a limited window. I think they said like a week you can get in and use GPT5 inside of Cursor and Cursor is going to default to GPT5 now when you log in because the founder of Cursor actually came on the live stream and said that this is the best coding model on the market right now.

And that's pretty much the recap of everything they covered on this GPT5 live stream. Now, they do have some blog posts here that have some additional demos that you can play with. They've got this pixel art tool that you can play with and, you know, draw pictures. It looks like a sort of Microsoft Paint, some sort of typing game, a drum simulator. They also have some examples where you can compare a GPT40 response to GPT5 cuz they claim that this is much better at writing than the previous model. So, you can actually see the difference here. So definitely check out the blog post over on OpenAI if you want to see some more demos and if you really want to dive deeper into like the benchmarks and things like that. They also have a separate blog post here on GPT5 for developers. So if you are a developer that wants to use the API, there's a lot more information here. It works in cursor windsurf forcell jet brains factory lovable gitlab augmented code github cognition. So it's pretty much being integrated into all of the vibe coding tools that are out there. And yeah, just lots of examples and things to look at on their website right now if you want to dive deeper into what it's capable of.

All right, so I have access to chat GPT5 here. And you will notice that up at the top they've removed all the other models. So you have GPT5, GPT5 thinking, and GPT5 Pro, which I believe these might only be available on the pro mode, but they got rid of all the other models. No more selecting which model you want. And there's one prompt I really, really want to test. I've been testing it with all the other models I've been playing with lately. And that's, "Make a Vampire Survivors clone. Make it beautiful and functional. Create it so that I can test it right inside of Canvas." Let's go ahead and submit that. It's telling me it's thinking longer for a better answer. Planning game development, coding game with React. Thought for 13 seconds. Now, it's starting to write the code.

All right. So, it spent about 3 minutes or so writing all this code. We can see it wrote about 565 lines of code here for this game. Let's go ahead and try it out. Controls are WD and arrows. Aim with the mouse. And let's go ahead and run it. Okay. So, we get a page that says neon survivors. Okay. Okay. This is really actually the best quick single prompt vampire survivors clone I've seen. It didn't use any images. It just used kind of shapes, but I mean, it's clean. It looks really good. The enemies are coming at a more rapid rate. I want to see what happens if I level up here. We're actually getting different enemy styles. We're getting bigger red dots and yellow dots with eyes on it. So, as it gets deeper and deeper into the game, we're getting different types of enemies. And yeah, they're starting to swarm. This is definitely the best one-shot vampire survivor clone I've created so far with AI. This is absolutely blowing my mind right now. It actually looks decent, too. Even though it's just circles, the way they put that like subtle gradient on the edges of the circles, and you've got the grid in the background with like the sort of purple colors. It all looks really, really solid. Yay, I finally leveled up. All right, let's try our orbit blades and see if it actually adds a new. Oh my gosh, look at that. It actually added a new weapon. This is all from one prompt. Now I'm leveling up faster cuz I got more weapons. Let's add a neon pulse here. Oh my gosh, this is really good. I mean, the swarm itself is sort of all merged into just like one, but you can see all the level ups actually work and it looks clean. I'm I'm just really impressed with this. So, as far as coding goes, GPD5 is from my final thoughts on GPT5 and this whole demo that they gave, and I do have a few thoughts on it. I'm personally a bit underwhelmed by the live stream, but that might just be because the leap from GPT3.5 to GPT4 felt so massive. If you remember the GPT4 demo, they were doing things where they were drawing out websites on paper, taking a picture of it, and then that website was just being built with code. Things we just never seen before and didn't even realize large language models were going to be capable of doing that quickly. This to me felt like, you know, pretty good improvements, but not that massive leap that we saw from 3.5 to 4. It felt more like the leap we saw from 4 to like the 03 models. But I might just have really high standards cuz I am somebody that is watching AI every day. I'm seeing the marginal improvements happen on a day-by-day basis. For the majority of the world, this may feel like a really, really big leap. I just feel like I was expecting a little bit more.

I was expecting maybe additional multimodality functionality like an improvement maybe in the image generator that we got in 4. Maybe Sora being directly built inside of GPT5. So now you can tell it to prompt images and prompt videos and have it explain a complex concept and then once it explains the complex concept say turn this into a video to explain it for me or something like that. I was kind of expecting like this mind-blowing moment of like, oh my gosh, we've never seen anything like this before. But I didn't really feel like we got anything like that from this live stream. It was more like, cool, it got a lot smarter. It got a lot faster, can think through things, and it can do more things with just one prompt. But like we didn't get that, hey, let's walk around the room with our iPhone with the camera open and see all this cool stuff that it can now do that it didn't used to do. Like we didn't get any of those mind-blowing moments from this live stream.

They showed off their agents, you know, last month and that was pretty cool. I was expecting maybe GPT5 to be shown off in collaboration with the agents and showing that the agents got way, way better. And I think that's coming. That'll probably be an upcoming announcement that GPT5 is now what's being used in these agents and the agents are a lot smarter. But I thought maybe that would come with the GPT5 announcement and it didn't.

I also wish they would update its ability to sound more like source material, right? Like one thing I like to do with Chad GPT is feed it content that I've written or feed it transcripts for my YouTube videos and say things like, "Help me write a script that sounds like this." And inevitably, the script never actually sounds like my voice. You know, it likes to put the word boom in there. Like you do this and then boom, this happens, right? There's always this like chat GPT feel to the writing and it never seems to really mimic my style of writing. Or if I wanted to pull in a video transcript from Mr. Beast and say, "Write in the style of this video," and get like a new script that sounds like a Mr. Beast video or something, and it never really manages to mimic the style of writing or voiceover or things like that from other videos. And I would love to see that come to life because I think that would make it so much more useful for people that do a lot of writing and things like that.

I was also excited when they mentioned that there's some new memory features, but the memory features were adding better integration with Gmail and Google Calendar and not actually like a lot better memory. I would love to have seen the memory get even better and last even longer and remember conversations from like a year ago and things like that. I'd also love to see memory being individual to projects. So, you've got the projects folder inside of chatgpt. I would love it if when I'm in one of these projects here and I ask it a question, it just uses the context of my previous chats within this project for the memory instead of the memory from my entire chat GPT account. Like, why can't we just get a memory inside of a project that just remembers what's going on in that project by itself? Or how come we don't have agents yet inside of a project? Or we can't connect this to custom tools like Google Drive or Calendar or Gmail or anything like that within a project. We can only do it from like a main new chat out here. These seem like little upgrades that we should be getting in chat GPT, but we haven't got yet.

But what it really seems to me that's going on is that the large language model community, the companies that are building these large language models, the Googles of the world, the Anthropics, and yes, the Open AIs of the world, they realized that coding sort of little bespoke apps that help one person solve a problem they're trying to overcome in their business is like the killer use case for large language models. And they're all just leaning into that. Like Anthropic leaned into that a while ago when the success of Sonnet happened and everybody started using that for coding. Anthropic seemingly went, "Let's just focus purely on our API and coding," and that is the main use case that Anthropic is going to have. Like Claude doesn't even get a lot of use on their website. Most of Claude's use is through the API from coding tools like Cursor and Windsurf. And OpenAI seems like they're realizing this as well. If we want to get the most possible mass adoption for our tools, we need to be the best at coding and let's create models that are amazing at writing code. Cuz if it can write code, it could kind of do anything. If it runs into a problem that it can't figure out how to solve, it can write code to then figure out how to solve that problem. But it's not as useful yet for like the everyday person who just wants help offloading email or managing their calendar or scheduling or things like that. But it's getting closer and closer because soon it's going to be able to write the scripts that do help you with that thing. So, it's like a nice stepping stone in that direction.

But overall, the live stream was interesting. I'm excited to get a new model. I'm excited to play with it, but I guess I was just hoping for more with all the hype we've been getting about GPT5. I thought we would see more, but again, hugely impressive. Still really, really exciting to get new models. I just don't think this new model is going to change the game yet for a lot of people outside of those people that are using it to write code. And yeah, that's my thoughts on the GPT5 announcement. I pretty much covered everything that they announced. It was an hour and a half live stream and hopefully this helped you speedrun all of the big announcements from the event and some of the demos that they showed off in the live stream and now you feel more looped in on what this whole GPT5 launch was all about.

If you like videos like this, make sure you like this one. Subscribe to this channel. I'll make sure more videos like this show up in your feed. Every Saturday, I make a news video where I break down all of the AI news that I think most people are going to care about. I also demo a lot of the tools if they're available to demo so you can see what they're capable of. And that video comes out every Saturday. On this Saturday's video, we'll actually have an opportunity to spend some time with GPT5, probably do a few tests with it. And yeah, again, like and subscribe if you want to see stuff like that. I really appreciate you hanging out and nerding out with me. We're entering a new era and things seem to be accelerating very quick. It was an insane week this week and hopefully I'll see you in the next video. Thanks again. Appreciate you. Bye-bye. Minimize hallucinizations. Hallucinations. Hallucinations.

Thank you so much for nerding out with me today. If you like videos like this, make sure to give it a thumbs up and subscribe to this channel. I'll make sure more videos like this show up in your YouTube feed. And if you haven't already, check out futuretools.io where I share all the coolest AI tools and all the latest AI news. And there's an awesome free newsletter. Thanks again. Really appreciate you. See you in the next one.