Transcription
The reception of GPT5 has been a mixed bag, to say the least. It garnered a lot of heat very quickly on outlets such as Twitter and Reddit.
Now, if you've been a fan of my channel for a long time, you know that I have been very critical of OpenAI at times, as well as pretty much every other AI company. Um, I'm a little bit of a Scrooge McDuck with that respect, as I'm always a little bit skeptical of new models. But I will say that with GPT5, I was rather impressed with its intelligence and other capabilities. But I want to be fair and show that there are plenty of people who were not happy with that.
So, to exemplify what's going on, this is one of the top posts on Reddit about GPT5, where they just said, "GPT5 is horrible. Short replies, obnoxious styling, less personality, yada yada yada." And then, of course, there have been a couple of updates. So, if you weren't aware of all this, the backlash was very swift.
One of the primary mistakes that OpenAI made was that they sunsetted every other model in ChatGPT at once and swapped them all to just GPT5. Now, anyone who's been in technology for any length of time knows that you do not do a rollout like that. You always give users a way to roll back. You make sure that it's phased, and so on. So, you know, they're still a startup technically. Um, they probably don't have the hard-won experience of serving business users for a long enough period of time to know that you have to let people keep using the legacy stuff, even if you don't want to support it anymore. Some people noted that I still run Windows 10. Why? I'm familiar with it. It works. If it ain't broke, don't fix it. I'm a good old southern boy like that.
So, with that being said, let's unpack some of the responses. This is a response from Sam Altman. I believe it was last night, um, where, you know, he kind of acknowledged like, yes, we kind of botched it, in not so few words. They underestimated how much people like GPT-4o. Now, this is really interesting because one of the things that they said during the GPT live stream, that Sam Altman even said personally, or maybe it was the interview with Cleo Abrams. Anyways, Sam Altman himself said that one of the things that was really touching to him about ChatGPT, particularly 4o, is that people rely on it for emotional support and those sorts of things. So, it seems like a really big oversight where they really doubled down on GPT5 for coding and scientific and research purposes, which I find it more than adequate for, and certainly smarter than Gemini and Grok are on those topics.
So, when he says, "Point one, we didn't realize people liked it that much," I don't really know if that's a defensible argument. Maybe they just were focusing on where most users are or where most money is coming from, and not realizing that they're actually serving many different market segments already. So, when you have 700 million daily users, over a billion users total, you're obviously serving different people with different needs. And when you have a general-purpose product like ChatGPT, you're going to have different niches, or niches depending on how you want to pronounce it, evolving. So, that just seems like, in hindsight, kind of like a, "Well, duh, what did you think was going to happen?" And of course, there are very different opinions, yada yada yada, so on and so forth.
Optimization. Some people are frustrated with usage limits, which, of course, every time OpenAI rolls out a new model, there are always usage limits and things going offline. You'd think that they would get better at launches by now, or would have learned the hard way to scale up slowly. I get that they do want to try and serve everyone as fast as possible, but let's dive into a little bit more of the data, because one thing is, on the internet, bad news travels fast, and negativity is echoed, whether it's rage bait, whether it's artificially enhanced or inflated rage bait, or genuine emotional distress, or that sort of thing, it all travels faster.
Now, obviously, take internet polls with a grain of salt, but on both my YouTube and on my Twitter, we see roughly similar proportions where, of the people that have used ChatGPT, or sorry, GPT5, the bad percentage are a very small minority. So, on YouTube, it's only 3%. Now, granted, you know, a third of people, or about half of the people that have used it, are nonplussed. They're meh, whatever. And that is the same over on Twitter as well, where only a very slim plurality of people, now granted, if you break it all down and remove the "not used, no opinion," a slim majority of people think that it's good. A middling number think that it's meh, and then a very small, not small minority, a small minority think that it's bad. So that's what the data shows.
Now, again, there's a lot of selection bias. You know, the selection first and foremost is people that follow me. Now, a lot of people know that I have been both a fan of and critical of OpenAI in the past. So, don't accuse me of being a shill or anything like that. They don't pay me any money. I don't get any special treatment. So, you know, they did some good, they did some bad, but I just wanted to share that this is what the data from my viewers objectively shows.
Now, obviously, if three to seven percent of your user base thinks that what you did was flat-out bad, and they prefer other products that you were offering, there is no reason to destroy the value of those products that you had that were working just fine for some of those users. Now, what does that mean? Maybe they can continue making 4o cheaper and faster. And so, then they just serve enough people that they say, "This product is good enough. It's what I want. It's what I need. I don't need anything better."
Now, there are power users like myself and plenty of other people that are always pushing the limits as to context size, intelligence, reasoning, planning, those kinds of things. But then again, not everyone is going to use it like I do or like other coders do. And one of the long-term impacts that I suspect this might have, I'm still not decided, is whether or not this will cause a bifurcation of the direction of frontier models. Right now, we're still seeing compounding benefits of putting everything into one model: math, reasoning, multimodality, that sort of stuff. So, I don't think that we're going to see a bifurcation in specialization yet, because generally speaking, what we're finding so far is the better a model gets at reasoning and math and writing, it gets better at all other domains. So, you have these spillover effects.
In the long run, however, it might make sense to have purpose-built models for very specific uses, you know, whether it's a bespoke model for just chat, for just companionship, or one that focuses just on text and doesn't have multimodal capabilities, those kinds of ideas, so that you can really optimize around efficiency, because remember, a lot of these AI companies are not profitable yet, so they're going to need to optimize their product to the point where it is cheap enough for them to run so they can operate at a margin, at a positive margin, rather than at a loss.
So, I wanted to close out by just showing what are the objective complaints that people are having. So, I used Perplexity Pro to just scrape together a bunch of sources. Now, I ran this a day or two ago, so there's probably some updated stuff, but we've got citations. So, the things that people are complaining about: short, robotic replies and reduced personality. You can change this with style prompts, but a lot of people probably don't know how to write style prompts or don't use the custom personalities. My chatbot has the exact personality that I want. I want it to talk to me like Commander Data and Spock. That is great. So that's what I wanted because I wanted a lot less fluff. The way that someone described it is, you want a very high insight-to-word ratio.
Next is reasoning and depth require prompt engineering. This has always been true, but I think what people were complaining about was the model routing. So, I'm not sure how the model routing works. I haven't looked into it. But basically, GPT5 will decide whether or not it needs to do any thinking. But I'm on the pro plan, which means that I can just say, "Use GPT5 thinking" or "Use GPT5 Pro." So, if I want it to really go overboard, I just select Pro. So that's not a problem that I had. But certainly, I could recognize that it would be very frustrating to people if you don't have the ability to explicitly choose, say, "I know what tool I want to use for this job." You don't get to pick. I pick the tool.
And that goes back to user agency, which is something that I've talked about for quite a while. But when, whether it's an image generator, a video generator, or a chatbot ignores your intent, and it knows what your intent is, but makes its own decision for business reasons or alignment reasons or ethical reasons, that's infuriating. So, if people felt like it was hijacking their intent, I can understand why people would be up in arms about that.
Next is no major upgrade, incremental changes only. This was my initial reaction as well, where, you know, it doesn't generate video. It doesn't generate images any better. It doesn't generate audio artifacts. It's basically just a smarter chatbot, which, to me, was very disappointing. So, agreed.
Next is bugs and problems with complex code generation. They say it's the best coding model so far, but some people are still complaining. I sometimes wonder if that is maybe inflated expectations, where the first thing that people do is hammer it with the hardest problems that they have. At the same time, it is the top-performing model on a lot of benchmarks, and they did show some one-shot coding challenges with no bugs live, although, as other people pointed out, you know, coding up a simple Python game in 400 lines of code, it's seen plenty of examples of that on GitHub and other places, so that's not necessarily particularly impressive in the grand scheme of things. I don't use, I don't do any coding anymore. I got out of coding and scripting a while ago. So, I don't have a personal opinion on that, but granted, you know, if Sam Altman bills it as, you know, "We're going to have software on demand," and then it still fails on basic scripts, then that means we're a little bit longer from software on demand.
So, granted, unpredictable and opaque thinking modes. I think this is kind of a duplicate of the other issue of reasoning and depth require prompt engineering, where it's just, you kind of don't know what it's doing. And this goes back to user choice. Going back to Meta and Facebook, one of the best things that Mark Zuckerberg did for software as a service, and particularly UX, is give users more choice and more power. One of the philosophies that they have for Facebook that made it successful is just put wherever you can put power in the hands of the users. Give them the ability to create events and groups and whatever else, and you just let them run wild. You give them a sandbox. When OpenAI removed user agency, that really rubs people wrong.
Increased restrictions and prompt limits. I already mentioned that, so I don't need to redo it. Fails on basic facts or current events. Some people have complained that its training cutoff data is still too old, which is interesting, particularly when Grok is being updated real-time, or at least that's what they say. But of course, Grok is also plugged directly into Twitter, so it's going to have the most up-to-date. And another thing that I noticed about Grok is that it is, by design, going to look for the most recent information possible on every topic. Now, for some kinds of work, that is going to be ideal. If you're a journalist or if you're trying to figure out what's going on with jobs reports and that sort of thing, Grok is a great tool for that. If GPT5 doesn't always know when to search, then that could be a problem.
Notable decline in agentic abilities. So, this is something that I kind of pointed out. This was a big disappointment to me as well from the live stream, where they said, you know, "It can do longer tasks autonomously," and that was it. They didn't talk about computer using agents or any agentic testing, nothing beyond that. So, it seems like maybe they underplayed agenticness because maybe, this is a reach on my part, I will admit, but maybe they're having a harder time making the agentic aspect better. Who knows?
Many people complain that it felt like cost savings, not advancement. It is a cheaper and faster model, and I will say that that increases UX quite a bit. One of the things that I don't like about Grok right now is that you can ask it a basic question, and it'll go search for way too darn long. Whereas I just go to Perplexity, and I'm like, "Ask this," and it uses their bespoke model, the Radar model, and it will start giving you answers in like three seconds. It's stupid fast, which is good. So, there, you know, that's not nothing, that if a model is fast enough and gets to the threshold of good enough, there is such a thing as throwing way too many resources at a particular query or problem.
So, with all that being said, I just wanted to share, like, yeah, it seems like the reaction to GPT5 was a mixed bag. Um, I still think it's a really powerful model, but that's because it suits my particular use case really well, which is frontier research on post-labor economics. GPT5 is by far the best model for that kind of work right now, hands down. But I understand that is just one opinion, and that is just one use case.
So, thanks for watching to the end. Have a good one. Cheers.