Transcription
The release of GPT-5 seems like it might be eminent, particularly if you go by rumors on the internet. Now, there have been some people from OpenAI, uh, on Twitter basically giving their usual cryptic hints. Uh, one of them was one of the, uh, a OpenAI researcher who said something to the effect of "it's wild watching people use chat GPT knowing what's coming." Um, so it seems like the release of GPT-5 is imminent.
So, let's look at all the facts, figures, and rumors, uh, that are out there. Now, one thing that I will say is that the content on the slide decks are consensus opinions. So, I've done a bunch of research, scraped it all together to try and figure out where the middle of the road is, but I will editorialize over that. So, yes, some of the facts and figures on here are relatively conservative. Some of them might seem a little bit outlandish, and I will make sure to make it clear when it's my opinion versus anyone else's.
So, moving on. Um, here's what we're going to cover today: Number one, release date; capabilities; modality; parameter count; agentness; benchmarks; training size; and then finally, timelines. Now, the timelines are where I disagree most with the rumors, but let's get right into it.
So, first and foremost, release date. July 2025 seems to be where people have settled. Uh, July or August maybe at the latest. Uh, at the same time, there are some people pushing back saying that we might, uh, not see it until December due to training reasons and the original cadence of around 33 months between major model releases. With that being said, earlier this year in February, Sam Altman said that GPT-5 would follow 4.5 in months, not weeks. And that's, you know, so months is like, okay, anywhere from, you know, 1 month, 2 months to 12 months. So maybe February 2026 at the absolute latest. Uh, but at the same time, the rumor mill is churning. It really seems like something is coming down the pipeline. Uh, so with that being said, you know, your mileage may vary. Uh, however, uh, you know, this is this is basically how OpenAI has done it, uh, recently, which is suddenly, you know, nothing nothing happens and then everything happens all at once. It never rains but it pours.
Moving on, raw capabilities and benchmarks saturated. We'll talk we'll take a little bit deeper dive into, uh, benchmarks in a later slide, but the overall, uh, take basically is that this is not incremental improvements; that we are looking at another paradigm shift in terms of capability, uh, beyond just human-level reasoning and reliability. So, uh, enhanced reasoning, obviously, they've been practicing with reasoning. 01 Pro, 03 Pro is out now. Even GPT-40 and 4.5 have been spotted reasoning in the wild. So, it looks like that OpenAI is bolting reasoning into everything, uh, which is a really interesting trend. Now, make of that what you will. Some of us out here have thought that maybe reasoning is going to go away, that it's just going to be embedded into latent space. But right now, it seems like they're actually doubling down on reasoning. So maybe all models will be reasoning models to a greater or lesser, uh, extent.
Now, one thing that's interesting though is some of the reasoning traces seem to be getting shorter, meaning maybe they're getting better at reasoning. Again, that is just my opinion. That is my observation. I don't know if anyone out there agrees with me, but that's just my gut instinct. Then reliability. So, uh, hallucinations have jumped, uh, markedly with models like 03 with hallucination, uh, rates jumping up into the 30% or more, uh, which is particularly problematic. Uh, with 03, I think 03 Pro hallucinates less because it spends a lot more time thinking and researching, but 03 vanilla will just make stuff up. Um, I was asking it to research some of the stuff that I had been working on with post-labor economics and it said that I had been like cited by Time magazine, and then I said like, "Can you provide a link to that?" and it just said, "Yeah, that link doesn't exist." I'm like, "Well, thanks." Um, so yeah, it still has it definitely has some hallucination problems. Uh, but we're expecting that to come back down. Um, as far as I know, I haven't followed the science closely, but we don't actually know why reasoning, uh, models tend to hallucinate more. It could be because it spends more time in self-referential thought. Who knows? But it's probably a solvable problem.
Next up is coding. Uh, so Sam has said at many, many, uh, public appearances over the last 6 months or so that, you know, their internal coding tools are amongst the best in the world, uh, and they're only getting better, uh, and I have I have seen some more tweets and leaks out there basically saying that that, uh, more and more anyone at these frontier labs uses their own models more and more for coding. Uh, so when you when you see tech people using their eating their own dog food, uh, that's basically a terminology we say like if your tools are so good you should be using your own tools. Well, they have very clearly passed that threshold at, uh, at OpenAI, at Microsoft, at Google. Um, all of the top labs are consuming their own tools first and foremost. Um, and honestly, when you're building any tech product, that's how you know that you've built something good is if it's something that you yourself can't live without.
Now, so we've we've certainly passed that stage and we think that there's probably a flywheel effect happening where it's already good enough at coding and every incremental gain in coding makes it that much better and more powerful. Uh, math is another big thing that, uh, is expected to be, uh, improved and we'll go over the benchmarks in just a moment. But then context, so, uh, people have talked about things like tribal knowledge, nuance knowledge, contextual knowledge as one of the biggest gaps in, uh, in agents. Now OpenAI has released several layers of memory. So there's project memory, there's scratchpad, um, there's, you know, uh, memory, uh, sorry memory across chats and those sorts of things. So memory is still pretty early, uh, in terms of in terms of, you know, how mature it's going to get, but we do expect GPT-5 to have substantially better memory management than the current models and platforms.
Now, many people are saying that it's very clear that AGI is not going to be a single model. It's going to be a platform. It's going to be an architecture of models and tools and and everything else, which is what I've been saying for literally the past four or five years with cognitive architectures. No one calls it a cognitive architecture now. They just call it an architecture.
So moving on, next is all-in-one modality. This is something that, uh, we have seen coming. Uh, it shouldn't be a surprise since, you know, uh, image generation, uh, went native into GPT-40. Uh, so when you add everything else that's coming down the pipeline, uh, you know, native voice generation, uh, native, uh, two-way streaming of audio, uh, with the with the latest advanced voice modes. So that is obviously going to be baked in. That's going to be a standard feature moving forward. The next thing though is audio and video. So will GPT-5 have the ability for real-time video streaming? I don't know if we're going to if we're going to see that yet. However, we should expect GPT-5 to have at the very least real-time audio streaming two ways, uh, high high-quality high-fidelity image processing and very likely video generation and video understanding if not video streaming. I doubt we'll see video streaming yet, uh, but that's probably right around the corner if I had to guess. Maybe GPT 5.5 or something like that. I would not be surprised if we see video streaming either by the end of this year or sometime in 2026. But basically, this is the direction that that, uh, Transformers have been going since what, uh, I think late 2023 when Nvidia announced their everything to everything. So, you might say like, "Well, this looks like everything." This is not everything. Three-dimensional data, visiatial data, um, uh, articulation data. So, for robots, in the long run, I suspect we're going to have an all-in-one droid brain that's going to be able to understand literally any format and output any format. As long as it's digital bits and bytes, it will be able to understand it. Basically, anything you can tokenize is where we're going. So, when we say all-in-one modalities, we're about actually halfway to all-in-one modalities because there's plenty of other file formats that it's not working with yet. However, what we should expect to see is GPT-5 will be a substantive leap in that direction.
All right, next up, parameter count speculation. Right now, GPT-4 is thought to be around 1 to 1.5 trillion parameters. Of course, GPT-4 is a family of models. Some of them have been distilled, some have been quantized, some of them we're not even sure, uh, but the raw GPT-5 people are expecting could be around one quadrillion parameters, a jump of a thousandx. Now this is not something that we have, uh, not seen before. We have seen jumps this large, um, but we've also been surprised in the past where sometimes they're able to squeeze a lot more performance out of far fewer parameters. So I'm kind of skeptical of this, um, and again, you know, most ambitious predictions. So this is like the high end. Um, the lower end, I wouldn't be surprised if it's only like 10 to 50 trillion parameters, uh, give or take maybe 100 trillion. I think I think, uh, a thousand trillion or quadrillion parameters is probably a little bit ridiculous, but that's a number that's out there. So just wanted to clarify. Most conservative predictions place it between 5 to 50 trillion parameters. I agree with that. That seems pretty reasonable. OpenAI has surprised us before, but also as we have seen parameter counts are not everything. Um, I mean Sam Altman even said last year the the era of of parameter scaling is already over. Now we're now we're scaling on, uh, inference time compute or test time compute, um, and those sorts of things. Mixture of experts architecture that is probably expected to to stick around for the foreseeable future. There might be new architectural innovations. I haven't heard any rumors of these, uh, of architectural innovations yet. Right now, it seems mostly can you tokenize everything and can you stream those tokens is the big thing. Uh, and then of course tool use and agentic behavior which we'll talk about in just a second.
So fully autonomous agents, OpenAI and others have clearly been experimenting with agents. So everything from operator to Codex and those sorts of things, which by the way if you haven't used Codex it's pretty good, um, if you give it a single, you know, coding task where you say, "Hey, you know, clean up this repository or fix this code or write me a script that does this and submit a pull request," it's generally pretty good; it does still have some really, really glaring hang-ups, uh, so for instance, um, and one time I was using it created some binary objects, uh, which it's not allowed to submit as a PR and I said, "You have you have to delete those binary objects and remove them from the pull request," and it kept saying I did that and it didn't, so I had to just like nuke that that whole like thing and start over. Um, so it's not perfect, but for a generation one tool Codex is pretty amazing. Uh, so what we're when we when we take a step back and we look at like, okay, what do what do we expect? We're basically expecting Codex plus, uh, performance on any, you know, any task, uh, so workflow management, API integration, real-world actions, um, Microsoft. So it's interesting that that people are expecting the integration with Microsoft to continue, uh, when you look at the most recent rumors and that is that the relationship between OpenAI and Microsoft is continuing to deteriorate, um, they're having disagreements over scaling, over cloud compute and those sorts of things, uh, at the same time, uh, you know, there is still a little bit of sympatico there in terms of, uh, what they what they gain from each other. However, getting a little bit off topic, when you look at the arc of of, uh, of agentic behavior. So, Deep Research was kind of the first fully agentic, uh, uh, tool that they released, um, and it's a search tool. So, that that kind of makes sense because you say, you know, "I want to I want to learn everything that there is to learn about X, Y, and Z." Cool. So then it'll go search and, you know, it'll it'll it has a really good pedagogy in terms of figuring out what to look for, uh, and then it can the next step is compress it into a report. Then Codex takes that another step further and now of course computer-using agents are the next big thing. Now computer-using agents haven't really taken off yet. TLDDR: I suspect GPT-5 will be the premier engine for computer-using agents in 2025. And this will be kind of the aha moment for a lot of people because everyone might start switching to say, okay, you know, computer use was kind of, you know, this derpy little thing that could do a couple things, but, you know, would frequently delete my entire codebase and those sorts of things. I think that we're going to see a paradigm shift in terms of computer-using agents because that really is the beating heart of the next epoch of AI being deployed; a, uh, KVM, so computer, uh, or sorry keyboard, video and mouse, that is the primary API to get to literally everything that people want to do on the internet; if you can do it with a keyboard, video and mouse, and and you have an AI agent that is a that is a master at keyboard video and mouse, it can do everything from playing video games to editing YouTube videos to writing books for you. Um, this the the the total addressable market of of computer-using agents cannot be calculated. I actually had someone reach out to me saying, "Hey, can you help me figure out what the TAM is for computer-using agents?" I said, "It's not possible to calculate; like how many computers are there on the in the world? How many people work in front of a laptop or a desktop? Like you cannot calculate it." And then when you put them into into the cloud into virtual space, what we're looking at is billions of computer-using agents within, uh, coming online within the next couple years.
So next is benchmarks. So, you know, obviously benchmarks are one thing. Benchmarks often do not represent exactly how useful or user-friendly a model is, but it is still useful as a target, particularly as the number of benchmarks proliferate. Uh, one of the primary things is MMLU, expected to saturate at about 95% accurate. Uh, SWEBench is expected to jump up to 85% from 32%, which would basically mean OpenAI knows how to solve all of coding and then it's just another another couple of iterations to get, uh, to the next level. Um, I'm personally skeptical of a jump this big, but we have seen jumps this big, um, recently in terms of from one epoch to the next or one generation of model to the next. Um, so it's not outside the realm of possibility, but this will a jump of this size will make a lot of people sit up and take notice. Um, so don't say we didn't warn you. Uh, factual reliability, so hallucination rates hopefully will drop back down to less than 15%, um, you know, wish in one hand, hopefully it gets that bad, uh, math, so the math the advanced math benchmark, uh, expected to improve to 40 to 50% accuracy; that would be great, uh, if we get that high, uh, obviously math is kind of like the big frontier, uh, because it requires kinds of abstract reasoning that, uh, even coding doesn't require, um, so I personally suspect that fully saturating all math benchmarks might be one of the very last things that happens. Um, we'll see. Uh, and then multimodal tasks. So, this is basically the success rate on larger and more complex task tasks. Visual and cross-modal benchmarks are expected to see 90%, uh, as unified architecture integrates all modalities.
One oversight that I didn't have on this slide deck is context window, um, and I didn't include it because some of the some of the guesstimates out there are, you know, "it's going to have a context window of 64,000 to 256,000 tokens." And I'm like, "You guys must be drunk." Like, it could easily be 1 to 2 million tokens because Google's already figured out gigantic context windows and others have as well. So, basically, I would be disappointed if it was anything less than a quarter million tokens, um, and I would not be surprised if it's 1 to 2 million tokens. So, I left that out just because like I don't really agree with the consensus and it seemed like a really silly consensus.
So, moving on. Um, training and post-training safety. This is something that, uh, that has popped up. I'm not going to spend a whole lot of time on this slide, uh, but basically, you know, it's using the latest and greatest, uh, uh, clusters from Nvidia. No duh. Uh, safety first, you know, one of So, this is this is probably the one part that's worth remarking on is that as these agents and models get more and more powerful, you might expect them to spend longer and longer on safety, uh, testing. However, I suspect that OpenAI has, uh, fully automated their their safety testing suite, um, and that that actually saves them a tremendous amount of time and that they've gotten really good at best practices for model safety. And you know, it's it's kind of like garbage in garbage out. Uh, if your pre-training and post-training, uh, steps are really good then you know you just check all the boxes to make sure that it, you know, you you it's not, you know, bio-risk and and, uh, you know, power-seeking and all those other things. If you train a model really great and you get you nail it on the first time and you you don't need to go back to the drawing board, you can actually save a tremendous amount of time. I mean, anyone who works in machine learning this has been true forever; you know, the quality of your data, the quality of your your training parameters and everything else that, you know, that that's where that's that's the heart of what you're doing. That's the hardest part. Hardest part number one is getting good data. Hardest part number two is, you know, feature extraction or reinforcement learning paradigms, uh, you know, RL RL specifications. You nail both of those. The rest is pretty easy. So, that's where I'll leave it on training and post-training safety.
Uh, next up is timelines. So, or almost finally timelines. Um, this is what I what I will say is this is the general consensus out there coming out of institutes like Brookings and McKenzie and MIT. I think that this is way way way too conservative. And I know a lot of you out there will agree with me. And when you look at the margin of error that every AI prediction has had, everyone has been too conservative. So, with that, uh, with that, uh, that qualifier already out of the way, 2025, the rest of this year, um, this is where we're expecting the dawn of agent deployment. Obviously, this is what I've been saying since the beginning of the year. 2025 is the year of the agent. That has borne out to be completely true. 2026 infrastructure scaling. So, agent SDKs rather than building your own agents, they're just going to come right out of the box. Um, obviously there are some some packages that are available out of the box, but it's all going to be API-driven by next year, um, with a billion deployed agents across enterprise workflows, market grows to 10 to 20 billion. Um, I would expect that the revenue, the potential revenue to be a little bit higher if you have a billion agents. But this is what I call the automation cliff. This is what we're heading towards with the automation cliff, which is basically once you have like, you know, there's there's a threshold, right? So there's a threshold for network effects and right now agents are below that threshold, and what that threshold represents is, uh, a battery of capabilities and reliability and value add and product market fit and price market fit and so on and so forth. Once you get above that threshold, you get to a tipping point where a a new product or service or tool goes from a novelty to indispensable. And that is basically it's almost binary, right? where it's like, "Hey, you know, we're we're we're in the, you know, it's a fun toy. It's fun to play with right now, but 2026 rolls around, as the integrations get better, I suspect it'll get to a point where it's like actually this product sells itself." And I I do suspect that OpenAI will be leading that charge. So 2026, expect that to be the year that, uh, agents proliferate. That I do agree with. Uh, 2027, autonomous ecosystems. Um, so next-generation agents, so agent Mark 2, uh, release capabilities, multi-agent collaboration and potential AGI breakthroughs, uh, approaching $20 billion market. This is where I think this is actually too conservative.
I would not be surprised if we see this by the end of 2026. There are already plenty of people working on multi-agent frameworks. There are plenty of people working on open-source hobby multi-agent frameworks. In fact, multi-agent frameworks seem like the best way to use agents. Uh, not OpenAI, Anthropic just released a paper where they gave their uh their their Claude agent a bunch of agents to work with. Um, so I mean we might even see multi-agent platforms be mainstream by the end of this year. Like, like I said, I think that this is way too conservative because again, once you solve that problem, once you have an agent that is really highly autonomous, one of the tasks you can give that highly autonomous agent is manage these other agents. It's you, it, it's it's not hitting two birds with one stone. It's like hitting a million birds with one stone.
Moving on. Uh, hybrid teams, 28% of managers hire AI workforce coordinators as agents automate 30% of knowledge work, nearing a $30 billion market. This is happening in 2027. I'm sorry, like 2026, 2027 at the latest. Uh, there's already hybrid teams out there. That's what Julia and I are doing with First Movers AI; we're building hybrid teams. That's happening this year for First Movers. So this is way, way too conservative at this point.
2029 guardian protocols. So oversight AI emerges to manage agent networks as regulation struggles to keep pace, reaching a $40 billion market. Again, you know, this is where the consensus is, is out there. This is coming 2026, 2027. And in fact, I know other agent platforms that are already building in supervisor agents. This is a default choice when you have multi-agent frameworks. Um, my friend who's the CEO and founder of Wayfound AI, they're already doing this this year. So the the the the cons, the um the overly conservative timeline forecasts out there are just spectacularly wrong when you're actually plugged into what's actually happening.
And then by 2030, uh, the consensus out there is that agents become the standard layer for digital work, automating 80% of workflows in a 47 to 70 billion market. Um, by 2030 it's going to be super intelligence and it's going to be like 95 to 98% of digital work. Uh, so that's my take.
Anyways, uh, as we wrap up, let's just take a quick look at OpenAI's stages of artificial intelligence. Right now we are very clearly at reasoners. Uh, OpenAI's 03 pro, which I use extensively for my research, is beyond me in terms of solving problems. Uh, I give it all kinds of problems to solve, and it goes and solves them relatively quickly. Um, and I'm able to then interrogate it and say, "How do you know? Help help me learn how you did that." Um, so it's making me smarter as we go as well. Uh, we are at the very cusp, the very beginning of level three agents. Um, I think that by the end of this year, uh, everyone will be convinced that we are very, very much at level three solidly. And I think that it won't just be OpenAI. I think it'll be OpenAI, Microsoft, Google, maybe Meta, you know, uh, Mark uh just created his crack ASI team uh or is working on building it. So, who knows? We might be on the cusp of level four by the end of this year. Certainly by 2026, uh, sometime during 2026, I think we we will be solidly in level four. Innovators, AI that can aid in invention, arguably they already can help. Uh, but if they can autonomously innovate on their own, choosing their own research objectives and those sorts of things, you might argue that that's that's kind of where we're at, and we're not there yet. Um, and then level five will come 20 end of 2026, 2027, 2028 at the absolute latest at the pace that we're going.
All right, with all that being said, I hope you got a lot out of this. This is what I expect. This is what the consensus out there is, and this is where I disagree with the consensus. Have a good one.