Transcription
Most people are using the wrong chat GPT model, and they don't even realize it. Open up the menu, and you're hit with GPT-40, 03, 03 Pro, 04 Mini, 04 Mini High, then even more models. It's confusing, but picking the right one can completely change your results. Faster answers, better outputs, and hidden features most users never touch. Whether you're just chatting or wiring the API into an app, let me show you which one to choose and when.
I'll make this simple because OpenAI didn't. Their naming conventions are just ridiculous. So, I renamed each model based on what it's best at, and I had ChadBT generate custom visuals to make them easier to remember. I will show the cost, speed, and best workflow for each one. We'll start with the core models most people use every day, then move into the ones you should switch to when needed. I'll be including power user tips for cost, API, and agents.
First up is GPT-40, the generalist. It's fast, and it's the default for most users, and honestly, it should be. 40 is great across the widest variety of use cases—like a Swiss Army knife. Need a quick summary of a blog post? 40 nails it. How about brainstorm 10 catchy YouTube titles about explaining the different chat models? Nailed it again. I'll probably use some version of one of these. Trying to describe what's happening in a photo? 40 actually walked me through my entire PC build just by taking photos along the way. It's conversational, quick, and surprisingly good at creative or open-ended tasks. This is my go-to assistant for probably 40 to 50% of my use.
If you're building something like a chatbot for customers, 40's fast replies make it a perfect model for that, but don't rely on it for accounting numbers or critical code. I'll use a simple, consistent test case across these first few models to demonstrate—explain the pros and cons of nuclear energy versus solar in detail with citations. 40 will give you a decent overview really fast, but it's surface level. It glosses over nuance. It may get details wrong, and once you push it, the cracks really start to show. It will sound super confident even when it's just hallucinating and making stuff up. For casual use, your daily driver for easy to medium tasks, 40 is the move. But when you need depth or precision, time to switch.
Let's run that same question—explain the pros and cons of nuclear energy versus solar in detail with citations. But this time, we're asking 03, the professor. If you're on the $200 per month plus plan, you will also see 03 Pro here now, which was just released. I'll cover that one a bit later since it will be used a lot less. But I will send this prompt with 03. Right away, you'll notice something different. It starts thinking, and you can actually watch its reasoning play out. It begins by considering the best way to approach the question and figures out the key points it needs to research further. Next, it searches for sources to support each one of those. The process feels almost human. At one point, it goes, "Oh, this is a solid source. That's a keeper." Pretty interesting to watch its thought process. Sometimes it discovers where it needs more information and finds new sources for that area. Finding sources for cost, then capacity, then carbon footprint. Then it refines its pros and cons, finds gaps, looks up new sources, and builds a full answer. In total, it thought for 1 minute and 41 seconds. And here's a cool part. I can open up this box and see every step it took. So, this is where 03 shines. So, this is real multi-step reasoning, not just an output.
And when you compare the results, the high-level structure might look similar to 40, but 03 goes miles deeper, backing up all its claims with actual sources. It ran its own calculations when needed, and it caught a lot of nuance that 40 completely missed. It's not just repeating general consensus. It's evaluating the question. And to test it further, I asked 03 to critique 40's answer. It flagged outdated cost numbers, oversimplified conclusions. It had a flat-out false claim. There was no mention of water usage, and there's missing context all over the place. And that's the danger with 40. It sounds smart, especially on topics you're not familiar with. But if you ask it about a subject you know deeply, you realize how often it misses nuance, relies on outdated info, or just makes stuff up. That happens much less with 03. Not zero, of course, but much less. Its reasoning quality is definitely on another level. And this matters obviously for like math or research-heavy prompts like combinatorial mathematics: 10 distinct people are seated around a circular table. A and B must sit next to each other. C and D must not. How many distinct seedings are possible considering rotations equivalent? Or philosophy: compare and critique the strongest arguments for and against pansychism. But it also helps with more like everyday logic problems like legal questions or business decisions or even something like a workout plan with a few added constraints that can benefit from a reasoning model too.
Now, sure, this output will take longer initially, but I often save time overall. With 40, I get fast drafts but end up in three to four follow-ups to get it right. 03 just nails it up front. I find that to be the case a lot, even for everyday questions. Not super simple, like I'll use 40 when my son asks me, "What types of food can we feed a snail he found?" But when I want to know all the bugs we can catch in Utah and keep in a terrarium that can all live harmoniously together while requiring minimal upkeep—yes, that's a real thing we're doing—and I save time by letting 03 think that through before the first answer. Then I just switch to 40 for the faster follow-ups. That tag-team workflow is something I use a lot.
Now, let's ask that same question one more time: Explain the pros and cons of nuclear energy with solar in detail with citations. This time, we'll open up the tools box, then select deep research mode, aka the scholar. I'll send that, and then it always comes back with a list of clarifying questions since it's about to go deep. Just answer those. And here's what happens. It disappears for a few minutes, usually 5 to 10 depending on how complex the question is. While it's gone, it's scouring the internet, studies, articles, public data. It analyzes what it finds, breaks down arguments, calculates trade-offs. Now, this works similarly to 03 because it uses 03 as its core reasoner. But deep research goes further. It does also pull in faster models for simple tasks—like it might call 04 mini to scrape tables. Then all the data gets handed to 03 for synthesis. And what you get back isn't just an answer. It's a mini literature review. You'll get a clearly structured breakdown, multiple perspectives, direct quotes and links to real sources, a conclusion that actually weighs trade-offs based on the evidence. Deep research is slow and capped, but when you want thorough, well-reasoned answers using extensive real-world sources, this is your tool. It's perfect for writing research-backed blog posts, prepping presentations or interviews, academic work, digging up real recent data from the web. Just think of this as your personal research assistant. This is not for fast answers. This is for big questions that need receipts.
Now, we're diving deep into all the different chatBT models, but when you're using each model, whether you're writing, coding, or building your own tools, understanding how to structure prompts is the most important next step to getting better results. I have a free resource provided by HubSpot linked below that teaches exactly that. It's called Advanced Chat GPT Prompt Engineering: From Basic to Expert in 7 Days. This guide isn't just a list of prompts. It's a full framework to transform how you think and work with AI. One of my favorite sections is on the ROSES framework. This shows you how to engineer prompts using role, objective, scenario, expected solution, and steps. It's incredibly useful for structured outputs. There's also a great deep dive into modular prompt systems. Basically, how to build prompt components you can use across projects. And that is something I've been experimenting with a lot. I use these types of techniques all the time. And if you've ever wanted to go from "I hope ChatGpt gets it right" to "I built a prompt system that delivers every time," this resource is exactly what you need. Download it for free using the link in the description. And thanks to HubSpot for sponsoring this video and providing resources like this to people who watch this channel.
All right, same question, three very different results. It's worth recapping these because they're by far the most useful across the most use cases, especially for casual users. GPT-40, the generalist, gave us a fast conversational overview. Great for quick, surface-level understanding, but it doesn't go deep, and it often hallucinates or reasons poorly. Not something you'd ever cite. 03, the professor, slowed down, structured the argument, and built a clear, logical comparison backed by real analysis. Deep research, the scholar, went full academic. It pulled citations, sourced real links, and wrote like a formal report. And one key point: 03 can search the web, too, especially if you ask it to. The difference is in how far it goes. 03 is often the better choice for reasoning because it's faster. It's easier to iterate with. Deep research is for when I want to really understand one thing. When I want chatBT to vanish for 10 minutes, dig through the internet, and bring me back the best thinking, data, and evidence available.
Now, let's talk about GPT 4.5, the wordsmith. This one's kind of a wild card. The 4.5 isn't the best at reasoning, coding, or research, but it shines in one specific area: tone. If you want something that flows naturally, something with rhythm, personality, or emotional weight, this is the model. The difference isn't huge, but it's noticeable. So, let's try this prompt: Create a persuasive product description for a futuristic smart pen using vivid sensory language. It's slower than 4, but still fast. As you grip its smooth ergonomic body, subtle vibrations gently confirm every input. All right, I said sensory, not sensual. But this part is pretty good. Whisper quiet, it harmonizes with your workflow, turning scribbled brainstorms into polished masterpieces. The Luminina Pro isn't just a pen; it's your bridge to the future. Feel your thoughts flow freely, captured vividly and effortlessly, and step boldly into a new era of creativity. That's solid copy, and that's where 4.5 especially shines: marketing, branding, and ad writing, but it works for other creative writing, too. So, let's try something different: Describe a quiet morning in a war-torn village, but make it peaceful and nostalgic. All right, this looks solid. A solitary bird trills softly from a gnarled olive tree, its song threading tenderly through empty doorways and windows draped in ivy. At the village square, an ancient stone fountain murmurs steadily, a comforting heartbeat in the silence. A gentle breeze rustles through cloth shutters, stirring curtains like ghostly hands reaching out to reclaim forgotten dreams. That's great writing, especially since I didn't really give it much direction. And 40 has come a long way in creative writing, too, but 4.5 just feels smoother. That said, don't ask it to do math, logic puzzles, or fact-heavy research. It's not the best at any of those. Use 4.5 when you want strong voice, tone, or emotion—you're writing something persuasive, descriptive, or stylized where you care more about flow than precision. Otherwise, you're better off with 40 or 03. Just one last thing: 4.5 is preview only and might disappear once 40 gets fully upgraded for tone. The 40 updates from May actually closed much of the gap already. So, this one may not stick around, but for now, it's your go-to ghost rider.
Before we jump into the next models, a quick note on where this video came from. It was inspired by a tweet from Andre Carpathy, one of the top minds in AI. He shared what he uses each chatBT model for, and there was a lot of debate and disagreements with his takes in the comments, and I didn't fully agree either. So I did a few things to help research this. I took his tweet. I extracted every comment from underneath and had those takes counted and categorized. I also had deep research analyze sources from all over the internet: Reddit, GitHub, blogs, everything. Then it compared all the data from those experiences too. The models I've covered so far were largely all agreed upon by everyone except maybe four or five. But for the next few models where there was disagreement, I ran fresh tests. So I do think this list is very accurate now. But if you've got a different take, drop it in the comments. I'd love to hear how you're using them.
So far, we've covered the core players: the generalist, the professor, and the scholar, plus the wordsmith, which is the biggest kind of random outlier. But now we're moving into some specialists, or what you switch to when you hit a wall. GPT 4.1, the coder. This is the model Carpathy listed as great for vibe coding. And he's right. 4.1 is excellent for coding and following detailed instructions. But its real superpower is context window. If you're using the chat GBT interface, it's capped at 32,000 tokens, same as 40 and 03. But through the API, 4.1 and 4.1 mini unlock up to a million tokens of context. That's massive for working with long transcripts, legal documents, or large code bases. 4.1 via API is often the best tool.
If you're new to the API, here's what that actually means. Prompting through the API bypasses OpenAI's default system prompt, the part that adds personality, safety filters, and tone. Instead, you define your own system prompt tailored to your task. That gives you more control, and fewer guardrails. It's more raw, flexible, and programmable. Now, there are different levels of how deep you want to go to use this. On the beginner level, there's tools like guey.ai. It's great for uploading long documents and chatting with them and just super simple. Next is kind of the intermediate level platforms like NADN, Make, or Zapier. Those let you build automations or agent workflows using the API. And I do have a full beginner tutorial on NADN if you want to start there. I'd also put a repo prompt in the intermediate tier. That's great if you're doing vibe coding or conversational dev work inside a real codebase. In the advanced tier, there's things like langchain or the openai SDK in terminal—give you full control. If you're in that last category, you probably didn't need this breakdown. Also worth noting: API usage is built separately from your chatBT subscription. So cost matters. Models like 03 Pro, which we'll get to later, are great but expensive. The standard 03 just dropped 80% in price. So now it's about the same price as 4.1, but the 1 million token memory via API makes it perfect for things like refactoring huge repos or reading entire legal bundles with strict instruction following. In this video, I'm just showing the chat GBT interface for simplicity. But this—using models like 4.1 through the API—is where you unlock their full power.
Here's an example of what 4.1 can do: Refactor this 800-line JavaScript project to TypeScript with full type annotations and explain the key changes. This is the kind of instruction-heavy task it handles really well. It's structured, fast, and clean. Perfect for dev work. 4.1 Mini, the intern, is like a junior version of 4.1. Same type of behavior, just faster, cheaper, and a little rougher around the edges. It still supports that million token context window through the API that makes it incredibly useful for working with long files, transcripts, legal docs, CSVs, entire code bases, but the outputs won't be quite as polished or reliable as full 4.1. Like, think of it this way: 4.1 is the senior developer—precise, thorough, consistent. 4.1 mini is the intern—quick, eager, mostly gets it right, but occasionally needs a second pass. That said, it absolutely can handle pretty complex prompts like that same one I just showed. And the result might need a little more editing or clarification than if you ran it through the full 4.1. If you're using the API and cost matters, which it often does, 4.1 Mini is a fantastic budget option for coding workflows, instruction-heavy tasks, long context use cases where accuracy isn't mission-critical. It's not going to reason like 03 or write like 4.5, but it's efficient, capable, and cheap.
If you're looking for speed and reliable reasoning, 04 Mini is your go-to. If you need extra accuracy for math or code, 04 Mini High is the upgrade. These two models are widely used by developers, researchers, and power users who need a balance of performance, quota efficiency, and task-specific strengths. They're not as publicized as GPT-40 or 03, but they're quietly becoming favorites on the back end. 04 Mini is the underrated workhorse. Its reasoning is surprisingly close to 03, and since 03 dropped in price, it's now only about twice the cost of lighter models like 04 Mini. But that difference still matters for a lot of workflows. It's a great fallback when your 03 quota runs out, especially for logic puzzles and STEM tasks. Using that same seating question from earlier, 04 Mini also got the right answer, but in 15 seconds as opposed to 1 minute and 5 seconds that 03 did. And if I did this through the API, it would have been cheaper. So, it's the perfect pick when you want something smarter than 40 but faster than 03. It's great for real-time chatbots or when you need snappy replies smarter than 40. 04 Mini High is the mathematician. The 04 Mini High is the same model as 04 Mini, but with more compute per token. Its accuracy and cost are close to 03, so it's a good option to switch to if your 03 quota runs out. Or pick it when you have STEM workloads and lots of simultaneous API calls. You'll breeze past 03's tighter rate limits. Prove the infinitude of primes using an analytic number theory approach. How about compute Eigen values of a 4x4 matrix and interpret them for a Markov chain? This one shines when you need STEM performance without burning 03 or high-volume tasks that still demand rigor. 04 mini and 04 mini high aren't flashy, but they quietly do a ton of work in real-world pipelines.
The newest and most expensive model is 03 Pro, the Oracle. It's 03 with more compute per token, and that extra reasoning comes with some big trade-offs. The price is much higher, but also it's extremely slow. Like 03 will give you quicker answers when the question is simple. But 03 Pro will think extensively every single time. Like that same seating question took 19 minutes and 45 seconds to get the same answer 03 got to in 1 minute and 5 seconds, and 04 Mini got to in 15 seconds. And I've seen people posting ones like saying, "Hi, I'm Sam Alman. It took 3 minutes and 54 seconds to respond" or "14 minutes and 17 seconds to decide there's three Rs in Strawberry." That's why I think the best use for this model is questions nothing else can get right. If 03 fails, escalate the question to 03 Pro, or things where the stakes are very high. You know, maybe for formal proofs or auditing tricky financial spreadsheets or a final QA pass in an agent pipeline, or maybe you want to give it just a ton of context and have it create an analysis and plan for your business. The situations it makes sense, but for most people 99% of the time it won't be worth the additional time and cost.
So, here is the bottom line. Every model has a purpose, but most people stick with just one. If you're using 40 for everything, you're leaving accuracy, creativity, and powerful features on the table. Use the generalist as your daily driver. Use the professor when you need logic. Use the scholar when you need citations. Use the wordsmith for tone. Use the specialists like 4.1 and the '04 series when speed, cost, or context become a problem. And when all else fails, consult the Oracle. Treat the model dropdown like a toolbox, not a default. And especially if you're building agents or longer workflows, real optimization comes from mixing models strategically across tasks, tools, and tokens.
If you want to go way more in depth on learning AI, on Futureedia, we have over 20 comprehensive courses on how to incorporate AI into your life and career to get ahead and save time. The one I just finished up is about making movies with AI, but we have specialists come in to teach the different types of courses. Another one that just came out is all about coding, and there's a ton of other courses in there. You can get a 7-day free trial using the link in the description, or check out this video that will take you from zero knowledge to building your first AI agent.