📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Yup, o3-mini is WORTH your money. Meta Q4 Earnings Prompt. Deepseek and Llama4 Insights

IndyDevDan23:44

Transcription

You've probably already been inundated with a lot of hype, alarmist, clickbait videos about O3 Mini, Deep Seek, and the US tech ecosystem falling apart. Let's push through the narrative and hype and focus on how O3 Mini is a differentiated model that can help us accomplish real engineering work.

Like every model, O3 has trade-offs. In this video, we'll break them down and I'll share my three-step process that I use to understand new models at a fundamental level: Vibe, Compare, Eval.

Instead of asking O3 Mini to solve riddles for us or to tackle arbitrary problems that don't actually create value, let's use O3 Mini to solve a legitimate problem.

Meta's fourth quarter and full year 2024 report were just published, and that means we can extract key information to help inform potential investment decisions. We'll start by taking a look at Meta's Q4 earnings call transcript.

Let's go ahead and open up our brand new reasoning model and compare it side by side. We're going to look at O3 Mini in high reasoning mode next to the previous state-of-the-art.

So, how are we going to extract information from this earnings call? We're going to use a prompt that I built here. This is a 12K token prompt in a highly structured XML format. You can see here we have the quarterly report, and then we have a unique structure that we'll break down in a second.

For now, I'm just going to copy this and paste it into O1 and O3 Mini in high reasoning effort mode, and we'll kick these off side by side.

While these are running, let's break down the stats of O3 Mini next to the previous state-of-the-art. The big headline for me when looking at O3 Mini is that it is effectively O1 at an eighth of the price. I think that's the most concise way to look at this model.

Running through some insights here, it's eight times cheaper than O1 and eight times more expensive than Deep Seek R1. It's important to note that O3 Mini maintains the 200K context window and the 100K max output tokens. This is absolutely incredible, especially when you compare it to Deep Seek R1, which only has a 64,000 token context window and a measly 8K max output tokens.

R1 is a great model for many reasons, most importantly the cost. But if you're an advanced user of generative AI with large prompts and lots of information on the input and output side, you likely ran into this issue, as I did, relatively early on with this model. 64K and 8K out is very limiting.

A couple of other insights, more qualitative: O3 Mini provides state-of-the-art instruction following. I've never seen a model that follows instructions better than O3 Mini.

A couple of other key features: reasoning effort lets us trade off speed and cost for performance. I love this feature and want to see them add additional settings for reasoning effort. I want more control over how much the model is thinking.

On the speed side, O3 Mini is fastest, specifically in low mode. When you use medium and high, it's about as fast as O1.

Again, guys, none of this takes away from R1; it's an incredible model. But we just have to look at the stats. My one-line summary for O3 Mini is that it is O1 but 15% better at an eighth of the price, with all the capabilities built to help you accomplish coding, math, reasoning, planning, and large context tasks.

Let's return to our results here and see what our Meta quarterly report parsing prompt has returned for us. We can see a nicely formatted markdown blob coming out of O1, and then on the O3 Mini high side, you'll see something interesting: this is more precise. We explicitly asked for a JSON response, and O3 Mini is being more precise.

Let's take a look at the result and the prompt we used to get this result. This is really cool; this is something that only powerful reasoning models can do. What we have here is effectively a list of prompts. You can see the purpose here: given a quarterly report, extract the information requested and the information to extract.

We have this dynamic variable that we filled out, and let's just start from the top. We're looking for total revenue for the quarter. If we look at this prompt search, you can see that information was parsed perfectly. We can copy this and go over to Meta's quarterly earnings call. You can see Q4 total revenue was 48.4 billion.

So, our prompt running on O3 Mini parsed that out for us perfectly. This prompt is effectively a JSON object where the keys and the values are themselves prompts that we want a reasoning model to pull information out of the quarterly report for.

If we go down the line here, you can see we have Llama 4 information, we have speakers, which is every speaker on the call, we have year-over-year growth, we have Deep Seek response. I want to see Mark's thoughts on Deep Seek. I have operator AI agent thoughts; I want to get Mark's thoughts on OpenAI's operator. Then we have the largest department growth.

In this list of prompts, we can basically query anything we want in one shot and get precise answers, thanks to the capabilities of the reasoning model. If we just ask for one piece of information, this is something that a powerful base model like GPT-4 or Claude 2 could definitely pull off.

In a benchmark, you'll see in a moment they can pull that off, but when you start adding additional information and you start asking for the response to be itself a specific type of structure, this is where reasoning models really shine because they can iterate over your instructions.

Anyway, let's continue. Let's learn about Llama 4. Llama 4 is an advanced development. Llama 4 Mini has completed pre-training, and its reasoning models, as well as the larger model, are performing well.

This is potentially really important news. If we look at the call transcripts and search for "Mini," we can find exactly where Zuckerberg is talking about Llama 4. We can scroll up here and see Mark Zuckerberg, CEO, responding about Llama 4, saying that Llama 4 Mini is done with pre-training and our reasoning models and larger model—notice the plural there—are looking good too.

Super interestingly, our goal with Llama 3 was to make open source competitive with closed source; our goal for Llama 4 is to lead. Very bold, very bold. We don't have a release date.

You can see here in the query I asked for three pieces of information: Llama 4 in production, expected release date, reasoning model. Just really quick off-the-cuff questions. We can see speakers, we can see year-over-year growth, and Deep Seek response.

Something really interesting about the list of every speaker is that I actually asked the model inside this prompt to respond in a JSON list with this format. This is a really interesting capability that you can tap into with these powerful reasoning models. This is not exclusive to O3 Mini; this is not an emergent capability that only O3 Mini can do. All reasoning models have this capability to respond in a JSON list, and then we give it a format here: name and role.

If we look at the response here, you can see speakers, and we have all the names and all the roles. This makes for a really interesting benchmark, as you'll see later on in the video.

We have all the names, all the roles. We even got the operator, right? Conference call operator. You have your growth: 21%. Deep Seek response: Mark noted that Deep Seek has introduced several novel advances that Meta is still ingesting.

So, some interesting information there. I also asked for sentiment. If we look back at the prompt, I wanted at least three items in objects that look like this. I want sentiment and thoughts, and O3 Mini understood this well and processed it.

He believes Meta can learn from Deep Seek's innovation and potentially incorporate similar improvements into its own systems. Good stuff. He mentioned that it's too early to definitively assess Deep Seek's long-term impact on infrastructure and capabilities.

This is part of the reason the markets are in complete shambles right now. Deep Seek came out at a very low price point, while US companies are putting up models at a much higher price point.

You can see Mark explicitly mentioning that. That's cool to see. We're not going to focus on that too much; we're here to get work done and understand model capabilities.

Then you can see here we have guidance for next quarter. Good stuff. His thoughts on the operator, a browser agent that OpenAI released, are very interesting. I've been talking about how we need new UIs, new UXs, new systems for interacting with our models that are not just a chat interface.

I'm really looking forward to this, and I'm working on a couple of my own. Very cool. Then we have the largest department percentage increase. I thought this was really interesting information: approximately 90% of the year-over-year headcount growth was concentrated in research and development, driven by aggressive hiring in technical areas such as infrastructure, generative monetization, and Reality Labs, so VR and AR.

This is a powerful prompt. I'm going to link this prompt for you in the description. It's going to be in a gist, so you can just try this out if you're interested. Find a quarterly report, paste it in, and then specify key-value pairs that are effectively prompts themselves. Then go ahead and run it against your favorite powerful reasoning model.

Feel free to use this prompting technique, where you effectively have a list of prompts within a prompt. You can do a lot more than you think with these powerful reasoning models, especially as they progress.

O3 Mini represents an improvement in what you can do inside of the prompt. This is one of many prompt engineering and AI coding techniques we explore on the channel.

Like, subscribe, and comment to let the YouTube algorithm know you want more actionable information like this.

Remember that we fired off both O1 and O3 Mini prompts, right? If we copy O1 and just take a look at this side by side—O1 and O3 Mini High—we see a couple of things here.

O1 missed some speakers. You can see it got the information about Llama 4 right, so we have Llama 4, we have Mini, we have reasoning, and we have larger. So that looks great.

It missed a bunch of speakers. You can see here O3 Mini found a bunch of speakers. It got our year-over-year growth correct; we have 21% year-over-year solid responses for Deep Seek in that exact object structure we were looking for.

Good to see, and then it filled out our last three prompts here, and it looks like they did a great job—nearly 90% year-over-year headcount growth.

When you see this, you can see that their outputs are fairly similar. We can see that O3 Mini High outperforms O1 just a little bit in this single case.

This brings us to the question you should be asking at this point: after you run a couple of vibe checks and the new model a couple of times, what happens next?

Now that we have an okay feel of the capabilities, it looks like it's doing work a little bit better than O1. How can we push this further? How can we understand the model at a more intimate foundational level?

This is where we move to the second step in the three-step process. First, we vibe check, and now we're going to compare.

We've already done a small comparison here, but let's scale this up. In the second step, the comparison step, when you're interested in the model and it has something to offer you that previous models did not, the vibe check is telling you there's something here, something more to learn.

Then you can take it to the next step, which is the comparison step. This is where we compare the model against a series of additional models.

We start with a single node, and now we're going wide. Now we have multiple nodes that we want to look into, multiple models that we want to compare.

This is where a tool like Thought Bench from Beni comes in. You can use any tool you want, but basically, you just need the capability of having multiple models running side by side.

We're going to do the exact same thing. We have our quarterly report here; we're going to paste this in and just run it.

Something really awesome about the O3 Mini series is that O3 Mini gives you three reasoning level efforts: low, medium, and high. By default, O3 Mini is going to run in medium mode, but you can see here our O3 Mini in low reasoning effort mode completed already.

Now we're starting to look at multiple model responses. We have O3 Mini in low, medium, and high mode. If we just want to take a look at these three, we can see that low mode dropped some of our speakers.

We have those three speakers, while medium and high gave us all of our speakers. If we look through this, we have all of our speakers, year-over-year revenue growth: 21% across the board. That looks great.

Total revenue: we have that exact same number here across the board. That looks great, and we can continue to look down the line. We can look at our O1 Mini; you can see O1 Mini giving us a nice response there, formatted.

We have O1 across the board. Multiple models can perform this task. It looks like multiple models are doing a decent job here.

Once we get to Deep Seek, we can see something kind of interesting. We are kind of falling off the bucket a little bit here. Deep Seek has split up our guidance and our operator AI agent thoughts into key-value pairs, and our last key-value pair here is not what we asked for.

We can see a little bit of deviation there. As we scroll down away from the powerful models, Deep Seek 32 just had to run this on my device just to check it out, and you can see it's got some NaN here. It's got multiple things wrong here, so obviously a much smaller model.

But we can see Gemini Flash thinking, missing speakers but putting out a good effort here overall. It nailed the structure and looks pretty good.

So, step two in the process is to compare models side by side. When you're looking for brand new capabilities in models, you need to compare them to the previous versions and across all capabilities.

Here we have low, medium, and high. This is pretty straightforward. I don't need to go on about this much longer; you can see exactly what we're doing here in step two. We compare.

Now, step three is where things get really interesting. After you vibe check and after you go wide, now we go tall. Now we're adding an additional dimension. This is where we eval.

This is where we actually run benchmarks. For benchmarks, of course, I've been building up Beni, a suite of interactive benchmarks. You can feel the link in the description. We've covered this in a couple of previous videos.

This is something I'm building up, putting a lot of effort into right now because it's helping me gain insight into these models in a more interesting, interactive way.

Go ahead and drag in a completed benchmark file. This is the final stage in the benchmarking process. Once you run a vibe check and compare the model side by side to other models, finally, you run multiple models against multiple use cases—multiple instances of the prompt.

Here, if we run Flash Benchmark, we'll autocomplete everything, and you can see we have multiple prompts running against multiple models.

If we scroll all the way down here, we can see our O3 Minis, as expected. As we would assume, we have O3 Mini in high mode getting 96% correct with only one incorrect answer.

If we click into this, we can see the prompt and the information. We have the expected result and the execution request. This is effectively the same prompt; we have the Meta quarterly earnings report pasted in here.

If we scroll all the way to the bottom, you can see the information we're extracting. Once you've gone to step one and step two, most engineers and builders stop at step one. Some make it to step two, but if you want to go all the way and dig deep to truly understand the capabilities of your models, you need to set up benchmarks and evals.

It doesn't matter how you do it; it doesn't matter what tool you use. I'm not trying to sell you on my tool. This is completely open source. I'm building this in public just in case there are other engineers that want to check it out, use it, and get ideas from it.

The third step for truly understanding models in the generative AI age is evaluations. You need to take a use case you're interested in and then run multiple versions, multiple use cases against multiple models.

This is super important. There's going to be a point where you can't do anymore if you don't benchmark what you're trying to accomplish. If you're not building a product, if you're not using level four prompts with dynamic variables, I'll link our prompt framework video in the description so you can understand how to write high-quality prompts if you're interested.

If you're not writing high-quality prompts with information—tens, 20, 50,000 plus tokens—benchmarking probably doesn't matter for you, and that's okay. But if you are, this is the most important step for scaling your impact, for repeat prompt usage, for prompt templates, for prompts that you're going to improve over time to solve a specific problem very well.

That's why we benchmark. That's step three. Let's go ahead and take a look at some data insights.

As I was scrolling down here, you might have noticed something really interesting. GPT-4 got 83% correct—25 correct, only five wrong. To be super clear, I have simplified the prompt, like I mentioned before, only asking for one piece of information. It's a lot simpler.

I'm not asking for eight pieces of information; I'm not running eight prompts at once with structured responses inside the prompts. This is a little bit easier. I've scaled down this benchmark here.

I'm in the process of adding some more complexity, some more problems for O3 Mini in high reasoning mode to handle because, as you can see, it's completely saturated this benchmark.

It's really cool to see that, and we can see here O1 preview and O1 performing really well. But look at this price tag. Let's just zoom in on that price tag. Look at this: to run 30 prompts of 12K tokens each, I'm burning $7 here, and even more here.

We likely had more reasoning tokens here, so the total cost is so, so important when you're looking at these models. This is how you can really understand the value.

Again, benchmarking just allows you to understand models with more dimensions. We can also see the average duration here: O1 preview, 15 seconds on average; O1, about 9 seconds on average.

When we look at O3 Mini, we are seeing our time savings come in here—both time and cost. Let's look at that. Look at the cost: $7 for O1, 40 cents for O3 Mini, 42 cents for O3 Mini medium, and 57 cents for O3 Mini high.

You can see the step increase in both low, medium, and high effort, and this is what we love to see. We love to see when these blog posts, these release posts, actually back up what they're saying.

You always want to see your benchmarks align at some level with what they're showing, even if it's a simpler use case data extraction benchmark like this. We can still see that relationship.

No difference here in performance; we paid a little bit more time and money here, but then when we went to high mode, our model turned on for us and solved three additional problems out of 30. That's significant.

Once you have a benchmark, you can really start to improve the prompt at scale and truly understand the capabilities of the model because you're seeing it side by side against models and against prompts.

If we go into this prompt that O3 Mini in high effort mode got wrong, if we scroll down here, you can see it made one very silly mistake. Out of all the things to get wrong, it returned 600 million instead of 700 million.

So, very silly mistake. I'll bet if I ran this again, it would get that right and get something harder wrong. Although, if we look across what was correct and what was incorrect, this is what you want to see.

You don't want to see 26 incorrect because that means something's probably wrong with the prompt or the prompt itself is really difficult. What you really want to see is the variance and the non-deterministic nature of the models come through, which looks more like this.

O3 Mini high mode only got this one wrong. Medium got 5, 16, 21, and 22 incorrect. But then we go to low; it got 1, 17, 30, and 23 wrong.

I think it's good to see that versus saying something like this: you see problem one wrong, O1 preview got problem one wrong, O1 got one wrong, O1 Mini also got one wrong, and so on.

This means that the models just cannot solve the problem. If we look at this problem, this is a really interesting one too. This is where we're saying, "List all the speakers," and this is where we're asking for that list of speakers.

I hope you can see why this is important. Benchmarking gives you a deeper understanding beyond model comparison side by side against one prompt and beyond the vibe check, which is just when we come in here and play, type a couple of things, and go, "Oh wow, it's amazing at five RS in Strawberry," or whatever people are running.

That's not how you understand model capabilities.

This is really interesting. I'm a huge fan of OpenAI's work. I do agree with the bold statement Sam put out a while back: it is much easier to copy than it is to innovate.

I'm not saying that R1 hasn't innovated; they definitely have. I'm not making a commentary on that. But just in general, overall, they took the Transformers architecture from Google and created immense value with it.

Mostly, 60, 70, 80% of everyone else is playing copy with a couple of variants. Again, I'm not saying that's bad; I'm not saying it's wrong. I'm just saying that's the way it is.

OpenAI, as I predicted so far, are the leaders. We obviously want to see the price point come down. I'm hoping we're going to see O3 base model come out at something like five bucks per input and ten bucks per output—something more affordable than O1.

But at the same time, if these models are giving us state-of-the-art powerful capabilities, I think paying the price is always worth it to understand what's coming next.

Don't look at where the ball is; look at where the ball's going. Prices are coming down, reasoning capabilities are going up, and models are becoming more accessible.

Deep Seek and O3 Mini are both testaments to that, and it looks like Meta is going to be following that up with Llama 4 very, very soon.

So, first, you run a vibe check. Then you compare models side by side to each other so you can see how they compare. Finally, you run additional models with additional prompts.

The more variations of your prompt you throw at your models, the better your benchmark is going to be and the better your understanding of the models is going to be.

Not everyone has a ton of money to blow. This is just one test, and I literally torched $7. Let me do the hard work. Let other leading engineers like Simon Willison and Paul, the creator of Aer, do the hard benchmarking work for you.

If you don't want to go all the way, I completely understand. All you have to do is drop a like, drop a sub, stay focused, and keep building.