Transcription
So, the headline is that the big Siri update we've all been waiting for is finally going to be coming in 2026, and it will be powered by Google. But this update, this headline is basically now Apple and [music] Google putting out statements confirming the hilarious news that Apple hasn't finally figured out AI. They will be partnering with a company who has figured out AI.
>> I think that MKBHD [music] is wrong that Apple is losing the AI race. His claim is that the Gemini announcement, the joint announcement [music] between Google and Apple that Gemini will power Siri is what we can point to to prove that Apple is behind in the AI race. But I think that that's a very narrow assessment of what the AI race actually is. And I think that his claim is more so talking about the LLMs that power the AI race because it only focuses on that one key component of these systems.
All the way back in 2024, I made a video about OpenAI's 03 model, which was their I think second giant thinking model as they call them. And the headlining news was that an AI model had made it into the top 200 programmers in the world. And so I was making a video about that. And my claim was, who cares? The way I talked about it was that I thought that LLMs were commodities, which was not I don't think it was yet conventional wisdom at the time. Kind of became more and more conventional wisdom in the coming months. And as the year progressed, as we tipped over into 2025, all I could do is continue making videos pointing back to that one, saying, "Hey, LM are still commodities." Grock made a new model that is benchmarking better than the rest. Commodity. Google made a model that's better than the rest, commodity. OpenAI's back out in front, commodity. Deepseek makes a new model, commodity. So everything continued to point back to that direction that this one single component really was not all that important other than the fact that it was a new technology and that it was a technology that was progressing at a rate that felt exponential.
Hopefully, this is starting to sound familiar to you because the parallel that I drew was between the AI race right now and all the way back in the days of the early PC race, in particular, the Apple 2, which I think is fair to say was the first real breakout hit of the PC race. So, back when Apple was making the Apple 2, what mattered the most? Was it the processor that they were using? Was it the hardware that they had built? Or was it the operating system that they had made? Well, the answer was actually that it was the apps. The apps were the thing that made the Apple 2 a success. Famously speaking, and this is again just a reiteration of what is in that video from before, the spreadsheet was invented for the Apple 2. And literally the spreadsheet was the original killer app that made Apple a success as a whole because it made Apple 2 such a big success as a whole.
In this era though, I [music] think that we should actually look at the Apple 2's competitor to see what mattered [music] the most there and then draw lessons from that era and bring them forward into this era. So let's take a look at the IBM PC. The IBM PC was positioned as an Apple 2 competitor. Both in the same product category, both competing for the same customers. And in order to get out very quickly, cuz IBM sat on its laurels for a very long time before making their own micro computer, they said, "Hey, Intel, will you make the microprocessor for us?" And they said, "Hey, Microsoft, will you make the operating system for us? We'll do the system integration. We'll do the packaging. We'll do the marketing [music] and sales, but you guys make these critical components of the actual system itself."
So in that era, what mattered the most? Who made the processors? Uh, to a degree, Intel was a big winner because of IBM. Was it who made the hardware? Well, maybe [music] initially IBM mattered, but over time we found out that IBM kind of fell by the wayside, at least [music] as it came to personal computers. And so, was it the OS? And I think that the answer is [music] yes. Microsoft was the unequivocal winner of that era. Despite the Windtel uh cabal, so to speak, from that period, Microsoft made the operating [music] system and they were the ultimate winners of the PC race. The OS mattered the most because it is what stood between users and their apps. And people don't buy computers. They don't buy microprocessors. They buy apps. This continues to be and will always be true. Apple didn't lose the processor race. It lost the PC race and it lost the PC race because it lost the OS race. Funny enough, this is also true for IBM. IBM didn't lose the processor race. It lost the PC race because it lost the OS race. And on the other side of things, conversely, Microsoft didn't win the processor race. That was Intel's responsibility. It won the PC race because it won the OS race.
And so I think when you take those lessons and you zoom back out and you bring it forward to where we are now, you can say that the most credit that I can give Marquez here is that he is right in a very narrow way. Computers are built in layers. The microprocessor is one layer. The OS is another layer. The apps are the pinnacle. And so if you're winning at apps, does it matter if you're winning in processors? Not really. If you're winning in the OS, who cares if you're winning in processors? The same is true for AI. If you are winning the application race, who cares about large language models? If you are winning the chatbot brace, who cares about large language models? LLMs like Gemini are the new microprocessors of the AI system race. I don't even think you could call it the AI race. I mean, maybe, but what MKBHD is describing is an LLM race, not an AI race.
And so the important the important question becomes if back in the era of the PC you had to look at the critical components the operating system and the microprocessor. What are the critical components of the AI era? So we're going to make it pretty simple. We're going to boil it down to three C's. Context, capability, and convenience. These are the three things that I think are going to define the real competition in what we're calling the AI race. Because in order to determine who's winning the game, you have to know the rules first.
Let's start with the most important, context. I have long thought that context is the thing that will decide who's going to win the AI race. And I'm going to draw you a quick sketch of how I see it. You can think of your chat as a kind of list of a bunch of lines of text that are going bit by bit. And then you've got your AI which looks at this window of text that you've sent it. You can imagine these lines continue throughout this middle section. the AI here. I'm going to draw over to the side as an eyeball cuz it's looking and it sees this window a little like that. So, this is the first important factor and this is the thing when it comes to what Marquez has said and what most people discuss when it comes to LLMs and the hardware that they're running on that actually is related to the LLM itself. The context window Context window is basically how much the AI can actually remember about your conversation, what it actually operates on. Because what happens is is that if you can imagine this being a really, really long amount of stuff further back, it gets harder and harder for the AI to remember everything. In fact, there are actual physical limitations on what it can even look at. But even if it is very long before that, the more context you give it, the harder it is to make meaning of what's actually in the text.
So there has been a lot of discussion about what makes up these uh chatbots. In fact, I've discussed it in the last video as I've talked about before, which is that um one of the things that turned LLMs into these chat bots is the system prompt. Now, the system prompt can start off very simply. It can basically say, "You're a chatbot. You're working to help this user." It's sort of a discussion of what it is that you're actually trying to achieve as the chatbot. And then the LLM, which is just predicting the next word, looks at this and thinks, "Oh, well, at the very top, I've been told to do this thing. And so as a part of my context, I'm going to include that in my context window here. Right? I'm going to include it up here at the top. So I'm always listening to the instructions that the person who's created me give me. Um these of course can get more and more complex. It can tell you uh what the kind of brand guidelines are, what the uh safety guidelines are. can give you the ins and outs of when to respond, how to respond, all of these sorts of things. Maybe not when to respond, but more so how to respond.
The second piece of a chatbot beyond just the LM and the context window itself is the chat history. Now, you may not think of it this way. Maybe you do, maybe you don't. But the actual history of the chat, everything that's included in that context window, the things that you've sent beforehand are part of the context of what an AI is considering when it responds. So if you imagine that you're about to respond right here, well, it takes into consideration the system prompt, everything that's been said before, both by you and by the chatbot. And then it takes into consideration the very last thing you said, the prompt, and then it responds with this information down here. So the chat history is actually part of the context itself. This core set of things is what the early chat bots looked like when you first tried out chat GBT when it was blowing up a couple of years ago. This is probably the summary of what uh your interaction with this chatbot looked like.
now pretty soon after that and as a part of 03 like we've discussed there was a fourth thing that was added which was thinking. Now I'm going to put a pin in that. Let's come back to it in a second because if you think about the cleverness of just these first three items in the past, you might have heard of this uh upcoming prompt and the way that you phrase it as prompt engineering. Now, some people sneered at this uh title, and some people thought that it's kind of a pretentious way to look at things, but it's true. It is what you were doing. You were engineering the next prompt such that the chatbot could respond in a way that was to your liking. I think that there's a better term that has come along. It has been less in vogue, less discussed in the broader media, although it has been quite well discussed in smaller circles, which is not prompt engineering because as we know, it's not just the system prompt and it's not just the prompt that you're including here that that's going on. It also includes things like the chat history and as we're going to talk about thinking here. The better word is context engineering, which makes sense because we're talking about this context window. And so thinking was a clever hack on this idea of context engineering because the I the thought was that well if we're including a system prompt and we're including all the chat history and we're including what the person's request was, what if there was a step here right after the fact in the chatbots context that included a bunch of hidden messages, things that it sent to itself. So these are not visible by the user but this is the chatbot basically taking pen to paper thinking through how it will respond thinking through what it will use to respond and then collating all of that into a final response at the very end. So thinking was the first real innovation when it came to a underlying methodology that helped out with context engineering be beyond just the system prompt and the LLM itself which brings us to the fifth item which I think is going to be the one that is most important which I'm just going to kind of bucket into a broad thing that I call external resources. You've definitely seen this in your chatbots if you've used them. This is things like web pages that the chatbot reads that come from a search that the chatbot does. You may have heard of this called RAG, retrieval augmented generation. The idea is that you have all of these documents and these web pages and this data on the outside beyond even the just the chat history. This is stuff from the broader world, maybe from your computer or from a database or from the internet. And that as a part of this step between your prompt and where the thinking happens and within the response, you can take some of these things and sort of suck in the appropriate ones, the most relevant ones to form a good response. So that now that they're in here in that thinking stage. And so this this kind of automated context retrieval is part of that context engineering. And this is why context is such an important piece because LLMs are really ultimately kind of dumb things. All they do is predict the next word. They have to feed the words that they're generating back into themselves in order to generate a complete thought for God's sake. And so all you can do as someone who's using this piece of technology is provide the right context within the context window in order for it to plausibly provide the correct answer at the very end of the day down here at the bottom. And so you can imagine that first of all in a scenario where you look at something like chat GPT versus Gemini, the search function is really really important because how well your search function works determines the quality of the information that's being provided back into this thinking stage. But there's another piece here that I think is even more important that we're going to come back to later, [snorts] which is personal data. These are things like your messages, [snorts] your emails, your photos, [music] your calendar. I mean, you can basically imagine that original iPhone home screen. And every single one of those things has [music] a lot of personal information in there that's important. Context. But there's other stuff, too, [music] that's less specific, but that's also important. There's kind of more contextual information. Your location, plus your location history, where you are and where you've been. uh your activity, what's in your health app. I mean, we've even seen [music] chat GBT recently, OpenAI, try and take in some of this information so it can provide a better service to its users. And I could go on and on. Literally anything about your life [music] could be included into this context with the right search function, with the right retrieval function.
So that's context. Just to summarize, context is the amount of information that an LLM takes in and considers before it provides a response. The better that information, the higher quality that information, the better it can sort through that information, the better the response will be and therefore the better the chatbot will be, the better the user experience is. So, here are the kinds of questions that you should be asking around context whenever you hear anything about an AI race. Uh, what data do you have access to? How good is your search function? What tone does your chatbot take? How does it behave? More and more lately, I've been saying this, and maybe I'll make a whole video about it another time. I think the company with the best search function is going to have a huge advantage over anyone else. And that's because basically the automated portion of context engineering is fundamentally search. And so if context is a very important piece of the puzzle here, the person who has the best context engineering is likely going to win.
That brings us to our second C. Because if context engineering is all about providing the right answer to display the right information, uh capability is all about taking the right action to manipulate information. I mean, think about it. The things that you want out of a chatbot are one, it listens to you and can show you the information that you request that you can retrieve the right information. But the second thing is that once you have that information, there's so many times on a computer where you want to manipulate that information to change it, to uh move it one way or the other to contextualize it, to add style. And so context gets you halfway there. It gets you the right thing. It provides the right information. But capability is where things really start to heat up. [snorts] Let's talk capability. I'm going to begin by making another diagram here, beginning with um what I would consider represents the LLM. I'm making it a black box for a reason, people. Now, the LLM is this black box that we feed a bunch of information to and then get something out of. Now, nice for us. If we look back at our context discussion, we know some of the information that we're going to put in this thing, beginning with the system prompt. Um, before we represented it as a little bubble at the top. So, we'll do that again. Here we have a little bubble here, which is the system prompt. >> [snorts] >> uh we will feed that in to our machine. Then beyond that we've got the actual chat history itself which we represented as little discussion squiggles that gets fed in as well. And then of course we can't forget that we also have a prompt that we've written which I'm going to put here at the bottom because it's what sits between the system prompt, the chat history and the thing we've typed in. So there you go. Those things get fed in to our LLM here again our black box. And beyond that what we end up doing is that we get out something else in return. And I'm going to make this a different color. So make sure that that is very clear which is a response. Now really what we actually get out is a token or a word. One way to think about this is a word. And so one word is often or one token really is not enough to get an entire response back. So what do we do? We take this and we feed it back into the system. We add it to the rest of the response which I'm going to likewise make red. This is sort of a concatenation of all of them together. And guess what? That response gets fed back into our system as well. And it continues and it keeps going and going and going. Okay. So when does it stop? Now, this is the core question that'll kind of illuminate what capability is really getting at because if this thing is just circling around itself over and over again, then there must be a point at which it knows how to stop. And the first clue is that really if you zoom out here, it's not the LLM itself that is doing this but a broader system overall because the LLM is actually running as a piece of software. Of course it is. Okay, that's kind of interesting. That software is running on hardware as we know. But the real question is not how the LLM knows when to stop because it doesn't with a small asterisk. It's how does the software know when to stop running the LLM? And the answer is that one of these words that the LLM knows is a special token. It's a special response that uh to the computer looks like and I apologize for my octagon here. It looks like a stop sign. And when this software encounters this stop sign, it says, "Okay, we're going to stop." And so really, it's the software here that is doing the stopping, not the LLM itself. The LM just sends a signal here. And so in a very rudimentary way, in the simplest versions of these chatbots, one of the capabilities of the LLM, of the chatbot, of the software is its ability to stop. I know, kind of boring, right? But really, this stop sign is just one of other tools that these chat bots have access to. And again, we'll find ourselves familiar with a couple of them. So, if a stop sign is one of them, the signal to the computer that it should uh stop here. Well, another one would be our thinking module. What you would have here is that you probably have a system prompt or maybe it's trained to do some thinking and it would spit out these things and it would spit out a separate special thing that says, "Hey, time to think or maybe even time to stop thinking. And guess what? The software would move it on to the next phase. These are distinct sections of the software. Continue the loop, move on, give the response in return, and then spit out the stop sign. The last version of these that you might be familiar with is search. Cuz guess what? Search is also a tool that these things have access to. So, how does the software know when to use search? Well, this one is special because it's not triggering the end and it's not triggering a different phase of the LLM. Uh, so I'll use a different color here. This one is triggering a actual programmatic web search of the worldwide web. And so it spits out this token which says to the software, "Pause. We're not going to go any further yet." And the software says, "Oh, okay. I'm going to go out to another piece of software over here which knows how to do a search, knows how to read the web pages that you're providing me. And then I'm going to provide those things back into my LLM as more context, more information for it to then [music] feed back of course into the system. That should probably be black. So we'll make it black.
So tools are really just signals to the computer that tell the computer to run other software, perform another function, do something else in return. And one of the interesting things is why does this have to stop here? Why does it have to stop with retrieving information? Why does it have to stop with uh creating a better uh thinking process for these chatbots? Why does it have to stop with literally stopping the computer itself? Well, of course, the answer is that it doesn't. As you know, on top of this hardware, there's plenty of other applications that you are running at any other given point in time. One of them may be search, one of them might be a photos app, one of them might be Tik Tok, who who knows. and all of these functions, all of these things, these bundles of software that this stuff has access to, as long as the LLM has some special signal that says, "Hey, computer, go out and do this thing. Report back with information, can become part of the capability of the system as a whole." This is super cool. Um, and you'd imagine that this would be something that anybody could take advantage of, but we're going to come back to it. I'm going to point to this as another reason why Apple has a very unique advantage over [music] pretty much anyone else with a couple of gigantic asterisks.
And so that's capability. If context was all about finding and displaying the right information, then capability was about actually performing the right thing, doing the right action on that information, manipulating that information in the way that the user wants. If the first two C's were about the chatbot itself, convenience is all about how you actually get the chatbot into the hands of the people that might want to use it. Another word for this might be marketing. Um, but I think it's also a good place because there's not a lot of diagrams to be drawn in in this one. I think it's a good place for us to begin our competitive analysis uh to explain why I think that uh MKBHD and the team there maybe miss the mark a little bit on on the common sentiment that people think that Apple's behind in AI. Um, after we do this, after we talk about convenience, we'll wind our way back through these first two C's. But let's go ahead and start with the first part of convenience which is distribution.
So take a look at chat GPT. How do they actually get chat GPT to you? Well, maybe you access it through a web browser. Maybe you use their app. Um, but it's 100% software in the way that they get their their chatbot to you. Now, how do they make you aware of it? Well, they market, right? They uh have been running a lot of ads lately. If you've been watching any TV at all, you'll see a ChatGpt ad, you know, most likely if you're watching any period of time on any type of broadcast TV and sometimes even on like something like YouTube. What about someone like Meta? Uh, just as a comparison point. Well, they have pretty good distribution, too. They also mainly get their products into your hands via software, and they haven't even had you download a separate app. In fact, they've stuffed their chat bots into the Facebook and Instagram apps for the most part um to get their distribution. So, that's kind of an interesting way to look at things.
Google is where things start to look the most interesting, I would argue, because guess what? They also have a way of getting the software to you, which is Google.com. If they wanted to tomorrow, they could switch Google to Gemini. They could just 100% do it, and every single person who has been using Google would get Gemini in instead. You've already kind of started to see that with AI mode and the AI summaries at the top of pages. You're seeing a lot of Google's prowess and their distribution coming together in order to make Gemini get access and the some of these AI features get access. But they've also got hardware products, too. They've got Google Home. They've got Android. And all of these things are stuff that they have 100% control over uh because they own the hardware and the software and the firmware. But they also have other applications like YouTube for example. You can chat with Gemini in any YouTube video. And I've already heard tons of people using these products. So Google knows their distribution advantage. They know the challenge that their existing business model combined with their distribution advantage gives them. But uh I think that they're doing an incredible job. And if the conventional wisdom right now is that Google is ahead, it's because they have used their technology combined with their distribution advantage, combined with their business model in a very stellar way over the past 12 months, I would say.
And that brings us to Apple. Um, they're probably the easiest to analyze out of anybody other than maybe Chat GPT because guess what? They've got the iPhone, they've got Siri, and they've got all the other devices in their ecosystem that can use Siri. So at any moment in time if they want to replace Siri with another Siri, they've already gotten that distribution advantage uh without even having to ask anybody. Which kind of brings us to a second but related piece which is accessibility. How accessible is this chatbot? Well, the web is one way to access it as we talked about with chat GBT. Apps are another way of doing it as well. But the lesson here is going to be that if you control the device, you have a huge advantage overall. It's just way more convenient to have a hardware button that is automatically by default mapped to a chatbot than it is to inform someone that they that the chatbot exists. Uh convince them that they're interested in it, get them to download it, sign up for an account, etc., etc., etc. It's just uh a lot of work to get somebody to do all those extra steps. And so if you can bake it into the device, you already have an advantage. So for example, you're not going to see Open AI get access to the Siri button on an iPhone without legislation. You're not going to see them get access to Google's voice activation on any of their devices [music] without legislation. [crying] And by the way, that's why they're rumored to be making a device, maybe a pen, because they know this. They know if you control the hardware, the firmware and the software, then [music] you can actually control the entire experience and uh you're not confined to the walled gardens of someone else's e ecosystem. So in this way, chat GPT, if you think that they're ahead in the AI race, is actually at a huge disadvantage when it comes to convenience. Uh Gemini I think is arguably the best in their position [music] because they have tons of touch points all over the place and they span all the way from hardware to the actual applications themselves. And I'd say that Facebook and Instagram are pretty decent too, although they're kind of weird. They are stuffed into those corners of those apps. And I think if you look at things as a whole, maybe Apple isn't beating Google, [music] but I do think that they're in a pretty good position here as well because they have every single iPhone that can support Apple intelligence as a touch point for their new version of Siri. So that's a huge advantage on their part. And I think that if you look at the big picture of things that this is just the first thing that shows you that Apple shouldn't be considered behind because these are hard modes that they have built over time that I think are only accelerated as long as they can execute on the actual chatbot thing which we'll talk about here at the end. [snorts]
Let's wind our way back through Apple's other advantages through the first two C's and then we'll close out. Beginning with capability. Well, if you'll remember, we talked about all of these extra functions that are built as software on top of the hardware of a device that these chat bots are ultimately running on top of, right? You'll remember that our LLM was spitting out these specific commands for the software to see, intercept, and then dispatch these tasks to other applications that are on the device. uh causing some action to be taken, causing some uh information to be retrieved or some combination of the two of them. [snorts] And so if we look at I'm going to go and say the three competitors here which is Chad GPT, Apple and Gemini. Maybe a better way to think about it would be Google uh OpenAI and Apple. Well, Chad GBT is kind of at a huge disadvantage in a particular way. They don't have any software that runs on their own devices. If you think about it, you need a list of functions or apps and they need a way of dispatching in order for this stuff to be useful that we talked about earlier. ChatGBT definitely can find a way to dispatch. I mean we know that we see that its tools are well done and that you know they even have a search function that they licensed or built that has uh buttressed their chatbots capabilities but nobody's building I mean to a first approximation no one's really building on chat GPT's uh software architecture building actual full applications beyond GPTs themselves and so chat GPT's best hope is that they will gain access to other applications out that are publicly accessible and indexed in other databases on the web. Um, so maybe this isn't on their own hardware, but somebody else's server, they can go and ask them for some application, something to be done, and then get that information back. By the way, this thing is basically called MCP. If you've heard any buzz about MCP lately, model context protocol, this is what all these people are talking about. If you think about all of the different uh servers on the web as a series of databases with functions that can be done with things to uh communicate between the two of them, then chat GBT hopes that well guess what their chatbot can sit in the middle and dispatch their software which will then go to this list of functions and say oh over on Google's server go ahead and do this thing to their email over on Apple's device go and ask it to do this thing. Now, this is gated by everybody else cooperating. So, guess what? Apple's probably not going to cooperate. I know people have said that they're going to use MCP, but it doesn't seem like they would cooperate with their own software stack, any of the apps on the app store, which we'll talk about here in a second. Um but for example uh another small news site might if they choose to participate or other types of small web applications might as well something like Canva might uh have some interest in integrating with chat GBT. So ChatGBT success is gated by the community's decision to buy into MCP and the quality of its ability to index those MCP functions.
What about Apple? Apple is interesting because it kind of goes the other direction. Now, it could take advantage of MCP as well, uh, for sure, but it also has its own software and hardware stack. So, it would make a lot of sense for them to be able to tap into apps at least uh that are first party, the ones that they're building. You can imagine it tapping into sending messages or uh adding calendars or creating notes. In fact, you already see Siri doing this as firstparty apps, right? But of course, they have a whole universe of other apps as well. They could basically roll their own version of MCP locally for third party apps on the app store or on somebody's device. And you can begin to see the beginnings of these things uh by what they call their intense framework. Or if you uh aren't interested in the kind of more nerdy side of things, well, if you've made it this far, hopefully you are. But if you aren't as nerdy, maybe you're thinking of something like Shortcuts. Uh Shortcuts is built on these little functions that Apple's system knows and has access to. And you could imagine a Siri that would know all of these things and go, "Hey, let's go ahead and take a look at this shortcut."
Which brings us to Gemini, which of course is a little bit of both. As we talked about, they have the advantage of both the hardware and the software distribution. And they're also very good. I mean, if Google's good at anything, it's at indexing the web. And so, you can imagine that they would also do well by MCP. They also have their own firstparty apps. If you look at Google Docs and Google Calendar, etc., Gmail, uh, and they also have third party apps as well on Android. And so Google seems very well positioned here as well, as long as they can get their um, third party apps to kind of play nice or as long as they can get MCP to work as well. And so to me, Chad GBT again seems to have the worst luck of the three. Google seems to be a mixed blessing in the sense that they're going to benefit as far as MCP goes. They're going to benefit from all of their existing app infrastructure. They'll benefit from to a degree the um third party apps on the Android app store. Apple also will benefit from MCP. They're of course going to benefit from their first party apps. And guess what? They have the most robust and uh most coveted third party app ecosystem on their built for their own devices. And so they [music] really benefit from this. So if I had to say, I would think that Apple has a slight advantage here overall, although Google is a very close second. And so you can start to see Apple and Google jockeying a little bit already when it comes to capability.
So hopefully by now I'm actually starting to sound a little bit like a broken record. When it comes to convenience, I think that Apple and Google have huge advantages. Chat GPD has an uphill battle and I think that if you had to say who's the number one, I would put Google. When it comes to capability, the actual manipulation of information, the access to applications and functions that are required to manipulate that information, chat GPT seems far behind. And I would say Apple and Google are relatively close to one another in terms of the uh amount of information that they have access to. If I had to say I would put Apple slightly ahead just because of their third party app support and their history of being able to get developers [music] to do their bidding uh just over the course of the entire second act of Steve Jobs's uh time at Apple and arguably since the Apple 2.
Let's go ahead and move on to the final piece here, which is context. So, when it comes to context, we'll look at everybody again. We're going to look at OpenAI. I'm going to do better this time. We're going to look at Apple and we're going to look at Google. There are two ways that I look at the context discussion overall and they're sort of separated by the two external resources that we've already discussed. Starting with OpenAI, um I want to look at this as it comes to world knowledge and as it comes to personal knowledge. So when it comes to world knowledge, I think the best way to look at this is that anything that you have access to that is out on the internet. If you've got a good search and can retrieve web pages well and you aren't blocked by a lot of these crawlers, then you're going to do a okay. And by the way, anybody that has access to the internet is likely going to do a okay here. So when it comes to world knowledge, I think it's relatively even playing field because everybody's got access to pretty good search and pretty good web pages. This is stuff that's accessible to everybody.
What about when it comes to personal information? Personal data that we talked about messages, photos, notes, location history, activity, contacts, calendar, emails. This itself I would like to separate into a couple of of things as well, which is stuff that is softwaredefined. and stuff that is limited by hardware. Beginning with OpenAI, well, as things go in terms of software defined information, they have a little bit, right? They've got a canvas feature. They've got uh somewhat access to uh the APIs of like Gmail. I think they call them connectors on uh uh chat GPT. I don't remember exactly, but really they're not producing much of their own data other than the chat history. That's sort of their unique advantage because they were there first because they're the biggest name in the space. They've got a lot of information about you recently if you've been using it a lot. That's kind of where it begins and ends for them. Now, what about hardware? They've got pretty much nothing. And so again, for OpenAI, other than their chat history, slight advantage, if you could even call it one, it seems pretty bleak.
What about Apple when it comes to a software advantage? Well, of course, they've got they will have all your chat history hopefully once you start doing it. But minimally, they've got every app on your iPhone. Everything in iCloud, everything on your iPhone, they've got access to. Messages, check. Photos, check. Notes, check. Emails, check. Calendar, check. Contacts, check. You go down the list, any other piece of information that's accessible and indexable. Anything that Spotlight Search can access, they've got access to. Good news for them when it comes to software because they're storing all of that information. What about hardware? This is more about that stuff like location history, activity, things that are kind of passively involved to some degree. Well, of course, they've got hardware as well. And they've got full access to it. In fact, they've got full access to pretty much all the hardware that they run on. And so for them, they will likely do well by the context discussion because they've [music] got both the personal and the world knowledge to its fullest extent. By the way, Apple's really embarrassing failure with more personal Siri. This was their pitch. They recognize their strategic lane is unique overall amongst everybody because they control so much of this information for the audience that they care about. Anybody that owns an iPhone is going to be doing pretty good by a lot of this stuff. And that's why they pitched more personal Siri this way. It's just a matter of execution at this point.
What about Google? Well, Google, of course, we know has pretty good world knowledge. They have all of their apps as well. Uh they've got photos, they've got notes on their Android phones, they've got messages, of course. They've got Gmail, Google Calendar, Google Contacts, Google Docs. I mean, you could keep going on and on. They're doing fine by software information. Any piece of data that you have put into Google, by the way, including search history, [snorts] is going to be ripe for the pickings when it comes to Gemini. What about hardware? Well, it depends to some degree. uh they their hardware is kind of in the middle of the road because it's limited by the people who they have access to. Now, depending on the reach that you're talking about, of course, Android is bigger than iOS, but when it comes to the economic battle, I believe that Apple has won. So, I'm going to give them kind of a middle of the road here because they've got access to hardware. They've got lots of hardware deployed is like a a positive thing for the most part, almost a check mark, but they I think have slightly less of an advantage when it comes to Apple. but they're doing really really well overall too. So I think that they get their big check mark overall at the top. So once again, for the third time straight, we've seen OpenAI fall by the wayside and Apple and Google, I would argue neck and neck.
So what's leading to this perception that Apple is behind? The answer is quite simple and this is where MKBHD nails it, which is that they haven't shipped. The Siri delay, the whole issue with JG and all of the Apple executive drama that happened after the failure to launch more personal Siri is the best evidence that that is the thing that you can point back to that says that Apple is behind on the AI race because they didn't ship the thing that they said they were going to ship. There's lots of nuances in these things that would prevent Apple, if they have high standard for themselves, from shipping a product that would take advantage of the advantages that we've discussed that Google has actually done a better job of shipping overall, by the way. But when you look back at the sort of Gemini powering the new version of the more personal Siri discussion, I don't think that that's really anything as it relates to Apple's position in the AI race. Because guess what? Apple cares about competing in user interface, in operating systems, in applications. They don't care so much about competing in the things that can be commoditized unless they have to. And so when you look at these three C's, context, capability, and convenience, Apple and Google seem far and away the best position to win the AI race. And they're the horses that I would be betting on right now. And the whole Gemini thing as it relates to Apple is just evidence of how lucky, I guess you could say, they've gotten with their strategic position overall. [snorts] Because guess what? All of these things as it relates to context, all of the personal data, the messages, the photos, the notes, and all of these applications, the third-party hopefully quality third party applications that Apple has at their disposal uh without a need of reliance on MCP are things that they have already built over decades by the way of time in the industry. So, if you can replace the core LLM, the tiny little dot in our diagram, or the tiny little eyeball that's searching here, and just buy it for a billion dollars a year from someone like Google, well, that seems like a pretty good exchange for them because guess what? If they don't have to build it, they have all of these other differentiators that they can care about that can put them ahead in the AI race. And that [music] is why why I think it's silly to point to this and say Apple is behind because this is actually more evidence in my opinion that Apple made the [music] right strategic bets and that they have a chance to vote themselves ahead.