📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Exposing Apple's Biggest Siri AI Myth (It's Not Gemini!)

Andru Edwards18:46

Transcription

What if I told you that Apple's new Siri AI is not Apple's at all? That behind the voice, behind the interface, you're actually talking to Google. For five months, that was the story. Every tech outlet ran with it. The leaks pointed one direction. Apple missed every deadline. Apple had no AI. And then Apple partnered with Google, and we suddenly have a significantly smarter Siri appearing at WWDC. So Apple wave the white flag and Siri is just gonna be a glorified Gemini rapper. But there's one problem with that story. It's completely wrong.

I was invited to a closed door briefing focused on Siri AI for media only during WWDC where Craig Federighi sat down with his team and addressed the Google question directly. So this is the traditional chatbot architecture. could be any of the things that you're using on your device today, whether from chat GPT, Claude, and so forth. Typically, there's a client app, or this could be something running in the browser that's running on your iOS device, or your Mac or iPad, for instance. It's a traditional application, and that application, of course, makes a connection out to a set of large language models running in someone's server infrastructure. Often there's some harnessing around those models, providing it with access to different tools. And then those models also will often ground their answers in searches of a web knowledge base.

Now listen carefully as Craig explains how a typical AI chat app like ChatGPT, Cloud or Gemini works. So if you take, for instance, the Gemini Assistant, You have the nice Google Gemini app as an iOS app. It reaches out and contacts one of Google's excellent suite of Gemini models. This could be Gemini Flashlight, Gemini Flash, Gemini Pro, or the Gemini image models. And then of course, when it comes to grounding these things in world knowledge, they'd reach out to the Google search service.

And what he said next made everything many thought they knew about this partnership fall apart. So then when it comes to our system, well, we use none of those things. Okay? So, of course, we don't have the Gemini app as our app. In fact, none of that client code is part of how we run an iOS. For these models, we use none of the models that Google deploys to their customers, nor we use the infrastructure and means by which they deploy models to their customers. And then when it comes to the knowledge base, we, of course, don't use Google search or anything like that as the foundation of our system. So I hope that's clear. This is the amount of the Google Assistant we use. The amount of Google Assistant we use is none.

And that broke the narrative. Zero Google code anywhere in the pipeline. And they're not even using Google search for Siri's world knowledge. Apple built an entirely independent infrastructure for that. And if you run a website like I do, you will even see the Apple bot crawling it just like Google and Bing do. So if there's no Gemini code running on your iPhone and Siri is not forwarding your requests to Google, what did Apple actually do? And what was the partnership between the two even for?

To answer that, you need to understand something Apple calls AFM, Apple Foundation Models. Apple built five third-generation foundation models, each built for a different job. Two of them run entirely on your device, two run on Apple's private cloud compute servers using Apple Silicon, and the fifth, the heavy hitter. Well, we'll get to that one later because it's the part of the story where things get genuinely interesting. But here's the thing, building five foundation models from scratch takes an astronomical amount of compute. It takes data at a scale that only a handful of companies on Earth have access to. And Apple, for all its resources, needed a way to accelerate that process, which brings us to Gemini.

Now, rather than going with a complex technical diagram, let's break this down with an analogy. Imagine you have a massive final exam coming up. The subject is impossibly broad. You could study alone in a library for years. But instead, you hire the best tutor in the world, someone who already possesses a vast understanding of the material and can help you grasp complex reasoning, correct your mistakes during practice runs, and accelerate your learning curve way farther than what you could achieve on your own. Then, exam day comes. You sit down at the desk. The test is in front of you. It's you who needs to take the test, not the best tutor in the world. You have to do it on your own. And as you go into the world with the knowledge you gained from your tutor into the workforce, it's you who has to show up at work to do the job. The knowledge, the refined reasoning, the ability to perform, it's all in your brain now, and you're using your own experiences along with what you've learned. And you have the ability to modify your methods and your approach, and even learn new things going forward.

The technology under Google Gemini was Apple's tutor. Apple took the high quality outputs from Google's frontier models and used them to train and refine its own developing models. You feed a student model the outputs of a more capable teacher model until the student internalizes that capability. As Apple put it, all of these are custom built for Apple Silicon, trained using Apple's proprietary data with reinforced learning and refined outputs from Google frontier models. Refined, not replaced, not copied, not wrapped.

And this isn't the first time Apple has played this exact card. See, back in the late 1990s, Apple was in a pretty similar position. Their operating system was aging. They needed to build something modern, and they needed to do it quickly. They didn't start from a blank screen. They licensed Unix, a robust battle-tested foundation that had been developed over decades. Then Apple's engineers took that foundation and built an entirely new system on top of it that includes the custom graphical interface, a completely unique user experience. They mutated it until it was unrecognizable from its origins. But nobody today looks at a MacBook Pro and calls it a Unix machine. But Unix is under the hood. Apple executed a similar playbook here with generative AI. They leverage Google's computational lead to accelerate their own development. Then they built their own custom system tailored entirely for their hardware, their operating system, and their privacy architecture.

But understanding this training story only gets us halfway there. The more important question is what actually happens under the hood when you use Siri AI? You see, when you ask a standard chatbot a question, your request gets packaged up and fired into the cloud to get you an answer. That's how ChatGPT, Claude and Gemini, that's how they all work. Siri AI operates fundamentally differently because of a piece of software Apple calls the system orchestrator. Think of it as a hyper intelligent traffic cop embedded deep into the various operating systems. When you ask Siri something, the orchestrator makes a split second decision. Can I handle this on the device or do I need to go to the cloud? And here's the key. Apple's primary objective is keeping as much computation local as physically possible.

So the on-device model is called AFM3 Core Advanced. On paper, it's a 20 billion parameter model, but it doesn't activate all 20 billion parameters at the same time. That would drain your battery and overheat your device. Instead, Apple uses something called sparse architecture or mixture of experts. The model only powers up and uses the parts of its brain that it actually needs for your specific question. So it might use one to four billion parameters as opposed to all 20 billion. So if you ask it to summarize a long document, it would only engage the language processing portion. Or if you're asking it to identify something in a photo, it would only fire up the vision pathways. That's the only reason it can run this 20 billion parameter model on an iPhone without absolutely obliterating your battery. And it only works on Apple Silicon. The A17 Pro chip and newer M series processors have a neural engine designed specifically for this kind of workload. This is not a generic model that Google allowed Apple to add into iOS. It's custom built for Apple's hardware from the Silicon on up.

Let's talk about more of a real world scenario that Apple demonstrated in that meeting I mentioned earlier. It's the clearest way to show how the orchestrator routes between on-device and cloud without ever exposing your data. So, you are heading to a neighborhood potluck. And you ask Siri, "What is everyone bringing?" Now, for a standard cloud chatbot like ChatGPT, that question is useless. You can't go into ChatGPT and type, "What is everyone bringing?" Hit enter and get any sort of usable response. It's useless because it has no access to your messages. and can't read your calendar, doesn't know what a potluck is in the context of your life. For Siri AI, the orchestrator immediately recognizes the sensitive nature of the request. So it flags it for local processing only, and then it securely scours your messages threads, your emails, and your notes, all on device. Not a single piece of data leaves your iPhone. Then Siri replies to you. Gloria is bringing watermelon feta skewers and Greg is bringing a summer pasta. It synthesized a cohesive answer by cross-referencing completely separate apps entirely on your device.

Now, you follow up. What drinks pair well with that menu? To complement the sweet and savory notes of the watermelon feta skewers, the richness of the summer pasta, and the sweetness of the berry crumble bars, you might consider a crisp Sauvignon Blanc or a dry rosé spitzer. For non-alcoholic options, sparkling water infused with lime and mint or a classic mint lemonade would keep things light and refreshing between bites. This is where the orchestrator kind of earns its name. It recognizes the request has shifted and you're no longer asking about your personal data. You're asking for world knowledge about what drink goes with this food. So it's going to route that query to the cloud. But here is what makes us impressive. Siri maintains the context of the watermelon skewers and the summer pasta. It packages that context securely and sends the query to Apple's cloud model and returns a suggestion without ever exposing who Gloria is, who Greg is, or what was in your text messages in the cloud request.

So then I did that follow-up request to ask about what drinks would pair well with that food. And of course, we have the history of what the interaction was on your device. So the system orchestrator, again, is constructing the prompt and including that context to go up and be processed. And now the model is deciding that it wants to do a search of world knowledge. And what's great about the way we do that is it only grabs the pieces it needs, again, for privacy reasons to go to world knowledge. So just asking what drinks would pair well with those two dishes. Just sense the pieces that are necessary to get the right answer. And that continuity between local and cloud without data leakage is not something you get by slapping a visual coat of paint on someone else's AI models.

And there's one more capability worth understanding, and it is arguably the most sci-fi thing in this entire system. It's called on-screen awareness. Your friend texts you a photo of a strange flat cloud formation rolling over a city. In a normal workflow, you would save the photo, switch to a search app, upload the image, type a query, and wait for results. Nobody does that. Nobody wants to do that. With Entrepreneur Awareness, you stay in messages. You just ask Siri, "Why are the clouds like this?" The orchestrator analyzes the visual data currently rendered on your screen and explains it in real time without you ever leaving the conversation. It understands text, user interface elements, and imagery simultaneously. And it does this because Apple controls the entire stack, the hardware, the operating system, and the models. In this instance, it's more than the model that's answering. It's also what the system can actually do inside the operating system.

Okay, I've been careful to frame this accurately, and that means I need to talk about the part where the story gets a little more complicated. Again, four of Apple's five foundation models run on Apple Silicon, two on your device, two on Apple's own private cloud compute servers. But the fifth model, called AFM3 Cloud Pro, is the heavy hitter. It handles complex reasoning, advanced coding assistance, and large generative tasks that are simply too demanding for on-device or standard cloud compute. This model runs on Google's servers with NVIDIA GPUs. use. And I know how that sounds. I just spent the last few minutes telling you how Apple built an independent system with zero Google code. And now I'm telling you the most powerful model runs on Google's hardware. So is that not exactly what the concept of a Gemini wrapper is? Well, no. And here's why.

Apple didn't rent standard server space from Google and then trust them to behave. They built something called private cloud compute and they extended it directly into Google's data Let me walk you through what actually happens when your request hits the AFM3 Cloud Pro model. Your prompt leaves your phone or your Mac sealed in an encrypted lockbox. It travels to Google's server. Your request gets processed inside a locked box that neither Apple nor Google can open. They physically can't see what's happening while it runs. And then, this is the critical part, the millisecond the answer is generated and sent back to you. is annihilated. The server forgets the conversation ever happened. It's like a self-destructing message. The data can't be stored, can't be logged, can't be siphoned off to train the next versions of Gemini or Apple Foundation models. Apple is so confident in this architecture that they opened the underlying code to independent security researchers. So third parties can continuously verify that there are no hidden back doors, no data leaks and no logging mechanisms that Apple forgot to mention. If Siri were just a rebranded version of Gemini, none of this infrastructure would exist. You would just send your query to Google's API and get an answer back. The private cloud compute architecture, the stateless computation model, and the Nvidia chip requirements, the third party verification program, all of that is overhead Apple would never build if the goal was simply to send everything to Gemini on the backend.

So looking at everything that Apple built here, we've got three things to consider. First, ownership. Apple's models are Apple's models. They were trained using Gemini as a tutor and the fifth model runs on Google's hardware, but the code running on your phone and in the cloud, along with the model weights, the guardrails, the behaviors, everything, all of that belongs to Apple. Craig Federighi was clear. There was zero Google client code in the new Siri AI. Second, the architecture. A wrapper just passes your question along. Siri AI decides where to send it. Most things stay on your phone while bigger tasks go to Apple's cloud. Only the most intense requests touch Google's hardware. And again, even then it's running Apple's code and your data is encrypted and destroyed the second it's done. Third, integration, on-screen awareness, personal context across messages, calendar, mail, notes, and other apps. The ability to follow up a local query about your potluck with a cloud query about wine pairings without losing the thread. These are all system level capabilities that only work because Apple controls the entire system top to bottom. A third party model wired into iOS can't do any of this.

By the way, the takeaway here is not that Apple won or that Siri AI is now perfect. It's currently part of the iOS 27 developer beta. It's English only for now. And the HomePod and Apple TV don't even get access to it. Oh, and Apple still has to prove this architecture works reliably under the real world strain of millions of daily users. So the takeaway is that the story everyone thought they knew that Apple couldn't get AI right, panicked, and just outsourced Siri to Google by putting a coat of paint on Gemini, that's simply incorrect. And it misses the most interesting part of what Apple actually built. That story went viral just because it is simple, cynical, and easy to repeat in a tweet. The truth is way more complicated and honestly, I think more impressive. Apple did the same thing they did with Mac OS X and Unix 25 years ago. They licensed a robust foundation, accelerated their own development, and then mutated it into something entirely their own. The real question is whether Apple can prove this actually delivers in the real world when millions of people are asking it questions it's never heard before. I'll be testing it and you will see that here on the channel, so be sure you're subscribed so you don't miss it. But for now, the next time you see someone confidently say, Siri is just Gemini with an Apple logo slapped on it, you will know exactly what they're missing.

Quick question before we get out of here. Does knowing this, does knowing how the architecture works change how you feel about Apple's AI strategy? Or did you just care whether Siri AI worked when you asked it to do something? Drop it in the comments and I'll meet you there for further discussion. Thanks for watching as always guys, I appreciate your support. I'm Andru Edwards and I will catch you in the next video.