Transcription
Many people have written Apple off in the AI race because they don't train models, build data centers, and invest small fortunes in GPU clusters. People say it's just a company that makes pretty laptops and overpriced phones.
But the critics have missed something very important. While Nvidia was building bigger and bigger, Apple was building smaller and smaller. They're running a counternarrative which I believe will reshape the entire AI industry. Today I'm going to show you why exactly the M series chip is one of the most underrated innovations in AI infrastructure and what that means for anyone serious about deploying AI that actually makes economic sense.
Before we get into the detail, I'd like to explain the architecture problem that Apple solved because it's not really obvious unless you've worked inside enterprise technology. Traditional computers have a bottleneck that most people don't understand. It's not the speed of the chips, it's the memory transfer. On a conventional system with one with a powerful Nvidia GPU, the CPU and GPU have different memory pools, which means that every time you ask AI a question, data has to move physically on a bus from a CPU to a GPU and back again. And that's latency, that's energy, and it's a ceiling on performance that no amount of raw GPU power will eliminate.
It's like having two offices on the opposite sides of town. And the only way that you can share files is to put them on a bus that crawls through traffic backwards and forwards. So, it doesn't really matter how fast the people work in the office. Everyone's sitting around waiting for the bus to arrive.
Apple looked at that architecture and said, "What if we just got rid of the separation entirely?" That's the genius of the M series chip. Unified memory architecture. The CPU, GPU, and neural engine all share the same memory, which means there's no need to move data around on a bus backwards and forwards.
What this means in practice is extraordinary. A Mac Studio can run a 70 billion parameter language model locally. That's the same size class as what's powering many enterprise AI deployments, but it runs on someone's machine sitting at their desk silently at around 8 to 12 tokens per second. Now, that's not great by any stretch, and you can certainly do a lot better with cloud models which use vastly more electricity. But the question is, do you really need to? As a rough benchmark, we might say that a Mac M4 would use 400 jewels of energy, whereas a cloud GPU might consume 10 times as much for the same task.
Now, I'd like to also mention the neural engine. Most people understand CPU and GPU, but the neural engine is a bit of a mystery, but it's a very important component in the M series chip technology. Think of it this way. Your CPU is a general purpose worker. It can do anything, but it only does one task at a time. Your GPU is a crowd of workers. They're great at doing thousands of simple tasks simultaneously. But AI inference has a very specific job called matrix multiplication, which means millions of numbers being multiplied and added together over and over in order to turn your question into an answer. Asking your CPU or GPU to do that is like asking a general employee to perform heart surgery. You could probably figure it out eventually, but you'd just rather have a surgeon who's been trained in that job.
The M4 neural engine delivers 38 trillion operations per second. The M5 went further, adding dedicated neural accelerators inside each GPU core. That's not an incremental upgrade. That's a generational leap in a single chip revision.
So, let's imagine a camera on a factory for production line watching every item that passes by, checking for defects. To do that reliably, the system needs to analyze every single frame in real time. If it's too slow, faulty products might get through. The M4 Pro can process over 90 frames per second for that kind of object detection problem. And that's three times faster than the M1 Max from just a few years ago.
Now, here's where the unit economics argument from my last video hits home. A H100 data center card can exceed 700 W, whereas a Mac Studio M4 Ultra runs at a fraction of that. But the power difference isn't just an environmental consideration. It's a direct operating cost, which means if you're running inference constantly, like in a factory or a back office process or in an edge deployment, the energy bill over 12 months versus Apple silicon is dramatically lower.
So, let me bring this back where it all started. Many people have written off Apple in the AI race because they didn't build data centers and they didn't make a GPU to rival Nvidia. But what they did instead was solve the inference problem. And that's the part that actually matters to business. They built a chip where memory doesn't have to move. Where the CPU, GPU, and neural engine share the same pool of fast unified memory, and where a 70 billion parameter model runs on a desktop that draws less power than your toaster. They did that at a fraction of a cost of a comparable GPU workstation. And the data never has to leave your building. That's not losing the AI race. That's running a completely different race. and absolutely crushing it.
The AI industry wants you to believe that intelligence lives in data centers and that you must rent a GPU by the hour to be serious about AI. But Apple just proved that the most powerful AI infrastructure that you can build might already be sitting on your desk. The M series chip is why Apple eats AI for breakfast.
If you find these videos useful, please subscribe and I'll keep breaking down these complex topics to simple bite-sized pieces. I'll see you in the next one.