📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Apple Just Killed Ai Subscriptions Forever ( Mac Mini M5)

Your AI Guy18:39

Transcription

18 months ago, this thing cost $599. It was the most boring computer Apple made. It sat in the back of the store next to the display cables. Nobody made videos about it.

Today, the configuration you actually want for local AI runs you somewhere north of $2,200. Apple has raised the price on this machine twice this year. They deleted the memory option that made it good. And on their last earnings call, Tim Cook told analysts, "The Mac Mini is sold out and might stay that way for months."

So, what happened? An open source AI agent happened and roughly 40,000 people decided at basically the same time that they wanted a computer that runs an AI in their house instead of renting one from a company in San Francisco for 20 bucks a month. Here's why you're here.

Last video, we cracked open a $2,000 AMD box and killed a bunch of AI subscriptions with it. The number one question by a mile was, "Okay, but what about the Mac Mini?" Fair question. So, I [clears throat] bought one. I ran it. And I'm going to be straight with you. My verdict on this machine is not the same as my verdict on the AMD box, and it's not what Apple's marketing wants you to think either.

I'm going to give you three numbers today. One of them is on every spec sheet on the internet, and it's wrong for this machine. One of them will tell you exactly which models you can run before you spend a dollar. And the third one is a date. A specific date on Apple's road map that should decide whether you buy this thing today or wait. Stick with me to the end for the date. It's the one that actually costs people money.

Let's go. You stop paying $20 a month times four subscriptions forever. You buy one box, the model runs on your hardware. Nothing you type leaves your desk. No usage caps, no you've hit your limit, come back in 5 hours. No company reading your prompts to train their next model. Last time the answer was the AMD Stricks Halo Box. 128 GB of unified memory for around two grand. Best dollar per gigabyte on the market. I stand by that call for the moment I made it.

Here's the problem. Memory prices did something in 2026 that nobody in this industry had a playbook for. DRAM contract prices jumped roughly 90% in the first quarter alone. The biggest quarterly move on record. PC memory more than doubled. Why? Because Samsung SKH Highix and Micron pointed almost all their wafer capacity at high bandwidth memory for AI servers. Every rack that goes into a data center takes memory out of laptops, desktops, and phones.

So, that AMD box I recommended, people who bookmarked it in October at around $2,99 came back six months later and found the exact same SKU listed north of 3,000. Corsair quietly added over $1,000 to their AI workstation. Nvidia's DGX Spark went from 39.99 to about 4,700. This isn't an Apple story. This isn't everybody's story. Apple just gets the headlines because Apple is Apple. Which means the honest version of this video isn't Mac versus AMD. It's given that every box on this list got more expensive. Which one still makes sense in July 2026?

Let's start with the number everybody gets wrong. If you Google memory bandwidth for the Mac Mini, you will land on articles quoting 546 GB per second. Some of them round it to 550. Some benchmark channels put it right on the thumbnail. That number is real. It just doesn't belong to this computer. 546 is the M4 Max. The Max lives in the MacBook Pro and the Mac Studio. It has never been in a Mac Mini and it is not in this one. The M4 Pro, the chip actually inside this box, does 273 GBtes per second. Exactly half, not a bit less. Half.

If somebody is benchmarking a Mac Mini and quoting 550, they either grab the wrong spec sheet or they're testing a different machine. Write down 273. We're going to use it in about 90 seconds.

Now, why does this one number matter more than the CPU cores, the GPU cores, the neural engine, any of it? Because of how these models actually generate text to produce one single token, roughly [clears throat] one word chunk, the chip has to read the model's active weights out of memory, every token over and over. So, your speed is not really about how fast the processor thinks. It's about how fast you can pull data through that pipe.

Here's the rule of thumb that will save you from ever getting fooled by a thumbnail again. Take your memory bandwidth, divide it by the size of the model file in gigabytes, multiply by about 0.7 for real world overhead. That's your rough tokens per second. Let's test it. A 7 billion parameter model at 4bit quantization is about 5 GB on disk. 273 divided by 5 is 54 times 7 call it 38 tokens a second. And that's almost exactly what this machine does in practice.

Now do a 70 billion dense model. That's roughly 40 GB. 273 divided by 40 is under seven times 7. You're at four maybe five tokens a second. Four tokens a second is slower than you read. You will sit there watching it type. It works. Technically, you will not use it twice. That formula is the whole video in one line. Everything from here is just applying it.

Okay. Money. And I want to walk the timeline because if you only look at today's price, you'll miss what's actually going on. May 1st, Apple kills the 599 Mac Mini. The entry price becomes $799. May 5th, Apple removes the 64 GB memory option from the M4 Pro Mini entirely. 48 is now the ceiling. They also chop the 32 gig option off the base M4. Same day, the Mac Studio loses memory tiers. June 25th, Apple raises the M4 Pro Mini by $200 across the board. 1399 becomes $1599. And this wasn't a Mac only move. HomePods, Apple TV, iPads, all of it went up the same afternoon. Apple's own statement to the Wall Street Journal basically said, "We've never seen a component get this expensive this fast, and we can't eat it anymore."

Meanwhile, the Mac Studio with the M3 Ultra went from $39.99 to $52.99. That's a $1,300 jump on a machine Apple didn't change at all. So, where does the 48 gig mini land? Roughly 2200, depending on storage. Check the configurator yourself before you buy. I'm serious. This number has moved three times in three months, and it may have moved again between me recording this and you watching it.

Now, the part that hurts. Apple loses. Not close. Even after the AMD boxes got hit by the same memory crisis, you're still paying roughly double per gigabyte to put an Apple logo on it. If your single goal is most gigabytes per dollar, buy the AMD box. I'm not going to pretend otherwise to keep you here, but dollars per gigabyte is one metric. It is not the only metric. And in about 6 minutes, I'm going to show you the metric where Apple wins so hard it's almost unfair.

First, something that changed everything in March and that a lot of buying guides still haven't updated for. Alma switched its Apple silicon engine to MLX. MLX is Apple's own machine learning framework built specifically for how these chips share memory between CPU and GPU. Alma used to run through Llama.cp's metal backend on Macs. In version 0.99 on Apple Silicon, they threw that out and rebuilt on MLX their own benchmark numbers on an M5 Max. Prompt processing went from about 1150 to 1,800 tokens a second. Generation went from 58 to 112. That's a 93% jump on the number you actually feel with no new hardware, no different model, no different quantization.

Now, here's the catch nobody mentions, and it's a buying decision. So, listen up. The MLX path requires 32 GB or more. Which means if you buy the 24 gig M4 Pro Mini to save 400 bucks, you don't just get less memory, you get locked out of the Fast Engine entirely. The 48 gig config clears that bar. The cheap one doesn't. That single detail is worth the price gap on its own.

Okay, the table. These are real world ranges on this box with the MLX backend, 4bitish quantization, ordinary context lengths, community numbers, vary quantization format, and context length. Move these around. So treat them as ranges, not gospel. Look at rows four and seven because that contrast is the entire strategy. A mixture of experts model holds all its parameters in memory but only activates a small slice, say 3 billion out of 35 for each token. So you pay the capacity cost of a big model and the bandwidth cost of a tiny one. Remember the formula. The divisor isn't the whole model anymore. It's just the active part. That's why a 35 billion parameter MOE model runs three to five times faster on this machine than a 32 billion dense one. It's not magic. It's the same math applied to a smaller number.

So what should you actually pull today on 48 gigs? The move on 48 gigs isn't one giant model. It's two specialized models loaded at once. One for agent runs, one for chat, and routing between them instantly. That's a genuine advantage of unified memory that people with a 24 gig graphics card simply cannot do.

Now, the wall GPT OSS 120B does not fit on this machine, period. It needs about 60 plus gigs, and Apple deleted the 64 gig option in May. On the AMD Stricks Halo Box, that same model runs somewhere in the low to mid30s of tokens per second because it's ane and its memory pool is 128 gigs. So, I'll say it plainly again. If 120BL class models are your use case, this is not your computer. Go buy the AMD box. I'll put the link to that video on screen.

Everybody [snorts] still here? Good. Because now we get to the part where Apple actually earns the premium under a genuine inference load, not idle. Actually generating this box pulls somewhere between 30 and 65 watts. The AMD Stricks Halo box, which is itself very efficient, runs roughly 45 to 140. A tower with a real discrete GPU doing the same job as at 7 to800.

Let's do the electricity bill because I've never seen a channel actually do this. Say your agent averages 50 watts and never sleeps. That's 1.2 kilowatt hours a day, about 36 a month at roughly 17 cents a kilowatt hour. Near the US residential average, check your own bill. You're at about $6 a month. Now, the RTX tower averaging 700 watts, about 500 kwatt hours a month. That's 85 to $90 a month every month. Over two years, that gap is roughly $2,000, which funny enough is about the price of the computer.

And this is exactly the use case Apple's own people keep pointing at. Doug Brooks, who runs product marketing for Apple Silicon, did an interview with the Deep View right before WWDC this year. He said Apple is seeing, his words, incredible demand for the Mini and the studio driven by people running AI agents. He described what those people want, a machine they control, separated from their main computer that can run around the clock. He also framed it as a whole chip problem, not a GPU problem. Tool calls, orchestration, all the stuff happening around the model, not just raw matrix crunching.

Here's the thing. Normally, I'd roll my eyes at an exec quote, but this one is downstream of something real. In late January, an open source agent framework called Open Claw went viral. It's an alwayson assistant with shell access, browser control, and critically on a Mac hooks into iMessage, reminders, notes, shortcuts. It hit 150,000 GitHub stars in weeks and kept climbing. And the hardware people chose to run it on was overwhelmingly the Mac Mini. Not because of a marketing push, because it's a silent 5x5 in box that draws singledigit watts at idle, sits on a shelf, doesn't need a Linux VPS, and can hold a real model in memory. That's what actually broke Apple's supply chain. On the Q2 earnings call, Cook told analysts both desktops were sold out and that customer recognition arrived faster than Apple predicted. Apple didn't engineer that moment. They got handed it.

One serious caution, and I'd be irresponsible not to say it. An agent with shell access and messaging permissions running unattended on your network is a real attack surface. There have already been token exfiltration vulnerabilities patched in that ecosystem. Keep it updated. Lock down permissions. Don't hand it credentials you'd cry about losing. That's a whole video on its own. Comment if you want it.

Now, let me push back on the Apple silicon optimism because there are cracks under the surface and most reviews won't touch them. One, the people who built the software advantage are leaving. In late February, Aoni Hanun, co-creator of MLX, the framework this entire performance story runs on, left Apple. He joined Anthropic, and he's not an isolated case. Apple has lost roughly a dozen AI researchers over the past year, including its head of foundation models. The reasons are pretty consistently reported. A secrecy culture that blocked researchers from publishing, and compensation nowhere near what Meta and Open AI were offering. MLX is open source and the remaining team is genuinely strong. But when the loudest advocate for a platform walks out the door, that's information.

Two, the company selling you on device AI is renting its own AI. At WWDC in June, Apple shipped the rebuilt Siri. It runs on a custom version of Google's Gemini in Apple's data centers for a reported billion dollars a year. Reporting says Apple evaluated Anthropics Claw 2 and by some accounts liked it more, but the ask was around one a.5 billion. So, Google won. Sit with that for a second. Apple pays Google a billion a year to power Siri. Google pays Apple around 20 billion a year to be the default search engine. They're running on each other's money. I'm not saying that makes the hardware worse. The hardware is the hardware. But when Apple's marketing tells you Apple Silicon is the future of ondevice intelligence, remember their own flagship AI feature phones home to Mountain View.

Three, finetuning on this platform is still rough. If your plan is to train or fine-tune models locally on a Mac, adjust expectations. The MPS backend is still the flaky part of the stack. This is an inference machine. Inference it does beautifully. Training right now is not what it's for.

Four, no CUDA. If any tool in your pipeline assumes an NVIDIA GPU, and a lot of research code does, you're going to hit walls that have nothing to do with speed. None of this makes it a bad machine. But Apple's software advantage is real and Apple's software advantage is fragile are both true at the same time and you should hold both.

All right, the date I promised you. This is the section that actually changes the buying decision. So put the phone down. Mark German at Bloomberg reported something in July that is without exaggeration the biggest Apple silicon story since the M1. Apple is skipping the M6 Pro, the M6 Max, and the M6 Ultra. All three entirely. That has never happened. Apple skipped the M4 Ultra once. They have never cancelled an entire high-end tier of a generation. The only M6 chip shipping is the base M6 later this year at around 200 GB per second, which notice is less bandwidth than the M4 Pro in this box already has. Instead, Apple is fasttracking to M7. Base M7 as early as the first half of 2027, M7 Pro and M7 Max by late 2020 CS7, M7 Ultra in 2028, and the internal framing on that chip is reportedly about getting into the conversation with Nvidia's Blackwell with an AI server configuration targeting 1 and a half terabytes of unified memory. That's not a typo, that's the gap. Apple is trying to close in one jump.

So, what does that mean for a Mac Mini buyer sitting here in July 2026? Near-term, the M5 and M5 Pros Mac Mini is expected late this year. October or November is the current read. It was supposed to land at WWDC and didn't because of the same memory shortage that's driving all these prices. The M5 Ultramax Studio is on a similar late 2026 track, and the M5 is a real jump for this specific job. In the MacBook Pro, the M5 Pro does 307 GB per second, and the M5 Max does up to 614. More importantly, the M5 generation puts neural accelerators inside every GPU core, dedicated matrix hardware, conceptually similar to Nvidia's tensor cores, which is exactly what those Alma MLX numbers were riding on.

Here's the date. Write it down. Late 2026 for the M5 Mini. And then nothing meaningful for Proclass Apple Silicon until late 2027 because there is no M6 Pro. There is no M6 Max. The gap between the M5 Pro and the M7 Pro is roughly 20 months. The longest stretch between prot tier chips in the entire Apple silicon era. That cuts both ways. And this is the nuance I want you to leave with. If you buy an M5 Mini this fall, you are not two generations behind next spring. You're actually buying into an unusually longived chip because Apple deleted the thing that would have replaced it. That's the opposite of what everyone's assuming.

One more caveat on all of it. Every one of those future configurations depends on memory supply. Micron has said tightness runs beyond 2027. A 1.5 TB Mac is a plan, not a promise.

So, three calls, no hedging. Call one, buy the M4 Pro, Mac Mini. Now, if you need an always on inference box this quarter, you're running 7 to 32 billion dense models or MOE models in the 30 billion range. Your workload is agents, automation, private inference. The software stack is mature. Alma and LM Studio just work. No ROCM wrestling, no driver archaeology, and six bucks a month to run a machine that never sleeps is very hard to argue with. Just go in knowing you bought a 273 GB per second machine with a hard 48 GB ceiling, not 550, not 64.

Call 2. Wait until this fall if you're not in a rush and throughput is what you care about. The M5 Mini is close. The MLX gains on M5 class hardware are legitimately large. And if that translates down to the Mini, the performance per dollar picture changes meaningfully. Given that there's no M6 Pro behind it, an M5 Pro Mini might be the longest lived Mac Mini Apple has ever sold.

Call three. Skip the Mac Mini entirely. If you need 120Blass models or dense 70B at real speed, AMD Stricks Halo Box is the right answer for capacity if you'll live with ROCM. The M4 Max with 128 gigs is the right answer for dense 70b if you'll pay Apple's tax. And if you can genuinely wait until 2027, M7 is where Apple stops apologizing.

Here's my actual position, and it's more conditional than last time. The Mac Mini M4 Pro is the cleanest, quietest, most power efficient local inference box you can buy. As long as your models fit in 48 GB, that as long as is doing a lot of work. The Apple Tax is real, the software advantage is also real, you're paying for both.

Back at the top, I said this was a $599 computer nobody cared about. What changed isn't the box, it's what we decided to do with it. A machine designed to be a cheap desktop turned out to be the best shaped object for a job that didn't exist when it was designed. A small, silent, always on brain that belongs to you and doesn't send your thoughts anywhere. That's the actual story of 2026 hardware, not benchmarks, ownership.

If this saved you from buying the wrong box or from buying the right box at the wrong time, hit subscribe. We do this breakdown every single time a new machine enters the conversation and the next one is already on the desk. Tell me in the comments which number surprised you more, the 273 or the $6 a month.