📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

The Great AI Migration (smart entrepreneurs are ditching cloud AI and going local)

Preston Rhodes40:05

Transcription

This is called the great AI migration: why smart entrepreneurs are ditching cloud AI and going local, while everybody else gets left behind.

I believe that there's a massive shift happening right now in AI, and a lot of the evidence is undeniable. There are three forces right now that I believe are creating an inflection point with AI. This is going to be one: open-source model performance. ChatGPT, for some of you guys that don't know, is like cloud-based and it's not open source. So you can't just like access the code online and install it on your hardware and run it offline. And a lot of, for the longest time, the best AI models have essentially—if you look here on artificialanalysis.ai—like the best AI models right now are pretty much cloud AI. So Gemini 2.5 Pro right now is the best; 03 Mini; DeepSeek R1. DeepSeek R1, though, is uh, open source, so this you can run locally if you have the hardware to do it. DeepSeek v3, same here; GBD 40, no; Llama 4 Maverick that just came out—yes, that's open source. So Zuckerberg and Meta, they even said, if you saw the other day, open source is the future. Um, Claude 3.7, no, that's cloud-based; Gemini 2.5 CL Flash, that's cloud-based; Llama 3.3, that's open source; Mistral, that's open source; um, Nova, I haven't even used that; Llama for Scout, that's open source; GPT-4 Mini. So what is the, you know, consistency that you're seeing here is, you know, the top—it was like when DeepSeek in China came out—how are they going to compete? Well, they're going to compete uh, by one, having an amazing model, but then two, open—like making it open source and making that the future. So that it forces the other companies to, you know, use open source in the future, because who's going to trust, you know, Google and OpenAI and all that with censorship? Uh, definitely not me. So that is essentially uh, the—the fact is that open-source model performance is getting extremely uh, better fast. And not only that, they're able to quantize it or break it down into a format that is going to be more available to everyday users. Especially even as the models become better, they're creating like a better form of essentially allowing you to install it, and this is going to be the future. So that's number one.

Number two is hardware efficiency breakthroughs. So previously, you wouldn't be able to run DeepSeek R1 on, you know, a normal computer that everyday individuals have, but now they are creating uh, uh, computing resources that can run these LLMs that are being quantized down to these smaller sizes. And a good example of this is what—recently I made an investment into—into my business as a systems integrator, building, you know, proprietary systems and processes with my clients. Mainly it's been on the cloud. But um, I have people coming to me in these industries that they can't run stuff on the cloud, and we need to be able to like also not break NDAs and all this stuff. So I figured what better an investment, because my community backs me and they know what they're, you know, investing into is that I'm going to be spending it on better technology and innovating. So um, that's what I did. I got the brand new M3 Ultra Max Studio, which is a state-of-the-art newer um, like resources on the M3 Ultra—trip—that on the M3 Ultra—trip—that is able to sustain the uh, the amount of uh, computing power necessary to run these AI models. Now you might be asking like what's the difference between running it on the cloud, like if I have the M3 Ultra and I'm running it on, let's say DeepSeek R1 on DeepSeek.com versus downloading the 70 billion parameter for quantized version of it, which I'll explain in a second, and um, you know, the difference is you're going to be relying on DeepSeek R1's like actual servers and computing resources themselves, and also the latency in the network and a bunch of people using it. Probably saw with the image generation, which I'll maybe even go into too: local image generation. Um, so there's a lot of reasons why not only like DeepSeek—you try and make a custom GPT on DeepSeek, you can't do that. You can do that locally if you're—or I—I want to stop myself there—you can actually do it on the cloud, um, and I'll show you guys a way here in a second, but it's just not as realistic um, in terms of at an enterprise level or at scale. So um, hardware is getting better, whether you're a Mac guy, OS guy, or uh, like a PC guy. I know Nvidia is like making even better and better chips, and like this is where all the money is going into is these computing resources. So one, open-source model performance is getting better; two, hardware is getting better and more accessible; and then three, the model quantization advances. So this is essentially what I was referring to with the, you know, open-source model performance. Not only is it getting better in terms of on a scale of compared to the other ones, but it's uh, because it's open source, they're allowing an ecosystem of other developers to extract certain parts of that and uh, like make it run faster for you, but not as smart. And so breaking it down into different uh, sizes so your computer can handle it. So the smaller the quantization level, the faster it's going to run, but the uh, worse output it's going to give you. So if you have a bad computer or a low computer, you can do everything we're going to be showing on today, you just have to run an extremely low quantized version of a lower parameter LLM. And so yeah, like if you don't have the hardware, using cloud is definitely going to give you better output and faster output than running locally, but you won't have all that ability to control a lot of those variables. So the quantization is continuing to advance, and with a platform like Hugging Face, which you might have been seeing recently, this is an ecosystem of people that are ma—taking the open-source LLMs, stripping them sometimes of their ability to decipher right from wrong, so you can even have uncensored versions of it now that you don't have on DeepSeek.com. Uh, they're also allowing you to fine-tune it and to create a data set that you can essentially fine-tune the local LLM on too. So there's a lot of stuff that we're going to be going over, but uh, let me break down what the research actually shows. And so here's an example of year 22, 23, 24, and open-source model performance and the size of the parameters and essentially the equivalent to um, you know, LLM at the time. So we have like the benchmark versus the normal one. So Bloom in year 2022 was 176 billion parameters, which is too big even at a 4-bit quantized level, broken down pretty far um, for like main consumers to be able to, you know, run. But then Llama 2 comes out at 70 billion parameters. Now if you have like, you know, decent computing, you can run Llama 2 at a 4-bit, maybe even 8-bit if you have like my computer. Quantization 8-bit is going to be better, just slower than 4-bit. Then you have year 24 where DeepSeek is now 67 billion parameters—um, well there is actually—don't get me wrong here because there's different parameter versions. So there's a seven—I think an 800 billion parameter DeepSeek—it's just not R1, but uh, they have essentially, obviously they have the different quantizations of them, but they also have the different parameter levels. So uh, DeepSeek R1 has like an 800 billion parameter one, and it also has a very smaller one. So I can't run that 800 billion parameter one at 4-bit quantization, but I can run let's say like the 70 billion one at a 4-bit quantization. So I know this a little hard, but stick with me here, and even ask AI these words—like I learned this from using Grok voice mode, going on walks all day, just asking what is the difference between quantization, GPU, all this stuff. So then from there, Llama 3 is now 70 billion parameters, and so you're seeing that these uh, open-source AI companies are fitting these into smaller sizes, and they are performing better at those sizes. Now compare that to hardware advances.

So the future is going to be running AI, and there will be a point where it is better to run local than cloud in terms of uh, speed and performance, I believe, and please correct me if I'm wrong if you're an expert in this, um, but this is based off of what I've been learning so far, as you guys can tell why I'm so excited and why we're going all in on this. Now this data tells an unmistakable story: the quality gap is effectively gone. And this isn't just theoretical benchmark performance; it translates directly to business applications. I've run systematic comparisons now across marketing content generation, customer service, data analysis, code generation, document—and in blind tests, the local models are indistinguishable from cloud versions in over 80% of the cases. You really cannot tell.

So now here's the hardware efficiency revolution, where people will say, "But don't you need massive GPU clusters to run these local LLMs?" Well, that's—this used to be true, but not anymore. The advances in model efficiency have been staggering in terms of the hardware required. In year 2022, Nvidia—this was running at 5 to 10 tokens per second; RTX 1520 tokens per second on seven billion parameters—the same variable across the board. A MacBook M2 in '24 can now run a 7 billion parameter model at 20 and 30 tokens per second, right? So this isn't marginal improvement; it's a fundamental shift in what's possible. Two years ago, running a good LLM required over $10,000 for a server, even renting GPU online from other people, specialized uh, machine learning knowledge, and constant maintenance. But today you can run a production quality model on a laptop you already own with one-click installation and zero technical experience. Now this trend couldn't be clearer. Here's the—here's where things really get interesting—the breakthrough that's enabling all this: quantization—that extremely confusing word I mentioned beforehand, which is the process of converting model weights from a higher precision to a lower precision format. Usually a 32-bit is like the highest essentially FP or PF, whatever it's called uh, precision format, and uh, okay, or awesome man, I'll see you. And then the precision formats are usually 8-bit, which people with better hardware are running, you know, a lot of these at 8-bit. That's going to give you that sweet spot, because if you look up the difference in uh, result or output versus speed um, from 4-bit to 8-bit, the drop off of speed is like not enough to compensate for the increase in quality. So it makes more sense if you have the hardware to run the 8-bit versus the 4-bit um, or even 3-bit for some of you guys that have, you know, lower hardware resources. 3-bit, I think there's even 2-bit, but uh, the results speak for themselves. Here's the precision or the quantization and the file size um, and memory usage, speed, and quality loss. So this is a 32-quant file size—like no one's really going to be able to run this unless you have uh, the—even the Mac I don't have—FP16, 70 GB; 70 GB. So I could run some FP16 models, maybe they're like, you know, uh, 10 billion parameters, whatever. Then you have uh, 8 quantization, 35 GB; 35 GB. Um, then you have 4 quantization file size—now uh, this is actually incorrect, guys, I think it is, so I don't want to correct—I don't want to have someone correct me, but I think for a quantization, I think the gigabytes are slightly higher than the memory usage for some of the quantization models, but I would look up online um, or use Grok to kind of compare your hardware and what you have to this. Um, and then we have the 4-bit quantization, which is this like GGUF—this is on Hugging Face. This guy named Barttowski will essentially quantize new models, and he'll offer—break them down in like different ones that some perform better on Macs, some perform better on PCs. And so this is an example of one that, you know, would just require 10 GB of memory usage, and it would only experience a 5 to 7% quality loss um, from, you know, running the uh, 0% quality loss at FP32. So this data—this is uh, just, you know, um, not actually statistically accurate, this is just like a representation of kind of what I've been learning, but um, this means a model that once required 140 GB of RAM can now run in less than 10 GB with minimal, but some, quality loss. And the innovation cycle isn't slowing down; it's accelerating. Now quantization techniques are emerging monthly, because everything's open source—like, you know, it's nothing's uh, centralized, everything's decentralized, people are going to go at their own paces, right?

So the multi-model strategy that's changing everything: one of the most powerful aspects of local AI that nobody talks about is the ability to use different specialized models for different tasks, and even doing a swarm-style—what you guys saw on my Instagram—of assigning a different LLM locally to a different task and running them in a swarm together. So that one is leveraging a different LLM versus another LLM; one's better at coding, one's better at writing, etc., or one's uncensored. Uh, with cloud AI, you're generally stuck with one model for everything, unless you're using Open Web UI, which I'll show you, and you can hop between models and compare them uh, even with cloud-based, but uh, with local development deployment, you can use a 3 billion parameter model for, like I was saying, simple tasks; a 7 billion parameter model for content and customer service. So like a different quantization level might be required for different tasks, and hopefully I'm not losing you guys here and you're fascinated by this, but this creates a massive efficiency advantage. Um, and then this is like—I helped a marketing agency, you know, implement a system; I've been testing this with a friend of mine, and we were able to get Mistral 7 to run for copywriting; DeepSeek for automation scripts; Llama for strategy development—this is the larger model—and they select the right task for each—the right model for each task automatically. The result: huge cost reduction versus running cloud AI, where you're going to be charged per API and per token. Uh, response time average—I'm going to have to delete this—like I said, this isn't like published anywhere yet. Um, 100% elimination of capability limitations—yeah, 100%—because uh, there's only certain uh, capabilities you're restricted to with the UI—running things on the cloud. But the selective deployment approach is something cloud APIs simply cannot match.

So the Hugging Face search data—everyone's missing—look at what's happening on Hugging Face, the GitHub of AI models, right? So this is where you can go and literally—one dropped yesterday—Meta Llama for Scout, 17 billion. I could literally install this right here. Here's uh, the quantization models of them are down here. So we got, you know, this guy has quantized them down into a version that you guys can essentially install—probably not Llama 4, it's pretty big—but I would be able to, or those with uh, like I think over 96 gigabytes of unified memory would be able to. So uh, this is the GitHub of AI models, as I was talking about, and it's even cooler—it kind of looks a little cryptoy. Um, and so here is the rates. And not only this, I've been comparing the search intent on Google versus AI automation, Zapier, Naden, and things like Olama LLM and Hugging Face, and there's like a—like I said, you guys are early if you're watching this, even if you're watching the recording uh, from here. January 2023: 1.2 million downloads of LLMs; 2024: 37 million downloads of local LLMs; 2024 April: 108 million downloads of local LLMs. Most popular models: DeepSeek, Llama 3. Before here, two years ago, I mean I don't even know GPT-J; I don't know anybody that was running this. But this shows you that the adoption curve isn't linear; it's exponential. And the search trends tell an even more interesting story. The top five uh, Hugging Face search trends are: quantization, local LLM, RAG, fine-tuning, and multi-model—what have you guys seen on my uh, in this video already and on my Instagram? Quantization, local LLM, RAG—which is essentially being able to interpret images—image to text instead of text to text; fine-tuning, where we can train these open-source models with our own data sets that are uh, our own IP without anyone—China—getting access to it; um, and then multi-model swarm—running, you know, multiple agents with different LLMs. Now this isn't just research in research or interest; it's developers and businesses actively implementing these tech—this tech now.

The cloud AI limitations that no one is admitting: let's talk about the limitations of cloud AI that most vendors won't acknowledge. Number one: the censorship reality. So you're probably asking still like why, you know, are you trying to convince me to run locally? What, you know, incentive do you have? Well, every business eventually hits what I call the capability wall with cloud AI. Here are some real-scenario businesses encounters daily: Analyze our competitors' market approach. Cloud AI: I cannot help with competitive analysis that could be used to… Local AI response: provides detailed analysis. Write persuasive sales copy for our weight loss program. I'm sorry, but I can't create content promoting weight loss as this falls into a sensitive category. Response on the local: creates effective sales copy. Help us optimize our pricing strategies to maximize profits. Now I don't know if that's legitimate that ChatGPT is giving you this answer, but I'll tell you one time I've tried to train Claude on creating extremely, you know, specific copy that speaks to the pain points and situations in a way that is cultlike, and all of a sudden I can't get it right. So what happens when I run a local LLM that's uh, uncensored? It's giving me it. So these aren't edge cases; these are everyday businesses—business needs that cloud AI increasingly refuses to address, and it's only going to get worse with censorship—being able to have Google and these China control, you know, what your AI is doing, right? So problem two is the cost structure that punishes success. Cloud AI pricing models are broken; it'll run you up, especially if it's involved in your workflow, as you can see right here—it's, you know, gets pretty—pretty expensive to run at an enterprise scale. Some API solutions—as your AI usage grows—so not only as integrators, you're going to be able to allow like cut costs from businesses—as your AI usage grows, cloud costs scale linearly while local costs remain nearly flat, except your hardware and your uh, electricity that you're using. This creates a—a perverse incentive: the more successful your AI implementation, the more punish—you're being—the more you're being punished with higher costs, even though your AI is getting more successful. Problem number three is the compliance nightmare for regulated industries. As I was mentioning, cloud AI—you really can't even use it. HIPAA—healthcare patient data and prompts violates PHI rules; you could lose your entire practice that's worth $8 million if a social security number goes through DeepSeek's API from an AI automation that all of you guys, including myself, we're building inside NADN, inside of Zapier. And I'll be able to show you guys—you can run local LLMs inside of NADN; it's—it's very easy, and you're not breaking HIPAA laws. Local AI never leaves your infrastructure. Finance—PCI compliance; you know, some financial guys you might be working with, they have data that cannot be sent to third parties, right? Legal—NDAs, client confidentiality. I just signed an NDA the other week, and if I break it, I'm liable to, you know, some pretty concerning things, right? So cloud AI—attorney-client privilege concerns with third-party processing, whereas local—full control of the information in the flow. Uh, education—FERPA—student data, right? Like for these institution—at institutional levels where like the real money is being made, guys—student data protections are incompatible with cloud terms. If you are helping any business or you are a business in these industries and you're setting up AI automations inside of—with the cloud, like you could have your personal asset seized. If you guys—especially if you guys aren't structuring your business right—like that is the result—is having your personal assets seized because you broke an NDA and somebody out there was harmed through their patient data being sent to China, and someone just lost their $8 million practice, and now your personal assets are being seized because of it. Well, one healthcare client had to abandon a $300,000 cloud AI implementation after their compliance team determined it would violated patient data protection requirements. They switched to a local deployment for less than 50k all-in. I was telling Josh—you know, if you guys—you got a big company—if you guys are running cloud AI—no, you guys have enough resources—get a supercomputer inside of the office that is running the LLMs and everybody has access to them and it's running them on the local network, right? And that's going to save them a ton of cloud AI implementations, and it's going to actually, you know, not violate any NDAs or anything that's going on with the clients that they have. Wow, okay, so we're getting into this—we got—let's see—I know you guys probably uh, like this—maybe this will be the YouTube video too—but the model size here—the trend that everyone needs to understand is that a—the critical trend that directly contradicts what cloud AI providers want you to believe—they literally want you to believe these things—um, that bigger models aren't always better for specific tasks—because like that's—that's not true—that bigger models aren't always better for specific t—or—or no, that is true—that bigger models aren't always better for specific tasks, because some models—if you're running a 70 billion parameter 8-bit on local and I ask it to do something simpler for me or write some type of copy that isn't like in-depth research client work—it's going to overcompensate, and it's going to ruin the copy. So even if you guys are thinking, "I don't have the right hardware," no—if you pick the right smaller model, it might perform better at certain tasks, and you know the cloud versions that you're using that are overcomplicating some things. So the research is clear: specialized smaller models often outperform general larger ones for business use cases. Here's an example of the model sizes: from 70 billion to 3 billion; the performance: excellent to adequate; the specialized task performance: good, excellent, excellent, very good; when specialized—inference cost: very high, moderate, low, very low—of the amount of computing resources it'll take. And this 3 billion model—can—let's say at an 8-bit—you could probably run 3 billion at 8-bit—for some of you guys, this could actually return you very good copy when you specialize it. And this is where those—these uh, settings that you guys have probably experienced that you never played around with on your AI—of temperature and like output tokens—all of that matters extremely with local AI, because like if you run—I heard—um, one of the QWQ 32 billion and you use it without a temperature and these different settings, it's going to look like horrible, but if you do it to the right settings—which I'm going to get those settings for y'all—that it—it runs amazing, even better than on the cloud. So this explains why we're seeing so many businesses now—or honestly, they're not—and why you guys are on here—adopting a multi-model approach—it's using smaller specialized models for routine tasks. And if you guys are a systems integrator—which is why we went through this phase of going from automation expert to systems integrator—because eventually you're going to be commoditized just building automation—so documenting the process was the most important part, which now plays into this local LLM stuff, because if you've already documented the process, you know what team members are doing what, you know which AI agent and which model to build a multi-model approach for and uh, all run it locally on your hardware. Um, this also means that using midsize models for general work is normal, reserving larger models for only complex reasoning. Let's say I'm running a full on—I'm taking the automated OS canvas and I'm created five different agents with five different 70 billion parameter LLMs that are better at documenting their process—one's better at identifying opportunities, right? Like I can build a multi-model approach with the Automatos canvas, the process I've documented, and I could essentially replace this with AI agents now. Cloud providers have a financial incentive to push you towards their largest and most expensive models for everything. I don't know about you guys, but I'm now paying $30 a month for Grok 3; I'm paying $30 a month for ChatGPT; I'm paying $30 a month probably for Claude; and now like even ChatGPT had that $200 a month version coming out, and I think even if you compare the $200 a month when they had just…

Uh, I think it was 40, or some type of 01. Like, if you look at, um, artificial analysis.ai AI, I'm pretty sure that locally running one of these LLMs is going to produce better results than paying $200 a month for ChatGPT. Right? So they have an incentive to push towards you, which is why you're now just hearing this, and why I was just now hearing this. Um, but some of you guys in here that were in on this early, like, we're going to need to rely on you guys for help in the community. So, uh, with local deployment, you can match the model to the task, creating massive efficiency gains. The right model for the right task—matrix here it is. One of my most important realizations was that different models excel at different business functions. So here are some examples: customer support, content creation, data analysis, code generation, document processing. Right? Document processing is going to be an image-to-text; there's different LLMs that are better at image-to-text. Code generation—there's, uh, QWQ 32 billion, there's QWQN code 2.5 coder, uh, Gemma 3 27 billion coder, uh, so some of these code Llama, DeepCoder; these are code-specialized. These range from anywhere from three to, or seven to 13 billion. So all of these have different LLMs that we're going to have to identify through us as a community—of essentially which LLM is performing the best and run comparisons of them in real time. Right? So the—because that could be the difference between you getting a client amazing result and a horrible result—is not only picking the LLM for the task, but also the temperature in these different stats. So like, guys, we're stepping into a whole new world here, which is also why I increased the price recently and let you guys know ahead of time, uh, for our school community because we're on year three now, and this is getting pretty insane.

So this selective deployment is impossible with most cloud services, where you're locked into one model for everything. The ability to select the right model for each specific task is a game-changer, as you know now, for effectiveness and cost. So there are two categories of an AI-powered business that are emerging as this shift accelerates. I'm seeing businesses separate into two distinct categories, and you have two options here: you can either be an AI sovereign. So, uh, so sovereign, where these organizations run their own models, customize the capabilities to their specific needs, deploy specialized models for different tasks, and maintain complete data control. This is where they're getting 70 to 99% lower AI operating costs, running at complete freedom without restrictions, full data and privacy and security, ability to fine-tune to their specific domain, customized capabilities that their competitors cannot match. And then there's going to be the AI dependent. These organizations remain tethered to cloud providers; they pay per token fees that scale with usage, and they're contending with increasing restrictions, and they're sending their sensitive data to third parties. This is 99.9% of people, and I'm sure you're seeing as well with a lot of the, uh, AI automation experts that are emerging; this is what they're building, and they're making their clients AI dependent. And so this used to be us, but as this stuff is emerging, we're going to become AI sovereigns. And so if you are an AI dependent, you're just going to experience continuously rising AI costs as the models get better and they, uh, release new ones, and, uh, inflation gets better, whatever the economy crashes, stuff gets more expensive, who knows. Growing frustration with capability limitations, increased workarounds for restrictions, vulnerability to pricing changes—this you don't want to be this guy.

And so the advanced systems being built today, while most businesses are still trying to get cloud AI to approve their marketing copy, forward-thinking companies are building remarkable systems with local, uncensored AI. Now here is some examples of a multi-agent architecture with local LLMs, and this also applies to cloud cloud AI; you could still build a multi-agent architecture, it's just not going to—you're still going to be an AI dependent. Um, but multiple specialized AI agents working together, as we said, with different models, different settings, different system prompts in the same script, uh, or in the same process. This is going to look like this: you know, a customer query comes in, there's one agent that classifies this query, and then this classified agent then sends it to one of five specialized agents that work in a process together, and then it generates it and sends what they've been working on to the response agent, who then comes up with, like, the final response. Right? And so if you had one long prompt, you would be restricted to, you know, one model. If you have, uh, one AI automation with multiple, let's say cloud, you would be restricted to one set process. But with the architectures that we have here soon, in the new research that's being done, essentially you can allow—with the video I made about swarm systems—is to where they will decide and choose who's, uh, you know, completely for you about what's going to happen between the customer query and the customer getting the response back—is that there's going to be a multi-agent framework. And if you run locally and you pick and you identify those LLMs, the parameters, the quantizations, and the settings for them, you're going to have technology no one can copy. Now these systems handle complex workflows with no per-token cost, too, if you're running them locally, and complete customization.

Now this might not be the best idea, but if I wanted an entire swarm inside of my Mac M3 Ultra to be running at 80 gigabytes of total unified memory for 24/7, it will cost me nothing. I could do that; it will cost me maybe an extra, you know, $10, $20 a month in electricity, but I could technically have AI agents running locally 24/7. Um, now, uh, this is another—let's see here—another advanced system being built today is going to be fine-tuned domain specialist models that are customized to specific business domains that understand industry-specific terminology, company product catalogs, brand voice and style, historical customer interactions, and competitive landscaping. Because if you are building some type of framework, it's going to have to—like with how far on a linear level we're breaking down the process as we do in systems integrator—you're going to have to do the same with—if you're offering this as a solution to other businesses or in your own business—it's going to have to be unique and built around your IP. And so niching is more important than ever with local, um, and these frameworks, these models achieve performance in specific domains that general models cannot match. Okay, so now another thing is reasoning-enhanced systems. So you're even seeing local LLMs, not just the normal ones, but the reasoning ones that think like Chat and Grock, uh, research mode before it returns you output; those are now also available as and have been as local LLMs. Advanced implementations combining multiple models for optimal performance, to where they have the ability to reason. So now imagine integrating reasoning models, too, with that multi-agent approach. And so I know we're going on here, but this almost done here, but, um, this is going to be the local AI adoption curve; we can map exactly now where we are. Uh, we have early adopters, early majority, late majority, laggers. Right now, like, we are here; the majority of people are here, and, um, like our goal is to become an early adopter obviously, but a lot of people are right here. And so this means that we're at the transition from early adopters to early majority, and, uh, this is precisely when the most significant competitive advantages are emerging. This is where your knowledge and expertise is going to be the most valuable; implementation is going to be accessible, but it hasn't yet been commoditized; you can't find somebody on Fiverr to do this for you. The cost-benefit ratio is extremely favorable; you're going to remove 10,000 API costs, and you're going to bring all those benefits to the business, right? And you're going to give them, you know, how to set up the hardware, and they don't even need to learn all the quantization, all that, you know, stuff, Python, to to install this stuff, all that; they're not going to have to learn; you build a protocol around it, right? And so by next year, guys, this will be mainstream; I'm afraid to tell you it will, and the window for significant advantage is right now.

So, uh, here's the implementation roadmap that actually works, based on multiple successful implementations. Here's the approach that's going to deliver you results. Phase one: model selection and testing. In week one, you're going to deploy Ollama or a similar local runner; you're going to test three to five different models on your specific use cases; you're then going to benchmark it against your cloud solution; then you're going to identify the optimal models for different tasks. Then, in phase two, weeks two to three, this can be initial implementation: you're going to deploy selected models for one specific business function, document its performance, costs, and capability improvements, train your team members on it, and create a standard operating procedure over it. Then, in months one to two, phase three: scale deployment; you're going to expand this to additional business functions, implement multi-model architecture where appropriate—don't hop right into that like I did—integrate with existing tools and workflows; you're going to find ways to add APIs, MCPs, all of the new stuff that's coming out; everyone's going to be doing it on the cloud, you're going to be doing it locally, and you're going to develop monitoring and management processes for it to make sure your computing is good; all that stuff is dialed in. Then phase four, month three: advanced optimization; you're going to explore fine-tuning for domain-specific performance; this is where we get really complicated, and we're going to be getting into this, guys, um, in in this program, um, and then implement RAG; this is where you're more advanced stuff of, uh, you know, taking documents and, you know, integrating that as well. Then develop, uh, you're also going to develop proprietary systems for competitive advantage that you're going to even license, and maybe you even want to open source that. I built a a framework inspired by the swarms and Crew AI framework; I made open source on GitHub that you guys have access to now as well, and that would be an example of, you know, something that could, if it was, you know, really something amazing, could give me a competitive advantage. And then you're going to create sustainability and scaling plans for this. So this approach minimizes risk while maximizing the speed to value. I wouldn't expect to learn all this overnight unless you're talking to Grock. Um, and so yeah, I think, uh, we yeah, the best thing to know now would be your hardware and what model you can run on your hardware. So definitely identify that; most businesses can run effective AI systems on hardware they already run, but for those requiring an upgrade, the costs are minimal, uh, compared to ongoing cloud fees. If you compare buying the brand new Mac versus, you know, what you're going to pay in five years in cloud AI fees. So, um, although the ROI isn't measured in years, it's measured in months or even weeks. Um, so I think that this is this is, uh, uh, good. I think the rest of this is essentially, um, well, yeah, let's end with this: the future is already here; it's just unevenly distributed, as William Gibson famously said. The future is already here; it's just evenly distributed; that's exactly where we are right now with local AI. Some businesses are already operating with capabilities that their competitors believe are impossible; some teams are already creating content, serving customers with tools that others won't discover for months or years. And this technology is available; it's accessible; the results are proven. So that is my VSSL sales letter, whatever you want to call it; it was more educational, really wasn't trying to sell anything in that; that's the whole goal of this is it's open source. So here's some resources to be important to know for this as well: artificial analysis.ai, real-time data and comparison of cloud and local LLMs; Hugging Face—this is where you're going to get your local models; you're going to even get your data sets; you can fine-tune them on and even test out, uh, apps built by other open-source users with these LLMs. Then you're going to have Ollama, where you're going to access all of your available models and different quantization levels that you can directly download and install in your terminal on your computer. You're also going to go inside of our school; let people know what's going on, what you're learning, because I can't do this all myself. So, uh, you guys educate each other and share what you're doing; make Loom videos; make share the open source; that's why this is open source; share all of it. Um, if you run on Apple—something I'm learning right now, I'm going to be documenting—is running it on an MLX framework, meaning you're going to get double the amount with the same quantization level of response time if you're running on Apple Silicon; this is a framework built by MLX. Then, if you're interested in multi-agent frameworks and swarms, there is a GitHub available by Kai Gomez called Swarms, which is his like research documentation; this is also associated to a cryptocurrency; this is not, uh, investment advice, nor do I even, uh, would say invest into this; this is simply a research framework, uh, that you can pull from. Then we have Crew AI, which is very similar and not crypto-based; more like, uh, Kai kind of came up with this from Crew AI. Then we have Open Web UI, which is going to be the front end that you can run your LLMs in that you can connect something like Naden pipe, Blender render, like all these—imagine like the tools in ChatGPT; you can take any, even if it's cloud or local, you can have your own ChatGPT, which I'll even show you; I've put inside of as a custom menu link in HighLevel, right? So if you're using HighLevel, boom, no more going to ChatGPT; you got Anthropic Claude, NAD pipeline, GPT-40, then all my local LLMs locally hosted, running on Web UI. Not only that, we have Naden also now locally hosted; you're going to be making a video on that integrated right into GoHighLevel as a custom menu link. So definitely check out Open Web UI; they also have GPTs that you can build with your local LLMs or your cloud-based LLMs, and there's a ton of stuff there on Open Web UI that I am now using compared to chatbots when I first got started and was showing you guys. As I mentioned, at we're going to be running locally and, uh, learning that. Then we have, um, Google Gemini 2.5 Pro, that until you're running locally, you're learning everything about local, use cloud AI to learn everything about local, to set up the local and use Gemini 2.5 Pro, as it now performs ex-way better on the last human on Earth exam, humanity's last exam; usually the delta was 2%; every single increase now Gemini 2.5 Pro is, uh, performing 17% accuracy on humanity's last exam; very interesting research that if you don't know what that is. And then another cool thing is like, yeah, while everyone else is stuck generating images on ChatGPT, thinking that's going to be the future, like, I'd stay away from that; if anything, use that as, you know, interest to say how can I go into Hugging Face and run an image generation model locally, and then from there. Yeah, that is everything that I have today; that's why kind of I want to make a document because it's so much; it's too many resources; like it would take me a month to probably bring a course to you guys in our group going over everything, and I know a lot of you guys just recently joined for this. So this is also recorded, and I will make all of this document available, and I'll, uh, do that. But yeah, I guess, um, I enjoyed presenting that to you guys; hopefully you learned something from that. Now we'll open the floor for discussion for around 10 minutes if anyone wants to ask a question, comment, or give feedback.