📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Deepseek R1 Explained by a Retired Microsoft Engineer

Dave's Garage10:07

Transcription

Hey, I'm Dave. Welcome to my shop. I'm Dave, a plumber and a retired software engineer from Microsoft, going back to the MS-DOS and Windows 95 days.

Today, we're tackling a seismic shift in the world of technology: the release of China's open-source AI model, Deep Seek R1. This development has been described as nothing less than a Sputnik Moment by Mark Andreessen, and for good reason. Just as the launch of Sputnik challenged assumptions about American technological dominance in the 20th century, Deep Seek R1 is forcing a reckoning in the 21st.

For years, many believed that the race for AI supremacy was firmly in the hands of established players like OpenAI and Anthropic. But with this breakthrough, a new competitor has not just entered the field; they've also seriously outpaced expectations. If you care about the future of AI innovation and global technological competition, you'll want to understand what Deep Seek R1 is, why it matters, whether it's just a giant scoop, and what it means for the world at large.

Let's dive in. To set the stage, here's the part that really upset the industry and sent the stocks of companies like Nvidia and Microsoft reeling. Not only does Deep Seek R1 meet or exceed the performance of the best American AI models like OpenAI's GPT-4, they did it on the cheap—reportedly for under $6 million.

When you compare that to the tens of billions already invested, if not more, to achieve similar results—not to mention the $500 billion discussion around Stargate—it's cause for alarm. Because not only does China claim to have done it cheaply, but they reportedly did it without access to the latest Nvidia chips. If true, it's akin to building a Ferrari in your garage out of spare Chevy parts.

If you can throw together a Ferrari in your shop on your own, and it's really just as good as a regular Ferrari, what do you think that does to Ferrari prices? So, it's a little bit like that.

And just what is Deep Seek R1? It's a new language model designed to offer performance that punches above its weight. Trained on a smaller scale, but still capable of answering questions, generating text, and understanding context.

What sets it apart isn't just the capabilities, but the way that it's been built. Deep Seek is designed to be cheap, efficient, and surprisingly resourceful, leveraging larger foundational AIs like OpenAI's GPT-4 or Meta's LLaMA as scaffolding to create something much larger.

Let's unpack that. Because at its core, Deep Seek R1 is a distilled language model. When you train a large AI model, you end up with something massive—hundreds of billions, if not a trillion parameters—consuming terabytes of data and requiring a data center's worth of GPUs just to function.

But what if you don't need all that power for most tasks? That's where the idea of distillation comes in. You take a larger model, like GPT-4 or the 671 billion parameter behemoth R1, and you use it to train the smaller ones. It's like a master craftsman teaching an apprentice. You don't need the apprentice to know everything, just enough to do the actual job really well.

Deep Seek R1 takes this approach to an extreme by using larger models to guide its training. Deep Seek's creators have managed to compress the knowledge and reasoning capabilities of much bigger systems into something far smaller and more lightweight. The result? A model that doesn't need massive data centers to operate. You can run the smaller variants on a decent consumer-grade CPU or even a basic laptop, and that's a game changer.

But how does this work? Well, it's a bit like teaching by example. Let's say you have a large model that knows everything about astrophysics, Shakespeare, and Python coding. Instead of trying to replicate that raw computational power, Deep Seek R1 is trying to mimic the outputs of the larger model for a wide range of questions and scenarios.

By carefully selecting examples and iterating over the training process, you can teach the smaller model to produce similar answers without needing to store all that raw information itself. It's kind of like copying the answers without the entire library.

And here's where it gets even more interesting. Deep Seek didn't just rely on a single large model for the process; it used multiple AIs, including some open-source ones like Meta's LLaMA, to provide diverse perspectives and solutions during the training. Think of it as assembling a panel of experts to train one exceptionally bright student.

By combining insights from different architectures and data sets, Deep Seek R1 achieves a level of robustness and adaptability that's rare in such a small model. It's too early to draw very many conclusions, but the open-source nature of the model means that any biases or filters built into the model should be discoverable in the publicly available weights.

This is a fancy way of saying that it's hard to hide that stuff when the model is open source. In fact, one of my first tests was to ask Deep Seek what famous photo depicts a man standing in front of a line of tanks. It correctly answered the Tiananmen Square protests, the significance of the photo, who took it, and even the censorship issues surrounding it.

Of course, the online version of Deep Seek may be completely different because I'm running it offline locally, and who knows what version they get within China. But the public version that you can download seems solid and reliable.

So why does all this matter? Well, for one, it dramatically lowers the barrier to entry for AI. Instead of requiring massive infrastructure and your own nuclear power plant to deploy a large language model, you could potentially get by with a much smaller setup.

That's good news for smaller companies, research labs, or even hobbyists looking to experiment with AI without breaking the bank. In fact, I'm running it on our AMD Threadripper that's equipped with an Nvidia RTX 680 GPU that has 48 GB of VRAM, and I can run the very largest 671 billion parameter model. It still generates more than four tokens per second, and even the 32 billion version runs nicely on my MacBook Pro.

The smaller ones run down to the Aura Nano for $249. But there's a catch. Building something on the cheap has some risks. For starters, smaller models often struggle with the breadth and depth of knowledge that the larger ones have. They're more prone to hallucinations, generating confident but incorrect responses sometimes, and they might not be as good at handling highly specialized or nuanced queries.

Additionally, because these smaller models rely on training data from the larger ones, they're only as good as their teachers. So if there are errors or biases in the large models that they train on, those issues can trickle down into the smaller ones.

Then there's the issue of scaling. Deep Seek's efficiency is impressive, but it also highlights the trade-offs involved. By focusing on cost and accessibility, Deep Seek R1 might not compete directly with the biggest players in terms of cutting-edge capabilities. Instead, it carves out an important niche for itself as a practical, cost-effective alternative.

In some ways, this approach reminds me a bit of the early days of personal computing. Back then, you had massive mainframes dominating the industry, and then along came these scrappy little PCs that couldn't quite do everything but were good enough for a lot of the work. Fast forward a few decades, and the PC revolutionized computing.

Deep Seek might not be GPT-5, but it could pave the way for a more democratized AI landscape where advanced tools aren't confined to a handful of tech giants. The implications here are huge. Imagine AI models tailored to specific industries, running on local hardware for privacy and control, or even embedded in devices like smartphones and smart home hubs.

The idea of having your own personal AI assistant—one that doesn't rely on a massive cloud backend—suddenly feels a lot more attainable. Of course, the road ahead isn't without its challenges. Deep Seek and models like it must prove that they can handle real-world tasks reliably, scale effectively, and continue to innovate in a space dominated so far by much larger competitors.

But if there's one thing we've learned from the history of technology, it's that innovation doesn't always come from the biggest players. Sometimes, all it takes is a fresh perspective and a willingness—or sometimes the necessity—to do things differently.

Deep Seek R1 signals that China is not just a participant in the global AI race, but a formidable competitor capable of producing cutting-edge open-source models. For American AI companies like OpenAI, Google, DeepMind, and Anthropic, this creates a dual challenge: maintaining technological leadership and justifying the price premium in the face of increasingly capable, cost-effective alternatives.

So what are the implications for American AI? Well, open-source models like Deep Seek R1 allow developers worldwide to innovate at lower costs. This could undermine the competitive advantage of proprietary models, particularly in areas like research and small to medium enterprise adoption.

U.S. companies that rely heavily on subscription or API-based revenue could feel the squeeze, potentially dampening investor enthusiasm. The release of Deep Seek R1 as open-source software also democratizes access to powerful AI capabilities. Companies and governments around the world can build upon its foundation without the licensing fears or restrictions imposed by U.S. firms.

This could accelerate AI adoption globally but reduce demand for U.S.-developed models, impacting revenue streams for firms like OpenAI and Google Cloud. In the stock market, companies heavily reliant on AI licensing, cloud infrastructure, Nvidia chips, or API integrations could face downward pressure as investors factor in lower projected growth or increased competition.

Now, in the intro, I made a little side reference to the potential of a scoop angle. While I'm not much of a conspiracy theorist myself, some have argued that perhaps we should not take the Chinese at their word when it comes to how the model was produced. If it really was produced on second-tier hardware for just a few million dollars, it's major.

But some argue that perhaps China invested heavily at the state level to assist, hoping to upset the status quo in America by making what is supposed to be very hard look supposedly cheap and easy. But only time will tell.

So that's Deep Seek R1 in a nutshell: a scrappy little AI punching above its weight, built using clever techniques and designed to make advanced AI accessible to more people than ever before. It's not perfect; it's not trying to be. But it's a fascinating glimpse into what the future of AI might look like—lightweight, efficient, and a little rough around the edges, but full of potential.

Now, if you found this little explainer on Deep Seek to be any combination of informative or entertaining, remember that I'm mostly in this for the subs and likes. So I'd be honored if you consider subscribing to my channel to get more like it.

There's also a share button down in the bottom here. Somewhere in your toolbar, there'll be a forward icon which you can use to click on to send this to somebody else that you think probably wants to be educated and just doesn't know about this channel.

So if you want to tell them about Deep Seek R1, send them a link to this video. If you have any interest in matters related to the autism spectrum, check out the free sample of my book on Amazon. It's everything I know now about living your best life on the spectrum that I wish I'd known long ago.

In the meantime, and in between time, hope to see you next time right here in Dave's Garage. Do it. Do it. Do it.