Transcription
Nvidia just released an absolutely game-changing supercomputer, and here's the thing: it's tiny, very cheap, and has the potential to deliver large language models to the edge more than anything else that we've seen. This means open-source models running independently anywhere.
Let me just have the Nvidia CEO explain. Watch this:
"Hi, welcome to my house! I'm living in a different house now. We're fixing the house that you guys saw the last time you were in my kitchen. Let's see, what were we doing? My hair was a lot longer, and I lifted a brand new HGX out of my oven while I'm cooking something up for you again today. Let me show it to you.
Okay, here we go, ladies and gentlemen, our brand new AI computer! Look at this! I think I might have cooked it a little bit too long; it shrunk the little tiny Jetson Nano, little Orin computer.
This thing, that's really amazing, is that a long time ago, starting with Xavier, you guys might have known that we created a brand new type of processor. It was a robotics processor; nobody understood what we were building at the time. We imagined that someday these deep learning models would evolve, and we would have robots of all kinds—everything that moves would be robotic.
And now here we are. We're seeing all kinds of amazing robots: robots on wheels, robots on legs, two legs, three legs, and of course, general humanoid robotics are nearly upon us.
This is a brand new Jetson Nano supercomputer—almost 70 trillion operations per second, 25 watts, and $249. It runs everything that the HGX does; it even runs large language models. I can't wait for all of you to try it. It's available everywhere; go get it, enjoy robotics!
It runs CUDA, cuDNN, and TensorRT. You could create an agentic AI that reasons and plans, so you can use it for a robot, you can use it for a workstation—it's an incredible computer! What do you guys think?
So, this is the Nvidia Jetson Orin. This is a micro supercomputer; it is blazingly fast at running inference on the edge. The edge means that it's not connecting to the cloud. You know that I'm very bullish on edge compute. I've talked about it in the past: chips in our phones, chips on our computers, and chips everywhere get better and better.
As large language models continue to get better and smaller, we're going to be able to run artificial intelligence pretty much anywhere. And now, with the Jetson Orin, this tiny supercomputer that is incredibly inexpensive, we are going to be able to put large language models into essentially anything—robots, Internet of Things devices, cars—literally anything this little device can just plug into and run inference using local models.
This tiny supercomputer can run 275 trillion operations per second. That is an insane number!
Thanks to Emergence AI for partnering with me on this video. If you've watched this channel at all, you know I'm bullish on agents, especially agents that can accomplish real-world tasks for you. That's why I'm excited to tell you about Emergence AI. Emergence AI just launched their enterprise-grade multi-agent orchestrator, and they showed off the first demo of a real-world use case where these agents can actually browse the web on your behalf.
This is enhanced web automation. This means multiple agents can go out and dynamically interact with different elements on the web under this intelligent orchestration to bring human-like interaction and navigation but at machine-level scale.
What's really cool about this is these agents can actually perform complex and sophisticated web interactions that previously required a human. They can navigate dynamic, early-loading menus, fill out forms, adjust settings, process embedded files, and extract relevant data from PDFs and HTML.
Emergence AI orchestrator offers a combination of design-time flexibility and runtime determinism. That basically means that these agents can heal themselves. So if it makes a mistake along the way, it can figure it out and then succeed on the next try.
Emergence AI has put a huge emphasis on privacy and security. They offer a fully hosted solution where you can access it via API, or you can host it on your own virtual private cloud. So if you're an enterprise business and you're looking to automate a lot of your processes, Emergence AI is a great solution. Integrate their agent API and seamlessly orchestrate multiple agents to accomplish tasks for you and your business.
This includes interacting with both modern and legacy enterprise applications. Emergence AI is just starting to invite developers to try out their platform. Definitely check them out; tell them I sent you. Of course, go to their website emergence.ai or simply email them at contact@emergence.ai. I'll drop all of the links in the description below.
So thank you again to Emergence AI for partnering with me on this video. Now, back to the video.
This tiny supercomputer has the potential to power the robots of the future, and those robots won't need to actually connect to the cloud to do their inference. Everything is going to be on-device. I could not be more excited about that future where we have more control over the large language models that we're running—more security, more privacy—everything on-device.
Look, there's always going to be a time where we need massive models like GPT-4, but if you've watched my channel at all, you know I believe that 99% of use cases can be accomplished by models that are 7 billion parameters or less. That is a small model and an extremely capable model, and now we have the device to run it.
It's kind of like the Raspberry Pi in a lot of ways. It is a very small, very modular device that can be built on and customized any which way you want. This device can run almost 70 trillion operations per second, and look at the form factor—it is tiny!
This is a nearly 2x improvement over its predecessor, and here's the killer feature: it's only $249! That is extremely inexpensive. Imagine stacking four of these up for $1,000, then you're running nearly 300 trillion operations per second, and it only draws up to 25 watts, which is so efficient.
It uses the Nvidia Ampere architecture with 1,024 CUDA cores and 32 tensor cores. It has a six-core ARM Cortex CPU, this is 64-bit, 8 GB of memory, but very, very fast memory—102 GB per second. It has an SD card slot, so you just load up the operating system you want easily. And again, all of this under $250—that is the killer innovation in my mind.
Let's take a look at the performance. So first, let's look at the large language models. We have the baseline of the previous generation, and we have all of these open-source models from Llama 3.1, Llama 3.2, Qwen 2.5, 3.5—even though 54 is out now—all of them perform extremely well.
It can power vision models; it can power vision transformers as well. This is the stuff of the future! Throw one of these into your car, and all of a sudden it can autonomously drive. Throw it into any of your Internet of Things devices that you have at your house, and all of a sudden they are powered by a large language model and become insanely intelligent.
Put it into any robot, and all of a sudden the robot becomes powered by a large language model or a vision model or anything. These devices just churn through the math required to run intelligence. I am so impressed with this release!
I plan on getting one myself. Should I create a tutorial showing how to set it up and get it running? Let me know in the comments.
Super impressive! Congrats to Nvidia on this release. If you enjoyed this video, please consider giving a like and subscribe, and I'll see you in the next one!