📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Predicting the next 5 Years of AI Agents with Nvidia’s AI Chief

The Next Wave - AI and the Future of Technology45:00

Transcription

Hey, welcome to the NextWave podcast. I'm Matt Wolf, and in today's episode, we're talking about AI agents. I had the opportunity to go spend a few days with Nvidia out at the Nvidia GTC conference.

And in this episode, I'm going to deep dive with Amanda Saunders from Nvidia. She's been at the heart of the AI agent revolution that's happening right now while working with Nvidia, the company that's enabling the AI revolution.

In this episode, we'll unpack exactly what Agentic AI is and how it's reshaping industries from healthcare and telecom to sports coaching, as well as why that's just scratching the surface of what's possible. You'll also discover NVIDIA's secret blueprint for easily building powerful AI agents, as well as get into some of the fears that people have around AI agents, like, "Is an AI agent going to take your job and replace you?" Yeah, we're going to go there.

And after the interview, I'm going to share some cool clips from Nvidia's GTC where I got Bob Petty, also from Nvidia, to break down exactly what Nvidia's DGX Spark is and how they're going to be putting AI supercomputers in normal people's homes. You can have an AI supercomput in your home by this time next year. We're going to get into that in today's episode.

So, without further ado, here's my discussion with Amanda Saunders, followed by my tour of the DGX Spark with Bob Petty. Enjoy.

Let's talk about Agentic AI, cuz that's a hot topic. It seems like uh, in 2025, everybody's saying this is the year of agentic AI. How would you define agentic AI?

So, agents are digital employees that help us augment our work. And what's really important about agents and and agentic AI is that it can perceive. So it sees the world around it. It sees the data that it has access to. It takes that information. It reasons about that information. It thinks about it. And then it can actually make actions based on that data. So it can perceive, it can reason, and it can act. And that's what really sort of sets it apart from generative AI that we had in the past is being able to take those actions, whether that action is alerting a human or actually actively using a tool and and making uh, you know, something happen.

What do you think makes agentic AI like powerful right now? Why why not a year ago? Why not two years ago? Why is now the time everybody's talking about it?

Well, it started with having to have the accelerated computing. You had to have the computing power to do that. And that's a problem that Nvidia has been working on for 30 years. And finally, we've got to the place where we had the computing systems. We also then needed the software systems. And so it started with the open models that became available. I think Llama was a was a huge advancement and step forward, and more and more models have come out recently. Reasoning models, in fact, are one of the the big pieces that have sort of driven that and and sort of added to that.

Yeah. And then just the the full ecosystem of tools and things that are required. AI is one of those really interesting fields where the more people use it, the more they want to use more of it, and they want to do more. And so it's really spurring this incredibly fast growth that we're seeing, which is, you know, breakneck speed. But that's what's really driving us from, you know, the early onsets of generative AI where you're just spitting out answers into these worlds now where they're actually thinking things through and making really smart actions.

Yeah. Or in my case, I got into AI, and now I just make content about it because I'm obsessed with it. That's all I ever want to talk about.

Because yeah. Once you get into it, you really that's all you want to do is is you're you figure out what it can do, you learn where it's you know boundaries are, and then you want to do more and more and more, and as it continues to, you know, get better, you can do more and more things with it.

So cool. Well, can we walk through like a agentic workflow? Like give some examples of like this is the types of things an agent can do.

Yeah, absolutely. I think, you know, some of the most basic things that we see from agents are really about being able to talk to data. I think that's probably the first thing that we see most people try to do is they have their own personal data, they connect an LLM to that data, and then they're able to ask questions. And I think, you know, more recently, they're able to ask more and deeper questions. Deep research has been a big topic. Um, it's a big topic here at GTC, and the more information you can get out of that data and the more quickly you can query that, that's a pretty standard agent that we would see out there.

Now, then they start to get a lot more complex. I think there's some really cool stuff going on in the the telecom space where they're actually using agents to um improve the network. So they can actually predict when there might be network outages coming, and they can make recommendations to the human employees who, you know, maybe you want to make these changes because maybe there's a large show in town like GTC and there's going to be a lot of a lot of traffic on that network, um here are the recommended changes so that you can actually, you know, still provide the the service that you're

Is that kind of thing happening at GTC right now? Are they are they using that kind of stuff?

I wish they would. Um, I think I think the the agents are just starting to be developed. So maybe next year we'll start to actually see that, but it's definitely something um I think we can we're going to see more of. Um, but yeah, so those I think are, you know, some, you know, varying levels of examples, and then there's obviously hundreds more in the healthcare space, you know, helping nurses and doctors become more efficient because we know they need the help um to, you know, everyday people like you and I. I mean, I I use it for almost everything I do, whether it's in my personal life or my work life.

Are there any like agents that that might surprise people? Anything that's like really sort of like interesting that, "Oh, I wouldn't have thought people would be using agents in that area?"

The most interesting one for me on the agent side is really as you started to get into video and other types of data sources like text, I think people have gotten the handle on um, but adding sort of video uh understanding and things coming through it. Actually, one of the really cool use cases that we recently did was uh Jensen got to throw out the first pitch at a baseball game. I saw the clip of that and and and we used a video agent to be able to critique his I remember which I just think is really cool. So you can imagine for, you know, athletes, whether they're amateurs or professionals, being able to use that is really helpful.

Yeah. Yeah, I can see that, like golfers and stuff like that, analyze my swing, tell me what I'm doing wrong, stuff like that. And it's probably too much for the AI on my swing, but you know it still would be helpful to know.

Yeah. Yeah. Well, let's let's talk about the NVIDIA blueprints, cuz when I was at CES, that was a big topic during CES was the NVIDIA blueprints. And to me, that's sort of like the the beginnings of agents, right? But there's sort of this pre-built agent. You could probably explain it better than I do.

So exactly. We tried to name them blueprints so that people would get the idea that these are reference architectures. And what they do is they take the different building blocks that Nvidia is offering to make these agents, and it helps you with a recipe on how to put them together. So it starts out with our NIM microservices, and these are packaged containerized optimized models. So all the leading open models that are out there in the market, we take them, we containerize them, we add a standard API call to them so that people, you know, developers out there can start building, and we we package those up. So now you've got the model, and now you need to connect that model to things, right? A model on its own only does so much.

So like a NIM would be like you've got a a video NIM that this one NIM can produce AI generated video. This NIM could produce text to speech. This NIM could do speech to text.

Exactly. Exactly. So we actually have a hundred NIM that are now available that you can download and use, and they do there are 40 different uh sort of domains or modalities that these different NIM can do. So yeah, so you can take a bunch of them, piece them together, and so when we built those, everyone was like, "This is great. We need the models. We need them to run really well on on, you know, our Nvidia GPUs." But then they said, "Well, but how do we build them into the next thing?" And that's where blueprints came about. So they are the recipe that allows you to take the NIM, start with this recipe, put piece this together, and you'll get to an agent. And then what's even better is you can then customize them. So you can, you know, add your own data sources. You can add your own pieces in. You can combine blueprints together. So maybe you started with a chatbot and you wanted to add a digital human on the front end. We have a blueprint for both of those. You piece them together, and all of a sudden you have a digital avatar who can, you know, talk to you with a full, you know, face and expressions and and language.

Your voice if you wanted to.

Yeah, absolutely. And you can you can start with pieces from Nvidia or you can start with pieces from the ecosystem. We try to make it really easy to piece it all together.

Let's talk. Let's talk a little bit about like security. I know when it comes to AI agents, they can be used for both good and bad. What kind of things can we do to to sort of protect and secure and make sure that people aren't going and using these agents for, you know, bad actor type stuff?

You know what's really cool about AI is where it raises concerns or security challenges, and it can actually also help answer them.

Mhm. So there's actually some really cool um there's a there's a blueprint out there for container security that does a lot of those pieces. Uh, but we also have other applications that can actually track activities and and make alerts based on detecting some anomaly, right? So that tends to be how you recognize that something's doing something it maybe wasn't supposed to. Um, and AIs are just really good at that because they can, you know, consume a lot of information. So I think where, you know, AI potentially can raise some concerns, it's also in some ways the answer to addressing some of those concerns, which I think is really helpful.

Right. So like the the sort of like sci-fi movie scenario of like the AI rising up against us, I mean, is this something that people should be concerned about, or is this something that that you feel there's pretty good guardrails in place for?

I I think I think, you know, the the builders of these apps need to make sure those guardrails are in place. But I think yes, in general, I think that the tools exist, and then it's just about us as the, you know, um application builders being smart about how we deploy them. Um, you know, in general, I think agents are really powerful. They're incredible tools that help humans. Uh, but on their own, they're not they're not about to to sort of go off the rails. It really takes a human to to take it there. So I think again, the tools it is a tool, it is something that can, you know, work alongside you. Um, but I I don't see the rise of the machines quite yet.

So along similar lines, when it comes to like, you know, the I think one of the fears a lot of people have is like, is an agent going to take my job? So, you know, what where what are your thoughts in that sort of realm?

I think yeah, I mean, the best thing I've ever heard is an agent's not going to take your job, but somebody using an agent might. And that's that's always the thing that I think gets framed. No, I think I think exactly that like we we as humans have the capacity to do so much more. There's just often so many hours in the day, right? And so we don't, you know, accomplish, you know, with with agents the work that we were trying to do and then just stop, right? No, it allows us to go and do more and try more. And so I think, you know, for me, it's not about replacing your job. It's about making you more effective at your job and being able to do more and and, you know, be more powerful. I think, you know, healthcare for me is one of those industries that this is, you know, driven out there. There's actually a great partner of ours um in the mental health space, and they're using agents to free up um therapist time from all the administrative work that they have to do. We know their jobs are hard enough as it is. If we can free that time up, it allows them to spend more time with their patients actually deeply understanding and working with them as opposed to worrying about calendars and taking notes and scheduling. So I think those are great examples where it's no, it's something there that's going to help your job. It's not going to take your job.

Yeah. No, I was actually talking, I bumped into somebody on the street while I was walking yesterday, and they had a a very similar concept for their company. They're they work with therapists, and they create chatbots specifically for the therapist so that their customers can go and have the initial conversation with the chatbot, but then the conversations get passed along to the therapist. The therapist can kind of, you know, detect when some bigger issue that they need to address comes up, but it's not replacing the therapist. It's just, "Hey, I can now handle more patients, you know, and give them more of my dedicated time." And I think there's something actually interesting as you start to see chatbots and particularly those with digital avatars is humans will open up in some ways to one of these, you know, chatbots or these avatars in ways that they may not necessarily feel comfortable opening up. So I think it's it's just giving us new avenues to do things that we were already doing anyway. And I think that's pretty cool.

Yeah. Do do you think um this is this is sort of switching gears a little bit, but do you think there's like a compute bottleneck for agents? Do you think we're going to are we going to be able to scale agents and get to agents as fast as people would want to? Cuz there's there was a narrative I I feel like the narrative sort of faded a little bit, but there was a narrative maybe 6 months ago that there's a wall and AI is hitting a wall. Do you think there's a wall? Do you think we're going to hit some compute bottlenecks?

I think, you know, this is this is our life's work at Nvidia is to make sure that we have the compute to, you know, drive the world's agents. And so I think, you know, a lot of the announcements we talked about today, both on the hardware and on the software side, are really focused on making these as efficient as possible. And I think what's really cool is as you're looking at this AI space, when the first comes out, it's usually big. It's boxy. It's maybe not the most efficient. And then over time that efficiency comes in and allows us to do the next big leap. And again, that starts out um, you know, big and and heavy, and again gets more and more efficient over time. So I think um I'm not seeing a bottleneck, but I do know it's an area we need to continue uh to drive because I do think the compute requirements are going to continue to grow.

Cool. Well, so let's let's extrapolate out a little bit. So if you're looking like 5, 10 years down the road, what do you what do you sort of envision with AI agents? Where where do you think this is all headed? And I know that's hard because it's it's sort of exponential technology, and it's really hard for humans to grasp exponentials.

It's incredibly hard. I mean, I can barely keep six months out, you know, and new things keep popping up. I mean, I think certainly the the biggest things that we're going to see are teams of agents, and that's going to be, you know, that we're already seeing it today. It's starting now, and it's going to continue to grow. And so, you know, the way that I see it is, you know, to when we first got agents and we first got, you know, chatbots and things, you could ask a question, and it would respond to a question, and then you could ask it to do a function, and it could now do that function, and then you can ask it to do a bigger job, and you can start to these these agents are just getting more and more sophisticated. So 5 to 10 years out, you know, we may be able to give something, you know, as simple as, you know, "design me my, you know, retirement home," and it will come back with all of the pieces involved in that without a human ever having to do another prompt to follow it up. And I think that that could be pretty cool. So that might even be just two or three years out, but but yes, I certainly see that. I think that's where we're going.

Yeah. So So you mentioned teams of agents, that that's really fascinating to me. Like, do you have any examples of like what sort of agents would team up to work together and what sort of task will accomplish?

Yeah, absolutely. So I I work in marketing, right? And so so a lot of my job is about doing messaging, uh writing blogs, uh getting imagery, building demo videos, things like that. And today, you know, you start with one, you know, uh app that can help you write, and then you go to another app that helps you build images. Then you go to another app that that helps you write the, you know, the the code for the website, and then, you know, right, and they're they're all separate and and each one of them returns to me, and I, you know, move to the next step. I think in the future, what we're going to see is those agents will all be connected through function calls, and we actually uh at GTC announced a blueprint that's going to help us do this. Um, it's called IQ. Um, and so by connecting all these together, again as a human, I'll be able to start put in my my overall request, "I need to build a new website for a new announcement that's coming out," and it will be able to do all those functions together. And I think what's really cool about this is it's going to allow us to design agents that solve specific tasks, but by combining them together in that sort of that composable way, they'll be able to do bigger and bigger job functions.

Very cool. Do do you see like what sort of like bigger world plot problems do you see like AI and AI agents solving for us?

I mean, I think networking is a really interesting one because anything that has that much data that humans can't solve I think is is amazing. You know, digital twins and simulation of those types of environments, whether it's, you know, uh the climate and weather, uh whether it's, you know, the businesses and and things that we run, I think all of these are areas where um agents are just going to add to the ability to work with these. And then I think businesses, of course, it's going to be about, you know, providing those tools so that their employees are just that much more effective, and that's that's not a big world problem, but it's a very common serious world problem.

Everyone everywhere is talking about AI agents right now, but here's the thing: most companies are going about it all wrong. This guide cuts through the hype and shows you what's actually working right now. HubSpot has gathered insights from top industry leaders who are implementing AI agents the right way. You'll discover which agent setups actually deliver ROI and how businesses are automating their marketing, sales, and operations without replacing their teams. Get it right now by scanning the code or clicking the link in the description.

Now let's get back to the show. And one of the things that I think was really fascinating, and I've heard I've I make jokes about how I've I feel like I'm on like the Jensen tour because I've actually seen his last like five keynotes. But one of the things that I've I've seen him talk about that's really fascinating to me is the Earth-2. Yes. Where it's got the they basically can map out the weather patterns and figure out weather events a lot earlier. So I'm really excited to see the sort of overlap of like the Earth-2 concept and agents and solving some of the more like bigger climate type issues as well.

100%. And I mean, I think one of the things when when we introduced Earth-2 is this idea that we we think a lot about these problems. We we have conversations about them, but it's really hard as humans to sort of visualize what, you know, some of the changes that we make in government and things like that are going to do in the next 5 to 10 years. We're very, you know, immediate creatures, right? So that's where we're focused. But I think, you know, imagine being able to go to Earth-2 and through an agent say, "Hey, can you, you know, if we made these changes, what would happen?" Or, you know, if we, you know, took these steps, how could that affect the world, and and what more can we do, uh maybe there are things that we're not even recognizing that um the agent could recommend. Um, because I think that's what's really cool about these agents is they're not humans. They don't think the way we do. And so by giving them all that data and giving them an Earth simulation, they might be able to uncover things that we've never thought of.

Yeah, I think that's amazing. What are some of the things in the AI world, you know, agents or otherwise, that have you personally really excited? Like, what what sort of stuff do you use? What do you play with? Like, what what's your AI stack that you use?

I use as many as I can get my hands on. Um, we have a lot of them in in Nvidia. I mean, one of my favorites and and and I know Jensen talked about it a little bit on stage is Perplexity.

Yeah. Oh, I love Perplexity. I love Perplexity because I think it takes it it changes the dynamic of of humans and how we interact um rather than searching for information and having to spend the thought process on that, it's really about who can ask best questions. And I think that's a really powerful change and and, you know, power dynamic that that perplexity has given to the users. Now, if you can ask great questions, you can find great information. So I love that tool. Um, we have AI agents in our company that help us with everything from, you know, our benefits and and understanding how to make the right decisions um, you know, for each employee, which I think is pretty cool, into, you know, how we do our jobs, whether that's video creation, image creation, um certainly content uh which which you have to write a lot of. Uh, so yeah, so all of those I think are great, and then, you know, I take them into my personal life in terms of organizing and planning.

Yes. Um, I'm very organized at work, and I don't always save a lot of that organization for my personal life, and so I can hand it off to, you know, any source of chatbots. I think uh ChatGPT is excellent for this. Um, it's just pretty cool. Um, and and lately deep research. So I've been using a lot of that and and again, we just introduced a blueprint that's going to allow um deep research on on your own personal things. And so I can imagine that's going to be pretty powerful.

Cool. So for people that want to sort of stay in the loop on AI that are curious, maybe they're worried about it, like what sort of advice would you give them to to sort of stay on

Top of things, my best advice on AI is: use it. It is, it, it sounds really obvious, but I think, um, one, I think it actually alleviates a lot of concerns when you understand how the technology works. By using it, you can see what works today, where the limitations are, and how it kind of functions. And I think it lets people see it as that tool versus, you know, something that might be scary. Um, so that's really the first piece of advice; I think that's really important.

And then I think it's about, you know, trying to identify where many of those, um, you know, concerns might be coming from—whether that's data or security or things like that—and understand how AI can also help with those. And so that tends to be, you know, one of my best pieces of advice.

Cool. Let's talk about, um, was it Llama 3 and Neotron? Yeah. So Llama Nemotron, Llama Neotron. Let's talk about Llama and Nematron.

Absolutely. So Llama Neotron was a really cool announcement. So Nvidia, we work with all of the leading model builders that are out there, including Meta, who released Llama. And what's really cool about Llama is there are—I think it's 85,000 derivatives of this model—right? And so, of course, in video, we've got a lot of really smart people who know how to optimize and make models more efficient. And when reasoning came out, and when Deepseek introduced this reasoning wave into the open source, we sort of said, "Well, how can we bring this and make it really efficient for people who want to deploy this on Nvidia?"

And so, by starting with the Llama model, which is an incredible model, we brought in our expertise to train it so that a model that previously couldn't do reasoning now could actually think through problems. And so we used—we found, you know, the data set to go be able to train the model on how to do this new task—and we trained it, and then we made the data source open, which I think is really cool. So if others want to do their own training, or if they want to train a different model or anything like that, that data source is available. Um, but it's just teaching a model a new skill; is I think a really powerful thing to see because it shows how quickly the space is evolving.

So is there anything actually happening like underneath Llama, or is it just you've got the Llama model, but now there's a new layer on top of it that knows how to think? Is there did any like new training happen to the model? Yes. So yeah, the model was absolutely—so we post-trained the model. So this is what a lot of, a lot of companies out there are doing today is they take a base model and they actually train it with new data. And in this case, sometimes you're training it so it has more information on a particular topic. Um, this is particularly popular when you've got domains and industries that have specific languages that they speak, things like maybe the finance industry. Um, but in this case, we actually were training it on a skill. And that skill is to think through problems.

And I think this is what's really interesting about reasoning is the way a reasoning model works is it starts by thinking through the question that you're asking it. And it breaks down that question into multiple steps and multiple parts. And then it actually goes through and comes up with answers for each of those parts. And then it checks to say, "Okay, now that I've got those answers, does this actually, you know, come back with the right question?" And it continues to do that until it gets to this highly accurate response, right? And so, not only are reasoning models really good for improving accuracy—which we all know that's when an, when an AI model is useful, it's when it's accurate—right—but it also allows us to solve problems we could never solve. Right. So a great example for me on this is: I love puzzles. I love all sorts of puzzles, but particularly Sudoku, because I hear it's good for your brain. Um, Sudoku is a problem that humans can solve, but actually traditional LLMs couldn't. There are actually almost 80 different decisions that go into solving a Sudoku puzzle, and each one of those affects the other decision because, obviously, based on the rules that you have to follow, reasoning models—and Llama Neatron is a great example of this—can solve Sudoku.

Oo, interesting. Where a Llama model on its own could? Wow. And so again, that's the type of thing—it doesn't sound like, you know, solving a Sudoku puzzle is going to change the world—but when you look at a Sudoku puzzle, there's actually a lot of things that go on in the world that are related to that. Yeah, I have a supply chain; I've got to ship, um, you know, items and get them to different stores around the country. I happen to know there's a snowstorm coming in, and I need to make sure that my trucks are taking the most efficient routes. All of those are steps that impact the other decisions that are being made. And so it's actually quite a complex problem that relates in some ways to Sudoku.

Yeah. Well, it's so interesting too, because you take a normal model, and it can't tell you how many Rs are in "strawberry." Give it a thinking model, and it will actually count and then double-check, "Did I do that right?" And then, and then, and then answer and get the right answer. It's so fascinating because it seems like such a simple problem to a human brain. But then that's—I think the thing is we, we look at models and we think of them like we, we, you know, personify them as humans, and they're not, right? And so, yeah, no, reasoning models are really cool for that, and we've seen a lot of those, um, you know, great examples that come out of what these reasoning models can do. And it, and it is—it's just—I think it's really interesting to watch.

Well, it's cool to—Yeah, it's so cool to actually see the thought process because you'll actually see the models think through something and then go, "Wait, that's probably not right. Let me think that through again." And you actually see that text come out of it thinking through. And that, to me, is fascinating.

It's fascinating. And it's a great example of the compute story we were talking about, which is, you know, we do need more compute so it can think. Mhm. The more it thinks, the more compute it requires. So it's this really, um, you know, interesting cycle that we're, we're watching these things go through.

One of my favorite things to do with reasoning models is to ask them to describe things to me like a 5-year-old. Yes. And it will come up with a description. It will check if a 5-year-old would understand that description. It will then make changes. And it's really interesting to see what models think 5-year-olds understand. So some point I'll have to go test it with a real 5-year-old. Yeah. Read this. Does this make sense to you? Does this actually make sense? Exactly. Exactly. Very cool.

Well, along the same lines, let's talk about real quick—Let's talk about hallucinations. Do you see like a path to zero hallucinations? Do we want to get rid of hallucinations completely? Like, what are your thoughts on that?

I think, you know, there's certainly things we can do to reduce hallucinations. And some of those are, you know, as simple as putting a guardrail in place that says, "If you're not 100% confident in the answer, or 99% confident in the answer, don't answer the question." Right? So that's a, that's, that's sort of—it stops the model from doing things that it shouldn't. Um, and Nemo Guard Rails, which is one of the products that Nvidia offers, helps with building those. But to your point about, "Do we want to stop hallucinations altogether?" Of course, we want it to give the right answers, but in being creative, we're asking it to come up with new things. So it's—you have to be able to distinguish between a hallucination and a creative generative response.

Right. Right. So I think that's where the this balance plays. So there are absolutely steps that can be taken, and depending on how targeted and focused you want the model to be, you can put more and more of those sort of guardrails or those policies in place that will keep it from hallucinating.

Right. Yeah. 'Cause I think in a lot of scenarios, hallucinations are a feature, not a bug. Right. If, if you wanted to write a short story for you, you want it to hallucinate that short story for you. Exactly. So I think that's where we have to understand what's the hallucination versus what's the model doing what it's supposed to do, and again, and that's where it becomes—what's the use case? What are you trying to have it do? And really think those through and then find the right tools from Nvidia or others to be able to actually, you know, go and do that.

Yeah. So if somebody wanted to get started with Agentic AI, they want to start playing with agents and testing the waters, what, what do they do? What are our steps?

Well, so from Nvidia, we have something called build.invidia.com. We made the URL super easy. If you're trying to build something, we have a one-stop shop for you. It's got all the models on there so that you can test them out. Whether it's the new Llama Neimatron model with reasoning, you can actually turn reasoning on and off. Um, it also has all the blueprints so you can actually test them and experiment from them. And then from there, there are also steps to go deploy them and test them out and build them yourself. So I think that's a great starting point.

Very cool. And for the more like technical people that maybe are trying to develop something, um, like is there a place they can go play with the APIs? Like, what do we do there?

Build.invidia.com. Same place. Same place. It's, it's literally whether you're a, you know, an enthusiast, whether you're a developer, whether you're actually trying to, you know, build something to put in production, this is the one-stop shop because everything's on there. You can take it as far as you want. You can play around with the UI. You can play around with the APIs. You can actually download and deploy these models on any, you know, Nvidia hardware. Um, it's, it's, it's your one-stop shop for everything.

Yeah, that actually—that brings me to another question. Um, does this stuff work on older Nvidia hardware? If you have a, you know, a 3080 or a 4070, can I use this stuff on those as well?

Absolutely. The only restriction is does it fit within the memory of the GPU, but if it fits, it ships. Um, so yes, this will run on GPUs that are out there in the market today.

Very cool. Awesome. Well, thank you very much. This is—this has been amazing. Fascinating. I love talking AI and nerding out, especially agents. Agents is the hottest topic. So really appreciate you taking the time with me.

Absolutely. I could also talk about the whole day. So yeah, have a good time. Thank you.

One of the things that I've been really excited about that Nvidia is getting ready to release is their project Digits, now renamed the DJX Spark. Well, while I was at Nvidia GTC, I got to have a chat with Bob Petty, one of the guys who's leading up that project, and he gave me a tour of the sort of personal units that you're going to be able to have in your own home or in your business to run your own AI supercomputers.

Hi, this is Bob Petty, vice president and general manager of enterprise platforms at Nvidia. A lot of the infrastructure that people are buying today in the cloud, or through some of our server partners, is Grace Blackwell. It was when it was Hopper Grace Hopper, now Grace Blackwell. Right. Grace being the ARM CPU. Right. Right. And that's kind of the same technology that all the big, big, like OpenAI, those types of companies are using. Exactly. And the beauty of that is the ARM CPU uses up so much less power. So, um, if the majority of the workload was in the GPU and the CPU is kind of a traffic manager, um, you, you'd want, uh, you want to be able to do what you needed to do for as little power as possible. So you can put more GPUs in there. So hence Grace Blackwell, right? Well, you can certainly develop AI on Windows workstations with RTX Pro or GeForce. That's all great. Mhm. Um, but if you need an ARM port of your software or getting familiar, that's, that's what, uh, that's where Spark fits the gap. And this is, this is the same thing as the project Digits that was announced this year. This is project Digits. Yeah, we, we, uh, we finally chose the name Spark. We had a people send in a lot of comments. Um, but it's, uh, as Jensen mentioned in the keynote, you know, what was a box that was like this several years ago—20-core CPU, one petaflop—is now in this little 5x5 by, you know, less than two, 2-inch box. Um, the, it's got the C-to-C memory between the Grace processor and the Blackwell GPU is like 256 gigabytes a second. You won't have that on a traditional workstation because you're going to go over the PCI bus, right? Um, so high, high memory bandwidth between the CPU and the GPU—128 gigs of memory available from us online—but that's really for the enthusiast who, who want the gorgeous bezel. And so what kind of things can I do with this now that I can't do with my 5090 at home?

Good question. So, um, from a frame buffer size, um, there are so many more models that you can run on this. You can do fine-tuning on 70B models. You can, you put two of these together with this cable—This is a Connectex Ethernet—put two of these together, you can run a 400B, 400-billion parameter model. Can't do that on the 5090, right? Uh, you can't do the 70B on a 5090, right? And so the size of the model is very dependent. The other thing with a 5090, your memory bandwidth between your CPU and GPU is throttled by the PCI bus. Okay. Right. And this has, you know, cache-coherent high-speed memory bandwidth. So it really enables you to, to, to test what your code might look like running on one of the data center providers, OEMs out there because same, same, uh, C-to-C memory cache coherency speed. You can do more than just say it works, right? You can, you can get it to the point where you can remove a lot of the bottlenecks, whatever, whether you're doing a vision language model or, you know, multimodal model. That's, that's the biggest benefit. Um, one is getting your code ready for what's the predominant AI infrastructure out there, but the other is testing in a way that simulates how it's going to run, you know, when you run it on a node there. And then the idea is you're, you're not wasting data center time or cloud time just debugging, right? Or eliminating bottlenecks when it, when you're here, you deploy and, and you scale from one GPU to N GPU. So, um, the sole purpose—and not the sole purpose—the main purpose of this was really to, to help, help, uh, spread the Grace Blackwell ecosystem, right? They're going as fast as we can make them, but they're not necessarily accessible to enthusiasts, right? Who want to use FP4 features of Blackwell, uh, which you can do on 5090, but want to use it with some of the more popular models that have higher, higher parameter sizes. So, Easy Box, the 1 TBTE version of storage—128 gigs—is $29.99. The 4 TBTE version is $39.99. Um, you can reserve on nvidia.com. Our initial go-to market partners are Dell, HP, ASUS. They've got their own branded, branded boxes without the gold foil. Uh, and then we'll expand that.

Yeah, I remember Jensen said something to the effect of, like, "Imagine it's a cloud computer just sitting on your desk." It's not going to the cloud. You know, you don't have to worry about internet connections, anything like that. It's just a cloud computer sitting on your desk that can do all of the inference right there on the bigger models. It's an AI supercomputer on your desk. And there's a bigger one that we'll walk to in a second. Cool. But the other thing about this is we're not suggesting that everybody just replace their existing, you know, laptop or workstation. You got a GeForce laptop or RTX Pro laptop. You, you plug this guy into it. So you might do all everything you need to on a 5090, run your games and everything. You're doing AI development or want to write AI inferencing that helps you on your 5090. Yeah. Plug that into it. It's a—and that's why Denton sold the, you know, the MacBook and the Spark, right? Because it's, it's really meant to be both a, um, a plugin to a Uber assist, right? You know, maybe less, less, uh, capable machines, whether they have a GPU in them or not. And it's not using all the processing on your computer. So you can be running AI models on this, Cyberpunk on your computer. Exactly. Yeah. Exactly. So, um, and you know, just form factor-wise, cost-wise, we think it would be easy for people to, to do it as an add-on, but certainly there's, you know, there's a great GPU in here. We've, we're running games on this. I, you know, I wouldn't want it as my GeForce laptop, but I'd probably want to connect it to my GeForce if I was, uh, you know, doing, uh, AI, AI tuning for game development or things like that. Right. Right. So that's that's DJX Spark. Um, if we just walk this way, we've got—this is the RTX Pro, um, line. Uh, our 6000 line is the one that's, you know, somewhat akin to the 5090 on the GeForce side. The reason we have this Pro line, um, you know, the, the manufacturing of it is a very precise bomb. It's not built by many different AICs. Um, there are, uh, computing benefits on here that we perceive the gaming community doesn't need. So some of the, some of the high-end compute, uh, performance here is going to be much better. The AI inferencing, uh, um, performance is about the same. Mhm. Uh, big difference is frame buffer. Right. And doesn't this one have like 96 gigs of video? 90? Yeah, 96 gigs of RAM. And, um, we, we've got the, uh, if you want to get full power, we've got the 600-watt version. We've got a 300-watt version, uh, that most desktops can take today. And then you've get the same technology in the server version, uh, that would go in a rack here. So, um, we used to call these, uh, in the past we've had A40s or L40s based on Ada Lovelace. The Ampere Lovelace. We're going to call that B40 based on Blackwell, but, uh, kind of aligned around, uh, RTX Pro Workstation Max Q. Max Q is that optimal power point and, uh, and the server edition. So again, same infrastructure, you can code and develop and then deploy on your RTX server in a rack in the data center meant to, um, save time.

Heading down this way again, we're doing a lot to generate. So, um, you'll see the brand here, but these are, these are our DJX stations, um, initial, initial partners, um, uh, Dell, HP, H, Asus, Lambda, and Super Micro. Uh, we show these two here because these are their boxes. This is what it will look like. Um, if you get look at here, this is a GB300 board. So a much more powerful, um, Grace processor and then an extremely powerful B300. B300 is the same GPU in the latest DGX. So, uh, 784 GB of memory again. So the benefit is you're doing ARM development. Um, a lot of memory and, and very, very high-speeded memory bandwidth like, like Spark. So you're, you're not just seeing that code works, you're seeing how well it works before you chew up time on your data center rack. Now does this need a, like, a separate CPU, like an Intel, AMD kind of thing, or it's, uh, the on both the Spark, um, and, uh, and the station, we're providing the Grace CPU. Okay. So Grace CPU and the Blackwell GPU there. Um, uh, graphics out. We, we don't put a big graphics card in here. So this one has a 4000. That's a small form factor. The reason being, and we want this to plug into a standard 15-amp wall outlet, right? Might need to be a dedicated one, uh, 'cause you know, it's at six, 15 be 1650 watts, I guess. Um, and we're going to come pretty close to that, uh, which is why the manufacturers are liquid cooling this. Um, and, uh, they started doing that in the gaming side. So some of the Alienware chassis you've seen to liquid cool, right? So they're well adept at that. Um, they will liquid cool this so we can stay, you know, thermally be good, um, and not need to take up, uh, a lot more power with a lot of fans. So, um, again, if you're, if you're an enthusiast just getting started, uh, even an enterprise where you're, you know, you know the size of your models and what you want to get done, uh, easily connect that if your full-time job is like prepping AI, developing AI to deploy to the data center, right? You have your choice of logging it to the data center, maybe you get a virtual workstation delivered back to you and you do your job or putting this at your desk. Okay. And this—It literally is—it's like one of those B300 nodes, right? It is a server node in a desktop, um, with graphics out right there. From a data privacy standpoint, from an IP protection standpoint, I'm not sending anything anywhere, right? It's right there. Uh, and that's just not—that's more important than just privacy and IP protection. It's just the time and the cost of transport of data, right? You, you might run a, you know, a 600-billion parameter model, um, but the data that you're running it on—whether let's say it's Cosmos, the VLM, all the videos that you're going to be processing—and the amount of data you'd want to just sit here versus upload all that data or even if your own dedicated data center and run it there and then download that. So you got ingress and egress cost of data transport all happening right here. Um, that guy will be available in the summer timeframe. Okay. Um, reservable today. This guy will be late summer. Um, we have a Founders Edition for that because it's, it's cool, and enthusiasts are going to want something on the desk, right? Uh, this one is all only available for the OEMs. Way for us to scale our enterprise businesses through the OEMs. And so, um, if you were to go to the Dell booth today or the HP booth or ASUS, you'll see their versions of these, which here you'll see their, their versions of the Spark with the Dell blue and the Nvidia green LED and the HP blue. Um, and that's, that's the way to expand the ecosystem for Grace ARM development, Grace Blackwell development, expand the access to technology that normally is only available if you've got the capital expense to put one of these racks in and now it's at your desktop.

Very cool. Yeah. Amazing. Well, thanks, Bob. This has been really, uh, informative, and I appreciate it.

Thank you. No problem. Thank you.

Once again, thank you so much to Amanda and Bob for having those conversations with me. Before I wrap this one up, I want to ask for a favor. The Next Wave podcast was actually nominated for a Webby Award in the business category, and we need your help to try to win that award. It's voted on by listeners. So if you enjoy these podcast episodes, we'll make sure there's a link in the description where you can go and let your voice be heard, and you can vote for the Next Wave podcast to win a Webby Award. Thank you once again for tuning in. We really appreciate you. Hopefully, we'll see you in the next episode.