📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

NVIDIA'S HUGE AI Announcements Will Change Everything (Here's Why)

Ticker Symbol: YOU24:54

Transcription

I'm excited to share this exclusive interview with the investing community. Most people think of Nvidia as a hardware company that builds chips to train massive AI models, but you're about to get an inside look at a very different side of the story. I'm joined by Joe Delair, product lead of AI infrastructure at NVIDIA. Joe spent the last four years deploying the hardware and software behind some of the most powerful AI models on the planet, and he shared a few surprising insights about where AI is headed next.

But that's just one of the many technologies that I'll be covering live at GTC in a few weeks. GTC is Nvidia's massive AI conference, showcasing the biggest breakthroughs in robotics and self-driving cars, AI agents and the chips that power them, and a whole lot more. And anyone who signs up for a free online session at GTC with my link can win an Nvidia RTX 5090 graphics card. Just attend any session, take a screenshot as proof, and send it to me after the conference using the links below. GTC should be on every investor's radar, and so should Nvidia's ecosystem for AI inference. Your time is valuable, so let's get right into it.

I'm so happy to be here with you. Thanks for taking the time, by the way. Jensen talked about a lot of awesome things at the keynote and one of the things that he talked about in detail is that Nvidia actually co-designed six different chips for the Vera Rubin generation. That's a lot to go through. So I'd love to go through all of it with you starting from the GPU itself and working all the way up to the rack scale system level if that's okay. So let's let's just start with uh Reuben itself. What's the difference between Blackwell and Reuben?

Oh, so there's several different things about Reuben that are are different than Blackwell. So we have the six chips that you talk about, uh, all of the code design together. So what we did is we looked at the data center requirements and we worked our way backwards and said, what do we need in all these six different chips to make sure that we get the best performance, the best energy efficiency, the lowest cost. >> Uh, so that's what the fundamental thing about Reuben is this extreme code design. All these chips manufactured together, designed together, uh working in concert for the best performance. >> And when you say you looked at the data center requirements, are those being driven by AI models today or what like what's driving those?

>> Absolutely. Models are definitely the thing that are driving this compute demand and models in particular mixture of experts where they're generating uh many many tokens factors more tokens because of the reasoning that they do. Uh also the model sizes are growing as well. So they're getting more intelligence from model size from reasoning. Uh so that is just generating a tremendous amount of compute demand >> and Reuben is designed to address that. >> Got it. So talk to me about the difference between Blackwell and Ruben the GPU specifically in terms of power and performance.

So in terms of power and performance uh for inference workloads we've see uh up to 10x better performance on Reuben versus Blackwell. >> Wow. 10x performance per watt. So that means that so at a given lat fixed latency you can see with those paro charts that we've shown in Jensen's keynotes >> at a particular latency a very you know high latency that's very uh good for users of the model. So yeah the 10x performance is across the rack scale >> is it at the rack scale >> rack scale architecture. So here we have the black wall ultra generation compute tray and uh I can show you what what we have here in terms of the components and their breakdown.

So we have uh two super chips. >> Two super chips. Okay. >> Super chips have uh two blackwell ultra GPUs um and then one uh gray CPU on on one super chip and then there's two of them together. So four GPUs, two uh CPUs. Uh and then we also have connect x8 super nicks that are also part of this super chip and that's going to be important distinction when we talk about ver rubin later and how those have been moved. Um but yeah, you can see that this is a hybrid cooled. >> Okay, >> so these are cold plates doing the liquid cooling on the super chips and all their components. And then on the bottom half of of the tray or the front half I should say, this is air cooled. So these are all what I'm actually looking at is the tops of all fans, right? >> Eight fans here. >> Got it. >> So eight fans and then we have a a blue field DPU uh that is part of this tray as well. That is for the north south traffic uh connecting to storage getting the data in to the the the compute rack so that it feeds the the GPU. >> Got it. So the yeah, the DPU brings data in and out and then all the proc all the magic happens in the super chips themselves. Got it. So there's two kinds of network traffic. North south is inside the same rack. East west is connecting multiple racks. Is that how we should think about it?

>> That's that's a proper way to think of it. Yes. Yes. >> Um I thought Nvidia was just a GPU designer, but Grace is a CPU, right? So what does the CPU do? >> So the CPU handles a lot of the management. So for example, like when you're doing uh you're trying to use inference and you want your your model to make uh some code for you >> and you want it to maybe it makes a little application a Python application, it needs to run that where a CPU can actually run that application. The GPU wouldn't run an application that's generated by by a model. >> Okay. >> Um but is also doing u other types of things like database analytics and those types of functions that are more CPU friendly. uh it's able to accelerate those types of effort. >> Oh, so really the whole idea is kind of like you let the GPUs do what they're the best at. Then you have the CPU to do things obviously that CPUs are much better at GPUs at so that you can sort of spread out the work over the right chip for the job. Right. >> That's correct. >> You also mentioned something called a DPU. Can you walk us through what a DPU?

>> So DPU blue field DPU data processing unit that's going to handle some of the north south traffic. >> North south traffic. Yep. And uh when you're connected to storage, that's on a different rack. Uh there's going to be compression, encryption. Um that's all going to be managed by the DPU that we have in Bluefield, Bluefield 3. >> And the goal for that is just to make sure the CPU and the GPU aren't doing those things. >> That's correct. Offloading offloading all those functions from the CPU and the GPU, accelerating those functions and hardware um so that you get the fastest data access to feed the GPUs. >> That makes a lot of sense. Okay. So those are three of the six chips so far, right? The CPU, the GPU, and the DPU. >> And the and the Connect X. >> Yeah. Talk to me a little more about that.

>> Connect X8. This is your east west connectivity. So this is your supernick for connecting east west. It also has uh inline encryption, those types of functions for the east west traffic that's going to be connecting between rack to rack of GPU racks. >> Got it. So we have the GPU, the CPU, the DPU, and the connect deck on this board. >> That's correct. >> Where are the other two chips? >> So the MVLink switch is the the other chip. Um and there there's two here on this switch tray. Uh this is MVLink 5 or the fifth generation of MVLink. And these are these are communicating to the NVLink network at 1,800 gigabytes per second. >> 1,800 G >> 1.8 terabytes per second. So, uh, very high speed. Um, and that's really going to be the the central nervous system of a Blackwell GB300 MBL72. >> Got it. So, so these are two completely different trays, right? So, this is the compute tray. That's where the magic happens in terms of crunching the numbers. And then this is the switch tray, which I think you mentioned earlier is all about just connecting all the GPUs together.

>> So, it connects all the GPUs together. Uh, there's several of these trays within a rack. Yeah. >> Uh all the GPUs are 72 GPUs. They have >> 72 GPU >> in a rack and it's all to all connectivity. So every GPU has to be able to talk to every other GPU at full bandwidth. And that's what the switches uh achieve. So 1.8 terabytes per second. Any GPU talk to any other GPU. Is that why it's called a compute fabric? Like when I think when I draw a network diagram of that's okay, got it. So >> So yeah, they call it compute fabric. Not just because it's connecting all the GPUs to each other. There's also some compute functions in our MVL link switch chips. So we call that uh all reduce or collective operations where in training when certain operations need to be shared uh across the network instead of sending it to all the GPUs it will do some of those operations within the switch. >> Oh wa. Okay. So the switch isn't just connecting things. It's actually also doing some >> some computation as well. That's awesome. Okay. So, I think we've covered five of the chips now, right? Is that correct?

>> That's correct. >> Where What's the sixth chip? >> Six is the uh Spectrum X. Uh >> what's Can we try to take a look at those racks? >> Yeah, let's go take a look. There's 10 trays up top. Those are the compute trays. Nine networking trays, nine NVLink switch trays, I should say. And their job is to connect all the GPUs in the 10 above and the eight below compute trays together. Right. >> So what's up there then? >> So that that is the top of rack uh 1 GB switch for telemetry. >> That's telemetry. >> That's just system management managing functions. It's low speed Ethernet. It's just a uh it's a just a management system for the rack itself. It doesn't it's not processing the compute data for AI. It's managing if a GPU goes down. It's like help me understand what telemetry means and what that >> telemetry means like I'm just looking at the the functions of the rack itself. I'm looking at uh its uptime. I'm looking at >> health and status. I guess >> health and status checking. Yes. Diagnostics would also >> and you mentioned that there's another kind of rack that would sit next to this. So yeah, you will have your your group of compute racks, GB300 compute racks and then you would also have racks dedicated to Spectrum X east west network switches. Um we don't have that here but uh that's how the the function it would be like a we call it a pod. You have uh maybe eight GB300 racks and then you'll have a a few uh switch racks with Spectrum X. >> Yeah. >> Yeah. So that's a great overview of the Blackwell system. Right now I want to understand how things changed from Blackwell to Reuben. >> Okay. >> Can we go over there? >> Let's go look at the at the trace.

>> So this is uh looking at the components up here on the wall. >> We talked about in the compute tray the Bluefield DPU Bluefield 4. Yeah. >> So that there you can see it uh on the wall that that that board is part of the module system that slides in and out of the compute tray for serviceability. >> Uh and then likewise the connect X9 is there in the middle. Uh and there's two connect X9s that are on that board uh for a total of eight in every compute tray. So every GPU is fed 1.6 terabs per second for the connect X9's. Uh and then we have the the SpectrumX photonix co-ackage optics. This is really really cool. >> Yeah. What is that? >> So instead of having uh SFP pluggable uh modules for for the optics, they're actually built onto the chip itself. >> Co-aged with it. So this has a a huge gain in energy efficiency uh reliability uh and its factors more in terms of of those two factors. So before we would have fiber optic transceivers, >> the fiber optic optical transceivers. >> Yeah. So the fiber optic cables on either end and those transceivers have lasers in them. >> That's correct. >> That need power, right? Like and that's what you're getting rid of >> and we're putting packaging on the with the chip. >> What does that actually mean in terms of like performance or power gains? So in terms of performance, the performance would be the same. >> Yeah. >> But it's going to be the uh the power reduction and the uh reliability improvement because uh those pluggable lasers can be very, you know, sometimes very unreliable. They have to be swapped out very frequently. But if it's co-acked here uh on on the chip, the reliability goes up like uh I think 10x better reliability. >> Wow. So that's a huge difference. And where in the rack does that live? So that would be in its own switch tray uh or a switch server and that's a separate rack. >> That's the side. >> That's the separate rack that's separate from the the MBL72. So that's the east west traffic switch rack. >> Awesome. >> So Quantum X uh there's also uh for Infiniband which is an alternative to Ethernet. There's also a co-acked optics for Quantum Infiniband as well. >> So those two chips are equivalent. One is for Spectrum X Ethernet, one is for quantum infiniband. >> That's correct. >> And then you also have a Spectrum X Ethernet photonix switch. >> So that is the uh the co-ackage optics chip is in there in the Ethernet photonic switch. So that's the photonix part is the co-ackage optics. >> Got it. But these these go in the sidecar. >> These go into uh switch racks. >> Yeah. Got it. >> As well as that one, right? If you're doing quantum infiniband as your east west traffic protocol then you would use the infiniban as a side rack. >> So these are sorry these are equivalents one for infiniband one for Ethernet. Right. >> Correct. >> Got it. >> Yeah that's right.

>> So what we kind of just talked about is what I would say is the current state-of-the-art for data centers. Right. Blackwell Ultra is the one that's sort of the best in data centers right now. And then Jensen announced Vera Rubin. The six chips we just talked about. We talked about the Blackwell versions. This is a substantially different compute tray than the one we just saw. Can you walk us through all the differences? >> Oh, yeah. There's plenty. So, uh, what we've done is overall it's a modular design. >> Okay. >> So, that means that there there's bays here and these can just slide out and slide in and just lock and latch. So, there's not a bunch of wires and cabling to do all the connectivity between all the components that are on the tray. Also the hosing as well that's been streamlined. >> Yeah. >> So, there's a manifold in the in the middle um and it manages a lot of the uh distribution of liquid. So overall on the GB300 there was 43 hoses. There was a bay of fans here cuz it was hybrid cooled. >> Uh the the bottom half of GB300 was was fan cooled. This we've eliminated that. Uh and because we're 100% liquid cooled now. So eight fans goes to zero fans. zero hoses >> and then there's a bunch of cables that have been removed as well. Uh so it's cable free. >> So this I I'm trying to even piece together what I'm looking at. So these would be where the two super chips were. >> So these are the super chips. They slide in and out. They latch in. Uh so you have the two Reubins uh one Vera on them. So, uh, one other important point is because it's modular now and we have all these bays that slide in and out and it's all connectivity with connectors instead of cabling, putting this together and doing assembly on it is like 20 times faster. >> Sure. >> So, something that would take 2 hours to assemble the GB300 rack, now you can do in 5 minutes >> on this particular rack. >> And that's and that's just assembly, right? Like if I have a maintenance issue and I need >> it's also for maintenance, right? the uh the amount of speed that you can do serviceability increases that many fold as well. >> No, it makes a ton of sense, right? If I don't have all these wires and hoses and I can just snap things out, fix fix whatever the issue is, snap it back in. >> And it's modular like so we'll talk about some of the other pieces down here. So, two super chips, Reuben Vera. Uh, we also have the CX9s, Connect X9, the next generation of that supernick are over on these in boards in modules. So, before they were connected to the bottom of the super chip uh on GB300, but now they're their own module and cards slide in and out. So, you can service different components now separately. >> Yeah. And then Blue Field 4, the new generation of the DPU is also a module here that slides in and out. >> Got it. >> So this is not just about performance, it's also about more uptime, right? So that's another multiplier on the overall output of an AI factory is how much up time you >> we call that goodput. Like you you want the >> the the amount of time that you're actually producing tokens. You want to maximize that. >> Yeah, >> that makes sense. So, okay, this is the equivalent compute tray and then there's also an equivalent switch tray, right? >> That's correct. >> And this looks a lot more streamlined, too. So, walk me through the changes here. >> So, in terms of the changes here, uh, you know, we have the the switches at the top, 100% liquid cooled. Uh, there's four switch chips. This is Envy Link 6, >> sixth generation Envy Link, twice the speed of what we had in the black level. >> Twice the speed. >> So, now it's 3.6 terabytes per second. Uh and that's just going to help us with our that performance I talked about 10x performance per watt or per megawatt per gigawatt whatever value you want. That's the increase in NB link speed is part of that contributes to that along with some other GPU features that we can talk about as well. >> And are there so is it the same number of total GPUs in a black wall rack versus a Reuben rack? >> It is. So it's NBL72. The 72 signifies the GPU count. So GB300, MVL72, uh, and now we have Ver Rubin, MVL72, same GPU count. Um, and it also makes it so it's very compatible for our customers to to move from one to the other. Um, and that's part of the goal of having the same GPU count, same kind of MGX rack architecture. >> Um, so that this just makes it easier for our customers. The ecosystem is, you know, been working with these racks for two generations now. Now we have a third generation. They're just going to be able to work very fast and deploy uh at a very high rate with our end customer. >> No, it makes total sense. Okay. Can we go look at a Vera Rubin rack now?

>> So this is the Vera Rubin. Uh, this is the Ver Ruben MBL72 rack. You can see that there's you know it's very similar in in form and and look to the GB300. The the most uh the biggest difference is on the compute trays. You'll see there's no vents. So, there was vents on the GV300 because the bottom half of the compute tray still had fans. >> Okay. Yeah. >> You know, we got rid of those fans. It's all 100% liquid cooled on the compute trays. Now, that's why you see in the face plate, you don't see those vents anymore. >> Got it. >> Uh but overall, still, you know, still the nine uh switch trays, still 10 compute trays on top and the eight on the bottom. Same kind of design. Still telemetry on top. >> Still top of rack telemetry with the one gig switch on top. >> Now here's the big question. Right from Blackwell to Reuben at the rack level. Talk to me about the performance gains at the rack level. >> Performance gain at rack level is the 10x 10x >> the 10x more uh tokens per second per megawatt or per watt. Um but and that's going to be a rack level kind of uh performance metric and that's with a mixture of expert model something like Kimmy K2 thinking uh which is a very large large model over a trillion parameters uh and that is going to fit and be uh optimized in a single rack uh with you know thanks to NV link switch the experts in a mixture of expert model are distributed across the 72 GPUs and uh that can uh factors more performance in tokens per second. So here we have the Kyber wrap. So this would be for the Reuben Ultra generation. Subsequent to Reuben, which is a 2026 product, in 2027, we'll have uh Reuben Ultra. >> Okay. >> So that's going to be a different rack architecture than we've had for the previous three generations. Uh we're putting much more compute. >> Yeah, I'm noticing a lot more trays in this one. >> So we have 18 compute trays in each of these canisters. So there's four canisters up to 72 GPUs in each of the canisters. So you would have 288 >> 28. So moving from 144 to 288 or is that 72? >> 72 to 28. >> Okay. So it's a 4x increase in GPU. >> So each of these canisters, the four I talked about is equivalent to the whole rack over here. >> So there's four racks worth of GPU NVL72s worth of of comput in here. >> So very uh high compute density. Yeah. Um and that's why the architecture is different. It's a blade type of architecture rather than a tray architecture. Uh so we have 18 uh compute blades uh in each of the canisters. >> Sorry, these are all compute. Then >> this is all compute on the front. >> Yeah. >> On the back is where the switch blades are for the for the NVLink connectivity. >> Got it. And so that's what is the performance leap that you guys are expecting from Reuben to Reuben Ultra in the Kyber Rack. >> So we haven't given any of the performance yet on Reuben Ultra, but it's going to be uh factors more performance as as usual between our generations. >> Yeah. >> Just because you're going to have in performance increases at the chip level, at the super chip level, at the rack level, and you're going to have four times as >> it's that extreme code design all again, right? Extreme code design all the chips being designed for for greater performance working in concert uh being designed from scratch together. >> Are we expecting extreme code design of all six chips for every generation from now on? We should expect to see six new chips. >> So for every generation uh there's going to be a new generation of GPU for for every year. >> Uh now whether all six are going to be co-designed every year that's that's uh probably not going to be the case. But you're going to see at the for the flagship start of each generation like Reuben, six new chips, >> uh some other new chips that go with Reuben Ultra, but not the entirety of all six. >> For example, we might see the Vera CPU, but the Reuben Ultra GPU. >> Exactly. Got it. >> Exactly. >> Got it. >> Yeah. >> I'm super excited for this. I can't wait to see what this looks like. When When can we expect to learn a little more about this? Is this something that we'll learn about this year, next year? So yeah, it will be something that Jensen talks about, you know, in the in the coming year. U I don't have a specific date, but yeah, >> I'm super excited for it, man. >> What are you looking forward to the most? Like what excites you the most as you see like this rapid evolution year over year and generation over generation. >> So the uh the amount of in innovation at the with the extreme code design, that's what's most impressive, >> right? So there's only so much and Jensen talked about this that you can do moving from one GPU generation to a next process technology can only improve so much. Um, you know, it's not factors more improvement in in the number of transistors that you can go from one generation to the next. So for example uh between Vera Reuben and Blackwell it's about 70% more transistors. >> Yeah. in terms of all the different chips that we we co-design, but we're getting the 10x more performance per watt. >> So, if you were just Morris law, it would only be a 70% jump, not a,000% jump from Yeah. Not a 10x. >> So, this kind of all these different chips being designed together, working together to maximize that performance. That's the most amazing thing about this generation and the future generations. >> That's really exciting. Thanks so much for your time. A huge thank you to Joe Delair for breaking down Nvidia's Blackwell ecosystem, giving us an inside look at Reubin and explaining how it will all make AI models faster, smarter, and more efficient. Not just language models, but everything from image and video models to medicine, robotics, and so much more. And if you want to really understand the science behind this stock, join me at NVIDIA GTC. You can register for free with my links below and jump into as many online sessions as you like. I'll announce the winner of that RTX5090 giveaway a few days after the conference, so make sure to enter. Another huge thank you to Nvidia for sponsoring my travel and my media access to cover GTC Live and to you for supporting the channel. Thanks for watching and until next time, this is TickerolU. My name is Alex reminding you that the best investment you can make is in you.