📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Vik Malyala, Supermicro | AMD Advancing AI 2026

SiliconANGLE theCUBE23:14

Transcription

Welcome back. We're the cube's live coverage here in San Francisco, California. AMD's advancing AI event. All the leaders are here, industry participants. It's a free event. So, a lot of practitioners and technologists here checking out the rack scale systems, checking out all the new gear and networking, servers, everything's here and power the next generation.

I'm John Furry, host of the cube with Dave Volante, my co-host Nick Mielis here. He's the chief business officer of Supermicro. Cube alumni. Recently, I interviewed, we interviewed, long time no see, >> Supermicro. Dave just interviewed him. Um, you guys had a big storage summit with the cube. >> Great to see you. >> Thank you for having me. As always, it's a pleasure to talk to you. >> It's lots happened since Supercomputing when we were riffing on the neo cloud growth, AI infrastructure. A lot of content we put, publish out together. But the big thing is just the massive buildout, the demand curve Dave was mentioning before we came on camera, and just the nature of these AI factories. They're kind of taking shape as a new architecture. It's not the rack and stack days anymore, although there's still a lot of that going on, but it's a whole different computing paradigm. Explain the current situation. Where are we today? November. We just had, a lot's happened since November.

>> I'm telling you, um, whatever we think is the pace at which the industry is moving, three months later, it, it is proving us wrong, right? Um, what I have seen is that, let's say last time in November when we were talking about, there was hardly much discussion about agentic AI. Was it, it was all about like, you know, RAG models and people are trying to develop these different applications. But then in a snap, it all changed. Now we are talking about uh, heavy use of agentic AI in uh, across the board. And we have a strong demand that is coming up, not just on a GPU, but also on the CPU because of that. Um, as far as the GPUs are concerned, now uh, enterprises are starting to pick up. It's not just all about uh, the training clusters, but also the application stack that is being used on them, which is basically making the enterprises adopt that. Um, so what I see is that uh, it's a very, very broad scale adoption of AI infrastructure, which is uh, propelling this massive, massive uh, demand.

>> And you guys have such a great history. I always kind of flex the Supermicro success over the many generations. You know, supply chain and right now, that is the number one thing people care about. Prices go up in memory, but people want solutions now, right? >> Talk about the impact of what that's done and also how's that's changed, how you guys think about building these large AI factories and these systems.

>> It's a complex equation. I mean, no, there's no ifs and buts about it. And uh, if you take a look at the like of uh, likes of AMD and Nvidia, for example, uh, for the rack scale solutions, they are bringing the predictability in the supply as well as pricing by them working directly with the memory manufacturers on the supply side, at least. And in in most of the cases, some level of pricing also. Uh, but as far as the standard uh, HGX platforms and everything else is concerned, uh, this is going to be a very tricky market for quite some time to come. And luckily for Supermicro, we've been working with pretty much every one of the industry leaders, whether it is memory or whether it's flash. Think of like Micron, Samsung, Hynix, Solid, and every one of them, uh, for the longest time. And as we build these platforms and solutions, we give them enough visibility into, you know, where the systems are, knowing who the customers that we are supporting and what problem we are trying to solve, which sometimes helps actually in getting, you know, an allocation at the right time. But, uh, you know, no matter how much planning that we have, we know for a fact, um, the supply is a lot more constrained than ever before. And uh, we continue to work with them to see how best we can uh, address. And most importantly, we set expectations to customers that it's no longer possible to have a system ship it in like, you know, a week or 10 days. It takes a longer time. And by taking the orders in, by having a proper forecast, it's also helping us to, you know, plan better.

This event is like, I won't say coming out party, if I want to say that word, but we've been following AMD. They've been in the x86, we've seen that growth. But I kind of been following, we can see them hiding the ball a little bit. They had FPGA GPUs coming. What is the new AMD like? You talk about the relationship you have with AMD, and how the, I won't say new and improved AMD, just the, now the AI high systems side of AMD starting to show, starting to see the results. >> What's it, explain what it means to the market and what people should know about it.

You know, I, I'll give like a parallel story here. Like, you know, back in 2015, 2016, when I was working with, uh, Dr. Lisa Su on the first generation of the EPYC platform. What ended up happening was the, uh, the hardware was fantastic. I mean, you have like many cores and the PCIe lanes and everything, but the software and other ecosystem wasn't quite ready. So while it's adopted in a very specific, uh, marketplace, whether it is HPC, whether it is, uh, hyperscalers, the general adoption took like, you know, an extra two cycles, you know, after Milan, then, you know, the next step, like Genoa, Bergamo, and all these things. So what I have seen at every generation, more people started adopting, the ecosystem started expanding, and to the point where they actually have the market leadership in x86, right? The same, the same situation I see happening with respect to Instinct also. Uh, initial units are like, hardware was fantastic, but the ecosystem is the one that needs to be developing. I was told that some 12,000 people signed up for this event, which is something that no one would have imagined two years ago. If you ask, in terms of the software developers, what I am seeing is that as people start figuring out a way to use the GPUs, then the adoption of the GPUs beyond the training clusters is going to happen. And we have seen a good demand build up, especially around the MI 350 and 355. And the demand is kind of extending into the Helios. There's a lot of people asking about it and people are trying to understand. Mind you, it's not just about the technology. People need to figure out a way to host them in their data center. It's not going to be, it's not going to be easy to just run, you know, pull a Helios.

>> It's like the GPU is the initiation. Wow, I love this value. Like, wait, I need more compute. So, so the discovery on the ecosystem becomes kind of its own progression.

>> It is. It is.

>> So, you guys hinted that you're going to, that your, your margins are up. Your backlog is, I think $60 billion was the number you, you published. Is it, is it as simple as just demand is so far outstripping supply?

>> I mean, we have done something right, right? So ultimately, uh, we need to take care of our customers' demands. And, um, as customers become successful, the business starts to grow. And that's precisely what uh, happened in this case. We have a growing uh, set of customers and wider adoption of the platforms. Think of like an enterprise, think of like banking and financial sector, as well as the traditional compute and uh, GPU environments. These are the ones that are actually creating the demand for us. And, uh, as far as the margins are concerned, it's a mix of various things, right? Uh, you're talking about data center building blocks that we are bringing, uh, to expedite the deployment phase of these GPU clusters. As part, part of it, service organization coming up to make sure that the things are going to be up and running and be able to, uh, stay, um, up and running for, you know, uh, at a, at a high percentage of time. And, uh, the relationship that we have with ISVs for the storage and what we are bringing in terms of the storage solutions to customers. So a combination of all these things is the one that is driving, uh, both the adoption as well as the impact on the margin. And margin is something that's going to fluctuate, especially because of, uh, how the whole, um, supply chain, uh, dynamics are working today. But we are certainly emphasizing on, uh, value and how we can actually help customers to bring some predictability into their equation.

Yeah, the Supermicro storage summit. I was there. It was, it was a great event. Some really good content. Um, you were talking before about how things change so fast. We actually earlier today, John and I came up with a list. It's like, oh yeah, RAG-based chatbots are, you know, going to add huge value. You got token cost. Who cares about token cost? And frontier models are going to dominate everything. You know, these other small models, you know, don't mean anything. Um, how are you and, and, and AMD engineering for these rapid changes and seeing around corners and how do you accommodate, uh, this change?

So none of us have a crystal ball, but what we can do is to bring the configurability into the platform design and be able to size it differently. So one of the things that we have, uh, done is, um, if you were to take a look at, uh, the GPUs itself, right? You have the MI 350 and MI 355. That is the HGX platforms predominantly being used, especially with AMD. And then we also have the, uh, inferencing point of view, which supports different accelerators. We have several accelerators that are actually based on AMD as a compute platform, and they have these accelerators as a PCIe. And AMD themselves are coming up with MI 350P, which is also a PCIe-based accelerator. Why I'm saying it is that for the guys who want, like, you know, think of like the Rolls-Royce, I mean, you have these, uh, platforms at the rack scale, and it's fantastic, right? But for the ones who are trying to use for different applications, they may not need that, uh, in order to have the right ROI for what they are trying to do. They can go with partially populated systems. Like, for example, I can have a system, uh, that supports up to eight of these PCIe GPUs, and people can actually start with two or four, or, you know, scale up to eight depending on what they are trying to do. And if it goes beyond eight and if they are trying to scale up, then we have the rack scale solution, and multiple racks, we can connect and create a cluster and whatnot. So there's no rocket science per se, but what we are trying to do is to give the option to the customers on what might work for them. And this is, uh, one of the things. And the second part of what we are seeing is, you mentioned you started with the RAG models and whatnot. That one is like, you know, single loop, right? I mean, you, you ask a question, the model is going to run it, spit out the answers, and you're good to go, whether it's right or wrong, and how much of it is hallucination and all depends on the data it's trained on, and, you know, how the models are developed and whatnot. They're getting a lot more accurate. But this agent is a completely different spin, right? I mean, you're talking about, it needs to think, it needs to reason, and it needs to go into multiple loops. It needs to memorize things, and, you know, and it spawns off like so many more, uh, sub-agents who are doing all their work. And what is happening with this is, uh, there is a significant demand for the compute, not the, not necessarily the GPUs. The GPU's point of view, they're doing a fantastic job. Think of like your, uh, what you mentioned about the RAG models, and when it gets to that point, it has matrix multiplications, it can do very quickly, it can spit out the answers, that's fantastic. But on the compute side of it, how do you make sure that the right type of queries are going to the GPU, and the rest of it is handling on this? In the RAG model, all it is doing is just sequencing, managing the pipeline, orchestrating done. But with respect to agentic AI, now we're talking about the workload that's happening on the CPU, which is impacting 50 to 90%, depending on, you know, where you read and what you read about in terms of the latencies. So if the impact is that much, that will be handled by the CPU, then it's a criminal waste not to have the best computation.

>> And the cost and the cost financially, the over and the technical costs are massive because.

>> You got to feed the GPUs, you got to do routing, so there's a whole another orchestration layer here.

>> But finish that thought. So it's, it's criminal to not have the best CPU architecture to orchestrate all of that, right?

>> Right, right. And that's where I think, you know, now you have the CPU, you have the GPU, you're connecting, bringing AMD into the equation, you know, having the Polaris cards right now, and then, you know, the Pensando part of the equation. So what, what we are trying to do is, um, you know, depending on what customers are trying to do, we can size it right. So have a PCIe-based accelerators or the standard, you know, HGX type of platforms, how much of pre-compute clusters that you need, and how much that you need it for agentic AI, what is the kind of bandwidth that is needed between these, uh, compute as well as the GPU platforms, how do you pair it with the appropriate storage so that, you know, when things go and go past the size of, uh, uh, available memory and storage available within the system, how it's going to expand into the storage for the KV cache allocation. Right? So the entire thought process of Supermicro as a building block is absolutely making sense right now in this kind of a scenario, because as things get more complex, ultimately what matters is how are we bringing the right value to customers and sizing it right and making it run most optimal way.

>> Talk about the AMD relationship when it comes to co-designing and engineering. A lot's going on to these rack scale systems. They're essentially me super servers.

>> As one. And then you have distributed computing also, other areas. They're going to have footprint. Got the nice cards that are going to actually target workloads. They all have to talk to each other. How do you guys partner now in the AI era and, and where is it the same and where is it different?

So take an example of, uh, the PCIe-based cards, right? Uh, it started with like, you know, 75W, 150W, 225W, 300W, now 600W. So AMD gives us a visibility, hey, my GPUs in future is going to be air-cooled, but I need a platform that supports 600W. Or I'm going to have 600W, but I would like to bring the maximum density. In which case, I will say like, well, if that's the case, let's see if that design can accommodate liquid cooling. So I can actually put those many cards. And the other part of it is, let's say, uh, AMD silicon for the GPUs is based on PCIe Gen 5 or Gen 6. Based on that, I can say, well, I can actually put in the Venice platform with the PCIe Gen 6. So you have much higher throughput that is going to be supported in the platform. So what we typically look at in this scenario is, uh, what are the different technologies coming from them? What is the time frame? And when I bring a platform, how am I going to make it run in the most optimal manner for that? Does it require liquid cooling? Does it need to be air cooling? What is the form factor? What could be the power delivery? And we look at where that product could be adopted. Like I mentioned, when people start adopting Helios, let's say deployment starting whenever it is, we will be ready with AMD to have our compute platforms to go along with it. And the public information is that AMD Helios is going to be running with the ORW form factor with the power delivery using the DC bus bar. The entire server product portfolio, we have transformed. When I say not entire, but most of them, like a Hyper, the Cloud DC, the Twin platform, the Flex Twin, all of them, we have started designing with the power delivery using the DC bus bar. Why? Because as they go into the data center, they don't need to figure out like, okay, now I need to have standard power supplies, I need to have a 63 amps or 100, uh, amp power drops and the PDUs. Helios doesn't need it. DC bus bar. So I'll just do the same DC bus bar on the compute also. That way, you don't have to go and rediscover. This is how we are looking at what would be needed for the customers in a given time frame and what kind of a platform that I need to develop, or platforms I need to develop in order for us to get that rolling quickly.

So where do you see Helios fitting that rack scale system? What's, where do you see the demand for that? What type of workloads is it going to support? What's the fundamental customer value prop?

>> So at this point, people, if they're doing the primary inferencing workloads, HGX platforms are doing a phenomenal job.

>> But when it comes to frontier models, that's where I see the immediate demand for, uh, Helios, because the models are getting trickier. That I say trickier is because it's no longer just about how many trillions of parameters, it's all about like, you know, what are the optimizations that they are bringing into that. And when it comes to inferencing, I mean, I think it was like a year ago or whenever it is, when we talk about the DeepSeek and everyone like, oh my god, the world is coming to an end. And after that, everyone has multiplied their numbers by like three or four times.

>> It's happening again now. Models of innovation. They're talking about the point is all of it is going to drive the democratization. So whenever people figure a way to run things more efficiently, it's not a negative thing. It basically makes it easy for others to engineer it. It's called engineering buying opportunity.

>> Well, no. This is where this is where I like the compute direction because when that gets smaller, faster, cheaper. That sounds familiar. Moore's Law. I mean, so you're starting to see a kind of a dynamic where the engineering is going to enable things to run cost-effectively. That should help the enterprise for sure.

>> For sure.

>> The neo clouds want to have versatility. They want to have great price performance. Get their margins because the margins they want to have higher margins.

>> And, and, and, and one thing that is happening, both this is not specific to, you know, AMD, it's even Nvidia and whatnot, is that, uh, every one of them have announced, when I say all the technology partners have announced the roadmap. Like this is the power delivery going to be, this is how the system going to be, this is how much cooling I would need, this, this is how heavy this system is going to be. So this way, the data centers are being either retrofitted or built. Otherwise, what happens is the traditional data centers, they have no way of taking this 240, 50 kW per rack, 1000 lb, and what, right?

>> There you go. And you need to change the elevators. And you need to, oh, by the way, um, we have seen cases even with the, forget about Helios or anything, even with the previous generation, the current generation platforms, we are shipping. Some data centers have retrofitted, everything was fine. We, we did an audit and everything, it was fine. When you start rolling in the racks, the tiles creak.

>> Yeah, because heavy. It's these are heavy. So that's the reason, um, it's not just about us working with AMD in developing the platforms. We also need to work with the data centers to ensure that they can actually handle this.

>> Otherwise, I'm going to have a very heavy paperweight.

>> Data center is the, the data center is the computer. It's.

>> Like data center is the computer.

>> It's like the sale, that's like the salesperson from the Xerox salesperson back in the day. Where are you going to put it?

>> The business is good. Obviously, love this. And, and the bubble conversation kind of dies down because you look at the backlog on all the top hyperscalers, neo clouds, they're not even recognizing revenue. They have backlog.

>> So there's backlog. So this bubble discussion, I mean.

>> I don't think it dies down. I think supply chain bubble keeps pricing okay. I see that. That's a supply issue. But the demand going forward, this.

>> I mean, like with any market in any type, there are going to be some fails for sure, but I don't see it as a bubble at all because I see this pent-up demand is so strong. I mean, that's part of the reason announced, right?

>> You don't think we're in a bubble?

>> Not a pop bubble. Not a top.

>> I think we're in a bubble. It just hasn't popped. That's it.

>> I don't think it'll, it'll settle. It won't bubble. It's expanding. The market's going like this, isn't it? That's.

>> If the application stack doesn't, Dave.

>> Yeah.

>> It's all good until, but, but the idea is, give me your perspective on that. Ultimately, all these tools and everything need to make our life more comfortable and better, right? Hardware alone is not going to do that. Firmware stack is not going to do that. It's ultimately the user experience and what value it's bringing into the equation. I mean, just with RAG models, we have seen so much of improvement in how we are able to handle. I mean, nowadays, you know, coding is great.

>> Exactly, right? And now.

>> Take it for granted. You're right.

>> Right. And then going a step further with agentic AI now. It's a completely different experience. And look at the different verticals it can make an impact. We hardly started seeing these things in any of these verticals. So my take is that there's going to be a whole bunch of software developers that are going to scale up and do different type of work, not just like writing bunch of lines of code, but actually architecting solutions that people can take advantage of. And that's what is going to make these, you know, hardware used. I think the metric, I think you're right. I think you're right, Vic, because one of the metrics that's coming out of all of our conversations over the past year here on the cube is whoever can bring the value to the user, and the value to them is agency, empower me to do new things. That's like the web when it came out. That's like all these new environments. Like, well, it's obvious, old ways not as good as the new way, but it's natural language. It's not GUI.

>> It's a whole another expectation experience that's driving the stack to be rethought. If you can't provide value that's easy to use and simple to execute, you're out.

>> Agree. So, take an example of enterprise customers. If a million tokens cost like 40 cents versus a million tokens costs $40, the CFO is not going to sign the check if it is going to cost like $40 to $80 for a million tokens because he's going to close the shop. Even if they're able to run the business very efficiently, the outcome is not going to be as great. Versus if they're able to bring the number of tokens, you know, the, the cost of tokens down to a reasonable level, then it will be adopted and the business is profitable. So businesses need sense at some point. I mean.

>> Yeah, I mean, it will be. But the idea again here is tokens, number one cost coming down, things running more efficiently. And more importantly, when people are using the models, get smarter. So they are not going to, you know, let's say boil the ocean to do simple things, but they are going to do only the very specific, vertical-specific application and tune to that, which is going to run more efficiently. So we have a long road ahead of us in terms of what all the things that can be done with it. And.

>> Yeah, I think so. The bubble's going to keep getting bigger and bigger and bigger and bigger and bigger. That's called TAM, Dave. It's called TAM. That's called market, market territory. Vic, great to have you on as usual. Great conversation. Uh, we'll see you again soon. Probably Supercomputing or.

>> Open Compute, one of these shows. Love what you guys. Just keep making the faster machines, systems for us. We need more tokens, cheaper, faster, make it.

>> Driving the cost down, Vic. That's your job.

>> Efficiencies.

>> Tech is deflationary. You're do, you've done a great job there. Keep it up.

>> And get as much memory and give it to the cube. So we have a side business of up-reselling memory because a lot of demand.

>> According to you, the bubble, it needs memory, isn't it?

>> No, no, bubble's a good thing. It's just when it pops, that's the bad thing. Chief Officer, Supermicro. Again, another example of the, the suppliers rethinking how they build the systems that are powering AI and advancing AI is a team sport up and down the stack. The ecosystem is a huge part of it. We want to see him go faster. We need more product. Doing our part here in the cube with the ecosystem. I'm John F with Dave Volante. Thanks for watching.