Transcription
Nvidia just committed $4 billion to two photonix manufacturers that most people have never heard of. And when I read the press release, I obviously went straight to Jensen's exact words because I really wanted to understand what he was actually saying about this, what he actually thought about this.
Now, he called it building the next generation of gigawatt scale AI factories. Gigawatt scale AI factories built on light. Now, if you know what that phrase means, you already know why this is one of the most significant infrastructure bets Nvidia has ever made. And if you do not know what it means yet, I mean, well, that is exactly what this video is about.
I went into this video thinking the story was going to be about chips. I mean, that is where every headline points. Transistors are shrinking more slowly. The 3nmter wall is real. And I did cover the whole dinard scaling breakdown in my Nvidia Blackwell video which is linked down in my description if you haven't seen it already. It's a really interesting I mean I think so I made the video but it's very interesting and really worth considering watching.
But the deeper I got into the engineering literature the more I kept hitting a different word and it wasn't chip it wasn't transistor it was interconnect. The connection between chips kept coming up as the actual binding constraint in the next generation of AI clusters. And when I finally understood why, it reframed the entire story.
So here's what is happening inside a modern AI training cluster. First off, you have tens of thousands of GPUs working together. Each one is doing an extraordinary amount of computation on its own. But training a frontier model means those GPUs are in constant conversation. I mean passing gradient updates, synchronizing parameters, moving activations across the network billions of times per second.
Now engineers call this collective communication. It is the invisible architecture underneath every large language model you have ever used, which is kind of wild to think about, but it brings up a question. What carries all those signals between the GPUs? Copper traces, electrical signals moving through the metal, and copper at the speed and density that the next generation AI clusters require runs into physics that it just can't frankly negotiate with.
Let's talk about the first problem. The first problem is resistance. Every electrical signal moving through copper converts some of its energy into heat. And I mean, at a small scale, it's manageable. At the scale of a 100,000 GPU training clusters that heat becomes the dominant engineering variable. The thing that every hardware decision downstream has to account for.
Now the second problem is signal loss. You see so electrical signals lose fidelity over distance and at high frequencies. So engineers face a real trade-off faster or farther, but not both without adding regeneration hardware that brings its own cost and its own heat.
Now the third problem, the one that really stopped me when I was reading these numbers is energy per bit. Moving a single bit of data electrically has an energy cost. Small but real. When you are moving pabytes of data per second inside a training cluster, the energy cost of the wires alone starts to show up frankly meaningfully on the operating budget.
The analogy I kept on coming back to was this. Moving electrons through copper is kind of like pushing water through a straw. You can widen it, you can push harder, but friction is a physical property of the medium itself. And no engineering decision removes friction from a physical medium.
Now, photons, what this video is all about, they work differently. Particles of light carry no electrical charge. They generate no resistive heat as they travel. They also do not degrade over distance the way electrical signals do. And through a technique that's called wavelength division multiplexing, you can send dozens of completely independent data streams through a single optical channel all at once. Each one riding a different color of light with no interference between them. I mean that is why a single strand of fiber optic cable can carry the internet traffic of an entire continent.
That physics has existed for decades in longhaul telecommunications. But what is happening right now is that engineers are building that same physics down to the scale of a chip down to the inside of a server rack. I mean frankly down to the connection between two GPUs sitting centimeters apart.
So before we go any further here is exactly where we are. First of all, the chips are doing their job. The wires, copper wires, they're the constraint and replacing those wires with light is a product road map with real money behind that. Keep that in mind because it changes everything that we are about to cover next.
A conventional GPU processes data electrically. Transistors switch on and off billions of times per second. And the data it moves through copper interconnects between compute cores, between memory, between GPUs in the cluster. And basically, every hop over copper costs latency. Every hop generates heat. And the more you scale the cluster, the more both of those costs stack against each other.
Now an optical processing unit engineers are starting to call these OPUs restructures this chain at a fundamental level. Instead of encoding information as electrical voltage, data is encoded as light, its intensity, its phase or its wavelength. So basically, instead of copper traces, signals travel through silicon waveguides, which are microscopic channels etched into a photonic chip that guides light the way a fiber optic cable does, but at scales measured in micrometers.
Now, here is the part of the MIT research that really made me stop and had to reread the paragraph. The core mathematical operation in virtually every neural network, the operation running inside every transformer, every diffusion model, every mean every large language model you have ever used is called the multiply accumulate operation. And it's basically matrix multiplication. Billions of these per forward pass. The intelligence of a model is I mean at the hardware level an enormous amount of this one operation repeating.
Photonic hardware can perform matrix multiplications in the optical domain. basically light passing through a series of tunable components and it's able to do this I mean at the speed of light without ever converting to an electrical signal. Now the energy reduction that MIT's research group published is over 90% for certain operations compared to electronic equivalents.
Now I really want to stay on that for a second here. Let's take it up a beat where we were really down in the weeds. I want to stay on that number for a second because I think it is easy to just kind of let it wash over you. We spend so much time in this community, I mean in Tiffen Techch talking about model architecture improvements that deliver 3% 5% I mean maybe 10% efficiency gains anyways and those matter. I mean people build careers on them but 90% is a different category on its own. That is a different physical substrate doing the same work entirely.
I do also though want to be careful with those because those figures come from lab conditions and specific workloads. Real world numbers will land lower of course as these systems mature but even at 50% or 60% consider what that means at a moment where an AI energy consumption is becoming a grid level policy concern as we just heard recently anyways for national governments the economics of inference change in a way that compounds every deployment.
Okay back to where the market is moving though now let's talk a little bit about light matter it's one of my favorite companies I'm sure you've heard me speak about it a hot. Now, light matter was actually founded out of MIT and it has built a photonic inference accelerator which is called passage and it replaces copper interconnects inside a server with optical ones. They have raised over $400 million and move from prototype to customer deployments. I mean they are from a university lab bench to production infrastructure and it happened in under 7 years which is pretty incredible.
There are some other labs as well that are working on co-ackaged optics. Another one is AR labs. it is working on co-ackaged optics. So basically integrating optical input and output directly into the chip package itself. So the data leaving a GPU travels optically from the moment it exits the die. Nvidia, Intel, I mean and several other major hyperscalers are investors.
And then recently with Nvidia, it was literally as I was finishing research on this this topic actually when they announced something that really just kind of puts a bow or a punctuation mark if you will around it all with Jensen committing 4 billion across two photonic companies, 2 billion into Lumenum and 2 billion into Coherent, two optics in laser component manufacturers and Nvidia just made them two of the largest infrastructure bets in its portfolio.
Now, Jensen's exact words were, "Together with Lumenum, Nvidia is advancing the world's most sophisticated silicon photonics to build the next generation of gigawatt scale AI factories." Now, what really makes this investment so credible to me and what I think gets missed most in the coverage of this all is that Nvidia has actually been cautious about photonics in specific parts of its own architecture. Jensen previously noted that sticking with copper on their rack scale GB200 systems actually shaved down 20 kowatts off of a 120 kW system. So this is a company that weighed photonics carefully against copper and chose copper where copper made sense. The $4 billion commitment tells you though where and when copper stops making sense. And markets read it the same way. Lmentum closed up to nearly 12% on the day it was announced. say with coherent at 15%.
All right, let's take another breather here. So that is a science and that is the market signal landing on top of it in the same week. It's kind of huge. It's it's wild. All right, but let me tell you why I think the combination of those two things matters more than either one of them alone because this is really where the real scale of what is happening becomes so clear.
We tend to think about technology improvements as incremental. A chip gets faster 20%. Bandwidth then in turn doubles. Important. I mean it's expected. We hear about it. We move on. Photonix though is a phase change. And the reason is that the problems it solves are compounding constraints that get worse together with every hardware generation. Here's exactly what I mean though.
Every generation of AI models has been larger than the last. The compute required to train and run them scales with parameter count, data volume, and cluster size. Under electrical architectures, so does the heat and energy cost. And these are problems that multiply across every dimension simultaneously. With pressure builds with every generation, there's a point on that curve where you simply cannot build the next model. The physics of moving electrical signals at that density and speed will not cooperate with the ambition regardless of how good the software is or how intelligent the architecture team really is.
Photonix though it moves that point. It removes the ceiling that copper imposes on how large a cluster can be and how efficiently it can operate. And that matters tremendously for where AI capability goes from here. Now, the map from the AI systems of say 2025 to the AI systems of 2028 may have less to do with the transformer architecture improvements and more to do with whether photonic interconnects reach production scale in time to support the training that runs the next generation of models and what it will require.
This is a hardware story, but it's also a capability story, an energy story, I guess, if you will. Really a national infrastructure story. Both Momentum and Coherent are now committing to expand US manufacturing capacity as part of their Nvidia agreements, which really tells you this is being treated as a strategic infrastructure play, not just a supply chain decision. And the fact that it touches all of those dimensions at once is exactly why the largest compute companies on the planet are funding it and taking it seriously. and also too why it deserves more attention than it is currently getting.
Nvidia just invested 4 billion into this. So let me put that into plain terms if you will. The models you will use in 2028, the ones that will feel qualitatively different from what exists today may owe a significant part of their capability to the fact that the wires finally got out of the way. That is what this transition means at scale.
Now, let me get specific about what this means for you. Every major hardware transition in computing history has created a new class of engineers who understood the new substrate before most people even knew it existed. And those engineers shaped what came next. For example, when computing went multi-core, the people who already understood the concurrent programming and memory models became disproportionately valuable. I mean when GPUs became AI infrastructure, it was once again the people who understood CUDA and kernel optimization who wrote the playbook. Everyone else eventually had to learn. When cloud replaced on premise infrastructure, it was the distributed systems thinkers who helped entire industries make the transition.
Photonix is the next substrate shift and really the window where understanding it early is a genuine differentiator that is open right now. But I want to be clear about something before I give you specific advice here. I'm saying this because I think this is one of the most interesting skill gaps to close right now. And I want to be very honest about what closing it actually can look like because well, first of all, it is more accessible than maybe it sounds.
The role that becomes increasingly valuable over the next few years is what I would call the AI systems architect. Someone who holds the full stack hardware and software together. Now don't worry, you don't need to be an expert per se in all of this, but someone who understands how light moves through silicon waveguides, how that affects the memory bandwidth and latency at the cluster level and how those physical properties cascade into how you structure a distributed training job or an inference serving system. That is the difference between being carried along by a transition and understanding it well enough to shape what comes next.
All right, as we wrap up this video, here is a thought I keep coming back to after spending the last few weeks really deep in this research topic. Now, every time computing has approached a physical limit, the story gets written as a dead end. Each one really described as the thing that finally stopped progress. No matter what story you're talking about, we are in such a monumental, such a pivotal moment right now. The engineers working on photonic computing are doing it because light is genuinely better physics for what AI infrastructure needs to become. No resistive heat, no signal degradation, no charge. The fundamental speed limit of the universe, which sounds crazy to say, carrying dozens of data streams simultaneously on a single physics path. These are structural advantages, the kind that compound across every generation of hardware that follows.
The infrastructure industry is really starting to act like it. I mean MIT research is in production. Startups are deploying and recently Nvidia put $4 billion behind it. This is what I genuinely find interesting and so exciting about this moment. We did not hit a wall. We found a better medium to think in.
So here's what I want to hear from you in the comments. As photonic interconnects change the data movement architecture of AI clusters, do you think it forces a fundamental rethink of how we write distributed training systems and inference pipelines? or does the software abstraction layer hold and most engineers never really need to care about it? I mean clearly I have a strong view on this. I think we need to care about it. But drop your comments down below. I always read every response that is left on these videos and a lot of the videos I structure come from what you suggest for me to do. I mean photonix was brought to my attention by some comments in my last video. So thank you for that. So let me know what other topics you're interested in. I love these deep dive tech videos that lets me do a ton of research and then get to share with you my findings and hopefully help you stay ahead in tech and better understand what is happening. All right, I'll see you in the next video.