Transcription
Over the past few years, we witnessed an incredible AI revolution, which has been driven by AI chips. In fact, the demand for computing power has never been higher. Meanwhile, the scaling of classical computer chips has slowed. So, what's next?
While graphene chips, probabilistic computers, and quantum computers are still in the making, light-based computers have already arrived. In this episode, I will break down a new light-based computer chip that is on its way to data centers right now, and I can't be more excited about this. Let me shed some light on it.
Photonic computers have been in the making for decades. It all started 60 years ago with the development of optical fiber for communication, and over time, we got excellent at sending information with light. Now, if it works so well, why not use light for computing? In fact, researchers have long been working on building light-based computers.
By now, you've likely heard the idea that light-based computers are faster than digital computers because light travels way faster than electrons. Well, it's true and not true at the same time. Let's take any conventional chip, like an NVIDIA GPU, for example. During computation, an electron travels through a copper wire, which acts as a conductor. This is how it always works.
In fact, the problem here is not the speed of the electron but the medium itself—the wire. Light travels at 300,000 km/second, while in this case, we are talking about mm/second. Again, here it's not a problem because the wire is a conductor, so it's full of electrons. We can reach speeds way faster than mm/second.
Now, you see, we can't simply say that photons are faster than electrons; it's way more complicated than this. In reality, the real reason why digital computers are slower than light-based computers is that in digital computers, we need to switch from zero to one and from one to zero. This switching requires us to charge and discharge a capacitor, and this takes time. This is where the real slowdown is coming from.
I explained this concept in much more detail in my previous episode on reversible computing—a great episode. Make sure to subscribe to the channel right now and watch it right after this video.
So, by now, we understand that the real slowdown comes from this switching, from charging and discharging a capacitor, which is slow. That's where the light-based chips save the day because nothing like this is happening in the photonic world. In photonics, we compute data without stopping it. Basically, we are computing as the data is flying by, and this computation on the fly happens in the range of femtoseconds, which is one quadrillionth of a second. So, it's very fast.
The main feature of light is not light itself, but the main feature of light is that you can realize an analog computer. This is the difference. It's not so much the light part when it comes to the math; it's more the analog nature of light that you can natively exploit. That's also why we call it native computing.
The main advantage here is that you can carry out complicated mathematical functions without digitalization, and that's very interesting. In fact, if we want to perform a simple summation on a digital chip to add up two numbers, we need roughly 200 transistors—those tiny devices that all digital computer chips are built off.
Then, when we want to do a square root of this number, we need another 7,000 transistors. And when we want to do a Fourier transform on this, we need roughly 1 million transistors. So, you see, the more complex the function you want to implement on a digital chip, the more devices, the more transistors, and the more chip area it will take.
What's so interesting is that when we want to implement a Fourier transform with light, we can do it on a single optical device. So, you get much higher computational density. You might be wondering how this is even possible.
You know, people who are wearing glasses—if you are wearing glasses, you are wearing a Fourier transformer on your nose every day. It performs this function using no energy at all. Once you understand this, you can use the same principle to implement such complex operations on a light-based chip using special photonic elements.
Just think about it: we can replace 1 million devices with just one optical device, and it's passive. This means light just passes through it, allowing you to do complex math without spending any energy at all. The same applies for multiplication operations, where on a digital chip we need roughly 1,500 transistors, but on a photonic chip, we can do it with just one device.
So, we get much higher computational density. That's the reason why the interest in light-based chips is growing at light speed. In practice, it took many decades from the time this concept of computing with light emerged until we figured out how to actually use it for computing purposes.
One of the main challenges is that light is really hard to control. It tends to spread out and scatter, and it has taken the industry a long time. But Q.ANT has finally built a fully functional commercial light-based computer. Their new computer chip is called NPU (Native Processing Unit), and it's powered by light rather than electricity.
We are already shipping the first servers to high-performance computer centers. We've decided that the first processor generations are coming on the standard interface of the CMOS world, mainly PCI Express. What we actually deliver to the customer are fully equipped servers that are compatible with x86 structures.
So, in the end, you get a server module. You plug in the Ethernet cable, you plug in the power plug, and the system operates. What's even more interesting is that their breakthrough technology relies on a special material they are using, called lithium niobate.
Essentially, they deposit a thin layer of lithium niobate on top of silicon dioxide, which sits on top of silicon. This particular material is Q.ANT's proprietary technology, which is fundamental for the success of their computer chip in several ways.
First of all, it's the only material that allowed them to build all the required optical components in the chip in one material. This is fundamental for avoiding losses—losses of light because losses of light result in a drop in accuracy in computations. So, we want to avoid it at any cost.
What are the fundamental features of lithium niobate? Well, the first thing is that the modulators—whenever you want to interact with the light, we can realize modulators that can operate in the gigahertz regime, so very fast. We can realize these modulators so that no light is lost in the modulators.
The last thing is the switching. In the end, at the technological granularity level, what you're doing is changing the refractive index of the material. This can be done only using a voltage. I know this sounds super technical, but it's elementary because when you only need to change a voltage, there is no electricity on the photonic part of your processor.
This means there is no heat, no heat dissipation, leading to a very clean signal. We already talked about how clean signals are fundamental to reach, for instance, an 8-bit precision. So, this is why lithium niobate is not just another material; it's basically the fundamental source of success for building a photonic analog computer.
In fact, the Q.ANT chip is the first photonic chip that is able to achieve the accuracy of 8-bit precision. Now, to be honest, what struck me the most about Q.ANT is that they have their own fab. They are manufacturing their own chips, and basically, they own the entire pipeline—from design to technology.
Then, they manufacture the wafers, dice them, package them, and write software stacks for them. That's a lot of work. This is a very untypical situation for a startup, especially owning manufacturing because this is very asset-heavy.
The question is how this startup is managing it all, and the most important: why do they need this fab? Light chips, by physical definition, are pretty large. You can't realize a photonic circuit with a 50-nanometer width because then the light would not be guided.
In that sense, we have access to a CMOS foundry—an old CMOS foundry from the 90s—and we repurposed it with strategic investments in a few tools to serve for the production of our own photonic chips. So, in that sense, yes, it's not cheap, but in comparison to what you need to invest in the CMOS world, it's easier.
This is a big advantage for the future as well because think about it: there are a lot of outdated CMOS foundries in the world that could be repurposed to build high-performing chips for the next generation of AI supercomputers, but using mature technology from the 90s. I mean, this on its own is a production paradigm shift.
This is indeed a paradigm shift—an interesting example of turning, so to say, obstacles into opportunities. Seeing all the investments governments are making into photonic technology and into photonic fabs, and also seeing all tech giants like NVIDIA, TSMC, and AMD going all in, these fabs might have a bright future.
Let me know your thoughts in the comments. Now, before we discuss how this new light-based chip works, what it's capable of, and how it compares to state-of-the-art GPUs, have you ever wondered how much of your personal private data is floating around online? Your name, address, even information about your family members—unfortunately, it all gets out there thanks to the data brokers that spread this information online.
This exposes you to risks of data breaches and, of course, personal security. You've probably heard about cases where databases containing information about millions of users are sold online, and sadly, this is happening more and more frequently. That's where Incogni, the sponsor of today's episode, comes in.
Incogni helps you remove your personal data from databases used by data brokers. I used Incogni to remove my personal data from those databases, and it's surprisingly easy. You create an account, give them permission to act on your behalf, and they send data protection law-compliant requests to these companies, forcing them to remove your information from their databases.
You can even track the progress of these removals day by day on your dashboard. As someone who values privacy, I highly recommend that you try out Incogni and put an end to annoying spam emails and calls. Use my code INTECH at the link below to get 60% off an annual plan. Thank you, Incogni, for sponsoring this episode.
Back to the bright new world. Now it's time to discuss applications and how it compares to state-of-the-art GPUs. Honestly, I spent quite some time looking into specs, trying to make an apple-to-apple comparison, but it's really challenging. One thing is clear: this analog photonic approach offers way better efficiency.
Just think about it: there are no wires, so no resistance, no heat generation. These chips require much less power to operate, especially at high frequencies. Here, we are talking roughly about 30 times better efficiency compared to conventional digital chips.
Now, with respect to scalability, I think the fundamental question is: can we compete with a GPU cluster? Because this is, in the end, our direct competitor in the present data center. To give you a bit of an outline for the future, in two years from now, we're going to have native processing units—processors coming on a PCI Express card that have the same performance as a graphics card in two years for AI-relevant functions.
On the same side, we anticipate that these systems will have a 30x smaller power consumption than a graphics card in the future. Now, what does this mean? If you look at a server today, you can bring eight graphics cards into one server rack, and then you're at the edge of what's reasonable in terms of power consumption.
We can bring many more cards into the same space, increasing the computational density in the server. Since we still have energy budget left, we can bring many more servers into a server rack, increasing density. This is the forecast of today, and I might be wrong, and it could be even better tomorrow.
But the forecast says that if we equip one of those servers and plug in the same electricity as they do today, we can exceed the computational density in the server rack by a factor of 10. What's even more interesting, according to Q.ANT, their chip is built for both inference and training of AI models.
This is very interesting. Typically, we distinguish between two different kinds of workloads: a more simple inference when we have a pre-trained model, and we apply new inputs to it, asking it to recognize an object, an image, for example, to recognize a fox. On the hardware level, this typically reflects into performing many multiply-accumulate operations in parallel.
We've decided that we are fully concentrating on AI inference and AI training. The layout of our chips is always such that a chip can serve both purposes. We can run AI inferences, which is fundamentally similar to vector-matrix multiplication.
For AI training, we are basically going a different route because we can, in contrast to how training is established using a CMOS equivalent GPU architecture, but the chip layout is always the same. This is very interesting because, in order to do training, we need to constantly update the model weights. We need to constantly adjust it to improve its ability to make better, more accurate predictions.
To do this in photonics might be really challenging. As we discussed in the beginning of the video, in photonics, this concept of capacitance or storing intermediate results does not exist. Nothing like in the von Neumann architecture, where we have local memory—in photonics, no storage is available.
In fact, it works entirely differently. The longer we can make the light propagate through the chip without stopping it, the more we can benefit from the properties of light. Let's say you want to train a neural network. First, you encode your weight into the phase of light, and as the light propagates through the chip, you modify it along the way, one by one.
At the output, you get the final value, and only then do you convert it back to digital and save it to memory. For that, the Q.ANT chip features a small electronic part on top of the photonic engine. Keep in mind that there is this special fundamental property of light: it can carry a wide range of frequencies within the electromagnetic spectrum.
To put it simply, we can encode many inputs, many data at once, at different colors of light, and compute it all in parallel. This is very attractive when we are dealing with large data sets, like in the case of AI applications.
We know photonic computing is new. We understood that this technology can really turn the AI world upside down. But on the same side, to be allowed to enter the ecosystem, you need to be compatible with the electronic interfaces. If we had our own proprietary interface, there would only be a minor change that would be adopted into this ecosystem.
The second thing that was also very clear from the very moment is that the programmers, the coders of the world, should not have to change their source code in order to experience the features of our technology, at least in the first instance. This is why we have a whole architecture that we call LENA (Light Empowered Native Arithmetic).
This includes the photonic world and the electronic world, which in the end is the processor that comes on a PCI interface. But at the same time, we also develop the drivers, the compilers, and the interpreters that are then seamlessly adoptable by the libraries used by all the programmers out there—from TensorFlow to PyTorch, from Keras to ONGs, you name it.
In that respect, it's the easiest way to step into this ecosystem because the customer doesn't have to change anything. To be honest, I'm really grateful to my channel for this opportunity to talk to the most visionary people out there. This is a very interesting chip and a very interesting startup with a big vision.
Of course, there is still a lot of work to be done, but to me, it seems like we are closer than ever to the light era in computing. Let me know your thoughts in the comments, and I would love it if you could share this video on social media with your friends and colleagues. I would really appreciate it.
Still, I felt this video wouldn't have been complete without mentioning another huge transition happening in the industry right now: using light for interconnect. Here, we are talking about interconnecting parts of the chip—chiplets—with photonics, as well as moving data between the racks in the data center.
At a large scale, as we've just discussed, light has a much higher bandwidth, or if you want, a much higher capacity because here we can access frequencies in the terahertz range, and that's a lot. Of course, this attracts a lot of interest from tech giants like TSMC, NVIDIA, Intel, and AMD.
Recently, NVIDIA and TSMC announced a collaboration in this space. They have together developed a silicon photonic-based chip prototype. Interestingly, TSMC is making this project, this innovation, a top priority among all their other R&D projects, and they call it COUPE, which stands for Compact Universal Photonic Engine.
This new technology will allow TSMC to integrate optical components closer to the processor core and combine multiple electrical chips with the photonic engine and fiber optic connections into a single package. With these, they will come to more compact and more efficient designs.
You know, modern data centers are quite complex. Generally speaking, there are two main parts to it: computing clusters and networking clusters. When we train a large neural network, we need to distribute this workload across the data center.
If we try to fit as much as possible into a single cluster, when we are talking about one of the latest GPT models, which is roughly two trillion parameters, this is really a challenge. The thing is, it simply won't fit on a single cluster. It means we will have to distribute it across many of them, and here the efficiency will come down to the wiring—wiring between clusters, the networking.
On this channel, I talk a lot about the computing power of a single chip or a single GPU, but at the scale of a data center, networking and wiring make a lot of difference. For example, according to Meta, about 30 to 50% of the overall elapsed time for AI workloads is spent in the network, waiting for the network.
Just imagine what if we could replace all this complex networking with photonic interconnects? Startups like Lightmatter and Ayar Labs are working on solving this problem by replacing all these networking switches with photonics. Ayar Labs, for example, is developing a solution that can be applied both to chiplets and data center networking.
Just last December, they closed a $155 million funding round that valued the company at more than 1 billion dollars. No surprise that NVIDIA, Broadcom, AMD, and Intel were among the investors.
So, in summary, all we discussed today points to a future where light will play a pivotal role in computing. For the moment, we are not focusing on quantum computing, if that is the question. It's not because I'm not believing in quantum computers; I believe the future compute ecosystem is going to have a multiple chiplet architecture, in my words.
This means you're having a CPU, a GPU, NPUs from us, and also QPUs—quantum processing units. But what I realized two years ago was that the time to a commercial product is way faster if we focus on these analog photonic computers because we understood them very well, and they're at the heart of a photonic quantum computer.
The mathematical operations that we are using in a photonic space are not so different from what we've been using when we build quantum computers. But for quantum computers, it's really unclear. You can't predict when you're going to have a system that has an economic advantage—not a scientific one.
We are all the time seeing scientific advantages with every new system, but it's hard to guess when there will be a system on the market that has a clear commercial advantage. On the same side, what we realized is that a lot of computations that were linked to quantum computers can already be very efficiently carried out by using an analog computer.
I think in the future, each of the computing paradigms I cover on this channel—whether it's analog, photonic, digital, probabilistic, quantum, or reversible chips—will find its own niche application. For example, for matrix multiply-accumulate operations for AI inference, we are likely to use photonic engines.
For quantum problems, we will rely on quantum computers, and for problems that require high accuracy and high precision, like banking transactions, we will still use our classical digital chips. I hope this video lightened up your day. Let me know, and now watch this video where I explain how computing backwards and reversible computing works.
This episode got a lot of attention, and I received a lot of positive feedback on this one, so check it out or watch another episode on probabilistic computing, where I explain how we can harness noise for computation. Must watch!
Thank you for watching till the end, and I will see you in the next episode. Ciao!