Transcription
There is a Colossus effect underway in computing. Something so massive it's reshaping the entire energy industry, computing, and even geopolitics. Let me show you.
Back in March, this was an abandoned factory. Just 6 months later, Colossus 2 stands here. A machine the size of a city, built not for people, but for intelligence. It's nearly a million GPUs under one roof, eating up a gigawatt of power. This is the story of the world's first AI gigafactory and the insane journey it took to build it.
I am an engineer who spent a decade building the most critical chips for this technology. That's why I'm so excited about this particular story and its impact. Subscribe to the channel right now, and let me explain.
Today, Amazon, Microsoft, OpenAI, and Meta, everyone is pouring billions into AI factories because the race isn't just about smarter models anymore. Every new generation needs at least 10 times more compute power. So, the hyperscalers are no longer just running datacenters. They are building power plants to fuel AI faster and cheaper. Amazon needs it to defend the cloud. OpenAI to push the frontier. Meta just to catch up. So, slowly but surely, AI labs are turning into energy companies.
And now comes xAI. Last September, they shocked the industry with a Colossus 1. When, just in 122 days, they turned an empty shell into one of the world's largest AI training sites. Then it took them 19 days, just 19 days, to wire and to deploy 100,000 NVIDIA GPUs. That alone was historic. The speed, the scale. No one in the industry had ever built like this before.
But today's story is about something far bigger. Colossus 2. And on paper, it might look like just another datacenter, but in reality, it's four giants rolled in one. Just imagine, at peak, it will draw up to 1.2 gigawatts of power. And that's a lot. That's enough to keep the lights on in more than 2 million homes.
So, what does it take to build a gigawatt datacenter? And here it's getting interesting, because you actually don't start with GPUs. You start with power, in every sense of this word. Without gigawatts of stable energy, the racks stay dark. And that's where Colossus 2 becomes fascinating.
Imagine you want to build a datacenter. The first challenge is simple and brutal. Where do you find a gigawatt of power? And even if you secure that much power, how do you push it into one building without frying the grid? At this scale, you have to actually build a new electrical backbone. You need substations, switch gear, transformers, backup, and distribution lines feeding straight into the racks. This means when you're building an AI datacenter, the first thing you have to build is a power factory. And that's where the drama of Colossus 2 begins.
Colossus needed at least a full gigawatt, with scaling up to two. The city of Memphis, Tennessee, could barely spare 50 megawatt. Colossus needed a thousand. Regulators said no. Communities pushed back. For a moment, the project looked dead, until xAI crossed the state line.
Here, they did something extreme. They found a solution just across the state line in Southaven, Mississippi. Just a few miles from Memphis, the site was a former Duke Energy power plant with gas pipelines and grid connections already in place, but gas turbines long gone. So, they decided to acquire this gas plant, which was perfect for a rebuild.
Then came the bold move, buying natural gas turbines overseas and shipping it to the US. Basically, they found the former power plant abroad in Europe and broke it into container-sized modules, shipped it to the US, and reassembled it in Mississippi. Sounds crazy, but it was far faster than building it in a conventional way. And it worked. Within just a couple of months, the turbines were spinning. Seven Titan class units, each over 35 megawatts. Together, they brought approximately 460 megawatts online.
What's interesting, solving power was only half of the bottle. The real challenge isn't just delivering electricity, it's in keeping it perfectly stable. GPUs spike in milliseconds, and such millisecond surge can drop voltage and, in the worst case, cascade through an entire hall. And restarting these runs isn't just annoying. It burns millions of dollars wasted in compute.
Ironically, NVIDIA advancements make this story worse. What actually makes it worse is density. Because these aren't just ordinary racks anymore. We're quite used to racks drawing about 50 kW of power. But with new NVIDIA Blackwell chips, power draw jumps from 50 kW to 130 kW per rack. And at this scale, each rack pulls as much electricity as a small neighborhood. And keeping this supply uninterrupted was one of the toughest problems Colossus 2 had to solve for.
Colossus 2 pairs its turbines with 168 Tesla megapacks. You can think of them as giant rechargeable batteries. They soak energy when demand is low and release it instantly when demand surges. What's essential to understand here is that power doesn't flow straight from the grid to the racks. It moves through a layered system built for stability. The grid and backup diesel generators feed into the megapacks, which smooth out spikes, and only then electricity is routed through power distribution and into the rack arrays. The result: steady, predictable power to hundreds of thousands of GPUs. And with that, the power factory for Colossus 2 came alive. Without that bold move across the border, Colossus 2 would simply not exist.
And power is just the first pillar of four, and the most invisible to most of us. But in effort, complexity, and cost, it's enormous. At Colossus 2, power infrastructure alone is up to 20% of the budget. Just think about it. It's billions spent long before a single server is installed.
Now, looking at the AI race more broadly, the real bottleneck is actually in power. Access to stable and abundant energy now will determine who can train the best models and ultimately who wins or falls behind. And that's why hyperscalers and AI labs are now quietly turning into energy companies. Microsoft is restarting the Three Mile Island nuclear power plant to secure about 850 megawatts of clean energy for AI. Google is funding three new nuclear power plants for the very same reason. And that's crazy, because just 10 years ago, the biggest datacenters ran on just tens of megawatts. And today, AI demands gigawatts of energy, enough to power entire cities, consumed just to keep models running. And the crazy part is that hyperscalers are now on track to own more nuclear power capacity than some nuclear nations. And I think long term, the one who can lock down the most energy will shape the balance of power, not just at the level of hyperscalers, but at the level of entire nations. And I'm really curious to read your take on this in the comment section below.
Now, the first challenge is solved: power. But the moment that electricity hits the servers, it creates the next enemy: heat. And the scale is staggering. A 1 gigawatt datacenter throws off one gigawatt of heat. Just imagine the output of an entire industrial power plant trapped inside these four walls. That's so much heat. It's almost hard to believe. Researchers estimate that a 1 gigawatt datacenter dumps enough heat so that if you capture it with a heat recovery system, you could generate roughly 15 megawatts of secondary power. In other words, the byproduct of keeping all these GPUs alive could power a small city. But in reality, almost all of it has to be cooled away. And just cooling itself eats up roughly 30% of the total energy bill. And this problem is getting worse and worse from year to year as AI racks get denser.
If cooling stops, even for 2 minutes, GPUs first slow down, then racks crash, and parts start to take permanent damage. I lived it from the other side. And when you design chips, you quickly realize that the performance isn't just defined by clever circuit design. It depends as much on keeping them cool, on the code that runs on them, and on the networks that tie them together. And the industry has already learned this the hard way. Meta demolished an entire datacenter mid-construction because its design couldn't handle the density required for modern AI. The techniques that worked just fine for social media and YouTube streaming collapsed instantly at Blackwell densities.
So, cooling became the second great battle of Colossus 2. They had to find a way to deal with all that heat. And at this scale, forget air cooling. What you need is water, and a lot of it. Massive AI datacenters are infamous for draining local water supplies. One average datacenter can consume millions of gallons per day, competing with cities, farms, and entire regions. In drought-prone areas, that's a nightmare. And many of America's biggest datacenters are rising in the worst possible places, in water-scarce regions like Arizona and Nevada, where every gallon is already contested.
xAI's solution was extreme, and I really liked it. They had to build a second factory, a water factory, right on the Colossus 2 campus in Memphis. They built the world's largest ceramic membrane bioreactor. Basically, a wastewater treatment plant. Instead of draining Memphis's underground drinking water, Colossus takes the dirty wastewater from the city and turns it into pure water for cooling. The scale is Colossus: 13 million gallons a day of recycled water. That's actually even more than the datacenter needs, and that's a rare case when an AI datacenter doesn't drain the local water supply but instead improves the balance. And I'm totally for that.
Next, we will uncover two more factories hidden inside the datacenter. Plus, the genius of xAI's approach because, honestly, making it all work together at the Colossus scale isn't just engineering, it's rocket science.
Now, we have about 90 days left in 2025. And if I look back at my goals, one of the biggest ones was simple: keep learning. And this year delivered on that. From building my tech startup to spending summer at Stanford, none of that was easy, but it was worth it. But the truth is, we all have to catch up on AI skills, and there is still time for that. Keeping up with AI is essential because already now, people who know how to use AI are replacing those who don't. That's why I recommend you to join me at this 2-day AI workshop by Outskill. Outskill is the world's first AI-focused education platform, and they're hosting a live workshop this Saturday and Sunday, 10:00 a.m. to 7:00 p.m. EST. Over 10 million people worldwide have already attended this training. Many boosted their careers or even built businesses. This course is rated 4.9 out of five on Trustpilot. This course normally costs $395, but I partnered with Outskill to provide 1,000 free seats for you. It's valid for the next 48 hours as a part of pre-Halloween sale. It's a fully hands-on AI training from experts at Microsoft and NVIDIA, people who've actually built this industry. In just 16 hours, you will cover more than 10 AI tools, AI in Excel, Sheets, and Presentations, building AI agents, automating your workflows, and more. You will also get an opportunity to connect with other builders, plus bonuses worth over $5,000 if you attend both days of the training. This includes a prompt bible and the personalized AI toolkit. This will enable you to enter 2026 ahead of the curve. Register right now through the link below or scan the QR code here. And thank you Outskill for sponsoring this episode.
How do you cool a million GPUs? Inside Colossus 2, cooling works like a factory assembly line, moving heat step by step until it's pushed out of the building. Here is how it works. First, every rack has its own plumbing. Coolant flows through thin channels called manifolds, feeding cold liquid directly into cold plates, bolted onto every GPU, CPU, and memory chip. Think of these cold plates as mini radiators that suck the heat straight off the silicon. Here, pumps at the bottom of the rack keep the liquid moving in a closed loop. As it passes through the chips, the liquid heats up to about 45° C or 113° F. Sensors in the rack constantly track flow and temperature, because a stall here could fry tens of thousands of dollars worth of GPUs in seconds. The warm coolant flows to a coolant distribution unit, which transfers the heat to the building's bigger chilled water system, and that's where the custom second stage kicks in. 119 massive air-cooled chillers outside the data halls. They work like giant car radiators, blowing air to strip 5 to 7° C or 9 to 13° F from the water each cycle. In this way, the water is cooled back down to around 38 to 40° C or 100° F, and the water is sent back inside this hybrid system. Liquid on the chips, chillers outside, allows Colossus 2 to keep GPUs packed tighter than ever before. At this scale, cooling isn't optional. It's survival. And roughly 15% of the total investment went into building up this cooling system, roughly $3 billion. As you can see, keeping this datacenter cool is just as hard as powering it. Building a power plant and a cooling plant are both colossal challenges.
And yet, that's still not what gives Colossus 2 its edge. The true advantage of Colossus 2 is hidden in the fabric, in the network. Because modern AI datacenters are not datacenters at all. These aren't just racks of GPUs. It's one massive, unified supercomputer. 500,000 GPUs mean nothing if they can't work as one giant GPU. And that's the power of the network. And that's where Colossus 2 pulls away from everyone else.
If you open a rack in Colossus 2, you will see NVIDIA NVLink stitching GPUs together. NVLink 72 packs 72 GPUs so tightly they behave like a single giant processor. But the real challenge isn't inside the rack. It's scaling beyond it. Picture this: thousands of racks spread across multiple halls, each packed with GPUs. But unless they move data in perfect sync, you get chaos. That's the edge of Colossus 2 in how these GPUs are interconnected. NVIDIA Spectrum-X Ethernet fabric links everything at terabit speeds. So, over 550,000 GPUs can act like a single brain. And forget the 1 Gbit per second Ethernet at home. Here, each link runs at 400 Gbit per second, with every GPU server pushing up to 3.6 Tbit per second of bandwidth. Smart traffic control keeps congestion down and throughput above 90% across the entire datacenter.
And here is why it matters. Because here, at this scale, just like in life, timing is everything. If gradients or parameters arrive late, even by milliseconds, their tail GPUs end up crunching outdated data, and efficiency drops by half. That's why Colossus 2 bets on Spectrum-X Ethernet, re-engineered for AI workloads. And under the hood, as always, the magic comes from chips. Here, Spectrum-X's special chip is co-packaged very close with optics, which cut electrical losses and keep latency predictable. At each server, special NVIDIA chips called Bluefield Data Processing Units handle the networking, storage, and security. So, the GPUs can only focus on actual computing. Think of it this way: GPUs deliver the compute, but the network is what keeps it all in sync, basically turning all this hardware into a single coherent supercomputer. So, from outside, it might look like a warehouse, but inside, it's the largest AI supercomputer on Earth.
And since high-speed datacenter interconnect is exactly what my startup, so my team and I are building silicon for this application. That's why I'm working on a deep dive episode on the key technologies that enable it. Make sure to subscribe to the channel not to miss this special episode.
As you can see, you don't just build a datacenter. You build four factories at once: power, cooling, networking, and finally, compute. Miss one, and everything collapses like pulling one leg off a chair. And when it comes to cost, compute is by far the biggest factory of all.
To power a modern AI datacenter like Colossus 2, everything starts with GPUs. And in this case, that means NVIDIA. Colossus 2 began with 200,000 Hopper GPUs. Now it's been expanded with another 350,000 of NVIDIA's latest GB200 and GB300 Blackwell GPUs. The GB300, known as the Blackwell Ultra, is built on TSMC's 4 nm process and pushes over 20 PFLOPS of FP4 compute per chip. Think of it as strapping a rocket engine onto every rack. At launch, Colossus 2 will deliver 50 exaflops of compute. That's about seven times more compute than the world's top 10 fastest supercomputers combined. A true Colossus of compute.
But raw compute still needs orchestration. That's where AMD's EPYC and Intel Xeon CPUs come in, handling control, scheduling, and the background workloads that keep hundreds of thousands of GPUs working as one. Then come the networking chips that we've just discussed before. And of course, none of that would work without high bandwidth memory. The high bandwidth memory used in Colossus 2 is predominantly from SK Hynix and it pumps terabytes of data per second into the compute cores, ensuring those GPUs stay fully utilized. And finally, this pipeline needs to be fed. So, petabytes of SSD storage stream training data at speeds fast enough to keep every rack busy. This is the silicon stack behind Colossus 2, and that's a lot of silicon.
Ironically, that's not yet the finish line. The long-term target is to scale Colossus 2 to 1 million GPUs, all stitched into a single working machine. Here, they are using NVIDIA GPUs. And you may wonder what happened to Tesla's DOJO supercomputer. I was watching this project very closely for years, and it was a genuine moonshot, but unfortunately, they shut it down. Building a supercomputer, silicon, the entire stack from scratch isn't an easy task. You need a lot of particular talent, which is not easy to find. You need a lot of investment, and you need a clear business case. Most likely, to justify this investment, very often you need to sell it externally. I admire Tesla for trying, and if you want to know more on what happened, I will share a link to my LinkedIn post in the description box below.
Now, let's move on. Colossus 2 is built to train xAI's largest models, starting with Grok, then powering the next generation of Tesla's full self-driving system and Optimus robot training. The strategy is simple: in AI, the lab with access to the most compute moves fastest and sets the pace for everyone else. That's why Colossus 2 is way more than a datacenter. It's a $20 billion strategic bet here. Roughly half of the total investment went into silicon and hardware infrastructure. And after discussing all these four pillars, it's crazy to think about this. That explains why NVIDIA is the most valuable company in the world. And the second biggest bill is power. Everything else is just to keep it alive.
Already, in some weeks from now, in October this year, Colossus 2 will light up. The pace is staggering. Despite all the battles, they managed to build it in just six months, which typically takes other companies more than 16. And that's exciting, but also terrifying, because facilities like this, they don't just crunch numbers and do math. They are consuming huge amounts of electricity and water.
And of course, Colossus 2 isn't alone. Switch's Citadel Campus in Nevada is a monster on its own: 650 megawatts across 1.3 million square feet, operating roughly 250,000 GPUs. But Colossus 2 isn't just bigger. It's an entirely new class of AI datacenter. And yet, the competition is right behind. OpenAI and Microsoft are building Stargate with total plant capacity reaching nearly 10 gigawatts of power, and the Stargate flagship campus near Abilene, Texas, targeting 1 gigawatt already by 2026. Apart from Colossus 2, hundreds more AI datacenters are breaking ground in the US, in China, across Europe, each one demanding more land, more steel, more turbines, and more energy. Actually, hyperscalers now command more energy than some of the countries.
The truth is, right now we are building a whole new industrial layer: power-hungry AI factories that depend on the same resource every community and every nation depends on. And the costs here aren't just dollars. It can be that eventually, we will decide who gets this energy: people or machines. So, whether you care about cheaper AI tools, or rising energy bills, or who wins the next technological race, what's happening inside Colossus 2 will touch your life. And that's exactly why I'm making this episode, why I'm telling this story, because the impact of this is thrilling.
I want to leave you with this one. Let's hope that all of this will pay off, and AI will help us to discover new energy sources, new drugs, and expand human lifespan. Whether it will or not, only time will tell. Now, drop your take in the comment section below.
If you enjoyed this episode, let me know, because this one took a lot of work, a lot of research from my side, and scripting, and we are obsessed with raising the quality bar with every single episode. So, if you enjoyed it, the best way to support us is by sharing this video on social media and with your friends, and subscribing to the channel. Remember to connect with me on LinkedIn and subscribe to the newsletter. You can find all the links in the description box below. If you enjoyed this episode, you will most likely love my breakdown on what it takes to build a semiconductor fab. I will link it here. Thank you for your support, and I will see you there. Ciao.