📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

(中英字幕EngSub)輝達核彈級新消息 會令股價再飆高?新產品B300系列稱能夠降低推理成本三倍!《蕭若元:蕭氏新聞台》2024-12-30

memehongkong24:16

Transcription

Dear netizens,

I was surprised when I saw the news last night. NVIDIA released news again. Its new products in the AI world are something like a nuclear bomb. It means it will bring about a new revolution. It’s like switching from horse plowing to tractor plowing. That kind of speed and power consumption also makes NVIDIA’s dominance even more difficult to shake. No one else can compare. I even doubt that it is an ASIC chip; it is difficult to have this update speed and catch up with its efficiency. You should think about it seriously whether to continue producing ASIC chips.

Alright, we have this offer now. Four days to go. It's very cost-effective. Many people buy this hair dye; needless to say, never tried it so cheap. HKD 248 for HKD 248, HKD 18 for a Korean beef burger, and HKD 68 for Iwate Prefecture A5 beef. How cheap! HKD 78 for hotpot, HKD 88 for 1kg of Korean oxtail. This bottle of wine we have sold many times. Buy it at TV Mall for HKD 720, HKD 255, and three boxes of soup for HKD 88. Never tried. What else?

And red ginseng essence gift pack, these are 50% off, HKD 220. Giving gifts is good. Naoxiaole nourishes the brain, letting young people concentrate. The elderly will not lose their memory, HKD 125. Many people take these medicines for three high blood sugar; no one is so cheap, HKD 75. Eye protector has lutein, carotene, etc., for HKD 90. We do it at any cost.

According to many media reports, NVIDIA’s GTC 2025 conference in March next year will launch a new AI chip. News has leaked about the B300 and its corresponding platform, the GB300 server platform. A complete set of launches.

We rely on industry authority SemiAnalysis's report. The B300 series uses TSMC’s 4NP process, which is a 4nm-level process technology specially tuned for NVIDIA. I tell you simply, last time it was also 4nm. The latest B200 and Blackwell are also 4nm this time, but this 4nm is optimized.

Actually, every time this thing was launched, you can improve the architecture to optimize functionality and efficiency. This is a 4NP process with improved functions. It has not entered the 3NP process yet. I think next time it will be 3NP. It is so eager to launch B300 because B200 had a period of time when it was said to be too hot and was late to launch. As a result, Blackwell was not released at the end of the third season. Launched in Season 4.

It is now launching the B300 quickly. This time, the release also corresponds to AMD’s plan in 2025 for the launch of the new MI300 series. The MI300 series has always been intended to correspond to NVIDIA's products. The trick is AMD improved: first, it’s cheaper; second, it improved memory capacity. It has higher memory capacity than NVIDIA.

So NVIDIA has greatly upgraded the memory this time. The GPU power consumption of B300 is 200W higher than B200, going from 1,000W to 1,200W. The Superchip of G300 is 1,400W. Each GPU module on the B300 platform is 1,200W, increased by 200W, which is a 10% increase in power consumption.

What is the most important thing? It is floating point operations. B300's floating point operations are 50% more than B200. Several things about it are making leaps and bounds. Why is it said to be nuclear bomb level? Its latency is reduced a lot, and its memory has been greatly increased. Add up, its capabilities are roughly tripled. Tripled in three seasons.

Such speed! Ten years later, just multiply it by three times. Computing power is improved tens of thousands of times. Especially it optimizes specifically for training and inference of AI models. I don't know about those ASIC chips; can it be as efficient as it is? Improved so quickly.

Jen-Hsun Huang is great because he has too much money and too many powerful people. His company has so much manpower and material resources to do these studies, with so many patents. The B300 is stacked with 12 layers of HBM3E memory. 3E is the third generation of HBM3. These cache memories are stacked together. Now it uses 8 layers; next, it will use 12 layers.

It is still not HBM4, but it is already testing HBM4. HBM3E has matured, so the memory capacity increases a lot. That increases by 50%. Memory capacity increased from 192G to 298G. Just how much? Add another 50%. Operation speed increased by 50%.

And it also has dynamic reallocation. Those computing abilities distribute power consumption between GPU and CPU. This saves electricity. You have to know that everything is in three parts: has CPU, has GPU, has NVLink. NVLink connects those things together.

It has HBM memory. The whole thing becomes a card. Upgrade to better expanded bandwidth, effectively improving server performance and efficiency. Connect more together, so the network card does better. As a result, the entire system improves efficiency.

Its memory upgrade is very important for OpenAI o1/o3 large inference models. Reduce inference costs three times; this is important. If the cost is reduced three times, it cannot be used. I don’t know if TPU ones can do this.

Now every operation has a cost. The most expensive ones are those from o3, up to 20 dollars. It becomes $7. A few cents becomes worthless. So OpenAI charges fees, or free ones like Gemini and Cloud, and its cost plummeted. The cost of reasoning has plummeted, providing more services to everyone.

Upgrade via memory, B300 enables longer thought chains, reducing latency for each thought chain, and be able to handle larger batches of data. I need to explain it to you first.

How to say it? Now o3, its method is to break the problem down, develop a way of thinking first, and then think about it step by step. Therefore, if you want it to think, sometimes the answer is two or three minutes late. Not needed now; it is much faster.

Reduced latency, and the chain of thinking can be longer. Thinking steps can be longer, and I think it's important; it can handle larger batches of data. So what?

We often feel this involves, for example, you can enter very long articles and call it a summary. This involves how much batch data it has each time. Let’s say how many tokens there are. It now works with more samples, so it will be more accurate.

The computing unit of G300 NV72 uses 72 GPUs, with extremely low latency and shared memory resources. Because you have to store it in memory, how much can be stored here? 72 GPUs can be done simultaneously and quickly in this memory, accessing data to calculate.

New generation computing unit containing 72 GP300 is evaluated as the only one that can make OpenAI o1/o3 large inference models reach 100,000 tokens at high batch size. I haven’t seen 100,000 tokens yet.

100,000 tokens can handle very long articles and very long videos. It still has memory. For example, the generated doll image is still the same. No need to calculate separately. You can take that thing from before, continuing a long-term calculation.

Let me give a simple example. Because it has 100,000 tokens, the length of the video that can be generated will be much longer. Now usually 20 seconds, 30 seconds; it may be possible to reach 1 minute. As long as the token is long enough, a movie can be generated in the future.

Let me tell you simply, just look at the number of tokens. Of course not now. Such progress again, for example, if 10 million tokens are reached, you can generate an entire short film, a 10-minute short film.

Jen-Hsun Huang is awesome; he changed his sales strategy. In the past, he forced you to buy from his dealer. You don’t just buy his card; you have to buy his entire system, and also buy Dell, Supermicro, Quanta, Foxconn to assemble your machine.

Now it’s not like that. No longer providing entire reference motherboards or servers; just sell core components. What do the core components include? B300's SXM Puck module, Grace CPU, and host-managed controllers. He sold these components to you; you go back and assemble it however you like.

Much greater creative flexibility. Of course, you can also buy the whole assembly from him, but you can also assemble it however you like. In this way, many companies will be able to participate in Blackwell's supply chain.

I think Jen-Hsun Huang is very smart. If you want others to package and sell them together, you can make more money, but they will be prosecuted by antitrust agencies.

I sell it in pieces; no more antitrust charges. This is the first. Although it doesn’t look so profitable on the surface, you give people more freedom. Those people are more likely to use your product.

In fact, it is a double-edged sword. Do you dare to do this? Anyway, if the assembly is broken, he is responsible for it. You often have to worry about his accessories being incompatible or burned out.

But you can give him advice; he has more flexibility. As a result, more people will buy your product. Blackwell-based machines are expected to be more readily available because you broke those who installed it.

Now they are mainly Dell, Foxconn, Mechaomei, and Quanta. The monopoly of these companies lets others compete with these companies. Their profits will fall. NVIDIA will serve as a cloud service provider and OEM, providing more design freedom, letting them build their own Blackwell systems.

I very much agree with Jen-Hsun Huang’s approach. According to reports, Samsung Electronics in the next nine months cannot enter the supply chain of GB200 and GB300. Too bad.

Hynix, who was once the second child all his life, now accounts for 50% of the HBM market. Samsung is still in the early stages, unable to pass verification for GB200 and GB300. Micron only has a few percentage points. Hynix almost has a monopoly on this matter.

Why is Samsung experiencing earthquakes? It’s because Samsung is the largest memory manufacturer in the world. Because it looked down on HBM, not many people bought it at first. As a result, it did not develop and lagged behind.

NVIDIA simultaneously plans to release graphics cards based on Blackwell architecture. This is important for gamers. He will use the GB202 chip, which is a big film, 744 mm². The graphics card is 22% larger, playing games with 8K imaging, and the number of frames per second has increased.

Due to U.S. export controls, NVIDIA can't help it. Without government permission, it cannot sell H100, H200, H800 GPUs to China. Initially, sales of H100 and A100 were not allowed. He just produces H800, but now H800 is not allowed to be sold either.

NVIDIA can only provide H20, but business is very good. Why? Do you understand? Once used, there are only two options: either use H20 or use Shengteng. Shengteng is said to be better than H20 because H20 reduces the functions to only one-fifth.

But Shengteng proves unreliable; actual operation is unreliable. Shengteng cannot supply it at all. It turns out that some of them are chips supplied by TSMC for computing power. Supply is now stopped, forcing the purchase of H20.

According to semiconductor industry analysts, the downgraded version of H20 has a quarterly growth rate of 50% in China. Scared you to death! This season is 50% more than last season.

Is it funny? A 50% increase in one quarter? What business? Autumn sells 50% more than summer, and winter sells 50% more. Is that okay? The growth rate of H100 is only 25%, which is already very powerful.

H20 performed very well in the Chinese market. Very funny. People say NVIDIA can no longer be sold in China. The performance of H20 is nowhere near as good as H100, but it will bring in tens of billions of dollars in revenue. Better than robbery!

It is estimated that the gross profit margin of B300 may exceed 75%. It is now 75%, and it is more profitable. Maybe 80%. Analysis points out that although China’s local Biren Technology is sanctioned, Moore Thread is also sanctioned.

If Huawei’s Shengteng cannot do 7 nanometers, it has to step back to do 14nm. Competing with 14nm and 4nm is not at the same level of measurement.

It’s like me boxing with Tyson. NVIDIA's stock reached a maximum of 150 yuan, 152 yuan. It has been wandering here for a while because there are two pieces of bad news.

One of the bad news is that China sanctions it. China is investigating it because it was purchased in Israel from Mellanox Corporation, which is the NVLink used in its chip. This violates monopoly.

So now it sells them apart. This is not a bundled sale. This is why we sell them apart. NVLink is from Mellanox.

The second is whether the U.S. government will further restrict NVIDIA's sales to China and whether it is investigating whether NVIDIA’s dealer secretly sells it to China. Now every chip needs to be tracked.

These are all restraints on NVIDIA’s price. But once new chips are announced, I think it's going to go up tonight.