📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

"터보퀀트가 문제가 아니다" 낸드 주식까지 급락? 시장의 치명적 착각을 '수익 기회'로 바꾸는 법

위즈덤투스15:52

Transcription

The market is currently in a state of confusion, and the market is falling significantly. Asia is the biggest victim, and the KOSPI is showing crazy volatility in the meantime. Our country has such a large proportion of semiconductors that it can be considered a semiconductor ETF due to HBM. With the addition of TurboCon, we are now seeing a market that is completely weak and abundant in potential for decline, which we have been observing for several weeks. On Wednesday and Thursday, in the US market, companies like Sandisk, Western Digital, Seagate, and Micron experienced sharp declines, with Samsung Electronics and SK Hynix leading the way.

The reason for the market's confusion is partly due to the market's turmoil, but also because of the TurboQuant, a compression algorithm announced by Google Research, which specifically affects the memory sector. It claims to reduce the memory usage of AI models by six times and increase speed by eight times. The market, seeing this memory news, has concluded that memory demand will decrease, leading to a sell-off in memory-related stocks. Furthermore, on Thursday, Samsung Electronics and SK Hynix showed long bearish candles. With no signs of the Hormuz Strait situation resolving, the KOSPI saw declines across almost all sectors. Therefore, today, we will set aside the discussion about that issue and investigate whether the reaction in the memory semiconductor sector is normal or if there is an error.

First, we need to understand what TurboCon is. We will identify the companies that are truly affected and those that are not, and also look for historical patterns of this type of event. Understanding exactly what TurboCon is will be the most important step. To do that, we need to understand what KV Cache is. When we ask ChatGPT or Gemini questions, the conversation can become very long and convoluted, with one question leading to another. In such cases, the AI needs to remember the previous conversation. The space where this memory is stored is called KV Cache, which stands for Key-Value Cache. We can think of this as the AI's conversation notepad. The problem is that as the conversation gets longer, this notepad grows exponentially. For a model with tens of billions of parameters to process tens of thousands of tokens, the size of this notepad alone can consume over 80GB of GPU memory. This means that the notepad itself, not the model itself, becomes this large.

To make it easier to understand, I will continue to use food analogies. Imagine serving food at a restaurant. Even the best AI restaurant will have people waiting in line if there aren't enough plates, which are the KV Cache, to serve the food. Even though there is a lot of delicious food, if only a small amount can be served on each plate, it's a situation of severe shortage. This has been the memory bottleneck in the AI industry. TurboCon is a technology that makes the plates for serving food six times more efficient. If previously, a plate could only hold one type of food, now it can hold six types. The size of the plate hasn't increased, but technically, by compressing the KV Cache data, which was stored in 32-bit, to 4-bit, there was zero loss in accuracy. Google ran benchmarks on its open-source models. These benchmarks are essentially tests. TurboCon outperformed or matched existing compression methods in all tests. Specifically, in a test that involves finding a specific text within a large volume of tokens, up to 100,000 tokens, which is literally called "finding a needle in a haystack," it achieved 100% accuracy. They also announced that in 4-bit mode, the processing speed increased eightfold.

Now, the critically important question is this: TurboCon does not require separate training. It can be applied directly without retraining or fine-tuning existing models, making it a quite powerful tool, which is why the market reacted so sensitively. Of course, what we are really curious about is not the TurboCon technology itself, but how it affects HBM and whether memory semiconductor demand will decrease. I will summarize the report from Morgan Stanley and then share my thoughts. TurboCon is a technology that is only applied to KV Cache during the inference stage and does not affect the HBM capacity that stores the AI model's weights. It is a separate area from the memory required during the training process. Let me explain this more simply. AI is broadly divided into training and inference. Training is like a chef developing a new recipe. It requires a large countertop where all the ingredients can be placed. This countertop is HBM. However, AI inference is the process of serving food to customers using the completed recipe. In this context, KV Cache is like the order slip where the AI writes down what customers have ordered while serving. If there are too many customers, the number of order slips increases. That order slip is KV Cache. Therefore, TurboCon is a technology that efficiently manages these order slips, not a technology that eliminates the need for a countertop. Explaining it this way might make it easier to grasp.

Looking at the fundamentals of the HBM market, SK Hynix holds a 62% share of the global HBM market, with Micron and Samsung Electronics rapidly catching up. All three companies have already sold out their production for 2026. What's even more important is that with the generational shift to HBM 4, the bandwidth will increase by almost three times compared to two years ago. The technical difficulty will increase accordingly, and Samsung and SK Hynix are aiming to advance the mass production of HBM 4 to the first half of 2026, as it is a next-generation, high-value revenue stream for them. Furthermore, NVIDIA's next-generation GPU, the Rubin architecture, will support HBM 4. However, even though Samsung and SK Hynix are experiencing sharp stock declines, and Sandisk and Western Digital have seen significant drops, the main businesses of these two companies are NAND Flash and hardware. And there is no relation between KV Cache and Sandisk. NAND Flash is a storage device, not a cache memory for computation. It's not for computation. Therefore, TurboCon efficiently creates plates for serving food, but the market seems to have misinterpreted this as eliminating the need for countertops and has judged that the demand for refrigerators, which are storage devices like NAND Flash in the kitchen, will decrease. However, changing the size of the plates does not eliminate the need for refrigerators. So, is the market wrong? I have never thought the market was wrong, but the stock market is governed by the logic of money. However, there is a logic that suggests the market is wrong.

To summarize that opinion, it is a typical case where the market, seeing only the word "memory," has broadly sold off the entire sector. Firstly, the name TurboCon includes "Con," which sounds like a revolutionary technology, but for your information, it is not "Quantum." You should not be confused. TurboCon is a type of technology that has been developed in one direction, making memory more efficient. Also, regarding the eightfold performance improvement, there are counterarguments that the comparison system used was not the one currently in use, but rather a 9-year-old 32-bit model. Some institutions have even issued buy recommendations for Google and Micron, expecting them to fall due to this issue. Instead, what should be noted in this issue is HBF. HBF has eight to sixteen times the capacity and consumes 40% less power. In AI inference, HBF complements the problem of insufficient capacity with HBM alone as the KV Cache notepad grows larger. It is a next-generation revenue stream. Ironically, TurboCon compresses KV Cache sixfold, which expands the range that can be processed with existing HBM alone, and some analyses suggest that the adoption of HBF might be delayed. However, as the amount of computation that AI needs to process continues to increase exponentially, if tokens continue to expand at this rate, even the efficiency created now will be quickly consumed. This will lead to a direction where further optimization is needed through other hardware, existing HBM, HBF, or other software.

There is a famous concept in economics called Jevons Paradox. We can learn from history. When the fuel efficiency of the steam engine improved, everyone expected coal consumption to decrease. However, the reality was the opposite. As efficiency improved, the cost of using steam engines decreased, leading to their use in more places, and consequently, coal consumption surged. Before the invention of refrigeration technology, storing food was difficult. Restaurants had to buy ingredients that could be consumed on the same day. Then, refrigeration technology emerged, and food spoilage decreased dramatically. Initially, people thought that since storage efficiency improved, the amount of food purchased would decrease. However, the opposite happened. Thanks to refrigeration technology, global food distribution became possible, and tropical fruits could be eaten even in winter. The convenience store bento box industry emerged, and ultimately, food consumption exploded at an unprecedented rate.

Looking at TurboCon, if KV Cache is compressed sixfold, the same GPU can process six times longer contexts. In today's AI market, rather than buying less memory, this extra capacity will likely be used to increase the context from 1 million tokens to 6 million tokens, or to attempt multimodal inference, which was previously impossible due to cost, or to enhance the simultaneous processing capabilities of AI agents. There is also the possibility of reducing token costs, which can be quite expensive for consumers. Looking back at the DeepMind incident last year, when GPU efficiency technology was announced, the market thought computing demand would decrease, and semiconductor stocks, including Nvidia, plummeted. However, the outcome was the opposite. As efficiency improved, AI services exploded, and overall computing demand increased to an all-time high. Some institutions have even referred to this TurboCon issue as Google's "DeepMind Moment."

Once TurboCon is actually deployed, the AI inference costs for companies will be significantly reduced. It is likely that this cost reduction will lead to demand creation rather than demand destruction. A characteristic new market that could open up is on-device AI. This means that AI performance on devices like smartphones, robots, and electric cars will improve, without necessarily requiring cloud-based operation.

The discussion continues. The market is currently in a state of confusion, and the market is falling significantly. As Saudi Arabia has emerged as one of the key countries in the Iran conflict, and the US is showing signs of trying to disengage, I view it positively. However, looking at the current situation, there is a lot that has been disrupted, and it will take considerable time to resolve. After all, it takes time for an economy that has already been broken, whether it's energy or the real economy, to recover. Asia is the biggest victim, and the KOSPI is showing crazy volatility in the meantime. With the addition of TurboCon, our country, due to HBM, has such a large proportion of semiconductors that it can be considered a semiconductor ETF. With the emergence of TurboCon, we are now seeing a market that is completely weak and abundant in potential for decline, which we have been observing for several weeks.

However, stocks always lead the way, and often, when the market situation seems to be at its worst, it is actually the bottom. For a stock market rebound, mere talk is no longer enough. Even if there is talk, the president continues to make dovish remarks, but they are no longer effective. A rebound will occur if there are signs of improvement towards a genuine ceasefire, with cooperation from mediating countries. Looking at it now, Saudi Arabia has become a major variable, making things quite complicated. If there are signs of improvement towards a genuine ceasefire, then there will be a rebound. However, unless more dangerous measures are used, most of the negative news has already been released. Therefore, I expect some changes over the weekend, and for the next week or two, I will ignore minor news. If I were to buy anything, I would look for a rebound with trading volume, and the chart might be more accurate.

We all have stocks that we have studied individually, and markets that we have studied. This could be an opportunity to fill in sectors where our holdings were small. In March, foreigners sold over 22 trillion won worth of KOSPI stocks. They were net sellers for 13 out of 16 trading days, while individuals were net buyers for the rest. Institutions are constantly buying and selling. Therefore, with a lot of margin trading involved, the market's resilience cannot be seen as support, but rather as a vulnerability. While the fact that individual investors hold a lot of stocks is positive, in terms of market strength, a market where foreign holdings decrease and individual leverage increases is highly vulnerable to volatility. However, when a rebound occurs, it can be very rapid, and if there is further decline, a chain reaction of forced selling can occur, leading to a margin call situation. For these reasons, if the situation in the Strait of Hormuz de-escalates, or if there is concrete progress in negotiations, the Korean stock market, including the Asian markets, could recover quite rapidly. However, it is too risky to bet your life on this expectation. Therefore, this is a period where leverage must be managed.

To summarize my thoughts on TurboCon, just a few days ago, there was news that memory semiconductor demand would continue for the next few years, and this was repeatedly extended to two or three years, and prices were also high. However, looking at the progress of technology, just like the advancement of medicine, it follows the logic of capital. It seeks out profitable areas or areas where costs are lower. Therefore, this can be seen as one of the trends in seeking efficiency. And in any industry in the future, if something is too expensive, new alternatives or complementary solutions will inevitably emerge, which is how the capital market works. Therefore, there will inevitably be a short-term impact because market sentiment is so low. Currently, the extreme fear index is around 21, and for cryptocurrencies, it's in the teens. However, the Morgan Stanley report's theory is plausible, and if it is correct, then with actual order volumes and demand, a new trend will emerge. For memory semiconductors to rebound, the war situation needs to reach a point of resolution, and all negative news must be released. We need to approach the market with a long-term perspective for a rebound. It's also boring, and it leads to disappointment and despair, and ultimately, we realize the value of labor, which is a recurring pattern for us individual investors. So, stay strong. In the future, whenever such efficiency technologies emerge, the market will first try to interpret it as a decrease in demand. However, history consistently shows that efficiency lowers entry barriers, and lower barriers lead to explosive demand. Ultimately, it requires more infrastructure. Have a happy and wealthy day. Thank you.