Transcription
The AI industry is heating up as DeepSeek rushes to launch its new AI model, going all-in on AI. If you're unfamiliar with DeepSeek's models, you likely haven't been in the AI space. This company completely halted the AI industry by creating high-quality frontier models at a fraction of the cost. Apparently, they're now launching their second iteration, R2, which is even more concerning given the current state of the AI industry.
This article states DeepSeek plans to release R2 in early May, aiming for the earliest possible release. This is interesting because DeepSeek recently leapfrogged other companies in price-to-output and quality. This situation may revolutionize the AI industry. If R2 matches the quality of recent models from Frontier Labs like Anthropic and OpenAI, at a fraction of the cost, it could seriously impact Western AI companies.
Currently, developers and users don't show strong brand loyalty. While brands like ChatGPT, Claude, and those associated with Elon Musk exist, developers prioritize cheaper prices to reduce costs for their apps and software. A key focus area is coding; this article mentions the company hopes the new model will improve coding and reason in languages beyond English. Details of the accelerated timeline haven't been previously reported.
I believe coding is one of AI's hardest tasks. The widespread use of Claude 3.5/Sonnet is significant. If DeepSeek surpasses Claude 3.5 or 3.6 in coding, it would be incredibly interesting, potentially eliminating Claude's market share. This wouldn't be the first time DeepSeek disrupted the AI industry.
Recent coding benchmarks, like the agentic coding evaluation using Devon (an AI agent framework for software developers working with large enterprises), show Sonnet 3.7 performing best. DeepSeek R1 scored 51%, significantly behind other frontier models. However, if DeepSeek R2 surpasses Sonnet 3.7, it would be pivotal. Coding is crucial for app and software development.
The Ada Polygot coding benchmark evaluates AI's ability to translate natural coding requests into executable code passing unit tests. It analyzes code, assessing coding ability and the capacity to edit and format code for integration. DeepSeek R1 plus Claude Sonnet 3.5, an agentic framework, is one of the cheapest options, costing $13 and performing relatively well compared to others costing $18, $100, and $36.
While accuracy matters, many applications will prioritize cost. DeepSeek's low price makes it attractive, even if accuracy is not its strongest point. The LMS Y Arena, a qualitative benchmark using human ranking, places DeepSeek R1 third or fifth, a relatively high position for a non-thinking model; even Claude 3.7/Sonnet ranks lower. This benchmark is subjective, favoring models like GPT 4.5. DeepSeek R2's performance here will be interesting.
These benchmarks aren't necessarily definitive, but they indicate quality. DeepSeek's rapid turnaround, potentially releasing R2 as early as May, is remarkable. A company releasing models quickly and cheaply has a significant advantage; its feedback loop for model improvement accelerates.
One worrying aspect of R2 is DeepSeek's recent Twitter announcement during their open-source week. They revealed a 545% profit margin, highlighting their profitability. This is surprising because many AI companies are loss-making, requiring large funding rounds (e.g., OpenAI's $5 billion loss this year on $3.7 billion revenue). While OpenAI's revenue is growing, speculation suggests Western tech companies overspend. I believe significant upfront investment in AI is necessary for later profitability. This raises questions about the stock market's reaction.
Models like GPT 4.5 require significant investment in computational resources and infrastructure; GPT 4 reportedly cost over $100 million to train. These expenses contribute to losses. OpenAI's investment in Stargate for massive data centers further increases costs. Despite losses, OpenAI continues to attract substantial investment due to its long-term potential and market leadership. Their fundraiser rounds are consistently oversubscribed.
Interestingly, OpenAI is losing money on OpenAI Pro subscriptions, despite higher-than-expected usage. Many overlook the decreasing price of tokens yearly. GPT 4 initially cost $36 per million tokens, then $40, $1, $7, $4, and finally GPT 4 mini at 25 cents. This trend will continue as the AI industry matures. Some question profitability with continued price undercutting, but this could benefit OpenAI, offering a nearly free service with a strong brand.
DeepSeek's R2 rollout might be shaky. The US government considers AI leadership a national priority; R2's release may concern Chinese authorities and companies already integrating DeepSeek models. Some countries, including South Korea and Australia, have blocked access for government employees, fearing it as a spy tool or national security risk; this could lead to a TikTok-like ban. China's rapid progress in AI, potentially surpassing the West, should be a concern.
DeepSeek's rapid advancement stems from a different management style. A former employee described a flat management style, with everyone working on the same level, enabling smoother and more effective collaboration compared to hierarchical structures. This style might become more common.
Are you excited about the competition? I think it's beneficial, forcing OpenAI to speed up releases. R2's potential release next month would be incredibly interesting, given the typical long release cycles for new frontier models. I'm eager to see what happens next.