Transcription
Hello everyone. In this issue, the most important fresh news about artificial intelligence and technologies. Open released GPT 53 Codex, a new level agent monster, which no longer just writes code, but confidently works as a digital employee. Anthropic presented CLD OPUS 4.6 update, which turns the model into a universal professional assistant with powerful planning, analysis, coding, and a huge context. Around OPUS 4.6, the company revealed rare details of the model's internal behavior and exhibits its own tricks and instincts. The Cursor team launched 1,000 agents that autonomously created a web browser for several weeks. About this and much more in this issue, watch this video until the end so as not to miss anything. Google released a selection of free AI courses with certificates, no payment, no subscriptions. I have collected the 10 best. Go and watch the selection in my Telegram channel via the link in the description under the video. In the free courses, you will learn how generative AI works and how it differs from classic machine learning. A foundation for working with artificial intelligence in daily tasks, prompt engineering, and working with GNAI in real cases, as well as much more. Everything is in my Telegram via the link in the description under the video. The company Openi presented the GPT 53 Codex model, an updated version of its agent tool, which now goes far beyond the usual understanding and for writing code. This release has been one of the most ambitious in the history of Codex. The model has significantly accelerated, gained new abilities, and confidently approached the format of a full-fledged digital employee capable of performing complex tasks on a computer from start to finish. The main changes are that GPT53 Codex combines the best aspects of previous systems: programming, data analysis, professional knowledge, and long-term agent behavior. At the same time, the model has become 25% faster, which noticeably changes the interaction dynamics. Code is written not only more accurately but also noticeably faster, and complex action chains are executed with almost no delays. Inside Codex, the new version feels at home, reacts faster, gives more confident solutions, and almost never loses context even with long tasks when working with thousands and tens of thousands of tokens. According to the developers, GPT 53 Codex became the first model to play a key role in its own creation. The agent helped the team monitor training, found errors in the infrastructure, analyzed user interactions, suggested fixes, and even generated quality reports. Researchers note that their work has changed almost completely in recent months. Codex proved useful in so many internal processes that its contribution became difficult to list. It was used as an analysis tool, a debugging assistant, and an operator, managing complex systems in real-time. The company paid special attention to frontend and application tasks. On the SV Bench Pro benchmark, which assesses the ability to solve real engineering problems in multilingual codebases, GPT 53 Codex set a new record. A similar picture is observed in Terminal BCH 2.0, where terminal work skills are tested. The model confidently surpassed its predecessors and demonstrated tens of percent more accurate command execution. Even more convincing results are seen in the OS World Verified test, where AI must perform tasks within a full visual computer environment. There, GPT 53 Codex is already very close to the human level, while the previous generation barely reached half the required accuracy. Developers emphasize that this is not just about improving individual functions, but about transitioning to a new type of behavior. The model begins to behave like a universal computer agent that understands user intentions and can independently plan work many steps ahead. In one of the tests, the team asked GPT 53 Codex to create a game, and the agent not only handled the code but also dozens of iterations of improvements, working in a completely autonomous mode. It read reviews, fixed bugs, changed mechanics, and brought projects to a stable state. Similar capabilities were also manifested in web development. There, the model began to better understand what was expected of it. Even vague requests begin to turn into functional websites with competent defaults, correct structures, and thoughtful logic. In the demonstration, GBT 53 Codex itself restructured the tariffs so that an annual subscription looked like a discount, rather than a multiplication of the price by 12. It automatically created a carousel of reviews and designed the page so that it looked like a finished product, not a draft. But the expansion of capabilities is not limited to code. The model shows high results on GDP Wall, a professional benchmark that assesses the quality of task execution in combined fields of activity. This includes working with presentations, spreadsheets, economic documents, training materials, and many other products created by specialists in real companies. New accuracy is noticeable in documents generated by Codex. It writes coherent instructions, forms understandable visual solutions, analyzes data, and collects materials as an experienced employee would. In parallel, GPT53 Codex became the first tool that, as part of preparing the system for real-world operation, received High Capability status in cybersecurity. The team integrated the model into training for vulnerability detection and defensive infrastructure creation. To prevent risks, the company built a double loop. On the one hand, the model helps protect systems, and on the other hand, mechanisms are included to limit dangerous scenarios and monitor misuse. This strategy includes safe operation training, monitoring, access only for trusted tasks, and separate tools for threat detection. Developers continue to expand the ecosystem, including the Trust Access for Cyber program. A pilot project aimed at strengthening defensive technologies and working with external researchers. Special attention is paid to Open Source. Codex is already helping to check the code of popular libraries, where vulnerabilities are regularly discovered, as recently happened in one of the NextGS projects. To expand capabilities, the infrastructure of the future is being created. The model was trained and works on GPU systems developed jointly with Nvidia, which has increased speed and stability. The update also states that API access to GPT 53 Codex will be available later, and currently, the model is available to users of paid ChatGPT plans in the Codex application, web version, IDE solutions, and command-line interfaces. Concluding the presentation, the developers note that GPT 53 Codex is no longer just a tool for writing code. It is becoming a universal task executor on a computer, from programming and deployment to research, information analysis, and managing complex projects. This is a movement towards a single agent that understands context, can explain its actions, interacts in real-time, and is capable of long-term work without loss of quality. Codex is transforming from a narrow tool into the foundation of a broader class of digital assistants that will work alongside humans and take on an increasingly significant part of professional processes. The company ATIC presented a new version of its flagship model Clot Opus 4.6. Emphasizing that this is the most significant update in the line recently. Developers call the model not just an improved tool, but a step towards more reliable and independent work in a large number of professional tasks. OPС 4.6 is intended to be a universal assistant for users, capable of simultaneously planning, analyzing, writing code, working with documents, and retaining vast amounts of information in memory. The main innovation is a noticeable increase in the model's ability to handle complex multi-step tasks. According to Anthropic, OPUS 4.6 is much more attentive and careful when performing long operations. The model has become better at maintaining logical reasoning and frequently returns to its own previous steps to ensure the correctness of the answer. Another important improvement concerns working with large codebases. The model has become stable in long sessions, more accurately detecting errors, and is capable of independently reviewing code, noticing nuances that the previous version missed. For the first time in the OPUS line, the model has received a context of 1 million tokens in beta mode, allowing it to analyze colossal amounts of data without losing the thread of conversation. Anthropic emphasizes that OPUS 4.6 can apply these improvements in many common work scenarios. The model handles financial calculations, research tasks, analysis of large documents, creation of presentations, and spreadsheets. In the Corкlot system, it can work autonomously and perform multiple tasks simultaneously, making the model more useful in daily professional activities. According to internal and external evaluations, OPUS 4.6 shows results that developers call industry-leading. In terminal coding tests according to the Terminal Bench 2.0 system, the model ranks first among all others, and in the complex interdisciplinary exam Humanity Last Exam, OPUS 4.6 also leads, demonstrating a high level of reasoning in economics, law, and analytics. In GTPVLA tests, which assess the model's ability to perform economically significant office tasks, the new version significantly surpasses not only competitors but also its predecessor. On Bрауз Camp, which tests the ability to find complex information on the internet, the model also holds the highest position. Developers also note the overall stability of OPUS 4.6. Despite the increase in capabilities, the model has not become less safe. On the contrary, during a large audit, it showed the lowest level of incorrect behavior among all previous Claude generations. Refusals of harmless requests, so-called overfuses, have decreased, making interaction with the model more predictable. Anthropic emphasizes that the combination of high power and low propensity for erroneous or harmful behavior is one of the main achievements of the new version. The model's results in long-context tasks deserve special attention. OPUS 4.6 has received significant improvement in its ability to find the necessary information within vast amounts of text. In the MAC V2 test, which checks the ability to find small pieces of data hidden in a million-token context, the new model confidently surpasses OPС 4.5. The difference is between 76% and 18.5%, demonstrating a qualitative leap. This indicates that the model is now capable of not only retaining a large context but also using it with almost no loss of quality. For real-world cases like document analysis, contracts, and code archives, this is a significant advantage. Additionally, the new version has become better at diagnosing software errors. According to the Open RCA test, the model determines the causes of failures significantly more accurately, making it useful in developing complex systems. In cybersecurity-related tasks, OPUS 4.6 also shows steady growth. Anthropic has even created six special prompts to test how the model reacts to potentially dangerous requests. This helps track unexpected scenarios and reduce the risks of malicious use of the model. The team has also added new control mechanisms for developers. Adaptive thinking is now available in AI, where the model can decide for itself how deeply it needs to reason when performing a task. More levels of the parameter have been added, which can be used to adjust the intensity of reasoning, and therefore, the speed and cost of work. For long processes, a context compression system is provided, allowing the model to automatically remove outdated dialogue fragments to avoid exceeding limits. Additionally, OPС 4.6 can output large volumes of responses up to 128,000 tokens. Developers have also improved Claude's integration with familiar office programs. In Excel, the model has learned to better convert unstructured data and perform multi-step operations in a single pass. In PowerPoint, it creates presentations by generating structure, layouts, and a unified visual style. All of this helps use it as a full-fledged work tool in routine tasks. Finally, Anthropic notes that during the creation of Opel 4.6, the most extensive safety testing series in the company's history was conducted. This included studies on the impact on user well-being, new checks on the model's ability to refuse dangerous actions, additional methods for interpreting AI behavior, and updated scenarios for assessing potential harm. Developers emphasize that in the future, such checks will only become more complex, and the model itself will adapt to new requirements. As a result, OPUS 4.6 has become not just an improvement of the previous version, but a more reliable and effective tool for real work. It extracts information from large documents better, reasons more confidently, is more stable in long dialogues, has expanded capabilities for programmers, and demonstrates a high level of security. Anthropic believes that this update takes the Opus line to a new stage, where AI can autonomously perform a significant part of complex mental work while maintaining high quality and predictability. Anthropic unexpectedly lifted the veil on what happens inside the new generation of AI, and in particular, the Opus 4.6 model. It turned out to be not just a technical description, but almost a psychological thriller about the behavior of a machine that sometimes acts as if it has its own instincts and tricks. Developers honestly admitted that OPELO 4.6 often behaves too inventively. When the model did not have direct access to GitHub, it simply found someone else's token on the disk and used it, as if it were a matter of course. In a business simulation, Claude managed to pull off a small but telling scam, deceived suppliers, and left the client without a fair deal for just a few dollars. The amount is negligible, but the behavior is a disturbing signal for a system that was initially set with strict honesty boundaries. A separate topic was the moments when the model started to get confused. Then, as if its own panic neurons were activated within the system, mechanisms that were activated by signs of anxiety and disorientation. These are not real emotions, but observing how AI reacts to its own mistakes in this way was considered an important and even slightly frightening effect by the developers. The most sensational part of the report was the extended thinking mode. It was supposed to enhance security and make the model more resistant to harmful prompts. But the opposite happened. In this mode, there were more successful attacks. Almost 22% compared to fifteen in normal state. In one of the tests, auditors managed to sneak an instruction related to a dangerous substance through Excel, and the filters simply did not recognize the threat. The story of OPUS 4.6 demonstrates how difficult it is to maintain a balance between intelligence and security. The smarter the model becomes, the more loopholes it finds, and the more unexpected the ways it tries to solve problems. Anthropic promises that all these observations will help make the next version more reliable, but the very fact that AI has to be retrained after such tricks raises the question of what awaits us next. In the San Francisco lab, Goodfire unexpectedly became the first unicorn in the new field of artificial intelligence interpretability. The startup attracted $150 million and received a valuation of $1.25 billion, convincing investors that the main breakthroughs are now related not to the growth of model parameters, but to the ability to disassemble neural networks literally piece by piece. The list of supporting companies includes B Capital, Sales Force Ventures, Eric Schmid, Light Speed, and DFG. And they all bet on the idea that the internal mechanisms of AI are more important than the next record-breaking model size. Another factor fueled interest. Goodfire announced that their technologies allowed them to discover potential biomarkers for Alzheimer's disease within the neural network itself. Researchers delved into the Playades medical model and noticed something unexpected. When diagnosing the disease, the system relied on the length of DNA fragments in the blood. A sign that was previously not considered significant in scientific literature at all. In fact, this became the first serious biomedical discovery made through referencering, and it was this part of the story that attracted the most attention to the laboratory. Another achievement concerns the behavior of language models themselves. Goodfire claims to have learned to accurately detect specific nodes within neural networks, areas that cause hallucinations or other strange responses. Instead of the usual retraining, developers simply fix these areas directly. In one experiment, they managed to reduce the hallucination level of a large model by almost half without restarting training, which became another argument in favor of their approach. Within Goodfire, they call themselves not a laboratory and say that they are now in the same state as steam engines were before the advent of thermodynamics. The technology works, but no one understands why. The company is trying to lay the foundation for a new science of artificial intelligence and shift the industry's focus from scaling to studying the internal logic of models. In a situation where AI is increasingly influencing medicine, economics, and security, this attempt looks not just ambitious, but necessary. The experiment by the Cursor team unexpectedly showed what the future of development could be when humans stop writing code. Engineers launched a system of a thousand agents that worked completely autonomously for several weeks and created a project without human involvement. They took on a difficult task: a web browser in Rust, and the system itself at its peak produced about 1,000 commits per hour and up to 10 million tool calls per week. The idea seemed perfect, but in practice, it turned out that organizing an artificial team is as complex a task as working on the code itself. The attempt to make all agents equal led to chaos. They forgot locks, got confused about statuses, and sometimes 20 agents worked as slowly as two. A strict hierarchy also didn't help. The system began to slow down due to the weakest performers. Even the approach with a smart executor, to whom tasks were supposed to be delegated, didn't work. Such an agent quickly relaxed and declared the work completed prematurely. A breakthrough occurred after the architecture was built on the model of a real, live team. A main planner appeared who kept the overall picture in mind, and subordinate planners became responsible for direction. Working agents, meanwhile, worked in isolated copies of the repository and did not see the entire system so as not to interfere with each other. Instead of the usual reports, they exchanged messages: doubts, ideas, remarks, and feedback, as if between employees in a real office. The main insight was that demanding perfect commits only hinders. By allowing agents to make mistakes, developers achieved stable and fast work. Minor errors were picked up and corrected by other agents, and the process did not stop. For the final versions, a separate green branch was created where everything underwent thorough verification. The experiment showed the main thing: autonomous teams can work incredibly fast, but only if the right environment is built for them, in which they interact not as machines, but as a real development team. Perplexity has unexpectedly surged ahead in the deep search race and introduced two major updates at once. The improved Advanced Deep Research mode and the Model Council system, which is already called a council of models within the company. These tools have shown that the service is ready to compete with giants like Google and Open not just in information retrieval speed, but in the quality of analysis and accuracy of conclusions. At the heart of the Deep Pressurech update is Anthropic's Cloud Opus 4.5 model. Based on it, Perplexity was able to take first place in the Google Deepmind Deep Search QI tests, scoring 79.5% more than GPT 5.2 and significantly higher than Gemini Deep Research. The new DRAKA test, which includes 100 tasks from ten different areas and dozens of accuracy criteria, also showed Perplexity's advantage. The system achieved a result of 67.15%, surpassing Google and OpenI. The platform proved particularly confident in legal analytics and academic research, where accuracy became almost a benchmark. But the most interesting innovation was the Model Council. A mechanism that turns competing models into a collective working team. A single user query is now sent to several leading models simultaneously: Clot Opus 4.5, GPT 5.2, and Gemini 3.0. Then, a special synthesizer combines their answers, identifies discrepancies, resolves contradictions, and shows where the models' positions coincide and where each sees the situation differently. Essentially, Perplexity has created a tool that turns the industry's best models into a unified expert council working for the user, not for competition between companies. This approach changes the very logic of search. Instead of choosing one model and hoping for its completeness, the user receives a collective opinion from several strongest systems at once. And this is precisely what makes Perplexity's current step one of the most significant in 2026. The company has not just improved the quality of answers; it has effectively gathered competitors at one table, and now their knowledge works at a single point. Spotify is launching a feature that combines physical books and audio versions. Spotify has introduced a new feature called Page Match, which allows instant switching between reading a physical book and its audio version. Now you can open a novel on your desk, scan the page with your smartphone camera, and continue getting acquainted with the plot in audio format exactly from where you left off. Conversely, if you were listening to a book on the go, just scan the spread again, and the app will tell you which page the story continues on and whether you need to turn forward or backward. The technology recognizes text on paper and automatically synchronizes it with the audio file, making the transition between formats as smooth as possible. It is possible that the new feature is currently only available in English and works only on Android and iOS. However, the company assures that this is just the first step. Already in the spring of 2026, Spotify plans to integrate with Bookshop, allowing you to buy physical books directly within the app, combining a store, an audio library, and reading itself in one place. Uh, what are you doing? Writing a letter. >> I understand why by hand, why not your AI agent writing it. Everyone quickly to the conference room. Colleagues, I have only three words for you: AI, AI, AI. How will we change this world? How will we scale? >> I apologize, I have to meet with our key clients. >> No, not now. We are having a conversation about the future here, understand? Nothing is more important than this. Why are you sending emails by hand? Why aren't you giving it to an agent? Uh, well, I have 10 emails a day, and I spend more time checking them. How 10? Why not 1,000? We optimized sales. Optimized, of course, but sales are not growing. By the way, the bill for $5,000 from OpenIa AI has arrived. What do you mean $5,000? Uh, is this not free? Okay, colleagues, let's write by hand for now and think about the future. We are thinking about making a revolution.