Transcription
Alright, friends, well, hello everyone. My name is Danil Simonov. Today, we will talk with you, and sum up the year a bit in agency development. We'll see what has changed over the year, what has appeared, from what point we have come, and discuss, perhaps, the most important essence of these changes, namely how we have moved from simple agents to ecosystems around large vendors. This will probably overlap a bit with Refat's report, but mine will be more about development. We will discuss context management a bit and how large players do it, and we will discuss a bit about how to measure the effectiveness of how development, development teams change under the influence of AI agents.
Well, I have been developing daily for about 17 years in total, and for the last 7 years, mainly in strategic roles. And for the last year and a half, I have been using AI agents for development almost daily, because I absolutely love it, it speeds things up incredibly. And I have already outlined the plan for you, what we will discuss. So let's just move on to it.
So, I decided to start with a quote from Dario Amodei of Anthropic, which went viral at the beginning of this year, where he said that we, as developers, have literally 6-12 months left until 90 to 100% of code will be written by AI. And some sources say that this is possibly true, but not from the perspective that everything that people used to write is now written by AI, but from the perspective that AI generates so much code that 90% of it is possibly indeed attributed to artificial intelligence. Although, in my intuitive feeling, it's still a bit less.
Let's quickly go through it, let's reminisce, remember what was, and reflect on this topic. I just want to remind you that the beginning of the year is when AI agents appeared in principle. That is, 2024 was more about chatbots and interactions with some conversational stories. Full-fledged autonomous agents that work in your codebase and do something, accomplish something, this is a story that only appeared at the very beginning of this year and began to fully develop into some production solutions. Therefore, if we briefly go through how it was, then January was the appearance of the Composer Mode in Cursor and in Warp. February was the appearance of Codeium, which marked the era of CLI agents, agents in the terminal. OpenAI quickly joined this, releasing its coding model Codex and, accordingly, also the console solution Codex CLI. Google joined in, releasing Gemini CLI for its just-released Gemini 2.5 and Gemini 2.5 Flash. After which July became a very serious dividing line, when the rules of the game and development began to change significantly under the influence of Cursor changing its pricing, it stopped charging per request and started charging per token, which led to the cost of working with AI agents increasing by five to 20 times for many people. Whoever was less lucky, so to speak. And this was the moment when everyone started urgently migrating to various C-solutions. I think this was the main moment when Codeium surged ahead. And at the same time, Codeium began to incorporate interesting new features that simply didn't exist anywhere else before. One of these features, for example, is sub-agents. And around here, the story begins about how tools are trying to roll out some unique functions that only they have. But so far, autumn, early autumn, is remembered primarily for the model race, when a new flagship model was released literally every week, when Grok entered the game, when all flagships competed for who would spend how much time in the top and who would lure how many audiences to themselves. And closer to the end of autumn, those very unique features began to appear in one place or another, which I was talking about. That is, Cursor 2.0 was released, which was marked by a very good integrated browser. And the agent works well with this browser. They released their own model, trying to enter this market from the model side as well. They started to roll out the features they missed while Codeium was outrunning them, so to speak, in functionality. And I placed Agent Skills on this slide, but in reality, the standard Skills only happened in December. That is, yes, Agent Skills are mentioned on this slide, but they only recently became an open-source standard for everyone. Nevertheless, November was the race for the frontend. Google AntiGravity appeared, which rolled out a lot of things related to frontend development, when the model can see the result, when the model generates the design first, when the model generates illustrations for your landing pages and so on. Google, in general, maximally exploited the banana. And at the same time, Cursor rolled out design editing directly in the browser, which the agent then applies to your codebase. And from here, the feeling begins to form that all players have become very strongly concentrated on the frontend. Previously, everyone just tried to generate code and ignored, in general, the visual capabilities of models, but for some reason, by the end of the year, everyone realized that this is apparently the golden nugget that needs to be grabbed.
And here I want to transition a bit to this reflection. What specifically has changed for us? If we look at it from a high level, what happened over this year, then, I think, three very clear stages can be identified. At the beginning of the year, all flagship companies, OpenAI, Anthropic, Google, competed with models. We had literally one or two key places where we did development, primarily it was Cursor. And the competition between flagships was that you simply selected one option or another in the Cursor model selector. And in principle, this is how the main loud news headlines looked, that here it says: "This model is out, let's switch, everyone, try it urgently, it's all great." Around the summer, the game turned around. The change in Cursor's pricing led to everyone competing for pricing. User migration between different systems became based mainly on limits and prices. People started fleeing Cursor to Codeium because Anthropic rolled out very good, cheap pricing. Then Anthropic urgently started tightening the screws, and Google released very generous pricing in Gemini CLI. Everyone started running there. And as a result, the main user migration, the main features that, in fact, flagships competed on, were simply banal prices. And still in mid-summer, somewhere around, I think, the release of a flagship model by quality was often a significant leap from the previous one. The release of each new flagship model marked a new step in the quality of how the model works. This is the edge that, by the end of the year, I think, has been blurred the most. That is, speaking now about December, the difference in quality of work between flagship models is so imperceptible that, well, even if you open an analysis and try to find a difference between, say, Gemini 3 Pro and OpenAI GPT 5.2, you won't be able to find it. And here, competition has moved into the realm of features. An ecosystem has begun to form around each of the flagships, where you are offered not only a range of LLMs that can perform work, but also unique features in each agent environment that exists. If it's code, then it's context engineering, it's skills, it's the ability to run sub-agents with isolated context, advanced tool usage, which we will also talk about a bit now, and all sorts of such things. If it's Cursor, then it's the UI, it's the ease of use in a graphical interface. Google has made a big bet on its graphical models, on Nanobana and the ability to understand visual information, screenshots from the browser, and so on. A very good ecosystem is forming around each of these tools. And what are the main differentiators here? I'll go through them quickly. It's the ability to plan and how work with plans, specs, and all that is organized. It's the ability for built-in debugging. The recently appeared debug mode in Cursor. Modes in Codeium, when people configured skills to debug errors step by step, and so on. It's general tooling, the ability to integrate into SDKs, GitHub pipelines, and CI/CD, and so on. It's memory, how work with summarizing facts, saving some information, and so on. And I wanted to highlight sub-agents as a separate point, because one way or another, everyone is playing with this. They are also trying to work with context isolation for different roles, more or less everyone. So, this is also an interesting point of differentiation, where Claude is, of course, the absolute leader. And we have already generally talked about ecosystems. I have already said what distinguishes each of them, but here I just want to repeat that the main players are Anthropic, Google, OpenAI, and Cursor. And Cursor is interesting in this regard because they are the only ones who entered this race from the side of the agent environment, not from the side of models. Cursor's own model appeared only recently. It's a cool solution because it's head and shoulders faster than all other models. If I'm not mistaken, they use Cerebras under the hood. Nevertheless, they entered this race from the model side as one of the last, and initially they were precisely an agent system.
Here I want to talk a bit about what I mentioned on the previous slides. These very skills, context isolation, and so on, because one of the key trends of this year was so-called context engineering. Why did this thing appear at all? For example, one of the studies by Nalim shows that simply stuffing a lot of data into the context is not a solution to all problems. Extra data, noise, and all sorts of low-quality information only confuse agents and actually hinder them from making decisions. Agents are very bad at prioritizing information and understanding what is important and what should not be paid attention to. Therefore, one way or another, throughout the year, all players tried to figure out how to ensure that only useful and necessary information gets into the agent's context window, and, accordingly, that unnecessary information does not get in. And this is precisely what we observed in general. And how everyone came up with ways to package context and what context engineering is in general, and what it consists of. I really liked this screenshot from one YouTube video. I brazenly decided to insert it into my presentation, the source is in the bottom left corner. This is about how we write context, how we save facts, extract memory, and other information. This is about how we extract important information from all the context. These are the same mechanisms of Anthropic, tool search, these are various mechanisms, well, the same skills, which, based on the skill annotation, can understand that they need to be loaded into the context, and so on. This is how context summarization and compression happens. And the last is how isolation happens. This is precisely the story about sub-agents, about agent delegation, and so on, which we will now discuss. And as a product, context engineering consists of, in principle, very understandable layers. On the first layer, you work with a certain global level, general rules for your system. Perhaps some security rules, general style rules, and so on. The second layer is the documentation of a specific project, a description of what we are currently doing within this system, what it represents, perhaps from a business perspective and all that. Then a plan, some global large story that we are working on right now within this project, it is broken down into the specific role of a specific agent that is doing it. In a situation where you have several sub-agents working on this task, each within a certain general plan may have its own role. One writes code, another reviews, a third tests the frontend, a fourth does something else. Within each sub-agent, there is a breakdown into tasks. These are, accordingly, to-do items. You have probably often seen in the interface how agents write a to-do list for themselves and then check them off one by one, performing these small, decomposed steps. And the last is the direct active context of what is happening here and now. The result of the latest tool calls, read files, something loaded, those same skills, and so on, and so on. And around each of these layers, a whole set of solutions has been built, what the market essentially offers us.
Here, before we move on to the next part, I want to separately mention what RAG is. Perhaps many have heard of it, but just in case. Many speakers today have talked about RAG, about how to load information into context on demand, how it's a somewhat similar story. It has several basic implementations. This is the idea that the context window has grown so much that we can, in principle, place the entire knowledge base directly into the agent's context window. But over time, it became clear that this is actually quite an expensive pleasure. Anthropic itself faced the problem that MCP tools, for example, heavily clog the context. And it evolved somewhat towards the idea that we can announce to the agent what pieces of information we have, and allow it to independently request this information if it deems it necessary. That is, it's a kind of retrieval, but based on the agent's own judgment, when it decides that it should obtain this information.
And let's discuss a bit about how documentation works. Well, in principle, I won't dwell on this for long, because everyone is probably more or less familiar with these stories. This is standard MD, Cloud MD, Cursor Rules, whatever they are called. It's some description, some text document in the root or a set of interconnected text documents. Usually, it contains general style rules for your project, where everything is located, and, let's say, the information that you would not want the agent to study from scratch every time. If you don't want the agent to have to go through the entire project anew in each new session and spend time and tokens to learn how everything is structured, you can pre-summarize this data and simply place it in a pre-fixed context. This greatly speeds up the exploration phase when the agent is simply getting acquainted with the project. If you give it this information out of the box, it, accordingly, doesn't need to get acquainted.
Here I want to mention something that was born this year. The spec-driven approach, when work on some projects turns not into work with the codebase, excuse me, but into designing how your project will look. Here, of course, one of the biggest resonant moments is the appearance of GitHub's Spec, when you literally interact with the agent to design how the project should look little by little, and it turns into a detailed decomposed pipeline of tasks, which can then be fed to coding agents that will implement all of this, fix it, and so on. And Memory Bank also famously implemented a similar concept, where you go from a high-level business task, break it down into subtasks, specific implementations, and so on, and so on. Spec-driven, essentially, tries to be an answer. If we return to this slide, to almost all things, starting with project documentation and specs, that is, global rules remain general, everything else from below can be determined by your spec framework, which helps you design each of these pieces.
And now I want to return a bit to the story about sub-agents. We have mentioned them several times. What are they? This is a feature that Codeium invented. They thought: "Well, suppose we are working on some tasks. For example, we need to fix some bug. Fixing a bug, for example, requires investigative activity. Inserting some logs into the code, seeing what certain variables are equal to at each launch. From this, drawing some conclusion about where the error is, writing the fix code, running it again, making sure the error is fixed. In short, a set of actions. And for the global pipeline of implementing your plan, implementing a bug fix in one item of this plan is, in principle, one small piece that is not very important, but all the context related to restarting tests, seeing test results, changing code, and so on, falls into the same context window. And this turns into each plan item growing into huge chunks of 30-40,000 tokens in the best case, when the agent implements it. Codeium, well, Anthropic had the idea: "What if we feed another agent the general status of current information about what's happening, the specific task it needs to do? And at the output, after it has performed all the iterations itself, this agent will return just a resulting piece of context, saying: 'The error is fixed, it was here and here, and in the future, you need to consider this and this.'" And this worked very coolly. One of the main problems of clogging the context window with all sorts of junk, irrelevant information, and wasting tokens, money, and time on it was very elegantly solved by the fact that subtasks within your plans can simply be passed to another agent. This is a very cool thing. It was subsequently implemented by more or less everyone. One of the strongest implementations I've seen so far was in Cursor. For some reason, they have now hidden it, put it back. Although in the beta, they had sub-agents that worked really well. But now we don't see it. Most likely, I think, there were some bugs due to which they decided to remove it, but I have almost no doubt that we will see the implementation of sub-agents in almost all solutions.
Well, what are skills? They have been talked about a lot today. I want to very briefly mention what they are and what they represent. A skill is a packaged piece of context, some documentation, some text plus a set of tools that are somehow relevant to this text. For example, a simple example. I have a Raspberry Pi at home. I really like to sometimes have an agent perform some settings on my Raspberry Pi. Deploying the agent directly on the Raspberry Pi on a Zero version is quite expensive and inconvenient. Therefore, I do it from my computer. And it's very convenient, for example, to package this into a skill for Codeium. I have a small piece of documentation that provides information about what board I have, what system version it is, what nuances and settings there are. And several bash scripts that allow me to execute some SSH command on this board, reboot it, do something else. And by packaging all this into a single skill, when I work with Codeium on some tasks, everything is okay. But if I start giving it tasks related to my Raspberry Pi, it says: "Oh, I see, I have announced information that I have this piece of data. It's not loaded now, but if I'm going to work with this, I should load it." And it turns out that initially, when you work with a session of Codeium or another agent that implements the skill mechanism, it doesn't load all the information about the skills that are available to it in principle. It only sees announcements, information about what skills exist, what they are for, what problems they solve. And if, in the course of work, it realizes that this information will be useful to it, it can load it into the context and then use the tools within this skill, this documentation, and so on, and so on.
Well, the last thing in context engineering that I want to mention briefly is Tool Calling. This is another interesting innovation from Anthropic, again related to MCP. They noticed that when you connect the output of one MCP and feed it to another MCP or call several MCPs in a chain, it leads to an unjustified inflation of context. That is, when one MCP gives an answer, whose only task is to be input for the next, this answer is, in principle, not needed in the LLM's context. And again, tokens, time, and money are spent on this. And they thought: "What if we create an isolated environment? where the agent will write a kind of pseudocode, which calls these MCPs as if they were ordinary functions in a programming language, and it can connect them to each other, nest them, write some ifs, loops, and anything else. That is, to create a kind of branching logic around calling these tools. And all this will happen in an isolated environment. And also, as with sub-agents, the task is received at the input, and the result is received at the output. Everything that happens in between is not a problem. This is a very interesting implementation. It is currently implemented in Codeium 2.0, and according to Anthropic's own statements, it significantly reduces context costs for tool calling. We'll see how it catches on in the future. My hypothesis is half-hearted. On the one hand, it's an interesting implementation. I generally like this concept. On the other hand, I rarely encounter situations where MCPs really need to be connected in such a chain. Therefore, how much it actually saves context, time will tell.
And speaking of all this evolution of models, ecosystems, the evolution of how we approach context design, one of the most important questions that arises in the business mind is: "What about metrics? You've come up with so many things this year, how useful is it for business? Does it affect it, does it not, how do we measure it?" We've decided to implement all this in our company, what do we do with all this? And here's one of the reasons, I think, why discussions about metrics became so heated and relevant from mid-year, is the Metri study, which has already been criticized a lot, I won't do that now, but nevertheless, it voiced the thesis that developers who work with AI agents are much more enthusiastic about working with AI agents than they actually benefit from them. And this generated huge doubts about whether these enchanted, romanticized developers understand what they are doing and how useful it really is. And from here, the discussion about metrics was launched.
Well, what metrics should you measure, globally assessing how development is changing as a result of implementing agents? In my opinion, a rather predictable set, but nevertheless, I will voice it. It's speed. If we only take the developer cycle, then time to merge or time to QA, that is, how quickly a task leaves the developer. And, accordingly, time to market, that is, the total time from when a task appeared until it went into production. Why is it important to measure both? I will talk about this in more detail later, but nevertheless, the fact that your developers have started generating pull requests faster does not guarantee that these pull requests are of high quality. Perhaps QA is overwhelmed and Time to Market has increased. Although time to merge could have tripled. And here is throughput, that is, if not just the time of one task, but how many tasks a developer can generate in general, because one of the problems, parallelism, conflicts, and so on, can reduce throughput while also reducing development time. It's quality, which I already mentioned, bugs, returning tasks from QA back to development, how much time is spent on review, and so on. How many rejections, pull requests have there been relative to what was before. An interesting point that is rarely discussed is predictability. One of the important business metrics related to a development team is to at least roughly estimate how much time and money it takes to create something, how you estimate features by their cost to your business. And AI agents greatly influence a developer's ability to predict how long a task will take them. Naturally, provided that they will do it with AI agents. This is such a new thing that the intuition of how an agent works in one case or another, for many, if it has developed even a little, it is still in a very nascent state. And therefore, monitoring predictability and how estimates change depending on whether you use agents or not, is important. And satisfaction, the change in the development process when you work with AI agents and when you write code, is very significant. It's a different behavior altogether. These are different things. In one case, you literally write code, you are an engineer, in the other, you are more of a manager, and you manage the agent. Not everyone likes to manage. Many engineers like to be engineers. Therefore, monitoring that your team is not burning out, so to speak, because you are forcing them to do activities that are uncomfortable for them, is actually important.
Well, let's talk a bit about the trap of perceived benefit. This is generally one of the main factors of AI, and it leaves no one indifferent, and everyone either hates it or falls in love with it. And therefore, when your developers fall in love with it, you need to be careful, because often, because you yourself spent much less time writing code, it seems to you that the task took less time in general. But in reality, perhaps, debugging, prompting, monitoring how the agent works, took three times longer, but psychologically, you feel that you have accelerated incredibly and you were not busy at all all day. Therefore, it is important to monitor this, and it is important to ground it in reality.
metrics. And in this context, the latest report from Anthropic is very interesting, which they wrote about how Claude is changing the atmosphere, let's say, and the work within their own team. There are several, in my opinion, cool theses. Firstly, people's self-assessment of productivity there has significantly increased by two to three times compared to the previous assessment at the beginning of the year. Secondly, they are increasingly using, and thus the share of artificial intelligence usage in working time is currently estimated at up to 60%. Interesting points from what they noticed are changes in social dynamics. For example, people touch each other much less on certain issues. It's easier for them to ask a question about how something works directly to the Claude code within the codebase, rather than going and asking a New York programmer. I was surprised that within their report, they positioned this as a negative point. I think many can argue with this, but nevertheless, such a change is happening. Another interesting thing is that one of the most important things that affects the quality of work with AI agents is the development of internal intuition. Will the agent cope with this task or not. And this is really a very important thing, because until you develop this intuition and try to give the agent everything, most likely, you will not only not see an increase in productivity, but on the contrary, you will feel a slowdown. You will repeatedly engage in this babysitting and situations where the agent is stuck, and you are trying to make it do something it cannot. When this intuition is developed, you find yourself in an ideal world where all the tasks on which the agent will not work, you simply do not give them, and they take up the same amount of time as they did before. And for those tasks where you know it will work, you offload them to it and get a tenfold acceleration on this spectrum of tasks, so to speak. And the composite acceleration you get is indeed significant, 30-40% and so on, depending on your specific codebase and your stack.
Well, an important thing that I want to talk about from the problems I've seen with my own eyes, the initial euphoria from implementing AI agents and the acceleration of code development time for some features, an increase in the number of pull requests, an increase in the volume of code that developers deliver, creates the feeling that wow, we have four times more pull requests. Incredible, we've accelerated four times. It is very important, very, very, very important to measure CHOSQA, because, well, I've seen many of these pains when developers start making pull requests without looking, thinking: "Anyway, someone will review it, if there's an error," they'll notice. Developers again start glancing at what the agent has identified for them, where the errors are, because they think: "Well, QA will check and so on." In fact, the developer doesn't encounter the problems of babysitting the agent because they implicitly delegate them to another department that comes next in the pipeline. And this is a problem, naturally, it needs to be measured. Naturally, the fact that development has simply started generating a lot of code tells you absolutely nothing. Don't forget to measure what's happening in QA, what's happening with rejections, what's happening with time to market. This is a very important piece. I truly see that this is sometimes forgotten, and they report how cool it is that we've implemented it and our development has started making features. And then they see that after 2 months, QA is overwhelmed and continuously rejects everything back to development.
Another very interesting symptom is the growth of "nice to have" things. Imagine there was some kind of thing, I don't know, you decided to visualize some data so that the whole company could easily look at it, some kind of beautiful graph. Before, you wouldn't even bother with it, because it would take, I don't know, 4 hours, and it's some kind of "nice to have" thing that's not mandatory. Now, with an agent, you can give it this task, and it will take 15 minutes. That is, the task of visualizing ready-made data to create a nice graph is an elementary task. And the agent will cope. It will definitely cope, and it will take you. Well, while you're setting up the task for a couple of iterations, well, 15 minutes, right? It seems like, wow, how cool, we've started doing a lot of additional pleasant, cool things. Anthropic also mentions this in their report, how many internal tools they now have, because it's now cheap to make them. What's the focus? Cheap is not free. If before these things took zero time, because you didn't even look at them, because they could take time, then now even 15 minutes, multiplied by 10 "nice to have" things, is actually a normal amount of time. If you spend a portion of your time every day on framing, visualizations, dashboards, and other such things, which, yes, are amazing, they are indeed 20 times faster than before, but before, in fact, you had zero, because you didn't do it at all. This is an important thing. Pay attention to it. It's very cool to make small, convenient little things that make you happy, but it's important to ensure that they don't eat up 50% of your working time.
And one more thing, close to the end, is one of the cases where AI works almost perfectly. A case that I highly recommend to everyone and advise you to spend time on. And agents can excellently improve code in terms of writing tests, automated testing. Documenting code, commenting, and so on, and so on. This is a case where it works very well in 95% of cases. Especially if you detail the test hierarchy, how they should be executed, then this is an ideal task for AI agents. This is a metric that is growing, probably for everyone. I have practically not encountered scenarios where a company started writing autotests with AI agents, and they didn't like it, and it was bad or somehow worsened their situation. This is a scenario where it works very well. And additionally, this is important because autotests are a catalyst for the quality of the AI agent's work itself. If the AI agent cannot automatically check what it has written, it will have to be checked by you more often, you will have to find bugs more often. If there are autotests written by the agent itself, then during the implementation of a feature, it tests itself and achieves an excellent result.
Well, the last thing I want to say is, speaking of metrics, people often ask: "Ah, ah, give me an example." That is, abstractly, we've talked, development has become faster, blah-blah-blah. But faster is what? Is it three times, is it one and a half times, is it 5%? What do we call faster? Here's the average result. I took it from another presentation of mine at another more business-oriented conference, where I literally gave examples of cases of how it works for different companies. And one of the examples is precisely for a large development team. What metrics did they have, what can you orient yourself by with some margin of error. So, time to QA for bug fixes and small features decreased by 23%, almost a quarter. At the same time, time to market by 15%. What does this tell us? Firstly, it's great, a 15% acceleration is actually very good. But at the same time, the difference between time to QA and time to market indicates that the load on QA has increased. That is, developers are delivering more features than ultimately reach production. And the slowdown of QA. Well, one can argue why it happens. Maybe it's just due to the increase in the amount of work that has fallen on them, or something else, but nevertheless. And this is precisely a good argument in favor of what I said earlier. Measure not only what the developers did, but also what specifically reached production. For average large features, the acceleration to market was very insignificant. At the level of error, it's hard to talk about statistical significance. This is mainly due to the fact that developers have developed this intuition. They understood that on their codebase, specifically in their tasks, it handles a certain spectrum of small tasks very well, but poorly handles large epics, so they simply didn't give them to it. Therefore, evaluating, let's say, a zero result is, in general, logical. If you haven't changed anything, then nothing has changed. And unit test coverage has increased dramatically. Well, that is, at the current moment, this is a slightly outdated presentation, I mean this specific screenshot. At the current moment, I was told that they have already reached almost 95%, where almost all these tests are written by Claude code. Integration test coverage has also increased very, very significantly. This is again an argument in favor of the thesis that autotests are simply a gold mine for AI agents. Unleash them and enjoy. The number of bugs at these QA stages has decreased. This is interesting from the perspective that, well, the quality of what developers deliver has increased. To say that this is directly the merit of AI agents, that agents write such amazing quality code, I probably couldn't. When we discussed how they thought this happened, we concluded that simply when your unit test coverage grows from approximately one-third to 95%, it's likely that a colossal number of bugs that simply reached QA now don't reach them due to, well, simply, code coverage, higher quality coverage.
Well, at this point, I want to thank you and, apparently, look at the chat, see what people are asking there. So. Thank you very much, Daniil, for your great presentation. >> Maybe there are some highlights, Max? Maybe you noticed something in the chat? >> I want to highlight what you said, because it resonates with me very strongly. I am also actively involved in implementing AI in development and in all technological teams. And it seems that not everyone understands that it's not enough to just code fast. It's much more important to deliver features to production quickly and deliver them with quality, so that there isn't a large number of rollbacks, a large number of bugs, blockers after the release, and so on. And I really liked this part too. I've written about it in my channel and I emphasize it very strongly at my work. From the questions from the chat, I liked the following. Can one agent, for example, feed code rules from another agent, Cursor or Kilocode? And does the efficiency change in this case? >> Specifically, Kilocode is one of the worst in terms of pulling information from other agents. Because they dominated the market at one time, and now they are at the top, and rather others pull rules to Kilocode. So I would probably advise you to write clot.md. Ah, for example, Cursor can work with them. In other agents, you can very easily configure through command-line parameters which specific file they will read, and so on. As far as I remember, I'm afraid to be wrong now, I might really be wrong. I think they still don't even support agents.md. So I've seen people implement pulling MD into Kilocode as a separate skill, although it's considered a universal standard for everyone. So, answering about Kilocode, it's difficult, it's not that easy to achieve. For other agents, everything is much better. They've all become friends, and Cursor can pull custom Kilocode commands. It can pull CL MD, AG MD, anything you want, and so on, and so on. Cool. People are writing in the chat: "Thank you, very interesting. I only came for you, I don't regret it at all." >> Thank you. Thank you. >> And in general, a lot of positive words about you and about our other speakers. >> Listen, can I briefly comment, if we have seconds, I don't know, 30 minutes, >> yes? We definitely have about 30 seconds left. >> I just saw that there was a micro-discussion in the chat between users on the topic that AI writes rather low-quality tests. I want to, well, since I dedicated a huge part of my presentation to how high-quality tests it writes, I want to briefly respond and comment. Yes, if you directly ask it to generate autotests for the codebase, it really, well, it can, let's say, it doesn't perfectly understand the influence in the code. Because of this, the tests turn out to be either redundant, when it covers the same branch 10 times, or vice versa, insufficient, when it misses important branches, important use cases. This is well solved by the stage of automatic planning of these tests, when it analyzes the branching in the codebase. And this is solved by good documentation on how you write tests. So I have a large block of documentation for this. It's actually three pages, four, where I describe in detail the hierarchy, that you should have this nesting, this nesting for classes, this for methods, this for this piece, and so on. Yes, these rails took me a normal amount of time at one point. I think I iterated and tested for two days, well, really no less, just two days of experimentation with the agent. But now it's complete autonomy, meaning I don't have to check these tests at all. They just, if they work and they are green, I know they are written well. >> Damn, cool, cool. Yes, I agree. Actually, autotests are where, I think, you can apply it first, because the risks are much lower than with production code, and there's quite high repeatability, a lot of boilerplate code, just like, for example, in frontend, but in backend it's already much more complicated. Well, that's my opinion, at least. >> I would really like, >> excuse me for rushing you, but I would really like to talk about trends for another 2 minutes or a minute, before we move on to the next speaker. What do you see, >> how do you see 2026 from a development perspective? >> Hmm, good question, how I see it from a development perspective. Let me think. Ah, well, I think there will be a strong trend of entering fields that haven't been entered yet. That is, basic agents write code well, but they have a big problem with feedback. And what's happening now, for example, with frontend, is essentially tooling. That is, the models themselves are changing little. Ah, the models are not being refined for frontend. Ah, models are given as many tools as possible for frontend tasks. Built-in browsers, editing capabilities, and so on. I think we'll see something similar in other segments. That is, mobile applications are currently very poorly covered. There are startups that are doing this, but I think major flagships will definitely get into this. Our desktop applications are very poorly covered. They have fewer problems with them. There are many desktop applications now. These are essentially web applications. But nevertheless, I think there will be a slight shift in this direction as well. Globally, I think there will be something interesting in terms of increasing context engineering. I think we'll see more solutions like advanced tool usage, various context isolations, context extractions, autonomous environments, but fundamental changes, honestly, it's hard for me to predict that there will be any fundamental change. It seems now that all the main pains that I encounter in agent development are precisely problems with tooling. I see that agents sometimes lack the ability to use existing tools. Debugging too. If you've seen how debugging is set up in AI agents, in Cursor, for example, it's a beautiful and elegant solution, but why not give it the ability to set breakpoints and simply literally see how my code works. That is, objectively, this problem has been solved for a thousand years. Cursor can do this. I mean Cursor as a code editor, but Cursor as an agent simply doesn't give its agent such an opportunity. We haven't figured out how to connect it yet. I think there will be more and more of these things.