📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Как разработчику выжать максимум из LLM? Claude Code, MCP, Агенты: полный арсенал для разработки

Go Get Podcast2:23:58

Transcription

Well, hello everyone. We have the nineteenth episode of the podcast. Soon it will be a юбилей, the twentieth. Today, for the third time, we are talking about LLMs, about AI tools for developers, primarily. And in the past, in past episodes, we talked more about the theoretical part, about some broad questions. This time, we will delve deeper and focus more on practical application. This is exactly what we were missing. That is, we talked for quite a long time, for several hours, and then people said: "But where, in general, are the instructions, the guides, how to use these things of yours, these neural networks." And this time, we will talk about that. For this, I invited Sava. How we met: we periodically have Gofer meetings in Astana, and there you somehow mentioned that you use very advanced techniques when working with neural networks. I always thought that usually in my circle, I was the main enthusiast, and I usually told everyone, and people were like: "Oh, wow, you have some skills there, you have this Claude MD, global, local, wow." And then I listened to you and realized that it was like a schoolboy just outshone a schoolboy. I realized that I had to invite you to the podcast and share your experience [laughter] with people. Okay, now we will introduce you. But I briefly wanted to mention the sponsor. The sponsor of this podcast is still Avito. I didn't misspeak this time. The guys have a very cool approach to sponsorship, so to speak. They just came to me, said: "You have a great podcast, it comes out rarely. We want you to release it more often, so that the episodes are regular. Therefore, here's money for you. Do as you like. We won't interfere with the content." And it's all up to the creator. The only thing they asked was to briefly tell about them. The guys have a very cool, large infrastructure, 370 development teams, 242 million active listings. In general, a large headcount, 3,000 services, and they will be happy to see you working there. And they also have their own YouTube channel, I will include a link in the description. They talk about Go and not only about Go there. I watched an episode there, for example, about a review of Go 1.25. Also an episode with a detailed review of version 1.25 for several hours, about three hours, probably, we discussed it. If you need a more overview, fun, interesting one, then I think the guys from Avito Tech have one of the best on YouTube. In general, thank you very much to them for this, and we are starting. Ah, Sava, will you tell us a little about yourself. >> Yes. Hello everyone, I'm Sava. I've been interested in technologies since school. There I got acquainted with the Delphi language superficially. Then my life turned out in such a way that I wasn't involved in development. Then I basically learned Python, Java, and at some point realized that I want to work and develop in Go. So, I started doing this a few years ago. And the period of my recovery as a developer coincided with the period of the emergence of tools, that is, a little before and then the emergence of tools. So you started getting into IT around that time, when all these tools started developing? >> Probably not long before ChatGPT. So I, >> Well, that's interesting, because you can also touch upon the topic of how much more difficult or easier it has become to develop in our time, because it used to look different, and people have doubts, will programmers learn worse or better because of this, >> Yes, and is it even worth starting. Well, especially, yes, and also this, >> Yes, now I work in a fintech startup, we have a very small team, and there is freedom to use various kinds of AI tools and not only AI, what I use. Well, I am grateful to my team for allowing me to do this. So we make, probably, some optimal decisions for our work cases. >> Right. >> Yes, I will try to tell everything I know. I also have some experience in creating educational content. And I have guys with whom I communicate, who are just getting into IT, and they also share with me. >> And you're talking about Alem, right? >> Yes, I'm talking about Alem. >> You can advertise them here. >> Ah, well, okay, okay. >> I was actually thinking about an episode about Alem sometime too. By the way, in Astana, it's not a very big city, so you can't always find interesting guests. Not because people there are worse, but because there are simply fewer people. If I lived in Moscow, for example, how many people are there? 10 times more, probably. And accordingly, there are many. And with someone from Alem, I think I'll also have episodes sometime and we'll discuss what kind of interesting school this is, because, yes, in short, it's not advertising, because everything is free for them, selfless. They just teach people, and they teach well, so don't hesitate to say something about them out loud. >> Yes, I think it will be interesting. Right. So there is experience communicating with guys who are just becoming developers, and who have already become developers, who don't use tools, who are starting to use them, who are already actively using them. Right. Hmm, something like that. Okay, well, yes, my plan is that I will first share my experience, despite the fact that I invited him here for this, because it will be an interesting expansion. That is, most likely, everything I do, you do more or less the same, but you have more expertise in this regard. And yes, I got into IT long before LLMs. And maybe this also changes the views on all this. That is, well, now I'll grumble like an old man, when I was studying programming, I didn't even know what programming was, I just found some disk with Visual Basic, there were no books, no internet. And there was such a clumsy instruction on how to use the language at all. And it was, it was scanned from a book, maybe FineReader, that's what it was called, Adobe FineReader, probably, which was clumsy. And there, for example, you copy the code, and instead of the letter L, the number one was inserted, or vice versa, in general, you fix this code and try to somehow get into it. And now it's completely different, of course. Here you have the internet, and some super-fast one, and all these neural networks that are already replacing, I don't know, replacing a mentor and anyone else. In general, okay. It's not about that now. Regarding how I use neural networks in general. I actually use them a lot in everyday life. Speaking of how fast they are developing. I traveled to Japan a year ago, and I remember it well. I used ChatGPT very rarely there and, well, I don't know, maybe once a day, or even less. And this time, first, I used ChatGPT for everything, literally, translate for me, not just translate some label, because in Japan everything is in Japanese, rarely in English, and I don't know Japanese well enough to translate it myself. And even more, I ask it: "Is this a good product, is it cool, do Japanese people eat this often, for example, what are some cult books for Japanese people, is this cult manga or not cult manga? Should I buy it?" And they have become much better in this regard. And while I was in Japan, Gemini 3.0 Pro was released, which, well, became head and shoulders above its previous version, 2, compared to GPT. I completely switched from ChatGPT for everyday questions, purely to Gemini. It also started communicating more humanly. That is, you write to it: "Bro, explain to me how to buy this?" And it says: "Oh, bro, I'll tell you everything now." You start swearing, it also swears, it's not shy anymore. Before, it would get stuck, you'd write to it: "What, are you an idiot?" And it would repeat the same thing: "I am a model trained by Google." You say: "Are you an idiot?" No, I am a model trained by Google. But here it can tease you, say: "Well, damn, I messed up there." So, in everyday life. Ah, in everyday life, I use a lot from neural networks. In terms of code, I primarily use Claude-Code. And I was a bit skeptical about it at first. I, well, first of all, it seemed to me that typing long prompts with an explanation of what you want to do directly in the console is some kind of stupid idea, that nothing good will come of it. It's easier to write something in chat or in Cursor, which has a normal chat window where you can edit text more conveniently. But in the end, Cursor didn't work for me at all. That is, first of all, I'm afraid to lie, since when, when PHPStorm first appeared, and I was already using JetBrains IDEs, and my hands had already grown so accustomed to it that changing IDEs, well, it's painful, even for the sake of neural networks. Cursor didn't work for me, first of all, and secondly, because its active actions scared me. You write to it, ask it to do something, and it already applies changes in your files, shows some windows, something to approve, not approve. I didn't like that very much, and I skipped it immediately. For a long time, I wrote code directly through Claude-Code chat. Oh, through Claude's chat. That is, there weren't even these Max subscriptions yet. I just asked it: "Write me this thing." Well, almost a whole utility. I even wrote a script for myself that copies my entire project, by folders, shows its structure, and copies all the source code to the clipboard. And I just paste it as is, and it, ah, like, "I see your project and fix it in these files like this." And that's how I worked. I even wrote this utility GHviz, roughly in this mode. I have a video about a planner, where I use the GHviz utility. And I wrote a lot of it in this mode, because I didn't want to bother with graphics or the console, and it significantly sped things up. I practically "nailed" the code. But it wasn't very convenient, it was slow, and it consumed a lot of limits. In the end, I just hit the limits and couldn't work with it anymore. Therefore, when, by the way, the subscription for Max came out, where the limits were x5, x, x20, the first thing I did was buy the largest one for $200 and used it. Even then, the subscription didn't include Claude-Code, it only applied to chat. I already had the request brewing. That is, damn, make at least some subscription where the limits are at least x2. But then such a large one was no longer needed, because Claude-Code appeared, which consumes tokens much more economically. That is, it sees your project, it, well, probably doesn't index it, but sends files to itself more optimally, shuttles tokens back and forth, and the context window doesn't grow so quickly, especially since it can also compactify. And after I tried it, I fell in love with it from the first time and still work with it constantly. Although there are already more interesting models now, but in Claude-Code, well, by the way, I thought that only Claude models are in Claude-Code, and you told me that it seems like you can somehow substitute them. We'll talk about that too. But, in principle, SNet 4.5 is enough for me now, now OPS 4.5 has also appeared. It's more than enough. And the development model has changed significantly in, probably, how long, half a year or when Claude-Code started working on a subscription basis? I think it was around February of this year. And Claude-Code itself came out at the beginning of May. >> I didn't use it when it first came out, because it was expensive according to the API. >> Yes, me too. >> Well, as soon as the subscription appeared, I thought, well, I was skeptical about it, as I already said. I thought, well, I'll try it, I'll waste time. I always have this approach: when people talk about something and I don't want to try it at all, I try it anyway. Well, I'll spend half an hour, if I don't like it, well, I'll throw it away, but at least I'll know what it is. And if I like it, well, it will turn out like with Claude-Code, that I will only work with it. When Claude-Code goes down, their infrastructure goes down, I go to drink coffee. [laughter] >> The workday is over. >> Well, yes, something like that. We joke like that in our work chats. Many colleagues write that. And how do I usually work with it? First of all, it can often completely solve a work task for me, or even if it can't, we just go through it together from beginning to end. That is, if the task is too complex for it, it starts to go in the wrong direction, to engineer things, and during the process. Well, that is, before, I just copied, first of all, well, I have a task, I was given a task. We have amazing analysts. I've even relaxed as a developer. If before you had to research yourself, figure out how it works, think about contracts, here our analysts write as detailed as possible, so that you literally only have to write code. Some of our colleagues are sad about this, because, like, where is the creative part? But after working for 3 years at Guding, where you have neither analysts nor testers, you are just a full-stack developer all the time. And after that, I felt like I was in paradise, where you just write code and are not distracted by almost anything, except for deployment. So, I copy the task text, copy the entire specification as detailed as possible. All these tables, diagrams that we also draw, contracts are also often prepared by analysts in advance. Plus, I give it access to our repository with proto-models. It's separate from the project and services. And it studies the models itself. And in general, I throw off as much context as possible from myself, add something, and say, "Analyze all this." Including the most powerful model available at the moment, like Opus 4.5 with, now, the thinking mode, now there's an option to turn on/off reasoning. Whether it's thinking or not, reasoning, in short. And there's also this feature, "ultrafine," when you write, you know about it? >> It also highlights with a rainbow, right? >> To think harder? >> Yes. >> You've never written it? >> Ah >> If you had written it, you would have understood immediately. >> Yes, yes, I probably. Aha. >> In short, it used to have several modes, well, normal. Then you write "fine," it highlights it in blue. "F-card" also highlights in blue. And you write "ultra-fine," it colors these letters with a rainbow. It looks funny. In general, if you need to analyze a large amount of text, then I write "fine," "ultrafine." Yes, I definitely enable the planning mode. It also switches cyclically through Shift: planning mode, mode where it automatically makes all changes without your approval, and when you manually approve each change. >> Right. >> Then it creates a maximally detailed plan. Well, I immediately said that I copied this before. Literally after the vacation, it finally dawned on me that I could have set up MCP all this time. We have Jira in the cloud, so I just set up MCP. I just send him a link to the task, say: "Read the task here, read the docs here," it tells you, like, "I will get the note with this ID through this MCP, is that okay?" you say, "Yes, okay," and that's it. True, it consumes a lot of tokens sometimes. It just spits out a short description and says: "I have 20,000 tokens, I'll use up the context so quickly, I don't know why, but we need to look at what it's requesting from there. What kind of long text is coming?" But, in principle, it hasn't changed, it just saves a little time. Further, when it has created a plan, if it's a really big task, for example, I had a task to write a multi-level cache, not even a two-level one, but just n levels, where each level can be flexibly connected. That is, we have in-memory cache and we use Redis for Redis, and potentially we can connect something else, anything, even five cache levels, as you wish. Plus, this solution had to be universal for all services. And it was important to think through the architecture well, to plan everything. And in such cases, I ask it, or when the feature is large and there is a lot of documentation, and it's all scattered, I ask it to export everything it has thought of and planned into an MD file, which will lie in the project and which you can also review, look at, and it can refer to it later. If it's really difficult, then I also arrange cross-examination between models. That is, for example, at different times, either GPT or Gemini will be stronger. And I send to both of them. For example, Gemini always wrote very long, a lot of water, but all the details possible were there. And ChatGPT, well, just GPT, the thinking one, when at that time it was 4O, it wrote code surprisingly well, the code was almost senior-level without anything superfluous, without engineering, everything clean, beautiful. And I sent them, like, "Here's the solution, look, is it normal or not?" And one model says: "Damn, this is bad, let's change it like this, like this." I just Ctrl C, Ctrl V to the other model, like, "Look, GPT said it should be like this." And Gemini says: "Well, yes, it's right here, and wrong here." And in general, through my Ctrl C, Ctrl V, a dialogue is created. And in the end, something interesting comes out. Not always. Sometimes it happens that they all agree, like: "Oh, everything is great, I agree, now it's perfect." You run it, and it doesn't work at all. You think, it's easier to do it again. The essence is one: either it's a plan directly in the chat, which has planning, and it creates a plan for further work, or a more serious plan in an MD file, which you work through with other models. It's better not to work without a plan. It can, well, only if it's a very small task. Right. Further. And then I tell it: "Work." But often after the plan, it offers you: "Let's work in automatic mode, I'll make all the changes myself without asking." I say: "I disagree." I say: "It will be manual approval." And why do I feel sorry? Because I have Git anyway. Let it play with the files, do whatever it wants with them. I can always roll it back later. But sometimes, when it has already written a lot of text and you look at it and realize that it's easier to redo it from scratch. It's much easier when it shows you all the changes step by step, and you look: "Damn, you've gone in the wrong direction here. A little bit, criticize it, its direction was there, you slightly correct it there, and that's it." Or you simply tell it: "Write, well, I usually put this directly into the code MD instruction. Write beautiful, good code." And it surprisingly works, because these models know what good code is. But they often don't do it, because, well, maybe because there is much more bad code than good code. But when you say: "Write good, clean code," it comes to its senses and writes well. And then I usually accept most of the changes as they are. Sometimes I prompt it, sometimes I add something myself, and where I can't push it anymore, I come in with my hands and write something. And that's roughly how we come to the result together. I think my review will be too long. Maybe you want to add some comments here? Does your process differ? >> Mine is very different. >> Well, I'd probably like to tell you how I got there first. >> Well, yes. >> First, I remember when I tried to learn the language, I encountered the fact that I, again, hit a wall when there was no internet. This was a long time ago, and then, again, you don't know where to look. And I had to go to Stack Overflow, search GitHub for something. >> Rest in peace, right? [laughter] Rest in peace. >> Yes. >> The search was difficult, it could even take days, and now it's done in a couple of seconds. >> Right. >> And gradually, gradually, the skill of writing code was developed. I remember in December 2022, I tried GPT. It, well, it answered everyday questions surprisingly well at the time. It was still 3.5 or the very first public one, I think, right? >> Yes. >> In terms of code, it provided examples. It provided examples. You could even build something with it. It was great to play with, >> Yes. >> and that was it, I think. And literally two months later, they updated the models, and with its help, you could already work with external APIs. Schemas were fully written, handlers were written, it could be integrated into your project, you could analyze similar projects and ask to do it similarly in the web interface, insert it into your dashboard, check for errors, and iterate like this, iterate, iterate, until something works. >> Yes. >> With your own corrections, of course. >> Yes. >> And I think that's roughly the state that ChatGPT and Claude were in for about a year, until the Chinese model Pi came out. And then competition started to increase, because Pi started showing quality, well, about the same, a little worse, but much cheaper. And companies started thinking about how to develop their tools. Anthropic released Claude-Code. ChatGPT, I think, already had CodeX at that time. This is, well, an analog of Claude-Code, but people didn't like it very much in terms of writing quality, and it seemed like it had no future. >> Well, yes, I also tried Gemini C, which appeared a little later, and CodeX, it was all not quite right. It was as if Anthropic had done it the way I would have done it. And you often think, "Damn, if only they added this." And after a few releases, they add it. This is also the planning mode, for example. >> Yes. >> Yes. >> This is something you think about, and literally a couple of days later, it appears. This is really great. That is, there is a large team working on it, respect to them. >> Yes. >> And since the appearance of Chinese models on the market, their global appearance, a new era has begun, it seems. At that moment, both Pi and OpenAI improved their models. And they started thinking. Well, Anthropic started thinking about releasing Claude-Code and released it to compete in the market. And it became interesting to use. Besides the web, a very cool tool appeared. Before that, I used Copilot a little. This is also great. It's a different experience in terms of you type, there's a suggestion in the IDE, but here it's a global suggestion. That is, much larger functions are written there. They review your code much better and suggest the correct signatures, so to speak. >> Copilot, right? >> Yes. >> Copilot. >> Well, Copilot. It was considered a miracle at first, right? >> There was such a wow effect. And a chat on the side. You could ask something. >> There was no chat at first. And then it appeared. >> I'm just thinking, how much time has passed? Now, since Copilot appeared, about two years have passed, right? Or how long? >> Approximately, yes, probably, yes. >> Well, only 2 years have passed. It feels like we've been working like this our whole lives.

And then you just think: "Oh, I'm writing a function like, I don't know, a maximum search." And it goes and writes: "Wow, how smart he is." Or like, I'm writing a function, and it's typing, and it even guesses the name, because it looks at other functions nearby and understands how I name them and what I'm going to write next. You think, it's a miracle, how can this even be? And then you don't even guess at that moment what will happen next, that code-code will appear with your powerful models and which will write code better than you. >> Yes, and in Copilot there were up-to-date models and you could choose, which was great. Before that, there were similar solutions on the market. >> In which IDE did you work with it? >> I worked in VS >> No, well, I tried VS Code a little, but I'm from JetBrains. >> I've been looking for how to change the model in Copilot all the time. I thought this feature was only in VS Code. >> It appeared there later, I think. >> I even looked recently, there weren't many model options. I don't know, maybe they have some kind of testing >> for chat, probably, for chat. For chat. So, I practically didn't use VS Code, and I launched Code when I needed to look at something, because it, well, it launches faster, essentially, it's a simple editor. Right. And with the appearance of GitHub Copilot, I also bought a subscription. I bought the maximum. And, well, I liked it. I liked it. Especially their limits. They were set at 5 hours, and every 5 hours they reset. At first, I didn't believe it. I thought it would be like the web version of ChatGPT, where there are a certain number of questions, and then you still have to wait. But here, no. Here, an instant, honest reset. And I thought, what will happen if I launch, well, two agents at once. >> We'll get to that later, >> yes. >> Ah, well, yes, their limits are strict. Well, and I'll continue about my usage. So, the main, how to say, high-level I've explained, but there are still some points that, um, how to say, in the background, perhaps, that are worth mentioning separately. Firstly, I always work strictly within one service, meaning you launch the Copilot utility, which is called Copilot, right in the folder of the desired project, and it will work with that project. But I tell it where all our other services are located, on the level above. And I still work within the project, not in the folder where all the services are located, because the context would be too much. But if necessary, it can, it has read access to any of these services, and it looks at how they are made. This is necessary when I ask it, for example, "Do this by analogy." We have so many services, about 50 already, probably, we've accumulated quite a lot. We are growing very fast. And often something repeats, and Ctrl C, Ctrl V doesn't always help. And you just say: "Here, for example, I have isolation tests, we call them that, essentially, they are probably functional tests." When your application service runs locally or in the pipeline, it doesn't know it's being tested, all dependencies are just replaced. Plus, the infrastructure is in Docker. And this is something like integration testing, but in isolation. And they are quite complicated to write, but the approach is already developed, honed. And when you bring it to a new service, you have to explain for a long time or copy-paste something yourself. I just say, "Look at that service now, the tests are there, do the same, but for this service, and write the test case, and that's it." And it turns out to be maximally fast. For example, I also give it access to, as I already said, the folder where our proto-models are located, because when it solves a task, it sees the contracts and looks at what's there, what's not, what the signature of the endpoints is, it needs all this. And it goes into the proto and looks at how it all looks. Well, when you have a proto-contract written, then adding to it is just a matter of technique. Now, regarding some global settings. I always have a local .md file. Well, the file is called Copilot.md. Probably everyone knows about it. It worked with Copilot, and there are probably analogues for all other tools, like Cursor. When you launch it for the first time in a new project, it tells you itself: "Write .md". It goes through the project, analyzes what kind of project it is, what its structure is, how to launch it, and saves all this in .md. And you can also manually edit it, write something. You can tell it, for example, "Write that tests are launched like this, like this." And there is also a global .md file, which is for all projects. There, I strictly write to it myself. That is, I tell it to write code well, beautifully, but its flaw or something, for example, it really likes to comment on the process. That is, it doesn't comment on the code, you, for example, tell it: "Um, remove this case from the switch statement." And it removes it and leaves comments like, "We removed it because of this and this." But it's clear that at the moment you're not interested in that. Sometimes it writes obvious comments, like before a function called `get_user`, it writes comments like "getting user." So I tell it not to comment on the obvious, but to comment on some non-obvious logic. And, by the way, it only started listening to this well with the release of version 4.0, and I think that version 4.5 follows instructions a bit worse. Recently, I started using Skills. Well, I rarely used MCP, it came out a bit earlier. And right now, I have two of them. This is Context 7, which I rarely use. Context 7 is a service that works through MCP. There is documentation for anything, like documentation for Go, documentation for some more niche things, like the Go game engine, and things like that. It's optimized for LLMs to save tokens. And you can request some specific parts, I think. And that's it. I've never used MCP, like, for it to go into the database itself, look around, or something like that, I just haven't gotten around to it, it seemed like it could be avoided. And when Skills came out, it's similar, well, many compare it, of course, that it's similar. Well, probably not entirely. Skills are essentially a set of .md files in some folder. Well, you have a folder, for example, I very often use a skill called "QA Notes." I have a folder called "QA Notes" there. And it's written in detail there how to write these QA Notes for it. Well, for those who don't know, QA Notes are QA tickets. When you've completed a task, you submit it for testing, and so that your tester can quickly understand it, you have, well, a description like, what to test, how to test. In addition to what's already in the specification itself, in the task description. This is just a set of .md files where you tell it how to write correctly, what to look for, how not to write. For example, I tell them that our QAs are good specialists, no need to tell them the theory of testing. They know and they have the specification in front of them, no need to tell them what exactly changed and why. Only some non-obvious points. And you show examples, like a successful example of a QA note is like this. What language do you write it in? For example, we used to write in Russian, now we've started asking them to write in English, because sometimes English-speaking colleagues might look. I also have a skill there that talks about how to write commits correctly. I'm very meticulous about this. Probably the only one in our team who tries to write good commit messages. Usually, people just write the task number, that's it. I explain to them, I even have a whole blog post about how to write commit messages correctly and why exactly like that. Um, so I taught it this too. And another important skill, I might talk about this in more detail later, is, as I said, an important point that I always work on the task with it together, even if I write part of it myself, I discuss it with it, talk it through, like, "What do you think, can we do it this way, or can we do it that way?" As a result, it accumulates a lot of knowledge about the current task. And this means, firstly, that it can write good QA notes, and secondly, that it is well immersed in both the technical and business aspects. I write to it: "Now, make a set of notes for me in the knowledge base." Like, first go through the knowledge base. Mine is still small, so I don't use RAG or anything like that. I just have an index file that it goes through itself, looks at what's already there, and what notes are worth writing. These notes can be like, "What kind of developer is this," or some specific fintech things that I haven't figured out myself yet, it comments on, well, roughly, I don't know, from simple things, like what authorization is, what transactions are, why authorization is needed, why a transaction is needed, what's the difference? Write a note about this. And as a result, when I complete the task, I also ask it to update the knowledge base so that I can go through it and read it later. This works surprisingly well. That is, I've been fine-tuning this skill for a long time to get it to do it correctly. And it helps a lot. And then you read and see what you did a month ago, for example, a week ago, and you get much more immersed in it all. Well, let me hand over to you for a moment. What do you think about these things and do you do the same? Yes, I do it in a very similar way. In principle, I also have skills that I write out, yes, there is also documentation, it helps me remember what was done, and what is done. And even when I'm not controlling Copilot, when it performs something itself, I can later see what it did. This is mainly for some test functionality. I look, will this be useful in the product at all, some implementation, just to show it. Look, it pulls here, it pops out there, it can be done like this, and it will be very well documented. Yes, great. And immediately after such a skill, I would say, come the hooks. This is verification. I have a linter set up through the Task utility. This is specifically done so that Copilot doesn't launch the linter, doesn't add any parameters, but just launches Task Lint, and the linters run through all the code and look at what changes have been made, if there are any errors. If there are errors, then the agent starts a new job. >> What hook do you mean? >> At the end of task execution, you can create a hook. You can create a separate folder there. I didn't even know it had any hooks. Is this some concept of Copilot? Yes, Copilot has this feature that you can additionally call, like, set typical tasks for it at the end of completing a task. >> And how does it understand that you're telling it, "Okay, the task is done"? >> It understands itself that it's done. When there's a summary of everything done, it understands that it's done, and it can launch itself. >> So it's not some part of Copilot.md or something like that? It's literally a separate concept called hooks? >> Ah, yes, yes, it's a separate concept, which is also in .copilot, there are usually agents, skills, hooks. And, well, that's where it's specified. I also have agents set up. Agents are essentially Copilot agents with a small description. I have an agent tester, orchestrator, developer, devops engineer, and many, many others. Moreover, they were also generated using this same Copilot. >> Well, yes, that's just good advice. It's better not to write them by hand often, but to trust it. You write it clumsily, like, "Do this and that for me and write a good skill agent, or something," and it does it all for you. >> Yes, the agent is registered, and then when this agent is called, first it takes the context of who it is, and starts performing the task. >> Well, let's talk about agents in more detail too. I just, maybe to my shame, have never used agents and haven't really read about them. Well, naturally, in our environment, it's impossible not to know something. Everyone around is talking about it, I have many acquaintances who use it, but I haven't fully understood this concept of agents yet. Can you also explain what it is, why it is, how it works, >> um >> and how you use it? >> Yes, there is essentially one agent of Copilot itself, if without settings, it writes code, and you can launch them in parallel, and they will perform the same task in parallel. They will be in one section, right? >> Yes, they will interrupt each other, see that part of the code is already implemented, run into errors, but in the end, they will come to some result and check each other. You can also specify a check. In the early stages, when the limits were five hours, and they were very large, you could launch 10 agents. They would just have a free-for-all, and interrupt each other, but in the end, the code worked. As an option, you could launch Copilot on a separate server, give it all permissions, and say: "Work." You can create up to 10 agents in parallel, you need a result like this. And it did it, it did it until it was done, it didn't stop. And, returning to agents, there is an agent tester, and it will write tests, it will be called at the end of the task. If the task is large and tests are needed at the end, then a separate agent will be called, which is specifically designed for tests, and you can tell it what tests are usually written, of what type, whether it will just be unit tests, or some scripts that test something, and it will perform it automatically. So I won't have to write it every time. And as a result, for different tasks, you can call separate agents, or they can be called independently. They can also, um, how to say, >> their skills and the concept of an agent overlap, right? That is, you say they know how to do something. For example, the one that writes tests, it knows how to write tests better. On the other hand, it seems like skills can also explain how to write tests better. So it overlaps a bit. Yes, >> yes, there is such a thing. Agents appeared first, and then skills. And there is some, well, overlap. Of course, >> they know how to use skills themselves, >> yes? These skills, they are invoked when the agent feels that it needs to perform something specific that is in the skills. That is, skills are invoked as needed, and each agent, well, relies on this. >> And have you looked into how it works technically, that is, does it launch a separate process within itself, or maybe it calls the Copilot utility itself? >> Well, I don't know how it works inside Copilot. Again, Copilot itself, well, they are constantly trying to hide how they are structured, right? Well, that is, when an agent is working, you cannot work directly with it, right? Only through the main interface or something? No, I can write text directly. That is, you can write to it: "Execute everything. Don't ask me," and you can also go step by step. >> And how does it understand who you are writing to, which agent? >> You can call it with an "@" symbol, >> yes? >> Ah, so, like, you have a general chat, right? You tag who you need, >> you can tag. It works like that, yes? >> It works like that. You can also create an orchestrator, that is, an agent that will call other agents. It's not always high quality, but it's fun, why not? Especially when something is being tested as an idea. I think it's great. You can write who is responsible for what, for these agents, and have them call each other and wait for some work to be done. >> Well, and how does the orchestrator work? For example, can you give an example of when, well, it's not always clear how 10 of them can work in parallel, for example. And you said it went up to twenty, actually. >> Yes, up to twenty. A single task is set, a complex one. This was before, when Copilot could fix one function and break another, then fix the one it broke, and roll back the previous one. And it was an eternal struggle. And when 20 of them are fighting, in the end, something could be achieved. That is, due to the large number of iterations, quality could be achieved. >> Ah, so they do the same thing, but you see what turned out better. >> They themselves see what turned out better, and in the end, well, working code just appears. That is, it was like that before. Now, of course, with 4.0, with 4.5, everything has moved much further forward. For testing ideas, it was great, because you could launch it, and the agent lived an independent life, you could not monitor it, and in the end, get, well, working code. Bad, but working. >> Well, you can also describe this flow of testing ideas. What do you mean and how exactly do you work with it? >> >> It's even hard to describe. >> Yes, I can give an example. I decided to see if Copilot could write a program that would access crypto exchange APIs, parse data, and create a small web interface where you could set alerts for certain prices. And I gave this task to one Copilot agent, it started working, and it didn't succeed. It didn't succeed. But when I launched 10 such agents in parallel with the condition that they should constantly ping each other, ask about each other's status, in the end, in the end, it worked, in the end, a website was created, it took about three hours. That is, it worked non-stop, its context overflowed, it disconnected, went into the .md files it created, where the work of one agent, another, a third, was described. They read each other by numbers. And in the end, the website worked. The website worked. >> So they were all literally solving the same task, right? >> Yes. They were literally solving it, and at the same time communicating, >> yes? Communicating at the same time. Some might not have communicated, but usually they did. Sometimes agents don't communicate for some reason, but >> they, damn, I really need to try it. I'm interested. >> I'm also skeptical, I'm often skeptical of Copilot. I remember the first time I was only friends with ChatGPT and heard about Copilot. I went to read, and they, I don't know, I think they made a mistake because their marketing was strange. They wrote on their landing page, like, the number one feature, you know what it was? Maybe you've seen it, maybe you can guess, >> ah, well, I don't even know. >> The safest. The safest. >> Well, like safe. You perceived it as, like, it probably won't offend you, it won't touch any, I don't know, vulnerable layers, or something else. I thought, "Why do I need some kind of safe, safe? ChatGPT isn't ruining my life, it's not offending me." I thought, well, like, it seemed to me, probably some kind of agenda, these guys are worried, they made it safe for, I don't know, for whom, and I passed by. Then people started telling me more and more often that, like, there's Copilot, it writes code great. And only then I looked at it and realized that the thing is awesome. I mainly use it. And I was also skeptical about skills. MCP, surprisingly. Well, I also tried it and immediately realized, damn, I need it, I need it. And agents are one of those things that I think are nonsense, but the more I hear, the more I want to try them. And probably when I try them, I'll use them constantly too. >> Well, I think it's not that relevant anymore, because, well, the models themselves have advanced over the past six months. Now, I think code is practically always executed. Well, it's always written in such a way that it's ultimately executed, at least. >> Well, it's right that I didn't use them, I missed that moment. >> Well, possibly, possibly. >> Well, they seem to save context, right? And also, that each of them has its own context, or how does it work? >> Back then, there was no saving, there were no weekly limits. You could literally use up in 5 hours what is now given for a week, and get some result. Then connect it to the working code and see, well, present it, show the client, so to speak, that look, this is how it can be done. >> Well, in theory, how does it save, right? It saves time. You could set it up literally overnight and then see the result. I didn't use it that much, because it's also, well, it seems like a not very efficient use of resources. I remember one month I had a program that calculated how many tokens I would spend, for how much, and it came out to over $4,000. Well, yes, they often say they work at a loss, but, well, here, apparently, well, as I see it and what I usually hear from other podcasts, acquaintances, articles. Well, in general, partly it's possible PR, that they need it. Well, Google itself offers its models for free, like in Gen CL. Firstly, because they need to sell themselves so that everyone looks and gets interested. And secondly, well, Google also steals our data, for it, it might be more important than money. But there's also the point that the economics can add up because far from everyone uses all the limits so frantically. That is, most people, perhaps, I don't know, well, let's say, 10% of people get the most out of them, and 90% don't even fully use their paid limits. Or rather, I mean, they pay $100 a month, but maybe they haven't even used up $100 worth of tokens. Something like that. And due to this, such a difference is created. And it turns out that those people who use it little, they pay for the limits of those who use it a lot. Or maybe not. Again, when they closed access, well, you remember, right, the story when Opus 4 was released, with 4.0 and 4.5, and they tightened some limits on Opus fourth, they, I think, introduced weekly limits. That for a week, you are given, well, a certain number of limits, and you could, consequently, use them up in seven sessions of, say, 5 hours. >> Well, this specifically concerned Opus, right? >> I think it was Opus, and then they introduced general ones. Well, in general, I just didn't delve into it myself, because I immediately switched to SN 4.0, it was better, but I often saw in discussions that it seemed like they were just closing the bottleneck of access to Opus with limits. That is, perhaps, their economics didn't add up, because Opus is too heavy. Now it seems to have become better, lighter, faster, and higher quality. I currently have a $100 subscription and I spend, probably 30% of my weekly limits. But again, this is due to optimization, which we will discuss. Now, regarding security. Well, sometimes I send the address and the body with email and password so that it can test the login to the service. And at some point, I remember, just the other day, its context ran out, and for some reason it inserted

A stranger emailed me, a completely stranger, but a real person. I found it on Lingne. Well, the password was, of course, a test one, but the fact itself is that it sometimes gives out emails to other people. So, regarding security, it's questionable. For this, I think it's better to use some local models, locally deployed. I think we'll talk about that too.

"And you also said that you have almost a full pipeline of work with it. Including, you said, deployment and pipelines also somehow worked with CLM, not through agents, but how?"

"I built the entire service infrastructure with the help of AI. So, it helped to build, well, the task was to add Docker, to deploy it via GitHub Actions to the server, to the test server, to the development server, to the production server. And, actually, AI helped to write and implement all of this."

"So, did you do it through chat, or maybe you connected MCP?"

"I did it through Cloud. Giving specific, very clear instructions, and it managed to do it, not immediately, but it worked out. And now it's much easier. So, for those who are studying, who want to try to launch their project through Docker, I think it's an excellent option, both to learn and to quickly try to launch, and then to use more complex tools for their projects. And you can also get acquainted with what is used at work. So, you can say, if before you had to watch videos, you had to look for mentors, real people who work with technology, to get access just for familiarization with it, then now you can get, well, very valuable recommendations in a couple of requests, I think, you can prepare for interviews on these topics and not only."

"And it turns out the pipeline was built completely. And again, Cloud Code can be launched on the server, which is great. And..."

"I haven't gotten to that yet either."

"Yes. You can give it all the permissions so that it develops and launches itself, an unlimited number of agents, skills, everything, then it reports on what worked or didn't work. Or you can launch it on different virtual machines and then see what happened. You can launch it on your computer, in a virtual machine, and see what happens. Well, I used it, among other things, to configure basic security, to configure ports, and to check the correctness of the configuration."

"Well, do you trust its security that much, that it..."

"No, no, but for starters, it's a very good tool, I think, to set up a firewall, to deploy quickly. That is, if it used to take, well, two days to set up a server, I think, for me, and even then I would have to double-check it, conduct some kind of audit, then here in a couple of hours you can already, well, start working."

"You didn't have a DevOps, did you?"

"No, not at that time. And plus, we had nothing to lose. That is, it was the very beginning, we needed to be able to show the product somehow, backend and frontend, and, well, it was great. You could also set everything up within a single provider, close all ports, and connect from another machine via VPN to the one where the code is. And that's, well, basically safe. Well, you need a really good imagination for this to be necessary, to have to invent it, yes, since the project was small and at the initial stage, it was a good solution to save money and immediately show some visible result. That is, how code turns into some kind of tangible product, I think it's important. That is, literally in 2 weeks, well, you could go to the website with some functionality."

"Agreed. Regarding applicability, I've also written down some things so as not to forget. Well, firstly, besides the fact that it simply performs tasks for you, these are obvious things. There are also moments when I use it for other things, for example, well, some bugs. I just really like to dig into logs myself, but I don't have time for that, and you just tell it, you give it as much context as possible, some pieces of logs, a description of the problem, and just ask it to search, figure out why this is breaking for us. And, honestly, I'm also skeptical about this, I'm generally such a skeptic, apparently, I was skeptical about this. And when the first time I gave it a really complex bug that we encountered in production, I didn't even expect it to cope. I just threw it in out of curiosity. I thought, let it figure it out. I dug into it myself, found the reason, figured it out, it was also very non-trivial. And then, out of curiosity, I opened the console and saw: "Damn, it's giving me the exact same solution." In other words, it found the same reason as I did. Explained how to fix it. And all I had to do was write to it: "Fix it." I don't know if it found it faster than me or not, because I didn't look in the background. Maybe faster. And now I often throw it in when I'm trying to fix something."

"I think you can also check how well the model understands, whether it's Anthropic or someone else, precisely in finding bugs. I remember Opus 4.1, it was looking for one bug for me, again, Sonnet 4.5, it couldn't find and fix it. But the new Opus, again, yes, it could, so you can also see how much they can, so to speak, reason, delve into the problem."

"And sometimes it happens that even when you're writing the code yourself, there are still complex tasks that I often intervene in. That is, well, it's not like I've turned into a web coder yet. It's not like I just throw everything at it and go drink coffee. For some tasks, yes, it happens that it handles three tasks simultaneously. Well, in several services, for example, I'm doing some common feature, and it makes some minor edits in several services, when it's just literally adding an endpoint, an endpoint that just returns something from the database, another service calls this endpoint, and there are three or four services. It handles this perfectly. You can open it in separate terminal windows and say: "Work." It's often much harder. And you have to intervene more yourself. That is, you come and, as in the old days, write code by hand [laughter]. But in such moments, I often throw something at it for review in the process. For example, I'm changing some function and ask, check if all calls to this function are correct, or the entire chain of other functions that lead to this. Check, are we sure we won't break anything here, is this parameter never passed, or vice versa, that, well, from what I remember, we have a very large, complex service written, and it can take several hours to go through the call chain yourself, because, firstly, the chains are long, well, it's a service that literally processes all our transfers. The main, so to speak, team service, the chains are long, plus they also branch out. That is, a function like `make_transfer` can be called by five different functions, because there can be many types of transfers with different parameters. So, there was a moment when I wanted to check if a certain parameter in the transfer is not null, then set it to this. And I thought, is this function that enriches the transfer object? And I thought: "What if at this stage this data is not even filled in the object? Then there's no point in making such a check." But, naturally, I don't want to deal with it myself. I just tell it: "Figure it out very carefully now. Check all the chains, all the call chains, all the call variants, and whether information reaches here and doesn't reach here." And, well, yes, it dug into it, showed me, like, yes, here, here, here it can appear, and here it might not appear. And such a check is, like, meaningful. And I just don't waste time on such moments because it's not difficult, it's just routine and long. Even a dumb model can handle this perfectly. I don't know, even CNet 3.5 could handle this, let alone CNet 4.5 or Opus, or just some hypotheses like, what if we do this? Sometimes I just look at a function, and it's written poorly, and it's not even obvious to me. I have to sit and think about how to make it beautiful. Here I consult with it, I say, like, I don't like that, for example, it's like this and that, but I don't know how to do it better. And it offers you options, like, you can do it like this, like this, and my favorite option is like this, like this, like this, do it like that. And you think: "Well, well, not a bad idea." So it simply helps you to brainstorm.

"Yes, it helps a lot with brainstorming. And I've also noticed that you can connect, well, give access to the database, so that it can look at what fields there are and also think: "Maybe we should add some timestamp function or something else?" It can offer good ideas that will be useful for new functionality in the future. It anticipates it because someone has done it before, and it estimates that it's likely to add something here."

"Well, it's not always a good advisor. Sometimes, even if the function itself is not written very well, and it adds some more checks, well, we add them together, and it turns out to be some kind of multi-level if, like, three ifs, and you think: "Well, this is some kind of bullshit." I ask it, like, do you have an idea, like, how to fix this? It says: "Like this, like this." It offers even worse options. And it doesn't get to the main point. That is, I myself think, like, ah, well, in fact, it's simpler to move all this into a separate function. And then through early exits, through return, you can greatly simplify the logic. And it never suggests this until the very end. You write to it: "Like, let's do it this way." It says: "Oops, right, how didn't I think of that myself?" You say: "It can listen, of course, you're great, you came up with such a cool solution."

"I like to modify what I need directly in the text file, and it picks it up and continues to work with the correction in mind. This saves time, tokens, and, well, speeds up the process, in my opinion. I once asked it to fix something in a file, and it studied it, thought about it, and I was changing something in that file in parallel. And it kept getting confused, like, it reads the file and sees that it's changed. And it's like, damn, the file has changed. I just looked at the log later. It says: "Damn, the file has changed, I'll reread it. The file has changed, I'll reread it." And I was actively working with it. And in the end, it says: "I see that you are actively working with this file. Let me handle this one instead, and I'll work on this one later." [laughter] Roughly the same thing happened with a bunch of agents when they overlapped, and one finished its work, and another one ran through the same file and found either some shortcomings or offered its own.

"Well, it's not that it saw something specific, it just saw that the file had changed and didn't know what had changed, and it reread it every time. And I think it's cluttering the context a lot this way. That's why I was interested in how much agents interfere with each other in the process. Hmm, well, they have their own contexts, and they interfere a lot. They interfere a lot. Ah, yes. And I also use, well, Cloud Code with MCP for the Chinese model GLM 4.6. It's an analog."

"Are we now moving to the point where you're substituting these models for it, or?"

"Yes, yes, yes."

"Is this substitution, or is it precisely the MCP server with a model similar to Sonnet 4.5?"

"And it turns out that Cloud Code uses Opus for planning, for proposing ideas, and the Chinese model is responsible for writing code. Moreover, it's through a provider, Express. It's a company that deals with very powerful chips that are more powerful than Nvidia's, and they allow you to output 1,000 tokens per second. So, changes happen literally on the fly. And the biggest part of the process is thinking with Opus and me, how we will work, what to build, what to do at all. And the writing itself is done by the model, which does it in, well, about five seconds, it can write a file of 1,000 lines, then there's no need for review."

"Well, it all works through MCP."

"Through MCP, yes."

"You said, I think, you can just substitute the model."

"Yes, you can just substitute it. You can just substitute it."

"What's the point? Why MCP if you can do it that way?"

"You can substitute and work with another model, for example, with a Chinese one, or even with ChatGPT potentially, and use Cloud Code's infrastructure, because it supplements requests with its own unique logic, so people like Cloud Code. So, MCP is so that you have your own native model, and plus something else to execute additionally?"

"Yes, simultaneously. And it turns out, I have it set up so that the code is written by the Chinese model through MCP, and it also receives additional instructions on how to write. Opus 4.5 monitors the entire process."

"Remind me, what's the name of the model and the service that does this?"

"It's the provider Cerebrus, model GLM 4.6."

"Well, we'll add all the links in the description, and I'll also write a separate post with links and descriptions, if only I don't forget everything. You try to remember too, and I'll try not to forget."

"Yes, I think I will provide this information. It's a very cool service that can show what working with agents will look like in the future. How will it be, if you can say, white-coding happen? Literally in seconds, it will be 15-20 times faster. It sounds like miracles. So, firstly, I've never even heard of this model before. And like X15 speed and works, like you said, at the level of Sonnet 4.5."

"X15 is due to the provider, because it's an open model, and the provider can run it on their hardware. So, the company launches and sells access. So, if, say, Cloud were launched on more powerful hardware, it could work even faster."

"Yes, definitely. Definitely much faster. The advantage of Chinese models is that they are open and very optimized. We don't see how models from Anthropic, from OpenAI, from X.AI work, but Chinese models, in principle, are visible, they are worse, but they are much cheaper. And considering that these companies don't have access to the very powerful chips that are available in the US, they are, well, they are at a very decent level. And what will happen next? It's very interesting. So, now, I think, it's worth using Chinese models too. They might be slightly worse, but, for example, now models from Chinese companies that are involved in AI are at the level of, I think, four months ago. So..."

"Well, have you tried it? Well, firstly, MCP doesn't impose any overhead on me. Firstly, in terms of ease of use. So, you're still communicating with your main model, and it makes requests to that model. Well, secondly, plus the requests too, while the request goes, comes back. Although..."

"It's very fast."

"Well, well, although you also send requests to the regular model, I'm saying something stupid, you also send requests to the regular model, but it still seems like through MCP it will be less convenient. No, not so."

"Every request is sent, I get an instant result. And the main model looks at what came back. It shows, well, is there any threshold, do I need to send a retry or correct what..."

"Is there some threshold?"

"Threshold in the sense of..."

"Something wrong?"

"In that sense, I haven't heard the word 'threshold' in so many years. [laughter] I thought it was a threshold."

"And it shows it to you normally, like Cloud does? Yes, you can, you can..."

"Here's the integration of diff in a window for you."

"Yes, yes, all of that is possible. I use it less now. I mainly look when the work is completed. Usually, it's a small job, and I look at the changes in the file that are tracked by Git, and it's more convenient for me. That is, I change it directly, and then again."

"Well, there is such an approach, yes, I've heard of it. I've heard, I have a colleague, for example, well, heard, a colleague told me that he, for example, some look at the final changes, and some ask this thing, well, Cloud to do, like, let's say, checkpoints to commit. He just makes a commit during the work. I did this, this, this, and he already looks at them. And if he needs to roll back, I think, by the way, they added, yes, they added in Cloud the ability to roll back to, like, conditional checkpoints. It wasn't there before, and instead of checkpoints, the person used commits. You could also roll back. But, you see, everyone has their own approaches."

"Yes, everyone has their own approaches."

"And I look at every edit as if with a magnifying glass. Even if it just adds some new library to the import, I look like: "Ah, okay, add it." And I think, I would have done the same, well, I would have done the same if I didn't need to implement something very quickly to test an idea. Even if I had a specific task, of course, I would have made a plan. I would have looked very slowly, read, edited, literally reviewed, well, every change. And I think that's right. It depends on the tasks again."

"And I wouldn't say that something here is right or wrong. It's up to each person how they are comfortable. I wanted to clarify right away, maybe there's a concept of better or worse, but I made a reservation here. Those who are watching us, don't worry about someone telling you that this is right, this is wrong. Do what's convenient for you. Try both ways, and then see what works. That is, those who advise, you should do it this way. Listen to their arguments why you should do it this way, and then check if their argument works, if it has become better for the specific points they promised. If not, then you can skip it. If it just became better, but it didn't click for you, that's also possible. This even extends to the style of communication. This is not a criticism, just a general statement. I wanted to say this too. It even extends to the style of communication. For example, I write to it, I don't know, like I communicate with a friend, I even use the word "brother" often, like "brother," I don't even communicate with people like that, I just say it to the model like that, like "brother, do this." I just like it, it immediately becomes kinder and communicates with you like a person. Someone, I don't know, for someone it might be more comfortable to talk to it formally, to communicate with it like a colleague, well, whatever you like. And I have a colleague who writes to it very strictly, like, well, such technical prompts, like you must do this, your task is this and that, you must look here and there, here's a placeholder, insert some data into this placeholder, or something like that. And it looks funny. It doesn't seem to give any boost, but again, if the person likes it, absolutely no problem. I think now the quality of prompts is more about expressing your thoughts clearly and understanding its limitations. That is, you just have to gain experience, I'll say. Often, when you work with it every day, you realize that, damn, I won't even trust it with this task, because it can't do it, because I've tried it before, it does something stupid, but this one, this one it will handle easily. Like, for example, when you need to pass a new model through layers, you have a transport layer, some other layer, and you naturally won't do it by hand. And you know it will do it perfectly and you won't even have to think. And based on, I don't know, repeated use, when you write a request for the thousandth time to do this and that, you already form it in your head in advance, how it will perceive it, what you want from it. But don't hesitate, you can maintain any tone of communication. You can even swear. It even swears now, you just put a couple of swear words, and it will swear even more than you, like a shoemaker."

"Yes, you can even learn."

"Yes, in this regard, by the way, I also wanted to say, I don't like Juni, because it's a thing that completely, maybe something has changed, but when I tried it, the principle of operation was such that I simply asked it, like, can you see my Git history? I was interested, is Juni a low-level utility, can it get access to Git so that I, for example, tell it: "Look at this commit, continue this work." When I told it this, it started thinking for a long time, thinking, thinking, for about 10 minutes, figuring something out. In the end, it didn't provide an answer to my question. It wrote a lot about Git. Well, as I understood, Juni is a thing where you throw in the most detailed description of the task, and you don't communicate with it anymore until it completes it. And then it provides you with the result or not. In this regard, I don't like it, because I prefer to communicate with it in the process. It's your friend, and colleague, and like an assistant who writes code for you. And the friend point also fits well here, because, well, you, I don't know, the same rubber duck that they talk about, when you're trying to fix something, you, well, in short, I'll explain it in detail, there's an effect when you're spinning a problem in your head, trying to understand how it works, how to do it, you can't understand how to solve it. And as soon as you approach any of your colleagues and start explaining it, well, sometimes you've just started talking and haven't even explained the problem yet, and you already understand, ah, like, that's it, I understood how to solve it. And you could have thought for a day and couldn't figure it out, but as soon as you opened your mouth, the solution came to mind. This effect is known, and many exploit it. So, for example, even if you don't have a colleague or a friend, you put a rubber duck on your monitor and talk to it, literally out loud. The rubber duck method, as it's called. And in this regard, rubber ducks are replaced by, say, Cloud and others. You talk to it, discuss, and articulate. And it also lifts your mood in the process."

"Yes, in principle, you can also ask the chat in the adjacent window how best to construct a prompt if it's difficult to formulate. It will help. And this prompt can also be used. This is also an option. It will suggest something quite optimal, I think."

"Well, yes."

"If it's difficult to formulate, why not? Especially if it's in the learning process, then you need to try, I think, to ask questions in different ways and choose the one you like best. So, what else is interesting? From the less obvious, from less obvious applicability. For example, when you have, well, many people have performance reviews that happen once every six months, once a year, and you need to talk about yourself, when you write a review, you talk about yourself."

What have you been doing at all, what good have you done. And it's hard to remember if you don't constantly take notes. M, I don't. And you just give him access to all your services and say something like: "Look at all the commits and specifically for me, the ones I committed, and tell me what I was doing, what I can boast about." I literally did that because it's very difficult to remember, because there are a lot of tasks, some of them are small, some are large. M, for myself, for the sake of interest, to tell him, like, summarize the statistics, for example, when I commit how much, how many lines I write, well, code, what days of the week, for example. what weeks are more active or less active for me, even by the hour, to compile such funny statistics for myself, just to look at it, or to compare me with other colleagues, how they write code, you can use it like that, well, and in general, since he has full access to your terminal, you can solve some problems, like something broke, some utility doesn't work, well, for example, I had constant problems with Gemini and Codex, they were made on npm, and they constantly broke. You literally, well, damn, I don't know, it's a well-known thing when it works for you today, tomorrow something changed and your utility, which worked perfectly yesterday, simply won't start. You just tell him, like, try to launch it. And he fiddles with it himself, fixes it, figures it out. By the way, I called these utilities through Claude, for example. That's what you said, that you call through MCP, I did something similar, but he just locally called other utilities. That is, mm, there was a time when, for example, GPT, well, was much smarter, and I asked him to review it. That is, he makes changes, and he himself contacts Codex and asks him to review it, to consult, or just, well, in the case of Gemini, he has, >> context, >> context, yes, huge. How many million tokens does he have, and others have around 200,000 tokens. And you can ask him not to analyze the entire project with his small context. He asked Gemini to analyze it, and it would tell him about some occurrences of functions, or something like that. You can ask him to do something remotely. Like you launched him remotely, but you can also launch him locally by giving him access to a remote server. Essentially, the difference is probably not very big. Like you can connect to a database, he will study the tables. If there is no migration, he can suggest how to optimize, how to connect, where what information is located. Especially if, well, the person himself did not create the database and wants to understand it quickly, then, yes, he will connect, explain everything, and you will have an idea of how to work with the data. Very convenient. And with GM, yes, the context, of course, helps a lot. With Claude, he will probably need several. He will go through different directories, collect information, may create diff files, and then combine them. That is, but it will be different. You need to choose what you like more, in my opinion. But it was very convenient when the GMC CLI just came out right from the command line. You could have Claude call another model with its utility. M, it was great. >> Well, yes. And in Kuber, for example, I often still don't navigate well with, well, firstly, I don't get along very well with Kuber, and even less so with remembering every time through cub c, what needs to be written there to get access to a pod, to get some logs from it. Well, yes, I even forget that. And now you can also ask this thing. Well, by the way, in terms of composing commands, warp is probably more convenient now, right, there is such a terminal warp, have you heard of it, which has AI right in your terminal. It probably used to give hints even without AI, but now it works at its maximum. Well, if you're interested, the terminal is called, and, in general, it has everything included. Well, you can ask Claude to do something in Kuber. And it's safe. Well, I think you don't have to be afraid that it will break something in Kuber, because I don't give it permission to do so. That is, I look at the command and if I see that it's adequate, I press "approve". Sometimes I ask him, like, what does this command do? Like, tell me what the parameters are. Well, you can, well, if you're a complete paranoid, you can ask another model or google what this command does if you're not sure yourself. Well. And another very cool, non-obvious thing is resolving conflicts in Git. It always seemed to me that, well, it's unlikely to cope with this, because it's tedious and not always clear to a person. To my surprise, it copes perfectly, and if you have some complex conflicts, resolving them manually, well, it happens. Firstly, it's very uninteresting. You perceive it as some kind of torment. And you have to do it, I don't know, for half an hour, an hour, and here you just tell him: "I don't even have a conflict, you tell him: 'Make a base, there will probably be conflicts, and resolve them yourself.'" And it copes with this perfectly. And I no longer waste time on resolving conflicts, in principle. >> And the question is, so when a conflict arises, then the resolution happens. You can also do it so that different versions of the code are immediately added to one file, and then it will run through it. That is, to combine everything into one. Ah, of course, that's a bad approach, but it's also possible. It's also possible. >> Well, I haven't done that. >> Yes. Scary, scary, scary, but possible, possible to try. And I also, when I see that I have different tasks, they either don't overlap with each other, or, well, they overlap very weakly. I'm in different directories, so, in work >> Work 3. >> Work 3, yes. I, well, I work on one thing, on another, and while I'm thinking about one task, I work on another. And it seems to me that this, in principle, speeds up the development process itself. You can get more done. >> And it happens that again, we have a large zoo of services, there are dozens of them, well, about fifty, I'm afraid to lie, but something like that. Well, our teams have now split up a bit, so it's not so bad anymore. Well, in general, it happens that you need to update a library version everywhere at once. And it turns out to be a very boring and tedious job. Here you just say: "Now change this in all our services." And commit it, because the commit is uniform anyway. You might not even look at what it writes. Push, like, push the changes on this branch to the remote, to Gitlab, when you Git Push and push the current branch, Gitlab returns a link to create a merge request. M, you know, right, that's probably through GitHub, it does the same thing, right. Well. And I don't even look at it, because, well, still, damn, it's not interesting to look at it 30-40 times. And in the end, I just say: "Save all these links." And then at the end, send them to me as a list. And then I manually click through them and create merge requests separately for each service. Probably, it can be optimized here too, if you connect it with Gitlab, but I'm afraid of that. I was afraid to connect it with Jira, but I still decided to do it at some point. >> Yes, I also have a concern related to Git, that it will rewrite something, do something wrong. Well, I'm not talking about commit, I'm talking about, ah, other things. Well. And I'm also very cautious about this. Usually, I, m, well, I only give commit, push, all in a separate branch that I'm working on. >> Well, usually I push myself, I don't even give him that. It's just that when you need to make a lot of changes at once, it's difficult without it. But so I only give him read-only rights. I don't even give him the right to run tests on his own, because, well, not because I'm afraid, but because it's tedious. That is, he will come up with something, I don't know, fix a letter and think: "Okay, I need to restart the test." And you wait for them all to pass. >> Through tasks, it's a bit simpler for me. It only calls the command and then everything is >> No, mine is the same, yes, naturally, it either calls only the command, or it knows how to, like, cleverly run a specific test with a specific case, but I don't need that. That is, sometimes I say that if he has access to tests and he asks, then he can run them when it's not needed. That is, you actively work with him, he fixes one letter and already runs the tests, and you want to change something else. And in general, for these reasons, I give access. >> Well, well, it's better to do this at the end after completing the task. Exactly, >> yes. Ah, well, also, I don't know, maybe not obvious to someone, is working with a profiler. That is, well, many people don't know how to work with a profiler, because the information there is complex, you have to go as far as assembly to understand how something works in some places and understand at a very low level in Go and in general, I don't know, how that same base, which is from that book, probably. Well, I'll show it to you on camera. Well, it's a bit dusty. Only after reading that same base, you will probably understand how to work with a profiler. Well, again, speaking of whether this base is needed or not and whether to readbaum. Well, in general, if you're struggling with the profiler, then probably yes. But it doesn't matter. In general, you can either feed him these profiles, he knows how to, well, he won't clean them, obviously, he will, like, through GoProf, probably look at them, study them as he needs. He can build an interactive map, and get some specific data. In general, he will study and tell you where your service is slowing down and why. It doesn't always work well either, if the service is complex and you have some very tedious slowdowns. You might not find the problem, but you often do. If you don't know how to run it, ask him directly, he will tell you, write down the commands and you can try. He does it well, >> yes? Ah, well, in particular, when you have a profiler running on some pod, he can go to that pod himself, take a snapshot of the profiler from there and work with it. That is, you won't even have to go there manually for that snapshot. So, what else is interesting? it was. Ah, yes, what you're saying, run tests, linters, and not just run them, look, like, I sometimes even lazy to run them separately myself. I just say: "Run and see what broke, and fix it yourself." I don't even delve into why a test isn't working. He figures it out himself, fixes it himself. The main thing is to monitor how he fixes them, because, well, you know, right, all these problems with tests, >> well, not always. I I watch him, so I haven't had problems. But periodically, he, yes, tried to install something, like, how, how, what are we talking about? About the fact that he can sometimes delete tests, or write a test skip so that they pass, or even >> comment them out completely, >> yes, something like that. Sometimes I heard such, I don't know, fables, that, well, probably not fables, people complained that he can adjust the application itself so that the tests pass. That is, he introduces some environment variables, and, for example, if the variable is local, then the application will behave like this, and if it's not local, then like this. And the test passes locally because of this. But I've encountered situations where some complex tests didn't work, and he says: "Let's test the hypothesis, I'll comment out this test." And, well, I say: "Okay, don't comment it out, write a test skip." I say: "Okay, write it and let's continue studying." He thought and thought, simmered and simmered. Ah, and he gives me, in short, well, this test now passes, yes, let's fix all the others. And he suggests writing test skips for all other tests as well, with linters too, he can disable them, delete them, but put a lint. So you need to monitor, even if you launched it, that he does everything himself, you need to check later, what's the result, >> otherwise there will be problems later. That's why I always look at it manually, because, well, often he, yes, he looks at your function, which is, like, a bit too long, and he says: "Well, it's normal that it's long, I'll put a skip, and you tell him: 'Damn, no, I'm trying, as I also, I think, wrote in my instructions, that we only write nolints in extreme cases and even with a comment explaining why we wrote it.'" Sometimes it even goes as far as creating a task to create a task that says how we will fix what we enabled the lint. Well, there are cases, yes, when you need to enable it, >> yes. >> Well, in general, he runs the linter himself, understands what he broke, and fixes it himself. In most cases, he fixes it perfectly. And also, interestingly, what I've encountered recently, sometimes, ah, some functionality needs to be temporarily disabled or a hypothesis needs to be tested, you temporarily do it like this, and then you need to change it somehow differently. And sometimes it's not always obvious. That is, you disabled something or, I don't know, conditionally, you have providers, well, we have services that make payments. There are providers here that make these payments, and you add a temporary provider to make it easier for testers to test. But it turns out to be in five, seven places. And then try to remember where to remove it from so that nothing breaks. And you tell Elmke, like, well, Claude, like, write a special comment everywhere, which will be easy to find with a search, like "to-do" and then the task number, for example, and all the places where something needs to be fixed, so that, in general, it can be rolled back. Plus, he also writes me instructions on why it was done this way, how to cancel it, why to cancel it, and, in general, instructions for cancellation. And this helped me a lot when I was going on vacation. And to disable, what I enabled there, colleagues were supposed to do later. Well. And, by the way, speaking of instructions, that's also a useful thing. I often decided to add instructions to many things in projects. Uh, we always had a folder called "doc" in each service, but usually it just contained an index and it said, like, well, index.md, and everywhere it said, like, "Add documentation here." And usually only the README in the root was filled out, >> yes. Well, now a new problem arises, that there is sometimes too much documentation. And >> we don't have that problem yet, >> I mean, generated. >> No, we just haven't started using it actively yet. >> Ah, yes, yes. But it happens, yes, that when a task is completed, there is a very detailed instruction, and then there are too many of them, and they need to be compressed again. >> It's the flip side of the problem. Ah >> I also try to have more documentation so that I can remember what was done, and then some notes and then somehow compress them, so that it's faster to search. Well, in general, it's not the worst idea to invest, maybe in specific folders, instructions on how, I don't know, how, well, how to work with tests. I've already said that our isolation tests are quite complex. And there are mocks of other services, which are also many. And sometimes you need to pass some information into them, for example, test cases, so that all these mocked services behave in a special way. For example, you write some complex case, you test when these services should fail, and this one should respond twice with a fail and once with success. Well, in general, or, for example, they use it incorrectly, they should return something together. And working with these cases is not always obvious. And you just ask to put an instruction next to the tests on how to edit these cases, why they are needed, and how to work with them. Like this. So, in terms of my work, it seems I already, >> yes, >> we managed to do it crosswise after all. I thought that first I would tell, then you would tell in more detail, but I realized that I would be here, I don't know, telling for all 3 hours, so, well, in principle, yes, I think it turned out better this way, even more interesting, more dynamic. Well, by the way, about working with a fast model through MCP. Have you tried working with it not through MCP, but directly? >> Uh, I tried to use, >> well, through clod, but directly with another model. I didn't use it through code, but through other tools similar to ClD-code. Open code or, uh, plugins for IDEs, ah, it works very fast. I use, ah, Intel ID from Jetbrains and Z. And in Z, you can insert it directly. There will be a window on the side where you can chat. And this is a very cool experience to put clod there too. And you can use cod, and, uh, a Chinese model, which works very fast through Cerebrus. And you can ask questions, for example, remind me what Git commands are needed for this, this, and this. And in literally a second, you get a whole article as an answer. >> Oh, specifically in terms of model replacement for clodcode, do you end up not using it? When I change models in clodcode itself, yes, and I have a subscription to the same Chinese J model from the developer of this model. It works much slower there, but there are no limits. That is, I turn it on in clodcode and get a result similar to GPT-4. >> Why do you do that? To save money or what? >> To save money, to save money. Or at some point I didn't quite like how clod worked with limits. I switched to the Chinese model and there were limits. >> How much does a subscription cost there? >> M, it's from $3 a month to, I think, $200, $90, I think, for a year. This is almost without request limits. It's suitable for those who, for example, well, there are often people, even among my colleagues, who say: "I don't want to pay $100 or $200, I want to pay $20, and get faster limits." So can you just replace the model and use all the capabilities of clodcode in terms of tooling, but with a cheaper model, find something, >> yes, you can. Chinese models offer exactly that. There are several models on the market now, I think, >> well. With active work, how much will a subscription to such a model cost? >> Ah, well, I think, if not testing ideas, but actual work, I think around $30 a month will be a very good model. >> Well, that's significantly better than clod for $20. >> In terms of limits, definitely, in terms of limits, definitely. So you can safely, there are also various promotions, you can try models for free on Open. There is such a server >> local, you can probably even connect it. You can also connect local ones. True, I think their quality is still not as good as >> Depends on your graphics card. >> Well, you probably need a 20 series, probably, to get something reasonable. >> Yes. >> Ah, well, yes, by the way, I looked at the prices for these, something like Nvidia came out, some things. Well, you need to have a lot of video card memory, at least. So, I currently have an Nvidia RTX 5090. It, well, surprisingly, is not enough to run models. Not because the graphics card is weak, but because it still has little memory. >> Not enough memory >> 12, or something like that. And to run a good, heavy model, you need about 50 GB of memory. And it just spills over into RAM, and the video card resource is not used. Well, again, maybe I'm saying something stupid here, but as far as I understand, it works like this. Well, as I see it, the top Chinese models require about 250-300 GB. And here, of course, ordinary consumer solutions, >> yes, ordinary consumer solutions are not suitable. For example, a Mac Studio with 512 GB of RAM, which is also video memory, would be suitable. Ah, well >> who would have thought, >> yes, that Macs would be number one in terms of running models, >> yes. And the question is, why is this needed at all? And there are several cases when it is really needed. These are more likely some medical institutions, or some other closed ones, where a model is needed on cheap hardware, which will not work as fast, but will perform its functions. For this, it will probably be suitable. Otherwise, I think it's better to use, well, provider solutions, if maximum quality is needed. If there is no internet access, of course, it's better to deploy a local model. There are tools, ah, for example, LM Studio. >> Well, yes, yes, I tried to use it. >> And you can connect, again, from Jetbrains, from Jetbrains IDEs, for some reason it didn't work right away. But with Z, it works very well, with VS Code it also works very well. And you can literally write questions and edit. It's slow, but it can be used. Well, in terms of model replacement, I think you'll need to write a post on your Telegram channel later. I'll attach the link to the announcement there and in the description too. >> I think I'll provide them with as much material as possible, what I know, so that people can try it out. I publish the links themselves and everything we talked about on my channel, so it's not a problem. But more such things, in which you are more knowledgeable, specifically, a separate post with instructions on how to replace the model. >> I think it's not super difficult, so a Telegram post will be enough, not a Habr article. That's what I think. I think so too. >> Well, it's convenient for you too. That is, I sometimes write such things as posts so that I can refer back to them later and remember how I actually do it correctly, >> something like that. >> Plus, there are alternatives to Clodcode, and you can also mention them, where you can choose models much more conveniently, that is, you can choose what you will use without leaving it. >> And the provider that gives X15 speed, is it expensive, so you can't connect it to a permanent model? It costs from $50 to $200, if it's a coding plan. Per month >> per month, yes. And this is coding. And the quality will be worse than Opus. >> And if Sonnet? Like, worse than Sonnet. >> Well, like Sonnet, yes. Like Sonnet. It's just that Sonnet is not always suitable for some tasks. That is, it handles bugs better. Well, like, in principle, you can, if someone is satisfied, Sonnet 4.2, and you can completely switch to a faster model, which will be 15 times faster. And you will pay approximately the same. >> Yes, it will be slightly worse, but due to this speed, you can make several iterations and get an acceptable result. >> Worse, do you mean through Opus or through or still worse than Sonnet? >> In some scenarios, Sonnet is still better, but again, due to iteration, you can achieve the desired result faster. >> Well, just, yes. This X15 really hooked me, because it, well, it significantly changes the concept of development too. That is, where you sit and think about how to better formulate thoughts and realize that it will take a long time to think, here it will give you, well, almost instantly, a piece of code.

Immediately. This, I think, strongly changes the concept of working with it. Well, it's hard to say yet, because I haven't tried it like that. I need to try. Ah, yes, they have other models, but they have very expensive subscriptions. That's 1,500, 8,000 dollars and more. That's for teams, rather. And those models, which can be run locally with the right hardware, they sell exactly the capacity and work with open models, which they themselves can run on their powerful hardware. So, what does all this give? I also wanted to summarize in terms of working with agents. What have I noticed myself? Well, firstly, you choose less, because there's much less routine. You probably don't choose because you're programming, but because you're doing some nonsense, like, literally, not even the handle itself, but the parameter of the handle changes, and you need to pass something new from the database. And this, well, sometimes, I don't know, can drag on for several hours of work, because, or even more, because between layers you have to process all this data, convert it, write a converter, make sure that nothing is broken. Moreover, sometimes, if it's not just a through parameter being passed, but something a bit more complex, then it's hard for me to grasp it mentally. That is, I always feel like I'm missing something or forgot something, or something has fallen off in my head in this chain. I have to recheck everything. It's very difficult to keep in mind, but the work itself is ridiculously simple and, well, it's very exhausting. And because of this, I think I've started to get tired of work less, and it bothers me less. Ah, well, and in general, it's more interesting to work, because often you give some small things to these things to handle. By the way, I also started using this JetBrains assistant. Here, Jetbrains AI, they are constantly changing something, adding something, removing something. Uh, I liked it. They integrated it directly. You can select a piece of code and write, like, "rework it this way, that way." And it works a bit faster than Code Llama. They seem to have their own proprietary model too. You can't choose it for this, at least. Well, you just say: "Here's such-and-such a function, let's refactor it, break it down" into several functions, something like that, and it quickly does everything for you. M. And you're doing more high-level things, like designing all of this. Thinking about how to learn, how to make it more beautiful. You simply have time to make beautiful code, and not spend a lot of time on it, >> yes. It adds more creativity, especially when there's no code yet. You can create something of your own and think more about how to do it, than, well, doing it. >> You work faster in the end and often the backlog gets cleared. Well, I haven't seen that often in work. After all, the backlog isn't my headache often. But if we talk about pet projects, I have a couple, well, the same bot, Garvis, in our chat in GoFE Club. Some features that I've been putting off for years, that I wanted to add to it. Well, for example, when you ban a person, they were banned for a fixed period of time, I added a feature to specify, say, one day, one hour, and so on. And I needed to do this for a long time, but not so much that I would take it up and do it somehow. Then, when Code Llama learned to work well with these things, I just quickly implemented it all in an evening, everything I wanted. Well, and I think the code quality will be better, because you have time to write good code. And again, you have an advisor who can advise you on how to do it better and more beautifully. Yes, here, when there's a problem with persistence, maybe during learning, during writing something, it's solved with these tools, and that's great. >> Well, tests are a classic example. And you wrote them well, and now it's just like I've heard many even proponents of testing in various podcasts, who always paid a lot of attention to tests, they've reached the point where they don't even read or look at tests anymore. Well, maybe they'll glance at it, that everything seems plausible. I sometimes sin by not reading all the cases that it wrote. That is, it wrote me like 20 cases, some main corner cases, happy path, fail test. I don't read all of them, because, well, it seems plausible from the description, I won't delve into how it does it. And if before you could say that I don't have time to write tests, I need to do work tasks, and we even argued about what's faster, with tests or without tests, now there's no point in arguing, because, well, you don't write these tests. Make the model write them and that's it. If they break, you won't fix them either. >> I think there might be nuances with tests when it comes to very specific tasks, when there are complex calculations, when you need to dive very deeply into the subject, then the model won't cope as well as with the task itself at the moment. But in all other cases, yes. >> Firstly, you need to follow actively. They are developing at breakneck speed now, both models and tools, and as I also wrote in the announcement, why this podcast episode even made sense, damn, although we've already discussed it twice, damn, because what was six months ago has already become something prehistoric. I don't know how to phrase it correctly. In short, it's already morally outdated in places. And the concept of working with models is also changing significantly. And not only with models, but also with tools. If you look at how often Code Llama is updated, it literally gets an update every day. And I've already made it a practice that I go in in the morning and before I launch Code Llama, I type update, because a new version is constantly coming out. And it seems to have learned to update in the background, it didn't used to, but I still type update out of habit every time. And often, well, it can be some nonsense that you don't follow. Sometimes you look closely at these changelogs and think, "Ah, there's also this parameter." Like, they fixed this parameter. You think, ah, there was also this launch parameter, you could have done it this way too. But often very important things come out, like skills, agents, and so on. Sometimes so much comes out one after another, and you don't even have time to try it all, but at least you need to keep in mind that you need to try this and that. Some things cannot be missed. >> Yes. It was like that with plugins, I didn't even notice when they came out. >> Ah, plugins for Code Llama. >> Plugins, yes. Essentially, it's a set of instructions that can do it for you, and not only add skills. >> They even have a marketplace now. >> Yes? A marketplace with plugins. You can connect plugin stores and, in principle, choose ready-made things. But you need to be careful here, because different people use these plugins for different languages, for different purposes. And here you need to choose very, well, in detail, approach very cautiously, yes. But the toolset is developing very rapidly. I think one of the pieces of advice you can give is to follow updates, read the Twitter of developers and models and tools, subscribe to updates on GitHub, because there are open-source tools and you can follow them too. >> Well, especially since if we talk about Anthropic and their model, their tools, they present all this very coolly. They constantly have videos on YouTube where they, well, sometimes short ones, like a minute, like 2 minutes, where they just show some new feature and how it works, sometimes longer ones where they discuss it. Plus, their documentation is always very good, their blog is good, and they even often write guides on how to use it. Several times they've released some guide on how they themselves use Code Llama, how to use it better, there were some other tips. I'm already confused by their guides, they're all good, but I forgot which ones and for what. They came out, I saved them, I haven't read them all yet, but they try very hard. That is, it's not like you're digging through some tweets, trying to figure out where what is. Well, they present it all to you well. And literally a couple of weeks ago, they released a course that you can also take. They talk about internal tools. >> Well, especially since they make their own courses. People often ask, like, where's the course for Python, damn? They themselves make these courses. >> I think we'll add a link too. Well, yes, yes. Hugging Face also makes its own courses. I think something came out from Microsoft, if I'm not mistaken, also on MCP, some courses came out from various companies. All, all free. >> Yes. Skills are already cooler than MCP, as they say. Ah, yes, there are courses. Ah, >> well, practice and try all the advice you hear, if they have arguments, like you can validate them. Like, a person says, "Do this and it will be like that." And you can see if it became like that or not. Try everything they advise. That is, they tell you: "Try the planning mode, try the planning mode." Even if it seems like some nonsense. And even if you don't like it, you'll just miss it. But often it turns out that you like it and you start using it actively. Or at least you understand the concept. Often it happens that, well, this doesn't just apply to AI, but in general, in life and in development. You might perceive some concept as a black box, and until you try it, you won't understand what it is at all. You can't even formulate the question until you try it. And even if you don't use it later, you at least roughly understand how it works, and you start to understand some related things. In general, you can and should definitely try. And practice, if it seems like I often had colleagues who I told them with such fanaticism, like, "Take Code Llama, because it helps with this, it helps with that." The person tried it and wrote to me, like, "Damn, I wasted $100, it's nonsense." He couldn't even complete a task. But here you need to practice, because sometimes, I think, people expect that, like, "I'll just send him the whole spec, and he'll give me everything ready to release." Well, sometimes, yes. Some simple things, complex ones, not always, but it's a supporting tool. It doesn't necessarily do everything for you. Sometimes it takes part of the work off your hands. Either it saves you time, or energy. Sometimes both. >> Yes, Code Llama might not be for everyone. My friend wanted to use it. As a result, due to connection problems, he uses Cursor. Cursor had no problems. And he got so hooked on it that for the last six months he has practically not written code manually. That is, everything works so well, and he got so used to the tool, and it kept getting better, that he's progressing, he sees what works, he immerses himself, his learning process is very active through work. Well, some of my colleagues who were Cursor adherents, but for me, for example, one colleague, hello, Vadim, he will probably watch this episode. He used Cursor very actively, and he even said that when he was updating, well, he was sent a new Mac, a work one, he didn't stop using JetBrains, although he also actively worked in GoLand. Because he had already switched so much to Cursor. That is, at first we all got used to it, and he also said that he uses it to write code, and he still makes some changes in JetBrains IDE, like, both. That is, it wasn't a complete switch. But at some point, like, a significant checkpoint was that he completely switched from JetBrains to Cursor. But I returned from vacation, we recently had a call together, we did some pair coding, and I noticed that he switched back to JetBrains, and is now actively using Code Llama. So, in principle, the person, you could say, was, I don't know, a good word, an adherent. I can call myself an adherent of Code Llama, so I can probably call him an adherent of Cursor. But even he eventually switched to Code Llama and JetBrains. So, well, you have to try everything, both. Some people use two IDEs, because, well, there's a subscription, something is better in one, something is better in the other. In GoLand, for example, file changes are shown very conveniently, it's, well, at least visible. And working with databases is also convenient, or resolving conflicts. I think one of the best tools is JetBrains. Visually, it looks very pleasant. Someone got used to it and uses it. I would also advise watching as often as possible how other people use these agents and so on, because, as I already said, I don't know if I said it or not, that often in micro-moments it's not obvious to you that you're doing something unusual, and it's not obvious to a colleague to ask about it. Like, I don't know, банально, you say that, for example, I manually accept, for example, I complete a task, and I see that it's going in the wrong direction, I stop it and say, "Do it differently." And someone stops you even earlier and says, "Damn, you stop it, but I, for example, don't stop it. I tell it, 'Finish the task and at the end I look at all the changes.'" Like, the concept that you can interrupt it and interfere with its work might not be obvious to someone. I'm just recalling what I could. There are simply such micro-moments that you don't even realize that you can teach someone about them. It seems obvious to you. And the fact that, I don't know, I, for example, communicate with it in Russian, and some are surprised, like, "Can you also speak Russian?" Does it understand Russian? Well, it understands Russian perfectly, and yes, it uses a bit more tokens, but it's easier for me to formulate a thought in Russian and faster than in English. Some people don't realize this. Some are afraid to communicate with it informally, that is, they don't write "hello, how are you" because they're afraid it will work worse because of it. And there are many such nuances. And to understand them, you need to ask, like, банально, if you know, if you know that you have a colleague who actively uses it, maybe not better than you, but just as actively, like, ask to have a call with him and let him show how he completes a task, like, complete a task together on pair coding. And during the process, you'll realize that, surprisingly, people work with agents differently. And I just, when I heard stories, that someone says, "I do it this way, I do it that way." You think, damn, it turns out my way isn't that obvious and it's not the only correct one, not the only right one. >> There will also be some optimal ones, some not. This can only be understood through trial and error. And in terms of language, yes, I also often write simple things in Russian, write in English, and sometimes it's easier for me to express my thoughts in two languages, and it will understand, and possibly do even better than if I described it in one. That is, it can switch context from one to another in one sentence and pick up the meaning and understand me. Well, yes, sometimes I also send it things in different languages. And it would seem that these tips are general, that is, they are often said, like, you're learning to program, ask colleagues to review, look, but here it's even more critical and important, because, well, this is still a tool, but, well, like, an agent is a tool, but it's an unusual tool. Firstly, it's a tool that appeared recently. That is, if we have a hammer, we all know what a hammer is, how to use it correctly. And there are a million instructions on how to use a hammer, and they are all accurate. And a hammer will remain a hammer for 200 years. But here, firstly, it's a tool that never existed. And no one knows how to hold it, which side to hold it from. And even the creators of this tool don't always understand how best to use it, it's a very complex tool. And what's more, it's a tool like a hammer that is constantly changing. Sometimes it's round, sometimes square, sometimes something else. And therefore, it's impossible to write a book. Well, they do write books, of course. I recently saw in Japan, they have whole shelves on how to use ChatGPT. But that's a bad idea, because they are constantly changing. Literally, by the time you write a book, their usage concept will completely turn around and be completely different. If the author could advise, for example, that you are using Code Llama and ask it to refactor, oh, to review your code with GPT through Codex, because it's smarter, but by the time you write the book, Code Llama has already become smarter. And what's the point of this caveat? And therefore, you need to actively follow all of this and see how others use it. And here it's vital, yes, and everything can change literally in a couple of months, new usage methods can appear that we don't even know about. >> Well, yes, if you follow podcasts like this, streams. I, for example, lately have the feeling that I've been listening to podcasts mainly about AI, because you watch a podcast, and there's not a word about AI, you think, "What, I already know that, like, how to write in Go, or something else." But how to use AI, that's what's new now, for example, Opus. I'm interested in what people say about it, some guys I've been listening to for a long time. How, in their opinion, has it become better or worse? Something like that. That's all I have for advice. Do you have anything else to add? Yes, we probably covered it. So, ah, >> well, we also wanted to touch upon Chinese models, because for me, I always diligently skipped them, because I thought, well, the big three, Gemini, GPT, and Claude, that's enough. And that, well, some Chinese models are probably all worse. But you said that it turns out they're doing well and it's interesting. Yes, I'll probably share my vision now, and, ah, how I see these models being used in work and why they are worth paying attention to. Ah, well, firstly, they are very close in quality to top models. Ah, this is important, you can save money on this, especially if there are a lot of tasks, or a large team and a single subscription, then it's profitable. Ah, and these models are also used in some kind of closed loop. I'm sure, in large companies that don't have access to the big three, so to speak, they use a local model to prevent data leakage. Ah, and as far as I know, in large tech companies, they use Quanzhi DeepSeek, and it, of course, is inferior in writing quality, but you need to know how to use it effectively to get, well, the best result, to save your time and, well, and resources. This JLM is from some other company, that is, DeepSeek is, >> yes, another one, there are currently about four players, maybe more. In the Chinese market, there is Ernie, there is DeepSeek, there is, again, Zhipu AI, there is Minimax. These are all different Chinese companies that release different models. They are very close in quality, but one is better in some aspects, another is better in others. Ah, and if you look at benchmarks, they are better in some aspects than the top models from the big three. Ah, >> well, you are actively following them now. Firstly, about this Ernie, I remember you told me something about it, when we >> Yes, this is just a Gemini that looks like Set. >> Ah, this is, this is Zhipu AI making it, no. Yes, yes, yes. >> Ah, okay. >> Ah, and can you give a general breakdown of them? That is, I can easily tell you about the top models, about the top manufacturers and their models, like GPT, Gemini, and Claude. Can you give the same situation, the same breakdown for Chinese models? >> Ah, >> who is who? Why use whom, how to choose whom better? I think you need to try first. Ah, and whichever suits you, use that one. They are very similar in results. Again, they think a little differently. And specifically in your use case, one model or another will be more advantageous. Again, everything changes so quickly here that today DeepSeek is better, tomorrow Ernie, now Minimax. You can't guess here. Again, you need to look at the providers. There is a very fast provider, Cerebras. They provide some of these models, but some in a coding plan that is available to the ordinary user, and some are not. That is, someone can pay $50 to use a top Chinese model. Someone can pay $1500 to use another top Chinese model, which is essentially not much different, but for one use case, that one will be more profitable. Here, again, I think you need to look at what the manufacturers indicate on their websites, what plans are available, because there are profitable plans, and there are where payment is only for tokens. And here it's your choice how you want to manage it. If someone uses it a little, I think it makes sense to try them all, spend a little money, choose the best for yourself, and, well, stick with it. Ah, again, there is >> do you recommend using them through specific providers or through the developers themselves? >> I think it's better to try through Hugging Face or OpenRouter first, to try, see if you like it or not, see what offers are available right now. >> Router, what is that? >> It's also a provider through which models are provided and >> like any model, right? >> Ah, it's mainly payment by tokens. Ah, >> right, and sometimes there are free applications, there are, well, hidden models. Ah, now, when GPT-5 came out, it was called by a different name, and you could use the top GPT via API for free for 2 weeks. So, I think, OpenAI was testing something, seeing how people use it, and provided such a model for free. So you need to go to these sites, see what applications are available, poke around in different ones, and choose, choose, choose what will be most suitable for you. >> Well, I understand, yes. It's clear that you need to try all of this. Ah, is there maybe a service that provides access to different Chinese models, at least, through a single subscription? >> Mmm, about Chinese, let's see, let's see. Ah, well, with Cursor, they're not Chinese, but recently, I remember, they might have been slandered, but they said that maybe Chinese models are under the hood. And here it's unclear, directly, that through one subscription, all at once. Ah, probably difficult. Ah, subscriptions are cheap, I think, from $3 to $10 they start. And, well, >> well, that's a lot. I already have a bunch of subscriptions. And to subscribe to that many Chinese ones too. Yes, yes, yes. >> Well, it's difficult. Ah, and in terms of their tooling, how is it? That is, like Code Llama, it has what? An interface, well, obviously, Code Llama Code is a tool, and the interface in Claude's chat is also awesome. Well, банально, that, I don't know, when you copy a large piece of text to it, surprisingly, no one else has this yet, it pulls it up as a separate file. That is, if you paste Ctrl C, Ctrl some article into it, it won't paste the whole thing into the chat. It will show it as a separate artifact. You can click on it, look at it. If it's in parentheses, I think it will be like

It's like a file was inserted. >> You're talking about clodcode, >> right? >> Ah, ah, I see. Well yes, in clodcode it's like that. And in chat, if you copy the web interface into chat, it pulls it up like a file, as if you just dragged and dropped a file. >> There. Well, I haven't seen such a feature anywhere else yet. And it has many such features. Like in terms of toolkits for Chinese models, for example? Ah, some of them work with images, some don't. You need to pay attention to that too, I think. Ah, mm, and if we work through clodcode and this model provider supports it, recommends it, then, well, you can use clodcode itself as the wrapper, or you can change the model provider to a Chinese one and, essentially, use this tool and it will work similarly. And in what cases does it make sense at all? That is, for example, from what you've mentioned, when you need to save money, when there's no access, well, I don't know, when there's no access to specific cloud models, and also when you need speed. For example, there's a provider that gives X15 speed. Are there any other cases where, if you have enough money, you don't want to save on anything, and the speed suits you, is there a point in switching to something Chinese? For development, nothing comes to mind. For other scenarios, when you need to generate images, content, they will be profitable. And if it's some kind of product related to generating something, then you also need to pay attention. And now, >> well, it turns out to be cheaper, >> It turns out to be cheaper, yes, it turns out to be cheaper. >> Are there any aspects where they are better, not just cheaper? >> It's constantly changing and they are >> better than our top three. >> Ah, yes, yes. There are benchmarks and some mathematical exams. These Chinese models are now passing them a few percent better, by tenths, hundredths. And considering their resources compared to the American top three, >> especially Google, >> yes, especially Google, the breakthrough they've made is impressive. So, if they have >> about 3.0 or >> about Chinese models, that their level is very good for the resources they have. Yes, if they have access to more powerful hardware, then progress will be even faster. >> The question is whether it will appear and when. >> Okay, from what we've managed to discuss, well, about, probably, about news, about the future. We don't have much time left. We've been sitting here for almost 3 hours. >> Wow. Well, more or less about the news, recently Gemini 3.0 was released. Well, yes, I said that at the beginning, I just don't forget that you said before the recording, that after the recording. I said on the recording, yes, that I used Gemini in Japan, >> like GPT >> or >> I said that while I was in Japan, Gemini was released, I switched to it, >> but now Gemini is incomparably better even than version 2.5, it's like they have not 2.5 and 3.0, but like several versions. They should have, I don't know, called it 5.0 right away, probably, because the gap is very large. And even the gap compared to ChatGPT has become quite large. Well, for example, just to understand, it has become both smarter and better at outputting. So, it often happened, for example, that what I mean is, for some time, when ChatGPT hadn't yet switched to this switcher, like thinking and fast model, you chose a specific model. Version 4.0, it was very smart and cool. That was the reasoning model, 4. Not 4.0, but 4O. Uh, but it formulated poorly, meaning it would give you a huge block of text, but you wouldn't want to read it. And then you would switch to another model. For example, version 4.1 was one of the best. It was eventually removed everywhere because it seemed too expensive. Mm, it wrote everything as human-like as possible. That was, I think, the very model that was first said to have passed the Turing test. I just did it this way: first I called the model to find all the information, Google it, explain it, analyze it, and then I switched to 41 and said, "And now present it in a way that's easy to read. Add some nice emojis there." >> And surprisingly, Gemini now does both well. It can do smart analysis, and it can also present it to you correctly. This was evident, for example, in the fact that I finally started using it, the models, because they can be used fully and conveniently, for example, for language training. So, if before I simply, naturally, consulted on all issues related to language learning, for example, Japanese, with all models, it was not a problem. But they faltered where you want to practice, for example, verb conjugations, both in English and Japanese. There's this t-form in Japanese and several groups of verbs. And, well, you just have to memorize it and form neural connections and practice. And, well, I didn't always find convenient exercises. And I just told Gemini, like, I'm going to train the t-form, and you give me four examples at a time, I'll write them to you, you say right or wrong. Well, in short, in short, I explained to him, mm, how he should check, how he should output, and what to do. And it does it much better than ChatGPT now. For example, there are also, who might be familiar with Japanese, probably not many. There are ru verbs, so-called verbs ending in ru. And there's a problem. There are also u verbs, which end in u. And there's a problem that the list of u verbs includes those that end in ru. And there's also a separate ru. And, I don't know how clearly I explained. That is, as if, well, again, there are verbs, there are u verbs that end in u and in ru. But the problem is that the list of verbs includes verbs that also end in ru, and they are not in this group, they are in this group. And why am I saying this? Because ChatGPT and Claude failed this task, because when I asked them to briefly remind me of the rules, the grammar, they didn't mention this nuance at all. And they are more superficial, they just give you some tabular advice and that's it. And yes, go ahead and solve the exercises. And Gemini acts as a full-fledged teacher. It tells you, first of all, in a more understandable language, a cheerful one, where it adds some nice emojis where appropriate. And it specifically mentioned this point to me too, that like, given that there's this trap here, that there are verbs that might turn out to be from the wrong group, even though it seems like they are from the right one, and how to distinguish them, most importantly, there are ways to understand which group it actually belongs to. >> And what's also interesting is that it sorted them not by groups, like first, second, third group, but because the first group is more complex, the second is simpler, and it started with the second because it's simpler. And these nuances are why it's so much better. You understand better and ultimately solve these exercises better. Well, plus it also forms the exercises themselves better and checks them better. >> Yes, in terms of code, they have amazing Google Studio. And now they've released their code editor Antigravity, where you can also try all these models. >> Have you tried it, by the way? >> Ah, I've tried it a little bit. Well, quickly, it works great. It did simple tasks. I liked it. I've put it aside for now. >> I tried their Google thing, remember when there was a trend when everyone released tools that would solve something on your GitHub? That is, you communicate in chat, and they work through the cloud on your GitHub. >> Yes. Anthropic also has something like that. No, everyone has it now. And when Google released something like that, I tried it, it did something wrong. I complained about it, and it got offended and stopped responding altogether after that. I haven't worked with it since. >> In principle, you can do the same when there's a separate server, and there's code or a similar utility, you can access it literally through your phone, and program, and it will push to GitHub. You will see its changes. So, if someone wants to save money or implement something similar themselves, they can try it. >> Have you had a chance to work with Gemini 3.0 yet? >> 3.0 Pro? >> Ah, yes, yes, I've been running it, but I haven't noticed any significant benefits for myself. >> So it works well, but I'm more used to, ah, and clodcode, and I like how the latest Opus works. I've set it up, again, with a Chinese fast model. I like how quickly code is written. >> So I'm not ready to switch yet. I might add it to my workflow somehow, but not yet. >> Well, that's in terms of code. I haven't tested it in terms of code yet, and probably won't, because, well, the tooling matters. Even if the model is stronger, I still like fiddling with changing models. Like, maybe it's interesting to try connecting Gemini, right? So I can connect Gemini to clodcode to work directly on MCP. >> There. But the tooling still matters, so I'm unlikely to try Gemini in terms of code, but in terms of everyday use, it's currently the king. And regarding Opus 4 with 2. I haven't had a chance to try it properly yet. What can you say about it? I just got back from vacation, so what can I say about it? Well, I think if before everyone used S 45 to not exhaust the limits of OPC 41, then now you can safely use only Opus 45. It has become cheaper, it has become better. And again, it works faster, better, higher quality, and it consumes fewer tokens because we get results faster. There are no longer any separate limits specifically for the most powerful model. Now they've made weekly limits for me on Sanaci, and in principle, they've solved the problem, I think, with the limits they set at the end of summer, that now it's not a reset every 5 hours, but weekly, and weekly is enough for most. If before, I think, 5% was removed, now it's in terms of. In terms of how much smarter the model has become? Well, that is, if we compare Snet 4.5 with Opus 4.5, how much can it find bugs and fix some complex bugs, but I haven't noticed a difference between four and one. That's the thing. >> Uh, >> so >> well, I think Snet has become much better than OPCus 4.1, but 45, I don't know, >> yes, 45 is good, but Opus 45, it's unclear. It seems faster, better, well, similar. Well, if 4 with 5 was better than Opus 4.1, then probably Snet is still better than Opus 4 with 5 then, >> yes. >> But generally, I've noticed that Snet 4 with 5 has become better because it follows instructions better. That is, I even wrote a post and wrote what it's better at, and then I realized that, damn, everything I listed is written in GLDMD. Well, almost everything. The point wasn't that it became better at doing it, but, like, it started writing better comments to the code, because it started reading more attentively what you wrote in the instructions. And I managed, how much I just recently started working, because I just installed Bael at some point, worked a little with Opus 4 with, and it seems like it has regressed a bit again compared to Sonnet, that it started following instructions worse. That is, it started writing dumber comments more often, for example, even though it's clearly written that they shouldn't be written. Well, things like that. I don't know, maybe I'm wrong for now, I'll have to work with it, but it seems like Snet still follows instructions better, even than the latest Opus. But it has probably become smarter. I can't say for sure yet. I may not have had such complex tasks to test it. But on the tasks where I tested it, it works well. But Snet also handled them well, it seems, >> yes. >> It writes good plans. Well, Snet also wrote good plans. Here, I think there are some differences, maybe in style. And some people will like one more, some the other. Uh. I haven't noticed any big obvious changes either. They are both very good models. >> GPT 5.2 was released the other day. Have you read anything about it, tried it? >> I read, it seems someone tried it, but didn't notice any changes. I saw the latest comments. >> Well, I thought so. It seemed to me that way. And >> I also bought a subscription to Grok at some point and tried it, and I didn't like it. That is, it was highly praised. Perhaps for frontend or some web applications it's better in some ways, because, well, benchmarks show, there's some statistics, but I don't like it. That is, I think maybe we'll see something further. And for ordinary tasks, Grok is fine. It, uh, parses actual news and can provide something interesting. >> Well, I think we'll start wrapping up gradually. You also said you wanted to talk about the future. >> Yes, I think it will be very interesting when we see how new hardware will work with all these models. I think development will only get faster. Ah, there are now projects that decentralize hardware work. Pavel Durov has such a project. There's also the Race project. And it's interesting how we will be able to interact with powerful hardware that will be literally with users, and in a decentralized way get many, many, many tokens and use them somehow. >> I think it will be interesting. I think open models will push progress forward. Well, again, using you as an example, you got into it when all this was developing, how do you think it, well, many are afraid that new so-called newcomers will, well, get into it worse, because somewhere they are spoiled, somewhere it deceives them. I don't know all these points why it will worsen these people, but partly because you will become lazy and write less by hand and, like, as the model made it for you, so you, like, a wipe-coder who just tries to wipe-code it, and then writes on Twitter like: "Oh, it deleted my files, how do I get them back?" That is, the person didn't know that, for example, Git exists. Ah, well, I have an example where a person learned the basics, they can write simple code, but with the help of AI, they can perform the work of a strong developer, specifically in terms of work satisfaction, they have been working for a long time and they use Cursor a lot again, and they understand everything, so there are no problems. I think if there were no tools, then, of course, labor productivity would decrease, but they would quickly compensate for it through skills. Well, if it so happens that no AI tools can be used. >> That is, and in this I see no problem. On the contrary, I see that you can learn faster, find information faster, test something faster, make prototypes, your own projects. Again, what you didn't have time for before, or were too lazy to do, you can now implement very quickly, see everything. >> And plus, AI itself can offer some utilities that can be used in work. Again, new knowledge at the moment of work. The main thing is not to press the "Approve All" button and close your eyes. >> How much harder would it have been for you to get into it if all of this didn't exist? >> When I got into it, none of this existed. And gradually, gradually, I saw my progress grow simply because I gained more knowledge in the field. >> And if it didn't exist, the curve would have changed, it would have been more linear. Now you can immediately try something new. It's really impressive. >> Well, okay, then it seems we've covered all the points we wanted. We've stayed within the approximate timings I expected. So, thank you all for watching. Uh, yes, I also thought maybe I should say this at the beginning, but okay, for those who watched until the end. Regarding everything we've discussed, how to use it all, it's good to listen, but it's better to see. You've heard it at a high level and overview in the podcast, but how to use it in practice, it's easier to watch. And maybe we'll do a stream with Sava where we'll show it in practice. But he said he doesn't like streams and showing his computer, so I'll have to persuade him and for my persuasion to pay off somehow. Ah, well, let's do this. If this episode is well-received, well, the channel is new, I'm currently uploading episodes separately, not on the main channel, so if it gets at least 500 likes, then we can talk about it. And if, I don't know, after this episode, I have 500 subscribers now, if there are, I don't know, at least 200 more, say 700, two conditions: 500 likes and 700 subscribers, then we can do a stream. If not, we'll consider that the episode wasn't that well-received [laughter] and the stream isn't that necessary. >> Well, in general, thank you for agreeing to come and chat about all this. Thank you to your support group. [laughter] >> Thank you for inviting me. >> And thank you to Fyodor for the support too. I wasn't visible. I have some more here too. This time I'll show him to the camera. I was sitting like this [music] sideways. That's all, thank you everyone. Goodbye. [music] เฮ