Transcription
Friends, hello everyone. It's Sergey. And today I want to show you such an interesting topic as creating your own agents based on Gemini. In one of my past videos, I showed how to create an agent based on ChatGPT that would generate images. And today, as you can see on the screen, there are things that can be done with Gemini. They are very visual, very promising. And I will show you today how to do it, how I came to this at all. Back at the end of last year, I stumbled upon one of the YouTube channels where it was shown how this is done. It inspired me, and I want to say a big [music] thank you to the author of this account, Setka Project. In one of these videos, I think, here it is, a free program for creating photos. We create an avatar with Google AI Studio. [music] This was about 2 months ago, where the author showed how you can create your own avatars, so that the face doesn't change with different [music] poses and so on. And he showed how he showed how you can create such a bot or application for yourself. And it inspired me a lot, and I decided to build my own. And today I want to show you how you can do it. In principle, it's quite simple, without any knowledge of coding, programming, and so on. Today it's called a very fashionable word, like "vibe coding," right? Well, like, uh, you program according to your mood, you sit, like, drinking tea or coffee in front of the computer, tell it what to do, and it does it for you. In principle, it [music] is. Ah, well, I can say right away that I've already tinkered a bit with these schemes, how to build all this. And in reality, it's not that simple. Sometimes, as many people say, right, it's easy to make simple applications, it's convenient, fast, and good, but if you're going to load, let's say, create something massive, then here, one way or another, [music] you need at least basic programming knowledge and, in general, how to understand what you're doing or be able to read code a little, what's written there. Well, okay. I created this thing. It has four such [music] tabs, pictures, I don't know what to call them. Cards, where each card performs, let's say, its own function. Here's the first thing, I called it photo generation. Here's just a simple, let's say, scheme. Uh, for example, first you choose the format in which I want to create the image. And here on the image, you just need to write something that you want to create. For example, I'll just write, for example, dog and cat, for example. And just click the generate button. And the work starts its, so to speak, [music] wheels. And here's a dog and a cat, please, here's a photo ready. Unfortunately, it made it for me not in YouTube format, in a wide format, but I can assume that because this application was created when I didn't have a Pro subscription to Google AI Studio, meaning I had the free version. And most likely, that's why it, so to speak, takes those old data that were based on Gemini 1.5 or Gemini 2, I think. Next, there's something like avatar creation, for example. Here. And it's already gone through all this. [music] Here we just consider, for example, again, choose the format, choose the gender, like female, male, androgynous, let's take male. Age young, let's say 40-55. Hair burning brunette. Wow. Well, let's take gray, silvery, and [music] studio style. Let's take, for example, let's take a fruity style. Here. Why not? Angle [music] close-up. Let's do full body, natural light. For example, cinematic lighting. Here we won't write anything and [music] just click create. Let's see. It will probably create a square again, because the application was created last year, at the end of last year. Well, look, what a guy turned out, right? Yes. And here, of course, you can, I also made a download button. So you can [music] download the image, for example, click, it selects your folder, it downloads. And here, please, here's the image [music] quality is also quite good. [laughter] The guy turned out a bit strange. [laughter] But oh well. So, then here I made a tab called "photo editing." That is, just like on that Setka Project channel. The author showed how you can do such things. Here [music] we insert a photo, for example, with an object. Well, let's take the same guy. And here we'll insert, for example, a landscape photo. Here I, well, let's take a bridge. I have some photorealism here. Well, since we've already started messing around here, let's take, for example, a plastic style. And here we'll write what. So, let me write in German. My keyboard is in German. Man flying over a bridge. Well, roughly speaking, over the bridge, if you speak German. Let's see what happens. Here are the things you can do. Look, see? It cut out the guy. In short, the guy is flying over the bridge. [laughter] Well, and the style is plastic, as we chose. See, everything is made of plastic. Well, in general, such a thing, yes. Here again, I made a button so that you can return to the menu. And the fourth button is to create a cover for [music] a video. Here we remove this reference. Here we, for example, upload, if we have our own photo, for example, we can upload it. Let's see what we have here. Uh, well, let's upload this photo, for example, someone else generated it for a video. We'll write the video title, for example, viral post. And we can write on the cover what we want to change. For example, uh, we are, for example, in a forest, let's say, or let's say, let's do something crazy. Pines are growing on the road. Uh, pines are growing on the road, in short. Well, I wrote it with mistakes, but it's okay. Visual elements can be added, for example, such. I put them in the menu so you can add red arrows, explosive effects [music] some stickers. Let's choose explosive light, bright, clickable style and click create. Generation error. It doesn't want to create for some reason. Let's try again. It doesn't want to create. Okay, we won't bother. I think the principle was clear to you. I'll show you another application quickly that I created, and then I'll show you the simple steps on how to make such an application yourself. So, here we have, genius, don't snore. So, where do we have it created? Such an application, by the way, I also created it on the fly, so to speak, just wrote when I was at work, when I had free time. And I created this thing for myself, that it generates images. This is already based on Pro, I did it, it generates, you know, an image in the format I want. That is, the idea was that [music] I just create several main categories, for example, architecture, technology, plant world, and so on. And in each of these categories, there are its own subcategories. Then these subcategories can be mixed. All styles are plugged in there, and in general, the fifth, the tenth. And in the end, you get a ready-made image that you wanted to get. This is done so that, uh, you don't sit there, don't think, what image should I create, [music] how to write a prompt correctly, or copy prompts from someone, or constantly be searching for these prompts. This program does everything clearly itself. For example, I'll just show you an example. It's still a bit raw. There are some unfinished things in it, but nevertheless, in principle, it performs its function. Let's choose, for example, architecture. And you see, I made it so that it's divided into regions, into, in general, Asia, then America, Russia, and so on. Let's choose, for example, the region Australia. And here there are already, so to speak, some of their main objects already entered into this program. For example, let's take the Sydney [music] Opera House. And you can click continue, or you can click mix, for example, and choose from another category, uh, for example, flowers, right, and continue. Or you can click mix again [music] and choose from a third category. That is, you can choose up to three objects, three categories, and so on. Then we click continue and go to the next step. And here, by the way, everything is written step by step, how, well, where we are. Here I have [music] two tabs: angle, for example, portrait, top view, macro photography, plan view, view, panoramic view, full body. You can go back. You can go in here. That is, in scenarios, for example, here it's already written, cut out on a leaf, frozen in a crystal, in a fruit slice, on the surface of a coin, inside a bottle, diorama in a cube, for example, let's take cut out on a leaf and go further. That is, you can choose both categories if needed. Well, let's take. No, we won't take anything here. Click confirm and go to the next step, number five. Also, there is a function here that we can go back at any moment. Confirm. We go to the next, so to speak, block. We have materials here. That is, from what material the object will be or should be made. That is, wood, stone, leather, foam rubber, denim, ice, metal. We can choose metal, or we can not choose a material. And the program will, let's say, adjust what's needed itself. Let's not choose it. We'll just choose a style. Style, let's take a pencil sketch, for example. We go, go to the next step. Here we just choose what format we need. For example, 16x9. Here. And here, if we want to add something from ourselves, for example, you can add to the window, write some elements in text, for example, I don't know, lighting. [music] Yes, let's not write anything, let's just start the synthesis and see. And at the output, we will get a ready-made photo [music] and there will be a prompt for it. And then we can, in principle, take this prompt, for example, and put it into any image generator, and it will generate, well, something similar. Well, and here, please, look, it wrote a full prompt. It also wrote in Russian what it's about. That is, we chose the Sydney Theater from the architecture category, then chose flowers in another category, a rose, right, then we chose that it would be done with a pencil and in styles, that it would be like on a sheet. And we got this photo. We can then download it. Let's open it and immediately look at it up close, what we got. [music] Well, look, a beautiful, beautiful drawing turned out, as if everything is drawn on a piece of paper. Here, please, you can color it or do whatever you want. Here. And, for example, there is a download prompt button. Let's go, for example, to Google Gemini. Let's choose banana here. Let's choose the Pro version, for example, let's put this prompt here, [music] click generate and wait. Ta-da! What did we get? Look, [music] here it turned out, in principle, the same thing, but here it even turned out more interesting. The main elements that we wrote in the prompt are preserved. And each generator, naturally, [music] interprets it a little differently. Well, and now let me show you just how to do this. So, we return to the beginning. We are on the Google AI Studio tab. And here there is such a thing. So, so, so it's called build. Let's go back. Here's the main menu. There's a menu on the left side called build or building, in short, [music] in English, I don't know how to say it correctly. Click the button here. We can click here, or we can [music] just click the plus button and where it says new application. Here [music] a text window appears. Here we can speak with a microphone or write text. Mostly I wrote text at work. Today I'll try the microphone. >> So, Penny, stop snoring. Let's create, for example, a similar simple generator that will generate, for example, an image. That is, in essence, any idea can be brought to life here. Not only related to generating some prompts, images. You can do it for, I don't know, recipes, or diagrams, or write music, or anything. Well, let me just show you the principle of how it's done. You are an application creation master, and I want you to create an application that will generate an image based on my requests. And at the end, there should be an image. Also, there should be a prompt written under it that I want to have. And, accordingly, a translation [music] into Russian. Let's start with the main category. In the first window, we will have two tabs. One tab will be objects, and the other tab will be styles. In short, I told it something here, uh, and it will start doing it now. But, let me tell you a secret, I don't do it that way. I just [music] open Google Gemini. Let's open this, we don't need this. Here I tell it, buddy, I want to build a new application in Google AI Studio. [music] Imagine you are a specialist in building such applications using vibe coding. We want to create an application that will generate pictures. [music] And I will also need a prompt in English at the output, and, accordingly, a translation into Russian. What it's all about. Before generating any code, [music] you listen to what you need to listen to, what I'm telling you, tell me how you understand it. And if we, I give consent, then we create the code. And here, because what I wanted to say here is that I can't always, for example, correctly express my request. And if I immediately, let's say, poke it here, right, then it, well, in short, this will be overloaded. Although, as they say, as developers say, it understands everything itself, what you can say or write directly here, but I adhere to a slightly different plan, so I'll remove all this. And Google writes to me: "Got it, buddy, great idea. Vibe coding is exactly about that, ttt." So, you want to create a tool that works according to the following [music] cycle. You give the system some idea or description, it creates, in short, the application turns your idea into a detailed artistic prompt in English. Simultaneously, a translation for this prompt is created. Generation after confirmation or automatic application. So, generation after confirmation or [music] automatically the application sends a request to generate an image. Output: you get the picture itself and a text block in English prompt. Where will the generate button be? We do this directly within Google AI Studio [music] as a system instruction or write a separate mini-service in Python. Look, we do all this within Google AI Studio as a system instruction. That is, you write a system prompt, what to do, where to put it, and so on. Let's create categories on the first, on the main screen first. [music] Let's make two categories. The first category will be uh objects. The second category will be uh [music] rather, the first category will be architectural objects, and the second category will be, uh, [music] technical objects. Let's see how it understood me, because, as I said, if you formulate it incorrectly, it will understand incorrectly, accordingly. Therefore, I'd better check [music] 10 times. So, my understanding of the logic. Interface selection. When the user starts the chat, Gemini should offer one of two branches: architectural objects or technical objects. If architecture is chosen, Gemini focuses on light, materials, [music] concrete, glass, wood. If technical objects are chosen, the focus is on [music] mechanisms, drafting accuracy, and so on. Well, that is, it even offers you some ideas itself, right? If you like such an idea, [music] you can take it. If not, you can simply say later: "No, I don't want this. I want to create a subcategory in each category later." Well, let's do as it wants. Okay. And it asks me what else, what details does it need. How will we choose the category? Do you want it to simply ask at the start, choose category one or two? So, how do we choose a category? That is, when the user opens this application, they see two cards on the screen. One card is architecture, the second card is technical objects. By architecture, I mean world objects that are very famous. For example, I don't know, the Sydney [music] Opera House, the Statue of Liberty, the Kremlin, I don't know, some mosque, Cologne Cathedral, and so on. By technical objects, I mean cars, robots, gadgets, steamships, submarines, and so on. For starters, let's do it this way. And the depth of detail, right. Do you need to, yes, do you need to specify the depth of detail, format, and, accordingly, although we will work on this moment later, so let's stop at categories for now. So, step by step, we are approaching the logic of how it will work. Greeting. Upon launch, Gemini immediately displays a stylized menu [music] stylized menu with these two categories. Logic: if you choose architecture, it connects to the knowledge base of world masterpieces. Aha. If you choose a technical object, it switches to an engineering mode. Authorobot, etc. The prompt will automatically [music] add technical parameters. Your system prompt looks like this. Let's do this for starters. We copy this code, go to our studio, paste this code, click build. Wait. We can see here what the code writes. Or here all [music] its steps can be seen, what it does. And while it's coding, I can say that you can simply, if you have a plan, it's better to write it down somewhere on paper so that you have a visual concept of what you want to do. And if not, you can just experiment. Well, look what we [music] got. We created two tabs or cards, whatever you call them. One is called architecture, the other is technology. Okay, let's click on it, and it, as it told us, what do you want to do here, for example, right? For example, what object do you want to [music] visualize? Well, let's write, for example, let's just write, maybe it will guess or not. Let's see how smart it is. Okay. And here it gave us a prompt and in Russian. So, a cinematic wide-angle architectural shot of Cologne Cathedral during the golden hour, showcasing its double spires. So, we can copy this prompt, of course, and go to any other generator, for example, the same Gemini, NanoBanana or wherever. Paste this prompt and see what it generates. [music] But it would be more convenient if it generated not only prompts but also images, for example. And we'll do that now. Here, please. We got such a cool image. Well, yes, it corresponds directly, as if, right? Let's ask it to do not only prompts but also to create photos itself. And we go here, to Google Gemini, to our chat. Where is it? I want this program to generate not only a prompt but also an image at the output. Also, before generating an image, I could choose the format 16x9, 9x16 [music] or 1x1. That is, there should be additional buttons before generating an image. So, this is an excellent addition. You have to say, it always praises you when you try to program something. It says: "You're so great, you come up with such things." Well, check, don't relax too much anyway. So, there are no physical buttons in the interface, but we can make the model simulate working with buttons through text selection. Well, it's struggling with something here, everything here [music] is. So. And here's another point, it always writes, sometimes it rewrites [music] the entire code, the system code from the beginning, right, and this, let's say, loads this whole program. We can do this. You are, of course, great, but let's agree on this. You should not [music] rewrite the system code from beginning to end every time. You should just write so that the program understands that this is an addition to the system code. And please, write me. These additions are already ready, so I don't have to figure out [music] where to insert it, into which block. You just write me this system prompt addition, so that I can copy and paste it into this application. Oh, man, no problem, buddy. Well, you're something else. In short, it started writing everything from scratch again. Well, [music] you understand, that is, you need to negotiate with it. Let's copy it completely and see what it's doing here. Here we take this window at the bottom and paste it. Here we click the arrow, wait, let's see what it's doing here. >> [music] >> If you are going to create such an application, try not to give it a big task right away, but step by step, for example. That is, first create the framework, then create the main category, or create one function, then add another function, and so on. Because if you give it a large volume of tasks, it immediately, in short, starts hallucinating, struggling, and then it's a complete mess. Let's see what it did for us. Let's choose technology, for example, let's write. Choose format. So, it offers us 16 by 9, 9 by 16, 1 by 1. Well, let's do 16x9. And let's go. It's making an architectural prompt. Architectural prompt and generating visualization. [music] That is, we wanted to get a ready-made image and prompt at the output. Well, look, it's already doing it for us. Here, please, a piece of an airplane. I don't know if this is right or not. Well, it should be. It says, prompt in English. Prompt, or translation into Russian. Cool thing. Cool. [music] And this is how you can make your own application step by step. In general, cool. It even gives hints here. The bot itself, for example, add, uh, how to download the prompt, download the image, for example. Well, let's add, let's add download, rather, download prompt. Here it says, let's write. So, can we download this at all or not? visualization. So, we can't. Let's write [music] here, add a function here so that I can download the generated image. Click, wait. Imagine how much time this saves for people like me, for example, who don't understand [music] code at all. If you, well, and programmers in general, it saves a lot of time. So, guys, let's see what we have here. Architecture, for example. Ah, a temple in the style of Zaha Hadid. Okay, let's write something to it. Ah, I'll write the city of Tashkent. The city of my childhood. Tashkent. What will it do for us? Choose format 16 [music] by 9. So, we wanted to add a function so that we could download the image. Please, I copied it. Here. Uh, well, about this, I don't know, I haven't seen anything like this in Tashkent. Maybe it invented it itself. Okay. And here, please, a download button appeared. Click download. Download. Everything is great. Everything works. Everything we need. So, guys, use it. Create, fantasize, you can do it. In principle, [music] it's all simple. Sometimes it doesn't work out the first time. Sometimes I started several times, then got confused, then realized, ah, you can't do it this way, because I was doing too many tasks at once or wanted to do too difficult a task. Uh, and I want to show this application again, the one I made. Personally, it still works for me, thank God, it's very massive because there are a lot of tasks done in it. There are even stories where we made images. And there's also a cool button. Here I made it, I want to brag a little. Imagine, you want some images, for example, and you have absolutely no idea, for example, and you just click the "Surprise me" button, and it goes through all the categories, subcategories that are built in here, it starts flying through and creates something crazy itself. Click, choose a format, launch, and that's it. Well, in short, I showed this to my daughter. She, in short, is wildly delighted. She's already messed up something again. That is, it goes through all these formats at once. And nothing worked [music] for some reason. Hmm, strange. Let's go back. Why didn't it work? Sometimes there are such glitches. I hope it works all the time. But it's always like this effect, you know, when you want to show it, it doesn't work. So, choose format [music] 16 by 9. We don't write anything here. Start synthesis. Again, it didn't work. Let me try to restart this thing altogether. It's happened a couple of times already. I don't know what the problem is. Maybe the code is just overloaded. So, what is this? No, this is not it. This is what we created. [music] Here's ours. So, let's try again. If not, we'll figure out what the problem is. We'll have a rhinoceros. Okay. Well, and here, please, a rhinoceros in a glass cube. Surreal [music] visualization where the power of nature merges with high technology. A rhinoceros made of flowing liquid metal. Oh, disassembled into layers in the weightlessness of a digital aquarium. Well, in short, it's like that. Let's download it, let's see what we got up close. That is, it assembled everything. Well, such a crazy thing, guys. Look, in short, it's made of liquid glass, and also disassembled into parts. These parts are displayed as these pictures on the glass. And it's all in a cube. Well, in short, wow. So, [music] once again, guys, I hope if it was useful, subscribe, like, don't forget to share this video. The algorithm loves it when you share videos, and I, accordingly, need your comments and your opinion on how you liked it and whether it's worth continuing to make videos on this topic. Bye. M.