Transcription
This is the ultimate AI image generator. It's free and open source, and the awesome thing about it is you don't even need an Nvidia GPU. You can run it with just a CPU or an Apple M1 or M2 chip, and it gives you total control. Not only can you do text-to-image, but also image-to-image, plus upscaling. Plus, you can control the poses of your character, or you can control the positions of objects in your image, or you can control the depth of your composition, and much more. You can also do face swapping and create consistent characters. You can also link to popular AI tools such as Mimic Motion, which makes your character dance, or Live Portrait, which lets you animate any photo of a face, or Turn Crafter, which allows you to input a start frame and an end frame, and it would generate an anime scene in between those two frames. The possibilities are just endless.
Now, the tool that we're going to go over today is called ComfyUI, and the interface looks like this. So, if you're seeing this for the first time, this might look very complicated to you, and that's exactly why I'm making this tutorial. I'm going to show you step by step how to install it, and then how to do text-to-image, image-to-image, upscaling, how to control the pose or the composition or for the depth of your images, plus how to download and use external tools like face swapping or Turn Crafter. I'm going to make it as easy as possible so that even if you don't have any technical background on AI or Stable Diffusion, you can still follow along easily. So, let's get started.
All right, first, let's go over how to install ComfyUI. You simply have to go to their GitHub, which I'll link to in the description below, and installing it is really easy. All you got to do is scroll down, and then you'll see this "Installing ComfyUI" section. And then, note that you don't need an Nvidia GPU to run this. You can also run it on CPU only, and it also supports Apple M1 or M2. Although, of course, if you do it this way or if you run it using CPU, it's going to be way slower compared to an Nvidia GPU. So, if you're serious about running image generation locally, definitely do get an Nvidia GPU. It's just going to make your life a lot more easier, and it's going to be compatible with a lot more other open-source AI tools. By the way, I'm using a Dell Precision 5690. You can integrate a powerful RTX 5000 Ada into this. Huge thanks to Dell and Nvidia for sponsoring this. Anyways, I'm using Windows, so under Windows, all you got to do is click on this direct link to download. So, this is going to install a .7z file, which we can unzip. Now, note that this is 1.4 GB in size, so it's going to take a few minutes to download depending on the speed of your internet.
All right, so once you've installed the .7z file, you can unzip it with 7-Zip or WinRAR, and then you're going to see this folder. Now, you can extract this folder anywhere. So, I'm going to extract this on my desktop, and this is a very large folder, so it's going to take a while for this to fully extract. All right, so once that is fully extracted, we can now open up the folder, and then depending on whether you have an Nvidia GPU or something else, double-click on the appropriate .bat file. So, in my case, I do have an RTX 5000, so I'm going to run this, and then you might see this message. So, we got to click "More info" and then "Run anyway." And so that's how you would install ComfyUI. Very simple to use.
Now, if you click on "Queue Prompt," you're going to see this error message. We don't actually have any checkpoints or models yet, so we got to download a model to use for our image generation. Now, there are thousands of models you could potentially use. A model basically defines the style of your image. So, for example, you can choose from realistic to Disney Pixar style to watercolor to anime. There are like thousands of checkpoints or models that you can choose from, and you can browse all of these models on a site called Civitai, which I'll link to in the description below. Now, be warned though, there's a lot of NSFW content here, so make sure you have your filters on, or it's going to be very not safe for work.
Now, because there are like thousands of models, how do you know which one is the best? Of course, you want to filter out all the bad ones and just use the best ones. Well, there's an awesome site called Imagecyst, which I'll also link to in the description below. This is basically a rankings list where users can kind of blind test different models, and then all these results are accumulated into this ranking list. So, you can see, for example, that RealVis has the most points so far, followed by Colorful XL, followed by Playground Version 2.5, followed by Juggernaut, which is also a very popular model. And note that most of the ones at the top are using Stable Diffusion XL, which is basically a higher quality version compared to Stable Diffusion 1.5.
So, let's say we want to use this one, RealVis XL. I'm going to copy this name, and in Civitai, I'm going to simply search for this. So, here is RealVis XL Version 4, and here are the results. You can see if you scroll down, here are some sample images from other users. So, let's just download the most recent one, Version 4.0 Lightning. So, let's go ahead and download this. It is 6.4 GB, just a warning, so it's going to take a while to download. And then you would place your download, which is a .safetensors file, in ComfyUI, and then models, and then checkpoints. So, let's save it directly in this folder.
So, once you've downloaded the checkpoint or the model file, make sure that it is located in ComfyUI, then in models, and then in checkpoints. So, notice that I have downloaded this RealVis XL .safetensors file in this folder. And so, the next time you start up ComfyUI, which I'm going to do now, you should be able to select the model. So, again, I'm going to open "Run Nvidia GPU.bat," and then wait for this to give me a link. All right, perfect. So, now if you go back here in the "Load Checkpoint," if it doesn't automatically select one for you, you can click on this, and note that I can select this RealVis .safetensors file in here.
Now, note that for other tutorials, they suggest that you download the SDXL base file and also the SDXL refiner file, but this is just the default Stable Diffusion XL model. Actually, this is not necessary, and right now, SDXL has gotten so good that you don't really need to use the refiner file anymore in most cases. And this base SDXL model, it's not great, so you don't have to download this, and you can save a few gigabytes of disk space. Again, if you look at the rankings from Imagecyst, the SDXL base files, at least the lightning model, seems to be in 12th place, so I wouldn't actually recommend you download this base model and the refiner model. So, you can actually save a few gigabytes of disk space.
So, anyways, back in here, all we need to do is download whatever model you want. And then here is the positive prompt, and then here's the negative prompt. You can use your mouse's scroll wheel to zoom in and out, and I'll explain what all these boxes mean in a second, but first, let's just generate an image to see if this all works. So, for the positive prompt, I'm going to enter "a castle in a forest," and then for the negative prompt, this is basically all the things you don't want to see in your image. For example, I don't want to see people, cars, knit, and that's pretty much it. I'll cover some more advanced prompting techniques later in this video, but let's just click "Queue Prompt" for now.
And you can see right now it's running this box, which is highlighted in green, and then it's running the positive prompt, the negative, it's going through the K sampler, and then now you can see the progress bar is up here. So, it's in the process of generating the image, and then it goes through this VAE decode, and finally, we get an image of a castle in the forest. Now, this image is not great, and we're going to go over how we can make this better.
Now, first, I want to really give you a solid understanding of what all these boxes mean, because I think it'll help you when you build more complex workflows. So, first of all, let's just ignore all of this and start from scratch. So, if I move down here, you can start a new box or node by right-clicking, and then you can see "Add Node," and you can select whatever you want. Now, there are so many options here, I wouldn't recommend this way. So, another way to do it is to double-click, and then after you double-click, you can actually search for whatever node you want. So, if I type in "checkpoint," for example, you can see that we have a "Load Checkpoint" node here. So, I'm going to select this, and because the RealVis model is the only model I have in this models folder, it selects this by default. So, again, this checkpoint basically defines the style of your image. For example, there's going to be some that work particularly well for anime, some that work well for realistic photos, and then some that work well for Disney Pixar like characters.
All right, so next, we need to add in our positive and negative prompts. Now, again, and we can either right-click and try to find the positive prompt node here. I would not recommend that, it's really hard to find a node in these options, or you can double-click and then search for the node here, in which case we see it here. Or another way is to drag from one of these connectors. So, the prompts should be linked to this CLIP connector. So, I'm going to drag a noodle out, and then once I release it, you can see that it gives me several options which it thinks is most relevant, and indeed the prompt window is this one, "CLIP Text Encode (Prompt)." So, we actually need to drag two of these, one for the positive prompt and one for the negative prompt.
Now, there are several ways to clone this node. So, one way is to right-click on this and then click "Clone," in which case we'll have another node here, but then we'll need to connect this ourselves. So, I'm just going to delete that. And then another way is you can just click on this and press Ctrl+C to copy it, and if you press Ctrl+V to paste it, this is what you get. But if you want this to have the same connectors as this original node, then instead of pressing Ctrl+V, you can press Ctrl+Shift+V. And so now, when I paste in this node, it's automatically connected to this.
Now, just to avoid confusion, you can also rename this. So, for example, you can set the title to "Positive Prompt" and then press OK. And then here, you can right-click and then set the title to "Negative Prompt" and then press OK. You can also set the color if you want. So, for example, I can right-click and then in "Colors," I can set this to red. And this is important because when you build more complicated workflows, this graph is going to be really complicated. So, if you color-code things and rename these nodes, it just keeps things more organized and helps you avoid confusion.
All right, next step is we need something called a K sampler, and it's basically an algorithm that takes in your prompts and takes in a latent image, which we'll go over in a second, and creates an image from that. So, if we click on this conditioning connector and drag it out, we should see "K Sampler" here. So, I'm going to select that. This is for the positive, and then for the negative prompt, we will connect this to here. And then for the model, we just drag this one all the way to here where it says "model." And then for latent image, let's drag a line out, and then here it would give us the option to create an "Empty Latent Image." So, let me just drag this over here to keep things cleaner.
All right, so what on earth does this mean? What exactly is an empty latent image? So, how Stable Diffusion works is it doesn't create an image from a blank canvas. Actually, what it does is it starts with an image of just random noise, and then at every step, it removes some of that noise. And if you remove enough noise, you get whatever image you prompted it with. So, this empty latent image is basically an image of random noise, and you can select the width and the height. So, for example, we can change this to 800 if you want, and then same with the height, 800. By the way, for SDXL, it's best to create an image of 1024x1024 because that is what it's optimized for. And then if you are using Stable Diffusion 1.5 or earlier, it's best to use 512x512. And then the batch size is how many images you want. So, if you set this to two, basically it will create two images at once. But for now, let's just set this to one.
All right, so next, let's go over all these settings. The seed is basically the starting point of this random image. This image is just random noise, but there could be an infinite number of images of random noise, each of them would be slightly different. So, the seed basically defines the starting point, and usually we keep the seed at random, so that's what this one does. However, if you keep it at fixed and you set the seed to a certain value, for example, 69, then if all of the other settings are the same, you're going to generate the same image every time because the seed or the starting number is the same. But for now, let's leave it at zero, and then for here, let's set it to "randomize."
And then the number of steps is basically, again, if we go back to how Stable Diffusion works, is how many steps of noise removal do we want. So, if it's just a few steps, you're only going to get something like this, you haven't removed enough noise yet. And after enough steps, you're going to get a very clean image of whatever you want to generate. So, generally, few steps would give you a lot of noise, and then after a point, like if you exceed 50 to 100 steps, then you're not going to get a better image. In fact, you might get some noise or artifacts because there's not much remaining noise to be removed. So, I'm just going to keep the number of steps at 20 for now.
And then CFG is basically how well do you want this algorithm to follow your prompt. So, if you set this to one, for example, it's not going to follow your prompt, and it's going to try to generate whatever it wants, basically, it can be more creative. However, if you set it to 15, for example, then it listens to your prompt very literally, and it tries to generate whatever you specify in your prompt, which is sometimes not what you want. Sometimes, if it's too literal, you're going to get some weird results. So, generally, a value of like seven or eight would work best.
And then the sampler name, this is basically the algorithm that's used to remove this noise and generate the image for you. So, Euler is a common one, this is one of the fastest ones. And then if you want to go for quality, I would suggest one of these, so DPM++ 2M or 2M Karras. But for now, let's just go with Euler. And then for scheduler, this basically gives even more quality. So, usually, if you want a really good quality image, you would select Karras or exponential. These settings are just very subtle differences you can add to the image. So, just play around with it so you can get a feel for what settings work best for your particular use case.
And then denoise, what this is, is because again, our latent image or starting image is basically an image of random noise, the denoising strength is basically how much noise do we want to remove from this initial image. So, if we're starting from scratch, then obviously we want to remove 100% of the noise, so that's why this is set as one.
All right, so the next step is we next need to connect this latent connector. So, I'm going to drag this out, and we need to connect it to what is called a VAE decode. A VAE basically encodes an image into a latent space, but right now we want to do the reverse of that. So, we need to decode this latent image into an image that we actually want to see. So, that's why we need this final step to decode this latent image. And then we need to drag the VAE from our checkpoint to this VAE connector that we see here. So, in most cases, your checkpoint should come built-in with a VAE, so all we got to do is just connect this VAE to this VAE connector here, and then we are almost done.
Right now, we just need to produce the image. So, once we drag this out, we actually have two options. We can either preview the image, so if we preview the image, it's not going to save the image on our computer. Or we can select "Save Image," which also shows us a preview of the image, but it also automatically saves the image. So, I'm going to select "Preview" first because I don't want it to save every image it generates, I only want to select the good ones to save on my computer. So, I'm going to show you how to save a previewed image in a second.
There are very few software tools that I use every day, but this is one of them, thanks to our sponsor TurboType. It's free forever and it saves me so much time. Basically, you can create custom keyboard shortcuts so that you don't need to keep typing out repetitive things. For example, if there's a prompt that I use in ChatGPT very often, I can make a shortcut here, and then when I go to ChatGPT or anywhere else, I just need to type in the shortcut and voila. Or let's say I have a very long email address, I can also make a shortcut for that so that whenever I need to enter in my email, I can simply type in the shortcut and it types out the email. Finally, it also supports rich text, for example, you can add in bold and italics and add links to your text as well. So, let's say I need to send out a lot of cold emails with the following template, well, I can just create a shortcut for that, and then whenever I start an email, I just need to type in the shortcut and voila, the text is already styled and linked for me. There are hundreds of pre-existing templates that you can choose from, including common prompts for ChatGPT, business, finance, medicine, and more. This tool saves me so much time every day. There's absolutely no reason not to use this because they have a free forever plan, so definitely check it out and download the free Chrome extension in the link below.
But that is pretty much it. So, for the positive prompt, let's put "a medieval warrior, realistic, 8K masterpiece." These are just some keywords that I tend to use a lot to give it more detail. And then also "ultra detailed" is another good one. And then for negative prompts, again, these are all the things we don't want to see in our image. So, for example, I don't want it to be a painting or a cartoon or anime drawing. I don't want it to have any copyright or watermarks. And I think we are good to go for now.
So, before we click "Run," note that I had this previously. This is the default, so I don't want it to run both of these at once. So, I'm going to select all of these nodes and delete them. So, to select multiple nodes at once, what you can do is hold down Control and then drag to encompass all the nodes that you want to select. And then if you want to move this group of nodes around, you need to hold down the Shift key, and then you can drag this wherever you want. And then to delete everything at once, all you got to do is press Delete.
All right, so moving back down here, everything else looks good. So, I'm going to press "Queue Prompt." So, note that it's starting here, it's loading the checkpoint, and then it's moving over to K sampler. Now, it's generating the image, and then it's decoding the image, and then voila, we have a medieval warrior. So, basically, this is like the standard workflow for a simple text-to-image generation.
And then let's do another one. So, let's say I want to generate two images at once. All we got to do is increase the batch size to two, and then press "Queue Prompt" again. And note that it starts off here, it starts in the K sampler, it doesn't start here or here, and that's because we haven't changed any of these other nodes in a previous step. So, all of this is already saved in memory. It only loads from here, and so this makes ComfyUI very efficient. And then so now we have two images. This is the first one, here's the second one. And remember, this is not saving your image, this is just the preview image node. So, to save it, all you got to do is right-click and then press "Save Image."
So, this is the most basic text-to-image workflow for Stable Diffusion. Hopefully, this gives you a better understanding of what a K sampler is, what an empty latent image is, and what a VAE decode is, because once you set up more complex workflows, you're going to need to understand what these nodes actually do. So, I hope this gives you a good understanding.
All right, before we move on to the next section, let's go over some tips and tricks for navigation and organization and productivity. So, first thing is, I'm not sure if you can see it on my screen share right now, but there is a very faint dark blue frame on your ComfyUI canvas. So, I'm hovering my cursor over that frame right now. I wish they could make the contrast higher so you can actually see the blue line, but anyways, there's this very faint blue line, and so whenever you start an interface, the default location would be within this blue frame. So, it's always best to have your workflow within this blue frame so that whenever you load up ComfyUI, your workflow will show up right away, and you don't need to like try to find it within this huge canvas.
All right, next thing you can do is, let's say this is a workflow that you want to use again in the future, you can click this to save it or press Ctrl+S. So, let's name this as "temp.json," and then you can save this wherever you want. I'm just going to save this in the ComfyUI folder. Another keyboard shortcut to be aware of is A, which is select all, and then Delete or Backspace, which will delete everything you've selected.
All right, so right now, this whole workflow is gone. Now, I can undo it by pressing Ctrl+Z, which would undo my delete of the workflow, or I can also redo the action by pressing Ctrl+Y, which deletes the whole workflow again. Now, to load up a previously saved workflow, you can press "Load" here or press Ctrl+O. So, if we go into our ComfyUI folder and then select this "temp.json," which we just loaded, you can see it has loaded our workflow back up.
Now, a few shortcuts for navigation. You can either press on your mouse and then move the canvas around, or you don't actually need to click on your mouse, you can also hold down the space bar and then move your mouse around, and it would still move the canvas. Now, you can zoom in and out by using the scroll wheel, and then to select multiple nodes, you just simply hold down Control and then select this one, and then let's say I want to select this one and select this one, so I'm holding down Control for each of these, and then let's say I want to delete these, I can just press Delete. And then let me undo that by pressing Ctrl+Z. You can also select multiple nodes by holding down Control and then dragging a frame around all the nodes you want to select, and then if you want to move these nodes around, simply hold Shift and then you can drag this group of nodes that you've selected wherever you want.
Another thing you can do is, let me delete this first. Now, every time you drag a connector out, there's also an option called "Reroute." So, if I click this, it's basically just an extra blank node which extends your connection further. So, it's basically the same thing as just connecting this VAE to this VAE, but the nice thing about this is, let's say you don't want this line to be hidden behind this node, well, you can drag it out like this, so you can clearly see that this line is being connected here. And then one more thing is, right now you see that for example, in K sampler, these values are set in this node, but what if you want to set this value somewhere else and then link it to here? So, for example, for CFG, if you want to set this value somewhere else and then link it to the K sampler, you can right-click on this, and then under "Convert Widget to Input," you can set any of these options to an input. So, let's say we want to set CFG to an input. Now, you can see that CFG has disappeared from here, and it is now an input connector. So, let's drag this out, and then we need to actually use the node called "Primitive." So, right now, this value is set to seven, it's connected to this CFG input, so the CFG of this K sampler is 7. So, that's how you would use it.
All right, one final thing I want to share with you, and this is really cool. Let me press Ctrl+A and delete all of these. Let's say I made an image previously using ComfyUI. Well, if I drag that image onto the canvas, what happens is it actually gives me the entire workflow that I used to create the image. How cool is that? So, like if you go online and other users have shared their ComfyUI generations, and assuming they haven't deleted that metadata, you can actually download their image and then drag and drop that image onto ComfyUI to look at the entire workflow that they used to generate that image. This allows you to learn really quickly. So, those are like the basic things you need to know for organization and productivity using ComfyUI. So, let's move on to the next section.
First of all, we need to install this plugin called ComfyUI Manager. It will make your life a lot easier for installing extensions and plugins and missing models. So, I'm going to link to this GitHub repo in the description as well, and if you scroll down a bit, they will give you some installation instructions. So, in our ComfyUI folder, and then in our "custom_nodes" folder, we just need to open Command Prompt here. So, in this bar up at the top, type in "cmd," and this will open up our Command Prompt, and you can see that we are now inside our "custom_nodes" folder. Next, all we got to do is "git clone" this repo. So, we will paste it in here. Now, you do need to have Git installed first. If you don't, here's how to install Git. If you already have Git installed, feel free to skip to the next section.
So, all we got to do is download the latest release for whatever operating system you're using. So, I'm using Windows, so I'm just going to click on "Download for Windows." I'm running 64-bit, so I'm going to click on this to download, and it's now downloading this .exe file. So, once that's completed, all we got to do is open that .exe file and then follow the steps. So, I'm going to click on "Next." I'm just going to go with the default install location, which is "Program Files/Git," so I'll click "Next" for that, and then I'm just going to leave this at the default, and then I'm going to click "Next" again, and click "Next" here. We're just going to use the default settings for all of these. There's a lot of settings that you need to go through, so I'm just going to click "Next" for all of these. All right, and then it should go ahead and install all the files. So, this might take a few minutes. Perfect, so now we have Git installed.
All right, so assuming you have Git installed already, you simply copy this line and then paste it in here, and then press Enter, and you'll see that now we are cloning this ComfyUI Manager into the "custom_nodes" folder. And then if you actually open up your "custom_nodes" folder, you can see this new folder called "ComfyUI-Manager."
All right, so next, we need to restart ComfyUI. So, going back in our Windows portable folder, we are going to run ComfyUI again, and you can see that it has detected that we have ComfyUI Manager installed. So, it's now installing dependencies. All right, so after you open your ComfyUI, now you should see this "Manager" button in the right menu.
Next, I'm going to show you how to upscale an image, and then we're also going to move on to some more advanced workflows like image-to-image and ControlNet, and installing other plugins. So, first, let's click on this "Manager" button, and then the main buttons that you will use is this one, "Custom Nodes Manager." There's also a button where you can like update ComfyUI or update all, or install missing custom nodes. For example, if you import a workflow that was made from another user, you might have some missing models or nodes or dependencies. So, clicking this button will just automatically install all of them so that this other user's workflow will work on your computer. And then here, this is "Model Manager." So, instead of going to Civitai and browsing through all the models, you can just easily download the model here through this interface. So, let's click on this, and then let's search "upscale," and you should see a lot of different upscaler algorithms. So, usually, the ones that I find work best are RealESRGAN x4, this is for realistic photos, as the name implies, and then there's also 4x-UltraSharp, which works pretty good as well. Just to keep it simple for this tutorial, I'm just going to download these two, but definitely you can install all these and play around with it and see which algorithm works best for you. So, I'm going to select these two and then click "Install."
All right, so after it is finished installing, note that you need to click the "Refresh" button on the main menu to apply these installations. So, we're going to click "Close," and then "Close" again, and then click "Refresh."
All right, let's dive in to see how we can upscale images. So, we're going to use the same prompts: "a medieval warrior, realistic, 8K masterpiece, ultra detailed." And then for the initial image width and height, let's set it to 512. Now, I understand that for SDXL, the optimum dimensions are 1024x1024, but in this example, I'm just going to show you a really blurry and low-resolution image, and then we're going to upscale it by four times so you can clearly see the before and after. And then for the batch size, let's leave it at one for now. Number of steps, we can also leave it at 20. And then here, instead of "Preview Image," let's delete that, and then we can just drag this image out, and then we don't see any upscale here, so we need to click "Search" and then we can search "upscale."
Now, there are a few options. One is "Upscale Image," which is kind of the same as "Upscale Image by." And then we have "Upscale Image using Model." Now, I'll show you what this one does first. "Upscale Image by," I would not recommend this because this is basically just upscaling your initial image, but it's not adding any details. It's basically just increasing the size. So, let's say we want to upscale this by two times, so instead of 512x512, it's going to be 1024x1024. And then for the image, let's drag it out, and we will have a "Preview Image" node. And also, I want to preview the image before we upscale it, so you can compare the before and after. So, over here, I'm also going to drag a "Preview Image" node.
All right, now let's run this and see what we get. So, I'm going to press "Queue Prompt." All right, so let's look at the initial image. This is only 512x512, so I'm zooming in quite a bit, and you can see the details of his face are very blurry. And then let's look at the upscaled image. Yes, this is 1024x1024, but you can see the details are the same, this is pretty much the same image, the same blurriness, we're not adding any details here. So, again, it's not recommended to use this "Upscale Image by" method because you're simply just resizing the image, but you're not adding more details to the image. So, I'm going to delete this one and also this one.
So, next, let's drag out another node, and this time I'm going to search "upscale" again, and then we are going to use "Upscale Image using Model." And then after that, we need to actually input an upscale model. So, I'm going to drag this node out, and then we should get the one and only option, which is "Upscale Model Loader." So, because I've downloaded 4x-UltraSharp and RealESRGAN, we should automatically see this over here. Now, both of these are 4x, so your output image will be four times the resolution. If you want 2x, for example, well, you can go into "Manager" again and then click on "Model Manager" and then find an upscale model that is only 2x, like this RealESRGAN x2.
All right, and then just one last step is we need to drag out another node to either save the image or preview the image. So, if I run this again, you can see that if I zoom in on both of these, this is 512x512, so the details are very blurry. But if you expand this image, which is like over 2000x2000, you can see that the details are a lot sharper, especially the patterns on his helmet and on his armor. Now, this isn't the best way to upscale. It's actually best to do one round of image-to-image first before we upscale. So, I'll show you that in a second.
So, one more thing I want to mention is that sometimes you don't want to upscale all the images that you generate, you want to decide which image you want to upscale. Because let's say for this initial image, you don't like the design, you don't like the composition, you don't want to proceed further with upscaling and waste computing resources. So, how do we decide if we want to upscale or not? Well, first of all, you can break off a workflow by selecting the node where you want to break it off. In this case, I want it to pause here, so it only generates the initial image, but it doesn't proceed further to the upscale unless I want it to do that. And then I'm going to press Ctrl+M, and this will mute the node, and basically everything that goes after this point is going to be paused until I unmute this node. Now, when I press "Queue Prompt" and it generates an image, I can decide whether I want to proceed further with this upscale method, and if I do, then I would unmute this node and press "Queue Prompt" again. However, one thing to note is in this case, you do need to set the same seed, otherwise, if you don't set the same seed when you press "Queue Prompt" again, it's going to generate a completely different image and then it's going to upscale that image. So, going back to here, we need to actually set this to "Fixed," so it's going to be a fixed seed. And then, so if we run "Queue Prompt" again, you can see that it's generating a new image, and that image is being fed into here. And then let's say I like this image, I want to upscale it, then I would click on this node, which is now muted, I'm going to unmute this by pressing Ctrl+M. So, now when I press "Queue Prompt" again, you can see it's actually proceeding from here, and then now it's upscaling this image, and we can now see our upscaled image. So, that's one way to do it.
Another even more efficient option is if you go back to "Manager" and then click on "Custom Nodes Manager," and then let's search for "Image Chooser," and this is a node that's created by Chris Orange. Let's go ahead and install this. All right, all right, so it says "Restart Required." So, let's restart this. All right, so we are back here. Now, let's say we want to generate four images, and we would choose one of them to go through the upscale. So, let's increase the batch size here to four, and then we can keep everything as is for the seed. Actually, we can set this back to "Randomize." And then, yes, this has to go through this VAE decoder to decode the latent image, and then instead of "Preview Image," let me just delete that. Let me also delete this node for now, and then let's drag this image out, and we will search for "Preview Chooser," which is down here. And then for any images that we select here, we can then proceed to upscale it. So, let me just run this for you first, so you can see what this does. So, right now, it's loading the checkpoint again because we've restored the interface. Now, it's inputting the positive, negative prompts. Now, it's generating four images through this K sampler, then it's going to decode these four images, and then so now we have four images.
All right, so let's say we want to select this one to upscale, so we would click on this. Or if you want to select more, you can always increase the count to two and then select this one, for example. But I'm going to decrease this to one and unselect this, so we are only going to proceed with this image through the upscaler. And then we simply click "Progress Selected Image," and then this would go through the upscaler, and it's using our 4x-UltraSharp model, and then voila, here is our upscaled image. So, in a nutshell, that's how you do upscaling.
Now, there is an even better method to upscale images which kind of uses image-to-image. So, next, we're going to go over how to do image-to-image, and then we're going to go back to this better upscaling method.
All right, next, I'm going to show you how to do image-to-image. So, what I'm going to do is, first of all, hold Control and then drag these nodes to select all these nodes. I'm going to copy it and then paste it somewhere here, and then I'm going to hold down Shift and then move it to where I would like. All right, now going back to this, another way instead of just deleting this workflow is to hold Control and select all of this and then press Ctrl+B to bypass. So, if you see the nodes highlighted in purple, that means it will be bypassed. This will not run, so only this will run.
All right, so this is our standard text-to-image workflow, right? We have our checkpoint, and then we have our positive prompt, our negative prompt. This goes into the K sampler. The K sampler takes a latent image, which is just an image of random noise, and then after going through this algorithm and going through this amount of steps, then the final step is we need to use our VAE to decode this latent space back into an image that we want to see.
Now, for image-to-image, we don't want to start with an image of random noise, right? We want our input to be an image. So, let me select this node and delete that, and then I'm going to double-click anywhere and then search for "Load Image." So, let me click on this. So, this is the default, but you can click this button to upload an image. I'm going to upload this image. Now, we can't just directly drag this image to this latent image connector. That's because we need to first convert this image to a latent image and then connect the latent image into here. So, let me drag this connector out, and you should see the first option here would be "VAE Encode," and that's exactly what we want to do. We want to use a VAE to encode this into latent space, and then drag the latent image onto this connector.
All right, and then where do we get our VAE? Well, we can just get the VAE from the checkpoint that we loaded. So, let's drag this to here, and then you can see that everything is connected now. We do need to set up some additional settings for the K sampler. So, one thing I forgot to mention is that if your checkpoint is lightning, it only takes around like 5 to 8 steps. You actually don't need 20 steps. That's the awesome thing about lightning models, and some models can even generate a decent image in as few as two steps. For us, let's set the steps to seven. Sampler schedule, we can leave it as is. Denoise, this is what we want to change. If you remember, for our text-to-image, denoise this is saying to take our latent image of random noise and replace 100% of that noise so that we get the image that we specified in our prompt. But in this case, because we're not starting with random noise, we're starting with an image, we don't want to remove everything. We want to retain some of this image. So, in this case, if you're doing image-to-image, the denoising strength means how much of this original image do you want to remove. So, if you set this to 100% or 1.0, in this case, it's going to remove everything from this image, you're going to get a completely different image. Conversely, if you set this to zero, then it's just going to produce this exact image, nothing would change. So, depending on how similar you want the image to be, it's better to have something in between. So, let's do 0.3 first, so it's more similar to this image, and then I'll show you an example of 0.8 for comparison.
All right, so now that we have everything in place, there's just one final step, which is to drag this image connector out, and then we are going to "Preview Image." Let me just reposition this here so you can see the entire workflow, and we are good to go. Let's click "Queue Prompt." So, it's our first time starting this workflow, so it takes some time to load in the checkpoints, and then it's going through the prompts, it's taking this image and encoding it, it's now going through this K sampler, and then it's going to decode the image and then preview the image for us.
All right, perfect. So, you can see if I drag this over here, just temporarily for comparison, this is image-to-image. So, here's our original image, here's our new image, and the denoising strength was set to 0.3. All right, let me now try one with 0.8 for example, and you should see that it would be less similar to the original image. So, I'm going to set the denoising strength to 0.8, press OK, and then run this again. And notice it's not starting from the beginning, it's starting from here because this is the last node where we changed the settings. So, again, this makes ComfyUI very efficient. And now it's decoding it, and you can see for this one, it's more different compared to the original image. So, basically, in a nutshell, this denoise value determines how similar do you want your output image to be compared with your uploaded image.
All right, now remember how I said there's a much better way to upscale images using image-to-image? Well, now that we've gone over image-to-image and specifically the denoising strength setting, I can now show you a much better way to upscale images, and this is one of the best ways to.
Actually, upscale images. In order for this to work, we need to download another node. So let's go into Manager and then click on Custom Nodes Manager. And this time, we are going to search for Ultimate SD Upscale. So let's click install here. And after it has installed, it says we need to restart ComfyUI. So let's click that.
All right, so let's start with the default text-to-image workflow, and I'm going to show you how to set this up. So instead of the K Sampler, this new node is basically going to replace the K Sampler. So let me hold Control and then select all three of these nodes and then click delete. And then I'm going to drag a connector out here and then search for Ultimate SD Upscale. And then we are also going to connect the negative prompt here, we're going to connect the VAE here, and also connect the model here.
This time, we are not going to use a latent image. So I'm going to delete that. And then instead, we need to upload an image. So I'm going to drag a connector out from Image and then select Load Image. And then I'm going to choose this image. Now, this is a 512x512 image that we generated previously. You can see it's very blurry. And then for the prompt, we are going to set this to the same prompt that we used before, which is "medieval warrior 8K masterpiece ultra detailed realistic". And then for the next negative prompt, again, same as before, we are going to type in "anime 2D cartoon painting watermark".
All right, so we are almost good to go. One last thing is we need to select an upscale model. So let's drag this connector out and then select the one and only Upscale Model Loader. And this will pull up the upscaler models that we've downloaded previously in this tutorial. So let's go with 4X Ultra Sharp.
And now, what this Ultimate Upscale does is called tiled upscaling. So let's say we upscale this by two. What this is actually going to do is break this into four sections, so it's 2x2, and then it's going to apply image-to-image for each quadrant. So it's going to generate image-to-image for this one, and then image-to-image for this one, and then this one, and then this one. And then it's going to stitch all four quadrants together to give you your upscaled image.
And now the trick is, if you use image-to-image with a low denoising strength, again, this is how much of your original image do you want to retain, then this method is actually a lot better than just upscaling with 4X Ultra Sharp. So again, the key here is to set a denoising strength to a relatively low value. I think 0.2 is a good start, or you could even go with 0.15.
All right, so all of these settings we've gone over before. Tile width and tile height just refer to the dimensions of each one of these quadrants. So this one would be 512x512, this one would also be 512x512, or whatever you set here. Usually, you would just set this to the width and height of your original image. And then mask blur and tile padding, this just refers to how well these tiles blend together after they are glued back together to form your final image. So I just tend to leave it at the default, but feel free to play around with these settings if for some reason you get a very obvious line in between tiles.
And then that's pretty much it. The final step is to drag this connector out, and then I'm going to select Preview Image. All right, so let's click Q prompt and I'll show you what that gives us.
All right, so if you now compare these two images, you can see that this is a lot more detailed, and the details are a lot finer. Now, previously, we did a clean 4X Ultra Sharp upscale, which looks like this, right? The details are not great. You can see the hair doesn't really look like hair, same with his facial hair, same with his face. It still remains blurry. Now, this is 4 * 512, so this is 2048 * 2048, right? Now, this is only 1024 * 1024 since we upscaled it two times. So to give you an apples-to-apples comparison, we can either set this value to four, which would upscale it four times, or let me show you another trick you can do to upscale this further. And this really shows the versatility of ComfyUI. You can literally just customize the workflow to whatever you want.
So let's leave this to two, and then get rid of this. And then we can actually plug in another Ultimate Upscale here. So I'm going to click search and then search Ultimate again. And then for model, we can drag model here. Positive prompt, we can use the same one. Negative prompt, we can also use the same one. VAE, we can drag that from the model. And then upscale model, we can also use the same one.
Or a better way to do this is, let me just select this node and delete it again. Is to click this, press Ctrl+V, and then over somewhere here, press Ctrl+Shift+V, and it would automatically link everything that is linked from the node that you are copying. Now, we don't want this original image to be linked here, so let me get rid of this linkage. And instead, we just want to pass our 2x upscaled image to here to upscale by 2x again.
All right, now the only thing that we need to change is because this image that is being passed here is now 1024x1024, we should set the tile width here to 1024 by 1024. All right, so now if we drag this out and select Preview Image, let me run this and I'll show you the insane quality that this can generate compared to if we just did a normal upscale method like this.
All right, so here is our preview image. Let me just save this first. And then I'm going to pull both of these side by side. On the left, this is only using the 4X Ultra Sharp to upscale my image of 512x512 to four times. And then this one is using Ultimate Upscale to upscale my image four times. Now, notice the insane difference. If I zoom in on this guy's face and zoom in on this guy's face, notice how much more details this Ultimate Upscaler is able to generate. His facial hair, his eyebrows, his eyes are super detailed. The lighting on his nose also super detailed. Whereas for this one on the left, even though it's the same resolution, his face is just blurry, and his facial hair does not look realistic. His eyebrows, his nose are super blurry. And then same with the crown here, you can see the details are really lacking in this left photo. Whereas for this one with the Ultimate Upscale, everything just looks super sharp and crisp.
Now, notice that because we are using image-to-image with a denoising strength of 0.2, there are going to be subtle differences from the original image. So, for example, you can see this dude is looking straight at the camera now, whereas this guy is looking slightly to the right. And that's because it doesn't just take the original image and upscale it. So if you really want to retain 100% of your original image, then I think this method is better. However, if you're okay with changing some of the details to get a much sharper and more detailed image, then Ultimate Upscaler is one of the best options out there.
So yeah, at least for now, Ultimate Upscale, or basically this is tiled upscaling, this is one of the best methods to upscale images. And it basically takes your image and breaks it down into tiles, and then for each tile, it does image-to-image, but with a very low denoising strength so that it retains most of the original image, but it just adds more detail to that image. And then it glues all these tiles back together to give you your upscaled image. This is one of the best upscaling methods out there right now.
So that covers upscaling. Next, let's move on to some more complex stuff.
All right, next I'm going to show you how to use ControlNet in ComfyUI. Now, what is ControlNet and why do we need to use it? Basically, it's a tool to really help you customize your image. If you're serious about image generation, you got to learn ControlNet. So, for example, you can really control the pose of your generation by using an OpenPose preprocessor. So you would upload a pose like this, and I'll show you how to do that in a second, and all your generations would all align with this pose. So here's another example, and you can see all these generations follow this pose to some extent. And you can also adjust, well, how much do you want your generation to follow your pose. Here's another example. It's a very powerful tool.
Instead of pose, you can also create a depth map and use that as a reference. So, for example, if you upload this depth map, all your generations would have the same depth map to some extent, including the lights, including the laptop. This is a really powerful way for you to control what objects show up in what areas in your image. And then here's yet another example. This is especially useful if you have multiple characters and the scene is very complex, but you really want to control where those characters are in the scene. Then again, this is a great tool to give you more control over those settings.
Instead of a depth map, you can also upload something called a Canny preprocessor, which is basically just lines. And you can see your generations would follow this Canny image. Here is another example. So ControlNet is very versatile. There are so many things you can do with this. Something that's very similar to Canny is line art. So it's essentially the same thing. You take an image and you break it down into line art, and then use that as a reference image for your future generations. So you can see all of these generations follow this line art to some extent. Here are some additional examples.
And if you want to generate anime, there's an even better line art preprocessor called Anime Line Art, and this is more optimized for anime. So as you can see here, if you want the same pose, the same outfit for your character, but maybe you want different colors, different backgrounds, well, you can use this option to generate those images. Here's another example.
And then similar to line art, there's also scribble, where you can draw in some lines, and then that would also influence your generations to some degree, as you can see in these examples. This is also a good one. So you can take any image and break down that image into different segments and use that as a reference. And you can see all your future generations would also follow the guidance of this image. Here's another complicated example, and you can see it's able to control for all these objects very nicely. You can see with this segment preprocessor, you're able to control the location of all these people, all these objects very precisely in your image.
All right, so let's jump right in. How do we use it? So first of all, let's start with a very simple text-to-image workflow. This is just your positive prompt, negative prompt, you're taking in an image of random noise, you're plugging it through this K Sampler, it's going to decode it and give you your final image. Now, let's start with adding ControlNet.
First of all, in this Manager section, let's click on that, and then we'll click on Model Manager. And then we'll search for ControlNet Unit. This is the newest ControlNet model, and it basically includes all of these options that we just talked about. So you don't need to go in and install all of those preprocessors separately. So let's go ahead and install this. Note that it is 2.5 GB, so depending on the speed of your internet, this might take a few minutes to download.
All right, after we've installed it, note that we need to click the refresh button. So let's click close here, and then close, and then click refresh.
All right, so how do we use ControlNet? It might be not intuitive, but we don't actually link ControlNet to the latent image. We actually link it to the positive prompt. So let me select this node and delete it first, and then I'm going to move up here for a bit. So let's drag this connector out and then search for Apply ControlNet. So again, it's not intuitive, but ControlNet is actually applied to the positive prompt before going into the K Sampler. So we are going to drag this conditioning connector back into the positive connector of the K Sampler.
All right, and then the next step is we need to select a ControlNet model. So let's drag this out, and then it's just the first option here, which is ControlNet Loader. So if that ControlNet Unit was the only thing you've installed, this is the only model you should see. If not, you can click on this, and it would have a dropdown of all the compatible ControlNet models that you've downloaded. But again, this is the only one you need. This is the newest one, and it contains all of these options for you, so you don't have to go ahead and download each of them separately.
All right, so the next step is the image. First of all, I'm going to double click here and then type in Load Image and then select this node. And then let's say I want to specify a certain pose for my generation. So I'm going to upload this image, and I want whatever I generate down here to follow this pose. So what we need to do is actually load this image to a preprocessor.
Now, in order to do that, we need to download another node. And you know, this is kind of messy. I wish we can just merge all of these nodes together into one node just to keep things cleaner, but anyways, it is what it is. Let's click on Manager and then click on Custom Nodes Manager. And then we'll search for Art Venture. And then you should see this one, ComfyUI Art Venture. Let's click on this to install it. And I'll show you what this node does in a second.
All right, so after we've installed this node, it says we need to restart ComfyUI. So let's click on restart and then click OK.
All right, so we are back after the restart. Everything is still here. So what we need to do is link our uploaded image to a preprocessor, which is the node we just installed. So let me drag this out, and then I'm going to search for ControlNet Preprocessor. And we should see this AV ControlNet Preprocessor. So let's click on this.
And then why did we install this instead of all the other options we could choose from? Because this node allows you to select from a lot of different options. So, for example, we can choose SDXL, which is what we are using, right? Our checkpoint is SDXL. And so the reason why we downloaded this node in particular is because it contains all of the preprocessors you need all in one node. So basically, you can select things like OpenPose, or Depth, or Canny, or Line Art, or all of these other examples that I showed you previously.
So let's start with the simple one. Let's start with OpenPose. So I need to convert this image into an OpenPose image. And then for the SD version, we are going with SDXL. And then resolution, let's set this to 768. And then let's set the width and height of our final image to 768 as well. Note that for SDXL, it's actually best to use 1024x1024, and it doesn't have to be square, but just to make our generation faster, let's go with 768. And then finally, we just need to drag this pre-processed image into the Image node of ControlNet. And then here, the strength is, well, how strong of an influence do you want this ControlNet, or basically this pose, to influence your final image? So 1 is 100%, 0 is 0%. If you set this to zero, then you're basically not using ControlNet at all. And so let's set this to something like 0.8, for example, and see what that gives us.
All right, so just a quick summary, how you would use ControlNet is it actually goes in between the positive prompt and the K Sampler. And then for ControlNet, what you need to do is upload an initial image, and you need to process that initial image into whatever preprocessor you select, and then it would turn that processed image and add it to the ControlNet. So actually, what I'm going to do, just to show you what this actually looks like, is I'm going to drag a connector out here and then click on Preview Image. So you'll see what this OpenPose image looks like. And then what I'm going to do actually is hold Control and select all these nodes and also select this one and then press Ctrl+B to bypass them, so that we are only running this. I want to show you what these steps actually do. So let me click Q prompt. And note that the first time this loads, it might take a while because you can see that it's downloading this OpenPose model from Hugging Face.
All right, so after everything is finished downloading, you can see the preview image here. So basically, this preprocessor is converting our uploaded image into this pose image and then feeding this into ControlNet, and this would influence the pose of our new image. So I'm going to hold down Control again and select everything, press Ctrl+B to unbypass, and then this time, instead of a medieval warrior, let's try a princess, arctic tundra, snowing. All right, and then let's click Q prompt.
Perfect. So if I drag my Load Image next to this final image, you can see that this princess is following the pose of my uploaded image to some extent. And if we want to follow her pose completely, then we would set this value, or this ControlNet strength, to one. So that's one example of how to use ControlNet to control the composition of your image.
All right, here's another example. So let's say I want an image similar to this composition of mountains, but I want this to be sunset instead of this lighting. So I can use a ControlNet. I'll upload this image, and then instead of OpenPose, I would select something like Canny, or you can also select Depth if you want, you could also select Line Art, it really doesn't matter, it really depends on your use case. And then we leave everything else the same. And here's just a preview image. So after the preprocessor, it looks like this. This is what the Canny preprocessor does.
All right, so we plug this into ControlNet, and then this time for the positive prompt, I just put in "mountains and sunset" and then all these other keywords. And then our final image looks like this. So again, if I drag my original uploaded image onto here, just for a comparison, you can see that our final image matches the shape of these mountains to some degree, but now it's sunset instead of this lighting. How cool is that?
Basically, there are so many different options you can choose from for ControlNet. Feel free to just play around with all of these preprocessors. There are just so many different preprocessor options you can choose in ControlNet to really give you maximum control over the composition of your image. So that sums it up for ControlNet. If you run into any errors or issues, just let me know in the comments below, and I'll try to help you troubleshoot as much as possible. But it should be fairly easy to install everything in just one click using this Manager button.
All right, next I'm going to show you how to install and use external AI tools on ComfyUI. The awesome thing about ComfyUI is that it supports a wide range of other open-source AI tools. For example, there's a ComfyUI node for Mimic Motion, which allows you to create dancing videos from a single photo. Or there's another ComfyUI node for Toon Crafter, which is a powerful tool for generating anime scenes. If you're not familiar with Toon Crafter, check out this video where I did a deep dive on how to install and use it. But basically, you just need to enter in a start frame and an end frame, and this tool will fill in an AI animation in between those two frames. Plus, there's another ComfyUI node for another tool called Live Portrait. In this tool, you basically input an image of a face, and then you input a video of another face talking or doing some expressions, and it can map those expressions onto your input image. If you want to learn more about Live Portrait, check out this video.
Anyways, today I'm going to show you the process of installing and using one of these AI tools in ComfyUI. For us, we're going to use this tool called Instant ID. At its core, it's basically a face swap that takes in a reference image of a person's face and then maps it onto your generation. So there are a handful of face swap tools you can use for Stable Diffusion, such as ReActor or Reactor, or this one, Instant ID, which I find to be the most realistic and best fidelity. So here's the original Instant ID page. You can see that it is really good for face swapping. As you can see, here's Taylor Swift, here's some Chinese actress. This looks really good. It really does preserve the details of that person's face, even across all these different styles of generations, like it doesn't have to be realistic. You can also do face swap for painting or drawings, and it even works with different angles. So even though you just have one image of Taylor Swift's face, this AI is able to miraculously kind of guess what that face would look like at these different angles.
So anyways, let's jump right into how we would set this up. On the GitHub page, which is called ComfyUI Instant ID, I'll link to this in the description below. If you scroll down a bit, here are the installation instructions. So you can either download or Git clone this repo into the custom nodes directory, or use the Manager. Of course, since we have Manager installed, we're just going to use the Manager.
So going back to our ComfyUI instance, I'm going to click on Manager and then click on Custom Nodes Manager. And then search for Instant ID. And then we're going to go with this one, the ComfyUI Instant ID Native Support. And the nice thing about this one, as it says in the description, is that it implements Instant ID natively and fully integrates with ComfyUI. So let's click on install. And then if you open up the command prompt while this is installing, you can see that right now it's cloning the repo, it's downloading all the files.
All right, so after that, it says we need to restart ComfyUI. So let's do that. I'm going to click on restart and then click OK. And then after clicking on restart, you can see in the CMD window that it's actually installing some additional dependencies, such as InsightFace. So depending on the speed of your internet connection, this might take a while to download.
All right, you can see now it's downloading ONNX Runtime GPU.
All right, so we've installed the nodes, we've installed InsightFace and ONNX Runtime, but we are not done yet. So next, we need to download these InsightFace models. So one of them is called Antelope V2, which we can download here. So it seems to be a zip file on Google Drive. I'm going to click download, and then you can download that anywhere. Once it's finished downloading, open up the zip file, and then it says we need to move this into ComfyUI/models/InsightFace/models. So let's go into our ComfyUI folder, and then in models, we need to create a new folder and call that InsightFace, and then within InsightFace, we need to create another folder called models. So I'm going to create a new folder, models, and then within models, we should have this Antelope V2 folder. So I'm going to just drag and drop this into here.
All right, so let me exit out of this, and you can delete the zip file afterwards. And then we also need another one. So we need this main model, which can be downloaded from Hugging Face, and then placed into this directory. And it says you also need a ControlNet, and you need to place it in the ComfyUI/ControlNet directory. Instead of downloading it from Hugging Face, there's a much easier way to download these, which is through the Manager.
So let me open Manager again, and then this time we are going to click on Model Manager. And then we are going to search for Instant ID. And you should see that down here, we have these two options, the IP Adapter and ControlNet. And this is for cubic Instant ID, so it's for this repo. It's basically downloading this model, which is based on IP Adapter, and this ControlNet. So we are going to select both of these and then click on install. Now, note that this one is like 1.7 GB, this is 2.5 GB, so it's going to take a while to download.
All right, so once we have these two installed, let's click close, and then close again, and then let's click on refresh.
All right, so let's start again with our very basic text-to-image workflow. So we have a checkpoint here, we plug in these prompts, and then these prompts go through this K Sampler, which takes in a latent image, and then it decodes that image and it gives us our final image. Let me just adjust the placements of these to keep things more organized as we add the additional Instant ID nodes.
All right, so where the Instant ID node goes is actually in between the prompts and the K Sampler. So if I drag this out here and then click search, I will search for Instant ID, and I should see this Apply Instant ID. All right, so let's drag this over here. I'm going to hold Control and then hold Shift and drag these nodes over here to keep things more organized.
All right, so let's remove this, and let's remove this. So the negative prompt goes to here, and then our model goes here, and then the model is then connected back to this K Sampler. The positive is connected to the K Sampler, and then the negative is connected to the K Sampler. Yes, there's a lot of connections that needs to be made, and I wish this process was simpler, but it is what it is.
All right, and then we need to also input an image. So I'm going to drag this node out and then click on Load Image. And then I will click upload, and let me upload this image of Will Smith. And then as the GitHub specified, this also needs an Instant ID ControlNet. So let me drag this out and then let's select ControlNet Loader. And in your options, if you search for Instant ID, you should only see this one, InstantID/diffusion_pytorch_model. So let's select this for the ControlNet model. And then for InsightFace, we also need to drag this out and select Instant ID Face Analysis. And then for the provider, since I have a CUDA GPU, I will select CUDA. And then finally, for Instant ID, we also need to drag a connector out, and the only option we see here is Instant ID Model Loader. So let's select this. And by default, it should just be this one, IPAdapter.bin, which is the file that we installed.
All right, so after all these things are connected, let's also look at these values. So the weight is basically how important do you want this face to be, or how much influence do you want this face to be in your final image. And then for start and end, again, for Stable Diffusion, it basically takes a latent image of random noise, and through each step, like right here, we've set it to 20 steps, and so for each step, it removes a bit of noise until it gets to step 20. So for these start and end values, it's basically saying, well, at what step do you want to start applying this face swap, and at what step do you want to stop applying this face swap? So let's say if we set this to like 0.5, or halfway, basically this would stop applying the face swap at step 10, since we set the number of steps to 20. If we set this to 1, then it's going to apply the face swap at all steps, right from step zero to the last step.
All right, and then just a few more things. We need to tweak for the width, let's set this to 768. For the height, let's set this to 1024. And then for the positive prompt, let's say "policeman realistic 8K masterpiece ultra detailed". And then for the negative prompt, we can say "cartoon painting anime blurry watermark".
All right, and if all is good, we can click on Q prompt and see what that gives us. So right now, it's loading the checkpoint, and then it's going through these prompts, it's going through this Apply Instant ID, which will take in our loaded image of Will Smith and all these other variables, and then finally, it's going to output this image of Will Smith as a policeman. So here we go.
Now, face swapping isn't just for generating deepfakes, right, generating fake images of real people. You can also use face swaps to create consistent characters, right? If you want a certain character to have the same face throughout your video, or throughout your animation, or comic book, or whatever, face swap is a really good way to apply the same face to all your generations.
Let's try something else. So the awesome thing about Instant ID is it doesn't just work for realistic photos. So you could also set this to, let's get rid of painting, and then instead of realistic, let's set this to "watercolor painting". And then let's click Q prompt and see what that gives us.
All right, perfect. So now we have a watercolor painting kind of of Will Smith as a policeman. So I hope this gives you a good understanding of how to install these different custom nodes of external tools. And the awesome thing about ComfyUI is there are a lot of other tools that it can support. So, for example, there's also an AnimateDiff node for ComfyUI, which helps you generate animations from images. And inside the GitHub, it should give you a demonstration of what a workflow should look like. Another person has created a ComfyUI node for Toon Crafter, and this allows you to take in one image as the input frame and then another image as the final frame, and it would interpolate an animation in between these two frames. So again, here's the workflow for the Toon Crafter node. And yet another user has created a ComfyUI node for Live Portrait, and this basically allows you to take one input image and then one video of a person moving their heads and doing some strange expressions or talking, and it would animate that input photo with this person's movements. And here they shared with you how the workflow would look like in ComfyUI.
So again, there's just so many custom nodes that other users have created based on a lot of different external AI tools, and that's what makes ComfyUI awesome. And with this Manager, you can easily search for all the nodes out there. There are like tens of thousands of different custom nodes, depending on what you would like to do. And that's what makes ComfyUI the most powerful free and open-source image generator out there.
All right, finally, some of you might be wondering, well, how do I use ComfyUI with the newest models such as AuraFlow or Flux? Well, for AuraFlow, actually, everything is the same. This whole text-to-image workflow is the same, and the only thing you need to change is the checkpoint. You just need to change the checkpoint to this AuraFlow safetensors file. See this video on how to install and run AuraFlow with ComfyUI. And then for Flux, the workflow is quite similar. You just need to tweak a few things like the model and a CLIP loader, plus the K Sampler. See this video where I go in depth how to install and run Flux on ComfyUI. However, if you've watched this tutorial, you should understand the basics of how to use ComfyUI and how all these nodes work. So changing between Stable Diffusion and Flux and AuraFlow is actually very easy.
And for this tutorial, I mostly used SDXL because it's still the most mature platform out there. There are hundreds of models and LoRAs you can choose from, plus hundreds of plugins and tools such as ControlNet, and all of these only work with Stable Diffusion. Whereas for Flux, it's still quite new, so there aren't a lot of tools and workflows built from the open-source community yet. But once we do have more of these tools, let me know in the comments, and if you want me to do an updated tutorial just for Flux, I'd be happy to make one as well.
So that covers my tutorial for ComfyUI. Like I said in the beginning, this is free and open-source, plus you don't even need an Nvidia GPU to use this, and the installation is super easy. If you followed all the steps as outlined in this video, you should be able to get ComfyUI up and running on your computer. Now, in this tutorial, we covered a lot of different topics, from text-to-image to image-to-image to face swapping to a lot of different workflows. So if you get stuck or you hit any errors along the way, let me know in the comments below, and I'll try to help you troubleshoot as much as possible. As always, I will continue to look out for the newest and coolest AI tools to share with you. If you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, we built a site where you can find all the AI tools out there, as well as look for jobs in AI, machine learning, data science, and more. So check that out at AI-Search. Thanks for watching, and I'll see you in the next one.