Transcription
Can you tell which image was created using prompt engineering? Hint, hint, it's the better one.
If you found this video, chances are you've already tried AI image generators, but they never seem to give you what you want. Well, don't worry, because in the time it takes for you to watch this video, you will be a prompt pro.
But before we get started, I want to introduce you to Sammy here. He is, isn't he cute? But he could be a lot cuter. I created these images by inputting the simple prompt, "a dog sitting by a lake," into Midjourney. My intention for the image is to create an atmospheric background for a camping supply website. And building the possibilities are endless with AI image generators. On one side, I can write a prompt that sells the calm, peaceful tranquility of retired life in the wilderness. But change a few ingredients of the prompt, and now I have a creepy, haunted lake that still gives me nightmares to this day.
Nightmares aside, in this video, I'm going to use our boy Sammy here at my camping supply store to demonstrate for you why prompt engineering is so important when it comes to text-to-image AI. I'm not only going to show you the breakdown on what to include in your prompt to help AI work with you, I'm going to show you how to use resources to help you, such as Gemini and ChatGPT. What settings to look for in your AI generator so that you're set up for success the first time and many more. I'll be primarily using Midjourney today, but I'll leave in differences you might see in other apps like Dolly 3 and Stable Diffusion.
[Music]
Welcome back to Learn with Shopify. I'm your host, Rachel Fischer, and I will be summarizing so many articles, YouTube videos, and Reddit posts that I watched and read. So trust me, watch this video, it will save you so much time and turn you into a master prompt engineer. This is part three of a four-part series that we have created here on Learn with Shopify about prompt engineering. Part one is hosted by Michelle and provides you with a general overview about what is prompt engineering and how to use it in all fields. Part two is about text-to-text prompt engineering for copy, hosted by Bridget. And part four is about text-to-video prompt engineering, hosted by Charisma. If you want to watch our prompt engineering masterclass, you can find the video links in the description sections below. Let's do this.
What is prompt engineering, you ask? Well, prompt engineering is the skill of creating inputs for AI tools that produce optimal results. A good way that I like to look at it is the difference between making a box cake and baking a homemade cake. Both avenues give you cake, but one gives you a more high-quality product because the recipes are a little different. This video will be about creating prompts for text-to-image AI generators. We are going to be mostly using Midjourney. I do want to say there's two versions of Midjourney. The first one you can access through Discord, and the second one is on their website. I'm going to show you both throughout the video.
So, the prompt "a dog sitting by a lake" is our box cake example of a prompt. It gives us a picture that definitely qualifies as a dog sitting by a lake. But let's see what happens when we increase the quality of ingredients we are using. How do we do that? We do that by increasing the specificity, creativity, and details of our prompt. Let me just say, the amount of videos I watch that claim to tell me how to prompt engineer and just ended up providing me with a bunch of adjectives was one too many. Okay, I will talk about it a bit, but I will not make your ears bleed, and I will not waste 10 minutes of your life.
To summarize, what they all basically say is that AI image generators won't ask you what you want, they will just show you what you asked for. So you have to be descriptive and provide key details such as who, what your subject is, what they're doing, adjectives about your subject, even style, lighting, inspiration, etc. Now, I've already mentioned a few key details in that list, but let's break that down step by step and flesh out our box cake prompt example together.
The first step before writing a prompt is intention. Ask yourself, what is this image for and what do I want the audience to feel? The biggest problem people have when prompt engineering is lack of specificity. So knowing what you want going into it is already going to bring you that much closer to your goal in image. So let's say I want this image of a dog sitting by a lake for my website that sells camping equipment. Which, by the way, if you don't have a website yet, you can have one in minutes. Shopify makes it super easy to build an amazing website, and we even have a free trial on offer. So check that out, the link in the description below. When people look at Lil' Sammy enjoying that lake, I want them to feel enticed to join him. I want them to feel peaceful, at ease, serene, and long to be by that very same lake. Keeping that in mind, let's continue.
So, what is the subject of your image? For us, it's easy, a dog. How would you describe the subject? We can do better than just a dog, because Sammy is more than that. Hint, hint, a thesaurus will be your best friend. But the first steps are easy. I'm just going to pick a breed, such as a golden retriever, because I'm partial to them, because spiritually, I am one. How about a middle-aged, scruffy golden retriever? But let's not stop at his appearance. What feeling does he give off? Perhaps he looks serene, at ease, lazy, and cuddly. We are not going to include all these words for now, just the ones that depict what we want in the most accurate way.
Details are important, but it's equally as important to be concise. If you make your prompt overly complicated, a lot of generators may get confused. So with that in mind, I'm going to keep it simple and just start with "middle-aged" for now. What is the action? He is lying lazily by the lake, watching the birds, patiently waiting for his owner to join him. It's important to make your subject active. Even if Sammy is being a lazy good boy, it doesn't mean they have to be moving, but it does mean they have to interact with their environment in some way. By including action, it tells the AI how your subject interacts with other elements in the image and typically results in a more cohesive product. Again, the more concise detail, the better.
What does the lighting look like? If you want your image to be lit well, it better be lit. I'm sorry, I had to. A lot of people forget about how important lighting is and how easily it can transform a mediocre, flat image and elevate it into something outstanding. It affects the mood, atmosphere, texture, warmth, and focus of the image, which is why it must be included in your prompt if you want a high-quality image. So, since it's a camping scene, we definitely need a fire, which will provide that cozy ambiance we're looking for. But I also want a little more magic. What's more magical than golden hour at night? So let's say we have a middle-aged golden retriever laying by a calm lake, watching the birds fly across the sunset at golden hour. Just to the left of him sits a smoldering campfire, casting a warm glow across his fur. Oh, I can smell the marshmallows already. I can also add detail about whether he is back-lit by the sun in soft focus to diffuse it, or in harsh light if I want maybe longer shadows in the image. The more you know about lighting, the more you can guide AI. Don't worry, we got you covered. Check out the description below for a link to a PDF of a quick list of lighting terms to use.
What lens am I seeing this image through? Lenses. This is something that people skip over all the time in their prompts. All images are seen through a certain set of lenses, which includes our eyes. By changing the lens, we can dramatically change the image. So let's take our box cake prompt. Look at what happens when we change the lens to a wide-angle lens or a fisheye lens. Completely different images. We also dropped a quick list in the description below of some lenses you can use. But if you really want to get crazy, you can put in your favorite make or model of a lens. Like, if you were doing a face portrait, try Panavision Primo 70mm. That lens is worth more than my car, and you can get it for free with AI. To be honest, a lot of things are worth more than my car at this point. For this prompt, I'm going to use a wide-angle lens because I want to capture the landscape, because of course, the whole reason people love to go camping is because of those beautiful views. I'll show you how I add that in towards the end.
All right, moving on. Angle. Angle is another ingredient you don't want to miss. Where is the camera in relationship to the subject matter? Right? Do we want to be looking up at the subject? At a low angle puts the character in a more dominant, hero position. Or do we want to be eye-level, which is like a natural perspective? Again, the quick list is in the description below. We got you. I'm going to be using an eye-level because I want the viewer to feel like it's their campsite that they are looking at.
What style do I want this image to be? Animated, or in a comic style, or do I want it to be cinematic, or hyperrealistic? Do I want this to be a digital image, or an oil painting? When it comes to style, feel free to use specific medium terms like I mentioned before. But also remember that a lot of AI models are trained from images on the internet. So the best way to get it to understand what you want is to include references it would understand. So feel free to include references to Van Gogh, or Da Vinci, or Monet, for example. Stable Diffusion not only loves descriptive words, but has a feature where you can upload a photo as a reference with your prompt. However, make sure it's an image that is in the public domain or is one you own. The same thing goes with D3 with the additional caveat that they will decline a request for a reference that refers to an artist who is currently living. Ethically, I would say avoiding terms and references to artists living and working is a good rule of thumb to go by when using AI image generators, especially if you plan on using your image for commercial use. You don't really want to get into copyright trouble. But AI image generators are great for using a pinch, but we also want to respect work created by real-life artists.
Finally, don't forget to include your desired aspect ratio and resolution of your image. Aspect ratio affects the height and width and is important so that you don't accidentally create a photo that is too small for where you need it and therefore becomes like a blurry mess when you blow it up somewhere else. Similarly, resolution affects the quality and detail of the image. The higher the resolution, the more detail. The lower the resolution, the less detail you get. The idea. Now, in Midjourney, the first variation of your picture is a smaller resolution. Once you find the one you like, you can upscale it to a higher resolution. Pretty handy. Once again, we have a breakdown of some aspect ratio prompts in the description below. Also, fun thing about Midjourney too is after you've picked your photo, you can also expand it. So you can either have it expand the view so it can become bigger, or you can shrink it down. So if you want to like narrow in on the dog, you can do that, or if you want a wider view of the landscape, you can do that too.
Having answered all those questions to the best of our ability, let's see what our prompt finally looks like. "A middle-aged golden retriever lying by a calm lake, watching the birds fly by across the sunset at golden hour. Just to the left of him sits a smoldering campfire, casting a warm glow across his fur. A green camping tent sits in the foreground by the fire. A hyperrealistic, wide-angle, comma, eye-level perspective. Period. D-AR 16x9." Let's see what happens when we put this into Midjourney. Already, there's a huge difference in detail, as you can see. We have a lot more going on in this image, but our prompt isn't quite perfect yet. Also, if you're wondering where all of your images are going, Midjourney puts all your outputs into your own web page with a reference code and prompts. So if you're ever wondering what your prompt was for that really great image you liked way long ago, it's all there.
Chances are, if you found this video, you've wasted time with bad prompts and spent all your available credits on your image generator of choice. Well, friend, I have a great resource for you, and you can thank me later. And that is the website Lexica. Lexica will allow you to view unlimited images and prompts without using up any free credits. So, for instance, let's input our box cake prompt: "a dog laying by a lake." Already, you can see tons of images showing this, all ranging in styles, details, lighting, etc. Hover over when you like, and it will give you the prompt that was used to generate that image. That way, you can see what kind of details have been included and how it was written, which is helpful if you're learning. Additionally, the prompts themselves may feature terms referring to style, texture, and lighting that you can click on, which will then bring you to another page of images specific to that term, which is great to learn about lighting and all that stuff. You can also directly copy the prompt used and apply it to any other AI generator. It may not create the exact same image, but especially if you find one you like on Lexica, it could bring you that much closer to what you're looking for. So put that in your back pocket.
But now that we understand the breakdown of what makes a good prompt, let's see what happens when we ask ChatGPT and Google Gemini to do the work for us. Because after all, who else should we ask to help us craft the perfect prompt for AI generators than AI? So I straight-up asked ChatGPT to create a prompt for ChatGPT that would tell it to work for me as a prompt engineer. The initial result was a little indirect, but it did end up working a bit. However, I tweaked it here and there, and I used it for both Gemini and ChatGPT. So here's what you can type in: "You will act as my professional prompt engineer for the AI generator software Midjourney, or Dolly, or Stable Diffusion, whatever you want. I'm creating an image for my business website, Instagram page, and I need an image that will capture beauty, peace, calm, whatever kind of adjectives you want to put in there. Here is the initial prompt you're going to put in the box cake prompt we talked about. I would like you to provide three variations, all high resolution, realistic, set in a wide-angle lens, or whatever other preferences you have. Just don't forget to include aspect ratio."
Okay, I did this twice with each avenue, once with our box cake example and once with our homemade cake example. Okay, so side by side, I think I like the Gemini responses a little better, just because, I mean, look at the way it brings everything down for you. Yeah. But let's start with the prompts we got when we asked ChatGPT and Gemini to enhance the box cake prompt we used. Okay, so "the dog laying by a lake." For the ChatGPT one, I'm going to delete this variation and just start the prompt at "a dog lounging," because again, we want to keep it active at the beginning. Side note for all the prompts, though, the only thing I'm going to edit is the aspect ratio. So I changed it to D-AR 16 by 9 at the end of the prompt, so it will read better. That just reads better on Midjourney, which is what I'm using. Okay, now let's see the ChatGPT image. Okay, now the Gemini one. Okay, so side by side, these look pretty awesome. But let's see what happens when we take the prompt ChatGPT and Gemini enhanced from our fancier one, the one we made together. Let's see if that makes anything better. Is more detail better? Let's see. Okay, so great. Now let's compare the image at the beginning of the video created with our box cake example alone with the ones that the enhanced prompts from ChatGPT and Gemini. Okay, so as you can see, while the first image is kind of flat, it doesn't really tell a story, it doesn't have really any nuance. Even though the prompts themselves see a little technical finessing already, you can see the difference having a clear intention and detailed descriptions can have on the output of your image.
The reality is, even with a perfectly pristine prompt, you may still be running into some issues. Don't worry, it might not be the prompt's fault. You may be underutilizing some key features. Some common features that you should look for and play with are negative prompts, weight, and seed number. So just like how regular prompts tell AI what you want, negative prompts tell AI what you don't want in your image. Let's use our box cake example for instance. We want the image to be warm, serene, calm. So I can put words that are opposite to that, such as "cold," "ominous," "depressing," into the negative prompt, and AI will emphasize the positive qualities we mentioned in our initial prompt. So it plays off of opposites, essentially. Stable Diffusion has a separate field for negative prompts, but for Midjourney, use `--no` followed by what you don't want, and it will pretty much work the same. You can also highlight an area of the image you don't want or change something like add a hot air balloon, like we did here, because who doesn't like a hot air balloon?
Weight, prompt strength, or prompt adherence, they're all the same, they're just under different names. Determines how heavily the AI relies on the prompt to create the image. The lower the weight, the more creative AI will get. The higher, the more specific it will be to the idea. However, if you put that setting too high, you may lose some of the variety AI can provide you with, or it'll take it like too literally. If it's too low, it might leave out details, such as our boy Sammy, which we don't want. We love Sammy. So it's important to check where your prompt strength is set so that you're getting the desired output from your prompt. This is something you have to play around with a little bit in Midjourney, but I recommend about a little less than 3/4 of the way up for your first prompt. This gives AI enough room to play while also remaining true to what you're asking for.
If you find the images that you're getting just need like small little tweaks, then check out the seed number. A seed number is a series of numbers that AI attaches to the image it makes from all the values determined by your settings and the prompt. This allows for an infinite amount of variations from one single prompt. So if you use the same seed number with the same prompt and the same values, it will turn out the same image every time. That's how it's supposed to work. If you want to see more variation, or even just slight changes, having a look at the seed number is a great place to start. Copying and pasting it next to the prompt you use and then changing one number at a time is a great way to see all those little variations you want.
Now, the sign of a real professional prompt engineer is someone who uses natural language enhancements in their prompts. This specifically includes dashes and brackets. They add more nuance to your prompt by adding emphasis, indicating range, styles, and altogether make your prompt easier for the generator to read. So dashes are great to include in Midjourney, Dolly, and Stable Diffusion software prompts because they separate elements, making it easier for the AI to interpret. An instance you would use this would be something like "a quiet lakeside -- a lush forest in the background." So there I'm separating the image of the lake and the description of the forest so it doesn't get lumped together. Another way to use this would be right before your preferred aspect ratio, just like we did earlier, so `--ar 16 by 9` or whatever one you want to use. You can also use it at the end of your prompt for like `--wide angle` all that stuff.
Brackets can help communicate where you want the focus to be in an image or add emphasis to the use of certain color theme, noun, etc. So say I want the lake to be extra whimsical or serene. Let's put brackets around either one of those terms. This effect applies to Midjourney, Dolly 3, Stable Diffusion, but remember to use them sparingly. If you use too many, you might confuse the AI. It's like, I don't know, which one's more important?
Here are some specific examples you can play around with when using Midjourney. So say you found a prompt that gives you what you're looking for, but you want more options. Type `--repeat` space and insert number. Let's use seven, for example, and it will use the prompt to create seven different images. Input `--chaos` space number, the number being between 1 and 1,000, and you've chaos mode. This will diversify the variations it creates and give you some more avant-garde ideas. The lower the number, the less variation. The higher the number, you get the idea. It's the same old game. You can interchange the word chaos with words like "weird," or "stylized," or "creepy," or "raw," and they will pretty much do the same thing with consideration of the word you use. But if you really want to get wild, don't be afraid to combine the goats. I know, I know, it's crazy, but why not? Type in `--chaos` next to `--weird` and dive deep down into the rabbit hole and see just how endless the possibilities are.
With all that being said, let's now add some final enhancements. The first prompt we made together, all the way back at the beginning of the video, feels so long ago. "A middle-aged golden retriever lying by a calm lake, watching the birds fly across the sunset at [golden hour], just to the left of him, sits a smoldering campfire, casting a warm glow across his fur. A green camping tent sits in foreground by the fire. --eye level --hyperrealistic --wide angle lens --ar 16 by 9." And finally, at long last, here is the final image. So let's take the first one and compare them side by side. Now, this image isn't perfect. It's great, but it's not perfect. However, I just saved myself a ton of time by inputting a high-quality prompt the first time. From here, the distance between this image and the one I'm after is much smaller than when I started with our box cake example.
As generative AI continues to evolve, so will our ability to communicate with it. And hey, since you made it to the end of this video, you're already that much better at it. If you want to see more like this video, let us know in the comments and make sure to like and subscribe. We love having you here. So why don't you stay? Thanks for being here today. I've been your host, Rachel Fischer, and I'll see you next time.
[Music]