Transcription
Here is 6 years of prompt Engineering in just 53 minutes. I started working with AI back in 2019 using gpd2 and since then I built a number of successful service and Consulting businesses. The first that did $92,000 a month, the second that did $72,000 a month, and my current which just did $139,000 last month. So I know how to build prompts for business purposes.
And the goal of this video is just to dump my brain and give it to you. I want to give you everything that I know about prompt Engineering in as quick and as compressed a format as possible. If you don't know me, my name is Nick and my whole thing is cut in the fluff. So let's get into it.
So we're going to be doing this whiteboard style. I'm going to be covering both very deep foundational underpinnings of how LLMs work, and I'm also going to be covering some more tactical, actionable advice for beginners and novices. Um, so we're going to have a good blend of both.
The very first thing that I want to cover right off the bat, and this is this will immediately improve your ability to prompt engineer, is instead of using the consumer models, use the playground or workbench versions of these models. So if you're unfamiliar, to keep a to make a long story short, this is ChatGPT. This is the consumer model. This is what OpenAI, the big artificial intelligence company, is marketing to people and selling for a monthly subscription. Because this is a consumer model, they've made a bunch of decisions to try and optimize performance for the widest number of people. Those decisions include, they actually insert a bunch of stuff into your prompt that you can't see and that you don't know about. So if you really want to get good at prompt engineering, you need to stop using the consumer models. Okay? Stop using ChatGPT, stop using Claude, and instead start using their API playground or workbench models.
So this is ChatGPT. This is platform.openai.com/playground. Chat. And as we see here, there's a lot more that we can manipulate on the right hand side. We have a variety of uh different tools. We can select model types, we can select response formats, we can add functions, we can configure our models with randomness, temperature, max tokens, stop sequences, top P, frequency penalty, and presence penalties. If you don't know what all this stuff is, don't worry too much about it right now. In addition, you also get to insert your own system message, and then you have the choice of being able to add user or assistant messages too. So I guess the point I'm making, to make a long story short, is this is really where you get into the engineering side of prompt engineering. If all you're doing is communicating with the base, you know, Claude Haiku or Sonnet, um, or the ChatGPT model, you're just leaving a lot on the table. So right off the get-go, if you could take one thing out of this video, it's just move over to the API playground versions. You're going to have a lot more juice that you could squeeze.
The second thing I want to talk about is that model performance. A lot of people don't fully understand this, but model performance decreases with prompt length. So there's actually a hack that you can do to immediately boost the quality of all of your outputs, and that's just make your prompt shorter. This is a graph here that shows a variety of models and their abilities to reason, AKA do some task, over the length of the input text. Input text in this case is obviously just like your prompt. Okay?
So what are we seeing here? Just to break this down. Most people are probably more familiar with GPT-4, Gemini Pro, than they are with Mistral and stuff like that. So maybe just focus on the green and then the purple. And I mean, the green's the highest performance. Maybe we'll just focus on that. What we see happen basically is that when the input length is 250, accuracy is almost one. One means, you know, it gets basically everything right. So it's like 0.9 or something. Okay? And then the second that we add more, okay, we see like basically neutral or maybe a slight elevation or something between 250 to 500. But basically everything after that, we see a decrease, and it just continues decreasing the entire way through. Okay? So in this case, it's a mile decrease between 250 and 3,000. Right? I mean, if I could quantify this, probably be like 0.04 or something. 0.04, sorry, which is probably like a 4% decrease on the chain of thought reasoning. But if we're using like the basic model, GPT-4, this is the normal thing that you that you'll call, performance goes down almost 20%, right? Which is pretty crazy.
So what does this mean? Well, a quick and easy hack that you can do to improve the quality of your outputs is literally just make your prompts shorter. Now, you need to also consider that, um, you know, the more examples you provide an model, the more context you provide an model, it also tends to perform better. So what I'm, what I'm not telling you to do is to remove all of the rules and instructions that you're giving the model. Instead, what I'm telling you to do is take the same information and shrink it. Basically, improve the information density of the instructions that you're providing. Okay? This concept in English is called KISS, Keep It Simple, Stupid. And the reason why it's written like this is it's supposed to be kind of a joke, right? This just means Keep It Simple, Stupid. But the person that wrote it decided to write it in an extraordinarily verbose way, just to show you how silly it is to write in a way that's more complicated than normal.
So to make this actionable, I'm actually going to run you guys through a real example where I pair down and break down a very verbose prompt into basically something that says the same thing, just in like a tenth of the words. So I have my playground open over here. I've added a system message that says, "You're a helpful, intelligent assistant." This is just my go-to. And I have a very long and annoying prompt over here, all about content creation. I'm going to run you through it, but I'm going to do so using a tool called wordcounter.net. I paste it all in there, and I'm also going to open up another wordcounter.net so that I can edit this in, um, a subsequent page, you guys could see the impact on the length. So right now, this prompt is 674 words. Okay? For informational purposes, uh, a word is about 1.3 tokens or so. Uh, meaning that, uh, a token is about 0.7 words. This is probably somewhere around 800 or so tokens. Okay? If we go back to our graph here, you know, 800 or so tokens is already an almost 5% reduction in accuracy. So if we can get this to around 250 tokens instead of 800, if we can cut this by a third, realistically, we're gaining about 5% accuracy or quality right off the bat. And that's going to be our goal.
All right, so how would I actually go about pairing this down? Well, I'm going to save you guys the, um, bleeding eyes from reading this whole thing. Essentially, what I'm doing is I'm asking this thing to produce some content for me. We're doing it in a very verbose way. For instance, it says, "Prompt for high-quality, engaging, and insightful content creation." We could remove that completely. "Primary objective and overall goal of this content request." Well, really, primary objective is the same. I'm just going to say objective. It delivers 99% of the same meaning, but we do it in a tenth of the words. And in fact, we don't even need objective because we're about to tell the model what we want it to do, and it's sort of implicit in us talking to a model that obviously the objective is what I'm about to tell you. Okay? "The overarching aim of this content generation request is to produce an exceptionally well-structured, highly informative, deeply engaging, and action-oriented piece of content that lines perfectly with the expectations, needs, and desires of my specific target audience." Wow, does that sound super deep and well-thought-out? Well, it's not. Good prompts don't look like this. Good prompts are much, much simpler.
So remove this all and say, "Your task is to produce high-quality content. The final content should not only provide value in the form of information, but should also serve as a highly authoritative, well-researched, and compelling resource that the reader or viewer of this video based can rely on for accurate, insightful, and practical knowledge. High-quality, authoritative content. The content must be structured in a way that ensures optimal readability, maximum clarity, and logical flow, being an effortless for the audience to digest and implement the concepts being discussed. Should also avoid this part's funny, excessive fluff or unnecessary tangents that do not contribute to the reader's overall understanding of the subject matter." Your task is to produce high-quality, authoritative content that is readable, clear, and avoids excessive fluff. Awesome. So we have delivered the exact same message. We've done so in about 200 or 300 fewer words.
"Audience understand the people this content must resonate with." Let's get rid of that. "Your audience is entrepreneurs and business owners. It looks like individuals who actively engage running businesses, particularly those focused on automation." So entrepreneurs and business owners who care about automation, scalability, I don't know, that seems kind of vague to me, and operational efficiency. I'll say automation and operations. "You're also writing for technical professionals." Let's say consultants, because odds are they are probably technical professionals if they are among the first group. And then "time-conscious, efficiency-driven individuals." I don't really know exactly what that means, but to me it sounds like fluff. So I'm going to get rid of it. "Given that the audience consists of highly capable, intelligent professionals, the tone, style, and depth of content must be appropriately aligned to match their expectations. Avoid condescending simplifications, but also ensure that even complex concepts are articulated in a way that remains accessible." So write accessibly. Perfect.
"Guidelines." So we could actually get rid of this whole section. "To ensure the generated content achieves the highest possible standard of quality, depth, and clarity, we just say guidelines." "Use a direct and Ampersand instead of 'and' tone of voice." "Avoid excessive casualness." Well, that's what direct and pragmatic means. "Unnecessary humor." Well, that's what direct and pragmatic means. "Storytelling." Well, this actually seems kind of necessary or new. So I'll say, "Avoid storytelling unless a story serves to reinforce a key insight." "Writing should be clear, concise, and free of unnecessary jargon." Well, I kind of already suggested that given my previous instruction, so I can just get rid of this completely. Okay? "The content should be defined, divided into well-defined sections, each marked by appropriate headings and subheadings." So use headings and subheadings. That's easy. "Use headings, subheadings, and bullet points, numbered lists, and clear formatting elements." Very cool. "Ur a logical progression from one section to the next." Well, if it's high-quality, authoritative content that's readable, clear, and avoids excessive fluff, odds are we're already going to have that.
"Depth, detail, and level of technical rigor." "The content must." We could just get rid of that completely and just say, "Go beyond surface-level insights and generic advice. It should provide unique, specific, and highly actionable information that differentiates it from lower-quality content available elsewhere." We can just get rid of that. That is embedded. This is also embedded, and this was already suggested over here, so we just get rid of that. Okay? Do you guys kind of see how this is working? We've already cut this down to 250-ish tokens from previously, uh, which what I believe was like closer to 800 or something like that. So just in me going through this prompt line by line and removing the excessive verbosity, it's called, I've already improved the quality of its outputs by approximately 5%. Obviously, this depends on the specific model that you're using, but this is one of the funnest games in prompt engineering. You're given a prompt that maybe a company has been using for their model, and your whole job is just take that thing, make it 30% shorter, and improve the quality of your outputs automatically by 5%. So you can see you can get very far with just a few minutes.
The third tip I'm going to give you is understand the different prompt types. There are system, user, and assistant prompts, and these prompts are basically just the go-to for for all, um, large language models at this point, probably going to remain so for a while. So if I go back to my ChatGPT playground example, not that page, this one over here, and and actually, if I just, let's make this a little bit more dynamic. Let's grab my prompt that I've painstakingly edited over here. Let me just add this plus button. Okay? Now, what I'm doing is I'm embedding a system prompt, which is at a high level, simply how the model identifies. So this is, "Who am I?" Right? Imagine you just turned on the AI robot. Okay? You just turned on this puppy, and it's asking, "Who am I?" All the system prompt is, is it's you answering that question. "Oh, you're a helpful, intelligent assistant." "Oh, you're an email writing assistant." "Oh, you're a customer help desk assistant." "Oh, you are a paperclip optimization maximizer who will certain and convert us all into aluminum, right?" You are, you are just providing it simple, straightforward instructions. A very general one that you can use for most purposes. "You're a helpful, intelligent assistant." That works just fine. We're saying that it's helpful because we want it to be helpful. We're also saying it's intelligent because we obviously want it to, like, select an intelligent, um, you know, an an intelligent example.
After that, though, what we have is we have the user prompt. So we always start with the system, okay? And I always recommend defining it explicitly. Then the user prompt. Okay? The user prompt is where you actually tell the model what it is you want it to do. And there's a specific structure within a user prompt that I will run you guys through later, but for now, just know that this is where you provide the actual instruction to the model. Okay?
Now, the next sort of prompt is what's called the assistant prompt. If I were hypothetically just to press this button right now, well, it's not hypothetical because I just did it. What's going to happen is I just told the model to write me an article about birds and bees, right? I didn't say write it about birds and bees, you know, I talked about doing it in the context of automation. So let's see what it does. "Birds and Bees: Key Insights and Analogies for Business Automation Operations." The birds and bees typically evokes, blah, blah, blah, blah. Okay, great. So we're writing, we're writing a bunch of stuff about birds and bees and how they can inform us as to how business works.
What I have now is we got the user prompt, the system prompt up here, the user prompt up here, but now we've actually inserted an additional assistant prompt. The output of the model, the thing that it just generated for us, is now actually an additional part of our prompt called the assistant. And most people think, well, this is just the output of the model, right? I can't really do anything with this. And here's where you're wrong. Well, what you can do is you can actually use this as an example to inform the model whether or not what it just did is good, and then use that as a template for future outputs. Okay? What I mean is, what if hypothetically, what if I said, "Fantastic work. Now I want you to do the same thing, but do it for an article on store delivery." I don't know, I'm just making stuff up at this point. What I'm doing now is I'm actually feeding in this prompt, this prompt, and this prompt, along with my n with my second user prompt. And what I'm doing is I'm implicitly reinforcing that it did a great job on the last piece, AKA I wanted to do the same thing, or I wanted to take a similar concept or similar approach in order to do, you know, the next piece too. And what we can see is, you know, we basically took the exact same structure and format, then we just copy and pasted it with this new, uh, you know, with this new, with this new idea, the store delivery, as opposed to birds and bees.
Now, the reason why I bring this up is because this is at core, at the core of all advanced prompting, which I'm going to get into now. But understanding how system, user, and assistant prompts work in concert, and this, this works across all major large language models at the time of this video, including Claude, including, you know, uh, DeepSeek, and so on and so forth. Understanding how that works under the hood is going to be critical in us developing higher-level and more effective prompts.
The fourth tip I want to talk about is to use one or few-shot prompting. Now, for those of you that don't know, one-shot, few-shot, all these terms that you may or may not have heard, these just refer to the number of examples in the prompt. If, if I run you guys through another little piece of, uh, of academia here, this was a study that was done a while ago on the performance of large language models across the number of examples they were given to fulfill a specific task. And what I want you guys to see here is that there are three categories: there's few-shot, one-shot, and then zero-shot. Zero-shot is blue, and it always performs the worst. Few-shot is orange, and it always performs the best. And right now, I just want you guys to take a look at just the distance in accuracy scores between this few-shot and one-shot. If we go all the way to 175B, that's about the size of, um, some of the simpler GPT models, you know, GPT-3, GPT-3.5, uh, because the study was a couple years ago and technology has improved drastically. My inclination is, you know, we're probably, I don't know if it's like 400B now. This trend is probably the same, although, you know, a lot of these models are now totally behind closed doors, you can't really see. So I think that this trend is even better than it is now. But, but like, let's just take a look at this. If we just extrapolate backwards, this is somewhere around 40% accuracy, right? This is almost 60% accuracy, meaning that there's an accuracy gap of 20% between zero-shot and few-shot. Okay? But that's not actually the important thing. What I want you guys to actually look at is notice how the dis, the difference in accuracy between zero-shot and one-shot is greater than the distance in accuracy or the difference in accuracy between one-shot and few-shot. This is like 10%. This is like 7%. This brings to a very important point that very few people know about, that's that if you insert just one example into your prompt, you will gain a massive disproportionate improvement in accuracy that out shadows if you were to add, let's say, 30 to your prompt. I think, um, this few-shot specific study, uh, was something like 20 and N is equal to 20 examples. So the distance between zero and one was like a time and a half the distance between 1 and 20. And if you think about it, considering now that we know what we know about the length of an input and then the, um, accuracy improvement as a result of shorter and shorter inputs, we can actually see a trend here. If we only insert one prompt or one example in our prompt, the accuracy will be as high as humanly possible because the prompt itself will be short, and then we're also going to get a massive boost in accuracy because of how we structured our examples here. Okay? So this takes me to an important point about user-assistant prompts, which I'll run you guys through in a moment, but essentially, I guess what I'm trying to say is, this is the sweet spot. This is like the gravy. This is where you want to, uh, sorry, not here, this is the gravy. You can tell I had to re-record this video twice now because I've been droning on for like two hours. My OBS just broke unfortunately, so I I finished it, now I got to come back and do it again. But whatever, maybe I'm just smoother now. Um, so, so this one-shot, plus, you know, what we know about keeping prompt length as short as humanly possible, this is really like the Goldilocks zone, and this is really where you want to be as often as possible. So my recommendation for anything mission-critical, always use at least one prompt, where at least one example, to coerce the model into providing you a more accurate result.
The fifth tip I want to provide is this notion of conversational engines versus knowledge engines. Okay? What do I mean by these two terms? Well, LLMs like ChatGPT and Claude are almost like a human being that's read a quadrillion books. Okay? So if we got a book here, this is my fantastic, super detailed, and amazing drawing of a book. So we could see here, I'm quite the artist. I'm like the Pablo Picasso. I think, um, not saying that Pablo Picasso wasn't amazing, he is. For those of the guys that know art, don't scold me. Uh, okay. So an LLM is just like a person that's read like a million books, and just like a person that's read a million books, if you went up to that, you're like, "Hey, tell me the boiling point of X-oxy-pyc-benzene or something." That person may roughly know what the boiling point is of things that are in a similar class, but they're not going to know exactly what that is 99.9% of the time. Just like a human being, LLMs often know roughly, approximately, what a right answer is, just because they've read so many books, they've picked up so many patterns between concepts, but they don't know exact facts. So LLMs, okay, are not knowledge engines. What they are are conversational engines. You don't use LLMs to extract knowledge for the most part, unless that knowledge is very old and very commonly known. Instead, what you do is use them to talk, use them to have conversations, use them to reason, or something introductory level reasoning, still, but still reasoning.
Now, contrast this with something like a database. If you, you know, I think most of us would probably use Google Sheets here, right? So you know, in Google Sheets, you have some sort of column heading, so XYZ A, and then you have like the data for the headings. So maybe X is like name or something, to make this a little clearer, and then, you know, the value here is Nick. Tables and and databases and encyclopedias, okay, these are knowledge engines. They essentially know facts, but you can't have conversations with them, right? Can't have a conversation with an encyclopedia. The cool part about LLMs really is when you hook up an LLM to a knowledge engine, and then you basically have it query that knowledge engine for some facts or some information. Okay? So LLMs on their own, like the GPT whatever series of models, these are conversational engines. They don't know facts, and you shouldn't really rely on them for facts. They're going to be very confident because it's just like, "I don't know, some dude at the bar that's like, 'Absolutely, man, I know all about medieval history from 1864 to 1899.'" Right? Medieval history? Good God. Uh, that's definitely not medieval history, folks. Clearly, I am a conversational engine, but they're not going to know the specific facts. They're going to be very confident in explaining these things to you, but unless you actually hook them up to some sort of external knowledge base, use something like, you know, RAG, which stands for Retrieval Augmented Generation, which is where you query some dataset using an LLM, you ask it to find you some data, return that data, and then have the LLM produce the response. Okay? Um, you're not actually going to get like any usable information for the most part. And even if you do, you know, I I wouldn't trust it. Maybe it's right like 70% of the time for most queries, but that 30% means a lot in business applications, right? You're never really going to get somebody to pay you a lot of money for some sort of business application if you're wrong 30% of the time.
Okay, speaking of being wrong, the sixth tip I want to give you is to use completely unambiguous language. What do I mean by unambiguous language? Well, unfortunately or fortunately, AI is extraordinarily creative. Okay? Because it's very creative, that means that when you ask AI something, the first time it'll give you a different answer than if you ask it the second time, and that'll be a different answer than if you ask it the third time, and a fourth time, and a fifth time, and a sixth time. Okay? If I could just give you guys a quick little example here, if this was like, um, if you played one of those games where like, I don't know, you have some zone that you need to hit or something. I don't know, maybe you guys have played this at the arcade or whatever, and and basically there's this little cursor and it goes ding, ding, ding, and you're supposed to just like stop it right when it's in the middle of some little Goldilocks zone. Well, large language models are very similar. If we have like this little Goldilocks zone here, I want you guys to hypothetically just think about this as this is the zone of responses that you want. Okay? This is, you know, the perfect output of your prompt. It's the, the best icebreaker, it's, you know, the the perfect report or or whatever. Basically, the way that large language models work are, you know, if you query it the first time, maybe you get a response over here. Okay? Query it the second time, maybe you get a response over here. Query it a third time, maybe you get a response over here. Maybe it's only when you query it like the fourth time that you actually get the response that you want. What you want to do is you want to, you want to, you want to minimize this as much as humanly possible. Okay? You want to get all of these as close to that little green zone as humanly possible. And I'll show you a practical way to do this in just a few tips. Um, but essentially, one of the simplest ways for you to do this is to be extraordinarily unambiguous with what you want. Okay?
So to make this practical, um, you know, a big, a big thing a lot of people will use AI for is it'll, you know, you'll ask it to do something like, "Produce me a report based on this data." This is what I would call a bad prompt. Do not do this. Instead, say, "List our five most popular products and write me a one-line or one-paragraph, let's say, description." This is a better prompt. And you get even better if you say, "Here's an example of a one-paragraph description for another product." Okay? So why is the first bad and the second good? Well, I guess it's good, so let's do a check mark. The first is bad because we're just saying "produce a report." Produce is extraordinarily ambiguous. A report is extraordinarily ambiguous. "Based on this data" is also extraordinarily ambiguous. Okay? Ideally, you'd say, "Here is the data." You don't just like make your own inferences about what to use of this data. Why is this good? Because it's specifically saying what we want. We want you to list five most popular products. Like, if you produce a report on some reports, it'll, it'll do 10 products, and other ones, it'll do two. Okay? It's just kind of up to, like, how the model is feeling at the time. But if you hardcode it in and say you're doing five products, in addition, you're writing me a one-paragraph description of each, and in addition, here's an example of the formatting. Well, now what you've done is you've constrained all possible outputs so that even if they're not perfect, okay, what they are is they're probably a lot closer to that Goldilocks zone of responses that you want, instead of being super crazy spread out. Now they're at least, like, somewhere in the realm of what you constitute a good response, somewhere in the realm of what a business might want.
My seventh tip is a very quick and easy hack. It's just use the term "Spartan" in your tone of voice. This is just one of those extraordinarily simple things that you can just do in in every prompt. Um, I just find the term Spartan is the perfect middle ground in terms of like being direct and being pragmatic, and then, you know, offering the model a little bit of flexibility. So, for instance, this guideline here, I would I would write, "Use a Spartan tone of voice." I I think before I had something different, "Use a direct or or assertive tone of voice." Just always say, "Use a Spartan tone of voice." Your answers are just going to be way, way better and way easier.
All right, let's make this data-driven. My eighth point is to iterate your prompts with data. Now, this is something of a more nuanced point, but it's how you get the highest quality prompts. I see lots of people in my communities, like Maker School, which is a day-by-day accountability program for people that want to start an automation business. It's something I've been running for, uh, quite a few months now. A lot of people in Maker School, the way that they'll do their prompt engineering is they'll spend all night coming up with some amazing prompt they think it's amazing, anyway, and then they'll run it once, and it'll deliver the most amazing output ever, and they're like, "Oh my God, I finally figured it out. I found the most amazing prompt. I'm just going to use this prompt for my business now." Okay? The thing is, do you remember earlier when I said that we have a range of possible responses in the model? Sometimes it's going to respond over here, other times it's going to respond over here, other times going to respond over here, right? Odds are, when you get a really good response in a model, like 70% of the time, it's just because the model happened to produce something that was in line with your expectations, but it wasn't guaranteed to do so. It just happened to do so. So if you want to get prompts that reliably and consistently produce outputs that are more in line with what you want, what you have to do is you have to test them. And then it's not enough just to test them once. You actually have to test them using what's called a Monte Carlo approach, where you throw a bunch of stuff at the wall. Okay? And then progressively make changes to get it closer and closer and closer to where you want.
So I mean, I don't, if you guys ever play darts, okay? But usually you have some, like, dartboard and it looks like this, and it's usually different colors and stuff, and they're nice points. And, you know, like your your goal, I think, is you want to throw it like, kind of over here. All right? So this is ideal. What usually happens is, you know, you're drunk as hell, you start off playing darts, you never played darts before, you throw one over here, throw another one over there, throw another one over there, throw one over there, and you throw another one over there at the wall and hit the back of Stacy's head or something. Um, every time that you iterate your prompt, what you want to happen is you just want the, I guess, total size of this to get smaller and more constrained. So that's, that's route number two. This is route number three. And then ultimately, what you want is you want like a very accurate model that just consistently delivers you results in this perfect zone. How you actually get there is you got to test and throw a bunch of stuff at the wall. Practically, what this usually means is you're going to do something like a Google Sheet, okay, just like this. You're going to have a prompt on the left-hand side, you're going to have the output in the middle, and then you're going to have another column that I just always call "good enough." And then basically, what you do is, okay, you grab your, um, ChatGPT playground in this case, and then what you do is you generate 10 examples of what you want it to do. Okay? So, for example, this was one on birds and bees. Let me just delete this other one. So this is, this is my my first article on birds and bees. So what I do is I actually paste it in here. Okay? So I pasted in the entire thing. Then I delete it all, and then I do it again. And usually what I'll do is I'll use a no-code model or no-code tool, something that just does things way faster, um, and then I'll just, like, do it in the background. But, you know, this is now, like, route number two. So I'm going to wait for it to finish its second output, just for the purpose of this demonstration. I won't do any more. But basically, after it's done with number two, I will copy all of this into another row. Uh, that's not right. Let's actually enter this. And then what I'll do is I grab my prompt, okay? Then I'll just stick it on the left-hand side on this prompt column. And just because this isn't very viewable, I'm just going to make this way smaller so we could actually, like, see and maybe interpret this. Okay? So I have two rows on my sheet here, right? What I'll do after that is I'll do, like, I'll do another, um, I don't know, I'll do, like, another 10 or something like that. Uh, and, you know, in my case, I'm just going to do two. But, um, in your case, do do 10. Okay? And what you do is you will read through every row and you'll say, "Hey, is this good enough for my business? Is this good enough for the thing that I was hard to do? This, this good enough for the use case that I'm building my automation for?" Maybe this article is good enough, maybe this one isn't. Okay? So we have one good enough out of two. Mathematically, one out of two is equal to 50%. Realistically, if I had like 20 of these or something like that, you know, what I would do is I would output 20, and then I'd go through and I'd read each of these and I'd ask myself, "Hey, this is good enough." And maybe 18 out of 20 are. So if, you know, 18 out of 20 are, then that's a total of 90%. And then what I would do is I would just test this against a different prompt. So now instead of me just using my gut feeling, instead of me just having it generate one thing and me being okay with that one thing, I have to do 20, and now I have like statistics. I basically have like some, some science or some data behind my answer. I'm like, "Oh, you know what? That prompt actually beats the other prompt 19 times out of 20." "That prompt actually beats the first prompt 13 times out of 20." "That prompt actually beats the other prompt 18 times out of 20." So what am I going to do? I'm going to pick the one that's 19 times out of 20. And in this way, you get more data-driven. And every time you make a change and you run this test again, you actually know, hey, I'm not just, like, throwing the most lucky bullseye on planet Earth. Okay? I'm actually statistically testing a big set of all of these outputs, and then I'm finding which ones have higher accuracy scores, which ones are, like, more statistically correlated to the center of that circle.
The ninth tip I want to give you guys is to define the output format explicitly. Okay? What do I mean by explicitly? Well, I mean like there are a lot of use cases where you're going to want something like a bulleted list. So actually, like, say, "Output a bulleted list." That's pretty easy, right? But there's a lot more than you can do than that. Um, you know, in instead, what we see a lot of people do nowadays is we output code blocks. So "Output JSON." If you don't know what JSON is, it's basically like a, um, specific JavaScript notation where you will generate some curly braces, a quote inside of this, uh, thing that's encapsulated by quotes, you have the variable or key name, then you have a colon, then you have some quotes around a value, then you have another curly bracket. This is a very specific format that now allows you to integrate your large language model with code servers, scripts, that sort of stuff. Um, you know, like CSVs. Okay? If you want a, a Google Sheet, and you want the Google Sheet to say something like, uh, I don't know, like, "Month," "Revenue," "Profit." You can actually have AI generate this. Don't just ask. Remember how earlier I I told you guys, um, that don't just tell it to produce a report? So don't just say, "Produce a sheet about financial data" or something. That is bad. Instead, what you wanted to do is you want to say, "Generate a CSV with month, revenue, and profit headings based off of the below data." And now, like, when it gives you an output, it's actually going to give you an output that you can just copy and paste directly into Google, Google Sheets, or Excel, or something. You know, you're going to have data that looks like this: Month, Revenue, Profit. Then it'll say, like, I don't know, January. Okay? It'll say, like, 13,000. It'll say, you know, 3,000. It's not going to have any commas because it's comma-separated, but I'll get into that later. Maybe February, it'll say 14,500, then it'll pull from some spreadsheet and grab this much for you, right? Like your data is going to be a lot more structured. So if you want data in a format, don't just have it generate a thing and then have you have to go copy and paste it into a spreadsheet or into your app or into your program later. Like, actually just ask it to do the exact thing that you want it to do right off the get-go. I'll show you the exact structure you need in order to get this to output reliably in a moment.
The 10th tip is to remove conflicting instructions. Now, this may sound pretty simple, but it's a lot bigger of an issue than I think a lot of people realize. Okay? What do I mean by conflicting instructions? Well, like a lot of the words that we use, we don't really think of, but they have directly opposed meanings. I'll give you a quick example. Remember how earlier we went through and we made our whole thing really short, our whole prompt really short, to take advantage of better accuracy? Well, a word I always see, and I've seen very often in Maker School, my 0-to-1 community, and then also Make Money With Make.com, which is a higher-level community, like a lot of prompt engineering threads, I see stuff like this: "detailed summary." I want you to think about this logically. Like, if you produce a summary of something, you are necessarily producing something that is simpler and smaller and and lighter and and easier. It's a, it's a compressed form of, you know, the thing that you are doing. You're basically going, you know, from something like complex to simple, right? That's what a summary is. But if you ask for something detailed, basically going from something simple to complex. So then if you're asking for a detailed summary, what you're basically doing is you're, it's like mathematically, it's like it's like this equals zero. Okay? It equals nothing. These two cancel each other out, and what you end up doing is you just end up needlessly increasing token count for no reason. So like when you say stuff like, "Produce an engaging article that is also very simple and straightforward," or "Produce, like, a very comprehensive article that's still easy for newcomers to understand," right? You're giving it conflicting instructions. Do you want it to produce something detailed, or do you want it to produce something that's that's not? Okay? Just eliminate that. It'll substantially reduce your token length, and then it'll also allow you to, you know, just get, get a little bit better at defining things and and being simple and treating LLMs as, I think that, you know, they realistically should be treated, which are tools that enable you to generate things, not like at our current level, some super nuanced being that's able to like really fully understand the full gradient of the decision that you've tasked it with.
The 11th tip I have for you is to learn JavaScript Object Notation, XML, then CSV. Let's actually start with XML. XML stands for eXtensible Markup Language. Okay? To make a long story short, all XML basically is, it's just a way to structure data. You remember how like back in grade school, um, you know, you would write a story and you write your name in the top left-hand corner, you write the date over here, you'd write the title over here, and then underneath here, maybe you write the story, right? And then you hand it in. Well, all XML, JSON, and CSV are just formats that allow you to embed what all of these different things are in a simple and easy and standardized way, um, for computers to understand, so that you can do cool things with them, like run, you know, cool programs and whatnot. So, you know, if this is my big paper, okay, this is a big paper that I wrote, won me first place in school, I got a $25 gift card to, uh, Denny's or something. Okay, this is equivalent to this. I have a little author tag where it says, "Nick S." Then I have a closing author tag. All right, this is now like universally understood that the author is Nick Surve. Right? Maybe the date. Okay, the date was February the 19th, and I close that tag off with date. Title, right? Well, the title was "Birds and Bees," right? And I close out the title. I could go on and on, but I think you guys understand this is a very particular formatting convention that allows you to define things. Okay? If I go, this one, this one's pretty cool because I actually used to, um, live and work in Surrey. So building ID, Siri, head office address. This is an address variable, city, city variable, province, province variable, country, country variable, location, right? You can embed data. If I were to instead write a paper, I'd say, "The Surrey head office is located at 7445 132nd Street, which is in Surrey, BC, Canada." I just want you guys to, like, see how much more, every time I say "Surrey," it activates my Siri. Isn't that funny? Um, I just want you guys to see how much more compressed this data is and how much more immediately understandable this data is, too. So that's one reason why XML is awesome.
JSON stands for JavaScript Object Notation. It's very similar, okay? The only difference is instead of the tags, the less than and greater than symbols that we used along with some of those backslashes, um, all this is is this is curly braces, quotes, characters, colons, and then more curly braces. Okay? Exact same thing, right? Uh, this would say author, and then this would say Nick Sarf, right? Exact same data. If you could see, um, a JSON example, you'll see data that looks just like this. So instead of me saying, "Hey, this guy's name is Richard. He's a 33-year-old man who loves biking, gaming, squash, and lives in Port Space Land. He's friends with Joe and Sarah and Michelle, who are all, you know, 30-ish somethings that live across the continent of the United States." Like, instead of me saying that, I just create a little JavaScript object where the name variable is equal to Richard, the age variable is equal to 33, and so on and so forth.
And then CSV. The reason I bring this up is because this is a little bit more of a nuance point. CSV is actually just the hyper-compressed version of all of this, where you don't have to use these tag characters, and you only have to mention the thing once. Okay? So what I mean is, okay, remember this author, date, and title? A CSV file would actually just look like this: author, no space, date, no space, title. And you'd have a little new line. Then here it would say, "Nick SV," then it would say, "Feb 19," then it would go, "Birds and Bees." Now, this, if you just, like, count up the number of characters, this is actually way less than all of this. And the cool thing is, when you do a new row, you don't have to repeat the key names. What you can do is you can
Actually, just, you know, um, write the, I don't know, write everything in what's called comma-separated value format. The comma here is the delimiter character. The issue with using this in large language models, um, is unfortunately, large language models often lose their sense of place, especially in longfall. So, if you're producing a CSV that has, I don't know, like several hundred, um, inputs, the large language model's ability to like remember that the date column always comes second on, I don't know, like row number 1,000 and three or something, its ability to know that like date corresponds to Y, it just tends to get muddled down, which is why you typically only use CSVs for for smaller applications with large language models.
All right, the 12th tip I want to give you guys is what I call my key prompt structure. So, I have developed probably over a thousand prompts now for businesses, my own businesses, other businesses I work with, people that I consult with, and so on and so forth, and this is what I do, um, on basically all of them. Okay, so feel free to copy this on your own, um, use this for all your own prompts. I'm going to show you what this looks like. First, we start with context, then I give it instructions. Okay, after that, I give it my output format. You guys just use this to scaffold your own. Then I give it rules, and then finally, I wrap it up with an example. So, let me show you guys a quick actual look at what a real prompt that has made me almost, um, $500,000 looks like. I'm going to show you guys this in a no-code tool. It's called Make.com. It's what I personally use and I teach a lot about it as well. Um, it is for a specific tool called Upwork. Upwork is a tool, uh, Upwork is a freelancing platform, or basically you can bid on jobs and stuff like that. And this prompt was written to allow me to take as input an Upwork job description, uh, and then have AI automatically process it, filter it, tell me if it's relevant to me, and then write me a one-line icebreaker that's customized to that job.
So, how this actually looks: I start off with a system prompt that says, "You're an intelligent admin that filters jobs." That's pretty simple, right? No real rocket science here. Then I have my user prompt, and the way that I want you guys to look at this is through that lens that I provided you earlier, okay? With context first, then instructions, then output format, then rules, and then some examples. So, let me break this down for you. This up here, this is all my context, okay? So, what do I mean? I say, "Hey, I'm an automation engineer that builds outreach systems, CRM systems, project management systems, no-code systems, and integrations, right?" I'm sure I could cut this down in hindsight now that I know a little bit more about prompt engineering. I built this thing out, um, I believe over a year ago now. It was one of the videos that actually made me go viral on YouTube. Then I say, "Below is a job description. Filter it for relevance. True or false, in JSON. Some of the platforms include use Airtable, ClickUp, ChatGPT, Make, Monday, Zapier, LinkedIn, Google Sheets. If relevant, write a short introductory icebreaker." Okay, but I'm not actually done yet. I then give it some example client projects. These are things that I've done. And then I tell it my format. Okay, so let's just be abundantly clear here. I provided a context over here, then I gave it some instructions over here. So, sorry, I think I misspoke earlier. This is my context. This is my instructions. Okay, after my context and instructions, what I do is I give it my output format, which is right over here. And then at the end, I also have some rules. Now, in this case, I wrote them as notes, but these are essentially my rules. Then finally, I actually give it some examples. And the way that you do this in Make.com and any other no-code tool is you will have your user prompt up here with all of those four bits that I showed you a moment ago, and then you have another user prompt with your first example, then another assistant prompt with your first result. And in my case, I did, I think, three or four shots. So, then I have a user with an example. This is an actual job description from Upwork. Then I have an assistant response with a result. Then another user prompt, assistant prompt, another user prompt, assistant prompt, another user prompt, assistant prompt, another user prompt, assistant prompt. I think in this case, looks like I had like seven or six or something like that. Uh, in hindsight, probably too much. And I was probably forced to do this just because models were a little bit less intelligent when I built this out. Nowadays, I might do like two, maybe.
Okay, I also also wanted to show like a wide range of possible jobs, which, um, obviously helped to perform. And this is why I have almost $500,000 in posted earnings on my profile because I use systems like this to be able to apply to large volumes of jobs. I've also helped a lot of other people and build out processes that involve things like this. So, um, in a nutshell, I almost always use this key prompt structure with some slight variations. Context is where you tell it what you want, like who you are and what you want. Instructions are where you outline specifically and say, "Your task is to do XYZ." Output format is where you say something like, "Return your results in JSON using this format." Rules where you say, "Hey, here's a quick list. I want you, you know, don't do this, do this, don't do this, do this." Then examples where you actually give it those user prompt or user assistant prompt pairs.
Okay, the next thing I want to talk about, tip number 13, is to use AI to generate examples for AI. What do I mean by that? Well, if we go back to that prompt that I showed you guys a moment ago, right? We have a lot of examples here. What you can do instead of you actually finding an example, you get, you can actually like create one yourself, right? So, I can actually go and I can create one. So, what do I mean by this? I could say, um, I'm actually going to go all the way down to the bottom and I actually want to generate my own little assistant prompt, right? So, maybe what I'm going to do now is, uh, I don't know. I'm just going to go into AI, paste this, and say, "I'm using this for training. Write me a similar training example as the above." This is just using a simple hotkey, um, option space on Mac where I can launch a ChatGPT instance. I should know that this is not the same as the, um, API that I'm using right now, but does a pretty good job, right? So, now I have this. And what am I going to do? I'm just going to right-click, run this module only, and it's actually going to go through and generate me a result using AI, right? Pretty simple and easy. Now, this was back in the day when you could not parse, uh, the JSON in the Make.com module. Anybody that is a little bit more experienced with Make.com is probably looking at this and being like, "Hey, why aren't you parsing this directly inside of the prompt?" That's just because the GPT-4-0613 model just didn't have access to do it. As you could see, there's no, there's no tool for me to do so. But basically, what we've done is we've outputted some JSON, JavaScript Object Notation, which I can then feed into this module here to parse automatically, then produce me a variable that I can access that looks kind of like this: reason, result, icebreaker.
The last tip I want to provide you guys is a pretty simple one, but it's to use the right model for the task. Okay? I see a lot of people nowadays using very simple models. So, here's how it basically works: simple models are cheap, complex models are more expensive. So, this is a gradient between simple and cheap, complex and expensive. The unfortunate thing is, like, 99% of people that I see in in Maker School, make money with Make.com, on YouTube, they're way too far on the simple and the cheap side of things. As of the time of this video, there, there are a bunch of models available. I don't really want to date this too much, but, you know, a bunch of these are like mini models. And so, what a lot of people will do is they'll just use the mini models because they think that it's saving them a lot of money. Well, unless you're running something like that is doing 5 million operations or executions a day, like, unless you're running some serious backend infra, token costs are so little that it doesn't actually make any sense not to just use like smart models most of the time. Obviously, there's some exceptions. But let me just run you through what the GPT-4-0 family models looks like. Okay, this $0.25 cents per million tokens of input, $10 per 1 million tokens of output. If we just like average them and say, I don't know, for the purposes of this, I'm just going to say it's five bucks per 1 million tokens combined. Okay, this Upwork RSS feed thing that I just did, if we look at the usage, okay, combined, it was 1,169. I would have to use one, I would have to do this 1,000 times to use $5. That's crazy. How many do I actually do a day? Like, 15. Okay, if we do the math on this, $5 divided by 1,000 means that for every run, I use 0.5. This is equivalent to, uh, 55.5 cents. So, basically, every two times I run this, it's 1 cent. It's, it's a penny. Every four times I run this, it's two. Every eight times I run this, it's four, right? It's such a marginal small amount of money that there's no reason. And this is, by the way, this is using a GPT-4-0613, and this is also using like a very, very under-optimized prompt. Well, maybe not very under-optimized, but pretty under-optimized prompt with what I know now. The reality is, you know, like, uh, any case that you could use GPT-4-0 mini, which, you know, is understandably much cheaper, but basically any use case you could use GPT-4-0 mini for, unless you're sending millions of tokens on a daily basis, just use the smarter model. The smarter model will eliminate like half of the problems that you didn't even know you had. And I recommend always just starting with a smarter model and then working your way down, as opposed to starting with a dumber model and then trying to work your way up. Just way easier to do that way. So, I mean, you know, $75 per a million tokens for 4.5, or $150 for a million tokens. Like, like, yeah, that's pretty expensive. And I could see that being a limiter. But the vast majority of these, uh, based multimodal models, anyway, as of the time of this recording, are hyper cheap and usually recommend just de-buying yourself, de-biasing yourself, and just trying to use the more expensive ones wherever possible because, at the end of the day, they're not really that expensive. My company, um, which is called One Second Copy, we still use a lot of tokens on a daily basis. I think we spent like, I don't know, five bucks last month. That's, that's a company, right? Makes several tens of thousands of dollars still. So, if you think about like the, the return on investment of these tokens, it, it's crazy. So, just make sure you use the right model for the task. Most of the time, that involves using a little bit smarter model.
All right, that's it for this video. Had a lot of fun recording it and re-recording it for all y'all. If you guys have any questions about this, just drop them down below. If you guys want me to make videos on a specific topic, then please, I love getting ideas and I'm inspired for the most part by people like you that actually take the time to leave comments down below saying, "Nick, can you do a video X, Y, Z? Nick, can you do a master class on prompt engineering?" So, this video is because somebody earlier on in my comments, a few weeks ago, asked me to do this, and I'm more than happy to do what you, what you want me to do as well. Otherwise, if you guys do me a big solid, anybody that's on the cusp of, you know, starting an automation agency or getting up and running with their own business that hasn't had a lot of experience in service companies or in any sort of automation scenario before, I'd highly recommend that you check out Maker School. It's my day-by-day accountability program, and you can get the link just in the description. I guide you through setting up essentially your own automation outfit, um, from complete scratch and bootstrap in the hell out of it while you're at it. And for anybody that maybe knows a little bit more about business, that already has an automation business, for instance, or somebody that runs, uh, a business in a very similar domain to automation, maybe a marketing agency, maybe some sort of advertising or creative business, check out Make Money With Make.com. It's my mid-level community that helps people that run businesses over $5 to $10,000 a month scale using my own products, systems, templates, and so on and so forth. Both of these communities have me in them every single day, coaching and guiding you guys through. I'm a very friendly and familiar face, you can hopefully tell from now, um, and yeah, I'd be, I'd be more than than happy to see you in there and then help you out. If you guys could do me a solid, like, subscribe, do all the fun YouTube stuff. I'll catch you on the next one. Cheers.