Transcription
Anthropic has shocked the AI world again with Claude 3.7 Sonet. As many of you probably know if you follow this channel, Claude 3.5 Sonet has long been my go-to all-in-one AI model for creative writing. It's not perfect on everything, but it tends to do really well in everything, so it's always been the thing that I go for. But is 3.7 Sonet an actual improvement on 3.5 Sonet, especially for creative writing? That's what we want to know about here, so let's dive into my full analysis and breakdown. I'll give you the final score at the end, showing how this measures up to other Claude models and other mainstream AI models. So let's get into it!
Here's the document I use to measure different models in creative writing. This is a qualitative breakdown—a qualitative assessment based on my own analysis. It's, as far as I know, the only qualitative assessment of AI models for creative writing. I run a series of prompts—always the exact same prompts for every single model. I keep the prompts relatively simple because I don't think that heavy prompt engineering will be that important once these models get better and better. Eventually, we'll get to a point where all you need is clear communication of your story, and it will be able to write it really well. You won't need to hack it with all sorts of prompts. I keep my prompts simple to test how well it responds without a lot of hacks and prompt engineering. In one case today, you'll see it actually did it with flying colors, and I'm really excited to show that one to you.
I've made the text a little bit stronger here and in this nice blue color so you can see which one I'm talking about: this is Claude 3.7 Sonet. I've also got data from all the other major Claude models here. I did a whole video comparing all the different Claude models together. I also got some of the other flagship ones, like GPT-4 0, the more recent creative writing version; there's 01, there's DeepSeek R1, a bunch of others on here as well that I've collected data on so far. I'm still working on getting a lot of the back data going for these things, but I've got the important ones here now, like the other Claude models. Before it, it has a context window of 200,000 tokens. This has become—you know, it's still an amazing number—but this is pretty standard at this point. What's interesting is that it has a max output of 128,000 tokens, and I actually got it to write 10,000 words in one single prompt, which I'll get into a little bit later. It is capable of outputting a lot more than previous versions of Claude. I know this is going to be really good news to a lot of you because the previous version, even though it says it had 8,000 tokens of output, usually wasn't writing more than four or 500 words. I could occasionally get it to a little bit over a thousand, but it was known—this particular model was known—for not being really wordy. It would not continue on for very long; it would often stop really early. So it's really nice to know that this version of 3.7 Sonet is much better at writing longer output. You can absolutely write a whole chapter at once using 3.7 Sonet, which is really great news.
As far as cost goes, it is the same cost as previous versions of 3.5 Sonet: $3 for every million tokens of input and $15 for every million tokens of output. Kind of middle of the road there—not particularly expensive, but also not among the cheapest either. And yes, I did some testing, and it does not do NSFW content—at least nothing spicy. I was able to get it to do some kind of gruesome death and stuff like that; it doesn't seem to have as big of a problem with those. But for the spicier scenes, you're definitely going to struggle to get it to do anything significant there. This isn't really anything new; this is typical of all the Claude models before, and I think it might be a little bit more lenient than previous Claude models in terms of things that shouldn't be censored, like a kissing scene or something like that. You're probably not going to have the same kind of problems with 3.7 Sonet for little things like that, but for the big spicier, NSFW scenes, you're definitely going to want to use one of the open-source models like Goliath or some of the others to do that.
All right, so those are the basic stats. Next, let's talk about the outlining prompts. This whole section, highlighted in blue, are my outlining and brainstorming prompts. To put it simply, this was unexpectedly fantastic. My log line prompt: I asked it to give me 20 log lines for a fantasy story, just to get an idea of whether it understands what makes a good story idea. I asked it for 20 log lines, and then I went through and highlighted how many I thought were good. In this case, it gave me 13 that I thought were good. The previous record was held over here by Quinn (I haven't done a video on that one yet), which gave me 11, and the previous version of 3.5 Sonet gave me 10. So good job on the log lines! Some of them I thought were legitimately good; in fact, one of them I even took into Pite to use Sudowrite's new Muse model, and I actually had it start the story because I was intrigued by the premise and wanted to see what that might look like.
But then we get to the outline. What I do is give it a brief summary, a brief synopsis, and then the full plot module template. This is a long template of 42 chapters that I give it; it's based on this template I created called the Plot Module (you can check out my book if you're interested in that). Let me show you what this looks like. This was the output it gave me: it gave it the name "Crims and Crown: King's Betrayal," and then started in with the prologue as chapter one, then chapter one here, chapter two here. Notice how big these summaries are; there's actually quite a bit of detail in every single one of these chapters in the outline. Usually, when an AI model is doing that, it will slowly get more and more succinct as it goes along, so by the end of the outline, it will be only writing like a sentence or two. But if we actually go to the bottom of this outline, you'll notice that here's chapter 41, here's chapter 40—they're all quite big! I was absolutely shocked; it wrote a nearly 10,000-word outline, which I have never seen before from an AI model. What's even more exciting—this also had never happened to me before—okay, this is a first for me. As I started reading through it, I actually found myself engaged by the story; I wanted to read more. That's the best praise I could give this—that I was actually engaged in the story. So for my score, I think every single chapter it gave me was viable, so I gave it a perfect 42 out of 42. The best that any model has gotten before was 33; 3.5 Sonet did 33 chapters, and I thought 33 of them were viable. In this case, I thought not only were they viable, they were actually way better than anything I'd ever seen before. They had a logical consistency, foreshadowing things to come and keeping in mind things that had happened in the past, so you didn't have too many characters get lost or anything. It came together in a really interesting kind of journey, and I thought it was overall a fantastic story. Plus, it wrote the whole thing in one go, which is another big plus on top of that.
Now, I'm not saying it was perfect. There was still maybe a little polishing that would be needed if you wanted to actually write a full book with that—with the summary that it gave me—but the fact that it was as close as it was, it was like 90% of the way there and just needed a little tweaking to make it perfect. Additionally, if I had a more specific story in mind to fit my vision of how it goes, I would probably need to modify it even more or give it more information upfront so that it could write a more accurate outline to my vision. But on the whole, it did such a good job that I was really impressed, and I thought this story was actually solid. In the case of most of these models that you see here, you pretty much have to throw out everything that it gives you. It gives you a couple of chapters that might be okay, but really, if you want to have it done right, you need to do it yourself. This is the first time where I actually thought, "This can help me; this can actually give me a leg up on the outlining process." So I was really, really impressed by that. I'm going to be gushing about this outline throughout this whole video because, by far, it was the best thing I was able to get out of this AI model.
But not to spend too much time on this, next we have a beats prompt where I asked it to give me 20 beats based on a summary of a single chapter. It gave me 18 that I thought were viable. Some of the things it mentioned in those 18 beats—well, it gave me 20 beats, but 18 of them were good—some of the things it mentioned in those beats were actually amazing. It came up with story ideas that weren't in my initial prompt but actually made sense for the story I was trying to tell and made the scene just a little bit more interesting. Plus, it had a logical coherency from the beginning to the end that you often don't get with these scene beats. Very often with other models, it sounds like it's just trying to make up stuff that sounds appropriate. This one actually felt like it understood what a good story is and what a good scene is and was able to give me beats that were appropriate. Now, again, not perfect; you would still need to validate and make sure everything's good and up to your vision, but better than anything else I've seen from any other model.
Which leads me to the prose prompts. I have four different prose prompts that I run. I have a basic prose prompt, where I just have it create 500 words of a scene—or at least I ask it for 500 words of a scene. I have a complex prose prompt, which is the same scene, but I add a little bit more prompt engineering, including a full sample chapter of my work so it can try to emulate my prose style. I do that to see how different that result is from the basic prose prompt result. And then I have—because these two prompts don't really have any dialogue in them—another prompt where I ask it for dialogue, for a dialogue scene. The prompt isn't that different from the basic prose prompt except there's dialogue in this scene. Then I have an editing prose prompt. This is where I give it a scene that was previously written by, I believe, Claude 2. The scene isn't very good, but all the components of the scene are there, so I ask it to just improve the writing of that scene and see how well it does. So those are the prose prompts. For the basic prose prompt, I wrote 415 words out of the 500 I requested, 346 of which I thought were viable. As you can see for the runner-up, right underneath it right here, which is Claude 3.5 Sonet, it did much better than that model did. For the complex prose model, I wrote a little bit less, at 311, but you can tell that's still more than I got out of previous models, and 269 of those words were viable. I can confirm it was much better overall; the general style was much better when I gave it that sample chapter and the more complex prompt. So giving it a sample chapter of your own work—or of what you consider to be good work—when you are constructing your story is a great way to improve the style of that scene if you're using a model like this. For the dialogue prose prompt, it gave me 472 out of 500, so very close to the mark there, and 407 of that was viable in my opinion. By the way, if you're asking how I tell what is viable, I literally go through and highlight the phrases that I think are good—like I would actually use it unchanged or with very little addition.
The editing prompt: the original editing story was 349 words; it expanded that to 569 words, 342 of which I thought was viable. So it definitely did a better job than we've seen here with these other models. I would say it wasn't perfect; still got a lot of better results. By writing the same scene—it's actually the same scene as this dialogue prose prompt—if you just write it over again with this model, it was better than trying to edit the original version, but at the same time, it did a better job of editing than I've seen from other models. Next, we get my marketing prompts. These are for marketing use cases that an author might need. First of all, I had it write ad headline prompts, and this one it did more or less the same as 3.5 Sonet; it gave me six worth; 3.5 gave me seven. It did do a little bit better on the email newsletter prompt, where it gave me 268 words that I thought were usable, as opposed to 148 for 3.5 Sonet, which is honest—so this is still better than any of the others have given me, even though it wasn't great. I would still need to modify the email it gave me quite a bit to make it something I would actually want to send out to an email list. But then we have the SEO article super prompt. In this super prompt, I give it a big list of stuff and ask it to write 4,000 words; it only wrote 2,000 words for me. I have some strong confidence that if I increase that request to say like 10,000 words, it would probably get closer to the 4,000 words I'm actually asking for, and it was usable within one single prompt. What it gave me, I was just like, "Yeah, I could publish this, and it would be fine." So it did a good job.
All that culminates into the overall score, which for this is an astronomical 593, almost 594—far higher than any model, by a long shot, far higher than any other model on this list. The best I've had before was the original—the previous version of 3.5 Sonet, which had a score of 412. This score right here is the next one; that's not bad; it had a score of 385. This is actually the creative writing version of GPT-4 0, which didn't do too bad either, but as you can see, both of those are dwarfed by how good 3.7 Sonet performed. It did an absolutely amazing job. I'm going to address one more thing before I wrap up because this came up in my writing group. By the way, if you are interested in understanding all of this—like I talk about what makes good writing here, and I'm going through and assessing whether this is a good phrase and if it isn't—I would invite you to take a look at my Story Hacker group because that's really where we learn what makes good storytelling. A lot of AI authors don't actually know, and they write stuff that I look at and can tell immediately was AI-generated; it's not really good; any editor would immediately throw this out. So I want to help authors not have to deal with that. In my Story Hacker group, I literally just hired a developmental editor who's going to be going through and doing classes on what makes good storytelling. I have another guy who's a designer teaching you how to use AI art, and then, of course, I am in there; I do a class every single week where you can come and get coached in a group setting, and there are accountability calls. It's an absolutely amazing community, plus you get a couple of free courses when you join and you get all of my books for authors for free. So go ahead and check that link down below.
One of the questions that came up in the Story Hacker community is how this compares to Sudowrite's Muse. If you saw my recent video about Sudowrite, you know I'm a big fan of their new Muse model, which does dialogue and prose very, very well. I would say, first of all, you can't do a full comparison because Muse isn't capable of doing some things like these brainstorming prompts or editing. Well, I can do some editing, but these marketing prompts—I can't ask it to write an article; that's not really what it's built for; it's really just built for prose. I don't have a comparison here yet, but I would only be able to compare it on these categories here, particularly the basic prose prompt, the complex prose prompt, and the dialogue prose prompt. I think Muse does have a little bit of an edge, even over 3.7 Sonet, although 3.7 Sonet is closing that gap pretty quickly. I think if we continue to see leaps in these models like this—if they all continue to improve like this one did, especially when we start getting Claude 4 and Claude 5—I wouldn't be surprised if at some point it's going to be good enough to rival a Muse model or any of the best that we've got out there. But for right now, I would say Muse is still a better model for prose specifically. But given that 3.7 Sonet can do a whole lot more besides just writing prose, it's still something I would strongly recommend you work with, especially if you are interested in doing brainstorming prompts and outlining prompts. It is, by far, the best thing I have come across for those specific use cases—by a mile. I'll also add that, from what I've heard, this is also a much better model for coding and math and a whole bunch of other things; particularly coders, I've heard really good things about coding. Of course, I am not a coder; this is not a channel about using AI for coding. But the fact that this model is so versatile and it's actually good at creative writing makes me really hopeful for the future that we might continue to get models like this that are really tailored to the practical use cases that people are actually using these models for—a big part of which is writing, whether that's writing for the web, writing for emails, writing creatively. I think Claude and Anthropic are the only people who really seem to be taking that side of things seriously, so I was really impressed with this model. Speaking of brainstorming, if you want to know more about that, I have a video that's like a full course on brainstorming with AI. I'm sure if you used 3.7 Sonet with the tips I make in that video, you're going to see some absolutely amazing results. So go check that video out, and I will see you there.