📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Anthropic’s Fable 5: A Warp Drive for Coding

Every16:38

Transcription

This is the infinite library of Babel from the Bourhees story. It contains all of the books in the universe because books are just strings. If you look, you can even go into bookmarks and I can click one of my articles after automation and it finds it in the library. It's truly infinite.

And look, I could go up the stairs. I can like look down. I can look up. This seems like it took a long time to make, right? Wrong. I made this entire thing in a single prompt with Fable 5, the new model from Anthropic. Like, like literally, let me show you. So, this is a prompt. It's from four days ago or so. I got this model a little bit ahead of time. Read Jorge Louie Bourhees's The Library of Babel and then plan and execute end to end a browser playable 3D game in which the player is dropped in blah blah blah blah loop until it's done. Just wrote that, pressed enter, and it just went off. It read the story and then it just ran and ran and ran. You can see it's like looping itself and it's checking its work. After 3 or 4 hours or so, done. We've got hexagonal galleries stacked endlessly. We've got 20 shelves, five per side. As one would expect, it's accurate to the book. It's got the mathematics right. It says the part I'm proudest of. This is crazy. It just made it in one shot in 3 or 4 hours running on its own.

Fable 5 launches today. Here's your day zero vibe check. But first, remember to never make any major life decisions within 30 days of a meditation retreat, a psychedelic experience, or your first encounter with a Frontier model. Cheers.

So, before we get into it, you're probably wondering, how is this video even out? My name is Dan Shipper. I'm the co-founder and CEO of EveryY. Every is the only subscription you need to stay at the edge of AI. You can kind of think of us as like an AI lab for the future of work. We spend all of our time testing new models, using them to do our work from programming to writing to design to business building to decision-making. We use them hands-on and we tell you about what works and what doesn't for real use cases. And I'm incredibly excited to be doing this because the first encounter with any new model could be crazy. But Fable, which is a mythosclass model, I think is like is a particularly big moment. It's like the most hyped model. When it leaked a month and a half ago, Anthropic said it was too dangerous to even release it. And now it's out. And I have a feeling that if you're like me, you might be excited, but you're also like a little scared. Because we've been using this model for about a week now, we get to pull back the curtain a little bit and show you what it's like to have lived with this model a little bit more. It does change things, but hopefully this can help alleviate if you're feeling a little bit of AI psychosis. I'm sure that that's going to be going around on X and YouTube and and the news and all that kind of stuff. This is a place for you to see how this thing might fit into your work and into your life and in a realistic way. So, let's get into it.

Okay, so Fable is a Mythos class model. Mythos is a model from Enthropic. It's the largest model that they make. There's Haiku, Sonnet, Opus, and then Mytho. As far as I can tell from talking to people internally at Anthropic, there's nothing special about it architecturally. It's basically the same thing as their other models. It's just bigger and better. In order to make it safe to release, they have put pretty strict safeguards on it. So, you can't use it for anything cyber related. You can't use it for anything biological related. That's what makes Anthropic comfortable releasing it to the general public. It's pretty expensive. It's $10 per million input tokens and $50 per million output tokens, which is about twice the cost of Opus. So, it's a lot, but it is just genuinely the most powerful coding model I've ever used by far.

To give you a sense, we have a senior engineer benchmark, which basically tests the model on its ability to act like a human senior engineer. We give it a vibecoded slop production codebase, a real production codebase, and we ask it, if you're going to rewrite this from first principles, how would you do it? And then we see how it does. We score it out of 100. The best model score is a 63 out of 100, which is Opus 4.8, which came out like two weeks ago. And right behind that is GPT 5.5, which is a 62 out of 100. Fable scored a 91 on this benchmark. 91 out of 100. That's the same score as a human engineer with just a just just one prompt. That's that's it's crazy. I like I I knew that this benchmark was going to get saturated, but I thought it would happen in like 6 months.

Look at this view of it when we break it down by what it's good at versus other models. This is Opus 47. The, you know, the orange stuff is what it what it does versus what it's going to do. You know, pretty spiky, not that great. GPG 5.5, like we're starting to fill out the the hexagon a little bit. This is just like, oh yeah, it just did it. If I try to like break down for you, okay, what is it really good at? Because it's not good for everything. I think it's fantastic at sustained autonomous execution. Like for example, the way to work with this model is to give it a task and then leave. Go do something else. Let it go for three or four hours. Set it up overnight. It's It's amazing. Like, it it just figures stuff out and it just does good work. It has good taste. It has good attention to detail. There's all these like little details that it does pretty well that I'll show you. That's that's really impressive. Even with a not very well specified prompt, um it it has it has more judgment. I think previous Cloud models, you'd be like, "Oh, do this thing." And it'll be like, "Oh my god, yes, I'm going to do it. I'm going to do it." And then purple accents, purple accents. It was like a little tryhard, to be honest. In this model, it feels like it it's going to go do it and it's going to think it through and think about how to do it well. And if it doesn't think it's it can do it well, it'll it might say something to you, which is which is really helpful. And it's also just incredibly good at like using a lot of context, like doing a bunch of research, digging into data, giving you a bunch of things from the data that that that you wouldn't have known beforehand.

I'm going to go through a bunch of really specific examples of exactly what this is and how it works. But the if I step back and I think about in particular for programming like what is this model? It's it's like a warp drive. Um, you know, like in Star Wars they, you know, you want to you want to like jump across the galaxy. You like, you know, you punch up you punch the coordinates into the computer. The computer does some calculations to see like, okay, we don't want to we don't want to like go through a star or whatever. Then you punch it and like the stars blur into blue and you're like punch it. >> Punch it. >> Just punch it, man. Just punch it. Go. >> You don't get there instantly with a warp drive. you're going across the galaxy, but it takes you like, you know, a couple hours or a couple days, which is pretty good because it used to take you like a year or two or maybe 20 or 100 years, but with a war drive, you get there in a couple hours or a couple days. And that's kind of like what this model is like. You can specify a destination for a big trip and it just like it compresses what normally would have been like years or months into like hours or days. But also like it's not really that good for getting around town. You know, you wouldn't use a warp drive to get around town. You need like you need more control. you need more feedback back and forth between you and the vehicle. And the same thing is true for this model. If you're using it for like true collaboration or quick questions or things that need tight back and forth, I don't think it's that good for that. I mean, it's very slow. It's very expensive. It's extremely token hungry.

One thing that's a sort of a pro tip that's not intuitive for this model is you can set it to lower reasoning levels, like, you know, instead of max or extra high, you can put it to medium or or low for more basic questions. That's not something that was very intuitive to me, but that's how people inside of Anthropic use it. So, if you want to use it a lot and you're trying for more of the like everyday thing, you can do this. But I just I I think it's I think it's a warp drive. It's really good for those big things, but you have to you have to have big meaty things to give it.

I want to show you some examples of some of the things that we saw in testing this that'll tell you some of the properties of this model and maybe give you some inspiration for where you might want to use it. Okay, so first example, there's this philosopher I love. His name is uh Hebrew Drifus. He wrote a book called what what computers can't do in the 70s which is like a famous explanation of why AI at the time was not going to work. Really like him. He's really into Haidiger a philosopher and he has these lectures on H Highiger from like 2007 but they're kind of hard to follow. The audio recording is not not that good. So I just told Fable hey can you like go get his go get his lectures and turn it into a little mini site for me so that they're easier to consume. That's literally all I did. I didn't even link it to the lectures. I just said go grab the lectures. I didn't know where they were. It went and grabbed the lectures. It wrote like look, it wrote like a little summary. Why does this matter? These are audio lectures. It broke it down into a table of contents. And then it has this like player experience where, you know, I can press play >> and it'll play him >> and and it highlights what he's saying as he's saying it. Like it synced the the the playback to the text. and it has this like player thing and it's got like, you know, minus 15 and plus 15 and, you know, you can do 1x or 2x and you can follow it or not follow it like this is from one prompt and that's what I mean when I talk about this model having really exceptional taste and attention to detail. There's a lot of different dimensions along which this is happening. Look at the font choices. This is all caps. There's a little bit more font weight on here. It does the drop cap here. These are not like the defaults that you would expect. It's not like the clawed purple highlight vibe slop thing. I assume some of this stuff will look sloppier in at some point because everyone's going to be using this model and using to do stuff, but for now it really looks good and the things that it builds feel much more thoughtful. These are things obviously you could have built with a model before, but now they're just available to you almost for free. And I I think that that's like a good test of a model is does it get you to just like try a bunch of new stuff because now stuff that used to be hard is easy and so it just opens up a whole realm of things. That's 100% true with this model. So this is my first one. This is the the kind of taste and and detail example.

Another thing that it does which I think is so cool is it's so good at using its context and doing research at every have a paid subscription. We've got like 10,000 paying subscribers. If you have not subscribed, you should subscribe every.subscribe. Uh you'll get a text version of this vibe check if you want more detailed insights. So check that out. So at every we've got maybe 100,000 free subscribers and we do surveys. So we want to understand what do people think about us. We have a bunch of recent survey data and so we fed it into this model and look at what it did. Like a team of us have been looking at this data for weeks now with AI and we have not come out with anything this interesting or succinct. This is from like hundreds and hundreds and hundreds of survey responses maybe in the thousands of survey responses. The punch line at the top you have a conversion merchandising problem. Your freeto paid conversion ratio is lower than it should be. Okay. Next, the falsifiable bet. If we ship pricing transparency in a trial offer, I think it's going to go up. Like, that is something that I would expect a really, really good growth person to do with a lot of time and thought and research. It requires it to think across so many different dimensions. It has to go look at all the data. It has to go look at all the survey responses. It has to go look at our analytics data. It has to then go through the site itself and put that all together in a way where it can tell you here's the punch line and here's what I would do next in a way that I could just scan. We've not seen a model that's able to do this and I think this applies to lots and lots of different challenges whether you're using it for coding or knowledge work or whatever you want to use it for.

So far I've shown you like a lot of cool oneshotted demos and stuff like that. The other thing I did is we have this we have this app called proof. It's a agent native markdown editor. This is this is proof. We we have a bunch of issues in the GitHub with proof because agents can submit issues. So as you're using it with your agent, it just submits a bunch of issues. A bunch of them come in every day. I just pointed it at our issues in GitHub and I just said like, "Okay, I want you to take all the issues for the last couple weeks and close out any that aren't relevant and write fixes for all the all the rest." And it just went boom boom boom boom boom boom boom and actually wrote fixes that we merged. Again, other models can do this, but it's much more okay, you go one at a time, you make sure it's doing well. You can't just be like, "Go do it." This is this is why it's it's a it's a warp drive. It just speeds through the backlog of simple things like this in a way that would be impossible with other models.

And now the question might be in your mind like who should use this? And I really don't think that this model as it is right now is for everyone. Again, it's slow, expensive, it's super powerful, but because of that, it's not for everyone. We have this article on every called the eight levels of AI adoption, which you should definitely read. You can actually throw it into your agent, and we'll put the link in the show notes. Your agent can just go through how you how you use AI, and it'll break it down into eight levels. everything from at the bottom level you're just using it essentially like a Google replacement to just like ask basic questions to the top level it's like you're orchestrating many many different agents where you're delegating work and and it's working 24/7 and all that kind of stuff. That's the kind of that's the spectrum and everyone falls on a different place in the spectrum. We probably had seven or so people testing it internally over the last week. Everyone from programmers to writers to editors to marketers. What we found is there's actually like a pretty high spread of who liked it and who didn't. And it's not like anyone hated it, but if you're if you're not at a certain level of AI workflow, you're kind of like, I don't know what to use this for because you don't you don't have a problem that's big enough where you need to speed through the galaxy. What we found is if you're at like a seven or eight on this scale, you know, you're you're using multiple agents and you're orchestrating them and all that kind of stuff, you've got big meaty problems where you're like, "Wow, this is crazy." And usually that's technical people. If you're non-technical, if you're a vibe coder, you you probably are watching this being like, "Holy I have so many projects I I want to do as long as you can afford it." I'll say it again, this model is expensive and it hogs a ton of tokens. At least for now, if you're vibe coding, I would be I would be careful with it, but I would definitely definitely try it. And then if you're kind of like if you're a knowledge worker and you're just using it to get your job done, I unless you're a very advanced knowledge worker that's using it in in this sort of way, like you're you're orchestrating multiple agents together and delegating a lot of your work, it's going to feel like overkill. And and I think that's that's a really interesting thing is like what we're finding is yeah, yeah, you have this AI, but using it is a skill. You need to be exposed to problems and working at a level of expertise where the problems come up in order for it to be useful. Internally, like all of us are early adopters of AI. I'd say probably about maybe half of us are are at a level where we can see the problems that this thing solves. And about half of us are still getting there. I think we will. I I actually think that this kind of workflow is going to be available to everyone soon or it's going to be useful for everyone soon, but it just depends on your on your where you are on the adoption curve.

And to get into a couple things the model doesn't do as well. It it didn't actually do substantially better at writing than Opus 4.8. So if you're using it for writing, we we found its sentences to be pretty dense, you know, like big blocks and they're pretty literary. So for some things, it can be good for that. And I I do think it's very good for thinking through writing issues, but for actually writing sentences for copyrightiting is probably not the thing you're going to want to use. If you're a cloud person, just use 4.8. If you're a GBT person, 5.5 is much better. I I personally prefer 5.5 and I still use 5.5 inside of codecs as my daily driver because most of the stuff I'm doing I want to go back and forth pretty quick and a model like this is going to is going to be overkill. So this this model like it increases my confidence with for big projects or maybe writing production code like that kind of stuff but uh for my dayto-day it's a bit overkill even for me.

Finally like let's talk a little bit about what is what is the meaning of fable? we want to go past. Oh my god, it's going to change everything. And to some extent, that's actually right. But it's not going to change everything in the way that I think people imagine. I just wrote this piece called after automation, which you should read, which is about what is it, what is work like after we've automated everything. It turns out automation actually creates a lot more human work. It's a very interesting paradox. I think the same is true here. What we're going to find is this model increases the floor of capability for non-experts, but it also raises the ceiling for experts. So, a vibe coder might be able to make a oneshot video game and an expert might be able to make like a true AAA game just by themselves. And I think that I think that is so cool. Obviously, the fact that things are changing this much and and I think it's really important for us to say this does change things. If you're someone who's used to typing code into your computer, this changes that a lot. And I think it's normal to be sad or angry or weirded out by that. And it also changes things even if you're using AI already. This this changes how you can expect to use it in the future. A lot of the skills, a lot of the things that you thought you might have to do are starting to change a bit because this model is so much more powerful. Change can be scary, but it's also an opportunity to be like, "Wow, what can I do now that I might be into that this now makes possible that I can just do? I don't need to ask permission. I don't need more money. I don't need anything." And and because this capability is out now, we can expect that even if even if it's too expensive for you to use right now, it's going to be pretty cheap soon. Like let's say within the next 6 months to a year, everyone will be able to have this. And I think that's I think that is incredible.

So if you like this video, you should really watch my video after automation which talks about what happens when we automate everything and read the article. You should also read our vibe check on every we go in depth through every part of the testing that we did for this model, all the benchmarks from coding to writing to knowledge work. We have takes from the entire team. We have a bunch of people testing this from different perspectives and different walks of life and different ways that they like to use AI. But if you're psyched about this, the thing I recommend most is go use your new warp drive and let me know what you