Transcription
Oh boy. Oh boy. This week in AI was wild. We got the biggest model release of the year, probably. Usually, in these model releases, I jump within 24 hours. I'm there reporting on what's happening.
Well, two things happened this week. First of all, I kind of looked at this and I was like, I want more information. I want to see what people think, and I want to actually run some of these test prompts that are going to run for 12 hours and let other people do it, cuz that's what this model does. And then two, I also just had a super busy week, and I didn't want to present you with like a half-baked take.
So here we are. We're going to cover Fable 5, which is Anthropic's biggest and, well, best model, especially on agentic and coding-related tasks and sciences, and actually kind of everything knowledge work. And then we're going to talk about some other quick stories that came up this week that are relevant. But honestly, main topic today: Fable 5. It's finally out. It's available on paid plans, and let's dive into it.
So, first up, the high-level. You probably heard it by now, but if you haven't, this is the best model in terms of benchmarks and user reviews across all of them now. And it's not by a little bit. Usually, these model releases, they inch up a little bit. This one's head and shoulders above.
If you look at the benchmarks, it tells the story really well. Look, especially agentic coding is one that I always look at, and knowledge work. It's just not even close. GPT 5.5 at 58, Claude Opus at 69. This is at over 80% on software engineering bench pro, an important one that usually all these model makers always publish. It's just the same story across everything else.
And then on the other hand, there's the user experiences and the user reviews, and those are excellent, too. I really liked this Reddit take over here, which says that Fable feels like a mature, calm, and down-to-earth programmer. It's very impressive. And in short, I'm going to show you a bunch of examples and talk about my experience and what I saw on the internet. But I agree. I like that it responds concisely and it just gets work done in a more thorough and in-depth way than anything else.
The cost of that, well, it's literally the cost of the model. If you run it on your Cloud Max subscription, the $200 subscription, you're going to find that within like 30 minutes, you can easily fill up your entire usage for the day. And if you run it through the API, well, it just straight-up costs twice the price of Claude Opus 4.8 today, which is already known for being notoriously expensive.
So, the big question: What can it do? What can you do with it that would be worth your time? Because this thing is available within your Claude subscription, but only until June the 22nd. If you go into your Claude subscription, you will find that, hey, they only gave this out for what is it, like 12 days from the release. And then you're going to have to pay for the tokens directly through the API, which just makes me think, in combination with the fact that they're prohibiting the use of your Anthropic subscription with agentic systems like OpenClaw or Hermes, which is happening also over the next few days. Anthropic is really moving towards a version of like, I don't know, I would call it AI capitalism, where you just need serious money to get the best models and agentic capabilities.
Anyway, as we have this window and everybody can access this on the paid plans, let's look at what you can actually do with it and what it performs better on. And there's a ton of examples on the internet. I'm going to show you some of the cool games and fun interactive and entertaining demos that people were creating. I want to start with the one that touched me the most and I thought was the most impressive.
So, I don't talk about this a lot on the channel. Actually, I've been thinking about how to talk about it more. But just because I don't make content on it doesn't mean I'm not all over kind of this agentic revolution that is happening right now. Over there, I have a Mac Mini running two versions of OpenClaw and a MacBook Pro running a Hermes agent, and then I also use Claude Co-work a lot.
Now, what I tried with Fable 5 is running a security audit on the whole system and the platform that I built for our education company, AI Advantage, inside of OpenClaw. You might not be familiar, and that is fine, and it goes way beyond this video to guide you through this. But I basically built out this entire hub. Alfredo is my agent with all of these different sub-applications for different parts of the business. And there's a lot there. There's apps, there's almost 50 cron jobs, there's a bunch of agentic loops running too.
But what I did is I ran a security audit on the entire Alfredo hub first with Opus 4.8. And I found some security audit prompt on the internet that was popular in some Reddit forum. And look, if you look into my Telegram chat from which I operate Alfredo most of the time, it did pretty well. It found some inconsistencies. It gave me some quick wins that I implemented right away. Fantastic.
Now, I switched the model over from Opus 4.8 to Fable 5. And I ran the same audit. Look at that. Alfredo Hub audit by Claude Fable 5. It just found a whole new list. And in one sentence, I would say that the first audit rated the platform as a B minus in terms of security. The new audit rated it as a C minus, and the C minus was received after I fixed the issues from the first report. So I don't want to go into the technicalities, but the different authorization levels that I had that were inconsistent, I missed. Opus missed, and I'm so glad it found it, cuz I have different access levels within the hub for different people in the company, and it cleaned all of that up. And then there's all the other technical things it recommended. I've been implementing all that, and I'm like, this is just really useful to me.
So that's my first take because people have been creating a bunch of games with this. Apparently, it's really good at coding and coding complex things like games, and it demos well. So people have been doing that. Look, one-shot prompt: make a Pokémon clone. And apparently, it one-shot well this version of Pokémon where you can walk around and fight others. Impressive.
If you follow the channel, you might know we also run this prompt of this 3D space shooter game and just see how it looks in different models. I mean, do you hear the sound? It did a soundtrack. This is by far the most impressive one. Look at that. As we begin here, approach here. Can collect these, go. Okay, we're getting it. And then if you crash into something, it's over. Nope, you lose life. This is the most robust game out of all of them.
Okay, some more amazing stuff. So here, Kumar on Twitter basically built a 3D map of Delhi. He said it cost around 1.5 million tokens. So this map cost him $75. But look at that. It's pretty amazing. It's a 3D map. Use various map libraries to actually create this.
And then there's all the websites that people have been creating. Here's my favorite example of Victor sharing a 12-minute tutorial if you want to check out exactly how he made it. Now, this is not exactly a one-shot website, meaning you give one prompt and it gives you an amazing result. He actually brought in visual elements that he then referenced and had quite a complex prompt with exactly what technology to use and what the outcome is supposed to look like. But look at these websites. I mean, this is next-level stuff, and all doable within a 10-minute tutorial. If there's a thing to try yourself, it's making one of these insane websites or any website. It's just really, really good at coding.
And what I always love seeing is that the sentiment across the internet plus my own experiences, that they align with also Reddit takes and Twitter comments, and it just does like there's a unified voice across the internet of this thing being just amazing, especially when you're building things, improving things, creating websites, creating anything.
This Reddit user puts it well that he used Fable for stuff he was working on today, and it did a great job. But yesterday he used Opus, and it also did a great job. It's hard to judge the relative quality when both models are smashing the tasks I give them, which is true. So, it's not like you should be using Fable for everything from here out. They're both still great, but there are differences.
I'll show you two more things. One of the test prompts we always also run on these models, you might know this, is creating a Death Star above Los Angeles. And yeah, I just have to say this is the most aesthetic and just objectively best one we have seen so far. The downtown skyline is accurate. You have the 405, little seagulls, and the Death Star above it.
But here's one note, and people have been reporting this too. Apparently, the guardrails on this are so tight that as the first model here, it had a problem with the idea of replicating the Death Star because that's a Star Wars-themed thing, and it didn't want to go there. So, it kind of did a version of it. I suppose the detail of it is way higher. But you can see that this is not the exact Death Star from Star Wars, whereas some of the other models that ran it exactly recreated, well, the Death Star with, I don't know, this laser-making ring that it has.
Okay, one last example to round this out. So, one thing that I like to do with these models is just accumulate a bunch of research and then sort through it. You might know the deep research functionality within ChatGPT, Claude, Gemini, doesn't matter where, where you can basically go into a chat, you know, click plus, and then you say research, and it's just going to accumulate, I don't know, 100, 200 links now, I believe, in chat. And this was the case a few months ago, but it didn't matter what model you picked, the research always used the same model. In Claude, what I found, it's different. It actually makes a difference if you select 5 or Fable 5 or Opus 4.8 today.
What I did is run two quick researches that actually found helpful on the competitive landscape in the AI education marketplace where we operate. And what I did is exactly what I teach. I ran this inside of my project with all of my context. And then I also enabled research. I didn't even need to fill in my market, cuz all that context is in the project. And voila, I got one result from Fable and one result from Opus.
Now, I did the following thing. I copied the result from Fable and put it into the Opus chat and asked, if you were to compare this report to a second version I have, what is the difference and which one is better? And it says, verdict: for actually making decisions, the version you have, the Fable version, is the stronger strategic document. Mine is more faithful to the narrow brief you gave me and goes deeper on that specific niche, but that narrowness is also its main weakness. The honest move isn't to pick one, it's to fold A's depth into B's frame. Even Opus is like, there's different strengths, but I think overall the Fable one, me reading, reading through it and looking at these differences, is superior in multiple ways. I'll just leave it at that without going into all the nuances where you would need context on our business to understand them.
I'll say one thing though, there was a mistake in my context. I need to update one of the files there that was still stating that our community is at, was it 300-something members? That's information from last year, as we merged with Dean and Tony and built this new community. Right now, we're at over, and you might not know this, but this is insane, 70,000 users in the new community. These are paid subscribers to the new community that we run, where I run the programs and teach the courses. If you want to learn how to set up a project, there's a course on that in the community. You can check it out for $1. We have a 14-day trial where you get everything in there. See if it's for you. Link in the description.
But my point here is that it really took that number with the 300-something subscribers super seriously, and it ranked us at fifth in the competitive landscape, whereas Opus took it less seriously and ranked us as number one. So I think I actually give that to Fable because it gave the critical context where it assumed that we're not able to deliver paid products that monetize and convert people and give deeper value from all our free offerings. It took that so much more seriously, and it baked it into the report and was way more critical and honest, which I always appreciate. So another golden star to Fable.
I think across everything that I showed you here is just superior. It's twice as expensive, and soon you won't be able to access this through the API, which, as I pointed towards, kind of shifts what's happening in AI. A lot of these best models are going to become premium, and then the competition is going to catch up, and China is going to come out with an open-source version, and then all the prices are going to go down. But the frontier of AI has gradually been getting more and more expensive for consumers. Remember, just two years ago, the most expensive plan you could get was a $20 ChatGPT subscription. Now, the $200 plans are not even going to include the best model available now.
So, go try it, build a game, run some research reports, and let me know what you think in the comment section below. That is my take on Fable as of now.
But that's not it. There were a few other stories this week that I found interesting. One of them is good news for all Claude users. Maybe they kind of did this to kind of soften the blow with Fable. They're essentially doubling your Claude Co-work usage for the next month. Simple as that. If you use Co-work, you can run more scheduled tasks, do more on your existing subscriptions. Goes for everything from $20, $100, $200 subscriptions. And double the usage also applies to you having Fable for the next 10 days. So, if you want to experiment, now there's a little sweet spot. Double the usage with the best model ever, which, you know, this makes up for the fact that Fable uses double the tokens. They kind of cancel each other out if you're using it, but it's good to know about. And hey, if you're enjoying this coverage, make sure to subscribe to the channel. It really helps us out. Now, see what's next.
Then, there were Apple announcements about an improved Siri. They're partnering with Google Gemini, and in July, we will receive the developer beta of what they actually cooked up. It's Apple at this point. They're a few steps behind in AI, and I'm not going to judge anything based on their announcements because they always sound good, and then Apple intelligence underperforms in practice. So, I'm really curious to see this to test this out for myself. I mean, everybody on planet Earth, everybody who interacts with tech, can see the vision of, you know, Siri actually working just as well as ChatGPT or Claude, but then it has so much unique data and applications that you use day-to-day, that the potential there is just unbelievable.
And then finally, I wanted to tell you about an interesting update to NotebookLM because they made it agentic. Now, this is unfortunately only on the Ultra plan, which is the $200 plan. Oh boy, they're making AI expensive. It's happening, you guys, as you can see it right here. But on the $200 plan, now you can have NotebookLM basically run an agent in the background that has its own virtual machine, and it can code assets that it hasn't been able to do before. And it can also research the internet and have a thinking loop that is more extensive than what it had.
Remember, NotebookLM basically works in a way where you just add a bunch of sources, and then it just looks at those documents and uses a large language model to produce results from them. It can transform them into various outputs that they have on the right sidebar. Well, now it can turn it into even more outputs because there's an entire agent in the background that can do things for Ultra subscribers, and it can find new sources by itself and kind of reason over everything and go more in-depth. They're really turning it into a research studio that is increasingly autonomous and flexible, but that has its price.
And yeah, there you go. That's pretty much everything I found really interesting for this week. I think it's really worth getting your hands on Fable and just getting a feeling for it. It writes well, too, by the way. Even basic writing, I didn't want to go there cuz that's so hard to judge, and there was a lot of other interesting stuff. But even the short stories we wrote with it were like, that's really good, and it has a different flavor than Opus and GPT 5.5. And if you have any sort of app, just running a cybersecurity audit on it is sort of a no-brainer thing that you should immediately do.
All right, there's your weekly roundup. My name is Egor, and I hope you have a wonderful week. AI news you can use.