Transcription
So, by this point, you've no doubt heard that Claude Code is just blowing everyone's mind. And a few hours ago, the brand new version of Claude Code is officially out.
If you haven't heard all the recent noise about Cloud Code, basically a lot of people are saying that the capabilities of these models and harnesses like Cloud Code, they reached a certain capability where things feel different. They got to a certain level where we're now going to really start feeling the impact. People are saying that learning Claude is the best upskill this year.
Here's someone who is a principal engineer at Google. She's saying, "I'm not joking and this isn't funny. We have been trying to build distributed agent orchestrators at Google since last year. So, she works at Google. Keep that in mind. There are various options. Not everyone is aligned. I gave Claude code a description of the problem. It generated what we built last year in 1 hour. It's not perfect and I'm iterating on it, but this is where we are right now. If you're skeptical of coding agents, try it on a domain you are already an expert of. Build something complex from scratch where you can be the judge of the artifacts."
At the same time, Ethan Malik of One Useful Thing gave a prompt to develop a business that generates $1,000 a month and where he, Ethan, doesn't do any of the work where Claude takes care of the entire business. Was Claude Code able to do this thing? It seemed like it did. Now, Ethan did pull it down. He wasn't comfortable putting up. You'll see what happens in just a second, but we're at a very different point in history right now. We're at a point where things like this are becoming possible.
But I do think it's noteworthy that this principal engineer at Google, that works at Google, they have their own coding models. They have their own ways of coding stuff up. You know, Gemini, you know, they have the Gemini coding models, everything else. Some of the best engineers in the world. Still, they're saying, you know, Claude Code did something pretty amazing. That really does feel like an inflection point of sorts, does it not? I'm going to give this one a follow.
Now, one thing that I do want everybody to keep an eye out for because, you know, I've been publishing this news for quite some time and it happens every single time and it, it's always hilarious and I know it's going to happen here again. And that is the shifting goalpost. As these models get better and better and are able to do more and more wonderful things, there's a certain percentage of the population that invariably says, "Oh yeah, yeah, of course it can build this insane thing that it couldn't 6 months ago and nobody could have imagined that it would be able to build it, but it can't build this other thing that maybe a handful of people in the world can build, right?" So the goalposts are always shifting.
In a minute, we'll see how Ethan Malik tried to get Claude Code to build a functional viable business. As you'll see, we're probably a lot closer to that than we've imagined. But just know when we cross that line, invariably there will be people that are saying, "Oh yeah, so what? Why can't it build a million-dollar business and eventually a billion-dollar business? It will just keep shifting."
Here's Mustafa Sullivan. So, he was one of the co-founders of Deep Mind. He's working with Microsoft at the time. He's saying, "The next big milestone I'm watching out for on our way to AGI, artificial capable intelligence, ASI, can an agent take $100,000 and legally turn it into 1 million?" To me, that's the modern Turing test. So, kind of keep that in mind. To him, that is sort of the modern Turing test. But you got to think of it this way, like how many humans can reliably do this? As AI safety memes put out here, you know, can AIs act humanlike for 10 minutes, right? And then you know once we pass that immediately they're like, uh, can AIs become millionaires? So his point here is that the goalposts, they have been lifted and now they are leaving the stadium and just, you know, moving down the street further and further away. This is the kind of the AI skeptics taking the goalposts of what AI should be able to do to be considered, you know, good or AGI and they're just, you know, moving across the country with it.
By the way, Mustafa Sullivan, he had this modern Turing test from a while ago. I think we covered like a year or a year and a half ago. So, he's he's not moving his goalposts. He kept his goalposts at the same place. He's not moving it anywhere. But, I did want to point out this article by Ethan Malik. So he's he's got a Substack at One Useful Thing because the thing that he's pointing out and something that I think a lot of people are missing is this. A lot of the efforts in the AI space, a lot of the kind of development efforts were focused on coding for whatever reason. I mean, engineers know coding, so if they can get something to code that's, you know, reducing their workload, obviously it's something that has a lot of financial sort of rewards if you can replace coders or a system augment them somehow. Definitely that's very economically valuable. But it's not like coding somehow unique or special. This thing is coming for for every profession. It's coming for marketers, for researchers, for office managers, just whatever you're doing for a lot of the knowledge work. This will become an indispensable part of your toolkit. It might become the entire toolkit.
I got to give credit to David Shapiro. He recently posted a video where he talked about this. He compared it to jet pilots because of how quickly things happen when you're flying a jet. They try to offload as much of their thinking to various automated systems as much as possible because there's just not enough time for the human brain to process information that fast. So it's this idea of cognitive offloading, right? So similar to how we can do work by moving stuff around like physical work, move stuff around, you know, hammer things to the wall, etc. There's only so much of that that you can do in a given day without recovering. You can think of it the same way with cognitive labor. Some tasks are light and easy to lift and some are very, very heavy and will just completely drain all your resources. And when you need to move fast, you know, with business, with work, if you're trying to get something done, you know, the more stuff you can offload, the better. So, if you're, you know, a manager, if you can offload it to your the people that you've hired, that allows you to do more. In the future, more and more of this cognitive offloading will happen to these machines.
Five years ago, if I wanted to start an e-commerce business, I would have to do everything by hand. You know, set up the the websites and install Shopify or whatever I'm using, you know, get the pictures and the descriptions and everything. Everything had to be done by hand. Already, we're in an era where not everything is going to be done by hand. I would never write out product descriptions by hand. I would use a large language model to help me write it out and I would massage it a little bit, change it. But like a lot of that would be written by a large language model with my oversight. Same thing with images. Most like if I need a quick HTML page for, you know, opt-in, for email capture, whatever, I might actually have an LLM also code that up really fast. But those are still me using it as a tool with oversight. We're now moving into the era of the agents where a lot more of the tasks it can actually complete autonomously. And I think that's why this is such an interesting article to look at because I think it kind of gives you a glimpse into what's happening, what's coming just around the corner. And in fact, this what he's describing here, it's already here.
So, he's saying, "I open Cloud Code and gave it the command, develop a web-based or software-based startup idea that will make me $1,000 a month where you do all the work by generating the idea and implementing it. I shouldn't have to do anything at all except run some program you give me once. It shouldn't require any coding knowledge on my part. So, make sure everything works well." So, so, so think about that. The human labor part of this equation is I just need to double-click on some program and that's it. and starts printing money. How insane would it be to hear this, you know, 5 years ago before the ChatGPT moment. This would be completely out of the realm of what we would think of as being possible.
So, the AI asked me three multiple-choice questions, right? The nerve, how how dare it ask questions. It should just do it. It should just read our minds, right? That's going to be the next step, the next goalpost, right? So, it decided that I should be selling sets of 500 prompts for professional users for $39. Without any further input, it worked independently for an hour and 14 minutes, creating hundreds of code files and prompts. Then, it gave me a single file to run that created and deployed a working website filled with very sketchy fake marketing claims. Whoops. But this website sold the promised 500 prompt set. You can actually see the site it launched. So here's the site and interestingly. So this is, as far as the websites go as like how well they're being generated by these large language models. This is even not the best ones. I've seen ones that were just incredible looking ones. This is very, very simple, but I know that it can be a lot better. Some of the stuff that Gemini 3 comes up with is just absolutely stunning. It's gorgeous. So here's your get instant access, onetime payment, limited access, no subscription download now, your free demo. Oh, and actually I did download something. So he said later that he disabled the the payment links because he didn't want people actually spending money on this because, you know, some of the marketing claims are probably a little bit sketchy. But I mean here it is the ultimate AI prompt library. So all of this was created by this large language model and you just fill in, you know, whatever like for your name for your company for whatever you're trying to do but all the prompts are here.
So as Ethan Malik is saying here, he removed the sales link which did actually work and would have collected money. That's an important thing to understand. So that code, that package that it gave him to execute built that website and had a working payment link. Now, while it's not super complicated to install those on a website, you know, there's a lot of people for whom it's it's a stumbling block. They they would not be able to do it or they would find it too difficult. They might think they need some professional engineer to come in and do it for them. So there's a lot of people for whom this would have been a block. Cloud code just did it. And so Ethan saying, "I strongly suspect that if I ignored my conscience and actually sold these prompt packs, I would have made the promised $1,000."
Now, at this point, I'm sure we've all seen some variation of this chart from Meter Research, right? Right? So, they're kind of testing how long these models can run and how well they're able to complete tasks that would take a human being a certain amount of time. So, notice here in 2023, we had a GPT-4. So, it could do a task that would take a human being, I don't know, 5, 10 minutes, wherever it falls in that chart. By 04 mini, we're getting well past an hour. So, for example, they say fix bugs in a small Python library. By GPT-5, we're getting past two hours, right? Exploit a buffer overflow. GBT 5.1 Codeex Max is just about 3 hours and Claude Opus 4.5. Well, it's crawling out, but it's getting pretty close to about 5 hours of work.
So I know a lot of people get tripped up about what exactly that means, what exactly that they're measuring. They're very, you know, open about it. They're not hiding it. So, they are looking at the task duration for humans, right? So, if it takes a human 1 hour to do, that's one hour. If it takes a human 3 hours to do, you know, that's 3 hours. So that's separate from how long Claude Code or whatever runs, right? If if Claude Code runs for 2 hours but completes a task that would take a human being 5 hours to do, you know, they mark that as 5 hours. It replaced 5 hours of human labor. In this specific instance, it's for software engineering tasks. And after modeling a number of the tasks at those sort of levels, they kind of find this to be when they think the AI has a 50% chance of succeeding. So I think a lot of people maybe complain about that because maybe just it's not completely obvious, but that's an important caveat. So this is where, you know, it's a coin toss whether it gets it right or not, but when it gets it right, you know, if it's at five hours, then it replaces five hours of human labor, right? At a 50/50 chance. Just want to clear that up.
Now, looking at this chart, you might say that, "Wait a minute. Isn't Claude Opus 4.5 doesn't seem to fit this line? Shouldn't be be here somewhere, right? It seems like it really jumped up dramatically." Some of you might have seen our interview with Adam Binksmith. So, he runs AI Village. That's part of AI Digest. So there's a lot of very interesting, you can call them benchmarks or experiments, whatever you want to call them, but there's a lot of very interesting research that they're doing. And in one of their blog posts that's called "A New Moore's Law for AI Agents," they actually talk about this specific thing. Though it looks like they've updated with the most recent information. I forget exactly when it was posted, but this isn't recent. They've just been updating with new information. But they've noticed this pattern a while ago. A while in AI years, so like six months ago, maybe 12 months ago, something like that. And what they're pointing out is that the old trend that the Meter Research has been showing, you know, you can see it here, this kind of orange line that goes like this, right? The problem is this orange line only works if you're including all of the older models, you know, 2023 and, you know, GPT-4 and GPT-3.5, 3.2, etc. If you're looking at just the more current models, that's this red curve here. That's an important point to understand. If you get nothing else out of this video, just see this thing. I think this is the most important thing to really grok, if you will. The reason why Claude Opus 4.5 looks like it doesn't fit neatly on this line is because this line also includes in it the old models, GPT-4 and earlier. If we just look at these models here, then guess what happens? That's this red line and Opus 4.5 like just nails it, right? It's it's almost exactly on the line there. Why is that happening? Because recently the trend has accelerated in 2024 to 2025. Time horizons doubled every four months down from every 7 months. Right? So, we used to have these abilities doubling every seven months. Now, they're doubling every four months. So the abilities of Opus 4.5, they're right on schedule. So an important point to understand.
But coming back to Ethan Malik and what he's saying here, and this is the thing I'm I'm so happy that he wrote this up because it's something that people have mentioned here and there, but I think he kind of like really condensed this information in a in a very good way. What he's saying here is unfortunately for most of us who want to experiment with AI, these new tools are built for programmers. And I mean, they are really built for programmers, right? So if you're an engineer, you're a software engineer, something like that. This is built by you and for you, right? Fubu for us by us, as Demon John would say. So if you wanted to run Claude Code, you would have to open up your CLI, your computer command line interface, whether that's PowerShell with Windows or whatever you're using, whether you're using Mac or or Linux, every system is going to have their own names. They're all kind of going to be similar, but there's going to be some variations that you have to know. You have to run Claude. You have to know which folder to run it from. I mean, for people that are in engineering and software development, this is brain-dead simple, right? They're not confused by any of this. But for like 97% of the population, this is brand new stuff that they haven't really experienced with, they haven't really experimented with before. You got to know the commands and then you got to run Claude Code. And even then, I mean, this is how it looks. It's very coder-oriented. This is the thing that's preventing this power from being used, you know, at every other profession. This is why most people aren't really understanding what's coming.
So, coming back to Ethan Malik, he's saying they assume that you understand Python commands and the programming best practices and they are wrapped in interfaces that look like something from a 1980s computer lab. I mean, come on. I mean, a lot of people would find this nostalgic, right? Some small percentage of the world's population. The rest of the people like, I don't know what to do with this. And he's saying they are also explicitly designed to help analyze, troubleshoot, and write code using approaches that fit into existing programmer workflows. In a lot of ways, this is a shame because these systems are actually broadly useful to knowledge workers of all types.
Now, I'll link this down below because it's got some fairly interesting ideas that I just want to read it all here. It's better if you go and visit OneUsefulThing.com and check it out for yourself or OneUsefulThing.org. Actually, I'll link it down below, so just check it out. But Ethan explains why this harness works better than just talking to a regular chatbot. He also talks about skills, which is kind of part of what Claude Code can do. And by the way, this recent today's new update for Claude Code looks like it's got automatic skill hot reload. Skills created or modified in Claude/skills are now immediately available without restarting the session. Boris Cherney, who's one of the creators of Claude Code, by the way, I also post a lot of information that would be interesting and useful if you're planning to learn how to use it. One of the things that he said in one of his recent tweets that really jumped out at me was the fact that most of the new code contribution to Claude Code was done by Claude Code, right? So Claude Code is Vibe Coding itself, so to speak. So definitely give him a follow. That might be a very interesting person to follow, like the person that actually built this whole thing or at least, you know, initially built it and has been adding to it since.
Here's Peter Levels or LevelsIO. He's saying yeehaw. Now I can talk to Claude Code on my VPS on my iPhone with WhisperFlow on iOS. It adds a keyboard. Basically, it's getting Claude Code to run independently on your phone or you can use a Raspberry Pi or whatever piece of computer technology that you're not really using actively. You can have this thing kind of running and doing stuff on your behalf without necessarily needing to use it on your computer when you're in front of your computer. You're able to use the microphone to give it commands, etc. So, if you think about where you would use Twitter from or check email from or check your text messages from, whether you're at the doctor's office waiting for something or you're in the bathroom or when you just wake up before going to bed, whenever you use your phone for whatever nonsense you're using it for, let's be honest, most of it is nonsense. But during that time, you can actually be giving your Claude Code agents some commands or updates for them to work on.
So, I'm sure everybody's familiar with those idle games or incremental games, right? Where you kind of just give it some commands, but then kind of the game plays without you and then when you come back, it's made some progress. You give it a few commands and then you go away. So, a lot of the progress takes place when you're AFK. When you're away from the screen, away from the keyboard, you just come in to kind of course correct and give it new orders. Very popular games. This to me feels just like that. But instead of playing a game and you know having the numbers go up, you're building something and we're getting to a point where these things can build some pretty impressive software. But again, that's just because it's been so concentrating books for example, for creating startups, researching startup ideas and then trying to execute them. for managing digital businesses, running email newsletters, just whatever you can think of. A lot of the knowledge work will lend itself to automation of this sort. Just being able to code and create software in of itself is a staggering thing that's going to change a lot of things. It's going to have a big impact, but that's just the point of first contact. And kind of keep in mind too that, you know, here if you're just seeing one agent doing this thing, there's no reason why you're not able to run 10, 20, a thousand different agents running simultaneously.
So I guess the big point that I'm trying to make here is number one, install Claude Code. If you don't know how to ask Claude or Gemini or ChatGPT, they'll walk you through it. Whatever system you're on, it might take a little bit of back and forth. There's some stuff that you got to install that depending on what computer you're using, it might be on it already or not, but just walk through it. If you run into any issues, tell the chatbot. It'll tell you kind of what to do. Just ChatGPT or whatever your favorite one is. Gemini 3 recently has been very, very good to me. So maybe start there or again ChatGPT or Claude itself, it doesn't matter. Install it, mess around with it, and as you see different people saying, "Oh, I use it to build this or that." see if you can replicate that for yourself.
Here, Peter Levels used Claude Code to one-shot a Karen bot. If you're wondering what that is, it's just if you need to report something to the local government. I think this makes it super, super easy. You just snap a picture and this thing drafts a letter in formal Portuguese because that's where he is. And it automatically sends it to the local government. So, you're basically able to complain about various stuff that the government needs to fix, you know, just with the snap of a finger.
The creator of Claude Code, of course, says that in the last 30 days, 100% of his contributions to Claude Code were written by Claude Code again. So he's creating Claude Code with Claude Code. This Google engineer used Claude Code to create what a bunch of Google engineers built last year, but it built it in one hour. Right? These aren't people that are like hyping and they just say stuff to get attention. And these are serious people at serious companies that are saying, "Hey, yes, we're engineers with tons of experience. We have teams and knowledge, but they're seeing the ability of these AI models to create stuff that's impressive. Maybe it's not yet better than what they're doing, right?" As she says here, it's not perfect, and she does have to iterate on it, but this is where we are now. Meaning that, you know, this is where we were earlier this year. This is where we were a year before that, right? This is where we are now. And if you follow kind of like the latest data, you only look at the more recent data, right? This red line that that that shows you where we're going. And if you're wondering, well, where is it going? Isn't there some projections? Don't worry, I got you. Or Adam Binksmith and AI Digest got you. Let's see.
So, by July 2027, so a year and a half from now, if we're looking kind of like that old scaling law, right? In these AI models, these AI agents will be able to do one workday or 12 hours of work. But if we're looking at the red line, kind of the the more current line that we seem to be following, well, then it looks like that intersects with the one work month line over here, right? So 167 hours. That's the July 2027th. So in a year and a half if this progress is maintained, then you're going to be able to give Claude Code something that would take a a human being, a human worker, one month of of work to do.
So I actually asked Gemini 3 kind of what the scope of a project like that would entail just to give me some ideas. So it's saying writing the first draft of a short book, creating a comprehensive brand identity package. So just not just the logo but the entire visual language for a new company, you know, research competitors, sketching logo concepts, selecting typography, the entire post-production for a short documentary, perhaps building a minimal viable product mobile app, migrating a small business website, right? So like from WordPress to something new, developing a full online course, etc., etc., etc.
So this is David Holz, he's the founder of Midjourney. The image journey, I think is still to this day my my my favorite AI image generation platform. There have been many, but that one just has a special place in my heart. I don't know if it was like one of the first ones that I've used that I really liked, but it just it stuck with me. But he's saying, "I've done more personal coding projects over Christmas break than I have in the last 10 years. It's crazy. I can sense the limitations, but I know nothing is going to be the same anymore." Elon Musk responded saying, "We have entered the singularity."
And certainly when you're looking at those charts and you assume that that progress continues, it's kind of wild to think about where we're going to be in 6 months, 12 months, 18 months, if in 18 months Claude Code can with one prompt do the job that would take one month for for a human professional to do. The explosion of productivity will be absolutely incredible. And of course, there's positives to that. There's also some scary, let's say, negatives or potential pitfalls, as in what happens to people's jobs, etc.
But whatever you do, if if you haven't been able to get to this screen yet, if you haven't been able to launch Claude Code yet, do so. Uh, I'll try to do a little bit more of an in-depth tutorial in one of my later videos, but you really don't need it at this point. The idea of needing a YouTube video to walk you through step by step, that's becoming rapidly outdated. Here again is a Gemini 3. I just asked how to install Claude Code on PC when it just came out. You couldn't run it on PC if my memory serves. I remember I had to set up a Linux machine just to run it on there because you would not be able to run it on PC again if I'm remembering correctly. Now you can. J3 walks you through it. If you don't understand anything, just ask. If you've never done this before, yeah, it might feel a little bit uncomfortable, but trust me, it's worth it. Start using it. Start figuring it out. If you have any of these chatbots, you have access to the most advanced artificial intelligence on the planet. Put it to work. Start learning because I'm I'm telling you, a year or two from now, you'll wish you had started earlier. Because yeah, it really does seem like we're entering the singularity of sorts and uh things are are going to get crazy. So, buckle up, get ready, and I'll see you in the next.