📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

This “Karpathy System” could 701x your AI Workflows (86,000 GitHub Stars!)

Dream Labs AI21:40

Transcription

Andre Kapathy, the godfather of modern AI, has just released a simple system called Auto Research, which the top AI labs in the world have been spending millions trying to create. Since releasing this auto researched, it's got over 85,000 stars on GitHub. And even Shopify CEO pointed at Shopify's code and sped it up by 53%.

But here's the thing, it doesn't just work for coding. In fact, auto research can be pointed at any part of your business. Your emails, your ads, your landing pages, your AI skills, AI agents, or even your organic content. And it forces each part of your business to self-improve at the speed of light while you sleep.

So, in this video, we'll break down Kapathy's simple order research system and show you exactly how to plug in any part of your business that you want to improve. And we'll do this together by going through three real-world examples that you can start using today. All I ask in return is you hit that like button down below. Grab your Slovakian flag and let's jump in.

Okay, so the fun started here with Andre Kapathy tweeting that got another 11 million views for a highly technical tweet, which once again shows the demand and how impressive this stuff actually is. He wrote, "I packaged up the auto research project into a new self-contained minimal repo if people would like to play with it over the weekend."

He says, "The human, which is me and you, will iterate on the prompt, giving the agent a set of instructions on what we want it to improve, and then the AI agent will iterate on the training code or whatever the asset it is that we want them to improve. And we give them, it's going to literally work all night trying to improve that asset while we're asleep."

Let me show you where Andre Kapathy started. He says, "The goal is to engineer your AI agents to make the fastest research progress indefinitely and without any of your own human involvement, which is why we can be asleep." In the image below, which this auto research image here is his LLM getting smarter and smarter with every iteration that it was running on a 5-minute loop until 83 experiments later, it was 11% faster than when Andre Kapathy left it. And he also says that he'd been working on this same LLM agent trying to make it smarter for a very long time and thought it was done. But his auto research tool found another 11% improvement in its IQ.

And this paragraph below the image might be one of the most fascinating things to read in terms of a mindset shift on how powerful this system is. Kapathy wrote, "One day Frontier AI research used to be done by meat computers or humans in between eating, sleeping, and having fun. And these meat computers would synchronize with each other once in a while using soundwave interconnect which is obviously speaking inside a group meeting." It's a facetious take on how slow us humans are to do these jobs compared to something like this auto research system. He says that era is now long gone. Research is now entirely the domain of autonomous swarms of AI agents. And he says that this repo or this system right here is the story of how all of this began.

And so the repo is up on GitHub, which is a crazy thing. Andre Kapathy just open-sourced it and made it free for anyone to use. It's called Auto Research. And you can see it has 85,000 stars at the time that I'm recording this. And so there's been hundreds of thousands of people who have been running this system, including, like I said in the intro, Toby Ludkkey, the billionaire CEO of Shopify. And he said, "Okay, this thing is totally insane." Before he goes to bed, he set it up with these experiments. And by the time he woke up, he woke up to a plus 19% score after 8 hours and 37 experiments. Literally improving his code while he was asleep. And Andre Kapathy replied and said, "Who knew early singularity could be this fun?" Showing more of the work that he's had auto research doing for himself.

And 4 days later, Toby Ludkkey was still playing with auto research. He says, "Okay, well, I ran auto research on the liquid codebase. It's now 53% faster with 61% fewer object allocations." He says, "This is probably somewhat overfit, but there are some absolutely amazing ideas in this," showing the iterations of auto research without any human intervention.

And even Gary Tan from Y Combinator says, "Kapathy just open-sourced auto research. One GPU, 100 machine learning experiments overnight while you sleep. You never touch the code. Just write a markdown file or a set of instructions on what you want improved. The bottleneck no longer is compute. It is your program MD, which is the instructions you're giving your agent."

But I personally don't want to apply this to codebases or AI agents. I want to apply it to my business and my marketing. Which is where we really start to see the brilliance of this auto research system. Andre Kapathy says, "You don't use it directly," talking about his GitHub and his system. "It's a recipe/idea. Give it to your agent and apply it to whatever you care about." Which is where our magic begins.

Because even Chamath Palihapatia had a potential use case which is kind of mind-bending to think about. He said the biggest threat to today's social media apps is an incredible video model. So, something that can take text and turn it into incredible-looking videos, plus TTS, plus auto research, which is where you can set up your AI agent with a TikTok or Instagram account, and have it constantly producing incredible quality content, learning from the results of each piece of content in this loop feature that I'm about to reveal to you, and improving and iterating a 100 times a night. Honestly, this thing in the hands of your competitors will be extremely scary.

So let me show you how this system actually works. It's a three-file system. Now Kapathy called these files program, train, and prepare. But to me, that's just confusing and unnecessary. We have the instructions file, which is locked to your AI agent and only used by me and you, the human actually setting the AI agent up for the task. Then we have the file or the asset that we want to optimize. This is the second file and this is the file that the AI actually gets access to because it's actually trying to optimize its performance. I'll give you a few examples in just a second. It takes that asset, tries to test a new variation of it, and then will compare this to the third file, which is a scoring mechanism. Again, the scoring mechanism is a file that is locked to the AI, because we don't want them tampering with it in order to score higher. We want the scoring mechanism and the set of instructions only accessible to us, the human, forcing the AI agent to actually do the task and optimize the asset.

So, let's get practical. Let's use a few examples to really understand this. This is the baseline of what Andre Kapathy first tested his auto research on. So he essentially wanted to improve the intelligence of an AI agent. We're going to use IQ here. He didn't use IQ, but I've used it just to keep it simple so we can understand it first. So in the set of instructions, he says, "I need you to improve the IQ of this AI agent." He gives Auto Research the AI file that makes up the current AI agent. And Auto Research will take that file and create a test variable. It'll change the code and make one test. It'll then take that test of that new AI agent and compare the IQ of it to the original file. If it is smarter, it'll keep that new file that it tested and replace the old file because you've improved IQ and it will loop it again. It will take that new improved AI file that has higher IQ, make another change to it, and test it. This is basically evolutionary biology and natural selection but in the machine world. Now if its test variation doesn't beat the original IQ of the AI file, it will revert back to the original file and try again. And this is done in 5-minute loops repeating indefinitely until it reaches a certain goal that you have or until a human comes and stops it.

Okay. So what parts of our business can we actually apply this auto research to to skyrocket past our competitors? Well, Eric Sue had a very interesting article on Twitter. He says, "Kapathy's autonomous AI can make you 701 times faster. It's the future of business, not coding, but business generally." He says, "Most marketing teams will run 30 experiments a year, but the next generation will run 36,500 experiments per year easily, and they'll run the experiments while they sleep using the auto research tool."

And so technically this order research tool could be pointed at any part of your business. But there are some criteria of what it works best for. So there's three must-haves and then three nice-to-haves. And if you fit this criteria, you can run order research on that part of your business. We're going to go through a lot of examples together in just a second.

The must-have rule number one is it needs to be scored objectively. So if you're like, "Make this page look better," there's no objective measure. If you said, "Come up with the best video idea," there's no objective measure. "Come up with the funniest joke." How do you measure funny? Well, is it the most laughs? Now, you're starting to get an objective measure, but that would be, "Make a joke that gets the most laughs and measure the decibel volume." You need that objective measure in order for AI to score it without a human in the loop. So, things like load speed of a website, excellent. Number of impressions a piece of content gets, excellent. Click-through rate on a page, excellent.

Then, rule number two is you need a fast feedback loop. You need the results in minutes or maximum hours, not weeks. For example, a load speed on a website. Once again, you can test that in seconds, which means you get more iterations and more improvements and it's going to actually work for you. Or email opens. How many people are opening that email that hour? That will pass. However, SEO rankings. You make a change to your website and be like, "Let's wait to see Google reindex this 10 days later." It's not going to work for you because there's too big of a feedback loop for the AI to actually get enough data to learn. Or pricing. What if I lower my pricing? Is that going to reduce my churn in 6 months from now? Really hard for an AI to actually have a feedback loop and iterate on that.

Number three, the AI obviously needs access to change it. So, if it's an HTML file or an API in a software you use, excellent. It's got access to the asset that you need. However, if it's a video that's been already published on YouTube and you're like, "Oh, change the intro." You can't log into a past YouTube video and change the intro because it's already published and AI cannot have access to that.

So, if you tick the box on those three things, then you want to have a look at the nice-to-haves because this will make your order research even more powerful. You want a high volume of feedback. If you have a website and you're getting 50,000 impressions per day and you're changing the ad copy on the website, incredible. You're going to get a lot of data, a lot more iterations, and therefore a lot more improvement in your website conversions. If you're only getting 50 impressions a day, it's going to be really tough. You're going to have to wait a lot longer to test it.

It does also need to be cheap to fail. So, if you plug an image model, so if you're looking to do graphic design and you're plugging an image model, say Nano Banana, into your AI and you're having it create images and then scoring them based on whatever rating system, we'll get to that in a second, you are using, that's great. That's going to be relatively cheap for you to generate those images. If you are having your AI literally go out there and hire graphic designers and then wait to see their work, it's going to be too expensive. You're going to have to pay thousands of dollars for these graphic designs, not going to work. Your iterations need to be fast, they need to be cheap, and they need to have a lot of volume.

Number six, you need a consistent measuring stick. So, we have a file. The scoring system is a file that the AI cannot touch. It cannot manipulate the goals to say, "Oh, yes, we did it. We improved it." What your definition of better is, which we said at the start, must stay the definition the whole time and AI cannot manipulate that. But also, the scoring mechanism that you put in that file must be objective and it must be consistent. For example, if you split-test an email to fresh audiences, you're going to have a consistent measuring stick. However, if you're emailing the same list over and over with new different titles, they're going to have list fatigue because they've had six emails already and therefore the seventh they're going to be way less likely to open because they're not going to open seven emails in a row in a period of 35 minutes from you.

Let's go through some examples together of things we can point this auto research at. So, coding efficiency, this is the obvious one. This is the one that Toby Luki and Andre Kapathy have also done. How do you make your coding, how do you make your website faster? The asset to optimize would be the source code of the actual website and the scoring system would be the runtime in milliseconds that it takes to load once booted up.

Cold email outreach. Nick Surv had an excellent example of this one where he has a cold email outreach company that is using auto research to test the titles of the emails and the body of the emails. What's actually in the content? The instructions would be, "Get more replies." This is what you're telling auto research to do. We need more replies. That is the metric that we want to score it on. So, a positive reply rate is the scoring system. The asset to optimize is the actual email, the subject, the opener, and the call to action in that, which it could just be spinning up tests and shooting it to 100 people in the first email, 100 people in the second email because these are cold emails. It's not to an actual list.

Instagram DM outreach. Your instructions to your order research could be, "Book some more calls, stay human, and don't spam." Of course, these instruction files, which I'll walk you through in just a second, how to build them yourself, are going to be much longer than this. This is just a snippet of an example just so we can start to get some ideas on how these would work. What to optimize? You'd optimize the DM script. What's the scoring system? How many replies you get and what the booking rate is per template.

Website load speed, an easy one. Sales page copy, an easy one. Video watch time, what videos are holding the viewer retention the longest? YouTube titles and thumbs. You could plug this into your YouTube dashboard and test your thumbnails and titles over time, as long as they're getting enough data to them. What metric would you use? The amount of views or the watch-through rate or the click-through rate is the one that I would actually use.

Sales scripts. Now, this is starting to get to the longer feedback loops. And Andre Kapathy says you need 5-minute feedback loops. But when we're applying this to business, we can be more lenient. It's just going to take more time to get that feedback loop connected. But you could do something like sales scripts and have the AI analyze which of these scripts are giving the best close rate. Sales funnels, app speed, cart checkouts, prompt engineering, what prompts are giving you the best results? Your agent intelligence or your agent AI effectiveness.

And so let's run through some of these examples actually in the real world together to show you how you would do it at home for your business. So really important to understand once again, we have an instructions file. We have an asset to optimize. We have a hypothesis that the AI makes and then tests, scores it to see if the actual test outdid based on the scoring mechanism the original asset to optimize. If it does, it keeps it. If it doesn't, it throws it away and goes again. And it repeats indefinitely overnight, which is critical that you have that as part of the instructions. So that literally works non-stop on your behalf.

And so I've made you a master prompt for it to walk you through setting up this exact system. All you have to do is copy-paste this prompt into your Claude code. And he will literally set up the three-file system for you and ask you what asset it is that you're optimizing. So when you walk through it, he's going to help you get all the files, the connections, the APIs plugged in so you can start your auto research on whatever asset it is you're trying to improve. It has pulled Kapathi's rules. It has modeled Kapathi's GitHub. It is the exact prompt that you need to start auto-researching in your business.

So we're going to walk through three live examples together. Starting with the easiest and the fastest iterative loop, making a website faster. So this is a local file I have. I got Claude code to mock up a website, but it's not fully optimized. It doesn't load as fast as it possibly can. And what a developer would normally do or I would normally do as a business owner is like, "I want my website to be quicker." Hmm, let me think about how it's going to be quicker or hire a developer. And the developer is like, "I have a hypothesis. What if I change this?" And test AI is handling all the hypothesis and all the testing. All we have to do is come to our Claude code, paste in the prompt that I'm giving you in this video, and Claude code will respond, "Hi, I'm now your auto research engineer. Here's the deal. In one breath, we pick one thing in your business and he turns is it good into a single honest number and then I sit here all night changing it, scoring it, keeping what wins and trashing what loses." Fantastic. That is auto research.

So, we want to improve the speed of a website. I said, "Cord says that's a great pick and an honest one. Website speed is a textbook auto research target. It is an objective measure. It is fast and is reachable. It can have the page and HTML and the JPEGs all in one place locally," which I'll show you in just a second. But remember, the fit check is for the ideal thing to auto research. We're kind of stretching the limits of auto research and applying it to business. And therefore, you may have a slower feedback loop or you might be trying something that, you know, even if you do have a 2-hour feedback loop or a 24-hour feedback loop, it might be better than you having to go do all the testing yourself. So, we're sort of on the cutting edge at this part, but at least for this example here, it does fit Kapathy's criteria of what auto research would be good for. So, I linked him to the website, which I just showed you, and he set off on his way. I let Claude pick scoring, which is any scoring. It is just the milliseconds, the time that it takes to load, and he's not allowed to touch that scoring file as we've discussed. And basically he analyzed it and did round after round after round making the website faster and faster. Now I also asked him to mock up a report to show me exactly how he went, which you can also ask your Claude code to mock up. This is a classic Claude code looking file. But you can see it went from 800 milliseconds baseline all the way down here to 90.5 milliseconds according to Claude code. And these are the rounds that it actually knocked off the milliseconds of uh and exactly what it did to knock them out. So now we have a pretty incredibly optimized fast website.

Okay. So the next example is cold emails. I saw Nick Serev who does a cold email business uh give an example of how he was using this. And so basically say I have 300,000 emails uh that are cold emails that I want to reach out to and do cold outreach, which cold outreach is very different to a warm list. You're not going to get your 40-50% opens. You're going to get your 1-3% opens. But a lot of people make money that way. I do not. But if you are in that camp, you can now use auto research to improve something like your email bodies, your email headlines, or just overall your click-through rate, your reply rate, or how many people actually buying your product.

So, I came back into a new Claude code window, pasted in our prompt. He says, "Consider me your one-person R&D department." Fantastic. He's going to run auto research. I sent him back, "I want cold email subject lines tested or improved, should we say. I want you to measure 24-hour open rates across your sent emails. Cross-reference the scoreboard or the scorecard of higher open rate." So, we're just changing the subject line of the email. We're going to send the same exact email. And how he's going to know what wins and loses is he's going to send a bunch of them, a thousand of them in this example, and he's going to look after 24 hours how many of them got opened versus the last headline, and he's either going to bin that or he's going to replace it and make it the best holy grail uh email subject line.

Now, if you use something like Smart Lead AI, you can plug it into that. You can also use auto research for something like calling. If you have an auto-dialer or if you have AI dialing agents, you can plug and play new scripts, new openers for that. Or if you're doing uh DM marketing, you could use something like ManyChat to do your auto research. If you do want more help integrating AI agents or something like auto research into your business, come and join our private community where you can get all the resources and step-by-step help for people like me or even our top AI researcher who is in the group 12 hours a day helping people out.

Now, for this example, 24 hours is a lot longer period of time and we can't actually wait weeks to show you this one example. So I asked Lord to mock up what it would look like if we ran this for weeks where it's starting to score and give open strength predictors based on what email subject lines that it is actually sending. And of course the only thing missing here is the actual how many people of those thousand people opened that email. Feed it back in to make sure that every change is making is actually making the email be opened more and more.

Okay. The last example, I once again pasted in the prompt and then asked it to improve my Facebook ads. Now, this is where things get really impressive and really scary. This is what Chamath was saying, how if you have something that can render images, can render video, can render really good copy. Go out there, post a live Facebook ad, testing that new copy or testing that new image, seeing how many clicks we got for like $10, for example, and then binning the ones that do bad and then testing a new one every five or 10 or 20 minutes or even an hour. Even if it takes you 2 hours to get to spend that $10, it doesn't matter because you're constantly testing new data in the background. And I saw a guy do this with his actual Instagram and TikTok short-form content where he had an AI rendering the short-form videos and then an auto research learning how well they did, coming back and giving him basically machine learning feedback on the next one. And he's doing really well. We're actually making a video over the next couple of weeks on a system like this. So make sure you hit that subscribe button if you haven't already.

But for Facebook ads, I said, "I want you to plug straight into my Facebook ad account and make variations of the ads I already have. I want you to test them with $10 each. The thing I want you to measure is the price per click of the ad." And so once again, this is going to be a longer tail thing to actually show. And we're going to be running these tests in our premium community if you want to see more of the actual results side of things, but I will be posting a lot more here on YouTube as well. We had some mock data. So I gave it my Facebook ad data. It started to create variations and it has a predicted cost per click for each one. It mocked up a directory and file structure for me. It's got its scoreboard, how it actually judged how well things go, and then it's got its experiment logs where you can see all the losing files. And this thing is literally going to run while I sleep.

Let me know in the comments below what you're going to be using auto research for. Thanks for watching. I'll see you in the next video.