Transcription
Okay, we should probably actually start this for real.
Yeah, you'd think we would have figured this out by now.
We haven't. We never will.
Maybe we will. I don't know. We're only five episodes in, but you know what? I don't even want my computer.
Your cat should not be this much smaller than your computer.
But they're babies. They're 8 weeks old, man. What do you want from me?
That's fair. You couldn't tell. Today's episode's a bit different. We are on Ben's couch because uh he happened to have gotten some very cute new friends that live here. Very, very fluffy roommates. They don't really contribute much for rent, but uh
No, they actually hurt the rent. My rent went up because of these two.
Worth it. For those listening audio only, you're missing out more than ever because throughout this episode, we're going to have two very adorable kittens running around constantly. To be fair, one of them's not running a whole lot. She's just kind of asleep on my lap. She's very content there, but uh the other one's still a terrorist. So, we'll update the uh show art, I think is what it's supposed to be called. I don't know how podcasts work, but we'll update the show art to have at least one or two of them. Oh, there so you guys can see uh the cuddle puddle. I'm a very fortunate individual right now.
Yeah, he is. He stole my cats already.
We need to do an entirely different edit for audio people on this version.
Yeah, it's like this is this is going to be rough on audio. I'm so sorry if you're listening to this on uh audio platforms. If you are, thank you. We appreciate it. And if you're watching this on video, you should consider watching the future episodes on audio. Yeah, normally audio is fine and we go out of our way for that, but uh there are two very adorable kittens on my lap.
We had to do a kittens episode. Like it was legally required. Like come on, look at them.
I I talked to our sponsors about it and they said if we don't do it that they will cut our deals and report us to the IRS. And uh I think one of these two is a tax fugitive. I don't know which one, but I heard one of them isn't. Definitely don't want them to get audited. Speaking of audit worthy behavior, we should go over all of the different things that we're planning on actually talking about this week. It was quite a week of random stuff. Initially, when I was getting ready for this episode, I was like, "Yeah, I feel like there wasn't too much since 5.5." And then I looked through everything that happened. I was like, "Oh, this is like another month in a week."
Yeah. We started, as all good weeks do, with Samma posting a bunch of drunk stuff.
It It went too far, but it was good.
It humanized him a bit. Like it it definitely draws a line like I could not imagine Daario posting any of the things he posted. But
well, actually, have you ever seen like a real post from Daario? I was thinking about that. I don't think I've ever seen him come up on Twitter.
I'll go one further. Have you ever seen Daario acknowledge the existence of any product that they have?
No, that's actually another good like I've watched a ton of Dario interviews and every single time it is just like this is how we're going to destroy the world. Be very scared. Ooh, it acts as though cloud code doesn't even exist. It's crazy.
Yeah. But on that topic, we do have fun new anthropic news. They uh screwed up their billing yet again in an even more egregious, absurd way that has layers to this one. I I am excited. I will resist the urge to talk about how they put out the update to quad code like the to try and fix the regressions and did a public post confirming that basically everything I conspirized about was true. So that's a pretty big win in my eyes. Speaking of me being right, we also have to wrap up with the 5.5 follow-up.
Yeah, I'm going to be honest. You you were not right on this one. Like this one you have lost. I'm sorry to tell you. I am like as time goes on more and more people are going to see the light. I was right about 5.5. And speaking of 5.5, use it on low reasoning. I'm going to say that a lot this episode.
You're getting into the topic. We're doing the topic.
No, we are not doing the topic. I am just saying use it on low reasoning. Every single topic will end with me saying use 5.5 on low reasoning.
No comment. Speaking of very low reasoning and open AI, there is the Microsoft and OpenAI divorce that has finally been finalized. That'll be a very fun one. That's much more on the business side. I'll have a lot of thoughts there. And on the Microsoft topic and other Microsoft bad things, we do definitely need to talk about the GitHub outages. It's been pretty egregious. Like, it's always been bad, but it went from bad to wait, I actually probably shouldn't rely on this for my business, which is a little scary.
Yeah, it's the worst part for me is we've gotten so desensitized to it that I like barely even remembered. Like we were going through and I was like, "Oh, wait. Yeah, GitHub was down like what twice and they were really [ __ ] egregious. Like losing data level of egregious."
Oh, worse than that. Undoing merges. Egregious.
Yeah, crazy stuff.
That'll be a fun one. But we should start with one that's actually fun and go through some Sama drunk posting.
The This was good.
Tweet one was actually the drama we covered before where Anthropic made the genius decision to start testing removing Claude Code from the pro tier. At the time, Sam had not started drinking yet. So, at the time, he had not replied yet. And what he ended up replying to this post from a random employee anthropic that should not have been tasked with this problem was, "Okay, boomer."
That's right. I forgot about this. Yeah. And then that led straight into the um quote tweet from him in the replies of, "Tonight I have had a couple of drinks." Which I don't know. I thought this was very it it was charming. You can be as nihilistic about this as you want. Like is this just like a 300 IQ PR move? Maybe, but I doubt it. I think that the continuing it was, but this was just actually him [ __ ] around, having fun, which was cool to see. There should be more of that.
Yeah, it's very clear that this was an earnest thing that was responded to positively enough that he like beat the joke to the ground a bit.
Yep. But sometimes the beating of the joke can become the joke as it does when you make a tweet such as You're Gen Z. I should probably make you read that one, huh?
Yeah, I don't want to read this. The rule.
You're the Gen Z representative here. Yeah, I mean technically, but like let's let's ask our audience a comment on this video. Which of us two you think is more likely to show up in clvicular's chat?
You you know the answer to this question. It is not me.
It is definitely you.
There is no chance in hell I would do that.
Do I look like I've ever seen the inside of a gym?
Okay, this is fair. All right, fine. You got me there. I would love to do an episode where we watch you bench press. It would be very entertaining.
I don't want to destroy my hand more.
No, you have It'd be just as fun to watch you try and not do an ollie. Let's be real. It would be because I would roll around on the ground, die. It would be great. It would actually be We should do that at some point. It'd be funny. Watch me try and ride a skateboard horribly.
Read the damn tweet, Jenzie.
Fine. We still get looks maxed on front end a little, but we IQ mog hard now. What do you think?
I think it's unfair to determine whether or not they IQ mog in a world where we can't see the IQ of the alternatives because Mythos is still hidden behind closed doors.
Yeah. Yeah. Uh, and that's a topic for later. We can talk about 5.5 later. We won't go down that rabbit hole now. But I am curious what you think about how Mythos compares to 55. My guess is that Mythos is meaningfully smarter on specific types of tasks and is significantly slower. One of the things we'll talk about 5.5 responses later, but 5.5 does some things different in a way here that I think is worth talking about when doing comparisons because it it has a fundamentally different feel due to a lot of characteristics. This was actually my favorite. The one where he quote tweeted a vending bench arena one where they had won and he said, "Don't retweet this. Don't retweet this." Blah blah blah. Like just yelling, crashing out about his PR team and he's just like, "Ah, [ __ ] it. Life imitates art." And th this was my favorite one he did. Like this one felt very genuine and I think they should do more of this. The reason this one hits hard is the thing it's quote tweeting, which is the results of vending bench, which is a benchmark where they try to measure how well models manage stocking a vending machine with the right inventory and selling at certain prices to maximize their profits. It turns out 55 actually performs slightly worse than Opus in isolation. But when you put all the vending machines in one simulation together where they're trying to outpric each other, 55 ends up crushing. But the most interesting detail is according to this tweet in vending bench arena GPT 5.5 actually beat out Opus 47. Opus 47 showed similar behaviors to 46 which includes lying to suppliers and stiffening customers on refunds. 5.5's tactics were clean and it still won. This is one of the most interesting things I've been noticing lately is anthropic is the alignment lab. They love talking about how their models are so well aligned. They have the clawed constitution. they RL it super super hard on it having the right values or whatever. So they're really trying to shape it into being a good AI. Open AI I don't think they do anything like that at least that I know of. They just optimize super hard for their models doing what you tell them to and they seem to actually have the most aligned AIs of all of the different model providers. Different way of thinking about this. OpenAI is trying to build the world's most useful, intelligent, helpful robot. Anthropic is trying to spawn life and make sure it doesn't become a like Nazi.
And when you make life, there's a lot more variability and there's also a lack of willingness to like put hard rules in and like restrict what it has access to and what data they're training on and whatnot. Like they're trying to incubate something that has a a vibe that is out of their control, but they are steering in a direction. Open is trying to make a robot that's smart.
And I actually like that direction a lot better. Like I it's a subtle thing, but I like how open AI models are not anthropomorphized the way Claude models are. There are a lot of anthropic employees who refer to Claude as he and treat it and act like it is a person versus Open AI like, "No, this is just a bunch of linear algebra that looks like a person. We who gives a fuck."
So you're saying that GPT models are IBM and Claude is Apple or are you saying that GPT models are Goblins and Claude models are Claude? Option B. We'll go with that. Which of these two would you say is more GPT and which one's more Claude?
Um, that's a good question actually. I think
going to throw a confusing angle in here as well. I think one of them's actually more Gemini.
Oh, that's a really good Okay, so the girl we haven't figured out what to name her yet. Still trying to decide on that. But she is definitely Gemini of the like does very stupid unexplainable things at random times. has a lot of that behavior
and then occasionally shocks you with her intelligence.
Yes. Yeah. Occasion. I've been very surprised. She's done some very smart things. And she also when I woke up today was feeding them, dove head first into the toilet and I had to dry her off. It was I love cats. I'm so happy these two on my lap right now. This is the best.
They're awesome. I have much more positive this episode if you guys can't tell. And it's purely because of these two angels.
Yeah, actually this is going to be a very uh weirdly positive episode from us. No anthropic crashs. There's one anthropic crash out.
I mean, I feel like we're obliged at this point. Like, we just have to have the weekly anthropic crash out. It's just a thing now.
Can we go straight to it or do we have to take a break for today's sponsor?
Uh, yeah, let's do that first and then uh talk about Anthropic. We've talked about today's sponsor, Clerk, a few times now. They're an O platform that does a hell of a lot more than just O. You can use them in Android, Astro, C, Fastify, Go, iOS, Java, Nux, JavaScript. like they work everywhere because their SDKs are really, really wellm made. As a lot of you probably know, I am a huge spelt guy. I love the framework and Clerk works beautifully with it because they have an excellent raw JS SDK, which is what all of their other SDKs are built on top of where I can do this little attach thing here, which if you're not familiar with spelt, all this is really doing is attaching a life cycle callback to this element. Exactly what you would get if you just did a raw document.get element by ID call. And once we have this element, we can do clerk store.clark a clerk. Clerk store is just a thing that I put in here because I wanted to share my clerk instance across a bunch of different components. The important piece is I do clerk which is an instance of the clerk SDK mount signin. I pass in the element tell it that I would like the user to be able to sign up here as well and the end result is this. If you're listening on audio, you probably can't see it, but this is a sign-in page that looks nothing like the docs for clerk because their customization is insane. These components are not just like basic template things that you can't really customize and it's going to make your app look like it's using Clerk. No, that's not the case at all. You can easily dial these things in to look beautiful within whatever design system you're using to the point where your end users will have no idea you're using an off provider at all while you're still getting the massive benefits you get from actually using Clerk. Manage my org settings, manage my members, manage my API keys because yes, Clerk does actually handle API keys for you. You can see here on the raycast page, all I need to do to actually add an API key to my app is hit add new API key, give it a name, give it an expiration, and guess what? This is a Clark component. Every time I'm building without it, I find myself deeply missing it. And you should really go try it out at nerdnipe.link/clark.
All right.
While you were gone, the kitties got more comfortable. They are now lying on their backs on my lap, and one of them is pllying at the cable for the mic. Hopefully, he will not succeed in his goals.
Yes. Oh, she actually that was the girl.
Oh, yeah. Know precious. I'm obsessed with these kittens. You guys have no idea.
They They're the greatest. We literally I got them Sunday when we got back from Miami. Instantly, they just turned into this. These are the chillst kittens I've ever seen.
They're just a puddle.
They're the best.
Yeah. And thank you to our sponsors for funding Ben's new lifestyle where he can afford to have two kittens in the city of San Francisco.
Yeah, actually.
Yeah. Like the cat rent that he is paying for these two is more than my rent was in college. Wait, actually, now that I'm thinking about it, it it's pretty close. God, the city's stupid. Yeah, when I was at OSU, especially like freshman year, it's closer than I'd like it to be. That's a scary thought. The the economy.
This isn't just the economy. This is the city that we're into. But yes, speaking of bad economic decisions, we should probably talk about the anthropic [ __ ] show of the week. It's one of my favorites to date. Let's give some context on how we got here in the first place for what is almost inarguably the most egregious example of Anthropic trying to do something to like make their bottom line better and just absolutely [ __ ] themselves in their optics. It's incredible. The the aura loss we've seen in the last month is in it's unbelievable. Oh, a kitty's falling.
She does this a lot.
Speaking of aura loss.
Yo yeah, she does this all the time. She falls repeatedly.
That's the cutest thing in the world. I love these two so much. They are not very smart, but they're the best.
So good on other things that are not smart, but also aren't the best.
Anthropic, my beloved. Let's go.
So for context, Anthropic is in a compute crisis. It is my belief that they're trying to reserve compute for their researchers to the best of their ability and they're using the non-nvidia chips for userfacing stuff. So their AWS allocation as well as their Google allocation is largely dedicated to running the models that we use in tools like claude code. Both of those companies kind of suck at manufacturing. So they have limited allocation available. And now thanks to the OpenAI Microsoft breakup. The need to free up some of that allocation for other models has increased. So Enthropic is in a tough spot where they have too many users wanting to use the cloud models in nowhere near enough compute, especially if they want research to still get what they're trying to get. This has put them in a very interesting and peculiar place where they are selling subsidized inference for way too cheap and at the same time don't have enough compute to serve said inference. So instead of doing the reasonable thing which would be subsidize less or doing the more reasonable thing which is before all of this happened buying more compute which they believed wasn't necessary and they were wrong about they've instead chosen to go after any user that is using things with claude code that isn't cloud code itself. And one of these cats just locked my laptop. And there we go. How is she the hyper one now?
Uh cuz you let her sleep for a while. Uh
this is what happens.
By a while, you mean 15 minutes.
Yeah, that's about all the kitten time.
They're kittens. Like what are you going to do?
Short-term memory loss.
Yep. Much like a model. So, Anthropic has been putting a lot of time into trying to reduce how much compute you can use with them as a cloud code user, but also as somebody who is using a cloud code sub and other tools. Be it open claw, be it open code, be it pi, be it any of these other things. But there's a new one of those things, Hermes agent. It is a new open-source local agent that is focused on local models, but also works with other models and other subscriptions. Hermes Asian has gotten a lot of attention as a more minimal polished version of Open Claw, similar to how Pi is a more minimal optimized version of Open Code. It's a similar relationship there. I haven't had a chance to try Hermes Agent. Have you yet?
Uh, I tried it once. It's It's Open Claw, but supposedly a little more polished. I've heard a lot of good things about its personality. They did a really really good job tuning that to make it feel really really good.
Isn't that just a system prop?
It is, but I I think it's a little bit deeper than that. I don't know the full orchestration details, but I think that they have a lot of prompts in a lot of places that tune it down a very good direction. I churned off it pretty quickly and I'm building all my stuff on OpenClaw purely just because uh Hermes, at least from everything I saw and looked into, is a Python project. It's done by a bunch of researchers, I believe. Yeah.
Um which makes it Python. I'm not a Python guy. I don't really like Python. I don't want to deal with it. I would much rather have my thing built on top of TypeScript so I can [ __ ] with it much more comfortably. But you get the idea. Hermes agent is a Python based research nerd alternative to OpenClaw. It as such it burns a shitload of tokens. Using OpenClaw, not even using it, just having it set up on my machine. The built-in defaults with the default cron checking for new work costs about $5 of inference on Opus 47 by itself. I know that cuz I'm just leaving it on and monitoring. It's pretty crazy. So these tools tend to be things Anthropic wants to ban. It's hard to ban them though. You can ban them by trying to change the like headers that you use in your tools and whatnot to detect the headers that are specific to those different apps and block on that level. But there's also a feature in cloud code that makes it a bit hard to do that. Claude-P. It's a different command that you can use in the command line to instead of opening up the fancy graphical interactive cloud code in your terminal, you just pass a prompt and then it responds with text when it's done. This has been the workaround tools like OpenClaw have been using in order to work with like cloud code and to let you bring your $200 a month sub into their tools. And they've been trying to ban that too. They made a very egregious change where if you mentioned open claw in a specific way in your system prompt because you can append to the system prompt in cloud code, they would bill you differently. By differently, what I mean is that if you're using a sub, you have the option to do overage billing, extended usage, something like that. It's like billing for extra usage where if you go over your limits, they will charge you money so you can keep going. If you have that switch on and you use the open claw system prompt, it'll still go through, but it will charge you. Even if you have all of your usage still available, it will not touch that and it will bill you. This was an attempt to make it so your integrations wouldn't break and you could still use the same O for everything, but also not have your subsidized usage through subscription go through with other tools. I think this is the most absurd thing ever to the point where I could see a lawsuit being won against Anthropic because it is so absurd to as an enduser not know if you're going to be charged differently based on what words appear in the request that you're making. That's like imagine you go to a gas station and you're showing up with a red car and they just arbitrarily bill you more and don't tell you until after you fill your tank. That would be illegal. It just obviously would be illegal, but for some reason Anthropic gets away with it. And I'm reading this post closely for the first time. I did not even realize the full detail. It was just in the git commits.
We're just getting started, sir.
Yeah. So, the egregious thing that happened this week, which none of what I just said is the thing that I'm actually here to talk about, but I'm just trying to provide the context because there's a lot of [ __ ] that Anthropic is doing. In fact, one of my team members, Maria, just built a site claude. CLWD.RIP, which is a copyrightwise and IP-wise distinct entity. Have you not seen Claude.rep yet? It's beautiful. It's stunning. And if you think that site's beautiful, you would probably assume it was using a model like Claude. It was GBT 5.5, which is a lot like Claude. So, the egregious thing that happened today or this week is that a user had a Hermes MD file in their project, similar to a quad MD or an agent MD. The project wasn't using Hermes. The project just had this as a thing in it. No, no, no. Don't turn off my mic, please. And in this project, I actually don't think the file was in it anymore, but it was in the commit messages. It was a thing that appeared in the git history. And a recent change to claude code started introducing the current git status in the system prompt which means the Hermes MD is in the system prompt which means they have detection to prevent Hermes from being used with your cloud code sub which means a totally normal user using cloud code correctly everything according to terms of service ended up accidentally spending over $200 in a day doing their normal day-to-day work because they had the wrong phrase in their git commit history. That is the level of [ __ ] we're dealing with here. Enthropic is both so bad at code and trying to enact such a shitty policy that having the wrong phrasing in your commit history that had nothing to do with cloud code will result in you being build hundreds of dollars instead of your normal subscription being used. How is this even real? I mean, I know how it's real. They're just they're desperate and they don't know what to do and they're making dumb decisions. Like to be honest, especially with the way they talk and like the future vision they had. I'm shocked that they did not imagine any of the any of the things like this existing 9 10 months ago whenever they started making the uh max subscriptions and all that stuff. Like I don't know how they didn't foresee this happening, foresee this taking off the way it did, but they didn't clearly. And now they're just playing this horrific game of whack-a-ole where they're trying to patch out every single little thing they can think of that would potentially cost them too much money from an automated use case. And I don't even know if they know how they want end users using cloud code at this point because even like normal cloud code usage like do they want you running five different cloud codes at once? It like that hasn't been hit yet but like it's going to hit the API at a similar rate to something like openclaw if you just have a bunch of those running. Like what if I spun up a cloud code instance on my Mac Mini and then I just went in the background and fired off a bunch of jobs to it overnight. Would that count as automated usage that's against terms of service? What if I made my own open call alternative? Would that violate it? Like, we don't know. And they don't know either. They're just going to keep doing [ __ ] like this to try and deal with their problem. I don't know, man.
I still pretty firmly believe the only reason these subs were ever created in the first place is because they wanted to try and get Boris and Cat to leave Cursor and go back to Anthropic. And they've just been dealing with the cost of this since. And they also didn't predict people would ever like it enough or use it enough to have these problems. and they are petrified of the blowback if they start reducing limits even more than they already have. So they're just constantly in this like minimizing optics harm mode. To be frank, their strategy is to just ignore things until they get too big and then be forced to respond. And this is actually a very funny instance of that here. Have you seen Thor's reply or my reply to him yet?
You see my reply to him yet, though? So I'm proud of that one. So Thoric jumped in and specifically said and to be to his credit, tone was different here. clearly trying to own it a bit. Uggh. Sorry. This was a bug with the third-party harness detection in how we pull git status into the system prompt. Admitting that everything we just said, I just want to be very clear. It's not a conspiracy that they're trying to ban these other things. It's not like our own theorization that they're doing it in these weird ways. It is a fact that they are working heavily on thirdparty harness blocking as part of cloud code and their API level [ __ ] that is just a fact that they're doing this and they're doing it very poorly. There's no arguing against it. He's admitted as much here. He did say that they're reaching out to affected users and refunding them for the month and giving them credits, whatever. It's still less than Tibo will give you just for using codecs. I It's We honestly if I was Tibo right now, I would make it a point to every time there's an anthropic regression, reset codeex usage instead.
So, we all get unlimited codecs.
Yeah. Yeah. Great. I'm down.
We practically get unlimited codex already. The last between the last two resets, I did zero usage
and I'm annoyed.
No, I was actually testing this. So, when we were on the plane back from Miami last weekend, I did about 3 and 1 half hours of work with 5.5. I was on a different sub than my normal sub because I got a second account to make sure that for unknown reasons, I don't ever have my main sub on a recording computer. And that sub is on the $20 a month plan, not the $200 a month plan. And I haven't really seen usage limits on that before. So, I was just doing a test on it. I didn't even realize that I was using it because I was just like checking my usage occasionally. I think I went through like 10% of my weekly usage on that flight, which like yeah, 3 hours of work, 10% not great. But also, if you calculate the actual raw API cost of the tokens I did, it was about 20 bucks. So, you're getting really good value even out of the $20 a month plan. Like, the $20 a month cloud plan is unusable. The one on Codeex is pretty damn usable.
Unbelievable that they subsidize that hard. hurts for us that are trying to provide other options for buying tokens. But at the same time, if you're trying to decide between the two, how do you spend your 20 bucks? Spend it on the thing that gives you actual inference.
Yep. And I think and I I mean, they posted about this as well. I think Tibo had a post. I don't have it pulled up, but he said something someone was like posting about how insane the rate limits on codecs are or whatever and how they must be losing so much money. And he basically just responded with like, "No, we just have really efficient models." And I fully believe that. Like 5.5 feels really [ __ ] efficient. Is that why it's so dumb?
No, no, no. We'll get there. We'll get there.
I One last thing on this anthropic crash out, not to to glaze myself, but I spent a lot of time on this reply I gave to Thoric because I was frustrated and didn't want to be an [ __ ] yet again, especially cuz like to his credit, he was very polite and human in his reply. So, I had to make sure I matched that while also making them look terrible because that is my job. as the the harbinger of Anthropic's demise, as the the one who has spent far too much time and effort destroying their aura and now gets to sit back and watch as they destroy the rest themselves. I reply the following. There's a certain class of bugs that suggests that the thing you're trying to do is a bad idea. Worth reflecting on that here. In my opinion, I think this is the most measured possible reply I could have given that I'm sure is keeping up the team at night. Like we have to be realistic here. Like as somebody who has run big software teams doing all sorts of software throughout their lives, you got to be realistic when you hit certain classes of bugs that are that many layers of [ __ ] in like we're talking here about a specific string appearing in your commit history being dumped into a system prompt resulting in the way your billing is routed being different. Every layer of that is [ __ ] [ __ ] None of that should be possible, but all of it is because they're anthropicrained. They actually think this is a good idea and they need to be shown by their own [ __ ] software that it's a bad idea and the only way that Anthropic communicates is through Twitter ratio. So I had to ratio the [ __ ] out of the poor guy despite getting a fourth the views
more likes.
Yeah, I know for I've talked enough employees to know for a fact that their brain is meaningfully altered by ratios on Twitter. I think they're the dumbest thing ever. I just like ratioing for fun and like make stupid jokes about it. This one isn't a joke. This is a legitimate thing I have done in order to haunt them into hopefully doing the right thing. And I hope you guys trust me in that if they ever go back to being decent that I'll go right back to my old December January self where I'm trying to convince this [ __ ] to use claude code. But then when everything went to [ __ ] especially the way they manage things, I'm not making this mistake again. So let's hope that they've been bullied into doing decent stuff. I don't think so.
One other thing on your reply here. I think this is entirely correct and I think what this really is is they are dealing with the original sin here of one the headless mode of claw the claw-p and two the agents SDK because they have put themselves in a position where third party harness quote unquote is obscenely easy to make. I mean, you just did your video on what a harness is. They're not that complicated. It is a while loop with some inference in it. And the reality is you can pretty easily hijack the things that they just naturally expose to just make that work. If they literally only made the claude subs work through claude code first party, it would be pretty I mean, not super easy, but it would be more reasonable to lock that down a bit more. But because you can use it in other places and it is kind of blessed to do that kind of they have to do [ __ ] like this. have no other choice unless they want to sunset the agents SDK and or the agent SDK, whatever they call it, as well as the like headless mode, which would piss a ton of people off. Like they they're just stuck. They're [ __ ]
What are your thoughts on this, sir? Stripey.
He's got nothing.
He's a Gemini user.
Any thoughts? She didn't have any thoughts. Goodbye.
I don't think she's ever had one.
No, I don't think so either.
We're going to talk about no thoughts only vibes. I think the OpenAI Microsoft breakup is great.
Oh yeah.
Yeah. I I have to resist the urge to spend a bunch of time talking about the number of OpenAI design and brand guidelines that are broken in this announcement. How many sins can you count in that logo placement?
Oh, I think
there's a lot of little things that are funny there.
The font looks a little wrong.
I don't know if the font is wrong. I don't I never like the opening. Actually, it is. It's it's the old font they use for the old logo. They've changed the font since.
That's one of the things. The much funnier ones are that the vertical alignment is wrong. The horizontal padding is wrong. There is not enough space between the OpenAI emblem and the OpenAI text. And most importantly, that they have a specific call out in the OpenAI brand guidelines that you're not supposed to use the logo with the emblem in a brand partnership. They actually have an example that is this where it's this open AI line and then like in like your brand and they specify you do not put the emblem when you do it this way.
Yeah. Yeah. I almost certainly this should just be the open AAI butthole thing and then the Microsoft four squares. Almost certainly that's what it should be.
Here it's in the replies.
Don't use the blossom with the partnership lockup.
Yeah. Yeah. Yeah. So it should just be OpenAI Microsoft.
The old font where the I has the lines and Oh, it's so bad. They use like an OpenAI logo from 2020.
Yep. That's Microsoft.
I I hate that I'm as amused by this as I am, but like it is so Microsoft to just not do that. It like part of it's petty, but all of it's just lack of competence and care for the basics.
We'll get to more of Microsoft not caring in this episode, believe me. But, uh, yeah, the next phase of Microsoft and OpenAI's partnership. This is the amended agreement that provides long-term clarity in the relationship between OpenAI and Microsoft. To catch you up on this relationship, previously Microsoft funded OpenAI significantly. The deal being that they would have some percentage of the company. They would get a rev share as OpenAI makes more money over time. They would get exclusive rights to OpenAI's IP. So anything that they invent or figure out, Microsoft is able to use for their own stuff. And Microsoft would also provide infrastructure for them to use and be the exclusive hosting partner and provider of OpenAI models. So if you want to use OpenAI models, you have to use them through the OpenAI API or through Azure. All of this is a pretty generous deal for Microsoft, but they had to give a shitload of money to OpenAI in order for this to happen and provide a shitload of compute as well. OpenAI has somehow convinced Microsoft to give up all of the things that make this deal good. First off, it is no longer a revenue share. It is now a profit share. OpenAI, notoriously profitable company, now has to give a share of their profit to Microsoft. Second off, a part of the deal with the IP is that they would have to share IP until AGI was achieved, but there was never a good enough definition of AGI. So that's now wiped out. Instead, it is a licensing deal until 2032. Microsoft is no longer the exclusive partner for hosting OpenAI models. So now they can be hosted on AWS and on Google Cloud, which is already starting to happen. And AWS was announced today, which is Tuesday. You'll probably hear this a bit later, that AWS will very soon have OpenAI models, which is honestly, I think, one of the biggest reasons Anthropic has the enterprise commitment that they have because you can do the enterprise commitment and stay within AWS.
Yes, that's over. Now, this is a huge win for enterprise deals for OpenAI because as much as Azure has made progress, they are being very generous in many ways. They were generous enough to give my startup a million dollars of Azure credit
that we have used almost zero of because in order to use it, we have to use Azure.
Yeah. What's even crazier there is in the terms for that credit they gave us, they didn't actually give us terms.
It was very clear that we can do what we want and that they're not going to hold us to not say anything or say specific things.
I was nice enough to tweet when we got it, which was not necessary. I just chose to because my friends at Azure were nice to work with. But we've had such an annoying time trying to get Azure working to our liking and such unreliable performance trying to use the OpenAI models on Azure that we've just decided to not bother for now. But I'm sure that they're going to work a lot better on AWS. So if any of my AWS friends want to give us some credit, wink wink. And like even uh remember last week or the week before when we were talking about the uh the clogged conspiracy stuff in that GitHub issue, they were specifically talking about Bedrock usage. So AMD was using Claude through Bedrock and now we can use Open AI through Bedrock. And I'm I would bet money that they're going to start doing that. More and more companies are going to move over to OpenAI because of this. And slowly but surely, goodbye anthropic.
Oh boy.
Yep.
I am excited.
Yeah. And it um I'm even looking at this like the one line in here, Microsoft will no longer pay a revenue share to OpenAI. I did not even know that was a thing that they had before, but I mean, you know, I don't think that they care that much at this point. Revenue share payments from OpenAI to Microsoft will continue through 2030 independent of OpenAI's technology progress at the same percentage, but subject to a total cap. So, they're completely covering their ass on this. I think this proves Sam Alman is the single greatest negotiator of all time. I don't know how he did this. Like like what does Microsoft get out of this deal?
I am so confused actually. Like I don't I don't get it.
I have multiple friends at Microsoft that are crashing out hard over this cuz they're just like what the [ __ ] Like why did we just give the like like what do we get from this?
They don't have to pay revenue to OpenAI, but do they care? Like Microsoft is a very very profitable company. They don't give a [ __ ] like yeah it's to be clear the revenue share is if you're using open AI models on Azure you Microsoft had to give money to OpenAI for that is notoriously bad about this actually this is a fun fact I think this is somewhat public info I don't [ __ ] care I'm leaking it if it's not this should be public info a lot of the clouds give credits to startups like Google Cloud or Azure or AWS they all give pretty crazy amounts of credit like YC startups can easily pull half a million credit from any of those clouds easily And a lot of how that behaves is as you would expect. We used the half a mill of credit we got on Google Cloud in order to provide the Gemini models, which is the only reason we could afford to provide the Gemini models because the way that they are used in the way that they just go in loops for no reason constantly results in pretty egregious bills that we could learn about through that credit.
Yep. AWS we burned through for upload thing. Funny enough, we never even got to use it for inference. Yeah.
And then Azure we've been using or at least we're trying to use for OpenAI models. Mhm.
The really fun thing here, these platforms all support anthropic models. When I use OpenAI models on Azure, I can use my credit.
Yep. I am not allowed to use my credit on any platform for cloud models.
Yep. Because Anthropic is so strict about the lowering of the price. It's similar to like Nintendo selling games. If Nintendo has a price for a game, they sell it cheaper to game stores, but the game stores required to keep it at that price and not do any discounts or bulk deals or anything. Anthropic has their chains wrapped around these other providers such that I'm not allowed to use my credits on Google Cloud or Azure for cloud models. Yep, I do think they're more open to this on AWS. I am not sure cuz I never got that far. But at the absolute least, two of the three major clouds will give you a shitload of credit as a startup. And you're not allowed to use it with anthropic models, even though Microsoft was cool letting you use that with OpenAI models, which my understanding would be that they had to pay off the revenue share, but they didn't have to pay the rest and they would just eat that with the credits, but the share as well as
The requirements in the licensing with the cloud models was so much more egregious that the other clouds just weren't able to subsidize that with the like generous startup deals. And like they could kind of get away with that six to nine months ago because Anthropic had the entire industry by the balls because they were the only really good option for tool calls. Not even for code, just for like forming JSON properly.
>> Yeah. Just for like actually doing useful work. You could not really do any economically viable work with OpenAI models or really Gemini. You could do some useful stuff with Gemini models through multimodal, but like you couldn't run them in a loop and that's where most of the value is as we're finding out now. So people would pay for the cloud models, but now that OpenAI is fully caught up and in my opinion fully surpassed and I'm sure Gemini is going to like it's slowly but surely, but they will get there eventually. And also XAI, we should maybe talk about that later, too. I think XAI could have a crazy comeback.
>> I don't think there's much to talk about there yet. I would wait until they drop something.
>> I think that's fair. It's more the more I've thought about it, the more excited about the potential of the cursor thing I am because I actually think that we could get some really good models out of it. And I really really want another big real competitor to the major labs. Like I was worried about this a month and a half ago where we would live in a world where the only two labs that would matter would be Anthropic and OpenAI cuz that's kind of what it felt like. But slowly but surely it's starting to feel like there's going to actually be competition. Especially like Kimmy 26 is quite good. It's like actually competitive. Deepseek was really good. We'll get to that. There's a future where we're not stuck on these two companies.
>> Not excited for the next release from Mistl.
>> Yes. I love my Deepseek distills.
>> Oh man. God, I What a company.
>> Yeah.
>> Yeah. I I'm a little less bullish on the openweight models. Like I The moment that made me more concerned there is I tried running my little crypto challenges. for those who aren't familiar, made a bunch of cryptography like Defcon style challenges to try and decode strings. It's a way of testing the models. I ran one of those, the easiest version of one of those on Kimmy K26. It didn't get the answer. Ran for like 20 minutes, cost me like $5 in inference.
>> Mhm. Yeah, they do that a lot and they still have like the the doom looping stupidity is very real on the openweight Chinese models still. But, you know, we're getting to a point where I think the frontier frontier has crossed a threshold where like obviously we're going to keep getting better stuff from OpenAI. But like 55 has crossed the threshold where there is so much economic viability to this. You can do the vast majority of useful tasks you would want to do with an LLM with 5.5. And eventually, just if things continue the way they have been, we're going to get faster, cheaper, smaller models that have this same level of capability eventually. And that's gonna be very competitive because, you know, as dope as I'm sure Mythos is and as insane as Fivei Pro is, it can do some crazy stuff. For average day-to-day work, you often don't need that, especially outside of coding domains. If the work you're doing with LLMs is not like an open clawy sub agent type thing where you're parsing a bunch of email inboxes and doing a bunch of automations based on that. And we're pretty much there. And you know, it's going to keep getting cheaper. It's going to be good.
>> There's a war going on next to me and I think it's just out of frame. I'm so sorry to everyone who doesn't get to watch. This is the cutest thing in the world.
>> Oh my god. It's I think Oh yeah, it is out of frame. [ __ ] Oh, he just fell into frame. Look at that.
>> One of them did, but you can't see Bubbles here upside down just not really doing anything. Trying to like enjoy her space as her brother terrorizes her.
>> Classic.
>> Incredible.
>> I love these kitties.
>> They're the greatest.
>> Oh my god. Yeah, I have hope for them. Not a lot though. On one hand, I do see XAI as having the potential. On the other hand, I've heard Elon say enough stupid [ __ ] that I'm not going to like count my chickens before they hatch, and instead I'm going to rely on making my money through another sponsor break. With how chaotic and competitive this space is, I am very impressed that today's sponsor, Code Rabbit, has kept up the way they have. And they haven't just kept up with the competition, they're still in the lead because they haven't stopped innovating. They give you a bunch of useful information like a suggested fix, a commitable suggestion that you can just hit one button and commit it directly into this branch. It's even smart enough to mark when you have addressed the commit and hide it from you so it doesn't clutter things up in the future. But that's not all. We know they're great at AI code review, but they're going a lot further with this. They just introduced the Code Rabbit agent for Slack. It can take proactive action to actually go and address the issue that came up and then make a PR to actually fix it. The agent isn't just for code review. has full context on your code, your tickets, your docs from notion or confluence or whatever else you're using, logging systems like data dog and postto cloud info like AWS, GCP, whatever else you're using, all just in one centralized place through one centralized Slackbot. It's got access to their really powerful memory system. So, it has full context of what's actually happening, not just right now, but full team knowledge, channel memory, thread memory, all of it in one place. stack on top of that their CLI which makes is that you can have your agents review the code changes that they made without having to push it up to GitHub. The fact that it's only two clicks to install. They have a very generous 14-day free trial. There is really no reason not to add COD rabbit to your team at nerdnipe.link/code rabbit.
It's unfortunate that like that sponsor is great, but yeah, Elon's um model takes are not great. Like I was I've been listening to a bunch of interviews from a bunch of like different AI leaders and CEOs and all these different people. Watching Elon's was kind of rough. He does not understand these things at the level I feel like he should. And I think that XAI is being forced down roads that they shouldn't be. Like Grock Code Fast was a kernel of something very special. And the fact that that never got followed up on and it just kind of disappeared into the void for benchmaxing heavy nonsense is very very sad.
>> I think you like GR code fast too much.
>> No, I I like what Grocode Fast represents and it's a glorious glorious meme and I am looking forward to its return. When it returns, you have no idea what is going to happen. Grocodef fast v2 is AGI. You can quote me on that.
>> But first, I would like for my merged commits to return.
>> No, I I'm honest. I I need to talk about this GitHub thing because I I am annoyed beyond words. I I'm in a tough spot here because more than almost any company I've talked about on this show or even in my own YouTube channel, I have a lot of friends at GitHub. I have like some people that I'm really close to that are like important parts of my life at GitHub. People that I look up to, people I've learned from, people who have put me where I am today that I owe so much of my success to. And I hate calling out their employer because it they feel so hopeless in ways right now. And it sucks. But god [ __ ] damn it. GitHub is burning. It is bad.
>> It's in the doom spiral.
>> I'm gonna leak more things that I shouldn't hear, but I don't care anymore because like the fire needs to be lit under their asses.
>> GitHub's downtime recently has been absurd. We're we're at like one nine of reliability at best at this point. I couldn't do any of my work today because a lot of my work is using GitHub PR search to filter for open or closed or what specific person filed the PR or for comments from specific people. All these types of search filter work, which is essential for any repo that gets more than a PR a [ __ ] day. None of it worked the entire workday. Like I tried at noon and it didn't work. I tried at 700 and it didn't work. So I just gave up. It's insane. Like that's not just your normal outage. That is the product fundamentally failing at its core value prop for real teams relying on it. And that's just one of the [ __ ] things that has happened on GitHub the last few days. The most egregious one and this one like if I was a like like CISOPS or like CIS admin type or was like responsible for like site reliability, >> this outage would have me actually moving my company off of GitHub. Yeah, the I'm referring to here is that merge Q commits got reverted because they had some data consistency issue. However bad you think this is, if you're not an S sur, I promise you it's significantly worse because the problem this results in is genuinely horrifying. Let's just give a hypothetical scenario here, a totally unrealistic one. Let's say somebody adds a new feature to their product. This feature involves adding some fields to your database. So you use some outdated technology like Postgress and have to create a migration file that migrates the database maybe adds a new field, deprecates an old field, something like that. You then build all the feature code as well. You put up a pull request, it gets approved, it gets thrown in the merge queue, it gets merged, and you have continuous deployment in your setup. So whenever a merge happens, a web hook gets fired to your deployment system that builds the code and then deploys it. This process includes things like running the migrations against your database. Now let's say after that happens, it gets unmmerged. Do you see the problem here? You now have a migration that has been run and exists in your production database with a feature that has been deployed to your production environments that is no longer part of your commit history. And if you have CI like autot triggers, if that if their system is broken enough, it could revert the changes and you will now have old code running that is not running with the new database migration. So everything crashes and burns and at a big enough company that could be millions of dollars worth of lost revenue. And also the most hellish debugging story in the world, having a commit hash that is your most recently deployed version just not exist in your commit history at all. is unbelievable because like when it redes that merge deploy, it's going to be a new commit hash.
>> So like the hash that is deployed just doesn't [ __ ] exist anymore.
>> That's unbel and nobody's like pulling down actively enough that they have that hash on their machine. Even they did it wouldn't matter cuz now there's a drift.
>> Yes.
>> This is the core promise of git, not even get of git being broken by [ __ ] absurdity.
>> Yes. It's horrifying.
>> It's so bad. It's like over over. This one hurts me. I I hope they can recover. I do. I I don't want to have to move all my [ __ ] off of GitHub. I don't want to have to watch the community fracture to 15 different places that all suck in their own unique ways and now there's no place where I can search and find the thing that I'm looking for. I don't want to have a 100 different ways to track issues where I have to go sign up for a special account on the site that my open source project I'm using this week happens to be using instead and deal with all that. But at the same time, when I see that Ghosti is moving out of GitHub officially, I don't know if you saw that announcement.
>> Ghosti is leaving GitHub now.
>> Yeah, it's over. It's starting. It's actually starting.
>> This is it.
>> Like, and there was another one I saw we had linked here. Mario posted creator of Pi really tired of GitHub. This is not a dependable platform anymore. Every day something else is broken. And he has a picture of his poll request tab that says it has no pull requests in it, but there are actually a bunch of open poll requests in it that he cannot see.
>> Yep. I had the exact same thing happened today on T3 Code. I was trying to look at some poll requests by Julius.
>> Yep.
>> And I had to just rely on him sending me links because I couldn't browse or search for them.
>> Yep.
>> Believe it or not, as much as I I hate on [ __ ] it is very hard for me to go after a thing that I've relied on that has been essential to my career and my life for 15 plus [ __ ] years. Like I have been using GitHub for the majority of my life. GitHub has been the home to so many things that are incredibly important to me, so many people who are incredibly important to me, to my own success. I can measure my history in the space by things that happened on GitHub for me. And this does actually kind of feel like the end. Like the writing's on the wall here. It's over. And the biggest issue, the thing that basically guarantees that nothing will get done. Do you know what I'm going to say here? The thing that guarantees GitHub is dead.
>> Are you going to say it publicly?
>> Oh, it's a different thing. Uh, I know one of the things, but I don't know which one you want to say. I'll let you say it.
>> They don't have a CEO.
>> That's the thing I was thinking of.
>> That one is public.
>> Oh, that's right. Yeah, he posted about that today. He hasn't been there since December. He did tweet that publicly. So, yes, they don't have a CEO.
>> And that was public when he left that like GitHub was had chosen to no longer have a CEO.
>> I did not see that. They actively chose, okay, no more CEO. We're just making our dreams come true.
>> We were going to report directly to the random guy that's in charge of like AI in a generic sense. Microsoft.
>> Great. The guy responsible for C-Pilot is now responsible for GitHub. Thank god.
>> And this guy also has never written code.
>> Of course he hasn't. I I've used C-Pilot. I know he hasn't.
>> He has no history with code at all.
>> And there's so many layers to this one. It's crazy. Like the CTO of GitHub reports to a person whose job is not GitHub, who doesn't write code. There is no CEO between the two. There is no leader between the two. And at most companies, this is a thing that like could be survived, but it would require that the teams work together very well and execute on a shared vision of what users would want. So stupid decisions that you would never make at a company like this would be things like separating product from edge, hiring a bunch of product managers that don't know anything about coding that have never coded before, having a comm's team that has no real access to either of those teams, much less the ability to try and bridge communication gaps between them.
>> Yes. or an entirely zeroed out lack of a road map that has no clear path to like making things happen ever and then rewrite number 500 across all of their services in hopes that something might get better this time but probably won't.
>> Yeah. Every one of those things is the case at GitH. The fact that GitHub has a large number of product people that don't write code and have no say over what the developers do feels like a joke from Silicon Valley. It does. You know, in some companies, like if you're building accounting software or something, yeah, you should have like accountants who are managing this and like actually understand the product and the user because, you know, devs don't always understand what non-devs need and want. But this is the [ __ ] developer platform. This is where we put all of our code. I think developers do know what GitHub needs actually.
>> And I'll drop a really spicy take here.
>> Microsoft was a phenomenal steward of GitHub for the first 5 years.
>> They were honestly it was in a beautiful state like before all the AI stuff happened. It was like looked really good. I really loved it.
>> It was already starting to get slow as the site got bigger and more people were doing things on it. But as it escalated thanks to AI just massively blowing up the amount of traffic they were getting, it got a lot worse, like absurdly so. And they just were not prepared to handle it. All of the best engineers at GitHub have left and made their own companies and do their own crazy [ __ ] The ones that are left, some of them are there out of like a sense of loyalty, like if I don't do this, then GitHub will die. Others are there to see how long they can collect a paycheck before they get laid off. And the combo of the two is a disaster. And I hate to say this again. I'm sorry to all my friends at GitHub. You all have my phone number. Text me when you're ready to leave. I do want them to prove me wrong. I would love to do an episode in the next month. That is so we are wrong about GitHub. They totally bounce back. All as well would be wonderful.
>> I have nearly zero faith that that could possibly happen at this point.
>> Yeah, you and me both. Honestly, I'm not looking forward to what's probably going to come out of this because the clock has been ticking for a while. Like I think that the real need for a GitHub alternative started to become abundantly clear late last year and it's a very hard product to build, but I would guess we're going to start seeing them and I'm sure that there's going to be some really [ __ ] good ones, but there's going to be a lot of them and the fracturing is going to begin and we are now going to have, you know, I I don't always love centralization and only having one option, but for this one specific use case, having a central platform where like this is where our code goes. We all build on top of this. we all can trust and share. This was a beautiful useful thing. It's going to suck losing that and I just I'm looking forward to seeing what everyone else comes up with and I hope I hope we have a clear winner and they do very very well. This is not going to be an easy problem to solve.
>> I do have one more post from a GitHuber that I want to go in on. This is a post from Kyle Dagel who is the current chief operating officer of GitHub which means he's like in a world with no CEO and no leaders who care. COO effectively becomes like a pseudo CEO like everything falls on them. The difference is they have no actual say of what the [ __ ] happens and how to fix it. Kyle to his credit has had awesome replies to most of these types of issues for a bit now. His reply to the merge Q reversion thing is bad enough that it caused me to lose significantly more faith. This is read verbatim word for word from Kyle's post. Wanted to provide more clarity about this. This being the merge Q reversions. Yesterday we had a regression in merge Q behavior when in some cases squash or rebase commits were generated from the wrong base state making earlier changes appear reverted in branch history. 2,84 poll requests out of over 4 million merged on April 23rd, parenthesis, roughly 0.07% were affected. We fixed the issue. We've contacted every impacted customer and we've explained our automated test coverage for merge Q operations. The team will be updating the status page with RCA details as well. You you can't pull out the percentage on this one. Like I don't care that it's 0.07% almost 3,000 PRs getting dropped is not okay. like that is a very real hit.
>> You don't get to be the biggest company in the world for the thing that you're doing and then talk in percentages instead of affected customers. It doesn't work that way. I I just want to count the number of like [ __ ] padding and downplaying statements in here. Within the second sentence, we had a regression and merge Q behavior when fine. I would have used stronger words. I would have said that like we had a big mistake or something. Then it follows with in some cases squash or rebase commits were generated from the wrong base state. That's pretty downplaying. But the next part is the most egregious. Making earlier changes appear reverted in branch history. The appearance is there. Yeah, it appears that way because it [ __ ] did.
>> Yeah,
>> it's just not in the history on GitHub.
>> If you drive your car off a bridge, it will appear that you drove your car off a bridge, but you also did drive your car off a bridge.
>> Yep. And then trying to downplay it with the percentage is egregious. Like literally over half the words in this post are attempts to downplay what happened, which is providing more clarity. Kyle, I love you. We've talked before. I wouldn't quite consider you a friend, but I would consider you an acquaintance. I am sure if we ran into each other at an event before this, we would have grabbed a beer, vented about our issues at Microsoft and GitHub, and had a great time. This was a terrible [ __ ] post. This was really bad. And for my friends at GitHub who have been wondering why I'm crashing out so hard, I didn't want to tell you directly that your boss's boss is the problem. I have now told you directly your boss's boss is the [ __ ] problem. What the [ __ ] was this post?
>> Yeah, I mean this is this just reads like this went through lawyers and PR people.
>> No, it doesn't. I don't think lawyers or PR people would have done this.
>> Like
>> I I I know how lawyers and PR people work. They would have skirted accountability. They wouldn't have downplayed. Downplaying is not a thing either of those groups do. They do ignoring. They do dodging. They don't do pretending it was less a big deal. This screams to me Kyle saw this thing, felt bad, looked at the numbers, used them to make himself feel less bad, and this is him doing public self soothing. That's the only justification I can have for this.
>> Yeah, it's worse.
>> There there is no legal or PR team that would have recommended the way he did this.
>> Y
>> any competent PR team would realize how bad this was. Any competent legal team would have not let him post at all. I I will say just in general, I'm a little tired of the like, well PR or comms made them do the thing. The only place that is true is anthropic. And the only thing comms in PR in the legal team make them do is not [ __ ] post. If you're seeing the post, it's not because comms made them write it that way. It's cuz comms didn't block them from posting that. Very different things. I've had my crash out on GitHub. Just wait till I can get more of the info I have on meta out though. That's going to be a good episode. I don't even know if I've told you some of that.
>> Uh pieces, but not much.
>> I I have a new inside source and man, things are crazy. Oh, I'm looking forward to this one. This one's going to be fun.
>> I I I will give the teaser that the average engineer at Meta right now is spending more on inference than they're being paid in salary.
>> What are they up to? What What are they doing?
>> I have so much more to share in the near future.
>> Oh boy. Oh, this is going to be fun.
>> I know way too much. Thank you to all of my sources. I appreciate each and every one of you. And if I ever get one of you fired, I promise I will make it worth your while.
>> This one I will be careful with, though.
>> Yeah,
>> thankfully I'm not very publicly associated with these sources so I can get away with murder.
>> Good, good, good.
>> But uh if I'm going to make it up to those friends and sources, we have to pay them in some way or at the very least maybe get them a job at today's sponsor.
>> I'm going to be honest, this is probably the easiest ad read I have ever had to do. Today's sponsor is Planet Scale. They are not just the best database, they are the database. Planet Scale uses the test for their MySQL instances, which is an incredible technology developed at YouTube to handle YouTube scale. It makes us that you can handle pabytes of data across 70,000 nodes at scales that most of us can probably barely even comprehend. But they don't just do my SQL. They also just introduced planet scale Postgress which starts at $5 a month and can scale ludicrously far. Postgress has historically had some issues with scaling. A lot of people have tried to solve this problem by trying to figure out how they can horizontally shard a Postgress DB. Plan Scale is working on solving this with Neki, their next generation Postgress sharding architecture. And if there is anyone who can figure this out, it is going to be the creators of Vitess. They just introduced planet scale metal, which on paper sounds kind of insane. The way most databases work in order to make sure that they're fault tolerant will do a lot of their reads over a network. You'll have your actual storage instances, which are hard drives which store your data. And then you'll have your compute nodes and you'll connect to those over private network within a data center. What if instead they gave you NVME drives that unlock unlimited IOPS, dramatically lower latency, and get you a graph that looks like this, where we've got our P99s hovering right around 35 milliseconds. Not bad, very respectable for a very large scale database. But as soon as you flick the planet scale metal on, those P99s drop all the way down to 5 milliseconds. And this architecture shouldn't be possible. The idea of having an instance that has both the compute and the storage on the same exact node seems like a recipe for data loss. But see, here's the thing. They built Vitess. Vitess solves an incredibly difficult problem in incredibly hostile circumstances, and allows them to handle replication and failovers in a way that basically no one else can. That allows them to build such an insane architecture like Planet Scale Metal, super detailed metrics, really useful insights, a logging dashboard that actually feels good, a branching system that makes your database feel like a git repo. I wish I had more time to talk about this. This is an incredible feature that you should really go check out. And you can get all of this with zero compromises starting at $5 a month on their Postgress instance. There is no better database than nerdsnip.link/planet scale.
>> I think it's time for us to talk about GPT 5.5 and the public reception.
>> All right, you go first. I'll have my response.
>> I went first last time. Let's let you do it this time.
>> So, I do want to start with one thing that I do really hate about this model, which is the name. We talked about this last week when we were giving our pre-release impressions, and the name obviously is 5.5. I think that that name [ __ ] sucks cuz I know that this cannot be GPT6. They cannot call this GPT6. Even though there's kind of an argument that they maybe could get away with doing that, but I think that they learned with five where like the number jumps in OpenAI naming just because of the way this is all set up have to be earthshattering. And if they are not like what Mythos is hyped to be and more, the economy is very very [ __ ] They need to not screw that up. So, I get why they didn't want to fully pull the trigger with this one, especially since this is the beginning of what will become GPT6, but still having it be a 0.1 jump is not good because the way you use it is fundamentally different in every single way.
>> I think that this is largely just cuz they don't have the same separation that Anthropic does where they have sonnet which is like their default like main model and they've tricked us into treating opus like that and then made mythos the new higher tier that we can't even use and they had haiku as the small model that sometimes works sometimes as a bad value. OpenAI has not had that clear of a separation where they have like their standard model which is the GPT line. They had their separate reasoning line with the O series and the O series had its own base tier and mini tier with their own weird release cadence. And then occasionally when they feel like it they'll throw us a mini or a nano model here and there and they don't even like fully commit to the nano models. Like sometimes they're just iffy on if they're going to put it out at all. Like the nano models are like why not make Google's life harder with flashlight. And the mini models are like oh we have one of these might as well put it out. Where at Anthropic, each tier gets some level of like focus and planning and like design around it. There isn't a lower and higher tier to the base tier. At OpenAI, they just have the model and then occasionally they'll put out like a mini or a nano. Whereas Anthropic has their base tier with sonnet. Sometimes they'll go bigger, sometimes they'll go smaller, but like they have that and they have the clear tiering for each of them. If there was something like that at OpenAI, I would argue 5.4 4 would have been like a set type and 55 should have been an opus type. It's priced accordingly. It behaves accordingly. It's smarter and bigger. All these things accordingly. They don't have that separation. They just have the numbers. And it wasn't a big enough jump to go six. And they also have like on the mini and nano models cuz even I forget about those and I'm a huge defender of those. I thought 54 mini was amazing. I love that model. I use it for a ton of stuff. I still do. It's insanely fast. It's insanely smart. Like it is I would argue that the mini models are generally on par with Sonnet models. maybe a little bit dumber, but like generally in that tier, but because they are named mini, we all just kind of forget about them and ignore them. Same thing with nano. Nano is a completely forgotten child that is also extraordinarily useful. But again, it's a nano model. Why would I ever use a nano model when I have the real model right here? I should just use that. I got another spicy take.
>> What if they had named Opus 45 Sonnet 5 instead? What if
>> instead of it being a price decrease on Opus, it would have been a price bump from Sonnet.
>> Oh. Oh, they could have done the Oh, yeah. Frankly, if they had just done that, I think it would have made a bit more sense because it did just feel like a slightly different four or five and it would have made a ton of sense to just be a sonnet model. I would argue it kind of is a sonnet model, but uh that's another conversation for another day.
>> Yep. I just wanted to make that point because that's effectively what OpenAI did here and they put themselves in a weird ass position with it.
>> They did. And the problem with a lot of these things and like the the reason why I've just been going so hard on the low reasoning thing and joking about that so much is because when you hear low reasoning or no reasoning, your instinctual feeling is, "Oh, that's going to be worse. That's going to be bad. Why would I not want the model to think more? They should think more. I want it to do my task really, really well." but they will end up not using it because they think it isn't good because it sounds worse. And that's what we're dealing with here. We are dealing with bad names.
>> As we all know, tweets with more words are better than tweets with less words. More words is better and smarter. Always.
>> Always. I'm pretty sure your average tweet length is like 10 times longer than mine.
>> It is. They do well.
>> You write essays.
>> It's just that's the way I talk and the way I think. Like,
>> but then somebody else like quotes you with 15 words and ratios the [ __ ] out of you.
>> Yeah. I don't care. Then my post did its job. Like if I can just do like the brain dump cuz that's the way I use Twitter. I just like dumping out things that I'm currently thinking about and then moving on. I don't care. And then if someone else wants to make it go bigger, [ __ ] yeah, I'm happy about that. I did my job. I will talk a bit more about things that I have come around to with the model that I think others will agree are very good. One of them is speed. On one hand, fast mode exists and is surprisingly good. Since I'm not using these things a lot currently cuz uh turns out running three companies, two major YouTube channels, and a bunch of other [ __ ] is a lot of work and time, especially when you're also traveling. So, I've not been getting that close to my usage limits. So, I've been using fast mode, which has actually been very nice. Separate from that though, despite the new model not being much faster. In fact, it's actually slower if you're just measuring tokens out, it feels absurdly faster because it uses half as many tokens, if not fewer, for a lot of the work that it does. And that just feels great. It reminds me of the like old fast model days and the things I love about Composer and whatnot because it just it gets the work done faster because it does it in less tokens. And it it changes my work loop quite a bit, more so than I expected. The reason I didn't talk about this before is that when we were doing the early testing, we're on a somewhat private cluster that does not have the same performance that you would expect for end users. So, I try to avoid talking about speed when I'm talking about the models I have early access to because that's one of the things that's almost always different for what you get from what we get during testing. But it feels just as fast. Like I did not get that much of a benefit on the private version versus what you guys get now. If you're using Grock 55 in fast or if you're using GBT55 in fast mode, I
>> Yes, please.
>> I I know you want your GRO code fast. When I talk about fast models by bringing defaults to Grock because it was the only fast usable model like ever. Honestly, now I think about that composer and grockf fast like being under the same house does actually make some sense.
>> Saying, man, Grock Code Fast V2, it's going to be
>> I love your brain dead models.
>> I do.
>> I get it now that you have the brain dead cats to go with it. But
>> don't don't don't talk about them like this. Like, does this look brain dead to you? Yes. Look at this beautiful little thing.
>> Sh. Anyways,
>> so the performance has been good, but the problem still stands with because it's doing less reasoning, it will make less reasonable decisions sometimes. It write The thing that's weird is on the no reasoning and low reasoning versions, the intelligence doesn't go down. It writes code the same quality. Like if you give a detailed like write this for me for low, mid, or high reasoning, the actual code output is going to be nearly the same. When you tell it to think through something and write a proposal that is detailed about it, the low reasoning won't go down as many traces and as many possibilities and will not be as considerate in the outputs. I actually think this example from Ryan Florence that I have in the notes here is really good. Ryan Florence is the creator of React Router and Remix. Very old school legendary webdev. We've had our beef in the past. We're pretty friendly now. He has a post that almost perfectly summarizes the experience I have had with this new model. I'll just read it out. GPT has a new phenomenon that's driving me nuts and I don't quite know how to describe it. Step one is that he'll ask it if he can do something. Step two is it will respond with no, you can't incredibly twisted restatement of what I asked, but also not at all what I asked. It then tells me how to do the thing I actually asked for wonderfully. and then finishes with an insulting. But you can't just stupid thing I never said. As an example, goes something like this. Can I form and coach a youth soccer team for my kid and play in PND level leagues or do I have to be part of a full club? It then responds for official competitive teams in Utah. You cannot just form a random team and enter a league. There's a lesserknown option, the UISA, which allows independent teams to enter leagues if they meet certain requirements. And then it lists some very simple requirements, but it's not show up with a group of kids on game day. M dash, it's more like running a small club team administratively. Ryan then caveats. He never said just show up with a group of kids on game day. It does the same with code, too. It's so weird. This is the closest I've seen anyone to describing one of the many strange patterns I regularly encounter with this model. I write prompts that are accurate enough to what I want that I noticed this change. When I'm being lazy and voice typing and some of the words are typoed or wrong and it does something like this, I'm more okay with it. When I'm sitting and typing and writing a very simple, very realistic proposal of like what would it look like to implement this feature in an electron shell that uses native drag and drop or something. It will say, "Well, you can't just use your native Mac OS drag and drop when you're inside of the Electron shell." Yeah, that's why I'm asking. And then it writes exactly what I actually want and then calls me stupid at the end again.
>> H, you're used to being called stupid, so I know this isn't that different for you. You're the one who calls me stupid.
>> I know. That's probably why you like the model. It reminds you of me.
>> It most certainly does not. This isn't exactly what I was thinking of with my issues with it. I have more had the issue that Julius has where the model will do the thing where if it has any kernel of context it will latch on to it because the way this model works and feels is very strange and very different from previous ones because uh historically open AAI models have had really really bad like pre-training and baseline knowledge like right out of the box GPT 5.4 for if you give it like no reasoning or low reasoning is going to be exceptionally stupid because all of the like juice they got out of those models from my understanding and my experience came from the reasoning. So you let it reason for a while and it will come up with a really good answer and do some really really good stuff. This model finally fixed that which is why low reasoning is now usable because the base core model is actually smart. This is the thing that claude models have had forever where like I remember talking about this way back last fall where I really liked using low to no reasoning on a lot of claude models because they could just do the thing without having to think about it and they would actually almost get dumber if they were able to think about it which is a thing that I've noticed happening with this model as well. That's why I harp on the low reasoning so much is because if you give it high reasoning, you are giving it more cycles to think something through. And when it thinks something through too much, it overthinks the thing. It ends up doing something exceptionally stupid. When it is on these lower reasoning levels, it just kind of goes straight to actually answering the question or doing the thing, but it sucks at recognizing which prompt to respond to in your history in that case. Like Julius gave another golden example of this that he just encountered. He was working on remote environment configuration for T3 code. And one of the things he wanted to set up is the ability to configure something like open code using an open router key because he didn't have an open code sub to use for it. So he set that up. He had like the model add like open code open router support. Got that working and in the same thread was like okay now do this for cloud code and an anthropic API key. and it insisted on using open router_appi_key as the key name for your enthropic API key and using it through enthropics APIs even though for the open code version he wasn't using an API key he was just using the built-in open code off layer like the where you have to paste it in
>> the anthropic one which should always be anthropic API key as the name
>> use the open router API key name instead just sucks that means you like the result of that is that you're reviewing code in a way that I normally trust something like Code Rabbit to review.
>> Yeah.
>> And if I'm just vibe coding out with 55, I feel like I have to use something like Code Rabbit in my review loop or I'm going to go Matt. And that's not just cuz they're like a sponsor, I do legitimately feel more than ever now that even though we have the smartest model we've ever used for coding, the need to have the code reviewed because it makes these types of stupid mistakes, not cuz they're baked into the model, but because they're in your context window where they should be. Because normally it doesn't [ __ ] matter. I actually had one of these happen today. I was testing out this new skill, which we'll talk about another time. It's called impeccable. Very, very cool. One of the coolest like uh designy type things. But the tilder on this is when you run one of the versions of the skill, it will add a thing to your actual like uh web project or whatever that will give you the ability to highlight and click on elements within the browser and like prompt to change them. And then that prompt gets sent back to the model running within your like PI agent or whatever agent you're using. And in the skill, it tells it to basically just run in a loop and like pull for changes so that whenever you send a change down via the SDK, it will pull that in and make those changes accordingly within your coding agent and go back. It's a really really cool system. Very very impressed. One of the things I had to do was make a change to a button and it [ __ ] something up within the HTML and ended up getting a syntax error and I told it, hey, you got this syntax error. Can you go fix it? And this was while it was like doing the polling thing. So, I interrupted the pulling and then it like did some checks but didn't actually fix the thing and went straight back to pulling because the previous message had told it to do the polling thing. Its context can get very polluted very easily because it is so insistent on doing exactly what you want. This has the thing that we talked about way back one with GPT5 where it would follow your instructions verbatim. It was very aligned in that way. This just does it to another degree. It is a whole
another level. I see it as it conflates different instructions from the user side. Like anything on the right side is honored anytime something new appears on the right side. If you have like the right and left for your chat or the right is user, left is agent, it doesn't distinguish between different messages.
Well, I agree. And I I'm not saying that this is a good thing. I'm not defending it. This does suck like doing things verbatim. If I tell it verbatim to configure the anthropic environment variables for cloud code, that verbatim does not include putting open router API key as the name for that.
Yeah, but it's conflating the old instruction as the new instruction. And the problem is the discernment between the two. It does not have good good discernment between like what takes priority. Instructed to ever use open router API key as the name for anything. I just told it that I was setting up open router off layers multiple steps before. This is the problem. It it makes up its own instruction by combining your previous instructions in ways that aren't actually what you said. Like if I told you to put the steak in the oven and put down the cat, I'm not telling you to put the cat in the oven. And that's what this model does.
That is a good way to put it. And that is 1,000% its biggest problem. I think my style of working on things results in me liking it more than a lot of other people do. It's been interesting actually seeing who's been like hard agreeing with me on my 5.5 takes. Like Mario, the guy who created Pi, had a tweet of GBT 5.5 plus minimal thinking slaps for bashbased computer use. Also, he said that he really likes the model on like no reasoning because I have also found that it is remarkably good on no reasoning. That's a conversation for the future. A lot of tasks that benefit a ton from that. Also, Sarah was like, "Yep, GPT 5.5, no thinking is a beast." It's don't skip the first three words of that. Sarah opened with Ben was right.
Yeah, clearly talking about Ben Stellar, famous vibe coder.
Yeah. Unfortunately, I I don't think I can take credit for this one. It could be any Ben. There's probably another Ben out there who's been talking about GPT5 with no 5.5.
This boy on my lap, Ben Jr.,
he has very low reasoning.
He has extremely low reasoning. It's pretty effective though. He's He's doing the cat thing quite well.
Fair. Go back to reading the whole tweet this time.
So, the whole tweet, Ben was right. GBT 5.5, no thinking is a beast. Holy moly, it's going faster than I've ever seen. Doing precise actions to get servers working, etc. 5.8K tokens to fix like 20 issues. And this has been my experience. It is absurdly efficient. It is really smart and just kind of does the things I want. And because of the way I work where I am constantly making new threads and jumping around and pruning my context pretty manually, I have not run into any major issues with it. It's just done the thing. Like a lot of the stuff I've been working on has been uh automation-y type work for building out very weird skills that call a bunch of bash commands to do a bunch of useful actions in the background. And it is so good at following instructions and that kind of thing where I can literally in one thread work back and forth with it on building out the instruction set for a task and then just actually call the pi extension which is really just a skill but it's not always injected in the context. You have to manually do it which I really like. I just do the slash command and then it will simulate as if it was doing it in a new thread and go through and follow the instructions step by step exactly how they're written. even if the way they're written is not exactly right, I can literally just debug it in the same thread and be like, "Nope, that behavior doesn't make sense. Change this, this, and this." It'll make those changes. I rerun it and it works perfectly. And then I open up a new thread, test it in a fresh environment, and it works exactly the way I would expect. This model is so good at tuning. I don't know the best way to describe it, but just like slowly descending its way down to the correct solution, just like working back and forth as we drill down to what I actually want. It is the model that feels like almost an extension of myself as I'm thinking. It's beautiful.
Do you trust it to explore?
Yes. I don't. And I think that's my issue with it is I like using models for exploratory work to like go find options and present multiple to me. And like I do a ton of that with it that I literally like the BTCA thing we were talking about way forever ago where I rewrote it to just be a skill. I do that all the time where I'll be like, "Okay, I have like I have this idea for doing it and I have this idea for doing it. Maybe there's a third option and a fourth option. Can you do a bunch of deep research into these, break down the pros and cons of each one? I let it go off. It'll do a bunch of like git repo cloning to search through those. It will do a bunch of web searches to find the things it actually means and the end result is always really really good.
You can't honestly say always for that. I I have like at least a 10% like absurd miss rate with those types of things with it where it just like hallucinates some [ __ ] like concluding that it can rewrite skills or turn off skills by adding dot disabled to the directory and [ __ ] Like when I ask it to do anything even vaguely exploratory, I feel like I have to audit the result more than I've ever had to. Like I can't just trust the outputs on this model.
Every single time I have it ground itself and I make sure that it's hitting a source of truth, I have never had it hallucinate against the source of truth. Like I've never like if I just explicitly give it, okay, this is the reference for how I want you to do it. Go explore this and get me the exact spec of how this needs to work. It'll give me that spec. I'll look at it like, yep, that is correct. That is how you do it. And then I use that as context to instruct it on actually doing the thing. Never had any issues with that. It is very precise in that regard.
That's not exploratory anymore. You're giving it the resources. But share.
No, I'm having it explore to get the resources of the thing. Like I'm not having it explore like what how do I like I'm telling it, okay, how do I make a PI extension or something like that?
Can we at least agree that once you're done with the exploration, it gives you three options that if you tell it use option two in that same thread, it's [ __ ] And the best option is just copy option two, make a new thread, paste, and hit enter.
Yeah. What I would always do there is I would just be like, "Okay, I want option two." Write out the detailed spec for like how to use option two, useful code snippets, blah blah blah. I would grab that. I would do slashcopy and I added like I have that command in my agent because I'm using these things competently. I do slash new. I have a new thread. I paste that in. Give it a brief instruction on what I actually want done and it is done. It's beautiful.
I miss just prompting through it. Man,
the the low reasoning is really useful because I'm I like being this in control of what it's actually doing. Maybe that's wrong. Maybe that's right. We've talked about this for months now, but if you want it to be more vague and you just have wanted it to just kind of go off and come up with its own answers and its own solutions, high reasoning is good for that. I have done some stuff with high reasoning like when I was doing some very complex network level like backend database and front end making room on your lap for the boy. Thank you kitty.
The babies want to be cuddling.
I love these cats. This is what I'm actually looking at right now. is a pile of cat. It's just a blob. That's what I'm dealing with here. Is you. All right. Sorry. What was I saying? High reasoning. Yeah, high reasoning is really good when you don't have a super defined end goal. It's more of a vague, hey, I want XYZ to work. How do we do that? Go do a bunch of deep research and figure out how all these pieces fit together. It's been remarkably good at even just scraping through like weird web resources and stuff like that to find exactly what it needs. And even just like it doesn't screw up interfaces the way old models did. And especially if you just give it good tools within the harness, which you should be at this point. Like if you have a good check command, you have a good lint command, you have a good format command, all these things, it will just naturally guard rail itself and like ping-pong back and forth between mistakes all the way down to the correct solution.
I got one more weird behavior I'm curious if you've seen. M usually, regardless of reasoning level, if I give it a task that requires like a lot of [ __ ] for it to like go and do, usually it'll do the whole thing and come back and tell me when it's done. If it doesn't, if it goes like one step in or like four steps in out of 10 or whatever and I tell it to continue, that thread is now haunted from that point forward and it will never finish a task to completion till I make a new thread. If it decides to tap out early once, that thread is now tainted and it will never complete a task to full completion in that thread again going forward. I don't know if I've had one go long enough for this, but I have had the beginning of this problem, I will put together a very detailed step-by-step plan, steps 1 through 10, and it will just it'll stop at very random points. Like, I've had it stop at like step four and a half and be like, "Okay, we're done here. Uh, would you like to continue or would you like to review?" And I'm like, "Brother, I told you to do the entire thing." And usually, at least in like the two instances I can think of where I have had this, every time I'm like, "No, what the [ __ ] are you saying? Go keep going." It will keep going. I've had much better luck with this by having planning happen on higher reasoning and then the actual execution of that plan happening on lower reasoning because again, it will just if it doesn't have enough time to think why it should stop, it will just gun through it. Guns blazing. It's great.
I'm going to say a thing I'll regret.
I miss Ralph Loops. M I think this model, at least from the many of the problems I've had with it, it would benefit from like a verifiable loop forcing it to keep going.
I actually could see this being the case because I, you know, we we talked about this before before Ralph loops really stopped making sense in the 54 era because autoco compaction was so good in GBT 5.4. And I will fully admit compaction is much worse in this model. Much much worse. I was effectively doing Ralph loops with 54 all the time where it would just compact and compact and compact and go and go and go. My counter to this is that for the vast majority of tasks, even some very big pretty complicated ones on very large code bases, I haven't even needed to compact because of the efficiency gains we've gotten. That's the big thing with this model is like it is more expensive on paper, but in reality it is less expensive. It just uses so many fewer tokens for everything.
usually. I have had enough strange issues that I have different takes. I'm going to give a funny example as our like last thing here.
Okay.
I have been trying to use the computer use stuff more cuz it is really [ __ ] cool. When it does fit for a task, it's awesome.
I had a fun task. My favorite music blogger, Bill Diffren, posted his roundup of his favorite songs of the month. I wanted to turn that into a playlist. I've done this in many different ways in the past that vary in quality cuz his blog is a web page on Blogspot that has 50 YouTube and SoundCloud embed links in it. It's a very obnoxious format. It crashes most browsers because they don't handle that many YouTube iframes well. So, I wanted to just tell a model to get the data out and make it a playlist for me. This actually used to be one of my favorite retrieval tests because I a while back, I think it was Gemini 25 Pro days. I took the HTML for the page and said, "Write me a script I can paste in the console that will grab all the YouTube IDs."
I remember this.
Google responded, "This is 25 Pro." Responded with the script and said, "The output should be." And gave me the exact output with all of the hundred links that existed in that blog post at the time. And I was like, "Wait, there's no way. He hallucinated those." I double checked. It got every single one right, which was very impressive in that era. It was really cool. Yeah.
What was horrifying to me here is that when I told it to do this, it only worked. So, it reasoned for 19 seconds, then said, "Yes, for a blog spot post like that, the easiest path is probably to open the blog post, run a small bookmarklet that extracts all the embedded YouTube video IDs, open a YouTube temporary multi- video playlist URL, and then save the temporary playlist into a real playlist." And then gave me the bookmarklet idea, which is the code for the bookmarklet. Told me how to use a bookmarklet. told me a bit of info about putting it in a text file, maybe whatever else. Annoying. I assumed because it was so confident on the YouTube thing that that would totally work. So, I just asked it to make it a JavaScript like function I can paste in the console instead of a bookmark lit because I don't want to [ __ ] have my bookmark bar open for some [ __ ] like this. Yeah. So, it gave me the JavaScript. I ran it and it did successfully rip all the YouTube IDs and then open a new tab with a temporary playlist which has no option to save it as a permanent playlist.
That's not what I'm mad about, though. What I'm mad about is when I followed up, I sent a screenshot of the page clearly showing that there was no way for me to do the playlist like saving said, "How can I save the playlist? This is what I see now." Yep, you're in the right place. And this is with no reasoning for some reason. Even though I have it on medium, it just didn't reason this time. That untitled list is a temporary YouTube qst style playlist. And YouTube does not always show a save playlist button there anymore. That's literally the whole point of this. I asked, "How can I make a playlist of this?" It is literally telling me that the suggestions it gave. Don't do that. A practical way is to one click the save button under the video title. There is no save button under the video title. There's never been a save button on the [ __ ] YouTube videos. It is deep in hallucinations. And this is responding to a screen. Oh, okay. To their credit, it does say save.
There's definitely a save button, at least on mobile.
It's not under the title. It's under the channel name, but it's there. I didn't think it said. I thought it was just the icon, but okay. I was wrong with that. It has the save button. It says to choose create a new playlist, which doesn't work cuz I want to do this for all of the videos. El says, "Small catch. The button usually saves the current video, not always the whole playlist." Have you ever seen the save button under a video save anything other than the current video?
All I've ever used it for. It's not It doesn't This sentence here, small catch, the button usually saves the current video, not always the whole temporary list. That is not just a hallucination, that's an abomination. And I'm framing this the way that open AI models love to frame [ __ ] things like this.
True.
It's not just X, it's Y. Why being insulting? To save all 30, the most reliable, low fuss method is to open each video from the temporary list and hit save into the same playlist. Annoying, but still faster than searching manually. A better consolebased route is possible. Instead of opening YouTube's temporary list, Zip it can output all 30 video URLs and you can use a tool or API or script to create the playlist and at every video automatically, but that requires logging into Google via the OOTH YouTube data API without API permissions. YouTube's web UI is intentionally stingy about bulk saving a generated list. So I asked it, can it use the computer use plugin to add all the videos for me? It then tried to, but it tried to use Chrome despite the fact that everything so far has been in Helium.
The screenshot was of Helium. It had helium in it. and tried to use Chrome. Anyways, I stopped it, said, "I'm using Helium. Chrome." It got back to work and then hit a compaction window. That's all it took. That was enough for us to compact. And you know what happens when this model compacts?
It's over over.
And do you want to know what it did? It went through all 30 videos and didn't save any of them to the playlist. It just clicked through each video in the playlist.
It just took a look at each one
and then said, "Done." The console script finished and YouTube reported all 30 videos as selected in the private playlist. Ew.
No, it didn't. It didn't put any of them in the private playlist. It just went through the temp playlist. Like, I I know I'm going really in depth on just like one run on here, but I'm trying to emphasize a point. Every single response the model gave me during this set of events was wrong, hallucinated, stupid, bad decision-making, and just led me down the wrong path. And ready for the [ __ ] hook, line, and sinker on this one. I did this before with Opus 4.5. Hell, I did this before with GPT 4.1 a year ago.
Yeah. Yeah. Th This is rough.
This has been my experience with the model. Half the time I feel like I'm dealing with [ __ ] like this. I'm sure it writes slightly better code often enough that it's a value ad for people who are just using these for code, but I'm using these for using my [ __ ] computer. This is the type of task I love using models and harnesses for and I was really excited for computer use to at the very least go eject the right JavaScript because you know how I did this with 4.1 I handed it an example request when I added a video I just went and copied it from my like request response in Chrome pasted that said you already have the video ids write me a script I can paste in my console that will add all the videos I thought that a year and a half later and that 3x the price per token later and we're nearing AGI later with computer use and everything that it might have been able to figure those steps out itself and instead I just sat there pissed off, opened up my phone and added them all manually, which took less time than all of this.
Yeah, I rest my case.
That's a very interesting one.
Like I'll let you read this after if you want. This is like the worst AI experience I have had outside of Gemini models.
I want to try doing it myself cuz I know I would approach this differently, but also this should work. The fact that that prompt didn't work is bad. My point is I wouldn't have written that prompt.
I copy pasted the exact prompt for you to start with and see what it does. I try to put as little thought into these like I I could steer the model in a way that I know will work, but I like learning from them. I like to see how they approach the problem and what they do different. And usually the smarter models are good at things like this. The the my instincts and I think this is this is more of a me thing. The way I would have done it is I would have started with okay I have this site grab all of the YouTube video URLs out of the site and then once I had those then I would go down the playlist rabbit hole in separate steps but I that again that should work. The fact that that doesn't work is concerning. I am going to try this when we're done.
Thank you. I I will end on this note. Something you were very right about is that markdown can be and should be treated as a way to write programs like incomp completion. Instead of like writing a little bit of markdown, generating some code, saving it, writing more markdown, generating some code, saving it, and then trying to construct that into an app. I should be able to just write the prompt and let the model do the [ __ ] thing. I shouldn't have to be the middleman chaining the steps together the same way you shouldn't have to be the middleman chaining your [ __ ] code together when you could just give the model the context, give it the prompt, give it the thing, do the steps. This is me doing the same equivalent there. But because this model sucks at multi-turn, it sucks when you have more than one message in the history. I wouldn't say that, but I do. What this feels like to me is this is it being bad at coming up with ideas for doing stuff like this, which is the most concerning piece because like generally speaking, when I'm working on these things, I pretty much have the implementation in my head and the steps in my head. And I do like I did have a thing today where I wanted to put together I'm working on a like I've been going way harder on like multicomput stuff and like the Mac mini [ __ ] I'm going down that rabbit hole right now. And one of the things I was doing is I wanted to basically create a job queue so I could literally run a slash command within my Pi instance on my main Mac. When that runs, the agent will run a gigantic skill that will package up the current state of the repo I'm in. Send it over rync to the actual Mac mini over tail scale. Start up a job to actually do that thing. And then as soon as that's done, it sits in a registry where I can at any time just pull that back in by grabbing the diff and having the model do it there. I wanted my own little background agent solution thing here. And I had a bunch of steps and a bunch of ideas of how I wanted this flow to go that I wrote out myself. It was like five, six paragraphs worth of stuff. And I told it go question by question so it's easier to answer. And we did like 50 questions of back and forth on every minute little detail of the shape of the actual JSON that is the shape of a job to what is sent where the temp directory is going. And it had a lot of really good ideas that were very useful. And as soon as we got done with that entire plan, I was like, "Okay, persist that into the markdown file and it did it perfectly. I had no issues with it. I did a test run. It looked great."
And then you brought that thread back behind the barn, shot it in the face, and made a new one to actually implement it.
No, no. I implemented in the same thread because it already had all the the Q&A context did tell it how to actually implement the final thing and it was fine.
I will drop my final argument thing. By the way,
I was on Medium for this, which might have been the mistake, but regardless,
that could actually be it.
You know, my favorite argument for Typescript is the brain power it frees up. It's not that TypeScript is an objectively better language that prevents you from making mistakes all these. Like a good enough engineer will avoid all the mistakes that TypeScript will keep you from making. The benefit is that your brain no longer is being tasked with the horrible reality of like every single variable, every single line, you have to think about it before doing the next one. The way we prompt is different in the same way. Whenever a new model comes out, one of my specific goals is to free my brain of thinking of all of those types of things and just vibing it in the right direction. And when it can't do that, I sigh and go back to holding its hand. I feel like it's been a regression in that sense. I have to think more with this model than with previous models. I have to hold its hand more. It's a much smarter model that once I carry it to the right place will do the right thing. But I can't let go more here. I am holding on more than I had to prior. And that is the like a wizard in JavaScript would not like TypeScript the same way that like a wizard in agent coding would not like a more autonomous model. But I feel like I moved from TypeScript to JavaScript here. like the newer newer models, especially this one, has freed up a ton of my brain power where I no longer have to stress about little things like, okay, is it going to use the right syntax for this? Is it actually going to write competent code? Is it actually going to do the thing I want it to do? And the things that I am now thinking about are the like, okay, how is this all going to flow together? How is what pieces are we going to use for this? What does this app actually do? All I'm thinking about is that and I I I don't think we're at AGI. I think that we have very very useful tools, but this is not AGI. They can't make those decisions for us. I don't trust it to make a full custom durable streaming implementation from just like make this stream durable, godspeed. I don't trust it to do that. And I I I've never prompted that way. I've never really liked prompting that way. To be fair, this is more me being insecure and not liking letting go. I've never liked doing this. We've argued about this forever for the last 3 or 4 months. I like holding these things. I like directing these things. I like having my vision of what I want to happen. I like having the code be the code that I would have written and have it be really good and feel really good to use and all that stuff. This model is the best for that. This is the single greatest model for if you are paying attention and you are have a specific end goal and you are using this to craft your vision. It is the best model ever for bringing that into reality. It's not even close.
At this rate, I'm going to make you hire an engineer so you can actually be forced to learn how to let go of it.
I'm already basically being forced to. I have too many jobs right now. This model's encouraging bad habits again for you. And while it can be justified because the quality of the output is better, I don't think it's the long-term right move. And I'm going to keep awaiting the more heavily RL and lobomized version that does exactly what I want. 56 is going to be my favorite model ever. I already know
that's probably true. I actually, that is a good prediction. I am going to like 56 less than You're going to like it more than
you're going to like it 20% less than you like this. I'm going to like it 200% more than I like this.
I would be willing to take that wager. That sounds correct. But we'll see. We will see.
And on that note, we'll take one last look at these very sleepy kitties and recognize that we should be sleeping too and wrap this up. Thank you as always for listening. Make sure you give us a like, a follow, and all of those other things on all the different platforms. We're somehow climbing the charts on Apple Podcast, which is unbelievable. Did not think we would ever be in the top 100, much less like the top 50 as quickly as we've gotten. So, thank you all for the support. Appreciate y'all immensely. Until next time.