📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

AI chat apps are driving me insane

Theo - t3․gg1:15:36

Transcription

Companies like OpenAI and Anthropic are really good at making groundbreaking models. It's a shame they're not as good at UI. I would know; I built my own alternative, and it was not easy. But man, whenever I go back to the other apps, I'm just blown away at how bad they are.

I was debating not doing this video because I'm basically handing a free win over to both the Anthropic guys and the OpenAI guys. This is an instruction set on what they need to do to make their apps not suck. But I feel like I'm going insane; I need to talk with you guys about the things I see when I use these apps because it's—it's not good. Let's descend into madness because, uh, we're talking about having three help buttons in a menu with five buttons. What? How? And this is just the surface; the more you go down this rabbit hole, the worse it gets.

I'm not going to sit here and pretend T3 Chat is perfect; in fact, we're going to roast it as well. But we have a lot of other things to roast first. So if you're interested in breakdowns of where UI does, and more importantly, doesn't perform as expected, this is going to be a really fun one. Stick around! But first, we have bills to pay, and the CLA bills are getting particularly nuts. So a quick word from today's sponsor.

Today's sponsor is one of those companies that's really, really easy to recommend because they just get it: it's Infinite Red. These guys are the best React Native devs you could possibly have working with you. If you want to make your mobile app better, or you want to make your mobile app in the first place, or you just want to get your mobile team to better understand how to make great software with React Native, you should give these guys a call. They know their stuff; they've helped so many companies get their mobile stack together, making it easier for web devs to make changes on mobile, or spinning up a mobile team that can ship to iOS and Android and other places as well, way faster than ever before.

In a world where everything's moving towards React Native—be it the Windows start menu, or the Xbox interface, or the PlayStation's entire UI, or even like desktop app-class stuff—everything's going in this direction. The entire Meta ecosystem, like the VR headsets, the whole OS is React Native; it's kind of crazy. These guys get it; they help a ton of those companies and more. Everyone from Zoom to Amazon has worked with Infinite Red in order to make their apps as good as possible. These guys make those React Native apps where you don't actually have any way of knowing it's React Native unless you look at the bundle. But at the same time, your company's web devs can contribute and add their features and make changes without having to go spin up the iOS and Android teams and fight with them back and forth.

Man, my only regret is that we didn't get Infinite Red to come help us over at Twitch because it would have made that app a hell of a lot less bad. If you manage to get a chance to talk to Jamon, tell him I sent you because I love that dude so much. He was one of my first people to really support me throughout this whole creator journey, and I was so thankful when he was one of the first to hit me up about the new sponsor program. These guys aren't just another company paying me; they're good friends of mine. I care dearly about them, and everyone who has chatted with them after my recommendation has had nothing but great stuff to say. The Infinite Red team is awesome, and if you do chat with them, you'll be in great hands. Thank you to Infinite Red for sponsoring today's video. Check them out today at soy.l/InfiniteRed.

I want to give these chat apps as fair a shot as possible, and doing that in a Firefox-based browser is not going to be it. I can already see banding in the gradient in the back; obviously YouTube compression is going to do that too. But out of my goal of being as fair as possible, I'm going to hop over to Vivaldi, which is Chrome-based, to maximize the likelihood that this goes okay.

First and foremost, I know this is like a silly thing to complain about: the fact that I can't hide my email on the sidebar is just the most annoying thing in the world. So, uh, my editor has to do a lot of extra work to hide that. We'll do our best anyways. This is the Claude UI. I'm putting it in a Chrome-based browser because I want to be fair; Vivaldi is Chrome-based; it should be peachy. But man, it's gotten cluttered; there's so much going on here, most of which I don't care about. Like, I'm never going to use Google Drive attachments; I'm not going to attach my GitHub. It's in beta; cool, I don't care. It's not the worst; it is still quite pretty. Oh, that's a fun one: the scroll area is a different background color, so when you scroll to the bottom of the container, it has that little gray there. It's annoying, but that's not even where things—oh, actually, that's a fun one: the scroll area. Notice something: once the chat hits the top, it's a sticky container, so everything else scrolls underneath it until you page scroll. That's God—oh God, I didn't even know that one.

Let's do a quick test: solve Advent of Code 2021 Day 3 in TypeScript. First favorite quirk: see "New Chat." See how it's still "New Chat"? I just hit stop; it took forever. The title will not come through until the stream has stopped. It's such a simple thing, but if you have multiple chats going, not generating a title for what it is until after you're done is the most annoying thing in the world. We generate as soon as we can, and on T3 Chat you'll get a title almost immediately. So if I say, "solve Advent of Code 2021 Day 3 in TypeScript," you got a title immediately, even when the rest of the response is still coming in. It's a small thing, but it's one of those things that is so annoying when you don't have it, and you go back to other apps that don't—or you do have it and you get the idea. Anyways, oh, that was another fun one: the artifacts don't pop in, so when you don't have enough screen space, it does this and just makes everything unusable. You can click that little back button, and it gets here: the world's least useful menu. My favorite game to play with Claude: guess what the button does. What do you think this menu button does? "Chat controls." Guess. Chat. I want to hear what you guys think that button does. It's chat settings, yeah, but what do you think it is? Chat settings. What do you think is in chat settings? The artifacts that have been generated and this "Learn More" button. None of these are settings. Previously, they let you select the font in here too, and they just changed that because I complained about it on Twitter because it was silly. The only thing you could control in "Chat controls" previously was your font; now there are no controls at all; there's nothing in here; I can't do anything in this; I don't know why this button exists. Also, why is "Share" a text button when these are both icons? Is it because these aren't behaviors people actually use and they shouldn't be there in the first place? Crazy.

Ah, let's have more fun; let's click this button. What happens when I click the down arrow on the chat? I could let chat guess again, but I'll show you. In the middle of it, despite the arrow being on the side, in the middle we get "Rename" and "Delete." "Rename" opens a modal. Here's where things get really fun. What happens if I save this with no text? The answer is you can't guess because it is non-deterministic, depending on the conditions of the page. Saving an empty chat name does different things, and those things can get really, really funny, so I will show you. We're gonna have to open the sidebar, which just put my email in it again—super fun. I'm just going to rename it with a blank right here, as is, save, change it back to "New Chat" and "Untitled" here. Annoying; those are different; that's the empty state; that's fine. Let's resubmit, and while it is submitting, let's rename to blank again. Okay, didn't do something weird that time; it would in the past. See if it updates correctly after—I need to pick a smaller prompt to test this with. Look at that; it didn't update the title, and if I refresh, then it did. It gets better; we're just getting started, boys.

What happens when I click "Delete"? When I click "Delete," I want to make sure you see what happens, so I'm going to do a weird thing: I'm going to inspect element and change the text up here to "Start another chat" because we know that is different. Watch what happens when I delete it: it refreshes the whole browser. That's how they handle delete, and it's not like they're handling an edge case because deleting the thread you're in—there's no other way to delete threads; you can't delete them from here; there's nothing you can do. The only way to delete a thread is to open it, go to the title, click the little arrow, click "Delete," and watch your browser refresh. What? This is a multi-billion dollar company; this is their primary product; how is it this bad? And the funniest thing about this is I'm convinced they're not using Claude at Anthropic because I don't believe Claude would write code this bad. It's just hard to fathom. There was another edge case, and I'm struggling to repro it. Oh, don't—that was my fault. Oh, actually, another fun one: when you hit "Rename," it doesn't target the input box, so I just hit "Delete" to start deleting text, and it navigates out of the thread because, of course, it does. So I'll hit "Rename," go here again, save it as empty, send another message—or how about "solve Day 1 in Python"? Give it an easy thing. Why is—oh, just thinking on—because on—in this chat already—no title change. I'm going to navigate—still no title change. I'm going to refresh—still no title change. Opening it, then the title changes. What? There are so many non-deterministic behaviors; it's insane. The only thing that you can guarantee with Claude is that the UI will not do what it's supposed to.

The thing I was trying to reproduce, and I don't know if they changed this since or if it's just that non-deterministic, is sometimes when you delete the title or it fails to generate a title, it makes a title that is just the input truncated with a chat emoji in the front. I didn't name this this way; why would I have ever named a chat this way? It didn't name the chat like this via an AI summary, and it's a small annoyance compared to things I just showed. Why do they default to putting a chat bubble in their copy next to a chat bubble on the—like what? And that's—I don't get it. But we're—we're just getting started, believe it or not. They did change the three help buttons, thankfully, but that doesn't mean they've changed the other stuff in this menu over here. "Learn More"—it's "Learn More," whatever. But "Settings"—oh boy. Also, I don't know if you noticed—I'll show this again because I think it's important to see—we're going to inspect, change this once again to "refresh test." Now we're going to go to "Settings." Why does that do a full-page refresh? Why does the whole browser refresh its content instead of navigating as a single-page app? "Profile"—what it is—"Appearance." They moved font here, which is nice because previously, as I mentioned in my thread, the only place to change the font in the chat styles was in the "Chat controls" I showed before; there was no other place to change this, and it was in the weirdest place. Now "Chat controls" just literally does nothing, which is way funnier in some ways. But man, they offer dyslexic-friendly, which I do like; I think it's awesome they offer dyslexic-friendly font. I've considered doing that ourselves for our stuff. They also had—do they still have those settings down here in "Profile"? Oh yeah, that's a fun one: this does not look like there's anything else; this looks like a completed page, and they're not showing a scroll bar, and they're not indicating at the bottom that there is more content. I have to know to scroll—hilarious.

For those who don't know, um, React Scan is a very, very useful extension; it's not officially on the Chrome Store yet, but it's a way to see how often renders occur on a site. There, now we have React Scan here, so watch what happens. Can I not even send the message? That was a funny state where it didn't have the send button for some reason. At least this does—that it puckers; you touch it a whole bunch; it gets mad at you; it's cute. Yeah, I can't get React Scan to scan the site. Okay, it appears that the extension is broken right now; that's annoying. Um, Aiden, Aiden, this is your job; fix it. Should I call Aiden? Yeah, you busy? I'm currently live; I'm trying to use React Scan, and the extension doesn't appear to be working at all. Yeah, if you wouldn't mind, appreciate it; feel free to hop in chat; we'll see you there. Thank you. Aiden's the best; for those who don't know, he's building Million.js and Million Lint and a bunch of things to make it easier for ACV to not make slow-ass sites. It's funny because, like, I'm an investor; I love Aiden, but he's also making it easier for companies like Anthropic to catch up with the performance that we have on T3 Chat. At the very least, you can see the frame drops right now; it's not dropping now—interesting; it was before. Yeah, we'll let Aiden go fix that. While we wait, I'm going to show off a few more fun quirks in the UI. I—I forgot to put in my Claude video, and I wish I did. You see there, I have it: "How many Rs are in the word 'strawberry'?" Close the sidebar. I don't know if they—was in the system prompter—if they hardcoded it or what, but it builds you a little React app that shows the Rs in "strawberry"—they solved the strawberry counter. They still know how to build a UI. Cool. Anyways, any more fun here? That's all—not too bad.

I want to complain about the sidebar though, because anyone who's using a modern browser like Arc or Brave or even Edge—most of y'all have the sidebar on the left because you probably should—mine's on the right, so this doesn't affect me. But for those of y'all with the sidebar on the left, it's very easy to try and trigger this sidebar and accidentally trigger yours or vice versa—to try and trigger the browser one and instead have this blocking you and getting in the way of your UI. It's so annoying to have like the stacked triggerable sidebars; it drives me mad. Another fun thing you'll notice: I only have like seven or eight chats here; I have a lot more than that; you can't scroll it; it doesn't paginate. The only way to see the rest: click "Chats," where it fetches all of them. It doesn't fetch some of them and then paginate later; it is your entire chat history. Oh, can you finally select and delete multiple at a time? Can I Shift-click to do—oh, look at that, real progress. You can't undo that, but there's finally another way to delete chats—good for them. Oh, that's funny: when you unselect everything, it kicks you out of select mode, but it still has the option to select things. Is that always there? That's janky; I hate that. I hate that whenever you hover one, it has this little not-indicated thing in the corner that you can select, but it doesn't give you any context on what it is or does. Let select mode be a thing; don't—yeah. Oh, what the—did you do, Aiden? How did you fix that live? What the—oh, okay, we have React Scan now, and we see it's certainly reacting and probably scanning too. God, it's a lot of rendering going on here. Make a new chat: quick, "solve Advent of Code 2021 Day 2 in Python." We're talking about thousands of rerenders here, and yes, every mouse move, the sidebar is rendering. Fun fact: the only way the sidebar stops rendering is if you stop generating things and you stop moving your mouse. Seems like we know what Anthropic is spending their money on: M4 Maxes for all their employees, so nobody ever notices the performance issues.

Just for a quick comparison, so I'm going to make a new thing: same message, get this out of the way, send. That's it. The sidebar rerenders three times, and as Aiden, the literal React performance wizard, says, that's really good. And I'm going to go further soon where each message renders what it changes, not the whole message thread. Right now, it's actually not too bad. The only catch is if you have a lot of code in one thread; it is triggering a render with the highlighter, which should auto-diff and realize, "Oh, I don't have to do anything"—it doesn't. And I've been putting a lot of work into fixing that. I'll show you guys this code, why not? I have a new beta highlighter. I just check if you have this flag set in local storage: `betaHighlighter: true`, and you'll know it's different because, at least right now, it has different colors; it's blue. So this new highlighter is significantly more performant; there's still a few hacks I want to do to make it even faster. But now every navigation is going to be instantaneous; you'll get like one or two renders as the sidebar realizes a different thread is open. Everything's instantaneous; you'll never see a frame drop, especially if you have the new highlighter on. We put a lot of work into making the performance not just good, but exceptional, because these apps shouldn't be that complex. The complexity is in weird Markdown rendering. But yeah, also, sidebar is virtualized, you know, because if you scroll really fast, you'll see the text disappears for a sec, but that's because we don't render all of your messages; we only render the ones that are visible. It—it's not easy, but it's not that hard either. Also, delete right there, and if you Shift-click delete, you don't have to go to another menu or a confirmation. If you delete the thread you're in, it just kicks you back to the "New Chat" screen; it doesn't refresh the whole browser. And if you want to change a title, double-click: new title. How hard is that? I don't get it.

We'll come back to T3 because there are things I don't like, and I do want to roast our own stuff. But I'm just in awe of the hell that is this. I still get like every mouse move when the sidebar is collapsed triggers that. Does it still trigger when the sidebar is locked in? No, but switching your mouse from one side to the other does. Just moving your mouse up and down to the sidebar triggers hundreds of rerenders. Let's see the CPU utilization; I'm so curious. I'm hitting 17% to like 20% CPU utilization moving my mouse in a circle on a four-to-$5,000 MacBook. It's fine; just—just move your mouse less. And for a fair comparison: 1% to 2% moving my mouse. Oh no, it bumped a little bit—chills at a significantly lower level. I do wonder if React Scan's affecting that at all. Yeah, React Scan has a penalty; like, I'm moving my mouse around, and I'm struggling to break 10%. There, if I go back over here and turn it off, still hitting the 20s—that's insane; that's actually insane. This is the thing that drove me to make T3 Chat. Watch what happens when I click a thread for the first time: it has to load; you can't just open a thread; you have to load the thread, and it's not even fetching it when you hover—just hovering over it. Click it, and it takes like a second to load, and as Chad is noticing, they're leaking DOM nodes. Just navigating between these, and I'm already over 100,000 DOM nodes; my RAM is spiking too. If I just move back and forth between threads long enough, I'm going to run out of RAM. We're already at like 200,000 nodes, 250 maybe. This is why they refresh all the time; maybe this is what's triggering the constant refreshing is the absurdity of the navigation causing endless leaks of new DOM nodes. Stop moving it; it constantly grows; yeah, it does; yeah, it's just going up. That's insane. Okay, it's not—not as bad when React Scan's off, but it is still just climbing by itself. And to those asking, "Can you crash the browser doing this?" Not only can you, I've probably crashed a billion browser sessions for users by screwing this up, both in my time at Twitch and with T3 Chat. You got to put time into the profiler and spend time in these tabs figuring out what's going wrong and hitting up people like Aiden to help you figure it out. We had a leak like this in T3 Chat where I was using the wrong version of a Dexie dependency that was really outdated and not maintained that caused for the event listeners to slowly grow constantly. And again, to compare: you got 10K nodes, and if I switch to different things—went up a bit, then it got stable—went up like 10—that's a big thread, so went up a bit for that, then it's just—once you've been to it, it doesn't create new nodes. Yeah, whole different ball game—like what? And also, like, the memory usage—are you kidding? 35 MB versus—and we're the ones holding it all locally, by the way. Your entire chat history is on your device; theirs isn't, and we are using less RAM.

Okay, we've covered performance; we've covered weird quirks; we've covered a bunch of the gripes I have with things. There's a bunch of weird stuff with file attachments and other stuff there; none of it's that big a deal. Oh, that's actually a fun one: there's a "1" here because I attached content, and that's all this menu for "controls" will ever do is show different content you've put in and the artifacts that are generated. There's no way to know what that number is unless you happen to know that's what you uploaded, and you've spent a lot of time—like, it's—I'm not going to say Claude's UI doesn't look good; it does. It's a huge inspiration for our UI; we actually use it as a design reference for a lot of the work that we do at T3 Chat because it's beautiful. Their designers know what they're doing; they have a vision for all of this. They don't have a leader owning the quality of the product, and the result is things just get stuffed wherever they fit, and they never fit the place they're stuffed; it's kind of chaos. And that's how—oh, do you guys see what I'm seeing? [Laughter] There—oh my God—I—I'll give you a [Laughter] hint. Oh my God. Z-indexing is hard, but it's not that hard. The X—for those who didn't notice—the X in this little window here appears over the chat box. I had to click it three times. Oh my God. Okay, we get it. Anthropic, a small indie company; they only have a few thousand engineers working on this. It's not fair to compare them to a really well-funded startup of elite engineers like T3 Chat, built by two people. Let's compare it to another indie startup: chat.com, otherwise known as ChatGPT, otherwise known as chat.open.com. They've changed URLs too many times. Let's explore.

First and foremost, I will challenge you guys to pay close attention to the sidebar on the left. Watch what happens when I refresh. Did you see that? The number of features changes, so the whole layout shifts a ton just in the middle of it. Even better is up until recently—I don't know when they made this change—your recent threads would also have a loading state and pop in either before or after, so you'd get like three things in the sidebar: you get "Projects," then "Previous Chat Messages" would come in, or Sora and Dolly would come in. Those buttons don't appear until it's loaded a certain bit of info. Now it's just "Operator," I guess the rest are hardcoded now. Wait, no, I think that—yeah, "S" is in either—if you Command-Shift-R, but if I just Command-R—wait, no, Sora is not in it every time. Cool, that's at least somewhat deterministic; it's bad, but it's deterministic. I'm going to do something I'll probably regret; I'm going to open the network tab because there are so many fun quirks with the networking in—in everything that ChatGPT does. Refresh once more. You'll see we just transferred six megabytes of content. But where it gets even more fun is what happens when I scroll. See that "conversation data" here that just came in? Took 700-ish, 600 milliseconds. Now I took 700; they're loading in more—that's a change that they just made. Yeah, the limit—the limit is 28; the limit used to be 10, which meant that it would paginate when you got to the bottom and not fetch enough data and then have to fetch again, and sometimes those fetches took absurd amounts of time. I have a video somewhere; it might even be in one of my old T3 Chat videos where these can take eight to 20 seconds each. And yeah, one of them just failed: "Sorry, too many concurrent requests." I was scrolling. What do you mean, "too many concurrent requests"? What? You have Grok right your back end? Oh, there was a nice slow one. How long did that one take? 1.24 seconds. Best part is it requests it twice; it doesn't request it once. Every time you hit one of those barriers, it requests the data two times. And that's a good point about offsets; that means if I was to go to another browser—another chat—just made a new chat, go here—if I scroll more—will finger—yep, it duplicates. It's using an offset for the pagination, not an identifier, so if a new chat comes in on a different device and you scroll, it will duplicate whatever was at the bottom. If twos come in, it will do it twice. So if I go switch over to a real high-quality model like 40-mini—make one—it's just not replying; that's fun. Three, four—oh, oh—"New Chat," "New Message"—really helpful. What are those named? "Clarification," "New Chat," "New Chat," "Clarification"—great job. Let's go back and scroll more. That load time—look at that; we successfully duplicated four threads just by scrolling. Are you kidding? And one of those took forever—yeah, 1.2 seconds, 1.6 seconds—insanity. The more you screw with the network tab open, the more insane you feel yourself going—kind of absurd. Why the—would you use limit plus—yeah, this is exactly how I feel; it's genuinely absurd how it's—uh, built.

The reason I'm frustrated here isn't because I'm like just trying to dunk on ChatGPT or Claude or anything; I'm frustrated because when I was building T3 Chat, I had a lot of these problems too, and they kept me up at night. I was obsessively solving these things because they were embarrassing, and I didn't like the fact that I was charging users money for things that had this pile of absurd, unjustifiable quirks. It's just absurd to me that a billion-dollar company with thousands of people isn't fixing these things. Yeah, also, it didn't fetch the new messages because it doesn't ever refetch—ever—just paginates. So the only way you get new messages is to refresh your—and then you see them all pop in because it is—it must be caching things locally now by default. We got to investigate; what do we got in here? They're not using Next.js, are they? There's no way. Okay, it's not there. "Session Storage"—got anything fun? Nope. Apparently, a lot of this depends on how old your cookies are, so I'm going to clear it—see if I get logged out by doing that. Okay, I didn't get logged out, so those weren't cookies. This all just appeared by loading the page, so they are caching things here—now some amount—just not necessarily really useful stuff like flare resources. This "resources" thing should prevent the weird popping I was seeing on the sidebar when we load the page, but it's not, cuz they only put Dolly in it; they didn't put the others in it. So they have a cache that should be solving this, and they're just not using it. What's in the shared storage here? Nothing. "Cache Storage"—just React Scan extensions. Where are they getting the old sidebar data from? If I go back here—also, how did that happen? How did this never get the new titles for these? Why are these chats in a bad state—just randomly popped in? That's fun; that's really cool. Oh, did you see that? One chat will like play its name out when you click it, and the rest don't. That one does, and then the content doesn't load. What's going on? I feel like I'm in the Twilight Zone of bad UI; I'm going insane. This is a different—this isn't because of C deletion; I'm in a different browser; I didn't delete cookies either; I only deleted local storage, and I deleted the local storage in Vivaldi. I then went back to—to Zen—Firefox because I'm trying to like show broken state when you hop between things. All of these messages were—it just changed back to "New Chat"; I didn't do anything. This haunted—OpenAI is haunted. Like, if this was a small Twitter startup, I would get it. They're competing with a small Twitter startup, and they're losing, by the way. OpenAI Sam, if you're watching this, I'm a YC alum; my phone number is in Bookface; we can fix this. There's a path to making the ChatGPT site the best AI chat site ever made; it might not be cheap, but my number is in Bookface, and if you want to explore that path, you know how to do it. Anyways, and when you click it, it reappears. That one scrolls; all of these have different behaviors. I sent all four of these at roughly the same time; I don't know what's going on anymore. Let's—let's show off more things that I think are stupid. How—my Corgi population—what—one of my favorites actually showing right now. Looks like they made a subtle change that makes it less bad. This one—all bit more of a personal gripe, but I don't care, cuz it's bad. This topnav cutoff is the worst; they're trying to do the ghost thing where it like doesn't have an actual topnav. The problem is that means this text gets cut off by an invisible bar. Oh, they blanked out again—yep, and a different one blanked out too. How are names not determined? How I broke this code one time in T3 Chat: I put in a `wait` in the wrong spot, and I felt so bad; I gave three free months to everyone who reported it, and that was because the title took too long to generate; it came in when the thread or the message was done. That was the only mistake. These guys can't keep a chat name. What happens if I rename it, by the way? If I rename this to [Music] "test"—I can't make this up—I—oh God, Jesus Christ, it just keeps getting better. We haven't even really gotten to the UI; we're just focused on the title still. Oh God, wait—wait till we pull out React Scan on this one. But I—I do want to complain about this navbar because I hate the ghost navbar that isn't a navbar. Just put a border on it, and if you don't have the balls to put a border on it, don't make it a navbar; admit your shortcomings. This sucks. The fact that your text or your content gets cut off by this invisible bar that is the same background color—it's just bad; it's just bad design.

I'm sure it looks great in Mockups, cuz this is how companies design. For those of y'all who haven't worked at a big company, this is the screenshot that gets made in Figma; it looks like this, and this here—everything on the screen at this moment in time looks totally fine and good. Hell, even if I like scrolled up here and had a different thread selected, all fine and good. The problem is what happens if I make it slightly shorter or wider? Now this overlaps with that; the design doesn't account for that cuz the design was this, where there's like a line here and here, and the designer probably thought, "Oh yeah, it'll just scroll through it." But when the engineers got it and had to implement it, they noticed that this would happen, and those would cover these, so they just made it a solid colored bar; it's not even a blur or something; it's just a solid colored bar. What the hell? And we are still just getting started. God, this—why is this thread so broken? Waiting—will it come in? What happens if I send another message? I pressed Enter; it's not sending; still not sending. Are we getting console errors? Nope. God, it's so slow too. I know it's like the stereotype, but like, just for a comparison: "solve Advent of Code 2021 Day 1 in Python." Okay, it's going faster than it used to; that's a good change. All these weird errors trying to load the wrong language model—that's funny; it didn't stream in "Python" as a single token, so it tried to fetch "language p," then "language pyth," and then it had the whole language tag "Python," at which point it could finally render "Python"—genuinely hilarious. Why is this JavaScript formatted? Do they just look for like specific keywords in JS and call it JS? If that happens, that's hilarious—if so. And once again, T3 Chat comparison: I'm going to turn off React Scan, switch over to the same model; I want to be fair; if they're using 4-mini, I'll use 4-mini. That's a little bit faster; that's about three times faster; I've measured it. Do we

Beautiful A+! You is hard. This one's just dumb. They have the floating chat bubble here; where is it going? The chat can't go below here, even though this area below it just makes this whole container feel really weird. Like, if you look at it in browser tools, the container here is strange. Close sidebar so we can see it better. We have the container. I God, I hate Firefox Dev tools too much. I tried; I genuinely did. I don't know what they're doing with this hierarchy at all. Is every message in an article tag? What are these even highlighting? Yeah, each message is a full-width article tag. What? That's hilarious because it doesn't have access to this whole container, so every single message is calculating its width itself with a WF. But they're using an article; are they using—they're not using Tailwind UI typography here, so this is all custom. Uh, where's the width style being applied for the content? Oh, that's just the screen reader. So they have the screen reader say "ChatGPT said," and then they have their section here: `m-auto py-8`. Does it change when the size changes? No, it doesn't. Okay, that's not as bad as it could have been. And here, like seven layers in, that's where they're applying the max width, and they're only applying the max width on medium, so they don't have one on smaller screens, which I—oh, they have the `py-8`, so that's probably how they handle padding on smaller displays, but that's still so weird. Is there a point where it gets narrower? No, interesting. Okay, that's 100% page zoom; the font's that big. I like big fonts; I'm like the Big Font guy. Yeah, individual article tag for every message is a not quite semantic HTML. Are they using data start and end tags to indicate where in the token count this starts? Cuz that's a hilarious way to do partial updates.

It's weird. I was just trying to find like the vertical scroll container, but it looks like it does just go over the area like that should be showing, but this area underneath—oh, this has a background color of its own. Then can I find it? No, no. Which of these elements has the background color that's covering that? Then how the heck is this covering that? The content is below; how is that not appearing here? Something has to be covering it; one of these elements has to have a background color. I want to find which one it is. `Pointer-events` is one thing, but it's not showing the text; like this text when I scroll, it should appear behind right there, and it's not. It's not even using the border color. Okay, that's hilarious. If you turn off a single `width: full`, the whole UI breaks. I'm sure they've done that; that's why they put the like article tag with `full` on everything. Why is there a scroll bar on the chat? That's weird. The scroll container is not smaller; I just showed that it goes the whole way, like this goes there; it should be showing. One easy way to check there; it might be that there is another scroll container within here, but I've hunted through this a bunch of times and can't find it. Yeah, `main` has that; there's presentation. Okay, here it is. It's an `overflow: hidden` on this box. This was a very hard-to-find box, but now it's actually highlighting the problem I have, which is that this box is a very weird section of the UI, and if I make that smaller, then it does this where it's—does that mean I can scroll too far? How does that get handled? What the hell's going on here? Cuz when I have enough width that the selector isn't in the middle, this becomes all the way to the top. Oh, I see it; they have `margin-top` that gets applied at the same container size, so that `margin-top` gets set and a `translateY` also gets set as their hack for changing the scroll container size at different resolutions. That's actually kind of hilarious. The amount of CSS hacks in here makes it hard to even navigate, but I'm not here just to roast the code. I was just curious; that became a deep dive in hell. I see why they're never fixing these things because that's not fun code to work in. Let's find some more fun edge cases. I'm going to show you a personal favorite of mine. I'm a paying user for 001 Pro. It's okay. Okay, 200 bucks a month, but yeah, here is one of my favorite quirks. Solve Advent of Code; let's find a really hard one from like 2023. So I'm looking for older ones; if I find older ones, they'll already be indexed. This looks nice and hard. Cool, go with that. Uh, 2023 day 9 in Python. So I'm using 001 Pro, the super expensive model. It's going to take us time to think, so because take a while, and it's not going to update the title in that time. I'm going to make another thread and ask for instead, so it's a little bit faster. Okay, it froze for a sec there, then it did a search; it's doing that. Cool. What's happening here? Still going. Interesting. They might have fixed this bug because it drove me mad before, and I complained a lot about it, but for the first month I had a ChatGPT Pro subscription, if you made a thread it would fail. I'll try 003 mini instead. I still hate the new title not coming in quicker; like it's so easy to do. Hasn't failed yet; that's good. Others are saying that it still does that when they try it. It's become a thing; like if you're using 001 Pro, you need to just let it do its thing. If you leave, there's a high chance it will fail. If I change tabs, maybe reasoned. Okay, but it's not going to show me the result there's—so I'm not even like intentionally hitting an edge case there; there's just so many of them. I'm not aware of what—what—how did it reason for a minute and 23 seconds and not even find it? My lord. Of the three AOC day 9 tests I just did, this one with the expensive 001 Pro model just failed outright; that one every time I click it, the title lags. Oh, is it selecting the wrong one right now? What's going on there? Did it just change the order randomly? That's what happened. Cool. This one it thought with 003 mini and just didn't get a result. The 004 is fine, but it searched the web; it probably—yeah, it found a Python walkthrough that solved it. Oh, that's not a clickable link, by the way. These two are—this one's not—what? Why are you citing your sources here and then putting sources here that are entirely different? I thought our search was bad cuz we just like kind of threw it together. I'm not going to like go out of my way to defend our search a whole bunch, but like I'm going to just take the same problem, hop back to 003 chat, switch over to one of the models that has search, uh, flash. I don't know who won the Super Bowl, grounding details, and here we have all these with confidence happening in seconds versus new thread, paste, search. Go. I wasn't that slow; people were telling me in chat that their search is really slow. It's not that bad; that's fine. Yeah, honestly, I would still say ChatGPT search is probably better than ours overall. Perplexity is way better; Perplexity wins for sure, but theirs is fine. I almost forgot to run React scan on ChatGPT. We got to give this a shot, obviously. There we go. Love that when you hover, this everything re-renders a few times. It's not as bad as Claude at all, but it's still not great. Well, let's actually make some messages. God, that's re-rendering way more than it probably should. Do a new one. I can't even see the UI because it's re-rendering so much. I won't do 001 Pro because it's so—001 Pro, new message. God, oh God, why is it rendering the dollar thing every single time? Oh, I know; that was the one that was in local storage and the rest weren't; that's why. That's hilarious. But look at the whole UI constantly; like the share button is rendering every time. The general experience isn't too bad; like the scrolls all handled properly; when you hit the bottom, it triggers some JS there to cause a render, but when you're not sending a message, it's not too bad. What happens if you're generating a message and you're in other things at the same time though? If I switch threads, it's still re-rendering a whole bunch while that's going on in the background. No, that's good; at least I kick it over the Dolly button and only the Dolly button rendering constantly. It's beautiful. Did the popup to let me know it finished pop up in the chat that I'm already in? Okay, I—I don't know what you guys—I hate this notification. I think this is a terrible—it disappeared, the little notification that pops up when something's done. It's the worst; I don't know why anyone ships that; I hate it. We need to do something better to indicate when our streams are done, but that's not it now. Do it in Python. Somebody ask what's the square that keeps re-rendering? That's a great question. Is that like a tracking pixel or something? I have no idea why that one little corner is being re-rendered constantly. It's like a neon lights show; it's like I'm at an EDM festival. The Popper provider for the stop button here has rendered over 1300 times so far. It, once again, is a reminder I picked a slow model accidentally. Yeah, significant look bad, and as I mentioned before, I want this one to not be necessary, and I'm ready to make it happen. As there are a couple small other things I have to tidy up first. Different worlds, different ball games. And if you look at like the CPU utilization, you'll understand even more so; it's kind of absurd. You guys get the idea. I got to try one more real fun one. This is a test I'm going to try on multiple different things. I'll grab the same problem, solve Advent of Code 2023 day 9 in TypeScript this time using 003 mini. We're going to hit send. I do something a little stupid; I'm going to hit refresh. Finished reasoning; something went wrong. This message is now in the ether; it might finish, it might not. Despite something going wrong, I can't send another message because it has the—the stop state here never got a title. Hop over to Claude; we'll try the same thing. Quick refreshed; the threads just vanished. Now I'm in this Untitled thread; new chat. Gone forever. See if it has that in its history still? Nope; it just lost the whole chat history. IDE is still the same, but it lost the history. Beautiful. The reason I'm showing this one is I'm actually insecure as hell about how we handle it in T3 chat. I've been talking to a lot of people about how we can solve this. Streaming in; I refresh while it's streaming in; it's immediately an error. I can still send messages; I even say "try again," and it will. React scans on; I can turn that off. There, it actually did it. Awesome. Look at that. And if I retry up here, it will clear the history, but I'm going to show you guys another really bad bug that pisses me off a lot. I'm going to fix this first thing after stream; I've just been busy and had to film this. Will pull back in all the broken messages cuz I don't delete them the way I need to be deleting them. Yeah, there's quirks here for sure; I will not pretend otherwise. I do not like how we handle if you terminate a stream while it's happening. I don't like the error states that are persisted incorrectly right now. They are persisted correctly if you don't retry, but if you retry, I am hard deleting, not soft deleting, which causes a weird UI state. If you refresh right after, it's an edge case on an edge case on an edge case, but it still pisses me off beyond belief, and I will fix it. But that's like the difference I want to emphasize is we're a two-person team that is haunted by these small things and fixes them as quickly as they can. OpenAI and Claude are giant teams that just don't seem to care about the craft anymore, and it's kind of funny to me how so many people say that like I'm not solving the real problems; I'm just a ChatGPT wrapper. They're not solving the real problem of how we consume the things they build either. And I'm not saying that cuz we're better than them or they're better than us; I'm saying because it's just a different focus. Their focus is on making the greatest LLMs ever made; our focus is making the best product interface with them. These are different things; you can't really do both at that level at the same time. I'm obsessed with the details of making great web experiences; that's why I built my own sync engine and my own highlighting layer in order to make our site as performant and reliable as possible because we care too much. I'm basically building my own framework at this point around it. I'm going to exit soon, either to TanStack Remix; I have working builds with both, but like the level we are going at with the details of making this a great experience isn't something we can reasonably expect a company like OpenAI to do. And even if they did hire someone who gave a—they need to let them have that level of ownership. And the problem with a company that has a single focused mission is anybody who is distracting from it to do something different that is their own mission is going to be ostracized; they're not going to be given the resources and the buy-in they need. It's really hard for a company to have two missions in parallel; you can only really do one thing. It's advice I give startups all the time; it's super common when talking about content stuff. CU, they'll be like, "Oh, you're doing well on YouTube; we want to make a YouTube channel for our company; how can we do that and get a bunch of users through YouTube? It seems like a great growth hack." My first question to them is, "Open up your YouTube watch history on your phone; scroll till you find a company YouTube channel." "Oh, you just had to scroll for five minutes, and the video you found was something from your friend's company that you watched out of politeness." Crazy. To which they immediately respond, "Yeah, I don't watch a lot of company YouTube channels, but that's cuz they're—they're bad. I'm going to make a good one." And I'm like, "Okay, that would make you the first company ever to do it; that'd be incredibly innovative. Aren't you currently trying to reinvent how we think about databases with AI? Aren't you currently trying to rebuild the IDE that we use every day? Aren't you currently trying to rethink how procurement happens at scale businesses? Why are you going to compromise your mission, the thing that actually makes your company different for any amount of time, to pretend you know how to do YouTube? You're just wasting your goddamn time." And I see a lot of people dropping examples; in every one of the examples that has been said in chat, I'm not going to share cuz I don't want to out companies. I have talked to the people who made the decisions to do those channels, and every single one of them regrets it and wouldn't have if they went back in time and advises others to not as well. It's just not worth it. And hiring an influencer is one of the worst things you can do because influencers aren't going to—just—people don't watch my channel because they want to hear all about T3 chat. If you are here, hi, I appreciate you. They watch my channel because it's a wide variety of different things. Your company channel can't have that same variety because your company channel is about your company, which means even if you put in more money, put in more effort, and make better stuff, the variety is low enough that an independent YouTuber is going to perform better. And if you want to do this, the solution isn't to try and build it internally yourself; it's to partner with the people who are doing it well. If you want to get growth from YouTube, my sponsor info is in the description. If you want to have the best UI possible for your AI chat app, my phone number is in Bookface. You shouldn't be trying to do this yourself. If you're fighting against somebody who's—it is the main mission of—you will never win a secondary mission competing with somebody where it's their only goal. This is my only goal; you can't beat me; join me. And there's a lot of companies that realize this. One of my favorites to work with is Grok; we use Grok with a Q, by the way, not with a K, with a Q. Grok is a provider for AI models; they are focused exclusively on making chips and infrastructure that make it so models can run really, really, really fast. They have a UI, and it's actually one of the better provider UIs; of use, it's really minimal, but they're working so, so hard to get us adopting stuff as early as possible, offering us crazy deals to make their models the defaults. They are trying so hard to effectively make our UI the best UI for AI chat apps right now, the best showcase for what they've done because they know they can't compete; they're not trying to; they're not pretending they even should. They are trying to find a way to align their obsessive mission with my obsessive mission, and that's the way to do it. And if I hop over to DeepSpeedAI, just to show it, same problem. Do you see how fast those tokens are coming in? It's stupid; it's incredible. And I don't think there was really a UI that properly highlighted how good it was before. Yeah, that's the way to do it. Your company should have one focused mission, and if that mission is "build the best AI models possible," you're not going to win against me; you're not going to compete against me in the UI the same way. If I was on the side trying to train a model, you guys would make fun of me. Like, forget everything I just said, guys. Hey, chat, I think OpenAI's models are bad; I'm gonna go make my own. I'm going to train it on all the data from T3 Stack users and make the best AI model ever, and I'm going to hire a whole bunch of people to come help me out. Give me your honest reaction if I was to sincerely pitch that to you guys; play into it. Goddamn it, nerds. The point I'm trying to make here is that if you express the intent to do something that's that far out of your mission, it would just sound dumb. If I was to come out here and pretend I'm going to magically make the best model ever, you guys would rightfully make fun of me for it, but for some reason when a company like OpenAI pretends they can make the best product and user experience when it's not their mission, people say I'm the stupid one for competing. It's just funny to me because to me, I—I see unfair competition. I agree when—when people say it's unfair and a suicide run to compete with and build a chat app for them, not for me. I genuinely feel like the competition is unfair; it is what it is; we will win. Speaking of winning, there is one more AI chat app I want to use because it is frankly significantly better than it has any right to be. In the mobile app, did a bunch of the things I plan to do for our mobile app that's coming eventually. This is Grok, like with a K, and Grok with a K, not to be confused with the other GRs; there's many at this point actually. Is a pretty good experience for doing—let's start with search. What football team has won the most Super Bowls in the rain? Deep search; give it a nice hard one. The scroll container here is a little weird because you get to that and then like that scrolls separately, so the page scroll containers are not the best. I'm going to zoom out a bit because I know I'm on a small screen. This is all pretty solid though; like the nice little dedicated search thing; I can expand it to see more. Don't love that, but it's fine. Show thinking separately on the side; fine, but everything like moves and behaves properly. Here's the crazy thing though; I just refreshed and it caught up fine. It's the only AI chat app I know of that resumes properly. Resumability of a stream is a very difficult challenge, and funny enough, it's one of the first ones GMO tried when he was playing with T3 chat because it is so hard to get right. They did it though, and they're doing crazy stuff behind the scenes with data persistence in order to make that possible, but it's actually pretty cool. If I refresh in the middle of that, it like resumes in a weird way, but it does resume properly. Give me all—okay, that was a weird scroll back to the top; it did edges, but it's one of the best ones I've used. The actual UI for Grok is—I would argue—close to competing with ours. It has its quirks; they all do, but it's pretty good. I also wanted to call it a message because I like this one a lot. Webdev is viewed by fundamental companies as an easy side quest when it's actually a nightmare. Yes, absolutely agree. Take a peek at how Grok search is implemented. He was playing with the network connection to see if it's inlight. There's a dedicated endpoint, rest app, chat conversations, reconnect response. They're using NDJSON, which is newline-delimited JSON, which means that new lines can come in individually. I—I don't even know if I have a dedicated video to HTTP streams; I bring it up in a lot of stuff, but if we go to like T3 chat and I use a slower model, 003 fireworks will be perfect for this demo. Hop over to the network tab and do a write five poems about JavaScript. You can hop over to the endpoint here and see what's the response is coming in as. What's interesting is the response comes in line by line over an HTTP stream and still fast enough that it like breaks this view a whole bunch, but you can see all of these coming in. The two prefix bees a JSON thing that I can parse; zero prefix means this is actual text content to render, and it's pretty good overall. It's like relatively reliable, and it's the main way to get messages down because that way you can stream the token as soon as you generate it and build a good experience that updates quickly. The problem is that's just one request, so if I refresh before the request is done, it can't resume the request that's half done. You have to either persist that somewhere and catch up or other piles of hacks. Another good point from Nan in chat here: "Maybe we're at fault as the users; we put up with annoying UI and refuse to pay for better experiences." Half agree. It's—to me, this is like—this feels like the piracy problem where game piracy exponentially declined after Steam came out because people are willing to pay if they're aware and it's convenient enough. And I think T3 chat being cheaper, better, offering more things, and also being much more transparent and like accessible has allowed us to pick up an awesome user base real, real early. And in the end, like Anthropic still wins; what's their incentive to compete anyways? Like, this is our month bill for Anthropic right now; they're making more money just letting us pay their massive markup than they would be selling the $20 a month subscription and pissing off their users. It's effectively like—this is a thing I've thought a lot about; the comparison I'm going to make will be a little bit weird. I want to talk about Fortnite. Fortnite is a very misunderstood product because Fortnite's goal was never to make Epic a whole bunch of money by selling like loot and items and V-bucks to kids. The point of Fortnite was to showcase the type of game SL experience, SL metaverse stuff that Unreal Engine and the Epic Games ecosystem was capable of. It's something they've invested in so heavily that externally it looks like that's the whole thing; like Epic is the company that makes Fortnite, and they also happen to make an engine for games. I genuinely believe Epic would be pumped, hyped beyond belief if somebody else was to use Unreal Engine to go make something bigger and better than Fortnite. Their goal was to show the world—specifically in that case—the horrifying experience that was PUBG, which was one of like the worst engines ever built; it was built on top of Unreal, but it was a total mess. They wanted to prove that a battle royale and Unreal didn't have to be that garbage, and they kept iterating on it and trying to showcase to the world you can build things like Fortnite with our tools. That's how Epic wins. Fortnite makes a lot of money, but they reinvested in all sorts of other things, including Fortnite itself, to expand how big that showcase is. But at the same time, like you can say they're making a lot of money on Fortnite; they let it not be on the App Store anymore; they're making like a billion a year off of it being on iPhone, and rather than cave and let Apple do their thing, they stood their ground because they believe if they can force Apple to have to lower their fees for game devs that they will be better off because more game devs can make more money and pay for more Epic tools. They were willing to let one of the biggest cash cows imaginable die in favor of proving to the world that you can use their tools to make incredible things and to fight for game devs to make more money; it was not making more money from Fortnite; they died on the hill. A lot of companies are able to and succeed in negotiating for better deals with Apple; Epic is fighting for the industry to have these deals, not for themselves. That's also why they're doing Epic Games Store with the crazy, crazy small cut they take on it. Yeah, this is Sweeney wants to do these things; to be clear, like Sweeney is choosing to use Fortnite on iOS as the sacrificial lamb to prove his point. I think that is awesome. Here—here's the theoretical I'm trying to—to frame for you guys to better understand. Imagine Apple decides to start banning AI chat apps if they have a voice feature; it's an absurd example, but it's a real one. OpenAI wants to keep making money off their iPhone apps, so they remove the voice feature; cool, awesome. Anthropic wants there to be more, better mobile apps, so instead of removing it, they put it in and at the same time, in parallel, plan a crazy PR campaign and lost suit so they can fight for everyone else to be able to have the voice feature in their AI apps because in the end, that means Claude can be more competitive with other things because their model can be an option in those other apps. The ecosystem is the thing that these companies should be investing in, and funny enough, I think Claude and Anthropic do a really good job of this, like with the MCP stuff that they just put out, Model Codex Protocol. Been hearing a lot about it; this is an Anthropic invention that they built; they could have made it a specific thing for Claude, but they didn't because they want the ecosystem to move forward because if the ecosystem moves forward, more people adopt more things that are AI-based, they can make more money long term. They probably put more time into MCP than they put into the entire UI for Claude. And if they didn't, then I have other separate concerns. This is a great example of fighting to move the ecosystem forward because you own a percentage of it, and that's a good thing, and we should incentivize that. I'm just annoyed that the UIs are this bad, and the response isn't, "Yeah, they should be better," or "Yeah, I'm going to go use a different app instead." The response is, "Yeah, but you're stupid for competing with them; you're not a real dev; you're not making your own models." Well, it turns out the model providers are probably not the ones who have the mo—it's actually the chat apps that we use and work with every day and the other interfaces. I'm more likely to switch between OpenAI and Claude than I am to switch off of Cursor, for example. Yeah, I go back to Grok for just a little bit, then we'll wrap up. Experience is pretty good overall. The history is only this like command K style view; they don't have a sidebar, which is funny because my CTO Mark really wants to—Grok was unable to load your last message; that's a fun UI state. If I refresh, will it work? We built the only stable AI chat—I didn't want this to be the case; I wanted to show Grok as a good one; I came here to do it, and I did. They all just break constantly. This is nice; like the icons, the scroll D appearing immediately is nice. I was—we need to change that; the current is a little weird, but I get it. If I delete the one I'm in, what happens? Leaves us open but navigates behind; that's fine. It's really not that bad; like if I was to tier list this, we'll start with everyone's favorite ChatGPT. I think C tier is actually pretty fair; it looks fine; it has some weird design decisions, but the—the bugs with title gen and the like lack of persistence was real bad. I think that—I'm leaving it in C for now. The mobile app's pretty good, so if we were including mobile, it'd be a little bit different, but for now, I think C is fair. Let's grab Anthropic Claude. If this was just looks, Claude would be B or A tier, but it's not just looks; I'm going to put it D tier. I would put it F tier; the reason I don't is they're actually somewhat responsive; not that like they're reaching out to me and saying, "Hey, I fixed the thing," but they're clearly paying attention and getting the fixed. There have been multiple times where I publicly complained, and they made the change like a week later, so it's—they're paying attention; they're doing it slightly faster than like YouTube is, but even YouTube makes changes when I—enough, which is unbelievable to me. xAI have one in here; cool, it does. xAI absolutely B tier; I would say it's—how do you guys feel? B or A tier for xAI? That one bug where the chat history disappeared; I had to refresh was weird. Y'all are saying A tier, and other people are saying that like it's not the right logo; it is the right logo for xAI. Everyone's saying A tier, even—okay, my CTO is saying A tier; I think that's fair. Also, the mobile app's really good too, but I—I think A tier is fair. I—I like having a sidebar, but I respect the balls to not. I honestly feel like the sidebar is more useful for demoing things rather than the actual UX of it, but yeah—yeah, I think that's where I'd put Grok; think that's fair. We can't forget to run React scan on Grok though; I don't want to give him too good a score until we know. For the most part, Grok seems really solid, which is why I feel bad stress testing it with React scan, but I have to do what I have to do; you guys know that. So let's do it; that's fun. Does every charact—yep, every character typed—that's a controlled input; that means if you copy-paste a big enough thing, like—it's just grab this guy and just paste it over and over in there; eventually, it's going to start getting real laggy because this is all being persisted in the JS layer. I hit enter accidentally there; oops. That's a lot of renders happening. I'll start at the wrong scroll location when you're doing too much at once. When you paste too much text, it starts at the top, which is weird. You don't have any indication the chat's actually going—not too bad, all things considered. I want to play a bit more though; I want to get this chat box to lag. Keep pasting the bottom of here. Oh, look at that; it's already starting to break. Testing; what if I type really, really fast? Okay, yeah, it's doing key lag. I typed that correctly, but it's sticky keying a little bit. Yeah, you can see I'm just going like one to nine on the keyboard; watch this. Okay, it's doing better this time, but there was a point there where it was putting numbers in the wrong order for a sec. It's not too bad; like it doesn't feel sticky keys, but if I paste the whole thing twice, can we get to sticky keys? Yep, we're at sticky keys now. Okay, our input box theoretically should never do that because we're not recalculating on every key press; we're not tracking the content until you hit send. Yeah, they're not even doing anything like—they should be locking me up because I have too much text here, and they're not—they're not doing anything about it. Yeah, it's not that bad. Let's pick a fun language for this one; um, let's do Zig. One thing I've heard about the Grok model that actually sounds cool if it's true; I'm waiting till I have an API to test it. Apparently, it's better at more obscure tools; like if you're using Spelt or Zig or these other languages, it's better versus something like Claude, which is really good at React and TypeScript and popular things but struggles more elsewhere. It is hitting some bad FPS here; we're at like 18 FPS. Is the CPU just getting hammered? What's going on here? Yeah, it just hung out at like 70% CPU for most of that, by the looks of it. Let's try again, just to be sure. Solve Advent of Code 2021 day 18 in TypeScript. Okay, this time it's not struggling as much; I don't know why it was struggling so badly prior. Still not great CPU utilization, but syntax highlighting is expensive; I get it. It is kind of funny; it's like if you're going to use 80% of my CPU, why not just run the model locally? But as soon as it's done, it flattens back to nothing. Yeah, the—the CPU and memory usage is not bad overall. I have no idea why that one thread got so laggy. If I go back to it and ask for it to follow up, "Can you do it in Python now too?" I don't like their scroll behavior either; I put a lot of time into ours. What we do on T3 chat is when you send a message, I create the new messages immediately, and I scroll up so there's enough space for it to fill, but it doesn't auto-scroll; there is no auto-scroll in T3 chat; it just on send. If you're already at the bottom, the container gets bigger and puts the two messages in, and then it overfills; if it goes longer, it's actually much better behavior than I've seen other things. How'd I go on that whole rant? It hasn