📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

George Hotz | mixture of experts (like deepseek) on tinygrad sovereign AMD stack | AMD YOLO

george hotz archive8:18:05

Transcription

Be good now. Get the microphone going. Turns out my keyboard was uh, was cheating and using blue—was cheating and using blue—was cheating and using blue—was cheating and using Bluetooth was cheating. Yeah, get that on repeat. Get that on loop. Was cheating and using Bluetooth—was cheating and using Bluetooth. It's a good loop. Uh, so where's chat? Why, why is the internet so hard to use? Like, how does anyone use this? I don't know, really. I have a lot of questions about like people on the internet; they like move things around. I can't wait till AI just uses the internet for me. All great. All right, good. We got some—we got some good bandwidth. We got some milk tea, which I kind of got rugged on. I thought it was coffee, but it's actually milk tea. We got an empty lemon Spin Drift. There's full lemon Spin Drifts downstairs. We got a shirt that reminds you to look left. Uh, you know, you can get great things in the—in the Bargain Basement, Hong Kong.

Um, I got—so I went to—I went to China for the week. We can—this doesn’t have to be subscriber only; you guys can—can all be in when we talk about going to China for the week. So yeah, went—went to Shenzhen. Uh, show you guys, I went shopping in Shenzhen. What did I buy in Shenzhen? This is a—more threads—see more threads—S80. Here’s the model. It reminds you that you use the CPU connector and not the GPU connector, which is actually quite smart. Um, the GPU connector is total nonsense, uh, cuz it only does 150 watts. This is an 8-pin connector that can do 300 watts—four 12 volts on top, four grounds on the bottom. It’s very sensible. The six-pin GPU connector is kind of nonsense. So good to see someone thinking from first principles there. Um, overall though, I hear it’s a crappy GPU. Uh, I just bought it because I’m excited to see one on the shelf that Marus and Guan B—here’s an AMD—this is a 60—this is a 56, 76, 50—gr Edition GPU. So this is like the—the cheap Chinese market gamer Edition. Um, and then I bought all sorts of stuff to plug that in over USB, but uh, you’ll be seeing plenty of that in the future, so have to show that now. I also bought this power supply. Guess how much? Guess how much? Guess how much? 500-watt power supply. Guess how much? $20. Yeah, pretty much—a little more—200 RMB. Uh, so I got that for 200 RMB, and then here’s all the USB GPU things that I bought in this little box here. It’s a lot of little thingies. We don’t need them all. They can be unboxed. We have the whole tiny gang coming—uh, you know, our crew is kind of like fsociety, you know, fsociety. Um, so we got the whole tiny gang coming uh, out to Hong Kong uh, for the month, and some of you are invited to—if you’re a Tinygrad contributor and you want to come on the Hong Kong trip—uh, yeah, we—we’ll see. I mean, this—how you get hired here, right? Kind of a—kind of a remote company. Show that you can produce insane value for the Tiny project. Tiny never really had an office in SD. We rent part of Comm’s office uh, in exchange—in the—say, as part of the joint comma Tiny Corp agreement.

I got—I got stopped by Hong Kong immigration yesterday. Uh, you know, the guy’s like trying to scan my passport. He scans it like five times, and then—and then the person comes out, and I’m like, oh, okay. Uh, but you know, there’s two types of power in the world: there is—there is rational power and there’s psychotic power. And the difference between these powers is when you comply with rational power, uh, they lift the boot; uh, when you comply with psychotic power, they push the boot on harder. Um, so if you remember like all like the—the protests, all the like the campus protests where they would go and like sit in—in the Dean’s office, and then people would comply with them—see, that’s the wrong kind of power; that’s psychotic power. Because when you comply with activists, their demands will only increase. Uh, so which—which is why you really just have to exclude activists from the conversation entirely. You can’t talk to activists. Uh, but—ir—rational power, if you comply, uh, it’s reasonable. So you know, I’m selected for secondary—secondary immigration, secondary screening. Uh, you know, in America is a terrifying experience. You know—you know, they’re gonna—you know, they—you know, they’re going to do some weird shit to you, man, like you don’t even know what it’s going to be. Um, so I’m like sitting there, I’m sitting there with like uh, three other Chinese guys, uh, and nobody looks nervous, which was interesting. Like in America, you’re in that same situation, there’s like a 10—yeah. Uh, you know, why—why are they bringing me here? Um, so I sat there for maybe 15 minutes, and then uh, this woman brought me into the uh, interrogation room, and uh, a few things—a few things about the—the interview really stood out. Um, the first was they explained to me why I was there. In America, they never explain—like they never explain to you—maybe a cop will tell you why he pulled you over, but sometimes it’s like some bullshit reason. They’ll never tell you in immigration why you were selected for anything. So they were like, well, you know, you’ve been selected because you’ve been coming in and out of the country a lot. Fair—fair. I’ve been coming in and out of Hong Kong a lot and following all the immigration rules. Um, and then she asked me a whole bunch of probing questions, but something I found interesting about the interview was in America this interview would be very adversarial. Uh, I would very quickly be thinking, okay, do I got to get my lawyer? Like, is this going to be a—you know, honestly, if I was put in that same situation in America—and again, I’m an American citizen, I’m a guest in Hong Kong—in America, I’d just be like, I’m not answering any of your questions. Uh, look, you have 24 hours to charge me with a crime, or you can let me back into the country. Uh, I’m happy to hang out for a day if that’s what you choose, but uh, you know, I didn’t commit any crimes, and uh, you—that’s just how it’s going to be. I’d rather not answer that question, sir. Um, they were very reasonable, even with my thing. So I always turn my electronics off at the border. Uh, you should always turn your cell phone and laptop off whenever you cross an international border. Uh, this is just like basic OPSEC stuff. Uh, so you know—and like, again, it’s—it’s an interview; it’s like, who are you? You know, what are you doing here? Um, but the whole time I was in there, it felt like—like compliance—the more that I comply, the more that I actually am not quite forthcoming, but the more that I’m—you know, like we’re all on the same team here in this interview, uh, the better things work out for me, and things actually—it—it really went well. And the main thing that they were trying to discern was, are you here to uh, bring—take money out of the country, or are you here to bring money into the country? And I really respect that because that’s the thing that you should be looking for in all immigrants, right? Whenever you’re thinking, should I let this person into my country as a visitor, as an immigrant, as a worker, whatever—are they trying to bring value in, or are they trying to take value out? Uh, so—so I—I think once I explained that, you know, I’m an owner of a business in America and uh, you know, I’m being paid a US salary and I’m coming here and I’m spending money at Hong Kong restaurants and so on and so forth, uh, the interview uh, you know, it became not a problem. Um, it was very quickly not a problem. You know, I didn’t violate any of the rules, so you’re absolutely happy to—they’re absolutely happy to—to have me here, and they gave me another—the normal 90 days and said that uh, you know, if I come in and out of the country a lot, it’s triggered by just an automated system. Uh, so again, very cool. Good signal for Hong Kong. Uh, and it just—it shows you what can happen in a society where there is trust, and it’s one of the things I really love about this place—that it is a high-trust society, and uh, I try my best to participate in a high-trust society in a high-trust way. If you’re in a medium-trust society, which America is, you participate in a medium-trust way. Look, I don’t think everyone in America is trying to kill you, but you know, be on your guard. No officer, I’m not interested in taking that field sobriety test. Um, you know, I—I—honestly, I—I would prefer to opt out. Uh, with all due respect, uh, I would—I would feel more comfortable having—having a lawyer present at this—interaction, and that’s just the attitude you have to have. It’s an adversarial attitude. It’s a medium-trust society; it’s not a low-trust society. A low-trust society is—is a whole different—like, you know, who’s—who’s—who’s pulling the Glock out first, right? Uh, so I don’t ever want to uh, go to or live in a low-trust society. I’m not—I don’t want to—the minute Glocks are out, you’ve already—you’ve already—you’ve already—you’ve already crossed the line, right? Even when everyone walks away from this—why are the Glocks out, right? Um, so—oh, someone should hire me as Hong Kong’s tourism master. I legitimately love this place. Um, again, there—there are so many wonderful things about a society being high-trust uh, that—that—that I really—yeah, I’m happy to be here. Um, yeah, I’m thinking about working on applying for residency—setting up a Tiny Corp subsidiary out of Hong Kong. Um, there was discussion in one of the group chats I’m in about Elon’s special economic zone, and I listed like the things I wanted. Uh, you know, what would you want in a special economic zone in America? Maybe you’d want something like uh, you know, easy high-skill immigration. Uh, you’d want there to be no crime. Uh, you’d want there to be cheap housing. Uh, you’d want there to be—you know, not no tax, but like maybe a 15% corporate tax rate. And then I realized at that point I was just describing Hong Kong. It has all those things. The immigration could be better. Uh, you know, I’m—I’m—I’m lucky. Passport privilege is a real thing. It’s nice to have a US passport. Uh, so it makes things certainly easier uh, on that front. And yeah, I mean, Hong Kong could be better about—about a simpler pipeline for—for high-skill immigration, but overall it’s better than America’s uh, on that front. From a crime perspective, perfect. I don’t want to live—I think Singapore—I mean, it’s a nice place, but I think that almost takes it a little bit too far. You know, I don’t think that—that’s—I think that—that’s a little bit uh, extreme. I’m not sure, you know. And then Dubai is another one. I think—I think there—there’s—you want a little bit of crime; you don’t want no crime, right? You don’t—like, or you don’t want just all these crazy things to be crimes, like chewing gum or, you know, that kind of stuff. It’s kind of a meme about Singapore. It is a lovely place, but uh, you want there to be a little bit of edge, uh, which I think Hong Kong’s pretty good about. But you obviously don’t ever want this to be violent crime. How does a society tolerate violent crime or like breaking-and-entering style property crime? Like, that’s insane. Someone broke into my—the comma office to try to steal a bike. Like, you know, that’s insanity. How do we have that? How—how—how is that—like—like—how can—how can a state—I mean, you’re not a state at that point if—if—if there’s people doing uh, violent crime. Let’s start with just violent—what the hell? A state is a monopoly on violence; that’s the definition of a state. You’re not a state if there’s people frequently committing violent crimes in your territory. You don’t control your territory, right? Think about OG—you know, there’s no difference between the government and a mafia. I think about a mafia—you can’t do a hit outside the mafia. That’s not how it works. We do the hits, okay? And if we don’t sanction it. So effectively all violence in America is sanctioned by the American state, and that’s what you have to realize, right? Or they’re not a state. There’s—there’s one of two alternatives here: one, all the violence you see is sanctioned by the American state, or two, the American state is weak and incapable of actually uh, you know, operating, so it has to be one of those two things. There’s no other explanation for it. Uh, no violent crime here. All you can look up the murder rate here, and you look up what the murders usually are; they’re either domestic or uh, you know, gambling, heat-of-the-moment passion kind of things—not all very much things that you can entirely avoid. Uh, and look, I’—I’ve never—you know, I’ve never had a—I’ve never—I don’t even like the phrase been a victim, but I’ve never—I’ve never like—it’s not like violent crime is super common in America. There’s a lot of places I wouldn’t go if I actually thought it was violent, but property crime is super common. I’ve had property crime so many times, uh, particularly in California, uh, and it’s just insane. It’s just insane. Like, you’re not a state. Um, you can’t—you can’t have that. Um, yeah, and I’ve been—I’ve certainly been like harassed and touched uh, by street people, and like—can’t—can’t have that. Um, and yeah, the tax rate—okay, now what do you—it’s not crime enforcement—what the hell are American taxes going to? I have no idea. Can’t run a state. Why should I pay for this?

And then the cheap housing—look, people are like, oh, robots are going to fix everything. Oh, robots are going to come in and build cheap housing. It ain’t robots. Housing is kept artificially scarce. It’s not that like we can’t build it; we could build it, especially if we bring in people from Bangladesh—do what Dubai does. You can build that; you can replicate that in America. Why not? Um, so yeah, uh, the Elon special economic zone: high-skilled immigration, no crime, 15% tax rate, cheap housing. What do you think? See, that’s innovation; that’s—we can get innovation. Uh, do I believe America is its people or its culture? America’s got a lot of things going for it, and if—if you—if you want to fix America, you really have to lean in on America’s strengths—against some of the strengths that are here. Like, you have—one of the most wonderful things about America, and I worry that it fades a little bit, but America’s civil religion—the idea of diversity—you can make diversity work; you just have to not be stupid about it, right? This idea of diversity—that—this idea of diversity—that people of many different races can work together is phenomenal; that’s just a huge talent pool, right? You know—any—you absolutely want this—people from different cultures, people from different races, people different sexes, genders, whatever—that’s who you want working together. Now, what these people will often launder in with that idea is that all ideas are equally good—an idea that—well, maybe in order to reduce homelessness, we should give homeless people a check. No, see, that’s just a fucking dumb idea, because wow—follow the incentives. We give homeless people a check; now there’s more homeless people. We give more checks; there’s more homeless people. Who give more checks? Wow. You see, now you got to wonder something, right? Are the politicians who propose those policies idiots, or are they not working for America? Traitors. It’s got to be—again, it’s got to be one or the other, right? They’re just either completely incompetent, which is possible; they can’t understand the second-order effects of their policies, or they’re misaligned; they’re working against you and your interests, and they want to destroy America, right? It has to be one or the other; there’s—there’s no other room for this, because these policies, over and over again, lead to the bad outcome. It’s obviously predictable that it will lead to the bad outcome, so are they too stupid to see it, or are they not on your side, right? I know—never attribute to malice what you can attribute to stupidity. They’re honestly probably just stupid—stupid. But uh, so yeah—no, I mean, America’s—America’s civil religion—and there is a sense of freedom and space that you just don’t get here. Um, you know, part of it is just by the nature of being a city; everything’s kind of crammed together, right? In America—yeah, I’m going to go get some land and a pickup truck and some guns—this good time—and you could do that. There is a sense of—but you got to get rid of all the regulations, right? I got—I got some land out in California. Can I go buy—and tariffs are regulations, right? Get rid of the regulations. Could I buy 100,000 solar panels from China? Get some Bangladeshi labor to come install them, wire them all up, and then have a lot of power in California and lower my electric bill? No, because it violates a whole bunch of regulations. But imagine I could. There’s nothing stopping that, and America wins, right? I own 3—3 acres of land about 30 minutes east of San Diego. Import huge numbers of solar panels from China—no import tariffs—just—just huge numbers of solar panels coming on boats from China, getting shipped to my land. Now I also have contract labor. I’m paying a minimum wage—what the hell is that? Get rid of all this bullshit, right? You can’t have this—you can’t have this. You want to be rich, or you want to be poor, right? That’s it. Oh, no, but you have to ensure a fair—no, no, no, no, no, no. You’re not thinking fourth-dimensionally. Make the pie big; make everybody rich. So I bring in—I bring in, you know, like you say Bangladesh, but I think there’s lots of countries where you can bring in uh, labor—labor from—I know Dubai brings in from Bangladesh. We could bring in—I mean, labor from Mexico is pretty good. Bring in—bring in—bring in cheap labor—bring in cheap, hardworking labor. Um, again, it’s a good—it’s not like people are like, oh, this is exploitation; this is slavery. No, it’s fucking not. No one’s forced to go; it’s an opportunity, right? You either take it or you don’t. Um, so bring in these people, have them set up the solar panels, right? Well, again, buy some cheap Chinese steel—I mean, we got to spin up that American steel too, but uh, get some—get some uh, stanchions, shove them in the ground, get my solar panels, get them all connected up. Um, yeah, and we’ve just raised the value and brought more literal power to America. So think about it, and then—oh—oh, you’re going to have to advertise that over the tax—over the blah blah blah blah blah—15% flat tax on profits. Well, your solar panels are actually only deductible of blah blah blah blah—no, no, no, no, no. How much money did you make, and we’re going to take 15% of that, right? That’s just fair. What happens if—if—if Cal—on that—look—like—like—you know how you bury your head in the sand, right? You bury your head in the sand even deeper. We can fix this. I don’t want you—I don’t want people looking on my posts about America and taking away an absolutely bleak perspective, but I do really think that the demoralization is just beginning, and in order—you know, the first step to change is realizing you’re an addict. The first step to—to making progress is realizing that there’s actually a huge problem that needs to be fixed. So uh, no, I’m in Hong Kong, but I’m an American citizen. I’m—I’m—I’m—again, I’m not an idiot; I’m not—I’m not blindly patriotic. I’m a value buyer, right? Is America good value, or is Hong Kong good value? I’m a value buyer, right? Oh, George, there’s no—you don’t have any loyalty. Yeah, that’s not how things should work, right? Think about brand loyalty. Nothing should be like, oh, aren’t you still buying Sony TVs? Don’t you have brand loyalty? No, I’m buying TCL. Oh, but that’s Chinese—no—it’s a nicer [Music] TV. Value buyer. Create value; don’t try to scam. Why is it so hard to see—for Americans to see that there’s a huge problem that needs to be fixed? Because a lot of the people are the problem. Remember when I did that—remember when I did that survey on the uh, on the stream? One-third of people actually are the problem. Needs to be fixed is in the—if you’re—if you’re raking in the dough from the scam, maybe the scam doesn’t need to be fixed. Uh, so there’s way too many of that. Someone’s really got to investigate—how are these politicians getting paid 300K salaries yet have net worths of $70 million? How is that happening? Someone’s really got to look into this. Um, no, it’s not—here, if there’s one article for you to read, uh, it’s linked from—it’s linked from one of my blog posts. Uh, there’s one article for you to read, and please sit down and read this article in full. It’s a phenomenal article. Um, it’s called “The Value of Nothing: Capital versus Growth.” Uh, so read and understand this. It doesn’t just attack traditional US thinking; it attacks Silicon Valley; it specifically attacks incumbent hearts. Uh, yeah, read this, and you understand what the problem is. We can fix it. And you know—you see Elon’s new Starling Factory—now that shit’s amazing, right? We got a pretty good factory going at comma—like we’re doing a lot of the same stuff; we have a lot of the same sort of ways of thinking. One of the particular things they talk about is how you have to bring the factory close to the design so you can iterate on the design to make manufacturing easier. So many things we found with comma, and the reason we’ve been able to bring the price down so much is just—if you design better, you can manufacture for cheaper—or not just better, but if you design with manufacturing in mind—well, shop teacher used to say that all the time—design for manufacturing. Um, Julius, I don’t know who this is; I should read all articles, but they’re good. The real class war I’m interested—you have it all figured out, George. Congrats. I don’t know if that’s sarcastic or not, but you know, something I find uh, when I look at the comments where people say like, oh, well, you know, Packer news—people are like, oh, George, really stick to reverse engineering or blah blah blah blah blah—like the thing about—that’s—that’s very low on the—on the tiers of uh, argument. You’re not making a coherent argument against me. Also, the gold standard? Yes, the fucking gold standard. Yes. Oh, but you could use Bitcoin; it’s just like gold. No, it’s not. Do you know why there’s 21 million Bitcoin? Because some GitHub repo says 21 million. What if I submit a pull request that says 22 million? Oh, well—no, no, no one will merge that. It’s 21 million, right? What if I bribe them? Yeah, well—no, they’re not going to be bought. Yeah, but what if I like slowly do an influence campaign over five years, right? Getting the key people who can actually get that pull request merged, and we’ll just merge the pull request—they’re 22 million—called the 22 million project, right? It’s like the 1619 project, but for Bitcoin—22 million project. The 21 million of Bitcoin is entirely socially constructed. There’s no fundamental reason that it’s 21 million; it’s just cultural. And eventually anything you do that’s cultural will be worn away. But you know what can’t be worn away? Literal physical gold. We’re gonna make 22 million gold. Yeah, how you gonna make it? You’re gonna—you’re gonna—you’re going to go mine, all right? Now we’re talking. All right, you’re going to get some trucks. Yeah, make them out of steel and aluminum. Yeah, you’re going to employ some people to make some trucks. You’re going to work on some advanced robotics to get the gold out of the ground better. See, now we’re talking. See, that’s an economy. And you know, it’s not an economy lobbying people to make it 22 million, and then—like the Fed’s even worse; you don’t even see the 21 million. There’s not a GitHub repo; there’s not—no code that says 21 million; it’s all fugazi, right? We got to go back to the gold standard. If you want money, dig it out of the fucking ground or convince other people who’ve dug it out of the ground to give it to you by providing value, right? There’s really only one way to get money. The only way to get money in the world is to convince other people to give it to you. Um, maybe in the gold standard there’s another great way too, which is you could dig a deep hole, and digging deep holes is very valuable [Music]—people’s time. I don’t know. I’ve seen AI—man, more faith in gold—just real—yeah, convince people, right? I convince people. Yeah, yeah, yeah, yeah. See, that’s what you’re gonna end up doing too, right? Right. The 22 million project? No, no, no, that would never happen. Yeah, but what if I had the backing of the US military, right? And again, we’re not going to do like a very clear—it’s cryptography; it can’t be changed. I’m going to kidnap the miner families. Yeah, 22 million bitches. All right, it’s 22 million project. That’s just real. No, really? How long you think it’s going to stay 21 million for? Jord, you can’t say that; that’s heresy to Bitcoin. See, the beautiful thing about gold is there’s no heresy. You can’t make heresy to gold, right? You can try to—oh, this guy’s a gold bug; oh, he should really read this economist. Let me tell you about economists, right? They’re kind of like Harvard professors when you ask them about race and IQ, right? You ask a Harvard professor about race and IQ, and then they—they think about it for a minute, and then they’re like, Harvard is committed to diversity, equity, and inclusion, and then they give you an answer, right? After they read that. Uh, so it’s the same thing, right? You ask an economist, right? It’s like, is printing money okay? I’m being a PO by the government who wants to print the money, and—oh, yeah, yeah, it’s called modern monetary theory. You see—you see—okay, modern monetary theory—monetary theory, right? And like, I just—I just want you to try to explain modern monetary theory to some like Roman Empire guy, right? Some like guy in a Roman Empire—no, no, no, no, no—the money doesn’t have to be backed by gold. You see—you see—you’re only limited in your money creation by inflation. Yeah, yeah, yeah, yeah. And inflation only accelerates when there’s full employment or—what the fuck does any of this mean? Gold, George. You just have to understand economics. No, this isn’t economics; this is politics; this is people who want to convince you that for some reason their fake money is better. Oh, but you see, you need to have lots of value because you have to support a growing economy—what—what are you talking about? Support a growing economy? Price goes up if there’s not enough gold; price goes up—supply and demand. No, no—you have to print money in order to—no, no—but you see, there’s going to be lots of recessions. Yeah, as opposed to massive inflation. So—so people are like, oh, well, you don’t want to be on the gold standard because of uh—let’s find like a—let’s find like a long-term—um, inflation over time—cumulative—yes, here we go—cumulative inflation since 1913. Okay, so you see when you have the gold standard, everything’s good. Then you start to go off the gold standard a little—and a little bit of inflation—okay, it’s controlled though, but you see like—like now, instead of a dollar being backed by $35 an ounce of gold, it’s backed by $145th an ounce of gold, and at least people can see what the inflation is, so you’re not—you’re kind of limited by what you can do, right? You can’t—you can’t—$145th—$35th an ounce of gold is now backed by one—now dollars backed by 100th of an ounce of gold. You know, everyone’s like, oh, you really rugged me, bro, but you know, they keep the rugging little, and then Richard Nixon comes along, and he’s like, rug the shit out of everyone, and look at what we get. And the trick is you have to look at cumulative inflation because they always talk about how—well, no, it’s 3%; it’s always been 3%. Yeah, but it also—sometimes used to be negative 3%, but now we avoid crashes and economic recessions—yeah—by making everyone fucking poor—by literally taking 30X of everyone’s value. Like, George, you just didn’t read enough economics. If you read more economics, you’d understand that inflation’s actually good, right? Remember—remember when—like—oh, God—I—I can’t get over like the stupidity of—are they brainwashed, or are they in on the scam? I think they’re—I think they’re in on the—I think they’re on the take. Remember—you’re—you’re an economist, right? You’re gonna—you’re going to—you’re going to talk about what’s best for the economy; you’re going to talk about what’s best for your career, right? You’re in social sciences at Harvard; you gonna talk about you know the truth, or you going to talk about what’s best for your career? Um, solution: vote for more Republicans, and they don’t solve the problem. Call them rhinos. No, unless—like the Republicans aren’t saying go back to the gold standard, but it’s—it’s not like this is a uh, you know—this is like bringing back—no—this is literally what—what happens when you have fake money. Oh, but we have new technology now. No, we don’t. The Romans understood how to make their money fake too, and eventually they did, and that’s what happens to every—what do you think happened to Roman dollars? I’m sure they had dollars; I promise you. And I don’t know; I’m no Roman Empire—you know, puff—but I promise you what happened to the Roman dollars is they got inflated out of existence because that’s what happened to every single currency ever except for the ones that happen to exist today. So what if we just like didn’t do that? Restored the gold standard. Be like, sorry, we rugged everyone for the last 50 years, but now we’re going to stop rugging people. Better late than never. Hey, but you know—okay—wow, we had a long—we had a long opening rant. We had a long opening rant. Uh, wait—no, something definitely happened in 1971. In 1971, America got taken off the gold standard entirely. To be fair, they’d been doing it for a while since someone had to pay for the stupid European War we never should have got involved in, but hey, that’s a political take; that’s—that’s a real ocean—big ocean—know big ocean—yeah, big ocean—we talking—you know—all right, good rant, Mr. Peter—sh. I didn’t realize—like, I can’t believe that I ever fell for crypto. Crypto’s fine; stable coins are cool. Uh, a lot of parts of it are fine, but the belief that like crypto was some fundamental backing store of money because there’s only 21 million Bitcoin—are you kidding? 22 million project, George? You can’t even say that. Someone might do it. Yeah, I know, but what—you think I’m the only guy who’s ever going to say it? And then why are we going to stop at 22? All right, let’s get down to the technical stuff. Modern mod—I’m just triggered—I’m just triggered by—okay, I want to be triggered; it’s a good day. W—it’s a good day. It’s j [Music]—um, let’s put that over here, so I’m routing—it’s a little bit faster—I’m routing this through my Tailscale network. Tailscale is really good, by the way. You guys use Tailscale? I mean, I’m sure they’re the kind of company who’s eventually going to rug you when they’re

Uh, want to stay in? You want to do stock buybacks with profit made from the business, on profit that you don't easily see how to redeploy quickly? I don't know, uh, but yeah, know I kind of like that. I think like that's if there was ever a way that I would take profit from one of my companies, it would be that. Um, I have zero interest in like selling to investors or selling to bankers. What are you going to do? So stupid. Um, but if you can make money through like this, if I can, if I can buy AMD stock, uh, with Tiny Corp work on better drivers for AMD, drive demand for AMD, drive the price of AMD stock, literally everybody wins except in video, but you know video wins too. Like everyone kind of wins here, I think.

The demand for AI has barely started. Um, I think Nvidia will likely, at least for a long, long period of time, have the high end. If you insist on the highest performing stuff, you will be buying Nvidia for at least the next five years. But if you're interested in, would you rather fight like Nvidia is selling you one horse-sized duck and AMD will sell you a 100 duck-sized horses? And if you're cool with duck-sized horses, then, then yeah, Nvidia is probably not the place to buy the duck-sized horse from. They sell you the horse-sized duck, and there's a premium. It's a, you know, it's a, it's a very large duck, right? So yeah. Um, no, stock buybacks with low-interest loans, come on. You do stock buybacks with actual profit that you've made by delivering value to the end customer. Uh, yeah, no, uh, very smart, uh, uh, quant guy explained this to me. I'm like, why don't companies pay dividends? He's like, dividends and stock buybacks are the same thing. Um, and actually like it's kind of better I think to do stock buybacks because again, there's so many ways you can scam with stock buybacks, but stock buybacks are not inherently a scam.

If you're using stock buybacks to try to juice your price, yeah, you're scamming. Look, we up all of this, all of this, uh, you know, orders off the order book, and now there's our price is higher. You got to stop looking at the price. The price isn't that important. You have to think about the depth of the order book. Come on, every shit-coiner should know this. Um, right, just because your, your, your Melania token is worth $3 billion when you multiply the last traded price times the total number of Melania tokens, like, no, no, that's not how it works. You got to look at your exit liquidity, but look at the volumes. The volumes are high. Yes, I understand the volumes are really high now. Sell a million dollars of it and see what the price goes to. Um, so George doesn't understand economics, came up in the crypto world, man. I think, uh, I might understand economics a bit better than, uh, you know, some, some, some copium folk. Got to look at incentives. The technicals of understanding it aren't all that hard. It's, it's mostly, it's like, like understanding, yeah, all right, we're done. Rant season's over.

Um, so what we're going to work on today is a mixture of experts, open source. So I think we can actually, we don't have to go to the tiny box yet. We can probably just do this on Mac. Let's take a look at which one we're going to use. So you guys understand what mixture of experts is. I think mixture of experts is probably how the brain works. Um, so you have a problem with, uh, me three. So you have a problem like with llamas, your llama 70b is accessing every single one of the 70 billion weights every time, and this is incredibly taxing to your memory bandwidth. It makes memory bandwidth and memory capacity basically the same thing, so you don't want this. Um, when you look at DeepSeek, DeepSeek is a mixture of experts model. To be fair, GPT-4 is a mixture of experts model too, so this is not like, it's not like a revolution by DeepSeek, but I think that most models in the future are going to go in this direction. So this is the total number of parameters in the model, but this is the activated number of parameters in the model. So each time you run the model, uh, it's also possible that none of this stuff's going to matter because we're going to move to a diffusion style architecture, but I'm less familiar with diffusion architectures, but I think even there, like mixture of experts is just a good idea. You should not be accessing every weight for every query. This just makes obvious intuitive sense. If I'm asking a query about archaeology, right, we shouldn't be an, we shouldn't be activating the neurons that relate to the Riemann hypothesis, and this just different. So you have a ratio, your activated params to your total params. Uh, so let's do it the other way around, so we can actually see the number. So yeah, DeepSeek is about, uh, 18x. So this is really large. We're not going to be able to load this on our Mac, but let's take a look and see what open mixture of experts models we can find. Oh, these are just, here's another one. I don't know which of these is the best. We support Mixl in TinyGrad, but Mixl's also rather large and kind of old now, but I think we'll be able to reuse some of this code. Uh, then we're going to have to try to get our mixture of experts models in our Jetson. There's a bunch of reasons that this is tricky, mostly because we have to access different weights based on the, uh, you're accessing different weights based on what your previous layer said. Looks pretty open, even more open than OpenMoE. This is also pretty open. Who are these people? Yeah, I was looking at this when I was in the coffee shop this morning. So we'll just pick one of these and get it to run. Mkay, they write Rags, uh, 64 small unique experts, of which eight are activated. Good, it's pretty big. Um, why didn't you use Mi-100x router Z loss and load balancing loss? Observe that they train 2x faster than dense LLMs equivalent active parameters. Let's check out this paper. Shoot, yeah, so this is, it's a good picture. I like, I like this project. Props to AlMo from Contextual AI. That's what we'll be using. So when you see this as a 1B, 7B, I think what that means is there's 7 billion weights, but only 1 billion of them are accessed each time. So I'm going to have to download this. Uh, cool. Yeah, so it clearly outperforms everything with 1 billion active parameters. A little aggressive to, I mean, it does have a lot more stuff there. Oh, this is also a mixture of experts. Wait, Quen is mixture of experts for, oh, I can register an Alibaba cloud account. Huh? So okay, it looks like they did one mixture of experts, now it's Quen 2.5. So they moved away from mixture of experts. I wish I, I use Grok for like everything now. Elon, you got to get rid of this. Do I have to log in? That ruins all my fun. Oh, oh, maybe I don't have to log in. It's going to make me log in after I type the question, which is an evil dark pattern. Oh, sick, it didn't make me log in. All right, now we're talking. Now Grok really, I use Grok for everything now. The killer feature of Grok that ChatGPT lacked was real RSS. Um, like Grok, OpenAI being a gear and a half out of date made it like unusable for so many things that I would use Google for, and now instead of Google, I'll very often go to Grok because it's up to date. Um, and you know, and you know what? I love it. I love it. And the reason it's like safety, it's like safety cocks. It's like, we got to spend six months with the model and make sure that it's not going to tell anybody to go kill themselves. Um, and then Grok's like, well, you know, have you really considered? Um, yeah, and that's like good, good, cuz we, this not like you, this isn't even what AI safety means. This isn't even what AI safety means, right? You've seen my legs. This one open, do interesting. Okay, so they kept doing mixture of experts just for the, uh, private one. I see. I really think mixture of experts is, is the future, especially for these bigger models. For 7B, it probably, uh, doesn't help that much, but once you start to get into these things, I mean, I imagine these models becoming, uh, like terabytes, and that's fine, just you can't access all the terabytes every time.

By the way, so this, these are the, the big project that I want to do with the, uh, the MI30 Xs that, uh, AMD gave to us. Also, I saw someone saying that AMD probably gave them to us based on some like non-disparagement clause or some signed contract or something like that. There was nothing like that, uh, and I appreciate them for understanding that I would never sign anything like that. It was literally just, uh, yeah, we're going to send you the boxes. What's your address? All right, here's my address. Here's a FedEx tracking number. Uh, so respect. When AMD wants to do something, they actually just did it. Uh, I was, I was, I was expecting, I gave, I gave like a 20% chance that they were going to be like, just sign this non-disparagement clause. I'd be like, fuck off. The hell, I didn't think you can buy me for that kind of money. No way, right? Um, but no, it was nothing like that. So don't, don't feel that I'm ever restricted in what I can say. Um, so I'm not like, yeah, I'm not like, like, uh, if it sucks, it sucks. I don't know, there's a chance I'll get these. Actually, this isn't right. I wish that were true. So say slower than that, but, um, yeah, no, if, if the hardware actually does end up sucking for some reason, uh, I will say the hardware sucks, but everything I found with the 7900 XTX, the 7900 XTX is phenomenal hardware. Um, we have, we have two PR requests that I'm very interested in. Uvn is working on, they have all this tracing stuff that seems to just work where you can trace like every instruction run, uh, in a shader. B1TG is working on, and these are just external contributors to TinyCR. It could be you, uh, is working on getting not using co-manager anymore, just calling right into LLVM. I'm not, LLVM is like TinyGrad does not replace LLVM. LLVM is kind of the last piece of external stuff that TinyGrad uses. Uh, you know, when I'm replacing kernel drivers that were written 20 years ago and hacked on and hacked on and hacked on. Okay, that's not that hard to replace. When I'm replacing new code that's written as copy Nvidia, but just rename everything to HIP instead of CUDA, yeah, that stuff's easy to beat, but, uh, can I beat Chris Lattner in [Music] LLVM? I don't know, uh, I'm not that, I mean, prob, who knows, right? Like it's not like, like, like when you're, uh, like playing up against people like that, or whenever you're playing, this is something I said about like always like Tesla versus Karma. There's no advantage that Karma fundamentally has over Tesla. Elon is thinking in the same cutting-edge ways, right? There's some companies that it's just, when you look at the advantages Karma has over something like Toyota, there's no contest. They're not thinking in new ways. There's no, there's no one at Toyota who's like capable of saying, well, shit, what if we just use a smartphone GPU and run AI? No one can say that. Oh, we have this approved thing from this vendor, and it's going to take 17 years, and you know, we'll, we'll have it part of, like, that's, that's the process of these companies. So it's easy to beat those kind of people. Um, then it's easy to beat like software development teams that you can see what their incentives are, but can I beat, uh, LLVM? Who knows? We're, we're, we're on a level playing field here. We all know the latest tricks. We all know the latest tools. That's true about PyTorch too. Like, can TinyGrad beat PyTorch? And the answer is, we can make different tradeoffs from PyTorch that may be more appropriate for certain users, use cases, but can we straight up dominate PyTorch? No, they're very intelligent people working on PyTorch. Some of them are in our Discord now, and like, I met a high-grade programmers. Uh, so I'm, I'm up against, I'm up against like tier programmers here. Uh, so you know, it's not like, um, yeah, there's going to be any clear domination, uh, of things like that, but I do think PyTorch made some tradeoffs that makes it tricky to do some of the things that TinyGrad will end up doing, uh, or TinyGrad can already do, and I think that's where a lot of the value will, uh, will accrue, but it's not that the PyTorch team was dumb. It's that the PyTorch team made tradeoffs that were appropriate at the time. So many times you're, you're, you're often, uh, with that, you know, thinking like that versus when you compare TinyGrad to like Qualcomm's SNP, yeah, we're just going to dominate them. Uh, we have better engineers who better understand the problem. I dominate like the poker charm, you know. Uh, so these things have tons of RAM. Why am I using Google? Don't ever use Google. How many free ones are they going to give me? Let's say if this really, if this really stays open like this, I will use Grok for everything. Please sell my data to every hedge fund or whatever, whatever, whatever. Oh, um, yeah, so this is, this is the theoretical memory bandwidth. Uh, so it's 42.6 terabytes per second. That's terabytes divided by 37 gigabytes. Theoretically, that's the max tokens per second doable with DeepSeek on one of those A300 boxes. All right, a thousand tokens per second would be sick. Have you been watching? I've been watching Claude play Pokémon. No, I bet on Twitch. Some guy wired up Claude to play Pokémon. It's cool. You know who's there like the OG Twitch Plays Pokémon and stuff. Um, [Music] So the big downside to it is how slow it is. Oh, the model is like brutally slow, and then it's like, I'm gonna, I'm in battle with a Zubat. Here's the status of my team. Okay, let me think about what I'm, I'm going to do. Okay, I'm here. Let me think about it. Press A. Okay, I pressed A. It went to the next screen. Let me think about what I'm going to do. Here's the status of my team. I am going to need to select an item because I'm selecting an item. Okay, the item is down. Press down. Right? Like, no, just make that go 10x faster. It would be so much more engaging. So I want to get some sick super-fast DeepSeek stuff up on the MI300X boxes, and we'll see what we can get it to do. Um, yeah, so the first step to that is getting mixture of experts to work in TinyGrad and be fast. It's in China. I always get a cough when I go to China. Something in the air. Oh, I also got the weirdest, freakiest thing. I went to a massage place in China, right? And like, you know, the White Lotus spa, spa, you know, U, but it wasn't like that kind of spa, spa. You lean back in the chair, and this woman took a stick with a very sharp point on it and cleaned my ears and like sticking it in the ear canal. It hurt. Um, another, another example where I'm just like, you know, for the most part, I trust this society. I see people moving around here. Everyone looks pretty healthy. Nobody's deaf, um, at least that I could see. So it's like, I trust whatever is happening here. It's okay. And it was okay. I also went to the, the dentist in Hong Kong. I got a root canal and a crown, um, and it was another like, wow, I kind of actually like, I just trust this a lot more than I trust the American medical system. Uh, you paid it a pocket. It wasn't cheap. Wasn't like crazy expensive. It was like a fair price, but like you do the thing, you go to the front, you get the bill, you pay the bill, done. I give him my credit card, I pay the bill. Why can't you do this at a doctor in America? Well, so actually, yeah, okay, you're going to pay us now, but that's just the office visit co-pay. Uh, so you're going to have to pay for the nurse, but that's to a separate contracting company. Don't worry, you're just going to get three easy and convenient bills in the mail. Um, you're not really sure which ones you have to pay. Maybe some of them you don't have to pay. Who has time for any of this shit? I'll pay you 20% more. Here's my credit card. Oh, well, you see, that's a PPO, and you're in an HMO, and that's going to be out of network. So we're going to be able to bill the nurse to your insurance, but the medication, well, that one is actually the doctor didn't write that. It could be generic on here, so you're going to get that filled now. If you do get that filled and you go to a Walgreens, uh, the Walgreens is going to have the, the Walgreens prescription discount if you don't have insurance. Now, I'm your insurance company will advise you to use the insurance, but of course, uh, if you're going into the Walgreens, I can't really tell you otherwise, but Walgreens can just, you can go out of insurance, and then they can change it to the generic, and that's going to be cheaper, but I can't really tell you to do what. Why do I have to know all of this stuff? Um, that's why I only plan to go to the doctor in Thailand. Yeah, yeah, I saw someone die in Thailand when I was there, uh, a couple of months ago. I talked about this on stream. I, I saw, I saw a motorcycle accident. I saw someone die. Uh, was a little traumatic, and you look up the road safety in Thailand, and like it's the worst in Asia by far, and like you can see why. So I don't know, I mean, I'm not saying like Thailand's bad or anything. I'm just saying that, uh, I wouldn't blindly trust the Thai medical system the way I'd blindly trust the Hong Kong medical system. That was a bit of a tangent. We got a lot of tangents today. Oh, okay. Well, I got three free Groks. Where were we? Oh, we're reading the AlMo paper. Um, okay, use dropoutless token choice routing. What does that mean? I want to see what we can reuse from Mixl and stuff. See if we can get this running. Oh, yeah, which browser are we switching to? I heard, I heard Firefox sold us out. I, we talked about this little last time. Google sold us out a long time ago. Is it Ladybird? Yeah, I'm using Brave on stream. I trust Brennand Ike at least there's like a name associated with, with it. So you know, it's unfortunate that they're in bed with crypto scams, but to be fair, like BAT was pretty early. It wasn't clear that all ICOs were a scam at that point. No, but this ICO is not a scam. You have to check out the new Dent coin. It's going to fix that problem you were talking about with healthcare, and it's going to unify them all on one simple coin. Also, sign up for the airdrop. You can get an airdrop. Get airdrop you Dent coin. All right, here's our Hugging Face. Which one do we want? Probably instruct. Why does this one have more downloads than the, why does the base model have more downloads? License Apache 2. What I really like about these, all right, so it's safe tensors. Why are there three of them? Like, how do I deal with this? How do I deal with this in like MixDraw? All right, let's go. So I think it's just straight up, we even have to say download equals true. Uh, why did that get downloaded like this? Whatever. So it keys it based on the, I don't think we need download equals true. We really need to make fetch, we really need to make this an HTTP tensor, so these fetches can happen in parallel. So fetch is just like, it's a little helper in TinyGrad. It's stupid, but, uh, actually, we don't have to use fetch. We should not use fetch. Fetch is the old way of doing it. By the way, TinyGrad has phenomenal docs. Watch the docs not have what I want right now. Uh, but yeah, that was over here from URL. So you can just do tensor.fromURL, but this isn't lazy right now, and it should be lazy. [Music] We'll let that download while we, uh, do other stuff. Oh, that's cool. I have a Wendy B, tell Allen Institute for AI. This Paul Allen. It's pretty cool. I've seen some good stuff from them, like Wendy B let you do notebooks like this. What they're running the ARC challenge crashed? Cool. Yeah, I think this is a good one. I feel good about this as a choice. Wow, this is new too, or did they just update the paper? Oh, they updated it with two. Which do I have a change list? What they change? Upside, what is Upside doing? It's so cool. I, if Hugging Face is really profitable, that makes me really happy because it's so cool that you can just download these huge things from this website, like no login, no nothing. All right, let's take a look at how this is actually going to work. Oh, it's just transformers. Interesting. And we could probably just do it through the PyTorch backend, but I don't love that. This is what like everybody builds on and stuff though. Now I want to get like really high token counts. Uh, so quick estimate of what our, what our max token count can be here. Uh, so we're on my Mac, and then is this going to be a, what, uh, so it must be FP32 or FP, FP16, I mean, it's got to be right because we're downloading, yeah, 16 gigabyte model is FB16 or is it BF16? I'm sure it just says this in the pap. It's fine. Okay, so it's probably BF16. You can quantize to eight pretty well. I like that the DeepSeek model was actually trained at eight because whenever you're quantizing, there's always loss, but if the model's actually trained with the quantization, uh, you can also do quantization-aware training, uh, which does the same sort of stuff, right? Like, but if you're doing this post-quantization is kind of, iffy. Okay, so this is BF16 entirely. Oh, mixed precision training. Okay, fine-tune in BF16. Well, we'll find out quickly. Actually, we can probably just see it if we go, [Music] here. Hugging Face should show me all the, here we go. Okay, yeah, so the BF16, that's a lot of weights. Cool. Uh, we should be fast with BF16 on Mac, but if we're not, this is a good time to fix it. Okay, so it's 1.3B active, 6.9B total. M3 is 400 GB per second. Max is 400 GB per second. 2.6, I don't understand why it's 2.6 and not 1.3. Who can do math? Nobody. That's why we use Google. So in theory, we could get 154 tokens per second. I would be extremely happy with 100. Actually, I don't know. Let's not, let's not undershoot ourselves. That's the max theoretical tokens we can get from this model. So this is just, you multiply this by two because we have 16, uh, so it's unquantized. In theory, we could get 154 tokens per second. How's our download going? I need this. It's actually, I'm just not a 5G hotspot. It's how a lot of people get internet in Hong Kong. It would be cool to have an actual fiber. I know this is one of the first places to have gigabit fiber. Look at this, thousand M home broadband, home WiFi. How do I actually get fiber to my home that work? Wow, aggressive. I'm not showing you guys my address. I don't know. Yeah, if I'm going to be here, uh, so I'm going back to San Diego for the summer. I don't want to be here for the summer. Sweltering 85 degrees with 95% humidity. Ah, got, we got interns coming for the summer too. It's always fun to be on comma when it's intern season. Okay. Um, we can start loading this now. So let's see if we can use Mixl and what we have here. I don't feel forward fun tools. These numbers are probably incorrect. Why does it say jit false stuff in the paper? All right, so the number of experts actually, oh, you see, yeah, this, this, this isn't going to work because it's assuming that there's only two, which is not right. So we're going to have to change that. Um, I mean, it'll still run. 64 experts, crazy high number of experts. Yeah, so Mixl uses eight experts per layer. Two of which, 64 experts. I like this, like bolding too. This is cool. This is a well-written paper. Okay, so how much layers is it? I think it's these. Here we go. Okay, so a hidden dimension, that FFN dimension, that norm, that rope doesn't matter. Num layer 16. So there's 16 heads. How many KV heads are there? Is that the same? Oh, using MQA. I think no. So is the default of that none? Yeah, default of that is none. So we actually probably don't need that. What's the vocab size? Some strange number. What kind of tokenize this thing? Us, who's got the tokes, toes, toes, toes, toes? Lots of toes. Oh, good, that one downloaded a little faster. Can I please have some fiber? Just need some fiber. Vocab press. This is like a meme. It's like, oh, yeah, it's only 1 billion, but then you need 500 million vocab params. I don't know why they're treated differently. Do you, the want a CPU? Is all the training code open for this too? I don't know. I mean, I think these things can be trained pretty well. Oh, see, this is throughput, which is a little bit different. I'm interested with the AMD stuff. I'm not interested in throughput, right? A lot of cloud fuckers are into throughput. Um, cloud fuckers, but I'm into, I'm into high individual TPS. How fast can my AI think for me? Talk about data. Is there like a common, like, there's got to be just common stuff everyone does these days. How big is the data? 17 terabytes? Eh, where's, uh, LibGen and his archive? Looks simple enough. Uh, how much tokens was this trained on? Where's the paper go? That's going to be fun to get to work. What's DeepSeek's number on that? So V3 and R1 are the same model. R1's just like fine-tuned. I'm pretty sure that's correct. It's a very different paper. I haven't downloaded DeepSeek or anything yet. We're also going to find out how much, what SSD is AMD gave me in the boxes. Uh, Corker is setting them up, uh, is going to be on Monday. So we'll have them up Monday. Um, so how many, uh, experts is this by the, all our code looks like, like hedge fund code. So interesting. They make things work, and they make things work fast at hedge funds because it actually like matters, like algorithmic trading code. Hedge funds not the right word. Okay, we use 32. No, that's parallelism. Where does it actually talk about the model architecture? Oh, one shared expert. Interesting. Okay, so I imagine a shared expert is just one that always runs amongst the routed experts. A will be activated for the token, and each token will be routed to be sent to most four nodes. Okay, interesting. Yeah, cool. Let's just look at the, about hedge funds and making things fast. Let's see what we got. Oh, yeah, I've seen this paper. It's a good paper. Uh, no, like, this is, this is what you really start to think about, and I want to get to this level with TinyGrad. Well, things are getting crazy. I like the MLX guy, and I like the project. Uh, I don't think there's a future in, uh, hardware-specific frameworks. I mean, if any company can pull it off, it's Apple. People are in their own weird Apple ecosystem. Never really coded in Swift. I wouldn't spend time learning it because it's not used in a lot of places. Um, but you know, to each their own, right? Like, for people who are in the Apple ecosystem, maybe this stuff is appealing. Okay. Um, that running anything. Why is that slow? What's slow? This initialization is slow. Why is that? We'll get there. Okay, whatever. Um, we have state, which is going to load. Um, see if this is fast. Should be cool. Glad that's fast. By the way, you can see it. So this is accessing the JSON from the, uh, safe tensor. So what we're going to do for a lot of this, uh, stuff to make loading faster is, I could just like blit those big files to the GPU and not do each one individually. In fact, if you want to see that, you should be able to do like two metal here. That should just work. Yeah, cool. See, see, we get item there. Great. We're loading it at 6.5 gigabytes. So I mean, we could do two metal there, and then that, uh, actually loads them from the disk, uh, which is cool. All right, now I got to go through and assign the weights. Tqdm, you can tqdm on a dictionary. Tiniest tqdm you can. I'm sure if Chinese can do it, 31 can do it too. We have our own tqdm in dyr because I didn't like it being a dependency. State, sure. CI, see if that works. All right, we got some weights. [Music] I'm going to find some more coffee. Amanda, do we have more coffee? Are you trying to rickroll me? Yes. Okay. All got that dry China cloth. Why is this slow? Let's try to figure it out. So TinyGrad has a useful helper called profiling, and you can just do with profiling. Let's see why it's slow. No, don't do that. All right. [Music] Um, it's just like randomness that's slow. It's just doing a lot of randomness that's slow. Wow, it's spending six seconds in Rand. That's annoying. Want to like work on optimizing this Python? I hate the idea of waiting six seconds. Who's got time for that six seconds? We don't know what the world's going to be like in six seconds. All right, chat, focus. Just make a database of all the rules, and then we can use all the rules and then have expert systems. No, none of that shit worked. Okay, build infrastructure. Wow, on topic. Why is this slow? All time we're wasting. Should we just spend the whole rest of the stream making this Python not slow because it pisses me off that it's slow. Be faster. The profile FL, I love how easy it is to just do profiling. Okay, so all the slowness is in rammed, and then we can like look at why this is slow. We can also throw a with timing around that, but it's profiling, so that impacts a little bit, but not too much. So we can actually just get a final verdict on the time. See if we're making things fast. All right, 1776 milliseconds. All right, that's a lot of milliseconds, yo. On topic, going to ban people. We have like so many layers of indirection here now. What is this crap? Mop, hello. Come in. Thank you. It's, I got rugged before, and it was tea, and then Amanda tried to gaslight me and tell me it was coffee. It was, yeah, this is coffee. No, that was tea. One was also coffee. I tried it. It was also coffee. She's trying to gaslight me. It was hi, no, it was tea. They're, they're, they're off topic. What are they saying? Some shit about like phonemes, like, like, what's a, it's because people want AI to be expert systems and not inscrutable weight matrices, but they're wrong. On it's inscrutable weight matrices, and you just got to have more FLOPS, bro. I'll let you guys get back on topic. That's good. Thank you for the real coffee this time. That isn't tea. I promise. Last one's also, enjoy. We're not even working on AMD drivers. Why is this shit so slow? It's because we have lots of code. Got indirection. The more code stuff has, the more slow it is. Shouldn't be slow. It's also a good thing I'm not actually doing the real, uh, three fry there. Okay, how many calls to Rand do we make? 900, 790. Yeah, they take milliseconds. That's crazy. We're making lots of calls to apply broadcasted uops. So this is all crap to deal with floats, and like we even made three fry its own UOP. No wonder this is slow. Look at all this code. Use NumPy. Think that's going to

I do not know. They might not have come with drives, too. I might have to go by drives, so how are these things attached? Um, we do have these nice—this interconnect is nice—uh, how can it fit 8X of these? There has to be some switch. Switch going to tell me there's a switch. You're telling me if there's a switch. Great, thanks, Super Micro. Uh, the box doesn't say Super Micro on it, so I don't know if it's Super Micro. Oh, it has this ridiculous board. Whatever, we'll understand this once we have the machines powered up. There's no reason to look into this now, uh, but we're going to have to, like, yeah, think about how to get the models off the disk as fast as we possibly can.

Um, onto the GPUs, and then we'll going to figure out how to make this fast. So there's a whole lot of things here that, like, is not exactly clear how to do it in TinyGrad right now. So maybe that's where we should be focusing, but let's first write the stupid code to make this load. Oh, GR could probably do it. [Music] [Music] We're going to get rid of the word model attention. It's called self-attention. That's fine; it's the same thing. [Music] FFN Norm. Are there multiple Norms? Well, this is the model, so I have Korm and Q Norm. Are those normal? Sorry, it's been a while since I've done any of these kind of models. I don't just know what's normal. You want to get good at something; you just got to do it 32 times.

Uh, okay, so why I have FFN norm and attention norm, and I have probably let input layer Norm. Let's mix—deal with this—all it kind of works. So Mixol just kind of has this—I don't have to do any Transformations on it. No, I just did torch load. So let's exclude these experts because there's too many of them. We'll deal with the experts later. Self-attention Korm, KQV, but I don't think I have norms for them. I just have attention Norm. I don't have a separate Norm for K and Q. Is that normal? Ah, you want the mi30 project, uvn. Uh, we could definitely do that. I think that sounds like a reasonable, uh, a reasonable project. Could put up a bounty for it and get you on one of those machines.

Uh, yeah, I'll I'll post when the machines are up. Uh, yes, so I think I think we'll have we'll have two bounties for that, um, and yeah, the cool thing about those bounties is I'm not going to give everyone access to the machines, uh, but yeah, no, I'll give you access to the machines. We'll have the runtime bounty and the driver bounty, and they're separate. You have to do the runtime first and then the driver. Uh, no Navy4 yet either. I could I could throw up a bounty for that one, too. Uh, my Nav4 is not here yet. I ordered one on eBay. N4 is already A4, uh, for those of you that don't know.

Um, yeah, so these Norms are just not—like, I don't have that for K. Okay, let's focus. Where is that? I don't have that. Where would that belong? That would be in my attention. So yeah, my attention just doesn't have Norms; it has weights. What I do with Korm and all right, do I have any more grocs? Come on, grock, give me more grocs. Oh, I'm not worried about SQT on RDNA A3 is fine for now, um, but we should add basic support for RDNA4. Uh, we should, yeah, at least start by adding support to the runtime. Uh, H, yeah, see, it's like it's not normal. Do they talk about it in the paper? There's only like six tricks that everyone uses for these Transformers, though, just kind of like know them. A QK norm. There we go. After the query and key, we find QAN Norm. All right, so I'm going to have to support QK Norm. [Music] Here, kind of annoying. So where are my other norms? Have to pass Norm EPS in there, too, or do SK normals nn.RMSNorm. Uh, the dimension is actually after the ection. Oh, just dim here is fine. Okay, fine. [Music] dimK, n. All right, got Q norm and Korm, so we apply them after the projection, like here maybe is right. At least we have wholesment now. Oh, we have to tell it that we have this, so I created in there to pass that all the way through. [Music] Transformer junk. Okay, uh, great. We have holes to STI them in now. Uh, no, we didn't, because we have to say QK nor equals 1, uus 5. Great, now we have Q norm and Korm. Now we're just going to have to write stuff that transforms that into that.

Okay, so first thing we want to do is say RC equal K dot—it's well model dot right sir K. 6al model dot—we have that's not part of the model LM head. Oh, well, where do I have the LM head? Let's get rid of the model. Um, so like feed forward.GateWeight becomes MLP.gateWeight. All right, which way are we doing this transform? Going from that to that, so we actually want it to be feed forward.GateWeight get to that other crab later. Um, like post-attention layer Norm weight becomes where's post [Music] attention post-tension layer Norm weight. Where's attention Lama? Where's Lama? Post-attention lay—yeah, I know I could use AI to do a lot of this faster, but you don't really learn if you do that. So is there a post-attention layer Norm? Oh, it's probably in here. Uh, it's called attention Norm, so that's—so FFN Norm is post-attention; it's probably that one, and attention Norm is which one? Input layer Norm weight. Okay, that's probably right. Input layer Norm becomes attention Norm, which is kind of—I guess a stupid name, but it's clearly on the input, right? So we're replacing it such that it becomes the, uh, like the model weights, so it should be good. Okay, see that's becoming that; that's becoming that. Uh, still the self-attentions. Jeez, man, the amount—the like people just don't understand anything. Who has—like, if you go—if you don't understand this stream, seriously NGMI. I understand when I go off the weeds into like, like crazy shit. I understand when I go off the weeds into like ials and drivers and yeah, fine, you know, whatever. Uh, that kind of stuff—I'm actually—I mean, that kind of stuff—like when you get to all that low-level stack stuff—I'm actually one of the best people in the world with that. Um, so I understand that that's hard for people to keep up with, but this shit—I'm a noob with this shit. So yeah, if you can't follow this, like, someone probably just actually knows the Hugging Face library that just does this. Inspirational. Thank you. Thank you. I really try to inspire people. I really try to inspire people and tell them to keep going, because you—keeping going is—I give a fuck if you keep going. That's right. Like, if you need somebody to encourage you, then you were never going to make it in the first place. I I think so many people like miss this when they talk about the self-esteem shit—like your self-esteem should be precisely calibrated to your actual skill. Americans rank like 30th in MCH math but number one in math confidence. Math confidence is bad; it's just bad. Like, you're just learning wrong things. You should be exactly as confident as your real skills are. Okay, embed tokens. We—well, it's probably talk embeddings. That's right. All right, we need an LM head. What's the LM head? Which one is that? Norm? Maybe there already is a norm.weight. Okay, so it's not that. Maybe LM head is output.weight. Those aren't even the right size. Uh, okay, embed tokens—if the one from the model is that. Okay, that makes me question the dimensions. These are all just wrong. Okay, so why do I think dim is that and hidden dim is that? Di is probably that. Maybe hidden dim is that. That's just backwards shape SMGE. Okay, so MLP gate is wrong, and that's because num experts is actually 64. 30th in math? Yeah, I'm taking my data from The Newsroom trailer. Nah, I'm I'm taking my data from The Newsroom trailer, but I've also seen Americans trying to do math. If you're telling me that's not true, I mean, my my lived experience backs that up, so you can tell me it's not true, but no, it's not true. It's not true. Feed forward.GateWeight. What is that? GP.app? I haven't heard of this. Cool. Okay, so we have our experts, and our experts are—so I guess our hidden dim is smaller than our dim, which is [Music] fine, I guess. Linear. So that's output features; that's input features. Uh, I—let's just do it the same way they're doing it, which would be putting hidden dim there. I think that actually is back where it's from. Uh, I'll also [Music] m—I mean, you guys want to see what the right thing to do was? You guys want to see what the right thing to do here is is actually just to print the whole list, print the other list, copy-paste that shit into AI and be like, write the code that changes these things for me, and like it would get it right.

Um, okay, so these ones are going to be a little trickier because we're going to have to get those experts out. Um, there isn't an up-down.gate kind of thing. These are all the experts junk. [Music] Junk layer junk junk expert uh name junk. Then we're going to want layers dot [Music] layer feed forward.name modelStateDict. So layer that—this—do that in F. Okay, model layer equals that uh stat. Okay, so model layer uh sub expert. I don't know if you can do that. Hope Tiny supports that. See what that should do. Okay, modelState.assignState. [Music] K uh assign device mismatch. Yeah, place. Let's copy these to metal. Should be fast. Okay, it's not that fast. Now we have assign dtype mismatch. Uh, yeah, because this should be replace. Let's get rid of those for now. Let's just move—do that—do to metal. Ian, it's a little stupid, but whatever. Uh, and this is replace, and we have the outside here to underscore metal. That'll QE it all. Right now, another value is to unpack CU. We got to split based on the dot. [Music] String indexing is not supported. There shouldn't be a string. Oh, okay. That might have loaded something. Maybe all those assigns are actually pending, which is interesting. Um, let's take a look at T at [Music] selfup. Yeah, it's not exactly right. Uh, let's take a look at where it comes from. Look at the lazy data for [Music] that. Um, so I never got assigned to, or does that not assign in place? That definitely should assign in place. Why is that not assigned to? See how it's not assigned to, but it should be just still zeros. That should work. Get it. Oh, that's broken. That's good to know that that's broken, but this is actually creating a new tensor. It's not a [Music] slice. Yeah, and then when I assign to it, it doesn't work. Huh. You see why it's broken? So that's actually creating a new tensor, and then this is never being realized. So when I realize it, it's probably fine. Uh, if we do realize.realize, that actually trigger the assign. Okay, now it like copies in all the stuff. Uh, wait, that's still zero. No, that should not be zero. Yeah. Huh. This gets into like deep TinyGrad issues. TinyGrad's not great at this kind of stuff, uh, and we have to—this is a good thing to focus on for mixture of experts models, and this is something that just works way better in PyTorch right now. So if we want to do this, what we do is [Music] this. Say W equals that. W sub uh experts.expert. W sub Z expert. Uh.cat uh state to metal uh unsqueeze Z uh wub expert one. Um, W.sign. I can do do replace there. Don't need any stupid assign. Should work great. Now you can see that that's actually aign, too. Takes a little bit of time, but that should be right. Does that make sense? Did you see what that's doing? And that's how like it has all that stuff there, so we can like print the tensor and see if it actually has shit. What? Why was that so slow? Okay, what's it doing? Okay, it took a long time, but it got it, sort of. Um, yeah, so you know why that is—like it's trying to copy in all the things, uh, but because if that's a metal bug—Alum doesn't support BF16, right? Is GPU? If we want BF16, we got to do this. Oh, this is absurdly slow. Is stupid. Oh, we have to just make assign work. So this is a different problem—like all this stuff works fine when you're dealing with small mixture of experts models, but once you get up to 64—like 64 is a large number—actually, almost—don't think—like I almost don't think it's saved correctly because instead of having the tensors look like this, it has the—all these as separate tensors, which don't tell you anything about how memory's laid out. [Music] Yeah, I mean, it is working; it's just very slow, which it shouldn't be. What's the name of this stream? Mixture of experts. Oh, I said it was going to be on the AMD stack. I'm not even using the AMD stack. This is just Metal. In fact, the AMD stack doesn't have this bug. I could do this over on AMD, and it would be faster. I also could have downloaded the model on that computer, and it would be faster. I mean, okay, this will fix [Music] it. [Music] Yay. Slowly copy in all the weights. You have 17 years to do this. [Music] Great. All right, there's all the way to the mixture of experts model, so like it does work; that's just really slow—SL slow. You know what would be faster? I think if we put the two metals there, it's fast maybe. So that's going to copy the whole thing into RAM. It's fine because I have a lot of RAM on this computer. Okay, it's not like super fast. I guess those cats are just slow regardless of—yeah, the cats really shouldn't be written like that. All right, you want to write the cats correctly? Whatever, let's slow—get rid of that for now. What we'll do here is we'll just—we'll hand-code some garbage. Uh, experts sub int layer, comma in expert. That shouldn't have t no equals that. Oh, we need name as well. Isn't exactly right. Uh, we want layer name here, and then—guess—actually, that doesn't really matter at all. It doesn't matter at all. Whatever, that's kind of more convenient to iterate over. Uh, 4KV in experts. Uh, okay. Uh, 4KV and experts modelStateDict. SubK equals uh—we want to extract them all, so do V sub I 4 I in range 64. SE when V = 64. Okay, tensor cats. Uh, no, we'll just do stack. Do stack, right? Yeah. Great guys, know the difference between a stack, a cat. That's garbage. [Music] Um, can't unsqueeze a list. Now can you? Why is that a list? Oh, splat. Interesting. That shouldn't be the shit. It's not right. Oh, what? ModelStateDict? No, that's not what I want. Uh, I want uh states. Okay, and we'll do to metal. I can't really stack them on disk, so that's not going to work. Oh, yeah, see, it needs a renderer. Uh, if we do 2 LVM and then we do 2 metal, that'll [Music] work. No, it won't, because it's BF16. See the problem with doing 2 metal here is this—this is another long-standing bug. People need to really work on the front end. You get this, but notice how much faster the whole thing is, so that's good. Okay, okay, okay. Let's put this [Music] back. Takes a little bit to load them, but at least that problem should be fixed. Great. Problems fixed. It's a little slow because it's got to load them. Is that printing [Music] garbage? It's actually loading really fast if that's real. Hang on. Uh, no, that's not the load. The load is there, and it's kind of slow actually. Uh, the load is here. Okay, the load is actually kind of fast. That's fine; it's a pretty good speed. So we're getting like 9.5 gigabytes per second, so it takes like—it's taking like three seconds to load the—no, sorry, taking like 1.6 seconds to load the model. I don't like having to sit here and wait for that every time, but there probably not gonna be in the way, so I mean, it has to—eventually it has to get onto the GPU somehow, and it's super nice how fast this is now because that's no longer—that's just like on the GPU, by the way—like something you guys don't realize here is just insane that no other library can do this. SafeLoad is actually operating on the GPU, so we're running the safe tensors on the GPU, uh, meaning like you can just copy the whole file in, and you don't have to copy the tensors one at a time. All right, now we have the shit. Now we actually got to make this work. The first thing we're going to do is run the gate. [Music] [Music] So now we on the gate, and now we get to another thing that TinyGrad is bad at. Why is it all ones? I mean, we can copy this. I don't know why that's all ones, but all right, so now like this is going to be slow, and then I guess we'll first make it slow, and then we'll do it fast. All right, so this—this is uh Norm. They can only be one. It's got to be like a softmax here. This code is [Music] buggy. Are you agree? There has to be a softmax, right? Oh, I guess that kind of is the softmax. Okay, that's fine. Uh, so this—the generalized version of this is some Lambda XA XA One landing x sub one 4X in top. Uh, right. So this one just implied it was two, but it doesn't actually have to be two. Just go crazy here. Stick that there. All right. Um, now we want to run let. Okay, okay, this is doable. Uh, self. Uh, so cell. I think maybe this works. See, this is again—this gets into like the indexing stuff. We're pretty weak. This is a—if someone can make this stuff work—like make this indexing stuff work well—it's a lot better. XLA needs to spend 10 years transposing tensors before upload. You see—you think this shit's bad in TinyGrad? Try to do it somewhere else. Slut sell you close. Um, let's run that. Why did that all just lowercase all my tensors? Don't do that. Okay, x times self.gate subcell x uh self.down subcell self.uppr shit. Uh, and then we're going to want to average it according to that. Okay, so uh always. Okay, so let's—um, probs divid equals probs.sum uh alls times probs. Is that going to broadcast how I want? Probably not, uh, but whatever. Times probs.x is zero, right? Something like that. Do right. So you see we're doing the sort on the CPU, and that's slow, so we'll have to fix that. Um, okay. Oh, of course, it's self. Oh, tensor objects is not callable. Uh, yeah, that's fine. Um, um, let's just like print that.shape and uh also print x.shape. We have to like map all them or something. [Music] Yeah, can I just map all it? Is it right? What's in linear? Have to transpose something. Sure. Looks like I could just—that's my inner dimension of the matrix, so instead of calling it, I just do that. [Music] Right. Okay. Why can I not dot those? Oh, I guess because it wants the inner dimension there. Okay, fine. Um, where is linear? Yeah, okay. I do a transpose there. Fine. Crap. I don't know why it's always done like this, but something about like how you put it—RAM was a good error message. Thanks, TinyGrad. Thanks for not giving me a terrible error message. All right, well, it's not right either. Why not? If you see what I want to do—right—is it not even do—like I just want to—whatever—we'll do with batching later. I think what I want to do is just literally multiply them and sum them on axis. [Music] To—let split this out a little bit—let be less clever. Okay, up channel cell. We can literally just multiply it by that, and then we can sum it on axis, too. Beauty of TinyGrad. That's plenty fast. It's a lot easier to think about. I'll sell you that one, and then—right—so that times that is going to broadcast on those dimensions, which is exactly what we want. Then for this—say x.down—that to x.down. That's probably not going to broadcast how we want. Reshape. Okay, so this should be self.activatedExperts 1 1 1. Um, so this one's going to be different. Actually, fast mixture of experts. All right, we're not trying to dot shit, boy. That one should be doable though. Can we dot those? Is that GNA dot? It should die, right? All right. Uh, fine, we can't dot them. Fine, we'll multiply them. I love TinyGrad. This is so easy. Dot unsqueeze one. Do sum two. Pretty sure that stuff's right. Let grock. I don't want to wait for this. All right, good. We're not dotting anything anymore. We're done with dot. Self.activatedExperts.one on x zero, right? Shape. Stop printing all that stuff. Why you printing all that stuff? Okay, something great. I think it's running. It's got to be the slowest LLM I've ever seen, but it's running. So this is another—this select is going to be hard to write. Who thinks it—who thinks it works? What's up PR? Okay, you think that's a good token? Okay, there's like 17 things we have to do to really make this FL. Why the jit's disabled? Because this can't be jitted. The loading all seems okay. It didn't give me a NaN either, which is good. Uh, so much shit here. Do need variable. Um, ar.temperature. Try zero. Or.count 10. Uh, let's figure out if this is the actual tokenizer we [Music] want. Be deleted. Okay, where's the tokenizer? [Music] It's made a JSON. Which tokenizer does it use? What am I going to load a J? Can this—can this load a JSON tokenizer? Tokenizer, not a string. Oh, this looks like a string to me. What you want? Doix cannot parse model proto. Okay, so that's fine. How do I load a JSON tokenizer? No, no, no, no, no, no. Llama—that Llama—the the Llama do have a JSON tokenizer, or is a model [Music] tokenizer? Tokenizer CH token. Here we go. All right, the appropriate format for pre-trained tokenizer. Fast the hell is this? Depends on the tokenizer library, of course. Not that—that would be too easy, and that would be too simple and convenient. Transformers that work. That's going to import PyTorch. Could not instantiate the backend tokenizer. You need to have SentencePiece installed. I installed SentencePiece. Okay, how do I load a [Music] JSON? God, I doesn't even use computers—a JSON. Okay, cannot parse model proto. Okay, can SentencePieceProcessor load a JSON? No. AutoTokenizer from pre-train. No. Okay, okay. We need to write custom code. Okay, I understand. Fine, we'll write custom code. Oh, it's unnecessary since you can import a huge massive ass library that'll do it for you. Okay, I understand. Yeah, yeah, yeah, yeah, yeah. It's totally reasonable. It's totally reasonable. Chas on object must be read text. Happy. Great. That looks good. [Music] Okay, delicious. Want's an apple? Who's honking? That doesn't sound good. Wasn't me. For this—looks difficult to use. Fine. What if I do that? Are you happy? All there you go. Using—I'm importing a library. I give in. I give—reporting the library is that it takes forever to load. Going to be faster? Slow. Let's shame it for being slow by putting it in a with timing block. No, no, no, but this is a fast tokenizer. See? 950 milliseconds. Who has time for this kind of [Music] shit? Then how do I even use the autoTokenizer? I use this. Am I out of grocs? Uh, what's BI? One sounds good. One good. We happy with one. It's extra slow right now 'cause I actually told it to do something. Oh, look how long it took to get me the word padding. That can't stay three seconds. This is why you can't use libraries; they take forever. Who has time for this? Time's only currency, right? All right, what token do we get next? Let's see our tokens. [Music] Oh, good. It gets faster. It's almost too fast. What was it doing? The select—so fast. Okay, well, that's what we got. I don't know what's zero. Maybe zero is better than padding. Padding is what upset it. Let's see how long that takes to decode like this. Look [Music] at—oh, IP address. Okay, well, that's great. Um, where's that token JSON? Can we read it in Hugging Face maybe? All right. [Music] Padding, spaces, phone number, end of text, Bosnia. [Music] [Music] Just using the KV cache; that's all good. Okay, probably have bugs somewhere. Where could the bugs be? Everywhere. Let's ask grock. No. All right, we'll have to go to stup grock it. Add a proper PyTorch module structure. That's good. Oh, yeah, it would be nice if we had a top K. Let's get our top K to work. Uh, well, I bet you that's wrong. Getting pretty sick. Uh, bandwidths—that looks good actually. Those look good. Routing to one of the experts. Normalize some across that axis. That's sick though. It's probably wrong. It can't be—the shapes match. The shapes match. How could it be wrong? That's like when a type checks in Haskell. That gets fast, too. The only thing that it's not doing there—I get item kernel. That works. It's crazy. Okay. Uh, why—why am I getting junk? Maybe is not supposed to be the first token. It's the only thing I can think of, unless this bugs—then I can think of a lot of things. All right, how do I get the first token from AutoTokenizer? Okay, tokenizer.bosToken. Try to BOS token. That is what I do in Mix. Right? Mix. Can someone write good—if it doesn't have a BOS token. That's why I used Gro2. Why does it take 13 seconds? String object could not be interpreted as an integer. Oh, BOS token ID. Oh, of course. Why does it take so long to import that? Someone needs to write TinyTokenizer. Oh, end of text. What? That's the BOS. What does BOS even mean? Go. I don't—where is end of text? It's like there. I mean, it looks stable. It's not—and it's not zeroing, so that's pretty good. Quasi seems quasi [Music] plausible. Com that out. It loads the weights in milliseconds, but it takes four seconds to import AutoTokenizer from Transformers on an M3. [Music] Hello in B—in bu. Okay, well, yeah, yeah. It's probably not right. Hello Indi and burn. Hello and being [Music] burn. Um, I forget any—less the model has extra ones. Go and think about it for a while. Did I put those Norms in the right place? Like, I can't really think of where else they go. I guess it could go in front of the thing—output—wait for the final layer. So that's LM head. Yeah, that's right. Oh, okay. Like we already kind of like wrote most of this, so it's all kind of right. Uh, input layer Norm becomes attention Norm. Embed tokens becomes talk embeddings. LM head becomes output, and model Norm becomes Norm. Okay, that's already done. Great. We're stacking the experts. That all looks correct. What else can to be doing wrong? Oh, I think the norm thing. Oh, oops. Korm typo. That could explain it. See if the AI finds it. This thing's very talkative. Ah, the Q Norm is there. We go. Hello and be [Music] burn. I think I'll—hello in—in—in—in—in the same—at—at—at Endor. Little temperature, maybe. K—probably fine if it works in—not jitting anything. Well, the temperature did nothing. I guess that's a pretty low temperature. Let's like just print before you go to a max and stuff. Okay, the first one looks very different. E—it looks like all equal probable for the first one is probably wrong. The thing about that code is like I know it works. Oh, deal with that—so it—without the KV cache—if that makes any difference, but because it shouldn't—because that code works, but this stuff all should work in BF16 on M3, no problem. Um, rel loading in the experts. Oh, well, that's annoying. Does anyone support that? Like, okay, what we really want to do is just transpose those two and then dot them. That should work, right? Oh, maybe we do it the other way. Let me read the code for the linear. Okay, so—so we'll try X.dot uh this uh transpose permute uh Z one to that. Oh, I having backwards or something. [Music] Uh, okay, so there—yeah, okay. That actually shouldn't be the size. Not waiting for—that's stupid. Token loser every time. Okay, so here relax.doDown, which is going to be do dot this. Should mostly be right. Okay. [Music] Um, yeah, okay. That's fine. The problem here is this uh one one I had originally. All right, it's going to be fine for the first layer; it's going to glitch out of the second layer, uh, because we're going to have to do—ex—that shape—is that the same math as that? I just don't really know. Yeah, it looks the same. I don't know why that's not—oh, why is it—that looks different. Okay, this stuff looks different. Okay, that's good. Don't think it's better; I think it's stupider, but looks different, at least. Um, yeah, that's just going to be one comma one. That's fine. So we're going to broadcast across both the batch dim and the uh the other dim. Then we're going to sum across the mixture of experts term. Do that. Normal space. That's fine; it's the same thing I did for mix. That could be wrong. There's another potential wrongness. Oh, let's wait forever to import the tokenizer again. Who's excited about that? I see—it's fast as long as I don't do anything with the tokenizer. No, the Apple. Hello in O—in O—in O. Okay, useless for. Okay, like something's broken if those are different. I hope that actually works. Looks pretty good to me. Okay, we're picking eight experts. Um, I mean, I'm pretty sure that's how you combine them. Oh, oh. Why don't we just import TensorFlow? Oh, no wonder Transformers is token so slow. Where could the bug be? Let's go back to the paper. Okay, we [Music] perform parametric RMS layer despite reducing C wow for—for. Okay, what else did I do wrong? V balancing loss for—know what looks right. We—the good tests. I'm going to just say that stuff's right for. Okay. I—what could be silently wrong? KV heads. Rope Theta. Our Rope Theta. Does this even use rope? That could be wrong. Let—if Llama-specific assumption. Rope. Rope th is 10,000. Okay, that's right. Attention heads. 16 layers. Every—what's DMoE? Decentralized mixture of experts. Oh, it sounds like crypto. Do they have a token? Have they considered a token? I would really trust this project more if it had a token. Okay, the only thing that it could potentially be is that there's an NKV heads thing or just some other like bug that I just don't understand. It's strange that those switch. Okay, the activation there is SquigGlo, so I'm using SquigGlo. This is SquigGlo or not. Silu is SquigGlo. Oh, no. What if SquigGlue is not Silu? Do I have SquigGlo? What is this? What is this? Are you going to build a Squig? You're going—oh, going—going camping in the snow. We're going to build a SquigGlo. Uh, what else uses Squig? Is it—could it be Squig? The hell is Squig? Does Llama use SquigGlo? Swish and GLU. Every new LLM uses it. Llama 2 uses that. Is it Silu? Is it the same? Someone has to know this, right? For—where's my SquigGlo? It's in here. No, it's not. It's in here

the normal boring shit that doesn't seem weird. Router prob per expert. See a long confusing routing weights. Okay, know any of that? Auil load balancing, what stuff shouldn't have changed? Okay, definitely gate prob is there, so that change is definitely right. Just like quickly, like print K here. This isn't wrong, right? Yeah, down P, up P, gate P, model State, dick replace. Okay, right. I mean, like if I comment this out, it should just behave really stupidly, right? No, it looks smart actually. Well, sort of. Never mind, it looks stupid. Okay, so it's definitely doing something. Um, and these are all like that split would fail if that just wasn't right. So all right, MLP gate is feed-forward gate. The gate's not one of the experts, so that's being set correctly. Here's the Norms, clip qkv, they're qor and Norman States. Oh, which is what we do to yeah, I don't think this can really be wrong. I want to put it before the reshape. Right? Yeah, you put it before the reshape. Okay, so there's the there's their reshape. Um, then you apply the shit whatever. K Norm applies to k, q Norm applies to q. That was a bug we had. I doubt it's that clip. We don't have biases. Oh, num key value heads, that is the same, right? Num key value heads is none. All right, that's fine. Um, okay, what is Norm topk prop? Whether to normalize the top K probabilities. Interesting. So you actually maybe don't do that again if it's really that oh, it behaves almost daily. I can't be that same crap. What could it not be? RMS Norm? Those are RMS Norm. Uh, interesting, so that's zero. Okay, so we'll get rid of that. Norm mean it's yeah, for I would never write this like this. That's going to ignore the top ones that prude. I'm pretty sure is correct. Pointwise multiply gate by up, go down by down, that goes there, zero. So we're not even getting into variable stuff. Temperature is zero, so that's not going to do anything. Dot item, it's not the problem. Okay, post-attention layer Norm is FFN Norm. I think we checked all that stuff. All right, who knows? Who can find the bug? What strange shit changed in this model call? On what? At that, see if that doesn't look stupid. Look fine. Why is this wrong? What was the bugs in the mixol? What did I have wrong in the mixol last time I tried to do it on the mixol? I apply the experts here and move everything around on devices. I manually divide those both by the norm. This terrible code. What if I say let's start with like toes equals z? How different does it look? Strange, really that different, or did I just break something? Interesting, almost working. Who can find the bug before me? How come if I do that? That's crazy that we get that. Like it can't be zeros, can it? I guess there's no biases in any of this. No, that still shouldn't do anything. That doesn't make sense, even if those are all zeros. Well, that just looks like numerical instability at that point. I mean, I have a whole bunch of questions. Why is the token embedding for zero zero? That doesn't make sense. Oh, that looks totally wrong. Why is there so with zeros in there? Or oh, maybe those are just fake tokens. Like it can't actually be zero. Never mind, that might be fine. And then yeah, like it makes sense if you put zeros in, you're going to get zeros out, which is like normal I guess. And that's just numerical shit. Is a silu have a fixed? What's silu? Zero, so z z. Okay, fine. This is this is all just like symmetric. So okay, if you give it a stupid token, it gives you a stupid token back. Is this going to give me buyer again? See, that's this is what's cool to play with though. You learn things like that. That just seems more balanced. Like I'm not getting an outside token there. 309 I, that's a very reasonable letter. I didn't fix anything, did I? I guess I did the soft Max, but I doubt that's anything. Oh, I fixed the K norms. We're going much more balanced. These look it's a little less balanced, but once we're in eye territory, it's thinking about itself. So come on, you're a reasonable token. It's getting slower. I period seems unlikely. Does the Llama have to put a boss ID first? I have so little experience actually like using these things, you know? Yeah, okay, so it does it does put the boss ID first. So this actually like a plausible thing to believe that like you have to actually put the boss ID first. That is what Mixel does, right? Oh, this isn't even the boss ID. This is some other stupid ID. It's called padding. Oh, okay, so like maybe this is literally just that. I hope it's that. Why is this so slow now? Did I like break something? Seems slower than ever. So this not even the b i like slow metal land for some reason. Sometime metal just gets slow. All right, we're putting in the bar ID, railing out to that expert again. Is that buyer? Stupid buyer. All right, who's like who knows this stuff? Um, what tokens am I supposed to put into this? Co no, what token goes first into an llm? Boss, beginning of sentence token. Why is that called end of text? But whatever, it should not be giving me that though. It does matter what I put in, which is interesting. There's no way like something could be transposed that I'm not doing right. I know like the GPT crap does like transposed stuff. I guess it could be missing a layer. Bed tokens way, that's right. What about the output now? Put that way, it's right. Can I see if any Rands are actually triggering? That would let me know that I'm missing something. Um, what am I missing? Freaks sis, it's probably okay. Yeah, I'm fine. Oh, maybe one of the reasons it's so slow now is because we moved everything to stupid floats. We don't actually need floats. That stuff can be keep that in bf16. Get the Ram faster. This copies makes no difference. I was really hoping I was on to something with the BS. Still print Logics. Print logs. Print the Logics. We don't print the Logics. Wow, it's way faster if we don't convert that to. Could this be right? Burn a former US. I mean, this seems coherent at least. Just loves burn. It's also fast if I don't make that float. Get rid of the other to float. Oh, I know why it's just slow cuz it's that's a lot of ram. It's killing my Ram. Oh, it's probably it was probably swapping to disc. No wonder that was slow. Okay, fine, fair. I won't ever make you go to disc again. I'm sorry. Gotta load the tokenizer. Burn a former US. Enter enter enter enter ENT enter enter enter. Uh, I don't know about that. Still some bug, but all right, I'm glad we're fixing the speed up. Where's that other float? Get rid of that one more floating only sinking. Uh, the problem with float is it was making the uh the thing really big, especially because I'm not deleting. I should delete these tensors when I'm done with them. Fros burn a former US. Enter enter enter enter. All right, um, for this one, am I giving it a temperature? Yeah, 0.7 is the default temperature, so we can try that too. I don't think I think it should still work without that, but you never really know. Delete print compile all the layers. Compile compile compile. Yeah, seems unlikely. Okay, importance of self-driving education. I mean, this kind of looks coherent. It's not a large model. No, okay, it's not this stupid. There's some bug somewhere still. I'm GNA guess BOS token ID is fine if it's instruct. It's even more important. Oh, oh, you're right, it's an instruct model. Oh, why don't we download the non-instructed end of the page for how to. Oh, that's the non-instruct model. Let's not use. Why did I downlo the instruct model? Can I just get rid of the word instruct? That work. Take forever to download. They hav an improved one. Wow, they popular model. All we're downloading this model. Bitcoin is we need also fast ways to just like quickly test it on one of these things and see if we're like that's the only way to really test if you're running these things correctly is going to be to actually put it in one of these uh just like one of these tests. Oh, and this is also a different Arc. Don't Norm them. Hiden activation is Silu. That's fine. Which what we're doing. Not doing any Dropout. That's my rope Theta. I don't really know what rope Theta is, but that is what it is. Uh, router aux lost coefficient. That's only for training, right? Number of key value heads is the same as attention heads, so that's just nothing. Number of experts per token is eight. Max position embeddings again fine. Initializer don't care about that. EOS token ID is that which is what we use. Yes. Um, so like yeah, there aren't there are like special tokens for these things. Does the instruct model have a different uh wait huh? Don't trust that. Yeah, I shouldn't do the instruct models. That's a good point. Leverage hugging phase to get the inputs. Yeah, well, yeah, I guess I could I guess I could just run through one token at a time. I don't have to sample from it. Still going to take 10 minutes to download. Sorry, my internet's slow. What's for dinner? Fishes. I love fish. There's this sample. Wow, M Moe girl. How good is how how good is that stuff? Now it's been a while since I've played with that. I remember being really excited about at the beginning, and I think a lot of people were, but then you realized it's just not that good. Talk about this a little while it downloads. I've read this before. So one of the downsides to Nvidia, one of the nice things about AMD is AMD documents they're uh documents their real instruction set. Nvidia does not document their real instruction set. Um, they have something called SAS which changes from generation to generation, and then PTX is transformed into it. Well, chat's quiet because when chat says some shit, they get banned every time you talk in life, it's just like you could get banned this time. Um, should we let let we let nonsubscribers talk? We'll let nonsubscribers talk while while the model downloads. You're not free, you're cheap because you don't subscribe. Who thinks it's just going to work? I'm going to regret it soon enough. What am I going to regret? To see chat format. What am I meeting an amhq? What am I going to do there? Why subscribe when there are ad bloggers? Add special tokens. Add generation prompt. What's the generation prompt? Indicate the start of a response. Subscribing is like paying for Windows. You see, I would say that, but you're subscribing to me, and that's the difference. You can pay Microsoft money. You can give you can give Sati and Nadella more money, or you can give me $2.50. Uh, no. Hey, hey, look, look, look, look, you. I used to tell nonsubscribers that they were bad people. I've I've come around on that. I've mellowed out a lot. You I moved to Hong Kong. Life's just pretty good. So you know, we don't got to we don't got to bash the this isn't some hyper-capitalist America. We got to bash the nonsubscribers. We just have to gently remind people that you know, social cohesion only works when we are all socially cohered, and if we let people drive a wedge in between us like not subscribing, and then we won't just get to talk, and it'll just be it'll just be less good because we all need to work together to make the society. It cost a dollar to subscribe in Russia. Uh, when I moved to Hong Kong. Uh, well, like as Hong Kong immigration will remind you, I do not live here. I am visiting. Uh, been here since like last August. I don't know if uh yeah, if you if you missed the beginning of the scream, I told you about my experience being interviewed by Hong Kong immigration, which was surprisingly reasonable. Um, you know, the CCP you people, it's actually it's it's so weird. Like I think the average American thinks about politics so much more than the average person in China or its periphery countries. Like like Americans' weird obsession with this shit. Like whose president is going to affect what you're doing tomorrow? Like it's not. Um, no, my internet could be faster. Yeah, if he starts World War II, as was in China as sitting at this Pub, and these British guys are talking. Uh, and uh they're talking about yeah, World War II and all the nukes and all that stuff. Like I don't know. I maybe maybe that's the vibe more in in Mainland, but I don't think that's the vibe. That's definitely not the vibe here. Uh, like you know, who who's nuke and Hong Kong? Nobody. It's just so let the let the people with the you know, I keep saying you know, I gotta stop saying you know, no reason to think about something people can't change. That's God grant me the serenity to accept the things I cannot change, the courage to change the things I can, and the wisdom to know the difference. There's a reason my phone's a blocked. Her apparently hasn't been the same. What's YH? Oh, oh, LKF, like the like the the bar Street. How I'm I'm 35. You think you think I'm going out like binge drinking and listening to like let's make a night you can remember, be the night you won't forget. You think like I'm I'm gonna like dance around to that kind of music and and uh yeah, no, that's uh maybe when I maybe 15 years ago. Uh, no, I saw I saw the clubs in uh Coco Park and in Shenzhen, and I was like, yeah, I'm too old for this. Wait, so it's like the same with that model. I mean, I don't know. Maybe that's just right. 29 is that a good token? A greater than sign? I don't think that's right. Maybe actually what is this? I think it's actually working. It's just like this is what it generates. I don't know. No, I'm sure there's bugs. I'm sure there's bugs that subtly degrade the performance. Must be real number, not n type. So this doesn't have a boss token ID, or this isn't the right tokenizer or something. None. Okay, okay. What's going on here? The current process just got forked. STP fork in my process and get kicked out of pon. Like it's still it loves that 30 on the third layer. It's got this got small model spell. No, it can't be this stupid. Three is that a good token? Hello quote. No, this isn't right. There's no way this thing's this dumb. There just B the first line of the file is the name of the file. I don't know why there's not a boss token ID anymore though because I went away from the instruct bottle. Uh, okay, let's fine. We'll download the torch one. Can even try torch with the tiny gr back end. Let's see what we got. Instruct not inst not instruct. Can anyone start work on jit Mo? Well, that's kind of what I'm doing. Great. We're going to have to redownload it somewhere else. Uh, did AMD? No, AMD sent boxes legit. What else could be wrong? Maybe the ll.cpp one is simple. For I just feel like that's wrong. It doesn't happen with all tokens. See like that token, it doesn't happen with. And if I pass in zero, it doesn't do anything at all. So you're thinking in tiny gr, you can always run with debug equals two, and then that'll show you the uh like how what it's actually doing. So we can look and think if like it's doing anything stupid. Fine, there's your qkv thing. There's your attention. Where's the mixture of experts? Well, actually, that has to happen up top here. Get item here. Be interesting. Let's see what it's actually doing. So we can look like these are the those things. I it's like probably right. Wait, what? What's calling numpy? Have a lot of questions about that. Like nobody has to realize those inner ones, then it gets the order for that to work. Those are all the slow ones are for just for sanity. That looks fine. 28 GPU servers. Or some crap broke. Why is this? What does this tied weight Keys mean? Like I feel like it's something like that that are potentially tied to another in state deck. Doesn't matter. This should affect yeah, I made it to the second file. Okay, well, that's good. By the way, if you want to get good at machine learning, I always say this. This is how you do it. Like you're like, oh, well, you're actually just sitting there, and you're just debugging stuff, and it's stupid, and there's one bug, and they someone just show you the bug. Yeah, but then you miss the whole point by sitting there and struggling and trying to find these bugs. You're reading all of the stuff. You're learning. I'm learning all the new tricks. It's been a while since I've hav one of these, so I'm like I'm learning all the all the tricks of what the latest stuff is. The norm is there. Is that the same? You know, we've also never really tested our uh like our stuff. Period. We've never tested our llm implementations to see if they're even I mean, they're not totally wrong. They print reasonable stuff, but that might not be enough. The FFN norm and we have feed-forward, and then we add before the norm to that or continuous. There figure excludes in details. Al uses qk Norm. Okay, see now we know what qk Norm is, and we did that. I mean, probably what you want to do to test them, and I'm not really doing that. We can go through. We'll just we're going to start comparing the layers to torch in a minute because I don't know what's going on uh on a scaler. Oh, really? Score one against torches thing. I don't know why it like doesn't work. Um, yeah, that's fine. F it's not wrong. Fine. I could have done this somewhere f with a faster. I could just test this somewhere that has fast internet. Okay, I have my suspicions around those tied weight keys. Ed tokens and talking about dinner's ready. We're streaming doesn't work yet. She wants I think we got to go to dinner and then come back, and then we'll find the B, and then we'll make it fast. That's all good. We got to go to dinner. I'll see you all in about all right. I got fish to go eat. Uh, I'll be back in an hour, and then we will debug this layer by layer, and then we'll make it fast. Uh, see bet. Should we leave leave it on and just just put a blackout screen? I don't know how how you usually do this on Twitch. We're not going to find this in a minute. This is going to be a slow methodical project. Why nobody on chat? Oh, wow, wow. Chat was lagging. Put a BRB screen. Okay, okay. Let's see how quickly this is going to load. All right, we'll eat quickly, and then we'll come back. It's still loading something. Oh, Jesus, it's crashing my computer. It's making copies. Uh, where's the example? We won't go to NPS on this whatever. Making that lag. Can we have waiting room music? Oh, I know. We'll put on seven hours of the Jeopardy song. Okay, is that deterministic? What do I say? Temperature. Okay, so that's clearly we know what we're supposed to get now, so we can just put this in. I'm try to find torches equivalent of that and print it. All right, you get nothing, but I'm going to eat fast. E e e e e e e e. I forgot I muted it. All right, I gotta go to dinner. We'll be right back. E e e. Okay, there was a negotiation. We got dinner being brought to us. It's the best of both world El. Okay, it looks almost entirely correct. Um, let's just check my sampler like it has like all this crazy stuff. Uh, what's the loss free? Thank you. We go out later. Yeah, when I done. Yeah, we at dinner. Look at dinner fish. This fish was recently alive. Why does this one look different? H. Well, when I divided by the some of the probs, it broke everything. I'm going to guess is something related to this. We're really close though. They dynamically change the bias to select the experts. What does that mean? You mean like this whatever this crap is? Do we have to do this? Yeah, it repeats. Oh, I mean, maybe oh, you know what it is? Probably actually, I bet their model generate has some improved sampler crap. Uh, reads from a Json file. All right, well, we definitely don't have a problem our sampler. That can be deleted. Um, a temperature. See what we got for it. Selects the most likely at each step. Okay, maybe it's not this. Yeah, no, there's still some bug in this. It's almost the same, but it's not exact. We're probably routing to the wrong experts. All right, maido, you're the reason it's going sub only. Stay on topic, please. Why is this not working? It's close. The other one is hello. I'm an newbie to this form too. I think we have to understand this switch Transformer thing. Compute load balancing loss function. Like what is that? Is it masking some of them out? Encourage a more even distribution. No, this is not a real okay. I don't have to do anything for this at all. Yeah, this is not I don't have to deal with this. It's just top K. I'm doing soft Max. It's the same as the Llama. Rotary embedding. Did I forget a norm? Do I need clip qkv? What happens if I comment out the Norms? So you see the problem like it's almost correct, but it's repeating. I am a newbie to this form, which means there's something slightly subtly wrong with it. Well, that breaks entirely, so it's not that wrong. Could it be KV cash? I mean, I guess we can run and Float. It's just really slow. I doubt that. I've seen no fixes at all from changing D types. And I think X is not even like I doubt it's accumulating in B flat. All right, where might a subtle bug be hiding? For I don't even know how that works. Should be that. All right, we have the first set of experts. Let's see what torch says they are. Let's see if torch says they're the same thing. Freezes. C routers. Yes, do that. I don't know for for for. Do they match? Yeah, identical. Oh, but okay, we have to be careful because it's possible that this isn't actually what's getting routed to the things. I wrote that logic myself, and if that logic is what contains the bug. Okay, that one all looks correct. See the next one. Oh, the next one's not right. 863 I get 68. Totally different probabilities. Okay, the first one matches exactly. The second one no longer matches. Is this some issuing the attention? Like KV cach has to be wrong. So we're going to totally different experts here. There's some similarities in them, but they're not the same, and the probabilities are definitely not the same. Okay, so it no longer matches once we get to even two. So let's just do max length three. Um, we can even just print for for. So it's either a KV cash issue or Norm issue. For I having a my uh he's got removed. I remov that. Okay, you guys see what should be the same there. So that's 8636. So the values must be different. Um, let's figure out what's different. Attentions. What are hidden States? So is that like the X that's being passed in here? Like PR this x got the same thing. Those are very different. Pass the MLP getting F in the MLP there. See if those are the same. Yeah, okay, cool. 41 41 all the way down to it looks right. Totally different once you get there. Yeah, it's minus two. That's minus 7. Okay, totally different. So it's got to be the attention. Well, or it could be this for the iner norm. So it depends. We'll see if this one's the same or not. Okay, that one's the same, but that one isn't. Okay, it's just a problem with the attention. But how this is the same attention I use for everything? What could it be? Is it the Norms? Well, okay, let's debug this. I we can debug this exactly. Um, it's crazy though that's easy. Kind of what the hell? SDPA attention scale do product attention. Okay, that's fine. That's what I'm doing. That shit. It's not clip qkv, is it? Rotary shit where key value cash here for it's kind of it's annoying because that's after the bullshit. Um, we'll do that. It's not clip qkv, right? Okay, it's not clip Q. There are my query states. Do they match or no? Particularly that one. There's like so many more methodical ways to do this that I should be doing, but no. Okay. So this one matches, and the second one doesn't. Okay, so that's not the problem. Where is the cashier? Where's the key value cache for? No data is available if shape is symbolic. Well, okay, that's annoying. Uh, for I don't understand how the key value cache is anything besides the absolutely normal key value cache though, which it works fine in all the other models. The only difference is that Norm and that Norm should be applied long before the key value cache. I guess we don't jit it, which is different. So there could potentially be a bug with the like that should work fine if it's bound. Wait, why isn't that taking in the sequence length? Okay, I don't know what could this be? Ty FL for for now we have two of them here. This is the key. I mean, clearly, it's added a new entry here to the cache, but that entry doesn't match torches. So this one you can see matches this one now. As soon as we go to the next one, this first one, which is the key from last time, matches, but the second one that's what that has to be doesn't match, which doesn't make a ton of sense. You assign the key and the value. Okay, is the problem here? Is the problem in apply rotary embedding? That should go right through. This is if it's the Rope. If it's the Rope. Okay, one before the one before is right. Now the one after. No, it's wrong. Right there is wrong. See that is wrong. Where am I putting that? Let's move it above the rotary embeddings. Has this just been right before we uh jump to conclusions. Let's check here. Doesn't look like it's the key value cach that's the problem. It's probably just some one tiny little word I didn't read in the paper. Well, actually, we switched the cosine and the sign because when you see when you put signs and then cosine. Don't try to do these both at once. Clearly, there's the one before, there's the one after. Wait, that matches and that matches. Huh. Okay, it's got to be the rotary embedding for huh. Where did that come from? For this looks like rug. Why is it that oh, maybe because that's the first one. Okay, that's fine. But then it's wrong for the second one. That's got to be wrong. I don't really know why it's wrong. Oh, no. Flip shit. This is frustrating for anyone to just watch. Favorite VSS code plugins. Oh, oh, oh, the the advanced rope type. Oh, yes, yes, of course. No, they look the same to me. This better be a good bug. This better be a good bug for how long it took to track down. I better learn and gain a lot of insight from this. All right, there's some bullshit numbers. They don't fucking match. Okay, well, at least there that's that that's that and that one matches that one, but look, that one doesn't match that one, even though the inputs to the Rope match. So why is it not getting the Rope correctly? I was I was like says it's the same as llamas for. Could this have just been wrong the whole time in tiny grad? Why are they different? They're clearly not the same inputs up the same. If feel like way different. So is that really the same as this for for for. That's sh. Has this been wrong the whole time? Sh. Okay, let's just try to write the other one. The other one even looks better. I don't complex malt garbage. It says this is copy from llama. Is it really? Looks identical. It's shorter code anyway. I hate that complex mold garbage. 8363 8636. Well, well, that's is not that yeah. What is supposed to be multiplied here? Wait, huh? Why is that repeated? That yeah, this doesn't really make sense. I feel like this has just kind of been wrong the whole time. Look at them. It runs twice. For for for. It's been wrong the whole time probably. 8636 is that the magic numbers? 863 629. Like I'm on fucking lost or something. You know you yeah, the numbers 8 23 6 29. Fu. That's probably been wrong the entire time. Look at that code. It looks like dog shit. I never want to see complex multiplication. All right, let's delete all these prints and get out of here. Jesus. That's probably been wrong the entire time. Does anyone even use tiny grad? No. We have to get much more serious about testing our llms. Like I see what the bug was too. It's interesting how little of a difference it makes. Let's just make that number bigger again. I mean, like if that's a real bug in llama. No, that's actually super worthwhile. That was so worth my time. That was so worth my time to track down, and I have to really understand. You're right. I have no understanding of what rope is doing, and that's what that just showed to me. Guys know my my one of my favorite lines ever, you know, never give up, never surrender. Um, I am a new be to this form, and I am looking for a cup people to talk about my new website. All right. Um, but I know Transformers. Does it match torch now? Damn, those libraries are easy to use. I see why people use them. It doesn't match closer, but it doesn't match. Wait, 6863. No, wasn't it the other way around? Yeah, we want 8636. Oh, this just has a the bug in it still. Did I delete it? I thought I got it. What? Oh, did some of that printing change something? That's super frustrating if that's true. 863 sex we want. See, 6863 is not what we want. It's so wrong. Okay, that's right, but it definitely said 8636, right? Yes, right there. 8636. Well, that's a whole another issue that turns out to be related. Now we are already doing that. What? 6863. No, there was one tiny gr that was correct. Now what changed? This shit is so troubling. Like what's going on here? All right, let's see if we can go back to it. Should I just copy and paste that? Wrong. I might have just copy and pasted that wrong. 8636. I think that one was right. Jesus. E 63. Try to get a better understanding of the different types of data that can be stored in. Okay, there we go. Matches geez. Well, the good news is it was not any tiny grad bugs. Should we do a whole stream tomorrow as punishment where we just go over rope over and over again because I really legitimately have no idea what rope is. If you get this shit subtly wrong, I think it's wrong in all of our llms. I think all of the tiny Grand llms have this one thing wrong that I wrote the first time that nobody checked because there's no real way to test it. It almost works, but you know what is this? Horseshoes or hand grenades? No, it's llms, and if you get it kind of wrong, you know what you get? Just just SL L stupider llm. Jesus, what a what a waste of a day, but not really. How many like libraries have this wrong? How are people validating this stuff? It had nothing to do with any of my Almo stuff. This is just the first time. I don't know. This model must be really sensitive to it because it's mixture of experts or something, and like the llamas can just take a lot more abuse. Don't abuse your llama. Um, yep. Okay, but I would like to point out to everybody that this matches this exactly, and you should never settle for anything besides exact. You should not say, oh, the words look kind of right. Um, great. You know, I kind of want to do tiny gr. apps, and I want to write like a super high quality. Did I commit this? So I want to do like a few super high quality llms in there. We can make this thing fly too. Got to make this thing fly. Get this all in the jet. All right, non subs, thank you for being here with me through this. You see, you know what you want to know something about greatness? There's got to be something broken in your head. There's got to be something broke in your head where you're just like I couldn't even I couldn't even eat dinner, you know? I couldn't even go downstairs for dinner like like a normal person cuz it wasn't right. There was some bug. Uh, and yeah, you just kind of got to be obsessed with that. There were times when I was going to give up. It was an arduous journey and adventure and turned out the whole time my rope was wrong. Notice how grock

Such that this is correct for Hugging Face, but not correct for... Do I do anything like that when I like convert from Hugging Face? What is this shit? Uh, yeah, okay. They're probably not wrong, and there was code in TinyGrad to fix them the whole time. How does that work? I'm sure someone else wasted tons and tons of time on this. Uh, no, it's not any of these. So like, I already like wrote this.

Oh, first, let's see if we don't need that for for [Music] [Music] [Music] No, I did save them. Is this right? Maybe our normal stuff is right, but I have an idea that you can't just do it to Q and K. Oh, no, it's okay. This isn't going to work. Yeah, okay, this isn't going to work because uh, you also have to permute the norms. I'm not dealing with how to do that. Okay, we got to the bottom of this. There wasn't... Oh my God, what what a stupid voice. I'm not... This is probably fixable by uh, yeah, I'm not sure this was right. Is this code like from somewhere? Is Hugging Face Transformers? Yeah, this is made by Hugging Face. For... Is this having store for training almost stuff is open source? Where's their training code? Transformers implementation is slow. We recommend using VM. All well, that was fun. Uh, we didn't get to work on making Mixture of Experts fast. Uh, I can just... So we complete what's said in the Stream title, let's move it over here and run it on uh AMD with the AM driver. Notice if we do rock msmi, we'll do our stuff. So we'll just pseudo rmod AMD GPU. Oh no, there's no more amdgpu. What we do to it doesn't [Music] matter. Uh, let's just run debug equal 2 to start. Oh, see that? What it means to have an internet connection. You know that's an internet connection. Oh, that was fun. That was so fun. Like Hugging Face stores something permuted. It seems to maybe permute it in such a way that the shape manages to match. Yeah, I think it does. Like usually when you see stuff like this, it's because it's uh mismatched, but really what... Okay, I mean the real thing thing to do here is just to update this function such that this function works. Like I shouldn't have reinvented the wheel. I should have just used it. There's nothing at all wrong with using this except for the fact that it's uh not always right with the Mixture of Experts, but it clearly marks gate down and up. Um, I would assume that those ropes are even probably correct, and all I got to do is flip around the uh I flip around the norms also. Wow. Who wrote this and why weren't there comments? Rearranging the weight matrices. Oh, I think we can do this because we flip both the key and the value. The rotary embeddings and multi-head attention would compute incorrect results. Do I have from metal [Music] somewhere? Why is it trying to use Metal? Oh, up here to Metal. All right, let's fix that. Oh, we're not we're not even... Why did I write this junk? I wrote this junk because I was on stream. I don't blame you. Blame myself. I should have used the library code in TinyGrad. The library code in TinyGrad is extremely high quality. This is just a standard format. It's just the Hugging Face format. [Music] We're just going to need to permute the stupid norm too, then it'll work. It's been a long road getting from there to here. Why is it still trying to put things on Metal? I have the word Metal anywhere here? Two Metal junk. Junk I wrote. Junk code. I wrote junk code. [Music] Why would they flip the rope? Okay, so in conclusion, the rope in TinyGrad was never wrong. Yeah, that was dumb of me to think that that would even be a thing. Um um All right, and here we are running that same thing on AMD. Uh, yeah, look how fast it reads to the GPU. It copies those 10 gigabytes a second. Yeah, we shouldn't have written this garbage. It should have just been convert from Hugging Face. We do have to deal with experts in a special way, but otherwise it's a standard Hugging Face model. The CPU in this is worse than the CPU on my Mac. That's why this is so slow, but you can look at the model time. It's not bad. H, we ran out of RAM. It's like not being freed. I wonder why that is. Is it freed on Mac? It has to be. No, it doesn't. Okay, and we have some memory leak too, to top it all off. There's a memory leak. I don't know. I'd be a lot less frustrated by all this stuff if I just like [Music] Hugging Face Llama interleaves the rotary embedding dimensions: dim one, dim two, dim three, dim one, dim three. And then let's look at scaled group attention um and confirm. Where's scaled group attention? So what that's going to do is like you have a different key value permutation there. It's all in the cache like that. That's all fine um when you're in [Music] key value, you multiply the key and the query matrix together with the inner dimensions. So all this stuff cancels out. Yep, so it doesn't matter. You can permute the qu the query in the key as long as you permute the query in the key inner dimensions identically, uh you'll get the exact same output from then on out. The only reason my permute and switching it back didn't work was because I didn't also permute the uh the the other thing. And let's just test that theory. Also, there's a memory leak that's got to be fixed. I don't know why that's happening. Oh, we got to make things up. There's so much to do. There's so much to do, but we're going to make DeepSeek go so fast on those MI30 machines. [Music] Um, there's a few things that need to really uh happen this year. One of them is uh I like how much simpler that one is. We can probably write something that looks like that though for... Oh, we have to give it heads too. [Music] The norm happens afterward. Um okay, wrong. Now I like vaguely even remember this from a bounty. For... You can also just make a flag called Hugging Face style rope. I am a newbie to this forum and I am trying to get a better understanding of the different types of data that can be stored in h. Okay. Uh, someone want to do this? Someone want to clean this up? It's a good it's a good it's a good project. Just like clean this up, integrate it beautifully. It's not a bounty. It can't be a bounty because if it's a bounty, I'm going to get I'm going to get 32 spammers who try to do it. Only do it if you're going to do it beautifully. But uh, yeah, someone do it. Integrate this function uh with that function. I should have looked for this function first. This function's way nicer than what I wrote. Um, this needs a fucking comment. Like whenever you see that, whenever you see permute, whenever you see that kind of thing, you always assume it's because the shape don't match, but permutes where the shapes match are horrifying. Uh, so yeah, that needs a comment to whoever wrote this. I don't know how you figured it out, but you could have left a comment for the next guy. I probably wouldn't have even checked it. Um [Music] Well, yeah, and this is this is a good lesson of where bugs are too. It's not like the bug is really in any one place. This explains it perfectly. Hugging Face Llama interleaves the rotary embedding dimensions differently. The target model expects a group in contiguous pair as the permute function fixes this mismatch. So there was never a bug with Llama. There was never a bug with Llama. There was a bug only in my stuff with the weight loading. TinyGrad was correct. I was wrong, and uh, I don't know. Took a lot of time. Should I be upset with myself? How did I do? And what did I do? Like no, I don't know. I tracked it down. I went through it methodically. Lessons were learned. Lessons were definitely learned. Uh, who did this? I did this one. This one's going to have good comments too because I'm sure there's like 20 people out there working deep in this LLM stuff and they're just kind of like... He fell for the Hugging Face stores the weights permuted differently bug. Um, I did I did I fell for the Hugging Face stores the weights permuted differently bug. Wow. So many of you are... Why we only let subscribers talk. Uh, the AMD one doesn't work because we have a memory leak, and I'm G to fixing that right now. Um okay, so todos. Todos. Somebody clean that cat up. Make a beautiful commit out of it and fix up that Hugging Face function. You can make all that code I wrote really be like just 10 lines. Um, we'll see if anyone does it. If no one does it, I'll do it. Uh, this whole thing could have been a 10-line diff. Uh, two, find the memory leak. Uh, where else does memory leak exist? It clearly doesn't exist in all LLMs. It probably has to do with some indexing stuff. Three, make that actually fast. So I know we have that top-k implementation coming, uh, so that's definitely going to help. And then yeah, make it fast. Get it in the jit. Get the whole MoE in the jit. It should be a lot more doable the way that I'm that my that code's a lot cleaner than the Mixel code. I know when I wrote the Mixel code, people really liked it, um and I know that like there's a lot of bullshit in here. Some I'll clean up tomorrow maybe, um but like this is a much more beautiful uh implementation of this because it supports this. Uh, I should also support the normalized probs as a as a flag um whether you want to normalize probs afterward or not, but this shouldn't go to CPU, and then this should all be jetable. I believe. Yeah, there's no reason it's not jetable. The hardest thing to jit is the KV cache. There was never a bug in our KV cache. There was a bug with with permuting the shit. Um, that can all be deleted. Can someone find out how to not use this tokenizer that takes 3 seconds to import? Um, that's fast now that we're using the new Mixture of Experts thing, which is a lot which is a lot nicer than the mixture one. Right, like the mixture one is doing this, and this is going to be slow, uh but this is not slow. I put the experts in the matrix itself. Um, so yeah, we need to not go to CPU here. We need to speed up uh these indexing functions, and we need to get this all in the jit. Uh, g g could someone write a for loop for that? Um, I don't know. It's not that bad. Whatever. This should be lazy still. It's even more reliable than the torch vatcher. Delete all that. Delete all of this. Um, something like that's going to have to stay, and then that's just a junk loop. But yeah, this can all be like 20 lines of code. Uh, you guys do it tonight. If no one does it tonight, maybe I'll stream again tomorrow. I think I'm going walking tomorrow. Uh, you know, I did this instead of eating dinner downstairs. I ate dinner right here. You guys own me. Uh, thank you for watching my stream. Does anyone have any intelligent comments? You will not be the last PL message. Is this open source? No, it's closed source. Keep it very secret. This is secret. This is high-quality shit, man. You think I'm just gonna give it away to anybody? Chat helped me today. Did it though? Did it? Were you guys talking about the Hugging Face thing? Did you know about Hugging Face permute? Um, you didn't know about Hugging Face permute. Wait, Hatch AI knew. Wait, you guys wrote it in chat. I should have paid attention to chat. You just didn't read check his logs. Yeah. Oh, man. Chat won. Oh, you knew it too. Why would they do that? Why would they do that? All right. Well, I didn't even spend that much time. I didn't spend that much time. Going... I'm gonna get a beer. I'll see you guys uh next stream. Thank you all for watching. I love some of you. I I am okay with some of you, and some of you... Well, you know, like like you know, being being a member member of the George Hotz streaming community isn't for everybody, um and we do hope that people who write stupid comments will self-select out. Um, you know, like what keyboard you're using? Why why do you care? You think you're G to get a better keyboard? It's going to make you a better programmer? You think that's what's going to happen? You know what's going to make you a better programmer? The gold standard. All right. Tell everyone you know. It's only real money out there. This isn't some gold bug shit. If you don't believe in the gold standard, you're hammered. That's it. That's the only explanation I have for people. Oh, but the modern monetary... Okay, thank you for watching the stream. Bye everybody.