Transcription
I don't know about you, but sometimes I feel like whenever some new model comes out like Kimi K2.6, it feels like everyone's trying to bench max. And this is kind of what I'm saying. My buddy over here, BridgeMine, he's a Vibe Coder, and he's been doing this as long as I've been doing it. And he's been testing out the latest Kimi K2.6, and he ran his same test to create a little lava lamp. And it's actually showing that here's Kimi K2.6, it can't even actually create this little lava lamp. And he has other types of tests and benchmark and he's actually testing it live right now probably as we're speaking because he's actually making products as well.
And in today's live stream I wanted to see if I can go ahead and explore, can we use Kimi K2.6, maybe not to make lava lamps today, but maybe to see if this could be a viable alternative to maybe Opus. Inside of Droid, it's one of the best agentic harnesses for open models right now and they have a lot of really great features around that and you can use basically Kimi K2.6 for a quarter of the credits. So that means I'm paying 200 bucks a month currently for a pro plan and so I can get four times as much. I can get $800 worth of credits if I just stick with the Kimi K2.5, Kimi K2.6 plan. And so there's some folks who have been testing it out and taking a look at it as we have as of now. And there's someone who's been using it kind of extensively over the last 24 hours. And I wanted to share kind of their thoughts because outside of like coding, there are the folks over at Notion. I also think that Kimi K2.6 has been pretty impressive and this is kind of where I'm much more interested in and this is kind of what this like live stream is going to be about us getting hands on a little bit with the model and seeing if it's going to be a right fit for some of our use cases and stuff like that. So yeah, even the factory team themselves say that it's pretty impressive open weight model and I think this can open the door for a lot of us to try to use this model in a way that will allow us to, you know, maybe take care of some other tasks that we're we're not familiar with.
Another thing that we're going to get into is I found this really cool context app from this guy called Arlan. This kid's 19 and he's insane. He's from Kazakhstan, I think, and he did an interview and he has this really cool context management app. So it allows you to grab documentation from your research agents. I'm not sponsored. I don't have any affiliates or anything with this guy. I saw him on a podcast. I thought, man, this is insane. I use EXA and REF extensively, so I want to kind of check this out and see if I can throw this into the mix as well and see how this performs. So yeah, it's called agentic search for external and internal data. This is something that I'm going to be taking a look at and seeing how this performs over the time there.
So in terms of what Krishna was saying, really quick, before we get into the joy things, as far as like this is an interesting flow that I thought was really interesting for a use case here. So using Kimi K 2.6 for brainstorming and then GLM 5.1 for coding. So this is really interesting. So the TLDR obviously is you can get a high return on investment if paired as like a good starting point. So I want to see if I can take my project here where I'm helping my friend. What I'm doing is basically I'm taking a look to see if like I can make a really basic agent and that can go through these different commercial properties and see if it can actually start to do some of the work. I gave this task to Opus and Opus created this little architecture for us. And basically it's going to do like a fetching, a diffing, analyst agent, and also do some notifying. So this would be a really good chance to try to test out Kimi K2.6. This is something I would normally give the Opus 4.7 to do for me. So if I just give it this plan and give it this rough architecture, can it really build it for me? So let's go ahead and get this going. So I can just do Droid. Oh, sorry. I'm just going to do Ghosty. So I'll load up Ghosty. I love Ghosty. So I'm going to just get out of here and then cd dot dot and then create a directory here. So let's go ahead and give a quick shout out right now. Roll call to the community. Big shout-outs to TechFriend, thank you so much for tuning in to the live stream. We got Foukos as well, yo. We got Z400, what's going on, man? Stay live until 12 o'clock for the OpenAI live stream. I got to dip right before that, so we'll see what's up. I'll have to come back on later. I like to look at LLM Arena and compare the Benchmax scores. I don't think Kimi K2-6 is up yet. Yeah, that's part of the reason why we just test what's going on. Yo, what's good, fam? What's good, what's good? The goats, the goats, show some love, people. Yeah, definitely smash the likes. Big shout-outs to Gael, AJ G, what's going on? Have you showed? I only seen one like. Oh, okay. Yeah, we got to get the likes up because we've only got like four up in here, so definitely get them up in here. So let's go ahead and kenny, M-K-D-I-R-D-S-H-P, and let's go ahead and make our permit tracker, SJ, I think. CD permit, okay.
So what I'm going to do is just load up Droid up in here. Let's go big, fam. Okay, so model is going to be, So I'm going to select Kimi K2.6. So this is at 0.25. I don't know if you've seen this in here, but Droid Core also has Minimax M2.7. This is 0.12 credits. That's crazy. So at 0.12, you can basically get 12, 24, 48, 37. Was it 12 times? Oh man, I'm just so bad with my multiplications. But basically quite a bit of multiplications more here that you can get from that model and then GLM 5.1 is a half a credit. That's pretty fire and then you have Codex up here at 0.7 and then GPT-5.4 at 1x. Okay, so let's just go ahead and select the Kimi K2.6. Reasoning Disable. We're going to change Reasoning Disable to Reasoning High and so they said that it eats up quite a bit of tokens. So if I do Usage, Cost, it shows credits usage for the current session. So cost should be zero basically. Yeah, perfect. All right.
And then as far as everything we got going on here, this is pretty cool. So I think what I want to do is I want to see is I'm going to paste in some of the project plan. So this was a conversation on a previous chat on how I can best like take a look at some of this information. And we did some web research for how I can look up some of the stuff for like a workflow here for getting the permits and everything. So I'm just going to go over here real quick and just pop this in, research, and then research. So the reason why I do these tags is because LLMs love delineation and that is really helpful for context. And then I can just give it whatever my goal is here. So I want to see if I can build this using Claude SDKs. Please help me abroad. Okay. So this is what I'm going to do. I should have ExaMCP and the other stuff already hooked up to this and I'm curious to see what this response is compared to what Opus 4.7 gave me because it should be able to have done similar tool calls and you can see it did like eight different web searches here before it gave me anything and so I just want to see what Droid's going to do with Kimi K2.6 high because it has high reasoning and it should be legit. Oh man, the new updates to the Droid are fire. Look at the CLI. Oh, this is really cool. All right. Yo, can you test it against the leak cloud code base and rewrite it? No, no, no, that's too much around here, bro. We can't be doing that. That's going to take a lot of time. Crazy. 4.7 at 1x, it still says 50% off. That's right. So yeah, if you do have, okay, also Half-Life Divided, why do you have those fake AI photos? Because I remember a lot of the fake profiles used to have the people working backwards, so it's kind of throwing me off a bit, but it's okay. The 4.7 at Half-Off, yes, that is very true right now. So okay, cool. So empty state, perfect for prototype. Here's a clean plan for the SJ Permit Tracker. So here's the prototype architecture. So it actually comes up with similar information, right? The fetch, the filter, the diff, the enrich, the draft, and deliver. So did it search through any of the other directories? It says thinking, let me check the current directory, see if there's anything there. It didn't actually use the tool calls, which I thought was really interesting. I thought it would, to be honest. Huh. Actually, they said, help me with the basic prototype of this. So they want a plan. Let me present a clear, concise plan for this prototype using Claude SDK. I mean, listen to my instructions. Okay, can you, so I should go back into spec mode, and then I should say, can you use EXA and ref tools to help enrich this research? And so now that I've bumped it into spec mode, it should go back into planning. But it's, it's, I mean, it was a pretty fast return here. I mean, like, there's nothing too complicated. This is also kind of the same here. It's basically saying the same thing. Opus gave me four pieces where in factory basically just said, you know, here's your core loop, that's it. Like, and then here's the tech stack. And I think Opus gave me an example. I'm going to get rid of this thing here. So Opus gave me, did it give me any code examples? Yeah, it gave me like some more details here, but that's, that's pretty much that. Minimal Layout ends. I mean, Kimi K2.6 High is not that bad, to be honest. I'm curious to see how it's going to do for planning here because obviously it's reached out to X. It ran four different web searches in parallel. Man, the new Droid UI is so fire. Oh, this is so good. I love all these little colors. So it's taking a look at these different resources for the patterns, similar projects. Now I should ask the user some clarifying questions. Yo, Kimi K2-6, I'm liking the way it does this conversation right now. Oh, okay, I have to try Kilo. Oh, wait, I see the console. Yeah, I'm using the Droid CLI. It's a picture of my wife from 10 years ago. Bruh, you gotta get a picture in the front, man. Oh, Kimi CLI, no, this is a Droid. Let's see, okay, so GERD Research Hall, there's some key findings before I unlock the spec. The data source confirmed, San Jose OpenData has two relevant CSVs. Perfect, similar project. Ask the user, which CLAUD integration do you prefer for the prototype? Simple Anthropic API call, straightforward, or CLAUD Agent SDK? Agent Loop, hands off, yeah, I want to do the Agent SDK. Okay, which data set should the prototype target? Last 30 days, active building permits. FullData said 80k rows, more comprehensive. Yeah, let's go comprehensive on this thing. Have I tried GitHub Copilot CLI? No, I have not. Is it worth it, though? I think it is. A lot of people are saying it's pretty good, but maybe for the cost is probably why they're saying that. How should the prototype deliver the results? Terminal, report, or JSON file, simplest? I should say HTML, shareable, yes. Yeah, yeah, yeah, perfect. Not anymore, LOL. Ha, okay, cool. All right, so we're just running it on this simpler task, and so I want to do something in parallel. So while it's doing this, do you guys remember, if you weren't here, I did a redesign of my blog, and what we did here is we gave the actual, like, we told Claude to give us a handoff prompt. This handoff prompt was pretty massive and we were actually extremely surprised to find out that the Gemini 3.0 flash model absolutely cooked, like cooked, cooked inside of the Stitch with Google. So for those who aren't familiar with Stitch with Google, basically what Stitch with Google does is they have their own like design platform similar to what Claude design is here. But what's cool about it is the fact that it really doesn't cost that much money. I think it was free for me and we gave it this entire prompt and it just went to town. And what was really surprising though was that the pro model was not even as detailed as the actual light model, which being, you know, Gemini 3.0. So the Gemini 3.0 flash outperformed it. So why don't we just go to like Kimi.com and I want to test a couple of things. I I want to see Agent Swarm here. Let's go see if I can send some swarms up in this bad boy. So I have the Kimi thing here and I'm just going to give it the same prompt with the Agent Swarm. Long scale search, long form writing, batch tasks. So let's try Kimi agent. I don't know, this is so confusing, but we're going to try this out. All right, let's just try it. Okay. So Agent Swarm beta, what does it do? I'm so confused. So this is long scale batch form batch tasks. So I'm gonna try both. Why don't I just try both? Okay, so let's just try agent, Kimi K2 agent. Yeah. And I'm just gonna paste this in here. So I think we have everything in here. Handoff prompt. Paste this entire chat. Okay, I'm just gonna just whatever. This is what we gave to Google and we're just gonna give the same thing here. So this is Kimi.com on the web and I'm gonna see how it does here. And then I'm also gonna try Swarm as well. Yeah, Andy says, do the swarm, do the swarm. Are y'all in the swarm right now? I don't see any likes up in here. Like, we only have 13 likes and we got how many people watching right now? We need some more likes if we're gonna get some swarms up in here, you know what I mean? This is gonna be some serious dough. And it's about to be really crazy though, right? So let's go ahead and check it out. Come on, let's go. Let's get the likes up. Swarm time. I need some more likes before we hit the swarms. Yo, Gary, what's going on, man? It's swarm time. I tried asking Kimi what it does and it was thinking forever. Even OpenCode has 2K token usage. Ooh, okay, that's interesting. How about a single file system Verilog source for a multi-cycle RISCV? Bro, I can't even read this, but I know what you're talking about though. CPU UR built up from primitive modules within the same design work of FPGAs. Bro, that would be so fire if you can give Kimi Agent Swarm some really low-level tasks like that. Bro. Okay, that's probably going to be a whole different stream. We've got to do that, something different. But yeah, Swarm. All right, we've got to get the likes up. We've got 24 likes in here. We've got to get it up. Once we get to 30 likes, I'm going to start a Swarm because right now we've only got 24. And I know there's some more folks up in here. It doesn't take that much more to get some likes up in here. So let's just go ahead and see how this is going right now before I get it up in here. We've got 25 likes. Jesus, that'd be too far without verifiable artists. Yeah, right? Like, that'd be interesting. thing. And then I also want to see Droid. So let's go ahead and also see this, what's going on in Droid. So in Droid itself, I'm going to just create mkdir-p. I'm going to do, let's see, RayBlog, Kimi K2-6, right? So, Rayblog, oh my god, I have so many of these, Blog, Kimi, okay, Droid, I'm just gonna burn all the tokens today. Model, and then we're gonna do Kimi K2-6, Hi, alright, we're gonna give it the same prompt in here, and I'm just gonna let it rip, right? So I'm just going to say go. I have no clue what's going to happen. I should have probably put it in spec mode but we'll see. I just gave it YOLO this command inside of Google Gemini 3.0 Flash and it knew what to do. So I'm just really curious to see. It's already going. Oh my gosh. So this is an empty workspace. This is a build from scratch. I need to create a comprehensive to-do list and start building this blog. Let me help the implementation plan. So here it goes. It's going to set up the Next.js project with AppRouter, set up the design tokens, build shared components, create 24 posts data, build page A, B, and C, tweak the controls, test everything. This is a massive task. Let me create a to-do list to start working. Yo, Droid is crazy. Okay, let's see. I think this is the best approach is create basic Next.js, UCSS variables, self-host for the Google fonts, and then use MDX for post. Let me initialize the project first. It's going to town. It's going to town already for me. This is crazy. So thinking, good, the project is created. Now I need to install the dependencies. Perfect. Okay, so it's going in. It's going to go all the way in. And I just gave that command to Droid and it's just chilling right now. So that is going to be interesting. I want to see how much is it going to cost us. And we're seeing all the thinking tokens fly by, right? So this is it just kind of doing its thing. By the way, the new Droid CLI stuff is Fire, so fire. Wait, is this Droid? I'm sorry, I missed that. Yes, this is Droid, and they have a new update for their CLI, and all of the new fonts and colors and stuff just look so beautiful. I'm using it inside of Ghosty, and that's what's up. Thumbs up, folks. We got 36 likes. Okay, now it's time for us to do the swarm. So while this is running, let's go ahead and go back. We have previously went to Kimi.com, we gave it the same plan, and I'm going to see, is there going to be a difference between what's on Kimi K2.6 website versus what Droid is? Code is doing, right? So the way the models are hosted, the instruction sets are going to be definitely different here, but I'm really curious on what the results are going to be, right? So we have a couple different projects in the fire. It could be very confusing if you're just hopping in right now. So we have a couple different projects. Project number one that we've launched right now is basically like a city permit, like lookup type of thing. So I have a friend who's in commercial real estate and asked me and said, hey, Ray, can you build like some type of agent or something that can help me go look this stuff up. And I said, oh yeah, this is a basic agent loop. Let me just hand this off to Opus and kind of see if it can kind of create some basic structure for this. So we gave that same prompt to Droid and Kimi K2.6 inside of there. And the results were pretty much a little bit similar. And so that's project number one. Project number two that we have is ongoing from the previous time. We already have a baseline, which is basically my blog. So rayfernando.ai is being basically redesigned. And we did that in Claw Design. So Claw Design gave us a beautiful handoff prompt and extremely detailed, very compact and succinct, but really, really well ready for agents to kind of take over. So that's the thing that we're now going to be testing with Kimi on 2.6 on the website, Kimi 2.6 inside of Droid as well. And then the third one that we're going to do right now, because everyone's been smashing the lights right now, we're going to do the Agent Swarm. So I'm curious to see what an Agent Swarm will look like inside of Kimi and I think everyone's ready for this because y'all been smashing the likes so we do Kimi.com this is gonna be crazy you know Kimi is basically I had a plan that was already activated a while back and I I think I'm I don't know how many I think I still have credits left I got some credits left it says upgrade your plan yeah I think I still have some credits left let's see okay Yep Let's see I don't know It could fail I just don't know Let's see Agent Swarm So we have Agent Swarm selected down here And I'm just going to paste this in right now And I'm just curious to see what the Swarm is going to do On the web, right? So if I mean, we're already seeing how good the model can think and behave Just inside of Droid And it wasn't even a Swarm I mean, this is going to probably go crazy When we gave this to Grok last time, we're going to compare because I paid $300 for Grok on my last live stream and we did Grok 4.3 beta and it sucked. Like it didn't even come out. So let's just try it again. They keep updating beta. I'm just going to give it the same prompt. We're going to give it the same prompt and just see what happens in Grok because this one like took forever. It just was just a pain in the butt. Yeah. Is Ollama and Kimi worth it? Right now, the model's so big to self-host, you'll need, I think, like six or eight, like, H200s or something like that? I don't remember. It's like a pretty big model to host. I'm a student, so I can only use the free plans. Yo, man, gotta get your dollars up right now. You know what I mean? Let's get it. Let's see. Is there a GitHub code base of people who have created it so you can possibly compare and check stolen code? I don't know. Oh, Paid Ollama. I don't, I haven't checked that out yet, actually. Claude is still king, right? Low key? Yeah, $300 was rough. I'm trying to put those $300 to work. Exactly, Maria. Exactly. And I'm also trying to get the update so I can just tweet Elon Musk and say, yo, fix your stuff, man. Y'all got to go get on this right now. So let's see. So I'm using the available. Okay, so it is actually working. And I'm curious to see if this version is going to work a little bit better. So yeah, that's it. So Agent Swarms. Let's check on our agents form right now. Dude, I have like 10 Claude accounts. Okay, Gary, are you one of my Korean friends who's basically spending 2 billion tokens a day? Let me know because I think you're on that path right now, bro. Do you find the Droid CLI better than Claude Code? Yes, 1000%. I think the couple of reasons is because of its ability to just do the long-running tasks, their agent readiness stuff, which you can also make in Claude Code, but also the performance is super fast. compared to ClawCode. ClawCode has a lot of regressions. Like every update, it goes slower or something happens and it's just so inconsistent. It's just kind of painful. And so that's why I like Droid. Yeah, I'm rooting for Elon. I hope Grok 5 is mind-blowing. We'll see. We'll see. We'll see. Yeah. Oh my God, Gary. You're that guy. Yo, Gary, the chat. Y'all got to give some respect to my man. He's got the 2 billion tokens a day. He says we should talk. Oh, I don't know if I'm ready for this. What about a Rust Agent system that runs tickets autonomously? We shall see. Can you use ClaudeSub with Android CLI? Yes, my buddy Corialis has something like that. Speaking of which, for those who are kind of tuning in right now, if you made it this far, I'm going to transition my community over to, I've been looking into this a lot, this new thing called School, S-K-O-O-L. and so what I like about school compared to Discord is the fact that it now has more of like a like forum style type of thing so that will allow me to actually participate every single day so if you ask questions in the morning I can log in and see all the questions and then other community members can also answer them because I find the the problem right now with Discord is that I can have all these channels and sometimes I want to point people to resources but I noticed it was a couple of months ago so I have to scroll back or I have to go to pin posts and and if I pin posts in different areas it's just hard to search overall and so in a community setting like especially like school I can organize a lot of the resources and I can have like videos and pieces of content that are dedicated so be on the lookout for that I'm going to be upgrading that stuff pretty soon and I'll be sending out a notice so yeah just just stay tuned I'll probably have some stuff in the description as well I'm gonna I'm gonna try to work on that today because I really have a lot to give and I feel like a lot of community members who are with me like Andy, like some of the other folks as well, have done a phenomenal job kind of, you know, hanging out. And recently we just did a little bit of a hackathon this past weekend. And before that, we were doing some product showcases and kind of going over what people were working on, which is really cool. So I want to kind of keep building the community. And I found that that's probably going to be a better platform than Discord for me as of now. So yeah, just FYI, school clone. Okay, so Gary, trust me, I'm already two steps ahead of you with this one. I just didn't want to tell everyone about this yet, but I don't want to, I want to test out school for real. So where is my school clone? I have, building a prototype with Claw. Check this out, bro. This is, oh no, is it here? Design skills, building a production, planning. Let's see. I, I have it in here. I, I have a whole, let's see, school. School oh gosh school school is it this past week Yeah I have a spec that just pretty massive I think it in here Yes architecture I think they called it AI Forge but basically this is my architecture plan for the entire school rewrite, and I'm going to try to see if I can build it for scratch just for fun, but it is a massive, massive, massive app and I'm realizing if I spend like a month doing that then I spend a month away from my community that I could possibly build so I might just build it in parallel slash on the side on the off time because you can see like all of the different features that I want to build into a school community require all of these types of things in code right like the marketing like that like actually sending emails right all the different resources for calendar onboarding members like all these relationships and I could do it. I could probably I could spend several billion tokens doing it. But I want to make sure that I dedicate the time to the folks who actually join the community, have questions and answer them and make little videos and stuff like that. So that's probably going to be a later project, but that'll be what's up with. So how do I connect with you? I'll pay for your feedback, bro. Yeah. So I think what I'm going to do, stay tuned, Gary. I'll have something up when I get the school community launch. I'll have a way for people to pay for some one-on-ones. Yeah. Build it live. That'd be fire. All right. Let's check up on our... This is Grok. It's returning back results already for us. And it's just going a little bit slower right now than I was probably thinking, because it's actually trying to build the HTML and everything for us. So let's check up on our swarm. Let's see how our swarm's doing. So our swarm went is... Let's see. What did it actually launch here? So it says thinking and it says first I need to write a plan and then load the Vibe Coding Web App Swarm skill and then execute it to build. Let me start building the plan MD and then read the skill file. So it looks like it did that here and it basically created its own computer. So the computer is basically it sounds like it's an isolated container that will basically do its own thing. And then it's going to looks like it's just filling itself with skills. and as of now basically it's going in here and it started to execute like the web app initialization and then it created a sub-agent called the Pro Designer. So the Pro, my sense of accomplishment comes from the moment of overcoming difficulties. Role description. You are a world-class web designer. You guys want the prompts right here. You create comprehensive detailed design documents for websites. You have deep experience in modern web typography, color theory, animation, and responsive layouts. Low-key I think that's the only prompt they give this Pro Web Designer if you know what I mean, we'll see. So this is the Pro Designer. It's basically going to work right now. It's doing its thing. And then this is like the agent's window, I guess. Switch to computer window. Okay. Switch to agent's window. Okay. So here's the thing. This is totally new to me, by the way. I'm just kind of navigating this right now from scratch. Okay. So thanks. Okay. Design complete. Now let me design document and proceed to page three integration. Okay. So here it is doing its design stuff and here's the React stuff. And then here's a create scaffold. Wow. So we've gone to design. So here's pro designer. My sense of accomplishment comes from overcoming difficulties. Yeah, that was the pro designer. And then it did its design MD. Let's see what the design MD says. So the design MD basically took my SPAC and then told it to create this. It created this from my SPAC, I guess you could say. And so this is around the design. Okay. Cursor animation too. That's interesting. So a scaffolded project and here's the scaffold builder. So this is the person who's, you are a React expert specializing. I think literally this is the system prompt. You write clean, concise code with CSS modules, follow design documents. Exactly. You build broadcast HUD aesthetics and you prefer CSS properties. Okay. Okay. Okay. Dang, okay, try math reasoning, maybe, maybe not, let's see, so let's see what's going on, as of now the Agents form has run the scaffolding script, React, create its own design documents, install the React router, now it's going to do the CSS stuff, and then it's going to go create posts, and then it's going to create the shared components, create some utility classes, or components, and then implement the homepage, and then the router, and then just get committed, so I have no clue how much longer this is going to take, but hey, it's still going. That's pretty cool. So that's the, and I didn't get kicked out because I didn't have credits or something like that. So that's cool. Wait, we're already done with the non-agent swarm one. That is crazy. Let's see. Let's see. Let's see. So damn, let's, is it, is it still working? No, it's done. This is it. Oh, this doesn't really work here. These little filters and stuff. Okay, but let's just click one of these. Bruh. Okay, let's, can I full screen? Oh my goodness. Oh my goodness. Oh my goodness. Wow. Okay. Damn. Let's see. Not found. Okay, okay, okay. It didn't do the, wow. Okay. This is weird. Like, mon days. I mean I get it yeah yeah yeah I mean let's see if I can switch to tablet mode or whatever or phone mode yeah there's a little bit of overlap here you got kind of lazy with the design this is kind of awkward and weird the font sizes are it's like definitely not optimized for mobile for sure so kind of but it still has a settings panel let's go light mode okay okay it's still kind of broken so it It's like 80%, you know, you get 80-20 from the web. This is just the web, non-swarm, right? So like, one-shotted from the prompt, okay. That's okay. I'm more impressed with Google's Gemini Flash 3.0, which was also ages ago as far as like AI terms, you know what I mean? But not bad, not bad. Let's see the streams page. Let's see about page. Okay, so we didn't get the streams page That didn't load for me If I go to about, I see this is my about page This is pretty fire And it says what I'm cooking And it's got all my different drops here Damn Wow, okay So, sort by newest This doesn't do anything Subscribe doesn't do anything Okay These don't work They're just there as kind of like a cute little preview Let's see So it did write a It looks like a Vite app or something like that, Index.html, yep. Okay, yeah, it's just like a React app. Easy to fix, though. Can you create old Renaissance art influence? Yeah, interesting. It seems like better than Opus models do 90% of the job and then you have to do the other percent. Yeah, exactly. Yeah, like Opus would at least finish it and like keep making it work or tell you that like bro you're out of tokens let's go ahead and like pay up but it's at this point am I gonna I don't know like how do I say it I'd rather use a Gemini 3.0 flash at this point you know because like this looks cool it totally followed my spec but it didn't like fully follow it right this is the now I have to spend a whole bunch of time fixing these colors where like it didn't have that problem with Google and go dark mode and like the filters and a bunch of other things that just really aren't worth it. It's also not optimized for mobile right out the gate. So yeah, I don't know. Like, okay. You know, I gave a lot of work with the prompt, but it's okay. It's okay. Like, I mean, if it's 0.25, the cost, are you going to spend the same cost eventually getting there with, you know, using the Opus model at 1.0? You know, I don't know. We'll see. I'm really curious to see what's going to happen with the Swarm because the Swarm apparently is supposed to have more agents that do this stuff here. so maybe there's going to be more attention to detail that these guys get for the swarm that the other guys messed up. So this thing's almost done already so it's already doing the git commit stage. My task is completed. Alexander. Okay. So scaffold complete. Now phase 5. Merge scaffold and create branches. Then phase 6 parallel agents. Yo, this is crazy. This is going to be interesting. I'm super interested to see what the results are here. Okay, Master Scaffold, K26 versus GLM51, that's interesting, I don't know, that would be a really good test too Okay, so now create the page branches, parallel agents, okay, while this is doing its thing there, let's just check up real quick on what's happening on the, okay, now hold on, I didn't do anything yet, I want to see I feel like we're almost done, right? Launching both page builders in parallel? Or does it still have more stuff? Like, I don't know. I really don't know. Assigning tasks? Huh. Are we almost there? Thinking, reasoning? I think it's just analyzing all the work that it's done to make sure it's complete, right? I don't know. Agents window. It says running the setup right now. Okay, iterating. Yeah, so I was trying to set this up right now so it could deploy the app, basically. Right? And there's two agents that are working in my swarm right now. Autopilot for the win. Run auto-research on swarm. You mean Karparthi auto-research? That'd be fire, though. I think that's actually what they did, Kimi folks. and they open sourced something, X.com, using this, the Carpathia Research and they, I don't remember what exactly they did, let's see, Kimi, Kimi, Kimi, Kimi, Kimi, Kimi, Kimi, let's go, here we go. We're open sourcing our high performance Cutlass based implementation of Kimi Delta attention kernels. It achieves 1.72 to 2.2 pre-fill, speed-up, deliver over the flash linear attention base on H20 and works as a drop-in back-in for flash. I thought somebody commented that they did this using the Karpathi auto-research thing. I'll have to hit up Crystal. Crystal! Flash, high performance. Let's see. Okay. Yo, Kimi's cooking, though. Not only, but the whole open source, yeah. 2x Prefill Speedup is, yeah, it's massive. It's ridiculous. These guys are just going brr. What's going on? Kimi Agents has a cool UI though. Yeah, yeah, yeah. It kind of low-key reminds me of the Manus stuff, right? Oh, we got a new GPT today? I don't know, man. I don't know. Gonna be insane UI. using auto research method to improve custom made skills is the heart fire. The more you work with the skill the more it learns. That's also true. Okay so I think it's done right so let's see what happens. No that's Grok. Okay let's let's go let's get back to Grok later. Let's see what's going on with our Agent Swarm. So our Agent Swarm is back to the latest. Let's see what's going on. Our Agent Swarm is still working right now so it's doing some type of build thing. Okay so I think I think it's setting up the articles for my, if I go to my agentic tasks, agents window, I think it, oh, execute tasks. Yeah, so I think it's still working on actually generating the blog posts that I need for them. So is Kimi2.6 worth it? As of now, it depends, right? We gave it one design task. I think it really good for conversation but I curious somebody mentioned that using Kimi K2 to kind of plan and do some of that stuff and then hand that off to like a GLM or something like that for coding could be really cool but we see like right now we having to do coding tasks too right so we having to do everything also in Droid we gave basically as part of Droid core Kimi K2 we gave it that same prompt and the single prompt and oh wow look at this thinking tokens It likes to do this, right? So now I need to create article client, now component, you know, like it just repeated itself over and over and over again. This is a characteristic that I saw on Twitter that people mentioned that it likes to like repeat its thinking tokens. So sometimes it does that. And we're kind of seeing it here firsthand. Thinking, thinking, thinking, thinking. And you can see it just repeating itself over and over again. So that's interesting. So here's the thinking tokens, and we're seeing it from Droid, and you can see how it just keeps repeating it over and over again. So it could just be an artifact of the way the model is behaving, and this is probably going to be tuned and stuff like that. So I'm really curious. So yeah, we gave it the same prompt in Droid, and it's basically now at the last step to fix errors and verify against the sanity checklist. This is pretty cool. So what has it said so far? Let's take a look. The build succeeded with no errors. Let me verify the sanity checklist. The top nav should have line beige working. Category chips, I need to check if all the pages look correctly. So it looks like Droid's probably launching something to make sure it could do that, or if the model's doing that itself. I don't really know. Actually, Lemurify, the DST output was created correctly. Okay, perfect. So we're almost done. Okay, then it just copied Gemini. That's a Gemini thing? No, no, no, no. At a quarter of the cost, oh, eight the cost, 10x the loops. Okay. Do you think professionals would use cost-efficient models like Kimi K2-6 and Minimax M2-7 and Promptit, right? Yes. So this is a big thing, Ahmad. a model, Kim6054. The reason why we're able to get such a cool looking design from just my 500 line prompt was because a smarter model like Opus 4.7 Thinking inside of the design of Claude design was able to create that really succinct prompt but have so much instructions baked in that even the Google like 3.0 Flash model was able to take that and just run with it and go to town with like a million token context window. So that's the power of having really good prompts. So if you get a bigger model even though you spend more money on Opus you know 4.7 thinking you know maybe spend those 2x tokens you 80K tokens, 100K tokens, 200K tokens to get whatever plan and results that you need and even specify that you're planning to hand this off to other models. It's going to be very helpful at getting that big prompt and then that big prompt that you have, then you can feed that into the smaller models like GLM 5.1, Kimi K2.6, maybe the Minimax 2.7 and say like, here is this plan. and you're taking chunk one and just going to just work on this and then have another model like GLM 5.1 verify the work or something like that because you don't always need the most expensive model doing every single task. You can even have like the codex model, right? The codex models are really good or even just GPT 5.4. 5.4 is really good. I mean, it's really, really good. You could just like have it do the verification work, right? Or even the implementation work for you. It's pretty, if you have one of those plans is fire. I mean, right? So like you don't always have to use like the top of the line plan and definitely switching them up is definitely going to be super helpful though. Yeah. Likely Droid with Kimi with another harness. They might need to update. I have a lot of harnesses. Usual splash aspect. Yeah. Autopilot loop with Cron, self-pacing, use agent teams, create agents, use swarms, create blocking tasks. Yeah. I mean, you could just go on and on and on, right? Gemini 3.1 taste in UI. I'm waiting for that Flash model though. That Flash model is going to be fire. Yeah. 5.4 sick. Everybody complains, but they forgot what six months ago was like. Yeah. Yeah. No, I appreciate y'all. I mean, I mean, I love doing live streams. This is kind of why we do it because we've got to figure out this stuff live. So yeah, stay tuned. I'm going to have a community launching in school, S-K-O-O-L. I'm going to have some of the links up in the description, in the comments coming soon. I'm transitioning stuff from Discord over to that and so that way we'll get stuff in and I'm going to have a special deal for those folks who are already joining in you know early on because the prices are going to go up with the because I want to focus on making sure I can answer everyone in the forums and some of these experiments and stuff I want to do live here but also do some stuff for the members in the community and stuff like that so yeah stay tuned for that that's gonna be really fun. All right so this is still verifying its errors and inside of the swarm it's still kind of doing its thing there. And in our other site here with Kimi, you know, this was just okay. Like it came out with a cool design because, you know, I'm not surprised anymore that it does this stuff, but it didn't have all the extra details that we were hoping it would, like the filtering, you know, changing from like dark mode to light mode correctly. It's kind of missing some of those details. Streams, the link doesn't work. If we go to about, it does work, you know, and it does key off some of these tokens and stuff which is pretty cool and you know these these are all in the instruction set so it's kind of how do you say like it's just hit or miss in terms of like what it what is it that is actually built for you like does it respond to all the font changes yes does it respond to list view edit view no it doesn't you know so that's just kind of you're like well okay now I'll tell the model again and then you know you're back to using the same amount of tokens that'd be really interesting which open source model do you think is the best in terms of design and front end I feel like you got to mix them all up, right? You have to try them all out. So people are liking Minimax 2.7, like just for execution. 2.6 is not bad for design too. I mean, some of the fire landing pages that they've made off the get-go,
You know, just from zero to prompt or whatever. I don't know what just happened.
Okay, so okay, we're back here. All right, so we're almost done with this site, but it sounds like our droid just finished. What task did it just finish with? Oh, we were supposed to finish this thing. Sorry. Okay, so it ran the build. I think it's kind of hung up, to be honest, right? I don't know. Is it still, are you still there, Droid? Run the build fix errors? Does it, let's see. We'll see.
The parrot effect effect effect is in harness. It's the model. Oh, okay, so Gemini calling it. Okay, so the model is the one that's basically doing that overthinking thing. What do you think about PyAgent? PyAgent is fire. Yeah, PyAgent is like, you want to go down a rabbit hole of really just getting to the core API layer and talking directly to the model. Pi is a really cool exploration. But another exploration that you should also check out in parallel is RepoPrompt. So RepoPrompt has a similar theory where you can control what goes to the model and then like craft your prompts because you'll get everything you see and then get sent right to the model. And so RepoPrompt has some new features that you should check out too, which is around agent stuff. If I had more time today, I would probably check it out because I want to take you all through the walkthrough of this. So here's the latest RepoPrompt thing. And so in ReproPrompt now you can basically have like an IDE mode and agent mode. So the agent orchestrator is really cool because similar to my friends who made the whole thing on my open codex where you can send off multiple agents, you can basically now do that in a nice UI inside of ReproPrompt and you can use your existing account. So I have a ChatGPT Pro account. I have a Claude Code account. You bring your own accounts and then you can use it within ReproPrompt so you don't have to pay for any API keys and stuff like that, which is cool. You can bring in the Gemini accounts, you can do everything. So more on that, that's probably going to be a dedicated thing later. So that's what's up.
So the static export is in my open, okay, to preview. Okay, so if I just do, let's see, open this, does it open it up for us? Is that it? Something's missing. Let's see. Is this? Hold on. Let's see. Yes, the blog builds cleanly. Okay. Okay, cool. Starts writing this down. Actually, it's more of like independent. So repo prompt independent from pyagent. PyAgent is like you have some workflows that you do over and over that maybe your existing agents can't really do with whatever the tool set you have. So you can kind of customize that. Like if you want to insert something at a thinking step or insert something at a build step or do something that any of the agents really don't offer right out of the get go, you can definitely start going to customize things. A lot of people say that they feel like the Opus models are smarter if they use them directly from PyAgent. Also probably very true. That's kind of why I use Droid, but Droid has the best, you know, as far as benchmarking. So yeah, I mean, just kind of figure out your own use case.
The CS isn't loading the next year static export uses absolute. Yes, okay, cool. So I haven't even checked the website. So let's see. LS, is it just see my app? Let's see. Oh, I should just basically just do a BunDev, BunInstall, BunDev, BunDev, right? Oh, my bad. I think it's, let's see. Here we go. Okay, so this is Kimi inside of Droid, and it kind of does work, kind of doesn't work. The filters, okay, did I just mess it all up? Yeah, I messed it up because the code changes. Okay, so. Try opening the file of the directory. Okay. Yeah, let me see. Bun install. Bun dev. Cool. Go back to here. Yeah, refresh. It doesn't really load. Okay. This is a signature here. Like, it doesn't really follow this. It's so funny how it figures out these design things, right? Huh. Weird. Okay. Let me just do mobile mode. Yeah. So Kimi doesn't like to do mobile mode. Oh, this is nasty. At least the other sites did it, like Gemini. Like, see, Gemini has that baked into its, like, instruction set. Yeah, this little thing to change the settings doesn't work either. Let me go to my about. Oh, no. Okay. Yeah. Hmm. Okay. Interesting. Very interesting. Okay. I'm not going to spend any more time with that.
So yeah, that's basically, like, I feel like I gave, I keyed up Kimi K2-6. Like, I gave it a fire, like, landing, like, here's a fire prompt that Google Gemini 3.0 Flash really just crushes. But look at all these details, right? So there's a lot of verifiable stuff, but it just didn't really pull through all the way. It just kind of got you 70% of the way there, I would say maybe, because you still have to do mobile view, right? We still have to do light mode, dark mode, we have to still implement a lot of settings. There's a lot of other pages that didn't implement. In fact, I feel like it did worse in Droid than it did on the web, right? Like, oh, I already did the other, if you check out our other live stream with Claude Design do. That's what we did, Gemini Pro, and it was not as good. It was like Gemini 3.0 Flash just crushed it. Yeah.
So let's go back to the Swarm. Here's the Swarm. All right. So does the Swarm perform better than all the other stuff that we've done so far? I don't know. Make sure you smash that like right now because let's just go ahead and get right into it because this is about to be really crazy. We're running Kimi K2.6, Agent Swarm, and so far, as of right now, I feel like it has been benchmarks. I mean, everyone's coming out saying, yo, this is fire. We gave it a really succinct 500-line prompt that has everything about our design detail. Google Gemini 3.0 Flash, which is basically from ages ago, from like, it feels like seven years in AI years, nailed that thing end-to-end completely, you know, with my whole design and everything. Kimi 2.6, just straight out the gate, not so much, maybe gets you 75% of the work done. And so we said, okay, let's just check it out inside of Droid and see how it performs in the agentic workflow there. And what was interesting is that we're seeing a lot of the model repeat out like repeated thinking and stuff like that. And when it comes out, it's maybe 70 to 65% of the work has been done. And there's still a lot of work to be done. And I was really surprised because I was hoping for it to actually yield a lot more results.
And so lastly, the agent swarm. Is the agent swarm what is supposed to be like the Kimi K2.6? Is this kind of what it was actually benchmarks against or is this what people are actually thinking is the model that is really, really good? Let's find out. All right, so let's go ahead and check this out right now. So let's go ahead and see. It should have popped me into a special page here. So let me just here. So here's our page. So the first thing that I'm going to do is I'm going to check mobile view to see if it does that. No! Why? Look at this. That's nasty. This is nasty. No, come on. Let's go. This is Agent Swarm. This is supposed to be the best of the best. And this is your mobile view. Trash. Trash. I don't know. What do y'all think in the chat? I mean, are you seeing what I'm seeing right now? Are you getting different results? Because I know that people who are saying like this is the best model, Opus Killer, y'all got to get out of here. Don't listen to those people. Y'all need to unsubscribe to them because they're just hitting hype just for the sake of hype. This is why I do these live streams because we need to see this firsthand. Who does this stuff, bro? Like Agent Swarm? I mean, come on, bro. Like this is not what are you? doing, bro? Opus is king. Low key, right? I mean, we just did Agent Swarm. We swarmed that ass. Oh, man. Because, I mean, I, okay, so at least these filters work, right? At least the Agent Swarms tested the filters themselves, right? At least they did that. Not a good first impression. Definitely Benchmax. Yeah. This is a disaster. Holy shh. This is a disaster. You gotta use it. All hail Opus. Opus, hey Codex. Codex, Codex, do not sleep on Codex. That's right, that's right, that's right, that's right. Gemma, do better than this, okay? Do we need to test Codex right now? I need to head on out super soon, but I feel like I gotta test Codex. Last time I tried Codex, everything was just all frozen up. Please don't be frozen. Please don't be. Yes. Okay, we're back. What happened? Okay. Sick. Okay. So can I. Okay. It's just GPT-5-4, right? Can we just GPT-5-4? I'm just going to do this. New chat. I'm going to go to. I'm not going to work in watchclaws. I had a new project. All right. So we're just going to create a new one. GPT, okay, RayBlog, GPT-5-4, okay? Let's go ahead and create this real quick. Go in here, and we're just going to do full access. We're going to do high or extra high? I feel like we got to go extra high, right? Yeah, we just got to do extra high. All right, let's do extra high, and then let's go back to this prompt. I have a feeling it'll do pretty well. Everyone says that, like, you know, designing inside of here is like a skill issue. Let me update. Do I want to update again? I feel like I have to update. Yeah, I got to update. I want to make sure I have the latest sauce so that way the Codex team and people don't come back to me like, oh, we fixed it in the latest update. All right, Codex, this is your chance. GPT, America, you need to prove me right. Can America? All right, I'm not doing these types of comparisons, American versus China. I think it's all about everyone coming together to win anyways, but I'm just kind of playing with y'all. Let's go ahead and just put this prompt right in. Let's go. All right. So we have GPT-5-4. We're going to use extra high because I want all of these details to be caught, right? That's the reason why I'm doing extra high. It should be able to reason within itself and know that it didn't finish things and just keep going. That's the whole point of extra high, right? GPT-5-4, we're going to give it full access. I'm just going to YOLO. It says Codex suggests... Okay. Let's go ahead and send it. We're going to full send it right now. So this is going to be GPT-5. We're going to check on this in a little bit. And this is, I'm just going to fully send it. Fully send it. Implementing Ray Fernando AI redesign. It's just going to look, there's going to be like, hey, there's an empty folder here. It's going to go ahead and create the stuff. I mean, it's going already. So let's go. The workspace path is in the repo. Can I just do slash fast now? Let's turn on fast. I should have done that to begin with. All right, it's okay, it's okay. I didn't want to use up double my usage anyway, so I think we're good. I think we're good. Send it! Can I wait for DeepSeek v4? Bro, we've been waiting for DeepSeek v4. Where are they, bro? Y'all gotta come out. Enable fast mode. Nah, nah, nah, nah. Flags in the chat. They probably use different metrics in China. Hey, I don't know. We need Broomstick. Extra high, overthink. No, I think this is what we gotta get what's going on. Team China, sorry guys. What do you mean bro? Where are you from? America, Canada, America. Extra extra High is better than extra high. We see right? Should I just pop in a new window and just do X high fast or something like that? I don't know. The browser is still the best. Oh, you know what we got to do? I want to compare this to Cursor 2 Ghosty. Let me make another one. CD dot dot. Okay, I'm making so many of these projects, it's so ridiculous. So I want to try Cursor Composer 2. So Cursor Composer 2 is supposed to be built off of Kimi 2.5, I think, but it's supposed to be really, really fast. So I'm curious, can Cursor Composer 2 step up into the mix up in here? We have like a Chinese-American mix. I don't know. I'm super curious now, right? So CD Workspace MK Deer Rayblog, Composer 2, Cursor, Composer, Rayblog, Composer 2. Okay, so let's open up Cursor. Let's get Cursor Bay up in here. I am super curious now. I'm going to open up the new Cursor 3.0 Agents Window thing. I am just absolutely just new agent. We got to go up in here. We have to change the project from here to Cursor, RayBlog, Composer 2. All right. And then we're just going to check Composer 2 Fast, Edit. Yep. Perfect. Perfect. So I have Composer 2 Fast. Will Composer 2 Fast cook for us? Ladies and gentlemen, this is the question of the day. So we have Cursor 3.0. Give it the big prompt, the big daddy prompt. And let's see. I am your father. Yo, what's going on? Ruben Garcia Jr. I recognize your face from all of the Alex Finns chats. Y'all better watch out with Ruben up in here. This guy is crazy, crazy, crazy, crazy good. So yeah, thank you so much for showing up, Ruben. I really appreciate you becoming a member. My boy always makes it rain over there. VS Code, if anything. Okay, okay, okay. So let's just try this prompt out. So this is Composer 2 Fast in Cursor. Let's see. I mean, have you ever seen this thing rip? This thing's gonna be ripped. So last time we did this, we were using Opus 4.6 Fast inside of Cursor, and that did a really good job. Like, Opus right out the get-go did a lot of stuff, and I'm super curious. Basically, take the older version of Kimi, but add the cursor sauce to it, and all the stuff that they do, and just like let it rip right now. So right, I didn't even tell it to use parallel agents or anything like that. I don't know if that would be the equivalent to Swarms, but now it's like, I mean, this thing is just gonna run crazy right now, so yeah, that's what's up. Ruben G in the building, yes, let's compare, high, X high, limits will be reset soon anyway, low key, I mean, it's hard to automate stuff in the Browser for raw coding. Opus is the go for sure. API pipelines, limit tokens somehow. I don't think they'll ever find a workaround. That'd be interesting. Opus is like, okay, I'm the boss. Browser takes you straight to their servers. Yes, correct. Implementing the design, shared components, organizing with the to-do lists, building files systematically. So here is Composer 2 fast. Like it is fast. We're seeing it rip through the stuff. And so this is going to be an interesting test of Cursor's harness because as you see here, Composer 2 only has a 200k token context window. And Cursor's latest harness is supposed to do compaction automatically, stay on track with tasks. So I'm going to be really curious to see what the results are here, right? Because we already got a thousand lines of code up in a blink of an eye. This model is fast. I'm telling you, this, this, is too crazy right now. Look at this. We're just... This is bonkers. Even my machine's going crazy too. So yeah, we already got 50k tokens. We used the 25k contacts. Swarm that ass. That's what's up. That's what's up. OpenAI is cooking live stream today. Yeah. Have you tried Opus and Codex working together? No, but I think that's where RepoPrompt is going to come in for us. Spud is coming Thursday and Thropic will be so cooked hard. All right, I have some theories on that. I just don't have enough time because I got to be out really, really soon. Have you tried Opus Codex? Yeah, this is the RepoPrompt thing. If you had to give up all paid API today, what would be your go-to self-hosted setup? I think I would self-host GLM5.1, I think. And then Minimax 2.7. Basically all the Chinese models, right? I would have Nemotron 3 Super up in the mix, especially for the agentic, like, searching and everything like that. I have a lot of problems with tool calling right now in Nemotron 3 Super, but I probably would write my own kernels for that and help, like, do some more optimizations. But then like my full-time job becomes like a kernel engineer or something that I don't want to do, you know, like I want to just consume the stuff right out the gate. So that's what's up. I don't know. Spud is a dud. How do you know this already? Gary, do you know something that I don't know yet? Do you? I think Gary is low-key an influencer and he's like up in here pretending to be like a no-name right with like little pink little like non non-face type of thing. And that's interesting thing. Hmm. Hmm. Hmm. I live for China AI. Oh, thank you for testing, taking this question seriously. No, for real. I'm for real. Yeah. I was wanting some. Okay, chill, bro. Chill, chill, chill, chill. Y'all need to chill in the chat. You know how you can chill right now? Smash that thumbs up. Smash that thumbs up. If y'all are digging the vibe right now for the live streams, smash that thumbs up because I'm going to do live streams now every Tuesday and every Thursday. My schedule is changing a little bit. So right now we're doing Tuesdays and Thursdays, 10 a.m. Pacific Standard Time is going to be the time until I get to Hawaii. I'll probably shift around just maybe an hour or two. But I like 10 a.m. It's a really good spot for me right now. And I'm going to be also getting into a school community. So I'm going to get that going in a little bit. So stay tuned for the links for that. I'm going to be dropping those. That way we can kind of participate. I like the forum style of things. It's easier to go back to posts and reference things as opposed to right now on the Discord. Like I have a whole bunch of really great content in there, but it gets kind of lost all over the place, especially when people have conversations. So you still be able to have conversations, you be able to do that stuff, but right now we're watching Composer 2 fast rip. It is ripping inside. 3,000 plus lines of code already inside of Cursor. It is just ripping. It's still on the, I think, the first task. Holy guacamole. This thing is ripping. Mama, I'm going to school. Yo, Coriallis, what's going on, man? Thank you so much for joining the stream. Yeah, 10 a.m. is fire, I love 10 a.m. Opus 4.7 is 7.5x for. No way. That's what they do in co-pilot. They just basically said, you still get 1x, but then it's like 7.5 for Opus. Sheesh, yeah. SolarWinds, what's going on, man? Where do you see Quen ranking? Oh, okay, so I forgot about baby Quen. I call him baby Quen because he's so packed full of power. Definitely got to get Quen in the mix. I think right now I'm trying to use Quen a lot for like parsing PDFs and just doing some like regular, here's a very well-defined task, go do it. And then, you know, verify your work type of thing. And I'm still trying to figure out the shape of Quen in terms of like repeatability there. But yeah, I'm working also on a project that deals a lot with like local inference and stuff. So stay tuned, stay tuned. Yeah, I'm trying to get some of these models, especially for like PDFs, like parsing and stuff like that. Quen, Quen, Quen, Lambo, bro, Quen Lambo. The issue with composer is fast as hell, but yeah, we'll see. I mean, I'm really curious. Corialis, it is, it is actually just vibe slop. I don't know. I don't know. We will find out. Build the Ray model, maybe, maybe, nah, nah, nah, nah, I'm good, I'm good. Wait, are we done already with Composer? Sheesh. Okay, so Composer just finished, and, let's see, let's see, let's see, okay, 3,000 lines later, 3,467 lines later, sheesh, we already finished. Can I load the website? Let's see. What's going on? Damn! Okay, so all I have to do to run locally is just do this. Command-J, I think, yeah, I'm already in here. Bun install, Bun dev, okay, moment of truth. Composer, okay, it got the filtering, okay, okay, okay, let's see, let's see. Can I click a blog post? It takes me to a blog post. Yo, Composer is, okay, and it has a little scrolly thing on top. Let's go to the About Mage. Let's go, let's go to like MobileView. Cursor, Cursor, where have you been, bro? Where have you been? Cursor. Wow, Composer like, and it's fast as hell. Oh my goodness, I need to talk to the Cursor team because they basically extended Composer to be free for like a week or something like that. Not Composer, but like limits. Let's see. Cursor Composer. I need to tweet at them. Let's see. Limits. Where is it? We're doubling the usage through the end of this weekend. Okay, so like, okay, this model absolutely cooks with Kimi. Okay, this model cooks. Cursor team, can we get another doubling this week, please? I like this model compared to Kimi K2.5-6. Okay, yeah, if they reply to this tweet, y'all saw this here, we're asking them live, I need you to like spread this news out further because I would take the Composer 2 model over Kimi K2.6 for a lot of reasons. One is it's super fast with Composer 2, right? Like the level of detail that we got there was crazy, right? Yeah, like and then let me just link the live stream studio next to YouTube. Hey Fernando, I gotta link the stream, cause like we're literally just, we cooked fam, we absolutely cooked. C stream for results. I think I'm gonna put this in, put this in the seconds for results. Okay. I'm going to post it here. Yeah. Post all. Yeah, we need this. This needs to be like immediately. They need to amplify this and like Cursor's team needs to put this thing on like double usage again. They just need to just let it rip. Bam. Retweet this. Do the whole, do the things. Do the Twitter things. Yo, this model rips! What? Oh my goodness, bro! 4,000 lines! In like that. Ugh. Hold up. I have a lot of bias right now. Safari. Loco Host 3000. What? See, it got mobile mode. It got the mobile mode correctly. Look at this. Kimi K2-6, no, not even the Agent Swarm, Agent Swarm didn't even get this. I got to the streams about, wow, I got the system tweak panel, I could do magenta, it doesn't look so hot, you know, with this little drop down, but, oh my goodness, let's go, let's go, what? What? Look at this. If I do editorial, it got this. If I do list layout, it got the list layout. Magazine style. What? I go light mode. Okay, you cooked on light mode too. Let flipping go. What? I would choose different colors, but still, still. So you can even see all the stuff up here, right? For the different fonts. Let go back to dark mode. About, let's go to Aboot, Aboot, for my Canadians. Oh my goodness, can I, I even can click a blog post. Get out of here, bro. What is cursor on right now? What the fuck? 100% harness matters. Composer made. Yeah, it, yeah, yeah. Oof, yeah, my bad, my bad, my bad. Have we tested Kami? No, no, no, bro, I'm not even gonna go to the CLI right now. We tested it on the web, like their own platform, directly from their own service, on their own China servers and everything with their own inference, and we did that. We basically did Agent Swarm, which probably runs their own CLI and their own skill set inside of their own thing, and we didn't even get close, bro. It didn't even touch mobile view. Mobile view wasn't even there. Like we got no mobile view. It was just nerfed. It only got us like 70% of the way there, maybe 65. Cursor one shot, bro. 4,200 something lines in less than like 5-10 minutes. We got the whole shebang. Dark mode, light mode, all the settings tweaks, all of the details inside of the prompt, even the actual blog posts themselves. We got the streams thing here too. Like we got it all. We got this. We even got mobile. Come on, bro. Come on, bro. Try their CLI. I ain't got time for that right now. We don't got time for that right now. We just gave them a whole bag. We gave them like Droid. We threw Droid at it, right? No, it didn't work. Kimi, I'll take the older model inside of Cursor with their sauce and everything like that. Can you try Kimi? Is Kimi 2.6 in Cursor? No, it's not. I don't think it is. Cursor really cooks with Composer 2, I'm telling you. Cursor just cooked it, bro. It's cooked. The latest was Cursor Composer. Yep. Yep, yep. Ray tried the Joyed CLI with Kimi and it didn't do as good as well. Yeah. Nobody got time for the CLI, bro. Exactly, bro. They have to earn my respect. They didn't earn the respect. They earned the respect of you. I feel like we just dabbled in all these different things here and it's low-key. I feel like Kimi K, I don't know where they got these benchmarks, but it is benchmarks. I mean, straight up, right? No disrespect, but whenever someone says no disrespect, they're disrespected, right? I feel like I feel like I've been disrespected because whenever I see benchmarks, that's why we do these things live. We do these things live because we want to see what's going on with these real results. But your favorite YouTubers probably aren't taking the time to run all these comparisons. And we did this whole thing live in what, less than an hour? Come on, fam. Come on, fam. Get to work. Get to Work, Fam. Yeah, seems like overkill for a blog page, but yeah, I probably do. That's what I'm saying, tech friend, like a blog page, right? Like we're like have, I think this is the perfect prompt for benchmarking in some ways because there's so much detail and so much richness in these types of things and it's small enough for a context window that like a 200k token context window with Composer 2, right? Cursor has a 200k token context window, right? I can't give it like this mega crazy prompt. It has to be very succinct with a lot of detail and that really just goes to show how well the team at Cursor has really fine-tuned that model to read through those details, to work with their harness, to make sure all those little details are included. I'm shook right now. And it's based off of Kimi 2.5. So when people say like, oh, it's just the Chinese model, da-da-da-da-da, we're seeing Kimi K2.6, which gets better benchmark scores, can't even do like 60 to 70% of the tasks that we give it from the same prompt in their own website. And then when we send some swarms at it, bro, I'm telling you right now, I'm telling you that when people make those like bad faith comparisons, have they tried it? Like show me your prompt, show me your prompt and I'll show you your results, right? Like show me your prompt, it's going to show me your thinking and show me your level of detail. Like do you really care about the levels of details? We care about that here. This is what we do. Everyone in this chat is about that life right now, right? So we care about those details because we want to figure out if we can use this for our work or not and see what it's going to do for us. Mind you, I probably shouldn't be the only one doing this and you should be doing your own tests too. So I want you to like take some tests, let me know what you've been cooking with and that's why I want to get the community up and running. So I'll get that community up and running so that we can start posting our results and so it's not just going to be me, myself and my own echo chamber. We'll have like a little forum talk and then we can have like a weekly discussion on like who's using what, what use cases, what models are working pretty well for certain things. So if you want to stay tuned, definitely take a look at that. I got to head on out and then we're going to have to check out some of that news with OpenAI, drop in a little bit later. But as of now, if you're thinking about Kimi K2.6, it's hype, bro. Just bench maxed. On the website, L. On the Agent Swarm, L. On the inside of Droid, L, bro. Just L, L, L, L, L, L, L, L, bro. Bruh, Bruh, Bruh, I can't. But, you get Kimi from last year, Kimi 2.5. And then you get some American sauce on it, with like the Composer 2. And put those little cracked kids on it, from Cursor. Sheee! That's all I gotta say, y'all. Peace out, y'all. See you soon. www.claw.com.au