Transcription
Here's the AI news you probably missed this week. In fact, for this one, let's change up the scenery a little bit. So, there wasn't a ton of AI news this week, but there were some really, really big stories that we need to talk about. Starting with this new launch out of Sakana AI, they released something called Sagana Fugu or Sagana Bugu. So this Sakanafugu model, it's not just like one new model. It's not like the new GPT model or the new Claude model. This is actually an orchestrator model or a manager model. A model that is in itself its own AI model, but this model figures out which model to route your prompt to and sometimes even routes it to multiple models. The beauty of this is if one model provider ever goes down, say for instance, one of the OpenAI models or one of the Anthropic models that it uses under the hood, well, it's just going to route it to a different model instead, and you're not going to lose any progress on whatever project you were working on. They launched this model with two variations, Fugu and Fugu Ultra. The regular Fugu model balances strong performance with low latency, making it a great default for everyday work. This one's going to be faster. It's going to be a little bit cheaper, but it's not going to get you the most insane output you've ever seen. And then they also have Fugu Ultra, which is tuned for maximum answer quality on hard multi-step problems. So this is the one that you're going to use when you really need it to think things through, work for a long time, and you know, solve complex stuff. And here's what's really interesting about this model. If we look at the benchmarks here, it actually beats Fable 5 and Mythos in some of the most important benchmarks. So in live codebench, Fugu and Fugu Ultra actually beat Fable 5 in Google proof questions and answers. It's better than Mythos here. It's on par with Fable 5 in Scode. Not quite as good on SWEBench Pro. But on most of these benchmarks here, you can see that this Fugu sort of orchestrator system is actually performing as good, if not better than Fable, the previous model that we had that was the best out there.
Now, as you've probably noticed, I'm not at home, but I did play around with Sakana AI and tested a little bit at home before I left to come on this trip, and here's my findings. If you want to test out this Sakana model for yourself and see how it compares to Fable, here's how you can do it. If you head over to console.sakana.ai, you'll need to create an API key. So, go to billing, add some billing details. I just set it up as pay as you go for right now cuz I'm just testing and I threw $20 in. But once you've got your billing information, you can generate an API key. Then, if you click on their get started button, it will walk you through how to set it up here. The easiest way is to use the codeex CLI. So if you have codec installed on your computer already, you're going to simply copy this command here. Open up your terminal, paste in the command, and it will install Sakana for you. Now I just need to type this little codeex fugu command here. CEX- fugu. And now I'm in the OpenAI codeex CLI using Fugu here. And if I want to switch the model, I can type slashmodel and change it to Fugu Ultra if I want because I want to see how it compares to Fable on something like making a Mega Bonk clone. I'm going to go extra high. Create a clone of the game Mega Bonk. And it's probably going to work for a while. While it's doing that, I'm going to open up a second tab here in my terminal window. Get codeex Fugu going on this one as well. And this time I'm going to ask it to make a working clone of the Future Tools website, but a little bit better and just see what it builds for us. I'll just say, "Build a working clone of the website at futuretools.io. Improve the site's UI and UX as you see fit." Once again, Fugu Ultra Extra High, just full-on yolo mode. Let's see what happens. And after about an hour, it finished both of my apps here. It created Bonk Storm and a Future Tools clone. Now, before I show you what it created, let me jump back to my Sakana console here. And it took 22 million input tokens and 210,000 output tokens. And when I look at my billing, well, it went through my initial 20 bucks and then used another 10 bucks after I topped up. So, these two things cost me $30.
[laughter] So, maybe don't use Ultra on extra high [laughter] right out of the gate. But, uh, here's what we got.
So, remember my first prompt was to make a Mega Bonk clone. It created Bonk Storm. I like that it built in meta upgrades right out of the box. We've got an upgrade shop where I could update some of the, you know, damage and padding and stuff. And then I can choose my hero. And I've only got one hero called the Mallet Sprout that I can use. But, uh, if you know what Mega Bonk looks like, this is not it. Everything does seem to work though, like the upgrades work when I level up. I just added like a lightning upgrade. But if you've ever played like Vampire Survivor, this is much closer to that game than it is to Mega Bonk. Mega Bonk's like a 3D game. This is more of a Vampire Survivor clone. Like it is a spot-on Vampire Survivor clone. It did it well, but I was hoping for that 3D 3JS look that we got when we made it with Fable. So, if I'm comparing the Fable version of my Mega Bonk clone with my Fugu version of my Mega Bonk clone, Fable actually did a lot better job with the graphics. It made it 3D looking. It used 3JS. I was able to change camera angles and rotate around the character, but it didn't dial in the level-ups and the meta upgrades and give me multiple characters. And I even think like the menus look better on this game. But the actual gameplay itself looks nothing like Mega Bonk here. It's just it's a Vampire Survivor clone. But if you ever played Vampire Survivor, that game is quite addictive as well. And when you do finally die, if I go back to the main menu, I have my upgrade shop where I could do some of these meta upgrades here. I have 87 silver, so I can upgrade my pocket gravity, but that's about it. Once I get enough, I can actually upgrade and purchase other heroes as well. Like I could be a tin can, apparently.
All right, let's take a look at our Future Tools clones that it generated for us. It looks like it generated the entire site. I wouldn't say it's a better design. I feel like it's probably more of a cluttered, confusing design, but you can see it added all of the details about the tools that are available. We could filter by pricing. So, if I do free, you can actually see it does filter our tools down here automatically. If I select, let's say, finance, it filters down to just finance related tools. So, it actually got all of the filtering done. Let's do Matt's pick and see if it's actually my picks. I mean, it looks like it's probably the tools that I have set as my picks right now. So, all of the filtering seems to work. I like that it added this little graph. There's more productivity tools than anything else, followed by marketing tools. You can also short list stuff, which I think is a cool feature, something that I probably should add to my site. Like, let's say I want to short list ChatGPT. I can add that to my short list. And let's say 11 Labs, add that to the short list. And then over on the right, we've got our short list. And if I copy short list, let's see, what does it do? It'll actually paste in the URLs, which is kind of cool. Now, what happens when I actually click on a tool? So, let's click on Leonardo AI. If I click view tool, it just redirects to the tool. So, it didn't make individual tool pages, but it made it so if you click on the tool, it opens to the redirect link for that tool. If I click into AI news, I mean, it's actually up to date as of the day that I'm recording this. You can see it's got the news in here. I don't think the design is better than my current Future Tools design. It's a little more cluttered and a little bit, uh, uh, rougher on the eyes, but it's all working and it links straight to the news. Newly added, we can see the tools that were recently added in here. Glossary of terms. Yep, that's there. Our FAQ page. Submit a tool page where you can submit the tools. That's here. Let's see. Light mode. Oh, I actually like light mode a little bit better than the dark mode on this version, which is rare for me to say, but it cloned the functionality almost perfectly. If I go to newsletter, I mean, it even cloned my newsletter opt-in page here. Now, it made that game and this website for $30, which is quite a bit to spend on something that looks like this. I do feel like Fable would have done a slightly better job, though. Anyway, let's move along. That's Sakana's new Fugu model, which has like model switchers built in for you.
Stop losing money on five different AI tools. There's a new AI platform called Nexos that lets you access a bunch of different models like ChatGPT, Claude, and Gemini all in one place, and it's basically an AI-powered work and productivity dashboard. Instead of paying for each of those tools separately, which could easily stack up to like 200 bucks a month, you pay once and get them all. So on any given day, my ideas and to-dos are scattered all over the place: voice notes, my note apps, meeting notes, emails, Google Docs. And one thing you can do with Nexos is take all of that messy information and summarize it. Pull out the action items, draft any follow-up emails, and build out a priority list for the whole week. Or if you've got a rough outline for a presentation, you can literally tell Nexos what you want in plain text and it'll turn it into a structured slide deck with talking points and even speaker notes. But the feature I keep coming back to is the Node Agent Builder. You don't need any APIs or code. You just describe what you want in plain words, and Nexos builds an AI agent that actually does it all for you. And it works for both one-off jobs and the stuff you repeat constantly. This all works so smoothly because Nexos is connected to the platforms you're already using, like Google Drive and Slack. So, the AI actually has context from your work instead of starting from scratch or needing you to copy and paste everything into your prompt every time. Save up to 200 bucks a month using only Nexos. Head to the link in the description and use my code for 50% off. And that code doesn't last forever, so grab it now. And thanks so much to Nexos for sponsoring this portion of today's video.
Next up, we got an announcement out of Anthropic this week. [cheering] No, it's not that we finally got access to Fable again. They introduced something called Claude Tag. So instead of using like a separate app to run your AI model, well, what this new Claude Tag does is you can just literally tag Claude inside of Slack, just like as if you were tagging another human inside of Slack. You could give it a project to go do. It's going to break down that project into multiple steps. It'll use your company's tools and continue to work on it in the background while you get to go work on other things. So you can tag Claude in a Slack channel that has multiple other members in it and all of the members can see it and interact with Claude and sort of tell it what to go do. It actually learns and remembers. You don't need to consistently copy and paste context in so that it would constantly have to remember things. As Claude actually hangs out in your channel, it learns and it remembers and it starts to understand your company better and better. It can actually take initiative. So, if this feature is enabled, it will actually watch for when it can be helpful inside of Slack channels that you gave it access to, and it will jump in and say, "Hey, I can help with that essentially." And it's also designed to be secure for your business. The admins of your business can decide what it has access to, what it doesn't have access to, and you can very, very tightly control it. Now, to give you an idea of how effective this Claude Tag is, Anthropic claims that 65% of their code is being written using this feature now inside of Anthropic. Now, typically I would demo something like this on my videos, but I'm not right now for two reasons. One, I'm not at home. I'm on my laptop. But two, it's actually only available on Teams and Enterprise plans as of the recording of this video, and I don't have a Teams or Enterprise plan to actually show it off right now, so I can't. But Andre Carpathy here seems to think this is a pretty big deal. He calls this a new paradigm for interacting with Claude that is significantly more in line with all other human activity or-wide. Once you do all of the under-the-hood engineering work to make it just work across tools, integrations, compute environments, memory, security, etc., Claude basically joins the team in a seamless way, you can talk to it just like you would talk to a person and it can help with a very large variety of workloads. He called this the third major redesign in how we use AI, right? The first one was these web chat apps, the second one was standalone apps, and the third one is inside the tools you're already using like Slack.
Now, the third really big thing from this week that I wanted to share with you was this news that came out on Thursday that is absolutely mind-blowing to me, like not in a good way. So, apparently the Trump administration is asking OpenAI to stagger the release of new models over security concerns. Here's the important bit right here. In a Thursday memo, Altman told staff that the government would be approving access customer by customer during this preview period for GPT 5.6. He added that he hoped there would be a more general release a couple of weeks later if all went well. Now, there's been rumors flying around that GPT 5.6 was coming any moment now. Originally, a lot of people thought it was coming this week and then there was rumors that it was going to be next week. And now it feels kind of up in the air when we might actually see this new model. The rumors that we've been hearing also put it on par with like a Fable level model where it is really just that good, which I guess sort of makes sense why they're being more cautious with the roll out seeing what happened to Fable, which to this day we still don't have access to again. Now, they did say that we've made it clear to the US government that this is not a preferred long-term model and will work with them and others in the industry to achieve a more sustainable approach for future releases. So hopefully we don't have to see this for every single new model release that comes out. A lot of people are saying that this could be the end of the sort of wild west era with AI and we may no longer just get random announcements where we wake up on a Wednesday and we have a new powerful crazy model that nobody was expecting. Like those days are probably kind of over unfortunately. But again, this isn't the preferred model. So hopefully that's not the case long term. As of right now, still up in the air when we're going to see Fable again. Still up in the air when we're going to see 5.6 again. But it sounds like the government wants to get more and more involved and basically say when these things can be rolled out and who they can be rolled out to. So I'll be following this story very, very closely cuz obviously it impacts like everybody that's using AI.
All right, those were the three main stories of the week. Again, I told you there wasn't a ton this week, but I do have a handful more. So let's jump into a rapid fire and I'll break them down.
First up, we got a sneak preview of a brand new video model in Seed Dance 2.5. There was a conference in Beijing this week, and during the conference, they teased this new 2.5 model. Now, this isn't a model that I don't think anybody really has access to yet, but it's supposedly pretty good. Here's what we sort of know about it so far. It can generate 30-second single segment native video output. Right now it can only do about 15 seconds. So this is double the length it can generate. It supports 50 full modal reference materials. So text, image, audio, video. You can give it up to 50 different reference assets to work with. That's insane! And they claim it has significantly more controllable video generation and editing. So better prompt adherence, finer control over motion, camera editing, and post-generation tweaks. But again, just expected soon. So we don't know exactly when this new model is going to roll out for us to use. But sounds pretty impressive. I mean, Seed Dance 2.0 is already kind of the most impressive video model already. So, this 2.5 just sounds like it's making the best even better.
We got a little bit of news out of Crea AI this week. They actually released the Crea 2 model as open weights, meaning that you can now download this model. You could run it in the cloud however you want. You can fine-tune it. You can, you know, train your own stuff into it. So, the same types of stuff you've been able to do with Stable Diffusion and Flux models where you can train your own likeness into them if you want or train it on your own art style, things like that. Craya just made their models open weight. So, now you can do that same kind of stuff with the Craya 2 model.
The news outlet The Atlantic created a searchable database of music used to train AI. So, you know, companies like Suno and Udio, they obviously used copyrighted music to train their algorithm to generate new music. And with this new free tool from The Atlantic, you can find out exactly what's in the training data. If you go to theatlantic.com/category/ai-watchdog, you can see we've got a little search box here. And if I search, you know, whatever Blink-182, you could see there's 150 songs in the training data set from Blink-182. One thing that I thought was funny was the YouTuber Moist Critical or Penguin Z0, he actually did a search and like his YouTube videos were [laughter] in here. >> Heaven help us. Lord have mercy on the absent souls of these AI models that are being Clockwork Orange'd having the eyelids pulled wide open to be trained off my videos here. 221 of his videos were trained into this data set which is just very interesting to me. Other YouTubers are in this data set as well. Like if I search Epic Meal Time, there's 212 videos from Epic Meal Time. 25 videos from Ryan George in here, the first guy to ever write fiction. Now, if I search my own name, it doesn't look like AI cares enough about me to train me into it. There's a Matt Wolf here, but yeah, these aren't my content. So, apparently AI is not training on me yet.
OpenAI and Broadcom announced this week that they've been working on their own inference chip. Now, this is interesting to me because they recently did a big deal with Cerebras, who makes inference chips, which make when you ask the model a question and the speed it takes to get that answer back, it works on that part. That's the inference phase. Well, the Cerebras chips already make that insanely fast. And they had this big deal with Cerebras, but it sounds like they're making their own chips, which are going to be competitive to the Cerebras chips that they had a partnership with. They also, you know, use a lot of Nvidia hardware for this. But a lot of the NVIDIA hardware, I think, is more used on the training side than the inference side. So, like I don't think it totally releases reliance on Nvidia because they'll still be training a lot of new models, but it will make their models a lot faster to use, a lot cheaper to use, and supposedly a lot smarter because these are purpose-built specifically for ChatGPT.
And finally, if you like the Meta Ray-Ban glasses, well, Meta just announced a whole bunch of new styles for their glasses. So, you don't have to just get that Ray-Ban style, they partnered with Kylie Jenner on this style here. There's the Meta Adventurer, Meta Fury, and yeah, just a whole bunch of different new, uh, designs. So, if you never liked the various styles that came out in the Meta Ray-Bans, well, they now have a bunch of new styles. According to this article, there is 26 styles across a range of colors, lenses, and frames.
Like I said, wasn't a huge news week, but some really, really big news between Sakana AI, the Claude Tag feature, and the announcement that the government wants to have OpenAI review customer by customer. Like, those are some pretty big deals this week. There just wasn't a lot of quantity of news, just some really big stories this week. But again, that's what I got for you. If you like stuff like this and you want to stay looped in on the latest AI news, I drink from the fire hose all week. I keep up with all of the news that's coming out in the AI world. I keep myself very, very overwhelmed reading it all, spending all day keeping in touch with this stuff so that I can make one video a week, share it with you on Friday, and break down all the news so you don't have to feel so overwhelmed. If you like that [music] kind of stuff, maybe consider liking this video and subscribing to this channel. I also make a lot of AI tutorials and show you how to actually get good use out of a lot of these tools. Those usually come out like midweek where the news videos come out Fridays. So again, if you like either of those styles of videos, uh, make sure you like and subscribe and, uh, that stuff will keep on showing up in your YouTube feed. Thanks so much for hanging out with me in, uh, uh, Big Bear, California in a completely different location than normal. Really, really appreciate you and hopefully I'll see you in the next one. Bye-bye. See is you asking aware since it might even opening screen. [music]