📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

AI News: 18 Breaking Stories You Missed This Week

Matt Wolfe28:10

Transcription

As with every week, a ton of stuff happened in the world of AI. And here's what you probably missed.

On Friday of last week, we got the long-awaited brand new model out of DeepSeek, Deepseek V4. Deepseek V4 is nearly state-of-the-art. It's open-sourced, and it's got a 1 million token context length. Now, here's the thing about this model. It's not quite as good as the state-of-the-art models, right? These are well kind of the last generation of state-of-the-art models. Here we now have Opus 4.7 and GPT 5.5. So, it's kind of comparing it to the last generation of models, but we can see it's really dang close in almost every single benchmark. It's right up there with the state-of-the-art models in math, right up there with state-of-the-art questions and answers. And if you look through all of the benchmarks, it's just really, really close to these state-of-the-art models here.

Now, this is important again because this is open-source open-weight models. Now, they're still a little bit too big to run on consumer GPUs, so you're still likely going to use them on DeepSeek's cloud or some other cloud service. However, the biggest differentiator and the reason all of the big model companies sort of freak out about this stuff and when models like this come out, it sort of has a little bit of a shock on the stock market is because of this pricing down here. This is basically almost as good as the best models out there. It's got the 1 million token context window, so you can put in a ton of text and receive a ton of text back from it. And it cost a $1.74 per million tokens input and 3.48 per million tokens output. If we compare that to GPT 5.5, GPT 5.5 is $5 per 1 million tokens input and $30 per 1 million token output. And even GPT 5.4, 4. It still looks pretty cheap compared to that. 250 per 1 million token input, 15 per 1 million token output. And this model is roughly as good as GPT 5.4. If we compare it to Claude Opus 4.7, $5 per million token input, $25 per million token output. And if we look at Gemini 3.1, that one's $2 per million token input, and $12 per million token output. But then it gets even higher when you increase the amount of tokens.

With that in mind, looking at their pricing, companies are able to get nearly the same capability at a significant discount. And because these models are open, they could theoretically put them on their own supercomputers in their offices and run them locally and have all of the security and safety and privacy benefits of having these local models. So, these open models are catching up and they're also being trained for a lot cheaper than the models here in the US are being trained for because due to export restrictions over to China, they're not able to get as powerful of GPUs to train these models on. So, they're finding more efficient, more cost-effective ways to train these models that are nearly as good as the models here in the US. They're making them open weight and then they're drastically undercutting the price of using these closed models. So every time Deepseek puts another one of these models out into the world, it kind of freaks out a lot of the Frontier Labs that are based here in the US, as well as typically has a little bit of an impact on the stock market as people realize that a lot of the spending that's going into AI here in the US maybe doesn't need to happen as excessively as it is because companies in China are figuring out how to get similar results for a lot less expensive.

But as these open models get better and better, bigger like enterprise companies that are spending insane amounts of money on their token costs to use these models, well, you might start to see more and more of them pivot to the open models for the security and privacy reasons as well as for the just insane cost-saving reasons. And while we're on the topic of open models, Nvidia released another open model this week as well called Neotron 3 Nano Omni model. Now, this model's Omni because it has vision, audio, and language, and it's designed to work really well with AI agents, which makes sense because Jensen's constantly talking about OpenClaw and how huge of a deal OpenClaw is to the world. So, Nvidia's building models that are much more efficient for that kind of use case. But unlike the models that most people are using for their agents, which are anthropic and open AI and Google models, this new Neotron 3 Nano Omni is an open model. So again, this is a model that can be run locally. If you are running it locally, the cost to run it is essentially just the cost of the electricity for whatever computer you're running it on. And it's capable of handling text, images, audio, video, documents, charts, and graphical interfaces as input. It's even capable of being run on a little DJX spark box.

A few years ago when we were talking about all these latest models coming out from Anthropic and Google and OpenAI, the general consensus was these open-source models, these open-weight models are pretty much never going to catch up. But I feel like we're getting to a point where they're getting really really dang close. And well, the Frontier models are so good now that most people's use cases don't even need the state-of-the-art anymore for most of what they're doing. If you're using it for like document summarization or you know finding correlations and patterns within data or helping you explain certain things or you know little AI agents for customer support, the best state-of-the-art models are almost overkill now for that. And the cheaper and open models are actually fairly good at that, but a lot less expensive to use. And if privacy and security are more of a concern for your company or what you do with your business, well, you can actually get devices and run these locally now. Pretty crazy to see how much and how quickly open-weight models have caught up with state-of-the-art to the point now where they're pretty usable for most people's use cases.

But let's keep this discussion on open going even further. This company, Poolside AI, here just released Laguna XS2 and Lagona M1. These are two new foundation models and the Laguna XS2 is an open-weight 33 billion total parameter model. Both models are currently free to use. So their M1 model is 225 billion parameters and it looks like it's designed to compete with some of these larger models here, but not necessarily the most state-of-the-art models. Their XS2, the one that they are releasing as open weights, appears to be designed to compete with like Gemma 4, Devstrol Small 2, Claude Haiku, models like that and seems to fall in line with the rest being pretty in the middle of all of them. The company Mistral released a new model this week for remote agents in Vibe and it's powered by their Mistral Medium 3.5 model. Now, this one's clearly designed to work inside of agent harnesses like OpenClaw and Hermes and things like that. And it merges instruction following, reasoning, and coding into a single 128B dense model. And it's also released as open weights.

Now, if you want to dive into the benchmarks and go a little deeper on what all of these models are really good at, I will link up every single one of them in the description so you could find out which model is right for you. I've just been really really impressed with all of these open models coming out lately that are like just that close to being as good as the state-of-the-art models. Alibaba also released a new model this week for images called Quinn Image 2.0 Pro. Now, on the Arena leaderboard here, it's ranked ninth. So, I'm not going to get too deep into this one just because I mean we've got plenty of image models that do quite a bit better than this one. But if I skim their thread here on X, we can see the types of images that it generates here. All of these were generated with this model. And it can do text and it can do memes apparently.

And here's another model that actually came out last week, but I missed it in last week's video, so I wanted to make sure I touched on it in this week's video. And that's the new Grok Voice ThinkFast 1.0 model from XAI. This is a new voice model that's designed to be really, really fast, have low latency, so when you ask it a question, there's not a ton of delay from the time you ask the question to when it responds. It just gets back to you really, really quickly. Stuff like this is kind of designed to be for like customer support. You can have your phone agents be AI and they'll actually respond a lot quicker than sort of having a delay between you asking it that question and it actually giving you that response.

"Yeah, I'm seeing a $65 charge for activation, but I haven't received my kit yet."

"Thanks for letting me know, John. Let's get this sorted. First, I'll send a six-digit code to your email address for verification."

"Okay, I see it in my inbox. The code is G4XM21."

"Got it. Let me verify that code for you. All set. One sec while I pull up your account's information. Okay. I'm seeing a charge for $65.10 for activation of the kit you ordered on April 20th. It looks like the kit is still in transit and expected to arrive on the 24th. Is that the one?"

"Yep, that's it."

"Great. Let's get this sorted."

You get the idea. It's pretty dang quick. That was all in real time. So, you can hear that this was a real person. And as he responded, the AI responded at a very natural sort of pace. It's also now what powers Starlink's phone sales and customer support experience.

I haven't used this animation in a while, but we've had a ton of drama in the AI world this week. So, let's break some of that down. Starting with the soap opera everybody's kind of paying attention to right now, and that is the legal battle between Elon Musk and Sam Altman. So Elon sued OpenAI, accusing them of abandoning its founding mission of developing AI to benefit humanity and shifting focus to boosting profits instead. That trial finally started this week. Musk, who is an original co-founder of OpenAI, claims that Altman and Greg Brockman tricked him into giving the company money only to turn their backs on their original goal. OpenAI is claiming that this is just a bid to boost Elon's own SpaceX, XAI, and X companies that have launched Grok, which is a competitor to ChatGPT. Musk wants Sam Altman and Greg Brockman removed from the company and to stop operating as a public benefit corporation.

I've been following along to what's going on with this trial over on this Verge Realtime Update page, which you know will be linked up. Every like 15 minutes or so, they're posting an update of what's been going on in the courtroom. And from what I've gathered so far, Musk has not been looking too great. They put him on the stand and during his testimony, he sort of refused to answer yes or no questions. He sort of argued with the attorneys and from my understanding has not really made himself look too good. There's obviously a lot of detail, so I'm not going to get into all of them. This particular writer for The Verge says, "After about 5 hours into Elon Musk's testimony, I typed the following sentence. I have never been more sympathetic to Sam Altman in my life. For hours, Musk refused to answer yes or no questions with yes or no. Occasionally forgot things he'd testified to in the morning, and scolded defense lawyer William Savit. I watched a few jury members glance at each other. During one testy exchange, one woman was rubbing her head." So, yeah, most of the trial has been them examining Elon Musk. From what I can gather, it doesn't sound like he made the best impression, but there's still quite a bit more to come from this jury. And if there's anything interesting, I'll probably bring it up in some future videos.

In some Anthropic drama, a lot of people are getting really, really frustrated by Anthropic's latest business practices. And that's the practice of them charging people more whenever they expect that they're using a harness that Anthropic doesn't want them to use. So, for example, Hermes or OpenClaw. Here's a tweet from Theo Brown. "Fun fact, if you have a recent commit that mentions OpenClaw in a JSON blob, Claude Code will either refuse your request or bill you extra money. This is an empty repo. I'm just calling Claude Code directly." Here's another case as shared from Patel here. Some guy was on Claude Max 20X, the $200 a month plan. Yesterday, Claude Code hit him with a "you're out of extra usage" out of nowhere. His dashboard showed 13% weekly usage, 0% current session, 86% of his plan was sitting there untouched. But he got $200.98 in extra usage, which should have been covered by his subscription plan. He started binary searching repos and commits mainly on his own until he found the trigger, the string Hermes.md in a recent git commit message. He then reported it. Anthropic support acknowledged the bug three times, called it an authentication routing issue, thanked him for finding it, and then refused to refund the money. Here's the original message that he was referring to over on Reddit. You can see in the code here, you've got adermes.md. When he reported to Claude, here was the actual response. "I sincerely apologize for the disruption you experienced with the billing routing issue. We take service reliability very seriously. However, I need to let you know that we are unable to issue compensation for degraded service or technical errors that result in incorrect billing routing."

So, yeah, essentially you can have a subscription plan with Anthropic. you know, pay $100 or $200 a month for one of the Claude Max plans, use it for coding, and if your code happens to mention OpenClaw or Hermes or one of these harnesses that isn't Claude Code or Claude Co-work, they're going to assume you're trying to use one of these other harnesses and either not let you use their product anymore or charge you extra. Like, that's kind of screwed up. Now, there is a slightly happy ending here. We can see Tariq here from Anthropic says, "Sorry, this was a bug with the third-party harness detection and how we pull git status into the system prompt. We're reaching out to affected users and giving them a refund plus another month's worth of credit." So, Anthropic did decide to actually refund its users after this sort of became a big deal. But I do wonder if Anthropic would have had this response had this post from On Patel here not gotten 1.4 million views and this post from Theo here getting 1 million views. They probably realize pretty quickly this isn't a good look. We're looking at people's code to try to analyze what they're using our tools for and then blocking them when we don't like some of the words that we find inside of the code. But there's some replies here that I kind of have to agree with. It's a good thing Anthropic tells us how ethical they are. Or we might worry about their morality, right? They're the ethical AI, the one that has all the red lines for the military and whatnot. Yet, they have practices where if they don't like that you have the word Hermes or OpenClaw in your code, they're going to go and then charge you extra for that or just refuse to like let you use their app. Or Theo's own response here on M's post, "There's a certain class of bugs that suggests the thing you're trying to do is a bad idea. Worth reflecting on that." The implication is, hey, maybe you shouldn't write it into your code that we should be blocking any time some of these keywords show up in the code itself. The fact that they're actively looking for things like Hermes and OpenClaw in the code is kind of the thing that people are taking issue with in the first place.

Anyway, I could rant on this for a while, but I won't cuz we even have more drama. There was recently a ton of drama around Anthropic refusing to get rid of their red lines and work with the Pentagon. And then OpenAI stepped in and said, "Hey, we'll work with you." But seemingly drew the same red lines. Well, Google just signed a deal with the Pentagon now for it to be used on classified information, and their agreement basically says you can use it for any lawful government purpose. Now, they do go on to say, "We remain committed to the private and public sector consensus that AI should not be used for domestic mass surveillance or autonomous weaponry without appropriate human oversight." So, they're basically saying, "We don't think you should use it for that, but you can use it for any lawful purposes. There's no like binding agreement that says you can't use it for those purposes. We're just saying we don't think you should." And this is likely to get them some backlash both internally at Google and externally because prior to this information coming out that this deal was happening, a bunch of Google employees asked Sundar to say no to classified military AI use. Over 600 Google employees signed a letter to the CEO demanding that Google block the Pentagon from using its AI models for classified purposes.

This is especially interesting because back when Google bought DeepMind back in 2014, there was actually an agreement between Google and DeepMind that they wouldn't use it for these purposes. When Google acquired DeepMind in 2014, DeepMind's founders secured a promise from Google that DeepMind's AI would not be used for military or surveillance purposes. DeepMind leadership only agreed to the acquisition after extracting a commitment that their AI technology would never be used for military applications or surveillance uses. This understanding was framed as a core condition of the deal and part of DeepMind's mission around ethical and responsible AI. So the fact that they even signed this agreement at all feels very counter to the agreement they made with DeepMind in the beginning.

Now to be fair to Google, I do think they're in a kind of no-win situation. All of their options are sort of bad. They either say no to a deal with the Pentagon and refuse to let their technology be used for whatever the Pentagon wants to use it for, but then they risk being sort of ousted by the Pentagon like Anthropic was, where Anthropic was basically deemed a supply chain risk. These AI companies are also needing legislation to sort of go in their favor to continue pursuing AI advancement. So they don't want to make too much waves with the government that's going to sort of create the laws around all of this stuff. So that puts them in a bad situation if they decline to work with the government. But then on the other hand, if they agree to work with the government on this stuff, it looks bad because well, their original agreement with DeepMind said they'd never do that. You've got 600 plus employees over at Google basically begging Google not to do this. And then you have the optics of what went down with Anthropic and OpenAI. It looked really, really good that OpenAI was kind of this hero that was putting a line in the sand against the government and a lot of people really appreciated that Anthropic took that stand. And then you had OpenAI come in who looked very opportunistic and said you can use our AI for that essentially. So Google's options were bad on both sides and they apparently are going for the option that's like let's make sure we don't get in the bad graces of the government cuz that hasn't really worked out too well for Anthropic so far.

And in the final bit of drama from this week, China blocked Meta's $2 billion acquisition of AI firm Manais. So basically, China is saying we need to keep our technology here in our country. And this is going to get really, really messy because from what I understand, Manais has already sort of like integrated itself into Meta. Now, Manais founders got their start in China, but relocated their headquarters and key staff to Singapore in 2025. Manais was Singapore Incorporated with founders based here, and it still got pulled back. It's still pretty unclear how Meta is going to unwind the deal. Meta's employees have joined Meta. Capital has been transferred and the startup's executives have joined the US firm's rapidly expanding AI team. Manais staffers have already moved into Meta offices in Singapore while investors have already received their proceeds. So, this is all sort of still developing. I don't know how this is going to play out, but China's basically demanding that this all gets unwound. And I don't know how that's even logistically possible at this point.

All right, I have a handful more little things I want to share that came out in the past week, but I'm going to break them all down quickly. So let's jump into a rapid fire.

This week, Microsoft and OpenAI once again restructured their partnership. There used to be a term in the partnership where once OpenAI reached AGI, then the deal would sort of be gone and Microsoft would no longer be entitled to profits from OpenAI. But that clause has been removed. So there's no sort of strings attached around AGI anymore. Instead, there's just an end date. Microsoft will continue to have a license to OpenAI IP for models and products through 2032. However, Microsoft's license will now be non-exclusive. That's kind of the main thing of this whole new deal. Microsoft will also no longer pay a revenue share to OpenAI. So, the money is not kind of flowing both ways as much anymore. But this non-exclusive thing is kind of the big deal. Now, OpenAI doesn't have to be exclusively on Microsoft Azure servers. And in fact, this news came out on April 27th. And the very next day, April 28th, OpenAI models, codecs, and managed agents come to AWS. So, as soon as that sort of unwrapping happened, OpenAI went and made a deal with Amazon to allow OpenAI models on AWS. Don't be surprised if you see OpenAI models popping up in more and more places. I mean, who knows? Maybe pretty soon you'll be able to use OpenAI models directly from within Google as well.

Speaking of Google, they rolled out a new feature inside of Gemini this week. You can now easily generate files in Gemini with just a prompt. Gemini can now create PDFs, Microsoft Word, and Excel, Google Docs, Sheets, Slides, and more directly in your chat. Meaning you can quickly move from a brainstorm to a complete file without ever leaving the Gemini app. So the new file types it supports PDF, DOCX, XLSX, CSV, LaTeX, plain text, rich text format, and markdown. So now you can basically upload an image of something and say take the data from this and convert it to an Excel spreadsheet or turn this into a PDF for me or convert this to markdown for me and it'll just quickly do that.

A new feature rolled out in Google Translate that actually now helps you pronounce the words. One of the toughest things about speaking a new language is getting the nuances of pronunciation just right. Now you can get instant feedback with pronunciation practice which uses AI to analyze your speech and help you improve. Right now it's launching in the US and India in English, Spanish and Hindi with more to come. So, that'll be helpful if you're trying to learn a new language using Google Translate.

Staying on Google for one more sec here, Google Photos launches an AI try-on feature for clothes you already have. So, this new feature actually looks at all of the images in your Google Photos of you to see what clothes it's seen you wear and then helps you generate new outfits based on clothes that it knows you own because it saw them in your Google Photos. The feature is rolling out to Android devices later in the summer and then is going to expand to iOS. Google I/O is coming up in a couple weeks, so I'm assuming we'll probably see this feature talked about a little bit more there.

This week, 11 Labs launched 11 Music, which is a new platform to discover, remix, create, and earn from music built on the 11 Labs music model. So, you can explore songs that other people have created on here. Listen to the songs >> and if you like what the song is, you can remix it and give it a prompt to change it more upbeat, slower, darker, lighter, or just, you know, describe how you want it changed and it will change the songs that are on here.

And while we're on that topic, there's a new Verified by Spotify badge that lets you know the artist isn't AI. Some artists will now have a Verified by Spotify badge and a green check mark on their profile indicating that the company has confirmed a real person is behind the music and the profile. AI personas or profiles that primarily upload AI-generated music are not eligible for the verification program. It did leave the door open to the possibility in the future though saying the concept of artist authenticity is complex and quickly evolving. So whatever that means. But for now if you see this green check mark you know it's not AI. But that doesn't mean everybody who's not AI will have a green check mark. So it's still going to be kind of confusing.

Honestly, if you do any sort of marketing, X is launching a new advertising platform. I've tried using X's old advertising platform and it's not great and I never really got great results from it. I don't really think people are clicking on ads as much as they used to on X. But the new ads manager will have an intuitive AI-powered platform to make campaign creation effortless and precise, giving advertisers full control and speed and superior AI-powered performance for stronger results, faster optimization, precise targeting, and AI-powered relevancy. The new systems understand user behavior at a much deeper level, enabling more precise, relevant, and dynamic ad delivery that aligns with what's happening on X in real time. The shift toward AI-driven, contextual, and semantic advertising will bring meaningful lifts in relevance, engagement, and ROI. I'm constantly looking for new ways to grow my Future Tools newsletter. I've tried advertising like this in the past. It didn't really work very well, but I'll probably give this new ad platform a try and see if it helps grow the newsletter at all, and I'll report back once I do.

And in the last bit of news I want to share this week, a Mayo Clinic developed artificial intelligence model can help specialists detect pancreatic cancer on routine abdominal CT scans up to 3 years before a clinical diagnosis. It identifies subtle signs of disease before tumors are visible when curative treatment may still be possible. Now, pancreatic cancer is one of the most deadly forms of cancer and if people are able to spot it 3 years ahead of time, that's pretty impressive. From my understanding, they went and sort of back-tested this model. They plugged in a bunch of scans from people who had pancreatic cancer and they plugged in scans from before they were able to tell they had pancreatic cancer. And the model was effective at spotting these old scans to tell them that these people would eventually get pancreatic cancer.

I love sharing stories like this when I come across them because well, this is what we all want AI for, right? We want AI to cure diseases, help us detect cancers way earlier. These are the real promises of what this AI is going to do for us. Whenever anybody talks negatively about AI or says we should just stop pursuing AI altogether, well, the counterargument is always, well, it's going to help us cure diseases and solve climate change and find alternatives to fossil fuels and all of that kind of stuff. Stuff that will actually sort of benefit humanity. Yet most of the time what we get is like, "Hey, making Excel spreadsheets is a little bit easier now." Or, "Look, you could try on different clothes that you have in your Google Photos, or you can make this animation of Mickey Mouse punching Taylor Swift in the eye or whatever." We need more stories like this to come out, like actually detecting cancers earlier and curing cancers and solving real actual problems. I love this kind of stuff. I'm hoping to see more and more of it.

But that's what I got for you today. Again, not a ton of like really really huge news. A lot of sort of little things, but my goal with this channel is to help you feel super looped in with what's going on with the AI news. Every week there's probably 300 plus announcements that come out. I drink from the fire hose. I consume all the announcements. I try to learn what they mean and the implications of them and then try to filter through the hype and all the noise and tell you what I think most people will want to know or hear about, the stuff that's actually moving the needle somewhere. My goal is to minimize overwhelm for people and just bring to you what I think is actually meaningful or useful. Hopefully, I did that in today's video. I do record these on Thursdays, publish them on Fridays. So, if there was anything that came out on Friday that I missed, it will likely be in next week's video. If stuff like this interests you and you want to make sure more of it shows up in your YouTube feed, make sure you like this video, subscribe to this channel. It really, really does help me out. Thank you so much for tuning in, nerding out with me today. Hopefully the week doesn't feel as overwhelming and you feel more looped in on what's actually happening in the world of AI this week. Thanks again. Hopefully I'll see you in the next one. Bye-bye.

Do you mind calling the cannons a try? You're going to give you day sure. All right, we are live. Welcome back to the stream everyone.