Transcription
Well, if you've been on X at all this week, you've probably seen a feed of Studio Ghibli images and pretty much nothing else. And well, there's a reason for that. ChatGPT just rolled out a new feature that's pretty wild, and it's only one of a ton of huge announcements this week, so let's get right into it.
On March 25th, OpenAI introduced DALL-E 2 image generation. It's the ability to generate images directly in ChatGPT. Now, you have been able to do that with Dolly, but the images haven't really kept pace with a lot of the other AI image generation platforms; Midjourney, Leonardo, Flux, Idiogram—all of those tools seem to have surpassed Dolly a long time ago. But this new model really, really closed the gap with realism, the ability to add text that's actually coherent and really doesn't make very many mistakes, the ability to edit images by just prompting what you want changed in an image. But really, the feature most people have been going absolutely nuts for is the ability to add any style to any image. And that's why we've seen so many of these Studio Ghibli style images. I fell down that rabbit hole myself, making a whole bunch of Ghibli-fied images.
I took this image of me and my wife getting a new vehicle and turned it into a Studio Ghibli image. And like I mentioned, you could have it edit the image with a text prompt, so I said, "Make it brighter and more colorful," and it gave us this version. I took this image of me with a bunch of people at a conference here and Ghibli-fied that image. You can give it two images and have it work both of those images into a single image. Now, it has this weird thing where it will crop off text sometimes, and when I told it to make it 16:9, it actually sort of squished the image a little bit, so there is some funkiness still.
I gave it this photo and said, "Turn me into a South Park character," and it made me a South Park character. Turn me into a Minecraft character. Here's me as a Minecraft character. Turn me into a pixel art video game character. Here's a pixel art video game character. A 3/4s view Voxel art style. Here's the 3/4 Voxel art style. It's really, really fun and addicting to take real images and sort of restyle them into anything you can imagine. I had it create this Venn diagram for me, and it's actually pretty decent-looking. And then I wanted to see how it did at infographics, and well, it made this infographic about how a neural network works. It's a little bit funky; I wouldn't say it's like 100% accurate, but it looks decent. It's a good start, and you can see I gave it a very basic prompt, so with some better prompting, I probably could have gotten this infographic even better. Here's one about how diffusion models work and one about how reinforcement learning works.
Another request I like giving it is, "Convert this into a style of GTA 5," so I took a picture of some Padres players, and this was the GTA 5 style image that it made. I think it looks pretty dang cool. I told it to take the same image and convert to the style of Rick and Morty, and this is what it made, which to me kind of looks more like The Simpsons or Futurama, but similar idea. I took this YouTube thumbnail that I made here and wanted it to put the text "Worth it?" across it with a question mark, and well, the first time it kind of messed it up; it put "Worth it" and then "Worthed," but then when I prompted it the second time, it did it just fine. It kind of changed the way my face looked a little bit; like it looks a little bit less like me now, but still pretty impressive. I mean, a lot of these little kinks are going to get worked out, I would imagine.
Here's another one where I gave it a thumbnail, and I wanted it to say, "Uh, WTF" instead of "Worth it," and it swapped out the text again. It sort of changed my face a tiny bit, but I actually thought it made the image stand out even more, so I rolled with it. I then asked it to take this image and make my mouth closed, and it gave me this brownie-face version, which, as of right now, is still my best performing thumbnail on that video. In fact, the thumbnail on the video you're watching might even have some sort of cartoony image in it, and this is how I did it. I had it take one of my normal little images here and had it converted to Studio Ghibli.
Another really cool thing you can do is say, "Remove the background and make it transparent," and you can see it actually made a transparent PNG image from this one above it. I also gave it this image here and the OpenAI logo and said, "Make this image in the style of GTA 5 and make the OpenAI logo floating above my hands," and it generated this. We're getting to a point where instead of using something like Photoshop or Canva, I could just go into ChatGPT, upload an image, and tell it what I wanted to change and have it make those changes for me.
Here's another one where I gave it this image of me with my hand on my head and the OpenAI logo, and I said, "Make it in the style of Studio Ghibli, add the OpenAI logo, and make sure it's in 16:9," and this is what it gave me. It hasn't quite got the 16:9 aspect ratio right, but it's getting the edits right pretty good. Here's another GTA 5 style one that I did, and another one that's supposed to be GTA 5 style. Here's a Studio Ghibli style and another Studio Ghibli. And then I started having fun where I took an image of me on this side and an image of me on this side and told it to make me high-fiving myself. And then I did the same two images, and I told it to put an explosion in the background. And here's another variation of that.
The final thing that I want to show off was a fun little experiment I did. I actually uploaded this image of myself and said, "Turn it into a 3D video game character in a T-pose," and it created this version of the image here. I was able to pull that image into this Triplanar 3D tool here, and it even rigged it up and made it so I can like animate the guy. Apparently, this is what I look like when I'm jumping.
Now, at the moment, this is available on Plus and Pro plans. Supposedly, it was going to come out on the free plan, but they actually had to delay it because it was getting so much use; everybody was jumping on it this week, making these Studio Ghibli style images out of old memes, posting them online, and bogging down the OpenAI servers, causing OpenAI to actually delay the release inside of the free version. But if you are paying the $20 a month or $200 a month version, you can use it right now. Here's some other examples of ways people have been using it that I found on X.
Alvaro Cus from The Rundown had it create YouTube thumbnails for him, had it create a product ad for him, took his logo and put the logo on a store for him, designed a landing page, redesigned an interior, took this image and it put the Mona Lisa in the background. Using the same image, it also added a different lamp, a movie poster, and of course, had it generate some memes. Riley Goodside here actually had it generate this Wikipedia screenshot. This isn't actually taken from Wikipedia; this is a fake screenshot generated by ChatGPT 4.0 of a Wikipedia article about the screenshot itself, with a copy of the screenshot in the article.
When all of this stuff was flooding social media, people started saying, "Whatever happened to Midjourney? Why is nobody talking about Midjourney?" Personally, I still have a special place for Midjourney. A lot of the videos that first helped me explode this channel were tutorials about Midjourney, and I still have my Midjourney account and use it from time to time. But I guess the Midjourney CEO, on his office hours, was a little bit salty about this release. According to Flowers here, who was on that call, said the Midjourney CEO just said, "DALL-E 2 image gen is slow and bad, and OpenAI is just trying to raise money and being competitive in a toxic way, and it's just a meme and not a creative tool, and in one week, no one will talk about it anymore." I don't think that's true. He also goes on to say, "Yes, he said that all in one sentence."
I actually think my buddy Baval says it best here when he said, "Everyone's saying ChatGPT's new image model is just style transfer; completely misses the point of multimodality. This is closer to an AI equipped with ControlNets, Lora's, IP adapters, and a graphic designer's mind. It's an expert compositor with world knowledge; a creative partner, not a filter. Canva, Figma, comfy UI workflows collapsed into a few prompts and image references. It's legit the closest thing we've seen to a graphic designer API. I'm dropping sketches for thumbnails and getting near final quality. I'm inputting 2D images and shifting the camera in 3D space. I'm literally prompting pixel-perfect text, interacting like I would with a skilled artist. And unlike diffusion models, this system behaves like a design partner with intuition, artistic sensibility, and the technical skills of a pro compositor."
This screenshot says it all. If you were using comfy UI in the past, this is kind of what it looked like. Now you give it a sketch and you get a result, and then you say, "Move the guy to the right," or, "Change the text color," and just by giving a text prompt, you get a brand new image that made the changes that you asked for. It's actually pretty mind-blowing, and I think everybody that thinks it's just AI-generated slop is really underselling the sort of leap that was just made with this. We kind of went from, "You have to be a sort of technical prompt engineer and know comfy UI and know how to use Loras and ControlNets and IP adapters and all of these things to get the image dialed in exactly how you wanted it," and now all you have to do is use ChatGPT and just have a conversation until the image looks the way you want it. It's the UI and the sort of process change that makes this such a big leap, in my opinion.
But I'm going to move on past this because you've probably been flooded with everything ChatGPT this week, and as great as it is, it probably wasn't even the biggest leap that we saw this week. OpenAI has this way of sort of overshadowing anything Google does. Whenever Google has a pretty decent launch or some new product, OpenAI almost always follows right behind with some other big announcement, and that's exactly what happened this week as well, because Google released Gemini 2.5, which is their most intelligent AI model. And not only is it their most intelligent AI model, it is just kind of hands-down the most intelligent AI model on LM Arena, a sort of user-blind voting leaderboard. Gemini 2.5 pretty much blew everything away, and then all the other benchmarks around things like science, mathematics, code editing, visual reasoning, and long context; it beat every other model out there. And what makes this even more wild is you can actually use this completely for free. So if you go over to AI Studio.google.com, over on the right side, there's a little model drop-down; you can actually select Gemini 2.5 Pro experimental and use it totally for free inside of this AI Studio.
And not only is it a super, super smart model that works really, really well, but it's got a million-token context window—that's about 750,000 words of input and output. And despite that huge context window, it is insanely fast. Now, the biggest downside of something like Google AI Studio is that it doesn't save all your previous chat history like something like ChatGPT or Claude, so you kind of got to use it, and then when you browse away and come back, you don't get to sort of pick up where you left off—kind of a bummer. But I'm sure they'll be rolling it into their normal Gemini chat soon where you do get all that functionality. But the speed of this model is just really, really good, especially with how much data it's able to consume.
So here's what I want to do to sort of show off how kind of impressive this model is. Here's a YouTube video called "Machine Learning for Everybody" from Free Code Camp. This video is almost 4 hours long; it's 3 hours and 53 minutes long. Let's go down here; we'll show the transcript, and I will go ahead and copy this entire transcript for this entire video here. Jump back over to Gemini; I'll paste this into our chat window down here. We can see it's using 43,266 tokens. Let's add a couple lines, and I will say, "Summarize this video into step-by-step bullet points." I'll go ahead and click "Run" on this one, and just like DeepSeek R1 and the OpenAI like 001 and 003 models, this is a thinking model, so I can expand to view the thinking, and we can see the whole thought process as it's reviewing the text that we just pasted in. And this is all real-time; I'm not speeding up this video. It's thinking really, really fast. Now, it is a long video, so it's probably going to take a little bit over a minute to process everything, but it's going to give us a full step-by-step breakdown of this 4-hour video. So it's done thinking now, and it's moved on to the actually giving me a response portion, and it's still moving very, very fast. But look, we can actually see this sort of step-by-step: introduction and setup, data loading and initial exploration, data visualization and basic concepts, supervised learning concepts, data preparation and collab, supervised classification models, supervised regression models, and we can see it took 62 seconds in total to look at this entire video transcript of a 4-hour video. We ended up using only 51,000 tokens; we used 5% of the context window. You can plug in like full-length books into this and get a step-by-step breakdown of a process in a book, and you can do it for free over at AI Studio on Google. Like, this is pretty mind-blowing. And again, it's kind of sad it got overshadowed by all of the Studio Ghibli memes that were spread everywhere. I think part of the problem is that Google does kind of have it buried inside of their Google AI Studio instead of on their front-facing Gemini platform where it can save the chats and everybody can use it really easily. And when you do look in AI Studio, it is kind of overwhelming, right? You've got all these settings on the side, and if you don't know what you're doing, you may get overwhelmed and feel like this isn't worth using, but it's really impressive.
The model was so impressive that even Sam Altman himself went and gave credit where credit was due and told Google that it was a great ship. A lot of people are claiming that Sam Altman is actually trolling Logan and Google, but Sam and Logan, I believe, are friends; Logan used to work at OpenAI. I actually don't think this is a troll; I think Sam was genuinely impressed by what Google shipped. Here are some more examples of what this new Gemini model is capable of. Someone created this super simple jump game here. I don't know if I'd really say this is like the best example of what it's capable of. This one's a little more impressive. Here's a snake game with some power-ups and some cool animations in it. We've got a soccer or football simulator here going on. I mean, this doesn't really seem like where the players should be. This one was pretty impressive to me. This one was actually created by Matthew Burman here, and it's actually a Rubik's Cube that when you rotate the cube, all of the colors stay exactly where they're supposed to stay. And according to Matt, this is the only model that he's been able to get it to keep this consistency where the colors actually stay on the square they're supposed to stay on. Here's a Three.js flight simulator, also made by Matthew Burman using Gemini 2.5. Another flight simulator from Philip Schmid. Yet another flight simulator from DV Meta. This one looks like he's playing it on a mobile phone as well. We got car simulators in Three.js, a block-building sim—I'm guessing this is sort of like a Minecraft style thing here—a building Tumblr sim. Somebody one-shotted a zombie game. I'm not actually seeing the zombies; I guess those are the zombies back there. A one-shot runner game, an asteroid game, a treasure hunt game in Three.js. Somebody remade Minesweeper. Somebody made Solitaire. And of course, more snake games.
Logan went to X to say, "So what have you done with Gemini 2.5 Pro today? Please show impressive things only." This was actually the one that Sam Altman commented on. Logan made this animation of a ball bouncing around inside of a square. AI Warper here made this particle simulator that simulates particles over like an existing image, which looks just wild to me. Look at that; that's pretty cool. Here's another one from Matt Burman: 3D bloodstream virus simulation. So yeah, some really, really impressive stuff coming out of Gemini 2.5. It's the smartest model we've got right now, and the largest context model we've got right now, and it's available for free, and the APIs are available so developers could build with it. Grok 3 is a really good model, but we don't have APIs for it yet. 001 Pro is also a really good model, but the API costs like $600 per million tokens, which is just sort of not feasible for most people to use inside of their apps. But again, all of it was kind of overshadowed by the ChatGPT thing. And if you like to Vibe code like me, WindSurf added Gemini 2.5 Pro in as well, so now you can use this model directly inside of WindSurf to code with. I also believe if you use Cursor, you can use it inside of Cursor now as well. WindSurf has kind of become my IDE of choice for coding, but you can use either one with Gemini 2.5 now.
Again, two massive updates from two of the companies that are sort of leading the way in AI right now. Real quick, if you're an AI enthusiast, creator, or someone who collaborates on tech projects, you probably know the pain of juggling multiple tools for meetings, messaging, file sharing, etc. It's a mess, right? Well, Microsoft just made it insanely easy to streamline everything, and it's free right now. AI is moving at breakneck speed, and collaboration is key. Whether you're testing new AI tools, brainstorming startup ideas, or just discussing the latest ChatGPT update with your team, you probably bounce between five different apps: one for meetings, another for chat, another for file sharing. It's clunky, it's inefficient, and it slows you down. Microsoft Teams Free gives you everything in one app: video calls, unlimited chats, shared docs, and even a community feature that's perfect for AI masterminds and startup teams. And it's built for early adopters like us, people who thrive on innovation and need tools that keep up. So here's what's in the Teams plan: 60 minutes of free video meetings versus Zoom's 40-minute limit, unlimited chat with no message deletion, 5 gigabytes of free OneDrive storage for easy file sharing, and it works across desktop, mobile, and web seamlessly across devices. It's perfect if you're running a side hustle, launching a startup, or just working with a group of AI-curious friends who want to stay connected. I've been testing it out, and honestly, it just makes sense: no more app switching, no more chaos, just everything you need for free. Check it out today at aka.ms/FutureTools. If you're serious about AI tech or just getting things done, this is a no-brainer. So thank you so much to Microsoft Teams for sponsoring today's video. Now let's get back to it.
We also got news out of Microsoft this week. They introduced their new Researcher and Analyst in Microsoft 365 Copilot. This is actually using OpenAI's 003 mini reasoning model, but they optimized it to do advanced data analysis work, and it uses the Chain of Thought reasoning, so it actually thinks through before it gives you a response. Now, let's say I work in product development, and we're entering a new market. I need help developing a product strategy for our expansion. I enter my prompt, and Researcher gets to work. It asks clarifying questions as a seasoned colleague would and uses my response to move forward. The agent takes my prompt, understands the task, and constructs a plan to reach an answer. You'll notice it reasoning over all of my work data in the Microsoft Graph, not just one file. Here we can see its Chain of Thought reasoning. You'll see it working through the problem in real time. It builds an understanding of my product lineup, references recent meetings, and even pulls industry updates from the web. This process takes a few minutes, so let's jump ahead to the result. I got a really thorough response in line with what I'd expect from a researcher on my own team. So pretty similar to what we've been seeing with things like Deep Research in ChatGPT and Deep Research in Google Gemini and Deep Research in Perplexity, but now it's built into Microsoft 365 Copilot.
Imagine I'm in marketing. I'm trying to understand our most loyal customers. I have this really messy data set. You can see thousands of rows and a few tabs across customers and their monthly revenue, and none of it has been cleaned or contextualized. I don't have to spend time writing the perfect prompt to get what I'm looking for. I'll just ask it to help me come up with an easy way to learn and visualize my customer base. You'll see the agent takes my question, understands the task, and constructs a plan to reach an answer. It also identifies the Python tools it may need to complete the task. It can take any complex data set, make sense of it, and execute Python code to reason through the questions you have about the data. If I want to see more into its thought process, I can always click and see its Chain of Thought reasoning as well as the Python code it's running in real time. This allows me to validate and trust in its thinking and approach. Now I've got the answer I was looking for. They also announced deep reasoning and agent flows inside of Microsoft Copilot Studio, which they describe as a platform to easily create, manage, and deploy agents for your unique business needs. So using Microsoft's suite of tools, you can actually build your own sort of mini agents that run on your own business's data.
And that was probably the biggest news of the week: the new ChatGPT image gen, Gemini 2.5, and Microsoft's new analyst features. But there was a whole bunch of other really cool AI announcements that sort of flew under the radar because, well, those ones really sucked all the oxygen out and got all of the attention this week. We had some more announcements out of OpenAI. Apparently, their GPT 4.0 model got even better. Now it's better at following detailed instructions, especially if prompts have multiple requests inside of a single prompt; improved capability to tackle complex technical and coding problems; improved intuition and creativity; and if you've ever had it do creative writing for you, you'll probably appreciate this: fewer emojis.
Another interesting move out of OpenAI this week: they've actually adopted MCPS, or Model Context Protocols. So early on, ChatGPT tried to do tool use functionality with their plugins and then with their custom GPTs. Well, late last year, Anthropic rolled out a sort of new standard for connecting large language models to tools on the internet called MCPS, Model Context Protocols. There's sort of a layer between the large language model and a software's API. All of these APIs that software companies give to developers to be able to use their tools are all kind of developed in different ways; no two APIs are exactly the same, so it's kind of difficult for large language models to communicate with software's APIs. Model Context Protocols act as a sort of layer in between to standardize the way that large language models communicate with tools on the internet. It was a standard again introduced by Anthropic, and it seems that OpenAI is getting on board with that standard.
There was some additional news out of Google this week. Google just keeps on shipping AI features. They added new features inside of Google Meet. The "Take notes for me" feature can now capture follow-up action items from meetings and give suggested next steps. When you turn on the transcript for the meeting, the meeting notes that Gemini creates will link to the relevant part of the transcript so you can see more details and verbatim quotes. And if you miss part of the conversation or need to refresh your memory, you can now scroll through the meeting captions while the meeting is happening. Google added a feature inside of Maps, too. Google can now save locations you screenshot in Maps to help with travel planning. So we can see in this little screenshot here, there's an image that says, "Little Island, I love New York City," and Google Maps is able to tell from that screenshot where that location is and actually show it to you on the map. It's rolling out this week in English on iOS and coming soon to Android. I just think it's interesting that it's going to iOS before Android, and it's a Google app, but never mind.
This week, Google also released TX-Gemma, or TexGemma—I'm actually not quite sure how that's supposed to be pronounced. It's a collection of open models designed to improve the efficiency of therapeutic development by leveraging the power of large language models. It uses DeepMind's open-source Gemma, but it's specifically trained to understand and predict the properties of therapeutic entities throughout the entire discovery process, from identifying promising targets to helping predict clinical trial outcomes.
Anthropic has been a bit quiet this week, but according to this testing catalog website, it's expected that Claude 3.7 Sonnet is about to get a 500,000 token context window—not quite the 1 million token context window we just got from Gemini 2.5, but this will be a huge upgrade to Claude 3.7 Sonnet, especially if you like to use Claude for Vibe coding in tools like WindSurf and Cursor. This hasn't been 100% confirmed, but people have seen evidence of this by looking inside of the code, so it looks likely to be rolling out soon.
If you're a fan of using Grok, you can actually now use Grok directly inside of Telegram. You have to be subscribed to both Telegram Premium and X Premium, but if you are subscribed to both, you'll be able to chat with Grok AI's chatbot within your Telegram streams. So if you're someone that just refuses to use the X app, well, Telegram is another option, although if you refuse to use the X app, you're probably not going to be a subscriber of X Premium, I'm guessing.
Perplexity rolled out a new feature this week inside of their web app where you can specifically search for images, video, travel, shopping, and more. So if I head over to perplexity.ai here and I search for "best restaurants in Kauai," I can see that it now added a "Places" tab, and I can click on this, and similar to what we would get out of Google, we've actually got the various restaurants, and it actually shows them on a map for us. Now, if I type in "Photos of a wolf," we can see it actually gives us a bunch of images right here in the top, but it also gives us a new "Images" tab. If I click on "Images," similar to Google Images, it brings up a bunch of images of wolves for me.
DeepSeek, the company that sort of freaked everybody out with their R1 model about a month ago, they just released an improved version of V3. Now, V3 was the underlying model of R1. When everybody was using DeepSeek R1, it was actually using DeepSeek V3 as the large language model, but R1 was the sort of reasoning element on top of it, the sort of Chain of Thought thinking add-on on top of V3. Well, V3 just got some improvements this week. According to Ani Hanun here, the new DeepSeek V3 0324, meaning it came out on March 24th, can actually run at 20 tokens per second on a 520GB M3 Ultra, which tells us that these large language models that run locally are actually becoming less reliant on Nvidia GPUs. Now, it did use 380.1GB of RAM, so yes, you can run it on a consumer Mac Studio with an M3 Ultra, but I mean, I have to put "consumer" in air quotes.
Cuz this is probably about an $8,000 to $10,000 max studio to get this to run. Alibaba actually had a few new models get released this week. They had a 72-billion parameter model and they had a 7-billion parameter model, but now they have a sort of in-between model at 32 billion parameters—their Qwen 2.5 VL model. This is an open-source model under the Apache 2.0 license, and it is a vision model, so it can actually see images and respond based on what it sees in the images. Alibaba also released QVQ Max; this is their visual reasoning model. It can understand the content of images and videos but also analyze and reason, so it actually does that Chain of Thought processing as well. Both these new models are available to test out over at chat.qu.ai. We can see if we click on the models; we've got Qwen 2.5 Max here, and we've got QVQ 32B here.
Now, even without OpenAI's Big Image Gen reveal this week, it was already shaping up to be a big week in the world of image generation. There was a new model that sort of came out of nowhere called Reeve, which, according to the Artificial Analysis Image Arena, beat out all the other image models, including ReCraft, Imagine 3, Flux One, Midjourney, etc. And similar to what we've seen from the new model inside of ChatGPT, it allows users to not only generate images from text but also modify existing images with simple language commands—things like changing colors, adjusting text, and altering perspectives. It also supports uploading reference images, enabling users to create visuals that match a specific style or inspiration. You can even try it out over at preview.re.art. And we'll give it a prompt like, "a wolf howling at the moon." It generates the images pretty quickly, and they're all pretty good. Now, if I select one of the images, I can actually instruct some changes and say, "Make the wolf have black fur," and we can see now I have images of a wolf with black fur. Now, it didn't maintain the actual composition of the original image, but I was able to create some new variations, and people were pretty blown away by this model for a day or two until OpenAI launched theirs, once again sort of sucking the oxygen out of every other announcement that happened this week.
Idiogram also released a new model this week; they released Idiogram 3.0, which is much better at realism, creative designs, and consistent styles—all in a powerful model, and it's super fast. So if we head over to idiogram.do.ai, we can use this one completely for free. Start creating with 3.0, and this one is also really good at text within the images. I can give it a prompt: "An image of a wolf standing on its hind legs holding a sign that says 'Subscribe to Matt Wolf.'" We'll click into the Creations tab here, and every single image that it generated pretty much nailed the prompt, nailed the wording, and is really impressive. I mean, in my mind, AI image generation is like a solved problem now—like you can generate anything you can imagine. Between ChatGPT's new model, the Reeve model, and Idiogram, anything you can imagine, you can do now.
We got some updates in the world of AI video generation as well. Luma AI showed off their new Magic Doodles feature with Runway's image-to-video. You can doodle an image and plug it in and have that image animated, and this one's really fun if you have kids that love to draw. My daughter is an artist; she loves drawing, so it's super fun to take some of her artwork, throw it into a tool like Runway 2 from Luma, and watch her little drawings come to life—super, super cool update. Dream Machine also rolled out a thread feature, which is a new way to organize your creative process—keep 720p, 1080p, 4K, and audio versions of the same assets all in one place. So basically, some better organization features inside of Luma Dream Machine. And Pika Labs rolled out a new feature—I mean, they're rolling out a new feature like every week it seems like as well—so they rolled out this Flashback feature where you can upload a video of yourself, upload a photo of yourself, and then the photo version of yourself will enter the frame. Like Pika has kind of been nailing the sort of like meme video generation. It makes me think that maybe Pika has realized that their way of standing out among the Soras and V2s and and Clings and Dream Machines of the world is to really lean into this like meme generator platform.
As a side note, I also find it really fun to take an actual image of yourself and a restylized image that maybe you created with something like the new ChatGPT image gen model and have Pika sort of animate between the two. So, for example, here's an image of the two Padres players here; let's set that as the first frame, and then here's the sort of Rick-mortified version; we'll set that one as the second frame, and let's animate between the two, and here's what we got out of that, where it sort of takes the initial real-life image and then sort of animates it into the cartoon version. I don't know why; I just think this stuff is like super fun to play with. This company, Earth AI, has an algorithm that found critical minerals in places that everyone else ignored. They claim that the actual real frontier in mining is not so much geographical as it is technological. Earth AI has identified deposits of copper, cobalt, and gold in the Northern Territory and silver, malum—not sure how to say that word—and tin in another site in New South Wales. Earth AI algorithms are trained to scan wide areas quickly and efficiently to find deposits that might otherwise have been overlooked. I just love hearing these like real-world use cases of world-changing uses of AI. They're actually using AI now to find critical minerals.
If you're into self-driving vehicles like Waymos—which I actually rode for the first time when I was in San Francisco last week—they're actually going to be launched in Washington, D.C., next year. So if you're out there and you haven't ridden in a Waymo yet, you'll get the opportunity pretty soon. And apparently, Lyft is going to start rolling out robotaxis in Atlanta later this summer. And finally, if you want to see how far these, uh, humanoid robots are getting and what they can actually do now, here's a pretty interesting one from Boston Dynamics, where we can see their robot running, and it actually looks like a human. We could see it kneel down, and we can actually watch it crawl on all four legs on the ground and even do these like barrel rolls or whatever you call those, and here it is doing whatever you would call this move right here. But I mean, if I saw these videos like two years ago, I would have been like certain that there was somebody inside of a robot suit doing these moves. This is just crazy that the robots can actually do this now. So much fun to watch; I love robots. There's one doing a cartwheel, and that's how I'm going to go ahead and end this video.
That Boston Dynamics robot video is awesome. If you like staying in the loop with AI news and the latest AI tutorials and all of the coolest AI tools, make sure you like this video and subscribe to this channel. I will make sure that you stay looped in and more videos like this show up in your YouTube feed. And if you haven't already, make sure to check out futurtools.com. This is where I share all the coolest AI tools that I come across along with all of the AI news as I come across it. There's also a free newsletter; join the free newsletter. I'll email you twice a week with just the most interesting tools and most important news for the week. You also get free access to the AI Income Database, a database of ways to make side income using various AI tools. It's all 100% free, and you can find it all over at futurtools.com. Thank you so much for tuning into this video and nerding out with me. It has been one of the most fun weeks in AI, between Google's announcements and OpenAI's announcements. I have been having a blast, and I've got so many more videos on the way where we're going to play with some of these tools and really, really challenge them to see what they're capable of. So if you want to see that kind of stuff again, make sure you're subscribed, and I will do my best to get it in front of you. Thanks for nerding out; hopefully, I'll see you in the next one. Bye-bye.