📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

AI News: Google's AI Can Now SEE Everything!

Matt Wolfe39:19

Transcription

Well, it's been a heck of a week for AI news, especially from Google and OpenAI. Both those companies have been shipping like crazy, and I'm going to do my best to break it all down for you right meow.

Google has entered what they call the Gemini era. Their new Gemini 2.0 Foundation model was released this week, and along with it was a ton of cool features and ways that this Gemini model is being used. They started with the release of Gemini 2.0 Flash, and the Flash model is a smaller model, but it actually performs much better than Gemini 1.5, which is a larger model trained on a lot more data. We can see by the benchmarks here that it pretty much outperforms in almost every single area compared to the much larger models that came before it, and it does it all at twice the speed.

You can actually play around with the Gemini model yourself right now for free. Just head on over to gemini.com. Now, I am a paying member of Google, so yours will probably look a little bit different, but if you click this drop-down here, you will see 2.0 Flash Experimental. This is a really good model, and you'll probably notice it's pretty comparable to things like Claude and ChatGPT and things like that. But there's an even cooler way to use it, which will open up even more possibilities, and that's by going to Google's AI Studio. If you go to AI Studio.google.com, even as a free Google user, you will see Gemini 2.0 Flash Experimental over here that you can use right now.

If we take a look at that pricing, well, it's currently free to use and test. You can also give it some extra features like structured output, code execution, function calling, grounding, and even adjust safety settings here, essentially allowing you to make it completely uncensored. But here's where it gets really cool: if we click over on Stream Real Time, we have some additional options here. You can talk to Gemini with real-time voice conversation, you can turn on your webcam, and Gemini can see what's going on on camera, or you can even share your screen, and Gemini will be able to help you with whatever you're seeing on your screen.

Let's go ahead and try to talk to Gemini. What are some of the best use cases of this new Gemini 2.0 Flash model? The Gemini 2.0 Flash model is great for various use cases. It's particularly good for tasks that need quick responses, like real-time data analysis or generating content summaries. It is also useful in chatbot applications or in any situation requiring rapid information processing.

Let's go ahead and try to show Gemini and see if it'll actually look at my camera, even though it's using my camera for the recording. Can you describe what I'm wearing right now? Sure, you're wearing a dark-colored T-shirt and a camouflage baseball cap. You also have headphones on and a beard.

Describe the background environment for what I'm doing right now. Okay, the background appears to be a room with dark blue walls. There's a cabinet behind you, as well as some hanging stringed instruments like guitars and a banjo. Additionally, there seems to be a shelf with some decorations.

Finally, let's go ahead and share our screen. Can you tell me what's going on on my screen right now? Sure, I can tell you what's on your screen right now. You are on a page titled Google AI Studio with a Stream Real Time section selected. The page has options for system instructions, a run settings menu, and other tools. There is also a chat box at the bottom with a start and stop sharing screen option and a video of your screen in the bottom right corner of the page.

How many tabs do I have open in my browser right now? Based on the image, it looks like you have 10 tabs open in your browser right now. Well, didn't quite get that right; there's quite a bit more browsers open, but it seems to work and be able to see my screen pretty well. So theoretically, if I am trying to work through something on my computer or I run into a coding issue, or even if I'm playing a video game, I can have my screen share going and have the AI sort of talk me through what I'm supposed to do next. That's kind of the idea here.

Now, this is an early prototype; it's still experimental, so as we saw, sometimes it can't do things very accurately, like count tabs. But in most of my experimentation so far, it's been pretty spot on. It's also got some additional functionality here, like spatial understanding, where you can upload a photo and combine 2D or 3D spatial analysis with reasoning.

So it's got some sample images here. Let's go ahead and use this image and say label each one with the type of food it is. You can see it's doing its little calculations here. It created these bounding boxes around it, and if I hover over it, you've got cheese, you've got a muffin, you've got a strawberry, etc. It labeled all these things in the image for me.

Here's another image I just uploaded from my computer. Let's try a 3D bounding box. Detect the 3D bounding box of, let's do the speaker. Okay, that wasn't the type of speaker I meant, but it did find the speakers on the stage. Let's do the person on stage, and there you go. It made a bounding box kind of where I am. It's not perfect, but you get the idea of what they're going for here. You can also upload videos and have it analyze the videos. The videos don't even need to have audio, and it will know what's going on in the video and describe the scenes and things like that. So you can theoretically find very specific points in your video.

There's also a map explorer. A lot of these features we'll probably play around with in a future video. However, the new Gemini 2.0 Flash is really, really impressive. But that was just the beginning of Google's announcements this week. They also announced Project Astra, which, if you pay attention to my shorts feed at all or you're following me on X or Instagram, I shared my experience with Astra because Google actually flew me out to London, put me up in a nice hotel, and gave me a brand new Pixel 9 phone with Astra pre-installed on it so that I can play around with Project Astra myself.

Basically, what Project Astra does is it's designed to be like an AI assistant on your phone, like an infinitely better version of Siri that can actually see what's going on around you and eventually connect to all of your Google apps. It should eventually be able to connect to Maps, Calendars, Gmail, and things like that. Right now, it's not connected to all that stuff; it's just got the sort of vision capabilities. It can search the web, it can answer questions for you, and it is really, really impressive.

In fact, let's go ahead and just demo it right here on this video. So I've got the phone open here. I'm going to go ahead and open Project Astra, and now if I click Start... Hey Astra, what am I looking at right now? Based on the view provided, you seem to be observing a setup with multiple screens, one of which is displaying a live feed from a smaller device, perhaps a phone or a monitor. There is also a smaller monitor mounted to the top of the larger monitors.

Can you tell me what kind of camera this is right here? Based on the logo, it appears to be a Sony camera. I'm unable to specify the model of the camera, though.

Can you tell me what kind of microphone this is? This is a Shure SM7B microphone. It is a popular choice for vocal recording and broadcasting.

Can you tell me where this picture on the wall was taken? Based on the appearance of the shoreline and vegetation, it's plausible that this image was captured in a tropical location with volcanic terrain, potentially in Hawaii or a similar region. It is only a guess.

Well, you're right; it's Maui. Glad to have gotten it correct. Another thing that this is really, really good at is you can take a book, like I've got this book here in my hand, open up to any page in the book, I could briefly flash the camera at the page on the book, and it will know everything about that specific page.

So check this out. Can you tell me what's going on on this page of this book? The page is discussing negotiation strategies, specifically focusing on the topic of never splitting the difference. It mentions a scenario involving a CEO and an employee who is frustrated at his negotiation.

Can you tell me the name of the book? The title of the book is "Never Split the Difference."

Do you know the author of the book? The author of "Never Split the Difference" is Chris Voss, a former FBI hostage negotiator. He outlines his strategies for negotiation throughout the book. Impressive!

And so that's Project Astra on the phone. Now, I am one of Google's trusted testers, so there's only a handful of people that have access to this, and it's not going to be rolled out publicly until after they feel comfortable that all of their tests and red teaming and things like that have gone well and they are confident to put it out into the public. But this will eventually replace Google's Assistant on your phone.

But that's not where Project Astra ends. The next phase of Project Astra is actually to put these inside of glasses, so I don't need to hold a phone up anymore. I can just look around; it will see what I see and it will respond in my ear in the same way we were just listening to it. I actually got to test the glasses early. I'm not allowed to talk about it too much yet, but the glasses do exist and they are coming at some point.

In fact, here's a little demo of what you might see while using the glasses, because they do have a little heads-up display in them. You can see an example here where it's actually giving directions right in front of your eyes as you're walking. Here's another example where they're looking at a menu, and right on the little heads-up display of the screen, it's actually translating what that menu is saying.

Here's somebody else using it to get instructions on how to hang floating shelves, and it is popping up right in front of their eyes inside of the glasses. It is wild, honestly; it is really, really cool.

And before I forget, I need to mention some of the things that make Astra really interesting. So again, it's using the Gemini 2.0 model, and the Gemini 2.0 model is multimodal from the ground up. So when you're using something like ChatGPT and you're talking to it using the voice mode, it's still using Whisper to transcribe everything you're saying into text, and then it's taking that text before it gives you a response. That's not how Gemini works. Gemini actually understands what it's seeing; it understands what it's hearing. It's not doing that transcribing underneath; it just understands the audio coming into it.

The other thing that makes Astra really interesting is its memory. So it's got a visual memory of up to 10 minutes. So I can ask Astra, "What book were we just talking about a few minutes ago?" As I recall, we were discussing the book "Never Split the Difference." And it actually remembers what we were talking about and what book we were looking at.

It also remembers things from previous conversations. So I was using this in London last week, so theoretically, it should remember what city I was exploring last week. In fact, let's see if it does. What city was I exploring last week? You mentioned a trip to London in our earlier conversation, so I assumed that you were talking about that. However, I apologize for assuming.

So it's actually got the context and memory of previous conversations. I had personally, I think everybody should have a website. I think it's great to have a place where you can express yourself online. That's why for today's video, I partnered with Hostinger. Hostinger is a platform where you can host your website, but you can also completely build your website using AI.

If you head on over to hostinger.com/slmtwolf and you claim the special deal, you're going to want to choose the Business Website Builder because this is the one that has all of these cool AI tools built into it. Select how long you want to set up your hosting for; the longer the period, the less per month it costs. Then see this little link that says "Have coupon code?" Click that link and type "Matt Wolf" and click apply, and you'll see you'll get an even larger percentage off the plan.

Once you're inside of Hostinger, you come over to Websites here, click on ADD Website, and choose Hostinger Website Builder. You can see this is an AI-powered drag-and-drop website builder. Let's say, for instance, I want to make a website all about 3D printers. Let's call the brand name "3D Print World," and for the description, I'll enter something like "A website that shares all the latest news and reviews on the newest and most advanced 3D printers."

Once I give it that little bit of detail, I just click "Create Website" and let the AI go to work for me. Within less than a minute, I have a fully fleshed-out website ready to go. It's got images of 3D printers, it's got pre-written articles already, it's got various page sections like blog and review. All I have to do is start swapping out the information with what I want on the site.

So if I click on "Edit Site" up in the top right, you can see I now have a drag-and-drop builder. I can drag my title around, drag these buttons around, and this whole website builder has AI baked into it. I can use AI to generate images. Over on the left, I have a whole AI tools menu here. Let's try out the AI heat map. If I click here, you can see it's creating an intention map, and it uses AI to predict where people's eyeballs are going to go on the website.

So as I scroll the website, I get a pretty good idea of what's going to draw people's attention in, so that I can strategically place the most important text and images in the areas that are going to grab attention. This has to be the easiest way to get a website online and running with no technical knowledge at all.

So once again, you can check it all out over at hostinger.com/wolf and use the coupon code "Matt Wolf" for an additional 10% off their already heavily discounted service. Thank you once again to Hostinger for sponsoring this video.

Google also announced Project Mariner this week. Project Mariner is kind of like the computer use that we saw coming out of Claude a few weeks ago, where you can actually give Google access to do things in your browser. I'm going to start by entering a prompt here. I have a list of outdoor companies listed in Google Sheets, and I want to find their contact information. So I'll ask the agent to take this list of companies, then find their websites and look up a contact email I can use to reach them.

This is a simplified example of a tedious multi-step task that someone could encounter at work. Now the agent has read the Google Sheet and knows the company names. It then starts by searching Google for Benchmark Climbing, and now it's going to click into the website. You can see how this research prototype only works in your active tab; it doesn't work in the background. Once it finds the email address, it remembers it and moves on to the next company. At any point in this process, you can stop the agent or hit pause.

What's cool is that you can actually see the agent's reasoning in the user interface so that you can better understand what it is doing, and it will do the same thing for the next two companies, navigating your browser, clicking links, scrolling, and recording information as it goes. You're seeing an early-stage research prototype, so we sped this up for demo purposes. After the fourth website, the agent has completed its task, listing out the email addresses for me to use.

So this isn't something I've had the opportunity to play with yet, but as one of the trusted testers, it might be something they put in my hands, and maybe we can experiment with it in a future video. They also showed off Jewels, which is an agent for developers. So Jewels is essentially an assistant that can help with all of your coding tasks. They also showed off agents inside of games, where you can actually be playing a video game on your computer and have it help you with the video game.

Another thing that's really interesting about this new Gemini 2.0 model, since it's multimodal from the ground up, is that it's got native image output. So you can upload an image like this car, tell it to turn the car into a convertible, and it will output an image of that car as a convertible. Now, that feature doesn't appear to be available yet; I haven't gotten it to work at least. But there's also some examples where they blend images, so I thought this looked really cool. You can see they uploaded an image of a cat and a pillow and wrote, "Create a cross-stitch of my cat on this pillow," and it created a cross-stitch cat on that pillow.

Here's another example where they put a detailed illustration sticker of my cat on my skateboard, and it put a sticker of the cat on the skateboard. So that is really exciting to be able to take two images and blend them together with a prompt like this. I'm super excited to get my hands on that feature, but again, I haven't gotten it to work yet, so I don't think it's rolled out just yet.

But there's even more news from Google this week. They also rolled out the Deep Research feature inside of Gemini, which is kind of like a more advanced version of Perplexity, where it will actually search the web and do research on topics for you. So if I head back over to my Gemini account, this one I believe is only in the paid plan, so I think you need to have Gemini Advanced to work with this one. But the Deep Research is borderline agentic; it will actually go and do multiple searches and deep dive on a topic for you.

So I can give it a prompt like "Research Quantum Computing and how far we are from being able to crack Bitcoin cryptography," and you'll see it will give us some steps that it's about to take. So research websites, find information on the basics of quantum computing, find research papers and articles on the potential of quantum computing, break cryptographic algorithms, etc. Then it's going to analyze the results, create a report, and it will be ready for me in a few minutes.

So if I go ahead and click "Start Research," it's going to go through all of those steps for me, and we can literally watch it as it goes through and does all this research. So it's researching 16 websites for me right now. Now it's up to 45 websites. We can see all of the websites that it's reading from. Like, it really goes deep. Perplexity might use like four or five websites when it's doing the research for you. You can see that Google's already up to 65 websites right now, and after several minutes, it came back with a really, really in-depth detailed report all about quantum computing and its ability to crack the Bitcoin cryptography.

What's really cool is there's one button here, and I could quickly open it in Google, so I can quickly save this entire report that it generated along with all of its source websites right inside of my Google Doc to refer to it later. Pretty dang cool!

And this video is already getting long, and so far, all I've talked about was all of the announcements Google made. As you probably realized by now, OpenAI has been shipping like crazy as well. They've got their 12 days of OpenAI announcements, and well, since I record these on Thursday, we know what four more of those announcements are this week. By the time you're watching this video, there's probably a fifth announcement that came out on Friday.

But let's get into the four announcements that OpenAI made this week, starting with the big announcement that on Monday, they released Sora. After demoing it and showing it off nine months prior, they finally gave us a version of Sora. Now, the version that's out is actually called Sora Turbo, which lets you generate up to 20-second videos if you're on the Pro Plan and I believe 10-second videos if you're on the Plus Plan.

Well, since OpenAI waited nine months between showing it to us and actually getting us access to it, a whole bunch of other alternatives have popped up, and I don't think Sora made people as excited when they released it as OpenAI had anticipated. Not to mention that OpenAI basically crashed their servers when they opened up Sora because everybody wanted in all at the same time; it couldn't handle the load. Well, they've since worked that out, and I believe people are able to get in again, but OpenAI's video generator is now finally available.

I did a whole video where I tested some prompts because it was so bogged down, and it was taking so long for the videos to load. I basically had to cut the video off shorter than I wanted to, and most of the prompts that I had tested to that point were all short prompts. What I've kind of learned since then is that Sora really likes long, detailed prompts. I couldn't get it to generate a wolf howling at the moon in my last demo, but after giving it a much more detailed prompt, I did manage to get a much better wolf howling at the moon video.

It really struggles with people doing things like dancing or gymnastics, as you can see in these examples here, but for some reason, it's pretty decent now at a monkey on roller skates. In the Sora video that I released earlier this week, I did go into more detail about some of the features that Sora offers, things like the ability to blend two videos together, like you can see here where I took a monkey on roller skates and blended it with clocks flying around.

It's also got a cool storyboard feature where you could try to sort of steer the video a little bit, but in my testing, it wasn't steering in the direction that I wanted it to. I still have to do some more testing. Again, if you want to see me do a little bit deeper of a dive on Sora, check out this video called "You Can Now Use Sora; Here's How." I dive deeper into some of these features in that video.

On day four of Sora's 12 days of announcements, they released Canvas. Now, if you've had a Plus Plan, you've had access to Canvas already. Now they've made Canvas available to everybody; it's on even the free plan. They also gave the ability to execute Python code inside of Canvas. So if I jump into ChatGPT here, you can see there's a new icon here that says "View Tools," and it gives us the option for pictures, search, reason, and Canvas.

Now I can tell it to write a simple "Hello World" code in Python, and just by giving it that prompt, you can see it will generate the code inside of this box. But if I switch to Canvas, give it that same prompt, you'll notice it changes the user interface completely and it puts my chat over on the left and it opens up a sort of code screen here. I also have a button that says "Run," so I can see the output, and down in the bottom, I can see in my console it says "Hello World."

Now, obviously, this is designed for much more complex code. You can also use this Canvas to do writing tasks and stuff. So if I create a new chat, make sure I've got Canvas selected, and say, "Write a poem about wolves on Christmas," you can see it opens up that same user interface with my prompts on the left and the poem that it's writing here on the right, along with some additional tools down here in the bottom right, like suggest edits, adjust the length, adjust the reading level, add some final polish, or for whatever reason, if you want, you can add emojis as well.

So again, Canvas was available already if you were on a Pro Plan; now it's available for free, and now it can also run Python. So that was the day four update. On day five, we learned that it is now available inside of Apple Intelligence. If you use Siri, you can actually turn on the ChatGPT extension inside of your Siri, and you can prompt Siri to use ChatGPT when asking it a question.

But not only did it roll out on iPhones—well, you have to have an iPhone 16 or newer—it also rolled out in the Mac OS as well. If you have a Mac computer, you can now use ChatGPT with Siri on your Mac computer. You can see they clicked a little Siri icon up in the top right of their Mac, and they can now type to Siri or talk to Siri. Apparently, it can also see your screen, so they were asking a question about this document, and it took a screenshot and was sharing the screen with ChatGPT using Siri on the actual Apple computer.

Now, I feel like this was kind of a cheater announcement from OpenAI just because it was really more of an Apple announcement, but OpenAI made it part of their 12 days of announcements because on the same day, the new iOS 18.2 rolled out into all of the Apple devices with Apple Intelligence, the new image playground, Gen Emojis, writing tools, and seamless support for ChatGPT and visual intelligence.

So it just kind of seems like they timed up their day five announcement with the same day that they knew Apple was going to be releasing 18.2, so it was kind of a joint company announcement. At this point, most people have probably already seen a lot of the new Apple Intelligence features, like the Gen Emoji, where you can actually have it create emojis for you, like a DJ octopus or a rainbow cactus.

It's got a new and improved note tool where you can use AI directly inside of your notes, new writing tools, and the ability to use AI to help you with your text. We've pretty much seen all of these features from Apple; they're just finally rolling them out in their latest 18.2 update.

And then finally, day six, which was Thursday, the day I actually record these videos, they announced advanced voice mode now has video capabilities. So this was a feature they showed off seven or eight months ago when they first showed off the advanced voice mode. They actually showed it being able to take pictures and see what you were looking at and have conversations with you about what you were looking at.

When advanced voice mode finally rolled out and was made available, this feature was absent. Now you've got it in your account. To access it, you can just use your ChatGPT account, go into the voice mode by clicking the little white circle there, and we can see some new options. We have the ability to turn on our camera. We've got the three dots here, which allows us to upload a photo, take a photo, or actually share our screen on our phone.

But let's go ahead and test it with some of the same types of stuff I was just testing Astra with. I'll turn on my camera here. What am I looking at right now? It looks like you're showing your desk setup. I see a monitor, some sticky notes on the bottom of the screen, and a black smartphone on the desk. There's also a camera on top of the monitor.

Anything specific you want me to focus on? Can you tell me where this photo was taken? That's a beautiful beach photo. It's hard to say exactly where it was taken, but it looks like a tropical or coastal location with those waves and rocks.

Do you remember where you took it? I do. I wanted to see if you can figure it out by looking at it. From the looks of it, it could be somewhere in Hawaii or maybe a rocky California beach. The landscape is definitely giving tropical vibes, so it got that right as well.

Let's go ahead and open to a page of a book and see if it can do the same kind of thing we got out of Astra. Same exact book. Can you tell me what this page of this book is about? The book appears to be discussing negotiation tactics. The title "Never Split the Difference" suggests it's about negotiation strategies, and the page mentions keeping emotions in check during negotiations. It's likely from a chapter about managing emotions and staying calm to improve negotiation outcomes.

Who's the author of this book? The book "Never Split the Difference" is written by Chris Voss. He's a former FBI hostage negotiator, and the book shares his negotiation techniques for business and everyday life.

Can you describe my background right now? Your background looks like a neat office or studio setup. There's a black shelf or cabinet behind you with some framed items and a few decorative pieces. It looks professional and tidy, perfect for video calls or recordings.

So I'll need to actually take that out into the wild and walk around with it like I did with Project Astra to really see how it compares. But on first impressions, first tests, it actually seems pretty on par with what I was able to do with the Project Astra demo. So maybe that's a future video. I'm just going to walk around with both of them out in the real world and see which one gives me the better responses on things.

OpenAI also had a somewhat gimmicky rollout today as well with the fact that you can actually talk to Santa Claus inside of your OpenAI app. How many houses do you have to visit this year? Every year, I visit millions and millions of houses all around the world on Christmas Eve. It's a magical night filled with joy and excitement as I deliver presents to children who've been good all year.

What if I've been really naughty? Well, everyone has their moments of being a little naughty from time to time; it's normal. The important thing is to learn from it and try to be kind and thoughtful. Remember, Christmas is a time for forgiveness and spreading joy.

So there you go! If you've always wanted to have a chat with Santa Claus, well, you can do it with GPT now too.

All right, so that's probably all the biggest news of the week. Most of it came from Google and OpenAI, but there's a handful of smaller items that I think are important to touch on just to make sure you're looped in with all of the progress that AI is really making.

If you use Anthropic and you're a fan of Claude, well, they now have a new 3.5 Haiku model, which is their smaller model, which is a little bit cheaper to use and a little bit faster to use. If you head on over to Claude, you have this drop-down to change models. You've got Claude 3.5 Sonic, Claude 3 Opus, and then under more models, you've got Claude 3.5 Haiku. Before, this was just Claude 3 Haiku, so it's a slightly improved Haiku model, which is their smaller model, and Anthropic hasn't even made any announcements about it.

So I expect pretty soon, Anthropic will probably do a detailed write-up with benchmarks and things like that, but they haven't yet. If you use Grok, it actually got a new image generator this week as well. Before, Grok was using the Flux 1.1 Pro model, which was really good, really photorealistic, but likely costing X quite a bit of money using the API credits.

So now X and Grok rolled out their own model. Originally, it was called Aurora; now they just call it Grok Image Generation. But if you head on over to your X account and you click into Grok, you can prompt it to create an image for you. Create an image of a wolf howling at the moon, and one thing you'll notice as it creates the image is it actually sort of scrolls from the top down. It's because it's not actually a diffusion model; it's not using the same technology that tools like MidJourney and Flux and Stable Diffusion and all the various AI image generator models were using.

This is a sort of new concept, a new way of generating AI images. Instead of a diffusion model, it is an autoregressive mixture of experts network trained to predict the next token from interleaved text and image data, and they look pretty good. Now, the realism isn't the same as Flux; Flux was much more realistic, but aesthetically, they're pretty pleasing. The colors are good, and all of the anatomy on the wolf looks good. Not a lot to complain about here.

It also has multimodal input, allowing it to take inspiration from or directly edit user-provided images. But from my testing, it hasn't been able to take images and modify images. Apparently, it's capable of that, but that doesn't seem to work quite yet.

Since we're on the topic of AI art, MidJourney got a little bit of an update this week with a new world-building tool called Patchwork. Patchwork is basically just a giant canvas where you can generate images, put them on the canvas, and collaborate with other people who can also see the images that are on this canvas. We can see some examples here of images and comments and text so that people can collaboratively, I guess, storyboard things together directly inside of MidJourney.

If you want to try it out, you can actually go over to patchwork.midjourney.com, and this is what it looks like here. So let's say I want to generate an image inside of this canvas here. I can click on this little image link here, click on my board somewhere, a half-wolf, half-man, and click paint, and it's going to use MidJourney to create some variations of this image of a half-wolf, half-man.

Let's say I like this one here, you know, give this a character name and sort of organize it with some notes here. It's just a way to sort of lay stuff out, organize it, add notes, and then give other people access through the little share button here, and then multiple people can collaborate on this giant canvas.

Adobe rolled out a new feature that allows you to take reflections off of images. So if you take a picture of something through a glass window and you've got the sort of glass reflection on it, it uses AI to remove that reflection to make it look like you took the picture on the other side of the glass.

Now, right now, this only works with raw images, where you'll be able to use JPEGs and HEICs later, and it looks like it's in Adobe Photoshop and Adobe Bridge. Now, this is something I really want to see in video because a lot of times I'll shoot video and I have to shoot it through a window, and I want to remove the reflection in video. So hopefully, they figure out this technology for that because that would excite me way more than photos.

YouTube's new dubbing feature has rolled out more widely now. This is really powerful because now any video in any language can be translated and dubbed into any other language, sort of opening up the possibility for anybody's YouTube videos to be put in front of a much wider audience than they used to be able to be put in front of.

This week, Cognition Labs finally released Devon, their AI coding assistant. This is something that was showed off months and months ago, got people really, really excited, and then didn't get released for a while. When they finally did release it, they released it at $500 per month. I guess if it could really do the same things as like a junior developer, then some people might find $500 per month worth it.

I'll probably pay for it for like one month to test it out and make some videos about it, and then I don't think I would stick with it. I don't code enough, and plus I do pay for Cursor, which really works well for me at $20 a month. So I have a hard time getting behind the pricing of Devon, but supposedly it is a lot better at seeing your entire code base, and it's apparently a lot more hands-off than tools like Cursor and WindSurf and some of those other tools.

I don't know; we'll see. I haven't tested it yet. We'll put it through its motions and see what it's capable of in a future video. Unfortunately, it did already run into an issue. I think this guy's name is the Primagen; I'm not sure if that's pronounced correctly, but he was actually on a stream playing around with Cognition and leaked a huge security concern.

I don't know exactly what the security concern was because he stopped the stream and removed the video on demand so other people couldn't exploit the issue, but apparently, there was some sort of big issue that would have been really easy to exploit Devon, which isn't great if you're paying $500 bucks a month for it. I'm sure whatever it is is patched up by now, though. Hopefully, we've got nothing to worry about with it.

I actually thought this was pretty funny. This AI company put up these signs in San Francisco that say, "Artisans won't complain about work-life balance. The AI of AI employee is here." So they're actually putting up ads around San Francisco encouraging people to hire AIs because they won't complain. I mean, the ad got attention; it got videos like mine talking about it.

I want to quickly talk about the world of virtual reality and augmented reality because there's been some new updates in that world that are pretty exciting to me this week as well, including the fact that if you have a Meta Quest and a Windows PC, you can now connect them together. Just like the Apple Vision Pro with Mac computers, you can have a giant workspace in front of you on your Meta Quest.

I haven't tested this yet; I do have a Meta Quest 3. I am excited to test this out, but you can sync up your Meta Quest with your PC and have monitors all around you in this sort of virtual workspace. Google is getting into the XR game on a much deeper level as well because this week they showed off Android XR.

Here's an example of somebody watching this immersive experience in VR inside of this new Android XR headset and actually being able to look around. This is an XR experience directly available on YouTube. Here's an example of somebody actually using Circle to search directly in XR, basically having this virtual desktop around them and circling things that they're seeing on their virtual desktop to be able to shop for them.

Here's somebody looking through their Google Photos directly in virtual reality or augmented reality and flipping through their photos and even, you know, adding depth to them, similar to what you get with the volumetric pictures on Apple Vision Pro. And here's somebody watching Google TV directly inside of this Android augmented reality experience, where you can just have a big floating screen in front of you.

So basically, it looks like this Android XR is trying to replicate pretty much what you get from the Apple Vision Pro, and they're working with Samsung and Qualcomm to do this. Qualcomm will likely be developing the processing chips that go into these devices, and Samsung will likely be developing the devices themselves that you actually wear. They look to go head-to-head with the Apple Vision Pro and the Meta Quest.

I'm excited to see how that plays out because that's another area that I'm really interested in. I go to all of the Augmented World Expos and keep close tabs on what's going on in the world of XR, so this is something that's really exciting to me.

And then finally, let's end with robots. Here's a video of a Tesla robot learning to walk down hills. In the first video, it showed it kind of slipping down the hill. In the second video, it got a little bit better, and then in the third video, it was able to easily and pretty effortlessly walk up and down bigger hills. So these humanoid robots are getting better and better at being able to traverse the world that we live in.

So I'm always excited to see that. I always like to end my videos with robots because robots are fun. And that's what I got for you today.

One little bit of housekeeping I wanted to share: I am going to start doing live streams. I've been getting a lot of requests on X to start streaming so I can show off tools like what I was showing off on my phone and so we can test things like Sora in real time. I get early access to a lot of various tools because companies want to give them to me to demo early, and I just thought it would be fun to start doing a fairly regular live stream to test out some of the tools, talk about some of the news, do an AMA, and chat with the community and things like that.

So starting on Monday, December 16th, at 11:00 a.m. Pacific, I'm going to go live on this channel, and the plan is to make that a regular thing. Just start going live every Monday at around 11:00 a.m. Pacific, talk about whatever tools are out that week, put them through their motions, take requests from the chat on what I should try, and we'll play with these tools in real time. You can see how they work, how fast they are. I won't be able to speed up the editing or anything like that; we'll be demoing this stuff in real time, taking your suggestions, trying to break them, trying to figure out what they're good at and what they're bad at.

You can ask me anything—ask me about a tool to solve problems in your specific business and things like that. I just really think it would be fun to get in the habit of doing some live streams and just nerding out about AI in a live environment. If you're interested, I will link you up to the sort of reminder page where you can be reminded of the live stream. I'd love to have as many people as possible joining me and having fun nerding out about AI live with me.

But that's what I got for you today. I hope you enjoyed this video. If you did enjoy it, I'd love a like and a subscribe, and maybe even the little bell notification thing so you see my future videos. I will make sure more cool AI news and tutorials and the various live streams and things like that show up in your YouTube feed, and it makes me feel good and really helps out my channel if you do.

Finally, if you haven't already, check out futurtools.com, where I curate all the cool AI tools that I come across. These are a lot of the tools we'll be playing with on the live stream. I keep the AI news page up to date pretty much daily. I talk about way less news in these videos than what I share here because every week there's just way too much news to squeeze into these videos. If you want to keep looped in on all of the latest AI news, you can check out the AI news tab and join the free newsletter, and I will email you just the most important news you need to know and the coolest tools that I come across, and it's totally free.

You can find it over at futurtools.com. Thank you so much for tuning in and nerding out with me on this one. I really, really appreciate you. I'm having so much fun diving deeper and deeper into these AI rabbit holes. Hopefully, you'll come along for the ride with me for these future rabbit holes that I dive down.

And thank you once again to Hostinger for sponsoring this video. This video was way longer than I wanted it to be, but there was so much cool stuff. Like, this was a week that really felt like it moved the needle. A lot of past AI news videos felt like really incremental improvements; this week felt like a leap.

So thank you so much for tuning in with me. I really appreciate you. I'm done rambling. I'll see you in the next one. Bye-bye!