Transcription
I think Google AI Studio is one of the most slept on AI tools out right now. It lets you do a ton of interesting things you won't find in other tools, and it's completely free. It's basically Gemini but with way more power and flexibility built in. So, I'm just going to jump right in.
This is at studio.google.com. This is a playground-style environment, so there's a lot more customization and tools, but that also means it can feel a little overwhelming at first. Don't worry though, it's easy to navigate, just not the prettiest interface. I'll walk through what all of this is throughout the video, but just know you can ignore half of it. And there are four main areas.
Chat, which is a standard chat interface, but with some really unique features. Stream, where you can interact in real time using your voice, your camera, or even by sharing your screen. Super useful. Generate media, which lets you create images, videos, and audio with prompts. And build, where you can create full apps using just natural language. Gemini writes the code in the background. You can make actual tools, or in my case, I built a full game in just a few prompts. No matter how many times I do that, it still blows my mind it's possible.
I'm going to start with chat. And before I dive into all the settings and customization options, I want to show you one of my favorite features that you won't find anywhere else: video input. A lot of models are multimodal now. They can handle text, images, and audio, but very few can actually understand full videos as input. This one can. I actually was going to make a whole separate video just about this ability, but the first use case for this that I really like is reverse engineering video prompts.
Here's one of those AISMR videos that went viral. It was a while back now, but I'll go with it. Drop that in. Then I'll type, "Give me a video prompt to create this exact video. It should include the appearance, camera style, actions, and audio cues." The prompt is for V3, which generates the video and audio simultaneously. So, I'll send that. And now here's the crazy part. Gemini actually watches the video. It sees everything that happens frame by frame while also listening to the audio. And here's what it gave back. And that would work, but I think it's a little long. So I'll say, "Condense that into a video prompt format." All right, that looks good. I will copy that. Open up Flow. I'll paste it in and then let it generate. [Music] That looks really close on the first try. But if it's not, here's what you can do. Download the new generated video. Drop it back into the same chat and say, "Compare this to the original. What's missing? Rewrite the prompt to fix it." It'll watch both videos, analyze what changed, and rewrite the prompt with any needed tweaks. You can iterate like that until it's spot-on. And this works with any video generator, not just V3, which is one of the most expensive options out there.
This video feature has a lot of use cases. It works with much longer videos than that last example, and you don't have to upload a file. You can also drop in a YouTube link. To show that, I'll use OpenAI's YouTube page. So, here's their recent announcement about study mode. If I open the video, you'll notice there's no audio. There's no narration, no voiceover. It's entirely visual. So, this is the perfect example to show that Gemini isn't just grabbing a transcript or pulling context from a few still frames. It's actually watching. And this is a fast-moving video with a lot going on. So, I paste the YouTube link here. It starts analyzing the video. Then I'll use a prompt like, "What is this video about? Explain it. Then tell me who this would be interesting to and why." Something that would help me if I was making a video about this. And in just a few seconds, it's watched the whole thing and returned a full analysis. It breaks down everything that happens on screen. And to prove it was actually watching, let's fact-check one of its lines. It says, "A user asks a broad conceptual question about why catalysis is a recurring topic in their chemistry class." I'll jump back to the video and, yep, there it is. Right there on the screen.
This is super useful, and it does work with much longer videos, but when you're dealing with hour-long content, you might not want it to watch the whole thing. So, back on their YouTube page, there's a podcast. I don't really want to watch this. So, I'll grab the link and drop it in. You can see up here, we're just over the token limit. So, this video is about an hour long. So, that means it could handle around 55 minutes before hitting the limit. But in this case, I don't need it to watch it. I just want the transcript. So, instead of wasting all those tokens, I'll just go back to YouTube, copy the transcript, and I'll leave all the time stamps turned on. Then, I'll paste that into the chat and ask for a summary with key insights and interesting quotes. Now, up here, since it's just text, the token usage drops way down from over a million to around 19,000. And looking through the quotes, yeah, I probably would not have made it through this whole episode, but this one's good. "Adaptation is the new job security," and it has a time stamp, so I can jump right to that moment in the video. "I think everyone's worried about jobs as well, but adaptation is the new job security." And this is helpful for me since I need quotes from videos sometimes. This way, I can get them faster.
That was just using the transcript, but there are a ton of other ways to use the video feature. For example, I use it to grab timestamps for my own videos so I can quickly add YouTube chapters. Or say you have like a presentation coming up, you could record yourself, drop the video in here, and ask for feedback on your delivery. Or maybe you're trying to document a task for someone else. So, just screen record yourself doing it, then ask it for a step-by-step breakdown of what you did. Now, you've got an SOP or a training doc without writing it all out yourself. Those are just a few examples. Everyone's going to have their own use cases depending on what they do. But once you get the hang of this, it's easy to see how powerful this can be.
A fast way to level up with Gemini is with this free resource provided by HubSpot. It's called Google Gemini at Work and it's a complete breakdown of how to use Gemini to speed up your research, improve your content, and build entire marketing strategies in a fraction of the time. My favorite part is the Gemini marketing stack. It shows how five tools like deep research, Notebook LM, and Gemini 2.5 Pro work together to handle everything from campaign planning to content creation to building interactive dashboards, all without needing a big team or agency. There's also a 4-week rollout plan if you want to actually start using Gemini in your workflow right now, plus some copy-paste prompt templates you can drop into the chat right away. If you're in marketing or just curious how far you can push AI tools, you'll definitely want to grab this. The link to download it is in the description. Thanks to HubSpot for sponsoring this video and making these free resources available to everyone watching.
So, I know I spent a lot of time on the video input, but that's a really standout feature that just isn't available anywhere else. So, I felt that was worth diving into. And beyond that, the chat tab can do everything you'd expect from a modern AI model, just like ChatGPT or Claude. You can do normal prompting, upload images, PDFs, all the standard stuff. But what sets it apart are the extra features and customization options that give you more control. So, I'm going to quickly walk through all the other settings in here, but I'll spend a little more time on the ones I find the most useful.
Let's start with the settings on the right-hand sidebar. First up is the model selector. You can ignore the API pricing unless you're building an app with this, but the "best for" section and use case description can be helpful when deciding. You'll see older models in this list, too, but most of the time you'll want to stick with 2.5 Pro and Flash. Those are the newest and best models right now. Below that is the token count. So, this shows how many tokens you've used in a chat, and every new prompt or question you ask in that thread adds to the count. And the context window is huge. It's a little over a million tokens. That's around eight times more than the web version of ChatGPT. That's especially useful for working with long videos or big files.
Then we've got temperature. This is basically your creativity knob. Lower temperature gets you more deterministic, safe, and accurate responses. You know, great for math, coding, or fact-based queries. Higher temperature gets you more random, creative, or surprising outputs. Good for brainstorming, poetry, or experimenting with writing styles. Next is media resolution. This affects how well the model understands images and videos. Higher settings give it more visual detail to work with, like it's watching more closely. In most cases, just leave it at the default, which is all the way up. The main time to lower it is with really long videos like that podcast earlier. At the default resolution, this goes over the token limit, but switching to low cuts token usage by about 2/3. That's really the only time I've ever lowered this.
Under thinking, we've got a couple options. Thinking mode is on by default for Pro and Flash. It improves reasoning and multi-step planning, but it also increases the token usage and slows things down a bit. With Pro, it's always on, but with Flash, you can toggle it on and off. Then the thinking budget lets you cap how many tokens the model uses to think. I usually leave this off and let the model decide.
Then we've got tools. Most of these I don't use much except for grounding with Google search. This tells the model to pull results from Google. It helps reduce hallucinations and adds real citations. You might use some of the others, especially if you're building apps. Structured output lets you constrain the model to output formats like JSON. Code execution lets it run Python directly inside the chat, then improve outputs based on the execution results. Function calling lets you connect to external tools or APIs. URL context lets you input a specific URL or multiple URLs for the model to read directly instead of going and searching Google. That is great for pulling data from specific pages or comparing multiple sources.
And one more I want to show under advanced settings is safety settings. So, click edit, and you'll see that it defaults to a more raw version of the model with fewer guardrails than what you'd expect in something like Gemini or ChatGPT's web interfaces. Since AI Studio is built for developers, you can adjust the safety filters based on your needs. Personally, I leave all the safety filters off.
Up top, there are a few more features, and a couple of them are actually really useful. The first is the system prompt. This lets you set the tone or role for the chat, like telling it how to respond, what style to use, or what type of assistant it's supposed to be, giving it background instructions before your main prompt. That way, you don't have to include it every time. Then, next to that, you'll see options for SDK code, prompt sharing, and save, but the one I actually use is compare mode. This opens up two chats side by side, and it's a great way to test how different models, settings, or system prompts affect the output. For example, I could use the same prompt on both sides, but turn the temperature way up on one and leave it low on the other. Then, "Give me five hooks for a YouTube short about Google AI Studio." On the left, I get solid, safe suggestions. And on the right, things get a little more creative. "What if you could build AI like magic?" "Stop scrolling, start creating." "Ever wonder how AI gets made?" Those are solid, and that's a simple example, but seeing them side by side really helps you get a feel for how temperature and system prompts affect the output and which settings might be best for what you're doing.
And the last thing I'll mention on the chat tab is this prompt gallery on the right. It has a bunch of preset prompts for different use cases. You can click each to see the example, but it doesn't look great within the UI right here. They also have all these on the ai.google.dev site as well, which looks much nicer.
Now, moving on to the stream tab. This is where you can interact using your voice, your webcam, or screen sharing. And each of these is useful in different situations. The screen sharing feature in particular has gotten a lot of attention. It went viral on its own. I actually made a separate video just about that. But first, I want to go over the settings really quick.
First up is the voice drop-down. There are about 30 different voices to choose from. "Ready to build something awesome today?" "Got a project in mind?" "What do you want to explore?" "Ready to make something amazing?" Then you've got some additional toggles. Turn coverage. This allows it to send audio or video input even when you're not speaking. Effective dialogue. This enables Gemini to react not just to what you say, but how you say it. It'll adjust its responses based on your tone of voice. Proactive audio. This tells it to ignore background speech that isn't meant for it. So, if you're talking to someone else while it's running, like in a meeting or driving through a drive-thru, it won't randomly chime in. And I'll turn that one on while I demo this, so it hopefully won't respond while I'm explaining things during the recording.
Starting with talk, this lets you have a full back-and-forth conversation with Gemini using your voice. For some people, this is the preferred way to interact. And in a lot of situations, it's just easier. "What's the difference between you and Chat GPT?" "I'm a large language model developed by Google, while ChatGPT was developed by OpenAI. We both have similar goals, which is to provide helpful and informative responses, but we're trained on different data sets." "Well, which one of you two is smarter?" "It's hard to say definitively which one of us is smarter. We both have strengths and are constantly being updated and improved. Our capabilities are continually evolving, so it really depends on the specific task or question." You get the idea.
Next is the webcam. I usually find this more useful on my phone rather than my computer's webcam. That's just looking at me, and I'd have to like hold things up to it. For the dedicated video on this, I had it help me with repotting a plant using my phone. "Yes, I'd like to repot it. Can you help me with that?" "Sure, I can help you with that. First, do you have a new pot and some potting mix ready?" "I do. Um, and this is the type of plant I have." "Okay. It's a spathofilm, also known as a peace lily." "That's good to know. Do you have your new pot and the potting mix ready?" "Yeah. Is this potting mix good?" "Yes, the Miracle-Gro organic indoor potting mix should work well for your peace lily. Now, are you ready to start repotting the plant?" "Yeah. Should I add some of this stuff to it?" "Yes. Adding perlite to the potting mix can help improve drainage and aeration, which is good for peace lilies. Do you want to mix it with the potting mix now?" "Yeah, let's do it."
The other option here is screen sharing. That's where Gemini can see everything that's happening on your screen while you're talking to it. This can be really helpful. So, I'll go ahead and turn it on now. So, I'll select this Premiere Pro window. "All right, I have this project open in Premiere Pro. And I've also got this little logo right here. I want the logo to kind of pop up on the screen and then do, I don't know, like a little wiggle effect or something. How do I do that?" "You can achieve that with key frames in the effect controls panel. First, select the logo clip on your timeline. Then go to the effect controls panel, find the motion properties, and set key frames for position and scale to create the pop-up effect. You can also animate the rotation property for a wiggle effect." "That aren't available." "All right, I have the logo kind of popping up. How do I make that a little smoother?" "To make the animation smoother, try adjusting the key frames. You can right-click on a key frame and select ease in or ease out to create a smoother transition. You can also adjust the spacing between key frames to change the speed of the animation." "All right, that's pretty good. How do I add that wobble again?" "To add a wobble effect, go back to the effect controls panel, find the rotation property, and add key frames with slightly different rotation values to create the wobbling motion. Adjusting the timing of these key frames will control the speed of the wobble." "Okay, that looks pretty good. Anything else I can do to make it better?" "You could also add a subtle bounce effect by adjusting the scale key frames. For example, have the logo slightly overshoot its final size and then bounce back. This can add a bit more emphasis to the pop-up animation." "Anything else?" "You could also add a sound effect to accompany the logo's pop-up, like a subtle chime or whoosh sound. This can enhance the overall impact and make the animation more engaging." "All right, so that worked pretty well. Just a quick little demo, but we've got the entire history of everything we did right here. Just showing every step of that."
And that worked perfect for learning new software. It can be helpful, but I wouldn't recommend using it to learn something entirely from scratch. In my experience, it gets too many things wrong to rely on as your only guide, and that can get frustrating really fast. Like, when I was having it teach me Blender, it sent me on this long chain of steps that was never actually going to help me do what I needed. It ended up being a total waste of time. What worked better was learning the basics first from a YouTube video or a course, and then use Gemini while working in the tool. At that point, you can just ask things like, "What's the name of the tool that does X?" or "Does this process make sense for this outcome?" And that's a lot more useful.
Beyond learning tools, there are a bunch of other great use cases for screen share. You could ask it to explain a confusing diagram or chart you come across online, could get help while live coding, could run a quick UX test, get critique on a landing page or website layout, or even just troubleshoot something that's not working. There are a lot of ways to use this, especially if you're already working on something and just need help as you go.
The generate media tab lets you create and edit images, generate videos, create great text-to-speech with multiple speakers, and a pretty interesting music feature. And a particularly useful part here is the image editing section at the end. But I'll walk through each of these features.
Starting with image generation. This uses Imagine 4, which is a solid model with great prompt adherence. You get a limited number of free generations for images and videos, especially videos, but for testing, it's enough. They've got a few example prompts down here. I'll click one. That looks great. Nice image quality and even a little text rendering. Let's push that further with a custom prompt. Over on the right is where you can change the aspect ratio. And then I'll add my prompt. This one's a Vogue magazine cover featuring a capybara. It's got specific text I want included. And that came back perfect. It nailed every line of text I asked for. And now let's try a more realistic style. This is a longer prompt, too, but it's a runway model wearing an octopus dress. Again, that's a really solid result. I'll do one more. Got a tiger eating at a restaurant using chopsticks. All right, nailed it. Again, I won't go too deep here. It works like most image generators, but it is a really solid option, especially for how well it follows prompts. And it downloaded those images.
So, let's move over to VO. It's a video generator. This currently uses V2, which doesn't support audio generation like V3 does. No sound effects or dialogue, but it still produces solid videos from image or text. I'll start by animating that runway image. "Let's do woman in octopus dress walking down the runway." Not bad. It's a little blurry on the tentacles, but good overall. Now, let's do the tiger image. Also, not bad for a first try. You can do text-to-video as well. So, I'll reuse the tiger prompt, but just modify it a little for video. I know it's missing the chopsticks, but it looks great. So, we'll try that octopus dress prompt again, too, but this time from scratch. No image. That looks good. Although, I liked the one where the dress looked more like a full octopus. And if I try to send one more prompt, yep, I have hit the limit. You get four video generations per day. So, not bad for free.
Now, next to VO is Gemini's image generation, which also includes image editing. And there's a lot of use cases for this, so I'll cover a few examples. So, I'll upload an image of my dog, and I thought this was a funny one that I came across. "Make a professional passport photo for this dog." Perfect. Now, Zuko can travel. All right, let's try one of me. "Give this man a face tattoo that says FutureMedia." All right, neck placement. Might need to book an appointment. And now, let's try removing people from a photo. And this is a particularly difficult one, and it handled that well, even with them covering most of the shot. And this type of editing works especially well on AI-generated images. For example, "Change the octopus dress to blue." It does a great job adjusting visual details while keeping everything else intact. There are lots of ways to use this one.
But moving on to the speech generation. This is a really high-quality text-to-speech system. You can use multiple speakers, customize styles, and guide delivery. And I'll just type in some quick, pretty lame dialogue and run it. "Hello. We're excited to show you our native speech capabilities where you can direct a voice, create realistic dialogue, and so much more." "Being able to use multiple speakers makes this a very useful tool compared to many of the others." "That's right. And we sound so natural, don't we?" "Indubitably." "Right." "That sounds great. Let's do another one." I'll switch it up this time. Over here, you can choose different voices. So, I'll do Atronar and Jedar as speakers. I'll give it new style instructions and change up the dialogue. "AI is getting out of hand. It just won't stop. There's way too many tools to keep track of. I need an AI to keep track of all the new AIs." "For real, these are really solid results. That is great text-to-speech quality, especially just built into all this other stuff."
Now, I'll go back again. The last one here is Laria Realtime. This lets you interactively create, control, and perform music in the moment. And if you look at the tabs over on the left, you'll notice when I click this, it switches over to the build tab. That means it was generated using this feature. I'll mess around really quick, then show you how you can build something like this yourself in here. So, you turn up and down any of the knobs, and it will guide the type of music that's being generated in close to real time. "Got to have some thrash. Throw in a hint of shoegaze and a little trip-hop. Then play." All right. I'm not hearing the thrash. Got to turn that up. [Music] "Down on the dubstep." [Music] All right. There it is. All right. That's interesting. All right. So, that's fun to mess with for a little bit. I'm not going to sit here and listen to this like an album, but it's really cool that this was built entirely in AI Studio. And it is surprisingly simple to create something like this yourself.
So, I'll click back to the main build screen. This is where you can create many apps and tools just by describing what you want in natural language. I'll show you how to build one from scratch in just a minute. But first, let's check out some of these examples. Scroll down and you'll see a bunch of featured apps built right here. Everything from games and dictation tools to music generators, a gift maker, map planners, and more. You can click into any of these to try them out, see what's possible, and maybe get some ideas. Let's try this co-drawing app. Looks like you sketch something and it generates based on your drawing. I'll try to draw something. Okay, there it is. It's hard to draw in here, but now I'll add the prompt. "Make this look realistic." Nice. That was pretty good. Not sure about the black background. I kind of want him underwater. Okay, now that completely changed the fish, but it kept the hat. All right, let's try something a little more practical. "Make this into an elegant, modern logo." All right, that was pretty good considering what I gave it. Anyways, I'll jump back out. Let's try one more. This is a flashcard maker. "I want to learn about music history." And then it generates a full set of flashcards I can quiz myself on. That could actually be useful. So, there's a ton of stuff you can do in here.
I'm going to try building a game. You just type directly into the prompt box using natural language. I will go with, "Create a game just like Pac-Man, except the main Pac-Man character is a picture of Ozzy Osbourne, except as a pixelized video game representation of him. Then instead of ghosts, they're bats." That's good. Normally, you'd want to think through the game a bit more before prompting, but I'm just going to send it straight like that. And from there, it jumps into planning mode, refining the concept, outlining the logic, defining the maze, game mechanics, and components. Once it finishes thinking everything through, it starts writing all the code. Does that for a little while, and then it will check for errors, fix what it sees, and keep going. In total, it took about four minutes to generate a fully playable game just from that one prompt. And that's still crazy to me. Let's test it. Okay, looks like it's working. I can move around. It plays like Pac-Man. Let's eat a power pellet. Nice. The bats turn blue, but it looks like eating them doesn't work yet. And I guess I only have one life. All right, no problem. I will just ask for the changes. "He should have three lives. Also, the bat logic doesn't work. When they are eaten while blue, they should turn into just the eyeballs and go back to the start, just like in Pac-Man." So, it's going to think through how to implement that and then it will update all the code. Edits usually run faster than the initial generation, about a minute.
All right, I gave that another test. It fixed some things, but I need to make some more changes. So, I'm going to just jump through these a bit faster here. The bat eating logic only worked sometimes. It needs to work every time. Also, don't pause the game after each death. Ran that. After that one, the eyes were just getting stuck in one spot. So, we fixed that. Then, I asked it to change the way the bats looked since it looked like they have four wings or something. Then, I fixed the part where Ozzy would go upside down and added a tracker for the lives that are doing the devil horns rock-on symbol. Then, for the final touch, we needed some music. I said it should fit the game, so "Ozzy style metal mixed with 8-bit Pac-Man type music." Then I tested that and needed one more change to have the music switch whenever Ozzy eats a power pellet and the bats turn blue. Then switch again when they turn back. And that was it. Now I'm going to do a test of the final completed game. [Music] "Fire." "Hey." "Hey." All right, that turned out pretty awesome. There's still tweaks I could make, but it works and is fun to play. And with these games or apps, you are able to share them. So, I'll leave a link to this down in the description. I gave it a couple tries. My high score was 3,140. Leave a comment if you beat that.
Now, I have covered all the features and settings, but that just barely scratches the surface of what you can do in Google AI Studio. And again, everything I've done in here was completely free, including those things you can't do anywhere else. So, I think this is a great platform to mix into your AI toolkit. And there is the caveat to it being free that I'll mention. Google will use the things you do in here to train their systems. This is the case with essentially all free AI tools. Similarly, with your data on any free software or website, you unless it's open source and you're running it locally. I assume most people know that by now, but if I don't say it, someone always brings it up in the comments. So, there you go.
And if you want to go deeper into learning AI, we've built a full course platform at Futureedia with over 500 lessons across over 20 AI courses. You'll find full learning paths on ChatGPT, prompt engineering, automation, custom GPTs, video generation, coding with AI, and a lot more, all included in one subscription. You can get a 7-day free trial using the link in the description. Or check out this video with an entire roadmap on how to learn AI.