📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Google AI Studio In 26 Minutes

Tina Huang26:29

Transcription

I learned how to use Google AI Studio for you, and here's the cliff's notes version to save you the hours and hours that I spent learning about it and experimenting, including actual realistic prompts that you can get started using to increase your productivity and to start building your own applications. Google AI Studio is honestly a really powerful tool, and I kind of regret not properly learning about it a little bit sooner. It's just because, like, the UI can feel quite overwhelming because there's just so many things that are going on. But trust me, like, once you actually learn how to use these tools and understand how it works together, it opens up a lot of things. It's like a whole new world of possibilities. As per usual, I have little quizzes throughout this video to help you actually retain the information I'm going to be talking about. And yeah, without further ado, let's get started. A portion of this video is sponsored by HubSpot.

Okay, so when you first go to Google AI Studio, this is what you're going to see. Do you see what I mean? There's just like a lot of stuff that's going on. It can feel pretty overwhelming, but don't worry, I'm going to be going through all of these step by step and also providing frameworks for how to actually use these properly. But first, I'm going to give you a quick overview. So here you have your chat interface where you can just type in your prompts. And because it's a playground environment, there's also going to be a little bit more additional features that you have control over, which we'll go over in a bit.

Then you have Stream Real Time, which is when you can interact with Gemini using text, voice, video, or screen sharing. Especially this one, I think is really cool. The screen sharing part, you're able to allow Gemini to see what you see on screen, and you're able to interact with it and talk with it based upon what's on your screen. Is there anything you would like to know about it? I want you to tell me what is happening on screen.

Okay, the video shows a football match between Arsenal and West Ham. The score is 0 to zero. Starter Apps is a great place to start where you can see the kinds of applications that you can prototype and build using Google AI Studio. And fine-tuning a model, you can actually fine-tune on Google AI Studio directly. And finally, there's a library where it shows the history of the prompts that you have and a prompt gallery that gives you some inspo for the kinds of prompts that you can try out. Google Studio is also free to use unless you're building like really large-scale applications.

So the most natural thing that you probably want to do is just to start typing something in as a general prompt, which is fine. You're directly telling the AI to do something very specific for you right now. But I actually want to spend some time talking about the system instructions here, or a system prompt, because I think it's something that is really slept on, is really underrated. A system prompt is more like setting an instruction guide for all of the interactions for that specific AI. You can very specifically define the AI's personality, capabilities, and limitations, making your results just much better.

The way that you define a general prompt is also slightly different from a system prompt. So to explain this, let me first just start with a tiny little crash course on prompt engineering in general. The framework that I like to use, I call it the "Tiny Crabs Right Enormous Iguanas" framework. It's a mnemonic that stands for Task, Context, Resources, Evaluate, and Iterate. And I promise you, if you can remember just this one mnemonic and what it stands for, you'll be better at just prompt engineering than like 90% of the population.

The Task is specifying what you want the AI to do. The Context is additional information and directions that you can provide to AI to have a better understanding of your task. Resources are examples or any additional information that you can provide the AI with that it doesn't already know. Then after you run the prompt, Evaluate is to see the results, and Iterate is to see what you can do to improve it.

Let me actually show you an example of this framework. So the task is that you want a marketing campaign for your new nail art collection. It's also important here that you specify that it should be an IG post caption, which is the output. Context is just any additional information that you can give to the AI. I've attached a picture of the nail art collection. I've also attached some inspo IG posts that I like. Here's the inspo. The inspo IG posts that you attach is going to be the resources where you're providing to AI with any examples. In this case, it's a few screenshots of IG posts, but it could also be like different PDFs that you're wanting to create a report or something like that, just anything that the AI doesn't already know and you're providing it additional information. Let's actually run it, and it gives us some options in the style of the examples that we provided.

Next is to Evaluate, and you might be like, "Okay, like, this is nice, but I want it to feel more personal." So we then move on to the final part of the framework, which is to Iterate. You might type, "I want the post captions to feel a little bit more personal, like a story." And there you go: "Remember those endless summer nights catching fireflies and dreaming under the stars? I poured that feeling into our new nail art collection. These glitters are like bottled starlight," etc., etc., etc. I'm like, "Okay, after this iteration, I quite like it now."

Okay, so that is a general good prompt engineering framework guideline, but let's actually go back to the system prompt. The framework for crafting a system prompt is not actually different, but there is a little bit more emphasis placed on different parts of the framework. Specifically, you want to put more emphasis on the Task and the Context. For the Task, it's not as direct as just saying exactly what you want it to do. You want to think about it more as the AI's overall role and interaction style. You want to consider the tone that you want the AI to communicate with. Do you want it to be formal, informal, humorous, professional? The length of its responses, should it be brief, should it be really detailed? The format of the output, should it be a list, a paragraph, code, prose, JSON formats? You also want to define a Persona, which is explicitly stating the AI's identity and characteristics. This is something that you can also include in your general prompts, but a lot of times you can get away with not doing that just because the task that you're asking it to do is so specific. But for a system prompt, you definitely want to include a Persona.

Some examples of personas would be like, "You're a helpful and friendly coding assistant." You might want to include some expertise, like, "You're an expert specifically in Python or JavaScript." Or you might want to put some restrictions, like say, "You're a medical expert, but you are not authorized to provide medical advice."

Now for the Context, this is the opportunity for you to provide even more additional information and instructions to the AI. You can specify like domain expertise, like, "You have access to a database of academic papers in the field of astrophysics." Give more details about what is allowed or not allowed to do, like, "You're allowed to access external websites to retrieve information, but you're not allowed to generate any responses that has anything to do with being sexually suggestive or exploitative." For example. The rest of the framework remains the same.

So let me actually give you an example of what a good system prompt will look like. Say you're interested in creating AI that is specific for providing recipes. Here is the system prompt: "You're a helpful cooking assistant specializing in simple one-pot recipes." So that's the Persona. "Your task is to analyze photos of ingredients and suggest recipes based on them." The Task. "Provide recipes in a clear, step-by-step format. Include approximate cooking times and how many people the meal should serve." This is specifying with more detail about the output. "Always stick to recipe suggestions and cooking advice. If asked about anything unrelated to cooking or nutrition, politely redirect the conversation back to their recipe ideas." And this is additional Context.

So let's set the system prompt, and I'm actually going to put a picture. This is an actual picture of what is currently in my fridge. The general prompt we're going to write is: "Analyze the photo and list the main ingredients that you see. Then suggest the meal, preferably as a stew or a soup. Um, not pictured, but I also have kimchi, Japanese curry cubes, and rice."

Okay, so it says that what I have here is tofu, eggs, lemon, soy sauce, and miso. That is correct. Um, and it says I can make a kimchi tofu stew with eggs. Kimchi jjigae inspired. Great, wonderful.

Okay, time for our first quiz. Please answer the question displayed on screen and put it in the comments. I really hope that by now you can see how important the skill of prompt engineering is. It can be the difference between getting the results that you want and just getting like, like some generic response. Hands down, prompt engineering is the highest return on investment skill that you can learn in 2025. So if you are looking to improve your prompt engineering skills, I highly recommend that you check out this free prompt engineering quick start guide that I made with HubSpot. It includes a step-by-step guide to creating great AI prompts and also tips to get better results. My favorite part is that there's a flow of bad to good to great prompts that exactly shows you how you can improve a prompt. If you're able to get yourself to start thinking in this way, you will become much better at prompting and also just be much more productive at everything that you do. So do check it out at this link over here, also linked in description. Thank you so much HubSpot for creating this free resource with me and for sponsoring this portion of the video. Now back to the video.

Okay, so now I actually want to talk about all the little features that you see over here that are actually really powerful. The first thing you can do is that you see how in the model, you can switch to models that are available. We're currently using Gemini 2.0 Flash, but we can actually click compare and be able to try out another one and compare the results, like Gemini 2.0 Pro, experimental, February 5th. Let's do that one.

Okay, so I'm just going to do the same prompt, click run, and we see that, um, Flash is faster, which is expected, and Gemini 2.0 Pro is a little bit slower in its response, but it is more detailed, like providing tips and notes as well. So I don't know if this is a fluke, but I actually identified the tofu is firm or pressed tofu, while here it was not specified what kind of tofu it was, and it actually is firm tofu. Also, I want to specify that because I'm in Hong Kong, the language that's being written, actually, I think some of it is Japanese, but the other parts of it is in Chinese. So it does work for things and labels that are not in English as well. Let's just stick to Gemini 2.0 Flash because it's faster, and I want to talk here about tokens.

So one of the biggest highlights of the Gemini models is that it has a huge context window at over a million. And just to give you an understanding of how big that actually is, this means that you can provide it with like five or six full-length books. People have tried doing that, and what's really cool is that it's also able to analyze video because video data is generally a lot larger. And I have managed to put in videos up to like four to five minutes. It's also really, really good, like better than any model I've seen, but I'm getting ahead of myself. I'll show you guys in a bit. Let's first get through the rest of these.

Temperature is the amount of creativity that you're going to allow the model to have. If you have higher temperatures, it's going to be more creative, and it's up to two. And if you have lower temperatures, it's going to be less creative. Structured output is when you want more control over what the result is going to look like, and you want it to be reproducible. Let me give you an example of when having a structured output would be helpful. So say you're a product marketer and you want to have AI help you a little bit. Here's your system prompt: "You're a product marketer targeting a Gen Z audience. Create exciting, fresh advertising copy for products. Keep copy under a few sentences long and use Gen Z lingo."

So if you were just to type something here like, say, "Create copy for a pair of men's black running shoes," and you run it, it would, you know, come up with like this kind of thing. Situation. But if you turn on structured output, for example, um, and you can actually come here and edit what it is that you want it to look like. And the reason why you want to do this is maybe you want the output to be a very specific format because you want to be able to plug this into a database or to an application that you're making. So this is especially helpful if you're a developer, um, because you can get the data to look a certain way to make sure that your app is consistent. So a certain, for example, here the code editor looks like this, but we can also use the visual editor, and you can add a property like here, we might want to say, um, I don't know, product name, and we want to make this required. And then you can also say the product copy, right? And it's also going to make this required. And you can, of course, change this into whatever it is that you want. Uh, we're going to keep it as string for both. And if you do that, you would get like this, like a JSON format, and you're able to copy that to a clipboard or to download it as well. And because it's in a JSON format, you can actually get the code for it too. Yeah, you'll be able to take this and then put it into your application. This is so helpful and really, really powerful.

So I'm not going to go and actually, like, make an application right now to, like, show you everything because this video is already probably going to be really long. But let me know in the comments if you're actually interested in me showing you how to build like a full AI application.

Hello, it is Tina from the future here. I just want to let you know that if you're interested into diving deeper into different AI topics, because I can't fit everything into just a single video, I also do AI workshops that are completely free. So if you're interested, you can sign up at this link over here, also linked in description. Okay, back to the video.

Anyways, in the meantime, let me go through the other tools that are available here. So code execution is pretty straightforward. You're basically allowing the AI to have the ability of executing codes, whether that be a calculation, whether that be if you're wanted to create an application for you, um, it can go and execute code to in order to do it. So a really simple example would be something like, "Just calculate how much money I would have earned if I put $1,000 with a compound annual interest rate of 9% for the past two years." You can put in code execution, and it would be able to do this calculation for you. There you go.

Next up is function calling, and this can be really, really powerful. It allows you to equip your AI with custom functions or tools so that it's able to either pull data or execute certain things. A simple example here is that you can have something called "get weather," and that can connect to an external API so it's able to directly pull information about a specific weather forecast directly from the API. Another example would be if you want your AI to be able to book restaurants for you, you might want to give it access to an API like OpenTable where it can be able to call this function and actually do the reservation for you. Another example, if you have a function that gives it access to the Google Calendar API, it would allow the AI to be able to schedule things into your actual Google Calendar. Function calling also really starts to shine when you're actually building AI applications. So when I'm doing that video, I will go into more detail about it. But moving along, the last tool available is grounding with Google Search. This gives your model the ability to browse the internet and to validate whatever it's outputting with results from Google. It's helpful for preventing hallucinations and also to give your model access to information that it doesn't currently have.

Here is an example. So this is actually true. I'm going back to San Francisco May 15, 2025 for a wedding. While I'm there, what are some interesting festivals, conferences, or events that I could also drop by and visit? I'll stick around until end of May, and I don't mind traveling within the US and Canada. Also, write, I'm interested in events related to education, tech, and just general cultural or music festivals. Then we're actually going to turn on grounding with Google Search so it's able to have information from Google to look up. It gave me some options over here, and it also shows what it is that it's searching, um, on Google Search as well.

There are also some more advanced settings like setting safety settings depending on the kind of responses that you want to get, a stop sequence, what it is that you want to potentially get the AI to stop if it reaches a certain sequence, the output length, how long you want it to be in terms of number of tokens, and the top K, which is another way of controlling the randomness of the output that is different than just adjusting the temperature. You can definitely play around with these, and it can be helpful if you're doing like very specific things, but I would say generally speaking, there's really not that much need to fiddle around with these unless, again, you're building an application or something like that.

Let's not talk about, let's now talk, let's not talk two hours later, let's not talk about video. Specifically, video stuff like video generation, video analysis, and things like that has been emerging for a while now, but it really is starting to shine, especially with Gemini. You can see why having such a large context window is able to allow you to analyze quite a lot of video content, and it does it better than what I have seen any other model do so far.

Here is a clip of a football match between Arsenal and West Ham United. I don't know anything about football, but it's a clip that I found as one of the highlights. So I'm going to take this 2-minute clip, put it into Gemini, and write: "Analyze the provided video highlight reel for Arsenal versus West Ham 0:1 Premier League. Could you first identify the timestamps of the goals and near misses, and then whoops, and then give a detailed commentary with timestamps?"

Oh my God, I actually even got the people's name right. I don't actually, I don't know if these names are right. So it says key moment timestamps, the near misses, the goal, and another near miss. I know that this is correct. That is correct. 47 seconds in, there's a goal. That is also correct, and it was in fact West Ham. And at a minute 49, there was another near miss. I don't know if it was actually these people that did it, but for anybody that knows football, are these the actual correct people? Yeah. And then it gives a detailed commentary about what it is that is happening as well. That is freaking cool, isn't it? Like, that is that just like blew my mind. I was able just to like analyze the video data, and that just unlocks so much potential for all the video content that we have available. Of course, in addition to video data, you can also upload like large files, PowerPoints, PDFs, upload all these resources as well to Google AI Studio, and the model does a pretty good job of extracting information from it. But before I move on to the next section, here is a little quiz to make sure that you're retaining all the things and information that I am talking about. Answer the questions in the comments.

Let's talk about Stream Real Time, and this allows you to interact with Gemini using text, voice, video, or screen sharing. So I'm sure that many of you guys have have tried out the voice functionalities, but using Google AI Studio, you can also show Gemini using your webcam and interact with it directly. And you can also share your screen so it's able to have context about what is happening in your screen and give you real-time commentary, directing you to things, giving you instructions, or just interacting with whatever it is that you're interacting with.

Okay, so we're going to do "Share your screen Chrome tab," and this one specifically. Okay, I see you're talking about a specific Chrome tab. Is there anything you would like to know about it? "What is happening in this clip right now?"

Okay, in the video, the players are moving around the field, and a player in black is on the ground having just been tackled. It looks like the player in red and white has gained possession of the ball. Also, the score is now Arsenal zero, West Ham one. So there is a little bit of a lag, but just the fact that it's able to watch that video with you and just comment on what's happening like that was so mind-blowing to me. You can probably see that this is so powerful as having like a very personalized specialist or personalized tutor just comment or teach you whatever it is and just constantly help you out and whatever it is that you're doing.

So another example here is I'm going to share a screen. Okay, I see that you are sharing your screen. Is there anything you would like me to do with it? "Yes, so currently I have this picture on Photoshop, and I want to remove the turtle at this very on this very back of the image. How do I do that?"

"Okay, to remove the turtle in the back, you could use the lasso tool to select it, then use the content-aware fill option to fill in the space. Would you like me to go into more detail about this?"

"Yeah, can you instruct me step by step on where to click and what to do?"

"Okay, first select the lasso tool from the toolbar on the left side of the screen. Let me, let me do a step-by-step. Hold on. So the lasso tool. So you want me to draw a selection of the turtle here?"

"Yes, that is correct. You will want to use the lasso tool to draw a selection around the turtle in the very back."

"Okay, I did that. Now what?"

"Now that you've made a selection with the lasso tool, go to the Edit menu at the top and then select Content-Aware Fill. This will remove the turtle and try to fill the area with what's around it."

"I don't see Content-Aware Fill."

"Oh, I see my apologies. Oh no, I see it here. It's grayed out over here."

"Okay, it seems the Content-Aware Fill option is grayed out. This means that you likely have an object layer selected. Could you please select the image layer and try again?"

"I do have the image layer selected. Then F and then the cont cont you to right and and go edit then fill and cont contare and press okay."

"Oh, okay. So now I can see it. So I now I do Content-Aware Fill?"

"Yes, now you can use Content-Aware Fill."

"Oh, nice. Okay, cool. And then I can just click okay. Nice."

"Yes, you can click okay to apply the content to Weare fill."

"Yeah, there you go. I played around quite a lot with streaming old time, and it is honestly like so incredible. Like it blows my mind that people don't know about this, and the fact that this is free. Like all the use cases that you can think of whenever it is that you want someone to just teach you something directly, whenever it is that you just wish someone would just like show you how to do something as opposed to watching like 50 tutorial videos about something. This this can do it for you. When you're coding, you know, you don't know how to do something, you don't know how to, whatever it is, you can use Stream Real Time to do that. Um, whenever you're looking at some Excel file, don't know what the function is that works as well. Even using like the, um, webcam style, like you're drawing something, you're asking it to critique what it is that you're drawing. It's just so many different use cases. Well, I'm curious to know what kind of use cases that you've tried out where you can suddenly imagine this being able to help you with. Let me know in the [Music] [Music] comments."

The next tab is some Starter Apps. So this is where Google provides some examples of some of the apps that you can prototype and build out using Gemini, uh, and using Google AI Studio. So I just want to show you like, for example, Map Explorer, which is really, really cool, and this is something that you can get the code for on GitHub as well. So for example, you can say, "Take me somewhere ancient," and it will tell you that you can go somewhere like, go, I don't know how to pronounce, in Turkey. This archaeological site is home to the oldest known megaliths in the world, and it gives you the information here. This is connected to the Google Maps API. And you can also type something like, "I want to go somewhere close, close to Hong Kong that has a lot of hiking trails and is relatively remote." And here it is: "Crooked Islands offers pristine beaches, lush forest, and challenging trails," and you can explore this area, zoom into things as well. Pretty cool, right? Yeah, definitely play around with some of these starter apps and just feel very inspired by all the things that you can build.

Okay, time for another quiz. Please answer these questions and comment in the comments.

Moving on now to the final really cool feature from Google AI Studio, which is Tune a Model. You can fine-tune a model directly within Google AI Studio, and you don't even need to know how to code. So why would you want to do this? The reason why you might want to fine-tune a model is if you want to increase the performance of the model for a very specific task. The way that fine-tuning works is that you provide the model with a training data set that contains many examples of a task. This is especially useful if you have a very niche task, and fine-tuning can allow to have significant improvements in the results. It's recommended that you have 100 to 500 examples, and you can create a structured prompt. So the data that you want to provide should be in two columns: the first one is the input, which is the user's input, and the output, which is what the model's response is. So one way of doing this is that you can just directly add test examples into the UI itself, or you can import a CSV. They also provide some samples of what this can look like. So we can try out one of the samples here. So say we're going to try out the news headline. So this is what the input and output looks like. So the input here is that you have like a huge chunk of things, blah blah blah, like all this information that is here, and the output, you wanted to say, "Just mobile games come of age." You can call this model whatever you want, so we can call it "News Model," and the model that we're going to choose here is the 1.5 Flash, and then we can just click tune.

So in your library over here, it is fine-tuning, and look, yay, it looks like that you can use your tuned model now. So you can use it in chat. Copy paste some new segment and let's run it. "Warning receipts are fast future contracts," which is in the style that the training data set is on.

Some other useful use cases for when you would want to fine-tune a model is say, for example, if you want to have some specialized ranking system, like maybe you have a bunch of comments and you want to rank it from like one to five, but in a very specific criteria sort of way, that can be useful by providing examples and fine-tuning. Maybe you want to translate medical note shorthands, or go through a lot of documents, like legal documents, and extract very specific components from each of these documents. It may also be useful in customer support if you're selling like a very specific type of product, and the way that you want to interact with a customer is very, very specific, that can also be useful. And actually, after you fine-tune a model, you can also easily incorporate that into an app. Like say, if you want to build, um, a customer support agent out of it from your fine-tuned model, you can just directly add API access where you can just go here and get your API key and have access to the model programmatically. Again, if you want me to do a video in which I'll show you how to build these applications, let me know in the comments. If you let me know in the comments, then I will know that there is interest and I will feel motivated to do [Music] it.

All right, I think that was a pretty comprehensive overview, tutorial, slash, I don't know, run-through of Google AI Studio, and I really hope that you're able to see how much potential this tool has and you feel inspired to try it out yourself as well. Before I end this video, I'm now going to put on screen the final little quiz. If you were able to answer all of these questions from all these quizzes, then congratulations, you are now quite educated on Google AI Studio. Thank you all so much for watching, and I will see you guys in the next video or live stream.