Transcription
So efficient and it's so fast and it's so smart and it's so flexible.
And within like 5 minutes, it had emailed, messaged, cleared, and managed everything for me. Imagine having this one product that literally can take all of your information, no matter where it comes from, and manage it in one place.
Right now, are the default on our team, everyone's 10x. So, we're not thinking about that anymore cuz it's it's the 10X is almost the new normal. What a cool time to be alive.
Welcome humans to the latest episode of the Neuron podcast. I am Corey Nolles, editor of The Neuron, and we're joined as always by our writer and fearless friend Grant Harvey. How are you today, my friend?
Great. Today we are diving into AI coding tools, specifically what happens when you give an AI agent full access to your operating system instead of just sitting inside of a text editor. Our guest today is Taylor Mullen, principal engineer at Google and creator of the Gemini CLI.
Before Google, Taylor worked on GitHub Copilot Visual Studio Code integration at Microsoft and his team now ships 100 to 150 features and bug fixes every week. But here's the wild part. They do it using Gemini CLI to build itself.
Taylor, welcome to the Neuron.
Thank you so much for having me. Yeah, it's uh I feel like when people hear that like, oh, using it to build itself, it's very inceptiony, right? [laughter] But yeah, I'm I'm so glad to be here though. This is like really exciting. What a cool time to be alive.
So, we're going to just do a little quick context at the top. So, we're going to dive into AI coding interfaces. And for folks who are not coders, please stick around because we're going to explain things uh explain all this stuff. Uh but then we'll also get into the weeds and uh for those of you who are coders, Taylor, we understand that this is your first in-depth interview about Gemini CLI and the origins story behind it. Is that correct?
That's totally correct. I have um I've had so many offline conversations and like customer conversations like behind closed doors, but yeah, I haven't really talked about it um publicly yet. So, this is kind of cool. uh very and I'm love to be here doing this on Neuron like I'm super stoked.
Let's let's kind of start there. How uh how did this all come together?
It's kind of crazy. So roughly I think it's almost two years ago now. Um I actually built an agentic terminal um at a hackathon and that's kind of the entry point. That's the starting point. People like wait what what is that like two years ago? This is 2026. Um, but it's funny like if if like if we take a second and like rewind ourselves, what two years ago was, this was the age of when um if you're building anything with Agentic AI, you were trying to use as few requests as possible, like maybe one or two in order to get a job done because it costs money every single time you spend it. You were trying to make things as fast as possible because in the age of Amazon, every like half a millisecond like resulted in a lot more of your user base just shedding. And so two and a half years or two years ago I built it and it worked really well but we scrapped it and we scrapped it because things were too slow. It took like 30 seconds or 30 seconds to a minute and a half to get an answer. It took like 30 requests to actually get an answer. Um and it was too expensive. It took too long. And then the last key thing is people didn't believe in a CLI at the time or a little bit too early. Um, and I went from hackathon to trying to bring it back into work, right?
Yeah.
And so it's so funny like thinking back it in that age um of oh my gosh, now look at us where this is the age of terminals, the age of CLIs to make it so you can bring agentic AI to everything on your computer. So that was kind of like the origin origin. But like fast forward to here at Google. Um those for those who don't know actually I just I think I just had my year anniversary yesterday.
Awesome. Congrats.
Thank you. That was so awesome. Yeah.
Yeah. I got I got a nice little email saying, "Hey, congratulations. You've been here a year." And I'm like, I've been here a year. That was the fastest year I've ever had in my entire life.
For real. [clears throat] Um, when you when you started at Google, did you already know you were going to build Gemini CLI or was that something that kind of came together after you landed?
Okay. Yeah, it's it's it's a good question. Like one of the one of my charters was think of what it means to like build the future of developer tooling. So, I started actually off at Microsoft building GitHub Copilot um for Visual Studio where I had built this foundation of what it meant to have more of an like more LLMs in the mix for having AI in your code editor, AI in your chat pane, AI in your errors and everything in between, right? Um and so I was very that that was my that was my prognosis there. But here it was like can you like what's next? And so with the onset of all these CLIs and um you look at cloud code like they have hit it out of the park like an amazing product.
Yeah. Um seeing some of these things just come to existence was really realized like oh my gosh I've done this before like I did this two years ago. It might be time to like take a tool like take a model like Gemini which is really good at a wide variety of tasks and bring that to developers. Um because for those who don't know like developers we write a lot of code.
Yeah.
But it doesn't take up a huge portion of our day, right? Like it's if like in like a big company setting you you spend far less than half your day writing code. A lot of it is just bottlenecked by human conversation and going back and forth.
Actually I what what what is the rest of the day? Is it conversations about what the code should be? Is it getting permissions? Like what what what do you do when you're not coding?
Yeah, it's it's a great question. It's alignment uh this is the biggest portion. So if you have imagine you have a whole bunch of people who can build things at the speed of light and that's kind of where we are today like everything is instantaneously buildable.
Well, if you have a person building a hotel and building different floors of the hotel at different times, what you go into one floor looks and feels one way and you look go to another floor looks and feels another way. It's the same thing with software is we like you want consistency, right? And like if you go to your Google Workspace like the docs and the calendar like everything that is offered has a level of consistency and the features have a level of consistency with them. And so that also goes into the coding world which is well what does it mean to build something that feels coherent? Because even as a user you don't want to relearn everything every step of the way, right?
Yeah.
And so it's like so it's it's kind of one of these like double-edged swords. you can build anything almost instantaneously now, but like you really need to make sure you're building the right things in the right way.
Yeah.
Um that's kind of a line. So that that's a big portion of it and like and that's handwaving a lot because a lot of this goes into well what should we do next? And to be frank, it's also somewhat new. Like it it this ability to build things so fast is relatively new for us. Like we've had like LM enabled coding for a while, but it's really leveled up in the past like year and a half or so, right? And the crazy thing about Gemini 3 Flash is that not only is it a faster model than Gemini 3, in a lot of ways it's a better coder, right? So now you have a faster, better coder that you're having to try and use to to like whenever you come up with an idea, a new idea to build, it's like I can build it insanely quickly.
It's so true. Like people don't realize it, but for Gemini with Gemini CLI, like the entire team, most of them anyways, prefer Gemini 3 Flash. Like that is like the thing and we use it where we're having like tens of tabs of just like things just chugging turning away in the background and it's so efficient and it's so fast and it's so smart and it's so flexible. Uh like even earlier this year um like I installed the Google work we have a Google workspace extension for Gemini CL life and it allows you to connect your terminal to your entire like workspace world. So think of documents, calendars, chats, yeah, everything in between. It's really cool.
And I was drowning in one-on- ones and I love my one-on ones. I love every like all like all the people I talk to. But I had too many and they were happening too frequently. Yeah.
So I went to the gym and I see a lot. Yo, can you help me clear my schedule starting in the new year, but for anyone you do, please DM them and let them know that they can always reschedule something. We can do things ad hoc.
Yeah.
And within like five minutes, it had emailed, messaged, cleared, and managed everything for me. And so, it's like these are just like some of the small little things of what it means to really use it for way more than just coding is because like it's it builds itself. We've already proven it's really really good at coding, right? And now we're trying to make it so it's well, of course, we're always trying to make it better at coding, but really make sure it can do the rest of the pie, the rest of what it means to actually be productive dayto-day.
For our newer listeners or people who are who are newer to this, less experienced, what exactly is the Gemini CLI? And maybe for people who've never even heard CLI before, what does that mean?
Maybe I could just share my screen and kind of show a little bit. Is that possible?
Yeah.
Okay. And I'll and I'll try and talk through this as well um as we kind of go here. Okay. So, I actually just had an instance here running. I'll go ahead and I'm going to just type Gemini. So, this is a console. It looks kind of intimidating, right? It's like just text, but if you boot it, um Gemini CLI kind of looks just like your average like your average prompt AI experience. You have a prompt box.
Very cool.
And you can do things Right. Um like I actually had this really interesting experiment where I had like it could say like hello and and it can go ahead and it can respond hello back to me. Right. And [clears throat] so like I had this experiment where I had it dig through all of my um like my emails, my calendars, my chats, my mails to create my own system prompt. Uh which is kind of a funny thing. So for people who don't know system prompt is like hey how do these things talk to you? So for instance, the way Gemini CLI talks to me is like hey I'm ready to assist you with your software engineering tasks or any other project needs. The way it responds if I said talk like a pirate it would tell me that in a piratey form. Well I had it create one say to talk like Taylor right so it talk like me and it could do those sorts of things. You asked what I said. This is a CLI though. It's truly uh it can do almost anything you can do because well it can probably do more than the average person, right? Because uh when I was a kid, I remember watching my dad pull up a terminal and do all these crazy commands in it like cd bash all this stuff. And I'm like what is what does that even mean? Why doesn't he just use the like menu interface like a normal person? But now with the command line terminal, you can tell the agent to do you know any kind of thing it would need a command to do and it'll write the commands for you. You don't even have to know the command.
So you don't even need to say necessarily like you know I need to know my directory. I need to know the folder hierarchy. You can say no get in my videos folder somewhere is a a video we recorded two weeks ago. Can you go grab that and email it to so and so or whatever.
Oh totally. So like okay think of it this way. So like it was just a prompt box. Pretty simple, right?
Yeah.
But what it what it can do is it's not constricted by the bounds of a typical application. So think of it this way. When you go to chat like chat GPT or Gemini or like cloud any of the AI models, you have to go ahead and copy and paste stuff into that world and you have to go ahead and like or they have to very intentionally build some level of connection to other data whether it's like for to connect it to your email. Well, when you have a CLI, like it's at the foundation of compute. Like CLI is it's actually a relative not it's an old technology. It's something that's been around for ages, but no one uses it because it's historically been hard to string together all the syntax. Yeah. Like, hey, how do you invoke these really arcane commands to accomplish really powerful like results? Well, we now have LLMs. They've been trained on the vastness of the internet. They know how to do all this and they know how to problem solve with all this and they know how to connect all this. So like imagine having this one product that literally can take all of your information no matter where it comes from and manage it in one place. Oh, like that that's like that's the value ad of of sealance why we're seeing this terminal like I just did a talk actually uh like a month or so ago the ter a terminal re renaissance honestly of it's really just able to do everything and when you're able to kind of give the right guardrails to AI.
Let's let's talk about the terminal renaissance because um you know obviously you worked on you you worked on Visual Studio Code at Microsoft and that's a different experience where that's called an IDE and and it's like a code editor almost like a do what would be a document editor for people who don't code but for code um yeah that's optimized for that. Um why do you think that the terminal is better? What can an OS level agent do that something like VS Code can't? And uh and I see I always see the battle. It's like, okay, we have a terminal agent that's really good and then now we're going to add a you know a UI to it. [laughter] It's like now we're we're just going back and forth. Where do you land on that?
It's such a it's it's a great question and I think like it kind of boils down to a few different factors. The first one is so developers utilize a wide variety of products. They don't just use an ID and they don't just use a terminal. um they'll use an ID for a lot of like a lot of variety of things from everything from committing code to writing code to interacting with AI to do other things. They'll use a CLI to do the exact same thing, but it's really the form factor that which like it's the form factor that feels familiar to them is the one that we want to make available. So like for Google, one of our goals is to make it so no matter what you prefer, no matter where you are, there's an option and a really powerful option at that. Um, the second one is that when you think of like hey what is the supported environments of like applications this is not just ids like any application you install like oh is is it supported on Mac is it supported on windows like that's usually a question you have to answer right terminals are super lightweight the ability to expand that matrix of what do you support is so much broader oh wow so like no almost like pretty no matter where where you are what you have there is an option in order to get into terminal environment And for a software developer, this doesn't always mean like I'm sitting in my keyboard and I'm typing. Um, this could also mean like there's things called a continuous integration like a CI/CD pipelines that we call them. And this is when you write code, you have to make sure the code is it works. Yeah.
And one of the ways you make sure it works is we give it to this like other system and this other system churns through it to build it to make sure it can compile. Well, we also integrate Gemini CLI at that layer. So we can have automated reviews. We can automated detection. We've automated changes to make sure that more things are happening within the right guardrails um on our behalf. That's a big that's a big deal with AI because of the way that I understand it is if you can't evaluate and and verify that something is true. It's really hard to train on it and and get the AI to be really good at it. Which is why they're really good at coding is because you can verify the code either runs or it doesn't.
Yeah. Yeah, I love that.
Yeah, it's it's it's a really good call out there like that the verification step. Um, we even find that's like there's a thing called um test driven development TDD is a thing where for developers what we'll typically do is we will write code with a certain goal in mind like we want to create a feature whether this is like a button to send an email.
Yeah.
Well, we want to make sure that when we click that button it sends an email. So we write something called a test and the test means verifies. is if I click the button, it sent the email. We do this so that when we write the next feature, we don't accidentally break the ability to send tests because if we accidentally break that, that's a problem, right? Turns out AI is really, really good at well, what if you were just to write the test to say, did it send the email and then write the feature? Like, what if you changed it? And this is really impactful with AI. Why? Because it enables the system to look at this like this problem to say, hey, it can't send an email. how do I make it send an email and it can iterate with that kind of in mind versus the backwards of like just trying to figure something out and then after the fact trying to write something that validates the experience.
Oh, that's fascinating. So, I guess can you give us like kind of a an example of what say I'm a developer and I I wake up in the morning, I've got a bug report. What does solving that look like with Gemini CLI versus the old way?
each developer is going to respond a little bit differently because of the surfaces that they interact with. Um, but I can give you a little bit insight. I can say, hey, how I actually interact with that. And so, one of my normal flows is very much I will open Gemini CLI and we're actually open source on GitHub. Actually, let me just sh I'll share my screen again. I'll show a little bit here. github.com, Gemini CLI, like we have uh we're literally building in the open every single day here. And when a bug happens or a feature request or enhancement um anything in between um actually are you able to see this picture move this off to the side um when we actually see this we end up getting these reports. So one of the flows that I do is I will actually just hand this issue over to Gemini seal and I say hey can you fix this? I'll give it I'll go ahead and take something and I'll go here and I will open up like CD GitHub get my CLI I'll boot this and I will ask um can you fix this and I'll hit enter right and then it's going to go
the URL right
then I threw it at the URL it's pulling this issue right here so it's using this whole command line to pull this URL it's going to figure out what's actually happening I actually haven't even looked at this bug see it's even solvable. [laughter]
But this is one of the flows how we roll.
Yeah. [laughter] But this is this is flow one of what we do here, right?
Yeah.
And so like the second flow that we do is I'll be having a chat like some level of um like just chat on the side with people on the team because in a meeting like oh people are reporting that something is broken.
Yeah.
Right. Think of how many see think of how many products you've used where something doesn't work as expected. Well, we talk about it as well. We want to fix it just as much as it sucks for that it's broken, we also want to fix it.
Um, so the other flow is we'll very much say, "Hey, can you just pull my chat with soand so and fix it [laughter] and it'll just kick off and we do and we do this all the time where um some folks what they'll end up doing is they they've taught their CLIs to spawn other CLIs." And so you'll they'll have like one me like kind of orchestrator which then spawns other ones and then so it will manage the life cycles of those and then before you know it you have all these CLI spawning CLIs and all these CLI kind of iteratively working through the problem. Um, but granted like we do this with the notion of um with great power comes great responsibility.
Yeah. Right.
The CLS they can do anything. And so we have guard rules in place. So when one of these actions occur that could be impacting in some way, it waits for the user to respond. So like you'll be told, hey, like you need to go ahead and approve this action or not approve this action.
And so we try and we we really try and play this game of spin up a lot of parallel activities as many as like you mentally can fathom.
Yeah.
And then wait for the notifications to start coming in saying, "Okay, this one needs some sort of human intervention. This one needs some sort of human intervention."
I have two questions. Number one, what are your best practices for conducting these like swarms of essentially sub sub aents or sublis? Uh and then number two, are you actually working more than you would uh because you're doing this for 100 to 150, you know, features a week now? Like how Yeah.
Oh my gosh, that's that's a good question. Okay. So my hot take on this is that I think that we are in this mental growing period as like in tech in developer industry of how many things can you keep going mentally at the same time right because when you're because for me like I will roughly have like I think it's around seven to 10 separate things going simultaneously.
I would have said seven for myself as well.
Yeah. But it's like they're just chugging, right? they're chugging and they're going through it and then eventually they stop. The question is when you go back to the thing that stopped, you have to kind of context switch back into that state of okay, what was I doing? What do I care about? Right? Because one of the big things that we like take very like strongly is we want to make sure we have human eyes on every single change that goes in because there things need to be blessed. It's very easy for AI to go off the rails and we very much don't want to be blindly making changes and kind of shipping it to the world.
Yeah.
So, um, so yeah. So, first off, like how we do this is we actually have things called policy files and these policies allow the agents to do certain actions without approval. Um, and we really guard these. So, like there's certain commands that can be run that we know you don't have to like you don't have to ask us like go ahead and just keep running. And the more of that we build out our policy files, the more we're able to go ahead and allow them to run for longer periods of time. Second is continuous like um improvement and iteration. So as these agents get better and hit faulty points like every every session that you have with a new agent is going to hit some sort of pain point.
Yeah.
The real question is what do you do from there? You could choose just to try another prompt which is one solution. Two, you could choose to is this a reasonable break? Is this a reasonable thing of why it made the mistake it made? And then how can I make it so it doesn't make that in the future? And to make it so it doesn't make it in the future, we have a thing called Gemini MD files. Um, and these are pieces of context that we continually grow with certain rule sets and guidelines to make so future runs for the entire team don't end up uh in a poor state. And these grow over time because they're very codebase specific usually. Um, and then lastly, uh one of the pieces that we we like to do is try to break things down and plan things ahead of time. So very much a flow that we do is we will plan something and then we'll implement something like we actually just released this thing called called conductor uh which is another it's a Gemini CLI extension um and it's a I actually like to call it planning plus where [laughter] it it's kind of silly, right? Yeah,
but the gist [clears throat] like and when you think of uh did you I'm assuming do you know do you know don't know what planning is in spec development by chance?
I yeah conceptually I I know what it is and I've also you know in other CLI tools I've used like there's like a plan mode where basically you come up with the whole spec and everything.
Yeah.
Hey well so we just went ahead and released kind of a an extension to experiment what would it mean if you dialed that plan mode to 11 like truly. And so actually I'll share my screen here and we'll go ahead and give a little bit of curious what that looks like.
All right. So this is the conductor. So conductor it's context driven development. I actually have behind the scenes I pulled up a thing here. Um this tab I'm scroll all the way at top here. I went this is me actually pulling it from a chat thread. I wanted it to build this app for me and I pulled it from a chat thread. It's kind of funny. Uh and then so I then said okay can you go ahead and um can you go ahead and start creating this feature for me? And this conductor went ahead and said, "Hey, welcome to conductor. I'm going to help you set up your repository." And it creates all the scaffolding for it to self-learn and iterate on itself and eventually starts asking questions. So think of planning where like you historically would say, "Hey, go and figure out the whole plan and suggest something to me."
Yeah,
this does that. But then it also it for any clarifying point it tries to give you detailed descriptions and questions for how you can make the right choices to make sure plans now and for the future go in the right direction. So it's like for instance it's like who are the primary audiences of this ex I think this one is actually having I was building a tool to make it so it would uh curate my emails and my chat messages so I didn't have to respond to everything. This is this is the purpose.
God I love that. [laughter]
Yeah. It's like I think it it called it an executive assistant, which is kind of hilarious.
Yeah. By the way, if you've successfully completed this, you should ship it. We'll we'll use it.
Yeah. Yeah. All right. I'm in.
Um I
What was that?
Oh, I said I'm in. [laughter]
Yeah. This is like it's can you clone yourself and can you make it so like your day-to-day is faster. This is like it's the I was trying to go for the holy grail. I got a little bit of it here which is pretty sweet. Um and so it's like who are the primary audiences and it asks like hey is it this executives engineers g you know power users and asks all these questions and eventually it comes down to this product guy like for a product what I'm trying to build it then goes a little bit deeper it's like okay like what are like like what are some of the details of how this should behave and it asks me more questions and then it has like even a deeper level guideline for it then it builds out and it keeps doing this where iterates with questions specification questions, specifications where finally it gets done. This is like this one I could finish the project set up where it had all these details where every single plan I would do from here on grounded in this information at the end of the plans it would self-improve itself to grow its if its knowledge base. So the plans would get better over time and uh this is just one example like so conductors is anyhow it's a really really cool thing that we built uh to make it so like you can take planning to you're not just like having it go through and give you a kind of an uninformed plan. It it will work with you to build the best possible plan and it will work with you to build something that can then evolve over time. Uh that not is that's checkinable which means it doesn't just help you it helps everybody on your team.
Wow. um improve the AI's intelligence.
Will it then help you with like not just the overall plan but like a build plan implementation plan as well along the way like okay we're going to start with these features we're going to work from here and
Oh totally. It's it's it's if I'm able to like behind the scenes I'll see if I can pull up um the exact I ended up like using that one to implement something and I'll see if I can pull it up real quickly behind the scenes. But yeah, it's is it's exactly that.
It went ahead and like it would go ahead and it' build a direct implementation plan of doing all the details in between which is pretty cool. Like it's just like very
And then you're and then you're just like go.
Exactly.
Yeah.
Um like here actually I just I just pulled one up um for so for instance like if I am to look at this this is the end of a plan here, right? And so it actually created, hey, build a sync and triage. This is actually for the same agent, a sync and triage support. So it could sync my emails and my chats and it could triage them. That was the whole purpose of this feature. It started tracking it here. And this is the very end of it where we move saying, "Okay, that's done. This is all self-improvement because it went from using Google APIs like Gmail to actually using Gemini CLI's workspace extension to do this." Like it's actually using Gemini CLI behind the scenes. So this is it's self-improving. Let's see if I scroll up a little bit here, I can show some of like where it actually has the implementation plan.
So yeah, so here task implement feature action execute handler and so it's literally checking off features bit by bit to go through here and writing code on my behalf and checking things off as it goes. This is so this piece here is it's actually trying to build it's trying to make sure the code is right. Nice. Um, and it's writing stuff and it's doing things like it's like it's saying, "Oh, this is in progress." And then it checks it off once it's done. And it keeps going back and forth. And one of the really cool things is it makes sure like when it validates that things are working as expected that those validations are actually at a level that is um coherent. So there's a thing called code coverage which we have and code coverage is something that in software engineering says have you validated every piece of code you've written to do the right thing and it says hey it must be 80% or higher and if it's not 80% or higher it goes back and it writes some more validation and it has like this really nice reinforcement loop um to make sure that it's building exactly what you want and doing it in the best way possible. It's using the policy document and it's doing all of this without asking you for permission, right? Because it knows what it can and can't do without.
So when it's when it does something that I haven't given so it doesn't it doesn't know what it can and can't do. But when it ends up doing something that it can't do, then it asks me if it's allowed to do that is the intention.
Cool. So we kind of because basically every single action that it takes we kind of run through this pipeline that says hey is this allowed or is this not and the is this allowed or is this not is tied to whatever this policy is that you've written.
Okay.
And the policy might say for those who are familiar there's a command called ls and the intentions it's like saying imagine like saying open your folder like a folder on your computer like what files are in that's effectively what it does. Well, I don't want it to ask me just to look at my folder I'm okay like feel free look at stuff. So, I've allowed that, right? And if I don't allow that, then it would ask me for permission like, hey, Gemini CLI is trying to look at my folder. Do you want to let it? And so, this is very optin basis.
Now, how about something go back and be like, it can do this now and and add a feature in there. That's okay. Uh or or occasionally even maybe be like, you know what, I need to be looking over this that I've given it. Either direction. Is that a thing that happens regularly?
Oh, totally like all the like so I actually have several p I have sever several policy files specifically for that intention um with different so I actually have one for the Google Workspace extension for Gemini CLI which allow lists all of the readonly operations which means feel free pull my calendar feel free to get my email free feel free to read my chats don't send anything don't edit stuff like I make it very clear what I'm allowing it to do Right. Um, and so I will use that when I use the workspace extension in that flow, but I will not use it in other times. So yeah, I'm very intentional for like the types of workflows that I have.
And is as far as the policy checks go, is that something that you built into the harness itself like it's like a hard check or is that something that the LLM is like having to remember and recall to do itself?
Great question. A hard check. So very much we take the perspective of the LLMs like the LMS will always do bad is the kind of the perspective we take and it's not true like it's it's super not true in the sense of like most of the time it is the job of the companies producing the models to make sure that they are mostly good. What's the phrase we heard recently Corey? Toddlers with machine guns that's LLM. Yeah, but it's so like it's it's one of those things where it's like we try and keep it so that like we're always assuming worst case scenario so that we can build for the best case scenario. Like as an example of this, right, we're open source like as we were showing before.
Um, the reason why we're open source is because with a tool as powerful as a CLI, we want to make sure we're building in the open because how else do you trust it? How else do you trust it that it's doing exactly what it's doing? Right? Like we're like I think I forget the numbers but like last I looked like we are the most popular GitHub open source CLI on the market. The reason why it's important to us is so we have enough eyes on what we're doing to keep us honest at all points, right? Because for us we like everyone makes mistakes.
Yeah.
But at the very least we don't want to ship mistakes ever, right? And so we build in the open. We make sure everything is there laid out on the table so that like when a big company decides to use us, they know it's tested, it's trial, it is it's gone through every possible sort of restriction it possibly could. And it's not just the people at Google. It's like is the million plus users using this and working with us to build it every single day. It's the reason why we can do 100 to 150 features and changes every single week is because we have a massive open source community that's helping us build this out.
That's so awesome. Yeah, as far as my personal experience, I have found that Gemini has the best like the longest context window where it can stay lucid or coherent. I don't know what the right word is, but but it can it can pay attention to details for the longest amount of any of the other AI models that I've tested. Um, so what's the limit for what uh how how big your plan file could get or how big your code base could get here? Is it a million tokens? Is it two million? Like where does it where does it lose track?
there's always like a where does it lose track and then where like and how much do you allow. So first and foremost was like um one of the things that was very like I felt very strongly about is we're Google we're building Gemini CLI we darn well be able to use a whole all one million of those tokens if you want to. Yeah. Right. And so like this is super important because a lot of products will just restrict the boundary because it's either more expensive or for varying reasons.
Right. Um, we allow users to now to restrict it further if they want to like but we all like we'll never restrict you not to be able to do it if that makes sense like you always have the a million token context sort of open to you um and so um where Gemini like falls over from our experiments is it's a little it's it's not actually a clear-cut answer. It's a little different for every single scenario. Um, so for some coding tasks it could fall over super quickly. For others, it could go forever and you'll never see a difference at all. Yeah.
Right. For if you give it a like several books of information because you can easily give it several books of information. Um, a lot of the times it can be coherent. I think like my personal hot take here is I think in industry a lot of people look at the context window and see these artificial limits um and like oh it falls over. It gets not lucid after a certain amount of tokens. when in reality there's been so much back and forth in the conversation. Like if you were to give this conversation to a human and say, "Okay, like what guidelines do you want to follow? Would the human even be successful?" Um, and a lot of the times the answer is no.
No.
Right.
Yeah. Yeah. Exactly.
Cory just made a whole video about this of like the you got to be really strict of what with your prompts and everything.
Yeah. Context context engineering, right?
Yeah. Yeah. I was kind of getting into the idea that, you know, there's this this idea that every every prompt you're using now needs to be like 3,000 words. And I'm like, you're throwing in redundancies. You're throwing in conflicting comments. You have all of these things in there, and it's spending all of its time trying to figure out what in the hell you're asking it to do instead of actually doing your thing.
Yeah,
totally. And and you see this too. So, it's like now a lot of the models will show their thoughts, right, of how they're thinking through what you've asked. And you'll see the thoughts go through every every little iteration of those 3,000 words. But
Yeah. [clears throat] Yeah. So, like I think it's not a clear-cut answer for where it falls over. I think everyone has a different limit for like their own workflows um for what makes sense for them. I vary amongst all them depending on what I do, right? I go everything from using a full million to going down to say like 400,000 um is I think the uh no I think I've had a few have gone down to 300,000 um but it really really depends um on the workflow itself. I think one of the really cool things is that like Gemini is so good at it where we have teams internally at Google who like who are using this and give huge swasts of data for it to pick through and it's able to do like Gemini CLI is able to do just instantaneous work for something that historically would be a week's worth of work.
Yeah.
And it's just it and people get a little scared by that sometimes. I don't know why they would cuz what this means is if your favorite site is down, this person is now able to bring it back up just that much quicker.
Or your agent can recognize and do it for you. That would be the ideal next phase, I think, is is is coming in in the morning and finding out your site was down. [laughter] Basically having down detector bot, which basically is always checking if it's down and then if down fix it. [laughter]
Yeah, it's it's funny. So internal to Google we have a thing like incident reporting which is because of course we have a a huge number of really really popular products like billion user products right and so we have a very robust incident reporting system out there and we have Gemini CLI hooked into like all of it and this is about saying okay when there's an error why what and who should be contacted yeah right and it's able to pull in all this information um to make that possible
we we read this really interesting uh paper recently I think it was from last October called RLM. And essentially the idea there is that instead of putting, you know, all of the tokens of context in, you know, running it through the model, you basically make the the context itself an environment that the agent can like go in and edit and and manipulate with like Python snippets and stuff.
Yeah.
Yeah. Yeah. And that makes me wonder, you know, if you're thinking about that certain tricks like that and and when you're developing Gemini CLI, if you're thinking about compaction or things that, you know, we've seen mentioned elsewhere, like if if any of that's on your road map or if you're already doing that and and what your thoughts are in that kind of
Absolutely. I I think I think [clears throat] that it's definitely a breadand butter um like thing that we're all bread and butter is me downplaying it too much. It is an important and vital thing that every product has to do, right? Right. And so for Gemini CLI like I actually went through um several iterations like early on when I was building it of how like how do you manage that much context
And what works most effectively. So in software development, there's a thing called embeddings, and um, there's an embedding-based approach, which is you take all the content that you care about and you index it effectively. You make it so that you can quickly look stuff up. And at the time, um, we had an approach where every time a user would ask a question, we'd try to say, "Okay, well, in our big index, what information should be pulled in?" Right? And we'd try to pull it in. It was pretty good. It, it was actually not bad at all. The problem was, when it was wrong, it was really, really wrong. It would like, if it pulled, if it pulled in a partial piece of information, the model would just dig in super heavily to that one. And it makes sense. It's like, if you're asking an authoritative figure an answer and they say something, you're going to give it more weight than a non-authoritative figure. Right?
Yeah. Right. And so that was a path which we said, "Okay, well, is there better?" And our limits. It's very similar to agentic search, which is what we have today, and this is where we landed on. So the intention is giving the model the ability to open files, read folders, control F like search, search through a document, search through text. Um, and it can reason about its own methodology for how it can get to a solution. Um, so as an example, when I, you said earlier, "How do you solve issues?" and I just pasted in the GitHub issue.
Yeah. Well, how it does that is it tries to look at what the person reported. It then tries to look at the codebase and it starts doing searches for to find the right files. It starts opening files to look at them in depth, and as it looks at one file, then realizes it needs to look at three other ones. So it goes, looks, and it keeps going until it finds some more reasoning. And so it's kind of relying on this innate piece of, um, the model to, to do better at understanding.
Yeah. Which is, it's kind of an interesting path here because like, as the models evolve, we're going to hit a point where there's like, they're basically perfect at deriving this information. We will hit a point where that's the case, but then the question is like, "Well, what's left?" And it's like, "How do you evolve products from there?" And that's an interesting thing, which for us is really important.
How do you, yeah, what's your, what's your, do you have a thesis on that, or is it, you just figure it out as you go? So, uh, so Gemini CLI, like one of our pillars is extensibility. Um, we released Gemini CLI, when did we release Gemini CLI extensions? This is a while ago. Um, we have a thing called extensions, and extensions, think of them as they could package your model context, protocol servers, MCP. It can proc, it can package your commands. It can package all the customization for your experience into one thing, which means you can say, "Gemini extensions install vision," and maybe you just give, you've now given your agent the ability to use your webcam to generate images, to generate video, and everything in between. Like, that's just an example.
Your webcam? That's terrifying. [laughter] Yeah. Well, one of the demos I will typically do is I'll, I'll say, "Hey, take a picture of me and give me long hair."
Oh, I love it. I have these, and I, and I get these wonderful walks that come out the other end. The number one thing that I use, uh, you know, command line tools to do is, uh, play with, uh, connectors. So, for example, you know, in GDAU, which is an open-source video game editor, uh, I'm having it write game code for me. But the problem with the current version that I have is that, well, it needs to actually be able to see the game in order to, you know, see if what it's doing is actually working. Like, is it actually creating something that works? So that's a big, you need to be able to do that.
Like, take a screenshot, record your screen, just like how I've been sharing my screen, right? And so, so like, for us, extensibility is that mechanism. It's like, how can you extendation to be customizable?
Because, and it's kind of funny, actually. We were, we were, um, definitely the first to market to have this. It's so sad that like, when we, we released it June 25th, like literally the day we released Gemini CLI, we actually already had extensibility by like extensions baked in. We then talked about it, uh, like a few months later, and, and like the day after, I think Claude Co released their own thing, which I'm like, "Oh, come on now." Oh, which is so funny. [clears throat]
Um. That's how it is now, though. It is. Everyone's so bad. Yes. But like, this, the idea of extensibility though, it plays into this. This is why I talk about it because in a world where the model is almost perfect at like, kind of doing the right mental model, it's how do you make it work for you? How do you, how does it tailor itself to your workflow? Like in our industries, we have lawyers, we have developers, we have teachers, and there's very specific sorts of ways you talk depending on what you're doing.
Yeah. Right. It's very like, so for an agent, the same thing applies. How does, how does it work better for each of those personas? How does it work better? And then there's even subsections of those personas where you really want to amplify it, which is, I'm an enterprise, and I tie internal to Google. Our entire code stack internally is custom, like everything is custom about it. There is not models that are trained on looking at that. But we have a huge user base of Gemini CLI users internally, why? Because we have made it extendable so that it could be extended to work for all of those customized experiences. Um, so yeah, so extensibility is like definitely our solution to this problem and where we think the world is going. Do you find that Flash can handle a significant amount of the tasks you're throwing at? What, what determines when you go to full three? Great question. So, honestly, right, so Flash, like always, it is like, it is literally my default everywhere. It is so freaking good. Uh, first off, um, but when I go to 3 Pro is in the moments when it's like, "Sorry, I'm stuck."
Yeah. So, I'll actually swap to Pro. Okay, like, your other sibling is stuck. You got to go ahead and work through this, uh, problem domain. Think outside of the box. Consider other alternatives. And then I'll, that's when I'll typically go fall back into three. But to be honest, um, over the past, I want to say month or so, I maybe have fallen back to Pro 10 or less times.
Nice. Do, do you tend to lean on it more for planning type type things? Three is great at planning. Three Flash is great at planning. I think like, to be frank though, it, it really, um, like people are gonna hate me for saying that, but it's true. It's, it's super, it's really how you prompt it.
Um, I'd say with a lot of these tools, it's a tool. How you use it really, really matters. If you don't, if you have not gotten your prompting ability up this enough yet, there's room for improvement. Like, you will see exponential gains of your own ability. Like, I think you mentioned at the beginning of a podcast, "Okay, what does it mean to be a 10x versus 100x engineer?"
Yeah. Right now, they're the default on our team. Everyone's 10x. Yeah. From what they were before, easily 10x. Yeah. So we're not thinking about that anymore because it's, it's the 10x is almost the new normal. Now you're 10xing the 10x. Yeah. Exactly. What's the difference between a 10x and a 100x then in your, in your? 90. It's, it is incredibly hard. And 90. Yes, it is 90 on VX's. No, honestly, the biggest difference is parallelism. Like, you're no longer, you're, you are no longer working with like a few things going like at a time. You are now mentally shifting yourself over and over again. And you're leaning on the agent's ability to cross-check itself, to optimize how much time you yourself spend on each problem, and you're willing to spend a little bit more to get there.
Yeah. So, there's a, a technique called, have you heard of the Ralph Wiggum technique? That's kind of going. I wanted to ask you about that. Yeah. Yeah. It's going viral right now. Tell. Please tell us. Yeah. It's super effective. Uh, first off, it's great. And then we have a person that made a Pickle Rick, uh, version of that, which is hilarious.
That's [laughter] good. I'm pro Pickle Rick. I like it. Would, would you explain real quick to, to folks who don't know what that is? Um, okay. So Ralph Wiggum is a, I believe, a Simpsons character. Um, and the intention is, is the way they talked in the show is like, very repetitive, like over and over again. And so how that kind of translates into software development is, imagine if you were to give your AI of choice a problem, and it gave you an answer. And then what if you took the answer and your original problem statement, you shoved it back to the AI, said, "Here's where I currently am. Here's what I just asked you. Do it again." You say, "Do it again. Do it again. Do it again. Do it again." You keep putting the same prompt back in the flow, right? Even though the state and the state itself is, it is slowly evolves. So your, your pro, like the output on the very first iteration looks a certain way, and you refeed the prompt. Then you get another version of the output that's maybe a little bit better. You refeed the prompt. You get another one, and it keeps going. Some exit criteria.
Yeah. That exit criteria is always the fun one. Honestly, I, that's what I always ask people. I'm like, "Okay, well, what is your exit criteria?" For me, I run it five times. That's my exit criteria. That's, I just, it's a, it's a numeric one. Um, for me. So, by exit criteria, you're talking about is the, the test that it has to do to make sure that it works, or exit criteria is just like, "I'm going to, I know I'm fairly confident it'll be done if I do it five times." Yeah. And when I say exit, exit is too strong of a word. It's when I'm willing to put my eyes and actually look at it is the intention. It's like. Got it. When they bring you in. Yeah. Yeah. Um. I like it. Yeah. Yeah. [clears throat] And so, but it's super powerful because these LLMs, they'll make subtle mistakes, but if they iterate over and over again, they can refine their own work pretty effectively.
Yeah. And so that, like, so that's, so that's the Ralph Wiggum technique is one of those examples of you're just letting it turn to improve its own answer. So is one of these things which I was saying is that's a way how you can let a thread run for a lot longer of a time and ideally get a better output at the end. It doesn't always work that way. Like there are times when the models will go off the rails and like they'll go into totally unrelated territory. It's a very fickle balance of do you want pure autonomy or do you want to keep the like the driver in the loop? We in Gemini CLI, we try and strike a good balance between the two so it's customizable so you can kind of do both.
Um, but yeah, always a hard problem. Yeah, because that's very much a variable thing. It's even that way when you're talking about using AI for writing. Like, uh, Grant and I have talked about this before about how we're kind of at a point where all of the tools are pretty good, like, like most of them aren't bad now, you know, and, uh, you wind up getting into this situation where there's a certain amount of preference and, and you sort of learn like, "Okay, I can trust this, I can trust this, I am never going to trust it doing that." I, uh, you know, so it makes me wonder when, when we're talking about like that Wiggum technique, is it doing it in a new chat every time? Is it doing it in a new, or is it refeeding it into the same window? Everyone has done a little. Like, I, I think there for us, we feed it into a new chat and new context every single time. It's the most, it's a, arguably I'd say it's the most pure version. There are other clients that don't do that. Um, like there other ones that maintain the current thread and they just let the history kind of carry forward. I think even Cloud Code does that. Um, but the intention is for it to be like a truly true mindset is what the, um, algorithm/pattern is supposed to be. So you have a fresh set of context, but the underlying state, like the, the files change, those stay the same. Command line interfaces are having a moment right now for people who are not coders discovering them and being like, "Oh my gosh, there's an OSA, it's a computer agent, it can, it can do everything I I want it to do." How would you recommend people who are not engineers and not coders use Gemini CLI, and Google has a lot of other tools, so maybe you could also kind of contextualize the other tools that Google has, if, and your thoughts on them, and maybe are there any resources that would help keep them from being terrified? We have tried. It's kind of funny. Google is a place of, we [clears throat] let a lot of blo, like, "What is a thousand blossoms bloom?" I think is our, is our motto, if I recall correctly.
I like that. Um, and I couldn't advocate for that enough, which is like, we have a lot of availability for products that let people utilize different form factors in different ways to get what works for them. And so when I'm thinking of what, like, works best for people, we're all trying to do the same thing. We're trying to make sure all of these can address a certain surface area the best it possibly can and give the certain sort of experience and certain sort of getting started and certain sort of flow that works for you. So, if you're getting started, find your product or your surface of choice and go to those getting started pages. This could be a simple download, and it could be an end product experience. Yeah, this could be walking through a website on the side of what you're doing, um, along the way. Uh, or use your AI of choice to do it for you.
Oo. That's another one I've been doing. So, a lot of things that I've been pointing people towards is, I [clears throat] actually encourage getting into the Gemini CLI ecosystem, because it unlocks a lot of other stuff, which is, again, it's a terminal. It can do anything.
Yeah. If you're trying to get started on a class project, if you're trying to install something, ask it to go ahead and figure it out. If you really don't know, if you're stuck, ask it. And if you run into problems, it'll troubleshoot it for you. If you're trying to use a new like code repository that's open, that's out in the open, have it pull it down and figure out how to build it. You might need to install stuff. You might need to troubleshoot things. It can figure it out for you. If you want a getting started that's truly tailored to you, ask it to create you one. Customized tools and being able to actually build things that is personal, um, is what these tools now enable. And so we're trying to tackle this, like I said, a lot of different ways. Um, but AI, I think is the AI, I think is the answer to this ultimately. And like, at the end of the day, like we're trying to make so all these are affordable, like efficient, like your subscription of choice should correlate to what makes sense for you. Like if you use it a little bit, or if you use it a lot, free tier could probably get you doing all this.
That's awesome. Can, can you use, this might be a controversial question, can you use like other models, like let's say an open source model inside Gemini CLI? So because we're open source though, um, we have seen a number of people fork our, which is intended. We actually love fork. We, we like honestly, being in the open is so rewarding. Our community is absolutely so amazing. Love them to death. Um, they, they forked it. They've made other ones that run Anthropic models and OpenAI models. Um, but we like, we try and make ourself very configurable, even down to the model lens, um, within reason. So within reason is like, light LLM is one example. So this is a product which you can route any AI basically to, and it will fan out to the right model behind the scenes of what you care about. So you're able to do these things indirectly through other products, and so we make sure that there's enough knobs and dials that doesn't jeopardize our ability to ship the product, but also enables other, um, other solutions like that. So in Gemini, we do stick by default with just the Gemini models, but there's a number of them you can pick from. Yeah. To get there.
Taylor, thank you so much. What is the best way someone can go try the Gemini CLI today? So, go to geminli.com, and there is install instructions right there for you, is was how I'd say it.
Very cool. Actually, let's, let's, uh, dig a little deeper there because I was actually installing it earlier today on a new laptop, and, uh, I wanted to do it with Docker. Um, and yeah, because I was like, you know, there's been some weird npm package situations lately, you know, with security, so I don't want to be installing a bunch of stuff. Um, so I was able to actually do it with Docker. Um, but then once you get to the point where you have to choose, you know, the API console and, and, uh, you know, the other stuff, it gets a little confusing. Could you maybe give, tell folks just what to do when they get to that part? Yeah, totally. So if you're doing it inside of a sandbox itself, I would first recommend configuring it outside of the sandbox to start. So you have like your API key set up. So it's like, "Hey, here is me." Because if you're doing it in a sandbox, every single time that you spin up said sandbox, it's going to have to save it and manage on that device. Good to know.
So that's, so that would be a recommendation for how to start there. Um, but yeah, sandboxing, that's actually a great call out. We have multifaceted sandboxing. We support Seatbelt, we support Docker, we support Podman, and we also auto-configure Docker and auto-configure Podman, uh, if you have it installed for you, uh, which I think it's definitely one of our unique value props. We've tried to make it so the environment itself is really curated.
Yeah, it was super easy. That was easy to log into, and it's a good call out to, uh, to set, set your, uh, API stuff up before you go in. Yeah. Well, Taylor, thank you so much for joining us today, man. It's been a blast. Yeah, I've had so much fun. This is so great. Oh, thank you for having me. That's good. We've learned a, I've, I've learned a lot, at least. I don't know if Grant has, but I sure have. Yeah. No, definitely. This is great. And I have a much better understanding of what I should now be doing with a CLI and, uh, really looking forward to going and diving into that. Well, to anyone who's watching, if you haven't yet, please take a moment to like and subscribe and stop by the neuron.ai and sign up for the newsletter. Join 600 and some odd thousand others who read it every morning. Otherwise, we really appreciate you. Thanks to Google and Gemini for sending Taylor over to chat with us today. And on that note, farewell for now, humans. We'll see you next time. [music] [music]