📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Code with Claude Opening Keynote

Anthropic1:40:21

Transcription

Hey hey hey. [Music] Welcome. [Music] [Music] [Music] [Music] Hey hey hey. [Music] [Music] [Music] [Music] [Music] Hey hey hey. [Music] [Music] [Music] Hey hey hey. [Music] [Music] Heat heat. [Music] Happy. [Music] Birthday. Heat heat. [Music] [Music] Doo doo doo doo doo doo doo doo doo doo doo. [Music] Let me do it. [Music] Hey come on. [Music] Heat. [Music] Heat heat heat. [Music] [Music] Happy. [Music] [Music] [Music] [Music] [Music] Hey hey hey. [Music] [Music] Oh hey hey you are. [Music] [Music] I'm. [Music] [Music] [Music] Hey over beh. [Music] What. [Music] Hey hey hey one one 1. [Music] [Music] [Music] 1.11. [Music] Here come. [Music] One one one one one one one. [Music] Hey hey hey. [Music] [Music] [Music] [Music] Hey hey hey. [Music] [Music] [Music] I. [Music] Love me. [Music] Come on hey. [Music] [Music] Hey hey hey. [Music] [Music] Hey hey hey hey hey hey hey hey hey hey hey hey. [Music] Party don't get. Get yourself. [Music] [Music] [Music] Hey down. Down down down down down down. [Music] Come on come on. [Music] [Music] [Music] Down down down down down down. Gabbit jump. [Music] [Music] Dick dick. [Music] D. [Music] Down. [Music] D. Get. [Music] Down. [Music] Hey hey hey. [Music] Down hey. [Music] Black yeah. [Music] [Music] [Music] Yeah yeah. [Music] Yeah yeah. Heat heat. [Music] [Music] Okay okay. [Music] Please welcome to the stage chief product officer of Anthropic Mike. [Applause] [Music] Kger. Good morning everyone and welcome to Code with Claude, Anthropic's first developer conference. I'm really happy to see you all here. I'm Mike Kger. I am chief product officer here at Anthropic. I just hit my one-year mark, which in AI years is about like three years. Um, but I'm having a blast. Um, and before this, I co-founded Instagram, um, and also an AI-powered news app called Artifact, which is where I first started getting exposed to a lot of these AI technologies. I joined Anthropic because of its founders' vision, building AI systems that are powerful as well as helpful and trustworthy. Today, that vision includes something immediate and concrete, a commitment to empower developers like yourselves to transform how work gets done and how companies get built. This transformation is about augmenting, not replacing, human creativity. AI agents are changing the way we work and the way we innovate. They're expanding what we can build by removing bottlenecks that have limited human productivity. Today, you'll hear from our product and engineering leaders, as well as some of our customers, about how they're pushing the frontier to give you a sense of what you can expect. Today at Code with Claude, you can attend three technical deep dives to transform how you build with Claude and five sessions from leading players already using Anthropic's platform to reshape their industries, and dedicated office hours and workshops for hands-on experience. But before we talk about some exciting new API API capabilities I have for you, I want to invite a guest on stage. Please welcome our CEO and co-founder, Daario Amade.

Hey everyone. Uh, I'm going to be back in 20 minutes for a fireside, so, uh, I'll be, I'll be really, really brief with, uh, with this appearance. Um, I'm not one to, uh, hype things up, so I'll just say this without any further fanfare. I'm happy to announce that as of exactly this moment, we're releasing Claude 4 Opus and Claude 4 Sonnet on all of our relevant product services. Now, I know that we haven't had an Opus model in a while, so just as a reminder, Opus is the most capable and intelligent model, and Sonnet is the mid-level model that you all know and love and have been using for the last, uh, approximately year. That's a, a good balance between intelligence and efficiency. Um, we tried to design both of them so that there are, there are use cases and times when it's optimal to use each one. So I will talk very briefly about the two of them and then, and then turn it back over to Mike, and then I'll be back for the fireside. Um, uh, first, let's talk about Opus. So, it is especially designed for coding and agentic tasks. It gets state-of-the-art on Sweetbench, Terminal Bench, some other things like, like that. Um, but I think in many ways, uh, as we're often finding with large models, the benchmarks don't fully do justice to it. Um, customers who we've previewed it to have found that it can do tasks that take humans up to six or seven hours autonomously. Um, within Anthropic, I've seen some of our most senior engineers be surprised at how much more productive it has made them. And for the first time, actually, when I've, you know, looked and and seen Claude written internal summaries, documents, and, and ideas, you know, in the past, the quality was often good, but you could never really quite mistake it for a human because it always had that, that specific style. This was the first time I actually got fooled where I actually get got back and then, you know, I just read by the name really fast and I thought it referred to someone on the team, and I'm like, no, the name was Claude. Um, uh, so I, you know, I think, I think there's a, there's a, there's a lot in Opus. Um, on Sonnet, um, I think this will be for many people a strict update for, uh, a, a strict improvement from Sonnet 3.7. Um, at the same cost and better intelligence. Many customers are simply, uh, uh, switching directly from one to the other. It actually does just as well as Opus on some of the, uh, coding benchmarks, but I think it's leaner and more narrowly focused. Um, I think in particular, it addresses some of the, uh, feedback we got on Sonnet 3.7 around overeagerness, the tendency to do more than you asked for, which is sort of the opposite of laziness, which was, which was an earlier problem, and, and some of the, some of the reward hacking issues. So many of our customers have been trying it out and view it as a strong upgrade from 3.7. For example, um, you know, Cursor, Cursor here has been, uh, one of our well-known customers, has been trying it and says, saying this is, uh, this is a state-of-the-art, uh, this is, this is, this is a state-of-the-art coding model. It's a leap, uh, forward in complex codebase understanding, and we expect developers will experience, uh, across the board capability improvements. Someone who was playing with the model in person, one customer said, "What the f is this model?" It's really, it's really amazing. So, um, uh, I'll, I'll leave the details to others, but, um, the last thing I'll say is we are going to continue to improve the Claude 4 series of models. We expect to periodically release perhaps minor version updates, ideally even more frequently than we, than we have for, um, for, uh, for, for Sonnet. So it should be out there. You should be able to try it on basically all, all the surfaces as of now, except I think free tier has, has Sonnet, uh, has Sonnet only, but all the other surfaces, uh, all the API surfaces have both. Um, so, uh, really hope you enjoy the model, and I'll turn it back over to Mike.

Thank you, Dario. Two new models, and you heard it here first. Um, we'll be seeing Dario again, as you mentioned, at the end of our agenda for a Q&A, where I'll get to ask him the questions that are likely on your mind, uh, right now as well. I'm personally very excited for our customers to try both Claude Opus 4 and Sonnet 4. Our teams have loved working with them, and we think you will too. Now that Dario has shared our big model news, I'll talk more about our detailed API roadmap. Our goals for building Claude 4 were clear from the start. Wanted to build powerful AI that safely introduces new model capabilities, continue to advance the frontier for coding and AI agents, and ensure that Claude becomes your virtual collaborator. And that's exactly what we've delivered with Opus 4 and Sonnet 4. Like, uh, Sonnet 3.7, both Claude 4 models are what we call hybrid models that have two modes: near-instant responses and extended thinking for when you need deeper reasoning. I've been surprised at how many customers use the deeper reasoning even for non-coding and non-math use cases. Opus 4 is great at understanding your codebase and planning additions. It's extremely effective, uh, and accurate with everything from migrations to code refactorings, and it's also the right choice for your most complex agentic workflows. If you found you've hit a wall with other models on your use case, I think you'll be really pleasantly surprised with what you can do with Opus 4. Sonnet 4, meanwhile, excels at everyday coding tasks, app development, uh, and pair programming. It's also ideal for high-volume use cases. It perfectly balances efficiency with performance. Think of it as your always-on coding partner. Both models are live today, as Dario mentioned, in Claude and Claude Code, as well as the Anthropic API, Amazon Bedrock, and Google Cloud's Vertex AI. These models bring critical new capabilities for building AI agents. They can use tools like web search during their reasoning process, which is new, handle multiple tools in parallel, and when given access to local files, it can actually maintain memory across sessions to build knowledge over time. And I'll talk with Dario a little bit about that memory feature too later. These aren't just incremental improvements; they fundamentally change what's possible for AI agents. Now, I know the term agents gets thrown around a lot these days. I have a personal, uh, joke, which is, how many minutes into a meeting can we make it at Anthropic without saying the word agents? I think I made it 17 minutes or something. Uh, but today, what we're going to focus on is agents beyond the hype. Um, I think what's really key is with the right underlying models and the right underlying platform tools, AI agents can actually turn human imagination into tangible reality at unprecedented scale. And that's especially important for startups and developers like yourselves. I've been a founder myself. When I think back to Instagram's early days, our famously small team had to make a bunch of very painful either-or decisions. We either explore, uh, adding video to the product or focus on our core creativity. Either focus on our mobile app, and at first, our single mobile app, or, uh, expand into the web. It was all very single-track. With AI agents, startups now can run experiments in parallel, learn from users, and build products faster than ever before, which is something I've heard from many of you all. And AI agents can give you, the founders of startups, uh, access to the kind of strategic thinking that you might get from a high-powered CFO. I see our CFO in the front row here, or head of product, uh, while they're still building towards those key positions yourselves. You're not ready to make those hires, but you can hire, uh, Claude for some of those roles for now. This transformation is no longer theoretical. I see it in my role and my work every single day. I personally spend a lot of time with Claude, maybe more time with Claude than my significant other. It's fine. Uh, in fact, uh, soon after I joined Anthropic, um, I, uh, sat down with Amazon's Alexa team, and they were eager to see how Claude might become part of their vision for the future of voice assistants. At first, my team planned on presenting some slides, talking points, kind of the plan we'd make for any other customer. But in the days leading up to the meeting, I had this, like, persistent thought: Why not use Claude itself to build a hands-on demo? I thought it'd make the conversation more interesting and bring to life the potential of Claude and Alexa functionality together. The challenge was building this demo without any access to Alexa's actual codebase. We needed to create a prototype of the core Alexa functionality while also integrating Claude's capabilities, all within a tight one-week timeline, really a tight one-weekend timeline. Claude was the only reason we were able to pull this off in such a limited time frame. Our three-person team, split between San Francisco and London, built a functional prototype that showed the potential. And thanks to Claude, the effort was a success. I even got to write some of the code. You can take the engineer out of the engineering CT, overall, but you can't take it out of me. And do a lot of the front-end development, uh, for the project itself. And of course, a lot more work went into the partnership after that first meeting. Um, but Claude is now one of the models that Amazon is using for Alexa Plus, which launched earlier this year and is now rolling out. And we were able, I think, to really show the potential thanks to Claude. I've been watching this evolution towards AI for years now. When I first got a demo, early access of GitHub Copilot back in 2021, I called it the single most mind-blowing application of machine learning I've ever seen. Back in those days, 2020, we called it machine learning instead of AI. That was generations ago, but it was really clear the potential for this early glimpse of agentic AI. I had an even stronger feeling last summer when we launched Artifacts. I could describe what I wanted for a mini app or visualization, hit send, go grab coffee, and come back to Claude having built what I'd imagined. And over the following year, it's become clear we're not just building better tools; we're creating genuine collaborators. And Anthropic's economic research, uh, confirms what I've seen firsthand: for the majority of use cases, AI is augmenting people's work instead of replacing it. It's much more about tasks than entire roles. And this is similar to the influence that your best colleagues have. The most talented people you work with don't just execute; they understand your context, they learn from experience, and they know when to take the initiative versus when to just check in. Great AI agents, like the ones you can build on our platform, should excel at three capabilities. They should have contextual intelligence: understanding you and your organization's unique context and continuously learning from experience, not just following instructions, but comprehending the why and the how. That means models that learn and personalize over time, acquiring not just contextual, but also episodic and organizational memory. The way I always put this to the team is, your hundredth task with an agent should be much better than your first, just like your hundredth day with an employee should be much better than your first, if you're doing the right things around training. Second, long-running execution: handling complex multi-hour tasks without constant management, coordinating with other agents and humans as needed, so you have the context and then you can execute it over a longer period of time. And third, genuine collaboration: engaging in meaningful dialogue, adapting to your working style, and providing transparent reasoning for their actions. The key insight here is that true agency doesn't mean uncontrolled action, and autonomy doesn't mean, uh, just YOLOing it. It means intelligent autonomy balanced with clear checkpoints, maintaining human oversight for critical decisions while delegating the smaller decisions that usually consume so much of our time.

Now, let's talk about those capabilities we're announcing to serve those three needs. We'll start with our new code execution tool, which is available on the Anthropic API today. The code execution tool gives Claude an environment where it can run code, enabling it to act as a data analyst that can transform raw data into visual insights. Claude doesn't just write code anymore; now it can execute it. It sees the results, and it can iteratively refine the results and the code to better highlight patterns in your data. Here, uh, we'll show Claude analyzing sales data to see how a specific type of product is performing. Claude can load your data set, clean it, generate exploratory charts, and drill down into anomalies, all in real time. As someone who started their career as a data visualization analyst, this resonates a bunch with me. And the code execution tool is even more powerful when combined with the intelligence of the Claude 4 models. This is what we mean by agency: the ability to take a complex task and see it through to completion. These are the first models capable of handling hours of tasks, saving you half, maybe even full days, when you work alongside them. And not just writing code snippets, but refactoring entire codebases or implementing complex features from scratch. To give you a sense of the kind of progress that we're seeing, back in the day when I started, you could delegate maybe minutes of work to Claude 3. Claude 3, meanwhile, could work autonomously for about 45 minutes without losing its thread. And now we're breaking into hours of work that Claude can take on autonomously. As you saw earlier, Rocketin mentioned that they ran Claude independently for an incredible seven hours with sustained performance. It can do it without losing the thread, especially as it's able to manage its memory and its own to-do list. We've already integrated this power where you work. Hopefully, you're all familiar with Claude Code, our agentic coding tool that we launched in research preview a few months ago. We're moving Claude Code to general access today. This actually started as an internal exploratory project by Boris, one of our tech leads. This is his announcement post, who wanted Claude to help him code directly in the terminal. Very early, we still called it Claude CLI internally. I think some of our best innovations, like Artifacts and Claude Code, have really come from this kind of bottoms-up experimentation. It's part of the culture we try to foster at Anthropic. Within just two days of launching it internally, our usage chart went vertical. People talk about product-market fit; we really often talk about product-Anthropic fit: like, are people internally dog-fooding? Are they using it? Today, most Anthropic employees rely on it for everything from routine coding to large-scale migrations. I've watched some of our most advanced coders run multiple copies of Claude Code across multiple terminal windows. They're moving from just being engineers to being managers of several autonomous agents, tackling everything from simple coding tasks to complex full-stack development projects across multiple codebases. I realized I was using Claude Code, and I would run one in our front-end repo, and one in our backend repo. And one of our Claude Code engineers is like, "You're doing it wrong. Just run it in the root. Claude can figure out where it's going to be able to do it across all of them." And it does it beautifully. And that's, that's changed how I've used it already. The vast majority of, of Anthropic developers use Claude Code daily. To give you a sense of the impact it's had on our team, it's shortened our technical onboarding time to get engineers up to speed from two to three weeks to two to three days. I've really seen it how it can help you build an understanding of the codebase, especially a large monolith like ours, very rapidly, as it's fantastic at navigating code. And today, we're bringing Claude Code capabilities directly into VS Code and JetBrains with full diff views and agentic workflow management built into the editors. And we're also introducing the Claude Code SDK, so you can build your own applications on top of the same core agent as Claude Code. As an example of the possibilities of the SDK, you can now run Claude Code in GitHub. You can tag Claude in a GitHub pull request or an issue, and it will respond to reviewer feedback, modified code, or implement test coverage. We're also focused on what we call closing the loop. So Claude Code is now helping build itself, and it demonstrates the power of self-improvement as it speeds up its own development. It's incredible how Claude Code empowers developers like yourselves to get more done. I think back to when I was building Instagram, our team was between two and six engineers, you know, before we got acquired, and we were supporting two mobile platforms. We would have been able to produce prototypes in days and not weeks if we had agentic coding products like this. We've talked a lot about building performant, reliable agents. Now, agency without responsibility is dangerous, especially when you're talking about something that is self-improving like our Claude Code, uh, product, and even more so in enterprise settings with stringent security and compliance requirements. I think widespread adoption of agents will require improving model discernment and judgment around confidentiality, decision-making, and coordination. So our models are already good at this, but we'll continue to improve, making sure that they know what's confidential, they know what to reveal, um, and that you can trust them in a production setting. That's why every feature that we build around our models incorporates what we call architectural safety checkpoints and controls, not just carte blanche agents, pausing on major decisions. While users can define which actions need human approvals, which we've also built into the model context protocol, they're robust against exploitation. We test them, we battle-test them, you know, uh, a lot around things like prompt injection, and they're also transparent by design with clear feedback loops and observable behavior. When you trust your agents to act autonomously, you're free to focus on innovation instead of mitigation. Another area we've invested heavily in is interpretability, the science of understanding exactly what's going on inside the minds of AI models. Dario recently wrote about the urgency of understanding how our AI systems actually work. If you read his essay, "The Urgency of Interpretability," what he calls the race between model intelligence and interpretability. Effectively, we want to be able to give our AI an MRI to see what, what it's thinking about and spot any potential problems like deception, so we can steer it in the right direction. When I joined Anthropic, I was excited about how our research pipeline could directly fuel our products. Take Golden Gate Claude. I pushed to ship that in our second, my second week at Anthropic because it didn't feel like it was just a good research paper; it would make a fantastic demo, a visceral demonstration of how interpretability works in action. When we amplified the Golden Gate Bridge feature inside the, uh, uh, Claude neural network, we saw that suddenly we could see what it means to manipulate the inner workings of AI and, in this case, make it deeply obsessed with our favorite bridge. The techniques that we used to create Golden Gate Claude could, in the future, help us reduce, uh, model harmful model behaviors or improve model performance for specific domains. And as we start employing virtual collaborators around companies, my hope is that we can lean on techniques like interpretability and auditability to be a cornerstone of their work, so we can figure out what they're doing at scale. These are the kinds of breakthroughs that, that are going to help us transform abstract research into tangible product capabilities. As you saw earlier, we're now at the point where AI models can handle hours of autonomous work, and that's a capability that's doubling every several months. But raw model capability alone isn't enough to unlock these multi-hour workflows. In practice, agents also need access to real-world information, a connection to your existing systems, and cost-efficient scaling. That's why we're launching four interconnected capabilities to help power agents with context and help them scale. So, first off, starting today, you can now connect the Model Context Protocol directly through our API. MCP is already being used by Microsoft, Google, OpenAI, Block, Atlassian, Zapier, Linear, and many more. This was the dream list when we started creating the MCP protocol and we open-sourced it. This was the dream list, like, maybe one day we'll get these companies to adopt it. It's less than a year, and they, they've all, uh, come on board. MCP acts as the universal translator and connector for AI agents, enabling seamless connection to your existing systems without needing to write a custom, bespoke integration every single time. This lays the foundation of what could become the agent economy, where specialized agents have access to the data and tools they need to tackle complex challenges. Second, web search gives Claude real-time access to current information. This is intelligent data augmentation that allows Claude to reason about current events, market trends, and emerging, emerging technologies. It's really powerful in combination with the MCP feature as well. You can imagine searching across an internal knowledge source, making some, uh, new, uh, insights, and then going off and searching the web to contextualize them. Third, the Files API is available today in the API to streamline how developers access and store documents, simplifying development workflows. We're also releasing a cookbook to help developers build that memory functionality that I mentioned directly into their applications. These new Claude 4 models have shown significant improvement in what we call self-managed memory. So you'll find that this works surprisingly well and it can be achieved with very little additional overhead by using the Files API, as we demonstrate in that cookbook. You'll see Claude both read and write to these memory files and maintain context over time. Last, power needs to be practical and scalable. We want to ensure that we can grow with you from prototype to production to millions of users, so that you're able to control cost and improve efficiency. We want Claude to work for you as you succeed and reach massive scale. That's why prompt caching was our most requested feature, one of our most popular API features. With prompt caching, customers can provide Claude with more context, uh, and background knowledge and example outputs, reducing costs by up to 90% and latency by up to 85% for long prompts. Now, every customer I talked to had one very clear request on prompt caching that we're delivering today, which is a longer time-to-live, TTL. So, in addition to the five-minute TTL we had out-of-the-box with prompt caching, today we're launching a premium one-hour TTL, which is a 12x improvement that dramatically reduces cost for long-running agent workflows. This infrastructure makes agent applications viable at scale. So, these capabilities all compound when we think about building features into the API. We don't think about them as one-off; we think about how do they complement each other, how do they form a cohesive story. Claude can now execute code, understand your systems, access current information on the web, creating the foundation for agents that operate with full context, even for long-running tasks. And it can use the Files API to maintain memory and context during that entire execution. Everything you've seen this morning is just the beginning. Our roadmap continues to build on three pillars. The first is industry-leading agentic tools and applications, so you can use Claude autonomously to handle hours of work, knowing it can use the code environment, uh, code execution tool to execute code in its own environment. Claude Code is now generally available, integrating with VS Code and JetBrains, so you can use the extensive SDKs to build your own custom workflows, including inside GitHub. We'll continue to push on integrating more context in the API. Our updates today allow you to bring this context via the Model Context Protocol, as well as build on real-time updates from the web and execute, uh, complex workflows across any data source and across anything in the API via MCP. And finally, efficient scaling. As of today, you can use expanded one-hour prompt caching to optimize performance and cost at scale. Each advancement builds on what we've discussed today with Claude 4 as the foundation: Opus 4 for your most complex, uh, agentic workflows, Sonnet 4 as your daily driver for everyday intelligence. We're enabling a new class of applications. Code execution expands the hours of work that Claude can do. MCP expands the comprehensive information that Claude can retrieve. And our platform updates ensure our models become increasingly efficient for every dollar spent. We're actively learning from developers like yourselves on how you use these tools, so please keep the feedback coming. I love API feedback. If you don't know this about me, like, absolutely, like, ping me. I love hearing the feedback and how we can continue to improve the API for developers. Or yourselves. And MCP is a perfect example of this. It started as an internal idea and then began and graduated to community experimentation, and now it's a core platform feature. If you watch the Microsoft Build keynote, they're building MCP into so much of their, uh, of their real infrastructure as well. We want to create an ecosystem of AI agents where we have the feedback loops to make them actually useful for you. Today, we stand at a major threshold. Our latest models, combined with all the latest tools that we've released, are giving the seeds of a new era. The future isn't about AI doing human work; it's about AI helping humans do superhuman work. And I'm really excited to build this vision together with you, and I can't wait to see the kinds of applications it powers for all of your companies. And to show you what's possible, I'm next going to hand the mic to Cat Woo from our product team to demonstrate how accessing our new models inside Claude Code transforms your development workflows, helping you ship complex multi-day tasks in a single conversation. Welcome again to Code with Claude, and thanks again. Hope you enjoy the rest of your day.

Hi everyone. I'm Cat Woo, product manager for Claude Code. As Mike mentioned, we recently launched Claude Code, our agent coding tool, in research preview. Claude Code gives developers direct access to the raw power of Anthropic's models right where they work, in their terminals. As of today, Claude Code is generally available. Throughout computing history, we've continually moved to higher levels of abstraction: from machine code to assembly to high-level languages. With Claude Code and increasingly agentic models, we're witnessing another step forward. Developers are shifting from asking for specific functions to describing entire features, guiding AI, and changing how software is built. Today, we're bringing the new Claude 4 models to Claude Code, making it an even more powerful and capable coding agent. And on top of new models, we're releasing several new features in Claude Code focused on making it a more versatile coding agent across your whole dev lifecycle. First, Claude Code now integrates with VS Code and JetBrains, bringing it to familiar interfaces for millions of developers. As Claude Code works, you can now see its proposed changes in-line in your editor. We're also releasing the Claude Code SDK, which allows developers to use Claude Code as a building block in your applications and workflows. The possibilities are endless with the SDK. To showcase these possibilities, we're releasing an open-source example of the SDK in action with Claude Code in GitHub. You can tag Claude directly on pull requests and issues in GitHub, and Claude Code will respond to reviewer feedback, fix CI errors, and add new functionality. With these additions, Claude Code now works everywhere you do, acting as a virtual teammate across all surfaces: in the terminal for deep development work, in remote environments like GitHub for automated workflows, built on the SDK, and in the IDE for seamless review, all in Claude Code. It's a versatile coding agent for accelerating development wherever you are, whether you're working directly with Claude Code interactively or using it asynchronously. Great. My favorite part: let's see what these updates look like in a demo. I'm going to show Claude Code tackling a real dev task in a product that many of you are familiar with. We'll use Excaladraw, an open-source whiteboarding tool, and ask Claude Code to implement one of their most requested features: adding a table component. How many of you have gotten that feature request that's been on your backlog for ages that you know your users would love, but you just haven't had the time to build? This is the kind of task that we can handle much faster with Claude Code. Normally, for a task like this, I would set Claude to work, make some coffee, catch up on email and Slack, and come back when the outputs are ready. But I only have 10 minutes with you all today, so let's show a sped-up, but real, workflow. Here's the Excaladraw repo open in VS Code. Let's write a prompt to tell Claude Code our requirements. We'll ask Claude Code to add a table component that supports custom dimensions, drag to resize, and all of Excaladraw's other styling options. Here's where it gets exciting. Claude Code will first create a to-do list for how it'll approach the entire problem. Then, we can see that Claude Code will start to explore the codebase, starting with the file that we already have open for context. The best part of the IDE integration is the ability to see diffs in-line in the editor. This way, you can see the surrounding code for more context, so you can accept changes with confidence or give Claude Code feedback. We can approve each edit as Claude Code works, or we can let Claude Code continue making edits with auto-accept mode, letting us balance visibility and control. In this demo, we gave Claude Code the ability to make edits, run lint and tests, and make PRs. So, Claude Code worked for 90 minutes on this task. I wish I could show you the whole thing, but we need to speed things up. What you're seeing is actual unedited output from Claude Code. An hour and a half later, and it's done. It added table functionality, wrote tests to validate the change, and iterated until lint and test passed. This normally required us to understand the codebase architecture and how every single other tool was implemented. In this case, Claude Code is literally doing hours of work for us. Pretty impressive, right? Now, let's run Excaladraw locally and just make sure the feature works as we expect. Let's check that we have a fully functional table component by making a three-row by three-column table. Great. We can reposition the table, we can drag to resize, we can change the border pattern and color, and we can add text to cells. This also integrates with Excaladraw's existing UI. All of this was done with one prompt in Claude Code. [Applause] Next, we'll ask Claude Code to use the GitHub CLI to create a pull request for this branch. Cool. Let's click in. Now we have our pull request. This is where the Claude Code SDK shines. It lets us build custom workflows on top of Claude Code, including through GitHub Actions. For this PR, I'd like to update the docs. Instead of going back to the IDE, we can just tag @Claude and ask it to update our documentation for us. Behind the scenes, this triggers a GitHub action that runs Claude Code. Claude comments on the PR as it works, and it'll, it'll make a commit for us when it's done. You can also tag @Claude on a GitHub issue, and it'll also make a PR for you there. With this feature, Claude Code meets users on even more surfaces where they're already working. Devs no longer need to context switch in their local environment, and you can even kick off runs on the go. This is all built on the Claude Code SDK. Beyond powering GitHub Actions, we've seen customers do incredible things with the SDK, including running many Claude Codes in parallel to fix flaky tests, increase test coverage, and even do on-call triage. Cool. It looks like the action is done running, and we can see Claude Code updating its comments to let us know what it did. Let's click into the commit and see Claude's changes. It updated the documentation for us in our PR and committed it without us having to do a thing. In just 10 minutes, you've seen Claude Code tackle a complex task that would have taken days to implement manually, writing hundreds of lines of code, integrating seamlessly with Excaladraw's existing features, and doing hours of work for us. All of this is available to you today. Claude Code in GitHub Actions powered by our SDK is available in beta, and you can install it by running a simple command on the screen. Within Claude, the VS Code and JetBrains IDE extensions are also live in beta. Just run Claude from your IDE to install. Last but not least, our latest models, Claude Opus 4 and Claude Sonnet 4, are available to Claude Code users today. Claude Code shows what's possible when AI can truly understand and work with code to build powerful agents. Whether coding assistants or applications in any domain, you need more than just intelligent models; you need the right platform. Please welcome Michael Gersonenhober, who will show you exactly how we're making that possible.

Thanks so much, Cat, and good morning, everybody. Thank you so much for being here. I'm Michael Gersonenhober, head of product for the API platform at Anthropic. How many people here use AI-generated code already to write their applications? Yeah. And how many of those are using AI at their core feature delivery, like everybody here? That's what I thought. Most applications in the world will be built by people already trying to solve the world's problems. Whether you pass Ble Code, whiteboard interviews, or getting started with Vibes, we're all software engineers now. But writing code is just the start. You need to more quickly build stable, secure, and maintainable AI applications. And that's why we built the Anthropic platform: a complete toolkit designed for building state-of-the-art AI applications and agents. Our platform is already powering most of the world's AI delivery in every domain. In finance, TurboTax helps millions of customers confidently file taxes with federal tax explainers. In healthcare, Novo Nordisk is using Claude to draft clinical study reports in less than 10 minutes instead of 15 weeks. And the world's best coding assistants run on our platform. Each of these companies took Claude's intelligence and turned it into something uniquely valuable for their users. At its foundation, our platform provides reliable access to Claude through our model inference service, which includes the Messages API and essential tools like prompt caching to optimize performance and costs over 50% of all input tokens are cached on the platform, doubling the effective context window for our models. Notion can put vast amounts of your documents in the context window but maintain snappy, real-time execution. This lets them adopt your voice for creative writing and virtually eliminate hallucination. Starting today, we're extending the cache time-to-live from 5 minutes to 1 hour. Your agents can now maintain complex context across the entire user session without breaking the bank. But that's just a foundation. To build powerful agents, our platform provides powerful building blocks. As Mike shared, we're releasing two new capabilities: the Files API and a code execution tool. Just like you and me, there are some problems that are easier to solve by writing a script. Our platform lets your agents write their own code in production, just like you would. These new features join existing components like web search for real-time information and citations for grounding responses in source documents. When Thompson Reuters provides analysis to attorneys in co-counsel, it's critical that they ground this in their legal research, in case law, not in the model's training data. Our platform also connects your agents and your data and business systems through Model Context Protocol. MCP has taken off within our developer ecosystem with over 3,000 integrations built by the community. Whether your agent is accessing application errors with Sentry, triggering Zapier workflows, or creating Asana tasks, the MCP connector enables the model to interact with any tool, data, or app your task requires. And today, the platform makes it even easier by handling all the technical complexity of tool and API calling for you. One thing that I want to emphasize about the platform is the composability of the APIs. They're building blocks that work together as well as they work apart, helping to solve unique problems that can't be coerced into a cookie-cutter shape. Think of Claude as the architect and general contractor for your agent. It doesn't execute predefined sequences or stack components randomly. Instead, it intelligently determines which materials you need, in what order, and how they fit together to create something far more powerful than any individual element. Let me show you what I mean. When you build an agent for complex financial analysis, Claude intelligently assesses the task and orchestrates the right tools: using MCP to access financial data, spinning up code execution for statistical analysis, searching the web for real-time market data, and grounding insights with citations for accuracy and compliance, iterating and refining based on results. No hard-coded workflow, no brittle scripts, just intelligent orchestration that allows you to build powerful agents and seamlessly adopt new capabilities as our research brings them to life. We understand that prompt quality can make or break an AI application, which is why we created dev tools like the Prompt Improver and Evaluations, along with new observability features that help you get to production and scale faster. Today, we're already helping developers build faster with resources like cookbooks and guides that show you how to implement features like memory into your applications. In the future, we'll adapt these for programmatic access and host them directly on the platform so you can build even more powerful agents that can research and remember on their own in production. Everything we've built centers on one goal: helping you ship better AI faster. The Anthropic platform isn't just tools; it's your path to building industry-leading agents. So thank you all for being here today with me at, at Code with Claude. I'll be on the floor the rest of the conference, but it's my privilege to welcome Mario Rodriguez from GitHub to show you exactly what this looks like in production.

[Applause] Thank you. Thank you, Michael. And I am here, um, thrilled to be with you all. We at GitHub are incredibly excited to be part of this energy and innovation and to share more about our deepening partnership with Anthropic. Um, this amazing team. Everything GitHub does is anchored on two core beliefs, right? Number one is giving developers choice, and number two is giving them the best developer experience. At GitHub Universe last year, we kicked off the relationship with Anthropic. We announced Claude Sonnet 3.5 support in VS Code and also in our conversational experiences. And we did this because we share a fundamental belief with Anthropic that AI can be a powerful force and a force multiplier for developers, augmenting their capabilities, not replacing, augmenting their capabilities, and freeing them up to focus on what they do best, which is imagination and creativity. Are of being a software developer is being a whiz. Since we haven't expanded since then, we have expanded the partnership and experiences across VS Code, GitHub.com, and our mobile app, just to mention a few. And today, I am delighted to announce that GitHub Copilot supports Claude Sonnet 4 and Opus 4, available right now. We just pulled the trigger right when Dario announced it. And every one of those services, that is what SIM shipping is all about. Let me tell you, it's really hard to do. I don't know if you've done it with every application that you have done, but it's incredibly hard to do. So thanks to all of the teams that make that happen. Now, as you all surely know, the future of code is what agent. An agent mode in VS Code is our autonomous pair programmer that can perform multi-step coding tasks based on your natural language commands. We've seen firsthand how having Claude's intelligence directly within the editor truly helps developers understand complex code bases, um, get faster code to production, and increase their productivity without ever leaving the environment they already know, love, and trust. But even that, right, even that is single-threaded. And in my opinion, the future is multi-threaded. You think about it, you're in your editor, it becomes a waiting room. You're, you're going faster, but it's still a waiting room. And that's why on Monday, we took one step further and announced GitHub's Copilot Coding Agent. Now, our coding agents, these are autonomous, asynchronous peer programmers. Not pair anymore; now it's your peer programmer embedded directly into GitHub Copilot's Coding Agent. Is currently powered by, you probably guessed it, Claude Sonnet. Uh, and, you know, the reason we chose that was very clear to me. So let me just walk you through three things that made that decision possible. Number one, our evaluation showed that Claude demonstrated three main strengths: right, strong software engineering and coding knowledge, powerful problem-solving, and that's very important because sometimes you have to go and look at the code and find the right place to make that edit. And then number three, excellent instruction following, and specifically when thinking about tools and MCP. So when you're building for ejected coding, dealing with these things and large code bases and system prompts, you also need something else, which is caching, right? And that prompt caching, bless you, that prompt caching support we get from the Anthropic API lets us build these experiences in a most cost-effective way. Every token counts, and every token counts also on the price side. So the more we save those, the better experience we could provide our customers. Now, on top of that, Claude was already the most frequently selected model in agent mode. So once we put all of those things together, it was very clear to us that Claude Sonnet was the right model choice for agent coding in GitHub scenarios. Now, with Claude Sonnet 4, we've seen improvement in all of these areas, not just aggregate benchmarks like Sweet Benchmarks, but more importantly on our real-world evaluation suites as well. Now, our collaboration goes deeper than this, right? It's not just about integrating models directly. We've been working closely with Anthropic to officially adopt and scale MCP. We're combining intelligence. If you think about this, like, these models are incredibly intelligent. You stack like three.

PhDs on them with knowledge. So how do you get knowledge into that intelligent model? Well, the answer to us is MCP and tools, and that really unlocks the next acceleration of developer tools.

Recently, Kevin Scott, that's Microsoft CTO, made the analogy that MCP is like the HTTP protocol of the web, and I completely agree with him. So, if you have not adopted MCP, do it today, right after this keynote, go and play with it. It's that important. It's the way you get knowledge into these intelligent models.

Now, as we step into this new era of software development, we're transforming GitHub's platform from an AI-infused into AI-native. From creation to deployment, we envision this SDLC powered by an agentic layer at the top of it that spans that inner loop where you are coding and that outer loop, those asynchronous experiences. And you are going to be an active collaborator every single step of the way. The reason why we say co-pilot is the human is at the center, and then there's agents helping you. That is why we're announcing a new partnership that integrates what Kai just showed you, Cloud Code and the extensible Cloud Code SDK, directly into GitHub's agent platform. This opens up new possibilities to customize Cloud Code, remotely invoke it from new surfaces that are embedded into GitHub and our workflows, again, all on the GitHub platform.

Now, we've already done a lot, uh, but the journey with Anthropic is still just beginning, in our opinion. We believe that by bringing together GitHub's deep, deep understanding of developers and Anthropic's AI capabilities through Cloud and the platform APIs, we will and we can unlock a future that is more intuitive, more efficient, more ultimately more human. That human power is important. So, I'm excited to see what we continue to build together and also what each of you builds with us. So, thank you so much, and please welcome back to the stage, Mike Kriger.

Thank you, sir. Hello again, and thanks again to Mario, to Michael, and to Cat. Um, I love the GitHub integration. The last project I did, I actually was like, "Oh, I can actually just install Cloud Code into a GitHub Codespace." And all of a sudden, I have Cloud Code against the repo that I've already been building. It was really great to hear from each of them and hear all about the exciting work being done with Cloud.

So, to close out the show, I'd like to dive a little bit deeper into Claude 4, our research direction, uh, and what developers can expect next from Anthropic. Um, so please help me welcome back to the stage, Dario, for our one-on-one conversation.

Welcome back, Dario. Hello again. This is great. This is like our one-on-one in front of the whole audience. This is great. Um, so Claude 4, uh, uh, is out. Claude Sonnet 4 and Claude Claude Opus 4 are available. Um, what excites you the most about the Claude 4 models, and how does it change your thinking about what's possible in the next 12 months?

Yeah, so, um, I, I think abstractly, the thing I'm most excited about is, you know, every time you have a new class of models, there's like more you can do with it, right? So, uh, uh, you know, we're, we're, we're going to be releasing, uh, models after Claude 4. There'll probably be a Claude 4.1 at some point, just like we did with, uh, with Sonnet, uh, uh, 3.5. And I think we're just at the beginning of, of, of, you know, what, what, what we can do with the new, the new generation of model. In terms of tasks, I think the autonomy is going to go, uh, uh, is going to go much further than it has already. Just the ability to give, you know, set your model free and and give it the ability to, you know, do something for, for a long period of time. I think we're, I think we're very much, very much still, still at the beginning of that.

Um, uh, I'm, I'm actually increasingly excited about the models for cybersecurity tasks. I mean, you can think of cybersecurity as like a, a subset of, of, of, of coding tasks, but they tend to be higher-end coding tasks. And so I think we're maybe finally hitting the threshold for that. And then, as a, as a former biologist, I'm, I'm always excited about use of the models for, uh, for, you know, biomedical and kind of, kind of, kind of detailed, uh, scientific research work, which I think Opus and Opus in particular is going to be good. Opus in particular, I think, is going to be particularly, particularly strong at that.

It really connects, I think, to Machines of Loving Grace. So how does Claude 4 fit into that trajectory overall? I like to joke that people think of Machines of Loving Grace as an essay, and I think of it as a product roadmap for the next few years. And curious how Claude 4 fits into that journey.

Yeah, it was sort of a product roadmap that I wrote without knowing how to, how to actually get to it, and and kind of said, all right guys, then this is your work, this is your job. Um, uh, yeah, uh, you know, we're, I think we're increasingly thinking about on the biology side of things, and and software is part of that, right? Where, and, you know, increasing amount because biology increasingly involves data, even involved data 10 years ago when when I, uh, when I was a biologist. Uh, uh, I, I think, I think, I think more and more of it is is going to be, okay, we have these models that know a lot about biology, and they can help write code. And so if you're a computational biologist, I think these models will will really accelerate what, what you can do. And, you know, we have a number of customers who are who are trying out the models for, for these tasks. I guess we'll, we'll get to that in a bit.

Yeah, I think, uh, one of the first hackathons we did after we, uh, released MCP, somebody hooked up MCP to one of those like plotters that so to do drawing. And so Claude could draw for. It's actually really fun to like see what Claude draws for itself. But it was like, the first one was like, MCPs don't just have to be connecting to digital systems, they could also be connecting to the real world. So like, when you'll be able to drive lab equipment, VMC P, I think is an interesting, uh, question for the soon we'll be able to test Claude by connecting it to a polygraph.

Yeah, I love that idea. Are you lying? Who needs interpretability when we have the polygraph?

Um, uh, you mentioned that moment where you were, you know, convinced that Claude, the Claude-written content was, was human-written. Um, any other breakthrough moments in watching us all, you know, uh, dogfood, uh, Claude 4, or even try it yourself, that made you realize this model felt different?

You know, I, I didn't actually understand, I didn't actually understand the details, but like there were several people in our side. There was a moment a few weeks before the model launched where someone said, "Oh my god, this model just like oneshotted this like incredibly difficult performance engineering task." And and no model had ever had ever done anything like that before. I, I will say that there, there's, there's this almost, almost like superstitious process in the model development where like, it, it, it somehow all comes together at the last moment, even if the training process is all planned out. Like, just some of the models' abilities, maybe it's something about their interaction with people, maybe it's something about like, just making it the last bit better matters, maybe it's people getting used to the model and prompting it. But, but you, you always find the, the early versions of the model, um, you know, people are struggling to figure out how to use them, and then, and then you finally get to a point and people are like, this works for me all the time. And there's that, there's that alchemy that happens somehow always the last moment. If you read, uh, Creativity, Inc. by Ed Catmull, he talks about the same process with all the Pixar movies. Like they're really bad until like two days before they're supposed to go out, and I feel the same way about our models. Not that they're really bad, but they're like, they're not quite there, and then suddenly they click, and we're like, I can't wait to get this out to people. It, it doesn't make, it doesn't make any sense because like the training process is uniform, and you know, you know, you would think that that it doesn't work that way, that it's all a rational process, but it's absolutely not. There's no point on the RL curve at all that that they come together. It comes together at the last minute. I don't know why. It's a real moment.

Many people in the audience are developers here. And a question that I know has come up internally as people, you know, think about, uh, how AI is developing is, which parts of the software engineering job will AI take over? And what becomes more important in a world where we have autonomous agents being able to do, do a lot of software engineering?

Yeah, um, so probably like many people here, I, I read with great interest, uh, Steve Yegge's blog post a couple months ago, "Revenge of the Junior Developer." He had some, uh, he had some similar blog posts. He had some similar blog posts around that, actually came in to visit us, even. Um, uh, uh, and and that laid out, I think, the vision of where things are going. Maybe, maybe even better than I could, which is that we're gradually going, we're gradually going to more and more autonomy of the models, right? We had this phase where you would do basically autocomplete. Now there's this thing that I guess people have called vibe coding, uh, uh, uh, and, and, you know, then, then we're going more to kind of like, you can dispatch the agents to to do things. And I think with, with Claude Code, we're going to go, go more in the direction of, you know, you can dispatch the agents to do things, and I'm sure we'll have other product surfaces that that that allow you to do that as well. And I think we're, we're heading to a world where a human developer can kind of manage a fleet of agents and say, you go off and do this, you go off and do this, you go off and do that. But, but I think continued human involvement is going to be, going to be important for the quality control, to make sure they do the right things, to get the details right. And so, you know, working together on both the models and the product surface around it to get the details right is going to be really important. I think it's also highlighted to me, it makes the stuff that is inefficient in your work way more painful because it's taking you away from like this flow of building. And so at least it's made me realize like where we're spending too much time on cross-functional alignment and, you know, road mapping when like we just should be trying to get more building. So it's, I've, it's become more painful as the engineering part has has been sped up as well.

So, there's endless debate, uh, you know, around the industry around, you know, uh, bigger models or smaller architectures, which will win in the long run. Um, you're famous for, you know, popularizing and and pioneering the scaling laws paper. What's your current take on, you know, the extreme being pre-training dead? Is pre-training all that matters still, and its role, you know, relative to to post-training?

I mean, without getting too specific, I would say that, you know, the Claude 4 models embody advances in both pre-training and post-training. Um, so we're continuing to see the pre-training scaling laws work the way that they've worked before, uh, uh, and we're also continuing to see continued advances in, uh, post-training, and, and they kind of, they kind of complement each other. Uh, and I, I think we're going to continue seeing advances in both of those. I think we're also going to continue to scale up. So we have these, these multiple trends, these multiple sources of exponential growth, and they're, they're all going to compound with each other, right? That's, that's why I think all of this is going to go very fast. One of the reasons I liked Yegge's blog post is that it was someone who was not me repeating the mantra of like, it's only going to be a year or two until until these things are like, you know, are basically peers to us. It's insane that 3.5 was just in February, right? It's, it feels like a year ago, but it was just three months ago. I, I, I know it, I know it feels like it's like, oh, this is, this feels like an obsolete model or something. And, you know, it was, it's less is like two and a half months or something. It's like the, the time scales are the time scales are compressing. And I often say that, uh, being in the AI field, I will go on a very brief digression, be being in the AI field, it feels like you're getting on a, a spaceship from leaving Earth at relativistic speeds. And, you know, one day you wake up, and, you know, it's like, you know, one day on your spaceship, two days on Earth. So you have to take in the news of two days. It accelerates. One day on your spaceship, three days on Earth. And, and, and, you know, that, that's just what it feels like being, being on this ride.

That resonates. I've heard the metaphor before, but it absolutely does. Um, maybe on the post-training front, one of the things that I got really excited about seeing developed in Claude 4 has been this concept of memory and having the the model being able to manage it. Memory, maybe talk a sec about why that's important and what that kind of enables.

Uh, sorry, repeat the question. Like for, uh, the model to be able to manage its own memory and be able to handle those long horizon tasks as well? Yes, yes. We have found that to be, uh, very useful. I think one, one place we found it to be useful is Pokemon, right? Um, uh, where the model's able to like remember its state. But, you know, presumably, it's, it's useful for many things other than just Pokemon. Um, uh, uh, but, uh, um, no, I think, I think it's great that, you know, the model, you know, just as a human would, like when I'm thinking, I'll write a bunch of notes, and, you know, then I'll like recall those notes at a later time. Or, you know, there's just a lot of, lot of intermediate work that I have to do that, that, you know, and models do that to some extent when they, when they reason, when they have, you know, like our, our reasoning traces. But, you know, not, not everything I do can be incorporated in one scratchpad, right? There's like presentations, there's, um, you know, individual documents that I, that I write. And so models are the same, right? The, the idea for them to kind of, you know, be able to create files, to do things with those files, to load data, and to kind of seamlessly interleave those things, right? The, the one of the new features that we have is this, this kind of interleaved re, interleaved reasoning and taking actions. And some of those actions can be storing data, recalling data. Again, the affordances that the models have are gradually converging towards the affordances that a human has, which I think is, is the way that it should be.

One of my mind-blowing moments in Claude 4 so far was we added like basically a to-do list scratchpad to Cloud Code. And just watching it turn through the to-do list, and then as it thought of more things to do, add to the to-do list, check things off, strike out what was no longer relevant. It really mimicked, I think, how people managed their own work and how they think about, uh, completion along the way. And then the interleaved reasoning, uh, and tool use as well. I saw a write-up this morning on MacStories where it was using a tool, it was an MCP, and it hit a rate limit with the backend MCP server. And because it was doing the reasoning, it was long, I was like, "Hmm, I probably hit a rate limit. Let me try this other approach to do this as well." And so like, that ability to reason and remediate as part of tool use, I think, is really powerful.

I'd love to touch on race to the top. So, um, uh, safety and and capabilities are often, you know, uh, thought of as being at odds with each other. And your thesis is exactly the opposite, and that these two things can move in tandem. I found that very inspiring and one of the reasons I joined here. But maybe touch on how you think of of race to the top.

Yeah, so, you know, I think, I think it, it, uh, applies to things, you know, from the, from the, uh, from the very mundane and simple and commercial to kind of, you know, the grand directions that, that, that, that, that, that AI is going in the future. Um, so, you know, I, you know, I think, I think when we, when we talk to customers, we have a number of customers who, you know, care a lot about making sure that the behavior of their AI models is predictable, that it's trustworthy. Um, uh, and I think that's aligned with what some of we're, what, what we're trying to do in the long term for, uh, you know, making sure that models in a more grand sense stay in line with human intent. Um, so there's, there's this nice, there's this nice synergy here. And, you know, I think whenever we're able to do so, whenever we think it's reasonable or responsible to do so, we do want to provide tools for the community. So, M, MC, MCP, MCP is an example of that. Um, I, I myself was actually surprised at the the pace at which everyone seems to have standardized around around MCP. I mean, it was, it was very strange. We released it in November. I wouldn't say there was like a huge reaction immediately, but then, but then within three or four months, you know, it kind of become, it kind of become the standard. Again, there's again this this feeling of like being on the spaceship accelerating from Earth and and and, you know, experiencing, you know, larger and larger time, time dilation constants. Yeah.

Um, where it's, you know, like, you know, think of like USB and other standards, you know, think of like standards in the '90s or the '00s, like, you know, this would take, it would take years for people to converge on something. Yeah. And even in talking to other participants in the industry around MCP, they're like, we don't want to slow down whatever is working on MCP. Like, we do want like some, you know, help on steering, but like, this is, you've captured lightning in a bottle. Let's make sure it becomes the new protocol and the standard by which we interoperate agents as well.

Um, uh, maybe tied together the race to the top. I loved your urgency of interpretability essay. You have a background in neuroscience as well. Can you talk a little bit about how you see the co-development of interpretability and, um, machine intelligence?

Yeah, so, um, you know, I think 10 years ago, uh, many people thought that neuroscience would tell us about how to do AI. Um, and indeed, there, you know, are a number of former neuroscientists in the field. I'm not, I'm not the only one. There, you know, there are other lab leaders, some who have that, uh, who have that background. And, you know, I found at a high level, there's some inspiration, but I wouldn't say I've said, oh, you know, this is how the, you know, this thing we know from the hypothalamus, we can use for, you know, for, for making these models. It's, it's all been pretty much from scratch. But interestingly, things have gone the other way more, which is that using interpretability, we're able to see inside models. And although of course, they're not ex made in exactly the same way the human brain is, at, at a, you know, the, a kind of superficial level, there's, there's a lot of differences. A lot of the conceptual patterns we have found inside models, sometimes they then get replicated in, replicated in neuroscience research. There was something about like high-low frequency detectors in vision, um, that, uh, was found via interpretability, via, via one of one of the people on Chris Olah's team. And then a couple years later, a neuroscientist actually replicated it in, in animal brains. The idea that, for example, vision models separate out, you know, they have one path that that tends to correspond to color, and, you know, another path that corresponds to, uh, you know, uh, brightness or to the boundaries between objects. These seem to be natural distinctions in the world, right, that are that are kind of there to be discovered. And anytime you have any kind of abstract learning system, whether it's artificial or biological, you kind of discover the same thing. So it's very interesting. I'm really curious how the Circuits paper ends up affecting neuroscience research as well.

Let's move into the five to 10 year time horizon. Um, to the extent that that is even possible in AI, as as we move relativistically, maybe relativistically, that's probably one year in real time. When do you think there'll be the first billion-dollar company with one human employee?

2026.

Yeah, I absolutely buy that. Um, do you have any advice for people building with Claude, um, for the next year, how to think about building at that frontier as well?

Yeah, um, I, you know, I think there's like a lot of very specific things you could say about like how about how to use the models, but I feel like because of this whole like relativistic time dilation thing, this like speeding things up, like almost all the advice is drowned out by like one sentence, which is, or maybe two words, which is just be ambitious. Like build something that's greater than you think is is possible. And even if it doesn't quite work yet, another model will come out in the next generation, which right now is three months, but like probably it's going to go down to two months, then one month. And, you know, then, then if I want to come up this year, maybe I'll be giving advice that's like, oh, you know, don't build anything today. You know, we're releasing something today, but by tonight, it'll be, you know, you won't want to be building with this tonight.

I talked to a founder who started a company two years ago in the sort of autonomous AI coding agent space, and he basically tried every single model, and his startup wasn't working. And then it was actually 3.5 where he's like, my startup works now. And it was the same thing of like, this thing that I was trying that was really hard, all of a sudden is now, um, possible. But hitting your head against the wall actually sometimes can be useful because you put all the other pieces in place, and and everything works except the model. And then when the model works, it's almost like you've built something that's like more robust than it needs to. And that can be like a positive property. Um, so, so, you know, as much as I joke about like, oh, you should, you know, you can just wait for the next model, actually hitting your head against the wall, as long as it's something that's like almost possible, if it's not like, you know, like three years out from from what's possible, um, I think it can actually be productive. We saw that even with advanced research internally. Like our, our research and Claude skills team had built a prototype of this. The model kind of lost its way, it wasn't good at using tools. And then with 3.5, especially with Claude 4, I think you'll find that it does advanced research really, really well as well. And it's because we were trying and kind of failing along the way as well.

Yeah, it's, it's almost as if you want to run your, you want to run your startup as like speculative execution against the next model, right? There's some kind of like, I don't know. I love that. Yeah, I think that's exactly right.

All right, so last question to wrap up. Um, for many of us today, um, who aren't Dario, we couldn't have imagined the progress that AI has made and the rapid pace of change. What are you most excited about for the coming year and in the next five years?

Yeah, so, uh, I think for in the next coming year, uh, we are going to see incredible things in in in code. I would refer again to kind of the, you know, taking where we are with Cloud Code and where we are with the coding models and going from there to kind of to kind of the agent fleets. Um, I think this will have an interesting effect in the world, which is I don't know that we've thought carefully like from an economic or business perspective about what happens when the cost of producing software goes down. It's kind of an assumption, an article of faith, that you only make software if it's only worth it to make it if millions of people use it, or at least hundreds of thousands, or maybe tens of thousands. Like you wouldn't make, you, you, you, you like wouldn't make a whole piece of software for this event, right? Like you might throw together something. But like when it just becomes really cheap, when it costs you 20 cents to like, oh, let's just, let's just throw together something that, you know, you know, changes, you know, changes my vision for this particular event or something like that. Um, uh, I think the world is going to be very different when these things can be made ad hoc on a on a one-off basis in like a few seconds for for less than a for for for less than a dollar. What are, what is the role of the developer there? What is the role of businesses? What is the role of startups? Um, and what is, what is the experience of the, of the, you know, of the, of the people using it? I think we don't know the answer to any of those questions. So that's very interesting. On the, on the five-year timescale, I will return again to biology. I think the biomedical stuff will not be revolutionized in the next year because it's, it's kind of, you know, slow to, slow to happen. But, uh, yeah, yeah, I hope that, uh, five years from now, we will have, uh, vanquished, uh, many of the diseases that now, uh, that now exist.

I love. We'll leave it at that. Unfortunately, we do have to wrap up. I feel like we could talk for another 40 minutes. So first, I want to thank Dario for spending time with us today. Thank you, Dario. I also want to thank all of you who are here in person and those watching via livestream. Uh, but before we close, I almost forgot one thing. Um, as a special thank you to everyone who joined us today at Code with Claude in person, I'm excited to announce that each of you will receive free access to Max 20X, our highest tier plan, for three months. So look out for that. I especially love using Macs with Cloud Code, so you'll be able to do that as well. So we can't wait to see what you build. Have a great rest of your day. With the different, um, sessions and welcome again to Code with Claude. Thanks for coming. Thanks for coming, everyone.

[Music]

[Music]

[Music]

Oh, yeah. Oh.

[Music]

Oh, hey. Hey.

[Music]

[Music]

Hey. Hey. Hey.