📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Anthropic's Jared Kaplan on the Future of AI Agents l TechCrunch Sessions: AI

TechCrunch26:39

Transcription

All right. Well, we're going to get right into it. Uh, so a couple weeks ago, Anthropic had your code with Claude event. Uh, you launched Claude 4. Um, and you also said something that kind of flew under the radar about kind of pivoting away from chatbots a little bit, more towards these agentic AI coding systems. Can you talk to me about that decision and how agents and AI coding assistants might be more aligned with Anthropic's mission?

Yeah. So, I guess when we first developed Claude, um, back in 2022, the motivation for us—we were just a research company; we didn't have a product—was to build a dialogue agent because we figured a tremendous amount of sort of what we do can be encapsulated by talking to someone, asking them to do things, and and sort of seeing how that goes. And that obviously turned into Claude, which is uh which is a chatbot. You can go to claude.ai and and and talk to it.

Um, but I think the vision was always that as AI advances, gets better and better, um, the challenge would be to have AI do tasks that take longer and longer, are more sophisticated. And fundamentally, I think uh the way that we get things done isn't just by giving you the answer, but it's by using tools to go out in the world, learn the information you need to do the task, do the task, iterate. Like if you're a coder, you write code, you run the code, you see if it works. If you're like me, it doesn't work the first time, and you have bugs. You you you see what those bugs are, though, and you fix them. And so I think that uh just the natural growth in AI capabilities and sort of AI utility comes from AI agents that that can do do more and more for you. And so I feel like a a paradigm where you just ask a question, get an answer, and you're done. Like it just feels very limiting. And so it's it just isn't kind of the focus that we see going forward, right?

But I I think a lot of your competitors like OpenAI and Google are kind of in this race to get this massive AI chatbot platform, like you know, to where people will come for their first stop on the internet, kind of to replace Google or maybe even social media. Some of them are kind of thinking very big picture with these consumer apps, and it doesn't feel like that's where Anthropic is competing.

Yeah, I would say that uh to a large extent we uh were trying to help uh businesses and individuals sort of integrate AI and use AI in uh the most useful ways possible. Um, I think that uh having a static platform like that does feel a little bit uh a little bit limiting. Um, that said, I don't think we know where AI is headed. So I think giving people giving developers tools to sort of integrate AI, use it in a lot of different ways, uh feels like the way forward. I think Claude code is is I think a cool example of that. We first built that as an internal tool that engineers and researchers at Anthropic were using uh to use Claude to to make themselves more productive. Then we sort of shipped it externally. And I think obviously the main use is developers, but I even see people sometimes using Claude code um to do things that are that are surprising, like just like almost like an operating system or something. And so I think uh we're all experimenting with what the best way to use AI is. And I think we're just going to keep experimenting and hopefully keep empowering developers to experiment uh and and build with Claude.

And on the topic of Claude code, I mean you guys have really emerged as a leader in the AI coding space. Your AI models uh have really shot apps like Cursor uh to, you know, great fame. Uh, you know, they they've really benefited from using your models. But now with Claude code, it feels like there's some tension between Anthropic and these AI coding applications. Are you competing with these companies that are powered by your models?

I think the way that I think about uh a lot of our business is, I mean, a lot of it's built on the API. We really want our customers to be as successful as possible. I think that um one reason why we, as I said, like we were playing with Claude code internally. One reason we decided, oh, we'll we'll we'll ship this, we'll make this public, was that it felt like it was at least a demonstration of uh how far you could push agentic capabilities. And that's something that we were we were experimenting with and leaning into on the research side. Um, I think we really our goal isn't to to compete with our customers. Um, uh our our goal is just to sort of make sure that people are exploring uh the the possibilities with AI. Um, and I always encourage people to sort of experiment with building applications, uses of AI that don't quite work because, uh, I mean, we found it at Anthropic because we expected AI to keep getting better and better and better. Um, we thought there were safety questions, uh, concerns around that, and that's why we're so focused on safety. Um, but uh because of that trajectory, because AI is getting better so quickly, I think a lot of things that uh don't quite work with the model right now are going to work in 3 months or 6 months, and it's a way of really kind of getting ahead of the curve.

Well, one of your customers, uh, Windsurf, was not happy with you this week uh because they said that Anthropic pulled uh back some of the direct access that they had to their models. Windsurf is a like Cursor competitor, but it's much smaller. uh, they were also reportedly acquired by OpenAI. What's the logic behind kind of pulling that direct access from Windsurf?

Yeah, so uh my understanding is that with Windsurf you can actually bring your own API key and and and and continue to use Claude, right? But it's much more expensive and like complicated. That's my understanding of that.

Yeah, it might be it might be more complicated. I mean, I think uh as you alluded to, I mean, I think that we were really uh we we've really been quite constrained in terms of supply. We're we're hoping to greatly increase the sort of availability of of tokens. We want to keep them flowing um for Claude over the next couple of months, but we really just are trying to uh uh enable our customers who are kind of going to sustainably be working with us in the future.

You don't think they're going to be sustainably working with you in the future?

Well, I mean, you you said it, not me. Like, uh I I don't know. I mean, I I I think it's it would be it would be odd for us to be sort of selling Claude to to to OpenAI. That would be weird. Um, but I I think uh you know that deal hasn't been announced yet, and I think in the short term I think you know what a lot of startups and kind of developers saw was that Anthropic can kind of just cut off access, you know, and because maybe you know your startup isn't competing with Anthropic today but maybe in the future you will, um it feels like AI coding is one of those spaces where you guys are investing a lot, so I think about your relationship with Cursor, which is a much bigger business and really relies on Anthropic, but they're building their own models now. So, how do you look at that relationship? I mean, it feels like that is something that might come to a head in the future, too.

I guess I I don't expect it to. Um, I mean, obviously, I I mean, there are many many many companies training their own models for all sorts of different purposes. Um, I I expect we'll be we'll be working with Cursor for a long time. I I I hope to be um they've been a great customer, great partner, um great tester of of of of new Claude models. So, uh yeah, I I I really think that if you look at us versus a lot of other AI companies, I mean, we're quite invested in our API business, we want people to be able to build with and on on top of Claude, and we're definitely not trying to limit that. We're not trying to compete with our customers. We're trying to empower them.

Got it. Um, and I'm really curious about um, kind of you were a really influential person uh, in the scaling laws paper. You were a key author on a key scaling law paper. Um, and I think there's been a lot of questions in the last year about the lasting power of the scaling laws, of how much compute and data you can throw at these models and keep getting gains out of them. Uh, or how much we're going to rely on reasoning. Where do you see kind of the value of, you know, continuing to add compute and data to AI models today?

Yeah, so I think scaling is continuing to pay dividends on both the sort of pre-training and RL side. Um, we've seen a lot of advances in the last year or two on uh on RL. Um, I think I mean maybe people here are all experts, but I mean the way we train AI models is that we do pre-training where we teach models to imitate um and understand predict the next word in human written text and and other data. Um, and then we fine-tune them we with RL both with human feedback and with uh uh AI being trained to do things like write code that passes tests. Um, the I think both of these show clean scaling laws. Um, uh the the work that I did and and many others contributed to uh maybe five or six years ago now was on the pre-training side and showed really really clear empirical trends where if you make AI models bigger, give them more data, more compute, then they they get better. Um, I think that's continuing. Now there's a question of is there enough compute? Is there enough data? Um, I think for now there is, but I do think that eventually um uh th those will be those will be constraints, certainly data. Um, on the RL side you can go a lot further, and I think RL is much less fully sort of even close to fully fully tapped, and on the RL side you're really training models with reinforcement learning to reinforce models doing useful safe stuff and not doing things that not making mistakes, and so very obviously that's the kind of thing that you're going to want to keep scaling because you want to make AI more useful um not just uh better at better at autocomplete. And so I think that there's uh there's a lot of evidence from the last five or six years that that you can get clean kind of log linear scaling with with compute in RL. And so I think that's a major source of investment as well as basically using new techniques like advancements on constitutional AI um and more complex agentic environments to train AI to to get better and better.

Yeah. I mean, I think I've seen with the pre-training scaling, like the bigger models that have been fed more data, I've seen that, you know, they're better at creative writing and they're better at like some of these kind of uh hard to measure tasks, but like on the core benchmarks, like they are not as like stellar as some of the reasoning models. And I'm curious um from your perspective like do you think that, you know, by scaling reasoning that is a reliable path to get to AGI? It feels like that's much better at getting to specific like tasks and skills.

Yeah. I I guess uh the way that I think about it is that at least for hard problems, I think that reasoning is obviously a great way to get more capability out of the model. Um, you give it more compute at test time and you can get a better response. So I definitely think that that's going to be uh an ingredient. I think that in some ways like agentic capabilities might be even more important. I mean, maybe they're just they're just both important. Um, in the sense that I think that when we do tasks, when we do work, there isn't a lot of the time that we step back and we like think for an hour about what to do next. I think a lot of what we do is we experiment. We try something. We see what our environment tells us like did we make a mistake? Um, did we write code that passes tests? Does does like the paragraph that we wrote make sense or can we make it better? Um, so I I I tend to think of uh scaling sort of agency, search, um, tool use as as like really important for that as well. But but definitely I think that reasoning can allow us to take models that aren't necessarily at the frontier as pre-trained models, but but make them better um, at uh, at at hard problems, solving hard problems. I mean, part of the problem with the pre-training scaling is that like we're just getting so all all of the AI labs are getting so compute restrained right now. It's hard to kind of, you know, level up by another factor of 10. How compute restrained are you guys today? I mean, going back to what you said with Windsurf, it's like you guys don't have unlimited compute.

So, I think it's scaling very rapidly, right? I mean, I think AI is on and has been on an exponential for a while. And I think that's that's continuing. And obviously you see that in the news with sort of the the fundraises in AI going up and up and up. And that's because uh the value unlocked by AI. As AI models get smarter is is scaling with it. There are more and more use cases. Customers can use AI to integrate AI more and more. Um, that means that the feedback loop is even faster. So, uh, generally speaking, I mean, we're always like every year increasing our supply of compute by a significant integer multiple. Um, uh, that's that's just continuing and continuing. Um, we're we've just basically started to unlock the capacity on our new Trainium 2 cluster, which is really really big um, and and continues to scale. Um so that's what's going to allow us to unlock more compute for customers and for sort of training the the next generation of like Claude 5.

Yeah. And I mean you mentioned Trainium 2. Uh, so I I want to talk about some of your partners. Uh, you know, Amazon is a big one. Um, like I want to you know think talk to you about how you're thinking about uh Alexa and and kind of the ways Claude might appear in kind of uh interfaces that are not your own. Like it it seems like Claude is powering some parts of Alexa plus their new Alexa. Is that right?

Yeah. I mean, it was it was sort of announced at at a at a joint event that that Claude is contributing to uh to to a lot of Amazon products and including Alexa. Um, the uh generally I mean as we were discussing earlier like we're excited about Claude being able to power all kinds of products from from from many different many different companies and developers.

Yeah. Uh, there's been some rumors that Apple might work with you guys. Do you think that'll happen at WWDC next week?

I mean, I can't comment on uh on on ongoing work, but but yeah, I mean, we're we're always I mean, we're always open to to working with a lot of different companies, a lot of different partners. Um, we work with Google as well. Google's great, too.

Apple, though, you know, I really don't like talking to Siri, and I do like talking to Claude. I would love to talk to Claude, and I'm sure I'm not the first person who's raised this to you. Um, but I mean I'm curious like when you look at, you know, working with people like Apple, like Amazon, you know, like Google, like what are the troubles with kind of integrating Claude into their voice assistants and and why haven't we seen that already?

Uh, it's a good question. I think it probably depends uh on all sorts of different factors. Um, I mean, I think for one thing, obviously if you're if you're a startup you can just move extremely quickly and ship and iterate, but I mean if you're a giant company with millions and millions of users then then then tend to move a bit more slowly. I think a lot of uh voice assistants um just just speaking generally um aren't just voice assistants, not such something you're chatting with. It needs to be integrated into some larger ecosystem of tools and other applications. You want to be able to like call something that accesses your calendar. Maybe you want to be able to turn the lights on and off in your house if it's like a physical assistant. And so getting really really high reliability across tens, hundreds, thousands of different tools and applications that are integrated into these assistants I think is complicated. It requires work and uh and the expectations are very high. You don't want to disappoint your customers. So I think I think those kinds of integrations are complex, but it goes back to the importance of tool use and and agency. Um, behind the scenes a voice assistant maybe needs to be an agent. Um, and maybe that's another reason to focus on those kinds of capabilities.

I I think so. Uh, I wanted to ask you about something that, uh, your CEO, Dario Amodei, uh, he wrote an op-ed in the New York Times that ran this morning. The title was, uh, "Don't let AI companies off the hook," which is funny enough, we were also considering naming the panel this. Um, but, uh, no, I I wanted to ask you about it because he argued against Trump's 10-year moratorium on uh, states regulating AI. Why are you guys picking a fight with the Trump administration on this piece?

I mean, my understanding, if you if you read the op-ed, the op-ed uh in in many ways is uh is saying positive things about a lot of the work that the administration has done in trying to make AI safer. Um, prevent, say, authoritarian governments from getting access to AI. Um, I think that uh this is really uh this is really sort of just a a question of uh states' rights. I think that like um if one state wants to say regulate self-driving cars, like I think it makes sense for them to be able to do that. Maybe maybe a self-driving car company tests their car in Arizona um and then wants to deploy in Michigan. Maybe Michigan is worried that like that car doesn't know how to drive very well in the snow. like that kind of that kind of use case. I think I'm not an expert. I'm not a policy person, but I think that that kind of that kind of thing could could be blocked. And so I think that uh uh many many folks um I don't think Dario is the most high-profile tech CEO to to criticize uh uh the legislation, but uh but yeah, I think I think that we we really want to be able for governments to be able to move quickly to respond to AI. We think AI is changing very rapidly. Um, we think states are one of our nation's laboratories for experimenting with different possible forms of engagement with industry with regulation, and so we think that in order to sort of enable government to act with the speed that it needs to given the speed that AI is moving um we we think flexibility is is great.

Yeah. I mean, one of the things he talked about is the need for transparency uh among AI model providers, and you guys call out, you know, OpenAI and Google and yourselves as, you know, having some transparent acts. Um, what kind of transparency standards do you think should be put in place for AI labs?

Yeah. So uh we have a responsible scaling policy that says that we want to think carefully about what kinds of risks there might be from AI in the future. Um, we want to evaluate pragmatically whether AI models really pose those risks or whether uh whether it's it's too soon, and then we want to sort of put in place various mitigations. Um, I think something that you could ask of AI companies is just to transparently discuss what risks they're considering, what evaluations they've done, what kind of safety testing they've done before deploying their models, and what kind of mitigations they have in place. And at least that would uh as kind of a bare minimum would sort of invite the public, regulators, academics, other AI developers to sort of understand uh what the risks are, what evaluations have happened, and and and how robust the mitigations are. So I think something like—I'm not a policy person. I'm not a legislator. I don't I don't have a specific policy proposal uh in mind, but I think that at least that kind of transparency seems like a good a good first step. Um, and as you mentioned, I mean, Google and and OpenAI are also have similar kinds of policies, but I think it would just be

Good sense, if that was something that that all advanced frontier AI developers, uh, uh, were doing. And and I think it would be reasonable for government to sort of ask us to do that and not let us off the hook.

Are you worried that about, you know, retaliation from the Trump administration? I mean, you know, they're an administration that notoriously, uh, can has made things difficult for people who disagree with them. I mean, we're we're trying to work closely with with all, I mean, state governments, the federal government, etc. I mean, I'm not on the policy team. I I I don't have any any detail, but I mean, we're generally engaged in sort of helping them to, uh, I mean, advanced American interests when it comes to AI.

Got it. Well, I mean, on the on the topic of transparency, uh, earlier this week, Reddit sued Anthropic. Uh, they claimed that you guys trained your AI models on Reddit's data and uh, you did not have uh, authorization to do so, is what they claim. Um, they say that in July 2024, Anthropic said you guys would stop training on Reddit and then continue to scrape their site over a 100,000 times. Is that true?

I can't comment on uh, new ongoing uh, litigation. Um, I can say that Anthropic puts a lot of effort into carefully uh, uh, obeying sort of requests in robots.txt and and kind of industry standard practices, right? But I mean, I think at at large there's a lot of publishers have kind of brought similar claims against uh, you know, Anthropic and OpenAI and Google and every company and just said, you know, like you guys are training these models and making billions of dollars. I think you guys are valued at $60 billion or more than that uh, off of you know, these free kind of works from the internet and and I think Reddit is kind of making its own case that that is uh, not okay and you have to pay them some money for that.

I mean, do you think that generally speaking, like, you know, people should be compensated for their you know, work when you when you train AI on it?

As I said, I think people can prevent uh, clawed and and and and Anthropic from accessing scraping their data just via robots.txt. We we respect that. I think that uh, more generally though, I think AI training is fair use. um, it's not it's not copying, it's not reproducing uh, data and uh, I think that's sort of the basis of of legally why we we think it's reasonable to train AI models um, on publicly available data that uh, uh, on the web that uh, has not that developers have not requested we uh, we avoid via robot.txt TXT, right?

But how can it be fair use if at the same time like like Reddit has like 60-70 million deals with Google and OpenAI, like the New York Times has deals with, you know, Amazon, like like like if people are paying for the right to train this content and then at the same time companies are taking it for free. How does that work? I mean, that's that seems like a…

Well, as I said, we we obey robots.txt. So if we uh, that's that's that's that's our policy. So if we've been asked not to not to scrape data, then that's uh, that's what we do.

Got it. I I mean, one thing from that lawsuit and I'll kind of move off of it, but I mean, Reddit kind of came off of they came off kind of saying that like Anthropic has always been this kind of, you know, cautious cousin of the AI industry, so so to speak. Um, but they kind of argue that that's not exactly true. And and I think that with the launch of Claude 4, a lot of people felt that, you know, you guys released a model that had a lot of safety concerns. There was some issues where Claude 4 in a, you know, scenario, one scenario, it blackmailed, uh, you know, developer who tried to take it offline and replace it with another model. Um, what do you say to people who think that, you know, Anthropic is not being as cautious as it used to be?

Um, I I I don't think that anything has changed. And I think that this goes back to what you said about transparency. Our I mean, you asked like what what would we like to see? We're so the the the scenario that you mentioned was not a thing that happened in the world. It was like we're doing hundreds thousands tens of thousands of examples of red teaming and our team actually calls it weird teaming where we try to find all sorts of weird scenarios where our model might do something problematic so that we can flag it and so that we can improve it in the future. And then we wrote a 120-page system card where we documented all kinds of these examples in order to encourage others to be transparent to sort of show what their models are doing and encourage others to to to do their own testing. So I think that uh, we're going to continue to do that. But I think that like in being transparent, I think uh, that opens us up to more criticism. um, we do a lot of other kinds of safety testing like with respect to our responsible scaling policy as well. We have mitigations in place there. So that may open us up to more criticism, but I think we're going to keep doing that because we think people should understand how AI models are be are behaving and they should be encouraged to sort of do their own tests on on other providers as well.

Just the last question here. I mean, just, you know, Anthropic's been up and going for a while now, a few years, and I'm just curious if, you know, you started this to be kind of, you know, put an eye on safety in AI. Uh, and have you seen things, you know, from your competitors out in the wild that have kind of reaffirmed your mission for why you started Anthropic? You said, "Yes, this is unsafe. This is why we need to exist."

Yeah. I mean, I think AI is getting more and more capable, and so I think safety risks are are becoming more and more salient, more and more real. Um, I think that I'm generally proud of us being transparent, of us being kind of the first to have something like a responsible scaling policy. I think a lot of our competitors want to do the right thing, but I think if if if we do we do something like that first, then uh, it sets a standard. It's it's something that uh, that that others can can experiment with, iterate on, imitate. Um, and I think we'll keep doing that. I'm happy that that we've done it so far.

All right, that's our time. Jared, thank you so much.

Cool. Thank you. Thank you. All right.