Transcription
Anthropic is not just a massive model builder; it's a massive product builder as well, with products like Claude Code and Co-work that have taken off like crazy over the past few months.
And Claude Code came out of a place that few of us know much about, called Anthropic Labs. And Anthropic Labs is an organization within Anthropic that is working on building the next level of frontier products with AI at the center. And so today, we are lucky to hear from the person running that lab, Mike Krieger, who is the co-founder of Instagram and now the lead of Anthropic Labs at Anthropic. We're going to welcome him on stage along with Lauren Goode of Wired, who will join me as a co-interview. Mike and Lauren, let's, let's, let's hear from both of you guys.
Hey, let me see. You're over here.
All right. All right, Mike. So, chill times in Anthropic land?
Nothing going on.
Slow week.
Slow week. You want to start, Lauren?
Yeah. Uh, first of all, I mean, we want to talk about what you're working on at Labs, uh, and explain your role to folks. But I want to ask you first, how close are you right now to the situation with the White House?
Less than in my CPO role. So I transitioned about, like, five months ago into this Labs role. I think in the CPO role, I think I would have been deep in it. Now I'm more, you know, obviously, you know, we want to restore it and and as a, like, product person, I want to make sure that that gets access to, but less close to it than, uh, you know, in that sort of C-level role that I had before.
Okay. Alex, do you have a follow-up? I have like eight follow-ups.
Definitely. Well, we, there was a guy named Ben on X who said, "I will mail Anthropic an original copy of my long-form birth certificate if they will enable Fable for me again." I sound like those lunatics who were obsessed with 4o now. Will you take Ben's long-form?
I don't know, uh, that we'll take his long-form, but it has been interesting. I mean, Fable was only available, um, for a few days, but I definitely, uh, every time I've tweeted since then, uh, they've not read whatever I was tweeting and they've mostly been like, "Bring back Fable." Which, like, in Instagram, where you got rid of Gotham, do you remember Gotham, the filter? This is like, and then for the rest of, like, the next eight years, all I heard was bring back Gotham. So, it, uh, it struck a nerve. But, Fable will, uh, will come back before Gotham did. But, yeah, it's, it's clearly the folks that had gotten to use it and started incorporating it. It's actually really interesting. Um, I've learned to not really trust day-of or even week-of model reactions. You don't really know until you've put it through its paces. And so, like, I almost just completely block out the noise in the first couple of days of any new model release, 'cause I don't know, everybody has maybe their, like, toy, like, example thing that they like to do with the new model, but it's hard to actually put it through its paces until you've actually had real work done with it. And I think people were just starting to do that and then, you know, we had to, uh, sort of pull back Fable. But, um, I remember in December when we put out Opus 46. Uh, it was like this interesting time where everybody went home for the holidays and a lot of people had that week off between Christmas and New Year's. And then they came back and were like, "Oh, I spent a lot of time and I really get why Opus is good and I'm going to do it." So, I don't think Fable has had that opportunity yet.
But, despite that though, I mean, this was a pretty big reaction from the Trump administration, and I think everyone here, especially if you were listening to the Alex Stamos, uh, session, uh, understands what's going on. But, this happened within a few days of the model release. Um, if I can give Wired some credit, Wired just reported last night that it was due to Anthropic's, uh, relationship with or having given, uh, access to the model to SK Telecom that could have raised flags within the administration. How surprised were you by how immediate that that backlash was?
Yeah, I think the, the, the sort of reaction decision was surprising and and when we were sort of immediately engaged with them to, you know, restore access as well. Um, and so, at the same time, you know, with the one thing that we, we think a lot about internally is, you know, there used to be a poster on the Facebook wall when we were still there that was like, "Every day feels like a week." And I think that's becoming true in AI, and I think a good thing to remind ourselves of in general in the industry is, uh, we're dealing with unprecedented times. We're dealing with new situations, and, uh, they can develop really quickly as well. And so, I think, uh, also developing the capabilities and the connections and and to make sure those conversations can happen quickly is really, really important.
I have a new motto suggestion for you for the wall. Move that, move fast and jailbreak things.
Move fast and jailbreak [laughter] things. I don't think they're going to use that, Lauren.
Okay.
Um, you know, we just heard from Alex that some of these capabilities have been, uh, available on, you know, previous models or, you know, you can, you can find bugs with other, with other models that are available today. So, why do you think Anthropic got singled out on this front?
Um, I don't know why, like Fable specifically was singled out, um, as, you know, again, Fable being the non-cyber intention, uh, model as well. Um, I think the, the thing that does change over time is there's, uh, you know, the capabilities if you knew what to prompt and knew what you're looking for, and then there's the capabilities I think, you know, uh, I'm not sure I resonate with the, like, uh, juicing high school athletes metaphor from Alex, but, like, you know, the, the uplift that you get, like, uplift is a thing that we think about a lot when we think about model, um, safety. So, if you look at our model cards, for example, one of the ways that we look at, uh, risk in the bio domain is comparing the uplift, uh, from sort of lay, lay person using the model versus, you know, an expert or lay person just using the internet and and seeing what the comparison there is. And so, that is one trend that is, you know, has been, you know, progressing as the models get more capable. Um, and so, it may be less why Fable is singled out and maybe more like what the overall trajectory is.
Interesting. So, we just had Our Karaziyan, the, uh, lead economist on Ramp on the podcast. And one of the interesting things that he said was, you know, we, we had during the Pentagon situation, there were these headlines, okay, this company will not use Anthropic anymore. But actually, the data from Ramp shows that spending actually increased to Anthropic models. It was apparently a good publicity moment for the company. So, I, I think it sort of plays into a debate that we have like here on the show about like how much of this is like real concern for the issues and how much of it is, you know, from Anthropic is marketing. And like, we have somebody from Anthropic here who can actually shed some light on that. We just said Themos talking a little bit about it from his perspective on the security side, but we're lucky to have you here today. So, so is it material or is it marketing or some combination?
I mean, I think it was one of the hardest things to like, really deeply believe that something is true and not marketing and then due to not even just Anthropic. I think like, you know, uh, uh, I think people are generally right to be skeptical of any company saying anything and you should like put it through your own filters as well. But for me personally, I'm always like, "No, but it's real and like, we like both deeply care about safety and are like, to the extent that we are being vocal about anything, it is to either sort of help paint the picture of what very likely is coming or we believe is coming or what we've already seen and spotted. You know, for example, in the, in the Mythos case, really just looking at like vulnerability scanning and bug finding and doing it in partnership with with companies that were in that kind of Project Glasswing initial announcement. And so, um, the technology is also not in the, like, like in the just, you know, it is, it is doing really incredible things and therefore even calling out what we see as what is happening, I think can seem hypey. I wish I could press a button and was and like, I could make everybody believe that we are not being hypey. I realize that's not the reality that we operate in, but at least from my perspective, we try to call it like it is.
My understanding is that within Labs in particular, in research and development at Anthropic, that you're using the, the AI models to actually build new products, to prototype new products, and sort of test out your thesis. So, is this, is this ban essentially now limiting your ability to do that within Labs?
Yeah, I mean, definitely Fable, uh, is the best model I've ever used, and, uh, it's not to say that, like, work has stopped, but it's definitely, like, less good than the other models that, uh, the, or sorry, or we were using models that are less good than that, um, as well. And it is also, I mean, maybe the reverse of the, uh, distrusting the first week, uh, you know, response is what happens when you don't have Fable. And like, obviously the Twitter reaction for the people that had already kind of like gotten, uh, into the model and and raising it was strong. Um, but I'd say even in my personal use, I'm like, "Oh, like, I'm on Opus 48." And it's good. Like, I'm still productive. I'm doing work. But, um, and we can go into sort of, like, how my work changed with sort of these, like, Fable or, like, you know, uh, the models of that sort of, uh, family, but it is noticeable for sure.
Yeah, I think we'd like to know that. I mean, you know, we, the public had access to Fable for, like, a half a minute. Um, groups in this, like, Project Glasswing have had access to Mythos. Um, we don't really know what the difference is between using a model that you can use today is and using one of these, you know, super Anthropic models. So, what actually could you do differently with a Fable or a Mythos?
I think for me, it's the sort of scope and scale of delegation. Um, and it, all these things are really imperfect. Like, you know, people say like, "Oh, is this now a level five software engineer or a level six?" But anybody who's used these models extensively knows that they're still spiky in capabilities, right? In some ways, they're, you know, in many ways, they're better engineers than me. Um, and in other ways, like, you know, I was complaining today that it had missed a descender, like, a the G that bottom part is called the descender. I'm like, "How did you put that in the UI?" And it's clipped. And of course, like, there's vision capabilities that need to improve, and there's, uh, debugging capabilities, and there's sometimes even just sort of human common sense that is, you know, more way better at than when the models are. But overall, I think the big shift for me working, and it was really interesting because it sort of coincided with me going back into a builder role. So I really got to see going from, you know, using these models as an executive, you're trying to do the most of them, but you know, like, not going to have it write all your email and, you know, you, I think strategy still needs to come from you, and then you can use the models to sort of test it. But going back into a builder role and going from, okay, I am delegating chunks, like, please fix this bug, or I'm thinking of implementing this feature, let's go back and forth, to something that ends up being much more, sort of, all right, like, I got this bug report from one of our users, or I have this notion of something that I want to build, like, can you sketch out two or three ways in which we could do it? All right, like, that seems plausible. Often, I find actually sometimes it'll give me the sort of explanation or proposal and be like, "Okay, that actually is over my head. Like, you are clearly way smarter than me. Explain it to me like I'm not five, but at least, you know, not you." And it'll sometimes, you know, explain it that way, but then go build it and gets it right. Getting it right, like, you know, at a very, very, very high rate. And I think that starts really changing how you operate. Like, I, I moved much more to before going to bed making sure that I had like queued up for Fable, like enough chunky work to last, I would call it the whole night, and I would check in later and it got it done in an hour and it was like the guest hanging out for the next seven hours, but like really like delegating, like much more of a goal than just a
So one, one task for example.
Yeah.
Like, give us an example of one task you would hand to it.
I mean, here's a kind of crazy one, which is for the programmers in the audience, like, I, I'd written one of our Labs projects in in Python. It's like the language I know well. Instagram was all Python. And for like some not super exciting reasons, we actually needed it to be in TypeScript to deploy it. And I was like, all right, that's going to be like, you know, at Instagram, we for years talked about moving from Python to PHP or Hack or the, the Facebook language after the acquisition. And then, at least when I was there, never did. Um, but I basically, like, we have a feature called dynamic workflows where you can have it like also break down the task into like a lot of subtasks. And I trusted it to sort of not just do the individual action, but here's like a whole language conversion of millions of, about hundreds of thousands of lines of code at that point, uh, go off and do it, go plan it, go execute it, go verify the work, double verify the work, and then I came back to the work being complete. So, that level of like, this is a big sort of chunky, uh, task.
Hm. So, that was, so you're basically saying it was faster. Did it in an hour? You're, you're sort of guessing compared to with Fable compared to what it would have been before.
I think the main difference is in the past, you would, it would be like, "Great, I did it." And you'd be like, "Well, did you? You kind of took a shortcut here, or this is not quite right, or I need to go verify it, or like, 'Oh, you cut this corner.'" Or
It's like the managing interns thing that everyone's been saying for the past year.
Yeah, it's, it's very
Offensive to interns, by the way, but yes.
Yeah, exactly. I don't know if you managed an intern.
Yeah, that's [laughter] true. And it was, and you're saying it was more correct. It was, so it's faster, it's more accurate, uh, more reliable, and then, according to the US administration, dangerous.
I think the other piece is it has like a, a greater theory of mind is the wrong word, but sort of like theory of project, so that it's less, you know, "Oh, I'm going to make this change." And it'll say, "Great, I'll make this change." But really, especially if you've done like software engineering at scale, the best engineers kind of keep in mind all the disparate parts of how this thing is, and they also see around the corners. Like, "I can make this change, but if I don't do it in this way, then the next change is going to be incrementally harder." And I think that's like been a significant difference I've seen in that kind of class of models.
Okay, cool.
So, I think when we talked about Anthropic Labs, right? People think of Claude Code because it is really your breakout product and, uh, and it sounds like you've been tasked with basically figuring out what the next Claude Code is. Would you say that's an accurate description of what you're doing at Labs? And also, why does Anthropic need Labs?
Labs, yeah. It's a, they're, it's also maybe worth thinking about why we needed Labs in 2024 when I arrived and why we need Labs today, 'cause I think that the answer kind of shifts. Uh, I started the Labs, the original Labs team with Ben Mann, who's one of the co-founders of Anthropic, um, in my third week at Anthropic, and it'd been something that had been bubbling under. And at the time, the reason was really different. It was all of our product engineering with team was 25 people, and we didn't have the models really. Like, the, we had Claude that when I joined, it was Sonnet 3, like, you, and Opus 3, like, those were for the time good models, but you weren't going to, they were not even interns, right? They weren't even IC3 engineers. Um, so if you have a team of only 25 or 30 engineers, they're working on like the next, you know, incremental thing. And we were feeling like the models are starting to get better, but we don't have any products that sort of show that off. Like, a good litmus test for me is when we get ready to release a model, do we have either a product or a demo or or, you know, some other illustration of something that is very different? And it gets harder over time. Like with with Fable, you know, even like illustrating like that weekend task or, you know, that this longer amount of work. So, really Labs at the time was, let's make sure we don't, like, our products don't fall behind the model exponential that's happening. Um, and, yeah, so Claude Code came out of that initial one because nobody in the rest of product work, people were thinking about coding, but nobody was sort of had the like space to go and think about, well, what if we totally change the form factor and we embrace the fact that the models were going to evolve in this way? And a lot of the, like, the two most useful thought exercises we do in Labs, one is like, visualize the gap between what the models can do today and how most people use it, and can we close that gap? So, that's one. And the other one is, imagine what the models are bad at now that they're actually going to be really good at in six months, and let's make sure we have a product ready for that by then. I think those are like the two guiding, uh, questions for for Labs. Um, and then al-, out of that first incarnation came computer use. Computer use was different though, because when we built it, it was really bad. Like, we tried a bunch of products with it, and this was around, you know, Sonnet 3.5, and you'd be like, "Claude, can you help me, you know, clean up my desktop?" And it would like, click the thing, and it would delete the file. You're like, "This is not, not safe for release. We're definitely not going to go and and and build this, uh, or to ship this." But we had that product so that every new model that we, we'd release, we'd first check it internally and say, "Did computer use get better?" And we'd tell the research team how it had gotten better or worse until the moment where we said, "It's good enough. We're actually going to put a product out around this." So it's also gives you this sort of, uh, sort of beacon into the future that then you can kind of measure your, your future products against. Um, but then compare it to now. So we have a driving product team. There's Co-work. There's, um, you know, Claude Code has grown a lot. We have our platform. Um, and now I think it's actually much, much less about, none of these product teams are doing this sort of thinking, and I think it's much more that the models are advancing really quickly, and even our capability to interact with them needs to evolve. So, one of the things we collaborated with, uh, Labs and Claude Code that we shipped today, um, is Claude Code Artifacts. So, having Claude Code not just be able to type back to you, but also sort of draw a picture or give you an illustration. And that partially came from spending a lot of time in Labs saying, "Just a text box and a big text response is not going to cut it anymore." Like, when I mentioned that the models feel like they're way smarter than me when they talk to me, sometimes I'm like, "Can you draw me a picture 'cause this is what I actually need to fully, uh, fully understand this." But it's really what we've been thinking about is, you know, yes, we have a lot more products. You know, we have, we are actually have a lot of consolidation to do in our products. Like, that's another initiative that we have. But within that, we still have an opportunity to make things much more accessible to a person that does not spend all of their time thinking about prompting and the exponential and the difference between high, low, and medium effort. Like, there's a lot we can still do there.
But Mike, so there's a, it puts people using Anthropic models in an interesting place, right? Um, there, you know, Cursor, I think just sold for 60 billion to SpaceX. And someone put this meme on Twitter that like, you know, Cursor would have sold for 300 billion if it wasn't for this guy. And it's a picture of Boris Tcherny, the person who created Cloud Code. Um, and so for companies that are going to build on top of Anthropic technology, you know, they're going to wonder, do I want to partner with Anthropic, or is Anthropic going to go ahead and and build the product that I'm going to want to build, potentially even after partnering with them?
Yeah. And I mean, we'll take the like agentic coding side, and I think the broader sort of aspect of of, you know, being both a platform and a product, I think is really interesting. Um, when we take on projects, it's the goal is often to sort of push that area of the industry forward. So, um, you know, there were AI coding editors, and some of them were really good, and, you know, uh, but nobody was quite thinking about it in as sort of free-form a way as we got to think about it with Cloud Code. And now a lot more products have that flavor than I think would have otherwise. And so, I think if we're ever, you can call me out on this, Alex, if like, if we're ever entering an industry where like all you're doing is the same thing everybody else is doing, but like you've got the Anthropic brand, I feel like that's a bad use of our time and a bad use of our either Labs or product team time. Like, if we're going in somewhere, it should hopefully be to say, "All right, we think that the direction of travel is this way. We can build a product of that." And then by the way, there's no world, nor should there be a world where like all the products are Anthropic products. That'd be a bad world, right? So, like, that is hopefully either creating new space for for companies or sort of showing the way where other products can incorporate that too.
It would almost be like working for a tech company that has like social, messaging, video. It's all too. Right?
Yeah.
Mike? Yeah. Okay.
Uh.
I'd leave.
Well, there was some question for example when, you know, Anthropic launched a product that was seen as competitive to Figma, and you had been on the Figma board prior to that. And I, I, you stepped down. Is that correct? Yeah. And so, um, it's a good question that it's Alex has brought up, I think, where Silicon Valley is known for this really healthy, vibrant, risk-tolerant startup ecosystem. And when the big start coming in with tons of venture capital and, you know, a lot of resources, people say, "Well, wait, are they just, are they essentially just going to steal my idea?"
Yeah.
Now, I think, I think our dual existence, and it's something that other companies have to navigate. We talked, I'll talked about Amazon a lot in the previous panel, like they have have to navigate this world where they are both infrastructure provider. They obviously have a very large e-commerce thing. They do video, but they also serve video. And and then, by and large, customers can live in that dual world of like, "Okay, I'm using their infrastructure." Also knowing that they, they are also using their infrastructure to do that. And I think the, you can talk to our customers and see how well we're doing at this. Like, the thing I always try to do is like, at least approach it with a lot of transparency. So, the Cursor example is an interesting one where like Michael and I talked a lot over the, you know, time around here's where things are heading. Um, and, uh, you know, with similarly with with the other products that we think about, like, can we, I think it's a couple things. It's transparency, and then it's shared building blocks. Like, I think, um, in general, and I actually don't think there's any cases where this is even true. Like, we're trying to build on top of the same capabilities that are available elsewhere. The last time I was here in the Commonwealth Club on the stage was our healthcare day at the beginning of the year. And we didn't ship like Claude healthcare only, we have it, like nobody else has it. We shipped a bunch of like plugins and skills and MCPs and like complimentary abilities. So, that's how I'm not claiming it's easy or that it's a straightforward thing, but it is how we're trying to navigate what is like admittedly a complicated sort of situation.
Speaking of startups, Anthropic is still technically a startup. But you're worth a lot of money. I mean, what, what's the latest valuation? Is it
965.
965 billion dollars or something like that. So,
They sold Instagram for a billion, right?
Startup.
In 2010 money.
Right. Right. Financials have changed quite a bit since then. And yet, Anthropic has positioned itself. It is, you know, a PBC, and it's positioned itself as sort of a more ethical company around building AI. And I'm wondering if you could talk a little bit about how you see that positioning and Anthropic's role in particular changing the culture of the valley. Like, I know I think back to how Google in the beginning of the 2000s really changed the culture of Silicon Valley in so many ways. And how do you see Anthropic's culture now dictating this next era?
Yeah, that's a really interesting question. Maybe I'll start like inside, and I think there's an external component too. I think the reason I joined in the first place, so I was winding down my second startup and knew I wanted to go work at a frontier lab because I'd started to use these models for coding, and they were bad at coding, but I could see that they were as bad as they were ever going to be at coding. They were going to improve. And I had started building on top of these APIs. So the startup I was doing was called Artifact, and we did sort of AI-powered sort of news recommendations, and actually read a lot of big tech via Artifact back in the day. That's one of the things we added. So,
But not Wired.
Not Wired. You know, you guys had a really hard paywall, to be honest.
[laughter]
Fair enough.
We didn't do very, we didn't do a great job on that.
Need a discount on my subscription 'cause I can get one for you? Okay.
It's actually really funny, like the making deals.
Yeah, making deals. It's the, it's the login cookies. It's actually really hard to keep people like, you know, so
I know. I just, please escalate this to Conde Nast. I know.
[laughter]
Um, uh, and email login is very hard to do in an app. Um, but so, but I was building on top of the APIs and be like, "Wow, okay. They're, they're able to do really interesting things." Um, but what ultimately made me go to Anthropic was like, they walked the walk, and they really like deeply believe in trying to make AI go well for humanity. And that is like in, that's like in the water internally, and I think has been why I think the company has remained as cohesive as it has even as we've grown. And I think that it's like a testament also to the, the co-founders there on how often they are talking about this as well. It's a surprise for me coming from a world where at Instagram, we did a weekly all-hands, and we talked about product 95% of the time, and maybe 5% of the time we talked about something else that was going on in the world around the company. Probably maybe underselling our like go-to-market. Maybe it was like 80/20, but it was definitely very, very heavy product. And I remember on topic about six months in, myself and Kate Jensen, who's one of the the leaders in the sales organization, did a joint all-hands where we talked about our like, you know, how we're doing product and go-to-market together. And people were like, "This is so great. I finally understand our product strategy and like what we've been doing." It's like, "Oh, right. This is not quote-unquote a product company, you know. It is a very mission-driven AI company with like a very strong sense of like why it exists in the world." I think in terms of the overall impact on the valley, remains to be seen. I think positive signs that I've seen or interesting signs that I've seen that I've seen are like a renewed interest in philanthropy across the board, and I think that's something that has been written about, and I think it will be an interesting sort of outflow. Again, who knows how all of this goes, but like depending on how it goes, it could mean a lot of interesting new sort of philanthropic deployment. And then I think the other piece is, you know, uh, the conversation around how AI could or should go is one that is happening in real time with the technology versus retrospectively, which I think has been the case for other technology waves, and I think that is a good thing.
Hi everyone, Alex Kantrowitz here. I want to tell you about a documentary I've made with Gravity to explore the future of AI agent security.
To find out if we're truly ready for autonomous agents, I sat down with MIT Professor Ramesh Raskar, former White House CIO Teresa Payton, Michelin's Group Chief Data and AI Officer Ambika Rajgopal, and Sharon Guy, a former executive at Alibaba. They each offer unique insights into this evolving landscape. We conclude with Rory Blundell, CEO of Gravity, to discuss the path forward. With Gravity leading the way, join us on this journey. You can watch the full documentary at the link in the show notes.
Mike, you know, you talked a little bit about, um, Anthropic has this gap that it sees between the capabilities of the models and where everybody is building products. And with Labs, what you try to do is get ahead of that so you can show people what AI might be able to do now and six months from now. So, please tell us. Please tell us like what you're building, um, where you see the, where you see the potential.
Mhm.
And and what people should be on the lookout for.
Yeah, sort of roadmap.
If I can throw in a, and and what you're building now, but also if you had, if you have a pie-in-the-sky, like Elon Musk, data centers in space type ambition, I want to hear about that too. Tell us, tell us, tell us everything.
You have 13 minutes left. Go.
Exactly. The rest of the 13 minutes [laughter] is the monologue in my product team.
Um, I think maybe two themes I'm really excited about that we've been exploring a lot. Um, the first one is giving Claude an environment where it has more agency and it also has more self-knowledge. And I'm going to unpack that 'cause that's like a lot of AI words, but, um, uh, I'll give you an example of where we are currently doing a bad job of this. Like, if you are in a Claude project and you make a file with Claude, you're like, "That's great. Can you add it to our project?" Claude will be, "No. You have to go download the file and go drag and drop into this day." And you're like, "What?" Um, uh, until yesterday, I would have said the same thing about Claude design and Claude code, where you're in Claude code, you're like, "Cool. Like, I need a design for this thing that we're building." Let's Or you're in Claude design and you make a mock-up and you want to go build it, be like, "Cool. Here's a zip file." And you're like, "What?" You know, and I think, um, so that's a little bit of interoperability, but in general, this theme of giving, if you give Claude a lot of notion of its environment, I was talking to actually a customer, like an API customer, and one of the things that they were experimenting with is actually even giving Claude, like, a secure version of their source code while it's running in the agent loop in their product, so that if it hits an issue, it doesn't go like, "I don't know, I hit an issue." It can be like, "Well, it's probably this thing, you know, at least when it's talking to one of the sort of maintainers of of the software." And so that overall theme, and of course, you have to do it with safeguards and be really careful about what you unlock with it. It sounds kind of obvious, but it's actually night and day in terms of how expressive these products end up being able to be, right? And you can even see it going from like maybe like core chat or classic chat and Claude AI and something like Co-work, where it's got a little bit more agency and it has a runtime and it's able to sort of like understand a little bit about its environment, but I think we are at like 10% of the journey about where we, we could go. Actually, one of the reasons I think people got excited about things like open Claude is seeing how a, you know, harness that is modifiable and you can talk to it about things and you don't ever get the sense of like, "Oh, sorry, I can't do that. You're going to have to go to this like settings screen and turn it on." It's just a thing it has access to, and hopefully with like the right guarding and permissions. So that's like theme one that I'm like extremely excited about, and I think, I, if we do it right, should actually like transform all of our products like from head to toe. Um, the other piece is, and I'll maybe like share like the, if not the internal product we're working on, but like the, the phrase I got out of feedback was like, I think closing the gap. I mean, I talked about closing the gap between capabilities and reality. I think it's also closing the gap between how people understand their own work and then how the actual day-to-day is to do that work. I was talking to somebody internally who's on our privacy team, and to move a ticket from like one queue through another one via the task tracker into another one was like eight different steps of copy and pasting of like manually moving a pretty, you know, like, uh, kind of annoying to have to do. Probably error-prone. I have to like keep spot-checking it. And we helped her with one of our Labs projects to basically like make that not a pain. And she's like, "Ah, this is the first time in my career." And she's like, "Been working for 30 years. We're like, 'What's in my head and what I am using is like now this.' Like, it is now closed." Um, and I want to like bring that feeling to everybody who like, you know, of course, Claude unlocked a lot of, you know, non-technical people being able to code. But like, it's, we're still asking people to understand way too many concepts of like, what is, what is, you know, the difference between like my sandbox environment and production, or like connected MCP as myself or others, or how should I store data? And of course, you can't abstract everything, but like, if you combine both of those themes, if you give Claude a lot of self-knowledge and you're creating an environment where you can actually solve complex problems for people in like repeatable ways, um, I think I get like very, very excited about that.
And your moonshot?
Yeah.
Not letting you off the hook. What's your moonshot?
Moonshot?
Yeah.
Uh, no, nothing in space. Although I guess we're, you know, we're talking to SpaceX about spacey things.
Wait, you're talking to SpaceX?
I mean, there's an internal.
Absolutely. You know, for a compute. Yeah.
Um, is it, sorry.
Was it, was it exploring extra, extra orbital? What was the phrase? Something about exploring like
Exploring, yeah.
Post-orbital world things.
Definitely not my department. But yeah, there's there's stuff, stuff in.
Are you, so the Labs isn't working specifically with the team on compute?
Right, exactly. Those are separate, separate compute.
Totally separate. Okay. Okay. So your moonshot.
Yeah.
Do you personally believe in data centers in space?
[laughter]
I had a conversation. I, I'm far from a data center expert, but I talked to somebody who is a, like, a person who sends things to space, who's not Elon Musk. Um, and that's what, that's what you would say.
[laughter]
Um, uh, and they're really bullish. And I was like, trying to talk about why, and it was basically like, I, I guess effectively infinite, um, power if you convert it well, and, you know, infinite land. And I was like, "Okay, you buy that, and then, then I think they feel good about the shielding you have to do." Again, clearly not my area of expertise, but after that talk, I was like, "Okay, I see, I see it, you know, even if it's going to take a few years." At first, I admittedly thought it was a crazy idea in general, but now I'm like, "Oh, I actually really can understand why this might make sense."
When, when you were talking earlier about the ways that the work in Claude is going to get compressed and all those steps, I couldn't help but think of tokens and how you [clears throat] know, maybe it's good for your business model in the short term if people have to take so many steps and use so many tokens, but tokens have become this unit of economics that we're using to describe the industry now, and people are token maxing, and now they're tokenizing, and and and one, I want to see, I want to hear how, where you sit on that spectrum, if you're a token maxer. And two, is there a near future in which the industry is not actually measured by tokens, you know, it goes the way of MIPS or dial-up or some other, you know, there's some other unit of measurement that actually defines the economics of this era?
Yeah, I think both of those are really interesting questions. It was interesting earlier this year when you started hearing about like companies that have like dashboards showing like who used it the most, and we of course have like internal metrics as well, and we found that there's not a lot of correlation between like the person who's using the most tokens and like the person that I would like it was interesting thought actually, as maybe do it your companies like write down your 10 most productive people that you think are most productive, and then like get your top 10 token users and see how closely they correlate. At least for us, it wasn't that correlated. It seemed dangerous to sort of like, uh, sort of purely glorify the like maximum uses. Obviously, it's like very gameable, but even beyond that, I think it's, it's, you know, yes, you can ask Claude to do 10 different variants on something, but if you thought about it deeply, maybe you would you choose two that you thought were most promising and a third one if you then had like some iteration on that as well. So I would not say like a token maxer. Actually, the tokeniest thing was that conversion thing I did was just like a couple million tokens. It's like a lot of tokens that it took to to convert the, uh, thing from from Python to TypeScript. Um, but I think people are being more thoughtful about, um, these different pieces. And and one of the things we look at whenever we look at a model launch is not just model intelligence, but we're also really thinking about model intelligence and effort and token efficiency as that combination. And I think that's a big lever we have to improve is how do we continue to be more and more token efficient for a given task so that you can also hopefully, you don't have to think very hard about this. We can do this automatically, but like we're able to tune the solution to the problem a little bit more. Um, and then to your second question, yeah, I, you know, when I was still in the CPO seat, like I was thinking a lot about sort of outcome-based pricing as something that would be really interesting to do if you could do it. And of course, if you talk to, like, the Sierras and Finns of the world that have like a really clear, like, we kept this, you know, we were able to solve this customer request and not have it go escalated. Like, that's really clear. Um, it gets so much fuzzier on these like tasks that we actually ask Claude these days. Like, I had a strategy document. I used Claude to critique my strategy document. Like, what was the outcome? It's like, well, I don't know. It's like, tell me how the strategy goes six months from now. It feels like it's going to be very hard to to capture that as well. But, um, I would like to see some more experimentation around, can you better capture like what it's worth, you know, to the individual, and then whether the company, and then can we find the best way to to do that as well. And I guess the, the most concrete thing we've moved towards that in, uh, we have a product, product called Claude Managed Agents, where we'll run all of the infrastructure for you in terms of doing all of the, you know, agentic harness and calling the tools, etc. And you can either do it in sort of the normal mode, which is you give it tasks, it, you know, will go through tokens, it'll tell you when it's done. Or we have an outcome-based mode where you can say, here's what good looks like, here's a rubric, go and do it, and, you know, you know, it'll go off and and make it more outcome. So, like, if everybody had moved on to that API, then I think maybe we could have a different sort of outcome-based, uh, pricing, but we'll see how that gets adopted.
John or guys in the back, can Do we have the random image? Can we show the random image? Um, if we can, great.
It's the random image.
Oh, here it is. Okay. It's just because we didn't have a good label for it, so we just called it the random image 'cause it might come up at any point. But this is a chart from the Financial Times, um, speaking of utility, where it shows the amount of app releases that have come out, um, which are skyrocketing, and then apps with significant usage that seem to be going down, and app reviews, which seem to be going down. Um, so Mike, I'd love to hear you respond to what we're seeing in the image here. Um, is it possible that like everybody's coding and releasing, but we're not really seeing a big boom in productivity?
That's really, I mean, I think there's definitely a parallel on in app usage in general. I'd be interesting to seeing
If any of those app releases became one of the apps with significant usage. Um, all right, we can take it down. Yeah, but go ahead.
Um, I think it ties into something I've been thinking a lot about, which obviously my background is in consumer, and I've been wondering what the consumer AI breakouts will end up being. And I don't know if we've seen a lot of them yet. Um, and I think part of it is, you know, I don't know how far back that chart goes, but when we were releasing Instagram, it still felt a little bit wild west in terms of the apps. Like people were excited about apps and like two kind of random people released an app and were able to get to like number one in photos and video within three months, right? I think that is much harder now when you think about how consolidated the top 10 is, and how much time spent is spent on like the TikToks and Reels of the world. It's a lot, right? And so I think getting that breakthrough consumer experience, I think is really, really hard. So I think that is, uh, as much a story about how sort of consolidated consumer products are these days, number one. Number two, how, um, entrenched or how powerful it is to have, uh, that sort of data, uh, people call it data gravity. Like the data gravity of something like your Google Docs are in your Google Docs. So even if somebody has a like two times better AI-powered, you know, doc editor, you going to move all your stuff up? Maybe, probably not. So I think that it speaks to, you know, the things that are sticky. I think I think about a lot is like the hard stuff is still hard. Like making something people want still really hard. You know, we have amazing bottles internally on topic. Not all of our products work, right? And so, I think that's a bullish sign for like product people like me because it means that I think we hopefully still add value. But I think that chart is maybe another place where it's like, uh, it's harder in many ways than ever to break through even if you can code more quickly. And we could have done Instagram in, you know, a month instead of three or four? Probably. But we got there after like a long winding turns and twists and turns process.
Had I, I think 18 people at Instagram when you sold it?
13.
With these tools, do you think you would have, how many people do you think you would have had?
Well, it's really interesting cuz of those 13.
Cuz like everyone's like, "Oh, one billion, one billion dollar, one person startup." How close could you guys have gotten?
I think we could have gotten there with like four to six. You know?
Okay.
Yeah. Um, or the thing that we would have done if we had grown up, we'd be able to do some things in more than a single track. Like Instagram was, if you ever watched my five-year-old play soccer now, by which I mean like the ball is there and every single person runs to the ball. Like that was our product team. It was like video, [laughter] go. And everybody like goes and works on the one thing and like we'd be able to like play positions. Like Android we built in about a month for Instagram. We could have done it probably in a week with the models. And but and to build Android we took everybody off iOS and we all like relearned to code in, you know, code Android OS. And then we went off and do that. And like for that whole month we were barely shipping updates on iOS. So, I think you can be a lot more. Actually, a really good example. There's a labs project I have internally that, uh, helps accelerate how Anthropic engineers like code and do code review. And that that project I am maintaining an iOS and an Android version of and I basically have the cloud that works on the iOS one basically like ping the Android one and be like, "Yeah, I implemented this." Sorry, Android users, it's still the second one even in AI on world, sorry. Um, and then the Android version is like, "Okay, I'm going to do this." Oh, that doesn't count because I, that feature doesn't make sense here. I'm going to drop it. And of course, we wouldn't have been able to delegate all of that on at Instagram, but we sure could have done a lot by having sort of platform parity. Like this dream of platform like close to parity is now actually quite doable.
Mhm. You're probably going to get calls now from the remaining six or seven people on your Instagram team going, "Who was I? Did I make the cut?" In the new era? Also, it sounds like you probably could bring Gotham back now if you really wanted to with all the.
I forget if we eventually, I think for April Fools maybe we brought it back one day. Yeah.
Yeah. Do we have time for one more question?
Last question. Yeah.
Okay, my, I mean, my last question for you is, um, you worked on a product that now as it has evolved is, is, um, in many ways ethically fraught because of some of the harms that people are concerned about with children. And when you talk about the fact that there hasn't really been a big breakout consumer app for AI, I think there has and it's chatbots, right? And chatbots have also led to some real dangers and harms for young people. And so when you are building in labs, how are you thinking about the, you know, the potential harms and the risks that come with just making this technology that much better?
Yeah, I mean, I think there are certainly products that we have either prototyped or conceptualized and been like, "This product is, it kind of sounds so hypey. I hate this." But like, this product if shipped would be bad for the world or would nudge people in the wrong direction. Or even if we did it right, the like wrong or like more morally fraught version of this would be actively we think bad. And so I think asking that question a lot internally, um, makes a difference, um, and having, it's a luxury to have core products and models that are doing really, really well so we don't, like that's, that in some ways an easy decision if we, if we think it could get a lot of, um, a lot of use. But yeah, I think going back to an earlier conversation, I think front-loading it is really valuable and really thinking through like, it is now more normalized to have people at a company and definitely at Anthropic does for like economists thinking about the impact of the thing that you're building on the world and that was [music] not the case in along the years on on most of social media, I think. Mike, it's always great speaking with you. Thank you again for bringing your insight today. And [applause] let's do it again soon. Let's hear it for Mike and Lauren. Thank [applause] you. Great job. Thank you.