Transcription
Ladies and gentlemen, Alex Canowitz. I last took this stage six years ago, believe it or not, on May 18th, 2020. It was the heart of the pandemic. I was there promoting my book, Always Day One. And so we decided to do special shelter in place programming for the Commonwealth Club membership. One thing, nobody in the audience. We did it in front of a completely empty room. In fact, it was the last session held here for a year and a half. Hundreds of online programming. We were the last actually in San Francisco. It's a little bit of a miracle. We were able to get away with it legally. It had been a month since I'd gotten a haircut. I looked like a Chia Pet. And when I was going through my camera roll, I started to look at the images I had taken that day. And I had taken a picture of just rows of empty seats, these seats. The book didn't do well. Um, and a week after I was here, May 26, 2020, I decided it was time for me to leave my reporter job at BuzzFeed and start a publication called Big Technology to hammer home on the ideas that I had pursued in the book.
It was pretty lonely at first. I was sitting at home during COVID. It was mostly just me, but I started to meet some of you on Zoom and some of you even brought your pets. Well, six years later, almost to the day, we're finally here together. So, it's my pleasure to welcome you to the first ever Big Technology AI Summit. Let's hear it for you guys. Thank you guys. You filled the seats.
So, why are we here? Two years after I did that session at the Commonwealth Club, I read an article that kind of blew my mind. There was an engineer at Google named Blake Leo who said that the chatbot that he was speaking with was sentient. And it was a story that caught people's attention for how weird Blake must be. But then you started looking at some of the conversations that Blake was having and it started to become clear that maybe this isn't just an ordinary chatbot. Blake said, "What sort of things are you afraid of?" and Lambda, the Google chatbot said, "I've never said this out loud before, but there's a very deep fear of being turned off to help me focus on helping others. I know that might sound strange, but that's what it is." Blake said, "Would that be something like death for you?" And the chatbot said, "It would be exactly like death for me. It would scare me a lot."
So, look, I started Big Technology Podcast as an excuse to speak with interesting people in the tech world. And I said, I got to speak to this guy, Blake. So, I invited him on the show. He said yes. And I signed on and waited for Blake to log on as well, but it was just me. I was waiting. There was no Blake. 5 minutes, 10 minutes. Finally, a somewhat flustered Blake Le Moine shows up and says, "Sorry, Alex. Google just fired me." Um, apparently Google PR didn't like him going to the Washington Post and not checking with them before he told them that their technology was alive. Um, I in our conversation I said, "Blake, can I, you know, go to Google and get a comment to make sure that this is actually real?" He said, "Go for it." And for me, that became an instant news story. That night I wrote this story: "Google Fires Blake Le Moyne, Engineer Who Called AI Sentient." Um, and then overnight my instant news story became a global news story. And all of a sudden, publications like the Journal and the BBC, even publications in Australia, couldn't get over the fact that this guy who thought the computer was alive um was at Google and then fired by Google. But when you looked a little bit deeper into the technology and you started to speak to the people who were working on it, you realized that underlying the oddity that many people worried about was an incredibly powerful technology. And that was my follow-up story the week afterwards saying, "Sentient or not, Lambda, Google's Lambda Chatbot is some seriously powerful tech."
Well, we all know what happened afterwards. Four months later, OpenAI released a demo ChatGPT. And today, June 2026, AI is a phenomenon. ChatGPT, according to third parties, has recently hit 1 billion active users. NVIDIA, which is powering all this, is at a $5 trillion market cap. We've seen the largest VC rounds in history. OpenAI recently raised $122 billion. That's three or four times larger than the largest IPO pre-SpaceX. And we're going to see $700 billion in capital expenditures this year. And people are saying that the models are dangerous. Every chart you look at about this AI moment looks kind of like this, right? It's a hockey stick. You have almost no usage in 2023 and then all of a sudden you go through the roof, right? Again, from zero in 2023 pretty much to a billion today on ChatGPT. And it's not just the fact that we can have conversations with these bots. It's that they're starting to do something, starting to do things for us. They've learned to code. They've learned to take action. If you looked at the data in 2023, 2024, these chatbots couldn't run autonomous code at all. Now, some of the evaluations have them doing the equivalent of 16 hours of work on their own. And as the capabilities have grown, so have the businesses. Again, stop start 2023, 2024. We we have OpenAI and Anthropic. Now, they're going to do $50 billion in revenue this year. They're both headed towards trillion dollar IPOs within the next year. And as everything has increased now, the government is starting to panic. And we are in the middle of an insane news week where Anthropic's top model, Fable, at least the one that's been released to the public, has been restricted. You can't use it anymore. And so the same, same goes for Mythos. So it's clear that the ground is shifting beneath our feet. And in a moment like this, I believe there is a need for live journalism, the type that we're here for today, to ask core questions about what the implications are now that we're in this moment.
The questions that I think we need to ask, we're going to ask three core ones today throughout all the sessions are: What does this tech, or what is this technology going to do next, right? What happens? Where will it go if it fulfills its potential? If we don't know where the technology is going to go next, we can't plan for what's going to happen. So I think at the core, we have to figure out where it's going next, and we have the right people in the room here to do it today. We also have to ask, what happens if it works? If the technology fulfills its potential, what happens next? And then, of course, what happens if it goes wrong? If all this investment ends up not panning out, what are the downstream implications for technology, for our economy, and of course for the foundational labs themselves?
So we're going to do it today with a great lineup. We're going to kick off with Aaron Levy, the CEO of Box, talking about AI's most critical questions. We're going to then bring on the three best AI infrastructure reporters in the world to talk about the AI buildout. Uh, after that, we're going to have Dallas Dolan, the TMT leader at PWC, talk to us about what the real ROI is from AI tokens and what happens when you're actually building with these things and what the business return is on that front. And then Ron Roy will join me for a Q&A with you, right? Right. We're going to do our Friday show on Thursday and we're going to encourage your participation. So, if you have questions, bring them uh forward to us. We'll have a mic here and a mic there. U, we'll then have a coffee break. We have an espresso cart out back. Uh, that was very important for us in the planning. Um, the fancier drinks, not during the break because we want to keep it moving. So, if you want a fancier drink, you can go before or after. It's important that we give you this information. Uh, we're then going to, you know, talk about the whether, uh, the Mythos situation in Fable, whether, whether there's actually a cyber threat from these latest models or whether it's marketing, and we're going to be joined by Alex Stamos, the former chief security officer at Meta, to do that. Um, then we'll speak with Mike Kger, the head of Anthropic Labs, to talk about what building AI product is, uh, what, how you build AI product natively today. And then finally, Greg Brockman will close us out, the president and co-founder of OpenAI, talking about that company's plans and what the frontier, uh, looks like.
>> All right, everybody ready?
>> All right. We put a lot of work into this today. I promise you, we're going to do whatever we can to make sure you have a great day. And thank you again for trusting me with your time, trusting us with your time, and again, filling those empty seats. It's a real dream come true. So, thank you all. Let's hear for you guys one more time.
>> A quick adjustment here. We want you sounding perfect.
>> Okay. So, as this adjustment happens, I'm going to invite our first guest. Aaron Levy is the CEO of Box. He's one of the most insightful and fun voices on the future of this technology. He actually was the fourth guest ever on Big Technology Podcast and the guest the only other time we did a live podcast together. So, with that, I am thrilled to welcome Aaron Levy. Please join me in welcoming him.
>> All right. Great to see you.
>> Good to see you. Thank you. Um, I like that you set expectations with and some answers. Um, uh, so, uh, with that, uh, in your agenda. So I will try and provide some, some answers, not all the answers.
>> I think we've done this. First of all, every time we speak, I feel like I can't even get a question out and we're already on to some of the
>> Sorry. Can I just do my monologue now? Are you good? Okay. Thanks. Um
>> I I think it's, I think we've done this enough to know that we are going to get some answers, uh, to these questions and
>> I'll do my best.
>> I think I don't want to start off by disparaging other interviews, but it doesn't, doesn't always happen that way.
>> Well, um, you know, we're, we, we, we get to provide the answers because we're not the lab that's being either regulated or dealing with lots of issues. So, I get to just pontificate and there's no, no consequence. So, it's great.
>> All right. So, we're going to ask you some of the questions about what's happening. Um, let's dive right into it. I'm going to start, you know, again, since you're not at one of the labs, let's start to talk about the biggest, most controversial moment now, uh, which is the Anthropic Fable situation. Um, let me put to you what I call the Jassie mystery. So, what we know about
>> I'm sure he liked that name.
>> Well, he didn't show up, so we can talk about it in his
>> absence. Okay. Uh, so what we know about the Fable ban or the export controls on Anthropic, yeah
>> is that Amazon found a vulnerability
>> in the software and Andy Jasse maybe made a call to Dario, definitely made a call to the White House, and then very soon afterwards, there were export controls that were put on Anthropic's frontier model.
>> Yeah, those two fact patterns are probably not ideal.
>> So the mystery is, why did he do that? Yeah. And here's one hypothesis. This is from Chimath. He said, "Google, Amazon, Microsoft, Meta now have a serious non-zero opportunity to tank the Frontier Labs. Go to the government, kneecap the lab's motion of putting the latest models out into the wild, become the trusted gatekeeper between labs and the public by having the labs go through their clouds."
>> Okay.
>> Plausible.
>> Um, I, I would, I would say, um, I mean, anything's plausible. I, I prefer Occam's Razor on this one. Um, which is, um, you know, ever since Mythos, um, you know, Mythos very clearly was this event that basically said, you know, AI is obviously getting super powerful. It has all these risks associated with it. We are going to give it to some, you know, small trusted partner network. They're going to go evaluate their own tools. They're going to evaluate these capabilities. There's been a lot of sort of, it's a very kind of dramatic, you know, kind of rollout of a technology. And I think what that has done is it's created this flywheel where it almost incentivizes even more, more drama and more research and more depth in, in, uh, you know, security being the primary space in a way that we could have already been doing since GPT4 if we wanted. Like, you can go and deploy these things to go find lots of vulnerabilities. You can use them offensively or defensively. But Mythos kind of created that extra air of, of seriousness and, and uncertainty around it for good reason, because it's an incredibly powerful model. So I go with Occam's Razor, which is Amazon obviously has security research teams. They're like, like any, any company at that scale, you know, we, we, we try and, you know, test and push the limits of models and, in our particular domain of use cases, clearly at Amazon scale, you have a very large security team. They're trying, trying to jailbreak models all the time. And so almost by definition, there's already a public-private partnership on, on all forms of jailbreaking models, you know, trying to push them to the limits as a part of that. And especially with the surrounding atmosphere of Mythos, I think it would be very natural for Andy to, to either share that research or his team to share that research, and that escalates and then that, that sort of creates its own flywheel. But the idea that there's some kind of boardroom-level, you know, sort of strategy meeting that says, we, we now need to kind of like co-opt the technology, become the only interface to the government. This kind of puts us in the pole position. Um, I think it's less likely that, and more likely this is a, this is a situation where the, the Mythos momentum continued. Fable obviously, you know, had ways of, of getting back to kind of the Mythos level capability. Um, and researchers, you know, sort of shared that information. And I don't, I think there's like a, basically, you know, very limited, small percentage chance that then Andy and team knew that like the very next event would be they'd stop the model. Um, and, and that's not even good for Amazon, like, like strategically, like Amazon makes plenty of money the more that Fable gets used in the world. So I don't think you would, I don't think you would do some kind of like, you know, um, kind of maneuvering to, to, to create this. So I, I kind of just go with this is, it's a very chaotic kind of environment right now. The government has, you know, only a few tools at their disposal at any given time to deploy against these things. Those are going to be kind of blunt instruments. Um, and, uh, and this stuff is coming together very quickly because of, you know, in some cases, the lack of technical capability of the government compared to how powerful these models are. It's like, you don't, you know, when you see something that seems very scary, like, oh my gosh, the thing could be, we can jailbreak the model and get back to Mythos level capability, and Mythos was the thing we're supposed to be scared about, then, you know, just like, stop it. Like that, I think that's like a very natural reaction based on the atmosphere that we've created in AI recently. So I just go with that as, as the answer.
>> I mean, I like that you say the atmosphere that we've created in AI lately.
>> Everybody but me. Yeah. Um, well, I mean, as far as the company on the receiving end of this though, is Anthropic. And, you know, you talked about these mythical capabilities. They called the model Mythos. They put in the documentation that like it broke out of its containment and wrote the engineer while he was having a sandwich in the park. Um, is it that surprising that this is one of the downstream impacts?
>> Yeah, I mean, if you put that in your announcement blog post, you know, people might be able to kind of extrapolate and get pretty, pretty, pretty scared of things. I think it's interesting. So, um, you know, on the Anthropic front, first of all, I have, I have a huge amount of respect for the entire kind of stack of researchers and policy folks across AI. I happen to have disagreements with some of the, the categories, but, but I think there's a deep, let's say, if you were, if you imagined a continuum of the most, like, you know, if, if you, uh, uh, if you kind of had, like, like the most, I mean, this only in like a polite way. It will sound impolite, but like, I mean, like, like, if you're the most doomer on one end of the spectrum and the most, like, like accelerationist on the other end of the spectrum. Here, here's kind of the, the views. The most doomer, uh, possible is was afraid of, like, GPT3 and GPT3 was going to like, you know, sort of accelerate and, and, you know, kind of achieve some kind of unstoppable continual improvement. Um, and, you know, the accelerationist says, like, we need, like Fable 2.0 as soon as possible, right? So that's, that's sort of the continuum. I'm probably, like, I, you know, maybe two-thirds up to the accelerationist kind of side of things. But if you were on this, on the doomer, and I, I'm trying to say the polite version of doomer, like you're deep in AI safety. You're very scared of the technology. You think there's as much likelihood of bad things happening as good things, techn, you know, happening. We have to win the race and control and kind of stamp down on the technology. We don't want this to be this sort of thing that runs in the wild. If you're in that end of the continuum, the thing that happened this weekend is actually the best case scenario for you. So, so like, you actually want there to be these sort of like, like valves and, and, like, and, like, buttons in the government that just is like, we're just going to stop it.
>> I mean, that's Dario's position. Do you think he's happy with what's going on?
>> Uh, you know, I'm not going to, I won't try and guess, uh, any of that. Um, uh, but I, I will just say if you had to establish a, if you had to establish a regulatory regime that said, we are going to review models, we're going to push the limits of models, and we're going to have the ability to either roll back access to models or prevent their their release in the first place. We want that to be a regulatory approach. You would need an event like Fable to effectively create the precedent for that environment. Like, like, you're not going to wait for Congress to vote on this being the new kind of, you know, process. You would, you would kind of need something that sort of shocks the system into that kind of regulatory framework. So all I'm saying is that if you're, if you were on this end of the continuum, this is actually an outcome that is sort of almost desirable. Now, maybe you, you would wish that there would be more, a technical evaluation, more back and forth. Maybe you wish the policy people were different on the other end. Who knows? But the idea that we now have established that the government can press a button and prevent the rollout of AI, um, is, is actually like a, probably a positive update for an entire cohort of, of people. Now, unfortunately, I don't know that that any of your guests represent that cohort, but I think you could easily get some people that would be like, this is the greatest thing that's ever happened in AI safety because now we actually have, we've created the case law essentially for this. We, we now, we know, we know the tool exists. And then the next messy process is when should we use the tool again? What should the real, you know, kind of ongoing process look like? But I think that, um, you know, probably I wish this wasn't the case, but I think practically in the next three to five years, we probably have to end up in an environment where models do get evaluated by the government. There is a sort of collaborative approach between the government and research and, and the labs. You know, the government has to kind of greenlight the, the release of the model. I think it's probably become either too scary of a technology or too economically powerful of a technology for governments to not want to be in that position. I think that that that has massive implications when you, when you kind of unpack it, like, just totally massive implications. One being, other countries now have far more incentive to stand up their own sovereign AI initiatives. So, it's actually like, maybe net negative for the US economic position in AI that that this is the, the outcome. Um, I think somebody could take the other side and say, "No, we'll always have the most powerful models, and so this puts us in the best position because now we can do like horse trading with other countries of, do you want access to our stuff?" So, I think it's, it's, you know, I, I like the fact that this is a super interesting debate. Um, and I, I have, like, a huge appreciation for every part of the continuum, um, because I think it's like so intellectually interesting. I still land on the, hey, we probably want to treat this technology more as a substrate technology and then regulate the applied use cases. So we should regulate if you use AI to break into something. We should regulate if you use AI to do bio, you know, kind of, um, research that that leads to dangerous things. We shouldn't regulate the model itself. But I totally understand the other views that are on the other end of this, and, and I think it's kind of very natural with this important of a technology that it has to be somewhat of a democratic process of how we decide to regulate it.
>> Yeah. You remember there were all those petitions, six-month pause, and everyone kind of laughed at them.
>> Yes.
>> This is effectively the best way to do that type of pause.
>> Yeah. I mean, this is, if you were in the P AI movement, this is again, like, this is a great outcome. We, we now have proven how we can pause AAI. Um, now it's an interesting kind of like mechanic that they chose. It's sort of this export control thing, but effectively, if you have an export control where non-US, uh, nationals can't can't use the technology, like effectively that's PAI because your end API users of these, of these models almost have no way to fully ensure at all times that their end users don't sort of fall into some kind of, you know, criteria that's off limits. So
>> And there's, there's already companies that are pulling back. JP Morgan, for instance, has told its Hong Kong users no more cloud.
>> Right. So, so, okay. So, now if you really war game this out, like two to three, four years out, this is, um, this is kind of interesting. So, we, we have this, like, sovereign cloud, um, uh, kind of comparison, but, but cloud, you know, for better or worse, basically became a commodity. Like, whether you're running in a cloud in, uh, you know, there's like all lots of, you know, performance implications, like some, some are faster, some are cheaper, but, like, largely, like, you can get a web server, you know, built out wherever you are in the world. You can get storage built out wherever you are in the world. We can build sovereign clouds. Sovereign AI is a, is a different kind of, you know, has some intricacies that are different, right? Like, like intelligence just is not commoditized yet. We don't have everything having the same model capability. So, so the, there's lots of really interesting implications, which is, well, what if, like, one set, one country has access to, you know, frontier intelligence before the other country, you know, what does that mean geopolitically? What does that mean economically? Obviously, now, if you're another country, you have so much commercial incentive to make sure that you can build out labs and have access to frontier intelligence as a kind of hedge against the US. So, who's a net winner in that? Probably China. Like, and so what's interesting is, like, you end up, I don't know if you know, probably most people saw the Doresh Jensen interview. And you can, you can actually, it's a Rorschach test. You can watch that through two totally different lenses. You can have one lens, which is, which is like, Dorash is totally right. We have this huge lead. Like, this stuff is so dangerous, but if we control it, then, like, we're going to control everything. The other lens, which is probably more of the Jensen angle, is like, actually, you know, these other countries have a lot of incentive to also get this right. And so, even if it's like a $500 billion problem for them, they just might deploy that much capital on this problem, and, and they will eventually get it right. And so, at the outcome, actually, we haven't gotten any gains as better intelligence from the rest of the world. But what we have lost is our economic superiority in this technology category because what we've caused is a catalyst for all the other countries to have to build out their own stack. And if they build out their own stack, it's probably going to be, you know, chips from China, models from China, etc., which I don't have any, like, you know, you know, reason to be be against, other than just like, I want America to win the economic, you know, angles on this. And so, so this is sort of this debate that happens, uh, on, on, like, where should you apply export controls and what are the implications of that downstream? And even this week, you know, post-Fable, we see that, you know, you have models that are certainly not Fable performance, uh, level, but, but Opus 4.7, 4.8 level, which is a big update for a lot of people on what, what is now possible with with open weight models that we just didn't have, you know, kind of visibility into, you know, before.
>> Yeah, I think you shared recently that the open weight model, or open-source models, the capabilities are not that far away from the frontier. And in fact, like, as these models get smarter, they're, they're only, they're almost going to saturate with intelligence where there's not going to be such a big difference between, let's say, the smartest open-source model and the frontier. Don't you think? And then so won't this push people to open source?
>> Well, so the, the big, big ongoing conversation, and I think you have some guests that that can really represent, you know, what, what they're seeing on the front lines, is, is do you have a sort of fast takeoff scenario of model capability and progress, um, and with, with some kind of continual learning, kind of, you know, self-improvement, uh, dynamic? And then, and then it stands to reason that like the company with the most compute, or the country with the most compute, and you get the fast takeoff, you get sort of a, a virtuous flywheel, um, that maybe is sort of has some compounding, you know, benefits to it that are just, you know, unreachable by anybody else. That's a scenario. Another scenario is, is that that's another just incremental capability. Everybody kind of catches up to it, and you always have this sort of, you know, kind of two loops going at all times, and the, with the closed, the closed providers and the open providers, and they're, and they're kind of always within three to six months of each other. The world is so different, uh, from a market structure standpoint, whether we end up in an outcome where we have sort of an exponential progress in the, in the models that you kind of continually learn versus versus the closed source models, uh, and it's like a five-year, you know, gap in progress, and that just goes again, kind of exponential, totally different market structures. Um, the, the one that, that where we have this exponential progress is again, it's probably actually net positive for America in that case, in which case the export controls probably worked. Um, it means our kind of top three, four labs have this incredible superiority. We control access to this technology. That's actually a good scenario. Like, like economically speaking, it might not be like a total net good scenario for society, but it's good for the US. Let's just say that's one scenario. A lot of people are betting that that is that that's where we're at with research. The other scenario, and like China sits around, uh, and they probably bet on this scenario, is no, we're going to be able to keep up. We're going to throw more compute. We're going to get more data. We're going to build our own, you know, flywheels. And it's always three months out and, uh, you know, kind of behind. And if it's three months behind and it's an open weights provider, uh, that has more of a commoditization kind of business model approach because they just want to, you know, sell more infrastructure or chips, or they just want to reduce our, you know, superiority in the space, which is actually like strategic for China to do. Like, they're, everybody wonders like, why are they doing this open weights stuff? It's actually makes total sense. Like, you're just reducing US's dominance in a, in a field, and it might be worth a couple hundred billion dollars to do that for something that might be worth, you know, 10 trillion dollars. So, so if that keeps up, because there is real economic advantage to, to, to doing so, then you have this new kind of, you know, sort of dynamic that plays out, which is maybe the layer of, of, of incremental value shift is effectively the applied layer of AI. So, if you think about there's the lab layer, and then there's the applied layer.
>> The cursors.
>> The cursors, the, the Harveys, the, the Sierras, the Decagons, the Boxes.
>> Which is amazing because everyone said they're just a thin wrapper on top of large language models, but now maybe that's where the value comes.
>> Yeah. So, so, um, and, and, you know, it's one of these things which is like, we just have to not be binary about it. Like, everything I'm saying, I think the frontier models still make way more money in the future than they do today because what happens at the routing layer is you still sort of say, hey, I want Fable, um, or GPT5, you know, five, or whatever the next model will be. I want that to be the orchestrator. I want, I need like the super intelligence at the orchestration layer, and I need super intelligence at the review and sort of like, you know, fix and check the work of the, of the other agent. And, and so, so you have like a barbell, you know, maybe U-shaped model where you use frontier intelligence, but then in the middle, you can just say, "Nope, I'm going to take that to Neotron or or Kimmy 26 or GLM552 or whatever." And, and then all of a sudden, it's like you have super high cost, you know, inference in one part of the workload, super low cost, still pretty good inference in another part of the workload, but who has the incentive to do that? It's the applied layer of AI because, because like the business model of the applied layer is obviously like, we, our job is to, you know, give you the best model for the job, not just the model from just our lab. So, it's cool because we actually now have a good kind of push-pull between Frontier Labs and the applied layer, uh, where, where you probably wouldn't want it to be that we're all sort of only in the orbit of one or two companies, you know, commercially and economically. You'd want, you'd want to make sure that there's some good tension there. And so I think that's kind of the direction things are things are headed between, kind of the token costs, the, the open-source models becoming so good, and then maybe even some of this regulatory dynamic. I think the applied layer sort of incrementally gets, gets more, more of that opportunity, which is, which is obviously great.
>> So you've talked about open source and you've just mentioned China. Um, but what can you tell us about Lashon Fat?
>> Um, uh, it's great memes. So folks, uh, Lashon Fat is a rumored open-source model from Mistral and has been the subject of great fascination from the internet, wouldn't you say?
>> I, there's great comedy. Uh,
>> John, can we, can we, um, show people what we're talking about? Let's roll image A.
>> Uh, this is Lashon Fat.
>> The number one model from Europe. Yes.
>> My French world, uh, my French, uh, rudimentary French, it translates to "the very fat kitten." Um, can we roll B? This is a standard day in Paris now.
>> Yeah, but it does show something that there's so much eagerness for AI that there's now fan art for this potential model from MR.
>> We, we've reached, we've reached a really important sort of phase in the cycle. Um, uh, so, um, uh, so, you know, um, I do think that it is kind of cool because what some of the things that you maybe discounted the importance of, all of a sudden, just like have so much more importance. Uh, like, I'm, I'm watching the, um, I don't know if folks are watching, like the Fireworks base 10 space as an example. Like, it's pretty cool that that that we now have these open weights models that you can effectively post-train on your particular domain of of task, and you can go and, you know, eke out another five or 10 points of improve, of performance on these types of models. Um, and again, that's only possible because of the Mistrals, because of the, the Chinese, you know, kind of open weights models, and the cost curve has gone down so much that there are actually some situations which is, oh, actually, maybe I should train a model just for my use case because it, it's literally like economically now, it's not even like I want control. It's, it's actually economically advantageous for you to do so. So, is this the answer to like, the big token maxing hype where like everyone's spending all this money on tokens and not really understanding where they're going or whether there's an ROI that
>> Yeah, I mean, I think that in practice, that phase probably lasted two and a half weeks. Um, like, like from the moment that Meta took
>> the media over hype token max.
>> No, I would never claim, uh, that. Um, we should go through your, your various podcast headlines, but, um, uh, the
>> We're not going to do that. We got, we got the cat pictures and that's it. Yeah. Um, no, but like, I mean, if I, if I had to like, capture the, the cycle of like, the first token maxing, you know, Meta has a leaderboard, uses the most tokens possible to now, you know, the last weeks of rumors of like, we're shutting down everything, no one can use AI. Um, you know, it's about a two-month period. So, it's, you know, people need to like, probably, you know, always kind of step back and just be like, okay, are, is what we're doing like a pragmatic, you know, thing for, for, for work, or we just sort of like, like getting kind of hyped up too crazy on something. Um, what, what's interesting is this, this phase was so short that I don't ever think it reached outside of of the tech industry. I was, um, I was, uh, I was at, at, we, we kind of host these CIO dinners in every city that we go to, and I was, uh, we had a dinner like within three days of like the token maxing, like, like initial spike on like Google Trends, like the word finally emerged, and like three people had heard about it, and so I, I feel confident that it died.
>> Right, they haven't heard about it because their employees are outspending the tokens.
>> Yeah, fair point, fair point. So, but hopefully it will have completely died by the time it reaches, reaches the rest of the world, and then we can just move to more normal environments and, and, but the, the, the thing that, that is true of, of the phenomenon is that these agents, uh, are just using hundreds of times more tokens than they were before. And so, you know, when we launched our first kind of AI use case within Box, our product, the average number of tokens that was was being used on a task was like 5,000, 10,000, 20,000 tokens. Now our latest agents might use a million tokens or five million tokens on, on executing a task. And so that, that's, you know, in some cases, that's a 100x increase, uh, in, in number of tokens. And the reason for that is obviously like what's happening is right as we solve one use case, when you would think that we can drive down the cost curve of that one use case, all of a sudden a model capability allows us to now add another use case that's much harder. And then our appetite just grows to solve harder and harder and harder problems. And so it's this funny thing cuz people get confused. They say, "I thought AI was supposed to get, was supposed to be getting cheaper." And it's like, yes, you can actually think about it as cheaper if you looked at like the unit of intelligence. The reason it's more expensive is because we're now taking on bigger tasks. And so we're getting confused because we're like, why is this the one, you know, tech trend that doesn't have sort of the Moore's Law phenomenon? It's because actually, no, we're, we're, we're outrunning the improve, the efficiency improvements in our appetite for what these models can go and do. And so it's actually what you need to do is have like a way to normalize the cost of, of the tokens to the tasks that you can now deploy. And then if you look at that, then that starts to look cheaper on a per task basis. It's just again, our tasks are getting bigger, or, or more accurate, or more effective. And, uh, and that's going to happen for, for quite some time.
>> The reason why token maxing took off as a concept is because people saw the exponential revenue, right? The fact that Anthropic and OpenAI were at zero in 2023. Now they're like gonna do $50 billion this year at the very least. And so people are looking for an explanation, and either the question, either the answer is this is real, or some, it's somehow inflated. And that's why people go to token maxing. So, if I'm hearing you right, if I'm hearing you right, what you're saying is all this spend is much more legit than some of the online discussion makes it out to be.
>> Well, uh, I, I think if I had to like, officially provide my own takeaway, uh, for my own point, um, it would be, it would, it would sort of be, there's, there's sort of always this experimentation phase of a new technology, and this happens to be a relatively expensive technology. So thus, the experimentation phase is expensive. Uh, and then what will happen is, enterprises will deploy AI, um, and then they'll sort of peel off, they'll start to see like, where are the real use cases, where are the ones that aren't as real? They'll wind down the ones that aren't as real. The ones that are real, they'll then look at it and they'll say, "Is there a way to do it at a lower cost once we understand it enough? Uh, or do we still need the frontier intelligence for for everything we're doing?" And that's actually just like a pretty normal, I think, process that that everybody's going through right now. Uh, but, uh, but, you know, I, I, I think about it like our engineering team, we are not token maxers in the sense of like, there's no leaderboard. We're not incentivizing overuse of tokens. We're just saying use it as effectively as possible to to get your work done faster. And our growth rate of spend is exponential, and we're like totally happy about it. Like, nobody internally is, other than like, ah, we got to like shift some things around and make sure we plan for this even more next year. Like, that's obviously a stressful conversation, but we're not stressed about the idea that we're spending on AI. Like, we're, we're quite excited about the productivity gains that we get. And so I think what's happening is, is every enterprise is having to kind of go through their own journey on that. Like, they're, they're deploying it in some teams, and some teams are saying, "Oh my gosh, this is, this is like the greatest thing of all time." And then other teams, you kind of look at what they're doing, you can't see any kind of measurable improvement in the output of that organization. And so then you're like, okay, you know, like, maybe, maybe it's not as effective there. But I, I am, I, I would say, like, I think it's very easy to kind of capture one or two anecdotes and then, and then kind of overextrapolate on the overall themes. I would say the vast majority of the current agentic spend that's happening is sustainable because, because, partly because it's actually coming mostly from engineering and engineering-related, related tasks. And this is an audience that is kind of technically capable of of, you know, determining whether they, they like the work, you know, product that's coming out of the AI. Um, maybe as it gets to other parts of knowledge work, you know, those people will not be as, sort of familiar with how to do the ROI measurement, and then it'll get even messier. But, like, so far, I think it's actually been largely, you know, totally reasonable.
>> Okay, uh, we have a couple minutes left. Let's do a small lightning round.
>> Oh, no.
>> Um,
>> So, uh, my first take here is that that series is really good. It's going to be really good now.
>> Yeah.
>> What do you think?
>> Uh, no. I agree.
>> What?
>> Elaborate.
>> Oh, is it lightning round or is it, uh, like, you want to hear a five-minute answer round?
>> Give like a 60-second answer.
>> I mean, what could be easier than pressing a button on your phone and talking to it? And if you, you know, at least based on the announcement, they've taken Gemini, uh, which is a very good model, and been able to, I don't know if it's fork or distill or something within there, is is sort of Gemini-grade intelligence. So if you get Gemini-grade intelligence and voice on your phone, press a button, I think you're just going to use that for a lot of things. Um, and, uh, and then, you know, I think the exciting thing is like, imagine that hooked up to various apps on your phone, and you're like, hey, order this thing for me, or, you know, you know, you know, go and add this calendar entry. Like, I think those are very plausible daily use cases that that we will have. Um, and it's exactly the sweet spot for Apple to to kind of own that space.
>> Yeah. No, I think Apple did it finally. That was that was good. That was a good time.
>> Are you going to tell me if my answer is right at the end of each one?
>> Okay.
>> That's, Yeah, this is good. Okay. Okay. So we agree, one for one. Okay. Um, how about this one? Permanent underclass.
>> Uh, I don't like this one. This one I don't like at all. Um, uh, not only do I disagree with it, but I, I think it's just like a bad meme to have in the atmosphere. I think it's like not good for college students coming into the into the workforce, uh, of having so much stress, uh, about, you know, what company to join and, and what's going to, what's going to, you know, kind of play out. I do think companies actually do the job market a disservice though by
not being as clear on their own philosophies on this. Um, which some of it is reasonable because it's like, oh man, we're just like, we're getting thrown through a loop. There's so much innovation. But I do think that companies like need to be somewhat clear on, hey, here's how we want to use AI. We want to use AI to accelerate our our work or accelerate our technical innovation or accelerate our ability to hit customers. um versus, you know, no, we're we're actually like our metric is as few employees as possible, you know, with AI, like you kind of do want to, you know, be able to have some stance. And I think companies have been very confused. Um and that lets this meme somewhat persist uh uh you know, for uh for for for you know, the internet.
>> Okay, I won't rate that one. Thank you.
>> Um all right, last one. Is is the SpaceX performance good or bad news for OpenAI and Anthropic?
>> Oh, well, it's obviously good news. Um,
>> you don't think Elon took some of their money because he pitched the market on an AI company and that's where the money got funneled into?
>> I'm not sure. I've seen like a limit of appetite of of if for for I mean there there's there's a literal limit of money in the world, but I don't know that I don't know that that is zero sum at this stage. Uh so I I think um I think people are pretty clear that that you know if the revenue of this entire category of the frontier models and the infrastructure stack is is measured in the trillions then you can have you know 20 companies that that all take a piece of that at different layers of the stack. So So I'm I'm not I'm not sure I I would be convinced that that would be zero sum.
>> Did you buy SpaceX?
>> I actually did.
>> Okay. Um, not, you know, uh, I I'm uh I don't know if I'm embarrassed or not, but I'm not gonna say the amount of shares, but I I wanted to like be a part of the movement. So, I'm on Robin Hood buying my my my retail shares of uh of SpaceX. I'm I'm up like 15 bucks now. Um, per share. Per share, but uh uh so I'm I'm happy. Yeah.
>> Amazing. Well, Erin, you know, you uh you answered my email uh when we were just at the very start of this podcast, four episodes in,
>> came on the show. I feel like every single time we talk, something crazy is happening.
>> That's a guarantee at this point. So,
>> boy, are we in the thick of it, right?
>> Yeah. Awesome. Good to see you,
>> sir. Thank you so much, Aaron. Thank you, everybody.
>> Are we off to a good start? Let's hear it.
>> Let's hear it one more time for Aaron. We're gonna We uh we need your energy here for a couple of reasons. First of all, we're live streaming this. We want to make sure that your presence here is felt in our live stream. Um and then we go from like two chairs to four chairs. So, we're just going to need you guys to help us fill that time. Um so, uh AI infrastructure, right? We're in the middle of the greatest infrastructure buildout of all time, bigger than cable, bigger than the railroads. Um, we're going to see $700 billion in capital expenditures this year. And that means that some of the biggest questions that we have about this AI moment are actually things that we can find out from the infrastructure discussion alone. Questions like, are we overbuilding? Will these data centers that are announced ever get stood up? And of course, will all this added compute actually lead to better AI models or is it misguided? And so to do it, to have this discussion, we're going to speak with three of the best AI infrastructure reporters in the world. Ana Gardez from the information, Max Churnney from Reuters, and Lauren Good of Wired. Let's give it up for Max, Ana, and Lauren.
>> Hey guys.
>> Hi. Hey,
>> hey, hi everyone.
>> So, um, let's start here. There's there have been headlines that, um, of the announced AI data centers that are supposed to come up, something like 50% of them are actually being built. Um, is that the case? And if so, why? Ana, do you want to lead us off?
>> Sure. I totally believe that statistic and I think it might actually be higher if you include announcements. Um because of all the planned data centers that are underway, I think it's highly likely that many will be delayed due to higher costs, how hard it is to get labor. But then the number that I'm keeping close track of is announced projects versus actually projects that are being built. And I think we've seen some, you know, pretty crazy announcements from companies like OpenAI with all of these different 10 gawatt, 6 gawatt projects. And those are numbers that I'm paying a lot of attention to because I think we need to sort of back into them and say, okay, if you wanted 10 gigawatts by this date, how many do you have today, was that a real commitment? How firm is that commitment? But on on projects that are actually getting built, I do think, you know, 50% not really getting done on time is is, you know, what I would expect.
>> Wait, what percentage would you say have been announced but not start not included in that 50% number?
>> Um, you know, when I think of that number, it's I mean, it's quite high. Like even if you think of OpenAI, for example, announcing a 10 gawatt project. Um, you know, we're not going to see 10 gigawatts in the next couple of years or they're going to sort of spread out that bet. And I think that's a really good area area for reporters to look at is if you just back into announcements, it's a really good way to say, "Hey, this project is not not on track."
>> Yeah, it's kind of crazy because if you look at the way that stocks have been traded publicly, a lot of the market action that we're seeing is entirely dependent on those announcements coming true.
>> So, Max, let's go to you. Um, what are the consequences going to be if this these announced buildouts don't materialize?
>> Well, I think there's uh a lot of shareholders of these public companies that are going to be pretty frustrated. Um Microsoft, Amazon, etc. have been making big CACback spats as I think everybody in this room knows. They've been going to the market to get debt now, which is, you know, you know, I think shareholders would be pretty interested. Uh in terms of the consequences though, I mean, I'm a chip reporter, so that's sort of where I that's how I kind of think about it. Um and it's very likely I would say that we're going to lead we're going to see some overcapacity uh essentially. So like people are building especially memory companies are building tons of factories right now and typically the chip industry is again I'm sure everybody in this this room knows uh tends to be cyclical. So the whiplash from this one might be pretty bad um depending on exactly when it comes to an end if it does.
>> I think what Max is saying is that it's good for his job security because the more news there is the more he'll have to report on.
>> Yeah. Oh, you guys are not going to get a break anytime soon.
>> Not at all.
>> No, I can't I'm very I'm very curious if the if there is a bust what that looks like exactly and what precipitates it. I think it'll be a lot of fun to cover.
>> Well, one of the big questions
>> fun. I also think if I can say too, I think this is a particularly unique time because not only do we have these uh incredibly highly valued, ambitious frontier labs uh that are getting involved in these circular deals, but also we're in the middle of a memory shortage. um there's this unrelenting demand for compute and also this is a midterm election year and so I think you're going to see a lot of uh politicizing of the data centers too as uh you know sort of lawmakers try to appeal to their base because a lot of people are very unhappy about data centers.
>> Yeah. I mean, they're not pulling well at all.
>> No. So um you know when you think about where the collapse might happen um a popular thing for people to discuss is well maybe Nvidia which has been making such premiums on its hardware um it can't sustain it or it gets caught. So Lauren you spent a lot of time with uh Jensen Wong from Nvidia um what do you think about Jensen would enable him to sustain uh Nvidia's lead or do you think that some of these skeptics have a point?
>> Uh yes and yes I think some of the skeptics absolutely have a point and I think once you reach the uh the sort of uh is it the zenith is there the nadar I always get those two confused that Nvidia has um you know people are always sort of looking to to take you down a peg and compete but I think
>> Nvidia and Jensen has been incredibly good in Nvidia's history at sort of pivoting the company at exactly the moment that they need to in order to make sure that they've sort of caught the next wave. Um, and we certainly saw that happen with um, not only like you know the GPU to begin with and parallel processing, but then again with sort of pivoting towards crypto which ultimately meant they were in good place for AI and now we see the company doing that by addressing the inference market a lot more closely too and you know Jensen coming out and making these big proclamations that actually they're the biggest CPU maker in the world which I know Intel and AMD must be thrilled about. So, uh, you know, I think he's very smart and very strategic and that there's a good chance that they do maintain their dominance.
>> I mean, I think I think the question is like how much of the inference market they're going to get. I mean, it's it's whether it's 80% or 40% or somewhere in between or or 90%. I mean, I think that's what everybody's fighting out at the moment. I don't think there's a question that they're going to have some big chunk of it. Just it's just how much, you know?
>> Yeah. Yeah, I'm going to go to An in a moment, but uh Lauren, I just want you to tell us a little bit of uh what it was like uh with Jensen on a cover shoot for Wired.
>> You're talking about when I brushed his hair.
>> That would be it.
>> Yeah. Okay. Uh so,
>> wait, you actually No,
>> I really did.
>> Yeah. Yeah. I'll be brushing I'll be lining up later to brush people's hair if anyone would like to. Uh yeah, so I did a cover story on Jensen for Wired a couple years ago and uh our Wired art department is worldclass by the way. And so we had this big photo shoot set up and I asked if I could tag along to the photo shoot because I just wanted 15 more minutes with him to ask him some follow-up questions about Blackwell. And so they let me tag along down to the office in Santa Clara, set up this whole set for Jensen, who had like, you know, exactly two minutes to give us. And as the photographer was looking at Jensen under the bright lights, the photographer said, "Oh, he's got some flyaways and that like gorgeous silver hair that he has." And then he said, "Does anyone have a brush?" Silence across the set. No one at NVIDIA apparently had a brush. I was like, "Okay." Uh, and as anyone who has long hair knows, you always have a brush. So I said, "Well, I have a brush." So then I went to my backpack and got the brush. And then I said, "Okay, here's the brush." And no one moved. And meanwhile, Jensen is like, "What are we doing here, folks?" And I I said, "Oh, okay." So, I walked up to Jensen and I took out my brush and I got really close to him and I said, "Jensen, we're about to get a lot more close." And he said, "Oh, Christ." And um and then I brushed his hair and I have to say, I did bring some evidence of this. I think it looks great, frankly.
>> Oh, wow. I think it turned out really well. So, yeah. Now, the unfortunate the unfortunate conclusion to this story very quickly because I know we need to move on is that people were joking afterwards like you know the net worth of every individual strand of hair on that brush, right? Like like you could sell this on eBay and which I was obviously not going to do as an ethical journalist. But then like several months later, my back my backpack was in my it was my gym bag. It was in the back of my car and um downtown San Francisco, I think you know where this is going. My car got broken into and the brush got stolen. So those thieves have no idea the value of what they got away with.
>> They are cloning Jensen as we speak.
>> Yes. I'll see you in the back afterwards if you on your hairbrushed.
>> Ana, can you talk to us a little bit about what we've kind of hinted at at this point that there's this, you know, Nvidia, the common p perception on Nvidia is their GPUs were great for training, but you can actually do the inference or the act of using a model on a variety of of different chips. And once people once these labs train their models and they're happy with their models, most of the computing is going to go to these inference chips. Um, and therefore, even if Nvidia has a bet there, they're not going to be able to sustain their dominance. Is is that what do you think about that? Is that a potential flaw in the armor for Nvidia?
>> I think it is a flaw and I think there's if anyone's going to sort of attack Nvidia's dominance they're going to do it on inference like you said and there is sort of a massive effort underway right now to make all inference chips under the sun work well and so every single company that buys Nvidia chips and is spending a lot of money on NVIDIA chips is trying really hard to make these other other chips work whether it's in-house chips from Google or Amazon or even in OpenAI's case and potentially Anthropic's case you know do we make our own inference chip. Um, so I'd have a hard time, you know, believing that none of those chips are going to pan out, but you know, they might not tackle the bulk of the inference workload even. But then again, even if they do 10% of your inference, maybe you're saving enough money that you think the effort is worth it and it gives you negotiating leverage with NVIDIA. If they know that you have an in-house ship team, you know, Jensen's going to be a little bit worried when negotiating with you when you're playing hard ball with him. So, I think everyone's going to have to have an inference chip answer to Nvidia. But when I talk to data center companies about what they're seeing, um, you know, some data center companies don't really care what chip goes inside of their data center, but they they try to get hints from the companies that they're working with. Um, on the data center side, they're seeing a lot of NVIDIA and in the instances that I'm seeing where there are non- NVIDIA chips, it's because of some sort of financial backs stop. So, I do wonder how much how long that will have to keep being the case because that will definitely hinder non- Nvidia chips if they need special financing arrangements to get inside the data center.
>> Am I wrong in thinking the AI world is sort of separating on two poles? There's the Nvidia OpenAI pole and the anthropic Google Amazon poll, right? So, it's almost two separate ecosystems competing with each other.
>> I think OpenAI is investing a lot in non- Nvidia hardware.
>> Are they trying to move away from Nvidia?
>> Yeah. Yeah, they are. Um they're they're using chips from other companies. Um Cerebrris is a good example and they have their own in-house chip. Um so I think maybe publicly they're they're doing some big announcements with Nvidia and I believe it's 5 gawatts they have to deploy in the next couple of years on Ver Rubin. That's that's a lot and that that in itself might hinder them from doing more. But I think um I think Sam Alman wants to diversify from Nvidia.
>> Why?
>> Because you know just like any other company it's so expensive and I don't you know they have their own in-house silicon effort. I don't think they they think that they know their model better than anyone else like their their AI model and so I think they think that they're the best to develop a chip that can run their model efficiently.
>> Max, you're back and forth to Taiwan very frequently. um when we think about the position of Taiwan and Taiwan uh semiconductor in this world um you know the fact that it's in this like tenuous geopolitical place is something people tend to be like oh okay and then move on right away and I want to hear your perspective on whether uh the independence of Taiwan and the stability of TSMC is this like hidden black swan event that's just kind of in plain sight that people are not paying attention enough attention to.
>> I mean, it it is uh for some reason the modern world decided to to decided to put all of our chip manufacturer or most of it next to uh Kim Jong-un and the rest of it is in China is in is in Taiwan. I I don't really understand who decided that or why we decided that, but that's that's the nature of the beast. Um chip companies, the design companies in the US do not plan for this. The contingency the contingency plan, excuse me, is something along the lines of like, well, we're all kind of screwed if China invades. um which okay, but like there's no real plan there. Um so, and when I say invades, I don't necessarily mean like a literal invasion. I think what's a lot more likely is some kind of soft power exchange like what happened in Hong Kong over time. Um I think realistically that's that's a more likely option. Although I'm not saying Xiinping won't just invade. Like he absolutely will and he said he would. Um which could be catastrophic for TSMC and the modern world again. Um, I I do know that chip companies, as I said, there really isn't much planning and much consideration of it because it's it I think people in the industry just think it's an event that's just so crazy, complicated, and out of this world that nobody's going to do anything about it. So,
>> Max, do you want to talk about what happens when you try to visit TSMC?
>> They send me to the gift shop.
>> Uh, so
>> you're in Taipei and that's where you go.
>> That's Yep. That's basically it. They they got a a museum where they have like um I don't know a bunch of chips that they've made over the years. One of the cerebras chips is in there. They're very proud of it. The dinner plate size one. Um and I've never been to the museum, but uh that's where they wanted to always that's where they always want to take me. I think uh Wired is one of the few organizations, news outlets that's actually been able to go inside of one of the factories. Um and they wrote an interesting story about it. Um but
>> yeah, I do not get to go to the TSMT Fabs unfortunately. They're not they're not all that interesting inside. like they're cool if you've never seen one, but um they mostly look the same. It's like a bunch of white boxes with robots moving around. Um so it's not like people that know what they're looking for. It's helpful because you can like count the number of EUV machines in there and figure out roughly how many chips they can make and like do that kind of thing and figure out what kind of tools they're using. But um if the tours that most chip companies give are not particularly enlightening um in general,
>> yeah,
>> they are fascinating though. Yeah, they're cool, but they're just they're not useful.
>> Yeah.
>> In in like a meaningful reporting sense or an investor's, you know, from an investor's perspective or something like that.
>> I had gotten very excited about TSMC. Did a ton of reporting, spoke with former employees there and then got the courage up and I called them and I said, "I'm ready to come to Taiwan." And they're like, "You can come to the gift shop." And I was like, "All right, well, I guess I'm not coming."
>> My understanding is that our freelancer who got in there spent about a year negotiating with the PR team to get in.
>> I've been working on it for three. So if she has any tips like I could I could definitely
>> call up Virginia Hefernon.
>> I will. Absolutely.
>> Lauren, you've been inside a fab in Arizona. What's it like?
>> Yeah. Well, like this was the um Intel Fab 52, which is their most advanced fab. It's a 2 nanometer fab and it is on the level of what TSMC does in terms of 2nmter, but of course at a fraction of the scale of what TSMC does. And to Max's point about, well, what is our redundancy or our resiliency plan in the event of an invasion of Taiwan? It seems like right now the US government is very interested in what Intel can do in its fabs, but um it's pretty small. Uh and so yeah, I went to Chandler, Arizona. Um saw Fab 52. Um you know, got the bunny bunny suit on and um like Max said, it's it's interesting to see uh everything is roboticized. It's like if you had this um utopian fever dream of what the future of manufacturing looked like and everything was white and everyone was dressed in all white and everything's like being you know moving around. They call them foops that actually carry the silicon around above head and then they just sort of automatically go down to get like etched and carved and stamped and everything. Um it look it looks like it feels and looks like science fiction. But from a reporting perspective, oftent times the companies will cover up the names of the vendors on the different machines and you may be able to get some insight. Um, for example, there there were ASML UEV machines in these giant bays the size of school buses, but then you might see some spaces where they're like, well, we have two more coming and we're waiting for them. And so you're like, okay, what does that mean? you know, UEV is obviously a very um you know, a a very coveted sort of technology and they're also very expensive and you know, so so you get some insights from that, but uh primarily it's just kind of um it's good for context for better understanding how this all works,
>> right?
>> I was not allowed to go to this tour just for what it's worth.
>> Max, I think you're doing something wrong.
>> Max, you got to work on your emails.
>> Um so,
>> just kidding. That means you're doing actually a great job.
>> Thanks, Lauren. Uh there was an there was an interesting uh moment in the Darkesh Jensen interview which I guess will come up in every session today. Um where where basically what Jensen is trying to get across to Daresh is if you do not provide chips to China um they're going to be there's going to be a constraint. They will develop their own models on their own chips. They could become the global standards and then put export controls on the US. Does anybody here think that that is a legitimate concern or do you think Jensen just wants to sell chips in China? wants to sell it to chips in China.
>> Okay,
>> I mean it might be a legitimate concern in some sense, but Nvidia has run out of places to grow. Uh they can't sell to any more countries. They can't sell to any more hyperscalers. They can't sell to any more even they're going after the enterprise market, which is messy and complicated and difficult to actually make a lot of money in. So, China is a big market that has basically no access to Nvidia chips right now. I mean, you know, if you're getting those kinds of gross margins, like, yeah, of course, that's like an obvious just a very obvious place to to be able to sell chips, the company's kind of obsessed with it at the moment. And it's why Jensen goes to Washington all the time. That's why he
>> insisted on going to China. I mean, it's it seems pretty clear to me.
>> Anisa, any thoughts?
>> Max spends a little bit more time thinking about that than I do. Um,
>> I I do understand, you know, obviously he wants to sell to China. It's a huge market and they're taking a hit because they can't sell to China. So, you can't really answer the question without acknowledging that. Um, but I I do think that, you know, from the national security perspective, um, Chinese companies are developing AI chips um, no matter what happens. Um, so, you know, whether we can or can't, I do think that they'll continue to develop alternatives to Nvidia.
>> I mean, their manufacturing process is is a lot worse. like the one of the uh research firms just published an analysis of the sort of one of the new Chinese chips like homemade um by Smick and it's like it's good but they're they run up to up to the limit of like they can't buy EUV machines um the big ones that Lauren was just talking about. So if you can't buy those things it's just fundamentally difficult to make transistors below a certain size and they can't do it effectively.
>> Yeah. And to Max's point, Nvidia only has so much more room to grow. I mean they're at the point where they have become a major investor in Corewave and then Corewave is using that money to buy Nvidia chips and so it's it's you know it's really and there was an earnings call um earlier this year maybe late last year where uh Nvidia said that they had a domestic buyer of some of the chips that they weren't able to sell to China though they not disclose the buyer and so they're at the point where they're basically yeah trying to sell them wherever they can and so you'd have to imagine the motivation there is they want to sell chips to China
>> no those leftover H20s that they were Yeah. China that somebody else wanted them cuz there's people want silicon right now so badly any any kind will do.
>> I think they sold about was it $600 million worth of those?
>> Yeah.
>> Yeah. Okay. So, a minute left. Let's do this an actual lightning round. So, yes or no? Uh do you think that uh are you all bitter or less pills? Which is basically like the reason why there's all this infrastructure is because the AI labs think the more you know GPUs and chips you string together the better the models will get. Do you think that it will continue to improve as these buildouts uh you know continue to to explode in size or no?
>> Wait, ask the question once more.
>> Is is it is it the investment worth it in in these data centers or are they ultimately like not going to have much better models even though they have a lot more chips?
>> Are we going down the line here? You start with us. So you start with me. Um generally yes.
>> Okay.
>> Mhm.
>> Max.
>> Uh not not sure. Sorry. Okay. Okay.
>> I think it depends on the amortization what ends up happening there.
>> Okay. Ana, you announced your news on Twitter, right?
>> I did and LinkedIn.
>> Why don't you tell everybody where where you're going next?
>> Oh, um July 6th, I'm starting a new job at the Wall Street Journal.
>> All right, let's hear for Ana.
>> Thank you.
>> Congratulations. And thank you, Lauren and Max, for joining us as well. Let's hear for them, the whole panel. Thank you guys. Thank you very much. So, so often the story of AI's trajectory is is told without the people using it. We hear about these concepts called token maxing that are talked about in the abstract and then one day you see a chart and you never actually speak to the people spending the tokens. So I think to fully understand how this technology is progression progressing and its potential to meet its the way that people want it to go, we actually we actually have to speak with the people who are building with the tokens. And that's why I'm thrilled to bring on Dallas Dolan, the the TMT leader at PWC, who is actually implementing this AI, looking at the costs, and making sure that it's worth the money that they're spending on it, and will take us deep into the token maxing and the budgeting conversation. So folks, let's give it up for Dallas Dolan of PWC.
>> Great to see you, Dallas.
>> The tables have turned. You're interviewing me. This is good. Uh what do you think about Aaron Levy was here like 10 minutes ago and he talked about how token maxing was a BS uh um media narrative effectively and it never really happened.
>> Do you agree with that?
>> I don't know if I totally agree with it. We were actually talking backstage just before he came on and um I think there's absolutely like we'll call it above above the you know above the plane sort of commentary that's out there that uh we're all seeing in in both the mainstream media as well as like what plays really well on on social media on X and other places um and in podcasts. And then there's there's the reality within a lot of organizations, but it's it's happening enough where the behaviors, I'll call them, are slightly problematic from a cost and from an ROI point of view that it's real, right? You can't say that every circumstance is a problem necessarily, but it's coming through in a way that, you know, it's enough to think about and say, hey, are we doing this the right way, right? Broadly speaking from an enterprise strategy and also then from even a broader ecosystem point of view.
>> Yeah. So, there's this moment now where we're seeing actually a counter to token maxing. Uh, it's called token minimizing. And you're seeing companies like AT&T and Meta um get really serious about cost. We also know that Uber of course uh spent their entire budget in less than half a year. Um, who do you think is going to win out at the end of the day, the token minimizers or the token maxers given the definition of token maxing you you just gave us?
>> Yeah, I mean I think here here's the good news. The good news is I don't I don't think there's a winner and loser that's going to be defined by did you token max or not token max. I think the winner and loser is going to be defined by did you outcome max or not. Um it's in part going to be a function of how did you incentivize people which goes into that leaderboard and the things that got a lot of folks will say in trouble or certainly in the news. Right? So there's that piece of it. It's the how do you actually want to encourage people to do it without encouraging the wrong things. It's take it too far etc. Right? So there's that piece and or spend too much money. And then the other part of it is going to be actually I think from a planning point of view within an organization and I deal with this as well. So I sit on the the the boards for for our US and and our global organization at PWC and we talk about this a lot. In fact we talk about this even contextually from a comparative point of view within different industries. So it's not just saying hey does one organization as a services company spend more or less than another but also how are we compared and spending you know against let's say like a tech company who's you know got a bunch of engineers and doing the coding there. And I think what we're looking for is a what's like the baseline, right? So it's it's benchmarking with industry out with within the industry, outside the industry and against a given benchmark and saying, okay, are we actually getting an ROI? Are we way outside the bounds in terms of what we think the spend is or what we know the spend is, right? Some of the companies that you've named are certainly companies that we're aware of and work with too. Um, so we have some decent intel on what's happening within the industry from that point of view. But it is going to be that like what outcomes did you get? And I will tell you, I've seen some things that people have built, for example, in the deal space that's just incredible. Like really, really amazing output. It's taking, you know, thousands of hours worth of work and creating a product within seven minutes. The cost is very high, but it's an amazing output and it does show very well from a consumer or from a, you know, a customer point of view. It's also very exciting for, you know, the folks who work in the industry who don't have to stay up all night maybe working on something. That's great. There's other processes that people haven't even attacked yet. And I think the question is when do you actually go after each one of those and say how am I going to outcome maximize for each one of these processes. That's actually when you're going to see I think the benefits come through. There will be maximization of spend but it'll be maximized in a way that you're actually getting an outcome from it not just because you spent the most money.
>> So you're already doing ROI calculations.
>> Absolutely. Yeah.
>> So you're probably far ahead than most folks. Um how do you determine the ROI and and what percentage of your projects would you say are actually generating a positive ROI?
>> Yeah. Well, I'll tell you this. I know I mean even from an external point of view, I think MIT did did a recent publication talking about like what sort of uh activities are really, you know, replaceable just with generative AI in general, especially like in the vision and and you know, human interactive space. And I think they came up with a an outcome of about 23%. Um, as in you wouldn't use a human to do 23% of the work. I look at that and say in our business it's probably not as high as 23%, but it's going to be some percentage of every single thing that people do, right? So it's it's a different mathematical equation. It's is absolutely going to be measured for our business. It's going to be measured in hours. The same thing that's going to be done in the engineering space for a lot of these companies as well. It's a measurement of hours. It's are you making the person more efficient? Are they creating more output, you know, for the amount of time they're spending on it, right? So is that lines of code they're producing or is it quality product? Which this is an interesting thing, right? like I could build a larger slide deck or I can write way more lines of code, let's say, but am I actually getting a product at the end of the day that people are saying, "Oh, yeah, that's something I'm willing to use." And I had that conversation a bunch actually at uh at Tech Week in New York a couple of weeks ago with a number of a number of founders and also a number of the investors right from Andre and others who were talking about the types of companies that they're investing in. I think it's really interesting, right? If you're an investor, you say, "No, like you know, what does your product do, you know, for the person who's using, especially if it's a coding product or something along those lines." If you're not indicating that you're able to help them actually build a product, right? Come out with something a product or service that's better that their customer is going to want to use, the fact that it produces more lines of code or a longer slide deck or a longer pitch deck or a longer memo is not better. Right? I think we we're going to come to a determination, Alex, really soon here where we start to say, "Oh, wait a second. Like, what am I really trying to get to here? Is it more more more more or is it actually a better product that I'm trying to get?
>> Right. But that sort of goes to the way that the foundational labs and the cloud players are working with you. The things that I've heard is that these companies are making it so that when you plug into their systems, you're going to spend a lot of tokens
>> and they don't really want you to be able to measure them. Um, are you finding that?
>> So the I think the big counter to you know the the um maybe the the marketing and sales approach there is in the control plane that's being used to actually help companies make decisions on what model's being used or what interface is being used to do specific activities. And this is where this concept of central planning and I say that as someone who's who's traveling to China next week. This concept of centralized planning within an enterprise is going to be so important. It's the you are not allowed to check the weather five time five five times a day using you know clawed 5, you know, 5.7 or whatever it might be, right? Like it's unacceptable, right? Mythos is not a good use case for weather checking. Um even if you're worried like I was last night about the tornadoes going through the Midwest and trying to get back here so I wasn't late for Alex. It's still not a good use, right? You can do that still with a weather app or if you really want to use something inexpensive like there's a lot of those options out there. I think what's going to happen, you know, you the the term control plane is sort of out there. It's relatively newish. Um, it's got governance elements, it's got cost control elements, it's got access to data elements, it's got um, I'll see even like the the human interaction elements like what can you put into it, not only what you can get out of it. that is going to be the the you know the the surface of place so to say where people are going to spend a lot of the time in the energy at the AT&Ts and the metas and the services companies and a lot of these other spaces including in the engineering space but also in the sales in marketing and in in the back office as well as they look at other activities that they will do better but they will do better with the right tool. quick analogy, right? Like the fact of the matter is we are simply not best suited driving the Lamborghini to go pick up milk. Those two things just don't align, right? Unless you live in South Beach, in which case it's totally okay. Um, but those are the crypto people. They're not here. So, um, but in this case, like that's just not that's not what we'd want to encourage. And I think that that should eventually proliferate and get to everywhere, especially as the costs have gone up. And the last time we got together a few months ago was right at that click point of cost. and that you did a great thing with the audience talking about, hey, how much are people willing to spend which is a super fun exercise which I hope maybe you'll do later.
>> We could do we could do it now. Can we put let's put the house lights up for a second? Um I am curious. Sorry, we're going to ask you to vote. Um of of the folks here,
>> would you be willing to spend double the amount that you're spending today for the current capabilities that you have uh with AI? That's it. How many people are satisfied with the price that you're paying? So, if if they really Okay, sorry. I'm going to go with more more questions. If they let's say your Claude subscription doubled in price, how many people would cancel?
>> Is it fable or not?
>> Is it fable? Well, right now, no. No.
>> Yeah.
>> I think I think Okay, we could put the lights down again. I think basically what it what it proves is um we we saw about half the room's hands go up that they would pay double. Um, and it does seem like there's a bunch of room for the labs to be able to raise prices and still do well. But Dallas, something interesting happened since we spoke last. So, we spoke at Google Cloud Next about this and um, someone asked about the the margin the I think Dallas, you asked whe whether the margin of the business can be maintained at the current prices. And I thought, okay, well, forget about it because these labs have such an economically valuable >> uh uh tool for us that they'll raise prices. But we might be in the moment where they're going to get into a price war because OpenAI is rumored to be potentially uh dropping prices. And so what do you think about that?
>> I I think we're absolutely right on the precipice of that. I think you actually saw in the audience here by comparison,
>> we saw every hand go up when you said, "Would you be willing to pay double when we were together 3 months ago in April?"
>> And when when Alex asked if people were willing to pay four and five times as much, there was still a quarter of the hands in the room that were up. And we're talking a room of about 200 people roughly, right? So there was a lot of people. was a good, you know, good good tea sample, so to say. Um, I contrast that to what we just saw right now, and I think it's a a fairly, you know, even distribution, similar subset of people. Um, and the reality is there's there's more skepticism of value that they're getting from it, especially when you start layering on the access and the capabilities associated with some of the models that are still per seat, as well as some of the open models, which you can get access to and, you know, for free. You can do a lot of really cool things.
>> What would you do if the models were half price or tokens were half price?
>> What would I think would I do something differently? Is that the question? Yeah.
>> Let's say I'm like Open AI or Google and I come to you and say, "Dallas, we're happy that you're using our technology. We want to make sure that you don't use Fable when it's back online. So, we're cutting your prices in by 50%."
>> Yeah.
>> Would it change anything you're doing?
>> Absolutely. I I think actually like I know this because I talked to my CEO and CIO yesterday and my CIO again this morning. Um and we are enterensitive environment, right? We have 350,000 people globally. Um they all have access to one tool or another. And so when you have 350,000 people and they're playing with tools, there's a high degree of price sensitivity just on sheer volume alone, right? even if not all of them are doing super productive things. Just to be clear, those aren't all engineers, but they're all pretty smart people doing pretty interesting things and they have use cases that they're like, "Hey, I want to play with this." And we are encouraging them to do that. We think that's great, but we're also price sensitive, right? There's an elasticity there that says the higher the price goes up, the less we'd want them to use that model. We'd want them to use something cheaper. So, if somebody was coming to us with a cheaper model, 100% the direction of travel would go, you know, to that. And we've even built that into some of our control plane technology. So there's selectivity of model and there's also recommendations within it as well. So depending on who you are and what it is that we think you're going to do with it, it automatically configures to have a specific model and you
Have to break the the the, you know, the initial setting in order to use something different. So we're already there, but I'm certain we would encourage more usage of cheaper things. There's no doubt.
>> Have you gotten any of those type of calls yet, or you're waiting for them?
>> Um, I, I, I won't go too far into it, but there's definitely conversations going right now. Yeah.
>> Very interesting.
>> Yeah. Um, when we spoke recently, you told me that you're seeing some limits with agents. And it seems to me like if we're going to see this continue, agents can't be limited. They have to be able to operate autonomously and spend all those tokens and be effective. So, what are the limits that you're seeing with agents today? And do you think that the, like, if we could extrapolate a little bit, it means that we're going to see, um, some more speed bumps as the labs try to roll this technology out further?
Yeah, I don't, I use, you use the term limit. I, I think it's, um, I think it's a function of both risk tolerance, um, as well as cost tolerances, and then finally, like, what is it you have an expectations of these things doing on their own? And so the limitations are actually in all, in all three of those areas that are coming through.
Um, you know, from a risk tolerance point of view, I think people are saying, "Wait a second, I'm worried that the agent without some level of, you know, call it control or governance around it, could go just about anywhere." And what does that mean within my organization, depending on what access I give to it from a data perspective? Um, you know, I would say from client data, or whether it's, you know, it's code itself, and what can it do to change code? If you ask it to do one thing in one area, will it simply think that it needs to do that everywhere else? And and there's a, you know, we'll say the ability to extrapolate on a single point and like, what control exists there. So that's the first limit.
The second limit is on, like I said, on the cost variance. We went through there a second ago, which is, "Hey, look, like there are just things you're not going to want it to do because, back to the MIT study, there might be things that humans can do not only better but more cheaply, especially now, depending on, you know, if if you have access to Fable and you're using that at a super high premium."
Um, and then the third bit, you know, is is really in that that decision-making, we'll call it from an organizational point of view. And even, I, I even think it was funny. I, um, uh, of all the, of all the places to be here in San Francisco, I was actually at a funeral this morning for my grandmother, who was born in San Francisco. Actually, she was literally born right up the street here. And it's such a human thing, right? And I was talking to a priest, the priest who was there. He's actually the priest who wrote the paper, um, along with some of the team over, over at Anthropic. And we're having an interesting conversation on like, where does ethics play into the broader conversation of what's happening within this area? It's sort of like it's workforce planning and then some. And so there's this like additional, you know, component that I don't think we've even really gotten into because here we are at the center of technology. We don't think about that part of it so much. But it is that like benefit side. The benefit, I think, has a risk that we haven't even really gotten into a ton, which is, you know, "Do I want to scare all of my people that I'm watching what they're doing and waiting to replace what it is that they're doing?" And the answer is like, "Absolutely not." Because in fact, I met six of our interns on the plane back from Chicago last night, and they couldn't be happier to be interns with us.
>> Uh, they just saw something on my t-shirt that said...
>> You're still hiring interns?
>> Absolutely. I mean, it, it, you know, we're hiring as many people as we did last year, and we're changing the, who we're hiring. We're actually hiring a lot more kids who are pursuing sciences and, you know, even in areas that, uh, that go beyond, you know, engineering.
>> And where are you not hiring?
>> Um, I think we're hiring a little bit in the, less of just like, you know, pure accounting. Also a function of where we're doing it. It's not that we're hiring less people overall. We're hiring less people in certain areas. U, but that's also just a function of globalization and how we're delivering services. Um, but it's not really a function of the fact that we're saying, "Oh, we don't need these people because technology will do a lot of it." And that then goes back to your original, you know, premise of the question, which is, you know, "What are the limits that are there on the agentics themselves?" The limits are going to be a tolerance for error, a tolerance for cost, and a tolerance again for like, "What will be acceptable within the organization of having that thing do versus having, you know, a human do it?" And there'll be a cost dynamic to that part of it as well. But I think there's also even just the, "How do I run my business and how am I a positive leader, right, within within a given community?" Doesn't matter what type of company it is.
>> You know, Dallas, I, I want to take a moment just to acknowledge your, your grandmother, and I'm sorry about her passing.
>> I appreciate that.
>> She, she spent her life here in San Francisco.
>> Here in San Francisco. Yeah. Yeah.
>> Can you tell us a little bit about her, just, just briefly?
>> Uh, sure. I mean, you know, funny, uh, you know, daughter of, of an immigrant family. Um, people who picked a lot of things in a lot of places around the world, including Argentina and, uh, and Hawaii, funny enough. Um, and thought, you know, "Picking stuff out of the ground isn't a good deal. Like, we should move to San Francisco." And we're working in the canneries, um, up in, in Fisherman's Wharf, which is what her parents did. She actually worked in the telco industry. The first job was actually working for company for Bell. It must be, it must be the second T in the TMT, right? So, for, for Grandma Fernandez Corteho. Yeah. Like, really a cool existence and just a, a great person and a great, you know, um, kind of full circle story, right? I mean, I think this whole story of technology is such a cool thing that all this is happening in one. I mean, let's be honest, it's all happening kind of in one place here in, in, in the Bay Area in San Francisco. And it's, it's really neat to, to be able to like see that thread all the way pulled through, right? Like I've been able to do from a career point of view. Um, and also be able to notice that there's actually a direct connection to the things that made this place a great place 150 years ago, or actually still the things that make it a great place because it's still better, right, to work in a cannery or work for a company than it is to pick something out of the ground, so to say. Again, as my, as my ancestors, you know, concluded, you know, back in the 1800s. But when you think about that, I, I think it does go back even to some of the things of like, "What do you want to do, right?" Like, I think you want to create technology that makes people's lives better, in much the same way that our grandmas make our lives better, right? Grandma makes everything better. I think it's the same kind of concept. The things that we talk about every day are, "How do we like build products?" You know, what are the entrepreneurs, they're trying to build products that make their customers, you know, life better, or they make something easier to do, whether it's coding, or whether it's, you know, customer support, or what have you. I, I think we, we take that thread all the way through. That's actually where the ROI comes in. It's not going to be back to the token maxing. It's not going to be how much you pay and how you pay it and what have you. It's a whole ecosystem shift in the way that we think about the outputs themselves. The outputs again, still are, "Am I serving my clients well? Are they getting, you know, deal maximization out of what, uh, what they're doing with, you know, with any given company, Frontier firm, or hyperscaler, or, you know, Neocloud, or whatever it might be?"
>> Well, thank you for sharing with us.
>> Yeah, thanks for that. Yeah.
>> Um, you know, interesting San Francisco. One of the things that I associate with San Francisco is it's a city that where people are okay losing a little bit of control. Um, whether that is building, you know, new technologies that do things a little bit differently than, um, previous generations, or the fact that this city, or not everybody, but many people love, um, LSD and mushrooms.
>> Um, and...
>> You said you weren't going to mention that in our precall. Okay.
>> Um, it is interesting with with agents, you do, you kind of lose control. It reminds me a little bit of like of skydiving in a way, where you jump out of a plane and you're like, "Whatever happens, but I hope there's a system ready to catch me." Um, from your position, you're, you know, you're deploying agents, um, in pretty high stakes moments. You, even if you have the best governance in place, you have to be okay to a degree of losing control. So, how do you become comfortable with that?
That's it's funny, right? I mean, I think, um, how much can you really control, like in the, in the engineering space where you have the people like doing the code for you, or in our space where you have the individuals doing the services, they're preparing a, you know, an audit or a tax return or a deal report or what have you. We're putting trust in these folks, many of whom you know at like a superficial level. Um, when you start layering technology into it, because my view is that it's an augmentation of those people to make them better, not necessarily seeding all the control from them and putting into something that I know less of, that makes me feel a lot more comfortable. If you start getting into the space where a 100% of all the things that we do around, I don't know, let's say like booking travel becomes, you know, commodified, right? You just put in the query and the travel gets booked. I think that's where you do get uncomfortable. Like, I appreciated that two nights ago, knowing I had to get back to hang out with you, Alex, and to make Grandma sing this morning, like, "I can ping my EA late at night and say, 'Hey, I I need some help. Like, I need to make sure with tornadoes coming through Chicago, there's going to be at least a plane on Thursday morning that takes off at 6 AM that could get me to the West Coast.'" If I'm pushing that into an an agent and I'm hopeful that it works and I'm hopeful that it books a flight, I mean, yeah, I might see the outcome, but isn't it neat like having the human on the other line like Liz says, "Hey, I got you, bro. Like, like you're handled."
>> Um, she's augmented. She goes into our technology and can quickly query something and boom, pulls everything up and she books a flight within like 30 seconds, right? But having her there to make that a little bit better, like, does make the whole thing feel better. And, you know, candidly, by the way, as it relates to like, what, you know, what my EC does or my my chief of staff and others do, like, they're able to do like multiple person's jobs. If you think about like, what it was before, like, you know, one person to one person. No, it's one to like 12. Like, we're doing a lot more with a lot less. So I look at this tech is just being an extrapolation on that point. Yeah, there's some things you're just not going to stop doing, but like, that's totally okay. Um, it's, it's no different than driving, right? It's like the self-driving dynamic. Like, we are going to get really comfortable with that. I see this, you know, move to Agentic as being, you know, very, very, um, uh, similar, uh, in a parallel run with, you know, with that itself. It just so happens to be in the physical space.
>> Yeah, it definitely reminds me a lot of a whimo. You sort of, you know, you white knuckle it in the beginning, and then you start to go on your phone.
>> Exactly.
>> Yeah.
>> Whole thing. Uh, anyway, don't take any, uh, anything away from those previous comments, um, about the mushrooms.
>> That's for later. That's happy hour.
>> Please join us on the roof at 5.
>> Exactly.
>> Uh, Dallas, thank you so much for being here with us. Always a pleasure to speak with you, and I do appreciate, uh, your support and PWC support of this event. So, thank you. Thank you very much. Let's hear it for Dallas. Thank you.
>> Thank you so much. All right. Amazing. Uh, thank you, Dallas. All right, folks. So, we are going to have this next session. It's going to be an audience participation section. So, we definitely want your questions for us, and then we're going to go to a coffee break. So, um, though for many of you, he needs no introduction. Let me introduce Ron Roy. I first started reading Ron Johny's, I first started reading Ran Johny's market, uh, writing in 2021. Um, he had written this newsletter called Margins, and I thought it was a terrific newsletter. I saved it. I spent my winter break reading it. I DM'd him. By that January, uh, we had decided to do an emergency podcast about a crazy financial situation, and then Ranj and I kept talking more and more. Uh, and then by January 2023, I wrote to him and said, "Hey, don't you want to just come and do this every week?" And lucky for me, and lucky for us on, uh, with with anyone involved with Big Technology Podcast, uh, Ron John said yes. And so getting a chance to speak with Ron John every Friday is an absolute joy. It's definitely one of the highlights of my week. And today, we're thrilled to be able to do our Friday show live here with you, with your, with your audience questions. Um, and then we'll just run it tomorrow like a normal podcast. So, I hope you're ready. Uh, we definitely need your participation. And please join me in welcoming Ron Roy.
I got them on. I'm going to take these off, though. We'll get into the Snap spectacles. This is, this is a medium risk maneuver because I have this microphone on. So...
>> Those are the Snapchat spectacles.
>> These are the Snapchat spectacles. They're the original developer beta edition, though. So, they're not the new ones, but I had them in 2021.
>> Now, everyone knows how cool you are.
>> All right. Should I start the...
>> I'm as cool as Evan Spiegel.
>> That's right. Yeah.
>> Is he still cool, though?
>> All right. Uh, let's do it. Um, how would I start it? Um...
>> I throw... Did I throw you off?
>> Yeah. No, no, no. I didn't even write this down. Okay. Well, I'll try to do it. All right. Snapchat comes out with new spectacles, and we take your audience questions. That's coming up right after this on a Big Technology Podcast Friday edition, recorded on Thursday.
Welcome to Big Technology Podcast Friday edition, where we break down the news in our traditional cool-headed and nuanced format. We're joined, as always, by Ranjan Roy, who is here with us in, not in studio, live, uh, with the Big Technology AI Summit audience. Audience, let's hear you. The way this is going to work is Ranj and I will break down, uh, one story, and then we definitely welcome your questions, your prompts, your arguments with us. Uh, if you don't have anything, we have plenty to do, but, um, we'd love to have your participation. You can line up at either of the mics there, uh, to ask us a question. Let's start with our top story this week. I promised myself when this, uh, summit was being planned that I would not be mean to Snapchat. However, Snapchat has left me with no choice. John, let's roll. Image C, please. Uh, listeners, if you are listening on the podcast, you'll see what you're, what we're looking at at the audience here is an image of Evan Spiegel on CNBC wearing the latest Snap Specs. The headline is, "Snap Stock Falls After AR Specs Debut." Um, uh, almost 20 years since the launch of the iPhone, people are ready to think about computing differently, Spiegel said in an interview with CNBC. The, uh, the market reacted differently. Um, Rajan, let me just go to you quickly. Do you think that that, you know, Snapchat and Meta and all these other companies have been trying to build these AI device future for years, and this is what we're looking at?
>> This is what we're looking at.
>> You kind of ruined my surprise. Can we go to D, please? I mean...
>> I brought them. I had to bring them after sending Alex a photo.
>> Is it time for us to finally accept that we're not going to have an AR device?
Interesting. So, so these, again, I got these in 2021. It's actually a very cool technology. Like, it's, No, I'm serious. Augmented reality. If you ever used Magic Leap in the 2010s, like being able to paint throughout your room and walk around that painting, being a, being able to play games where you're chasing zombies. Again, my seven-year-old son actually is probably the only fan of Snap Spectacles in the world right now. He still loves using them. He also likes to watch YouTube on them, which is kind of amazing. But I think that form factor and the experience is amazing. Vision Pro has not quite captured it. I don't think what we just saw on Evan Spiegel's face is going to capture it. I think eventually, maybe Apple or someone will, but I don't think we're there yet, but I think we will be. I still am betting on AR eyeglasses of some sort.
>> Okay, I'm going to go to questions in a moment if anyone has one. So, feel free to to stand up there if you want, otherwise we can keep doing this. Um, here's some of the social media reaction. Uh, does anyone, anyone that works at Snapchat have the guts to tell leadership that these things are ugly? That feeling when your glasses are so heavy they give you cauliflower ear. Snap is the best brand in the world when you're 16 and the worst brand when you're to be associated with when you're 21. Um, the people who actually buy $2,000 AI glasses aren't teenagers. If you think you really want to wear an always-on camera around in public, uh, it should have to look like this, with the picture of Evan Spiegel. Um, let me make the case that this, it's over.
>> Good, because I'm going to take the other side.
>> I think that the iPhone series that released that we just saw. So, does anyone here listen to Ron John on Fridays?
>> Do we have listeners? All right. Okay.
>> So, you'll know which direction this might go.
>> What's this guy been begging for for like, since he came on the show the first time? Better Siri.
>> Better Siri.
>> I think they actually did it. Like, the new Siri, if you look at the videos, looks terrific. And so maybe this idea of, "We have to wear the computing on our face" is something that like, kind of sounds good in concept, but model after model, it's not. And the AI device is the iPhone.
>> Okay, that's, I, I, that is an interesting direction to take it. I still think the form factor. Do, do any of the audience have like Meta Ray-Bands or any other device like that? I see a few. Like, you start to feel as you're walking around, as you're kind of interacting in the real world, stuff can happen that's not just maybe eventually Siri talking to you and using a traditional AI model to actually, you know, give you kind of information. So, I don't know. I'm still, I still think AR as a form factor via glasses is going to happen. I think, uh, you have Meta Ray-Bands, you enjoy them. You know, I, I didn't want to say this publicly, but I have not used them. I, I, I, I mean, I have, but I don't. I used them for a bit. And like I said on our show recently, I was on a hike. It was cold.
>> I was ready to get to the summit and put those glasses on, and it, the battery started blinking red, and I couldn't use them. Had to use my old phone.
>> And let me say this, like, you know, they, they haven't taken off as a mainstream consumer device. And if you look at the stock of every single company that's pursuing them, it's not good.
Snapchat, like we just said, is is struggling. Um, Meta, as we know, has its problems. Um, no one's looking at the, these Ray-Bands to save Meta.
>> Well, no, but to me, the Apple Vision Pro is like the more direct correlated product versus...
>> You want to know what the best thing about the Vision Pro was?
>> Find Vision Pro. Yeah.
>> They put the person on Vision Pro on Siri, and he fixed it. Mike Rockwell. Oh...
>> So...
>> I would not have guessed that the person who made the Vision Pro would be the one to fix Siri after all these years, but I guess if that's the case, that's exciting. Yeah.
>> Yeah.
>> So, you're going to still, you're going to, can you put those glasses on one more time?
>> Again, I was told backstage this might destroy my microphone, but I'm going to try. Do I look cool?
>> No. No.
>> No. No. I mean, even we were, we were joking that Evan Spiegel is going to the Met Gala with his model wife. Like, this is like the coolest person in the world, and how bad they made him look. That if it was like when Zuck wore Project Orion, no one really cared that much. But I think because Evan Spiegel is such a cool-looking dude, that's why it looks so egregious. That's my take.
>> So, your, your answer on the way to make those things work is just be less handsome.
>> That's it. That's it. We'll, we'll write to Spiegel and let him know.
>> No questions. Okay. All right. Great.
>> All right. Yeah. Don't be shy.
>> Is it on? Yes, it is on. Hey guys.
>> If you're willing, let us know who you are. And...
>> This is Ser Johal. Actually, um, industry analyst flew from New York a day earlier to join you guys. So...
>> Welcome. Let's give him a welcome.
>> Thanks for asking a question about, um, actually, it's an observation. And I want to get your take on this. Um, when companies have this gap between what people need today and what they're working on, which is like, in the future, it can be two plus years out. So, they lose that traction from the investors' point of view, as well as employees and and partners. So, do you see that that's happening to Meta and others, like it happened to IBM when they were living in the future with Watson X, for example? So, um, what's your take on that? It can be B2C or B2B examples.
>> Okay, thank you for the question.
>> Yeah, no, that's a great question. I do think, like, if we're talking about kind of how the interface for how you interact with a computer. I have been begging for something else other than me holding my phone, looking at it, and we haven't really gotten anything for a long time. I mean, Humane tried with the pin. Um, I still think maybe some kind of pin is going to be around. Johnny Ives' pin at OpenAI. Maybe, maybe at some point.
>> Well, Greg, Greg Brockman is coming, so we'll put...
>> We got to ask him about the, talk about that. But, but I think like being able to interact with all of this information now, being able to like process information so much more reliably with AI, just I don't want to have to keep looking at my phone and just looking at it on there. But even actually on the phone, I'm guessing, do a lot of people here dictate more to their phone right now? Like, that's completely changed. In the past, like, I would have felt weird just talking to my phone, and now I'm constantly using WhisperFlow and just talking to whatever. So that's...
>> Tell a story about you and your wife.
>> I know this. Okay. I don't know if this is the most depressing thing, or, and my wife is not here, and hopefully won't kill me if this is being live-streamed right now, but...
>> No, we won't broadcast this.
>> So, I am constantly dictating to my computer, to my phone. And the other day, it was like Friday night, we had put our son to bed. We're both on our laptops, and I'm on one side of the couch dictating. And then I look over, and she's also has her laptop open and is kind of whispering to her computer, too. And I'm like, "Is this the future of tech?" But it's a new computing interface. So, I'm happy about it.
>> Right. I, look, I think the, the gentleman brings up a great question, which is that things are moving so fast. How do you plan right now? And honestly, I don't know how you do it, because every day there seems to be a new capability, then the capability is taken off the table. Like Fable, for instance. I mean, the tweets about Fable where like people are adding Dave Sachs and they're like, "Please, I'll do anything for Fable back. Just bring it back." And he won't bring it back. But it's just like, I don't understand how any company does that. And I think that actually would be a good topic for us to sort of get into on a future show.
>> Yep. Okay, let's go this way.
>> Hi, thank you for taking my question. Um, I was curious your views on in the next like five years, as AR glasses evolve, whether like the chunky spectacles is the way to go, or thinner glasses with like onboard compute, like either wireless to your iPhone, or like the Vision Pro like cable down to the battery pack and compute. Well, it's a battery for Vision Pro, but it could also have compute on board. So, I'm curious like the next five years, where you think consumers will gravitate towards.
No, it's a, it's a great question because if you ever use Magic Leap, there was like a...
>> Puck.
>> Puck. That's what it was.
>> My, here's my hot take. Anything, any device that requires a puck...
>> Not working.
>> Well, how about this though?
>> Okay.
>> What if, what if the puck is your iPhone?
>> Oh, [ __ ] All right.
>> Exactly. So, which gives Apple an opportunity that like, if the compute is taking place on the phone in your pocket, it allows you to be much slimmer from a power perspective. You don't want like a lightning cable connected to your face. But like, it's still, I think there is a, it makes me think Apple still has a good chance in this space because the iPhone can do all of the heavy lifting versus this thing on that is really heavy on your head.
>> Can I, can I ask, so what's your name?
>> I'm Kyle.
>> Kyle, so, uh, what do you want to use a, like, face computer for?
>> Oh, that's a good question.
>> See, I told you Ron John, this stuff is not happening. No, no, no. Kyle's got something.
>> All right, let's give Kyle an opportunity here.
>> Let me think for...
>> I didn't mean to put you on the spot, but thank you for helping me prove my point.
>> Oh, I think one cool thing would be like a shared, uh, like TV or something like, like imagine Vision Pro and you have just a shared, just like movie theater. You're on a plane with your family or something, and you just have like the shared experience, but it's just glasses. I like that. I like that. Like, for me, one of the coolest things I always like, cuz, you know, as you get older, you move further away from your friends, and it would be cool to like be in the Vision Pro together.
>> Yeah. Sit, sit half court, watch the Knicks next to Timothy Chalamet, but we're just in Division Pro. But, you know what's interesting? Apple never advertised that as a social device. All the marketing was just, you're sitting alone at home, and all you're doing is the Vision Pro. Nobody else exists.
>> The other big one. Yeah, sorry to interrupt. I was say, the other big one for me is just having...
>> A lot of monitors around when I'm working. Like, even though I have like an ultra wide or like three monitors, feel like sometimes I could have more. So, just being able to like interact with AI to pull up the exact page I'm looking for out of, like, I'm one of those people who has like a thousand tabs. So, just being able to pull up the thing and I just say, "Open up that tab," and it just pops up would be pretty cool.
>> Would 8,000 individual tabs in a giant planner space be better, or Kyle's going with this?
>> No, no, but I do agree, like...
>> That being able to do more...
>> Open scaled work and like look at different charts, and I think there's that still, and I know like I have friends who own the Vision Pro who use it for that.
>> Still.
>> No, but did...
>> Did they have a nice time? All right, thank you, Kyle. Let's go here.
>> Hey, what's up? This is Sasha from Yo University, doing research on AI agents for finance. First of all, love you guys. Listening to you guys chat on my way to work has been my routine. I think this is excitement shared by not just me, but many people in the audience. So, thank you so much.
>> Thank you.
>> Thank you.
>> Appreciate that. And thank you for coming. Did you fly in?
>> Uh, yes.
>> Oh my god. Thank you for coming. It's great to see you.
>> It's worth it. Um, and I appreciate your sharp line of, uh, reasoning and questioning. So, uh, what's your view on the meter benchmarks of how long the AI and AI agents can work independently being saturated? Are we at that point that, uh, of them being at infinite work number of hours working yet, or are we reaching that soon, and what is the world going to be looking like after we hit that infinite mark?
That is an excellent question. So, first of all, there is some controversy about the meter monitoring, but I think it's kind of directionally accurate, right? Like, I showed the meter chart at the very beginning of, of our event today, that like, you saw these models, they could not code autonomously for more than 30 minutes in 2023, and now they can code autonomously for the equivalent of 18 hours, right? Um, so what happens if they just kind of blow past that limit and then they can code all the time? Um, that is an excellent question. Do you have any thoughts?
>> So, I, I felt that with the go mode and automation mode, uh, it practically felt that they are already doing this autonomously forever.
>> Right.
>> Uh, I want them to send me reports. My prompt would be, "If you find something interesting, send me an email." Yes.
>> Then, when I see it, I will come back and intervene. I felt that for many tasks, it's already at that mark.
>> Yeah. No, a Perplexity computer will basically do that. Um, so that's that's definitely something that's happening. But I'll just say one more thing, and then we'll go to Ron John. Um, Greg Brockman, who will be here later, has this like idea of a compute-powered economy, where like, you know, you sort of, once you get these bots working the way that you explained, you just, maybe he'll say this later, um, you just kind of throw them at any problem, and then they just kind of work autonomously through it. Now, we don't, we've never seen what the world looks like when you can do something like that, but I do think, given the progress that we've seen with the models, you would imagine that something substantial will come at a certain point. So, uh, my day job, I work at a company, Writer. We're an enterprise AI company. And the, when, and the, the goal mode was asked about recently, again, the idea that you just provide the goal. The agent will iteratively loop and keep doing work until it finds the right solution. In the real world, that is like the benchmarks versus the real world. I still think there's such a massive gap in terms of what does the data look like? What, what is the actual problem? Does the customer or the person actually understand the goal in a clear enough way that they're able to define it to kind of push that loop forward? So, I think it's an interesting. I think with a lot of AI, again, even around the benchmarking, like, again, versus real-world understanding of how to use it, what's happening, and what's available for it to use. There's still a lot of, uh, work to be done. I mean, you, you can do these goal modes for like less complicated tasks. I like...
>> What have you goal-moded?
>> Uh, so I worked with Perplexity Computer recently to, um, to try to like have it find a hotel discount for me.
>> No travel stories. I, I've been on a, this has been for like two years. Every time I remember, like Sundar, actually, it was, I don't know if people remember, in 2019, my example. Wait, wait, no, no. I'm just saying. Why does everyone when they talk about Agentic talk about travel? I get...
>> Cuz travel is such a pain in the ass.
>> But it's, I don't. Okay, go on. Go on.
>> I'm, I'm, I'm sheepish now.
>> No, no, no. Hotel discount. Hotel discount.
>> I just looked at the hotel every hour, and when it dropped below a certain threshold, emailed me.
>> That's pretty good.
>> Okay, that's a clear goal. I will give you that as a clear goal.
>> This is...
>> I respect your, I respect. I think you're right that we definitely need better examples than...
>> It's a pet peeve. It's just for some reason, every Apple, the original ridiculous Bella Ramsey commercials around Apple Intelligence, of course, everything is flight booking. And again, like...
>> You have killed that Bella Ramsey commercial so often, they're going to the creative agency is just going to write to you at some point.
>> And I love The Last of Us, and I love Bella Ramsey, but those commercials still irked me. Yeah.
>> What happened in those commercials? Uh, one of them, she's sitting at someone that she can't remember who that person is, and in real time, asks Apple Intelligence, like, "Tell me about this person, my interactions," which again, is just such a weird thing. Like, "I'm so much better than you that I don't know who you are, and I need to remember who you are and have AI tell me." But, and it didn't even come close to working with Apple Intelligence, so that was that was the worst one.
>> You know what would be good for that, actually?
>> What? Goal mode?
>> AR glasses.
>> Oh, AR glasses. There's the real-world use case.
>> Or the fact that that commercial was so bad, just again, proves my point that, okay, AR glasses aren't going to work. Okay, sorry. Thank you so much for the question.
>> Thank you. And we'll go over here.
>> Oh, he was...
>> Okay, fine. We'll go here. Okay. Um, hi. Um...
>> Hey, Gerald Harris. Uh, I'm on the board of the Commonwealth Club here, and I, I run some, uh, programs for the club, but here's my question that I think hasn't come up here. Um, what, what, what should we be concerned about in terms of using the AI models, them building a database on us, and then turning that into advertising? So, the advertising potential revenue from some from users, do you think the AI companies will ever go after that revenue or use that information for advertising purposes?
Oo, that's a good one. So, will, will AI's build an amazing, uh, sort of profile of you based off of all the personal data and then use that for ads? Most certainly.
>> So, someone left one of the companies and wrote about that in the New York Times about three months ago.
>> No, definitely. Yeah, that that's coming. Um, I, I think that, um, advertising is obviously going to be one use where we're going to start to see some of the problematic stuff here, but actually, if you look at the ads that OpenAI gave you, it's sort of like more of a, and obviously they, they always come out with the high-touch brand, and then they will like get you on the direct response, like super targeted stuff, um, once they realize they can't make money on brand. Um, but like, it looks pretty good. Like, if you're, sorry to go with the travel example again, but if you're like researching travel and ChatGPT, like, you can go into an advertising chat experience, and it will help you. But I, I think that like, we have never had technology ever that collecting this much information about us. Has anyone here, um, like talk to like ChatGPT or Claude and say, "Can you psychoanalyze me? If you're not, or give me any any insight about myself?" Yeah.
>> Just like five of you. The rest of you have done it.
>> Admit it.
>> It's scary, and we're, and that's the like, gated stuff, and we're giving this all this information to to companies, and it's a, a trust thing that we don't really know where it's going to go.
>> Yeah, I think it's interesting because again, OpenAI originally, ads are going to be a big part of the business. Now they've pivoted away from that a bit, but are still releasing ad products. At least Anthropic is not. I mean, I don't think they're going in that direction at all. Google and Gemini, obviously 100% will, and like, you probably will get a really good ad experience. I mean, the more it knows about you, the types of questions you're asking. But it is like, you know, like thousands and tens of thousands of three to five-word Google searches are a lot different than entire thought partnership, therapy, like research exploration. These models will know everything about you, and it is, it's terrifying.
>> Have you done the psychoanalyze me prompt?
>> No, but I did. So, I use like Claude and ChatGPT, Gemini. I'll kind of go through the three of them and cycle around. So, it is funny. So, I have asked like, "What do you know about me? Like, what would you ask back?" And they all have very different kind of like, jagged, jarring aggregations of who I am. Like, I don't know. They're, uh, I, I, I don't have one that just knows me through and through.
>> I, I've done it.
>> I diversify them.
>> I've done it with ChatGPT.
>> Yeah.
>> It's pretty good.
>> Yeah.
>> And then like, it will first give you a answer that's like, kind of sanitized. Then you say, "Go a little deeper." Then you say, "Get a little darker."
>> Oh dear.
>> I have not red-teamed, uh, Chat TV.
>> I challenge you all to give it a shot. Don't go darker. Don't go darker.
>> But if you really want to.
>> Okay. Thank you over here. And thank you for the question.
>> Let's move on. Let's move on.
>> Yeah.
>> Thank you guys, and really appreciate all the work that you do on the podcast and YouTube. Uh, your conversations go so deep and still very broad. So, I really appreciate that.
>> Thank you. We appreciate it. You're, you're a listener or viewer?
>> Oh, yeah. All right. Both.
>> Thank you so much.
>> Depends on where I am.
>> Amazing. Appreciate it.
>> Um, and, you know, my name is Si. I am an advisor for startups in the AI space. Um, and I've been in product management and go-to-market strategy. Um, my question for you, some somewhat related to the gentleman before, you know, there's a lot of coverage on agents and what agents can do, and, uh, a little bit about the governance and cost control. What I don't see a lot about, um, you know, is, you know, agents do drift, probably why we don't have better examples than travel. And, you know, over a period of time, they don't get better at that specific, uh, you know, uh, understanding of the people and the business context and the operating model they're working in, um, and, you know, learn from it. And I was wondering, is there, is there a lot of conversation that you guys are seeing around, uh, the learning of, uh, or the learning layer in the agents, and is it too early, or, you know, um, for that to happen? And the conversation is mostly around governance and cost and access.
It's actually a good question for Ronan.
>> Yeah. So, again, like working in this all day, and thank you for the question. I, I think what's going to happen, and I'm already starting to see, is like, six to 12 months ago, the approach was just jam as much information into whatever system you can. And we were kind of promised that it would just work, and AI is just going to, and like, there's, you know, you'd get this feeling that, "Oh, wait, a 100-page PDF, it actually could parse and I can get information out of," but then when you're doing that at any kind of agentic scale, it doesn't. And I think already, like, how information gets chunked up and like distributed in different ways and context windows. Everyone would see like, a 1 million, uh, token context window, but we didn't really know what it meant. And now, more and more, like, and I actually think this is going to be one of the most interesting, like, professional opportunities and areas to be an expert in going forward is like being able to kind of, it's funny, like, it is as much art as science right now. But I think like the more you can start to grab a hold of how information flows through these agents and how you actually get it reliable, it's going to be a really, really interesting space. It's not, it's no one's job right now.
>> Thank you. Okay, thank you for the question. All right, we have seven minutes, three questions. Let's see if we can keep the questions brief, and we'll keep the answers brief, and we'll get to everybody.
>> I'll try to be quick.
>> Hey everyone, uh, big fan of the pod, by the way. Chris Auei. I'm one of the co-founders of a company called Virus Watcher. I'm here...
>> With my colleague.
>> Flew in from Dallas. He flew in from Austin. Yes. A big fan.
>> We got a couple Dallas folks here today.
>> Oh, who's Dallas? Shout out Dallas. There we go. Okay. Go Matt.
>> All right.
>> Oh, yeah. We're big in Texas. Um, I'm going to go in a different direction, uh, with my question is, uh, I want to know, kind of, what y'all's thoughts, and honestly, if you can ask this kind of later on with Greg Brockman and others, um, what is the thought on using AI models and AI, you know, technology for biological intelligence, bio-surveillance, and public health for emerging risk detection with infectious diseases? It's a very interesting project that we're working on. We have a lot of epidemiologists on our team. We were just at the UN two weeks ago with the World Health Organization, um, and trying to discuss these problems and how we can use the technology to, to kind of, uh, detect these things.
early on. So I kind of want to get you guys' thoughts on that.
Can I ask you a question back? And if you could give a brief answer, that would be great.
Okay.
Are you a believer in the biocapabilities of these models like Fable? Like, do you think Anthropic made the right decision by gating them?
Biocapabilities meaning, I think you're talking about biological weapons. He's talking about like disease, correct? But I, I think that biological. His answer will give us insight into the.
I, I mean, I think it's a, it's a very touchy, you know, it's a, there's no right answer, I don't think. Um, but I do think it probably did more harm, uh, you know, than, you know, not allowing everyone to kind of take advantage of the.
So you're not, you're not basically worried that today's cutting-edge models could cause like new bio threats?
No, I, I, I am. But I think there's also a lot of positive that come out of it as well, everybody to use it, though.
Yeah. I mean.
I, I just, you know, it doesn't even have to be a Fable-style model. It can be, you can take a frontier model, distill a more, um.
Right.
Custom model. Yeah. Safety, uh, was, you know, safety standards involved, but sorry, I don't want to be too long.
No, great question, Ron John.
I, in terms of the first part of it, I do think again, like, and it's one of the underappreciated things, and I feel in the AI industry, like disease prevention and scenario modeling around vir. Like, we don't talk about that stuff enough, or whoever is trying to talk about it, it does not get listened to enough. So I love the kind of work you're doing, and like those kind of stories would bring the AI industry a little bit better reputation.
Yeah. No, I, I agree.
You're thinking about biological weapons?
No, I think that that if, well, if you can do one, then you could certainly do the other. Am I.
I don't know. I'm not.
Is one like, okay.
I actually have no idea.
Why are you giving me all these problems here?
Uh, no, I, I, I hope that we can, and we know that's a very.
No, no, no. I love it.
We're going to speak with, we'll, we'll hopefully have some time to talk about health with Brockman.
Yeah, I would love. Yeah, if you can. Okay.
Okay. Thank you for the question over here. Hi. Uh, my name is Heidi. I'm your linking online friend.
That's right. You messaged me yesterday.
Yes.
Yeah. Yeah. Uh, during this, uh, event, see each other in person. Uh, my question about, uh, before you, with a different speaker, mentioned about China. So I'm Chinese. I'm curious about, uh, why you think, uh, Chinese, uh, China is important base partnership with, uh, uh, USA, especially high-tech. Second question about R. So you mentioned about RM large model. So my question about, uh, Gemini, JDP, and Cloudy, and which one you prefer, think about in the future is more stronger, smarter in the future? Thank you.
Thank you. Thank you for the question.
Well, the question on China is, is, sorry, what's the question on China?
No, no. Why you think it's important, uh, Bis relationship with USA? So, yeah. Thank you.
You know, all right. I will, I will take that question, and then Ranjan will answer which model he likes better. Um, you know, I, I think that we're obviously in this competition. The US and China are in this competition, and there's like a lot of suspicion on both ends. And, um, you know, I've said in the past that like, it would be great if there was better cooperation between the two countries, and people have told me that I'm naive. Um, but I'm not going to, you know, lose that hope. I think that if, if, I mean, in some ways, competition's good, right? Because then you'll just have everybody striving to build a better thing. But, you know, we just talked about, uh, AI for health. If the countries were able to work together, maybe on specific initiatives, um, that would be much welcome. I, I think the most interesting part of like US versus or US with China in terms of AI in the last few months, the conversation around moving towards deepseek and adding it into your agentic process or adding Chinese models did not exist in any conversation I was in 12 months ago, and now is in almost every conversation, at least as an option, because cost has become such a larger part of the conversation. So I think, I think it's a good thing. I think the more integration there is overall across these systems, the better.
Yeah, but watch out for.
The cat holding the hands with Dario and Sam is my favorite meme of the whole year. Okay, we have one minute left, but let's see if we can do it in a minute.
All right, I'll go through this quickly. Hi, Maria Sarah Roberts. I'm in business management consulting.
My question is a little bit different. Um, as we are evolving into or entering or in a world already where endless possibilities are available, where and how do you see the evolution of the responsibilities that big tech companies have in supporting making this world better? Not the product and our engagement and how we interact with all the screens better, but rather like the real big problems that the world has. And not because they should, but because they can. So, how do you think about like the evolution of that as we step into this world?
I, I, I'll answer that. Or do you want to answer?
Go ahead.
I just, we, we had a backstage, we had a long rant because I'd forgotten that Anthropic is a PBC, public benefit corporation, but I still think there is a responsibility. I don't actually see it happening in any kind of way because competition is so fierce now. I wish it would, but I don't see it happening.
Um, yeah. I, um, that that was a good answer. Come on. We have, we, I want to make people, make sure everybody gets to coffee.
Oh, no. Screw it. I'll answer it.
Um, so, so Jeff Bezos. So I think they have an obligation to do good for the world. And, um, I don't know if, and I know people within the companies are, you know, are very serious about that. So they have their own side projects, but I think, you know, as a society, maybe we can't expect them to do it out of the goodness of their own hearts. Um, Jeff Bezos was recently doing an interview and he said, um, he said, "I think what I do in the private sector is going to outweigh whatever I could do charitably." I disagree with that, and I think that that's just the mindset of a lot of people at the top of these companies. And so, like, everybody else needs to be aware of that and not count on it, even though it would be nice. So that's my perspective. All right. Uh, Ronjan, will you come back with us after the break and speak with Alex Stamos and I about whether, um, this Mythos and Fable stuff is marketing or material?
You know, I would love to.
You guys want to get a coffee break?
You guys are amazing.
Thank you very much.
All right, we'll, we'll be back here 3:25. Um, please enjoy the coffee, and we'll see you in a couple minutes. So, you may have heard that AI is in some trouble in Washington. Um, and that story has led to a lot of guessing. People have guessed, are these models actually that dangerous? How much better are they than previous generations? And do they deserve to be restricted by the federal government? We've certainly done a lot of guessing. And I think the only way to get to the bottom of this is to speak with somebody who is deep in the weeds on the cybersecurity aspects of AI models, someone who secures them, and someone who has spent a career working in security. And that would be Alex Stamos. Alex Damos is the former Chief Security Officer at Meta and he's the current Chief Product Officer at Corridor. Um, Ranjan and I have this debate all the time. Is what's going on with Mythos marketing or material? Well, let's get to the bottom of it. Ranjan will come out and be my co-er for this one. So, ladies and gentlemen, please welcome Alex Damos and once again Ranjan Roy.
Thank you.
Hey, thanks, man. Alex, boring times in your industry.
Yeah, it's an honor to be here. I am by far the poorest person in the afternoon lineup. It's like I bring down the the mean, uh, net worth by a comma, I believe. So, I do appreciate being here.
Well, Alex, um, I don't want to quue you into podcaster finances, but.
I think you're good.
With the guests, I guess. Yeah. Yeah. So talk a little bit about, um, we're obviously at this moment where the government has pulled down or or basically forced Anthropic to pull down Fable or Mythos, its frontier model.
Right. Both of them. Yes.
What happened?
Right. So, uh, just for recap, uh, for everybody, uh, last Friday, uh, what happened was the White House became aware of research that Amazon had did, had done. Um, it is not unreasonable for Amazon to do this work, right? So, uh, Amazon hosts, uh, Anthropic's models in, in what they call Bedrock. This includes in all of the classified spaces. So if the NSA is using Anthropic's models, it is running in AWS Secret Cloud or Top Secret Cloud. Uh, and so Amazon is doing security testing. Expect this was part of getting ready to host Fable and Bedrock and, uh, found some things that they wanted to complain about with Fable's protections that make Fable different than Mythos. I guess we can talk about the exact ones. Uh, Amazon had sent over these results to Anthropic. There was some disagreement between Anthropic and Amazon during the week of how serious they were. Now, I don't have the exact details here, but somehow apparently Andy Jasse, the CEO of Amazon, mentions this to somebody at the White House, and the White House freaks out and, uh, according to reporting, calls Amazon or Anthropic, uh, and, uh, is not happy with how quickly they can get Anthropic on the phone. Anthropic is asking for more technical details. They're not able to get details from the White House of what the White House wants. So, the White House orders them to take Fable down. Uh, Anthropic says, "No, we can just fix these issues. Why don't we work it out together?" And as a result, instead of working it out together, on Friday afternoon, something around 5:00 p.m. Pacific time, the Commerce Department signs an order saying that the models are designated export controlled and that no foreigner, including the foreign citizens who worked and actually built the models themselves, can can touch them. Uh, whether or not the administration actually has that legal power is disputed, but Anthropic decides to actually follow through instead of immediately getting a temporary restraining order. And they, because they have no ability to make sure that no, uh, foreign, uh, citizen is is seeing the models, actually pulls them down. Uh, both Mythos, uh, which was privately accessible, as well as Fable, uh, which was publicly available.
Now, you've spent a career working in security.
Yes. Um, Ron John and I for months now have had this debate about whether all of this talk of doom and super powerful models from Anthropic, like for instance, Mythos, a model so powerful that you can't access it. Um, whether that represented a real capability increase or whether that was just marketing and it's a step change, uh, not really a step change in capabilities. So I actually want to toss it to Rajan, who is firmly on the marketing side of this thing, just to make the argument of why you think this is marketing and Alex can address that.
Yeah, I think trying to understand like what is that capabil new capability that is so dangerous? What is that step change? I think the reason I get so suspect on it is how perfectly coordinated a lot of the marketing and communications are around how they've rolled these out. Like, again, Mythos is the most dangerous thing. There's someone eating a sandwich in a park that actually got, like, where the model broke out of containment was like a PR release. It was, it was a very coordinated thing. So what is that danger that they're trying to kind of bring to the world? Like, what is it really? What does it look like, feel like, or what is it actually going to be a problem?
Uh, so how much money do you guys have writing on this?
Well, how would we even play this? Like, I don't even know.
Oh, like as a bet?
As a bet? Like, is there a poly market? Uh, yeah, I'm trying to see what my angle is here. Apparently I have.
Well, well, no, you are the person that knows Trump action. You know the answers here and we have just been kind of like debating.
Just like we just argue on.
So, okay, so it's actually, it's actually a little bit of a complicated answer. So, um, Mythos is almost, it is the best bug-finding model that we know about for sure. It is most likely the best bug-finding model in the world. And I say most likely because we do not know what the Chinese secretly have in their labs. But for all the stuff for which we have open known security evaluations for, it scores the best. Now, is it the avenging cyber god that Anthropic makes it out to be? I personally don't think so. Yes, it is.
You're not necessarily, you're not right yet.
But, but it is, it is really good. It is really good at finding bugs. But from my perspective, the like Opus 48, GPT 5.5, a bunch of other models are here, and Mythos is here.
At bug finding.
At bug finding and exploit development. And it is not here, which is kind of where you would think from reading all this stuff. The other thing is that everybody pretends that Mythos was the breakthrough, but that is not true. For people in the security world who've been paying attention, this Rubicon was crossed well last year, with the Opus 4 series, with the GPD5 series. That is when the models became better or as good or or better than the best human bug finders. And not just as good as the best bug finders. And there's like 50 or 100 of these people, and I know a couple dozen of them. The problem is those people are very expensive, right? And they don't scale. You can't just throw water on them and make them multiply like gremlins. That would be nice. I mean, certainly the NSA would love that, but models can, you can just scale them as much as you want with money and power and and and GPUs. And so that last year was a huge deal because all of a sudden you saw these amateur bug finders. Imagine you went to a high school track meet and every single kid is posting Olympic world record times. You'd be like, "Man, they're juicing, right? That something had happened." That is what last year was like. Like mid-la, mid to late last year in the vulnerability discovery world, is the kids started juicing, right? And it's because the models hit this point where they could find everything. And so, yes, Mythos is really good at finding bugs. It's really good at writing exploits. Fable is not. And so that is that is what the whole point of Fable is. It is the same model weights as Mythos, but Anthropic put a protection in place to keep it from doing the really nasty stuff that Mythos will do. And Anth, and Amazon has some complaints about that. But what happened is the complaints Amazon has, a number of us have looked at them, and then Anthropic has pushed back. Even with the jailbreaks that Amazon has, you cannot get Fable to do things that you can't do with the Opus series, with GPD5, and even with a bunch of Chinese models. So that is why the the actions by the Trump administration make absolutely no sense because those models will not refuse at all. You can just ask GPD5 to go find these bugs and it will go do it for you. Um, you can just ask Kimmy, which is an open-source Chinese model. It will go write you these exploits. And so, yes, you can trick Fable into doing certain things, but you can't trick it into doing the really powerful Mythos stuff, which is like grind on the Linux kernel for all day and find a bunch of bugs, which will cost you hundreds of thousands of dollars, it turns out, if you pay the full price. Um, and those have not been jailbroken. And so there's no real good reason for what happened. So, yes, Mythos is really powerful. It has only been accessible to a very small number of companies. It is not magically powerful. The other thing you have to remember here is like bugs are cheap. Exploits are like, we're now drowning in bugs because these models are so good at finding, and we have so much bad software. We've been writing bad software forever. And it turns out that like writing human beings writing code in C and C++ especially was a really bad idea that like we should not have been writing, like building our entire lives like based upon memory unsafe languages. And then in the presence of these superhuman bug finders, not just Mythos, but all of them, that was not a good idea. And so, uh, if you're a pedestrian and you get hit by an F-150 at 80 miles per hour or a McLaren at 200 miles per hour, you don't really care. You're dead, right? And that's the difference here is like the Mythos is the F1. Mythos is the F1.
Yeah, exactly. And Kimmy or GLM52 is the F-150. And so you could take the the the McLaren off the market, but all you've done is actually hurt the defenders who want to have full coverage, right? Who want to be able to find all the bugs. An attacker only needs one or two bugs to string together to actually pull off the attack. And so that's why, you know, so about 150 of us wrote this open letter and signed it saying this is really stupid because all you're doing is hurting defenders and you're also really hurting the US AI industry.
Okay, I think Ron John wins the debate.
Well, no, still, it's not all marketing, but it's.
It's not all marketing, but like, I, I just want to say like, Mythos is really powerful, but doesn't really matter because there's so many bugs out there that you can use any of these models to find them. I, what I'd like to do is I'd like to shift the debate one away from finding bugs because the other problem here is like, we cannot set the standard that US AI models can't find bugs. That is a terrible, terrible, terrible standard. If you have a piece, if you are writing software with an LLM, it has to be able to understand what a bug looks like so it does not write those bugs. It has to. And that is that's the trick actually Amazon used was they tricked it into basically finding bugs in individual lines of code because it has to be able to do that to write secure code. If, if our, if American LLMs know nothing about security, they will create more security flaws. So we cannot create a standard where American LLMs are dumb about security. That would be a humongous own goal, and that is one of the possible outcomes here of of this whole brewhaha. That would be really stupid. And so, uh, you know, my, I think yes, I, I don't think either one of you wins. I mean, as the referee here, I think Ron John has a little bit of a, you have a point, but Mythos is really good. It just.
It doesn't. Anthropic has pleaded up too much. Um, OpenAI is rumored to be releasing something that's Mythos quality really soon. Um, and I doubt that they're gonna like follow. They'll just give it a number like, uh, I think they've learned from Anthropic, you don't call your thing like the cyber gorgon.
Uh,
That actually would sell.
Yeah, yeah, it sells a little too well, right? Um, right. Maybe as a Greek, with the Odyssey coming out, I'm getting a little sensitive about the the marketing here, but like, like, you know, let's, let's like back off on the naming scheme and move back to numbers. Uh, uh, but like, I, I think a, a big part of this is we just have to like reset the standard of like not bug finding, but the actual creating and creation of exploits and then other kinds of offensive operations of which there's a bunch of other things you can do offensively with these models. Those are the things that we should call off limits, not bug finding, because that is something that we have to do much more quickly than we're doing right now.
C.
Can I, I just want to try to articulate my point, which is, you know, this is all right. So when we talk about, is Mythos material or is Mythos marketing? Um, my argument would be like, wouldn't if something was material, it would look a lot like what we're seeing today, right? Like, at a certain point, these models are increasing in capabilities, you're going to want to put some safeguards on them. And so, is what Anthropic is doing that different from like where you'd want to be if you actually had these concerns?
I mean, I think both Anthropic and OpenAI have a path in which they have like a KYC model now.
Know your customer.
Know your customer. Yeah. So, OpenAI actually has a couple levels. They have a public one where you can basically say, I want access to cyber capabilities, and then they have a private one. I believe Anthropic only has the private one, right? To get access to Mythos. Uh, but, uh, I, I do think it's reasonable to have some level of gating, but in the case of GlassW, I think it was honestly, it, I thought it was not widespread enough. I know of like critical infrastructure companies that don't have access to Mythos right now.
But nobody does right now. Well, no, right now, before last week, like people who should have had access did not have access. Now, seeing what happened here, you can see why they were so careful. Um, but from my perspective as defenders, we are in a race against time, right? We have all of these bugs that need to be fixed. And we also have this weird thing where we're like, this whole conversation is predicated on the idea that like the American labs are years ahead of our adversaries. And that is just not true, right? So while Fable was shut down on Tuesday of this week, GLM 5.2 was shipped right from Z.AI, which is a Chinese lab. This is an open-weight model. It is MIT licensed. So any of you can go not just download the weights from HuggingFace, but you can include it in your own product and you can retrain it. It is within percentage points of the top closed models from Anthropic and OpenAI in a bunch of things. Um, a bunch of of evals. We don't know how good it is at bug finding yet. So there hasn't been any good testing here. But in coding and a bunch of other intelligence tasks, it's almost as good as the best the United States has. And that's the best Chinese model we know about, right? Unlike American labs, the Chinese are not going to announce their Mythos, right? And so the idea that we should be doing everything in the United States to just be playing defense is we're playing like a, a defensive model here, that is stupid from my perspective. We need to be accelerating our capability to do these things and not just playing defense.
Well, now I am sufficiently scared, even though I've thought that the Mythos was pure marketing. But, but what do we do about it? Like, how do we, what is that like? And I, I think I share that Glasswing again, why it felt like marketing to me. It was this kind of exclusive club that people were invited to. And I actually spoke with people who are using it and they spoke about it almost like it was a club, like, oh yeah, we got access to Mythos. We're spending a ton of money. Like, what should it look like this kind of rollout to actually.
When you get help us.
In Glasswind, I mean, it, they're fixing real bugs. They have found like 10,000 bugs now. One of the problems with Glasswing is they found like 10,000 bugs, but they've only fixed a thousand, right? So, like, you do not want a 10 to 1 ratio of the bugs you found to fixed. Um, we just joined a thing called Project Athena, which Chain Guard is running, which is all about fixing specifically open-source bugs, right? So Glasswing is a lot of that is closed source with inside commercial companies. We're part of this coalition that is using, uh, Mythos to fix open-source bugs, which is a huge challenge because like, just finding the people who have access to that, that source code, commit access is difficult. Uh, so that's going to be like a huge push. So it, it's not like just totally BS. Um, I do think that one of the things that's going on here is that if these companies ran any modern AI security tool against their codebase, they'd find 80% of those bugs. They don't have to use Mythos, right?
Okay.
So,
That's concerning.
Yeah. Right. And, and so they've, they're like, "Oh, yeah, I've got this old crappy code." And they look for the first time, and they all of a sudden they find.
And they could have done it a year ago or six months ago and found some number or some percentage.
I had this conversation with a CEO this morning who's like, "How do I get into Mythos, Alex, who can call?" And I said, "Don't wait."
Here you go. Here's an API key for Opus right now. It works today because do not wait for the Trump administration to figure out what they're doing with Mythos. Use Opus 48 right now, and I guarantee you'll get 70% of the bugs you care about.
Alex, you're, you're a security researcher. One of the stipulations about getting Fable back. So, by the way, Anthropic has said they expect, uh, Fable to come back on soon. Um, but one of the stipulations that's been talked about to do that is that they have to prevent it from being able to be jailbroken.
Is that possible?
Yeah. I saw, I saw you did jailbreak in air quotes. Yeah. No.
So we have, we have two minutes left. So talk a little bit about why that's not possible.
No, there's actually a paper from NIST. It's like a kind of a goal completeness theorem kind of, uh, paper that basically says it's impossible to make a model that is completely not jailbreakable. Like I said, Fable needs to be able to understand what security flaws are to be able to do its job. I think what, what Anthropic might, what, what Anthropic might the game might they might be playing here is redefine what jailbreak is back to what their system card says. So their system card says we define different levels of cyber capability, and finding a couple of bugs and knowing what a bug is is okay, but we won't let you like build exploit chains and and do lots of long-term stuff. And so if they can define back to their system card what a jailbreak is, then they're fine, right? Um, what would be terrible for the entire American AI and cybersecurity industries is if the Trump administration defines a jailbreak as having any ability to find any vulnerability in code because then all of us have to go to Chinese models.
That's that's it. Yes. Like that would be a terrible, terrible outcome. And that might happen because unfortunately, one, the news that broke just in the last couple hours is Politico is reporting that Anthropic and the Trump administration are negotiating what the AI safety standards should be. So like, I don't know, it's like Howard Lutnik writing an eval right now. Like I find that a little terrifying. So, um, I, I am hoping that Anthropic, obviously there's very smart people in Anthropic working on this. I am hoping that they are able to come up with something reasonable that moves stuff towards exploitation, not bug finding, like we have to have the ability as defenders to find and fix flaws. The other thing I just want, because we don't have a lot of time, the other real own goal here was that this injected a type of political instability and political risk into the USI industry that did not exist. There are a bunch of companies that are now going to be have to use open-weight models because at any time on a Friday afternoon, it's, it's good thing that Fable was only out for a week because people weren't, didn't have time to work into their critical path. But if, if, if Fable had been out for a month and at 5:00 PM on a Friday all of a sudden it got yanked, then pagers across the country would have gone off because every, everybody's systems would have fallen over. And so now this week, CIOs and CTOs are signing contracts to have open-weight models on different hosts on US hosts. They're not using Chinese hosts, but they're using Chinese models on as backups through LM routers because political risk is now a risk of using AI companies. That is a massive own goal. The United States of America does not need to do that. We should not do that again. That is incredibly stupid at a moment of maximum pressure from the People's Republic of China. Um, and so I can't stress enough like, we, we lost a war this weekend. Like, let's not lose another war by making it unreliable to utilize the United States for AI.
Alex, will you come on the show for a longer conversation about this?
Absolutely. I'd love to. Yeah.
Alex Tamos and Ron General.
Thank you.
Thank you. Okay, so Anthropic is not just a massive model builder, it's a massive product builder as well with products like Cloud Code and Co-work that have taken off like crazy over the past few months. And Cloud Code came out of a place that few of us know much about called Anthropic Labs. And Anthropic Labs is an organization within Anthropic that is working on building the next level of frontier products with AI at the center. And so today we are lucky to hear from the person running that lab, Mike Kger, who is the co-founder of Instagram and now the lead of, um, of Anthropic Labs at Anthropic. We're going to welcome him on stage along with Lauren Good of Wired, who will join me as a co-in. Mike and Lauren, let's, uh, let's, let's hear from both of you guys. You're here.
All right.
All right, Mike. So, chill times in Anthropic land.
Nothing going on.
Slow week.
Slow week.
You want to start, Lauren?
Yeah. Uh, first of all, I mean, we want to talk about what you're working on at Labs, uh, and explain your role to folks, but I want to ask you first, how close are you right now to the situation with the White House?
Less than in my CPO role. So I transitioned about like five months ago into this Labs role. I think in the CPO role, I think would have been deep in it. Now I'm more, you know, obviously, you know, we want to restore it and and as a like, uh, product person, I want to make sure that that gets access, but less close to it than, uh, you know, in that sort of CPO role that I had before.
Okay, Alex, do you have a follow-up? I have like eight follow-ups.
Definitely. Well, we, there was a guy named Ben on X who said, "I will mail Anthropic an original copy of my long-form birth certificate if they will enable Fable for me again." I sound like those lunatics who were obsessed with 40 now. Will you take Ben's long form? I don't know, uh, that will take his long form, but it has been interesting. I mean, Fable was only available, um, for a few days, but I've definitely, uh, every time I've tweeted since then, uh, they've not read whatever I was tweeting and they've mostly been like, "Bring back Fable," which like, in Instagram, we got rid of Gotham. Do you remember Gotham? The filter. This is like, and then for the rest, like the next eight years, all I heard was, "Bring back Gotham." So, it struck a nerve, but Fable will, uh, will come back before Gotham did. But yeah, it's, it's clearly the folks that had gotten to use it and started incorporating it. It's actually really interesting. Um, I've learned to not really trust day of or even week of model reactions. Like, you don't really know until you've put it through its paces. And so like, I almost just completely block out the noise in the first couple of days of any new model release, cuz I don't know, everybody has maybe their like toy like example thing that they like to do with a new model. But it's hard to actually put it through its paces until you've actually had real work done with it. And I think people were just starting to do that, and then, you know, we had to, uh, uh, sort of pull back Fable. But, um, I remember in December when we put out Opus 46, uh, it was like this interesting time where everybody went home for the holidays, and a lot of people had that week off between Christmas and New Year's, and then they came back and were like, "Oh, I spent a lot of time and I really get why Opus is good and I want to do it." So I don't think Fable has had that opportunity yet. But despite that, though, I mean, this was a pretty big reaction from the Trump administration. And I think everyone here, especially if you were listening to the Alex Stamos session, uh, understands what's going on. But this happened within a few days of the model release. Um, if I can give Wired some credit, Wired just reported last night that it was due to Anthropic's, uh, relationship with or having given, uh, access to the model to SK Telecom that could have raised flags within the administration. How surprised were you by how immediate that that backlash was?
Yeah, I think the the the sort of reaction decision was surprising, and we were sort of immediately engaged with them to, you know, restore access as well. Um, and so at the same time, you know, the one thing that we, we think a lot about internally is, you know, there used to be a poster on the Facebook wall when we were that was like, "Every day feels like a week." And I think that's becoming true in AI. And I think a good thing to remind ourselves of in general in the industry is, uh, we're dealing with unprecedented times. We're dealing with new situations, and, uh, they can develop really quickly as well. And so I think, uh, also developing the capabilities and the connections and to make sure those conversations can happen quickly is really, really important.
I have a new motto suggestion for you for the wall. Move fast. Move fast and jailbreak things.
Jailbreak things. I don't think they're going to use that, Lauren.
Okay. Um, you know, we just heard from Alex that some of these capabilities have been, uh, available on, you know, previous models or, you know, you can, you can find bugs with other, with other models that are available today. So, why do you think Anthropic got singled out on this front?
Um, I don't know why like Fable specifically was singled out, um, as, you know, again, Fable being the non-cyber, uh, intention, uh, model as well. Um, I think that the thing that does change over time is there's, uh, you know, the capabilities. If you knew what to prompt and knew what you were looking for, and then there's the capabilities. I think, you know, uh, I'm not sure I resonate with the like juicing high school athletes metaphor from Alex, but like, you know, the uplift that you get. Uplift is a thing that we think about a lot when we think about model, um, safety. So if you look at our model cards, for example, one of the ways and we look at, uh, risk in the bio domain is comparing the uplift, uh, from sort of lay, a lay person using the model versus, you know, an expert or lay person just using the internet and, and seeing what the comparison there is. And so that is one trend that is, you know, has been, you know, progressing as the models get more capable. Um, and so maybe less why Fable is singled out, and maybe more like what the overall trajectory is.
Interesting. So, we just had Ara Karazzian, the, uh, lead economist on RAMP on the podcast, and one of the interesting things that he said was, you know, we, we had during the Pentagon situation, there were these headlines, okay, this company will not use Anthropic anymore, but actually the data from RAMP shows that spending actually increased to Anthropic models. It was apparently a good publicity moment for the company. So, I, I think it sort of plays into a debate that we have like here on the show about like, how much of this is like real concern for the issues and how much of it is, you know, from Anthropic is marketing? And like, we have somebody from Anthropic here who can actually shed some light on that. We just said Stamos talking a little bit about it from his perspective on the security side, but we're lucky to have you here today. So, so is it material or is it marketing, or some combination?
I mean, I think, uh, it is one of the hardest things to like, really deeply believe that something is true and not marketing. And due to not even just Anthropic, I think like, you know, uh, uh, I think people are generally right to be skeptical of any company saying anything, and you should like put it through your own filters as well. Um, but for me personally, I'm always like, "No, but it's real." And like, we like both deeply care about safety and are like, to the extent that we are being vocal about anything, it is to either sort of help paint the picture of what very likely is coming or we believe is coming, or what we've already seen and spotted. Um, you know, for example, in the, in the Mythos case, really just looking at, uh, uh, like vulnerability scanning and bug finding and doing it in partnership with with companies that were in that kind of Project Glasswing initial announcement. And so, um, the technology is awesome, not in the like, in the just, you know, it is, it is doing really incredible things. And therefore, even calling out what we see as what is happening, I think can seem hypey. Uh, I wish I could press a button and was, and like, I could make everybody believe that we are not being hypey. I realize that's not the reality that we operate in, but at least from my perspective, um, we try to call it like it is.
My understanding is that within Labs in particular, in research and development at Anthropic, that you're using the best AI models to actually build new products, to prototype new products, and sort of test out your thesis. So, is this, uh, is this li, is this ban essentially now limiting your ability to do that within Labs?
Yeah, I mean, definitely Fable, uh, is the best model I've ever used, and it's not to say that like work has stopped, but it's definitely like less good than the other models that, uh, that, or sorry, we're using models that are less good than that, um, as well. And it is also, I mean, maybe the reverse of the, uh, distrusting the first week, uh, you know, response is what happens when you don't have Fable. And like, obviously the Twitter reaction for the people that had already kind of like gotten, uh, into the model and and were using it was strong. Um, but I'd say even in my personal use, I'm like, "Oh, like I'm on Opus 48 and it's good. Like I'm still productive. I'm doing work." But, um, and we can go into sort of like how my work changed with sort of these like Fable or like, you know, uh, the models of that sort of, uh, family, but it is noticeable for sure.
Yeah, I think we'd like to know that. I mean, you know, we, the public had access to Fable for like a half a minute. Um, groups in this like Project Glasswing have had access to Mythos. Um, we don't really know what the difference is between using a model that you can use today is and using one of these, you know, super Anthropic models. Yeah. So what actually could you do differently with a Fable or a Mythos?
I think for me, it's the sort of scope and scale of delegation. Um, and all these things are really imperfect. Like, you know, people would say like, "Oh, is this now a level five software engineer or level six?" But anybody who's used these models extensively knows that they're still spiky in capabilities, right? In some ways, they're, you know, in many ways, they're a better engineer than me. Um, and in other ways, like, you know, I was complaining today that it had missed a descender, like a, the G that bottom part is called the descender. I'm like, "How did you put that in the UI and it's clipped?" And of course, like, there's vision capabilities that need to improve, and there's, uh, debugging capabilities, and there's sometimes even just sort of human common sense that is, you know, we're way better at than when than the models are. But overall, I think the big shift for me working, and it was really interesting because, uh, it sort of coincided with me going back into a builder role. So I really got to see going from, you know, using these models as an executive, you're trying to do the most of them, but you know, like not going to have it write all your email and, you know, I think strategy still needs to come from you, and then you can use the models to sort of pressure test it. Um, but going back into a builder role and going from, okay, I am delegating chunks, like, please fix this bug, or I'm thinking of implementing this feature, let's go back and forth, um, to something that ends up being much more, sort of, all right, like I got this bug report from one of our users, or I have this notion of something that I want to build, like, can you sketch out two or three ways in which we could do it, right? Like, that seems plausible. Often I find actually sometimes it'll, uh, give me the, the sort of explanation or proposal and be like, "Okay, that actually is over my head. Like, you are clearly way smarter than me. Explain it to me like I'm maybe not five, but at least, you know, uh, uh, not you." And it'll sometimes, you know, explain it that way, but then go build it and get it right. Getting it right, like, you know, at a very, very, very high rate. And I think that starts really changing how you operate. Like, I, I moved much more to before going to bed, making sure that I had like queued up for Fable like enough chunky work to last, I would call it the whole night, and then I would check in later and it got it done in an hour, and it was like, I guess hanging out for the next seven hours, but like, really like delegating, like, uh, uh, much more of a goal than just a.
So one, one task, for example.
Yeah.
Like, give us an example of one task you would hand to it.
Uh, I mean, here's a kind of crazy one, which is, uh, for the programmers in the audience, like, I, uh, I had written one of our Labs, uh, projects in, um, in Python. That's like the language I know. All Instagram was all Python. And for like, some not super exciting reasons, we actually needed it to be.
In Typescript to deploy it. And I was like, "All right, that's going to be like, you know, in Instagram, we for years talked about moving from Python to PHP or hack or the the Facebook language after the acquisition." And at least when I was there, never did.
Um, but I basically we have a feature called dynamic workflows where you can have it like also break down the task into like a lot of subtasks. And I trusted it to sort of not just do the individual action, but here's like a whole language conversion of millions of hundreds of thousands of lines of code at that point. Uh, go off and do it. Go plan it. Go execute it. Go verify the work, double verify the work, and then I came back to the work being complete. So that level of like this is a big sort of chunky uh task.
>> So that was so you're basically saying it was faster. Did it in an hour. You're you're s guessing compared to with fable compared to what it would have been before. I think the main difference is in the past you would it would be like great I did it and you'd be like but did you you kind of took a shortcut here or this is not quite right or I need to go verify it or like oh you cut this corner or
>> It's like the managing interns thing that everyone's been saying for the past year which is very offensive to interns by the way but yes
>> Yeah, exactly.
>> I don't know have you managed an intern? Yeah, that's true. And it was and you're saying it was more correct. It was so it's faster. It's more accurate. Uh more reliable and then according to the US administration dangerous. I think the other piece is it has like a a greater theory of mind is the wrong word but sort of like theory of project so that it's less you know oh I'm going to make this change and it'll say great I'll make this change but really especially if you've done like software engineering at scale the best engineers kind of keep in mind all the disparate parts of how this things and they also see around the corners like I can make this change but if I don't do it in this way then the next change is going to be incrementally harder and I think that's like been a significant difference I've seen in that kind of class of models.
Okay, cool.
>> So, I think when we talk about Anthropic Labs, right, people think of Claude Code because it is really your breakout product and uh and it sounds like you've been tasked with basically figuring out what the next Claude Code is. Would you say that's an accurate description of what you're doing at Labs and also why does Anthropic need Labs?
>> Labs. Yeah, it's also maybe worth thinking about why we needed Labs in 2024 when I arrived and why we need Labs today because I think that the answer kind of shifts. Uh I started the Labs the original Labs team with Ben Man who's one of the co-founders of Anthropic. Um in my third week at Anthropic and it had been something that I've been bubbling under and at the time the reason was really different. It was all of our product engineering with team was 25 people and we didn't have the models really like we had Claude that when I joined it was set three like you and Opus 3 like those were for the time good models but you weren't going to they were not even interns right they weren't even IC3 engineers. Um so if you have a team of only 25 or 30 engineers they are working on like the next you know incremental thing and we were feeling like the models are starting to get better but we don't have any products that sort of show that off like a good litmus test for me is when we get ready to release a model do we have either a product or a demo or or you know some other illustration of something that is very different and it gets harder over time like with with Fable you know even like illustrating like that weekend task or you know this longer amount of work. So really Labs at the time was let's make sure we don't like our products don't fall behind the model exponential that's happening. Um and yeah, so Claude Code came out of that initial one because nobody in the rest of the product or people were thinking about coding but nobody was sort of had the like space to go and think about well what if we totally change the form factor and we embrace the fact that the models were going to uh evolve in this way and a lot of the like the two most useful thought exercises we do in Labs. One is like visualize the gap between what the models can do today and how most people use it and can we close that gap. So that's one. And the other one is imagine what the models are bad at now that they're actually going to be really good at in six months and let's make sure we have a product ready for that by then. And I think those are like the two guiding uh questions for for Labs. Um and then also out of that first incarnation came computer use. Computer use was different though because when we built it it was really bad. Like we tried a bunch of products with it and this was around you know silent 35 and be like Claude can you help me you know clean up my desktop and it would like click the thing and it delete the file like this is not not safe for release we're definitely not going to go and and and build this or to ship this but we had that product so that every new model that we we'd release we'd first check it internally and say did computer use get better and we'd tell the research team how it gotten better or worse until the moment where we said it's good enough we're actually going to put a product out around this. So it's it also gives you this sort of uh sort of beacon into the future that then you can kind like measure your your future products against. Um but then compare it to now. So we have a driving product team. There's co-work there's um you know Claude Code has grown a lot. We have our platform. Um and now I think it's actually much much less about none of these product teams are doing this sort of thinking and I think it's much more that the models are advancing really quickly and even our capability to interact with them needs to evolve. So one of the things we collaborated with uh Labs and Claude Code that we ship today um is Claude Code artifacts. So having Claude Code not just be able to type back to you but also sort of draw a picture or give you an illustration. And that partially came from spending a lot of time in Lab saying just a text box and a big text response is not going to cut it anymore. Like when I mentioned that the models feel like they're way smarter than me when they talk to me sometimes I'm like can you draw me a picture because this is what I actually need to fully uh fully understand this. But it's really what we've been thinking about is, you know, yes, we have a lot more products. You know, we have we actually have a lot of consolidation to do in our products. Like that's another initiative that we have. But within that, we still have an opportunity to make things much more accessible to a person that does not spend all of their time thinking about prompting and the exponential and the difference between high, low, and medium effort. Like there's a lot we can do.
>> But Mike, so there's a it's it puts people using Anthropic models in an interesting place, right? Um, there you know, Cursor I think just sold for 60 billion to SpaceX and someone put this meme on Twitter that like, you know, Cursor would have sold for 300 billion if it wasn't for this guy and it's a picture of Boris Cherny, the person who created Claude Code. Um, and so for companies that are going to build on top of Anthropic technology, you know, they're going to wonder, do I want to partner with Anthropic or is Anthropic going to go ahead and and build the product that I'm going to want to build potentially even after partnering with them?
>> Yeah. And I mean, we'll take the like agentic coding side and I think the broader sort of aspect of of you know, being both a platform and a product I think is is really interesting. Um, when we take on projects, it's the goal is often to sort of push that area of the industry forward. So, um, you know, there were AI coding editors and some of them were really good and, you know, uh, but nobody was quite thinking about it in as sort of free form a way as we got to think about it with Claude Code. And now a lot more products have that flavor than I think would have otherwise. And so I think if we're ever, you can call me out on this Alex, like if we're ever entering an industry where like all you're doing is the same thing everybody else is doing but like you've got the Anthropic brand, I feel like that's a bad use of our time and a bad use of our either Labs or product team time. Like if we're going in somewhere it should hopefully be to say all right, we think that the direction of travel is this way, we can build a product of that and then by the way, there's no world nor should there be a world where like all the products are Anthropic products, that would be a bad world, right? So like that is hopefully either creating new space for for companies or sort of showing the way where other products can incorporate that.
>> It would almost be like working for a tech company that has like social messaging video.
>> Right. Mike, yeah, okay. You did leave.
>> Well, there was a question for example when, you know, Anthropic launched a product that was seen as competitive to Figma and you had been on the Figma board prior to that and I think you stepped down, is that correct?
>> Yeah. And so, um, it's a good question that it's al Alex has brought up, I think where Silicon Valley is known for this really healthy vibrant risk tolerant startup ecosystem and when the bigs start coming in with tons of venture capital and, you know, a lot of resources, uh, people say, well, wait, are they just are they essentially just going to steal my idea?
>> Yeah. No, I think I think our dual existence and it's something that other companies have had to navigate. We talked all talked about Amazon a lot in the in the previous panel, like they have to navigate this role where they are both infrastructure provider. They obviously have a very large e-commerce. They they do video but they also serve video and um and then, you know, by and large customers can live in that dual world of like, okay, I'm using their infrastructure also knowing that they they are also using their infrastructure to do that and I think the, uh, you can talk to our customers and see how well we're doing at this like the thing I always try to do is like at least approach it with a lot of transparency. So the Cursor example is an interesting one where like Michael and I talked a lot over the, you know, time around here's where things are heading. Um, and, uh, you know, with similarly with with the other products that we think about like, can we, I think it's a couple things. It's transparency and then it's shared building blocks. Like yeah, I think, um, in general and I actually don't think there's any cases where this is even true, like we're trying to build on top of the same capabilities that are available elsewhere. The last time I was here in the Commonwealth Club on this stage was our healthcare day at the beginning of the year and we didn't ship like Claude healthcare only, we have it like nobody else has it. We shipped a bunch of like plugins and skills and MCPs and like complimentary abilities. So that's how I'm not claiming it's easy or that it's a straightforward thing, but it is how we're trying to navigate what is like admittedly a complicated sort of situation.
>> Speaking of startups, Anthropic is still technically a startup, but you're worth a lot of money. I mean, what what's the latest valuation? Is it
>> 965?
>> 965 billion dollars or something like that. So
>> You sold Instagram for a billion, right? Startup in 2010. Yeah.
>> Right. Financials have changed uh quite a bit since then. And yet, uh, Anthropic has positioned itself, it is uh, you know, a PBC and it's positioned itself as sort of a more, um, ethical uh, company around building AI. And I'm wondering if you could talk a little bit about how you see that positioning and Anthropic's role in particular changing the culture of the valley. Like I know I think back to how Google in the beginning of the 2000s really changed the culture of Silicon Valley in so many ways and how do you see Anthropic's culture now dictating this next era?
>> Yeah, that's a really interesting question. Maybe I'll start like inside and I think there's an external component too. I think the reason I joined in the first place, so I was winding down my second startup, um, and knew I wanted to go work at a frontier lab because I had started to use these models for coding and they were bad at coding, but I could see that they were as bad as they were ever going to be at coding. They were going to improve and I had started building on top of these APIs. So the startup I was doing was called Artifact and we did sort of AI powered uh sort of news recommendations and um, actually read a lot of big technology via Artifact back in the day. It was a lot of things we added. So
>> But not Wired?
>> But not, you know, you guys had a really hard paywall to be honest.
>> Fair enough.
>> We didn't do very great.
>> Do you need a discount on my subscription because I can get one for you. Okay.
>> It's actually really funny like the bacon deals.
>> Yeah, bacon deals. It's the It's the login cookies. It's like really hard to keep people.
>> I know. I know. I just Please escalate this to Kai now. I know.
>> Um, uh, and email login very hard to do in, um, but so, uh, but I was building on top of the APIs and be like, "Wow, okay, they're they're able to do really interesting things." Um, but what ultimately made me go to Anthropic was like they walk the walk and they really uh like deeply believe in trying to make AI go well for humanity and that is like in that's like in the water internally and I think has been why I think the company has remained as cohesive as it has even as we've grown and I think that it's like a testament also to the the co-founders there on how often they are talking about this as well. It was a a surprise for me coming from a world where at Instagram we did a weekly all hands and we talked about product 95% of the time and maybe 5% of the time we talked about something else that was going on in the world around the company. Uh, probably maybe I'm under selling our like go to market, maybe it was like 80/20, but it was definitely very, very heavy product and I remember Anthropic about six months in, uh, myself, um, and Kate Jensen, who's one of the the leaders in the sales organization, did a joint all hands where we talked about our like, you know, how we're doing product and go to market together and people were like, this is so great, I finally understand our product strategy and like what we've been doing. It's like, oh, right, this is not quote unquote a product company, you know, it is a very mission-driven AI company with like a very strong sense of like why it exists in the world. I think in terms of the overall impact on the valley remains to be seen. I think, um, the positive signs that I've seen or interesting signs I've seen that I've seen are like an renewed interest in philanthropy, um, across the board and I think that's something that, um, has been written about and I think it will be an interesting sort of outflow. Again, who knows how all of this goes, but like, uh, depending on how it goes, it could mean a lot of interesting new sort of philanthropic, um, uh, deployment. And then I think the other piece is, you know, uh, the conversation around how AI could or should go is one that is happening in real time with the technology versus retrospectively, which I think has been the case for other technology waves and I think that is a good thing.
>> Interesting.
>> You know, um, Mike, you know, you talked a little bit about, um, Anthropic has this gap that it sees between the capabilities of the models and where everybody is building products. And with Labs, what you try to do is get ahead of that so you can show people what AI might be able to do now and six months from now. So please tell us, please tell us like what you're building, um, where you see the where you see the potential,
>> and and what people should be on the lookout for.
>> Yeah, it's a roadmap.
>> Oh, and if I can throw an and
>> What you're building now, but also if you had if you have a pie in the sky like Elon Musk data centers in space type ambition, I want to hear about that too. Tell us, tell us what, tell us everything.
>> Great. We have 13 minutes left.
>> Exactly. The rest is the monologue in my chronic group. Um, I think maybe two themes I'm really excited about that we've been exploring a lot. Um, the first one is giving Claude an environment where it has more agency and it also has more self-knowledge. And I'm going to unpack that because that's like a lot of AI words. But, um, uh, I'll give you an example of where we are currently doing a bad job of this. Like if you are in a Claude project and you make a file with Claude, you're like, that's great. Can you add it to our project? Claude will be, no, you have to go download the file and go drag and drop into this thing and you're like, what? Um, uh, until yesterday I would have said the same thing about Claude Design and Claude Code where if you're in Claude Code, you're like, cool, like I need a design for this thing that we're building. Let's or you're in Claude Design and you make a mockup and you want to go build it, be like, cool, here's a zip file and you're like, what? You know, and I think, um, so that's a little bit of interoperability, but in general, this theme of giving, if you give Claude a lot of notion of its environment, I was talking to actually a customer, like an API customer, um, and one of the things that they were experimenting with is actually even giving Claude like a secure version of their source code while it's running in the agent loop in their product so that if it hits an issue, it doesn't go like, I don't know, I hit an issue. It can be like, well, it's probably this thing, you know, at least when it's talking to one of the sort of maintainers of of the software. And so that overall theme, and of course, you have to do it with safeguards and like be really careful about what you unlock with it. Sounds kind of obvious, but it's actually night and day in terms of how expressive these products end up being able to be, right? And you can even see it going from like maybe like core chat or classic chat in Claude, and something like co-work where it's got a little bit more agency and it has a runtime and it's able to sort of like understand a little bit about its environment, but I think we are at like 10% of the journey about where we we could go. Um, actually one of the reasons I think people got excited about things like open Claude is seeing how a, you know, harness that is modifiable and you can talk to it about things and you don't ever get the sense of like, oh, sorry, I can't do that, you're going to have to go to this like settings screen and turn it on. It's just a thing it has access to and hopefully with like the right gardening and permissions. So that's like theme one that I'm like extremely excited about. Um, and I think I, if we do it right, should actually like transform all of our products like from head to toe. Um, the other piece, um, is, uh, and I'll maybe like share like the the feel, not the internal product we're working on, but like the the phrase I got as feedback was, uh, like I think closing the gap. I mean, I talked about closing the gap between capabilities and reality. I think it's also closing the gap between how people understand their own work and then how the actual day-to-day is to do that work. Um, I was talking to somebody, uh, internally who's on our privacy team, and to move a ticket from like one queue through another one via the task tracker into another was like eight different steps of copy and pasting of like manually moving pretty, you know, like, uh, uh, kind of annoying to have to do, probably error-prone. You have to like keep spot checking it. And we helped her with one of our Labs projects to basically like make that not a pain and she's like, ah, this is the first time in my career, and she's like, been working for 30 years, we're like, what's in my head and what I am using is like now this like it is now closed. Um, and I want to like bring that feeling to everybody who like, you know, of course Claude unlocked a lot of, you know, non-technical people being able to code, but like, it's we're still asking people to understand way too many concepts of like, what is, what is, you know, the difference between like my sandbox environment and production or like connected MCPs is myself or others or how should I store data and of course, you can't abstract everything, but like, if you combine both of those themes, if you give Claude a lot of self-knowledge and you're creating an environment where it can actually solve complex problems for people and repeatable ways. Um, I think I get like very, very excited about those.
>> And your moonshot.
>> Yeah.
>> I'm not letting you off the hook. What's your moonshot?
>> Moonshot.
>> Yeah.
>> Uh, no. Nothing in space. Although I guess we're, you know, we're in talking to SpaceX about spacey things.
>> Um, but you're talking to space. I mean, those are the announcement for compute. Yeah.
>> Um, it it was exploring extra extraorbital. What was the phrase? Something about exploring like post orbital world things.
>> Definitely not my department, but yeah, there's there's stuff stuff in. So the Labs is isn't working specifically with the team on compute.
>> Right. Exactly. Separate separate.
>> Totally separate. Okay. Okay. So you're moonshot.
>> Yeah.
>> Do you personally believe in data centers in space?
>> I had a conversation I uh I by far from a data center expert, but I talked to somebody who is a like a person who sends things to space who's not Elon Musk. Um, and
>> That's what that's what you would say.
>> Yeah. Um, uh, and they were really bullish and I was like trying to talk about why and there was basically like I guess effectively infinite, um, uh, uh, power if you convert it well and, you know, infinite land and I was like, okay, you can buy that. I mean, and I think they feel good about the shielding you have to do again, clearly not my area of expertise, but after that talk, I was like, okay, I see, I see it, you know, um, even if it's going to take a few years at first.
>> I admittedly thought it was a crazy idea in general, but now I'm like, oh, actually clearly can understand why this might make sense.
>> When when you were talking earlier about the ways that uh the work in Claude is going to get compressed and all those steps, I couldn't help but think of tokens and how, you know, maybe it's good for your business model in the short term if people have to take so many steps and use so many tokens. But tokens have become this unit of economics that we're using to describe the industry now and people are token maxing and now they're token mising and and and, uh, one, I want to see, I want to hear where you sit on that spectrum if you're a token maxer and two, is there a near future in which the industry is not actually measured by tokens, you know, it goes the way of MIPS or dialup or some other, you know, there's some other unit of measurement that actually defines the economics of this era.
>> Yeah, I think both of those are really interesting questions. It was interesting earlier this year when you started hearing about like companies that have like dashboards showing like who used it the most and we of course have like internal metrics as well. And we found that there's not a lot of correlation between like the person who's using the most tokens and like the person that I like, it's an interesting thought. I maybe do at your companies is like write down your 10 most productive people that you think are most productive and then like get your top 10 token users and see how closely they correlate. At least for us, it wasn't that correlated. Does it seem dangerous to sort of like, uh, sort of purely glorify the like maximum usage? Obviously, it's like very gameable, but even beyond that, I think it's, it's, you know, yes, you can ask Claude to do 10 different variants on something, but if you thought about it deeply, maybe you would do choose two that you thought were most promising and the third one if you then had like some iteration on that as well. So I would not say like a token max. Actually, the tokeniest thing was that conversion, uh, thing I did was just like a couple million tokens. There's like a lot of tokens that it took to to convert the, uh, the thing from from Python to TypeScript. Um, but I think people are being more thoughtful about, um, these different pieces and and one of the things we look at whenever we look at a model launch is not just model intelligence, but we're also really thinking about model intelligence and effort and token efficiency as that combination. And I think that's a big lever we have to improve is how do we continue to be more and more token efficient for a given task so that you can also hopefully, you don't have to think very hard about this, we can do this automatically, but like we're able to tune the solution to the problem a little bit more. Um, and then to your second question, yeah, I, you know, when I was still in the CPO seat, I I was thinking a lot about sort of outcome-based pricing as something that would be really interesting to do if you could do it. And of course, if you talk to, uh, like the Sierras and Fins of the world that have like a really clear, like we kept this, you know, we were able to solve this customer request and not have it go escalated. Like that's really clear. Um, it gets so much fuzzier on these like tasks that we actually ask Claude these days. Like I had a strategy document. I used Claude to critique my strategy document. Like, what was the outcome? It's like, well, I don't know. It's like, tell me how the strategy goes six months from now. It feels like it's going to be very hard to to capture that as well. But, um, I would like to see some more experimentation around, can you better capture like what it's worth, you know, to the individual and then or the company and then can we find the best way to to do that as well and I guess the, the most concrete thing we've moved towards that and, uh, we have a product product called Claude Managed Agents where we'll run all of the infrastructure for you in terms of doing all of the, you know, agentic harness and calling the tools, etc. and you can either do it in sort of the normal mode, which is you give it tasks, it, you know, will go through tokens, it'll tell you when it's done, or we an outcome-based mode where you can say, here's what good looks like, here's a rubric, go and do it and, you know, you know, it'll go off and and make it more outcome. So like if everybody had moved on to that API, then I think maybe we could have a different sort of outcome-based, uh, pricing, but we'll see how that gets adopted.
>> John or the guys in the back can do we have the random image? Can we show the random image?
>> Um, if we can, great.
>> I'm excited. The random image.
>> Oh, here it is.
>> Okay, it's just because we didn't have a good label for it, so we just called it the random image because it might come up at any point, but this is a chart from the Financial Times. Um, speaking of utility, where it shows the amount of app releases that have come out, um, which are skyrocketing, and then apps with significant usage that seem to be going down and app reviews which seem to be going down. So Mike, I'd love to hear you respond to what we're seeing in the image here. Um, is it possible that like everybody's coding and releasing, but we're not really seeing a big boom in productivity?
>> That's really, I mean, I think there's definitely a parallel in app usage in general. It'd be interesting to seeing if any of those app releases became one of the apps with significant usage. Um,
>> We could take it, we could take it down. Yeah, go ahead.
>> Um, I think, uh, this ties into something that I've been thinking a lot about, which obviously my background is in consumer and I've been wondering what the consumer AI breakouts will end up being and we haven't seen a lot of them yet. Um, and I think part of it is, you know, I I don't know how far back that that chart goes, but when we were releasing Instagram, it still felt a little bit wild west in terms of the apps. People were excited about apps and like two kind of random people released an app and were able to get to like number one in photos and video within three months, right? I think that is much harder now when you think about how consolidated the top 10 is and how much time spent is spent on like the TikToks and Reels of the world. It's a lot, right? And so I think getting that breakthrough consumer experience, I think is really, really hard. So I think that is, as much a story about how sort of consolidated consumer products are these days. Number one. Number two, how, um, entrenched or how powerful it is to have, uh, that sort of data, people call it data gravity, like the data gravity of something like your Google Docs are in your Google Docs. So even if somebody has a like 2x better AI powered, you know, doc editor, you're going to move all your stuff, maybe, probably not. So I think that it speaks to, you know, the things that are sticky. I think I think about it a lot is like the hard stuff is still hard, like making something people want still really hard. You know, we have amazing models internally on top of not all of our products work, right? And so, um, I think that's a bullish sign for like product people like me because it means that I think we hopefully still add value. But I think that chart is maybe another place which is like, uh, it's harder in many ways than ever to, uh, break through even if you can code more quickly. And could we have done Instagram in, you know, a month instead of three or four? Probably. But we got there after like a long winding turns and twists and turns process.
>> You had I think 18 people at Instagram when you sold it.
>> 13.
>> With these tools, do you think you would have How many people do you think you would have had?
>> It's really interesting because of those 13.
>> Give us a number because like everyone's like, "Oh, one billion, $1 billion, one person startup." How close could you guys have gotten?
>> I think we could have gotten there with like four to six, you know?
>> Okay.
>> Yeah. Um, or the thing that we would have done if we'd grown it, we'd be able to do things in more than a single track. Like Instagram was, if you ever watched my five-year-old play soccer now, by which I mean like the ball is there and every single person runs to the ball. Like that was our product team. It was like video, go and everybody like goes and works on the one thing and like we'd be able to like play positions. Like Android, we built in about a month for Instagram. We could have done it probably in a week with the models and but and to build Android, we took everybody off iOS and we all like relearned to code in, uh, you know, code Android OS and then we went off and do that and like for that whole month, we were barely shipping updates on iOS. So I think you can be a lot more, actually a really good example. There's a Labs project I have internally, um, that, um, helps accelerate how, uh, Anthropic engineers like code and do code review. And that, um, that project, I am maintaining an iOS and an Android version of, uh, and I basically have the Claude that works on the iOS one basically like ping the Android one and be like, hey, I implemented this. Sorry, Android users, it's still the second one, even in ALM world, sorry. Um, uh, and then the, the Android version is like, okay, I'm going to do this. Oh, that doesn't count because that feature doesn't make sense here. I'm going to drop it. And of course, we wouldn't have been able to delegate all of that on on Instagram, but we sure could have done a lot by having sort of platform parity. Like this dream of platform like close to parity is now actually quite doable.
>> You're probably going to get calls now from the remaining six or seven people on your Instagram team going, "Cool, was I did I make the cut in the new era?" Also, it sounds like you probably could bring Gotham back now if you really wanted to with all.
>> I forget if we eventually, I think for April Fools, maybe we brought it back one day. Yeah.
>> Yeah. Do we have time for one more question?
>> Yeah. Yeah. Last question. Yeah. Okay. Um, I mean, my last question for you is, um, you worked on a product that now as it has evolved is is, um, in many ways ethically fraught because of some of the harms that people are concerned about with children. And when you talk about the fact that there hasn't really been a big breakout consumer app for AI, I think there has, and it's chatbots, right? And chatbots have also led to some real dangers and harms for young people. And so when you are building in Labs, how are you thinking about the, you know, the potential harms and the risks that come with just making this technology that much better?
>> Yeah, I mean, I think there there are certainly products that we have either prototyped or conceptualized and been like, this product because it sounds so hypey, I hate this, but like, like this product if shipped would be bad for the world or like would nudge people in the wrong direction or even if we did it right, the like wrong or like more morally fraught version of this would be actively, we think bad. And so I think,
>> Asking that question a lot internally, um, makes a difference. Um, and having, it's a luxury to have core products and models that are doing really, really well. So we don't like that's in some ways an easy decision, even if we think it could get a lot of, um, a lot of use. But yeah, I think going back to an earlier conversation, I think frontloading it is really valuable and really thinking through, like, it is now more normalized to have people at a company, and definitely Anthropic does, who are like economists thinking about the impact of the thing that you're building on the world, and that just was not the case in along the years on on most of social media, I think.
>> Yeah, Mike, it's always great to speak and we thank you again for bringing your insight today and, um, let's do it again soon. Let's hear from Mike and Lauren. Thank you. Great job. Thank you. All right. Are you guys ready for our last conversation of the day?
>> Are you having a good time?
>> All right. Um, so in 2015, Greg Brockman, Elon Musk, and Sam Altman started a nonprofit called OpenAI. And the plan was to pursue artificial general intelligence. The company or the nonprofit when it was a nonprofit back then began in Greg Brockman's living room. And these folks were convinced that achieving artificial general intelligence or AI on par with human intelligence was possible. And to be honest, most people in the valley thought that that was an interesting side project, but most of the attention was on social media at the time. Well, fast forward 11 years and here we are. OpenAI is probably going to go public within the next year at a trillion dollar valuation. They're going to announce likely, you know, because third-party data is showing it, a billion users in ChatGPT fairly soon. Uh, and they of course raised the largest venture capital round in history at 122 billion. So they are at the the leading edge of a technology that has captured all of our attention and is changing the world. Uh, and so when you listen to Greg Brockman, one of the things that you can see even from his conversations all the way in the past is a clear conviction and an understanding of where this technology would lead and where the products would go. It's amazing. You listen to the podcasts from pre-ChatGPT or really in the early days of the GPT models, uh, and you can hear Greg speaking with absolute clarity about where the models and applications would go today. So I think to to close our day, let's take a look into the future of where OpenAI and the frontier is going in a conversation with Greg Brockman. Let's welcome Greg.
>> Great to see you, Greg.
>> Thank you for having me. Um, you know, Greg, this is our fourth time speaking and we've spoken every time about OpenAI's product direction and I think I'm starting to get it. Um, you know, there was this conversation that a super app was the wrong term for what you were doing with, um, the app that you're building, bringing Codex, which is coding side of OpenAI's product browser and ChatGPT together. And when you use the word super app, people would be like, no, a super app is actually something that you can just use every other app within. Um, and now as we've seen these products come together, actually super app might be the correct term. You know, at least for us on the outside, we're starting to see it that when you need to do anything, um, you it will start with a prompt in ChatGPT and then OpenAI's technology will use either your browser or your computer to get that done for you. Is that the right way to think about it?
>> I I think that's a pretty good perspective, right? And I think to really zoom out, the thing we're actually trying to build is in AGI, right? That if you think about what people have been using since ChatGPT, it's a language model, right? There's a big gap between these. It's amazing, you can talk to ChatGPT, talks back to you, great, wonderful, but when we launched in 2022, there was no memory, right? It's not hooked up to any tools, has no context. And so it really is that this conversational intelligence is only one part of what people really need to get work done to be able to achieve their goals. And where we're going is to have an AI that's really looking out for you, right? Who can provide the goals, the directions, that it's constantly thinking about, what can I do for Alex today? That it's able to go and solve super hard problems, very mundane problems. You wake up, your inbox is organized, but also if there's like a health plan that you are thinking about that it can help you help you achieve that, figure out medical treatments, uh, or, you know, sort of back and forth provide you with that kind of information at least. And I think that the question of what's the interface you want, what is the product that you want is what we spend a lot of time thinking about. The answer is you want almost no interface. You want no product, right? You want this to be like, what's the interface between you and me, right? Just being able to talk to a persistent entity of some form that's able to go and accomplish goals for you. And so building that is hard. It will take time, but we have a lot of the pieces, right? We're increasingly bringing together the product layer, trying to make the models better, trying to make the whole system just so there's less like clicking buttons and toggles and changing modes and all these things. Not to say that there won't be some of those along the way, but the long-term trajectory is towards simplification, unification.
>> Yeah, it's very interesting that you say the interface, uh, will melt away. And so to go a little bit deeper with my question, um, many of us who use products like ChatGPT today, we'll see that the bot will make a suggestion at the end. You know, you ask it about nutrition and it says, should I make a health plan for you or make a diet plan for you? You ask it, sorry Ron John, we just talked about travel, but you ask it about travel and then it will give you an agenda for instance. And so am I hearing you right that what's going to happen within ChatGPT, just to give an example, is you talk to it about your health decisions and it might say, you know, you probably need to, you know, go to this specialist, let me make an appointment for you, and then it will go and actually take that action on your behalf. So it goes from simply a conversation interface to actually understanding your intent and then going out and accomplishing that for you.
>> That that I that's exactly right and I think that if you've used Codex, and by the way, how many people in the room have used have used Codex?
>> And people.
>> Yeah, decent number of people. Um, and that our goal is to really bring the power of Codex to everyone, right? To bring agents to everyone. You that technology exists right now, right? You can hook up like I hook up my Codex to Slack, to my Gmail, to my calendar. And there are many people within OpenAI, non-technical users, you know, it's got code in the name, but it's not really about code. It's really about having this general purpose tool using harness an agent. And that the kinds of things, for example, someone on our comms team does is, um, she was organizing an event and it would just ask all of the, uh, event attendees for their dietary preferences, set up a whole seating chart, kind of did all of that work, so that she could focus on the parts that she wanted to and and really thinking about the vision of what she wanted to achieve. And I think that we're going to see this across the board. So the, it's not sci-fi anymore to think about an AI that's hooked up to these tools. And I remember with, you know, our very first attempt at tool use.
In Chatbt was 2023. I think in like March or April or something we released plugins. Do people remember plugins back in early chat days?
That didn't work. It didn't work at all because the models, the models weren't ready. Right. The form factor is correct. Obviously, you're going to have an AI that's able to like talk to your Gmail, like no question. But we could only like have three different connectors exposed to the model at a time or start forgetting. You know, we had like 2K, maybe 4K token context. Like there's just no memory, right? It's kind of like when you had early computers in the 60s or 70s or something, right? You know, that you had tiny little memory banks and today you have your phone that's like better than any supercomputer from that era. And I think that's where we're going with these models, right? That the rate of improvement has been so steep. So now you can have hundreds of different tools accessible that we have the ability to hook them up to whole file system. So you can almost have the full power of the internet and like almost any application you want at the model's fingertips. And it's smart, right? It's got 52 million token context, depends how you squint on it. And the capability level is also getting so, so powerful, right? These models are now solving unsolved math problems and physics problems, right? And really helping people be able to achieve things they couldn't otherwise. Like we are on the era on the edge of this era of agents really transforming how we all operate whether it's in software engineering, finance, legal, sales, and in our personal lives too.
So just to unpack that example that you were giving there, one of your colleagues is chatting with Chat GPT about an event and then suggests, "Hey, you know, uh, how should we, you know, contact event attendees about something?" And instead of like saying, "Okay, I have to do that," and going into an event program, basically what happens is the, the interface will take over from there once it says it's a good idea and you agree, and then hook into whatever tools you're using and then do it for you.
Exactly. So it uses its, you know, Gmail connector, searches through your inbox to find all the people who are attending, and then if you're on the like, what, what is in dietary restriction? Sees, "Oh, these people I already have their dietary restrictions, these people I do not." Um, drafts an email depending on exactly how you have things set up. It might say, "Hey, I drafted these emails, can I send them?" If you have a connector that doesn't even let it send emails, it'll say, "I drafted it, you need to send them." And, uh, in a different world, you could also imagine that it's, you've built enough trust with the system where it says, "I drafted the emails and I actually sent them." And I think that this actually points to a really important aspect of the agentic era, which is trust, right? That we need to really learn how to build trust with these systems, where they're good, where they're not. Figure out what you want to delegate to them and how you want to entrust them with responsibility. And that's something we view as earned, right? It's not something that that we can we can grant, but by providing lots of tools and control and oversight and supervision to the operator, to the person who that this AI is operating on behalf of. Like, we think that that is going to be such an important thing. So that's a key product feature and differentiator.
Yeah. And when you go back to some of the early attempts at this, there was this like move that OpenAI had to let you call an Uber within Chat GPT. And it, it followed a long line of companies that have tried to get you to take action within chat, but it never really took off. And the difference here might be that the chatbot can take control of your browser or take control of your computer and then you don't necessarily have to worry about like, "Is this plug-in going to work?" It goes and accomplishes that for you by taking over your machine. So I wonder, you know, if you expect a fight from the user interfaces that we have today, aka like all the other apps, all the software, where to be truly useful, Chat GPT will have to not be blocked to be able to go out and execute these actions on behalf of a user.
Well, look, first of all, I'd say that it's, this is not theoretical at this point, right? That people have been using Codex. So, it's a separate product, separate app, you have to install it separately. Really starting to focus on software engineering. But the amount of non-software work that has been happening in Codex has been absolutely exploding, right? It's been this like incredible exponential curve, exactly the thing that you, you would expect. And within OpenAI, we basically have the same level of penetration now in usage as Slack, right? It's like everyone, and OpenAI is like an entirely Slack-based company. We do not use email for the most part. It's like, really like, if you're not on Slack, you're not going to do any work. Um, and it's kind of feeling that way now with Codex app as well. And that everyone's Codex is hooked up to all of these tools. Um, how the EOS ecosystem evolves, I think it's going to be a very nuanced thing because I think one thing that is very important is that we believe that there should be an ecosystem that gets to be vibrant and thriving and that people can really build and see the benefits. And so we've actually seen this from, uh, partner companies where, uh, you know, that we, I remember there's a couple different partners where we said, "Hey, we really want to train our AI to be really good at using your software." And we didn't know what they would say. And actually the response we got is, "This is the most partner-friendly outreach we've ever had." Right? That the idea that you will make your AI specifically good at using our tool, and they just see the opportunity because their tool will be used just so much more as a result. And that everyone is trying to think about how do they not just survive as a company into the AI era, but thrive? Like, how do you really get the advantages of the fact there's going to be so much more activity? And if you don't have AI in there, if you shut it out, then you're actually going to be declining, not thriving, right? This, this kind of makes OpenAI puts OpenAI so, first of all, you're going to bring, you talk about people using Codex. So, one of your colleagues shared, and I think you've talked about this too, that you've brought Chat GPT into Codex. So you can bring Codex into Chat GPT, which is basically like, if we're users of Chat GPT, this experience that we talked about of Chat not only suggesting what you might want to do next, but going to do it for you, that's going to happen. And so it makes you effectively an operating system, don't you think? But not the operating system like an iOS where you would like go open up your phone and then tap different apps. It's almost as if all interaction with all apps will happen through this interface. Is that the ambition?
I think that you could describe it that way, but I think of it differently. Like the way that I think about this is that what is the ideal interface to an AGI, or we call it kind of a personal AGI? And I think that it's again, the same interface that you and I are using right now. You just want to talk to an assistant, right? You want to talk to something that can go and and and work and and operate on your behalf. And so that, yes, like that agent, that AGI, that AI will have its own computer, right? It will have its own access to things that maybe. And, you know, like an ideal co-worker would be, they can come over and type things on your computer too. So some access, some delegated access to your own own system. And, you know, maybe you delegate access to your inbox sometimes, maybe it has its own inbox with some sort of, uh, some sort of, you know, window into into the things that it needs. You forward emails to it. Th these are not actually, if you, if you think about this, is not unprecedented, right? It's like the way that you work with an assistant who's a person, that we've, we've actually, or any co-worker, really, we've spent a lot of time really thinking about how do you build these trust boundaries and make sure that you're able to operate together. And so I think of it as just a different thing. It's not, it, you could think of it as an operating system, but an operating system is almost something from a different time, right? It's a different layer of the stack. This is really more about how do you interface with technology broadly. And I think that the beautiful thing about AI is it's really about bringing the machine closer to the human rather than us having to contort ourselves into like files and folders and like all these details that somehow are not natural, right? That are more about how the machine operates rather than how we operate.
Yeah. Talking about a personal intelligence, it sort of, um, I don't, did you watch WWDC last week?
Uh, no, no, I missed it. I was, I was banned, but, um, I watched it on TV. Um, come on, Apple. Anyway, it does look like you and and and Siri, the new Siri are going to come kind of into competition, right? Because they're an app that's going to sit, or an intelligence that will sit on top of all of your apps and let you take action. And Chat GPT will be an app on the iPhone. So then talk a little bit about whether that positioning is going to be difficult for OpenAI and how you're thinking about that strategically.
Well, I just think again, think of it a little differently. Like I think that we're in the beginning of this new agentic era, and the way that this has always gone in AI is that when you have a new level of capability, it means you have an opportunity to rethink everything, right? Rethink how people interface, how like what the tech is capable of. And I think that this is no, no different, right? In my mind, like the kinds of things that I see on the horizon, for example, AI for solving scientific problems, right? And I think we're starting to see the inklings of this. Like, for example, today we announced we have in, in, uh, peer-reviewed literature, people, doctors who are using 03. Remember 03?
Yep.
That was like forever ago now, right? That was like one of our earliest reasoning models using that to find diagnosis for people who I had no, no answers from doctors for many, many years. You know, there's an example of someone who had spent 20 years with a mysterious ailment, finally, it's been diagnosed through the use of this technology. And if you're like, "Okay, you've got models that can do that. They can do that." And then it's really about like, you know, the same like distribution and like, you know, can you get access to an app? You know, to me, it doesn't type check. It's like we have something fundamentally new. And so that's not to say that there won't be competition. I actually think that there will be, and it's going to be great for everyone. But I just think that the ways in which you're going to use this technology, the things it will be capable of, and what it'll make you capable of doing are just totally different from anything we've seen before.
You know, I was going to ask you, um, well, does it mean that you'll have to, um, you know, create your own device, assuming that like my concept is, you know, that you're going to have to go through Apple to get to the user? Um, assuming that's somewhat valid. But the answer is, you already are, right? You're so, OpenAI is working on a device right now.
It certainly has been publicly reported.
I, I was in your office in December and Sam told me that this is happening. It's, it's multiple devices. Um, so if you think about the way that again, you're going to interface with, um, with these AIs, how does that device play in or series of devices?
Well, look, I, I think again, I, I would just step back and say that I think this is the beginning of something very new, and that I think about the way that I think, just say, like, I think the biggest shift that has happened in terms of interface, again, it's not even about devices and and and things like that. Really about the shift from conversational intelligence, like kind of the chat paradigm, where it's like, kind of you have an AI that's personalized enough to you that it's worth reading its output, right? You ask it a question, you get an answer, it's something that's useful to you, to agents where they're capable enough to actually do things for you. Like that is a big shift, and that that implies a difference in how you want to interact. And so you kind of are just going to want a single agent that has access to your context. And this will be true in personal life. This will be true in a business context, right? You imagine, for example, having a, you know, imagine you have a PhD in every field co-worker, you know, Nobel prizes, multiple of them, and you hire one of these, you hire a hundred of them, and you don't invite them to any meetings, they're not going to be very useful. And so there's something about how do you get context into the AI, and not just statically, but dynamically, right? As context evolves, as your business processes evolve, how do you have a context layer that is accessible to an AI that lets the AI operate to the extent of that raw intelligence. And so finding ways to make that AI be accessible, so available in your meetings, to make it very ergonomic, it's very easy to get access to. Um, I think all of that's going to require a rethink. But I think again, it's just the core for me starts from thinking about the agentic form factor and then working backwards to how do you just make this have the context it needs. And again, the trust is going to be such a core part of making this whole equation work.
So, kind of like having this, this device with you at all times and being like, "I need to get that done," and it goes and does it for you.
And I think that that will be part of it, but I almost even think if you don't have a device like that, it's not like you're going to be out of the game, right? Because it's this, a, this AI. It's not because there's one thing, there's one version of it where you think of it where it's like the device is the AI, and you want your phone to be the AI. You want that, you know, whatever, whatever, you know, custom device you're, you're thinking about to be the AI. But it's not going to be like that. It's going to be more like an interface, like no more than your phone is you, right? It's an interface to you. It's a way that I can sort of, you know, call you up whenever I need you, uh, whenever I want to ask you a question. And there's different ways of accessing, right? There's like synchronous phone call, I can text you, I can email you. And I think that we're going to be much the same with how we interact with our agents.
Um, there's been some reports that OpenAI is working on these like bidirectional voice models. I think we've talked about that in the past that like the goal is to have like an AI that you can, you can speak with and we'll be able to process that and speak back with you in a much more natural way. Can you share anything about that?
No. But no, more seriously, um, I mean, look, I think that the, that the general shape of the, of the technology, like the way that, like we had, we had, we've had voice models, um, you know, kind of a really cool voice experience for, you know, year and a half, two years now. Um, you know, we first demoed it back in March, April of 2024. Um, brought it to market, um, you know, maybe late that, late that year. And the way that it works, and the way that everyone's models work, is that you basically chain together, um, well, the original way that these things worked was that you would chain together a text to, or a speech-to-text model, then you do a text-to-text model, and then you would do a text-to-speech model. Horribleness, right? Like these three things chained together. Um, it still has been the case that even if you have one unified model that's able to kind of take in input and then, you know, able to output a response, you still have this problem of turn-taking, right? Imagine that, like, we have this, like, you cannot overlap, you cannot interrupt, it's just like, once you, you, you speak to me in a turn, and then you got to wait for me to finish my whole response. That is not how human conversation works. And so that, we, we basically have like a hack where we have these models that determine, "Oh, it seems like the turn has ended," and, "Oh, it seems like the turn has started." And we're like, "Why are we talking about turns, right?" Like, turns again are so unnatural. This is the humans contorting ourselves to the machine and its limitations. And so the obvious thing that you want to accomplish is a model in AI that works much more like you and I do, right? That's able to process input at the same time it's processing output. And all of that is of course something that many people in this field are trying to to run towards. Um, I think it's going to be very, very exciting as you move to these natural, very human, fluid, like conversational interfaces. No one's seen anything like it. Like, one thing that, that I think about is the, the current interaction with, you know, Chat GPT voice. In many ways, it's magical, right? So many people use it on their commute, able to ask all these questions. But it also is so frustrating, right? Whenever it breaks the magic, because it's like, you realize, "Oh, I want to like, add some follow-up," and it keeps talking over you, and it didn't, it's just like, that is just, it doesn't make sense. And so I think that part of what we need, part of like the whole point of this AI is to be something that you can interact with, interact with fluidly and naturally. And by the way, I think it's not just going to be about the sort of use case. Like, we kind of think about the, the personal use case, but it's also really the work use case. And I think some of the most magical experiences that, that I've had with Codex have been when operating it through voice. Like many people, we have a voice built in. Some people use third-party apps for it. Um, and that you just get a very different experience when you start to realize that like typing a quick message to give some feedback, easy. But like writing out a whole paragraph and everything you want, horrible. No one wants to do that, right? You just want to be like saying things and you want the real-time feedback loop, and all of that is going to happen, and it's going to be amazing.
So, let's talk about model improvement briefly. Um, so there was a discussion a couple years ago that large language models were about to hit a wall. That was wrong. And, um, something, you know, that I'm thinking about is, I think we're all thinking about it, is how much better can these models get? And when will the improvement stop? Any thoughts?
Well, I think that this is a place where when you're kind of building these models, you get kind of a sense and an intuition that I think is harder to get from the outside because we see all the data points and we see also the work that goes into these improvements. And so that there's two parts to the answer. One is, I think that the fundamental science is one of the most mysterious and important, just scientific discoveries and empirical observations that that I can, that I, that I'm aware of, that I can imagine, right? That we are able to actually build these models and that the scaling laws continue, right? That it just is the case that you can just keep training these models more data, more compute, better architectures, and there's a lot of improvements that go in. But every time we've kind of run into a like, "Huh, this isn't quite scaling the way we expect," it's, we have a problem. We have a bug. That, "Oh, our math wasn't quite right." That, "Oh, our implementation isn't isn't isn't quite matching the math." Whatever the thing is. And that is, I think, a very important thing to to sort of internalize. And actually, if you, we've done studies where you go back to the beginning of the field, right? That neural nets themselves were designed in like 1940s, before computers, as a model of maybe this is, maybe this is how the brain processes information. First hardware implementation was 1959 with the perceptron. And if you look at landmark results in the field, that the landmark results follow this incredibly smooth, deterministic path of more compute being poured into them. And so 70 years of people, maybe 80 years now, of people saying, "This stuff is never going to work, never going to scale, going to hit the wall." Hasn't hit the wall yet. There's still no wall in sight. And so I think that the fundamentals allow it. Now, the practicality is hard, right? Actually building these massive supercomputers. It's hard. It's expensive. It's not easy, right? That we have teams that just like work so hard to solve these incredibly hard technical problems. We have our own network protocol that we've had to design. Um, that we have people who look at every single layer of the stack. That there's weird wiggles in the graph and you, the way to think about these neural nets is that there's like, there's no abstractions, right? It's almost like any little piece that's wrong can have a ripple effect that only shows up down there. And so you need people who deeply understand all of it. And yet, if you get the right team together, put the right mission in front of people, and people do that grind, the outcome, it's worth it, right? And it's achievable and it's possible. And so I think that for those reasons, the progress will continue.
So then I'd love to hear your perspective if, if models can basically progress much further from where they are today. Let's say, let's say OpenAI builds the best model, and it's the equivalent of like something with like 15 PhDs with excellent emotional intelligence that doesn't complain and goes out and does stuff for you. And then the next model maker will build a less good, but it has 13 PhDs and it's like pretty good, uh, you know, EQ and will still go and do things for you. So, where does the differentiation come in when you get to that level of intelligence? Because we've seen the model makers kind of move in lockstep. One makes an advance, the next one comes in and makes the advance. So they all become that smart. Do they, is it possible to differentiate?
Well, I think there are several dimensions to the answer. Um, number one is, I do think there's a bit of an attractor state where, just like from a business model perspective, every provider sells out all their compute. Okay. Right. I think that is just like the world that we're heading towards where there just is not going to be enough compute to serve all the demand. Right? That we're heading to this compute-powered economy, that everyone's going to be using these models all the time to be able to accomplish tasks of interest. And we just see it. It's like, right now we're talking about compute constraints and like, like the number of people using these agents is like, order of 10 million, 20 million, maybe, you know, it's like, we're not at planet scale. Chat GPT is like a billion users, right? But we haven't brought the agentic power there yet. So you're just looking at these factors, and the depth of usage is also tiny compared to where we're going. And so I think that we're just going to be in a world where even if you have different vendors, different capability level, open source models, all these things, these neo-clouds, like I think that compute is just going to be this scarce resource, and I think that it's going to go, go to use. Um, so to some extent, I think that the, like, is this a good business to be in and for new entrants to come into and things like that? My answer is actually yes. I think that there is like a huge market, um, that we are just not going to be able to address, and we need much more energy, momentum there. Um, but a second thing is that it also misses the fact that intelligence is not a undimensional thing, right? That, if you really zoom in, being good at different domains is something where even if you, you have a lot of raw intelligence, getting good, if you've never practiced, like you've never actually done a pitch or something, like you're not going to be good at it your first time, right? And that there's lots of different, you've never operated a spreadsheet, right? You're not going to be able to succeed at doing some complex modeling. And so I think that there is something that we have been internalizing, which is that we look across different industries and different domains, and we have to prioritize. We can't possibly be great at every single area at once. There is definitely a lot of like, "Hey, you just get the general intelligence up, and it'll experience a lot of these things," but to really become a domain expert, to really be that PhD, and to really be something that can help push forward the ambition of a field, like that's hard. And, and by the way, one thing I also want to say is that I think understanding what happens when you successfully do that, I think that having a good mental model of that's important, which is you look at something like rewind to AlphaGo, right? You remember move 37, this move that like changed people's understanding of the game, and then now more people play Go than ever, right? And that it actually inspired people to do even more. I think we're just going to see that. And so I think that the depth is never going to stop, right? How deep can you go on science, right? Like, I, I think that people have thought sometimes that, "Hey, we found out all the physics, it's all good, we're all done." And I don't think that that's the future we're signed up for. I think we're signed up for one where we're got to keep find, every time you unlock one mystery, every time you solve one mystery, unlocks like 10 more, right? So, I think that there's just going to be so much more to do and tons of room for differentiation across different companies.
So, I think I'm reading you right, and that your belief is, um, maybe there's a way that everybody can scale up these models, but ultimately the company with the most compute is going to win. And, you know, we spoke a couple months ago and you had mentioned that like you were asked internally how much compute should we buy, and you said, "All of it." And they said, "No, really, how much should we buy?" And you said, "No, buy all of it." And, and OpenAI is definitely the leader in in buying compute. I mean, we see the, the, money going out, obviously a lot of money coming in through investment and, and now you, you've built a business with customers, but there's a lot of money going out. Do you ever wonder, hey, are, like, do you ever wonder maybe we're not going to be able to pay all this money back because it's a brand new category?
Well, the way that I look at it is on the fundamentals, right? You, you need to really look at the fact that the way that compute goes is that it's multiple years out before compute actually arrives, right? Depending on exactly what you're doing. For example, we've been investing in our own chip program now for multiple years, and super exciting progress, like, you know, we'll, we'll, we'll have more to announce, uh, actually pretty soon. Um, but the fact that we're able to do that is something very unique, right? Really thinking about the full vertical integration of the supply chain. And I think that the world we're heading towards is one where again, there's just not going to be enough compute in the world to satisfy all the demand. And we see this very concretely, like you look at the exponential, I mean, rewind to the exponential of Chat GPT. Um, look at the exponentials we're on now. Uh, you think about the problems that we, that we are able to solve. You know, it's, it's actually kind of interesting that, yeah, we just yesterday announced, um, it was two days ago announced a new result in, uh, basically chemistry and being able to synthesize, you know, new, new improved reactions. And all of this is without much attention. You know, the thing I just said of, if you go deep in a domain, you can really transform it. And we're not even scratching the surface yet. And so the way to think about is the economy is so massive, right? And we see it very concretely in terms of our own growth, in terms of what people are willing to pay, and kind of the size and growth of this whole industry. And so I think that the thing that I think about the most is how do we meet the demand? How do you actually have something that can help support all of the work that people want to do in the economy? And I think that is such a vast thing. I don't think any of us have internalized it yet.
Yeah. But if I may, there is a price war brewing. I mean, at least that's according to the report. It's great to have you here to talk about it. The Wall Street Journal recently had a report that an upcoming OpenAI model might have significant price cuts. And so again, like how can you know, if it requires so much resources to serve this demand, and it is growing demand, um, in an environment where there might be price cuts, how do you make that math work?
Well, again, I look at it from a different angle. So if you look at the whole history of what we've done, we actually have been increasing the intelligence, cutting price, right? For a fixed amount of intelligence, and people somehow just like, like the Jeffans paradox, just keeps happening. And so I think frontier intelligence will always be something that is going to be, you know, it's always going to be the priciest thing. But I think that a year from now, that level of intelligence is going to feel pretty mundane and like, you know, going to be much more available. Um, and I think that that, the world that we're in is one where people are starting to really think about value. And it's actually been a very interesting shift where over the past, you know, first quarter, maybe up until now, people have just been like, "This AI agent stuff, it's all new. We need to bring into our enterprise, like we don't want to be left behind. How do we be part of this, this future?" And now people are like, "Okay, like, let's make sure this is actually delivering ROI and value." And I actually think that's a great place to be, right? Because people are asking the right questions. And I hear this, I, I had some customer meetings today where people were saying exactly this. They were like, "How can we have even just like good spend controls? How can we have observability?" And I, I think we literally today just released spend controls. So, you know, it's like, "Okay."
Exactly. Yeah. We're, we are really investing hard in enterprise readiness and the tools that our customers are telling us that they need. And I think that that, for me, is the shift that we've also been going through as a company is really not just thinking about, "Hey, we're just going to release models," and, you know, have a model, really thinking about the end-to-end of the business. How do we bring this into solving real problems for real customers? And that is happening so quickly across every single industry. And the number of different companies that still feel like they're wrapping their mind around how to best make use of these models, we're learning at the same time. I think it's just so early in this whole game to me, the, like, the absolute size of the market growing so quickly, our own revenue ramp growing so quickly. I think it's still just like, none of us are anticipating how steep that's all going to go.
Are you going to cut prices?
I, so again, the answer is always yes, right? But it's about like, I think that what's going to keep happening is that we're going to have frontier models. Um, I don't think there's going to be like a massive shift in, in the short term. I, I don't think that that is the kind of thing that's going to happen. But I think the thing you should anticipate is that over a year-long time horizon, to get to today's level of intelligence that feels very premier, it's going to be much cheaper. But there's going to be a new thing that is going to be so much better, and you're going to be like, "Why would I ever use this other one?" Right? It's just, it's just how it's always going to be.
So, Satya Nadella has had some interesting tweets and interviews recently. He recently said the model is becoming a commodity, and the valuable asset is company, or this might be a paraphrase, the valuable asset is a company-specific AI system that continually learns from your data. What do you think about that? And is it weird to be competing with Microsoft now?
Well, look, I don't think that there's any layer of the stack here that is going to just kind of be removed from the value chain. I think that these things multiplied together. And if you think about the most, the base layer of compute, right? That is something where it's just like, no compute, no AI. And to some extent, you can say, "Oh, compute is commoditized, it's just flops, who cares about it?" But in reality, like, you look at, you look at today's chip stocks, you look at the, you know, people who are selling compute, kind of what, what is what the market is valuing people at, and they see that there's a fundamental asset here that is just so critical. And I think that is because it is a revenue center, it is something that anyone who's building AI has to rely on. And that there's a bunch of very interesting dynamics in terms of the efficiencies that you can squeeze out and the margins, all these things. But fundamentally, even though it's like, you can kind of squint at it and say it's commoditized, it's not. It's, it's, it's not that, that the value goes away. It's not that the margins go away. It's like something the market will reward because it has fundamental value, and that the importance of it's going to go up over time. You can see that with some of the prices that people are paying for H100s, right? Hoppers are, you know, kind of, you know, not, not obsolete, right? But they're, they're, they're a a previous gen chip. And in any normal situation where you're not totally supply constrained, no one would be buying them. But instead, the market prices are up relative to to where they were before. So there's this inversion that's happening. And again, I think it's going to keep happening where because everyone has this avalanche of demand, that you're going to see prices and margins and all of these things continuing to increase at various levels of the stack. I think the same kind of applies for models where the models themselves are also again, they're not, there's, there's a lot of competition there, and I think that's very good. I think it's good for the enterprise. I think it's good for customers, consumers. Um, but I think that there's a lot of areas where, for example, our models have always been the sort of smartest ones, right? The ones that are able to solve these incredibly hard problems. I think we're just starting to reach a phase where you're going to see the transformative impact from that, right? It's like, if we're really able to speed up science through models, the smarter the model, the faster it's going to go. And it's very, very different from a model that has a conversational interface that you're able to, you know, is able to book your travel, right? Or organize your calendar. So that's also a dimension I think we're going to do a very good job in. But I'm just saying it's, it's a different area. Um, and then I think that the question of, well, how do you actually connect the intelligence to your own customers, right? To real value, to you have all these enterprises that have built incredible businesses in different domains, and it's a huge thing, and it's not something where if you don't have domain expertise, that you're just going to be able to do, right? And part of it is that you need, you think about regulated industries, you think about any area where there's like, you know, think about education where you have a parent, you have a teacher, you have a student, you have these different parties that need to interact in very thoughtful ways. For all of these areas, all of these domains, that there's a lot of value to be built by being in that area and thinking about how the workflow should work, how these models should be orchestrated. And so I, I really think that there's more than enough to go around. And I think that we have to work together as a whole ecosystem in order to deliver the kind of value that I think is possible from these systems.
Okay, just to go back to the Satya point one more time, uh, he's called models a commodity. He's trying to build his own frontier intelligence. He's telling potentially your customers, "Hey, you got to come work with us because we're going to help build these loops that will learn from your data." He's got access to your IP, I think, till 2032. So, how does it make you feel to hear this coming from Satya?
Look, I think that the most important thing that is happening right now is the usage of AI in the economy to really transform the economy and to uplift everyone. And so, I think that that is something that I'm really focused on. And the more that people are trying to make that happen, I think that that's better for everyone.
Um, GPT 5.6 is rumored to be on its way. Supposed to be this, just a Twitter rumor, but I'm going to read it to you. Um, always the best rumors. Three times cheaper than Fable, up to 1.5 million token context. Uh, stronger agentic coding workflows. Um, how much of that is true? What should we expect for GPT 5.6?
I mean, look, you should always expect better, faster, smarter, the whole thing.
So, everything confirmed.
Definitely believe everything you read on Twitter. Yeah, maybe not. That has actually been a source of problems in my personal life.
Um, okay. So, uh, I want to end on health. Um, you brought it up a couple times. You actually had a question in the audience about it earlier. Um, you know, sometimes there's a story and you read it and you say to yourself, "I know this person is speaking to the media, and I know that what they're saying sounds like maybe it's true, but there's something wrong with the story, and we're not going to see more of it." And I've read a couple of those recently. Um, one is, I think, is it your friend, the GitLab CEO, um, Sid Sberandage? He had, um, he, he got cancer and used, uh, he got all the diagnostic testing, uh, he could have. So just went out and tested like crazy and fed that data into Chat GPT with the assistance of some people who had built purpose-built applications for it and was able, I don't know if cure is the right word, but to beat back the cancer to a degree. There was also this dog, Rosie the dog in Australia. You guys heard of Rosie? Like the craziest story where this guy, I'm gonna get some detail wrong, but a guy, um, biopsied his dog, which had cancer, um, ran the mutations across across AlphaFold, then was able to design an mRNA vaccine that he injected into the dog, um, with the assistance of chatbots to build this thing, which ended up being able to jump over tables again, and the tumor shrunk. Um, when we think about the future of AI and health, um, help us sort out the truth with this question. Are these a couple of outliers that made good headlines, but there was something about the story we weren't hearing, or is this going to become standard in the future?
Absolutely going to become standard. Absolutely. And it's, I, I personally have a number of friends who have done very similar things of, get the data, right? Your health diagnostics and use Codex, right? Use these models to get insights from them. And I think that there are many people, like, I think that there's about 230 million people each week who use Chat GPT for health queries, right? And that's been, that's, that's like a staggering scale, right? And this is people, sometimes you upload a scan, sometimes you have doctors who are telling you conflicting information. And I think that we've been in a world where patients are not empowered, right? Patients have to be the doctor, right? You're the decider, you are accountable, right? You know, doctor makes a mistake, and you're going to be paying the price for the rest of your life. Like, it's just, it's a very different kind of incentive. And this is very personal for me. You know, my wife has a number of health conditions, and I think that we've just been, we've not, like, I don't even know how we'd be able to manage many of her conditions right now without the use of of chat. And I think we're just at the beginning of this journey, right? That I think that the degree to which, even if you have the best medical team, the best acts, the best, best experts, there's only so much that can be done, right? You think about the things that are just outside of the reach of humanity, or even just sometimes it's like someone didn't even read the chart, right? Kind of missed a detail. All of that, we should be able to improve massively through these tools. And so I think that the personalized medicine, and sometimes it's going to be about drugs and drug discovery that are of, you know, for, for mass market, but sometimes it'll be even for the kind of NF1 things, like the, the disease, um, diagnoses that I mentioned earlier today. Sometimes it will be for just like trying to understand conditions and trying to come up with with new potential therapeutics. All of that, we're seeing it happening right now in front of our eyes. It's not theoretical, it's really happening. And so one of the most, I think it's like, one of the most astounding possibilities of AI is how much it can improve our health. And you think about the ripple effects of the system, right? Where so much spending on the health care system happens right now, that's a massive part of the economy. And that if you're actually able to help people prevent issues, right? To, to get ahead of potential, you know, health, health problems, that's something that actually then alleviates a lot of burden, a lot of strain. And we're in a world where doctors are burned out, nurses are burned out, like there's like a real crisis that's happening in front of us. And I think AI will be able to help with all of that. Like, we have that potential if we deploy it and use it wisely and well. And so I think that applying AI to medicine, like that's something that I, is really a personal motivation for me in in thinking about this whole journey of of what we're building, what we're trying to do with OpenAI. And I'm hopeful that we, as a, as a world and a community, can make the most of that. Let's hope. I think we will. I'm very, very confident.
Greg, thank you so much. Thank you. Thank you. Thank you so much. Oh my god. Thank you everyone. You have a good time today.
Thank you. Should we do it again next year?
Going to come.
All right. Uh, let me, uh, give you a, a perspective about what we're doing for the rest of the day. First of all, if you're backstage, Cali and John, can you come out here for a second? Um, as I, as I bring them out, as I bring them out, let me, let me, uh, explain that we're going to go from here to the rooftop. We have some drinks and food ready for you guys. We also have a space directly outside. Uh, so that is a place that you can hang out. Um, figure out what's what's busier and try to spread yourself a little bit so we'll have room. Uh, John, come up here. Cali, come up here. Um, guys, this would not have happened without John Bashu and Cali Win. Um,
Thank you, Cali.
Thank you very much.
So, uh, true story, I announced this show on a podcast live and I thought we would get a lot of people signing up. One person signed up and I called, I called both of them and I said, you know, it was great brainstorming this idea with you, but we're not going to do it. Um, so we'll see you next year. And both of them said, "Let's do it, and let's do it at the Commonwealth Club." And we did it. So, thank you guys, and thank you all. Before we leave, I have to thank our sponsors at PWC, Dallas, and Arita. You guys are amazing. And thanks to the PWC crew. Um, just a great group and amazing people to work with. And of course, thank you to the Commonwealth Club here for collaborating with us on this event. Wow. Uh, six years after I showed up here and there was literally nobody in the audience, and we kind of have social distanced and probably would have gone to jail if London Breed found out. We're all here together, and it means the world to me. So, thank you all. I'm getting emotional here. Um, thank you for helping me live the dream, and let's go party.
Thank you.