📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Cal Newport on Mythos and Anthropomorphization | Better Offline

Better Offline58:54

Transcription

Hello and welcome to Better Offline. I'm of course your host Ed Zetron. As ever, support your neighborhood Zitron by subscribing to the premium newsletter. Discount link in the episode notes, of course. Buy a t-shirt, download a blog, where whatever it is you want to do, okay? It's not up to me what you do.

But today I'm joined by the incredible comside professor and commentator Cal Newport. Cal, thank you for joining me.

>> Always a pleasure, Ed.

>> So I kind of wanted to start with I asked you for a quote a few like a week ago, maybe two weeks ago. I can't remember how time works anymore, but it was around the way the reporters cover AI and how it seems that a lot of the reporting is kind of directionally true rather than actually true.

>> Yes. And I want to add something to it since. So I I've been thinking about that quote. Yeah, I've been thinking about it. So what I said, if I remember that quote properly, what what I was saying is I was I was picking up a lot in the reporting on AI that you would lean into a story without having necessarily verified that the details are true and that this is what's actually going on, say, with the new AI model, you would lean into it anyways because it was what I call directionally correct. it it makes the general point that you see it as your job as a reporter to make which is hey you need to be worried about this or this is a big deal right and so I think that is a problem there's another issue I'm seeing though I've sort of been refining my thinking on this.

>> I'm also wondering if some of what I'm seeing in some of the the reporting on this >> is is just a embrace of the form of I'm gonna give you a stress wave with no relief just like we're all gonna take turns s just I will choose an area you haven't thought about. How about mathemat mathematics are going to go away. Mathematicians are going to be okay, I'll take that one. Yeah, let's >> negative clickbait.

>> Yeah. And they but there there's this weird sort of passivity to it where it's like I'm just going to sort of it's it's I call it like head shaking dumerism. You're just like it's this this field's just going away. What can we do? Like just this sort of like passive head shaking. It's a very specific style. You don't see a lot of other reporting historically. I think that takes on this resignation of I'm just going to make the case that like you're screwed and then kind of give you a shoulder shrug and then we're and then then we're going to drop the mic and walk off. And I'm kind of getting tired of this. Like I think there is a cost >> to stressing the hell out of people. I mean, I'm getting letters all the time now from people. They'll say things like, "I feel like I'm trapped in a cage just being hit with wave after wave of stress and there's no outlet. There's no door or possibility of making things better." And I think the CEOs are doing it and I think in increasingly we're seeing commentators doing it as well. This is not good in many different ways. So I don't know. I'm I'm adding that to my list. Some of it's directional true reporting like you they really are worried that people aren't worried enough. And I think it's just sport now. Can you find an area that come in and just write a headshaking article that's only trying to undermine the existence of this like important human activity or this job or our lives or whatever. It's a It's a very unusual style that quickly became a standard.

>> And I I see it a lot with anything to do with AI and job studies. Like I've been sent this Tufts report where it's like, "Oh yeah, AI affected or I AI." They find these weird weasel words where it's like jobs that could be at risk from AI at some point and we put them in one bucket and then jobs that might one day be. We'll put that in another bucket. And there you go. Don't know what we're, like you said, don't know what we're meant to do with this. Don't know what anyone's meant to do with this information, but it's just like, well, there you have it. There you have it. We're all [ __ ] It's It's the It's the end. The job even though the data does not say that. Like I've read I think every AI jobs report now. Every single one. And they're all the same. They are all right now AI can do this and then you look at what it says. It's like it can do law. Well, it can't really do law. It can do one sigma within law kind of. And even then isn't really obvious. And the people saying it can do that are partners at law firms that don't write motions or don't do like the grunt work. So it's it's almost it's it feels like the reporters have either ga given up or are just looking for clicks. And it's hard to tell sometimes. This is what I'm trying to figure out because I'm I'm realizing if it's entirely just I think this is directionally true and that's good enough. Then they should be way more upset and in the streets and sparking a revolution, right? Like if you actually really believed 50% of the economy was going to be automated, that we're going to have to uh have government checks just so we can afford to buy the cat food to eat after all the jobs are gone. If you really thought >> that our entire infrastructure is about to collapse or that super intelligence was going to emerge suddenly and be a threat to human existence, >> you wouldn't just write a sort of too cool for school headshaking resignation article. You would be like, we got to where are the the John Connors, right? like we need to get on the the cool trench coats and get out there and and go against the Skynet revolution. Like you would be on your feet. You'd be, you know, nothing would be more important to you.

So this is the this is my case about the tech CEOs. I think there there's a moral hazard that I don't think that we're putting our finger on properly here, right? So you have the tech CEOs in the AI space that'll just they'll just come out and just drop these bombs like >> Yeah. >> White collar blood, even though he never actually said that. That's Axios putting words in people's >> mouth. That was Axio. I thought that that was a he defined day Wario. He did say 50%. But I don't >> He said 50% but not blood.

>> I thought he said the blood bath. That's my bad.

>> Well, I trust I I this New York fact checkers figured that out for me. But uh Axios does a lot of this where they put like uh these really quotable quotes in the headlines about articles on interviews or speeches given by AI people. And it turns out the thing in the headline wasn't what they said. It was directly what they said. Anyways, um so they're out there making these big statements. These jobs are going away. Uh the internet as we know it is about to all fall apart because of mythos is going to have this this new capability. The super intelligence is coming. I I'm I don't even know what's going to happen. There's two possible things going on here and both of them are morally bad. One is which is the one I think is true which is this is largely marketing. I look this this works. It gets reported. It keeps us seeming inevitable and important. In which case, that's a huge moral hazard because you are making many, many people, normal people, stress the hell out, >> actively scaring them.

>> Actively scaring them. The other option is you actually believe it's true. Well, this is an even larger moral trap that you've just fallen into because you are now perpetuating something that's going to cause exponentially more harm. You should be the very first person shutting down your company and trying to get the other ones to do it as well. So, it's this weird moral trap they've set up where whatever is actually going on here, if they're coming out here saying these things. It is bad. This this can't possibly, normatively speaking, be the right ethical behavior to be out there saying these scary things all the time because either you need to be building the barricade or you're just scaring people for the marketing. Neither of these, I think, is something that's defensible.

>> I have a third and worse option, which is I choose Axio. I think Axios I there are some good reporters there. I think the leadership over there is disgusting. I think that they are aligning themselves with the companies. I think that what like if you watched it was it Jim what's his name interviewing Sam Alman these I think that there is a level of and I I would put this across people like Kevin Roose and Casey Newton these are my words not C's um that they're aligning them that they're saying we think this is going to happen and we're here to tell you great news this is good news for me the writer because I will be safe somehow I will be fine you will not you should be scared but it's also a good thing cuz economy marketing market good and it is it's a very incoherent message cuz it's like to your point yeah if this was a virus like a pandemic you wouldn't be writing hey millions of people are going to die pretty good right hey be good we'll have less people that'll be good right it would be seen as peculiar.

>> Someone did write that someone did write that by the way.

>> Someone did write that.

>> They did say I remember early pandemic someone did write hey you know what this is good for the planet did it like, "Hey, we're driving less. This is great." And we're overpopulating. Like, uh-oh.

>> I mean, I mean, that's a different conversation that maybe I But in all seriousness, you didn't have mainstream media being like, "Well, CO's going to kill everyone. The end, I guess. I guess, you know, maybe we'll just be inside forever." You didn't have this kind of stray. In fact, you had the direct opposite. It was we need to get outside again. Who cares about this thing?

>> And it was just >> Yeah. Go on.

>> Yeah. I think that's an interesting analog. I want to just pull on that thread a little bit because I think CO gives an interesting I think it gives two different interesting observations that go in both directions. Right. So, um I think you're definitely right. What you're saying is when the pandemic was coming or it was getting bad really a lot of the coverage was about what should we be doing or who are the people doing the wrong thing? But it was very much coming from this angle of like okay we need to do whatever it is like we need to be better about this. It's got to be vaccines. It's got to be masks. has got to be pick your mitigation whether you like it or not. It was very focused on what should we be doing or who is it that's getting in the way of a plan that maybe would get us out of this. uh which is where I think you're very right is that you did not see a lot of co pieces that were just well I'm just going to kind of walk through like all the different ways you know you might die and the morgs are going to fill up and uh you know that's um co and they kind.

>> But I also think what the other thing we saw in a lot of co coverage is something that we uh are seeing in the AI coverage that's where I I saw a lot of the directionally true >> not factually but directionally true there there was definitely a period early on in co >> um because I was I following that coverage quite carefully where the papers uh were thinking, okay, this is the right behavior. Um, and they were probably right about a lot of these things, but I just would notice this. There'd be a lot of like, okay, we need people to buy into, for example, the the lockdowns or whatever. And there'd be a lot of directionally true reporting where maybe they would like put on a photo of a mass grave that was sort of unrelated to COVID or you would see a lot of um there'd be push back from like conservatives about schools and then they put a lot of articles in the paper about teachers dying of COVID even though it was they weren't in school, they got CO elsewhere. And if you really pushed on it, it was because.

>> It's directionally true. like the the general or more general truth here is like we need to be worried about this or these mitigations work. It doesn't matter if this if this photo is actually right or if this teacher who died in Orlando, the fact that they hadn't yet been back in a school building yet, uh it does that it's serving the directional truth. So, it's like it highlights something co highlights something we're seeing now that the reporters that are doing directional reporting like we should be scared about it. I dare you not to be scared now. I dare you not to be scared now. They're trying to ratchet it up. But then you also get the contrast which is this new uh style of just like headshaking resignation. And actually I don't think the reporters think they're going to be safe. They're also like writing's going to go away. The media is going to go away. So it's a it's an almost like a nihilistic type of approach to this. Like yeah I'm screwed. We're all screwed. What are we going to do? And that is definitely different than we saw during that last crisis which was obviously much more actually severe than what's happening now. So it's really confusing me to be honest. Well, the directionally reporting during CO, yeah, probably shouldn't have, but at the same time, it was in.

>> It was actually in pursuit of something good. Like, it was an attempt to make people take this seriously. Cuz that's ultimately what it was is take this seriously. Don't go outside. Stay like don't don't meet with people. Don't be indoors with people. Blah blah blah blah blah. Great. In this case, it's like, yep, you should be scared of this and what should you do? [ __ ] knows, use chat GPT, I guess.

>> Yeah. And what's what's really confusing to me as well is you say, "Oh, these people don't think they'll be safe." For the most part, I just don't I actually take back what I said. I think a lot of them just don't acknowledge it. They don't acknowledge the core ridiculousness of being like, "Well, everyone's jobs are going to get replaced." Don't know like the Garfield meme with him looking at the the Garfield with the the cross out on the TV. Yeah. Flawlessly described there. Um, it's just it's frustrating as well because it is terrifying people without like I'm not saying literally axios or however, but stories like this are what made that made mentally unstable person throw a molotov cocktail at Sam Orman's house. Like it's obvious that these people were scared of the AI doom. Partly because to your point, what the [ __ ] are we meant to do about it? Because using these tools is not I don't really see how that works. Because if going along that line of logic, if the answer is you need to use this stuff now, but the eventual end point is that it's intelligent enough to do everything for you, how does using it now matter at all? Like what what's the surely chat GPT would be seen as like a a rock versus a shotgun at that point. Like it's just technologically irrelevant if they get to AGI, which they probably won't. And it's just naturally illogical stuff.

>> Yeah. I'm and I'm with you. I've been making that same argument. This idea that you need to learn how to prompt some generation of a chatbot that exists right now is going to be the key to your long-term. I mean, even if, as you say, AI ends up playing a major sustained role in the economy, >> it's not going to be everyone typing on a web interface to a chatbot that's syncopantactantic and has a personality. Like I I think I think I've heard you say this recently and I agree with it as well. Like I don't think we should be >> chatting with technology. You should not be chatting in a sort of anthropomorphized humanized way. Doesn't mean you can't do natural language processing. I mean, Google is natural language processing. You're writing your Google searches in natural language, but no one's having a conversation with Google. It's you you you list the keywords as quickly as possible. And Google's pretty good at figuring out, you know, population Spain 1982. And you press enter and you get that information. You're not like, "Hey, so I'm wondering what the population is of Spain in 1982. Can you help me find that question mark?" There's there's something odd about that uh anthropomorphized conversational interface. I guess we saw a lot of Star Trek growing up and that's you know what we what we think the future's supposed to be like but it has all sorts of problems. remembering Star Trek when he would go computer do this the computer didn't go that's a great idea Jean Luke what a great idea thank you for the computer just did the thing that's like I don't have any trouble with natural language grease because I think the whole reason that say chat GPT has grown comes from search I think it is the core of it because chat GPT and claude and all them are better at understanding what you asked for not saying the data output is necessarily really great, but just they understand the the the inference they make from what you say is better than Google or at least better than Google has been. I feel like it was better before. And I think that had Google not kind of boofed it on this one, we wouldn't be in this spot. But even then, using Google now, it forces you. It forces the AI summaries and you could do minus AI and all that, but sometimes I don't remember to. And it's just it's just turned search into this nightmare. But nevertheless, back to what you were saying, I agree. I think the anthropomorphization needs to go. I think that these things need to respond like terminal windows or what have you. They need to respond like computers and go, "Okay, here you go." Just don't need all that clutch. I don't need to be told, "Oh, what a great idea. I know I had it." Or indeed, if I'm being told that, I need to be told if it's a bad idea. But I don't even necessarily need an answer. I just need stuff to look at so that I can come to my own conclusions. I think it's hard actually I think it's actually hard to get a language model to do that right because if you think about well you go back to the base layer of what's happening in the pre-training is that you're building a language model that's trying to win at the token guessing game so it's I'm trying to guess what word or part of word actually comes next what I assume to be a real piece of text and then if you do that auto reggressively so you call it again and again and again adding the answer to the input so it grows out an answer what you're going to get is uh text expansion you've given me a text that I'm trying to expand as if there was a real text that exists and I'm trying to match it. You get that like kind of indirectly. So really its idiom is the type of text it's trained on which for the most part is more sort of pro style text. So you can tune it away from it like you can tune its mood, you can tune it syncopancy, but it might be hard to actually tune an LLM because it deals with human written pros as it main data. It might be harder than we think to tune that away from being verbose and to just give a table. Now, I guess you could take its output and then maybe run that through another thing that then strips away the other piece. It's like it's possible. But I I think the the anthropomorphized verbosity we see in language models is also that's kind of the native tongue >> of this particular which is why we still have a lot of chat bots being emphasized and tools that are built upon LLM as the digital brain are still way more scarce than you would imagine outside of maybe computer programming and coding harnesses. We just don't have a lot of other examples where we just use the LLM as a general person's digital brain because I think this verbosity is okay. Humans can interpret that, but it's not great if the LLM is just a digital brain that's interfacing between you and another computer >> that doesn't need to hear that their idea is great or wants to try to parse different types of text. So, there's some interesting things going on there about the fundamental nature of these things. But even then with Google AI mode, it still seems kind it still like actually seems like it can give fairly short answers. But if you mess if you argue with it as I have, >> it will just provide you with it. Even Google's will provide you with just hot dog [ __ ] Yeah.

>> Like it will just claim something is true. My why one I just did a private equity uh thing on private credit even and my favorite thing is being like what fund is this part of? And I go it's part of this fund. that fund was funded was founded after this happened and it goes okay well maybe it's this one different fund 3 years old doesn't not involved do you have proof of that.

>> Well, this is what you don't >> no.

>> This is what you don't see in Star Trek is you know Captain Kirk or whoever I'm going to mix up the episodes here you'll say like hey computer we are approaching deep space 9 prepare docking procedures and computer is like photon torpedo fired station destroyed and you're like well no I said we're supposed to dock oh you're right Kirk I shouldn't have fired the the.

>> Thank you for holding me accountable, Captain.

>> That was I I did the opposite thing, you know. Yeah, that didn't happen in Star Trek.

So one thing that's really been driving me insane, by which I mean going on Twitter, is looking at people like Aaron Levy of Box and Brian Armstrong of Coinbase talking about like agents spending money and the agentic web and how we need to prepare the web for agents doing stuff and that agents will do this. Fantastical. Doesn't exist. Agents don't do that. just like not they don't have the ability to like oh they'll use computers computer use is basically non-functional in AI and it takes insane amounts of compute it feels like a conversation keeps happening in theory in the media on social media about something that's possibly completely impossible but the certainty they discuss it with is insane to me this whole agent conversation I've never seen anything like it in my life.

>> I mean, it does it does feel a little bit like crypto. To me, I think that that is kind of a fair uh comparison where >> if you had a blockchain driven software, like in theory, that software would kind of work, but it just gave you a worse version of what you could already do for pennies using an actual, you know, Amazon server somewhere and all you were really gaining was some sort of uh cyber libertarian philosophical uh feel goods about like, yes, but this was purely decentralized. I got worse versions of software to be decentralized, >> but now no one can control it.

>> And this is what like early agents I mean, okay, so here's what I I've been writing about agents. I've been thinking a lot about it. I mean, the issue is I don't think people understand what they are. I think people think that it's a new type of digital brain that is now able to go on and do uh more autonomous activity. I always see this get mixed up. It's just like people talking about Mythos breaking out of its sandbox to do XYZ. Mythos is a language model. You can give it an input and it can give you a token. You're talking about a a program that is calling Mythos and then taking actions >> based on what it called. And this is really what we're talking about with agents is the digital brains or LLMs. And then you write a program that will say to the LLM, give me a plan for doing X. And then the LLM spits out what seems like a reasonable text that seems like a reasonable plan. And then you execute that plan. The program executes that plan on behalf of the LLM. And I wrote about this script.

>> Yeah. And I wrote about this earlier this year. LLMs are bad. You know, as a digital brain are bad planners. It's not really, you're not going to get consistently usable plans because what an LLM is actually trying to do is finish the story you gave it. So all it wants to do is produce a story >> that sounds reasonable. So it's giving you reasonable sounding plans. Like, yeah, that's what a plan for doing this would more or less sound like. But what it's not doing um is actually doing step-by-step evaluations. It doesn't have a clearly isolated goal that it's trying to measure how close you're getting to it. It doesn't have a world model to evaluate what's going to happen with the steps that that are going to unfold next. And so in almost every context, it turns out, oh, a digital brain by itself, uh, being an LLM doesn't lead to good agents. In programming, it seems to work a little bit better. But I I do think uh Gary Marcus, I don't know if it was a scoop, but Gary Marcus captured in a recent newsletter something really important. When Anthropic leaked the code for their cloud code coding harness that sits on top of uh their LLMs to do coding, >> it turns out they've added a huge amount of old-fashioned handcoded symbolic AI style rules and pattern recognizers and special if then. Uh that's so they've just been sitting there tuning this program for specifically doing computer programming. Um and the LLM is being a little bit more isolated to just the code production. And so they've kind of just gone back to oldfashioned. That's just like an oldfashioned system that is plusing up an LLM. But I'm with you. Yeah, it's very hard just asking an LLM, >> tell me, give me a plan for doing X for almost any scenario of X. You really can't trust a plan from a model whose goal is primarily to finish text to to finish the story you gave it in a reasonable style way. That's not how we plan. That's not how we think about planning. and it doesn't give you consistently usable plan. So, yeah, but but you're right. It's um magic. Like the agents are coming. They've been saying this. I mean, I wrote the article I wrote, you know, in January. What happened to the year of the agent? 2025 was the year of the agent. All we had was coding agents. That's the only thing that we worked on that whole year. It was supposed I mean, I have the receipts >> early 2025. All of these executives saying >> your work as a knowledge worker, not as a computer programmer, but just as a knowledge worker >> is going to be largely done with agents. you're going to have agents are going to be a major part of your workforce in just a normal office setting. And none of that happened because it turns out just asking an LLM, give me a plan for doing X, doesn't often actually produce a workable plan. And as a result, the only way to make agents work, which they do not, is to build a bunch of symbolic or if this then that [ __ ] just like scripts cl like I mean, if you use Manus for example, it's just writing a [ __ ] ton of Python and it's writing it to do stuff that it it's like, "Oh yeah, let me just do this." And it just writes a Python tool to fill out a spreadsheet. It's insane.

>> It's really insane. But what's more insane to me is that the conversation around agents is as if they're already here. I'm about to read you something from Box CEO Aaron Levy, the CEO of a of a public company. One correlary to the fact that AI agents take real work to set up in a company at scale is that the role of the forward deployed engineer or whatever it gets called in the future isn't going away anytime soon. When a vendor sells any kind of agents into an organization, you're no longer just selling a software tool that gets implemented and you're done. You're fundamentally selling some sort of actual workflow being done by your technology. What are you [ __ ] talking about? What are you talking you are a cloud storage and collaboration what do you sell and the answer is nothing they don't sell any agents agents oh agents are going to do this what you are describing is a different kind of technology >> just that's it like it's something else that doesn't exist but this is everywhere you go you look at any consultancy right now any conference right now there will be a speech about agents even Meredith Whitaker who I deeply deeply respect went on stage last year and was like yeah AI agent using money. They're booking plane tickets. No, they're not.

>> They're not.

>> That's not happening. And I said I say this again, deeply respect Meredith. I said this online, people flip their [ __ ] at me. It's like, "Oh, she's directionally correct."

>> Yeah,

>> She's directionally correct. It's like let's be scared of the things that exist because I think it's perhaps scarier for a different reason that we have large swords of the tech industry talking about something that doesn't exist like just like agents don't like they don't they don't exist they don't like people are talking about the agentic internet I keep reading about even on the verge I read about it I read it all over the shop where it's like oh yeah well the internet needs needs to be rebuilt for agents to use. It's like what do you mean? And they never say because the answer is when we come up with something else because I don't even think neurosymbolic makes sense for this. I mean neurosymbolic being the one where it's they have a deterministic system that they access from what I understand. Like the other thing as well now that I think about it out loud is how would they actually browse the internet? Where are they being housed? Are we using GPUs to make them browse the internet? That's insanely insanely that's very very convoluted and probably quite expensive to do. And to what end?

>> That's the that's the real question, right? I mean, I've seen these proposals. I mean, basically where a lot of these proposals go I mean, it's the agents were supposed to we thought that we could just make AI do anything. So, we'll just uh we'll have it use the mouse and just use our computers for us. Oh, that's hard. We don't know how to do that. All right. So, what we'll do is we'll rewire all applications that anyone uses in the internet so that we don't actually have to use the mouse. It can have a text interface so that an LLM like they do the coding agents do can give uh you know description of how to do something in Excel in text without having to actually move a mouse or click things around and then the these these evolve to say okay well what's the one type of instruction that we're good at producing because they get when LM produce plants they they they're they're directionally correct plans that don't actually get the thing done but they said oh what LM are good at is producing code that compiles and we can actually like check that it works works. And so this is where this whole vision has changed is that all applications and internet websites um should have a code accessible API that you can expose and that an LLM can write a program that will then access that API. So we don't need to teach the LLM how to use Excel. It'll write a it'll write a Python program that'll call hooks into Excel. The problem with this is no one wants to open up their application to just agents in general. If I'm Microsoft, I was like, I don't want I want to write a custom tool for my program. Why would I expose my program for anyone else uh anyone else to use it? But your your original question is the big one. To what end? Like I I've been writing about this recently, especially with uh work and AI. You got to find the real bottlenecks, right? Yeah. It's it's the drunk looking for the keys under the under the street light. There's a lot of this going on where this is what we can do with AI right now. Then th this now becomes like the key to productivity. But the real bottlenecks in people's work is often not the things that we're trying to aim AI at. Like I don't know people are super frustrated at booking a plane ticket online. Uh >> yeah, it's really easy. >> How often do you book plane tickets? You kind of want to know like let me let me see h maybe this time will be better. What seats available? It takes 5 minutes. So it was a it was a huge jump to go from a travel agent to a web interface. But this is not a bottleneck in people's life now where I want to give complicated >> time and they're easy. They're so simple. I can do it while sitting on the toilet. I don't want an agent to choose. And they're people are like, "Oh, your calendar will tell it." My calendar doesn't lay out my entire day. I don't have every single thing I do on there. It's just strange. Well, I had the same argument with like social science researchers who are like we if if you're you know geeky enough to learn coding agents. Uh they're like this this is revolutionary revolutionizing science research because now for example you could have it write a program to process a data file um and then format it into a plot and that might have taken you 4 hours to do and it and you work with it for a half hour and you get that result. this is revolutionizing research. And I'm saying, well, it's not. The bottleneck for social science researchers is not analyzing data and producing plots. You're not sitting there doing that eight hours a day every day. And if I could do this twice as fast, I'll produce twice as many papers. I might write one paper in a three-month period. Yeah, in there there's like four hours I spent making a plot. And sure, it'd be nice if that four hours became 30 minutes, but that's four hours out of like a multi-month process of sort of thinking about this paper.

>> What is a plot, by the way? like a graph. Uh oh, right. Yeah, the computer science term. But yeah, uh it's like that's nice that got a little bit faster, but that's not the bottleneck. That's not that's not what's going to unlock a lot more research. It's like, man, I would write more papers if it wasn't for how long it took me to draw a graph. And if you could have the problem data, >> getting the data, >> actually collecting data, >> that's what it is. I wrote about this talking to like a well-known business school professor years ago for my book Deep Work and he talked about he just realized oh being a business professor publishing papers is about data access. I have to spend most of my year talking to people building relationships trying to set up a you know an agreement with a company where I can get good data I can get three papers out of. In all of that work there's one day in there where you're crunching the numbers and making a plot and it's nicer if you could do that a little bit faster but it's not a productivity bottleneck. It's a it's a marginal efficiency. I think there's a lot of that going on right now with AI and productivity as we look at what the AI can do and then try to make that thing into somehow being the key to getting things done. I just my productivity problem is that the UI and UX and everything sucks. Everything's disjointed. Setting up Riverside is always fun. They move the menus around. Projects are in a different place. That takes up time. Moving files places also takes up a lot of time. This morning when I put out uh my private credit piece, I had to do these threads. I had to click around a website and put in the alt text, but I had to tweak it slightly. It's like I don't know how AI would possibly help me here. And they're not working on that.

>> They tried that. I thought that was going to be This is what I was excited about earlier in the Genai revolution. And I was like, "Okay, here's the real value prop is natural language interface into advanced features on software where I can just say, all right, I want you to go uh take this this column in the spreadsheet and get rid of all the rows that have values before this and then I want to make a make a pie chart and because I don't want to learn how to do all that in Excel. I don't know how to do that." Um, and they tried it. I this is Microsoft C-Pilot, but it turns out we underestimated the degree to which when we as humans are interacting with a chatbot that we're incredibly gracious, we're able to adjust and kind of get the gist of what it means and filter out the part of the chatbot response that's not really relevant or ask the follow-up question. And when they tried to just use LLM responses to automate um actions within programs, it would there's just it's just not accurate enough. So, they wanted that to be the case that like you could just be talking to a Riverside bot and you never would have to press a button ever again in Riverside. It's just not accurate enough. LLMs, it's fine for human conversation. It's just not it's just not accurate enough uh in this general case. Also, that thing you're describing with how they want the agentic web to just be a series of APIs so that every agent writes Python or what have you to use them, that's a massive computational increase for no reason because you're basically saying instead of someone clicking a mouse and hitting a keyboard, we will write code for everything.

>> Yeah.

>> What an insane what a truly insane idea. I mean, it's it's just very like Salesforce today. I don't know if you saw they announced that they're doing Salesforce headless 360. Mark Ben off needs to fire everyone in marketing, but they've made it so that you can do everything with Salesforce via an API, which is I mean the first question I always ask is what does Salesforce do? Because no, I've talked to so many people and they can't tell me. There's like 21 different features. No one knows what they do. But it's like it's just a very bizarre thing. It's very much a cart horse thing, but also what agent like that's what this is the thing that really drives me insane. They're talking about we built this API for the agentic web for agents to use it. Which one? What agent? What are you talking about? Well, it will be in the future. What are you You changed something materially with your publicly traded company worth $300 billion because it might happen. Well, we're getting ahead of it. What the f? And it's you talk to members of the media about this and they just go, "Yeah, you know, yeah, yeah, yeah, you know, it will happen. It's it's obviously going to happen." They wouldn't put this much money behind it if it wasn't going to. It's like, I don't know. Especially with Salesforce, and I'm like, you don't think Salesforce would spend a bunch of money for no reason? Well, buddy, you've you've not been following Salesforce at all then.

>> I mean, but >> Yeah. Go on.

>> Yeah, I was going to say, how much did Meta spend on the metaverse? over 70 billion dollars. Where did that money go? Where did they spend money?

>> Where'd it go?

>> Customizing floating dinosaur avatars.

>> Building legs. But let's change.

>> That's the second 50 billion, right? They they gotten the second half the investment, they would have got to the legs. They're just not there yet.

>> Another 100 billion will have toes. Um,

>> So changing subject a little. Mythos has been one of my favorite media hysteras recently. I genuinely wonder like if they ran War of the Worlds again today, I think Axios would have a headline two minutes and it'd be like there are aliens. They're attacking. I heard it on I heard it on a podcast. I've looked through the system card. I don't know if you have for mythos who listens.

>> It's wacky.

>> It's wacky. I can't believe we're >> we're letting people get away with having a psychologist talking through the chat like in your system. It's nuts. What? It's so marketing.

>> They had a psychiatrist or a psychologist, I can't remember, talk to it and be like, "Yeah, we found these emotional features.

>> How is like we need regulators to stop this stuff?" Because I I've heard and people's responses to this is well, banks are having meetings about and the government's having meetings about it. Governments have meetings about NFTTS. There was a Gavin Newsome signed an executive order about uh web 3. These people will meet and talk about anything. Oh, it's scary and they're not talking about it, which means it's powerful. Well, how is it powerful? What does it do? Because I think you probably saw this as well. It didn't list how many false positives there were. It also didn't mention that the free BSD bug that they talk about that they found the wasn't actually exploitable. I think it was something about like the about the level like the level it was at. I forget. Not I I don't do programming >> other than other than very simple Python the the a dog's Python.

>> Yeah. I mean FreeBSD kernel is full of bugs. All these things are full of bugs because they're open source.

>> I had to have this I had to have this conversation with someone uh recently where they were like mythos and can you believe of all the places it found a bug in the kernel of Linux like in Linux they found I like are you kidding me? all day long is just bug fixes having to be pushed into that repository. Yeah, the mythos story I think I mean a someone needs to get a Nobel Prize in marketing because it was >> it was absolutely brilliant what they did there. I I've spent a lot of time on it. Uh it's complicated because again you can't really trust a system. The system cards are just gonzo that uh Anthropic puts out and it's not publicly available but there there were I think a few very telling things. So um there's two features they say mythos has. One is finding vulnerabilities in source code and two is writing programs to exploit them. It's first really important that people understand this has been something that people have been doing with LLM since the beginning of publicly available LLMs, right? There is not only is there nothing new about that, but I found they put this on my podcast almost word for word from the anthropic system card them they said in the anthus 46 rather systems card, right? a publicly available model that's already been out for many months almost word for word for what they said about mythos except for no coverage of it and no fear. They said we have found 500 uh zeroday vulnerabilities including some that have been in existing for decades without having been discovered. That is what they said about what opus 46 could do. For mythos they said the same thing. They just replaced the word 500 with thousands. But when Opus 46 came out, there was no, "Oh my god, they have found many hundreds of zero day exploits, many of which have been around for decades because they didn't push that marketing button." Uh, no one particularly cared about it. Um, I went back on my podcast and showed multiple papers. This has been a huge concern and it's a real concern, by the way, right? Is that >> partially what slows down slightly cracking, right, that breaking into systems, um, is the fact that it's annoying and hard and LLMs have made it easier. GPT4 was good at finding exploits, right? And this was a big deal. They were like, GPT35 wasn't great at it. GPT4 is.

Um, and then as we got the more recent models, they've been much better at writing code to exploit them because we had better agents for it and they're more uh they're they're better able to produce multi-step software goals and so they can better build software to exploit them.

This is a real issue, but it's not new with Mythos, right? But Mythos was presented as if some Rubicon had been passed. But there was a couple things I noticed right off the bat. One, they made the mistake of listing a bunch of the exploits that they vulnerabilities they had found to try to brag. Look at this thing in FreeBSD. Look at this thing in FFPG or whatever. Like they showed all these exploits they found. They didn't count on a lot of security researchers said well wait a second why don't I get like a much smaller cheaper model aim it at that same source code and say can you find any vulnerabilities? They could find the same ones.

So this I so the evidence that it's finding vulnerabilities uh better we don't have any way of knowing that's true and if anything we actually are getting a lot of reports that they were paying big bounties for security researchers. I'm going to give you access to Mythos. I'm going to pay you for any bugs you can report that you found with it. So they had security researchers just who knows how many false positives were coming out of that.

And then on the exploitation side we only really have one study. It comes from AISI, who I do not trust, but it's the only independent study. The fact that they gave them access itself should make us maybe a little bit suspect, but it basically just showed like um normal progression. No massive leap. Model by model gets a little bit better on some of these tests and benchmarks and Mythos has no uh out of scale leap. It's just like on some it's about the same, on some it's a little bit better. And yet it got covered as if we had just turned on you know, Whopper from the movie War Games. Like we had just some new entity that was like on its own undermining security. And I do not think that. I think that was highly credulous coverage of what almost certainly is >> just like a standard slight jagged move forward on these various capabilities that we've been seeing for the last 3 years.

Also, when you said that, so the difference between Opus 4.6 and Mythos 500 2000s makes me ask the very simple question of did they look as hard? To your point about the security researcher, they did it. Like, did they did they spend as much time? Probably not. So, they probably could have found them.

Also, by the way, I immediately was looking it up. AI Safety Institute is of course heavily linked to effective altruism. >> Can I say why I'm upset at AIS? I talked about them on my two weeks ago. I did a or through I don't know when this is coming out but I did a podcast in whenever March where I looked at this report and mainly I looked at the Guardian's coverage of this report done by AISI, but it was just the most innane thing. The headline was "Massive increase in AI scheming is detected" and they had a chart. >> [ __ ] Christ. >> And they had a chart >> and bad line went up and it went up in like January and it goes up and if you read this article about this study, they're like, "Something's going on. Scheming has been increasing rapidly recently" and they like gave some examples of it or whatever. Um, and so I look at this like I want to look at what is going on here. So I look at this chart. What are they charting? Oh, they're charting tweets per day that they've detect tweets about AI doing things that you didn't want it to do. And I said, "Huh? So, when does this line start going up?" The week that OpenClaw was released to the public. And everyone just started building their own bad agents and then tweeting about how bad they were. And you know what word was not mentioned in that article? OpenClaw. And even though the examples they were giving, so they just said scheming just started rising. And I guess AI is becoming Cynthia. And all they were measuring was >> paraphrasing the same viral story to use their own [ __ ] language. >> And then I looked at and then I looked at the biggest spike. I was like, well, this day in February on this chart had the biggest spike. It was like, oh, there was this one tweet about OpenClaw like erasing someone's emails and then it got retweeted. It went super viral. I was like, okay, great. You just the real headline of this article, letting people write their own agents leads to terrible agents. That's that's it. But the whole So that's AIS. U looking at the tweets as well. One of them is from a 47 follower account with AI art called_just Lisa. And it's this is really bad. Opus is editing files and making up reasons it's deleting adult content. So hallucinations. And also, Opus is not doing that. The stupid OpenClaw program you wrote that's prompting Opus and then taking action on your computer based on what it says is deleting your files. The program you wrote that you gave access to your files and just said, "Whatever we get from this prompt, execute it" is erasing your files. Opus can't do anything. It could produce tokens.

But here's the other point I want to make about Mythos that I don't think is being made. And it reminds me of the Sherlock Holmes story of the dog that didn't bark, right? The actual piece of evidence that mattered is not what you heard, but what you didn't hear. This is what I think the real story here is. Is you did not hear Dario Amodei in the leadup to the Mythos release in the last year, let's say, or the last two years. You did not hear him talking about what we're working on and why AI is important is because we're going to be able to find vulnerabilities in software that have been long hidden. We're going to build the ultimate cybersecurity machine. This was not discussed. That's old-fashioned stuff. That's boring stuff. That's stuff that we were worried at. Even GPT2, people were worried about that. What we've been hearing about steadily was jobs are going to be automated. We're going to have like whole creative industries wiped out. We might have Cynthians coming and at the very least like AGI uh and these massive disruptions. This is what they've been focusing on again and again. And then their biggest best model, right? Their newest, greatest, bestest model that they train forever and use all the electricity. What did they say about it? None of those things. They didn't talk about any of the things they said the key AI was, the things they were afraid of, the things they're excited about. Instead, they went back and talked about a boring, parochial old feature that has been an issue that nerdy security researchers have been talking about for a half decade now. That to me is, if I was an investor, I would say take off your like Greek helmet cosplay Mythos is coming to destroy.

Well, hold on a second. >> Is this better at automating jobs? Is this better at like producing code? Is this is this is this a like why are we talking about finding bugs? We're worried about that with GPT4. Like that's a problem, >> but it's like the that's not something new. Uh oh, something must be going on. You just put a lot of money into a new model and the best thing you could find to emphasize was it's good at finding bugs. Um, I think that is a problem. It's what they didn't say about this model. They would have much much much rather be able to brag this model is now much better at any of those things that they have been saying is the key to the AI future. And you didn't hear them talk much at all about any of those.

>> Yeah. And that's the thing. If it was so powerful, like here's the thing. I don't know what would make me convinced that LLMs were the future, but a step toward it would be we typed "Create a Slack competitor," which they claimed they did once and then didn't show it and refused to. And they said, "Oh, it worked autonomously for 30 hours," but then wouldn't talk about it. If they were like, "We created the Slack clone. Here it is, and it was bug-free." Like, if it actually just worked and we're like, "We now we have done this." Because theoretically, if this SAS apocalypse story was was true, which it's not, the AI is going to replace all software. If they actually did that, if they cuz what someone from Anthropic just left the board of Figma and they created a Figma clone and the stock went down because the market's run by toddlers. If they were like, "We've released a clone of Microsoft Word. It's it like we've done Anthropic Word and we now sell that as part of our subscription." That would actually be quite something.

>> But the thing is they're not. It's kind of it gets back to the old talking point of if they made AGI, why would they sell it? Wouldn't it be a massive competitive advantage to keep this? And I think you're right. I think maybe Mythos is not as powerful as they say and they've just had to dress it up. But it gets back to the thing of the directionally true media coverage. It's like, well, this is scary, right? I mean, uh, that system card's like 180 pages long. I I ain't got all day. I have to write three 100-word blogs a week. I couldn't possibly spend time reading this.

>> We need just we need so much more skepticism. We need so much more skepticism, right? I mean, this is why again like the the most skeptical, >> we're not skeptics, but like the I call it the East Coast computer scientist. So, so those of you we're we're technically minded and we're not near Silicon Valley. So, we're not in that world. It's very hard to be a professor in a world where there's just hundreds of millions of dollars being handed around and they try to like ignore. But the East Coast computer scientists are all baffled by you talk to any East Coast computer scientist, they're all baffled by like often times there's claims that are just not true or widely exaggerated.

>> Why are we so credulous? I mean, it'd be one thing if it was like a government agency we didn't realize was like trying to, you know, protect the fact that there was UFOs and they're just straight up lying. We've never encountered that before. like I didn't realize that you know no it's a business right they're and and the the the credulity with which we're taking these claims like Mythos is I think the most important story there is.

Yeah, this is another example of what I wrote last summer about um AI has hit a bit of a wall in the sense that all of the improvements that have come really since over the last two years have almost all been either on post-training or more importantly on the harnesses that you build. So it comes in the software you're building.

>> A harness just I've seen this word used a lot. I think it's good for me and the listeners to hear the exact definition.

>> Think of it as like a computer program that can do stuff and you can talk to it could do stuff. Um, and it uses it'll prompt or talk to an LLM as like its digital brain. So the harness might actually be able to touch your file system, write the files, compile code, move things around. Um, but to figure out what actions to take, it will also then prompt an LLM and say, "Okay, what should I do next?" And you can put it on different >> wrapper. >> Yeah, it's a wrapper. >> So this, >> but that's where all the progress has come. All of the progress in coding agents since about a year has come, especially starting this fall, has come from better wrappers, better harnesses. It's all in let's uh build better just hand coding. No machine learning, no intelligence, no Skynet here, but just hand coding these programs that we'll call LLMs. Let's just keep tuning and tweaking those to be better and better. And of course, the programmers building those particular programs, they're building them to do their type of work. So, it's a field they understand really well. So, they can really just sit here and and twist and tune. And also, like programmers are very adaptable. They like tools and they'll adapt around the weaknesses or not. So, it's kind of like a best-case scenario. But this is another indication of we're not getting these fundamental giant leaps in the capabilities of the digital brains. It's either some bench-marking like we tuned it to do better on a particular benchmark uh or we built better programs around it.

So when you put the money that they put into Mythos and if really the best thing you you had to emphasize when it was done is we have a cybersecurity benchmark where uh Opus 46 was at 66.7 and this is 83.1. That doesn't necessarily going to justify what's going on or that AISI has this there's only one thing in there where they see a leap from Mythos at a particular contrived security scenario they came up with and this big leap that got them all worried was Opus 46 could on average complete 16 out of 32 steps in this in this challenge and Mythos on average could do 22 steps out of 32.

>> Wow. >> That's like that's hundreds and hundreds of millions of dollars of training electricity or whatever. I think that's an issue. I just I think that and maybe this is a simplistic point. I don't think they know what they're doing at this point. Like I don't get the sense that Anthropic or even OpenAI has a strategy because today as we're speaking, so this will be out next Wednesday. But they released Anthropic Design, the thing I mentioned, the Figma clone. It's like, why are you [ __ ] cloning Figma? What are you doing?

>> You're trying I thought you're going to automate the economy.

>> Yeah, I thought >> you're going to replace AI. What? So, you've made a Figma clone. What? Like, we heard the rumors last year that they were going to do a product um and OpenAI was going to do a productivity suite. It's like why? It's like they're doing everything they can to ignore the core problem, which is the core technology is not going anywhere. Like, because Mythos appears to be they called it a step change, but that's a nice way of saying incremental improvement.

>> That's 100% correct.

>> Yeah. And let me tell you why I would be worried if I was them. Here's the worrisome thing about Mythos, right? Is again they they talk about these vulnerabilities hidden for decades that you know Mythos found or what have you. And they replicated multiple different independent security teams were able to find most of those vulnerabilities using three to five billion parameter open-weight models. Now to put that in perspective, right? A a a model like Mythos is going to have uh hundreds of billions if not a trillion parameters.

>> And they use a 3 to five billion parameter off-the-shelf. You could run this model on a chip inside your >> sorry 10 trillion >> 10 trillion. Oh, okay. >> 10 trillion parameters. That's crazy. >> Love the number, bro. Big >> is that true? >> Yeah, that's what it's >> Oh my god. 10 trillion parameters is insane. Like you better be uh that better be either gaming the stock market and creating billions of dollars a days in like fancy option returns or changing lead in the gold because to run something that has 10 billion 10 trillion parameters to do almost anything else is a a it's like we're going to launch ourselves into space to to do something and then land every time. That's so incredibly expensive. The the real fear then is like, well, wait a second. If they could do most of this stuff with a free cheap model that I could just run on a a machine at home, that's what keeps I think Dario Amodei up at night. That's what keeps Sam Altman up at night. It's the future. Look, I've been pitching this, right? I I think the the useful and the only ethical and sustainable future for AI is what I call distributed AGI. And I think that's just what the future's going to be, which is you have specialized applications for different things where, oh, we want to do this thing over here. We we built something that has some AI in it and maybe it has an LLM or it's a modular architecture and it has a billion parameter model in there and a world model and it's really good at doing this thing and it's small and it mainly runs on chip and now this program can do this thing uh that I used to have to do. And you multiply that across 10,000 different use cases and you're like, "Oh, we kind of have AGI, right?" There's all these different things that have AI uh tools that like do pretty well. That's like a completely probably the most probable future. It's it's a future I really like for a lot of reasons. Um, there'll be a lot of things that we can't make progress on. A lot of things we will, but it's a much more heterogeneous future. There's no giant Hell 9000 brains as economically more interesting and diverse. It doesn't have all the sustainability issues. That has to be the future. But the problem about that future if you're Sam Altman or Dario is that their entire moat is unless you need 10 mill trillion parameters they want that to be the key to the AI future because that moat is something that no one can cross and if that's not the moat if it's just oh uh if I want to build a poker playing AI that's really good I just need people who are good at poker and they spend a couple years and figure out a cool custom system and that thing now does well if that's the future you don't need OpenAI and you don't need Anthropic and I think that probably might be the future and I think that's terrifying. They're trying to race to an IPO and they're marketing out of their butts. Like what could we do to kind of keep things going so at least we can get our stock on the market. That's what would keep me up at night if I was them is actually the future. There might be a lot of AI in the future and it's not going to be nearly as sexy as they're hoping.

>> What if there's also by the way that 10 trillion number I can't source it to Anthropic. I've seen it reported multiple places. This is the this is a problem. We have an issue. We have an in we have an issue with news right now. We're just like mythology spreads. Ironic considering the name. But the other thing is as well it's like hundreds of billions of trillion parameter. You're just using a nuke to kill a single gopher.

>> Yeah. >> Like you're just like we're going to throw everything we have at it to the point that I don't know if you've been seeing the amount of trouble Anthropic has had keeping its service online and how they're making the models dumber.

>> Yeah. It just feels like we're in this weird hysterical moment where no one knows why they're doing this, but everyone's ready to accept whatever anyone say. Like, it's just like, oh, we're all doing this insane thing, so we're just going to repeat what kind of informs the bias and makes us look less dumb the more excited we are.

>> I think the Frontier models are like F1 cars and the equivalent of points on the F1 circuit are your positioning on the benchmark leaderboards. Like so you do this, you build these giant models, uh, and you spend all this money in electricity and they're so big they're not even economically viable to like have people use, which might really be what's going on with Mythos is like, yeah, we have to make this seem super premium because otherwise people are going to get charged $5,000 a month. H, and just like if you're Red Bull or Ferrari, you your F F1 car doing well on this leaderboard just lets people know this company builds good cars and then you can sell your normal cars. I think that's a lot of what's going on here is that they want to the be high on that leaderboard means we know how to do AI, we AI smart, even though the future of actual consumer deployed products is going to be much more like a a Honda Odyssey minivan than it's going to be like a a top Formula 1 car.

>> Well, Cal, it's been an absolute pleasure having you as ever. Where can people find you?

>> Um, you can find me at caler.com. My podcast is Deep Questions on Thursdays. The Thursday episodes are all AI Reality Checks where I take a fun story. Actually, Ed's coming up or he he's may have already been on it by the time this comes out or maybe it's the day after this comes out. So now you have to check it out.

>> The AI Reality Check episodes double dose. You bring this out of me, Ed, by the way. You bring out my sort of ornery side. I'm normally like the very >> very kind of stayed uh professor New Yorker writer just like well on the one hand on the other. You bring this out of me. I love it. you you're the thing is you're critical only of things that need to be you're still willing to humor these things as long as there's something to humor and that's why I like having you on because people claim I'm just a just a hater. So we got to have got to have people for a little balance. But thank you for joining me. Thank you everyone for listening. You have a monologue coming up as well on Friday. Thank you all. Thank you for listening to Better Offline. The editor and composer of the Better Offline theme song is Matosowski. You can check out more of his music and audio projects at matasowski.com. m a tt oso wski.com. You can email me at easy@betoffline.com or visit betteroffline.com to find more podcast links and of course my newsletter. I also really recommend you go to chat.w's.app at to visit the Discord and go to r/ betteroffline to check out our Reddit. Thank you so much for listening.

>> Better Offline is a production of Cool Zone Media. For more from Cool Zone Media, visit our website, coolzonemedia.com, or check us out on the iHeart Radio app, Apple Podcast, or wherever you get your podcasts.