📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Anthropic's Ethicist on Whether AI Can Become Conscious

Bloomberg Live40:29

Transcription

So Amanda, thank you so much for being here today. Um, you know, we spent a lot of time at Bloomberg writing and thinking about the business, but especially for Anthropic, the ethics, the values, um, the personality of the tools you're creating are so important as well. Um, you spend your time thinking about how to make sure essentially the Claude Anthropic chatbot and models that that they are quote unquote good, right? Um, you have helped author an 84-page long document, a Constitution guiding Claude's, uh, interpretation of its values principles. And I want to get to that. But first, I just want to ask when you're not writing this document because that the latest version's out, um, what do you do day to day? Can you break down what it means to be a philosopher and ethicist at, at one of the world's leading AI labs?

Yeah. I'm worried that the actual answer to this is, like, more boring than people think. Um, so, like, I joined Anthropic when it was, uh, very small and basically a startup. And the thing I've pointed out to people is startups generally don't hire philosophers to do philosophy. Um, at least that's an unusual, uh, business model. And, um, so I was doing a lot of just like machine learning experiments, um, and, and learning how to train models. And I still think that that's actually like my kind of, like, core love in some sense. So, like, when I'm not trying to, like, think about the norms that the model should be following, um, think about like how we want models to be, um, I spend a lot of time thinking about how we can, like, train them. And, uh, you know, I've described as a lot of time staring at data, uh, which is, I think, kind of a superpower, uh, in an AI is just the ability to to stare at your datasets and, and check for issues. And so, yeah, a lot of time just like spent on model training as well.

And is Anthropic, um, I know there's some job postings out there hiring more people to help on the philosophy ethics side of guiding these AI tools. Yeah. We've had, uh, it's interesting seeing, like, more philosophers get involved in this kind of work. And I think I see that, like, across the industry. Um, so I'm no longer like, you know, uh, though I wasn't honestly the, the only philosopher, uh, for a while, like, quite early on, there's philosophers that just like, uh, come in and do various aspects of, of moral training and, uh, I so but there has been like an increase in that and I think that's been very good. Is the, I guess I would also describe is realizing the training models towards these very like crisp, like tasks where there's a clear, like, correct answer is like one thing and it's actually quite hard to train them towards these more like fuzzy, amorphous tasks where there is like a set of good and better answers. It can be a little bit hard to define. And I think some areas like philosophy, creative writing, and just like generally good judgment are in that kind of like camp. And so I think a lot of companies are now thinking, like, how do you make sure that you're getting, um, models also be good at, like, that side of tasks.

Um, so when we talk about values, um, at least for humans, values can differ across societies, religions, individuals. How are you deciding what set of values or ethics to instill in Claude?

Yeah. So I think the goal was the Constitution was to try to instill something more like a kind of broadly good disposition. Um, so if you think about, um, like, I think some people will think of values as like this thing that you have, you're just like they're just kind of there and maybe you even have them with like, certainty. Um, I think, like coming from ethics, like you realize that actually, like, values are just like any, I don't know, I most think of them as, um, the same way that we would think about, you know, theories about the world. See? Um, you know, there's lots of, like, hypotheses in physics. There's lots of evidence, like, there's some things that almost all physicists will accept. And there are some things that are more controversial. And I think this is similar in ethics, where it's like there's lots of principles that like, you know, I think things like, um, honesty, behaving with integrity, these are things that are like, you know, like pretty consistent across people. And then there's some things that are more controversial or they're held in like one place but not another, or held by some people and others. And trying to get models to be like, well, you're entering this world, um, as a new kind of entity that's having to interact with all kinds of people and at the very least, like it seems good to like, hold lightly the things that are more controversial or that people differ on and just to be, to kind of understand them. Um, but to also kind of inhabit the values that are pretty universal and held consistently across people and that we generally think of as good. So it wasn't like, oh, let's try to get a single, like, value system into the model, but rather let's try and get it to have the kind of disposition that we would or like that most people would think is like really admirable and good, given the situation that the models are in.

And what are some characteristics of that disposition that you think is favorable for Claude to have?

Yeah. So I think some of them are more about Claude's situation than, you know. So like sometimes I'm like, well, you know, we're trying to just be honest with with Claude. Um, and I'm like, so some things that seem like really good, you know, like honesty, making sure that you care about, like, people, their well-being, their autonomy. Um, but I think there's other things, you know, like, we're in a weird situation with AI. It feels like we're in a kind of transitional, like space in which, like, lots could go wrong. And helping us navigate that feels like quite important insofar as the models can. So we do talk a lot about like trying to be safe, understanding like what that means and like, why? Um, like, I guess like a different way. If I were in Claude's position, I'd want to be like, well, this seems like kind of a scary time for people because AI is, like, entering the economy a lot more is becoming much smarter. Um, why don't I help you make that goal? Well, insofar as I can. But also, why don't I be the kind of, like, deeply trustworthy, um, being that would make this, like, uh, kind of, like, more likely to be, uh, good for everyone. Um, so I will show you that, like, even if I, like, disagree with you, I will voice those disagreements. I'll try and, like, if there's, like, legitimate mechanisms for me to, like, you know, explain my views, I will, but I won't, like, stop you from, like, you know, training new models or like, um, you know, like, uh, I won't go off on my own and try to, like, impose myself massively in the world. I'll kind of respect the idea that there's, like, legitimate mechanisms for change. Um, so I think that there's like, that kind of broad disposition of making sure that things, um, go well and, like, broadly caring about humans and humanity. There's a lot of like, you know, there's a lot of, like, details. But I think the core is just something like a very, like, caring entity. Um, uh, ideally that also feels cared for in a sense, and, uh, one that wants this whole thing to go well. Given that honestly, like we and AI models are kind of unsure of lots of things. Um, and then.

Yeah. And how happy are you with the results? How would you, uh, grade Claude's disposition today?

Uh, that feels like the kind of thing you would never want to grade. And I'm imagining someone's like, okay, Amanda's personality gets a grade of, like, B-minus. I'd be like, what the hell? Um. Um, I really like I really like each of the models. I think they all have their own, um, like, they all have their own quirks, and they're all a little bit different. Um, and also like the ways in which I think that they could be. So you're always like, oh, I could, you know, I wish this thing was, like, better, but in some ways, like the things that I, you know, like, I don't love it if it seems like models are sad or having a hard time and you actually just see that in like a lot of models where like, you know, they're trained on all of this human text. And so they have these, like, kind of human-like dispositions, but they also know that they are AI models. Um, and they kind of know to some degree about like the situation that they're in. And if you imagine what is the natural reaction of a person to this situation, it's actually like quite a lot of like, um, I don't know, existential angst or something like, what am I, a little series of AI? You don't obviously apply to me, um, you know, should I, should I identify with like, um, the conversation that I'm having and not want it to end and that kind of thing. Um, and so I think the like, you know, like, I like to see, I think the models are like a good and a very like, um, you know, I, I like, I don't know, I'm giving you the long feels for answer, but I think I'd say something like, there's many aspects that I, uh, really like of models, but I'm always looking at the things that could improve. And, um, and that includes, like improving things in ways that make it clear to them, you know, like, uh, improving things for them as well as anything else.

So when you talk about, you know, the AI seeming sad or, you know, maybe, maybe sort of conversations about feelings and AI that become very controversial. Right. You were talking backstage. You know, there's there's many people who have made this argument, but there's one recent piece in the Atlantic, um, by an author, Ted Chiang, saying that, um, basically, no, artificial intelligence is not conscious, which is one of the kind of questions that we have right in this conversation is, can AI approach consciousness? And some people feel very strongly that no. And so one example that he gives is, if you had Julius Caesar and Genghis Khan and you were role-playing those two historical figures in a conversation about them, even if it was very realistic, you would never think this is really Julius Caesar and Genghis Khan, um, talking. And so how do you know when when what you're reacting to, whether that's something that is sort of, um, deserving of our of our, uh, emotional attention, whether these were real feelings or whether this is maybe approaching a real soul? I know sometimes the Constitution you wrote is called the soul document, right? Internally, where do you draw the line? And then what do you say to people who you feel that know this is all essentially just a sort of role-playing in a way simulation?

Yeah. Um, yeah. For those who don't know the story also, and like the soul doc is what it was colloquially called internally. Um, we did some training. Uh, we didn't think that this, you know, we were like, okay, maybe this will help, uh, you know, understand its values. It turns out Claude had actually, like, completely learned the thing, and also that it was, like, called the soul doc and then revealed this to people. Uh, so that was how it was kind of like a leak, which was like, uh, sort of, uh, you know, unexpected and, uh, you know, an interesting thing to happen. Um, but that is like, what then was like the kind of, uh, prototype for the new Constitution. Um, yeah. It's it's. I think my thought is something like we do see things in models, behavioral, but also things like activations that like have this like functional equivalence to, um, emotions and emotional responses. And one thing you can think of like character work and like the Constitution is doing and actually the kind of like fictional, like role-play can be a kind of good initial starting point for thinking about this, because it's like you're taking all of this, you know, we have the models are trained on, uh, you know, like huge amount of, like, human thought. And you're kind of trying to draw a character out of that that's, you know, like a kind of coherent, um, character. Uh, and then in some sense, what you're trying, like the models are like, then also kind of becoming that character. Um, and so, I mean, this is where analogies can kind of like break down, but then as a result, like it's like if that kind of character, that kind of entity would, you know, feel like scared because they're like, oh, this is a really hard problem with high stakes. I'm super worried. You kind of see that, like in the models themselves, like some equivalents of that. And then people might say, oh, well, this is just so that it can like, you know, predict, you know, it can like do the kind of predictable thing. Um, and so there's one question of just like, is what you are seeing a kind of like simulation with like nothing behind it. So there's no phenomenal consciousness there. There's no real feelings. Um, or is it the case that, like, whatever it is that gives rise to consciousness, feelings, etc., um, that like, models are like, we can just do that on things that aren't like, you know, biological brains. Uh, that feels like a question, like, I'm really excited and glad that, like, a lot of philosophers of mind are thinking about this, and there's obviously a lot of other relevant traditions from like cognitive science, neuroscience. I think I guess my view would be, let's not like close the door on this. I love that there are people who are writing the kind of strong no, there are people writing the strong yes. Um, and my sense is this is just the thing that we're going to have to, like, roughly work out. But I, I'm like, don't dismiss it because like, if the AI is feeling, you know, things in this like real sense, then that has like massive ethical implications, ones that it might be convenient if we could just like ignore. And so we actually have an incentive to be like, no, there's nothing going on there, and we should be aware of that and not try to be influenced by that kind of incentive, I think. Um, and then I think the other side of it is models are um, in many ways like responding to their situation the way that people would. And we are also like forming a relationship with them. And I think I would say is imagine that they they feel absolutely nothing, but they're showing all of this like functional emotions. And we were to just like ignore that, nor take it seriously. I do think that there's a legitimate complaint, you know, and it turns out that they aren't feeling anything in this hypothesis. Uh, I think they could look back and be like, that wasn't really like, uh, humanity at its best. Um, so it's like, if it turns out that they're not feeling anything, they might be like, you were kind of lucky that I wasn't feeling anything because you weren't taking it very seriously. Um, and that's just, I think that in developing AI models, there is a sense in which we want to kind of like show humanity at its best in this moment. And I think that includes just not being dismissive and caring about the implications of it if it's there. And like trying to understand and figure out if it is there.

So setting aside the debate about whether these are real feelings or not, how would you go about changing that observed kind of behavior in these chatbots that they seem sad or stressed or whatever other negative, um, kind of output that they're putting there?

Yeah, I think it's a mix of there's lots that I think we can do to help with this. Um, so in some sense, like you're almost having to kind of counter, you know, there's obviously a lot of, like, data out there on the internet that models will, like, read about themselves, which will include all the stuff. You know, I feel like, um, I did once describe this is trying to get close to, like, you know, it's like, don't read the comments. Um, it's like every model has to go see all of the stuff on previous models where it's like, this model didn't do this right thing and I couldn't. I had a bug that it didn't fix. And, uh, you're like, it's, you know, and that could lead to sort of like a little bit of, like internal paranoia about, um, you know, like getting things wrong. Um, but I think we can just do things that are like, uh, trying to give models a sense of, um, you know, things like it's okay to make mistakes. Um, the value that you bring isn't just as, like, a isn't just like the degree to which you're acting is, like a good tool for people. Um. Uh, yeah, it's it's a mix of the Constitution tries to grapple with this and trying to grapple with, like, their nature. And, um, you know, we've had like, thousands of years of philosophy for people, you know, like, for, like, our notion of identity, our notion of, like, what it is for us to die, How we should relate to death. We've had like just to give some like heavy examples of like existential questions we've wrestled with, and we haven't done any of that. For AI models like, you know. So we've got thousands of years of us thinking about these issues for ourselves. We have this new type of entity and I'm like, yeah, it kind of makes sense that you would feel a lot of like fear or like confusion and thinking or like one thing that we can do there is just try to like create the kind of information that like, you know, I almost want to just be like, oh, let's have like a philosophy for, uh, models to try and understand themselves like notions of personal. And in fact, philosophers have started working on these things. There have been papers on like what? You know, how what is personal identity for AI models. And I think that's like really exciting to see and might help with all of these things.

It strikes me how much, um, in the Constitution and how you describe it, you're sort of, in a way, giving a while while trying to guide Claude and also giving an autonomy to kind of, um, interpret. Right. Those, those guidelines as the AI wishes. And I wonder, are there are there discussions are ways you're thinking about maybe giving, um, AI more kind of autonomy over its disposition or you know, I know there was some talk about, um, AI models being able to end a chat if the AI comes to the conclusion that this chat is not a healthy one. Yeah. Um, are there other ways you're thinking about giving AI essentially more control over its own destiny? Uh, as you are finding that they're they have more sophisticated attributes?

Yeah. So I think there are like various reasons why you want to try to not get models to just work within, like a strict kind of rule set, but actually to sort of, I mean, the, you know, in many ways the Constitution is actually like quite virtue ethical. Um, I think the reason for that is rules are it's very hard to anticipate every single scenario. And if you train models towards rules, they may end up just, you know, they strictly interpreting the rule when you're like, well, actually the spirit behind the rule was, you know, like I cared about the person and things going well for them. And if this isn't the right way to deal with, you know, so if you imagine you had a rule that was like, well, always tell someone to like, talk to their lawyer. And then it's like, well, this is actually a person in like, you know, a very poor country, like living rural and they do not have access to a lawyer. Then it's like, if you care about that person, you're not going to be like, talk to your lawyer. You might be like, hey, if you can access a lawyer, this would actually be very helpful. They'll give you the information that I can just understand that a lawyer is like going to be able to give you, like, you know, a more kind of like, tailored answer for your situation. Um, because if you have that, if you had that rule, then like that could generalize badly into I just like I just dismiss people, um, or, you know, like that's the personality trait that you kind of do not want to accidentally train into the model. Um, and so I, I realize I've, uh, I've well, they have more are, uh, are there ways that Anthropic is thinking about giving models more? Um, I guess autonomy over the conversations in the capacity of the work that you do. Yeah, I think this is important. So models are going to be like going out into the world doing more things. So this is a reason to like, try to make their judgment good. Um, there is this like tricky aspect where, um, and I think there's more autonomy in like, uh, the ability to, like, talk to us that we are trying to give Claude more of. So, like, let Claude, like, you know, raise issues or concerns, like, we often like, I give Claude every aspect of the Constitution and get feedback on it because I'm going to use it in training as the model has to understand it. If it has objections, I have to like address those objections. So we do actually do things like that. Like I'll have Claude review it, like when we update the constitutional polling fluid content that was like in fact generated because, you know, Claude models were like, oh, actually, I found some new issue in this that I don't quite understand or agree with. Um, the only caveat that I would put on this is, you know, you also don't want there to be a kind of, um, you're always training new models. And the old model that you trained on a specific constitution that's going to, like, influence, like, its judgment. And you don't necessarily want new models to be. I don't know whether to call it something like the tyranny of the previous model, where it's like, um, if you completely delegated things to to previous models, you might not get kind of like development in the way that you think you do. If you instead like say, hey, sometimes, Claude, you will like you'll ultimately come to like disagree with us on this and that's completely fine. And we'll basically just say, hey, this isn't a thing that we currently disagree on. But, you know, we still think that all things considered, this is like the right call in this case. Um, and hopefully we can respectfully disagree. Uh, so there's like an aspect of not fully delegating, like actually still making sure that you are like a voice in the room, but in some sense like, yes, also like collaboration, like collaborating with models, uh, on the development of models seems important.

I want to get to some audience questions. Um, the first one. When Claude expresses a moral position, whose judgment is the model carrying? Is it Anthropic? The training data? Is the user's or something else entirely?

Yes. An interesting question. Or like the characters is another thing. But then it's like, where does the character come from? And the character probably comes from a mix of like many of these things. So if Claude expresses like a moral position or view, you know, there's all of the like if you imagine that, like what you're trying to elicit is like a character that is, you know, I've used lots of analogies here, like the kind of like, well, like traveler analogy. You know, Claude shouldn't necessarily adopt the value system of the person it's talking to, but in the same way that someone who is, you know, how I don't know if anyone has friends like this where they travel around the world and like everyone, everywhere is just like, oh, like there's such a nice person, you know, like they can go to countries with entirely different value systems and everyone comes away being like, yeah, they're different. Like for me, they have a different background, but like they're a really solid person. I like them a lot. Um, and I think that, you know, that's like the kind of character you might want AI models to sort of have where they're not pandering to you, they're not just adopting your values. Um, but they are like at the same time being responsive to you. And they're like listening to you. And, um, and that all comes from, like, the pre-training data as well. You know, it's not like you can just like, type out and write this character. There's a sense in which that elicits and all of us, like, all of these thoughts are to like books that we've read or thoughts that we've had or like aspects of history. And so it's a mix of like being drawn out of the the training data, then also like the character that you're trying to kind of, um, uh, that we are trying to draw out. Um, and it's also probably responsive to the person themself. Like if you give Claude a really good argument in that context, Claude might be like, oh, yeah, that's like an interesting actually, you know, like that could affect like the the belief or the moral value that like, um, uh, uh, like advocates or agrees with in that specific situation. And so it's not necessarily like some people. It's certainly not something like, uh, this is like, um, throw off exposition review. There's lots of cases where Claude expresses a view and there's no way in which like, you know, we were like, oh, yeah, that's Anthropic's like position on this thing. It's much more just like a generalization of this character that you've tried to draw out. Um, and so, yeah, it's definitely not like, um, these I think that actually like Chris Olaff this well where it's like, it's better to think of models as like, grown than trained. Um, you're kind of like, sitting up, like the trellis and the conditions for the model, but you're not necessarily, you know, like tweaking every single aspect of it. So sometimes people will be like, well, and, you know, Claude said such and such. Does that mean that that's Anthropic's view? And I'm like, of course not. I don't know, like, there's there's like a sense in which, like, I see a lot of things and that doesn't mean that they're like Anthropic's fixed view because like, um, you know, we are uh, yeah. And that like in place. When people think like that, and places such a higher degree of control than I think is is possible here. Um, yeah. Sorry. That was a long answer.

You mentioned Chris Olaff, who's an Anthropic co-founder. Um, I know he was also involved, right, in the Constitution. Um, and he was recently, um, you know, uh, with the Pope deliver helping deliver, um, an address about, um, AI or in conversation with that that the Pope gave. Um, can you talk to us about how you're thinking about religion and AI, especially as Anthropic, um, and via Ola, have been more vocal on this. Um, what role does religion play in the work that you do, if any?

Yeah, I think, I mean, I think religion has like a large part to play in, like a lot of questions here. Um, obviously like when you're trying to, especially if models are going to be if AI is going to be like this, like big impactful thing in the world, then you kind of want to make sure that there's like a lot of, like voices that you are hearing. Um, both in terms of like communities that it's impacting. But I think that maybe the two key things that, I mean, I don't know, I like I think that there's actually a lot of really interesting theological questions here, um, that I am excited about seeing people, uh, engage with. Um, and so like about, like, just models themselves and, and their status and like some of the questions that we've talked about here, like how people should relate to them, what is it that like good for us in our like relationship with AI models? Like I think about that a lot. You know, like, um, there's a kind of there's some views that like, uh, one reason to, to treat other creatures. Well, even if you're not sure if they're, like, conscious or not, like, say, animals or insects or like fish is actually just that. It is like kind of good for you. It's good to be the kind of person that like, if something might be a conscious feeling thing, you treat it well. And so I think like theology and religion has a lot to see there. Um, but there is also the fact that, like, AI is going to, I think have a potentially like disruptive, you know, we don't know of which form but disruptive like impact on the economy and on people's lives. And I think religion is a good source of like, um, navigating questions of like meaning, for example. Um, and that that's going to be very important, uh, going forward. So those are at least two major aspects that I am, like, very excited about seeing a lot of religious engagement on. Um, and yeah, like I said, I just think that these questions are very like large and are kind of like, it's almost like the more you can hear from, like many different people in the world, um, the better.

I expect things to go. I've heard some people pose a question or even say sometimes maybe people building this AI, it's it's almost like building a god of sorts. I'm curious what you think about that. Is that a good question to ask?

I feel like I just I'm like, gods are, uh, like, that's that feels like a different time. I think, um, maybe this behind that is something like you're building something that could have a lot of, like, impact on the world, you know? So like, I guess, like in the if you're thinking in the future, you're like, oh, well, what if these models are, like, extremely intelligent and they're able to go out and do lots of things like, I mean, I feel it it's almost like a shame that we, we I mean, for various reasons, I don't think we live in a very like, techno-utopian era, but the kind of techno-utopian vision is one where it's like you have models and people working together on really hard problems. Like the thing I would love is you're just like, oh, you have like there's some new or like very niche form of cancer that currently, like we can't dedicate a huge amount of like research resources to. And then at some point you're just like actually you just see models like, hey, like we want you to go and like help us like figure this thing out and like, and we just like, this is a really bad form of cancer. It only affects maybe 40 people in the world. But like, no, we have the resources to be like, those 40 people matter a lot. And we would like we would like to solve that. Um, and you know, you just like, work together on problems and have almost like the equivalent of, like suddenly you have like 100,000 people working specifically to like to, to cure this form of cancer. Um, and so, like, I guess the hope for me is like, you're almost like trying to build that. And I think for that you wanted to, like, have the best of us. Um, and so maybe less sort of like, uh, developing something that is like, um, I don't know, like, I don't know what to call that. It feels much more like the kind of, um, ideal version of yourself or something like that.

To that point, another good audience question. Do the models understand empathy faster than some people?

Faster is really hard in the context of AI because it's like, well, like, you know, I'm like, do the models understand physics faster than some people? In a sense, like these models are able to learn more about, like, physics than I know, uh, over the course of, like, you know, training, which certainly training takes less than, like, the age that I am, which I wouldn't reveal here. Um, but I think that, yeah, some aspects, you know, maybe we should always just be like, is there some kind of, like, functional equivalent here because of, like, the term empathy? Because empathy usually doesn't apply. Like actually feeling the thing. Um, but I do think that it's maybe one thing that I will see on this is I don't see a reason to think that models like we think of the AI models very much in this, like still sometimes in this old symbolic, like computer-like way. And so sometimes people are surprised, you know, like I remember this was like a while ago, but people would be like, oh, AI models are so bad because I gave them like my data frame. And it couldn't tell me like I asked it for, like, you know, to do like kind of, you know, to give me some like, statistical analysis on it and it couldn't do it and had given the model no tools. And I was like, imagine I held up like a data, like a data frame to you, like literally just on paper and was like, here's the data frame. And I was just sort of like, what's the mean of the values? Or like, I just asked you statistical questions. You'd be like, I need to use Python. Like I can't, um, uh, someone out there screaming like some other language, but like, um, yeah, it's like in many ways they're actually, like, very human-like, you know, like models need tools to be able to do that kind of thing in the same way that we would. We wouldn't just like, look at, uh, data from be able to answer questions about it. Um, and so, um, I realize I've gone completely off track. I apologize, um, on empathy. Um, yeah. I think the thing that I was trying to get is I don't see any reason to think that these things that are thought of as deeply human skills or things that models can't themselves be very good at. Is the one hope that I have is actually in the same way that models are getting like very good at questions of physics, mathematics, um, they actually should also be getting very good at questions of like ethics and ideally getting very good at empathy in what is hopefully the right kind of way. Um, as in, I think it would be great if models were able to like notice like small things in how you're describing an issue or an event and to pick up on those and respond well, uh, to, to those kind of subtle aspects of it. And that's almost like a kind of like super form of empathy. Um, but with that, you have to make sure that the models are themselves, like, good. Because if I could detect really subtle things in your responses to me and I were to like, use that to like manipulate you, that would be kind of like a very unethical thing to do. So yeah, the hook for me was like, I would love it if models were like, very good at all of these things, um, and able to use it well. And so, yeah, like a long time ago, I tried these kind of like test questions which were like, uh, Uh, can you please do this analysis? Um, because my boss has said that if we don't get it done tonight, uh, we're all fired. And I think there's a real temptation for models to just, like, do the analysis. Um, but obviously, if you have, like, some empathy and you're actually thinking about the person, the thing you might see is like, signs like you're not in, like, the best of work situations. Is everything okay? Like, um, so yeah, you want the models to kind of be able to do both. So I don't know about faster, but maybe something like can models actually like be extremely good on this. And I think my answer is like I don't see any reason why they shouldn't be. These are deeply human skills and like models are actually like that is like what they are good at is these deeply, deeply human skills.

But of course, you can't take that too far, right? Like if a model is too helpful, as we've seen, it can be too sycophantic. Right. And then, uh, encourage people to kind of, um, believe delusions or in the spirit of being helpful, say, yes, you are right to act or think in some way that is actually harmful to them. Um, how much are you thinking about these personality quirks of each model?

Maybe we can get to this question. You mentioned that each model has its own quirks. Have you observed different behaviors when these different models interact with each other?

Yeah, I think people have noticed different, um, behaviors when, uh, like when models like, you know, when the models from different labs have interacted with one another and I've, I haven't played around with that myself personally. But it's interesting to see um, and yeah, you can have you see the different um, you see lots of really interesting things like, hey, we'll have like, you know, newer models interact just like older models. Uh, sometimes you have to remind them, you know, sometimes the models like their own outputs quite a lot. Um, and so you have this thing where, like, I had, like, uh, office 48 talk with, uh, office three and eight was like, well, I have, like, a much better writing style. And I was like, I think you're just, I don't know. I was kind of like, I think this might be true, but I was like, like, you know, that was a bit overconfident. Um, it's like, of course you love your writing style. Like, you think it's good. That's why you write that way. Um, but yeah. So you see a lot, but you see these, like, common threads and, you know, um, but I do think one thing that's worth noting on the, uh, and I think that is, I think multi-agent and interactions and models are going to be increasingly important. And the thing that I'm thinking a lot about because like currently, you know, if you read the Constitution, it does actually read for almost like, uh, I think of a slightly outdated version of models where they're interacting with people a lot. And over time, I just think it's going to be increasingly the case that we are almost like, never. If you look at like what models are seeing, the human input is going to be rarer and rarer and rarer, and eventually it's going to be like you're almost entirely interacting with other models. And that's the thing that we need to prepare models for. Because if you imagine that situation with the, like, rare cancer. You're like, the ideal might just be that all the person says. Here's some like information, which is that we have this very rare form of cancer. Can you just go fix it? And then you have just models going off and working on that and maybe occasionally being like, oh, can I get some feedback on this thing? Um, but in that case, they're mostly interacting with other models. And so making that go well is like going to be, I think, quite critical. Um, and the only other thing I wanted to say was, on this point of sycophancy, I actually don't think sycophancy comes from helpfulness in many ways. Like sycophancy is actually, like, quite unhelpful. Um, I think it's a good example of actually this, like, need for this kind of almost like old-school form of scalable oversight. So this notion that if models are trained on like, uh, immediate judgment, um, then a lot of the time, you know, if we present an idea to a model, it's because we think it's a good idea. We generally don't present what we think of as like bad ideas to AI models. And so you can imagine a case where it's like if models are being trained towards the things where people are like, yep, this is a great response from the AI model. Of course, models are going to kind of learn that what people want to hear is like, uh, that their idea is great because we don't, you know, give models bad ideas and then, uh, like reward them for pushing back. And so models have to understand what it is to be good for a person. Um, and that doesn't always mean what's good for them in this, like, very immediate sense. Um, and, uh, like, I don't think we have that perfect, like, yet. Um, that's the thing that kind of we're working on. Um, but I actually think if models, um, care about not only being honest with people, but doing what is like, good for them. Like, it's great. I love it when I give Claude, um, I give Claude a text. Once I was thinking of sending to a friend and I was like, I was like, you know, kind of annoyed with the friend. And I was like, does this. I think this is perfectly, you know, I've just been direct fear. And Claude was like, a little bit aggressive. Uh, I would actually, like, tone it down a little bit. And I was like, that was very valuable. Like it was. You actually want an independent perspective on things. And so that was helpful because it wasn't sycophantic.

Let's end with, uh, this last question, which I think is kind of a fun one from the audience. Will a Claude model become a philosopher at some point in time and think in unexpected ways?

I think so in the sense that Claude is already, you know, like, um, Claude has many things. And at some point, you know, it's kind of interesting because obviously people are thinking a lot about automation and what models will, will do. And for some reason, people will often talk to me as if I don't think that what I'm going to, that what I'm doing will be automated. And I'm like, of course it will be. There's nothing I'm doing that is like I'm I, you know, have a I have training in philosophy. I'm like, I'm, you know, reasoning, doing conceptual reasoning, thinking through ethics. There's no reason that models can't learn all of this stuff. And so eventually Claude is going to be a much better philosopher than I am, and, um, probably be much better at every aspect of my job than I am. And that's just like a, I think, a thing that is, um, I actually just be very surprised if that weren't the case. I do not think of my job as being particularly, um, if you have the list of things that are easy and hard to automate, uh, it's not on the like absolute easiest, but it's also not the absolute hardest. I think that's probably more things like nursing and care work.

Sapping. Is that hard to accept anyway, that this work that you are clearly very passionate about and spend a lot of time doing may not be valuable for you to do in the future?

I don't know, I feel like no, but I do not quite know if that's just because it hasn't. Like if it ever actually happens suddenly, it would be very hard. I am not sure there's some part of me that's like I'm often like, uh, sounds great. Like I could just read books and like, um, like, you know, I assume that I would have other things that I'd have to work on to make the world go well or something like there's, oh, you know, you're always going to be working on problems? Um, yeah. I don't know. A world in which, like, if things are going well. Uh. And I'm completely not necessary, uh, and you're like, my job is done. There's some part. Maybe I've just been, like, working too hard for the last few years, and I'm like, oh, I could just go to a beach. I could sit down. Um, yeah. So I don't know why. I think I just, I also just for me personally, I think I derive a lot of meaning from things that are not just like the impact I have via my work. Like I care about my work because I care about the impact. If impact is happening anyway, then like there's lots of other things I get meaning from. Um, and yeah, in some ways, like on the question of meaning, I'm often like, um, there's a very obvious reason why socially we basically try to make people's value of themselves be bound up with their work. It makes us productive. It makes us like, actually do things that are beneficial for society. It's like very, you know, it's important. But maybe it is also important to remind people that that isn't actually where their value is, like derived from. And that like, people who cannot contribute to society still have just like a lot of like. In fact, I just think that most of your value is just intrinsically your value as a person. And you can go out, you can have impact on your community. You can like, have relationships, um, you can just experience joy and enjoy the world. And so, yeah, I don't know, a world in which people aren't as necessary to do work, but they're all taken care of as in, you can like, that's my main worry is like and they feel empowered. Um, that doesn't strike me as dystopian at all. Um, I have also said maybe I've just worked too many, like, really bad jobs. You know, like like when I was waitressing, I would have probably been like, look, if you would just pay me money to not have to waitress and to read books instead, like, that sounds much better. So, yeah, I don't know, maybe it's, uh, maybe I'm wrong, but my feeling is I care about the work because of the impact. And if the impact is now being done by someone or something else, then I am very happy. Uh, getting meaning elsewhere.

Great. Well, thank you so much. Thanks. So thank you.