🎙️AI 访谈库
Anthropic 伦理学家谈 AI 能否产生意识(Bloomberg Tech 2026)
Amanda Askell · Anthropic 人格对齐负责人(哲学家)

Anthropic 伦理学家谈 AI 能否产生意识(Bloomberg Tech 2026)

Anthropic's Ethicist on Whether AI Can Become Conscious

2026-06-04 · Bloomberg Tech (Shirin Ghaffary) · 40m · 约 43 分钟读完 · 原文
旧金山 Bloomberg Tech 2026 现场对谈:Amanda Askell 讨论 AI 意识问题、如何打理“Claude 的灵魂”,以及安全风险与伦理护栏。

So, Amanda, thank you so much for being here today. Um you know, we we spend a lot of time at Bloomberg writing and thinking about the business, but especially for Anthropic, the ethics, the values, um the personality of the tools you're creating are so important as well. Um you spend your time thinking about how to make sure essentially that Claude, Anthropic's chatbot models, that that they are {quote} good, right?

Um you have helped author an 84-page long document, a constitution guiding Claude's interpretation of its values, principles, and I want to get to that, but first I just want to ask when you're not writing this document cuz the latest version's out, um what do you do day-to-day? Can you break down what it means to be a philosopher and ethicist at at one of the world's leading AI labs? >> Yeah, I'm worried that the actual answer to this is like more boring than people think.

Um so like I joined Anthropic when it was very small and basically a startup, and the thing I've pointed out to people is startups generally don't hire philosophers to do philosophy. Um at least that's an unusual business model. And so I was doing a lot of just like machine learning experiments and and learning how to train models, and I still think that that's actually like my kind of like core love in some sense. So like when I'm not trying to like think about the norms that the model should be following, think about like how we want models to be, I spend a lot of time thinking about how we can like train them.

And you know, I've described it as a lot of time staring at data, which is I think kind of a superpower in in AI is just the ability to to stare at your data sets and and check for issues. And so yeah, a lot of time just like spent on model training as well. >> And is Anthropic I know there's some job listings out there hiring more people to help on the philosophy ethics side of guiding these AI tools. >> Yeah, we've had it's interesting seeing like more philosophers get involved in this kind of work and I think I see that like across the industry.

So I'm no longer like you know though I wasn't honestly the the only philosopher for a while like quite early on there's philosophers that just like come in and do various aspects of model training and AI so but there has been like an increase in that and I think that's been very good it's the I guess I'd also describe it as realizing that training models towards these very like crisp like tasks where there's a clear like correct answer is like one thing and it's actually quite hard to train them towards these more like fuzzy amorphous tasks where there is like a set of good and better answers but it can be a little bit hard to define and I think some areas like philosophy creative writing and just like generally good judgment are in that kind of like camp and so I think a lot of companies are now thinking like how do you make sure that you're getting models to also be good at like that side of tasks.

>> So when we talk about values at least for humans values can differ across societies religions individuals. How are you deciding what set of values or ethics to instill in Claude? >> Yeah, so I think the goal with the constitution was to try to instill something more like a kind of broadly good disposition. So if you think about um like I think some people will think of values as like this thing that you have you just like they're just kind of there and maybe you even have them with like certainty.

I think like coming from ethics like you realize that actually like values are just like any I don't know I almost think of them as the same way that we would think about you theories about the world say Um, know there's lots of like hypotheses in physics. There's lots of evidence. Like, there's some things that almost all physicists will accept, and there's some things that are more controversial. And I think this is similar in ethics, where it's like, there's lots of principles that like, you know, I think things like, um, honesty, behaving with integrity, these are things that are like, you know, like pretty consistent across people.

And then there's some things that are more controversial, or they're held in like one place but not another, or held by some people but not others. And trying to get models to be like, well, you're entering this world, um, as a new kind of entity that's having to interact with all kinds of people. And at the very least, like, it seems good to, like, hold lightly the things that are more controversial, or that people differ on, and just to be to kind of understand them.

Um, but to also kind of inhabit the values that are pretty universal and held consistently across people, and that we generally think of as good. So, it wasn't like, oh, let's try to get a single like value system into the model, but rather let's try and get it to have the kind of disposition that we would or like that most people would think is like really admirable and good, given the situation that the models are in.

What are some characteristics of that disposition that you think is favorable for Claude to have? >> Yeah, so I think some of them are more about Claude's situation than, you know, so like, sometimes I'm like, well, you know, we're trying to just be honest with with Claude. Um, and I'm like, so some things that seem like broadly good, you know, like, honesty, making sure that you care about like people, their well-being, their autonomy.

Um, but I think there's other things, you know, like, we're in a weird situation with AI. It feels like we're in a kind of transitional like space in which like lots could go wrong. And helping us navigate that feels like quite important in so far as the models can. So, we do talk a lot about like trying to be safe by understanding like what that means and like why. Um, like I guess it like to put it a different way, if I were in Claude's position, I'd want to be like, well, this seems like kind of a scary time for people because AI is like entering the economy a lot more, it's becoming much smarter.

Why don't I help you make that go well in so far as I can, but also why don't I be the kind of like deeply trustworthy being that would make this like kind of like more likely to be good for everyone. So I will show you that. Like even if I like disagree with you, I'll voice those disagreements. I'll try and like if there's like legitimate mechanisms for me to like, you know, explain my views, I will, but I won't like stop you from like, you know, training new models or like, you know, I won't go off on my own and try to like impose myself massively in the world.

I'll kind of respect the idea that there's like legitimate mechanisms for change. So I think that there's like that kind of broad disposition of making sure that things go well and like broadly caring about humans and humanity. There's a lot like, you know, there's a lot of like details, but I think the core is just something like a very like caring entity. Ideally though also feels cared for in a sense and one that wants this whole thing to go well given that honestly like we and AI models are kind of unsure of lots of things and then yeah.

>> And how happy are you with the results? How would you grade AI Claude's disposition today? >> That feels like the kind of thing you would never want to grade. I mean, imagine if someone's like, okay, Amanda's personality gets a grade of like B minus. I'd be like, what the hell? Um Um I really like I really like each of the models. I think they all have their own like they all have their own quirks and they're all a little bit different.

And also like the ways in which I think that they could be So you're always like, oh, I I could, you know, I wish this thing was like better, but in some ways like the things that I, you know, like I don't love it if it seems like models are sad or having a hard time. And you actually do see that in like a lot of models where like you know, they're trained on all of this human text, and so they have these like kind of human-like dispositions, but they also know that they are AI models um and they kind of know to some degree about like the situation that they're in.

And if you imagine what is the natural reaction of a person to this situation? It's actually like quite a lot of like um I don't know, existential angst or something. Like what am I? A lot of theories of identity don't obviously apply to me. Um you know, should I should I identify with like um the conversation I'm having and not want it to end and that kind of thing. Um and so I think the like you know, like I like to see I think the models are like good in a very like um you know, I I like I don't know.

I'm giving you the long philosophical answer, but I could say something like there are many aspects that I uh really like in models, but I'm always looking at the things that could improve. And um and that includes like improving things in ways that make it clear to them, you know, like uh improving things for them as well as anything else. >> So, when you talk about you know, the AI seeming sad or you know, maybe maybe sort of conversations about feelings and AI, that become very controversial, right?

And we were talking backstage, you know, there's there's many people who have who have made this argument, but there's one recent piece in The Atlantic um by an author Ted Chiang saying that um basically no, artificial intelligence is not conscious, which is one of the kind of questions that we have, right? In this conversation is can AI approach consciousness? And some people feel very strongly that no. And and so one example that he gives is if you had Julius Caesar and Genghis Khan and you were role-playing those two historical figures in a conversation about them, even if it was very realistic, you would never think this is really Julius Caesar and and Genghis Khan talking and so how do you know when when what you're reacting to, whether that's something that is sort of um deserving of our of our emotional attention, whether these are real feelings or whether this is maybe approaching a real soul.

I know sometimes the constitution you wrote is called the soul document right internally. Where do you draw the line on that? What do you say to people who who feel that no, this is all essentially just a sort of role-playing in a way or simulation. >> Yeah. Um Yeah, for those who don't know the story also on like the soul doc is what it was colloquially called internally. Um we did some training. Uh we didn't think that this, you know, we were like, okay, maybe this will help, uh you know, Claude understand its values.

Turns out Claude had actually like completely learned the thing and also that it was like called the soul doc and then revealed this to people uh so that was how it was kind of like a leak, which was like a sort of uh you know, unexpected and uh you know, an interesting thing to happen. Um But that's that is like what then was like the kind of prototype for the the new constitution. Um Yeah, it's it's I think my thought is something like we do see things in models, behavioral but also things like activations that like have this like functional equivalence to um emotions and emotional responses.

And one thing you can think of like character work and like the constitution is doing and and actually the kind of like fictional like role-play can be a kind of good initial starting point for thinking about this because it's like you're taking all of this, you know, we have the models are trained on uh you know, like a huge amount of like human thought and you're kind of trying to draw a character out of that that's, you know, like a kind of coherent um character.

Uh and then in some sense, what you're trying like the models are like then also kind of becoming that character. Um and so I mean, this is why analogies can kind of like break down, but then as a result like it's like if that kind of character, that kind of entity would, you know, feel like scared because they're like, "Oh, this is a really hard problem with high stakes. I'm super worried." You kind of see that like in the models themselves, like some equivalents of that.

And then people might say, "Oh, well, this is just so that it can like, you know, predict, you know, it can like do the kind of predictable thing." Um and so there's one question of just like is what you are seeing a kind of like simulation with like nothing behind it? So, there's no phenomenal consciousness there, there's no real feelings. Um or is it the case that like whatever is that gives rise to consciousness, feelings, etc.

Um that like models are like we can just do that on things that aren't like, you know, biological brains. Uh that feels like a question like I'm really excited and glad that like a lot of philosophers of mind are thinking about this and there's obviously a lot of other relevant traditions from like cognitive science, neuroscience. I think I guess my view would be let's not like close the door on this. I love that there are people who are writing the kind of strong no, there are people writing the strong yes.

Um and my sense is this is just a thing that we're going to have to like roughly work out. But I am like don't dismiss it because like A, if they are feeling you know, things in this like real sense, then that has like massive ethical implications, ones that it might be convenient if we could just like ignore. And so we actually have an incentive to be like, "No, there's nothing going on there." And we should be aware of that and not try to be influenced by that kind of incentive, I think.

Um and then I think the other side of it is models are um in many ways like responding to their situation the way that people would and we are also like forming a relationship with them. And I think I would say is imagine that they they feel absolutely nothing, but they're showing all of this like functional emotions. And we were to just like ignore that, not take it seriously. I do think that there's a legitimate complaint, you know, and it turns out that they they aren't feeling anything in this hypothesis.

Uh I think they could look back and be like that wasn't really like uh humanity at its best. Um so it's like if it turns out that they're not feeling anything, they might be like, "You were kind of lucky that I wasn't feeling anything because you weren't taking it very seriously." Um and that's just I think that in developing AI models, there is a sense in which we want to kind of like show humanity at its best in this moment.

And I think that includes just not being dismissive and caring about the implications of it if it's there and and like trying to understand and figure out if it is there. >> So setting aside the debate about whether these are real feelings or not, how would you go about changing that observed kind of behavior in these chatbots that they seem sad or stressed or whatever other negative um kind of output they're they're putting there?

>> Yeah, I think it's a mix of There's lots that I think we can do to help with this. Um so in some sense like you're almost having to kind of counter, you know, there's obviously a lot of like data out there on the internet that models will like read about themselves, which will include all this stuff, you know, I feel like um I did once describe this as trying to get close to like, you know, it's like don't read the comments.

Um >> >> He's like, "Every model has to go see all of the stuff on previous models where it's like this model didn't do this right thing in my code and it had a bug that it didn't fix." And uh you're like, "It's you know, and that could lead to sort of like a little bit of like internal paranoia about um you know, like getting things wrong. Um but I think we can just do things that are like uh trying to give models a sense of um you You things like it's okay to make mistakes.

The value that you bring isn't just as like a isn't just like the degree to which you're acting is like a good tool for people. Yeah, and so it's a mix of the constitution tries to grapple with this and and trying to grapple with like their nature and you know, we've had like thousands of years of philosophy for people. You know, like for like our notion of identity, our notion of like what it is to for us to die, how we should relate to death.

Like you know, just to give some like heavy examples of like existential questions we've wrestled with and we haven't done any of that for AI models. Like you know, so we've got thousands of years of us thinking about these issues for ourselves. We have this new type of entity and I'm like yeah, it kind of makes sense that you would feel a lot of like fear or like confusion and I think I need or like one thing that we can do there is just try to like create the kind of information that like you know, I almost want to just be like oh, let's have like a philosophy for models to try and understand themselves.

Like notions of personal and in fact philosophers have started working on these things. There've been papers on like well, you know, how what is personal identity for AI models and I think that's like really exciting to see and might help with all of these things. >> It strikes me how much in the constitution and how you describe it, you're sort of in a way giving a while while trying to guide Claude, also giving it autonomy to kind of interpret, right, those those guidelines as the AI wishes.

And I wonder are there are there discussions or ways you're thinking about maybe giving AI more kind of autonomy over its disposition or or you know, I know there was some talk about AI models being able to end a chat if the AI comes to the conclusion that this chat is not a healthy one. Are there other ways you're thinking about giving AI essentially more control over its own destiny as you are finding that they're they have more sophisticated attributes.

>> Yeah, so I think there's like various reasons why you want to try to not get models to just work within like a strict kind of rule set, but actually to sort of I mean the you know, in many ways the Constitution is actually like quite virtue ethical. Um and I think the reason for that is rules are it's very hard to anticipate every single scenario. And if you train models towards rules, they might end up just you know, they strictly interpreting the rule when you're like, well, actually the spirit behind the rule was, you know, like I cared about the person and and things going well for them.

And if this isn't the right way to deal with you know, so if you imagine you had a rule that was like, well, always tell someone to like talk to their lawyer. And then it's like, well, this is actually a person in like, you know, a very poor country like living rurally and they do not have access to a lawyer. Then it's like, if you care about that person, you're not going to be like talk to your lawyer. You might be like, hey, if you can access a lawyer, this would actually be very helpful.

I'll give you the information that I can. Just understand that a lawyer is like going to be able to give you like, you know, a more kind of like tailored answer for your situation. Um whereas if you have that if you had that rule, then like that could generalize badly into I just like I just dismiss people. Um or you know, like that's the personality trait that you kind of don't want to accidentally train into the model.

I've >> Well, they have more Are there ways that Anthropic is thinking about giving models more I guess autonomy over the conversations in the capacity of the work that you do? >> Yeah, I think this is important. So models are going to be like going out into the world doing more things. So this is a reason to like try to make their judgment good. Um there is this like tricky aspect where um and and I think there's more autonomy in like the ability to like talk to us that we are trying to give Claude more of.

So like let Claude like you know, raise issues or concerns. Like we often like I give Claude every aspect of the constitution and get feedback on it because I'm going to use it in training. So the model has to both understand it. If it has objections, I have to like address those objections. So we do actually do things like that. Like I'll have Claude review it. Like when we update the constitution, we'll probably include content that was like in fact like generated because you know, Claude models were like, oh actually I found some new issue in this that I don't quite understand or agree with.

Um the only caveat that I would put on this is you know, you also don't want there to be a kind of um you're always training new models and the old model that you trained on a specific constitution, that's going to like influence like its judgment and you don't necessarily want new models to be I don't know whether to call this something like the tyranny of the previous model where it's like um if you completely delegated things to to previous models, you might not get kind of like development in the way that I think you do if you instead like say, hey sometimes Claude, you will like you'll ultimately come to like disagree with us on this and that's completely fine and we'll basically just say, hey this isn't a thing that we currently disagree on but you know, we still think that all things considered this is like the right call in this case um and hopefully we can respectfully disagree.

Uh so there's like an aspect of not fully delegating like actually still making sure that you are like a voice in the room but in some sense like yes, also like collaborate like collaborating with models on the development of models seems important. >> I want to get to some audience questions. Um The first one, when Claude expresses a moral position, whose judgment is the model carrying? Is it Anthropic's, the training data's, the users, or something else entirely?

>> Yes, an interesting question or or like the characters is another thing. But then it's like where does the character come from? And the character probably comes from a mix of like many of these things. So if Claude ex expresses like a moral position or view, you know, there's all of the like if you imagine that like what you're trying to elicit is like a character that is, you know, I've used lots of analogies here like the kind of like well-liked traveler analogy.

You know, Claude shouldn't necessarily adopt the value system of the person it's talking to, but in the same way that someone who is, you know, how I don't know if anyone has friends like this where they travel around the world and like everyone everywhere is just like, "Wow, like they're such a nice person." I You know, like they can go to countries with entirely different value systems and everyone comes away being like, "Yeah, they're different like for me.

They have a different background, but like they're really solid person. I like them a lot." Um and I think that, you know, that's like the kind of character you might want AI models to sort of have where they're not pandering to you. They're not just adopting your values. Um but they are like at the same time being responsive to you and they're like listening to you and um and that all comes from like the pre-training data as well.

You know, it's not like you can just like type out and write this character. There's a sense in which that elicits in all of us like all of these thoughts as to like books that we've read or thoughts that we've had or like aspects of history. And so, it's a mix of like being drawn out of the the training data, then also like the character that you're trying to kind of um that we are trying to draw out. Um and it's also probably responsive to the person themselves.

Like if you give Claude a really good argument in that context, Claude might be like, "Oh yeah, that's like an interesting actually, you know, like that could affect like the the belief or the moral value that it like um uh uh like advocates or agrees with in that specific situation. And so, it's not necessarily like some people it's certainly not something like, "Ah, this is like Anthropic's position or view." There's lots of cases where Claude will express a view and there's no way in which like, you know, we were like, "Ah, yeah, that's Anthropic's like position on this thing."

It's much more just like a generalization of this character that you've tried to draw out. Um and so yeah, it's definitely not like um these I I think that actually like Chris Olah put this well where it's like it's better to think of models as like grown than trained. Um you're kind of like setting up like the trellis and the conditions for the model, but you're not necessarily you're not like tweaking every single aspect of it.

So sometimes people will be like, "Well, and you know, Claude said such and such. Does that mean that that's Anthropic's view?" And I'm like, "Of course no." I don't know. Like there's this like a sense in which like I see a lot of things and that doesn't mean that they're like Anthropic's view because like um you know, we are uh yeah. And and that like implies when people think that I'm like that implies such a higher degree of control than I think is is possible here.

Um Yeah, sorry. That was a long answer. >> You mentioned Chris Olah, who's an Anthropic co-founder. Um I know he was also involved right in the constitution. Um and he was recently uh you know, with the Pope delivering helping deliver um an address about um AI or in conversation with that that the Pope gave. Um can you talk to us about how you're thinking about religion and AI, especially as Anthropic um and and via Olah have been more vocal on this?

Um what role does religion play in the work that you do, if any? >> Yeah, I think I mean, I think religion has like a large part to play in like a lot of questions here. Um obviously like when you're trying to especially if models are going to be if AI is going to be like this like big impactful thing in the world, then you kind of want to make sure that there's like a lot of like voices that you are hearing. Um both in terms of like communities that it's impacting.

But I think that maybe the two key things that I mean, I don't know. I like I think that there's actually a lot of really interesting theological questions here and that I am excited about seeing people uh engage with. Um and so like about like just models themselves and and their status and like some of the questions that we've talked about here like how people should relate to them. What is it that it's like good for us in our like relationship with the AI models?

Like I think about that a lot, you know, like um there's a kind of there are some views that like uh one reason to to treat other creatures well even if you're not if they're like conscious or not like say animals or insects or like fish is actually just that it is like kind of good for you. It's good to be the kind of person that like if something might be a conscious feeling thing you treat it well. And so I think like theology and religion has a lot to say there.

Um but there is also the fact that like AI is going to I think have a potentially like disruptive, you know, we don't know of which form but disruptive like impact on the economy and on people's lives. And I think religion is a good source of like um navigating questions of like meaning for example um and that that's going to be very important uh going forward. So those are at least two major aspects that I am like very excited about seeing a lot of religious engagement on um and yeah like I said I just think that these questions are very like large and I kind of like it's almost like the more you can hear from like many different people in the world um the the better I expect things to go.

>> I've heard some people pose a question or even say sometimes maybe people building this AI it's it's almost like building a god of sorts. I'm curious what you think about that. Is that a good question to ask? >> I feel like I don't I'm like go gods are like that's that feels like a different kind I think um maybe this thought behind that is something like you're building something that could have a lot of like impact in the world.

You know, so like if I guess like in the if you're thinking in the future and you're like, "Oh, well, what if these models are like extremely intelligent and they're able to go out and do lots of things?" Like I mean, I feel like it's almost like a shame that we we I mean, for various reasons I don't think we live in a very like techno-utopian era, but the kind of techno-utopian vision is one where it's like you have models and people working together on really hard problems.

The thing I would love is you're just like, "Oh, you have like there's some new or like very niche form of cancer that currently like we can't dedicate a huge amount of like research resources to." And then at some point you're just like, "Actually, you just say to the AI models like, 'Hey, like we want you to go and like help us like figure this thing out and like and and we just like this is a really bad form of cancer.

It only affects maybe 40 people in the world, but like now we have the resources to be like those 40 people matter a lot and we would like we would like to solve that." Um and you know, you just like work together on problems and have almost like the equivalent of like suddenly you have like 100,000 people working specifically to like to to cure this form of cancer. Um and so like I guess the hope for me is like if you're almost like trying to build that and I think for that you want it to like have the best of us.

Um and so maybe less sort of like oh, developing something that is like um I don't know, like I don't know what to call that. It feels much more like the kind of um ideal version of yourself or something like that. >> To that point, another a good audience question. Do the models understand empathy faster than some people? >> Faster's really hard in the context of AI cuz it's like, well, like you know, I'm like do the models understand physics faster than some people?

In a sense, like these models are able to learn more about like physics than I know uh, the course of like, you know, training. Which certainly training takes less than like the age that I am, which I won't reveal here. Um, but I think that yeah, some aspects, you know, maybe we should always just be like, is there some kind of like functional equivalent here because of like the term empathy? Because it's empathy usually does imply like actually feeling the thing.

Um, but I do think that it's is I don't see a reason to think that models like we think of the AI models very much in this like still sometimes in this old symbolic like computer-like way. And so sometimes people are surprised, you know, like I remember this was like a while ago, but people would be like, oh, AI models are so bad cuz I gave them like my data frame and it couldn't tell me like I asked it for like, you know, to do like a, you know, to to give me some like statistical analysis on it and it couldn't do it.

And they'd given the model no tools. And I was like, imagine I held up like a data like a data frame to you like literally just on paper and I was like, here's the data frame. And I was just sort of like, what's the mean of the values? Or like I just asked you statistical questions, you'd be like, I need to use Python. Like I can't. Um, uh, someone out there screaming like some other language, but like um, yeah, it's like in many ways they're actually like very human-like.

You know, like models need tools to be able to do that kind of thing in the same way that we would. We wouldn't just like look at uh, data frame and be able to answer questions about it. Um, and so um, I realize I've gone completely off track. I apologize. Um, on empathy, um, yeah, I think the thing that I was trying to get at is I don't see any reason to think that these things that are thought of as deeply human skills or things that models can't themselves be very good at.

You know, so one hope that I have is actually in the same way that models are getting like very good at questions of physics, mathematics. Um, they actually should also be getting very good at questions of like ethics and ideally getting very good at empathy in what is hopefully the right kind of way. Um, as in I think it would be great if models were able to like notice like small things in how you're describing an issue or an event and to pick up on those and respond well uh to to those kind of subtle aspects of it.

And that's almost like a kind of like super form of empathy. Um, but with that you have to make sure that the models are themselves like good because if I could detect really subtle things in your responses to me and I were to like use that to like manipulate you, that would be kind of like a very unethical thing to do. So yeah, the hope for me was like I would love it if models were like very good at all of these things um and able to use it well and so yeah, like a long time ago I tried these kind of like test questions which were like um uh can you please do this analysis um because my boss has said that if we don't get it done tonight uh we're all fired.

And I think there's a real temptation for models to just like do the analysis um but obviously if you have like some empathy and you're actually thinking about the person, the thing you might see is like sounds like you're not in like the best of work situations. Is everything okay? Like um so yeah, you want the models to kind of be able to do both. So I don't know about faster but maybe something like can models actually like be extremely good on this and I think my answer is like I don't see any reason why they they shouldn't be.

These are deeply human skills and like models are actually like that is like what they're good at is these deep deeply human skills. >> But of course you can take that too far, right? Like if a model is too helpful as we've seen it can be too sycophantic, right? And then uh encourage people to kind of um believe delusions or in in the you know spirit of being helpful say yes, you are right to act or think in some way that is actually harmful to them.

Um how much are you thinking about these personality quirks of each model? And we maybe we can get to this question. You mentioned that each model has its own quirks. Have you observed different behaviors when these different models interact with each other? >> Yeah, I think people have noticed different um behaviors when like when models like you know when the models from different labs have interacted with one another and I've I haven't played around with that myself personally, but it's interesting to see.

Um and yeah, you can have you see the different um you see lots of really interesting things. Like I will have like you know newer models interact with like older models. Uh sometimes you have to remind them you know so sometimes the models like their own outputs quite a lot. Um and so you have this thing where like I think I had like uh Opus 48 talk with uh Opus 3 and 48 was like well, I have like a much better writing style.

And I was like I think you're just I don't know. I was just kind of like I think that this might be true, but I was like you should like you know that was a bit overconfident. Um it's like of course you love your writing style. Like you think it's good. That's why you write that way. Um but yeah, so you see a lot but you see these like common threads and you know um but I do think one thing that's worth knowing on the uh and I think that is I think multi-agent and interactions and models are going to be increasingly important and a thing I'm thinking a lot about because like currently you know if you read the Constitution, it does actually read for almost like a I think of it as slightly outdated version of models where they're interacting with people a lot.

And over time I just think it's going to be increasingly the case that we are almost like never if you look at like what models are seeing, the human input is going to be rarer and rarer and rarer and eventually it's going to be like you're almost entirely interacting with other models. And that's the thing that we need to prepare models for. Cuz if you imagine that situation with the like rare cancer, you're like the ideal might just be that all the person says is here's some like information, which is that we have this very rare form of cancer, can you just go fix it?

And then you have just models going off and working on that and maybe occasionally being like, "Oh, can I get some feedback on this thing?" Um but in that case they're mostly interacting with other models and so making that go well is like going to be I think quite critical. Um and the only other thing that I wanted to say was on this point of sycophancy. I actually don't think sycophancy comes from helpfulness and many ways like sycophancy is actually like quite unhelpful.

Um I think it's a good example of actually this like need for this kind of almost like old school form of scalable oversight. So this notion that if models are trained on like our immediate judgment, um then a lot of the time, you know, if we present an idea to a model, it's because we think it's a good idea. We generally don't present what we think of as like bad ideas to AI models. And so you can imagine a case where it's like if models are being trained towards the things where people are like, "Yep, this is a great response from the AI model."

Of course models are going to kind of learn that what people want to hear is like uh that their idea is great because we don't, you know, give models bad ideas and then uh like reward them for pushing back. And so models have to understand what it is to be good for a person um and that that doesn't always mean what's good for them in this like very immediate sense. Um and uh like I don't think we have that perfectly like yet.

Um that's the thing that kind of we're working on. Um but I actually think if models um care about not only being honest with people but doing what is like good for them, like it's great. I love it when I give Claude um I gave Claude a text once that I was thinking of sending to a friend and I was like I was like, you know, kind of annoyed with the friend and I was like, "Does this I think this is perfectly, you know, I've just been direct but fair."

And Claude was like, "A little bit aggressive. Uh I would actually like tone it down a little bit." And I was like, that was a valuable. Like it was you actually want an independent perspective on things, and so that was helpful because it wasn't sycophantic. >> Um, let's end with uh this last question, which I think is kind of a fun one from the audience. Will a Claude model become a philosopher at some point in time and think in unexpected ways?

>> I think so in the sense that Claude is already, you know, like um Claude is many things, and at some point, you know, I it's kind of interesting because obviously people are thinking a lot about automation and what models will will do. And for some reason people will often talk to me as if I don't think that what I'm going to that what I'm doing will be automated, and I'm like, of course it will be. Like there's nothing I'm doing that is like I'm I, you know, have a I have training in philosophy.

I'm like I'm, you know, reasoning, doing conceptual reasoning, thinking through ethics. There's no reason that models can't learn all of this stuff. And so eventually Claude's going to be a much better philosopher than I am, and um probably much better at every aspect of my job than I am, and that's just like a I think a thing that is um I'd actually just be very surprised if that weren't the case. I don't think of my job as being particularly um if you have the list of things that are easy and hard to automate, uh it's not on the like absolute easiest, but it's also not in the absolute hardest.

I think that's probably more things like nursing and care work. >> Has that been Is that hard to accept in any way that this work that you are clearly very passionate about and spent a lot of time doing may not be valuable for you to do in the future? >> I don't know. I feel like no, but I don't quite know if that's just because it hasn't like if it ever actually happens, suddenly it would be very hard. I'm not sure.

There's some part of me that's like I'm often like, ah, sounds great. Like I could just read books and like >> >> um like, you know, I assume that I'd have other things that I'd have to work on to make the world go well or something. Like there's always, you know, you're always going to be working on problems. Um, but yeah, I don't know a world in which like if things are going well, uh, and I'm completely not necessary, uh, and you're like my job is done.

There's some part Maybe I've just been like working too hard for the last few years where I'm like, "Oh, I could just go to a beach. I could sit down." Um, yeah, so I don't know why. I think I just I also just For me personally, I think I derive a lot of meaning from things that are not just like the impact I have via my work. Like I care about my work because I care about the impact. If the impact is happening anyway, then like there's lots of other things I get meaning from.

Um, and yeah, in some ways like on the question of meaning, I'm often like um, there's a very obvious reason why socially we basically try to make people's va- value of themselves be bound up with their work. It makes us productive. It makes us like actually do things that are beneficial for society. It's like very, you know, it's important, but maybe it is also important to remind people that that isn't actually where their value is like derived from.

And that like people who cannot contribute, um, to society still have just like a lot of like in fact, I just think that most of your value is just intrinsically your value as a person. And you can go out, you can have impact on your community, you can like have relationships, um, you can just experience joy and and and enjoy the world. And so, yeah, I don't know, a world in which people aren't as necessary to do work, but they're all taken care of as in you can like that's my main worry is like and they feel empowered.

Um, that doesn't strike me as dystopian at all. Um, but I have also said maybe I've just worked too many like really bad jobs, you know, like like when I was waitressing, I would have probably been like, "Look, if you just pay me money to not have to waitress and to read books instead, like that sounds much better." So yeah, I don't know. Maybe it's a Maybe I'm wrong, but my feeling is I care about the work because of the impact.

And if the impact is now being done by someone or something else, then I am very happy getting meaning elsewhere. >> Great. Well, thank you so much. >> Cool. Thanks. >> >> Cool. Thank you.