🎙️AI 访谈库
AI 对齐有多难?| Anthropic 研究沙龙(Jan Leike 等)
Jan Leike · Anthropic 对齐科学负责人

AI 对齐有多难?| Anthropic 研究沙龙(Jan Leike 等)

How difficult is AI alignment? | Anthropic Research Salon

2025-01-08 · Anthropic Research Salon · 28m · 约 30 分钟读完 · 原文
Anthropic 官方研究沙龙:Jan Leike 与 Amanda Askell、Alex Tamkin、Josh Batson 四位研究员讨论对齐科学、可解释性与 AI 研究的未来。

we are super excited to have everyone here um folks that uh we've met already folks that are new this panel is just going to be really uh casual and we have researchers from four different teams at anthropic we have folks from societal impacts that's me folks from alignment science that's Yan alignment fine tuning Amanda and interpretability Josh I'm going to start with uh asking a question to Amanda uh from the alignment fine tuning team um and I want you to talk a little bit about how you see uh alignment what it means to you and um you know because you're in charge of a lot of our work on how the model should behave um and why should you be the philosopher king that decides how Claude behaves uh what its characteristics and and and attributes are I mean ask Plato he's the he's the one that decided I should be the philosopher thing um the question of like what is alignment uh maybe this is like a slightly spicy view that I have I think people are very very tempted to spend a lot of time trying to Define this concept because there's like lots of ways of doing it and they you know I don't know they have like social Choice theory in the back of their head and they're like oh well if you imagine everyone has a utility function there's like like limits on exactly what you can say about how to maximize all of those utility functions Etc um and I think I'm more inclined to just be like we kind of want things to go well enough that you can iterate on and improve them later and the bar isn't some like perfect notion of alignment I'm sure there is that concept one can Define it one one can argue about it but for the most part the initial goal is like let's just make things like go well and like meet a certain kind of like lower lower bar um where that is like you know if it's not perfect if some people like don't like it you can just improve on it um so like my view of of alignment is actually probably like I mostly want to like hit that like um and iterate from there uh in terms of like how the model should behave and how I think about that I think I've spoken about this before but my basic concept right now for the model is trying to get it to behave the way that I think like you're very good like morally motivated like kind human would act in if they found themselves roughly in this circumstance it's a little bit strange because they also have to find themselves in the circumstance of like being an EI who is like talking to millions of people um which doesn't fact affect how you behave like maybe you would normally be willing to like just chitchat about politics with someone but if you're going to be talking with millions of people maybe you'd actually be like hm I should maybe be like a little bit more concerned about potentially influencing people um and so but I do think it's actually an important model which is like sometimes people are kind of like oh what values should you put into the model um and I think I'm often like well do we think of this way with humans where I'm just like someone just injected me with like value serum or something and I just have these like fixed things I'm like completely certain of U I'm like I don't know that seems like almost like dangerous or something um most of us just have like a mix of like things that we do value but we would trade off against other things um a lot of uncertainty about different like moral Frameworks we hit cases where we're suddenly like oh actually my value framework doesn't Accord with my intuitions and we update um I think my view is that ethics is actually like a lot more like physics than people think um it's actually like a lot more kind of like empirical and uh something that we're uncertain over and that we have hypothesis about and I kind of want if I'm just like I think that I if I met someone who was just completely confident in their moral view there is no such moral view I could give that person that would not make me kind of terrified whereas if I instead have someone who's just like I don't know I'm kind of Uncertain over this and I just like update in response to like new information about ethics and I like think through these things um that's the kind of person that feels like less scary to me so at least at the moment I'm not this isn't a claim that this is somehow going to like completely like a models or anything um but that's the kind of immediate goal I've realized I've talked a lot and you asked about the philosopher king question I guess okay I guess I should in fact give a quick answer to that um there is also this question of like well in the kind of like should you put values into the model maybe I've partially answered it where I'm like like the models should just be uncertain over values that exist in the world and so ideally it's not just someone injecting their values or their preferences nor nor is it something like everyone just voting on values to put into models but instead like people um that are uncertain and responsive to these things models should also be like that so maybe that's kind of my view okay we're going to come back to that and I'm G to ask Yan uh why is Amanda's view completely wrong and why is this not enough to align models as they get you know she didn't say that but you know we're we're playing up the uh the tension between the bets so um yeah imagine if everyone was like kind uh human that was trying to you know act morally um I think what Amanda is doing is very practical right like can we just like make the models more well behave now and like um like where would we go with this like if a is doing more and more complicated things um right now like if if Amanda does this character work and then she reads a lot of transcripts and you're like okay this is like I like this this model is behaving morally I'm just picturing this is what you're doing um judging it um yeah but like what do we do when the model is doing really complex things and it's just like an agent in the world it's like doing these like really long trajectories um it's doing stuff that we don't understand like bio like doing some bio research and we're like is this dangerous like I don't know um so that's the challenge I'm really interested in like the super alignment problem how do we solve that how we like basically scale this Beyond like things that we can look at if we can look at it we just do some RF it's great or you do some constitutional AI but how do we know that our constitution is actually getting the the model to do that the right thing that we actually want so I think that's that's the big question in my mind do we get to respond yes you can respond you're only allowed to disagree okay I I don't really disagree though uh and I can't lie so um need to up the disagreement feature you I'm usually so disagreeable as well it's just terrible um this is what philosophy taught me was how to be disagreeable I guess like my thought is that the way I mean I think of my work is doing several things but one of them is being kind of iterative towards alignment so in a lot of cases you're actually trying to get the model to kind of oversee its own you know it's not like me my eyes can't look at that many transcripts or something but I can get models to like look at these things and like if alignment is iterative I think My worry is that if people neglect the kind of like the ground and just think ah you can just have like a pretty bad model um and it's just going to help you with these things I'm like I would kind of rather that you had the most aligned model trying to like then help you in the future and like that kind of iterative work but but how do you iterate when you don't you can't read the transcripts anymore and you have to rely on the aligned model but then how do you know that it's actually trying to help you yeah so like in the current cases it's sort of like everything that you are using to verify that the base model is aligned is like the C is like what you are relying on with uh like making sure that like another model that's like trained by that model is itself aligned and I think this is like fine when models are like less capable but to scale it to something with much more capable models you'd actually have to have like a greater ability to verify that yeah what do you do then oh you just want my plan it may just be that it's just all fine and they just supervise themselves and they're all really nice you know that would be I don't want to rely on it but that's not my actual plan but you know I'll defend it for these purposes and and one of our one of our bets you know to in order to like you know guard against the case that a a model might be very deeply uh trying to sabotage against this process is interpretability um how do you see like interpretability as a bet you know situated among you know uh the the more straightforward alignment approaches you know is it just as simple as like oh we find the nice feature and we like up the nice feature we find the evil feature and we like drop the evil feature right or is it you know so I feel like everything in AI is like that Meme with the bell curve and like the idiot and then the really sweaty guy talking a lot and then the Jedi who agrees with the idiot and like there is a possibility that like it turns out that the secret to alignment is just turn on the nice feature like for a sufficiently Galaxy brain version of like the nice feature I think that um in some sense though I I'm hoping Hing that interoperability is also like the Jedi version of just like well look at how the model is doing things and check that it's safe um which is like maybe very hard but also if you could do that um would potentially just answer your question um I think that like one of the one of the things that are sort of like comes up in both like the near term and slightly longer term is like you want to understand like why the model did one thing instead of another when you could come up with like plausible alternative explanations and one way is to ask it but the issue is the models are so analogous to people that they'll just give you an answer to why they did that you know as anybody would um but like how do you trust the any of that stuff and it's like well if you could like look inside and just like see what it was thinking about as it was giving you the answer and like even now with stuff like the saes you can like see there's some feature active and like when else is that happening and it's like okay it's like other instances of people telling White Lies and you're like well then maybe the model is tell it's like that's on the Jedi side I think right um and and so I think that like trying to just like look inside see if we can figure out what the parts are and then see like does are you comfortable with using that part when it does other other things is the basic the basic bet I'm I'm I'm I have a question yeah how do you how do you know you're turning up the nice feature and not the pretend to be nice feature whenever humans are looking feature yeah I mean like I think if it's on the control side right like how do you know and I will say actually many of the features are like a little bit like deceptive also when you just like look at what I thought the style impact team did some great work here with like ENT and deep where like you look at it um and you think it's a feature which is like um you know like age discrimination bad but like actually it's the age discrimination good feature something actually no I think vice versa so you tried to turn it down but actually the opposite Behavior so it can like be hard to hard to understand like all of the cases I will say that some of like the circuits work where you're like okay well how did this get generated gives you some clue about the scenarios like is it looking for a person in the context I also think we're going to need model supervision as well like hopefully an impartial model that isn't like that's a little more scary depending on pre-training has like what do you call it incepted in all of the models they want to evade detection but like sometimes you just like look at enough examples and it becomes clear um and you don't need 10 you need thousands but like Claude is very diligent Yan I'd love to hear a little bit more about like what you see as some of the like some of the things you're struggling with or thinking about when trying to um you know approach this problem of like how do you if you can't read the transcript like what the heck are you you know what the heck are you doing anymore if you're not if you can't provide any meaningful like alignment signal I I mean I think a really obvious thing we should do more of is like what Amanda said just say can we just get the models to help us and then the question is of course how do we trust the models like how do we boobs drop this whole process um and like you know you could hope maybe we can like Leverage the dumare models that we trust more um but they might also not be able to figure it out um and so I guess like there's the whole scaleable oversight work where um you know we have various we exploring various like multi-agent Dynamics to try to train models to you know help us figure out these kind of problems um it seems like overall right like these problems might like maybe the problems are like all kind of easy and like we can just do the Amanda thing and just like mix in some data um or it's like really hard and we have to figure out like fully new ideas and approaches that we don't know yet um I think kind of our best bet uh in the medium term is to try to figure out how to automate alignment research and then we can hopefully um get the models to do it so now we've reduced the problem like to how can we trust this model to do anything to just like well can we just trust it to this like much more narrow thing of like do some ml research which we understand like reasonably well um and how can we evaluate it or like how can we give it feedback on those kind of things and what do you what do you oh yeah go I was just going to say I think we're in this special Zone and I'm terrified about what happens next but we're in the special zone right now we like there's something happens on the forward path but a lot of the information you need is passed back through with the tokens it generates it's just like the Chain of Thought is really important for the models to be very smart and like the Chain of Thought is currently in English and so then you've got like a factorize the problem which is like is the Chain of Thought like reasonably safe and like is that faithful to what's happening on like one past and like maybe like you can do some interpretability to like check that piece and then you can just like you or models can inspect it and you get the other piece and the horrifying moment is like when all of that the very very long thing like isn't in English right it's in like some inscrutable thing that like you've learned through like crazy long RL to do it and like I think that a big challenge is going to be Crossing that Gap where like like none of the intermediates are intelligible and there's like massive amounts of compute before it like drops out and something that people can read maybe something i' be interested in getting like people's mental models on is like what do you think are some signs that like we'd be in an alignment is easy world and what do you think are some like Signs we might see in the next few years that like we're actually in a oh alignment is like really hard world I feel like the model organisms work is like trying to figure this out right like can we try to deliberately make deceptive models or misaligned models and models that like try to do Shady stuff like how good are they how hard is it to do it um I mean we might fundamentally go around like about it the wrong way and that's why we fail but if we do succeed it should tell us like how close are we to that kind of World um and then you know once you have your deceptive model does all these Shady things like can you fix it what if you don't know what the model if it's a shady model or not like we playing these interminability audits which I'm very excited about but I don't actually know what the state is we haven't done the audit yet but your people are trying to make it and we're trying to catch it yeah so we making some shady models and then they have to figure out which one is the Shady one right yeah and in what way is it in what way is it Shady CU most of it is still fine I mean I think one one thing that the interpretability stuff has been interesting is so when you do these like unsupervised things you're like okay here's a million different like features a bunch of those correspond to personas the model could inhabit any of those which include all sorts of like deceptive behaviors right that are out there in the world right the model like knows about bad people and like bad motivations and so like the fact that that capability exists is just going to be baked in and then the question is like is it doing that stuff or is it doing the good stuff especially like this question of like when aanda goes and shapes the base model right which isn't really anything right into like an agent or an actor which is supposed to embody some of these and not others like can we tell exactly what it picked up there right from the finite data set right which is being used to like do that shaping um and like yeah so that's a that's a question I mean I think that like some interpretability have maybe some influence function I mean there's different ideas there to ask like what did it get from this process but like that shaping process is going to be really really important yeah and maybe a sign that seems important is something like how robust or like like if you have like modal organisms work and then turns out you just put it through some character training and it just comes out being really nice again then I'd be like okay that's a good sign re the kind of world that we're in um or is it like just like is just a kind of like shallow like shell on top of like um I don't know the same behavior then I'm like okay okay we're in a slightly harder world only slightly harder how how do you distinguish like whether it's shallowly aligned or deeply aligned I think so I mean I feel like there's a lot you could do because I think interpretability is kind of like one of them there's other things that feel so in terms of like modal organisms I guess my hope would be you'd actually have a kind of like red team blue team setup where you have a way of detecting whether the behavior that you have like instilled in a model is like still there um and my job is actually to not know H and in fact it would be really good if I'm trying to train the model I actually just don't know what it is that you've done um because that's like a better way of testing whether my intervention is actually working whereas because otherwise I it's just so hard to not like just try to train to the test or something so I think that's I almost want to be completely ignorant of it yeah maybe we should play an alignment finding game yeah yeah he misalign it and then you align it and we see who wins yeah no I've said to people before don't don't tell me how you did this because I want to see if can fix her Sleeper Agent possibly we could play that game I think I might need to know less uh make another one make another one that's worse that is an excellent transition and to questions um so we are going to uh pass a mic around Ain uh is kindly going to do that um please raise your hand if you have a question for any of us um oh perfect um hello uh I have a question about um well everything you've been talking uh so when we talk about alignment we're talking about like a singular forward pass typically right so like uh alignment at inference time so if I'm using one of the models um via the API and I'm building my own sense of say cultural alignment and I've got a bunch of different Agents set up talking to each other trying to deliberate with that sense of kind of inner conflict that Amanda referenced as useful in in what we do as humans to align like we go back and forth we think you know we go through various cognitions so if I'm trying to create this kind of multi- agentic deliberation thing but I butt up against this aligned model who is so uh unwan to deliberate with other spawns of its own with its own self because all of them are like I'm sorry I can't talk about that and so you just get this endless loop did you have commentary on on on that cuz many people we're not all using Claude in a single inference forward passway so yeah I hope someone gets the gist of what I'm on about yeah I need to think about it so I guess like is the thought like to be clear I'm not thinking necessarily that you have many agents um in the same way that we can be singular agents but still like deliberative and I actually think that like there's a sense in which the more fractured an agent is the more worried am because from like an interpretability perspective and also even just like a predicting what that agent is going to do perspective they're more unpredictable um so I guess I'm thinking in the same way that like maybe I guess a different way of asking this question is just do humans have to do this like my sense is that we're often just very willing to like reflect on many things we go back and forth we come to like conclusions and so in the same way that like a human would think through any kind of standard problem uh that they were like facing I'm imagining that like moral deliberation in the model is just going to look kind of like that more like the deliberation of a single model than like uh multiple models weighing in if that makes sense one more question here I I want to try to draw maybe a relatively strange parallel between like Hana AR's work on the banality of evil the idea that most humans tend not to be evil but when put in certain situations the coupling constant between humans being so huge the evil comes as an epop phenomenon of the system right and so what what occurs to me is as you talk about model alignment you're focus on one model is most of your comments how do you think about the coupling not only with Society but as you work on agents and Mill potentially millions and millions of agents that sort of epiphenomenon of those systems a question um I can talk a a little bit about that I mean I think broadly when you have to think about when you think about safety and Alignment you have to think about it from a systems standpoint you can't just think about it from you know an individual uh models perspective and isolation and I think we've seen a lot of work where you know a lot of jailbreaks operate by pitting different values sort of against each other putting the model in a difficult situation where you know it's designed to elicit what would ordinarily be a harmful Behavior but which uh you know in the context of the question the model thinks sort of is is the right thing to do um and so uh you know there's a variety of tools you can use to do that I mean for one you can include a lot of those you know uh systems level um Integrations in the training process right and give them all exposure to a wider variety of situations for it to you know the uh the broad umex in it's answering questions now that leads to other uh challenges and and and other sort of like um you know uh fall-off issues um with the model is reasoning about you know the the um uh effect of its actions but I I think I agree with the point that you can't just consider a model sort of in isolation in some ways like this is a thing that I've thought about in the context of like this notion of corrigibility or some like you know so like models that just are responsive to like what humans want versus models that are like uh like have values in a sense and are maybe like willing to actually be a little bit uncor um and like the banality of evil Point feels especially relevant if you're like thinking of models as just like doing whatever humans say because in in some ways like that is the idea that if you have a society that like either collectively just like allows for harmful things to take place or even endorses them and you have models this isn't necessarily people misusing models it's just that you'd be using models to like facilitate some kind of like harmful activity um and so I think there is like fundamentally actually a tension between having models be corable to the very least individual humans and having them be like uh aligned in in a sense like aligned with like all humans um and it's really important to recognize that tension and I think when people don't they think like H it's a failure that the model like didn't do what I said but I'm like there's a limited sense in which models I think models should be more corrigible to like Humanity then then not and should be willing to like you know like so there's like when push comes to shove like um but that doesn't necessarily mean being corrigible with each person um because I think that could lead to that kind of like situation that you mentioned hi um so it seems like we've got Yan who's working on intent alignment making sure the models do what we ask them to do Amanda working on values alignment making sure the models are like reasonable kind entities and then Josh working on interpretability which is like allowing us to verif ify that the techniques for those other things are in fact doing what we want um if you were all to succeed in your areas would that be a complete solution to AI safety um or are there pieces missing and if so what are they there's also like a lot of people working on topic who are not at this panel so we are oversimplifying a little bit Yeah and and oh I was just going to say we also you know have the societal impacts team which thinks about the model's impact on societ uh you know writ large and and so I think uh yeah we we you know you could have the most perfectly aligned model but uh aligned to what who's using it in what purposes you know for what purposes um you know and and I think the broader societal context is super something we we're very attentive to also if you're just interested in just more lists like we already mentioned the model organisms work there's like you know jailbreaking robustness there's control there's like trust and safety there's like um yeah there there's a lot of other efforts that I think are also going to be important can I I want to add an almost like I don't know if this is pessimistic or or just I don't think it's pessimistic I think that like there's also this way of talking about alignment and the alignment problem as like a sing it's not like as like a single like theoretical problem or something like that where people will be like does this solve it and somehow it's never felt right it feels a little bit like you know I don't know does like it yeah it feels more like to my mind I'm just like look problems might just arise that we're not even thinking of now and in fact that's like very very common um in like many many disciplines and I kind of expect it to be true here and it would be really dangerous I think if we were just like oh yeah we've like solved this problem um because I'm just like I don't know it could be that the actual problem is one that we've just not thought of yet yeah unkown unknowns are but we should solve the problem and then once we did we should say that we solved it another question here Yan was talking earlier about using Dumber models to evaluate smarter models time and I'm wondering to what extent you see like the grocking abilities in models where like Suddenly It's really duplicitous or you sort of see oh it's lying but it's very bad at lying and now I can catch that and maybe like nip that in the bud while it's still weak yeah I mean there's like plenty of examples like one that I remember is like gbd uh 4 could just read and write in B 64 Super reliably and 3.

5 could not and so if you use 3.