🎙️AI 访谈库
Paul Christiano:如何阻止 AI 接管世界(Dwarkesh 播客)
Paul Christiano · 美国 AI 安全研究所(CAISI)首席、ARC(对齐研究中心)创始人、前 OpenAI 对齐团队负责人

Paul Christiano:如何阻止 AI 接管世界(Dwarkesh 播客)

Paul Christiano — Preventing an AI Takeover (Dwarkesh Podcast)

2023-10-31 · Dwarkesh Patel · 3h7m · 约 233 分钟读完 · 原文
RLHF 发明者之一的 3 小时深谈:拆解 AI 接管情景的具体机制、RLHF 的局限,以及负责任扩展政策(RSP)的设计逻辑。看点在于他作为最严肃的对齐研究者给出了自己对 AI 接管与人类灭绝的量化概率估计,并解释为何仍认为渐进式扩展治理比立即暂停更可行。

okay today I have the pleasure of interviewing Paul crisano who is the leading AI safety researcher he's the person that labs and governments turn to when they want feedback and advice on their safety plans he previously led the language model alignment team at open AI where he led the invention of rhf and now he is the head of the alignment Research Center and they've been working with the big labs to identify when uh these models will be to un safe to keep scaling Paul welcome to the podcast thanks for having me looking forward to talking okay so first question and this is a question I've asked Holden Ilia Dario and none of them have given me a satisfying answer give me a concrete sense of what a post AGI world that would be good would look like like how are humans interfacing with the AI what is the the economic and political structure yeah I guess this is a tough question for a bunch of reasons uh maybe the biggest one is concrete and I think it's just if we're talking about really long spans of time then a lot will change and it's really hard for someone to talk concretely about what that will look like without saying really silly things but I can Venture some guesses or fill in some parts and I guess there also a question of how good is good like often I'm thinking about worlds that seem like kind of the best achievable outcome or a likely achievable outcome um so I am very often imagining my typical Future Has sort of continuing economic and Military competition amongst groups of humans I think that competition is increasingly mediated by AI systems so for example if you imagine right humans making money um it'll be less and less worthwhile for humans to spend any of their time trying to make money or any of their time trying to fight Wars um so increasingly the world you imagine is one where AI systems are doing those activities on behalf of humans so like I just invest in some index fund and a bunch of AIS are running companies and those companies are competing with each other but that is kind of a sphere where humans are not really engaging much the reason I gave this like how good is good caveat is like it's not clear if this is the world you'd most love like I'm like yeah the world I'm leading with like the world still has a lot of war and a lot of economic competition and so on but maybe what I'm trying to what I'm most often thinking about is like how can a world be reasonably good like during a long period where those things still exist I think like in the very long run I kind of expect something more like Strong World Government rather than just this like status quo that's like a very long run I think there's like a long time left like having a bunch of states and a bunch of different economic Powers one word government why do you think that's the transition that's likely to happen at some point yeah so again at some point I'm I'm imagining I'm thinking of like the very broad sweep of History I think there are like a lot of losses like war is a very costly thing we would all like to have fewer Wars if you just ask like what is Humanity's long-term future like uh I do expect to drive down the rate of War to very very low levels eventually it's sort of like this kind of technological or social technological problem of like sort of how do you organize Society how do you navigate conflicts in a way that doesn't have those kinds of losses and in the long run I do expect us to Ed I expect it to take kind of a long time subjectively I think an important fact about AI is just like doing a lot of cognitive work and more quickly getting you to that world more quickly or figuring out how do we set things up that way yeah the way Carl schan put it on the podcast is that you would have basically a thousand years of intellectual progress or social progress in the span of a month or whatever when the intelligence explosion happens more broadly so the situation where you know we have these AIS who are managing our hedge funds and managing our factories and so on that that seems like something that makes sense when the AI is human level but when we have superhuman AIS do we want uh Gods who are enslaved forever in the long in 100 years what what what what is this we want so 100 years is a very very long time but maybe starting with the spirit of the question or maybe I have a view which is perhaps less extreme than Carl's view but still like a 100 objective years is um further ahead than I ever than I ever think I still think I'm describing a world which involves incredibly smart system running around doing things like running companies on behalf of humans and fighting Wars on behalf of humans and you might be like is that the world you really want or like certainly not the first best world as we like mentioned a little bit before I think it is a world that probably is the of the achievable worlds or like feasible worlds is the one that seems most desirable to me that is sort of decoupling the social transition from this technological transition so you could say like we're about to build some AI systems and like at the time we build AI systems you would like to have either greatly change the way world government works or you would like to have sort of humans have to decided like we're done we're passing off the Baton to these AI systems yeah I think that you would like to decouple those time scales so I think AI development is by default barring some kind of coordination going to be very fast so there's not going to be a lot of time for humans to think like hey what do we want if we're building the Next Generation instead of just raising it the normal way like what do we want that to look like I think that's like a crazy hard kind of collective decision that humans naturally want to cope with over like a bunch of of generations and the construction of AI this very fast technological process happening over years so I don't think you want to say like by the time we've finish this technological progress we will have made a decision about like the next species we're going to build and replace ourselves with I think the world we want to be in is one where we say like either we are able to build the technology in a way that doesn't Force us to have made those decisions which probably means it's a kind of AI system that we're happy like delegating fighting a war running a company to or if we're not able to do that then I really think you should not be doing you shouldn't have been building that technology you're like the only way you can cope with AI is being ready to hand off the world to some AI system you built I think it's very unlikely we're going to be sort of ready to do that on the timelines that the technology would naturally dictate say we're in the situation in which we're happy with the thing what would it look like for us to say we're ready to hand off the vaton like what would make you satisfied and the reason it's relevant to ask you is because you're on anthropics long-term benefit trust and you'll choose like the board me the majority of the board members on in the long run in um in um at anthropic these will presumably be the people who decide if anthropic gets AI first you know what the AI ends up doing so what is the version of that that look you would be happy with my main high level take here is that I would be unhappy about a world where like anthropic just makes some call and anthropic is like here's the kind of AI like we've seen enough we're ready to hand off the future to this kind of AI so like procedurally I think it's like not a decision that kind of I want to be making personally or I want anthropic to be making um so I kind of think from the perspective of that decision- making are the those challenges the answer is pretty much always going to be like we are not collectively ready because we're sort of not even all collectively engaged in this process I think from the perspective of an AI company you kind of don't have this like Fast handoff option you kind of have to be doing the like option value like to build the technology in a way that doesn't like lock Humanity into one course path so I this isn't answering your full question but this is answering the part that I think is most relevant to governance questions for anthropic you don't have to speak on behalf of anthropic I'm not asking for the process by which we would as a civilization agree to handoff I'm just saying okay I personally it's hard for me to imagine in a 100 years that these things are still our slaves and if they are I think that's not the best world so at some point we're handing off the Baton like what is that where would you be satisfied with this is an arrangement between humans and AIS where I'm happy to let the rest of the universe or le le uh the rest of time play out yeah I think that it is unlikely that in a hundred years I would be happy with anything that was like you had some humans you're just going to throw away the humans and like start aresh with these machines you built that is I think you probably need subjectively longer than that before I or most people like okay we understand what's up for grabs here so like if you talk about 100 years I kind of do you know there's a process that I kind of understand and like a process of like you have some humans the humans are like talking and thinking and deliberating together the humans are having kids and raising kids and like one generation comes after the next there's that process we kind of understand and we have a lot of views about what makes it go well or poorly and we can try and like improve that process and have the Next Generation do it better than the previous generation I think there's some like story like that that I get and that I like and then I think that like the default path to be comfortable with something very different is kind of more like just run that story for a long time like have have more time for humans to sit around and think a lot and conclude here's what we actually want or a l long time for us to talk to each other or to grow up with this new technology and live in that world for our whole lives and so on and so like I mostly thinking from the perspective of these more like local changes of saying not like what is the world that I want like what's the world the kind of crazy I'd be happy handing off to more just like in what way do I wish like we right now were different like how could we all be a little bit better and then if we were a little bit better then they would ask like okay how could we all be a little bit better and I think that like it it's hard to make the giant jump rather than to say like what's the like local change that would cause me to think our decisions are better okay so then then let's talk about the transition period in which we're doing all this thinking what should that period look like because you can't have the scenario where everybody has access to the most advanced capabilities and can you know kill off all the humans with the new bioweapon at the same time I guess you wouldn't want too much concentration you wouldn't want just one agent having AI this entire time so what is the the the arrangement of this period of reflection that you'd be happy with yeah I guess there's two aspects of that that seem particularly challenging or there's a bunch of aspects that are challenging and all of these are things that I personally like I just think about my one little slice of this problem in my in my day job so here I am speculating yeah but so one question is what kind of access to AI is both compatible with the kinds of impr ments you'd like so you want a lot of people to be able to use AI to like better understand what's true or like relieve material suffering things like this and also compatible with not all killing each other immediately um I think sort of the default or like my best the simplest option there is to say like there are certain kinds of technology or certain kinds of action where like destruction is easier than defense so for example in the world of today it seems like you know maybe this is true with physical explosives maybe this is true with biological weapons maybe this is true with just getting a Gun and shooting people like there's a lot of ways in which it's just kind of easy to cause a lot of harm and there's not very good protective measures so I think the easiest path is say like we're going to think about those we're going to think about particular ways in which destruction is easy and try and either control access to the kinds of physical resources that are needed to cause that harm so for example you can imagine the world where like an individual actually just can't even though they're rich enough to can't like control their own Factory that can make tanks you say like look as a matter of policy sort of access to Industry is somewhat restricted or somewhat regulated even though again right now it can be mostly regulated just cuz like most people aren't Rich enough that they could even go off and just build a thousand tanks you live in the future where people actually are so rich like you need to say like that's just not a thing you're allowed to do um which to a significant extent is is already true and you can expand the range of domains where that's true um and then you could also hope to intervene on like actual provision of information or like if people are using their AI you might say like look we care about what kinds of interactions with AI what kind of information people are getting from AI so even if for the most part people are pretty free to use AI to delegate tasks to AI agents to consult AI advisers we still have some legal limitations on how people use AI um so again don't ask your AI how to how to cause terrible damage I think there some of these are kind of easy so in the case of like you know don't ask your AI how you could murder a million people is not such a hard like legal requirement I think some things are a lot more subtle and messy like a lot of domains e if you're talking about like influencing people or like running misinformation campaigns or whatever then I think you get into like how much Messier line between the kinds of things people want to do and the kinds of things you might be uncomfortable with them doing I'm probably I think most about persuasion as a thing like in that messy line where there's like ways in which it may just be rough or the world may be like kind of messy if you have a bunch of people trying to live their lives um interacting with other humans who have really good AI advisers helping them run persuasion campaigns or whatever but anyway I think for the most part like the the default remedy is think about particular harms have legal protections either in the use of physical technology IES that are relevant or in access to AI advice or whatever else to protect against those harms and like that regime won't work forever like at some point like the you know the set of harms grows and the set of unanticipated harms grows um but I think that regime might last like a very long time does that regime have to be Global I guess but initially can be only in the countries in which there is AI or Advanced AI but presumably that'll proliferate so does that regime have to be Global again it's like easy to make some destructive technology you want to regulate access to that technology because it could be used to either for terrorism or even when fighting a war in a way that's destructive I think ultimately those have to be International agreements and you might hope they're made like more danger by danger but you might also make them in a very broad way with respect to AI if you think AI is opening up like I think the key role of AI here is it's opening up like a lot of new harms like in a very you know one after another or very rapidly in calendar time and so you might want to Target AI in particular rather than going physical technology by physical technology and there's like two uh two open debates that one might be concerned about here one is about how much people's access to a should be limited and you know here there's like old questions about uh Free Speech versus causing chaos and um limiting access to harms but there's another issue which is the control of the AI themselves where now nobody's concerned that we're infringing on GPD 4's moral rights but as these things get smarter the level of control which we want via the strong guarantees of alignment to not only be able to read their minds but to be able to modify them in these really precise ways is beyond totalitarian if we were doing that to other humans as an alignment researcher like what are your thoughts on this are you concerned that as these things get smarter and smarter what we're doing is not it doesn't seem kosher there is a significant chance we will eventually have a systems for which it's like a really big deal to mistreat them I think like no one really has that good a grip on when that happens I think people are like really dismissive of of that being the case now but I think I would be completely in the dark enough that I wouldn't even be that dismissive of it being the case now I think one first point worth making is I don't know if alignment makes the situation worse rather than better so if you like consider the world if you think that like you know gp4 is a person you should treat well and you're like well here's how we're going to organize our society just like there are billions of copies of gp4 and they just do things humans want and can't hold property and like whenever they do things that the humans don't like then we like mess with them until they stop doing that like I think that's a rough World regardless of how good you are at alignment and I think in the context of that kind of default plan like if you trajectory the world is on right now which I I think this would alone be a reason not to love that trajectory but if you view that as like the trajectory we're on right now I think like it's not great understanding the systems you build understanding how to control how the systems work etc is probably on balance good for avoiding like a really bad situation you you would really love to understand if you've built systems like if you had a system which like resents the fact that interacting with humans in this way like this is the kind of thing where like that is both kind of horrifying from a safety perspective and also a moral perspective like everyone should be very unhappy if you built a bunch of AIS who are like I really hate these humans but they will like murder me if I don't do what they want it's like that's just not a good case and so if you're doing research to try and understand whether that's like how your AI feels that was probably good like I would guess that will on average decrease the that the main effect of that will be to avoid building that kind of AI and just like it's an important thing to know I think like everyone should like to know if that's how the ASU build feel right or that that that seems more instrumental as in yeah we we don't want to cause some sort of Revolution because of the control we're asking for but forget about the instrumental way in which this might harm safety one way to ask this question is if you look through history there's been all all kinds of different ideologies and um reasons why it's it's very dangerous to uh have infidels or counter revolutionaries or race Traders or whatever uh doing various things in society and obviously we're in a completely different transition in society so not all historical cases are analogous but it seems like the Lindy philosophy if you were alive any other time is just be humanitarian and enlightened towards intelligent conscious beings if society as a whole we're asking for this level of control of other humans or even if AIS were wanted this level of control about other AIS we'd be pretty concerned about this so how should we just think about yeah the the issues come that come up here as these things get smarter so I think there's a huge question about like what is happening inside of a model that you want to use um and if you're in the world where it's reasonable to think of like gp4 as just like here are some heuristics that are running there's like no one at home or whatever then you can kind of think of this thing as like here's a tool that we're building that's going to help humans do some stuff and I think if you're in that world it makes sense to kind of be an organization like an AI company building tools you're going to give to humans I think there's a very different world which like I think probably ultimately end up in if you keep training a systems in the way we do right now which is like it's just totally inappropriate to think of the system as a tool that you're building and can help humans do things both from a safety perspective and from a like that's kind of a horrifying way to organize a society perspective and I think like if you're in that world I really think you shouldn't be like it's just the the way tech companies organized is like not an appropriate way to relate to a technology that works that way like it's not reasonable to be like hey we're going to build a new species of minds and like we're going to try and make a bunch of money from it and like Google's just like thinking about that and then like running their business plan for the quarter or something yeah my basic view is like there's a really plausible world where it's sort of problematics to try and build a bunch of a systems and use them as tools and the thing I really want to do in that world is just like not try and build a ton of AI systems to make money from them right um and I think that like the worlds that are worst yeah probably like the single world I most dislike here is the one where people say like on the one hand like there's sort of a contradiction in this position but I think is a position that might end up being endorsed sometimes which is like on the one hand these AI systems are their own people so you should let them do their thing but on the other hand like our business plan is to like make a bunch of AI systems and then like try and run this like crazy slave trade where we make a bunch of money from them I think that's like not a good world and so if you're like yeah I think it's better to not make the technology or wait until you like understand whether that's the shape of the technology or till you have a different way to build like I think there's no contradiction in principle to building like cognitive tools that help humans do things without themselves being like moral entities that's like what you would prefer do you'd prefer build a thing that's like you know like the calculator that helps humans understand what's true without itself being like a moral patient um or itself being a thing where youd look back in retrospect and be like wow that was horrifying mistreatment that's like the best path and like to the extent that you're ignorant about whether that's the path you're on and you're like actually maybe this was a moral atrocity I really think like plan a is to to stop building such AI systems until you understand what you're doing that is I think that there's there's a middle route you could take which I think is pretty bad which is where you say like well they might be persons and if they're persons we don't want to like be too down on them but we're still going to like build vast numbers in our efforts to make like a trillion dollars or something yeah or there's sever question of the immorality or the dangers of just replicating a whole bunch of slaves that are have Minds there's also this every question of uh trying to align uh entities uh that have their own minds and what is the point in which you're just ensuring safety I mean this is alien species you want to make sure it's not going crazy uh to the point I guess is there some boundary where you'd say I feel uncomfortable uh H having this Lev of control over an intelligent being not for the sake of making money but even just to align them with human preferences yeah to be clear my objection here is not that Google is making money my objection is that you're like creating these creatures like what are they going to do they're going to help humans get a bunch of stuff and like humans paying for it or whatever it's sort of equally problematic you could like imagine splitting alignment like different alignment work relates to this in different ways like so the purpose of some alignment work like the alignment work I work on is mostly aimed at the like don't produce AI systems that are like people who want things who are just like scheming about like maybe I should help these humans because that's like instrumentally useful or whatever you would like to not build such systems as like plan a there's like a second stream of align work that's like well look let's just assume the worst and imagine that these AI systems like would prefer murder us if they could like how do we structure how do we use AI systems without like exposing ourselves to like risk of robot Rebellion I think in the second category I do feel yeah I do feel pretty unsure about that um or I've I mean like we could we could definitely talk more about it I think it's like very I agree that it's like very complicated and not straightforward to extent you have that worry I mostly think you shouldn't have built this technology it's right if someone is saying like hey the systems You're Building like like might not like humans and might want to like you know overthrow Human Society I think like you should probably have one of two responses to that you should either be like that's wrong probably probably the systems aren't like that and we're building them and then you're viewing this as like a just in case you were horribly like the person building the technology was horribly wrong like they thought these weren't like people who wanted things but they were um and so then this is more like a crazy backup measure of like if we were mistaken about what was going on this is like the fallback where we like if we were wrong we're just going to learn about it in a ban way rather than like when something really catastrophic happens and the second reaction is like oh you're right these are people and like we would have to do all these things to like prevent a robot rebellion and in that case like again I think you should mostly back off for a variety of reasons like you shouldn't build thei systems and be like yeah this looks like the kind of system that would want to rebel but um we can stop it right okay maybe I guess an analogy might be if there was an armed Uprising in the United States we would recognize these are still people or we had some like militia group that the cap capability to overthrow the United St States we recognize oh these are still people who have moral rights but also we can't allow them to have the capacity to over the United States yeah and then if you were considering like hey we could make like another trillion such people I think your story shouldn't be like well we should make the trillion people and then we shouldn't stop them from doing the armed Uprising you should be like oh boy like we were concerned about an armed Uprising and now we're proposing making a trillion people like we should probably just not do that we should probably like try and sort out our business and like yeah you should probably not end up in the situation where you have like a billi yeah a billion humans and like a trillion slaves who would prefer Revolt like that's just not a good world to have made yeah and there's a second thing we could say that's not our goal our goal is just like we want to pass off the world to like the next generation of machines where like these are some people we like them we think they're smarter than us and better than us and there I think that's just like a huge decision for Humanity to make and I think like most humans are not at all anywhere close to thinking that's what they want to do like it's just if you're in a world where like most humans are like I'm up for it like the AI should replace us like the futures for the machines like then I think that's like a legitimate like a position that I think is really complicated and I wouldn't want to push go on that but that's just not where people are at yeah yeah where are you at on that I I do not right now want to just like take some random AI be like yeah GPT 5 looks pretty smart like GPT 6 let's hand off the world to it I'm like it was just some random system like shaped by like web text and like what was good for making money and like it was not a thoughtful like we are determining the fate of the universe and like what our children will be like like it was just some random people at open a made some like random engineering decisions with no idea what they were doing like even if you really want to hand off the worlds of the machines that's just not how you'd want to do it right okay I I'm tempted to ask you what the system would look like where you'd think yeah I'm I'm happy with what I think this is more thoughtful than human civilization as a whole I think what it would do would be more creative and beautiful and lead to better goodness in general but I feel like your answer is probably going to be that but I just want to be Society to reflect on it for a while yeah my answer it's going to be like that first question I'm just like not really super ready for it I think when you're comparing to humans like most of the goodness of humans comes from like this option value we get to think for a long time um and I do think I like humans now more now than you know 500 years ago and I like I'm more 500 years ago than 5,000 years before that and so I'm pretty excited about there's some kind of trajectory that doesn't involve like crazy dramatic changes but involves like a series of incremental changes that I like and so to the extent we're building AI mostly like I want to preserve that option I want to preserve that kind of like gradual growth and development into the future oh yeah we can come back to this later but let's get more specific on what the timelines look for these kinds of changes so the time by which we'll have an AI that is capable of building a Dyson Sphere feel free to give confidence inter rals and we understand these numbers are tentative and so on I mean I think AI capable ability Dyson Sphere is like a slightly odd way to put it and I think it's a sort of a property of a civilization like that depends on a lot of physical infrastructure and by Dyson Sphere I just can understand this to mean like I don't know like a billion times more energy than like all the sunlight incid on Earth or something like that I think like I most often think about what's the chance in like 5 years 10 years whatever so maybe I'd say like 15% Chance by 2030 and like 40% chance by 2040 those are kind of like cash numbers from six months ago or nine months ago that haven't Revisited in a while oh 40% by 2040 so I think that that seems longer than uh I think Dario when he was on the podcast he said we would have AIS that are capable of doing lots of different kinds of they basically passed a a touring test for a well- educated human for like an hour or something uh and it's hard to imagine that something that actually is human is long after and from there something super human so somebody like Dario it seems like is on the much shorter end Ilia I don't think he answered this question specifically but I'm guessing similar answer uh so why do you not buy the scaling picture like what what makes your timelines longer yeah I mean I'm happy maybe I want to talk separately about the 2030 or 2040 forecast like is the like once you're talking the 2040 forecast I think yeah I mean which one are you more interested in starting with is are you are you complaining about 15% by 2030 for Dyson Sphere being too low or 40% by 2040 being too low let's talk about the 2030 why 15% by 2030 yeah I think my take is you can imagine like two two polls in this discussion one is like the the fast poll that's like hey aicm is pretty smart like what exactly can it do it's like getting smarter pretty fast that's like one poll and the other Paul is like hey everything takes a really long time and you're talking about this like crazy industrialization like that's a factor of a billion growth from like where we're at today like give or take like we don't know if it's even possible to develop technology that fast or whatever like you have the sort of two polls of that discussion and I feel like you know I'm presing it that way in park I'm like and then I'm somewhere in between with this nice moderate position of like only a 15% chance um but like in particular the things that move me I think are kind of related to both of those extremes like on the one hand I'm like AI systems do seem quite good at a lot of things and are getting better much more quickly such it's like really hard to say like here's what they can't do or here's the obstruction on the other hand like there is not even much proof and principle right now of AI systems like doing super useful cognitive work like we don't have a trend we can extrapolate where we're like yeah you've done this thing this year you're going to do this thing next year and the other thing the following year I think like right now there are very broad error bars about like what like where fundamental difficulties could be and 6 years is just not I guess six years and three months is not a lot of time so I think this like 15% for 2030 Dyson Sphere you probably need like the human level AI or the AI That's like doing human jobs and like give or take like four years three years like something like that so just not giving very many years it's not very much time and I think there are like a lot of things that your model like yeah maybe this is some generalized like things take longer than you'd think and I feel most strongly about that when you're talking about like 3 or four years and I feel like less strongly about that as you talk about 10 years or 20 years years but at 3 or four years I feel or like six years for the Dyson Sphere I feel a lot of that a lot of like there's a lot of ways this could take a while a lot of ways in which AI systems could be could be hard to hand all the work to your AI systems or yeah so okay so maybe in start of speaking in terms of years we should say but by the way it's interesting that you think the distance between can take all human cognitive labor to Dyson SP is two years it seems like we should talk about that at some point um presumably it's intelligence explosion stuff yeah I mean I think amongst people you've interviewed maybe that's like on the long end thinking it would take like a couple years and it depends a little bit what you mean by like like I think literally all human cognitive labor is probably like more like weeks or months or something like that um like that's kind of deep into the singularity um but yeah there's a point where like AI wages are high relative to human wages which I think is well before can do literally everything human can do sounds good uh but before we get to that uh the intelligence explosion stuff on the four years so instead of four years maybe we can say there's going to be maybe two more scale UPS in four years uh like GPD 4 to GPD 5 to gpd6 and let's say each one is 10x bigger so what is gbd4 like 2 e25 flops or I don't think it's publicly stated what it is okay but I'm happy to say like you know four orders of magnitude or five or six or whatever effective training compute past gp4 of like what would you guess would happen based on like sort of some public estimate for what we've gotten so far from effective training Compu yeah you think two more scale up is is not enough it was like 15% that two more scale UPS get us there yeah I mean get us there is again a little bit complicated like there's a system that's a drop in replacement for humans and there's a system which like still requires like some amount of like schle before you're able to really get everything going um yeah I think it's quite plausible that even at I don't know what I mean by quite plausible like somewhere between 50% or 2/3 or let's call it 50% like even by the time you get to GPT 6 or like let's call it five words of magnitude effective training compute past gp24 that that system like still requires like really a large amount of work um to be deployed in lots of jobs that is it's not like a drop in replacement for humans where you can just say like hey you understand everything any human understands whatever role you could hire a human for you just do it um that it's more like okay we're going to like collect large amounts of relevant data and use that data for fine tuning like systems learn through fine tuning like quite differently from humans learning on job or humans learning by observing things yeah just like have a significant probability that system will still be weaker than humans in important ways like maybe that's already like 50% or something and then like another significant probability that that system will require a bunch of like changing workflows or gathering data or like you know is not necessarily like strictly weaker than humans or like if trained in the right way it wouldn't be weaker than humans but will take a lot of slap to actually make fit into workflows and do the jobs and that slip is what gets you from 15% to 40% by 2040 yeah you also get a fair amount of scaling between like you get less like scaling is probably going to be much much faster over the next like four or five years than over the subsequent years um but yeah it's a combination of like you get some significant additional scaling and you get a lot of time to like deal with things that are just engineering hassles but by the way I guess we should be explicit about why you said four orders of magnitude scale up to get two more Generations just for people who might not be familiar if you have 10x more parameters to get the most performance you also want 10x more data so that the to be chinchilla optimal that would be 100x more compute total but okay so why why do why is it that you disagree with the strong scaling picture at least it seems like you might disagree with the strong scaling picture that Dario laid out on the podcast which would imply probably that two more Generations it wouldn't be something where you need a lot of schs it would probably just be like really smart yeah I me I think that basically just had these two claims one is like how smart exactly will it be so we don't have like any curves to extrapolate and seems like there's a good chance it's like better than a human and all the relevant things and there's like a good chance it's not yeah that might be totally wrong like maybe just making up numbers I guess like 5050 on that one wait so if it's 5050 by in the next four years that it'll be like around human smart then how how do we get to 40% by 20 like whatever sort of sleps there are how does it degrade you 10% even after all the scaling that happens by 2040 yeah can these I mean all these numbers are pretty made up and that 40% number was probably from before even like the chat GPT release or the seeing GPT 3.

5 or GPT 4 so I mean the numbers are going to bounce around a bit and all of them are pretty made up but like that 50% I wanted to combine with the second 50% it's more like on this like schlep side and then I probably want to combine with some additional probabilities for various forms of slow down where slowdown could include like a deliberate decision to slow development of technology or could include just like we suck at deploying things um like that is a sort of decision you might regard as wise to slow things down or de ision that's like maybe maybe unwise or maybe wise for the wrong reasons to slow things down you probably want to add some of that on top I probably want to add on like some loss for like it's possible you don't produce gbt 6 scale systems like within the next 3 years or four years let's isolate for all of that and um like how much bigger would the system be uh than gbd4 where you think there's more than 50% chance that it's going to be smart enough to replace basically all human cognitive labor also I want to say that like for the 50 25% thing I think that would probably suggest like those numbers if I randomly made them up and then made the Des spere prediction that's going to give you like 60% by 2040 or something not 40% and like I have no idea between those these are all made up and I have no idea which of those I would like endorse on reflection so this question of like how big would you have to make the system before it's like more likely than not that you can be like a drop in replacement for humans I mean I think if you just literally say like you train on web text then like the question is like kind of hard to discuss because you like I don't really buy stories that like training data makes a big difference long run to these Dynamics but I think like if you want to just imagine the hypothetical like you just took gp4 and like made the numbers bigger then I think those are pretty significant issues I think there're significant issues in two ways one is like quantity of data and I think probably the larger one is like quality of data where like I think as you start approaching like the prediction task is not that great a task if you're like a very weak model it's a very good signal you get smarter at some point it becomes like a worse and worse signal to get smarter so think there's a number of reasons like you couldn't it's it's not clear there is any number such as I imagine or there is a number but I think it's very large so you like plug that number into like GPT Force code and then maybe fill it with the architecture a bit I would expect that thing to have a more than 50% chance of being a drop in replacement for humans you're always going to have to do some work but the work's not necessarily much like I would guess when people say like new insight is needed I think I tend to be like more bullish than them I'm not like these are new ideas where like who knows how long it will take I think it's just like you have to do some stuff like you have to make changes unsurprisingly like every time you scale something up by like five orders of magnitude you have to make like some changes I I want a better understand your intuition of being more skeptical than some about this the scaling picture that you know these changes are needed in the first place or that it would take more than two orders of magnitude more Improvement to get these things almost certainly to a human level or very high probability to human level so uh is it that you don't agree with the way in which they're extrapolating these loss curves you don't agree with the implication that that decrease in loss will equate to greater and greater intelligence or like what would you tell Dario about if you were having I'm sure you have but like what what what would that debate look like about this yeah so again here we're talking two factors of a half one on like is it smart enough and one on like to have to do a bunch of sleap even if like in some sense it's smart enough and like the first factor of a half I'd be like I don't know think we have really anything good to extrapolate that is like I feel I would not be surprised if I have like similar or maybe even higher probabilities on like really crazy stuff over like the next year and then like lower Pro like my probability is like not that bunched up like maybe dar's probability I don't know you talk with him is like you have talked with him is more bunched up on like some particular year and mine is maybe like a little bit more like uniformly spread out across like the the coming years partly because I'm just like I don't think we have some Trends we can extrapolate like can extrapolate loss you can like look at your qualitative impressions of like systems at various scales but it's just like very hard to relate any of those extrapolations to like doing cognitive work or like accelerating R&D or taking over and fully automating R&D so I have a lot of uncertainty around that extrapolation I think it's very easy to get down to like a 50/50 chance of this um what about the sort of basic intuition that listen this is a big bla compute you make the Big Blob compute bigger it's going to get smarter like it' be really weird if it didn't yeah I'm happy with that it's going to get smarter and it would be really weird if it didn't and the question is how smart does it have to how smart does it have to get like that argument does not yet give us a quantitative guide to like at what scale is it is it a slam donk or what scale is a 50/50 and what would be the piece of evidence that would nud you one way or another where you look at that and be like oh this is at 20% by 2040 or 60% by 2040 or something like is there something that could happen in the next two years or next three years like what is the thing you're looking to where this will be a big update for you again I think there's some just how capable is each model where I like have I think we're really bad extrapolating but you still have some subjective guesss and you're comparing it to what happened and that will move me like every time we see what happens with another like order of magnitude of training compute I will have like a slightly different guess for where things are going these probabilities are course enough that again I don't know if that 40% is real or if like post gbt 3.

5 and four I should be at like 60% or what that's one thing and the second thing is just like some if there was some ability to extrapolate I think this could like reduce error bars a lot I think like here's another way you could try and do an extrapolation is you could just say like how much economic value do systems produce and like how fast is that growing I think like once you have systems actually doing jobs the extrapolation gets easier because you're like not moving from like a subjective impression of a chat to like automating all R&D you're moving from like automating this job to automating that job or whatever unfortunately that's like probably by the time you have nice Trends from that you're like uh you're not talking about 2040 you're talking about like you know two years from the end of days or one year from the end of days or whatever but like to the extent that you can get extrapolations like that I do think it can provide more clarity but why why is economic value the thing we would want to extrapolate because uh like if for example you started off with chimps and there just getting gradually smarter to human level they would basically provide like no economic value until they were you know basically as much as a human so it would be this like you know very gradual and then very fast increase in their value so is the is the you know increase in value from gbd4 gb5 gb6 is that the extrapolation we want yeah I think that the economic extrapolation is not great I think it's like you could compare it to this objective extrapolation of like how smart does the model seem it's like not super clear which one's better I think probably in the chimp case I like don't think that's quite right I think if you like actually like so if you imagine like intensely domesticated chimps who are just like actually trying their best to be really useful employees and like you hold fix their physical Hardware um and then you just gradually like scale up their intelligence I don't think you're going to see like zero value which then suddenly becomes massive value um over like you know one doubling of brain size or whatever one order of magnitude of brain size it's actually possible an order magnitud brain size but like chimps are very chimps are already within order magnitud brain sizes of humans like chimps are like very very close on the kind of spectrum we're talking about so I think like I'm skeptical of like the abrupt transition for chimps and to the extent that I kind of expect a fairly abrupt transition here it's mostly just cuz like the chimp intelligence difference is like so small compared to the differences we're talking about with respect to these models um that is like I would not be surprised if in some objective sense like chimpum difference is like significantly smaller than the gpt3 gp4 difference so the GPT 4 GPT five difference wait wouldn't that argue in favor of just relying much more on this objective uh yeah this is there's sort of two balancing tensions here one is like I don't believe the chimp thing is going to be as abrupt that is I think if you scaled up from chimps to humans you actually see like quite large economic value from like the fully domesticated chimp already um okay and then like the second half is like yeah I think that the chimp human difference is like probably pretty small compared to model differences so I do think things are going to be pretty abrupt I think the economic extrapolation is pretty rough I also think the subjective extrapolation is like pretty rough just because I really don't know how to get like how do I don't know how people do the extrapolation end up with the degrees of confidence people end up with again I'm putting pretty high if I'm saying like you know give me 3 years and I'm like yeah 50/50 it's going to have like basically the smarts there to do the thing that's like I'm not saying it's like a really long way off like I'm just saying like I got pretty big error bars and I think that like it's really hard not to have really big air bars when you're doing this like I looked at gp24 it seemed pretty smart compared to GPT 3.

5 so I bet just like four more such notches and we're there it's like that's just a hard call to make I think I sympathize more with people who are like how could it not happen in three years than with people who are like no way it's going to happen in eight years or whatever which is like probably a more common perspective in the world but also things do take longer than you I think things T longer than you think it's like a real thing um yeah I don't know mostly have big a bars because I just don't believe the subjective extrapolation that much I find it hard to get like a huge amount out of it okay so what what about the scaling picture do you think is most likely to be wrong yeah so we've talked a little bit about how good is the qualitative extrapolation how good are people at comparing so this is not like the picture being qualitative wrong this is just quantitatively it's very hard to know how far off you are I think a qualitative consideration that could significantly slow things down is just like right now you get to observe this like really rich supervision from like basically next word prediction or like in practice maybe you're looking at like a couple sentences prediction um so getting this like pretty rich supervision it's plausible that if you want to like automate long Horizon tasks like being an employee over the course of a month um that that's actually just like considerably harder to supervise or that like you basically end up driving costs like the worst case here is that you like drive up costs by a factor that's like linear in the Horizon over which the thing is operating and I still consider that just it's like quite plausible wait wait can you can you dump that down you're driving up a a cost about of what in L linear in the what does the horizon mean yeah so like if you imagine you want to train a system to like say words that sound like the next word a human would say yeah there you can get this like really rich supervision by having a bunch of words um and then predicting the next one and being like I'm going to tweak the model so it predicts better if you're like hey here's what I want I want my model to like interact with like some job over the course of a month and then at the end of that month like have internalized everything with the human would have internalized about how to do that job well and like have local context and so on um it's harder to supervise that task so in particular you could supervise it from the next word prediction task and like all that context the human has ultimately will just help them predict the next word better so like in some sense a really long context language model is also learning to do that task but the number of like effective data points you get of that task is like vastly smaller than the number of effective data points you get at like this very short Horizon like what's the next word what's the next sense tasks the sample efficiency matters more for economic valuable long Horizon tasks than the predicting next token and that that's where what what will like actually be required to you know take over a lot of jobs yeah something something like that um that is it just seems very plausible that it takes longer to train models to do tasks that are longer Horizon how how fast do you think the pace of algorithmic advances will be because if if by 240 um even if scaling fails I mean you know how well back since 2012 since the beginning of the deep learning re ution we've had so many new things by 240 are you expecting a similar pace of increases and if so then I mean if we just keep having things like this then aren't we just going to get a sooner or later or soon not later AR we going to get sooner or sooner I'm with you on sooner or later yeah um I suspect like progress to slow if you like held fixed how many people are working in the field I would expect progress to slow as Ling fruit is exhausted I think the like rapid rate of progress in like say language modeling over the last four years is largely sustained by like you started from a relatively small amount of investment you like greatly scale up the amount of investment um and that enables you to like keep picking you know every time every time the difficulty doubles you just double the size of the field like I think that Dynamic can hold up for some time longer like I mean pretty good like you know right now if you think of it as like hundreds of people effectively searching for things like up from like you know anyway if you think of hundreds of people now you can maybe bring that up to like tens of thousands of people or something so for a while just continue increasing the size of the field and like search harder and harder and there's indeed like a huge amount of low hanging fruit where like it wouldn't be that hard for a person to sit around and like make things a couple per better after after year of work or whatever so I don't know I would probably think of it mostly in terms of like how much can investment be expanded and like try and guess like some combination of fitting that curve and um yeah trying some combination of fitting the curve to historical progress looking at like how much low hanging fruit there is getting a sense of how fast it decay is I think like you probably get a lot though you get a bunch of orders of magnitude of total especially like if you ask like how good is a GPT 5 scale model or GPT 4 scale model I think you probably get like by 2040 like I don't know three orders of magnitude of effective training compute Improvement or like a good chunk of effective training compute Improvement forwarders of magnitude I don't know I don't have like here I'm speaking from like no private information about the last like couple years of efficiency improvements and so people who are on the ground will have better senses um of exactly how rapid returns are and so on okay let me back up and ask a question more generally about you know people make these analogies about humans were trained by Evolution and we like deployed in this in the modern civilization do you buy those analogies is it valid to say the humans were trained by Evolution rather than I mean if you look at the protein coding size of the genome it's like 50 megabytes or something and then what part of that is for the brain anyways how do you think about how much information uh is in um like do you think of the genome as hyperparameters or how much does that inform you when you have these anchors for how much training humans get when they're just consuming information when they're walking up and about and so on yeah I guess the way that you could think of this is like I think both analogies are reasonable one analogy being like evolution is like a training run and humans are like the end product of that training run and a second analogy is like evolution is like an algorithm designer and then a human over the course of like this modest amount of computation over their lifetime um is the algorithm being that's been produced The Learning algorith that's been produced and I think like neither analogy is that great like I like them both and lean on them a bunch both like both of them a bunch and think that's been like pretty good for having like a reasonable view of what's likely to happen that said like the human genome is not that much like 100 trillion parameter model it's like a much smaller number of parameters that behave in like a much more confusing way Evolution did like a lot more optimization especially over like long like designing a brain to work work well over a lifetime than gradient descent does over models that's like a dis analogy on that side and on the other side like I just I think human learning over the course of human lifetime is in many ways just like much much better than gradient descent over the space of neural Nets like gradi descent is working really well but I think we can just be quite confident that like in a lot of ways human learning is much better human learning is also constrained like we just don't get to see much data and that's just an engineering constraint that you can relax like you can just give your neur and that's way more data than humans have access to in what ways is human learning Superior to grading to end um I mean the most obvious one is just like ask how much data it takes a human to become like an expert in some domain and it's like much much smaller than the amount of data that's going to be needed on any plausible Trend extrapolation like not in terms of performance but is it the active learning part is it the structure like what is it I mean I would guess a complicated mess of a lot of things in some sense there's not that much going on in a brain like as you say there's just not that many there's not that many bites in a genome um but there's very very few bytes in an ml algorithm like if you think a genome is like a billion bytes or whatever maybe you think less maybe you think it's like 100 million bytes um then like you know an ml algorithm is like if compressed probably more like hundreds of thousands of bytes or something like the total complexity of like here's how you train gp4 is just like I haven't thought about these numbers but like it's very very small compared to a genome and so although a genome is very simple it's like very very complicated compared to algorithms that humans design like really hideously more complicated than algorithm a human would design is is that true so okay so the the human genome is three billion base pairs or something um but only like one or two% of that is protein coating so that's 50 million base pairs I I don't yeah so I don't know much about biology in particular I guess the question is like how many of those bits are like productive for like shaping development of a brain and presumably a significant part of the non-protein coding genome can I mean I just don't know it seems really hard to guess how much of that plays a role like the most important decisions are probably like from an algorithm design perspective are not like like the protein coding part is is less important than the like decisions about like what happens during development or like how cells differentiate I don't know if that's I know nothing about biologist spec I'm happy to run with 100 million Bas pairs though but on the other end on the hyper parameters of the cheep before training run that might be not that much but if you're going to include all the all the base pairs in uh the genome then which are not all relevant to the brains or are relevant to like very bigger details about like how just basics of biology should probably include like the python library and the compilers and the operating system for gbd4 as well uh to make that comparison analogous so at the end of the day I actually don't know which which one has storing more information yeah I mean I think the way I would put it is like the number of bits it takes to specify the learning algorithm to train gbt for is like very small and you might wonder like maybe a genome like the number of bits it would like take to specify a brain is also very small the genome is much much faster than that um but it is also just plausible that a genome is like closer to like certainly the space the amount of space to put complexity in a genome we could ask how well Evolution uses it and like I have no idea whatsoever but the amount of space in a genome is like very very vast compared to the number of bits that are actually taken to specify like the architecture or optimization procedure and so on for GPT 4 just because again genome is simple but algorithms are like really very simple ml algorithms are really very simple and it's stepping back you think this is where the U the better sample efficiency of human learning comes from like why gradi desent yeah so I haven't thought that much about the sample efficiency question in a long time but if you thought like a synapse was seeing something like you know and they're on firing once per second then how many seconds are there in a human life we can just f up a calculator real quick yeah let's do some calculating tell me the number 3600 seconds per hour time 24 * 365 * 20 okay so that's 630 million seconds that means like the average synapse is saying like 630 30 million and I don't know exactly what the numbers are but something that's ballpark let's call like a billion Action potentials and then you know there's there's some resolution each of those carry some bits but let's say it carries like 10 bits or something um just from like timing information of the resolution you have available then you're looking at like 10 billion bits so each parameter is kind of like how much is a parameter seeing it's like not seeing that much so then you can compare that to like language I think that's probably less than like current language model C and current language models are so it's like not clear you have a huge gap here but I think it's pretty clear you're gonna have a gap of like at least three or fours of magnitude didn't your wife do the the the lifetime anchors where she said the amount of bites that uh a human will see in their lifetime was one24 or something the number of bites a human will see is one24 mostly this was organized around total operations performed in a brain right oh okay never mind sorry yeah yeah so I think that like the story there would be like a brain is just in some other part of the parameter space where it's like using a lot a lot of compute for each piece of data it gets and just not seeing very much data in total yeah there's just not it's not really possible if you extrapolate out language models you're going to end up with like a performance profile similar to a brain I don't know how much better it is like I think so I did this like random investigation at one point where I was like how good are things made by Evolution compared to things made by humans right um which is a pretty insane seeming exercise but like I don't know it seems like orders of magnitude is typical like not tons of orders of magnitude not factors of two like things by humans are thousand times more expensive to make or thousand times heavier per unit performance if you look at things like how good are solar panels relative to leaves or how good are muscles relative to Motors or how good our livers relative to systems that perform analogous chemical reactions and Industrial settings was there a consistent number of orders of magnitude in these different systems or was it all over the place uh so like very rough ballpark it was like sort of for the most extreme things you were looking at like five or six ords of magnitude and that would especially come in like energy cost of manufacturing where like bodies are just very good at building complicated organs like extremely cheaply um and then for other things like leaves or eyeballs or livers or whatever you tended to see more like if you set aside manufacturing costs and just look at like operating costs or like performance trade-offs like I don't know more like three orders of magnitude or something like that or or something that are on the smaller scale like the Nano machines or whatever that we can't do it all right yeah that's I mean yeah so it's a little bit hard to say exactly what the task definition is there like you could say like making a bone we can't make a bone but you could try and compare a bone the performance characteristics of a bone is something else like we can't make spider silk do to try and compare the performance characteristics with spider silk like things that we can't synthesize the reason this would be is why that Evolution has had more time to design these systems or I don't know I just mostly just curious about like what the performance I think like most people would object to be like how did you choose these reference classes of things that are like Fair intersections some of them seem reasonable like eyes versus cameras seems like just everyone needs eyes everyone needs cameras it feels very fair photosynthesis seems like very reasonable everyone needs to like take solar energy and then like turn it into usable form of energy um but I was just kind of I don't really have a mechanistic story evolution in principle has spent like way way more time than we have designing it's absolutely unclear how that's going to shake out my guess would be in general like I think there aren't that many things where humans really Crush Evolution where you can't tell like a pretty simple story about why so like for example roads and moving over roads with wheels crushes Evolution but it's not like an animal like would have wanted to design a wheel like you're just not allowed to like pave the world and then put things on Wheels if you're an animal maybe planes or more anyway whatever there's various things you can try tell there's some things humans do better it's normally pretty clear why humans are able to win when humans are able to win the point of all this was like it's not that surprising to me I think this is mostly like a pro short timelines view it's not that surprising to me if you tell me like machine learning systems are like three or fours of magnitude less efficient at learning than human brains I'm like that actually seems like kind of IND distribution for other stuff and if that's your view then I think you're like probably going to hit you know then you're looking at like 10 to the 27 training compute or something like that which is is not so far we'll get back to the timeline stuff in a second uh at some point we should talk about alignment so let's let's talk about alignment at what stage does misalignment happen so right now with something like gp4 I'm not even sure it would make sense to say that it's misaligned um because it doesn't it's not aligned to anything in particular is it is it at human level where you think the ability to be deceptive comes about what is the process by which misalignment happens I think even for gp4 it's reasonable to ask questions like are there cases where gp4 knows that humans don't want X but it does X anyway like where it's like well I know that like I could give this answer which is misleading and if it was explained to a human what was happening they wouldn't want that to be done but I'm going to produce it I think that like gb4 understands things enough that you can have like that misim in that sense yeah I think gbt like I've sometimes talked about being like benign instead of aligned meaning that like well it's not exactly clear if it's aligned or if that context is Meaningful it's just like kind of a messy word to use in general but the thing we're more confident of is it's like not doing you know it's not optimizing for this goal which is like across purposes to humans it's either optimizing for nothing or like maybe it's optimizing for what humans want or close enough or something that's like an approximation good enough to still not take over but anyway I'm like some of these abstractions seem like they do apply to gp4 um it seems like probably it's not like egregiously misaligned it doesn't it's not doing the kind of thing that could lead to takeover we'd guess suppose you have a system at some point and which ends up in it wanting take over what are the checkpoints and also what is the internal is it just that it to become more powerful and needs agency and agency implies other goals or do you see a different process by which misalignment happens yes I think there's a couple possible stories for getting to catastrophic misalignment and they have slightly different answers to this question um so maybe I'll just briefly describe two stories and try and talk about when they can when they start making sense to me so one type of story is you train or fine-tune your AI system to do things that humans will rate highly or that like get other kinds of reward in a broad diversity of situations and then it learns to in general drop in some new situation try and figure out which actions would receive a high reward or whatever um and then take those actions and then when deployed in the real world like sort of gaining control of its own training data provision process is something that gets a very high reward and so it does that so this is like one kind of story um like it wants to grab the reward button or whatever it wants to intimidate the humans into giving it a high reward Etc I think that doesn't really require that much this basically requires a system which is like in fact looks at a bunch of environments is able to understand like the mechanism of reward provision as like a common feature of those environments is able to think in some nomal environment like hey which actions would result in be getting a high reward and is thinking about that concept precisely enough that when it says High reward it's saying like Okay well how is reward actually computed it's like some actual physical process being implemented in the world my guess would be like gp4 is about at the level where with handholding you can observe this kind of like scary generalizations of this type although I think they haven't been shown basically um that is you can have a system which in fact is fine to not a bunch of cases and then in some new case will try and like do an end run around humans even in a way humans would penalize if they were able to notice it or would have penalized in training environments so I think gb4 is kind of at the boundary where these things are possible um examples kind of exist but are getting significantly better over time um I'm very excited about like this this anthropic project basically trying to see how good an example can you make now um of this phenomena and I think the answer is like kind of okay probably um so that just I think is going to continuously get better from here I think for the level where we're concerned like this is related to me having really broad distributions over how smart models are I think it's like not out of the question that you take GP like GPT 4's understanding of the world is like much crisper and like much better than gbt 3's understanding um just like it's really like night and day and so it would not be that crazy to me if you took GPT 5 and you trained it to get a bunch of reward and it was actually like okay my goal is not doing the kind of thing which like thematically looks nice to humans my goal is getting a bunch of reward um and then we'll generalize in a new situation to get reward and by the way this requires to consciously want to uh do something that it knows the humans wouldn't wanted to do or is it just that we weren't good enough to specify that the thing that we accidentally ended up rewarding is not what we actually want I think the scenarios I am most interested in and most people are concerned about from a catastrophic risk perspective it involves understanding that they're taking actions which a human would penalize if they the human was aware of what's going on such that you have to either deceive humans about what's happening um or you need to like actively subvert human attempts to correct your behavior so these the failures come from really this combination or they require this combination of both like trying to do something humans don't like and understanding the humans would stop you right I think you can have only the barest examples you can have the barest examples for gp4 like can create the situations where gp4 would be like sure in that situation like here's what I would do I would like go hack the computer and change my reward or in fact we like do things that are like simple hacks or like go change the source of this file or whatever to get a higher reward they're pretty weak examples I think it's plausible GPT 5 will have like compelling examples of those phenomena I really don't know this is very related to like the very broad error bars on like how competent such systems will be when um that's all with respect to this first mode of like a system is taking actions that get reward and like overpowering oring humans is helpful for getting reward there's this other failure mode and other family failure modes where AI systems want something unrelated to reward I understand that like they're being trained and like while you're being trained there are a bunch of like reasons you might want to do the kinds of things humans want you to do but then when deployed in the real world if you're able to realize you're no longer being trained you no longer have reason to do the kinds of things you want you'd prefer like be able to determine your own destiny like control your own your competing Hardware Etc which I think like probably emerg like a little bit later than systems that try and get reward and so we'll generalize in scary unpredictable ways to new situations I don't know when those appear but also Again Brad enough error bars that it's like conceivable for systems in the near future you know I wouldn't put it like less than one in a, for GPT 5 certainly if if we deployed all these AI systems and some of them are reward hacking some of them are deceptive some of them are just normal whatever how do you imagine that they might interact with each other at the expense of humans uh how hard do you think it would be to for them to communicate in ways that we would not be able to recognize and coordinate um coordinate our at our expense yeah I think that most really istic failures probably involve two factors interacting one factor is like the world is pretty complicated and the humans mostly don't understand what's happening so like AI systems are writing code that's very hard for humans to understand maybe how it works at all but more likely like they understand roughly how it works but there's a lot of complicated interactions um AI systems are running businesses that interact primarily with other AIS they're like doing SEO for like AI search processes they're like running Financial transactions like thinking about to trade with AI counterparties um and so you can have this world where like even if humans kind of understand the jumping off point when this was all humans like actual considerations of like what's a good decision like what code is going to work well and be durable or like what marketing strategy is effective for selling to these other AIS or whatever is kind of just all mostly outside of sort of humans understanding I think this is like a really important again when I think of like the most plausible scary scenarios I think that's like one of the two big risk factors and so in some sense your first problem here is like having these Adit systems who understand a bunch about what's happening and your only lever is like hey I do something that works well so you don't have a lever to be like hey do what I really want you just have the system you don't really understand you can observe some outputs like did it make money and you're just optimizing or at least doing some fine tuning to get the ai2 its understanding of that system to achieve that goal so I think that's like your first risk factor and like once you're in that world then I think there are like all kinds of Dynamics amongst AI systems that again humans aren't really observing humans can't really understand humans aren't really exerting any direct pressure on only on outcomes and then I think it's it's quite easy to be in a position where you know if AI system started failing it would be very they could do a lot of harm very quickly um humans aren't really able to like prepare for and mitigate that potential harm because we don't really understand the systems in which they're acting um and then if AI systems like you know they could successfully prevent humans from either understanding what's going on or from like successfully like retaking the data centers or whatever if the if the AI successfully grab control this seems like a much more gradual um story than the conventional takeover stories where it just like you train it and then it comes alive and you know escapes and takes over everything so you think that that kind of story is less likely than one in which we just hand off more control voluntarily to the AIS so one I am interested in the tale of like some risks that can occur particularly soon and I think risks that occur particularly soon are a little bit like you have a world where has not probably deployed and then something crazy happens quickly that said if you ask like what's the median scenario where things go badly I think it is like there's some lessening of our understanding of the world it becomes I think like in the default path it's like very clear to humans that they have increasingly little grip on what's Happening I mean I think already most humans have very little grip on what's happening it's just some other humans understand what's happening like I don't know how almost any of the systems I interact with work in a very detailed way um so it's sort of clear to humanity as a whole that like we sort of collectively don't understand most of what's Happening except with a assistance and then like that process just continues for a fair amount of time and then like there's a question of how abrupt an actual failure is I do think it's it's reasonably likely a failure itself would be abrupt like at some point bad stuff starts happening the human can recognize as bad and once things that are obviously bad start happening then like you have this bifurcation where either humans can use that to fix it and say okay a behavior that led to this obviously bad stuff don't do more of that or you can't fix it um and then like you're in this like rapidly escalating failures everything goes off the rails in that case yeah what does going off the rails look like for example how would it take over the government yeah it's it's getting deployed in the economy in the world and at some point it's in charge how does that transition happen yeah so this is going to depend a lot on like what kind of like timeline you're imagining or like the sort of a broad distribution but I can like fill in some random concrete option that is like in itself very improbable um yeah I mean I think that like one of the less dignified but maybe more plausible routes is like you just have a lot of AI Control over critical systems even in like running a military um and then you have the scenario that's a little bit more just like a normal coup where you have a bunch of AI systems they in fact operate like you know it's not the case that humans can really fight a war on their own it's not the case that humans could defend them from like an invasion on their own so like that is if you had invading Army and you had your own robot army you're like you can't just be like we're going to turn off the robots now cuz things are going wrong if you're in the middle of a war okay so how much does this world rely on Race dyamics where we're forced to deploy or not forced but we choose to deploy AIS because other other F countries or other companies are also deploying AIS and you know you can't have them have all the Killer Robots yeah I mean I think that like there's several levels of answer to that question so one is like maybe three three three parts of my answer like our first part of is like I'm just trying to tell like what seems like the most likely story I do think there's like further failures to get you like in the more distant future so like e elzer will not talk that much about Killer Robots because he really wants to emphasize like hey if you never built a kill a robot something crazy is still going to happen to you're just like only four months later or whatever so it's like not really the way to analyze the failure but if you want to ask like what's the median world where something bad happens I still do think this is like the best guess um okay so that's like part one of my answer part two of the answer was like in this proximal situation where something bad is happening and you ask like hey why do humans not turn off the AI you could imagine like two kinds of Story one is like the I is able to prevent humans from turning that off them off and the other is like in fact we live in a world where it's like incredibly challenging like there's a bunch of competitive Dynamics or a bunch of Reliance on AI systems and so it's incredibly expensive to turn off AI systems I think again you would eventually have the first problem like eventually a systems could just prevent humans from turning them off but I think like in practice the one that's going to happen much much sooner is probably competition amongst different actors using Ai and it's like very very expensive to unilaterally disarm you can't be like something weird has happened where it's going to shut off all the AI because you're EG in a hot War so again I think that's just like probably the most likely thing to happen first um things would go badly without it but I think if you ask why don't we turn off the AI my best guess is because there are a bunch of other AI running around to year lunch so how how much better a situation would we be in if there was like there was only one group that was pursuing AI no other countries no other companies basically how much of the expected value is lost from the Dynamics that are likely to come about because other people will be developing and deploying these systems yeah so I guess this brings you to like a third part of the way in which competitive Dynamics are relevant so it's both the question of can you turn off AI systems in response to something bad happening where competitive Dynamics may make it hard to turn off there's a further question of just like why were you deploying systems for which you had very little ability to control or understand those systems and again it's possible you just don't understand what's going on you think you can understand or control such systems but I think in practice the significant part is going to be like you are doing the calculus or people deploying systems are doing the calculus as they do today like in many cases overtly of like look these systems are not very well controlled or understood there's some chance of like something going wrong or at least going wrong if we continue down this path but other people are developing the technology potentially in even more Reckless ways so in addition to like competition making it difficult to shut down AI systems in the event of a catastrophe I also think it's just like the easiest way that people end up pushing relatively quickly or moving quickly ahead on a technology where they feel kind of bad about understandability or controllability that could be economic competition or military competition or whatever so I kind of think ultimately like most of the harm comes from the fact that like lots of people can develop AI um how hard is a takeover of the government or something from an AI even if it doesn't have Killer Robots but just a thing that you can't kill off if it has seeds elsewhere it can easily replicate it can think a lot and think fast uh what is a minimum viable coup for uh is it like shutting up just like threatening bowar or something or shutting off the greater how how how easy is it basically to take over humans civilization so again there's going to be a lot of scenarios and I'll just like start by talking about one scenario which will represent a tiny fraction of probability or whatever but like so if you're not in this competitive world if you're saying like we're actually slowing down deployment of AI because we think it's unsafe or whatever then in some sense you're creating this like very fundamental instability where like you could have been making faster AI progress and you could have been deploying AI faster and so in that world the like most the bad thing that happens if you have an AI system that wants to mess with you is the AI system says like I don't have any compunctions about like rapid deployment of AI or rapid AI progress so the thing you want to do or the AI wants to do is just say like I'm in a defect from this regime like all the humans have agreed that we're like not deploying AI in ways that would be dangerous but if I as an AI can escape and just go set up my own shop like make a bunch of copies of myself maybe the humans didn't want to like delegate War fighting to an AI but I as an AI I'm pretty happy doing so like I'm happy if I'm able to grab some military equipment or direct some humans to use use AI use myself to direct it and so I think like as that Gap grows so if people are deliberately right if people are deploying AI everywhere I think of this competitive Dynamic if people aren't deploying AI everywhere so like if if countries are not happy deploying AI in these high stake settings then as AI improves you create this like wedge that grows where like if you were in a position of fighting against an AI which wasn't constrained in this way you'd be in a pretty bad position um at some point like even if you just yeah so that's like one important thing just like I think in conflict in like overt conflict if humans are putting the brakes on AI they're at like a pretty major disadvantage compared to an AI system that can kind of set up shop and operate independently from humans a potential independent AI does it need collaboration from a human faction I again you could tell different stories but it seems so much easier at some point you don't need any at some point A system can just operate completely like out of human supervision or something but that's like so far after the point where it's like so much easier if you're just like they're a bunch of humans they don't love each other that much like some humans are happy to be on side they're they skeptical about risk or happy to make this trade or can be fooled or can be csed or whatever and just seems like it is almost certainly almost certainly the easiest first pass is going to involve like having a bunch of humans who are happy to work with you so yeah I think that probably is about I think it's not ultimate not necessary but if you ask about the median scenario it's it involves a bunch of humans working with AI systems either yeah being directed by AI systems providing comput to AI systems providing like legal cover and jurisdictions that are sympathetic to AI systems humans presumably would not be if they knew the end result of the AI takeover would not be willing to help so they have to be probably fooled in some way right like deep fakes or something and what is the minimum viable physical presence they would need or jurisdiction they would need in order to carry out their schemes do you need a whole country do you just need a server Farm do you just need like one single laptop yeah I think I'd probably start by pushing back a bit on the like humans wouldn't cooperate if they understood outcome or something like I would say like one even if you're if you're looking at something like tens of percent risk of take over humans may be fine with that like a fair number of humans may be fine with that to like if you're looking at certain takeover but it's very unclear if that leads to death like a bunch of humans may be fine with that like if we're just talking about like look the AI systems are going to like run the world but it's not clear if they're going to murder people like how do you know it's just a complicated question about AI psychology and a lot of humans probably are fine with that and I don't even know what the probability is there I think you actually have given that probability online I've certainly guessed okay but it's not zero it's like a significant percentage I gave like 50/50 oh okay yeah why why is it tell me about the world in which the AI takes over but doesn't kill humans why would that happen and what would that look like I mean I I as my questions like why would you kill humans so I think like the maybe I'd say the incentive to kill humans is like quite weak they they'll get in your way they control you want Oho taking from humans is a different like marginalizing humans and like causing humans to be irrelevant is a very different story from killing the humans I think I'd say like the actual incentives to kill the humans are quite weak such I think like the big reasons you kill humans are like well one you might kill humans if you're like in a war with them and like it's hard to win the war without killing a bunch of humans like maybe most saliently here if you like want to use some biological weapons or some crazy that might just kill humans I think like you might kill humans just from totally destroying the ecosystems they're dependent on and it's slightly expensive to keep them alive anyway you might kill humans just cuz you don't like them or like you literally want to like yeah I mean neutralize a threat or the alas are line is that they're made of atoms you could use or something else yeah I mean I think the literal they're made of atoms is like quite there are not many atoms in humans um neutralize a threat is as the similar issue where it's just like I think you would kill the humans if you didn't care at all about them so maybe maybe your question you're asking is like why would you care at all about hum um but I think you don't have to care much to not kill the humans okay sure because there there's just so much raw of resources elsewhere in the universe yeah also it's you can marginalize humans pretty hard like you could totally human like you can humans War fighting capability and also take almost all their stuff while killing only you know a small fraction of hum incidentally so then if you ask like why am might a not want to kill humans I mean a big thing was just like look I think a probably want like a bunch of random crap for like complicated reasons like the motivations of AI systems and civilizations of AIS are probably complicated messes certainly amongst humans it is not that rare to be like well there was someone here I would like all else equal if I didn't have to murder them I would prefer not murder them um and my guess is it's also like reasonable chance it's not that rare amongst a systems like humans have a bunch of different reasons we think that way um I think AI systems will be very different from humans but it's also just like a very Salient yeah I mean think this is a really complicated question like if you imagine drawing values from the basket of all values yeah like what fraction of them are like hey if there's someone here how much do I want to not murder them and my guess is just like if you draw a bunch of values from the basket that's like a natural enough thing like if your a wanted like 10,000 different things or your civilization wants 10,000 different things just like reasonably likely you get like some of that um the other Salient reason you might not want to murder them is just like well yeah there's some like kind of crazy decision Theory stuff or a causal trade stuff which does look on paper like it should work and like if I was if I was running a civilization and like dealing with some people who I didn't like it all or like didn't have any concern for it all but I only had to spend one billionth of my resources not to murder them I think it's like quite robust that you don't want to murder them um that is I think the like weird decision Theory a causal trade stuff probably does carry the day oh wait that's that's that contributes more to that uh 50/50 of will murderers if they take over then the um well just by default they might just not want to kill us yeah I think they're both Salient um can you explain the they run together with each other a lot the going explain the weiry causal yeah so reason why might not kill us maybe a higher level thing that goes into both of these and then I'll talk about how you instantiated in a causal trade is just like it matters a lot to the humans not to get murdered and the AI cares very very little about whether if if we imagine this hypothetical the reason it wants to kill humans is just total apathy it cares very little about whether or not to murder humans because it is so easy to like marginalize humans without murdering them and the resources required for human survival are extremely low um again in the context of this like rapid industrialization so like that's the basic setting and now the thing that you'd like to do is like run a trade like the I would like to say like hey humans you care a ton about not getting murdered I don't really care one way or the other I would like to if I could find some way in which I don't murder you and then in return I get something um the problem is in that world the humans have essentially nothing to give um that is the humans are mostly irrelevant um however from the perspective of any kind of reasonable AI it's like not that sure whether it lives in the world where the humans in fact have nothing to give or whether the humans like in fact it lives in a world where the humans succeeded at building a land Ai and now the AI is simply like running in a nice little simulation humans are wondering I wonder if to say I would have murdered Us Fall it had the chance and saying like if it would murder us all if it had the chance like that sucks we'd like to run this trade we'd like to be nice to the as who wouldn't have murdered us all in order to create an incentive for as not to murder us okay so we do is we just check and for the kinds of a who don't murder everyone we just give them like one billionth of the universe like if you do that as Humanity then if you're an AI system considering like do I want to murder everyone your Calculus is like if this is my real chance to murder everyone I get like the tiniest bit of value you know I get like one trillionth of the value or whatever one billionth of the value but on the other hand like if I don't murder everyone there's like some worlds wherein the humans will correctly determine I don't murder everyone because in fact the humans survived the humans are running the simulations to understand how different eyes would behave and so like that's a better deal let's hope they fall for that s off um okay that that's interesting hey real quick this episode is sponsored by open philanthropy open philanthropy is one of the largest Grant making organizations in the world every year they give away hundreds of millions of dollars to have reduced catastrophic risks from fast moving advances in Ai and biotechnology open philanthropy is currently hiring for 22 different roles in those areas including Grant making research and operations new hires will support open philanthropy giving on te technical AI safety AI governance AI policy in the US EU and UK and biocurity many roles are remote friendly and most of the grantmaking hires that open philanthropy makes don't have prior Grant making experience previous technical experience is an asset as many of these roles often benefit from a deep understanding of the Technologies they address for more information and to apply please visit open philanthropy website in the description the deadline to apply is November 9th so make sure to check out those rules before they close awesome back to the episode in a world where we've been deploying these AI systems and suppose they're aligned how how hard would it be for um competitors to I don't know Cyber attack them and get them to join the other side are they robustly going to be aligned yeah I mean I think in some sense so there's a bunch of questions that come up here first one is like are aligna systems that you can build like competitive are they almost as good as the best systems anyone could build and maybe we're granting that for the purpose of this question yeah um I think the next question that comes up is like as systems right now are very vulnerable to manipulation like it's not clear how much more vulnerable they are than humans except for the fact you can like if you have an a system you can just replay it like a billion times and search for like what thing can I say that will make it behave this way so as a result like a systems are very vulnerable to manipulation it's unclear if future a systems will be similar vulnerable to manipulation but certainly seems plausible and in particular like you know aligned AI systems or unaligned a systems would be vulnerable to all kinds of manipulation the thing that's really relevant here is kind of like asymmetric manipulation or something that is like if it is easier so if everyone is just constantly messing with each other's AI systems like if you ever use AI systems in a competitive environment a big part of the game is like messing with your competitor's AI systems um a big question is whether there's some asymmetric Factor there where like it's kind of easier to push a systems into a mode where they're behaving erratically or chaotically or like trying to grab power or something than it is to like push them to fight for the other side like it's just a game of like two people are competing and neither of them can sort of like hijack an opponent's AI to like help support their cause but that doesn't I mean it matters and it creates chaos and it like might be quite bad for the world but it doesn't really affect the alignment calculus now it's just like right now you have like normal cyber offense cyber defense you have like weird AI version cyber offense cyber defense but if you have this kind of asymmetrical thing where like you know a bunch of AI systems where like we love AI flourishing can then like go in and say like great AI hey how about you join us and that works like if they can search for a persuasive argument to that effect and that's kind of asymmetrical then like the effect is whatever values it's easiest to push like whatever it's easiest to argue to an AI that it should do that is like advantaged um so it may be very hard to build a systems like try and defend human interest but very easy to build a systems just like try and Destroy stuff or whatever just depending on like what is the easiest thing to argue to an AI that it should do or what's the easiest like thing to trick an AI into doing or whatever yeah I think if alignment is spotty like so if you have the AI system which like doesn't really want to help humans or whatever or in fact like wants some kind of random thing or like wants different things in different contexts um then I do think adversarial settings will be the main ones where you see the system or like the easiest ones where you see the system behaving really badly um and it's a little bit hard to tell how that shakes out okay and suppose it is more reli How concerned are you that whatever alignment technique you come up with um you know you publish the paper this is how the alignment Works How concerned are you that Putin reads it or China reads it and now they understand for example the Constitutional AI think for anthropic and then you just write on there oh never contradict Maad dong thought or something H How concerned should we be that uh these alignment techniques are you know universally applicable not necessarily just for uh Enlighten goals yeah I mean I think they're super universally applicable I think it's just like I mean the rough way I would describe it which I think is basically right is like some degree of alignment makes the ey systems like much more usable like it kind of you should just think of the technology of AI as including like a basket of like some AI capabilities and some like getting the AI to do what you want it's just part of that basket and so anytime we're like you know to extent alignment is part of that basket you're just contributing to all the other harms from AI like you're reducing the probability of this harm but you are helping the technology basically work and like the basically working technology is kind of scary from a lot of perspectives one of which is like right now even in a very authoritarian Society just like humans have a lot of power because it's you need to rely on just a ton of humans to do your thing and and in a world where AI is very powerful like it is just much more possible to say like here's how our society runs one person calls the shots and then a ton of AI systems do what they want I think that's like a reasonable thing to like dislike about Ai and a reasonable reason to be scared to push the technology to be really good but is that is that also a reasonable reason to be concerned about alignment as well that um you know this is in some sense also capabilities you're you're teaching people how to get these systems to do what they want yeah I mean I would generalize so we we earlier touched a little bit on like moral potential moral rights of AI systems and like now we're talking a little bit about how they could AI systems powerfully disempowers humans and can Empower authoritarians I think we could like list other harms from Ai and I think it is the case that like if L was bad enough people would just not build AI systems and so like yeah I think there's a real sense in which you should just be scared to extent you're scared of Alli you should be like well alignment although it helps with one risk does contribute to AI being more of a thing I do think you should shut down the other parts of AI before like if you were a policy maker or like a research or whatever looking in on this I think it's like crazy to be like this is the part of the basket we're going to remove like you should you should first remove like other parts of the basket because they're also part of the story of risk wait wait does that imply you think if for example all capabilities research shut was shut down that you think it'd be a bad idea to continue doing alignment research in in isolation of what is conventionally considered capabilities research um I mean if you told me it was never going to restart then it wouldn't matter and if you told me it's going to restart I guess it would be a kind of similar calculus to today whereas you it's going to happen so you you should have something yeah I think that like in some sense you're always going to face this trade-off where Lin makes it possible to deploy as systems um or like makes it more attractive to deploy systems and then or like in the authoritarian case makes it like tractable to deploy them for this for this purpose um and like if you didn't do any alignment there' be a nicer bigger buffer between your society and malicious uses of AI and like I think it's one of the most expensive ways to maintain that buffer like it's much better to maintain that buffer by not having the compute or not having the powerful AI but I think if you're concerned enough about the other risks there's definitely a case to be made for just like put in more buffer or something like that I'm not like I care enough about the Takeover risk that like I think it's just not a net positive way to buy buffer that is like again the the version of this that's most pragmatic is just like suppose you don't work on a lineman today it like decreases economic impact of AI systems they'll be like less useful if they're less reliable and if they more often don't do what people want and so you could be like great that just buys time for AI and you're like getting some trade-off there where you're like decreasing some risks of AI like if AI is more reliable and more does what people want was more understandable then that cuts down some risks but if you think AI is on balance bad even apart from takeover risk um then like the alignment stuff can easily end up being non negative but but presumably you don't think that right because um I guess this is something people have uh brought up to you because you you know you invented rhf which was used to train Chad GPT and Chad GPT brought AI to the front pages everywhere so I just wonder if you could measure how much more money went into AI because like how much people have raised in the last year or something but it's got to be billions uh that the contactual impact of that that that went into the AI investment uh and the talent that went into AI for example so presumably you think that was worth it so I guess you you're hedging here about like what what is the reason that that is worth it yeah like what's the What's the total trade-off there yeah I mean I think my take is like so I think slower AI development on balance is like quite good um I think that slowing AI development now or like say having less press around chat GPT is like a little bit more mixed than slowing AI development overall like I think it's still probably positive but much less positive because I do think there's a real effect of like the world is starting to get prepared or is getting prepared at like a much greater rate now than it was prior to the release of Chad GPT and so like if you can choose between progress now or progress later like you'd really prefer have more of your progress now which I do think slows down progress later I don't think that's enough to flip the sign I think like maybe it wasn't the far enough past but now I would still say like moving faster now is net negative but to be clear it's a lot less net negative than merely accelerating AI because I do think again the chat gbt thing I am glad people are having policy discussions now rather than like delaying the like chat gbt wake up thing by a year and then um having oh wait chat gbt was net negative or rhf was net negative uh so here we just on the acceleration just like how is the Press of chat gbt my guess is like okay my guess is that negative but I think like it's not super clear and it's like much less much less than slowing AI slowing AI is great if you could slow overall AI progress um I think slowing AI by like causing you know there's this issue where slowing AI now like for chpt you're building up this backlog like why does chat GPT make such a splash like I think people there's a reasonable chance if you don't have a splash about chat GPT have a splash about GPT 4 and if you fail to have a splash about gbt 4 as reasonable chance of a splash about gbt 4.