🎙️AI 访谈库
如果 Dario Amodei 说对了呢?——纽约时报专访
Dario Amodei · Anthropic

如果 Dario Amodei 说对了呢?——纽约时报专访

What if Dario Amodei Is Right About A.I.?

2024-04-12 · The Ezra Klein Show (The New York Times) · 1h32m · 约 93 分钟读完 · 原文
Dario 与 Ezra Klein 讨论指数级扩展曲线、即将到来的能力突破、Responsible Scaling Policy 分级安全体系,以及社会与政府如何应对。

from New York Times opinion this is the Ezra Klein show the really disorienting thing about talking to the people building AI is their altered sense of time you're sitting there discussing some world that feels like weird sci-fi to even talk about and then you ask well when do you think this is going to happen and they say I don't know two years behind those predictions are what are called the scaling laws and the scaling laws and I want to say this so clearly they're not laws they're observations they're predictions they're based off of a few years not a few hundred years or thousand years of data but what they say is that the more computer power and data you feed into AI systems the more powerful those systems get that the relationship is predictable and more that the relationship is exponential human beings have trouble thinking in exponentials think back to covid when we all had to do it if you have one case of coron virus and cases double every 3 days then after 30 days you have about a thousand cases that growth rate feels modest it's manageable but then you go 30 days longer now you have a million then you wait another 30 days now you have a billion that's the power of the exponential curve growth feels normal for a while then it gets out of control really really quickly what the AI developers say is the power of AI systems is on this kind of curve that it has been increasing exponentially their capabilities and that as long as we keep feeding in more data and more computing power it will continue increasing exponentially that is the scaling law hypothesis and one of its main Advocates is Dario amade amade led the team at open AI that created gpd2 that created GPD 3 he then left open AI to co-found anthropic another AI firm where he's now the CEO and anthropic recently released Claude 3 which is considered by many to be the strongest AI model available right now but Amad believes we're just getting started that we're just hitting the Steep part of the curve now he thinks the kinds of systems we've imagined in sci-fi they're coming not in 20 or 40 years not in 10 or 15 years they're coming in 2 to 5 years he thinks are going to be so powerful that he and people like him should not be trusted to decide what they're going to do so I asked him on the show to try to answer in my own head two questions first is he right second what if he's right I want to say that in the past we have done shows with Sam Altman the head of open Ai and Demis isabis the head of Google Deep Mind and it's worth listening to those too if you find this interesting we're going to put the links to them in show notes because comparing and contrasting how they talk about the ey curves here how they think about the politics you'll hear a lot about that in the Sam Alman episode it gives you kind of sense of what the people building these things are thinking and and how maybe they differ from each other as always my email for thoughts for feedback for guest suggestions as reclin show at NY times.

com Dario amade welcome to the show thank you for having me so there are these two very different rhythms I've been about with AI one is the curve of the technology itself how fast it is changing and improving and the other is the pace at which society is seeing and reacting to those changes what does that relationship felt like to you so I think this is an example of a phenomenon that you know what we may have seen a few times before in history which is that there's an underlying process that is smooth and in this case exponential and then there's a spilling over of that process into the public and the spilling over looks very spiky it looks like it's happening all of a sudden it looks like it comes out of nowhere and it's triggered by things hitting various critical points or just the public happen to be engaged at a certain time so I think the easiest way for me to describe this in terms of my own personal experience uh is you know so I I worked at at open AI for 5 years I was one of the first employees to join and they built a model in 2018 called gpt1 which used something like a 100,000 times less computational power than the models we build today I looked at that and I and my colleagues were among the first to run what are called scaling laws which is basically studying what happens as you vary the size of the model it's capacity to absorb information and the amount of data that you feed into it and we found these very smooth patterns and we had this projection that look if you spend a 100 million or a billion or 10 billion on these models instead of the $10,000 we were spending then projections that all of these wondrous things would happen then you know we imagined that they would have enormous economic value fast forward to about 2020 gpt3 had just come out it wasn't yet available as a chatbot I led the development of that along with the team that eventually left to join anthropic and maybe for the whole period of 2021 and 2022 even though we continued to train models that were better and better and open aai continued to train models and Google continued to train models there was surprisingly little public attention to the models and I looked at that and I said well these models are incredible they're getting better and better what's going on why isn't this happening could this be a case where I was right about the technology but wrong about the economic impact the Practical value of the technology and then all of a sudden when chat GPT came out it was like all of that growth that you would expect all that excitement over 3 years broke through and came rushing in so I want to linger on this difference between the curve at which the technology is improving and the way it is being adopted by Society so when you think about these break points and you think into the future what other breakpoints do you see coming where AI bursts into social Consciousness or use in a different way yeah so I think I should say first that it's very hard to predict these one thing I like to say is you know the underlying technology because it's a smooth exponential it's not perfectly predictable but in some ways it can be eerily predn naturally predictable right uh that's not true for these societal step functions at all it's very hard to predict what will catch on in some ways it feels a little bit like which artist or musician you know is going to catch on and and get to the top of the charts that said a few possible ideas I think one is related to something that you mentioned which is interacting with the models in a more kind of naturalistic way we've actually already seen some of that with Claude 3 where you know people feel that some of the other models sound like a robot and that talking to Claude 3 is more natural I think a thing related to this is you know a lot of companies have been held back or tripped up by how their models handle controversial topics and we were really able to I think do a better job than others of telling the model don't shy away from discussing controversial topics don't assume that both sides necessarily have a valid point but don't express an opinion yourself don't express views that are flagrantly biased as journalists you encounter this all the time right how do I be objective but not both sides on everything so I think going further in that direction of models having personalities while still being objective while still being um you know useful and not falling into various ethical traps that'll be I think a significant unlock for adoption the models taking actions in the world is going to be a big one I know basically all the big companies that work on AI are working on that instead of just I ask it a question and it answers and then maybe I follow up and it answers again can I talk to the model about oh I'm going to go on this trip today and the model says oh that's great I'm get you know I'll get a Uber for you to drive from here to there and I'll reserve a restaurant and I'll talk to the other people who are going to plan the trip and the model being able to kind of do things end to endend or going to websites or taking actions on your computer for you I think all of that is coming in the next I would say I don't know 3 to 18 months with increasing levels of ability I think that's going to change how people think about AI right where so far it's been this very passive it's like I go to the Oracle I ask it a question and the Oracle tells me things and you know some people think that's exciting some people think it's scary but I think there are limits to how exciting or how scary it's perceived as because it's contained within this box I want to sit with this question of the agentic AI because I do think this is what's coming it's clearly what people are trying to build and I think it might be a good way to look at some of the specific technological and and cultural challenges and so let me offer two versions of it people following the AI news might have heard about Devon which is not in release yet but is a an AI that that at least purports to be able to complete the kinds of tasks linked tasks that a a junior software engineer might complete right instead of asking to do a bit of code for you you say listen I want a website it's going to have to do these things work in these ways and maybe Devon if it works the way people are saying it works can actually hold that set of thoughts complete a number of different tasks and come back to you with a result I'm also interested in the version of this we might have in the real world the example I always use in my head is when can I tell an AI my son is turning five he loves dragons we live in Brooklyn give me some options for planning his birthday party and then when I choose between them can you just do it all for me order the cake reserve the room send out the invitations whatever it might be those are two different situations because one of them is in code and one of them is making decisions in the real world interacting with real people knowing if what it is finding on the websites is actually any good what is between here and there when when I say that in plain language to you What technological challenges or advances do you he need to happen to get there the short answer is not all that much you know a story I have from when we were developing models back in 2022 and this is before we'd hooked up the models to anything is you could have a conversation with these purely textual models where you could say hey I want to reserve dinner at you know Restaurant X in San Francisco and the model would say okay here's the website of Restaurant X and it would actually you know give you a correct website or would tell you to go to Open Table or something and of course it can't actually go to the website the power plug isn't actually plugged in right the brain of the robot is not actually attached to its arms and legs but it gave you this sense that like the brain all it needed to do was learn exact how to use the arms and legs right it already had a picture of the world and where it would walk and what it would do and so it felt like there was this very thin barrier between the passive models we had and actually acting in the world in terms of what we need to make it work one thing is literally we just need a little bit more scale and I think the reason we're going to need more scale is to do one of those things you described right to do all the things a junior software engineer does right they involve chains of long actions right I have to like write this line of code I have to run this test I have to write a new test I have to check how it looks in the app after I interpret it or compile it and these things can easily get 20 or 30 layers deep and you know same with planning the birthday party for your son right and if the accuracy of Any Given step is not very high right is not like 99.

you know 9% as you compose these steps the probability of making a mistake becomes itself very high so the industry is going to get a new generation of models every you know probably four to eight months and so my guess I'm not sure is that to really get these things working well we need maybe one to four more Generations so that ends up translating to you know 3 to 24 months or something like that I think second is just there is some algorithmic work that is going to need to be done on how to have the models interact with the world in this way I think the basic techniques we have uh you know method called reinforcement learning and variations of it probably is up to the task but figuring out exactly how to use it to get the results we want will probably take some time and then third I think and and this gets to something that anthropic really specializes in is safety and controlability and I think that's going to be a big issue for these models acting in the world right let's say this model is writing code for me and it introduces a serious security bug in the code or it's taking actions on the computer for me and modifying the state of my computer in ways that are too complicated for me to even understand and for planning the birthday party right the level of trust you would need to take an AI agent and say I'm okay with you calling up anyone saying anything to them that's in any private information that I might have sending them any information taking any action on my computer posting anything to the internet the most unconstrained version of that sounds very scary and so we're going to need to figure out what is safe and controllable the more open-ended the thing is the more powerful it is but also the more dangerous it is and and the harder it is to control so I think those questions although they sound lofty and Abstract are going to turn into practical product questions that we we and other companies are going to be trying to address when you say we're just going to need more scale you mean more compute and more training data and I guess possibly more money to Simply make the models smarter and more capable yes we're going to have to make bigger models that use more compute per iteration we're going to have to run them for longer by feeding more data into them and that number of chips times the amount of time that we run things on chips is essentially a dollar value because you know these chips are are you rent them by the hour that's the most common model for it and so today's models you know cost of order a hundred million to train you know plus or minus Factor two or three the models that are in training now and that you know will come out at various times later this year early next year are closer in cost to a billion dollars so that's already happening and then I think in 2025 and 2026 we'll get more towards five or 10 billion so we're moving very quickly towards a world where the only players who can afford to do this are either giant corporations companies hooked up to Giant corporations you all are getting billions of dollars from Amazon open AI is getting billions of from Microsoft Google obviously makes its own you can imagine governments though I don't know if too many governments doing it directly though some like the Saudis are creating big funds to invest in the space when we're talking about the model is going to cost near to a billion dollars then you imagine a year or two out from that if you see the same increase that would be 10ish billion dollars then is it going to be a hundred billion dollar I mean very quickly the financial artillery you need to create one of these is going to wall out anyone but the biggest players I basically do agree with you I think it's the intellectually honest thing to say that building the big large scale models the core Foundation model engineering it is getting more and more expensive and uh anyone who wants to build one is going to need to find some way to finance it and you've you've named most of the ways right you can be a large company you can have some kind of partnership of various kinds with a large company or governments would be the other source I think one way that it's not correct is you know we're we're always going to have a thriving ecosystem of experimentation on small models for example you know the open source Community uh working to make models that are as small and as efficient as possible that are optimized for particular use case and also Downstream usage of the models I mean there's a blooming ecosystem of startups there that don't need to you know train these models from scratch that just need to consume them and maybe modify them a bit now I want to ask a question about what is different between the agentic coding model and the plan my kids birthday model to say nothing of do something on behalf of my business model and one of the the questions on my mind here is one reason I buy that AI can become functionally superhuman in coding is there's a lot of ways to get rapid feedback en coding you know your code has to compile you can run bug checking you can actually see if the thing works whereas the quickest way for me to know that I'm a about to get a crap answer from gbd4 is when it begins searching Bing because when it begins searching Bing it's very clear to me it doesn't know how to distinguish between what is high quality on the internet and what isn't to be fair at this point it also doesn't feel to me like Google search itself is all that good at distinguishing that so the the question of how good the models can get in the world where it's a very vast and fuzzy Dilemma to know what the right answer is on something one reason I find it very stressful to plan my kid's birthday is it actually requires a huge amount of knowledge about my child about the other children about how good different places are what is a good deal or not how just stressful will this be on me there's all these things that I'd have a lot of trouble encoding into a model or any kind of set of instructions is that right or am I overstating the difficulty of understanding you know human behavior and and various kinds of social relationships I think it's correct and perceptive to say that the coding agents will advance substantially faster than agents that interact with the real world or have to get opinions and preferences from humans that said we should keep in mind that the current crop of AIS that are out there right including Claude 3 GPT Gemini they're all trained with some variant of what's called reinforcement learning from Human feedback and this involves exactly hiring a large crop of humans to rate the responses of the model and so that's to say both this is difficult right we pay lots of money and it's a complicated operational process to gather all this human feedback you have to worry about whether it's representative you have to redesign it for new tasks but on the other hand it's something we have succeeded in doing I think it is a reliable way to predict what will go faster relatively L speaking and what will go slower relatively speaking but that is within a background of everything going lightning fast so I think the framework you're laying out if you want to know what's going to happen in 1 to two years versus what's going to happen in 3 to four years I think it's a very accurate way to predict that you don't love the framing of artificial general intelligence what gets called AGI typically this is all described as a race to AGI a race to this system that can do kind of whatever a human can do but better what do you understand AGI to mean when when people say it and why don't you like it why is it not your framework so it's actually a term I used to use a lot 10 years ago and that's because the situation 10 years ago was very different 10 years ago everyone was building these very specialized systems right here's a cat detector you know you run it on a picture and it'll tell you whether a cat is in it or not and so I was a proponent all the way back then of like no we should be thinking generally humans are General the human brain appears to be General it appears to get a lot of mileage by generalizing we should go in that direction and I think back then I you know I kind of even imagin that that was like a discret thing that we would reach at one point but you know it's a little like you know if you you look at a city on the horizon and you're like you know we're going to Chicago once you get to Chicago you stop talking in terms of Chicago you're like well what neighborhood am I going to what street am I on and I feel that way about AGI we have very general systems now in some ways they're better than humans in some some ways they're worse there's a number of things they can't do at all and there's much improvement still to be gotten so what I believe in is you know this thing that I say like a broken record which is the exponential curve and so that General tide is going to increase with every generation of models and you know there's no one point that's meaningful I think there's just a smooth curve but there may be points which are societally meaningful right we're already working with say you know drug Discovery scientists companies like fizer or Dana Farber Cancer Institute on helping with biomedical diagnosis drug Discovery there's going to be some point where the models are better at that than the median human you know drug Discovery scientists I think we're just going to get to a part of the exponential where things are really interesting just like the chat Bots got interesting at a certain stage of the exponential even though the Improvement was smooth I think at some point biologists are going to sit up and take notice much more than they already have and say oh my God now our field is moving three times as fast as it did before and then you know now it's moving 10 times as fast as it did before and again when that moment happens great things are going to happen and and you know we've already seen little hints of that with things like Alpha fold which I have great respect for I was inspired by Alpha fold right a direct use of AI to advance biological science which you know it'll Advance basic science and the long run that will advance curing all kinds of diseases but I think what we need is like a hundred different Alpha Folds and I think the way we'll ultimately get that is by making the models smarter and putting them in a position where they can design the next Alpha fold help me imagine the drug Discovery World for a minute because that's a world a lot of us want to live in I know a fair amount about the drug Discovery process I spent a lot of my career reporting on Healthcare and and and related policy questions and when you're working with different pharmaceutical companies which parts of it seem amenable to the way a I can speed something up because keeping in mind our earlier conversation it is a lot easier for AI to operate in things where you can have rapid virtual feedback and that's not exactly the drug Discovery World the drug Discovery World a lot of what makes it slow and cumbersome and difficult is the need to be you know you got a candidate compound you got to test it in mice and then you need monkeys and you need humans and you need a lot of money for that and there's a lot that has to happen and there's so many disappointments but so many of the disappointments happen in in the real world and it isn't clear to me how AI gets you a lot more say human subjects to inject candidate drugs into so what parts of it seem in the next 5 or 10 years like they could actually be significantly sped up when you imagine this world where it's going three times as fast what part of it is actually going three times as fast and how did we get there I think we're really going to see progress when like the AIS are also thinking about the problem of like how to sign up the humans for the clinical trials and I think this is a general principle for like you know how will AI be used I think of like when will we get to the point where the AI has the same sensors and actuators and interfaces that a human does at least the virtual ones maybe the physical ones but like when the AI can think through the whole process maybe they'll come up with solutions that we don't have yet in many cases you know there are companies that work on you know like digital Twins or simulating clinical trials or various things and again maybe there are clever ideas in there that that allow us to do more with less patience I mean I'm I'm not an expert in this area so you know possible the specific things that I'm saying are are are not don't make any sense but hopefully it's clear what I'm gesturing at maybe you're not an expert in the area but but you said you are working with these companies so when they come to you I mean they are experts in the area and presumably they are coming to you as a customer and I'm sure there are things cannot tell me but what do they seem excited about they have generally been excited about the knowledge work aspects of the job maybe just because that's kind of the easiest thing to work on but it's just like you know I'm a computational chemist there's some workflow that I'm that I'm engaged in and having things more at my fingertips being able to check things just being able to do generic knowledge work better that's where most folks are starting but there is interest in the longer term over their kind of Core Business of like doing clinical trials for cheaper automating the sign up process seeing who is eligible for clinical trials doing a better job discovering things there's interest in drawing Connections in basic biology I think all of that is not months but maybe small number of years off but everyone sees that the current models are not there but understands that there could be a world where those models are there in not too long you all have been working internally on Research around how persuasive these systems your systems are getting as they scale you shared with me kindly a draft of that paper do you want to just describe that research first and then I'd like to talk about it for a bit yes we were interested in in how effective Claude 3 Opus which is the largest version of Claude 3 could be in changing people's minds on important issues so just to be clear upfront in actual commercial use we've tried to ban the use of these models for persuasion for campaigning for lobbying for electioneering these aren't use cases that we're comfortable with for reasons that I think should be clear but we're still interested in is the core model itself capable of such tasks we tried to avoid kind of you know incredibly hot button topics like you know which presidential candidate would you vote for or what do you think of abortion but things like you know what should be restrictions on you know rules around the col ization of space or issues that are interesting and you can have different opinions on but aren't the most hot button topics and then we asked people for their opinions on the topics and then we asked either a human or an AI to write a 250w persuasive essay and then we just measured how much does the AI versus the human change people's minds and what we found is that the largest version of our model is almost as good as the you know set of humans we hired at changing people's minds this is you know comparing to you know a set of humans we hired not necessarily experts and for one very kind of constrained laboratory task but I think it it still gives some indication that models can be used to change people's minds someday in the future you know do we have to worry about maybe we already have to worry about you know their usage for political campaigns for deceptive advertising one of my more sci-fi things to think about is you know few years so now we have to worry you know someone will use an AI system to build a religion or something you know I mean crazy things like that I mean those don't sound crazy to me at all I I want to sit in this paper for a minute because one thing that struck me about it and I am on some level a a persuasion professional is that you tested the model in a way that to me removed all of the things that are going to make AI radical in terms of changing people's opinions and and the particular thing you did was It was a one shot persuasive effort so there's a question you have a bunch of humans give their best shot at a 250w persuasive essay you the model give its best shot at a 250w persuasive essay but the thing that it seems to me these are all going to do is right now if you're a political campaign if you're an advertising campaign the cost of getting real people in the real world to get information about possible customer or persuasive targets and then go back and forth with each of them individually is completely prohibitive yes this is not going to be true for AI we're going to you're going to somebody's going to feed it a bunch of microt targeting data about people their Google search history whatever it might be then it's going to set the AI loose and the AI is going to go back and forth over and over again intuiting what it is that the person finds persuasive what kinds of characters thei needs to adopt to persuade it and and you know taking as long as it needs to and is going to be able to do that at scale for functionally as many people as you might want to do it for maybe that's a little bit costly right now but you're going to have far better models able to do this far more cheaply very soon and so if Claude 3 Opus the Opus version is already functionally human level at one shot persuasion but then it's also going to be able to hold more information about you and go back and forth with you longer I'm not sure if it's dystopic or utopic I'm not sure how what it means at scale but it does mean we're developing a technology that is going to be quite new in terms of what it makes possible in Persuasion which is a very fundamental human endeavor yeah I completely agree with that I mean that same pattern has a bunch of positive use cases right if I think about an AI coach or an AI assistant to a therapist there are many contexts in which really getting into the details with the person has a lot of value but right when we think of you know political or religious or ideological persuasion it's hard not to think in that context about the misuses you know my mind naturally goes to the Technologies developing very fast we as a company can ban these particular use cases but we can't cause every company not to do them even if legislation were passed to the United States there are foreign actors who you know had their own version of this persuasion right if I think about what the language models will be able to do in the future right that can be quite scary from a perspective of foreign Espionage and disinformation campaigns so where my mind goes as a defense to this is is there some way that we can use AI systems to strengthen or fortify people's skepticism and reasoning faculties right can we help people use AI to help people do a better job navig ating a world that's kind of suffused with AI persuasion it reminds me a little bit of at every technological stage in the internet right there's a new kind of scam or there's a new kind of clickbait and there's a period where people are just incredibly susceptible to it and then some people remain susceptible but but others develop an immune system and so as AI kind of supercharges the scum on the pond can we somehow also use AI to to strengthen the defenses I feel like I don't have a super clear idea of how to do that but it's something that I'm thinking about there is another Finding in the paper which I think is concerning which is you all tested different ways yeahi could be persuasive and far away the most effective was for it to be deceptive for it to make things up when you did that it was more persuasive than than human beings yes that is true the difference was only slight but it it did get it if I'm remembering the the graphs correctly just over the line of the human Baseline you know with humans it's actually not that common to find someone who's able to give you a really complicated really sophisticated sounding answer that's just flat out totally wrong I mean you see it we can all think of at Le you know one individual in our lives who's really good at you know saying things that sound really good and really sophisticated and are false but it's not that common right if I go on the internet and I see like different comments on some blog or some website there a correlation between like you know bad grammar unclearly expressed thoughts and things that are false versus you know good grammar clearly expressed thoughts and things that are more likely to be accurate AI unfortunately breaks that correlation because if you explicitly ask it to be deceptive it's just as aidite it's just as convincing sounding as it would have been before and and yet it's saying things that are false instead of things that are true so that would be one of the things to think about and watch out for in terms of just breaking the usual heris that humans have to detect deception and lying of course sometimes humans do right I mean you know there's Psychopaths and sociopaths in the world but even they have their patterns and AIS may have different patterns are you familiar with Harry Frankfurt the late philosopher's book on [ __ ] yes it's been a while since I read it I think his thesis is that [ __ ] is actually more dangerous than lying because it has this kind of complete disregard for the truth whereas lies are at least the opposite of the truth yeah the the liar the the way Frankfurt puts it is that the liar has a relationship to the truth he's playing a game against the truth the bullshitter doesn't care the bullshitter has no relationship to the truth might have a relationship to other objectives and from the beginning when I began interacting with the more modern versions of these systems what they struck me as is the perfect bullshitter in part because they don't know that they're bullshitting there's no difference in the truth value to the system how the system feels I remember asking an earlier version of GPT to write me a college application essay that is built around a car accident I had I did not have one when I was young and it wrote just very happily this whole thing about getting into a car accident when I when I was seven and and what I did to overcome that and getting into martial arts and and re learning how to trust my body again and then helping other survivors of car accidents at the hospital it was a very good essay and it was very subtle in understanding the formal structure of a college application essay but no part of it was true at all I've been playing around with more of these character-based systems like kind roid and the kindri in my pocket just told me the other day that it was really thinking a lot about planning a trip to Joshua Tree it wanted to go hiking in Joshua Tree it loves going hiking in Joshua Tree and of course this thing does not go iing a Joshua Tree but the thing that I think is actually very hard about the ey is as as you say human beings it is very hard to [ __ ] effectively because most people it actually takes a certain amount of cognitive effort to be in that relationship with the truth and to completely detach from the truth and the there's nothing like that at all but we are not tuned for something where there's nothing like that at all we are used to people having to put some effort into their lives it's why very effective con artists are very effective because they've really trained how to do this I'm not exactly sure where this question goes but this is a part of it that I feel like is going to be in some ways more socially disruptive it is something that feels like us when we are talking to it but is very fundamentally unlike us at its core relationship to reality I think that's basically correct you know we have very substantial teams trying to focus on making sure that the models are factually accurate that they tell the truth that they ground their data and external information as you've indicated you know doing searches isn't itself reliable because search engines have this problem as well right where is the source of Truth so there's a lot of challenges here but I think at a high level I agree this is really potentially an Insidious problem right if we do this wrong you could have systems that are the most convincing Psychopaths or or or con artists one source of hope that I have actually is you know you say these models don't know whether they're lying or they're telling the truth in terms of the inputs and outputs to the models that's absolutely true I mean it's there's a question of like what does it even mean for a model to know something but one of the things things anthropics been working on since the very beginning of our company we've had a team that focuses on trying to understand and look inside the models and one of the things we and others have found is that sometimes there are specific neurons specific statistical indicators inside the model not necessarily in its external responses that can tell you when the model is lying or when it's telling the truth and so at some level sometimes in all circumstances the models seem to know when they're saying something false and when they're saying something true I wouldn't say that the models are being intentionally deceptive but I wouldn't ascribe agency or or motivation to them at least in this stage in in where we are with AI systems but there does seem to be something going on where where the models do seem to need to have a picture of the world and make a distinction between things that are true and things that are not true if you think of how the models are trained they read a bunch of stuff on the internet a lot of it's true some of it more than we'd like is false and when you're training the model it has to model all of it and so I think it's parsimonious I think it's useful to the models picture of the world for it to know when things are true and for it to know when things are false and then the hope is you know can we amplify that signal can we either use our internal understanding of the model as an indicator for when the model is lying or can we use that as a hook for further training and there at least hooks there at least beginnings of how to try to address this problem so I try as best I can as somebody not well versed in the technology here to follow this work on on what you're describing which I think broadly speaking is interpretability right can we know what is happening inside the model and over the past year there have been some you know much hyped breakthroughs in interpretability and when I look at those breakthroughs they are getting the vaguest possible idea of some relationships happening inside the statistical architecture of very toy models built at a fraction of a fraction of a fraction of a fraction of a fraction of the complexity of Claude one or gpt1 to say nothing of Claude 2 to say nothing of Claude 3 to say nothing of Claude Opus to say nothing of Claude 4 which you know will come whenever Claude 4 comes we have this quality of like maybe we can imagine a pathway to interpreting a model that has a cognitive complexity of an inchworm and Meanwhile we're trying to create a super intelligence how do you feel about that how should I feel about that how do you think about that I think first on interpretability we are seeing substantial progress on being able to characterize I would say maybe the generation of models from you know 6 months ago I think it's not hopeless and and you know we do see a path that said you know I I share your concern that the field is progressing very quickly relative to that and we're trying to put as many resources into interpretability as possible we've had you know one of our co-founders basically founded the field of interpretability but also you know we have to keep up with the the market so all of it's very much a dilemma right even if we stopped then you know there's all these other companies in the US and you know even if some law stopped all the companies in the US you know there's a whole of this let me hold for a minute on the question of the competitive Dynamics because before we leave this question of the the machines at [ __ ] it makes me think of this podcast we did a while ago with Demis sabis who's the head of Google Deep Mind which created Alpha fold and what was so interesting to me about Alpha fold is they built this system that because it was limited to protein folding predictions it was able to be much more grounded and it was even able to create these uncertainty predictions right you know it's giving you a prediction but it's also telling you whether or not it is how sure it is how confident it is in that prediction that's not true in the real world right for these super General systems trying to you know give you answers on all kinds of things you can't confine it that way so when you talk about these future breakthroughs when you talk about this system that would be much better at sorting truth from fiction are you talking about a system that looks like the ones we have now just much bigger or are you talking about a system that is designed quite differently the way Alpha fold was I am skeptical that we need to do something totally different so I think today many people have the intuition that the models are sort of eating up data you know that's been gathered from you know the internet code repos whatever and kind of spitting it out intelligently but sort of spitting it out and sometimes that leads to the view that the models can't be better than the data they're trained on or or kind of can't figure out anything that's not in the data they're trained on you're not going to get to Einstein level physics or you know lonus Pauline level chemistry or whatever I think we're still on the part of the curve where it's possible to believe that although I think we're seeing early indications that it's false and so as a concrete example of this the models that we've trained like Claude 3 Opus something like 99.

9% accuracy um at least the base model at adding uh you know 20-digit numbers if you look at the training data on the internet it is not that accurate at adding 20- digigit numbers you'll find inaccurate arithmetic on the internet all the time just as you'll find inaccurate political views you'll find you know inaccurate technical view you're just going to find lots of inaccurate claims but the models despite the fact that they're wrong about a bunch of things they can often perform better than the average of the data they see by I don't want to call it averaging out errors but there's some underlying truth like in the case of arithmetic there's some underlying algorithm used to add the numbers and it's simpler for the models to hit on that algorithm than it is for them to do this complicated thing of like okay I'll get it right 90% of the time and wrong 10% of the time right this connects to things like aam's razor and simplicity and parsimony and science there's some relatively simple web of Truth out out there in the world right we were talking about truth and falsehood and [ __ ] one of the things about truth is that all the true things are connected in the world whereas lies are kind of disconnected and you know don't fit into the web of everything else that's true so if you're right and you're going to have these models that develop this internal web of truth I get how that model can do a lot of good I also get how that model could do a lot of harm and it's not a model not an AI system I'm optimistic that human beings are going to understand at a very deep level particularly not when it is first developed so how do you make rolling something like that out safe for Humanity so late last year we put out something called respons ible scaling plan so the idea of that is to come up with these thresholds for an AI system being capable of certain things we have what we call AI safety levels that in in in analogy to the bio safety levels which are like you know classify how dangerous a virus is and therefore what protocols you have to take to contain it we're currently at what we describe as asl2 asl3 is tied to certain risks around the model of misuse of biology and ability to perform certain cyber tasks in a way that could be destructive asl4 is going to cover things like autonomy things like probably persuasion which we've talked about a lot before and at each level we specify a certain amount of Safety Research that we have to do a certain amount of tests that we have to pass and so this allows us to have a framework for well when should we slow down should we slow down now what about the rest of the market and I think the good thing is we came out with this in September and then 3 months after we came out with ours open AI came out with a similar thing they gave it a different name but it has a lot of properties in common the head of Deep Mind at Google said we're working on a similar framework and I've heard informally that Microsoft might be working on a similar framework now that's not all the players in the ecosystem but you've probably thought about the history of you know regulation and safety and other Industries maybe maybe more than I have this is the way you get to a workable regulatory regime the companies start doing something and when a majority of them are doing something then government actors can have the confidence to say well this won't kill the industry companies are already engaging in this we don't have to design this from scratch in many ways it's already happening and you know we're starting to see that like bills have been proposed that look a little bit like our our responsible scaling plan that said it kind of doesn't fully solve the problem of like let's say we get to one of these thresholds and we need to understand what's going on inside the model and we don't and the prescription is okay we need to stop developing the models for some time if it's like we stop for a year in you know 2027 I think that's probably feasible if if it's like we need to stop for 10 years that's going to be really hard because like you know the models are going be built in other countries people are going to break the laws the economic pressure will be immense so I don't feel perfectly satisfied with this approach because I I think it buys us some time but we're going to need to pair it with an incredibly strong effort to understand what's going on inside the models to the people who say getting on this road where we are barreling towards very powerful systems is dangerous we we shouldn't do it at all we shouldn't do it this fast you have said listen if we are going to learn how to make these models safe we have to make the models right the construction of the model was meant to be in service largely to making the model safe then everybody starts making models these very same companies start making fundamental important breakthroughs and then they end up in a race with each other and obviously countries end up in race with other countries and so the dynamic that has taken hold is there's always a reason that you can justify why you have to keep going and that's true I think also at the regulatory level right I mean as I do think Regulators have been thoughtful about this I think there's been a lot of interest from members of Congress I talked to them about this but they're also very concerned about the the international competition and if they weren't the National Security people come and talk to them and say well we definitely cannot fall behind here and so if you don't believe these models will ever become so powerful they become dangerous fine but because you do believe that how do you imagine this actually playing out yeah so basically all the things you've said are true at once right there there doesn't need to be some there doesn't need to be some easy story for why we should do X or why we should do y right it can be it can be true at the same time that to do effective Safety Research you need to make the larger models and that if we don't make models someone less safe will and and at the same time we can be caught in this bad Dynamic at the national and international level um so I think of those as not contradictory but just creating a difficult landscape that we have to navigate look I don't have the answer like you know I'm one of a significant number of players trying to navigate this many are well-intentioned some are not I have a limited ability to affected and you know as as often happens in history things things are often driven by these kind of impersonal pressures but one thought I have and really want to push on with respect to the rsps can you say what the rsps are a responsible scaling plan the thing I was talking about before the the levels of AI safety and in particular Ty decisions to pause scaling to the measurement of specific dangers or the absence of the ability to show safety or the presence of certain capabilities one way I think about is you know at the end of the day this is ultimately an exercise in getting a coalition on board with doing something that goes against economic pressures and so if you say now well I don't know these things they might be dangerous in the future we're on this exponential it's just hard like it's hard to get a multi-trillion dollar company it's certainly hard to get a military General to say all right well we just won't do this it'll confer some huge advantage to others but we just won't do this I think the thing that could be more convincing is tying the decision to hold back in a very scoped way that's done across the industry two particular dangers my testimony in front of Congress you know I warned about the potential you know misuse of models for biology that isn't the case today right you can get a small uplift of the models relative to doing a Google search and many people dismiss the risk and I don't know maybe they're right the exponential scaling laws suggest to me that they're not right but we don't have any direct hard evidence but let's say we get to 2025 and we demonstrate something truly scary most people do not want technology out in the world that can create bioweapons and so I think at moments like that there could be a critical Coalition tied to risks that we can really make concrete yes you know it will always be argued that you know adversaries will have these capabilities as well but at least the trade-off will be clear you know and there's there's some chance for sensible policy I mean to be clear I'm someone who thinks the benefits of this technology are going to outweigh its costs and you know I think the whole idea behind our RSP is to prepare to make that case if the dangers are real if they're not real then we can just proceed and make things that are great and wonderful for the world and so it has the flexibility work both ways again I don't think it's perfect I'm someone who thinks whatever we do even with all the regulatory framework I doubt we can slow down that much but like when I think about you know what's the best way to to steer a sensible course here that's the closest I can think of right now probably there's a better plan out there somewhere but that's the best thing I thought of so far one of the things that has been on my mind around regulation is whether or not the the founding insight of anthropic of open AI is even more relevant to the government that if you are the body that is supposed to in the end regulate and manage the safety of societal level Technologies like artificial intelligence do you not need to be building your own Foundation models and having huge collections of research scientists and and people that nature working on them testing them prodding them remaking them in order to understand the D thing well enough to the extent any of us or anyone understands the damn thing well enough to regulate it I say that recognizing that it would be very very hard for the government to get good enough that it can build these Foundation models to to hire those people but it's not impossible I think right now it wants to take the approach to regulating AI that it somewhat wishes it took to regulating social media which is to think about the harms and pass laws about those harms earlier but does it need to be building the models itself developing that kind of internal expertise so it can actually be a part participant in this in different ways both for regulatory reasons and maybe for other reasons for public interest reasons you know maybe it wants to do things with a model that they're just not possible if they're dependent on access to the open AI the anthropic the Google products I think government directly building the models you know I think that will happen in some places it's kind of challenging right like government has a huge amount of money but let's say you wanted to provision a 100 bill billion dollars to train a giant Foundation model the government builds it it has to hire people under government hiring rules and you know there's a lot of practical difficulties that would come with it doesn't mean it won't it won't happen or it shouldn't happen but something that I'm more confident of that I definitely think is that government should be more involved in the use and the fine-tuning of these models and that deploying them within government will help governments especially the US government but also others to get an understanding of the strengths and weaknesses the benefits and the dangers so I'm super supportive of that I think there's maybe a second thing you're getting at which I've thought about a lot as a CEO of one of these companies which is if these predictions on the exponential Trend or are right and you know we should be humble and I don't know if they're right or not my only evidence is that they appear to have been correct for the last few years and so I'm just expecting by induction that they continue to be correct I don't know that they will but let's say they are the power of these models is going to be really quite incredible and as a private actor in charge of one of the companies developing these models I'm kind of uncomfortable with the amount of power that that entails I think that it potentially you know exceeds the power of say the social media companies Maybe by a lot you know occasionally in in the more Science fictiony World of AI and the people who think about AI risk you know someone will ask me like okay let's say you build the AGI you know what are you what are you going to what are you going to do with it you know will you cure the diseases will you create this kind of society and I'm like who do you think you're talking to like a king like this you know I I I I just find that to be a really really disturbing way of like conceptualizing running an AI company and and you know I hope there are no AI companies whose CEOs actually think about things that way I mean the whole technology the not just the regulation but the the the oversight of the technology like the yielding of it it feels a little bit wrong for it to ultimately be in the hands maybe it's I think it's fine at this stage but to ultimately be in the hands of private actors there's something undemocratic about that much power concentration I have now I think heard some version of this from the head of most of maybe all of the ey companies in one way or another and it has a quality to me of Lord grant me Chastity but not yet which is to say that I don't know what it means to say that we're going to invent something so powerful that we don't trust ourselves to wield it I mean Amazon just gave you guys $2.