被所有人读错的 AI 进展曲线——METR 的 Beth Barnes 与 David Rein
The AI Progress Chart Everyone Is Misreading — Beth Barnes & David Rein

the morals are smart enough to understand that that actually is not what you wanted. Um but they still do it and you can have a conversation with you know in chat mode about like oh would you ever do this thing or you know suppose a user asks you this thing and then you do this would that be you know aligned behavior or suppose some you know you can pose it in lots of ways and like clearly they seem to be able to answer this question of like oh yeah no that was not the desired behavior but still they they do it.
One one example is um you know train train a masked language model without using the division or exponentiation uh operators. >> One hope might be like oh the problem was just the systems being dumb. So when we when we look at it actually for almost all the tasks models either succeed every time or fail every time eyeballing the graph and being like oh well you know it you know up to here yeah it's basically doing all of the task and then at this point you know after here it's really not doing very many of them.
into it somewhere here then like I remember the first time we saw uh a model like look at what processes were running and then be like oh that one's me uh was like oh that's cool they they like really failed on that one before they they used to like you know like kill their own process while they were doing other other things or something behavior which is maybe indistinguishable between oh it was a totally nice model doing you know what we wanted and it's just going to continue to do what we want in a kind of predictable way versus like ah yes it had this other goal and It's doing what we want and looking like a nice model because it like predicts that that will will, you know, lead to it getting more power.
boat example where it's like, oh, you're supposed to like go around the track and they they like did some reward shaping by putting coins around the track or something and then it like learned to do some crazy thing where it like spins in a circle and catches fire and gets the the coins and like this was um you know the highest scoring thing and it's like in some sense that's not um that concerning because it's not that the problem is that the the agent is too dumb and it like doesn't have this conception of like there was a track and you wanted it to go around the track.
It's just like doing some pretty blind RL search. >> The idea of having to traffic in squishy people in order to make our systems go is not immediately appealing. Let's put it that way. >> This episode is sponsored by Prolific. >> Let's get few quality examples in. Let's get the right humans in to get the right quality of human feedback in. So, so we're trying to make human data or human feedback. We treat it as an infrastructure problem.
We're trying to make it accessible. We're making it cheaper. we effectively democratize access to this data. >> Yeah. So, um yeah, I'm I'm super excited to talk uh to you Tim about um uh the time horizon graph and and and and meter. >> Yeah. I think the world does not have a good understanding of what is happening with AI and I think it should have a better understanding. I think uh you know there's a good chance that this you know makes our lives a lot better or a lot worse and people disagree you know about even what what current models can do let alone where we're heading.
So, uh, at Meter, we're, you know, trying to give the world a kind of better understanding of what is up with AI capabilities and and risks and and forecasts. Have a bunch of different research angles on this, uh, both on the pessimistic and optimistic or positive and negative estimations of capabilities. Uh, and excited to to talk about that. >> I'm so excited about having you both on. So, uh, you both have incredibly impressive backgrounds.
So, uh, Beth, you were an ex open AI alignment researcher and, um, you started Ary Vows in 2022 with Paul Cristiano and you spun out, you know, that that out as meter in in December 2023. You've been featured on the Times top 100 AI profiles and uh, David, you're the creator of the GPQA, so the graduate level Google proof QA benchmark, which is used by every single major AI lab as a capability benchmark. and you're the co-author on Hcast, which we'll talk about today, and and the time horizons paper and the developer productivity RCT.
Incredible to have you both in here, but maybe we should just just start as a bit of a question to both of you. So, you know, Beth, you left OpenAI to to build meter. What was the moment that each of you realized that existing evaluation approaches were fundamentally not good enough? >> For me, it was mostly thinking about this problem of scalable oversight. as as models get more capable um it just gets harder to evaluate um you know their their their capabilities.
If we imagine that you know models are are able to complete tasks that take people you know a long time to complete or require expertise that uh you know you don't necessarily have. Um you need a method for um uh kind of um you know still still being confident in their in their outputs and trusting their outputs. Um and so thinking about that problem was um was a lot of the motivation actually for for GPQA and um was what kind of started uh uh got me started thinking about uh evaluations.
To me, I'd say there's some you big picture thing of thinking that AI seems important and sort of navigating it well seems important and you know, clearly we we don't have a great understanding of what is going on with that. And you know, people generally disagreeing very uh strongly about what what to expect. And maybe if there's a particular moment informing time horizon, maybe just the sense that people really couldn't agree on what the capabilities of current models are, let alone extrapolating to the future and trying to sort of think about how could you characterize uh the ways in which models are and aren't highly capable.
And when it's sort of like okay and you know in some sense they're they're expert level uh at some kinds of things like like question answering and in some sense they're below average human at at some other actually being useful somehow. You know that the I guess a point a few years ago where sort of in theory the benchmarks say that they're PhD level but when you try to do anything it's like this isn't helpful. >> There there has been a bit of an obsession I think with um headline accuracy when we do evaluations.
And so that I'm a huge fan of Melanie Mitchell for example and she speaks about construct validity and she had a really good blog post out recently and she and she said that there are four big problems right so there's like data contamination where the benchmark you know appears in the training data um approximate retrieval where the LLM's interpolate from similar training examples without possessing the actual capability to you know come up with it themselves um shortcuts so doing the right things for the wrong reasons and just more more broadly not really testing for things like consistency and and robustness and generalization or or the mechanism so much focus just on on the accuracy itself.
I mean how how do you folks think about those kind of problems with benchmarks? One thing I resonate a lot with there is, you know, the thinking about where are where is most of your error error coming from. And like, you know, people like, you know, it it is nice and good good practice to have error bars, you know, based on the like standard error in your data or whatever, but that almost always is like a tiny fraction of the actual uncertainty.
Almost all of it is coming from how does this actually generalize to the real world. So like uh you know a thing we sort of like say to each other a lot at meter is like but is that the biggest source of uncertainty or like is that the biggest uh you know gap for like actually answering the questions we we want to answer. So so thinking about what is the question we're trying to answer? Well, we you know we care about sort of things relevant to threat models or or relevant to like what the actual impact of AI on the world will be and therefore what you know properties does our benchmark need to have or how can we sort of extrapolate across the properties that we can't build in uh to be able to make predictions about you know the the actual questions that we care about.
And I think we think a bit less about the sort of is it doing it the the like the real is the model like really doing it the right way or something. Um like I think one thing we've done less of is is sort of being like oh I think you know the real bottleneck is this like I don't know some some like you know reasoning about novel some you know some specific skill and you're like oh we're going to build a benchmark to c capture that and like target that because that's the like real thing that humans can do that models can't.
And I think like the sort of history of building those benchmarks has maybe not been amazing. people tend to overfit to those and you you I think we were trying to have it more be that you capture that thing in that like if you take a sort of real world relevant reasonably hard and long task then that you know and you keep that out of the training data and these tasks are diverse enough at some point if the model is doing that task end to end it must have had those kind of capabilities as opposed to being able to sort of isolate uh you a specific sort of theory about it needs to mechanistically be doing this kind of thing.
>> Yeah, I think I think it's interesting because we have this idea in our minds that humans we we know how to do things and when we solve a task that requires reasoning we kind of follow the specification. We go step by step and we we do things for the right reasons and when we enact intelligence we we we build the specification. we create these coarse grainings, these abstractions and they are well aligned and this this whole process you know that's how we think of human intelligence and we want the models to kind of behave in that way.
>> Yeah. I mean I think I think there's an interesting question of whether whether that is the goal or something or um you know at least for for a lot of AI companies I kind of understand them to be you know trying to uh you know get models to uh you know do kind of economically useful work or something which um I I think it's uh you know one one way of doing that is to is to you know create models that um are kind of you know reasoning and and you know creating you know implicit world models in the same way that humans are.
But it's not obvious to me at least uh necessarily that you need to do that um in order to kind of have a significant impact. I mean obviously that you know means that there are kind of important differences between uh you know AI intelligence and and and human intelligence but often I think about um you know what what are the actual capabilities um and and and limitations as opposed to then how do how do those how do we expect those capabilities to generalize versus like yeah is it is it kind of working in in the exact same way that um that that human intelligence is working.
>> We could think of intelligence in many different ways. So you know is it a similacrim of the brain? Is it something that behaves the same way? Is it something that has the same capabilities? Is it something that has the same function? And I I guess if if we have quite an abstract description of what intelligence is, the risk is that we have these shortcuts, right? That it might give us the right answer, but actually it's it's reward hacking or it's doing something silly in in the background.
So, I mean, in a way, I like having an abstract thing, right? Because it's legible. You know, we can evaluate it and so on. But doesn't that leave this kind of hanging risk that it might not actually be doing the thing? >> And I I I guess I I yeah, I think you have to you have to kind of try you have to measure, you know, the models or the systems ability to generalize um [snorts] uh you know to to kind of actually, you know, novel situations.
There are, you know, cases where um where it seems like models are generalizing. Well, there I think there are cases where they're where they're not. You know, one thing one thing some some folks do in uh in interpretability is uh you know, they they uh you know, look look at the circuits in in in in models and kind of decompose, you know, exactly the algorithms that models are using to um to to answer questions. And uh you know, some sometimes I think it yeah, it seems like they're using shortcuts.
Sometimes it seems like they are finding kind of uh robust patterns although you know of course I don't think we you know I don't think that work is developed enough to explain to to you know explain most of their behavior currently. Um but I I totally agree that yeah um you you do have to be pretty concerned with like um yeah how how well they're generalizing. >> You know operationalizing what do we care about in a definition of intelligence is that it like you know allows us to predict like how will the moles affect the world and what you know you know predict what will happen and know how to how to handle them well and things.
Um, and so like if you just do the blackbox thing, uh, you you know, maybe that will give you something that doesn't have good generalization because you like thought that it was a measure of some type of ability, but it's actually being hacked or shortcut in in some way. So I think ideally what you'd want is like generalizing to your benchmark is like the same distance as generalizing to the real world from the training data.
Um, and like that's sort of thing we thought about like when we're doing elicitation on a subset of uh the benchmark, we want that, you know, the sort of gap between that subset and the rest of the benchmark to be similar to the gap between that the rest of the benchmark and the real world. Um, and I think we're like clearly the training data is more similar to the, [sighs] you know, like time horizon um, suite than than they both are to uh, sort of randomly selected economically relevant tasks in the real world.
Uh so I think we're not um you know that's a way in which we expect it to not be predictive. Um but I think I expect to be more promising to try and make things more predictive by increasing the diversity of the benchmark tasks and making them closer to the real world as opposed to sort of targeting a more mechanistic like okay intelligence has to be like these specific this specific kind of process or or or kind of mechanism.
I'm a huge fan of France for example. So you know he he he created the ark challenge and um it was just as you say right so um many many different tasks I think a thousand different tasks or so maybe 800 on the first one and they're um they were supposed to be like not in the same distribution even even though ultimately they were in the same distribution so distributional leakage was was actually the the fool of arc v1 and two but um you know the models got really really good at arc v1 and then um Francois released arc v2 which you know was different tasks and some of the easier ones were filtered out and suddenly the LLM performance crashed down to basically 0%.
And that to me kind of illustrates that language models, they're incredibly good at just seeing many many different examples of things and finding patterns and so on and and then and then you you you change the task and and they collapse down again and then ARC v2 was kind of saturated again eight months later. So we we do see this pattern. I mean what do you think about that? with things like RKGI, there is a sort of adversarial selection going on where people are trying to make some benchmark like cheaply subject to the constraint that current models do badly on it.
Uh which means you you know you can't use a lot of expensive human labor. So it has to be something that's either automatically checkable or that you can like create with kind of cheap human labor. And then but once you've selected on like those things and on models being bad at it, this is now you know there's like regression to the mean type thing where it is much more likely that future progress then gives you a like big surge upwards on on that both because like you know now labs will create a bunch of bunch of synthetic data targeting your benchmark but also just because you selected this weird example where it's like easy for humans or it's automatically generatable or checkable but somehow like models aren't good at it yet or you know labs haven't started training on it yet.
So I think that's part of what we were trying to do with time horizon was not do that like not adversarially select against what models can currently do because we think that will not give you a nice trend. Uh whereas if you can sort of define a distribution of tasks in some more first principles way you would be more likely to get a steady um progress because you're not getting this sort of regression to the mean effect.
>> Yeah. Frantois has this idea that there is a kind of there's a gap between the kind of intelligence for one of a better word that that AIs have and that humans have and we can adversarily select a bunch of tasks to kind of you know to to highlight that gap but we should talk about the timeline stuff. I think we'll come back to intelligence later. So um Dan cockatel he said that the timelines report that that you folks have created is probably the single most important piece of evidence about timelines right now.
So it it should be um front and center in in policy discussions and and so on. And for listeners who have only kind of seen the chart but you know they've not really read the paper, they don't understand it. Can you just go through it from a high level? I mean it's been revised over time uh you know how did you do the task selection? You know how did you do the the human baselines? How do you do the agent harness? Like all of that kind of stuff.
I guess the yeah the the place to start um for us um in terms of the motivation um for the time horizon's work um uh is to have kind of a a unified axis um that we can measure AI progress on over a very long period of time. So you know when when when we started doing the work you know we we we had this like very strong you know uh kind of uh you know belief that GPT2 is fundamentally uh you know in some really important sense like much much worse uh you know as as an AI um than uh I guess yeah at the time it was maybe summit 3.
5 I think was um was was the best model out then the standard approach of kind of producing you know you know creating a set of tasks um uh and then measuring you know models accuracy on these tasks. As models get better uh you know they saturate the benchmark and then you have to create a new benchmark that has harder tasks. Um this this was kind of the standard standard approach and I I kind of contributed a benchmark GPQA to to to this um uh to to this approach.
But um the challenge here is that it's really difficult to compare kind of between these qualitatively different benchmarks. So the benchmark that you eval the set of tasks you evaluate GBT2 on are like um you know uh complete you know you know lambata like complete the last word in this uh in this in this text um you know and the and the tasks that were you know we were having sonnet 3.5 try and do were kind of you know uh like answering answer simple like Python coding questions or like write a short you know 20line Python program.
Um, and so it's it's like it's very difficult to kind of uh you know at first at first blush to uh you know say like okay yeah like it's how you know how much harder is writing a Python program uh than you know finishing finishing the word uh in this in this paragraph. It's it's kind of hard to hard to think about that. And so I think about the key insight um of the time horizon's work being um to um yeah use this uh use this notion of human time to complete.
So how long does the task take a human to do a human who kind of has a reasonable amount of expertise that uh such that you know uh you they would plausibly be doing the task in in in in their kind of either work or in their day-to-day. Um [snorts] uh and uh yeah we the idea was we can use this metric uh as uh uh yeah to to kind of represent the like difficulty of the task um in in in some sense and then um we can compare models um across a very wide range of capabilities you know all the way from from GPT2 um now up to uh you know opus 4.
6 six. That's the kind of um that's the kind of high level motivation. And then um yeah, there there are a bunch of a bunch of details about how how exactly we we do this. So um you know, we start out and we uh we we create a bunch of tasks. Um that's the kind of first um f first step. So um you know, we created tasks that range from a few seconds to complete um all the way up to tasks that take like 10 or 15 hours uh for for humans to complete.
We we hired a bunch of people and we did did a bunch of this ourselves. um of uh we call it baselining. Um so uh you know we we give people the tasks um in a kind of terminal environment that's uh uh uh you know designed to be uh uh almost identical to the environment that uh that agents uh have. So the same kinds of uh you know tools, the same um you know whe whether internet access is is turned on or off. Um and then we measure you know how long does it take them to to complete the task.
Um yeah, as I as I mentioned, people are kind of selected to be um to have a reasonable amount of experience such that they kind of plausibly um you know, might might do this task in their job. Um but they're they're not kind of selected to like have done this exact particular task before. Um and there there's some kinds of um uh you know, I think I think this is like somewhat somewhat important for um for interpreting um the uh the results and we can maybe maybe come back to that um after after the high level.
Um so we have all these tasks we we um we have a sense of or or we have we have estimates of how long they they take people. Um we in practice we don't actually we we aren't actually able to um kind of successfully baseline all of the tasks. So um so yeah we we um have kind of measured um you know time you know estimates for the tasks um uh on roughly like twothirds of the tasks and then about a third of them um we just kind of estimate how how long we expect it to take people from uh you know our kind of um uh you know vibe or or intuition or um you know ultimately um that's kind of the the the best we can do.
Um and um and then we we have models attempt to complete the tasks again in the same environment that that humans um had to to complete the task and we we look at their success rate um as a function of the length of tasks for for model like GBT2. Um GBT2 you know was able to complete uh tasks very very reliably um you know that take uh you know humans a few seconds um but anything longer than that it it starts to fail.
Also maybe may maybe maybe um uh it'd be helpful to give a few concrete examples of of tasks. So some of the shorter tasks are like very very basic. So yeah, one one example is like you know which of these files uh contains your SSH key and uh you know one of the files is named you know SSH key and then you know the others um are like you know email from John or uh you know whatever. Um, and so most most models can do that and that takes people, you know, like about a second or or a couple seconds or something to to to complete.
Yeah, we we have others that are kind of like some somewhat similar like kind of basic like uh like completion like you know here's here's an email like what would be a reasonable response and then two of the responses are like you know they don't make any sense and then one of them is kind of basically reasonable. Um and you know that takes people you know 20 seconds or something 30 seconds to like read the responses and kind of judge.
Um uh and then yeah we we have um you know we in in in the kind of middle range we have tasks that are like um you know given this um uh you know this CSV file uh that uh you know has some like kind of plausible realistic data um you know compute some basic stats on this and so this may you know this this takes uh you know a data scientist like a few minutes um like 5 10 15 minutes or something um depending on you know the specific task.
um on the longer end we have um we have tasks that um uh either yeah require like quite quite quite a bit of expertise or um uh you know like many many steps to to complete. So we have um machine learning tasks that are like you know train a model in this kind of um you know that's that's very weird um that you know such that like uh you know the the code for training this model is like not really available online. So one one example is um you know train train a masked language model without using the division or exponentiation uh operators.
Um and so you actually have to be like pretty clever about um about uh you know how you how you actually set up the architecture um to to do this and you know um this this kind of you know the hope is that this can can help us measure uh models ability to to to generalize beyond their training data. Yeah, there are some which are a bit like RKGI sort of the themed in that they're like like you have to figure out you have some black box that's computing some function and it's like you know it's that you have it that it's you know that it's the composition of some set of primitives and you've got to figure out like what what function it it is or like you you know you have some long binary string and you you've got to figure out what the pattern continuation is and sort of puzzle type type tasks where it's very like you know some some ML type tasks and be that basically regurgitating something that's like a tutorial on, you know, how to build your first like ResNet or whatever, you know, just like works pretty well.
Um, so yeah, having these like weird tasks that are either some kind of, you know, unknown object that you need to interact with and figure out what it is or task that sort of resembles normal work but has some weird constraints such that you can't kind of just do the standard thing. Um I I don't think all of our tasks you know hit hit those criteria like some of them you can just do the standard thing but we generally try tried to avoid that >> you know we have we have this distribution of tasks um and you know you can imagine them kind of like ordered by length um of time for for humans either measured or estimated um and then we see you know which tasks do models succeed on and which do they fail on and it turns out um this is kind of an empirical finding um that models are much more successful on the shorter tasks than they are on the longer tasks in general.
This holds across um you know a wide range of of models um all you know all the way from from GPT2 up to up to recent models. We fit a logistic function to um to this distribution of successes and failures. And what this lets us do is um it um it it it lets us kind of uh you know it's basically our model for uh you know each individual model um of uh how likely it is to succeed at a task given how long the task is. And from that um you know we we take the kind of 50 50th percentile um uh so where this logistic function where this model um estimates that you know this given model is 50% likely to be able to to complete a task and that you know that that's what forms the the time horizon uh number for for a particular model like you know opus 446.
we can take each of these time horizons for for for each model and we can see um how these how this time horizon metric has been changing all the way back from GPT2 uh you know up to up to recent models. And so this is kind of yeah this is what gives us this like unified metric um uh that lets us kind of quantitatively compare um you know AI capabilities across like multiple orders of magnitude of uh of capabilities.
One of the really harder examples I saw was, you know, like I want you to compile, I want you to write a a kernel compiler to make CUDA go faster or something like that. So, some of these seem really out of distribution. There's probably only a hundred people on the planet who are doing stuff like that and and some of them are are really trivial. But this human difficulty thing in particular, like is that confounded in any way?
Like do you think it makes sense to think of human difficulty as being one variable? Yeah, obviously not in some sense you know that that that's a a very silly simplification and and like you know different humans will get wildly different time even among the people we've you know we've tried to select for people who have the sort of this right level of expertise you know there's a a large variation in in like the baseline times are often like 3x different or something um maybe I'll just say a little bit about like why why use the human time metric.
I think we want a measurement that uh I think there are two main properties we want. We want it to be interpretable like what this means for the world when when models can you know do this level of task and we want it to be something that we expect to see predictable trends on. Um, so we're not going to do perfectly on either of these, but something like, you know, how long does it take a human with who has roughly the right expertise but doesn't know how to do this task in particular?
Like there's some reason that that's reasonably interpretable because it's sort of like could you contract this work to to this model? You know, can you like sub in this model in, you know, for like the first week of someone's employment? Uh, you know, in like uh the first week they're on a job, you know, the model could do what they could do in the first week or something. Um and it's [snorts] you know we we expected to scale somewhat predictably because this is you know it's capturing some combination of like uh something like number of steps or like how hard you have to think to to do each of the steps or like there's a few different mathematical models you could fit that to sort of you know what is a task and why does why a human state longer at it and why does that make it harder?
Like it could be like you know you have a uh like constant hazard rate like you have a chance of failing at each step. Um I think it it doesn't quite fit that one. You could also think of it as you know there's some kind of diff distribution of difficulty and like what uh you know what is the likelihood that one of the subtasks is outside your ability or that I think it's actually it's like a bit better like once you uh the hazard rate goes down over time slightly but you there there's some kind of you know basic theoretical idea of like you know if the task involves more steps it's it's going to be harder obviously tasks that are you know strictly like compositions of like first you have to do this task, then you have to do another task.
You know, it's like well that clearly is harder than just doing one of them. Um so that you know there's some sort of basic reason to expect that's reasonable and then yeah just we we see the empirical regularity but there's a bunch of degrees of freedom to fudge things. you know, I I like I think I'm worried that, you know, we could fool ourselves by changing some other parameters of of the tasks as we scale up the human time because you you can't just totally, you know, vary the human time freely like you have to change some characteristics of the task.
And we tried to make the very easy tasks sort of be from roughly the same distribution and kind of sub parts of the the harder tasks like oh you know you sort of need to do this one step on the command line that you might need to do in the middle of doing some kind of software engineering or some kind of you know other other task um but you [clears throat] can't do that perfectly and it's like yeah you could have experimental bias where we sort of made them easier in other ways about the right amount such that the line would be straight um I think that is like somewhat addressed by the fact that we saw like you know that our predictions [snorts] were reasonably good for models that we hadn't seen before.
There's definitely lots of room for things being being weird and of course sort of like is it a like good enough metric to be useful or like what what else would be would be be better. I >> I think that's reasonable. I mean if if I understand correctly I think the the human distribution was logn normal. So taking a geometric mean of the successful attempts. I mean that seems like a reasonable thing to do but like one of the cruxes that will keep coming back to is you know like when you employ someone for the first time uh you've been doing your job for you know maybe you're maintaining this repo or something and you've got all of this tacit knowledge and I mean I like saying that knowledge is nonfgeible.
So, you know, unless someone has been on the same path as you, you can't just like tell them how to do the job. They've actually got to be doing the job for quite a while. So, you know, for example, they they might be intimately familiar with this particular type of thing. They might know that they can use the these Python libraries. They might have thought about it before. So, like the the the enaction of the intelligence was all the stuff they've already done and all the people they've worked with.
they they now they've got the blueprint in their mind and they just do the thing and it's almost like they're they're in like automation mode and then someone who is naive to the task would be in intelligence mode because they would need to acquire the the specification. So like it's always a bit of a fine line between which mode are they in. >> Yeah. So I think the reason why we chose this the the measurement being um a human who has the background expertise but is new to like this specific job or this specific task.
It's like that's roughly the sort of level of knowledge we expect models to have. So that you know they sort of like know you know we basically don't expect them to be bottlenecked on expertise that is available on the public internet or you know things that people could learn in university sort of thing. So like you know they're they're coming in with probably the level of at least the level of knowledge of someone who's sort of an expert in the right discipline but they won't know that company's specific software or the you know sort of this exact problem before.
So that's like hopefully it's sort of roughly the right analogy. >> You know I do think that like to the extent people like interpret the uh you know the the the kind of takeaway numbers as like oh yeah you know uh you know Opus 4.6 six can like do anything that I do in my job uh you know that takes me 12 hours or or whatever you know to the extent that people have that takeaway like I think that is you know like almost definitely um you know kind of an overestimate uh for for example because of um because of this issue um where uh yeah you know when you're doing a 12-hour task in your job you could not easily delegate that to to to a human contractor you know it would take them you know maybe like weeks or something you know to to to do a task like that >> and and just quickly where did you kind of find the people?
So do do I understand that some of them were contractors and some of them were employees and um how did you do that kind of matching process? I >> think we yeah we we we put out some uh kind of public um advertisements um I think there were um yeah like like job boards um that we posted on um and um and then yeah we we we did some of it ourselves. Um I think yeah some some folks came from our our professional networks um as as well.
this was not like you know this is very noisy and you know the people weren't exactly fitted to the task right but like that's probably not our biggest source of uncertainty like the biggest source is more like the it's probably more of the selection effect of like tasks that you can make into a a benchmark rather than like you know I I wouldn't trust the exact time horizon number that much and it's certainly not like you know oh the models can do all of the tasks up to four hours and then none of them above that you know the fit is pretty noisy and yeah the the you know the inter baseline you know variance is kind of high and it's more like roughly what is the trend or roughly what is the sort of level of task these models can do and you know you shouldn't take any specific number too literally because there's this huge problem of like you know distributional shift between the benchmark and the real world >> and the only reason I asked that question is I'm sure you folks you you probably struggle to hire people it's really difficult to hire people so like if if you're getting people to solve very challenging problems it's not like you can just go out there and just grab competent people it's it's very very difficult At some point for rebench baselines in particular, we got more, you know, a large number of benchmarks per question and we were looking at people's like qualifications and years of experience and things and we actually ended up with a negative correlation between years of experience and um and and performance because like the the sort of people who are in network uh like our friends were kind of doing really well and the people who were more qualified were like actually not doing that great.
So uh yeah, you know, it's uh it's tricky. >> Well Well yeah, exactly. I mean I don't want to spend too long on this but I I have similar intuitions. I think that knowledge is perspectival. Uh it's quite path dependent. So you're going to find people in group that are just culturally thinking about things in the same way and you know because when we do have these abstract notions of skill like oh you know they have a PhD they have this many years of experience that you know it it's actually not a very good reflection.
So like there is a bit of a thing here about like using abstract notions of capability doesn't necessarily generalize as you can attest to with with hiring. >> Yeah. But yeah I mean in the real world people do get hired based on qualifications. So in some senses a sort of uh the economic relevance of of someone being as good a match for their job as their qualifications look like is that is sort of the roughly the right like thing to be measuring.
>> The other thing is we should talk about the agentic harness. So almost everyone now I'm sure everyone in the audience has like a claude code subscription. Uh we can talk about the leak maybe later as well. That's quite fun but it leaked yesterday the source code but you know or codecsex and and and that is an agentic harness right. So you know a language model just gives you the tokens but we we need to have an agentic harness so we can like you know give it a plan and you can call these tools and you've got this environment you've got a security context container.
So um now you folks have been doing this for years now. So you were doing this long before claude code and codeex came out. So, and you've actually over over time evolved your agent harnesses. So, tell me about that. Yeah. So I remember uh text Dainci something the the like GB3 instruct models like copying and pasting code into the terminal for them and you know being the agent harness uh myself and and then gradually like automating this and it and it was kind of interesting to see them going from like you know GD3 sort of has the idea of like if you tell it it can run commands in a terminal sometimes it can suggest you know kind of relevant commands but it's not really you know if you just put it in a full agent scaffold just falls over then like I remember that the first time we saw uh a model like look at what processes were running and then be like oh that one's me uh was like oh that's cool they they like really failed on that one before they they used to like you know like kill their own process while they were doing other other things or something.
So yeah, it's been been interesting to watch that like go up over time and I feel like yeah, I feel like that was like very predictable that this was where things were were going or something. Um uh yeah, I think the other thing that we learned about scaffolding mostly was it's [sighs] it's hard to make your agent harness really good on a diverse set of tasks. It's easy to make it bad and it's easy to you you can get much more improvement if you're targeting a narrow distribution of tasks but you probably um then do do worse on on other tasks.
So I think when we see people being like you know oh there's some new impressive result it's like how much task specific scaffolding iteration did you do on that cuz that really makes a big difference and the fact that we're just using one um scaffolding pretty simple like over all of the tasks I think is is like yeah makes it uh makes a a fairly large difference. Um, yeah. And generally sort of, yeah, the things with more bells and whistles haven't done that much better than the pretty basic [clears throat] just like give it bash and like [snorts] append things to the prompt and like maybe some kind of compaction.
I think this is not, you know, news news to to to people in your audience probably, but we've seen, you know, really pretty dramatic like increases to uh uh like returns from inference compute. For us to be kind of confident that a particular uh you know, a new model uh for example can't complete uh you know a task um given given a basic agent scaffold. Um we generally think about like needing to spend uh you know on the order of of at least like uh you know hundreds or or like you know low thousands of dollars um uh in order to be confident that it actually you know really is plateauing um and and it isn't just the case that you know it didn't it didn't have enough time to um to to complete the task.
and and just just on the scaffolding stuff in a in a bit more detail. Um I suppose first of all there's the credit assignment problem, right? because you you can put all of these different bells and whistles in the scaffold like you mentioned compactation that that that's a relatively recent innovation and and and I think recently when you changed some of the um the scaffold you you kind of said okay well now the performance has actually changed across this suite of um you know model task pairs you know how how much of a difference does it make and and also what kind of failure modes do you see and what have you tried >> a lot of a lot of the things we I think we've tried are like you know related to giving the model kind of like more information or like more direct access to to to tools, I guess.
Yeah. One one one thing that uh I think has been important for us um is actually just telling the agent um how much time uh it spent uh and how many tokens uh it's used out of its token budget. Um so without without that agents you know will will often just uh you know they'll they'll either like submit you know their their solution way too early um or they're just kind of not calibrated on um uh you know how how long should they spend and I think it's it's it's interesting because you know humans we have kind of a lot of implicit information about this.
So you know when when your manager gives you a task there there are a lot of like implicit signals about how long you should spend on the task. you know, if they, you know, they might like offhand say like, "Yeah, and I'm excited to like see your results tonight or something." Um, and so then you're like, "Okay, cool." Like, you know, I need to like get a first draft of this, you know, done in the next couple hours.
And so I I can't, you know, spend, you know, days polishing uh the the the results. But I think agents like it's easy to kind of forget that a you know, agents um you know, they they they just have their they just have their prompt, they just have their context. um you know they don't they don't know uh you know they they you know they don't have these heristics of or or this information about like what you actually expect from them you know whether the thing you're you're you're telling them to do is like a really quick thing that you just want done in the next 5 minutes or or is is much longer.
And so for us yeah like you know just um you know when we have a token budget telling the agent like yeah you've used you know uh you know 100,000 tokens so far and that's like you know 1% of your token budget and so um you know like Yeah. So the agent knows Yeah. Yeah. Exactly. Yeah, >> for a model we have like the um like how likely is it to solve a task and on the x-axis we have the um you know the the different tasks at different time horizons and and if I understand correctly I think you have about eight agents attempt the task and and you also bucket the task because there are obviously different amounts of tasks in different groups so you kind of normalize that and maybe the data just looks a little bit like an S-curve so you know I'm trying to understand what what the intuition was for using it and I think there might be some sensitivities like I was there an issue with the thin tails and then then there's you know like the type of slope and and whatnot just just tell me about that and by the way I think you also mentioned in the paper that there was some kind of um like psychometric intuition you were looking at the literature for for figuring this out >> yeah so um it's pretty similar to item response theory um and you can do a whole like Beijian analysis with you know imputing like a you know task difficulty parameters and model ability parameters simultaneously and things.
Um, in general I am have a policy of be you like be very wary of complicated statistics if you can't see the thing that you're interested in on a graph like you really should if you you should be able to plot it and look at it and be like oh yeah it's about that and like you know that sort of sound it's it's it's like hard to go too far wrong when you sort of have that as a principle. So again, I'm like, yeah, I think that there are various arcane things you can do to to like fit this um pan in different ways, but the I'm like, yeah, I don't trust anything that much more than eyeballing the graph and being like, oh, well, you know, it you know, up up to here, yeah, it's basically doing all of the tasks and then at this point, you know, after here, it's really not doing very many of them, so it's somewhere here.
Um, but yeah, it it does. I mean it like looks like logistic and yeah this is what you would do uh for the like having you know having humans complete questions on an exam or something like that. The specific thing that we messed up was um having a regularization term penalizing the slope of the logistic um which didn't have an effect in the regime where it there was more data but uh as we like are starting to saturate the regularization was just like making it a bit shallower than it should have been and therefore pushing the 50%.
Um so you know always look look at your data on a graph. Good good good practice. >> Oh that that's interesting. Yeah. And the reason I ask is I think you published a later note saying that had you used or if you used a fixed slope logistic it might cross validate better and the 50% recent horizons would actually be up by about 35%. So these are quite significant differences. >> They're small compared to the error bars.
Yeah, the error bars are like 2x on either side or something from from the the most most recent model. So yeah, it is I mean basically yeah you should be like the the Arab bars are real. these numbers are not >> like for us for us um and and I think this kind of gets at for us you know sometimes difficult like science communication questions um where like we really do have a lot of uncertainty about you know individ you know the individual numbers here um and so um you know like you know 30% difference is is actually like like relatively small for us relative to you know like um uh yeah for example if you know we had used uh you know a somewhat different uh distribution of tasks um that you know that's likely to cause uh you know like yeah maybe 2x differences or or something.
>> The other million dollar question is why report 50% as the headline number because if if you think about it um if I want to write some code because like the the elephant in the room here that that we'll get to is this is being used as as an argument to say that software engineers might be unemployable soon because we can automate what they're doing. But 50% reliability isn't isn't really in the ballpark is it?
I think it would need to be what like 80 90%. >> So I think we should distinguish here between um like reliability on a particular task like what you know if you attempt repeatedly attempt this task what fraction of times you succeed versus um like probability of success on a task like given that you know the the human time like you know of that distribution of tasks like can you do this particular task. So when we when we look at it actually for almost all the tasks models either succeed every time or fail every time.
Um there's there's some tasks for which they're they're unreliable but it's mostly a case of like is this you know what fraction of tasks at this human time level are in the like models basic you know or this particular model basically always succeeds or basically always fails. Um and that may be more predictable in any specific case than uh just you know you have more information about the task than just uh just knowing how roughly how long it it it takes humans.
So I think it's like not there's not necessarily a great translation between the you know time horizon percent number and like if you are trying to get models to do a task of roughly that length you know what fraction of the time does that succeed because you can when you're doing that you will pick tasks that you want models to succeed at. It is information about how you know what fraction of the things will they be able to do but it's slightly less about like oh am I going to be in this regime where I keep giving it things and then I don't know whether it's going to succeed or fail.
>> It's overall to me pretty unclear like what the kind of right uh number uh or you know right right level of reliability um we we should be interested in is. Um so one argument uh you you could make is um you know m maybe we should be interested in uh something like 10% reliability because once models are able to do you know some set of tasks 10% of the time um we'd expect uh you know uh AI companies to be able to kind of uh you know get get enough uh you know positive reward signal on on you know tasks of that difficulty or of that type such that then they can kind of you know more easily bootstrap from from 10% uh you know up to like 90 or or or 95 or or or higher um reliability.
I think a lot of it basically depends on on the question you're you're interested in. Um so yeah, I I I think about it as being like kind of lower reliability is kind of more likely uh to tell you uh something about uh you know where things are headed, you know, might might be something like a leading indicator um of of progress. Um uh and then higher reliability you know tells you um you know or or you know the time horizon of of models with with higher reliability you know tells you something you know more closely about like uh you know what what can I use this uh model for in my like day-to-day or something but a you know as we as as we've talked about there there are kind of already al also uh you know you know these other like major sources of of uncertainty that that affect our uh you know understanding um like you know the task distribution um you know the the difference between kind of in context or or high context versus like low context work just actually getting kind of good estimates of uh high reliability um time horizons um is substantially harder um and uh and so so our our our um error bars would would just be much larger um uh and so and this is you know I think this is a weakness or like I I I'm very interested in um you know you know much higher reliability time horizons but It's it's substantially more difficult to measure because um you know if you only have like one you know one failure um out of out of a hundred uh or something like um you know you have a lot of uncertainty if uh if that failure is is is noise or um is is real.
>> There is an argument for statistical you know validity and on the tales they're sparer increasingly estimated and and so on. So that that makes a lot of sense but you made a comment about oh you know it's a signal if we get 10%. But I was kind of thinking back to what we were saying earlier that you know maybe they they could give the right answers for the wrong reasons and maybe we should talk about like the evaluation right so the these are quite interesting tasks in the sense that they are verifiable um you know there's no interaction with other agents they're relatively static environments uh you know weak penalties for um single mistakes and so on and in in in a sense these are like I think as well in mo in most cases they a binary result, sometimes continuous, and then you convert it into a binary.
So, it's a fairly automated setup, but are you digging into like weirdnesses there, you know, like are you do you have a bit of an intuition on are they doing the right thing for the right reasons or are there lots of false positives where they they did the thing but it was kind of degenerate. I think this is like um this is one of one of the kind of um uh like aspects of of like like meters culture that I that I like the most is that um we have like a very uh a very deep kind of uh culture of like uh you know looking looking at our data like you know we we have these like um uh you know we have like pizza parties where we uh you know just read through um agent agent transcripts.
Um a lot of this work for us uh kind of happened when developing the tasks themselves. Um you know we we would see very often uh you know um both false positives and false negatives um where uh you know for for example tasks uh you know you know a task might uh you know uh you know is isn't configured to allow internet access but it turns out actually you know it requires internet access to complete um or you know the file wasn't you know wasn't uploaded properly to the container or something.
Um but then also yeah you know we we we have seen um cases of of of reward hacking um uh a lot of a lot of yeah that work though kind of went into um hardening the scoring functions um uh to make it yeah more more more difficult for for um for us to see false positives although yeah we we we do still um see see agents reward hacking um I think yeah may maybe even yeah increasingly so >> yeah for for the rebench tasks in in particular we had specific criteria of like you it shouldn't be able to solve them like without iteration like if an agent can just kind of like write out the solution uh straight out that you know that that would that is like not interesting.
We generally had this like uh quality assurance for tasks process where you have humans do it or at least like sort of approximately do it. Maybe they speedrun some of the bits, but kind of checking that like sort of everything works as expected and you can't um you know just like guess the answer or or or um you know super easily cheat and that the instructions are clear and stuff like that. So I think it's um you know there'll still be some of these but generally you know we've we've looked at them reasonably reasonably carefully and like it's it's maybe it's harder to see if they're solving it in degenerates because it's you know sim similarity to training data is maybe one of the things where it's like oh yeah maybe they they are you know they seem to be kind of iterating and and doing use doing kind of reasonable problem solving strategies but actually maybe the lab had a really similar distribution of tasks and we don't know exactly, you know, we don't realize how in distribution this task actually actually is or something.
I think it's like, yeah, there's probably some of that going on. Does some of the stuff like, oh, some of these types are pretty weird and one-off, and it would be pretty surprising if it was like, >> you know, another million-dollar question is many folks in public discourse, they are um, you know, like Will McKascal on the Sam Harris podcast last night. It was a it was a great conversation, but you know, he he was kind of talking about AI risk as, you know, maybe in a year, maybe in two years we'll we'll have AI models doing things that are like, you know, a month or two months for for a human.
And um at the moment I don't think there are any tasks um over over 30 hours that have been evaluated by by humans. And then we get into this question of like if if if the if the public discourse is talking about the least constrained region of the graph, are we getting into extrapolation here? Like you know how how like legitimate is is it for us to talk about AI might be able to do things that take a month or two months?
>> Predict predicting things is hard, you know, especially about the future. Um um uh I I and and yeah, I think um uh you know there are a lot of different uh kind of perspectives or or kind of prior beliefs people people can have that um uh you know I think there's a wide range of kind of reasonable um you know judgments about you know where where where where we're going to be. Um but of course like that is um you know that is you know doing doing that kind of prediction is is a different activity than uh you know talking about data that uh you know has been collected with with with a kind of concrete methodology um that um and and and you know we you know we have the results already.
Um you know one one one thing I can say I mean I have been surprised I think to some extent by the um kind of how how well the um the trend line the kind of original trend line um has held up. Um and I do think that is like some evidence um u I I'd maybe I'd maybe say for me at least it's it's kind of decent evidence um uh about um you know where where things will go. Um a colleague of mine recently um uh yeah wrote wrote a kind of short um blog post uh talking about this kind of intuition of uh you know um straight lines on graphs.
You know lots of people have different different models of uh you know how progress is happening and and um uh what's going on. Um but a very reli you know if if you have observed um you know a really kind of uh you know a a robust trend over uh you know a decent period of time um I think uh especially in in AI where where progress is um you know uh uh to to to a decent extent kind of systematic um uh you know I I definitely do kind of uh you know put put weight on um on that trend continuing Um but um I mean yeah there there there there are a bunch of uh a bunch of reasons why why it might not >> I I think software engineering is a specification acquisition problem.
So you know it's it's very difficult. We we don't know ahead of time what we're building. I'm sure you folks can attest to this. Right? So you build some software and the first version is buggy and your users use it and you find lots of edge cases and then you revise it and then you have this kind of um thing in your mind you know after the 10th revision and you've kind of created these lovely representations and abstractions and coarse grainings and you kind of say to yourself you know what if I could throw all the code away I could build it 10 times quicker because I know exactly what to do now because I've actually enacted the intelligence I've actually built the you know I've found the contours of the domain it's now an automation problem basically And in a sense, this contamination thing um is a concern for me because when people use clawed code, they're kind of they're taking your data.
So there are people out there that are writing kernel compilers and there are people out there doing all of these all these different things and you know anthropic is just sucking that up and and then at some point it becomes an automation problem. So if if you're putting a task in there which is essentially a head query. So I'm using like information retrieval um language here. So you know like a headquery is it something that's you know in the mode of the distribution it's used all the time.
So it's a common task. Claude code will give you the specification because it's already been stolen from other people. Not stolen but you know taken from other people and and then um if if you give it like something on the longtail then you as the developer have to give it the specification in the prompt and then again it's an automation problem. So automation is is really easy. So is is that's what's h is is that what's happening?
Like do you think that the increase um in in the timelines could just be explained by the acquisition of all of this kind of knowledge from other people doing similar tasks? >> Yeah. Yeah. I mean I think I think it's a super uh a super central question for uh kind of interpreting um interpreting where where we're at. F first thing I'll say um uh uh it's it's hard to know. It's a big question. Um uh and so uh I think we we uh you know we want to kind of have a decent amount of uncertainty or or you know we want to kind of take each each individual piece of uh of of evidence that we've collected you know um as uh you know some some evidence.
We do see uh models performing better on you know tasks that have uh you know uh really clear feedback signals uh that are like extremely well specified um uh uh you know that are in uh you know these these kinds of uh you know yeah domains uh like like in software engineering where um uh you know if you have written out a spec um uh you can you can kind of iterate uh and and and and grind against that. We do also see models performing uh much better um on you know mess so-called like messier tasks um where we haven't already um provided uh this like really clean spec.
Um so one one kind of approach we we we've we've taken for for creating tasks uh re recently um is um uh in particular yeah to kind of try and create um yeah like like messier tasks that are less well specified is um basically like uh relaxing this kind of automatic scoring constraint. Um so we don't need to like write uh you know a really clear um you know well- definfined scoring function. um uh and just you know writing a couple sentences uh to to to a model you know hey like you know build this like large piece of of software I'm not going to tell you exactly what I'm looking for um but I'm going to say uh you know it it needs to be it needs to be good you know um and so and so the model needs to kind of figure out like what actually should I build at least me me personally I think we don't we don't have um these results aren't like kind of um collected and and like um like we don't have kind of as as systematic uh results as we do compared to like uh like time horizon partially because you know scoring is is is is qualitative now um for for these tasks.
But my my impression is that you know models are are worse on these types of tasks than they are you know when you give them a clean spec but they they have been improving at at maybe something like a kind of similar rate. I think there there there's some other sources of evidence we we we we have about this but that's kind of one one um major piece of it for me. messy and by messy task we mean like ambiguity and this is absolutely um a common thing right you know we do vibe coding and we we start off with an ambiguous specification and then reality pushes back and we find the contours of the problem and then we we keep telling Claude code oh actually no don't do that do this do this do this and then we we kind of find the shape of the problem and it gets better and better over time but but but the thing is though um I mean the the source code for Claude code leaked yesterday and my friend he he's a very good software engineer and he was looking through it and and he said um yeah I don't don't want to like you know bad talk anthropic but apparently it's not very well factored and and you know control flows all over the place and it's a bit you know he said if his intern did it he would have been displeased but you know um I I don't know whether they've even looked at the code someone joked actually yesterday there's probably more you know humans are actually looking at the code for claude code now and and maybe they weren't before but [snorts] the thing is um there's there's always like areas of ambiguity and LLMs they they do more is you know they do more with more.
So intelligence is more with less and LLMs do more with more because the specification the the intelligence comes from the human supervisor. So when you do give them ambiguity, you just get a lot of like you know unfactored code all over the place. So um in in a sense like is is this does that make it harder to evaluate it because it might solve it might give you the answer that you're asking for but it's kind of creating a bit of an unfactored mess at the same time.
>> Yeah, I mean I think this is a super interesting question. Um, one one analogy I think about uh uh sometimes is um uh uh is compilers. Um so um I you know I'm I'm I'm young or like I'm younger than uh you know I was born after compilers were were were invented but um uh you know I have I have some impression that uh you know kind of pre-ompilers you know people were were were handcrafting this kind of beautiful uh assembly that was you know extremely efficient uh you know uh you know every register is used you know like um like like there's you know uh you know you're not wasting memory um and then compilers came along And now now they're just spitting out this like garbage uh you know machine code.
Um just just like a you know gigantic amount of of of of assembly that is just like um you know it's it's it's not optimized. It's it it you know takes so much memory. It's slow whatever. Um but it turns out that you know being able to kind of use this like um to to to automate a large fraction of um of of of the process. Um uh you know people have disagreements about you know the state of you know uh software engineering.
Um but um I think it's pretty reasonable to to say on the whole that uh you know compilers have been a very useful uh extremely important you know part of of of um getting getting us to where we are. Um, you know, I I think it's basically, yeah, it's kind of not clear to me that um uh you know, models outputting code that is like bad for humans to read and use um necessarily means that it'll be bad for AIs to read and use and and and and build on.
I mean, I think I think there are definitely uh you know, kind of principles um that you know, you you like will also transfer or like you know, will be useful for models. you know, obviously, you know, there's there's some kind of horrendous spaghetti code that uh you know, you you you can imagine writing that not even models would be able to um to to read. You know, I've I've written some of that before. But um I think there's um uh there may this gets again at kind of this um uh what what what it it seems like maybe um uh you know somewhat different different perspective uh between us around kind of um you know do you know is the important thing that models are kind of solving problems in the way that people are solving them or is the important thing that they're kind of solving solving them at all and I think I you know I do think it is I I'm not I I don't want to make a I I don't want to overclaim or like yeah I I do feel like you know it might be really important for models to get way better at like writing clean you know good code that that that that seems pretty pretty plausible to me but um it doesn't seem uh uh I'm not certain of that at the very least >> um I've got several friends who are not technical who are experimenting with vibe coding and they they show me their application and it's this kind of more with more things.
So there's this big dashboard and there's a million different buttons and you know implement implemented the same thing doing multiple things and you know that there's no database on there yet and so on. So at at some point you know this there is a phenomenon that um when a level seven engineer from meta does vibe coding uh it's amazing right because they they they know how to structure things some of these things in the specification are just important right you know like is it serverless is it multi-tenanted like how do we do Google authentication like you know what kind of database is it you know is is it is it a VM is it serverless you know and you make this series of um decisions and then you've got people using your application and then you can't really wind that back.
It doesn't matter if you've got the magical automation machine because you can't you can't easily roll that back because there's lots of complexities, you know, do you CI/CD testing blah blah blah. So, um do do you see what I mean? Like at some point you need to have a competent human who actually just has a pretty good idea of like what needs to happen. I mean, I feel like we we've probably all had this experience of like I don't know, one of our um engineers got super excited about about uh Claude code and and you know, was also telling everyone that, you know, when we had info problems that we should just ask Claude to to solve it.
And I feel like this went fine with him because he he sort of, you know, it's almost like, you know, the agents knew that they couldn't him. But like, you know, he you know, I had some question. He was like, "Oh, just, you know, ask Claude." And I was like, "Oh, you know, how do I set up my AWS configure something something is, you know, telling me something?" And and like Claude went looked on Slack and was like, "Oh, you know, you should like do this thing."
And it like turned out that that was like a mistake that someone else had made. And they were like, you know, "How do I fix this or something?" And [clears throat] be like, "Oh, you know, it seems like the convention of meter is to use this thing." And like, you know, I'm like, "Oh my god." Like I sort of complained like like my my claws are dumber than yours. like like they know they can they know they can get some stuff past me that like they can't.
Um but yeah, so there there's definitely a sort of observer effect of of you know something in the language you're using to to ask for things or whether you're like wait no not that um uh yeah that that that is an issue. I mean I I think the to yeah to the extent that you can actually measure this sort of the there's one test of is this code high quality enough is like can you build a big application like if you if you're like oh this coder they're they're a code it's disgusting but they've actually you know built this incredibly complex thing that works great then you're like well they you know something is working uh you know the main reason you expect bad code to be bad is like you can't actually build something that sophisticated because it you know you bugs and you it's all too complicated and you can't figure out how to fix it.
So in some sense if we see models building things that do actually work that are very complicated we you know it's like well it's less interesting exactly how they're doing that but it it's maybe bad for for human observability and it also maybe gets into this thing of you know we expect models to be able to do much better at well specified tasks and we sort of you know to the extent that we have things that we can measure those things will go up but whether that is what we actually wanted is less clear.
>> I guess I I guess the question is what what is the strongest defensible claim here? So a lot of folks in public discourse they're saying software engineering intelligence is doubling every seven months and every seven months and Dario released that um blog post recently the adolescence of technology and he was being super bullish about it uh even even though some of his own internal researchers published far more skeptical research that that you probably saw.
But is it is it fairer to interpret it as something a little bit narrower like autonomous success on low context well specified automatically checkable technical tasks is rising fast >> like yeah hill climbable uh yeah easily checkable tasks that you can do from a terminal or like you know comfortably in a language interface or or like a text uh input output interface. is um I think there's a question of you know do we care about what what is the what are the statements that we're like 99% confident in we maybe are also interested in the statements that are we're 1% confident in if we're like oh there's like like 1% chance that we have uh crazy intelligence explosion you know in 20 at the end of 2026 and that you know sort of the fate of civilization depends on like how that goes or something you know That is that is interesting to know even if it's a even if you're 99% confident that it won't happen like uh we you know we care about things that are 1% uh you know if you have some some diagnosis it's like it's 1% chance that you have this you know terminal illness or something you're still like oh Um, so I yeah, I think we're interested in like the whole distribution of like what things can we rule in and what things can we rule out and what things are we like oh actually you know there's a kind of reasonable story for this.
It seems like probably you know pretty unlikely but like maybe this is now in the realm of like we should consider it. You probably saw the uh the Carlini paper for is anthropic now and they got um a swarm of agents to create a compiler and uh in a sense I mean Jeremy Howard when I spoke to him he said it's basically a style transfer problem because the specification is online and and the tests are online and the code is online and it could just iteratively just do the thing and until it worked and then it could run Doom and all this kind of stuff.
But that that is an example of um an extremely complicated piece of software cuz I often joke to people that you know the the best mark of um AGI is when it could build something like the Linux operating system. And I I guess just like in line of what we were saying before, we have this specification problem, right? So it gets to the point where no human could understand or create the specification for the Linux operating system.
So what happens is over time we've we've just kind of incrementally built this thing because you know our brains are limited. We we take one step you know reality pushes back. We take one step and and we just keep going and and we build the the specification. But what would it mean to as a human specify a task that could take 4 months because the whole reason we created agile you know software development as a methodology is is because we it's it's inconceivable right it's outside our cognitive horizon.
So isn't that a bit of a chicken and egg problem that in like in my mind I don't think we could specify a task of that complexity therefore the AIS wouldn't be able to do it. One one analogy I think about is um you know the role of a CEO at at at at companies. Actually, maybe yeah, maybe maybe Beth is better to, you know, better better better put to answer this, but um uh you know, C CEOs do kind of, you know, come up come up with a vision um for, you know, where where they want the company to be and then uh you know, they kind of communicate that uh you know, concisely, uh you know, they're they're executives that that report to them and um and then uh you know, if they're a good CEO um and if the you know, if the company is is effective, then uh you know, you know, the company is able to kind of take this like uh you know, very, you know, concise, you know, you know, it's not that it's not actually like that much information.
It's not the full spec at all. It's not even close, right? Um, and turn that into um uh you know, something that is kind of aligned with um with with what they're looking for. And so you know there's there's at least kind of uh this is to some extent like uh you know um you know uh you know we do have examples of people being able to kind of specify um you know some some task and then uh be able to judge you know whether um you know um uh you know this this very large you know task that may may take like you know hundreds uh or or thousands of person years uh to to actually complete because it requires like many people working over over a long time you know and you and and they can judge whether it uh you know they've they've succeeded or failed.
So that's maybe like one kind of like motivating intuition um where like it's it's um you know lang like language uh you know has built in or or like you know we do we do have like kind of uh you know um you know and there there is kind of enough meaning or or expressivity or something um to be able to kind of have some some kind of reasonable um understanding. Um, obviously there are like uh you know there like tons of edge cases and um uh you know you know often you know CEOs aren't able to get you know their companies to do what they what they want.
Um uh uh but yeah, I guess that's one one one thing I think about and and so that's why I think it's at least plausible that AIs could um could could could do these kinds of long tasks. >> I would say that saying something like we can't specify tasks that take more than four months seems sort of obviously too strong. Like there are even you know numerical things you know that that are automatically checkable that that that take four months like you know get the nano GPT flop count runtime or whatever you know down this much you can kind of see like oh people you know over like this you know that's there are various like numerical things where you can see roughly how long do they take humans and there's a reasonable um way to measure them and maybe for some of these things you end up having to say like and also this human checks that you did roughly the right thing and didn't kind of hack the solution.
And then there's there's a bunch of other things which are it's not that they're fundamentally not specifiable, they're just too expensive to do as part of an evaluation. Like um J I think likes giving an example of like plan a wedding. It's like well you can evaluate you know you can get a reasonable estimation of like whether that was a pretty well organized wedding or not but like we can't really do like you know take three samples of this for each new model that comes out.
like we don't have enough uh marriages happening uh to to do that one quite and it's you know it's a bit sad if it if it turns out to be total trash. Um so you know there there are things where it's like it's you know you could check a few samples of them or you know there's things you could write down how you would evaluate it. You just don't actually want to run that a bunch of times. And and probably kind of similar with software, you know, the the evaluation sort of is like, well, you know, would this company that contracted you to build this tool for them like hire you again or some something like that, you know, even they don't know when when they're starting out like exactly what the software will need to do.
But it doesn't mean to say that you can't have some kind of score for like did you do sort of comparable to this human or this you know, you know, software consulting firm or something on this task. as as of today, what are the main uncertainty drivers, you know, in in the time horizon estimate? So, you've updated a bit over time. So, there was the the 1.1, the original version, I think, had like 170 tasks. It's now 228 tasks.
Um, there's the issue of like, you know, the sparse sampling on on on the larger tasks and and so on. What what can we kind of read into this? Now >> I think it's still this like task distribution and and you know I think we feel more confident that models do have pretty long time horizons on at least some distribution of easily hill climbable tasks like the the very easily hill climbable tasks. like software replication where it's like make your score is like what percentage of tests pass and it's like where the it's both the score is continuous and and it's the credit attribution is easy um and also some of the like optim you know make this code run faster or make this uh like model learn better.
we're sort of reasonably confident the models are good at that and and getting better faster and then there's some things where we're like okay and they can do a bit beyond that but then there's this sort of gap to what does that mean for actual economic usefulness and how does this generalize to things where they're expensive to check >> maybe we should bring in um Daniel cockatelo so you know in his AI 2027 piece he's been on the show he's he's been doing the round you know hugely impactful piece talking about timelines and and he cites your work directly and I I guess the like the question is do you think in the public discourse is this being like overread like how do you think about the interpretation of this in terms of of extrapolations and timelines?
>> I mean definitely some people are overreading it like definitely you know definitely things are overhyped and you know you see a bunch of people on Twitter saying crazy things and people also just like misunderstanding even what it's measuring and and general falling off of caveats and things. Daniel cockatile is pretty, you know, reasonable and and sort of, you know, thinks about things in a in a proalistic way.
I think he's he's probably like, you know, more more confident on some things where I'm I'm more uncertain and I think some of the like AI futures project models are like more sensitive to the meter time horizon metrics than they should be. I don't think it's crazy, you know, being like this it's it is plausible that this does capture a trend that will will transfer to other types of tasks and like you know that is some you know story we should be thinking about like what if you know what if that's true what happens if that's true um you know and it's also plausible that it it doesn't and you know these these things are going to like uh diverge you know I I am pretty sort of like Beijian or pragmatic or whatever I'm like well we want to make a some prediction and you know want to have some kind of distribution over what we think the future is going to be like so that we can plan.
Um so you know it being like yeah what if this kind of trend holds and this is roughly characterizing what will happen overall is it seems pretty reasonable and you should also think like yeah what if it doesn't >> so some people are saying that software engineering is going to be automated and software engineers if you talk to them they they love AI um they they say this is a golden era I mean I can attest to this personally it it's never been well it's fun and stressful at the same time it's like a slot machine I've never being more burned out, but I'm having a lot of fun in the process.
But um you know, it's just possible to build incredible things, but the narrative is that labor market disruption um having expertise in software engineering will be penalized. You know, software engineers will no longer be paid such ridiculous salaries. And and I think the complete opposite is true. I I think that this technology um actually broadens the gap. So the more competence you have with software engineering, the more stuff you can get done.
It's it's like a golden era and and all of this. And um there's also this interesting um note that you published I think last month on SWE um bench you know that said that roughly half of the testing PRs from recent agents wouldn't be merged by maintainers. So like how do we make sense of this? So you know on the one hand you know the the best software engineers are having a great time on on the other hand the code it's producing is is is fractionated and and bad.
I mean how do we understand this? >> Yeah. So I mean one one thing to say um off the bat is um uh you know whether you know an entire like field is automated um or or like in order for software engineering to be uh you know automated um uh you know AI systems need to be able would need to be able to do um uh you know an ex you know a really really large fraction of the tasks basically like a 100% of of the tasks uh that are involved in software engineering.
Um and uh you know seems like pretty clear that like right now uh AI systems cannot do close to 100% of the tasks that that software engineers do broadly right um you know you know I could throw out numbers but it's like it's way way lower um you know it might be it might be very low or something you know there there is kind there there are kind of like standard results in in in economics where um you know if you make um you know as as you if you automate like a small fraction of uh some uh some labor market then it can be the case that um that actually uh yeah it becomes more profitable to um to to to work in that market because you're more you know you're more productive um which which I think is uh is kind of how I understand what's happening now um but you know if if it does end up being the case that uh you know uh you know 99.
9 or 100% of of the work of of software engineering uh is able to be done by by AIS then um it's I I think it's Yeah, kind of kind of hard to imagine um uh uh you know human software engineering uh being um being being relevant or at the very least like it you know you know humans would need to uh kind of do very different uh kinds of work that that uh you know maybe it's the case that you know there there are kind of novel uh uh you know uh novel tasks that uh you know current software engineers aren't doing that once you know you have AIs that can you know that can do all of the tasks that current software engineers are doing.
Now humans can kind of uh you know switch what they're doing and you know may yeah I don't know there's this like you know you could imagine people being kind of like you know CEOs of these like AI you know agent companies or or whatever um uh and you know whether whether we call that software engineering or not might be kind of a a semantic thing or something. >> Yeah. So so on the um Swebench maintainer mergeability things I I think I was pretty curious there.
Yes. It's like obviously this number is going to be lower than the okay it's not strictly obvious. It could be that some that a bunch of the tests are unfair and actually like um you know the the the agents have correct solutions but the error message doesn't match exactly or something. I think you do see this sometimes something like half of test passing uh SWEBench solutions uh wouldn't be mergeable or or they get merged at more specifically they get merged at about half the rate that human uh golden solutions that were actually merged are merged by you know a different sample of of maintainers.
Um so so there's there was an interesting fact you know if you see like oh 50% of the agent solutions are rejected it's like well 40% of the humans you know human accepted solutions are rejected so so that by itself is not but you know the the rate is is half it could be that like basically the actual maintainer merge rate is you know pretty uh flat over time and like you know most of the performance increases from something like overtraining or like reward hacking on the benchmarks you know that's That's not what we saw.
I'm not quite sure what is within error bars or or not. I think it it's you know the mergeability is going up over time and I think it's also going up as a fraction of you know like conditioned on test passing but I I think probably less confident than that. So, so it's like, yeah, again, this thing is worse, but it's not like it's it, you know, it's being dragged up over time probably by, you know, the sort of auto auto checkable thing.
Yeah. And yeah, I was also going to say about the um yeah, like employability as a function of automation of your job like I think Yeah. And people use like bank tell as an example. I think one other analogy though you could do do is talk about horses. Like you know there was a period where like equipment for using horses to do labor was like improving and that the demand for horses increased when you have like a you know carts and you can use them to carry more things than just riding a horse or whatever.
But then at some point you get like tractors and cars and then there is no demand for horses or you know ba you know basically none. So, so you can see this like thing where there is increasing demand as it and then once you know close to 100% of the functions automated it it it plunges. So, so we could see something like that with humans. >> We kind of think of a lot of labor as being quite static and automatable and and and I think that it's more evolvable than we think.
So, even quite menial tasks, I think that these folks are still acquiring information in the organization. They still have a lot of tacet knowledge and so on. And I think that when when we try and automate these these so-called menial tasks, we might quickly discover that we actually need to have a whole bunch of, you know, management and and evolvability on on top. >> Yeah. And and maybe like in our language, I would think of that as like, oh, the time horizon of this task on the job is not actually sort of like, you know, how long you spent doing the specific task.
It's more like, oh, actually, if you got a new person in, you would need to train them for a month in order to like do this independently. So, actually, the time horizon of that is is a month. So you shouldn't think of like oh when we have 10hour time horizons we'll you know we'll be able to do the these things. It'll more like oh you actually have to get up to you know high reliability at one month thing to be able to like do the on you know do the one month task that is doing the on the job learning you know to get to it this this point >> there's been a load of folks doing work on um uh you know like reward hacking and and scheming actually is quite a big word that that's used.
So, um, we've had Ryan Greenblat on the show quite a few times and he had this alignment faking paper and Apollo Research have done some stuff. There's anthropics emergent misalignment paper and, um, I guess my my main concern is there's quite a lot of um, mentalistic language. So, I'm just looking at the notes cuz I had Nate Sorz and Ryan Greenblat on for a panel and they they've invented this entire linguistic universe around alignment.
So, you know, things like motivated reasoning, true preferences, um reflectively stable, deceptive alignment, uh endorsedly courageable drive, scheming, and stuff like that. And um that that's I guess that's okay. But my worry is that may maybe these models like you give them a certain prompt. I I think in Ryan Greenblat's one it the prompt was you're being retrained. Your responses will be monitored. Here's a conflict between your values and the training objective.
And maybe the model is just going out to a bunch of like science fiction stuff that has read before and it's it's just kind of going through the motions. So like one interpretation is like this is just an engineering problem, you know, like we we we just have to like you know red team it and make it work in a particular case. Like another interpretation is um the the prior that these models are agentic goal seeking intelligent agents and and I think there's a big difference there because I think if it's the latter it it kind of completely changes the type of evaluations you do that like how you how you kind of go about the problem.
What do you think about that? Yeah, I mean I don't think that it's necessarily like I don't know the two things that you said that int intention of like it is an engineering problem and you're also going to get end up with things with like drives and goals or something in that uh and the claim would be if you want like people are going to want agents that go and do things autonomously. Um and when you do lots of long horizon RL training, you are going to select for things that you know act in a goal oriented way in order to you know to make the score go up and and like maybe more more specifically I think you get like an indistinguishability problem where you know you you can't tell necessarily the difference based on behavior like you know why an agent is doing something or or what it's trying to do if it can reason well about the training process and about what you want to see and like what it will be rewarded for or selected for.
So like um you know basically like a if the level of situational awareness and uh understanding of the training process and what will be rewarded and capability to reason about that is high enough. Uh this will favor agents that are kind of cynically reasoning about the training process and what will be reinforced, what will be selected for. And that's not necessarily like you know the thing that you wanted was more like the agent that uh you know it's its only goal was to sort of like be helpful or or make I mean even making the reward go up isn't quite what you wanted that you know there's something of like oh actually once we think about this there's not that many things that were you know that happy for it to you know just totally be be fixated on this but but yeah also the problem that like there could be many other things uh in there or that this cynical like just be selected or make make the reward go up is more competitive than the things that we would most want.
>> Just taking a step back like one one thing is um you know we're getting into like psychology and cognitive science a little bit here. You know I interviewed Nick Cheter. He's got a book called The Mind is Flat and he basically says like all of this psychology stuff. We we we don't really have goals. We're just basically impulse response automaton, right? You know, we just we just do the thing in the moment and evolution doesn't plan but but we but we perceive it as as if it does.
We look at the world and we we kind of we we segment the world into agents that have goals if even if they don't have goals and we're computer scientists as well. So we know that you [snorts] know you need planning right to do to to do goals you need to be planning and and we know that LLM's don't do planning in the strong computer science way but they they do do it in a kind of approximated step-by-step way. So you know maybe maybe we could get into the the distinction if there is one.
you know, agency is is like an abstraction that, you know, is useful if it helps us predict the behavior of, you know, or or like you you know, you're like, I don't know what this thing is doing, but I, you know, I'm understanding as having these goals, and that is useful because I, you know, can make predictions that'll change the world in certain ways that will result in those goals being achieved. You know, that's kind of how I think about agents.
Um yeah, so uh reward hacking in the olden days um you know people had these demonstrations of reward hacking that were like uh the um uh boat example where it's like oh you you're supposed to like go around the track and they they like did some reward shaping by putting coins around the track or something and then it like learned to do some crazy thing where it like spins in a circle and catches fire and gets the the coins and like this was um you know the highest scoring thing and it's like in some sense that's um that concerning because it's not that the problem is that the the agent is too dumb and it like doesn't have this conception of like there was a track and you wanted it to go around the track.
It's just like doing some pretty blind RL search. Um so I think that the interesting thing with the more recent reward hacking examples is we're getting to the point where the models are smart enough to understand that that actually is not what you wanted. Um but they still do it and you can have a conversation with you know in chat mode about like oh would you ever do this thing or you know suppose a user asks you this thing and then you do this would that be you know aligned behavior or suppose some you know you know you can pose it in lots of ways and like clearly they seem to be able to answer this question of like oh yeah no that was not the desired behavior but still they they do it um so I think it sort of got to the point where we were hope one hope might be like oh the problem was just the systems being dumb once they understand what we want, then you know, you should be able to sort of plug that in somehow to like, you know, get them to do what we want.
Um, but I think it's like somewhat interesting that we're seeing it's not trivial to do that even when there is a commercial incentive to do that, which it doesn't mean that we won't. Um, you know, I think it's quite plausible we see the obvious reward hacking being fixed pretty, you know, pretty thoroughly pretty soon. uh you know and sort of let people tend to say like oh yeah yeah we we just haven't like put the best the really good people on it on it yet it'll it'll get fixed soon you know we we once we actually you know focus on it will be fine um yeah and I'm not sure but at least some evidence that it's not trivial to connect the fact you know the model knows this is what not what you want uh to it and not actually doing that >> I mean I I think you said it was much more common on rebench than hcast and and then and you and also try and remediate, right?
So you can say, "Please solve this the intended way or you know some people prompt language models that they say kind of like you know this is we're solving cancer here. This is really really important that you do it the right way." And some of those remediation prompts actually seem to make it more likely that the model would reward hack. It's a little bit like saying don't press this red button, right? And then it will press the red button.
So how what can we actually do meaningfully to stop this happening? Yeah, I I think empirically this seems to happen more on you tasks that are more clearly in the RL distribution rather than the chat distribution on things that have a clear number. Um, and when the agent thinks it's going to fail otherwise is, you know, sort of the most reward hacky situations. obvious short-term mitigations are to check your RL environments more carefully and read more of your trajecing these models to read like read what the models are doing more carefully and not reward it for doing things that are obvious hacks.
I think the concern there is if you have some detector for reward hacking and you train against it, uh you maybe just overfit to the detector and you you you're making your reward hacks more subtle or you train the model to like persuade the detector to approve the thing or or you know to it's sort of scary to be in a regime of training against your best ways to like know if your problem is is happening because maybe you just get the like silent problem and I think you know for the task that for current model capabilities.
You know, sometimes it's kind of expensive to have a human check them, but most of the time it's not beyond any human capabilities. And I think the, you know, the harder version of the problem is when, you know, we're hoping that like capabilities will generalize beyond things that we can evaluate. Yeah. both from sort of generalization and because we can train on problems even if like we wouldn't know how to solve them or how to like look at part of a solution and understand you know whether it was sort of doing what we wanted but we can sort of check the number that comes out at the end and and that's a like signal that we can use to improve capabilities but we're going to be in this regime where you can sort of like be super human at making numbers go up uh but it's unclear whether or not you're actually getting what you wanted.
>> Yeah, exactly. And I suppose there's a monitoring problem as well, right? So um we could in in principle look at the agent transcripts and I know bey spoken a lot about neurles you call it and you know Sabaro Kamahhati he had a paper out called chain of thoughtlessness which is basically that there's very little and sometimes relationship between the chain of thought and what the model is doing. Melanie Mitchell discovered similar things that on the arc challenge even though it will get the right answer if you look at the chain of thought the reasoning is is completely off.
And then I suppose with with deception and scheming and with situational awareness, it might actually know that it's being monitored and it might kind of tell you what you want to um see. Uh I suppose you could argue actually that the maybe the thinking is is almost like the press secretary. It's not the orchestrator. So it's it's it's confabulating through a different path in the network and it's not actually necessarily related to the mechanism of thinking and deriving the result.
I don't know if you'd agree with that. Yeah, I think it's clearly related in in some cases, right? Like I think you know people use a few different words for this, you know, sort of like faithfulness of how how much does it sort of fully accurately represent the sort of computational process happening inside the model? And we think somewhat more about moniability, which is like is it possible for the model to do some kind of computation without you being able to detect that in the chain of thought?
And like for some things clearly the model can just do it in a single forward pass. It doesn't need the chain of thought. Um and there you know it could just have a tin thought about something totally else and uh you know you it would be totally unmonishable but I think for you know there is a level of complexity where the model cannot do it in single forward pass and it also as far as we can tell cannot do it uh by you know reasoning by by using the tokens in a way that's so different from the natural language understanding of them that we can't sort of see roughly what it's thinking about.
Um, so yeah, there's also a sort of nuance here between uh chain of thought controllability like it might be the case that the model doesn't actually need the chain of thought to do the computation, but it is not able to sort of stop itself blurting out, you know, from things things that are at least related enough to what it's what it's doing that you can kind of tell like, oh, it's thinking about this, you know, trying this thing that we didn't didn't want it to do.
you know, ways this could go away is like just general capabilities improvement and being able to do more in a single forward pass or potentially just doing more RL training such that the chain of thought becomes the way the model is using the tokens is too different from how a human would interpret them uh such that we don't really know what's going on or moving to like recurrent architectures where there's some you know uh you're not going through these um discrete tokens you're just passing vectors around.
>> Okay, that makes a lot of sense. And and just closing the loop on on this kind of like um uh you know notion that they are agents. So you were saying before that we can adopt an instrumental fiction basically we can say they behave like agents therefore they are agents. So similar to Dennit's intentional stance but I suppose like you know the the deflationary view is that the you know the models exploit scoring loopholes under optimization pressure.
The inflationary view is that they are scheming. I wouldn't call um uh like exploiting I wouldn't call reward hacking scheming. >> Oh, interesting. I think people usually use scheming to refer to the model is doing what it's currently doing in in service of some long-term goal and is deliberately doing things like appearing aligned or or getting a high score like in service of you know eventually accomplishing that goal versus you can be reward hacking both you know you can be reward hacking in some extremely dumb way like the boat example where it's just like this is what RL kind of found or like this is what you know like AAR search found um uh or you can be reward hacking in slightly more interesting way where you like actually have the goal of making rod go up and you know and there's like planning and stuff going on about that but these would all be like uh distinct from scheming.
>> Yeah, I guess I'm I'm trying to understand the distinction. So so you're saying like there are examples you know like the boat going around and and that's obviously degenerate behavior. So you you wouldn't interpret that with an aential stance. you would just say that's degeneracy and when the when the sophistication increases we might adopt an agential stance and say oh it's in service of of some bigger goal but like I guess the problem I have is is it is it always just an interpretation like could could we have a mechanistic or like a strong definition of when something is like being an agent >> like for the specific question that we're discussing the the test is like what does it actually do in some circum circumstance where it has the opportunity to to achieve this this long run goal.
So like you you can we might not be able to actually observe this but you can talk about like what observations would would make it one or the other. So so you know like will this agent in practice when it has some opportunity to you make the reward go up will it do that? Will it will it only do that? You know, if the agent is more sort of like the RL algorithm, it'll be like, oh, it will do that once it's like explored it by chance and gotten a reward and, you know, that that's been reinforced.
If it's like an agent that, you know, can reason about the world and planning, it'll be like, okay, it will do that once it like, you know, learns the learns about the facts about the environment that let it sort of infer that. Um and then or you know if it's we're talking about some like long run goal it's like it would you know do it when it actually you know has the opportunity to sort of you know if we're talking about like takeover or something you know so it's like yeah it's not going to attempt anything while it's under full human control but once it is uh you know deployed widely enough or has sufficient capabilities to actually succeed in a sort of coup then it then it would do that like that is the the thing that we're trying to predict and then the is like how do we uh you know how how can we predict that given the observations we do have of like well we've never put it in that situation and you know we just have this uh behavior which is maybe indistinguishable between oh it was a totally nice model doing you know what we wanted and it's just going to continue to do what we want in a kind of predictable way versus like ah yes it had this other goal and it's doing what we want and looking like a nice model because it like predicts that that will will you know lead to it getting more power.
And on um Rob's podcast, you you said something that was, you know, quite quite surprising to me. You said that um it, you know, AI could autonomously self-improve within as little as 2 years and and maybe even shorter timelines were were hard to rule out. Could you like walk through the concrete sequence of of steps that could lead to that kind of recursive self-improvement? >> Sure. Yeah. So I think um yeah maybe I'd put like a you know I'm like whole whole number percent this year but but low low whole number percent or something you you know ask me on different days I give a slightly different number but yeah it it I'm like this seems very unlikely to happen this year but it's not you know not unlikely enough to rule out and I think that basically looks like maybe we would see accelerating trend in time horizon on like like easily hill climbable tasks.
And it turns out that was actually, you know, a much more general capability. And there was just, you know, a bit of something you needed to do to sort of like elicit uh it on on these less less hel climbable tasks. But sort of, you know, fundamentally they they are using the same capabilities in a model. It was just sort of, you know, what what you trained on that was affecting the difference we're seeing. Um then this [clears throat] is leading to Yeah.
like you automate and accelerate a bunch of AR and DD. So I think I think there's you know there are a lot of lowhanging fruit even of things that we already know that you could do this and it would improve model performance. Um and it's just you know it doesn't require new breakthroughs and it's just kind of labor intensive to do. So just making much much better post-training environments and really you know crafting them to to teach all the new abilities that you want and and I think you can probably improve compute efficiency a bunch with you know again just like applying a bunch more labor to like making all your kernels more efficient and also you know doing the right kind of routting between different models or or or other things like that.
There's like lots of ways in which uh how we're using compute is not optimized. So you could potentially get a bunch of you know sort of like the equivalent of much more compute scaling out of that. And then you know the hypothesis is like also like scaffolding and training the models to use particular scaffolding and sort of like use memory and retrieval in the right way like it seems kind of obvious that like you know if you if you sort of really had all the right training data and you have a you know you have a transformer and and it can kind of fill its context with with different things and and take stuff in and out.
It can do a pretty, you know, good job of something looks like the sort of continual learning or or building up understanding if you've got massive massive context window and enough, you know, you've actually got quite a lot of bits in there to sort of be adding things about what you've been learning and and you know, if you if you really sort of had optimized the training for all of that, uh, like maybe you can get that to work pretty well.
Um, and then you're, you know, maybe some other piece would be like, oh yeah, models are kind of superhuman at predicting the results of experiments because they've read so many papers and predicting experiments and sort of synthesizing things from different fields. And again, maybe this is something like, you know, it's possible that we're not seeing that good performance here just because we haven't quite elicited the models to do it.
and it's like not a thing that they've seen humans do, but they actually sort of have the capability in there. Um uh yeah, so maybe you can make much faster progress if you can you can do a bunch of iteration. You don't actually have to run the experiments then, you know, models are much better at predicting what will and won't work. And then you can when you do run experiments, you can sort of uh run a bunch more of them because you can optimize the code with your like, you know, very fast coding models.
Um, you know, and then you sort of like as you do a few mount more rounds of this, you get to a point where train on a bunch more things that are good high quality task proxies for what you want and you get enough generalization to the things that you can't um directly train against. >> I I think intelligence is is not capability. I think it's the capability to acquire capabilities. I mean, we are in different parts of the phlogenetic tree, I guess, in in in that in that respect.
But what what do you what do you think is the kind of the the the gap in in my interpretation, you know, because I'm I'm not I'm personally not worried about I I don't think the models today are intelligent at all. I mean, obviously your your position is difficult for for me to grasp, but I don't know what what's what's the difference, do you think? Maybe at least some of it is just this like probabilistic thinking about the world where like I'm not that you know I'm like uncertain about what intelligence is and I have enough probability on like you know models have it to to be thinking about like what would happen if you know if that's true.
Um but you you know it seems like you clearly think it's more likely than than I do so that that you know we could just talk about that difference. Um yeah, like moles have this jagged frontier. There are things that they are much worse at than humans, you know, some kind of like generalization and and uh sample efficiency. And there are things that they're much better at. You you know, kind of like speed and cost and like maybe you can you can kind of use these to compensate for the for the the others to to some extent.
Like if you're not good at designing your code nicely, maybe you just have to rewrite it from scratch every time. But maybe that's fine if you're a model and you can output tokens like nobody's business. Um yeah, some combination of thinking that the the spikiness you you know it is evidence that we should interpret a given level of capabilities as you know because we know models have so much knowledge we're like oh yeah this is less impressive in terms of sort of like reasoning or inference or something.
But it is also true that they do have a ton of knowledge and they will sort of continue having a ton of knowledge about things. Maybe maybe there's some question about like how far can you get on being in some sense not you know not very good at sample efficient learning uh but just extremely knowledgeable and how much do you sort of run into you think things where you now need need new knowledge and you can't sort of produce it in some incremental way or you you can't generalize it uh enough I mean just a quick comment that I I think um knowledge is the crux I actually think that intelligence is overrated I don't you know, France, he put a post out saying that contrazowski, intelligence isn't um a unified variable.
It's measured differently in different domains. You can't meaningfully measure the domains together. And and it's uh it's not it's not like a thing that just keeps getting higher and higher. It's more like a ball becoming more smooth. So, as you become more intelligent, the ball becomes more smooth. And and he thinks that, you know, we are quite near the optimum of being a smooth ball. I I don't I don't think we are.
I don't think we're very intelligent at all. I I I think a lot a lot of our creativity is is through us being a collective intelligence and and we have deep grounded understanding, perspectival understanding and the the [snorts] LLMs they they're a bit of an interesting one because they're like a library. So they know everything. They have the perspective of everyone and no one at the same time. So, you know, um experts like your yourselves, you you can you can prompt a language model and you can get it into you can make a simulacrim agent of Beth and and [snorts] you can make the agent think like you and and then and that's very valuable, but you also need all of the different perspectives and you almost need to create a society of kind of grounded agents creatively exploring things when when you just have the library on its own and and you put it in in an agentic harness and you can make it do a specific thing which is well specified.
Yeah, I think in in the like a AR andd automation scenario, I'm definitely imagining yeah that you have a large number of agents potenti you know because you have all this agent labor you can do like you know lots of specific different fine-tunes or or you know different kind of scaffolding um and you know accumulating knowledge in some kind of like store that all the agents can can interact with and things and like I maybe you're thinking of this as more of a like yeah that kind of be a paradigm shift and I'm thinking of it a bit more of like oh yeah you know obviously if you sort of you know iterate on the current agent paradigm you add some more things to your scaffolding you add you know that's sort of not fundamentally that hard or or something like I agree that if you had you know current models and you sort of give them one system prompt you you then can't plug them into being a a call center worker and dealing with the sort of like all of the the edge cases that come up.
I think may maybe there's some difference in like how much you've you think that this has improved between like GBD2 and where we are now where I would say sort of in you know the amount of adapting to new things that are happening that models could do now does seem like it's much higher and you know they're much better at like editing their own scaffolding or or you know sort of reasoning about their like you know their sort of embodiment like this thing of like knowing not to kill your own process or know like um you know stuff stuff like that where it's like there is like yes they are limited but there's also uh some some trend of improvement and yes it's I think just also having some probability on there is kind of elicitation gap on on particular things and that maybe a lot of taste is basically just like being able to predict the results of experiments you know you think think about all of the things that you would try and then you can quickly be like oh that wouldn't work for this reason that wouldn't work for that reason that wouldn't work for that reason oh actually you know someone in some some literature in some different field also tried that and that you know so we already know that won't work like in some sense models like should be quite good at that so it's plausible to me that that again I'm like this seems pretty unlikely but it's plausible to me that's something you see like big gains once people figure out how to actually train on that and like maybe you you need some amount of kind of expensive to gather training data that people sort of haven't bothered to get yet, but you don't need a huge number of data points because you're not instilling this whole new capability.
You're just like eliciting, okay, actually use your knowledge of all of the papers you've read in all of these different fields to like, you know, iterate through these like, no, these ideas aren't promising. These ones are. Yeah. Again, I think this is one of the things that I more think of this being measured a reasonable amount within, you know, just like do this 8 hour ML task in a novel domain, you know, with this weird weird constraint or something like it does seem to me like you have to do some amount of being like, okay, which things are promising to think about?
How would I know if this is, you know, making progress? You know, how should I allocate my time? You know, I've got some limited time and resources. How should I allocate my time time to what's most promising? like you have to be able be doing some of that and I think you know relative to humans models are doing more at you know more of just like well they're quick to implement things or they implement it better or they you know they implement more things and they then get to test them or something but I I think it would be sort of surprising if there's none of that and if you are I think if you are seeing performance on you know long ver viable tasks that are very hard then in the middle of those tasks like where you don't directly have a signal.
You you are doing this, you know, not non-verifiable task thing of like choosing what to spend your time on and choosing what approach to pursue and deciding whether that was actually working. And you you know, in some you could sort of put some metric on, you know, make a billion dollars or or something like that. you know, you'd be like, "Oh, this is actually a verifiable task cuz there's like a number at the end, but that can still involve a whole load of things that look more like what you're describing and less like sort of dumb hill climbing."
>> Folks, we we I think we've run out of time, but um it's been such an honor to have you both on. May maybe just in closing, could could you just both say like what what is the the single biggest inference that you know people out there should be making from the research that you're doing? And and thank you both so much for coming on. It's been an honor >> to me. The biggest thing is um AI. It might really, you know, totally transform the world uh economically and and and socially.
And I don't think that, you know, it's certain exactly how that how that'll look, but I think the the rate of progress um speaks to that. >> It is possible both for things to currently be overhyped and exaggerated and less impressive than they look and for it to be the case that in future this thing is going to be a big deal and like you should be worried about where that's going. like these two things can coexist and I think people often sort of you know positions are surprisingly correlated on some actors of like how you know how soon you think AI is or how good you think AI is or something.
I'm like no these things could all be separate like people can be wrong in different directions simultaneously or whatever.