🎙️AI 访谈库
什么在阻挡 AGI?——Jerry Tworek
Jerry Tworek · OpenAI

什么在阻挡 AGI?——Jerry Tworek

Whats blocking AGI? Jerry Tworek

2026-05-13 · ARC Prize · 27m · 约 27 分钟读完 · 原文
ARC Prize 对谈:什么在阻碍 AGI——推理模型之后缺失的能力、测试时学习与更有表达力的架构,以及 ARC 类基准对衡量泛化的意义。

>> Thank you very much for joining. We have uh Jerry Tworek uh who recently left OpenAI and is now CEO of his new stealth startup that he won't even tell us the name of it's so stealthy. Uh welcome to Y Combinator. >> Can't wait to be here. Thank you very much for the introduction. >> Yeah. Um so we have a bunch of things to hit. Uh the first one that I really I I think as a research community we're not even really settled on um is a definition for intelligence.

Um there's a bunch of literature in the last 100 years on from Spearman 1904 to Legg yeah, 2007 that was on this paper uh universal intelligence and then Chalais's uh on the measure of intelligence obviously which is kind of why we're here. Um what do you think about these? How do you how you define intelligence? >> Thinking about intelligence and trying to define intelligence is something that faultily very close to me and very personal to me.

I've been working at for for for many years of my life. And as I am thinking about it like how do we how how how would how do I define and what what what it means for me is a lot about the unknown and the and being able to adapt to unknown. If we think about the computers for example like playing chess or playing Go. As as kind of smart as it is it's worth asking ourselves a question is that intelligence? The computer doing something that is hard for us is that enough?

Like is the computer summing 10,000 numbers in in 1 second is that intelligence or not? And like in some it is it is computation but but this is a computation that is easy to be programmed. So for me like the intelligence is not that it's not doing something hard. The intelligence is about being able to adapt itself and being able to learn on the fly and in some way like intelligence is a process that can technically over time if you have enough enough ability to focus on hard problems to kind of like unravel them and solve something that wasn't able to solve before and if intelligence is put in a new place and your environment it is able to solve it and it is able to figure out which is the program that plays chess or can sum 10,000 numbers in a fly.

It cannot really like the the program that plays chess cannot play go. It cannot it cannot sum numbers. It cannot it cannot like uh work on figuring out new mathematics. But even if there is a program that can figure out new mathematics but cannot figure figure out other types of problems and still it it isn't that intelligence like you know there's probably a spectrum. There's some levels of of like middle intelligence if if you can adapt to like new domains in mathematics and at least like work with in an open-ended fashion.

I think we as humans like ascribe intelligence to ourselves and we think our our intelligence is open-ended and has been at least like open-ended so far to this moment where we are in the world we are able to expand and learn and build new technologies but there's this only some something that feels limited like feels like it isn't it isn't intelligent in in some like very strict sense of that word. >> And what do you think about the need in the definition of intelligence or at least measuring intelligence to control for uh prior experience, priors in the architecture, um and training data, number of training iters.

Um are those important? >> Yeah. Yeah. It's it's it's a great question and I think for me personally and I I I don't want to over over-generalize it here but for me it isn't. For me like if we see someone and and and and I guess it's it's humans thinking of intelligence. If we see someone incredibly successful solving incredibly hard problem one after another, and and then and then like adapting quickly and going to a to a challenges of higher and higher magnitude very successfully in a row.

Like we don't think, "Oh, he had more data. He had more books than the other people." We we we just say this is this this is intelligence. Intelligence is like practical applied problem solving in a new situations that we didn't like foresee before and and plan for before. And in that case, like it doesn't matter what that that person do. If we If we think like Do we think that that some other organism on Earth could be more intelligent than humans because just they had less training data and then and then that's that's how we did.

Like I don't think I don't think it's a it's a really fair thing. All that All that training what happened training results in intelligence in the end and and I personally at least at least don't really don't really discount that that much except for may may may maybe there's Mhm. Maybe there there there there's one angle in which in which it actually makes sense. Which is like there's definitely something like rate of learning.

If there are If there are If there are two kind of like players in some way in a game where they kind of start from the same place, but you need to you need to make sure they they they they start from the same place. And at that moment, like one just learns much faster like like kind of like kids in school in a way. One learns much faster. I think I think we would say that that kid has like is is more intelligent because they are learning faster and eventually they will they will overcome that.

So maybe may may maybe there is actually an argument there that um like the the slope of your learning is is in many ways intelligent, not only the not not not only the actual level of skill that you have. >> Yeah, the way we always said this is at Focal was I care a lot more about M than B. I care a lot more about slope than Y intercept. And so like um because the world changes and you want something that's going to adapt quickly.

Um we look for that at YC in our app in our applicants. How quickly are you moving? Um and so uh I think if I were to try to make this argument very tight, would you be more impressed with a 5-year-old that got a 35 on IMO um versus a 35-year-old? >> Yeah. Yeah. That's That's That's a good question. And there can be many different interpretations of this fact. I don't think it's it's like a super quick thing. Like technically 5-year-old doing that is much more rare event.

So in that case it is it is something to be to be impressed about. But there is also like a few see like a 10-year-old kid that's like memorized all the IMO problems ever and kind of like solves new problem just by pattern matching versus old problem versus versus like a season mathematician who understands a lot of theory. Understands all the all the links and kind of like solves problems in a more more creative way.

That That also could happen in a way. So there is there there there there is something deeper there beyond just age of of like how how fluid that intelligence is and how how how how dynamic it is and how like how that person would appear how would solve problems that are not necessarily IMO because IMO also by itself is kind of like very well-defined domain. >> Right. >> And kind of like a box to be put in in some in some way.

>> Is the same like 10 strategies. If you just follow the strategies you kind of like do answer most of the questions. Um yeah. So you can just like chess too. You can kind of like cheat and then you can like as Chalet would say, buy arbitrary levels of intelligence by studying the work of works of others and learning these heuristics that someone else figured out. And then you just kind of like really it's cheating.

It's fine it it's it's bootstrapping your intelligence from some other thing that was more intelligent, perhaps. So, the other angle on it would be let's say that we have two people. You get 1 hour to take IMO and I have 8 minutes and we both get the same score. And that would be like, you know, iterations or efficiency of the compute. And so, like again, that's a slope of like intelligence per joule. And so, like if I take less joules to get to the same to sort the list, like I should be a smarter algorithm, right?

I should I should get some points for that. It shouldn't be like I shouldn't be an imbecile. >> That's hardly the case. Hardly the case. >> Any any thoughts on that? Any to further expand? >> Yes, one of the works of my life of the recent years has been scaling test time compute and maybe you have you have seen those curves and and we worked a lot with the previous version of the of the ARC-AGI benchmark to try to instill even this methodology of learning like what kind of performance do we get per various amounts of test time compute and usually denominated per dollars because this is this is the easier to compare, for example, across intelligence providers of this of this day and age.

And like saying only like right now in this in this in the level of technology we have oh, my model is able to get this performance without stating the cost of getting to that performance. Actually, it's like doesn't doesn't give you that much information because today's models already they can spend more time thinking like they they can spend more time trying to solve problem and they usually get better results to a to a certain degree, but that degree actually is pretty pretty far.

They are it's it's much farther than most people feel comfortable spending on solving those problems. So, so that is that's a very very important input. >> Yeah. And so, we'll get to this I think later on, but along the same lines of like let let's say the the the very contrived example of like sort only and I have um infinite examples of unsorted lists and sorted lists and I have a um trace traces of bubble sort for example.

Um there is no amount of test time compute that will get me to a merge sort. The only thing you'll get is the more you traces you train of unsorted list to sorted list following some thinking tokens in between of like doing the bubble sort procedure you would get a crappy implementation of bubble sort. Maybe you'll get the actual implementation of bubble sort, but if I scale that up and just do more test time compute, I don't test time compute my way into merge sort.

Yes. And so like that's a fundamental issue I have with test time compute where like if you have if humans have already solved it um then we have the trace, but if humans don't have to solve it, there is no trace to train on. And so then you actually like if we didn't know about maybe we know about maybe there's an exist an even more efficient sort that we don't know about. Um and we can't test time compute our way into solving it.

How do you think about that? >> I think it the question here relies largely on your training data and if you are training on the sorts of like on the traces of a bubble sort, you'll only learn bubble sort, but that what's model makers are training models today is kind of a combination of all human knowledge which are largely not only itself the algorithms, but also like ways to think about algorithms and ways to derive algorithms which when combined together compressed in a model in the sec slightly fluid like some people call it crystallized, I call it fluid fluid intelligence because in in in models is like we take the data and you kind of like liquefy it into some some some weird amalgamation of everything they have seen and in that frame like the models have learned not only those algorithms, but have learned some ways how to derive algorithms.

Like it's it's clearly not perfect. It's clearly like not They are not uh absolutely superhuman at it, but they have they have various parts of like ways how humans think about about deriving algorithm about thinking of it. And and then and then kind of it's it's in them sometimes the question is how do we how do we get it out? How do we how do we elicit those capabilities? But if you were able to say like today's like best LLMs they will never be able to like figure out a new sorting algorithm.

I would say I don't I I I I don't have that confidence. I think I think I think they may because we actually learned taught them a lot about algorithmics, about algorithm design, about how people think about algorithms. I think I think this this can happen. >> But in the same way that like um the IMO uh strategies, these heuristics to um win to to get before you two on IMO or to play chess extremely well where, you know, I a pawn is one point and, you know, a rook is five points and all these other schemes that someone told us, right?

Um that's not like that's basically taking intelligence from someone else that and these heuristics that it's learned to do this thing called algorithmic um creation. Like to search over the space of algorithms and like do a better job. Um but it hasn't discovered its own strategies and heuristics. It was it's kind of bootstrapped and and on our own on our COTs that we gave it. And so like can it Do you think that you can test time compute your way to sample meaningfully outside the distribution?

>> Yeah. I think I think fundamentally you can if you trained all the ways Like the the the the the the the the the main question here is your training data, which is is your training data just the results of a the algorithm itself or kind of like patterns of thinking and patterns of of deduction and patterns of of how to find new things. Cuz in some way like space of like chess heuristics is is is is is is a space of some kind, space of algorithms is space of some kind.

And humans, even when we think about algorithms or or about space or about chess heuristics, we usually like have some ways how do we structure those things, how do we how do we think about them. And in many ways like those already are encoded in today's today's models. I think models of that of the day are very good like search operators over spaces that's that that kind of like at least at least defy brute force searches that that that that all their all their algorithms would do if you if you try to think oh let me search for over a space of all the search algorithms all the sort algorithms, it'll be it'll be extremely hard thing to do.

Like it's a it's a it's a gigantic space with with tons of things. If you ask LLM to start exploring that space, we can we can suddenly narrow it down very much. And and I think eventually they would they would come out with something that that that makes sense that is interesting. And in terms of like AlphaGo move like 37 or whatever that move is called was was also that which was found through through reinforcement learning, through models trying a lot of different moves, through searching through them.

And it was also not not a brute force search. It was was kind of informed search for the for for for the value function of that model. And then and then it it developed that that new strategy. So I think I think think the models it's like given given how we train them, they're not completely incapable of they are may may may maybe not yet not yet as as good at it as we would like, but but it's not impossible. Yeah. Um yeah, no I I agree.

So so now we discussed um a definition of intelligence. We've talked about skill acquisition. We've talked about um skill acquisition efficiency, which is like I think still the efficiency, the M is actually a very important thing. Um I we have we've talked about normalizing for prior experience and itters. What do you think is the best practical way to measure and quantify intelligence? >> That's a really really great question because I am personally a little bit bearish in a way.

I think we have been consistently measuring wrong things or at least I think that I think maybe maybe saying it a little bit more like being cautious to to all of us in the past. We were we're measuring things that we had a hard time solving at a given moment and until chess was solved people were thinking of chess as a good benchmark for intelligence until other things were solved were thinking those all those things were good measures but what happens is we solve all our past benchmarks.

We solve all our past measures of intelligence and we still think this is not AGI. So we realize whatever we were measuring it wasn't like it was kind of adjacent. It was kind of close to intelligence but it wasn't it wasn't it. And to me like the really real best ways to measure whether whether the models can be smart is like in a in in kind of like some infinite source of task problems where they have some tendency to be fresh.

Can the model continuously do them well versus versus other alternatives whatever whatever that will be in some way the the biggest weakness of benchmarks is the Goodhart's law which is like as long as benchmark is known people will target it in training. They will solve it. For any benchmark right now in the world there are already very good algorithms that can solve this specific benchmark and then they can they can maximize that.

So unless you have a way to generate fresh data that is unseen the the unseen part is is the important part of the end intelligence and I think you you you agree with that with me. If someone takes a benchmark no one has seen yet and takes today's models and and and tries to measure them in some kind this probably is a good benchmark. As soon as the first frame runs after this benchmark has been has been released I I think it it stops being a good good measure of intelligence because suddenly the models the models have been have been trained on it and um it's it happens to all of it.

So so only only real really like future tasks irreducible tasks that we can know beforehand are they really really really good ones and they also would be nice if they are different enough and and interesting enough to be to be worthwhile. >> Yeah. >> And that's why a lot of people like like inventing your science is the following one of those things. >> So now we have discussed some way to to do this.

There's the Val PPL which I think we both disagree with as a measure of intelligence. There's the cacophony of random benchmarks that you can just reward hack that everyone has done already. Um back in like I was actually just talking to Tim Sheehy one of the original OAI folks um talking about ProcGen uh you know the coin run these world of bits things that were happening in OpenAI way back when pre pre-Elon era. And um it was gameplay.

And like even DeepMind was back then very much like pilled on gameplay. And now we have Arc AGI 3 which is gameplay. And there's holdout games. And each game requires new skill acquisition where like the skills to to the hard uh deterministic code to win one game would never win the next game. Um what are your thoughts on that? >> Yeah. >> Gameplay is very interesting. I I particularly like the game and I think it's worth study them because games were meant and designed for human intelligence to be interesting and then engaging.

At least it's like if the game was very simple for us it wouldn't be super interesting to play. There there there there definitely some entertainment elements but I think I think in many ways like games were were were meant to be challenging to human intelligence, and that's what makes it really really cool. And a lot of games test things that we are struggling with right now, resource allocation, long-term planning, like multimodal perception.

A lot of a lot of of those things current models are not very good at, and I think models will like with general game distribution, can the models play the latest Halo? Uh it's it's it's kind of it'll take a while for us to for us to really really get there. I don't think I don't think they will those benchmarks will fall anytime anytime soon. There There is like a bit of a question, and and I'm I'm asking it myself, like how how how possible it is to to kind of like good heart games as well.

Generalized to to unseen games. I remember like in the in the back in the day we were researching a lot of Atari games, and I don't think anyone like we were each Atari game ever anyone would be able to solve. Montezuma's Revenge was that was that was kind of the hardest one, but no one I think was able to get generalization between games. It was was very very hard, but arguably the uh the the distribution of the data set still wasn't reaching an algorithm wasn't there.

I personally believe that today's LLMs would have a much better time um generalizing between games because the representations of the world are already built well enough to kind of to try to be able to make those jumps. If I if I if I'm good in some games, I will be able to good be good in the other games. But on the other hand, they are just just struggling with things like um low latency like actions and multimodal perceptions, and like they they are big enough, and like sometimes playing games requires being being fast and and responding quickly, which may be may be a little bit a little bit of a problematic thing there.

Um but No, I think I think I think there defaultly we can see some progress there and whether whether we think if if there is a model that solves the games and and then we start thinking is it is it is this the AGI? Can it can it create reliably new science? I think I think this will be this will be a thing to to still see. >> So, there's two more questions here. So, the um which So, now we've talked about the benchmarks.

Uh which of the shots on goal to hill climb on that measure do you most agree with? There's like what I call Ilyaism, which is the famous N to B is enough. All you need is more tokens, bro. Um Lacunism, which is this world modeling Jet Pa, self-supervised learning, um you know, latent state latent space predictive coding kind of thing. Um Noamism, which I would call Ilyaism plus generator verifier gap uh to give you infinite data.

Um Chillaism, which is this program synthesis plus neurosymbolic methods. Thomas Parr stuff with the active inference. Uh you know, Chelsea Finn Abialism uh is which is like metal learning these metal learning RL approaches. Which one do you most align with? >> That's a good question. I think all of them have some point and all of them are are defaultly like getting some part of the truth. And and in reality we need more tokens.

We need more metal learning. Like probably if you if you ask me like which one is the most right I would I I I think I would side with the metal learning angle. I just I just don't think we've been doing metal learning well enough, but in many case even if you are looking at transformers like they are they are metal learning a lot. Like the the in-context learning transformers exhibit this metal learning and it's it's it's the best metal learning algorithm we have ever created as as as a humanity by far, but but but it's still it's still not AGI.

So, like the question is like what is what is the bottleneck in the current models to to get to versus versus the models earlier? And is it more tokens? Like, I think not really, although more more tokens always help. Is it more metal learning? Like, I think I think like metal learning kind of happens naturally in in in in in in through through the process of optimization and through the process of deep learning. Is it is it generator verifier gap?

We already are doing RL. We are and we are doing a lot of RL already. Like, is do we do we just need more RL? Maybe just just just just one more order of magnitude of scaling grow and that will be good. Uh that's that's kind of like I think so. Like, I am like myself, I believe we haven't tried to iterate enough in the recent years on the architecture of the model itself. I think people are like we thought oh transformers scale and that is that is really great.

Let's just just just keep scaling transformers. And people are forgetting how much of the prior is built into the architecture and how much of the structure of thinking and maybe maybe this is the layer we should be trying to change more because we are already having a lot of tokens. We are already doing RL. We are already doing metal learning. So, what's what what what what what aren't we doing that is still missing and that and that's that's one of the things I'm thinking about a lot these days.

>> My my point of view I wrote this whole blog post about the importance of train time recurrence. And so, like, if you want to we're trying to build an architecture that is Turing complete. To be Turing complete, you need unbounded recurrence. You don't have that at train time. So, it never actually learns to be a be Turing complete. It actually only has one forward pass available to it in at train time. Uh and then you're snapping back to some teacher forced uh trace.

And so, it doesn't allow it to like, you know, develop its own latent space representations along the way. You've forced it through it's every single step. And we saw that the most successful AGI 2 were was HRM and TRM, which, you know, the last scaling law was test time compute. I think the next the next one is train time compute, train time recurrence. What are your thoughts on that? >> I am pretty enthusiastic about approaches like that.

I think it makes sense. Personally, I really like that the TRM results and there's a pretty big chance recurrence will make one one way or another come back in the coming years. >> Cool. Last question and then I'll let you go. Sorry. Um OAI versus Anthropic. I understand you just laughed. I don't know how well you are to to think about it, but it does seem like Anthropic is kind of catching up very quickly. There are some plots that say that like by mid-2026 they'll actually be beating in revenue.

Um what do you think? >> That's a charged question anyways and I have good friends in both of those companies and I think both of them are defining the the reality of the artificial intelligence we live today and both are very very successful companies that have been growing, maybe one of the fastest growing company in the current times. The only thing what I think about them is that in the current world of being a little bit too caught up in the competitive landscape of of like oh, this company is catching up, this company is is catching up.

Like they are slightly like gotten off the path of innovation and I think very much like in the in the way of like we need to squeeze in from existing way of doing machine learning as much as as much as possible and I think the the machine learning research is not done in any way. There's still much more much more to be done and I think I think that makes space for for for for like you know, for others also to try to come back and make some some new developments and new innovations in the field to happen.

>> Very exciting. Thank you so much. >> Thank you. >> I'm excited to say that for me it's a pleasure to do so.