🎙️AI 访谈库
前 OpenAI 研究员:离开的原因、真实的 AGI 时间线与 RL 扩展的极限
Jerry Tworek · OpenAI

前 OpenAI 研究员:离开的原因、真实的 AGI 时间线与 RL 扩展的极限

Ex-OpenAI Researcher On Why He Left, His Honest AGI Timeline, & The Limits of Scaling RL

2026-01-29 · Unsupervised Learning (Redpoint) · 1h03m · 约 65 分钟读完 · 原文
o1/o3 与 Codex 核心架构师离职后首批深访之一:坦诚的 AGI 时间线、RL 扩展的真实极限、离开 OpenAI 的原因与下一步打算。

As VP of research at OpenAI, Jerry Torque was a part of many of the biggest advances in AI these past years, reing models like 0103 codeex. He recently left OpenAI, citing that he wanted to pursue research areas that were harder to pursue within a big lab. I'm Jacob Efron, and today on unsupervised learning, I got to sit down with Jerry to talk about everything that's happening in AI right now and where the space is headed.

We talked about reinforcement learning, what's required to further scale it up, and the origin of the ideas for 01 and 03. We talked about continual learning and other research approaches and how Jerry is thinking about their promise and the problems to go solve there. We talked about his reflections on seven years at OpenAI and some of the key decisions that had to be made as well as how he sees the race between the foundation models playing out from here.

And we talked about what makes good researchers and ultimately why researchers choose to work at different labs. It's just a fascinating conversation to get to ask everything that's top of mind in the space today with someone who's been so close to building these models and is really zooming out to think about what's next. I think folks will really enjoy this conversation. Without further ado, here's Jerry. Well, Jerry, thanks so much for coming on.

I've been super excited for this episode for a while. I feel like you've been behind some of the biggest breakthroughs in AI over the past years, you know, 01, 03, Codeex, uh, and now have some really exciting next steps. Uh, I know you've left OpenAI to pursue something new. So, I really can't think of a better person to talk about the state of where we are with AI today and kind of all the the potential future directions.

>> Thank you. Thank you very much. Very happy to be here. And AI is one of my favorite topics to talk about. So let's let's do that. >> Amazing. Well, you know, I I I think I'll start with you obviously played a key part at OpenAI in introducing reasoning models and then scaling reinforcement learning. And so maybe we'll start with the existing scaling paradigms and I'd love to get your temperature. How far do you think we have scaling the current vectors we have of pre-training RL?

Like where does that get us from model capabilities? >> It definitely gets us somewhere. The the question is how do we how do we call that place? How do we how do we name it? I figur you could tell me. >> But it is a very very real and for for for most practitioners pretty striking thing that the benefits to scaling are real and predictable and nice. And whenever we scale up pre-training, we get better pre-trained models which fundamentally know more about the world, which fundamentally understand language better and fundamentally build the the linguistic world model of everything around them.

And in the same way, scaling up reinforcement learning makes the model better into acquiring skills that we want to do. And in in many ways, in both those cases, you get what you train for. If you if you want to do next token prediction, you can pre-train models very heavily and get really really great models in next token prediction. If you want a specific set of skills, you train reinforcement learning models and then you get them really really great at whatever you are training for.

Like there there are basically no limits. Everyone everyone knows these days that if if you care about a skill, you just just just do reinforcement learning on that skill and you get them all. That that is excellent. It's it's it's kind of that simple and works. What people hesitate sometimes where where the m moment of holups are how do those models generalize? How do those models perform outside of what they've been trained for?

How do how do models do the models do knowledge that is not in a pre-training corpus? Probably not. How do models do tasks that you don't reinforcement learning on? And probably not that great. And this is this these are basically the remaining questions in the in the world of AI because what we train for we can we are getting really really good at. And it feels like there's kind of two schools of thought here. One is you know look we're still early in this reinforcement learning paradigm.

You know as we scale this up we'll start to see you know more signs of generalization. And so maybe these two scaling vectors alone take us to to most of what we want from AI models. And then the second you know uh being you know really we need something new and different to to kind of continue. uh where do you kind of fall on that >> in some way? I almost think it's a question of the economics. It is pretty clear that if you again if you if you like scaling means in many ways adding data that scaling doesn't work that much without data.

So if you if you add more data that you want then you get better models doing doing what you want. It's what you see right now every quarter, every lab releasing a better model. It mostly means they both together with scaling more compute. More importantly, they added more data to the model and most importantly it was data targeted to what the previous model was was was bad. It's a it's a really powerful really powerful methodology of training better and better models and in that case obviously iterating through it will give you a model that to do everything that you want if you keep adding the data to do the skills that you want that you want the model to do but this this loop is slow in some aspects and then the fundamental question can be can be faster because I I I I believe it is very natural that that with with the paradigms of the methods of training models today.

If we keep adding the data that we want the models to be good at, they will be good at it. And there there there is some amount of generalization within it. But the main question is there is there is there more is there is there some kind of research that would give us more results with less data in in in some ways or more fundamentally better uh better ways of of of most generalizing from what they've seen so far and what they what they've learned so far.

And I know we'll hit on some of those, you know, future potential directions later. Maybe in the reinforcement learning world, just to kind of set the context for our listeners, how do you kind of characterize today like what works and what doesn't? You know, what's your what's your own mental model? Is it literally anything we have uh data for? I know a lot of people talk about the difference between spaces that have kind of are easy to verify versus those that are a bit more difficult.

What's your own mental model for what we can do reinforcement learning on today effectively? Uh the question about the easy and hard to verify often comes very close to what is easy to even get a signal on the quality of like if you if you'd like and at some moment we've made a pretty good progress at OpenAI at training models that that are meant to become great writers and it is is possible reinforcement learning can really been done on a on a lot of great things sometimes meaningfully.

It is very hard thing to know what is what is good or what is not good or you need to wait a lot of time. If you write a book, there are some easy ways to tell if it's a good book or not. But most likely um getting a signal on that you have to try to sell this book and see how many people will like to read it, how many people want to buy it. And sometimes even that is not a good signal because a lot of critics says oh this book is so good just no one no one did no one did buy it because marketing didn't didn't work well.

So how do we how how how can we do reinforcement learning on writing good book? It's hard to say how do people learn how to write good book. It is is a very very hard thing to say. Similar with starting companies that there are lots of companies started early on and how do how do we know which ones are good or not only only five to 10 year down the line we will see that someh entrepreneurs have succeeded and started a really successful companies and some of them failed.

Was it was it because of some actions started early on were good or bad or maybe it was it was a stroke of luck. Uh doing reinforcement learning on that very directly is a is a very very hard thing to do. Everything where you get feedback of any kind you can you can you can use that to do to do reinforcement learning on. >> People have been blown away by the results that you know the models you worked on have gotten in coding competitions and you know math competitions and other things.

And I think a lot of people are still trying to figure out their intuition for do most tasks look like coding and math or do most tasks look like books or starting companies or things that are actually very hard to build uh rewards and also just test uh tons of times. You know, maybe we'll take things like accounting or medicine or legal. Do you have a gut instinct on if those are more like the the former or the latter?

I I think fundamental question is how how easy to tell if you did a good job or not because like arguably even for humans it's a very hard thing to tell if if you did did a if if you did write a good book or not with with with a lot of other work if if if if you can be let's say a manager of of accountants and be able to tell which which accountant is doing good good job or not if there are rules Then with those rules you can you can train whatever whatever you want really for for for medicine I've been I've been thinking a lot about surgeons and there clearly is a number of rules and there are there clearly sense of feedback of human have survived the operation and that that that is uh that that that is an um success criteria that is very good but sometimes and that is that that is a very interesting way where really expert, really skilled humans go against the rules to do something differently that has ever been done before and they succeed through that.

If a if a surgeon really for their experience sees that they need to do this particular operation differently than than all the others before they go against the established practice and through that they do something completely new that suddenly succeeds and makes makes the operation successful. I think the models would be able to do that too with the with with some time and with enough enough ability to try. The question is the question is how much how much time would it get the today's malls to really tr to really get to something like that.

>> As you think about the problems that need to be solved to kind of you know make RL more and more generalizable for tasks that people care about like you know maybe help our listeners understand what are what are kind of the next frontiers for these other domains of RL. Yeah, generalization I think fundamentally is is a property of a model. So, so the the the story is whenever whenever you train, you really affect your training objective and that's and that's kind of it.

And the training objective you can like you get what you train for and the question is like how many other things you are getting for free. There are there there there are clearly methods of learning various bits even nexttoken prediction that generalize poorly. there there is like nearest neighbor classification very very very classific ML algorithm you theoretically can use it for any any ML problem there is it just generalizes poorly because if it has um very very simple representation it builds of the world neural networks how they how they work magically is that through large scale training for large scale pre-training they manage to learn very interesting very useful representations of the of the world around it and kind of it sometimes may seem like we get it for free.

Why how has it happened that through training large transformer on the internet it learns really really to understand well a lot of the concepts around the world. Where does it come from? It's it's a magical part that comes from a transformer and a lot of parameters that we that we hammer over with gradient descent repeatedly. Um and you know this this is the type of generalization we get from this model. There's a question is there a different model that will that will generalize better and almost surely there must be.

The question is how does it look like? >> I've heard you talk before about you know say that you kind of had updated your timelines maybe to be a bit longer for some of the you know different aspects of AGI after working on scaling RL. Um why why was that? I was definitely a very optimistic person in a sense of thinking we do reinforcement learning on our model and we'll we'll get to AGI and maybe we did maybe it already is AGI like the definition of AGI is very personal >> exactly it's a very personal thing and in some ways it is everything we still don't have.

So, so the models that that can solve uh basically any any olympia any competitive problem the models that are meaningfully right now proving new solving new mathematical problems that no one has solved before h we get with with the latest GPT 5.2 to examples of that every every week of of someone achieving that. Is it is it an AGI? Like many people would have would say yes. At the same time, I I'm a big fan of using coding models and they still they still make mistakes.

They they in some places they do things I could take me really really long time to do well. So they can be they can be extreme force multiplier on on doing on doing programming work. But at the same time there there are clearly places where they fail. And I would say the biggest limitation of the malls today is that if they if they fail, you get kind of hopeless pretty quickly. Sometimes you can do a bit of back and forth between pasting an error message.

Hey dear model, this didn't work. Like try harder. Sometime we need to do words of encouragement. But fundamentally there isn't a very good mechanism for a model to to update its its its beliefs and its internal knowledge based on based on failure which like this is this is this is probably the biggest update on me. Unless we get models that can work themselves through difficulties and get unstuck on a on a on solving a problem.

I I I I don't think I would call it AGI uh because because of this of this feeling of hopelessness that if you if you hit a wall with existing model and they cannot solve it, they just cannot solve it that they either try a different model or or or do it yourself. Uh the intelligence intelligence always finds a way. Intelligence works at the problem and probes it until it solves it which which the current models do not really.

Well, I mean it's it's kind of a great transition uh to some of the other research areas maybe beyond kind of the pure scaling of of pre-training and RL and a lot of what you're talking about sounds, you know, uh in a similar vein to continual learning. A lot of topics that people have have been discussing more and more uh you know uh in in public these days. I'm wondering like how do you maybe at the highest level for our listeners think about the set of problems that need to be solved like to make something like continual learning um you know actually possible?

>> Yeah. And um at a very core thing, being able to continuously train a model means being able to have the model not collapse and not go into the into the weird mode or error. Like there are many ways in which training a deep learning model fails horribly. And a lot of all of the work is h happening in the big labs these days is about keeping those models so-called on the rails and keeping the training healthy.

And it's fundamentally a fragile process. It is it is a process that that you have to make effort to go well. And if you if you don't make that effort, it's it it explodes. It's you you just don't get a good mole in the end. And that's seems fundamentally different to how humans learn. I I think human learning is much more much more anti-fragile in a way. It can it can get itself again it can it is fundamentally robust.

It can get itself unstuck throughout learning models with reinforcement learning. I've often really marveled at how infrequent it is for humans to crash out and and and then then start talking gibberish and then and then the brain to and the human brain after getting some new information to to spiral into some into some weird state while the AI models do that they do that pretty frequently and it's it's something that researchers try to find both their theoretical and practical solutions of of of how to how to fight it.

And I I think this this fundamental robustness of a training process is something that is necessary for continual learning. >> How much of the ideas for you know the the maybe some of the interesting ideas for continual learning? How much of it does it feel like has been around maybe for a bit or has been discussed uh versus you know entirely net new research problem? >> The the the main question is worth asking yourself specifically as a researcher.

That's something that I am asking a lot why why why hasn't it been solved yet and that is that is that that must be the number one like when someone starts working on problems like continual learning which I think pretty clearly hasn't been solved yet the main the main question is why why not what what was the particular path that no one has taken so far and why there there are many researchers in the world that are very smart are having a lot of brilliant ideas and so far no one has really cracked uh continual learning and there are there there are many many hypothesis for it but one fundamental one I I think is that most likely it is a research that needs to happen at scale at least at a certain scale and there are only so many well-funded research labs in the world right now that can only do so much research and so many few research projects that that that is most likely One of the one of the really big reason if there is a research that you could do at a at a small scale and fundamentally discover something you know it's probably there were there were a few but either it would be something very complex or very theoretically difficult to do or or or it just requires already models and levels of compute that are available to very few and it's very likely that that the very few labs didn't go in a in a particular direction yet because because they were they were busy doing other things.

I >> mean, I've heard you talk about before this idea that like there's ideas in AI whose time isn't right, but there's still good ideas and certainly we saw this with reinforcement learning and you know, uh it it becoming much more effective after having large pre-trained models to be on top of. So, it sounds like, you know, your your maybe intuition is that there are some really good ideas out there that maybe uh if actually applied at scale would uh would be really helpful toward this domain of problems.

>> I definitely think so. >> I've heard you talk before about, you know, the labs really converging on on working on pretty similar stuff, right? And and I I don't know how if that feels like that's been common over your you know past 2 three years but it seems like when you were leading a lot of the work in you know one that that was a genuinely new thing that you know a lot of the labs were caught maybe flatfooted on.

Talk a little bit about this convergence that's happened maybe over the last year and uh was that surprising to you? Yeah, even when training models with reinforcement learning, there is there is this this well understood and well doumented trade-off of exploration versus exploitation and you wonder when is the right time to try different things than you've been doing so far and when is the right time to try to optimize very very well what you already know well how to do and it's a it's a tradeoff that has no real solution because you don't know what is what is the unknown like whether the exploration is successful or not fundamentally is is there any way that is different from the current one that would give me a lot and unless you know the landscape of what you are doing it's fundamentally a very hard thing to do so uh so it's like I don't think there's a fundamental question about it but I I I remember someone telling me at some in my life how do why why do all the planes all the commercial planes look the same even though there are there are a few companies building those because in the end this is the most economically uh economically efficient design in a in In a way why what all the all the labs are doing today that the forces of economics are fundamentally very strong in that and if you want to if you want to compete you need to have the best models off of the lowest price and the competition is pretty pretty efficient there in terms of in terms of customers can can switch whenever whenever they want and it's really really the customers that are winning for that in many many ways but that is that is one thing that drives the labs to go for higher and higher efficiency and to go for produce better and better models in a in a pretty predictable way.

And then there there there is there's a question of exploration versus exploitation. Uh should we try to go should we try to sail over the sea and see what is out there? Should we try to train a model that is that is completely different? It would it would probably be um it would probably walk away and lose some focus on trying to get to get the current thing better. it you lose some focus in trying to get the current thing more efficient, but maybe there's something 10 times better.

Maybe there's something 100 times better there. There there's a question of of a belief and conviction at the heart of it of of how much do we want to try those other things versus not. >> And to your point, I mean, obviously there's such a clear path forward on adding more and more data, you know, uh to, you know, to RL and different domains and that improving models for economically valuable tasks. there's kind of a clean road map to uh you know to how each of the labs can you know continue to uh you know improve their underlying models that it maybe makes it harder to to go out and make that big bet.

Um whereas when it felt like maybe pre-training was slowing down, it's easier to go out and explore a bunch of different things. >> Yeah, I think I think there are also just just different times uh in in in history. Sometimes there's there's more appetite for it and a little bit of a more of a freedom to explore various dimensions. Uh the more the more competitive the landscape becames slightly slightly becomes harder because it it almost is something like a prisoner's dilemma.

>> Yeah. uh situation where uh where trying to do things things differently uh like exposes you very heavily to to uh to to to to losing market share to other players. >> Yeah. Do you think it actually you know even matters for the labs if the next big breakthrough is is discovered there? I mean one thing I've been struck by is just the dissemination of a lot of you know uh of a lot of these you know advances.

Obviously you know 01 you were kind of the pioneer on the reasoning side. There's a few labs now that have great reasoning models and I almost wonder if if the the labs would be just fine if if the if the uh breakthrough happens somewhere else because you know these ideas diffuse and and eventually they'll be able to plug it into to the existing business >> ideas diffuse and and that's that that's a good thing but at the same time the lead it gives you to do some to be the first I don't think I don't think it's something to really uh to to really discount um in Anyways like we have seen the fact that like you know if you if you believe that fundamentally there would be no place in the world for open AI to succeed and did succeed because it went into pre-training transformer large scale much better than anyone else and that lead made it one of the largest and most successful company in the history of the world and in the same way because of open AAI being the first one to figure out how to do large scale reinforcement learning I think I for a long time and still until today.

I believe that it has the best reinforcement learning research program out of all labs, which allows it to to do things better and more ambitious than than most other labs that had had to catch up to this to this much later. And while while the ideas diffuse, the lead can be a very very powerful thing that if maintained, it will it will stay with you for maybe maybe potentially forever. In many ways, the I've been reading book about about semiconductor manufacturing which many many of the core initial parts of the invention were done in the in the United States and through that they they slightly disseminated through the world through to various places but at the same time there has been like moments and and and space of lead that a lot of other countries couldn't ever match the they continued good compounding of advantages over time for some of the countries that bet on it early and really really start hard to build it.

And it's not that there's only one country building having having successful semiconductor business, but also not not every country. It doesn't it doesn't exist everywhere in the there always is a place whenever there is business shift there. There there are some newcomers that will be successful. There are some newcomers that will be unsuccessful and some old companies that will that will stay and manage to turn themselves around and some old companies that will die.

That's the that's the Darwinian part of of progress. >> No. And I feel like consumers and enterprises always remember the first company that that introduces them to some, you know, pretty magical experience. I mean, certainly you guys experienced those with chat GBT. One thing that's fascinating about you is you obviously made all this incredible progress in in RL and you know we're help pioneering a bunch of this and RL progress is still alive and well and and going along and and we're making lots of progress across domains and and you decided to leave OpenAI um and I think you cited kind of different research areas you wanted to explore.

Um I'm curious like when when did you kind of begin to know that might be something you wanted to do? Um and and you know how did you ultimately make the decision? >> It's it's it's definitely not something that happens very quickly. It's a it's a just something that slowly grows in a human and opening eye is not not an easy place to leave because I have many friends there a lot of whole shared history a lot of I I I I've built a lot of a lot of my life there and for a long time really really tried hard to make it to make it work and try to uh try to see what are what are the various place to do it but at some moment if you spec specifically as a researcher If you if you wake up and if for for for any any reason figure out you don't love your work anymore, you are not incredibly incredibly excited about what you are doing, then it's a good moment to try to explore and try to do something something else.

Like it it is basically impossible as as a researcher to do your your best work if you are if you're not 100% excited and if you don't go with with full enthusiasm. There there were many many days at OpenAI when I was when I when I had basically infinite enthusiasm for the work I was doing and was believing I was doing exactly the right things but somewhere somewhere around at the end it was was getting it was getting harder and harder.

Um so so like you know that's that's that's kind of kind of long story short. >> Yeah. What what are some of the things that are giving you uh energy today? Um I think on on the most fundamental level what I did and what I what I've done I when I started open AI I believe reinforcement learning as a necessary element of a path to AGI and I really wanted to make it happen and I and I really really did and like introducing reasoning to the world reasoning models to the world has been like you know to me a tectonic shift and a paradigm of of how how we train the malls and in some ways I want to chase that high again and try to do something something similar.

Try to find something that is missing in how the world is training models so far and try to try to try to make it mainstream as again in one way in one way or another. But once you once you did something like that you you know [snorts] it's it's it's hard to hard to get get another another similar hits doing something else. So that's that's kind of what I would like to do and I would like to have a bit of a bit of freedom thinking about how to how to explore it and try to attack the most the most core the most the most important problems there are.

>> How much are you feeling like hey I've got dozens of hypotheses or or how much are you zooming out and being like you know I've been so heads down in open AI let me zoom out and kind of see what else has been going on. you know in in general the the the real important hypothesis and the real important problems are most likely not something that will that will be new and that will appear to you if you if you worked on machine learning for seven years.

It's very likely you know what the what the important problems are. The main the main question is what I what I alluded to before which means how do you solve it differently than everyone else because because it means no one no one has solved it yet. So what is what is different and what is what what what can be done what can be done differently than that that than than people in the past. >> Well I mean I definitely want to hit on you know uh OpenAI and obviously you had such an incredible run there and and and time.

You know one thing I've heard you say before was you you know I think you've been at OpenAI since 2019. you said every year felt like a different company in in some ways. Um and so I'd love if you could just walk through that evolution and kind of you know the the how you kind of describe the narrative of the past uh six seven years there >> starting with a small lab that that is like 30 40 people where we were always open from the very beginning was extremely ambitious believing this is this is the place that will build AGI and that will that will create benefits of of digital intelligence for for the world.

But starting from a few people trying to do a few kind of cool projects that were that were incredibly ambitious to to where it is today which is which is again one of the largest companies in the world with a product that everyone knows and everyone uses and it's it's almost hard to imagine not using it. It's been it's been it's been a wild ride and as you as you also are are are aware the execs at OpenAI have shifted throughout the years pretty significantly.

So the type of people you work with every day has changed a little bit. The size of the company has changed. The the themes of research have have changed. It at some moment in in in the old days there was no pre-training at all. Then for a while it became kind of the pre-training company. Then for a while I think it became pretty much largely our old company. And now now it's it's a little bit more in a in a balanced way uh of of of a mix of those two with with many many people living OpenAI and building both businesses and then that their their life outside and many new fresh great people coming in and doing doing incredible incredible research be research inside it's know it's just a company that manages to reinvent itself and manages to to grow through that through all the stages trying to imagine I always always had this weird hope thinking of all those big successful companies how incredible it would be to live through such a story of of being there for those for through for that through those stages and I feel like I've been through for through quite a bit of that open AI and it's it's an experience that I think is very hard to compare with something else >> you know I think everyone's eagerly awaiting for the uh the definitive stories to be written about this this chapter of open AI when those stories do get written I feel like people always like to focus on the kind of you know the difficult you know 51 149 decisions that could have gone either way that really like moved the company forward.

Are there any kind of of those pivotal ones that that stick out to you? >> Good question. Um, you know, I was only only central to some of them. There were probably many that I didn't that I that that I like only only was a was a background character and you know probably even even this discussion about releasing Chpt to the world or not as as you may have heard and many people did this popularity and its its virality total didn't like what wasn't expected internally by at least like no one I heard about it.

Uh I I think in the end with with TG GPT and GPD4 released soon thereafter we created a bit of a moment and a bit of a momentum that was um incredibly hard to predict but made opening eye largely what it is today. Um that was that was definitely a very very very important call on on on many axis and many decision on many many dimensions like the decision to pull a lot of resources to train GPD4 at the moment it is and with that with a lot of trade-offs that went there again remained very very important and very critical in the in the history of of open AI turned out to be really good decision and this in the same way betting and and saying reasoning malls are our future in a world where it was completely unsure of and just just a bit of first principles thinking and and a bit of a a bit of a intuition that this is the right thing to do allow the open AI to completely reinvent itself and say we are we are doing recently models right now even though there is no product market fit even though they seem to be kind of cool with puzzles if if you look at a one like it was was a smart model but it it wasn't really good for anything practical ical except for except for just destroying around hey we have a model that that is kind of smart really only with with 03 and with a little bit of more investments into into tool use with those models we're able to start building something that started being incredibly incredibly useful for research for coding and from there once you have the first signs of real product market fit then humans are really really good at optimizing something that exists and that they can see it works but getting to that moment has and great great journey and something something to study because it was it was not an easy thing and open at that moment really really passed the exam.

>> Yeah. No, I mean I think what you describe is so interesting of this idea of you you know have to keep scaling and investing in something you know uh not quite knowing if it's really going to work or it's working. It's obviously very relevant to some of the things you're thinking about in the in the future too, you know, with the with the reasoning models like, you know, uh was it clear after 01 that this was going to be more than kind of fun in games or what was like the the moment the spark for you that you're like, "Oh, this is really going to work and and we can really scale this."

>> I I kind of believed in it from the very beginning just just because I believed in reinforcement learning. Again, my my core belief from my my very first days of open AAI is that reinforcement learning is a necessary part of getting to AGI. And that was the the main way was how rather than rather than if. The question is when are we ready to do it and how exactly do we do do we do it? And I fundamentally well just just silently I started I know I know this is this is the this is what we what we need and over time and over over research there there came various experimental results that that have informed us this this is the right way of doing it.

One thing that's so interesting about open AI is obviously um I think people liked I think Ben Thompson coined like this phrase like the accidental consumer business. The idea that like obviously you were a research lab pursuing AGI and then kind of almost accidentally stumbled upon this this incredible consumer product uh that that was uh you know immediately picked up by the broader world. Um and I'm wondering you talked about some of the kind of different chapters of of being there.

I think a question everyone always asks about OpenAI is the company's doing so many different things. you know, they got the consumer product, core research, you know, Sora, the enterprise product, like how does it actually work internally and and um you know, how how do you kind of uh do you feel that tension at all on the research side of getting pulled in a bunch of different directions? One thing pretty clear is that opening eye research operates very separately from the product and has from the very beginning almost almost like you said which is that open eye's goal and mission is to build intelligence and that is I think the main goal and motivation of majority majority of research there is there's like one team specifically built and we've directed towards product research and that's that that that is the part of research that that optimizes uh for for for whatever whatever product metrics it is.

And the rest of the rest of the research mainly focuses on how do we how do we make our models more more intelligent and and that that tension really doesn't exist there. What I think this is is real and interesting is that openi is at the center of probably the biggest technological shift of our lifetime which means there is so much opportunity to do various things. It is almost it feels wasteful and it feels imprudent to not try to do all those things because basically everything in the world will be disrupted by AI.

But it has a downside which is very real and very very problematic here of focus. Companies are very bad at doing multiple hard things successfully. [snorts] um like companies are well known for succeeding at one very hard thing and then doing others similarly well to others. There are very very few places in the world that can do a few of those things and that is very very hard thing to do and I think this is a very very very big risk for open AAI try to do everything and then and then not succeed at it.

uh but but like it will we will see if if if like open AAI in a then half focused state can execute well on all those side bits or there will be other companies famously and I think I think it's a little bit sad OpenAI um really lost focus on coding for quite a while when it focus on the on the consumer product and that has costed a bunch of bunch of markets sure that it is uh working working very hard on regaining right now and I coding open ice, coding malls are are really really but lost focus and lost lead definitely definitely has a has has a cost here.

So um there are there are like in some way like when you are a company doing AI in the world right now you feel like a kid in a candy store because there are so there's so much potential of of extremely valuable things that can be built for the world that it's hard to prevent yourself from doing all of it but but like for for everything there is competition and there are there are question who will who will do each one of those things in exactly the right way.

Yeah. Well, I mean, I think that's a great uh, you know, point to transition to just like the general ecosystem today. Um, and you kind of alluded to coding, which I think has been a really fascinating space to watch play out. Why do you think Anthropic's been so successful at coding? >> It's focus. I I think focus can explain 95% of things. What are what are what are companies why are companies succeeding in in things?

I know I know Antropics founders from even the time when they were at OpenAI and they were always always extremely fond on coding and they always believe that it's a it's a necessary and critical part of AGI and I I think I can I can only imagine how focused they have been on it on it over the years and they they definitely managed to get their their vision very far with with with the latest model with cloud code and coding agents and I I I I they they are saying truth when they say very few people in entropic type code themselves these days >> and do you think that's kind of a foreshadowing almost of you know pick the the major labs that are out there like each can kind of focus on different things and so you end up with uh you know models that are particularly good at at at some things in in different labs it's it's a good question and I think there are multiple worlds in which we can be there are there's a world where data matters and in that case data is a very zero sum game where like you put the data into the skills that you want and your mall is better at that skills and in that case we see splintering of the the market into >> basically you can shift the data mix of what you're putting in but ultimately it's at the cost of some other skill.

>> Exactly. and and a little bit about about shifting the data mix but mostly shifting the work. Data is a is a labor intensive and like slightly more less money intensive but but it's it's it's largely very much labor intensive work in terms of in terms of research engineering let's let's call it that way and so so it's just just you have so many researchers that can work paralle in parallel doing doing doing doing useful useful work on preparing the next data set.

So if if if data drives the improvements then we will see different labs being being better at different things in natural specialization trade-offs. But if research is king, I think I think research has this magical property that is is hard. It is is high risk, high reward. But at the same time, if you have a good idea in research, it could improve your model in all domains at the same time. And you can you could leaprog everyone in all domains naturally by just by just training better models and and like which one which of the world which of the futures we are in it's it's very hard to say right now.

Uh but we'll we'll we will see we'll see if with with in the in the future which one >> I mean >> this current paradigm feels almost like very anti- bitter lesson right where everyone's going off into specific domains and like you know specializing and and really putting in those and you're you're right that it feels there's this intuitive feeling almost that there must be something else that is that is a bit more generalizable.

I'm I'm pretty sure there is the question is like how easy it is to find. So there there is one one version of the world where which is not completely impossible uh although although slightly pessimistic on humans that says coding agents are so good right now. Let's first get them to the moment where they can automate AI research and then have the models like like research better models because because maybe we are at the at the last mall that human have humans could have figured out and no it's not not a completely impossible uh framing it's maybe it makes sense with all those GPUs we have and all those all those really really capable models and their tenacity maybe maybe they should be uh researching future models but but you know maybe maybe there there are still a few things humans can do that we can that that that we can put our put our heads to work.

>> I feel like a lot of the top AI researchers Yeah. are working on coding for that very reason, right? This this belief that it will will speed things up. Um you've obviously worked on coding too. I mean you spent time uh you know a lot of time on codecs. Uh how do you think about like the next frontiers for these AI coding products and it seems like we're on this like just crazy exponential hill climb right now.

How do you even conceptualize like what these products will be able to do in a year? Fundamentally the whole the whole story of coding is that we are able to program computers at a higher and higher level of abstraction in many ways and we we need to know less details. I need to track less details and coding agents in some ways can be thought of as a higher level programming language that has very different semantics from all the all the all the other programming languages.

And um you know I I think it's a trend that will I think it's pretty unlikely that in the future we will be typing code ourselves. Very few people already do and it's it's a it's it's a one-way ticket. But at the same time software is important. Software needs to be reliable and there will be more and more progress in how do we get certainty about software doing the right things if we are not the ones typing it and maybe not even the ones reading it.

Fundamentally I think those are all solvable problems and I I'm very excited about them. I think it'll be interesting whether the the core skill set to work with these agents, you know, looks somewhat similar to to a software engineer or really everyone just becomes a PM. Uh, and it's really just about knowing, you know, having some idea of what to go do. Uh, and and and the models going for it. >> Yeah, I think it's it's an interesting question because I think almost the most important skill right now is being a skill of a good manager of junior software engineers.

A lot of a lot of software engineers have been have been like a little bit reluctant of going to management and really really like to uh specialize in in in in narrow domains and I think like for for a long time it was it was really really the right thing to do and understanding deeply systems is incredibly important. You need to combine deep understanding of computer systems with being able to give back at least a little bit of control which which what I think really the best managers are the best managers understand extremely deeply the work of of people on their team but at the same time are able to give back some control for the people to drive their own their own destiny and that's probably the best way how to work with malls these days.

You you you you you kind of everything that the malls do you you understand the trade-offs and you understand what they are doing but you allow the malls to make their own choices and and live with the consequences of them one way one way or another >> you know around coding one thing people are also trying to figure out is just where the various applications kind of fit in and I think you know there's uh I guess a few questions like one being um you know hey like you have the codeex the cla code team sitting you know right by the research teams they build these great harnesses and products like you know how do you think about the you know the opportunity for companies like cursor and cognition and to to what extent is it like a disadvantage to not be you know sitting next to the researchers of the labs.

>> It definitely is a disadvantage. I think I I think the fact that the most successful companies of the world are training their own models is is telling about something. And I I kind of think just like the future of big AI companies is to become hyperscaler and run their own data centers instead of renting compute, the future of successful AI application companies to start training the models themselves. This is this is just how the how the stack works.

But you got to start somewhere. So, so it's it may be the path that you start with with the AI application that is successful and then you first start post-training your own models and then you start pre-training your own models if you are if you are more and more successful and then you start building your own data centers. Uh those are those would be just just the natural natural paths and stages of of success of being a good being a good AI business.

So you think it makes sense for these companies to kind of do reinforcement learning on their own user data, you know, post-train a bunch and I mean do they have any hope of catching you have, you know, endless compute going toward the big labs, you know, endless collection of talent like if you if you focus on a very, you know, specific domain, do you kind of have a hope of building a better model there or uh is is it kind of a almost a hopeless task in some ways?

>> Well, not nothing ever is hopeless task. The future the future is not determined. But some of that is what I what I hinted before which is is is the data important or is the research important. If the data is important, you can always try to differentiate yourself with the data. But it is is not clear we are we are really in that world. Maybe there's a there's a world where research is important. But that also allows the smaller companies to do some research that they think is better [snorts] and maybe maybe win through win for research in the market.

>> But it seems like it it almost requires a world where you know you have specific types of models that are focused on things being better, right? and not generalizing, right? So, we get into a world of generalized models. It feels like it becomes hard for an application company focused on any specific task to have like the better pre-trained or large model. Sometimes innovation comes from constraints. I think it is possible for a company focused on a specific domain through seeing the deficiencies of the models on this domain to create a model that is generally better.

I I don't think it is possible and that could be like the next layer of of success of this company. If you try to do really really good model for X and suddenly through that you make the best model for everything else and then you grow and then you become another another big successful company. >> Feel like in the past the problem's been you get better for like one second and then the next generation of models come out and you're like oh man we're we're way behind again.

competition is is difficult and definitely we have seen seen for a while in the in the US tech landscape how big companies have tons of advantages and this is this is true but at the same time there are there are new big successful companies coming up so uh it's not it's not hopeless it's just hard >> well I want to you know uh shift gears to the talent ecosystem and maybe you know research itself because obviously you've uh both are an incredible researcher and have worked with amazing researchers maybe to start obviously researcher hiring is very competitive today.

I know you were probably on the forefront of of bringing folks into OpenAI. What determines what companies researchers join today? >> Damn, it's a good question. In the end, people are very complex, definitely more complex than the models these days, which means everyone's incentives are different what what they want. And I I I honestly cannot really generalize. Whoever is hiring people should not think about how do I convince the most of the people to join me and how how how do I become the most appealing place for for researchers to do it's probably a good question to be to be asking yourself but I think there's a second one which I think is much more important what what what what type of researchers like would really want to work here and then and then find those because it's it's kind of impossible and and very very difficult to try to appeal to everyone just just just because of of diverse diverse preferences and diverse viewpoints and diverse ways of working.

So it's it's much better to try to build team that has like you know some shared values, some shared approaches because it's pretty clear that the that the teams that are that are aligned and I have that have the same the same goal move faster and work better that the teams than the teams that are not. So like you know it's it's really should be should be kind of filtering on both sides and trying to just find the right people for the right for the right group and that makes everyone happy that makes the group successful and that makes that that that makes the group more attractive over time.

>> Yeah but there's been some interesting experiments around this right I think you had meta famously offering you the first offer these like mega packages um you know what's your kind of reaction to that? Well, like you know there there are different strategies how to build a research group and you know Meta at some moment I think like you know there there are supply and demand curves but they needed to make offers really attractive to start bringing people back after a few missteps in the in the in the world and you know the momentum is a very hard thing to stop if if if like if for at any moment there is a perception in the industry that you are not doing very well then that then then you will not hire people and then and then then it can it can roll itself.

So you know in many ways I think it was a really really good strategy to try to change that dynamic and change the negative momentum in a place where AI is so important for every every ever every large scale business and Meta has a very much new team built out that is that is training a new model these days and a lot of a lot of people in the industry are watching will it will will be successful and how how it will be successful and it will be determine the future of this lab But but it was definitely like, you know, a good moment to bring in some some new life into the into the meta efforts.

I mean, you've obviously done a a ton of groundbreaking AI research. You've worked with other great AI researchers. What makes a great AI researcher? It's a it's a good question and it's a it's a difficult one because in in many ways being successful AI researcher in many ways it is about just being in the right place at the right time but >> but like there there is something to that I think the fundamentals of being a great AI researcher these days but but in the in the span of my career are one thing being extremely good at both systems and engineering levels, understanding how computers work and how neural networks are trained together with the theory of of neural networks and optimization.

Um, like it's very very hard to be very successful doing only one of those things well and suddenly if you are at least okay in both of those things, it makes you easily 10 times more productive in any in any research endeavors. And I think this is this this this is an important part. Um like the the other one it is like if you want to be a successful researcher you very necessarily need to have some ability to think independently uh be able to be able to get away from from group think.

um people have kind of some natural tendency of converging on a median viewpoint of a group which is which which kind of kills research. I I have a saying that you have 100 researchers that think the same thing. You essentially have one researcher. Being a researcher means being slightly contrarian all the time because you want to work on something that is that is not working yet and that by default people don't really believe in.

And that's being contrarian means like there there there's something that a lot of researchers that are extremely brilliant and extremely hardworking are like unfortunately missing is something you can call it courage. In some way it is about standing up and saying let's try to do something different. Let's let's do something in a way that most people don't believe in yet. And it's it's an extremely hard thing to do especially in a world where where experiments are so expensive as they are and they involve so many things in a world of of machine learning experiments being the cost of Hollywood movies.

uh directing one is like you know it it does require a lot and you don't know like similar in the movies you don't know if the movie will be successful but with a big budget you risk at buying stars and and and doing the best the best CGI you can in the world of ML you also try to try to risk and bring as many as many things as you can stack as many advantages on your side but in the end experiments are experiments and they are meant to be go into the unknown and and either either will be will be successful or not.

>> But like I think I think like if I were to say this is like knowing both systems and theory very well. It is about not following group thinking that much and then and then having a courage to to state that to other people. >> Well, we always like to end our interviews with a quick fire round where I stuff like all the other questions I couldn't fit in elsewhere into the end here uh and get your your rapid reactions.

And so maybe to start uh what's one thing you've changed your mind on in AI in the last year? >> You know, probably the last thing I I meaningfully updated on is that I don't think a static model can ever be can ever be AGI like that that continual learning is a necessary element of of of what we are what we are pursuing. Is that just because of something it won't be able to accomplish or is it more that uh you know that that it doesn't just meet the like the definition of AGI needs to have a continual learning aspect to it?

>> It's mostly about uncovering what our models are still missing and then in many ways you you go layers and layers deep because because our models are already good into so many things but them being as as as brilliant as they are it is clear that that without it it will never feel to me that that intelligent. will still be still be a tool that needs to be supervised by someone who has the ability to continuously learn.

>> Yeah, there's obviously a lot of AI progress happening in in spaces we didn't talk about today. Like what's your kind of timeline for I don't know a chat GPT like moment in the robotics space. [groaning] >> Probably around 2 three years from now. >> Uh that's pretty bullish. Uh I feel like everyone's still trying to figure out if there are scaling laws in robotics or if the stuff will will actually you know if there's enough data.

I I think like you know honestly between you and me I think things are slightly better than most people realize. There are tons of companies making tons of progress but but always the progress needs some time to play out and some more investments to happen but I think robotics will be doing pretty well over the next few years. >> Yeah. And in a similar sense on biology >> I think I think biology will take longer.

>> Why why longer than robotics? M like ju just thinking from a perspective of how much intelligence there is needed, how much precision to to really manipulate biology successfully. It's a it's it's a harder problem and require more fundamental investments to to to start to work. >> Yeah. I guess most like three four year olds figure out how to manipulate things in the world, but they're not like worldclass biologists and so >> something like that.

Yes. Um, what's one impact of like this continued model of improvement we're on that you think maybe we're underestimating or not talking about enough as a society? I know it's hard to say but like fundamentally widely deployed work automation will be reality over the coming coming decades and in one hand we are talking about it but on the other I think we are not talking about it enough and seriously enough because because the world will change the world will change very drastically from where it is today.

If if if there are some people who it is still not obvious today like it is it is obvious to me at least and like societal changes are slow and I think I think it will be extremely weird. I think it will be probably painful in some ways and we should try to figure out how to make it the least painful but but like we need we need to think about how how does the world look like where where they the job market is looking very very differently than it is today.

And I guess related to that, one thing I'm uh interested in obviously been thinking about is you know whether you know your kind of thinking about this has impacted the way at all you you know you act as a parent right um and how you think about I guess uh you know raising kids and and if AI has changed that at all. the the interesting thing and I I think I differ here to most parents, but I don't know. I'm I'm not I'm not my my daughters are very very young and very very very small, so there isn't there isn't really that much, but I'm definitely not pushing them very very hard to study these days.

It's kind of hard to imagine how how hard you can push seven-year-old to to study. But I don't think I will be the parent saying, "Oh, you have to be the best in maths in your class." I don't think you will be like oh you have to win all those competitions. I don't think you have to be like you have to read all those books because like you know being a specialist like when they will be grown up probably will like look very different and it doesn't make sense to really really bet on like trying to find for your place in in a job market anymore in that in that way.

I like there clearly the the ability to think critically will be always important because if you don't if you don't think for yourself no one else will even even even AI uh will not will not fully advocate for you in the way you want but I given given how how many unknowns there are I like at the moment just just want them to have happy childhood and have have a good time as much as they can and I hope I hope they will also have have happy adulthood if we don't if we don't screw up this this this AI deployment thing.

>> Yeah, I guess that you know everyone likes to talk about the existential risk side. Have like has your worries about that gone gone up or down these past years? I fundamentally think and hope that no human wants the the humanity to go extinct which is a pretty good like the incentives are aligned for everyone in the world for the existential risk not materialize which which I I really am not like subscribing to that that it is that easy to just make a mole that will that will that will paper clip all of us if I thinking there are enough humans in the loop through all throughout all the process that just because it is in no one's interest that we will manage to be to be successful enough and the people controlling the biggest clusters in the world will be will be responsible enough because it's also not a good business to kill everyone so uh capitalism is also is also fully fully aligned here I think it's much more dystopian and much more worrisome if we like push entertainment so far that that is more interesting in the real world and the humans will want to live in VR and only only play virtual games which is like not not impossible.

I think much more much more realistic than all of us going extinct but but probably probably similarly bad. >> Yeah. Now I feel like Ready Player One and Wall-E and all these kind of like depictions of of of that world before. >> Yeah. So that that that to me is much more much more worrying. But this one we can only solve ourselves really and this this is not an AI problem. this is a human problem and what's what what what what what should our preferences be?

So that's that's that's kind of how I've been how I've been thinking about it. >> Yeah, I love that. I guess anything you're you're comfortable sharing at this point about what what's next for you? >> This is still still very early, but I am thinking about it. >> Awesome. Well, Jerry, this has been a fascinating conversation. >> It has been nice chatting with you and thank you very much for inviting me here.

I'm Jacob Efron and you've been listening to Unsupervised Learning, a show where we probe AI's sharpest minds on what's true in AI today, where the space is going, and what it means for businesses in the world. I love doing this podcast alongside my day job as the managing director at Redpoint, where I've led investments in companies like Lora, a bridge, and physical intelligence. If you enjoyed it too and found today's episode valuable, please subscribe on YouTube or follow us on whatever platform you're listening on.

It's the best way to support the show, which helps us to continue to grow and get access to the best possible guests. Thank you so much for your support and listening. We'll see you next episode.