🎙️AI 访谈库
Begin Proof——对话 Noam Brown
Noam Brown · OpenAI

Begin Proof——对话 Noam Brown

Begin Proof — Noam Brown

2026-06-04 · Begin Proof (Bain Capital Ventures) · 44m · 约 49 分钟读完 · 原文
Begin Proof 首期节目:OpenAI 模型推翻 Erdős 平面单位距离猜想的幕后,大规模测试时计算对数学发现的意义,以及推理模型下一步。

The models have reached a point where they have surpassed my ability to really understand what they're capable of. The sense that I get from talking to people that are more familiar with the space is that the models are extremely good at combining a lot of different areas of mathematics. It's arguable that they're not superhuman yet, but we're also seeing rapid progress on that side. But over time it will just become superhuman in basically every respect.

And I don't know how long that transition period is going to take. It could be It could be a couple of years. It could be could be 10 years, but I think we'll get there eventually. >> Welcome to the Game Theory. I'm Slater with Bain Capital Ventures, and today we're talking with Noam Brown. At OpenAI, Noam focuses on multi-agent reasoning and test time compute scaling. He was one of the key people behind O1, OpenAI's first public reasoning model.

Last year his team swept the math and CS contest circuit. Their model won gold medals at the IMO, IOI, and ICPC. Today we'll talk about OpenAI's recent disproof of the Erdős unit distance conjecture and about AI for math research in general. Noam, thanks for coming today. >> Of course. I'm great to be I'm happy to be here. >> Well, you know, we're so we're sitting here today. It's about a week after OpenAI announced you sort of disproof of this Erdős conjecture on the planar unit distance problem.

And it's the first thing that I want to talk about. And you know, maybe my first question is why is this happening now? You know, I think the IMO results, which are already incredibly very incredibly impressive, happened about 10 months ago. You know, sort of why are we seeing the first sort of major results in math research today as opposed to 10 months ago? What's the different about the models or the systems? >> I mean, in many ways it's it's not that surprising.

When we got the IMO result about a year ago, we expected You know, we said that progress is going to continue, that we'll see um more significant results in math. And eventually we're going to reach a point where the models are proving things that humans can't prove. And there were some initial like signs like this over the over the course of the year when the model started proving things that were unproven. Um, I think there were questions around like, well, was that just because humans hadn't paid much attention to it or was it doing things that were just like, you know, not that interesting.

And um, I think what's cool about this result is like it's it's perhaps the first case of like a model proving something that was like really interesting to mathematicians and also a very significant problem. And we we kind of felt like this was going to happen eventually and it just, you know, the models kept becoming more capable and um, it just finally happened. And I should say this wasn't like, you know, it wasn't like me or my team that like enabled this result.

It's just a natural consequence of the models becoming more capable. Um, and then, you know, some researchers at OpenAI decided to like, okay, well, is this at now at the level where it can actually prove a problem like this? And they ran it on that problem and it turned out the answer was yes. >> Yeah. Were you um, one thing I wanted to ask about is sort of like domains in math that you think are you're sort of most optimistic about?

Do you feel like combinatorics is an area where, you know, you were especially interested or you felt like rapid progress um, with AI was sort of more possible or, you know, sort of is it that behind the scenes you're sort of trying lots and lots of open problems and this just happened to be the one um, that fell first? >> Yeah, I don't think there's anything um, special about combinatorics. I don't think I mean, basically I'm a result of anything I was like a little bit bearish on combinatorics.

>> Yeah. >> But um, it it just it just so happened that yeah, we just tried a bunch of different problems and this one the answer came back like, oh yeah, it seems like this one's correct. >> Yeah. Well, one thing I was interested in is um, you know, the construction itself and it looks it's it's very interesting because it's the kind of thing that I think a human mathematician would have really struggled um, to come up with.

Um, you know, it's a very complicated construction. It has like some analogies to the original um, Erdos argument, but it's it's quite a lot more complex. And you know, you guys released this kind of companion piece with commentary from mathematicians on the results. And you know, I think it was Jacob Zimmerman who had remarks of like, man, you know, I tried something like this a while ago, but it just was sort of too complicated.

You know, sort of the waters were too treacherous, I think was his phrase. And it sort of raises this interesting question about whether AI systems are going to be better at math in ways that are really different than humans. Like they sort of might be maxed out in different dimensions. In other words, you know, in the specific case, like maybe even a human mathematician given the prompt of try to find a counter example to this might never come across this.

And I am curious for your commentary on that. Do you kind of feel as though the models are better at math in a way that's different from humans as opposed to saying being like a human, but just, you know, some factor some factor better. >> You know, it's interesting like the models have reached a point where they have surpassed my ability to really understand what they're capable of and like what they excel at just because like already the you know, these kinds of problems like understanding the proofs so is just beyond my ability.

I'm sure it's beyond most people's ability, but it's also beyond my ability. And so it's hard for me at this point to say like, oh yeah, these are like the weak points of the models and this these are like the strong points of the models in the field of math. But I the sense that I get from talking to people that are more familiar with the space both at OpenAI and external mathematicians is that the models are extremely good at combining a lot of different areas of mathematics.

And they're also like in this in this respect I would say that they're superhuman at this point. >> Yeah. >> And then there's the separate question of like reasoning through the the the complex problems. And this I think like the argument it's arguable that they're not superhuman yet, but we're also seeing rapid progress on this. And so I think we're going to see this kind of spiky intelligence where um the models are going to be superhuman in some respects and then also not quite superhuman in other respects and the point but over time as the models become more capable they are going to become like more capable across the board and that's what we're seeing.

It's that it's not just like oh they're getting better and better at ingesting a lot of mathematics and combining it. It's that they're getting better at that but they're also getting better at just like the fundamental reasoning capabilities. So I think we will in the short term see more of these more more results where it's like combining different areas of mathematics and in novel ways that that nobody thought to do before and then having it be like really cool proofs in this kind of way but over time it will just become superhuman in basically every respect.

And I don't know how long that transition period is going to take. It could be it could be a couple of years, it could be could be 10 years um but I think we'll get there eventually. >> Yeah. What do you That is one thing I was curious about is if you speculate that there is this kind of like centaur period for sort of like human aided math or you know if you think that that period is too short to be sort of interesting in the long run.

Curious what you think about that. >> Yeah, I think there there definitely will be a centaur period and this the classic analogy which I think is appropriate is to things like chess and go where in chess you have Garry Kasparov losing to Deep Blue in like 1997 and then there is this like 10-year period where okay chess AIs were better in some ways but you know they weren't like the human plus the AI was actually better than the human alone or or an AI alone.

And at this point the AIs are so good at chess that like the human doesn't really add anything and so it's like clear that it is beyond the human's ability to add anything to the to the table. Um but and and I think we'll see something similar with mathematics where okay the AIs are like extremely effective in some ways but they're not and and we might reach a point where they're just like if you had a choice between a human mathematician or an AI mathematician, you would choose the AI mathematician.

But even when we reach that point, you could argue that the human plus the AI would be more effective than either alone. >> Yeah. >> Now, how long that period is, it's hard to say. It it it could be very short, it could be it could be long. I I think a lot of it depends on uh on the trajectory of mathematical progress. I think also this is where the analogy to things like chess and Go break down because chess and Go became superhuman because of self-play and the ability of the AI to just like get collect infinite data and like play against itself and get arbitrarily better.

And we don't really have a similar version of self-play in mathematics because it's not a two-player zero-sum game. So, you it could take longer for the capabilities to improve. Now, that said, I mean also the capabilities have been improving very rapidly and so I don't know, it's a trade it's like kind of a trade-off. It's kind of I'm trying to balance the fact that like AI progress has been extremely fast so far um versus the fact that like there is the possibility that progress slows down as you reach superhuman capabilities.

>> Yeah, I've um you would know much more than me, but I sort of heard like speculations around the self-play idea of like, you know, you should have someone propose one model proposes problems, the other tries to solve them. You can kind of adversarially make the problems more difficult in some way, but maybe it's it's like very hard to get to work. It's it seems way clear less clear how you would do that than in like a like a sort of very um sort of boxed setting like a like playing Go.

>> Yeah, this is like I think people under appreciate how unique two-player zero-sum games are when it comes to self-play. Like in in these two-player zero-sum games, you have a very clear objective, which is you want to converge to a minimax policy. And you can do like you just have the AIs play against each other. You don't need any external data. You don't need any human data. Just have the AIs play against each other and figure out how to beat each other and then learn from those mistakes and they will converge to the minimax policy.

When you go outside of a two-player zero-sum game, you don't have this property anymore. So, an example I like to give is the ultimatum game. Are you familiar with the ultimatum game? >> No. >> Okay, so the ultimatum game, it's a non-two-player zero-sum game. One player, Alice, has $100 and they have to choose how much to give to Bob. They can offer anywhere between $0 and $100. >> Mhm. >> And Bob chooses to reject or accept.

>> Mhm. >> And if Bob rejects, then both players receive $0. If Bob accepts, then the money is split according to what Alice proposed. So, like let's say Alice offers 20 bucks. >> Right. >> Bob accepts, then Alice gets 80, Bob gets 20. And then if Bob rejects, then they both get zero. >> Mhm. >> And so, you know, you could imagine what self-play would converge to in this game. If you had to guess, like what what do you think it would converge to?

>> Like always accepting. >> Yeah, always accepting and then like Alice offering basically like a dollar or something or a penny. >> Yeah. >> If you were to do this with humans, it would not go very well, right? Like if I were to offer you like, okay, I have 100 bucks, I'm going to offer you a penny. You have a choice of accepting or rejecting. Like >> Yeah. >> how would you react?

>> Yeah. Yeah, I would >> So, so like once you go outside of two-player zero-sum games, this this idea of like self-play converging to this like beautiful, elegant, like perfect strategy that is objectively correct no longer holds. >> Yeah. >> And you run into a similar problem with things like math. You know, you could say like, okay, we'll do self-play in mathematics where like one agent is trying to propose arbitrarily difficult problems and then the other agent is trying to solve them.

The problem is the proposer agent could for example propose >> Right. >> impossible problems. >> Yeah. >> Or they could propose problems that are like difficult but not but difficult in a way that's not interesting to humans. So, for example, it could be like 50-digit multiplication. Like, yeah, that's going to be difficult for an LLM to do if it doesn't have a calculator, but is it interesting?

Like, not really. >> Yeah, it's very hard to regularize, right? Like, even things where there's a short solution or sort of like a known solution, it's hard to come up with problems of the sort that um like would sort of push you in the right direction for training. >> Yeah. So, it's not to say that self-play can't work. I think I think self-play is going to be extremely important, and I think it will probably be the thing that propels us past superhuman performance to just like, you know, unimaginable levels of like of intelligence for for these models.

But, it's not as easy as it was done for for Go and chess and these two-player zero-sum games. And that And that's the main reason why it's unclear to me how long the Centaur period will be because I think it depends a lot on how long does it take to like really unlock self-play in a proper way. >> Super interesting. Well, coming back to um the Erdős problem for a minute, you said something um a little bit ago that was really interesting, which was um you know, the models uh have gotten so good that it's sort of hard for non-expert humans to sort of tell whether a solution is correct or not.

And so, I was curious like at a micro level what this specific result was like. You know, I imagine that you guys have some automated testing and maybe like a team of in-house mathematicians. But, uh you know, I've I'm just sort of curious like about the play-by-play of, you know, we think that we have a solution to this problem, and um what does it take to actually sort of, you know, believe that we've we've gotten all the way there?

>> Yeah, and basically what happened was we trained a new model. Um and it wasn't designed for math or anything. It wasn't um it's a general-purpose reasoning model. And um afterwards, I mean, we have people with, you know, math PhDs at OpenAI, some people that were formerly math professors that um are interested in math. And they just decided like kind of as a as a side thing, like, "Okay, well, let's let's see where the capabilities of this new model are.

Let's run it on a bunch of unsolved problems and just see like if any of them you know, if if it thinks that any of them are correct. And of course like one of the challenges is like actually verifying these things. So So basically the model was so they ran it through a bunch of problems um and this is one of the problems where it came back and said like actually um it's it's solved. Um and then there's like this very difficult process of like verifying that that's correct because you know, you need not just an expertise in math but like an expertise in several areas of math in order to determine that it's actually correct.

So this is like actually pretty interesting week-long period where like we they thought it was correct but like they weren't really sure. Um And so they had to like go to a bunch of external mathematicians and discuss like okay, well, we have this candidate proof. Can you like help us figure out if this is actually a valid proof? >> Yeah. >> Um And then like yeah, like once it became clear like okay, this is actually correct, there's a question of like what to do.

Um and you know, honestly like for me, you know, they they told me like yeah, you know, we think this model like prove this result and I'm like okay, what does that mean? Like I don't I don't know. I don't know if this is actually a big deal. Like you know, I heard about like, you know, dozens of Erdős problems being solved before. Like why is this one different? Um but I I think it what would really helped convince me and I think I think a lot of people was getting the um talking to the external mathematicians and and hearing from them that like actually this is a very significant result.

>> Yeah. >> Um and and and so yeah, that was basically the the trajectory of things. So that's also why we put a big emphasis on like getting the feedback from the mathematicians um these external mathematicians because it it's one thing for us to say that this is a significant result. I think it's another thing to say for external mathematicians to say that it's a significant result. >> Yeah. Yeah, I agree.

I mean, I think it's something that um you know, like I remember hearing about this result um as an undergrad and um it's it's certainly like in the flavor of things where mathematicians spend a lot of time, you know, sort of working on it. It's not like one of these areas that's been attention starved. Um and yet I think like this approach like it's probably received like very little attention. Um and you know, I think that's one of the things interesting about AI models is um you could potentially like pour a level of of sort of attention and effort into these problems that sort of like far exceeds, you know, sort of even human society's capacity to work on them.

Like so many areas in math might only have a couple of people working at the absolute forefront and so it's easy to imagine kind of like a nearer time field when um the vast majority of the effort, just measured by like hours of time spent on these problems, comes from AI systems as opposed from human mathematicians. And I've wondered if, you know, that alone should actually propel a lot of discovery. >> I think that could be the case.

And but one one thing I want to emphasize is that we did not put a lot of effort into solving this problem. Like it was not a lot of compute at the end of the day. Uh we later like showed this plots where, you know, the x-axis is the amount of test time compute and the y-axis is like the probability of solving the Erdos problem. And like to put that in perspective, to make that plot, we had to run the model on the problem 100 times for every single data point.

>> Yeah. >> And so like if we're doing that 100 runs for every single data point on that plot, it's like okay, you you you can kind of get a sense. It's not the most expensive thing. Now one of the really cool things from that plot was that we showed as you put more test time compute into uh solving the problem, the probability of solving it goes up pretty pretty significantly. And I mean, my takeaway from that is like there is uh probably a bunch of other problems that we have not solved that this model can solve.

And we just haven't put the amount of compute into it that's necessary. >> Yeah. Are there any areas you're a optimistic about sort of like, you know, given this result, you should all have a belief update that, you know, sort of like true sort of like annals of mathematics level um you sort of novel discoveries are possible. Um you sort of have this new instrument. Um you know, where do you want to point it? >> It's a good question.

I mean, the truth is for we at OpenAI like have a huge amount of leverage to make the models more effective. And we're in this weird spot where like we have we have access to these frontier models like a few months before anybody else does. And it's really tempting to then go through all of the unsolved mathematical problem mathematics problems and and try to run the model on them and try to see what can be proven and what can't be proven and to spend basically all of our time doing that because we can get some incredible results.

Um but I don't think that is actually the highest leverage for us. Like I think the highest leverage for us is making the models better, more powerful, uh safer, getting them out to the world as quickly as possible, and enabling mathematicians to use these models to solve all of these unsolved problems. So, for us I I it's really tempting to to try to solve all these problems and I actually try to discourage um researchers that I work with at OpenAI from from going down that route because I think that the most important thing is for us to focus on improving the models themselves.

And that's actually what I'm most bullish about. I mean, I I the fact that this was not a math-specific result, that this was not a math-specific model uh or or algorithm or anything, um it was just a general-purpose model. It's not the capabilities are not limited to mathematics. So, I think it's also um going to be very effective at doing things like machine learning research and making the models more effective and enabling us to basically do recursive self-improvement.

>> Yeah, I remember um I think you were one of the first people that I talked to that really believed in this idea of using the general purpose reasoning models for these types of things. Like so for contest math, for example, I think like a year ago there was this big debate of, you know, do you have to go through the sort of lean formalization path? You know, can you use a general purpose reasoning model? If I remember correctly, the model that you guys used for all of the contests, like uh IMO, IOI, ICPC, was effectively the same model, if I'm remembering correctly.

>> Basically the same model. I I think for ICPC and IOI it was like um a scaffold of a few models, um but the most important model was like the same model as was used in the IMO. >> And so um why why was that so important to you? Is this this feeling of like I want these improvements to be generally useful for things like recursive self-intelligence, uh self-improvement? Um is it that, you know, you sort of you cared about sort of transfer from these sort of contest domains to the model that, you know, hundreds of millions of people use every day?

Um you know, why I think you were sort of very early to sort of making that decision. >> For me it it it was always, you know, what what was extremely impactful was like the the really general purpose systems and um trying to like just break things down to the actual bare minimum that's needed, um and then and then figuring out what could be scaled. >> Mhm. >> And so, especially at a time when like the field is progressing so quickly, it's it's really easy to get nerd sniped into making these like custom-built models that are very, very good at, you know, solving the IOI or whatever, um but ultimately will not be useful 6 months down the road.

>> Yeah. >> And I think it's important for us to recognize that like there is a goal, and that goal is like AGI or superintelligence or whatever you want to call it, and we don't want to take detours from that route. And so like I think it's I think it's great to do things like the IMO as as a milestone along the way to AGI, but I don't think we want to get sidetracked by it too much. >> Right. It should be like a byproduct of ascending the curve as opposed to, you know, sort of putting resources just to sort of claim claim the post.

>> Yeah. And it is it is a constant tension because like there's such a temptation to to get these really impressive results to present it to the world and and get recognition and all this stuff and and we we have to like resist that temptation, I think. >> Yeah. So, if the path is sort of releasing the models to human mathematicians, and and so the idea being that, you know, they'd be the ones that sort of spend time you sort of playing with them and producing producing these results.

You know, one thing I wondered about is who's going to be really good at that and if it's sort of different in any ways than, you know, sort of the current set of human, if you want, like manual mathematicians. And so like the analogy here is, you know, Richard Hamming had this interesting thing where he I think he was like an early proponent of simulations in physics. And you know, he has this thing where he's talking to the president of Bell Labs and he's like, you know, today like 90% of the experiments are being done, you know, sort of in the physical world and in the lab and it's like 10% on the computer.

Like that's going to switch immediately. It's going to be 90% on the computer and maybe 10% in physical reality. Then he goes on to kind of describe how like the people who are really good in the sort of, you know, simulation regime are not necessarily the same people who are really good in kind of the experimental regime. Are the people who are going to be really good at proving new math with these models different than the people who are really good at doing math manually, so to speak?

Like should we expect, you know, lots of impressive new results to come from like Fields Medalists or should we predict them to instead come from people who are really really good at using these models, which are maybe a different set of people? >> I think it's possible that it is a different set of people. I think It's a feel like this is always the case with new technology where it requires adapting and approaching things very differently.

And um and it's also a question of if the models are extremely good in some respects and below human performance in in other respects then you want like the people that are going to be most successful are the people who are the complements to the models. >> Yeah. >> So that profile could be very different from what makes a great mathematician today. It's a little hard to say because I think it depends on I mean it depends on like where the AI models are are going to be strong and where they're weak and what makes a great mathematician today um but I I think the people that will be most successful will be the people that complement the AI models really well.

>> And do you see any early hints of who those like what that might be? Like what are the ways in which you know like one thing I can imagine for instance is like um math has become very specialized. You know a lot of these domains people I remember when I was an undergrad um you know some I asked about this result and I was like is this this really important result about elliptic curves and I said you know if I really spend time trying to understand it could I?

And um you know this sort of uh very well known math professor is like well no because um you know I spent like six months trying to do it as a fifth year grad student and like I don't feel like I understood it. It's just sort of so specialized. One thing I can imagine is that like maybe the generalists have like another you know sort of day in the sun where connecting a lot of these different areas is valuable. Another thing I can imagine is it's like totally different than that.

It's it's sort of like more about um sort of point having intuition for where the models are good and pointing on the right problems. I'm just curious if you have like any early hints of like you know who who the sort of like new mathematicians are. Like where are those sort of complimentary skills? >> I my personal experience is you know in using these models to do AI research and in my experience they tend to not have very good research taste.

And so for me this has actually been been amazing. It's been amazing period because like you know, I was like okay at the you know, the kind of like the IC kind of work but it was I was never like exceptional in this respect and now you know, I I feel like I'm such a great compliment to these models. I'm able to be so much more productive. And um I wouldn't be surprised if it's something similar in mathematics where the ability to ask the right questions and know what to investigate becomes um the most valuable thing because the models probably not very good at that right now.

And I yes, I I I could see that being people that are extremely good at that might be a good compliment to the to the AI models in math. >> So maybe somebody who sort of has good research taste and has is like good at program construction. Like I'm going to work on Langlands or I'm going to you know, sort of like be Grothendieck style. Like you have kind of like this larger program that you know, sort of is the right path to go down.

>> Mhm. >> Um interesting. You know, one thing I was curious about is this kind of construction um you know, sort of versus versus sort of like a different kind of argument for the Erdős problem specifically. So in your companion piece the Field Medalist Timothy Gowers you know, sort of has the story where he says, you know, at first he thought that what OpenAI had done was sort of prove a much tighter upper bound.

And then he said he sort of like really lost a lot of sleep after that. Like if that's true then like we're all completely cooked you know, because that that that sort of like I can't even imagine how progress would be possible there. And then you know, he says that when he sort of learned that instead it was that it was a sort of much it was um you know, sort of the construction of a of a counterexample that sort of disprove that you know, Erdős's upper bound.

You know, that felt to him like more plausible. Like that was the kind of thing that you could at least imagine like an AI model coming up with. And I guess I'm curious like if you would draw that distinction like if there's any anything that like you take from that, like again, I'm the spirit of this is the flavor of like places where the models are going to be better than humans. Um, you know, maybe I'm curious if you sort of view that as like a real philosophical distinction.

>> I mean, I I think I would expect separate from the Erdős problem, like if I were to if I were to predict what the progress of AI in mathematics would look like, I think it would be, okay, you get something like IMO gold, and then you start proving some like unsolved problems that aren't that, you know, significant or that researchers haven't paid much attention to. And then you get to something like you prove a significant res- like significant result, but in a way that's like, you know, not, um, you know, causing a field of mathematicians to like lose sleep saying like this is like beyond human ability in like every respect.

And then the progress will continue, and eventually you'll get to that point where, you know, yeah, the mathematicians do lose sleep. But I I think it's just natural to expect that this is the progression of things, and we're just seeing another point uh along that progression. And that's kind of what I would expect from this level of capability, and a year from now, I think the capability would be much higher, and I wouldn't be surprised if it's like doing things that are unimaginable to mathematicians today.

>> Mhm. >> Yeah, and it sort of already feels like, um, I think a lot of the mathematician commentary too is like, um, this is such like a sort of deep way of improving sort of the original construction. Like there's sort of like a very There's the original Erdős sort of, you know, like relatively elementary construction. To improve it requires all of these different ideas, um, from algebraic number theory.

And again, like the number of people who sort of have all those things in their head, and who could sort of spend serious time in that on the problem, um, um, is very small. And so, you know, maybe like another way that I've wondered about AI systems having a major advantage, um, over humans is this idea of, um, you know, the existing research literature is vast. Um, most professional mathematicians only know like a small fraction of it.

Um, there are probably lots of problems where if you, you know, sort of put five papers in front of them and said like, read these five papers, you know, you can't leave this room, you know, and sort of until you've proved some result using them, um, you could make progress, but you just can't do that for like many different areas. And so, I I sort of wondered about this idea of like this kind of one-time dividend that we might be very close to achieving where it's like just sort of putting together ideas from very different research areas to produce something new.

You maybe that's one of the things that's more near field. Um, and I'd be curious what you think about that. >> I do think that's going to be short-term where a lot of these breakthroughs come from. I mean, the fact that the model's like trained on so much different on so so many different data sources, so many different papers, it it's able to connect ideas that, I mean, it's it's impossible to keep up with like every single field in an in-depth way.

And so, this is already something that it's pretty clear the models have been superhuman for a while. >> Mhm. >> And, um, and so, I do think that well, this is where it begins. And I think it just progresses from there. >> Maybe you're kind of bumping one level up. Um, you know, what are the reasons why you think math is a particularly good domain? Um, you know, sort of I'd put it inside of AI for science sort of writ large where sort of making new scientific discovery, um, with AI models.

And, um, you know, I'd I'd be curious for your take on if math is a particularly, um, you know, sort of fertile domain for doing this. And then sort of what other domains you're optimistic about. Is it inside of physics or particular sort of uh sort of branch of physics? Maybe the models are um, you know, sort of very good at at sort of particle physics problems or something. Curious for sort of where you're most optimistic.

>> I think math is special because it is purely bottlenecked by reasoning. >> Mhm. >> Like in in physics, my impression, I don't know I'm not a physicist, so I don't know for sure, but talking to physicists, it sounds like a lot of the field is bottlenecked by experimental results at this point. And so, you know, you could you could make all sorts of crazy theories, but ultimately you need to put a bunch of money into the actual physical experiments to be able to validate them or get more data to inform like what the next steps are.

And you don't really have that problem with mathematics. You can just sit in a room and think for a really hard for a long time and come up with something like amazing and that's how you make progress. And um And so, I think we'll see the biggest benefits in the short term from these kinds of domains where you're bottlenecked not by actual physical experiments or experimental data, but from just pure reasoning ability.

>> Mhm. >> So, wet labs, for example, are going to be a bottleneck. I think in physics, collecting the experimental data could be a bottleneck. Uh mathematics won't have that and the question is like what other fields also don't have that bottleneck. I actually think in some ways like things like AI research, you know, it's debatable. Um on the one hand, you do like a lot of research is bottlenecked by just having a ton of GPUs and running large-scale experiments.

On the other hand, there is a lot you can do with small-scale experiments and small-scale resources. And so, I am overall like pretty bullish on the ability of AI models to be able to like make progress on actual AI research. >> Do you think um so, you know, we're I think we're about like 2 months out from the next IMO. Um Do you sort of view this as the year that it's you know, that sort of saturates, so to speak?

I think last year um the scores were already like 34 out of 42, like clearly at the sort of gold medal level. Um you know, I could imagine sort of soon getting to the point where the IMO becomes less interesting as like a problem set to solve because you just expect to kind of like max out um on every problem. And, you know, I guess I'm curious what you view as like the interesting frontier in contest math. Um, you know, sort of if there is still one left or if you sort of think that to find interesting problems now you kind of actually have to go to math research.

>> I remember when we got the IMO gold last year, um, there was somebody from, you know, another lab that texted me and was like, "So, do you think you'll get a perfect score next year?" And I told him like I'll be disappointed if we don't have a model out by next year that anybody can use to get a perfect score on the IMO. I I think I would be surprised if we don't get a perfect score with like our latest internal models.

I would also be I would be disappointed if if it's possible that I'd be disappointed if we don't get a perfect score with like the released models at this point. Um, so I don't think there's anything I I think either now or pretty soon competition math competition coding is not going to be interesting anymore. And the frontier really is actual unsolved problems. Like actually doing real research in the real world. >> Um, yeah, maybe the right, you know, sort of evaluation is just like can you literally just upload the PDF of problems to the publicly available, you know, GPT-5.

5 and, you know, sort of max out. Maybe maybe that's sort of that's the interesting thing um this year. Um, >> Yeah. I should also say like I don't know we we never really mention it, but we did like our models have been able to solve problem six from the IMO last year for a while now. Uh, we decided not to like advertise it because we just felt like the field has already progressed beyond competition math, but like we've been able to get perfect scores on P6 for several months.

>> I was going to ask about that because um you know, so historically problem six is like one of the harder problems and I think even last year amongst sort of like human participants maybe only six people um got full marks um for it or something. Um so I was curious if it was if there was anything specific about problem six um you know, that sort of um you know, made it less sort of a like sort of it's harder to solve for an AI system specifically or is it just that it's uniformly harder for humans and AI systems together and we should just expect problem six to be hard cuz it's sort of I remember correctly it was this kind of combinatorial tiling problem.

I've heard the argument in the past that like maybe geometry sort of combinatorial geometry problems are sort of particularly hard because humans have so much visual intuition, but yeah, I'm curious if you if you think there's anything special about problem six. >> Uh it's a good question. I mean I I I think it's kind of interesting how much correlation there is between like difficult for humans and difficult for for AIs.

So I think that is a major factor that it's just like look, it was a really hard problem for humans and so in that respect it's actually I guess not too surprising that it's like a very difficult problem for AI models. I think it's also the nature of the problem that I do think the models lag a bit when it comes to like geometry um and just like geometric understanding. Uh but you know, the models just get better across the board and they get better at these things and so like it wasn't surprising that you know, eventually the models were just like suddenly able to solve it.

>> Do you feel like there's sort of like a different in kind um chasm to cross between contest math and research math? Like did you feel it sort of you know, in theory like the contest problems have to be somewhat self-contained, you know, there it has to be possible for a human to sort of write a solution to them like in an hour or something like that. Um but maybe kind of and you know, also I think they tend to sort of favor um the sort of like certain known techniques or tricks.

Like if you get really really good with the pigeonhole principle, you know, there's sort of always a couple of contest problems that you can solve. Um but I was curious if you felt like any difference in kind on the AI research side between sort of pointing the model at contest math versus pointing it at open research problems. >> I I think there is a difference in horizon I mean there's there's a difference in horizon of the problem.

So, the IMO you have an hour and a half to solve a problem. And one of the things that my colleague Alex Way pointed out, which I didn't realize at the time but I think is really true in retrospect, is that if you look at the progression of AI models when it comes to math, it it kind of follows this trajectory of the model getting better at uh at like doing problems that would take a human mathematician longer and longer.

So, if you think in like 2023 we were doing GSM8K. And GSM8K takes a human mathematician like you know, 5 seconds. >> Mhm. >> And then you in 2024 we're doing math. And um you know, that's those problems they're like high school level um maybe like early college level and that takes a human mathematician like maybe a minute. And then you get to the AIME that could take a human mathematician maybe 10 minutes.

And then you get to the IMO and it takes a human mathematician about an 100 minutes, about an hour and a half. And you can just like basically every year you progress like an order of magnitude in this respect. Like the horizon how long it would take a human mathematician to do it. And like there's reasons to think maybe that trajectory doesn't continue but it's like been continuing pretty regularly so far. And so the difference between competition math and research math is that research math is over a much longer horizon.

>> Mhm. >> But it's also and I guess not surprising that yeah, the models eventually are able to do things that would take a human mathematician 15 hours or or a week and it's just the models have to become more capable. I there is also a difference in that in many ways the IMO was like adversarially hard for AI models because, you know, we discussed there's different there's different strengths and weaknesses of these AI models.

Like they're really good at knowing a bunch of different areas of mathematics and being able being able to combine results from different fields. And then there's the like being able to sit and work on a difficult problem for a long period of time and and just reason through it. Where they're they're good, they're getting better, uh but they I don't wouldn't say that they're superhuman in this respect yet. And the IMO really emphasizes the second part.

It's really about It's not about knowing a bunch of different areas of mathematics, it's about sitting down and reasoning through a very difficult problem for a long time. And and and that in that respect it's actually like pretty interesting that the model was getting a gold medal at the IMO. And I for that reason I think it's not surprising that less than a year later we're getting these like you know, this the starts of proving results that mathematicians cannot prove because we already know, okay, it's a I it was at IMO gold level in the ability to like reason through hard problems a year ago and it's or it was probably already better at knowing a bunch of different areas of mathematics and combining results between them.

So in that respect like yeah, being able to do research math it probably isn't too surprising that it happened so quickly. >> Yeah, it's like you're sort of not able to use things the models are best where they have the most comparative advantage relative to humans in the contest setting. And well, okay, maybe speculating a little bit, what are human mathematicians going to be doing in five years? Like how does the job change?

How do you imagine people you know, sort of will be spending their time? Is it things like, you know, sort of coming up with new research programs and then just letting the model crank in the background to see if there's a there there over the next six hours? Um you know, is it is it something different? Yeah, I'm curious what you would speculate, you know, sort of human mathematicians are doing in whatever time horizon you think is predictable.

>> Yeah, five five years is so far down the road at this point that it's like I don't know what the world looks like in five years. Uh I I can't speculate what anything is like in five years. Just because like progress is so fast. >> Yeah. >> I think a year or two from now, um I think a lot of mathematics is it it kind of looks like the way coding is today, where it's it's a lot of working with the models.

Um the models are doing a lot of the work and you're kind of like steering them um and pointing them in the right direction and challenging them and like working with them to to do these kinds of results. I In many ways, I'm kind of surprised that it didn't happen in mathematics first. Um that's what I would have expected actually. Um but I don't think it's far behind where where coding is today. >> Um what happens next?

You know, do you sort of think um that maybe there's some limited release of the model to sort of human mathematicians to sort of make progress in these areas. I'm curious how you think about It makes total sense to me that OpenAI's focus would be on making the models better. Um but you know, sort of if you were looking just at math as a domain um and sort of asking, you know, as beyond OpenAI, just sort of as a research community, um you know, how are we going to make fast progress here over the next year?

What what sort of um what's what's sort of about to happen? What do you think are the next steps? >> We after the IMO result, we actually did discuss like should we make this model available um in advance to mathematicians. And ultimately, it was just a matter of like the bandwidth and and the difficulty involved in doing that kind of like custom deployment. And also, I mean, progress is so fast. We we want to get the model out to everybody.

>> Yeah. >> Uh as quickly as possible. And so, fast-tracking it for a particular group, it introduces a lot of complexity uh that I just don't think is really worth the overhead. I think it's better for us to just focus all that effort into getting the model out for everybody as fast as possible. So, um I don't think I mean, maybe maybe we'll do some like special cases, but I think for the most part we're just really focused on getting the model out to everybody as quickly as possible.

Um and allowing everybody to use this thing to to prove all sorts of crazy results. >> Awesome. Um well, you know, we're almost at time. Is there anything that we didn't talk about that you wanted to talk about? You know, sort of anything that you think it's important to to sort of understand um or that you just want to call attention to? >> You know, one of the things I was I've been thinking about ever since this result is like I it the proof is at a level that like I'm not I mean, maybe if I really spent a lot of time trying to understand it, I would be able to understand it, but I haven't really spent that that much time.

Um and I'm sure there's like a lot of people that can't understand the the proofs um for these kinds of results, let alone proving it themselves. And I think we're as the models like become superhuman in many respects, we're going to have this problem where like the proofs themselves are going to be too difficult for human mathematicians to verify. >> Mhm. >> And I think this is a classic problem that's basically scalable oversight.

How do you how do you get the models to prove that the proof is correct to to human mathematicians when the proof itself is beyond what humans can really comprehend. >> Mhm. >> I think I'm excited for mathematics to be like on the vanguard like be the vanguard for for this kind of question of like how do you do scalable oversight? And I think it's going to be uh it's going to be like an interesting problem and um one that I hope we can we can figure out because I think it's going to be important for a lot of other domains as well.

>> Yeah. No, thanks for doing this. This great. >> It's been great.