🎙️AI 访谈库
Terence Tao 与 Mark Chen 炉边对谈(UCLA IPAM)
Mark Chen · OpenAI

Terence Tao 与 Mark Chen 炉边对谈(UCLA IPAM)

Terence Tao and Mark Chen - Fireside Chat with James Donovan - IPAM at UCLA

2026-03-09 · IPAM (UCLA) Fireside Chat · 1h01m · 约 65 分钟读完 · 原文
陶哲轩与 Mark Chen 炉边对谈(2026-03-04 录制):AI 如何改变数学研究——文献检索、代码生成、证明尝试与协作工作流,前沿研究仍依赖人类判断与目标设定。

Well, thank you very much. Um, before we get going, just a a massive thank you obviously to the Institute for hosting us today. Beautiful space. Um, and also for all of you for turning up. Uh, I know you're not here to hear me talk, so I won't talk for too much longer. Uh, so really just to say also a massive thank you to you both for coming. Uh, it's rare that you get two such great minds in the same same place. Uh, so we really appreciate the time uh, going into this. It's our third time. Yeah, and and a little pattern is starting to build up. Uh, and maybe actually that's a good place to

to start the conversation. You guys had a conversation be almost a a year ago or so uh, to the day. And at the time, uh, Terry, I think your prognosis for where uh, GPT was for mathematics was something like uh, a very ineffective grad students, which which remained with me because I I'd heard that feedback myself as a human being. So, it was a clear benchmark. Yeah, okay. Yes. Um, why don't we start with uh, how you think things have changed since then and then Mark give your your side of the story. Okay, yeah. So, a lot has happened in the last year. Okay, that's a not just in uh, in AI, but um,

yeah, so uh, yeah, these tools have definitely become a lot more powerful. Um, I think uh, there are now capabilities that basically uh, uh, are now normalized and like we just use them all the time. So, so so deep research tools. So, so so literature search has become really really good. Uh, it it has surpassed traditional um, uh, searches. Um, code generation of course is is a big thing, but so uh, um, as a pure mathematician, I I'm not as as heavy a user of code, but it has changed the way I um, I approach a math problem. I I will I will I will plot something. I will just if if there's an inquiry

I think is true, I will I will ask um, an AI to try to prove or disprove it. Um, it's yeah, so I I I I already use it. Yeah, if there's a lemma that I don't um, I think I know how to prove, but I just I just can't be bothered doing the pen and paper calculations. I will just outsource it. Um I've not yet found it to be useful at at um a sort of at the um deepest level of you know, when I'm trying to solve a problem and and um at um with pen and paper or with a colleague. Um I can't sort of interact with it on a conversational level

[snorts]

quite um at the level I I need yet. Um but maybe uh maybe in the future. Um but I think also socially, I think we're beginning to uh the uh mathematician community as a whole is beginning to understand that this that this is these tools are here to stay and we have to actually start adapting how we um do our research. Um so certain things that were very tedious um and maybe we would uh force our graduate students to do um you know, we can offload to to AI and this opens up uh lots of new ways to do mathematics. Uh lots of research projects um especially at scale that we just we could not

dream of doing. Um so while I think we can use AI to assist our current workflows, it it's a little bit awkward still to do that, but I think uh the um maybe much more miles are in creating new workflows which are optimized for AI. Um it's like when we um invented the automobile, we started changing the way we built cities. Um to um you know, and of course you could say that maybe not all the the changes were good, but um but um yeah, we're sort of in this intermediate state where it's somehow um our roads are still built for for people and horses and and and we now have automobiles. Yeah. So would

it be fair to say that we've got to the point where occasionally helpful collaborating, uh but maybe more interesting than all the bigger open spaces how you change the way you do maths with these tools coming. Um Mark, would that be true to what you're seeing and what you're building for? Yeah, honestly I don't blame Terry for saying it's an ineffective grad student a year ago. I think that's largely the state that we were in back then. And you know, I I really do think of the backdrop of AI progress as hill climbing this what we call meter plot internally of uh the models doing autonomous work for longer and longer periods of time. And

I think last year we were in the category of minutes. And you saw that, right? It would just the model would hallucinate, it would kind of fall over when you gave it significant chunks of work. But I do think the last year has been a transition for a lot of us in that we've seen the mistakes go down and therefore you can trust the model to do longer periods of work in general. And that's you know really kind of allowed us to do away with a lot of the scaffolding that we might have needed to use before and really start to attack you know bigger problems and and truly orchestrate with the model. And yeah,

I just think of a year ago we were in the world where we were kind of roughly achieving a bronze medal at the IMO. I think this summer you know across all kind of high school mathematics and programming competitions we are achieving gold medal performance and I think we've just kind of run out of these human written benchmarks and that's why you do see people evolving to this sphere of doing mathematical research. And fundamentally that's always been the goal. We don't find any pride at OpenAI just kind of solving you know IMO problems or anything like that. The the real ambition is to push the frontier of science and finally the task horizon has caught

up to a point where we are actually able to go do that work. And again it's not there yet. I I think the trend in the structure is strong but yeah, I I do think you know it's true that a lot of people are finding utility in it today. I mean I'd I'd like to come to maybe first proof and that transition as we go into more frontier mathematics. But maybe to stay with the capabilities right now I think often the IMO problems are seen as a way of getting a litmus test for where the models are. Um and that that maybe is a representative set in that some of those problems are maybe not

as complex as other parts of those problems and we're we're better designed that way. And that you might say that the success of the models has been in doing lots of the easier ones quickly versus necessarily moving towards the the kind of acorn level problems. Is that a fair depiction in your mind, Terry, where they are today? Yeah, so I I've I've been heavily involved in tracking the progress on the Erdos problems in particular. Um, so I mean, yeah, and it it is basically largely uh what what would you say? These these problems range widely in difficulty. There there are some that we desperately want to have solved and and uh they've been worked out

for decades. You know, I have papers you know, making tiny progress on on on some of these these problems. Um, um and today AI has not really helped with with the ones that have been um that we've already poured a lot of attention to. Um, but there was this very long tail of problems. It was a Erdos posed a thousand problems. Um, they weren't all winners. Um, but he understood that that um, you know, the the important thing was to stimulate and and and and interest and um, you know, and um, you know, and he kind of knew that the problems that that that that were um going to be important would have taken a

life of their own. Um, but there was this long tail of small problems um where there's maybe almost no follow-up literature. Um, and that's where the AI tools have have really made a lot of spectacular progress. You have maybe 20 30 of of these um problems have have been solved with fairly minimal uh human supervision. Um, um, but these AI tools are I mean, and we were able to verify them, too. Uh often with with some other um AI tools in form of verification. Um, and we've kind of worked out a a a kind of a workflow for for for doing this without being overwhelmed by by by AI slop, incorrect solutions, or whatever. Um,

so yeah, it it it is um it's a new capability that we hadn't had before. Um, because we can now attack attention bottleneck problems. Um, and so um I think what this suggests to me is that we need to start creating more and more broad challenge sets of problems for AI tools also but the general public. So, actually in the same period many of the of these earliest problems were also solved by amateur mathematicians, sometimes with AI tools sometimes without. The same mechanisms that that the same kind of workflows that enable AI to be successful also actually enable um um amateur mathematicians to be to be successful as well. Um So, I foresee a change

in our culture where instead of only working [snorts] on on on a small number of of really hard problems and not sharing a longer list of of other things we'd care about we'd we'd all start as mathematicians we'd all start releasing um problems of as of things that we want to to get answers to. Um and 100 problems and maybe this AI can solve 10% of them and and maybe this this other high school student can solve another 5%. But it um we can get a much much much more uh sort of community driven way of doing mathematics. Um so, I I think this is what the earliest problems are sort of an early harbinger

of. Yeah, and it's it's maybe interesting to contextualize that other domains of science uh at least in my my own world of biology in which the number of people collaborating in any given paper is just exponentially risen over time and that that seems to be the trajectory that science is much more of a team sport. Maths is maybe in and to some degree theoretical physics the outlier in that domain. Um when you're thinking about this Mark, is it is it always just a question of uh how smart can we make the models and the ever more difficult questions they can answer or is it also a question of how do you empower humans to work

collaboratively on on these problems? Yeah, I mean right now we really do see heavy engagement with the community. That is a necessary part of driving progress in all of these scientific fields. And um Kevin here he runs our Open AI for science program and part of this is that kind of like you said, you know, these experiments like first proof or the Erdős problems, um it really is an engagement with the community on figuring out what problems are actually important to tackle. Um we've done this kind of exercise in physics as well, right? Brought in kind of expert physicists to kind of lay out a program of here are the really important things that feel

like they're amenable to AI. And that also helps us shape the AI in turn. Um and it it allows us to kind of find the deficiencies, right? We can look at where our models fall over and and really shore those things up. Um what we hope to build is this platform where scientists around the world can just accelerate themselves. And we want to empower that community mathematician. Um we see people like that today uh empowered, you know, you have these 20-year-old, 21-year-old uh kind of kids using the models to solve some of these Erdős problems. It um may not be, you know, these sophisticated and and very significant leaps, but they're able to do a

lot of self-directed work. Um I I kind of had this thought when you were asking the question of I know that you've organized a lot of big community initiatives in math before. Um

[clears throat]

I don't know how you think AI is changing that world or, you know, does it enter that world in a significant way? I think it I think it combines very well, actually. Um so, I think what AI will enable is finally a way to use use division of labor, which is something that um like all industry, you know, since the industrial revolution, like every industry has managed to become more efficient through a division of labor, except mathematics. Um so,

you know, so to to do mathematics traditionally, um there's there's several different tasks to, you know, yeah, there's there's problem look generation, there's strategy um generation, and then strategy selection among all the strategy you generated, and then execution of the strategy, verification of the strategy, communication of the results. Um and um so, um you basically need we've trained our mathematicians to sort of be somewhat good at each of these tasks.

So we specialize in field, but but we have to have some idea of where problems come from, what are good problems, what are good strategies. You have we have to have some technical skill. We have to write and we have to explain.

are better at some of these than others. So we have been able to benefit by from collaboration because of this. But but we can't really specialize um the same way that that in the sciences, you know, you can have you can have technical staff and you can have people who are project managers and things like that. So but now with AI and other modern collaboration tools and full verification, it has become possible to have a to run math projects where individual participants specialize in just one of these areas and and maybe maybe there's some gaps among your collaboration. No one knows how to to do the technical thing, but but AI can plug in plug

some of the gaps in a in a collaboration. So but you can't but you still need the humans because the the AI performance is very very jagged. So maybe some of these inputs can now be be be be automated, but if you automate that too much, you know, for example, if if you can automate strategy generation, okay, but you can't automate verification, then then you just get this whole hundreds and hundreds of possible strategies. But anyway, AI generated that you can't do that. But if if the verification also keeps pace, then suddenly you you you you have a a new style of doing mathematics that is like extremely effective. Yeah, just one quick comment on

that, too. I I actually absolutely agree that AI capabilities are super jagged today, and so you see this really fruitful collaboration with humans. It's also interesting to explore the flip side of that, which is that some of these AI systems are more human-like than you imagine, and you have to like pump a lot of RL in the right way to like not have the problem have the models give up in the same way a human would. You know, if you give a too hard problem, often times the model can just, you know, it it would like run a couple of test or, you know, prompts in in its own train of thought and be like,

"Ah, you know, this problem's too hard. I'm I don't think I can actually do it. Let me pretend to the user like I tried really hard." So, we've seen with the Erdos problems that you you you get an AI to to try to solve an Erdos problem. It will First thing it'll do is it'll go to the Erdos problem website, look it up, say, "It's an open problem. It's too hard. I'm not going to try." So, you have to say, "Do not use the internet. Try to solve the problem yourself." It's actually pretty easy. I swear.

It's good to know that it's, you know, frontier research is actually just about coaxing the models into behaving the way you want to. I That vision right now is probably quite a compelling vision for this room and and beyond where we're sort of saying the technology fundamentally empowers more people to collaborate on these problems. But is this just a stepping stone to a world terrain which you're only collaborating with many AI agents and slowly but surely they come to dominate the the space? Um I think um yes and no. I mean, I I think the type of of math that we do today uh might slowly kind of move in that direction. Um but there

could be very new types of doing math that we can't even envisage right now, um which uh would I I I think I so so math is infinite and you know, the the difficult the the problems in in math the difficulty levels are unbounded. Now, there are even problems in math that are unsolvable. We know they're unsolvable. Um Uh well, okay, with an asterisk, but I don't want to talk about it. Okay. Um You know, so like there's certain things that even the most powerful AI we can't There's even right now there's there's certain cryptographic challenges that that, yeah, we AI cannot, you know, mine all the Bitcoin right now and whatever. So, I think

there will always be a frontier um and um and I I I'm pretty I'm pretty sure that just because how complementary human and and AI as as at least current generation LLMs are with with human skills that the best combination is always going to be a convex combination of humans, but the nature of the combination may change over time. So let's assume even just philosophically there is this frontier beyond which at least the current paradigm of AI wouldn't be able to cross and some central like human AI collaboration is is needed. Getting to that frontier in your mind Mark, is that a question of um much smarter RL training or is it actually just a

question of raw computation? If I could give you an infinite amount of compute today, would you be able to accelerate your way to to that frontier? Yeah, I mean I think when I think about the open AI research program overall, it is really fundamentally about how do we improve the algorithm such that they scale to the level of compute we have you next year and the year after. So it's grounded in the reality of what compute we actually have and I think all the algorithms we know they are simple and they scale, but they they take a lot of engineering and you know, fine-tuning to make sure that they truly scale to the next order

of magnitude and the order of magnitude beyond. One really great thing is this is a very multi-dimensional problem today. There are many axes by which we can scale model intelligence. We can you know, scale up the model and you know, build these bigger brains with just more core knowledge and this kind of captures the intuition that just like kind of maybe the more math you know broadly like just internalize deeply like it's easier to make these connections and these jumps. There's also a reasoning axis which we scale and this is the ability to take all of that base knowledge and chain it together to create new insights. And we have a couple people in this

room kind of working on this thing we talked about in the GPT-5 live stream which is kind of connecting this a little bit and having the models just generate new knowledge for themselves and really kind of amplify its knowledge in certain domains. So I think [snorts] there are a lot of different axes that are going to play into bringing the models to the next frontier. Um yeah, but overall like all of these things are grounded in in this meter pile. Like we are aggressively uh hill climbing towards more and more autonomous longer horizon tasks and we see that trend continuing. The um the term hill climbing, if I've learned anything working with the research team

is that we always must find and define a hill to climb and perhaps that's where these two worlds come together as defining the right hill. So, maybe we focus on that uh for the next segment. First proof would seem like an obvious example of where we've tried to codefine a hill to climb. In your mind, Terry, is that um representative of what you're thinking could be the new emergent mass to come or is that the final form of this more classical mass that AI has been working on? There'll be a a spectrum. Yeah, so first proof is is is very very interesting experiment um and um the the proofs that the the various um people

with AI tools generated uh were were quite good. Um what we saw actually that there was a definite verification bottleneck. Um so, we had a lot of proofs generated. Some were terrible, some some were quite good, some were similar to uh things in the literature, some were similar to the proofs that the um the authors themselves had. There was a couple who were which were actually different from the the uh the official proofs and so that was interesting. Um but uh um yeah, but to evaluate carefully exactly how novel and how um um interesting each each proof was, we actually don't have a um uh um a way of doing that effectively. So, I think

that the first proof team are going to create a more structured competition later where they will have some some mechanism for verification. Um so, we um in order to take full advantage of the new capabilities AI's have, uh we do need to um create challenges that are easily verifiable. So, somehow the level of automation and AI power that you can you can probably use before it becomes slop is roughly proportional to to how stringent your verification is. Um So, yeah, so I think initially you're going to see a lot of um um progress in in areas either which are sort of elementary enough that they they can be relatively easy to formalize. So, combinatorics I

think you're going to see So, the earliest problems definitely fall in this category. Um there's some numerical type challenges where you you want to to find a configuration mathematical object that that obeys certain properties. Once you have the object, verifying them are very easy. We saw some examples actually here in in in physics there's some similar type problems. I think there we will see a lot of progress. But, there are other parts of mathematics where the object is is not to find an object that obeys a certain property, but to find a good um overarching theory to explain something or a good definition or and those we have a lot harder time to verify like

you know, if you want to propose a new way to a new conjecture or a new strategy to find an unsolved problem to you know, maybe AI can generate a hundred of these possible strategies, but only a human expert can verify or can give an informed opinion. And so, that will be a bottleneck. So, even if AI drives the cost of sort of creating solutions down to zero, there there are still other huge bottlenecks which we didn't which were not front and center in our minds, but yeah. Um I think we also need to become very much better at stating goals precisely. Um so, AI is almost too good at fulfilling a goal to the

letter. So, you ask I I want to to solve this this problem. I want to I want to prove all this theorem. Um and maybe AI of the future just runs for now future proof. Um, but actually what you wanted was you wanted people to work hard to fail, to find examples, to connect to a literature, um, and to communicate all the partial results. Um, and um, and that was actually the value of of of solving a particular problem. Um, and there's a danger that if you specify your your your your your goal to an AI too narrowly, you miss out on on most of the benefit. Um, so we'll have to be more careful

about about goal specification. Yeah, just two really quick things to add on. Um, I I do kind of think of this offline version of first proof. And we are actually discussing this a little bit, too. Where you can imagine that you just train a model with knowledge up to some specific, you know, very very detailed, like like this day, this time. And you can imagine like what a first proof would be at that point in time. And now you have the benefit of hindsight. You're you know kind of what the techniques you're after might be, what creativity in the model might look like. And I think those are very interesting thought experiments to run. Um,

I I think you know, there's a thought experiment of what day you would choose as a cut off um, to get maximum signal on that experiment. Um, but yeah, I I also do kind of think about um, yeah, just, you know, the process of mathematics isn't just answering or proving a theorem. It's all this partial progress that you assimilate somewhere. And um, I I do kind of think about, you know, we have AI systems at at OpenAI where they're kind of just like central repositories for information. You could imagine that kind of serving a function in mathematics as well. Like you have just this kind of you know, global, you know, library in some sense.

As um, I I know I think Daniel Lit published something online a while ago. Kind of it it just is this agent that mathematicians can interface with. And it kind of like fills out this convex hole of, you know, math mathematical results. And um, you know, you can always kind of use it as the source of truth for like what people are exploring. And it'll it'll kind of connect a lot of the dots for you. And um, you know, kind of yeah, just be be this place which stores what we know. Yeah. Right. Well, it may sometimes be useful to turn that that that off, you know, so [snorts] I mean I've worked sometimes on

a problem where I know too many techniques, okay? And and I there's a powerful thing I know will solve the problem, but it's it's it requires a lot of technical skill to use and so I I I I do it and I solve my problem and then I publish or something and then and then someone points out actually if you had used this much simpler tool, you have a much simpler proof. Yes. I I I do worry a little bit that that sometimes having access to every single technique known in the literature is not necessarily the the best way forward. Um but but having a diverse array like multiple AI tools and and there I

think there will still be people who who will take pleasure and pride in sort of doing things old school and and finding sort of yeah, more more human ways to solve problems too. So Well, I wonder if that that pattern even plays to the case study you gave Mark where it's true if if we could go back in time just before a particular paradigm shift in whatever domain of science and then see whether or not the model would predict it. That could be one verification tool. But I guess the kind of Kuhnian vision of paradigm shifts could also mean that in fact there is a future paradigm shift to come that would invalidate the prior

paradigm shift. So you don't actually want the model to guess the previous one because it might take it off a pathway that doesn't get to the next one as you go. Um and it throws into relief I think this question of of verification validation. It it's both a philosophical and a practical question. Of all the domains you could argue that the maths and disclaimer here I did work a little bit on Lean so big shout out to that team has the the actual capacity to do automated verification in a way that very few other domains do. Not perfectly and and not without its own drawbacks. Is your instinct that that structure where there'll need to

be a sort of separate validation tool will need to come into existence for all the other domains of knowledge that that we want to work on that it will mirror what's happened in maths or it will need some other type of of paradigm. I I as I said again I I definitely believe that the there's an upper bound on how much AI you can inject into a workflow before before it it it it it becomes a net loss that it is causing more errors and and and problems than it is solving and one of the biggest upper bounds is the ability to verify. So yeah so in in math I think we have the best

shot at getting really high levels of automation being able to effectively use excuse me high levels of automation in a way that you couldn't do in less trustable and less verifiable domains because we have a a high verification bar at least for the specific task of proving things which is not the only thing we care about but but and proving things that we've already specified we want to prove. Yeah but although even formal verification does have weaknesses language itself can be exploited by by malicious agents so and so and AI may sort of attempt to be helpful and try to prove as many things as possible just add secretly add some axioms to to to

the formal system and things you can try to to shut them down but but if the AI is too powerful actually at some point you have to sort of limit how capable your AI is or have periodically humans involved in the So you know there are other in the other sciences you can do some of this so for example numerical simulation can be used as a verifier in some cases but again you can't rely on I give you to say what you want to model the weather and you have a you have a supercomputer that that that that that predicts the weather and you haven't trained an AI to mimic the the numerical simulation. Um

it is possible that at some point they will just exploit some feature of the numerical simulation that is not part of of the ground truth. Um so um it will work up to a point um and then and then it will stop. So you um we do need to get a lot better at knowing the limits of our verifiers. Um so you know a lot of verification systems that we have, they work just fine if if they're used non-adversarially. Um but you know if you're training an AI specifically to maximize it output based on um using this verifier, it will find the exploits. Yes, AI is so good at that. Um yeah, it's it's it's

a ruthless cheater. Um so yeah yeah yeah

you um we we do have to to be aware of that and and um just because a human you know a verifier so passes a human test, it it it may not be suitable for AI use. That makes sort of sense and and intuitively to AI cheating, the easiest way to make something measurable is to design it to be measurable from from day one from step one. Um Mark, do you do you think like that when you're trying to make the models ever smarter? You think in terms of first principle, what what would have to be true to be measured as being smarter? Or do you rely purely on generalization to try and get ever smarter models?

Yeah, so I really think when it comes down to it, why do we care about attacking math and physics at at a place like OpenAI? And it really comes down to we are out of good evals, good human written evals and science doing science is the eval now. And um math is particularly exciting because you can you know attack some kind of theorem, you can verify it um in many cases and you know you feel confident that you're legitimately pushing the frontier forward. I know there are initiatives in physics, too. You I know in physics there's a little bit more hand waving around oh you know this is constant too small and so you

know but you can still you know build pretty formal systems, right? And then I think um and so you know it allows us to kind kind really push push the frontiers in both math and physics. Um but fundamentally one of the reasons we cared so much about reasoning in in informal language is we care about generalization, right? We want to be able to do deep reasoning in fields like biology, too, and create breakthroughs there. Um even if it's kind of fuzzy what a breakthrough means, right? I think in in math it's much more clear. It's like you solve an obvious Stokes, yeah, that's a big breakthrough. Um if you kind of the model says, "Hey,

here's your next breakthrough in machine learning." I mean, I I don't know how to verify if if that's true. And um I think it's it's just so empirical and you're kind of time tells with a lot of these things. Um so I I think what we care about is this fundamental generalizable reasoning layer. Um natural language is feels like a good way to express this in a way that kind of falls less into this trap of like you have a tool bag of techniques and you just center on the known techniques. I feel like in in natural language we are able to kind of express, you know, um these these new techniques at least uh

I I think um yeah, we we've been able to do that so far. So um yeah, we we really deeply care about generalization and um and I do think kind of these formal fields, um they give us, you know, a really rigorous way to test that we're pushing the frontier.

[clears throat]

Beyond the structural nature of of maths and therefore the the ways that you can formally verify is there some other practical benefit to pursuing ever greater capabilities in that space or or is it in your mind really more just a equivalent of an email? Well, I think one one positive feature for using math as a test bed for other use cases is that um so you know, we we had this this quote earlier today of Vladimir Arnold that mathematics is the place where experiments are cheap.

It's also the place where failure is cheap. Um so you know, it's it's related. You know, so you know, if you're an engineer and you're asked to build a bridge and the bridge collapses, that's an expensive mistake. You know, if you're a surgeon and you're asked to and you cut the wrong thing, that's an expensive mistake. You know, but in in math if you if you try to prove a theorem and you and your proof actually doesn't work, that's not an expensive mistake.

[clears throat]

So, I think it's it's we have this freedom to fail which is more so than in other disciplines. And because of this we have a culture of learning from our mistakes a lot more than in in other other disciplines. And so, it's a relatively safer place to experiment with AI than you know, let's say bridge building or heart surgery. Okay. Yeah, I I love that you say that. It's exactly the way that we think about things that at OpenAI as well. Because I think fundamentally what we care about developing AI for, the the really really inner core goal is to use it to develop stronger AI, right? We want to design better experiments to build

a stronger models, more intelligent models, which you know, by extension will do you even better math, but you know, it will build an even stronger model and you you get that flywheel. And that is an expensive thing. You know, if you screw the system up in any way, you know, you compute that state, right? Like you you run the wrong experiment, you you burn a lot of money, a lot of compute. And so, I do think about math, physics as you know, safe domains to push the frontier. Yeah, it makes a lot of sense. I mean, I even wonder if you can push that in in Kevin your work touches on this. Assuming a world

in which the models may be discovering things that really are beyond the frontier of human knowledge or even the ability for a human to really conceptually follow. Presumably there needs to be some way to re-represent those findings into a logical chain that at least we can follow the steps if not the actual constituent parts. And that math more than any other domain seems to have invented a a workflow that accommodates that. Certainly compared to say a biology or chemistry which doesn't have as much by way of kind of formal axioms that you can work against. So, potentially it's a necessary precursor to any truly frontier science advancement to have this capability. Uh regardless though, it

it does to a point Terry suggests that the way that we think about doing math in the future will change. That we're emphasizing maybe uh creativity, collaboration, different skills perhaps than than what have been happening the last 100 years. Does that filter then through into how you teach maths? Yeah, this it's um an answer [snorts] open problem still how to um Yeah, so I in the very short term yeah, some things have had to change. Okay, so yeah, like I homework weekly homework assignments have been the first casualty. Um yeah, um for instance, you know, it's um but you know, um I think I think we can push our students to do more ambitious things

now. Um uh so I I I I've switched much more to a project-based type of assessment. Um in smaller classes you can you can do some oral assessment. Um um The skills we need to teach will be different. Uh yeah, so so validation for independent the ability to independently verify um AI-generated output will become essential. Um Yeah, um

[clears throat]

softer skills um yeah, how to work with people. Um mathematicians have not been not uniformly good at that in the past. Um but uh we'll have to get better. Um [snorts] Yeah, it it's um um we yeah, um the pace of change is such that uh education systems are not catching up as rapidly as as as but I think by necessity we'll be forced to I mean with COVID for example, we we did do some emergency changes to our our curriculum and it it kind of worked. It was not a great experience. Um So, uh hopefully this time we can do with a bit more planning. Um but but we won't it will be on

that level of change, I think. Yeah. Yeah, um I I think the analog of that is, you know, our interviews became busted very quickly, too. You know, I think you know, people if they have, you know, time to do some kind of take home or some kind of, you know, written type of interview, it's very hard. I I do think kind of moving to a world where um, you can have a model also just kind of interact with you and teach you things um, and the model itself can judge, you know, how much are you are you learning? Are you uptaking? Um, you know, that that actually kind of feels like a directional really good

update. Um, I've thought about revamping interviews in the form of, you know, you convince the model that you have the skills necessary to work at OpenAI. So, you know, I mean I I I and of course, you know, you have to prevent hacking and jailbreaking and stuff like that. But um, yeah, I um, yeah, I I I I really do think, you know, fundamentally uh, yeah, teaching has to change in in some way. Um, I I am kind of curious to get Terry's take on a couple things. First, you know, um, I've heard from some other professors that, you know, it's Yeah, you really do see this divergence of like, you know, it's like the

worst uh, the best homework and like worst live exam scores in history. I don't know if that's a trend that that you see as well. Um, I um, and I think the second thing is like, do you actually see this um, divergence of students who are like very motivated to learn that get really good using the tools, um, and is there this cohort that really feels accelerative? Yeah, so definitely have noticed homework scores going up and in-person scores going down. Um, not so much it's not not a collapse level. Um, I mean, um, it's um, I don't have hard data. Um, I I do get a sense that the weakest students are using AI to

sort of get to an immediate level. And the brightest brightest students generally tend to avoid using AI just because they're worried about um, they can notice um, that using it too much atrophies their scores. Um, the weakest students I think they they they feel like they have less to lose. Um, but um, what so it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it

it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it it Yeah, I mean once you are have a certain of expertise these tools are great. Um, and so maybe the equilibrium is is is to actually discourage their use or or use them in specific um, ways. I I I can certainly see um, home because I'm in the future where uh, the solution

is not the point because anybody can can get into it. But for example, what what prompt did you use to get to the solution? Um, and and that might be the more interesting um, assessment tool. So yeah, we we we have to figure it out. Um, yeah, and actually it's important you know, I just just just like AIs will optimize whatever reward function the reward function we give to our students will actually make a lot of a lot of difference. Yeah, we have to think this carefully. Yeah, it's um, in some ways the extreme cases are easy to understand. A total cognitive offloading and therefore no learning is occurring etc. Maybe the more nuanced uh,

there would be something where it is a productive use of AI. And I think you used this metaphor in a recent interview uh, which is to say you've been helicoptered to the destination as opposed to taking the the scenic route there. There's something about the change in workflow corresponds to something uh, at the cognitive level for humans. And that we don't know yet what it may be that you lose if you do that, but we need to be a alive and alert to it. And do you have a thesis about what we might lose if we if we start doing that? Um, I think we'll see empirically pretty soon. Um, so um, yeah, I think

we I think we we just need much more um, awareness of all the different facets of um, research or any other task. So somehow um, so as I as I said before AI allows for a decoupling of of many things which which can be good for division of labor is more efficient but it does mean that that goals that previously it was okay to to set very fuzzy goals because any human attempt to reach these goals would sort of also hit all the nearby goals as well. Um So yeah, so as I said with with an AI you know like yeah if you want to go see a a nice waterfall or something along on

on [clears throat] on on a on a mountain you know you take a hike and sometimes you see some interesting wildlife or you you get a glimpse of an of an even even nicer location that you might want to go to someday and and and maybe you meet some some other hikers and you have conversation there's all serendipity which just naturally happens and so in the past we we would just say it's a good idea to to to go visit this waterfall but we didn't we didn't sort of unpack that carefully enough to see why we do that and what are what are the actual values what what what what what's the actual benefits? Um

And so but now we have this alternate way to to get to this waterfall as I said you can you can get an AI helicopter to drop you off there and so yes you you get your your little Instagram photo but maybe that's that's that's not the only thing that you wanted. Um so yeah I think unfortunately we're going to have to learn this by experience. It's it's hard to I mean you can talk romantic about about the the the journey and and things but I think it's only when we see what what happens when we don't have that that will really understand what we're missing. Yeah, I mean I should say just today we

we released our learning outcomes measurement suite which is how we use the models to assess whether humans are learning when when they use them. So agreed it's sort of a live research question. But I I wonder Mark for you um serendipity and the idea of an inexact answer from the model in order to create space to explore uh is that a uh quality of the models that you're interested in exploring? Is that a model behavior question,

[snorts]

personality question? How do you grapple with that? Yeah, so actually one of the biggest initiatives we have this year in terms of building a new primitive and a new interaction paradigm with the AI is we started an interactive agents team. And I think it's not sufficient that you just ask an AI a question and then it just comes back even let's say like a day later with its best attempt at a solution. I think you you know humans are collaborative. You know they they work in these these constructs and you know you take like a completely non-math example, right? You want to create some kind of let's say PowerPoint or some kind of artifact like

that, right? You that's not the way you operate, right? You don't just tell some AI like just make me a PowerPoint and should be perfect and should kind of address these things. Like you you want it to kind of come back and you shape the direction it's going and and there's like multiple rounds of interaction with with this agent. And I think that that truly is what it's like to co-work with with a very intelligent agent and we want to build that deeply into the model. Just you know something that's very steerable. It feels like a thought partner and I I do hope you know within a couple months at least within a year that's

the way the AI looks. Yeah, it's much harder to do reinforcement learning on collaboration. Yeah, yeah. How do you score how how good you're vibing with your co-workers? Yeah, yeah.

My thesis is that it is possible. Yeah. Yeah. I agree that it might it might not be as difficult as as you imagine. I I there there are pretty quite clear like biological signals of what what vibing looks like in the real world that you can bring back into the machine world. Okay, yeah. So embody your AI with body language. I think they're they're vibing too. Yeah, there we go. I'm glad that I've got a quote that will be yeah, kept after this. I will open up the floor for questions uh just after this. I'll ask one more just to give you all time to think about it. Uh perhaps we will meet again

in a year and I'd be interested to know what what your predictions look like for what that what where we'll be in a one more year's time. Um I really hope um we're going to see a lot of um new types of methodical projects that are kind of challenge-based where um you know, so like first proof type things where um where some group of mathematicians for example will will will create a really good creative set of problems that um they would like some proportion at least one to solve. Um and and they have a very good gradation of difficulty, they have a very good verification protocol, and uh they would just open it up to

to the to the the community. Um and so it's um just just sort of taking full advantage of of um well not just AI, but also just like the internet and and the Metcalfe's law. You know, if there's if there's if there's n people who can produce problems and people that can that can um solve problems, then there's n squared possible connections. Um and and mathematicians [clears throat] have been very bad at using that sort of making this large-scale network. Um so I I think we will see a a different almost like a marketplace for style of of doing mathematics. Uh so um and there I think we'll see AI shine. Uh so so this

is one thing I want to see and and maybe in a year we'll start seeing that. Amazing. Yeah, I I sure think um of ML's kind of foreshadowing math here. Like um when you look at how Frontier Labs operate today with research scientists, um we are moving into this world where the strongest research scientists, they're able to kind of pursue a lot of ideas in parallel and just really act as orchestrators. So, you know, they can you know, think of about this idea, uh think about a bunch of variations in the experiments, and just kind of have the model go and execute and implement that. Um I I hope there is that kind of similar

paradigm in in math where you know, people like Terry and and yourselves feel empowered to just go, you know, explore you know, broad set of ideas and strategies um with very little hand-holding. I I do think kind of the the very little hand-holding part will also become more true. The task horizon will continue to to elongate. Just like we were at minutes of you know, in terms of the horizon a year ago, I think in a year from now we're going to be in multiple days where you know, you can actually trust the model to do tasks that would take you that long. And then I think beyond that, yeah, it's just making [snorts] sure

the interaction is seamless, right? These things should just feel like they interact very naturally with groups of humans and and with the communities that you guys operate in. And finally, I really do hope we have some really big breakthrough whether it be in math, in physics, in in biology. I think you know, today um you know, the things we're proving are they're good and all, but I I do think there's a potential for this to actually you know, produce something that's very beneficial for humanity. Fantastic. There there was this metaphor in software developments of bazaars and cathedrals. A bazaar being a self-organizing thing that springs up and is very diverse. A cathedral being one great

mind architects and is therefore very elegant. And that the idea perhaps will be that the mass we get both those phenomena occurring

[clears throat]

and hopefully that is a flourishing. So, I do want to open up to questions. Does anyone have a a burning question? Yeah. I wonder if you know, if you can talk about the world models which instead of predicting the next token predict the next state and what I'm reading about the video party ways is that it can actually self-correct thing because the hallucination can be avoided. Is that all true because whatever I could run on my match contest model is just a hello world thing. So, I I don't know much about it. For for the benefit of those watching, the question was about how do we feel about world models? Are they a paradigm shift

and would they be specifically useful for maths and solving hallucinations? I think I think it's it's it's potentially a very promising alternate direction. LLMs are great and and in some ways they're they're too great, actually, in that we've kind of routed our entire AI infrastructure around making the the LLMs as as as as as powerful as possible, and it could crowd out some other very complementary ways to um to to um create AI assistants that have a are jagged in a completely different way. Um uh um So, uh definitely support research into world models. Um I think for a long time they will underperform the LLMs because uh we have just because of it's all

the momentum and infrastructure LLMs have. Um it's it's like we have built our cities around the automobile and and and and and gasoline, and we have this entire infrastructure, and it's actually making hard for for alternate motor transportation to break through. Um but there are definitely people who are are pushing that, and and and I wish them a lot of luck. Um yeah. Yeah. I I do think um when you think about a pure video, you know, generative video world model, um we still seem pretty far from that. You know, I think the the existing video models, they're pretty good, you know, physics simulators, but I do think with a little bit of our own

pressure they also fall apart. Um I I do imagine that'll get more and more robust over time, but, you know, you just um it's not quite there yet. We are pushing fairly hard on that. I think, you know, there's many spectrums of world models. You can see an LLM as a world model, too. Um But, you know, I think digital world models where, you know, we're interfacing with with computers, you know, there's all the rules, and and the feedback of a computer, um that's a very important and interesting system, and I do think we'll tackle and and and really get a lot of value from that very soon. I wonder on that if um there

is a middle ground in as much as you can construct our LLM environments that are based on on the laws of physics or follow the laws of physics, which to some degree confers the benefits you otherwise get from a world model to a LLM or anything else. So, perhaps it's an intersection rather than two alternate pathways. Yeah. So, AI in science, many fields, well, is effective when it's predicting very well. For example, protein folding, predicting weather accurately. But in mathematics and theoretical physics, we're asking for something different, like we want to kind of we can [cough] get a formula, get a proof. Mhm. But is it I do think it's potentially too limiting? Will it

be easier to get a, you know, an AI to potentially you I have a proof for my hypothesis, but your your brain is too narrow to understand it. So,

I I can teach other AIs about it and do more progress with it. Yeah. That's definitely my relationship with AI already. So,

some of us have hit that frontier. Um again, for the benefit of those watching, the question was in some domains of science, I think we're satisfied with pure simulation. If if it can do the thing, we consider the thing to be proven, even if you can't formally verify it. So, if it can predict weather accurately, even if we don't know how, we're kind of happy with that that outcome. We hold in maths and physics to a different standard, which is that it must be uh verifiable in the way we already discussed. Is that somehow limiting or or a mistake to to put that restraint? Yeah, I I I think there'll be different types of mathematical

tasks that we don't do nowadays, which we would trust AI to. And and and and they could be quite complementary to the task of of getting a formal proof of of of a problem. So, [clears throat] to give you an analogy, so, in chess nowadays, all chess players train using these chess engines. And one thing that a chess engine does is it gives you this score at any time in this position, you know, it's white is is is three points ahead or whatever. And it's a really great signal to train human chess players. You know, you get instant feedback, oh, that was a really bad move. Okay, I'll I'll take try to do this instead.

Um I could imagine um an an AI which, you know, you're trying to human is trying to do a proof, and like every time you say, I'm going to try to prove a contradiction, you know, your your score goes down. Wait, okay, that's a So, bad idea. Okay, you should back up and do something else. So, Maybe a small electric shock just to

So, that's that's a pro the pro model. Yeah, okay. So, uh Yeah, um so, yeah.

Yeah, a good math tutor can do that, but yeah, but maybe so, we have to be creative about the type of tasks that maybe AI could could help with that we just don't think about today. Yeah, I mean, I do think uh verification is important. It doesn't have to be formal verification. I think we deeply want to know why something is true. Um and I think that's just actually part of a deeper alignment problem, right? When the AI is attacking, you know, actual real-world impact tasks, you want to know why it made a certain decision, right? Let's say you decided here's the best strategy to um I don't know, like uh you know, grow a

business or something, right? You you don't want it to to do that without having a good justification. Um and and so, we have a lot of alignment techniques like debate, right? Where even if you don't get necessarily an airtight formal thing, you can kind of understand the outline and kind of like interact with with the proof and and kind of question it. So, yeah, I do think kind of investment into techniques like debate and and alignment, they'll they'll really help us in the future. So, I'm I'm not

Please can you get anything in the sort of the latent space stuff like see what the interview layer is thinking Yeah, yeah. So, um that's something we we look into a lot, right? Um I think the the top-order thing is, you know, just being able to monitor the the reasoning of the chain of thought and uh you actually get a lot of insight from there. You can actually like do a lot of even with many attempts on a problem, kind of get a sense for like what strategies the model gravitates towards and and and and just kind of glean a lot in of insight into how how how the model brain works. Um yeah, I

think you can go a level deeper like look at the activations and like try to find mechanistic circuits and things like that, but um yeah, definitely a a very deep field of study. Um these two questions may link up in as much as um for interpretability, the way that we collapse the latent space does limit uh the kind of associations that could come out of the model. At what point do you decide that it's better to maintain the latent space because of the theoretical new connections it can make at the cost of interpretability? Or do you just think that we need a different interpretability paradigm that doesn't require that kind of collapsing? I I think

the reason we operate in text space today is interpretability buys you so much, right? I think you can debug so many things that go wrong with the model by just being like, "Oh, well, clearly it's like reasoning wrong here, so there's you know, we we can go and and debug." When you're doing something in in just like pure in uninterpretable latent space, it you lose that. And um I I don't know that we would switch to something like that in the long term. Well, ideally we should have a diversity of models. So so so maybe there are some applications where you just want the answer, you don't care about interpretability, and then you just turn

the dial one way. But [clears throat] but there would be other applications where you you really want to see the process, you really want to see a human readable um a channel thought or whatever, and you turn the dial the other way. Yeah. I mean, that it it does feel intuitively like having to compress things into language does come at a cost. So even a alongside a priority of models, you might also want a priority of expressions or or verification methods so that you don't somehow force it into a shape that doesn't make sense. Sorry, please go ahead, Carl. I have a question about attribution and and something that was a little bit as we

look to this kind of future of of science. So one thing with AlphaFold, you know, uh sort of the the the world I think largely thinks like AI came and solved that problem, and of course AI in some sense did, but it was sitting on the protein data bank and decades of of effort. And then you see the protein data bank loses its funding immediately after various things like that, and I'm just curious that you know, as as we think about like these these large-scale math problems and all of the human effort that's going to need to go into producing the sets of problems and verifying them and etc. It does seem like there are

for in many of these things, theoretical physics problems, all of them, that there's a lot of danger that in some sense while AI was maybe like the critical enabler that made something happen, it didn't happen on its own. It's really this ecosystem. And that so on one side it's how do we control that narrative? And the other part is that in some sense like that like for OpenAI and for the big companies there's a lot of uh well, you know, responsibility in some sense about how how they navigate that. And I'm just curious about, you know, how do we avoid it going in a bad direction? Um well, um just yeah, this is a this

is an important point. Um a partial solution I I so um as I said, I I do envisage the rise of like challenge problems where people will create these data sets of of tasks that they want solved. And there it's kind of win-win because the the the people who who create these data sets, they will get the problems some fraction of them solved, which is what they want. And then but these data sets could be very useful to to calibrate AIs and and and and so there there are some cases where it can be win-win. Um but yeah, there there are definitely cases where people have built a data set at great expense for not

for this reason and then it it gets absorbed into various AIs and um yeah, I I don't know how well we can track that. Um yeah, so I I I it leads into like yeah, intellectual property right law and yeah. It is a very tricky problem. Okay, which which I will toss to you to deal with. No, I think as of now, I think the AI doesn't want your credit. So, you know, I think uh um but yeah, I I do think um the vision we have for for OpenAI for science, it's really not about us claiming the credit here. I I I do think um I mean, we certainly have the ambitions to

to move science forward, but uh you know, Kevin, you're you you wants to build a platform where mathematicians around the world can just accelerate the field in in in its totality. I think like we don't know the right questions to ask. We aren't the orchestrators within OpenAI. I I do think kind of the credit should just go to you guys. I I know that's not exactly the question you're asking. I think that it's not like it's not that AI is going to claim credit, but I think the public perception is AI solved this problem and then comes somehow like humans and

[cough]

experiments [clears throat] and data and all of this is not like that. So what the earliest problem is what we've seen is that the weak times when there's an open earliest problem that no one has looked at and then some AI solution just gets a solution and then this hits social media. AI solved an unsolved problem. Okay. And in many many cases, you know, like 24 hours later someone often armed with one of these deep research tools is uncovers that this this this result was actually already proven by a very similar method in the literature and we can't say for sure whether the the AI

[clears throat]

used that that solution or indirectly was aware of it, but it it happened so frequently. I mean we have we have this whole table of of of AI contribution. It's it's like a whole section which is just this. Um and um so to some extent um we we at least have the capability to detect at least some of this that that because we also have these these these research tools. So we can recover attribution sometimes. Um it's not perfect. Um but yeah it it could be that the same technology that allows us to to to um use the literature to solve problems can also use the literature to attribute solutions. Yeah, just one thought on

top of that. In in general it is a very hard problem just data attribution and you know, when you generate something, you know, how inspired is it from which data points. One interesting thought here is perhaps like the novelty and contribution is somewhat, you know, correlated to the amount of time the the models spend thinking about something. Um, maybe it's not always true, but um um yeah, I I I think, you know, modulo stuff it rediscovers in literature. Um I think there's there is a PR element though that's just like like DeepMind could have gone out of their way to say more about uh protein data bank and the narrative like there's things that are

not about the AI model and data so much as about just how we talk about it and PR and all that.

Yeah. Yeah, that's that's a very fair point. I I understand the incentives. Um I I do hope, you know, anyone here at OpenAI can can can verify this, but like we we care a lot about integrity. I think um you know, it um at least I I would very much strongly fight for the correct narrative there. Maybe just uh to add a flavor from my my previous life. Um maybe one paradigm shift is that the general public tends to underestimate the degree to which the progress of science relies on improving tools fundamentally. You need better microscopes so on and so forth. These are not glamorous jobs. It's not typically what what theoretical scientists like to

work on. Um and that that's a narrative we should be saying is that by building ever better tools for science, that's what enables humans to accelerate the entire field. Often, really big steps forward. AlphaFold is really a tool. It's not per se scientific research so that it does do some uh research component things there too. And if we can stay on the emphasis of it being a tool, that might also flow back to your your point about where does the investment flow in society because it's to be things that need to surround the tool to make it effective, but ultimately in service of the human researchers or scientists that will use the tool to solve

the problem. Uh we are just in the edge of time so I'll take one more if that's all right.

Yeah. Okay, perfect. We could go to the back. Yeah, so uh I think one of the really interesting things that's happened is you know, it's accelerating math and physics using AI and I think there is a certain amount of kind of synergy that comes from like being able to solve AI or math and then how that can help solve like other physics problems. So, I'm curious to hear your thoughts on like additional synergies you see coming out of Open AI, you guys [clears throat] kind of tackling math, physics, and then kind of what what else you guys see out there in the future. Yeah. Yeah, so I I think Kevin would probably be the best

to speak to this, but we we do care about tackling domains outside of math and physics as well. I think some that we have explored are biology where we have had the AI work on just making the biological procedures in the wet lab much more efficient. I think with one of our partners Ginkgo Bioworks, we iterated on a lot of their core processes and made the cost for synthesizing proteins I think 40% more efficient. And I think that's just the underlying primitive that will will drive more progress. So, a a lot more you could imagine doing in things like material science and and other domains. Yeah. And so the IPAM, this institute, I mean our

entire core mission is basically to find these synergies. With you know, Institute for Pure and Applied Mathematics is basically in the name. Yeah, we we bring together even like this you know, where different communities talk talk together. Yeah, so and a lot of it is this serendipity. I mean we have some idea you know, we we don't just randomly smash together fields. Okay, but but yeah, we we we we do pick ones where we do believe there will be a lot of unexpected fruitful collaboration. I'm glad you leave randomly smashing together particles to the physicists.

yeah.

That's a good moment to to close it and also to plug that Kevin is about to give a lecture which covers many of these questions. I mean what needs to be true to accelerate all parts of science and the tools to build for it. So, I hope you'll come to that, but thank you all SO VERY MUCH.

MHM.