🎙️AI 访谈库
ChatGPT 内幕、AI 助手与在 OpenAI 做产品——OpenAI 官方播客第 2 期(Mark Chen 与 Nick Turley)
Mark Chen · OpenAI 首席研究官(与 ChatGPT 负责人 Nick Turley 同台)

ChatGPT 内幕、AI 助手与在 OpenAI 做产品——OpenAI 官方播客第 2 期(Mark Chen 与 Nick Turley)

Inside ChatGPT, AI assistants, and building at OpenAI — the OpenAI Podcast Ep. 2

2025-07-01 · OpenAI Podcast (Andrew Mayne) · 1h7m · 约 77 分钟读完 · 原文
Mark Chen 与 Nick Turley 同台:Turley 回忆发布前夜团队临时把 'Chat with GPT-3.5' 简化为 'ChatGPT'、上线前的内部争论与病毒式爆发,两人还谈及谄媚事件与 RLHF、ImageGen 的突破时刻及智能体编码的下一步。

Hello, I'm Andrew Maine and this is the OpenAI podcast. My guests today are Mark Chen who is the chief research officer at OpenAI and Nick Turley who is the head of chat GPT. We're going to be talking about the early viral days of chat GPT. We're going to talk about image gen, how openi looks at code and tools like codeex, what kind of skills they think that we might need for the future, and we're going to find out how chat GPT got its totally normal name.

Even half of research doesn't know what those three letters stand for. You know, you're going to have an intelligence in your pocket that it can be your tutor, can be your adviser, it can be your software engineer. There's a real decision the night before. Do we actually launch this thing? First off, how did OpenAI decide on that awesome name? It was going to be chat with GPD 3.5 and we had a late night decision to simplify.

Wait, wait, say that. Say that name again. It was going to be chat with GPD 3.5 which rolls off the tongue even even more nicely. That's uh and you said that was a late night decision meaning like weeks before you finally decided what to call it. Right. Right. Right. No, weeks before we hadn't started on the project I think. Oh goodness. But you know I think we we realized that that would be hard to pronounce and um came up with a great name instead.

So that was the night before roughly. Might have been the day before. It's all kind of a blur at that point. I would imagine a lot of that was a blur. And I remember here uh I remember being in a meeting where we talked about the low-key research preview which like really was like we really thought like oh this is cuz it's it was the 3.5. 3.5 was a model that had been out for months. And from a capabilities point of view, when you just look at the eval, you're like, "Yeah, it's the same thing, but we just put the interface in here and made it so you didn't have to prompt as much and then chat comes out."

And when when was the first sign that this thing was blowing up? I I'm curious for everyone has their slightly own recollection of that that era because it was a very confusing time. But for me, day one was sort of, you know, is the dashboard broken? Classic like, uh, the logging can't be right. Day two was like, oh, weird. I guess like Japanese Reddit users discovered this thing. Maybe it's like a local phenomenon.

Day three was like, okay, it's going viral, but it's definitely going to die off. And then by day four, you're like, okay, you know, it's going to going to change the world. Mark, did you have any expectation about that about No, honestly, I mean, we've had so many launches, so many previews over time and yeah, this one really was something else, right? The takeoff ramp was huge and yeah, my parents just stopped asking me to go work for Google.

Wait, so wait, wait, wait a second. Up until Chat GPT, your parents were asking like what you're doing here. Yeah. No, I mean um they just never heard of OpenAI. Um I think for many years thought AGI was this pie in the sky thing and I wasn't having a serious job. So it was a real revelation for them. Yeah. Uh what was your job title at the time? Um I think just member of technical staff. Member of technical staff and then then that blew up and now you're head of research I guess.

So yeah so all right. Yeah. Actually on on the GPT name um I think even half of research doesn't know what those three letters stand for. It's kind of funny, you know, like half of them think it's generative pre-training, half of them think it's generative pre-trained transformer. And and what is it? It's the latter. Okay. All right. Yeah. Those people, they don't know the name of it. Yeah. Uh it is it's weird how just a silly name like that um all of a sudden becomes a thing.

But you see that with like you know Google, Yahoo, Kleenex, things like that. Xerox and sometimes they were some of those were names by intention and this was really just a a silly sort of name. For me the moment that I felt like after watching the launch, watching it accelerate, I knew was going to happen. And then when it did it was when it was on South Park and remember that when South Park made fun of the name and that was the first time I'd watched South Park and let's just say a while and that episode I still think it's magic and it was obviously profound to watch and see you know something you you helped make show up in pop culture but there's the punch line at the end where it's like ah this was co-written by Chachi BT I think they took that off though I think I think in later episodes cuz it used to say I think written by like uh Trey Parker and like Chacht and then No, it is and then I think later I think they may have pulled that off at some point.

I don't remember like well I strongly feel that you shouldn't have to give credit to it's your business whether you're using the I had to give credit to Chad GPD for every aspect of my life. Um well might as well just say Chad GPD maybe with Andrew. So do you use it for prep for your interviews? You know one of my my co-producers Justin probably uses it. I haven't asked him yet cuz I'd like to think that he's handcrafting every single question that we're thinking about here, but I am sure.

You say it was a bit of a blur and I'll tell you like a standout moment for me at the launch of Chad GPT was I don't know if you remember this, but the Christmas party and we'd had several weeks of Chat GPT out there and Sam Alman went up and said, "Hey, it's been exciting to watch this, but the internet being the internet and I think we all felt this way. It's going to die down." Spoiler alert, it did not die down and it just kept accelerating.

What were the things you had to do internally to sort of keep this thing up and running as more people wanted to use it? We had, you know, quite a few constraints and if if for those of you who remember, you know, I think you guys remember touch was down all the time. Yeah. In the beginning. Um and that was, you know, we' said, "Hey, this is a research preview. No, no guarantees you maybe goes down, but the minute you had people loving and using this thing, that didn't feel super good."

So, you know, people were certainly working around the clock to keep the site up. I remember, you know, we obviously ran out of GPUs, we ran out of database connections, we had, you know, um, we're getting rate limited in some of our providers. Nothing was really set up to run a product. So, in the beginning, we just built this thing. We called it the fail whale and it would just tell you kind of nicely that the thing was down and made a little poem.

I think it was generated by GPD3 um about being down and and it was sort of tongue-in-cheek and that got us through the winter break because we did want people to have some sort of a holiday and then when we came back we were like, okay, this is clearly not viable. You can't just go down all the time. Um and eventually we got to something we could serve everyone. Yeah. And I think you know the demand really speaks to the generality of Chad Grippy, right?

um we had this thesis that chatbt embodied what we wanted in AGI just because it was so general and I think you know you're seeing that demand ramp just because people are realizing you know any use case that I want to to give or to throw to the model it can handle we were kind of known as the company working on AGI and I think prior to chat GPT the API was certainly the first time we had a public offering where people could go use it and do it but then it was more for developers and stuff and I think that as long as people were sort of thinking AGI I that seemed to be the point at which people thought these models would be useful but we saw GP3 we saw that that was useful and then we saw that we could do other things were useful um was everybody at OpenAI on board with chat GPT being useful or being ready to launch?

Yeah, I don't think so. you know, um, even the night before, I mean, there's this very famous story at OpenAI of, uh, you know, Ilia taking 10 cracks at the model, you know, 10 tough questions, and my recollection is maybe only on five of them he got answers that he thought were acceptable. And so there's a real decision the night before, do we actually launch this thing? Is the world actually going to respond to this?

And I think it just speaks to when you build these models in house, uh, you so rapidly adapt to the capabilities and it's hard for you to kind of put yourself in the shoes of someone who hasn't kind of been in this model training loop and see that the there is real magic there. Yeah. Yeah. Yeah. I think to build on that like the contray internally about, you know, is this thing good enough to launch, I think is humbling, right?

Because it's just a reminder of how wrong we all are when it comes to AI. why, you know, frequent contact with reality is so important. Could you elaborate more on that contact with reality? What does that mean? Yeah. I mean, when you think about iterative deployment, uh, one way I like to frame it is, you know, there's no point everyone agrees where it's suddenly useful, right? Um, and I think usefulness is this big spectrum.

Um, and so, you know, there's not one capability level or one bar that you meet and suddenly, you know, the model's useful for everyone in. Were there any hard decisions about what to include or what to focus on? We were very very principled on chatbt to not balloon the scope. Um you we were adamant to get feedback and data um as quickly as we could. Slack telling you things like Nick add this add this. I remember actually there was a lot of controversy about like the UI side.

For example, we didn't launch with history even though we thought people would probably want that and you know guess what that was the first um um request. I also think there's always the question like can we train an even better model like is you know with two weeks more time. I'm glad we didn't because you know we I think got a ton of feedback as as we did. So um yeah there was a ton of those scope discussions and you know the holidays were coming up.

So I think we had this kind of natural um forcing function for getting something out. Yeah there was this this habit of things that if it's going to come after a certain point in November it's not going to come out like February. You know there's a sort of window where things would fall on either side. Oh that would that that would be the classic in a big tech company. I think we're we're definitely a bit more flexible than we should.

I I I felt like one of the big impacts was once people are out using it, it felt like the rate of these things improving was tremendous. I don't know if that was something that we really had in the calculus. We could certainly think about training on larger site more data scaling compute, but then the idea of actually having them the signal you would get from that many people using it. Yeah, I think over time, you know, feedback really has become an integral part of how we build the product and it's also become an integral part of safety and so you always feel the time cost of losing out on feedback.

You know, you can deliberate in a vacuum, right? Uh are they going to respond to this better? Are they going to respond to that better? Um but it's just not a substitute for just bringing it out there, right? Um I think our philosophy is let the models have contact with the world. And if you need to revert something, that's fine. But I think uh there's really no substitute for this fast feedback and it's become one of the big levers for how we improve model performance too.

It's sort of funny like I feel like we started with shipping these models in a way that is more similar to hardware where you make like one launch very rarely and it has to be right and you know you're not going to update the thing and you're going to work on the next big project and it's capital intensive and the timelines are long and over time and I think chatbt was kind of the beginning. It's looked more like software to me where you make these frequent updates.

Um you have kind of a constant pace the world can adopt. Something doesn't work you roll it back and you sort of lower the stakes in doing that and you lower you increase the empiricism and and of course just operationally too you can innovate faster and in a way that is more and more in touch with what users want. Yeah. One of the examples we had of that was the the model becoming too obsequious or sophantic. Could you explain what happened there where that was where people all of a sudden say, "Hey, it's it's telling me I've got 190 IQ and I'm the most handsome person in the world," which I had no problem with personally, but other people did and what was going on there.

Yeah. So, I think um one important thing is we rely on user feedback to improve the models, right? And it's this very complicated mix of reward models which we use in uh a procedure we call RLHF, right? uh using human feedback to use RL to improve the model. Can you give me just like a brief example what that would mean? Yeah. Yeah. So I think um one way to think about it is you know when a user enjoys a conversation you know they provide some positive signal um thumbs up.

Yeah. A thumbs up for instance and uh we train the model to prefer to respond in a way that would elicit more thumbs up. Right. And this may be obvious in retrospect but um stuff like that if balanced incorrectly can lead to the model being more syophantic. right? Um you can imagine users might want that kind of uh that feeling of you know a model saying good things about them but um I don't think it's a very good long-term outcome and actually when we look at kind of our response to uh sency and and the roll out that resulted there um I think there were a lot of good points about it you know this was something that was flagged just by a small fraction of our power users it wasn't you know something that a lot of people who generally use the models noticed And I think we really picked that out fairly early.

We responded to it, I think, with the appropriate level of gravity. And um yeah, I think it it just shows that, you know, we really do take these issues quite seriously and we want to intercept them very early. Yeah, it felt like there was maybe 48 hours since the model came out and then Joan Jen had a response explaining exactly what happened. And I think that that's the that's the hard part. How do you navigate that?

Because the problem with social media is you're basically monetized by engagement time. You want to keep people on there longer so you can show them more ads. And certainly the more people use chatbd, obviously there's a cost to open. The idea is maybe use it once and stay around forever, but that's not practical. How do you weigh that? The idea of making people happy with what they're getting versus making the model, you know, be broadly more useful than just pleasing.

I feel very lucky in this regard because we have a product that's very utilitarian and people use it to either achieve things that they do know how to do but don't feel like doing um faster or with less effort. Um or they they're using it to do things that they couldn't do at all. Um you know first example is maybe you know writing an email that you've been dreading. Second example might be you know running a data analysis that you didn't actually know how to do um in Excel you know um um true story.

So, so you know those are very utilitarian things and fundamentally as you improve you actually spend less time in the product, right? Because you know ideally it takes less turns back and forth or maybe you actually delegate to the EI so you're not in the product at all. So for us, you know, time spent it's very much not the not the not the thing we optimize for. You know, we do care about um your long-term retention because we do think that's a sign of value.

Um if you coming back three months later, that's clearly means we did something right. Um but what that means is you know um I always say um show me the incentive and I'll share the outcome. We we have I think the right fundamental um incentives to build something great. Um that doesn't mean we'll always get it right. Um the sycopency um events were were really really important and good learning for us and I'm proud of how we acted on it.

Um but fundamentally I think we have the right um the right setup to build something awesome. So that that brings up the the the challenge. I want to know how you navigate that is that one of the things early on when you know chatb came out there was like the the allegations it's woke it's woke and people are trying to promote some sort of like agenda from it and my argument always been like you train a model on you know kind of on corporate speak you know average news and a lot of academia that's going to kind of follow into that and I remember Elon Musk was very critical about it and then when he trained the first version of Grock it did the same thing and then he's like oh yeah when you trained it on this sort of thing and did And internally at open eye there were discussions about how do we make the model not try to push you not try to steer you could you go a little bit how you try to make that work yeah so I think um at its core it's a measurement problem right and I think it's actually bad to downplay these kind of concerns because they are very important things right and um we need to make sure that the model the default behavior that you get is something that's centered that you know doesn't reflect bias um on the political spectrum um or in in many other you axes of bias.

And at the same time, you know, you do want to allow the user the capability to, you know, if it you want to talk to a reflection of something with more conservative values to be able to steer that a little bit, right? Um or liberal values, right? And and so I think the thing is you want to make sure that defaults are meaningful and they're centered and that's a measurement problem. And you also want to give ability some flexibility, right, within bounds to steer the model to be a persona that you want it to talk to.

I think that's right. Um I think you know in addition to neutral defaults ability to sort of bring your own values to some extent. I think you know being transparent about the whole thing is I think really really important. I'm not a fan of of you know secret system messages that you know try to like you know hack the model into saying or not saying something. Um what we've tried to do is publish our spec so you can go look at you know if you're getting certain model behavior is that a bug?

um you know is it in violation of our own stated spec or is it actually in the spec in which case you know who to criticize and who to who to yell at or is it just underspecified in the spec in which case that allows us to improve it and add more specificity into that document. So by sort of publishing the rules of the AI that it's supposed to be following. Um I think that's an important step to have more people contribute to the conversation than just the people inside of open air.

So we're talking about like the system prompt, the part of the instruction that the model gets before the user puts the input and well I think it's um system prop is one way to steer the model but you know it goes much deeper into that right you know yeah we have a very large document that um outlines across a bunch of different behavior categories how we expect the model to behave. And just to give you an example here, right?

Um you can imagine if there's someone who comes in with just like a incorrect belief, just a factually incorrect kind of a point of view. Um how should the model interact with that user, right? And should it reject that point of view outright or should it collaborate with the user on kind of figuring out what's what's true together? And you know, we take that latter point of view. And I think there are a lot of very subtle decisions like this which we put a lot of time in.

Yeah, that that's a hard one because I think some things you can test for and you can try to figure out advance, but when you're trying to figure out how an entire culture is going to adopt something, it's challenging. Like if I was someone who's convinced that the world was flat, you know, like how much should the model push back against me? And some people like, oh, it should push it back all the way, but it's okay.

What if you're one religion or not another? And yeah, it turns out rational people and well-meaning people can disagree on how you know the model should behave in these instances. And you're not always going to get it right, but you can be transparent about what approach we took. We can allow users to customize it. And I think, you know, this is our approach. I'm I'm sure there's, you know, ways we can improve on it, but I think by being transparent in the open about how we're trying to tackle it, we can we can get feedback.

How are you thinking about as people start to use these models more and more? Um, regardless of whether or not that's some dial you're trying to turn, it's just the more useful it becomes, the more people want to use it. You know, there was a time when nobody wanted a cell phone and now we can't get away from them. And how are you thinking about relationships people are forming with with their systems? Um, obviously, you know, we um, you know, I mentioned this earlier, this is a technology you have to study.

It's not designed in a static way to do XYZ. It's it's highly empirical. So you know as people adopt and the way that they use the product it's something that we we need to go understand and and um um and and act on as well. I've been um observing this trend with interest where I think you know increasing number of people especially Gen Z and you know younger populations are coming to Chhatri as a thought partner and I think in many cases that's really helpful and beneficial because you've got someone to brainstorm on a relationship question you've got someone to brainstorm on a you know um a professional question um um or or something else but in some cases it can be harmful as well and I think detecting those scenarios and first and foremost having the right model behavior is very very important to us Um um so actively monitoring and in some ways it's one of those problems we're going to have to grapple with because with any technology that becomes ubiquitous, it's going to be dual use.

People are going to use it for all this awesome stuff and people are going to use it in ways that you know we wish they didn't and u we have some responsibility to make sure that we um um handle that with with the appropriate gravity. I I find myself having longer conversations with it. I like the memory function. I like the fact you can turn it off if you don't want. And I think about like, you know, what's this going to be two years from now or three years from now when it has a much longer memory, much more context with this.

I like the idea to have these sort of like, you know, momento anonymous modes, too, where it's not going to store this. But I I kind of wonder how much you've been thinking about two years, three years down the road, what what's that going to be like when Chad GBD knows way more about you? Yeah. I mean, I think memory is just such a powerful feature. In fact, it's one of the most requested features when we talk to people externally.

Um it's like this is the thing I really want to pay pay more for. And I think um you know you liken it to if you've ever kind of had a personal assistant, you know, you No, I'm not. Well, you do need to build up. Sorry, guys. I'm sorry, guys. But you know, it's Yeah, it's just like it's um kind of in any kind of relationship that you have with a person, right? You you build up contact with them over time. Mhm. Um and I think just the more they know about you, right, the richer the relationship, the more you know, um they can also help you, right?

You can work together to collaborate on tasks together. I I do become self-conscious of the fact like it knows everything about me when I'm grumpy. And I've I've argued with it recently, by the way. That's good. Yeah. You should be able to argue with it. You understand a lot about yourself and having a thing to argue with. Um and I think you spare others of that experience, which which can also be beneficial, but Don't argue on math and science.

You're not gonna win those. Yeah. No. Increasingly very unlikely. Yeah. Yeah. I think memory's cool. And to Mark's point, it's been part of our vision for for a long time because, you know, we said we were going to build a super assistant before we really knew what that meant. Um Chad GBT was sort of the early demonstration to that idea. But if you kind of think about, you know, real world intelligences, you even they are not particularly useful on their first day.

Um, and I think being able to solve that problem or begin to solve that problem has been profound. To your earlier question though, you know, it it really does feel like, you know, if you fast forward a year or two, chat GBT or things like it are going to be your most valuable account by far. It's going to know so much about you. And that's why I think giving people ways to talk with this thing, you know, in private is very important.

Um, we make, you know, this like temp chat thing very like literally on the home screen because we think it's, you know, increasingly important to talk about stuff sort of off the record, too. So it's an interesting question and and uh um I think privacy and AI is going to be be an interesting one um for the next coming years. I want to switch gears, talk about another release which again kind of caught people by surprise and blew up was Image Gen and uh I was here for Dolly Dolly 2 and then then Dolly 3 came out and I thought Dolly 3 I thought was a very capable model but it seemed like it preferred a certain kind of image and a lot of the utility and the capabilities for variable binding was sort of kind of hidden away and then image gen was kind kind of just this breakthrough moment that it caught me off guard.

How did you guys feel about the launch of that? Yeah, honestly it caught me off guard too. Um and this really props to the research team, you know, um Gabe in particular did a ton of work here. um Kenji, many others on the team did phenomenal work and um I think it really spoke to this thesis that when you get a model just good enough that in one shot it can generate an image that fits your prompt that's going to create immense value and I think we never quite had that before right um that you just get the perfect generation oftentimes on the first first try um and I think that's something very powerful you know like uh people don't want to pick the best out of a grid I think uh yeah you just got very good prompt following and you know just great style transfer too, right?

Um this ability to kind of put images um as context for the models to modify and to change and the fidelity that you could do that with um I think that was really powerful for people. I think I think this image and experience um it was just kind of another mini chat GBT moment um all over again where you know you have kind of this you've been staring at this for a while like yeah it's going to be cool. I think people are really going to like it.

Um but you kind of you know you're launching like 20 different things and then suddenly the world is going crazy in a way that you you kind of only find out um um by shipping. Like I remember distinctly you know we had like 5% of the Indian internet population try image genen over the weekend and I was like wow we're reaching new types of users who we wouldn't even have thought you know who might not have thought of using chatbt that's really cool and and um to Mark's point I think a lot of this is um because there's just this discontinuity where something suddenly works so well and truly the way you expected um where I think it blows people's minds you know and I think we're going to have those moments in other modalities these two.

You know, I think voice, you know, it it hasn't quite passed the touring test yet, but I think the minute it does, people are going to um I think find that immensely powerful and valuable. You know, video is going to have its own moment where you it starts meeting the expectations that users have. So, I'm really excited about the future because I think there's so many of these magical moments coming that are really going to transform people um people's lives and and also you change um sort of chatbt's relevance for people because um you know there's um I've always felt like there's text people and there's image people and like some of them are a little bit different um and now they're all using the product and discovering the value across the board.

the moment when it launched, I think it kind of illustrated the the problem that been with image models before. And you know, when Dolly came out, it was super exciting because you're like, I'm I'm like doing pictures of space monkeys and all these sorts of things. The moment you try to do a really complex image, and that's the the phrase I brought up before, which is variable binding, you start to see these things drop off.

And that was when I realized, oh, there's going to be a challenge for other image systems that don't have kind of the scale and the compute of like a GPT4 under the hood. And now was it just was it basically that like taking like a GP4 scale model and say now you do images that made breakthrough? Well, I think there are a lot of different parts of research that made made this such a big success, right? I think um with a complicated multi-step pipeline, it's never just one thing, right?

It's like very good post training. It's very good training. Um and um I think it's just all of that coming together, right? Um variable binding definitely was one thing that we paid a lot of attention to. I think one thing about the image launch is is a launch that was very deep. I think people, you know, they started by working on, you know, creating anime versions of themselves, but you realize when you play with it more, you know, the infographics they work with like you actually create charts, you comic book panels, you can mock up what your home would look like with furniture in it, different furniture.

Exactly. I've heard all these things from from users that are like completely surprising about the way that you we did the the the podcast setup by literally taking some photos of chairs in the room and just putting it in there and saying, "Create a better setup." And it was amazing. So, we've seen kind of a lot of the, you know, there was a lot of the anime style images which kind of like for some it was just sort of the weird thing where it was just just better than what we'd seen before and I don't think anybody is ready to be really surprised by an image model in that way.

I think obviously internally and externally. What were some of the things that surprised you or some of the new things you saw people doing? Yeah, I'll I'll tell you a quick story there, too, because um you know, up until the day of the launch, we're trying to figure out what's the right use case to showcase, you know, like uh and I think I'm so glad we ended up on kind of anime styling. It's just everyone looks good as an animated.

So, that's true. I mean, it's funny with original ChatBT, I thought it would be strictly utilitarian product and then I'm surprised that people use it for fun. In this case, it was sort of the opposite where I was like, okay, this is going to be really cool for memes. People are going to like have fun with this thing. Um, but then I was like really surprised by all the like genuinely useful ways of using image gen whether or not it's you know planning your home project um as I mentioned earlier um you know of of uh um you're doing a construction you want to see what things would look like if you know you had this remodel or this furniture or whatever um to um you're working on a slide deck um for this important um presentation and you just want to have really useful consistent illustrations um that are on topic and um and and and get it.

So, so I I really have been been kind of personally surprised by the utility in this case because I knew it would be fun. That was not a question. Yeah, I think I used it to generate a tier list of AI companies and it put opening at the top. You win model. Um, what good post training? Yeah. Yeah. it just happened you know who knew um what has been the thinking in it's changed because I remember originally with Dolly the idea of like okay we have to be a lot of very controlled about what it can do what it can't do originally I remember when we first launched you couldn't do people which was not a very useful model and then finally was trying to roll back how much of that was cultural shift how much that was the technological ability to control for things and how much of that was just saying we've got to push the norms I would say it was both cultural shift and an improvement in our ability to control things.

The cultural shift, you know, I'm I'm not going to deny it. I think when I joined OpenAI, there was um a lot of conservatism um around, you know, what capabilities we should give to users. Maybe for good reason. The technology is really new. Um a lot of us were new to working on it and you know, if you're going to have a bias, you know, biasing towards safety and being careful, it's not a bad, you know, in your DNA to have.

But I think over time we we learned that there's so many positive use cases that you um effectively prevent when you make arbitrary restrictions in the model. What about faces? Why not why can't I make any face I want? Um so this is a good example of of um a you know capability that that's got pros and cons and you can air on one side or the other. you know then when we um first shipped um image uploads um into chatgbt uh we had some debates you know about what what capabilities do you allow versus where are you conservative and I think one debate that we had is like do we upload allow the upload of images with faces or rather when you upload an image that contains a face do you um you know um should we just like gray out the face because you avoid so many problems right you can make inferences about people based on on their face um you could say mean things to people based on um their face.

Um um and and you know, you would just take a giant shortcut on all the gnarly issues if you didn't allow that. But um I've always felt we need to um air on the side of freedom and we need to do the hard work. And I think in this case, you know, there's so many valid ways. You know, if I want feedback on makeup or on my haircut or anything like that, I want to be able to talk to chat about it. That those are valuable and benign use cases.

And I would prefer to allow and then study, you know, where does that um um fall short, where is that harmful? And then iterate from there versus taking a default stance on disallow. And I think that's one of those ways in which our stance and posture has changed a bit over time in terms of where we set, you know, where we start. Yeah, we we're very good, I think, imagining worst case scenarios. What if I use this these faces to evaluate hires for a company or whatever, but also it's like, hey, is this eczema?

You know, like, you know, there's a lot of utility there. And and honestly, I think there are certain domains of of AI safety where worst case scenario thinking is very appropriate. Mhm. So I think that is an important way of thinking about risk when it comes to certain forms of risks that are existential or even just very very bad. You know, we have the preparedness framework which helps us reason through some of those things.

You know, um can the AI let you make a boweapon? It's good to think about the worst case there because it would be really really bad. So you kind of have to have that way of thinking in the company and you have to have certain topics where you think about um safety in that way but you can't let that kind of thinking spill over onto other domains of safety where the stakes are lower because you end up I think making very very um um um conservative decisions that that block out many valuable use cases.

So I think being sort of principled about different types of safety on different time horizons and with different levels of stakes is very important for us. I think I want a blunt mode sometimes and just because like right now where it actually roasts you. Well, I mean like yeah because I'll ask the model like with the the the voice in speech out model be like do I sound tired and it's like well you know I don't really want to you know and be like yeah you know just are trying to get it to be honest you know um I I I think I think there's many cultures that would prefer a blunder chat.

So very much on the radar. Yeah, just to piggyback off Mick's answer, um I think it's the iterative deployment that gives us the confidence, right? Uh to push towards user freedom and you know, we've had many cycles of this. We know what users can and can't do. Um and that gives us the confidence to launch with the restrictions that we do. One of the other capabilities, one of the other generative capabilities that's been very interesting has been code.

And I remember early on GPD3 we saw that all of a sudden it could spit out entire react components and we saw that oh wow there's some utility there and then we went we actually trained a model more specifically on code and that led to we had codeex then we had code interpreter and now codeex is somehow back and you know a new new form uh same name but the capabilities keep increasing and we've seen code work its way first into uh VS code via copilot and then uh cursor and then I windsurf which I use all the time now.

What uh how much pressure has there been in the code space? Because I'd say that if we ask people who made the top code model, we might get different answers. Yeah. And I think it reflects that when people talk about coding, uh they're talking about a lot of different things, right? I think there's coding in a specific paradigm like if you pull up an ID and you want to kind of get a completion on a on a function that's very different from you know agentic style coding where you know you ask uh you know I want I want this PR and you know um and I think we've done a lot of focus could you sorry could you unpack a little bit what you mean by agentic coding yeah yeah so I think um when you you can draw a distinction between more kind of real-time response models um you can think of chachki uh to first order a as you ask a um a prompt and then you get a response fairly fairly quickly.

And then a more agentic style model where you give it a fairly complicated task, you let it work in the background and after some amount of time it comes back to you with what it thinks is something close to the best answer, right? And I think we see increasingly that the future will look like more of a async kind of uh you know where you're asking it very difficult hard things and um you're letting the model think and reason and come back to you with really the best version of what what it can come back with and we see the evolution of code in that way too.

I think eventually we do see a world where you'll kind of give a very high level description of what you want and the model will take time and um it'll come back to you and so I think uh our our first launch codeex really um reflects that kind of paradigm where uh we are giving it PRs units of fairly heavy work um that encapsulate you know a new feature or you know a big bug fix and we want the model to spend a lot of time thinking about how to accomplish this rather than kind of give you a fast response.

And to get to your question, you know, there there's there's there's coding is such a giant space. There's so many different angles at it. It's kind of like talking about knowledge work or something incredibly broad. Uh which is why I don't think there's one winner. I don't think there's one best thing. I think there's um so many options and I think developers are the lucky ones because they have so many choices uh right now and I think that's fundamentally exciting uh for us too.

But to Mark's point, I think this agentic paradigm has been particularly exciting for us. One framing I often use when thinking about product here is I I want to build products so that have the properties such that you know the model gets 2x better product gets 2x more useful and I think you know chat has been a wonderful thing because for a long time I think that was true but I think as we look at you know smarter and smarter models I think there's some limit to people's desire to talk to like a PhD student um versus you know they might value other attributes about the model like its personality and you know what it can actually do in the real world But um experiences like codecs I think they they they they create the right body such that we can drop in you know more smarter and smarter models and it's going to be quite transformative um um because you get the interaction paradigm right where people can specify this task give the model time um and then then then get get a result back.

So I'm really excited where it's going to go. It's an early pre research preview but just like with chat GBT we felt like it would be beneficial to get feed get feedback as early as possible and um excited where we're going to take it. I was using sonnet a lot which I love. I think sonnet for coding is fantastic. But with 04 mini medium setting in wind surf I found was great. I found once I was started using that I was really happy because one the speed everything else like that and I think that and I I think there are very good reasons why people like other models and I don't want to get into comparison but I found that for me for the kinds of tasks I was using this was the first time I I was very happy you guys put that out there because absolutely yeah and um you know we feel like there's still a lot of low hanging fruit in code it is a big focus for us and I think we'll find in the near future you'll find many more good options for uh the right code model tailored for your use case Yeah, I find often if I just need a quick answer like how to write something in Dart, I'll just get a 4.

1 and say it but yeah, for something bigger and I think that's going to be the harder part is because yeah, these eval are some ways saturated but also everybody has their own criteria that we look at and that's going to be kind of a you know a question to sort of see you know how are we going to adapt to all that right yeah I mean specifically in code right I think there's more beyond did it get you the right answer with code you know people care about the style of the code.

They care about you know how verbose it was in the comments. It cares about um you know how much proactive work did the model do for you, right? Um on other functions and so I think you know there's a lot to get right and users often have very different preferences here. Yeah, it's funny. I used to I used to you know people used to ask me what domains are going to like you know be transformed by AI you know fastest I used to say you know it's code because like similar to math and other things it's very very verifiable and testable and I think those are the domains that are particularly great to do RL on and you know you're therefore going to see all this this awesome you know agentic stuff just suddenly work.

I still think that's true. But the thing that surprised me about code is that, you know, there is still so much of an element of taste in terms of what makes good code. And there's, you know, there's a reason that, you know, people train to be a professional software engineer. It's not because their IQ gets better because they, but rather because they learn, you know, um how how to build software inside an organization.

What does it mean to write good tests? What does it mean to write good um documentation? How do you respond when someone disagrees with your code? Those are all actual elements of being a real software engineer that we're going to have to teach these models um to do. Uh so I I expect progress to be fast and I still think code has a ton of nice properties that make it very ripe um for the Gentic um products. But but but I I do think it's very interesting to the degree that you know the the the element of taste and style and um real world um um software engineering matters.

It's interesting too because with chat GPT and and the other models, you're kind of dealing with having to bridge the divide between consumer and pro. I open up chat GPT and I tell my friends like, oh yeah, because I'll plug it into whatever code model I'm working because I can actually connect it to there. And I think about, you know, well, that's a very different use case a lot of other people. Although I've shown people like how to go in and use, you know, uh an IDE and actually have it just write documents for you and create folders and stuff, which people don't realize like, yeah, you could do that.

you have chatbd actually control it and do that which is cool but then you think about like okay we've got a tab now for images there's the codeex tab so if I want to connect to GitHub and have it work through there and there's a Sora into there so it's kind of interesting to see how all of these things are coalesing into there how do you differentiate between a consumer feature a professional feature and maybe like an enterprise feature look um we build very general purpose technology and it's going to be used by a whole range of folks And unlike many companies which have this kind of founding user type and then they use technology to solve that user's problems we do start often times with the technology observe who finds value in it and then iterate for them.

Now with codeex um our goal was very much to build for for professional software engineers knowing though that there's sort of a splash zone where I think a lot of other people will find value in it and we'll try to make it accessible for those people as well. Um there are a lot of opportunities to target non-engineers. I'm personally really motivated to create a world where, you know, or help help build a world where um anyone can make software.

Codex is not that product, but you could imagine those products existing over over time. Um but, you know, as a general principle, it's really hard to predict exactly who the target user is. Um until we made some of these general purpose technologies um available because it gets back to the empiricism I was talking about. Um uh we just never exactly know um where the value is going to lie. Yeah. And I think even um to to dig deeper into that zoom like you know you could have a person who's mostly using chatbu for coding right but 5% of the time you know they might just want to talk to the model or like 5% of the time they just want a really cool image right and so I think you know um there are certainly archetypes of people who who use the models but in practice we see that people want this exposure to different capabilities.

Yeah, with codeex and watching the launch of that it kind of struck me there's some tools you see that's there's a lot of excitement about because there's a lot of internal demand for that how much are you using it internally are tools like that more and more okay I' I've been really excited um to to see the internal adoption it's it's everything from you know exactly what you'd expect you know people using um codeex to awful their tests to you know we have a uh um um analyst um workflow that will look at, you know, logging errors and automatically flag them and slack people about it.

Um, so there's all these these ways that or I've actually heard heard some people are using as a to-do where like future tasks they're they're hoping to do, they're starting to fire off COD CEX tasks. So this is the perfect type of thing that I think you can you can dog food internally. Um, and and you know um I I'm very excited about, you know, the leverage that engineers are going to get out of a tool like this. I think it's going to allow us to move faster um uh with with with the people we have and make each engineer that we hire um you know like 10 times more productive.

So so in some ways internal uh usage is is a very good predictor of of where we want to take this. Yeah. I mean we don't want to ship something to other people that we don't find value in ourselves. And I think you know leading up to the launch laundry buddy Laundry Buddy Laundry Buddy is an essential partner. Okay. Sorry. Sorry. Um I mean yeah I mean we had some power users though that you know hundreds of PRs a day um that they were generating personally right so I think you know uh there are people internally finding a lot of utility from what we're building also though if you think about internal adoption it's also a good reality check because you know people are busy you know adopting new tools tools takes some activation energy so actually um the thing you find when you try to dog food things internally is is is is some of the reality component of how long it takes people to actually adjust to a new workflow.

And it's been it's been humbling to to watch, right? So, so I think you learn both about the technology, but you also learn about some of the adoption patterns when you're trying to get a bunch of busy people to change the way they write code. As you build these tools internally, people have to learn how to use them and are having to adapt. And there's a lot of question now about kind of what kind of skills do people need in the future?

You know, what kind of skills do you look for on your teams? I've thought about this a lot. Um, hiring is hard, especially if you want to have a small team that is very very good and humble and able to move fast, etc. Mhm. And I think curiosity has been the number one thing that I' I've looked for. And it's actually my advice to you students when they ask me what do I do in this world where everything's changing because I mean for us there's so much that we don't know.

There's a certain amount of humility you have to have about building on this technology. Uh because you don't know what's valuable. You don't know what's risky until you really study and go deep and and try to understand. And when it comes to working with AI, which you know, we obviously do a lot, not just in code, but in kind of every facet of of our work. Um, it's asking the right questions that is the bottleneck, not necessarily getting the answer.

So, I really fundamentally believe that we need to hire people who are deeply curious about um the world and what we do. I care a little bit less about their experience in AI. Mark presumably feels a bit different about that one, but for the product side, uh it's been curiosity that I've I've um found the most the best predictor of success. No, I mean even on research I think increasingly less uh we index on you have to have a PhD in AI, right?

I think uh this is a field that people can pick up fairly quickly. I also came into the company as a resident without much formal AI training. And I think correlated to what Nick said, I think one important thing is for our new hires to have agency, right? Open is a place where you're not going to get so much of a oh here's today you're going to do thing one, thing two, thing three. Um, it's really about being kind of driven to find, hey, here's the problem.

You know, no one else is fixing it. I'm just going to go dive in and fix it. Um, and also adaptability, right? It's a very fast changing environment. That's just the nature of the field right now. And you need to be able to quickly figure out what's important and pivot what you need to do. The agency thing is re is real. You know, I think we often get asked for, you know, how you how does open up, you know, keep shipping and, you know, you it feels like you're you're pushing something out every week or something like that.

It's a funny because it never feels to me. I always feel like, you know, we could go be going even faster. Um u but but you know, I think fundamentally we just have a lot of people with the agency who can ship. Um that comes to private, that comes to research, that comes to policy. Shipping can mean different things. uh we all do very different things at OpenAI but I think the ratio of people who can actually do things um and you know the lack of red tape except where it matters you know the couple areas where I think red tape is very very important but you know um um I I think that is what makes open very unique and it obviously affects the type of people who we want to hire too.

I was brought into the company because I was originally given access to GPD3 and I just started showing all these use cases for it and making videos every week for it. Yeah. And I was annoying people I'm sure but I was It's not. It was really fascinating. It was exciting. It was an exciting time. I I described it to people like they, you know, they I think they built a UFO and I get to play with it, you know, and then I make it hover and like, "Oh, you made it hover."

I'm like, "Well, they built it. I just pressed the button and got to do that." But that was just what I found very empowering was the fact that I I I'm self-taught. I learned to code by Udemy courses and stuff and then to be a member of the engineering staff and be told just go just go do stuff. Um, nothing too critical. I didn't break anything. anybody. Um, and that's good to know that that kind of spirit is still there.

And I think that is part of the reason why OpenA is able to ship even though, you know, it was like 150 200 people worked on GP4. I think people forget about that, you know, totally. And and honestly, this is how and even ch this is how how it came together. You know, we we had a research team, they'd been working, you know, um, for a while on instruction following and then the successor to that and, you know, post- training these models um, to be good at chat.

Uh, but the product effort came together as as a hackathon. I remember distinctly we said like who who who who who's excited to you know go build consumer products and we had all these different people like we had a guy from the supercomputing team um who you know was like I'll make an iOS app I've done that in a past life where we had you know researcher who wrote some backend code and it was just convergence of people who were excited to do stuff and I think the ability to do so and I think that's how you get the next is is is running an organization where where where that is possible and continues to be possible as a scale hackathons were my favorite thing because one being a performer and loving showand tell but it was just neat to be able to things that you knew were going to be a product or something later on because when you're playing with the technology this advanced and all that, do you guys still do them?

Yeah, absolutely. Yeah. Um, we've had some fairly recently and they are typically tied last week actually. Uh, can't say what it was about, but it was an exciting campaign. Sure. Um, and and it's it's how you find out what's possible. Yeah. Yeah. Uh, I'm excited to hear that. I I I do have a question which is how much as it grows again like when when I started I think like 150 people on the company now there's like 2,000 and now you know I see a video with Sam talking to Johnny IV and how much is that going to change the character the spirit of bringing in all this I think all the outside expertise has been great we've seen this great sort of run of products but do you see it changing the culture?

Well, I mean, I think probably in the right way, right? It's like um I think when we look at AI, we don't think of it as some fairly narrow thing. And we've always been kind of enthralled by just the potential and all the different things you could build with AI. And yeah, to to Nick's point, right, this is why we're able to ship so quickly because people imagine all these different possibilities. They imagine the future with AI and they try to bring it about, right?

And I think these are facets of that imagination, right? It's like what does AI look like? If you imagined a AI first device for instance, yeah, when you go from 200 to 2,00 you'd think a lot would change and um yeah, maybe in some ways it has but but I think people often underestimate you the number of things that we're doing. I always feel like being at open feels much closer to being in a university where you know you got this kind of common reason to being there but everyone's doing something different and you'll sit down at dinner at lunch and you'll talk to someone and learn about their thing and you're like wow that's so cool that you're doing that.

Um and so it feels much smaller because I think of the sort of broad range of things we're doing and and and therefore each individual um effort whether or not that's something like chat GPT um or something like like Sora or etc is actually staffed in a in a very very conservative and lean way that continues to keep people very autonomous and and make sure they have resources etc. So um and I think it's part partly that that has made it feel um very very similar in the good ways to to when I started here.

We talked a bit about one of the things you look for is curiosity and and Mark said that's helpful too. If I'm somebody outside of AI, okay, if I'm 25 or I'm 50 and I'm looking at the advancement of technology and maybe having a little bit of fear because I see copywriting is one of the things that Chad GPD got great at. Uh, writing code is great. I personally of the opinion that we will never have enough people creating code because there's more things code can do in the world than we can imagine.

And even the thing replaces the copy. My wife showed me the other day on her um her uh skin block or sunblock lotion bottle. Showed me on her sunblock lotion bottle like some very funny copy about like the ingredients. I said, "Oh, this is not a place I expected to see this." But that's one of the tiny little places that all of a sudden that you can put more thought into it. That being said, I know that I'm a bit of an optimist because I see all these opportunities or places to go in there.

What advice do you give people, you know, where at whatever point they are in life about preparing for or adapting to or being part of the future? You know, I like how Mark just looked right to me. You take this. I can go. Okay, I will jump in right now. Yeah. No, I think the important thing is you have to really lean into using the technology, right? And you have to see how your own capabilities can be enhanced, how you can be more productive, more effective by using the technology.

I fundamentally do think that the way this is going to evolve is you will still have your human experts but what AI helps the most is the people who don't have that capability at a very advanced level right so if you imagine right like uh as these models get much better at healthcare advice um they're going to help people who don't have access to care the most right uh image generation right it's not producing you know an alternative for you know experts or you know professional artists it's allowing people like me and Nick to create creative expressions, right?

Um, and so I think it's kind of rising the tide that allows people to be competent and effective at a lot of things all at once. And I think that's kind of how we're going to see a lot of these tools bootstrap people. The world's going to change a lot. And I think truly everyone has a moment where the AI does something that they considered sacred and human. Um um I know a guy that got vested and felt very threatened about his achievements in codeabilities.

Well, that happened for me a long time ago. Let's be talking about someone else in the room. Oh, yeah. I mean, yeah, it's definitely better than me at a lot of code problem solving for sure. Yeah. Right. So, I think it's deeply human to to to feel some level of um awe, respect, uh and maybe even fear. And I think to Mark's point, be actually using this thing can demystify it. I think we all grew up or you know learned about the word AI um in a world where AI means something pretty different from what we have today.

You've got these algorithms that you know try to sell you things, try to do things and or you've got movies you know where the eye takes over etc. And like that term means so many things to different people that I'm entirely unsurprised that you know um there's fear. So actually using the thing is is I think the best way to have a grounded conversation um about it and then I think from there the best way to prepare I I think there's some degree to which you need to understand the products and keep up sure but I think things like prompt engineering or sort of understanding the intricacies of this AI they're kind of not the right direction.

I I I think sort of there's fundamental human things like learning how to delegate that is incredibly important because increasingly you know you're going to have an intelligence in your pocket that it can be your tutor can be your um adviser it can be your software engineer um it's much more about you understanding yourself and the problems you have and how someone else might help than a specific understanding of AI.

Um so I think that's going to be important. Curiosity. I mentioned it earlier. I think asking the right questions, you'll get you only get what you put in, right? That's important. And I think fundamentally being ready to learn new things. I think the more you learn understand how to pick up new topics and and and domains, etc. Um the more you're going to be prepared for a world where, you know, the the nature of work is shifting much faster than has ever shifted before.

So, um I'm prepared that my job, you know, in product is is going to look different or not exist at all. But um I am looking forward to picking up something new and and I think as long as you you bring that perspective um you're well set up to leverage AI. Yeah, I think we we sometimes overindex on you know sometimes certain jobs go away because like you know we don't really need a lot of you know typewriter repair people anymore, right?

And then certain kinds of coding jobs are probably going to go away but like I said I think there's way more opportunity for coders or people to create code however it's done. Um, and you mentioned like the health field. And that's one of the things I hear people like, "Oh, when you know when we replace everything with AI, like well, I mean, I would be very happy having an AI diagnose me, operate on me, and probably do everything else, but I do want somebody there to talk me through the procedure and hold my hand, but also I want people asking questions like like, you know, every day I take a bunch of vitamins.

Is this the right time of day to take it?" You know, I can't bother my doctor with all these silly little questions. I I really don't think you end up displacing doctors. you and you'd end up disposing not going to the doctor. You end up democratizing the ability to get a second opinion. Very few people have that resource or know to you know take advantage of a resource like that. You end up bringing medical care into pockets of the world where that is not readily available and you end up helping doctors gain confidence.

You know I think I have often heard from doctors that you know they already talk to existing colleagues to get a second opinion. In some cases that's not possible and I think you'd be surprised by the number of doctors that use chatbt. Um, now on things like medicine, there's work to make the model really, really good, and we're excited to do that work. There's also work to prove that the model is really good because I think you're not going to trust it until there's some degree of sort of legitimacy.

And then there's work to explain the areas where the model might not be good because increasingly once it gets to human and then super human level performances, um, it's hard to frame exactly where it will fall short, which is also um, hard hard to sort of reckon with. But nonetheless, I think that opportunity is one of the things that gets me up in the morning. education might be the other one and um I think there's a tremendous opportunity to help people.

What do you think is going to surprise us the most in the next year to 18 months? I honestly think um it's going to be the amount of research results that are powered even in some small way by the models that we've built. And um one of the kind of quiet things that's taken the field by storm is the ability of the models to reason. And you already see some research paper. I'm gonna make you explain when you say reason.

Yeah. So, this fits into the I want you to reason through Yeah. the question as you explain reason out loud. Yeah. Think out loud. Tell us your your your traces. Yeah. This um this really fits into this egentic paradigm that we were talking about earlier. And um the way that the models approach solving a problem that takes some time to solve is that it reasons through it. Much like you or I might, right? If I give you a very complicated I think you reason probably much better than I do.

No, I mean um I think this with a yeah like a a complicated puzzle, right? You might think to yourself, uh for instance, let's just use a cross word puzzle, right? Like you might think through all the different alternatives and um what's consistent? Um you know, is is this row kind of consistent with that column and you're searching through a lot of alternatives. You're backtracking a lot. Um you're trying out a lot of different hypotheses and and then at the end, right, you come up with a well-formed answer.

And so the models are getting a lot better at that and that's what's powering a lot of the advancements in math, in science, in coding. So this has reached a level where today in many research papers people are using 03 almost as a sub routine, right? There's sub problems within the research problems they're trying to solve which are just fully automated and solved through plugging into a model like 03. Um I've seen this in several physics papers.

um talked to physicists even where they're like wow like I had this expression that I couldn't simplify but 03 made headway on it and and these are coming from some of the best physicists in the country so um I think you're going to see that happen more and more and more and more and we're going to see just acceleration in in progress in fields like physics and mathematics. It's a hard one to beat because, you know, I I I would swap many things we do in exchange for making a true, you know, significant, you know, scientific advancement.

Um, but I think we can we we we we can have multiple of these things. I I think for for me it's it's the fact that any wellscribed problem that is intelligence constrained I think will be solved in in products. And I think we're we're fundamentally just limited by our ability to do that. So what that mean is like you know in companies in the enterprise there are so many problems that are fundamentally hard the models are not smart enough to do yet.

Um whether or not it's software engineering, whether it's not running data analyses, whether or not it is um providing amazing customer support, there's all these problems that um the models fall short at today that are very very um easy to describe and evaluate and I think we'll make tremendous progress at those. Um, on the consumer side, these problems exist, too. They're a bit harder to find um just because consumers are um um um worse at um telling us exactly what they want.

That's the nature of building consumer products. But I think it's very very worthwhile where you know there's many hard things we do in our personal life. Whether or not it's doing taxes, whether or not it's planning a trip, whether or not it's um searching for a high consideration purchase, whether or not that's a house or a car or a piece of clothes. Um all of those things are um problems where we need just a little bit more intelligence.

um and the right form factor. So I think the other thing that's going to happen in the next year and a half is you'll see a different form factor in AI um evolve. I think chat is still incredibly useful interaction model and I don't think it's going to go away but increasingly you're going to see more of these sort of asynchronous um workflows. Coding is just one example but for consumers it might be sending this thing off to go find you the perfect pair of shoes or to go you know plan a trip or for to go um um finish your taxes.

And I think that's going to be exciting and we're going to think of AI a little bit differently than um just a chatbot. One of my favorite examples both from a utility point of view capability and then uh UI was deep research and deep research is probably the best example we maybe have of probably agentic sort of model use right now because it used to be you would ask for a model to tell you about a topic. It would you would either get the data or just do a big search the internet and then it would just summarize all of that where deep research will go find some set of data look at it ask a question then go find some new data and come back to it and keep going on and I think the first time I used other people used like wow this is taking a while and then you added a UI change so I can actually go away and go do something else and then the lock screen on my phone will show me this is working which was a paradigm shift and I talked to Sam here about that and Sam said that was a surprise to him was the fact that people would be willing to wait for answers and now I've seen uh a new metric for models is how long a model can spend trying to solve a problem which is a good metric if it ultimately solves it and that's has this been an update to you and how you think about these things the idea of like oh we don't just want and I guess you talked about this before about agentic and the idea that it's not just give me the answer it's like take your time get back to me I think you know to build a super assistant, you got to relax constraints.

Like today, you have a product that is, you know, entirely synchronous. You have to initiate everything. Um that's just not the maximally best way to help people. Like if you think about a real world um intelligence that you might get to work with. Um it has to be able to go off and do things over a long period of time. It has to be able to be proactive. Um so I think there's like we're we're sort of in this process of relaxing a lot of the constraints on the product and on the technology to better mimic a very very helpful um entity.

Um the ability to go do five minute tasks, you know, five hour tasks, eventually five day tasks is like a very very fundamental thing that I think is going to unlock a different degree of value in in the product. So I've actually not been that surprised that people are willing to do that. Um like I I don't really want to be sitting around um waiting for my coworker either. Um, and I think if the value is there, um, I' I'd gladly be doing other stuff and come back.

Yeah. And we really don't do it just because, right, we do it out of necessity. The model needs that time to solve the really hard coding problem or the really hard math problem. And it's not going to do it with less time, right? You can think about this as I give you some kind of brain teaser, right? Your quick answer is probably like the intuitive wrong one. And you need that actual time to kind of work through all the cases to like are there any gotchas here?

Um, and I think it's that kind of stuff that ultimately makes robust agents. We we've seen kind of there's like the the paper of the moment where somebody comes out and says, "Ah, I found a a blocker." And I remember there was one a month or so ago and they said models couldn't solve certain kinds of problems. And it wasn't hard to figure out a prompt that you could train into a model and it could solve those kinds of problems.

And we had a new one that talked about how they would fail at certain kinds of problem solving ones. And that was kind of quickly, I think, debunked by showing that, you know, the paper kind of had flaws in there. There are limitations. There are things that there might be some blockers and thing or things we don't know are going to be there. I think brittleleness is one of the things. There is a point where models can only spend so much time solving a problem.

We're probably at a point we're only having the model, you know, maybe two systems watch each other and we have to think about how a third system stops, you know, to wait for things to break down. But do you see kind of any blockers between here and where I'm getting the models that are going to be solving, you know, doing things like coming up with interesting scientific discoveries? I mean, I think they're always technical innovations that we're trying to come come up with, right?

Um, fundamentally, we're in the business of producing simple research ideas that scale and the mechanics of actually getting that to scale are are difficult, right? It's a lot of engineering, a lot of research to kind of figure out um how to kind of tweak past a a certain roadblock. And I think um those are always going to exist, right? Every layer of scale gives you new challenges and new opportunities. So um you know, fundamentally the approach is the same, but we're always encountering new small challenges that we have to overcome, right?

Just to build on that, I mean, the other business we're in isn't building um great product with with with great product with these models. And I think we shouldn't underestimate the um challenge and amount of discovery needed to really bring these um ever intelligent models into the right environment whether or not that's giving them the right sort of action space and tools whether or not that's really being proximate to the problems that are hardest understanding those and bringing the eye there.

So I I think there's, you know, the the technical answer. Um um but I think there's also the the the you know real world deployment and I think that always has challenges that are like very very hard to predict yet you know worth worthwhile and part of our mission to solve. All right, last question and I'll begin. It's uh what's your favorite use or tip for chat GPT? Mine is I take a photograph of a menu and I'm like help me plan a meal or whatever if I'm trying to like you know stick to a diet or whatever.

See, I really want that use case, but like I've been trying it for wine lists and that is my eval on multiodality. It still doesn't work. Like really, it keeps embarrassing me with like hallucinated wine recommendations and I go order it and they're like, "Never heard of this one." So, I'm glad yours works. But for me, that's the that's still a use case. Well, I mean, maybe the wine lens is too dense. That was a problem.

That was a problem with operator was it like originally was the the vision models that too much dense text it just loses it placement. Yeah. I mean, speaking to deep research, I love using deep research. And, you know, when I go meet someone new, um, when I'm going to talk to someone about AI, right, I just pre-flight topics, right? I I think the model can do a really good job of contextualizing who I am, who I'm about to meet, and what things we might find interesting.

Um, and I think it it really just helps with that whole process. Very cool. I'm a voice believer. I I it's still got um I I don't think it's entirely mainstream yet because it's got it's got many little kinks that all add up, but for me, you know, half of the value of voice is actually just having someone to talk to and forcing yourself to articulate um um yourself. And I I find that to sometimes be very difficult to do in writing.

So on my way to work, I'll use it to process my own thoughts. And like with some luck, and I think this works most days, I'll have sort of a structured list of to-dos by the time I actually get there. So, uh, voice for me continues to be the thing that, you know, I both love using and want to see improve, um, over the next year.