🎙️AI 访谈库
Codex 团队如何用自己的编码智能体——Sottiaux 做客 Every《AI & I》
Thibault Sottiaux · OpenAI Codex 负责人(与 Codex App 工程师 Andrew Ambrosino 同台)

Codex 团队如何用自己的编码智能体——Sottiaux 做客 Every《AI & I》

OpenAI's Codex: This Model Is So Fast It Changes How You Code

2026-02-18 · AI & I (Dan Shipper, Every) · 47m · 约 49 分钟读完 · 原文
超级碗广告播出后'4 点刚过系统即遭巨大流量冲击'的第一手复盘:为何在终端之外做 Codex 桌面 App——CLI 在多模态与并行智能体场景'上限太低'、App 是运行中智能体的'指挥中心'且上线首周下载破百万,以及智能体已在代码之外接管 Linear、Slack 等工作流。

The first time I showed it to someone, they were like, "No way. This is like a fake demo. Like, this cannot be this fast. This will change everything." Especially because it's not yet the fastest that we can actually get it to be. >> My experience was trying the app. I didn't really want to go back to a terminal. What I realized is actually gooeys are great. IDEIDes are just the problem. There's something that's a guey for programming that's not an IDE.

And it it seems like you're figuring that out, but I don't even know what that's called. >> It's called a codex. >> Dan here and I want to take a second away from the episode to tell you about Granola. Granola is an AI noteaker for your meetings and I use it pretty much every day. That may sound a little bit weird or a little bit creepy like transcribe all your meetings. Well, for me, it's actually kind of indispensable as a leader.

Every is about 20 people now, and it's really important to me that I understand how decisions get made, how I'm showing up in meetings, and how I can help my team the best way I can. Granola acts a little bit like a leadership log for me, so I can see how I've done in meetings, what situations came up in a particular week, and how I can do better next time. If you're trying to improve as a leader and scale your company, try Granola as your AI powered notepad for meetings.

Head to granola.ai/every, code every to get three months free. And now back to the episode. Tibo Andrew, welcome to the show. >> Hey, thanks for having us. >> Thanks for having us. >> Great, great to get to chat with you. So for people who don't know, Tibo, you are the head of Codex OpenAI and Andrew, you are a member of the technical staff um on the Codeex app at OpenAI and uh you are the people of the moment.

They just ran a Super Bowl commercial about Codeex uh OpenAI did. How are you feeling? >> Yeah, that that Super Bowl ad was quite surprising, wasn't it? It it really was. I think the the the core thing and I think the way the reason the place I want to start this conversation is um it feels like that that is a strategic shift. You would expect OpenAI to have run a ChatgBT commercial during the Super Bowl and maybe not especially if you looked at Codeex's positioning like three or four months ago for professional engineers, maybe not have run um a an ad targeted at a much more at a much broader audience.

Um it it felt like for a long time there was this divide where codeex was for professional engineers and if you wanted to do vibe coding you do that in the chatbt app and it seems like that that has shifted a lot um over the last like month or two. Can you tell me about that? Yeah, I think especially like in you know we can talk about last week right so like last week on Monday we released the codex app uh immediately we saw like a ton of downloads like more than a million downloads in the first week and then we knew that we were releasing like an extremely strong model like you know 53 codecs on Thursday that just made I think this it very visible that you know we're here to you know put incredible experiences out there we're very committed to codeex and like also agents are really starting to work and be able to create these things, you know, even if you're like a little bit less technical.

Uh, I think like the app really showed that, you know, it's like it much more inviting for people to just try it and like know run multiple agents, you know, with our models being like very um very good at sort of like allowing for multitasking and being reliable for longunning uh longunning sessions. So, like allows you to create a lot more. So, it just felt that, you know, maybe we can inspire more people to build and then show that agents are here, right?

It's like it's not it's it's coming. It's going to be mainstream. Um you know why don't you try and like create something new and you know inspire people. I felt like the right thing that we wanted to reinforce. >> Yeah. While we were designing and developing the app, one of one of our like internal mandates to ourselves the whole time was that we had to make something that we love to use and that we used for like all of our work.

Um and if we couldn't do that, then we weren't going to put this out. And this was back when we started. And I think that we surprised ourselves a lot with how fun it was. And especially as you know, we started to build this app um before we started to build agent skills and then once we kind of paired them together, it became this really rich interactive experience where you could open the browser or you could connect to these various services.

And so all of a sudden we started to feel this like really connected interactive experience and um wanted to share like I I kind of see the the ad as like a a love letter to builders, right? I have never seen a Linux CD in a Super Bowl ad and so you know like that was really cool to watch. >> What what was the impact of the ad? >> Um we're still to measure that. uh we'll see like you know how how it uh plays out over the long term but we saw a giant surge of traffic actually like remarkably like you know very very quickly after 400 p.

m. like PST when it aired like the surge and like our systems were like under heavy load. So it was it felt kind of weird to me like you know people are watching the Super Bowl and then going and like you know installing the app and they just like trying it out right there and then but that it happened. Um, and a lot of a lot of people reached out and like saying they were really inspired u by by by it and just just wanted to build afterwards, which is, you know, what we're like aiming for.

So, >> Tammy back, I I still want to talk a little bit about the the strategic shift. So um codec app moving from or codecs in general moving from something that is really for professional developers to moving to something that has a more a broader audience um and and and maybe moving some of the vibe coding from chat into the codeex app. Tell me about that. >> I I don't think we're trying to move uh vibe coding you know from from Chad into the the codex app.

We're very much you know two things are happening like one we're pushing the frontier on like professional software development like 53 codeex like beats every single other model like on the top benchmarks for coding so it is a very very capable model um and you know it's also like at the at the speed and cost it's like you know it's it is a top performer rather um I think the app the second thing is like the app does make things more accessible and so like it does appeal to like a wider audience but internally We're also seeing the app, you know, just it is very much used within research, within our own team, like the entire codeex team uses the app.

It makes people more productive. So, it's like very much leaning in into, you know, how we think agents are best used, the patterns that we were seeing, you know, that were making people like very uh productive here at the company and outside and then just sort of like going all in on that. It does happen that at the same time also it's like, hey, it's just delegation is finally here. it works, you know, and it's like much more accessible and we're gonna try and see like how we can package that and actually ship this to like, you know, much much wider audience, but that's that might not be the codeex app.

>> I mean, you use that all the It's like you just build in there. >> 99% of the code that I write is using the Codex app. >> Same. I mean, I live in there now. >> Yeah. >> Um Okay. Well, that's that's actually really interesting. I I definitely want to talk about the app in particular, but I want to go back to to the thing you just said, which is um maybe if I if I'm reading you right, you're you're kind of like we're pushing the frontier.

We're seeing lots of people who are um maybe broader than just like senior engineers using this. However, the overall idea of like who is doing what in which app like maybe you haven't totally figured out yet and it's it's not as clean of a line as like no longer vibe coding in CHBT or really vibe coding in CEX. like you can do it in both but we haven't figured out exactly like which thing you're going to do where >> yeah I think credex is like the most powerful experience right out there so you should be fairly technical so that you understand like hey you know code is actually getting written you know it's going to get executed on your machine by default it's executed in the sandbox um but you should probably be able to read code uh in order to use you know codeex to like its its fullest we will bring a similar experience to GGBT at some point which will have like different uh properties in terms of like the sandbox and like how concepts are represented.

Maybe we won't be showing like you know hey this like this scary terminal command like thing is like running and you know you should probably approve it. It's like you know of course you shouldn't do that to someone who is not technical and Codex is really there to like appeal to you know just all coders builders you know technical like people who are close like either technical themselves or like technical adjacent you know like data science these kinds of things.

Yeah. And you know, if you use the Codex app for any amount of time, you can see the inspirations from chat. You know, the layout's very similar. We autoame your your conversations. We've got contextual actions, but it's pretty clean, right? The composer looks very similar. Um, and you'll see some of that inspiration back in chat for for other types of things. Um but we we still believe that uh some you know when we set out to make something that was for the professional software developer and for us that it deserved a dedicated experience that could could really showcase the power of the models and the way that the models could change the development life cycle.

And so we you know we made something very tailored to that. Um and we've we've had a lot of success internally with research teams with product teams. Um, and so, you know, we're we'll look beyond, but I I think we're really happy with where we've ended up on that kind of tailored the tailored approach to this. >> Can you tell me about the decision to invest in a guey over a 2y? I feel like twoies are are so hot right now.

And obviously you have one for codeex already and you you could have said okay we're going to double down and and just make the make the terminal terminal experience even better than it is now and really invest in that versus okay we're going to go you know I think uh yeah making a gooey is a little bit of like a counterintuitive or uh uh like counternarrative thing to do. So tell me about that decision process. >> I think it wasn't counterintuitive.

it's more maybe it's not mainstream right and so we we experiment with a lot of different approaches like I very much consider that we're still in the experimentation phase uh and you know we're responsible primarily for two things is like you know building the most powerful entity out there you know that's capable of coding and then you know increasingly this will become like a multi- aent system and it will become like more and more capable and you know you will have to figure out like how to steer and supervise like its outcome and its behavior you know that's like one thing that we're building.

And then we're also building like how do you even you know interact with this? It's like you know what is the optimal way to have visibility into what this like very capable entity or like system of entities is doing. How do you steer them? How do you supervise them? And so we we're very much still experimenting with what that is. It's like you know sure you can do it in the two. It's like at some point it starts to feel like very limiting you know especially on like multimodal like you know actually like the models can like draw little diagrams and generate images and you know or you know you can talk over it you know using voice um maybe you have like many of them going in parallel and so you start to lose track.

So we felt like we needed to start experimenting with something else and it is only when you know we saw it become like super super popular internally we were like we have to ship this externally like this is kind of like this has come to a point where it's like too good to sort of like just keep it to ourselves. Um I mean that was like the journey that you went you know you were not building in the app although like when did you start building in the app?

That was actually like fairly quickly like when the app was building itself >> that that was yeah that was pretty quickly and yeah because I was starting with the TUI and with the the IDE extension and I think that my goal personally was how can I get to fully building the app on the app as fast as possible >> right it's like it's really easy when building this stuff to slip into the mode of like oh this will be good for somebody like somebody will love this a certain type of like they will love this right so we really wanted to get quickly to like I want to be able to build the app on the app.

I want it to be able to run itself with skills. I want it to click around on the app that it spawned. Um, and I I want this to be like part of my workflow as soon as possible. and they're um like I I still use the TUI sometimes when I want to fire something quick, but I think that like there is something about the flexibility of controlling UI and being able to have some pains be persistent and others be ephemeral um and be you know we we shipped voice with the app um so you can prompt with with voice um we have mermaid diagrams in the app we have full image rendering so all those things I think are like the tip of the iceberg on what we want to do with a dedicated UI.

Um, and it's it's pretty simple and it's simple intentionally, but I think we're going to do a lot with um, dynamic stuff there. >> I mean, yeah, the the ceiling is just much higher. >> Yeah, it's interesting. Um, my experience was trying the app, trying the app. I didn't really want to go back to a terminal and I had been coding most mostly in cloud code and and some and some codecs in the terminal for the last like for several months before that and I think what I realized is um actually guies are great ideides are just the problem.

Um and like there's some there's something that's a guey for programming that's not an IDE and it it seems like you're kind of in that figuring that out but I don't even know what that's called. It's called a Codex app. >> You know, there there there was a moment during the development of this where everybody and their mother was forking uh the same IDE and we um we kind of looked at each other and we were like, "Hey, should we have done a fork of VS Code as well?"

Like it like very seriously. I I remember exactly which day it was. And I think I don't know if I don't know if I would say that idees are the problem, but I go back to like the truck analogy sometimes with them, which is that like I will open an IDE here and there. Like I opened one today. Um it was something very specific that I wanted to do that um I don't even remember what it was, but then I closed it and I went back to using the Codex app.

And I think that there there is something there with like the codeex app being a great um daily driver and like occasionally you need an IDE or occasionally you need like a really complex terminal setup but that this should be your home base. It should be your command center for the agents that are running and a place that you can come back to and track all this stuff. And you know there were a lot of design decisions around like do we allow free form panels like an IDE.

And we kind of came to the conclusion that a lot of what these models are great at is knowing what is needed in the moment for what type of task and and so we wanted to have kind of more full control over what was able to show at what point, right? And you can see that in plan mode where you're not necessarily getting a composer. You're getting a really quick way to answer questions. Um you can, you know, and you've got your plan and you can edit your plan and um I think we only want to do more with that as we go.

>> It seems like you were surprised that you didn't want to go back to the the after. >> I was. >> Yeah. Is that >> were you like a a like like Greg Greg did an interview and Greg was like I am a twoe power user. I thought I would never leave the terminal. Like >> Greg lives in Emacs. >> Are you like a >> I was a 2 power user for like six months starting with starting with like when cloud code first got really good and I was like holy [ __ ] this is so much better than being in cursor when surf or whatever and now I feel like we I speed I speed ran my two era and I'm back in back in gooies.

like I'm I'm kind of flipping back and forth right now, but I can I sort of see the light where um it just if you're especially if you have a bunch of them going at once, the affordances of guey are just like make it much nicer. >> Yeah. And there's a lot more to come there. And it's it was a very intentional thing for us like we sort of see, you know, agents will act and you know are already acting on like much more than code, right?

And so they need to be a companion you know to like every single app and every single thing that you can do on your computer. Um just like we integrate with like linear slack and you know of course you know also need to be able to like you know read the code and like produce code but maybe you know it can do like a deploy as well like are you going to do all these things from your IDE that would sort of like feel very odd.

Um and so it's like it's like this command center for your agent. We optimize the entire uh experience around that you know around the idea that you have a very capable intelligent entity that you're like controlling steering and supervising and you know you you never need to like sort of like go in there and like you know do the things yourself it's like you know the thing is very capable of like you know being delegated to like I think you know when when you accept that that is like you know what we're headed towards and like you know with five codex is like you know just feels like you know we're getting you're like almost there right Then you're like, well, you know, it's the same with with you, right?

You know, like when I talk to you about like a feature idea or something, it's just like, you know, you just, you know, you go and you get inspired and you go and do it. You're just like, you know, I don't suddenly jump into your ID and like, you know, just go and like implement it. >> You could. >> Yeah. I mean, I think you would find it disturbing, right? It's like I mean, so that's the way that you will, you know, everyone will work with agents.

It's like you just talk to them. How has your your workflow changed with 53 codeex versus 52? >> I was surprised at how much faster it was. Um and sort of like I had to adjust on I had been optimizing a lot more for like long running sort of like multitasking. Um and you know I sort of like had an expectation of like okay this type of task will take like you know 10 15 minutes. I'm going to kick like you know four like you know different things and then come back.

So I'm able to like you know maybe do a little bit less multitasking and like you know be more in the flow. So that you know felt really good [snorts] and then it just feels now very satisfying as well like you know to kick off like automations with it using skills. It's like it's it's a more generally capable model. it's like less sort of like super focused on code, right? And so I find it like much more reliable at like you know sort of like going through like Twitter replies and like you know summarizing like the important teams or like filing bugs in like linear and then you know coming back to that and using automation so that you know things are like implemented like daily.

Um feels it's like much more robust for these things. Um, I mean, but you're really like the superpower user here, Andre. Like, you know, it's just like the [snorts] kind of stuff like, you know, he does is just like, you know, it's like I have very vanilla usage of codeex compared to Andrew. >> Uh, [snorts] no, I mean, well said. Um, I had a a series that I I had intentions to run this for a while and I only ran it for three days on on X Twitter.

uh which was that I was I was setting up a prompt to basically add a feature to the Codex app like some random like non-shippable feature to the Codex app. I had this long prompts like about the the quality bar that we had to do and um once I switched it to 53 CEX the results got actually much more interesting. Like we did a uh subway surfers panel on the right was one of them. Um like a little Minecraft UI for the sub aents was another one that we did that I don't know maybe maybe we'll ship it.

>> I was like get back to work. >> Yeah. Yeah. Um >> why do we have Minecraft in the codex out now? >> Yeah. But uh got to explore. No, I mean 53 codex like it's it's it's neat, it's fast, it's capable, it's multimodel. Um, >> what are TBO says you have a lot of cool use cases like what are what are the like more interesting ways that you're using the CEX app that maybe people should try but haven't thought of yet.

>> Andrew came up with automations and I think that sort of like shifts the way that you know you're thinking about these things when they can just like sort of like hop it in the background you know on a specific trigger at a specific time >> and then you know just you can sort of like program it yourself. >> Yeah. you're using that a lot. >> There are a lot of things that I use um the app for that are a little bit outside of just like coding features.

Um I keep it to uh I I use it to keep my PRs mergeable with automations and so it'll resolve merge conflicts. It'll keep them updated. It will fix like build issues so that basically like as soon as they're ready to go like they're ready to go. There's no like oh hey somebody merged a big thing and there's a conflict now. Um so I do that. So you said like so at what point is the is the automation trigger because I thought the automation triggers like at a certain time schedule but it sounds like there other triggers I didn't know about.

>> Um I yeah I we're looking a lot of things. I have it right now just on a time schedule and I use our GitHub skill and um some internal skills for our CI and that that runs hourly or every two hours and kind of just cleans everything up. >> I see. So it's like it just looks through all you know there are any changes on main and it just looks through any PRs and just like make sure that they're all up to date so that whenever you're ready to go it's never like that's actually that's good.

I like that. >> Yeah, it's it's actually really helpful. It's surprisingly helpful. Um I have one that every day at like 9:00 a.m. uh I get sent all of the contributions that have merged to the Codex app over the last day. And so it'll do like a nice report of who merged what and it will um I have it group it by theme so I can be like all right like three people worked on this part of the the composer, two people worked on automations like here's what happens so that I can at least be like knowledgeable what's happening because things things get chaotic uh uh right before launch and yeah >> one automation I have is uh it's I run it like multiple times a day and it's like pick a random file and uh find and fix like a subtle bug and then it's kind of funny because it actually does pick a random file.

Uh so it will run like you know Python like rand and then it will like you know find a random file and it will start from there and so it's like every time like explores like a new one um >> has it caught anything? >> Oh yeah yeah like we we catch like it's often laten bugs that you know are not triggering actually like on the critical path but you know they're actually bugs. uh and then you know just it's like trivial to fix it like merge it um takes very little of time and it's a thing that you know I would have never found myself uh found like an issue in like constraint sampling like the other day.

Yeah, >> that's really cool. Do you have other other automations that are worth sharing? >> Let's see. I feel like I have 60 that are running at all times >> like some for testing and some for real. Um, some of the members on the team really like this one that looks at the PRs that you've done in past day or so and quietly cleans up any bugs you shipped. Um, and kind of like looks at a few of the observability platforms to see and and like tries to basically ship a fix before anyone's noticed that you shipped a bug.

>> That's cool. >> And one that's not coding related, which is like marketing research. It runs daily and then it's just sort of like it's uh prompted with like a specific skill to do like deep marketing research um which have like sort of like tuned over time and then that uh just goes and like searches the web on you know any sort of like new things that sort of like came up in terms of like how you know just like how users are like perceiving talking about uh CEX then I just receive that little report.

Um and it always makes for like an interesting read. Um yeah we can just go on. And it's like these are just examples that you know we do rely on like you know they run. Yeah. >> Yeah. [snorts] Yeah. Do you have any particular skills that you guys like that are beyond the normal kind of you know I have a GitHub skill and that kind of stuff. I love Andrew's uh yeet ye skill which um it just like takes like the change and then you know does the commit does the PR writes like the draft um puts it in draft and like you know publishes a PR with like a PR title and body.

>> Yeah, it's very satisfying. >> Yeah, it just does everything. Um that one is like makes definitely makes people like productive. What are the top used ones for you? Um, image gen is a cool one. >> Yeah. >> Um, for both like silly automation purposes like, "Hey, make me an image that characterizes my last day of work." Um, not my last day of work, my previous day. [clears throat and laughter] >> Um, yes, Andrew.

>> I, you know, the, um, the image gen skill was actually really cool, uh, for I I used the CEX app to make a book for my daughters. Um, and so I had I like, you know, put together this prompt for teaching it about like a script that I wanted written. Said like 24 pages. Here are my daughter's ages. Here's like where we've lived in the past. Like we were in Boston and moved to New York and then moved over here.

Um, and then I said like after that, we went through that. I agreed on the script and then we went through and I I said like all right now it's time to use the image gen skill and it made um like it prompted for every page in the book based on the script it prompted for the image and then it kind of put them all together and use the PDF skill to put together the book's PDF and then I printed it um and so we've got like a super custom book that you know I read to my kids and it's it's really cool.

It's just this awesome thing when you can combine like the intelligence of like the agent and then it's like works in a programmatic way like you know by [snorts] using skills and then you can just combine them in like novel ways and like yeah I think the the PDF and image gen one is like a it's a common combo that we see. It feels like the codeex model it obviously it's gotten faster which makes it much more usable and and it also feels a little more obesey like it's a little more has a little more emotional intelligence but it still has a little bit of that like it does exactly what you say thing in a way that is a little can be annoying.

How are you guys thinking about how you shape the way the model feels and which way you're pushing it? >> It's something that we obsess over. So we we definitely want the model to excel at coding and be really good at instruction following. At the same time when we optimize a little bit too much in that direction it can like overindex on like you know specific words or you know sort of like misunderstand the intent you know in ways that you know humans wouldn't.

Um, like sometimes I will just like [snorts] have a typo and then you know the typo like you know actually find its way into like the file and I'm like you know obviously you know I didn't mean like you know the typo is like you know I meant like this name of this class. Um so that's something that we're um you know definitely continuing to push on. But like the thing that we're pushing on the most right now is like really efficiency you know speed and then also like what we now refer to as like personalities like you know how supportive is it?

And then we understand that not everybody has the same preferences there. Like the previous default, you know, was definitely like super blunt like pragmatic personality. Now we've also introduced like a more supportive like friendly personality and you can just like pick between those. And I think for things that don't have like sort of like a universal like um accepted you know thing that you know everybody that you know should just use is like you know we're probably going to introduce like some way for you to just make it your own, right?

You know, you should feel like you have your own little personal codeex um that you know works in exactly the way that you want it to work. Um do you use the friendly or the pragmatic one? >> Pragmatic. Yeah. Okay. I'll say use pragmatic. >> Yeah. >> Um [clears throat] interesting. I think um you guys recently put out a model that is so [ __ ] fast. I was testing it before it came out and I was just like uh I c I can't really keep up with this thing.

So I'm curious how that changes how you think about um what is now possible with coding with a model like this and also the affordances that you need in order to manage models that are so quick effectively. >> Yeah. The fir Yeah. The first time we used this model in the app um we had kind of that same thing happen where all of a sudden there was just like this wall of text and we were at the bottom of the scroll and we were immediately like all right we need to smooth this thing out.

uh coming in. And so we actually do slow it down ever so slightly just so that you can see the words come in like a little bit smoother. >> So funny. >> It's it's like a really funny problem. Um but this thing has been super fun. And I think I think what I'm most excited about is what sort of capabilities we can start to add to the app that are really really dynamic that we couldn't with a model that wasn't this fast.

So yes, this model is going to allow you to iterate really really quickly, but it also opens up a lot of new opportunities to how like how you code and and how you interact with the codeex app. Yeah, the the first time I I showed like the the the very first prototype uh when we hooked everything up and like you know obviously like the model is like powered by uh Cerebras and it's like you know we've talked about the the the partnership there and like we're very excited to put like you know the first model that we're serving through that um you know out there like it's it's you know obviously like uh still like very early.

It's like literally the first time we hook it all up and uh we're just like so excited that we want to share it. But the first time I I showed it to someone um they were like, "No way. This is like a fake a fake demo." It's like, you know, this is not real. Like this cannot be this fast. And then they tried like a few prompts that were just like, [snorts] "Oh, I can just I literally cannot keep up." It's like this is insane.

Um, and yeah, I think this will change this will change everything, especially because it's not it's not yet the fastest that we can actually get it to be. Um, like with the preview, we're putting it out like, you know, quite early. We're actually going to layer a number of optimizations on top of it. Uh, which should be able to like make it, you know, maybe two to threex faster than the experience that you have experienced.

[snorts] So, that's going to change things. And we're thinking about this also from a point of view of like delegation you know like we think this model has a huge role to play um as part of like a system of like you know multi- aent systems and as a way to like speed up you know maybe the the slower more intelligent agent as well. Uh so we're going to be experimenting that uh in that way. >> H and do you do you expect the same hardware speedups on like the more intelligent agents to come out soon?

Um so we a lot of the things that we worked on were interesting sort of like distributed systems and like infra problems that we uncovered because the model we were able to sample from the model at like unprecedented like unprecedented speeds, right? And then if you're if you're getting tokens back this fast, you need to go and like optimize, you know, the entire set of bottlenecks that you sort of like uncover on the critical path of serving.

All of those benefit you know like the current uh they benefit like GPT53 codecs and like you know all future models and there's one thing that we've been doing as well which we're I'm sure we're going to put in like a more detailed blog post at some point which is we rewrote the entire service stack to be based on like websockets and like a persistent connection and to do things like a lot more incrementally and like statefully and that decreases like the overall latency like you know across all models like we haven't chipped it like by default yet but it's you It it is something that you know we are making the default for uh this new like super fast model and then we're also going to enable like on the other models and like it makes things decreases like overall turn latency by like something like 30 40%.

Um I can we can look into the exact numbers like >> Yeah. What are the most surprising things that you've seen using the model internally in terms of like what a what a speed speed up like this enables? >> It just allows you to be super super in the flow. Um and you know just you're almost like just in real time you know sculpting the experience or like the code. Uh it's just a very different feel to it.

um it's it's very unsettling at first and then you know and then once you get into it it's very hard to go back to any other model. Uh that's like the feedback that we've seen that's like what I have felt myself. Uh and so it's it's like this very it takes like five minutes to adapt and then and then you sort of like know okay it's like this is how I'm going to use this thing. >> Yeah. I also don't think that we've poked at the full extent of what we could do with it.

Yeah. Like it's very early. We haven't had it for very long. >> Yeah. Someone on the team like Channing was just showing like, oh yeah, it's so fast and it can actually like play Pong, you know, not very well, but it's like the model is able to react to things like, you know, almost like real time, right? >> It it like you start to see how it might replace some deterministic steps. So we have we have in the codeex app a set of git actions, right?

And as everybody knows with Git, like certain configuration of things or certain states that you can be in can make it really hard to run those without a ton of error handling and like all sorts of like error messages and guidance and it's really hard to create a good Git experience, which is why like nobody ever has. Um, but if you have a model that's as almost as fast as running these scripts, then you can imagine a world where these things turn into skills or something like that.

And you can have your operations run a little bit differently with some like some intelligence and and not have the same latency that you have today when you're asking it to go track something down the code base, right? You can kind of like vaguely gesture and be like, hey, like send this up and and have that be fast enough for a for a button. What I'm very excited about is like when it's going to come together with, you know, one thing that we shipped with 53 Codex as well is like this thing that we call like midterm steering, you know, where you're you're just you start with your prompt.

It's like it it got to work and then you send another prompt like while it's still working and it adapts like in real time as well. Like it will just sort of like receive that message, acknowledge it and then you know continue its work. Like if you start to think about okay what would this look like with voice and then with a model that is as fast as the one that we just shipped um then that's like a whole other experience that you know we would be very excited to bring um you know hopefully very quickly >> because you can easily interrupt as you're talking.

>> Yeah. If you're just talking and engaging with like you know lateral language and then doing the midter steers and then the you know the implementation happens like almost instantly because of the speed is like it becomes like a very uh pleasant thing to use like right now you can sort of emulate it you know with like voice uh voice dictation and then send it and mturn steering and then you know watch the model implement and it's like a very cool thing.

I think we're going to have a step change in that experience when we just like really just polish it. >> If speed as a bottleneck is like close to being solved, what do you think is the next bottleneck? What is the next limit on making the thing you want? >> Uh I the bottleneck that is very apparent is like you know how fast can you verify that uh things are correct. So like we we I mean we can generate like code faster than ever before.

Uh we can implement entire features and you know we like I saw like someone just based on a description of you know the Codex app if you sort of like synthesize that into a plan uh just based on screenshots like the models are very much capable of like reproducing 95% of the features and just rebuilding the app from scratch. Um now is it going to be bug free? Is it going to you know is everything like implemented to like you know perfection in the same way that you know the actual app is like that takes like a lot of time still like you know for like a human to go and click and verify and you know make sure that um you know it's like it's the designs are like consistent and that you know there's like no bugs here or there um that the settings panel like you know when you click that button is actually does the thing that you expect.

I think verification, you know, definitely becomes a bottleneck. Like we have people on the team that complain, you know, like there's too much code to review. It's like, you know, that's what we're trying to solve for. >> I mean, you you complain about that. >> I complain about that. There's so much code to review now. >> Um, both that like co like on your own machine and like from another peer. It's it's like >> we're going to have to figure that out.

>> Yeah. you're already reviewing, you're reviewing the code the first time because the agent is just presenting it to you and then you have to review, you know, the code produced by your peers, you know, who are like there's like these two runs of reviews and >> Yeah. I mean, this is something that we're working on. Uh, a lot of us still do have to review code. Um, and we want, you know, we're taking a look at what that experience should look like with the model involved, right?

Um, we've got a review mode in the Codex app that works really nicely and kind of annotates your diffs on the side with findings and stylistic things and um, lot to do. Yeah, it's one thing I'm I'm sort of like also excited like about you know like making the models faster and then this like you know this one that we just put out is like you know which is mind-blowingly fast is like you can also use it you know you can imagine using it like in a way to understand code understand features um you know helping you with code review like helping you understand like you know the code that appear wrote and it's it's like much more pleasant because this is something that you want to do like you know you want to be there in the flow it's like something that has to be like synchronous uh it's not something that you delegate.

You cannot delegate understanding, right? It's like you know you're trying to like you know get to uh understand something and so like speed there like is a real advantage. So it sort of like helps offset as well like you know the fact that models are like producing more and more code is like you know speed helps you understand you know this code faster as well. >> Yeah. I mean, I definitely think I've found this already with this with this new model is speed, especially for end to end end to end testing is faster because if you're having it do end to end testing, like manual integration testing, often there's like a toast that pops up that pops up for like a second and if the model's not fast, it's not going to get it.

Um, and it seems like it's better for that because it it it the cycle times are much much shorter. So, um, and and I definitely find this too. It's like I can produce so much code but when I see a PR come in or when I make a PR my first question is like is there evidence that you've actually tested this and this actually works like not just unit test like you've gone through it end to end >> how do you how do you handle this >> I [snorts] mean I've seen a lot of peers that I have the same question about um it's like it's so easy to to code things now right um yeah I I we have gotten the Codex app to be pretty good at uh through some skills that we have of running itself uh clicking around screenshotting itself for evidence and uploading it to the PR.

There's there's like a lot that's pretty interesting there especially when we make this like more async or when you know the models get really fast at this stuff. like I don't know exactly what it looks like yet, but there is a lot there around like, hey, here's a bug fix. This is exactly like what it looked like when it was happening and here's exactly what it looks like now with the same exact click path. And so like maybe that's the turning point that code review becomes >> less important when it's like you can verify that part instead.

So you have to kind of like do less through the code as a proxy. Um there but there's there's definitely more to explore there. Um, last last couple questions. I'm curious. What have you guys learned from Anthropic and Cloud Co? Cloud Code and how do you think about your positioning in the market versus them? Like how how do you think about the differences? >> I I think they were first to put something out there and that was interesting to us because we had been working on similar ideas for a bit.

Um but I think our models were a little bit at the time not ready like you know they were not like reliable like on long horizon tasks like you know they were not able to like do like reliable tool calls and like you know stay on topic and so as soon as like we started to really invest on that and you know especially with GPT5 is like you know we were like okay the models are there we know how to make them even better 52 like brought you know even like better like long context long horizon like reliability, non context understanding and what we were seeing is that Entropic was sort of like you know through us like losing a little bit of steam um when it came to the model and we were in this fortunate position where like the way that we run CEDX is like you know we've got like product we've got engineering but we've also got research and we just like all work together and sit together and solve problems together and it's like a highly creative space where you know at times times we decide to like solve problems in the product in the harness.

But at times we also we're like hey how can we actually improve the model and like let's just you know talk about it and like you know idea together and then like research will come and be like hey you know we've got this like breakthrough that we're sitting on. It's just like would this be like sort of like something we can ship and then it was just sort of like get excited about that. One of the examples was we had a lot of complaints on compaction.

um you know compaction was like something that people felt like whenever you would hit compaction you know people would complain it's like it's losing too much uh context and so we sort of like solved that uh end to end and like you know we decided to do like end to end RL training and um you know introduce compaction like within research and then you know make the model like you know itself like you know very familiar with the concept of compaction and like producing like optimal like sort of like delegating to itself like across time and you know once we had that and we had solved it at model level like sort of like the hardness problem became like so much easier because it was just like oh just let the model do it and it's going to be like very reliable.

So through that and like through that collaboration, it just felt like like the momentum has been like very strong and that we're so like able to improve like models and like you know ship a model like roughly like on a weekly a monthly cadence and then we took like a bit of a different bet and like a different approach with the codeex app. Um which turned out to be like you know an awesome thing you know to just try and do is like not just like sort of like force ourselves you know and like trying to cram everything into the two.

Um, I mean it was like it was like a great challenge, right? You know, you were like, I'm you know, it's like let's build an app. Like just like where do I get started? And then, you know, just like you just got obsessed by it. >> It's hard not to. >> Yeah. I mean, it's like how was it to just like, you know, build something that was quite contrarian, I suppose. >> Yeah. I mean, I remember you and I talking about whether or not like early on we were like, we don't know if we'll ship this.

>> Yeah. like we're we we'll try it out. We'll see if we can get there with something that we love and see if we can get I remember saying like let's get some PMF internally. Let's let's get everybody at OpenAI to want to use this thing without being forced to use it. Let's see let's see if we can do it, right? We did and it was like adopted very quick. I mean the the minute it was barely usable, the research folks like put dev boxes on it, right?

like which was like this crazy hack at the time. >> Yes. Yes. >> Um but now they use it like for everything. >> Yeah. Yeah. It's like including in training like 53 codecs and so like I think I feel really good about having hit the point where like you know like everyone technical at the company like almost everyone at technical at the company like uses codeex but like the people who use it the most are you know actually building codeex and building the models and so you know we're just able to like you know improve things at like crazy crazy speeds and you know there's like no signs of it slowing down.

>> Amazing. Well I'm excited for what you ship next. Um, thank you guys for your time. I really appreciate it. >> Thank you. Thank you for having us. >> Thanks. >> Oh my gosh, folks. You absolutely positively have to smash that like button and subscribe to AI and I. Why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard, but instead of gold, it's filled with pure unadulterated knowledge bombs about Chat GPT.

Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat, craving for more. It's not just a show, it's a journey into the future with Dan Shipper as the captain of the spaceship. So, do yourself a favor, hit like, smash subscribe, and strap in for the ride of your life. And now without any further ado, let me just say Dan, I'm absolutely hopelessly in love with you.