发布 Gemini 2.5:Tulsee Doshi 谈史上最强模型
Launching Gemini 2.5 (Google AI: Release Notes)

We're here in Mountain View. Tulsi's back. We launched Gemini 2.5 Pro today. It is extremely strong in reasoning. It's an amazing coding model. It is the best model we've ever built. My head is somewhat spinning. We just shipped Gemini 2.0 originally a few months ago and now we're already at 2.5. We really have to get this model out the door. We have to get it in the hands of developers. We have to see what people are doing with it.
Shipping is a team sport. Long context, tool use, like so many of these things that we ship with our models are truly end-to-end innovations. What's coming next? I think a lot of what we're trying to do is just build like the most capable models. So I'm actually really excited to see what what people do people do with it. Hey everyone, welcome back to release notes. We're here in Mountain View. Tulsi's back. We launched Gemini 2.
5 Pro today. It's been a ton of fun. Everyone's been sort of celebrating the the state-of-the-art milestone for us. Can you sort of give us a rundown on on what we launched today and and why folks should be as excited as we are about Gemini 2.5 Pro? Yeah. So like you said, we launched Gemini 2.5 Pro today. Super excited. I think the really big thing about this model is that it is the best model we've ever built. And I think maybe even going further than that, I think it is one of the best models we have in the industry right now.
It is extremely strong in reasoning and so it is actually state-of-the-art in a number of the common reasoning benchmarks. It's an amazing coding model. It's particularly good at creating really fun web applications. It's great at agentic kind of code applications. And it also actually like is particularly good at code editing and transformation. So I think it just actually makes for a really good coding partner. And so those both are really strong.
And then I think also it builds on all of the great stuff we already have of Gemini Pro. It's multimodal so it's great for video understanding and image understanding. It comes with long context and so this 1 million long context window that we have which allows you to process like long videos or documents. And so really overall this is just an incredibly strong model. And we're just like really excited to actually put it in the hands of developers and customers and actually see what people are going to build with it.
I think one other thing I'll say about Pro which I think is really cool is I think it's actually a really well-rounded model. So I think a lot of models when you push for them to be really strong reasoning models, they're really smart at the benchmarks. But what's really cool about this model is it's also really good at you know, style. And it's like a fun model to actually talk to. And I think that's also why we see it doing really well on leaderboards that are useful for like user preferences, right?
So we actually have a 40-point jump in LM Arena ratings and Elo compared to the next model and I think a lot of that comes from the kind of ability for the model not just to be really smart but also to be um yeah, I guess well-rounded is maybe the right word. Yeah, I feel like that the balance of like passing the academic evals and the vibe test is like actually hard to do and it actually is kind of weirdly unintuitive that that's the case but I feel like it generally ends up being true from what we see.
Like I think you're you're closer to this but like as you hill climb some random eval like it doesn't actually translate towards like something that users are interested in the model doing. And it is this weird yeah, it seems like there's a weird disconnect between the Yeah, I think you used the right word which is like vibes of the model. So like I was playing with the model we shipped today over the weekend and I've been doing that, you know, with every model we train just like trying some things out and seeing how it goes.
And there's What are your go What's your go-to personal benchmark for this? Yeah, so I mean I think I typically go through like three different things. One is I'll just like try the basic like, "Hey, how are you? Like how's it going?" kind of just like to see how the model like responds to like really simple basic like, "What's going on?" I'll do things like, "Write a poem about walking on the beach where the sun is actually shining but also make sure that it references a sunset and also please make sure that it's about the month of March."
And like trying to see what the model does when you give it a set of very interesting instructions that kind of cover a different range of things. And then what's been really fun more recently as our models have gotten better at coding is to also have the model try to make like games. And so actually giving the model one-shot prompts and seeing what it can do, right? So I spent an unreasonable amount of time on Saturday playing snake because I had made a web app using Gemini 2.
5 Pro. I was like, "It's awesome. Like I can play snake." And to be clear, snake itself is not a complex game but what was awesome to see is like being able to take a single prompt and then actually have the model build something like visually aesthetic. So actually like colorful and with the right effects and actually being able to supply the JavaScript to really make the game more engaging. And I think the thing with the vibe check is you actually not only see the model doing the right thing and actually being able to like follow your instructions but also when you look at the the thoughts of the model and you actually look at the the response itself, it actually feels more engaging and I think that's important, too.
Yeah, I love that and I have so many random threads about this but let's come back to some of those examples cuz I think there's just like threads now all over the internet of like really really cool one-shot use cases. We put out a bunch of stuff. Developers have already put out a bunch of stuff. So hopefully we'll put them in the show notes or something like that. totally do that. I mean like Jack today was talking about a great example that I think a lot of people use to vibe check models which is and Jack is our our lead for thinking.
But he talks a lot about like the ball bouncing around the square. And I love that as a vibe check prompt cuz it tests not only like the ability for the model to generate graphics but also like understanding physics and being able to actually like manage the physics of that in some very simple use case but one that is actually like very evocative I think in in what you can do. Yeah, I put out a tweet exactly with the like ball bouncing use case and I used like the prompt that I thought was the standard prompt and somebody was replying to me like, "Hey, it doesn't have gravity or something like that."
And I was like, "The prompt didn't mention gravity." And then and then I think Jack actually shared an example somewhere of with gravity and it's just like it can do all types of crazy stuff. So yeah, hopefully we'll we'll have a link somewhere in the in the show notes and folks can look at some of these examples. But we made the jump from Gemini 2.0 Pro to Gemini 2.5 Pro. And we were talking about this jump and what that means.
Can you sort of talk us through like why this is so substantive, why we added that next version. I feel like, you know, my head is somewhat spinning. We just shipped Gemini 2.0 originally a few months ago and now we're already at 2.5. believe we only shipped Gemini 2.0 3 months ago? feels like a year ago. Feels like a year ago. Yeah, so why did we make the shift to 2.5? So I think for us these 0.5 increments, so we shipped Gemini 1.
0, then 1.5, then 2.0. I think 2.5 signifies two major things. One is a shift to what our models are going to represent going forward. So I think going forward all Gemini models are going to be thinking models. And that's going to be a fundamental part of how the models approach problem-solving. And I think that's already a big shift in thinking through the capabilities of the models and what they will represent. I think the second one is really the step change in performance.
Right? And so and that one is critical, right? If you look at the 2.0 series, that was already a step change from the 1.5 series and here we see this like significant jump. Tulsi, we were having this conversation earlier with with Sab who leads pretraining, with Melvin who leads post-training, and with Jack who leads the reasoning efforts about sort of what you described as like this across-the-stack improvement. Which it really does feel like this is true and and we sort of talked about how, you know, the pretraining benefits translates to better reasoning capabilities.
Just at the like from your perspective as the person orchestrating to make sure that we actually get a model that the world likes and developers can build with. Is it just like this moment where all of those three paths converge together in like a really like natural way and it happen or like is the was the intent like let's make sure that this pretraining thing lands at the same time that all these other improvements land so that we can get the like full breadth of the like how how does that process actually happen?
I think you know, I probably like any good system, there's a little bit of like centralized organization and then individual like ambition and like execution, right? And so every part of this stack is doing innovations of their own, right? And they're driving progress to really think through research and technical breakthroughs. And so you have the pretraining side really thinking through what is their sort of scientific method of essentially testing pretraining improvements and growing.
There is the post-training side that is thinking through specific capabilities and how they tune the model. There's thinking that is trying to drive new algorithmic innovations and kind of how we do reasoning. And so all three are driving those innovations. But I think what is really cool is one, they are all trying to think through composability, right? So the pretraining side is trying to think through how do we best train this base model such that it is most composable and most, you know, uh compliant I guess to like downstream post-training.
And so all of these pieces are actually trying to think through like how they be one puzzle and like how they fit together. I think the second part also is like intentionality of goal setting, right? So because we knew for example that code is an area we really do want to make progress on, that was something that we prioritized across all parts of this stack, right? So from a pretraining perspective, we thought about, "Okay, what are the what is the data that for example would be required in pre-training to be awesome at code.
From a post-training perspective, we thought about okay, if we wanted to for example, build better web apps, what would that look like? And then from a thinking perspective, we've also been thinking through how do we help the model reason about code? And so those three things then also come together well to push a single domain forward. And then I think that also translates across the board, but I think that's also been really critical is kind of driving towards a common force.
Yeah, that makes a ton of sense. And and actually to provide the counter example of this, like with 2.0 flash thinking, was that just like reasoning innovation? They're like we we didn't do a bunch of post-training or pre-training work? Or like what what was it the full breadth of the model? With 2.0 flash thinking, we of course benefited from the 2.0 flash model, right? So the 2.0 flash model was of course also had pre-training innovations and post-training innovations.
And then we were thinking about okay, how do we take that model and introduce reasoning into that and and build on top of what 2.0 flash is. So I think it still had all parts of the the stack, but I think the difference between 2.0 flash thinking and now we've introduced with 2.5 pro is I think with 2.0 flash thinking we proved that we could uh even on a smaller model uh that is sort of more cost-efficient and and things like that, actually make it perform extremely well on reasoning and complex prompts with the introduction of thinking.
But I think what we've done with 2.5 pro is to the point about well-rounded is we've taken all of those innovations and we've A heightened those themselves, but we've also introduced this idea of like how do we also make sure that it's also just a really um that it keeps the other aspects of what makes flash and pro great models, which is like the the strong multimodal performance, the long context performance, the style, the the vibes, if you will.
To use. Yeah, like all of those those aspects. Yeah, I love that. Um something that is exciting about this is like I think historically a lot of the reasoning narrative right now is like test time compute is the only thing that matters, like, you know, stop pre-training, stop post-training, there's no value in more compute at the end of this process. Yeah. And somehow magic is going to happen. And I think this is actually a great like verifiable example of like what what that's not actually true.
Like the there really is like the the hard work on pre-training like is paying off in having a better model. Is there any like yeah, how how like Yeah, I think actually Sub gave this example maybe earlier, I think, when we were talking with him about pre-training innovations. And I think like one of the examples he talks through is like you can have a model be really good at reasoning, but if it doesn't actually know the theorems, the the underlying knowledge to then build out its reasoning, its reasoning will only go so far.
And so I think the idea is that for example, pre-training gives you a base foundation of knowledge. And that base foundation of knowledge can then be like better customized and tuned and tweaked um in post-training for different use cases and, you know, built out. And then you have uh inference time on top of that is or as a part of that rather really um that can build that out even further. And so I think about it is like test time is obviously important.
Like inference time clearly is important and that's been proven both in 2.5 pro, but in a series of models that have shown that um extending the way that a model can think allows it to better produce outputs, but I think that is uh you can do more if you have the strong foundation to build off of. Yeah. Tulsi, I think one of the the most interesting threads of this launch is is sort of back to this like pace of how how quickly we're able to to make these models happen.
I think we we obviously behind the scenes we had a bunch of other, you know, model candidates. We were thinking about how we want to bring a a new version of this model that thinks to the world. But like what was the rundown in the series of events that led to today when we just shipped the model? Yeah, it's a great question. Like you said, I mean, this is obviously been something we've been working towards, which is to have a 2.
5 pro candidate really strong on reasoning, bringing thinking to our models, you know, actually having a Gemini series that that is really built on that reasoning capability. And so that's what the team had been working towards across pre, post, thinking. And I think one of the challenges we were encountering was to something I said earlier about the models being well-rounded, we were both trying to make models that were good at the vibes, if you will, right?
Like good for users in our products and um you know, consumers who actually want to want to play with the models, whether that's like a developer or an enterprise customer or consumer. We also wanted the models to be really good at the benchmarks and and at thinking and at reasoning. Um and so we were pushing on on both of these fronts and we were having trouble getting to a model that was really doing both of these things well.
We were finding candidates where we felt really good about how it was engaging or how it was working on certain tasks, but then couldn't maybe push it as much on code or weren't seeing the gains on reasoning. Or we had another candidate which was like really great at code, but we weren't necessarily seeing the impact elsewhere. The team had been, you know, um hill climbing on both of these goals and really pushing sort of innovations to to build on both of these.
And we got back some of our eval results, we were looking at the uh numbers, um we were playing with the model and going through examples and we're like, this is a really good model. And we're really excited about it. We did a bunch of our own vibe tests to kind of check the model on a variety of different prompts. I maybe built like 50 random web apps. I just tried to see what the model would do. And we were like, we're really excited about this.
Like this is a really fun model to test. And so we're like, we really have to get this model out the door. We have to get it in the hands of developers. We have to see what people are doing with it. I think one thing that's also really interesting when you're trying to take a model through this is how do we also make sure that we're being intentional about what we're releasing, right? So we have a model that we're excited about.
What are all the pieces that we need to work through? So for example, one area that I obviously think a lot about is safety um and how do we make sure that the model we're putting out is not just a good model, but it is also a safe model. And so one thing that's also I really like about how we've adapted our process to build quickly is safety is actually embedded in the development process itself. So every time we build a model checkpoint, we are evaluating that model checkpoint for safety when we're building it.
And so when we're looking at the model's numbers and we're looking at how it's performing, we're also looking at safety numbers um to say, hey, you know, what is it how is it performing? And then we ask our teams to actually red team the model. The team has been bashing at the model trying to, you know, identify issues. And it's funny actually because sometimes in doing so, they don't even uncover safety issues, they just uncover other random issues.
And so, you know, we had the team reaching out and being like, hey, so we're doing some safety red teaming and we found this other weird thing. What should we do about that? And so it actually also helps us patch kind of issues with the model overall, which is actually pretty cool. Yeah, I think the safety as a feature of the model development process is actually super important because I think it's like a key note of how we're able to move quickly.
Like I think if you if if you decouple safety from the model development process, you end up like, hey, we've got a great model, we want to put it out the door. Now we need to wait five weeks or however long it would actually take to go through the whole process. So Yeah, it's also not fun for anybody involved, I think, right? Because then it actually makes safety like a blocker in the process. Um and that doesn't lend itself to innovation, right?
Whereas I think when you actually enable safety to be a part of the process, what you're actually saying is how do we build a model that is also helpful? Right? Um and ultimately you change the notion of safety from being like um a a wall to being like actually a way to make the model better and more useful to people. And I think that's actually a a really nice change, too. No, I love that. And Tulsi, so one of the one of the big threads of this launch is the model's very good at multimodal understanding.
Video understanding is is something a bunch of folks have been talking about. Like why is there like something special we did to do this? Or like why like what's the the highlight of why video understanding is so much better with with 2.5 pro? Yeah, I think again it goes back to the whole stack piece of things where with video understanding video understanding is interesting because it's a combination of good multimodal understanding, right?
So you need to understand vision. It also requires often long context because for example, if you want to look at um a multiple hour uh match. So for example, my parents are big cricket fans. Um some of these cricket matches are hours. Like we're talking hours of content, right? And so to actually be able to put that in process that, you need long context. Um and then the last part of it is also actually strong reasoning, right?
So if you imagine for example, taking a video of a cricket match and then being able to say okay, I want to can you help me identify all the points in the video where actually a wicket was taken, right? And so you actually then want the some you want the model to be able to analyze the video, be able to pull out critical timestamps of that video, reason about those timestamps and give you explanations. Um and that's the kind of thing that brings together, I think, a lot of the magic of what Gemini models do particularly well, like the multimodal understanding, the long context and the um the reasoning pieces.
And so we're seeing that I think actually play out with this model. And so I'm actually really excited to see what what people do people do with it. What I wanted to ask earlier when you were mentioning sort of this trade-off between the vibes being right and academic evals. Like what does that actually look like in practice for like you're you're closer to the model development process than I am, but like you have a model that's really performant on academic evals.
Is it just that it's like harder to get the model to do like it's less like not less good at instruction following? Or like what does it actually look like for like the vibes to be off? it's less I mean, instruction following is important I think even for academic evals, right? I think there's some certain foundations to a model actually that are important that build up everything else, right? So like instruction following and steerability I think are just kind of critical tenants to the model being able to do a lot of other things.
But I think one way you can think about it, too, for example, is like maybe like another way to think about it is model behavior or like persona of the model, right? So anytime as you're trying to improve a model, you can think of it as hill climbing some sort of goal, right? And often that goal is set by an eval. So often you can hill climb academic benchmarks because you can you have a metric and you can sort of hill climb towards it.
And just to clarify, we care about at like academic benchmarks for us matter a lot because it's just like something that people universally agree upon or like is Yeah, it's a good question and actually like we don't look at all academic benchmarks because we sometimes look at an academic benchmark and say actually we don't agree with what this benchmark is testing or maybe it's a really leaked benchmark and so it actually like being really good at it doesn't give us much value.
And so there are academic benchmarks that we actively kind of are like, okay, that's nice but we're not going to pay attention to it. But I think some of these academic benchmarks so like I'll give the example of humanities last exam which actually like 2.5 Pro is soda on which is awesome. Um we got 19% without tool use or 18.6% I think it's without tool use which is which is awesome. Um and And what does humanities last exam actually cover?
Yeah, so so humanities last exam is the example of like I think 3,000 prompts that are supposed to represent really hard questions compiled by researchers and industry experts. And so humanities last exam is the kind of example of an academic benchmark that is really interesting to climb because it represents the kinds of questions we want Gemini to be really good at, right? And so then actually moving that metric is also moving towards a meaningful goal.
Right? Another example of of an evaluation like that is SweBench. Uh so SweBench verified um is an example of sort of like agentic code and that's another case where like we really want the model to be good at these type of sweet agent tasks. Being able to, you know, push on SweBench is a good way to to veri- verify and validate that. So I see academic benchmarks as kind of a way to A motivate a specific destination of progress and then also help communicate to developers what the model is good at, right?
And and kind of where it has excelled um which I think is interesting. And I also think it's it's important to accompany any academic benchmark also with our own internal evals, right? So a big part of our efforts within Gemini are to building the right evals um and making sure that we're actually measuring, you know, the the goals that we have. Yeah, I love that. That's awesome. Um this this launch has been a ton of fun.
There's been lots of of interest. Uh I'm happy that we have a state-of-the-art model. What's coming next? I think there's we we sort of, you know, have a sense of what the road map looks like at least for the next few months. Um what can folks uh look forward to? Yeah, uh so much to look forward to. So first of all, uh 2.5 Pro released today experimentally but one thing I know we've been talking a lot about is is actually getting people uh access to the model in a way that they can build with it at scale.
So one of the things we're really excited about is actually pricing uh the 2.5 Pro model and actually releasing it for use in production and use at large scale. So that's one thing I think we should all be really excited about. Very soon, hopefully. Very soon, hopefully and I I think this is hopefully us really listening to developer feedback, right? And actually like acting on it in a useful way and so I'm ex- excited to see what people do when we can give them more scaled access.
So that's kind of more tactical. I think in the more kind of immediate term, we of course want to bring the 2.5 kind of series to more models and so, you know, Flash is is of course the next on our list to come forward. Um and then we're also thinking about how do we make these models more usable, right? So one of the challenges with thinking models is that they think often for a long time. Um and that's helpful to make the model more performant but models don't need to think as much necessarily for simpler prompts, right?
So going back to like my sort of different levels of vibe checking, for a web app that is really complicated, maybe the model needs to think longer. But for like hi, how are you? It probably doesn't really need to think at all. And so how do we make sure that the model kind of like better learns, you know, how how to modulate? Um and then also how do we provide developers more control and what does that look like, right?
Um especially when we're thinking about cost and latency and different types of applications. So that's like a lot of what we're thinking about in the near term is kind of like how to bring that to bear. Um we're also thinking about things like image generation and how to, you know, bring that uh into the fold of these models. And so there's a lot of really exciting pieces coming. I think if you then look like even farther out, I think a lot of what we're trying to do is just build like the most capable models um to be able to help you whether that's in coding, whether that's in building agents and awesome kind of agentic applications.
You know, when we talked about 2.0 Flash in December, we were talking about Mariner and UI control and I think there's a lot of really cool things that we can start to do as these models get more and more capable in terms of like building sort of end-to-end experiences that I think will be really fun. Yeah, and just one quick follow-up. So today's 2.5 Pro model um does do a little bit of dynamic thinking. So it's it's maybe not the like full version of how we want that experience to look like but if you if I ask a simpler question like the model's been trained to know like hey, these are Yeah, absolutely, right?
I think like for simpler prompts the model will definitely think less than for more complex prompts but I do think it is true today that the model probably overthinks uh quite a bit. Um and I think that's actually a good place for us to start, right? What we really wanted to do was kind of push how the model could use reasoning to solve awesome problems and now I think we want to take a look and say, okay, how do we continue to make this more and more useful for developers and for customers who should be more cost-conscious of like kind of how how the model actually modulates that.
Yeah, I love that. Tulsi, this last 3 days has been a ton of fun. I think the most fun I have in in my job is when we're when we're sort of sprinting through the chaos to get the little ship emoji. Yeah, the little ship emoji and getting stuff out the door. So this was a ton of fun. Huge thanks to your team for making all this possible and and the rest of the Gemini. I mean shipping is a team sport. I think that's one of my favorite sayings these days.
Like it really does take so many people to make these launches possible. So It's kind of amazing. Like I feel like we have we talk a lot about like the parts of the model training but I think we don't always talk a lot about the parts that happen after. Like um you know, we have Madvi and and the team building demos and actually like really testing the model. We have like our marketing and comms teams. We have our deployment team like actively figuring out how to serve the model efficiently and like trying to figure out all the bugs and the kinks and getting the model stood up and it's Which is not easy.
It's not easy. We talked to Emma. Emma was the first person to come on this on this show or whatever we call it um and like he's, you know, him and and the serving folks and and Evan and everyone Joe and everyone else in the and Alvin in the trenches like doing all this work. It's it's a ton of it's like a beautiful orchestration of chaos Yeah, absolutely and I feel like the like I feel like even in the last 3 days but as we've been doing more of the shipping, I've been learning so much more about like deployment and serving and just like how complex that is to get a stable experience set up for for an end developer.
And I yeah, I feel like it's it's been quite the quite the team team effort. Yeah. Yeah, I need to trademark that long context is as much of a model innovation as it is an infrastructure innovation. I think we actually feel this in a lot of cases. There's there's challenges with making long contexts work because it's it's an infrastructure innovation and it takes it takes effort to make it happen. context, tool use, so many of these things that we ship with our models um are truly end-to-end innovations and yeah, long context is a great example of that.
Getting it right is is hard um and it's a huge team effort. Tulsi, this was a ton of fun and you are now becoming the resident model launch co-host of this. So hopefully we'll we'll have you on to do to do this again for going to ship a ton. So we're going to be doing this a lot. I'm excited. Hopefully we'll be in person in Mountain View again. Deal. I love it. All right.