Tulsee Doshi · Google DeepMind Gemini 模型产品负责人

Gemini 2.0 幕后:Tulsee Doshi 深度解读

2024-12-11 · Google AI: Release Notes (Logan Kilpatrick) · 35m · 原文链接

→ 在 AI 访谈库中阅读(可切换中英、记录进度)
深入 Gemini 2.0 的多模态能力、函数调用与多模态智能体等原生工具使用。看点是她解释 Google 的'实验性模型'发布策略——先 experimental 快速放出、再走向正式版的节奏是如何设计的。

I'm super excited about 2.0. What's the sort of big picture? Like why why should we be excited about 2.0? Gemini 2.0 actually can enable you to build these awesome multimodal agents. I think we're going to start talking more about the kinds of agentic experiences we're building as opposed to just talking about agents. The fact that we're shipping 2.0 flash first is the thing that makes me the most happy. It basically performs better than our 1.5 Pro model. Flash is a really practical model, which I love. What are the new modalities, new capabilities that are coming in this model? Screen understanding, spatial understanding, native search tool use. And when you combine these things together, you can do really awesome things. I'm actually honestly so excited. It feels like this is just the start of the 2.0 story. On today's episode, we're talking with Tulsi Doshi. She's the head of product for Gemini models, has led a bunch of different teams at Google, and is one of the many co-conspirators in helping ship all the experimental models we landed recently. Tulsi, welcome. Thanks for having me. It's great to be here. Tulsi, I'm super excited for this conversation. Um we're launching Gemini 2.0 today. It's been a long road to get here, but I'm actually curious before we talk about all the cool new stuff that we just shipped with 2.0. Can you talk a little bit about sort of where we've come in the last year? Gemini released December of 2023. What what's the progress look like in the last 12 months uh to get us to today of of shipping Gemini 2.0? Yeah, it's kind of amazing. I think um if you think about the fact that we're really at the 1-year anniversary of Gemini, and it has only been a year, um it actually just shows the rate of progress that we're having both at Google and around the industry, right? So a year ago today, we were shipping 1.0 Gemini. This was the first time we were shipping. Uh a large model in this context, in the context of like an API and an external experience for developers. We were learning so much even just about how to organize a team to do this work. What we wanted to ship, what kinds of experiences developers wanted to see, what kinds of experiences enterprise customers wanted to see. And you fast forward a year and we are shipping regularly. We have new versions of these models um that we're training on a regular basis. We've developed more of a rhythm. We've developed more confidence, I think, in what we're building and why it's important. We've also just shipped so much into our products, right? So, if you look at Google Search, if you look at the Gemini app, if you look at YouTube, if you look at our APIs, we're actually bringing Gemini to all of Google's products, Workspace. And that's been also really awesome to see just the progress from where we were last December to now. It's amazing and it's really exciting to be a part of. Yeah, I love that. And does anything feel like starkly different? I wasn't at Google when when the last Gemini model shipped. I was I was watching from the sidelines, but does anything feel like fundamentally different this time or does it really just feel like, you know, another another cycle of craziness? No, I I actually think it really does feel fundamentally different. I mean, so to be fair, I also was not in the current role that I'm in right now when Gemini shipped last December. But I was in Google and I was part of part of the Gemini launch. And I think when we shipped last year, we were shipping something for the first time. It sort of felt like we were walking into the unknown and really kind of coming together do something that was fundamentally new. I think what's really exciting about this year is we are shipping fundamentally new capabilities and we are making Gemini 2.0 more awesome, but it feels a lot more like we've built the muscle around how to ship these models, and that's a very different feeling, right? Like the feeling of a more a more well-oiled machine, I guess, than the one that we had when we started. I also think we have a lot more to the point I made earlier, like conviction in what we're shipping. I think we were really excited about what we shipped last year, and it was amazing, but it was also the whole industry was learning kind of what GenAI would mean for the world and how people would use it and how people would use the models in different ways, whereas I feel like this year when we're shipping, we also have a lot more clarity on the kinds of use cases we're trying to optimize for, what metrics we really are excited about, what progress really means and what good really means for these models. There's just also a lot more in the industry. We see more people testing out the models and trying them out, so there's more community around Gemini, which also makes it more exciting to ship because we have people who are actually really on the other side of testing and trying these models. And so it does actually feel like a meaningful shift in in kind of how we build and how we ship. I'm assuming it's probably just as chaotic, but but I think that's Yeah. That has not changed. I love it. So, 2.0, what's the sort of big picture? Like why why should we be excited about 2.0? What are the new modalities, new capabilities that are coming in this model? I'm Yeah, tell it tell us a little bit more from your perspective. I'm super excited about 2.0. So, you know, when we initially announced Gemini a year ago, we talked about this kind of universal multimodal model. And I think with Gemini 2.0, you're really actually seeing that realized in so many ways, right? Gemini 2.0 actually can enable you to build these awesome multimodal agents. And you see that in Project Astra, which is being powered by Gemini 2.0. You're seeing that in another project we're announcing called Mariner that actually enables um actions to be taken on on your computer screen. And I think a big part of that is that Gemini 2.0 is fundamentally natively multimodal. So, it can actually output images, it can output audio in the form of text-to-speech, it can actually help drive actions through really awesome spatial understanding and reasoning. And so, the bringing together of these capabilities and how much better they're getting in the model actually empowers a whole new set of use cases, which I'm really excited about. I think it's also worth calling out too that 2.0 Flash is also just still an extremely fast model, which is awesome. Because you can actually do incredibly complex reasoning, code, multimodal tasks with a model that can still perform at lightning speed. And I think the combination of those two things means that the model is super good for like real-time applications. The model is really good when you actually need to get things done quickly or large numbers of tasks. Um, and I think that is a huge delta, right? It basically performs better than our 1.5 Pro model. Yeah, I think it's uh the fact that we're shipping 2.0 Flash first is the thing that makes me the most happy, because I think like 1.5 Flash has really struck the cord with developers. Like if you go ask developers, it's like clearly the de facto choice for a lot of people starting to build with with generative AI and even others who have been using other models. And to like continue that with like the narrative of you know, better, faster, cheaper, all the things the developers were already excited about with 1.5 Flash, more of that I think is just like such a such a win for us. Yeah, and I think it really like speaks to ultimately what we really want to build at Google is models that are really smart and really capable, but also that you can use, right? And that are practical. And I think Flash is a really practical model, which I which I love. As a product manager, it feels like that's the kind of thing you want to anchor on, right? Like, what can real users actually value out of our product? Yeah, 100%. Um, so this was not the first new model that we've shipped recently. We released those experimental models, uh, the 1114 model and the 1121 model. Yeah, the numbers are confusing in my head right after rapidly one after another. Can you give and we you, me, a bunch of other folks internally talk about this a ton. Like, why why are we doing this? Like, why do we release those experimental models? Like, from your perspective, like, as the person pushing send on putting new models out in the world in many cases, like, what what's the feedback loop look like from a developer perspective, as well? Yeah, um, so we started shipping experimental models this summer, and we've recently shipped 1114 and 1121. And, honestly, it's been so invigorating to ship these models. I think what is the So, the motivation behind them is, you know, I used to be a product manager on YouTube, for example, and we would run live experiments constantly, right? We would put out new changes, uh, and actually get real feedback from our users on how those changes were perceived. And, on the model building side, it can be hard to figure out how you get that fast feedback loop, right? Because if you put a model out and you actually have a company start building on it, you can't just pull that model back and be like, "Oh, actually, psych, we changed something. Here's a new model." Right? But, you also want to be able to give developers and enterprise customers a chance to give you feedback and actually tell you what they want to see in the model, what kinds of use cases they're actually excited about. The other cool thing is, because GenAI is still so new, every time we ship a model, we see people doing things with that model that we didn't even come up with ourselves. Right? There are use cases that people are discovering and trying and finding that actually open up new doors and get us excited about new things we should be be Um, and I think that only happens if we create this experimental feedback loop. And so, the big motivation behind putting this model or putting uh experimental versions of these models out is that we can get a model out in front of developers, we can get real feedback, we can see what people are excited about, we can take that feedback and put that into the next version of the model and keep making it better so that by the time we ship a model that is for production use, we've actually gone through the process of figuring out what are the bugs, what's working, what's not working, what are people excited about, and I don't know. I I think it motivates us to ship um in a way that's more agile, which I think is really good. Yeah, I think the most fun that I have is when we end up with a new model that we want to ship. It's experimental and then it's like everything is on fire for the next 48 hours until we get this thing out the door. It is um it's chaos, but it's I actually think it's one of the most fun things to be a part of uh getting those out the door. So, hopefully we'll see more in the future. I'm hoping this actually becomes like a part of our muscle. Like I'm hoping that that and I think we're already seeing that that experimental models is just becoming a way that we ship Gemini and a way that we actually engage with our Gemini community. And so, to anyone watching also, we really actually really want you to try these experimental models and we really want your feedback because that is actively going into the models that we're shipping um both to our products and eventually to developers and enterprise customers, you know, in global availability. Yeah, what one of the questions and some of the feedback that we've gotten is like, you know, experimental models are great, we love them, they're awesome and ton and a ton of fun to play around with, but like we actually want to build stuff in production with these models, they're really good. What what what is the answer or what would you say to like folks who are just trying to get their hands on a model that they can actually put into production? Yeah, I mean, we want that, too. So, we we also really want you building with Gemini and and building in production environments. I think um and we are working to get you something soon. Um I think that like the big thing I would say is I hope that us shipping experimental models is not a sign of us shipping production models slower, right? I hope it's actually a sign of us just shipping more faster. Um and so we're trying to put more models in your hands to get your feedback, see how these models can build, um keep building the momentum, but we plan to keep shipping productionizable models on a on a regular cadence and make sure that we're holding ourselves to that. So, um I actually excited people are asking that question because it means people are excited about the models we have coming, um and I hope we can keep that keep that excitement going. Yeah, we we need like a public benchmark for ourselves, which is like the the time between GA models and like the number of experimental iterations it sometimes take. And I'm hopeful that the data will will say exactly what you're saying, which is like this is really a a path to help us get the best models into the hands of the world so they can really build the things that they're excited about. Yeah, and we will have a GA model coming soon. We really are are working to make sure that there's something exciting that you can build with. I love it. 2.0 three sort of main features in my mind, all the native tool use stuff, all the multimodal stuff, and and maybe, you know, I don't I don't maybe agentic nature of the model as well as a as a third class if if you want to separate that from tool use, but maybe we'll start with with native tool use. What's the what's the story here? Like why why should people care that native tool use is is now available in the model and and what's the sort of one level deeper of like what's actually happening at the model level if you can talk at all about that? Yeah, um so there's two things I'll say about kind of what native tool use is in this model and why it's exciting, and then we can talk a little bit about like what why this matters. Um so one of the things that we're introducing with uh with 2.0 is native search as a tool. And what's really cool about that is it actually we're training the model to know when it should go and call search, um, to validate a response or to get information. And one of the things that models struggle with, and we've known this for a long time, is like hallucinations and factuality, right? Because the model doesn't have all of the information, and that's especially true with freshness, right? So, if something happened yesterday, that wasn't in the model's training data, um, and therefore the model is likely to hallucinate or make up a response. And so, when you train the model to know when to call search or to realize when it doesn't actually have the information to answer a question, then you're training the model to call from search, um, and actually answer in a way that is much, much more accurate. And so, we can see significant increases in factuality of the model, um, with native, uh, search usage. And the question you might ask is, well, okay, how is that different than just calling search? Like, what does it mean, um, to natively call search? Well, you don't actually want to call search on every question, right? There are certain questions where you might actually want to write a creative story, and you don't actually want search, um, or certain questions where it's not necessary because the model can actually answer well. And so, training the model to be smart about when it calls search, uh, means you actually just have a much richer overall model experience because when you call search, you do it well, and when you don't need to call search, you also have a great model response, and so, your overall quality of the model stays really good, right? And that's kind of true for any sort of native tool usage, um, someone today, when I was talking to them, gave me this analogy that I think is a really helpful way of thinking about it. Um, if you're learning a new language, right? Function calling, which is how a lot of models call tools, is basically telling the model, here's a new word, like, here, go, learn this new word, right? Native tool use is all about how to use words well, right? So, how do I use that word in the best possible way, in the best possible sentence structure, in the most creative ways, but also, how can I chain multiple words together, um, and actually do multiple tool calls, uh, that can actually combine together? So, for example, uh what we want Gemini to be able to do is know that it has to call search, and then also know that it has to call code, right, a code interpreter. And so, it has to both take the search information and then maybe generate a graph. And so, it actually has to call Python. Um and so, the ability for the model to actually know that it has to use multiple tools together and be able to do that seamlessly, um is really where the power of native uh tool use comes from, which I think is really awesome. I love that. That uh that analogy is incredibly helpful as someone who's been trying to wrap their head around uh native tool use. What Yeah, I can't actually take full credit for it, but I I will I will share it cuz I think it's awesome. I love that. What is um in in situations where the model natively wants to be able to call a tool, is it actually is there a way to say like can you disable tool use? It's like, you know, your your Yeah, passing the tool to the model to say it's enabled. And then what happens when the tool isn't enabled? Does the model still like say I'd like to be able to search here? So, when the tool isn't enabled, the model won't be able to call the tool, and so it'll answer with what it has at its disposal. And so, you will still get a reasonable answer to the extent that the model can answer reasonably, right? But I think what you find in the case of search, for example, is just the factuality is so much higher when the model can access information, especially in the context of like recent things, you know, things that happened yesterday or the day before or something like that. Yeah, that that's super interesting. I'm I'm also curious if we have any um info from like internal benchmarks about like how this impacts like human preference. Like I know that like on a pure factuality basis, like you would people probably want the truth. But in some cases, like I'm not sure like how did from like a style perspective, from like a like just way in which people interact with the model, do we see the like model behavior like shift when native tool use is enabled versus like a traditional sort of chat experience? Yeah, it's a good question. Um I think we've done some testing, although again I'm hoping actually that our 2.0 release will give us even more feedback of like real usage um from from people. I think to to your point, one thing that's tricky about factuality type cases is um I don't often know what is factual. And so if you're asking me between example A and example B, which one is better, um I'm not necessarily going to make a call based on what is correct because I might not actually know what is correct or I might be biased by my own personal opinions. And so I think it's hard to like fully disin- like disentangle when we're trying to be factual versus when we're trying to provide like a user preference um experience. And I think we want to balance both. Like we want to especially in certain domains, we care a lot about factuality, right? Like there are domains where I think it is actually like medicine where it's like it's really important to get the answer right. Um and I think that's like one way that we've been thinking about evaluation. Um on the whole though, I mostly would say we've seen positive feedback. And I think part of that Logan is because of the conversation we were just having about like chaining multiple tools together. Like native tool use gets you on this path towards agentic behavior. Like it allows you to do just really cool things, right? Like you can actually say, "Hey, um help me go get this information and then help me plot it." Right? Um and now you can actually do a combination of things because of these tools being able to work together. Um and I think we've seen like a lot of positive user feedback on like, "Whoa, that's something I didn't realize the model could do or I could use Gemini to do." Um and I think that leads to uh to a lot of positive feedback in and of itself. I'm curious how we make the decision about whether to make something a native tool. Obviously, search makes a lot of sense for us because we're Google, but what are the other sort of on the horizon from a native tool capabilities? Is the plan to like make everything a native tool capability or should developers be thinking about like if my, you know, the tool I really want this thing to be good at isn't a native tool, like what is the on-ramp for people look like in the in the future from that perspective? Yeah, it's a good question. I mean, really fundamentally we want the model to be very, very good at function calling. So, my hope is that as a developer, the fact that we're introducing this kind of native tool use and making the model better at chaining multiple tools together and enabling like compositional function calling, will actually just enable you as a developer to to use any of tool that is that is valuable to you. So, I think we don't want to be limited by a small set of native tools, right? I think of native tool use and specific tools as areas where we really know we as Google can add value to you as a developer and really double down. So, if you look at Astra, for example, Astra uses lens, it uses maps, it uses search, right? So, it uses a lot of the Google magic to provide a kind of more holistic experience. And I think one thing we're thinking about is, okay, how do we bring in the magic from some of those tools? What does that look like? When does it make sense? And so, I don't have a great answer of like these are the five native tools we definitely want to have. I think a lot of what we're looking at is what are developers trying, where do we as Google like have tools that we can really add value to developers, and where does also the fact that we're training something natively make a difference in performance. So, I think like one thing about search that like really makes a difference is is the ability to know when to call search versus not, right? Because you don't actually want search for every prompt. And and you also don't as a developer or as a user want to have to know when to call search, right? Like it's not always intuitive to be like, ah, this is an answer question for which I need a factual search motivated response, and this is a question for which I don't. And we want to be able to like abstract you away from having to think through that um for something that feels as core to the model like factuality. And so we're trying to figure out what are other examples of that. Code is another one where we really do want the model to be able to do native code execution because there's a lot of cases where again, like you shouldn't have to think like ah this is a case where I need code, right? For example, if you ask the model um what is 1052 + 47? Right? Um and the model says, "Okay, I need to actually run code to to compute that because that's the way I'm going to get it right." Um you don't want to have to be as a user knowing that you need to call code. You want the model to know natively that it should execute code for that kind of example. And so how do we um how do we find those cases where we should actually abstract that away and make it easier for a developer or user? I'm super curious if we if you can talk at all about how the model makes the decision uh whether or not to call search. Is it just like we have a bunch of training data based on human preference where an annotator says, "In these cases, you know, there's a factual question and therefore we need the model to go and call the search tool." Like what does that actually Yeah, can you give us an extra level of detail? I think it's it's actually like really interesting in terms of how to actually train the model to be good at these things. What's also really interesting is like helping train the model to understand when it needs to use multiple tools, right? Like when it can't get like all of its answers from search and therefore would potentially need to use maps or lens or or something else. Um and that I think is also uh like an interesting thing when you think about Astra and it trying to actually the model needing to understand when to call lens as a tool versus when to actually do a map search versus when to go to Google search, for example. Yeah, that makes a ton of sense. I'm going to I'm excited to spend uh the next few weeks over the holidays looking and trying to build some stuff with native tool use. The the next sort of thing that is new in 2.0 is this additional multimodal story. Uh we're launching the ability for the model to be able to generate images natively and then the model is also going to be able to generate audio natively. I don't know if you have anything like more you want to add from that perspective, but I'm also curious like what the you know, what what else is on the right? Like is the model going to be able to smell in the future or all these are like what what are the other modalities that we'll hear If the model could smell, that's like a whole dimension I have not thought about. Um but I think yeah, first of all, I'm so excited about the multimodal generation. I think one question people ask me is like why does it matter if Gemini can generate uh audio or images because there's also imagine which can generate images and there's text-to-speech APIs which can generate speech. And so why is it so cool or different or interesting that Gemini can do it? And I think what's so powerful about where we are now at the frontier of Gemini being able to generate multimodally is that you can actually combine like Gemini's real-world knowledge with um with its ability to generate, right? So and like two different examples of this, one in the image generation case um is we have a fun example that we've been playing with where, you know, you put in an image of a cup and a book sitting on a table, right? Uh and the cup is on a saucer. And you say, "Hey Gemini, add a spoon to this image." Right? Now, if you weren't a model that actually had real-world understanding and understanding kind of of where a spoon should go in that image, you could put a spoon anywhere, right? But because the model actually understands that the spoon is probably connected to the cup and the saucer, that's probably where you want it in the image, not on the book or randomly on a plant, right? Um the the Gemini actually does a great job of sizing the spoon the right way, placing it in the right place. Um and that's the kind of magic you can do when you actually bring these things together, right? Another example is localization, right? Like when you say, "Hey, generate a person on a bench." If I'm in India, generating a person on a bench might look different than if I'm in Seattle, right? Um then if I'm in France. And because again, Gemini has real-world knowledge, you can actually generate images that much more map to uh the real-world context, right? An example I love to give is breakfast is just different in different countries. So, generate an image of breakfast is just not the same, you know, depending on who you are and where you're eating. Um And so, I I think that like combining those things together is like really where the magic comes from, and like that's really what I'm excited to see people play with and see like get feedback on. Um and we're seeing this play out with native audio, too. Like, what you can do with native audio generation is actually give it styles, right? So, like, say this in the style of a pirate, or say this in the style of, you know, and um that is again something that you can do because of Gemini's real-world understanding. You can actually bring these things together and create, I think hopefully, more magic in the process, or at least complementary magic in the process, which is great. So, Gemini 2.0 Flash, our first natively agentic model. What does that actually mean? Why should people be excited? personally, I'm really excited about like the future of where efforts like Project Mariner go. Um I think there is something really powerful about being able to actually automate certain tasks that otherwise feel so manual, right? So, being able to like take a recipe and then actually say, "Hey, put the ingredients in my shopping cart." Like, that's the kind of thing that I feel like um they're simple, right? They're simple tasks, um but they make such a difference. Uh I think I'm also just really excited about audio continuing and like dialogue being a continued form factor with the model. Um and I think that's what Astra really capitalizes on because it actually just makes the interaction seem so much more natural uh than I think we're typically used to when you you know, type in a query or you type something on your phone. Um in terms of like what does it mean for Gemini to be able to power these kinds of agentic experiences? The way I would think about it is Gemini is like an engine that has a lot of these core capabilities like screen understanding and spatial understanding or um native search tool use, right? And when you combine these things together or you can do really awesome things. Right? Because if you take for example Gemini's core reasoning capabilities and screen understanding capabilities, you can then help the model understand how it should actually navigate across a website. Um and those kinds of things enable then this notion of what is agentic which I think to me really means um the model being able to actually complete actions in the real world on your behalf. Um and I think we're now at the tipping point of really being able to achieve those with the capabilities that we've brought together here. Tulsi, this is a hot take about agent stuff. Um I think a lot of the demos that I see are around like the model doing things that actually humans get value out of. I was talking to my girlfriend about you know, AI agents shopping and stuff like that. Shopping? Yeah. She's like I love shopping. Like why would I want agents to be go and do that? I also love shopping. shopping, too. I mean it's uh it's a ton of fun. So, I'm I'm curious like where A, what you think the actual agentic use cases like perhaps like for a developer right now who wants to build something, where the value can be created. Um but also like is that going to change in the next 12 months? Like does Gemini like fundamentally change that? Uh yeah, I'm curious. Look, I think that we do want Gemini to be augmentative to humans, right? And so, ideally you want it to help you in the places where you need help. Ideally you want it to take things off your plate that you didn't want to have on your plate in the first place. And different people are really different, right? Like I love shopping and I personally get a lot of value out of it. My husband hates shopping and would probably very much love if Gemini helped him buy a bunch of black t-shirts, you know, for him to wear for the next 6 months. But that being said, I think there is like this is why like for example the shopping cart use case I find like particularly valuable as an example because it's the kind of shopping that I don't enjoy doing, right? And I think there is like an it's like interesting to see that these models will probably have different value to different types of users in different types of contexts, right? And so I think our job with Gemini should be to create a number of opportunities for developers to create experiences for their user bases that maximize their user bases entertainment, productivity, potential, however we define that. And I do think we have to be careful to make sure that we do so and build the experience in a way that still gives people agency over those experiences and enables them to find the joy in the parts of the experience that they're excited about, right? And do so in a way that is like safe and intentional. But I I don't know. I think that because there's such a diversity of people around the world who want to automate different parts of their experience, I think that developers will will find that they can find different niches of where that actually plays out. Tulsi, we've seen sort of a bunch of other trends over the last 12 months sort of end up getting bundled under the umbrella of AI and I feel like agents has the propensity to potentially be one of those things. I hear agents every single day. I hear AI every single day and I feel like they're eventually going to, you know, potentially merge together as the same thing. I'm curious what you think about that and also whether or not 2.0 is sort of a a step in that direction. I kind of hope Gemini 2.0 is a step to that, right? I think um the minute you start using a word to refer to a lot of things, it starts losing some of its meaning. Um and so I think right now agents have a lot of meaning because we're still trying to define what they are and still trying to build kind of the foundations of what they are. Um but I think as we have more um I don't know, as we have more hopefully models like Gemini and then kind of platforms and infrastructure built on top of them that enable um more developers just to build agentic experiences. I think we're going to start talking more about the kinds of agentic experiences we're building as opposed to just talking about agents um as a whole because we're going to get a lot more nuanced about the kinds of ways that you might want to interact with the world. Um and then I think the conversation will shift and the type of vocabulary used will also shift. Uh so yeah, I do think agents as a term is probably going to become less I don't know if less used, but at least less meaningful over time and is then going to be replaced with like more specific definitions of what we actually mean. Tulsi, I have a bunch of rapid-fire questions uh before we wrap things up. Uh these can be your your personal opinions, not sort of officially Google-endorsed positions um as as much as you you want them to be. If you had to have named this Gemini model, the 2.0 model, something different, what would you have named it if you had the chance? Any anything goes. You don't have to follow the brand guidelines. Oh my god, I am such a horrible namer. Like that's That's a terrible question. Um Ooh. All right, I don't know actually. I really don't know what I would name We're good. can call it Tulsi. That's not a bad option, but I love that. Um this is one of the you know, one of the questions we've we've shown sort of success making small models. Um I'm curious if we'll if you think we'll get a large model eventually. No, I I do still believe that there is a lot of value in scaling up. I think we need to do both the small models and the larger models. And so we are thinking through what that looks like. I think we need to first think through like what are the real use cases where a developer wants that level of power or reasoning is valuable. Um I also think with the growth of kind of inference time um efforts to actually power really strong reasoning and capabilities, we also need to think about where does uh size of a model add value versus where does inference time add value, and that's some of the you know research we're looking into right now ourselves to kind of better determine what the path going forward is. I love that. And last one is um any AI use cases that like you're personally extremely excited about that like just don't work well enough for them to be like daily parts of your life, but you're hopeful like another couple turns of the model crank or even 2.0 like actually end up enabling uh for for you? Ooh, good question. Um I mean, I do think like going back to kind of what I said earlier around like the computer control, I think we're like at the beginning of that journey, but I think there's so much more work to do to really get that to be a like safe, intuitive, useful experience. Um I'm also a huge fan of just like dialogue. Uh and I think right now multilingual dialogue is a space that we are getting better at as an industry, but really to be able to like speak to the model in Gujarati and Hindi and English interchangeably and actually have the model respond in a kind of clear um aligned way to my mental model um is a space that I think we're going to see continued improvement on, which I'm really excited about. Um and then one thing the model does not do well right now that I would love to see more of is I'm a dancer, and I feel like um I keep trying to use the models to like help me with choreography. Um and that is an area that I think we uh we can still do some work to get better at. So, that's one that I'm like secretly holding out hope for. I don't think anyone else cares about that use case other than me though, so you know, we'll we'll get there. Tulsi, th- this was a ton of fun and awesome conversation. Thank you for all the hard work and all the hard work from the rest of the Gemini team the last you know, six months to get this model out the door since May and I'm actually honestly so excited because even though this you know, we've been in a full out Sprint to get here feels like this is just the start of the 2.0 story like there's so much more stuff that we have coming and yeah, it's going to be a fun next six months. Yeah, I'm really excited. 2025 is going to be great and thank you Logan too. I feel like this has been quite a ride but it's been so awesome. I'm excited for the next year. That's it for this episode of Google AI release notes for more about Gemini 2.0 scaling and some of the new research prototypes check out the Google deepmind podcast with Oriol Vinyals. They'll be diving deep into the future of generative AI. Until next time, I'm Logan Kilpatrick.