🎙️AI 访谈库
AI 与编程:对话 Google DeepMind 的 Varun Mohan
Varun Mohan · Google DeepMind

AI 与编程:对话 Google DeepMind 的 Varun Mohan

AI & Coding: Interview with Varun Mohan of Google DeepMind

2026-04-03 · MIT Schwarzman College of Computing (Armando Solar-Lezama) · 45m · 约 51 分钟读完 · 原文
MIT 教授 Solar-Lezama 专访加入 Google DeepMind 后的 Varun Mohan:AI 编码系统的演进与 Windsurf 经验对 Antigravity 项目的启发。

So, for for the next for the next event for today, so we're going to have a discussion {slash} chat {slash} interview with uh uh with Varun Mohan, who is one of our alumni who graduated here a few a few years ago and has gone on to do some great things in the world of of AI for code. So, maybe what do we start just telling us a little bit of just your history of how you went from you know, being a student here at MIT to driving the uh driving the development of the anti-gravity tool at at Google.

Yeah, so I'm happy to happy to give like a pretty quick summary. So, I'm originally from the Bay Area. I actually started the company with someone with another MIT alum that I knew since since middle school. So, I started a company called Windsurf and later that that a lot of the folks from there ended up going to Google afterwards. So, basically went to MIT. I did a lot of systems work at MIT actually. Like I did very little machine learning work I would say at the time.

I think computer systems was like my big passion. I did a lot of work at PDOS. Worked with like Franz and Rob and and also Sam Madden on data base stuff. I did my super year up and then and then my image with Sam Madden actually. Um I think at MIT I just like performance engineering a lot. Actually, I was a TA for 6172, which is the performance engineering class two times and and really really enjoyed enjoyed that work.

Immediately after graduating from MIT I joined a company called Nuro, where I worked for four years. It was a self-driving car company. It's interesting in that like every five years it feels like everyone thinks robotics is going to make the next huge breakthrough. Um and you know, I guess I guess you know, in 2017, it seemed like robotics was going to make large strides. I think we were still a little bit early there.

Like Waymo is just now starting to hit a good amount of scale. Um but but that was because of I guess 15 years of work before then, right? So, I worked at Nuro for for 4 years. Um I think the the biggest feeling that kind of like bothered me was it was taking quite a while to productionize. And the reason for productionization was it was just there's a lot of edge cases for self-driving cars. Um and I guess the big workload that we felt was going to be big.

So, a lot of the people that that we worked with at the company was originally called ExaFunction. So, actually the company went through a couple pivots before even getting through getting to to Windsurf. But originally, we felt that GPUs and accelerators were going to be very powerful. And the the the sort of workloads that we saw in autonomous vehicles, we thought were going to be the huge workloads. Um so, we built a GPU virtualization system at an accelerator virtualization system, which basically allowed workloads that would previously run on accelerators to allow them to run on CPUs and we transparently offload the compute to remote machines, remote accelerators.

And if the accelerators died, we'd reconstruct the state of what was on that machine on another machine. We ended up doing this for a couple years. Um what ended up happening was these autonomous vehicle companies ended up kind of going bankrupt. I mean, I'm getting being very honest with you in terms of the problems that startups face. And along the way we realized, "Hey, it's probably not a good idea if you build a company and while you're building the company, the start the companies you are selling to are are going bankrupt."

Um like this isn't a good property. So, I guess at the time like all of us were programmers and we really enjoyed kind of like software engineering and we thought AI would fundamentally transform many, many industries, right? Um and we were early adopters of GitHub Copilot. Um so, we actually pivoted the company in the middle of 2022 to what was called Codium. Um I don't know if anyone here has even heard of that company.

We basically built We trained our own autocomplete models from scratch. and I guess this will tie into the future of where I think this space is going, but it started with just auto complete. We like took out a bunch of TPUs and we actually trained an auto complete model. And we also trained our own chat models, too. We were doing like basic kinds of like SFT DPO on top of models to basically get them to chat. And the idea is we sold them to enterprises and self-hosted the capabilities inside their VPC.

For companies that didn't want to like send their code outside, we gave them AI capabilities. And we had like hundreds of thousands of users using that product itself. Um, but I guess the the big worry for us was was kind of like where where I'm going to get to, which is that the capabilities were improving so quickly. So, while we're doing that, ChatGPT launches. And then we see the writing on the wall that like probably auto complete is not going to be the thing that people are going to be thinking about in a couple years, right?

This is probably not it. And in the in the beginning of 2024, we assumed that, "Hey, the models would have agentic capabilities." In other words, the models would be able to take actions themselves. At the time, the latest and greatest technology was for the models to basically be able to slightly chat with you intelligently. It was like a personalized Stack Overflow. That was what you were getting at the time, right?

But instead, we thought the models would be able to actually take agentic actions on your behalf. And we bet we kind of like bet that the system would be able to do this. And that's why we actually built our own IDE, which is Windsurf. Right? We built our own IDE because we thought the surface would need to get reimagined if agents would need to be modifying your code. Right? So, we forked VS Code and built built an IDE.

And we were the first kind of agentic IDE. Now, after that, this became a very obvious decision and many other products started uh started doing it. But I guess the the biggest takeaway for us was we thought the capabilities of the models would go through a like this crazy exponential curve, right? If today was auto complete and tomorrow was chat and the next day was agents, probably in the future what's going to happen is people are going to be managing many tens of agents in parallel um all the while like, you know, the the developer's going to spend less and less time looking at code.

And obviously, we went through this process of of of moving over to Google, like the core R&D team ended up moving over to Google. And and we built this product Antigravity that basically summarizes this, right? Like like developers are spending We have an agent manager on the product, and developers are spending more and more of their time purely in the agent manager multiplexing between agents rather than even going into the IDE.

Like the IDE is becoming a thing that people are using for debugging rather than the default state. Um right, in in a lot of ways. So, by the way, there are a whole host of other problems that come because of this because imagine you have so much of IDE code at software, and and people cannot even understand what the code is doing. So, there are a whole host of problems, but I will just end with one thing before before turning it back to you, which is the the kind of crowning moment for me in the last couple weeks was Linus Torvalds recently actually like built a repository, and recently he said he IDE coded it with with with Antigravity.

So, he actually like built some his most recent Git repository he did. And I guess the the thing is is not like a flex at hand. You gravity's a good product. I think it's more so like where is the industry going? You have a person who was a staunch advocate that like I need to use Emacs, right? You will never see him using an IDE. And he skipped auto complete, he skipped chat, and now he's directly using Antigravity to IDE code applications, right?

So, if like the best programmer of all time is doing this, we probably can see where this is going. Um so, yeah, I'll just I'll just kind of end with that. Thanks. Now, this this is really this is really interesting. And you know, maybe to to build on that. So, you know, I would say even over the past 6 months, right? We've seen this dramatic improvement in the capabilities of of these models, right? And of the the programming tools more more broadly.

And and I wonder if you have some some insight as to you know, what has gone into that, right? Is it simply, you know, the models got better and everybody else is just uh cruising on the rising tide of better models or you know, what does it what does it taken? Yeah, I think you know what people will be very surprised by is it's it's not a simple answer, but I think the data kind of is maybe helpful. So if you were to if you if you go back to early 2024, there was this kind of like this product Devin.

It was mostly like a Twitter proposed that hey, we're going to be able to do software engineering is automated entirely. That was completely overblown and for for the most part, but the key the key benchmark that they took was this benchmark called SweBench, okay? And people here might have heard of SweBench, but the idea here is this is you take you that's a 500 problems, you take some popular repositories and issues there and you actually see if the model is able to solve the issue, right?

And the way you solve the issue is you you take a look at the code that was generated, you have a golden unit test and you see if the unit test passes, right? And at the time, people were very excited about 15%, right? Now the models have gotten to north of 80% and I'm very sure by the way, like even of the north of 80%, some of the reason why it's not even higher is because the problem is itself ill-posed, right? It's a vague problem statement.

You actually look at some like a couple of examples of SweBench, the unit test is even not a good unit test. So there's some there's some weird stuff here where fundamentally like like we have cracked this. We have cracked this benchmark. I think there's a couple of reasons why. So if you were to just look at like just the the LLM space right now, there's this whole thing of pre-training, right? And pre-training, you would think it's only scaling, right?

Like we are scaling it and 100% every lab is scaling, you can see there the capex spend that's going into this, the amount of [clears throat] spend going into pre-training is increasing, but also algorithmically we are making large progresses in every single one of these axes, pre-training, post-training, like every flop is taking you further than it was the previous generation. You're seeing something that that like much faster than Moore's law even in how quickly how much efficiency you're getting from every flop.

And and the terminology used in the labs is computer efficiency, right? Which is which is what is one flop getting you today compared to what was one flop getting you like a year ago. And you know, obviously we have folks like Noam Shazeer who invented the Transformer, at least at Google DeepMind, that are just spending all the time thinking about how do we get maximal intelligence for the amount of flops that we sort of have.

So that's like on the pre-training side. On the post-training side, you can imagine that task complexity is also increasing and the amount of reinforcement learning that's happening is also increasing, right? Like if you were to think about it, pre-training, you it's like a form of imitation learning, you're kind of letting it memorize the way humans think. And then with the reinforcement learning, you're giving it real cases and real examples.

And as you ratchet up the task complexity, the model's coherence is increasing. But there's like a lot of other details that need to get done right correctly, right? Cuz these agents themselves, like there's so many problems that need to get solved that are not just like increasing scale. I'll give you a a simple example here. Let's say you want an agent to like to migrate an entire code base from one language to another, right?

This probably takes many tens of thousands of steps, and the context lengths of these models are not long enough to do this. So you need to come up with mechanisms, whether it be sub-agents, uh context compression, all these other things that need to also make improvements to hit this. So if I were to say this, there's like many many exponential curves. There's a pre-training exponential curve, maybe there's a post-training exponential curve, there's a tool use exponential curve, there's an exponential curve for agents.

And and all of these are sort of getting ridden up concurrently. Right? So I I know that that was not like an a clean answer, but that is the fundamental reason why in the last like year, I think ever since it was obvious, and I think in middle of 2024, that agentic capabilities of these models would unlock tremendous amounts of value. And I do feel like when we started was one of the first applications to kind of do this.

Um it has become an obvious sort of li- direction that all the labs are pursuing to kind of like continually improve. Yeah, I know, and it makes sense given the scale of investment that is going into this, right? That it's not going to be one simple answer, you know, one small trick that that makes everything It's the same thing as treated Exactly. Treat it similar to what was going on with Moore's law with Intel, right?

With the whole tick-tock phenomenon. Maybe it's like much more compressed because obviously you can make a lot of these improvements in software, right? You don't need to tape out a brand new chip of sort of every time, but it's it's a bunch of things that happen, right? You you know, maybe you shrink down the transistor, but not only do you do that, the ALU itself you you get some optimization there, right? So, it's it's like a you're stacking on many many gains to to see the each improvement.

So, given where we are today, what do you see as the the biggest gap, right? The biggest open problem that that we really just haven't got a clue how to how to solve and you know, might be cracked tomorrow, but also might to be you know, we might be discussing it five years from now still trying to figure out how how to get around it. Yeah, I think I think one of So, there are maybe a couple things. Let's go Maybe one thing is like just very tactical.

I'll give you one thing and then I'll also say one thing that you know, I I've spoken to Demis a good bit about this is is his perception on like what kind of what AGI actually is and why we're actually you know, I think some companies again that are probably trying to make claims that hey, we're a year or two away. We're probably actually you know, the the the more kind of like commonly held view by folks like Demis is maybe like five to 10 years away right in in which by the way already feels very aggressive, right?

Like let's let's imagine what that actually means. Um But okay, so first thing obviously for the models is it doesn't feel like this model learns the same way someone like you know, Isaac Newton did to make progress. And what I mean by that is like first of all, there's this ability to kind of like make a deductions and then from that almost take those deductions and the next day wake up and use that to pursue it to make much much deeper sort of insights, right?

This idea of continuous learning, right? Like like it feels like we don't have this, right? The models can go on and on, but the underlying value function of the model or the the the knowledge of the model has not been increasing. Now, you could argue maybe the current paradigm already can support this, right? Like I'll give you some examples. Maybe all it takes is the model needs to get very good at distilling its knowledge and retrieving it very efficiently.

That's maybe one approach you could you could possibly take. Another approach you could possibly take is maybe as the model is kind of operating, it actually does need to change its own weights, right? Maybe the weights of the model actually do need to get changed dynamically. Maybe the right way to operate it and maybe an example is let's say I'm a student at MIT versus I'm a student at Harvard. Maybe this is not the best example, but let's say I'm an engineer at Meta versus Tesla and I work for Mark Zuckerberg and instead I work for Elon.

Probably the way I succeed in one organization is different than the way I succeed in another one. And maybe you could read some some distillation for probably actually there's some internal like just like value function of the model that needs to fundamentally change if I'm a Tesla employee versus I'm a I'm a Meta employee. I'm just giving you an example here. And and I don't think I don't think anyone has like fully cracked what that actually looks like.

I think at a higher level sort of what what Demis's sort of viewpoint on on on what feels like we don't have is how do we make the model ask the harder questions, right? Like right now it feels like the model can make deductions that are complex but within the realm of like what we understand, right? Like I never look at it and I'm like wow, that's like a crazy deduction you kind of made. Um there's I'll give you an example of something that it's like already amazing at that I've like sort of been telling the team we no longer need to do, but that jagged form of intelligence where you you have like superior retrieval and you can make the deductions that all of humanity has made, but actually ask the one one like most important question to solve humanity's hardest problems.

It doesn't feel like we are there yet, right? Um and I don't know I you know, obviously all of us are thinking about this all the time, but but this this could be another breakthrough that that that needs to happen for that. So, one of the things about you know, one of the big questions that I think we're going to be contending more and more in the next few years is how to collaborate with this with these tools, right?

And you know, I think that the tools themselves have made a lot of progress in terms of how they integrate into our workflows, right? And everything from you know, monitoring our GitHub repositories to you know, now even uh integrating with Slack, right? And and reading your your messages. But you still don't have the sort of uh relationship that you can have with a colleague, right? Who you know, you go out for drinks after work and you build up this shared language and these shared assumptions and it makes communication efficient, right?

I don't know, I I'd be interested in hearing your thoughts around this question of of the interaction model and and what do we need to do to actually get the the models to the point where uh we can collaborate with them in a more human way. Or maybe this is the wrong question, right? Maybe we shouldn't actually be thinking about it. No, I think it's the right question, which is like what is fundamentally the reason why like the models are not able to even do simple tasks sometimes.

And I think actually like what you said was a little bit of the reason. You know, there's coffee room conversations between people. And there's so many conversations that are not actually like within the realm of what the models actually see. Not to mention actually even if the models can see it, there's such a large space of things that might need to get retrieved that that itself is like a fairly complex problem. Granted, we could we will solve that problem.

Like I don't think this is an unsolvable problem. Like I think with the with large with a enough reinforcement learning and enough enough sort of training, I think this this kind of problem actually can be solved. But this idea of not every piece of information is actually distilled within knowledge that it has access to. I think it's like a big piece here. I think what's going to fundamentally happen is in spaces where AI is critical in improving productivity, people will spend more time making sure that the not that the model has the information it needs.

I'll give you an example of something I was trying to do recently. So, a lot of folks we've gone through the process of trying to automate work, right? Automate tedious work on the team with with antigravity, right? Antigravity um and and automations, right? Around the antigravity agent. And and what we found was we were just like, "Okay, fine. There's a Google Doc of of how you can like run our evals." Right? And we we use this Google Doc.

The problem is the Google Doc is actually ill-specified even if you have the Google Doc in the entire codebase. Right? You actually need to work with the like if you actually need to work with the model to to create a more machine-readable Google Doc because there are parts of it that assume conversations that we've had that have not been recorded anywhere. Right? And I think that effort that work for us to do that the the the sort of work we do to do that is valuable enough now since it's going to eliminate a lot of boilerplate work we're doing on the team.

Right? [snorts] So, we will do that upfront work. So, I think what's maybe my point to that is if the work is very important, I think organizations will make the effort to make sure the agent has the information here. Right? It's it's too valuable to not do this, right? Like we have folks like I don't know, I'll give you an example for the eval system on why it's so valuable. You know, the eval system will break for a variety of reasons at at the scale it's operating.

And and I don't like, you know, we have people that babysit this, right? And it's by the way, the the reality is like there's so much more we would want them to be doing. This isn't a thing where like we're trying to cut the size of the team. We're actually only hiring more people on the team, right? Um but I think I I don't know if that answered the question there. I think, you know, if the value is there, I think we will find ways to get that information to the agent.

Right? And I don't I don't know how dystopian this future looks, but it does feel like some of these companies like OpenAI and and so on and so forth are trying to also build devices to to kind of like listen to everything. Now, I don't know I don't know exactly the if that's a form factor people like, but, you know, some companies are trying to solve this even. >> [snorts] >> Yeah, I know. I mean, I think we have uh you know, there's uh um we have evolved to generally uh like interacting with fellow human beings in a way that uh uh you know, interacting with a box doesn't quite uh tickle the same uh the same nerves uh but uh but yeah, I mean, I think if uh if I understand what you said correctly, the uh uh you know, this this question of uh of communication and getting the right context is is a big one and you think that in specific domains that are sufficiently valuable and important, you can just train that background knowledge into into the model.

But uh especially as you move to more niche domains or newer domains, there is always going to be uh you know, in in some ways the model is always going to be working at a disadvantage, right? It's uh like the the guy who never gets invited to to the uh water cooler conversations. Exactly. I think, you know, some spaces that where this is just fundamentally a lot harder is you take something like the hard sciences, right?

Where the iteration cycle is a lot longer. This is actually like, you know, the reason why you can do all this stuff is in in software, you're much more willing to just do do brute force. You can just do 10 10 times as many things, right? Now, granted, if you're hardware limited, that's that's true, but in a lot of ways, let's say it's running a unit test until you can just the cost of trying something is is low, right?

[snorts] And the interesting thing in the hard sciences that like that means the value of your intuitions it's still high, but it's less valuable. Because brute force exists, right? So, and and not to mention the models are getting smarter and and I think anyone that thinks the models are just brute forcing now it's they're putting their head in the sand, right? Like the intuitions of the model are improving. But in the sciences, there's so much information locked up in the head of the of like the the scientist and the cycle time is high like the cost of like using up the lab for like the next 3 months is high.

Like you don't want to just randomly do it, right? So, some of these This is why like different domains will have different times at which AI is going to kind of like help a lot. I think in the hard sciences going to take longer. Right? And that's why we're seeing so many gains in like the bits world rather than the atoms world. Yeah, I know that that makes sense. Now, you know, we in the in the academic side, right?

Traditionally, sort of our our job has been a to produce research and to you know, ask questions that you know, maybe are a little bit beyond the horizon of of industry, but also to train students, right? And to make sure that that they know what uh what they need to know when they go out on the workforce and to to be productive. And you know, recently there has been a lot of conversation around how both of these might be disrupted by by AI, right?

And so, I'm curious how you see from sort of your your side in in industry, you know, on the one hand of uh insights and results between industry and and academia as well as you know, maybe the the perspective on you know what what should uh students coming out really be knowing? Yeah, so I guess we're still hiring like new grads from MIT. So I think I think that hopefully that that kind of explains that we're not saying hey like you know people coming out of college are not useful anymore.

I think the expectation though is fundamentally been that you're going to be like much less productive unless you use these tools, right? I think this is not like a very interesting insight. Um people that understand the fundamentals and can use these tools are much much better. Like I think you can clearly tell when you talk to someone if they both understand the fundamentals and they're using the tools as a way to accelerate accelerate their work, right?

Um I think I think it's it's it's clear to to me and I think it's it's basically like what I would just say is like the playing field has has basically just shifted up, right? Obviously you are no longer on the playing field there. I think if you don't use these tools anymore, but the person that was like very knowledgeable before these tools is also like strictly better than the person that was like not knowledgeable in using these tools now.

I don't know if like that that should be like um I I don't know if that that sort of answers that um that point there. I don't think this is a case where where people should be stop should stop learning the fundamentals of kind of computer science, right? Cuz I think the fundamentals of computer science is just the fundamentals of logic, right? Like in in a you know I explained the world in which you're live coding everything.

Like if you build a critical application anywhere, you're going to need to understand the impact of the software you do, right? Even when you talk to AI, very quickly you will notice that sometimes it makes trade-offs that don't make sense. Once again, because actually it doesn't have all the information in your head. And this isn't like a bug. It's like there are trade-offs that were made in a coffee room conversation and and and you actually need to optimize for these trade-offs.

It's actually hard to specify every trade-off under the sun. There's like some trade-offs that your co-worker explicitly thinks that you understand, right? Like when they say hey let's rewrite this to use X Y and Z." They're not meaning let's rewrite the entire codebase. Right? They're probably just meaning like a a small chunk of of of software kind of needs to get rewritten, right, in that case. Um so, I would just say like, yeah, I mean, the the students should be like definitely be using the tools.

I guess this is a hard thing from like a piece of standpoint. Like, I guess this is a this is a that's a sort of a tricky question. On the research side, I mean, this comes back to the whole jagged intelligence piece. I don't think the models are in a state right now where they are actually doing fundamental research very well. Right? Like, they are actually not able to do that. Um I granted I think a lot of research work, again, is is probably like uh is probably requires a lot of boilerplate work.

So, I would assume researchers are actually actually able to be also a lot more efficient with with AI. So, um yeah, I don't see I don't think this is going to create a lot of a lot of changes in in academia. Maybe this is like a controversial take. So, maybe building also on on some of these. So, uh you know, I think we're already seeing a lot of uh change in how uh you know, the typical developers they looks like with with some of these tools, right?

Uh uh you know, extrapolating maybe 5 years down down the road, how do you see uh you know, the the typical developer uh work day uh being uh different from from what it is today? Yeah, I think it's probably I mean, the expectations on productivity is just going to be a lot higher in terms of like the amount a singular developer can do. I think this person is managing like maybe [snorts] maybe somewhere on the order of 10 to 20 agents in parallel.

And the reason why they're able to do that, some agents are very, very long running. We do need good interfaces to actually communicate back with these agents um when the agents do intermediary work, right? It's probably a bad idea for an agent to go and do work for 3 months and and and, you know, I mean, this is like any any report, right? You have a report and you have a team of people and you do work for 3 months.

If the work that's getting done at the end of 3 months is something that you don't understand and thought was the wrong direction, you probably want to check in with what's ultimately happening. But but yeah, I think ultimately like the job of a developer is to have taste on what needs to get built. This is, by the way, the hardest problem. I think this is actually why, even though AI is accelerating software development so much, it isn't that companies are getting faster at the same speed.

It is actually a complex decision to decide even what to build. I'll give you an example. Let's say I'm building a user application, and the user application faces some users. You could build like a thousand features, but there's a limit to how many features the user can even consume. Right? There's some like There's some bottleneck there. And then there needs to be some taste that actually the three features I do need to build should be built like this, and the tradeoffs for those should be X, Y, and Z to make the user experience good.

You have some number of shots on goal. Like it kind of becomes like the hard science exper- hard science uh kind of case all the way uh all the way there. But that is what developers are spending more time thinking, right? When I think about what software development is, it's like maybe three things. It's like building it, how should I build it, and what to build. And I think people are spending much more time thinking about what should I build.

Right? Um and I I don't know if this is like fundamentally different than in the past. Like I think, you know, let's go back to when people were writing assembly. Uh you know, the the way people build and a bunch of design decisions they made about, let's say even doing networking across machine, was probably like people were taking a lot of time doing this, and now people in Python can just do like socket.send. Right?

So, I think this is just another level of abstraction on top of that, and people will be thinking much more closely to what is the What are the problems I need to get solved to make my technology succeed. Thanks. So, I think it's a good point to maybe open the floor for uh questions from uh from the audience. Uh I don't know if you can see uh the audience uh from from where you are, but so we yeah, let's let's open the floor.

First of all, very proud that two of the best AI assistant coding tool are designed by MIT alum, Wing Serve and Cursor. Be part of the community. I would like to know after you join Google, you know, Gemini has been selected by Apple to be the default model and then also Colab has been widely used in the academic world. What's the next breakthrough or product features you guys are thinking of in terms of the core assistant tool at Google level and what kind of for example is more better human in the loop, better product design or better back end system that can integrate everything.

Thank you. Yeah. I think I think the mission the mission fundamentally I think a lot of it is the model, right? Like the product, the way we think about products at Google, at least like on on our team is and we thought about it this way at at Wing Serve, too. It's how does the product maximally showcase the capabilities of the model, right? And the reason why we think about it that way is because every generation as the model improves, if you don't think about it that way, a lot of your product needs to get deleted.

And I'll give you an example of this just because I think I think people might remember. And I think Cursor actually had some features here before the agent capabilities sort of became very popular. But they had a feature called like at docs or you can at a directory, right? You could actually do the at mention of of the directory. >> [snorts] >> And what happened is that was very useful when the models actually couldn't like do anything agentic.

But then as the models actually ended up getting agentic capabilities, the model was able to read documentation itself. It was actually much better to have the model actually browse the web because at least that was up to date versus indexing a particular version of the of the documentation. So, what I really want to just get at is the way we think about products is how does it maximally showcase the capabilities of a model?

Now, where are the model capabilities going? The model capabilities are going much more multimodal. Actually, if you look at a product like antigravity, it can even actuate the browser. And we think if the model is supposed to automate a lot of work and tedious work that developers and knowledge workers do, developers and knowledge workers spend a lot of time in the browser because they might be using a code review tool.

They might be looking at logs. They might be looking at I I many, many systems that only exist outside, right? Like even Drive and Workspace, right? And and also also knowledge workers and developers, the reason why I'm I'm distinguishing knowledge workers and developers and why I'm calling it something the the same thing is I think I think automation and and using code is something that everyone is going to be expected to be able to create, right?

As part of their work, right? Like code is like the most atomic unit of automation, right? In some ways. And the fact that antigravity can generate code just means it can automate things on a machine. That's like all it really means uh down the line. But imagine, okay, if I was to be very concrete, we're going to be able to go like much more powerful on more surfaces that includes the browser or the operating system.

The model is going to be able to be coherent over many, many more steps. And it's going to it's reasoning is also going to be able to be much, much stronger over over many, many data sources. So, I know that that sounds like uh that sounds like kind of crazy, but this is like it's just better at every at every dimension you kind of think of is what's being pushed. Also, on top of that, what I think is very exciting about what Google is doing in particular, we should ship the model called Flash.

Right? Gemini 1.5 Flash. The coding capability is actually on par with Pro. So, what this actually means is we are also trying to drive the cost of intelligence down for our users, also. That is the active work of effort that we are always thinking about is how do we make intelligence much faster and cheaper? Right? And Flash is significantly faster than Pro. So, like I know that it's a cop-out, it's better in every dimension, but like this is this is what it takes, right?

Like the sheer amount of capex, if we're not doing this, like what are we actually doing? It's cool. Um I think I think he was first. Oh. Yeah. >> [clears throat] >> Thank you very much. Um you have a unique insight on how users are interacting in the model. So, I'm curious to learn what you have learned on uh seeing how users interact with the model. It could be in form of experiments you ran that you expected that failed or maybe some surprises on how your users use some of the uh anti-gravity or or your previous products and uh ways that you've never thought about, but how it was how users adopted your product.

Yeah, I think okay, couple things that I thought were pretty interesting. First of all, we put the browser in. We thought people would largely speaking be using it for UI testing, right? In other words, you make a full-stack application, you use it. But people, as the browsing capabilities actually increase, people have actually been using it for for general browser automation work and look at like arbitrary insights that then help them write software.

So, it's like, "Oh, like open up a system design from X and Jira and then after that write some software write some code." People have actually been doing this. And the benefit of doing it this way versus like all this like MCP shenanigans is you are at least then you are no longer reliant on on on an API existing. And the reason why that's helpful is why did the people at Jira, for instance, build their product? They built it for people to use it on a website.

So, like fundamentally like like the website is the source of truth, right? The API is something that lags the website in some way. So, that was an interesting thing that I also found. The second thing that I thought was also really interesting is we have so many users that are running so many agents in parallel. I did not think we would have a person running 10 agents in parallel. Like, I talked to a user. We we give very, very generous flash capacity for our ultra plan, very.

And this person ran out and I was like, "How is this even possible?" Like it literally doesn't make sense. Like, you would have to be a monster to do this. And I found out what the workload was and it was insane. Like I I did not think software development was changing this much. Right? People were doing a migration on one code base while doing X on another and Y on another. And I think this is just the future, right?

And this is why we built out the agent manager, but it's cool to see that pull already coming naturally from the product. >> [snorts] Thank you. Varun, thank you for the talk. Thank you for building such a sensational product. I use it all the time. I just had a question actually based on, you know, I firmly believe in what you're saying. I'm actually a medical doctor and I've been using antigravity a load and building all kinds of stuff way beyond my intrinsic capabilities to do so.

But I'm curious like so my own workflow has changed a lot. I was using Codex and VSC before and now I've shifted over entirely to antigravity, but still this problem of figuring out what to build, what problem in the real world is it solving, and what am I going to do after it's built to like get people to use it, what's the business model, all that kind of stuff. I'm still doing elsewhere, and then write getting Gemini or ChatGPT or whatever to write a prompt that then I'm using in the antigravity world.

So to what extent are you looking at bridging that gap or is that always going to be two separate worlds? So so look, I think I think my feeling is there's always going to be this like kind of consumer use case, right? And the consumer use case is is an endless product free base. And the consumer use case is like I want to answer an a one-off question that is very topical to me. Like, "Hey, like tell me a little bit about this bed and I want the latest and greatest bed in X, right?"

And this is like very specific and I think Gemini app and and ChatGPT are going to be great products for this. The idea you brought up is something we genuinely care about. Like I think you will see we are going to flex more and more to be able to do more powerful things in workspace, be able to use all the data sources you have internally. Like we want this to be the the most powerful knowledge assistant, right? That guy that is able like the fact that you can write code is a feature.

Yeah. fundamental feature, right? Now, I still think you will need to apply a lot of your knowledge being an actual in the medical space to to actually figure out what problems you need to solve, but I think we want to help with with this with being able to piece together the knowledge from all these different sources. Like that is something we we definitely want to do and we will improve with it. I think the browsing capabilities at its very very early innings and frankly speaking right now for a lot of tasks it is still very tedious.

I think we have a lot more improvements to be made. This is once again, I know everyone will say this is this is the worst the technology will ever be. Right? Right now. Sure. Thank you. So oh hope if people can hear me. For me it's great to see you again. Met you in the fall. Hey, so last year I was using what we are Windsurf in the summer to I do auto complete like you mentioned the code. It's really cool to see the new product come out.

So this kind of question is about the shift from like a auto complete to the new anti-gravity IDE platform. When you think about when the agents are taking the intermediate logic versus like the developer where kind of coding the logic. How do you think about repeatability of the flow and how do you think about eval? Like is the current benchmark still right way of thinking about it or you're thinking about having to have new benchmarks in order to evaluate these agents and also making sure what the work they do even with artifacts you have on anti-gravity can be repeated.

Yeah. So so we are we we care a lot about this. Actually a lot of the work on our team is actually eval because I think otherwise we don't know like we don't know if we're actually making the product fundamentally better, right? And we have a lot of evals that are held out that are held out and actually represent the hardest task our team can do. And we the hardest task within Google the benefit of it being within Google is we have the entire environment of Google.

So like we can in theory the agent can pull drive documents and all of this stuff internally, right? And and it has a self-consistent system, right? And we also have anti-gravity usage internally. So, we actually can do a lot of fancy things on the quote-unquote determinism side, which is to say, let's say the agent didn't do something good, right? For an internal Google user. We can actually take those trajectories and replay them afterwards.

Replay them afterwards with a newer version of the agent and validate, this is more in line with what the user ultimately did want. So, we can get very, very sort of deep in the weeds. We have a lot of other sort of tasks on the team that are just very custom and specific to what we're doing, and we have made those evals, too. And they are like outside of the training regime, and we know, hey, like the thing I'm the most excited about is when we have an eval where the where the score is like 5% or 10%.

Because that just means the headroom is so high. Like, there's so much more for us to fundamentally improve and and for us to climb. So, um eval is the primary way for us to improve, and and we take eval very seriously internally. Okay, and repeat a little bit repeat ability uh of the agent. >> repeatability is okay. I think the idea of repeatability fundamentally comes through if I Okay, the idea of exact repeatability is just not possible given like the stochastic nature of these models, right?

That's like That's 100% true. But, I think you can get behavioral repeatability. In other words, you can run evals and validate, hey, across a wide swath of examples, the number of critical negative cases is X, the number of the average score is Y, and you can across a bunch of dimensions, and we do measure this, and that is like our closest proxy of repeatability. Cool. Sweet. Thanks. Uh our dog fooding question, what percentage of your time or your team's time on the computer is spent in the agent mode versus more traditional coding mode or not using it at all, I guess?

Oh, not using it at all meaning not using the product? >> Yeah, I assume that would be zero, but just asking. Yeah, yeah. So, I I think I think like literally everyone and anyone who like even people that manage people the expectation is like you're you're actually building a lot, right? Otherwise the intuition is like busted. I think we have we have a lot of paper cuts in the product in terms of like managing many agents because that workflow is like a little bit of a novel workflow.

And I think by the way, this is just to say one thing, this is why I think some of these products like Claude code have have have taken off in a lot of ways because I think managing many agents in a terminal is natural for a lot of developers, but I think you'd be surprised most people want like a a clean UI to like manage a lot of things and that form factor is something that we're constantly iterating on. Like I think we are probably the people that give the most feedback for the product and we have like a large dog fooding channel where every day you could probably see 50 50 50 issues we are talking about, right?

Um, and that's because everyone is beating up the product. Um, so I I will say like uh we are our biggest critic right now. Uh hi. So, I have more execution point of view questions. So, I understand for example when you're um when you're Surf like you can have the flexibility to train your own model like based on whatever like use cases, what I whatever like emerging capabilities you need like when you're trying to build your products.

But after like plugging into like this D-Mine bigger chapter, like what's your uh collaboration relationship with the foundational model side, especially for example I also noticed like uh browser user has become a topic because it is so efficient like on uh UI automation on the on the like a UI testing. But I'm just wondering for example when you're building like anti-gravity right now, like do you see there are some like new use cases that probably we haven't thought about when we're trying to train like the foundation models.

Then are you plan to like train your own model like to lead such capabilities? For example like probably a lot of customers right now are trying to use anti-gravity to build uh agents. Then do we have the model like to automatically spin up like a agent for for the customer like computer use agents etc. Yeah. Yeah. So, those are two questions. We are going to make it much easier to build these kind of agents you are talking about uh in in the product and and otherwise.

Like that is a there will be announcements about that. To the former is like this was the reason why we also came to Google. Like our entire research team moved over and we've even made contributions to like the past two generations of Gemini 3 Pro and and and Flash as well. I think yes, of course if we like we are contributing to the fundamental progress of the model and that's what makes me feel so good, which is twofold.

It's not that we are contributing, we are but we are also getting the benefits of the baseline also growing tremendously. Right? As pre-training continues to improve, post-training recipes continues to improve, and so many other improvements that have coming down the line, we get the advantages of that. I think one thing that I really appreciate about the way Google is doing things, unlike other companies like I like that have like 100 model skews, we are pretty much pretty clear about the fact that we want a very targeted set of model skews, and we want a generally intelligent model that that is great across the board, right?

Like some models like obviously like Anthropic, they have taken they've made the decision that they are going to they're going to regress on a lot of evals that are not that are not coding related. But I think at Google we want an across-the-board great model and we think a model that is ultimately the best assistant is one that is like capable at at many many different spectrums, and we're going to soon see this with computer use.

Multimodal capability, we don't just think it's there to make cool images. I think it's actually a core capability of a model to be able to understand images. It's actually very important if you want to be great at browsers. So, couple things. A, yes we are contributing to the model, and B, it is actually very valuable that Google as a whole is is so is so diverse and and sort of strong in many many realms in in its model.

And I think this is going to become patently obvious. Like okay, maybe just to explain explain like a last point here. When I look at like the software development life cycle, there are many things you do. You you write software, but that's only one piece. You review software, test software, debug software, deploy software, right? Design software. Yes, you do so many different things and I think this requires you to have reasoning capabilities on many surfaces that are not just like your underlying code base.

Like it's quick like we we got the quick wins of being able to quickly write software. But if I were to say this, if I were to take every developer and instantly allow them to write software, we would not save software development time tremendously, right? And I think this is where we're going to start seeing the advantages and I really hope Gemini Gemini's able to like push forward here and I'm like very very confident in sort of the direction we're taking right now.

Thanks. So, I think we are out of time for this for this session. So, we're going to have to jump to the next final session of the day, but let's thank Varun for spending some time with us. >> >> Thanks a lot, guys.