hi listeners and welcome to no priors today we're talking to Orel vineel the VP of research at Google deepmind and Technical co-lead for Gemini his story career in machine learning includes leading the alphastar team which built a professionally competitive and pioneering Starcraft agent all the way to today and we're really excited to get his historical perspective on where we are in machine learning welcome to the show Oriel yeah amazing thanks Sarah for the invitation and like thanks a lot for for hosting me yeah thanks for joining last year was an eventful year at Google in Deep Mind um you know how is that research effort organized now and what do you what do you think of the mission as internally yeah so sure I mean I'm happy to obviously discuss the different phases that research uh organizations have gone through in the last many years but focusing on last year two major events happen one was that uh the Gemini Project was formed um as a result of having two sort of parallel efforts on llms uh mostly led by uh Google brain and and what we now call Legacy Deep Mind so uh earlier in the year there was uh an effort to merge the two the two projects and that's when sort of Jeff Jeff and I came together and brought the two teams together to create the very first Gemini model which was eventually released later in the year then the second big event was to uh take the all the organiz izations uh that were doing uh AI research or AI research and also form a sar organization that's what today is called Google Deep Mind um and it comes from uh Google brain and Legacy deep mind coming again together Under One Roof obviously Gemini being a very large and very important project within that organization um and really the goal uh of Gemini itself is to create an awesome core model to power uh the technology that of course llms today are powering all around the world and we obviously expect this to un increase how do you interact with the rest of the company and like Google is a business and like I I feel like I have to ask you know does AI replace traditional search so even running that from a research standpoint um is super interesting right there's there's um two major centers one in California one in London given the organizations that we come from so that in itself is very interesting in a way we we have the project run in 24/7 which is helpful when you train these large models and then you have to do a few things right one of the things we do of course is trying to build state-of-the-art technology showing from sort of a research knowing where the field is coming from and when it's going to trying to really um showcase from our our own sort of intuitions and ambition what might come next right so a prime example of these was for example the long context that we released earlier in the year right millions millions of uh tokens now are been able to be processed by by our models but then of course we also um sort of take into consideration all the different needs right from the different products that we work with Google has a lot of product areas so we try to focus of course initially especially to form the project we try to focus on critical projects and you see that very much um by how Gemini is first surface to to users or to Enterprises right so obviously Cloud um and Enterprise is very important developers as well uh super cool to put these models in the hands of creative minds that are going to do things you you didn't even anticipate these models could do um and then very important Formerly Known B Now jini app which is sort of the chat bot surface of our models and then maybe the last uh very important piece indeed is is search which uh is trying to integrate of course this technology into their product um and of course has a lot of users so it's extremely exciting to to think well the decisions you make at modeling uh eventually and eventually means just maybe a few couple months after or so we'll make it into into the users that maybe are signing up for a beta Etc so super exciting and and it's obviously connected it's the core of of the company really especially for the products that require very intelligent um AI systems like the ones we're creating today how do you think about um the various types of use cases that fall under chat based model versus search based models cuz I remember I was at Google many years ago at this point and at the time a lot of the different types of search queries were um broken into different chunks there's navigational queries you're trying to get to some other site and it just kind of helps direct you there um there were uh sort of strong intent Commerce queries you know you could kind of break it down there's medical queries and so you can kind of map out the world of different types of things that the user is actually trying to do with the searches of the interface to get there what do you think will move more towards chat and what do you think remains in the domain of more traditional search based approaches we might not know the answer like we're sort of experimenting in a way right so you can can kind of think an llm first experience like that's the chat bot right so so in there search has a role to play because it can be seen as a tool that enhances the the answers um of your chat experience right it provides citations uh a bit more Reliance we know that of course language models can and do hallucinate so there's that sort of llm first point of view which we are building up sort of from it's kind of more of a new product and then searge itself which is obviously embracing sort of enhancing some of as you said different query types that could be useful um to be sort of do a bit more like summaries like AI summaries but there's a lot more coming I think IO actually showcased this year quite a lot lot of a good breath of vision of what search is trying to do with with language models right so some of it um is not ready but it's being tested and and and obviously like feedback is very important but I think for now it's hard to to think of a convergence even or that one sort of dominates the other I think both right now seem useful in different ways I as a user you use both certainly uh but but I think what is very though is that um if even if search is sort of the initial point where you have a query and you want to research something um it just it just feels like that that experience uh will tremendously be enhanced by these models so the search product itself is trying to figure out right how to integrate uh the llm answers or their capabilities uh reasoning capabilities Etc so I think we're going to see a lot of that and I I kind of call like from search uh integrating l and of course that vice versa is very obvious as well and then as a core model project like Gemini Project is um we have an open mind we don't really need to you know decide one or the other and one of the things of of operating at the scale that Google operates uh is that yes we are the research team that builds the models but then of course the Pas um are the ones that are going to drive the strategy of course influenced by what the models can do and with input from many many of us that have been dreaming about this world but um it's kind of cool to also um have the very experienced folks right sort of iterate with their users and their use cases that makes a lot of sense I mean I think one of the things that's very understated or forgotten about Google is the degree to which it really was the world's uh first AI first company or ml first company as it used to be called right I mean the entire product both in Search and ads was very AI driven in the very earliest days and part of that I guess is also inferring user intent through action and then algorithmically adding that as an outcome and so I guess to your point you could continue to use search as a primary UI or interface or entry point and then if it routes to more of a chat based thing or a chat bubble could pop up or something else it could just happen organically based on the type of intent that the user ex is exhibiting and so um I think you're raising really interesting points about the capabilities of of Google and um and all the rest of it I guess related to that what has been the most surprising thing about how users or companies have been interacting with g I yeah there's quite a few right I guess um maybe the ones that surprised me the most because initially I even thought this was just a number that you could report is is the fact that you know infinite context length is coming eventually um so I thought look I mean this is interesting right we come from a world where we had recording neural networks and lstms that actually had infinite memory although it was not very capable right you the models in in practice they never remember more than a few hundred words or so so that was kind of first that we could make the context length so long and then seeing the use cases just emerge even internally when we were first just trying the model that I don't know that that seems very trivial now in hindsight but you know putting a whole one hour video and just just ask anything and it feels super human right you just literally put the video in and after 10 seconds 30 seconds I mean it does take some time to process the context but you just can't ask anything right and I mean thinking of computer vision as a field or video question answering some of these data sets that we come from I mean they all that seem very dwarf then compared to the capability that was in our hands and then we put it in the hands of developers and we saw like I mean amazing obviously demos and and things that people could do um even as mind-blowing as you point the camera to the screen directly and not even the code right it's it's just the letters that appear on the screen and you can debug like that right so you can imagine how future interfaces will be effortless I mean you just need to point the camera ask a question and you get answers it's an interesting research problem of course but maybe it's not going to be that useful to now thinking wow that's amazing yet it's not very mainstream either right so we are kind of still trying to discover right what what does this enable we showed project Aster where of course memory you get the phone and you just interact with it as as if it is an agent memory is very important but it's still not very clear in a few years what this might be although you could imagine well the whole web is in the context or um all your personal data because it's remained in in your device because it's the working memory of the model rather than the weights in lots of applications but still fairly early days right so it surprised me yet of course it it in a way it hasn't taken off fully um gone fully mainstream although it does feel magical when you start interacting and realizing you can ask anything you get an answer like this without watching any of these movies or um many books you can upload Etc what do you think is a time frame for very long context Windows really being in Broad based use not not this is not a specific to Google question but it seems like multiple people are on a trajectory to add this at scale and you're also seeing it crop up in biology models and other things like that so I'm just a little bit curious in terms of the time frame for which it actually hits either large Enterprise or consumer use cases at to your point you could imagine ones where the company is doing something on behalf of the consumer by adding all the context uh from their device or from their interactions into the context window or it could be an Enterprise uploads a large folder including a bunch of legal documents and then it gets incorporated into some query or something if there is a comparing use case the technology is not far away from being able to be deployed that scale and of course Hardware is also being updated based on what I guess the research developments are right so um certainly like one two years um this might be quite the context that is a commodity will be definitely enhanced um by a factor I mean I don't know 10x or so everywhere and then I mean extremely long context uh I think it's going to be a motivational drive for for definitely from a research perspective and then deploying it that scale some techniques like many have been explored already like hierarchical memories and so on um and even rack is pretty common right like um so we're going to combine this and probably thanks to the use cases um I expect I mean order of one two years you might see wow it we went another order of magnitude from both state-ofthe-art and of course what's considered commodity so I'm pretty certain about this that's again modul finding use cases that will be compelling to serve the model that requires I mean more memory there are some certain limitations of course uh but these will be figured out from a technological standpoint if there's the motivation for sure in the um very quickly coming era of infinite context like what is the relevance of retrieval architectures and and more hierarchical memory um and like I you know I think you can make an argument for this being continually relevant just from an efficiency perspective but you know what how do you think about it yeah I think definitely the efficiency argument to hierarchical memories to make context even longer make a lot of sense and even from just efficiency of learning and and of course of of retrieving the memory um in a sort of course toine manner like like you know an intelligent being like ourselves might do make sense so I think that even the quality uh will motivate this sort of solution regardless and we do have a lot of experience of course of retrieval based methods definitely Google and then combining them with neural based methods U I think it's a matter of time and finessing the details and the use cases of how much uh the the problem with of course retrieval based methods is they tend to simplify uh things to say hey like this whole book is just a single Vector um whereas if you just upload a whole book into Gemini and ask questions it can really reason about every single word right so probably finding the middle grounds for different use cases um is needed but I think to me it seems like a featured not a back that we do have a bit of a hybrid mode perhaps going into the future and research will be driven like this how do you contextualize this moment in time just in terms of you know what the biggest limitations are for current State ofthe art llms and like what's worth working on I mean there's one reflection that even many years ago with friends who were kind of early quote unquote in the game they said well get ready lots of billion people as this gets mainstream will enter the field and you certainly see this right with open sourcing and a bit of a random search EV right it's not like I mean you just selection bias like someone does something random but people actually want that and then that becomes sort of viral in a way so I think there's this the sheer size of the field that is one aspect that I think we were sort of anticipating but to me that's one of the biggest changes that I've seen that there's more brains more different backgrounds coming into the field and that is combined I usually tend to assign credit uh what's with what has happened to of course the scale of data and compute algorithmic advances that you you can simplify but there's certainly been some that have been important in the last 10 20 years or so and then actually the accessibility right the software the open sourcing efforts um those have been quite critical to then create these sort of exponentials or linear Trends in log scale that we're seeing now a bit more into sort of how I see the field from maybe like I I tend to call maybe the 200 let's say 10 to 20 like deep learning era so what that era did right is it took a set of algorithms that were General right the the algorithms are are like stochastic GR in the s deep learning neural networks right reinforcement learning and you could think of these are ingredients they're common you expand them over the years but they're certainly uh the same and then you just apply them to a domain and you get extremely good uh at that domain right so we we have you know mastering the game of Go beating the imag net challenge becoming state-of-the-art speech recognition state-of-the-art image Generations right so so that decade is kind of the models themselves C are not General but the algorithms are General right you could just take the same algorithm and then change the data set and of course tweak a couple of things and voila you get like protein folding um really enhanced right from the traditional methods and I think then the greatest Insight of course came from realizing that and I think that's a lucky factor for us because it makes communication with these entities much easier but modeling language turns out to be such a powerful abstraction for generality so of course the GPT uh two paper especially post that sort of in the abstract it says look you can solve every task not one task by modeling language and then perfecting that the whole field of course building up from a lot of years of of research created what is not just the algorithms are reusable but actually the model now becomes General so that's why I think like okay AGI is getting closer that you know we have powerful models first in 2010 20 now we have General models and I think multimodality has been more recently another amazing breakthrough that these techniques expand not to language but also to to Vision sound videos Etc so that means that we have very powerful General models that from a AGI definition standpoint it starts to take many boxes reasoning capabilities of the models are there but I don't think we've perfected sort of making the reasoning very crisp and accurate so that these models would not hallucinate or would not you know the model might solve an you know a Olympia mathematical problem yet failed to then discern a very basic puzzle right and I think what you do with the model uh this reasoning step um there's a lot of ideas and a lot of experience and a lot of algorithmic advances we've also done in the last few years over like Surge and so on but we haven't quite per effected this and then of course the question is you know push the frontier and and move forward in certain domains at least but yeah we've come long way I think um the investment and luckily these models are finding usefulness so the resources that go into the training the models um right now there's a good sort of uh feedback loop right of Revenue and then reinvesting certainly from the biggest players so as a researcher of course I mean we welcome that what is the difference between having reasoning capability and having reasoning capability that's crisp and accurate sometimes the distinction between probabilistically solving something right so right now you know you get these models they they assign they still assign probability uh Mass over every sequence of let's simplify and think of not multimodal but words so every every single sequence of words it will assign a probability distribution over those so then you of course are absorbing all the knowledge on the internet and then sharpening those models around being following instructions being aligned with humans but you still have this probability distribution that will assign nonzero probability to certain things that would not deem to be correct although in language of course there's so many ways to say the same thing correctly so that's why these models shine in the end of the day they are very efficient ways to integrate over all possible sequences right now it's POS possible that let's say now you predict um when it's a hard problem right a tricky question that requires deep knowledge um you might you might be at 95% accuracy but of course these will create errors even if it's small this is deployed to the whole world so certainly you will get to see the the mistakes so one thought would be like you just keep making the models larger you keep improving the algorithm and you're going to hit a point where the probability of a mistake vanishes it's possible we will obviously explore that but to accelerate that that sort of progress you you want to start sort of really exploring what's the reasoning the model has and by making it more redundant more logical by iterating more on these kind of ideas you could imagine um generating a a very small progam program right that runs slowly with the language model at the center and then getting that 99. many ni% faster and of course you you're going to do both as as an ambitious like lab and so on so but but that the crisply means that this probability of error diminishes and of course we can always put more compute power we humans will make mistakes we get tired Etc but I mean these models are powerful we can put more hard work inference so the hope is that in that sense they become at least as good as humans you know one way to like frame this problem is that even Deep Mind or any large lab or you know the human race we have some limited amount of compute to Port toward this problem right and like there's there seems to have been a strong shift of how much of that compute should be scaled at training time versus at inference time with you know test time search or some of the um techniques that you describe in in system to what is your prediction of like what that mix is of compute at training versus inference um inference time compute uh you know let's say two or three years from now what we're kind of all aiming to to discover is to make the bitter lesson from reach out and true right the bitter lesson states that you skill learning and you skill search and then that's all you have to do as a computer scientist I mean it's it's controversial I certainly like simplicit I'm a deep learner at heart so um I don't disagree with that certainly the learning scalability part has been tested and proven quite quite heavily recently the search at least when you do not have access to a perfect reward which is um this is the current I think current problem we have in language that is you know one of the key research areas right how do you assign reward fily to statements that I mean not even again you and I might agree to be true right right I mean is the sky blue I mean I don't know at night it's not right I it's it's quite an interesting kind of point that um assigning truth or one or zero which is required in games um is not so applicable here now historically if you look at Alpha go which actually followed quite closely the recipe of you pre-train your model on all human data you then use RL to make it better and then you do some search at inference time time the compute there was very skewed for the middle step the reinforcement learning step right so I don't even know the exact numbers but certainly the majority of compute was spent on the training of the selfplay as as we called it at the time uh self-improvement loop and if you look at today that's clearly not the case most Compu spend in pre-training and in fact you will see over and over that you over fit so you have to actually stop the reinforcement learning process to overfit to these imperfect reward functions otherwise you start sort of doing a bit of adversarial search against a reward function that might be imperfect because um you have a data set of human preferences and all of a sudden the model might discover hey you can issue lots of emoticons and this reward function think that's great right and clearly like we have a problem of reward functions not being as accurate as in the game of Go or chess and whatnot and then there's the third component which is now you have your model trained how much do you let it Ponder which in in alphago again using the same example well the rules of of course say don't quote me but let's say roughly a game must last four hours of compute time from Human perspective so we know we had quite a limited inference time because we obviously couldn't go over time so of course there was parallelism and so on involved but the compu there was certainly not as big as the one you used to train so to me that balance feels correct like some on pre-trading and here we we're trying to learn every task so certainly that's going to be you know let's say it can be as high as 50% not as high as over 90 like today and then the rest mostly on reinforcement learning or or if we get access to good rewards that would be kind of the next maybe piece of compute but much bigger than it is today and then inference time I mean system to is a low but it's not terribly slow so unless you ask the model I mean solve protein folding then probably you can go for a vacation for a month and the model will come back with a solution um I think a few seconds of compute is okay uh for inference so that would be the rough like so that's a small percentage probably compared to the to the amount you spend in training of course you're serving many queries for billions of people and then that means research is needed especially on the middle buet of uh reinforcement learning step that currently feels still reasonably in the early stages of research how do you think about scaling the reward function Beyond games once you really hit superhuman performance of the model I'm thinking of things like the older like Med Palm 2 models and things like that where they outperformed human physicians in terms of output relative to physician expert panels and then obviously you could then do some post trainining with physician experts is that is the key but at some point the machine will be better than that and so how do you keep scaling reward functions traditionally right you you just get it's supervised learning right reward function means good or bad so we can scale that process as much as we we have so far um I mean obviously many many players are are realizing the power of human annotation and in fact deep learning comes thanks to amazing like fath fa and lab effort to to label a data set of a million examples right so so that way of scaling is one uh but then I strongly believe that there might be a bootstrapping effect of the models that become better at judging their own outputs right and so maybe and that really is probably maybe even the main hypothesis of reinforcement learning as I see it I mean I'm not a huge expert in RL but if checking that something is correct is easier than creating the solution then we're in business because the language models will be able to evaluate their own samples more accurate than to generate them and then we have a sort of reinforcement learning Loop because we can reinforce the ones that seem more promising and then the model gets better right so that using the model itself as a reward um which incidentally uses language which is already fuzzy is one area that I mean I'm excited about there's a leaderboard of reward models um some of them are I think the the name that they use is maybe generative reward model think that area goes beyond this kind of need for specific task annotations we might need specific task annotations then the question is how many labels will we need and the hope is that in the limit maybe you need as many labels as the user will provide the system when the user wants to teach it something new not to abuse it but I you know a friend of mine used to use um nyis Channon sampling theorem as sort of a proxy for how much an intelligent person or machine can actually extrapolate the intelligence of something smarter than itself and you know it feels like you're almost falling into some version of that where you know I think the the basically states that um you know you have some frequency on a wave and you're sampling it and you need to sample uh above a certain rate to be able to actually reconstruct the wave right and you could argue that that's some form of learning or intelligence or something else and so you need to be smart enough to actually tell how smart you can be in some sense yeah yeah I I love that analogy I I my undergrad degree I studied Nyquist theorem quite heavily which is funny because we broke it so much right I mean let let let's okay sorry aside on on nqu but um you know what it says roughly is like look I mean if you want to let's say output certain resolution or frequency in the fer domain you need to sample at half the frequency you want to reconstruct right otherwise the information is simply not there and you can see if you take a CID and new sample too little you will not rec recover the original frequency but then look at this super resolution generative models right you input like a 32x 32 pixel image it fills in the details in a way that completely violates that principle so I I talked to my signal processing teachers um about this and and of course it is violating and it is inventing it's hallucinating but of course it's cheating because it turns out that the world has certain structure that you can learn which is essentially what all these generative models of images do going back to your point I mean I agree there's there's that's a bit of like this sort of argument that there's these emergence properties right at some point the model might have the capability let's say to self-correct let's call it self-correction and as soon as it hits that capability you can see how wow boot strapping now is Trivial and of course it's not going to be that blanket across all the but certainly we're we're going to see some of this it's not going to be that dramatic it's going to be more like hey this one now emerg and so on but um of course once that capability is there you need to have the algorithms to exploit it especially these reward models that will be mostly driven by the model itself right and then how many labels how are you going to correct the the mistakes that it might still do those are very interesting questions and even from a product standpoint they're quite cool right like I mean we all play with these models I mean if they if if they fail I mean you just could say look I didn't like this and the models should adapt and that's where long context also plays a bit into into into the equation right you have long interactions you you finesse what the model is doing for you that sounds good a prior so um how do you en enable those capabilities and and so on is one one of the many kind of exciting future directions in the field we have this uh progression of the field from General algorithms to uh increasingly General models we'll see how far we get that uh in in that vein uh Deep Mind has also done really amazing work in particular domains right um protein folding Material Science whatever um where are there are we fully in the era of General models or are you know is is starting with language and going to audio video as such does that solve the rest for us or what what other data or domains do you think are unique that are not well represented in that Corpus the way you explain is perfect but I guess there's there's a sort of a time and and also what level of performance question here right so could current models do ass you the reasonable job at folding proteins perhaps right they might even use tools on the internet download a bit of a piece of software and figure it out um so one way I I kind of characterize General models is they are I mean the level of performance is irrelevant but they're like 20% good at everything okay so that they that's the level of performance they reach but it's general which I mean it's powerful so you want that 20% to keep going up relatively uniformly across the board so you don't want to over specialize the models or the research but the world has very important challenges that are worth solving with specialization right so I think what I tell people in the team and and around me is like look if the problem is worth it then let's obviously specialize probably we can and we might see more and more bootstrapping from from Gemini and these models to to then a specific solution that that the model is going to be throw away except for you know maybe cracking protein folding or you know figuring out nuclear fusion or modeling the weather right some of the projects that are currently active in in the Deep Mind portfolio well you better choose the ones that will matter uh because the timeline is such that well we'll get there sooner if we specialize um on those domains and then perhaps eventually like the general model will overtake but I mean that's just probably far away or further away so what it's worth it do it and that's I think we're going to still see this hybrid mode although again more and more taking make maybe Gemini and then doing something amazing in a more narrow domain and I I think that to me seems like a a good situation to be in because that directionality of taking a generalist model and doing something uh by fine-tuning it there's going to be a loop back as well right then you're going to do something amazing let's say in math right we recently did that and then well the data or as a reward model or something that model will do will will help the main model get better as well so it's quite a synergistic thing to do as well but but if it's important for the world it's okay I'm okay with still Tas specialization that's for sure there is a vein of criticism of math and computer science as with games or any other constrained domain like this that uh you know it's it's and and tell me if you think this is like a really Niche view but that it's say dead end versus like General reasoning advancement um even though has all these attractive attributes like the you know ability to generate and uh self validate in many ways like how do you react to that criticism there's a validity that that again going back to the reward function sort of question right you want to create the most General model or or agent or intelligence um turns out that reward is never perfectly defined by the environment um even if you go oh no surviving is look it is it is extremely complicated to compute a reward so you could argue like look when you do get these rewards um from certainly somewhat artificial you know interesting but artificial domains that might not generalize to the real problem and in that sense it could be at that end then how you do the research is important right let's use math as an as as an example right we do have access to the reward but even there it's not that simple right you know if if I mean if I'm doing a simple calculation okay yeah 4 + 4 equal 8 that I can check simple but now you start thinking well prove this theorem um the proof is either correct or not but that starts to be more complex to get a creas reward signal sometimes unless you can formalize it and then there might be backs in the formalization process or in the um maybe in the engine that checks the math underneath means and even then I mean if I say 4 + 4 and instead of saying eight I say 4 + 4 equals 8 is that correct or not right how do you check correctness I mean it's you start to interl language explanations so I think by trying to solve these problems not in a strictly hey the reward is sharp and crisp did you win or did you not but you start also say saying hey the reward model just needs to understand what's correct or roughly correct and look at what you wrote and assess then you might start generalizing and going away from what would be this Niche Niche World which is reward is perfectly defined to hey like what's truth what's not true I mean that's quite complicated right so I think in that sense um depending how you attack the problem it's definitely not a dead end and I think even because it's so hard to have a perfect reward when you interl language I think by even by accident the field will move forward as we get to discover these more General reward functions strain them better uh and and they will themselves then push the models and hopefully some sort of self-improvement loop will appear uh more than it has so far I I think you already said uh you feel AGI within grasp in the next in the next few years like what is your most um contrarian take on on this or contrarian under discussed take on this or AI in general I sort of didn't like the term AGI too much uh which is funny because Shane is a co-founder had a lot of I mean incredible foresight right Shane I I I I just had a discussion with him about AGI timelines recently and I mean in 2009 he predicted it's 2028 and that seems again depending on how you take the definition and how stct and what's the test that was quite you know a a long time ago that that he claimed that and I think that still metac calculus Etc probably suggest the world agrees with that prediction a bit maybe a bit still pessimistic rather optimistic his prediction versus the world estimate is that 2030 or something the contrarian view would be I'm not sure it matters that that we achieved AGI I think it might not look like hey like it's going to be exactly like the cognitive task we can do and then we reach parity um it's going to be a distribution of things that these models can do or can't do and that to me feels like what still worth pursuing right and I mean you see plenty of examples of as I as I was saying earlier the model will crack this impossible puzzle in math and then it will just contradict itself trivially somehow right so I think we need to be ready to not be too fixated on aggi um still fix the most obviously agous errors that's very important because I think they they they are something profound that is wrong in the models but I think it's not the exact goal probably it's it's it's just I I think more of a distributionally rather than the one point that before it wasn't and now we have it and honestly it's going to be also impossible to get agreement so it's going to be quite a cloud of moments that people might feel it has happened now but it doesn't doesn't matter because the models might be used for amazing things and products and and research itself right bootstrapping research and science is one of the things of course we're excited about um Google deep mes Mission has science very much present in in the in the mission statement that's what sort of motivates at least myself so I don't really care maybe to build a in in the strict definition uh but I understand it it's a good goal and I think um I appreciate uh to have a single number to aim for but I think it's going to be hard to agree with with everyone as usual maybe one last one for you would just be uh do you it sounds I'm guessing no because it sounds like the mission continues quite a bit beyond that but do you live your life any differently believing in 2028 I think I was I was reflecting on on like sell like I guess smartphones right so so yes with kids like how do you how do they do you present with the option of smartphones and I mean I we don't have that many samples or data points but then I mean obviously like that's obsolete right there's there's these Technologies um I have my kids are young young so I don't get to the the kind of oh you can try like this Jam I CH GPT whatnot but I think that is woring in the more human of being a that sense I think I think you you adapt as well sort of of how you do with like scaling up right you know scaling up is very important as you progress in your career you go from Individual contributor writing codes to helping others sort of figure out what their their you know their path is so I think the scaling up thanks to this technology is pretty there's a huge opportunity so I've been of course trying to use this to you know figure out what has happened in the endless chat rooms that I'm in I'm in London so so when I wake up I mean California has like given me like a lot of tokens and long context to process so so I think there's of course personally you also try to figure out how to best use the technology to scale yourself and I talked to quite a few people that are not that much into technology and all I say is look try to figure out how you could collaborate or use it as a tool um because I think that that change is definitely coming no matter if AI or not and I might s apply it but then yeah for for from a intelligence creation standpoint which is having kids that is much more complicated sadly it's a bit zero shot learning so we'll see um I'll tell you next podcast probably how that's going can I ask you one more question if you were to give advice to somebody with kids today what do you think their children should study they go to college in 10 years what what what should they mature in or what should they be doing to prepare for the future world honestly from a studying perspective I feel like there's always your passion element that cannot be sort of I mean many of in the early days right you you did deep learning not because it was like the thing to do I mean it was just what many of us like to do and I feel personally I cannot give advice that says you know find the top professions and choose based on that uh so I'll modulate my answer saying definitely find which aspect right is your true passion or you know you admire someone Role Models Etc now what I will say and I've been saying this for actually quite a few years is that project how that thing changes with AI exploit that of course not everyone will understand basics of ma of Technology but to be honest as I was saying language is the driver so it's not that hard to kind of understand how this works I was talking to my sister who is a teacher and I mean I just look she just used it to create like a summary of a kids's homework uploading all the homeworks and okay you know that that that way of thinking that needs to basically go into quite a few profession so so I guess first order beat find your passion follow that and then second really Embrace some of the tools that that are present today of course if you're into technology computer science then another very fruitful direction is to find the corners of the space that AI hasn't gotten into maybe you you you will train a specialized models because again it might be worth doing there's still quite plenty of opportunities I was just chatting to someone about um climate modeling weather modeling is kind of cracked with deep learning but climate is quite different we only have one sample which is one planet so it's pretty tricky uh but then yeah think about that as a kind of what areas could be enhanced and and of course otherwise if you like the technology um I think there's quite a lot of llm research and related to to be done for I mean 5 10 years at least so um that's probably still a worthwhile way to to investigate and if you go into research there's definitely lots to do well thanks for doing this Oriel it was a great conversation yeah thanks for joining yeah likewise great questions and hopefully next time I I'm in the bay we can we can do one where we're in 3D find us on Twitter at no prior pod subscribe to our YouTube channel if you want to see our faces follow the show on Apple podcast Spotify or wherever you listen that way you get a new episode every week and sign up for emails or find transcripts for every episode at no- pri.com
Oriol Vinyals · Google DeepMind 研究副总裁、Gemini 技术联合负责人
Oriol Vinyals 谈 Google DeepMind 的 AI、搜索与 Gemini 愿景
→ 在 AI 访谈库中阅读(可切换中英、记录进度)从领导 AlphaStar 打星际到组建 Gemini 项目的完整职业历程,兼谈长上下文 LLM 的进展、模型推理能力与 AGI 时间线。价值在于亲历者视角:他讲述了 Google 如何把全公司 AI 力量整合进单一 Gemini 项目并铺进各产品线。