Hi listeners and welcome back to No Priors. Today's guest is Dr. Fay Lee, a pioneer in computer vision and deep learning. She created ImageNet, the groundbreaking data set that helped spark the deep learning revolution. Fay is a Stanford professor and the co-director of the Stanford Institute for Human-entered AI. She's also led AI at Google Cloud, advised international policy makers, and recently co-founded World Labs, a company dedicated to developing spatially intelligent AI. Fay Fay, thank you for joining us today. Well, thanks for inviting me. This is going to be fun. So, you have made extraordinary contributions to um science and policy over the past two two decades. I'll start with the biggest question like why start a company now? Because in my heart I want to build. I see this as such a critical and fun and exciting moment to build some extraordinary technology that everybody can use and I believe so much in spatial intelligence and the kind of 3D world models that can empower so many people as well as so many use cases and I think that's just it's going to be really um exciting and I can do that with an extraordinary very extraordinarily brilliant group of young technologists. I want to come back to you know the people you're working with because I I uh know some of your co-founders and was uh you know trying to convince them desperately to start a company a while back and then they were like oh no we have a bigger mission now with Fay. What is spatial intelligence? Can you define it for a broader audience? Spatial intelligence to me is the ability to um to understand reason and interact and generate 3D worlds because our world fundamentally no matter how you say we can project it fundamentally is 3D and it's 3D because physically it's 3D and digitally if there is a true 3D representation then we can make a lot of things happen more easily whether it's designing more creation or navigation or or simulation or or the experiencing of um uh AR VR all this to me is part of spatial intelligence and again I think it's what really excites me is humans have spatial intelligence we are it's part of our uh core intelligent capabilities animals have a spatial intelligence the the entire entire journey of evolution also is um deeply intertwined with the evolution of spatial intelligence. So it's so fundamental without spatial intelligence AI would be incomplete. How does that translate into what you're doing with your company or is there anything you can share in terms of what that means relative to what you're building? Yeah. So work fracking one of the hardest problem in AI which is actually um making world models that are fundamentally 3D because once you can um crack that problem you can unlock a lot of spatial intelligence problem. So we are the first company we know of that is solving this uh the 3D generation uh foundation model problem. I have many questions but since you um are you know describing this uh first as you know the you know 3D's criticality to just sort of understanding the world um does that imply you you feel that the world models that you know world labs will create or or others in academia or in uh companies will create will someday be like you know realistically accurate like represent physics and understanding of the world that we can many more things with. Yeah, it should it it should be realistically accurate or a plausible. So you can create a fantastical world, but it should be plausible because the geometry and the physics of it need to be plausible and uh um and that is fundamental to spatial intelligence. Does that imply you have a particular um point of view um from a like a neuroscience perspective of like you know how fundamental visual you I mean you've always been a leader in um uh computer vision right but in how important visual intelligence is versus let's say like large language models and textual intelligence. I actually do. I think from a neural and cognitive science point of view that spatial intelligence is a really hard problem that evolution has to solve for animals. And what's really interesting is I think animals have solved it to an extent but not fully solved it. It's one of the hardest problem because um what is the problem animal has to solve? Animals have to evolve the capability of collecting lights in something which we call eyes mostly. And then with that collection of eyes, it has to reconstruct a 3D world in their mind somehow so that they can navigate and they can do things and of course they can interact. For humans, we're the most capable animal in terms of manipulation. We can do a lot of things and all this is spatial intelligence. To me, that's um that's just rooted in in our intelligence. What is interesting is it's not a fully solved problem even in animals. We uh for example uh for humans, right? Um, if I ask you to close your eyes right now and draw out or or or build a 3D model of the environment around you, it's not that easy. We don't have that much capability to generate extremely complicated 3D model till we get trained. You know, there are some of us, whether they're architects or or designers or just people with a lot of training and a lot of talent, and that's that's a that's a hard thing to do. And imagine you do it at your fingertip much more easily and allow much more uh fluid uh interactivity and editability. That would just be a whole different uh world for people, no pun intended. Are there other big areas like um spatial intelligence that you feel haven't been as developed as they could be from a model perspective or other sort of missing gaps that you think in general as we think as we build this sort of AI future we should um focus on over time or people should build out. I was just wondering in addition to sort of 3D and world generation and are there other big problems like that because it feels like there are a few big things that we've solved for over time and other things we're working on. We're sort of solving language. I would say language is solved to a huge extent and uh 3D to me is as you know critical and and difficult as language. So what else that's not solved? I mean the entire space of emotional intelligence is something that um I don't even know how to begin to solve. I know a lot of people who haven't solved it. So that's when AGI is achieved. I can tell you the training data for that is not going to come from Silicon Valley people. Don't underestimate Silicon Valley. Yeah. So, I'll put myself in this bucket, but I I I think we probably need a broader set of people. Yeah. No, that I agree. But these are the three three big buck buckets to be honest. That's I don't know. What do you think Elon and Sarah? I think it depends a lot on um what you en encapsulate in each model. So I agree with your framework in terms of those three and then certain things like um you know the spatial intelligence I'm assuming also delves into different types of physics simulation and simulations of the world and that you know like those are big areas that I think a lot of people aren't working on that I think are really interesting or important. So um and there's there's sort of the macro and the micro scale of that. the microscale eventually becomes material sciences and other very different types of things from what you're talking about where it's more molecular modeling or yeah right and also somewhat goes out the current definition of AI which I do think they'll be empowered by of course there's robotics but robotics is very much a system integration problem as much as a um you know even if you look at animals it's not just uh the compute in the brain per se right yeah a lot of these things seem much more distributed in terms of spatial intelligence relative to specific systems that animals have and in some cases it's to your point not not as centralized as one would think. So it's it's very interesting to start thinking in terms of those models of more distributed intelligence across an organism uh versus a CNS but um yeah I think I think it's very interesting stuff. You've also done work in in this field of of robotics and like physical intelligence. I think of the data hierarchy for you know robotics foundation models and actuation as you know people want to of course use video right because that is what is available to us there's a big question on like simulation and how much you can get from that today perhaps people are do not see the future of like the quality and the physics that are going to be available to us um and then there's you know close to embodied uh like different forms of tea op and then like embodied data collection is that the hierarchy you have in your mind or do you think people underestimate simulation and world models for the future? Yeah, great question. First of all, I like you say, I do work in robotics, especially in my lab at Stanford. I have no doubt that humanity will move into an age where we cohabit with robots and also the world the world robot is not humanoid per se. Robots taking all kind of forms and shapes. Actually, a few years ago, my lab wrote a really fun paper about morphological intelligence is where the the morphology of a a an agent actually can change by optimizing the tasks they're trying to achieve. So, so we should be a little more imaginative than just human humanoids. Having said that, how to train robot uh you mentioned this whole data some people call it data pyramids or data cakes or whatever. I agree. I think it's going to be a hybrid of uh many different forms of data. I also think uh simulation is underrated. It's um actually it's not underrated by a lot of experts and people in the field. If you look at a lot of robotics companies, they are working on simulated uh simulation and synthetic data. I also think we have to be um also aware that unlike language models or even unlike um spatial intelligence foundation models, robotics is a highly multimodal um um uh system that I think what is truly underappreciated in my opinion is haptics is there's so much especially if we want to do manipulation not just navigation. I think haptics data and the ability to really integrate haptics into vision and perception and spatial uh data is is absolutely critical. One thing that you said that I thought was really interesting is um how many different what what are the different morphological forms that a robot may adapt or adopt. And um there's sort of two counter arguments people make in terms of the potential future. One argument is that um from a supply chain perspective and managing builds and scale of manufacturing, you're going to have many fewer form factors. And the other argument is the economic value of specialization is very high and therefore there'll be, you know, thousands and thousands of different form factors as we move to sort of a robot-driven future. Do you have a point of view on sort of where we're likely to land between those two viewpoints? I think we're going to gradient descend into optimization of productivity and efficiency. My hypothesis is that the requirements of different tasks are so vast that having very few form or or sticking with one form is energy energy inefficient and a lot of tasks can be done and should be done by much more energy efficient form factors. Just an extreme and and trivial example. If we put robots underwater, they should not be in the shape of humans. They better be in the shape of fish, right? Just think about energy efficiency. And the same with flying. I don't think human form is our airplanes are becoming more and more robots. And so I I I do think there's going to be diversity. Robotics is one potential application for the future. You're scientist first um but also you know did the Twitter board involved in startups. Um what are the near-term commercial applications that you can imagine for generating 3D worlds? I believe creativity is a vastly um exciting area where humans can be superpowered by uh by uh AI and by spatial intelligence. And here I draw an analogy with software engineering. If you look at today's success of LLMs in software engineering including applications like cursor and windsurf and all that what you see is is a lot of collaboration between AI and and humans and then the collaboration comes in different levels of skill sets and all that and I think creativity will be similar is that whether we're talking about designers 3D artists VFX ex artists or even marketing talents and and and game developers. There's so much need in co uh in designing and creating 3D space and this is fundamentally such a hard problem even for the trained skilled uh people that having a collaborator will be uh extremely um fun if if we do it right. And so I see creativity as an area that is really exciting. I also do think um that um a lot of what we're waiting for for metaverse or XR uh ARVR is content creation. I understand hardware itself needs to continue to evolve, but I also think software uh we're we're looking for content creation and that lends itself so naturally to um uh 3D modeling and 3D uh uh or generative spatial models and that's another interesting area to to look into. Do you have um a strong point of view on whether or not world models are like an interesting answer to scalable RL for like more generalizable agents? I actually do think this this is like I said AI is not complete uh without spatial intelligence because uh um humans interact in um in uh 3D worlds and in the digital world we need all kinds of interaction you know take design as an example it's a deeply you know it has um when we are thinking about design there's so much we are optimizing for in our mind's eye whether it's beauty or efficiency or optimization or or whatever it is and that lends itself pretty naturally to RL um settings. What are the biggest challenges in um I guess trying to go down this path of you know designing and training world models. I imagine one is like you worked on images, you worked on video but we we have images and we have video and we don't have lots of you know 3D worlds like in in a format I assume you're building. Yeah, data is absolutely a challenge. You're totally right about that. Um, you know, to create uh world models, 3D foundation models. Uh we we require more and more sophisticated data engineering, data acquisition, data processing and data uh synthesis. So um uh I am envious of my uh NLP LLM colleagues that the the data is so abundant on the internet and we don't necessarily have that luxury. So that's definitely one um one challenge. Another one challenge is that um 3D it's this is kind of um ironic right every one of us use 3D every day like in so many settings basically you open your eye and and and the whole life that you experience is 3D okay even when we type on the computer or stare at a screen all the time yet it's still not as easy a form factor to deliver in the hands of people compared to language. The language is just so easy and uh it's also a very active form of it's not a passive consumption of viewing. Nobody wakes up and say I'm just going to sit here and watch 3D, you know. So, um that uh creates challenges for for for productization and how to do it in the right way. Were you ever a like a Second Life player or any anything? I'm not a gamer, but my kids love Minecraft. I was going to ask you if there was like a world that you want to experience or imagine. That's a great question, Sarah. You know, I would love to see worlds. I love seeing worlds I don't see. for example, like zooming in and in and into like microscopic worlds or, you know, um going into the inside of a engine, you know, knowing how the the the actual engine is is I know, of course, I know theoretically how it works, but seeing it with my own eyes, experiencing it, or even um you might laugh at this, I want to be inside a dishwasher and just experience what what that is. All this can be done in a in a virtual way if we manage to create you know role models of anything. Okay. I I I think a lot both a lot and I both want to talk a little bit about your past career and maybe some uh insights for anyone doing research or trying to you know have an impact within AI. Right before this I asked Andre Karpathy what I should ask you and he said you know um FE is really magic about ambition and thinking about data. you should ask her about her PhD like a and the creation of that 101 data set with phro um because it's instructive. So I have to ask you about that. You know first of all I have to say it's always really the greatest thing when you're a student is is u more well known and and achieving so much more than than you can. It makes me so proud. So very proud of Andre. Um, I was I'm surprised he remembers my PhD work. So, yes, it's true. It's u well gosh, it goes back to 2003ish and the world was just barely scratching the surface of uh internet and data was not much of a thing. But doing computer vision, we were my PhD work was really trying to get object recognition to work. That's the that's the problem of calling out cats and dogs and microwaves and chairs and all that when you're presented with a picture and u and we were beginning to hypothesize that data matters. But we had no idea there's no scaling law. We had no idea um you know how far data can go. All we wanted is if we have a machine learning algorithm whether it's a neur neuronet network or baset at that time was very popular or support vector machine we need some data to train and there was no data to train and as a PhD student you want to you know graduate and uh and Petra was like well fa curate a data set and I and you know I was thinking yeah I do need to curate a data set because every data set out there is so tiny I'm just not convinced. And Petro and I were just talking, you know, is it 15 different things or 30 different things? And then, God forbid the PhD advisor said the three-digit number 100. And I was like, you know, that's a lot of work. But I deep in my heart, I know he's right from a mathematical point of view is pushing the the model to generalize. We need enough data at least. So, you know, um I did write about this process in in my book uh the worlds I see that I stumbled upon a dictionary somehow and it really was for my own English study that the dictionary I think it's the Webster dictionary if I'm not wrong. It just kind of randomly has depiction of a visual depiction of some words. I don't even know what rule they follow to be on uh to be honest. Some are flowers, some of bicycles, some are dogs. And I was like, okay, this is actually, you can call it a cheater or a tool. I grabbed 101 of those words. Um, and that really made my PhD advisor kind of chuckle because he's like, "Ah, yeah, you just want to do one more than I asked for to, you know, dare me." So that's what I did. And I gotta say that I still remember I downloaded or you know tried you know from Google and Google was so new at that point and the Google image search were so terrible at that point you know compared to today and I had to do so much cleaning at some point I got so desperate I just asked my mom to do clean the the the image cleaning because I I wrote a little interface on the computer she doesn't know computer but at least she knows click click. So she helped me to to do some of that. I mean you've had one of the most storied careers in AI and to your point many of your students have similarly gone on to do really great things across the field, across industry, across you know the world. Um what are two or three moments that you think of uh when you think back on your career today? And obviously there's still a lot of career to come but I'm just sort of curious. I mean obviously there's a lot of things that you did in terms of sort of uh image and visual recognition related systems and all sort but I'm just sort of curious like when you think think of the last 20 years what stands out the most just given everything that you've done oh thank you for asking that question of course image that is one of those uh image that is consists of multiple moment from the early struggles and being told I will not get tenure to um to actually realizing Amazon mechanical turk comes to rescue you to the moment of Alex Net winning and also to a couple of years ago I was at an event in Toronto with Jeff Hinton and he said publicly like how that was so defining and he he was almost a little bit um apologetic that image that was not as recognized as neuronet network. So that journey is very validating and for scientists the validation is not about recognition or awards. It's that you made a difference like that conjecture that no one believed in that hypothesis that no one believed in we were make able to make it happen. So that's one thread just to make sure for any like you know people from the business world that are not familiar with it. that imageet was a large scale is a large scale data set with millions of labeled images across thousands of categories not just 10 and one right 15 million label images 15 million labeled images thank you fay that you know led to um uh amazing breakthroughs in deep learning in particular AlexNet and lots of progress in the field of um yeah driven a lot of machine vision forward and I actually remember um in 2016 or 2017 I used to show a slide which was the history of AI or you know Back then it was CNN's and RNN's and just GANs were you know kind of going and I had image net and AlexNet is like one of the seinal moments of you know this very small number of events that really defined AI progress and obviously now we have transformers as part of that and maybe diffusion models or something but it's it was uh such a big breakthrough. Yeah, thank you. Another moment I'm very proud of was actually Andre and um and also also Justin Johnson and and their dissertations. It's where in my opinion the first time that language and images converged by captioning and writing stories of of uh the visual world. It was significant for me for two reasons. One is that I literally thought I kid you not at the end of my PhD I thought when if I can live till 100 year old that was the problem we might be able to solve which is storytelling of pictures. So I entered my my my career like my first year um uh assistant professor thinking okay I'm going to do image that to solve object recognition and then I'm going to spend the rest of my entire career solving this problem of uh storytelling and then by the time Andre and then a little later Justin Johnson entered my lab that was around 20 um 13 2014 14 the beginning of deep learning and then suddenly the combination of um sequential model at that point is uh LSTM it's not transformer models but LSTM and CNN just had this blasted open the image captioning um um work and my work were the first together with Google's that was out of the door and that was really to me I I I almost had it was made me so proud. I almost had a crisis which is like what am I going to do for the rest of my 70 years or 65 years it's so that was really exciting how how fast the field has uh has u you know evolved. Can I ask you one more question about this just because you have um uh you know made this amazing progress like very efficiently right like you and I have uh offline talked before about how um you you feel it's really important for there to be you know moonshots and creativity in AI research beyond like very large funded corporate labs let's say and you know you you pointed to several moments that they come from like uh creativity and research in academia. What advice do you have for people about whether or not there's still opportunity for that or you know it's all just 10 billion dollar training runs from here? My singular advice and I still say that in my company in my lab um is be fearless. I think scientists and technologists and entrepreneurs have to be fearless. You know, eventually you have to figure out, do you need 10 billion dollar runs or then you come to Sarah to ask for funding? Probably a lot for both. Yeah. Yeah. Um or you have to figure out, you know, I don't know, data. Sometimes fearless is this very interesting position where you're somewhat delusional and crazy but somewhat just rationally bold and and it it kind of is in between because if you're too rational, it's not courageous enough. You're not identifying problems that are are big enough. But you're if you're completely crazy then I don't know there's some many things um that can go wrong. So, so be fearless, be courageous. I to me that is um you know even as old as I am that's how I feel I started my uh startup world apps is I want to be fearless and solve this problem of spatial intelligence as part of problem solving you've worked with some of the best uh AI researchers in the world over time and best engineers um how do you think about that in the context of your company like what sorts of people are you trying to hire are there open roles currently undoubtedly it's an amazing team I'm just curious like what sorts of folks you want to add and how you're thinking about that over time. Yes, we have open roles and we would love to hire the best engineers as well as product thinkers um at this point for for our company. So if you're a engineer or AI researcher or product talent out there passionate about joining the the most talented team and and solving this problem, please join us. So who do we hire? First of all, we really do hire in diversity of thinking. And this is where you know you call us a AI company but if you look under the hood we've got computer graphics experts we've got computer vision experts we've got data experts we've got you know uh generative AI experts we've got uh machine learning infra experts we've got optimization we've got so it's actually really important to uh to to hire um a diverse group of really talented people because a problem as hard as spatial intelligence is not a homogeneous problem like it takes talents of all kinds of background to solve it and and then I also just you know like I look for fearlessness how do you do that like how do you identify if somebody has fearlessness in their background or in their thinking processes it's in their background you talk to them you can sense someone is fearless you know you can sense what drives them You know, you can sense the questions they ask if they are if they start to asking you a lot of things about I don't know how to get this done. And I mean, of course, you have to ask those questions because you want to get it done. But if if you sense that it comes from the the the the the point of view of uh being scared of solving that, then that's not fearlessness. But um but those fearless people, they are creative, they're ambitious, they they they they can they're not afraid of um the uncertainty or the unknown. And I really love that. Well, I think a lot and I you know, we try to make make a business of uh doing business with fearless people and hopefully those that are technically creative. Um uh one one last broader question for you because I think an important part of your work um has been also um thinking how how to bring more people into AI you know co-directing the um Stanford Center for um human- centered and artificial intelligence. uh what is your most like if you picture you know not to use a pun on the book but if you picture the world like several years out from your your last set of predictions what's your most optimistic view of uh what human- centered AI looks like yeah thanks for asking in fact that is another point of my career I feel very proud of is the founding of uh human- centered AI institute hai and also the continue movement towards that uh way of thinking I think I want to build a world that AI collaborates and superpowers people. I still believe our world, our human world needs to be human- centered, you know, where love, relationship, um, just prosperity across, you know, all communities. These are really important. justice and all these are really important values and I don't think any piece of machinery whether it's AI or airplane or or biotech should take those away but with that those critical values in mind having AI to superpower us is is really really important because there's so many unsolved problems um one application area I um I had worked on is um is uh healthcare for example at Stanford, right? If you look at health care from drug discovery to cure diseases to diagnosis that can reach all people in the world to treatment that can be accessible to all people in the world to the whole health care delivery, how to make aging better, how to take care of chronic diseases, how to deal with mental health, all of this. We do not have a issue of excessive humans or anything. We're lacking help. You know, we are lacking scientific discovery. We're lacking diagnosis. We're lacking precision medicine. We're lacking safer and more effective ways of healthcare delivery and aging help and all that. And I that's what I believe. I think AI is a tool to help people. Yeah. I think a lot and I are um collectively invested in a series of companies that I I hope will be useful here from a bridge to open evidence to LA and but as you said there's a there's a huge spectrum of problems and honestly I've been less optimistic about the adoption of you know generally technology and healthcare for the last 15 years but it does feel like this time it's different and actually it's just massively net good here. Yeah, I actually started a digital health company before this and my hope is finally a lot of the things that people have been talking about for decades will come to fruition and it seems like AI is a great delivery mechanism for that. So totally totally well thank you so much Fay. It was fantastic. This has been inspiring and uh great to hear a little bit more about World Labs as well. Thank you. Thank you a lot. Thank you Sarah. Find us on Twitter at no prior pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way, you get a new episode every week. And sign up for emails or find transcripts for every episode at no-briers.com.
Fei-Fei Li · World Labs 联合创始人兼 CEO、斯坦福大学教授、斯坦福 HAI 联合主任
教 AI 理解物理世界——对话 World Labs 李飞飞(No Priors 第 117 期)
→ 在 AI 访谈库中阅读(可切换中英、记录进度)李飞飞讲述从 ImageNet 到 World Labs 的历程、3D 世界建模的技术难点,以及她挑选和培养出 Karpathy 等顶尖学生的方法。她强调今天的 AI 系统仍缺少对物理世界的感知与交互能力,智能的下一步必须建立在具身与环境交互之上。