🎙️AI 访谈库
World Labs 的 Fei-Fei Li 谈打造大型世界模型
Fei-Fei Li · World Labs / Stanford

World Labs 的 Fei-Fei Li 谈打造大型世界模型

World Labs' Fei-Fei Li on Creating Large World Models

2026-06-04 · Bloomberg Tech 2026 (Emily Chang) · 约 14 分钟读完 · 原文
李飞飞与 Emily Chang 对谈:称 LLM 是「黑暗中的文字匠」,阐述世界模型三分法(渲染器/模拟器/规划器)、Marble 商业化与创办 World Labs 的原因。

Everyone is focused on LLM's, chat GPT, Claude, large language models, but you have raised a billion dollars to build something different. Large world models. Make the case for us. What is the bet you're making that others aren't? >> Right. So, um this is my uh co-founded startup, World Labs, and uh we are uh all in in spatial intelligence. And uh the means to spatial intelligence is building a large world model.

So, what is the case for us? The case for us is a 500 million year story. Is that animal intelligence starts with seeing and moving in a physical world. That uh evolution began with us as animals knowing what the world is, knows knowing who we are, knowing how to move around it, interact with it, and uh much of life, human life, human work life, human private life, has a lot to do with uh perceiving, understanding, reasoning, interaction with the world, including imaginary world of creativity, of uh of uh productivity, uh as virtual worlds.

So, unlocking that capability in machines, unlocking the capability of generating from any 3D, 4D worlds, unlocking the capability of reasoning within any world, unlocking the capability of uh teaching agents or robots or or assisting humans to interact with the world is what spatial intelligence is about, and that's what we are focusing on. >> So, what can world models do ultimately that LLM's will never be able to.

Can words put out fires? Can words cook an omelet? I think there's so much, right? So, we for example, creativity. People design. People whether we're designing interior space, we're designing machines, we're designing we're designing homes, we're designing stories. So much of that is beyond words. We also use agents. Whether we use agents in virtual world, whether it's for entertainment like gaming, or for more serious industrial industrial applications, whether it's digital twin design or or inspection or or or what kind of many kind of optimization tasks.

Or we build robots and to help us to do a lot of things from putting out fire to helping health care scenarios to manufacturing. All these are application downstream applications of unlocking spatial intelligence and building world models. >> So, what's the what do you think the chat GPT moment for world models will be? Like how will we know this has arrived? >> Yeah, that's a great question, Emily, because chat is such a consumer behavior that chat GPT moment tends to be used to describe a viral public consumer moment of getting so close to what AI can do.

In a in a world of world models, um the kind of spatial intelligence we're trying to unlock, um I'm still trying to figure out if there is a corresponding consumer moment because the kind of applications we are talking about um tends to be first going to the professionals, professional creators, professional designers, professional developers, uh professional researchers and engineers who will use it for robotics and industrial design and all that.

So, maybe we will not necessarily have a consumer moment. But maybe we will. And you know, I I would love to design my home in a much easier way and just change the color of the curtain, you know, with a click. >> All right, that sounds pretty cool. So, in the last 6 months, Yann LeCun left Meta to work on world models, Google shipped Project Genie, Nvidia has its own world models, Cosmos. Nvidia's also one of your investors.

What do you have that they don't? And which competitors out there worry you the most? >> Yeah, so first of all, we started World Labs in 2024. I still remember when when we were out talking about world models and spatial intelligence, it was just a year after ChatGPT. People were still totally talking about LLMs. So, we we really had a head start and understanding that this is going to be the next frontier of AI.

I'm very excited by that. So, what do they have we don't? Well, first of all, I think we have an incredible team. We have the conviction. >> They don't have the godmother, that's for sure. >> Um but but the the world is big and and I think this is just like LLMs. I think there will be many companies doing incredible work in world models. Just as 24 hours ago, uh I we kind of got fed up that the word world model has been so uh confusing and being used so in so many different ways that we actually put out a blog just explaining what a functional taxonomy of world model is instead of mushing everything together.

And the way I see it is right now there are three ways of calling world models when it comes to spatial intelligence. One is what I call a renderer when the model puts beautiful pixels on the screen. Mostly like video generation model. And the consumer is mostly human eyeballs. And while the model commits to beautiful pixels on the screen, it doesn't necessarily commit to physics and dynamics and geo geometric correctness because that's for just consuming human eyeball consuming not necessarily for computation and and other other tasks.

Then another kind of world model is what we call a a planner. That is more for machines, more for robots where it outputs whatever the input is the state of the world or the action. It outputs a correct action to take to the next step. And you see that kind of world model a lot for robotics applications and you hear that in that context. The third kind which I think is the lynchpin of the the three is a simulator. Is that it actually is consumed by humans as well as machines is trying to respect the structure, the physics, and the dynamics of the world and really simulate the 3D and 4D information of the world as well as well as the semantic information.

And a simulator could become a renderer. The simulator could become a planner. But this layer is a huge critical path in my opinion, to unlock spatial intelligence. And that's what World Lab is working on. >> All of this rolls up into robotics, so I want to get your take on the field and humanoids in particular. Funding for humanoids hit $6 billion, but, you know, they still can't load my dishwasher as fast as I can.

They still can't go get my Amazon packages. Will world models, World Labs, close the gap between hype and reality? >> That's a loaded question, Emily. First of all, >> That is my job. >> I get it. First of all, robotics is going to be one of the most important revolution in human industrialization. $6 billion is too small. Right? If you look at self-driving cars investment, if you look at language models investment, it took way more than $6 billion.

I'm not saying we now I think it will take time to invest, and it will also hopefully not take the hype, but take the thoughtfulness to invest in the right effort. And, for example, unlocking world modeling and spatial intelligence and simulation layer, all this is part of that that important effort. Um are we going to close the gap? I do believe World Labs is working on one of the most critical technology in spatial physical intelligence.

And obviously, that's the that's the hope. >> You've been more measured on AI safety, skeptical of the doom narrative, but also of heavy-handed regulation. When you look across the industry, where do you feel real safety work versus safety theater? Is anyone getting it right? >> So, in general, I've been just more measured on every every rhetoric makes me very boring, to be honest. Um I think there's just so much hype.

There is so much hype. Um obviously, we need to build the right technology. We need to guardrail the technology. Whether you use the word responsible, you use the word safety, you use the word um um trustworthy, um building the right technology and product so that it can empower, enhance, augment humanity, and not harm them is the goal of any any work we do, whether it's AI or not. So, where is it doing right? I really hope every company, every um every product that's being built, that the people behind it are being mindful of that, and are thinking about, you know, what data are we using?

What system are we building? What evaluations are we conducting? What guardrails are we putting in? How do we communicate with our with our users and customers? How do we work with regulators so that when the rubber hits the road, that we are um you know, being responsible. I do believe a lot of this work is happening. It's not happening in a theater, to be honest. For example, so many pharmaceutical and health care um industry uh companies are incorporating AI.

Literally, I just came from the hospital to come to your to to your panel because I have a family member uh about to get a surgery in in the next 1 hour or so. And I was just in Stanford Hospital looking at where AI is already being used and where AI could be used. And it's already happening. Doctors are using AI to to to help them with charting. Radiologists are using AI to assist them reading the the MRI and the CT scans.

I do hope that we have more AI to help our nurses, to help family members. I got this long radiology report last night and the first thing I did is send it to a AI so that it can help me to explain it. So, all this is happening. Um safety measures are happening. Um but there needs to be more in a right way, in a in a scientifically grounded way. Um and that's the conversation that should be taking place instead of what you say the theater.

>> Well, thank you for coming and I hope your person is okay. We all we all do. Um the backlash is we all it's being called the AI hate wave. I'm sure you've seen the video of former Google CEO Eric Schmidt getting booed at a college graduation. You spend a lot of time with students. What are they saying? And if they're scared, are the fears justified? >> Yeah, I do spend a lot of time with students. Uh to be fair, my students are pretty privileged cuz they're Stanford students.

I think it's I think it's even more important and I try to do it myself that we spend time with our teachers, with our nurses, with our our parents, grandparents. And that's actually something I try to do. I try to talk to K-12 educators. I try to go to places and talk to people where they feel that they're not part of the conversation. And even Stanford students reflect some of this mixed sentiment. There is anxiety.

There's sense of hope. There is also excitement. There is also confusion. There's also um simultaneously a sense of dignity and agency when AI can help me do things that I couldn't do before and a sense of loss of dignity and agency if AI is is going to take my job. So, I think I think the sentiment is mixed and I really want to point out a lot of this sentiment happens when there's a vacuum of thoughtful public discourse.

Right now, the oxygen, the air is all sucked into the polarized extreme of doomerism or total utopian. And when hype takes all the oxygen in the room, that void brews the kind of anxiety and it's actually that void we really need to care about because that's where real people live. That's where real people are seeking answers. And I think it's uh As a scientist and a educator and a entrepreneur, I'm on ground zero with students, with educators, with entrepreneurs and I really do believe it's is one of my responsibility to not hype and try to speak with with both science and humility and and in inspire people to to recognize this is a technology that can truly empower a lot of our work and life, can truly help us, you know, have a better health care system, have better scientific discovery, have better uh uh better environment, better education if we do the right thing.

>> Mhm. We're both moms. We both have young teenagers. How do you think AI will change learning in the college experience? >> AI must change learning. AI must change K to 16 learning. I think this is one of the biggest opportunity for humanity in the next decade to come is that what the most precious resource of our entire world is human capital. And when we have gotten a technology that can answer standardized tests, whether it's it's a common core kind of test all the way to international Olympiad math exams, when AI can do better than average human, it's not about humans are bad.

It's about we need to change the education system. We need to change how we evaluate. We need to change the way we empower teachers to teach to to educate the next generation of students where they can use these tools, be empowered, and do things that we can never imagine. >> So, do you think our kids will still learn? >> Absolutely. If we teach them right, if the society prepares them right, they should not be all of the kids today should not be scared of AI.

They should feel the human agency to to lead AI, to use AI in the right way, and to use AI to make the right to make the impact that they want to make for the world. >> Anthropic CEO Dario Amodei has suggested AGI is 2 to 3 years out. We'll get there by scaling the current paradigm. Demis Hassabis says we're at the foothills of the singularity. You've said you don't even engage with the term AGI. Are they wrong, or is the disagreement about what we're calling the goal?

>> I don't engage with the term AGI because the founding fathers of artificial intelligence as a scientific field had this dream of thinking and doing machines and that is a scientific quest and that quest has been my lifelong career and I'm still on that quest. Now I'm combining that scientific quest with making products that can make people's life better and that is the field called artificial intelligence and um I'm okay people call it whatever they want they can call it an Apple that's fine.

>> >> Um I'm focusing on building a technology can that can truly that can truly make a difference in people's lives and at work. >> What's the one thing you'll have shipped this year that we'll be talking about next year? >> I hope that we will be shipping a model for spatial intelligence that will inspire incredibly exciting product opportunities that people haven't seen before.