🎙️AI 访谈库
炉边对谈 Yann LeCun:AMI Labs 执行主席
Yann LeCun · AMI Labs

炉边对谈 Yann LeCun:AMI Labs 执行主席

Fireside Chat with Yann LeCun, Executive Chairman of AMI Labs | RAISE Summit 2026

2026-07-16 · RAISE Summit 2026 (Bloomberg's Tom Mackenzie) · 28m · 约 25 分钟读完 · 原文
LeCun 离开 Meta 后以 AMI Labs 执行主席身份与彭博 Tom Mackenzie 对谈,阐述超越当今 LLM 的路径:世界模型与对 LLM 规模化路线的反向押注。

Take your pick of chairs, Yan. >> All right. >> Here we go. >> Yan. Godfather of AI. >> You know, Godfather is uh I live in New Jersey. Godfather in New Jersey means you are the you work for for the mafia. >> Yes. Okay. Well, maybe we use a different terminology than that. I should start by wishing you a happy birthday. It was your birthday yesterday. >> Yesterday. Yes. >> Okay.

So, there we go. Happy birthday, Yan. >> >> I but I turn I turned 26 for those of you who are older. >> Yes, 26-year-old Yan has achieved has achieved quite a lot. Uh formerly of course um chief AI scientist at Meta for 12 years and am I thinking correctly [snorts] and so consequential with the world we live in right now is the convolutional neural network CNN which you pioneered in late 1980s that underpins really how we use our phones.

image recognition every time you go for a scan at a hospital, autonomous driving that is still having a worldwide and global impact. And seven months ago, am I right in saying seven months ago, you launched Ammy Labs, founded, co-founded Ammy Labs. So give us an update. How is how is that going? The focus of AML is to build world models. So physical AI that can interact in the world, learn, react, and take in those physical physical cues.

How is all of that going? >> Right, it's AI for for the real world uh for real signals. So, LLMs handle discrete sequences of symbols quite well. >> Mhm. >> But when it comes to real signals, video sensors from industry uh or whatever. Um LM are completely useless essentially. Um and so it's a new type of AI that we we're developing. Uh our ambition is to really bring about the next revolution in AI.

Um the you know we have LLM and derived system systems from LLM that can uh you know pass the bar exam, demonstrate like prove mathematical theorems, write code, you know all of this is great, very useful, very productive, improves productivity of a lot of people, can write long emails which the recipient will feed to an LLM to produce a shorter version of um um etc. It's great for that but but not not for understanding the physical world and you know that's why we don't have level five self-driving cars despite commercial net and we don't have domestic robots.

We don't have robots that can do what a you know what a 10-year-old can do or even what a what a cat can do. So, so there's a huge gap uh in performance between the type of tasks that we think of as you know intellectual task and the tasks that we take for granted but turned out turn out to be really really difficult for for machines to accomplish. >> So just explain for us why scaling LLMs at the frontier the anthropics and the open AIS cannot do this.

>> Okay. Um so there's a lot of complexity about the physical world that lens really cannot capture because the representation of the world that they have is through text descriptions of it. Of course, you know, modern lens have vision pipelines, right? So they you can feed an image or a video to it and they produce basically what amounts to a tokenized version of that uh inputs which is you know merged with the the flow of text and then the the system to some extent can interpret image images but but there is a lot more in to the world than than just a textual description or even a description um you know by by token.

So, um, there's a lot of intuitions we we develop about the real world. If I if I drop this bottle here, you can you can predict what's going to happen, of course. And of course, you ask a LM, you know, I'm holding a bottle. I'm going to let it go. What's going to happen? It's going to tell you it's going to fall. And if it's above the floor, it's going to hit the floor. Maybe it's going to explode and splash water everywhere.

You can predict exactly in what uh you know what the details of this are but most of the details in the world in the physical world are not predictable and so the very idea of a generative model does not work in the real world if you it's relatively easy to predict the word that follows a sequence of words which is really what LLMs are doing um whether it's natural language or code or whatever right you show the system a long sequence of of of words and you train it to predict the next word.

It can never exactly predict the word that follows but it can produce a score for every possible words in your dictionary or a probability distribution. Right? So then you pick a word from that probability distribution. That's the word the system is going to produce. You shift that into the input and now you can produce the second word. Shift that into the input third word. That's called auto reggressive prediction.

It's a very old idea and that's what edens are based on. uh if you want LLMs to reason a little bit um you have it produce lots and lots of those sequences and then you have another neural net that selects the best answer among all of the ones that are produced which is ridiculously expensive. Uh but this is how those you know coding mathematics kind of system work more or less. Um, now if you want a system to understand the real world, why not use the same technique?

Take a piece of video and train the system to predict what's going to happen next in the video. And if you train the system to predict at the pixel level, it simply doesn't work. Um, because most of the details in video are not predictable. Most of the information contained in video is entirely unpredictable. If I take a video of this room, I point the camera in this side. I slowly rotate the camera. I stop here and I ask a system, fill in the blanks, tell me what's going to happen next in the video.

Probably going to predict the camera is going to continue rotating. Uh there's absolutely no way it can predict what every one of you looks like. There's just no information from which that he can use to make that kind of prediction. So if you train a system to make that prediction at a pixel level, you kill it because you basically are asking it to solve a test that is completely unsolvable. Now, of course, you're going to say, "But wait, we have systems that can do video generation."

Yes, they can, but they're going to make a prediction that is false. They're not going to predict what what you look like. They're just going to put random people there, right? um and those system don't really construct a good understanding of the underlying uh physics or reality of that. So the technique we've been working on for a number of years when I was at at fair and and since we created AM labs um is called JEA and it consists in training a system not to generate and predict all the details of a signal but to construct or learn an abstract representation of that signal of a video and then to make predictions in that abstract representation space.

uh and what the system does is that it eliminates all the information from that representation that is unpredictable and uh and that's the way we actually can apprehend the world as as humans or or as scientists. We we always have a a mental abstract representation of reality that allows us to make prediction and ignores all the >> all the underlying details. >> When do you prove that out? When do you release your first model from AMI?

>> Okay, so we know this idea works because we've been working on this for I've been working on this for the better part of the last 15 years, but I had a keynote at the the NIPS conference. It wasn't called NIPS yet um in 2016 where I said the future of AI is world models. So things that can predict given the state of the world at time t. So an abstract representation of the state of the world. Given an action that you imagine taking, a world model will predict the next state of the world that results from this action.

>> But you made that forecast in 2016. >> 2016. Yeah. I mean, I had this idea for a long time and it's this basic idea is not particularly new, but um but I said this is the way we should build intelligent systems. Have a system that can predict the consequences of its actions. Because if you have a system that can predict the consequences of its actions, it can plan a sequence of actions to accomplish a task, right?

Um you can say, well, you know, I want to move this uh uh this bottle from here to here, right? It's not a very complex task, but you can have an objective which is the bottle is in this location and then figure out like what sequence of action should I accomplish so that it actually puts the bottle here. And of course, there's an infinite number of possibilities. I could do this in my left hand or my right hand. I could, you know, turn the bottle around.

I could follow a funny trajectory. I could throw it or whatever. Probably not a good idea. So, can you plan a sequence of actions that you know will result in the result in the outcome that you want? And if you have such a mental model that predicts the consequences of your actions, you can do this planning. This is very classical in optimal control. This is not a new technique. It's called MPC model predictive control.

Okay, >> what's new is can we train this world model from data from observational data and that's basically how you know humans and animals train themselves. You know you justiculate when you're a baby and then you know you randomly hit objects and you realize you can hit an object. You realize also uh that if an object is not supported, it falls. In humans, it takes about nine months. It takes about nine months for infants to learn that about gravity and inertia.

Um so can we do the same with machines and we have we already had models for the last two years that basically already do this models called VJA VJA 2.1 VJ 2 VJA 2.1 they understand video they can tell you if something impossible occurs in a video so they have a little bit of common sense um the representation of video that they extract uh can be used for all kinds of downstream tasks with very minimal fine tuning so we know it works.

How much of a bottleneck is visual data? >> No, it's not a bottleneck. Not nearly as much as um as as text. When you take the totality of all the text available on the on the internet publicly, it's about uh 20 trillion words. You turn this into token, that's 30 trillion tokens. Um a token is two or three bytes. Let's say three. So that's 10^ the 14 bytes. That's a total amount of publicly available text on the internet.

Of course, the big AI companies are also training with synthetic data and licensed, you know, non-public data, but it's not a lot actually compared to um all the internet. So, 10^ the 14 bytes. >> Now, 10^ 14 bytes is just about the amount of bytes that gets to your visual cortex in the first four years of your life. Okay? Whereas the text, the publicly available text on the internet would take any of us at least 400,000 years to to read through, right?

Which is completely impractical obviously. So a 4-year-old has seen as much training data as the biggest >> except is video. So video is very redundant, you know, it's not as rich and and and sort of dense as as text. Uh but you learn a huge amount of basic knowledge about how the world works from this. You you you learn that the world is threedimensional. You learn that objects don't cannot disappear from one place and appear in another place.

That you know trajectories are continuous. You learn about inertia. You learn about gravity. You learn you learn about all kinds of stuff. You learn that if you push this bottle from the bottom it's going to slide. If you push it from the top maybe it's going to depending on friction it's going to flip. If I push on the table with the same strength, it's not going to move. Like all those things we take for granted, we think, you know, we don't need to be particularly smart for this.

And in fact, you don't. You know, a rat can figure this out. A cat, too. And a cat only has >> 800 million neurons. It's about 100 times less than the human brain. Um, but it's still a challenge for for computers. >> So, if if you succeed, let's pick one industry and see how it changes this industry. If if this succeeds, let's take defense because we're seeing companies like Andrew and Helsing in Germany, Stark in Germany, they are already in embedding AI and intelligence into hardware that's being deployed in theater as we speak.

What does success at world models mean for an industry like defense? >> Okay. So, what they're doing at the moment, a lot of uh uh defense or industry more generally, not just defense, but are examples of similar to what you cited at the beginning. So things like you know image uh understanding uh for like you know autonomous driving or or driving assistance or or collision detection for a car right your car stops automatically right uh so a lot of it is just perception.

>> So the next step is more than perception is action. Can we plan a sequence of actions to achieve a particular goal? plan a trajectory for a robot or a a sequence of mo motions for a robot to I don't know fry eggs or whatever. Uh that's a bit harder. And for for defense there is something that's relatively easy which is planning a a trajectory for a drone or something like this avoiding obstacles that already exist.

Um but there are a lot more kind of complex uh decision- making uh tasks to to accomplish in that context where world models really will play a role. >> You were a meta for 12 years. Why were you not able to do this with Mark? >> I totally was. I could I could have been I could have stayed at Meta and do this at at Meta. Um the the project which we called internally AMI advanced machine intelligence. Okay, it's the name of the company.

uh we we kept the same name. This had the support of Mark Zuckerberg and Andrew Bosworth the CTO and a bunch of other people in the leadership. So they they were really um you know kind of believed in the idea that eventually AI systems maybe of this type you know with that kind of blueprint. Um but what happened in the LA in 2025 was that you know a lot of the sort of AI organization within meta kind of focused itself on catching up with the rest of the industry for LLMs and there was bunch of people who really believe that you know AGI whatever you mean by that will emerge from scaling up LLM and training them on more data which I don't believe and I was publicly very vocal about this Um so there's a bit of a disconnect.

Um so despite the fact that you know the project had the support of of Mark um this probably was not the optimal place to do this. Also we were starting to get really good results with those those those system with the VJA etc. And and so we knew there were applications. Most of those applications would be for industry, industry process control, you know, controlling a complex system like a a power plant, a jet engine, uh, you know, pharmaceutical manufacturing uh plant or whatever or or a patient something like that, a human cell like you know basically modeling the and predicting the behavior of complex systems.

uh most of those applications are B2B and Meta is not a B2B company. It's a company that is entirely focused on connecting people with each other. Um so it made sense to leave start a new company and and go into high gear and for that project and then make it real. I know you think about AI sovereignty a lot as well and you're a leading member of the AI alliance and there's a program called tapestry and certainly for the industry in Europe there was a kind of gut punch moment with the export controls around mythos and fable talk to us about what sovereignty around AI should actually mean in practice for Europe where the focus should be >> so I think uh I mean clearly Europe is is behind in terms of the the race for LLMs that's basically in the hands of you know a few companies in Silicon Valley you know who have labs everywhere in the world, right?

Including in Paris and London. Uh, and of course a bunch of Chinese companies. Uh, and that ship has sailed. Uh, at least from the sort of traditional proprietary platform type view. The Chinese models are open. We don't know how long that's going to last. It may or may not last. Uh, and currently the Chinese models are the best open source ones. uh there's a lot of people uh in the world outside the US and China but also in the US who would like to have access to first of all an open and free um foundation model from which they can build a custom product through finetuning right um and those don't exist anymore there was llama from from meta that kind of took a of a nose dive and now basically it's not really uh pushed by uh by meta anymore.

Meta is becoming a little more clammed up like like Google couple years did a couple years ago. Uh so no credible provider of open AI platforms from the west anymore. um they're all Chinese and a lot of companies in the US, in Europe, in Asia are reluctant to adopt Chinese models. First of all, because they could be shut down tomorrow uh by by political decisions, but also because of issues of biases and things like that.

uh but we see a movement from the user proprietary platform towards open platforms uh mostly Chinese engines finial models because of you know recent actions of the US government and entropic and the know people realizing like we don't have any control on this if it's a basic infrastructure is just too risky to depend on it so project tapestry is an idea I've had for three years I talked about it internally at meta It didn't go very far.

I talked about it at the United Nations a few years ago again just a few months ago at the United Nation Open Source Week [snorts] and talk to various governments around the world about this idea. The basic idea is as follows. countries, regions, universities, private entities, nonprofit or for-profit could contribute to training a large scale foundation model L&M style um without having to release their data. They could use their own data, their own computing infrastructure and contribute to training a global model without actually communicating their data ever.

So maintaining sovereignty on that data. [snorts] There's a huge amount of training data that is not publicly available that would improve greatly the performance of uh foundation models. Think of u let me take a a non-random example India. India has 22 languages official languages and about 300 languages total plus some dialects. There's no way OpenAI, Anthropy, Google or Meta is going to train a model that speaks all of those languages, right?

So what is India to do? What they did a few years ago is that they set up an organization called Barat Gen which takes open source models and fine-tuned them so they they speak Indeek languages. uh in the past they were using lama now they're using one of the Chinese models uh they're not happy about it >> and they would like open platform same problem in Europe same problem in various countries Vietnam Kazakhstan Africa various various places Korea Japan all of those countries want a sovereign uh sovereign model >> and this is moving towards implementation now this isn't just a conversation >> that's right so two months ago we had a a kickoff meeting in Paris of the Tapestry project.

Um, and and the the it's it's it's a kind of a bottomup thing. Anybody in this room can if you have any kind of technical ability, you can sign up and contribute to the GitHub. Uh, the basic thing here is to build a software infrastructure to do distributed training, right? So every country in the world would you know scan their own digitize all of their cultural uh information whatever it is that resides maybe in their national libraries and things like this not publicly available yet and and have computing resources and then would train locally a a model on their data also using uh commonly available data and then periodically would send their parameter vectors to a central server if you want that would compute the average parameter vector of all the contributors.

And if you let this go enough with various tricks, you can get all of those systems to slowly converge towards that consensus model >> that uh is as good as if it had been trained in a single data center on all the data, right? >> But the the participants do not actually exchange that data. So now what you have is a is a financial model that actually could be much better than all the proprietary system because it would be trained on more data would be open and free and the valability of an open and free financial model in AI is indispensable.

This is something that will have to happen. And the reason it will have to happen is because all of our information, all the information we consume eventually is going to be consumed through the mediation of AI assistant. AI assistant like the one I'm wearing on my nose right now. These are the Rayban meta, right? Um there there is a assistant in this. I'm not going to ask it question, but I can take a picture of you guys.

All right. Okay. another one for this. Okay. Um I could have asked it to just take take the the picture. Um so all of our information diet is going to be mediated by AI systems. If those systems are a handful of proprietary engines from a handful of companies on the west coast of the US or China, this is the end of local culture. You know, the French are really attached to their language, right? I've heard >> and not not just the French.

>> Uh so >> it's that dramatic if we don't get this right. It's absolutely dramatic because uh you know it's loss of linguistic uh uh not just not just language but but all of culture right so the the the systems the people who can test in fact that's part of the tapestry project you can test to what extent an AI uh system is is culturally biased >> and every system is necessarily culturally biased there is no unbiased system that's impossible >> can I ask will will Amy you have offices in Singapore, SF, San Francisco in Paris.

Obviously, no >> I'm wrong. You don't you you you you're you're building yourself as an international. >> It's international. It's global. >> Will you have to pick a side at some point? >> No. >> I mean, unless you know, the US government goes goes all nutty on us, but >> Well, we're not Some would argue we're not too far. I mean, President M is concerned. >> No, it's a concern.

Uh it's you know possibly one of the reasons that Amid Labs is uh established in headquartered in in Paris. So Amid Labs has offices in Paris, New York. >> Okay. >> Montreal and Singapore. >> And we're really a global company. Our investors are from all over the world. Our industry partners/investors are from all over the world. 40% of our investors are from Europe. 33% from the US and 27 from Asia with one fund from from UAE.

So we're really a global company. Uh but we have zero presence on the west coast of the US, zero presence in China. >> Okay. And that's by design. >> That's by design for now. >> We are six minutes over time. I can't believe I have to wrap this up. But lastly, I want to squeeze this in. Um we started with the economic and social impacts that your first really key development back in the 1980s late 1980s uh CNN uh convolutional on neural networks had and still has and we're still benefiting from if you succeed with world models briefly what is the potential economic and societal impact at the at the higher end >> all the things that you see you expect AI should be able to do and doesn't do like domestic robots like level five self-driving cars which we still don't have uh like you know systems that have common sense that are perhaps as capable as smart as as humans in in many ways uh but as you know augmentations assistance to uh to uh to humans so AI for the physical world I don't think we can have AI systems that are you know generally uh useful unless they can understand the real world I don't think you can have agentic systems that are reliable unless unless they have the ability to predict the consequences of their actions.

In other words, unless they have world models. And so it's going to be a a new AI revolution. That's what we're building on. And Europe has not lost that that race. It's that is going right now. >> Europe's lost on the frontier, but it can win this race. >> Europe can absolutely it's not going to be a win. you know, you you never you know, it's always a constant, you know, constant progress. >> Okay.

>> Uh so it's not like you you win because you discover one secret that nobody else has. >> We will have a seat at the table. We'll have a seat at the table. >> We have more than that, I think. Uh and right now, Silicon Valley is a bit in a kind of a trench that everybody is digging. They're all kind of working on the same thing. Nobody can deviate from the the mainstream LLM because if they do they run the risk of falling behind and so um that's an opportunity to seize for people who have alternative approaches.

>> Okay. >> Like us. >> Yan Yan Lon, thank you very much indeed. I won't call it the godfather. The OG. The OG. Thank you. Thank you. >> Thank you both.