Yann LeCun:信息瓶颈|EP20
EP20: Yann LeCun | The Information Bottleneck

Hi Anne, and welcome to the information bottleneck. Um, I have to say this is a bit weird for me, like I've known you for almost 5 years and we have worked closely together, but this is the first time that I'm interviewing you for a podcast, right? Um, usually our conversations are more like, "Yann, it doesn't work. What I should do?" Um, okay, so uh, even though I'm sure all of our audience knows you, uh, I will say Yann LeCun is a Turing Award winner, one of the godfathers of deep learning, the inventor of convolutional neural networks, founder of Meta's fundamental AI research lab, and still their chief AI scientist, and a professor at NYU.
So, welcome. Pleasure to be here. Um, yeah. And it's a pleasure for me to be anywhere near you. Um, I have been, you know, in this industry for a lot less time than either one of you and doing research for a lot less time. So, uh, the fact that I'm able to publish papers somewhat regularly with Rafeed uh, has been an honor and to be able to uh, start hosting this podcast has been even more of one. So, it's really a pleasure to sit down with you.
Awesome. Um, yeah, so we thought I congra- congratulations on the the new startup, right? Uh, you recently announced that they after 12 years at Meta uh, you're starting a new startup, Advanced Machine Intelligence, uh, that you'll focus on world model. Um, so first of all, how does it feel to be in the in the other side? Going from a big company uh, to starting something from scratch. Well, I co-founded companies before.
Uh, I was, you know, involved more peripherally than uh, than this new one, but uh, but I, you know, I know I know how this works. What's uh, unique about this one is a new phenomenon where um, there is enough hope from the part of investors that, you know, AI will have a big impact that they're willing to invest a lot of money essentially, which means now you can create a startup where you know, the first couple years are essentially focused on research.
Uh, that just was not possible before. Like, you know, the only place to do research in industry before was in a large company that was you know, not fighting for its survival and basically had a dominant position in this market and had a, you know, long enough view that they they were willing to to fund long-term projects. Um, so from, you know, history the the big labs that we we remember um, like Bell Labs belonged to AT&T, which basically had a monopoly on telecommunication in the US.
Uh, you know, IBM had a monopoly on the computers essentially, you write and they had a good research lab. Xerox had a monopoly on photocopiers and that enabled them to fund PARC. Did not enable them to profit from the research going on there, but that profited Apple. Um, and then more recently Microsoft Research, Google Research and FAIR at at Meta. Um, and and the industry is, you know, shifting again. Uh, FAIR had a big influence on AI the AI research ecosystem by essentially being very open, right?
Publishing everything, open sourcing everything with and and with tools like PyTorch, but also like research prototypes that a lot of people have have been using uh, in industry. So, we caused other labs like Google to become more open and and other labs to also kind of publish much more systematically than before. But what's been happening over the last couple years is that um, a lot of those labs have been kind of clamming up and becoming more secretive.
Uh, and uh, that's certainly the case. I mean, that was the case for OpenAI some years ago and uh, and and now now Google is becoming more closed and possibly even Meta. So, um, yeah, I mean, it was it was time for the the the type of uh, stuff that I'm interested in to kind of um, do it outside Meta than inside. So, so to be clear then, does uh, AMI, Advanced Machine Intelligence, plan to do their research uh, in the open?
Yeah, I upstream research. I mean, in my opinion you can't really call it research unless you publish what you do because otherwise you can get easily fooled by yourself. Um, you know, you you come up with something you think is the best thing since sliced bread. Okay, if you don't actually submit it to the rest of the community, you might just be delusional. Uh, and I've seen that phenomenon many times, you know, in lots of industry research lab where there's sort of internal hype about, you know, some internal projects um, but that kind of realizing that other people are doing things that actually are better, right?
So, so if you if you tell the scientists like, you know, publish your your your work uh, first of all, that is an incentive for them to do better work that is more you know, where the methodology is kind of more thorough and the results are kind of more reliable, the research is more reliable. Um, it's good for them because very often when you work on a research project the impact you may have on product could be months, years, or decades down the line.
And you cannot tell people like, you know, come work for us. Don't say what you're working on and maybe there is a product you will have an impact on 5 years from now. Like in the in the meantime, like they can't be motivated to really do something uh, useful. So, if you tell them that they tend to work on things that have a shorter impact, right? So, if you really want breakthroughs you need to let people publish. You can't do it any other way.
And this is something that a lot of the industry is forgetting at the moment. Does AMI, like what what products if any does AMI plan to to produce or make? Is it research or more than that? No, it's more than that. It's actual products. Okay. Um, but uh, you know, what things I have to do with with, you know, world models and, you know, planning and and basically with the ambition of uh, becoming kind of one of the main suppliers of intelligent systems down the line.
We think the the current architectures that are employed, you know, LLMs or, you know, agentic systems that are based on LLMs um, work okay for language. Um, if agentic systems really don't work very well. Uh, they require a lot of data to basically clone the behavior of of humans. And they're not that reliable. Um, so we think the proper way to handle this, and I've been saying this for almost 10 years now, is uh, is have world models that are capable of predicting what would be the consequence or the consequences of an action or sequence of actions that an AI system might take.
And then the system arrives at a sequence of actions or an output by optimization, by figuring out what sequence of actions will uh, optimally accomplish a task that I'm, you know, setting for myself. That's planning. Okay? So, I think and the central part of intelligence is being able to predict the consequences of your actions and then use them for planning. Um, and that's what we're that's what I've been working on for many years.
Uh, we've been making fast progress uh, with, you know, a combination of projects here at NYU and also at Meta. Um, and now it's it's time to basically make it make it real. And what do you think are the missing parts? Like and why think it's taking so long? Because you're talking about it as you said, like for many years already, but it's still not better than LLMs, right? It's not the same thing as LLM, uh, modalities that are high dimensional, continuous, and noisy.
And LLMs completely suck at this. Like, they really do not work, right? If you try to train an LLM to kind of learn good representations of images or video, they're they're really not that great. Um, you know, generally vision capabilities for for AI AI systems, right? Are trained separately. They're not part of the whole LLM uh, LLM thing. So, yeah, if you want to to handle data that uh, high dimensional, continuous, and noisy, you cannot use generative models.
You can certainly not use generative models that tokenize your data into kind of discrete symbols. Okay? It's it's just no way. We have a lot of empirical evidence that this simply doesn't work very well. What does work is learning an abstract representation space that eliminates a lot of details about the input, essentially all the details that are not predictable which includes noise and make predictions in that representation space.
Uh, and this is the idea of GEPA, right? Joint Embedding Predictive Architectures, which, you know, you are as familiar with. Yeah. >> With as I as you worked on this. Yeah. So, uh, Also, Rondel was a hosted in the past cuz Rondel was also in in the podcast. We probably talked about this at length. So, so there's a lot of ideas around this and let me tell you my history around it. So, okay. Um I um I've been convinced for a long time uh probably the better part of 20 years that the the proper way to building intelligent systems was through some form of unsupervised learning.
I started working uh on unsupervised learning as the basis for you know, um making progress uh in the early 2000s. mid-2000s. Before that, I wasn't so convinced this was the way to go. And and basically this was the idea of uh you know, training autoencoders to learn representations, right? So, you have an input, you run it through an encoder, it finds a representation of it, and then you decode, so you guarantee that the representation contains all the information about the input.
Does that That intuition is wrong. Like insisting that the representation contains all the information about the input is a bad idea. Okay. I didn't know this at the time. So, what I worked on was um you have several ways of doing this, you know, Jeff Hinton at the time was working on restricted Boltzmann machines. Um uh Yoshua Bengio was working on denoising autoencoders, which actually became quite successful in different contexts, right?
For NLP among others. And I was working on sparse autoencoders. So, basically you you know, if you train an autoencoder, you need to regularize the the representation so that the autoencoder does not trivially learn an identity function. And this is the information bottleneck. Uh podcast, this is about information bottleneck. All right? You need to create an information bottleneck to limit the information content of the representation.
And I thought high-dimensional sparse representations was actually a good way to go. So, um so, much of my students used did their PhD on this. Uh Koray Kavukcuoglu, who's now a chief architect at DeepMind uh at uh Alphabet, and also the CTO at DeepMind, actually did his PhD on this with me. And uh you know, a few a few other a few other folks, Macaire Onzato and uh Ian Goodfellow and a few others. So, so this was kind of the idea.
And then as it turned out, and the idea the reason why we worked on this was because we wanted to pre-train very deep neural nets by pre-training those things as autoencoders. We thought that was the way to go. What happened though was that we started like, you know, experimenting with things like uh normalization, uh rectification instead of hyperbolic tangent or sigmoids. Uh the gradients. Mhm. Um and that ended up you know, basically allowing us to to train fairly deep network completely supervised, so self-supervised learning.
And and this was at the same time that data sets started to get bigger. And so, it turned out like, you know, supervised learning worked fine. Um so, the whole idea of self-supervised or unsupervised learning was put put aside. Um And and then came ResNet, and you know, that sort of completely solved the problem of training very deep architectures, right? Uh in 2015. Um but then in 2015, I started, you know, thinking again about like, how how do we push towards like human-level AI, which really was the original objective of FAIR really, and and my objective my life, you know, mission.
Um and realized that, you know, all the approaches of reinforcement learning and and things of that type were basically not scaling, you know, reinforcement learning is incredibly inefficient in terms of samples. And so, this is not was not the way to go. Um And and so, the idea of world models, right? A system that can predict the consequence consequences of its action and can plan. I started really seriously playing with this around 2015-16.
My keynote at at uh what was still called NIPS at the time Mhm. uh in 2016 was on world model. I was arguing for it. Um That was basically the centerpiece of my talk was like, this is what we should be working on, like, you know, world models, action-condition. And a few of my students started working on this on video prediction and things like that. I we had uh some papers on video prediction in 2016. And um I made a the same mistake as as before and the same mistake that everybody is doing at the moment, which is training a video prediction system to predict at the pixel level.
Um which is really impossible. And you can't really represent useful probability distributions on the space of video frames. Um and so, those things don't work. I knew for a fact that because the prediction was non-deterministic, we had to have a model with latent variables. Mhm. To represent all the stuff you don't know about the the variable you're supposed to predict. And so, we experimented with this for years. I had a a student here who's now a a scientist at at FAIR, Mikael Andrawes, who developed a video prediction system with latent variables.
Um And it kind of solved those problems we were facing slightly. I mean, today the solution that a lot of people are employing is uh diffusion models, which is a way to train a non-deterministic function essentially, or energy-based models, which I've been uh advocating for decades now, which also is another way of training non-deterministic functions. Um but in the end, I discovered that this was all a bad idea, that the really the the way to um get around the fact that you you you can't predict at the pixel level is to just not predict at the pixel level.
It's to It's to learn a representation and predict at the representation level, Mhm. eliminating all the details you cannot predict. Uh and I I wasn't really thinking about those methods early on because I thought there was a huge problem of preventing collapse. Um so, I'm sure Rondel talked about this, but um you know, when you train uh let's say you have an observed variable X, and you're trying to predict a variable Y.
But you don't want to predict all the details, right? So, you run both X and Y through encoders. So, now you have both a representation for S for X as X, and a representation for Y as Y. You can train a predictor to produce, you know, predict the representation of Y from the representation of X. But if you want to train this whole thing end-to-end simultaneously, um there's a trivial solution where the system ignores the input and produces constant representations.
And the predictor's problem now is trivial, right? So, if your only criterion to train the system is minimize the prediction error, it's not going to work. It's going to collapse. I knew about this problem from a very long time because I worked on joint embedding architectures, we used to call them Siamese networks back in the '90s. >> Those are the same because people have been using that term Siamese networks even recently.
That's right. I mean, the the concept is still, you know, up-to-date, right? So, you have you have an X and a Y, and think of the X as some sort of degraded, transformed, or corrupted version of Y. Okay? Uh you run both X and Y through encoders, and you tell the system, "Look, X and Y really are two views of the same thing. So, whatever representation you compute should be the same." Right? Um So, if you just train a neural net, um you know, two neural nets with shared weights, right?
To produce the same representation for slightly different versions of the same object, view, whatever it is, uh it collapses. It doesn't produce anything useful. So, you had to find a way to make sure that the system, you know, extract as much information from the input as possible. And the original idea that we had, you know, it was a this paper from 1993 with Siamese net was to have a contrastive term, right? So, you have other pairs of samples that you know are different.
And you train the system to produce different representations. So, you have a cost function that attracts the two representations when you show it two examples that are identical or similar, and you repel when you show it two examples that are that are dissimilar. And we came up with this idea because someone came to us and said like, "Can you encode signatures um of someone, you know, drawing a signature on the tablet.
Can you encode this on less than 80 bytes?" Because if you can encode it in less than 80 bytes, we can uh write it on the magnetic uh tape of uh of a credit card. So, we can do signature authentication for credit cards. Right? And so, we came up with this idea. I came up with this idea of training a neural net to produce 80 variables that were quantized one byte each, and then training training it to kind of do this the same.
And did they use it? Uh so, it worked really well, and they showed it to their, you know, business people who said um "Oh, we're just going to ask people to type PIN codes." Uh We have even less than of like that, like, how you can integrate that technology. Right. And you know, I knew this thing was kind of fishy in the first place because like, you know, there were countries in Europe that were using smart cards, right?
And there was a much better solution. But they just didn't want to use smart cards for some reason. Anyway, so so we had this uh technology in in the mid-2000s, uh I I worked with two of my students on to revise this idea. We came up with some new objective functions to train those. So, these are what people now call contrastive methods. This is a special case of contrastive methods. We have like positive examples, negative examples, and you train, you know, on positive examples, you train the system to have low energy, and for negative samples, you train them to have higher energy, where energy is the distance between the representations.
So, I had two papers at CVPR in 2005-2006 by Raia Raia Haddad Sal who is now the uh DeepMind Foundation, the the the sort of fair-like division of DeepMind, if you want. Uh and uh Sumit Chopra, who is actually a faculty here at NYU now, working on medical imaging. And so, um this gathered a bit of interest uh in the community and sort of revived a little bit of work on this on those ideas. But it still wasn't working very well.
Those contrastive methods really um were producing representations of images, for example, that were kind of relatively low-dimensional. If we measured like those, you know, spec the eigen value spectrum of the contrast matrix and the representations that came out of those things, it would fill up maybe 200 dimensions, never more. Like even training on ImageNet and things like that, even with data augmentation. And so, that was kind of disappointing.
And uh it it did work okay. There was a bunch of papers on this. And it worked okay. Um There was There was one paper from DeepMind, SimClear, that that demonstrated you could get decent performance with um contrastive training applied to SimSiam. Um but then, about 5 years ago, um uh one of my postdocs, uh Stefan Deny at at Meta, um tried an idea that um at first I didn't think would work, which was to essentially have some measure of information quantity that comes out of the encoder, and then trying to maximize that.
Okay. And the reason I didn't think it would work is because I'd seen a lot of experiments along those lines that Geoff Hinton was doing in 1980s, of trying to sort of maximize information. You can never maximize information because you never have uh appropriate measures of information content. That is, a lower bound. If you want to maximize something, you want to either be able to compute it or you want to a lower bound on it so you can push it up, right?
And for information content, we only have upper bounds. So, I always thought this was completely hopeless. And then, uh you know, Stefan kind of, you know, came up with a technique and which was was called um uh Barlow Twins. Uh Barlow is a famous uh theoretical neuroscientist who came up with the idea of information maximization. And uh and it kind of worked. That was wow. Uh so, then I said, like, we have to push this, right?
So, we come up with another method with uh Simon and I, um Adrien Bard, and Jean Pons, who is affiliated with with NYU, too. Uh technique called VICReg, variance-invariance-covariance regularization. Uh and that turned out to work to be simpler and work even better. And since then, we've made progress and Randall recently uh you know, I discussed an idea with him that he kind of pushed and made practical. It's called SimReg.
Uh the whole system is called a Jepa. Okay, he's responsible for the name. I don't know. Uh The latent latent Euclidean Jepa, right? Yeah. Um Uh and and and and SimReg has to do with sort of uh making sure that the uh distribution of of vectors that come out of the encoder is an isotropic Gaussian. That's the I in the G. Mhm. So, uh I mean, there's a lot of things happening in this domain, which uh are really cool. I think there's going to be some more progress over the next uh year or two.
Uh we get a lot of experience with this. And uh and I think that's kind of a really good um promising set of techniques to train models that learn abstract representations, which I think is key. And what do you think are the missing parts here? Like, do you think like more compute will help or like we need better algorithms or like it's kind of like do you believe in the bitter lessons, right? Like, do you think Well, and and furthermore, what do you think about, you know, the data quality problems with the internet post-2022, right?
I've heard people compare it to low-background steel now to refer to all that data before LLMs came out. Like, low-background tokens, I mean. Okay. >> Yeah. I think I'm totally escaping that problem. Okay, here is the Here is the thing. And I've I've been, you know, using this argument publicly uh over the last couple of years. Uh training an LLM, uh if you wanted to have any kind of, you know, decent performance, requires training on basically all the available freely available text on the internet, plus some, you know, synthetic data, plus licensed data, etc.
So, a typical LLM, like, you know, number three, you know, going back uh a year or two, is trained on 30 trillion tokens. A token is typically three bytes. So, that's 10 to the 13 bytes for pre-training. Okay, we're not talking about fine-tuning. Um 10 to the 14 bytes. And for the LLMs to be able to really kind of exploit this, um they need to have a lot of memory storage. Because basically, those are isolated facts.
There is a little bit of redundancy in text, but but a lot of it is just isolated facts, right? Um and so, you need a lot of you need very big networks because you need a lot of memory to store all those facts. Uh and we regurgitate them. Okay, now compare this with uh video. Um 10 to the 14 bytes, if you count uh 2 megabytes per second uh for video, for, you know, relatively compressed video, not highly compressed, but a bit.
That would represent 15,000 hours of video. 10 to the 14 bytes. If 15,000 hours of video, you'll have the same amount of data as the entirety of all the text available on the internet. Now, 15,000 hours of video is absolutely nothing. It's 30 minutes of YouTube uploads. Okay. It's the amount of visual information that a 4-year-old has seen in uh his or her life. The entire life, waking time, is about 16,000 hours in 4 years.
Um it's not a It's not a lot of information. We have video models now, uh V VJepa, VJepa 2, actually, that just came out last summer. Um That was trained on the equivalent of a century of video data. It's still probably data. Okay. Much more data. But much less than the biggest LLM, actually. Because even though it's it's more bytes, it's more redundant. So, you say, "Okay, it's more redundant, so it's less useful."
Actually, when you use self-supervised learning, you do need redundancy. You cannot learn anything in self-supervised or anything, by the way, uh if it's completely random. Redundancy is what you can learn. And so, um uh so, there's just much richer structure in uh you know, real-world data like video, than there is in text, which kind of led me to claim that we absolutely never ever going to get to human-level AI by just training on text.
It's just never going to happen, right? So, there's It's It's a big debate in philosophy of whether AI should be grounded in reality or whether it could be just, you know, in the realm of symbolic manipulation and things like this. Well, and when we talk about for world models and grounding, I think, you know, there's still a lot of people who who don't even understand what the idealized world model is, in a sense, right?
So, for example, I'm influenced by having watched Star Trek, which I would hope you've seen a little bit of, and you're thinking of the Holodeck, right? I always thought that the Holodeck was like an idealized perfect world model, right? Even so many episodes of going too far, right? People walking out of it, right? But it, you know, it even simulates things like smell and physical touch. So, do you think that something like that is like the idealized world model or do you think like a different model or or like way of defining it would be?
>> Okay, this is an excellent question. Uh and the reason it's excellent is because it pushes it goes to the core of of really what uh you know, what I think we should be doing, we can doing, and you know, how wrong I think everybody else is, okay? So, so, people think you know, think that a world model is something that reproduces all details of what the world does. They think of it as a simulator. Yeah. >> Right?
And of course, because, you know, deep learning is the thing, you're going to use some deep learning system as a simulator. A lot of people also are focused on video generation, which is kind of a cool thing, right? You you produce those cool videos and they're wow, you know, people are sort of really impressed by them. Now, there's no guarantee whatsoever that when you train uh a video generation system, it actually has an accurate model of the underlying dynamics of the world and it's learned anything, you know, particularly abstract about it.
Mhm. >> Um so, the idea that somehow a model needs to reproduce every detail of the of reality is wrong and hurtful. Mhm. And and I'm going to tell you why, okay? Uh A good example of simulation is uh CFD, computational fluid dynamics. It's used all the time. People use supercomputers for that, right? So, you want to simulate the flow of air around an airplane, you, you know, cut up the space into little cubes, uh and within each cube, you you have uh a small vector that represents the state of that cube, which is uh you know, velocity, density or mass, and uh temperature, and maybe a couple of other things, right?
So, and then you solve Navier-Stokes equations, which are which is a differential uh partial differential equation. Uh And you can simulate the flow of air. Now, the thing is this does not actually necessarily solve the equations very accurately. If you have chaotic behavior like turbulences and stuff like that, simulation is only, you know, approximately correct. Um but in fact, that's already an abstract representation of the underlying phenomenon.
The underlying phenomenon is molecules of air that bump into each other and bump on the on the wing and on the airplane, right? But nobody ever goes to that level to do the simulation. That would be crazy. Right? It would require an amount of competition that's just insane. And it would depend on the initial condition. I mean, there is all kinds of reasons we don't do this. And maybe it's not molecules. Maybe it's, you know, at a lower level we should simulate particles.
And like, you know, do the Feynman diagrams and simulate and you know, all the different paths that those particles are employing because they don't take one path. It's not classical, it's quantum, right? Yeah, yeah. Um so, at the bottom it's like quantum field theory and probably already that is an abstract representation of the underlying reality. So, um so, you know, everything that takes place between us at the moment, in principle can be described through quantum field theory, okay?
We just have to measure the wave function of the universe in a you know, a cube that, you know, contains all of us and even that would not be sufficient because there are entangled particles at the other side of the universe that, you know, we have. So, it wouldn't be sufficient. But let's of uh of the argument. First of all, we would not be able to measure this uh wave function. Um and second of all, the amount of competition we would need to devote to this is absolutely gigantic.
It was would be some gigantic quantum computer that, you know, is the size of the earth or something. Um so, no way we can we can describe anything at that level. And it's very likely that our simulation would be accurate for maybe a few nanoseconds. Yeah. You know, beyond that we'll we'll diverge uh from reality. So, what do we do? We invent abstractions. We invent ab- abstractions like particles, atoms, molecules.
In the living world, it's proteins, organelles, cells, organs, organisms, societies, ecosystems, etc., right? Um and basically every level in this hierarchy ignores a lot of details about the level below. And what that that allows us to do is make longer-term, more reliable longer-term predictions. Okay? So, we can describe the dynamics between us now in terms of the underlying science and in terms of psychology. Okay?
That's a much, much higher level of abstraction than particle physics, right? Um and in fact, you know, every level in the hierarchy I just I just uh mentioned is a different field of science. A field of science is essentially defined by the level of abstraction at which you start making predictions, right? That you and I have to use to make predictions. Um in fact, physicists have this down to an art uh in the sense that um you know, if I give you a a box full of uh gas, um you could in principle simulate all the molecules of the gas, right?
Uh but nobody ever does this. But at a very abstract level, we can say, you know, PV = nRT, right? You know, pressure times volume equals, you know, number of particle times you know, temperature, blah blah blah. And uh so, you know that at a you know, global emergent phenomenological level, if you increase the pressure, the temperature will go up. Or if you increase the temperature, the pressure will go up, right? Uh or if you let some particles out, then the pressure will go down and and blah blah blah, right?
So, um so, we all the time we build phenomenological models of something complicated by ignoring all kinds of details that physicists call entropy. Um but uh but it's really systematic. That's the way we understand the world. We, you know, we we do not memorize every detail of uh we certainly don't reconstruct it of what we perceive. So, world models don't have to be simulators. >> No. Well, they are simulators, but in abstract representation space.
And what they simulate is only the relevant part of reality. Okay? If I ask you, where is Jupiter going to be 100 years from now? Uh I mean, we have an enormous amount of information about Jupiter, right? But within this whole information that we have about Jupiter, to be able to make that prediction where Jupiter is going to be 100 years from now, you need exactly six numbers. Three positions and three velocities. And the rest doesn't matter.
Mhm. So, you don't believe in uh synthetic data sets? I do. No, it's useful. Uh you know, data from games. I mean, there's certainly a lot of things that you learn uh from synthetic data from, you know, from from games and things like that. I mean, you know, children learn a huge amount from from play, which basically are kind of simulations of you know, the the world a little bit, right? But but in conditions where they can't kill themselves.
Mhm. But I I I worry at least for video games that, for example, the green screen like actors doing the animations, they're doing extremely it's designed to look good, you know, for like an often badass, I guess, for an action game. But these often don't correspond very well to reality. And so, I I worry that like a physical system that's in, you know, been trained or through with the assistance of world models might get similar quirks at least in the very short term.
Is this something that worries you? No, it depends at what level you trained them. So, for example, I mean, sure, if you use a very accurate robotic simulator, for example, right? It's going to accurately simulate the dynamics of an arm. Uh you know, when you apply torques to it, it's going to move in a particular way. There's dynamics, no problem. Now, simulating the friction that happens, you know, when you grab an object and manipulate it, that's super hard to do it accurately.
Friction is very hard to simulate. Okay? And so, those simulators are not particularly accurate for uh manipulation. They're good enough that, you know, you can train a system to do it and then you can do, you know, sim to real uh with a little bit of adaptation. Uh so, that that can work. But but that's not I mean, the point is much more important. Like, for example, there's a lot of completely basic things about the world that we completely take for granted, which we can learn at a very abstract level, but it's not language related.
Okay? So, the fact, for example, and I've used this example before and people have made fun of me for it, but it's really true. Okay? I have those objects on the table. Mhm. And the fact that when I push the table, the object moves with it. Like, this is something we learned. It's not something that you're born with. Okay? >> Mhm. The fact that most objects will fall when you let let them go, right? With the gravity.
Maybe it's learned this around the age of 9 months. Uh and the reason people make fun of me with this is because I said, you know, LLMs don't understand this kind of stuff, right? Uh and they and they absolutely do not even today. But but you can train them to give the right answer when you ask them a question, you know, if I put an object on the table and I push the table, uh what will happen to the object? It will answer, the object moves with it.
But because it's been fine-tuned to do that, okay? So, it's more like regurgitation than sort of real understanding of the underlying dynamics. But if you look on items so uh like nano Yeah. um a nano banana, they they have a good physics of the world, right? They are not perfect, but >> That's not physics, yeah. They have some physics. But so, do you think like we can't push it farther or do you think like uh it's a one way to to learn uh physics?
So, all of those models actually make predictions in representation space. They use uh diffusion transformers. And that prediction that um the the the competition of of the the video snippet at an abstract level is done in representation space, okay? Not always And then there's a second diffusion model that turns these abstract representations in into a nice-looking video. And that might be mode collapse. We don't know, right?
Because we can't really measure like the coverage of such systems with reality. Um but but like, you know, the to the the the previous point, I can train like, here is another completely obvious concept to us that we don't even imagine that we learn, but we do learn it. A person cannot be in two places at the same time. Mhm. Okay? Like, we learned this because very early on we learn object permanence, the fact that when an object disappears, it still exists.
Okay? And we we are peers as the same object that you saw before. Mhm. Um how can we train a an AI system to learn this concept? So, object permanence, you know, you just show it a lot of videos where objects, you know, go behind the screen and then reappear on the other side or where they go behind the screen and the screen goes away and the object is still there. And when you show 4-month-old babies scenarios where things like this are violated, their eyes open like super big and they're like super surprised because reality just, you know, violated their internal model.
Uh same thing when you show a scenario of like a little car on a platform, you push it off the platform and it appears to float in the air. Um they also look at it, you you know, 9-month, 10-month-old babies look at it like really surprised. 6-month-old babies barely pay attention cuz they haven't learned about gravity yet. Uh so, they haven't been able to like, you know, incorporate the notion that every object is supposed to fall.
So, this kind of learning is really what's uh what's uh what's important. And you do this you can learn this from very abstract things. You know, the same way uh babies learn about like you know, social interactions by, you know, being told stories with like simple pictures. It's a simulation. Uh an abstract simulation of the world. But it sort of learns them, you know, particular behavior. So you could imagine like training a system from, let's say, an adventure game, like a top-down 2D adventure game.
Mhm. Where, you know, you you you tell your character like, you know, move north and he goes to the other room and he's not in the first room anymore because he moved to the other room. Right? Now, of course, in adventure games you have Gandalf that you can call and he just appears, right? So that's not physical. But um but like when you pick up a a key from a you know, from a treasure chest, you have the key, no one else can have it, and you can use it to open a door.
Like there's a lot of things that you learn that are very basic, you know, even in sort of abstract uh environments. Yeah, and I I just want to observe um that some of those adventure games that they try to train models and one of them you might have know about is NetHack, right? And NetHack is fascinating because it is an extraordinarily hard game. Like ever ascending in that game without cheats is like 20 years without, you know, going to the wiki.
People still don't do it from playing. And my understanding is that uh AI agents, the very best agent models we have or even world models are pathetic. Totally. Yeah. Yeah. Yeah. So to the point people have come up with sort of you know, dumbed-down version of NetHack. It's called MiniHack. MiniHack. Exactly. MiniHack. They had to dumb it down just for for AI agents. So some of my, you know, colleagues have been working with is actually one of my master's students who is working with me.
So uh and and you know, my colleague NAF, who I mentioned earlier, has been also doing some work there. Now, what's interesting there is that uh uh there there's a type of situations like this where you need to plan, okay, but you need to plan in the presence of uncertainty. The problem is you know, all games and adventure games in particular is that you don't have complete visibility of the state of the system. You don't know the map in advance.
You need to explore and blah blah blah. You can get killed every time you do this and you know, uh but the actions are essentially discrete. Yes. Okay. The finite number of possible actions is turn-based. Uh and so in that sense it's like chess except it's it's not, you know, chess is fully observable. Go also is fully observable. Stratego isn't though. uh poker is not. Uh and so it makes it more difficult if you have uncertainty, of course.
Uh Uh but those are games where the the number of actions you can take is discrete. And basically, you know, what you need to do is is uh do tree exploration, okay? And to do, of course, the tree of possible states, you know, goes exponentially with the number of of moves. And so you have to have some way of generating only the moves that are likely to be good and basically never generate the other ones or select them down.
And you need to have a value function, which is something that tells you, okay, I can't plan to the end of the game, but even though I'm planning only sort of nine moves ahead, I have some way of estimating whether evaluating whether a position is good or bad. It's going to lead me to, you know, a victory or a solution, right? So you need those two components basically, something that guesses what the good moves are and then something that uh you know, essentially uh evaluates uh ends.
And if you have those both of those things, you can train those functions using something like reinforcement learning or or um or behavioral cloning if you have uh data. I mean, the basic idea for this goes back to Samuel's uh checker players from 1964. It's not recent. Um but but of course was, you know, the power of it was demonstrated with, you know, you know, AlphaGo and and and AlphaZero and things like that. So that's good.
But that's a domain where humans suck. Humans are terrible at playing chess, right? And playing Go. Like machines are much better than we are. Yeah. Um because of the speed of tree exploration and because of the memory that's required for for tree exploration. We just don't have enough memory capacity to do breadth-first uh tree exploration. So we suck at it. Like, you know, when AlphaGo came out, uh you know, people before that thought that the best human players were maybe two or three stones handicap like below an ideal player that you call God.
Does that No, like, you know, humans are terrible. Like we, you know, the best players in the world need like eight or nine stones to to get Well, I I can't believe I I get the pleasure to talk about game AI with with Y I I just have a few follow-up questions on this. The first one is this example that you talk about around um humans being terrible at chess. And I I'm familiar a bit with the development of chess AI over the years.
Um I'm Do you you know, I've heard this referred to as Moravec's paradox and explained as, you know, humans have evolved over billions or millions, sorry. It's a large N number of years to uh physical locomotion and that's why babies and humans are very good at this, but we have not evolved at all to play chess. So that's one question. And then a second question that's related is a lot of people today who play video games, and I'm one of them, have have observed that it feels like AI, at least in terms of like enemy AI, has not improved really in 20 years, right?
That some of the best examples are still like Halo 1 and Fear from the early 2000s. So when do you think that, you know, advancements that we've been doing in the lab are going to actually have real impact on like gamers, you know, and and in a non-like generative AI sense, right? I I used to be a gamer. Never a like a addicted one, but um my family is in it cuz my I have three sons in their 30s and they have a video game design studio between them.
So um so I was sort of, you know, embedded in that culture. Uh But uh yeah, no, you're right. Uh and you know, it's it's also it's also true that the, you know, despite the accuracy of physical simulators, a lot of the a lot of those simulations are not used by studios who make uh animated movies because they want control. Uh they don't necessarily want accuracy, they want control. And in games it's really the same thing.
It's a creative act. What you want is some control about the course of the story or the way that, you know, uh in NPC kind of behavior and all that stuff, right? Uh And and AI kind of you know, it's difficult to maintain control at the moment. So I mean, it it will come, but uh um you know, there's there's some resistance from the creators, but Right. I think Okay, Moravec's paradox is is very much still in force. So Moravec I think I think formulated it in 1988 if I remember correctly.
Mhm. He said like, yeah, how come you know, things that we think of as uniquely human intellectual tasks like playing chess, we can do with computers or or, you know, computing integrals or whatever. Um But the thing that we take for granted, we don't even think is an intelligent task, like what a cat can do, we still can't do with robots. Mhm. Mhm. And even now, um 47 years later, we still can't do them well. I mean, of course, we can, you know, train robots, you know, by imitation and a bit of reinforcement learning and, you know, by training through simulation to kind of locomote and, you know, avoid obstacles and do various things.
But they're not nearly as inventive and creative and and uh you know, agile as as a cat. It's not because we can't build a robot. We certainly can. It's just we can't we can't make them smart enough uh to do all the stuff that a cat or even a mouse can do. Let alone the dog or monkey. Right? So uh so you have all those people gloating about like, you know, AGI in in a year or two. It's completely diluted. It's just complete delusion.
Cuz the real world is way more complicated. And you're not going to get it. You're not going to get anywhere by tokenizing the world and and using LLMs. It's just not going to happen. Um so um So what is your uh timelines? When will When will we see like, I don't know, AGI, whatever it means, or like And and also where are you on the And and where are you on the optimist-pessimist side? Because, you know, there's some doomers among or doomerism amongst like Gary Marcus and and I think well, I guess he's a critique he's critiques it Sorry, the doomer would be Is it Joshua?
Yeah, there you go. Like where do you fall on all these things? Okay, I'll answer the first question first. Okay. Uh So first of all, there is no such thing as general intelligence. This concept makes absolutely no sense because it's really designed to designate human-level intelligence. But human intelligence is super specialized. Okay? We can handle the real world really well, like navigate and blah blah blah. We can handle other humans really well because we evolved to do this.
And chess we suck. Okay? So And there's a lot of tasks that we suck at that where a lot of other animals are much better than we are. Okay? So what that means is that um we are specialized. We think of ourselves as being general, but it's simply an illusion because all of the problems that we can apprehend are the ones that we can think of. Yeah. And vice versa. And so we're general in all the problems that we can imagine.
Okay? But there's a lot of problems that we cannot imagine. Um And there's some mathematical arguments for this, which I'm not going to go into unless you ask me. But, um so there is So, this this this concept of general intelligence is completely BS. Um We can talk about human-level intelligence, right? So, are we going to have machines that are as good as humans in all the domains where humans are good or better than humans?
And the answer is, you know, we already have machines that are better than humans in some domains, like, you know, we have machines that can translate, you know, 1,500 languages into 1,500 other languages in any any direction. No humans can do this, right? Uh and and, you know, there's a lot of examples in, you know, chess and go and various other things. Um But, will we have machines that are as good as humans in all all domains?
The answer is absolutely yes. There's no question that at some point we'll have machines that are as good as humans in all domains. Okay. And and and that leads But, it's not going to be an event. It's going to be very progressive. We're going to make some conceptual uh advances, maybe based on, you know, GPT-3 models, planning, things like that. Uh over the next few years. And if we're lucky, if we don't hit an obstacle that we didn't see, uh perhaps this will lead to kind of good paths to human-level AI.
But, but perhaps we're we're still missing a lot of basic concepts. And so, the most optimistic view is that perhaps this you know, the you know, running good world models and uh and and, you know, being able to do planning and and, you know, understanding complex signals that are continuous high dimensional noisy. If we make significant progress in that direction over the next 2 years, the most uh optimistic view is that we'll have something that is close to a human intelligence or maybe dog intelligence within, you know, 5 to 10 years.
Okay. But, that's the most optimistic. Um it's very likely that as what happened, you know, multiple times in the history of AI in the past, there's some obstacle we're not seeing yet, which will, you know, actually kind of uh require us to invent some new conceptual uh new things to continue go beyond. In which case, that may take 20 years, maybe maybe more. Okay. But, no question it will happen. No question. >> And do you think it will be easier to get from the current level to a dog level intelligence compared to a dog to humans' levels?
>> No, I think I think uh the hardest part is to get to dog level. Once you get to dog level, you basically have most of the ingredients, right? And then, you know, what's missing from Okay, what's missing from like primates to humans uh beyond just size of brain is language, maybe, okay? But, language is basically handled by the Wernicke's area, which is a tiny little piece of brain that's right here, and the Broca's area, which is a tiny piece of brain right here.
Uh both of those evolved in the last, you know, less than a million years, maybe two. Um and it can't be that complicated. And we already have an LLM that do a pretty good job at at, you know, you know, encoding language into abstract representations and then decoding uh thoughts into into text. So, maybe we'll use an LLM for that. So, an LLM will be like the Wernicke's and Broca's areas in our brain. What we're working on right now is the prefrontal cortex, which is where our world model resides.
Well, well, this this gets me into, you know, a few questions about safety and the destabilizing potential impact. So, I'll I'll start this with something a little bit funny, which is to say if we really get dog level intelligence, then the AI of tomorrow has gotten profoundly better than any human at smell. And And something like that is just, you know, tip of the iceberg for the the destabilizing impacts of AI tomorrow, let alone today.
I mean, we have Sam Altman talking about super persuasion uh because AI docks is you. So, it it figures out who you are through the multi-turn, so it gets really good at kind of customizing its arguments towards you. We've had uh AI psychosis, right? Like people who've uh done horrible things as a result of uh kind of believing in a sycophantic AI that that is telling them to do things they shouldn't do. Uh happened to me, by the way.
Whoa, whoa. You You've got to tell us about that, too. What? Uh one day, a few months ago, um I was I was I didn't want you, and I walked down to get lunch. And there was a a dude who was surrounded by a whole bunch of police officers and and uh security guards. And I walked past, and and the guy recognizes me and says, "Oh, Mr. LeCun." Huh. And the police officer kind of whisks me away outside and tells me like, "You don't want to talk to him."
Uh turns out the guy had come from, you know, the the Midwest by bus over here. Wow. And he he's a kind of emotionally disturbed. He, you know, he had gone to prison blah blah blah uh for kind of various things. And he was carrying a bag with, you know, like a a huge wrench and and pepper spray and a knife. And so, the the security guards got alarmed and basically called the police. Well, and then the police realized, "Okay, we, you know, this guy is kind of weird."
So, they, you know, took him away and had him examined, and eventually he went back to the Midwest. Hm. Um But, I mean, he didn't feel threatening to me, but the police wasn't so sure. So, so yeah, it happens. Um I had, you know, high school students writing emails to me saying, "I read all of those, you know, piece by by doomers who said like, you know, AI is going to take over the world and either kill us all or take over take over jobs.
So, I I'm totally depressed. I'm not going to go to school anymore." And blah blah blah. So, you know, I answered to them saying like, you know, no. I don't believe this. All that stuff, you know, that humanity is still going to be in control of of of all of this. Now, there's there's no question that, you know, every powerful technology has, you know, good um consequences and bad side effects that sometimes are predicted and corrected sufficiently in advance, and sometimes not so much, right?
And it's it's always a trade-off. That's the history of technological progress, right? So, uh let's take cars as an example, okay? Cars crash sometimes. And initially, you know, brakes weren't that reliable, and and cars would like flip over, and there was no, you know, seat belts and blah blah blah, right? And eventually, kind of the industry made progress and, you know, started putting seat belts and and crumple zone and and uh and and, you know, automatic kind of uh uh controlling systems for so that the you know, the car doesn't go sway and doesn't flip or whatever.
Um so, cars now are much safer than they used to be. Now, there's one thing that is now mandatory in every car sold in the EU, and it's actually an AI system that looks out the window. Uh it's called um it's called AEBS, uh automatic emergency braking system. It's basically a camera shot there, all right? Uh And it looks out the looks out the windshield, and it detects, you know, all objects. And if it detects that an object is too close, it just automatically brakes.
And uh or if it detects that there's going to be a a collision that the driver is not going to be able to avoid, uh it just stops the car and or sways, right? And um that one statistic I I I read is that this reduces frontal collisions by 40%. And so, it became mandatory equipment in every car sold in the EU, even low low end. Cuz it saves lives. Okay. So, this is AI not killing people. Saving lives, Yeah. >> right?
I mean, also same thing for like medical imaging and everything. There's a lot of life being saved by AI at the moment. And like so, but do you think so, you, Jeff, and Joshua right? Like both of you won the the Turing uh award together. And like and you have a different opinions about it, right? Um and and and Jeff says like he's regret regrets, and Joshua works on the safety, and you trying to to push it forward. Do you think you will get to some uh some level of uh of intelligence that you will say, "Oh, this become too dangerous.
We need to work more on on the safety side." I mean, you have to do it right. Um I'm going to use another example. Uh uh jet engines, okay? Uh I find this astonishing that you can fly halfway around the world on a two-engine airplane in complete safety. And I really really say halfway around the world, like you said, 17-hour flight. Okay. Um You can fly direct to from New York to Singapore, right, on an Airbus 350. Um it's astonishing.
Uh and when you look at a jet engine, a turbofan, it should not work, right? I mean, there is no metal that can stand the type of temperature that takes place there. Uh and and the kind of like efforts when you have like a huge turbine, like, you know, rotating at 2,000 RPM or I don't know what speed. Like, the the the the force that puts on it is just insane. It's you know, it's hundreds of tons. So, it it should not be possible.
Yet, those things are incredibly reliable. So, what I'm saying is you can't, you know, build something like a turbojet the first time you build it. It's not going to be safe. It's going to run for 10 minutes and then blow up. Okay. Uh and it's not going to be uh fuel efficient and it's, you know, etc. It's not going to be reliable. Uh but, you know, as you make progress in engineering, in materials, etc., there's so much, you know, uh economic motivation to make these good that, you know, eventually it's going to be um the type of reliability we see today.
Um the same is going to be true for AI. We're going to start making systems that, you know, have agency, can plan, can reason, have world models, blah blah blah, but we you know, they're going to have the power of maybe the a cat brain, right? Which is about 100 times smaller than a human brain. Um and then we're going to put guardrails in them to prevent them from doing, you know, taking actions that are obviously uh dangerous or or something.
And you can do this at a very low level, like if you if you have, I don't know, a domestic robot, right? That uh um So so one example that uh Stuart Russell, for example, have have used is um is to say, "Well, you know, if you are a robot, a domestic robot, and you ask it to fetch you coffee, and someone is standing in front of the coffee machine. If the system wants to fulfill its uh its goal, it's going to have to, you know, either assassinate or smash the person in front of the coffee machine to get access to the coffee machine."
And obviously you don't want that to happen. Now, it's like the paperclip maximizer, it's kind of a ridiculous example because it's super easy to fix this, right? You put some guardrail that say, "Well, you know, you're a domestic robot, you should stay away from people and maybe ask them to move if they are in the way, but not actually kind of, you know, hurt them in any way or whatever." And you can do like, you know, you can put a whole bunch of low-level conditions like this, like uh if you're a domestic robot then it's, you know, it's it's a cooking robot, right?
So, it has a big knife in its hand and it's, you know, cutting a cucumber. You know, don't flail your arms if there are if you have a big knife in your hand and people around. Okay. It could be kind of a low-level constraint that the system has to satisfy. Now, some people say, "Oh, but you know, with LLMs, we can fine-tune them to not do things that are dangerous, but there is always you can you can jailbreak them.
You can always find prompts where they're going to kind of escape their condition, you know, the the all these things that we stopped them from from from doing. I agree. That's why I'm saying we shouldn't use LLMs. We should use those objective-driven AI architectures that I was talking about earlier, where you have a system that has a world model, can predict the consequences of its action, and can figure out the sequence of actions to accomplish a task, but also is subject to a bunch of constraints that guarantee that whatever action is being pursued and whatever state of the world is being predicted uh does not end end endanger anybody or does not have, you know, negative side effects, right?
So, there is it's by construction the system is intrinsically safe because it has all those guardrails and because it obtains its output by optimization, by minimizing the objective of the task and satisfying the constraints of the guardrails, it cannot escape that. It's not a fine-tuning, right? It's it's by construction. Yeah, um and and I'll there's a technique, you know, that that for LLMs for constraining the output space, where you say that you ban all outputs except whatever you want, like maybe zero to 10 and everything else.
And they have that even for diffusion models. Do you think that tactics like that as exist today significantly improve the utility of those kinds of models? >> Well, they do, but they're ridiculously expensive because the the way they work is that you have to have a system uh proposals for an output and then have a shelter that says, "Well, this one is good, this one's terrible, etc." or rank them and then just put out the the one uh rating, essentially.
So, it's it's insanely expensive, right? So, unless you have, you know, some sort of uh objective-driven value function that kind of drive the system towards producing those those high, you know, high score uh outputs low low toxicity outputs, um it's it's going to be it's going to be expensive. Yeah, and I want to change the topic just a a tiny bit off. We've been very technical for a moment, but we, you know, I think our audience in the world uh has a few questions that are maybe a little bit more more social related.
Um you know, the person who appears to be trying to fill your shoes in at Meta, Alex Wang, where I'm I'm curious as to do you have any thoughts or or anything about, you know, kind of how how that will play out for for Meta? He's he's not he's not in my shoes at all. He's uh he's uh he's in charge of all the uh R&D and product that are AI AI related at Meta. So, he's he's not a researcher or scientist or or or anything like that.
He's more kind of, you know, uh overseeing the entire operation. So, within uh Meta Superintelligence Lab, which is his organization, there are four um kind of divisions, if you want. So, one of them is FAIR, which is long-term research. And one of them is TBD Lab, which is basically uh building frontier models, which is mostly entirely LLM focused. Um The fourth organization is uh AI infrastructure, um software infrastructure.
Hardware is some other organization. Um and then the last one is products. Okay, so people who would take the frontier models and then turn them into actual chatbots that people can use and, you know, disseminate them and, you know, plug them into WhatsApp and everything else, right? So, um so those are four divisions. Uh he oversees all of that. Um So, and there are several AI scientists. There is AI scientist at FAIR, that's me.
Uh and I'm I really have a long-term view and basically, you know, I'm going to be at Meta for another, you know, 3 weeks. Okay, so. Uh Um and and FAIR is led by uh our NYU colleague, Rob Fergus, uh right now, uh after Joelle Pineau left some months ago. Um FAIR is being pushed towards kind of working on slightly, you know, shorter-term projects than it has done in the in the traditionally, with less less emphasis on publication, more focus on sort of helping TBD helping TBD Lab with LLMs and frontier models.
Uh and and, you know, less publication, which means, you know, Meta is becoming a little more close closed. Um And TBD Lab has a chief scientist also. Um but which is really focused on LLM. Uh and uh and the the other organizations are more like infrastructure products. So, you you know, there's there's some applied research there. So, for example, the group that works on SAM, Segment Anything, >> Yeah, yeah. that's actually part of the product division of uh Meta.
They used to be at at FAIR, but because they worked on kind of relatively, you know, kind of outside-facing uh kind of practical things, they were kind of moved to the product. And and do you have any opinions on uh like some of the other companies that are trying to move into world models, like Thinking Machines or even I've heard Jeff Bezos and and some of his It's not clear at all what Thinking Machines is doing.
Not at all? Maybe you have more information than me, but Or maybe not. Sorry, maybe I'm mixing it up here. No, Physical Intelligence. Physical, sorry. Feffe leaves, yeah, sorry. I and and then I I mix them up with like SSI as well. They're all kind of like So, SSI, nobody knows what they're doing, including their own investors. Okay. At least that's the rumors. That's the rumor I've heard right here. I don't know if it's true.
It's getting becoming a bit of a joke, but uh but uh yeah, Physical Intelligence uh Feffe's company is is focused on uh on, you know, basically producing geometrically correct uh videos, okay? Where, you know, there is uh persistent geometry and, you know, when you look at something and you turn around and you come back, it's the same object you had before. It doesn't change behind your back, right? Uh It so it's it it's generative, right?
I mean, the whole idea is to to to generate pixels, okay? Which I just spent, you know, a long time arguing against that it was a bad idea. Uh There are other companies that are have world models. One one is Wave. Wave? Uh W A Y Uh W A Y V E. So, um it's a company based in Oxford and they I'm an advisor, full disclosure. Uh and they have they have a world model for autonomous driving. Uh and the way they're training it is that they're training a representation space by basically training a VAE or VQ-VAE and then training a predictor to do temporal prediction in that abstract representation space.
So, they have half of it right and half of it wrong. The piece they have right is that you make predictions in representation space. The piece they have wrong is that they haven't figured out how to train their representation space in any other way than by reconstruction and I think that's bad. But their model is great. Like it works pretty well. I mean, among all the people who kind of work in this kind of stuff, they're they're pretty far advanced.
Um There are people who talk about similar things at Nvidia, a company called Sandbox AQ the the the CEO of it Jack Hillary talks about quantitative models you know, large quantitative models as opposed to large language models. So basically predictive models that can deal with continuous high-dimensional noisy data, right? Which is what also a bit kind of talking about. Um And Google of course has been working on you know, on world models mostly using generative approaches.
There was an interesting effort at Google by Nija Hafner. So he built models called Dreamer Dreamer V1 2 3 4. >> Yeah. Uh that was on a good path except he just left Google to create his own startup. Um uh do you have So I'm interested so you were really criticized about Silicon Valley culture that they are focusing on LLM and this is like one of the reason that now you started the new company is starting in in Paris, right?
Um so this is something do you think that we will see more and more or do you think this is something will be very unique that only a few companies will will be in Europe? Well the company I'm starting is global, okay? It has an office in Paris but it's a global company. It has an office in New York, too. Um a couple of other places. So um Okay, there is an interesting phenomenon industry which is that everybody has to do the same thing as everybody else because it's so competitive that if you start taking attention you're taking a big risk of falling behind because you're using a different technology than everybody else, right?
So basically everyone is trying to catch up with the others. And so that creates this herd effect uh and it kind of monoculture which is really specific to Silicon Valley where you know, uh OpenAI Meta, Google, Anthropic, everybody is basically working on the same thing. And all you know, sometimes like what happened a while back uh another group, you know, like DeepSeek in China comes up with kind of a new way of doing things and everybody is like, "What?"
Right? You mean like other people in Silicon Valley are not stupid and can come up with original ideas? Um I mean there's a bit of a you know, superiority complex, right? Mhm. Um but you're basically in your trench and you are you have to move as fast as possible because you can't afford to kind of you know, fall behind the the other guys who you think are your competitors. Uh but you run a risk of being surprised by something that's completely out of the left field that uses a different set of technologies and um or maybe addresses a different problem.
Um so um you know, what I've been interested in is completely orthogonal because the the the whole J idea of world model is really to handle data that is not easily handled by LLM. So the the type of applications we're envisioning they have tons of applications in industry where the data comes to you in the form of continuous high-dimensional noisy data including video. Our domains where LLMs basically are not present where where people have tried to use them and totally failed essentially, right?
Okay, so if you don't want to be Okay, so the the expression in in in Silicon Valley is that you are LLM pilled. You think that the path to super intelligence you just scale up LLMs, you train on mostly synthetic data, you license some more data, you hire thousands of people to kind of find you know, to basically school your system in post training, you invent a new tweaks on RL and you're going to get to super intelligence.
And this I think is complete [ __ ] Like it's just never going to work. Um and then you add a few you know, kind of reasoning techniques which basically consist in you know, doing like super long chain of thoughts and then having the system generate lots and lots of different token outputs, you know, from which you can select good ones using some sort of evaluation function the second LLM basically you know, that I mean that's the way all of these things work.
Um this is not going to take us there. It's just not. So um so yeah, I mean you need to escape that culture and there are people within all the companies in Silicon Valley who think like this is never going to work. I want to I want to do LLM and J blah blah blah. I'm hiring them. So uh yeah, so escaping the monoculture of Silicon Valley I think is important. This yeah, this this uh this part of the the story. And what do you think about the competition between like the US, China and the and Europe?
Like now that you are starting a company like do you do you see more uh um I know that some there there are some places that are more attractive than others? Well in the we're in this very paradoxical situation where all the American companies up to up to now not Meta but all American companies have been kind of becoming really secretive uh and to preserve their competitive what they think is a competitive advantage.
Um and by contrast the Chinese players companies and others have been completely open. So the best open source systems at the moment are Chinese. And that causes a lot of the industry to use them um because they want to use open source system. Uh and they hold their nose a little bit because they know those models are kind of fine-tuned to not answer questions about politics and stuff like that, right? Um but they don't they don't really have a choice and and certainly a lot of academic research now you know, uses the the best Chinese models.
Certainly everything that has to do with like reasoning and things like that, right? So so it's really paradoxical and a lot of people in in the US in industry are really unhappy about this. They they really want a serious non-Chinese open source model that could have been a platform but a platform was a disappointment for various reasons. Um maybe that will get fixed with you know, the the new efforts at Meta or maybe Meta will decide to to go close as well.
It's not clear. It's not clear. Mistral just had a model released. Yes, which is cool for Kujin. Yeah. >> Yeah. Yeah. That's right. No, it's it's cool. So yeah, they they maintain openness. Yeah. Uh uh no, it's really really interesting what they're what they're doing. Yeah. Wow. Okay. Um let's go to more personal questions. Yeah. >> Um yeah, so like you are uh 65, right? Yeah. Uh you won a Turing Award, you just got a Queen Elizabeth Prize.
Basically you could retire, right? Yeah, I could. That's what my wife wants me to do. So why why to start a new company now? Like what keep you up? Because I have a mission, you know. Um I mean I always thought that uh either making people smarter or more knowledgeable or making them smarter with the help of machines. So basically increasing the amount of intelligence in the world was an intrinsically good thing. Okay, intelligence is really kind of the commodity that is the most in demand.
Uh certainly in like government Okay. Uh So but but like in you know, every aspect of of life we are limited as you know, as as a species as a planet by the limited supply of of intelligence, right? Which is why we we we spend and things like that. Um so you know, increasing the amount of intelligence at the service of humanity or the planet more globally not just humans is intrinsically a good thing despite all the what the doomers are saying, okay?
Of course you are dangerous and you have to protect against them. The same way you have to make sure your jet engine is safe and reliable and your car you know, doesn't kill you with a you know, small crash, right? Um but that's okay. That's an engineering problem. We It's not It's not like fundamental issue with that. Um also political problem but not it's not like insurmountable. Um so that's an intrinsically good thing and if I can contribute to this I will.
And basically all research projects I've done in my entire career even those that were not related to machine learning uh in my professional activities were all focused on either making people smarter. That's what that's why I'm a professor. Uh And that's why also I'm communicating publicly a lot about AI and science and things like that and I have big presence on social networks and stuff like that, right? Because I think people should know stuff, right?
Um but also on machine intelligence because I think uh machines will assist humans and make them smarter. Okay. People think there is a fundamental difference between trying to make uh you know, machines that are intelligent and autonomous and blah blah blah and and it's a different set of technologies so I'm trying to make machines that are assistive to humans. It's not It's the same technology. It's exactly the same.
Uh and it's not because a system is intelligent or even a human is intelligent that it wants to dominate or take over. Uh, it's not even true of humans. Like, it's not the humans who are the smartest that want to dominate others. We see this on the, you know, international political scene every day. Uh, it's not the smartest among among us who want to be the chief. Uh, and probably any of the smartest people that we've ever met are people who basically want nothing to do with the rest of humanity.
They just want to work on their problems, you know. Like Yeah. I mean, I'm I'm, you know, kind of stereotype you here. That's what Hannah Arendt talks about the vita contemplativa, right? Versus like the active life or the contemplative life, right? In in her like philosophical analysis and like making a choice kind of early on on what you work on. Right. But you can be, you know, somebody who's totally kind of you know, a dreamer or contemplative but have a big impact on the world.
Yeah. >> Right? By, you know, your scientific production. Like, think of Einstein or something. Yeah. Or even Newton. Like, Newton basically didn't want to meet anybody. Right. Famously. Or Paul Dirac. Paul Dirac was kind of uh you know, practically autistic. Well, well, um is there like a paper or idea you haven't written or or something else that you you're, you know, nagging that you want to get to or maybe that you don't have time or any regret?
Oh, yeah. A lot. Oh, my my entire career has been a uh, succession of me not devoting enough time to express my ideas and writing them down. Mhm. Uh, and mostly getting scooped. What is the the the most significant one? Uh, I don't want to go through that. The backprop is a good one. Okay. Okay. I I published some sort of early version of some algorithm to train multi-layer nets which today we'll be called target prop.
Uh, and I had the uh, backprop thing figured out um, except I didn't write it before. Um, you know, they were more hot than Jeff Hinton. They were nice enough to cite my earlier paper in their in theirs, but uh So, there's been a few of those, yeah. Recurrent nets, you know, very similar things. But, um uh, and things are a mark uh, past more recent, but but then you know, I have no regrets about this. Like, you know, this is life.
Like, you know, I'm not going to say, "Oh, you know, I invented this in 1991 and I should Like somewhere. I I don't know. I shouldn't say the name, but you know, we all know. If you know, you know. The the way ideas pop up, you know, is is rather unique and complex. It's rare that someone comes up with an idea in complete isolation and that you know, nobody else comes up with similar ideas at the same time. Most of the time they appear simultaneously.
But then there is various ways to there's having the idea and then there is kind of writing it down, but there's also writing it down in a sort of convincing way and a clear way. And then there is kind of making it work on toy problems, maybe. Okay? And then there is making the theory that shows that it can work. Uh, and then there is making it work on a real application, right? And then there is making a product out of it.
Okay? So, there's this whole chain and, you know, some people in the extreme think that the only person who should get all the credit is the very first person who got the idea. Mhm. I think that's wrong. There's there's a lot of really difficult steps to get this idea to a state where it would actually works. So, this idea of a world model, I mean, it goes back to the 1960s. You know, people in optimal control had world models to do planning.
That's the way NASA you know, planned the trajectory of the rockets to go to to orbit. Um, basically simulating the rocket and sort of prioritization figuring out the the control law to to get the rocket where it needs to be. Um, so that's an old idea, very old idea. The the fact that you could do some level of training or adaptation in this is called system identification in optimal control, very old idea, too. Goes back to the '70s.
Something something called uh Yeah, system identification or even MPC where you adapt the model as it goes, like while you're you're running the system. That goes back to the '70s. To some of us here in France. Um, and and and then the fact that you can just learn a model from data. Um, people have been working on this with neural nets since the 1980s. Right? And and not just Yann LeCun. There's like a whole bunch of people who've been uh, working people who came from optimal control and realized they could use neural nets as kind of a universal function approximator and use it for direct control or feedback control or world models for planning, blah blah blah.
Um, and like a lot of things in neural nets in the 1980s and '90s um, it kind of worked, but not like to the point where it took over the the the industry. Um, so it's the same for, you know, computer vision, speech recognition, there were attempts at using neural nets for that back in those days. Uh, but it started really really working well in the late 2000s where it totally took over, right? And then early 2010s for vision uh, mid-2010s for NLP and for robotics it's starting, okay?
But it's not Why why I think it's only like in the this time started to get over the Well, it's a combination of like having the right state of mind about it and the right mindset, uh, having the right uh, architectures, the right machine learning techniques, like, you know, residual connections, real use, whatever. Uh, then having powerful enough computers and having access to data. And it's only when those planets are aligned that you get a breakthrough, right?
Uh, which appears like a conceptual breakthrough, but it's actually just a practical one. Uh, like Okay, let's talk about convolutional nets, okay? Um, lots of people during the the '70s had the idea or even during the '60s, actually, had the idea of using local connections, like building a neural net with local connections for extracting local features. And the idea that local features is like convolution, like in in image processing is like it goes back to the '60s.
So, these are not new concepts. The fact that you can uh, learn adaptive filters of this type using data uh, goes back to the perceptron and adaline, which is early '60s, okay? But that's only for one layer. Now, the concept that you can train a system with multiple layers, everybody was looking for this in the '60s. Nobody found a lot of people made proposals which kind of half worked. Uh, but like none of them was convincing enough for people to say, "Ah, okay, this is a good technique."
One technique that was adapted is what's called polynomial classifiers. So, now we would turn this into kernel methods, but it's, you know, basically you sort of have a hand-crafted feature extractor and then you train a basically what what amounts to a linear classifier on top of it. Um, that was kind of common practice in the '70s and and certainly '80s. But the idea that you could train uh, a non-linear system I mean, a system composed of multiple non-linear steps using gradient descent.
Uh, the basic concept for this goes back to the Kelly-Bryson algorithm, which is optimal control, was mostly linear in, you know, from 1962 uh, and people in optimal control kind of wrote things about this, you know, in the '60s. But nobody realized you could use this for machine learning to do pattern recognition or to do, you know, natural language processing. That really only happened after uh, you know, the Rumelhart-Hinton-Williams paper in 1985.
Even though people had proposed the very same algorithm uh, a few years before. Uh, like Paul Werbos, you know, proposed, you know, what he called ordered derivatives, which turns out to be backprop, but it's the same thing as the adjoint state method in optimal control. So, like those ideas I mean, the fact that an idea or technique is reinvented multiple times in different fields and then only after the fact we knew about this before, we didn't realize we could use this for it for for this particular stuff, right?
Um, so all those like claims of plagiarism, let's just put it it's just a complete misunderstanding of how ideas come about. Okay, um, what do you do when you're not thinking about AI? Um, Um, I have a whole bunch of hobbies that have uh very little time to actually partake in. Um, I like sailing, so I go I go sailing in the summer. Uh, I like sailing multi-hull boats, like trimarans and catamarans. Um, I uh I have a bunch of boats.
Um, I like uh building flying contraptions. So, A modern DaVinci. I wouldn't I wouldn't call them airplanes cuz a lot of them don't look like airplanes at all. But they do fly. Okay. I like the the sort of, you know, concrete creative uh, act of that. My dad was uh he was face engineer and he mechanical engineer working in the aerospace industry and he was, you know, building airplanes as as as a hobby and like, you know, building his own radio control system and stuff like that.
He got me and my brother into it. My brother who works at Google at Google research. In France. In in Paris. And uh and and that became kind of a family activity, if you want. So, my brother and I still still do this, but um, and then uh in in the COVID years, I picked up astrophotography. So, I have virtual telescopes and take pictures of the sky. Uh and I build electronics. So, I uh since I was a teenager, I was interested in uh music.
I was playing Renaissance and Baroque music. And also some type of folk music. Um you see, uh playing wind instruments with winds. And but I was also into electronic music. And uh my cousin who was slightly older than me was inspiring. Electronic musician. So, he had like analog synthesizers. And because I knew electronics, I would like, you know, modify them for him. But I was still in high school at the time. Um and uh and now in um in my home, I have, you know, a whole bunch of synthesizers and I build electronic musical instruments.
So, so these are wind instruments. You blow into them, you know, there's fingering and stuff, but what they produce is signals for a synthesizer. Oh, that's cool. Very cool. Heard a lot of people in tech are into sailing. Like in Uh-huh. Yeah, gotten that answer surprising amount. I'm going to start trying to sail now. Okay, so I tell you something about sailing. It's very much like the world model story. Uh to be able to uh you know, going to control the sailboat properly to make it go as fast as possible and everything.
You have to anticipate a lot of things. You have to anticipate the motion of the waves, like how the waves are going to affect your boat. Um you know, whether a gust of wind is going to come in and have to, you know, start, you know, the boat is going to start healing and things like that. And you basically have to run CFD in your head. Because you have to figure out like the uh you know, fluid dynamics. You have to figure out like what is the flow of air around the around the around the sails.
And you know that if the angle of attack is too high, it's going to be turbulent on the back and the the lift is going to be much lower. So, uh blah blah blah. So, like, you know, uh uh tuning sails is basically requires running CFD in your head. But at an abstract level, you're not solving the, you know, Navier-Stokes, right? You have to really good intuitive. So, what That's what I like about it. Like the the whole thing that you have to build this mental, you know, predictive model of the world to be able to do a good job.
The question is how many samples you need. Uh yeah, probably a lot. But but, you know, you get your running time in uh you know, a few a few years of practice. Yeah. Um okay, um you're French and you lived in the US for many decades already. Do you still feel uh French? Like does that perspective shape your view of the of the world of the American tech culture? Well, inevitably, yeah. I mean, you you can't completely escape your your upbringing and your your culture.
So, I mean, I feel both French and American uh in the sense that, you know, I've been in the US for 37 years. Um and uh in North America for 38 cuz I was in Canada before. Uh or or children grew up in in the US. Uh And so, from that point of view, I'm American. But I have a view certainly on, you know, on various aspects of science and society that probably are you know, a consequence of growing up in France. Yeah, absolutely.
And I feel French when I'm in France. I'm curious. I did not actually realize that you had a brother that also worked in tech. I'm fascinated by this because um Yoshua Bengio's brother also works in tech. And I always thought that he was the only Serena Venus Williams situation in AI. But you you two also have a brother. So, how many more AI research Like like is it that common that it just runs in families? >> no idea.
I also have a sister who, you know, is not uh in tech, but she's also a professor. Uh my brother was a professor before he moved to Google. Wow. Oh. Yeah. Uh he he doesn't work on AI machine learning. He's very careful not to. He He is my He's a younger brother. He's 6 years younger than me. Um and he works on operations research and optimization essentially. Which Now is actually also being invaded by machine learning.
So, there's no Yeah. Um okay, one more question. So, like if the world models work in 20 years from now, what is the what is the dream? Like how what how does it look like? How does I know, like our lives will be? Uh total world domination. Okay, no, it's a joke. IT'S A JOKE. THE THE UH I I I SAID THIS sentence because this is what uh Linus Torvalds used to say. He said like, "What's your goal with with Linux?" And he said, "Total I thought that was super funny.
And he actually succeeded. I mean, basically, you know, to first approximation, every computer in the world runs Linux. Um there's only a few desktops that don't and a few iPhones, but, you know, everything else runs Linux. So, uh No, really like, you know, having, you know, pushing towards like a a recipe for training and building intelligent systems, perhaps all the way to human intelligence or more. Uh and and and and basically building AI systems that would, you know, help people and humanity more generally in their daily lives uh at all times.
Amplifying human intelligence. We'll be their boss, right? It's not like those things are going to dominate us. Because again, it's because something is intelligent that it wants to dominate. Those are two different things. Um in humanity, uh you know, we are hardwired to having to influence other people. And sometimes it's through domination. Sometimes it's through prestige. Um but um but we're hardwired by evolution to do this because we are social species.
There's no There's no reason we we would build those kind of drive into into our intelligent systems. And it's not like they're going to develop those kinds of drives by themselves. So, um so yeah, I'm quite optimistic. Me, too. So am I. All right. >> Um okay, so we have final question from the audience. Um so, yeah, let's start. Um if you were starting your AI career today, what skills and research directions would you focus on?
I get a lot this question a lot from young students or parents of future students. So, what what, you know, area should I I mean, I think you should learn things that have a long shelf life. And you should learn things that help you learn to learn. Because technology is is, you know, evolving so quickly that you you want kind of, you know, the ability to learn really quickly. Um and basically that can that is done by, you know, learning very basic.
So, in the context of STEM, right? Science and technology, engineering, mathematics. Uh and I'm not talking about humanities here. Um this is, although you should learn philosophy. Um this is done by learning things that have a long shelf life. So, the joke I say is that if you First of all, you the things that have a long shelf life tend to not be computer science. Okay, so here is a computer science professor. You know, are you against studying computer science?
>> Don't come. Don't come to to study. Uh And I have a terrible coefficient to make, which is I I studied electrical engineering as a as an undergrad. So, I'm not a real computer scientist. Okay. Um but what what um we should do is learn kind of basic things in in mathematics, in modeling, uh mathematics that can be connected with reality. You You tend to learn this kind of stuff in engineering. Uh in some schools that's linked with computer science, but sort of you know, electrical engineering, mechanical engineering, etc.
Engineering disciplines. You know, you learn in the US calculus 1 2 3. That gives you a good basis, right? Uh computer science, you don't you know, you you can get away with just calculus 1. That's not enough, right? Uh you know, learning, you know, probability theory and uh linear algebra, you know, all this stuff that are really kind of basic. And then, if if you do uh electrical engineering, things like, I don't know, control theory or signal processing, like all of those methods, optimization, you know, all those methods are really useful for things like AI.
Um and then, uh you can you can basically learn similar things in physics. Um because physics is all about like, what should I represent about reality to be able to make predictive models, right? And that's really what intelligence is about. So, um so, I think you can learn most of what you need to learn also if you go to a physics uh uh uh curriculum. Uh but obviously, you need to learn enough computer science to kind of program and use computers.
And even though AI is going to help you be more efficient at programming, you still need to know like how how to do this. What do you think about live coding? I mean, it's it's cool. Uh it it makes it's going to cause a funny kind of thing where a lot of the code that will be written will be used only once. Right? Because it's going to become so cheap to write code. Right? You're going to ask your kind of AI assistant like, you know, produce this graph or or like, you know, do this research blah blah and it's going to write a little piece of code to do this.
Or maybe it's an applet that you need to play with for, you know, a little simulator and you're going to use it use it once and throw it away because so cheap to produce, right? Um so the idea that we're not going to need programmers anymore is false. We're we're going to I mean, the the the cost of generating software you know, has been going down continuously for decades and it's that's just the next step of of the cost going down, but it doesn't mean computers are going to be less useful.
They're going to be more useful. Um okay, one more question. So what do you think about the connection between neuroscience and the and machine learning? That there are some ideas a lot of time and that AI borrows from neuroscience and the other way, right? I don't know, predictive coding for example. Um do you think it's useful to use ideas from Well, there's a lot of inspiration you can get from uh from neuroscience, from biology in general, but neuroscience in particular.
I certainly was very influenced by, you know, classic work in neuroscience like, you know, Hubel and Wiesel's work on the architecture of the visual cortex is basically what led to convolutional nets, right? And and you know, I wasn't the first one to use those ideas in uh in in artificial neural nets, right? There were people in the '60s trying to do this. There were, you know, people in the '80s building locally connected networks with multiple layers.
They didn't have ways to train them with uh backprop. You know, there was the the the cognitron and neocognitron from Fukushima uh which had a lot of the ingredients just not proper learning algorithm. And there was another kind of uh aspect of the cognitron which is that it was really meant to be a model of the visual cortex. So it tried to reproduce every quirks of uh of of of biology. For example, the fact that in the brain you don't have uh positive and negative weights.
You have positive and negative neurons. So all the neurons all the synapses coming out of an inhibitory neurons neuron have negative weights. Yeah. Okay? And all the synapses coming out of a uh non-inhibitory neurons have positive weights. So uh Fukushima implemented this in his model, right? He implemented the fact that uh you know, neurons uh spike. So he didn't have a spiking neuron uh model, but but you cannot have a negative number of spikes.
And so so his function was basically rectification like a ReLU except it had a saturation. Yeah. Um and then he knew from, you know, various works that there was some sort of normalization and he had to use this because uh otherwise the there was no backprop so the activation in his network would go haywire so we had to do like divisive normalization. That turns out to actually be um uh correspond to some theoretical models of the visual cortex uh that some of our colleagues at at the Center for Neuroscience at NYU have been pushing.
Like David Heeger and people like that. So um yeah, I mean, I think neuroscience is very a source of inspiration. Um you know, more more recently the sort of macro architecture of the brain in terms of, you know, perhaps the world model and and planning and things like this. Like how how is that reproduced? Like why do we have a separate um module in the brain for for for factual memory uh the hippocampus, right? And we see this now in certain neural net architectures that there is like a separate memory module, right?
Maybe it's a good idea. I think we're going to have I think what's going to happen is that we're going to come up with sort of new AI neural net architectures, deep learning architectures and a posteriori we'll discover that the characteristic that we implemented in them actually exists in the brain. And in fact, that's a lot of what's happening now in in our science which is that there is a lot of feedback now from AI to neuroscience where the best models of human perception are basically convolutional nets today.
Yeah. Um okay, do you have anything else that you want to add to say to the audience? Uh whatever you want. >> we covered a lot of grounds. Uh I think you you know, you want to be careful about who you listen to. Uh so don't listen to AI scientists talking about economics. Okay? So when some AI person or even a business person tells you AI is going to put everybody out of work talk to an economist. Basically none of them is saying anything anywhere close to this.
Okay? Uh and uh you know, the effect of technological revolutions on the labor market is something that a few people have devoted their career on on on doing. None of them is predicting massive unemployment. None of them is predicting that radiologists uh are going to be all unemployed uh you know, etc., right? Also realize that actually fielding practical applications of AI uh so that they are sufficiently reliable and everything is super difficult and is very expensive.
And in previous waves of interest in AI the techniques that people had put a big hope in uh turned out to be overly unwieldy and expensive except for a few applications. So there was a big wave of interest in uh expert systems back in the 1980s. Uh Japan started a huge project called the fifth generation computer project uh which was like computers with CPUs that were going to run Lisp and, you know, inference engines and stuff, right?
And the hottest uh job in the late '80s was going to be knowledge engineer. You were going to sit next to an expert and then turn the knowledge of the expert into rules and facts, right? And then the computer would be able to basically do what the expert wants. This was manual behavior cloning. Okay? Uh And it kind of worked, but only for a few domains where where economically it made sense and it would and it was doable at the level of reliability that was good enough.
Um but it was not a path towards kind of human level uh intelligence. And so the idea somehow like the delusion that people today have that the the current AI mainstream you know, fashion is going to take us to human intelligence has happened already three times during my career and probably five or six times before. Right? You you should you should see what people were saying about the perceptron. I wrote a New York Times article.
People were saying, "Oh, we're going to have like super intelligent machines within 10 years." Marvin Minsky in the '60s says, "Oh, within 10 years the best chess player in the world will be a computer." It took a bit longer than that. Um and, you know, and you know, this you know, this happened over and over again. Uh in in 1956 or something when Newell and Simon uh produced the the general problem solver very modestly called general problem solver.
Okay? What they what they thought um was really cool. They said, "Okay, the way we think is very simple. We pose a problem. There is a number of different solutions to that problem, different proposals for solution, a a space of potential solutions. Like, you know, you do like traveling salesman, right? There is a number of you know, factorial you know, n factorial uh paths, possible paths. You just have to look for the one that is the best, right?
Um and they said like every problem can be formulated this way essentially for a search for the best solution. If you can formulate the problem as an objective by writing a program that checks whether it's a good solution or not or gives a rating to it. And then you have a search algorithm that search through the space of possible solution for one that, you know, optimizes that score then you solve AI. Okay? Now what they didn't know at the time is all of complexity theory that basically every problem that is interesting is exponential or NP-complete or whatever, right?
And so oh, we have to use heuristic programming. You know, kind of come up with heuristics for every new problem. And basically, you know, the our general problem solver was not that general. So like this idea somehow that the latest idea is going to take you to you know, AGI or whatever you want to call it is very dangerous and a lot of very smart people fell into that trap many times over the last seven decades. Do you think that um the field will ever figure out continual or incremental learning?
>> Sure. Yeah, that's that's sort of a technical problem. Well, well, I thought I thought catastrophic forgetting, right? Because your weights that you trained so much money on get overwritten. Sure. But so you train just a little bit of it. I mean, we don't already do this with SSL, right? We train a foundation model. Uh like for video or something like Video BERT, you know, produces really good representations of video.
And then if you want to train the system for a particular task, you train a small head on top of it and that head can be, you know, run continuously. And even your world model can be trained continuously. That's not not an issue. I don't see this as like a big a huge challenge, frankly. Uh in fact, uh Raya Hadsel and Samy Bengio and I and a few of our colleagues back in 2005, 2006 built a uh learning-based navigation system for mobile robots that had this kind of idea.
So, it was it was convolutional net that was doing semantic segmentation from camera images. And on the fly the top layers of that network would be adapted to the current environment. Um, so it would do a good job. And the labels came from short range uh, traversability that were indicated by stereo vision, essentially. So, yeah, I mean you can do this. It's particularly if you have a multi-modal. Um, yeah, I don't see this as a big challenge.
It's been a pleasure to have you. All right, it was really a pleasure to Thank you so much. Thank you. Thank you.