AI 到底是什么?Yann LeCun 对话 Nikhil Kamath(People by WTF #4)
WTF is Artificial Intelligence Really? | Yann LeCun x Nikhil Kamath | People by WTF Ep #4

Heat. Heat. I thought we could use today to figure out a what is AI? How did we get here? What likely next? As an Indian 20-year-old who wants to build a business in AI, a career in AI, what do we do today? Today, like right now. Yeah. Hi, Yan. Good morning. Thank you. Thank you for doing this. Pleasure. The very first thing we like to do is get to know you a bit more. Uh how you came to be what you are today. Uh could you tell us a little bit about where you were born, where you grew up leading up to today?
So, I grew up near Paris in the suburbs. Um, my dad was an engineer and I learned almost everything from him and um [clears throat] um always was interested in in science and technology since I was a little kid and [clears throat] and always saw myself as uh perhaps becoming an engineer. I had no idea how you became a scientist, but I became interested in this afterwards. What's the difference between an engineer and a scientist?
Well, um it's very difficult to to define and uh very often you have to be a little bit of both. [clears throat] But uh um scientists you try to understand the world. Um engineer you try to create new things and very often if you want to understand the world you need to create new things. The progress of science very much is linked with progress in technology that allows to collect data. you know the invention of the telescope allowed the discovery of planets and the planets are um rotating around the sun and things like this right the microscope open the door to all kinds of things so um so technology enables science and for the problem that really has been my obsession for a long time is uh discovering the mysteries of uncovering the mysteries of intelligence um and as as an engineer I think the the The only way to do this is to build a machine that is intelligent.
Right? So there's both an aspect of a scientific aspect of understanding intelligence what it is um at a theoretical level and more practical side of things and then u of course the consequences of building intelligent machines could be could have you know could be really important for humanity and school in Paris studying what? So I I studied electrical engineering. Mhm. Um but as I progressed in my studies, I became more and more interested in sort of more fundamental questions in mathematics, physics and and AI.
Um I I did not study computer science, right? Uh of course there is always computers involved when you study electrical engineering even in the 1980s and late 70s actually when I started. Um but um but I got to do a few independent projects with mathematics professors on on the questions of AI and and and things like that. And I really got hooked into research. I I was uh um you know my my uh my my favorite activity is to to build new things, invent new things and then understand new things in a new way.
When somebody says godfather of AI, the term, how does it make you feel? What do you think about it? Uh, I mean, I don't particularly like this term. You know, I I live in New Jersey. Godfather in New Jersey means you are you belong to the mafia, right? I mean, it science is never a sort of individual uh pursuit. You you you make progress by by the collusion of ideas from multiple people and you you you make you do make hypothesis and then you try to show that your hypothesis is is correct by demonstrating that the idea you have the mental model of what should work um is correct by demonstrating what that uh that it works um or doing some theory and things like that.
um and um it it's not an isolated uh activity. So there's always a lot of people who have contributed to uh to progress but then because of the nature of how the world works, we only remember just a few people. Um I think a lot of the credit should go to a lot more people. It's just that we don't have a good you know memory for attributing credit to to a lot of people. So how does it feel to be a teacher today? An when you were at NYU, are you the celebrity at NYU?
Um, let's say over the last uh several years, uh, students come up to me at the end of the class, I want to take selfies. Yeah. So, so there's a little bit of that. I I think um if you are in the same room with someone, I think it's important to sort of uh make the session interactive because otherwise you can just watch a video. Um so so that's why that's what I try to do really sort of engage with the students. Do you suspect being a hero in academia in research is much like being a hero in sport or entrepreneurship or do you think it's harder?
Okay. There's something I'm I'm happy about the fact that there can be heroes, you know, from in science. Um, argue there was Newton and Einstein and all these people, right? Well, Newton was not really kind of a public figure I think. Um, I mean, he was a Cambridge, but Einstein certainly was. Yeah. Um, and to some extent, you know, other some other scientists also were were sort of minor celebrities. Um so I mean I think uh some of that comes from you know scientific production but frankly there's a lot of people who have made scientific contributions that are completely unknown um and which I find it a little sad but um I I think a lot of people who have become prominent in in in science and technology is not just because of the science they've they've produced but also because of their public stands and um out there.
You know, one one thing that perhaps differentiates me from other scientists who are a little quieter is that I'm very present on social networks and I give public talks and I have strong opinions about not just technical issues but also uh policy issues to some extent. So I that that I think amplifies a little bit the the popularity or unpopularity in certain circles. I'm seen like a complete idiot. I've watched a lot of your interviews over the last Fortnite, last month in fact.
If you were to state three problems with the world from Yan's lens, what would they be? Um, so as a scientist, you try to establish causal models of the world, right? So there are effects that we're seeing. And then the question is what is it caused by? And almost for almost every problem that we have the the uh the cause is really a lack of knowledge or or intelligence by humans. We're making mistakes. We're making mistakes because we're not smart enough to figure out we have a problem because we're not smart enough to figure out solutions.
We're not smart enough to organize ourselves to find solutions. Right? So things like I mean climate change is a huge issue right? Um and you know they might there are you know political issues with that and questions of organizing uh the world the governments etc but also uh potentially technological solutions to to climate change and and I wish we were you know smarter so that we could find solutions faster. So, are you saying humans don't know why we do what we do and that's the problem?
No, I think the mistakes we're making is because um uh if we were if we were a little smarter, if we had a bet better mental model of of how the world works. Uh and that's a central question in AI as well. I [snorts] think we could we could solve our our problems better. We would take decisions that are more rational. Um and what I the the the the big issue I see in the world today is is is people who um are not interested in finding the facts are not interested in educating themselves.
Um or maybe they are but they don't have the means to do this. They don't have access to information and knowledge. So I think the best thing we can do and maybe that's why you know I became a professor is to uh make people smarter and to some extent that's the best reason also to work on AI because AI is going to amplify human intelligence I mean the the overall intelligence of humanity if you want. So I I think that uh that's the key to solving a lot of the problems that we have.
So just to preface this conversation uh I'm an idiot when it comes to anything around AI or technology. There isn't much that I know and I've tried to learn uh over the past over the very recent past and uh I have a lot of curiosity for it but I don't know enough about it. A lot of the people watching us today are wannabe entrepreneurs primarily based out of India. A lot of us have heard conjecture around AI. We have heard about the edge cases both on the positive side and the negative side.
I thought we could use today to figure out for all of us a what is AI? How did we get here and what likely next? If I were to break today down into three parts should we start with what is AI? Okay. Um that's a good question. What is intelligence even? Um, so in the history of AI, I think the problem of what is AI feels a little bit like the the story of the blind man with the elephant, right? That there are very different aspects to intelligence and over the history of AI, people have addressed one view of what intelligence is and and and basically ignored all the other all the other aspects.
So um one of the early aspects of uh of intelligence that people addressed with AI in the 1950s was you know intelligence is about reasoning um how do we reason so how do we reason logically um how do we search for solutions to a new problem and in the 50s people figured out um when we have a problem let's say that's become a standard problem in uh in AI it's or computer science Now, um, you know, I give you a bunch of cities and I ask you, you have to go through every single city and what's the shortest path, the the shortest circuit to go around the city.
That's called a traveling salesman problem. Um, and they say like every reasoning problem can be formulated in terms of searching for a solution to a problem. There's a space of possible solution. There is something that tells you whether you found a good solution or not or or some number that tells you the length of the path and you just have to search for the shortest path, right? and and to some extent you could reduce every reasoning problem to a problem of of this type.
In mathematics we call this optimization. Okay. So you you have um a problem you can you can evaluate whether your problem is solved or not with a number that indicates you know it's low if your length of your path is small and it's high if it's longer and you search for a solution that minimizes that uh that the so it's finding solutions related to intelligence. If you were to ask me what is intelligence, I would be like dumbfounded in trying to define it in a sentence, right?
I mean, so that comes back to the elephant analogy. Can you explain the elephant analogy? Well, so you know the the blind the the blind man and the elephant, right? So you the first blind man you goes to the side of the elephant say that sounds like a looks like a wall and then one goes to a leg that looks like a tree. um and another one you know touches a trunk that's a that's a pipe and nobody has a complete picture of what an elephant is right and you you see it from from the various angles so this aspect of intelligence as being a search for a solution to a particular problem is you know a small piece of the elephant it's it's one aspect of intelligence but it's not it's not the entire thing but in the 50s um one branch of of AI was basically only concerned by this Um and and that branch was essentially dominant until until the 1990s.
Um that that AI uh consists in searching for a solution for plans. You know, if you want to, you know, stack a bunch of objects on top of each other and some objects are bigger than others. You have to sort of organize the order in which you're going to stack the objects. You know, you search for a sequence of actions to arrive at a goal. That's called planning. Mhm. Or even um let's say u you have a robot arm and you have to grab an object but there is there's you know obstacles in front of it.
You have to plan a trajectory for the arm to grab to grab the object. Um so all of that is planning that's part of this searching for a solution to a problem. Um but that part of AI which again was started in the 50s and was dominant until the 90s uh completely ignored things like perception like how do we understand the world? Mhm. Um how do we recognize an object? Um how do we separate an object from its background so we can identify it?
Um um and um you know how how do we think um not in terms of logic or or or search but perhaps in more abstract uh abstract terms. And so that was essentially ignored. But there was another branch of AI also started in the 50s um that said well let's try to reproduce the mechanisms of intelligence that that we see in animals and humans. And animals and humans have brains. The brains basically organize themselves. They they learn, right?
They they're not spontaneously smart. And the intelligence is sort of a emerging emergent phenomenon of networks of very simple elements in large numbers that are connected with each other. Um so in the 50s or 40s people started discovering that intelligence and and memory comes from the strength of the connections between neurons in a sort of simplified manner. And the way the brain learns is by modifying the strength of the connections between neurons.
So, so some people came up with sort of theoretical models and and actually electronic circuits that reproduce this, you know, can we build? So, intelligence you're saying was largely the ability to solve a certain problem. So, that's the first view, right? To solve particular problems that that we're given. The second one is the ability to learn, right? Okay. And that created those two branches of AI, right? Um so the the the one that started with the ability to learn um there was some success in the late 50s early 60s and it died in the late 60s because uh the type of learning procedures uh for for those neural networks that people devised in the in the 60s turned out to be extremely limited.
You there was no way you could use this to produce truly intelligent machines. But it had a lot of consequences in uh various parts of engineering. A field of engineering called pattern recognition. Um so you're saying now that intelligence is the ability of a system to learn as well to learn and and the simplest situation in which you need machines to learn is for perception interpreting images interpreting sounds.
And what did computers use to do that? So for that it's it's basically what caused the emergence of what we could call classical computer science. Okay, you write a program and that program basically internally searches for a solution and has some way of checking whether the solution it it proposes is good or not. Um people had a name for this in the 60s they call this heristic programming because you can't you can never exhaustively search all solutions for a good one because the number of solution is ridiculously large.
You know at chess for example right you you can play a certain number of moves but then for every moves that you play your opponent can play a certain number of moves and then for every of those moves you can play a certain number of moves. So you get this exponential explosion of the number of possible trajectories basically or sequences of moves and uh you cannot possibly explore all of them until the end of the game to figure out which move to uh to play first.
So, so you have to use what what you know was called huristics to to basically not search the entire uh graph or tree of possibilities. So, we'll put up a graph explaining this. But what you're saying in heristics AI is you would have a user who would put in an input. There would be a bunch of rules and you would use like a tree search or an expert AI which would run a function like if this then that if not then this to try and get to an end state.
Yeah. So something that would but a defined end state be defined and and the the the program would be completely written by a person. Um and uh and the the difference between a good and a bad system would be in how uh smart the system is in in searching for a good solution without doing exhaustive search. Okay, that's the heristics part of it. Um a slightly different approach is the the one that's based on logic, right?
So you have rules and facts. what other facts can you deduce from the from the existing fact and the rules which would be logical formula and things like this that was you know pretty dominant in the 1980s um and that led to a um an area of uh of AI called expert systems or world-based systems um to some extent is very connected with this idea of search okay and then in parallel to this there is the bottomup uh approach you Let's try to reproduce the to some extent get inspiration from the the basic mechanisms of intelligence in biology implement um allow machines to learn and basically organize themselves um with the idea that how would you do that?
So so it's based on the same idea that neuroscientists figured out was going on in the brain which is that the learning mechanism in the brain proceeds by modification of the strength of the connections between neurons. Right? um and and and people had imagined that know this type of learning could actually be reproduced in machines. So first there was the idea that you could you know that neurons were simple simple computational elements.
There were proposal around those lines in the 1940s by mathematicians like Mullock and Pitts and people like that. And then in the 50s um and early 60s people proposed a very simple algorithm to change the strength of the connections between neurons so that they could learn the task. Um so the first machine of this type was called a perceptron and it was proposed in 1957. It's a very simple thing and it's very simple to understand.
Um let's say you want to train a system to recognize simple shapes um images. Okay. What is an image for a computer or for a artificial system? It's a it's an array of numbers. Um we know that today because we're familiar with digital cameras and pixels, right? So um let's take a a black and white camera. A pixel um is if the pixel is black, it's a zero. If it's white, it's a one. Okay, so it can take only two values, black or white.
Um, if you want to build this with 1950s technology, you would put an an array of photo sensors, photo cells, right, with with a lens in front of them and you would show an image very low resolution, maybe 20 by 20 pixels or something like this or even lower. Um, so now that gives you an array of numbers that you can feed to a computer, but what they did in 1950s, computers were incredibly expensive. So they actually built electronic circuits.
So the the pixels were voltages um coming out of the photo sensors. Um and then you want to train a system to recognize simple shapes. Let's say distinguish the shape of a C from the shape of a D uh drawn on this uh on this array. Um so you show an example of a C and then you let the system produce an output. This output will also be a voltage. And the way the output is going to be computed is a weighted sum of the of the values that come in of the pixels that are one or zero.
The weights are connections to a simulated neuron which is just an electronic circuit that computes, you know, if it's uh if it's a one or zero, I'm going to multiply this one or the zero by a weight, which is like a resistor that you can change the value of. Okay? And then um all of the pixels with their weight are going to be summed up. If the weighted sum is larger than the threshold, it's a C. If it's lower than that threshold, it's a D.
All right. What era was this? Which year did you say 1957? Um so now how do you train this? So training consists in changing the value of those weights. You can have positive or negative weights. Um and what you do is you show a C and the system computes the weighted sum. So for C, you want the weighted sum to be large, larger than zero, let's say. Okay. And let's say it's smaller than zero. So the system made a mistake.
So you tell it no, it should be larger. Okay? You press a button basically and you tell it I really want the output to be to be bigger. So what the system does is that it changes all the weights that get a one. So it increases them a little bit. If you increase all the weights that get a one, the weighted sum increases, right? And if you keep if you keep doing this changing the weights just a little bit every time eventually the weighted sum is going to go above zero and then the system will recognize this as a C.
And what did we use this for back in the 50s and 60s? So nothing really very practical other than recognizing simple shapes. Okay. Um so you you you repeat showing a C and a D and for for the C you say increase the weighted sum for the D you say decrease the weighted sum. So decrease the weights that have a one, increase the weights that have a zero. And then eventually the system settles on the configuration of weights so that when you show a C, it's above the threshold.
When you show a D, it's is below the threshold. So it can distinguish the two. And what it's going to do is, you know, give a positive weight to the pixels that only appear for the C and a negative weight to the pixels that only appear for the D. And that will sort of discriminate between those two. So we had we had heruristics AI, expert AI, trying to mimic biology, all of this in the 50s and 60s. In the 50s, yeah.
Starting in the 50s and then you know two different branches basically competing with each other and and they tried to kind of um so one person a prominent figure in um in AI in the pioneering days is Marvin Minsky. He was a professor at MIT. Here's a Marvin there is a I remember reading about this there's a Marvin clause or debate or something like that right um well he was uh he had pretty strong opinions about things so there was a lot of discussions um and he's interesting because he started his PhD in the 50s trying to build neural nets and then completely changed his mind and and and became basically a big advocate for the other approach the the more logic based and search approach And in the late 60s or mid-60s he wrote a book co-wrote a book with Seymour Peppert who was a mathematician at MIT um whose title was perceptron and the whole book was to do some theory about perceptron and to show that the cap the capabilities of the perceptron was limited.
So the people who were working on neural net at the time kept working on neural net but they changed the name of what they were doing. They called it uh statistical pattern recognition which sounds much more serious or adaptive filter theory which also sounds very serious and those had enormous applications in the real world. In my world it's always been I I work in finance and hedge funds and fund managers have always been attempting to pump a lot of data into a neural network to recognize patterns.
Right? Is it the same thing that we're talking about an evolution from the 50s? Yeah, absolutely. Um, I mean the process I describe of changing coefficients, you know, up or down to get the output you you want. Uh, you could think of this as a iterative process very similar to linear regression which if you work in finance you probably know about. Uh so it but what I've realized Yan is it's even today it's very easy to tweak data that you have collected retrospectively to make something appear like it makes sense but financial activity tends to be so random that I don't know if you can build a model based on that right so the um well that that addresses a bigger issue when you when you train a system this way right so the the The generic principle which is called supervised learning is u you give an input to the system it produces an output if the output is not the one you want.
Uh you adjust the coefficient so that the output gets closer to the one you want. Okay. And there are efficient ways to figure out how to tweak the parameter so that the output gets closer to the one you want. And if you keep doing this on hundreds, thousands, millions, billions of examples, eventually the system, if it's powerful enough, will figure it out. Now the problem with the perceptron is that the type of functions input output functions that was accessible to perceptron was very limited.
So there was no way you could take a natural image uh you know a photo uh uh and and train the system to tell you whether there is a a dog or a cat or a table in it. There was just not possible. The system was not able to um was not powerful enough to really compute this kind of complex function. uh this is what neural nets and deep learning changed in the 1980s and the what just before you get into neural nets if I'm trying to paint the entirety of the picture would you say there is intelligence on top artificial intelligence and below that is machine learning and neural nets are a part of machine learning yeah so in terms of fields and subfields AI is more of a problem than a solution it's a field of investigation and then there is different techniques you can use for that right so there is something that jokingly is referred to as good old-fashioned AI.
Goi which is using logic and search and heristic programming and things like this which is this is what you will find in sort of standard textbooks on on AI then there is machine learning so there the idea is you don't completely program a machine to do something you just you train it from data that means you need data within this there is a subcategory called deep learning and this is what the reason why we hear so much about AI in the last dozen years is because of deep learning and neural nets is really the ancestor of deep learning deploying is a new name for it if you want.
Mh. Um um and and then there is you know application areas. Um so below that so and and they can use combinations of those techniques right so so big applications are computer vision interpreting images uh speech recognition natural language understanding um and maybe speech synthesis also can be viewed as part of this although it's more connected with signal processing and then you know various other applications so in you know time series prediction or financial modeling and things like this you know could be seen as part of this if you So I'm I'm breaking it down.
AI has goi under it which is traditional in nature like you explained then machine learning. Can you define goi in a simple uh oneline definition? So goi is the the descendant of uh the the what I was describing earlier as searching for solutions right this idea that re it's all about reasoning reasoning is all about search uh you know looking for a solution to a problem and having a way of characterizing whether you found a solution so you mean the rule-based thing input and an output based on what applies the what rule applies like that um the the um yeah I mean any any rulebased system anything that uses logical inference.
Mhm. Um deducing facts from rules and previous facts. Searching for a solution like finding the shortest pass in a you know in a graph or something. Um those are good old fashioned AI and under machine learning. What are the different types of ML? So so there is uh so-called traditional machine learning. I'm not sure that deserves the term. And this is basically derived from uh statistical estimation. So things like linear regression would be part of it.
And then there are other methods slightly more uh sophisticated boosting um uh classification trees, super vector machines, kernel methods. I mean there's there's a bunch of methods of this type and bas inference that are part of machine learning in the sense that they they obey that model of you know you you you build a program but the program is really not finished. It's got a bunch of tunable parameters and the input output function is determined by the value of those parameters.
And so you train the system from data using this iterative adjustment techniques that I described before. Show examples. If the answer is incorrect, adjust the parameters so that it it comes closer to the answer you want. So machine learning is supervised in a way. So that's supervised learning. Okay. You you tell the system here is an output here is the desired desired output. Mhm. Um but there are other forms of learning.
So one one different form is reinforcement learning. So in reinforcement learning you don't tell the system the correct answer. You just tell it whether the answer it produced was good or bad. You give you give it a single number that tells it your answer was good or was bad. And what happens next? Say I'm a reinforcement learning engine and you tell me an answer was good or bad. What do I do next? Well, so if your answer was good, um you don't do much.
If your answer was bad, then um you have to figure out which answer among all the possible answers that could have produced, which one would be a better one. So maybe you try another answer and you say, what about this one? Is it better or is it worse? uh if the environment tells you it was better then you kind of deemphasize the first one and emphasize that one by tuning the parameters inside of a neuron net or something like that some sort of learning uh learning machine.
So what is self-supervised learning? Okay. So, self-s supervised learning is what has become very prominent over the last five six years and um is is really the the the main component or the the main contribution to the success of things like chatbot and natural uh language understanding systems. They don't fall under reinforce reinforcement learning. No, it's more similar to supervised learning. But the difference is that instead of having a clear input and output and training the system to produce the output from the input um you basically only have things that can either be input or output.
Let me take an example. Um you take a a piece of text and you corrupt that text in some way. So by removing some words, right? So now you have a partially masked uh text where some words are missing and you train a machine to predict the words that are missing. So the technique you would use for this is supervised learning because you tell the system here is the correct word that you should predict at that location.
Um and the system can use all the words that it can see to predict the words that it cannot see. And this is an example for supervised learning. Self-supervised learning. It's self-supervised because the there is no differentiation between input and output. It's really kind of the same thing. Mhm. And if the input is for example an image, the the way you um you would train a self-supervised learning system is that you would corrupt or transform the image in some way and then you would train the system to recover the original image from the corrupted or transformed version of it.
Okay? So there's no supervision. You don't need someone to go through a few million images and labeling them is it a cat or a dog or a table or a chair. Um, it's it's it's a task of basically understanding the the input uh the internal structure of the input by being able to filling in the to fill in the blanks. Forgive me for asking maybe a really stupid question. I'm trying to picture this. Let's say I have X amount of data.
I have 10 lines that say cats are black, dogs are white, whatever. 10 lines. I remove a part of it and then I tell the model to fill it in. Yeah. Are you saying at that point of time I also tell the model the answer saying this should be the answer? Yeah, you you you tell it here is the answer that I removed like can you predict this missing? Can you arrive at the answer which I removed and I'm telling you that this was the answer right?
But you can only use the thing that you can see. So you don't see the answer on the input. You have to predict it. But I'm telling you when during training and tell you what it is and so the system can adjust its parameter to its parameters in a supervised fashion. So the the only different the difference is not in the algorithms themselves. It's basically supervised learning but it's in the structure of the system and the way the data is uh is is used and and and produced.
You you don't need to basically have u you know someone going through millions of images and telling you uh this is a cat or a dog a table. Um you just show an image of a dog, a cat or a table and you corrupt it partially change it, change the colors maybe or something. um and then ask the system to recover the original one from the corrupted one. Okay, so that's that's uh one particular form of self-supervised learning and this is what's been incredibly successful for natural language understanding.
So things like so chatbots are or LLMs large language models are a special case of that where you train a system to predict a word but you only allow it to look at the world the the words that precede it. Mhm. Um you know that are to the left of it. Mhm. Um and that requires kind of building the neural net in a particular way so that the connections that that predict one word only look at the the words that precede.
So then you don't need to corrupt the input you just show an input and through the structure of the system the system can only predict can only is trained to predict the next word from from from the context and these are all examples of neural networks in a way these are all underlying this are particular way of connecting neural network neurons with each other simulated neurons right uh or or simple elements that compute a very simple mathematical function something like a weighted sum and what's adjustable are the weights or in the case of uh transformer architectures which are are very uh popular at the moment um uh they consist in basically comparing every input to each other and and producing weights.
I I could explain this is a little more complicated but what is a transformer? So okay so there are several architectural components uh which from which you can build a neural net. So let me start with um very simple idea. Let's say you want to build a neural net that recognizes images. Okay. So again an image is an array of numbers indicating the brightness of every pixel, right? Um you can build a neural network with a single layer.
So let's say you want to distinguish um 10 categories, okay? Cats, dogs, tables, and chairs and cars and whatever. Um or let's say it's simpler. We want to recognize the 10 digits. Okay, 0 to 9. Someone drawing a digit. It's drawn on a 16x6 pixel area. So you have 25 256 inputs and you have 10 outputs. You can have a single what's called a single layer neural net uh which basically each output is a weighted sum of of the pixels and you try to train those weights in such a way that when you show a zero the output zero is the most active and the other ones are less active and and so forth for all the categories.
Okay, that may work for simple shapes like like printed uh digits. it won't work for handwriting because there's so much variability in the characters that you cannot reduce the classification to a simple weighted sum. Okay, so the breakthrough that occurred in the 1980s was to um stack multiple layers of neurons. So each neuron computes a weighted sum and then passes this weighted sum through essentially a threshold function.
So if the weighted sum is below a threshold, the the neuron stays inactive. The output is zero. And if it's above a threshold, it's active. Okay, there's various ways to do this. Um, but it's nonlinear and that's very important. So you stack two layers where um the the middle layer you could think of as detecting sort of basic motifs on the inputs and then the second layer sort of integrate those motifs to figure out okay this is a a C because it's got two end points you know the the uh the shape of the C kind of stops there and I can detect that.
So if there's two of them, that's a C. And the D doesn't, but the C has two corners. Maybe I can detect that. The system learns to do this from end to end. And and the way it works is through an algorithm called back propagation. Um and what this back propagation algorithm does is that when you show an image of a C and you tell the system this is a C, so activate this output neuron does not and do not activate the other ones.
It knows how to adjust the parameters so that the output gets closer to the one you want. Um and that's done by propagating signals backwards to um to basically figure out the sensitivity of each output to um to to each weight so that you can change the weights in such a way that the good output increases and the bad outputs decrease. Right? So that's back propagation that um um algorithms to do this. So the back propagation algorithm popped up in the 1980s.
Conceptually it existed before but people didn't realize they could use it for machine learning and there was a wave of interest in neural nets starting in the mid mid mid 80s lasting 1015 years um to kind of exploit this idea of multi-layer networks and this was crucial because it it lifted some of the limitations that Minsky and Papert in the 60s said were you know the perceptron was the was subjected to um so a big wave of interest but then people realized that to train those neural nets um you need a lot of data and this was before the internet there was not much data you need fast computers and computers were not that fast um so so people kind of lost interest a little bit in this but one thing that I worked on in the late 80s early 90s is um if you want a system of this type to recognize images you kind of have to connect the neurons to each other in a particular way that facilitates the system sort of paying attention you know being able to detect motifs for example right local motives.
So um I got inspiration from biology again um a classical work in neuroscience that that went back to the 1960s to basically organize the way the neurons are connected to each other into layers. Uh so that they bias towards kind of finding good solutions for image recognition. So that's called a convolutional neural network or comb. Um so just just to come back to this where you are like so you broke down machine learning I'm sorry I keep going back sure or I'll get confused.
Yes. So under machine learning the really popular pathway right now let's say self-s supervised which has chap GPT and a bunch of other things. What's happening in the reinforcement learning space? So not so much anymore. Um there was a big wave of interest in reinforcement learning um about you know a dozen years ago and companies like DeepMind Mhm. set themselves up with the idea that reinforcement learning was going to be the the key element towards building truly intelligent machines.
Mhm. Can you again like define reinforcement learning once more in a line? So reinforcement learning is a situation where you don't tell the system what the correct answer is. You just tell it whether the answer you produced was good or bad. Right? Okay. Okay. So there are many possible answers. It's very inefficient because the system has to try many things before it gets the correct answer. Um and so it's very inefficient.
It requires many many many trials. And so it works really well for games. You know, you you it's very efficient. If you want to train a system to play chess or go or things like that, poker, reinforcement learning is great because you can have the system play millions of games against itself or copies of itself. U and it can adjust it, you know, it wins or loses a game. So it can, you know, it knows which policy, which flavor of the neural net won the game and sort of reinforces that and deemphasizes the one that lost and so the system basically can train itself, right?
And what did you say a transformer was? Okay. So I was coming to this, you know, through commercial net, right? So there is this uh um particular way of connecting simulated neurons with each other to bias it towards doing a good job for certain types of data. And uh conial nets are really good for uh data that comes from the natural world whether it's an image or an audio signal um which are things that are um where where nearby values in the array of numbers that come to you in an image or audio signal.
Nearby values are generally very similar to each other. So if you take a picture, any picture natural image and you take two neighboring pixels, they're very likely to have the same the same color or the same intensity. Now what I'm talking about here is the fact that the you know natural data like images and audio and just about any natural signal has some natural underlying structure to it. Um and if you build a neural net in a particular way that that can take advantage of this structure, it it will learn faster.
It will learn with fewer samples. So we started doing doing experiments with this in the late 80s and and and build those convolutional nets. They are inspired by the architecture of the visual cortex really. Um and there's some mathematical justification for it but the the basic idea is that um each neuron in a commercial net only looks at a small area of the image and and you have multiple neurons looking at multiple areas of the image and they all do the same thing.
They all have the same weights. um it's a basic concept which connects with the mathematical concept called convolutions and so that's why those things are called convolutional nets. Okay. So that's what's called an architectural component or a module. A convolution is something that has an interesting property which is that if you show it an input it's going to produce a particular output. If you shift the input the output would be shifted but otherwise unchanged.
And that's a very interesting property for audio signals images and various other natural signals. Okay. Now a transformer Mhm. is a different way of arranging the the the neurons uh if you will in such a way that um the the inputs are a number of different items. We call them tokens. They're really vectors which means list of numbers. Okay. And the property of a uh the layer or the block of a transformer is that if you permute the inputs, the output would be permuted similarly but otherwise unchanged.
Um when you say otherwise unchanged, you mean you mean I what I mean is that um if you you give a bunch of tokens, you run through the transformer, you'll get a bunch of output tokens. Okay. The same number generally as the number of input tokens. There'll be different vectors. Um if you now take the first half and the second half of of your sequence of input tokens and you flip them what you will get is the same result that you got previously but it would be flipped exactly the same way.
Okay. So the input output function is uh technically we call this equivariant to permutation. So it basically views the inputs as a set in which the order of the object does not matter. Okay. Um, convolutional nets on the other hand view the input as as something where an object could appear at any location on the input and it shouldn't make any difference to the output or the output should change but otherwise I mean should shift but otherwise stay unchanged.
That's equivariance to translation. Now when you build a neural net you basically combine components of this type uh so that you get the the property you want out of the entire neural net. So you combine things like convolutions and transformer blocks. What is a convolution? Yeah, I'm sorry. I'm going to ask you like to simplify every single term. Oh, absolutely. So, convolution is this component for a convolutional neural net.
So, the idea of it is that you have a neuron that looks at at a part of the input and then you have another neuron that looks at another part of the input, but it computes the same function as the first neuron. And then you replicate that same neuron for every location on the input. So that um you can think of each of those neurons as detecting a particular motif on a part of the input and all the neurons detecting the same motif at different parts of the input.
So that now if you take an input and and you shift it the you're going to get the same output shifted because you know you're going to have the same neurons looking detecting the same motif just at different locations. So that that's what gives you this shift equivariance. Um that's a convolution. Mathematically there's something called a convolution that mathematicians invented a long time ago and that's basically what this what this does.
When you say neuron in all of this can you explain the basis of just that term what is it? So uh we we we use that term it's an abuse of language because those neurons are not really neurons like in the brain they're they are to real neurons as an airplane wing is to a bird wing. Okay. So it performs the same it has the same concept and what a neuron does in a neuron net is computing a weighted sum of its inputs and then comparing that weighted sum to a threshold activating the output if it's above the threshold and and producing zero if it's below the threshold.
That's the the basic norm. Now there are variations of this and in a transformer it's a slightly different type of mathematics. you're kind of comparing vectors to each other and things like that. But uh but that's kind of the basic functionality of a neuron. It's a combination of a linear operation where you have coefficients that you can change the value of through training and then a nonlinear uh function a threshold or something like that that uh you know detects something or not.
Right? We looked online and while we were researching we could not find a good definition for neural network language model and how it works in simple terms. Okay. Um so the idea of a language model goes back to the 1940s. A gentleman called Claude Shannon. He's a very famous uh mathematician who used to work at Baz where he used to work um although he wasn't there anymore when I joined. uh and he he came up with a theory called information theory and then was fascinated by the idea that you could discover the structure in data.
Right? So he invented something where you take a text um and and you say I'm giving you a I'm giving you a sequence of letters and I'm asking you what is the next letter that comes afterwards. Okay? So let's take a you know English word um or whatever in in a sort of let's say you know Roman language if you have a series of letters and the last one is a Q, it's very likely that the next letter is a U. You almost never have a Q without a U behind it unless it's an Arabic word or something that's been transliterated, right?
Um so for every every letter that you observe you can you can build a table of the probability that the next letter will be an A a B a C. This is where the word generative comes from. Yeah. So it's generative because if you have this table of conditional what we call conditional probabilities right given the previous letter what is the next what is the probability of the next letter? You can use this to generate text.
Mhm. You you start with a letter, let's say Q. Okay. And then you look through the table of probability. What's the next letter that is most likely? You just pick that one. That's going to be U. Mhm. Um or you don't pick that one. You you you pick the next letter with the probability that is, you know, you flip a coin or you generate a random number in the computer and then you you produce the you know, the following letter according to the probabilities that you measured on real text.
Um, and you keep doing this and the system is going to just generate letters. Um, it's not going to look like words. Uh, it's probably not even going to be pronouncable, right? Um, but that if instead of a context of one letter, you take a context of two letters, then it becomes kind of more readable. It's still not words, right? If you take a context of three letters, then it it becomes, you know, even nicer. And as you increase the size of the context that determines the the probability of the next letter, it becomes more and more readable.
But you have an issue there, which is that um uh the the size of the table you need if you have if you look [snorts] at the first letter and and and and you have to figure out what's the probability for the next letter. You need a table of 26 rows and 26 columns for each first letter. What is the probability for every possible second letter? Right? So it's table 26 x 26. Now if you the context has two letters. Now the number of rows in your table is 26 squared because you have 26 squared sequences of two possible sequences of two letters, right?
And for each of those you need 26 probabilities. So it's 26 cube the size of your table. As you add characters um the table increases to 26 to the^ n when n is the length of the of of the uh of the sequence. So that's called an engram model and that's a language model. You can do this at a level of characters. It's more difficult to do this at a level of a word because you might have 100 thousand possible words, right?
So now your table is gigantic. Mhm. So you can train a word model or or language model by just filling up this table of probabilities by training on a large corpus of text. uh but it becomes impractical uh above a certain length of context number because of the amount of compute and work required. It's it's also it's the memory of storing all those tables and uh also the fact that um those tables are going to be very sparsely populated because you can have billions of words of text.
Most combinations of words don't appear. Some of them are extremely rare and so you cannot estimate the probability properly. Okay. So is this a part of self supervised learning? So you could think of this as an instance of self-s supervised learning because you only need sequences of symbols and it doesn't matter where they come from and you don't they don't necessarily come from human uh production if they're not text if it you know it could be for example uh a sequence of frames for a video right I mean you would have to turn it into discrete objects which of course doesn't is difficult but um uh but it's you know whatever data comes to you so uh in the late 90s um some people had the idea that in particular your Benjio the idea that you could use a neural net to to do this prediction instead of filling up tables with conditional probabilities that you measure from text.
Just train a neural net to predict the next word. Okay, give it a context of words and just train it to produce a probably distribution over the next word. And he experimented with this with uh you know neural nets that were big for the time but small by today's standards. And one difficulty was you cannot exactly predict what word is going to come next. So you have to produce a probability over all the words. And there's maybe 100,000 words in a typical um uh language.
And so that means you're you need to output a 100,000 scores, one for each word that indicates with which probability that word follows the the previous sequence of words. Um so he demonstrated that that that could work and you know even with the computers of the time it was kind of challenging but but it but it could work and then the idea was kind of revived more recently. Um, and it turned out that if you use those transformer architectures, which I didn't explain, uh, and you you train them on basically the entirety of all the publicly available text on the internet.
Um, and you you build the architecture of the system so that it's trained to, you know, take a context of words and predict the next word. Um, and if you make the context potentially very large, something like a few thousand, a few tens of thousands or even a million words, then you get systems that seem to have emergent property that they can answer questions. They can, you know, if you make them really big, they have so many parameters that are adjustable.
They have they may have tens of billions or hundreds of billions of parameters that gives them a large amount of memory and they they they seem able to store a lot of knowledge about the the data they've been trained on. If it's text, they will regurgitate solutions to puzzles. They will, you know, give you answers to um to questions you may have. It's mostly retrieval. There's a very tiny bit of reasoning, but really not much.
Uh, [snorts] and that's an important limitation, but it's still surprising how well those things work. And it's um, you know, what people got really surprised about is that those systems can manipulate language in ways that are uh, very impressive, right? I mean, humans have pretty limited in the way we manipulate language and and those things seem to be really good at it. I mean, they capture grammar and, you know, syntax and everything in multiple languages, right?
That's pretty amazing. So if I were to like go back and paint a tree. So let's say AI on top, machine learning under it. I'm talking about what is making the news today and what everybody's so excited about. Machine learning has different things, different neural networks under it. There's a reinforcement one like deep mind. There is a self-supervised generative chat GPT because using it as a placeholder as it's the most popular one right now huh LLM auto reggressive LLM really that's what it should be called auto reggressive LLM yeah I mean the the the proper organization is yeah there is AI at the top uh machine learning is a particular way of approaching the AI problem under this is deep learning which is really the the the foundation of pretty much [snorts] all of AI today.
Mhm. Um so basically neural networks with multiple layers, right? The idea of this goes back to the 1980s and back propagation. That's still the the the basic foundation of everything we do. Under this there is several families of architectures, convolutional net, transformers, combinations thereof. Um then there is under transformers there is several um flavors of it. Some of which can be applied to image recognition or audio.
Some of which can be applied to representing natural language but not generating it. And then there's a subcategory large language models which are auto reggressive transformers. So transformers have a particular architecture that uh allow them to predict the next word and then you can you can use it to just generate word because you know given a sequence of word has been trained to produce the next word. So given a text you have produce the next word and then you shift the input by one.
So now the word it generated is part of its input and you can ask it to generate the second word shift that third word shift that fourth word that's auto reggressive prediction. It's the same concept as auto reggressive models in finance and econometrics and stuff like that. Same stuff. And these work best for text but not for pictures, videos or any of that. That's right. And the reason it works for text and not for other things is because text is discrete.
So there is a finite number of possible things that can happen right there's a finite number of words in the dictionary is a you know so if you can discretise your signal then you can use those auto reggressive prediction systems and the you know the main the main issue is that um you're never going to be able to make an exact prediction. Mhm. And so the system is going to have to learn um some sort of probability distribution or at least you know produce scores that are different for for uh for different potential outputs.
So you can output a list of probabilities if you have a finite number of possibilities which is the case for language. Um but if you want to predict what is going to happen in a video the number of possible frames video frames is essentially infinite right you you have you know let's say a million pixels right an image th00and by thousand pixels the pixels are in color so you have three values so that's three million values um that you have to produce and and we don't know how to represent a probability distribution over the set of possible images with 3 million pixels.
But this is what everybody's very excited about. This is what a lot of us consider the next challenge in AI. So basically have systems that can learn how the world works u by watching videos. And if you were to say videos learn from videos and pictures which will be the next phase, where does that fall in this entire equation? Does it come under where LLM sit today? No, it's completely different from LLM, which is why I've been u pretty vocal about the fact that LLMs are not the path to human level intelligence.
Um, LLMs work for discrete worlds. They don't work for continuous high dimensional words, which is the the case for video. And this is why LLMs do not understand the physical world. Um, and and cannot be used in their current form to really understand the physical world. And so we have I mean LLMs are amazing in their ability to manipulate language but they can make very very stupid mistakes that reveal they really don't understand how the world works right the underlying world.
And um this is why we have systems that can pass the bar exam or or write an essay for you. But we don't have domestic robots. We don't have self-driving cars or completely autonomous level five self-driving cars. we we you know we don't have systems that really understand very basic things that your cat can understand. So I've been you know kind of vocal saying that you know the smartest LLM are not as smart as your house cat and it's really true.
So so the the challenge for the next few years is to build AI systems that lift the limitation of limitations of LLM. So systems that understand the physical world are have persistent memory which LM really don't have at the moment. Persistent memory. Persistent memory, which means, you know, they can remember things, right? Store facts in a in a memory and then retrieve them when it's interesting. Um, can't LLM remember stuff?
Now, the only memory that an LLM the only two there's two types of memory that LM has, the first type is in the parameters in the coefficients that are adjusted during training, right? So, they will learn something. they they don't they it's not really kind of storing a piece of information. If you train a LLM on a bunch of novels, it cannot regurgitate the novels, but it will remember something about the statistics of the words in that novel and it might be able to answer questions, you know, general questions about about the story and things like this, but it's not going to be able to regurgitate all the words, right?
Um kind of like humans, right? You read a novel, you can't you can't remember all the words. you unless you spend a lot of efforts trying to do this. So that's the first type of memory and then the second memory is the context the the prompt that you type. Mh. And since the system can generate word and and and those words are or those tokens are injected in its input, it can use this as some sort of working memory. But it's a very limited uh form of of memory.
What you want is a memory that would be more similar to what we have in our brains. What mamleians have called a hypocampus. Um hippocampus is a kind of a brain structure in the center of the brain in our cortex. And if you don't have a hypoc campus, you can't remember things for more than about 90 seconds. And if you were to draw a path from intelligence that we described on top all the way down to self-supervised learning, how do you suspect that path will look towards us getting to the point where we are learning from videos and images and more humanlike intelligence?
So the the path that um I've been trying to to plot um is u discovering new architectures different from those auto reggressive architecture used for LLM that would be applicable to to video so that self-supervised learning could be used to train those systems. And this type of self-supervised learning basically would be here is a piece of a video predict what comes next. Um and and if a system can do a good job at predicting what's going to happen next in a video that means it probably has understood a lot about the underlying structure of the world similarly to a large language model learns a lot about you know language by just being trained to predict the next word right not like I will understand but if you had to give us a line on how that architecture might look okay so here is the issue because as I told you um those auto reggressive architecture work for text because text is discrete and you can never predict the what what comes next.
We can produce a prob probability distribution over what comes next. You cannot do this for images and video because it's just too complicated mathematically and you can show that it's intractable and blah blah blah. So predicting all the pixels in a video that follow a particular video segment basically is not possible or not possible to a degree that would be useful for the problem that we're interested in. You know what we want is a system that has the ability that uh the the ability to predict what's going to happen in the world because that's a good way to for a system to be able to plan.
If I can plan that if I approach my hand, you know, to this glass um and I I close my hand and I lift it up, you know, I got to grab the glass and I can drink. Um I can plan a sequence of actions to arrive at a particular result, right? Right? So I have a good model of the world that says the state of the world at time t is this the glasses on the table. The action I'm going to take is close my hand around it. Okay.
Um and lift. What is going to be the state of the world at time t plus 3 seconds after I close my hand and lifted my arm. And the state of the world is going to be I'm going to have that glass in my hand. Um so if you have this kind of world model um state of the world action next state of the world then you can imagine you you can predict you you can predict the outcome of a sequence of actions. You can imagine taking a sequence of actions and then predict in your mind what the outcome will be.
You can predict if this outcome is something that satisfies a goal that you want to accomplish like drink a little bit of water take a sip and what you can do is through search. So now we're connecting with old AI search a sequence of actions that will actually satisfy this goal. Um so this is the type of um reasoning and planning that psychologist call system two. Okay Daniel Kaman is a late um Nobel Prize winning psychologist.
Um and he he makes this distinction between system one and system two where system one is actions you can take without thinking subconscious it's just reaction reactive and the system two is what you have to deliberately plan uh and think about to be able to uh produce an action or a sequence of actions. So Yan will memor memory eventually be the answer because as humans from biology we learn through memory right well it depends what type of memory I mean we also have multiple types of memory we have the hypoc campus that I I mentioned so hypoc campus is used to store long-term memories like um you know things that happened to you when you were a child and things like that uh basic facts about the world like you know when your mom was born or something um you know also So you know which way you came in here.
So where's the door? So this is more recent short-term memory. Episodic memory, working memory. So if you're thinking about something, you're kind of manipulating things in your head. You have to kind of temporarily store uh data. That's the hypoc campus. Um and and your cortex does the computation and basically reads from this memory and updates it. Okay? It's very much like a a bit of a computer where the cortex is the CPU and the hypoc campus is the memory, right?
um that you read from and and write into u but the current design of AI systems is not like that. So L&Ms do not have a separate memory other than the prompt that you can generate token in uh and they don't have this ability to search through a set of answers for which one is the correct one although they're starting to have that to some extent. So you may have heard of 01 from OpenAI and there's kind of similar work um at Meta and other places where um this sort of very basic forms of of reasoning that consist in having an LLM produce lots of different sequences of uh of of words and then having a way of searching through this list of word which one is the best but it's very inefficient.
So ultimately that's not what what you want. So going back to the question of how do we get machines to learn by observing the world from learning from video, we cannot use the architectures that are generative that just produce every pixel in the video. That's just completely impractical. And I've tried to do this for almost 15 years. Okay. Um and five years ago we came up with a different way of doing things that I called Jeppa.
So it's a different architecture and that means joint embedding predictive architecture. What it means is I watched this for a long time on your Lex Freedman interview when you spoke about Jeppa and I still don't get it. Okay, here's a [clears throat] basic idea. Tell me if you don't understand because I can explain it in different ways. uh instead of taking a piece of video and training a big neural net to predict all the pixels of the continuation of that video, you you you you take the the video and you run it to an encoder which is going to be a big neural net that's going to produce an abstract representation of the video.
Okay? And then you take the reminder of the video, the the the future, you know, the second half of that video, run it through the same encoder, and then you you train a prediction system, which is similar perhaps to much like LM where you delete a part of the data to train the model. That's right. So, you know, an LLM, you you take a piece of text and you train it to predict the reminder of the text, right? And you do this word by word, but you could do well, you could predict multiple words.
So, here we're going to do the same thing. We're going to take a a video and then train a system to predict the reminder of the video. But instead of predicting all the pixels in the video, we're going to run those videos through encoders which are going to compute abstract representations of the video. And we're going to do the prediction in that space of representations. So instead of predicting pixels, we predict abstract representations of those pixels where all the things that are basically unpredictable have been eliminated from the representation.
So is that a bit like also predicting tomorrow? Cuz if I were to video my life up until now and run it through the encoder, it will give me some kind of representation of tomorrow. Well, yes, but at an abstract level, right? So you can predict um you're based in Bangalore and I I heard so at some point you're gonna fly back to Bangalore. uh and you can predict how long it's going to take to go back to Bangalore but you cannot predict all the details of what will happen uh during your your journey back to Bangalore exactly how long it's going to take given traffic how far can you extrapolate what will happen 3 months from now if I have data video data of the last 10 years of my life so here's a trick the interesting question you can predict very long term but the the longer in the future you can predict predict the more abstract the representation level at which you can make the prediction.
Let me ask you a question. If you were to extrapolate 50 years forward all of our lives, you figure out how to build this architecture and it's implemented and it's working where video of our life up until now has been programmed into it and we're trying to predict 50 years forward. What do you suspect you will see climate change and world war? So what I see is okay so there is a a plan for the next few years to build systems that can understand the world from video.
Um perhaps what they'll be able to learn are those world models which are action conditions. So they they will be able to imagine what the consequence of an action or sequence of action will be. They'll be perhaps able to plan complex sequences of action hierarchically because those world models will be hierarchical. They will have world models that can predict really short term make accurate prediction but only in the short term like if I move my muscle in this particular way you know my arm is going to be in this particular location 100 millisecond from now that's really short range but very precise and then longer term prediction um if I go to the airport catch a plane I'll be in Paris tomorrow morning or you know if I study uh and I get good grades in in college you know I can have a good life or something right uh so so you can make long-term prediction and and design plans that would satisfy certain criteria that that you have.
So if we can build if AI were to predict the future, would it be utopian or dystopian? It would be utopian because it would be just a an alternative way for predicting the future than our brains and for planning action sequences to satisfy certain conditions to achieve goals. that is uh alternative to using our brains perhaps accumulating more knowledge to be able to do this and perhaps having abilities that humans don't have because of the limitations of our brain right computers can calculate and stuff like that right so so the the future is that if we succeed in this plan which may succeed within the next five or 10 years you know five to 10 years we have systems that as time goes by we can build up to become as intelligent as humans perhaps.
So reach human level intelligence within a decade. That may be optimistic. All right. Um 5 to 10 years would be if everything goes great. All the plans that we're we've been making will succeed. We're not going to encounter unexpected obstacles. But that is almost certainly not going to happen. You don't like that, right? Like AGI and human level intelligence you think is far far away or unlikely. No, I I don't think it's that far away.
I I don't think my opinion about how far it is are very different from what you will hear from Sam Alman or Deis Sabis or things like this. Um it's you know quite possibly within a decade but it's not going to it's not going to happen next year. It's not going to happen in two years. It's going to it's going to take longer and so you don't want to extrapolate the capabilities of LLM and and say we're just going to scale up LM train them on with bigger computers on more data and you know human level intelligence intelligence is going to emerge.
This is not going to work this way. we're going to have to have those new architectures, those jas systems that learn from uh from the real world um and can plan hierarchically uh can can plan a sequence of actions, you know, as opposed to just producing one word after the other essentially without thinking. So system two instead of system one. LM are system one. The architecture I'm describing, which I call objectived driven AI, is system two.
I'd love to come like do a course at your college and learn if you'll have me as a student. I don't know if I qualify. I'll have to go back and finish high school, but would would love it. Just to finish the LLM loop. So, because it's in the news and everybody's talking about LLMs. So, you define a problem, you find a large data set. Most of the time goes in cleaning the data. You choose a model, you train the model, and then you execute the model.
Uh before that, you fine tune the model. Before that, you fine-tune the model. Yes. what will change here? Um, so there's still going to be a need for collecting data and uh and filtering data to to to keep high quality data and basically get rid of junk. That's actually a pretty expensive part of the whole thing. But I think what's going to need to happen in that respect is that currently the the you know LM are trained with a combination of publicly available data and licensed data basically um but it's mostly publicly available data you know publicly available text on the internet right and it's extremely biased in many ways that um the the u a lot of it is in English um you know this you know significant amount of of data in in commonly spoken languages like Hindi but not so much in all 22 official languages of India and certainly not in all the 700 dialects or or whatever the number is particularly since most of those dialects are not written so only spoken so um what we need in the future is uh data sets that are more encompassing so that the the systems that are trained with it understand all the world's languages all the world's cultures all the value systems you know everything and no single entity I think it would be able to to do this.
Um which is why I think the future of AI is AI is going to become a a kind of common infrastructure which people will use as a repository of all human knowledge and this cannot be built by a single entity. It's going to it's going to have to be a collaborative uh project, right, with training being distributed all around the world so that you can have models that is trained on all data around the world, but you don't have to copy the data anywhere.
And a private digression, I was reviewing a data center business to invest into. Uh a lot of people tell me that compute as a commodity will soon be sold outside of the data center and not inherently in it. Is it a good place to focus energy and time on like building data centers out of India? I'm taking the sovereign AI model where every country will probably fight to retain their data a bit more than they're doing currently.
Yeah. So in in that kind of future which also I I alluded to with the distributed training of models. Uh having local computing infrastructure I think is very important. So yes I think that's kind of crucial. It's crucial for two reasons. one is uh have having local ability to to train models. Okay. Uh and the second one is um having very low cost access to inference for AI systems because if we want AI systems to be used by I don't know 800 million Indians right um I know there are more Indians than this but most people you know uh not everybody will use AI systems but um it's a lot of computing infrastructure it's actually much bigger than the infrastructure for for learning um and there is this scenario for which there is a more innovation than than training.
Training is dominated by Nvidia at the moment. There's going to be other players, but they they have a hard time um competing because of the software stack. Basically, their hardware may be really good, but the software stack is uh is a challenge. For inference, though, there's a lot more innovation there and and that innovation is bringing down the cost. I think the cost of inference for LLM has gone down by a factor of 100 in two years.
I mean, it's it's amazing, right? It's way faster than mors law. Uh and I think there is still a lot of room for improvement and you need that because you basically need the inference for a million tokens to be a few rupees. Um so so that's a big future if you want to deploy AI assistance widely in India. I want to I want to use the time Yan because I realize we're running out to bring it into the Indian context. Are people watching this like I said are entrepreneurs in play or people trying to be entrepreneurs as an Indian 20-year-old who wants to build a business in AI a career in AI what do we do like as we sit today a 20-year-old u today I would cross my finger so that when I graduate at 22 uh there will be good PhD programs you know in India outside of the academic lens I I mean more no no but that that's that's what I need to train myself to innovate you know doing a PhD or graduative studies it trains it trains you to invent new things and and also make sure that the methodology you use um prevent you you from from fooling yourself into thinking you're being in an innovator but you're not okay so you learn this what if I'm an entrepreneur a 25year-old entrepreor you still want to do a PhD if you're an entrepreneur or at least a masters because you want to really sort of learn deep I mean you might be doing this by yourself you don't have to but it's useful because you you learn more about you know what exists out there what's possible what's not possible what uh uh you get more uh legitimacy in hiring talented people I mean there's a lot of advantages you know particularly in a complex deeply technical uh area like like AI um you might succeed If you don't, you know, that's not the issue.
But, but it gives you kind of a different perspective. Okay. But now, you you know, you're doing your PhD, you're you're doing a startup. It might be easier to raise money if you've published a few papers where you've invented something new and you say, "Well, this is a new technique that really may make a difference." You know, you go see an investor. What What if I were to even like go one step further? Let's say intelligence.
I'm I'm going to leave the AGI side of it. Let's say narrow intelligence, self-driving cars, robots, all of that. What should I build in? If I have to pick a subset where I can use narrow intelligence through any of the models that we spoke about, what would I start which has a capitalistic leg to it? today like right now uh the the most likely business model uh that has to do with AI is taking a open source foundation model like LMA which is the open source system which is used everywhere now right every almost every startup uses uses it even large companies uh so take an open source platform um whether it's an LLM or an image feature extraction system or a segmentation system whatever and then uh fine-tune it for a particular vertical application and become an expert in that vertical application and which vertical should we pick the guy any vertical right so but I want to know like give me like top three we did gates we interviewed him recently and he said focus on building that layer around law because the legal processes arrived for disruption that's a good example right if you had to pick one or two more well there's there's I mean in sort of B2B there is there is legal, accounting, business information, right?
I want some report on the competitive situation in a particular segment of the market. F, you know, fintech, finance, I mean those are obvious obvious areas. Um, uh, you know, LLM's information system that give you all the private information inside a company so that any employee can ask any question about anything, administrative or whatever. Uh, and and you get the answer. So you don't have to plow through like multiple internal websites and information systems.
So that that's um that's certainly a good thing. Um and I think there is a lot of work there in [clears throat] uh basically companies that can fine-tune models for particular verticals and then there is you know other markets that are more consumer um assistance for various things uh you know in in for education there's not a huge amount of money there unless you can get contracts from the government but education certainly is a web application probably the the the other big one is health right so there are a lot of companies particularly in developing world that are being formed for using LLM to provide assistance, medical assistance essentially.
You know, you call up your LLM and you say, well, you know, I have those symptoms, you know, should I go to the hospital or or like, you know, this issue I have and it's much easier than to get an appointment with a doctor. There's certain parts of the world where it's basically impossible to get to see a real doctor. Um, you have to, you know, travel to a city or something. Mhm. So um so I think they would be useful, you know, other applications in rural areas particularly things that are enabled by AI assistant.
They can speak local languages u and and serve people who are not particularly comfortable with literacy that don't write very much, don't read very much. So interacting in your own language through speech with an AI assistant I think opens up a lot of applications um in agriculture in you know all kinds of are and if I were to switch the lens from an entrepreneur to an investor what would a investor benefit from investing into AI would it be Nvidia Llama Meta Chat GPT open AI okay uh so I think the first order thing is imagine what the future is going to be five years from now.
And basically that's going be that's going to be dominated. I suspect you would do a much better job at imagining the future than I can you pick can you depict a future five years from now. So five years from now the world is going to be dominated by open source platforms u for the same reason that the world of you know embedded devices and operating system is dominated by Linux. Mhm. Uh, you know, the entire world runs on Linux and it wasn't the case 20 years ago, 25 years ago.
Um, and it's become so because open source platforms are more portable, they're more flexible, they're more secure, they, you know, they they're cheaper to I shouldn't take credit for this, but we have somebody called Kellash who's our CTO who's a big proponent of this and everything we do is open sourced. We have a fund which gives grants to open source companies and stuff like that. Right. Okay. So, the world is going to be open source.
We're going to have open source AI platforms uh in a few years. They'll probably be trained in a distributed fashion. So they're not going to be kind of completely controlled by a single company. Uh the proprietary engines I think are not going to be nearly as important as they are today because the open source platforms are catching up in terms of performance. And then what we know is that a fine-tuned open source uh engine like Lama always works better than a non-fine generic uh you know top performing model.
So but if everything is open sourced it'll also be democratized for a investor to invest into then what is the differentiation? Well it enables the ecosystem. If you are a startup, you're much better off using a source engine and fine-tuning it for a vertical application than you are to you know using an API because you can build a tailored uh product for for your customers in a in a much much better way. So that's the first thing.
Second thing is if you really want this technology to be democratized and used by everyone eventually using smart glasses and stuff but but you know at first just smartphones. Do you think the form will change soon? the form of how you interact with technology will move from smartphones to different kind of devices soon. Smart glasses. Yeah. I mean, yeah, there's almost no question you're using one. So, I don't have them right now, although they're in my in my bag right here.
Um uh I use them all the time. Yeah. I I find them really really uh useful for all kinds of stuff. Even if you don't use AI for it, just taking pictures or listening to music or whatever. But but then then you have the AI assistant and I could be you know sitting in a restaurant with a menu in a foreign script and a foreign language and they could be translating it for me right so what happens to intelligence in society with all of this changing what becomes forget computers and AI for a second for humans what is intelligence in that world so people's intelligence will be moving to different set of tasks than the one we are trying to do today.
Um because a lot of what we're trying to do today will be done by AI systems and so we will focus on other tasks. So things like not doing things but deciding what to do or figuring out what to do. Okay, those are two different things. Like think about the difference between a low-level employee in a company. Mhm. That is told what to do and just does it and then you know a high level manager in the company that has to figure out like strategy and think about like what to do and then tell people below what to do.
Okay, we're going to we're all going to be a boss. We're all going to be like those uh high level managers. We're going to tell our AI assistants what to do, but we're not going to have to do it ourselves necessarily. Okay. So we need lesser [clears throat] people to tell something more efficient than us what to do than we need today for them to actually do the task, right? So what happens to everyone else? Well, I think everyone is going to be in that situation is going to have access to AI assistance and um and and be able to delegate a lot of tasks u you know mostly in the virtual world but eventually in the real world.
We're going to have at some point domestic robots and and self-driving cars and things like this once we figure out how to get the system to learn how the real world works from video. Uh but um so so the the the type of task on which we're going to be able to concentrate oursel are going to be more abstract. The same way you know nobody needs to do like super fast mental arithmetics anymore. we have calculators. Mhm.
Or or you know solve integrals of differential equations. We we have to learn the basics of how we do this. But we have you know we can use computer tools to do this right. Um so so it's going to lift the abstraction level at which we can place ourselves and basically enable us to be more creative uh be more productive. Okay. And and there are a lot of things that you and I have learned to do that our descendants would not have to learn to do because that would be taken care of by machines like go to school.
No, no, we'll still go to school. We'll have to educate ourselves. We'll have to there's still going to be the the you know the the competition between humans to kind of do something better than the others or something different, more creative always. Right. Innately we want to compete with our beer group. Yeah. Yeah. So we're not going to run out of jobs. M economists that I talk to tell me we're not going to run out of jobs because we're not going to run out of problems but but we're going to find better solutions to problems with the help of AI.
Maybe we can end today Yan trying to define what is intelligence really. I had written down intelligence is a collection of information and the ability to absorb new skills. It's a collection of skills and an ability to learn new skills really quickly or an ability to solve problems without learning. This is called zero shot in the in the AI business. You know, you're faced with a new problem and you can think about it for a while and you may not you may never have faced similar problem before.
You can solve it by just you know thinking and and using your mental model of uh of the situation. That's called zero shot. You're not learning a new skill. you're just solving a problem from scratch. So the combination of those three things, you know, having already a number of skills that you know, experience with solving problem, accomplishing tasks, being able to learn new task really quickly with a few trials. Um, and then the next step is being able to solve new problems zero shot without having to learn anything new.
Um, that's the combination of those three thing things really is intelligence. No, thank you, Yan, so much for doing this. Uh, I'm going to try and figure out how I can do like a course under you wherever you're teaching. Maybe you can recommend me uh to the college to give me a seat so I can attend some lectures. But I'd love to know, you can do better. Uh the 2021 edition of my deep learning course, yeah, is fully available on the internet for free.
It's all on YouTube. All the problems, all the um the uh homework and everything. I I feel like I'm going back to old school. I feel like being in front of you and learning first person has a innate value of its own. So I'll I'll try and do this and that. Wonderful. Thank you so much for doing this. Thank you. Pleasure. A real pleasure. Thank you. That was fun. That was fun. Yeah. You didn't get bored? No. you ask you ask a professor to speak like you know that's the job but I I guess like when you're talking to people who know so much lesser than you it can't be fun all the time it's it's an art uh I mean I I don't I don't claim to be I don't claim to be particularly good at it but I I try hard so you know trying to kind of simplify concepts and stuff like that.
Yeah. So, but I think we needed it because so many Indians are speaking about AI, but so few of us actually understand what went behind where we are. Well, that's true across the world. It's not just India. I mean, in fact, I think it's kind of the opposite in India. There's like way more people who are who are kind of educating themselves like particularly among the young people. So, we wanted to focus today on that like just to get to telling our people a lot of young people watch this, young bright people.
Yeah. how we got to be where we are today. The most questions around that I'm glad that Yeah, I think it's important because it it it helps convince people that uh they can do it regardless of what like you know I went to I studied engineering in France but I didn't go to like one of the top you know equivalent IV league or anything. I went to like a regular school, right? Um um and I didn't do my PhD was a famous person and you know all that stuff and I was in France.
I was writing papers in French that were terrible. nobody read but you know kind of managed to do something and and people are are sometimes telling me like you're you know you helped convince me that I could do something impactful even though I didn't go to Harvard or MIT or Stanford. Thank you.