🎙️AI 访谈库
驯服硅谷——Gary Marcus 谈新书与 LLM 批判(MLST 播客)
Gary Marcus · 纽约大学名誉教授、《Taming Silicon Valley》作者

驯服硅谷——Gary Marcus 谈新书与 LLM 批判(MLST 播客)

Taming Silicon Valley — Prof. Gary Marcus (Machine Learning Street Talk)

2024-09-24 · Tim Scarfe · 1h57m · 约 122 分钟读完 · 原文
配合新书《Taming Silicon Valley》的两小时长访谈:Marcus 讲 LLM 的技术缺陷与硅谷 AI 公司应受的监管约束。他的一贯尖锐立场:幻觉与分布外泛化失败是纯 LLM 路线的架构性缺陷而非工程小毛病,出路在神经符号混合方法。

why do you do what you do what after all of these years you you keep persisting I mean why did you write this book well the reason I wrote the new book is um because I sensed a moral decline in Silicon Valley that's in fact one of the chapter titles is the moral decline of silicon valy um I particularly sensed it when Microsoft um released this product called uh uh Sydney and Kevin Roose had this conversation with it in which it told him to get a divorce and this stuff and instead of pulling the product Microsoft just put some Band-Aids on it and that was a sign to me that things had changed and within a few days of that SAA said that we're going to make Google dance the whole culture changed like overnight um you know the antecedent condition precipitating condition was the popularity of chat gbt suddenly you know people thought there was real money to be made here and their postures changed entirely and that worried me because I do think the technology is premature I do think a lot of harm can come from it and I kind of dropped what I was doing research-wise um and really moved full-time into policy and in some ways the book is actually a memoir it's not couched that way at all but it's really A Memoir of that time when Gary went to the senate had a great conversation with the Senators and then realized that nothing was going to happen so you know I I had this peak experience talking to the Senators and feeling like they were going to do something and then gradually bit by bit and I was warned you know that this might come down this way but I had to see it for myself I'm naive I guess in that at least that one respect um seeing that nothing was actually happening and seeing you know learning close up how the lobbying works you know like one time I was in DC and I took a meeting with Google and it was a time when I got the closest hotel I could get to the capital um because I was going to be like in and out and I had to have a bunch of meetings and I couldn't you know be late for the meetings and like Google had an office next to that hotel which is like right on top of the capital like they're so there embedded in it in every level I mean that's just a sort of metaphor um and it's not just Google it's it's open Ai and and meta and all once you learn how all that lobbying is working and once you see how like good ideas that have big support go to die never even get voted on like it's incredibly disillusioning it's the disillusionment really led to this book called taming Silicon Valley I realized that we are heading very quickly towards an oligarchy I mean imagine the data that these guys are getting like open AI is like getting access to everybody's documents and wants like you know everything about you which they're of course going to weaponize they're going to sell in some you know various ways including to targeted political advertisers probably you know I mean they'll say they won't but like we've seen this movie before um they're getting an enormous amount of power a lot of that is because people think that they're going to get rich which may not not actually be true ironically so they've been given power in advance of actually delivering the goods like people think they're going to make AGI and therefore they should be powerful but in fact they've made llms which are not AGI that are not actually commercially useful but they given enormous power they're like on all these committees and whatever um and so Silicon Valley has suddenly got a lot of power they've taken a lot of power from the government and the government should be like hey hold on guys like you know prove yourselves first and you know you can have some power but not infinite power government's not doing anything about it um EU is is different but in in you know North America not so much um and I watched all of this happen and I thought kind of what the consequences were going to be and I realized we just could not count on the government I had this tension as I wrote this book um this crazy tension which was if the world went the way that I wanted to I would have entirely wasted my time writing the book and like anybody else I don't want to waste my time writing a book that nobody's going to read and is out of date and I had this fear as an author um that the book would be completely undermined because suddenly Washington would get its act together but of course it didn't um you know I would have been happy for the world and sadden for my book it was a weird position to be in sort of like betting against yourself or something I don't know um so so I was afraid maybe things would actually get be done right we wouldn't need the book I need not have had that worry for a second because Washington in fact mostly abdicated um abdicated uh Chuck Schumer in particular you know had the power to do something here as the Senate Majority Leader and put some strong legislation forward and he didn't he he took eight months of listening meetings and put out a white paper rather than an actual law so he kind of ran out the clock now as we record this um you know I guess you know no very little is going to happen because the elections soon and nobody you know the way Dynamics work in Washington like nobody wants to stick their neck out because it might hurt them in the election so like basically this period of great excitement about how we could handle this stuff that started around the time when I appeared in the Senate which was May 16th of of last year has entirely dissipated it's It's Gone With the Wind opportunity was completely squandered so the point of the book is we can't trust these companies to self-regulate they don't do the things that they promised they you would know better than me for example the things that they promised about pretesting to the UK government and then didn't deliver um I think there a big scandal in the UK people in in us may not know about it um but that's one example where they have not done you what they said they would do with self-regulation um and of course they'll you know weasle down anything that they said so that you know it's less invasive to what they're doing and so forth and then the big issue with governments is regulatory capture or doing just doing nothing and that's you know I read the writing on the wall and realized that nothing was going to happen and that really logically leaves only one possibility to get this right if we don't get it right it's going to be bad you it's going to be social media but much worse move fast and break things and so the only thing left is to directly appeal to the people and so that is what I am doing or trying to do we'll see if anybody cares but I I am out there talking about this stuff trying to get the citizens of the United States and some other nations to speak up loudly do things like boycotting for example and say look if if the government's not going to take care of the artists the government's not going to take care of the writers we're not going to use software that steals from artists and writers because we know we're next you know it's like Pastor Nemo or first they came for the for the Jews and gays um the companies want to take everything they really want the whole ball of wax they you know they're going to take your keystroke loggers whatever it is that you do if you do something that's on a computer and and they're going to try to replace you that is the game now and so if we don't stand up together with coordinated action and say look we want a more Equitable AI here um this doesn't mean Equity like everybody gets the same outcome but Equitable like everybody gets a fair chance and like if they use your IP you get some compensation and so forth and there's a million different aspects to this um if we don't stand up and say we want democracy to function we're not happy with the Deep fakes deep fakes is the one place maybe something will happen legally but if we as the people don't stand up and say this is really important and we don't do it soon and this is really important we're going to be screwed just the way that we were with social media but possibly worse so the problem with social media is things got entrenched and we can't fix them now I mean yeah the the the child act just just passed but by and large like social media is what it is now there's nothing we can do about it and it's not good you know it it's probably a net drain on society it's fun I use it but you know um it has a lot of problems and AI we're just giving so much power to these companies and if we don't set the right precedence in the next I don't know 12 24 36 months or something like that we are going to be stuck with whatever comes up which is probably going to Bean basically anything goes like we have section 230 says that these companies are not liable for anything on social media basically I think we could go into the story but was maybe well intentioned but the world changed the law did not keep up with the difference between um being a carrier of information and being a company that prioritizes social media feeds and makes more money if it makes things more polarized like the laws didn't keep up in their bad laws we were going to be stuck in the same position and so the choices that we make as a society right now are going to have an effect for the next decade maybe the next Century I wrote this book to wake people up Gary it's an honor and a pleasure to have you on mlst I love the show I'm glad to be back wonderful so um you gave the keynote this morning at the AGI conference and it was it was fabulous so by the time Folks at home watched this you would have seen the keynote and there there was about a 10-minute section towards the end which was absolutely hilarious but um yeah why didn't you tell me about that uh about the the talk as a whole or about the talk as I don't remember what was last um well look it was an interesting way of giving a talk I rewrote the entire thing um but I it was almost like something borrowed something new uh so the the context is I actually gave a keynote at this same conference three years ago and it's not that often that you go to conference twice in a three-year period and typically if you give a keynote like it's once every 10 years or something like that um and so I given a Kino three years ago and it's been such an interesting and yet such a disappointing 3 years in AI you know most people are excited about it I'm a bit disappointed and I wanted to explain why and so I thought about I looked at that old talk and I was like almost every word here is still true I thought about it a little more and I realized that there was something I missed before and so the talk was kind of divided into two parts one was all the stuff that really hadn't changed despite you know billions of dollars and enormous excitement enormous press and then the last part was what I really missed the first time around and that was interesting too so so the the first part was basically I went a few years ago and I said everybody's excited about these large language models this was before they were really big but they just started being called Foundation models everybody was excited and they said well what should a foundation be Ernie Davis and I said this together well a foundation should be like something robust that you can stand on that is what a you know a foundation of a house is or building and these models aren't that they make all kinds of dumb errors and and you can't really trust them that was the talk I gave a few years ago and I pointed out like why there was a lot at stake like you know telling Radiologists they shouldn't train anymore like hyping these things actually has a cost for society so I wrote that whole talk before and I looked at it I'm like this is all still true the examples of changed I don't know how many people will know Welcome Back Cotter the names have all changed since you hung around but it's all still basically the same that was about a high school the names have all changed little details on the errors but basically we still have unreliable AI you know since large language models came on the scene we have something that looks General but it's not that intelligent it's not that reliable you can't really count on it it's not nearly as reliable as a calculator is for example right I mean calculator you type in 3 * 17 and you get 51 and you're good to go on a large language model you never know what you're getting and that was true in 2021 when I gave this before and it was true you know today in 2024 and in fact because the the privilege of writing a talk yesterday is we wake up and you add another example um like these things are just they're just not trustworthy they're interesting but we see the same problems as before so that was the first part of the talk that was the larger part of the talk and it literally went like Slide by slide this was still true this is not true and most of it was true for fun I used like orange letters where something was new and not that much was new and then there was the part that I missed in 2021 I had a little hint of it but really not wasn't clear in my head it was clear in some other people's heads but the thing I missed then in my kind of critique of large language models was excuse me the thing that I missed in my critique in 2021 was what was going to happen to society and to the tech industry um more cynical people than I may have seen it coming I didn't quite um I grew up in the kind of era of Google um I mean I wasn't let me say that again I I started thinking about the tech World a lot in the age of Google and Google had its problems with surveillance capitalism but I think was genuine in saying don't do evil they really didn't want to be evil and I think the companies now really don't care they've put in all this money and all this chips and they need to make back the money and that's just driving so much in so many ways it's driving the hype it's driving decisions about copyright law and exploiting people and so forth and the last part of my talk was really about what I would call the moral decline of Silicon Valley which is actually the name of a chapter in in my new book um I think it's been precipitous I think there's really been a change like I don't see Steve Jobs being happy with what's going going on right now like he didn't build Apple to be this kind of company Apple still I think is not so much um but so many of these other companies it's all about the surveillance capitalism it's all about making as much money as possible it's about screwing artists which I think jobs would not have you know done I think jobs really cared about the artist and it's not that anybody wants to screw the artist but they're completely indifferent to it at some level right they'll make a licensing deal if they're forced to it but it's not like they want to you have these people talking about Universal basic income and yet they don't want to pay artists and nickel if they don't have to the courts force them to they'll pay the nickel but they're really trying to not pay the artist not pay the writers and so forth um you have a lot of people that I think are really just in it for the money and don't really care about society and that's having consequence and so the last part of the talk was like really what do we do about that like can we trust these companies to self-regulate no we can't um the book taming Silicon Valley is is also about the fact that we can't really trust governments to do the right thing either the governments are lobbied constantly by these companies there's so much money behind the scenes and so like here in the United States hardly anything has happened you probably know I testified in the Senate a year ago and that was kind of one of the highlights of my life it was amazing it was this historic moment Sam Altman was there the Senate was there it was the first time the senate had a um a full hearing on on artificial intelligence policy and I it really like I remember walking to the capital the night before and seeing it um you know Twilight it just kind of blew me away and it was amazing to be there and it was also amazing because all the Senators seemed to understand the urgency of the moment how important it was that we regulate AI in a right way not too strong not too weak um that we get it right now they all realize that we had screwed up with social media really bad that they had screwed up with social media they were incredibly humble I it was amazing watching the Senators Who as a lot of people said to me afterwards were on their best behavior they really seemed to get it and I was so excited and I've been so disillusioned ever since because you know that's over a year ago and nothing has passed the Senate hasn't even voted on any you know serious AI regulations a little bit that's coming up soon but by and large like nothing has happened on that though you you've spoken about the apparent um you called it the Messiah myth of of Silicon Valley and is is that what's going on I mean the open Ai and anthropic they they seem to be doing a lot of things around safety is that just theater to a large well mean it's hard to get into other people's heads but I would say that there's more theater than not that you have companies like open Ai and anthropic publicly saying they're for AI regulation and then they're out there trying to weaken whatever regulation is proposed um you know they all tried to block SP 1047 and and ultimately anthropic apparently was um instrumental in weakening it a good bit at the last minute you know open AI was you know Sam Alman was telling the Senate while I was sitting next to him how important AI regulation is and behind the scenes Billy Parago reported this in Time Magazine the the lobbyists for open AI were trying to weaken the eui K act and probably succeeded some so there's definitely like a two-sidedness to it where there there's public statement that they're supportive and then in private you know they really aren't why is there such a Divergence between the perception of this technology and the capability and I often look at people's Twitter just before an interview and uh you posted a beautiful example which um reminded me of an old example of of yours from a couple of years ago um it's an astronaut riding on a horse but of course it wasn't supposed to be that was it yeah so I I actually almost went epiplectic when I tried that one so so so I wrote a whole paper a whole substack essay um called horse rides astronaut and it was really a riff on something that goes back probably to Chomsky but I kind of knew it through Pinker he had this old example of um man bites dog as opposed to dog bites man so it's not really news it's a you know J journalists I think have used this expression right it's not news if a dog bites a man but it is news of a man by a do so I kind of riffed on that when do 2 came out so you might remember when do 2 came out it was a big deal it was the first of these really good image generation things 20 minutes after it came out or maybe was an hour after it came out Sam Alman posted AGI is going to be wild and a lot of people thought wow this is like the AGI moment and I looked at this stuff and I realized it's not there nothing to do with AGI it's really nice Graphics the these systems reconstructing images in a really interesting and Powerful way but they don't really understand language I did some experiments with Ernie Davis and Scott arenson and then later with Aina uh levada and and Elliot Murphy showing that these systems don't really understand the compositionality of language which is to say that language is made up of Parts you put them together in larger holes and you're mapping a syntax onto a semantics and you're deriving it from there um fragga is the philosopher the we most associate with that concept um so philosophers have been thinking about this for a really long time and uh formal semanticist people like that in linguistics um have thought about it a bunch and it was clear playing with dolly for a few minutes really that he didn't really understand compositionality in fact in linguistics uh computational Linguistics there's an old idea of a bag of words model and with a bag of words words model is is you just take the words in a sentence and you scramble them up you as if they were in a bag and like how much can you explain with that or whatever and a bag of words model is almost like a control group it's not a very good control group um you know if you can't do better than a bag of words then it says you don't really understand the structure of the sentence the way that the meaning relates to the parts of the words and I could tell playing with Dolly you know even for a few minutes it had a lot of that flavor it wasn't literally that but it was more like that than a system that really understood the the components of the sentence and so I thought about that in pinker's Old example so I wrote this substack essay called horse rides astronaut and it was riffing off the old man bites dog example so you know we have lots of astronauts riding horses but we don't have many horses riding astronauts and I showed in this substack essay called horse rides astronaut that do tended to have problems with it that if you said it's hard to even say it right if if you said horse rides astronaut it would tend to give you the more canonical astronaut rides horse and then I went through all of the kind of stupid or not stupid that's not the right way to say all the kinds of defensive objections the people who love uh this kind of AI would make and they would say well that's because the system has enough common sense to know this is impossible and yada yada and what I showed is actually if you prompted it the right way it could actually do this so it wasn't that it couldn't draw the graphics of a horse on top of an astronaut and it wasn't because it thought it was because the system thought it was literally impossible but just it didn't understand the relation between the words in that sentence and what it was supposed to do which unfortunately is it it's one job you know the hashtag on Twitter you had one job your one job of your dolly is to understand the meaning of the sentence and draw the picture and it couldn't do it for these kinds of cases and then I showed later another example I wrote a whole another essay about I can't remember the title um but I showed an example the um NPR covered where one of these systems couldn't get um a black doctor with white children as patients because it wasn't canonical as some of these things get fixed up some of the time people train on more data or whatever that particular one I think got fixed up but today I was working with grock on literally in the cab on the way over in the Uber on the way over I was like I should try that one and see if it's any better I tried a bunch of other things and generally the kind of like challenges that I give grock was not doing so well we could talk about some of the others but so I did horse rides astronaut and I got it wrong I'm like there's so much discussion about this one this was a popular essay and people came at me and lots of different ways and so for and still like you know this allegedly state-of-the-art system managed at least on the first try I didn't try multiple times um managed to get that wrong and I was just like we are back in 2020 to everybody has been saying for the last two years you know especially these influencers saying every day they're like look at this new amazing thing that came out we live in this time of exponential you know Bounty and and whatever but on the things that count on the things that matter at least from my perspective as a cognitive scientist who's spent his career studying Intelligence on the things that matter we really have not made that much progress but Gary I thought these things learned abstract world models and the reason this is interesting is that there's a there's a dichotomy between um Al alteric um uncertainty and epistemic uncertainty you know there's there's a difference between actually understanding something and being able to reason being able to do this deductive closure to deduce new knowledge about the world well there's a couple of different things in that um to unpack so one is is this a question about uncertainty and it isn't really a question about uncertainty so I mean you can have something that's very clear like horse rides astronaut there is no uncertainty the horse is supposed to be on top of the astronaut um so there are all kinds of interesting things about alator versus epistemic uncertainty They Don't Really apply here um you know this is a perfectly deterministic phrase and it's just getting it wrong now there is probabilistic uh outputs in these systems so they're the systems themselves are not deterministic and if I ran that same prompt 10 times I might get 10 different answers and maybe six of them would be right and four of them would be wrong um you know we're the other way around or who knows uh so there's that kind of uncertainty but the fundamental is you are supposed to map your semantics onto this description of the world you asked about world models they don't really have World models I think that that's actually easier to see though in Sora because Sora has changed over time and World models are in part about understanding the Dynamics of the world that's really why you want to have a world model um and you know humans have World models so I like I have a model of the room that we're in where there are lights it might not be perfectly specific I might not know where all the lights are um and I can do updates so there's somebody else in the room if that somebody else in the room um suddenly says fire then you know we're going to do something different and maybe stop having this interesting conversation and so I have a model of all the things are going on or many of the things are going on and all humans do that all the time if you watch a movie you make a model of the characters or I'm watching the bear so I have um I haven't quite finish catching up and so I have a model of you know the lead chef and the person who works with them and um his girlfriend who's maybe not his girlfriend anymore and I'm watching all that stuff and and things will unfold and I'm maybe trying to put flashbacks and try to order the sequence so I just saw a wonderful episode about um how somebody first got her job and I don't want to give away too much but um working at the restaurant and like for the first 10 or 15 minutes you're sitting there how do I relate this thing that I'm watching is this in the now of the film or is this before and eventually we realize it's a flashback and it's a kind of origin story it's a really beautiful story um uh and and sad and powerful and so and I'm sitting there trying to make a mental model of how this piece of this narrative fits in with this other piece and what kind of person like I've seen this character before but now I see much more I'm learning more about her I'm more learning more about her partner and the relationship I'm building a model of all this stuff and current systems just don't really do that what they do is they build a model of the words and maybe some other stuff that have been said in the sentence so far and try to guess what would happen but they don't have like the equivalent of index cards if you remember those where you like like write down notes or databases where you you know have records you know this is your your phone number and this is your address and so forth they just don't have that um people don't understand that they also don't understand that these systems don't do sanity checking that they don't look look things up in encyclopedia or Wikipedia whatever they just don't have meth models world world models cognitive models or what have you um when when they're trying to do horse rides astronaut it's not like they have a world model of what horses usually do and what astronauts do they have a bunch of pictures and their their pictures are kind of clustered in space and they're going to some cluster trying to find the nearest cluster I'm oversimplifying a little bit but they're going to this cluster of words that have been around this thing before it's not the same thing as as understanding so back to Sora which I think is actually a better example um so with Sora you see frequently weird things happen like for example there was one where people are carrying some stuff and one guy goes behind another and then the camera moves and the guy is just gone right so if you have a model of the world then you know you know which people are there there's another one where there's like four dogs and then the angle changes the camera and then there's suddenly three dogs and then the camera changes and there's five dogs so like people would find that weird there are circumstances where people would miss it but fundamentally you can see that Sora does not really have the notion of object permanence which is one of the basic things that I believe that we're born with based on Liz spelly and Renee bayan's cognitive development work and so forth um a lot of people know the old P stuff saying that only 8 month old do you learn object permits that's been shot out of the water by much better experiments using more sensitive and so forth my best guess is that stuff is in eight but Sora never gets it um and it never gets lots of other stuff too like it doesn't learn that a chessboard is 8 by eight it sees a bunch of chess boards but it might draw a 7 by seven chessboard it doesn't understand that there are conventions there was a a SORA ant um Yan Lon and I both posted about this him an hour after I did this ant that had four legs right like it seen God knows how many ants and it still doesn't understand how many legs and ants typically has um so there are lots of things that an ordinary person's model of the world would have and these systems just don't have it but it really comes out in the changes over time where just a bunch of impossible things happen and they only make sense in terms of the statistics of pixels rather than the statistics of the world so they make sense from the statistics of pixels from this Frame that the next frame might you know not have anything behind it or whatever um you know every pairwise bit of pixels one frame to the next sort of makes sense if you don't understand what is an image of or another example is there's a SORA video where somebody goes into a building it's a fly through and it is amazing the first time you watch it um it's like a museum but the second time you watch it you realize that like the outside and the inside don't actually match so pixel by pixel everything in that panning shot makes sense but if you go across you know it's only like a minute film if you go across the minute there's enormous an enormous number of really massive discontinuities that make no sense whatsoever except on the you know frame by frame from this Frame you could see a frame like that but if you look at frame one and how you got to frame 100 it don't makes any sense at all and that's cuz there's no stable World model there's not a stable model of like what the dimensions are of the building and so you wind up I forget what it is I think it's that like the exterior shot you know is something small and the museum is massively bigger there's like an empty Courtyard but then it's not empty anymore so there's all these inconsistencies yeah I mean this is where I was going with the epistemic risk because we can verify so so we have a world model we have facts about the world and we can verify and sometimes that's a binary we can we can say in certain situations that an astronaut is on a horse or a horse is on astronaut there is some vagueness around the boundaries perhaps but it but but it's binary but I'm interested in why we anthropomorphize these models that they're are kind of adversarial attack on our perception in in many ways take the classic Turing test what it really turned out to be was a measure of human gullibility so passing the Turing test doesn't actually mean you're intelligent it means you can fool humans it turns out we're very easily fooled um the most dangerous version of this is you can see a few minutes of a driverless car and conclude that he drives basically like a person that everything is good and in fact that driverless car may have a lot of serious problems in a lot of different contexts um and so something that superficially looks like a person for a few minutes may not actually be like a person and so what has happened with large language models is they superficially produce humanlike output and people are willing some people not all are willing to cut them some slack when they make an error once in a while but they actually attribute intelligence to these systems in a form they don't have it and one of the ways that they do that is they attribute intelligence they think it's like me it would do the things I would do in in such and such context and it doesn't so you know one example I used in my TED Talk was U the um the Lost Galactica um the the meta system that was pulled that Yan Lon is still bitter about um Galactica uh said that Elon Musk died in a carc the sentence was in March of 2018 Elon Musk uh was involved in a fatal uh automobile Collision I think it was and then it continues on and it's clear that it thinks that musk in fact died um any person would say oh wait a minute in 2018 I mean if there is I mean we we think about data we're not perfect with thinking about data but an Evidence but if there's evidence that any human being is alive right now it is Elon Musk because he is on X every day he's in the news every day there the amount of evidence that Elon Musk is alive is greater than for any other person on the planet and so this is a obviously false assumption you could also go to Wikipedia and so forth so if I was an editor of a you know newspaper and somebody gives me Elon mus died in a car crash in 2018 i' probably fire them um be like 2018 that just does not make sense it does not check out and even if you said now I'd be like well can we get some extra sources and you know if you said that he died today um you know i' want some sources can we get confirmation on that we don't want to run it yet um and so you or me or certainly any you know editor of a newspaper something like that is going to fact check things especially ones that seem you know Prima facially to be implausible but llms don't do that and it's very hard for the average person to realize that because they see this small sample of data and in the environment in which our brains evolve the kind of evolutionary history we didn't have this problem of you know impostor chat pretending to be people we had other problems like that lion is it going to eat me and so we're pretty good at looking at motion is that thing coming close to me or not how big is it you know we make a lot of judgments about the world that are really good but we don't make judgments about AI that are really good unless we took cognitive science classes in college or something like that which most people didn't and so we see this tiny sample and we wildly overgeneralize because it said a few sentences and it like types the word out one by one which was a stroke of Genius by open AI it gives this illusion of being personlike it's not actually personlike at all it's a statistical you know autocomplete on steroids that is trained on a lot of data to look like a person but it is not reasoning like a person ever I mean there there's two other things here as well first of all there's the phenomenon that people want to believe that it's intelligent um especially when publishing newps papers that's what sabaro said to me the other day that you don't want you want people to think the llm did it because then you get a Europe's paper but also um I read a blog post by Nicholas khini from Google and he said he used to be in the camp that thought that llms are just databases and he started playing chess not databases either but we come back to we'll get we'll get to that we get to that he started playing chess with gp4 and he was he was blown away with um the sophistication of of the moves and it was often making correct moves but he said something quite interesting which is that if he played chess like a bad chess player it would reflect the bad behavior back and it's the same thing with generative code if you write code with sophistication it gives you better code back because it's almost adopting a role player well it is a mimic just as sidebar on chess gbd4 doesn't really play chess that well so um it makes a lot of illegal moves there there's a very good I I'll try to give you a link to put in the show um the very good I'm blanking on the guy's name um has done a very good analysis of of GPT um for in chess playing and it plays like I don't know like A600 game which is like you know better than the average person in your high school but not anything like world class and it makes a lot of illegal moves um like 6% of the time or something you know some crazy number like that which no you know the chess compare with a chess computer that I bought in 1979 where you could stick the little pieces in the hole I think it was called Sargon um was was the software underlying that never made an illegal move never ever right we're talking about Chess software from I mean really even further back 1969 never made illegal moves like dpt4 is not not you know state-ofthe-art and chess and we we should understand that but people don't they're they're amazed that it can do it at all and there's some reason why you should be amazed that can do it at all but you have to realize that it is not you know searching a tree the way that a proper chess program can do it is doing mimicry to come back to the other part of your question and you do get these like weird mimicry effects because essentially what you're doing is a little bit like what humans do when they're priming you're directing the system to a particular part of its Corpus so you can direct it towards the more sophisticated language or the less sophisticated language or whatever and it's going to try to replicate language like that and so that is why you get some of those um kinds of effects I don't know whether um what Carini described you know really Bears out in a systematic study but if it did that would be my guess for you know why it would is like the database of lousy chess games is going to look different from the database um of good chess games yeah it's strange isn't it though as we memorize more of the long tail um the the reliability sort of goes down a little bit so around 5% 4% and a lot of people argue that that that's just fine we can engineer our way out of it do do you think that it depends on what the problem is the first thing I would say so um large language models are not like calculators calculators give you 100% correct answer and large language models in very few domains give you 100% correct um in some domains they're just completely outmatched so you know floating Point arithmetic they're going to be lot less than 95% correct um especially with large digits they're going to be much worse I would suppose I don't know if anybody's done exactly that study um and something that matters probably in all forms of AI but particularly in large language models is the cost of error so uh large language models are best suited I would say to things that are like brainstorming or autocomplete so coding is a kind of autocomplete where the coder is still at the wheel right it's not a fully autonomous SE ity the system is not actually writing the whole code or whatever so it's writing little bits you drop in and coders have spent their entire lives learning how to debug bad code partly because most coders don't type that well um and also because coders forget things and whatever I I've done a bunch of coding in my life and I know how it goes and so like if you don't know how to debug things you're just not going to become a coder like it's just not the profession for you so everybody using those tools can tolerate a certain certain percentage of error and they're trading off how long does it type take me to type this out to look it up versus how long to debug it I think people are initially excited some people are less excited now because they realize sometimes like they did a bunch of tests to make sure that the code works and up passed all of them um they wrote the code in an hour and then three days later they don't remember why this code is there because they didn't actually write it themselves and it takes them like 24 hours to debug it and then they're like I don't know if this trade-off was worth it or not so there's there's still some open questions there about security and so forth there's been um some academic literature suggesting their problems but at least in principle you have a coder who is picking you know the outputs they're taking some not others they're they're they're fixing it and so you can have a high amount of error but it might still be worth it um there are other domains where also a high amount of error might be worth it like brainstorming so I'm trying to think of a commercial and it gives me 30 ideas I reject 29 but I'm happy that I got the one and so um you know I that might be great right so it could be 90% error rate but you're still happy and then there are domains where like any error is probably going to like kill somebody it's like a medical domain and you know you have the system treating a child like an adult and giving the wrong answers because it's not really trained enough pediatric data and like any error might like actually like kill a child or send in the hospital or whatever and so you need to be much more accurate so there is no blanket statement you can say about like what percent error matters if you want to use a thing as a calculator probably you shouldn't get anything wrong like you should just use a calculator um so it does depend what you're applying it to yeah coding is an interesting example because there's a verification step so does the code compile and then there's the behavior does does it does it run the way yeah I'm going to pause you right there though the worst stuff is going to come from people who think that's the only verification step I'm not saying you think that but I think some people do so for those who are not programmers um in a language like C for example um C++ you write the codee you compile it which turns it into machine language but that does not prove that there are no bugs but a a beginning programmer actually might labor under that assumption so they think if it compiles I'm good to go now a sophisticated program realizes that's not the case at all that bugs can emerge in code that does correctly compile but there's still some assumption that's been made that's wrong and so forth but you know there's a certain amount of bad code that is seeping into the code base because people do think that's the only verification step another verification step that good coders know about um would be unit test or some some kind of testing like I now that it compiles that the journey has just begun I'm going to make sure that this code actually does what I want it to do and a good programmer understands the logic of what they're programming and they know what a good test might be they know what wacky user input might come they want to make sure that it handles that input they understand you know when the circumstance of their assumptions might be tested they test that um and a bad programmer doesn't really do that they do a couple tests and they call it a day they say this is good they will go on to the next thing and then it all falls apart three weeks later when one of those assumptions was violated it's very true I mean with Gen coding I've been doing a lot with Claude 3.