🎙️AI 访谈库
走进 NotebookLM:Raiza Martin 与 Steven Johnson
Raiza Martin · Google Labs 高级产品经理(NotebookLM 产品负责人)

走进 NotebookLM:Raiza Martin 与 Steven Johnson

Inside NotebookLM with Raiza Martin and Steven Johnson

2024-11-26 · Google DeepMind: The Podcast (Hannah Fry) · 46m · 约 48 分钟读完 · 原文
与编辑总监 Steven Johnson 一起向 Hannah Fry 解析 NotebookLM 背后的技术进展:'什么才算有趣'与生成自然语音的难点。看点是用户创造力的规模:人们用简历、私人日记、销售材料等素材生成了成千上万档个性化播客。

welcome to Google Deep Mind the podcast I'm Professor Hannah fry now I want to start today unusually perhaps with a clip from another podcast I've listened to this what's the overall message here is it social commentary artistic expression or just a really elaborate joke that's the beauty of this piece I think it defies easy categorization it exists in this FAL space between language and non- language between art and absurdity this is a very interesting discussion which as you might have guessed is AI generated but what is notable about this particular clip aside from the fact that neither of the two podcast hosts have ever existed is that their conversation a mini treaties on human nature and our relationship with art was generated from the most unusual of prompts the podcast itself was created by a new feature called audio overview part of notebook LM a personalized AI research assistant from Google abbs Now notebook LM it's powered by Gemini and it lets you upload your sources anything from PDFs to videos to generate insights explanations and of course podcast we often think of AI as just crunching through data and spitting out answers but notebook LM draws on expertise from storytelling to present information in an engaging way and we wanted to see what happen s when you ask Nota LM to analyze what most people would consider to be nonsense a single document containing Just Two Words repeated a thousand times over cabbage and puddle and here is the result so I have to admit at first I was like what is going on here but uh the more I think about the curious I get you know it is fascinating isn't it we're like dealing with this one piece puzzle right and we're trying to figure out does this piece tell us what do you think what's your first impression honestly it's uh it's almost like hypnotic or something like if you were really staring into a puddle and all you saw were these gabag floating around I can see it it's a little unsettling but but also kind of funny several minutes of intellectual analysis packed to the brim with seemingly relevant ideas that are nowhere in the original document that quite impressive really well I am joined today by two people who are deep involved in writing notebook LM story joining us from San Francisco is Steven Johnson notebook lm's editorial director and also a New York Times best-selling author and in Mountain View California ryza Martin is a senior product manager for AI at Google Labs who leads the team behind notebook LM welcome to the podcast both of you okay now I I want to start with a feature that everybody's been talking about um this this audio overview and uh well I understand that you've got a little clip that you want to play me yes let's let's play the clip I think you will enjoy this Hannah okay here we go welcome back everyone ready for another Deep dive today we're shrinking down way down microscopic you might say exactly think about those tiny little droplets of water you know like the ones you see on a freshly washed car oh yeah but imagine those droplets clinging to an airplane wing or on a plant leaf right being sprayed with pesticides the way those droplets behave is actually incredibly important for all kinds of things making planes safer more efficient farming even figuring out how rain forms wow that's fascinating we're diving into some serious research today that was my PhD the first page of my PhD thesis extraordinary I mean frankly there is no good stuff apart from heavy equations in okay lots of things to notice about that for starters they made it sound much more exciting than it actually is that's the point but also though the're sort of back and forth I mean the two the two voices there are finishing each other's sentences it felt very fluid to pdon the pun um very very natural imagine Defending Your dissertation now you could just play the podcast and kind of leave it at that I think if you'd only had that at your disposal back then Riser have you been surprised by people's reaction to this because I mean it's had really quite serious uptake hasn't it yes and I think the most surprising thing to me and and really equally delightful is how people are using it I think I imagined how they might but I think the beautiful thing about launching something with this much sort of excitement around it is you see a whole new universe of what everybody has been trying from things that are funny things that are entertaining things that are inspiring or really meaningful it's just been incredible I actually probably spent a good chunk of my day a third of my day just listening to these really you set up a Discord server didn't you just to to let people share uh stories about the ways that they're using it what what kind of things have have come up so I mean that was an interesting example playing your dissertation because one of the things that I think genuinely surprised us is people would put their CVS and their resumés in there and it was almost like a little like hype machine like if you were feeling down about yourself you would you would listen to like a 10-minute audio conversation between two very enthusiastic hosts you're like wow stepen has really done a lot in his career it's very impressive but actually a more serious version of that I mean that's kind of fun and and playful but um people are using it like you can kind of Workshop things you're you're working on so you can upload a short story you're working on and say Hey you know give me some constructive criticism on this and you get you know you listen to people talking about your work and they're very good at pulling out the kind of interesting twists or focusing on the characters that are uh particularly compelling or not um and so it's a way of getting a little kind of like it's almost like a little Focus Group group for stuff that you're working on which is which is really amazing I guess also hearing people actually talk about it out loud adds that kind of extra layer of I don't know objectivity of whyer I would say um it's been really surprising because if we think about it a lot of the content or content generation if you just render it in text is not new right it's like if I upload my CV and then I have an llm spit out something that says like oh here's R his career right a summary of sorts maybe there's a few interesting tidbits that it pulls out here and there that was novel two years ago everybody was excited by that but I think adding that new layer or that new modality of just very humanlike voices I think it connects with people in a very different way right I think like personally like I call this I call this type of Technology humanik where you sort of recognize it as being very similar to you and it resonates with you in a different way as a result and I think the first time I listened to my CV I knew what to expect but when I heard it it I still felt that like that bubble inside of me like that and I think that's the magic of of new modalities I think that you know the other point on this is that like human beings have been learning and exchanging information through conversation for hundreds of thousands of years we've been learning by reading structured text on a page for you know 500 years and structured text on a screen for you know 30 years um and so when you activate that sense of like a a genuine humanlike conversation um it it's just a deep ancient kind of ancestral part of who we are um that I think that's you know that's one of the reasons why just it lights up people when they when they hear it for the first time also interesting I think that you you decided to have two hosts rather than just one person sort of talking into space as it were which I guess it speaks to the point that you're making see yeah it's just a very different format if you just have one person it feels like text to speech right we've heard text to speech before you're just like the computer is turning the text that it just wrote Into You Know into something I can listen to which is great and you know we we're interested in trying to figure out ways we can do that in other formats but to get the conversation right and we can dive into this in more detail like there all these like subtle things that you have to like make work nobody wants to listen to two robots talk to each other like that it that will fail um and be unlistenable like after 30 seconds you have to master all these very subtle weird things that people do in conversation PR it to work to make it human like exactly as you said right I I want to come back to the to th those features a little bit later to to the audio overview um because I also wanted to discuss the origins of this of notebook LM um how did it come about Risa for one I think a lot of people think that notebook LM is new because of the audio overview feature we had such a massive influx of people and people were like wow what is this a brand new thing from Google but actually we've been working on Notebook LM for over a year we first announced it at Google IO last year as project tailwind and before then we actually had been incubating it inside of Google labs and um it's actually how Stephen and I met Stephen was brought in what was your original title Stephen I was visiting scholar yeah yes then I then I became editorial director that's right he was promoted and uh at the time Josh Woodward who now leads uh Google Labs he's the vice president uh told me he's like I want you to build a new AI business and I I thought to myself what does it take to actually do that but what I'll say is one of my early Inspirations was just watching Stephen work honestly just like understanding how he does what he does I was like wow that could be a real superpower if you could give that to people it was a mix of Stephen is abnormal uh in his his research abits but maybe we could turn this into a main dream Pursuit somehow yeah it was interesting cuz we we I had had this um long history writing books and Josh had read some of those books and had read some things I was writing about Tools For Thought basically like how do you use software to help you think and help you develop your ideas and research this is middle of 2022 so language models were at the top of the list then and so he kind of reached out to me and said hey any chance you would want to come to Google and help build the tool that you have always wanted to help people learn and and organize their ideas now built on top of language models and and what rise and I kind of like right from the beginning like I think I met rise like day two at at Google we were like let's build something new this came about at a time when large language models were at the top of the agenda in those early conversations how did you see this as being fundamentally different to just I don't know like uploading a document on Gemini and getting it to summarize it for you from the very beginning we we call it Source grounding that's the way we describe it like you supply The Source information that you want to work with and might be the story you're writing it might be the book you're researching it's it might be your journals it might be the marketing documents you're working on and uploading that to the model then creates a kind of personalized AI that is an expert in the information that you care about and that was not no one was talking about that in in the middle of 2022 so that was like the first thing we built was like I mean we like we uploaded part of like one of my books and I could like have this very crude conversation with the model that was not at all like what you see now in Tex store with audio but you you could get a little taste of like what it would be like to have all the ideas you were working with instead of just talking to an open-ended model that just had its general knowledge actually have that personalized knowledge and it was great because it also like reduced hallucinations it made it more factual you could fact check it um you could go back and see the original Source material that's a big part of the whole notebook LM experience um that was the beginning of it and everything we've done is built on that platform and audio overviews is just okay take that Insight of I Supply my sources and now I turn it into something else in this case it's a audio conversation because I guess the real key difference here is that it's very focused on the sources that you're giving it and and and anything that's connected to that rather than just as you say this General model yeah I think that uh I'll say too that what we've seen is I think it's a little bit harder to get started with this Paradigm because it's it's so new right the idea that one you're talking to an AI two you have to bring your own stuff so I think there's a little bit of a layer where it's like okay you have to convince somebody that it's worth doing but once you can get somebody over that hump it's just massively useful because you know I think about the work that I do every day the work Stephen does every day and many people around the world that work on computers every day we are working with very specific sets of information shared context that we have with others right like we do research we pull it in we want to sort of extract our own insights from it I think that's a that's what makes notebook LM really special and has made it's metal from the beginning so it does include these text elements too then because as you say the podcast part is is the bit that's sort of most most notable that's right so the podcast thing is the most recent development in Notebook LM but we actually launched a year ago where it was primarily a chat feature so you're chatting with the system using your sources and it's always referencing back to exactly what pieces of your content that it used so give me some more mundane examples of how people are using this like on a day-to-day level then stepen yeah so I mean we actually see a huge amount of usage of the product just with the text features right and suddenly you you have this like amazing like resource that can answer any question about all you know hundreds of pages of documents um and in the text version you get citations and everything it's a very scholarly thing actually an you would appreciate it like you get your answers back and every fact that the model says has a little inline footnote and you can click directly on that footnote and go and read the original passage writers journalists obviously are using it this comes a little bit out of my like my involvement for the project I have like one notebook that has thousands and thousands of quotes from books that I've read over the years plus a lot of the text of books that I've written and that's that notebook has basically like my brain kind of captured in the AI and so whenever I work on anything that's just have a new idea for something I'll go into that notebook and be like hey what do you think about this idea and the notebook you know the AI will say Hey Stephen you read something related to that like seven years ago what about this passage and so it's a true like extension of my memory um so so that kind of stuff and the other thing last thing I'll say is like we're not training the model on this information so you your your information is secure it's private it's not going to get into the general knowledge of the model and be used by somebody else so so you can put private information in there and when you put you know a couple of years of your Journal um in a in a large context model like this um you can get these amazing insights and you can turn them into audio overviews and kind of listen to two people talk about yourself or you can just be like what was I thinking about like last may you know give me an overview of like all the stuff that was going on and this you 20 seconds later you'll have this amazing kind of document of of your own life rather than just recalling stuff though can it actually be insightful in terms of your own journals I would say yes because I've used it for that purpose and um one of the things that I like to ask it after uploading I do these weekly journals is I say you know how much have I changed over time and it's really remarkable um it's been able to pull out for me really interesting nuances that I haven't been able to observe about myself um it's been able to say things like hey you know you tend to associate a lot of negativity with this particular topic you associate a lot of positivity with this topic and it's just really interesting because I think de earlier question around the mundane mundane use cases I think we see a lot more of those which is like just people trying to take the work that they're doing every day for example like sales teams use this a lot to share knowledge with each other makes a lot of sense there's a lot of technical complex changing documentation so it's really nice to have an AI partner I think that's really different from how a lot of AI systems work today right like you know I use everything I use everything that's out there and the prompts that I write are massive right like the first thing that I write is you are a blah this is what we are doing here are the documents that are relevant and I think for Notebook LM it sort of just shortcuts it it's just a project space it knows what you're talking about you can have a conversation forever it takes up to 25 million words it's just sort of contextually quite massive I think one of the things that was interesting and maybe a little bit distinctive about it was so many of the questions about like what makes this product work or not work are not so much technological questions as they are editorial stylistic questions like what is the right kind of answer when when you get an audio overview that works you know what what what's the style what's the house style for those conversations what's what level should they be pitched out and those are not technological questions those are those are language questions and that's that's that's the crazy reality of of the language model age is that all these things that you know used to be just mostly a question of kind of like let's get the programming right now become more about the rhetoric of it all well I do actually I want to dig into um some of the house style a little bit more I guess why did you decide to go into audio overview what what was it that inspired that I mean there are already quite a lot of podcasts let's f wa audio overviews really began was is a great example of the lab's structure I think really working well um because it was it was another small team inside of labs that were just kind of focused on the audio version of this and uh we had and and part of the idea of it was was not so much to like compete with podcasts but rather that there was a whole universe of content that you would never the economics of generating a podcast for it would never make any sense but um if you could generate One automatically you might have you know five people that would want to listen to it or like one person who would want to listen to it or 20 people but not you know 200,000 and so that's like you know we want to create a podcast based on our like team meetings from the last week so we can review them like that's not going to be a Comm business like no one's going to ask you to host that Hanah but but actually might be useful for that team and so they had started developing this thing and and rise and I heard it I don't know in probably in like March or April of this year and as everyone who's heard an audio overview initially were just like wow what what did I just hear that was amazing but we realized like pretty early on that part of our mission with notebook LM was was to build a tool that helps people understand things suddenly we were like oh wait people really understand and remember and you know pay attention when they hear something in the form of a of an engaging conversation between two smart people we released it internally to googlers over the summer and that was I think when we started to think this is going to be a hit like because you could just see the Delight that people had with it um so while we were surprised that it was went quite as crazy as it did um we knew we knew we were on to something now I remember in the last season we got to hear a demo of waver which of course is one of the first AI models to generate this humanlike speech and I mean it was quite impressive back then but I mean presumably there have been technological advancements that have happened since that have been necessary to make something like audio overview possible I think there's you know the the underlying model for Notebook LM is Gemini 1.