🎙️AI 访谈库
BG2 第 17 期:欢迎黄仁勋
Jensen Huang · NVIDIA

BG2 第 17 期:欢迎黄仁勋

Ep17. Welcome Jensen Huang | BG2 w/ Bill Gurley & Brad Gerstner

2024-10-13 · BG2 Pod (Brad Gerstner & Clark Tang) · 1h21m · 约 74 分钟读完 · 原文
谈通往 AGI 的智能扩展、推理与训练并重、NVIDIA 护城河、xAI 孟菲斯超级集群与开源闭源之争。

what they achieved is is singular never been done before just to put in perspective 100,000 gpus that's you know easily the fastest supercomputer on the planet as one cluster um a supercomputer uh that you would build would take normally three years to plan right and then they deliver the equipment and it takes one year to get it all working yes we're talking about 19 days Jensen nice glasses hey yeah you too it's great to be with you yeah I got my ugly glasses on just like you come on those aren't ugly these pretty good do you like the red ones better there's something only your family could love well it's Friday October 4th we're at the Nvidia headquarters just down the street from imer welcome thank you thank you and we have our investor meeting our annual investor meeting on Monday where we're going to debate all the consequences of AI how fast we're scaling intelligence and I couldn't think of anybody better really to kick it off with than you apprciate um as both a shareholder as a thought partner kicking ideas back and forth you really make us smarter um and we're just grateful for the Friendship so thanks for being here happy to be here you know this year the the the theme is scaling intelligence to AGI and it's pretty mindboggling that when we did this two years ago we did it on the age of AI and that was 2 months before chat gbt and to think about all this change so I thought we would kick it off with a thought experiment and maybe a prediction y if I colloquially think of AGI as that personal assistant in my pocket if I think of AGI as that colloquial assistant get used to it exactly yeah you know that knows everything about me MH um that has perfect memory of me that can communicate with me um they can book a hotel for me or maybe book a doctor's appointment for me when you look at the rate of change in the world today when do you think we're going to have that personal assistant in our pocket soon in some form yeah yeah soon in some form and that that uh uh that assistant will will get better over time that's the be of Technology as we know it and so so I think in the beginning it'll it'll be uh uh quite useful uh but not perfect and then it gets more and more perfect over time like all technology when we look at the rate of change I think Elon has said the only thing that really matters is rate of change it sure it sure feels to us like the rate of change has accelerated dramatically is the fastest rate of change we've ever seen on these questions because we've been around the rim Like You on on AI for a decade now you you even longer is this the fastest rate of change you've seen in your career it is because we've reinvented Computing you know a lot of this is happening because we we drove the marginal cost of computing down by 100,000x over the course of 10 years Mo's law would have been about 100x yeah and and we did it we did it in several ways we did it by one introducing accelerated Computing taking taking what is work that is uh not very not very effective on CPUs and put it on top of gpus we did it by uh inventing new numerical precisions we did it by new architectures inventing the tensor core uh the way systems are for are formulated mvlink um added uh insanely insane insanely fast memories hbm and uh uh scaling things up with uh mvlink and infiniband and working across the entire stack right basically everything that that I describe about how Nvidia does things uh led to a super Moes law rate of innovation now the thing that's that's really amazing is that that as a result of that we went from uh human programming to machine learning and the amazing thing about machine learning is that machine learning can learn pretty fast right as it turns out and so as we as we reformulated the way we distribute Computing you know uh we we did a lot of parallelism of all kinds right tensor parallelism pipeline parallelism parallelism of all kinds and and uh uh we became good at good at uh uh uh inventing new algorithms on top of that and uh new training methods and and all of this Tech all of this invention is compounding on top of each other exctly as a result right and back in the old days if you look at the way uh Mo's law was working the software was static right it was pre it was pre-compiled as shrink wrapped put into a store it was static and the hardware underneath was growing at Mo's law rate right now we've got the whole stack growing right innovating across the whole stack and so I think that that's the now now all of a sudden we're seeing scaling that is that is extraordinary of course but but um we used to talk about pre-trained models and scaling at that level and how we're doubling the model size and doubling therefore appropriately doubling the data size and as a result the Computing capacity necessary is uh increasing by factor of four of a year right that was a big deal right but now we're seeing scaling with post trining and we're seeing scaling at inference isn't that right and so people used to think that pre-training was was hard and inference was easy now everything is hard right which is kind of sensible you know the idea that that all of all of human thinking is is one shot is kind of ridiculous and so there's there must be a concept of fast thinking and slow thinking and and reasoning and reflection and iteration and simulation and all that um and that now it's coming in yeah I think to that point you know one of the most misunderstood things about Nvidia is How deep the true Nvidia Moe is right I think there's a notion out there that you know if some as soon as someone invents a new chip a better chip that you know that they've won but the truth is you've been spending the past decade building the full stack from the GPU to the CPU to the networking and especially the software and libraries that enable application n so I think you you know to that but you know when you think about nvidia's today yeah right do you think nvidia's mode today is greater or or or smaller than it was three to four years ago well I I appreciate you you recognizing how Computing has changed in fact the reason why people thought and many still do that you designed a better chip it has more flops has more flips and flops and bits and bites you know what I'm saying and you you see that you see their keynote slides it's got all these flips and flops and you know bar charts and things like that and and that's all good I mean look horsepower does matter yes so these things fundamentally do matter however however um unfortunately that's old thinking it is old thinking in the sense that the software was was uh uh some application running on Windows and the soft software static right which means that that the best way for you to improve the system is just making faster and faster ships but we we realized that that machine learning is not human programming machine learning is not about just the software it's about the entire data pipeline it's about in fact the flywheel of machine learning is the most important thing so how do you think about uh enabling this flywheel on one hand uh and enabling data scientists and researchers to be productive in this flywheel and that flywheel is is uh starts in the very very beginning a lot of people don't even realize that it takes AI to curate data to teach an AI and that AI alone is pretty complicated yeah and is that AI itself is improving is it also accelerating you know again when we think about the competitive Advantage right it's combinatory of all these system exactly exactly and and I was exactly going to lead to that because of smarter AIS to curate the data we now even have synthetic data generation and all kinds of different ways of of curating data presenting data to and so before you even get the training you've got massive amounts of data processing involved and so so people think about oh uh pytorch that's the beginning end of the world and it was very important but don't forget before pie torches amount of work after P hor is amount of work and and that the think about the flywheel is really the way you ought to think you know how do I think about this entire flywheel and how do I design a Computing system a Computing architecture that helps you take this flywheel and be as effective as possible it's not one size slice of an application training does that make sense that's just one step yeah okay every step along that flywheel is hard and so the first thing that you should do instead of thinking about uh how do I make Excel faster how do I make you know Doom faster that was kind of the old days isn't that right now you have to think about how do I make this flywheel faster and this flywheel has a whole bunch of different steps and there's nothing easy about machine learning as you guys know there's nothing easy about what open AI does or X does or Gemini and the team that in Deep Mind does I mean there's nothing easy about what they do and so we decided look this is really what you ought to be thinking about this is the entire process you want you want to accelerate every part of that you want to respect amd's law you want to amd's law would suggest well if this is 30% of the time and I accelerated that by a factor of three I didn't really accelerate the entire process by that that much does that make sense and it you really want to create a system that accelerates every single step of that because only in doing the whole thing can you really materially improve that cycle time and and that flywheel that's that that that that rate of learning is really in the end what causes the exponential rise and so our what I'm trying to say is that our perspective about you know a company's perspective about what you're really doing manifests itself into the product right and notice I've been talking about this flywheel you know entirey yeah that's right yeah and we accelerate everything right right now right now the main focus is video uh a lot of people are focused on on on physical Ai and video processing right just imagine that front end right the the terabytes per second of data that are coming into the system give me an example of a pipeline that is going to ingest all of that data right prepare for training in the first place yeah so that entire thing is Cuda accelerated and people are only thinking about text models today yeah but the future is you know this the video models as well as you know using you know some of these text model like 01 to really process a lot of that data before we even get there yeah right yeah yeah language model is going to be involved in every sing it took us took the industry enormous technology and effort to train a language model to train these large language models now we're using a large language model in every single step of the way it's pretty pretty phenomenal I don't mean to be overly simplistic about this but again you know we hear it all the time from investors right yes but what about custom as6 yes but their competitive mode is going to be pierced by this what I hear you saying is that in a combinatorial system the advantage grows over time so I heard you say that our advantage is greater today than it was three to four years ago because we're improving every component and that's combinatorial is that you know when you think about for example as a business case study Intel right who had a dominant mode a dominant position in the stack relative to where you are today perhaps just you know again boil it down a little bit you know compare contrast your competitive advantage to maybe the competitive Advantage they had at the peak of their cycle well Intel extraordinary um Intel is extraordinary because they were probably the first company that that was uh incredibly good at manufacturing um process engineering manufacturing and that one click above manufacturing which is building the chip right and designing the Chip And and architecting the chip um uh in the x86 architecture uh and building faster and faster x86 chips that was their Brilliance and they they fused that with Manufacturing um our company is a little different in the sense that and we recognize this that that in fact parallel processing doesn't require every transistor to be excellent serial processing requires every transistor to be excellent parallel processing requires lots and lots of transistors to be more cost effective I rather have 10 times more transistors 20% slower right than 10 times less transistor 20 % faster does that make sense they would like the opposite and so single threaded performance single threaded processing and parallel processing was very different and so we we observe that in fact our world is not about being better going down we want to be very good as as good as we can be but it's our world is really about much better going up parallel Computing parallel processing is hard because uh every single algorithm requires a different way of refactoring and re architecting the algorithm for the architecture what people don't realize is that you can have uh three different Isa CPU Isis they all have their own C compilers you could take you could take software and compile down to the ISA that's not possible in accelerated Computing that's not possible in parallel Computing the company who comes up with the architecture has to come up with their own openg GL right so we revolutionize deep learning because of our domain specific Library called qnn without qdn nobody talks about qnn because it's one layer underneath p and you know and and tensor flow and back in the old days Cafe and theano and and now Triton and there's a whole bunch of different different Frameworks and so that domain specific Library CNN a domain specific Library called Optics we have a domain specific Library called K Quantum um Rapids uh the list of you know areal for for uh industry specific algorithms that sit below you know that pie torch that everybody's focused on like I've heard often times well you know if llm if I didn't if we didn't invent that uh no application on top could work right you guys understand what I'm saying so the mathematics is really what Nvidia is really good at is algorithm right that in that Fusion between the the science above the architecture on the bottom that's what we're really good at yeah there's all this attention now on inference yeah finally um but I remember you know two two years ago Brad and I had dinner with you and we asked you the question you know do do you think your moat will be as strong in inference as it is in training yeah and and I'm sure I said it would be it would be greater yeah yeah and and and you touched upon a lot of these elements just now just you know the composability between or you know the we don't know the total mix at one point into a customer it's very important to be able to be flexible in between that's right um but can you just touch upon you know now that we're in this era of inference it was you know inference training is inferencing at at scale I mean you're right and so so um uh if you if you INF if you train well it is very likely you'll inference well if you built it on this architecture without any consideration it will run on this architecture okay you could you could still go and optimize it for other architectures but at the very minimum since it's already been architect you know built on Nvidia it will run on Nvidia now the other aspect of course is just kind of you know capital investment aspect which is when you're training new models you want your best new gear to be used for training right which leaves behind gear that you used yesterday well that gear is perfect for inference and so there's there's a there's a trail of free gear there's a trail of free infrastructure behind the new INF structure that's Cuda compatible and so we we're we're very disciplined about making sure that we're compatible throughout so that everything that we leave behind will continue to be excellent now we also put a lot of energy into continuously Reinventing new algorithms so that when the time comes the hopper architecture is two three four times better than when they bought it so that that you that infrastructure continues to be really effective and so all of the work that we do uh improving new algorithms new Frameworks notice it helps every every single installed base that we have Hopper is better for it Amper is better for it even Volta is better for it okay and and I think Sam was just telling me that that they had just uh decommission the the Volta infrastructure that they have at open a recently and so so I I think it's uh we we leave behind this trail of install base just like all Computing install base matters and nvidia's in every single Cloud we you know on Prim and and at all the way out to the edge and so the the the V you know Vision language model that been created in the cloud works perfectly at the edge on a robots MH without modification it's all Cuda compatible and so so I think this this idea of architecture compatibility was important for large it's no different for iPhones no different for anything else I think the install base is really important for inference but the thing that that I really really um uh we really benefit from is because we're de we're working on training these large language models and the new architectures of it uh we're we're we're able to think about how do we create architectures that's excellent at inference someday when the time comes and so we've been thinking about about uh iterative models for for uh reasoning models and how do we create uh very very uh interactive inference experiences for this personal agent of yours you don't want to say something have go off and think about for a while you wanted to interact with you quite quickly so how do we create such a thing and what came out of it was MV link right you know MV link so that we could take uh these systems that are excellent for training U but when you're done with it the the inference performance is ex ex exceptional and so you want to you want to optimize for this time to First token right and time to First token is is um uh insanely hard to do actually because time to First token requires a lot of bandwidth but if your context is also Rich then um you need a lot of flops yeah and so you need an infinite amount of bandwidth infinite amount of flops at the same time in order to achieve just a few millisecond response time and so that that architecture is really hard to do and we invented uh Grace Blackwell MV link for that right in the spirit of time I have more questions about that but but don't worry don't worry about the time hey guys Hey Hey Hey listen Janine yeah look let's do it until right let's do it until right there you go I love it I love it so you know I was at dinner yeah um with Andy jasse earlier now we don't have to worry about the time earli with Andy jasse earlier this week and Andy said you know we've got trainum uh you know coming and INF frenia coming and I think most people um again view these as a problem for NVIDIA but in the very next breath he said Nvidia is a huge and important partner to us and will remain a huge and important partner for us as far as I can see into the future the world runs on Nvidia right so when you think about the custom Asic that are being built that are going to go after targeted application maybe the inference accelerator at meta maybe you know tranium at Amazon uh you know or Google's tpus and then you think about the supply shortage that you have today um do any of those things change that Dynamic right or are compliments to the systems that they're all buying from you we're just doing different things yes um uh we're we're trying to accomplish different things you know what Nvidia is trying to do is build a Computing platform for this new world this machine learning world this generative AI world this agentic AI World we're trying to we're trying to create you know as you know in what what's just Prof so deeply profound is after 60 years of computing uh we reinvented the entire Computing stack uh the way you write software from programming to machine learning the way that you process software from CPUs to GPU uh the way that the way that uh uh the applications from software to artificial intelligence right and so uh uh software tools to artificial intelligence so so every aspect of the Computing stack and the technology stack has been changed you know what we would like to do is to to create a Computing platform that's available everywhere and this is really the the the complexity of what we do the complexity of what we do is if you think about what we do we we're building an entire AI infrastructure and we think of it as one computer right I've said before the data center is now the unit of computing to me when I think about a computer I'm not thinking about that chip I'm thinking about this thing that's my mental model and all the software and all the orchestration all the Machinery that's inside that's my M that's my computer and we're trying to build a new one every year yeah that's nobody has ever done that before we're trying to build a brand new one every single year and every single year we deliver two or three times more performance as a result every single year we reduce the cost by two or three times every single year we improve the Energy Efficiency by two or three times right and so we ask our customers don't buy everything at one time buy a little every year okay and the re reason for that we want them cost average into the future all of it's architecturally compatible okay now so that building that alone at the pace that we're doing is in incredibly hard now the double part the double hard part is then we take that all of that and instead of selling it as a infrastructure or selling it as a service we disaggregate all of it and we integrate it into gcp we integrate it into AWS we integrated it into Azure we integrate it into X we inte does that make sense yes and so everybody's integration it's different we got to get we have to get all of our architectural libraries and all of our algorithms and all of our Frameworks and integrated into theirs we get our security system integrated into theirs we get our networking integrated into theirs isn't that right right then we do basically 10 Integrations and we do this every single year right now that is the miracle that is the miracle why were I mean it's Madness it's Madness that you're trying to do this every year so so so so what drove you to do it every year and then related to that you know Clark's just back from Taipei and Korea and Japan when meeting with all your supply Partners who you have decade long relationships with how important are are those relationships to again the combinatorial math that builds that competitive Moe yeah that's that's um when you when you break it down systematically the more you guys break it down the more Everybody Breaks down the more amazed that they are yes and and um how is it possible that the entire U ecosystem of electronics today is dedicated in working with us to build ultimately this cube of a computer integrated into all of these different ecosystems and the coordination is so seamless so there's obviously apis and and methodologies and business processes and design rules that we've propagated back backwards and methodologies and architectures and a apis that we propagated forward that have been hardened for decades Harden for decades yeah and also evolving as we go and um but they they these apis have to come together right right when the time comes all these things in Taiwan you know all over the world being manufactured they're going to land somewhere in in azure's data center they're going to come together click click click click click someone just calls an uh open ey API and it just works that's right yeah yeah exactly it's kind of craziness right and so that's what we invented that's what we invented this this this massive infrastructure of computing yeah the whole planet is working with us on it it's integrated into everywhere it's you could sell it through Dell you can sell through HP it's hosted in the cloud it's an uh it's all the way out at the edge uh people use it in robotic systems now Rob and you know human robots they're in self-driving cars they're all architect compatible pretty kind of craziness it's it's craziness Clark I don't want to I don't want you to leave the impression I didn't answer the question in fact I did um what I meant by that when Rel to your Asic is is um the way to think about we're just doing something different yes um as a company as a company we want to be sit situationally aware and I'm very situationally aware of everything around our company and our ecosystem right I'm aware of all the people doing alternative things and and what they're doing and and and and sometimes sometimes it's adversarial to us sometimes it's not I'm I'm super aware of it but that doesn't change what the purpose of the company is yeah the singular purpose of the company is to build an architecture that a platform that could be everywhere right that is our goal we're not trying to take any share from anybody Nvidia is a market maker not share taker if you look at our company slides we don't we don't show not one day does this company talk about market share not inside all we're talking about is how do we create the next thing what's the next problem we can solve in that flywheel how can we do a better job for people how do we take that that flywheel that used to take about a year how do we crank crank it down to about a month yes you know what's the speed of light of that isn't that right and so we're thinking about all these different things but the one thing we're not we're not to we're situationally aware of everything but we're certain that what our mission is is very singular the only question is whether that mission is necessary does that make sense yes you know and all companies all great companies ought to have that at its core it's about what are you doing for sure the only question is it necessary is it valuable right is it impactful does it help people and and I am un certain that you're a developer you're you're you're a generative AI startup and and you're about to decide how to become a company the one choice that you don't have to make is which one of the A6 do I support if you just support a Cuda you know you could go everywhere you could always change your mind later right but we're the onramp to the world of AI isn't that right once you decide to come onto our platform the other decisions you could defer you could always build your own basic later you know we're not against that we're not offended by any of that um when I work with when we work with um all the gcps uh the gcps aure we present our road map to them years in advance they don't present their Asic road map to us and it doesn't ever offend us does that make sense we we create we in the if you have a sole purpose and your purpose is Meaningful and your mission is is is dear to you and is dear to to everybody else then you could be transparent notice my road map is transparent at GTC my road map goes way deeper to our friends at Azure and AWS and others um we have no trouble doing any of that even as they're building their own as6 I think you know when when people observe the business you said recently that the demand for Blackwell is insane you said one of the hardest parts of your job is the emotional toll of saying no to people in a world that um has a shortage of the compute that you that you can produce and have on offer but critics say this is just a moment time right they they say this is just like Cisco in 2000 we're overbuilding fiber it's going to be boom and bust you know I I think about the start of 23 when we were having dinner the forecast for NVIDIA at that dinner in January of 23 was that you would do 26 billion of Revenue uh for the year 2023 you did 60 billion right the the 25 people let's just let let let the truth be known that is the single greatest failure of forecasting the world has ever seen right right can we all can we all at least admit that what what what to me to me that was my takeaway I just got and that was and that was we got so excited in November 22 because we had folks like Mustafa from inflection and no him from character coming in our office talking about investing in their companies and they said well if you can't pencil out investing in our companies then buy Nvidia because everybody in the world is trying to get Nvidia chips to build these applications that are going to change the world and of course the Cambrian moment occurred with chat GPT and not withstanding that fact these 25 analysts were so focused on the crypto winner that they couldn't get their head around an imagination of what was what was happening in the world okay so it ended up being way bigger you say in very plain English the demand is insane for Blackwell that it's going to be that way for as far as you can you know for as far as you can see of course the future is unknown and unknowable but why are the critics so wrong that it that this isn't going to be the Cisco like situation of overbuilding in the in in in 2000 yeah um the best way to to think about the future is reason reason about it from first principles correct okay so so the the question is what are the first principles of what we're doing number one what are we doing what are we doing um the first thing that we are doing is we are Reinventing Computing do we not we just said that the way that Computing will be done in the future will be highly machine learned yes highly machine learned okay almost everything that we do almost every single application word excel PowerPoint uh Photoshop Premiere you you you AutoCAD you you give me your favorite application that was all hand hand engineered I promise you it will be highly machine learned in the future isn't that right and so all these tools will be and and on top of that you're going to have machines agents that you help you use them right okay and so we know this for a fact at this point right isn't that right we've reinvented Computing we're not going back the entire Computing technology stack is being reinvented okay so now that we've done that we said that software is going to be different what software can write is going to be different how we use software will be different so let's let's now ackn knowledge that th those are my ground truth now yes now the question therefore is what happens and so let's go back and let's just take a look at how's Computing done in the past so we have a trillion dollars wor of computers in the past we look at just open the door look at the data center and you look at and say are those the computers you want doing that doing that future and the answer is no right right you got all these CPUs back there we know that what what it can do and what it can't do and we just know that we have a trillion dollars wor the data centers that we have to modernize and so right now we speak if we were to to have a trajectory over the next four or five years to modernize that old stuff that's not unreasonable right sensible so we have a and you're having those conversations with the people who have to modernize it and they're modernizing it on GPU that's right I mean well let's let's make another test you have you have you have $50 billion of capex you like to spend option a option b bill capex for the future right or build capex like the past right know um you already have the capex of the past corre it's sitting right there it's not getting much better anyways Mor's law has largely ended and so why rebuild that let's just take $50 billion put it into generative AI isn't that right and so now your company just got better right now how much of that 50 billion would you put in well I would put in 100% of the 50 billion because I've already got four years of infrastructure behind me that's the the of the past and so now now you just I just reasoned about it um from the perspective of of somebody thinking about it from first principles and that's what they're doing smart people are doing smart things now the second part is this so so now we have a trillion dollars worth of capacity go bill right trillion dollars worth of infrastructure wor about you know call it $150 billion into it right okay so we have we have a trillion dollars of infrastructure INF foral bill over the next four or five years well the second thing that we observe is that the way that software is written is different but how software is going to be used is different in the future we're going to have agents isn't that right we're going to have digital employees in our company MH in your inbox you have all these little dots and these little faces in the future there's going be little icons of AIS isn't that right I'm going to be sending them I'm going to be I'm no longer going to program computers with C++ I'm going to program AIS with prompting isn't that right now this is no different than me talking to my you know this morning I I wrote a bunch of emails before I came here I was prompting my team course right yeah and I I would describe the context I would describe the the the fundamental constraints that I I know of and I would describe the mission for them I would leave it sufficiently uh I would be sufficiently directional so that they understand what I need and I want to be clear about what the outcome should be as clear as I can be but I leave enough ambiguous space on you know a creativity space so they can surprise me isn't that right absolutely there's no different than how I prompt an AI today yeah it's exactly how I prompt an AI and so what's going to happen is is on top of this infrastructure of it that we're going to modernize there's going to be a new infrastructure this new infrastructure are going to be AI factories that operate these digital humans right and they're going to be running all the time 247 right we're going to have them for all of our companies all over the world uh we're going to have them in factories we're going to have them in autonomous systems is that right so there's a whole layer of computing fabric a whole layer what I call AI factories that the world has to make that doesn't exist today at all right so the question is how big is that right unknowable at the moment probably a few trillion dollar right un unknowable at the moment but as we're sitting here building into the beautiful thing is the architecture for this modernizing this new Data Center and the architecture for the AI Factory is the same right that's the nice thing and you you you made this clear you've got a trillion of old stuff you got to modernize you at least have a trillion of new AI workloads coming on you're give or take you'll do 125 billion in Revenue this year you know there was at one point somebody told you the company would never be worth more than a billion as you sit here today is there any reason right if you're only 125 billion out of a multi- trillion Tam that you're not going to have 2X the revenue 3x the revenue in the future that you have today is there any reason your Revenue doesn't no yeah yeah the as as you know it's not about it's not about um everything is you know companies companies are only limited by the size of the the fish pond you know a goldfishing can only be so big and so the question is what is our what is our fish pond what is our pond and that requires a little of imagination and this is the reason why market makers think about that future without creating that new fish pond um it's hard it's hard to to figure this out looking backwards and try to take share right you know share takers can only be so big for sure market makers can be quite large for sure yeah and so you know I I think I think the the good fortune that our company has has is that since the very beginning of our company we had to invent the market for us to go swiming that Mark and people don't realize this back then but anymore but you know we were at the at the at Ground Zero of creating the 3D gaming PC market right we we largely invented this market and all the ecosystem and all the the graphics card ecosystem we invented all that and and so so the the the the need to invent a new market to go serve it later is something that's very comfortable for us exactly exactly and speaking to somebody who's invented a new market you know let's shift gears a little bit to models and open AI open AI raised as you know $65 billion dollar uh this week um uh at like $150 billion valuation we both participated yeah really happy for them really really happy they came together they did a great s and the team did a great job yeah reports are that they'll do five billion is of Revenue or run rate Revenue this year maybe going to 10 billion next year if you look at the business today it's about twice the revenue as Google was at at the time of its IPO they have 250 yeah 250 million weekly average users which we estimate is twice the amount Google had at the time of it IPO and if you look at the multiple of the business if you believe 10 billion next year it's about 15 times the forward Revenue which is about the multiple of Google and meta at the time of their IPO right when you think about a company that had zero Revenue zero weekly average users 22 months ago Brad has an incredible command of history when you think about that um H talk to us about the importance of open AI as a partner to you and open AI as a for foring kind of driving forward you know kind of public awareness and usage around AI well this this is one of the one of the most consequential companies of our time uh the um uh a uh a pure play um AI company uh pursuing the the uh uh the vision of uh AGI right and whatever its definition right I I almost don't think it matters fully what the definition is nor do I um uh you know really believe that that the timing matters right the the one thing that I know is that that ai ai is going to have a a road map of capabilities over time and that road map of capabilities over time is going to be quite spectacular and um uh along the way long before even get to anybody's definition of AGI we're going to put it to Great use right um all you have to do is uh right now as we speak go go talk to uh digital biologists uh climate Tech researchers material researchers um uh physical sciences astrophysicists Quantum chemists um you go ask uh uh video game designers um uh manufacturing uh um engineers uh roboticist pick your favorite whatever industry you want to go pick right and you go deep in there and you talk to the people that matter and you ask them has AI revolutionized the way you work right and you take those data points and you come back and you you then get to ask yourself how skep skeptical do you want to be right right because they're not talking about AI as a conceptual benefit right someday they're talking about using AI right now correct right now act Tech material Tech climate Tech you you pick your Tech you pick your field of science they are advancing AI is helping them advancing their work right now as we speak every single industry every single company every High every University unbelievable isn't that right right it is absolutely um going to somehow transform business we know that right we I mean we we it's it's so tangible you could it's happening today it's happening today it's happening today yeah and and and so I I think I think um uh I think the the The Awakening of AI chat GPT triggered uh it is completely uh incredible and and I I love I love their uh their uh uh their velocity and their their singular purpose of advancing this field and and so really really consequential company and they build an economic engine that can Finance the next you know Frontier of models right and I think there's a growing consensus in Silicon Valley that the whole model layer is commoditizing llama is is is is is making it very cheap for lots of people to build models and so early on here we had a lot of model companies you know character and inflection and and cohere and mrr and go through the list and a lot of people question whether whether or not those companies can build the escape Velocity on the economic engine that can continue funding those Next Generation my own sense is that there's going to be that's why you're seeing the consolidation right it's open AI clearly has hit that escape velocity they can fund their own future it's not clear to me that many of these other companies can is that a fair kind of review of the state of things in the model layer that we're going to have this consolidation like we have in lots of other markets Market leaders who can afford who have an economic engine an application that allows them to continue to invest um there's a first of all there's a different fundamental difference between a model yes and artificial intelligence yes right yeah a model is an essential ingredient correct for artificial intelligence it's necessary but not sufficient correct and so and the and artificial intelligence is a capability but for what right then what's the application right the artificial intelligence for sof driving cars is related to the artificial intelligence for human robots but it's not the same which is related to the artificial intelligence for a Chad bot but not the same correct and so so you have to understand the taxonomy yes of Stack yeah of the stack yeah and at every layer of the stack there will be opportunities but not infinite opportunities for everybody at every single layer dis sack right now I just said something all you have to do is replace the word um model with GPU in fact this was the great observation of our company 32 years ago that there's a fundamental difference between GPU Graphics chip or GPU versus accelerated Computing and accelerated Computing is a different thing than the work that we do with AI infrastructure it's related but it's not exactly the same it's built on top of each other it's not exactly the same and each one of these layers of abstraction requires fundamental different skills somebody who's really really good at building gpus have no clue how to be an accelerated Computing company I can I there are a whole lot of people who build gpus yeah and I don't know which one came to came you know we invented the GPU but you know that we're not we're we're not the only company that makes gpus today correct you know and so they're gpus everywhere and but they're not accelerated Computing companies and and there are a lot of people who you know there's they're they're accelerators accelerators that does uh application acceleration but that's different than an accelerated Computing company and so for example a very specialized AI application right could that could be a very successful thing correct that is mtia that's right but it might not be the type of company that that um had brought reach and brought capabilities and so so you've got to you've got to decide where you want to be there's opportunities probably in all these different areas but like building companies you have to be mindful of the the the shifting of the ecosystem and what gets commoditized over time recognizing what's a feature versus a product right versus a company for sure okay I just I just went through okay there's a lot of different ways you can think about this of course there's one new entrant that has the money the smarts the ambition that's x.

a right and um well there are reports out there that that you and Larry and Elon had dinner talked you out of 100,000 h100s they went to Memphis and built a large coherent super cluster in in a matter of months um you know so first three three points don't make a line okay yes I had dinner with them causality is what do you think about their ability to stand up that super cluster and there's talk out there that they want another 100,000 h20s right to expand the size of that super cluster you know first talk to us a little bit about X and their Ambitions and what they've achieved but also are we already at the age of clusters of 200 and 300,000 gpus um the answer is yes and then the um uh first first of all uh acknowledgement of of achievement where is deserved from the moment of concept to um a data data center that's ready for NVIDIA to have our gear Gear there to the moment that we uh uh powered it on had it all hooked up and it did its first training yeah okay so uh that first part just building a massive factory um liquid cooled uh energized permitted uh in the short time that was done I mean that is that is like super human right yeah there's and and as far as as far as I know there's only one person in the world who could do that you know I mean Elon is singular in this understanding of engineering and and construction and large systems and um and and and marshalling resources um yeah just it it's unbelievable and and then and of course then his engineering team is extraordinary I mean the the software team's great the networking team's great the infrastructure team is great you know Elon understands this deeply and from the moment that we decided to get to go um the planning of with our engineering team our networking team or our infrastructure Computing team the software Team all of the preparation Advance um then all of the infrastructure all of the logistics and the amount of technology and equipment that came in on that day nid nvidia's infrastructure and Computing infrastructure and all that technology to train training 19 days just you know do you don't want did anybody sleep 247 no no question then nobody slept but but first of all some 19 days is incredible but it's also kind of nice to just take a step back and just do you know how 19 how many days 19 days is it's just a couple weeks yeah right and and the mountain of Technology if you ever to see it is unbelievable all of the wiring and the networking and you know networking in video gear is very different than networking hyperscale data centers okay the number of wires that goes in one node the back of a computer is all wires it's just getting this mountain of Technology integrated and all the software incredible yeah so so I I think I think what Elon and the ex team did and um and I I'm really appreciative that he he acknowledges the the engineering work that we did with him and and the planning work and all that stuff um but but what they achieved is is s never been done before just to put in perspective 100,000 gpus that's you know easily the fastest supercomputer on the planet as one cluster um a supercomputer uh that you would build would take normally three years to plan right and then they deliver the equipment and it takes one year to get it all working yes we're talking about 19 days wow what's the credit of the Nvidia platform right right that it's the whole processes are hardened that's right yeah everything's already working and and of course there's a whole bunch of you know X algorithms and X framework and X stack and things like that and we got a ton of integration we have to do but the planning of it was extraordinary just pre-planning of it to you know n of one is right Elon is an N of one you but you answered that question by starting off saying yes 200 to 300,000 GPU clusters are are are here yeah right um does that scale to 500,000 does it scale to a million and does the demand for your products depend on it scaling two millions that part the last part is no um my sense is that uh distributed training will have to work right and my sense is that that uh distributed computing will be invented right and and some form of Federated learning and and distributed you know um asynchronous distributed computing uh is going to is going to um uh be discovered and I'm very very enthusiastic and very optimistic about that um the the uh uh of course of course the um uh the thing to realize is that the scaling law used to be about pre pre-training now we've gone to M multimodality we've gone to synthetic data generation right um post trining has now scaled up incredibly synthetic data generation reward systems reinforcement learning based and then now inference scaling has gone through the roof right the idea that that a model before it answers your answer had already done internal inference incredible 10,000 times M it's probably not unreasonable and it's probably done research it's probably done reinforcement learning on that it's probably you know it's probably done some simulation it's surely done a lot of reflection it probably looked up some data looked some information isn't that right and so his context is probably fairly large I mean this this type of intelligence is well that's what we do right that's what we do isn't that right and so so the ability this this scaling if you just did that math um and you compound it with you add you compound that with 4X per year on model size and Computing size and then on the other hand demand continues to grow in usage uh do we think that we need millions of gpus No Doubt yeah yeah that is that is a for certainty now yeah and so the question is how do we architect it from a data center perspective and that has has a lot to do with you know are are there data centers that are gws at a time or they 250 megawatts at a time and and um uh my sense is that you know you going to get both I think analysts always focus on the current architectural bet but I think one of the biggest takeaways from this conversation is that you're thinking about the entire ecosystem and many years out so you know the idea that you know because Nvidia is just scaling up or scaling out it's to meet the future it's not to you know not not such that you know you're only dependent on a world where there is a 500,000 or a million you know GPU cluster it's you know by the time there's dist distributed training you'll have written you know theof to enable that that's right remember without Megatron yeah that we developed some seven years ago now yeah the scaling of these large training jobs wouldn't have happened right and so we invented Megatron we invented nickel um GPU direct right all of the work that we did with our dma um that made it possible for easily to do um uh P pipeline parallelism you know right and so you know all the all the model parallelism that's being done you know the breaking of the distributor training and all the batching and all that all of that stuff is is U uh because we did the early work and now we're doing the early work for the future future generation so so let's talk about strawberry and 01 I want to be respectful of your time so I got all the time in the world well you're you're very generous yeah I got all the time in the world but first I think it's cool that they named 01 after the 01 Visa right which is about recruiting the world's best and brightest uh you know and bringing them to the United States it's something I know we're both deeply passionate about um so I love the idea that building a model that thinks and that takes us to the next level of scaling intelligence right uh is is is an homage to the fact that it's these people who come to the United States by way uh of immigration that have made it made us what we are bring their collective intelligence to the an alien intelligence certainly you know it was spearheaded by our friend noan Brown of course he worked at pabus and Cicero when he was at meta how big a deal is inference time reasoning as a totally new Vector of scaling intelligence separate and distinct from right just building larger models it's a huge deal it's a huge deal I think the um uh a lot of intelligence can't be done a priori right you know and a lot of computing even a lot of computing can be reordered I mean just you know out of order execution can be done a priority you know and so a lot of things can only be done in runtime right and and so so whether you think about it from a computer science perspective or you think about it from from a from a intelligence P perspective uh too much of it requires context right um the circumstance right uh the quality the type of answer you're looking for uh sometimes just a quick answer is good enough depends on depends on the the the uh um the consequential you know impact of the answer right you know depending on the nature of the usage of that answer and so so uh uh some some some answers uh please take a night some answers take a week yes is that right so I could totally imagine uh me sending off a prompt to my Ai and telling it you know think about it for night right think about it overnight don't tell me right away right I want you to think about it all night and then come back and tell me tomorrow what's your best answer and reason about it for me and and so I think the the U uh the the quality the the segmentation of intelligence from now from a product perspective right there's going to be one shot versions of it right for sure yeah and then there'll be some that take five minutes you know and the intelligence layer that Roots those questions to the right model for The Right Use case I mean we were using advanced voice mode and 01 preview last night it I was I was coaching my son for his AP History test and it was like having the world's best AP History teacher sitting right next to you thinking about these questions it was truly extraordinary again my tutor is an AI today right of course they're here today which comes back to this you know over 40% of your Revenue today is inference but inference is about ready because of chain of reasoning yeah right it's about up by a billion times right by by by by a million x by by a billion x that's right that's the part that most people have you know haven't completely internalized this is that industry we were talking about but this is the Industrial Revolution right that's the production of intelligence that's right right and and it's going to go up a billion times right and so you know everybody's so hyperfocused on Nvidia as kind of like doing training on bigger models yeah right isn't it the case that your Revenue if it's 5050 today you're going to do way more inference in the future yeah right than I mean training will always be important but just the growth of inference is going to be way larger than the growth in training we hope it's almost impossible to conceive otherwise yeah we hope that's right that's right right yeah it's it's good to it's good to go to school yes but the goal is so that you can be productive in society later and so it's good that we train these models but the goal is to inference them you know are you already using chain of reasoning and uh you know tools like 01 in your own business to improve your own business yeah our cyber security system today can't run without without um uh our own agents uh we have we have uh agents helping design chips Hopper wouldn't be possible black wall would be possible Ruben don't even think about it U we have digital we have we have ai chip designers AI software Engineers AI verification Engineers um and we we build them all inside because you know we we um uh we have the ability and and we rather we rather use it use the opportunity to explore the technology ourselves you know when I walked into the building today somebody came up to me and said you know ask Jensen about the culture it's all about the culture I look at the business you know we talk a lot about Fitness and efficiency flat organizations that can execute quickly smaller teams um you know Nvidia is in a league of its own really um you know at about 4 million of Revenue per employee about 2 million of profits are free cash flow per employee you've built a culture of efficiency that really um has Unleashed creativity and Innovation and ownership and responsibility You' broken the mold on kind of functional management everybody likes to talk about all of all of your uh direct reports um is the leveraging of AI the thing that's going to continue to allow you to be hyper creative while at the same time being efficient no question U I'm hoping that that someday Nvidia has 32,000 employees today right and um we have four we have 4,000 families in Israel I hope they're well I'm thinking of you guys yes and uh um I'm hoping that Nvidia Sunday will be a 50,000 employee company with a 100 million you know AI assistants wow and and they're in every single group all right um uh we we'll have a whole directory of uh AIS that are just generally good at doing things we'll also have our inbox is going to full of directories of AIS that we work with that we know are really good specialized at our skill and so so um AIS will recruit other AIS to solve problems right AIS will be in you know slack Channels with each other and with humans right and with humans and so so we'll just be one large you know employee base if you will uh some of them are digital in AI some some of them are biological and and I'm hoping some of them even megatronics you know I I think from a business perspective it's something that's greatly misunderstood you just described a company MH that's producing the output of a company with 150,000 people but you're doing it with 50,000 people now you didn't say I was going to get rid of all my employees you're still growing the number of employees in the organization but the output of that organization right is going to be dramatically more this this is this is often misunderstood um AI is not it's not AI will change every job right AI will have a a a seismic impact on how people think about work let's acknowledge that right um AI has the potential to do incredible good it has the potential to do harm we we we we have to build safe AI yes we let's just make that foundational yes okay the part that is the part that is overlooked is when companies become more productive using artificial intelligence it is likely that it manifests itself into either better earnings or better growth or both right and when that happens the next email from the CEO is likely not a layoff announcement of course because you're growing yeah and the reason for that is because we have more ideas than we can explore and we need people to help us think through it before we automate and so the automation part of it AI can help us do obviously it's going to help help us think through it as well but it's still going to require us to go figure out what problems do I want to solve there are a trillion things we can go solve what what what problems does this company have to go solve and select those ideas and figure out a way to autom automate and scale and so so as a result we're going to hire more people as we become more productive people forget that you know and and and if you if you go back in time obviously we have more ideas today than than than 200 years ago that's the reason why gdps are larger and more people are employed and even though we're automating like crazy underne I it's such an important point of this period that we're entering one almost all human productivity almost all human Prosperity is the byproduct of the Automation and the technology of the last 200 years I mean you can look at you know from Adam Smith and shumer creative you know destruction you can look at chart of GDP growth per person over the course of last 200 years and it's just accelerated which leads me to this question if you look at the 90s our productivity growth in the United States was about 2 and a half to 3% a year okay and then in the 2000s it slowed down to about 1.