transcribe

Ep18. Jensen Recap - Competitive Moat, X.AI, Smart Assistant | BG2 w/ Bill Gurley & Brad Gerstner

Bg2 Pod · 54m · transcribed Jun 2026
More from Bg2 Pod Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:00 you may also be running up against the even for the mag 7 the size of kappo X deployment where there CFOs start to talk at higher levels for sure [Music] [Applause] [Music] totally Sunny Bill great to see you guys good to see you good to be back thanks man it's great to have you we literally

0:31 just finished two days of the altimeter annual meetings I mean we had hundreds of investors CEOs Founders and the theme was uh scaling intelligence to AGI uh we had nikesh talking about Enterprise AI we had Renee hos talking about AI at the edge we had noan Brown talking about you know the strawberry no1 model and inference time reasoning we had Sunny talking about you know uh accelerating inference and of course we kicked off with Jensen talking about the future of

1:02 compute you know I did the Jensen part uh talk with my partner Clark Tang who covers the compute layer in the public side we recorded it on Friday we'll be releasing it as part of this pod and man was it dense I mean he was you know he was on fire he told me I asked him at the beginning of the Pod what do you want to do he said grip it and rip it and we did uh 90 minutes we went deep I

1:25 shared it with you guys we've all listened to it I learned so much playing it back back that I just thought it it made sense for us to unpack it right to really uh to really analyze it see what we agree with what we may disagree with things we want to further explore Sunny uh any highlevel reac reactions to it yeah you know first it's the first time it you've I really seen them in a format

1:49 where you got all that information out in in one setting because you kind of get the you get the tidbits and the ones that really struck with me was when he said Nvidia is not a GPU company they're an accelerated compute company I think the next one you know which we'll you'll touch on is where he really said the data center is the unit of compute I thought that was that was massive and you know sort of just

2:12 closing out when he talked about he thinks about using and already utilizing so much AI within Nvidia and how that's a superpower for them to accelerate over everyone they're competing with I thought those were kind of really awesome points and him you know eating the dog food as they say it is incredible you know there's this this thing we'll talk about later but he said he thinks they can 3x you know the top line of the business while only adding

2:39 25% More Humans because they can have a 100,000 autonomous agents doing things like building the software doing the security and that he becomes really a prompt agent not only for the his human direct reports but also for these agents which uh you know really is is is mindboggling Bill anything stand out for you well one I mean you should be pleased that you're able to get his time um you know this is at points in time the largest market

3:10 cap company in the world um if not one two and so it was so I think kind of him to sit down with you for so long and during the part he kept saying I can stay as long as you want I was like does he have something to be doing like it was it was is incredibly generous and it's fantastic but my other big my my I mean two I had two big takeaways one I

3:34 mean it's obvious that this guy's you know rolling on all cylinders here right like you have a company at a 3.3 trillion market cap that's still growing over 100% a year and the margins are insane I mean 65% operating margins there's only like five companies in the S&P 500 at that level and they certainly aren't growing at this pace you bring up that point about getting more done on the increment with fewer employees

4:04 where's this going to go like 80% operating margin I mean that would be unprecedented there's a lot of that's already here that's unprecedent but obviously Wall Street is fully aware of the corre unbelievable performance of this company and you know the multiples reflected and and the market cap reflects it but but it's super powerful um how they're executing and you can see the confidence in every answer that he gives we spent about a third of the pod

4:35 on nvidia's competitive mode really trying to break it down really trying to understand this idea of systems level advantages the combinatorial advantages um that he has in the business because I think when I talk to people around the investment Community despite how well it's covered bill right there's still this idea that it's just a GPU and that somebody's going to build a better chip they're going to come along and displace the business and so when he said again it can sound

5:03 like marketing speak sunny when somebody says it's not a GPU company it's an accelerated compute company you know we showed this we showed this chart where you can see kind of the Nvidia full stack and he talk about how he just built layer after layer after layer of the stack you know over the course of the last decade and a half but when he said that sunny I know you had a reaction to it right even though you

5:28 know it's not just a GPU company when he really broke it down it seemed like you know he did break new territory here yeah like what was great to hear from him and really you know positive for you know folks thinking about where Nvidia lives in the stack right now is he kind of got into details and then the sub details below Cuda and he really started going into what they're doing very particularly on mathematical operations

5:56 to accelerate their partners and how they work really closely with their Partners you know all the the cloud service providers to basically build these functions so that they can further accelerate workloads the other little Nuance that I picked up in there he didn't Focus purely on llms he talked in that particular area about how they're doing that for a lot of traditional models and even newer models that are being deployed for AI and I think just

6:21 really showed how they are partnering much closer on the software layer than the hardware layer alone right I mean in fact you know talked about you know the the the Cuda library now has over 300 industry specific acceleration algorithms right where they deeply learn the industry right so whether this is synthetic biology or this is image generation or this is autonomous driving they learn the needs of that industry and then they accelerate the particular workloads and that for me was also one

6:53 of the one of the key things this idea that every workload is is moving from kind of this deterministic you know uh you you know handmade workload to something that's really driven by Machine learning and really infused with AI and therefore benefits from acceleration even something as ubiquitous as data processing yeah and I shared this code sample with Bill as you know we were just preparing for this pod and you know U I knew Bill processed it

7:23 right away and and ran it which was it really showed like every piece of code that's out there now that's related to or not every piece of many of the pieces have this like sort of if device equals Cuda do X and if it's not do y and that's the level of impact they're having across the you know entire ecosystem of services and apps that are being built that are related to AI bill I don't know what you thought when you

7:44 when you saw that piece yeah I mean I I think there is a there there's a question for the long term that relates to Cuda and I want to go back to the system point you made later Brad but while we're on Cuda um is what percentage of developers will touch Cuda and is that number going up or down and I could see arguments on both sides uh I you could say the a models are going to

8:12 get more and more hypers specialized and performance matters so much that the one the models that matter the most the deployments that matter the most they're going to get as close to the metal as possible and then Cuda is going to matter the other side you can make is those optimizations are going to live in pie torch they're going to live in other tools like that and the marginal developers not going to need to know that and I don't I could make AR both

8:39 arguments but I think it's an interesting question going forward I mean I I just asked chat GPT how many Cuda developers there are today just to be on top of three million Cuda developers right and you know a lot more that touch Cuda that you know aren't specifically kind of developing on so it is one of these things that has become pretty ubiquitous and his point was it's not just Cuda of course it's you know it

9:02 it's really full stack you know all the way from data ingestion all the way through you know kind of the post training I think I'm on the ladder of your point Bill like I think there's going to be fewer people touching that and I do think that's a point where there the moat is not as strong as a longer term as you say and think about like you know the way the analogy that I I would go with is like think about the

9:22 number of iPhone IOS like developers working at Apple building that versus the number of app developers right and I think you're going to have a you know 10 to1 or 100 to1 ratio of people building at layers above versus people building down closer to the bare metal that would be something to watch we can ask more people over time um obviously it's a big luck today for sure you know and I think Bill to your point you know I reached

9:45 out to Gavin actually before I did the interview Gavin Baker's who's a good buddy and who obviously knows the space incredibly well has followed it uh at a deeper level for a longer period of time than I have and you know like when when I asked him about the competitive Advantage he really said a lot of the competitive Advantage is around this algorithmic diversity and and and and innovation in why Cuda matters he said if the world standardizes on

10:12 Transformers on pytorch then it's less relevant for G gpus um you know in that environment like if you have a lot of standardization right then then Advantage goes to the custom Asic but I'll tell you this you know and and I've had this conversation with a lot of of people when I asked Jensen I pushed him on you know custom Asic I was like hey you know you've got you know accelerated inference coming from meta with their

10:37 mtia chip you know you got infinia and tranium you know coming he's like yeah Brad like they're it you know they're my biggest Partners I actually share my three to five year road map with them yes they're going to have these these Point solutions that are going to do these very specific tasks but at the end of the day the vast majority of the workloads in the world that are machine learning and AI infused are going to run

10:59 on Nvidia and the more people I talk to the more I'm convinced that that's the case despite the fact that there'll be a lot of under winners including Gro and cerebrus Etc and they're acquiring companies they're moving up to stack they're trying to do more optimization at higher levels um so they want to extend obviously what C is doing don't go to inference yet that's a whole G get I'm actually on that bit about the Deep

11:24 Integrations right because you know really that's a Playbook that I I think Microsoft really had done well for a long time in Enterprise software and you really haven't seen that in Hardware ever you know if you go back to say Cisco or the PC era or you know the cloud era you didn't see that deep level integration now Microsoft pulled it off with Azure and it when I heard him talking all I could think about was man

11:48 that was really smart what he's done is he's gotten together really understand what the use cases are and build an organization that deeply integrates into his customers and does it so well all the way up into his road map that he's much more deeply embedded than anyone else is I when I I heard that part I I kind of gave him a real tip of the Hat on that one but what did you you know Brad what was your take on that you and

12:12 I had this conversation after we first listened to it and you know if you really te telescope out you know he talks as a systems level engineer right even if you hear like people PE you know people went to Harvard Business School say how can this guy possibly have 60 direct reports right but how many direct reports Does Elon have right these systems level and he said I have situational awareness right I'm a prompt engineer to the best

12:39 people in the world at these specific tasks I think when I look at this the thing that I deeply underappreciated a year and a half ago about this company was the systems level thinking right that these art that he spent years thinking about how to embed this competitive advantage and how it really it goes all the way from power all the way through application and every day they're launching these new things to further embed themselves in the

13:05 ecosystem but I did hear from somebody over the last two days who you know Renee hos the CEO of arm right Renee was also at our event and and he's a huge Jensen fan he he worked eight years at Nvidia before becoming the CEO of arm in 2013 and he said listen nobody is going to assault the Nvidia Castle head-on right like the main frame of AI right is entrenched and it's going to become a lot bigger at least as far as the eye

13:35 can see he said however if you think about where we're interacting with AI today right on these devices on edge devices he's like our installed base at arm is 300 billion devices and increasingly a lot more of this compute can run closer to the edge if you think about an orthogonal competitor right uh again if he has a deep competitive Mo in

14:05 the cloud what's the orthogonal competitor the orthogonal competitor peels off a lot of the AI on the edge and I think arm's incredibly well positioned to do that clearly nvidia's got arm embedded now in a lot of their uh you know in a lot of their Grace Blackwell Etc um but that to me would be one area like if you looked out and you said where can their competitive Advantage uh you know be challenged a little bit I don't think they

14:29 necessarily have the same level of advantage on the edge as they have in the cloud you started to pod by saying you know the everyone's heard this in the investment Community it's not a GPU company it's a systems company and I in my brain I think had thought oh well they've got four in a box instead of you know just one GPU or eight in a box at the time I was listening to the podcast you did with Jensen I was reading this

14:55 uh neocloud Playbook and Anatomy post by Dylan Patel yes that's a good one he goes into extreme detail about the architecture of some of the larger systems you know like the one that X that AI that we're going to talk about that was just deployed which I think is a 100,000 nodes or something like that and it literally changed my opinion of exactly what's going on in the world and actually answered a lot of questions I

15:24 had but it appears to me that nvidia's competitive Advantage is strongest where the size of the system is largest which is another way of saying what Renee said it's flipping it on its head it's not to not to say it's it's it's weak on the edge but it's super powerful when you put a whole bunch of them together that's when the networking piece thrives that's where nvlink thrives that's where Cuda really comes alive in the biggest

15:53 systems that are out there and some of the questions that answered for me was one why why is demand so high at the highend and why are nodes available on the internet you know single nodes available on the internet for at or below cost and this starts to get at that because you can do things with the large systems that you just can't do with a single note and so you can those two things can be simultaneously true why was Nvidia so

16:21 interested in cor weave existing now I understand like like like if if the biggest systems are where the biggest competitive advantage is you need as many of these big system companies as you can possibly have and and there may be if that if that trajectory remains true you could have an evolution where customer concentration increases for NVIDIA over time um rather than going the other way depending on how you know if if Sam's right that they're going to spend 100

16:51 billion or whatever on a single model there's only so many places you're going to they're going to be able to afford that um but but a lot of stuff started to make sense to me that didn't before and I clearly underestimated the scale of what it meant to be a non-gpu company to be a system company there's there this goes way way up yeah and you know again bill you touched on something that I think is is

17:15 is really important here and this is this question of whether their competitive Moe is also as powerful in training as it is in inference right because I think I think that there's a lot of doubt as to whether their competitive mode is as strong as inference but you know uh let's just you want to flip to that what no but but but I asked him if it was as strong and he he actually said it it was

17:43 greater right to me you know when you think about that in the first instance right I think it didn't make a lot of sense but then when you really started thinking about it he said there's a trail of infra behind the infrastructure that's already out there that is Cuda compatible and can be amortised for all this inference and and so he like for example reference that open AI had just decommissioned Volta so it's like this massive installed base and when they

18:10 improve their algorithms when they improve their Frameworks when they improve their Cuda Li libraries it's all Backward Compatible so Hopper gets better and Amper gets better and Volta gets better that combined with the fact that he said everything in the world today is becoming highly machine learned right almost everything that do he said almost every single application word excel PowerPoint Photoshop AutoCAD like it all will run on these modern systems Sunny do do you buy that do you buy that

18:41 you know when people go to replace you know compute they're going to replace it on these modern systems so when I was listening to it I was buying it but then when I he said one thing that kept resonating in my mind which he said inference is going to be a billion times larger than training and if you kind of double click into that these old systems aren't going to be sufficient enough right if you're going to have that much

19:06 more demand that much more workload which I think we all agree then H how is it that these old systems which are being decommissioned from training are going to be sufficient so I think that's where that argument didn't hold just didn't hold strong enough for me if that grows as fast as he says it is as fast as you know you guys have seen it in their numbers then it's going to be a lot more net new inference related uh

19:30 you know deployments and there I don't think that that argument holds on the the transfer from older Hardware to newer Hardware well let's you said you said something pretty casually there right let's underscore this right we were talking about the strawberry and the 01 preview and he said there's a whole new Vector of scaling intelligence inference time reasoning right that's not going to be single shot but it's going to be lots of agent to agent interactions thinking time as noan brown

20:00 likes to say right and he said as a consequence of that inference is going to 100xx a million x maybe even a billion x and that in and of itself right to me was you know kind of a wow moment 40% of their revenues are already inference and I said over time does your inference become a higher percentage of your Revenue mix and he said of course right but again I think conventional wisdom is all around the size of

20:29 clusters and the size of training and if if models don't keep getting bigger then their relevance will dissipate but he's basically saying every single workload is going to benefit from acceleration right it's going to be an inference workload and the number of inference interactions is going to explode higher yeah one one technical detail which is you need bigger clusters if you're training bigger models but if you're running bigger models you don't need bigger clusters it can be distribut it

20:57 can be distributed right and so I think what we're going to see here is that the larger clusters will continue to get deployed and as Bill said they'll get deployed for folks maybe a limited number of folks that need to deploy it for hundred billion dollar runs or even bigger than that but you'll see inference uh clusters be large but not as large as the training clusters and be a lot more distributed because you don't need it to be all in the same place and

21:22 I think that's what'll be really interesting it was interesting he um he simplified it even more than you did there he said think about a human how much time do you spend learning versus doing and he he used that analogy as to why this was going to be so great but I in a little different way than than sunny I thought the argument that the reason we're going to be great at inference is because there's so much of

21:48 our old stuff laying around wasn't super solid and in other words what if some other company Sunny's or or or some other when um decided to optimize imprint it wasn't an argument for optimization it was an argument for cost Advantage um because it might be fully distributed or whatever and and of course if if if if you had maybe poked him back on that he might have had another answer about why for optimization but but there are clearly

22:19 going to be people whether it's you know other chips companies some of these accelerator companies they're going to be people working on inference optimization which may include Edge techniques I think some of the accelerators may look like AIC cdns you know if you will and they're going to be buying stuff closer to the customer so it that all TBD but but the just the argument that you've got it left over didn't seem super solid to me and the

22:46 three fastest companies in inference right now are not Nvidia right so who are they Sony show we'll we'll post the leaderboard yeah it's you it's a combination of grock cerebrus and Sova right those are three companies that are not Nvidia that are on the leaderboards of all the models that they run you're talking about performance performance is what performance yeah yeah yeah and I I would argue even price yeah and make the argument why why are they faster why are

23:15 they cheaper in your mind but yet notwithstanding that fact Nvidia is going to do let's call it 50 or 60 billion of inference this year um and these companies are you know still just getting started right why is their inference business is just because installed base yeah I think it's a combination of installed base and I think it's because that inference Market is growing so incredibly fast I think if you're making this decision even 18 months ago it would be a really

23:43 difficult decision to buy any of those three companies because your primary workload was training and you know the first part of this pod we talked about how they have such a strong tie-in integration to getting training done properly I think when it comes to inference you can see all the non- envidia folks can get the models up and running right away there is no tie in t Cuda that's required to go faster that's required to get the models running right

24:06 obviously none of the three companies run Cuda and so that moat doesn't exist around inference yeah cud Cuda is less relevant an inference that's another point you know worth making but but I wanted to say one other thing to what Sunny just said if you go back to the the early internet days and this just an argument that optimization takes a while all all of the startups were running on Oracle and Sun every single one of them were running on Oracle and sun

24:35 and five years later they were all running on Linux and MyQ like in five years yeah and so and it was it was literally it went from 100% to to 3% not and I'm not making the I'm not making that projection that that's going to happen here but you did have a wholesale shift as the industry you know they went from developing and and building it for the first time to optimizing which are really two separate motions it seems to

25:04 me I pulled up this chart right that we shared we made Bill way earlier this year for the pod which showed the trillion dollars of new AI workloads expected over the next four to five years and the trillion dollars of effectively data center replacement and I just wanted to get his updated you know kind of reaction or forecast now that he's had you know six more months to think about whether or not you you know he thinks that's achievable and

25:32 what I heard him say was yes the data center replacement is going to look exactly you know like that of course he's just making his his best educated guess um but he seemed to suggest that the AI workloads could be even bigger right like that once he saw strawberry and 01 that he thought that you know the amount of compute that was going to be required to power this and you know the more people I talk to the more I you

25:55 know I get that same sense there is this insatiable demand so maybe we just touch on this you know he goes on CNBC and he says the demand is insane right and I kept trying to push on that I I was like you know yeah but what about mtia what about custom you know inference what about uh all these other factors what if models sto getting so big I said well any of that change the equation and he consistently pushed

26:24 back and said you still don't understand the amount of demand in the world because all compute is changing right I thought he had a you know one Nuance that answer which was when you asked him that he said um look if you have to replace some amount of infrastructure you know whatever the number was was really big and you're you're part of that and you're a CIO somewhere task with doing this what are you going to do

26:50 what what are you going to replace it with it's accelerated compute and then immediately once you make that choice because you're not going to traditional comp compute then Nvidia is your number one choice so I thought he kind of tied that back together in that like are you really going to you know get yourself in trouble by having something else there or you just gonna go to Nvidia the yeah it to when he said it I didn't want to

27:12 say that bill but it felt like the old IBM argument yeah look I mean one thing Brad is this company's public when a private company says oh the demand's insane I I you know I I immediately get skeptical this company's doing 30 billion a quarter growing 122% like it the demand is insane like we we can see it there there's no doubt about it and and part of that demand was a conversation about Elon and x. and what they did and I thought it

27:43 was also just incredibly fascinating right I I thought it was funny I asked him a question about the dinner that he and Elon and and Larry Ellison apparently had and he's like you know just just because that dinner occurred and they ended up with 100,000 h100s don't necessar you know connect the dots um but listen he confirmed that his mind was blown by Elon and he said he is an N of one superhuman that could possibly pull off

28:13 that could energize a data center that could liquid cool a Data Center and he said what would take somebody else years to get permitted to get energized to get liquid cooled to get stood up that x. a did in 19 days you know and you could just tell the immense respect that he had for Elon it's clear you know he said it's the single largest coherent supercomputer in the world today that it's going to get bigger and if you

28:42 believe that the future of AI is tied closely together with the systems engineering on the hardware side you know what hit me in that moment was that's a huge huge Advantage for Elon yeah I think he I forget the exact number but like he talked about how many thousands of miles of cabling that were just in there um as part of the task um look you coming to it from a you know doing a lot of that ourselves right now

29:11 building data centers standing them up racking and stacking you know our nodes it's impressive it's impressive to do something at that scale in 19 days you know it doesn't even include how quickly they built that data center I think it's all happened you know within 2024 and so um that's part of the advantage the interesting thing there is he didn't touch on it as much as what when he talked about it doing the integration with cloud service providers

29:39 what I'd love to kind of double click into is because you know Elon is in a unique situation where he's obviously bought this cluster he has a ton of respect for NVIDIA but he you know is building his own chip building their own clusters with Tesla so I wonder how much um you know crosscorrelation or information there is for them for them to be able to do that at scale and you know you guys look at this what what

30:01 have you kind of seen on their clusters I don't really have a lot of data on uh on the non- envidia Clusters that they have I'm sure Freda on my team does I just I just don't have it off the top of my head if we have it I'll you know I'll pull a chart and I'll show it sun you you said you now think the xai cluster is the largest Nvidia cluster alive today I'm saying because I I I

30:22 believe Jensen said it in the Pod that he said it's the largest supercomputer in the world yeah I mean I I just want to spend 30 seconds on what you said Brad about Elon I'm steering out my window at the Giga Factory in Austin that was also built in record time star Link's insane when we were walking in Diablo I just kept thinking you know who I'd love to reimagine this placeon right and I I don't the world

30:48 should study how he can do infrastructure fast because if that could be cloned it would be so valuable not really relevant to this podcast but worth noting the the other thing that I thought about on the Elon thing and this this also were these pieces coming together my mind about these large clusters and how important that was to Nvidia he got allocation right this is supposed to be like the hottest company the hottest product backed you know

31:20 backed up for years on demand and he walks in and takes what a quate sounds looks like about 10% of the quarters availability and in my mind I'm thinking that's because hey if there's another company that's going to develop these big ones abolutely I'm I'm gonna let them to the front of the line and that speaks to what's happening in Malaysia and the Middle East and any one of these people that are going to get excited

31:48 he's gonna he's going to spend time with them put them at the front of the line you know what I tell you you know I pushed him on this I said you know elon's gonna you know rumor is he's going to get another 100,000 you know uh H2 200s add them to this cluster I said are we already at the phase of two and 300,000 cluster scale and he said yes and then I said and will we go to

32:12 500,000 a million and he's like yes now I think these things Bill are already being planned and built and what he said is beyond that beyond that he said you start bumping up against the limitations of base power like can you find something that can be energized to power a single cluster and he said we're going to have to develop distributed training and he said but just like with Megatron that we developed to allow to occur what

32:44 is occurring today we're working on the distributed stuff because we know we're going to have to decompose these clusters at some point in order to continue scaling them you may also be running up against the even for the mag the size of Capo X deployment where their CFOs start to talk at higher levels for sure totally and and there's a super interesting article in the information just now where it came out today where Sam Alman is questioning

33:15 whether Microsoft's willing to to put up the money and build a cluster and it may have been um that may have been kind of triggered by elon's comments and or elon's willingness to do it at x. yeah what I will say on like the size of the models like we're going to push into this really interesting realm where obviously we can have bigger and bigger training clusters that naturally imposes that the models are bigger and bigger

33:40 but what you can't do is you can't take a single like you can train a model across a distributed site and it may just take you you know a month longer because you have to move traffic around and so instead of taking three months takes you four months but you can't really run a model across a distributed site because that inference is a in like real thing and so we do you know we're not pushing it there but when you start

34:01 to get to models that become way too big to run in single locations that may be a problem that we want to be aware of and we want to keep on thep in our minds as well on this question of scaling you know uh our way to intelligence one of the things I asked noan Brown uh you know today in our fireside chat he made very clear his perspective although he's working on inference time reing which is

34:26 a totally different vector and a breakthrough Vector at open AI um which we ought to spend a little bit of time talking about he said you know now there are these two vectors right that again are multiplicative in terms of the path to AGI he's like make no mistake about it like we're still seeing big advantages to scaling bigger models right we have the data we have the synthetic data you know we're going to build those bigger

34:51 models and we have an economic engine that can fund it right don't forget this company you know is has over 4 billion in Revenue scaling probably most people think to 10 billion plus in Revenue over the course of the next year they just raised 6 A5 billion they got a $4 billion line of credit uh from City groups so among the independent players bill right like Microsoft can choose whether or not they're going to fund it

35:14 but I don't think it's a question of whether or not they're going to have the funding at this point they've achieved escape velocity I think for a lot of the other independent players there's a real question whether they have the economic mod uh model to continue to fund the acity so they have to find a proxy because I don't think a lot of venture capitalists are going to write multi-billion dollar checks into the players that haven't yet caught

35:35 lightening in a bottle that would be that would that would be my guess I mean you know um I just think it's hard you know listen at the end of the day we're economic animals um you know and I've said before you know if you look at the forward multiple most of us underwrote two on open AI it was about 15 times forward earnings right if chat gptt wasn't doing what it was doing if the revenue wasn't doing what it it was

35:57 doing right this would have been massively dilutive to the company it would have been very hard to raise the money I think if mrr or all these other companies want to raise that money I think it'd be very difficult but you know you never I mean you know there's still a lot of money out there so it's possible but I think this is you you should 15 times earnings I think you meant revenue or 15 times uh revenue for

36:18 sure which I which I said you know when Google went public it was about 13 or 14 times revenue and and meta was like 13 or 14 times Revenue so I do think we're on the on the precipice of a lot of this consolidation among the new entrance what I think is so interesting about X is you know when I was pushing him on this model consolidation pushing Jensen on it he was like listen with Elon you

36:41 have somebody with the ambition with the capability with the knowhow with the money right with the brands with the businesses so I I think a lot of times when we're talking about AI today we often times talk about open AI but a lot of people quickly then go into all of the other model companies I think X is often left out of the conversation and one of the thing that I things I took away from this conversation with Jensen

37:05 is again if the I I if scaling these data centers is a key competitive advantage to winning an AI right you you like you absolutely cannot count out x. a in this battle um they're certainly going to have to figure out you know something with the consumer that's going to have a flywheel like chat chpt or something with the Enterprise but in terms of standing it up building the model having the compute I think they're uh uh you know going to be one of the

37:32 three or four in the game you um you you touched on maybe wanting to close out on the on the strawberry like models you know one one thing we don't have exposure to um but we can guess at is cost and that chart that they showed when they released strawberry the xaxis was logarithmic so the cost of a search with with the uh with the new preview model um it's probably costing them 20x

38:05 or 30X what it does to do a normal check GPT search and so which I think is fractions of a penny but figuring out which and it also takes longer so figuring out which problems um it's acceptable and Jensen gave a few examples for it to take more time and cost more and to get the cost benefit right for that type of result is something we're going to have to figure out like which problems tilt to that

38:31 place right and you know the one thing I feel good about there and again um I'm speculating I don't have information from open AI on this but what we know is that the cost of inference has fallen by 90% over the course of last year what we you know what Sunny has told us and other people uh you know in the field have told us that inference is going to drop by another 90% uh over the cost of

38:53 the next uh you you know period of months if you're if you're racing logarithm needs that's you're going to need that to happen right and I you know and and and here's what I also think happens bill is in this chain of reasoning you're going to build intelligence into the chain of reasoning right so that you know you're going to optimize where you send these you know each of these inference interactions you're going to batch them you're going

39:19 to take more time with because it's just a time money tradeoff right at the end of the day I also think that we're in the very earliest Innings as to how we're going to think about pricing these models right so if we think about this in terms of systems one systems 2 level thinking right systems one being you know what's the capital of France right you're going to be able to do that for fractions of a penny using pretty simple

39:45 models on on on on chat GPT right when you want to do something more complex if you're a scientist and you want to use 01 as your research partner right you may end up paying it by the hour and relative to the cost of an actual research partner it may be really cheap right so I think there going to be consumption models um you know for this like I think we we haven't even scratched the surface to think about how

40:08 that's going to be priced but I totally agree with you um that it's going to be priced very differently again I think this puts I think open AI has suggested you know that the 01 full model may even be released yet this year right I one of the things I'm kind of waiting to see is I think you know listen having known noan Brown for quite a while now he's an N of one right and he wasn't the only

40:34 one working on this for sure at open AI um but you know listen whether it's pbus or or or winning at the game of diplomas he's been thinking about this for a decade right it was his major breakthrough on how to win the game of six-handed poker and so he brought this to open AI I think they have a real lead here which which leads me back to this question you and I talk about all the

40:58 time which is memory and actions right and so I I I have to tell you this funny thing that occurred at at our investor day so I had nikesh on stage and you know obviously Nikes you know was instrumental at Google for a decade and so I wanted to talk to him about both consumer AI as well as Enterprise Ai and I asked him I said I want to make a wager with you I knew of course he would

41:22 take a bet um and I said I want to make a wager with you over under I'll set the line at 2 years until we have an agent that has memory and can take action and the canonical use case of course that I used was that I could tell my agent book me the Mercer hotel next Tuesday in New York at the lowest price and I said over under you know two years on getting that done I said I'll start 5,000 bucks I'll

41:50 take the under he snap calls me he says I'll take the over and he said but only if you 10x the the BET and of course we're doing it for we're doing it for a good cause uh so I had to call him because I you know I I I can't I can't not step up to a good cause um so we're we're taking the opposite sides of that trade now what was interesting is over the course of the next couple days I

42:14 asked some other friends who took the stage you know uh where they would come down on uh on the same BET right our friend Stanley Tang took the under um a friend from Apple uh who will remain nameless kind of took the over and then noan Brown who was there pleaded the fifth he says I know the answer so I can't say um and so yeah I was kind of provocative and I I you know I texted

42:42 Nash and I said I think you better get your checkbook ready um you know uh so coming back to that bill you know strawberry 01 is an incredible breakthrough something that thinks of this whole new Vector of of of of intelligence but but it kind of makes us forget about the thing you and I focus so much on which was memory and actions right and I think that we are on the real preus of not only these models

43:08 think you know can spend more time thinking not only can they give us less hallucinations uh you know and just scaled compute but I also think I mean you already see the makings of this I mean use these things today they already remember quite a bit um so I think they're they're sliding this into the experience um but I think we're going to have the ability to take simple actions and I think this metaphor that people

43:31 had in their minds that they were going to have to build deep apis and deep Integrations to everybody I don't think is the way this is going to play out and let me just what do you think it's going to play out well I mean the Easter egg that I thought got dropped last week is they did this event on you know their voice API right and it's literally your GPT calling a human on the telephone and

43:54 placing an order so why the hell can't my GPT just call up the Mercer hotel and say Brad gersner would like to make a reservation here's his credit card number and pass along the information there is a reason for that I mean look scrapers and and form fillers have existed for how long Sunny 15 years like you you could write an agent to go fill out and book at to Mercer hotel 15 years ago there's nothing impossible about

44:22 that it's the corner cases and like the hallucination when you're credit card gets charged 10 grand like you you you just can't have failure and how you architect this so that there's not failure and there's trust I'm sure you could demo this tomorrow I I have zero doubt you could demo it tomorrow could you provide it at scale in a trustworthy way where people are allocating their credit cards to it that might take a little longer okay so over under Bill on

44:50 two years I mean I I'm gonna get I'm G to get your action either way but what's the test the demo I think you can do it today no not this not the cheesy demo you just said I'm talking about a release that allows me you know at scale to uh to book a your credit card and not just you but everybody full release well yeah we'll call it a full release just because I know that's the only way I can

45:12 t you to take the BET which today is October 8th 2024 I mean Sun you're already know what he's gonna say where you'll take the over right bill yeah yes okay so Bill's in cash Camp Sunny where do you come down over under on two years no don't start hedging Bill don't start hedging go I already said it demo today is 15 years ago you let me let me comment on your what you're worried about Bill and

45:39 I think people still are still working their way through it you don't need a single agent right now to book the Mercer and deal with all the scraping stuff you're talking about you can have a thousand agents working together you can have one that's making sure that the credit card charge is not too big you can have another one to make sure that the address is right you can have another one checking calendar and so all

45:58 of that's free so I'm on the under and Brad I'll even go under one year wow wow yeah W so we got a little side action you and I sunny I don't I'm not going to go under a year but I I think it could we could have limited releases in a year but Sunny you and I now have action with Bill what do you want Bill a thousand bucks sure to a good cause Okay thousand

46:20 bucks each to a good cause and uh I'll just assume Sunny that we'll get action from NES as well and you know our friend Stanley Tang is definitely in the tank for some so we're going we're going to give uh some good money to a good cause and listen I think this is the trillion dooll question I know we're all focused on you know scaling models and I know we're all focused on the compute layer but what really transforms people's

46:43 lives what really disrupts 10 Blue Links right what really disrupts the entire architecture of the app ecos system right is that when we have an intelligent assistant that we can interact with it gets smarter over time that has memory and could take actions and when I see the combination of advanced voice mode voice too API strawberry 01 thinking combined with scaling intelligence I just think this is going to go a lot faster than most of us think now listen they may pull on the

47:14 RS right they may Rel slow down the release schedule in order you know for a lot of business reasons that's harder uh to predict but I think the technology I mean even no I'm said I thought it was going to take us much much longer to see the results that we have seen can I hit on one other thing this is you know we started the Pod a little bit talking about it I just want to get your

47:36 impression bill this idea that Jensen can scale the business two or three times with you know increasing the headcount by you know 20 or 25% right we know that meta has done that over the course of the last two years and you and I've talked about are we on the eve of just massive productivity boom and massive margin expansion um like we've never seen before right Nesh said we ought to be able to get 20 or 30% you know

48:05 productivity gains out of everybody in the business first of all I think Nvidia is a very special company and it's a company that's it's even even if it's a systems company it's an IP company and the demand is growing at such a rate that they don't need more designers or more developer Engineers to create incremental Revenue that's happening on its own and so their operating margins are are record levels um for the majority of companies you know I've

48:36 always just held this belief that you know you evolve with your tools and the real the real answer is the companies that don't deploy these things are going to go out of business yeah and so I think margins get competed away in many many cases I I I think it's ridiculous to imagine oh every company goes to 60% operating no no no no I mean listen Delta Airlines is going to do all of these things with AI and immediately

49:05 because it's in a commodity Market it'll get competed Away by Southwest and United bad Industries remain bad Industries yeah yeah yeah so so but but there might be some you know that that figure it out and and the I have another theory that uh that always keep in mind which is hyper growth tends to delay what you learned in microeconomics class you know I I remember when I was a PC analyst and there were five public PC

49:32 companies all growing 100% And so in in moments of hyper growth you will have margins that may or may not be durable um and you'll have a number of participants in a market that may or may not be durable um during periods of hyper growth I have two more things on my mind Sunday do you do you have any reactions to that I mean I I just have to get to a couple of these topics no I

49:56 like this going to be a Lex Freeman link podcast once you the no I I look I I really you know been thinking a lot about Jensen's point in the Pod about you know how much AI they're using internally for design design verification for all those pieces right and I think you know it's not 30% I I actually think um sort of that's an underestimate I I think you're talking you know multiple hundreds of percent Improvement in productivity gains and

50:26 the only issue is that not every company can grasp that that quickly and so um you know I think he was kind of holding some cards back at that point when he made that comment and it really got me thinking about like how much are they doing there that they don't want everybody to know about and you kind of see it now in the model development because they you know if you've noticed the last couple of weeks they put some

50:49 models out there that are models trained on their own and they don't get as much um noise as you know ones from meta and um you know the other players that are out there but they're really doing a lot more than we think and they they I think they have their arms around a lot of these very very difficult problems Brad why did they put their own model out well it's related to to this topic of

51:11 open versus close so um Bill you know I hope you're proud of me you know I went back and I said I have I have to ask this question right and you know I thought Jensen you know I thought he gave a great answer which is like listen we're going to have company that for economic reasons right push the boundary toward AGI or whatever they're doing and it makes sense to have a closed model that can be the best and they can

51:35 monetize but the world's not going to develop with just closed models we're going to you know he's like it's both open and closed and you know he said because open he's like it's absolutely a condition required it's going to be the vast majority of the models in the industry he's right now if we didn't have open source how would you have all these different fields in science you know be able to be activated on AI he

51:57 talked about llama models exploding higher and then with respect to his own open source model which I thought was really interesting he said we focused on right something that a specific capability and the capability that we were focused on is how to agentically use this model to make your model smarter faster right so it's almost like a training coaching model that he built and so I think for them it makes perfect sense why that may they may uh you know

52:27 put that out into the world um but I also you know a lot of times the open versus closed debate you know gets hijacked into this conversation about Safety and Security and you know and I think he said you know listen these two things are related but they're not the same thing you know one of the things he he commented on that is just he said there's so much coordination going on on the Safety and Security level like we

52:50 have so many agents and so much activity going on on making sure you know just look at what meta is doing uh you know on this he's like I think that's one thing that's under celebrated that even in the absence of any you know platonic Guardian sort of Regulation right without any top down you already have an extraordinary amount of of effort going in by all of these companies into AI uh Safety and Security that I thought was

53:19 uh I thought was a really important comment thanks for jumping in guys kicking this one around it was a special one to yeah congrats on having that opportunity that's pretty that's pretty unique and uh and now we got a little wager so I I mean listen I am so looking forward to like doing a live booking at the Mercer on the Pod right and then Sunny we can just drop the money from the sky we can just collect we can just

53:47 collect exactly exactly good to see you guys we'll talk soon all right peace take care [Music] as a reminder to everybody just our opinions not investment advice

Summary

The podcast discusses insights from a recent interview with Jensen Huang, CEO of Nvidia, focusing on the company's strategy and advancements in AI and accelerated computing. Key themes include Nvidia's positioning as a systems-level company rather than just a GPU manufacturer, the importance of scaling intelligence, and the competitive landscape in AI inference and training.

- Nvidia is not just a GPU company but an accelerated compute company, emphasizing its systems-level advantages.
- The data center is considered the unit of compute, highlighting Nvidia's focus on large-scale deployments.
- Jensen Huang believes Nvidia can triple its revenue with only a 25% increase in headcount due to the use of autonomous agents.
- The company is deeply integrated with cloud service providers and focuses on optimizing workloads across various industries.
- Inference workloads are expected to grow significantly, potentially outpacing training workloads, with Nvidia's existing infrastructure providing a competitive edge.
- The podcast discusses the potential for massive productivity gains in companies using AI, with Nvidia leading the charge.
- There is skepticism about the long-term sustainability of Nvidia's competitive advantages in inference as other companies emerge.
- The conversation touches on the importance of open-source models alongside proprietary ones in advancing AI technology.
© transcribe · For agents Built with care and craft by Gokul Rajaram