transcribe

The Neocloud Boom: State of AI Compute 2026 | Stephen Balaban

The MAD Podcast with Matt Turck · 1h 14m · transcribed Jun 2026
More from The MAD Podcast with Matt Turck Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:00 It's pretty clear that we have an amazing system that can take in money and output software. The people who are the naysayers, you're going to throw these GPUs out in 5 years are completely wrong. They're completely wrong and they've been wrong the entire time. We continue to be generally underbuilding. Most people that are sort of in leadership positions at NeoClouds or within the market have been recognizing this insatiable amount of demand for large language models to do everything from being an assistant to code generation. We continue to see no end to the scaling laws.

0:36 >> Hi, I'm Matt from First Mark. Welcome to the Mad Podcast. My guest today is Steven Balaban, co-founder and CTO of Lambda, one of the top neoclouds powering the AI boom. This episode goes deep on the physical layer that everything else in AI runs on. We get into why GPU compute was never actually a commodity, how you finance billions of dollars of data centers and chips, why a 2023 H100 can be more expensive to list today than when it was bought, and what it actually takes to stand up a gigawatt scale AI factory. We also cover Lambda's wild origin story. From a facial recognition startup to a baseball cap with a camera in it to a near billion dollar cloud business today. Please enjoy this amazing and very educational conversation with Steven.

1:21 There was a moment in time in Silicon Valley a few years ago. If you had asked most people, they would have said that Neoclouds were going to be a commodity in particular because GPU compute was going to get commoditized. Uh and if you fast forward to today, it seems to be exactly the opposite. Both Lambda uh but several of your competitors seems to be uh absolutely ripping. So what is it that uh naysayers got wrong then and continue to get wrong today?

1:49 >> The big thing is that cloud compute is not a commodity service. It is a very complicated, highly vertically integrated type of service that spans everything from land land entitlement, construction, HPC, high performance computing, design, software, virtualization, cloud services on top. And there's a reason why the biggest companies in the world, these multi-t trillion dollar market cap businesses, whether it's Amazon, Microsoft, Google, Oracle, are all in the cloud computing business is because it's a great business. And so I think that's like probably the fundamental thing that was misunderstood is that, oh, this is somehow a little bit different than a normal cloud service. Um, but really what it was was it's a cloud service designed for the age of AI. But there is some element of commoditization right the price of rental of a GPU is going down but uh what you're saying is that to some extent it doesn't matter because it's only one layer of the cake.

2:57 >> Yeah. So when you look at for example I I think it's like actually worth doing is to try to like kind of dig into some of the methodology on for example an index like there's there's the there's the index that's on Bloomberg for H100 rental prices. And what we're actually seeing in the market is that first of all there's two different rates. There's a public cloud on demand rate and then there's a long-term rental rate. And I think that some of these in indices don't properly take that into account because what we're actually seeing is a very consistent if not increasing long-term rental rate and very consistent and increasing ondemand rental rates. And so what happens is if if if the index mix for example if the methodology in the index biases towards long-term contracts being a bigger part of the volume that will look like a decline in the index when the reality is it's just a decline in the mix that the index is meth you know the index is covering.

4:02 >> Fascinating. So I'm curious about your thoughts as a key leading player in the Neo cloud ecosystem about how you see the market evolve. How much of uh the competitive advantage that uh you guys are building and other players are building is based on technology versus financing uh race. >> There's a few different layers on it which is there's a lot of differentiation and work that's being put into for example the cloud software orchestration layer which allows us to for example take a very large scale GPU cluster and partition it up for our customers. So we've got, for example, our one-click cluster product that allows us to do that. And that's something that's like quite unique in the Neocloud space. Most of the other Neoclouds either don't have the ability to launch a cluster from their website or max it out at say 32 GPUs. Whereas Lambda's designed a piece of software that allows us to give you anywhere from 16 up to, you know, 4,000 GPUs in a web interface. And um then there's innovation on the data center construction and design side of things which is also really important right because that's like the physical layer um underneath the high performance computing equipment and you know we're working on a lot of different ways to dramatically reduce the time it takes to construct and stand up new megawws. And then there's as you mentioned innovation on the finance side of things where you know we're coming up with new and unique ways to uh finance underwrite package these uh these these largecale capital projects really and um so I think it's like innovation's happening on every layer of the stack and it's a it's a very complex coordination style business.

5:57 >> Yeah. And do you think that ultimately the NeoCloud ecosystem becomes a winner take all or is there room for multiple very large players? No, >> I I think it I think it's absolutely room for multiple very large players just like the uh traditional cloud business has shown that there's room for multiple large winners and multiple large players. And I think that the fundamental reason for that kind of going back to well what drives I guess market structure and I'd say generally speaking when you have an industry that is um has like technology modes and capital formation modes and economic modes that tends to be oligopolistic in its market structure. When you have markets that have more um sort of network effect modes, uh those tend to be a little bit more, you know, single winner take all.

6:48 >> What are the various scenarios uh in your head as you think about the the future about how it all play out? Are we are we overbuilding? Are we underbuilding? Nobody knows. How do you think about it? >> Well, I I think that we continue to be generally underbuilding. Um and and the most people that are sort of in leadership positions at NeoClouds or within the market have been recognizing this, you know, sort of insatiable amount of demand for large language models to do everything from, you know, being an assistant to code generation.

7:26 you know, you you can kind of look back to some of the talks that I've that I've given in the past around I kind of called, hey, in a couple of, you know, months to years, we're going to be at a point in time where you can put money in and get software out the other end. And now at that at that point in time when you were predicting when I was predicting that, it was maybe not w not as widely held of a belief. Um but now with you know let's say the release of Opus 45 I think it's pretty clear that we have an amazing system that can take in money and output software and the I think the the part which makes me feel so confident that there's going to continue to be demand is that we continue to see no end to the scaling laws which are like the underlying idea that you put more compute in and you get better intelligence levels out of your models. You know, as you increase the capacity of the model and train it with more compute, train it with more data, you get more intelligence out. And as long as that continues to hold, I think that we still have in store for us, it's hard to predict exactly when scaling laws might start to uh reach sort of a diminishing marginal return type of part of the curve. But um right now it's very clear that we're going to continue to see more and more and more capable models that is kind of expanding the cone of the addressable market, right?

8:59 Like originally the cone of the addressable market was all right, this is going to be helpful for customer support. It's a sort of substitute good for Google search and for other search um online. And then now it's like well this is a substitute for a lot of software engineering roles or a huge augment to software engineering roles. And so as that cone expands the total market and the demand for compute expands and I think that we're continuing to underestimate it.

9:26 >> Do you worry about model training and model inference becoming I don't know 10x more compute efficient and what that would mean in terms of the buildup. I think that generally speaking what you're seeing is that if if let's say you do become 10 times more efficient, I think that that just means that everybody is able to process 10 times more tokens and there's there's still the same fixed amount of compute in the world at any given point in time. And so in the early days, it's funny, we used to talk a lot about this back in let's say 2017. Oh, well, maybe there's going to be some new type of model, let's say, that will look more like a random forest model, which the audience might some some members of the audience might know.

10:05 You can kind of train a random forest model on a MacBook, right? And there was there was this concern that was kind of persistently raised around like, well, okay, what happens if you have this sort of like um adjacent disruption on the model side of things? And so far, we haven't seen that. And again, everything that we're building towards is sort of based on these scaling laws, which is really about scaling up this architecture. So, um I don't really foresee a very likely outcome where we have this huge model disruption that would cause a a decline in the demand for compute.

10:44 >> Where's the main bottleneck these days that you're experiencing building lambda labs? Is that GPU, power, electricity? So um I always say that bottlenecks are always like kind of local before they're global in terms of you know one one development might be bottlenecked on let's say generators or on UPS systems as a function of like the sort of idiosyncrasies of the site but broadly in the industry the thing that is the main bottleneck is basically land powered shell which is basically land that is entitled to have a certain amount of megawatt commitment from a utility and um then of course the data center and the mechanical, electrical and plumbing equipment, the MEP equipment that goes into that data center. Um and so that's the main bottleneck that we're seeing in the industry right now I'd say across the board.

11:38 >> How real is the movement against data centers from the the global community and how do you think about uh how to respond to it? >> Well, it's certainly it's like very popular in the news right now. I'd say that um it's definitely very real. I mean, I think that rightfully communities that host any type of large capital project, whether it's a power plant or a uh solar farm or a data center or a distribution center, right?

12:11 Those communities want to have a seat at the table. I'd say in general though, I spend a lot of time reading through a lot of the comments from communities and people want jobs. They want tax revenue. Any major capital development is going to bring a lot of tax revenue and it's going to bring a lot of jobs and it's going to bring investment into their community. And what they really are voicing I think is one is having a seat at the table while while this stuff is you know being developed. I think that's an important thing just to have their voices heard and that that the developers coming in and actually understanding the community. The other thing to kind of I think keep in mind is that there's a lot of misinformation out there. So for example, every single modern deployment of let's say a Blackwell class or a uh Reuben class GPU, you know, the VR GBN VR GPUs, um these are oftent times in a closed directtochip liquid cooling system that's connected to a dry cooler, which means that there's almost zero evaporation. It's not using evaporative cooling. It's using a dry cooler system that does not consume a lot of water.

13:27 Um, and on top of that, most of these data center developments are bringing a ton of power to the grid. They're either standing up behind the meter power. They're standing up and bringing battery electric storage systems to the grid. And they're bringing all these like sort of ancillary benefits that strengthen and fortify the grid and also, you know, eventually in the long term will maintain the costs that are being experienced by the community. And so I actually think that there's a very clear path towards um you know maybe spreading more of that fact around what does a data center bring um because there's just a lot of misinformation. You'll see people talking about how data centers consume a lot of water. Well, an evaporative cooling tower might evaporate a lot of water, but practically no new builds in the United States are using evaporative cooling uh for doing the these closed loop direct to chip liquid cooling systems.

14:23 >> Do you think we do a terrible job as an industry explaining this to the broader world because like those things keep coming back and they seem to be accelerating but then when you have the discussion there from a technical standpoint a lot of this is just simply based on museum information as you just said. I think that everybody's trying to get better at that kind of communication. Um, and it just takes some clear thinking, writing down what are the benefits, writing down what are the costs and presenting that clearly and plainly to a community so they can make a good decision about, you know, what kind of jobs and what kind of development they want in their communities.

15:00 >> Let's open the the hood for a minute. People talk about things like flops and GPU hours and tokens and MFU. what what is the the best way to think about a a compute unit? >> Yeah, it's interesting. You know, you said a few different terms and I always like to kind of break it down from like a physics perspective into like the SI terms. So okay on the the left hand side is all of the energy production and then on the you know my right hand side is sort of tokens being consumed by by by somebody and um you know maybe you can even have the application layer on on the far right of that that's using the token. So on the left hand side you've got either photons coming in per second or molecules of natural gas coming in per second and then that through a power plant or a solar farm gets converted into jewels per second which is um a measure of electrical power production and then the jewels per second obviously in engines there's a level of efficiency and that's a engine efficiency. It's interesting because like the MFU percentage is kind of like an efficiency up on the higher end of that chain. The power plant or the solar plant then converts that into jewels per second which is watts which is consumed by the entire data center. the data center itself um you know needs to cool itself and that's the PUE and that's actually the the efficiency metric that you can use to measure a data center on and then you put the servers and all the different networking and storage gear in and that's producing floatingoint operations per second or flops per second. Okay, that is what gets consumed. The flops per second capacity is what gets consumed by let's say a model builder when they're training a model or when they're inferencing a model and that gets turned from flops per second into the tokens per second. Then on top of that tokens per second you might have some level of efficiency that the end customer is actually you know turning those tokens into real actual intelligence. That's like the entire pipeline I I would say from end to end.

17:11 uh super helpful. If two companies have the same chip fundamentally, how do they extract more value from it? What what needs to happen to maximize the usefulness of that chip? >> If you look at the cost structure of let's say one GPU hour of time, you know, we we're talking about H100s. The the largest part of that cost structure is the depreciation that is associated with that GPU hour. And um basically you can think of a utilization metric as being like kind of a multiplicative factor on that. So one over the utilization. So if you if you use your capital asset 50% of the time, you will have on a per hour basis twice one over 0.5 the amount of perh depreciation expense associated with that. And so I think that the number one way that companies are, you know, sort of gaining a unique advantage is well, how can I build a cloud product that is beloved by people that is going to drive a high utilization and um you know in addition to that the market as we mentioned earlier for ondemand compute basically The retail pricing is obviously much higher than the wholesale pricing. So the retail is like on demand, spin up a GPU, spin down a GPU, normal cloud service. The wholesale is sort of buying 10,000 GPUs for 5 years for example. And so one of the things that we do at Lambda is really try to figure out, hey, how can we sort of get the most dollar utilization and percentage utilization out of the capital deployments that we do and that's that's by making great cloud software that makes it easy for somebody to spin it up and down. So for example, if you don't have that cloud software, you can't rent you can't extract a retail pricing, right? you know, you cannot rent it out to somebody for an hour because you just simply don't have the means to be able to do that. And actually, a lot of NeoClouds are in that position where they they don't even have the infrastructure to be able to run a real cloud service.

19:25 >> So, you have GPUs, but like a big part of how those data centers work is transforming GPUs into networks of GPUs. Do you want to explain at a high level how that works? The general idea is that you've got a large scale high performance computing cluster of a bunch of you know let's say Nvidia GB300 NVL72 racks that's 72 GPUs allorked together via um Nvink and then there's a connection between the racks uh that's either infiniband or you know high-speed Ethernet and um that is a essentially what's called the spine leaf topology which is basically a way to say hey this is a completely non-blocking every port on every GPU can talk with every other GPU in the network it's fully connected and it's uh um able to provide maximum bandwidth between every individual GPU and that cluster is useful for training large models it's also useful for inferencing so frontier inference as we sometimes refer to it at lambda is basically you know very much a distributed inferencing problem where they actually will you know fragment or shard the model there'll be some sort of sharding strategy for the model uh where it can be um essentially run on multiple GPUs and it uses and that high-speed infiniband or Ethernet interconnect to to do that communication >> and so what is a frontier inference is that in France for the most advanced reasoning models like the more demanding.

21:11 >> Yes. >> Well, well, you know, it's not necessarily associated with reasoning models so much as like just a very large frontier model that is, you know, kind of the domain of let's say three companies in the world or four companies in the world when they're doing their inference. It's a very complicated thing that is is fully utilizing all of the interconnection that's available. >> And what you describe for Frontier in France is that conceptually the same thing as what happens for training. This concept of just distributing a task massively across a bunch of GPUs. What happens during a training run from a compute standpoint? Generally speaking, when you're doing a training run, you might think there might be some sort of split between the backwards pass and the forward pass on the model. And the backwards pass might be, let's say, 2/3 or more of the compute. And the forward pass, which is basically the same thing as inferencing, uh, is, you know, the remainder. And one of the realizations that I think has been made over the last bit of time is that the type of infrastructure that you'd want for uh doing a large scale training run can be reused to do the inferencing of that model and um what I mean by the sort of frontier inference and the fact that the inferencing is being done in a distributed way you know you'll have like a mixture of experts model and there'll different basically sharding strategies for how you put those experts onto different servers and to different GPUs.

22:46 Um, and you know the models can be very large. They may not fit on one single uh rack or you know they may not fit on one single server. They they might they might need to be distributed across different servers to even just do the forward inference pass. And so that's where sort of distributed frontier inference kind of comes into the picture, right? Because like if you're doing a small model, let's say llama that the users might be, you know, familiar with or uh you so some of the quantized small models can fit on a single GPU.

23:19 >> Mhm. >> Okay. Well, let's just say that like Opus and Chad JBD 5.5 can't fit on a single GPU. And when uh we think about compute costs, what what costs the most money? Is that model size? Is that memory bandwidth? Is that latency? Uh does context window and like those very very large context window do they change anything to the compute cost? What what costs the most money? As I mentioned, like the biggest component of the unit cost for a cloud service like this is the depreciation expense. And um within that, you know, is basically some sort of bill of materials for the servers that are in the data center, which is by far and away the biggest portion of the cost. If you talk about the capital stack, let's say you can go back down to power generation,2 to3 million a megawatt.2 to3 billion a gawatt for a power plant. Um the data center is between 10 and 15 billion a gigawatt for um building the data center. And then the compute the servers can be anywhere from 35 to45 billion a gawatt and um within that that so so you can see the server portion is obviously by far and way the largest and that's like a big part of the depreciation expense and then the um within that obviously you have um the sort of server and cluster bill of materials which is prim primarily the GPUs. Um, if you were to kind of break down Nvidia's uh bill of materials, then you know, you can kind of get better allocation towards uh where those costs are coming from. But certainly in the most recent period of time, memory expenses, you know, memory has gone up a lot in price and uh, you know, there's there's very few vendors, right? You know, for HBM memory, but Samsung, Highix.

25:21 >> So, you guys are a big Nvidia shop at a precise level. You mentioned some of the names, but like which chips uh do you uh use mostly? What's your kind of chip stack? >> Yeah, so Lambda really loves Nvidia's products. I mean, they're the the only server provider, the only chip provider that is available in every single major cloud platform, which is a huge platform advantage. And we've stuck with the Nvidia sort of ecosystem for all the chips we've deployed. And we've got everything from V100s, A100s, H100s, H200s, uh, B200's, GH200s or GB200, B300s, and VR200's coming soon. And so we, you know, use everything in the in the in the ecosystem. Do you think that today or in the near future we're going to be in a multi-silicon kind of world? Is there like room for different players beyond Nvidia?

26:28 >> Well, I mean, I think that we're already in a world where there's a huge amount of competition from massive massive multi-trillion dollar companies and they're all trying to fight for the same thing, which is to be the best chip in the world for running and training neural networks. Essentially, Nvidia's built a great product that has gotten a lot of distribution and has a great platform of developers who love what what they do. And you have to take into account not just the cost of the chip, right? The price of the chip is one aspect, but you know, you have to take into account the entire software ecosystem and what's been developed. So, one of the big people talk about like what's Nvidia's moat? One of the big moes they've got is just the CUDNN stack. It's not just CUDA. You know, CUDA is sure that's like the water we all swim in, but like KUDNN has got so many, you know, matrix multiplication routine optimizations baked into it.

27:22 >> What is KUDNN for everyone to understand? Yeah. Okay. So, CUDNN is the it's CUDA deep neural network library and it's basically Nvidia's you can think of it like a highly tuned engine for matrix multiplication. And basically, if you were to just sort of naively implement the matrix multiplication algorithm, you would maybe get a certain level of floating point opt floating points uh per second. But they've gone and tuned every single aspect of it and you know come in and do winrad filtering or you know a bunch of different algorithms that you would apply to speed up matrix multiplication and KUDNN means that you know you don't have to go and do the optimization yourself and so like that's that's one aspect the other one is u nickel NCCL which is their networking optimization library where it will sense the topology and the connected nature of your um network, your Infiniband or your Ethernet network and it will suggest an optimized sort of routine for doing uh you know reduce all and broadcast the different what what are called open MPI primitives which are used for that sharding that we were talking about for both training and for inference. And so that's like the kind of software stack that I think really is hard for a lot of the new entrance in the chip space to overcome. I think we're already, like I said, we're already in a world where there are multiple options for silicon.

28:53 You know, the biggest labs in the world are using multiple different types of chips to do their inferencing and training on. >> What would be a plain English definition? We talked about the chips, but like the rest of the stack, the networking and the storage. Just walk us through how it works. when you're running a cloud service, uh, one of the things, you know, you'll you'll train your model or you'll upload your train model and you you're ready to start doing large scale inferencing. Well, you're going to need a place to put your data whether it's the data that you're using to train uh with or whether it's the data that's coming in and streaming in uh from your end customers. And uh so having high-speed storage is like a really important part of it. So, Lambda offers the the basically AI optimized file system service that is um significantly faster than like your standard let's say cloud file system which is maybe more of a traditional NFS type of thing. Um this is like a highly optimized parallel file system that's designed for high performant read and writes and mostly high performance reads. That's like the kind of most of the workload >> and that's something you build inhouse completely. We have I mean well in-house completely right it's like you have to ask the question like what what is the definition of in-house completely right you know like we've never spun a PCB at this company we have not authored for example you know we use KBM Kimu for our virtualization for example right and so um you know what we we have both commodity offtheshelf hardware that has software on installed on top of it for our some of our storage we have some storage partners that we work with as well. But you know generally speaking everything that we do on the cloud I would generally say is is something that is like we we rolled it ourselves with the help of the broader ecosystem because again there's no such thing as rolling it yourself unless you're like you know mining um you know ultra pure silicon from some like you know and then coming up with your own ASML uh you know it's like it's uh it's funny.

31:02 >> Yeah. Yeah, that's the highly optimized uh storage. What else? The networking part and what other pieces? >> Yeah. So, so I I I was talking about this one cloud cluster product that we've got and the the way just for everybody to think about this is like okay well look you've got a bunch of GPUs. Let's say you've got a cluster of 10,000 GPUs. Well, I want to partition that cluster up. And so what it is is it's a bunch of GPUs. some CPU servers as well because you need to have an orchestration uh fleet as well and then you've got some storage and um all of the CPU servers and the storage servers and the GPU servers are interconnected with the the storage so they can quickly read and write from it and um so there's and and that that communication happens over what's called you know the inband network and then there's the compute fabric which is where I was talking about where all of the sort of weights and uh feature activations are being shared uh throughout that compute fabric and then there's an outofband monitoring network where you've got access to whether it's BMC or uh some of your DPUs um and when you are trying to create a subpartition of a 10,000 GPU cluster you need to simultaneously partition the inband the out of band and the compute fabric okay so like that complex coordination between we've got a bunch of bare metal systems to hey we've got a virtualized system that has you know what's called RDMA you know RDMA remote direct memory access that allows them to read and write quickly not just from the disks but from each other's memory the GPU's uh sort of HBM memory and allow them to do that um that sort of direct memory access allowing it to go directly ly from a GPU to another GPU without getting copied to the CPU for example.

33:05 Having that all work is a immense immense software undertaking and um and and and this is going back to the original question like well what are people not getting about Neoclouds? Well, first of all the answer is that most Neoclouds don't have this kind of technology. Most Neoclouds have not made the really it's like kind of high tens to hundreds of millions of dollars of software investment that you need to make to build a real cloud system that can partition a high performance computing environment like this and um so like that is um and then to have it all work with the storage. Anyways, I guess that that sort of summarizes the steps that you need and kind of you can think about all the different moving parts of a modern like how does an AI data center work? People talk like AI data center, but really you have to kind of go down that one next level down, which is because if you were to, by the way, if you were to ask an AI data center landlord, a traditional one, >> yeah, >> what what's going on inside of the data center, they'd be like, well, look, we're real estate people, and you know, we well, you know, we we really outsource this to the GC, but like the GC doesn't know, of course, anything that's going inside. And this is then they it's their tenants who know. So this is what's actually happening inside of an AI data center and then it serves the the result like also going back to the community stuff. If people knew a lot more about like well this AI data center is actually just serving the chat GBT request that I'm that I'm giving it, right? Like sometimes they don't even realize that that's actually what an AI data center does.

34:46 >> So you mentioned tenants. Do you rent them? Do you also own some building some? and where does that fit in the overall strategy? Yeah. So, um you know, initially we we started off as being primarily a renter and we've actually started to get into the business of uh financing some of them uh the construction of them ourselves as well as you know we're going now into full vertical integration where we are identifying land coming to the table with a basis of design which is basically all the engineering diagrams to construct the data center financing and constructing that data center, putting the servers in, and then associating that with like a a long-term offtake agreement with one of the major uh compute uh consumers in the world and financing it all. So, like we're getting into full vertical integration at Lambda and it's been it's been great because we've been able to kind of again bring that engineering mindset to this problem which was historically mostly run by people in real estate. In your own data centers, are you the sole tenant or is part of the idea that you can also rent some to others?

35:58 >> In a lot of our data centers, we are the sole tenant. In terms of the data centers that we're planning on constructing, we don't yet have any plans to lease that space to others. So, we're not trying to get into the the leasing data center business. Um, maybe that's something that you can imagine down the road. I wouldn't write it out completely, but you know, for now, we have to focus on providing Lambda with the compute that we need to service the market.

36:24 >> How international are you, by the way? >> You know, I'd say that we're very much focused on North America. And so we have data centers in Canada, United States, and Mexico. We're very much, like I say, primarily focused on North America, but really within, you know, that the United States obviously. And um we we haven't had this desire internally to try to go and expand into Europe or or uh too far into Asia. We've done some partnerships with some of our great investors like uh SK Telecom and we have a data center uh that we've operated in Korea in Seoul. And so we have some experience with international um but right now we're just like look let's focus on the US market. It's where the opportunity is.

37:15 >> Do you need for performance reasons to be close to the customer the way you you need to have regions in cloud? >> You know, it's super interesting. A lot I get this question a lot and people they're like well does latency matter does so I'll tell you what what matters and what doesn't matter. You can look at your own utilization of whether it's chatg or claude or grock or gemini and you can see hey a lot of the things that I'm doing I kind of shoot it off I come back later and there's a research report for me maybe it's a longunning agent workflow in those cases latency doesn't matter at all the only thing that matters is your cost per token that's all that matters and um so that that's been a really interesting change. I think that you know the old school traditional legacy cloud business was so latency focused because of some of the applications but this new fleet of AI applications are far less latency sensitive. So that's one. But there is the caveat which is this governance and data governance is becoming an important thing and a lot of countries are wanting to have the AI compute that their citizens are using be run out of their own country so that they can you know at least have their own per their perception of control or whatever and the you know that is that is another that is an element to it but I'd say that from the latency and there's no technical reasons. Let's talk about the financing stack. So presumably it's a commission of equity and and debt. How does it all work?

38:52 >> Yeah. So um the way that it works is that you you know you could really fragment it into these two parts which is like financing your on demand cloud versus financing an offtake agreement which is like a longer term commitment. And on the ondemand cloud you're kind of looking at Lambda's credit quality. on the offtake aaker groom you're kind of looking at the credit quality of the the end customer who's paying the bill and so what you do is you just you know take uh your offtake agreement you take this chunk of GPUs that you're deploying you take a lease or the the property and you kind of put it into a box and you can go to the private credit markets and you can come up with you know an assetbased loan you can you can get a a variety There's a variety of different methodologies uh for financing it. Um most of it is just some sort of like special purpose uh vehicle that's designed to finance this particular deployment with a very known and easy to underwrite which is basically just a fancy way of saying the you know finance uh term for just assessing the risks and the downsides of a particular credit investment. And uh there's there's a there's a vibrant private and uh you know there's a there's a there's a vibrant credit market for that. On the on demand cloud side of things, it's you know not quite as mature as when there's a for example an investment grade offtaker agreement.

40:27 Um, but it's becoming more and more mature and in general creditors and lenders are really starting to understand the value of an Nvidia chip because, you know, you actually look at the chips that we deployed in 2023, H100s, we're now leasing those out at a higher rate now than we were originally in 2023. So, so these creditors are starting to look at these assets and say, "Wow, this is an asset that is very valuable and also easy for us to underwrite." And of course, while they are underwriting towards the actual cash flows that are coming out of that agreement, just as an asset class overall, people are realizing that this is a really great opportunity. And so, creditors are starting to flock to these deals.

41:16 >> You're renting an 800 at a higher rate because why? because the demand for comput is so rapid that uh people will take any or the technical depreciation of the of the product is slower than people thought. What what drives that? >> Well, what's driving it? I mean certainly it's the the de the demand being high increases the price that you're able to get in the market. There's no question about that uh fundamental law. Again, going back to what people didn't understand about this market, there was people who were saying, "Oh, well, there's a, you know, there's a fiveyear lifetime or threeear lifime." I even heard some people say three year lifetime for these GPUs. This completely false. You know, we have GPUs that we've commissioned and we're one of the earliest Neoclouds. In fact, we we're we're probably the only one only Neocloud that actually has GPUs in our fleet that are fully depreciated from an accounting perspective, right? which is most people are adopting around a six-year accounting depreciation schedule. But that's not the usable life. The usable life is longer than the accounting depreciation schedule. And what really matters is the economic usable life. And so what we're starting to see is that like the people who are the naysayers, oh this is going to be you're going to throw these GPUs out in 5 years are completely wrong. They're completely wrong and they've been wrong the entire time. Do you think there is going to be or do you already see um happening some kind of financial market for compute units you know with trading and derivatives is that is that happening >> I'm starting to see some people you know start to examine what a maybe vibrant spot market you know first you need to have a spot market for something before then you can establish you know um a derivative like a future or or other other more exotic things. Um, I'm starting to see that. But fundamentally, I think that the the the asset class is just starting to mature and creditors are starting to become very comfortable with in investing in the credit side of buying Nvidia GPUs and deploying them into data centers and uh we don't need to get too fancy with it. That's kind of like my part of my opinion is that like I think that the that market is starting to mature that that that may be an eventuality is having more complex securities that surround GPUs. But um uh I think for for right now people are starting to realize that it's a it's a great credit investment and that's that's what's changed I'd say over the last year is that people have started to really uh treat it like a a more mature asset class. Maybe quickly just go back to the very origin because I think you've been in the effectively in the AI world the whole time but are coming from a very different angle uh with multiple pivots. What what did you start with and and when?

44:07 >> Well, you know, with the complexity of the business, you can now, you know, the complexity, the capital intensivity, just the sort of not fitting into a box and you can see why we've oftentimes not had a lot of traditional venture investors in in Lambda. and uh you know all all of our investors have done exceptionally well but but they've they've kind of come from more often than not outside of traditional let's say mainline Silicon Valley VCs and um so just going back to the origin story I started Lambda in 2012 and we were a facial recognition software company so I was training convolutional neural networks to do face and image recognition And we eventually hosted that on an API. I was training those convets on a 4X Nvidia uh uh GTX 580 workstation that I had uh that I had bought from a friend who had built it actually. And um and this was, you know, really pretty avanguard stuff at the time. Most people didn't really believe in what was called the field called deep learning at the time. And that was inspired by the imagenet 2012 moment or that was even before that >> the Imagenet moment. You know I I pulled the CUDA connet repo off of Google code.

45:28 That's how you know how old Lambda is is that Google code was still around and I pulled the CUDA connet codebase and was like playing around with it. I got very lucky that the AlexNet paper had been published the same year that Lambda was founded. It's not a coincidence at all. It's not a coincidence at all. We launched this space recognition API, got a couple thousand users, but it wasn't really generating a ton of cash. And um uh sort of as part of that the complex story of startups, you know, in parallel, I was sort of I found these guys who had just graduated from their PhD programs um the gentlemen Zach and Nico and uh they had said, "Hey, we're going to start a company." I said, "Hey, you know, let me help you guys out. I'm going to work with you for a year. I'm going to learn a little bit more about neural networks." And um uh we I helped him out on this company, helped them get a company called Percepio started. And I was the first employee there while I was running Lambda. And uh we were we were running these convets locally on the iPhone. And again, this is 2013. So we were we were using uh the GPU image library and just straight open GLES shaders like the shaders that are used for rendering. We were using those to run the confinets on the iPhone and um eventually I kind of left to go continue to work on Lambda full-time and probably about a year or so later they got uh acquired by Apple. And so if you know the feature on your iPhone where you swipe up on an image and you can, you know, recognize faces and search through your library, that's maybe some of the stuff that eventually got integrated into iOS through that acquisition. And then Lambda, you know, we continued on.

47:16 I had we had a variety of different products. Everything from Lambda hat, which was a baseball cap that took a camera every 10 sec took a picture every 10 seconds with a camera embedded in the tip of the brim for gathering data sets for image and face recognition, >> which is fascinating because fast forward to today and that's a whole segment, right? Like capturing everyday life to train the AI. >> It goes to show you have to, you know, it's one, it's important to be able to see the future. It's also important to get your timing right as well. Right.

47:46 Um, and now it it all worked out, right? Despite maybe that Lambda hat product not being great, but it taught me a lot about how to build hardware. I was I lived in Shenzhen for a little bit, uh, working on the PCB and spinning the PCB and designing the actual hardware product. And, you know, it taught me how to make consumer electronics. And that was actually a huge huge skill because it totally opened my mind to new ways of doing business that aren't just making apps, right? And um eventually we had this product called Dream Scope which became really popular in 201 um 15 and 16 and it was basically using the Google Deep Dream uh methodology of using a confet to generate images. It's like an early version of midjourney or whatever.

48:37 and um DeepDream uh and the Leon Gate style transfer algorithm allowed you to turn a photo into a painting basically. And we got like a million users on that, processed tens of millions of images, maybe 15 million images or something like this. And um that caused us to have a huge AWS bill. It was like $40,000 a month or something. And so to replace that, we'd ended up building a little cluster out of workstations.

49:07 And uh then that was a $60,000 capex that we were terrified to make, by the way. We were so so scared that doing this capex was going to put us out of business. We made it out of workstations because we thought, oh well, worst case scenario, we can just sell them. And so lo and behold, we did end up, you know, turning it online and it brought the bill down to zero. So it paid itself back in a month and a half. And we thought, wow, this is like we're we're saving more money than we're making.

49:32 Maybe we should be in the business of providing compute to other AI researchers. And thus, we started selling workstations and servers and started developing our cloud platform. Maybe did $3 million of revenue in 2017. That first year, selling workstations, then 10 million in 2018, then $30 million in 2019. Um, we grew the hardware business over the next couple years to probably about $200 million run rate. And then the cloud business we really started in 2019 and um you know we started development before then but we started really um marketing it and it kind of was slow to grow to be honest because you know not a lot of people in 2018 and 19 and 2020 wanted a bunch of AI compute. There was a pretty niche market for it. But eventually you know our cloud business continued to grow and now it's at you know a little bit under a billion dollar revenue run rate. we've fully exited the hardware business and uh so yeah, Lambda's got a absolutely wild uh founding story to summarize.

50:34 >> Are some of the people that were there at the beginning still around? I think you started the company with your brother. Is that right? And your brother is is still at the company. >> Yeah. And so in terms of like the early people um basically it's not I mean not even basically of of the four people who are making Dream Scope me, Michael Balaban, my co-founder and fraternal twin brother, Tran Lee who's our chief scientific officer and then Steve Clarkson who's um an engineering leader at the company and you know has a bunch of folks reporting into him. Uh now you know they're all still at the company.

51:09 Um the next hire, one of those uh gentlemen named Mateesh uh Agaral who's one of the the next hires in that team. Um he was with the company for maybe eight years or something like this. um five yeah something like 8 years and uh then he eventually uh left and and joined another former Lambda team member Thomas Summers to start Posatron which is um uh an accelerator company and they're like now valued at over a billion dollars and uh so uh not only has like the original team stuck around but we've already started to kind of see what like a lambda uh alumni a lambda a mafia network looks like in in the world lambda lab member alumni >> how did you keep the the band together during the difficult times >> just when you're running a startup company that's this capital intensive working capital intensive as well like it's just you get a lot of shocks to the system as you're growing um and then co what co I mean in April of co software companies were feeling great because they could ship software and there was so much de more demand. Hardware companies the docks were closed. You couldn't ship revenue in April and March and um so I mean I remember all these things really distinctly. I I think I remember just getting in front of the team like, "Hey, look, it it's it's really tough right now."

52:45 And um you know, there's certainly a feeling that like we're not sure if we're going to make it through this, but the only thing to do is just to like suck it up and enjoy the pain, run through it, and uh come up with the solutions to the problems that you're presented with all in the service of delighting customers. Because fundamentally, I mean, the the big thing is just aligning people towards the only reason we're all here is to build something that people want and they love so much that they tell their friends about it and they give you money. And then everything else, it just fall follows from that customer experience of delighting customers with what you do. When we do onboarding, for example, I used to do this thing in what called Lambda 101 and we would show a picture of like a Linux penguin and he was like on a Lambda workstation and he was reading the GPT uh 2 paper and training had a loss curve which is like what you see and you look at as you're if you're doing machine learning research. I was like, look, just put yourselves in the shoes of the penguin who's training using this workstation or cloud service to train a neural network and just think about what's going to delight them. You know, whether it's, you know, uh people on our shipping team who said, "Hey, let's put some t-shirts inside of the boxes." And so every workstation came with a a Lambda t-shirt, you know, or members of the data center operations team said, "Hey, you know what? we should do a white rack because that that'll kind of like set us apart and make make everything look good and we'll be really proud to showcase that. And you know those are the types of things that as you kind of imbue your company with the kind of delight the customer first mentality that I think helps you get through the hard times.

54:30 recent evolution in that journey is that you just brought on uh a new CEO and my fellow French countryman Michelle K to run the business. Uh walk us through the thinking and what led you to make the decision and how that equips the company for the next chapter. >> It's a huge honor as a founder to get to the point where the company can um afford to bring on amazing talent like Michelle, you know, in that seat, right?

54:59 is because if you think about it, I'd say most companies I it's not uncommon for somebody to say, "Hey, look, a lot of people sometimes maybe there's a comp there's a component of ego involved where they have to be the founder CEO." I've never really personally had that. I I care about the technology, as you can tell. I care about like building a great generational company and uh I think there's so many different seats to do that from. And so, you know, getting to the point of maturity where we could afford to bring on a a CEO like Michelle who has experience like he, you know, obviously previously Soft Bank International CEO, Sprint CEO, >> Alcatel, >> um, Alcatel, he's on the board of some like really amazing companies. Um, you know, including McLaren, which is a kind of a fun one. I always did the sort of like fundraising and capital formation and day-to-day business management as a necessity and not like because that's what I really love doing for example right and I think there's plenty of founder CEOs who absolutely love every aspect of their CEO job I think that privately and it will be very hard for you to get this out of any founder CEO oftent times but like secretly When I talk with founders CEOs, I'm always like, "Yeah, so like how much do you hate?"

56:27 >> I find it shocking that people don't find speaking to VCs all day exciting, but I will take your word for it. >> It's been like an amazing experience for me that I to be able to form a team around the company and to just see everybody flourishing in the things that they love to focus on. So for example, now that I'm the CTO, one of the main things I'm focused on is what does rapid data center deployment look like at the company and you know kind of working to say like hey I want Lambda to be this sort of vertically integrated high velocity powerhouse that so when you look at the world you say all right there's two people in the world that can and two companies in the world that can do high velocity deployments SpaceX AI and Lambda where we're just extremely focused on how do you cut every little piece out of the process to stand up compute faster. Um, and that's like something I've just been diving into and really enjoying with my new uh time as uh uh CTO.

57:33 >> What was XAI's record like when they launched? >> I think it was like 200 and something days. >> Yes. And you think that can be matched or exceeded at a repeatable pace? >> I think it it can be matched or beat. Yeah. >> And that's process mostly. I think it's it's everything from like the site selection process, the set of constraints that you use in a site selection process, the the the MEP pipeline, the way that you construct the data center, um you know, h how do you make it so that the end customer will consume that compute, you know, and and how do you cut out a lot of stuff out of the process? Because often times, you know, the people who've been designing these data centers have really kind of been real estate people, as I've mentioned, who've been kind of grabbed by the scruff of their neck by a hyperscaler. They're like, "Go and build this design. Here, go go go get a GC run off." And they don't know anything about what goes inside of it. And so, and and the hyperscalers on the other hand have been really building towards traditional cloud services. I mean if you look at a modern region in any of the clouds they have hundreds of services. I mean everything from satellite base stations to tape storage to spinning disc to face recognition APIs. I mean these are all the services and each of those services requires a different skew and has different parameters about what you're kind of servicing and and in fact you know you might have somebody who's trying to run an ATM backend on one of these things that's a pretty different design space and design constraint than an AI data center that could maybe you know have a lower availability and uptime right and so that's kind of where I think Lambda is able to build a lot of really unique um value and you know through this kind of targeted AI first approach >> you had a a quote where you said that AI won't write software it will become the software what what do you mean by that >> so uh that's in my sort of like idea around what I call neural software and or you know a neural computer neural operating systems and the best way to kind of get this experience is to go to your Chad GPT or your Claude and say, "Hey, just um you know, render for me an ASI art desktop interface." Okay, so you're you're working in purely in the domain of text and I want you to just pretend to be an operating system for me. I'm going to say click on this, you know, open up this and I want you to just behave like a computer. So give it that prompt. Okay? And um what you're going to see I think is that you're going to see that uh that sort of future of the large language model becoming the software and not generating the software. And um this results in an extremely sort of squishy and flexible way of interfacing with a computer where it's not possible to have a bug, only a misunderstanding about the prompt and what you've asked for. And I think that for a lot of the pieces of software on your computer, you might see that taking over where you know you can get the glimpse of the future with this ASI art and then eventually it'll also have a multimodal network that's generating every pixel on your screen as well as every audio waveform that comes out of your speakers. The advantage to this is that you can really sort of dream up software. the only the part that is being experienced by you and uh you know is actually implemented right if that makes sense um you know it could have whatever feature it is if you ask it and that's like a really powerful way to interact with a computer I think >> so it's not like you you give simple instructions to the LLM so the LLM is the software >> I guess make the analogy like you know vibe coding takes in a prompt and then outputs It's human readable, writable, compilable code that runs on normal human software programming language substrate, right? You know, it outputs C code which gets put through a compiler.

61:55 It outputs Python code which gets put through a Python interpreter. That software is static. Once it's been generated, it can't change, right? You can vibe code it again and maybe, you know, vibe code on the fly. There's like a couple different stages of the gradient between traditional human written software and then you go to like maybe vibecoded software. But then you go to just in time vibecoded software where like it's a live creation of the software application >> but still software >> but still software but then you go to the next step which is just you're interacting with the LLM and it is emulating kind of how software might behave and that's that's the difference um between vibe coding and a neural operating system or neural software.

62:41 Neural software there is no code that's running. It's just modifications of the feature activation space and the context in the mind of the neural network. >> And how far do you think we are from that? Is that something >> I mean we we have prototypes of it today. So we have prototypes of it today. >> And when you say you is that is that lambda or is that is that >> yeah lamb lambda's developed a prototype. There are multiple other companies have developed prototypes of this. Um there's academic research that is has you know outlined what this might look like and um you know how far are we from you know mass adoption. I would say that generally speaking when I'm early on something I tend to be about a decade to a decade and a half early. So I would say that between a decade and 15 years we will see mass adoption beginning or otherwise happening for neural software.

63:42 I mean you you already have it. So here here's another example by the way you already have so you can think of a Tesla self-driving car or you know any type of endto-end neural network and then you know and then large model that's doing autonomy as a form of neural software right you know people understand that aspect right which is it's seen video it's making decisions about what to output now the user experience is the driving experience that said that is an example of neural software I would I would argue and so we already see that today now the question is when is your everyone's computers going to adopt that I'd say a decade >> do agents change anything from your perspective as a compute provider and if so in what way >> to understand what needs to change on the compute layer we understand need to understand what's changing with the user so when you're doing vibe coding with agents one of the things you'll notice is that your wall clock time, you know, in in the world is mostly spent on running tests, gathering data, searching through a codebase. A lot of the time is spent act not just inferencing a neural network, but it's actually spent doing other things. And it's actually very much not it's very much similar to how software engineers spend some of their time, right? You know the old XKCD cartoon of compiling where they're sword fighting on the office chairs and someone says, "What are you guys doing?"

65:17 Well, compiling. And so, so now there's a bunch of time spent compiling. There's a bunch of time spent running tests because the part of the way that that the agent 24/7 loops really work well is when you are constantly banging against a nice suite of automated tests to make sure that the code you're writing is good. And so, well, what does that mean? It means that every single um cloud service needs to start doing a lot more um traditional CPU workloads. They need to do uh focus on a great uh environment, a secure environment to host your clawed code instance on. Um and then you need to think about security from the perspective of you need to think about how this massive influx of new applications are going to be secured.

66:14 >> How do you use AI agents internally? >> Well, I mean a lot of the engineers at Lambda are already you know doing a fully agentdriven workflow. I mean if you just go to cloud code and say hey use advanced workflows or you know spin up agents um you can do that. So that's like step one. Um I've demoed internally and some folks have adopted what I kind of call self assembling software. And so self assembling software is this idea where you, you know, kind of tie in to a 247 running agent fleet um product requirements and constant user feedback that's coming off of the system. So you have a very clear and tight loop to go from submitting, hey, this is a bug or this is a feature request and there's a fleet of agents who are implementing that live for you. Okay. And that sort of cycle I call self assembling software is because you kind of say hey this is what this software is for but most of the development for it is going to happen after the software is launched and the users start to interact with it and customize it for themselves collectively. And I think that that is kind of maybe the future paradigm of where a lot of the agent driven development is going to go towards. um the the the the other side of that eventually once the models get smarter, I think that they're quite not not quite there yet. But, you know, tying that back into, hey, I need help. And I'm not talking about the human. I'm saying the agent's going, "Hey, I need a human to help me. Like, I need you to plug in a thousand GPUs for me, or I need you to um give me an API key to a particular service. I need you to go sign up for something for me. can you go please negotiate this? And I think that that's actually how you're going to start to see it happen, which is product user feedback gets implemented by the agents. The agents then also ask the people at the company to go and do things for them all in the service of you know delighting customers and making money.

68:18 >> You've talked about gigawatt scale factories. Is that what you were describing earlier around like setting up beginning super good at creating data centers very quickly but it's also making them bigger. What is that concept? It's a an AI factory which is a basically land data center servers inside that is generating tokens and a gigawatt scale means that it's consuming a thousand megawatts or a billion watts which um is a lot of power sort of like maybe you can think of it for context New York City is something like five gawatt. You also talked about one person, one GPU. Is that your your vision for the future? Unpack that for us.

69:05 >> So, um, you know, before people really believed in the AI thesis, I, you know, when I was pitching our series B and C, I would kind of talk a a lot about the similarities between, let's say, the computer industry and the AI industry. I really felt like AI was forming a set of generational companies. Um, and there was going to be a set of generational companies that got minted with the changes that were coming with AI. And this is like in 2020, 2021.

69:35 And if you read about the history of Apple, for example, in the early days, the motto and the the sort of the credo at Apple was one person, one computer. One person, one computer. And you know, there's a sense of humility that's embedded in this one person, one GPU, which is the one person, one computer. You think about how visionary Steve Jobs was. That was, you know, Apple was, you know, what, founded in 1976 or something like this. The Macintosh came out in 1985 or 1984, excuse me. 1984, you know, whatever. Um, 8 years or so after founding. Is that one person, one computer yet? No, not even close. All right. So, 1984 to 1994. All right.

70:19 Well, is it one person one computer? Well, we're just starting to have the internet boom. So, we're we're we're I we mean not quite there yet. 2004, we finally have broadband internet access. And maybe for the first time in the United States, there's not quite one person, one computer, but there's certainly like one person, one family, one computer, you know, or you know, something like this. It's like getting close to it. You don't have until 204.

70:46 So 74 84 94 2004 2014 40 years after one person one computer do you have probably truly one person one computer and you get actually beyond one person one computer because people have laptops and cell phones and I would consider a cell phone a computer so and then and then finally you didn't even have e-commerce penetration until 2024 50 years after the founding or so of Apple computer when e- e-commerce starts to actually penetrate because of co I think that the the reason I really wanted to choose that one person one GPU is because one I believe that in the future everybody in the United States will need the computational power of one GPU or more to just do their daily work you know enjoy life whether it's getting access whether it's do getting entertained whether it's being productive whether it's being creative And I also recognize that it took Steve Jobs and Apple, one of the best companies in the history of capitalism, half a century to accomplish their goal.

71:56 And so I I think this is not it's not just like an overnight let's, you know, quickly get to one person, one GPU. Um, so that's that's that's what that means to me. >> To close, are you ready for a couple of uh quick hot takes? >> Sure. >> What is one idea in AI that is overhyped? I think a lot of the sort of agent agentic workflows for things that are not software engineering I think tend to be overhyped and I'll tell you that the reason for that is because one of the ways that you get an agentic workflow working really well is that it needs to have very concrete feedback mechanisms which are done brilliantly through automated testing. It's not at all done brilliantly for like go and buy a site. there's no there's no traction to give a model to go and iterate over a long period of time on. So I think agentic workflows for things that aren't readily verifiable. Now I wouldn't say as far as everything that's not software engineering because there's plenty of readily verifiable fields CAD uh computer AED manufacturing um finite element analysis computation fluid dynamics there's a bunch of fields where you can really do in a great agentic workflow I mean and and simulate it and then go and iterate it's not the case for hey Claude make me a billion dollars make no mistakes you know >> sadly or maybe not okay fascinate inflation would be well it wouldn't be inflation it would actually be just value creation in the economy deflationary even >> what is one idea in AI that is underrated >> yeah I I really think that the neural OS thing and you know also some of the aspects of self assembling software like I still do think people you know the funny thing is I'll give the same answer agentbased workflows for software development I think that most people don't understand they literally don't understand because they've never tried it. They've never gone to cloud.

73:48 Go to cloud to go say maximum effort. Uh use the latest model and then go and build whatever you wanted to build and say, you know, spin up 10 agents to go and do it. I think a lot of people still haven't done it yet. >> Well, Stephen, it's been wonderful. Thank you so much for spending time with us. >> Matt, thank you so much for having me. >> Appreciate it. >> Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks, and see you at the next episode.

Summary

The podcast features Steven Balaban, co-founder and CTO of Lambda, discussing the evolving landscape of AI and cloud computing. He emphasizes the growing demand for GPU compute, the complexities of building AI-focused data centers, and the misconceptions surrounding the commoditization of cloud services. Balaban also shares insights into Lambda's journey from a facial recognition startup to a leading player in the AI cloud space.

- The demand for large language models is insatiable, leading to a continuous need for GPU compute.
- Cloud compute is not a commodity; it requires complex integration of hardware, software, and infrastructure.
- Lambda's unique software allows for efficient management of large GPU clusters, differentiating it from competitors.
- The market for cloud services is expected to remain oligopolistic, with room for multiple large players.
- Financing for data centers involves innovative methods to manage capital-intensive projects, with a focus on long-term contracts.
- The biggest bottleneck in the industry is securing entitled land for data centers.
- Balaban predicts that AI will evolve to become the software itself, moving beyond traditional coding methods.
- Lambda is focused on rapid data center deployment, aiming for gigawatt-scale AI factories to meet future demand.
© transcribe · For agents Built with care and craft by Gokul Rajaram