Section Insights
Understanding GPU Depreciation and Data Center Costs
What are the implications of GPU depreciation on data center investments?
The depreciation of GPUs significantly impacts the overall costs of maintaining data centers. Unlike traditional real estate investments, data centers require continuous financial input due to the limited lifespan of GPUs, which typically last around three years. This necessitates ongoing investment rather than a one-time expenditure.
- 50% of data center costs are attributed to GPUs, which have a limited lifespan.
- Data centers require continuous investment to replace aging hardware.
- Traditional NPV calculations for real estate do not apply to data centers.
Demand for AI Tokens and Hardware Longevity
How does the demand for AI tokens affect the usage of older GPUs?
Despite the decreasing price of AI tokens, demand for their outputs has surged, leading to increased usage of older GPUs. Companies are willing to pay higher prices for renting out older models due to the growing capabilities of AI models.
- Demand for AI outputs has increased significantly, sometimes by 10-30 times.
- Older GPUs are being utilized more effectively as demand rises.
- Companies are willing to pay more for older hardware due to increased demand.
GPU Lifespan and Usage Context
What factors influence the lifespan of GPUs in data centers?
The lifespan of GPUs is heavily influenced by their usage context. GPUs used for intensive training have higher failure rates compared to those used for inference tasks. Understanding the specific usage history of GPUs is crucial for assessing their reliability.
- GPUs used for training have higher failure rates than those used for inference.
- The context of GPU usage significantly affects their longevity.
- Data centers have a mixed population of GPUs with varying failure rates.
Market Dynamics and Price Competition in AI Models
How does the convergence of AI models affect pricing and competition?
As AI models converge, price becomes a primary differentiator, leading to a decline in premium pricing. This situation reflects Jevons paradox, where cheaper tokens lead to increased usage, but the expectation of massive growth in token demand may not be realistic.
- Price competition is intensifying as AI models converge.
- Cheaper tokens may lead to increased usage, but growth expectations may be unrealistic.
- Historical tech bubbles show that price declines do not always lead to increased demand.
Challenges of Declining Token Prices
What are the implications of declining token prices for AI companies?
Declining token prices pose significant challenges for AI companies, as they struggle to maintain profitability while facing increased competition. Companies are shifting their focus to higher-margin markets to offset the impact of falling prices, which may lead to ethical concerns and market instability.
- Declining token prices create pressure on AI companies to find new revenue streams.
- Companies are moving upmarket to seek higher returns amid falling prices.
- The shift in focus may lead to ethical concerns and market instability.
Transcript
0:00 So, the depreciation that you're talking about is because 50% I think you've said 50% of the cost of data centers is in the GPU. And the GPU has a lifespan, you know, someone some people say 3 years, right? This is the typical depreciation argument. You put, let's say, a Nvidia H100 in there. 3 years later, it either fails, like you said, or you have to replace it with a black whale or a Rubin, whatever it might be. so so and so therefore these these expenses in the data centers aren't just like you invest $100 billion in a data center and you get to live off the land for 20 years. The investment is, you know, you have to continually, you know, feed that data center with more money in order to make it work.
0:46 >> Right. >> So so >> Which doesn't work in the context of the NPV calculations that underlie a typical real estate project, obviously. That's completely different from the kind of math that we use to justify a multi-tenant apartment building, for example. >> And and then, you know, further on, what you what you mentioned is that the tokens, right, with the things that these GPUs produce, they are depreciating. They are getting cheaper. No one will argue with that.
1:11 The the counter argument >> of just I'll just say they're not really depreciating. They're actually staying the same value. They're just deflating. There is a difference. >> Okay, right, deflating, right? The what you used to pay for a token is is much cheaper now than it was previously. Okay, so here's here's what the counter argument would be wrapped up into one. The counter argument would be, you know, as tokens have gotten cheaper, people have wanted more of them because the the AI models that used to use them have become A more powerful and B more capable. So, therefore, you know, even if those tokens are cheaper, people just the demand for the outputs of these factories have grown by a magnitude, sometimes, you know, 10, 20, 30x than they were previously. And as you do that, you know, as that demand has grown, you know, people are willing to use even the older chips at rates that would be higher than when they initially came out with the less powerful models. So, I was speaking with CoreWeave at the end of the year last year, beginning of the year this year, you know, around New Year's time, and they said they were actually renting out H100s for higher prices than they had previously. So, all of what you said is true. The counterargument that they would make is, yes, and their demand for the tokens is higher, and the old hardware is working well beyond that chip that typical 3-to-5-year estimate that people expected. So, what is your thought when people say that?
2:41 >> So, there's a whole bunch of nested arguments in there. so, let's take them on kind of one at a time. the lifespan of a GPU, in terms of just looking at it from an MTBF standpoint, I mean time between failure standpoint, depends very much on what it was used for in the in its in its in its adolescent years inside the data center. The analogy I often make is if you could buy a used car, both two if two used cars, one of them both has they both have like 5,000 mi on them. One was driven in a, you know, 72-hour nonstop race across the country. The other one was driven That was the only That's where all the 5,000 mi came from. And the other one was driven to church on Sunday for a year.
3:17 Which car would you buy? Well, I think we would all buy the car that was driven to church on Sundays. I want nothing to do with the one that was raced in some kind of, you know, bubblegum rally across the country. So, the the in the context of GPUs, what we have is a generation of GPUs that were largely used for very intensive training purposes. And so, the the failure rates of of GPUs used so intensively for training purposes are much higher than inference-specific usage. So, yes, there's no question that if a chip is used exclusively for inference, which is to say token completion in response to prompts, then the lifespan will, all else being equal, likely be longer. And if I have if I have a chip that didn't fail during training, then I can I repurpose it potentially to be used for inference? Sure, there's no reason. But in aggregate, there is this problem that the failure rates of chips that were used for training is very different from the failure rates that were used for inference. So, we have this kind of mixed population of chips inside of data data centers with very different failure rates. And people have a tendency to conflate this and just pretend that it's all the same thing, and it's not. And that's not very helpful cuz if you actually talk to people who are running data centers, they will say, "This is exactly what we're seeing." Is we see much higher failure rate. So, there's a there is this sort of blended problem that you have to understand the nature of what the chips were actually used for. And that's only going to become more profound in future because increasingly I've often I often joke that the the the frontier model company that will the most valuable frontier model company in future will be the one that stops pretending to train models and actually just moves on to harnesses and moves up the stack because what we're seeing increasingly, if you look at things like the Epic composite index and other things, is that while models are still improving, they're improving at a much slower rate. And I often do this kind of Pepsi Coke test where I'll put a couple of different models in front of people using some kind of a harness like open code and ask them to tell the difference. And everyone thinks they can tell the difference, and the reality is no one can tell the difference. And so, we're at this point of convergence that we're increasingly the thing that differentiates differentiates models outside of marketing is price, which is one of the reasons why on tables like like a router or whatever else, it's now dominated by Chinese models. So, we're rapidly seeing this move away from any kind of premium pricing in terms of the models themselves, which is and I'm trying to get to the point about this kind of Jevons paradox, which is really what you're pointing to, this idea that as models get cheaper, we use we our tokens get cheaper, we use more of them. And this is a this is a common idea and and you know, we've seen it repeatedly play out in different ways over the last 150 years.
5:46 But I think this mostly speaks to the innumeracy of people that they don't understand what a compounding price decline of 80% means in terms of what you would have to see in terms of growth on the other side. You have to see around a 100 million-fold growth over the next 6 years in terms of tokens. Is it possible? Absolutely it's possible. Is it likely? No, it's not likely, but it could happen. But let's not pretend that it's probably that it's one of the most probable outcomes. To throw out to throw out this insane but Jevons paradox, but people will use more is to really dodge the core problem of the geometric decline in the price, which will only continue and get faster now that we've got increasingly price-based competition because of the convergence of models.
6:29 So, that problem not not only doesn't go away, it gets even harder in future. And, you know, now you're competing with sovereigns who have state-subsidized token prices. as a China is the good probably the canonical example. And so, all it just becomes increasingly difficult to make the kinds of returns that your investors expect given the comparable cap rates that they're they're comparing them to. So, this idea that but it will work itself out because prices will continue to decline and magically we'll just use enough is is both historically naive. This argument gets made all the time, has been made repeatedly in prior tech bubbles, and people wave their arms and say this, and it almost never works that way. And it's worse this time because at the core is this deflating commodity called tokens that is being used to pay a fixed cap rate in terms of what the expectation is from the investors who fronted capital for these instruments. So, is it possible? Sure, but think about some of the carnage that it's already creating.
7:25 You know, Kar Alex Karp was complaining on CNBC the other day, I'm sure you saw it, that that these companies are increasingly marching upmarket and trying to eat other The reason why they're marching upmarket is because they see this coming, and they're looking for higher return places to be because they see the collapse in the fundamental commodity that they're selling. No different than, you know, gold miner deciding they need to start making jewelry. This is the same phenomenon playing out. So, they're marching out and so that's going to have collateral damage in terms of them being seen as fair and unbiased players, which will then play into the likelihood of companies going down the path of, you know, sovereign data centers and doing token inference generation inside of inside their own organizations as that becomes increasingly possible. So, I think there's no doubt that we'll see this continuing growth, but whether or not the growth will be large enough to compensate for what will be essentially be a an asymptotic collapse to zero in terms of the price of tokens is mathematically a very hard argument to make.
Summary
- GPUs account for about 50% of data center costs and have a typical lifespan of 3-5 years.
- The failure rates of GPUs used for intensive training are significantly higher than those used for inference tasks.
- Demand for AI tokens has increased dramatically, sometimes by 10-30 times, despite their declining prices.
- The speaker emphasizes the difference between depreciation and deflation in the context of token prices.
- Historical patterns suggest that simply expecting increased usage to offset declining prices is often naive.
- The competitive landscape is changing, with state-subsidized token prices from countries like China affecting market dynamics.
- Companies are shifting their focus to higher return markets due to the collapsing prices of AI tokens.
- The future growth of token demand may not be sufficient to counterbalance the rapid decline in token prices, posing risks for investors.
Questions Answered
What are the implications of GPU depreciation on data center investments?
The depreciation of GPUs significantly impacts the overall costs of maintaining data centers. Unlike traditional real estate investments, data centers require continuous financial input due to the limited lifespan of GPUs, which typically last around three years. This necessitates ongoing investment rather than a one-time expenditure.
How does the demand for AI tokens affect the usage of older GPUs?
Despite the decreasing price of AI tokens, demand for their outputs has surged, leading to increased usage of older GPUs. Companies are willing to pay higher prices for renting out older models due to the growing capabilities of AI models.
What factors influence the lifespan of GPUs in data centers?
The lifespan of GPUs is heavily influenced by their usage context. GPUs used for intensive training have higher failure rates compared to those used for inference tasks. Understanding the specific usage history of GPUs is crucial for assessing their reliability.
How does the convergence of AI models affect pricing and competition?
As AI models converge, price becomes a primary differentiator, leading to a decline in premium pricing. This situation reflects Jevons paradox, where cheaper tokens lead to increased usage, but the expectation of massive growth in token demand may not be realistic.
What are the implications of declining token prices for AI companies?
Declining token prices pose significant challenges for AI companies, as they struggle to maintain profitability while facing increased competition. Companies are shifting their focus to higher-margin markets to offset the impact of falling prices, which may lead to ethical concerns and market instability.