Section Insights
The Trade-off in Hardware Specialization
What is the impact of hardware specialization on performance and flexibility?
Specializing hardware for specific workloads increases performance and power efficiency but reduces flexibility. The durability of the workload is crucial for making such a specialization worthwhile.
- Specialized hardware can lead to faster and more power-efficient performance.
- Flexibility decreases as hardware becomes more specialized.
- Understanding workload persistence is key to effective hardware design.
Doubling Capacity: A Combined Effort
How does Google plan to double its serving capacity every six months?
Google aims to double its token generation capability every six months through a combination of hardware improvements and software optimizations, rather than just increasing hardware performance.
- Capacity improvement relies on both hardware and software advancements.
- Model and runtime optimizations play a significant role in enhancing performance.
- Continuous incremental improvements are essential for maintaining growth.
Collaboration Between Teams at Google
How do Google's AI Infra and DeepMind teams collaborate on model development?
The collaboration involves close partnership where hardware and model designs are co-developed, allowing for optimizations that enhance efficiency and performance.
- Collaboration between hardware and model teams is crucial for innovation.
- Joint efforts lead to significant improvements in training and serving efficiency.
- Real-time feedback between teams can drive rapid advancements.
Advancements in Data Transmission Technology
What is optical circuit switching and how does it improve data transmission?
Optical circuit switching transmits data in the optical domain without converting it to electrical signals, allowing for faster and more efficient data movement between racks.
- Optical circuit switching enhances data transmission speed and efficiency.
- This technology reduces the need for electrical processing of data packets.
- Implementing optical solutions can significantly improve data center performance.
The Lifespan and Utilization of Data Center Hardware
What is the useful life of chips in data centers and how does Google manage older hardware?
Google's older TPUs are still utilized at 100% capacity even after several years, indicating that the useful life of chips can extend beyond typical depreciation periods.
- Older hardware can still be highly effective and utilized in data centers.
- The depreciation lifetime of chips is approximately six years.
- Maximizing the use of existing hardware can be a cost-effective strategy.
Transcript
0:00 In hardware again as you know there is this opportunity where the more you specialize to a particular workload the less flexible it is the faster the more power efficient the hardware is going to be. So it is this art and it's this projection of what are you designing to and how persistent is that workload. In other words if it's going to go away after a month or two months or 3 months even if it's big for those three months you got a really narrow window to intercept it. So it has to be somewhat durable. and you have to be able to project ex exactly what win you can get for specializing to it.
0:52 >> Thrilled to welcome Amin Vad to the show. Amin thank you for joining us for today. It's really exciting to be here. I'm really looking forward to it. >> I am very excited for today's topic because we are in the middle of the biggest capex buildout in human history. You are in the middle of it. Google alone is expected to spend more than $200 billion on capex this year. Most of which is going into building data centers. you're at the center of it all. You were named the head of Google's AI Infra at the end of last year.
1:22 so you are the man spearheading the efforts of one of the most capital intensive buildouts in in human history and so I am really excited to get into it with you today. >> It's a huge year for sure of course across the industry including at Google really have I don't think we've ever seen anything like this honestly certainly not at Google but as you said I think in the history of humanity in terms of the buildout and the pace of transformation really unbelievable.
1:46 Before we get into it, maybe just just crash course for the audience. What is an AI data center and how how is it different from a non-AI data center? >> It's a very good question in that AI data center and a non-AI data center have do have a lot of similarities actually and they're they're not radically different. I mean it consists of concrete right an enclosure it consists of electrical yards mechanical yards cooling you have row after row of power being distributed.
2:16 there's a huge amount of network infrastructure there. In other words, we connect large amounts of computes to one another. a fair amount of storage infrastructure goes in there. I think that the big difference that we're seeing with AI infrastructure is specialization. In the past, when we're building a data center, it really is a 20 25 30-year building investment. And we're thinking about how it's going to be evolving over that 20 25 30-year period. You might have servers go into it. You might have networking storage. You might have some accelerators, GPUs, TPUs, whatever they might be going into it, but it's a 25 30 year planning horizon. So we have and the lifetime of the hardware might be six years. So many generations that we have to plan for. An AI data center often times is going to be more purpose-built. So in other words, we are oftenimes co-designing the building with the hardware that might go into it. We might be saying, you know what, actually we're not going to put a lot of storage into this building. Why? Because a storage rack might have 10, 20, 30, 40 kilowatts of power. You put that next to a TP rack or a GP rack that easily is hitting hundreds of kilowatts today. Might hit me. People are talking about that in the next few years. Designing a building that might take 30 storage racks in a row versus one or two AI racks in that same row. Very very different design. Just think of it in terms of the size. Think of it in terms of the power and how that would be distributed across the building etc.
3:56 Networking would be the same. The amount of networking that you would need for a storage rack tiny. I mean, if especially if it's hard drives compared to an AI rack. So, if you're trying to make it fully fungeible, you'll probably make it too big and too overbuilt. Fungeible over certainly a 30-year period. An AI data center likely is going to be much more purpose-built, co-designed with the hardware, even to the point of cooling and power distribution, etc. So, a lot more co-optimization. You all delivered a big Ver Rubin cluster to one of my portfolio companies, Ineffable Intelligence, and I saw the photographs of it as it went out. That is a thing of beauty.
4:38 >> It is. It is. >> It's it's I mean, that's one of those things where, you know, similar to looking at the biggest construction projects in human history, you look at that data center, it's like, wow, that is a monu monument to what what mankind can do. No, we got a nice picture of it and it and this is just you know a couple of racks. You know the the cabling the fiber distribution for it is is really be I mean you got beauty is not the beholder but for folks like me and probably yourself and others in the audience as well. it is a thing of beauty. We put this picture up on social media and people love the fact that ineffable by their huge fan of what what they're doing. Fantastic team that got a lot of attention. Okay, great rack for an effable. But actually, people just loved seeing the the pictures of the fiber and the sort of the fractal nature of the fiber. One of the most popular posts that we've ever had actually was around that. So, so yeah, absolutely. We're really really excited about that.
5:29 >> Yeah, I had goosebumps seeing the picture. >> How do you measure and hold yourselves accountable in this in the middle of this big buildout? we were talking right before the show about how flops is a vanity metric and you prefer an alternate metric. Can you can you say more? Yeah. So I think one thing to note is that you know whether it's flops or pick your other favorite chip centric metric these are in theory. In other words under some conditions for whatever chip you have this is the maximum amount of flops that you can deliver. What we really care about in the end is what's the performance delivered by workload.
6:05 That's that's what we're looking at. And rarely it it happens but actually relatively rarely is that performance determined by a single chip. It's flops, it's HBM, it's SRAMM capacity etc. Those are all super important but it might be how 2 4 8 16 a thousand 10,000 more of these chips composed together with not just the accelerators again whether TPUs or GPUs but then the CPUs that might feed them the data the network that connects them all together.
6:36 So what is the workload that you're running and what is the performance of that workload? One measure that is interesting is what's your flop's utilization. So for example for a particular workload if you have in theory teraflop or a pedlop capable for your workload what fraction of that are you delivering? >> Now that's a measure of goodput. >> What is goodput? >> So you can think of throughput which is a well-known term. Yep, >> that's the throughput that is possible.
7:06 But now let's consider some other considerations. One is what is the slowdown of the workload? Again, just inherent to the workload, but other aspects of it that really hits our reliability. So in terms of how we hold ourselves accountable, if we have a chip failing for a synchronous workload and many of these workloads, whether training or serving or agentic workloads, they're synchronous. There's many many components that are working together simultaneously. Now if you have let's say a thousand 10,000 100 thousand of these components working together simultaneously and they're really needing to coordinate at microscond or millisecond granularity. One of them fails it might actually bring the whole thing to a stop right because everyone is counting on everyone else to do their part of the job in order to come up with the answer to a really tough question.
7:55 One of them stops. Okay now we have to figure out what happened which one stopped. What's a checkpoint that we have of the computation at some previous state in time? How do we restore that checkpoint? How do we restart? Worst case, we have to restart from the beginning. That would be really bad. But that that can happen in certain cases, especially on more on the inference side. The point here is that if you now then have to go back and redo a bunch of computation, if you then have to pause and wait to figure out what happened for the failure, all that work is work. It's not actually helping you get the answer.
8:28 Right? the in terms of let's say that you're working out a problem on paper. Step one, step two, step three, step four. If you have to go back to step one because you you made a mistake, yes, you're still doing work. That's throughput. You're doing work, but what is the goodput in delivering your answer? It's basically the total amount of time it took you to solve the problem. That's that's what matters. And if you have failures, if you have failure recovery, if you have whatever it is that's interrupting your work, that's that's part of the issue. So now, how do we hold ourselves accountable?
9:01 It's delivered goodput, not theoretical benchmark throughput or goodness in theory, but actually for a workload for real failure conditions, what's happening? And the unfortunate reality is that at a 100,000 accelerator scale, >> I was going to say at that scale, something's filling all the time, right? >> All the time. and and these each one of these as you're aware each one of these chips is a u wonder of nature. I mean that they're at the very bleeding edge of what's possible to manufacture, right? And so again to to me it's stunning. And now actually these chips aren't just one chip, they're packages made up of in many cases two, four, eight, maybe more chiplets, right, that are composing together. And then of course HPM off to the side, network connectivity, maybe things like co-ackaged optics. No criticism. Lots of things can fail. And now if you got 100,000 of them, something is going to fail.
9:56 You have to be prepared for that. Detect that in near real time. Recover from that in near real time. Doing all that it's it's really like finding the the telemetry problem is massive. It's like finding a needle in the haststack continuously across certainly seconds and minutes. But for some of these jobs, hours and days and even weeks like just continuous and online. How we hold ourselves accountable is what's the goodut that we deliver for the workloads that actually matter in the data center.
10:24 Is goodput a Google term or is it an industry term? >> It's a Google term but I think more and more people across the industry are starting to pick it up. >> And then just to calibrate me 100,000 accelerator scale are we talking it's going to fail once a minute once a once a day like how how >> 100,000 accelerators. So let me see I think that it is definitely going to be at the scale multiple times a day and perhaps depending on the exact configuration multiple times an hour something is going to fail. And what are the most common failure reasons?
10:55 >> This is the problem. Actually, it's a very good question. if there were a common failure reason, we'd be we'd have figured it out and and fixed it. It is a long tale of constant discovery. You know, whenever we have a new product that's being introduced, there's going to be something that hits us. in many cases, frankly, because again, it's at the very bleeding edge, it might be network related. It might be how we connect these things together at super high speed. that might be something to do with the hardware. Those things we work through. but then a lot of the issues could be software. So this is the other aspect of how we hold ourselves accountable. Again, the chip might be capable of a certain level of flops. But if you have a compiler bug, a runtime bug, a model issue, something else, operating system issue, it doesn't matter. That that's going to impact the antenn system performance. So you you might have perfect hardware fully reliable but then you might have software issues that hurt you.
11:52 >> Is there a standard reference stack from the accelerator companies from Nvidia or from the TPU team of you know here is the optimal system to build around our accelerators and as long as you build to that system you're you're good or how much of your own kind of data center design do you have to do above and beyond what the what the semiconductor companies give you? right. So there's a there's a reference stack and I mean I would say that Nvidia is a incredible whole systems company that I mean they're obviously semiconductor company but they're not just a semiconductor company. They give you a really really strong reference stack but I would say that many we find that the way I'd put it is most of our customers leverage that reference stack but many also specialize. So in other words, they might find that as as natural for their particular use case, they have an optimization opportunity or something different that they need to do and they're going to do that. Similarly for the TPU side, we have a reference stack.
12:50 But then and most people will leverage that. Many will specialize it as well though. >> Got it. And then I guess in terms of overall capacity, you've told your teams that Google has to roughly double the serving capacity every six months or so. Is that right? Well, so to be clear, this is in terms of effective available capacity from the way in the end that we look at it is from a serving perspective. It's token generation capability. So capacity and this is where I would say it's a combination of software and hardware. The hardware might have some quote unquote inherent level of flops.
13:26 What I'm saying here is not necessarily you have to double the number of flops every six months. That's one one path. You have to double the capability of that hardware to generate tokens every six months. And as much or more of that is going to come from software as it is from hardware. So in other words, it could be a model optimization that delivers that. It could be some runtime optimization that delivers that. It's really probably going to be dozens or hundreds of individual optimizations that are just landing again and again and again to make it all possible. But yes, the rate the rate of capacity improvement is incredible.
14:03 >> I I guess we've had a few years of data center buildout now and a few years of model progress and software progress. What has been the empirical breakdown of how much capacity me as measured by intelligence per watt? How much capacity increase has come from the silicon itself versus the models versus other software and and any other kind of big components? Yeah, it's a really good question. I don't have the exact breakdown, but I would say that in our experience, most of the benefits come from model side improvements in terms of intelligence per watts. And by the way, the intelligence per watt is a fantastic metric. really focusing we we also for us it's goodput per watts and we can come back to why the watts need to be the denominator. Goodput could well be a measure of the intelligence delivered per watt. In other words, goodput is a workload specific metric.
14:53 But I would say most of the gains often come come from the model side software system quite a bit. Why? Because they're able to actually make sure that the hardware that you have is being used effectively. And then the hardware, it it is pretty stunning. In other words, we're living in a world right now where 2x or more year-over-year performance improvements is absolutely possible. And so the hardware does does support it. And it it is a if you want it's a free multiplier.
15:24 Not free but everyone else above it can count on it in terms of that being a multiplier that lifts everyone else up year-over-year. >> I'd love to talk a bit about the TPU program and and codeesign. so Google started building custom silicon more than a decade ago. It was contrarian call at the time. how has the TPU program evolved? >> Quite a lot. I mean when the program started 2013 it was really a contrarian call and I mean I think that at the time it's hard to put yourself back into the moment 2013 conventional wisdom you know all the smartest wisest people would say you don't build a customuilt accelerator for a single workload why because Moore's law is there doubling of performance every whatever it is 18 or 24 months is there you get to leverage standard programming models all your C++ code etc java python whatever it might be like the bitter lesson of chips.
16:17 >> Yes, exactly. It's like the bitter lesson of chips is that specialization never wins. But in in this particular case, it was we have this one application or or a few a small number of them that would benefit tremendously and that would require an unimaginable amount of general purpose CPU to support. So really in 2013 it it was a bet and there were quite a few people even within the company I would say who were not sure that it would work out. So it was a bet. It turned out to be a massively successful bet. So the first chip was all all about inference.
16:54 the second chip was hey actually we can take the same idea and build a training chip. From from there it got picked up for more and more use cases. Around the time the second chip came out transformers were invented. I mean this this was a major major moment and that actually completely shifted the TPU program. we of course learned that recommener systems could run really really well on the TPUs and so ads and related use cases came along. So I would say that change and evolution has been one of expansion in scope and impact. In other words, it started with a really impactful use case of a couple of applications for inference serving. This was language translation primarily and voice recognition to then training to then transformers recommener systems and continuing to generalize as I guess the genai moment took a hold to ever larger ever more scalable systems as well.
17:55 >> You mentioned moment the transformer came through being a a major moment. I'm curious to explore the TPU was like a specific architecture but it wasn't you know specific to the transformer specific. how do you kind of straddle the fine line of you know how specific you want your chip to be for the workload? >> This is a great question and I think that it really comes down to the applicability. I mean I think that of of the chip to a particular workload.
18:21 So in other words in the end we've been thinking about this for every generation. Would we further specialize? The question that we were facing let's say two years ago or a little bit more was do in 2026 should we have two chips or one? We could have one chip that could do inference and training simultaneously and do both quite well or we could have two chips. One that was further specialized for inference and a second one that was further specialized for training. This analysis and this work led to the release of two chips this year. 8 I for inference 8 T for training and in the end what we realized is that it wasn't the case necessarily a couple years ago but by 26 we saw inference and serving really taking off and so having a chip that would be significantly faster for serving that we thought might be 30 40 50 60% of the market in its lifetime started making a lot of sense for us relative to if if inference if we projected inference to be 2% or 5% of the market even if that specialized chip is let's say 2x faster it might not make sense right because you'd actually go for the one general purpose chip that's not fully optimized but that that's okay because you now have the uniformity etc. So really it does come down to a calculus question of okay would you further specialize? How big is that workload? How big is that workload projected to be in a couple of three four years? Is this a sustained growth for a particular workload in hardware? Again as you know there is this opportunity where the more you specialize to a particular workload the less flexible it is the faster the more power efficient the hardware is going to be. So it is this art and it's this projection of what are you designing to and how persistent is that workload. In other words, if it's going to go away after a month or two months or three months, even if it's big for those three months, you got a really narrow window to intercept it. So it has to be somewhat durable. and you have to be able to project exactly what win you can get for specializing to it. So the rough trade-off is there's a fix a large fixed cost to support a new program.
20:40 >> Yes. >> And you basically have to think that there's going to be enough demand for that specific program to justify the cost. >> Exactly. And to some to some extent, you know, for example, for our 8i and 8t chips, both chips can do the other workload. This is also key. If 8i could only do inference and it couldn't do training at all. In other words, it had zero performance for training. And then vice versa. If it was amazing at training and had zero performance for inference, that would have also been a key limitation. Why? Because we would have to predict ahead of time over a six-year period lifetime of the hardware exactly how much we need of each. In this case, it was nice because both chips are better at what they're specialized for. But if you needed to, if you had leftover capacity in one place or the other, both can actually do the other other's job. depending on the specialization you do, you might get so specialized that you actually limit yourself from being flexible, funible.
21:34 >> And if the rough trade-off then is kind of size of program on the other end, like it seems to me that so much of the modern AI market is transformer-based. And so maybe just asked a provocative question, why not just burn the transformer architecture into the chip? >> Yeah. So I think that it's transformer based but then the next level of question would be transformers in the end are about vector and matrix multiply operations and you know a soft max. So there's a range of essentially linear algebra primitives.
22:06 We have those and others do as well roughly baked into the hardware. It's all transformer based but then it's your model architecture is okay how many layers do you have? How do you go through the layers? How do you go across the layers? how for each dimension what exact shape of matrices and vectors are you applying you you could go further and specialize to not just the transformer but your model that's the next level of specialization I think there are a number of companies out there that are thinking about that I think it's a very very interesting direction as well >> okay so at this point in time do your customers view TPU and GPU as roughly fungeible equivalents or is is there a certain set of problems that is better suited for one or the other.
22:48 >> There's for sure a better set of problems that are better suited for one or the other. I mean GPUs are more general purpose than TPUs for for one. That that is clear. I there we there are incredible products at Google and Google Cloud. we we sell a lot of GPUs. We use GPUs internally. But I think it really then comes down to the specifics of your problem. I so I think that our customers there there is overlap between the two but our customers then evaluate their workload and evaluate their options. What we like to do at Google is give our customers choice. In other words, we want to have the right solution for their needs and of course provide the solution that that best meets their needs for them.
23:30 >> What is the case for code design from the chip to the network to the software and then what is the case against codeesign? Yeah. So there's huge optimization opportunities with codeesign. And so what you can imagine is that if you have a end layer stack and you want to be able to pick and choose whatever component you want, let let's say you're running across many clouds or you're running across many pieces of hardware, many pieces of software. You could design abstraction layers that would basically say I can run on any hardware. I can run on any software. I can run on any network topology any amount of network that you give me no problem and my system is going to be fully adaptive very likely. So now you have a great capability. You can move anywhere like new capacity becomes available overnight. You're up and running because you've designed your system that way to actually be able to take advantage of anything. You have not hardcoded or specialized at all to anyone's particular infrastructure. Downside is you probably leave a lot of performance on the table if you become fully funible. If you become fully flexible to anybody's hardware, software, network, storage, compute, etc. stack. So the case for code designing is between each one of those layers there's a big impedance mismatch if you try to get full generality. And so there might be 10% 20% 2x across each of these layers.
25:00 You start multiplying those optimization opportunities through and all of a sudden you're left with a big endto-end opportunity in terms of whatever you want to say intelligence per watt or goodput per watt that you can leverage even again all the way down to power delivery and power availability software optimizations etc etc. So pros are you can run anywhere anytime you have no lock in. Cons are you're leaving significant and in all likelihood significant amount of performance on the table.
25:35 >> So my understanding is that if you take an open AI and enthropic OpenAI was kind of primarily building on a homogeneous comput stack and Enthropic was building on a roughly more heterogeneous comput stack. Do you think that's part of the reason that code design is part of the reason why they converged on what is rumored to be very different architectures for their models? I can't I can't speculate. I can't I don't want to speculate in terms of what open and anthropic are doing.
26:00 It's is one one possibility. but I think that without knowing the details of what they're doing, >> I would imagine that there could be many reasons for why they wind up with different architectures. >> Well, what about for Google then? I'm curious what the working relationship looks like between your team and DeepMind. Kind of who's in the room at what stage of model development and how are you making decisions together on code design? It's one of the most fun and frankly gratifying parts of being at Google is the opportunity to work really shoulder-to-shoulder with the deep mind team in terms of co-design of our hardware and models.
26:36 there's a third element to it in in terms of that we also get to extend that with the consumer services and cloud. I'll put that aside for a moment. I can come back to that. But with respect to deep mind, it it really is a deep partnership. So let me I mean from the past I can give you examples where they have come up with model optimizations let's say to transformers or to particular math that they might want to do and we might have a chip in progress not quite done but then they say oh my gosh if we had hardware support for this our endto-end training or serving might get significantly faster more efficient what would it take to actually now change our hardware definition that we might be an execution on to accommodate.
27:23 And this then leads to our engineers from and researchers to getting together intensely in the same room for a few days, a week, two weeks saying, "Okay, yes, we can do this. More likely, we couldn't quite do what you wanted, but we can do this other thing." And then maybe you can change your model architecture in this other direction that gives you 98% that gives us 90% of what we were looking for. And yes, we can then go back and intercept the hardware. Andor we can say, you know what, we're going to delay the tape out by a week or two weeks because wow, for this level of benefit, it's it's totally worth it. You know, similarly, when we're projecting our road map out and it's time for, you know, we we we have many generations of chips that are essentially in progress at any point in time. We have the chips that are in production. That's one. We have the chips that we're working on getting into production. They're already back from the manufacturer and we're debugging them, making them work. We have the chips that are in implementation that are about to tape out and go to the manufacturer. We have the chips that are in design phase and then we have the chips that are in concept phase. So it's really this five six stage many many year pipeline from in production to in your mind etc. The collaboration with deep mind for in production is significant because we get to work together in maximizing delivered intelligence or delivered goodput per watt and we know exactly what's happening in the model and we know exactly what's happening in the hardware and everything in between. We also get to collaborate deeply on though the chips that are actually just about to tape out. Why? Because we can intercept and we can make changes to the chip literally the chip architecture in flight which would be somewhere between hard and impossible to do if we were working across company boundaries. not not impossible. It would be much harder, right, for us to say, "Oh my gosh, we're a few weeks or a few months away from getting this chip done and now let's get in the same room shoulder-to-shoulder and figure out if we should disrupt the program." It's possible, but but harder is what I would say. And then of course for the chips that are in design, we have many architectures we can evaluate together. We can then ask our colleagues in deep mind, where do you see model architectures going in 2 three years time? Here's the paro of things that we can do. Here's the paro of where model architecture is going. We have actually deep and significant simulation infrastructures that can predict how the workloads are going to map to different hardware architectures.
30:02 Deep deep and fast iteration. It really isn't the teams are working separately. It's in in same building, same rooms many times. Deep daily interaction. I talk to whether that's Cororey or Demis multiple times a week etc. So it it really is a a super fun aspect of the work. >> That's awesome. And eventually their models will help with with chip design. >> We're using Gemini to design hardware for future Geminis as well. >> That's really cool. I want to come back to that a little bit later on the collaboration itself. I guess is there an impedance mismatch still of just I figure the cycle times are probably just slower when you when you're dealing with hardware than than than your deep mind folks get to deal with on the software side and I figure your planning cycles are much further ahead >> for sure.
30:50 >> and so how much room do you actually have to adjust? >> Yeah, it's it's a really good question. I we are planning hardware two three four five years in advance. No, no question. In other words, we talked about TP8 high and at just now that we announced, but you can imagine that 9 10 and maybe some others are are in in concept execution to to something else. And so they and they might be years out if you're working with model architecture by default you're not going to be thinking years out.
31:20 >> Yeah. But I think this is also the great thing about how the company has grown up together and you know there was words Google research and deep mind and transformers being invented all of this was happening in in these overlapping rooms as well. So in other words, there's a whole generation of researchers who've grown accustomed to being able to influence the hardware and knowing that the hardware is operating on multi-year cycles and they also know that look, if they have a small tweak that's going to deliver like 1% or.5% or something like that for a chip that's about to tape out, they're probably not going to come to us because they they they know enough to know that actually it's not like software where you can just do a change list and it's going to ship to production in two weeks time.
32:00 like it actually stopping a tape out is a big deal, but they also know if they've got a big like a really good opportunity. Yes, we're absolutely going to work together to figure out if we can get it in there. So, as I said, it's multiplicative in terms of where the benefits come from. The hardware does lift all all the tides, right? And so, significant portions of the Deep Mind team and we're so grateful for it. It's an amazing team, but significant portions of it are thinking about how do I influence the road map because it's actually a pipeline. Like the idea I had two years ago, like it's in production now and it's helping all the workloads at Google go faster. Like that's a good feeling.
32:42 >> Totally. It seems like an impossible task though to predict in five years what workloads will be most common and what algorithmic breakthroughs will have occurred. doesn't seem like an impossible task. I hear you that it would seem like an impossible task, but here's the awesome thing and actually we're working on writing this in great great detail and it's a lot of fun to do it. The stunning thing is that the TPU architecture at a medium level of detail, not at a super high level of detail, at a medium level of detail hasn't really changed since TPU v1. Like one way to look at it is the instruction set architecture for a CPU. like you've got loads and you've got stores and you've got ads and you've got subtracts and branches and like and yet what has the software on top of it done right over over that period of time. Same thing with TPUs. We have some fundamental instructions and fundamental primitives. Of course, we've extended it. It's not like the instruction set hasn't gotten changed at all, but the primitives there in terms of specializing the numeric very large matrix multiply units a sparse core that manages vector operations and scatter gather operations etc. There's five or six things that really define a load remote load store.
34:00 That's another one actually. We can read and write remote memory through our ICI network. the fundamentals have been there and they've extended just to many generations of models, many generations of even deep neural network algorithms and model structures. >> Okay, speaking of shifting workloads, it seems like one of the biggest changes in workloads over the last year or so. I think it really started at the beginning of this year, this calendar year, was kind of the rise of the long horizon agent.
34:27 >> Yes. And I would guess that's a very different shape of workload than kind of the quick turn LLM conversations of of years past. what does that mean in terms of data center needs? >> Yeah. So I think one there's two huge aspects to this. One is it's no longer human to whatever you want to say accelerate interaction. In other words, when you are typing at prompt on a on a web browser or on your phone or whatever, of course, in response to your prompt, there's going to be a bunch of work that happens, but then when the response comes back, you've got to read it. You've got to think about it, and then maybe you have a follow-up that's going to be multiple seconds of interaction time. Now, in Deep Horizon Agents, it's not there's no human in the loop that is going to naturally rate limits how quickly requests are going to go to the model, right? So this what went from you know seconds maybe tens of seconds in terms of inter interaction time is now going into perhaps milliseconds right as soon as I get a response back I can parse it I can reason about it perhaps a bit and I can figure out what my next request is going to be. So that's big change one. Big change two is all that reasoning and all that parsing is going to probably happen on a CPU. And that CPU is then going to have to probably think about, okay, what other state do I need to go gather before making my next prompt back into the model? Like in other words, I' I've learned something from this response.
35:50 I'm going to take another step, but I actually need to go grab some context. maybe from DRAM local to me, maybe from someone else's DRM on another CPU, maybe on SSD or maybe on HDD somewhere else. So now huge amount of orchestration has to take place as well. So the design actually is changing pretty significantly where the demand for accelerated compute is going up but the demand for CPU and networking and storage the you know traditional data center CPU etc is also going through the through the roof.
36:24 >> does that mean you're putting more CPUs alongside your G your GPU racks then? So this is yeah GPU TP racks. This is the key question is going back to this question of optimization and specialization. If we start putting lots of CPU racks and it is a question next to the GPUs and TPUs that means that actually we can't fully specialize to the density and network requirements let's say of a TPU rack relative to a CPU rack. TP rack is going to be more dense than a CPU rack. It's probably going to need more networking than a CPU rack. So in other words, now our building design is going to change.
37:00 Another option is maintain your uniformity. Put your TPUs or your GPUs all in one building and the building next door, but maybe put your CPUs and your maybe your and maybe by the way the hard drives have to be in another building on the other side because they have yet another set of requirements. Now you need networking between these buildings at a pretty significant level. So once you leave a building the networking complexity goes up significantly from a reliability perspective from a cost perspective latency goes up probably acceptably but still now you might go to hundreds of microsconds potentially more with queuing between the components. So the considerations do change in pretty interesting ways.
37:47 >> Your job is hard. >> That's fun. >> It's fun. Yeah. what's happening on the networking side and I've heard that Google's always been the at the forefront of the newest in networking including optical. can you say a word on what is like this the state of of optical networking? You know, we at Google, this was probably 15 16 years ago, were among the first to bring essentially what's called wave division multipplexing where you could put u multiple signals on a single fiber within the data center and we actually leveraged that for all of our communication between racks.
38:25 At the same time, we introduced a technology along with the wave division multiplexing called optical circuit switching. And essentially what optical circuit switching does is in contrast with traditional electrical packet switching, what it does is it transmits and moves data entirely in the optical domain. The great thing about that is and so let me describe the technology that underpins it in a in a packet switch. You would take a packet that has a header. You would look at it in the electrical domain. Figure out where the packet is headed in its header. It might have an IP address that says, "Okay, where do I head it?" You look up, okay, for that destination, you look up in a table, what ports do I forward it along? So basically, billions, billions and billions, perhaps trillions of packets coming through at super high speed per second. You're forwarding them along. Optical circuit switching says I'm not touching these bits in the electrical domain. What I'm going to do is I'm going to figure out for an input port which output port to send the light to. Okay. And now I have there's multiple ways to do this. The one that we started with is called MEMS switches, micro electrical motors that basically control mirrors in 3D. So we can now programmatically take a box that might have you pick 128 ports, 256 ports, some number like that. And we can configure every input port for a fiber that comes into it, map it to an output port, and then change the rotation of mirrors where the light literally shines down on these mirrors and gets reflected to the right output port. Initially, the reason we did this twofold was to create locality between groups of racks. Let's say I had a compute cluster and a storage cluster and both of them were in support of again making it up search. We knew that these two clusters would talk to each other a lot. So we would configure the mirrors to create shortcuts purely optically between those two clusters, clusters of racks. That was reason number one. Reason number two was we wanted to be able to expand the network and contract the network without actually moving any fiber. And without working into the details, I could draw this on a board. The optical circuit switch would allow you to actually reconfigure the spine of the network to expand it or to shrink it without a human being having to do anything other that' be literally a controller that would manage that. Now fast forward to TPUs. a TPU has a Taurus topology that connects all the TPUs to one another directly.
41:03 I I talked about throughput and goodput earlier. One of the things that we can do is if we have a TPU rack that fails, we can replace it with another TPU rack without moving any fiber. Again, I'd have to wave my hands or draw a whiteboard. but essentially, we can say we have a spare rack available at all times and when a rack fails, we redirect the light to that new rack >> and that can be done in milliseconds.
41:31 >> Why do you have the fiber at all then >> as opposed to fully free space? Yeah, it's a very good question. attenuation loss and the bandwidth would drop dramatically. F furthermore, across the range of a very large building, aiming everything without f without the benefit of fiber to connect it all together in 3D would be would be challenging to possibly impossible. we've talked about it. We've talked about it actually and there have been some really interesting discussions.
42:01 but but yes it's mostly in fiber but then when they hit the optical circuit switch essentially the light then literally shines down on these tiny chips. So that's one big direction of networking. There's but there's there's a lot I mean I would say networking is exploding in in terms of its u capabilities and need frankly in the data center. >> So cool. We could have an entire conversation on that. >> Yes, it is very very cool.
42:25 >> Even like the going back to the picture from ineffable it was all the the cables that that cut everybody's eyes. >> Exactly. Yeah. And those cables eventually some of them wind up at optical circuits which in our data centers. >> Makes sense. Okay. I want to flip to talking about power. You've been talking about goodput per watt and other units per watt. Makes me think per watt means power is kind of the the binding constraints or the scarce constraints or the expensive constraints in some way.
42:47 >> You know, so I I I get asked this question, what is the biggest constraint that we face? And the reality is there is no single biggest constraint that we face. They're all constraints. They're all super hard. It could be and then they shift continuously. But if I had to answer fundamentally I would say that power is the single most fundamental constraint that we face like everything else seems like we know how to solve them and it's a question of solving them over some period of time. Power yeah I think the way that you put it is really nice.
43:18 It's a it's a binding long-term issue that we do not have a I mean there you know nuclear you know perhaps nuclear is going to abundant clean energy which would would solve a lot of problems for for sure when that happens and at when it happens at scale still still unknown >> and so how does it work in practice you you know you're standing up a new data center it needs a gigawatt of power I imagine you can't just call PG& and say hey please send a gig a lot of power.
43:48 So, h what does provisioning power actually look like and are you having to actually vertically integrate all the way down to like, you know, doing your own turbines or how do you solve the power bottleneck? >> Yeah, it's again it's a it's a big question, important question. We do our our preferred model at Google is always to be utility connected to be good connected. So, it's the equivalent of calling your >> Oh, so you can't just call PG&E. You exactly you call ve very politely your favorite utility wherever it is that your data center is and of course you give them many many years of notice.
44:22 So in other words, if we're talking about a gigawatt scale, it's not something that you can say, "Hey, I need a gigawatt tomorrow. When when can you start billing me?" it's something that we co-l plan together. You know for for us it's something that we also take very seriously from the perspective of ensuring that when we work with these utilities the the costs of putting that infrastructure in place the we could get into a long conversations long conversation just on this topic as well that we cover those costs because with the way that billing works it actually could be that by the act of the utility building out capacity for us let's say other people's rates could go up in theory what we ensure is that actually let's say the transmission lines that have to be built upgraded ated additional utility based stations etc that we pay for those as well but so it is a long planning process it can absolutely be the case that in let's say that we need a gigawatt in I'll make up a date 2028 and the utility can get us a gigawatt in 2029 that can get us let's say 700 megawatts in 2028 so now we might be left with a question of how do we cover those 300 megawws well one answer is just Another answer is to say okay well would we figure out how we generate some subset of that power ourselves and maybe that would be with solar cells or with with batteries as a backup etc. Would it be for other sources? So then we again work with the utility. So it might also be a combination where we might maintain some power generation local and have the capability even when the utility comes online fully at let's say the gigawatt scale where we could actually provide power back to the grid. So having that when they need it as well, right? So in other words, there might be the one way to look at it is there might be the two weeks of the year maybe it's the hottest two weeks of the year where there's huge amount of residential demand. If we have some local generation, we can then provide that back to the grid as well. So it really is working over multiple years with the utilities for us. Why is your preference to do that versus to kind of go like to work with the utilities as opposed to to kind of vertically integrate yourselves?
46:32 >> Yeah. So, the main main reason is flexibility and uplift on both sides. So, one one way to look at it is statistical multiplexing or if you want the law of large numbers. If we need to have a gigawatt of power, let's say we want to have that with 99.99% plus reliability, probably means we have to build two gigawatts of power, right? At that level at 99.99 or 99.999, you have to have oneplus 1 redundancy.
47:02 That that gets expensive. And then making sure that that's going to be ideally a clean energy source probably right next to our data center, that can also be challenging. Now if we partner with the data center again maybe we have some some amount of that that we can bring ourselves that we can give to the grid when they needed less that we can take their power. So in other words by leveraging that statistical multiplexing over a much larger base actually everybody wins we win the the grid wins residences win etc. We prefer that and in rare rare cases we we will do it let's say behind the meter but even then we're doing it under the assumption that we're going to work with the utility where it might be a year later we we want to be connected to the grid.
47:46 So again it's an uplift for us it's an uplift for the grid. >> Makes sense. How do you decide how big to make a data center? >> Yeah that's that's an art and it's a a source of significant debate. You know, at some point, you know, we had a debate. I remember even 10 plus years ago, 15 years ago, big debate at Google. Should we just put everything into one data center? Like back then, it was going to be a gigawatt, which was >> simpler times.
48:09 >> Yeah. Different times, like 105 years ago, we're going to have a gigawatt data center. So, one obvious concern there is single point of failure, right? If you're looking at things from a 30-year perspective, that's a a huge concern. You know on the other hand if you're talking about training workloads today bigger is better in other words having more because from a networking perspective actually you want to have things concentrated in as small a distance as possible but then again two issues single point of failure but then also now power availability right while a gigawatt was huge 10 or 15 years ago but imaginable there's no way anywhere in the country or the world we're going to get the total requirements of Google built in in one place. it's just just not going to be possible. So now what's the optimum size? again we have models for this simulators etc. But this it it also depends on look in some places we're going to be at the edge of the network or we're going to be in country that might be tens of megawatts. We actually partner with ISPs that might be Iraq like literally it could be Iraq. This is a training cluster. Okay maybe that's going to be closer to a gigawatt.
49:20 some other sites might be hundreds of megawatts etc. >> Interesting. I'd love to understand how you think about the portfolio life cycle management. so I guess my guess would be that you know the biggest newest clusters are used for training the latest frontier model and then you kind of recycle the older stuff and run inference on it. Is that is that like a fair framework? Are you are you also standing up inference specific clusters? How does that all work? very reasonable framework that that you're thinking about that and I think makes a lot of sense but I would say that the demand for inference is such that we can't just rely on whatever older training clusters that are no longer being fully used just for training as the basis for inference and also if you think about it we might centralize again in a particular year into a small number of large sites our training clusters and there is benefit to keeping the network distance between them small so now let's wherever they're located in the world is they're probably going to be let's say on the same continent or they might be even in the same portion of a continent in a particular year. So now you might be left with not enough serving capacity on the other continents etc. So then okay actually we then have to go build the specialized inference elsewhere across the world. So I I think your intuition is spot on but not wholly. In other words, it really does have to be okay here are the training clusters. Yes, probably they're going to be used for serving in some number of years. but then we're also having to build the serving clusters as well.
50:50 >> And your serving clusters are they different than the training clusters? Like are they smaller? Is there are they are they cheaper per per megawatt? not necessarily cheaper actually because for serving there is more need to colllocate compute networking and storage and so that then we get to the inability to specialize. So for training you actually have this big density uniform deployment etc. whereas for serving you're going to have to have the mix of storage compute and accelerators.
51:21 Furthermore, for it's a very interesting aspect of this that we could go into more detail on as well, but for serving, you actually don't want to have too much in one place. You want to be serving your workloads from all over the the planet, but now we wind up having individual model endpoints and that actually can be variations of models. So now we have to spread these models out across the world also accounting for locality. So they will be smaller. They'll actually be less vertically integrated probably on the infant side.
51:55 >> Super interesting. So if you made a data center 5 years ago, you know, state-of-the-art accelerator 5 years ago was very different, >> vastly less efficient than the ones that are being made today. >> Yes. >> Do you actually go back and like swap out the the chips in those in those old data centers? And I know this kind of relates to there's been an ongoing debate, I think, of like what is the actual useful life of a chip?
52:16 >> Yeah. Yeah. So, you know, I've been on on record of saying this and one one of my more I was surprised by how how much pickup this had, but but you know, I said that our seven and 8 year old TPUs are still at 100% utilization. >> Wow. >> and and >> I think I saw this. I think that was >> Yeah, I didn't mean for it to be a a meaningful statement, but it was a meaningful statement apparent. Yeah. So, our older TPUs GPUs, but our older TPUs are seeing significant utilization.
52:44 We do replace them in the end. there's a question of okay have they depreciation lifetime is approximately 6 years so once they're fully depreciated and given the power efficiency of newer generations etc it does make sense to replace them and upgrade them it's not that we I do replace the chips but it's really the the the systems so in other words we think in terms of pods so a TPU let's say 8 pod might be 9600 chips and it might be 140 something 152 to racks etc. So then we would say, okay, we're going to pull out that pod and we're going to replace it with a whatever it is, TPU 12 or 13 or 14 pod. It might not be a perfect fit and then we have to, in other words, the new pod's footprint might not be a perfect fit for the old pod's vacancy. We we then just have to account for that. We have to figure out how we would retrofit. We can't plan that because we don't know what that many generations of TPUs are going to be. so then it's again a bit of an art to figure out how we would and it's actually hard work figure out how we would decom and then replace as quickly as possible with with new TPS.
53:53 >> You've written publicly about open standards. can you say a word on that? >> Yeah. So I think that's the we we talked about this interoperability question and so for us while we support vertical integration and we make it possible to you know extract just as much performance as you want for us it's really important to not force lock in and not force a sort of walled garden in terms of how the system works end to end. So you know for mean I'll give one example at Google we've developed a model development framework called Jax we like it a lot we think it's really really good and we use it extensively internally many of our customers like PyTorch right and so one one thing we could say is hey if you want to run on TPUs you got to use Jax it's the best I'm being facicious a little bit but maybe it is maybe it isn't we think it's the best and that's only choice because it's it's so great or we could say look if you like Jax we like Jax if you like Jax you can use it but if you like PyTorch we have torch TPU where your unmodified models etc can can run we've and you know we've seen this throughout history actually with many many examples I've used the example of IP why did the internet protocol win there were actually many competing protocols to IP this is ancient history in the 70s and early8s piece. Why did IP win was because it was open standard and it was this narrow waist to the hourglass where anything could run on top in terms of software and anything could run on underneath it in terms of hardware. So open standard interoperable anyone who brought IP could plug into a router port and become part of the internet like that. It was beautiful and that's what allowed the internet to explode in its growth and reach across the world. So we really do believe in those open standards in terms of plug-in points. Now if you want to plug in something highly specialized, you can right if if you think that you've got a better thing to plug into our our framework, you absolutely can.
56:03 But we want to really support open standards ideally open source around it as well. >> It's too big and too important of a buildout to be you know any one vendor's closed proprietary st. Yeah, we really we really believe that and it's got to be one of choice and that's also why for example we fully support and have TPUs, GPUs, other accelerators etc. >> Yeah. How has day-to-day life for for your team changed with AI? Where is it changing your your function the most?
56:31 >> So I think the the whatever you want to say the easiest answer is on the software engineering side where I mean it's been documented externally quite a bit. so I think that we are at Google and my team using it to great effect in in terms of our software development capabilities but also frankly test rollouts even helping with design etc. But there have been on the hardware side which has maybe gotten less external coverage also significant change. In other words, my my hardware engineers, I was just looking at the numbers earlier today, in fact, are using AI using whatever the the token counts if if you want, not the best metric, but it's still a metric.
57:16 Our hardware engineers are using AI as much as the software engineers are. And productivity has gone up. Time to you know, time from design kickoff to tape out is shrinking. Time for bringup is shrinking. So another where's productivity is going up significantly but then other maybe less expected changes how we do data center design has changed significantly. In other words how we and how we evaluate you you mentioned hey do we do a gigawatt building or 100 megawatt campus or 200 megawatt campus. In the past, this would be very detailed, very sort of spreadsheet driven, human-driven processes, and it still is to some extent, but there's a lot of AI now involved that really streamlines the process in terms of planning and development as well.
58:09 >> Interesting. So, >> like reasoning models or >> I wouldn't say quite yet reasoning models. It's not it's not replacing human judgment, but it's making it much easier to bring the necessary information together like in in in one place and basically put the right information in front of the humans making the decisions. >> Okay, I'm going to bring us home with two kind of out there kind of fun questions. >> Yeah. >> question number one, orbital data centers. I've seen that you know Google is thinking about this quite seriously and you want me to comment on that but I mean does the fact that people are seriously running the numbers on orbital computes mean that there are kind of serious binding constraints on Earth and and what do you think of of orbital data centers?
58:53 >> Yeah, it's an exciting direction. We are actually pursuing it and we've referred to it with no tongue and cheek as a as a moonshot as one of our big big efforts that we're excited about investing in. goes back to the earlier part of the conversation where you raised I think correctly that in terms of fundamental constraints energy and energy production is a key key challenge and so really bottom line is from a fundamentals perspective in space you have something like 40% more power available because of lack of attenuation in the atmosphere etc. In other words it's just more solar capacity what 1.4x 4x. So that's one part of it. But but in a sun-synchronous orbit, you have 98 to 100% coverage of sunlight on your solar cells relative to 28 30 maybe 35%.
59:49 on land, right? So in other words, just this enormous amount of power, obviously the sun, you you take the 1.4x 4x and you take the 3 to 4x in terms of number of hours per day, you remove batteries more or less from the equation and now you have the potential for something that can really and and of course carbon-f free lots of benefits, lots of challenges. >> Yep. >> Right. Lots and lots of challenges. So in other words, okay, now what about cooling? you know, naively you might think cooling in space is easier. It's actually harder. reliability. We talked about how these things fail sometimes. So repairs is becomes harder not not impossible but it becomes harder in space robot up with the with the cluster >> and exactly so perhaps there'll be this model is a is a promising direction. redundancy will probably be your friend. your pre question on free space optics is now going to become reality. We're probably not going to be stringing fiber between these components. So it will actually be free space and we're going to have the lasers pointing at the receivers and calibrating in in real time. These are there are no showstoppers here. No fundamental showstoppers.
61:00 >> Last question. I've had this image of the ineffable big supercomputer that you built them >> in my head this whole time. And so here's a question. >> 10 years from now, what will the most frontier supercomputer look like? >> Oh my gosh. Yeah. 10 years is right at that edge where it's I would say in all sincerity probably some people are thinking about what the 2036 computer looks like. We might have a few thinking that far ahead but wow I the cone of uncertainty there is is way way too wide. Like in other words if you look at the trends the level of integration is going to be tremendous.
61:45 So, you know, we saw the beautiful fiber and the VR200 racks that we put together for ineffable. My my guess is it's going to be more integrated, less fiber to I won't say no fiber in 2036, but I think that probably from a rack perspective the rack will look much more tightly integrated and the amount of fiber there'll be maybe a small fiber bundle coming out of the rack. And then so in other words, one one view you might have from a modular manufacturing perspective is these racks likely in 2036 are going to be manufactured centrally. And so whether they have 72 or 144 or 288 or probably more you know 576 52 pick your favorite multiple of GPUs since we talking about the an ineffable case but then be the same for the TPUs deeply integrated into a rack would a rack be multiple megawatts you know we don't know exactly how but 2036 we get to we get to imagine things could be multiple megawatts in a single rack and now you bring water and you bring power and you ing fiber, right? You you you wheel that rack in, you plug those three things in and you're off off to the off to the races as well.
62:59 >> Okay. So, it's not going to be like a big alien orb then >> up in space by 2036. I wouldn't say that it's not possible for 2036 for it to be then directly launched into space and then maybe taken by a robotic arm at some space station and then plugged into the right module. Yeah, maybe that for 2036. >> Fun >> fun fun fun to think about. >> I really enjoyed this conversation where this is >> the most intense biggest capback spilled out in history, but I also think it's it's a technology >> revolution >> revolution that is a thing of beauty and I think >> you have such a grasp over you know the beauty of the technology and and the binding constraints and and how to balance this this all. Google is in good hands. Thank you for taking the time to share what you're doing.
63:48 >> Ton of fun. Ton of We get to live. We get to live through this and we get to define it. So So thank you very much. This is a great conversation. >> Thank you. >> Yeah. Thanks.
Summary
- Specialization in hardware leads to increased performance and efficiency but reduces flexibility.
- Google is investing heavily in data center infrastructure, with a projected $200 billion in capital expenditures.
- AI data centers differ from traditional ones primarily in their purpose-built design, focusing on specific workloads.
- The concept of "goodput" is introduced as a more relevant metric than theoretical performance benchmarks, emphasizing real-world workload performance.
- Collaboration between hardware and software teams, particularly with DeepMind, is crucial for optimizing AI models and infrastructure.
- Power remains a critical constraint in data center operations, requiring long-term planning and partnerships with utility providers.
- The rise of long-horizon agents changes data center requirements, increasing the need for both CPU and networking capabilities alongside accelerators.
- Open standards and interoperability are essential for avoiding vendor lock-in and ensuring flexibility in AI infrastructure.
Questions Answered
What is the impact of hardware specialization on performance and flexibility?
Specializing hardware for specific workloads increases performance and power efficiency but reduces flexibility. The durability of the workload is crucial for making such a specialization worthwhile.
How does Google plan to double its serving capacity every six months?
Google aims to double its token generation capability every six months through a combination of hardware improvements and software optimizations, rather than just increasing hardware performance.
How do Google's AI Infra and DeepMind teams collaborate on model development?
The collaboration involves close partnership where hardware and model designs are co-developed, allowing for optimizations that enhance efficiency and performance.
What is optical circuit switching and how does it improve data transmission?
Optical circuit switching transmits data in the optical domain without converting it to electrical signals, allowing for faster and more efficient data movement between racks.
What is the useful life of chips in data centers and how does Google manage older hardware?
Google's older TPUs are still utilized at 100% capacity even after several years, indicating that the useful life of chips can extend beyond typical depreciation periods.