Section Insights
Investment in Space and AI
What do the investments in the Apollo missions and AI buildout signify?
The $338 billion spent on the Apollo missions and the projected $765 billion for AI infrastructure in 2026 highlight the scale of investment in technology and innovation. This comparison illustrates the potential for significant advancements through substantial financial commitment.
- Historical investments in technology can lead to groundbreaking achievements.
- The scale of AI investment is comparable to major historical projects like the Apollo missions.
- Future investments in AI infrastructure are expected to exceed past technological endeavors.
Reliability Through Engineering
How has engineering improved the reliability of systems over the past year?
Real-world operational data has shown that the reliability of systems is better than expected due to rigorous engineering practices. This includes optimizing components based on failure data and enhancing redundancy.
- Real-life data collection has led to improved system reliability.
- Engineering practices have evolved to anticipate and mitigate failures.
- Continuous testing and optimization are crucial for maintaining operational efficiency.
Scaling Challenges in AI Training
What are the implications of scaling AI training jobs on system reliability?
As AI training jobs scale up, the likelihood of individual failures increases, which can significantly disrupt operations. The data indicates that larger GPU clusters experience more frequent failures, necessitating robust system design to handle these challenges.
- Scaling AI jobs introduces compounding failure risks.
- System design must account for interdependencies to maintain job continuity.
- Understanding failure rates is essential for managing large-scale AI operations.
AI's Impact on Software Safety
How has the introduction of AI affected software safety and incident rates?
The integration of AI has led to a significant reduction in incidents per code change, despite an increase in the volume of changes. This suggests that AI can enhance safety in software development.
- AI can improve safety in software changes, reducing incident rates.
- The frequency of code changes has increased dramatically, complicating incident management.
- Organizations must adapt to higher volumes of changes while maintaining safety.
Investment Projections and Future Outlook
What does the future of investment in technology look like?
The projected $1 trillion investment in technology by 2027 indicates a significant escalation in funding for innovation, comparable to the entire U.S. military budget. This suggests we are at the beginning of a transformative era in technology.
- Future investments in technology are expected to reach unprecedented levels.
- The scale of investment reflects the importance of technological advancement.
- We are in the early stages of a major technological revolution, akin to the space race.
Transcript
0:01 last talk, so we're going to finish strong here. I got a quick question to start this off. So, the first question is actually just a number. And that's $338 billion. Does anyone know what this represents? All at once, come on. Big energy, we've got to finish last. Any guesses? GDP of something, that's true. Any other guesses? What was that? Data center buildout, pretty close. Okay. This is actually something that's near and dear to my heart and something that I find actually very, very exciting. And that's the amount of money that we estimate that it took over the course of a decade to put a person on the moon. In the 1960s, over 500,000 people came together and we spent about $338 billion to put a person on the moon and bring them back safely. And what that tells us is with a large investment and with a great technology and great engineering, we can do some incredible things. We can put people on the moon.
0:56 I've got another number for you. $450 billion. Any guesses what that is? What was that? Manhattan Project. Ooh, that's an interesting one. Okay, no. This, despite being about 50% more than the cost of the Apollo missions and the Gemini missions combined, is the amount of money that we've spent in 2025 on the AI buildout. That's pretty massive. And interestingly, somebody said, you know, the GDP of a country previously. Well, just for some perspective, $450 billion is about the GDP of Hong Kong, home to 7 and 1/2 million people. And that's how much money went into the AI buildout in 2025.
1:37 But that was last year. In 2026, Goldman Sachs estimates that we'll spend $765 billion on infrastructure throughout the course of the year. Not in a decade, in 1 year. Twice the cost of the Apollo and Gemini missions inside of 1 year. Now, I want to spend some time going through what these investments have gotten us over the course of the last year, as well as what we're seeing in the industry and in Meta and what we've heard from many of our colleagues and peers throughout today. But one last show of hands, is anyone in the last couple weeks here used Claude Code or Codex or anything like that?
2:09 All right, I'm going to guess cuz I can't see anything based on the lights that we're at like 70-80%. And that's actually a pretty big difference because if you remember just over a year ago, the last time I was on stage at @Scale, things were quite a bit different. You know, the code Claude and Codex are sort of, you know, these things we almost take for granted today. But a year ago, the state of the art was quite different. In fact, it kind of looked something more like, I don't know, auto complete. You typed a few characters, you get, you know, a couple more characters after that, maybe get a few lines of code for free.
2:37 And things changed pretty dramatically over the course of those 12 months. In fact, it wasn't until the end of that month in May of 2025 that the first version of Claude Code was released to the public. But things changed really dramatically around October. And I think that's when all of us felt that seismic shift as Opus 4.5 came out and we really shifted away from something that looked like in-line code completion to agentic driven development. So in the last 12 months, we've gone from something pretty basic to a powerful tool that is now in daily use for many of you in the room.
3:06 And it's doing everything from planning my vacation that comes up after this to research, to writing, to code generation and more. It's building our infrastructure. But that's just scratching the surface. And I want to dig into what takes what it takes to make this all possible, the infrastructure, the engineering, and the products. Hello everyone. I'm Peter Hues. I lead production engineering at Meta and we're responsible for the performance and reliability of a fleet of millions of servers serving some of the largest products on planet Earth to billions of people every single day.
3:36 Srujan started us off earlier today talking about going from theory to real value. I'm going to close this out on going from theory to reality and all the things along the way that make that happen. Now let's do this by taking a look behind the scenes at that infrastructure development and build-out. I'm going to start from the ground up with our data centers. I'm going to follow this by talking about the compute that's involved in this, and lastly, I'll connect this all together with the systems and servers involved in that.
4:01 Now, let's start with data centers. Last year I was up on this stage, I started talking about some of the cool new things that we were doing to be innovative, to move faster, to bring us optionality. I might have been a little bit early in that. So much early that in fact we had to remove that from the recorded version of the talk. True story, I was talking about our investment in intents. Intents at the time, we had just deployed our first model of this into production. We had a few servers in this. We weren't even actually using these live yet. It was that early along the way. It was purely a theory.
4:31 Now, if you fast forward to today over the course of the last year, this has become a critical part of our overall strategy. Well, not only being fast in the absolute sense, it allows us to take advantage of our investments in land and power and network and use those much more quickly and on the order of months instead of the order of years. Now, over the course of last year we actually have a whole lot of real life information as well. Now, we have weather data, we have wear and tear data, we have all the sort of operational experience that you gain over a year of running these things. And one of the things that really surprised us is the reliability is actually far better than we expected.
5:07 And that wasn't by accident. That was the result of real hardcore engineering. Electrical, mechanical, software, systems, everything coming together to optimize that system and make it better. We used our DR planning and fault injection and systems to disrupt it and try these things out to try to create failures. We experienced weather events. We also experienced construction and accidental events. All of these things gave us data on how to understand which components would fail the most frequently and with the largest blast radius.
5:35 That meant we could tune and optimize the things that we repaired, that we increased redundancy on, as well as our systems that improve this is on top of that, whether there's a control plane, our training systems, and so on and so forth. And all of this is actually built on top of the history that we've all done in building large-scale distributed systems. You know, whether that's component failure or system failure or services failure, we've designed these things to fail and built around that.
5:58 With these tents, we assume that they'll fail, we test it, and we iterate forward. Now all of this engineering and investment means that now, 1 year later, that theory is reality and serving hundreds of megawatts of production capacity every single day. But to really move fast, tents weren't enough for us. We needed to build into the cloud. We had to build one of the largest hybrid clouds in the world, and you heard from Sarupa and Dinesh and Sargun about our efforts there. And inside of 1 year, we went from just a small cluster with a few thousand GPUs to tens of billions of dollars invested with our partners in this. It's massive.
6:32 And the exciting part about this is when that capacity is handed to us, we can bring it online and start serving real production traffic not in months or weeks or years, but just in a few days. But it wasn't all easy, it wasn't all for free. In fact, one of the things that we found that was quite unexpected is for many years we've actually taken a lot of the things in our homogeneous environment for granted. One of these is the cross-sectional bandwidth we have between our data centers. In effect, it's essentially unlimited. But as we go to the cloud, we find that bandwidth to any given provider, between those two given providers or and the to those providers and the internet, obviously varies quite a bit. So all these systems and services that have optimized for that large cross-sectional bandwidth that's consistent region by region, data center by data center, had to be changed and optimized.
7:16 We also went from a single homogeneous hardware platform to something that's unique per provider. Each provider has their own systems, their own software, their own components, their own firmware, and all of these things add up to increased complexity that we have to modify and iterate and change our stack to accommodate. But we didn't stop there, cuz that wouldn't be enough. In July, Prometheus, our first large-scale multi-gigawatt cluster, came online and is now training some of the largest models in the world.
7:44 And all this together actually gives us a really unique strategy. And key to this strategy is abstractions. Cuz whether that capacity is in a tent or in a cloud or in a multi-gigawatt facility, we abstract that away from the end user. Regardless of that capacity is using GPUs from AMD or Nvidia or own custom silicon like MTIA or it's using CPUs from AMD or ARM or Intel, all of these things fade into the background because for the products and services that use it, it just looks like compute.
8:16 It's pretty mind-blowing. Now, across the fleet, all this has added up in this all-in-one approach has given us over 1.3 million H100 equivalent GPUs at the start of 2026 and that is increasing at an increasing rate every single day. Massive stuff. Speaking of GPUs and servers, let's dig into that a little bit as well and see what we've seen over the last year. Starting with compute, this is our GB200 rack. It's called Catalina. There's two compute racks in here and then combined they have 72 GPUs and 72 CPUs and they're surrounded by four air-assisted liquid cooling racks. And I think if you heard from the folks at Nvidia earlier today, these things are quite complex, but they also have a lot of new novel technologies in them. NVLink, one of the things they talked about in just that specific example, has thousands of custom copper interconnects inside of these racks and backplanes. And many of this has new feats of engineering that really means it means that we have to change the model of how we think about reliability in a system like this.
9:12 And if you go back for the last 20 years or so of how we thought about systems and reliability, we've designed these things to fail. If any given single component fails, we don't really care. The large cluster continues to operate and things move forward. Only the impacted machine and jobs on it are failed, everything else keeps running and those jobs move somewhere else. The impact is generally pretty isolated. But with large-scale training that changed pretty fundamentally.
9:36 Because it's not just a single In fact, it's many systems that are interconnected. All of those things are interdependent on each other. All of those thousands of cables across thousands of racks, any failure at any of those, if you don't optimize your system, can take that entire training job down. Now, to accommodate this, we had to redesign our software and think of these systems not as individual servers, but as a large cluster and a large collection of machines with complex interdependencies.
10:01 The model for this starts to look a lot more like a network topology instead of a typical distributed systems topology. But why does that really matter? Why is it so important that that job keeps running? Well, fortunately, Alibaba just released a paper about a month ago on their train mover solution for this, and it had some really excellent data points in it. In fact, I'd encourage you all to go read this. And what they found is as a job scales from hundreds to thousands to even tens of thousands of GPUs, those individual failures compound and increase the interruption rate. For some perspective on this, GPT-3 ran on about 1,000 GPUs for about a month.
10:34 Llama 3 at Meta ran on about 16,000 GPUs for a little over a month and a half. As each new model comes along, it generally requires more GPUs and longer run times. Today's frontier models are running on something like 100,000 GPUs, and that's where these compounding really effects really add up. Because at 1,000 GPUs, you expect to see a failure once every eight or so hours. At about 16,000 GPUs, that drops down to about 2 and 1/2 hours. As you start to approach hundreds of thousands of GPUs, you would expect to see a failure every eight minutes. If that job fails every eight minutes and doesn't restart immediately, you're dead in the water. You're not doing anything.
11:10 And their data shows this pretty clearly. In fact, as that job scales up, you start to lose the effective training rate of the system. And you can in effect end up with idle GPUs on the order of hundreds to thousands to even tens of thousands of GPUs. As you approach 100,000 GPUs without optimizations, without changes to how you restart, without making this better, you can expect to have over 30% of your capacity idle at any given time.
11:34 In a neighborhood of 40,000 GPUs sitting and doing nothing. If you just do the rough math on the assumed CAPEX of that, it's like close to a quarter billion dollars doing nothing. If the dollars weren't enough, I told you in the beginning this is a race. That means that we're falling behind. Every time that job starts is time lost. We had to improve this. We had to get a lot better at it. So, we went back and we started thinking about this from the very basic level.
11:58 We had to reduce the frequency of failures. We had to reduce the time it takes to recover from a failed job. We had to rethink everything and all those component level failures from memory to CPU to GPU to control to IO and so on. And we work closely with our partners to stress test the software and firmware and all the components involved in this. And we work closely with them to improve their diagnostics along the way.
12:19 But we found that even our basic things like our own repair methodology was interrupting jobs as well. Because that same mindset of hey, if we detect a failure, we need to get it out of the fleet as soon as possible. Well, that in and of in of itself would actually interrupt a job. So, what do you do? How do you reduce failures when they're hardware and real and causing problems? Well, and I think this is an area where we have a unique advantage. And that's that we control our own destiny from the silicon to the server to the model to the training software and everything in between. And what this means is that we can do something that few others can.
12:51 Because we can see the issues that occur at the lowest level of the system and we can see when it's degraded or when it's impacted. We can also see from the job level when it's been interrupted and it's having problems. And we can marry these things together so that we can understand that when a system is having issues and if that affects the job, then we want to do something about it. But if the server is having an issue but it is not affecting the job, we can just leave it in the fleet and let it continue to run until something else naturally occurs like a cable failure and takes out the cluster. At that point, we can go in, we can identify those broken systems, we can swap them out with new hardware, and the job can keep going. Or I can repair them. Whatever we need to do.
13:27 Now, the collective result of all these optimizations from the software and the firmware through improved diagnostic in partnership with an NVIDIA to optimizing our operational systems to doing this across many generations and building on that compounding knowledge from A100 to H100 to GB200 to GB300 and everything in between, it all adds up. In fact, what we're finding is that we've been able to reduce the amount of failures and job interruptions by something on the order of like, I don't know, 95%. And yes, there's fireworks because it's almost July 4th, so there you go.
14:00 This is pretty awesome stuff. Okay. Now, as we move up the stack from the compute and into the software, you heard from a lot of folks today. You heard from Serupa and Gaurav and others who mentioned that we're moving from a world where software is written by humans and for humans to one that's written by agents for agents. And most of the code in our systems today is actually being written by AI. And I think all of us have experienced this massive increase in developer productivity and output as a result. I think Joe earlier showed that graph that, you know, was telling us that there's about an 8x improvement in developer output. That's something that looks like it can get to us 10x and maybe in a 100x over time. It's pretty amazing.
14:34 But interestingly, and I think actually counter to the popular narrative on this, we're actually finding that each of those changes inside of software is actually safer as a result of AI. Not more dangerous and not causing more incidents. In fact, if we look at this over time, in the course of 2025, the amount of incidents caused by changes in our code and software was relatively flat. Meaning for any given, you know, change that we made, the amount of incidents that we had was pretty constant over time. But that started to change dramatically, especially in October and November and the beginning of the year as we introduced more and more AI. And what we found is that every given change became dramatically safer. Over the course of just 1 year, we're seeing an over 50% reduction in the amount of incidents per any given change.
15:17 There's a caveat here, and there always is. While this is on the order of any given change, the amount of diffs and changes has gone up dramatically. So, even though safety is increasing, the frequency has increased at a far faster rate, which means incidents are up overall. And this is not just a meta-specific issue. In fact, GitHub had a post about a month ago about how they had planned for a 10x increase in capacity as a result of AI-driven development, and they are finding that they need 30x increases in capacity to keep up because the volume of code and the volume of changes is that great.
15:49 But, it's not just the volume of changes making this more difficult. It's the time that it takes to deal with any one of these given incidents. And we can track the time our developers spend on what we would call toil. And that's dealing with the operational results, whether that's incidents themselves or the follow-ups from this or the code that's necessary to fix it. And what we're finding is that over the course of the time that every one of these incidents is actually taking humans more time to diagnose, to resolve, and mitigate. In fact, year over year, it's increased over 120%.
16:17 So, you're spending an hour before, spending an hour and 20 minutes, 2 hours, 3 hours now at this point on all the same incidents. And that's why the work you heard of them from David in terms of agent guardrails, the work from Gaurav on agentic operations, the work from the folks at Microsoft and increasing their debugging is actually so important because we need to drive that down and reduce the cost of this. All right. So, with that, I've given you a pretty good summary, hopefully, of data centers, of compute, of the software systems involved, and where those investments are going across our industry. I started this talk also talking about Codex and cloud code and the kinds of things that we're seeing in reality. And at Meta, you can actually see that reality in our products as well, whether that's Threads or WhatsApp or Instagram, you're starting to see all that agentic behavior and all those new AI capabilities. What I'm particularly proud of is Ray-Ban Metas, and a new program that we just launched that Sheryl mentioned earlier before, and that's to give Ray-Ban Metas for free to all the vision-impaired vets in the the States.
17:16 Now, these glasses can take advantage of the cameras and the microphones and they can effectively see the world around them. They can capture images and audio and video of the what these people are looking at. And they can send that data back to our inference systems, which goes back to our servers, which eventually goes back to our models and it can process that data and then it can give those people a clear description of what they're seeing and what they're interacting with and in a sense it gives them an ability to see again.
17:43 That is pretty awesome and that is something is only possible when you have that control of your destiny from silicon to the products and everything in between. And that's not just technically exciting to me. I think that's really meaningful and important to people in the world. Now, before I close this out, I wanted to kind of come back to the start of this. I wanted to come back to that picture of Neil Armstrong on the moon because I used, you know, an example in comparing this to the Apollo mission and I explained that we're spending, you know, almost twice the amount of money that we spent on the Apollo mission in 2026 and I talked through all the things that we're getting from this, you know, from one single prototype of a 10 to hundreds of megawatts in this.
18:19 From a single hybrid cloud provider to tens of billions of dollars invested in this. From 100 megawatts of capacity to multi-gigawatt facilities all around the world. From primarily H100 and A100 GPUs to AMD and GPU or sorry, GB200 and GB300 and significant improvements in job reliability and things across the board. From in basic inline code completion to cloud code and code X and really identified software development life cycles and ultimately to products like Ray-Ban Meta that gives people the power to see again.
18:52 Now, all told last year was pretty incredible and the reality, the results, the impact is clearly coming into picture. And going back a bit further, last year I ended my at scale talk just asking where we were on that sort of overall AI timeline. And I said, you know, what? We're not at the middle of this curve or the end of this curve. We're just at the beginning. It's the next 20 years where we're going to see this really massive change in our industry, in our technology, and in the lives of all the people that we interact with. And you would think that given how much we've progressed in the last year, and how much faster that was than I think all of us in this room thought, that we would have moved a much further on that line.
19:29 But I think the reality is, if anything, I was probably being a little bit too optimistic. In fact, we are not at our moon moment. We're far from it. I think the best analogy, in fact, is we're at the beginning of the space race. We're at our Sputnik moment. All right, this is just the very start. And what that means is the coming weeks and months and years, is we're going to see that investment really take off. And we're going to see the results of that increase dramatically over time. In 2027, the estimate for CapEx is close to $1 trillion. That's That's one followed by 12 zeros.
20:02 Massive, massive amounts of money. Almost impossible to think about. In fact, just for one more perspective on this, and just to give you some context to the size and scale of what a trillion dollars is, that's the combined net worth of exactly one person. Jokes aside, I couldn't handle it. >> >> But in reality, it's the size of the entire budget of the entire United States military, all branches and all 1 million service members. That's the amount of money that we're about talking about here.
20:34 And well, throughout the course of today, you heard about the progress in, you know, fault detection and remediation. You heard about agentic operations. You heard about building infrastructure for agents. You heard about all the real-world experiences and some of the challenges that people have dealt with with outages and incidents and problems that have been caused. And then you think about that and the progress we've made and where we need to go, and I think back to Sputnik. And the distance between Sputnik and the ground is far less than Sputnik into the moon.
21:01 And that's the work that we still have in front of us. Because when I think about training and fault tolerance, yes, we've done a ton of work, and yes, we've done a ton of progress, but guess what? Those failures still interrupt those jobs. We have to get much better at fault tolerance in our training algorithms. We have to get much better at our detection and remediation systems. We have to get much better at our gigantic operations, as well.
21:20 And all of that changes with every new GPU platform that comes out. We In many ways, we've cracked the software coding problem, but as we all know, code is only part of our job. There's deployments, there's operation, there's mitigation, there's remediation, there's long-tail migrations. I was just talking to somebody about at lunch. These are all parts of our job that agents have still yet to solve and still have a long way to go to help us.
21:42 I spoke about the challenges of hardware constraints, and I think many of you who have been involved in this for the last couple years can talk about how GPUs have been holding us back and how we haven't been able to find, you know, all the hardware that we need. And far from helping us, that $1 trillion investment is actually going to make this even worse. We're going to move from a place where we have just GPU constraints to the everything constraint, from CPU to memory to disk, and we're already seeing that today, and it's only going to increase as we go further. And to add to that, that means we're going to have to keep hardware in our fleet longer, which means we're going to have increased failure rates.
22:12 We're going to have to take capacity from wherever we can get it. We're going to have increased heterogeneity, which means more complexity and more problems to solve. Pretty interesting stuff, and I'm super excited about it. Now, in closing, and on a bit more of a more personal note, I've been in this industry for something like almost 30 years at this point. About half of that time has been at Meta. From the early dot-com boom to early internet to the advent of mobile and streaming video, both on your phone and on your desktop and long-form and short-form, a lot of things have changed dramatically, but never, never in that 30 years have we moved more faster than we are right now.
22:47 And the crazy part is we're actually doing a pretty darn good job of this. Things are actually generally to to work. And I was trying to understand what is it that makes that happen. And this is where it comes back to those moon missions, because I think there's a lot in common. The first is we have control of our destiny. The second is we allow people to move very fast and reduce the amount of red tape that gets in the way. And the third one is the people themselves.
23:09 And to go in a bit more detail of what I mean by that, on the first topic, it's about giving people the autonomy and the speed and reducing the red tape to allow them to move really fast. On the second topic, it's the ability to give them the the products and capabilities to really move things quickly and go back to the point where you say, "Look, the only thing that is up for debate is something that is not bounded by the laws of physics." And if you can get to that point, we can move very quickly.
23:36 And finally and most importantly, like I said, it's the people. It's all the people in this room. It's all of you collectively coming together and bringing your experiences and your talent together to move the needle forward. It's these ingredients and these investments and the technology that's going to come together to allow us to make more meaningful and valuable products in the world, to increase the capabilities of people around the world. And that's why it's such an exciting time to be right where we are today in this industry. And that's why it's so exciting to work on systems and reliability. Cuz this is right at the heart of some of the hardest problems in the largest infrastructure build-out in human history.
24:11 With that, thank you.
Summary
- $338 billion was spent on the Apollo missions; $450 billion on AI buildout in 2025, with projections of $765 billion for infrastructure in 2026.
- AI development has evolved from basic code completion to advanced agentic-driven development, significantly enhancing productivity.
- Meta's infrastructure has seen improvements in reliability and efficiency, with a focus on fault tolerance and rapid deployment of resources.
- The introduction of AI in software development has led to a 50% reduction in incidents per change, despite an increase in the volume of changes.
- The complexity of managing large-scale systems has increased, necessitating better diagnostics and fault detection.
- The speaker draws a parallel between current technological advancements and the early stages of the space race, indicating that the journey is just beginning.
- Future challenges include hardware constraints and the need for improved operational systems to handle increasing complexity.
- The importance of collaboration and innovation among teams is highlighted as essential for driving progress in technology and infrastructure.
Questions Answered
What do the investments in the Apollo missions and AI buildout signify?
The $338 billion spent on the Apollo missions and the projected $765 billion for AI infrastructure in 2026 highlight the scale of investment in technology and innovation. This comparison illustrates the potential for significant advancements through substantial financial commitment.
How has engineering improved the reliability of systems over the past year?
Real-world operational data has shown that the reliability of systems is better than expected due to rigorous engineering practices. This includes optimizing components based on failure data and enhancing redundancy.
What are the implications of scaling AI training jobs on system reliability?
As AI training jobs scale up, the likelihood of individual failures increases, which can significantly disrupt operations. The data indicates that larger GPU clusters experience more frequent failures, necessitating robust system design to handle these challenges.
How has the introduction of AI affected software safety and incident rates?
The integration of AI has led to a significant reduction in incidents per code change, despite an increase in the volume of changes. This suggests that AI can enhance safety in software development.
What does the future of investment in technology look like?
The projected $1 trillion investment in technology by 2027 indicates a significant escalation in funding for innovation, comparable to the entire U.S. military budget. This suggests we are at the beginning of a transformative era in technology.