transcribe

AMD Advancing AI Keynote

AMD · 2h 6m · transcribed Aug 2026
More from AMD Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:56 Trust has always been at the heart of progress. It's how ideas take flight, from test to triumph. But trust also asks us to believe in one another, and that's why trust has to be earned. It's earned by relentlessly working together to solve the most important challenges, and drive results.

1:28 That's why, as we advance into the AI future, trust has to lead the way. So yes, we've built a roadmap we deliver on, on time. Yes, we build for the open ecosystem. And yes, we've built the broadest AI portfolio with CPUs, GPUs and FPGAs. But of all the things we build, trust is the most important.

2:04 Please welcome to the stage, Dr. Lisa Su, chair and CEO. Good morning. How's everyone doing? Woo! Woo! It is great to be back here in Silicon Valley with so many friends, press, analysts, partners, and especially all of the developers who are here today.

2:38 And a big welcome to everyone who's joining online from around the world for our Advancing AI 2025. Now, it's been an incredibly busy nine months since our last Advancing AI event. We launched lots of new AI data center PC and gaming products, but today, we have so much exciting news to share with you. I'd like to go ahead and get started. Now, you guys know us well. At AMD, we're really focused on pushing the boundaries of high performance and adaptive computing to help solve some of the world's most important challenges. And frankly, computing has never been more important in the world.

3:17 I'm always incredibly proud to say that billions of people use AMD technology every day. Whether you're talking about services like Microsoft Office 365 or Facebook or Zoom or Netflix or Uber or Salesforce or SAP, and many more, you're running on AMD infrastructure. And in AI, the biggest cloud and AI companies are using Instinct to power their latest models and new production workloads, and there's a ton of new innovation that's going on with the new AI startups.

3:52 For example, life sciences company 310.AI uses MI300X to train a model that turns simple text prompts into novel proteins to really help accelerate drug discovery. Our versatile AIadaptive SOCs are being used to build more efficient 5G networks and improve automotive driver safety. And Ryzen is bringing AI to PCs, enabling more intuitive, responsive, and more powerful experiences.

4:22 Now, since ChatGPT launched a few years ago, the pace of AI innovation has been unlike anything I've seen in my career. And in 2025, it's only gone faster. We've seen the emergence of more powerful reasoning models, the rise of agents, really growing momentum in real world use cases that are actually starting massive scale deployments, and it's clear that we're entering the next chapter of AI. Now, training is always going to be the foundation to develop the models, but what has really changed is the demand for inference has grown significantly, driven by more capable models and new use cases that are increasing AI usage.

5:07 We're also seeing an explosion of models. So o of course you have, you know, the new frontier models from folks like OpenAI and Google, but you also have open models from Meta and DeepSeek and many others. And we're also seeing now a surge in new specialized models that are built from everything from healthcare to finance to coding to scientific research. And when you look over the next few years, one of the things that we see is we expect hundreds of thousands and eventually millions of purposebuilt models, each tuned for specific tasks, industry, or use cases.

5:44 And as AI does more complex tasks, like reasoning, you expect agents to become more capable. It drives significantly more compute, which frankly is great for all of us. Now let me talk a little bit about agentic AI. You know, agentic AI actually represents a new class of user. One thing that is always on, constantly accessing data, looking at applications, looking at systems to really make decisions and really work autonomously. They need high performance GPUs to generate insights in real time, but that's really only part of the story.

6:20 What we're seeing now is as agentic AI activity increases, all of those agents are now also spawning a lot of traditional compute going to high performance CPUs. And just think about it, what we're actually seeing is we're adding the equivalent of billions of new virtual users to the global compute infrastructure.All of these agents are here to help us, and that requires lots of GPUs and lots of CPUs working together in an open ecosystem.

6:52 Okay, so let's talk a little bit about the market. when we were here last year, we said that we expected the data center AI accelerator TAM to grow more than 60% annually to $500 billion in 2028. And frankly, for many of the analysts and folks, you know, at the time, that seemed like a really big number. People were like, "Do you really think it can be that big, Lisa?" And I said, "Well, you know, that's what we're seeing." and what I can tell you, based on everything that we see today, that number is going to be even higher, exceeding 500 billion in 2028. And most importantly, we always believed that inference was actually be the driver of AI going forward and we can now see that inference inflection point. With all the new use cases and reasoning models, we now expect that inference is gonna grow more than 80% a year for the next few years, really becoming the largest driver of AI compute.

7:56 And we expect that high performance GPUs are gonna be the vast majority of that market because they provide the flexibility and programmability that you need as models are continuing to evolve and really algorithms are moving so fast, you want that programmability in your compute infrastructure. Now, the other thing that we see is AI is also moving beyond the data center, from intelligence systems at the edge to PC experiences, and we expect to see AI deployed in every single device.

8:29 Now, to enable all of this, you don't have any one architecture that is the right answer. So I like to say there's really no one size that fits all. What you need is the right compute for each use case and that's exactly what we're focused on. Our strategy is really focused on three key principles. First, we're delivering a broad portfolio of compute engines so customers can match the right compute to the right model and the right use case.

9:00 Second, we're investing heavily in an open developer first ecosystem, and you're gonna hear us talk about open a lot today. We're really supporting every major framework, every library, every model to bring the industry together in open standards so that everyone can contribute to AI innovation. And third, we're delivering full stack solutions. We're building, we're forging partnerships. You're gonna hear from some of our partners about our ecosystem today to really put all of these elements together.

9:31 So now let me just give you a little bit of color. From a portfolio standpoint, we offer the most complete suite of computing elements end to end for this vision. That includes CPUs, GPUs, DPUs, NICs, FPGAs, and adaptive SOCs. No matter where AI runs or how much compute you need, AMD has the right solution. Next, let's talk about open. There are a lot of developers in this audience and online, so this is really talking to you. Thank you for being here. Thank you for coming today.

10:07 And we believe an open ecosystem is actually essential to the future of AI. AMD is the only company committed to openness across hardware, software, and solutions. And when you just take a look back, some of the most important breakthroughs in tech actually started out closed if you think about things like, early networking protocols, Unix operating systems, and even mobile platforms. But the history of our industry shows us that time and time again, innovation truly takes off when things are open.

10:46 Linux surpassed Unix as a data center operating system of choice when global collaboration was unlocked. Android's open platform helped scale mobile computing to billions of users. And in each case, openness delivered more competition, faster innovation, and eventually a better outcome for users. And that's why for us at AMD, and frankly for us as an industry, openness shouldn't be just a buzzword. It's actually critical to how we accelerate the scale, adoption, and impact of AI over the coming years.

11:22 Now, we also recognize that these AI systems are getting super complicated and full stack solutions are really critical. So to deliver full stack AI solutions, we've significantly expanded our investments over the last few years, both organically and through strategic acquisitions and investments. We're very happy to say that we recently recently closed our acquisition of ZT, giving us new capabilities in rack and data center scale design that are becoming extremely useful for what we're doing next.

11:52 And we've also strengthened our software stack, acquiring leaders like nod.AI, Mipsology, Silo.ai, and in the last several weeks, we announced adding the Breum and the Lamini teams to AMD. And we're also investing broadly in the AI ecosystem. Over the last few year over the last year, we've actually done more than 25 strategic investments that have been a great way for us to build new relationships and also support the AI software and hardware leaders of tomorrow.

12:26 So let's talk a little bit about customers. We have tremendous momentum in the data center. Since launching in 2017, EPYC has transformed the data center. Today, EPYC is trusted by the world's largest cloud providers and businesses to run their most important workloads.EPYC powers everything from hyperscale services to enterprise data centers, supporting the most important workloads with leaders in financial services, healthcare, media, and manufacturing.

12:59 And our momentum is just accelerating. We exited the last quarter with a record 40% market share, and we believe with AI and high performance compute, there's a lot more room for us to grow. In AI, MI250X and MI300A enabled the exascale supercomputing era. I'm very happy to say actually this week, there was a new top 500 list that was released, and AMD powers the two fastest, supercomputers in the world. So that's pretty cool.

13:32 And... Thank you. And with MI300X and 325, we've extended that leadership to GenAI with large scale internal and cloud deployments at Microsoft, Meta, Oracle, and many others. And I'm happy to say we've added a lot of new Instinct customers in the last nine months. Today, seven of the top 10 model builders and AI companies are using Instinct in their data centers. Leaders like OpenAI, Meta, xAI, and Tesla.

14:10 Innovators like Cohere, Luma, and Essential, and many, many more. You're gonna hear from several of them. They're our guests here today, and they'll tell you a little bit about how we work together. Now, as powerful as our hardware is, it's truly the software that enables their full potential. And I hear from lots of you as developers on what we can do better in software. I can say that I hear you and our ROCm software stack continues to make just incredible progress.

14:39 We're really focused on broadening the coverage for AI models, accelerating the pace of our releases, and really setting a north star of a developer first mentality with ROCm. When you hear me talk to our engineers, what I say is, it is all about the developer experience. It's all about what you guys say, and this is our guiding principle. So you're gonna hear a lot about that from Vamsi today. And then for those of you, who are going to be able to stay with us this afternoon, we have, a ton of developer content to just show you how you can really use, AMD and ROCm.

15:15 Now, to give you some perspective about what it's like to use AMD, I'd like to bring out my first guest. One of the newest partners who is running Instinct in their production environment is xAI, and here to share more, please welcome Xiao Sun. Hello, Xiao. Hi, Lisa. How are you? I'm good. How are you? We have a decent audience today. What do you think? that's a great audience. Xiao, we are super excited about the work that we are doing with xAI. You know, you guys are really at the forefront of developing stateoftheart AI models. you're going super fast.

15:57 can you share a little bit about what your team does and, you know, how are you managing all this? Sure, sure. Yeah, at, xAI, we have a very small team, and then we're moving very fast, and, we're following first principle. Basically, you know, we are advanced like a Grok fa family models. And then for, maximum truthseeking. And, to have that, we actually, you know, basically need to go with the first principle thinking, which is like we always challenge, like status quo. And we also always ask a question that why do things, have to be done like this, and could we do it better.

16:34 And yeah, we also apply that into like our computer infrastructure, which is very important for us. Yeah, absolutely. Look, we've been part of some of that, first principle thinking and, and how you, you know, really are focused on speed. look, we're super thrilled of the work that we've done together, on, MI300X at xIA AI. I asked you guys to give us a shot. You know, can you talk about how you're leveraging the MI300 infrastructure? Like, how has it worked? How, how did it come up for you?

17:03 Yeah, yeah, yeah. so if I use one word, right, that word is basically effortless. So as I mentioned Can you say that word again? Yeah, indeed, it's effortless, to use Yeah. To use the MD GPU in your product. so yeah, so as I mentioned, we are a very small team moving very fast. So for us, right, the most valuable resource is engineering time, right? So the opportunity cost is a mess, right? So, with your team's help, we, we basically can, you know, do not need to like spend too much time. Basically just a few of us engineers and your team, we successfully pushed one very important product, Grok family model into product. And I remember, you know, when we first start to, collaborate together, right, I look at, oh, there's a meeting on Friday. I say, "Is this very important? Can you just... can we just meet now?" Right? So after that, your engineers, adapted to our pace. So I always get like a phone call at like 9:00 PM or midnight, you know, and then my partner was like asking me like, "Who is calling?" I was like, "Oh,? Calling to ask about some question kernel." And he's like a violinist, so he's like, "Oh, that almost never happen in orchestra. And then, and what is a kernel?" Yeah.

18:19 so because of that, we collaborate very closely, and then we can actually, you know, in few months we can push something, you know, into product. That's really impressive. Well, I I've been impressed because, you know, our engineers are always, reporting to me, you know, where are we on the Grok model performance? And you guys have moved super, super fast. So, Xiao, you know, the other thing is, we're talking a lot about open ecosystems, and I know that you guys are a strong believer in open ecosystems. Can you talk a little bit about how ROCm and all of those community efforts have actually helped you? Sure, sure. So as you know, right, our infer, inference structure is based on SGLAM, which is like a open source, very popular open source platform.And, also the major contributors and also, are also in, xAI. So while they are advancing, you know, most, optimized, inference system, they are also contribute a lot to the open source community. We upstream, you know, our innovations, to SG Lang public ripple. And at the same time, we also benefit a lot from the open source community, right? They make... They find bugs, they fix the codes, and then we merge that into our, like, production, infra. And that really helps a lot. I think it's very essential and we'll continue to com commit to, you know, contribute and work together with the open source community.

19:40 Yeah. No, that's great. I think, the SG Lang progress has been just a great example of how fast, things go. So look, I know, you guys are always moving ahead. And I have, you know, a lot of products to talk about, with this audience today. Can you share a bit about your perspective of our collaboration and, like, what are you excited about? You know, what do you think about MI350 series and and just all the work we're doing together?

20:02 Sure. I'm actually very impressed by your, you know, yearly cadence about the new hardware. so, thinking about the future, right, I think we will continue to go back to first principles thinking. I mean, you are a pioneer also, in semiconductors. So you know that, what essentially we're doing is basically like a a fancy waterworks here, except that it's not a water molecules. We are basically manipulating, like, electrons, right? Like, we're pumping electron in very high energy level, and then we guide it through the, you know, the channel of transistor to the gate of transistor, and then dissipate it to, you know, the ground, right? This is how we do compute. But, you know, I think that this is not the end of it. This is the start of it. There are a way to do it like, probably 1,000 or if not one million times more efficiently. And also on our side, right, one way of thinking about the computer is basically compression of data, right? The data is like all the the text a human has ever written. and I think now and in future will be like all the realities, all the truth in the world.

21:02 And probably even further future, there will be like all the state of affairs that has not yet happened, but it could happen, right? So we compress them all and then put them into like, you know, your USB disk or something like, you know. And then when you need to use it, you retrieve it and decompress it, right? This is how we think about it. But both sides have many innovation to do, but, we cannot, like, do it separately. So this is basically, from my point of view, from silicon to product, this is like the largest, you know, codesign of human history. And then, you know, at xAI, we are very happy to collaborate with vendors and AMD, right, to, you know, do this large codesign together, accelerate the iteration. And I hope that, you know, all the talents from the world should join on both sides. That's fantastic. Xiao, thank you so much for joining us today. Thank you for your partnership, with, with us on Yes... MI300. And we look forward to doing a lot more together. Thank you, Lisa.

21:55 Thank you, Xiao. Well, look, we have a full lineup today of new announcements across hardware, software, and solutions. So let's go ahead and jump right in. Now, since launching MI300 less than two years ago, we're on an annual cadence of new Instinct accelerators. With the MI350 series, we're delivering the largest generational performance leap in the history of Instinct. And we're already deep in development of MI400 for 2026 that is really designed from the grounds up as a rack level solution.

22:32 So today, I'm super excited to launch the MI350 series, our most advanced AI platform ever that delivers leadership performance across the most demanding models. This series, you'll hear us talk about the MI355 and the MI350. They're actually the same silicon, but MI355 supports higher thermals and power envelopes so that we can even deliver more realworld performance. And thank you, Drew. My favorite part.

23:03 Here is MI355. This is our flagship product, and I'm showing this to you. It's powered by our latest 4th Gen Instinct architecture. It supports new data formats like FP4. It uses the latest HBM3E memory, and it has 185 billion transistors across 10 chiplets all integrated with our Leadership 3D packaging. So what do you guys think?

23:36 All right. Thank you, Drew. Look, the MI350 series delivers just a massive 4x generational leap in AI compute to accelerate both training and inference. With an industryleading 288 gigabytes of memory, we can now run models up to 520 billion parameters on a single GPU. The MI350 series also uses the same industrystandard UBB8 platform as MI300 and MI325. This is actually really important because it actually makes it super easy to deploy MI350 series into existing data center infrastructure.

24:16 Now, if you look at the specs compared to the competition, 355 supports 1.6x more memory and delivers higher FLOPS across a wide range of AI data types. And especially if you look at FP6 and FP64, we're double the throughput. Now, what does that mean? That means that you have leadership performance at both ends of the spectrum, whether you're talking about leading edge AI models or large-scale scientific simulation or engineering applications.

24:49 Now, at the platform level, an MI355 X server has massive memory capacity and compute relative to the competition. We're talking about 161 petaflops of FP4 compute.And 2.3 terabytes of HBM3E memory. And we have it in both air cooled and liquid cooled configs, giving customers the flexibility to meet their specific thermal, power, and density needs. Now let's look at some of the performance.

25:20 We set an ambitious goal with MI350 series to deliver a 35x generational increase in AI performance, and today I'm proud to say that we've delivered that. On Llama 3.1, MI355 delivers 35x higher throughput when running at ultra low latencies, which is required for some realtime applications like code completion, simultaneous translation, and transcription. We also deliver significantly higher performance across a wide range of AI applications, things like chatbots, or content generation, or summarization, or conversational AI. We can see performance up to 4.2x higher gen-on-gen.

26:05 And now when you look across a wide range of models, we see great performance as well. In DeepSeek and Llama 4 Maverick, we're seeing things like triple the tokens per second gen-on-gen. So that level of performance drives faster responses and the ability to serve more users with much, much greater efficiency. Now let's take a look at the competitive performance. When running DeepSeek R1 or Llama 3.1, MI355 delivers leadership throughput using open source frameworks like SGLang and vLLM.

26:41 We're generating up to 30% more tokens per second compared to B200, and actually matching the performance of the significantly more expensive and complex GB200, even when the competition is using their latest proprietary software stack. This is actually pretty cool because it tells you a co a couple of things, right? It first says that we have really strong hardware, which, we always knew, but it also shows that the open software frameworks have made tremendous progress, to the point where they are outperforming a closed vendor specific ecosystem.

27:16 And when you combine all of that performance with lower CapEx, what we're seeing is MI355 can deliver up to 40% more tokens per dollar than competing solutions. 40% more tokens per dollar. That means higher throughput, greater efficiency, much better TCO for cloud and enterprise, and it really makes MI355 the best choice in the industry for inference at scale.

27:51 Now, one of our earliest partners to deploy Instinct broadly was Meta. To share more on our work together, please welcome Meta VP of Engineering, Yee Jiun Song, to the stage. YJ. Thank you for having me, Lisa. Hey, thank you so much for being here. We're so excited about the work that we're doing together. you know, look, we have so much respect. Meta has been an incredible leader in AI across, infrastructure, models, and services. I get to talk to YJ a lot. He gives us good feedback, good feedback. you're delivering, all of this capability at at amazing scale, and so it's it's been our privilege to be your partner across, EPYC and now Instinct. Can you talk a little bit about, our collaboration? Well, first, thank you for having me here, Lisa. I'm thrilled to be here and excited to see all the progress that you and the AMD team are making. we're seeing incredible advancements that started with your EPYC products and now extending to all of your AI offerings. I think AMD and Meta has always been strongly aligned on vision, roadmap execution, so this means that we have very close coengineering, performance tuning. We troubleshoot problems together, and we're able to deploy, optimized systems at scale. So our teams really see AMD as a strategic and responsive partner, and someone that we really rely on.

29:13 Well, we love working with your engineering team, and you know that. we love the feedback, and you know, you also hold a high standard. you were one of our earliest partners in AI. you know, as your demands continue to scale, you know, you've talked about your usage of MI300. just what are you doing today and and what are your plans with, MI350 going forward? Yeah. So AI has been core to many of the user experiences across all of Meta's products for a long time, across Facebook, Instagram, WhatsApp. And now, of course, with Llama and Meta AI, AI has become even more important than before. It's been fantastic to see our collaboration on MI300X come to fruition. So MI300X accelerators today are a key part of our infrastructure. We've deployed this quite broadly for Llama 3 and Llama 4 inference due to its high performance and excellent performance per TCO. As we gain experience with MI300X, we're also expanding the workloads that we run on them. So we're now t today using MI300X both for training and inference of the different ranking and recommendation workloads which are critical to our business. We're also quite excited about the capabilities of MI350X. We like that it brings significantly more compute power, next generation memory, and support for FP4, FP6, all while maintaining the same form factor as MI300 so we can deploy quickly.

30:30 Thank you, YJ. Thank you, by the way, for your trust. I know that, you know, to to deploy us on more workloads requires, a effort on both sides, so we really appreciate that. you know, Meta has really been a leader developing, frontier models, if, you know, think about, you know, AI across all of your, applications. tell us a little bit about what you're seeing. You're like at the front there. What are you seeing in applications? What are you seeing in that, your compute investments? Yeah. So I think the AI work at Meta that gets the most attention is probably the Llama models, though we are committed to developing frontier models with our Llama efforts. But that's really just the tip of the iceberg. We're seeing growth of AI workloads across all of our different products. AI is not only improving our existing products, but also allowing us to develop entirely new products. All of this is driving investments in computer infrastructure at a scale that's quite unprecedented. We're building data centers and filling them up faster than I've ever seen. Our entire infrastructure team is sprinting to ensure we have the capacity to build the next great AI models and then take advantage of those models to deliver value to our users.

31:35 The result of this incredible capacity buildup is that we really care about per pertissio and making sure that we get the best bang for the buck for our investments. Well, we've been part of a few of those sprints. Yeah. Just a few. So, Waijie, you know, we've also worked very closely on the software side, Yeah ... you know, PyTorch, ROCm, the open hardware environment. Like, how do you think about open in your strategy? Yeah. So I I think our collaboration has spanned both software and hardware for many years at this point. I think most recently, since 2021, we've partnered closely to enable ROCm through PyTorch to ensure that developers can leverage AMD GPUs with PyTorch's, ease of use right out of the box. Thank you.

32:14 beginning last year, we've worked closely on improving ROCm's communication library, RCAL, which is critical for AI training. We ve we really appreciate that partnership, by the way. So... o o us too. Beyond PyTorch, Meta contributes heavily to the open source, community. An example here is to work on compiler frameworks, such as Triton, which allows us to write code once and then write o run on different accelerator families. Now, of course, we also collaborate on optimizing Llama models to run well on AMD GPUs. operating at scale also require that our accelerators, the accelerators that we buy, be compatible with our network and data center infrastructure.

32:51 Here, we rely on our common infrastructure hardware racks to be the integration point between our accelerators, the network, and the data centers themselves. This is one of the reasons why we've been able to introduce MI300 into our production environment so quickly. Yes. No, absolutely. I think the, the OCP work is fantastic. look, you know, we're we're super excited about our partnership, everything that we're doing together, and it feels like we're just starting the ramp of MI350. Yeah.

33:17 But in our business, we're always talking about the future. Absolutely. You guys are always asking us what's next. Absolutely. I'm actually gonna preview MI400X a little bit, later in the show. So can you just talk about where you see AI going in the future and, you know, how does that shape what you need from, you know, partners like us? So, so AI is driving massive growth and infrastructure demand, but actually, it's not just about the size of the demand or the amount of capacity. The type of workload is also changing very rapidly. as an example, not so long ago, as an industry, we were very focused on pretraining.

33:49 Yes. But towards the end of last year, we start to see the emergence of test time inference and reinforcement learning and other new workloads that demand huge amounts of computation. We also started to see the rise of mixture of expert models that placed high demands in network interconnection speeds and the performance of, the collective communication primitives we just talked about. And beyond generative AI, we are also finding that our recommendation systems are also getting more complex. The direct implication of these, rapid changes is that Meta and AMD will have to work more even more closely together to define our accelerator and network roadmaps for the future. I can't wait to hear what you're about to share.

34:26 Well, you know everything I'm about to share. but but I will say that I I do remember sitting in your office and asking you, "So YJ tell me what's gonna happen when the work loads." And you're like, "Well, you know, be flexible." That's absolutely true. Thank you, YJ. Thank you. It's it's, we really, really appreciate the partnership. Yeah. it's wonderful working with you and the entire team, and, and and thank you for, you know, all that we're doing together.

34:50 Thank you. All right. So now let's turn to training. in addition to all the work we've done to improve inference, we've also made a lot of optimizations for training, and I'm happy to say we've seen some fantastic results. MI355 delivers significantly better training performance than, MI300. And when you look at, pretraining, for example, where foundational models are built from the grounds up, 355 is delivering up to 3.5X higher throughput across a range of models and data formats. And in fine tuning, we're delivering up to 2.9X more performance gen-on-gen, which enables just faster iteration cycles and reduces the time from model development to model deployment. Now, comparing to the competition, MI355 pretraining performance is actually on par with B200 across a range of model sizes and data formats. And I actually think that's very good considering, how new MI355 is. Now, in fine tuning, we just saw some of the latest MLPerf benchmarks that were out there, and MLPerf is largely considered kind of the gold standard for training benchmarks. We see that MI355X actually outperforms both B200 and GB200 when we're talking about completing the benchmark up to 13% faster compared to the latest published results. So that just tells you how much progress we've made in training.

36:28 And now, as we talk about solutions, we said we want to make this super easy to use. So with the MI350 series, OEMs and ODMs are launching racks built entirely on AMD technology for the first time. We're combining fifth gen EPYC CPUs, Instinct MI350 GPUs, and our Pensando NICs in an integrated solution. And these are all OCP compliant designs, so they drop right into existing infrastructure. And in some of the densest environments, we have liquid cooled racks that can scale up to 96 or 128 GPUs that deliver up to 2.6 exaflops of FP4 compute and 36 terabytes of HBM3e memory. And on the enterprise side, we can do aircooled systems that support up to 64 GPUs and integrate all of that into an existing infrastructure. So this is the kind of flexibility and range that our customers really want. What they want is to be able to take the technology, get it into production, get it into the data centers as quick as possible with as little work, as little, disruption. And that's exactly what we can do with MI350.Now, one of our most strategic cloud partners that's building with AMD across the stack is Oracle. To share more about our work together, please welcome Mahesh Thiagarajan, executive vice president at OCI.

37:55 Hello, Mahesh. Nice nice to meet you, Lisa. Hey, it's wonderful. Thank you for being here. we so appreciate the partnership with OCI. You guys have been with us, across the board. That's right. you know, you guys are frankly at the center of AI compute, tremendous momentum. Doing our best. Powering some of the largest training and inference clusters. So, look, as you look through what's facing you in all of these deployments, you know, what's most important to you?

38:20 Look, I I I'm truly honored, first, to be actually working so closely with AMD and building these AI infrastructure, at relentless pace together, right? Fundamentally, to solve the next frontier of challenges at the intersection of cloud and AI, we need to do a deep integration across the entire stack, starting from power, to compute, to network, to storage, to truly use every last ounce of the performance that's available. So, let me break that down a little bit. So, when we talk to customers about compute, what we see is that the most demanding training and inferencing workloads need the exceptional reading of the CPU and GPU memory. This is where I think the AMD Infinity Fabric actually truly enables offering the performance between moving the data sets really clo you know, close to the the acce the accelerators at AI speeds, and we're seeing customers seeing tremendous value. The second thing, which I think is is super important, I'm very passionate about what AMD is doing here, is really around the high performance networking that comes close to delivering these large training clusters, right? And and fundamentally, when you think about an AI supercluster, it is about... It's operating as one giant supercomputer, really looking for that ultralow latency, extreme high performance bandwidth, and truly operating as one supercomputer. And so, that performance across these nodes really matter to complete that task. And what we partner with and we work with AMD a lot on is on the networking technologies. And for example, the Pensando work that we've been doing for a while truly enables the security of these AI workloads and actually is powering the performance that we're seeing. Yep. Yeah. No, that's fantastic. I mean, I I I think... Thank you. We, we love that vision, overall of of putting all these pieces together. By the way, that's exactly our philosophy as well That's right ... that you need CPUs, GPUs, networking coming together. now, OCI was one of our, early adopters of MI300X. you know, it's been great to see some of the customer response. Can you talk a little bit about, you know, howamd instinct looks in Oracle Cloud? Look, MI300X on Oracle Cloud Infrastructure is a very deep integration. We've seen massive demand from both AI native companies and large frontier model companies actually doing work on top of OCI today. Now, the support model is actually like where we work partner well together to offer that fantastic experience, where when a customer comes looking for a cluster, in a moment of minutes, they're able to get their amd instinct machines, truly with what AMD brings to the table. The latest and the greatest ROCm innovation, the updates, everything's available out of the box, step one. Two, what we try to do is also ensure that customers are getting the latest performance benefits of everything that you guys are doing on ROCm very regularly on a monthly update. So, the customers actually get the best of AMD immediately. And third, is, the latest edition of your PyTorch and vLLM support. We've actually seen acceleration of some of our customers Thank you ... who've been waiting for that support. So, they're like, "Oh man, this is exciting, so let's go use that platform." And and so, you know, look, some of the largest names, largest customers, who you've all heard about are all running amd instinct on OCI, and they're having a great experience.

41:43 it's, wonderful, Mahesh. Look, I I really have to thank you and your team. I know that this has very much been about getting exactly what customers need and That's right... and you guys have been super, super agile in that. Now, you know, when you think about, Oracle and your adoption of technology, we talked about 300. you're actually now leading with us on 355. I know there's a lot of interest there. Can you talk a bit about just the evolution of our partnership and what you're seeing?

42:08 Look, I think our partnership probably started about a decade ago, right? I think we've been partnering for a long time Yes ... but I think it started heating up around a decade ago. And it started with us actually using AMD EPYC for our databases. Yes. Right? Our Or Oracle likes the database machines now, like back, you know, back since then, we've been using AMD EPYC. We've seen tremendous performance where we're seeing 3X higher transaction throughputs and 3.6X faster analytic squares. And the great news about that is that's actually available not only on OCI, but on premises, other cloud partners that are actually supporting Oracle's databases, including our latest Oracle Autonomous Database. So, it's everywhere. And then, you know, something that's very personal to me is is our Pensando partnership that truly enables a hardwarebased network virtualization in our entire cloud, which is a fundamental unique innovation that offers security and high perf for every customer that runs on OCI today, and that's with AMD.

43:03 and obviously, you know, we, 18 months ago, we did the AMD Instinct partnership. Lots of customers are growing. And, you know, we think there's tremendous demand, you know, project about a 10X growth over the next year, really trying to drive the AMD Instinct platform on OCI. And the most exciting part for me is that we're announcing our partnership on MI355, truly bringing the latest, out to the cloud with support for zettascale clusters. But I think in a couple By the way, did you guys hear that support for zettascale clusters?

43:35 So, I I'm I'm truly excited about that, because I think when I talked about that deep integrated compute network storage, we're gonna support AMD MI355 and more importantly...We're actually gonna go live in a couple of months with over 27,000 GPUs in a single cluster available on OCI in two months. That's fantastic. Thank you. Now, I'm asking everybody here, this is also about the future and, you know, Oracle has actually been leading on this, you know, concept of building giga scale, you know, data centers and, you know, we're talking about, what we have to do from a systems standpoint. Tell us what it means to build giga scale data centers and and how we can participate in that. Yeah. No, look, I'm I'm an infrastructure nerd so I'll talk about power or start. look, gigabyte scale, I think if you talked about this phrase two years ago everybody would've been like, "What are you talking about?" Right? It's insane. but I think the biggest challenge that we see is still power. Power is the r you know, fundamental bottleneck that still exists and the speed at which we can build them. but I think Oracle's partners come here where we've been a big, you know, partner with all of the utilities and other businesses and industry verticals so that actually has been very helpful. And second, we are making investments in sustainable energy, be it green energy around geothermal, wind, solar, and we're also looking at new small nuclear reactors that can power this. The second thing is about time to market.

44:59 One of the things that OCI pioneers in actually bringing our cloud infrastructure in, like, three racks. So we spend a ton of time tuning time to market. And second, bringing that price performance value to our customers. And for me, today we're operating at over 100 regions so we're able to reach a farreaching audience. But I think going back to the price performance message, I really think that's why AMD and Oracle work together. Yeah. We're we're bringing value to customers. That's that's what we do.

45:23 It's it's it's all about the price performance and today, you know, with, that price performance value, you know, customers actually enjoy the AMD plus Oracle partnership. And lastly, we're really excited and looking forward to your 450X platform. I'm seeing some of the specs and I think it's gonna be truly special. Fantastic. Mahesh, thank you so much. Thank you for the partnership overall. Any time. And I look forward to everything we're gonna do together. Absolutely. Thank you so much, Lisa. Thank you.

45:47 Appreciate it. Thank you. All right. So, look, customer excitement for MI350 is very strong based on the performance and, you know, cost per token advantages. I'm happy to announce that MI355 production shipments actually started earlier this month and we have the initial wave of partners on track to launch platforms and public cloud instances here in the third quarter. So really, really excited about MI350.

46:19 Now, another important focus for us is sovereign computing. Around the world, we're partnering with national governments and research institutes to help build the high performance computing and AI infrastructure that is really critical for their economies. And the goal really goes far beyond just building domestic compute capacity. It's really about using AI to power public services, research, and national programs that create societal impact. To get there, governments are actually prioritizing resilient infrastructure.

46:51 They want open standards, they want flexible architectures, and they want a diverse ecosystem of technology partners. Today, we have more than 40 active engagements globally powering critical public agencies, national computing centers, and sovereign AI activities. From the world's fastest super computers in the US to the rapid expansion of high performance computing across Europe, Asia, and the Middle East, to a wave of sovereign AI deployments around the world. This is a growing part of the market and we are increasingly spending more time helping nations build their computing strategy and infrastructure.

47:28 One of the best examples of our progress is in Europe with our Silo AI team. Silo is our A AMD AI lab, but they're also a solutions factory collaborating closely with governments, industries, and research institutions to develop models and applications aligned with national priorities and optimized on our hardware. Silo is working across Europe collaborating with companies like Allianz, Nokia, Philips, and Unilever, advancing open multilingual LLMs with the European Commission and pushing frontier model research on the AMD powered LUMI supercomputer.

48:05 They're also playing a very important leadership role in the open source AI community contributing to models and partnering with leading AI innovators like Aleph Alpha, Mistral, and NXAI. Now another extremely exciting example of our sovereign efforts is our work with Humain, a new company with an ambitious vision to build advanced, locally developed AI in the Middle East. To share more about our work, please welcome Tareq Amin, CEO of Humain.

48:37 Good morning, Lisa. Hello, Tareq. Good morning. Good morning, everyone. It is great to have you here. Thank you so much for joining us. we're so excited about our partnership, together. Well, first of all, thank you very much for inviting me here. I don't need to tell you this but AMD is really an important partner for Humain, important partner for Saudi Arabia, and also important partner for the entire larger ecosystem of AI companies. Well, look, you guys are on an exciting mission. I had the, the pleasure and honor to be with you in the kingdom just last month and the vision that you're laying out, launching Humain, you know, it's such an important moment, for, you know, the kingdom and and just taking sovereign AI to the next level. So, can you share with us. Just tell me about our vision, your plans, everything. So so Lisa gave me four minutes. By the way, this is the biggest challenge You can take four.

49:26 This is the biggest challenge I have. But, I I wish I could share with you what we have done, last month. I'll take a perspective just to tell you the entire story and the partnership that we're doing with AMD to redefine the entire AI infrastructure ecosystem. in the US, I had the opportunity to build digital infrastructure across 22 cities. I moved to India where I learned how to scale things.I moved to Tokyo where I built technology that was in research paper, realize it, and the story gets completed with Humain.

49:59 Humain and Saudi Arabia came together through the consolidation of enterprise, e various enterprises in the country, and also one government entity that was developing large language models. Our obsession is about disruption via technology. The way we pick partners is not based on what I call transactional selections. When we met Lisa and her team, we really hit it off, because we both agreed that we're gonna co-own the outcomes. It was very, very important that co-owning the outcomes and having a skin in the game, taking a risk to build something that is good for humanity was a very, very important mission. So today, you know, in front of all of you, though we've talked about this, the announcement about the joint venture with AMD, I am really thankful for your support for what we need to do. But I wanna give you a glimpses of what this really mean.

50:57 We are committed, and when we looked at the advantage of what Saudi Arabia could really do, by 2030, the deficit in power is estimated to be around 100 gigawatt. No matter what you do, you will still need power to build the capacity that we need, for AI. This is an added advantage that we thought we could really help and we could participate into this AI global ecosystem. We have an abundance of land, an abundance of power, a mixture of renewable as well as traditional energy, and a really very young society that is hungry to learn. So we thought this could be great. our commitment for all AI developers and AI companies, what if, if we reduce your cost of ownership by 30%?

51:46 From whatever you could achieve as the lowest worldwide cost, I'm committing to make that together with Lisa the lowest cost That, that sounds like a good commitment. So, so we're really, really happy. This is a, a game changing moment. We're really privileged that this joint venture is gonna be a game changer. 2030, 1.9 gigawatt, 2034, six gigawatts. It starts in Riyadh, but it doesn't stop there. We will go and look at other global opportunities to build our infrastructure.

52:17 You know, Tareq, I want to just point out some of the things that you said, right? We've talked about, the need for power, we've talked about, you know, the need for speed, and we've talked about the need for efficiency in what we're doing. I think the, you know, what impresses me the most about the work that we're doing together is, you you really have like a clean sheet of paper to talk about what's next. So we've talked about a lot of AI infrastructure both in the kingdom and outside of the kingdom. Can you just talk about how, you know, some of the milestones that we have in place? So, so I think, as soon as we really crafted this agreement, I mean, the timing of Lunch Humain was not also coincidental.

52:54 you know, we were really happy that it was coincided with a presidential visit into Saudi Arabia to talk about relationship and partnership they're doing with technology companies such as AMD. We have already started the construction of two large campuses, 11 data centers, each one of them of 200 megawatt capacity each. I will tell you, Lisa, almost on a weekly basis, "Tareq, we need to move faster, we need to move faster." So I really appreciate your spirit.

53:20 I, I, I have some MI350s for you that need data center. No. So, so, so we're, by, by this year, I mean, our entire build is to get our first 50 megawatt done, and then we start scaling up on 50 megawatt modules every quarter. So my entire obsession now is about the infrastructure layer. one thing that I think all of you saw when, Lisa was talking about the new generation, I mean, I would tell you, congratulation on MI350.

53:50 I could not even be more excited about what 2026, I think the MI400 series is a game changer for our industry. But realize that what we are doing with AMD is not just on buying chips. Lisa and her team have enabled us to really disrupt the TCO. Second is about openness. We talked about this. We said we need an inclusivity. The world working together is a better place than us being fragmented. And the idea that we build an open ecosystem, inviting many other to participate, including AMD, including Cisco and many other financial partners that are gonna come and take this hopefully as a blueprint of what we need to do to address the gap that the world have in energy. That's fantastic. Tareq, thank you again for the incredible partnership. we are super excited about what we're doing together. I think we're super excited about what we're gonna do for this AI ecosystem going forward. And, Thank you, Lisa. Thank you very much.

54:49 Yeah. Thank you. Thank you. Thank you. Really appreciate it. Thank you. All right, so you can see there's just a lot of excitement on, MI350 and our roadmap. Now, as exciting as the hardware innovation is, it is really the software that unlocks the full potential of AI. So to share more about everything that we're doing in ROCm and the developer ecosystem, please welcome AMD Senior Vice President of AI Vamsi Bopanna to the stage.

55:22 Thank you, Lisa. Thank you, Lisa. Good morning, everybody. AI innovation is advancing at an unprecedented pace, reshaping compute and redefining what's possible. Our vision for ROCm is simple, to create an open, scalable software platform that unlocks this AI innovation for everyone everywhere. And over the past year, we made tremendous progress realizing this vision. By partnering deeply with the open ecosystem, we are delivering a credible alternative that the industry can trust. ROCm is now powering AI platforms at scale, delivering some of the most demanding workloads on the planet. So today, I'm so excited to show you how far we've come and why this is just the beginning.

56:15 Now, last year around this time, we were super focused on delivering leadership inference performance to our largest customers. Since that time, we have significantly expanded our customer base, accelerated our inference capabilities, and now added training support across key models and frameworks. We have been relentlessly focused on what matters most, making it easy for developers to build with better out of the box capabilities, easy setup, more collateral, stepping up community engagements. We've been running hackathons, contests, meetups, and more.

56:50 And our customers are deploying AI capabilities at unprecedented pace, and that's why we've significantly accelerated our release cadence. New features and optimizations are now shipping every two weeks. Leading models like Llama and DeepSeek work on day zero. We've also responded to asks from the community for more industry benchmarks, starting with inference, and now for the first time just last week, we've demonstrated leadership training performance at MLPerf.

57:20 Our collaboration with the open source community is deeper than ever before. Over 1.8 million Hugging Face models now run out of the box on ROCm. PyTorch now has a performance CI in addition to functionality. We've added vLLM, SGLang CI pipelines on our latest hardware. A great example of our collaboration is the work we are doing with Triton. After achieving functional enablement last year, we've been laser focused on delivering performance in recent releases.

57:54 And now, in the last year, we've added significant support for JAX. With libraries like maxtext, we're seeing increasing adoption of JAX in our lead training engagements. Now, as we look ahead, the world of AI never sleeps. The pace of innovation is only accelerating at every layer of the stack from hardware to algorithms to models and applications. And all of this is happening at scale. Our customers continue to need feature velocity and performance gains to stay at the forefront of AI.

58:26 So today, I'm super proud to announce ROCm 7. ROCm 7 is bringing exciting new capabilities to address these emerging trends and brings support for our MI350 series of GPUs. It continues our relentless focus on usability, performance, introduces the latest algorithms, advanced features like distributed inference, support for large-scale training, and new capabilities that make it easy for enterprises to deploy AI effortlessly. Within ROCm 7, inference has been the largest area of focus. We've innovated and invested at every layer of the inference stack. From the latest framework enhancements in vLLM, SGLang, implementing serving optimizations, supporting advanced data types, to delivering extremely high performance kernels, to implementing the latest algorithms, like flashAttentionV3. We made it easy to author and integrate kernels with Pythonic abstractions, and we've done significant work in our communications stack.

59:34 This is how ROCm 7 delivers over 3.5 times the performance of ROCm 6. And when it comes to inference serving frameworks, it's becoming more and more clear that open source feature velocity and performance is in fact outpacing proprietary alternatives. Just look at what's happening in frameworks like vLLM and SGLang. They're actually setting the pace on commits and have both enabled FP8 optimizations and support ahead of closed alternatives.

60:05 Working closely with these open source communities, MI355 is today delivering up to 1.3x better throughput on DeepSeek FP8 when compared with B200. That's the power of open collaboration, moving fast and delivering more. One of our earliest partners that's innovating at scale with ROCm is Microsoft. So to talk about our work together, please join me in welcoming Eric Boyd, CVP, AI Platforms from Microsoft.

60:42 Eric, so good to see you. Thank you for joining us this morning. Yeah, really glad to be here. Yeah. You know, we've been very close partners for a long time. Can you tell us a little bit about how that partnership has evolved, and particularly around Instinct? Yeah, sure. I mean, as you know, we've been using several generations of Instinct. it's been a key part of our inferencing platform. and we've integrated ROCm into our inferencing stack, making it really easy for us to take and deploy new models on the platform.

61:13 That's great. Now, tell us a little bit about the type of models, what kind of work our teams are doing together. Yeah, so at Microsoft, you know, the customers that come to AI Foundry, or even our internal customers, are looking for the cuttingedge leading models, and so models like GPT4o or 41 from OpenAI. And, you know, the Instinct chip is, really gives us great performance on top of that platform, really enabling us to scale and perform at the, you know, tremendous scale and low latencies we need.

61:43 That's great to hear, and we've been super lucky to have collaborated with your team over the years. Tell us a little bit about the role AMD plays in enabling performance, efficiency, and what flexibility does it provide in your infrastructure? So when you're serving these large language models, one of the big challenges is...... taking advantage of all the memory on the chip, and so the models have tons of parameters and they have caches and things. And so the more memory you have available and the better bandwidth, the better performance you get and the better latency that you get out of it. And so the Instinct chip brings a, a large memory footprint along with really dense compute across it, and, that all combines to give us really great TCO benefits as we use these, these chips to serve our platform.

62:26 That is so great to hear because that's exactly how our engineers have been thinking about it when designing these features. They did a good job, yes. And, and, you know, you've expanded from, the original set of models now to actually working with more open models, so can you share a little bit more about the work there? Yeah, of course. At, at, at AI Foundry, we're committed to making sure customers get the most advanced models from, you know, OpenAI, Mistral, Cohere, other companies like that, but we have over 11,000 models in our catalog and most of those are open source. I think one of the interesting things over the last few months has been the emergence of DeepSeek, as an open source model that provides really great quality in it. And, we inference the DeepSeek model on, SGLang, which is an engine, that's open source that we've contributed to, you know, adding things like predictive sampling and the like to it. And, being able to use that sort of open source framework has really accelerated the development in this space. And of course, ROCm's integration with open source makes all of this really easy for us to deploy at scale.

63:27 Yeah, that's been so refreshing, all the work that we have done in the open. Now, as you look ahead, you've again expanded the footprint of activities and now we're looking at training, so that's super exciting, so maybe share a little bit about what we're doing there. Yeah, it's really interesting. As we look forward, we've seen such tremendous growth in inferencing, and we don't see any signs of that slowing down, and the Instinct looks to be a key part of our platform on inferencing going forward. but it's also great that it works really well as a training chip, and so we've been able to train, you know, on 2100 MI300Xs, you know, a stateoftheart multimodal model, in our research team. And, you know, really being able to use the same platform for inferencing and for training gives us tremendous flexibility in our data centers. and as we look forward, we're really excited to continue partnering with AMD on our inferencing and our infrastructure solutions.

64:21 That's awesome. And actually, this afternoon, there's more, information. There's actually a nice talk on the work, around training, so I encourage you to go, hear about that. Thank you so much, Eric. Of course. It's been, great having you and, wonderful Thanks so much, Razi ... partnership here. Microsoft has been an incredible partner, right? With e atscale deployments, running everything from closedsource GPT models to now open source DeepSeek, and extending the work now to large-scale training.

64:52 To talk about training, it's an increasingly important area of focus for us, and ROCm is making big strides there too. ROCm now supports all major parallelism strategies with functionality across major frameworks, and libraries, including PyTorch, JAX, TorchTune, and TorchTitan. And look, we're just not enabling models, we are also building our own. Training on ROCm internally at AMD is helping us improve performance, reliability, and the developer experience.

65:22 And just like inference, training performance has also taken a big leap. ROCm 7 delivers three times the performance of ROCm 6. More importantly, our users are actually telling us that they are now scaling confidently with ROCm. And one of those users is a tierone leader in AI models. Please welcome Aidan Gomez, CEO and cofounder of Cohere, to stage.

65:54 Aidan, thank you for joining us. It's so good to have you here. Thanks for having me. Tell us a little bit about Cohere and, your vision for where you're heading. Actually, before I do that, I, I, I actually should introduce you. Everybody knows you as a famous AI person, but there was this seminal paper, Attention is All You Need, and Aidan was one of the authors of that paper. Thank you. Thank you. Tell us a little bit about Cohere. Yeah, it'd be my pleasure. Yeah, thank you for having me. So, Cohere, what we do is we build highly secure and private AI specifically for enterprises. And our focus on security and data privacy means that we can serve large global en enterprises in some of the most highly regulated industries, like finance, healthcare, manufacturing, the public sector. And our products, in particular, our AI Workspace North, it gives AI agents the tools that they need to carry out extremely complex tasks securely. And so that spans the normal stuff, like emails and calendar and docs, but also the much more sophisticated stuff, like ERPs, CRMs, and even custom internal tools that are secured behind firewalls.

67:06 And with our models and our Product North, we're giving enterprises control, to really let them customize it to their needs and leverage all of their data in a secure and private environment. And most of Cohere's use cases rely on secure links to internal data, and that lets employees at large enterprises automate tasks around HR, customer support, finance, and even the supply chain. That's great. Now, you've been working with Instinct, running your models, inferring on them, and, running training on them. Tell us a little bit about how things have been going. Yeah, it's been going great. The partnership has been accelerating massively. So we were able to port our most recent model, CommandA, over to the AMD platform super easily, very quickly.

67:53 And our stack and models are now actively deployed on AMD and even at leading enterprise customers and global leaders like Fujitsu.... and we're extremely excited to start training at scale on AMD GPUs. Instinct's compute and memory characteristics make it a great platform for training our next model, and we're very pleased with how things are going and looking forward to all the innovation that's been announced here, and we're excited to get access. Yep, we're equally excited as well. Our teams are collaborating super close together. tell us a little bit about how you're taking advantage of the memory system in Instinct, particularly as you serve large models and more complex models like reasoning.

68:31 Yeah, so for agentic systems and complex reasoning, they really depend on the context window, that our models are able to support, and that can apply a lot of pressure to the, the memory, that's necessary to serve these models. and so that's because for agents and for reasoning, they spend a lot of time at inference, consuming tons of external data and putting that into the context, as well as reasoning over that data, and thinking in their heads before they actually respond.

69:03 So each one of these, increases, the computational demand on the hardware, and so the higher higher memory capacity and the strong memory bandwidth of AMD's chips have let us fit longer context onto the GPUs. And I think most importantly for us and our customers, it helps lower the overall footprint that's needed for our models, and that drives down the total cost of ownership for our customers. That's great. Again, you know, super delighted that the memory system is proving to be extremely valuable for you.

69:33 Now, as you look ahead, what do you foresee as the next set of things coming for enterprise AI, and what breakthroughs do you envision? So on the future, I'm extremely bullish about AI agents. I think that they're going to be deployed and used at scale, and we'll see a huge impact to both productivity and the types of work, the employees spend their time on, what their daytoday work looks like. So agents are gonna allow people to go beyond just augmenting work and towards actually fully automating tasks which take hours, days, or even weeks. and so an example of that would be doing research, over the course of weeks to answer some sophisticated question. Can we compress that down into a matter of days, or even hours?

70:19 So I'm really excited about AMD's roadmap with the MI350 series, and the rackscale MI400 solutions. it's a great choice and offering for our customers, and we can't wait to team up with you on it. That's awesome. Thank you so much for joining us. Thank you. Cohere is training and serving on AMD. We are so excited that we've been able to earn their trust at every level of the stack. Now, as inference becomes more computationally intensive and gets pervasively deployed into applications across industries, it is critically important to drive down its cost. And one of the most exciting new opportunities to drive down inference costs is distributed inference, so let's talk about it. Let's talk about distributed inference. In any LLM serving application, there are two phases. There's a prefill phase and there's a decode phase.

71:12 While it's simpler to deploy, in traditional inferencing serving applications, these two phases of the model are typically handled on the same GPU. But now if you apply it on the same GPU, it often becomes this bottleneck for large models or when demand spikes happen, and you can get limited in performance or flexibility. We can significantly improve throughput, reduce cost, and boost the responsiveness by disaggregating the prefill and de decode phases. Prefill and decode can be now assigned to specialized GPU pools, which can be independently optimized and with sparse MoE models and expert parallelism, there's even more room to optimize.

71:52 We have a great solution coming for distributed inference on AMD platform. Staying true to our strategy, we are embracing an open approach, building alongside an ecosystem of vLLM, SGLang, and LLMD. New technologies like GPU Direct Access and DPP deliver significant performance gains. Together, this stack enables a truly open and performant foundation for next generation distributed inference workloads. Now, as AI is moving into real world enterprise deployments, ROCm is evolving to meet those needs. Enterprises need more than just raw performance. They need end-to-end applications that helps teams hit the ground running, enabling easy and secure data integration for compliance and trust, and supporting robust workflows for ease of deployment.

72:44 To make all of this possible, today, I'm excited to announce ROCm Enterprise AI. ROCm Enterprise AI makes it easy to deploy AI solutions. With new cluster management software, it ensures reliable, scalable, and efficient operation of AI cluster, and our MLOps platform allows fine tuning and distillation of models with your own data, and a growing catalog of models will come for specific industries.

73:16 We partner closely with our ecosystem to deliver end-to-end applications that integrate with existing workflows and data systems, sometimes structured and sometimes unstructured. And to show how all of this comes together in a production enterprise stack and also discuss our strong collaboration on distributed inference, I'm excited to welcome to stage Chris Wright, CTO of Red Hat.

73:49 Hey, there. It's go so good to see you, Chris. Thanks for joining us. You bet. now, Red Hat and AMD, we've been collaborating for a long time, starting with, our x86 64bit architectures, but now we're extending it to AI. Tell us a little bit about sort of, what's exciting. Where do you see AI getting traction in enterprises today? Well, man, I love that you brought up 64bit X86 because we started there and it's been a long time. Yes.

74:14 we actually followed that up with virtualization, and that that support and effort, these things aren't static, right? So, fast forward to today, and that virtualization support is more important than ever as customers are looking for options to really virtualize their data center, and you guys just shared some amazing numbers at Red Hat Summit a couple weeks ago with, 77% OPEX savings and 71% power reduction, AMD and Red Hat together powering the the virtual data center. So that that's really cool.

74:54 Now, as for AI, quite a few things are happening. first, we you've you've seen it here today, we talked a lot about it, the surge of open. And some of that is open source software, the frameworks, things that we're more familiar with, but also open LLMs. And today, they have capabilities that are, on par with the really proprietary large-scale models including things like reasoning, so they're there or even outperforming in some cases. second, the emergence of vLLM. this is something really important for Red Hat, work that we're doing together, and this makes high performance inference deployments of open models easy.

75:33 And then third is bringing this vLLM support to a broad set of accelerators like AMD's. And so all of this together creates this ease of use to generate real efficiency and then choice for companies today. That's so good to hear. Now, we're not stopping there. Together, we have announced LLMD and open source distributed inference framework. tell us a little bit about why it is so significant for AI. I mean, we you've seen it here today already talking about, reasoning, talking about agents, talking about token production and and driving down the cost of token production. So, a key challenge for the data center today is lowering the cost of of token production. It's not just, dollars, tokens per dollars, but it's also tokens per dollars per watt. So really thinking about the overall efficiency to meet the GenAI demands of reasoning models in agentic workflows. reasoning models literally produce more tokens as they effectively think to produce results. And so our LLMD project is trying to address this need. How y how do you distribute and saturate these amazing instinct processors with requests to to respond to inference? you you mentioned a little bit earlier, the disaggregated prefill and decode, and and these are the the low level technologies that we're building into LLMD. LLMD builds on vLLM, and then extends that into a distributed environment with Kubernetes.

77:04 so we're so thrilled that you're joining us together in this journey and bringing your experience so that we can create this critical kind of industry initiative. Yeah. Our weight is behind vLLM and the open communities, and now with LLMD we can get to extend that further. So let's shift a little bit to OpenShift. OpenShift, AI is playing a key role in simplifying AI for enterprises and making it easy to deploy. How are we working together with OpenShift and, you know, what role do our platforms play in, that product?

77:34 Yeah. Well, broadly, Red Hat AI and AMD's processors, you know, CPUs, GPUs together bring this efficient production ready AI environment. so vLLM and LLMD are a key part of the Red Hat AI portfolio, which includes OpenShift AI. It includes the Red Hat, inference server specifically. And then the AMD Instinct GPUs are fully supported within OpenShift AI. So, a lot of work that goes into bringing that to life, and then this delivers this powerful AI processing across hybrid clouds. You heard Lisa talking about cloud, data center, edge, even even consumer devices, so that we we can deliver something for our customers to efficiently use these precious resources. OpenShift AI's both predictive and generative AI support needs smart CPU and GPU choices, and our work with AMD ensures this flexibility and maximizing the customer investment so they're getting the most out of the hardware that they're procuring.

78:35 Yes. Super exciting and, you know, very, very happy with the collaboration that's, we've had around OpenShift. So now as you look ahead, right, it still feels like we are in the very early innings of enterprise AI, right? So, you know, what excites you about what's coming and, you know, the work that we can do together? Yeah. Early and yet moving so fast that things change fundamentally day to Every day. Daily. Yeah. I think it's clear that gen AI is gonna deliver huge value, both in terms of efficiencies or net new value for enterprises.

79:07 I think the pressure is on each and every one of us to help get from those those pilot projects, those POCs, into production. And so our mission is to make that as efficient and accessible as possible. And much in the same way that Linux brought to life all these applications across different kinds of infrastructure, we're doing that. We're entering the same era with AI. And so to me, I think it's happening right now with Red Hat AI and AMD and what we're doing together to really unlock that AI value for enterprises across every different kind of industry vertical.

79:48 That's so good to hear, Chris. we are super grateful for all the work we are doing together. Our teams love working with each other. Thank you for joining us. Absolutely. Thank you. With OpenShift AI and ROCm, we are now enabling enterprises with GenAI workflows. I'm especially excited with the joint work we've done on LMD to slash the cost of reasoning and now agentic based inference.So now, none of this happens without developers.

80:19 So let's talk a little bit about what we are doing there. We are deeply, deeply committed to delivering an exceptional developer experience. We've significantly stepped up our efforts to make the out of the box experience better, and deliver great collateral. From videos, blog posts, tutorials, we're helping developers ramp up fast. And with frequent meetups, hackathons, contests, we're building a community. I was actually so excited to see that our recent contest developed GPU kernels generated huge interest with thousands of submissions, including a high schooler who wrote high performance Triton kernels. That was just so good to see.

80:54 And over the last year, as we have enabled the cloud access to AMD GPUs, there's been a big ask from the development community for an AMD developer cloud. So today, I'm super excited to announce the AMD developer cloud. Instant access to AMD GPUs, no setup, pure development velocity. Every developer in this room has a 25hour free GPU credit email in your inbox. No strings. Just launched, go.

81:32 Now, to show it in action and to tell you about all the collateral we are going to bring to you as part of this dev cloud, please join me in welcoming Anush Elangovan and Sharon Zou to stage. Anush. Hey, Sharon. Anush is responsible for a number of open source software efforts here at, AMD, and is actually well known for his huge passion working with developers. he was previously the CEO of nod.AI, a company that was famous for their open source compiler contributions.

82:07 I'm also thrilled to welcome Sharon. Sharon is also very well known in the AI community. A former, a former Stanford faculty, she was the CEO and founder of Lamini. I am delighted that Sharon and her talented Lamini team joined us recently, with a focus of delivering rich content for developers. So Anush, tell us a little bit about the dev cloud and all the goodies. Thanks, Vamsi. Developers, developers, and developers. That is the new mantra of ROCm.

82:43 We are serious about bringing ROCm everywhere, and to everyone, from client to the cloud. In AI, speed is your mode. Access to compute is paramount. We've been delivering on speed, so now, let's get you access to compute. Today, we are announcing the general availability of the AMD developer cloud. With the developer cloud, anyone with a GitHub ID or an email address can get access to an Instinct GPU with just a few clicks.

83:21 All right, let's see how easy it is to get access to an AMD GPU in the cloud. Go to devcloud.amd.com, say hello to a legal friend, and sign up with GitHub. That's it. You can choose between a one GPU VM or an eight GPU VM, and you select the operating system that you'd like to use. One of the cool new features of ROCm 7 is that we've made it really, really easy to install. Just pip install ROCm. In case you forget, we've also printed it in a Tshirt, and it's in your goody bag.

83:59 We've also included a lot of, easy to use frameworks, like vLLM, SG Lang, PyTorch, et cetera. You just select one of those frameworks, add your SSH key, and then create, and you're set. We've also spent a lot of time building a lot of Jupyter Notebooks, making it easy to use. And if you've been tracking the latest attention algorithms, the log linear attention came out a few days ago. You could try something like that on the MI300X just in a few minutes.

84:31 And we're just getting started. ROCm is open, proven, and now, really, really accessible. Sharon? Hi, everyone. I'm Sharon. I've taught AI to nearly a million people, many of you, at Stanford, as well as on Coursera with my startup, Lamini. As Vamsi and Lisa just shared, I'm super excited to announce that Lamini has now joined AMD.

85:01 I'm personally very excited to be part of AMD's AI mission. We're just getting started, as Anush said, alongside extremely talented teammates from Lamini. We're here to make AI and AI compute easier to use and scale for you, the AI developer, you, the AI researcher, you, the AI leader in this audience. What you may not know, many of you, in fact, tens of thousands of you, have already run on AMD GPUs over the past year. And that's through Lamini courses with myself and Andrew Ng, who you'll hear from later today. And that's on prompting open source LLMs, LLM fine tuning, and improving LLM accuracy in partnership with Meta.

85:48 And we're gonna amp that up further here by creating a huge set of intuitive engaging courses, from LLM post training and reinforcement learning, to vibe coding agents, to GPU programming. All of this humming on powerful AMD Instinct GPUs on our developer cloud that Vamsi just announced, and there will be hands on tutorial in this afternoon's developer track, to get you started. We'll also be out in the community, ears to the ground, listening to your feedback at top AI conferences. So whether you're at a foundation model company, an AI startup, university lab, hacker house, or just someone attending their first AI hackathon, don't be shy. Come say hi.

86:33 Thank you, Anush. Thanks, Sharon. Thanks, Vamsi. So, you just saw how easy it is to access our dev cloud, but what if you want to develop locally on your own machine with your own data? That's where we're going next, because ROCm isn't just for the cloud anymore. We are expanding ROCm to Ryzen laptops and workstations, so you can build anywhere using the same software stack from cloud to client. Whether you're on Linux or Windows, cloud or client, ROCm is there. Coming to you in the second half of this year, ROCm will include... will be included directly in major distributions. Windows as a first class OS, fully supported and production ready.

87:19 And you can do that on the best AI client portfolio in the industry, capable of delivering breakthrough AI experiences all locally. So, we've talked about all the exciting capabilities in ROCm. We've talked about empowering developers everywhere. Now, it's time to hear from the builders themselves. So, we have a fantastic program for later today. Join us this afternoon for the developer track featuring leaders that are driving the shift to open and scalable AI.

87:53 Hear from them how they're enabling their communities to build on AMD. So as I close, let me leave you with this. We built ROCm to empower the world with an open software platform that unlocks AI innovation for everyone everywhere, and we made tremendous strides in just the last year. Our strategy of combining forces with the open source ecosystem is paying off. Together, we are delivering a credible, high performance alternative.

88:24 ROCm is delivering some of the most important AI workloads on the planet today, but this is just the beginning. We are going to push forward with urgency, with focus, and with a deep, deep commitment to developers, because the future of AI is not closed. It is open, it is collaborative, and it is for everyone. Now, to deliver AI at scale, we need to bring system level solutions together that integrate computing, networking, software into a unified AI platform. To tell you about all the progress we are making over there, it is my pleasure to invite Forrest Norrod, EVP and GM of our Data Center Solutions Group, to stage.

89:16 Thank you, Vamsi. As Lisa said to start this morning, we're moving into the next phase of AI. From a period where chatbots were interesting curiosities to an era where AI drives business and innovation. And agentic AI, as we've heard, is a leading driver of that change. AI u agent usage is exploding across use cases and industries, not just automating manual labor intensive tasks, but optimizing and automating complex workflows with planning, analysis, and creative problem solving.

89:51 So, not just streamlining processes, but driving innovations across business, science, and product development. Just as information technology revolutionized the paper based economy into a digital one, agentic AI brings about another revolution, an innovation revolution where new ideas can be implemented at an unprecedented rate. And so, agentic AI has the power to impact workflows across many fields. The key in agentic AI is connecting the power of the LLM models to the business, to the organization, to its datas, tools, and applications. agentic flows will employ many models, including specially trained models, each performing their own roles, but working together to execute complex tasks.

90:43 These AI agents execute multistep processes, many of which will need access to enterprise tools, datas, even humans. So these agents are not simply running isolated on a few GPUs. Each agent accesses many different resources, applications, databases, unstructured data from social networks. The list could be endless. And they map onto real hardware, onto the GPUs, of course, but also onto a host of CPUs running the applications and processing data going into and out of the GPUs. And, onto the network infrastructure providing secure access to that data.

91:25 Now, agents challenge the GPU. They do more than chat, and they need higher performance inference and more memory for larger reasoning models and larger context windows, things that you've heard about earlier today. But equally, CPUs are at the heart of agentic execution, running both enterprise applications as well as managing and orchestrating AI systems. And the data fueling all of this flows across the networks connecting everything, but that data includes the crown jewels of any organization, and hence it must not just be accessible quickly, but above all, it must be secure.

92:07 So thus, agentic AI will increase the demands on every part of the data center, not just the GPU, but the CPU and networking as well. At AMD, we build the technology powering each one of those elements. Our Pensando NICs to securely access data, EPYC CPUs, the industry's best, to process the data and manage the GPUs, and of course, the Instinct GPUs to power agentic model execution.Beyond that, the scale-up and scale-out networking for AI scalability allows you to go from small enterprises to a gigawatt data center.

92:48 AMD has worldclass technology in all of these elements, and we have the ability to put it all together. But we also believe firmly in the principle of open. We have taken the lead on helping the industry develop open standards, allowing everyone in the ecosystem to innovate and work together to drive AI forward. We utterly reject the notion that one company could have a monopoly on AI or AI innovation.

93:19 History shows the most vibrant ecosystems are open. Now, another key belief at AMD is the principle of programmability is critical. AI is evolving so quickly that having fixed function devices or limited accelerators is the wrong approach and will slow down progress. Software innovation for many, including folks like DeepSeek, has shown time and time again the value of flexibility. Putting all of those elements together now in an open, holistic, programmable design results in the optimal pro platform to power the age of agentic AI.

93:59 So let's look at each element. The frontend network connects the compute nodes to the rest of the world. It's the bridge to the AI node. With agentic AI, as I said before, data is evermore important and security is paramount. But security is a layered discipline. With AMD's advanced DPU technology, we support encryption, authentification, and eastwest firewalls on every connection. The key to all of this is Pensando's flexible third generation P4 engine that delivers data with security and performance. Turning to compute. Some will naively tell you that CPUs are less important in the age of AI, but that's not correct. With agentic AI, we see an explosion of autonomous agents accessing data and enterprise applications. This increases the needs for efficient, high performance x86 compute across the data center.

94:58 Then within the AI server itself, the CPU serves the demands of preprocessing and workload orchestration to keep the GPUs working efficiently. Our EPYC CPUs with boost frequencies up to five gigahertz and the highest server CPU performance available, period, are perfect to feed those GPUs. But just as importantly as performance, the CPU needs to be able to seamlessly integrate into a user's environment.

95:29 Our x86 EPYC CPUs not only bring trusted enterprise reliability, but provide architectural consistency across the data center, increasing flexibility and performance, enabling workloads to move seamlessly to wherever they can get the best levels of latency and throughput. Now, let me show you a few examples of how the right CPU can make GPUs work better, and how choosing poorly can create bottlenecks that strand valuable resources. As you can see, across a range of models and use cases, our fifth generation EPYC CPUs can boost the inference performance of the entire system from 6% to 17%.

96:13 That makes a huge impact on the overall TCO and performance of the AI deployment, and it's a critical element in designing the best possible AI system. So, get much more out of your GPUs with the right CPU. And as AI gets more advanced, particularly with new model architecture innovations like mister, mixture of experts, or as MCP becomes ubiquitous, the right CPU will continue to be critical in delivering AI performance.

96:46 Well, so the CPUs drive the GPUs. And for five generations, AMD has perfected our Infinity Fabric architecture, connecting the CPUs and GPUs together in a low latency, high speed coherent interface. As part of our belief in open standards, we donated key IP from Infinity Fabric to the Ultra Accelerator Link Consortium. UALink expands the protocol, scaling well beyond eight interconnected GPUs up to a thousand coherent GPU nodes, enabling AI systems to ramp, deliver GPU performance for training and distributed inference, and, and for whatever innovation software develops next.

97:29 Ultra Accelerator Link 1.0 specification has been released. It's a modern load store architecture engineered for the demanding needs of scale-up AI systems, including low latency and high bandwidth. Now, importantly, it leverages the physical interface layers of ethernet enabling standard components such as connectors, cables, and retimers to be leveraged by the ecosystem and drive favorable economics and reliable interconnect. And UALink isn't just optimized for performance, it's engineered to scale.

98:04 This open standard allows customers to build and support tailored systems, scaling up GPUs spread across racks, enabling pod partitioning for efficiency and security, delivering rock solid resiliency, and accelerating performance going forward with support for in network collectives. But one of the most important features of UALink is it is an open ecosystem. It's a protocol that can be used in a system regardless of the brand of CPU accelerator or switch. It is thus fully opened rather than being shackled to one company's systems or technology.Again, AMD firmly believes in the power of an open, interoperable ecosystem that accelerates innovation and protects customers' choice while still delivering leadership performance and power efficiency.

98:57 The consortium is store, steered by some of the largest scale users and suppliers in the world, hyper scalers and leaders in the semiconductor industry. We are excited to invite some of the contributors to the Ultra Accelerator Link Consortium to the stage. Please welcome Jitendra Mohan, CEO and cofounder of Astera Labs. Jitendra, thank you so much for joining us. I, I know we're both excited about UALink.

99:30 can you tell us from your perspective what makes this so exciting and why Astera has chosen to focus on it? Absolutely, Forrest. But first, those 5 gigahertz CPUs are cool. They make our chip simulations run faster. Fantastic. So, thank you for the partnership. I'm, I'm really stoked to be here. we founded Astera Labs seven or eight years ago with a mission to eliminate AI infrastructure bottlenecks throughout the data center. That's what we've been doing. From the beginning, we have been laser focused on delivering solutions that meet our customers' demands.

100:05 In fact, we partnered with AMD on PCI5 before the spec was final. Mmhmm. We have a strong track record of taking cutting edge open standards and delivering market leading products. At Astera Labs, we know an open approach works. It spurs innovation, builds robust ecosystems, and results in wide adoption. Today, we provide a comprehensive portfolio of connective solutions for the entire AI rack. scale-up connectivity is a particular focus for us, because it is the most critical element of AI rack architecture, and UALink is purpose built from the ground up for scale-up.

100:43 There is no baggage, no backward compatibility. UALink is designed to be efficient, fast, robust, and it combines the best of many protocols. UALink for scale-up completely aligns with our mission, our expertise, and naturally fits into our roadmap. What is more, our customers are asking us to deliver UALink products to take the next step forward in deploying a truly open rack scale AI platform based on a vibrant ecosystem. And, Forrest, in this case, I must say, the customers are coming. We just need to build it. Absolutely, completely agree. I'm hearing the same from, from, the, particularly the key hyperscalers. Now, what do you plan to build on UALink?

101:25 Great. our vision is to provide complete connectivity infrastructure for the entire AI rack. This includes purpose built silicon, hardware, and software to support AI platforms based on custom ASICs and merchant GPUs, including AMD's Instinct Solutions. Fantastic. We are at the forefront of scale-up connectivity innovations with our Scorpio Xseries fabric switches and our Aries eTIMES. As a UALink consortium board member, we are working with AMD and industry leaders to advance UALink.

102:01 We have a closeup view of the features and timeframes needed by our customers to realize their vision of deploying UALink based open rack architectures. We are working shoulder to shoulder with AMD and XPU partners. We plan to offer comprehensive portfolio of UALink products to support UALink deployments at scale, smart fabric switches, signal conditional controllers, and many more. All of these solutions are built on our Cosmos software that provides an unparalleled view into the health of the entire rack.

102:33 Our cloud-scale interop lab provides a robust validation environment for ensuring interoperability at rack scale and accelerate time to market for our customers. Together with AMD, we are excited to bring UALink to scale-up AI infrastructure. Fantastic. Amazing. We're, we're just as excited to be working alongside you and the whole team at Astera and the, the whole UALink, consortium to drive it forward. Thank you so much for joining us here today, and thanks for your partnership. Thank you, Forrest. Thank you, everyone.

103:04 Thanks, Forrest. Now, I'd like to welcome another guest and fellow member of the UALink consortium, Nick Kucharewski, SVP and GM of Network Switching BU and Cloud Platforms at Marvell. Good morning. Nick, thank you so much for joining us. glad to be here. you know, Marvell is well known as a leader in custom, custom ASICs and custom solutions for hyperscalers, and, and you're engaged on many networking topics as well. Tell us what your customers are telling us, or telling you about, UALink and scale-up.

103:42 Yeah. As you know, Marvell is deeply involved in infrastructure technology for cloud and AI data centers, including high speed electrical and optical connectivity, switching, storage, compute, and custom silicon. And in that, in that process, we've developed partnerships with customers who are really operating at the forefront of cloud compute infrastructure and AI technology. And one of the questions we hear often is that, what standards based options exist for building a large scale-up AI cluster that enables high bandwidth, low latency, high reliability, and the capability to scale beyond today's rack level implementations to clusters with hundreds of connected accelerators?

104:22 Now, UALink is at the center of that conversation, because it enables all of those attributes, and it also carries with it the promise of an ecosystem of interoperable components from multiple suppliers. Yeah, totally agree. Now, you've got a pretty broad portfolio already, but tell us, what, what are your specific plans around UAL? Yeah, sure. So we've been involved with UALink from the beginning, and Marvell engineers are active in the working group supplying our expertise in high speed interconnect, low latency fabrics, high layer packet processing, and the networking software stack.

104:54 This week, we announced UALink is part of the Marvell custom cloud platform for system, system designs and silicon. Now, this solution can enable next generation scale-up fabrics and endpoints, offering interoperability between GPUs and switches for next generation AI infrastructure.UALink joins the broader Marvell offering for custom AI silicon, which is rooted in decades of expertise and billion transistor design, and our portfolio of design IP, including networking cores; high speed SerDes for rack scale connectivity; copackage optics for row scale; and our family of connectivity and switching for scale out networks. But with UALink, Marvell customers can deliver a platform comprised of their own custom vision, working literally side by side with interoperable silicon, GPUs, and fabrics from UALink partner companies.

105:46 Nick, that's a compelling vision. Customers want choice, and they want the ability to innovate freely. I think, together, we're gonna give that to them. Thank you so much, and thank you Yeah ... for your, for coming to visit us today. Yep. Thanks very much for having me here today. Take care. Thank you. So UALink enables scaling up coherent GPUs, soon to over a thousand, but the most complex AI systems need to scale out way beyond that, to truly gigawattscale deployments. That level of scale drove the Ultra Ethernet Consortium standard. UEC leverages the complete ethernet stack, but it's more than ethernet. The UEC standard defines a whole new transport layer, addressing the challenges of efficient data center wide deployments. The result, an unparalleled scaling capacity of a shared memory fabric to over a million GPUs. UEC delivers a set of capabilities well beyond InfiniBand.

106:43 AMD is proud to be a founding member of UEC, and we're excited that the UEC standard 1.0 got to full release yesterday. And we're proud, as well, to have the industry's first UECready NICs. We introduced the third generation Pensando P4 engines last fall to drive frontend networks, but their incredibly flexible and performant P4 packet processing technology allows them to match the rate of innovation and is ideally suited for the unique needs of backend AI networks.

107:22 Pollara 400 supports advanced transport and congestion control innovations from multiple standards and multiple custom solutions for customers, including shortly UEC 1.0. We've seen Pollara improve AI performance while reducing network costs for customers by up to 22% through higher fabric utilization and more uniform and simpler switch deployments, while also improving system reliability and resiliency by up to 10%.

107:55 That improvement in resiliency and availability is ever more important as AI evolves into mission critical agentic applications. With a backend network, we complete the end-to-end AI platform needed to support agentic AI and drive AI forward. And at AMD, we know that agentic AI isn't just a vision or a concept. It is emerging here today. Our customers want it. The industry is demanding it, and we are enabling it with a leadership portfolio of products and our OpenRack infrastructure.

108:29 To develop that leadership p performance at scale, again, you need more than a powerful GPU. You need a modern Open Rack architecture pur purposebuilt for AI. You get that with Selina 400 DPUs for frontend networks, the fifth generation AMD EPYC CPUs, the AMD Instinct 350 series GPUs, and scale-out networking solutions with AMD Pensando Pollara AI NICs. All integrated together into an industry standard OCP design, fully supported with O UEC NICs and offering unprecedented performance.

109:06 The industry thrives on, it requires an open ecosystem. Open done right enables fully optimized rack level infrastructure without proprietary lockin, and enables innovation across the industry. To show us how we take these principles to the next level, please join me in welcoming Dr. Lisa Su back to the stage. Thank you. All right. So, look, you've heard a lot from Vamsi and Forrest and a bunch of our customers and partners about all the momentum we have across, you know, hardware, software, and solutions. But now let's talk about the future and how we're expanding our rackscale solutions portfolio to essentially deliver compute performance, efficiency, and density that customers need over the coming years.

109:59 Today, I am super excited to give you a first look at the next big step for our AI roadmap, the Instinct MI400 series. You may hear us call it MI400 series. You may hear us call it MI450. MI400 series is really bringing together everything we've learned across silicon, software, and systems to deliver a fully integrated AI rack platform. And this guy was built from the ground up for leadership for both large-scale training and distributed inference.

110:31 Let me now introduce you to our Helios AI Rack. Helios is truly a gamechanger. For the first time, we architected every part of the rack as a unified system. That's combining our CPUs, our GPUs, our Pensando NICs, and our ROCm software all together in one platform. And it's really purposebuilt for the most demanding AI workloads, from training, to the largest, the largest inf frontier models, to scaling inference across thousands of nodes. But Helios has more than just lots of compute.We also have leadership memory capacity, leadership memory bandwidth and leadership interconnect speed. And all of that is delivered in an open OCP compliant rack that supports both Ultra Ethernet and UALink.

111:27 And when Helios launches in 2026, we believe it'll set a new benchmark for AI at scale. So think of Helios as really a rack that functions like a single massive compute engine. It connects up to 72 GPUs with 260 terabytes per second of scale-up bandwidth. It enables 2.9 exaflops of FP4 performance, and that is a great number, but Helios goes even further.

111:58 Compared to the competition, we support 50% more HBM4 memory, memory bandwidth, and scale-out bandwidth, and these are big advantages. I mean, this is our sweet spot. We've always had this memory architecture, and what this translates in is faster training, higher inference throughput, and the ability to really handle massive models. Now, let's take a look at each of the components that make Helios possible, starting with our next generation EPYC processor, codename Venice.

112:31 Venice extends our leadership across every dimension that matters in the data center. More performance, better efficiency, and outstanding total cost of ownership. It's built on TSMC's 2 nanometer process and features up to 256 high performance Zen 6 cores, and it delivers 70% more compute performance than our current generation Leadership Turin CPUs. And to really keep feeding MI400 with data at full speed at, even at rack scale, we've doubled both the GPU and the memory bandwidth and optimized Venice to run at higher speeds. And you heard from Forrest how important the CPUs are.

113:11 Now, we just got Venice back in the labs, and it is looking fantastic. Now, at the heart of Helios, though, is the MI400 series. This is truly the most advanced accelerator we've ever built. It's really engine for the next generation of AI, and it's designed to run trillion plus parameter models. We deliver up to 40 petaflops of FP4 performance. We have 432 gigabytes of HBM4 and supports 300 gigabytes per second of scale-out bandwidth to connect across racks and clusters.

113:53 And now, as you've also heard from Forrest, we need a high performance networking fabric to connect all of that, and that's why we're also introducing Volcano, our next generation scale-out AI NIC. Volcano is fully UEC 1.0 compliant. It supports PCIe and UAL interfaces to connect directly both GP CPUs and GPUs, and it delivers 800 gigabits per second of line rate throughput to scale for the largest systems. Now, with Helios, every GPU in the rack is connected through the high speed, low latency UALink tunneled over standard ethernet.

114:30 Now, when you look at our AI roadmaps, you know, every generation is always special, but Helios is truly a giant step forward. With MI355, we're taking a big step forward. You've heard some of that this morning. We're delivering 3X more performance across a broad range of workloads, extremely competitive versus the state of the art today. And with Helios, we are bending that curve further. The MI400 series is expected to deliver up to 10X more performance for the most advanced frontier models, making MI400 the highest performing accelerator.

115:09 I think 10X is a good number. Is it a good number? Yeah! Yeah! Look, if I sound excited, that's 'cause I am excited, and, as you might expect, customer excitement for the MI400 series and Helios is really high. Like, these are the types of programs you don't just start today. I mean, we have been working with customers for the past few years to really, you know, really, like, just jump ahead of the curve and see what our customers really need.

115:40 One of those customers who has been a very, very early design partner, who has given us significant feedback on the requirements for next generation training and inference, is OpenAI. And we have a very special guest today. I am so happy to say that, you know, this person is a great friend, someone who is really an icon in AI. To hear more about our work, please welcome OpenAI founder and CEO Sam Altman to the stage.

116:14 Can I call you an AI icon? I don't think so, but that's okay. I like 'Cause you know what? It's your show. You do whatever you want. Sam, look, we are truly so happy and excited to be your partner. OpenAI has truly been at the center of the universe. Everyone listens to what Sam Altman has to say, when it comes to GenAI, and, I think I think they just listen to ChatGPT at this point. Actually, I listen to ChatGPT.

116:39 We'll take it. I listen to ChatGPT. you know, some of the numbers I've seen, like over 500 million weekly active users, just amazing growth. can you just give us a little bit of a landscape. Where are we today? What's the state of play? What are you seeing? you know, what, what's most exciting right now? It, it's definitely been, for us and many other people, just an explosion of usage over the last year. I think the models have gotten good enough that people have been able to build really great products, text, images, voice, all kinds of reasoning capabilities. we've seen extremely quick adoption in the enterprise now. Coding's been one area people talk a lot about. But...I, I think what we're hearing, again and again in all these different ways, is that these tools have gone from things that were, you know, fun and curious to like truly useful Doing real work people's personal lives, and the, the fact that you can now like ask a system like Codex to go off and auto do some like work for you autonomously over minutes or hours, it's, it's like pretty remarkable. Yeah. I mean, I, I think the, the key point that you said is, is really enterprises are seeing lots and lots of value. you know, I think the other thing that's been amazing is, man, I mean, the rate and pace of what you guys are putting out, it seems like every week you have a new model. you know, workloads are just changing so fast. like what are you seeing? Like h how are things changing, and, you know, most importantly for us, like how are you seeing compute demands changing?

118:05 I mean, tons have changed all the time, but o one of the biggest differences has been we've moved to these reasoning models. so we have these very long rollouts where a, a model will go off and think about a problem and come back with a, a better answer, or y you know, like in some cases, like a whole PR ready to go. But this has really put pressure on model efficiency, and long context rollouts. we need tons of compute, tons of memory, tons of CPUs as well.

118:33 I've seen that actually And our, like our infrastructure ramp over the last year and what we're looking at over the next year has just been a crazy, crazy thing to watch. Is there ever enough GPUs? I mean, like theoretically at some point you can see that like a significant fraction of the power on Earth should be spent running AI computes. and maybe we're gonna get there. Yes. Yes. No, that's, that's definitely true. look, we have been honored. Really we've, we've really, really appreciated the partnership and collaboration, between, you know, OpenAI and AMD over the last few years, you know, working together in Azure, you know, working on some of your research stuff and particularly, you know, the deep design on MI450. I think, you guys were really early in just some of the important insights. can you just tell us a little bit about, you know, how that's evolved and, and sort of like how we can do more for you. It's, it's been amazing working with y'all, obviously. And, you know, we're, we're already running some work on the, 300X, but the, the MI450 series, I, I think, and the, the work we've been able to do there as you've worked on that over the last couple of years, and we're very grateful for you listening to our input. hopefully we'll be a good representative for what the industry as a whole needs. but we are extremely excited for the MI450. The memory architecture is great for inference.

120:01 Believe it can be incredible, option for training as well. And the... When you first started telling me what you were thinking about for the specs, I was like, "There's no way." That just sounds totally crazy. That's too big. but i it's really been so exciting to see you all get close to delivery on this, and I think this is gonna be an amazing thing. Well, first of all, thank you for saying that. I appreciate that very much.

120:22 One, one of the things, that really sticks in my mind is when we sat down with your engineers, you were like, "Whatever you do, just give us lots and lots of flexibility because things change so much." And, you know, really that that, that framework of, of working together has been phenomenal. now, you know, Sam, look, this is a, this is a moment here where we have lots of folks in AI wanting to know like, you know, what's next. So, help us with big picture. Like, what do you see in the future? You know, perspective on where things go, you know, how did, how did the workloads evolve? What happens with, you know, quote unquote "AGI" and, and really, you know, how do we as AMD and we as, you know, the computing industry kind of help enable all of that for you?

121:09 At the beginning of the 2020s, we didn't have We didn't kinda have AI as we think of it today yet. We had a bunch of other systems, but that, that was still, you know, the preGPT3 era just by a little bit. And now as we sit at the sort of like halfway mark through the decade, it's really been remarkable progress from, you know, not even a GPT3 model to, GPT4.5 and 03. These models that, that really feel smart and helpful and can give these sort of, you know, real utility experiences where people would look at this, if they could go back in time and say, "That feels, that feels almost impossible." Like if, if, if you went back to 2020, said, "By halfway through the decade, we're gonna be at this system that you can talk to, and it's really smart. It's like a smart person. It can do work for you."

122:00 I think we're gonna maintain the same rate of progress, rate of improvement in these models for the second half of the decade as we did for the first. I wasn't so sure about that a couple of years ago. There were new research things to figure out, but now it looks like we'll be able to deliver on that. So if you think forward to, to 2030 and the systems that we can have, these systems will be capable of remarkable new stuff, novel scientific discovery, running extremely complex functions throughout society, and things that we just couldn't even imagine was possible before.

122:28 To, to get there, to be able to deliver on this, it's really gonna take... You know, these are huge systems now, very complex engineering projects, very complex research. And to keep on this curve of scaling, we've gotta work together across research, engineering, hardware, how we're gonna deliver these systems and products. and this has gotten quite complex. But if we can do on that, if we, if we can deliver on that, if we can drive this collaboration across the whole industry, we will, we will keep this curve going. And so we're tremendously tremendously excited about the work that we're doing with AMD, and what you all are gonna deliver to, you know, we'll keep delivering great models.

123:07 Sam, I can say that, we really, really appreciate the work with OpenAI. You guys push us. You guys push us hard. but at the end of the day, we all want to deliver that vision. So thank you so much for being here. Thank you for the partnership. Thank you very much for having me. Yeah. Thank you for the partnership too. See you. All right. Take care. Thank you. All right, so as you can tell, we are super excited about what MI400 brings to the market. there are lots of active customer engagements already. You know, this is about, really co-optimizing, together. But it really doesn't stop there. We're already deep into development of our 2027 rack that will push the envelope even further on performance, efficiency and scalability with our next generation Verano CPUs and MI500 GPUs. So lots and lots of stuff to come from AMD.

124:01 Now that brings us to the close. It's truly been an amazing day. We've covered a lot, from the launch of MI350 series to our next generation MI400, to the Helios rack scale solutions, to all of the incredible momentum that we have building our open software and hardware ecosystems. And I really want to say a big thank you to all of our partners who joined us today on stage. There are a number of partners who have helped us, with, putting together this event. There are a number of breakout sessions I hope you guys get to later on in the day. And hopefully what you've gotten from today is that we're moving faster than ever before to deliver the best AI solutions for the market.

124:40 But let me just end with a few personal thoughts. You know, when I think about this past year, it's really redefined what progress in AI looks like. It's really moved at a pace unlike anything that we have seen in modern computing, frankly, anything that we've seen in our careers, and frankly, anything that we've seen in our lifetime. You know, we in this community, I call this community the AI ecosystem, we're really at the center of everything that matters.

125:09 And isn't that just an incredibly phenomenal place for us to be? I think of it as a journey. I've always said, you know, this would be a journey, and I'm incredibly proud of how far we've come. But more than that, I'm actually really proud of how we're bringing together the technology, the talent, and the partners needed to make AI more powerful, more accessible, and more useful for everyone. The future of AI is not gonna be built by any one company or in a closed ecosystem. It's going to be shaped by open collaboration across the industry. It's gonna be shaped because everyone is bringing their best ideas, and it's gonna be shaped because we're innovating together.

125:53 So on behalf of all of us at AMD, we look forward to changing the world with you together. Thank you for joining us today.

© transcribe · For agents Built with care and craft by Gokul Rajaram