Section Insights
The Evolution of Processing Speed
How is the perception of processing speed changing in AI?
The perception of what constitutes fast processing is evolving, with traditional speeds of 100 to 200 tokens per second now seen as standard. This shift highlights the growing demand for faster prompt processing and parallel workloads, although it is acknowledged that this alone is not sufficient for the market's needs.
- Processing speeds are rapidly increasing, changing industry standards.
- Faster prompt processing is essential for handling parallel workloads.
- The market requires more than just speed; it needs comprehensive solutions.
Transformational Hardware and Model Integration
What is the significance of combining advanced models with faster hardware?
The integration of intelligent models with high-speed hardware is transforming the industry, enabling the resolution of real problems and the development of new capabilities. The focus is on scaling this integration to meet user demands and pushing the boundaries of performance.
- Combining advanced models with fast hardware is a game changer.
- The industry is transitioning from demos to practical applications.
- Scaling capabilities is crucial for meeting user needs.
The Impact of Jalapeno on Throughput and Latency
How does Jalapeno influence the AI hardware landscape?
Jalapeno is setting new standards for throughput and latency, which benefits the entire industry. Its capabilities will allow for a fast inference portfolio that surpasses existing technologies, emphasizing the importance of speed in AI applications.
- Jalapeno enhances both throughput and latency, benefiting all players in the industry.
- The upcoming CS5 will provide significant advancements over current GPU capabilities.
- Speed is a critical factor in the future of AI hardware.
Heterogeneous Inference Solutions
What are the key considerations for heterogeneous inference solutions?
Heterogeneous inference solutions require a balance between speed and capacity. The right architecture mix is essential to serve diverse use cases effectively, and the technical perspective emphasizes the need for flexibility in hardware deployment.
- Speed and capacity are crucial for effective inference solutions.
- Finding the right architecture mix is key to serving various applications.
- Flexibility in hardware deployment is essential for scalability.
Innovative Directions Beyond Chip Design
What future directions should AI hardware development take?
The future of AI compute lies in innovative integration beyond just chip design. This includes advancements in rack, node, and interconnect technologies, as well as optimizing algorithms for scheduling and pipelining to handle larger models and applications.
- AI hardware development must focus on integration beyond the chip.
- Innovative solutions in interconnect and memory packaging are vital.
- Handling large models requires a holistic approach to compute architecture.
Transcript
0:00 what used to be considered fast at like 1,000 100 or or 200 tokens per second is quickly becoming the new batch mode. And so I think that you know there's a place for that. It's great for prompt processing. It's great for very very parallel workloads. >> but my emails take 20 hours. I can hit a button and run it in two hours. >> Exactly. Exactly. Right. But but quickly, you know, being able to do prompt processing isn't isn't quite enough. Right. That's a big part of the market, but that's not quite enough.
0:34 >> Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis, but fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way.
0:57 But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you. And it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If you do it, I promise you we'll never stop working to make this show even better. Now, let's get into it.
1:22 Okay, we're here at Crisis HQ with CTO Sha Lee. Welcome. >> thank you. Thank you for having us. >> Yeah. No, thank you for having me. >> and it is the day after hot chips. Lots of things launching. Even like Apple launched the M6 stuff. but you guys obviously had we were at Supernova. You talked about CS4. We're going to talk a little bit about CS5. OpenI talked about Jalapeno. All the all the hot stuff in hot chips. What's your take this year having been in this industry?
1:53 I think I think well, first of all, thank you for having me. and and and you're absolutely right. It's a super exciting time right now. hot chips is a is is a is really good, you know, time when the whole community comes together. and you know what what used to be, you know, I still remember going hot chips 10, 20 years ago when it's a bunch of like, you know, computer architect, you know, geeks, you know, geeking out on, you know, speeds and feeds and and now it's like this is where, you know, industry changing hardware is is being revealed. And so, you know, it's come a long way. It's really exciting, you know, community.
2:28 And I think one of the main takeaways that that that I had after, you know, after leaving Hot Chips yesterday is like what an exciting time it is for the the chip design and the hardware industry. It's you know, this is really the golden age of of hardware right now. and you know, like you said, I've been in this industry for some time. this is a very very unique time not just because you know AI is taking off and you know and there's you know obviously a lot of people doing their own hardware but the amount of innovation that's happening now across multiple fronts we've never seen in the history of of the semiconductor industry this level of innovation across all the levels right in the chip the interconnect the system design the software the optics the you know the methodology the tooling people are pushing across the board and so like you know as a technologist and and you know a computer architecture geek at heart it's like really really fun to see this community and this industry at this stage right now and when we and we felt that like at at hot chips and everybody I talked to it's like that that buzz is there it's amazing >> we want to go through the chips let's you know talk about your chip first what's what's most exciting CS4 came out some crazy numbers here 30x faster the GPUs but you know walk us through what's new with the new CS4.
3:53 >> Sure. So, you know, we we we we designed our next generation CS4 architecture with a new brand new system platform with the goal really to make wafer scale mainstream to make it to bring it to hypers scale and to solve a lot of the the the high density data center problems that you know we're facing every single day. So we've designed this modular platform that provides twice the amount of power to the wafer than we have in our previous generation. twice the amount of interconnect bandwidth, half the latency. and all of it is done at the system architecture level. And what you see here is the result is that by being able to provide significantly more power to the to the wafers by you know packing more density into the rack we're able to drive up the performance even further. you know we like to say that Cerebrris is with our current product already the undisputed leader in you know ultraast inference. and and what's amazing is this CS4 product will actually push that frontier even further by taking it another two times faster.
5:13 and this is a a perfect example right we here in this in this demo that we gave at hot chips. we're showing GPTOSS running at over 4,400 TPS which is just mind-blowing. It's like it's it it it almost feels like it's fake, right? and and and we think that this is going to really revolutionize the entire you know industry because you know not only are we able to you know continue to push the frontier on how fast these models can run it will enable all sorts of new applications you know user experiences start to become extremely different right what used to be batch and offline applications start to become real time And you know all of this you know growth in new agentic flows and agentic frameworks now also mean that you know you're you're sitting there waiting for these agents to go over and over and over in their agent agentic loop and all of a sudden if you're running your model at over 4,000 tokens per second. Now the you can do you know more agentic loops you can do more reasoning ultimately you get significantly more capable more intelligent agents. And so this is what really excites us about this. Not just the fact that you we show these really really big numbers which is also really cool. but but the fact that you can you know really start to do things that you can't really do otherwise and and and start to enable brand new capabilities. that that's that's what makes me so excited about this this reveal that we had yesterday.
6:51 >> Yeah. I think it's also very rewarding that you guys started on this journey the big chip journey. you know wafer scale everything everything that you've always said it just like people didn't take it that seriously until they had to and then now they're they're really really taking it super seriously right >> I I I do feel like as you know co-founder CTL the the architect of this whole thing how are you guys approaching the new generations post like AI boom like I imagine CS123 was a bit more sort of calm >> now like literally open AI is launching soup the the ultra fast mode with you guys and like it it is a matter of like I don't know like you're you're you're co-designing your model with the chip almost.
7:35 >> No, absolutely. I I feel like the main shift that has happened over the last few years for us is that you know the the first generation chip primarily was a technology demonstration, right? We had to demonstrate to the world to ourselves that you know you can actually build such a thing. and you know we we we told a lot of people that this is the future and most of them were just like you guys are insane this is not possible. and so you know at the at at the beginning it was really about you know figuring out the foundational technologies to make it possible. And what we're now seeing is now that we have in fact demonstrated that this can be built and it can have the kind of performance that you know that this type of architecture can have what are all the things that it can enable right and you you know you alluded to you know our recent ultraast launch with open AI that's a perfect example right we're running you know frontier level one of the most intelligent models right open's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds. And it's becoming this completely transformational thing when you can combine the most intelligent models with the speed that the hardware brings. And this transition from technology demonstration to you know solving real problems enabling new new real capabilities is this transition that we as a company are going through.
9:15 And how we're coping with that right now is we're really shifting our focus to ensuring that we can make this scale. We can bring up bring the capacity that we need. we can run the models that you know the users ultimately you know care about. and then we continue this flywheel by making sure that we continue to put out faster hardware and and keep pushing the frontier. And that's that's kind of you know how things have shifted for us. And it's I would say like you know unfolding in a way that is you know almost better than we could have imagined right it's this perfect confluence of like you know having the hardware with the capabilities demonstrated now matching with the capabilities of the models and the applications. and and so our goal right now is just you know scale this thing as fast as we can.
10:10 >> And so maybe we'll talk about yesterday. >> Yeah. >> previewing CS5. can you recap for people who are maybe not not yet caught up and you know maybe are new to the story for the first time as well. >> Yeah. No, no problem. So I mentioned that our CS4 system that we just launched is based off of a brand new RAS platform that we call the Nexus platform. And as I mentioned the what's special about this platform is the modularity. it has a a modular power supply in the front. It has a modular server in the back that we call the backpack. And this is a platform that enables that 2x performance that we just all saw the numbers for. But it's also the foundation for multiple generations of products. We've designed it from the ground up to support multiple generations of our products. And in particular, we designed it together with our next generation wafer chip that will be coming next year. And with that chip, we'll be pushing the performance even further. So, we just saw a 2x improvement this this year with CS4. We're going to push it even further with another 2x improvement in performance. And what this ultimately means is you'll be able to run, you know, medium-sized models like GPOSS or Gemma at speeds up to 10,000 TPS and even frontier level models like Kimmy or DeepSeek and GPT56 Soul up to 5,000 TPS. Completely gamechanging. all again enabled by this brand new platform that we've designed for multiple generations. And so, you know, this is the trajectory that we're on right now.
11:56 And and you know, now that the world sees the value of ultraast inference, our bet is that even this is just the beginning. >> What's the rollout process like? So, you have a deal with OpenAI over the next few years. >> when do we get access, you know? So, everyone wants ultra fast right now. Sure. There's Gemma OSS. When do we get the big Kimies, the Deep Seeks? when can we start to you know actually use >> that's a great question. So right now we are basically you know sold out of everything that we're building right and we are very strategically making sure that we're deploying every single megawatt in the most strategic way possible and in particular quite a bit of it is going to open AI right and we've been very public about this you they're our big biggest partner not just because of the the the commercial arrangement but also because of the fact that we have this co-design you know spirit which you guys have heard open talk about this a lot as well right which we believe will basically allow us to kind of continue this flywheel of not just making models you know faster but using that speed to make them more intelligent and so on so so today a lot of that capacity is going into openi and within openi they have very strategically decided to use quite a bit of it for themselves. so they're using it right now. That's correct. Right. So right now internally they're using it for a lot of really critical use cases where the speed really really matters. Like they're using it in like their incident response teams, right? When there's an outage in their service, for example, every single second, every single minute matters. and so they're getting a tremendous value out of that. they're using it in some of the most critical research applications where the extra reasoning is is enabling significantly you know more intelligent responses and so right now that's quite a bit of the capacity is going there and then as as we mentioned in in in our launch with OpenAI we are now also making it open to you know to enterprise customers who are you know who are who are able to again use it for some of the most demanding and you know the most high value applications over time openai and cerebrus we have committed to you know to bringing enough capacity to make ultra fast inference available to a much much wider audience and you know and our our CS4 announcement our CS5 announcement these are very much part of that commitment to continue to drive more faster tokens and you know basically more throughput to be able to satisfy you know all the various use cases out there >> by the way you know the way that you framed it I realized that they didn't have to expose it to enterprise customers they could have just kept it for themselves >> nobody knew that that's an interesting like business decision almost I don't know if you have any weighing on that that's more like a business analyst point of view >> I would say I would say that you know obviously I can't you know explicitly comment about their thinking right ultimately it's open AI's decision how they want I use this. But I think if you look at it from the outside, it kind of makes sense, right? I mean, their mission very much is to continue to, you know, push what's possible and by using it internally that helps them do that.
15:31 But they're also a business now, right? And there's, you know, talks about them going IPO and all this and so very much even externally, you can see they're balancing both of these, right? And I think we very much see see that playing out in our >> you know, in in the ultra fast space as well. I mean they're balancing a lot you know most recently their hot >> we got to talk about this as well we got to we got to go into them so what are we thinking prefill jalapeno decoded cerebras what's the >> what are your takes >> I think that's I think that's that's a very you know rational you know conclusion I I think of all of the hot chips announcements probably jalapeno was the most exciting but to me, but not maybe not for the same reason as everyone else.
16:20 >> I guess to to double click on that, you're the expert in chips, right? There's people see token per second. What do you see as interesting when you going to say I think that like they you know they pushed a lot on the performance and the fact that they're significantly better than you know better performance than than the GPU than Nvidia. But what I see is that they've built a significantly better GPU and that in its own right is is very is a is is a big achievement. Nvidia knows what they're doing, right? They they they own the market for a reason.
16:58 They're not dopes. and so to be able to come out of the gate and and build a significantly better GPU is a big achievement. But to me, the reason why Jalapeno is so exciting isn't even these all these paratos. It's really the design methodology behind it, right? They they very clearly took a very drastically different approach to building this chip, right? Having an AI first methodology enabled them to, you know, build the chip faster and achieve some of these very impressive results, right? And that is 100% the future of our industry. And it's not surprising to see that OpenAI is kind of leading the way here because this is you know very much their their MMO. But coming back to you know what you said about you know how we're going to use this right. I think that you know Jalapeno is is pushing the boundaries for both boundaries of what's possible for both throughput as well as latency. And for for me, even if I take the Cerebrra's hat off, I think that's awesome for the industry, right? This is going to lift all the tides. Everyone is going to benefit.
18:09 you know, they've they've also now been able to, you know, push the latency into regimes that that traditional GPUs can't can't hit. And and that's also great because again, we believe in speed. There's a lot of value there. And what's going to be really interesting and what I'm super excited about right and you know OpenAI is our biggest customer. So this is this is one of the things that you know I think is very strategically important to us is that when Jalapeno is available next year when our next generation CS5 is available together right we will enable a full fast inference portfolio right that is substantially different and better than what's already available today which is already substantially different than than your baseline GPU and then on top of that there's opportunity to integrate even further like pre-fill and decode disagregation for example or other >> for they didn't actually specialize jalapeno for right >> they did not specialize jalapeno for that but they specialized it for throughput right and so they get a tremendous amount of throughput and so just as a computer architect there's like so many different things you can start to do with that right and then we have you know the you know insane latency right like I mentioned up to 10,000 TPS in in CS5, you can start to imagine some some really really cool products that we can build together, right? And that that's what really excites me about having them as a partner. And we're also excited about collaborations on the AI tooling front because much of the the benefits that you know that that they're seeing from the AI tooling infrastructure everybody probably says this now but like we're obviously doing you know a lot with AI but we're also collaborating very closely with OpenAI right to use their tools to help us also continue to push what's possible in our chip design and our software and all that. So both of those together I feel like you know this is a really unbeatable combination.
20:11 >> I think one thing that people are talking about like you know I'm trying to get to the disagreements of the hot takes now. So people focusing on you haven't really mentioned like power and I I do think that something that seems to be a consistent theme is you know performance per watt rather than than tokens per second. any variation of of this theme or what are people sort of talking about offstage that you know is more contentious? Well, I think there's a there's a few things that, you know, in terms of, you know, some of the the more contentious things like I I would say one of the the themes that came out quite a bit in my discussions at the conference was around the the Grock announcement, right? You know, Grock, obviously, not surprising they have a new chip. In general, I think it's it's awesome that SRAMM designs are becoming, you know, more mainstream now.
21:03 you've >> been here the whole time. We >> we we've been talking about it for a long time. And it's amazing to see, you know, the the industry starting to embrace it, right? >> do they feel different post acquisition? I mean, you've been competing with them for >> a while. >> to first order, no. I think it's really awesome to see that, you know, SRAMM architectures are becoming, you know, more accessible. you know, even the biggest of the big guys here, Nvidia is embracing SRAMM design, acknowledging that, you know, that the traditional GPU designs really can't hit the ultra fast, you know, regimes. A lot of what was being discussed offstage was I mean the natural obvious questions is like why did they launch on a 30 billion parameter model? and you know, how come when Jensen spent so much time at GTC talking about attention FFN disagregation, there was no mention of that. And so I think there's, you know, I think that's that's pretty telling, right?
22:06 >> I mean, there's a separate Reuben, you know, strategy. >> Well, so there's there's there's Reuben, but you know, the LPX itself >> that that was supposed to be where it was. >> It was supposed to be Reuben LPX together, right? if if you guys recall, I mean, the Jensen spent like >> GTC >> GTC like half an hour explaining attention runs here and and you know and and the movies run here and so on. And I don't know if this is like a a hot take per per se, but you know, it it it's it's very suspicious that they're that their product that's in full production, they've only shown performance numbers on a non-disaggregated 31 billion parameter model, right? And I think to me what this this shows is that there's definitely some challenges in running on a nonwafer scale SRAMM design because there's just not enough memory in each of the chips, right? And I think that's you know that that's what's happening and and and we're seeing the evidence of that and in many ways I think it's very much validating kind of the design choice that that that we had right if you think about it to run a frontier level model of like let's say a few trillion parameters you need thousands and thousands of Grock LPUs just to hold the the weights.
23:29 >> Yeah. >> Right. And so when you start to think about it that way, it's like, well, is it surprising that the only performance numbers that they're showing are, you know, on 30B, right? >> so they're going to gradient this into this. >> Well, I I think I think it's I think it's the other way, right? I think what's what's going to end up happening is they're going to end up focusing on significantly smaller models. >> All right. you know, if you if you have that limitation in your architecture, then I I think that's that's what ends up happening, right?
23:58 whereas in our case you know we're running the world's largest models and in some ways kind of simple because you know one of our chips has you know order 100 times more memory than one of their chips right so you have two orders of magnitude difference in scale kind of for free right and so I think that's definitely one of the things that was a topic of discussion again independent of of cerebrus just kind of odd that you know you'd launch a brand new product on on performance numbers of such a small model. But when you peel it back a little bit, it kind of makes sense. I mean, I've been living in this space now for a long time. There's a reason why we needed the wafer scale integration to be able to aggregate enough SRAMM to be able to actually make it useful for large models.
24:44 >> Yeah. And look, the market's large, right? You have a different market than them. And you know that you you you clearly are the longest running incumbent now in this space. >> No, absolutely. I mean the the the market's large there's a lot of different opportunities for you know different different hardware to to play different roles. In fact, in general, you know, at Cerebrus, we we we believe very very strongly in a heterogeneous disagregated, you know, ecosystem, right? Not just pre-filled decode disagregation, but, you know, I think we're just at the beginning of what's possible in this space. And with the scale of of of these deployments and of these models, right, inference is no longer just like one workload. There's many many kind of subworkloads within it. And you know, you really want to use the right, you know, tool for the for the problem. You really want to use the right hardware for the problem. And so, absolutely, I think there's there's a there's a spot for all the different types of architectures out there. And, you know, we we've we've chosen to to target, you know, the frontier.
25:50 >> yeah. Frontier ultra fast. >> Exactly. Exactly. >> I we was going to go into some of the other companies that were, you know, top of town this year, >> but it gives me this gives me opportunity. I want to follow up on one thing which is how do you think about the classes of workloads right to me the ultra fast frontier workload is just like very clearly like one of the fastest growing segments of the entire infrance market right like the fact that like I want it and I cannot pay for enough for it I it might have been a mistake for openi to offer it even right like but it's good for anyway so just in your experience right you said you said you're strong believers in a heterogenous inference solution okay what are the buckets and how do how do you see it from your talking to your customers?
26:35 >> Well, so I I think there's probably two different views here. you know, one is the the product view and then the other is kind of the the the technical computer architect view, right? From the product view, it's actually just very simple is that bringing more speed opens up significant opportunities, applications, different use cases, different capabilities, more intelligent models, more intelligent agents. And so the further you can push that, the more and more that you can enable. In fact, there's probably all sorts of things that we can't even imagine that you can build, you know, when you're even faster than what we are calling ultra fast today, right? Which is why we keep pushing that. And then the other angle really just comes down to capacity, right? That's so it's like yes, I want the fastest, but then you need enough to actually be able to to to serve your use case, right? And so it really just is that simple, right? And then finding the right architecture and the right mix of architectures, right, to be able to provide that is is really the name of the game. Now, from the kind of computer architect's point of view, in a lot of ways, it's it's even a little bit simpler, right? Like I was having a conversation with somebody actually at hot chips about this and they were asking me well you know if you're doing disagregation like doesn't that mean you have to like partition the data center and you have to deploy a certain amount of this kind of hardware and you know different type of a certain amount of another type of hardware isn't that restrictive and I'm like I mean you know when we design >> modular it's modular and when we design a chip every single day we're deciding like am I going to use the silicon real estate for memory or if I'm going to use it for compute or if I'm going to use it for IO. and so from a computer architect standpoint, what I see, right, is there's significantly different parts of the workload, right? The simplest ones are, okay, there's prefill, you know, decode, but then you go one level deeper. It's like, well, what's actually happening during these phases? Oh, well, there's, you know, there there's loading the actual KV cache. There's doing the actual, you know, attention. There's spreading the experts across the various different hardwarees and balancing them.
28:45 All of these things can be now addressed with kind of more tailor made, you know, hardware solutions and and architectures. And at small scale, it doesn't really make sense to bring together a bunch of different hardwares just to solve like one problem. But at the scale we're all talking about, the hundreds of megawatts to gigawatts, the multi- gigawatt scale, it easily pays off. And then at that point, it's as if you're like thinking about the entire data center as if it's like one computer. It's like one chip that you're trying to figure out, okay, well, I want this amount of this capability so I can run, you know, attention fast. I want this amount of capability so I can run prefill really, really efficiently. And being able to piece all these things together is like the computer architect's like dream to have all of these tools in our toolbox, right? So that's how I how I think you know where you get the value from this heterogeneous disagregated ecosystem.
29:41 >> How much of that plays into model design, model architecture? So working with open AI, you get to hopefully see what's going on on models like there was the whole were thinging models. >> It's not enough. There's also voice diffusion. what what if what if there's more interactive models like the thinking machine stuff? >> Absolutely. And I think as as we start to get into more interactive use cases that we enable through ultraast we're even seeing a lot of this right so if you're trying to interactively for example do graphic design now you mentioned diffusion now you know image generation or video generation might be kind of now inline in your you know creative like inflow interactive use case for with with the thing and so there's so many of these different use cases that are coming and what I You know, you mentioned about code design with opening eye and so on. what what I think is the most untapped opportunity right now frankly for Cerebras but frankly for the entire nonvidia environment right and the non-envir sorry community is that you know we're all running models that were designed for Nvidia GPUs right and and and it's not just Nvidia GPUs but like usually they're designed for like one particular Nvidia via GPU, right? Like, okay, this thing was designed to run on B200, GB200, MBL72, right? And you know, and so as a reverse, for example, here we were showing that you're running the model like 14 times faster already, but it was a model that was designed for a completely different architecture. And so if you start to then open up the possibility of adjusting that model architecture even slightly, you can get massive gains. And if you start to do even more code design, I think, you know, the the the opportunities are almost limitless. And that's really what excites me a lot about, you know, about working closely with customers and partners like OpenAI.
31:45 >> Yeah. And we we don't have time for this because we want to move on, but the people under are really sleeping on AI codegen for kernels. which makes actually which is very good for you guys. >> Yeah. No, absolutely. I'm I'm a huge believer in that for sure. I think before getting into etched and super spicy chips, any takes on AMD, Nvidia, what they're doing? Could they be doing anything different? Any hot takes there? >> I think that Nvidia is doing exactly what we all expected Nvidia to be doing, right? like I said, they're no dopes.
32:17 They know exactly what they're doing and they're continuing to push through and in many ways, it's absolutely the right thing, right? Cuz we need more tokens. We need them cheaper. That's a that's a reality. However, you know, what used to be considered fast at like 1,00 or or 200 tokens per second is quickly becoming the new batch mode, right? Quickly becoming the new like overnight is true and and so I think that you know there's a place for that. It's great for prompt processing. It's great for very very parallel workloads.
32:54 >> but my evals take 20 hours. I can hit a button and run it in two hours. >> Exactly. Exactly right. But but quickly, you know, being able to do prompt processing isn't isn't quite enough, right? That's a big part of the market, but that's not quite enough. And so I think all the the the traditional architectures I feel like are very much all on that treadmill, right? In a good way, right? Not not necessarily a bad way. It's in a good way, right? And so I AMD, Tranium, in many ways TPU, like all of these in my mind are all trying to build a better Reuben, right? And there's a huge amount of value in that, but I think that there's also a lot of value in trying to push, you know, the boundaries in other vectors as well, which is obviously, you know, what we're trying to do here at Crisis.
33:44 >> Okay. spicy spicy chips. There's this page that you have on your website >> that this company Etch just doesn't have. No, no numbers, no benchmarks, any any takes. That's that's their one criticism. But that's etched. >> Well, so they have said very little about what they're doing. but what they have said has changed also a lot compared to what they originally said. >> which they acknowledge >> which they acknowledge and and you know the the environment is changing a lot and so that makes sense right they're going to pivot they're going to try to find their their space I think my reaction when I see pictures like this is that it's very impressive graphics design but I also don't see them building anything beyond just or trying to build something better than just you know a traditional GPU right you know they've made claims about having a distributed storage that is somehow faster because it's distributed, which I don't quite understand. You know, it' be very helpful if they can explain those kind of things. I think, you know, I I it's hard to see where the the differentiation is that they're trying to push. I think they had they had a rack or something at Hot Ships, I think. And so clearly they're they're building something, but it feels like it's a little bit, you know, far off for now.
35:14 >> Yeah. But at the same time, like $20 billion now, it's like, you know, not not a pushover. I I think maybe the the steel man for this is like what interesting direction would you want to see them pursue, right? that in in the best case >> in the best case I would love to see them or others frankly push the boundaries of more innovative solutions outside of just the chip right you know we I I really believe that the next kind of frontier of of taking AI compute to the next level is all about better ways to integrate outside the chip right >> the rack node, >> the rack, the node, the package, the >> the the interconnect, >> the factory, >> the the the factory. they may be doing something like that. It's hard to tell, right? but I think that this is not just a a point about edged. I think for the entire industry, right?
36:16 even if I I take my cerebrous hat off for a second, right? What excites me the most is our ideas and and creative thinking outside of the single chip. We all know that the applications, the problems, the models are now so large that you can't just do anything on a single chip. So everything comes down to the integration. Everything comes down to how do you bring together more compute? How do you bring together more memory? How do you bring together more interconnect, right? And so, you know, >> scheduling, pipelining, all those algorithms, >> all of those things, right? And and so, you know, of the of the, you know, the reveals yesterday, I think probably one of the the the the more interesting ones to me, I mean, not not it wasn't new news, but was was Dmatrix, right? the fact that they're really leaning into into into DRAM, you know, 3D DRAM packaging and like this is the kind of stuff that I think we as an industry need to be doing much much more of.
37:18 >> I was going to say like your backpack stuff reminds me of what the memory people are doing. Is that analogy that >> there there's a there's a little bit of of that, right? Because you know what we're doing with our backpack and and power delivery is very much you know a a very compact 3D package right where we can bring power directly to to the face of our our chip of our wafer right and in in many ways like you know the the DRAM 3D stacking integration is doing something similar but not with power but now with with memory right yesterday we also announced that you know But we also started two years ago our DRAM stacking program for the very same same reason, right? And cuz I think that figuring out ways to to do this kind of integration is the key to to unlock the next major step forward, right? And so I don't know exactly what Grock is doing or sorry what Etched is doing underneath the covers but you know when you ask like what is the stuff that excites me the most about what we're seeing in the industry it's really more creative you know integration more creative packaging more creative outside of the chip thinking.
38:31 >> Yeah. let's leave some hints for people other than Datrix. a couple other names that stood out to you just >> along these lines. I I think you know Samsung has some discussions around their ZHBM right >> all memory guys >> all I mean look in the end there's only three main things to to building a you know a an AI chip right there's the compute there's the interconnect >> and then there's memory >> right and so and and more and more importantly as we all know for at least in the low latency space the memory is the key right the memory bandwidth is the key >> and power and cooling not as major still.
39:12 >> No, absolutely. I mean, you just did a big redesign. >> Those are those are what's powering all of this, right? >> Yeah. I'm just saying, you know, like physical limits, >> but that's actually a really really good point, right? I think that when we when we originally started with wafer scale, most people thought the biggest challenges were, okay, how do you connect all of these dye together? How are you going to yield this thing? And those were absolutely kind of fundamental challenges that we had to solve. But what turned out to be, you know, the biggest enablers was ultimately how do you power it? How do you cool it, right? How do you build it reliably at scale? And I think that a lot of the the 3D DRAM technologies are going to go through the same thing.
40:00 And it's, you know, one thing to draw, you know, a PowerPoint slide with a DRAM and and and logic and then it's another to actually make it work to figure out how you actually going to power it, how you actually cool it. And that's why a lot of what we've been doing is, you know, is trying to leverage and use all of the expertise that we've built up there and put it into our DM stacking approach. And so, you know, we already have solved yield at scale, for example. But we've already solved how you can actually package in a three-dimensional way and and that's you know problems that Samsung that Dmatrix and and everybody else are also going to have to solve over this time. Thank you for those comments. last topic because we got to go. US supply chain semiconductor supply chain stuff you know like China is slowly becoming completely independent of us. vision GLM like really literally just like they're they're bragging >> Huawei and what what have you >> hyped outside of knowing what it was on >> just model is good doing good in you know open router came out not even running on US chips >> so the TDR is are people talking about it what's what what we doing >> I mean I think that absolutely I mean of course people are talking about it right because like the the the we're in a very difficult you know situation right now the open- source model market is 100% Chinese right >> not 100% but >> almost okay 95% right most of the most of the big models most of the big open models that are you know qu high quality are coming from the Chinese labs in some ways it's great that sharing is happening and and you know the the global community is benefiting but al obviously that's a very strategically challenging place to be to have such dependence on on them on the model.
41:54 separately as is evidenced by this is like you know we have independently the all of the infrastructure the hardware infrastructure slowly being built up in the background to support all these Chinese models right I think it's absolutely a a problem that globally as well as as you know the the the US and as a as a national interest that we have to continue to push the boundaries so that you know we and continue to to compete and and continue to be ahead.
42:26 it's not easy. These guys know what they're doing. but it's something that I believe is not solved by, you know, one company or or one fab, but it's needs government level natural national interest level kind of initiative to be able to make this happen. >> so I I imagine that should take place as at Hot Ships. It's like the secret room of all you guys. You're leaders of our industry, >> right? that you would be the guys to >> I mean we we have absolutely been been pushing for this and we've been very supportive of of you know any initiatives that will you know continue the US dominance in this space 100%.
43:04 >> Okay. you got to go but you've been very generous with your time. Congrats on all your success. I mean enormous you IPOed. >> Yeah. Yeah. Yeah. Sometimes I still have to pinch myself to remind me that we IPOed and it was like the largest like you know semiconductor IPO in history and it's like >> still early. >> That was not bad. That was not bad. >> no congrats and look forward to meeting up with you in the future year.
43:27 >> We got to ask you know give us ultra fast. Give us bigger give us more. >> No no I mean give give the people ultra fast too. Like >> we but we but we can probably get you in line at at open eye. Oo, >> I hope you didn't bias OpenAI too much to, you know, just keep it internally. No, people need it. We need better models. >> It's OpenAI's decisions. We're we're just providing the infrastructure.
43:47 >> Hey, but give us give us GLM. Give us Dec. >> Just give us our own rack and then we'll run it. okay. Thank you again for coming. Really really appreciate it. Yeah.
Summary
- Traditional benchmarks for speed (1,000-200 tokens per second) are becoming outdated as new architectures push for ultra-fast processing.
- Cerebras' CS4 architecture aims to revolutionize AI inference with a modular platform that doubles power and bandwidth while halving latency.
- The CS4 can run models like GPTOSS at over 4,400 tokens per second, enabling real-time applications and more intelligent agents.
- The transition from technology demonstration to practical application is crucial for Cerebras, focusing on scaling and meeting user demands.
- OpenAI is a key partner, utilizing Cerebras' technology for critical applications, indicating a strategic collaboration to enhance model performance.
- The industry is seeing a shift towards heterogeneous architectures, where different hardware types are optimized for specific workloads.
- The importance of memory bandwidth and power management is emphasized as critical factors in AI chip performance.
- The competitive landscape includes companies like Nvidia and emerging players like Etched, with discussions on the need for innovative integration beyond traditional chip design.
Questions Answered
How is the perception of processing speed changing in AI?
The perception of what constitutes fast processing is evolving, with traditional speeds of 100 to 200 tokens per second now seen as standard. This shift highlights the growing demand for faster prompt processing and parallel workloads, although it is acknowledged that this alone is not sufficient for the market's needs.
What is the significance of combining advanced models with faster hardware?
The integration of intelligent models with high-speed hardware is transforming the industry, enabling the resolution of real problems and the development of new capabilities. The focus is on scaling this integration to meet user demands and pushing the boundaries of performance.
How does Jalapeno influence the AI hardware landscape?
Jalapeno is setting new standards for throughput and latency, which benefits the entire industry. Its capabilities will allow for a fast inference portfolio that surpasses existing technologies, emphasizing the importance of speed in AI applications.
What are the key considerations for heterogeneous inference solutions?
Heterogeneous inference solutions require a balance between speed and capacity. The right architecture mix is essential to serve diverse use cases effectively, and the technical perspective emphasizes the need for flexibility in hardware deployment.
What future directions should AI hardware development take?
The future of AI compute lies in innovative integration beyond just chip design. This includes advancements in rack, node, and interconnect technologies, as well as optimizing algorithms for scheduling and pipelining to handle larger models and applications.