Transcript
0:00 Hey everyone, welcome to this month's AI performance engineering meetup. And we've got a lot of folks and usually people trickle in here. It's top of the hour, but today we have awesome awesome awesome talk actually is this this is April. Today is 420 as some of you that may mean something to you. Hopefully you're still you know, in the right headspace to absorb this content. All three of us hosts are based in San Francisco. So it's going to be interesting when we step outside there might be big big clouds of smoke.
0:35 Special day. Okay. Um quick updates from our side. We continue to roll through. I think we're on almost 10 years of this meetup or we're actually past 10 years and we've been focused exclusively these days on AI performance engineering. This comes in a in a wide array of different things like yeah, everywhere from storage optimizations which I think we're going to have on next month. We're supposed to have them on today, but they they had to focus on some production issues.
1:08 So we'll be getting insights at all the layers of the stack and I first saw Natalie speak at a recent Swix conference out in New York that was focused on coding agents and it was it was the first talk I've seen out of the the Swix AI engineer conference that actually went into hardware and I was so excited and I I contacted Natalie and she's been very busy raising money and everything. So welcome to Natalie.
1:38 We're going to hand it over to her in a minute and she's then going to explain her background how she got here. She's going to talk about the recent fund raise which is always exciting. So yeah, we love startups on this. Auntie, do you have any updates as well too? Maybe from the Amazon AGI side. What's going on there? Yeah, just a quick hello from my side as well. And you bought for those of you who don't know me yet. I'm helping Chris here to co-host this meetup and if I'm not doing that I am a member of technical staff at the Amazon AGI lab here in San Francisco and for folks who are interested we're building agents that can work like humans. So becoming digital teammates and most recently we just actually launched a bunch of like cool developer tools. We have if folks have heard of Nova act it's kind of a UI browser automation tool. We launched MCP server skills Kira powers. So a whole ton of cool stuff.
2:35 If anyone is interested yeah, send me a message after. Awesome. Let's get right into it. Natalie, I'm going to stop my share and the floor is yours. Awesome. Thanks to you both and just checking that my screen is visible. It is visible. Amazing. Okay, it wouldn't be a zoom meeting without that question, would it? Yeah, someone's likely going to ask if you can maximize it or close your toolbars or bookmark bars, but do whatever you need to do. Is that better?
3:08 >> Yeah. Yep. That looks great. Thanks Natalie. No problem. Okay, yeah, thanks everyone. Thanks Chris Natalie for having me. It's awesome to be on here. Yeah, as Chris mentioned it's been a crazy few months. I think not just for us at Gimlet, but for everyone in AI. So just a quick background about me. So my name's Natalie. I'm a co-founder of Gimlet Labs. We'll get more into what Gimlet Labs is doing, but what I'm here to talk to you about is why we think that the future of agentic inference requires heterogeneous hardware and talk more about the kind of hardware needs of agentic inference specifically. So Gimlet we have raised we did a recent round of our series A with Menlo Ventures. We just closed 80 million dollars.
3:57 We are an inference cloud. We work with very large scale inference workloads. We're working with a major frontier lab and our goal is to make the fastest most efficient inference possible to serve the needs of what we see as not to sound cheesy for a second, but the future of software which is agents. So just a little for the agenda here for this talk. You know, I think first we want to start with why why do we need heterogeneous hardware systems for agentic inference?
4:32 Then follow with if you accept that premise, what are the technical problems that we need to solve to deploy these workloads across heterogeneous hardware and then finish with what results do these systems provide? Like you know, it's a lot of work to do. So what are the results that we get for it? Just to start with some industry context. Has anyone else noticed this that as Bloomberg put it, you don't hear much about the AI overbuild anymore.
5:04 I think that it's pretty clear to a lot of us at least that we are in a period of significant compute scarcity and that we may have if anything under invested on the data centers and capex that we need to serve the agentic workloads that are and continue to be exploding. And part of the issue here is that inference is growing across almost every dimension right now. And this is leading to significant pressure on our infrastructure and our capacity.
5:41 So if you think back to the advent of generative AI initially it used to be a simple chat model. You would chat with it. It would give a response. But over time these workloads have evolved and now we see them doing more complex things like coding agents, tool calls, voice agents. These agents have become multi-step. They do more than one thing when they are computing their response and so there's more stages and more parts of the workload that are happening.
6:12 And what we're also starting to see and I think we'll start to see even more of this this year is the full application which may actually be multi-agent. So it's not just me talking to my agent who's doing things, but it's me talking to an agent who is collaborating with other agents in the background doing lots of different tasks and this could be over a long time horizon. And so we see that the complexity size and quantity of these workloads is just exploding.
6:45 So these workloads they demand a lot of compute and as mentioned are growing. If you look at MUTTR or I don't know if people call it meter. I've never actually said it out loud just read their stuff. If you look at the time horizon of what these agents are capable of operating over. It's gone from a few seconds all the way to 14 hours in the most recent report. These agents use 10 to 30 x more tokens than chat models. We've seen context windows explode. Frontier models are in the trillions of parameters. And something's happening that people have been predicting for a long time which is that inference is actually going to overtake training in terms of the share of the workload.
7:28 So for a long time training has been the dominant workload and now in 2026 Deloitte predicts that this will shift to inference. So to meet this demand we're investing ridiculous amounts of money in this capacity, but are we getting that expected return on that investment? I see a question. I'm not sure how we typically, but happy to address Mario. Feel free to go on Natalie. We'll get to the questions at the end.
8:01 >> later. I'm asking people to to put questions in the chat. Okay. Okay, great. Sounds good. Okay. Um So yeah, fun fact the estimated spend on AI capex this year which by the way, I think I need to adjust this cuz it has grown is more than what was spent on the US highway system. And yes, that is inflation adjusted already. So we're spending ridiculous amounts of capital on this, but at the same time what we see is that for inference especially the utilization of these GPU fleets is not where you would hope it would be given the amount of capital being poured in.
8:37 So I think there's been two primary solutions to this. The first one is Okay, we have huge workloads huge workload needs. It's not enough. Let's deploy every single chip we've got for inference to solve this problem. And if you name a major vendor and you name a lab, they probably have a partnership. We're actually at the point now where CPUs are all sold out because they're a critical part of agentic inference. And I don't know if anyone's tried to rent an H100 on AWS lately. Auntie, I don't know if you can help me with that actually.
9:11 But you know, everything is sold out like crazy. And it's in the gigawatts. Like I don't know if people realize like how much a gigawatt is, but it's a lot. It's billions and billions of dollars of spend on the hardware. The other thing that people are doing is investing in specialized silicon. So it's almost like all the time you see a cool new chip company making something that is specialized for inference or for transformers specifically. And it's exciting because we're seeing like what we think of as like a Cambrian explosion in the hardware right now to serve this workload because this is like a really big deal and completely shifting the way software is.
9:52 So there's a lot of great companies here. Obviously very notably was Nvidia's acquisition of Groq's assets for $20 billion. This was an amazing turning point in the space because up until then it was unclear what was going to happen with the SRAM centric stuff. And now with Nvidia buying Groq's assets and incorporating it into LPX, I think everyone sees that this is now something that even Nvidia is looking at.
10:22 So, I think the question that kind of grounds this is is it enough? So, GPUs have been dominant for a long time. Is the solution going to be that they're going to be replaced by custom chips for inference? Is the solution going to be that we just keep throwing all the compute we have at the problem? Like what is going to happen here? So, this is sort of where I'm going to pause and just give a little bit more background on Gimlet.
10:48 So, Gimlet, what we're building is an inference cloud built for agents. So, the idea is that developers can easily deploy not just models, but entire agents through our API and scale it on our cloud. Our special sauce is that and I'll give away the you know bit a little bit, but our special sauce is that we do all of this on a multi-silicon multi-hardware inference serving stack where we automatically optimize the workload to use the right hardware for each subtask and allocate the workload and scale each of these phases independently. So, you might say, "Okay, I'm going to take this piece of the workload and run it on this type of architecture and this piece of the workload and run it on this type of architecture."
11:41 And we do this because it's really really efficient. And all of this actually started, just to get to the kind of origin of the company, we were doing research at Stanford into efficient AI on different types of hardware. And I would say that that kind of research aspect of the company has been part of our DNA from the beginning. And that's just because as we all kind of see these days, these systems are being invented and built across our industry and deployed to production before they even make it to archive. We really think research and engineering in this space have really merged. And so, it's important to continuously be trying to discover new techniques to solve problems that haven't been solved before.
12:28 So, what we what we did was we wanted to kind of take a step back when we were looking at this and say, "Okay, from first principles, what does an agentic inference workload actually need from a hardware perspective?" And so, if you kind of break it down, inference isn't just one thing, right? It has different phases. It has LLMs, but it also has tool calls. It might incorporate other types of models, data processing. You have to handle the KV cache. These are just a few example stages of inference. The list would be like many many slides, but what we see when we kind of visualize it is that this is not a uniform workload. This is not a workload where every part of it has the same needs from the underlying compute. And this actually gives us a clue into why the utilization isn't where we would hope it would be for these homogeneous fleet serving inference today. And that's because each of these parts of the workload actually need something different. And when you look across your workload, there's no one chip that is going to serve all of these optimally.
13:41 And these needs, like I gave like a view of like pretty high level stages of inference, but these needs get even more granular if you dig in. There was a great paper that you can check out kind of trying to characterize this where basically these authors were trying to characterize the different parts of inference and what their compute needs are on real hardware. And if you look if you break down individual parts of the inference workload and you even like break it down by prefill versus decode, you see just dramatically different usage of DRAM, dramatically different cache hit rates, and you know, fundamentally different profiles on roofline charts for not just prefill versus decode, but actually individual portions of those workloads within prefill and decode. So, every single point here actually would theoretically have a different piece of hardware that was optimized for it. And those are different than each other.
14:45 And with these different types of hardware, you see like significant differences in the programming model and the specialization. So, take two examples that are very different, an Nvidia GPU versus a Cerebras chip. Nvidia GPUs are cache based with a memory hierarchy. So, when you're programming an Nvidia GPU or any GPU, you're always trying to think, "How can I write this workload in such a way to maximize data reuse and really use my cache to the best of my ability?" And so, you're thinking about it in those terms. Whereas with something like Cerebras, it's a more spatial layout where you have to plan out this part of the memory lives on this part of the chip, which is connected to this part of the chip. So, it's like a very different model of programming and it's a very different model of optimization. And these chips are good at different things. And when you look at what the actual like footprint of the chip area is, you see that Nvidia devotes that B200 devotes more of its kind of area to compute and Cerebras on the right devotes more of its area to memory.
15:51 And neither one of these is right or wrong. They're just different. But with their differences, they offer different characteristics. So, we see as mentioned, Nvidia's B200 has more memory or has more service area devoted to compute and then Cerebras to memory. So, like SRAM centric chips like Cerebras and Groq and D Matrix and folks like that, they're making a very conscious choice, which is they're saying, "I want really fast memory access and I'm going to put more of my die area to SRAM even at the cost of other things that I could put there, even at the cost of cost and other things that also affect the overall chip. But I'm doing this because I really want to optimize for accessing more data really quickly." Whereas with a GPU, you are saying, "I'm going to put as much data as possible in the high bandwidth memory on the like just off the chip and I'm going to be limited by the perimeter area getting in, but overall this is going to be optimal for my workloads and I'm going to have very high compute density.
16:56 So, instead of saying like one of these is correct or incorrect, what we asked is what if we actually tried to leverage all of these together within individual workloads or individual models to get better performance? And there's different ways of splitting the workload. So, I'm just going to go through three different types of disaggregation that we typically see. So, the classic like the the kind of OG of disaggregation is prefill decode disagg. So, this is where you would say take all of the reading of the input context and the production of the first token and put it on one set of hardware.
17:36 And then taking decode, which produces each subsequent token in an auto progressive loop, and put it on a different set of hardware. And one thing that's cool about this scheme is that even when you use the same hardware for both, it's actually better performant than aggregating them both together from a throughput perspective. And this can get even more dramatic when you use hardware specialized for each one instead. Another common way of disaggregating is doing disaggregation of speculative decoders. So, in that case, if you have a speculative decoder, which is essentially making guesses about what token should be with a smaller model or a smaller architecture, you could basically take that piece of the workload and run it on a completely separate chip than your prefill or your verified pool, which is responsible for verifying the output of the speculative decoder.
18:28 And we'll dive a little bit more into speculative decoding later, but this is like another way of splitting the workload across different hardware. Another one that people are starting to talk about more and implement more is attention FFN disaggregation. So, obviously with attention, it's very memory intensive. You're doing a lot of looking across different memory. And then the FFN is like basically like a very compute bound operation. So, if we took these two pieces and run them one on a very memory heavy hardware and then the other one on a very compute heavy hardware, like we can get better performance by doing this.
19:11 And we do this because this pushes the frontier of like inference performance. And this is the thing that Jensen talks about a lot and I think it's a good framing. One thing is that when you look at different inference setups, you basically are trading off two things against each other. The first one is the throughput and it's usually normalized by something like per dollar, per water, per GPU. So, the total throughput of your setup, how many tokens it can produce.
19:42 And then the other axis is interactivity, which is if I'm a user interacting with that system, how many tokens am I personally seeing per second? And you would think at first glance like, oh, throughput, more tokens, wouldn't that mean I as the user get the tokens faster? But it's actually not true because producing lots and lots of tokens is a different problem than producing individual tokens very quickly. And so there's this constant tension between these two things and each hardware architecture will basically have a different Pareto frontier for this trade-off. So you could say something like a GPU might be the blue line where it's really amazing at doing very, very high throughput per watt. It's optimized to do that. That's its sweet spot. But the purple, something maybe more like an Astramcentric chip, it isn't actually as good at the total throughput per watt, but what it does give you is very, very good interactivity behavior. So instead of saying let's take one of these or the other, what we can do is leverage them both together in the same workload and that will produce something like the green line. And I'll show a more um like I'll show another real Pareto frontier later on in an example.
20:56 But the idea here is that we don't have to choose. We don't have to pick. We can leverage each of them where they most make sense. But in order to do this, we have to build an orchestration system that actually knows how to break up the work. And that's what we're building at Gimlet as part of our inference cloud. So with Gimlet, when a request comes in to us to be routed to an agent, we split that request up. We actually have already split the work up across the different types of hardware that are optimal for that particular workload. That can include GPUs, specialized accelerators, and CPUs, and other types of chips.
21:37 And we basically say, okay, this portion of the workload's going to run really well on a GPU, so let's run that on a GPU, and so on on the other types of chips as well. And the result is that we get much better throughput for the same latency or much better latency for the same throughput depending on what you care about than a homogeneous stack. But one thing that people always ask is, okay, cool, I buy that in theory, isolated parts of the workload are going to perform better on chips that are optimal for them. But what about the overhead of transmitting data between all of these systems? And that's obviously something you have to model because I don't care about the isolated performance of individual components of my workload. I care about the performance of the end-to-end request.
22:23 So in order to do this with optimal performance, you actually need to take all of these different hardware platforms and figure out how to connect them together in a co-located data center with a high-speed fabric. And so that's why Gimlet's a little bit unique where we're not just leveraging, you know, hardware that already exists out in clouds, we're actually composing new types of systems together because we have to. We have to take chips that have never been connected to each other before and connect them in our data centers because that's what gives the best performance for running a workload across both of them.
23:03 So what are the questions that need to be answered to solve this problem? So one thing I'll say is, you know, it's hard to cover all of this in a short time, but we have a paper um on archive that you can check out called Efficient and Scalable Agentic AI with Heterogeneous Systems that uh please check it out if you're looking for more details, but we'll try to kind of walk through some of the high level at least.
23:27 But that paper gives more information about, you know, kind of the architecture and like how we solve the problem. But any heterogeneous inference stack needs to address these three questions. The first one is, how do I know how to split the workload and map different tasks to different hardware? So like what's my way of knowing how to take that thing and split it up because it might not be as simple as run model A on hardware A and model B on hardware B. We might actually want to decompose model A into multiple parts and maybe even merge part of it with model B depending on what the hardware needs are.
24:04 Then once you've solved that problem, you need to figure out how to lower that slice of the workload for each slice in the workload to the hardware that we've selected for it and optimize it for that platform. And so for that, you know, what we do at least is we work really closely with this with the vendors that we partner with and we work really closely with the software stacks that they have and try to target the work to those low-level frameworks that already exist. So we're not trying to like make some unified software stack for every piece of hardware, we're more trying to take, hey, I have this attention function and I need to run it on this chip. What libraries or what runtimes or what frameworks can I use on that chip to do that well?
24:46 This is also an area that we look at kernel generation, which is a really hot topic in this space. How do I generate optimized kernels for different types of hardware uh using AI? And then the last question is, once this work is compiled, lowered, and scheduled on all of the target hardware, how to dynamically route the work between all of this hardware with minimal overhead? How do I make sure I'm not doing too many hops? How do I make sure that I'm leveraging RDMA and things like that? We run into all kinds of interesting things here like different vendors' hardware don't agree on the KV cache format. So we have to do things like translate it as it's being sent over so that the other hardware gets the format that it expects. So there's a lot of things that come out out of trying to connect all this hardware together that has never been connected before.
25:43 Um just a bit about the kind of orchestration system and the architecture. So um basically like the way that our system works is it's very Kubernetes-based, so all of uh kind of our stack will deploy across any Kubernetes cluster and there's kind of two phases. There's planning, scheduling, and allocating the workload and then there's the requests and serving that workload as it comes in. So with the planning and scheduling, it's making decisions like, okay, I've ingested this like let's say a PyTorch workload. How do I split it up? What hardware do I have? What hardware should I run it on?
26:21 Then once it's allocated, then we have, you know, the ability to serve that workload. And so we have, you know, obviously like an API server, a load balancer, cache-aware routing, and then we have a runtime. And that runtime is running on all of the different hardware platforms within the data center. And so each one of those will be able to execute subgraphs of individual models or pieces of the workload, you know, manage memory, communicate with other uh you know, devices, and all of this is connected together through high-speed Ethernet.
27:00 And um we expose a lot of metrics so that we can monitor how it's performing. Actually, most of us came from um a uh most of the like founding team came from an observability company in the past. So we really, really believe in building out very robust observability systems so you can monitor performance and make sure it's tuned appropriately. So I just want to close with giving a little preview of uh some of the results that we're talking about publicly. A lot of our customers are very secretive um but um for this one we did like a kind of public uh case study of this particular example just to give you a flavor of the upside of doing this, right? It's a lot of work to run models and workloads across different types of hardware at once and we only do it because the benefit is so great.
27:52 So just to get into this particular example, a quick primer on speculative decoding. Um I covered it a little bit before, but when you're doing traditional inference, as mentioned, you'll have the prefill phase, which injects the question or context, and it will basically do the processing of all the input context and produce the first output token. And tokens are not exactly the same thing as words, but it's not like a bad unit of comparison. So in the top chart we have like the prefill stage outputting the word the.
28:24 And then this goes into the decode loop and each decode pass generates another new output token. So one of the problems with this setup is that the decode can't actually run until it gets like the new last input token. So we're always waiting on the generation of the last token before we can start work on the new one. So one of the optimizations that have been widely adopted in this space is speculative decoding where we're basically saying, hey, what if okay, we know that decode is inherently auto-regressive, but what if we tried to shrink that work by using, for example, one way of doing this is using a smaller model for it. So instead of doing a full decode, we'll do what's called a speculative decoding pass where I have like a smaller model potentially making guesses about what that next token should be. And the idea here isn't that it's always right, it's just that it's right enough of the time to be useful. And when it's made its guesses, we can then go verify multiple guesses at once on the main model in batch. And this is more efficient because instead of doing one token at a time, we can actually process multiple guess tokens at once and then figure out which ones are correct and we can keep and which ones need to be recomputed.
29:46 And so this is basically a way of trying to shift the workload to be more compute bound to be a better fit for high compute hardware like GPUs. So, one of our partners is D Matrix and D Matrix makes an SRAM centric accelerator and it has actually 2 GB of SRAM per chip. And so this is like a very very large amount of SRAM compared to traditional hardware and that means it's really really fast at memory bound stages of inference. So, when we were initially working with the D Matrix team, we thought, "Hey, what if we took this hardware and actually ran the entire speculative decoder on it?
30:31 Because even though it's a smaller amount of work, it's still a very memory bound operation. This could enable incredibly fast speculative decoding." So, we worked with them to do just that. So, we basically compared three different systems. So, the first system is a traditional prefill decode disaggregation GPU to GPU. The second system was a speculative decoding setup. Uh so, the model that we used for the main model was GPTOSS 120B and we used a 1.6 billion parameter speculative decode model.
31:12 So, in the second setup configuration two, we said, "Okay, what if we did the uh like GPU setup for all three of these pools, prefills on GPU, speculative decodes on GPU, and verifies on GPU? So, it's still disaggregated, but we're using homogeneous hardware for all three." And then the last one that we compared with was the same type of disaggregation as configuration two, but instead of using GPUs for the speculative decode part, we used the D Matrix Corsair for the speculative decoding part and maintained GPUs for the other two parts of the workload, which are very compute heavy and very good fit for GPUs.
31:56 And so, these are some of the results. Um I can decode this chart really quick. So, this is the same kind of Pareto that I showed earlier where we have throughput per kilowatt, so tokens per second per kilowatt in the Y axis and then we have interactivity, so tokens per second per user in the X axis. So, we show three different configurations here like the three I just showed. Green is pure prefill decode disaggregation.
32:27 Blue is the speculative decoding setup on the homogeneous GPU deployment. So, we can see that that does provide a significant improvement over the green. And then the red is the speculative decoding setup using GPUs for prefill and verify and D Matrix Corsair for decode. And we can see a really really dramatic change here. So, going from just PD disagg alone to spec decode on GPUs gives you somewhere between like a two to four x speed up depending on where you draw the lines.
33:04 But you can actually get an equivalent speed up by switching out the speculative decoder to the very very memory efficient Corsair. And so this was incredibly exciting for us because this is a really common type of workload and it just shows how much more performance we can get out of the investments in power and hardware that we're making for agentic inference. So, if you break down where the time in each individual request is being spent, I think this is also helpful to see the user perspective.
33:40 So, in the first setup, you can see the huge huge cost of the decode house for this workload. So, the workload that we were looking at was 8,000 input tokens and 1,000 output tokens and you can see we spend the vast majority of the time in the prefill decode disagg on the decode phase, this auto regressive loop that's very memory bound. Then we can see a major step up for the speculative decoding setup GPU only. This shows the power of speculative decoders. They do take time to train, but they can offer very significant performance benefits if you can, you know, make them well.
34:18 And in that setup, we see, you know, nice shrinking of the work that is dedicated to decode and then you add on a little task to verify the results in the GPU afterwards. And then the last one is another time shrink that we see for the GPU plus Corsair setup. We can see just how small that blue section has gotten and now most of the time is actually being spent on the other stages and so we're using the hardware much more optimally here.
34:50 And one thing that's pretty interesting too for this is that one problem today with speculative decoding is that you are sometimes limited to shorter sequences because of the fact that it's still kind of expensive time wise and resource wise to produce each of those auto regressive tokens even though the speculative decoder is smaller. But what we see with running that part of the workload on the Corsair is that it's so fast and cheap to generate those tokens that we can actually speculate way more tokens at once. So, instead of only speculating five tokens at a time, we can speculate 20 tokens at a time and even if a bunch of them are wrong, it didn't really cost us that much to do it and they might be right.
35:34 And so for that setup, we see an even more dramatic speed up over the homogeneous speculative decoding setup than we did before. So, just to kind of, you know, summarize the main points here before we close out the main part of the talk, um you know, we believe that heterogeneous hardware is needed for making agentic inference be what we all need it to be, which is very fast, very efficient, and leverage our hardware and data center investments well.
36:07 Um because different chips are different in terms of their cost and performance characteristics just like different parts of agentic inference, we can use all of this compute better by intelligently deciding what piece to run where and not just throw the same Ferrari at every problem or the same, you know, I think uh Jonathan Ross said you need both long haul trucks and delivery vans when you're orchestrating like logistics and I think it's a great analogy. Neither of those two things is better or worse than each other, they're just used for different task. And then the result is that these systems deliver much better performance in terms of both throughput and latency for the same resource footprint. This matters because we're all scrambling to figure out how we're going to serve these workloads at scale as they continue to grow and by the way, building out data centers, that's a very physical process and that's hard to do in the blink of an eye. So, we need to figure out how to make better use of all the compute that we have in order to make it so that I don't have to wait a single second waiting for my coding agents to give me the response.
37:13 Just a joke there, but I think we all like to have faster inference. So, if you're excited about problems like this, uh would love to chat. Uh three areas, obviously if you're looking to speed up your own inference workloads, you should reach out. If you are a hardware provider and you're interested in partnering with us, we love to partner with hardware providers and figure out how to best use what they have to speed up workloads with our customers. And then the last is we're hiring, so we're going to triple this year. We need uh talented uh performance optimization experts like this group, so we're based in San Francisco. Uh scan that QR code if you're interested or feel free to give me a shout.
37:53 Thanks. Awesome. Super exciting, Natalie. Thank you. A lot of great great uh compliments, comments in the chat here. This is definitely a lot of information and yeah, Anja, um I just got got back on the call from from my other call. So, there's questions you have queued up. Take it away. >> Yeah, we we have a bunch of questions here in the chat. Um one I want to pick here because it's kind of interesting.
38:23 Are you planning, Natalie, to also design, build on the semiconductor level? I think for now you're partnering, right, with with chip vendors, but are are there any plans to go a level deeper on your side? Yeah, it's a great question. I mean, a lot of us come from a hardware space and uh like folks have been architects of chips before, but I think that we think that the problem we're biting off on the software level is already quite significant. Um I don't know who caught the Jensen Huang podcast, but one thing Jensen said repeatedly is to do as much as possible and as little as necessary.
38:58 And so, I think that's a really good framing. We want to do as much as possible to make these workloads fast, but as little as necessary is to leverage the amazing innovations that are already happening. So, we plan to partner very closely with our partners and future partners on the hardware side rather than make our own chips right now, but you know, never say never, I guess. Right. Cool. Um and then another kind of not related, but like similarly on the hardware level. Um how do you handle the hardware failures? Like, let's say individual GPUs or networking delays. Is that something you tackle on your software layer or would you expect like another layer to handle that?
39:36 Yeah, I mean, you know, when we're serving these workloads, it's on us to make sure that things are working. You need redundancy, you need good observability, and you need to plan for it, right? Like even like very battle-tested GPUs can fail, and when you're dealing with emerging hardware, that problem can be even more common, right? Just because some of these platforms are new and haven't been rolled out at scale before. And so, it's something that you have to design your system to anticipate, because otherwise it's going to really affect your end customer experience.
40:08 Yeah, agree. And then I think we had um earlier on when we were when you started talking about the how you disaggregated different components, the prefill decode, speculative decoding, um there was a question how much does the KV cache transfer affect the overall memory pressure? Yeah, yeah, that's a great question. So, um you know, with these systems, it's like we don't care about isolating any one thing in isolation, right? You need the full thing to work well, and if it doesn't, then it doesn't really matter.
40:35 That's just a science experiment. So, where KV cache transfers, um there's a lot of things you can do there to hide the latency. So, with prefill decode, one benefit is you only have to do it once. So, you can pipeline that transfer and stream that transfer and start processing on the decode side as it's streaming, but also it's only one single time. And so, that helps because you're not continuously doing it for that type of disaggregation. For other ones though, you are you are regularly going back and forth, and that's where ensuring that these chips are connected to each other becomes really important.
41:09 And that means we have to design systems that are really good fits for different kind of subsets of the workload, so we know which ones need to be connected together, because you can't actually connect every single chip to every other chip. Right. Yeah, Chris, did you I know you had a couple questions. Yeah, on a on a related note, there's there's a lot of attention these days, at at least from my lens, on the network fabric, and making sure that So, like have you folks thought about um tuning that that layer as well?
41:50 Like tuning it Like it Do you mean in a particular sense or I mean, you definitely need to you definitely need to tune it and make sure it's doing what you need, but I'm not sure if you have a more specific I guess I'm I'm I'm I'm thinking at it from like a PyTorch level. You've got, you know, uh you're using nickel collectives, distributed you know, training, distributed inference, you've got MOEs, you've you have different parallelism strategies and such, and um a lot of that, you know, it's it's more than than just the kernels, obviously, and it's it's the system-wide. So, I was just wondering, cuz I do notice most PyTorch folks Mhm. don't don't really think about those layers. I mean, once they they start tuning, then they do. Then they realize, oh, there's you know, there's a whole 'nother thing here, but even like just having insight into the fabric and and how those collectives are performing.
42:49 Yeah, I think there's a few things there. Like the first is knowing what the performance is, right? Just knowing where things are at is half the battle sometimes, because if you know the problem, a lot of times the solution is more obvious than the problem identification itself. So, obviously you need to have like a really good sense of what's happening and be empirical about it, because sometimes you can speculate all day about what the problem is, and then you're completely wrong. Um I think another thing that is interesting is the communication kernels. I don't know if you saw, but there was a really cool paper about co-optimizing kernels with like compute kernels with communication kernels that came out, I think like a month or two ago.
43:29 And figuring out how you optimize these things not just in isolation, but together in the context of the workload, I think that's something that's under-explored in the networking side right now. Definitely, definitely. Any Okay. Uh did I see that paper? I No, I don't think I did see that paper, and that's exactly what I'm I'm asking about. Okay, cool. Yeah. >> Yeah, I'll send you a link once I dig it up. Yeah, yeah, I can find it here. Um Let's see. Yeah, a lot of great feedback here.
44:04 Um the other layer that I've been paying attention to recently is the storage layer, uh you know, and it it it on the surface appears to only impact training, right? Cuz you have to load your data sets, you have to do distributed samplers, all that kind of obvious stuff, but inference-wise, um when we're moving, you know, it there's like CPU offloading, there's like KV cache offloading, there's like a whole bunch of like offload strategies.
44:37 Yeah, can you talk about some of that? Yeah, I mean, I think storage is a really big part of inference, too, especially agentic inference. Like you said, KV cache, I want to come back to my chat and be able to resume, right? That means it needs to be stored somewhere. Um but also things like sandboxes, right? And more server-side execution of tools, um you know, those typically need storage, too, and I also want to resume those interactions. And so, I think that you know, it's one of those things where the problems kind of get solved as the workload develops further, but yeah, we we do see storage as a critical part of this.
45:13 Totally. And it there's really only a few companies that are, you know, there's there's the Wekas and the Vast Data and things that are really hammering this stuff yet. Like uh so, obviously NetApp and, you know, folks like that, but yeah, I think it's an interesting space to watch, cuz we're pushing the the boundaries, not just compute anymore, right? It's like all the different layers, it's the network and storage, so. And you have to, because like I mean, you know, we aren't seeing the same transistor scaling than that we had been before, and so you have to look for other ways to optimize if you want to keep delivering these crazy speed-ups that we've seen over the past whatever many years.
45:57 So, kind of a fun question. Yeah, Anshul, like you and I stopped like stopped asking this question. Maybe it was too personal and probably wasn't good, but what what kind of music do you listen to, Ben Natalie? Yeah, how do you keep yourself entertained when you're in the zone? When you're in the flow? So, I have the wonderful characteristic of being locked into the music I listened to when I was 20. So, [laughter] when I was 20, the Glitch Mob came to my college to perform. Uh-huh. And I have Pavlov'd myself into whenever I'm working, I listen to the Glitch Mob, and now sometimes I'm like, why am I not being productive? And I realize I didn't put on my headphones and put on the Glitch Mob.
46:41 >> Totally. And I get mad, cuz I'm like, I wasted two hours not in the zone. That's hilarious. Yeah. Um yeah, like we've had everything. So, yeah, Harrison Chase is a big Taylor Swift fan, so we we thought that was interesting. Um I'm a big country fan myself. Um Anshul, who do we have on that liked Bob Marley? And yeah, we've had the whole spectrum. >> remember, yeah. We had every kind of music, I think, over the years, which is good. Is this your like pull at the end?
47:10 Yeah. I mean, my dad, I have to say, he is uh someone that is like fundamentally cooler than all of his kids, I think. And we didn't actually realize it when we were teens, and then he will like source all of these artists that are way cooler than anything the rest of his kids listen to. I think he I think he like discovered Billie Eilish before the rest of us and >> [laughter] >> like all this stuff. So, yeah, I think I need to take some lessons from him.
47:43 Hilarious, hilarious. Yeah, Anshul, uh what do you listen to when you're in the flow? You're a Jack Johnson fan. I'm a Jack Johnson Bob Marley fan, yeah. I need like because my stress levels go up with the work, and then I need music to kind of bring me down again, so to balance in the middle. Yeah, I'm a I think for the for the flow, I'm best with Grateful Dead, just cuz it's it's stuff I used to listen to, you know, and it just it's like jam bandy type stuff, right? But yeah. Oh yeah, I can I I really like the Dead, um gosh, there's this one song, but I can't think of it, but China something, China sunflower something.
48:26 >> Oh, uh yeah, China Oh god, China Grateful Dead. Yeah, yeah, that's one of your favorites. Yes. Oh, China Cat, right? >> China Cat. Yeah, yeah, yeah. I love that one. Yeah, it's great when when that song leads into like three other songs, and you're like, wait, this is the same song, you know, but and it's been like 35 minutes or whatever. >> I love what artists do that, and they blend the like melodies across different ones. I feel like obviously the Beatles were known for that.
48:57 Totally, yeah. Awesome. Uh yeah, Natalie, this is, you know, one of the best talks we've had based on the the feedback here, and you know, yeah, my own observation. So, very happy to have you on. We would love to have you on again, um maybe after the next fundraise. Hey. Yeah, or whenever, yeah. Let's let's do it. This is fun. Uh I think like whenever you give a talk, you always like have a panic the day before, and then, you know, it goes better than [laughter] you think.
49:24 Yeah, and I chose Mondays like 15 years ago when I started this beat up, and um and so, fortunately, the panic happens when you have a little buffer, like on a Saturday, Sunday, and you have to put it together. Um quick uh actually, yeah, one quick question about the fund raise. Yeah, how did that process go? Was it smooth? Was it Did you talk to like 50 folks and they, you know, like 49 said no kind of thing? Yeah, so our CEO Zane definitely bore the brunt of this as the CEO. So he could speak better to this, but what I will say is that I think that it's one of those things where when you get that believer and for us, um, that was Tim at Menlo.
50:05 Everything else falls into place and it's about finding that person who really believes in what you're doing and he he, you know, he believed in us before a lot of the really exciting stuff from 2026 happened for us and that is something that you always really value in an investor partner because it's easy to jump on a bandwagon, but it's hard to see when it's not proven yet. And so with all of our investors in our series A, I think they're always going to hold a very special place for us because of that.
50:36 Awesome. Yeah, did you go to Sand Hill Road and all that? You got to put in the >> a lot of them in SF these days, too. Yeah, it's a mix, but yeah, I think there were a lot of trips down there. I think though the thing with, you know, pitching to investors is that it's kind of like writing like a blog post or something where you may have things in your head, but forcing you to articulate it to an outsider, I think is a very good exercise because you kind of realize that as you start to make the slides like, "Hmm, like this isn't sound like as good as it is in my head. Like, why is that?" And then through that process you often discover that there's better ways of communicating about what you're building.
51:14 Yeah. And they they ask questions that force you to think of the value proposition and, you know, those those key slides like team and and things are like one thing. You Yeah, it sounds like like you guys have that nailed down pretty well, but then it's value prop. Yeah, how are you going to spend the money? So it sounds like you folks are just hiring like crazy, right? For the rest of the year. We're hiring like crazy. Yeah, I think we're in that phase now where everyone's annoyed that how they have to do so many interviews.
51:43 >> So many interviews. Yeah, but [laughter] I mean it's important. It's important and we need I mean, what we're doing like all the platforms we need to support, the scale of the workloads we're deploying, like we need we need, you know, great people and a hire fast. Got it. Okay, so everyone check out the Gimlet. It's what? Like gimlet.ai/careers or jobs or Yeah, there was a QR code there. >> I I can also like re-flash the QR, but um, let me do that actually.
52:14 Where does the name come from? I mean, what Yeah. Good question. So, um, once upon a time I was looking up words that had ML in them. And I gave the short list to my co-founders and I was like, "I kind of like Gimlet. What do you think?" So, yeah, it's not that uh, but it is a Rorschach test in a sense because people come up and they're like, "Oh, is it for the drink?" And it's like most of the people at the company don't drink, which is funny.
52:44 >> [laughter] >> Or they'll think it's the um, they'll think it's like the tool. Like there's like a tool called a gimlet. Right. Or they'll think it's the podcast. Like they'll be like, "Are you a podcast company?" And it's like, "No, but maybe one day." >> [laughter] >> Yeah, didn't Open AI just just buy a podcast or something, which no one could figure out why? Oh, yeah, VPN. Yeah, yeah. So Gimlet, the podcast company was actually, um, they made some of my favorite podcasts, so maybe in the back of my head it was like a positive association when we selected the name. We also liked how it ended in LET because it felt like a little uh, reference to Kublet and we all had worked a lot on Kubernetes systems.
53:25 Gotcha. Crazy. Yeah, when did you start using Kube? Oof. I guess back in 2017. I know, right? Yeah. >> We're like the old-timers now with Google >> Yeah, it it's funny how the hot new thing becomes like, uh, you know, the thing everyone uses and then becomes like, "Oh, yeah, that was like a while while ago we were all working on that." But I think it has stood the test of time. Yeah, like Linux and, you know, Kafka, right? Like Kafka is just like part of every, um, yeah, large scale system, so.
53:58 Yeah, 2017. I don't like I don't even think Kubernetes supported like GPUs really back then, right? Maybe 20 16 or 2017. >> don't think it If it did, I think it was experimental. Yeah. Yeah, there was a custom runtime. There was a custom Docker runtime. Uh, I think it was Nvidia-Docker that could mount the GPUs and there was all kinds of security holes and all kinds of crazy stuff, but yeah, they tightened it up over the years, Yeah, it's one of those things where you you make this thing, you prove it out and then you just, you know, whack-a-mole all the problems over time.
54:30 >> Totally. Totally. Yeah, you remember KubeFlow and all that? Oh, yeah. Oh, yeah, there's been [laughter] Yeah, so much cool Kube stuff and yeah, for us like I said, we build everything on Kubernetes because it's just like the foundation for what we do. And, you know, I've been going to KubeCon since I think 2016 like when it was real small and there were no ML talks, zero. And this this started, I think, 2021, maybe there was eight talks and then 2022 KubeCon had, you know, 10 or 15 and now it seems like every other talk is, you know, like yeah, AI and and Kube, so. Yo.
55:07 Awesome. >> Yeah, well, you need distributed systems for AI, so it makes sense. >> [laughter] >> Great. Thank you so much. It's been awesome. We will see you soon. All right, thank you. You too, Natalie. Bye, everyone.
Summary
- Natalie introduced Gimlet Labs, which recently raised $80 million for developing an inference cloud focused on agentic AI.
- The talk emphasized the importance of heterogeneous hardware to efficiently handle diverse AI workloads, which have become more complex over time.
- Inference workloads are evolving, requiring different hardware capabilities for various tasks within the same model.
- The discussion included the concept of disaggregating workloads across different hardware types to optimize performance and resource utilization.
- Gimlet's orchestration system intelligently splits workloads and routes them to the most suitable hardware, enhancing throughput and reducing latency.
- The presentation showcased a case study demonstrating significant performance improvements when using specialized hardware for specific tasks within a workload.
- Natalie encouraged collaboration with hardware providers and highlighted the importance of continuous research and development in the AI hardware space.