transcribe

Inference Engineering (The infrastructure of AI) with Philip and Ben

Ben Dicken · 56m · transcribed 1d ago
More from Ben Dicken Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Introduction to Inference Engineering

What is the purpose of this live discussion?

The discussion is centered around inference engineering and the new book on the topic.

  • The hosts are excited to discuss inference engineering.
  • The book on inference engineering has gained significant attention on social media.
  • The conversation will cover insights from the book and the field of inference.
# 11:15

Understanding Inference

What is inference and how does it work?

Inference involves using a GPU to run a model quickly, which includes capacity provisioning and managing dependencies.

  • Inference requires a GPU, a model, and optimization for speed.
  • There are complexities in provisioning capacity and managing model dependencies.
  • Understanding the basics of inference is crucial for engineers.
# 22:30

Challenges in GPU Reliability

What are the reliability challenges faced in inference systems?

GPUs can fail frequently, affecting uptime, and cloud providers can experience outages, complicating reliability.

  • GPU failures are common and can lead to significant downtime.
  • Achieving high uptime is more challenging in inference than in traditional databases.
  • Active monitoring and reliability strategies are essential for managing GPU failures.
# 33:46

Learning Path in Inference Engineering

How did the speaker learn about inference engineering?

The speaker learned through hands-on experience at Base 10, starting before the AI boom.

  • The speaker's journey in inference engineering began with no prior knowledge.
  • Real-world experience and curiosity drove the speaker's learning.
  • The rapid growth of AI created a unique learning environment.
# 45:01

Optimizing AI Models for Specific Tasks

How do companies optimize AI models for different tasks?

Companies may fine-tune multiple models for specific tasks to improve performance and reduce costs.

  • Fine-tuning models can lead to significant improvements in cost and reliability.
  • Companies are increasingly adopting specialized models for various use cases.
  • Investing in model optimization is becoming a standard practice in the industry.

Transcript

0:00 No, you're fine. Yeah. Okay, we are live. >> We are live. Let's go. >> Live. We're live to talk about inference. Let's do this. >> Well, folks, we're live to post on Twitter and we post on Twitter. >> Yes. Yes. Yes. Yes. Keep it simple. Boom. There we go. Okay. >> My my number one Boomer habit is still calling it Twitter. >> Twitter. Yeah. I feel like about 80% of the time I still call it Twitter.

0:29 Hopefully Elon's not watching because I'll get banned probably. But >> all right, I'm gonna I'm gonna quote tweet this. independent sources confirm that we are in fact live. >> We are. The sources might be biased, but >> Hey everybody, >> dependent independent sourcing. We've got we're just going to let people file in for a little, but we've got Philip with us today. As you all you all are used to me talking about database books. You're all used to me talking about either this book or everybody's favorite, this book, but we're talking about inference engineering today. New book is >> Yeah. Oh, yeah. Get it up here.

1:17 >> This book. Boom. >> Yes. Which my copy is in the mail. >> Doesn't have it. Yeah. So, it's it's in the mail. in the mail. But yeah, we if you if you exist on Twitter in tech Twitter, you probably saw this book because it kind of took over Twitter on what Tuesday it >> Monday. Yeah. >> Trending. This is the first time I've ever been trending. And I didn't have to become an NFL player and commit some kind of traffic violation to do it.

1:46 >> The funny thing is I saw myself trending today too for different reasons though for like all these database benchmark things that came out. That's probably not even in your world. That probably just shows up on mine. >> That's awesome. >> But yeah, so we're gonna we're gonna chat about this, but we're going to let people file in a little bit, think of questions for Philip. Basically, Philip's going to teach me what the heck inference engineering is. and then maybe after we finish our database internals book in a couple weeks, maybe we'll do some more streams where we learn more about inference infrastructure. Might be fun.

2:19 >> That would be cool. So, Ben, while we're waiting for everyone to come in, I have a very important question for you. >> Okay. >> How are you doing your lighting to look so good all the time? >> Thank you. Yes. So, a part of it, the secret weapon is a guy named Steve Tonudo. Twitter. >> Yes. >> Yeah. Yeah. So, he we overlapped at Planet Scale for a couple months and he taught me a lot of tricks for how to set this up and then I took it from there and improved it and all that. But anyway, there's some lights, there's a nice camera. and it's just been like fine-tuned over time. So, >> but yeah. Yeah, if you if you get that that recording setup at your office, hit me up and I'll help you pick out the right equipment or just message Steve.

3:04 >> Yeah, that's that's necessary. I I'm currently on a laptop in the windowless room in the back of the office that I was allocated to turn into my book production factory. >> It happens. It happens. Yeah. Let's see. Oh, wait. It did. Oh, hold on. Hold on. Hold on. I double posted it. It did autopost it. Okay, cool. Well, >> I mean, two tweets is better than one. >> Exactly. >> Nice. All right. Well, let's we'll start with a softball question here >> because you're an AI guy.

3:41 >> Yes. >> I Okay, I actually I hear everybody on Twitter talking about how codeex 53 is like the best model ever and is better than Opus, but I don't agree with that. I don't think I don't think Codeex is better than Opus. What do you think? What models are you using, Philip? >> Look, I am number one Koso Shell. and so I use a lot of composer because it's really really fast and I'm impatient and I want my code now.

4:09 but I >> Okay. Is that your because again some people have these styles of like write a prompt and leave for 10 minutes and come back. Are you a lot more of a like in the loop coder? >> Very in the loop guy. Like I I am a I am bad at delegation and good at micromanaging. And so I do that to my agents as well. >> >> yeah, >> that's what I found when the first when the first composer came out. I was using it like crazy because I was like it was like right after it came out right before I went on a plane flight and so I was on a plane just coding the whole time. I had good internet star link thankfully and yeah it was like so nice to be like wow I'm getting results back in like three seconds to my prompts right like this is incredible.

4:50 >> Yeah especially because I'm the sort of very lazy developer who will prompt it to like you know change pixel size or you know change some padding or something when I could just go in there and you know control all edit I'll just be like yeah make everything a little bigger so it's nice to it's nice to have some some quick stuff >> that said when I am doing something a little bit more sophisticated larger context I do toggle over to opus that's my that's my go-to >> okay cool >> I really like 4.5 4.6 six. It's also nice. and >> yeah, I mean, there's a there's there's a reason it's a it's a popular model.

5:31 >> To be honest, I feel like I didn't even notice a difference between 4.5 and 4.6. To me, it was like feels like the same thing. They're both good. Like, >> yeah. 45 was a big jump, though. >> Yeah. Yeah, that was a big jump. >> We We think about like benchmark saturation, right? you know, it's like, are you going from 99 to 99.9% on some benchmark and is that even valuable? And >> my sort of hot take is that it actually is like maxing out those last few things because if you think about it one way like say going from you know again like 99 to 99.9% correctness. It's like on the one hand you've only gained less than like 1% of of quality but on the other hand you've reduced 90% of errors.

6:24 It used to for every 10 errors you now have one error. Yeah. >> And so I actually think that is very material to the end user experience. and so sometimes these like perhaps on paper smaller jumps can actually have really big impacts on your day-to-day. >> Yeah. Yeah. Cool. Well, the the big question I want to ask and probably the audience wants to know too is like I I feel like Okay, so I work at Planet Scale. We're a database infrastructure company. you work at an AI infrastructure company.

6:58 >> So like I'm pretty familiar with like what it's like to run heavy infrastructure that companies rely on, but what I'm not super familiar with is like the way AI infrastructure companies set up their infrastructure, right? And obviously there's like training and inference. Those are the two big parts of it. So just give us like the high level kind of one like why you wrote the book, but I think that also ties into like what is inference and why is it important for not just like AI engineers, but for like any engineer?

7:24 Why should every engineer learn this stuff? >> Let me let me see. I've kind of forgotten what the answer is. So let let me let me see if I there's like some some book that I that I can read this from. >> Interesting. >> Yeah. So the thing is like there are a lot of people running like using models. You know, if you look at all of the AI applications, however many trillion, potentially soon quadrillion tokens per day are being generated across the world. there's there's a ton of demand, but relatively few engineers actually know how to do inference. Like maybe a few years ago, it was like a hundred, I don't know.

8:13 if you if you think about like Google and OpenAI and Anthropic and a couple other big labs having a a small team of you know mostly systems engineers people who are familiar with like Kubernetes and and web infrastructure setting up these these inference systems. today there may be a few thousand. and my my sort of premise behind the book is that I think depending on a few different factors in terms of demand and in terms of like how much it remains a very like human-driven process to set these things up potentially like hundreds of thousands to even like a million really good engineering jobs where inference is at least a part of the work.

8:59 >> and this is both I assume like on the hardware side like building the data centers and managing them and on the software side like writing the systems that manage these and all that, right? >> Yeah. Yeah. Because I'm thinking about like every application layer company is going to need people at it who understand influence. even if they use a managed service provider like Bayen, which they absolutely should, like you still need people on your team who know what to ask for and know how to check whether or not we did it right. and so, you know, I I think that that inference is one of those skills that every single developer has the opportunity right now to to add to their repertoire and then like if you really like it, become an early expert in it before everyone else gets there.

9:50 So, that's why I wrote the book is because I wanted there to be a lot more influence engineers. And we hire at basin, we hire amazing engineers from all kinds of backgrounds. And even they sometimes come in and struggle to like really grasp the whole scale of this problem because it is a large problem. Like inference requires dozens of different technologies working together to make a a model run fast, have high uptime, have high throughput. and so you know it's it's it's really hard to become a expert in the full stack and it's hard to you know if you're an expert in just one piece like see the rest of the puzzle end to end but inference is like a very integrated problem like you might have to go from thinking about like kernel fusion to thinking about you know your K native setup from from one hour to the next if you're setting this whole thing up yourself.

10:52 >> Yeah, you're going to have to explain what all those words mean. But I like what you're saying, which is like, and I kind of say similar things about the database space. Not as new as the inference space perhaps, but similar of like a lot of people are like, oh, I don't have to learn anything because AI is going to write all my code. I'm like, database systems are very complicated from the software, including the hardware layer and all the infrastructure. So like if right now you can like learn that super well, you're going to like set yourself apart as an engineer, right? And similar to like probably even more so with AI inference, right? If you could like become an expert about it in 2026 and get a job at a great inference company or whatever, like you're you're going to do great things, right?

11:30 >> and okay, so but the the background too, so like your audience probably is more familiar with even like how inference works, maybe my audience a little bit less. Like give us the couple minute like what is inference? Like what do you mean when you say that? So, I've done a lot of trade shows, like Reinvent and GTC and that kind of stuff, and I this question a lot. and the simplest answer I give when I'm tired and I can tell that I'm talking to someone who's never been inside of a you know, CS degree or anything. is I say, well, inference is really simple.

12:04 You got to do three things. Number one, you got to get the GPU. Number two, you got to put the model on the GPU. Number three, you got to make it run fast. >> and while that is a a very very oversimplified e explanation of of how inputs works, it's not wrong. you have first like the problem of capacity provisioning and and securing capacity and then you have the infrastructure problem of getting you know a terabyte of model weights onto whatever infrastructure you've provisioned along with the images and you know setting up the many dependencies of you know that that are constantly having breaking changes and overnight builds in them. and then once you have everything set up and you can run your inference engine, you can run the the various programs that actually like go through the model iteration by iteration and generate output tokens, generate output images.

13:09 you need to figure out how to make that process faster because >> users demand extremely low latency applications. and then you have to figure out how to keep it fast and keep it online when you start slamming these things with the sort of viral usage spikes that AI apps get. Like AI applications, a lot of a lot of, you know, applications in this industry grow at like a 10x year-over-year rate. And within that, that's not just steady.

13:41 That that's like big stair steps that you have to figure out. And I imagine AI is pretty spiky in the sense of like during the during the day versus during the night and all of this kind of stuff, right? Or like new model drops. Yeah. >> Spikes and then declines after that, whatever, right? Yeah. >> And then you're on to the next spike because another model dropped the next day or sometimes. >> Yeah. >> Sometimes even the same day at this point.

14:04 >> Yeah. And that's okay. That kind of ties into a question I've like personally had and I have wanted to research this but just haven't had the time. Maybe your book has the answer. One of the things I've always been curious about is like okay you've got a model whatever I'll just use the the anthropic ones as an example but take for anything right like sonnet versus opus right one of those you get your results back a lot faster but it's a dumber model right it's not as sophisticated it's it doesn't have as many parameters so in terms of like the infrastructure what is actually the difference like why is opus or whatever bigger fancier model take longer is it because there's more cycles or is it because it needs more GPUs and needs to be more parallelized. Like what actually makes it take longer, >> you know? That is a great question and a framing that I haven't really heard before. It's like it's like well obviously intuitively a smaller model is faster. but but why why precisely is that?

15:04 it's one of those sort of laws of the universe I've never thought about. It would be for a few reasons. So when you look at the architecture of a model, this is why like mixture of experts models generally are a little bit faster. you have when when you sort of break it down all the way to the konal level. you have a lot of matrix multiplication going on.

15:35 >> and multiplying big matrices is is slower than multiplying small matrices. so everything that you see around you know model size and and speed is downstream of of that reality. There are architectural things that change how fast a model is. It's not like when you look at models with different architectures they larger models are generally slower but not like it's not it's not perfect in that direction. It's it's it's like it's directionally correct but it's more like a scatter plot with a line of best fit than it is >> because I also imagine yeah you could just parallelize more in theory and make it just as fast as a smaller model.

16:27 Right. parallelism sometimes makes it faster and sometimes makes it slower. this is >> probably again all depending on how you architected it to begin with. Right? >> So so there's there's a couple different types of parallelism. So let's say you have a model first off like most frontier models I'm saying frontier to mean like a model that is among the top most top say like 10 most intelligent models out there. any any frontier model is going to require parallelism because it can't fit on one GPU.

17:01 >> Yeah. >> So if you look at like >> if you look at like Kimmy K 2.5 that's like a trillion parameters. and if you look at that in a 4bit quantization, where you're using half a bite per parameter, you have 500 gigabytes of weights and >> and that has to fit in a GPU's VRAMm or many GPUs VRAM. >> Exactly. So if you look at a B200 which is you know kind of the top of the line for inference right now all the GB300s are are coming online quickly a B200 has 192 gigabytes of VRAM so you need you know and and and you can have one two four or eight of those and so you need at least four just to hold all the parameters and then you need more headroom on top of that to hold the KV cache and to actually like do computations during inference. so then you end up needing eight GPUs and and anyone who's done parallel programming before knows that like parallelism makes things harder and makes things slower unless you are unless you're doing it right. So you can't just like send it from one GPU to the next to the next to the next like a pipeline because then you get bubbles you get you get a lot of slowdowns.

18:19 >> instead you have to do something called like tensor parallelism or expert parallelism for mixture of experts models. and you know very very generally tensor parallelism makes things lower latency and expert parallelism makes things higher throughput. Although like the most serving things you'll find will have a mixture of the two. And these are again like if if one of the actual model performance engineers was in the room hearing me say that they'd be like whoa whoa whoa whoa whoa. There's like a hundred factors you haven't considered there. But you know from my observation so far it seems like tensor parallelism is latency one and expert parallelism is throughput one.

19:04 >> Yeah right cool. Okay. Yeah, that's again maybe yeah, maybe others don't think that way, but maybe because I'm in the database world, I'm constantly thinking about like query latencies and why is one database slower than the other? Why is one query slower than the other? So to me, that's an obvious like why is it that one model can give me an answer in two seconds and another takes 20 seconds of inference, right? yeah are are you so base 10 is do you guys have like host your own GPU infrastructure or are you kind of like managing GPU infrastructure that's provided by like the other big cloud providers like AWS >> exactly more more like the second one we're on like a dozen different clouds but we can also like if you bring us some GPUs we can run on those too where you get new GPUs I think of as much more like that's like a a capacity that's almost like an accounting question. then once once you have them the influence looks the same no pretty much no matter where they are, >> right? Yeah. Interesting. Okay.

20:08 What what are the well so I don't know how familiar you are with again I get I don't know your your background super well but like I'm very wellversed in the whole database infrastructure space and like how you set that up and like what the best practices are for like okay you have an app and you're connecting to your database and what makes for a good HA database but like at least largecale inference is a pretty new problem like inference as a thing has been around for a while right but like it's never been used as much in the past couple years right so what are like the biggest challenges right now in scaling it is it literally just we don't have enough GPUs, Nvidia needs to make more GPUs or like are there like optimization challenges or just like or for example why do like I'm not saying base 10 but like other providers sometimes you see their status pages and they're like going down all the time and it's like why why are they going down every day like why is this so hard? So what makes it hard to like run a big infrastructure of compute?

21:02 >> Well, there's there's there's a there's a lot of challenges. I'm going to start with the last one because your last question. I think that's that's one that's very interesting. these machines that you run your your databases on, they're CPU based, right? >> Yep. >> How often do they fail? Like how often does the machine itself just like die? >> Not very often. Yeah, there there is. So when when you have some very large databases run like you have thousands of servers all working together.

21:34 >> Yeah. And when you have thousands, one will die frequently and you just replace it, right? But like a single server, like it's very rare, right, that will >> and if you if you were to like estimate like one failure per n like CPU hours, >> yeah, maybe one a year, like that's that's order of magnitude, right? Yeah. Like might be more or less, but yeah, one a year. >> Yeah. So if if you look at this this chart is from adapted from like the llama 3 training paper >> and training and inference are different problems but they still run on GPUs and GPUs fail for the same reasons like the GPU itself can fail there can be an infrastructure failure there can be a network failure there could be like dependency breakage but if you look at like pure hardware failures alone based on those results you might expect one hardware failure every 50,000 GPU hours. So if you run a single node of H100 AX H100 for a year, you should like more than likely like more than 50% chance expect that one of those GPUs is just going to straight up die during that year. and if you're trying to hit, you know, four nines of uptime, which I imagine for like a database provider is not actually that good.

22:58 >> yeah, that's like the bare minimum. 49. Yeah. Yeah. >> For influence, 49s is is best in class. yeah, that's a few minutes of downtime per year. 59s is a few seconds of downtime per year. which we have achieved in like >> I think it's per month. Actually, I think it's like a few minutes per month versus a few seconds per month. Something like that. But but yeah, I get what you're saying. Yeah. >> Yeah. So like >> I guess the point is it seems harder because obviously there's a million very smart people working on it, but yet they still seem to like go down a lot more often than a database does. Right.

23:29 >> Right. So so the first the first problem is that like the what I'm trying to illustrate here with all these figures is that the the ground that we are building on is not very stable. so right >> the GPUs themselves fail relatively often. and when a single GPU fails oftentimes the whole node has to be taken out for it to get fixed. so the the the first problem is like how do you detect GPU failures and then automatically like adjust based on that.

24:01 so we do a lot of stuff around like active active reliability active passive reliability. The second problem is the cloud providers themselves. while of course everyone is working super hard and and you know following every possible best practice, the the reality of this world is that like your cloud providers will have regional outages as well as complete failures. and sometimes those can take hours to recover from. >> And again that's like you could just tell your customers like hey our SLA is like exclusive of cloud provider downtime. or you could again figure out how to like fail over so that you've got some GPUs running in AWS, some running in GCP and like >> exactly >> put put stuff wherever it needs to be.

24:49 so it's the same honestly like it's a lot of the same engineering work from a systems perspective that went into making distributed databases the sort of bedrock of reliability that they are in our industry today. just like starting with something that's orders of magnitude like more flaky and then applying the same sort of rigorous process but like more of it I guess >> because yeah in the database world that's the same thing right we run on the cloud providers but these nodes can fail at arbitrary times so we're we gotten very good at doing failover and all these kind of things to make it seem like there's never a problem even though there sometimes are problems and we just mask over them right >> yeah and Then everything just like takes longer because you know node acquisition time like actually spinning up a GPU takes longer than spinning up a CPU m much longer. then you have to stream your image onto it and your image is much bigger. I don't know how big a Postgress image is but I would imagine it's like less than a gigabyte.

25:56 >> I don't know off the top of my head but yeah probably less than a gigabyte. Yeah. >> Yeah. oftentimes our images would be like 10 to 50 times that size. >> Okay. Yeah. >> And then you have to stream the weights on and again those can be like 500 gigabytes. and then you have to like start the service that's actually relatively fast. so yeah, recovering after an issue, even if you like get a node right away, if you think like one gigabyte per second is great bandwidth, which it actually is in like almost all cases, it takes you 500 seconds to load your weights over that. That's just, >> you know, you can't you can't can't be like that.

26:38 >> yeah, that actually that's a very good point. Yeah. It's just like you expect this super low latency, but there's all these like even with good hardware, there's all these like really slow operations that have to happen for it to work. Yeah. >> And then and then another thing is like the the majority of the capacity out there in the world is like hopper capacity or even earlier. Blackwell came out Blackwell's the latest generation of Nvidia GPUs. came out over a year ago, but they're still like much scarcer than than hoppers and and previous.

27:14 >> And the other thing is that like >> the actual kernel level engineering work changes almost entirely with every release, especially with you know Hopper and then more so Blackwell introducing like brand new asynchronous programming paradigm CCUDA. building like a highly optimized kernel for every single step of a you know transformer forward pass is like a big challenge and with export controls being in place the majority of the open source research coming out of China even today is still built for hopper GPUs.

27:52 >> >> interesting. So becoming really great at Blackwell inference and Blackwell kernel programming like remains a meaningful challenge even when we're already like halfway into the Blackwell life cycle. I saw this play out by the way with Hopper three or four years ago. Like it took year plus for the the industry to sort of >> catch up to the the hardware. and I imagine we're going to continue seeing that cycle. Like I'm very very excited for Reuben. When you look at the timeline of how long it takes to architect a chip and do tape out and manufacture it at scale like Reuben is the first GPU that was designed in a world where largecale language model inference was like more than theoretical, >> right? Yeah. So it should be really good at it.

28:44 >> Yeah. You can see in like the the even in the early ways they're talking about it that like they're leaning hard into disagregation they're you know leaning hard into VRAM bandwidth which is very very critical and into you know VRAM capacity into like low precision FP4 FP8 tensor core like everything that that has proven to be effective for inference is like coming way more but again it's going to take a long time for the entire industry to like figure out how to use these chips to their maximum potential. Yeah, I wanted to we got a question from Ahmed who he does infrastructure at LinkedIn.

29:28 So he knows database infrastructure and other kind of infrastructure really well, but he's curious does the your book talk about GPU health checks, remediation, scheduling, orchestration. Is that something you cover? >> I sort of touch on it at a high level. Like to be perfectly honest, I'm not so much an infrastructure guy. I know much about what happens on the GPU itself versus how we get the GPU and how the weights get to the GPU. So on the GPU is like six or seven chapters of the book and all of the infrastructure stuff around it is one chapter at the end.

30:05 which is you know something I'm definitely excited to to learn a lot more about. but honestly I think for for someone what was this guy's name? Amit. >> Attit Yeah. good good to see you Amed by the way. Thank you for joining in the stream. For someone like you, I I think the value of of a book like inference engineering is like you already know probably a lot of the infrastructure level stuff and so you can match that with like the understanding of what's actually happening on the GPU after you provision it.

30:37 >> Okay, cool. I have another one here too. I'm trying to figure out maybe you understand this better than me so it might be a dumb question but if we're doing inference in parallel >> Yes. How do these parallel GPU hold the context so they don't send generate the same result? >> That is >> I don't know if I quite follow the question but maybe you do. >> Yeah, that is actually an interesting question. So it I if I'm understanding it correctly, it's like you send one request to eight GPUs, how is it that you get one answer back instead of eight back? and that's that that comes from the the parallelism strategy where you actually shard the weights between the GPUs but you share the you shard the weights but you share the context back and forth. So on a node of 8GPUs you have a very high B bandwidth interconnect called NVLink which allows you to send the the data back and forth. It's still like much slower than say VWAM bandwidth. but it's it's fast enough that you can send like hidden states back and forth and and do your all reduces and stuff. so you know I think that the easiest way to visualize parallelism is actually expert parallelism. tensor parallelism is a little bit like more dependent on a strong grasp of the actual like the the actual architecture of the models. But if you think of a mixture of experts model and let's say you have a simplified one where you have 16 experts and you have eight GPUs basically two experts live entirely on each GPU and the way mixture of experts inputs works is each each pass through the model every time you generate a token it's like hitting a a few different experts at at every level. So maybe it's like two experts per level. So you can imagine in the forward past you're kind of like bouncing from GPU to GPU. You could you could even like sort of visualize it as like a student in a college and there's all the professors in their offices and they want to figure out an answer and they're running in between all the different offices to get that answer a little bit at a time and then at the end they walk out of the building knowing the answer. so yeah so that's that's kind of how the that that's how that's how the the system is able to use the the parallel the parallelization effectively without just like repeating work. and then all of that stuff is also going to be stored in something called a KV cache. so if you think about KV as like key values, and you like, "Oh, keys, cash, damn."

33:21 >> Well, I read, not the whole book, but a couple parts of the book this morning, and that's this is kind of one of the components of it that I was reading. Yeah. Yeah. >> So, so that's how that's how both within a request, you kind of remember where you are in the inference process. and then also like in between requests, you can actually hold on to that KV cache and then on the next request like potentially skip part of pre-built and use that to accelerate yourself as well.

33:50 >> Nice. Yeah. I'm I'm curious. This is changing the subject slightly, but like this book covers so much stuff. How did you learn all of this? Because like again a part of what you're published this book for is like you want others to learn more about this become become inference engineers. How did you learn about it? Because you didn't have this book to to teach you how to do it all. Right. >> So I've been working at base 10 for more than four years.

34:12 >> Oh wow. Okay cool. >> When I showed up here I didn't know anything about this. but fortunately like kind of neither did anybody else in the world. >> well, so four years, that's like even before the AI hype really began, right? Right at the beginning of it. >> Yeah. I joined B 10 almost a year before Chat GPT launched. >> Wow. Okay. Wow. >> So, you know, I had back then like I was super obsessed with AI and I had this really strong thesis that AI was going to be the sort of like massive thing that it is. No, I'm just kidding. I had no idea. I knew back then I would have bought a ton of Nvidia stock and just like you know played video games full years.

34:57 >> We wouldn't be talking because you'd be on a yacht or something right now and >> we'd still be talking. I I'd come on your show even if I had yachts. but what I what I was you know I I had no idea like back then I I had no idea that that it was going to be big like this. I was just like I had worked at a couple different jobs out of college and like hadn't really like found a ton of traction at either of them and I was looking for like some place where I felt like I might stand a chance of sticking around for a year and I, you know, met some cool people who were working at BAS. I like cold emailed and was like, "Hey, can I like interview here to be like a technical writer and write your documentation?" and I I I got in and then I just kind of stuck around and and again like my first job here was a technical writer. So I spent a couple years like just writing documentation for the platform.

35:56 >> >> yeah, which you have to then really understand the platform to do. So yeah. And then and then I switched to Devwell and a lot of what I was doing was like blog posts and conference talks where basically I would sit down with an engineer and they would just like explain everything they worked on and then I would package that up in a way that was digestible to a more general engineering audience. >> And so over that time over those four years you know I I just kind of picked up all of this stuff. And then about six months ago when I had the idea for this book, like we really started inflecting in our growth and especially at our hiring and I I saw people coming in like really smart, experienced people >> and being like, well, I don't actually know how to teach you this stuff because you don't have four years to just like sit around learning it at at, you know, at at the feet of the engineers who are who are doing it. You you need to know.

36:53 >> You got to hit the ground running. Yeah. Yeah. Exactly. >> so that that's part of what, you know, inspired me to be like, well, may maybe I could write down everything I know in one place. >> Yeah. And and definitely >> what's the most like interesting part to you like of all the stuff you you do your conference talks and all that like is there a certain aspect of inference or the education that that you like the most? I think the model performance techniques like the applied research is so cool because if you look at like >> when you say that you mean like the benchmarking and stuff or what do you mean?

37:25 >> I mean like like the the techniques that you use you know if you go look on artificial analysis allow me to show for a second. If you go look at artificial analysis and you look at like >> our performance for an open source model and you know there are a couple other companies out there who have model performance figured out pretty well as well. but if you look at that and then you look at like 10 other providers you know you you can look at something like Kimmy and we'll be at like 300 tokens per second and the longtail providers will be at like 50. So like why you know what what what are we actually doing?

38:04 >> yeah is that at the same cost or like comparable cost but just way faster. >> Yeah generally little lower actually. >> Oh wow. Okay. >> But unless un unless anyone feels like you know running unless unless other people are just like selling below cost. >> but the cuz yeah cuz with with better optimization you get like more throughput and more latency and you can kind of like >> move wherever you want on the on the sort of front efficient frontier between the two. anyway, so like the the in in something like medicine, you know, you might do research and then 10 years later some commercial drug will come out. And in something like, you know, maybe physics or or math, like maybe it's it's a multi-year delta between frontier research being published and like actually seeing it.

38:59 And here it's like >> people do research and then it's running live in production for the biggest companies in the world like a monthly. Yeah. >> And that's super cool. So you know when I sat down to write this book like I actually was only going to write chapter five which is about >> which one's chapter five >> performance techniques. It's the it's the quantization, speculative decoding, KV cache reuse, parallelization, disagregation, like batching, like all this like >> all the the cool stuff. And then I realized like well actually you need to know everything up and down the stack for like this is all this all to like have context and make sense.

39:37 >> But that's definitely my favorite part. is I love like the the applied research piece of it. and figuring out how to use these techniques to like make models fast in production. >> Yeah. What what do you think about and again that's that is more like actual model performance, but what is your take on the like evaluation benchmarks, right? There's all these benchmarks people publish, right? And sometimes it seems like they're kind of useful and sometimes it seems like they're meaningless because it's like, well, this one scored well, but then when I use it, >> it it feels terrible, right? So like >> what's the best way to evaluate a model?

40:16 >> You know, benchmarking is like maybe the biggest unsolved problem right now. >> I'll give you I'll give you an example. It's like the emergent behavior of language models does not always match expectations. You would think that language models would be amazing at proofreading. you know, like finding objective arrows in text seems like the sort of thing that a language model would be very good at.

40:46 >> Yeah. >> and I tried to get it to proofread my book. and you know, if if you know even a little bit about how to build AI applications, you would say, "Well, Philip, that's a dumb idea because there's like context degradation. Like, if you throw an entire book, which is 47,000 words, which is maybe depending on your tokenizer, 60 70,000 tokens. If you throw that in a model all at once, it's going to get confused." I'm like, "Yeah, absolutely." So, let's build a harness. Let's pass the book in a page at a time with like a very clear prompt about few shot examples what errors look like.

41:26 >> Let's make sure we run it through every frontier model and not just one or two so that we're you know seeing it from from all different angles. and you know, either I suck at writing agent harnesses, which is entirely possible because I'm a make the model fast guy, not a make the model useful guy, >> or proofreading is something that they're actually pretty bad at because it caught, you know, maybe a couple dozen errors in the book and then I sent it to a human proofreader after fixing those who caught another like hundred.

42:01 >> Right. Yeah. So we need cursor bugbot but we need it for authors you know bot or something. Yeah >> exactly. So so again you know potentially just like a an an issue with the way I orchestrated the system but also kind of the sort of thing you would expect a language model to just be able to handle. >> and like it's you know so so potentially I've discovered some kind of you know novel economically valuable task that language models are not good at yet. And >> right >> if so like I would want to build a benchmark for it. But like how do you actually go about doing that? How do you actually like build proof writing eval because >> here's what you do. You publish the draft of your book as the benchmark. You say here are the 150 errors and then see how many it catches. Right.

42:50 >> Yeah. Yeah. But generally you want to have your evals have like thousands or tens of thousands. >> And so that's one test case. Yeah. Am I am I asking LLMs to generate like synthetic errors, but then they're going to generate the kind of errors that they already know are errors? are going to a bunch of authors and asking for like, hey, can I get rough drafts of your books marked up by your editors so that I can put you editors out of business? And >> that's not going to go over well.

43:22 >> yeah. Yeah. >> Yeah. So the the the big problem in in evaluation is often like setting up the evaluation. Like once you have the data set and once you have a mechanism for figuring out whether or not the output is correct, like you can very quickly hill climb toward a model, including in many cases like a smaller fine-tuned model that is amazing at that class of problems. That's something we do a lot here now. We we added a post-training team. we do now a lot of like a lot of data collection and and eval work with customers. and >> oh to make to make like custom model tunings basically for their use case.

44:05 >> Exactly. Because like the the way to think about switching to open source is not like well I've been using Opus for everything and now I'm going to plug in DeepSeek and I'm going to run everything on that and now my whole application is going to be 10 times cheaper >> and you know and and five times faster. That's that's not actually how it works. the first thing you kind of have to do is like decompose your application into tasks and say like actually you know I have a task that's like search and I have a task that's like outline generation and I have a task that's like writing chapters. I have task that's proof reading and you you kind of create all these tasks and maybe a louder system that understands what task is happening and and selects the appropriate model and then you go build a data set around that task. You take model, you fine-tune it to be really really good at that task. And in some cases you can get your your model fine-tuned to a point where it's better than a frontier model at that specific task.

45:08 >> and then you you know you can do that with SFT, you can do that with RL like there are a lot of mechanisms for that. once you have again like the data and the the verification mechanism and then you you put that out there and all of a sudden that specific task in your in your product is way faster and way cheaper and way more reliable and then >> so are there some companies that will like let's say they have within their application like 20 different use cases of AI will some actually tune like 20 different special modules models just for being good at this and this and this and this Okay.

45:44 >> I mean, if you're spending tens or hundreds of millions of dollars a year on inference and you can cut your costs and improve your your quality and and reliability like that much, then >> then it's worth it. >> Like like it's it's 100% worth it. And and these companies do that and and they do it like more and more like this was this is kind of like a a frontier strategy. There's a handful of of companies who maybe a year ago were thinking like this and now like many many more are thinking like this.

46:14 >> Yeah. Yeah. That's cool. I have a question. It looks like it's a two-part question which we kind of talked about but maybe you can touch on more. So, how do you practice in this field? How do you practice in this field considering the huge cost of GPUs which maybe it's not so much about that but like what is your advice for people who want to become like you or who want to get into this industry?

46:36 Yeah. well, it's kind of a chicken and egg problem, right? Because once you get a job at any s sort of AI company, you should have basically effectively unlimited access to compute. but you need access to compute to get the skills to get the job at the AI company. So, like how does it how does it go? H how do you do it? the first thing is that like you do not need eight B200s to like practice setting up SG lang and practice like figuring out parallelism like you can run this stuff on A10s. you can you can go there's you know >> how much does that like if someone wanted to set up a super minimal home inference setup? Like what's the minimum cost for someone to get into that? I mean the minimum cost is like you can get Olama going on your laptop and like fi you know play around with that. I also think that like a lot of domains outside of language models are much cheaper like image generation comfy UI and and that whole like home image gen lab is like a big open source community for a reason because like a lot of consumer hardware can even run these models. There's like you can do I mean you can do audio and video on your oh sorry audio in audio out on your phone.

48:01 >> wow >> you can do you can do embedding models in a browser. You can do tokenizers in a browser. Like there's a lot of interesting stuff that you can build and learn. you can go like look at some of those like build a GPT from scratch type of courses which generally you're building like a GPT2 type model that you can again build and run on like an average consumer laptop and then there's also like a lot of cheap cloud GPUs out there Ampio series GPUs A10s like love lace L4s T4s from from even a generation or two before that like can >> so are those in lower demand, so they're easier to like person get one.

48:45 >> and they can often be rented for like well under a dollar an hour. >> so yeah, it's like it's definitely not a cheap industry to be in by any means. but there's a ton of educational resources out there and you know, again, like unless you're if you're at the point where you're like, "Oh man, the kernel I'm writing has to run on Blackwell because I need this specific tiling pattern, like you are well past the point where some company with Blackwell GPUs is going to be thrilled to you and give you as many as you want."

49:26 >> Yeah. What are like at base 10 like what are what are people looking for? What are the characteristics of like an engineer that you guys want to hire? Because I imagine like you hire some people who are already experts and maybe you hire some people that aren't but have the potential to become one, right? >> Yeah. You know, fortunately like we're in a position now where we we do get to to work with like pretty much premier experts in in whatever field we're interested in.

49:53 >> which was like, yeah, definitely not the case when I There's no way that the the guy I was four years ago would get hired at base 10 today because that guy >> you had very good timing there. >> I did. I did. But, you know, you don't have to be an expert in everything. Like we we recently hired some folks who are open source contributors to like some container frameworks and container libraries to help us with some of our application security stuff and to help us with some of our you know cold starts and and those sort of like infrastructure challenges. we hire a lot of out of water and who come in like in many cases with some experience from from companies like Nvidia and have you know worked on some of these model performance challenges before. but like more more than like credentials or you know experience like we look for sort of evidence of exceptional ability.

50:57 If you've done one amazing thing in your life before, that makes you much more likely to do another one. like a big part of my application to base 10 four years ago was that I had written a book and published it to like moderate commercial success and that sort of signaled that I was able to do a lot of other similarly difficult things until such time as I like did literally the same thing again. >> Yeah, you did the same thing. Yeah. But yeah, so so it's it's more about like have you done really cool stuff and do you know just one of the technologies that we care about exceptionally well and if those are both the case then like everything else can be learned with a you know good attitude and strong work ethic and like all the normal table stake stuff.

51:43 >> Yeah, that's something at Planet Scale we talk about we use the term a P99 engineer, right? which is like the 99th percentile. And it doesn't mean you like are the 99th percentile in terms of your your background or whatever, but it's like do you have that like drive to just like I'm going to learn things. I'm going to work hard. I'm going to like get jobs done excellently. Right. Yeah. >> so that's kind of the it's it's somewhat tangible but sort of semi- untangible that you have to figure out when you're interviewing somebody. for fortunately if if you're looking for a 99th percentile engineer let's say there are maybe 20 million software engineers in the world that means you have 200,000 to choose from that's actually quite a large >> a lot yeah >> and if you're out there and you want to be one then like you know it's one of the top 200 thousand is like not that hard to crack >> yeah if you start working now start learning working hard yeah you absolutely can be in that group okay let's see there's one more question we're getting close to the top of the hour so we'll wrap soon but I see this here What do you think about Rust inference libraries like candle burn tract? Do you see see a future for Rust in inference engineering? I can't answer that question at all, but I don't know if that makes a good question for you.

52:52 >> I've absolutely never heard of these guys. I very much in the Python plus C++ CUDA stack world. >> Okay. Yeah. I will say that like one thing I've been increasingly seeing and even experimenting with myself a little bit which is fun is like >> using AI coding tools to like write code in languages that I don't know because I am very much just like a run-of-the-mill Python and JavaScript type of developer. and I do know that there are languages out there like Rust and Go that have massively better performance characteristics. and are in many cases just held back by like a lack of general knowledge of the language. So I would certainly not be surprised. one one thing we we did release a client library actually for embeddings that leverages a lot of Rust stuff. so >> yeah to to to what the the bottleneck was like CP Python based clients were like CPUbound and couldn't actually like send >> requests fast enough to saturate the server. I'm sure actually that's something that you're familiar with as well.

54:08 >> Absolutely. Yeah. It's a great language but it's a slow language. Yeah. so yeah, I am I while I don't know about like these particular libraries, I do definitely imagine a future where like the best programming language for the job is able to be used more than just like the programming language that everyone kind of agrees on. Like I feel like Python and JavaScript are kind of like English in that like they're a reasonably functional language that is successful because of like large scale coordination more than like you know maybe a a more obscure language that has better poetic characteristics or is more information dense or whatever.

54:50 >> but just like doesn't have the the global coordination mechanism behind it. >> Yeah. Nice. Well, I think we will we'll wrap up here. do you want to plug your book again or leave the audience with anything? What do you want to say? >> I want to plug my book. I have been I've been chilling books my my entire July. >> show us the book. Show us the book. >> Here it is. It's called Infidence Engineering. It's got a very shiny cover. how do you get a copy?

55:19 you don't. I'm sold out. all of these are getting shipped out today, but I've got a lot more coming. so we've got we've got a big shipment coming in next week. you can go to b10.com/influence engineering. get on the wait list for a copy. we've got PDF and EPUBs. They're the really nice files. They've got like good layouts and stuff. so you can download those in the meantime. probably the the best way to get your hands on a physical copy is if you're going to be at GTC, come to the base 10 booth. That's going to be like the first large scale in person and giveaway that we're going to be doing.

55:56 but yeah, we're punching a lot of these things and we want to get them in people's hands. >> Nice. Sick. Okay. Well, thank you, Philip. Appreciate you being on and maybe we'll chat again sometime soon. Yeah. >> Yeah, absolutely. Thank you so much for having me. It was great chatting with you, Ben. >> Cool. Thanks, everybody. See you later.

Summary

The discussion centers around inference engineering, a field gaining importance as AI applications proliferate. Philip, the guest, shares insights from his new book on the topic, emphasizing the need for engineers to understand inference systems to effectively leverage AI technologies.

- Inference engineering is essential for AI applications, with a growing demand for skilled engineers.
- The book aims to educate engineers on the complexities of setting up and optimizing inference systems.
- Key challenges include GPU failures, infrastructure reliability, and the need for low-latency applications.
- Inference involves provisioning GPUs, loading models, and optimizing for speed and uptime.
- Parallelism in inference is crucial, with different strategies like tensor and expert parallelism affecting performance.
- Benchmarking models remains a significant challenge, as real-world performance may not align with theoretical evaluations.
- Engineers should focus on practical applications and performance techniques to improve model efficiency.
- Learning resources and cheaper hardware options exist for those looking to enter the field of inference engineering.

Questions Answered

What is the purpose of this live discussion?

The discussion is centered around inference engineering and the new book on the topic.

What is inference and how does it work?

Inference involves using a GPU to run a model quickly, which includes capacity provisioning and managing dependencies.

What are the reliability challenges faced in inference systems?

GPUs can fail frequently, affecting uptime, and cloud providers can experience outages, complicating reliability.

How did the speaker learn about inference engineering?

The speaker learned through hands-on experience at Base 10, starting before the AI boom.

How do companies optimize AI models for different tasks?

Companies may fine-tune multiple models for specific tasks to improve performance and reduce costs.

© transcribe · For agents Built with care and craft by Gokul Rajaram