transcribe

Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud

No Priors: AI, Machine Learning, Tech, & Startups · 42m · transcribed Jun 2026
More from No Priors: AI, Machine Learning, Tech, & Startups Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:05 Hi listeners. Today Al and I are here with Tu Hin the founder and CEO of Base 10, the AI inference cloud. We're here to talk about capacity constraints for AI compute, why inference is the last market, how the workload is changing, the open-source and perhaps multi-chip future, and what 30x scale in a year looks like. Tu Hin, welcome back. Hi. Good to see you. Thanks for having me. All right, you are in one of the uh craziest markets, AI inference. Uh it's very important, there's a lot going on. You guys have grown 30x over the last year.

0:42 And I think I can say you're expecting to do more than a billion dollars in revenue this year. Mhm. What's going on? Tell us about scale. Yeah. No, it's been it's been nuts. I I think what's happened over the last I'll only say 24 months, but this kind of keeps getting bigger and bigger, is that I think everyone is re- realizing that you can put AI everywhere. Um you have you have all these great options available from closed-source open-source models.

1:09 The open-source models have crossed some sort of chasm in terms of their baseline capability, and then I think RL techniques and post-training it's um for specialized models um has become mainstream enough, and you know, there's enough examples of it work of it working. The customers are realizing they can, you know, kind of own their inference um more and more, and what that's meant for us is um more, you know, the long-tail models coming through, customers in-housing a lot of that intelligence themselves, and as the application layer just gets, you know, bigger and bigger and bigger, and that's growing, we we are we are just someone index on that, and we've been around to be able to collect the demand.

1:55 There's an existential question in here that I think everybody is uh continually asking of does the independent application layer get to exist at all versus the labs? Like how do you You have to believe this. Why do you believe it? Yeah, look, I I I think it'd be it'd be a sad thing if it didn't exist in general and I think that's like my but you know, sa- sadness is fine. Um the >> [laughter] >> Sad all the time Uh yeah, sa- sadness is fine. Um but like that that that's not the reason why I think the application layer will exist. I think the application layer will exist for a number of reasons. One is because you know, I think this idea um that what it what is valuable to a company um is you know, the the user signal that they can gather that only they can gather.

2:42 Um and to the extent that that is encoded um in a model, I think a lot of their business will um be at risk. But to the to the extent that it is encoded in workflows, um that is where they will be able to develop mode. So, a good a good example of that is say a company like a Bridge where the clinician's edits of the notes and what they do with the notes after the fact in the um the thing that happens in um inside the EMR three steps down, you know, that becomes a workflow that only Can you explain what a Bridge does? A sorry, a a Bridge is a um ambient um scribe um that is used by physicians in um or like, you know, almost all hospitals in the US. I think it's lot of investor um Greg Shives is amazing. Great great company, great team, um great product.

3:30 Um and you know, they they've basically uh you know, got this very very deep integration into into hospitals into clinician workflows. And my argument would be here is that actually, you know, it's very very hard for um a frontier model company to be able to eat up where that cuz they just don't have access to that user signal. And what will happen over time is that the folks who have access to that user signal can start to post-train models on that reward signal and and start to get long long horizon agentic models running then. And I think to the extent that that is possible and that signal is differentiated and unique and and and is um somewhat um rare to to get access to, there will be an application layer. And I think you know, it's like support companies is another example of that where you know, a support a support task isn't one-shotted.

4:24 Usually at a company like Basecamp, when a ticket comes in, there's like what, like one, two, 10, 20 actions that get taken. And that is where, you know, someone can develop a specialized model. So there's almost two versions of this then. There's the new companies like a Bridge or Decagon or some of these other things that you mentioned that are doing these new types of applications that are using AI and they sell it to customers. The other is that enterprises building things in-house or building their own models. What proportion of the market today do you think is um these new application companies versus enterprises just adopting AI and how do you think that looks in a couple years?

4:57 >> Yeah, I I I think that's a that's a you know, we did I think you asked me the same question two years ago. I hate to be repetitive. The answer is just that it's crazy that the answer is still I think I think if you look by inference count, it'd be 99% the former. Yeah. Um and that is the kind of represents the scope of the opportunity here is that the majority of the market hasn't come online and and and um and added AI into this market.

5:26 >> of enterprise adoption is well ahead of us and I think that's one of the very exciting things about AI. Yeah, this is just so much still to come and people are underestimating that, I think. 100% and I you but but what's cool is that we're seeing the transition happen. Right? Before it was like, "Hey, are they are they using AI tools?" I don't think that was immediately obvious two years ago. I think that's obvious now that yes, they are.

5:44 Are they using closed-source model APIs? Um I think they're starting to get there. And then once you do that and then you kind of see what is possible, then comes the whole custom model adoption. I think that is all that is ahead of us today. So if the majority of your customer base today is as you described the the former like application companies AI natives the um fast-growing and some of them are at considerable scale now like the Abridge Cursor Open evidence. Open evidence is of the world. Uh what you know what do they what do they teach you? What does that push the company to do? How do you think about serving them versus evolving for the enterprise?

6:21 Yeah. Um I I think firstly like you just learn a lot by building with the companies the greatest scale doing the most interesting things. Um we we think of it two ways. Like I think there's like the the the most obvious way which is just build for the highest scale um you know most the customers that will push you the most from technologically and everything kind of will fall into play. I think that's the Stripe evolution as a company showed that which is like Stripe now like serves like so many enterprises but 12 years ago that wasn't the case. But they just built for the frontier and kind of went with them. Um and the second way we think about that this is to just think about building for companies that are serving enterprises. So yes, we don't serve the enterprise but our customers serve enterprises. Um Abridge serves enterprises Open Evidence Deck of Cards um all these Writer Gamma all these play all these companies serve enterprises in mass and what we actually get is like a translation of the requirements from them which is like you know they're like hey we need this sort of data retention. We need this type this web models need to be deployed. Um this is the types of GPUs or the latencies they're okay with. This is the model requirements from like a transparency perspective that they care about. And so I think that is actually the more nuanced answer is that if you listen to what their needs are, we actually get a full translation of what the enterprise would require. Like I would say that by serving companies like a bridge and Open Evidence, we're probably pretty well suited to go serve the health care system given that they are selling and latent health given that they are selling to them. How how much of a shift are you seeing in terms of the types of open source models that are being used?

7:59 And so I think we've seen an evolution where 2 3 years ago, I think the main thing was kind of Mistral and a few other things and then Meta kind of came along with Lama and then it kind of really shifted. >> Yeah. In terms of the best performing models are of Chinese origin in different ways. Do you see that sort of mix reflected in terms of what's being used by our customers? Yeah, I I I think customers, at least the customers we are serving, are very and these are like the fastest growing AI companies in the world that are very forward thinking. They they want to use that small. And they they are optimizing. I think there is there there are a there's a subset of tasks which I think is small today where people really start to start with cost.

8:38 Mhm. Um but everyone comes for capability first because that's really where the economic growth is being unlocked, where the value is being delivered, and then they optimize. And I think that's like we actually been you know, and so with that in mind, you know, you you you name like you name it everything from GPT-3 OS all the way to um Moonshot models to Deep Seek um to um Canopy with Orpheus which is like really good text to speech Mhm. um models um Customers generally want to use whatever is at the frontier. And and I think the um the difference is just being um I think we have a lot more visibility into how to run these and how to run these really well. And secondly, that they're good now. There've been a number of different concerns raised about the use of Chinese models, in particular security or is there something embedded in the models or, you know, Trojan horses or other things.

9:31 Um A, do you think there's any real concern there? And B, you know, people often talk about how there should be like US counterweights to this. From a geopolitical perspective, do you think that's something that's legitimate or something we should be worried about or how do you think about the Yeah, sort of origins of these models versus their uses? Yeah, look, I I I think these these models firstly are fantastic. They're amazing. We work with these teams. They're truly awesome. Um I'd say Look, I I don't it is hard for me it is hard for me to see and I I could be wrong, but like, you know, if I if I net- if I network bound these models that they're not magically, you know, going to be able to cross the network boundaries and to data is data and, you know, I don't and we I've never seen any real evidence um except from some very early models that I think people picked up on very quickly that there is some agenda or bias built in these. Um I do think that to some um to some extent is I I think there is importance to the US that we develop our own models. I think that that would be a massive loss if that there are five companies you know, five different labs in China that are creating open source models um and we're struggling to get one set up.

10:47 It's so it's necessary. Um I also think it's inevitable. Um and you know, like the the to Deep Deep The Deep Seek moment a year ago, um I remember someone saying to me and I thought I thought it was like a very well said, which is like and the world's changed a lot, but they said, "Hey, you know, we should kind of just forget Mhm. >> that this is a Chinese model. We should just act like this came from Mhm. from Meta and and build and build with that in mind."

11:13 >> Mhm. It's like, you know, I think you're kind of missing the forest from the trees. Like, there's there's two there's two scenarios, right? Either America does not ever come up with good open source models and there's probably a fundamental problem there or we will get there and we need to be ready for that world. Yeah, that makes sense. It's interesting because um you know, like you I I think it's very important for the US to have a strong open source footprint here.

11:36 Um at least for now it looks like effectively the Chinese government is subsidizing at least a large subset of these models. And that subsidy or surplus is effectively just being passed on to US enterprises who are adopting these models. In other words, it's a way for the Chinese government to effectively subsidize US enterprise Yeah. in an indirect manner and I think that's a little bit lost right now. Um but you know, it's always interesting to weigh that against some of the other concerns that are raised. I appreciate your your comments on this. Well, yeah, and I I think the concern also just there just becomes like what happened if we aren't able to Like if if it is fun like I I think if you think of the economics here, which is Deep Seek by most Deep Seek is a good a very good model. Mhm. Um you know, like and like you can argue whether it's at the absolute frontier or not, but like let's let's go back three months. Mhm. Ends there. Yeah. And so think about everything and we're doing a whole lot of things three months ago.

12:25 And so let's just think about that. Well, um you know, if it you could run Deep Seek probably 20% of the cost Mhm. of running of Anthropic models um in production um with comparable better latency, probably better reliability. If we don't have access to that intelligence in that form, I think it's just a massive loss. Um and as a country that we won't be able to innovate as fast because like the the cost of intelligence going down and control of intelligence, what we have seen just means more intelligence. Yeah.

12:54 Intelligence being embedded in more places. Yeah, an important note uh here that we didn't mention explicitly is that the state-of-the-art models, the ones that are most far ahead on the frontier, are actually still the closed source Anthropic, Open AI, Google, etc. >> Yeah. What has been actually maybe you can just characterize like workload a little bit, like how of tokens being served on Base 10, like how many of them are uh from custom models of some kind versus like vanilla open source today?

13:20 It is all custom. It's basically Okay. So like 95% plus. 95% like and I I think that's really cool to be honest. Like look what we have. We have two businesses that we have we have three business we have we have three businesses right now. And like we have no no dedicated dedicated inference which is basically custom on inference. Um your SLA is your SLA and we have shared inference which is shared inference at some point. Um shared SLAs um and then we have a training business. Um um I'd say 95% of the tokens today are on the the the first business and almost all of them um there's probably a for almost all of them the customers making some modifications to the model with their with their own data um specialized for the use case and I think what's even more important is um they might be compiling in different ways. No one is just running the vanilla open source weights. Like you you might be customizing it for quality but you must might be customizing it for performance.

14:22 You made an acquisition of a research team a few months ago. You've mentioned post training customization. Uh what was the rationale behind the acquisition? What is that team doing today? Yeah um so the the rationale around the acquisition was you know we we are infrastructure and product people. Um we are product people um and now are really good infrastructure people and the um and we didn't have much of a um research capability ourselves and and what we saw was um the market moving heavily and heavily like that we could accelerate the market itself um with post training um resources um either product title or or even just as resources for that market. Um so PaLM was a a company that was a based in customer. So they were post training models and running them on base on base 10. And I think what they realized was um that they would eventually need to become an inference company.

15:27 Um and what we realized was like, "Hey, we we really needed that expertise to because it is it represents a way for us to get closer to the customer earlier um and you know, be able to support them more." And it just made sense um as a hit like pairing them together and um just as as I said in the opening statement here which is, you know, as more and more post-training models have come up, we've realized that the the demand for people to for people to either um for software loops to do post-training or for post-training expertise is very high and then we we're really really investing um in that. Um there's also a bunch of Australians um you know, we I like to think that we had a bit a bit of alpha there. Um but yeah, that that's been fantastic. They're working with all sorts of customers um and it it's also very interesting when you when you um start, you know, we were doing a lot of research on the performance side and less so on the post-training side. Um it's interesting as we've started to do a lot more research on the post-training side, um you start to see how linked inference and post-training are. And like, you know, even e- even when you think about stuff like quantization and then when you should do that and like, you know, how how how how training um how how you train the model affects how you need to quantize for inference and how paired these problems are. Mhm. Um has become like very apparent and more and more we realize that the post-training and inference are kind of both sides of the same problem. So cuz inference will ideally will beget more post-training.

17:00 Where inference creates data, you do evals, you can now post-train post-train on that um on that reward function that you that you found with those evals and and hopefully this will play entirely out. Plenty of folks from uh Ant and OpenAI, Sam, Greg, etc. have said in recent months that like inference is super strategic, inference talent is strategic, capacity is strategic. Uh so, between that and post-training, these are uh very uh difficult to gather like capabilities.

17:30 >> Yeah. Um I imagine that lots of your customers go to you guys for advice on like how to do this progression of moving to custom models. Like, what do you tell people about the life cycle and when they should invest in that? Yeah, I I think it's hey, go find go to prove to yourself with the best-in-class model that you have something worth optimizing. Um and and I think, you know, a lot of you know, if a customer comes to us, you you you there there was what's that meme which was like it was like 2 years ago.

17:59 It feels like there's no GPUs pre-product market fit. It's like no post-training pre-product market fit as well. It it what I'd say that >> Some people that you're working with here are very very at scale first. >> Yeah, they they they they have they have a user signal that they know how to optimize. And they've shown that they can, you know, they can serve customer value and that value and that they have something special around that value. And once you have that value, it's like, "Okay, now how can I do that better, faster, and cheaper?" With the idea being that, "Hey, if you need to be very good at customer support, you can you maybe don't need to be that good at coding and that a specialized model might be a better fit for that problem and you can do it better, faster, and cheaper.

18:35 What about the capacity side? You uh started with, you know, unifying capacity across all the clouds and new clouds. How do you think about this when everybody keeps talking about a a supply crunch and a multi-year supply crunch? I think um you know, there there's so much narrative around the supply crunch. Um and no matter like as much as we hear about it, I don't think people realize how bad it really is. Like, there is you know, there's very very little slack compute available. Like, you know, we we run pretty large clusters ourselves and we run them in like uncomfortably high utilization. You know, we we're not saying we're like mid-90s utilization um most of the time. Um there is we have made we we have we sit in 18 different clouds now.

19:26 We have 90 clusters around the world across 18 different clouds and like, you know, initially we started we like built this technology to be able to like kind of create one runtime fabric that spans all these different clouds and try to abstract that away from our customers um as a way to think about reliability, latency, failover, all these things that we think are going to be very important for very mission-critical use cases. Um that same technology like just our ability to get compute wherever humanly possible um has been really really helpful in our ability to get supply and and what what I mean by that is we can be introduced to a new provider in a different country um and have it up and running with the whole base 10 inference stack as part of the fabric. Part of the fabric in I don't know, half a day?

20:14 Half a maybe less. Uh um even for and that gives us enormous flexibility. Um even for us it is hard for us to grow. Like we have a we have a I think it's yeah, I can't say. Um we have a a 4:00 p.m. standing meeting for the company where we basically go like like how do we like how do we how do we manage capacity for the demand right now. I think the second part which people don't really um the two the second part that people don't really understand is um that there are also a lot of um suppliers right now um that's it's kind of grifty.

20:57 You know, like the I I I think, you know, they haven't run they haven't run data centers before. You know, they don't understand SLAs especially for inference. Um and so like, you know, even when there is capacity available, um there's a lot of dough like there's probably we we run a lot more than this than we've ever done in Seattle, so it's fine. But if you, you know, there's probably like a dozen good like clouds and I probably like put like three or four of them in like the the gold tier. Mhm. Um and I think that just means that like supply like not only are we supply crunched, we're supplier and operationally crunched onto people who can who can run these data centers as well. How how far ahead can you actually buy capacity right now? In other words, like is there any any slack in the market if you buy 2 years ahead or 5 years ahead, you know, what you mean like actually like contract length or actually the K I want this in January 2028.

21:54 >> Yeah, either one, yeah. I mean it's more the I want this in January 2028 or at least I have some visibility into my future supply. Yeah. Um you could buy that, but you got to also remember how quickly the market is how quickly the market is moving. Um and like, you know, that gets balanced somewhat off like the fact that the H100 is such a great chip. Yeah. Um and like and then so, you know, it's crazy it's 4 years 4 and 1/2 years old, the price is going up still. Yeah.

22:20 Maybe it has a useful life of 9 years. Yeah. Um so, um you know, that that's that's good, but at the same time at the same time, you know, yes, you can do that. Um but, you know, you're making a lot like you're making a lot of bets. Yeah. Um as part of that. And then in terms of I think that's the big thing that's changed over the last 6 months is that the term length that people want um has just gone up. So, like if you if you wanted um a thou- a thousand thousand 24 B200s, Mhm. um which is, you know, um from a good cloud right now, you're not getting that less than a 3 to 5-year contract.

23:01 Right now with a probably 20% TC TCV prepaid. Um so like actually what becomes important when acquiring capacity um is you need to have enough demand to supply it so then you also need like a low cost of capital which is which is actually changing the dynamic pretty significantly. Does that Does that impact how you think about going public as a company because arguably Yeah. I think you'd go sooner. Yeah, exactly. Yeah, I think you need like I think the I think there was demand for that.

23:31 Um but I think you know the pool the It will It also you know one one of our one of the one of the realization that we had recently and with with software people um and so we don't we don't think like this all the time is that you know our our business has like very interesting working capital. Mhm. Um requirements you know and and I and I think even and that as a result of that it has very interesting financing Yeah, yeah. um requirements and we're not at least right now we're not even going down to the Yeah, there's also also things you could do in terms of debt or other structures that yeah.

24:06 Yeah. And yeah, I've learned a lot about debt Yeah. um recently. Given the supply crunch Inference being one of you know the top couple markets you're going after you have plenty of people who understand this problem and therefore you know some competition. How do you How do you think about like what are the factors that create a dominant player here or a winning player? Is it as you mentioned cost of capital? Is it access to supply? Is it software? Is it demand?

24:36 Yeah. Just being excellent at everything. Yeah, it's um Look I I I think what's so interesting about Inference is Is it operations? Like a special cloud? I think so. Yeah, I think like GPUs as a service is not sticky. I think that's been seen. Like customers generally just see that as as commodity. Um Inference with the software layer included is incredibly sticky. Um, you know, like just just like you know, none of our top 30 customers have ever turned. You know, we're talking like 400% annual NDR. Mhm. Uh uh um around our business. And so, it's like very it's um it's very very sticky.

25:15 So, I think that software layer is very important. The optimist in me is like, "Oh, there's so much value in the software." And like I we will build the best software layer for inference that exists, I think. You know, as I think it's becoming clear now access to inference compute is Yeah. is a strategic strategic advantage. And I think that is like the I think that is the um strategy that even the labs are going after, which is like, "If we have the If If we have all the compute, good luck running inference." Yeah, yeah. In in a world of constrained compute, the number one thing to own is compute. Yeah.

25:48 >> And so, you know, just owning it in and of itself is an asset, and I think people under appreciate that. >> Yeah, you can't you can't make a good hot chocolate without milk. You You know what I mean? They aren't. >> [laughter] >> Unless you're vegan. Unless you're vegan. So, no one wants the vegan inference. So. >> [laughter] >> Well, I got to ask you, people might want um they want they might want alternative milk, right? So, like when you the H100 is a great chip. People, you know, want a B200, they want GB200, they want, of course, tons and tons of Nvidia.

26:16 Um when you think about making a bet, you know, several years in the future, do you believe that there's a like multi-chip world? Like what do you What do you think happens from a compute perspective? Um on the chip side. Yeah. Um I think I think you know, like diversification everywhere is a Mhm. Same way I want a world of many models. I think, you know, we want a world of many most things. I don't know. I think It'd be sad if it didn't happen. Yeah, and and and I I think I think everyone would be sad. I I I will say um to some extent, which is um Yeah, and I think there will be inference-specific chips. I think you have like decode specific chips. I think and we're looking at it.

26:56 Yeah, yeah, I mean that that was a whole crock. LP thing. It's like you know, I I think I think that is um very straightforward and and makes sense. I think people really, really, really underestimate supply chain stuff with Nvidia. Like how good are they are at that. Cuda, how good Cuda is. The developer ecosystem around it. Um and you know, we it the ability like to me like one of the most important things as an infrastructure company in this moment is how fast you can move and you can move fastest with Nvidia today.

27:30 Um and I think that is the reality and like it just like given the scale that they operate at. Given the scale that they operate at, it's um it's hard to it's hard to see um a it's hard to see the the I'm not I'm not saying it won't happen like the short-term like in the next couple years how anyone can be able to compete with that. Espe- especially with you know, so much of the other the other players. Like what you need um to be able to compete here is the ecosystem to form around you. And if you tie up all your supply with one buyer, which you know, a bunch of the other chip providers have done, it's actually hard for that ecosystem to form. You know, like if you if you think about if if you're a big lab and you have a proprietary deal with one chip type where you get 90% of the supply, it's actually in your best interest to make sure you get 95% of supply and everything just get built for you and no one else can ever use it. When you think about reacting to the market, um what do you think is like happening with the actual workloads that you have to go invest in, right? Like obviously code agents and long horizon agents over time have become a big deal. People talk a lot more about CPU compute. Video inference is different. Um I don't know if it's that. Sandbox is like what what's important for you guys to invest in now? Yeah, look I I I think this for for us all the runtime stuff is obviously very important and what that means is like what chips we run on, how we run, what kind of workload we support, like do we get very good at diffusion transformers?

28:54 Yes. Um coding agents need sandboxes, which could be called sandboxes. Um there's all sorts of new speculation techniques to get faster inference. We need to do that. Um even stuff like um KB cache aware routing and you know, that stuff is a bit old now, but like getting continuing to be very good at that and um somewhat disentangling prefill and decode and starting to treat them as separate problems. I think that's you know, something we're very focused on and we're seeing massive gains there.

29:22 That's at the runtime level. Um I'd say you know, beyond that, you know, everything we think about is how to create more of that loop between inference post training because we think that just begets more inference. Um and so like we will we will build a partner in almost everything. There. So like you know, we're going to work with, you know, the best eval companies in the world to make sure that's very well well integrated like Brain Trust um into and around Base 10, you know, we will partner with all on the sandbox side build build the best sandboxes experience um that will exist. Um and then we'll create the the best training APIs to make it so continual learning becomes somewhat of a solved problem.

30:06 It's not just like a discrete thing. That's I think the core Base 10 product thesis is like how do we build that loop and then everything after after around that becomes how do we make sure that we can do everything we can to um ensure that gets the biggest possible. That's access to compute. That's our infrastructure. Make sure we can get compute anywhere. Make sure we have access to our own compute. Um and then I think it's all the primitives that come after that just that just become incredibly like margin accretive both for us and our customers um which is, you know, stuff like you know, sandboxes and like the um async batch inference. Like how do we drive utilization by having a first-class batch inference experience. To me, this is like what an inference cloud looks like. It's that you are very good at inference and then you you start to do all the things tangential or that loop into inference and partner in when necessary and build when necessary. Um but we really do want to own like start with our core inference story and then go down to unlock supply and create margin and go up the stack to unlock value. What would surprise people about some of the issues you discover only at scale? I'll give you an example. I was surprised when you guys ran into scale limitations, like fundamental limitations with some of the hyperscaler products that you were consuming. And I because I kind of think of, you know, the AWS GCPs of the world as supporting infinite scale.

31:31 >> Yeah. I mean, I I think you just And like again, like I think very, very large companies will that run services at big scale is probably the same stuff. Mhm. >> The all the edge cases um just become >> actually experience them. >> You experience them. And like, you know, and you I like I'll give you a few examples here. Like you you see you know, you start seeing, you know, yesterday we had for the first time ever we saw some kernel panic. Um that only happened because some um uh fluent bit worker was creating too many logs and and the scale was too big and it was all into one node and it was happening to two two times at the at the same time by two different workers. Um so you see all like the systems level and kernel level problems.

32:15 Um but then you start to see I think with the craziest stuff is that you start to see with with LLMs um that these runtimes are pretty immature. Even how we use KB cache is, you know, um you know, probably a little less sophisticated than most people see than most people see. You know, and we we we are starting to see the the limitations of the current and the next set of primitives that need to be built from a scale security or performance perspective. But I I think it's really at the runtime level and the systems level and then but the edge cases are I'd say a lot more systems level than they are L M specific.

32:54 What are the things that keep you up at night? Capacity. Um I think I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I I >> Quick answer. Yeah. Yeah, I think capacity. I I I I I I I I I I I I I I I I I I I I I I I I I I I I I Probably just this market so big and and so like it it represents um a moment when you should be as aggressive as possible.

33:12 Um and you know, really you know, we've we've grown a ton of this year the last 12 months the last few months. But the answer is always just go you know, go bigger go faster and I think that's really really fun. It's also a little exhausting and it's also like we are we are all in somewhat uncharted territory in terms of how fast and how big you can go and how things can get. But I I I I think the big one is compute. I think like there's no world in which there's enough compute to you know, get the amount of the the amount of value that we want to get out of L M's in the next 5 to 10 years. Or we have to invent a lot of new stuff.

33:47 >> Yeah. To me, yeah. Maybe if we just talk a little bit about uh what you're learning scaling, you know, 30 X is like an aggressive thing to go through as a company. Uh you've brought in a lot of um really amazing talent like um Danny and Samir and Stephen Dave, folks on both the the technical and the um go-to-market side. Like what do you what do you think is working about how you are recruiting and scaling or what's your philosophy on that? We were very very flat like until I don't know, 8 to 12 to 18 months ago.

34:24 I remember when I walked with a lot actually. And a lot of it's like you just need leaders. And I and like it's actually like so contrary to um everything, you know, and the engineers you're like, "Oh, It's all overhead. It's all It's all Everything is overhead. >> [laughter] >> Everything is overhead. Um and I and I You once told me, I think, that you you didn't You're like, "Hey, Sarah Sarah, what about we just have engineers instead of sales people?"

34:47 >> Yeah. Yeah. Yeah. Bad. >> [laughter] >> Everybody learns that at the same time. And we all we all we all learned that but I remember like, you know, you you you said it so clearly at the time a lot and I and I think that's what we noticed was just like actually having a leadership team um that you can trust, that you can trust um is is is so important. I think I think the the two or three things that I'll say is like, you want people where you can give them whole problems.

35:14 And so like, you know, if if if you are if you feel like you are micromanaging, if you feel like you need if you feel like, you know, you you have to be involved in everything, I think that's a bit of a cop-out as a founder cuz you're just like, "I just need to be involved in everything." It's like, "No, you you probably just don't have the right people." Um I think the second thing is um be very very clear what you're optimizing for as a cuz I think when you're very very clear what you're optimizing for, the people on and like if it's something generic like, "We want the smartest, hard-working people."

35:44 Like, you can't do much with that. Like, what that's what we cared about was, "Hey, actually we don't care about a lot of people who've done this before. We care about first prin- people who think from first principle prin- first principles. Work has to be um a high priority, but they also have to be very kind and nice and you know, care about the collaborative environment. We don't have a hero culture. Um you know, very low ego.

36:06 Um and you know, if you need if you need a manager, like um it's probably not >> [laughter] >> it's probably not the right place um to be, but I think once you have when you have that clear rubric, the the people become very apparent that will fit into it and the people that don't um fit into it also become very apparent. I think what's more like we've hired amazing people like you mentioned, but I think what's a lot more interesting is like I think we've we haven't had a ton of like turnover there unnecessarily. Like people tend to work cuz we cuz we have a very we have very clear on what we want. It it took us a while to get there though.

36:44 What about the idea of like an operations culture? And we were talking to Alyssa and Henry about this and she's like well the hard thing about cloud is actually just operations. I slept with a pager under my pillow for a decade. I don't think I've seen you detached from your slack channel Yeah. >> for My phone is buzzing right now. The um the um the as I I'm getting I'm I'm I'm getting anxious. [clears throat] So um >> And and you've you've been concerned before like do people get it? Like it you know, what is distinctive about that? I I think I think like one, I think if you have worked at an infrastructure company like we were once in a meeting with a bunch of AWS execs and this was you know, like very senior AWS folks all of their pages went off multiple times um during our 45-minute meeting. You know, like it's a I I I think like it's it's it's very much like it's a cultural thing. Um but yeah, like I I don't you know, our like infrastructure can't go down and like you know, we you know, you you you you learn to like you know, what was it like I think me and my co-founder when his pager goes off his 7-year-old said is that a P0?

37:51 >> [laughter] >> Oh, is that is that is that a P0? And so you know, I I I think that is you just have to get used to it and that's the culture you live in and it just changes the speed. Um but it also it's it's you know, becomes like a you know, a cultural thing. I think it very very it rejects people that don't fit into it very very quickly. Like engineers who avoid pager Yeah, you know, when we when we have pagers like everyone on the call.

38:15 Like you know, like there's been a joke that they may as well be a siren that goes off in the office. And when they when when the siren hits so So people have been talking ad nauseam in the AI community about Jevons paradox. Um, where if you decrease the cost of it's a it's really a question around price elasticity and availability. If you decrease the cost of a good, say intelligence as a good, um, people actually consume more of it.

38:41 Um, like the personal or business ROI of it, the demand for it goes up, not down. Um, do you see this and are you are you working against yourself to make these models more efficient and people just use them more or less? Yeah, I I think you have to think about this from a developer perspective and a consumer perspective. I think like I think consumers just want the best answers and the and the and the best experience that's somewhat um, governed by yeah, more intelligence to some extent.

39:08 I think when you go to the developers from the developers perspective, um, they would insert more intelligence if you make it cheaper. Like that's yeah, and they would they would they would insert more intelligence anyway, but if you make it more cheaper they'll they'll insert a hell of a lot more intelligence. You see this with agents and stuff. Agents are just longer running now, and I think that's what we have seen with the cost of inference going down, which is, you know, folks are just like, "Okay, we can we can run this for longer. We can make it do a bit more work and we'll get to a um, a larger end." I think like if a compute scales from an inference perspective as well, um, and you know, I think we are seeing that with almost all our customers, which is you know, they either they either start with like this is the quality of inference I need to get I need to get to and this is the amount of inference like I need to do to get that or this is the base level model that I can start with or that I can work with to get there and I think the more we drive down the cost, um, what they realize is um, more intelligence just means better user experience.

40:11 >> want a better answer. Better answers, better experiences, more dollars, more dollars, even more revenue, so yeah, I think I think inference going down just begets more. But I I I it it is truly like I think we're kind of in a world that is you know, it is the last market, right? Like even if there's AGI, all that's left is inference. Yeah. So, you do not see in your customers a uh like this this answer is enough and this action is enough dynamic. No. No. Yeah, it's going to keep going for a long time, it looks like. Yeah. How do you view all this kind of evolving towards the future? So, basically, this is one of the It seems like it's going to be one of the biggest markets of all times.

40:47 We have this massive shift where we're moving from software and seats and digitalization into actual intelligence, selling units of cognition. Yep. Selling agentic workflows. What does this all look like in a couple years? Like what is your view of this future world? I think for the for consumers, it's the best possible thing, right? Like every everything is somewhat smarter, you know, your doc you get better care cuz your doctors have access to better um better tools. Um there's more, you know, like there's all this stuff about there being less software engineers. I think we just build more software. Mhm. I think we just build a ton more software and like, you know, I I see, you know, we're not slowing down hiring of software engineers, we're just building more things. Um and in that for the consumers, that just means better tools, more software. Um all those all those good things.

41:34 >> It's almost like everybody has their own team for everything, right? You have an agent which helps with your doctor. You have an agent that helps you learn stuff. You have an agent that helps you organize your life. >> It's concierge. It's concierge everything. Yeah, concierge everything for everyone. Yeah, and and and and I think like what that means what that that's amazing. I think that's great. And I think that the the and the education same thing, you have concierge education. Like you get personalized access to everything. I think then you go one step back in how it um affects developers, I think, you know, um and and companies, I think if you don't embrace this, Mhm. I think it's the extinc- extinction moment Mhm. for for a bunch of folks, which is like, you know, you everything needs and I I don't think that means that you know, Ford design needs to figure out. I don't think that's a thing. I I think like what what's more what's more interesting is just like, you know, all these workflow and software companies need to figure out what is the intelligent or intelligent inserted versions that that drive the amount the all that user value for those and consumers that we talked about. Yeah, very exciting.

42:34 Thank you so much for joining us today. Yeah, thanks, guys. Find us on Twitter at No Priors Pod. >> [music] >> Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts [music] for every episode at no-priors.com.

Summary

Tu Hin, founder and CEO of Base 10, discusses the rapid growth and evolving landscape of AI inference, emphasizing the importance of custom models and the challenges of capacity constraints in the AI compute market. He highlights the shift towards in-house AI solutions by enterprises and the necessity for companies to adapt to the changing demands of AI workloads.

- Base 10 has experienced a 30x growth in the past year, with expectations of over a billion dollars in revenue.
- The AI inference market is evolving, with a significant shift towards open-source models and in-house AI solutions.
- Custom models are becoming the norm, with 95% of tokens served on Base 10 being customized by clients.
- The demand for AI is expected to grow as costs decrease, leading to increased consumption and more intelligent applications.
- Capacity constraints in AI compute are significant, with high utilization rates and limited slack available in the market.
- The integration of post-training techniques is crucial for optimizing AI models and enhancing performance.
- The future of AI will likely involve personalized, concierge-like services for consumers, while businesses must adapt to remain competitive.
- Base 10's strategy focuses on building a robust software layer for inference, ensuring reliability and performance across various cloud providers.
© transcribe · For agents Built with care and craft by Gokul Rajaram