transcribe

Baseten Founder Spotlight | Tuhin Srivastava

Altimeter Capital · 37m · transcribed Jul 2026
More from Altimeter Capital Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Introduction to Base10's Journey

What led to the founding of Base10 and its evolution?

Base10 was founded in 2019 with a focus on machine learning, initially exploring various ideas before settling on building tools for the machine learning trend. The founders aimed to create a 'picks and shovels' business to support this growing field.

  • Base10 started with a diverse set of ideas before focusing on machine learning tools.
  • The founders believed in the potential of machine learning as a significant trend.
  • The company's evolution involved deep exploration of ideas to find a viable business model.
# 7:25

The Role of Inference in Base10's Strategy

How does Base10 prioritize inference in its offerings?

Base10's strategy centers around enhancing inference capabilities. The company aims to own more of the model training and post-training processes to improve inference outcomes, recognizing a gap in effective post-training tools.

  • Inference is the core focus of Base10's business strategy.
  • The company aims to integrate training and post-training processes to enhance model performance.
  • Base10 has identified a lack of quality post-training tools in the market.
# 14:51

Components of Model Serving Infrastructure

What are the key components necessary for serving models effectively?

Effective model serving requires fast performance, robust infrastructure, and a positive developer experience. This includes managing compute resources, optimizing deployment, and ensuring observability and monitoring.

  • Fast performance and reliable infrastructure are critical for serving models.
  • Developer experience plays a significant role in the successful deployment of models.
  • Observability and monitoring are essential for maintaining model performance in production.
# 22:17

Understanding Model Workloads

What distinguishes the workloads handled by different models?

Different models handle various types of workloads, with the frontier model managing more complex tasks like synthesis, while open-source models can handle simpler tasks. The synthesis layer often requires more advanced models for final outputs.

  • Workloads vary significantly between models, with frontier models handling more complex tasks.
  • Open-source models can manage simpler tasks effectively.
  • The synthesis process often necessitates advanced models for optimal results.
# 29:43

Inference vs. Training Infrastructure

What are the differences between infrastructure designed for training versus inference?

Infrastructure for inference requires clean data lines and effective partitioning to prevent interference from training jobs, while training infrastructure can tolerate more noise due to its batch processing nature.

  • Inference infrastructure needs clean data lines to ensure performance.
  • Training infrastructure can handle more noise and data influx at once.
  • Effective partitioning is crucial for maintaining inference performance in shared environments.

Transcript

0:00 You can run, you know, 3month-old intelligence at 20% of the cost. One of our customers said they were using open code with the GLM 5.2, but it was 15% of Frontier. >> So significant reduction in spend. And with our most sophisticated customers is that it looks like something between 30 and 50%. Goes to open source and closed source models. >> Duhan, it's great to have you. Every time we chat, I learn a lot. Glad that we're going to open source this conversation. Yeah, amazing. Great to be here.

0:31 >> There's a lot going on in your world, AI, inference, open source, a lot to chop up. but before we get into it, I thought I'd ask you about the scenic route to what B10 is today. Bin obviously started in 2019. Yeah, >> you've had lots of scenic stops along the way. Tell us through them and and and what got you here. >> We started the company with a very different mindset. I'd say it's like you know we were we done a bunch of machine learning stuff.

0:57 >> Mhm. And not many people know this but the way we started the company which you know may be inspiring maybe uninspiring everyone can decide which is that we created a list of things we wanted to work on >> and then you know the then we would go and basically go really really deep on this each of these ideas and come back discuss and be like all right should we keep going or not? >> Mhm. And we very quickly found what became base 10. You know, we start when we started the company, we'd been doing a lot of work in machine learning from 2012 to 2019. And the idea really was how do we start a pix and shovels business along some large secular trend.

1:43 >> Mh. >> And you know, we we were pretty convinced the machine learning was going to be a big deal. We didn't know what it would look like. So like all right let's just start building tools for you know those people who were doing machine learning. The people doing machine learning in 2019 were then weren't like the people who were messing around with models today. There were a lot of like data scientists >> working on like classification systems and recommendation systems >> and content moderation systems. and so we were building tools for those folks and funnily enough we were building serving as part of that and alongside a bunch of other things. serving small models is actually trivial >> because you can basically you know load them and run them in memory and they run on CPUs. It's very quick.

2:31 They're tiny. We we would go around saying which is funny now in hindsight we'd go around saying well model serving and is a commodity. There's nothing there. we'll give it away for free. and and so that was really the genesis, but I think the bigger genesis was, you know, spending time with my two co-founders, Amir and Phil, and then Punkage, who was, you know, we found right after that was like we just want to hang out and build some cool things together. and from 2019 to 2012, I'd say like not much was happening. you know, we we we'd go and talk to a bunch of customers and everyone was pretty exploratory.

3:14 All these use cases were predominantly internal. >> Mhm. >> And so there wasn't much budget behind it. And I I still remember this moment in 2021. We went to talk to the head of AI or the head of innovation at a financial services company and we asked and we were like, "All right, so what do you want to do with AI?" and and he and I remember he was like couldn't he didn't really have an answer and he just said but I was like okay we're really early if we're really in the buzz word territory but then I think in 2022 the market start to shift a bit and I think a few key moments come to mind I think obviously like stable diffusion >> was a big moment whisper was a big moment >> transcription >> transcription And honestly, Chad GPD was a big moment just in general because I think it set the standard very very early >> for the the experience that developers and consumers would come to expect of AI.

4:18 >> And so, you know, like even small things, right, like you know, the AI blinks at you when it's thinking, >> you know, like that was, you know, it was clearly something that had been designed with the user in mind. And that was really the the moment that everything started, you know, moving as an industry. and then it kind of just kept like the breakthroughs just kept happening. not many people know this also was like the first time where it felt like inference was a market was we had a buddy called Seth who started this company called Refusion. And what he had done was him and his his co-founder Hike they had fine-tuned stable diffusion to generate histograms like m music sorry spectrograms music spectrograms and so they created like the first text to music model and they came to us and he'd been working on this for six months. This is like November 22. And he said, "Hey, I've got this I've got this model that I'm going to put out tomorrow." He's like, "I'm going to need a lot. It's going to pop. I'm pretty sure it's going to pop." I'm like, "Sure." Like, "We're going to need a lot of GPUs." We're like, "Yeah, what?

5:30 Like five 10 ATS at the time." And then you know it went to the top of hacking news and it felt like a lot of the time but it ended up needing like 40 or 50 >> ATGs to serve the traffic which is very small today but it was very big. It felt very big at the time and we ended up spending like 48 hours >> like kind of readjusting all our infrastructure to be able to service it.

5:53 And that's kind of when it felt like our inference towards a proper market because this would be if you believe there were going to be a lot more such challenges >> and a lot more such like consumer enterprise experiences a lot more people would need all that help >> being able to serve the traffic and that's when I'd say like >> this version of base 10 was really formed and that was the end of 22. >> End of 22. Yeah. Say the name of the business one more time.

6:15 >> It's called Refusion. It became Producer AI. It was acquired by Google maybe like three months ago, four months ago. >> So end of 2022. So this is December 2022. Chad GPD's come out. It's been like one or two months. >> Things are blowing up. The OpenAI Microsoft deals just happened or happening maybe >> based in back then was you know you guys had to make a decision to pivot to inference. >> Mhm. >> Actually the the other big insight is your first signal was not a LLM model.

6:42 This sounds like text to music. Is that a diffusion model? >> A diffusion mod because all stable diffusion cuz stable diffusion was open. So the entire ecosystem was being built >> around table fusion. That was the first thing. Also >> I I wouldn't say it was a pivot. I'd say it was more just like a focus a refocusing on inference which like we already did inference. >> We just thought it wasn't interesting >> and it all of a sudden became very very interesting >> and then I think as LM got really like got better over the time as the llama 3 moment >> happened you know all these things became really important >> for us.

7:15 >> Fascinating. You are now synonymous with the inference cloud. but shocker, you're doing more than inference, you're also doing training. >> Yep. >> talk about that for a second. Where does training fit in the portfolio? What led you to it? And why does the independence vertical independence matter for the customers that we serve? >> The we the way we think about everything at base 10 is >> is as a function of inference. So like you know the northstar of the company >> is how do we power more inference >> over time >> and so >> all roads must lead to inference.

7:51 >> Yeah. So even to what we realized over time though is like the sure there's all these open source models and custom models >> but you know what those models create is you know real user data >> that you can then get some signal out of >> that then you can in turn use to train better models that ideally you would then go and do inference on >> on base 10. Mhm. >> And two things that we realized, one is like that that if we the more and more that we could own of that loop >> Mhm.

8:22 >> the better. And secondly that there weren't that many great post-training tools. >> And so as we started to think about that loop, we're like, okay, let's own more and more of them. And that's when we really created you know, we first launched our training product about a year ago >> based in training. And then we acquired a company in November or December could pass through a post- training basically a post- training Neolab. >> Mhm. >> More or less. And now we have a product called loops which you know allows you to do post training you know manages rollouts allows you to do all this work >> but also all all of it is meant so it feeds back into inference >> on base 10.

9:00 >> That that makes sense. And >> you know launching training puts you arguably close to the labs themselves. how do you think about that competitive line? at what point should should a customer come to you guys for for their post training versus versus go to go to a lab? >> Yeah, it's a good question. I look I I I think it it does and it doesn't, right? and I'd say it does in the sense that it allows all of a sudden, you know, customers have control over their models.

9:33 >> Yeah. So it is competitive in that before you were renting intelligence and now you're owning intelligence and you know if all owners are competing sure I I still think that a lot of these like most of our customers do and should start at you know using the easiest thing that they can possibly use >> which a lot of the times >> is a close source model >> right >> because it comes with the entire ecosystem around and we are building out that ecosystem through our model API product >> and like that will become more and more fleshed out >> over time, especially as open source models continue to catch up, there'll be like real viable alternatives right out the gate.

10:10 >> Got it. >> To use, but a what I always say to go as a customer until you have product market fit, >> right? >> Figure out how to use the easiest thing to use. >> Yeah. >> They they were saying that, you know, a lot of it's about like they right now they're spending time collecting collecting water by the riverbank, but there's a waterfall over there. It's like go to the waterfall. >> Yeah. collect the water and then yes everything will break.

10:35 >> Yeah. >> Once you're collecting that water from the waterfall but then you can start to optimize. >> Got it. Got it. So go to the clo the the API for discovery for figuring out your use case for iterating >> even value. Yeah. Yeah. >> But then once you have product market fit and the waterfall is coming down your bucket and it's overflowing. >> Yes. >> Then you show up to us >> 100%. And I took that from someone today. I don't remember who it was.

10:58 but I I will find it and we can credit in the comments. >> Perfect. Perfect. >> The but but yes, that's right. Which is, you know, go find product market fit, then start to optimize. >> Yeah. >> And some of that optimization might even mean not using open source models or custom models. It might just be like go figure out how to use those other APIs in better ways or use cheaper models from from there them. But a lot of the time when you need specialized in intelligence and more importantly when you need control and want independence >> that's when it will make sense for you to switch over >> to these models that are you know being run hosted and post-trained on base 10.

11:34 >> Is using your models or owning your intelligence kind of a power user use case or and and and and the vast majority of customers will will stay on the bank collecting the water. >> no I don't think so. I I think I think it's a it's not a power user. I think it's just like a look at some point like what do you have as a company, right? So you have you have your data, you have your workflows, you have your UI, >> you have your reward signal and so you know what are the things that >> you and your customers like. At some point like if I if I am one of these companies, >> I want to hold on to that with dear life. That is all the IP I have left. M >> and so at some point you know and this is one of the challenges I think that a lot of these a lot of companies have is that first you start by quering the APIs and just using them >> in path to production. Soon you just start giving them all your data so you can get the best possible answer >> out of them.

12:32 >> Yeah. In the context or whatever >> in the context eventually maybe you even start giving them all the reward as well and try to figure out how to encode that in your in your in your in your context as well. And then >> as a result what will happen is two things. One you know you're giving >> someone else's API someone else all the data they >> they potentially need to that you find to be near and dear to you but you're also adding dependence yourself >> more and more dependence yourself on those APIs.

13:00 >> Mh. >> So in a lot of ways you are eroding away both your kind of customer data mo >> Yeah. >> and your financial mode as well. Mhm. >> And so at some point at scale you need to move away from I talked to a company that was you know at scale recently. This was an atscale company that said kind of hit their >> gross margin by 20 25 point 25 to 30 points. >> Wow.

13:25 >> That doesn't have to be the case. You know and like that's when I think a lot of companies start to think about >> how do how do we do this ourselves? How do we post train models ourselves? How do we >> Yeah. >> own more and more of the stack. >> Yeah. >> so you don't have to give all the data away. >> Yeah. Yeah. >> And so and you can also hold on to the business edge you have >> with these models.

13:44 >> So now if I'm a if let's say I'm a customer I'm you know I graduated from being two guys in a laptop and a dream to a product market fit use case that is the the waterfalls flowing. Explain to us in simple words what the counterfactual looks like. because in the mind of many it's a it's too much work. >> Yeah. What would that process look like in in in simple words having to own your intelligence and and maintaining your remote?

14:09 >> Well, look, I I think there's like three two or three things that you need to have one is that you have to have a very good understanding of what is good. >> Mhm. >> So, like you know, and this this is a problem that we have we see with a lot of our customers too. is that you know they're like, "Hey, can you help me post train a model?" I'm like, "Sure." >> But we can't tell you what your users like. That's like ideally what you are good at, >> right? Look, if if I'm creating an app, like I know I know the utility >> that my users care about. I have the data on that. So, one, >> you know, you have to have a good understanding of the of the thing you were trying to maximum.

14:43 >> Yeah. >> The maximum. >> But once you have that, it's actually not so bad. So, like, you know, one, you obviously need to have a bunch of data. >> You need to have >> a way to evaluate how a model is doing. M >> and then like you know through products like ours >> a product called loops it's actually pretty straightforward to get started with >> recipes for post training models >> and so we have a loop SDK we have a forward deployed function that can help you >> there as well but once you have that >> Mhm. You're actually also in in pretty short order. You'll have the ability to have a model that you can run yourself >> and that's when you need serving infrastructure. And so serving infrastructure is you know comes in many forms.

15:24 >> What we believe is that it's kind of like the three or four different key components to serving models. >> One is the ability to serve them very fast. So you you need to have some knowhow in terms of what it takes to you know get you know choose there's like kind of two variables here which is how how much throughput do you have >> and how much speed do you have and like you know our goal is to move that curve as >> as far to the top right as possible and then you choose where you sit on top of that. That's the performance piece. The second piece is >> infrastructure. And so that's not only just having access to compute, >> but also knowing how to wrangle that compute and run these performant models on that compute in a reliable way.

16:05 >> And the third thing is just developer experience and the ability to do like the tooling required to get these deployed but also run them in production. So what that might mean is a SDK to get these deployed to optimize them to configure how they scale in production but then also like all the observability all the CI/CD all the all the infrastructural care to AB test these things and monitor billing and that's the three things and I think all that surfing infrastructure you kind of get for free when you work with base 10 or someone like base 10 >> super helpful can I double click into the forward deployed function >> you you mentioned as a forward deployed engineer anytime I hear the word I want to learn more. why is that valuable?

16:44 Are there customers of yours that that find it valuable and and and and talk about that a little bit? >> Really two core reasons. So, one is it can just ex you know at no point when we through our FTE function are we trying to do the work for you. What we trying to do is accelerate how quickly you can get it done. >> Yeah. So what it means is like hey if you did this yourself in base 10 through our software we hope you'd be able to do it and it would just take you let's say n days or y weeks you you know that's fine hopefully we can just do it in a third or a quarter of that time >> with the help of our FD function who just has lots of reps in this and what they're really good at is kind of setting up that infrastructure and scaffolding so you can move >> from I have a model to I need that serving a lot of traffic at scale very very fast >> right also presumably a lot of your customers are very very busy and resource constraint.

17:36 >> Totally. >> And so this is helpful in getting them some capacity to make the transition >> 100%. >> Search capacity. >> Search capacity. Exactly. and then the second piece is that you know when you're running these things at scale you need an infrastructure team >> and we're kind of like infrastructure team on loan in that way >> but you know and what we believe is that you know we can provide that included in the base 10 experience >> and what we get out of it is happier customers.

18:03 >> Happier customers. what our customers get out of it is, you know, they don't need to stop a massive S sur and infrastructure team. >> And the third one is is that, you know, what we have seen is it scales pretty well. >> This is a plug to everybody out there who tells me that they don't have time or capacity. >> B10 will send you the capacity. >> Yes. >> To to make this transition. >> Back to the, you know, the question that we were going down is how does B10 help these customers? You know, one of the other things that I thought I'd get your pulse on is what percentage of traffic for your, let's say, your most mature customers is going through an open-source model or a fine-tuned model versus a closed >> Yeah.

18:41 >> model. how is that traversed over time and where do you see that headed? >> Yeah, it's a good question. I mean I I think this actually just like goes can I just zoom out really quickly and I think like this is what you need to believe to believe that base tenon inference is going to be a massive massive business is like fits into this which is right like that AI usage is going to be massive I think we all agree on that >> trillions of dollars of spend in 2030 >> if you think about the app layer running at 50% gross margin and we think in in three years there's going to be three to five trillion dollars of AI app spend what that really means There's one and a half to two trillion dollars of inference spend >> model. Y >> of and and then the question just becomes how much of that is in closed source inference and how much of the open source and custom model inference >> today it's probably like 3 to 5% is on open source and custom model inference >> dollar spent >> dollar spent. What's what's interesting and this is definitely being driven by the rise of really good open source models and a lot of researchers who are working on post- training technique techniques is that you can run >> you know three-month old intelligence >> at 20% of the cost and so what that means a lot of people are starting to invest >> three-month old intelligence at 20% of the cost >> yeah I saw I saw something on on Twitter about this one of our customers said they were using open code with the GLM 5.2 to >> and it was spent it was 15% of frontier of and they were using frontier I don't everything's frontier but of of of using you know the closed source APIs >> so significant reduction >> in spend and if you believe that so what what we end up seeing with our customers is with our most sophisticated customers is that it looks like something between 30 and 50% goes to open source and closed source models >> 30 to 50% of volume goes to >> go to open source and customs and custom models no not of volume of spend >> of spend And okay, >> but they're actually doing like 3 to five like it's like 3 to 5x the volume, right, >> than the closest.

20:44 >> And so you end up with with somewhere where you can do way more tokens at a fraction of the cost of a way more relative tokens at a fraction of the absolute cost. >> Yeah. >> Where it's like win-win win-win win-win. What we think is that over time the market will mature towards that. But I think you know going back to my pri prior like the thing I just said right before that says today if it's like 3 to 5% >> of spend >> even if you assume it's 10 to 15%.

21:11 >> Yeah it's a big number. >> It's a big big number. >> Let me rephrase all of that to so you said for your most advanced customers the volume is 3 to 5x closed volume. So at let's say at mid4ex. So the ratio is 80% tokens going on >> open or or or fine-tuned open 20% on closed. >> Yeah. And then in terms of dollar dollar spend, it's 30 to 50% of total. So that's 40% let's say at midpoint. Yeah.

21:35 Of total spend going on. and and the reason that's the case is actually a key insight that you could have the where the frontier was 3 months ago for 15%. >> I guess the other way to say it is like if you could roll back time 3 months ago, today's best open source model would have been at the frontier. >> Yes. >> 3 months ago. Obviously, you can't go back go forward in time, but if you're okay, >> you know, like I buy >> last season shoes from on running, if you're okay with that, >> yeah, >> you are at the frontier.

22:02 >> And what and what I would argue is that and and I I would argue this even goes to >> the real like the adoption of certain open source models and custom models themselves, which is like not every use case needs GLM 5.2. >> Yeah, >> that that's it's a fantastic model. You can use a you can use a model for three months ago. You can use Kimmy K25. >> Yeah. >> And Kimmy K26 >> and get a lot of the value. And we see that from our customers who >> you know who are trying to figure out for a given use case what is the best model?

22:31 >> Can I double click into the 8020? >> Yeah. >> So think about one or two of your leading customers. >> and no worries if you can't name them. What portion of the traffic is going to the frontier model the 20%. >> Yeah. >> What is the nature of it? What is the shape of it? What is it like planning or or or some kind of prefrontal cortex kind of work? And and then what's the 80%, what is the nature of that workload?

22:53 >> Is it planning and execution or or what's the split? >> No, actually planning ends up getting being fine-tuned to some extent. So you can actually fine-tune routers to be like where where should I give this? >> Mhm. >> you usually it's stuff that's a bit more open-ended. >> Mhm. >> And so like or synthesizing I'd say. >> Mh. >> So a good example would be like hey you you break some you there's a query >> the you first break that query up into a bunch of tasks. you solve a bunch of those tasks >> and then you synthesize them all >> into like the final answer >> or the final response for the user. And what stuff I've seen is that actually everything up until last point >> can go through custom and open source models and like that synthesis layers goes to >> the most frontier but even that is being you can post train we call it main agent work.

23:42 >> you can post train models really well just it's just like a bit more of a bit more of an involved >> postraining >> to to take over those main agent tasks. >> Now you are serving some of the fastest growing AI companies. Yeah, >> that's the who's who cursor, open evidence, a bridge, Harvey, >> Sierra, others. There's a lot that the best of them have figured out. This list that we just named, they're probably the best of the best. They've figured out a lot of things. Talk about that a little bit. What modality in this is growing the fastest? Is it text? Is it audio? Is it video? What comes after coding? I assume is number one. So, like what is after coding?

24:15 >> Yeah, look, we we work with a bunch of really great customers. everyone's a bit different. Everyone's a bit different and I don't think there's one size fits all. There's a lot of audio, there's a lot there's a little bit of video, there's a lot of text. I think more more and more things are going to like LM just dominate so much spend. >> Yeah, >> that that more and more is that >> I think >> is that 90% plus.

24:35 >> Yeah, I think so. probably I I don't know off the top of my head, but yeah, that's what I'd guess. >> And what percentage of that 90% that is text is coding? Like 99%? >> No, no, no. >> Oh, more or less? >> Less less. Yeah. I I think coding is a big part of it, but it's not it's a lot of the tokens, but not necessarily all the spend. if that makes sense. I think what we're seeing more and more is customers just want to post like the way that customers just want specialized models for all the specialized tasks.

25:04 >> So I think these models just do less and less >> over time >> and and that is definitely changing. >> And then I think the other piece >> that is changing is I think we're going to see a lot more biological use cases. >> Fascinating >> to be honest. to be honest, which is like there's a bunch of really cool companies here. Obviously, Kyomorphic and Bio >> and Chai, but it's one of these open-ended generation tasks where more is more >> to some extent where you could we could we could generate, >> you know, biological structures forever.

25:33 >> Mhm. >> It's kind of like code, like we can just generate code forever and we we wouldn't get bored. >> And so, it's it is one of these open-ended generation tasks where I think we're going to see a lot more >> coming there. And then I think on the tech side like I think we're just going to see a lot more custom main agent. So that's customers taking care of that planning stuff >> more and more themselves.

25:55 >> I'm glad I asked the question. Bio was an unexpected second pick. U I'm sure we'll look back at this conversation. >> And by the way, we're not seeing that in our >> like it's it's still a really small percentage of of our usage. >> It's just something I see which is like >> the the computational requirements there. >> I see >> for us to truly generate everything we want to do. If you think about like how many drugs we want to test >> is high >> is very very high and again and the dollars behind it are very high.

26:20 >> Yeah. >> Think about like you know pharma is a really good example. >> Are there common mistakes or common pitfalls that companies make as they start this journey? What advice would you give the next set of startups who are thinking about this journey to to keep in mind or to start doing early to start putting their data checkpoints in place? I I'd say create a very you know invest in eval I think that's one thing which is like you know companies that invest in eval or have a very good sense >> you know it's kind of like why why coding is such a good thing is because we just have a very we have a natural we have a natural eval >> you know we obviously have all the sweet ventures but you have doesn't compile you know so we have this natural eval >> there >> so I'd say invest in eval >> have all that checkpointing in place I I think >> be able to switch between models quickly >> you know and the routing layer becomes more and more important >> over time and then I think the last one is >> don't be scared I think it's just like you know a lot of these things there are a lot of companies that are trying to make all these problems easy >> for you to some extent for you to some extent >> and yeah like invest invest early especially now where there's not this massive gap between closed source models and open source models >> it's like you know Once you have something that works >> start testing stuff out because what will happen otherwise >> Mhm.

27:42 >> is that you know you you we've all seen this right which is like how many times have you opened up your credit card bill >> and you're like it's too expensive >> and then all of a sudden you start looking into what all the things you're spending money on >> and it's really really hard to start to cut stuff back. right. And but it's a lot more easier when you're when you're intentional and you're like >> even understanding how much we spending on certain models >> like invest in that infrastructure >> proactive. Yeah.

28:07 >> Yeah. >> And like I think the best companies that we talked to the ones that are like really it's not that they're necessarily very very profitable but they have a very good idea what they're spending money on >> and and and and how that will evolve over time. >> That makes sense. Yeah. >> I think every AI company needs three things, right? And it's a core AI asset that is sticky. >> And that that is a question of do your product market fit or not.

28:35 >> then you need a talent strategy. >> And you need a comput strategy. And your comput strategy might be use other people's compute. >> Like it might be like hey we we're never going to run our own models >> but like you know I'm only going to build this on top of anthropic models. That might be your comput strategy. But if you're not planning around that like and you're not communicating that and you're not >> your financial plan is not built around that, you just might be left in a place where there's no >> capacity left for you.

29:01 >> I love that. >> Product market fit base 10 check. Compute strategy. Let's talk about that next. >> And then talent strategy, we'll talk about that next. >> Great. >> How do you think about your compute strategy? >> But now today we sit on top of I think 18 or 20 clouds and >> like 87 or 90 different clusters. >> Mhm. and so we have a lot of experience now running this across clouds and to some extent you know there just like a lot of apps are like how do we own our future to some extent we need to figure out >> hey over the long period of time what is both the >> the way that we are going to financially make sense and also like how do we guarantee ourselves access I think the big thing really is about like >> two things one is >> hey do we make are we in a place where we will always have enough >> yeah your destiny >> supply yeah >> and the Second thing is like I think a lot of a lot of GPU clouds today aren't necessarily built for inference and like we feel like we have a lot of experience running running inference clouds and >> you know how the networking stack should be >> designed how the storage stack >> should be like which ISPs how do you fail over between ISPs and which ISPs are good which ISPs are bad >> you know we have a good a lot of understanding around how these things should be run and a lot of opinions and you know we think we can maybe even do a really good job >> of doing that. What is the difference between a traditional cluster or a cluster designed for training versus designed for inference?

30:23 >> Well, I mean, you think about, you know, as a good example to think about like network storage, noisy neighbors is a real problem >> for one of those and less of a problem for the other one. >> Mh. >> So, like, you know, with training, with training, you often times are pulling down all the data at once. >> You're doing a bunch of runs and you know like whereas inference you just need clean like clean data lines. our networking lines >> so that you know if someone else in the cluster >> is running a big training job it doesn't affect >> inference workloads >> got >> and so it it just requires a lot more what's the right way to think about it like >> clean partitioning >> clean partitioning and also the communicating from in to out >> in inference because it's constant >> fascinating well you've clearly thought about this at many different layers now building your hardware obviously aggregating you said 87 to 90 different locations.

31:15 >> How much of your edge is in the in the software inference stack and optimization versus you know securing GPU capacity and and and wrangling that sort of multicloud multicluster setup or or great inference engineering with customers. >> Yeah, I think all three I think the first and two of both I think all three are software problems. >> But I think look we we are best-in-class in running these models fast. We are >> very good at the software to run this reliably across clouds. Mhm.

31:42 >> And treat the underlying compute unit as truly fungeible. >> Mhm. >> So, and that's a lot of software work. That's, you know, that's the thing we've been doing the longest that we've been doing for seven years, more or less. >> and then when it comes to customers, you know, >> we're pretty good at just being in the boat with them. >> and I think, you know, that's you and if you ask me which part of the problem motivates us as a team the most, it's probably the latter.

32:05 >> Nice. >> You know, there's no point of doing all that. Yeah. >> If you don't have happy customers. >> Yeah. >> At the end of it all. >> Makes sense. The number one question that I hear about base 10 is is the next part of your compute strategy which is why will the clouds not just do this >> AWS Azure Google talk about that a little bit how do we differentiate against u the big clouds the neoclouds as we deliver the service to the customers >> I mean they probably will and they already do like I don't I don't think you know I think every single cloud and most neo clouds now have a managed inference service Yeah, >> to some extent I think there's a few things interesting here. One is like you know I think there most of the big clouds are very good at being partners good partners you've seen this time and time again right there's >> you know red shift existed but also snowflake >> right >> you know data bricks and click house of partner most clouds >> mh >> and so like you know I think there is an ecosystem where you can have best-in-class software >> that operate on clouds >> and you know the clouds win anyway because that's where we we are running I think where we differentiate really though is like >> again >> we're just like a you know you could think of us as a boutique peak inference shop.

33:15 >> Mhm. >> You know, like we're we're just going to be a lot more laser focused on on the customer and like what we trying to scale >> is that customer experience and less so, >> you know, >> we're we're not running these behemoth businesses just yet. >> >> and instead, we're just focused on how do we get our customer to value and >> we're never going to make the fact that we have to operate at an enormous scale the customer's problem.

33:40 >> so that's that's one. The second one is that this market is incredibly fast moving and we're just going to back ourselves to move a bit faster >> than the rest of the market >> here and you know we as a 300 person team >> that's you know well capitalized and has the right resources and and a >> has a strong customer base. We just think we can do that faster >> than most. >> Makes sense. Let's go to the third part of it which is your talent strategy.

34:04 Mhm. >> You know, one of my mental models is management teams scale from, you know, three stages of a business, 0 to 100 in revenue, 100 to 500, and then 500 to whatever the natural asmtote of the business is. This is a big asmtote for you guys in inference. >> Yeah. >> You've scaled from stage one to stage three in less than a year, maybe even six months. >> That is a lot of growth. >> That is a lot of growth for for for the leadership, for the team. What are the growing pains? How have you changed your role as as a CEO? Yeah, look, I I I think what happens is that you just end up having to figure out how you can as an early stage company run by a bunch of engineers, you're often a bit arrogant and think, hey, you know, structure and leadership is and I think I think I think you can get humbled by that when you have that view very quickly.

34:51 And you know, >> for us, what that means is, >> you know, how do we hire >> create a bit more structure, create leverage, hire really good leaders. >> Yeah. figure out how to how to scale what that thing that is special >> that we can you know do one to one and make it one to many >> to some extent and like you know what that means is >> you know making sure we have a really good we have a really good culture we >> we we we're fast at recruiting >> we're fast at you know taking care of things when they don't work out >> and really just again like having having the business kind of keep pace with the market >> to some extent. But I think a lot of it just comes back to hiring.

35:33 >> Yeah. >> You know, it's like, you know, we we have we've had awesome people on the go to market side. >> Yeah. Yeah. >> Danny, >> Danny, Matt, >> Matt, Amy, >> Amy, >> on the engineering side, >> Samir, >> you know, Samir's great and Joy's awesome. Edit some miracles on the comput side. >> Yeah. >> And so, yeah, there's there's a ton that just goes into hiring the right leaders for >> if all goes per plan. Yeah, >> your dreams, your wishes go to plan the next three years. what does the world look like? if all goes per plan, >> I wish I knew to be honest. Like, you know, if you told me this was the world three years ago, I like I don't know what you're talking about. If you told me this was the world a year ago, I'd be like, I don't know what you're talking about.

36:13 >> Right. Right. Right. >> I think I think what one thing that's been interesting is that this market's evolving faster than any of us could ever imagine. >> Yeah. >> I think a lot of it just like, hey, we think there's going to be a ton of inference. let's just make sure we are the default place where where it happens. >> Yeah. >> And whatever that means in terms of adjusting to new models, adjusting to different compute strategies.

36:34 >> you know like one thing that's been very different for us is that thinking about financing this business is very different. You know like we just raised a massive round last week. >> You know if you told me a year ago that we're going to go raise a billion half dollars. I'd probably be like why would we need that? >> But right now it's very clear to me that we need that. We probably need a lot more.

36:50 >> Yeah. And so, you know, I think a lot of it is just like how do we keep pace with the market know that, you know, the final and largest market is probably inference. Yeah. >> And how do we be the default for all the world's inference to happen? Well, >> awesome. Well, thank you so much for the time. This is fun. >> Thanks for having me.

Summary

The conversation focuses on the evolution of Base10, a company specializing in AI inference and training, highlighting its journey from initial explorations in machine learning to becoming a key player in the inference market. The discussion emphasizes the significant cost savings associated with using open-source models compared to closed-source alternatives, the importance of owning AI models for companies, and the strategic decisions that have shaped Base10's growth.

- Base10 began with a focus on building tools for machine learning, evolving into a company centered on AI inference.
- Customers are seeing cost reductions of 30-50% by utilizing open-source models compared to closed-source options.
- The company pivoted towards inference as the market matured, particularly after the launch of models like ChatGPT and Stable Diffusion.
- Base10's training products are designed to enhance inference capabilities, allowing customers to own and control their AI models.
- The importance of understanding user data and model evaluation is emphasized for companies transitioning to owning their AI intelligence.
- The conversation highlights the need for a robust compute strategy and the benefits of having a dedicated infrastructure for inference.
- Base10 differentiates itself from larger cloud providers by focusing on customer experience and agility in a fast-moving market.
- The future vision includes positioning Base10 as the default platform for AI inference, adapting to evolving market demands and technologies.

Questions Answered

What led to the founding of Base10 and its evolution?

Base10 was founded in 2019 with a focus on machine learning, initially exploring various ideas before settling on building tools for the machine learning trend. The founders aimed to create a 'picks and shovels' business to support this growing field.

How does Base10 prioritize inference in its offerings?

Base10's strategy centers around enhancing inference capabilities. The company aims to own more of the model training and post-training processes to improve inference outcomes, recognizing a gap in effective post-training tools.

What are the key components necessary for serving models effectively?

Effective model serving requires fast performance, robust infrastructure, and a positive developer experience. This includes managing compute resources, optimizing deployment, and ensuring observability and monitoring.

What distinguishes the workloads handled by different models?

Different models handle various types of workloads, with the frontier model managing more complex tasks like synthesis, while open-source models can handle simpler tasks. The synthesis layer often requires more advanced models for final outputs.

What are the differences between infrastructure designed for training versus inference?

Infrastructure for inference requires clean data lines and effective partitioning to prevent interference from training jobs, while training infrastructure can tolerate more noise due to its batch processing nature.

© transcribe · For agents Built with care and craft by Gokul Rajaram