transcribe

Inside the a16z Startup Attacking AI’s $400B Spending Problem

Will Phillips · 18m · transcribed Jul 2026
More from Will Phillips Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

The Transformative Power of AI and Inference

What is the significance of AI's growth and how does Inference fit into this landscape?

AI is experiencing unprecedented growth, with enterprises spending billions on models that may not justify their costs. Inference aims to address this by creating specialized AI models that are cheaper and more efficient, allowing customers to own their models outright.

  • AI's growth is accelerating, with significant financial implications for enterprises.
  • Many companies are overpaying for AI capabilities that don't require advanced models.
  • Inference offers a competitive advantage by providing tailored AI solutions at lower costs.
# 3:36

From Seed Stage to Venture Scale

How can Inference transition from a seed stage startup to a successful venture?

Inference raised an $11.8 million seed round and aims to leverage this funding to scale by addressing the inefficiencies in AI spending. Their strategy involves demonstrating a clear path to significant valuation growth by offering cost-effective solutions.

  • Inference's seed funding is a stepping stone towards achieving a larger market impact.
  • The startup's previous pivot from crypto to AI highlights adaptability in a competitive landscape.
  • Understanding the financial dynamics of AI usage is crucial for attracting further investment.
# 7:12

The Role of AI Agents in Modern Applications

What is the emerging trend of AI agents and how are they changing the landscape?

AI agents are evolving from simple task performers to autonomous systems capable of handling multiple tasks simultaneously. This shift necessitates new tools for managing and optimizing their performance in various applications.

  • AI agents are becoming more complex and capable, leading to increased demand for management tools.
  • The proliferation of AI agents signifies a major shift in how businesses utilize AI technology.
  • Companies need to adapt their strategies to effectively integrate and monitor these advanced AI systems.
# 10:48

Building Observability in AI Applications

Why is observability important for AI applications and how is Inference addressing this need?

Observability allows companies to track and analyze their AI usage effectively. Inference is developing a product that integrates observability into AI platforms, enabling better insights and performance metrics for users.

  • Integrating observability is crucial for understanding AI performance and usage.
  • Inference's approach to observability can enhance customer engagement and retention.
  • The ability to monitor AI applications is likened to fitness tracking, providing actionable insights.

Transcript

0:01 People underestimate how technology compounds over time. It was pretty clear early on the LMS were going to be pretty transformational and AI would be big, but it wasn't totally obvious what 10x year-over-year and now month overmonth growth really looks like. AI is going to be even bigger than maybe people thought at the beginning of all this. AI has never been cheaper to use and yet running it has never been more expensive. Enterprises spent over $8 billion on AI models last year, and a lot of them are waking up to a hard truth. They're paying frontier prices for tasks that don't require frontier intelligence, and the output isn't justifying the spend. But the startup I'm visiting today has built its entire business around that gap. It is such a competitive advantage if you're a startup and you have some technical engineers. Like, you can build things that are so specific to your use case.

0:53 so quickly and it's just like a superpower. >> Today, I'm going behind the scenes with Inference, a startup that takes the biggest AI models in the world and shrinks them into small specialized models that do one task at a fraction of the cost. It's a process called distillation, and the customer owns the finished model outright. They've raised an 11.8 million seed led by A16Z, and their founder, Sam, has declared war on what he calls Big Token. the handful of labs that control access to Frontier AI.

1:26 >> It's a growing dynamic market. We want to help people who build AI apps build them better, build them faster, build them cheaper. Our mission is to power the world's inference. >> The long-term bet here goes deeper than just saving money. When you own the model, nobody can repric it, change it, or take it away from you. So the real question I want to understand is how fast a startup can wedge itself into one of the most funded and crowded layers in all of AI.

1:58 >> This is a very very very cool office. >> Thank you. >> So it's funny man. I actually had heard about you on my last trip. They were like the biggest fan of your ex posts. I don't know why because apparently you have really hot takes on X. Is that like the case? >> You should post a lot on X. Yeah. Is there a formula? Is there a strategy to that or >> No, no, definitely no strategy. It's all just like SF tech culture discourse.

2:22 >> I think you're some people's daily news. That's what I kind of hear. >> Oh god, >> there's some insightful stuff mixed in there occasionally, but a lot of it's just like wake up in the morning and like rip some random tweets. >> That's how it should be. Maybe I'd like to quickly touch on you as a founder. You were with Founders Inc. for a while. How did you actually start in this process like moving to SF and doing the whole thing?

2:45 >> Yeah, so Founders Inc. just wrote an angel check into what at the time was use Context Incorporated and was spending a lot of time on Twitter. There's not much of a tech community. So really found like-minded folks online, started engaging there, posting a lot of like open source projects I was working on. This was all around the time chat GBT came out and one of the founders in guys for con reached out to me and was like, "Hey, what are you working on?

3:07 This is cool." And so they ended up writing a small check. We raised a preede round for what was effectively fairly generic. So that was kind of our entry point into working on open source models, working on lower levels in the stack on infrastructure all the way down to, you know, now training and inference. >> What do you what's what's going on? >> That's our that's our office little like test rig for for training models, running models.

3:39 All right. It's pretty cool to see their office and also just inference from the inside. But I want to find out how a company like this gets from seed stage to a genuine venture scale outcome. So, we're going to pick up a pen and some paper and we're going to map it out live. Let's go. So, inference just raised an 11.8 million seed round co-led by Multicoin Capital and A16Z startup accelerator CSX. And if you know those names, you'll notice they're both crypto investors. This isn't a coincidence.

4:09 Inference actually started life as a venture-backed crypto startup before pivoting into what you see today. But what I want to look at is how they turned an 11.8 million seed round into a venture scale outcome with the startup they've got now. So the deal set up. They don't make the valuation public. But I'm going to make an educated guess here for a seed round about their size at $1.8 million raised. It values the company at around 50 million. 11.8 million buys roughly around a quarter of the company's equity split across those investors. And by now you know the math. A check like that only makes sense if there is a genuine path to the company being worth billions of dollars later on. So where's that path? The current problem is companies are spending billions of dollars a year calling the big frontier AI models and wasting a ton of money by massively overpaying. Let's make this super concrete. Take a simple repetitive task like reading an email and pulling the order details. Claude's top model, Fable 5, charges $10 per million tokens in and $50 per million tokens out.

5:05 Claude's cheapest model, Haiku, $1 per million tokens in, $5 out. Now, for that same task, we're producing the same result at 10 times the price. And as we know, the cheapest open source models will do it for cents, up to 100 times cheaper. Multiply that gap by millions of tasks a day and multiple people in the organization running those. That is where the overpaying comes from. Because that gap exists, there is now an arms race to own the layer that decides which model executes on which task. For example, RAMP, the finance company, just the other day opened its internal model router to the public. Open router, arguably the biggest name in routing, is reportedly fielding acquisition offers in the billions. By the time this comes out, things could have changed. And also, the big labs keep shipping cheaper and cheaper models to keep you in their ecosystem. So, everyone's fighting for the same seat.

5:56 And inference's angle on this race is a little different. They don't just help you route between models that you rent. They replace the rented model with a small one that you own trained for your task. Now, of course, some of you may ask the obvious question, why don't you just grab a free open-source model? Because a lot of the ones, especially coming out of the Chinese labs, are starting to close most of the quality gap. And on some tasks, we already see that they beat the closed source models in the US. But like any hosted model, you've got to fine-tune it, serve it, test it, all while keeping up with a leaderboard of new models that is consistently shuffling week by week. The trouble is some models can even get retired bid project. Now this takes someone who's really technically savvy or a machine learning team that most companies don't have. So the tax isn't really quality anymore. It's the operations. Enterprises can't build on ground that's consistently shifting.

6:45 Inference's pitch is we do all that work on a small model that you own and no one can tweak and retire a model that you own. So how does a bet like this actually pay off for investors? It's the same three doors as always. One, they get acquired. two, they keep growing and early investors like CSX and Multicoin sell their position to bigger funds. That's called secondaries. Or three, the big one, an IPO. But here's the thing about this corner of AI. Companies in this exact layer have a habit of walking through door one bought by bigger players before they ever get the chance to go on it alone. Whether inference stays independent long enough to own its own category is the real question. And of course, nobody knows the answer yet.

7:25 Now, I want to be super clear in this breakdown that I'm putting on the hat of a generalist VC doing a bit of digging. I'm not an expert in distillation or inference or any of that corner of the AI world specifically. But that's also kind of the point. This is roughly how an investor without a machine learning PhD would size up a check for a company like this. So, for anyone who's a bit closer to this than I am with technical knowledge, I'd love to hear from you in the comments below. Now, that breakdown took me a ton of research time to distill a bunch of complex points into a simple breakdown, but for most of my week, I don't actually have that luxury.

8:00 It's backto-back founder calls between travel, and sometimes I only get like 2 to 3 minutes between calls to recap on what's coming up. This is where our episode partner, Granola, has been genuinely useful for me. It's an AI notepad that runs in the background of my calls. One of the features that I use most is the preall brief. Before each meeting, it's somehow pulled together who I'm sitting down with, what the startup does, what they've raised, who the founders are, links to their profiles, literally everything in a few concise sentences. And if I've had prior meetings with these contacts, it'll pull me some relevant context from previous meetings that I may find useful. So those 2 minutes I get between meetings, I can quickly brief myself on what's going on, so I can focus on being present in the meeting and not just playing catchup the whole time. I'm using this tool every day, so it's a really easy one to recommend. I've got a link down below for your first month's free. Thank you so much to the legends at Granola for making this episode possible.

8:54 >> So I think there's kind of like two pieces to Asian traces. There's all the like backend work or maybe there's three. So there's like all the ingestion work that's just like purely on our side. So what do we need to do in order to support that? I think a lot of the existing inference storing infrastructure should work if we annotate things correctly. There's the client side where we actually send traces from the agent application and then there's like UI level which is just like trace exploration and however we want to render those data structures.

9:25 >> How much do we want to divide our current inference flow from traces? Like right now we're showcasing inferences, build data sets from inferences. Do we want like everything that's in a trace also show up in inference and being able to build a data set off that or are we just dividing this into two different forms? A trace is not going to include just inferences. >> It's tools all the same. >> Yeah. Everybody I talk to who uses like Lenfuse or these products like they find it very valuable that they see a trace that includes both their LLM calls and other things executed by their agents in that loop. So you actually you're probably going to receive traces that have nothing to do with language model calls. But the this is an agent execution. Like the trace itself is the agent executing. It might make a vector DB query. It might do a database insert.

10:14 It might do a web search. And then it does an LLM call. And then it does something else something else. >> You want to just fill me in, Sam, on what we're trying to achieve here in this meeting? >> well, we're doing some competitive research. We have some features that allow you to like store and organize your inference calls in different ways. The industry sort of moved to like one kind of layer of organization on top of what we current currently offer, which is more agentcentric. And so we're planning all the tools that will allow us to support kind of full agentic tracing end to end which is it just gives you a bunch of kind of really nice ways to like organize all the data that your LLMs are producing. For like a random viewer who's trying to figure out like what is an agent trace like rewind to like 3 years ago when AI was just kicking off most people use chat GBT right and like your thread you can think of that as like that's an agent trace.

11:08 It's you and you know the AI having a conversation and it goes back and forth and there's turns. You take a turn, it takes a turn, you take a turn, it takes a turn. And then most products that were being built on top of these AI models were one turn, right? So like the most popular app is like let's say Cal AI. You would take a picture of some food and like get back an answer from the AI and you're done. And like that thread like that chat thread is done. but now these agents AI is starting to get used autonomously in the background where it's like it's basically doing a bunch of things itself. So you went from like Uber driver that goes from like destination A to destination B to like an Uber driver that's doing like 10, you know, drop offs in one route. And so now these agents are starting to do many things. Like you know, you'll you'll give it one instruction and it will go off and do five, six, seven, 10 different things. A bunch of people will say it's the year of agents, so this is it. Like we're kind of getting ahead of the fact that we think agents are going to start to proliferate in a big way amongst companies and we need to like build tooling for them. And we're seeing proof of it because all these companies that we talked to say like, yeah, we're deploying agents now and we're running a crapload of them and we don't have a good way to like figure out what they're doing. And so this this product will help them figure out what they can do.

12:39 >> >> I'm working on data set uploads all day. a few meetings this afternoon. Aiming to have new data set creation flow in prod by the time I leave. >> Training flow is all working to end no. Evolve are working. checkpoints working. We got the loss graphs. Metrics all look really good. next up I need to add a new provider. So, we're going to add Runpot on there. maybe Lambda as well. We'll see. and then just focusing with Declan on testing. Sweet.

13:07 >> We're going to be testing a lot of training stuff today. We're going to be testing out multimodal long context agent. Basically just having trainings over and over again, seeing what breaks. And it seems like there might be a few things that are buggy right now. So, we're going to go over that. Hopefully, the goal is my today is get something stable working on the training. using all the model identifiers right now. So, I'm working through that and hopefully get it up to production today.

13:31 >> Nice. >> All right. Training. >> You guys want to go in the back? >> Yes. >> Let's do it. >> What are the problems you're tackling right now? Where is the product? Where's raising the team? Give me like a snapshot of where you guys are. >> Probably the most important thing right now is just getting in front of as many customers who are actually consuming AI tokens. and that's like pretty straightforward to do because there's so many of them out there. and so for the last couple months, we've been trying to pitch custom model training to customers, but this it's a lot of it's a big commitment. It's like buying a house, right? And so we're trying to switch the the pitch now to like, hey, just just rent this tent for a bit and like see the value you get from using our products versus like jumping straight to buying a house, which is training a custom model. And so we're trying to deliver value faster with less effort for our customers. And we think the first voyage into that is observability, right? So before we tell you how we can make your stuff faster and cheaper by training a model that might take like days or weeks, like hey, just plug us in to your app and we'll immediately start giving you insights that we've been able to build through, you know, just working with a ton of customers and knowing what people what questions people ask when they're trying to analyze their AI usage. And so the biggest most important metric right now is how many customers we can get to integrate this observability product into their platforms. and so far I think we've had about a dozen and we haven't even launched it yet. And so I think it's like you know I think we're capturing like over a few million inference calls an hour right now.

15:02 just from customers who like didn't have an observability product before but do now. It's almost like putting an Apple Watch on, right? Like now you start actually counting your steps and your calories and you know your heart rate. Yep. you got to whoop on. we're going to be getting as many people to put these on on their AI apps as possible. And once we have that, I think that's a great top of funnel to to custom model training, which that's our that's one of our more higher conviction bets is that we bet we're betting that training specialized models is way better for most teams than just using an offtheshelf model.

15:38 >> >> A was telling me it's about getting top of funnel in so you can start proving value without them having to have much uplift like as a customer, right? >> Yeah, there's there's a couple different strategies that we are trying to employ. We we actually do have this whole go to market engine that we built inhouse thanks to cloud code. but what that kind of system does is it's very personalized to like our data like internally and then also just like the criteria for what we are looking for in leads. So it it does a bunch of different things like when people come to our websites onto our website we actually track like who those people are through the cookies and so we know who they are through this platform called RB2B which is really interesting. We track all the interactions across all our social platforms. We search for news to find out like which companies have recently gotten funding or starting like new AI initiatives and all that. And then so we have this whole system to kind of source and qualify leads. And from there we kind of reach out to people that are already kind of like warm and kind of nurture those relationships and hopefully get them on board to the platform.

17:05 This is for the younger founders out there. >> Is there anything that you would tell them now, you know, guys and girls currently getting into AI, currently get like, you know, learning this new technology? What would you say to them right now? >> >> Yeah, I mean I think one other thing I've been reflecting on a lot recently is like you know there's this whole corpus of knowledge of startup knowledge from the last generation which is like all the you know YC videos and 0ero to one the Peter Teal book which is like you know mythos at this point that advice is valuable and I think can still give you a good foundation but the world has changed a lot and the rate of change is also increasing so don't be afraid to try things that seem outside the box that maybe you're antithetical to some of the things that you'll read online because things that didn't used to work will work now and things that used to work probably won't work anymore. So, be weird. Do do weird stuff.

17:58 Nice.

Summary

The discussion centers around the transformative potential of AI technology, particularly through the lens of a startup called Inference, which specializes in creating smaller, task-specific AI models at a lower cost. The founder, Sam, aims to disrupt the dominance of major AI labs by offering enterprises ownership of customized models, thus addressing the inefficiencies and high costs associated with current AI solutions.

- Technology's impact compounds over time, with AI growth accelerating rapidly.
- Enterprises spent over $8 billion on AI models last year, often overpaying for unnecessary capabilities.
- Inference specializes in distilling large AI models into smaller, specialized versions that are cost-effective and owned outright by customers.
- The startup recently raised $11.8 million in seed funding, positioning itself in a competitive AI landscape.
- Inference's approach addresses operational challenges faced by companies in managing AI models.
- The startup aims to provide observability tools to help customers analyze their AI usage before committing to custom model training.
- The founder encourages new entrepreneurs to embrace unconventional strategies in the evolving tech landscape.
- The ultimate goal for Inference is to facilitate better, faster, and cheaper AI application development for businesses.

Questions Answered

What is the significance of AI's growth and how does Inference fit into this landscape?

AI is experiencing unprecedented growth, with enterprises spending billions on models that may not justify their costs. Inference aims to address this by creating specialized AI models that are cheaper and more efficient, allowing customers to own their models outright.

How can Inference transition from a seed stage startup to a successful venture?

Inference raised an $11.8 million seed round and aims to leverage this funding to scale by addressing the inefficiencies in AI spending. Their strategy involves demonstrating a clear path to significant valuation growth by offering cost-effective solutions.

What is the emerging trend of AI agents and how are they changing the landscape?

AI agents are evolving from simple task performers to autonomous systems capable of handling multiple tasks simultaneously. This shift necessitates new tools for managing and optimizing their performance in various applications.

Why is observability important for AI applications and how is Inference addressing this need?

Observability allows companies to track and analyze their AI usage effectively. Inference is developing a product that integrates observability into AI platforms, enabling better insights and performance metrics for users.

© transcribe · For agents Built with care and craft by Gokul Rajaram