transcribe

AI’s Next Race: Cost, Control, and Compute

CNBC · 1h 0m · transcribed Jul 2026
More from CNBC Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

The Shift to Open Models in AI

What is the significance of open models in the current AI landscape?

The AI industry is transitioning from a focus on proprietary models to open models, which are expected to dominate the market. Open models will provide cost control, accessibility, and empower developers, challenging the dominance of a few frontier labs.

  • The model is no longer the sole product; it's about harnessing AI capabilities.
  • Open models are projected to account for over 90% of AI tokens created in the near future.
  • Affordability and accessibility of AI are crucial for widespread adoption in businesses.
# 12:09

Future of Model Scaling and Deployment

How will the scaling of AI models impact their deployment in businesses?

There is still significant potential for scaling AI models, and both frontier and open-source models are expected to advance. However, the focus should be on deploying these models effectively to empower small businesses and enterprises.

  • The gap between frontier and open-source models may shrink significantly.
  • Deployment and orchestration of AI capabilities are critical areas for development.
  • Hybrid compute solutions may help reduce costs and improve efficiency.
# 24:18

The Importance of Open Source for Cybersecurity

Why is open source software important for cybersecurity?

Open source software, like Linux, allows for rapid patching of vulnerabilities, which can enhance cybersecurity. Advocating for open-source models can similarly improve defenses against cyber threats.

  • Open source enables faster security updates and vulnerability management.
  • Regulatory discussions should consider the benefits of supporting open-source models.
  • Affordable AI through open-source can drive economic growth and innovation.
# 36:27

Enterprise Adoption of Open Models

What trends are emerging in enterprise adoption of AI models?

Enterprises are increasingly recognizing the potential of open models for managing costs and achieving performance comparable to frontier models. This shift indicates a growing acceptance of open-source solutions in business applications.

  • Enterprises are waking up to the benefits of open models for cost management.
  • Performance improvements from fine-tuning open models are significant.
  • The landscape of AI is shifting towards open-source solutions for better efficiency.
# 48:36

Corporate Concerns and Open Source

What are the main concerns corporations have regarding open-source AI models?

Corporations are primarily concerned about the origin of AI models and their operational security. They seek assurance that open models can be run safely and effectively within their geographic and regulatory frameworks.

  • Corporations prioritize where AI models run and how they are managed.
  • Security and performance are key factors in the adoption of open-source models.
  • Providing a secure environment for open models can facilitate corporate adoption.

Transcript

0:00 AI is entering its postfrontier era where the model is no longer the whole product. The value will be in routing cost control and compute and openw weight models. They may start to squeeze the frontier labs. The >> model alone is no longer the product. It is the harness. >> 90 plus% of the tokens created will come out of openw weight models over the next 18 to 24 months, possibly even by the end of the year. You know, the average Fortune 500 business may have thousands of internal applications and they're modernizing all these with AI. And for that to truly be productive and in the best interest of the business, it's to use open models.

0:37 >> If you want the benefits of AI to be widely distributed, then you really need AI to be a lot more affordable. >> The idea of developers empowered is what fuels our industry. And we resist early oligarchs igopies. >> This is the open model squeeze. Cheaper models, more control, and a real challenge to the idea that a few frontier labs will own the AI stack. The full conversation is here.

1:09 For the last couple of years, the AI race had a pretty clean scorecard. Bigger models, better benchmarks, more GPUs, the right to claim, the right claim to the frontier, at least until the next launch. Now, that became the easiest way to understand the boom. Open AI would ship. Anthropic would answer. Google would close the gap. Maybe a Chinese lab comes out with something surprising. And everyone recalculates where they stand. That scoreboard though, it's getting a lot more complicated. The frontiers still matter.

1:35 Certainly, OpenAI releasing GBT56, Chinese Labs, they're shipping competitive models at a pace that is hard to ignore. But companies, they're now far enough into actually using AI that the conversation is moving from model rankings to deployment reality where cost control, where their data is actually going, it matters a lot more. And increasingly customers, they're not waiting for the labs to tell them which model is best. Door Dash, the food delivery platform, it had a good example of that this week. Instead of relying on public benchmarks or model launch claims, Door Dash built its own test for AI code review using its own codebase, its own pull requests, and its real engineering problems. The result was a pretty good snapshot of where AI is right now. No single model won everything. A basic one pass AI reviewer missed a lot. The better answer was a system, multiple models, different roles orchestrated around the actual job that Door Dash needed done. And this is really an incredible reversal and maybe a sign of what is to come for more companies. I had a few people reach out when I was talking about this like G2 Patel, president over at Cisco, who said this has already been happening for months. Public benchmarks, they don't always map to real world results. For years, this is interesting because for years the labs benchmark the models and the rest of the market reacted. But now the customers, they're the ones benchmarking the labs. Once companies do that, they don't have to take anyone's word for anything. they can decide which model works for them. Now, Perplexity is coming at the same shift from another side. This week, it previewed a new system for its AI browser and computer product that starts with a cheaper open model from China's ZAI, also known as GPU, then calls in a stronger, more expensive model only when the task needs it. And an analogy that many people are using these days, you don't use a Ferrari to go to the grocery store. Now, both moves, they point to the same bigger idea. The product is no longer just the model. It is the system around the model. And that's why this next phase of AI might look different from the last few. The postfrontier era is what we're going to call it. Today, Perplexity CEO R. Vancas joins us to talk about the company's new orchestrator model, why he's building on open-source Chinese AI, and his argument that token value per watt may become one of the defining metrics of the next phase. I also spoke with benchmarks Peter Fenton and Oblama CEO Jeff Morgan a few hours ago. We're going to run that conversation in full later on in the show. Peter made a striking prediction that you're going to want to stick around for. He said that over the next 18 to 24 months, maybe even by the end of this year, more than 90% of tokens could come from openweight models.

4:16 That's a huge claim that also explains why Benchmark was an early investor in Olama. This is a company making it easier for developers and enterprises to download, run, and manage open models. Let's get into it all. First up with Arvin. Arvin, it is great to see you. Thank you for joining. I always say that you were kind of Perplexity was the OG model router. You have seen this shift from the very beginning. What is it like right now? Did you expect sort of everyone to jump on model routing this quickly and this kind of momentum that we've seen this year?

4:49 I think the importance of model routing definitely increased exponentially since the advent of agent harnesses. our positioning with the perplexity computer product which is essentially an agent harness is that one model alone is never going to be good at all things. This was actually told by anthropic CEO Dario himself that models are actually a lot more differentiated than cloud. Clouds are actually pretty undifferiated and different models are pretty good at different capabilities at different cost performance trade-offs and so the value of orchestration and model routing automatically increases as models starts specializing in different kinds of capabilities and so many different enterprises even within an enterprise different teams within an enterprise and across enterprises different companies have so many different kind of capabilities that they want out of AI and so the the the the leader in any category keeps changing pretty fast and it's so hard to keep up in terms of how the public benchmarks they release actually translate to real value for your use cases and what your customers want and so the value actually lies in every company having their own eval harness model and loop to iterate on this right so that that there in lies the value that that's how you can throw your own tacet knowledge that your company and your customers have into a platform that you can control and bring down the costs and see the ROI on the AI spend instead of being worried about token maxing and not knowing exactly what to do with >> it, >> right? It's not just like better models for better tasks or cheaper models for cheaper tasks. It's more complicated than that, which is I think what you're saying. You need to use the >> There's another way of saying this, which is the model alone is no longer the product. It is the harness, the orchestration system that puts the model inside a very capable harness and pairs the model with a lot of tools.

6:47 >> The tools could be bringing in valuable context that exists within your enterprise uniquely that allows you to serve unique value to your customers. So that is essentially the product and whatever allows you to serve the best performance cost trade-offs here is what you should be optimizing for rather than how many tokens your engineers are spending. >> Right. And I think we're totally aligned there. It's no longer just the product. Things like cost, control, compute matter. I want to dig into all of those.

7:15 But first, when you said that, you know, Dario Amode, the CEO of Anthropic said he saw this coming. He acknowledged that you wouldn't just use one model for everything. Do you think that was in a different environment? Do you think he expected did you expect sort of the explosion and the competitiveness of these open source Chinese models? I think it was obviously like you know on the horizon and I think we've had a conversation about this maybe one or two years before this where you you you asked me whether it was predicted that deepseate would catch up with 03 and then I made a prediction at that time that yes they'll catch up but then the frontier will again make progress and then this will keep going for a while and that's basically what ended up happening. there was a point of time when people thought Opus was too too far ahead to catch up with. now you know you've seen multiple different model makers saying they caught up with it at a fraction of the cost. and so that that's basically very important and Perplexity is is the only company that can claim to actually benefit from competition at all layers of the AI stack and also contribute to all layers of the AI stack. So when open models make progress, it's actually amazing for for for us. And when frontier models make progress with between each other, competing with between each other in terms of offering maximum intelligence for the lowest cost, we're the benefits of that flow the most to a company like us that is able to orchestrate whatever they do to the customer with minimum cost for maximum performance. which is what I meant when I said maximum token value per watt per user. Because at the end of the day if data center buildout is bottleneck by power what really matters is you provide the ROI on token spend with minimal power consumed and and so that's where our positioning is unique and that's also why we are engaging in post training that is also unique to just like fine-tuning on a corpus of data. Instead, the specific form of post training that we are working on is to embed the skill inside the open source model to escalate to the frontier model which could serve as an advisor model to the open source model and do that with a specific agent harness that allows you to serve production workloads for knowledge work. If you handle this, if you do that, you never have to be worried about like whether to use open source models or frontier models. The answer is always use whatever is the best for the task and embed that routing skill deeply inside the model weights itself instead of building like standalone routers that have no knowledge of what the model is capable of inside the agent harness. So this is actually the biggest paradigm shift. The model is alone is no longer important. the it's very important to view the model in the harness in a vertically integrated fashion and the frontier model itself would serve as a sub agent or a tool inside the harness and that allows you to imagine so many different forms of inference compute in paradigms including the entire orchestrator model and the harness running on a local device that you own and control and don't have to pay for any token that's consumed on that device and so there are so many different forms of imagining this with hybrid compute so That's what like we are very excited about and Nvidia has this amazing architecture of DGX Spark that puts the GP10 GPU chip with unified memory on like one local device and that's going to allow you to serve like really efficient local models that orchestrate frontier models on the server and at the same time like you can also imagine serving much more efficient models like GLM as the core orchestrator and still have them escalate to an adviser like Opus 5 or you know GPD56 whenever necessary. That's that that so that you you want to position the product and the platform in a win-win fashion between open source and Frontier and not be caught up in this sort of like debates about like what's actually going to win, >> right? And I feel like over the last few weeks as Wall Street and enterprise pays more attention to open source models, the line has been okay, this is how you do it cheaper. But when I was talking to Peter Fenton and Jeff Morgan earlier today, we're going to run that full interview. Peter said something really interesting. It's not necessarily cheaper. Sometimes like a GLM will be better for the task, not just cheaper.

11:29 And that is something sort of that I learned. Do you think that back to your prediction a few years ago, which was incredibly precient? >> Do you still think that there's the same gap between the frontier and Chinese open source models, or has that been narrowing? Is it three to six months as someone else told me? I think it's fair to say >> I think it's fair to say the gap's like six months maybe that's I mean obviously it's a ballpark estimate but maybe it's going to take another six months for them to produce a model that's as good as MTOS or Fable or or GPD 5.

12:04 >> Will the labs always >> will the labs always be six months ahead or do you think that as time goes on like three years from now what does that gap look like? >> to me it doesn't feel like we're hitting the limits of scaling. I feel like the GB GB300 paradigm of training models has just started and so we're going to benefit see advances from models that have been post-trained on top of even larger pre-trains that have been pre-trained on GB300 chips. And so I I do expect both Frontier models and open source models to keep making progress in a similar fashion for the next year or two. And I think it almost doesn't matter because what really matters is taking all these capabilities and grounding them in the context that will power small businesses and enterprises and letting them own the harness eval loop. There's so much work to be done on the deployment side that even if the capabilities actually stop advancing and open source and frontier stay at the same level, there's so much work to do on the deployment and orchestration side which is what like we want to focus on and like deliver value. But just purely as a from an intellectual prediction standpoint like I do expect like more progress to happen on both sides both open source and frontier >> it's possible that the gap could shrink from 6 months to even you know a month or two or it's possible that there may be no gap at all u but nothing is technically free right you still have to serve it on like GPUs on server which is why we're excited about like hybrid compute as well where some of the token spend can actually go towards local models on devices you own Right. And this new sort of metric that you're calling intelligence per watt. I want to get into that a little bit because I want to expand this out too from just like a purely AI or tech company conversation. Why should other companies that are not in tech? How should they think about hosting their own models?

13:59 who is Nvidia Spark for? Is it for a consumer, a developer? Is it for the enterprise as well? everybody u enterprises, consumers, proumers, small businesses you want the AI value to diffuse to the society in in a wide way right like people are obviously annoyed and almost to the extent of like for example I don't know if you saw the news about this bacteria thing in the seas I don't know how true it is that it was caused by the metadata center but in general the public perception of AI is so negative and so you're not going to do any good if AI is just locked up in frontier labs in a to large companies that can afford to pay for inference for these tokens.

14:39 It needs to diffuse and the only way for it to diffuse in a wide way is people should be able to control where it runs, what model they run, what knowledge it has access to and shouldn't necessarily be worried about it getting shut down because some model gets pulled off from the data center or the data center gets attacked in some manner. and so that that's that's a problem that local compute solves. It allows you to have the freedom and and cost control and the flexibility and sovereignty. And DGX Spark is such an amazing piece of hardware because it unifies the memory across the CPU and the GPU inside that local device unlike a data center where the memory between the CPU and the GPU are not unified. So most of the power bottlenecks is because of constantly having to move data around. If that problem is solved in a specialized way for a local device, then it's a lot more power efficient. it can run in few hundred watts and that's honestly like the only way to scale a lot of token spend across people without having to build more data centers >> and and that's why I'm very excited about that and u the models are getting there like a quen 35B can actually be pretty efficient at running local workloads >> and if you train it with a post- trainining strategy that perplexity is adopting to be able to escalate to a frontier model on the server you get a win-win situation where whenever it feels like it can accomplish the task of the user, it can still use a frontier model as a remote tool through the local device, through the local harness. And that's like a different paradigm from what's happening today. And it's going to completely change the landscape of how people spend tokens.

16:17 >> Let me make sure I'm understanding this as well. Arvin, are you saying that if more developers, companies, governments, etc. do sort of let's call it like deskside computing or local computing, you don't need as big of an infrastructure buildout. Is that what you're saying? >> Not not not not as much as you do today. Definitely because >> that's really interesting. >> What what is one thing you you hear all the time when people ask why is AI not like why are you not making more revenue? They they say that they don't have enough compute. Okay. Then you ask the question why don't you have enough compute? then say we don't have enough power.

16:54 >> Okay. Like so clearly a power bottleneck. So what is a lot more power efficient solutions to use the power we all have at home all the time >> and build something more power efficient architecturally at the chip level with unified memory across the GPU and the CPU which DGXR clearly solves. that I can I can imagine talking to like the CFO of a non- tech company though and saying that just sounds too complicated. We don't want the risk of hosting our own compute. What do you say to them? And then the second question >> actually are we building too much?

17:33 >> Okay. >> Well, you're not actually hosting a data center. you're it's it's like buying a PC for each of your employees. Why is that risky? You can still ad you can still have admin control over each of those PCs just like how it's being done on Microsoft server across enterprises. Large banks do that all the time. All all the PCs they give to each employee is still centralized through admin controls. You can do all that stuff with DGX parks on every desk. No problem there.

18:01 >> Right. That's really interesting. That's an interesting analogy that makes it sort of more tangible or understandable maybe. So, >> AI is the new computer. >> PCs are like outdated and legacy. Like they're just display monitors. >> Like you can actually have a DGX park with a Dell display monitor and not have to use Windows at all >> because most of the windows remotely too. >> You just use it through connectors. >> Yeah. The agents would use Excel and PowerPoint and Word through connectors inside an agent harness like computer perplexity computer and you just delegate tasks to that and so the compute runtime becomes the computer more than the OS itself. AI becomes the operating system that's what we want to realize through perplexity computer. So then the whole enterprise landscape changes entirely. People start doing work through AIs rather than doing it manually.

18:59 and the local compute becomes your computer that your an enterprise buys for every desk. So this thesis part of it rests on local compute and the other part of this is open source models or openw weightight models right you have to have the ability to actually download and put that on the system I know that you built your latest orchestrator based on GLM52 can you explain why GLM was the right model for you >> so GLM 52 post train is what we're doing on the server side not not on local but it's a pretty efficient model 700 billion parameters in total and approximately 40 40 billion active. I could be wrong on this number exactly but roughly around that and it's almost an opus grade. It's definitely a sonnet grade model and with post training and pairing with an advisor tool, it is opus grade at one/ird of the cost. So and and unlike a lot of models that get published that claim to be opus grade but only good on certain benchmarks that have been academically like like hill climbed on GLM actually generalizes pretty well to like lot of production grade tasks.

20:06 So it's such an incredible model that the company has trained Z AI and we you know congrats to them and we took that model we post trained it on task that our customers care about in our product and our agent harness and then deployed it in production today and have lowered the cost like one-third to what we're paying Opus >> and every enterprise should do this and every enterprise is going to do this like every enterprise is going to create a bunch of evals that they care about that their customers care about that's relevant to their business and and harness harness will be purpose-built for that those eval trained to be good inside that harness measured on those evals and then deployed on production. This is going to be a flywheel. So once it's deployed, you're going to collect more data and keep improving iteratively and you're going to own and control the cost of intelligence and deployment yourself.

20:56 >> And that's the future. That's how the ecosystem actually expands. >> That's why I found that Door Dash story and I'm sure so interesting. I'm sure Perplexity does something similar. it feels like only a matter of time when companies are going to have their own benchmarks because, you know, we've always said that the public benchmarks or the traditional ones, you can benchmark max, they don't necessarily relate to your own company, but it's hard to trust them. Yeah.

21:20 >> Yeah. And I know we've we've been talking about this for years, Arand, how the Chinese models when you download them, they're secure because they live on your own servers. But I feel like that is still something that's a little bit misunderstood or maybe you know, companies are still a little bit skeptical of. Do you worry about how that's perceived and also like we've seen how Washington can see these things sort of black and white, not very nuanced. Are you worried about any backlash towards these Chinese models, especially as they gain so much adoption and maybe take market share away from American labs?

21:54 >> I'm not worried about it. I think Nvidia will take care of it for America and Neimatron 3 Ultra is pretty close to the quality of these Chinese models and you know give it like two or three more generations of pre-trains and post trains and they're going to get there and we're working with them on the post- training side as part of the Neatron coalition. and and so I can totally see a version of Neatron either 3 Ultra or the next model that will be as good for American open source. And here's the thing like the open source model doesn't ha have to be like sonnet 5 or fable five level right it just actually even has to match sonnet 4 six >> and you it's pretty clear that if you have that capability you can post train it inside a harness pair it with an advisor tool and make it as good as opus or sonnet fi with a fractional cost. So, and and I think I'm pretty confident that like it'll get there. So, even if there's a situation where the Frontier Labs lobby Washington DC to ban the use of Chinese open source models in America, Nvidia is probably going to be very well positioned to take care of that.

23:08 >> Is it risky to put all of our hopes onto Nvidia? You know, especially when China has so many different AI labs working on open source. and I guess I'm also wondering whether Washington should be thinking about open source as, you know, strategic infrastructure. We talk all the time about chips and data centers as national priorities, but if open models are how developers, companies, governments, how they actually get to build with AI, do you think that the US should be doing more to support that ecosystem as well?

23:39 >> I think so. That's my opinion. it's and Jensen made these points in some interview recently that when a company when when a country has a lot more power than America like actual incremental power generated per year in the grid it's actually good for us if they build on our chips and our ecosystem and so it's better to be more open than closed and compete on capabilities rather than export controls. So that's my bu my my opinion is that and I I also feel like most of the capabilities that feel scary are stuff that you can build guard rails to fast enough if it's done in a more open manner and in general like all the software we take for granted today Linux particularly every whoever is using an Android phone is basically thank should be thankful to the fact that Linux exists and that and and and nobody's worrying about security on their phones. because even though it's open source software because Android and Linux are both open source and that allows people to patch vulnerabilities pretty fast and the same kind of capabilities can be made possible if they're open weights and people can post train it to be like u pretty good in terms of cyber security defenses.

24:55 It feels like when we talk about Washington and regulation, it feels very defensive when it comes to open source. Like what happens if we ban or restrict Chinese open source models over here? What about the idea of going on the offensive? If you were talking to lawmakers, what are some of the ideas you might float in terms of advocating or supporting or even investing in open source? If you want the benefits of AI to be widely distributed to small businesses in America and American allied countries, then you really need AI to be a lot more affordable. And open source is the only way to do that because when the models are open weights and can be served at the price of compute rather than the markups the labs charge on top, that's the only way small businesses can afford that. And that's the only way the economy truly moves forward. So it is in America's interest to actually support open source and open meets.

25:50 >> How do you do that? Could there be more investment in the university system? Do you think you know open AI is reportedly floating a stake for the government? Should the government be taking stakes in Nvidia or Reflection AI if that helps those open source ambitions? What do you think? >> I think Nvidia is doing the job of ensuring this is possible. so the and as far as the the Chinese relationships go I actually think we should continue supporting open source models using them here and competing with them on open source arena like and I think that's the competition is good like the best things usually happen when multiple amazing people compete with each other and and then out of the competition affordability problem is solved and businesses can flourish and there'll be a lot more AI enabled businesses business going forward and existing small businesses will actually be able to use AIS and you know like make themselves a lot more efficient and go out and and a lot more entrepreneurship will automatically flourish in America.

26:53 >> I want to get to this question of data. We touched on it a little bit earlier but you're hearing more people come out like Alex Karp was a high-profile example. Peter Fenton earlier today was very blunt that enterprises should be careful about giving away the context and workflow data that makes their businesses so valuable. does that apply to consumer AI products as well? And I mean you did give a good case an example of exactly how you can do that.

27:20 but like do you agree with that louder faction now that you need to protect your data? >> 100%. I mean the value of any enterprise is in the unique proprietary tokens you have. tacet knowledge is not even documented in the form of tokens that you should figure out a way to somehow pass it on to like your agent harness loop. So it's actually not in your interest to just rely on Frontier Labs. You have to have your own weights. You have to deploy it. You have to have your harness and you have to have this loop to keep hill climbing on this. So you have to work with companies that actually help you do this and be sovereign and and this applies to consumers too at the end of the day like why do you want to remain you know you you basically do you really want a repeat of the internet future where like all our data just lies in the hands of two or three companies or do you want a lot more control and privacy and sovereignty over that and that's what local AI enables for you once you set up your own model inside it and give it access to everything including audio at home, your like photos, whatever you want, right? you're not worried about it because the data is not even stored inside a data center and you could imagine future applications like physical and robotics at home which actually need a high frame rate of inference and can should probably be processing a lot of private data at your home. And do you really want that to be streamed to a server and like some frontier models making decisions for you there and surveilling you or do you actually want control and flexibility and like you know sovereignty and like have your own hardware that host local models. I I think there's like enough people who want control and privacy and so and and it's less about just the privacy argument. the cost argument is like pretty favorable for people who don't even care that much about privacy.

29:13 And that's actually why I see a future where everyone's going to have some some form of like data center at their home. >> Right. I want to get your take on something also I heard earlier Peter predicted that 90% plus of tokens could come from openw weight models over the next 18 to 24 months. what do you think? >> I think it be it could be amazing if that happens. I actually think what matters more is the valuable tokens more than the number of tokens. For example, AI search on Gemini probably serves a lot of tokens, but they're not necessarily as valuable as agents doing real work for you.

29:52 >> So, I think it's more important that we figure out a way to make agent loops be run with open weight models. That's actually where we should put a lot of energy into getting openweight models to be part of like most agent harnesses that serve most workloads right now inside enterprises and that's how you can maximize the token value produced by the output tokens from open weight models and you know that that's the metric that I care more about rather than the raw number of tokens >> and does that value to the frontier though I mean there is also this maybe misconception that open models can't be monetized Okay, >> value should acrue to the user and whichever company helps do that will win and so value acrewing to the frontier is not necessarily aligned with value acrewing to the user or the business.

30:46 Right? So a business can only flourish if users actually benefit in a long-term basis. If if the companies benefit more than the users, those companies are not long-term sustainable. the market forces will automatically ensure that and and so u when we talk think about AI and like who's going to get the value like fundamentally whoever maximizes the token value per watt per user will get the value and if the frontier is not in alignment with that objective it's not going to obviously get the value >> how do you measure that value >> you measure that by cost performance curves you want to you want to be on the to optimality there.

31:30 >> Arvin, I know like Perplexity started as a consumer AI company, but a lot of the things you're describing, you know, goes over to the enterprise as well. And like I said, you were kind of like the original model router. Are you seeing that business grow? are what are companies asking you about open source and model routing and harnesses and orchestration? Do they understand it fully yet? Even outside of tech, >> I think I think it obviously takes more education to be clear. Like first of all, harness, eval, agent loops, orchestration, sub aent, these are all like words that most people still don't fully understand. So it takes more education. So you have to simplify things and there's a lot of value you can create in the economy by simplifying things. as for like this consumer versus enterprise question I think there's a whole world where like there's a lot of small businesses and entrepreneurs solopreneurs are using our products and in fact the way they started using our business products our enterprise product is actually they were already consumer they were on the consumer version and they were paying for it and they converted themselves into like teams and business owners and now they're paying for it even more.

32:41 The amount of money we make from a company with sub 10 employees is sometimes even higher than companies where we've sold hundreds of seats. This is possible because even with a few seats, people are spending a lot of tokens for valuable business workflows, especially when they see the ROI. One particular customer, for example, their yearly revenue is around 3 million a year. It's a single person company with three or four contractors and the spend on our enterprise computer product is around $800,000 a year. So they're actually making that 2.2 million. And you know like that's pretty amazing if that business grows 10 times. that amount of spend like around eight million a year from like sub 10 people companies was never possible before and and that's why the SAS model is outdated and legacy and like like what matters in AI is truly delivering output value for the tokens.

33:34 If you can do that and help that business flourish is probably a better thing for you to do focusing on those people than trying to sell to legacy companies who don't even understand AI and who don't know what to do with it. And there's like a lot of small businesses. Last time I I looked at this, there was $3 trillion worth of payroll that's being spent by small businesses and entrepreneurs. And that's why it's very important for us to diffuse the output value of AI to all these people rather than the value of AI being locked in frontier companies and a few large mega corps.

34:06 >> So it's kind of a bottoms up approach and really interesting. I think that's how you know a lot of new trends especially in tech you know take off. They start more grassroots and you know I love the idea of looking at small business owners. >> I mean Apple is exactly like this. People don't think of Apple as an enterprise company. People think of it as a consumer devices company but a lot of adoption of Apple happen at the grassroots level from small businesses and they power a lot of American small businesses and you know that that's an inspirational example for us. Nvidia is an example like this. Nvidia initially sold to small companies and like graphics companies and gamers and even like AI labs that were really tiny, right?

34:47 >> And so it's a it's by actually delivering the value that they became much much larger than Intel or AMD. >> Well it's a great note to end on. Arvent your comments, your insights always seem to age very very well. so thanks for coming on and you covered a lot of ground. So thank you and hope to talk to you again soon. >> Thank you. Thank you for having me. >> So earlier today I spoke with benchmark general partner Peter Fenton and Alama CEO Jeff Morgan. Alama just raised a $65 million series B round. Benchmark and early investor and the timing here is pretty telling. Open models are moving from developer experimentation into real enterprise adoption as we also heard from Arvin just now. Now, Peter also made this really interesting prediction about how much AI usage could move to open weight models, what that could mean for the frontier model margins. Here is that conversation.

35:44 Jeff and Peter, thank you so much for making the time to chat with us today. Jeff, let's start with you because Alama, my understanding, is basically a way to download and run open AI models, open AI models yourself on your laptop server, inside your own infrastructure. I guess more broadly on this moment for open source, we've been talking so much about it. Did you expect it to have this kind of moment or momentum this quickly? What do you think made it possible and why is Alama sort of positioned to come out right now?

36:16 >> Yeah, thanks so much for having me. you know, Alama and and the open mall ecosystems grown way beyond our expectations. but you know from the beginning we saw Spark with developers where they could truly take control and ownership of their AI journey from you know that first kind of chat application they're building all the way to these longunning long horizon agents that thousands of businesses are talking to us about building with open models >> right and so I guess this certainly we saw this happen with developers you see the open router data you see that there's been a lot of adoption of these really good models that are reaching the frontier. But Peter, it feels like a moment right now where enterprise is waking up and realizing, hey, we can use these models. We can manage our costs a lot better and they're good enough or almost even as good as the frontier model. So do you think things have changed? What does it mean going forward for Alama and just this open source space?

37:13 Yeah, I think a maybe contrarian view that is becoming consensus is our belief that 90 plus percent of the tokens created will come out of openw weight models over the next 18 to 24 months possibly even by the end of the year and the forces behind that are twofold. one is you talk about cost but the inference margins generated by the frontier model companies I think are going to come under pressure when you can run those without the markup that they're providing when you have good enough models from open weights. but the second I think probably more important driving force of this is that the fine-tuning of the models and using smaller models gets you lower latency and higher performance. So we see a companies like Sierra and and you know the major applied AI companies an embrace of open models for performance reasons. So they're getting orders of magnitude reduction in time to to first token through the use of these open weight models that are tuned for the specific tasks. So that that force the inexurability of open weight models I think in the ecosystem suggests that that a for as far as the eye can see a super majority of the tokens generated will come out of these and then the question becomes which model and having a place of llamas is the destination for developers to discover which model is the best suited for their task and I think that's a major innovation in in empowering developers to give them choice and give them control >> right And I think that that is well understood in a place like the Bay Area and among applied AI companies. But Peter, if you think that 90% of token usage by the end of this year is going to come from open source, how do we get there? How do you get beyond the Sierras, the AI companies into the Fortune 500 because there's still hesitation to really embrace these models? Maybe you go first, Peter, and then Jeeoff, I'll get your your thoughts on this.

39:13 Yeah, I think that it's one of these cases where both models will thrive. Meaning I think the frontier models are continued to pursue the most complex use cases where you need to use the latest and greatest most capable model. But as the enterprises and and they do that through developer tools and through relationships with companies like fireworks are understanding that there are use cases that are good enough and that they can train and and then more importantly I think take internal confidential information where they have total control over the provenence and the you know where that data goes and how it gets used. I think that those use cases get lit up over this coming phase where we go from discovery, which is what the frontier models have have opened up the the possibilities of what you can use them for to optimization and ultimately you know driving excellence on on total cost which is an inevitable shift towards open models where they're good enough and and they can be run without the markup of the the foundation models on their inference.

40:16 Jeff, are you seeing that shift happen in your data? What kind of companies are making that shift? Is it still mostly tech or are you starting to see non- tech companies wait into this space and start to use more of these open source models? >> Lama's adopted by over 85% of the Fortune 500 and that's from some of the most regulated industries whether it's aviation, insurance, and a lot of that comes down to the trust relationship they can build by having these models run in their environment. and being able to do that securely and safely.

40:48 and what we've been most surprised by is that extends even to these largest cloud models that are are becoming extremely powerful. And you know being able to provide cloud computing environments in the US and Europe where these customers can take that trust and extend it to the data center is a motion we're really seeing. And so, you know, a customer might get started with a small model that they're running colloccated with their data and then as they further their journey with open models they're able to go and adopt these larger ones. And so we're seeing that across a surprising number of these industries you know even the most regulated ones like healthcare and it's truly one of the most powerful parts of open models.

41:32 So you're telling me that they're actually using open source models right now that many of the Fortune 500? >> Absolutely. And you know what's amazing is that even from the beginning we saw adoption quite early. Obviously the use case evolves right we went from a world where you know customers were setting up their own internal chat GPT type experiences for for their teams and their customers. and that's extended all the way to these really longunning coding coding agents. And that's all happening in every everywhere from, you know, a hospital to power plants, >> right? That's fascinating. So, you know, I guess over the last few years, has it been mostly built on top of the American labs open-source models, the GBT or the Google Gemma options? Do you see that shift moving to more of the Chinese models?

42:26 I think we see a really healthy mix of models from different geos. Obviously the Gemma models from Google deep mine are just incredible multimodal capabilities and then you know as of recent we've seen extremely solid longunning agentic coding models come out of some of the Chinese model labs. OAMA partners with all these model labs and so it's kind of our job to help, as Peter mentioned, bring models to developers, help them match it to a use case and we do see different use cases driven by different labs that that have their own focus.

42:59 >> Peter, when I look at sort of the adoption numbers, it feels like Gemma and you know, chat GBT's open-source model has you don't really see them on the adoption. It's really been taken over by these Chinese models. Have the Chinese run away with this leg of the race? >> I think just measured in tokens, they have a dominant share. No question. I mean GLM and I think people are highly anticipating the new deepseek model and I think it was the the the GLM the ZI model that had for most folks approximated what they were getting out of the Opus models and that that proximity for the agentic longunning more complicated reasoning tasks that the Chinese models have delivered I think has put them in if not you know absolute par within a you know large percentage of the use cases good off and you know one of the constraints in all of this is the compute and I think that's one of the variables that you have to talk about when you talk about open weight models is where does the inference compute sit and there'll be an interesting almost you know thermonuclear battle to to provide the compute to support the demand that's emerging for the openweight models given that they are delivering them at a significantly lower cost but again I think more importantly higher performance performance. So we see many of our applied AI companies forming long-term relationships with companies like fireworks together and and others to NIBUS, you know, the Neocloud constellation. Alam is also working with those providers so that that when it's good enough in the Chinese models that the developer doesn't have to then go through the brain damage of finding an inference provider, then standing it up and seeing if it works or doesn't work.

44:45 And and there's still so much complexity to making the openweight models effective. which is why when I say 90% of the tokens are generated by open weight models, it still may be that 90% of the revenue is going to the frontier models. So so the the pricing difference between the two is something on the order of three to five times more expensive today to generate a token on a frontier model than it is an openweight model. and I think over time that just puts real economic pressure on the system.

45:14 puts economic pressure on which system the >> build for the the frontier models to continue to deliver the the value that they're >> and now it may be having capacity is enough economic pressure release and I think right now since it is a capacity constraint environment we're able to see clearly where does this all settle down and capacity constraints create weird perversions in long-term economic models and they they would you believe they evaporate? Did do they evaporate in 2030 2035? It's it's it's anyone's guess.

45:52 >> Are you getting at the idea that infrastructure is scarce? Everyone needs compute. So maybe like the meta SpaceX playbook, we've seen them actually host compute for others instead of just their own models. Are you saying that could apply to open AI and anthropic the labs? Yeah, I think that the the the constraint on compute and power will be the mediating factor to the adoption rate of AI more than demand and >> right and then the second sorry the second part of what you're saying too Peter like I just will go back and say that you said that GLM and I you're not the first person to say this I hear this from tons of people in tech that I talked to that it's almost as good as in some cases better than the latest you know opus model from anthropics So, like I'm kind of wondering if as that pricing pressure weighs on the labs, do they cash in on the compute that they've already bought or rented to host for say open source models?

46:52 >> I think that's unlikely. you know, there's an era where Microsoft fought forever Linux and then they embraced it. So, it would require their business models to evolve from being pursuing AGI and pushing the frontier. I think what's more likely to happen is that they move into vertical applications is that they go proverbally up the stack to places where the margins are more defensible from the generalized intelligence moving you know into mainstream and going deep in places that they've already announced design financial etc medical or biotech so so I think it's more likely given their strategies that they pursue more differentiation then they try and you know leverage their compute access over the mid to long term. You have to wonder will they deliver the lowest cost token given their scale or will there be some interesting fragmentation where there's certain use cases that are just delivered with a more almost bare metal you know kind of open weight mode approach the the thing which I think is sort of again under underappreciated is that the applied AI companies aren't going to open weight models for cost reasons exclusively I think most of them Sierra has said this publicly it's driven by performance so So being able to have direct control over the model to fine-tune it and then once a task becomes understood and bounded, it wants to go to a high performance low cost you know of inference model and and there'll always be the next frontier I think of intelligence where you need the latest and greatest but you know a a large bulk of the market's going to be served by good enough open source where performance is going to be a driver which then ultimately will show up as cost.

48:36 >> That's really interesting. So cost is almost the bonus. Whereas I think there's a tendency especially on Wall Street to still think of you just go to the open source model for cost not necessarily performance. You have to give something up. But what you're saying I think Peter is that that fine-tuning is really valuable. Jeff, do you see that? Do you see I last week we talked a lot about the misconceptions of open source? because they come out of China. I think there's still a lot of companies that are worried about security issues. You solve that, right?

49:05 because you're hosting them here in the United States. what kind of questions are you getting from corporations that are now interested in open source and maybe looking to shift some of their workloads? >> Yeah, that's exactly right. What we see is that, you know, there's one thing is where the model's from and where it was created and trained, but the more important thing to these businesses we speak to is where it runs and how it runs. And what they're looking for there is first off a way to run it, you know, in a a geographic location that's where their business operates, whether that's US or Europe, as an example. And the second thing is then tooling to run it safely. And, you know, this works into the performance piece, which is, you know, can can you bring the model closer to your data? C can you then monitor it, make sure it's running correctly?

49:52 these are all things that, you know, the Frontier model labs provide as part of their product. But for an open model where you receive the weights, that's something customers are left to do on their own right now. and that's where Alama's cloud comes in and and really where, you know, we can provide a secure environment where you can access the power of these open models and customize them. but also get, you know, the safety that you get from a frontier model >> lab.

50:15 Peter, I've talked a lot about sort of the implications of building on Chinese models and sort of the longer term, right? And if the whole world is going to be building on essentially a Chinese model, then it's going to be a Chinese ecosystem and there's risks involved with that. Do you think that American AI American companies are paying enough attention or should be paying more attention to developing our own open source models? >> Well, I think the open weight model ecosystem benefits from heterogeneity from multiple geographies. I would love to see even some of the frontier model companies take more aggressive stance with this where they recognize that you know you have an ecosystem that will put pressure to have control and access for developers which is something that you know OAMA is enabling. It's why you have 9 million developers that are coming there discovering and and building. And if you built your business where there's a wall behind that that that prevents, you know, critical mass and ubiquity, I think it ultimately positions you to be, you know, a closed system. And I I think in the history of technology, those strategies can leave you isolated in the high end of the market. So I think it's a mistake to not have a strategy that gets you the super majority of developers. And if you aren't achieving that, that suggests you will you'll lose some major part of the market.

51:36 And you mentioned Microsoft and Linux and maybe that's sort of how how this evolves. But I guess Peter, do you think that there's more that washing kit can do to create those incentives? And I also wonder if you think that like openw weight AI is as important as say chips or infrastructure in the AI race and it feels like an area that you know policy makers don't know what to do with. They want to regulate it potentially if it's going to keep China out. Are there things they could or should be doing to actually encourage that development if it's so important as you kind of laid out?

52:11 >> Well, I I think it's in part to not pervert the system by blocking access. There's some discussion that that might be a coherent strategy to limit access to open weight models. I think that'll have the exact opposite effect. It'll disadvantage the US technology market. it will globally those those those models will not be blocked from other regions that end up creating critical mass that I think isolates the United States. So my hope is that there's a there's a pro-active, you know, embrace and and support open and try to be at the same level of multiple regions including chi China.

52:49 And I think that's a at risk. the the risk we've talked about of of what might happen in the regulatory capture which I don't think will happen but I think there's a risk that the frontier models are able to encourage a strategy that would be counterproductive even for them because I think ultimately they're going to want to have the super majority developers strategy that doesn't achieve that I was a losing strategy >> right and you know I think there's also another misconception that you can't monetize open source AI which I think the Chinese labs are kind of proving you can through API access in other ways.

53:22 Jeeoff, I wonder if you agree with what Peter's been saying and also do you think like what would you tell a policy maker in Washington that's looking at this wave of momentum for open source and seeing it come out of China? >> I think the most important thing to do is to to really get to know the customers, especially the US-based ones, and what they're trying to accomplish with their AI journey. And in doing that, learning, you know, to to Peter's point, what is truly productive for these businesses and how can we help them adopt AI as fast as possible in a way that's safe and secure and I think a lot of the points Peter touched on are right and they they may be counterproductive. And I think the first step is to really get get to know these customers and talk to them and understand their needs.

54:06 >> Right. And then I guess this is a question from Jasmine, my producer Jeff. is it important that this developer tool, what Alama is, is it important that it's American and does that help when you talk to companies that are worried about security in Chinese models? >> Yeah, Llama's part has partners with companies around the world, whether it's US, Europe, >> you know, we helped launch a Canadian coding model a few weeks ago. And, I think it's really important that, you know, back to why these businesses are adopting open models, it's trust. and that there's there are there is an entity that that can be trusted to access these models. I think it's very important and that's a common, you know, piece of feedback we hear from customers all the time.

54:48 >> Peter, I wonder if you can weigh in on this kind of debate that I feel like has broken out into the open just even over the last few days that really centers around data. I don't know if you saw Alex Karp from Palanteer was on CNBC and he sort of railed against the frontier models and said why would enterprise want to give you our valuable data so that you can train off of it. We want to own more of it and I wonder if you think that that's right and if sort of companies should be more protective of their data from the big labs.

55:18 Well, I think there's a clear translation from open and closed models into economic value within enterprise which requires context and orchestration and the context layer of the enterprise provides where investors in a company Merkore which provides that sort of task mapping of the internals of a company their their workflows their proprietary data onto the models and the idea that you would you know put that back into the frontier models is incoherent. So, so a major mandate to apply AI to productive use cases requires capturing the internal workflow data in a way that you you have ownership of it and also portability. So again, one of the major benefits of open weight models is that you can pour the the and this is one of the one of the long-term challenges for margins I think in in AI in general, which is the switching cost from model to models only a function of rehydrating your context.

56:17 So so the idea that enterprises would seed that context to a closed model company seems incoherent to me. So ownership of that of of your internal workflows of your proprietary knowhow and expertise is sort of where you think about why is an employee more valuable to one company versus another. Well, they they've trained them, they've given them context, they've orchestrated them. And I and I think that's the mandate for enterprises is to develop an ownership and and invest in that because it's it's what allows them to get the benefit of those models applied to their business as as opposed to giving it to those models and then losing it to their competitive product down the line.

56:54 If a portfolio company asks you, should we, you know, work with forward deployed engineers from OpenAI or Anthropic? What would you tell them? Peter, >> I would have to argue with them that there may be better choices. Specifically on this question of who will own the data, how portable will it be to other models? And I think there's a series of questions I would ask that would likely get them to say, we don't want to engage in a contract that requires us to do that. I think anthropic and openi would resist those those protections and it's it's why there should be a robust independent ecosystem that's able to then ride on top of all these heterogeneous models that are going to continue to come out that can allow you to pick the right model for the right task as opposed to you know vertically integrating them.

57:41 >> Okay, last question. Peter, you already kind of answered this so I'll go to you Jeff. Three years from now, what does the enterprise AI stack look like? is it one frontier model plus a bunch of tools? Is it a messy portfolio of open, close, local, hosted, cheap, expensive, specialized models? what do you think the majority of companies are on and specifically non- tech too? >> Yeah, I think like Peter said, it'll be a mix. but that super majority of the the the tokens will be from open models. And if you think about what these businesses are doing, you know, the average Fortune 500 business may have thousands of internal applications and they're modernizing all these with AI. And for that to truly be productive and in the best interest of the business, it's to use open models and increasingly open models that are colloccated with the data. And so I think we're going to see a transformation of these applications into AI augmented agents and and I think in in you know three years from now these new AI agents are going to be processing hundreds of millions of tokens again a super majority of that being open models just because they're so efficient. They're lower latency and they're higher performance for their use cases.

58:51 >> Right. Okay. So Peter, last one for you then. can we expect you to see can we expect to see you do more of these deals leaning in? what's exciting you right now about this space? >> Absolutely. I think the idea of developers empowered is what fuels our industry and we resist early oligarchs igopies I should say. And so I think one of the the predicates of entrepreneurship is creative destruction. So that means giving tools to the creators, the developers to enable things we couldn't even imagine.

59:27 And I I am very optimistic that instead of being a igopoly, open weight models and open source is enabling a constellation of these use cases to emerge from the creativity of developers. And so we we are maximally interested in that. And then one other thing which I think comes out of this is the idea that we're now opening up a whole set of use cases around long horizon agents. agents that'll be working over many days, weeks, and even months. And and part of that requires the the sorts of optimizations that open weight models give developers. So we we are extremely excited about that opportunity set.

60:02 >> Right. Well said and very missionoriented. I'm so glad that we were able to talk to both of you and have that wide ranging conversation. I hope you will both come back and help us and help our audience figure out what's going on. It is moving very very quickly. That's why I love it and I could do this every day. So, thank you so much, guys. Talk to you again soon, I hope. Some great conversations today on a postfrontier world. We're going to post this, as a new video, so make sure to catch that. And thank you as always to Sammy, Jasmine, Robert, Evan, Jordan.

60:34 I will be on vacation for the next few weeks. I'm going to be in Canada, so we'll put the live stream on pause, but can't wait to get back and we'll see you then. Thanks for watching.

Summary

AI is transitioning into a "postfrontier" era where the focus is shifting from merely developing larger models to creating efficient systems that utilize various models for specific tasks. This change is driven by the need for cost control and the increasing adoption of open-weight models, which are projected to dominate token generation in the near future.

- The model alone is no longer the product; the orchestration system that integrates multiple models is key.
- Over 90% of tokens may come from open-weight models within the next 18-24 months, potentially by the end of this year.
- Companies are moving from relying on public benchmarks to creating their own evaluations based on specific needs.
- The importance of local computing is rising, allowing businesses to run models securely and efficiently without relying on external data centers.
- Open models are becoming increasingly trusted, with significant adoption among Fortune 500 companies across various industries.
- The competitive landscape is shifting, with open-source models proving to be cost-effective and high-performing alternatives to frontier models.
- Companies are encouraged to retain control over their proprietary data and workflows, rather than ceding them to large AI labs.
- The future of enterprise AI will likely involve a mix of open and closed models, with a significant emphasis on open-source solutions for efficiency and performance.

Questions Answered

What is the significance of open models in the current AI landscape?

The AI industry is transitioning from a focus on proprietary models to open models, which are expected to dominate the market. Open models will provide cost control, accessibility, and empower developers, challenging the dominance of a few frontier labs.

How will the scaling of AI models impact their deployment in businesses?

There is still significant potential for scaling AI models, and both frontier and open-source models are expected to advance. However, the focus should be on deploying these models effectively to empower small businesses and enterprises.

Why is open source software important for cybersecurity?

Open source software, like Linux, allows for rapid patching of vulnerabilities, which can enhance cybersecurity. Advocating for open-source models can similarly improve defenses against cyber threats.

What trends are emerging in enterprise adoption of AI models?

Enterprises are increasingly recognizing the potential of open models for managing costs and achieving performance comparable to frontier models. This shift indicates a growing acceptance of open-source solutions in business applications.

What are the main concerns corporations have regarding open-source AI models?

Corporations are primarily concerned about the origin of AI models and their operational security. They seek assurance that open models can be run safely and effectively within their geographic and regulatory frameworks.

© transcribe · For agents Built with care and craft by Gokul Rajaram