transcribe

MFU Is the Most Important Metric in AI — Here's Why | Gavin Baker at Aria Networks Launch

Aria Networks · 9m · transcribed Jun 2026
More from Aria Networks Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:01 [music] [applause] >> Well, first I want to Well, first I wanted to welcome everyone here. I'm I'm thrilled to see all of you guys here participate in our launch event. This has been an exciting week for for us at at Arya. We finally can talk publicly about the details of what we're doing. And so, you know, I want to thank our customers, our partners, our investors, our board members, everybody is here.

0:37 And we look forward to showing you more about what we've been building today. I'm going to invite to the stage Gavin Baker. Gavin, it's good to see you. Good to see you. Thank you for having me, Mansour. Of course. Well, first thank you I know you flew here for this event, so really appreciate appreciate that. And you just joined our board, so I want to welcome you to our board and look forward to be working with you.

1:05 So um what I wanted to do here is is explore the economics of AI factories. You've talked extensively about MFU, a model flop utilization. You've talked extensively about a token efficiency. You've stated that for AI factories to be successful, they need to be the lowest cost producers of tokens. Can you break this down for us? Yeah, absolutely. And and I actually think it's really it it it's interesting.

1:37 And it's awesome for everyone in this room. But at no point in my career has an investor in tech has a big winner been because they were the low-cost producer of anything. It's not like Apple is worth 3 trillion cuz they're the low-cost producer of smartphones. Nvidia is not worth 5 trillion or whatever it is today because they're the low-cost producer of accelerators of silicon. They're just the best. AI is fundamentally different in that um at some level, the amount of intelligence you get out of an answer is determined by two things. The token efficiency, how many tokens you need to get that answer.

2:15 Um other people call that intelligence density. We'll We'll see which term wins. Um I like token efficiency and then cost per token. So, for a given level of intelligence, the person who can get there with the fewest tokens and the lowest cost per tokens is going to win. Because and this is this is something that is I I think uh I started actually um in aggregates, rock pits. And in any extraction industry, oil and gas, metals and mining, the low-cost producer always wins. And we We haven't had that in tech.

2:50 >> [clears throat] >> And if you were the low-cost producer, what's up, Thomas? If you were the um If you were the low-cost producer of intelligence, you can choose between actually generating more tokens and having a higher-quality product or you can choose >> [clears throat] >> to have higher margins and invest that money in user acquisition, customer service, all these things. Now, token efficiency, that is fundamentally the job of the labs. Full stop. But everyone in this room, presumably, is here lowest cost per token. Um and there is a trade-off as Dylan referenced. You know, if you can make really, really fast tokens, maybe maybe cost is less important. Um but strangely enough, the way you get the lowest cost per token is you spend a lot on everything around the accelerator, which makes perfect sense from first principles. If these data centers are AI factories, you know, in in my days following manufacturing industries, factory utilization is a key input.

3:56 In airlines, it's how much you are the seats full? And the truth is is that GPUs, TPUs, Traniums, they are rarely running at full utilization. So, we have spent hundreds of billions of dollars on these AI factories and they are running at low levels of utilization. And the network is an one of the one of the biggest culprits in this.

4:26 >> Right. So, that's why I'm that's why I invested and I'm very excited. >> [laughter] >> Yes. Well, that's indeed actually, if you think of about the utilization of these accelerators, uh it's around 30 to 45% maybe for training, even lower for distributed inference. And so, like you pay for all the flops in the accelerators and you're only getting, you know, a very small fraction to, you know, towards the job. And the larger the cluster, the more uh inefficient it is.

4:56 Absolutely. >> the network is connecting all of it. And as the world Maybe maybe you can comment on this. As the world moved from training to more inference at scale and especially distributed inference like in these high inter activity use cases that uh Dylan mentioned, you know, the network becomes even more of a of a bottleneck. Absolutely. Yeah, I mean, especially with, you know, the expert parallelism and the disaggregation of prefill And the codes, right? and decode. I mean, the network is more important than ever.

5:28 And just if you take your utilization from 40 to 80, you've dropped your cost per token in half and you've literally and you have dramatically increased your margins. And it's crazy and on top of all of this, not only are we have we spent hundreds of billions of dollars on these AI factories that are sub-optimized often because of networking. Um but we are increasingly in a what constrained world. So, you know, seven states are have bills to make it illegal to build data centers.

6:00 And so it you know, you may you want to get if you have a data center in the ground, it is crazy not to run those GPUs as close to 100% GPUs or TPUs. Um >> XPU's XPU's XPU's >> XPU's XPU's is the word. XPU's XPU's as as close to I I see uh Hassan from Broadcom. So, XPU's not GPUs. Um It's crazy not to run them as close to 100% utilization as you can. And then you can make all sorts of trade-offs between interactive interactivity and and latency >> right? and hit hit the right um target for your workload and your users. Right.

6:38 And then just I'd like to more on the network. Uh for training traditionally it was the back end only scale out, but for distributed inference it's also the front end. Because that's where a lot of when the when the query comes in it's being fanned out into all these agents that are converting back and forth into the token space. And then it and then you have you know, to to the point we made the decode and then the the pre-fill decode and transfer of KV cache. And and even oh even a level up when you're like training costs they're kind of obfuscated, but you feel it has a user inference costs. these >> So, I just think training, you know, if you're if you're Google or if you're one of these resource rich resource rich mag seven companies, maybe you can live with a low utilization of your AI factory during training cuz you really don't care. You have some huge monopoly. I mean, you care you care a little bit. But during inference, you are out there competing for users every day.

7:40 Most applications have a router. They're constantly routing to where they get the most intelligence per dollar. And so inference makes these tokenomics, which I believe Dylan coined, really explicit and essential for the economics of everything. [snorts] So just a level up, just inference makes costs important in a way they never were before. >> Right. That's indeed. Actually, it's all it's about the utilization of these accelerators. And also, just from an interactivity standpoint, time to first token, a lot has to happen. And that involves the network to get to that first token.

8:18 And >> And we know the history of the internet is every 100 millisecond delay Right. in Google search results cost them 1% of revenue. Exactly. So I can't think of the number of times that I've like, you know, like Claude has a good trick now where they start they they have some generic words they say at the beginning every time. They really have they sometimes have to do with your query, but the number of times I've abandoned something because I was waiting. I guess. Yeah. So it's it's everything.

8:48 >> Indeed. The two numbers are token to time to first token and then inter-token latency, right? From every token to the next. And then again, and if it stalls at some point it stops. That's you know I would say likely the network, right? So Almost certainly the network, yeah. I like seeing that answer fast, that first token. I like seeing it populate quickly. And then I want to get pay the lowest cost I can for the um you know, the intelligence I'm getting.

9:13 Excellent. Well, again, we're going to bring you back on stage to wrap up at the end, but thank you so much. Thanks Mance. Thank you. >> [applause] >> Who do I give this to?

Summary

The launch event for Arya focused on the economics of AI factories, emphasizing the importance of token efficiency and cost per token in achieving success in the AI industry. Gavin Baker highlighted that unlike traditional tech companies, the low-cost producer model is crucial for AI, where maximizing the utilization of AI accelerators and optimizing network performance can significantly reduce costs and improve margins.

- AI factories must prioritize being the lowest cost producers of tokens to succeed.
- Token efficiency (intelligence density) and cost per token are critical metrics for AI performance.
- Current AI accelerators operate at low utilization rates (30-45%), leading to inefficiencies.
- The network plays a significant role in the performance and cost of AI, especially during distributed inference.
- Increasing accelerator utilization can halve the cost per token and improve margins.
- Inference costs are becoming more critical as competition for users intensifies.
- Delays in response times can lead to significant revenue losses, highlighting the need for fast token generation.
- The transition from training to inference requires a focus on network optimization to enhance performance.
© transcribe · For agents Built with care and craft by Gokul Rajaram