transcribe

Why The Current AI Infrastructure Is Built For the Wrong Future | Sail Research CEO Neil Movva

AGI House · 59m · transcribed 23d ago
More from AGI House Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

The Importance of Efficient Chip Utilization

What does it mean to have idle GPUs in a data center?

Idle GPUs represent wasted potential and resources in a data center, as they are designed for high performance but not being utilized effectively. The speaker emphasizes the need to maximize GPU usage to recognize the engineering effort behind their design.

  • There are no bad chips, only bad prices.
  • Creative solutions can allow large models to run on chips with limited memory.
  • The AI industry is facing a paradox of falling token prices and rising enterprise AI costs.
  • Maximizing GPU utilization is crucial for efficiency.
# 11:58

The Shift in AI Infrastructure Needs

Why is it important to rethink AI infrastructure?

As AI workloads evolve from quick responses to long-running agents, the infrastructure must adapt accordingly. This includes rethinking chip selection and data center design to support sustained operations.

  • AI infrastructure takes time to build and must evolve with changing workloads.
  • The shift from human intervention to autonomous agents requires a new approach to AI infrastructure.
  • Investing in infrastructure today is essential for future AI capabilities.
# 23:56

Navigating Model Routing Challenges

How can users effectively choose AI models?

Users face challenges in keeping up with the rapid release of new AI models. Companies like RAMP are emerging to assist by automatically selecting the most suitable model for specific tasks, alleviating the burden on users.

  • The pace of new AI model releases can overwhelm users.
  • Model routing companies are developing solutions to simplify model selection.
  • Automated model selection can enhance efficiency for enterprises.
# 35:54

The Future of Chip Development

What are the challenges and opportunities in chip development?

Chip development is a complex and lengthy process, often requiring multiple iterations before success. However, there is growing interest and investment in specialized chips, indicating a potential for innovation in the industry.

  • Chip development is challenging and time-consuming.
  • There is a rising appetite for specialized chips in the market.
  • Successful chip launches can lead to significant advancements in technology.
# 47:52

Revolutionizing Memory and Compute with Sparse Attention

How does sparse attention impact chip design and memory usage?

Sparse attention changes the dynamics of memory usage in AI models by decoupling capacity from bandwidth. This allows for the use of alternative memory technologies, potentially reducing costs and improving efficiency in processing long sequences.

  • Sparse attention allows for more efficient memory usage in AI models.
  • Decoupling memory capacity from bandwidth can lead to cost savings.
  • Innovations in memory technology are essential for scaling AI capabilities.

Transcript

0:00 We have this motto at sale which is that there are no bad chips only bad prices. >> Sure. >> So you can always find another chip that can do the task you want but you you might have to get creative with how you serve the model on that chip. I'm talking in particular about fitting large models on chips with small memory which increasingly dominates the cost of these chips. And so if you can get clever with that and you can work around differences in network topology as well, you can actually get to a really efficient serving regime on chips that are not traditionally used.

0:34 Here's a strange fact about the AI industry. P token price keep falling and yet enterprise AI bills keep tripling. My guest today says that because of the entire inference stack was built for a wrong future built for the chatbot that answer in milliseconds when the real workload is becoming agents that run for hours, days, even weeks. He and his co-founder just came out of stales with $80 million from client Perkins and Sukoya to rebuild that stack from the CUDA kernel up. Neil Moa, co-founder and CEO of sale research. Welcome to Ajas.

1:09 >> Thank you for having me. All right, before we get into any of it, you said public that you feel emotional pain when you see a GPU s idle. Walk me into that. What does a waste GPU actually look like physically in a data center right now? >> Yeah. Yeah, absolutely. So I I think I have the privilege of having seen some of these GPUs be designed and I know how much work goes into squeezing every piece of silicon to deliver the most performance that it possibly can at the design stage when you're at Nvidia or Apple or any of these other companies making trips. And so to me it's very important that all that work be recognized by pushing that GPU to its full limit when it's being deployed in a data center. That means running at full power utilization. means running and doing the most work per cycle that's possible and getting the peak flops out of that machine.

2:02 >> Got it. Well, let's go back to the original story. you met your co-founder Samir on your first day at Stanford. You jokingly he got a slightly better grades. What did you actually see him say a decade later? This is the person I want to build a company with. >> Yeah, I think building a company is an extremely extremely intense and high trust process. there. We're going to be pushed in ways that we can't even imagine today on the horizon. And you know, the only way I can imagine doing that is with someone I've known for as long as I've known Samir. in college, we got along very well because we were basically very similar to each other in the in the freshman dorm we were assigned to. We were two kids out of 80 in that freshman dorm who just really liked computers and really liked talking about them all the time. So, we took to each other very quickly. And outside of of course our classes and work, we were just great personal friends. We shared a lot of friend groups and we ended up living together actually all four years through Stanford. we're very close to each other and we've lived together in SF ever since graduation for about six years now. and I think you know what he's really strong at is the software side. he continues to be one of the best programmers I've ever met.

3:12 That was true on day one at Stanford and it's still true today. and I always lean more hardware and I think together we have a good mix of the whole stack from chips all the way to APIs. >> Amazing. I know from the background he runs a security engineering at Apple and you did a vision silicon and the GPU kernels. Well, how do you guys split the the work at the sale today? >> Yeah. Yeah. Well, first Samir CTO, so really all the hard problems fall to him. He he does an incredible job the whole stack through. but you know historically and kind of like what our I think natural affinities are for everything that's close to the chip I would say everything that relates to squeezing chip performance that's really something that I'm excited about and I spent most of my time thinking about and and Samir thinks about how to take those chips this fleet of chips that are thousands across the world and how do we turn that into a reliable service that can actually serve trillions of tokens a week. that is no small feat and he does an incredible job. He and the team do an incredible job of making that reliable.

4:12 >> Amazing. Trillions of tokens per week. That's a big number. Yes. >> Yeah. >> What's your worst co-founder fight so far, if there's any? >> How did you resolve it? >> That's a great question. So, one of the things I I will absolutely say about starting a company with your best friend, it's very good to do that because you've already had lots of fights. That's the thing you're buying is you you've already gone through all the fights you could possibly have.

4:33 And so anything that you have to fight over in the course of the company, you already have a good way to deal with it and you know you're going to be able to make up after. So Samir and I have mentioned for 11 years now. We fought a lot over a lot of different things. usually today we actually don't fight much anymore. It's like, you know, we put all of our energy and emotion into the company. Now we're really each other's support system. But in the past, you know, we were hotaded young nerds.

4:55 We would fight over the stupidest things. Usually some technical topic. you know, I would say something like, "Oh, the chip is capable of this in theory." and he would say you know it's impossible that you achieve that you know here's all the things that actually stand between us and that achieving that global performance and so we'd have these very theoretical debates but anyway it's really good to work with your best friend >> amazing good to hear that yeah your resume is a tool of the efficiency frontier Nvidia right as a gaming priv to AI in 2016 to 17 and Apple building what do you call the world's most efficient computer vision algorithms on iPhone silicon and then the first colonel high at a together AI. What are those jobs most shapes how you build a sale?

5:37 >> That's a great great question. I think throughout all those jobs, I think there's one really northstar vision that I carried forward onto sale, which is this drive to push efficiency to the frontier. We care about doing the best we possibly can per mill, per per millimeter of silicon, or per millisecond. like in all cases there's a very scarce resource, time, energy or silicon and we want to do their best to squeeze it to the limit. and Nvidia actually has this term that I like a lot and I I say all the time at sale. It's this idea that we're chasing the speed of light. We don't just ask what's good enough. We ask what's physically possible of the system and we want to go to that limit of physics and nothing short of that. and that's a tough mandate. it means we are always squeezing a little bit more performance.

6:24 It means we are when we design a system we think about what is it going to take to achieve the happiest path for this whole system. What does it take to have every cycle utilized at 100%. in particular Apple I would say is also they are a bit of a trial by fire type company. You know there's a fixed budget for what's able to be shipped in an iPhone every year on a cost basis on a power basis and on a timeline basis too. You have to make these things happen on a unrelenting cycle. Every September there has to be a new iPhone.

6:54 And so those constraints really focus a roadmap in a way that I think was really challenging for the pace of AI development. But I learned a lot from some really great engineers there. >> What did you learn about that told you a different company need to exist? >> Yeah. So I think I think together has always been focused on the speed frontier. They they've always been you know producing the highest tokens per second on any new model that comes out and there's immense work that goes into that. It's it's a very challenging job.

7:21 That is what I describe as kind of the frontier of low latency inference. and so you have companies like voice companies who who absolutely love the work that together does like decagon for example. they really need the speed that together provides. But what I noticed over the course of my time there is of course in the early days of AI we had mostly chatbots. We had mostly interactive use cases. the human was very much in the loop and it made a lot of sense to push on latency and driving that down was the only way to get great applications out to users. Over time, we have gotten to the point where agents can run autonomously for much longer.

7:56 And the simplest thing that means is that they need a lot more background throughput. They need to do stuff working for hours without human interaction. and in that world, efficiency is the most important thing. It's not so important anymore that you ask a question and get an answer in 5 seconds. It's more important that the work be done really well over the course of hours or days just like you'd expect from a human employee. And if that's the goal with agents, I think the infrastructure we need for that is going to look quite different.

8:24 >> Well, that's a great segue to actually the second question. Yeah. that's stay the CC's plainly. the fortune headline that was saying a former Apple engineer thinks AI infrastructure is beautiful for the wrong future and what's wrong with the current future and what's the right one? I guess you are leading try towards that. >> Yeah. You know this is a fundamental trade-off in any system. Whenever you design a chip or a data center or a software engine, you have to make a choice between are we optimizing for throughput >> Yeah.

8:53 >> as in tokens per hour the aggregate efficiency of the system or are we optimizing for low latency. >> You can choose but you must choose. >> Yeah. >> And I would argue that so far we've chosen low latency across the stack. We really really focus on using really great chips from Nvidia that can produce lots of tokens per second per user. And with agents, the bigger constraint I believe is actually more tokens per dollar. It's the tokconomics that allow you to actually spend way more tokens on a hard problem and have a chance at getting great results. I liked Jerry Tourk's talk a few weeks ago at AGI House where he >> basically laid out the thesis for test time compute scaling. Yes, >> you give more tokens to a problem, it will it will get solved. That is the future we're building towards. And in that world, the only limiter on your ambition is scale, efficiency, economics. We seek to deliver the best economics, the best scale for test time compute.

9:52 >> Amazing. So yeah, like you mentioned, agents need scale basically the the main core things like reliability, sustainable cost. I guess that's my big issue, right? Humans need a speed, I guess, when we interacting with people. that work for us or doing the work for us. So when did that click for you? Was there a specific workload or customer conversation or is that something that makes you wondering like why is okay this is the future of building and then wait a minute I think that's something we should go this direction >> you know when we started the company I don't think it was super obvious we started the company December of 2025 >> this is just around the time that claude is really getting good at at aentic workflows but still very human in the loop this is around the time that cla 4.5 is released 4.5 and everyone has a great winter break that just doing amazing things with cloud code and it was 4.5 but that was very human in the loop still. It was amazing what agents could do and we could start to see the affirmative vision of the future where agents are doing a lot more interesting work but it wasn't quite a background agent world yet or it wasn't very much like a proactive autonomous agents running for hours type world yet >> and so it took a leap of faith it took some amount of >> introspection at that moment in time and saying all right let's play this out we are excited about where we are today but what's what are we going to be talking document in three months time and then a year's time and ideally two or three years time from and I think the main axis on which I could project anything out was going to be the amount of time between human interventions. At that point in time, December 2025, I would argue that the typical cloud code session had a user interaction every few minutes. You would type a command, claude would go do something, take a few turns, come back to you in two to five minutes.

11:35 >> over time, that number has gone up and crept up slowly. Now it's like 10 minutes. You can go give a task to Claude. And in some cases, you can give Claude an hourlong task and you'll be happy with the results. That number has to keep going up. Yeah, >> that is what success means. It's what autonomy is all about. And so what does that look like when you have Claude working on a task for a day? That's what we're very excited about.

11:56 >> Amazing. Now looking back looks like a very obvious but at that time definitely you need to take a leap of faith to thinking about that because I'm a heavy user of the AI. I use the agents like at that time because humans still in the loop so you expect to run faster, right? But when you are getting slowly phase out human intervention takes longer and longer. So the whole philosophy starting to switch around. Right. >> Absolutely. And it's very important that we started early because infrastructure great infrastructure takes a lot of time to build. Yes.

12:24 >> We are trying to build the entire stack from as I've mentioned the low-level software to maybe even choosing chips differently than how you would pick for low latency and maybe even choosing different data centers. >> Yes. >> And so that whole stack the physical world does not move at the speed of software. And if you have an opinion on how that should look, better express it today and start making progress there. >> Amazing. I love the way how you think of it. Like even you need to rethink about a data center, the whole stack to rebuild.

12:51 >> Absolutely. >> Yeah. We we have a strong bias at Asia House for that. We're trying to rebuild the future of society around this foundation model AI technology innovation here. >> So so enterprise AI bills are tripling while to a token price fall. So unpack that paradox for the audience. >> Yeah. Where's the money actually spent? Yeah. >> Yeah. I think there's a few forces at play here. Number one is that this is the first wave of major AI adoption in the enterprise. Sure. frankly, last year there wasn't much. It was very much it was only startups and AI native companies that were actually seriously consuming tokens. I think cursor was still the largest consumer of tokens for the open source model community like fireworks together and base 10 last year.

13:32 this is the first year that we've had a shot at getting into the enterprise and I think where enterprise is today is still mostly experimenting but in the way that is basically adopting cloud code and similar workflows as like a first pass of getting AI into the organization. I still don't think we've seen really a substantial amount of deployment yet. We're we're in for a lot of growth here. but that said, most enterprises adopt something like claude early on.

13:58 >> It is the easiest thing to deploy. It is a familiar interface. Everyone's used chatbt or cla at this point and so they start with that and that's part of the reason the token bills are so high. Cloud is not a economy option. They do not try to be a >> not yet at least. >> Yes. Exactly. They they don't work as hard at the let's say the costefficient frontier. OpenAI has done a lot of work in this direction. But the real solution to great great economics on tokens are open source models. Yes, >> I believe open source models have always been a story about cost and efficiency.

14:31 The way to win is to continue to specialize models. For example, you're going to see more open source models that are fine-tuned and pushed in directions that are unique to certain verticals. for example, Harvey's done great work recently on postering their own models for legal work. and that's one great path that they have to improve their economics. And of course on the serving side, now that the models are open and I have access to the weights, I can get really creative on how I would change my chips and data center choices in order to serve those models really efficiently. and that's the kind of innovation that we're excited about. more open models give me more options in how to serve them and that lets me drive the ROI per token to be dramatically better than it was last year. so we'll see improvement in the enterprise just >> amazing. Yeah. Cool. Looking forward to that. a follow on question that was advocate latency still matters for the human reviewing the agents that work right we just mentioned that >> where does latency still win and where does where do you refuse to compete >> yeah so an obvious one is voice inference I think voice inference is something that you will demand the lowest latency possible it's critical to the experience that's not going to change and I think voice inference is still growing quite well for example Sierra just acquired a great company called takeoff which is building the excellent long horizon agents that interact with the customer over many touch points.

15:51 >> and in that world there's a good amount of voice interaction too because people like to talk on the phone with businesses. >> Yeah. >> and so in that world I think you're going to continue to see good demand for for voice models when they touch the customer. But there's going to be even more work in the back office and that's what we're excited about is what can we do in the back office to make that more streamlined, efficient and intelligent.

16:11 >> Got it. Yeah. Basically when there was like a more interactivity like there that's demands like a low latency right yeah so serving agents efficiently is a deep system problem not a faster endpoint problem give me a concrete example of a decision that only takes makes sense if you believe that something you build as a faster endpoint company never would >> yeah I think the most important thing that focusing on asynchronous or long horizon agents specifically does for us is it lets us be very creative with chip choice. There are a lot of chips that have a lot of efficient compute available that just cannot achieve the interactivity levels that Nvidia silicon can today. so I'm thinking about for example some variants of TPUs. Some TPUs are small on a per chip basis but can be chained together in a long pipeline for example that make it possible to fit very large models on a very efficient chip but will not achieve the tokens per second that people are used to. It will never be 100 tokens per second, but it can be 20. And that to us is very interesting because it allows us to push models on the efficiency frontier and deliver again really great cost per token.

17:20 >> Cool. Well, on Kleiner's memo that argues about inference like a hundred billion dollar market is fragment into workload specific platforms. Do you think we will end up with like a five specialized inference clouds, 10? Or maybe do you think they're going to all reconsolidate? >> inference is too big of a market. >> Yeah. >> For there to be just one. I I absolutely believe that it's going to look a little bit like the cloud market. Yeah.

17:45 >> But even more specialization. I think that inference is the new unit of compute and compute is going to come in many different shapes and sizes. We've already talked about a few cases. Of course, everything interactive is going to have specialization. We're going to have voice specialized. We're going to have Still chatbots are going to be a big fraction of the market too. maybe the reason I'm bullish on long horizon agents is simply that they grow unboundedly.

18:08 >> Sure. >> It's not that I think chatbots are going away. Absolutely not. Just that there is a new sector that is emergent that is growing very quickly and we think it's going to require or demand new infrastructure. so yes, we're going to have more inference companies. >> Okay. >> Now I think the question there is what each of them specialize in. We're going to we have a specific thesis around the way we build the lowest levels of the stack.

18:31 >> Yeah. >> Differently to service the high throughput long horizon scenario. >> we're also going to see sovereign AI clouds as well. That is going to be a big factor in the future of >> right now most of our customers are in the US and Europe and so we special we focus in those countries. That's where we've seen a lot of deployment. But >> every country is going to want AI and every country is going to want that AI to be in a jurisdiction that's favorable to it. And so we're going to see a geopolitical angle to inference clouds as well.

18:58 >> Yeah. So we expect multiple platforms. So sure >> a little bit. >> Okay. that's getting a little technical. So take us through the stack. You customize open source engines like VRM with page attention >> and what have you changed the offtheshelf world that hasn't caught up to. >> Yeah. Yeah. I I hinted at this a little bit with the TPU comment, but basically we think that the parallelism schemes that you want for different chips are going to be the area where you change things the most. In particular, the world so far engines like BLM and SG lang have focused very much on things like tensor parallelism >> to trade off a little bit of efficiency in exchange for absolute speed. Yeah, >> this is the best way to deliver the lowest lowest latency per token and they're exceptional in that regime. But other forms of parallelism such as really wide expert parallelism or really deep pipeline parallelism. These are less first class citizens within VLM and SG lang. It's not the focus of those engines. and in particular this compounds when you think about supporting non-traditional chips.

20:01 >> If you look at chips like yeah again TPU or even more exotic chips like Intel has this line of chips called gaudy that we're quite excited about. you see those chips being fairly poorly supported in the frontline inference engines and that's where we see the most benefit to customizing and building our own building our own inference engine and innovating on the parallelism schemes. >> Got it. I see. So follow up question completion windows is the feature that I find the most original. the developers tells you how long they are willing to wait and your cut token cost 5 to 80% depends on the urgency. Mhm.

20:37 >> How did you scad how does the scadger actually exploit that slack? Yeah. Stack. >> That's the most interesting part of the company. I mean the whole job of sale is to open up that efficiency and speed frontier. >> You can always have more efficiency at a lower speed, >> but not everything can tolerate the lowest speed. There's going to be a spectrum. I would say we're not really going to be able to compete in low latency inference fundamentally. We're not we're not focused on that at all right now. Yeah. but for medium latency and beyond let's call that any response that takes seconds to hours but not sub-second >> that is an area where we have something to say we are a capable of serving efficiently at all points along that frontier and we expose a few of those points today in our API but that's the main novelty is I think we we are we treat the problem as a continuous curve and we're still serving at discrete points on that curve but we want to open up more points over time >> got it is through batching off peak hardware model routing or any of those or all of those.

21:39 >> So, so batching is the most obvious and important thing you absolutely need to use. I mean the economies of scale and in LLM inference auto reggressive inference in particular are absolutely in favor of you must batch users together and the more you batch especially with these mixture of experts models you're going to slow down the tokens on a per user basis and that is the correct choice to make for a throughput optimized service and maybe less correct to make if you're thinking about more of the medium latency tier.

22:05 That said, you can also flex on the chip choices I as I mentioned. and there are some chips that again are not going to be able to compete in the medium or low latency tier that are actually very good in the more long horizon more latency flexible tier. >> There's another point which I want to mention which is you mentioned off- peak hours. So off- peak I think is kind of the traditional interpretation of batch inference. Wait until chips are cheap overnight and then run them at really low cost. I don't know if that's actually super obviously available right now. It's not a lot of AI demand does not follow obvious dial cycles or even if it did well there's a lot of users across the world now and it's not so easy to find chips that are idle overnight because there's always demand coming from somewhere else when you're sleeping somebody else is awake and working. So AI is not just a US phenomenon anymore. Okay. So what does that mean? Well, we don't actually think that telling people to just wait 8 hours for their inference to start >> Sure. that is not super exciting. It's it's batch inference. And I don't use the term batch inference for that reason. We call it asynchronous inference because the goal is not to make users wait hours. The goal is to make them wait. Maybe there's some like flexibility for our longest latency tiers where the scheduler has the leeway to slot a given request in not over the course of hours, but maybe there's like a fiveminute window in which you can start the work. And that gives me enough time to go find GPUs somewhere in the world, spin up the inference engine that we run and just schedule on demand instead of having a GPU warm and waiting for the request. and so that simple difference, just being able to go allocate capacity or schedule in on demand and not being held to a super low or tight P99 latency bound, that buys me a lot of efficiency.

23:49 >> Got it. It's a relaxation of of worst case lat latency but maintaining average throughput. >> What about routing model routing? That's >> yeah model routing is very interesting. I think that you certainly see a lot of users kind of feeling fatigued at the pace at which open models are coming out. It's really hard to keep track of who's best today. GLM, Kimmy, Deepseek, Gemini, like who knows who knows what's popular today. And so I think that that is currently our users problem. I don't really help them with it today.

24:21 >> but there are great companies like not diamond an inference routing company that do a great job of selecting the right model for the task. I think we're going to see a lot more of that in the near future. And you have companies like RAMP launching their router product who are going to be very very good and well positioned to help larger companies not worry about the details of which model they're using but just RAMP will automatically be able to find the right model for the task.

24:45 >> Yeah. like open router they just acquired by strive right yeah right >> the second one is about sale sale boxes yes >> so for Linux virtual machines where an agent can leave for weeks v position state and you pause them during v states what breaks first when a agent runs for three weeks like contacts file systems or developers patients >> yeah yeah so we absolutely want to create a world in which it's practical to run agents for 3 weeks. But if you look at the existing sandboxing solutions, none of them are welld designigned for that. Most of them have runtime limits. They don't want you to run for more than 24 hours >> and they optimize very much for fast startup time. They don't optimize for bursts of work sustained over the course of a a long period of time. and so in particular, what we really wanted to fix was the kind of the cost of running an agent should not be linear with its just time it's act the box is active. We want the agent to only have to pay for work that it's actually actively doing, meaning the time it's scheduled on the CPU and doing CPUbound work. Yeah.

25:51 Because that's the true cost to the system. And so we we built sandboxes as a way to kind of have such an intelligent scheduler that we can pack multiple sale boxes onto a machine overcommit both the CPU and the memory and use a really intelligent to track which machines are using what resource at any given time and then move them around such that >> any given sandbox has the illusion of being essentially one or a few of a few tenants on the machine when in reality there are so many other sandboxes on the same machine that are all decorrelated in their actual usage patterns such that we're every given sandbox is only paying for exactly the time it uses. It only pays for exactly the CPU and memory critically that it uses. And that encourages you to just throw up a salebox, let it run for as long as it wants, and it's most of the time not going to be doing anything. It's going to be waiting for the for the next request.

26:40 >> we actually came up with this to cover some of the weaknesses of our inference product. Meaning that some of our early customers, they picked up sale inference and they started building these longer horizon agents. But the inference, you know, by definition, they were moving to more like flexible inference tiers. That means that the inference trajectory is going to run a little bit longer. And they noticed that their sandboxing costs were going up because they're spending more time in inference.

27:03 >> So we built sandboxes as a way to make waiting on inference actually free. >> When the agents waiting on inference, it's basically almost by definition not doing anything. It's they can't make any tool calls. They're all it's just waiting on the inference step. And so we want that to be free. >> Cool. you posted that 90.72% score on a browser comp plus for the long horizon research tasks and what actually moved that number the model the harness or the infrastructure >> great question so it's absolutely a tight interplay of all three in particular I think the harness being aware of the fact that inference is now so much more scalable you can get 10 times as many tokens for the same dollar that encourages you to build the harness in a way that does much more parallel fan out for a task like search or deep research. in fact, we're going to keep seeing this. Search as a problem is very paralyzable and scales extremely well with more compute. More intelligence and more intelligence per dollar are the determining factors in getting great search. And so some of our customers like Parallel really see this.

28:05 They have actually been great. They've used the term background agent for a long time. We we we very much see the future the same way they do on that. and they use sale inference to build really good indexes of the web. They're always scanning the web in the background building the most rich and robust index of all the pages on the internet which is a crazy task and that's running LLMs at web scale is not a thing that we considered possible even two years ago. Yeah.

28:29 >> But by combining efficiency of our stack with their great harness which is able to detect when changes are made to web pages and only filter down the the changes the web pages that are worth rescanning with an LLM. they're able to actually build one of the best indexes of the web. >> Amazing. give me ideas. So what do we need to do at AJ house? you were saving GRM 5 5.2 deepseek v4 ki gboss Qin. What are your audits reads on all the open race models versus the frontier APIs for agent work in 2026 and where's the real gaff now? Yeah. Well, it's been a very good few months for open models. we we've seen great releases from Kimmy first, then Quen, and Deepseek as well had some major updates in the last few weeks. So, it's a very good time to be an open source. I I think we it's always hard to say what the future holds. Why are these companies so good at matching the frontier? How are they doing that with what appear to be smaller models than what the frontier closed models are serving? I think that there is a challenge though here where educating users on the pace of model changes is getting harder. The models are changing every two weeks. It is as I mentioned a little fatiguing for users to be told like >> you need to move your GLM to Kimmy now.

29:51 That's that's just a little too much turn for a company focused on building great product and scaling revenue. >> and so I I I do believe that routers are going to have a larger role in the future. I think routing is going to become increasingly important. It could be the abstract layer, right? So users, any user don't need to worry about all the updates. >> I think the challenge before for routers was that you know the agents themselves did were not super high success rate. So why would you add another variable which is now you have a somewhat stochastic choice on which model is going to be chosen for the task. and that means that you're even less likely to hit the right result on on the first try. U I think that's improving substantially.

30:28 Agents are getting better across the board. And so overall, I think that routing is going to be a solution here. >> But yes, there there are a lot of models. There may be too many new models. >> Too many new models. Okay. >> But it also creates great competition. It means that there's a there's a fierce competition to >> make prices competitive and just make >> models perform as well as they possibly can on a per parameter basis.

30:48 >> Yeah. >> Great. Yeah. >> you support a lot of fine tunes and RL rollouts with Tinker integration. Is RL on your own agent about to become the table stakes for serious teams? >> I wouldn't call it table stakes. I think that I think that fine tuning is actually very difficult and most companies >> are probably not in a great position to do this on their own, but they can work with great companies like trajectory or applied compute to help them fine-tune their models and that I think we're going to see a lot more of. For example, you see companies like Harvey going out and training their own model with the help of applied compute and they are getting great results by specializing on the task that they actually care about most. They know their task very well.

31:30 >> but that's actually the constraint for most companies is that >> having them basically define a very crisp eval around their actual task is not easy for most enterprises. It's good for the best AI native companies like Harvey, but >> yeah, >> I think we're going to see a few more months, maybe a year before most other enterprises, most serious teams are able to actually do that really well. >> So that process might be able to democratize it over time.

31:53 >> Yes. Our role in that of course is to make sure the infrastructure is really efficient for customization and you know our position on that is that Laura adapters are actually a very good idea. They're very efficient to serve. we can service many different companies on the same physical hardware if they're all using Laura adapters. and thinking machines we think has done a great job building an API around this. they focus on the tricky hairy parts of training, >> which we don't want to get into for now. they're very good at that piece, but a lot of the cost is coming from the rollout step. That's a new part of how training works in this day and age is, you know, we're almost always talking about reinforcement learning, RL fine-tuning. And in Ral finetuning, it's fair to say that about 80% of the cost in customizing a model is in the rollout step.

32:40 >> Yeah. >> And we focused because we're an inference company, we focus on being the most efficient home for the rollouts and we leave the tricky complex parts to experts like the game machines on the training side. >> Agree. Yeah. actually we we do think that's like the well the big trends on the enterprise side. I think recently there's discussion from about those enterprise they might starting to train models own intelligence right instead of yeah the all the the frontier lab they own everything. Yes world owning your intelligence is I think >> this is important >> yeah it's it's existential I mean you cannot outsource the most important part of your business and I think that we're seeing an increasing reason on that front for businesses to customize models to sovereignty at the moment >> totally when you run open source you have many vendors to choose from my job is to be the best one of course but you'll always have many choices and with choice comes freedom >> amazing the claim is 3 to 10x cost improvements rough roughly one10 the inference cost of computing services on some workloads. Where does this number come from? Well, and when should a skeptical CTO CTO not believe it?

33:52 >> Yeah. Yeah. I I like to break down our advances on efficiency into two axes that are complimentary and actually orthogonal so you can stack them together. One axis that we talked about is this idea of completion windows. When you can trade off latency optimization for efficiency, you make that trade and you get to run basically more efficient inference. You're getting more flops out of the same GPUs. You're getting more utilization out of the same expensive chips. And that is a great path to getting anywhere between two to 3x improvement in cost per token on its own.

34:23 >> The second factor is to be more clever about what chips we choose to run with. like I mentioned, we are we have this motto at sale which is that there are no bad chips, only bad prices. >> Sure. >> So you can always find another chip that can do the task you want, but you you might have to get creative with how you serve the model on that chip. I've been talking in particular about fitting large models on chips with small memory which increasingly dominates the cost of these chips. And so if you can get clever with that and you can work around differences in network topology as well, you can actually get to a really efficient serving regime on chips that are not traditionally used in they may not be the latest Nvidia chip. You might use older Nvidia chips. You might use a combination of Nvidia chips plus some other vendors chips and you can specialize these inference systems into the comparative advantage for every chip. Some steps in inference need lots of memory bandwidth. Some need a lot of compute throughput and you want to specialize those two things. And so that a lot of complexity there of course but it's an always moving industry. but we we do a lot of work there and combining that with the flexibility of our customers demands. We opened up that efficiency and time curve that allows us to deliver 10x speed up or 10x improvement in in efficiency.

35:38 >> Amazing. Yes. we didn't get to talk a lot about those the chips because there's quite a few chips we're talking about nowadays. You have edge the Cros Gro and all the TPUs. Have you guys like played around most of it or >> Absolutely. I mean we are I'm a chip person at heart. I love chips. I love talking about chips. I love learning about chips. I think there's there's a ton to be said about specialization of these chips and how some of these newer players like Etched are making really specific bets that are very ambitious in a way that I didn't think there was appetite for 5 years ago.

36:11 Yeah, >> building a chip is a hard and difficult like hard and long process. It's there's no guarantee you're going to make it work the first try. It's it's really incredible that Ash is able to make their first chip so successful. that's not the standard in the industry. industry is usually you have two or three chips that are not good and then you can start to execute. >> and so >> I think that's going to be the other area that we're going to see a lot of growth in over the next three to five years.

36:38 >> It may not be this year or next year even. chips take a long time to come to market come to scale but yes the chip economy is growing really well right now. >> Got it. Okay cool. we'll move a little bit to the customer dem the the market side of things. so there's a detailed dev runs code review agent for three to four hours straight. >> Mhm. >> And parallel jack and jail quadrillion I think you mentioned parallel web systems are on the site. What are most surprising things a customer have done with sandbox?

37:13 >> Yeah with sandboxes we have actually a broader mix of customers on saleboxes. They they do very interesting things. I I think by giving them the primitive that the sandbox doesn't have to be ephemeral. It doesn't have to be turned off. You can simply let it run as as long as the agent wants to because as soon as you're not doing any work, you're instantly not being built for it. That is a critical new feature in sandboxes that is not traditional in the sandboxing or agent VM market. and so that has really encouraged some people to build these like always on agents.

37:44 That's maybe a new category that didn't exist before. Previously, agents would always work in response to some explicit user trigger. Yeah. >> And now the agent can >> live in the background and just wait for users to send traffic to do something. we use it a lot internally. We use for example sandboxes to keep track of all the various GPU providers that we have. We work with hundreds of different data centers that are basically al have dynamic availability of compute and we're very good at basically having boxes monitor the actual supply at every given site and route us or give us feedback for ouruler as to where to go allocate compute next where to go pick up a GPU. That's a very internal use case. We also use it for more fun things like scheduling our lunch every day. We had to make sure with things like the Door Dash CLI that everyone has actually ordered lunch and make sure that the order goes out by 11:30 so that we have a chance of eating around the noon time.

38:39 >> and so fun things like that. I think I think we're going to see a lot more very personalized software running in boxes. It is the easiest way by far for an agent to build a small app for a user base that could just be one person >> and have it run with no further thoughts about cloud infrastructure or deployment. >> The simplest thing is, for example, deploying software with no database. just run a sandbox and have the database be a file in the sandbox. That's the easiest thing you could possibly deploy.

39:04 Agents find it very easy to deploy software in sandboxes. >> Amazing. I wish we talked about this earlier so that we can have the salebox for the developers. We have 100 people here, builders here today. Yeah, maybe next time we'll set it up. Yeah. >> So, you were serving trillions of tokens a week as we just talked about launched only in March 2026. It's like a few months ago with no rate limits and no talk to sales.

39:28 What did removing a sales gate do to your growth curve? >> Yeah. Well, you know, we're I'm an engineer. almost everyone on the team is an engineer. we we're very much like we want people to be able to feel the product and use it immediately. that's really important just I think just like spiritually to me. U besides we're selling a product that is quite well understood by people at this point. We're selling tokens and I think the proof is always in the pudding. There's a lot of people who are talking about inference at scale, but you know, at the end of the day, you have to serve the tokens to be a real company. And I knew that having allowing basically anyone to sign up and send us traffic is the ultimate test of reliability, of scale, of speed. And honestly, we have a lot more work to do there, too. Like, we're continuing to improve every single day on reliability. But the way to do it is not by, I would say, having a long sales process than getting into enterprise agreements. that will come.

40:19 But for me personally, I think, you know, we're we're proud of our work and we want to stand behind it and we want our customers to hold this to a high standard. So, we say welcome all people. Please sign up, use sale, and give us your best. >> Cool. what's the pro profile of the team that gets 10x value from sale versus the team that shouldn't bother solution yet? >> I think that we're seeing basically the market expand to almost all AI builders at this point. I mean again tokens are a very hot commodity at this point.

40:48 I think the question for us has been how do we meet that demand at all layers that can exist. Like I mentioned there's a line where we we currently don't cross which is low latency inference. We cannot deliver 100 tickets per second to to most people. But for everything else we're finding that increasingly you don't need the lowest latency possible for an interesting agent that is running for many turns in the background. and so we want to be best-in-class for those users. they will see immediate ROI. They will see that their agent can do a lot more when you don't have to count the tokens out carefully. When if I can give someone 10x more tokens for a problem, that's a new class of product. They should hold it a different way and they should they should push it much further than they have before.

41:30 that's the goal. That's people who are going to see the most impact from sale. >> Got it. Amazing. let's talk a little bit about competition. So together AI >> your former employer and a fellow clienter portfolio actually also is the obvious comparison and you've said that biggest threat is the frontier labs building proprietary infrastructure. How do you outrun open AI and Google on their own inference? >> Yeah it's a great question. So you know first on on together and and you know base 10 and fireworks I think these are all great companies. They've done an amazing job of serving low latency inference demand. They will continue to be best-in-class I think at the at the low latency piece. They are doing great work there and that's not easy to build.

42:14 I know because I I built it. >> Sure. >> and I think we will continue to see variance of inference like we mentioned before there may be diversity of inference use cases and there will be specialization in all these. >> in particular though I think if you want to make a big swing for efficiency the way we are it makes you take some different choices than than what I think together base tenant artworks have done. we have to be a little bit more creative with the way we use chips. We have to be more creative and opinionated about where and how we host those chips as well. And that will I think turn out to be one of the areas in which we start to diverge a little bit more over time.

42:49 today it's where we focus on the you know having a really great async product and we focus on having great economics at many points along that speed and and cost curve. the larger question about the role of open source or all of us inference players against open athropic now I don't want to make it sound too adversarial obviously they created this field they have they continue to lead and create new ideas I mean they're the ones who invented reasoning they have immense impact on the field I don't think they see open source models as competitive with where they want to go they are focused on the >> longest risen task of all building AGI >> for this very house and they >> they're focused on that goal Got it.

43:31 >> And they continue to focus on things like pre-training models. Pre-training models drives a lot of their comput strategy. They build these huge massive data centers focused on training because training requires concentrated compute. Yeah. >> And that is top of- mind for their comput supply teams. My compute supply team thinks quite differently. We we are very curious about new ways to deploy inference compute. We think inference compute can be smaller, more modular and use a very different set of chips than what you need for training. And so that's why we're very excited about building a different set of decisions from the data centers and power level all the way up.

44:07 >> Got it. Okay. So I guess you also have another thing that we didn't mention much yet is the compliance posture about sock 2 he support and zero data retention by default those things like an enterprise play. So is that end game the agent the cloud for the fortune 500? Absolutely. I think that we're I mean every company is going to want to deploy artificial intelligence internally and in their products. like we're very early in the deployment to the Fortune 500.

44:38 there's a lot of work to do there and a lot of that comes from trust and and building a really secure foundation from the earliest days. So even if right now most of our mix is not enterprise demand we know that we had to lay the foundation today so that we're ready for that future. >> Got it. Okay. Amazing. Looking forward to it. so you emerged from STS in June with 80 million across seed and series A at a 450 million valuation.

45:05 Kleiner SEOA leading with red point and a few other as co-investors. When Intel CEO and the chairman of Alphabet both write checks, what are they each betting on? >> Yeah. Well, I can talk about maybe John Hennessy first. I knew him for longer. So, John Hennessy taught me computer architecture at Stanford. Great, one of the greatest privileges of going to that school is getting to work with great just legends like John Hennessy. I I I've loved his work, his books for so long and and just meeting him in person as a student was was a dream come true, but imagine my excitement when I got him to write Angel. I mean, it's just it's a he's a hero of mine. He's done so much for the industry and yeah, he's a real titan. I mean, and he he's so he's still still so sharp on this. he he's very familiar with all the AI inference companies and all the chip companies as well and his opinions on where we're going. I think we saw a lot of alignment there and we're really happy to have him on the team.

46:02 >> Lip Bhutan. >> Yeah, >> similar story. I mean Intel is plays a has a had a major impact on my life. My mom worked at Intel for 25 years. Wow. >> I grew up in Santa Clara just 5 minutes away from the Intel headquarters. I spent a lot of time in those Intel cubicles while my mom worked. and you know Intel is a great American silicon company. I love that history of Intel and I am so excited for Intel's new direction to come back. I think Glip Bhutan is the greatest CEO Intel has had in a long time and he's going to he's the right person to bring Intel back to its its really really powerful state that it that it once had.

46:38 >> Amazing working with your heroes. That's what a great feeling about that, right? Yeah. >> so the company name sale research the name promise research and your mission is the pursuit of a token abundance. what does the research agenda look like beyond serving tokens to >> that's a great question. Yeah. What what do we research at sale research? we research how to make GPUs more efficient. >> Sure. >> That is our northstar at all times. It is chase the speed of light for efficiency.

47:06 >> And I think that has taken us in a lot of interesting directions. Number one, as I keep talking about these these other chips, I really think that, you know, pushing the frontier of other chips has has has really been so productive for us and it it's the most exciting thing for me right now is we have all these ideas about heterogeneous serving, mixing different chips that are great at different things. For example, Nvidia has by far the best rack scale system today, the GB300.

47:31 >> but it's not the entirety of inference. there there are phases of inference that you'd want to offload to other chips and making a GB300 rack be the center with some satellite chips that work well with it that is going to be a very interesting way to build inference specialized compute over the next couple years I think other research that we do we think a lot about how to efficiently serve or what what the changes in model architecture kind of mean for how we should think about chips differently for example we've seen a large move in the industry towards sparse attention alone and >> I think actually profoundly changes the way we should think about the role of chips or like the role of memory in chips specifically. We had this bottleneck before where attention was linear in memory capacity and memory bandwidth. Meaning that as the se sequence length gets longer, you have to every new token you're decoding has to read linearly more tokens from fast memory. And that was a huge challenge because HBM is very expensive these days and it's not getting any cheaper. So sparse attention what it does is it decouples capacity from bandwidth. For the first time you still have you need high capacity memory. You need the ability to store long sequences but every new token that you decode is only going to read a subset a small constant size subset of that full sequence KV cache. And that means that maybe we can use alternative memory technologies like flash. Yeah, >> in order to read at lower bandwidth but the same large capacity and that's how we're going to scale sequence length.

49:09 >> Yeah, especially you're solving a problem with compute and memory they can balance right. I think there's a lot of research on algorithmic kind of directions as models as well and also the other thing I don't I would love your opinion is >> we maybe it's a buzz word kind of thing people are talking about self-improving kind of auto improvement of kernel optimization do you do you guys look into that >> absolutely yeah I think the job of a kernel engineer which is what I trained in for many years that has changed substantially in the last few few months alone >> I think the main difference is that writing a kernel used to mean opening your editor and starting with code.

49:50 >> At this point, I rarely start with code. We may look at code, but ultimately I think the kernel design is done on the whiteboard these days. >> I get my top kernel engineers in a room and we we sit down and we think about okay, what do we want the machine to be doing? >> Yeah. >> On every cycle, >> what is the path to get there? what is the pipeline we need to build from data movement to compute to communication to finally writing the data out to memory again that pipeline that four- stage pipeline essentially is the fundamental >> loop in all kernels and it's our job to map that onto the pieces that Nvidia AMD Intel give us that we had to map the kernel's logical pipeline onto the physical hardware conceptually based off what we know about the chip and then the agent tends to write most of the actual low-level to actually make that happen.

50:39 >> Yeah. >> So that that has been a profound change. I would say that we're still for kernel engineering although it seems like it's the number one target for auto research. It's a verifiable problem with the smoother reward function. What's not to like? >> For some reason we have not seen dramatic impact from AI generated kernels. I think >> we're still talking about >> centaur systems where the human plus the human expert plus the machine is doing a really good job.

51:02 >> But I'm quite excited for the future where it is fully autonomous. It makes perfect sense to me that it should be automated. for whatever reason, it's not as well represented in the pre-training distribution, let's say, of these models, but the post- training is very good at given a small scope task, absolutely delivering the best possible performance. We can write way more kernels per day than we used to. >> Yeah, I'm looking forward to that. Yeah, that's one part of the things we've been trying to put a lot of attention on nowadays. Yeah.

51:29 >> who are you hiring right now? Like, and what does a sales shaped engineer look like? Yeah. we're always hiring very curious engineers who want to push the frontier of performance. I think what makes a good performance engineer is ultimately like I mentioned before what I feel which is >> almost an emotional desire to make the chip work at its peak efficiency. you have to want that at the end of the day making the number go up just in this abstract sense is probably not going to be enough motivation to do really really excellent work. We're going to dig through a lot of >> frankly sometimes painful interaction with these chips. like things are going to break in weird ways. The only thing that'll keep you going is is a desire to see the machine really sing at its peak efficiency. and that to me is just something that you you meet someone and you know that they they feel that way and they're a member of the team already essentially. great engineers.

52:20 >> Yeah, >> we don't actually hire that many researchers right now. I think we're we're mostly focused on people who really focus on who want to >> make the system better, a little bit better every day and go after some very hard problems that conceptually feel exotic and new. That's maybe the most researchy thing that we do. But at the end of the day, it's really important that we deliver on the ideas that we have. Got it. Okay, cool. that's we're almost to our end. I would like to get a little bit about the future.

52:51 Paint 2029 for me. Yeah. >> If abundant intelligence happens, if token cost for for another 100x and agents run for monsters. >> Mhm. >> What does a software company even look like? >> Yeah. Yeah. So, one area that I think we should start with is we're already kind of seeing the the the cusp of here is, you know, what does cyber security look like in that world? we're already getting to the point where some security researchers are calling modern security basically a proofof work problem in the sense that your software is as secure as you're willing to spend tokens trying to break it. you should red team everyone should be red teaming their software at all times. Okay. And you have great companies like Armadin or Expo leading the charge here on pentesting applications before actual attackers do.

53:38 And so I think that is the first area of software development that's going to change dramatically in the next six months. it's a little scary honestly the the pace of which people have to adapt here. But abundant intelligence will change the way we think about security substantially. It's going to be this is going to be an always on cat-and- mouse game and >> we're going to be deploying so much intelligence simply to secure and defend all software. the larger piece about how software development will change.

54:07 Well, I think we should look at the pieces of the stack that have not changed enough yet. there I feel that computer use is the one of the frontiers in the software ecosystem that has been a little slower than software engineering or like basically code generation. Computer use tends to be one of the parts of the stack that is harder to accelerate. for example, if you want to build great consumer software, a big part of that is testing the software and enumerating all the features that you've added in code and make sure that they work well. Well, it's not super easy to do that today. There are some companies working on autonomous testing of iPhone apps, but >> yeah, >> it still requires a level of computer use and vision language understanding that I would say we're still developing and working on. and so we'll have to see some improvement there before you can really essentially have all of the software that touches people's lives be be more automated. But the progress is very promising.

55:03 >> Amazing. What breaks first at at that scale like power, memory benefits or human oversight? >> good questions. I mean 2029 scale I think we were still talking about power at that point. Power is still one of the limiting factors to AI deployment in the United States. that's why I'm, you know, I'm I'm mentioning that we we want to build power in a more distributed fashion. I think smaller data centers that that are each maybe like sub megawatt in scale. that's a much more attractive and tractable problem that we can actually make progress on in the next year and two years and three years.

55:39 The US does have a lot of power. It's just distributed. and if you want to find a 10 megawatt or 100 megawatt site, that will continue to be very difficult. I also think that, you know, kind of public perception of data centers is very poor right now. No one likes data centers. And I think the trick there is, yes, it's going to be always hard to build a 100 megawatt site in a way that's seamlessly blends in with the community. But small sub megawatt sites, those seem much more reasonable and possible to build throughout the United States in ways that create lots of distributed economic productivity, can use things like solar power and wind power much more effectively than the large sites which need effectively gas power at this point. and so we'll see a lot of diversity in how these small data centers sprout up and we hope to lead the charge on that.

56:26 >> Interesting. Okay. you build a pathfinder because technology wasn't reaching people who need it. Does abundant intelligence reach everyone or does it pull at the top? what's the assistive tech version of this agent economy? >> Yeah. Yeah. Well, okay. I think there's two interpretations of that. One is kind of how do you bring the benefits of a of intelligence to to basically the entire world? And if you look at, for example, users of open source models, a lot of them tend to be from countries where, you know, a $20 chat GBT or a cloud subscription is a substantial amount of money. And they, you know, it's just not sustainable for individuals to access that. And so they turn to things like open code, which is really great at giving people access to abundant open models or they'll go use the like Kimmy models directly, for example, and get frontier intelligence at prices that are much more sustainable. we see our role in that in that world very importantly as well. I think it's absolutely the case that we don't want intelligence to be something that only customers in the United States can afford to access. We think it's very important that intelligence be open as well as priced in a way that is delivering the maximum value to not just companies but also individual users.

57:43 >> Cool. Well, before we close up, we have a few rapid fire questions for you. >> most underrated open weight models right now. >> Underrated open weight. yeah, good question. Probably Neatron. >> Neatron. All right. Okay. One cool kernel trick you wish I would know. >> I hope that people have to learn lesson. Frankly, Mention should take over. But >> yeah, if you know what swizzling is, you need to know a lot about swizzling for >> yeah, for working some with some of these chips. All right, cool. Apple, Nvidia together is one word to describe this.

58:24 >> Well, there's they operate at very different sizes of scale, but yeah, >> you know, these are all companies who care about essentially chips. Chips are the lifeblood for all three of these companies. >> Cool. the agent work reload. Nobody is building yet. That's someone listening should. agent workload. yeah, I think you know, we we kind of talked about this a little bit, but the web the web is now accessible at a scale that never has been before. Use parallel web search and >> like build an agent that really tries to understand the entire web or a large corner of the web. that's possible at this point with the cost of intelligence where it is, the scale of intelligence where it is, and great companies like Parallel who are pushing the frontier of making that index searchable.

59:14 >> Amazing. Okay, cool. now where where should the people go to if they want to try your like sale API credits and career or >> Yeah. Yeah. >> Yeah. salresearch.com. We are always hiring great people. We are always happy to support hobbyists, companies, enterprises, anyone on that scale. We're open for business for anyone. >> Cool. Well, Neil co-founder and CEO of SE Research building the inference cloud for the agent area. Thanks for coming for AJ Hus. Thanks for the time, Rocky. Cool.

Summary

Neil Moa, co-founder and CEO of Sale Research, discusses the evolving landscape of AI infrastructure, emphasizing the need for efficient serving of AI models on diverse chips. He highlights the paradox of falling token prices alongside rising enterprise AI costs, attributing it to a shift from low-latency chatbot applications to long-running autonomous agents that require different infrastructure.

- Sale Research's motto: "No bad chips, only bad prices," encourages creative solutions for model serving on various hardware.
- The AI industry is transitioning from low-latency chatbot models to long-running agents, necessitating a rethinking of the inference stack.
- Sale Research aims to optimize chip performance and efficiency, focusing on throughput rather than just latency.
- The company has raised $80 million to rebuild the inference stack, targeting the enterprise market.
- Moa emphasizes the importance of collaboration and trust between co-founders, noting his long-standing partnership with CTO Samir.
- The future of AI infrastructure will involve specialized inference clouds tailored to specific workloads and geopolitical considerations.
- Sale Research is committed to making AI accessible and affordable, particularly through open-source models and innovative scheduling techniques.
- The company is focused on research to improve GPU efficiency and explore new chip architectures to enhance AI model performance.

Questions Answered

What does it mean to have idle GPUs in a data center?

Idle GPUs represent wasted potential and resources in a data center, as they are designed for high performance but not being utilized effectively. The speaker emphasizes the need to maximize GPU usage to recognize the engineering effort behind their design.

Why is it important to rethink AI infrastructure?

As AI workloads evolve from quick responses to long-running agents, the infrastructure must adapt accordingly. This includes rethinking chip selection and data center design to support sustained operations.

How can users effectively choose AI models?

Users face challenges in keeping up with the rapid release of new AI models. Companies like RAMP are emerging to assist by automatically selecting the most suitable model for specific tasks, alleviating the burden on users.

What are the challenges and opportunities in chip development?

Chip development is a complex and lengthy process, often requiring multiple iterations before success. However, there is growing interest and investment in specialized chips, indicating a potential for innovation in the industry.

How does sparse attention impact chip design and memory usage?

Sparse attention changes the dynamics of memory usage in AI models by decoupling capacity from bandwidth. This allows for the use of alternative memory technologies, potentially reducing costs and improving efficiency in processing long sequences.

© transcribe · For agents Built with care and craft by Gokul Rajaram