transcribe

The 0x-alpha Mystery, Qwen's 27B Beats Opus 4.8 & Legal Is AI's Breakout | This Week In AI

Mastra · 24m · transcribed 16d ago
More from Mastra Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Stealth Release and Agent Preferences

What are the community's thoughts on using agents versus tools directly?

The community is divided on whether to use their own agents to access tools or to rely on built-in agents within those tools. Some prefer the convenience of integrated agents, while others want to maintain control through their own agents.

  • The concept of stealth releases creates excitement but also confusion.
  • Users are debating the effectiveness of using personal agents versus built-in tool agents.
  • Preference may depend on the user's familiarity and expertise in the specific domain.
# 4:50

Speculation on Token Capacity and Models

What are the implications of the recent speculation about token capacities and models?

There is significant speculation regarding the capabilities of new models, particularly concerning their token processing capacities. This has led to various theories about the origins and performance of these models, with some believing they may be tied to major companies like Google.

  • The community is intrigued by the potential of new models and their token capacities.
  • Speculation can lead to conspiracy theories about the origins of these models.
  • Performance benchmarks are critical for evaluating new models, but results can be misleading.
# 9:41

Growth of Open Weight Models

What trends are emerging in the use of open weight models?

There has been a significant increase in the share of open weight models being used, indicating a shift in preference among users. This could be due to either a migration from frontier models or an increase in the overall token usage of open models.

  • The share of open weight models has dramatically increased from 28.4% to 62% in two months.
  • Users may prefer open models for cost-effectiveness and variety.
  • The growth of open models suggests a potential decline in the use of closed models.
# 14:32

Real-World Applications and Evaluation Practices

How are companies leveraging data for model evaluation?

Companies like Harvey are focusing on real-world applications and have developed robust evaluation practices to ensure their models perform effectively. This includes using extensive data and evaluation suites to refine their models.

  • Real-world applications require rigorous evaluation methods to ensure effectiveness.
  • Companies are increasingly collecting data to improve their models over time.
  • The ability to train smaller models using collected data can lead to cost savings and better performance.
# 19:23

Scaling Challenges in Git and Technology Evolution

What challenges does Git face in scaling, and what solutions are being proposed?

Git faces significant scaling challenges due to its architecture, which was designed over a decade ago. New solutions like Cursor's approach aim to address these challenges by enabling infinite scaling, highlighting the need for modern technology to handle current demands.

  • Scaling issues in Git are partly due to outdated architecture and unprecedented growth.
  • New solutions are being developed to address these scaling challenges effectively.
  • Understanding the limitations of legacy systems is crucial for future technological advancements.

Transcript

0:00 This like stealth release stuff bewilders the community. Takes them by storm. Will the Frontier ever give it out in stealth? >> I feel like they don't have to. Their marketing is it got banned by the government. That's their marketing stunt. >> >> Welcome to Agents Hour. Today we have a pretty short episode considering some of the past weeks we've had, but there are some big juicy topics that we do want to dive into. Let's start with this post.

0:39 This is from someone named Ali and she says, "I really don't want to use your agent. I want to use my agent to use your thing." And this got a ton of engagement. A lot of people, you know, were commenting, resonating with it. Some disagreed, a lot of people agreed, but the question is in this new age, do you just want to use your agent and connect to someone's Elsa's tool maybe through MCP or whatever, or do you want that tool to just build the agent themselves and you use that?

1:06 >> I agree with this statement for things where like the boundary of the product is also on my end, right? Like if it's a coding thing or infrastructure whatever, but here is where I disagree. I would rather have an agent in Google Slides than have my agent talk to Google Slides. You know what I mean? Where these products are encapsulated, like Descript. Maybe I want to be have my own agent, but no, I'd rather just be in Descript just making moves in this thing. But where I think this comes into play, if I were to make it more general, if I also have industry experience in the product, I want my agent in there.

1:43 If I have no business doing that in the first place, like making music or anything, I want to use their agent. That's where I kind of draw the line. >> Yeah, so you're saying because they would conceivably be better at building an agent in the domain that you're not experienced in. And you wouldn't be able to guide the agent to the result. It's probably better if they just build it. They control the context rather than you try to control the context.

2:04 >> Correct. >> I think both use cases are valid. And I will ultimately think it depends, very similar to you. For instance, you know, for an example is Posthog. Like I don't use Posthog every day. I use it quite a bit, but not every day. And I've noticed that the agent in Posthog is actually pretty good. It does what I need it to. And now I don't have to connect the Posthog MCP that probably has 50 tools to my agent. I I do worry that over time your agent might get worse, and you're not validating it or you don't know it cuz you're just slamming a bunch of MCPs in there, and which adds to the context. And every what the average person doesn't understand, if I have 100 MCPs connected to my agent, that's not very good, right? Like that's going to impact the results. It's going to slow down the agent. There's more tokens. All that stuff, you know, factors in.

2:49 And the average person doesn't really understand that. So I think there is this risk of, at least today, and maybe this changes over time, but if you just connect your agent to every tool possible, you now your agent has to try to do every single thing. And I think that you might get worse results, and you might not know why your agent seemingly got dumber over time. You might think it's the model. You might think it's, you know, the way you're using it. But maybe it's just your tool context is getting polluted because you're connecting so many damn MCP servers to it.

3:16 >> Yeah. >> So that's why I think the specialized agents are useful because you don't have to worry about that. You only go in there some of the time. I don't use Posthog every day. If I did, I would much rather connect it to my agent and just use the MCP. Another thing is like Notion. Notion's agent, I think it kind of sucks. And I use Notion almost every day. So yeah, I'll just use the MCP, and it works, right? I understand what I want it to do. I can guide it. But I I'm taking on that, you know, all those extra tools in my agent, right? If especially if I just like connect it to the my agent, and I don't have different, you know, settings or profiles or whatever, or projects that have like different configurations. So, I think in general, I agree with Ali for a lot of the tools, but not for everything. Like they're specialized tools, just let let them build the agent. They're the experts. Make the agent really damn good. You know, maybe someday my agent can talk to your agent and not have to care about the context, just like passing messages across. But for now, I'll I'll use the agent when I need to.

4:12 >> If your agent is built in Mistral, I'll use it. >> There you go. There you have it. All right, we got to talk about I don't know if it's 0xAlpha. I've been saying OXAlpha, but we got to talk about 0xAlpha because it's a bit of a mystery and it kind of just hit the timeline later last week around what is this model that just seemingly popped up within OpenRouter. So, this is from MTS-Live. It says situation explained, a stealth model on OpenRouter is beating Fable and Soul on coding and nobody knows who made it. OXAlpha has a 1 million token context window, text, image, and video input free for a week with capacity for 100 trillion tokens a day. And I think that's what blew people's minds cuz they didn't understand where could 100 trillion tokens a day come from. That compute doesn't just exist lying around. There's only a few companies that in theory could have that capacity.

5:01 >> I love the conspiracy theories that have come from this. >> Yeah, there have been a lot and we'll talk about some of them. Brandon Carl says, "Whoever is offering 0xAlpha had the ability to offer the entire token capacity of Google immediately, 100 trillion tokens a day." Which led to a lot of speculation that was this a Google model? The Gemini team was posting some kind of obscure things making people think maybe this is a Gemini model. Could this be the new Google? How good is this model? People were starting to run benchmarks on it.

5:26 This post says, "WTF, Ben ran this mystery model through 10 deep suite tasks and it scored over 80% versus 65% for Fable and 52% for GPT 5.6 Soul. This is insane. Probably a Chinese company, either a new GLM or Kimmy model, I reckon." I did see some posts that said they ran more of the tasks. It didn't score this well on everything. I think this was a bit cherry-picked. I think overall it seems like what I've seen it actually performs not as good as Fable in 5.6. So, maybe maybe it is Gemini. I don't know.

5:57 >> >> But, it's not. At least not according to this post which says breaking the stealth model 0x alpha is z.ai. So, it's a GLM. Everyone else is guessing it from tokenizer vibes and emoji rates. I made the server say its own name out loud. Said it one malformed request and it threw a Java stack trace naming its own internal. What do you think? Is it GLM? >> It'd be good for the narrative that it is, but I'm really hoping it is Google.

6:24 I think it might be zai. >> And this is all like some speculation, right? But, it seems that it's a smaller model, but it's doing it does really well on at least the benchmarks that people have run. Not Fable level though, but their theory is that this is almost like a GLM flash type model or a smaller model that maybe is on par or better than the last like GLM 5.2, but is a smaller model and so that is how z.ai can serve 100 trillion tokens a day is that it's just significantly more efficient because it's so much smaller.

6:58 >> There are other indicators that would lead for this to be Z. So, like open code offered this through open code go for free or whatever and they were tracking the token usage. People are using it. If it was Google, anybody would have to sign a bunch of NDAs. We have done this. I don't know if I'm not allowed to say that I've done it, but okay. But, you have to sign NDAs to use future models. So, I don't think it's Google for that reason alone.

7:24 They're not going to give this away in stealth that other people can use without tons of like bureaucracy. So, it's got to be an open model. That's what my >> >> my guess is. And you know, if this evidence is damning enough, then it probably is Z. >> And I think we will eventually find out. I think the hype has died down from the model a bit because I think people realize it wasn't as good as what they originally thought. So, I don't think it's Frontier. It's not a new Frontier model, but if it is, you know, a GLM model and it's much smaller, you know, if you think back to, you know, the Qwen model that came out, which we will talk about and how that's, you know, like a small, basically, on-device, you know, you can run on your own computer model and it benchmarks very well. Well, then maybe GLM has done the same type of thing, where it's a smaller model. We don't know the size yet, but maybe they can compete close to the Frontier without being Frontier-sized.

8:14 >> This like stealth release stuff like bewilders the community. It takes them by storm. So, I'm curious how many stealth models we can go through before it like loses its luster. Or is it always going to be a really hype way to debut a model, to give it out in stealth first? >> Yeah, I mean, I think there's just a marketing aspect to it, for sure. >> Will the Frontier ever give it out in stealth, I wonder?

8:36 >> I feel like they don't have to. Their marketing is it got banned by the government. That's their marketing stunt. >> >> Qwen breaks the size curve. We talked about this a little bit last week, but if you look at the artificial analysis agentic index, Qwen 3.8 27B is actually ahead of models like Opus 48, ahead of 56 Luna. So, it is, you know, ahead of GLM 52. It's ahead of these much larger models, right? And this is only a 27 billion parameter model.

9:09 >> Yeah. >> If you told someone 18 months ago, right, not even you told someone 6 months ago that they could have, you know, an Opus-level model on their device, like on their computer running, you know, obviously still needs to be decent-sized, right? You still got to have some hardware to run that, but I think people would freak out. Running an Opus-level model on consumer hardware is pretty insane. Makes me wonder like how many tasks do we even need to pay for tokens anymore? Will teams just have like their own reasonably good model on their own machine that they run most tasks do and then the the hard task gets sent to the frontiers. I don't know.

9:44 >> Dude, did you subscribe? >> Dude, I host the show. Did you subscribe? >> Did you subscribe? >> Subscribe to Agent Saur, every Monday noon Pacific. Let's talk about open weight models a little bit more. So, this is from Guillermo from Vercel. Says today is a record day for open weight share of tokens on Vercel AI Gateway and OpenRouter's posted similar things. So, this is not just Vercel, this is across a lot of the popular AI LLM routers. And so, August 22nd, 62% of tokens were open weight models and in June 24th, so about 2 months ago it was only 28.4%.

10:20 That's a pretty >> Huge. >> wild increase. I am curious on is it cuz there's two ways you can kind of read this. Maybe, you know, a lot of people moved from frontier models to open models or maybe there's just a lot more tokens being sent through and a higher proportion of those tokens are now becoming open models. So, >> Yeah. >> they don't really share like total token growth count or anything so you can't really determine are closed models in, you know, are they total are they did the total count go down? I doubt it, but it probably just didn't grow as fast as the open models. And I think if you're using frontier tokens, a lot of people, the majority people are going right through the providers, right? Not through a router.

10:58 >> I mean, I would suspect this too. Like, it's going to be more open cuz, you know, you are not paying max coding plans on these gateways. And there are a lot more open models to choose from for price. >> We got to talk about Slack code. This is from Benioff on August 19th. Says, "Don't code alone. Slack code is live. Humans and agents, same channel, same work. Launching today with agents from Anthropic, GitHub, Cognition, Vercel.

11:29 This is real multiplayer coding. See it at Dreamforce." What do you think of this? >> I mean, if their website was up that day, I would have been more happy about it, but it was down for like the whole day. >> Yeah. I mean, they said see it live and then people were saying it wasn't quite live, but if you watch the video, there's like basically a code panel that pops up in a conversation, right? You're basically you're prompting in Slack and you're just having it write code. You know, we have mentioned that coding, you know, from your mobile device is the dream and Slack has a pretty good mobile app, so now you can in theory code from anywhere within Slack. I'm not convinced that Slack's going to win this on its own. I feel like if you ever ever used Devin, Devin Slack agent is pretty good.

12:09 I know Linear's investing heavily in their agent as well within Slack. I think obviously Slack is trying to compete with all the apps that integrate in with it and try and just do it themselves. They want to Maybe Slack wants to own a piece of the coding pie in some ways, but we'll see how many people actually use the the Slack version or there's always those things where there is a version in the app itself, but people use third parties because third party ones just end up being better. And so, I think we'll probably see some of that as well.

12:35 >> Could you imagine a world where Slack bans third-party agent apps just so you can use Slack code? >> I feel like they can't, right? It's like Salesforce is kind of built on this ecosystem of like you can build sales I mean, they have a proprietary language, of course, but it's all built on like this marketplace and all these different apps. It's like so complicated they have all these different companies that just provide Salesforce type services. I feel like they wouldn't want to go back and change that specifically for Slack. I think they'd want Slack to be open, Slack to be the communication tool, but maybe not. Maybe they're looking at all the, you know, the revenue from all these coding agent companies and saying maybe we can grab a piece of that.

13:13 >> Is this GA yet? Can we use this? I don't see it in our Slack. >> Yeah, well, we didn't go to Dreamforce this year, so >> Oh, that's why. >> I mean, maybe we should reach out and see if we can get it turned on. I'm guessing there's a way to get it, but maybe you have to know someone right now. >> >> So, there's been a a lot of talk recently, and especially the last week, around just the legal as a breakout use case for AI. So, Harvey has introduced Tenent, which is the first model post-trained for legal. So, Harvey's actually not just using frontier models, they're actually training their own or post-training their own. So, Tenent is a Kimmy K3 base that we post-trained with Fireworks on a corpus of publicly available legal data, synthetic data, and human expert data simulating long-horizon legal work. This is pretty wild. That means there are actually companies that started as just you could call them like a model wrapper, right?

14:05 >> Yeah. >> They were just an application on top of a model, and now, I would say Cursor was the same way, right? Like, Cursor didn't start as building its own LLMs, it just was kind of like an agent that used whatever model you gave it. And I think Harvey was kind of similar, right? It just >> Yep. >> picked a model under the hood or whatever. But now, I think they're seeing one, they don't want to compete with Claude. They know if they send all their things, all their data through Claude or OpenAI, that they might eventually be competing with them. So, they kind of want to own their own destiny. And now, with some of these open-weight models, it's becoming easier and more possible for teams to do that.

14:38 >> on a panel with a principal engineer from Harvey, and we were talking about the eval loop. And Harvey has so much data now on different cases and different users. They have so much, they have huge eval suites, they have this whole like discipline in making sure their evals are good, because they're what we're what we call real-world agent. They impact the real world through, I mean, obviously, law, I guess, you know, does impact people. And I think anyone who is a real-world agent that has a lot of data, eval suites, and actually takes observability seriously, will be doing the same thing.

15:16 >> Yep. And I think there's different methods for doing it, right? Are you just doing like traditional like fine-tuning, post-training, reinforcement learning? I mean, there are different practices for how you could pull this off, but I do think that it is getting easier and especially if you have a real-world use case where you're collecting data first, eventually you can use that data and train potentially a smaller model or a cheaper model to do it as well or maybe, you know, again, a of close to frontier model that can then, perform better and you own the the outputs, right? You own that now.

15:47 >> Yeah, Kim e K3 base is not a bad model to start from to then train with your company, let's say your company swag in there, and then you can now make more money, right? You obviously you put some money into training, but now your inference costs are way lower than they what they used to be if you're using Fable or something. I would say this is kind of a threat to frontier usage in companies with data. Obviously, you have to have a practice in your company to harness this data and do things with it.

16:17 You probably have to have AI squad like Harvey does. There are a lot of costs that go into this, but it's not about what happens now, it's about the horizon of once you have this, your margin will get better over time. >> And if you think about models getting smaller, right? We talked about Qwen 27B. The smaller the model, the easier it is to train it, right? So, the price of training your own model is going to continue to be compressed downwards. So, I think over time you're going to see even more of this. And Harvey's has been, you know, kind of on the leading edge of a lot of a lot of things, right?

16:47 As as being a prime example of a successful agent out in the wild in a specific niche or specific vertical. But continuing to talk a little bit about law and the legal use case, this is a post from a16z. AI power users are showing up outside of tech. The fastest growing Codex adopters since February, legal is 108x. So, basically they're just figuring out, enterprise job title when they sign up for Codex, and in the legal use case it's 108x since February. Sales is 41X, recruiting is 41X, marketing 26X, healthcare 24X, and then coding, which, you know, is what we're all probably using it for, is still impressive. It's 5X, but that's tiny compared to 108X.

17:31 >> Yeah. Dude, I want to see all of these even go even higher. Like healthcare needs to be in the 100X. Actually, I don't really care about sales, to be honest, but law, healthcare, real-world they need to be in the 1000X. That's where we should put our energy into. >> Yeah. I mean, I think the reason coding is lower, the reason sales is lower is because those are the com- those have been kind of really common use cases.

17:54 >> Yeah. >> Sales agents have been around for a while now. A lot of people they've already been using them, so they're the growth rate isn't as high, but it still is wild to think that even in coding if it's been it was 5X. It means that most people aren't using these tools on a daily basis. There's still a lot of room for it to grow. We have not hit the peak by any means, and I think it's going to continue to expand outside of just engineers and developers into other places.

18:22 So, let's cover some additional quick hits. This is another post from A16Z. Again, some play on the show today, but it says humans are the minority user of AI. Agents burn nearly five times the tokens that people do, up 14X since February. So, just like, you know, we've said in this show, you know, you're not writing docs for humans, and then agents are consuming docs more frequently than humans. Agents are consuming websites. I think Cloudflare announced, right, more than humans. Agents are now using tokens more than humans.

18:53 >> Wasn't it always the case though? Like don't we use AI through our agents? >> Yeah, but I mean I think a lot of cases it was like single turn call and response for a lot for the longest time with ChatGPT. Like then you had some tools, right, but it still was just like user message, now you get a response. And I think what this is probably charting is that an agent is deciding to like use tokens on its own, right? It's like calling to the LLMs on its own. So, again, don't know you know, you can read the post and see exactly how it's measured, but I think it's because the agents are running for longer, they're doing more autonomously, they're calling sub-agents, right? That are doing work on the humans' behalf. Humans are becoming further disconnected from the initial query to the result that you get, right? Where in the past it was want the agent to run more than 20 seconds because you knew if it did it probably was going to go way off track. Where now people are trusting it for longer horizon tasks. This post came out on August 18th. This is from Cursor. We talked about last week Cursor Origin, but they wrote this really detailed blog post called Git at any scale. And it's really about why it's so hard to scale Git the way it was built and then their solution for how they have basically said they they solved Git's scaling problem. They can infinitely scale it with this approach.

20:11 Did you read through this? >> Yeah, it's very interesting. >> It definitely is a very detailed post and it it's very well done. So, if you if you're curious on like scaling Git and it's not just about it's just scaling in general, I think. It's a really instructive post. >> Yeah, the main thing after reading it, the main thing main takeaway is it's kind of like what we said on the show where it's not entirely GitHub's fault that it's had these issues. Mainly because the volume is super high and their architecture started in 2008. And so, to evolve over the years to then suddenly change your architecture to supply this mass scale that they've never seen before, which is an exponential growth, it just doesn't make sense for them to take all the blame.

20:56 Obviously, the the market is huge factors, but it also makes this case where maybe you do need to have something like Origin or something built with current technology that's meant for this scale cuz it breaks down like every architecture piece of GitHub or Git in general and then what you need to do to to scale that when it gets load. And so yeah, I I had more empathy for GitHub after reading it. >> Yeah, and I think if you read it, you'll you'll say that Cursor definitely doesn't blame GitHub, but I think this is Cursor's like write a really detailed post around why Git is hard and show that you have a very clear understanding and you can handle the scale.

21:34 >> It's honestly a marketing tactic as well, right? It's like showing your expertise, showing why Cursor Origin has to exist, showing how you solve the problem so people can't just say, "Oh, like yeah, it's easy now because you don't have the traffic." But when you get to GitHub scale, but they're trying to prove ahead of time, "No, this actually is the architecture that will scale and you know, unfortunately for Git, they've they kind of have this legacy architecture that's been set up for years now." And so Cursor wants to show itself as really have deeply thought about this problem and solved it in a way that can inspire confidence for people who want to make the shift or who are considering moving off of GitHub.

22:12 >> It's not only not to say that Cursor Origin has proved that they can scale yet, but they think with architecture that they've done, they will. >> Yeah, I mean the theory is there. Now it's does it work in practice? That's always the question. >> Cuz they're trying to make distributed file system work which GitHub tried and decided not to do. We'll see. >> And then this is Tobi Lutke from Shopify said based on that last post, Git at scale has been one of the most interesting blog posts I've read in a while. It came right when I was frustrated with Shopify's internal Git system. As an exercise, I've implemented over the weekend as open source. It's a single Rust binary that you can point at any S3 type object store. It uses WAL and CAS primitives and requires no other data store. It also implements bundle URI so large Git repos like the Shopify mono repo are very fast download as a chain of static bundles. So there you go. They basically, you know, open source their own kind of version of it.

23:05 And again, it's spin off is a vibe coded over a weekend. I don't know if I'd trust it, but >> in production right now. >> Yeah. Yeah, I don't think Shopify's running in production on it. But, it would be cool if someone did, you know, take something like this, open source it. So, I a cursor origin of sorts could exist. >> The a funny comment on this post was like, "Yo, can you just join cursor origin? Like, why are you CEO of Shopify? They don't need you anymore."

23:30 >> >> And that's the show. You can follow us on X at Mostra. You can subscribe on YouTube mostra.ai. You can follow me on X at smthomas3. Follow Abi at Abi Iyer. We do this show every Monday, usually every Monday at least, right around noon Pacific time. We do the news. We bring on guests, talk some smack around some AI concepts or SF-isms occasionally. You know, we we have some fun. If you are tuning in live, thank you. We do appreciate all the comments, all the the love, and sometimes the hate that you you give us. Any parting words before we close out, Ali?

24:04 >> Don't be a meat proxy. >> Don't be a meat proxy. And with that, let's get out of here. >> Peace. >>

Summary

The episode discusses the complexities and community reactions surrounding AI agents and their integration with various tools. It highlights the debate on whether users should rely on their own agents or utilize specialized agents built by tool creators, alongside speculation about a mysterious AI model called 0xAlpha and its implications for the industry.

- The community is divided on whether to use personal agents or rely on specialized agents for specific tools.
- Concerns exist about the performance of agents when overloaded with multiple connections to various tools.
- The stealth release of the 0xAlpha model has sparked speculation about its origins and capabilities, with theories suggesting it may be linked to Google or a smaller company.
- Open-weight models are gaining traction, with a significant increase in their usage compared to closed models.
- Slack has introduced a coding feature that allows collaboration between humans and AI agents, raising questions about its effectiveness compared to third-party solutions.
- The legal sector is rapidly adopting AI tools, with significant growth reported in AI usage for legal tasks.
- The episode emphasizes the importance of real-world applications and data in training specialized models, as seen with Harvey's new legal-focused model, Tenent.
- Cursor's blog post on scaling Git highlights the challenges of legacy systems and positions Cursor as a solution for modern scaling needs.

Questions Answered

What are the community's thoughts on using agents versus tools directly?

The community is divided on whether to use their own agents to access tools or to rely on built-in agents within those tools. Some prefer the convenience of integrated agents, while others want to maintain control through their own agents.

What are the implications of the recent speculation about token capacities and models?

There is significant speculation regarding the capabilities of new models, particularly concerning their token processing capacities. This has led to various theories about the origins and performance of these models, with some believing they may be tied to major companies like Google.

What trends are emerging in the use of open weight models?

There has been a significant increase in the share of open weight models being used, indicating a shift in preference among users. This could be due to either a migration from frontier models or an increase in the overall token usage of open models.

How are companies leveraging data for model evaluation?

Companies like Harvey are focusing on real-world applications and have developed robust evaluation practices to ensure their models perform effectively. This includes using extensive data and evaluation suites to refine their models.

What challenges does Git face in scaling, and what solutions are being proposed?

Git faces significant scaling challenges due to its architecture, which was designed over a decade ago. New solutions like Cursor's approach aim to address these challenges by enabling infinite scaling, highlighting the need for modern technology to handle current demands.

© transcribe · For agents Built with care and craft by Gokul Rajaram