Section Insights
Introduction to Agents Hour
What is Agents Hour about?
Agents Hour is a weekly show hosted by Shane and Abby, focusing on AI news, discussions, and problem-solving. They aim to engage the audience with relevant topics and invite participation.
- The show airs every Monday at noon Pacific.
- It features news, guest interviews, and discussions on significant AI developments.
- Listeners are encouraged to provide feedback, especially positive reviews.
The Concept of 'Meet Proxy'
What does the term 'meet proxy' refer to?
'Meet proxy' describes a person who forwards AI-generated content without validating it. The term gained traction after a viral tweet, sparking discussions about the responsibilities of sharing AI outputs.
- The tweet highlighting 'meet proxy' received significant engagement, indicating a strong interest in the topic.
- There is a concern about the lack of validation when sharing AI-generated content.
- The irony of the term itself being AI-generated adds a layer of complexity to the discussion.
Using AI Agents in Various Domains
When should one use AI agents versus relying on personal expertise?
The discussion revolves around using AI agents in domains where the user lacks expertise, suggesting that it's better for experts to build agents in their fields rather than the user attempting to guide them.
- Expertise in a domain can significantly enhance the effectiveness of AI agents.
- Overloading an agent with too many connections can degrade its performance.
- Understanding the balance between personal input and AI capabilities is crucial.
Trends in Open Weight Models
What are the current trends regarding open weight models in AI?
There has been a significant increase in the share of tokens used for open weight models, indicating a shift in preference from closed models. This trend reflects broader accessibility and usage of open-source AI technologies.
- Open weight models accounted for 62% of tokens in recent data, up from 28.4% two months prior.
- The growth suggests a movement towards more open and accessible AI solutions.
- The implications of this trend could reshape how AI models are developed and utilized.
Future of AI Tools and User Empowerment
What is the potential future for AI tools and user capabilities?
As AI tools become more accessible and user-friendly, there is potential for widespread adoption beyond engineers. The ability for users to train their models could lead to revolutionary changes in how AI is utilized.
- The growth of AI tools is still in its early stages, with significant room for expansion.
- Mainstream training capabilities for users could democratize AI development.
- Open-source datasets may become more prevalent, enhancing collaboration and innovation.
Transcript
0:02 Find me. Heat.
0:51 Hey, heat. Hey, heat. Heat. Heat. N. Heat.
1:22 Heat. Heat. Heat. N. >> >> It's Monday noon. The time is here. Shane and I be loud and clear. Pacific vibes, we bring the heat. AI agents can't be beat. AI agents, let's go.
2:03 Losing guest, the big show. Solving problems we do alive. Staying focus that work for you. Making moves we see it through. Every week a brand new show. Tag along and watching knowledge grow. AI agents out. Let's go. News and guest the big show answering questions. We arrive every Monday. Come alive.
2:37 We want reviews, but only if it's a five. We share the drama. We got the drive. Stay in the loop. It's the place to be. Shane and I'll be setting you free. We're here to stay. Tune in Mondays make your day. From the news to the problem solve, get involved.
3:37 Every week in AI, something insane happens >> and there's so much drama. Every Monday, we break it down live. >> We do the news. We bring on guests building in the space >> and we go deep into the stuff that actually matters. >> Agents Hour every Monday, noon Pacific. >> Follow. Don't miss it. Peace. This is Agents Hour with Shane Thomas and Abby Ayer.
4:20 Hello everyone. We are live and all right, we're here. Hey everyone, welcome to Agents Hour. I'm Shane. I'm here as always with Obby. How's it going, dude? >> What's up? And if you've been following the show for a while, you know that we often mention producer Yan. He is not here today. He has other things going on in his life because, you know, who knows, you know, why he has a life outside of the show. I don't know, but he's allowed to, I guess. And so, he's not here.
4:54 >> And so, for the first time in what seems like, you know, I guess like eight months, I'm back behind the control board. So, if things are rocky like that start, you know who to blame. It is not >> We need you, Yan. It's not producer Yan's fault. It is my fault. But we're here. We're live. We're going to have a good show today. It's going to be reasonably short. we we just got some news to cover. A few other topics we want to talk about. You know, quite a bit of news, but it's not quite as jam-packed as always, but there are some definitely juicy topics for us to dig into.
5:30 Well, >> you can tell you can tell how far we've come because we didn't have a control board before or any like ad reels or whatever and so and we don't actually control them, you know? So, it was like where where are things at? >> And we also started late so I was trying to find all the buttons and the videos and and clearly I'm in over my head here. But hey, you know, we we do this every week. The show must go on.
5:56 >> Yeah. >> So, we're here. yeah. How was last week, dude? We were We were in person in San Francisco. >> It was fun, dude. Busy ass week. We had a hackathon on Wednesday with Elastic. We had our bunch of internal things going on as well. Pretty awesome week. Pretty awesome week. What about you? What did you think? >> Yeah, it was fun. E, John, one of our engineers was in town as well, so we got to hang out >> with with YJ. That was fun. It's always a good time. I think really it was it's always just nice, you know, I'm not in San Francisco all the time. I'm there quite a bit, you know, at least every month for, you know, a week usually. But it's just a different kind of excitement. I'm always happy to leave, but I'm glad to be there, if that makes sense.
6:45 >> Yeah. if you can feel the energy in it. It's like I don't I want to be part of that, but I don't want to be part of that every single day >> because >> as someone who does it more often, you definitely get tired from all the the energy, let's say. >> Yeah. I mean, it's good, but also can be distracting, too, right? It's a balance of you you want to be tapped into what's going on, what people care about, >> and you also need to, you know, get things done. And so, there's a tough balance sometimes when there's so much going on.
7:14 I I was excited about the hackathon though, just seeing you some people who hadn't ever used Ma, some people that had used Maestra, but able to ship something in a reasonably short time. And hackathons are just they hit different when you have AI because you can actually show something real where before >> you couldn't even do a 4-hour hackathon. You couldn't show anything if this was two years ago. And so it's always more fun, I think, today to see. You know, also, you know, a lot of it's fake because it's not real. You didn't actually build anything real in four hours, but it does look pretty damn real. And that that's pretty cool.
7:48 >> Yeah. I mean, four hours is quite short, but you know, we've been to these 24-hour hackathons, the multi-day ones, and you have these like mini startups that get created after it essentially, you know, for the really good ones. But yeah, there's definitely a hackathon culture in SF >> for sure. So, funny enough, so I'm in a different location today, y'all. I'm at my childhood bedroom. I'm visiting my parents and I was going to get on the plane yesterday and I found this site. So, I'm going to show you guys something. it's called, and we're talking about SF. So, there's this site called SFMS. I'm going to share it right now.
8:32 SF isms. It's made by actually a friend of the show and a YC founder from Coyal AI. His name is Mahul. If you guys remember this stunt that happened last year where the only way to get more traction on X was to take turn yourself into a a girl. he's the one who started that with the account called Mihoo where he turned himself same picture into a girl and got crazy traction. So this is the person we're talking about. He started this website called F SFISMs and there's a lot of good SFISMs. I'm not I'm not I'm going to go to the ones that are like let's say safe for work, but 996 is on here. AI Psychosis.
9:17 Balboa Cafe. this one's funny because we were talking about Balboa Cafe last week. The mythical marina bar where SF goes to finally meet someone. Nobody ever does. The espresso martini is real though and everyone will be back next Friday to check again. >> 100%. So true. >> That is the vibe. You got to get the, you know, the the lantern of the chandelier espresso martini or whatever they call it, the rack of espresso martinis. That that's the staple from Beloa.
9:44 >> We have a founder mode. okay. All these things. I found one that we were playing. >> Goblin mode. You got to have >> goblin mode from high agency. Like all these dumb words we use all the time. and I honestly we find ourselves using them whether someone told us or not. Maybe it's just like the the the energy of the area. So I just cuz I I found So we were talking last week about how we want more users to give us issues than pull requests.
10:14 Issues are way better than pull requests. And then so I did a tweet about that last week. Sam also did a tweet which was like essentially the the conclusion was like don't be a meat proxy. So when I saw this on SF ISOMS, I was like, "Oh shit." Like that's exactly what we were talking about last week. And so I posted it yesterday. meet proxy. This is in SFSMs. I love the design here. It makes it look like legit. U a person who forwards AI generated text code or other output without reading, understanding, or validating it. the the person acts only as a relay between the AI system and the intended recipient.
10:55 So, this tweet has gone viral. which I did not intend. It was just for me to post, but my phone has not stopped buzzing since yesterday. it's over almost at 300k views, 7.8K likes, etc. But it has definitely started a discussion, right? There's a lot of comments here, a lot of quote tweets. A lot of people that I respect from the past are like, you know, they are taking this post very seriously, which is cool. That was not my intention though. I meat proxy everybody, right?
11:30 This is an AI generated thing. I meet proxied to everybody else. It's also written wrong. There's three tri two triagrams in the sentence. so anyway, like this was generated by AI as well, and I meet proxy y'all. So, I guess it's the ultimate irony and meta thing. but I'll leave it to you, Shane. What do you think about this concept or term? >> I mean, okay, so first of all, this has come up because I've seen too many examples of this where someone says, you know, the the worst thing I think the thing that always makes me the the most mad is when someone will will say something and say, you know, Claude said this, I haven't validated it.
12:11 >> Yeah. And my thought is I'm not an idiot. I can send it to Claude myself if I wanted to. You know, it's like I thought, you know, the point is like you should validate it before you send it because you're just all you're doing is you're throwing work over the wall to someone else because AI generated text is hard to read. Sometimes it's dense. It's actually like a taxing to read output from an LLM and try to understand it because of just how it typically talks. Not every model is the same, but most follow a similar pattern and it's actually hard to read. It does take effort to read that type of writing. So, if you're going to send me a wall of text, I hope you've at least validated it first. Now, there are some times where I think it's totally fine to to send over something that the LLM generated if you've read it and it makes sense. I have an example today where Tyler was asking on our team about something that a customer wanted and I pulled up the transcript. I had an agent like summarize it. I read through every line of it and I said, "Yes, this gets I was in the call. This is roughly what the customer wants."
13:19 >> Yeah. >> I cut out some of it and I sent it over. And I think that kind of thing is fine because I did some actual work, but I think the whole idea of like throwing it over the wall is not is not fine. >> Yeah. >> Now, I did see one comment in here that I thought was interesting because >> one person said, "Actually, I think it's okay because we used to have this site. Let me Google that for you."
13:44 >> Yeah. >> You'd send someone the thing and it would Google it. And I thought that was interesting. But the caveat here is I have done this before with you know someone who's less technical who I know doesn't know how to use AI just like someone doesn't you know at one point maybe didn't use search engines very effectively sure send them some basic information again I've validated it but I I'll send it over but anyone that I'm working with on a day-to-day basis we're all using AI every day we know how to use it there's no doubt that I know how to hook my >> agent to your MCP server and get the information, right? Like you I can do that. So, you should expect if you're dealing with a technical person that you should be doing some work or you shouldn't do it at all. If you don't have the time to actually do the work, don't just send over the result because you're not saving them time. You're actually >> wasting time because they have to put in effort now to read your response >> and they have no context of how you got there. At least if they did it themselves, they'd have the context.
14:44 There was another comment about how this is not a new a new term and I totally agree with this example. So back in the day you would be in let's say you are one step away from your your manager. So you have like a step level manager. So there's like you know your manager step level you. And often times if that person is not good at their job, they are a meat proxy between the upper manager the upper manager's question to you and back. Right? So these these have existed let's say forever, but it's way more prevalent now when people are just throwing clawed sessions over the over the wall for you.
15:28 >> Yeah. So, Lemi in the chat says Reed Jordan did a video describing this, but called it boneless meat layer for LLM. Yeah. All right. Similar concept, similar idea. >> Well, we're just a flush that's like passing LLM tokens between each other at this point. >> so yeah, don't be a meat proxy. You know, use LLMs, use the agents, >> but understand it. Please understand it, parse it, do a little bit of the work so you can save the person on the other end the time and the effort if it's important to you. Otherwise, don't do it. That's my my feedback.
16:05 >> One trick I like one trick I like to use is if I do have this dense amount of text and I understand it, I read it. I ask it for just like bullet points, one sentence bullet point to summarize this stuff and I'll send the summary. And then if you want to double click into anything, we can talk about it. But this is like just the highlevel thing like this is what we're we're going to do or whatever.
16:28 >> Yeah. Put a put a little work into it. Put the effort goes a long way. otherwise >> you look a lot smarter if you do that. >> Yeah. Because you can tell when someone's just sending AI generated things or you know I've also seen you know I had my agent wrote this PR. I haven't reviewed it yet. Can you review it? >> Yeah. >> What the No. >> Yeah. Don't do that. If you think it's a valuable PR, review it. You know, it's otherwise don't do it. And I I can have my agent do it if I if we want to. So, I think there's a again, it's a it's a brave new world. There's a little bit of tact and how you should be using these things. Don't throw work on other people. Use it to save this. These are meant to save time, right? But you still have to do some of the work yourself.
17:14 You can't just offshore or offshoot it to someone else all the time. >> Yeah. Like another way we we use it is like if we're like an incident response right now, we'll say like, "Oh, my agent said this, but I need to verify because that information could be useful for someone else who already knows what like who could use that as like a clue for what's going on cuz you obviously don't know. So, it's nice to have caveats, too. Like, hey, this might be something, but I'm gonna go double check, but here's what my here's what I'm looking into." We used to do this before, too.
17:46 we'd have thoughts. We would say, "Hey, I'm going to go explore these three things, but I don't know if it's true or not." And then you'd go. Now, you can just get that from your LLM, paste it, and then go validate it. So, there's there's ways to be a good me proxy. >> Yeah, exactly. There there are certainly times where if someone else knows the area of the codebase and you're not sure, but you you have a reasonable level of confidence that it could be good, you should have at least done some of the work, though, so you can have an intelligent conversation about it rather than just saying, "Can you validate this? Does this look good to you? I don't know. So, I think part of it is we still have to keep our critical thinking skills. You should still be doing some of that work. I know it's it's tempting to just send over, you know, sling PRs around and sling text around, but don't be a meat proxy.
18:34 >> Don't be proxy. >> All right. and with that, you know, this is a live show. Yeah. Thanks, Lemi, for dropping a a comment in YouTube, I guess. But we're on X, we're on YouTube, I think we're on LinkedIn today. So feel free to drop in a comment. We will pull some of them up live. We will discuss it. If you have not already reviewed the show, please go to YouTube, give us a like, give us a subscribe, go to Spotify, go to Apple Podcast, wherever you get your podcasts, live streams, whatever. We do appreciate the likes, the subscribes. That helps other people like yourself find the show.
19:12 And with that, I suppose we should probably just get into the news. Should we do it? >> Let's do it. Let's do it. All right. Like every week we do the news. Welcome to Agents Hour. Today we have a pretty short episode considering some of the past weeks we've had, but there are some big juicy topics that we do want to dive into. So, why don't we go ahead and get into it.
19:59 All right, a little preview for the show. >> Let's start with this post. This is from someone named Ally and she says, "I really don't want to use your agent. I want to use my agent to use your thing." And this got a ton of engagement. A lot of people, you know, were commenting, resonating with it. Some disagreed, a lot of people agreed. But the question is in this new age, do you just want to use your agent and connect to someone else's tool maybe through MCP or whatever or do you want that tool to just build the agent themselves and you use that? What's your take?
20:38 >> So I I agree with this statement for things where like the boundary of the product I like is also on my end, right? like if it's a coding thing or infrastructure or whatever. But here is where I disagree. I would rather have an agent in Google Slides than have my agent talk to Google Slides. You know what I mean? Where these products are encapsulated like dscript. Maybe I want to be have my own agent. But no, I'd rather just be in dcript just making moves in this thing. But where I think this comes into play, if I were to make it more general, if I also have industry experience in the product, I want my agent in there. If I have no business doing that in the first place, like making music or anything, I want to use their agent. That's where I kind of draw the line. What about you?
21:30 >> Yeah. So, you're saying because they would conceivably be better at building an agent in the domain that you're not experienced in. and you wouldn't be able to guide the agent to the result. It's probably better if they just build it. They control the context rather than you try to control the context. >> Correct. I think both use cases are valid and I will ultimately think it depends very similar to you for instance at first you know an example is post talk like I don't use postto every day I use it quite a bit but not every day and I've noticed that the agent in Postto is actually pretty good like I hadn't used it very often and it was pretty good it does what I need it to and now I don't have to connect the PostG MCP that probably has 50 tools to my agent. I I do worry that over time your agent might get worse and you're not validating it or you don't know it because you're just slamming a bunch of MCPs in there which adds to the context and every what the average person doesn't understand. If I have a hundred MCPS connected to my agent, >> that's not very good, right? Like that's going to impact the results. It's going to slow down the agent. There's more tokens. all that stuff, you know, factors in and the average person doesn't really understand that. So I think there is this risk of at least today and maybe this changes over time but if you just connect your agent to every tool possible now your agent has to try to do every single thing and I think that you might get worse results and you might not know why your agent seemingly got dumber over time. You might think it's the model.
23:10 You might think it's you the way you're using it. But maybe it's just your tool context is getting polluted because you're connecting so many damn MCP servers to it. >> Yeah. >> so that's why I think specialized agents are useful because you don't have to worry about that. You only go in there some of the time. I don't use PostG every day. If I did, I would much rather connect it to my agent and just use the MCP. Right. Another thing is like notion. Notion's agent. I think it kind of sucks and I use notion almost every day. So yeah, I'll just use the MCP and it works right. I understand what I want it to do. I can guide it, but I I'm taking on that, you know, all those extra tools in the, you know, in my agent, right? If especially if I just like connect it to the my agent and I don't have different, you know, settings or profiles or whatever or projects that have like different configurations.
24:01 So, I think in general, I agree with Ally on a lot of like for a lot of the tools, but not for everything. Like they're specialized tools. Just let let them build the agent. They're the experts. Make the agent really damn good. And, you know, maybe someday my agent can talk to your agent and not have to care about the context. Just like passing, >> you know, messages across. But for now, I'll I'll use the agent when I need to.
24:23 >> If your agent is built in Ma, I'll use it. There you go. >> There you go. There you have it. >> >> All right, we got to talk about I don't know if it's 0x alpha. I've been saying oxal alpha, but we got to talk about 0xal alpha because it's a bit of a mystery and it kind of just hit the timeline later last week around what is this model that just seemingly popped up within open router. So, this is from MTS Live. It says, "Situation explained.
24:54 A stealth model on open router is beating Fable and Soul on coding and nobody knows who made it. Ox Alpha has a 1 million token context window, text, image, and video input free for a week with capacity for a 100red trillion tokens a day. And I think that's what blew people's minds because they didn't understand where could a 100red trillion tokens a day come from. That compute doesn't just exist lying around. There's only a few companies that >> in theory could have that capacity. I love the conspiracy theories that have come from this.
25:26 >> Yeah, there have been a lot and we'll talk about some of them. So, this Brandon Carl says, "Whoever is offering 0x alpha had the ability to offer the entire token capacity of Google immediately, 100 trillion tokens a day, which led to a lot of speculation that was this a Google model?" You know, the Gemini team was posting some kind of obscure things making people think maybe this is a Gemini model. Could this be the new Google? How good is this model?
25:53 People are starting to run benchmarks on it. >> Yeah. >> This post says, "WTF Ben ran this mystery model through 10 deep tasks and it scored over 80% versus 65% for Fable and 52% for GPT56 Soul. This is insane. Probably a Chinese company. Either a new GLM or Kimmy model, I reckon. I did see some posts that said they ran more of the tasks. It didn't score this well on everything. I think this was, you know, a bit cherrypicked. I think overall it seems like what I've seen it's actually performs not as good as Fable in 56. So maybe maybe it is Gemini. I don't know.
26:38 But it's not at least not according to this post which says breaking the stealth model 0x alpha is Z.AI. So it's a GLM Everyone else is guessing it from tokenizer vibes and emoji rates. I made the server say its own name out loud. Sent it one malformed request and it threw a Java stack trace naming its own internal. What do you think? Is it GLM? >> Dude, I mean it' be very it'd be good for the narrative that it is, but I'm really hoping it is Google.
27:12 but I think it might be ZA Zai. And this is all like some speculation, right? But it seems that it's a smaller model, but it's doing it does really well on at least the benchmarks that people have run. Not fable level though, but the theory is that this is almost like a GLM flash type model or a smaller model that maybe >> is on par or better than the last like GLM 5.2, but is a smaller model. And so that is how Z.AI AI can serve a 100 trillion tokens a day is that it's just significantly more efficient because it's so much smaller.
27:54 >> Yeah, there there are other indicators that would lead for this to be Z. So like Open Code offered this through Open Code Go for free or whatever. and they were tracking the token usage. People are using it. If it was Google, anybody would have to sign a bunch of NDAs. We have done this. I don't know if I'm not allowed to say that. I've done it, but okay. But you have to sign NDAs to use future models. so I don't think it's Google for that reason alone.
28:27 They're not going to give this away in stealth that other people can use without tons of like bureaucracy. So it's got to be an open model. That's what my my guess is. And you know, if this evidence is damning enough, then it probably is Z. >> Yeah. And I think we will eventually find out. I think the hype has died down from the model a bit because I think people realized it wasn't as good as what they originally thought >> once benchmark started coming out.
28:56 >> So I don't think it's Frontier. It's not a new Frontier model. But if it is, you know, a GLM model and it's much smaller, you know, if you think back to, you know, the Quen model that came out, which we will talk about, and how that's, you know, like a small basically on device, you know, you can run it on your own computer model and it benchmarks very well. Well, then maybe GLM has done the same type of thing where it's a smaller model. We don't know the size yet, but maybe it maybe they can compete with some of close close to the frontier without being frontier sized.
29:32 >> This like stealth release stuff, it it takes the community by it like bewilders the community, takes them by storm. So, I'm curious how many stealth models we can go through before it like loses its luster, you know, or is it always going to be a really hype way to debut a model to give it out in stealth first? >> Yeah. I mean, I think there's just a marketing aspect to it for sure. >> Yeah. Will the Frontier ever give it out in stealth, you know? I wonder.
30:00 >> Feel like they don't have to, right? They get the marketing just by, >> you know, No, their marketing is it got banned by the government. That's their marketing stunt. >> so Quen breaks the size curve. We talked about this a little bit last week, but if you look at the artificial analysis agentic index on the Quen model, so it's a little hard to see here, but Quen 3.827B 827B is actually ahead of models like Opus 48, ahead of 56 Luna.
30:37 So it is, you know, ahead of GLM52. It's ahead of these much larger models, right? And this is only a 27 billion parameter model. >> Yeah. The idea is like you can if you told someone you 18 months ago or not even you told someone six months ago that they could have you know an Opus level model on their device like on their computer running you know obviously still needs to be a decently you know has to have be decent sized right you still got to have some hardware to run that >> but I think people would freak out you know running an Opus level model on consumer hardware is pretty insane and it >> makes you know makes me wonder like how many tasks do we even need to pay for tokens anymore.
31:25 >> Will will teams just have like a you know will everyone just have their own like reasonably good model on their own machine that they run most tasks to and then the the hard tasks get sent to the frontiers? I don't know. We have a thing in the chat here from Lemi if you want to pull it up. >> Yeah. So, Lemmy says OpenAI did stealth models previously cursor as well and said maybe it's Grock 47. Some speculation there.
32:02 Yeah, we did have stealth models from OpenAI cursor, but I feel like the distribution has changed. All these routers that exist now, but we'll see. So, let's talk about open weights model, open weight models a little bit more. So, this is from Garmmo from Verscell says, "Today is a record day for openw weight share of tokens on Verscell AI gateway and open routers posted similar things." So, this is not just Versel. This is across a lot of the popular AI LLM routers. And so August 22nd, 62% of tokens were open weight models. And in June 24th, so about two months ago, it was only 28.4%.
32:46 That's a pretty >> wild increase. I am curious on is it because there's two ways you can kind of read this. Maybe the, you know, a lot of people moved from Frontier models to open models or maybe there's just a lot more tokens being sent through and a higher proportion of those tokens are now becoming open models. >> So, they don't really share like total token growth count or anything. So, you can't really >> determine if it's, you know, are closed models in, you know, are they total are they did the total count go down? I doubt it, but it probably just didn't grow as fast as the open models. And I think if you're using Frontier tokens, a lot of people, the majority of people are going right through the providers, right? Not through a router.
33:32 >> Yeah. I mean, I would suspect this too, like it's going to be more open because you know, you are not paying max coding plans on these gateways. So, and there are a lot more open models to choose for for price. So, cool. We're gonna talk about Slack code. >> This is from Mark Benny off on August 19th. It says, "Don't code alone. Slack code is live. Humans and agents. Same channel, same work. Launching today with agents from Enthropic, GitHub, Cognition, Verscell. This is real multiplayer coding. See it at Dreamforce.
34:13 What do you think of this?" I mean, if their website was up that day, I would have been more happy about it, but it was down for like the whole day. >> Yeah. I mean, they said see it live and then people were saying it wasn't quite live. But if you watch the video, you know, you can go check out the tweet. >> There's like basically a code panel that pops up in a conversation, right? You're basically >> prompting in Slack and you're just having it write code.
34:40 I think you know we have mentioned that coding you know from your mobile device is the dream and Slack has a pretty good mobile app so now you can in theory code from anywhere within Slack. I don't I'm not convinced that Slack's going to win this on its own. I feel like you know if you have ever used Devon Devon's Slack agent is pretty good. I know linear is investing heavily in their agent as well within Slack. I think obviously Slack is trying to compete with you know all the apps that integrate in with it and trying to just do it themselves. They want maybe Slack wants to own a piece of the coding pie in some ways. But we'll see how many people actually use the the Slack version or if if it's like, you know, you know, there's always those things where there isn't there is a version in the app itself, but people use third parties because third party ones just end up being better.
35:38 >> And so I think we'll probably see some of that as well. If I were >> Could you imagine a world where Slack bans third party agent apps just so you can use Slack code? I feel like they can't right like Salesforce is is much is kind of built on this ecosystem of like you can build sales I mean they have a proprietary language of course but they >> it's all built on like this marketplace and all these different apps and all this different like it's like so complicated they have all these different companies that just provide Salesforce type services I feel like they wouldn't want to go back and change that specifically for Slack >> I think they want Slack to be open >> Slack to the communication tool, but maybe not. Maybe they maybe they're looking at all the, you know, the revenue from all these coding agent companies and saying maybe we can grab a piece of that.
36:29 >> Is this GA yet? Can we use this? I don't see it in our Slack. >> Yeah. Well, we didn't go to Dreamforce this year, so >> Oh, that's why. >> I mean, maybe we should we should reach out and see if we can get it turned on. I'm guessing there's a way to get it, but maybe you have to know someone right now. I don't know if you're in the chat. Have you are you going to use slack code? Would you use it? Let us know.
36:56 So there's been a lot of talk recently and especially the last week around just legal as a breakout use case for AI. So Harvey has introduced tenant which is the first model post-trained for legal. So Harvey's actually not just using Frontier models or they're actually training their own or post-training their own. So Tenant is a Kimmy Kimmy K3 base that we postrained with fireworks on a corpus of publicly available legal data, synthetic data, and human expert data simulating long horizon legal work.
37:32 So this is pretty wild. That means there are actually companies that started as just you could call them like a model wrapper, right? Yeah. They were just an application on top of a model. It now >> I would say cursor was the same way, right? Like cursor didn't start as building its own >> LLMs. It just was kind of like an agent that used whatever model you gave it. I think Harvey was kind of similar, right? Just it picked a model under the hood or whatever. But now I think they're seeing one they don't want to compete with Claude. They know if they send all their things all their data through Claude or Open AI that they might eventually be competing with them. So they kind of want to own their own destiny. And now with some of these openw weight models it's becoming easier and more possible for teams to do that. I was on a panel with a principal engineer from Harvey and we were talking about the eval loop and Harvey has so much data now on different cases and different users.
38:34 They have so much they have huge eval suites. They have this whole like discipline in making sure their evals are good because they're what we're what we call a real world agent. They impact the real world. through I mean obviously law I guess you know does impact people and I think anyone who is a real world agent that has a lot of data eval suites and actually takes observability seriously will be doing the same thing >> yep and I think there's different methods for doing it right are you you know you just doing like traditional like fine-tuning post- training >> reinforcement learning I mean there are different practices for how you could pull this off. But I do think that it it is getting easier and especially if you have a real world use case where you're collecting data first eventually you can use that data and >> you know train potentially a smaller model or a cheaper model to do it as well or maybe you know again a close to frontier model. Yeah.
39:38 >> That can then perform better and you own the the outputs right you own that now. >> Yeah. Kimmy K3 base is not a bad model to start from to then train with your company, let's say your company swag in there and then you can now make more money, right? You obviously you put some money into training, but now your inference costs are way lower than they what they used to be if you're using a Fable or something before. I would say there's kind of a threat to frontier usage in companies with data.
40:14 Obviously, you have to have a practice in your company to harness this data and do things with it. You probably have to have AI squad like Harvey does. There are a lot of costs that go into this, but it's not about what happens now. It's about the horizon of once you have this, your margin will get better over time. >> And if you think about models getting smaller, right? We talked about Quen 27B. The smaller the model, the easier it is to train it, right? So, it's going to only continue, the price of training your own model is going to continue to be compressed downwards, right?
40:49 >> Yeah. >> So, I think over time you're going to see even more of this. And Harvey's has been, you know, kind of on the leading edge of a lot of a lot of things, right? as as being a prime example of a successful agent out in the wild >> in a specific niche or specific vertical. >> But continuing to talk a little bit about law and the legal use case, this is a post from A16Z and it's, you know, a little bit small here. Maybe you can can't really zoom in, but I'll kind of read the results. and that is that AI power users are showing up outside of tech. The fastest growing codeex adopters since February. Legal is 108x.
41:34 So basically they're just figuring out enterprise job title when they sign up for codeex. And in the legal use case it's 108x since February. Sales is 41x. >> Recruiting is 41x. Marketing 26x. Healthcare 24x. And then coding, which you know is what we're all probably using it for, is still impressive. It's 5x, but that's tiny compared to 108x. >> Yeah, dude. I want to see all of these even go even higher. Like healthcare needs to be in 100x. actually don't really care about sales to be honest, but law, healthcare, real world they need to be in the thousandx, you know, that's where we should put our energy into.
42:20 >> Yeah. I mean, I think the reason coding is lower, the reason sales is lower is because those are the common, those have been kind of really common use cases. Yeah. >> Sales agents have been around for a while now. a lot of people, they're a known quantity, so maybe less people are, you know, they've already been using them, so they're the growth rate isn't as high, but it still is wild to think that even in coding, it's been it was 5x, right? So, there's still it means that most people aren't using these tools on a daily basis. There's still a lot of room for it to grow. We have not hit the peak by any means and I think it's going to continue to expand outside of just engineers and developers into other places.
43:04 >> Yep. >> Justin says related to our last topic once training becomes mainstream I think we'll see a lot of crazy stuff like when the normal user can train their models it'll be revolutionary. I agree. I think that >> just like it is now pretty easy for someone to build an agent, right? That that was a hard thing 18 months ago to do. It's still not easy. I mean, I think anyone who's built a real world agent knows that it's actually still incredibly difficult, but the tools are there to make it quite a bit easier. I think once the models become small enough the you know people get hands on good enough GPUs to do it I think it's the costs are going to go down the time to do it's going to go down and people are going to do some pretty amazing things >> and I hope people open source data sets >> yes >> but that is very proprietary today but it might not be tomorrow >> so let's cover some additional quick hits this is another post from A16Z we're getting some play on the show today.
44:06 but it says, "Humans are the minority user of AI. Agents burn nearly five times the tokens that people do, up 14x since February." So, just like, you know, we've said in this show, you know, you're not writing docs for humans and then agents are, you know, consuming docs more frequently than humans. Agents are consuming websites. I think Cloudflare announced right more than humans agents are now using tokens more than humans.
44:37 >> Was that always the case though? Like don't we use AI through our agents? >> Yeah, but I I mean I think a lot of cases it was like single turn, right? You it was like single turn call and response for a for the longest time with J GPT. Like then you had some tools, right? But it still was just like user message. Now you get a response. And I think what this is probably charting is that an agent is deciding to like use tokens on its own, right? Is like calling it calling to the LLMs on its own.
45:16 So again don't know you know you can read the post and see exactly how it's measured but I think that it is true that a it's because the agents are running for longer they're doing more autonomously they're calling sub aents right that are doing work on their be on the human's behalf humans are becoming further disconnected from the initial query to the result that you get right >> where in the past it was you know you didn't want the agent to run more than 20 seconds because you knew If it did, it probably was going to go way off track where now people are trusting it for longer horizon tasks.
45:56 This post came out on August 18th. This is from cursor. We talked about last week cursor origin, but they wrote this really detailed blog post or called git at any scale. And it's really about why it's so hard to scale git the way it was built. and then their solution for how they have basically said they they solved get scaling problem they can infinitely scale it with this approach did you read through this >> yeah it's very interesting >> any big takeaways from I mean I read through about it's pretty long so I did not read this whole thing I read through about half of it and I understood about you know half of what I read so maybe I understood a quarter of it but it it definitely is >> a very very detailed post around like how like how and it it's very well done.
46:50 So if you if you're curious on like scaling git and it's not just about it's just scaling in general. I think it's a really in instructive post. >> Yeah. The main thing after reading it, the main thing, main takeaway is it's kind of like what we said on the show where it's not entirely GitHub's fault that it's had these issues mainly because their volume is super high and their architecture started in 2008. And so to evolve over the years to then suddenly change your architecture to supply this mass scale that they've never seen before which is an exponential growth. it just doesn't make sense for them to take all the blame. Obviously, the the market is huge factors, but it also makes this case where maybe you do need to have something like Origin or something built with current technology that's meant for this scale because it breaks down like every architecture piece of GitHub or Git in general and then what you need to do to to scale that when it gets load and so yeah, I I had more empathy for GitHub after reading it.
48:06 Yeah, and I think if you read it, you'll you'll say that cursor's, you know, cursor definitely doesn't blame GitHub, but I think this is cursor is like write a really detailed post around why Git is hard and show that you have a very clear understanding and you can handle the scale and it's honestly a marketing tactic as well, right? It's like showing your expertise, showing why, you know, cursor origin has to exist, showing how you solve the problem. So people can't just say, oh, like, yeah, it's easy now because you don't have the traffic, but when you get to GitHub scale, but they're trying to prove ahead of time, no, this actually is the architecture that will scale.
48:45 And you know, unfortunately, forget they've they kind of have this legacy architecture that's been set up for years now. And so I think you know cursor wants to show itself as yeah really have deeply thought about this problem and solved it in a way that can inspire confidence for people who want to make the shift or who are considering moving off of GitHub. >> Yeah. But it's not really not to say that Cursor Origin has proved that they can scale yet, but they think with the architecture that they've done it'll it will.
49:19 >> Yeah. I mean the theory is there now. That's theory. >> Does it work in practice? That's always the question >> because they're trying to make distributed file system work, which GitHub tried and decided not to do. So, we'll see. >> And then, this is Toby Lutkkey from Shopify said, you know, that based on that last post, Git at at scale has been one of the most interesting blog posts I've read in a while. It came right when I was frustrated with Shopify's internal git system. As an exercise, I've implemented my as an exercise I've implemented over the weekend as open source. It's a single Rust binary that you can point at any S3 type object store. It uses WA and CAS primitives and requires no other data store. It also implements bundle URI. So large git repos like the Shopify monor repo are very fast to download as a chain of static bundles.
50:14 So there you go. They basically, you know, built their own open sourced their own kind of version of it. And again, it's spend it's vibe coded over a weekend. I don't know if I'd trust it, but >> put it in production right now. >> Yeah. Yeah. I don't think Shopify is running in production on it, >> but it would be cool if someone did, you know, take something like this, open source it, so a a cursor origin of sorts could exist.
50:38 >> Dude, a a funny comment on this post was like, "Yo, can you just join cursor origin?" like why are you CEO of Shopify? They don't need you anymore. And that's the show. As we said before, it was going to be a short one. You can follow us on XMRA. You can subscribe on YouTube, MRA-I. You can follow me on XM3. Follow Obby at Abby. We do this show every Monday. Usually every Monday at least, right around noon Pacific time.
51:12 We do the news. We bring on guests. We, you know, talk some smack around some AI concepts or SFISMs occasionally. You know, we we have some fun. If you are tuning in live, thank you. You know, we appreciate all the comments, all the the love and sometimes the hate that you you give us. Any parting words before we close out? Abby, >> I think we have to end where we began. Don't be a me proxy. >> Don't be a me proxy. And with that, let's get out of here.
51:47 >> Peace. >> Agents out with Shane and I be on the throne. Did you give us that review only if it's a five? Jump on the tube. Make sure to like and subscribe. Do so fresh. Yeah, we keep you in the loop. Get so fly. They bring the whole troop. AI on the rise. Don't miss this power. Welcome to the show. It's AI Sour. Did you just drop in? Is this your first time? Make sure to follow us on next and go like and subscribe. Yeah, learn the principles and patterns in our books.
52:16 The master.ai site. Give it a look. New so fresh. Yeah, we keep you in the loop. Guess so fly. They bring the whole troop. AI on the rise. Don't miss this power. Welcome to the show. It's AI sour. This is the end. We all wrapped up. Another showdown. Another one coming up. AI agent is done, but the news doesn't cease. Shane and Abby, we out of here. Peace.
52:47 >>
Summary
- The show features a catchy intro highlighting the weekly AI news and problem-solving.
- Shane and Abby reflect on their recent experiences at a hackathon in San Francisco, discussing the excitement and challenges of AI development.
- They introduce the concept of "meat proxy," cautioning against simply relaying AI-generated content without validation.
- The duo discusses the emergence of a mysterious AI model, 0x Alpha, speculating on its origins and capabilities.
- They highlight the increasing adoption of AI tools in legal sectors, noting a significant growth in usage compared to other fields.
- The conversation touches on the challenges of scaling Git and the potential of new solutions like Cursor Origin.
- They emphasize the need for users to engage critically with AI tools and not rely solely on AI outputs without understanding their context.
Questions Answered
What is Agents Hour about?
Agents Hour is a weekly show hosted by Shane and Abby, focusing on AI news, discussions, and problem-solving. They aim to engage the audience with relevant topics and invite participation.
What does the term 'meet proxy' refer to?
'Meet proxy' describes a person who forwards AI-generated content without validating it. The term gained traction after a viral tweet, sparking discussions about the responsibilities of sharing AI outputs.
When should one use AI agents versus relying on personal expertise?
The discussion revolves around using AI agents in domains where the user lacks expertise, suggesting that it's better for experts to build agents in their fields rather than the user attempting to guide them.
What are the current trends regarding open weight models in AI?
There has been a significant increase in the share of tokens used for open weight models, indicating a shift in preference from closed models. This trend reflects broader accessibility and usage of open-source AI technologies.
What is the potential future for AI tools and user capabilities?
As AI tools become more accessible and user-friendly, there is potential for widespread adoption beyond engineers. The ability for users to train their models could lead to revolutionary changes in how AI is utilized.