transcribe

Anthropic IPO, OpenAI DevDay Recap, And Why Are Decision Models Everywhere?

Mastra · 1h 26m · transcribed 1h ago
More from Mastra Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Introduction to Agents Hour

What is Agents Hour about?

Agents Hour is a weekly show focused on AI agents, featuring news, guest discussions, and problem-solving in the AI space.

  • The show airs every Monday at noon Pacific.
  • Listeners are encouraged to engage by sharing their insights and experiences.
  • The hosts aim to provide valuable knowledge and updates on AI developments.
# 14:21

Importance of Security in AI Agents

Why is security crucial for AI agents in enterprises?

Security is essential for enabling stable agent identities and enforcing policies across various environments, which is necessary for enterprises to leverage AI effectively.

  • Enterprises face challenges with agent delegation across different clouds and systems.
  • Without security measures, the potential of AI agents cannot be fully realized.
  • Establishing a secure identity layer is critical for the deployment of advanced AI solutions.
# 28:42

Customization and Security in AI

How does customization relate to security in AI systems?

Customization in AI systems must be balanced with security measures like tamperproof audit logs and just-in-time authorization to ensure safe experimentation.

  • Security foundations enable trust in AI systems, allowing for greater customization.
  • Observability and policy enforcement are key components in managing AI agent interactions.
  • The evolution of standards for AI agents will impact how systems are designed and implemented.
# 43:03

Emergence of Decision Models

What are decision models and their significance?

Decision models, like the ones introduced by OpenAI, enable rapid decision-making and have created a new category in AI, impacting how organizations approach classification tasks.

  • OpenAI's decisions API allows for fast, constrained decision-making.
  • The introduction of decision models is reshaping the landscape of AI applications.
  • There is ongoing competition among companies to define and dominate this new category.
# 57:24

Consumer Adoption and Market Positioning

How does consumer adoption impact AI companies?

Companies like Google are well-positioned to leverage their existing consumer base and tools to create competitive AI solutions, particularly in workplace settings.

  • Consumer adoption is crucial for the success of AI products.
  • Integration with widely used tools like G Suite can enhance market competitiveness.
  • Predictions suggest that Google may lead in model performance in the near future.
# 71:45

Managing Contributions in Open Source AI

What is the best approach for managing contributions in open source AI projects?

Encouraging detailed issue submissions rather than pull requests can streamline the contribution process and ensure faster resolution of valid issues.

  • Issues should be prioritized over pull requests for efficiency.
  • Clear communication and triaging of issues can enhance project management.
  • The community's discipline in following guidelines is essential for effective collaboration.

Transcript

0:24 Hey, hey, Heat. Heat. Heat. Heat. N.

1:09 Would you like to be a guest on the show? Visit masterra.ai/guest. Do you have a hot takeache, a problem you cracked, a cool product, or a demo worth watching? Share them right here on Agents Hour. It's Monday, the time is here. Shane and I be loud and clear. Pacific vibes, we bring the heat. AI agents can't be beat.

1:45 AI agents now. Let's go. Losing guess the big show. Solving problems we do alive. Staying focus that work for you. Making moves. We see it through. Every week a brand new show. Tag along and watch your knowledge grow. AI agents out. Let's go. News and guests, the big show answering questions. We arrive every Monday. Come alive.

2:22 We want reviews, but only if it's a five. We share the drama. We got the drive. Stay in the loop. It's the place to be. Shane and I be setting you free. AI here to stay. Tune in Mondays make your day. From the news to the problem solve world, get involved.

3:23 Every week in AI, something insane happens >> and there's so much drama. Every Monday, we break it down live. >> We do the news. We bring on guests building in the space. >> And we go deep into the stuff that actually matters. Agents Hour every Monday noon Pacific. >> Follow. Don't miss it. Peace. >> Peace. >> This is Agents Hour with Shane Thomas and Abby Ayer. >> Hello everyone and welcome to Agents Hour. It's Monday, October 5th. It's right around noon Pacific time. For those tuning in live, thank you. And if you're watching this after the fact, thanks for you know, tuning in and leave us a comment, a review, a like. We appreciate that. We have a great show today. We are bringing back Security Corner. It's been a while. It's it's time. You know, security is important.

4:13 We got to give it we got to talk about it. So, we'll be talking about security. And be bringing on Ally, who many of you probably have seen here before, but it's been a little bit. So, we'll see what she's been up to. And of course, we will do the news because there's always some drama. There's new model drops every week. We got to talk about them and we'll go deep into all that. What's up, dude?

4:35 >> What's up, dude? How are you? >> Good. So, it is tech week in San Francisco. What's going on? >> Oh, man. It's going to be quite a busy week. I mean, compared to last tech week, we are a part of a lot of events this time around, which is nice. it all started yesterday with the personal agent hackathon. probably one of the coolest hackathons I've been a part of. the production value was very high. it was at the Terara Center on Harrison Street here in SF. There was over 200 people on a Saturday Sunday morning 9:00 a.m.

5:14 There was so many people. I >> I saw there were 700 registered, which of course not everyone always shows up, right? >> Yeah. But 700 people registered is insane. And then 200 plus actually showing up on was it was it on Sunday? >> Sunday. Yesterday. >> Dang. >> And dude, it got so it was so there were so many people there that the Wi-Fi had problems. There just too many people connected to the same Wi-Fi. And then they had to turn people away. Like people who didn't register. It was just you know just to hit capacity.

5:47 They unlike a lot of hackathons, they did breakfast. It was like almost like a conference hackathon where they did like the breakfast, the coffee throughout the day, lunch, and happy hour. it was super cool. we sponsored with Neon. Other people there were Fly.io, IO kernel. freaking agent mail. code rabbit. exa >> was >> itz was Simon. Simon was there, right?

6:18 >> Simon was there. Assistant UI. >> Yeah. >> if I'm forgetting someone, I'm sorry. But, yeah, just a a full crowd. All the homies were there. >> Great prize pool, too. >> Yeah. So over like I think the final number was like $279,000 in value for the credits or credits and other things. And I think the number first place got like a 100k in value of the the prize. So I think maybe that's why a lot of people showed up, right?

6:47 It's just the the ticket number or the the value of the prizes was so high. >> you know, you get some data bucks from data bricks for coming, you know. it was great. So, that started off tech week. That was like hitting tech week off. Now, I'm going to share with y'all some other events. There's so many events this week, but these are the ones that we're going to be at. And let me just share what those look like, and it all starts today. so today we actually have two events going on at the same time. I cannot clone myself, so we'll have two different people. I'll be going to this one which is Ship It and Sip It Software Factory Night at the Corgi Cafe. it's a whole whole mess of software factory meetup type stuff. So if you're interested in that, please come by.

7:38 it's at the Corky Cafe. So if you've wanted an excuse to go there, then here it is. And they're going to have boba and So if you like that, definitely. And if you and if you're not in San Francisco and you want to try to participate in some of this stuff, you can just try Monster Mactory, you know, and do it in spirit. Just, you know, play around the software play around with software factories in spirit of not being able >> to be at tech week. I'm missing out this year. So, you know, I'll be getting some FOMO.

8:04 >> Yeah. And this one comes comes to you from Copilot Kit and Corgi Corgi Insurance. So in the same day, same time, same everything, we have agents in production, shared memory. So we're going to be talking about some things. We're going to be spoil we have some spoilers coming ahead. We've been running a lot of memory benchmarks. So we're going to be talking about that and our research engineer Amish will be representing us here with AWS and serial DB. Very cool location at the builder loft. so you know if you're interested in that please go then so that's only today right Tuesday we have a break thank god because Wednesday we have in it work OS's conference which is pretty much all day and so it's a conference for builders be a lot of AI engineering things there you'll see a lot of product releases I wasn't originally going to go to this but I want to talk to Michael about something so I'm going to go and see if I could just bother him while I'm there.

9:10 but you should see, you know, just like a look at the stacked lineup. You know, you got Aaron Levy, who I mean, I respect that guy a lot. Swixs will be there as of course as he always is. And, you know, Dex will be there, Nikita, you know, Claire Vo, like cool cool lineup of people. So, and you can just I think you can still sign up, but go for it if you can. And then that evening, if if it weren't events already, that evening, Code Rabbit is throwing factory reset. So there's a big theme this year on software factory. So this is with Code Rabbit on, you know, what are some hot takes on SDLC. So I'll be representing Monster there. so that's at Code Rabbit office. Really cool rooftop. You can see the Bay Bridge. if for anything, you know, you get some access to all these company offices during tech week just to see what people are doing. And I'm sure there'll be a lot of pizza for everyone every night. So, and then finally, rounding off our take on tech week, we are going to be doing an MCP for GTM workshop meetup type thing. that's with MRA, Elastic, and Arcade. So if you're MCP is back so if you want to come learn about that we'll be we'll be you know sharing stuff. This is at the elastic surge office very cool office again. So yeah, happy tech week to all those celebrate >> and I guess on that topic of MCP if you are looking to you know learn more about MCP see what it's you know what it's been up to. I am hosting a workshop with Alex from our team.

10:55 Thought maybe I should RSVP. I'm in. All right. I'm hosting a workshop basically how you can build a chat GPT plug-in which uses can use MCP apps. So we're going to be doing it from a master perspective but how you can you know build an MCP app deploy it you know what it might be useful for. We'll talk a lot about MCP the spec all that and it'll probably have a lot of overlap with your conversation that you're having that night as well. So they'll, you know, if you can't make it to tremendous overlap.

11:24 >> Yeah. If you can't make it to Abby sock on Thursday, you can get a glimpse of what it would have been like with a little le, you know, maybe not quite the same fire, a little different fire from the workshop. >> So do register for that. >> But I >> less cursing for sure. >> Yeah, probably. I think that's a safe bet. I I would I would agree. But why don't we go ahead and bring on our first guest, our only guest today.

11:48 So we're going to bring on Ally. We're gonna do security corner. Let's get into it. >> Let's do it. All right, we have Ali How here on the show for the first time in probably like six months. It feels like I don't know if that timeline's right, but it feels like it. How are you doing, Ally?

12:22 >> I am good. Yes, great to see you all. I know it's been a while. Thanks for having me back on. yeah, hope things are going super well. Looks like they are. So, super excited to chat. Security corner again. First time in a bit. >> Yeah. And for those of you that, you know, haven't been watching the show for a long time, we, you know, have occasionally done security corners where we bring up interesting security topics.

12:43 You know, technology is changing, writing code is changing, building applications is changing and security is at the forefront of a lot of it. And so we like to have conversations around what should be should we be thinking about in terms of security as you know the whole world and the technology under our feet is shifting. >> Yeah. And Alli is the founder of Security Corner for the show. >> Yeah. Yeah. We've had other we've had other we've had some other guests on, but Alli is the original the OG as some would say. but what have you been up to, Ally? It's been a while.

13:16 >> Yeah, it's been a while. Been busy. I was I was chuckling to myself when y'all were talking about tech weeks. I feel like every week for me is is tech week. I do a lot of events these days and I still run the Insecure Agents podcast. it's powered by Keycard. I also work for Keyard. I'm a member of technical staff there. We do runtime off for agents. so I help out with a lot of like Devril stuff, both like workshops and the podcast and lots of events we're doing. we're also on the software factory train and doing an event next week around AI NYC around software factories. We're calling it software factory like show and tell but like the secure version. So what does it take to run agents not as pets like your personal open cloth for example but as like agents as cattle where they can be spun up spun down run in the cloud serve multiple users once you started getting into the actual architecture for secure software factory which I'm sure as you know we know you've done amazing work in this area so we'd love to chat with you all about it too but yeah it's it's it's one of those easier said than done kind of adventures and everyone's trying it right now and excited to do demos and a panel discussion on Tuesday at AIE NYC next week. We're on like the side events page, so you can find it there. Yeah.

14:34 >> Very cool. How do you like how's the key card life? >> It's good. super excited about like the problem we're solving. it's really interesting because like once you start talking to a lot of like the enterprise folks they have agents that are making delegation hops between u multiple different agents different clouds different environments so like if one of like if in a large enterprise like if someone's like talking to cloud code and they're like oh cloud like you know like what meetings do I have for the day for example or you're going to talk to like the Google MCP server or something and maybe also just like a legacy app that helps like manage like events or something internal.

15:12 you know those two applications could be hosted on two different like clouds or different environments and that legacy system who knows what that runs on. And so navigating and having an stable agent identity across all of those different delegation tops and be able to enforce policy at every single step is something that like enterprises need to be able to roll out something like a software factory or a highly agentic workflow. And so we have this amazing like all these amazing like model advancements. And some people are like like oh my gosh we have like AGI.

15:41 Well, it's like that's great but without security without that like identity layer. It's really hard to harness the potential and the capability and the autonomy that these agents did have. So it's been really rewarding to bring that to folks that thought they couldn't have it. >> Dude, without security we're all dead in 10 years. So I mean it's like really >> I mean if you follow the pedum narrative Yeah. >> Yeah. What are your thoughts on Pedum?

16:05 I'm curious. >> Yeah, >> I don't think we're gonna die in 10 years. >> Yeah, we talked about it on the show. I mean, I think I think we're going to be okay. You know, I'm not I'm not at the labs. I'm not on the, you know, I don't see what they see, but >> it feels like maybe if if that's the immediate risk, we should just not do that. I I think the only ones we really have to worry about are the people with all the compute. So, as long as we can as long as they control themselves, we should be all right. But, you know, I think it's a little overblown is is my hot take.

16:39 >> Yeah. But I think a lot of enterprises are especially we talked to a bunch there. There's just three big topics that I'm seeing out in the world right now with regards to security. One is agent identity which you just mentioned that your you know key card does. Then there's this whole concept of agent access like are they accessing as you is this identity part of your identity? do they have their own service account? That type of stuff. And then lastly, is there like a vault, like a one password for agents? and I think with all those three things together, you maybe that's what you need. but you know, no one's agreed yet so far.

17:20 >> Yeah, there's a lot of situations where you want like sort of the intersection between the two permissions between the the agent and the user. And it's really interesting when you start thinking about like intent based off where like your agent should be have a design. We have this like reference architecture called like agent baseline. It's got six outcomes team should work towards and one of the first ones is like discover and it talks about like designing your agent and thinking through like what permissions like what's the goal like what's the purpose of it and so like if a user has an intent to do something completely different than was the intention of how the agent was designed or what purpose it's trying to solve then like that even that agent shouldn't follow that user's intent. So, it should be some sort of like a union between the two permission set, which is something that's possible if you're able to give your agent its own identity and not use like the user's identity.

18:03 >> Yeah. And I I think that's where a lot of people are struggling right now is is should it just be identified as me? It's working on my behalf. It should get my permissions or, you know, it should if it makes a wrong decision, it should fall back to me. or is the you know is this agent and this is becoming more so the case running relatively autonomously it's not necessarily tied to any one user on your team you know we have a number of agents in our Slack that you know I might have built it or Abby might have built it but everyone uses it and sometimes it works autonomously when triggers come in you know whose permissions should it be using and and why and and how do you validate you know just like you would with any new you know team member I might want to give this you know that team member permission, but certain things I would want them to like run by me before they actually did, right? Or they before they had access to. So I think there's like the real time like do you want to allow access and how long should it just be right now give them access to run this one command and then lock it down again they have to request every time or do you want to give them ongoing access to certain resources? So I think there's a lot of things that people are still trying to figure out, especially in the enterprise space where there's a lot of systems, a lot of different permission levels that you have to figure out.

19:20 >> 100%. And yeah, it's like those two different modes you talked about where it's like is it working on behalf of a user or is it if it's like doing a chron job in the background responding to an asynchronous event like it should be working as its own identity in itself and have a different set of permissions. Yeah, I have a good segue into, you know, Ally, you have a something you wanted to talk about today. I have a really good segue to that. because I was I was noticing from our homies at work OS, our homies at resend where the UI like most things are going headless and even most initialization of products are your agent is signing up for you, you're redeeming credits, whatever. But I don't know if anyone's really like putting the security in the in in the in the corner, let's say. but yeah, like why don't you give us your opinions there and you know, let's take this conversation that way.

20:13 >> Yeah. No, great segue. yeah, it's really cool that like most things seem to be coming headless, but then once things become headless, seems like chat becomes like the primary interface, which I don't know if that's like the best interface for all apps, which is like a whole another like thing that's up for debate, but when things become headless, you're losing the UI being sort of like part of like the security layer and now you're relying on the security of like MCP and agent identity.

20:44 authorization policy providence and approvals that effectively be become like the new interface. so having your agent security posture and thinking that through as like headless replaces some of these legacy applications and and frontends becomes like critically important and making sure like the security is in place but also just like from the usability standpoint like people want to use it. like I don't know that I'm always going to want to chat with my agent for example. so like what does the future look like and how do we make sure that yeah security is in the corner like you mentioned.

21:20 >> Yeah I I mean I I would say if you think of everything as headless right it ultimately you should have had the right security on your APIs the whole time. you should have had all this posture in place but so often especially when you're building apps you just rely on the fact that you know a user has to click this button to perform this action and yes it could be scripted and yes you might have some like rate limits but ultimately UI does obscure a lot of security issues I think for most people because a lot of sub security is just you know a bit of obscurity right and I think that's been a lot of people's posture for a long time but now agents will find things that a human probably never would even an attacking human, right, that was trying to be, you know, trying to do bad bad actions on your site. So, I do think that it's going to push people to think about you have to secure every aspect because you don't know what an agent might try to do and it's going to try to do it in ways that a human probably wouldn't unless they were being, you know, a bad actor.

22:23 And oftentimes these agents are just trying to accomplish a task. They're not trying to be a bad actor. they just see that, oh, this API is available. If it's a browser agent, I can actually just hit this API directly and get what I want done. So, I'm going to do that rather than go click this button. And so, I think you have to think about things from a different perspective. And you have to bake security in a bit more than you even not that you shouldn't have before, but I think people could get away with not doing it as much. And now it's harder. It's going to be harder for people to do that. Well, do people publish open a o open API specs now because they want agents to read it to then access their API? But before you would probably not release those types of things because your service area was just your app using it or the end user.

23:07 Maybe you had a CLI or whatever. But like now an agent can then crawl every API endpoint and try a bunch of different out and then see what happens. you know like it's not out of the the threat vector has now increased right and then if you don't put tool approvals or if you don't put any approvals on your agent then you know you got these anonymous hits to your your API and now you're screwed. >> Yes. Because it's like when you open everything up to be headless versus having a UI in the past like you assumed the people that were ask like accessing your UI were humans but now it's these people that are accessing the APIs are the agents and the humans. So what does that like look like in terms of like the permissions, the rate limiting? I know we just had like the resi incident recently where an instinct agent made 200 API requests per hour and actually got the users resi account banned. so how do we make sure our APIs are prepared for that? We might not have to have done that before.

24:05 but also how do we make sure that agents are going to be good thoughtful consumers of those downstream APIs? And like there's a whole another problem here of like, oh, the web just wasn't made for agents. so those issues would have to get fixed as well. >> I mean, this kind of brings up the question, do are websites and people who are building apps and APIs, are they going to have to enforce way more strict limits as far as just rate limits to prevent these, you know, and a human's probably not going to do that? Of course, you could script it, but that's kind of hard to do. But now imagine every human on the planet could write any script to do this to any website because that's basically what an agent is, right? It could it can script it itself. And so I'm I'm thinking like at minimum you're going to have to like everyone's going to have to have strict rate limits on their APIs where before that maybe wasn't a problem or as much of a problem.

24:57 >> Dude, websites are even at risk if you expose webmc or webmc right on your website and it's actually a prompt injection. So now like your agent isn't trying to do something benign. It's crawling web pages and then activates webmc and then you start screwing around. Like that's another threat vector that has been opened up as well. So even websites, not just APIs anymore. It could be like the MCP server that's exposed or the web MCP that's exposed on the website. which I think honestly doesn't it all come back down to permissions and actions and stuff?

25:33 >> 100%. Yes, it does. And yeah, thinking through what permissions each agent or user should have and then potentially having like an escalation like pattern where you said like having that approval in place like maybe you should have like more permissions or at least like we should know who the caller is so we can assign a certain like subset of permissions to them. >> But then that also makes me think like when you have no if you're only headless product let's say and you do need to do an approval where does that approval go?

26:04 Is that like an email or a text message? Is that an MFA like Google off a authenticator thing? Like you kind of like cut out the UI that was supposed to grant these things. I'm curious how people do it or will do it in the future. >> Yeah, that's a good point. It's like also with approvals like a set few is okay, but of course you hit that consent fatigue as a human pretty quickly. so how do we use security such that like we're not relying on the approvals in the in the loop to be like our primary like source of security. I think like part of that too is having the right eval and like audit log in place as well where you're able to see over time like exactly like what actions an agent took. So if you're going to do an investigation and see like okay, you know, I want to tweak my policies or make changes, maybe it's super permissive, maybe it's not like productive enough, like you can actually do that. and start to understand like some of these use cases as you gradually like work up that like autonomy curve. You might not, you know, want to open everything up like right away.

27:16 >> So maybe the suggestion is to not do things through APIs. >> True. this like MCP would be would be better. >> Yeah. And I do you know I want to go back to kind of something you said at the beginning though is you know I feel like if you move to MCP then you're moving to okay what's the interface to access it? Is it just a chat window? Is that a shitty UI or is that the right UI? And how do you how do you think because UI is still important for humans to understand. I I need to understand what my agent is doing. What is the best way to actually understand that? Is that some kind of MCP app? Is that some kind of generative UI system where it can it can generate the the application I'm using can generate the UI as my agent is interacting with it? I'm curious what and I know this is not necessarily 100% security but it is security adjacent and like what is that going to look like in your opinion? I know I have some thoughts, but >> yeah, the MCP apps one is kind of interesting. but that's just like one part of like the ecosystem. Like you might, you know, MCP app might be good, but I also have seen people like on Twitter online be like, "Oh, I'm just going to render like HTML dynamically wherever my users are and like generate the UI like on the fly." And I think we're just also in this age of like customizable software that everyone's like really excited about. so having your UI be unique to you or somehow like serving the user in that way. I think that's where it could be going. and I think like letting people have that customization, letting them be able to do what they want goes back to that permissions problem as well. As long as you trust the person's like or the agent's level of access and you also are able to like authorize it just in time.

29:01 You're also able to have like a tamperproof audit log that's outside of the agent's hands. We saw like the hugging face incident. They tamper with their transcripts of their chain of thought. so when you have all of that in place, it gets you more prepared to be able to go and experiment with whatever the right UI is, whatever the right consumption pattern is, you're able to trust that because those security foundations are in place. you're able, you're ready for whatever that future looks like. And I don't know what it looks like, but I'm really interested to see where this like where this heads. I think the benefit of using abstraction like MCP is like there's observability baked in and then you can add more primitives on top because it is like a standard for agents as opposed to a rest API existed before agents and now it's existed after agents. So then like the quality of or the standards of that API surface area is not necessarily all agentified or not which could be such a good use case for just putting everything through MCP because that will evolve with agent standards as opposed to like open API specs have existed for 20 years and we'll see if they will change and people will have to change but like when you do an MCP you are making a like an action towards an agent future.

30:16 So, and maybe I think that's probably the move. >> Yes, I agree with that. And MCP also creates like a policy enforcement point if you're using a gateway for example between you and the MCP be able helps you like filter down like the tools that are allowed and also be able to see exactly like what traffic's coming in and out. >> Yeah, absolutely. I I think some of the recent additions to the spec make it more palatable for everyone, you know, everyone for everyone that has an API should probably have an MCP if they want agents to interact with it. I think before when it had to be stateful or felt like it had to be more stateful, it it was it kind of turned some people off, but it it feels like today everything's going to be I I think agents probably drive more traffic than users to most sites and apps now, or it very likely will if they aren't already.

31:12 >> So having a more well- definfined MCP if you have an API seems like the right decision to make. >> Yes. And we also have ID Jag which is a new protocol that enables cross app access which is what enterprise managed o by quad provides and essentially allows you to log into an MCP server once and then you've got access to all of the other ones you're supposed to log into. So you don't have to like log into every single one. So it creates a better user experience.

31:43 but also lets your IT and governance people at your company be able to revoke your access like all at once to everything versus like before you used to have these like OOTH islands where you'd have like an OOTH token for every single service you would be aing into and then your IT governance folks would have absolutely no visibility or ability to revoke and so those were stolen that would create a whole like bunch of problems and incidents there. So yeah, ID JAG is a huge improvement for MCP just like for security and for usability.

32:15 >> How do folks Yeah, how do folks learn about IDJ? You know, I'm asking for a friend, you know. >> Yeah, I'm asking for myself. >> Yeah, I'm surprised I didn't get more coverage. there is an like a podcast I did over the summer with Carl McInness who's an identity expert. used to be like chief identity principal or something like that at Octa, but he's had an amazing super long impressive identity career. he helped author ID Jag. so if you go to insta agents.com there, if you scroll down long enough, you'll find one with Carl McInness and we did a podcast on that. But you can also Google Enterprise Vantage O quad and you can see that there and learn more about that, how it gets applied.

33:01 But yeah, it's episode 41 in regions. >> Hell yeah. I'm g have to watch that. >> All right, I'll pull it up here for those of you that maybe want to see it. This is the episode if you are curious and want to learn a little bit more about it. I do think I've also seen other types of MCP solutions like MCP governance solutions. I think that there's, you know, it's it's not finished. There's a lot of innovation that's going to happen here.

33:28 there's going to be some winners >> and I think that's an important thing for enterprises is to try to figure out how do they manage all these different connections that their not that their humans are just going to access but their agents are going to access as well. But I I heard you also have some things that some events that you you have coming up this week as well. We talked about at the beginning when you came on, but did you want to share some of the events that you potentially are going to be at or going to be hosting?

33:55 >> Yes. Yeah, thank you. yeah, we're doing a lot of amazing events next week. So, when you're done with all of the amazing MRA SF, software factories events and you want more next week, join us at AIE NYC on Tuesday for software factory show and tell. We'll have demos and a panel panel discussion about that. I'll drop the RCP in the chat, but yeah, that's event. I'm super excited. We'll have like the founder of AOTH and co-creator of OOTH 2.0, know Jack Hart will be there on our panel plus friends from Brain Trust and Socket and Docker. So super excited to talk Secure Software Factory that night. And then later that week two days later on October 15th we are running AOS night for the second time.

34:39 We ran AOS night the first night July 1st during AI World's Fair. we had an amazing turnout and we're going to get the gang back together and SF to talk about how AOTH, which is agent off, who's going to replace OOTH eventually, the progress that's made and one of those things that's made progress on is this new concept called AOTH budgets, which allows you to put a restraint on how many times a agent is allowed to call a downstream resource. So, we talked about that Resi incident of like how this agent's making, you know, 200 API requests an hour to resi's APIs.

35:12 Well, with AOT, you'd actually be able to limit your agent. So, you can still have the benefits of all the autonomy capability. Beat somebody else out to that restaurant reservation. but, you know, don't go too far. You get your account banned. >> AOTH. That's interesting. It doesn't roll off the tongue. I got to get used to it. AOTH, you know. >> Yes. Yeah. No, security folks aren't like, we're not the best at naming things. >> Yeah.

35:36 >> Yeah. There's another one called like ARM. I don't know if you've heard of it. It's like ARM. It's a framework for runtime authorization. And people used to pronounce that like a AARM. ARM. We weren't sure. So, yeah, it's a it's a community issue. >> Yeah. >> Yeah. Eventually the A's just combine. It's just off. >> This is off. >> Yeah. This is the off. >> It just becomes off. >> Off. >> Oh, dude. I actually like that better.

35:56 Off, >> you know. >> All right, Ally. We We appreciate you coming on and doing Security Corner for the first time in a while. We'll definitely have you back on when there's more, you know, security. There's always going to be security things to talk about, but go check out some of the events they have going on in New York or in San Francisco next week. how how do people follow you, Ally? What's the best way? I am on Twitter and my handle is VT Ahow. you put it in the chat. It's kind of like Well, yeah, VTA. And then Instagram.com, that's the podcast. that's a great way to follow. And then if you're looking for our events, luma.comkeycard, find all the events I help out with. But yes, that's the handle. So that's probably the best way. You'll see everything there.

36:43 >> Yeah. So follow Alli on X, Twitter, whatever we call it these days. And yeah, we'll we'll see you next time, Ally. Thanks for coming on the show. >> Awesome. Thank you very much. Bye. >> Dude, did you subscribe? >> Dude, I host the show. Did you subscribe? Did you subscribe? >> Subscribe to Agent Hour every Monday noon Pacific. Tune in Mondays make your day from the news to the song.

37:26 >> All right. >> Good seeing old friends. Yeah, it's nice to >> maybe we should go to that AOTH event next week when we're all together. >> We could Yeah, we we are having a mini offsite for a few folks from Mastro are going to be there. So, we'll have, you know, not a bunch of the team, but a few members of the team on site in San Francisco. We'll try to do a few things together. So, maybe yeah, maybe we should attend a meetup or an event.

37:51 >> And we have inerson show next week as well. So, >> yeah. And we speaking of that, we got to figure out what are we where are we gonna do it next week? We're g >> Should we do a code rabbit or victory? >> It's a holiday on Monday. Are they going to let us in? >> Oh, victory hall it is then. >> Yeah, I don't know. I mean code C code rabbit if you are listening can we use your studio on Monday even if no one else is there? Maybe. I don't know.

38:15 We'll see. >> but we will see. We'll probably have an event or or probably do the show on Monday still I think. I don't see why not. But it is a holiday. so maybe that means you can watch it while you're off work hopefully or maybe you can watch it later if you are not at your normal desk that you are multitasking and watching this and getting some work done or eating lunch or whatever you're doing.

38:37 But before we move on to the news, which we're getting to next, if you are just watching the show for the first time, you're tuning in, thanks for watching this, it's live. Go to YouTube, go to Spotify. It's, you know, available on all the podcast networks you normally listen to, but live is, you know, the best way because you can drop chat messages in, you know. So, if you have a comment along the way, especially as we get into the news, let us know what you think.

39:06 All right. Well, should we do the news? >> Let's do it. >> Let's do the news. you're tuning in for the f if you're just tuning in. I'm back. I was I was muted. We'll cut that out. Hey everybody, welcome to Agents Hour. I'm Shane here with Obby. It's October 5th.

39:42 We're doing the news. We do this every week. Let's see a preview of what we're going to cover. We have a bunch to cover today. There's a lot of little topics, a few big ones. So, we're going to fly through it pretty quickly. The first is Anthropic released or at least had leaked that they were had filed for their IPO, I think a while back at this point, but now the the prospectus of it kind of leaked out over this last week. and they talked about their revenue. They talked about their operating losses and I think some people are given are giving them pretty hard times because they're going targeting a very high valuation and their revenue isn't even close to what some other tech companies have >> with you know I think they're like >> targeting metal level valuations with a tenth of metal you know level revenue or something like that right but >> yeah what do you think >> I think you know every the multiples are crazy now. But I think a lot of people are giving them flack, too, because of where the source of revenue is coming from. It's not really distributed, right? You have outliers that are making up a ton of your revenue. And if you ever lose those to Muse or OpenAI, then you'd be screwed, right? so it's all based on if they can continue the trajectory of getting whales to to, you know, to keep the revenue high. I don't know. I I think it's going to be hard. but you know what? The lame in person in the market will probably buy on IPO day for a premium like >> crazy premium.

41:22 >> Yeah, I can't imagine they get to two trillion. That seems I mean that seems wild. I I do know if they can show you know the growth rate though is tremendous, right? That's what people buy on is that >> it is a tremendous growth rate never before seen. And so if it continues, it becomes, you know, potentially, you know, you could argue potentially becomes the biggest company in the world, right? Or very close. Now, I think is that going to happen? I think it's still a long shot, but that's what people potentially buy on. And we will see, you know, I don't think there's any more news on when the date would be, whether it gets you held back. I've I've heard rumblings that they might push it back a little bit. I think all these new model launches actually hurt them. This dumerism I think probably doesn't help them, but it's coming from Daario, so maybe he thinks it does. I don't know.

42:10 It's It's like a bit of a game. >> I don't know what game they're playing. >> But I do think that they they have a lot to show to hit that valuation. It's going to be very hard. It's it's an uphill battle for them to actually reach that valuation. But if people buy, you know, that's all you need. >> Well, not financial advice. don't buy at IPOs. just that's my rule of thumb personally because market will correct you know you're operating in a private market where your valuation it's like whose line is it anyway all the numbers don't matter and stuff but then once you get into actual public market you'll get screwed on that IPO price so you know if you're really an investor's weight personally >> not financial advice >> yeah the only thing I didn't get burned on is that Cerebrris IPO so what's up but other than that you You didn't follow your own advice is what you're saying.

43:04 >> Yeah, but I was screwed for a while though. So then, you know, finally it came back. So, >> all right, let's talk about decision models. This is a new genre. We've been talking about it on and off, basically on for the last couple weeks. This was at OpenAI Dev Day, which we're going to talk about OpenAI dev day a little bit on this show. We did a whole episode on it last week. So if you want to know, you know, deeper, you know, what we what we think about that what they launched at OpenAI dev day, go back, watch that episode. It's on YouTube, Spotify, all the places. But OpenAI released the decisions API for lightning fast constrained decision-making powered by Luna. Supports visual inputs and tuned to be able to make decisions in less than a few hundreds of milliseconds end to end.

43:48 to which the founder of Jev said begun the Clone Wars has because if you've noticed there's a ton of now they've Jeb basically created a new category. >> Yep. >> On the one hand that's good. On the other hand, you know, we'll see if Jev makes it out of their own category that they defined. >> Yeah. And you know it's interesting that they it's not only Luna as like you know a decision model but you know a lot of people are training a Quen model. So they're trying to get this similar concept going. yeah, >> Open Router now has an entire decisions model category in, you know, in open router. So they're building out this Yeah, these options. So if you wanted what they're calling decisions, you know, we've heard people call them like we've called the classifier models, people calling decision models. I I do know the creator of Jev says decision model or as a using decisions as as the only classifier for it is a little shortsighted, I think. But we'll see what they have cooking. They might have more things coming or maybe they don't.

44:54 I don't know. >> But then how do people understand what a system one is, right? Exactly. You need some word for it. Yeah, they call it a system one model which just means if if you are just hearing system one model for the first time, all it means is you don't actually think, you just act on like simple decisions. >> So that's why it's good for simple things. If you need it to like reason about and make a decision, it's actually not not good at that. So you got to break down a complex problem into simple decisions and then these models are good at making very quick fast decisions on without having to do a lot of deep thinking.

45:30 Cloud fair cloudflare released cleft and it's you know they say today we're releasing two fast and accurate decision models that top the benchmarks for quality and latency use them hosted on workers AI or grab their weights from hugging face. So this was on October 1st. So Cloudflare is getting in their own decision models. They're getting in the game. >> so SG lang also has decision models which they you know you mentioned like the Quen models. That's what they have available. So you can again answer quick yes no questions get probability. It's just more proof that everyone's building decision models into their products, gateways, frameworks.

46:13 There was something called glide was the first decision decision model that thinks I guess you know which is interesting. So rank >> system 1.5. >> What's that? >> System 1.5. >> Yeah. Yeah, it's a system 1.5 model. This one is a decision model that can think also known as just a normal model, but I guess it it is fast, right? Smaller. It beats Jev by 6.9 points. So it uses adaptive thinking. So it calculates, you know, fast probability distribution. If it's uncertain, then it does some reasoning. So again, it you the the founder of Jev says he doesn't like benchmarks, but people are going to benchmark and figure out where where does Jev fall short, where do these other models potentially beat it, and then that's how people are going to make decisions on what models to use.

47:00 >> Oh, there's a couple before we move on to the next thing, we have to mention the homies at Respan have a decision model, span one. >> Yeah. >> And then Devon made Kevin or Kev which was trained on Quen as well. So this space is heating up for sure. >> Yeah, it does feel like the one thing I will say about this is I think it's a good thing for the industry because for so long there was this push to get everything into your skill file, your prompt and just let one big model do it all. And I think what we've seen and what we've probably known, but we've kind of had to push against the narrative a bit or people have pushed against the narrative is that if you can break down a complex problem into more deterministic chunks, it's you're going to have better results long term, which is for if you're using MRA, it's workflows, right? If you can break something down into a workflow rather than just like a prompt, that's probably going to give you better results. Now, sometimes a skill file is is fine. Just put it in a skill, let it do its thing.

48:03 it's going to be accurate 98% of the time and it's good enough. But if you need it to be accurate 99.9% of the time or whatever, a more, you know, directional workflow with maybe certain decision points is actually the better way to architect something like that. If you can define the process and you need to make smart decisions throughout the process, but you can break down those things to where you know a decision model or a system one model could could decide between, you're probably going to get better, more reliable results at least.

48:31 >> Yeah. Got a question in the chat. Are we getting a rename for classifiers in Maestra? Maybe. I mean, AISDK called the model evaluation model. So, like we're all mixed up on names here, but classifiers in MRA are primitive. So, I don't know if we'd call them like decision or something, right? >> Yeah. >> But maybe the model parameter becomes a decision model versus evaluation model, which it is today. Yeah, we will that this, you know, it's one of the things about being early is the space has to play out a bit and we have to see what do you know what do users expect it to be called? What's the >> word classifier? That's what I like. But all the all the old heads are like, oh, we've had classifiers forever. It's like, man, shut up, dude. It's just a classifier.

49:21 >> All right, we're going to breeze through OpenAI dev day because again, we did a whole session on a whole live stream on it last week. Go check it out. It's been launched. It's on YouTube, all the places. OpenAI introduced DOTS. There was a pretty big demo failure because of some Wi-Fi issues. So, you know, kudos to you the, you know, the individual that did the demo. She did a great job of working through it. But dots are, you know, just like Grockbot, like Muse, there's the you know, it's very much a trend as well.

49:52 They re released GPT61 soul and they introduced plan a new plan a more expensive plan and then they also made changes to the two to the $200 plan which basically cut your limits a bit. They one other thing that is not on here is they also change kind of improved the chat GBT extensions right so you can build and kind of extend chat GBT almost like you're extending VS code or you know you basically build a plugin you can extend the UI you can make changes to the app and you can use codeex or chatgbt right on top of it so that was also a cool launch anything else any other things we should highlight besides telling people just go watch the full episode.

50:37 >> Yeah, go watch the full episode. I did play with dot and yeah, it's pretty cool. I don't think I'll ever use it after playing with it, though. That's just me personally, though. I'm just going to stick to my one chat GPT thread that I've had for two years. Just going to keep That's my dot right there. >> I I do if you have used dot and you use Grockbot or Muse, I'd love to know how they compare. We have people on the team that have used, you know, some of these different tools. So, I'll be anxious to see how it levels off where, you know, where people are like migrating to after they've had some time to settle, they've tried out the different options. I do think, you know, time will tell. I'm sure all these teams are shipping fixes to make their experience better, but it there is going to be something where like personal agents are going to be figured out. These lab, you know, these labs or these large tech companies are all competing for that same space.

51:34 And Hashim in the chat says, " dot manages my other threads for me." So, using Dot as your supervisor or your manager. Okay. >> I just have one thread, man. Me and my homie in one thread. >> OpenAI did release a new Dots demo after the first one on stage didn't go well. So, if you want to see what it is, you can go check it out. This was from October 1st. >> This got so much controversy, though, which is not fair. but I think the industry as a whole is sick of seeing booking flights as demos.

52:08 and then there's a lot of people like arguing like what what airport you should be going to, whether if you're trying to go to Silver Lake, should you go to Burbank or LAX? Like that's like some minutia BS. I think the bigger point that was made was why are we always doing simple things for these demos like booking a flight and then in the chat or in the threads just like yo I want I want to see dot do my taxes like I want to see dot do some real stuff. so I think that's the bigger point.

52:39 >> Yeah. A lot of people also hated on it because it took quite a bit of explaining in and you know of how to actually do it the first time. And I would just you know encourage you to say like yes you have probably have to explain yourself at least once. You know if me and Abby are working on something together and Obby works a certain way he has to explain it to me once but then hopefully he never has to tell me again or he doesn't have to tell me every time. So you know I would give you know any demo like that. Imagine you could explain it once and then it knew your preferences. a new, you know, I like window seats, right? A new I want to sit by the window. So, I think, you know, it'll be figured out. I don't know if it's going to be OpenAI or Muse or, you know, Grockbot that that does it best. I do agree though. Needs to be better.

53:24 >> Yeah. >> Better demos. >> Needs to be better, dude. Also, the funniest roast on this post, which is so funny, the camera was like shaky like this, like like this all the time. And then someone said like, "Oh, yeah. It's like watching the Blair Witch Project." >> Oh, I mean, >> internet never ceases to amaze me. >> Why would they Why would they have Shaky Cam on when they like the first demo didn't go well? You think the second one they'd want to make it like as professional as possible, not like there's an indie, you know, camera crew holding the camera? All right. Sam Alman came out on October 2nd and said, "There's some speculation about our partnership with Cerebris. Cerebrus is a close partner and we have a deep engagement pushing on the frontiers of speed." That was the entire tweet. That was it. So then >> not financial advice, but >> yeah, but that led me to think, okay, what's going on here? So I did a little digging and apparently it's because 61 soul is you know using Nvidia GPUs rather than Cerebrus for inference or something like that. And so it caused a lot of speculation of okay why are they not using Cerebrus for this? How are they you know are they going back on that? Is that are they not caring about that partnership anymore? But it sounds you know according to Sam he's trying to assuage the fears that they're not going to use Cerebris. Maybe they just haven't yet. Maybe they will later. I I don't know. But I thought that was interesting that the newer models are just using Nvidia GPUs.

54:54 >> I mean, stock the stock went up since this tweet. It was trending down. The tweet comes up and now the stock's trending back up in a month overview. Not not financial advice, though. >> All right. So, probably the biggest news from last week, and we kind of hit it here in the middle, is Google dropped a new model, Gemini 4 Argon. And this is from September 30th. It says, "Introducing Gemini 4 Argon, our new frontier model. It's built for complex workflows across coding, enterprise knowledge work, and cyber security defense. Rolling out today to a set of trusted testers through our Fairwind program. So again, not rolled out to everyone yet, but if you look at the benchmarks, they're very good.

55:41 >> Very good. So we can look at, you know, I'll pull try to pull it up here a little bit bigger. Why don't show the full screen so we can look at some of the benchmarks. You can see valves index, automation bench, deep suite, terminal bench 4. and it doesn't win on everything, but for the longest time, you know, Gemini was just releasing new flash models because it couldn't really compete on the Frontier or at least they didn't want to talk about their benchmarks. So, their benchmarks always seemed to perform pretty poorly. I think this was the first Gemini model in a long time that caused people to stop and say, "Damn, they might actually be back in this race."

56:31 Now, the caveat here is I haven't tried the model, Obby. I don't think you've tried the model yet, have you? >> We're not in the Fair Wind program. >> Yeah. Not yet. >> Not yet. >> But we have not tried the model. And I'm sure most of you listening have also not tried the model. So, do you know, are the benchmarks true? Does it actually live up to the hype? That's to be determined. But if it does, then I do think Google has kind of reinserted itself back into the frontier conversation at least.

56:59 >> Yeah. And I'm curious how they'll position it marketing wise and you know will they talk about Muse and like pay pay back on some of the talk from Muse that if everything is good Google is in a really interesting position to like capitalize on on the the model itself. >> Yeah. Yeah, we already talked about this in the past when we were kind of hating on their the fact that they didn't perform well in the benchmarks, but that they have such huge consumer adoption, right? There's no company better positioned. I mean I I would say Facebook is is pretty well positioned with Muse but as far as like getting consumer market adoption you know they way way better positioned than X I think because everyone's using you know Gmail and and Google and you have access to all these things if you could build especially for the the workplace like a personal assistant for an organization and you tie it into G Suite and and all the tools that work workplaces use you could probably win that space you could win the consumer space. I think that if they can be competitive on the model layer, they can win a lot of other places.

58:09 >> Yeah. I mean, we've been talking so much about Gemini. I want them to have a win. >> Yeah. I think it I saw some kind of prediction that now at the end of the month, they think that Google will still will have the best model. It's like flipped. It went from like 1% to like 60%. >> And which is, you know, I wouldn't even that's an insane jump. Sebastian in the chat says, "Let's hope they don't make it exclusive to their CLI or anti-gravity."

58:34 I hope not. >> We need to start putting some of these bets on Khi. >> That's how we're going to make our money. >> That's how we'll really make money. >> so Val's AI released that Gemini was on top of the Val's index. So you can see the accuracy rating compared to beating Sonnet 55, Opus 55, you know, GPT6 Astra, GPT61 Soul. So, it's beating those on the valves index. And then someone came out and said, "Wait, did I read that right? Gemini 4 argon has a 1 million token output limit." So, not a 1 million token context window, but an output limit. So, a lot of frontier models cap it at, you know, like in the hundreds to to 200 range for like output tokens, right? So, it's essentially like five to seven times larger output. You could basically write entire books almost in one shot, which is insane.

59:30 And then I thought this response was very funny. They named it that because after four prompts, all your tokens are gone in that because you can just output a ton of tokens. You're going to spend a lot of money potentially. >> So, it'll be interesting to get access to it to play around with it and to see how good it does on various tasks. Of course, everyone's going to test it with coding, I think, right? That's the first task we probably all do if you're watching this show. But I would love to test on some of the other tasks as well.

59:55 Just, you know, writing, things like that. Any comments on closing comments on Gemini? >> I'm stoked. You know, it'd be nice if Google's back in the game. I hope it's not a flop. I really hope. >> We will see. So, let's talk about the White House AI lunchon. I don't know if it's a summit, a lunchon. I don't know what they're calling it, but this is a post from David Saxs. But to set the stage, President Trump got together a bunch of leaders in AI for this kind of lunch dinner thing and they came together and came up with kind of some resolutions that they're all going to like hold themselves to. They signed it. So, we'll read through what it exactly entails here in a bit, but I think there's some critiques of it that it was just like all the, you know, net worth billionaires in one place like agreeing, you know, kind of backslapping and agreeing that they're going to do this thing whether they do it or not.

60:50 there's really no like guarantee that they can do it, but I think it's a step in the right direction. What were your thoughts on it? And then I'll read through what they actually, you know, signed here in a second. >> I mean, I think it was cool. given the doom week we had and people saying different things and the president also like disagreeing. So, it was nice to see them all together. The amount of memes that were created from the pictures is comical. It's like pure comedy. so I think that that blessed it blessed us with that too. So two good things happened.

61:24 >> And and like if you saw the video at the press conference with Daario in the background and he's just like moving around in the video, you're just like, "What are you doing, dude?" >> Someone comparing comparing Daario to Mr. Bean Bean in the background. >> I thought that was pretty funny, >> you know, like Dar's just in the background of this. And I but I do think one thing I will give the president credit on is I don't think you know you could tell like Daario would not traditionally like President Trump just based on politics I'm imagining but being able to get them in the same room.

61:57 President Trump said a lot of nice things about Daario. I think Daario kind of said some reasonable things back like night he he I think they they played nice and they now it's hopefully are getting some of the resolution that I think Daario wanted in the first place was which was at least like this audit >> this like at least getting all the labs to agree to this audit. So I feel like Daario hopefully can hang his hat on that he made some progress in that which sounds like it was always very important to Daario but I feel like everyone else was kind of in agreement and then you know Sam Alman wasn't there I think because of OpenAI dev day so I think Greg Brockman was there >> but I I feel like it was like this whole thing was kind of targeted towards like Daario being a bit more of the doomer narrative and like Sam was a little bit but then the other labs are a bit less doomerism coming out of them.

62:47 So, I feel like it was very much like how do we get everyone to agree that that we're going to take ownership, but we're going to also play nice and we're going to have kind of independent audits. So, if we read what it says, it says White House Accord on super intelligence, joint commitment on frontier responsibilities in order to build a positive future for the American people in the world. We believe every company is responsible for developing its own technology safely and in a way that builds trust with customers and the public. starts with every company that is training and deploying frontier models having robust internal processes and controls to ensure that their technology behaves as intended and that any issues are promptly identified and resolved. Therefore, in addition to any other precautions, we believe each company should implement the following four layers of controls and audits. So, these are like the four things that people are agreeing on. robust internal controls to monitor the capabilities and alignment of its models during training and deployment around areas like cyber security, biocurity, and chemical threats and to ensure that its models do not hack or access technical systems in unintended ways. Two, empower an internal team to ensure all the controls, monitoring, and detection are operating as intended and that any issues are remediated. Three, partner with an independent external auditor or evaluator to carry out independent assessments of whether the controls, monitoring, and detection are operating as intended. So that was a big one is like you have an external party that has access to validate and audit. And then this is the fourth one that I think makes it potentially give it a little bit of teeth where you can hold people accountable, but designate an independent committee of the board of directors to oversee and receive reports from the teams operating the controls and the internal and external auditors and evaluators as well as to ensure any issues identified are remediated. So essentially getting it to like board level visibility that if there's issues the board has to at least sign off that they're going to release this model knowing that there's these potential issues right so I think it puts more responsibility on the board and which means there's you know potential for them you know lawsuits and all those things right that they have there's a lot of like responsibility that comes with that and so >> potential for them to disagree on release >> yeah exactly so I do think it the goal is like hopefully it's enough to slow the teams down from shipping something they shouldn't have, but not so heavy-handed that there's this huge process that needs to be in place where people we have to slow down so much that we can't get new models out the door.

65:18 >> Yeah, dude. Startup idea model exaluator. >> Yeah, I mean there's got there's gonna be a lot of those, right? >> Make some money, y'all. >> All right, more drama. Factory verse cognition. So for those of you obviously probably have heard of Devon, you've heard of factory droid. This is from Matan from factory says, "We're terminating Chris Dgnen for unethical conduct involving cognition. The last few months have seen incredible progress in AI capabilities. San Francisco has flourished as new companies that solve new, more ambitious problems are finding great success.

65:57 Generally, it's a wonderful time to be building." So basically goes on, I'm not going to read the whole post, but person that has been like an adviser to factory was also starting to talk to Cognition and apparently got hired by Cognition without Factory knowing and they're obviously very direct competitors. >> I think it was after he got removed then he joined. >> Yeah. But then there's speculation that he wasn't actually removed, he just resigned and now but Matan saying he was removed. But this also came out, you know, the or basically the same time after.

66:33 >> Yeah. Like a half hour or hour after >> said, "I'm super excited to join excited to join Cognition as its chief revenue officer. Going from the first sales rep at Snowflake to CRO at a hundred billion dollar public company was a thrill of lifetime. Basically wants to do it again. Excited to join Cognition, loves its growth, all those things. So the issue is should he if he was an adviser to factory you know at like board level advisory should he be able to should he have done this? Was was this like clearly factory thinks he was doing something wrong? He thinks he you know had disclaimed everything he had not shared any secrets with cognition but of course factory is accusing him of sharing sharing trade secrets of what they're doing what they're building. And then this is where it gets really wild because that's like kind of interesting from the outside perspective. Then you have an investor in both companies. So this is Fenad Kosla from Kla Ventures who has investments in both factory and Cognition picking a favorite here. This is like picking a favorite between your two kids publicly. It says you are this is targeted at factory was a response to their message. You are a struggling second tier competitor that is more unethical in lying just because you have no decency or sense of proper behavior and shows your desperation. Straight out lying about if Chris being fired. I thought it would be below even you.

68:00 >> Shots fired. >> This was shots fired. And then of course then everyone saying you should never raise money from Kla because they'll literally tear into you in public >> and just a lot of drama on the timeline. And then Sean Magcguire says, "What Cognition doesn't understand is that we're sitting on a nuclear weapon to end them. Hopefully, we don't need to use it." >> Okay. >> Cause so much more drama. >> So, so much drama on the timeline if you are paying attention to what's going on in the AI world and as as far as like coding agents.

68:33 >> Yeah. You know, it's true though like people in the valley, they move differently based on who they've invested in. We've met people who've invested in our competitors and treat us different ways because you know their allegiance lines which kind of whack right as like a human that's super whack cuz maybe you are indifferent. You want to be friends with homies and like in every company but typically that doesn't necessarily happen. You know just your your your investment is your alignment.

69:02 I don't act that way so y'all can be friends with me but yeah it's just whack. We've seen it happen. This is another example of that happening. >> And it's why you'll often hear us say, you know, talk about this as like the Game of Thrones because it sometimes feels like that, right? Allegiances. >> The red wedding. >> Yeah. Yeah. There's allegiances and there's, you know, falling outs and it's all happening in public and, you know, X is free, you know, or you can pay, but X is free for the most part, right? So, you can view all this, grab your popcorn, and read the thread like, >> you know, like we're seeing in the chat.

69:36 >> Yeah. Yeah, like when you're a smaller stage like seed round a, you know, you don't really see this you don't really see it much because at at that stage investors are making multiple bets and you know, unless they're like taking big swings at certain companies. but once you get into serious rounds like these companies in particular have raised hundreds of millions of dollars now and so things get a little Game of Thronesy. >> Yeah. Yeah. They're they're thrown around big numbers and they want to protect their investments. Yeah.

70:07 >> All right. Let's talk about open source. You know, master's open source. We love open source, but there's this trend of open source closing the door. So, this is from Cynus. Says, "Due to AI, what's that? >> This guy's a legend, by the way. >> due to AI, I've disabled external pull requests on all my repos. Open source as we have known it was fun while it lasted. It's I will still maintain projects and handle issues.

70:37 So, what are some of the open source? >> PMAP, my favorite freaking thing. Execa, my other favorite thing. And the list goes on. Cyn has done so much for TypeScript and open source in general. but those are like my two favorite, but he's done a bunch of stuff. >> Then we also see Yuki Wada says, "Sad news. We disabled PRs from external contributors on Hono.js. Hono has not stood here without PRs. I will never forget the PR you created for Rejax router. A damn fast HTTP router we have never seen. But PRs don't work in this era. Contribute in other ways. Thanks.

71:17 >> Use copied us. But no, I'm just kidding. yeah, just another example. >> Yeah. So yeah, I guess we can say we concur, right? In some ways. I don't believe open source is dead but I do believe and we've said this the contribution model to open source is changing. I do believe that quality issues with you know ideally reproductions and you know enough details without like trying to enforce a specific decision is the right way and then that way the team that is you know behind the open source project can use their models with their context to actually fix them in their way. so I think issues are the new PRs, but I do believe that people should get contribution credit for issues.

72:07 >> Yeah. And both of these authors have in their posts, they're pretty much saying the the problem is your agent opens up a slot PR and then the maintainer is using an agent to review it and that's going back and forth of models of unknown quality on the other end. and you know, the question you ask yourself is why does I could just do all this stuff myself and it'd be faster, less less lead time, etc. We've been talking about this for weeks, though. Just want to say we've been talking about this stuff for weeks. We have not closed PRs, but we might, right? We might do this, too.

72:44 We've been not wanting to and hoping that the community has discipline to listen to our wishes. >> For the most part, they have. We we do let PR let them open PRs if they have issues, right? Like there is some requirements. You have to have an issue. It has to be triaged as like a valid issue. >> Yeah. >> But we typically don't, you know, I would much prefer you just open the issue and and make it detailed. Follow up if we have requests or questions about it and then >> we will prioritize it and get to it. And it's probably going to get done faster than if you give us especially if it's a big PR >> because you got to review it. It takes more back and forth time. You got to get through our CI, right? There's a ton of CI.

73:25 >> Good luck >> with approve view to run it. It's like, no. >> Yeah, you're you're better off just opening an issue and making a detailed. The more detailed the better with, you know, reproduction and that's going to get it fixed much faster. >> Like we probably should close PRs. I just feel like that is an extreme thing. Ideally, and it's been working so far. community knows that hey we'd rather have issues and we do say when you when you do open a PR for an approved issue oftentimes I close it and say hey factories is about to handle this anyway so don't worry about it then people got mad at us you know oh you know you're not letting us contribute but the issue was the contribution all along so >> yeah I think that's as long as we can change the model where we value issues as much as we used to value PRs especially good issues I People should still feel good that you're contributing to open source >> but has changed a little bit.

74:21 >> Dude, you know what I'm seeing? Last point on this, you know, people think I'm seeing this a lot lately where we get slop reviews. So instead of contributing an issue or a PR, people are sending their PR or agents to review code, which is even the worst thing that you could do. You know what I mean? So if you're doing that, don't do it. All right, so let's go through the quick hits. We're going to go through these rapid fire. Terso is joining Superbase.

74:51 So Superbase raised a pretty good sized round. They acquired Terso. So if you've used Terso or you you know you now are going to be using Superbase, I guess, because they're going to be they're going to keep it going but within Superbase. It sounds like >> many people have used Terso because of Maestro. >> Yeah, Maestro uses Terso. >> The homies. >> Yeah, they are the homies. So congrats. so this is from co Ty from co-pilot kit says introducing open dots self-hostable always on AI co-workers that works with any agent harness includes computer use browser terminal bring agents to slack teams spaces pages for project voice calls web and mobile.

75:29 So if you want to try to build your own open or dot rather than using open AIS you want to use open dots that's cool. Congrats. so Charlie, friend of the show, been on a few times. Introducing Conduct Conductor Mobile. Run a team of cloud agents from your iPhone live now in the app store. So you can code from the beach. Kua says, "We're reimagining what it means for your agent to work with all your computers." So they're excited to share KUA spaces built on KUA driver and our virtualization stack rolling out to Mac OS today. It's free and open free and source available so you can control your desktops from anywhere.

76:13 Yeah, I think what do what do you think of computer use in general? >> I think it's the next frontier, let's say. >> You think so? I mean, I I do think that computer use is really cool when I can watch the agent control my computer and so I would know how to do it or I can like interact with I feel like that's one way to even like learn how to use an application. It's like just to watch the agent do it.

76:40 >> Yeah. >> But you know all the new models are the new models are like saying specifically that their computer use is a is better right. So this is all like a natural progression of the frontier. Kua friend of ours YC like they're doing it open source. So >> yeah and related to the issues conversation we got some chat. Sebastian says, "As long as it's clearly communicated so people don't get the wrong expectations, everything's fair game." Yep.

77:09 >> this individual says, "Props so far with how fast you close issues. I've had five plus so far." >> You're welcome. Thank the factory. >> Yeah, factories making moves. All right. introducing ETE, the agentic testing framework for any app. NPX ET in it. So it's again a way to better write in into intest for your application. Let's talk about this Andre Carpathy post. I'm going to share the whole full thing. This one's going to take we'll go through it really quick, but it is pretty interesting. Carpathy's kind of been silent for a while. So >> he got neutered.

77:47 >> Yeah. But but he's he's back. He at least was back to post this. I think there are a few good tricks in here though. So we'll be spending a lot more time trying to understand the outputs of language models. a few thoughts, tips, and tricks. So, it says, "Ask your LLM to explain something in ASD ST 100." Okay, if you wanted to have better writing, use diagrams in images instead of writing, ask your LLM to create a diagram. I think, you know, I I've do that quite a bit. I think, you know, a lot of us probably do, but that's a good tip. Ask for output in HTML, so you get rather than just text. this is the one that resonates with me because this is what I want. explainer videos. The output format I am most bullish on is fully custom bespoke explainer videos generated on any arbitrary topic.

78:34 And I here's why I think this is interesting because if my agent would know my preferences, it could make a video based on my learning style. So I I go back and someone's probably built this. I haven't seen any good examples, but sometimes I want like I'm looking at this PR. It's a lot of code. It's 400 changed files. I've read I have this wall of text to read of like explaining what it is because my agent condensed it. But if you could just explain it to me like a you know on a two-minute YouTube video in terms that I understand because you know the parts of the codebase I've touched that would be pretty cool. You could explain it to me and show how to interact with the things I already know about. So I think like there's room for personalized explainer videos as these models get better, which is, you know, I'm using code as an example, but I think there's a lot of examples of like I want to learn this topic. I have this specific, you know, camera. I want to learn how to do it.

79:27 Build me an explainer video and show me how to how I can actually like use this camera to do this one task. And it can customize that video specifically for you. I don't know, things like that. Cool. All right, let's continue on. You can now mod Claude code, change how it behaves, customize the UI, swap in your own features, write one with a few lines of TypeScript, or even have Claude build it for you. Mods ship inside plugins, so you install them with slashplugin in the CLI or desktop app.

80:02 People like mods. Did you see this thing called Griffin? >> Yeah, dude. All right. So, introducing Griffin, the first model to pass the video touring test. 48% of people who talked to it live thought it was a real human. >> This is wild. >> Yeah. I mean, if you go ahead, we're not going to do this now. We're running low on time. Watch the video. It's pretty wild. you know, it's it's not so much better than what was there before.

80:30 Like, I feel like I can tell, but only because I know now. If I was in the moment, maybe I wouldn't have been able to tell. The latency is pretty good. It looks pretty real. Like you I know I it it's pretty hard to pick out. If you weren't paying attention and you just had a conversation, you would probably think it was real unless someone told you ahead of time. >> I mean, if watch the video because in our YC batch, we had a company called Pickle that was trying to do this in Zoom meetings and I thought it was pretty impressive back then.

81:00 >> But the Yeah. The thing about Pickle though is it was kind of like it was still your voice, right? Like you would talk, but it would just be your avatar matched your voice. This is like 100% synthetic. Yeah. Which is a which is a different level of game. >> Yeah, dude. That's the next level. Not attending standup in the morning and staying in bed. >> Yeah. >> so 11 Labs on September 28th said, "Introducing 11v4 and 11v4 Turbo, our fastest and most emotive voice models yet." You can watch the video. They're very good. They've always had really good voice models. I think they're getting even better. I saw some comments after both this and both the Griffin and 11 Labs is like you should be now is the time where you talk to your parents and your grandparents around like having a passphrase or knowing knowing that it might not be you.

81:51 >> I I do worry about that like you know our voices are out there on the internet. If you've ever recorded anything live, it's out there. Like it doesn't take a lot to like clone someone's voice. Clone someone appear someone's appearance and you can definitely like scam people. So >> password. >> Yeah. Yeah. Have a passphrase with your family. Be a recommendation. All right. this post came out from Killian says, "Can LM discover and fix bugs if you don't tell them what went wrong? To proactively maintain a repo, agents need to find problems before users do. So, S swep benchmarks this on 100 repos, 22 languages, 4,000 real bugs. But anyways, that's a a way to just like scan your codebase for bugs. What do you about what do you think about this as a topic?

82:38 Like using some kind of external tool to scan scan your codebase, try to find bugs proactively before users even do. I think it's amazing and we are starting to use one from our homie Dan Robinson. Detail. If you guys want to try Detail, go to detail. Actually, let me pull it up. Detail. It can run a scan and on your whole codebase and then start opening issues that with the proposed fixes. So, I think I'm very bullish on this because we ship a lot. We also need to fix a lot of bugs. So it'd be nice to have like early detection on a lot.

83:18 >> Yeah, it's like proactive monitoring. It's like pre-monitoring before users see it and then you can evaluate which ones are going to be the biggest impact, which ones users are most likely to hit and which ones are just like edge cases that no one will probably ever see. >> Yeah. >> So yeah, I think it as a category it makes sense, especially if you if you think of a world of basically unlimited compute, which we're not there, right?

83:40 Like we're definitely far from it. But if you had that, well then why wouldn't you want a bunch of agents to try out or look through your code, try your application out in all different kinds of ways and find all the bugs before the user did. >> Yeah, >> that's what I would want. And I think like tools like this are going to be steps to get there. Tools like you know Sweet Sweep, Detail and and others will help.

84:03 >> Yeah. >> And that's the show today. This is Agent Hour. You can follow us on XRA. Follow us on YouTube-ai. I am SM Thomas 3 OnX Obies at Abby. Any parting wisdom before we close the show out this week? >> I think I'm gonna start using the word super intelligence. >> No more AI. >> Just throwing it out there. I'm going to give it a try. >> See how it goes. >> Should we rebrand the show Super Intelligence Hour?

84:30 >> Probably, you know, Monster.si. You know, maybe. Well, I guess in California you're not allowed to say Super Intelligence, but who cares? That's >> We'll give it a try this week. >> Yeah. just yeah throw it throw super intelligence around at some of the meetups and yeah >> see if see how it lands. >> I might do that tonight and be like, "Hey, you guys are all interested in super intelligence, right?" >> All right, everyone. Thank you for tuning in to Asian Hour. We'll be back again next week. We should be doing it on Monday. It is a holiday in the US, but I think we'll still be here. So hopefully you will be as well.

85:04 We'll see you next time. >> Peace. Yo yo that show's a rap. We were live in the zone agents out with Shane and I be on the throne. Did you give us that review only if it's a five? Jump on the tube. Make sure to like and subscribe. New so fresh. Yeah, we keep you in the loop. Get so fly. They bring the whole troop. AI on the rise. Don't miss this power. Welcome to the show. It's AI sour. Did you just drop in? Is this your first time? Make sure to follow us on next and go like and subscribe. Yeah.

85:38 Learn the principles and patterns in our books. The master.AI site. Give it a look. New so fresh. Yeah, we keep you in the loop. Guess so fly. They bring the whole troop. AI on the rise. Don't miss this power. Welcome to the show. It's AI agent sour. This is the end. We all wrapped up. Another showdown. Another one coming up. AI agent sour done, but the news doesn't cease. Shane and Abby, we out of here. Peace.

Summary

The Agents Hour episode covers a range of topics in the AI space, including the significance of security in AI development, recent tech events in San Francisco, and the latest advancements in AI models. The hosts, Shane and Abby, engage with guest Ally to discuss the importance of security measures in AI systems, the implications of tech events, and the evolving landscape of AI models and their applications.

- The episode emphasizes the critical role of security in AI, particularly as AI agents become more autonomous.
- Ally discusses her work in security and the challenges of maintaining identity and access control for AI agents.
- The hosts highlight various tech events happening in San Francisco, including hackathons and conferences focused on AI.
- The conversation touches on the recent release of new AI models, including Gemini 4 Argon from Google, which shows promising benchmarks.
- The episode discusses the implications of the White House AI luncheon, where tech leaders committed to responsible AI development and oversight.
- A significant portion of the discussion revolves around the evolving nature of open-source contributions in light of AI advancements, with some developers disabling external pull requests.
- The hosts also cover the drama surrounding AI companies and their competitive dynamics, highlighting recent controversies in the industry.

Questions Answered

What is Agents Hour about?

Agents Hour is a weekly show focused on AI agents, featuring news, guest discussions, and problem-solving in the AI space.

Why is security crucial for AI agents in enterprises?

Security is essential for enabling stable agent identities and enforcing policies across various environments, which is necessary for enterprises to leverage AI effectively.

How does customization relate to security in AI systems?

Customization in AI systems must be balanced with security measures like tamperproof audit logs and just-in-time authorization to ensure safe experimentation.

What are decision models and their significance?

Decision models, like the ones introduced by OpenAI, enable rapid decision-making and have created a new category in AI, impacting how organizations approach classification tasks.

How does consumer adoption impact AI companies?

Companies like Google are well-positioned to leverage their existing consumer base and tools to create competitive AI solutions, particularly in workplace settings.

What is the best approach for managing contributions in open source AI projects?

Encouraging detailed issue submissions rather than pull requests can streamline the contribution process and ensure faster resolution of valid issues.

© transcribe · For agents Built with care and craft by Gokul Rajaram