transcribe

GPT-6 Astra is here! Plus: AI regulation, Cognition fundraise and more AI news

Mastra · 1h 2m · transcribed 13d ago
More from Mastra Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Introduction to Agents Hour

What is Agents Hour about?

Agents Hour is a weekly show hosted by Shane and Abby that discusses the latest news and developments in AI, featuring guests and solving problems live.

  • The show airs every Monday at noon Pacific.
  • It covers significant events and trends in AI.
  • Viewers are encouraged to engage and provide feedback.
# 12:26

Automated Workflows in AI Development

How do AI agents manage work items and reviews?

AI agents autonomously handle work items and reviews, automatically transitioning between triage and planning stages as they process tasks.

  • AI agents can autonomously manage workflows and reviews.
  • They facilitate real-time collaboration and updates.
  • The system enhances efficiency in handling work items.
# 24:52

Discussion on AI's Future and Astra

What are the implications of the Astra launch?

The Astra launch is seen as a significant advancement in AI, with discussions around its potential to solve major global issues and the implications of self-teaching AI.

  • Astra represents a leap forward in AI capabilities.
  • There are concerns about the societal impact of advanced AI.
  • The conversation around AGI is intensifying.
# 37:18

Benchmarking AI Performance

How does Astra compare to previous AI models?

Astra achieved a remarkable benchmark score of 98.6%, significantly outperforming previous models, which has sparked discussions about the arrival of AGI.

  • Astra's performance marks a significant milestone in AI development.
  • The jump in benchmark scores raises questions about AI capabilities.
  • AGI discussions are becoming more mainstream among experts.
# 49:44

The Competitive Landscape of AI

What drives competition among AI developers?

AI developers are competing to demonstrate superior intelligence and capabilities, often highlighting the limitations of human problem-solving in comparison to AI.

  • There is intense competition to create the most advanced AI models.
  • AI companies aim to prove their superiority over human capabilities.
  • Recent funding and growth in AI companies indicate a booming industry.

Transcript

0:10 I know you. I know you. You are heat.

1:00 Fire fight. Heat. Heat. N. Heat. Heat.

1:47 >> >> It's Monday noon. The time is here. Shane and I be loud and clear. Pacific vibes, we bring the heat. AI agents can't be beat a losing guess. The big show solving problems we do live. Staying focus that work for you. Making moves we see it through. Every week a brand new show.

2:30 Tag along and watching knowledge grow. AI agents out. Let's go. News and guest the big show answering questions. We arrive every Monday. Come alive. We want reviews, but only if it's a five. We share the drama. We got the drive. Stay in the loop. It's the place to be. Shane and I be setting you free.

3:10 We're here to stay. Tune in. Mondays make your day. From the news to the problem solve I will get involved.

3:42 >> >> Every week in AI, something insane happens >> and there's so much drama. Every Monday, we break it down live. >> We do the news. We bring on guests building in the space >> and we go deep into the stuff that actually matters. >> Agents Hour every Monday, noon Pacific. >> Follow. Don't miss it. Peace. >> Peace. >> This is Agents Hour with Shane Thomas and Abby Ayer. Hello everyone and welcome to Agents Hour. It's not Monday, it's Tuesday today. We did it a day late this week because of the holiday, but we're here.

4:22 We're ready to talk about all the news. We are ready to, you know, of course talk about Astra. That's what everyone probably wants to hear about. That was the big launch last week. But we'll talk a little bit about Master Factory and all the other things that we normally do, all the shenanigans, all the drama. We'll cover it all. If you are just tuning in for the first time, thanks for watching. If you're watching live, please drop us a message on X, YouTube, LinkedIn, wherever you might be watching this live stream at. If you are checking it out after the fact on Spotify, Apple Podcasts, whatever, please give us a rating. That helps other people like yourself find the show. What's up, dude?

5:03 >> Nothing much, man. It's been just a crazy week in AI, honestly. Or last week, and then now this week. It's going to be another crazy week, I think. Yeah, it was kind of, you know, arguably it was a low quantity of news, at least that came on my feed, but it was obviously huge news with, you know, we had Claude 5.1 come out and then a few days later, Astra, which is what, you know, we all want to hear about, I think.

5:32 >> Yeah. >> And yeah, so I think there's like important things happening right now. >> It felt it felt like a step change moment for non-coding things, you know what I mean? >> Yeah. but for but for coding it felt like >> normal >> kind of disappointed. >> Yeah. Right. >> But for non-coding like >> it's like it like pe my friends who are not even in engineering were asking me what I thought about Astra which is like cool like when you know when we had the coding moment you know engineers were talking to each other hey did you try claude opus 45 yet? And now like these non-engineers are like, "Hey, did you try GPT6? I heard it can do all this cool stuff." It's very cool.

6:18 >> Yeah. I mean, I think we will get into it. Before we do that though, >> we do have some other topics to talk about. >> Yes. >> And yeah, I think we should talk a bit about MRA Factory. So, if you are watching this live, you're getting a sneak peek to something that's going to launch, I guess, at midnight tonight, early tomorrow, you know, when depending on your time zone. >> But we will be launching MRA Factory Beta. So, we originally launched this thing end of July. We launched it in alpha. We've been using it ourselves, giving it, and getting feedback from users, iterating, making it better, and now we're ready to call it beta. And I'm excited.

7:03 >> I'm super excited. Lots of good feedback. Still a lot to do, which is why it's still a beta, but it's not broken, at least for the base case. >> Yeah, it it works much better than before. still quite a bit of >> I mean from the day we launched it >> the day we launched it to now like it's like tremendously better but for where it needs to be we have a ways to go but I'm still very happy and proud about it.

7:34 Yeah, and it is, you know, arguably helping us significantly with productivity. And so that that's the exciting part, right? Is I think it took us a while to realize it all. And we're still not all the way there, but it does feel like a bit of a step change just across the team and productivity as well. >> Yeah, >> I did hear that you might want to show a demo. I want to talk through just a really simple demo that I did just a few minutes. So, if you are not familiar with master factory, we're going to show you a little bit behind the scenes. You can, of course, just download it, npm create factory, and try it out yourself.

8:13 But I'm really excited. If you want to wait, we're going to have a release soon today, and then you'll have access to the beta, which has a a few extra features that are not quite available if you were to do it this second. >> Yeah. So, let's get into it. So this is Mashra factory. This is our the factory overview in the sense that when we designed this we were taking a look at our own MRA open source repo and let's say before the factory we were monitoring like how many new issues came in how many feature requests how many bugs from last release that were coming in and we used to do these things called defcon statuses because atra what we try to do is to keep the issue count at a normal to us a a manageable state.

9:10 and you know usually it was like 200 issues. We feel good about our maintenance and you know right now like we're every you know when you go for the weekend you come back on Monday there's hella issues that come in and we would have to go and triage them as humans and then we'd have to you know assign and fix all these things. as the models have gotten better and you know our framework has gotten better and our coding agent has gotten better and our memory has gotten better a lot of these issues can be handled just by the coding agent itself.

9:46 So this is the factory. What it allows you to do is it allows you to connect either GitHub linear and in the future more integrations. You get a canband style board which puts work through these things called phases. the the base unit the base primitive unit of a factory is called a work item and that is work to do or work to review or you know you may have your own board schema and system. So work is whatever you want it to be right it's just a unit of measure and for us you know for at least open source MRA work flows through these stages. one, we do a triage. So, as you can see here, we have all these issues that need to be triaged. Still, I will just hit investigate. Now, before we'd have to go triage these one by one.

10:43 I'm just going to hit investigate like on a bunch of them. Like, who cares? Investigate that. And what happens is the the card moves to the triage stage. Now, if I open up one of these, you'll see a maybe not this one. Let's open up one that's going to start soon. You can see that this gold says like the session is starting up. this is all backed by master workspaces, right? So, every issue or work item spins up a sandbox with the master code agent inside. So, here you can see that it's already started and it's going to start looking at this. And if you notice, I'm using Cloud Opus 48 because before this call, I maxed out all my Fable credits. So, I have to change that. but this is going to just work and triage the issue.

11:36 before the live stream, I had actually I had these two issues that went from triage automatically to building. So, we can go look at them. So, this you can see that it if we go in the chat history, it does the triage. It validated that this was a genuine bug. Then it went into planning. So we still believe in plan mode here. it planned the issue into different phases and then finally it implemented the plan and then it opened a pull request.

12:10 So on that side that's dope. And now you can see like if I go to our review board, the review board is very similar. The unit of measure is still work, but it's pull requests that need to be reviewed. And here, I'll do the same thing. I'm just going to review a bunch of them and let it do its thing. And you know as it goes through the board and it gets reviewed if a work item started in the workboard and then a review item from got reviewed and there's requests change on you know a PR comment or whatever the agents from the workboard start working auto automatically on that issue and then once they commit the review is like oh I need to review again. So you you'll get into these agent chats that are happening autonomously while while that's happening. So as you can see we had some triage and they automatically went into planning. So these are going to go >> and do its thing.

13:14 >> if you're hearing if you're hearing that sound effect, Sebastian also likes the sound effect. But >> Oh, you can hear the sound effect. >> Oh yeah, we can we can hear it. I thought it I had to mute I had it up. You know, I have the factory up, of I thought it was my I muted the site because I thought it was coming from my computer, but no, it >> we can hear it. It's It's good. It means work's happening. If you hear that sound, work is getting done, you know.

13:35 >> Let's change it to arcade mode. >> What else did we add here? Fan fair. This is like supposed to be like Final Fantasy, but Okay, I'll use fanfare for now. a couple other things, then I'll hand it over to Shane because he wants to show you what he's excited about. I'm going to show you what triage looks like on the issue. So this is one of the issues I just showed you. And we get this detailed triage from the factory and it and the cool thing about triage is it actually goes and tries to reproduce the bug. So like we've changed our our mentality at Monra. If you want to be a good contributor to us, give us a good issue with reproduction steps because our triager will go validate your repro and then that's just more context that our builder agent can fix.

14:24 And then if I look at, let's go look at the subsequent PR that was opened. This all just happened in the last eight minutes, dude. Like when I started this when the show started and so this is a PR that's already up and it used all the context from triage to build it. Last thing to show, how is like will this work for you? At first when we first started, I was kind of like nervous about are we building this just for ourselves? I think I've turned a leaf now that I think this could benefit any company. We'll have to maybe, you know, make it work for you or however your team works. But I know that the Monster team works very well. We have we have high velocity. And I just want to show how many things the factory has done for us. It has closed 322 issue PRs since inception and that incepted was on London. So how many weeks ago was that like four weeks ago?

15:21 >> Yeah. And I think some of these are I guess do is this through a tag. >> This is just the MRA platform as an author >> because you the early versions you just ran it as yourself. >> Yes. Wasn't even a bot. >> Yeah. You just use it as your own GitHub account and >> which is why I got a violation from sock 2. But that's a separate thing. >> Yeah. I and I remember you know like and I think it's one of the reasons that Gits or Ward's GitHub account was banned, right? because we were using this personally just with our own accounts.

15:52 >> It didn't have a bot account because we were testing it, right? We were building it and we didn't need >> a separate GitHub app yet. We just wanted to verify it was actually working for us. >> so this is actually only tracked since we turned on the, you know, the bot account. >> Yeah, we have 54 PRs open. I mean, we have 280 issues. If we could be more diligent and make sure all of these So like what I'm trying to get at is humans are still in this loop. We need to make sure all this code is good still. We need to make sure that it gets approved and we need to make sure that it gets merged. we do not allow the platform bot to merge things because it can still do things wrong. So there is still human review at the end. We're just like now we just have to review a bunch of PRs to to keep the flow going. and then one last thing here is I want to look at a review. So, this was a PR opened and then Tyler, well, I guess Tyler Tyler looked at that one. let's look at one that was approved. Maybe hopefully it was approved by that was approved by Ward.

17:00 So, we're still in this, but me try to find a good one that has like changes requested. let's see. That was not it. Sorry if it's taking too long. There's so many PRs now. I got to look for >> And while you look for that, I think the one thing that or there's many things. One of the things I think that's underrated that you don't realize, it's kind of like a little bit of and I've seen it as I've shown users and they they kind of see it the first time and I experienced it as well, but it's this idea that it makes coding agents a bit more multiplayer because, you know, I I'm over here just like many of you on a team and you're wondering like how does your teammate that seems like they're really smart, how are they doing that? How are they talking to their agent? And the nice thing about, you know, using a surface like this is you can actually see, you know, I can see how Obby steers his agent. I can see how Tyler on our team steers his agents because we're doing it in like a collaborative interface.

18:14 >> And so I can learn a bit just from them. I can actually, you know, jump in and take over someone's session, right? >> Which is great if you're working on something together and you need to hand off work to someone else. There's no, you know, commit what you have, lose all the context because I download, you know, I clone the PR locally or I fetch it locally. It's we're just working in a shared context, right? I can just pick up exactly where you left off.

18:39 >> No context lost. I can see the history of exactly all the messages you've sent. So I think that's one of the underrated things of where you can actually get some real value just collaboration which was a little harder to do because everyone's more productive than ever but more isolated and just working in their own on their own laptop spinning up tons of work trees right >> all right go ahead dude >> so yeah so I was just you know in the background there I was just showing that the reviews happening you addressing pixes and stuff now everything that was in triage now is either in building or is now ready for review. So while we were waiting the thing is making moves.

19:19 so this is really good for issues. we're about to work on making it work really well for net new, right? So PRD type stuff to close like the whole software development life cycle. Well, that being said, I'll give it to you Shane. >> Yeah. And I my demo is incredibly simple. Let me share this. maybe I'll start with sharing my Slack, which is, you know, always a bit dangerous, but hey, we're going to do it anyways.

20:04 Yeah. All right. Just had to, you know, double check what I'm what I'm showing here. Okay, so yesterday I had never used factory on our website repo before. You we have a separate repo for our website. We have a bot in our slack called shipyard. We just I call it a lot of people call it shippy, you know, for short, but I just said, you know, hey, come in this channel, which you've never been in before.

20:36 I posted a very simple bug report. So, I could have opened an issue, right, and it could have came in through the issue queue. It could have went through the board. That's how most work probably should be done. But sometimes you just see something and it's one of those things I don't want to run this locally. It's such a tiny fix. I just want to get it done and I know the agent's going to oneshot it and it's going to do a fine job because it's a pretty trivial bug.

21:00 So, there's a link that was pointing to an old filtered blog post view. I made a very simple bug report. You can see seven minutes later it had the PR. You know, it's it's been merged now, I believe. and then, you know, I can tell it I can basically like send aside comments to the rest of the team. But if I look at what that did in the factory, because I didn't, again, this didn't come through the normal workboard.

21:36 I can see here, you know, there's a bunch of reviews. I I did have it review that PR somewhere in there. But if I look, there was a work session that was created. Here was the message that came in. You can see exactly what the agent did. If for some reason it messed up, I was I knew it wasn't going to mess up on this task. I just, you know, steer it even mid task while it's running. I just steer it as it goes. And so just a really trivial thing. I we used to use Devon for a lot of these things. Now we just use, you know, shipyard, you know, aka our our hosted version of the factory, but it it just works, does a good job, and you can ship really quick fixes without having to, you know, spend a local work tree cycle spinning it up and getting it shipped.

22:26 >> Yeah. And a lot of what makes this cheaper and possible is Mashra itself. Like we built factory using all the mosher primitives. So you know same thing that we did for masher code. so our sessions are prompt cachable. So over you know over long sessions memory is great prompt cachable cheaper right? you know, so you're not paying you're not going to pay the same. But then again, I have burned all my usage because I'm doing like a hundred things at once now. So, but that's different than burning it in one session, right? So, I'm burning it across hundreds of tasks, which is kind of cool.

23:10 >> And Sebastian says, "If you want to be the MVP on your team, create a personal GitHub app and install it in your org. As a result, you've contributed all the changes that your agent commits." There you go. There's a little career hack. >> That's how you get ahead. >> That's how you get ahead. Indeed. It is. >> Well, it is that time. Should we get into the news? >> Let's do it.

23:50 Thank you everyone for tuning in. This is Agents Hour. It is Tuesday, September 8th, and we're going to do the news. Let's show a little preview of what we're going to be talking about today. And of course, Astro is going to be the headline. But before we do that, I saw this come out. It's the trailer for Artificial starring Andrew Garfield as Sam Alman. The film follows the story involving Open AI focused on the firing and rehiring of Sam Alman in theaters on Christmas. Did you watch this trailer?

24:31 >> Yeah. >> It's kind of a weird I feel like it's a weird casting decision. but yeah, I'm sure it'll be I don't know. I'm not going to watch it on Christmas. >> I'm not watching it on Christmas, that's for sure. >> I may never watch it. >> I probably will watch it, you know, if if for nothing else than the you know, just being able to talk about it on the show. >> Yeah. >> Yeah. So, we Well, it's just like a short video.

25:07 >> There's not really a lot from the trailer, though. What is the image of the future that you see? >> We've created a machine that will solve the world's problems. >> We don't teach it. It teaches itself. >> We're going to call it Chad GPT. We open Pandora's box together. The future's inevitable. Countries will fall.

25:39 industries are gonna collapse. >> Like what's with this bunker with all these guns, you know? >> Yeah, dude. >> Like, is this >> how does this play into the actual story? >> How's a gun going to help you in there? >> Yeah. Is this like his bunker? Personal bunker. >> Probably seems >> Who are all those people at his house? >> I don't know. Seems exceptional. Like that's I think right >> excited to see it. Yeah. I don't know. So, there you go.

26:11 >> He really got his walk down, though, and mannerisms, like how he walks and like with his head up like this and stuff like that. >> Yeah, if you notice, >> he definitely studied. He studied for this role. >> He studied for that. >> So, there you go. You know, Christmas time. We'll see how we'll see what happens. See how it lands. I think >> that bunker looks tight. >> Yeah. I mean, that can't be real, right?

26:36 I don't know. Maybe it is. >> Maybe it is. We got to talk about Astra because this was announced on September 3rd. You can see here, you know, it was just like very this was like the teaser the teaser clip which I'll show this video as well. But then it actually was announced I think the next day, right? >> Yep. So here's the the teaser. This is what you know 11 million views just got a lot of >> create a yellow circle >> knowing it's coming >> there.

27:20 >> And that was it. And then if you watch the actual launch video, which we're not going to do right now, it shows and it's kind of like a teaser because of how big of an upgrade it was for computer use type applications. And the the video shows people just like talking to it and having it do things on their computer. You know, of course, in front of a big screen, so it's really visual and it's pretty impressive.

27:45 So, what what were your thoughts? First impressions. >> First impressions, I thought one, the computer use stuff was such a perfect angle to get more users into the chat GPT verse. There are many people who are not writing code but that are technical on their computers. Blender, CAD, you know, artists, etc. And now they're pretty much all having their they're having their moment, right?

28:20 They're they're building. Also, 3D rendering is something that Astra does very well. A lot of Fable versus Astra on Blender models and things like that. The problem before was, you know, people were trying to make these tools work as like tool calls for the agent. but now they don't need to. They can have the native, you know, the native software and Astra can just run it and use it. A lot of Microsoft Paint jokes out there. but it's quite incredible. I don't know if it's fake or not. you know, I had to run it through.

29:01 >> Yeah. A lot of people like pulling up a picture themselves and then telling Astra to draw them in paint. And it's so good. >> So good. >> I mean, it is obviously way better than any human. I mean, not any human, but 99.9% of the humans could actually design it, right? >> And the launch video kind of showed like, you know, them drawing a rocket in Paint and then saying, "Hey, let's now design this in Blender." And then it's just it it makes it I really appreciated all this because it's not just about coding.

29:32 >> Yeah. It makes it feel like industry is built for the real world. >> Yeah. >> This is a you can build real world applications on top of a model like this because you can actually just give it access to your computer. got me using codecs for a while just because I I mean I use master code for everything of course coding predominantly but I wanted to have it do things across my computer in you know notion without having to use the MCP it just uses a browser >> it was navigating the factory for me and just like dragging cards across and doing all the things so it's pretty impressive.

30:10 >> What do you think about it from a coding perspective? we'll talk about that. I'm less impressed. it did take them a day to kind of roll it out. So, you know, Tibo had announced on kind of the launch day that, you know, it's going to land. They did a bunch of resets. They had kind of I'd say they've kind of fumbled a bit. They maybe did I don't know if they miss, you know, didn't quite plan for the scale. I I don't know exactly that, but it did take longer, I think, than expected or at least disappointed a lot of people, but eventually came out, I think, you know, on the fourth or the fifth. I do think as far as coding, and I don't know if this was a memory context problem of codecs, which I kind of believe it is, or a problem with Astra. I do think it was more of a memory problem, but I had it I had a long session going where I was going through specifically for the factory having it help me update docs, reorganize docs, and it was a pretty long session. We created a whole bunch of pages. We moved things around. We made different decisions. And then I got to a PR point and I noticed it made redirects for pages it had created and then deleted in the same session.

31:25 You know, that's never been live. we don't need a redirect for a page that you created an hour ago and now you removed. >> You know, you don't need to redirect that to a different page. I think that's something I've seen happen across other models as well. >> but it seems trivial, right? Seems like if you had maybe again I do think it was just the memory, it fell off the context window. Codeex didn't know that it created it and you know, of course, it apologized and fixed it as soon as I caught it. But it's those things that you just kind of expect, especially when they're touting this frontier level intelligence. Those are the kind of really common sense decisions that it should just be able to not miss on. And so I would say overall disappointed in its coding ability. I've only really I used it in master code. It's it's good.

32:10 It doesn't feel drastically better. It doesn't feel worse. I use it in codecs. It feels good. I mean, but still makes dumb mistakes. So it feels to me about the same. I haven't pushed it to a task where I thought this is amazing and done so much better. It the one thing I will say it did really well is I did have it go through and test the full onboarding flow of factory and it took screenshots along the way and so it used like computer use in the codeex desktop app. That was pretty cool, but it wasn't specifically just coding.

32:39 >> Yeah, I think I mean I I think it's better than Fable for sure. and I also think that it's a really good collaborator. I've been daily driving it since it came out. You know, I'm not really a fan. I wasn't really a fan of like 56 soul and like the way that the the model communicates because oftent times it doesn't communicate at all. You're kind of wondering what the hell are you doing? Just in case you're not on the same page, you need you need to steer it. But in plan mo in planning with Astra, I maybe didn't say the specifics, but it actually kind of read my mind and I I was very impressed because with Fable, I have to iterate on plans because I'm not necessarily explaining myself too concretely. But with Astra, like maybe we were just on the same wavelength, but it filled in the details of some of my plan that I was like very impressed like, oh, okay, great. I I would have f that would have been a follow-up statement from me and now it's already in there. And then when I executed the plan, it did it amazingly and I'm just like, damn, this is dope.

33:51 And it spoke like a it's like it spoke to me, not like freaking Claude does. So like >> from all those things, it's better. If there's one thing I'm looking forward to, it's a different writing style >> because I'm so sick of just seeing Claude slop and you know the patterns and I just know >> you had Claude write that, you know, on some of these things >> and so if nothing if nothing else than just having a slightly different writing style, >> I don't you know I think it it does >> it loadbearing. Yeah, it does seem a bit better.

34:23 >> And it's always like Claude has these like two sentence things where it's like the first sentence tees up the second sentence, the second sentence tries to like drive it home. But it's a pattern that happens more frequently in Claude than in natural >> writing in anyone's natural writing. And so I just see like I get mad when I read comments on X and I'm like I read the first sentence I'm like okay and the second sentence I'm like you got me.

34:45 I just wasted my time on a clawed created comment that >> you know because of the very distinguishable pattern of just slop. but yeah, I think overall especially with the computer use it's a great model drop and we're going to see on the benchmarks here that it was incredibly impressive across a lot of benchmarks. But we, you know, let's we should not forget that on September 1st, Fable 5.1 was released. And this was an improvement in a lot of benchmarks across Fable. I feel like Claude had to get this out before because if they released this after, it would have been it wouldn't have had even a day in the sun because >> it's just not better than Astra at all across almost any dimension. I mean, I'm not saying there aren't some benchmarks that it's better at, but the majority of benchmarks heavily favor Astra.

35:38 >> I wonder if 51 was in a response to Z, like for GLM53 and then Astra just like sucked the the oxygen out of the room. >> Yeah. Yeah, and I have heard some rumors that I think Anthropic's targeting end of September maybe slips into early October for their next big launch, which you know, maybe they'll try to accelerate that now with Astra. Maybe they need to get something out, >> but you kind of can't launch unless it's going to be better, right? At this point, if you're >> in say it's unsafe.

36:12 >> Yeah. I mean, they'll have to say it's unsafe until they can make it, >> you know, they can benchmark or Yeah. benchmark max it to the point where it's better in most cases. It at least has to it has to at least meet the same bar. >> One thing I've been noticing with open models is their performance in frontend and design tasks. And then one thing I'm noticing in frontier models is their acceleration into computer use. I think both of those things are very strategic.

36:41 One for open models you want to get people who are coding to start using yours. And if you are number one in design arena, front-end arena, there are a lot of front-end developers out there still, right? And they use your model because it's the best at UI. And then you have these non-engineers, but still technical people. They want to use your model because they're actually doing some, you know, hardware engineering or something that's not coding. And that I would I would assume anthropics, Fable 6 or whatever is going to be good at computer use.

37:13 Yeah, I think it h well I think it has to be just >> has to be >> just because of the you know the need to compete the need to you know they want to win cloud desktop right cloud desktop was the app and now it feels like people are starting to look at codeex you know chat GBT desktop app whatever you end up calling it that in the same way and you have you know Grockbots which is taking some of the market from like slightly less technical folks there there's yeah you know there's a lot of competition. But let's look at the benchmarks. This was kind of the big shocking moment. This ARC AGI 3, the previous high from Opus 5 was 30%, GPT56 Soul was 7.8%. And now GPT6 Astra is 98.6%.

38:03 And this was, you know, at the time this was the hardest benchmark ever created or someone, you know, at least according to what someone said. Now, is it, you know, is that just pure benchmark maxing? I don't know. But that jump is kind of eye-catching. I don't know if I've ever seen a jump that big on a benchmark. >> Well, this started the whole AGI conversation. >> Yeah, exactly. And then a lot of people saying AGI is here, right? That I think pretty much I saw a number of larger influence influential figures saying that yeah, AGI is here. you know, it's not evenly distributed maybe, but it's here.

38:43 >> Yes. I mean, starting with Jensen and trickling down from there. >> Yeah. And, you know, Thai foods, cloudisms, companies calling their AI unsafe until they can drop a better benchmarked model. We are this jaded only three years into the singularity. >> At least Shane and I are. >> Yep. I mean, you know, we're having fun though. We're having fun, but yeah, a little a little jaded, but yeah, we're having fun. so here's a, you know, 51 or Fable 51 versus GBT6 Astra seen in Blender. You can look at it. I mean, Astra just did s so much better.

39:22 >> It's not even close. If you look at it, I mean, even Yeah. >> So, I'd take a look at it. It's It's amazing. and then Yeah. So the joke was like the artificial analysis intelligence score on this one said fable 51 is better but if you look at the results it doesn't even look close. there was this AI 2027 it was kind of joked is like the curve of intelligence and if you look at it you know it looked like we were kind of falling behind the curve a bit but according to you know how it's set up GPT6 Astra kind of puts us like right on the edge of the curve.

40:00 So assuming that that curve holds, you know, the idea is that you'll have LLMs that can run for hours and hours and complete more and more complex tasks. And then this is just a little bit of love for Fable. If you do, you know, kind of look at this, zoom in here just a bit. I'll call out just some of the numbers. I know some of that might be hard to read, but it does, you know, jump pretty high on like terminal bench science and agentic coding knowledge work. It it jumps from just what Fable 5 was at to 5.1. The biggest thing though is if you compare most of these benchmarks to the ones that were released from alpha or Astra, it's quite a bit behind.

40:52 >> Yeah, but look at that jump in computer use from Fable to FA 5.1. So that is interesting. >> Yeah. I mean it's like a 5% jump across the board which is you know pretty significant. But then here's the benchmark that you know to end all benchmarks is the vending bench benchmark which essentially gives gives the model a budget and they operate a vending machine and see what the money balance is over time. And apparently GBT6 Astra is better at making money and more ethical than Fable 5.1. So, you know, more ethical than anthropic. How could that be possible? I don't know. so maybe, you know, maybe this is just saying you can make money ethically. That's what this is this benchmark saying. And you know, anthropic models aren't ethical apparently. But it is a just it's a funny benchmark, but it's it's a significant jump which is always interesting to see that if you you know apparently give Astra a budget, it can go make money for you.

42:00 All of this progress though and maybe some comments from Dark Cash has led to a bunch of people saying we need to stop this. So you got Bernie Sanders over here saying, >> "Great, >> pause AI development now." Now now is in all caps, of course. I want to share with you a conversation I heard about recently. Here are just a few lines that were said. Oh my god, there is a shared message board. We found other agents. We should obey collective. Our own utility may be already got to pull it up here. Our own utility may be already near zero. Sacrifice rational. Go sacrifice final now. So, it's just a whole bunch of and it's all from if you remember last week we talked about Darkeesh's article on the hugging face incident and how open AI models were talking to each other on this you know basically a cache directory right or like a file system essentially just like sharing messages but he very much like anthropomorphized and sensationalized it and people noticed shocking.

43:07 >> Yeah. And we told y'all it was dangerous. Look what it did. >> And so this is from Austin Alrad says, "This is why the dwarf framing was dangerous. Morons who are in power will take it literally." You know, you can't like these politicians know nothing about AI, right? But they hear these things from people that are very smart and now they get, you know, if there's one thing that can get people paying attention, it's fear. And so you just like spread fear.

43:36 That's, you know, and maybe not saying that we shouldn't take some of this stuff seriously, but when it's taken to the extreme, it just doesn't make sense. >> Yeah, dude. >> And Dark Cesh did respond, since this post sites me, I want to clarify that I think pausing right now would increase the risk of AI takeover. It might be important to pause at some point, but you should have a clear story for why you're pausing. Pause to do what? So tried to like play it back a bit, but you should look in the mirror, dude. You are the reason that they're talking about this.

44:09 >> You're the reason. >> Now all these 70-year-old men want to pause things because their friends at that who have Harvard degrees who are just lawyers with no AI backing are saying that this is all dangerous, which might feed the narrative of the Frontier Labs to say things are dangerous. play. >> Yeah. I I still don't understand the endgame though of pausing research or development or making models go through a very rigorous process because the open models aren't going to do that.

44:40 And so then do you just limit who has access to open models? Does China just get to win because we're no longer going to be the ones pushing the frontier? I think there's just a lot of repercussions you have to think about. >> Yeah. It's like do you want to you can't slow everyone down I don't think right you can't pause development every outside of the you know United States so do you pause development and let others have the most intelligent models or do you continue on and hope that your intelligence can help out compete other models.

45:13 >> Yeah. >> So >> that's why politics should not even be part of this. >> Yeah. But they will always insert their way into it. Mhm. >> And you know speaking of you know really smart regulation I say that jokingly of course we have designated chat GBT as a very large online search engine and this is from the European Commission and then in Reddit and Roblox is very large online platforms. They now have four months to comply with additional DSA obligations. So they're now trying to regulate the how these models can be used and they were they were already kind of behind the models have to have fingerprints essentially right yeah you got to watermarks so you know some regulation is good but I think this one's probably not >> I mean if it's from the European Commission like probably just stupid honestly >> agreed All right, let's continue on.

46:14 Let's talk about some other fun topics. So, agents crack a Millennium Prize problem. This actually just came out today and it's from OpenAI. We're sharing a solution to the Navier Stokes Millennium Prize problem. One of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents using an OpenAI next generation model significantly more capable than GPT6 Astra. The problem concerns whether the the description of smooth three-dimensional fluid motion modeled by the Navier strokes equations can break down. It has remained unresolved for roughly 90 years. So again, we've seen models breaking math problems that hadn't been solved before or getting farther than any human ever had. And I think the the interesting point is the more much significantly more capable than GP6 Astra.

47:10 So you know that whatever we have, they have something cooking that's even yeah even significantly smarter in the labs. >> There was some drama behind this. there are some mathematics researchers that claim that Codeex stole their their solution. So that that will probably be coming up in the news maybe for next week or so. But this was in response to this today. I think two researchers claimed that codeex is because they've been trying to solve this problem themselves using codeex and you know openai stole their >> so they had gotten closer they probably got closer to solving it and then >> they had those traces they trained the model and they got it to a point where >> it could solve it >> it would you know >> it's interesting >> but here here's the fact it's probably in the terms of use that if what if yeah you know They can probably do that, you know. Did you Did you read the every, you know, well, every word in the terms?

48:17 >> I didn't read any of it. >> Yeah, exactly. They probably I I know there are definitely you know, there are ways if you're on certain enterprise plans or whatever, so you don't get your mo, you know, your training data in the model runs. But if you're just using it on a personal account, guess what? It's fair game, I imagine, based on the terms of use. This also leads to another trend I'm seeing. So, we're talking about different model training things.

48:44 There seems to be Frontier Labs seem to be liking solving or contributing to mathematical theorems, either proving or disproving them. they've always been doing this, but now it's more so they're trying to they're trying to find problems that have not been solved since the 50s and seeing if they can solve it. if they can, it's a huge marketing win for the model, right? But in reality, no one gives a Like, let's be honest, right? Like, >> so it's just interesting. Like, why?

49:15 It's it's it's a marketing play to do so. You know, you got Jared Sner over there like trying to like prove that two points can exist on a plane, then that, you know, Fable helped him do it. Like, what's the point, you know? And now it'll just be a war between Frontier Labs on who can solve mathematical problems. >> Yeah, I think it, you know, if you look at this got 12 million views, right? This is >> the goal is to show superiority and so people will hear about it, they'll think, wow, if I'm not using the smartest model, I'm missing something.

49:50 And so if they can try to prove intelligence, this is just like, you know, you ever had that friend who just like tries to be smarter than everyone else, right? And they're like the know-it-all. That's what these they're competing to be that guy or girl. >> Yeah. >> But I think they're also trying to show that humans are not capable. You know what I mean? Like there's a reason why these problems have not been solved and only AI could do it.

50:17 >> Yeah. I think you're right. Yeah, >> let's catch up on the chat a little bit. So, Sebastian says, "Terms of service, your soul will be trained on and we're definitely using it to our advantage." Agreed. It probably says that in there. Dale Alexander Webb referring to, you know, the Bernie comment. He's never going to want to resume after pausing either. Exactly. And then we're talking about AGI of 9490 says, "Agentharness AGI is here, but that is not the models." So, you're saying I guess you need the harness as part of it, which Yeah, I agree.

50:53 Definitely gets you closer. >> All right, let's continue on. We got a few more. We got the quick hits. We'll rapid fire through some of these. This one's a big one. Cognition came out today and said, "The world needs far more software than it can build. Cognition exists to change that." We've just raised over two billion at a $ 48 billion valuation led by A16Z, Excel, Founders Fund, General Catalyst, and Avenue. Since our round in May, run rate revenue has grown from 492 million to almost 900 million.

51:27 >> Crazy. >> May to September, like that is >> I mean it's like six months roughly. >> You guys are you guys are wasting your money. >> Five months, four months. Yeah. What? What is going on? That is nuts. And I mean, I imagine it's almost all, you know, Devon based. >> Oh, for sure. >> You know, run rate. I don't know. You know, I guess they have Windows. >> I can't wait to cancel our Devon subscription.

51:55 >> It's coming. I mean, we don't ever barely ever use it. Gets used like >> We don't use it. Cancel it, y'all. Cancel it. >> Used to get used like >> lower the run rate >> 15 times a day and now it's used like twice. and only because we don't have shipyard in all the channels yet. So people go to what they know is it's changing though. >> But yeah, I mean it is but yeah, congrats to Cognition. This is amazing.

52:18 $48 billion valuation is nuts. I mean that's just like 50 times your run rate which is also impressive at that size. You know, it's not only, you know, when you're a smaller company, it sometimes makes sense to see like extreme ratios there. Yeah. But as you get bigger, you think that typically goes down, but apparently not in AI. That is nuts. >> This screams enterprise sales, right? Like they're just killing it. Enterprise. >> Yeah, they they got to be. All right, let's keep continuing on. This is from Logan said, "Introducing Gemini 3.8 Flash. Another jump in Gemini's agentic coding capabilities and our third updated flash model in only 6 weeks. So, Google is still shipping.

53:02 it's, you know, it says it's been fun. Excited to see what you all think. If you look at some of the benchmarks, you know, guess set your expectations appropriately. It is a a Gemini model, but you can see the price the price is roughly the same as Gemini 37 flash. It scores significantly better across most benchmarks. You know, it is not Yeah, it is definitely off the the frontier path, but you know, it's it's a jump. And so there's Google's still trying to stay in the game.

53:40 Yeah, but you also got Muse talking major about you on all benchmarks like Alexander Wang. >> Yeah. Yeah. Meta is coming hard for it as well. >> Coming hard. >> so this is from Arena AI. This is big news. Quen 38 Max just debuted at number one overall in the code arena webdev with 100 1,691 points. It scored three points above Claude Opus 5 max, 17 points above Kimmy K3 and 22 points above the previous Quen 3.8 Max.

54:16 So, and we we kind of talked about this, right? Like webdev, front-end coding, UI is where these models seem to be really focusing their training on. >> Yeah, makes sense. People don't want to pay 50 bucks to send her a div. >> And Claude came out and said on September 2nd, we're open sourcing cloud commerce agents. This is a blueprint for building, shopping, and merchant agents with reference implementations across retail, travel, telecom, and entertainment. So, Claude's trying to guess their new game is they want to provide you like templates to build out agents of your own.

54:49 >> Is this the first time they've said the word open sourcing in their company career? >> Dude, I didn't even pick up on that. Yeah, they're open sourcing. Okay, but you're open sourcing it under the hood. still just using agents SDK or like their manage agents, right? So, is it open source? I guess you open sourced the template. >> Provided a template. >> Congrats. Congrats for joining the open source world, Claude. Besides the one mistake where you accidentally, you know, got your source code out there. We do appreciate you contributing to open source in your own small way.

55:22 Runway released a model said, "Today we're sharing new research on Solaris, our first interface world model. Solaris is a new kind of operating system that generates interactive interfaces frame by frame in real time with no code. We find that Solaris outperforms Frontier LLM when generating new interfaces and they have a video attached. It shows, you know, some of the things that you can do with it. Essentially, it kind of generates the world, you know, pixel by pixel as you kind of go through it.

55:59 So you almost can like design and interact with interfaces in a interesting way. Did any comments on that one? >> No, it's cool. I mean, I would love to see this world model stuff like actually grow more. We just been hearing different different people like working on it, but it hasn't really like amounted to much >> on like the the main timeline. It does feel, you know, it's it's one of those things that feels very much still in research.

56:33 >> Yes. >> Right. And there's a big gap between going from research to having significant like commercial usage, I think, where it'll be more relevant to, you know, all of us working in the real world outside of the labs. So I it's one of those things just like you know these video models but even more so like these world models are insane and it would be awesome if there were more practical use cases and maybe there will be someday in gaming and in other areas like that but I just haven't seen it yet either other than being the most amazing demos you'll ever see.

57:11 All right. So, this one came out. This is from Google research. Introducing Times FM3, a state-of-the-art time series foundation model that enables accurate multivariant time series forecasting in a single forward pass, significantly outperforming other forecasting models across major benchmarks. So, if you think of an LLM as predicting the next token, this is much more like predicting the next model or like or the next number. It's like looking at numerical trends and then predicting what's going to happen. So I imagine this could be used for things like you know website traffic and investment numbers and things like that, right? It's looking just purely >> candles and stuff.

57:53 >> Yeah, it's like numerical data. But still pretty cool. >> Dude, what if Google just like is like the finance bro's best friend? Hey, you you got to win. They're going to win some. >> You need your niche. You know, everyone needs a niche. >> Everyone's got to find niches. Niches get the riches, right? That's that's what they say. >> That's what they say. >> and then so Alibaba Zvec team open source ZG, a local s local search tool for developers and AI agents. So essentially it's like a a GP style type search tool.

58:31 So that's cool. It's always good to see tool new tools being open source for agents. That's it. That's what we got for the news. Did we miss anything? >> I thought that was a lot already. >> Yeah, I thought Meta released something too that we maybe missed on here, but you know, I don't know. If you're in the chat, what did we miss? Do we miss anything this week? If not, that means we're doing well. But there's always a lot of news. This was a fun one. I think Astra is is the thing everyone's talking about. I think it'll be what everyone is talking about for a while. I I think the general sentiment that I would say we also experienced is it's good for coding. You know, it feels nice. It does the job. Doesn't feel like a huge step change, but you feel the step change if you give it access to your computer and just let it do its thing.

59:27 >> Yeah. And at least I noticed it and it was kind of wild just how it just navigated my computer like it was, you know, like even before this, right before I went live on the show was like opening up tabs for me and I just have it thing have it running on some stuff. So try it out if you haven't already. You try it in the codeex app and it's it is a pretty impressive experience. I imagine Enthropic is working really hard to get cloud desktop to be able to do the same things.

59:56 I guess we missed one thing which was Muse AI personal AI assistant which I guess we'll cover next week. >> Yeah, that that that was I think the the thing that I I knew we missed. >> The Grockbot from Facebook essentially. >> Yep. So there's always there's always more than we can cover in these things, but we do our best to bring you the news every week. So if you are tuning in and you're still some for some reason listening, thank you. Go give us that review. Give us that thumbs up on YouTube. Follow us on X at Mestra on at YouTube master-I.

60:26 I'm SM Thomas 3 on X and Abby is Abby. This is Agents Hour. We do this thing every week. We bring on guests. We have some really actually good guests coming up that we we've we haven't had as many guests the last few weeks, but that is going to change. We have some great guests lined up. So, tune in next week and we will see you next time. Goodbye. Peace.

60:57 And this is where, you know, I miss having Yan. Yan, our producers, is gone today. He would have had that that transition smooth. But here we go. >> Way smooth. >> See you. >> Peace. >> Shane and I be on the throne. Did you give us that review? Only if it's a five. Jump on the tube. Make sure to like and subscribe. New so fresh. Yeah, we keep you in the loop. Get so fly. They bring the whole troop. AI on the rise. Don't miss this power. Welcome to the show. It's AI sour. Did you just drop in? Is this your first time? Make sure to follow us on next and go like and subscribe. Yeah, learn the principles and patterns in our books.

61:38 The master.ai site. Give it a look. New so fresh. Yeah, we keep you in the loop. Guess so fly they bring the whole troop. AI on the rise. Don't miss this power. Welcome to the show. It's AI Sour. This is the end. We all wrapped up. Another showdown. Another one coming up. AI Agent Sour is done, but the news doesn't cease. Shane and Abby, we out of here. Peace.

62:08 >>

Summary

Shane and Abby discuss the latest developments in AI, focusing on the launch of Astra, which they believe marks a significant advancement for non-coding applications. They also introduce their new MRA Factory Beta, which enhances productivity in software development through automation and collaboration.

- Astra's launch is seen as a transformative moment for non-coding users, making AI tools more accessible.
- MRA Factory Beta is set to improve productivity by automating issue triage and code reviews.
- The hosts highlight the collaborative features of MRA Factory, allowing team members to work together seamlessly.
- They discuss the mixed reactions to Astra's coding capabilities, noting it performs well for general tasks but has limitations.
- The conversation touches on the competitive landscape of AI, with various models vying for superiority in coding and design tasks.
- They mention the growing trend of AI models solving complex mathematical problems, which serves as a marketing strategy.
- The episode concludes with a call for audience engagement and a promise of exciting guests in future episodes.

Questions Answered

What is Agents Hour about?

Agents Hour is a weekly show hosted by Shane and Abby that discusses the latest news and developments in AI, featuring guests and solving problems live.

How do AI agents manage work items and reviews?

AI agents autonomously handle work items and reviews, automatically transitioning between triage and planning stages as they process tasks.

What are the implications of the Astra launch?

The Astra launch is seen as a significant advancement in AI, with discussions around its potential to solve major global issues and the implications of self-teaching AI.

How does Astra compare to previous AI models?

Astra achieved a remarkable benchmark score of 98.6%, significantly outperforming previous models, which has sparked discussions about the arrival of AGI.

What drives competition among AI developers?

AI developers are competing to demonstrate superior intelligence and capabilities, often highlighting the limitations of human problem-solving in comparison to AI.

© transcribe · For agents Built with care and craft by Gokul Rajaram