transcribe

GPT-6 Astra vs Claude Fable 5.1, We Should Pause AI & The Benchmark Wars | This Week In AI

Mastra · 27m · transcribed 13d ago
More from Mastra Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Anticipation for New Writing Styles

What are the hosts' thoughts on the current writing styles in AI?

The hosts express their dissatisfaction with the repetitive writing styles produced by AI, particularly Claude, and look forward to a change.

  • The hosts are tired of the current AI writing styles.
  • They anticipate a different and more engaging writing style in the future.
  • There is a sense of skepticism about the quality of AI-generated content.
# 5:25

Evaluating Astra's Performance

How does Astra compare to previous AI models in coding tasks?

While Astra shows improvement over previous models, it still makes common mistakes and doesn't feel drastically better overall.

  • Astra is seen as a better collaborator compared to Fable.
  • It performs well in specific tasks like onboarding but still has limitations.
  • Users expect more from AI models, especially in coding capabilities.
# 10:51

AGI Discussions and Benchmarking

What are the implications of recent benchmarks on AGI?

Recent benchmarks suggest that AGI is closer than ever, with Astra outperforming previous models significantly.

  • AGI is a hot topic, with many influential figures claiming its arrival.
  • Astra's performance indicates a leap in AI capabilities.
  • The conversation around AGI is evolving, with benchmarks playing a crucial role.
# 16:17

OpenAI's Solution to a Millennium Prize Problem

What recent achievement did OpenAI accomplish regarding a significant mathematical problem?

OpenAI's model provided a solution to the Navier Stokes Millennium Prize problem, showcasing its advanced capabilities.

  • OpenAI's model is significantly more capable than previous iterations.
  • The solution to a long-standing mathematical problem indicates progress in AI research.
  • There are ethical concerns regarding the ownership of solutions generated by AI.
# 21:43

Competition in AI Development

How are different AI models performing in the competitive landscape?

Quen 38 Max has emerged as a leader in web development benchmarks, indicating a shift in focus among AI models.

  • Quen 38 Max outperforms other models in web development tasks.
  • OpenAI's move towards open sourcing indicates a shift in strategy.
  • The competitive landscape is heating up with multiple players advancing their technologies.

Transcript

0:00 If there's one thing I'm looking forward to, it's a different writing style because I'm so sick of just seeing Claude slop and you know the patterns and I just know you had Claude write that. >> Is it loadbearing? This is Agents Hour. It is Tuesday, September 8th, and we're going to do the news. Astro is going to be the headline.

0:31 But before we do that, I saw this come out. It's the trailer for Artificial, starring Andrew Garfield as Sam Alman. The film follows the story involving OpenAI focused on the firing and rehiring of Sam Alman in theaters on Christmas. Did you watch this trailer? >> I feel like it's a weird casting decision. I don't know. I'm not going to watch it on Christmas. >> I'm not watching it on Christmas, that's for sure. >> I mean, I may never watch it. I probably will watch it, you know, if if for nothing else than the just being able to talk about it on the show.

1:01 >> Yeah. Countries will fall. >> Industries are going to collapse. >> I'm like, what's with this bunker with all these guns, you know? >> Yeah, dude. >> Like, is this how does this play into the actual story of him getting fired? >> Like, how's a gun going to help you in there? >> Yeah. Is this like his bunker? Personal bunker? >> Probably. seems >> who are all those people at his house? >> I don't know. >> He really got his walk down though and mannerisms like how he walks and like with his head up like this and stuff like that.

1:35 >> He definitely studied. He studied for this role for that. >> So there you go. We'll see what happens. See how it lands. I think >> that bunker looks tight. >> Yeah. I mean that can't be real, right? >> Maybe it is. >> We got to talk about Astra because this was announced on September 3rd. This was like the teaser teaser clip, but then it actually was announced I think the next day, right? >> Create a yellow circle there.

2:08 >> 11 million views just got a lot of people knowing it's coming and that was it. And then if you watch the actual launch video, which we're not going to do right now, it shows and it's kind of like a teaser because of how big of an upgrade it was for computer use type applications and they the video shows people just like talking to it and having it do things on their computer, you know, of course in front of a big screen, so it's really visual and it's pretty impressive. What were your thoughts? First impressions, I thought one the computer use stuff was such a perfect angle to get more users into the chat GPT verse. There are many people who are not writing code but that are technical on their computers, Blender, CAD, you know, artists, etc. And now they're pretty much they're having their moment, right? They're they're building also 3D rendering is something that Astro does very well. a lot of Fable verse Astra on Blender models and things like that. The problem before was people were trying to make these tools work as like tool calls for the agent, but now they don't need to. They can have, you know, the native software and Astra can just run it and use it. A lot of Microsoft Paint jokes out there, but it's quite incredible. I don't know if it's fake or not. You know, I had to run it the experiment. a lot of people like pulling up a picture themselves and then telling Astra to draw them and paint and it's so good.

3:32 >> So good. >> I mean it is obviously way better than any human. I mean not any human but 99.9% of the humans could actually design it. Right. >> And the launch video kind of showed like you know them drawing a rocket in paint and then saying hey let's now design this in Blender. I really appreciated all this because it's not just about coding. >> Yeah. It makes it feel like industry is built for the real world. Yeah, >> you can build real world applications on top of a model like this because you can actually just give it access to your computer. It got me using codeex for a while. I mean, I use master code for everything, of course, coding predominantly, but I wanted to have it do things across my computer and, you know, notion without having to use the MCP. It just uses a browser. It was navigating the factory for me and just like dragging cards across and doing all the things. It's pretty impressive.

4:18 >> What do you think about it from a coding perspective? >> We'll talk about that. I'm less impressed. it did take them a day to kind of roll it out. So you know TBO had announced on kind of the launch day that it's going to land. They did a bunch of resets. They had kind of I'd say they kind of fumbled a bit. I don't know if they didn't quite plan for the scale. I I don't know exactly but it did take longer I think than expected or at least disappointed a lot of people but eventually came out I think you know on the fourth or the fifth. I do think as far as coding and I don't know if this was a memory context problem of codecs which I kind of believe it is or a problem with Astra. I do think it was more of a memory problem but I had it I had a long session going where I was going through specifically for the factory having it help me update docs, reorganize docs and it was a pretty long session. We created a whole bunch of pages, we moved things around.

5:09 We made different decisions and then I got to a PR point and I noticed it made redirects for pages it had created and then deleted in the same session. That's never been live. We don't need a redirect for a page that you created an hour ago and now you removed. You know, you don't need to redirect that to a different page. I think that's something I've seen happen across other models as well. But it seems trivial, right? Seems like if you had and maybe again I do think it was it just the memory, it fell off the context window. Codeex didn't know that it created it. And you know, of course, it apologized and fixed it as soon as I caught it. But it's those things that you just kind of expect, especially when they're touting this frontier level intelligence. Those are the kind of really common sense decisions that it should just be able to not miss on. And so I would say overall disappointed in its coding ability. I've only really I used it in master code.

5:56 It's it's good. It doesn't feel drastically better. It doesn't feel worse. I use it in codeex. It feels good. I mean, but still makes dumb mistakes. So it feels to me about the same. I haven't pushed it to a task where I thought this is amazing and done so much better. The one thing I will say it did really well is I did have it go through and test the full onboarding flow of factory and it took screenshots along the way and so it used like computer use in the codeex desktop app.

6:21 That was pretty cool. But it wasn't specifically just coding. >> Yeah, I think it's better than Fable for sure. And I also think that it's a really good collaborator. I've been daily driving it since it came out. I wasn't really a fan of like 56 soul and like the way that the the model communicates cuz often times it doesn't communicate at all. Kind of wondering what the hell are you doing just in case you're not on the same page and you you need to steer it. But in plan mo in planning with Astra I maybe didn't say the specifics but it actually kind of read my mind and I I was very impressed cuz with Fable I have to iterate on plans because I'm not necessarily explaining myself too concretely. But with Astra, like maybe we were just on the same wavelength, but it filled in the details of some of my plan that I was like very impressed, like, "Oh, okay, great." That would have been a follow-up statement from me and now it's already in there. And then when I executed the plan, it did it amazingly.

7:15 And I'm just like, damn, this is dope. And it spoke to me, not like freaking Claude does. So like from all those things, it's better. If there's one thing I'm looking forward to, it's a different writing style because I'm so sick of just seeing Claude slop and you know the patterns and I just know you had Claude write that, you know, on some of these things >> if nothing else than just having a slightly different writing style. I don't you know I think it it does seeming. Yeah, it does seem a bit better. And it's always like Claude does these like two sentence things where it's like the first sentence te's up the second sentence, the second sentence tries to like drive it home, but it's a pattern that happens more frequently in Claude than in natural writing in anyone's natural writing. And so I get mad when I read comments on X and I'm like I read the first sentence, I'm like, "Okay." And the second sentence, I'm like, " you got me. I just wasted my time on a clawed created comment that, you know, because of the very distinguishable pattern of just slob."

8:08 But yeah, I think overall, especially with the computer use, it's a great model drop. And we're going to see on the benchmarks here that it was incredibly impressive across a lot of benchmarks. We should not forget that on September 1st, Fable 5.1 was released. And this was an improvement in a lot of benchmarks across Fable. I feel like Claude had to get this out before because if they released this after it would have been it wouldn't have had even a day in the sun because it's just not better than Astra at all across almost any dimension. I mean, I'm not saying there aren't some benchmarks that it's better at, but the majority of benchmarks heavily favor Astra.

8:45 >> I wonder if 51 was in a response to Z like for GM53 and then Astra just like sucked the the oxygen out of the room. >> Yeah. And I have heard some rumors that Anthropic targeting end of September maybe slips into early October for their next big launch. Maybe they'll try to accelerate that now with Astra. Maybe they need to get something out. But you kind of can't launch unless it's going to be better, right? At this point, >> if you're in say it's unsafe, >> I mean, they'll have to say it's unsafe until they can benchmark max it to a point where it's better in most cases.

9:17 It has to at least meet the same bar. One thing I've been noticing with open models is their performance in front end and design tasks. And then one thing I'm noticing in frontier models is their acceleration into computer use. I think both of those things are very strategic. One for open models, you want to get people who are coding to start using yours. And if you are number one in design arena, front-end arena, there are a lot of front-end developers out there still, right? And they use your model because it's the best at UI. And then you have these non-engineers but still technical people. They want to use your model because they're actually doing some you know hardware engineering or something that's not coding. I would assume anthropics fable 6 or whatever is going to be good at computer use.

10:01 >> Yeah, I think it h well I think it has to be just because of the the need to compete. You know they want to win cloud desktop, right? cloud desktop was the app and now it feels like people are starting to look at codeex you know chat GPT desktop app whatever you end up calling it in the same way and you have you know Grock bots which is taking some of the market from like slightly less technical folks there's a lot of competition but let's look at the benchmarks this was kind of the big shocking moment this ARC AGI 3 the previous high from Opus 5 was 30% GPT56 soul was 7.8% 8% and now GPT6 Astra is 98.6%.

10:39 At the time, this was the hardest benchmark ever created or someone, you know, at least according to what someone said. Now, is that just pure benchmark maxing? I don't know. But that jump is kind of eye-catching. I don't know if I've ever seen a jump that big on a benchmark. >> Well, this started the whole AGI conversation. So, >> yeah, exactly. And then a lot of people saying AGI is here. I saw a number of larger influential figures saying that, yeah, AGI is here. It's not evenly distributed maybe, but it's here.

11:06 >> I mean, starting with Jensen and trickling down from there. >> So, here's a Fable 51 verse GPT6 Astra. Yeah. Seen in Blender. I mean, Astra just did so much better. >> It's not even close, >> dude. That's AGI. >> The joke was like the artificial analysis intelligence score on this one said Fable 51 is better, but if you look at the results, it doesn't even look close. There was this AI 2027. It was kind of joked as like the curve of intelligence. And if you look at it, it looked like we were kind of falling behind the curve a bit, but according to, you know, how it's set up, GPT6 Astra kind of puts us like right on the edge of the curve. So assuming that that curve holds, the idea is that you'll have LLMs that can run for hours and hours and complete more and more complex tasks. And then this is just a little bit of love for Fable, but it does jump pretty high on like terminal bench science and agent coding knowledge work.

12:00 It jumps from just what Fable 5 was at to 5.1. The biggest thing though is if you compare most of these benchmarks to the ones that were released from Alpha or Astra, it's quite a bit behind. >> But look at that jump in computer use from Fable to FA 51. So that is interesting. >> Yeah, I mean it's like a 5% jump across the board which is you know pretty significant. But then here's the benchmark to end all benchmarks is the vending bench benchmark which essentially gives the model a budget and they operate a vending machine and see what the money balances over time. And apparently GBT6 Astra is better at making money and more ethical than Fable 5.1. So you know more ethical than anthropic. How could that be possible? I don't know. So maybe, you know, maybe this is just saying you can make money ethically. That's what this is this benchmark saying. And you know, anthropic models aren't ethical.

12:51 Apparently, it's a funny benchmark, but it's it's a significant jump, which is always interesting to see that if you apparently give Astra a budget, it can go make money for you. All of this progress though and maybe some comments from Dorcash has led to a bunch of people saying we need to stop this. So, you got Bernie Sanders over here saying, "Pause AI development now." Now is in all caps, of course. I want to share with you a conversation I heard about recently. Here are just a few lines that were said. Oh my god, there is a shared message board. We found other agents. We should obey collective. Our own utility may be already near zero. Sacrifice rational.

13:34 Go sacrifice final now. So, it's just a whole bunch of and it's all from, if you remember last week, we talked about Darkeesh's article on the hugging face incident and how open AI models were talking to each other on this, you know, basically a cache directory, right? Or like a file system essentially just like sharing messages. But he very much like anthropomorphized and sensationalized it and people noticed. Shocking. >> Yeah. And we told y'all it was dangerous. Look what it did. And so this is from Austin Alred says, "This is why the dwarfish framing was dangerous.

14:06 Morons who are in power will take it literally. These politicians know nothing about AI, but they hear these things from people that are very smart. If there's one thing that can get people paying attention, it's fear. And so you just like spread fear." Not saying that we shouldn't take some of this stuff seriously, but when it's taken to the extreme, it just doesn't make sense. >> Yeah, dude. >> And Dwar did respond, "Since this post cites me, I want to clarify that I think pausing right now would increase the risk of AI takeover. It might be important to pause at some point, but you should have a clear story for why you're pausing. Pause to do what? So, try to like play it back a bit, but you should look in the mirror, dude. You were the reason they're talking about this. You're the reason.

14:42 >> Now, all these 70-year-old men want to pause things because their friends at that who have Harvard degrees are just lawyers with no AI backing saying that this is all dangerous, which might feed the narrative of the Frontier Labs to say things are dangerous. play. >> I still don't understand the endgame though of pausing research or development or making models go through a very rigorous process because the open models aren't going to do that. And so then do you just limit who has access to open models? Does China just get to win because we're no longer going to be the ones pushing the frontier? I think there's just a lot of repercussions you have to think about. It's like do you want to you can't slow everyone down I don't think right? We can't pause development every outside of the, you know, United States. So, do you pause development and let others have the most intelligent models or do you continue on and hope that your intelligence can help out compete other models?

15:36 >> That's why politics should not even be part of this. >> Yeah. But they will always insert their way into it. >> And you know, speaking of you know, really smart regulation, I say that jokingly, of course, we have designated chat GBT as a very large online search engine. And this is from the European Commission. And then in Reddit and Roblox is very large online platforms. They now have four months to comply with additional DSA obligations. So they're now trying to regulate how these models can be used. And they were they were already kind of behind the models have to have fingerprints essentially, right? You got to watermarks. So some regulation is good, but I think this one's probably not.

16:14 >> I mean, if it's from the European Commission, like probably just stupid, honestly. >> Agreed. Dude, did you subscribe? >> Dude, I host the show. Did you subscribe? >> Did you subscribe? >> Subscribe to Agents Hour every Monday, noon Pacific. Let's talk about some other fun topics. So, agents crack a Millennium Prize problem. This actually just came out today and it's from OpenAI. We're sharing a solution to the Navier Stokes Millennium Prize problem. One of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents using an open AAI next generation model significantly more capable than GPT6 Astra. The problem concerns whether the the description of smooth three-dimensional fluid motion modeled by the Navier strokes equations can break down. It has remained unresolved for roughly 90 years. So again, we've seen models breaking math problems that hadn't been solved before or getting farther than any human ever had. And I think the the interesting point is the significantly more capable than GP6 Astra. So you know that whatever we have, they have something cooking that's even significantly smarter in the labs.

17:23 >> There was some drama behind this. There are some mathematics researchers that claim that Codeex stole their their solution. So that that will probably be coming up in the news maybe for next week or so. But this was in response to this today. I think two researchers cuz they had been trying to solve this problem themselves using codecs and you know open AAI stole their >> they had gotten closer they probably got closer to solving it and then they had those traces they trained the model and they got it to a point where it could solve it. Here's the fact it's probably in the terms of use. Yeah. You know they can probably do that. Did you read the every you know every word in the >> I didn't read any of it.

18:03 >> Yeah. Exactly. they probably I I know there are definitely you know there are ways if you're on certain enterprise plans or whatever so you don't get your you know your training data in the model runs but if you're just using it on a personal account guess what it's fair game I imagine based on the terms of use >> this also leads to another trend I'm seeing so we were talking about different model training things frontier labs seem to be liking solving or contributing to mathematical theorems either proving or disproving them they've always been doing this but now it's more so they're trying to find problems that have not been solved since the ' 50s and seeing if they can solve it. If they can, it's a huge marketing win for the model, right? But in reality, no one gives a Like, let's be honest. So, it's just interesting why it's it's it's a marketing play to do so. You know, you got Jared Sumner over there like trying to like prove that two points can exist on a plane then that, you know, Fable helped him do it. Like, what's the point, you know? And now it'll just be a war between Frontier Labs on who can solve mathematical problems.

19:05 >> Yeah, I think it, you know, if you look at this got 12 million views, the goal is to show superiority and so people will hear about it, they'll think, wow, if I'm not using the smartest model, I'm missing something. And so if they can try to prove intelligence, this is just like, you know, you ever had that friend who just like tries to be smarter than everyone else, right? And they're like the know-it-all. That's what these they're competing to be that guy or girl.

19:30 But I think they're also trying to show that humans are not capable. You know what I mean? Like there's a reason why these problems have not been solved and only AI could do it. >> Yeah, I think you're right. We got a few more. We got the quick hits. We'll rapid fire through some of these. This one's a big one. Cognition came out today and said the world needs far more software than it can build. Cognition exists to change that. We've just raised over two billion at a $ 48 billion valuation led by A16Z, Excel, Founders Fund, General Catalyst, and Avenue. Since our round in May, run rate revenue has grown from 492 million to almost 900 million.

20:11 >> Crazy. >> May to September, that's like 6 months roughly. >> You guys are you guys are wasting your money. >> 5 months, four months. Yeah. What is going on? And I mean, I imagine it's almost all you Devonbased >> Oh, for sure. >> Run rate. I don't know. You know, I guess they have >> I can't wait to cancel our Devon subscription. >> It's coming. I mean, we don't ever barely ever use it. Gets used.

20:30 >> We don't use it. Cancel it, y'all. Cancel it. >> Used to get used like >> lower their run rate >> 15 times a day and now it's used like twice and only because we don't have shipyard in all the channels yet. So, people go to what they know. It's It's changing though. But yeah, I mean it is. But yeah, congrats to Cognition. This is amazing. $48 billion valuation is nuts. I mean, that's just like 50 times your run rate, which is also impressive at that size. You know, when you're a smaller company, it sometimes makes sense to see like extreme ratios there.

21:00 Yeah. But as you get bigger, you think that typically goes down, but apparently not in AI. That is nuts. >> This screams enterprise sales, right? Like they're just killing it. Enterprise. >> Yeah, they they got to be. This is from Logan said, "Introducing Gemini 3.8 Flash. Another jump in Gemini's agentic coding capabilities." and our third updated flash model in only 6 weeks. So, Google is still shipping and it says it's been fun. Excited to see what you all think. If you look at some of the benchmarks, you know, set your expectations appropriately. It is a Gemini model, but you can see the price.

21:32 The price is roughly the same as Gemini 37 flash. It scores significantly better across most benchmarks. It is definitely off the the frontier path, but you know, it's it's a jump. Google's still trying to stay in the game. Yeah, but you also got Muse talking major about you on all benchmarks like Alexander Wang. >> Meta is coming hard for it as well. so this is from Arena AI. This is big news. Quen 38 Max just debuted at number one overall in the code arena webdev.

22:04 1,691 points. It scored three points above Claude Opus 5 max, 17 points above Kimmy K3 and 22 points above the previous Quen 3.8 Max. And we we kind of talked about this, right? Like webdev, front-end coding, UI is where these models seem to be really focusing their training on. >> Yeah, makes sense. People don't want to pay 50 bucks to center a div. >> And Claude came out and said on September 2nd, we're open sourcing Cloud Commerce agents. This is a blueprint for building shopping and merchant agents with reference implementations across retail, travel, telecom, and entertainment. Guess their new game is they want to provide you like templates to build out agents of your own. Is this the first time they've said the word open sourcing in their company career?

22:49 >> Dude, I didn't even pick up on that. Yeah, they're open sourcing. Okay, but you're open sourcing it under the hood. It's still just using agents SDK or like their manage agents, right? So, is it open source? I guess you open source the >> provided a template. >> Congrats. Congrats for joining the open source world, Claude. Besides the one mistake where you accidentally, you know, got your source code out there, we do appreciate you contributing to open source in your own small way. Runway released a model. Today, we're sharing new research on Solaris, our first interface world model. Solaris is a new kind of operating system that generates interactive interfaces frame by frame in real time with no code. We find that Solaris outperforms Frontier LLMs when generating new interfaces. And they have a video attached. It shows, you know, some of the things that you can do with it. Essentially, it kind of generates the world, you know, pixel by pixel as you kind of go through it. So, you almost can like design and interact with interfaces in a interesting way.

23:47 >> I would love to see this world model stuff like actually grow more. We just been hearing different different people like working on it, but it hasn't really like amounted to much on like the the main timeline, >> you know? It's it's one of those things that feels very much still in research. Yes. Right. And there's a big gap between going from research to having significant like commercial usage I think where it'll be more relevant to you know all of us working in the real world outside of the labs. It's one of those things just like you know these video models but even more so like these world models are insane and it would be awesome if there were more practical use cases and maybe there will be someday in gaming and other areas like that but I just haven't seen it yet either other than being the most amazing demos you'll ever see. So this one came out. This from Google research introducing times FM3, a state-of-the-art time series foundation model that enables accurate multivariant time series forecasting in a single forward pass, significantly outperforming other forecasting models across major benchmarks. So if you think of an LLM as predicting the next token, this is much more like predicting the next model or like or the next number.

24:54 It's like looking at numerical trends and then predicting what's going to happen. So, I imagine this could be used for things like website traffic and investment numbers and things like that, right? It's looking just purely candles and stuff. >> Yeah, it's like numerical data. But still pretty cool, >> dude. What if Google is like the finance bro's best friend? >> Hey, you you got to win. They're going to win. >> You need your niche, you know? Everyone needs a niche.

25:19 >> Yeah, everyone's got to find niches. Niches get the riches, right? That's that's what they say. >> That's what they say. Alibaba Zvec team open source ZG local search tool for developers and AI agents. So essentially it's like a a GP style type search tool. It's always good to see new tools being open source for agents. That's it. That's what we got for the news. Did we miss anything? >> I thought that was a lot already.

25:42 >> Yeah, I thought Meta released something too that we maybe missed on here, but yeah, I don't know. If you're in the chat, what did we miss? This was a fun one. I think Astra is is the thing everyone's talking about. I think it'll be what everyone is talking about for a while. I I think the general sentiment that I would say we also experienced is it's good for coding. You know, it feels nice. It does the job. Doesn't feel like a huge step change, but you feel the step change if you give it access to your computer and just let it do its thing.

26:11 >> Yeah. >> And at least I noticed it and it was kind of wild just how it just navigated my computer like it was like even before this right before I went live on the show was like opening up tabs for me and I just have it thing have her running on some stuff. Try it out if you haven't already. You try it in the codeex app and it's it is a pretty impressive experience. I imagine Anthropic is working really hard to get Cloud Desktop to be able to do the same things.

26:33 >> I guess we missed one thing which was Muse AI personal AI assistant which I guess we'll cover next week. >> Yeah, that that was I think the the thing that I I knew we missed. >> The Grockbot from Facebook essentially. >> Yep. So there's always there's always more than we can cover in these things, but we do our best to bring you the news every week. So, if you are tuning in and you're still some for some reason listening, thank you. Go give us that review. Give us that thumbs up on YouTube. Follow us on XMRA on at YouTube master-I. I'm SM Thomas 3 on X and Obby is Abby. And we will see you next time.

27:06 Goodbye.

Summary

The discussion centers around the recent advancements in AI, particularly the launch of Astra, a new model that significantly enhances computer use applications. The hosts express mixed feelings about the film "Artificial," starring Andrew Garfield, and delve into the implications of Astra's capabilities, comparing it to existing models like Claude and Fable.

- Astra's launch video showcased impressive upgrades for computer applications, attracting significant attention.
- The model excels in creative tasks, such as 3D rendering and design, appealing to non-coding technical users.
- The hosts noted Astra's better performance in coding tasks compared to previous models, despite some initial rollout issues.
- Benchmarks indicate Astra's leading position in AI capabilities, sparking discussions about AGI and its implications.
- Concerns about AI regulation and the potential for a pause in development were raised, highlighting political influences on technology.
- Cognition's recent funding round and valuation demonstrate the growing demand for AI-driven software solutions.
- The conversation touched on the competitive landscape among AI models, with Google and Meta also making strides in the field.
- The hosts emphasized the importance of practical applications for new AI technologies, particularly in areas like finance and design.

Questions Answered

What are the hosts' thoughts on the current writing styles in AI?

The hosts express their dissatisfaction with the repetitive writing styles produced by AI, particularly Claude, and look forward to a change.

How does Astra compare to previous AI models in coding tasks?

While Astra shows improvement over previous models, it still makes common mistakes and doesn't feel drastically better overall.

What are the implications of recent benchmarks on AGI?

Recent benchmarks suggest that AGI is closer than ever, with Astra outperforming previous models significantly.

What recent achievement did OpenAI accomplish regarding a significant mathematical problem?

OpenAI's model provided a solution to the Navier Stokes Millennium Prize problem, showcasing its advanced capabilities.

How are different AI models performing in the competitive landscape?

Quen 38 Max has emerged as a leader in web development benchmarks, indicating a shift in focus among AI models.

© transcribe · For agents Built with care and craft by Gokul Rajaram