Section Insights
Introduction to GPT-6 Astra
What is the significance of the release of GPT-6 Astra?
OpenAI's GPT-6 Astra is being touted as a generational leap towards AGI, excelling in various professional and personal assistant tasks. The model is designed to perform tasks quickly and efficiently, positioning itself as a powerful tool for users.
- GPT-6 Astra is a significant advancement in AI technology.
- It is aimed at personal assistant use cases and excels in computer-related tasks.
- OpenAI claims it may represent a step towards AGI.
Initial Reactions to GPT-6
How do experts feel about the implications of GPT-6?
There is a mix of excitement and skepticism regarding GPT-6's capabilities. While benchmarks suggest significant improvements, the true excitement about AGI's arrival is tempered by the practical applications of the model, such as creating presentations or booking tables.
- Initial reactions are mixed, with some excitement about benchmarks.
- The practical applications of AGI are still a focal point of discussion.
- Experts are cautious about fully embracing the AGI label.
The Emotional Aspect of AGI
What metaphor is used to describe the announcement of AGI?
The discussion compares the announcement of AGI to expressing love, suggesting that while there is a strong belief in its arrival, there is also a hesitance to fully declare it without certainty. The metaphor emphasizes the emotional weight of such a declaration.
- The announcement of AGI is likened to a hesitant expression of love.
- There is a desire for clarity and certainty in declaring AGI's arrival.
- The emotional implications of AGI's announcement are significant.
The Importance of AGI as a Goal
Why is it important to hold AGI as a goal?
Holding AGI as a distant goal prevents disillusionment with current AI capabilities. If AGI is perceived as merely achieving basic tasks, it could lead to disappointment. The discussion suggests that the pursuit of AGI should remain aspirational.
- AGI should be viewed as a long-term goal to maintain motivation.
- Achieving basic tasks with AGI could lead to disappointment.
- The pursuit of AGI is essential for future advancements.
Benchmark Performance of GPT-6
What do the benchmark scores indicate about GPT-6?
GPT-6 Astra has achieved impressive benchmark scores, outperforming competitors while also reducing API costs. This performance is crucial in the competitive landscape against companies like Anthropic, indicating OpenAI's strong position in the market.
- GPT-6 Astra shows significant benchmark performance improvements.
- Lower API costs enhance its competitive edge.
- The model's performance could impact the valuation and perception of competitors.
Transcript
0:00 Open AI, as you may know, if you're listening to this show, released GPT-6 yesterday. And remember the big wait for GPT-5, when's it coming, and the expectations? This kind of came seemingly out of nowhere. We have GPT-6 Astra, and not only does it crush on the benchmarks, Open AI is saying that it might be AGI, and it's obviously geared towards a lot of sort of personal assistant use cases. So, let me just read a little bit from what The Verge reported on this. So, The Verge says, "Open AI's next big AI model has entered the AGI era." the next big model is here. It's called GPT-6 Astra. The company calls it a generational leap in capability in capability for areas like cybersecurity, professional work, software engineering, science, and computer use. The actual This is I'm just going to read from Open AI's branding here. They say it's GPT-6 Astra. Anything you can do on a computer, Astra can do for you fast. And they, of course, released a, you know, a snazzy video, as they tend to do on these releases, with showing people in front of a big computer asking it to do things for them, like book tables, build a presentation, create legal drafts, and basically the idea here is that their new model is going to excel at computer use, and be able to get things done for you. And it's very interesting that they use I know they they highlighted voice as the interface to get it done, almost like the personal assistant computer in in Star Trek. So, Ranjan, your thoughts about the release of GPT-6? I mean, I'm I'm curious to hear your your initial reaction, but also like we talk a lot about like the fact that even recently I talked about how, you know, it looked like Anthropic was opening up the gap between itself and Open AI. maybe that gap has shrunk or or closed closed completely. What do you think?
1:52 >> All right, let's separate out those two questions. What this means in the the AI race and then first, you know, like is this an exciting launch? Again, I always any of these new model launches try to wait until I've actually had access to it and I'm unfortunately not part of the day break platform and an open AI cybersecurity researcher, so that'll have to wait. >> Right, so those are the people that have gotten initial access.
2:17 >> access. >> There's already a little controversy because like a handful of people can already use it and it's supposed to roll out to everybody else soon, but it hasn't yet. But that that will come in time. Go ahead. >> But it is certainly rolled out to every ex-influencer who has now built some virtual world or recreated a video game or whatever else and has posted about that, but I do love that like my favorite part of the launch announcement was AGI is here. Build new world models like the crushing benchmarks. But create nice decks and book a restaurant table. That it still always comes back to that. I love that the test of AGI in the end is going to be can you actually book a restaurant table or create a good PowerPoint deck?
3:02 I think like I don't know. >> >> Do you do you have an opinion on how big this is already? Are you excited? It's hard for me to try to gauge on that. I think on the benchmark side, I think it's really interesting and I think in the Anthropic context, it's even more interesting, but on is this really exciting? I don't know yet. >> Right, so it's a great question and I'm kind of on two minds about it. So, you know, the way to sort of think about these releases is, you know, I do think to some degree you can't use all of what these companies say about the releases as gospel when they come out, but you can sort of take some signals cuz they are putting their reputation on the line to some degree. And you know, earlier this year I was at OpenAI with Craig Brockman and he said that he thought the company was about 80% of the way towards AGI. very different comments with this new model. So, he says, if we fast forward it a couple years and we look back and say, "When was it really that AGI was created?" I think it's going to be about this time. I think it might be about this model. For me personally, I do think we're there. I think it's not unreasonable to feel that we are now in the AGI era. Okay, I read this and it sort of was like, you you know, you know when you want to tell somebody you love them, but you don't want to like take the risk and, you know, you say something like, "Well, if I knew what love felt, I think this is what it would be." I think that's what Greg Brockman is saying about AGI. Like, I think he's a little fearful about coming out and saying it, but the dude's in love. It's AGI and that's effectively what he's saying in these statements.
4:41 Wait. Sorry, is that describe the entire feeling again? Or what what the statement is? This I I want to I want to work through this scenario quickly. Young lovers, when they're in love, the words I love you are very difficult to say because of the stakes involved. So, you say something and I I'll admit, like, I've been in scenarios like this in my early years when I didn't know anything, where I would, you know, sort of like you have these strong feelings for someone and you dance around it and you're like, "Huh."
5:16 And this won't be foreign to I I I think this won't be foreign to some of our listeners. Where you say, "Huh, I wonder what love feels. Is this it?" Where you really want to say, "I love you." to somebody. And that's I think to to a degree like Greg Brockman is saying, if we fast forward a couple years and we look back to say, "When was it really that AGI was created?" I think it's going to be about this time. It's the same thing. It's the except instead of like a young lover telling the other that they love they love they they love their, you know, person. What Greg is basically saying is this is AGI. Wait, but just to confirm, the first part of that about you're not saying out loud to the other person.
5:53 >> No, you say it out loud. You say that >> If I knew what love felt like, if I knew what love would feel >> You say these things. You say, "I wonder is this is this love?" You know? >> Okay. >> to you, Ranjan? >> I'm trying to think. >> straight out You just when you you straight out say >> Just said it. >> You know, just like matter-of-factly, "Listen, I love you." >> Yeah, that's it is what it is.
6:15 >> I respond with All right. Just imagine telling somebody that and being like, "Listen, I need to tell you something. I love you. It is what it >> what it is. And you know what? I would appreciate if Greg Brockman would just say that. And I think if OpenAI issued a press release and said AGI is here. What's interesting I I read somewhere that every contractual obligation around the term AGI and mainly the Microsoft one does not exist anymore now. So now he should just say I love you. AGI is here. But but but it is even I think they're so trained to because do you know what what to me what actually the greatest danger in the world to OpenAI is?
6:58 Is to say AGI is here and then everyone goes to ChatGPT, types in something and gets a lukewarm response that isn't quite right. And then suddenly I'd actually think that is like a just massive threat to the overall story and hype cycle because the whole the whole beauty of AGI is it's this thing that's dangled in front of us on an ongoing basis to promise this future. So as long as you don't say it's here and you dance around it in a teenage romantic sort of way, it's pretty effective and I did I think that's what's happening here and that's that's why he's hedging. I don't think he thinks it's here.
7:42 >> Ooh, I I really >> Otherwise, he would say it. He's a I mean, these guys like a Greg Brockman, they are believers. I believe they are believers. So, if they believed it, they would say it. It's too important. >> They got I mean, they did get all the headlines, but I think you're right that it is it is worth holding AGI as this sort of like goal that you're never going to reach, holding it out that way because or maybe that's what superintelligence will be at at a certain point. because there was people that were like, you know, if we reached AGI, what do we have to look forward to anymore? And you're right, if it's AGI and it's just like it can't get some stuff done for you, you're going to be like, what was the wait for?
8:21 >> No, think about how like disheartening that would be. You get just kind of like a slop deck with bad formatting and some overlapping like chevrons. >> Damn it, AGI. >> The most The most basic stuff, that you got a video that where the motion isn't quite right, and then that's it. Like, what do we do from there? Then I guess we wait for ASI and superintelligence. Yeah. >> exactly. >> But, do you believe Do you believe he believes it's here? You started the thread with you do believe he wants to say he wants to say it.
8:57 >> Yeah, I do think that he thinks it's there. I just think that, you know, obviously OpenAI is also seeing like one of the ways that you can parse his words is OpenAI is also seeing even more powerful models internally. Of course, we're going to get into the hacking side of things with the Hugging Face situation. but they see this stuff internally. And they're probably saying, "Okay, yeah, we're definitely entering that moment." and you know, you can also even look, and this is sort of the second part of the of the discussion, you can look at some of the benchmarks. remember the Arc AGI test?
9:27 Right? This is sort of like the way to show whether the AI can can generalize. it saturated GP6 GPT-6 Astra saturated the test, scored 99% on the test. And even the Arc-AGI folks were like, well, they're like, this was just one marker, doesn't mean if you, you know, saturate the test, you've reached AGI. It's like, why do you call it the AGI test anyway? But yeah, this is from the OpenAI blog post. Arc-AGI 3 test how well agents learn as they solve unfamiliar interactive tasks, and GPT-6 Astra saturates the eval, scoring 99%. Every human scored 48%.
10:08 All right, so that's kind of like where you start seeing this. You also, I mean, there's a bunch of other evaluations, but, you know, even for doing science, there's this called terminal bench science, 0.1 eval. And GPT-6 Astra scores 64% on scientific research tasks using code and terminal tools, that's what the evaluation test for. Whereas Fable is at 52.6% and OpenAI says Astra hits this higher bench higher mark with 31% lower API costs.
10:45 So, that's what we're looking at benchmark-wise. >> I think that is the yeah, the most like important part of the announcement are those benchmark scores. And I think like, I don't know, again, I'm going to need to use it so I can feel what AGI feels like, but like what in terms of the competition against Anthropic, I actually think this is a very big deal. Like we've already seen over the last two to three months, you know, some major rumblings again, none of this Well, certainly there's been like ramp data, but around Codex starting to close the gap again with Claude code, frontier like OpenAI getting back into the race in a bit.
11:32 So, I think I think especially in the IPO backdrop context, I think this actually anything that kind of creates any doubt on the Anthropic story could be very harmful to them given it's a very tightrope they're walking in terms of that $2 trillion valuation. So, I think in that way if this starts getting rolled out, we all feel magic in a What would you say like What what were the models that made you feel magic?
12:05 GPT-3 certainly. >> I've always been a an 03 guy. I mean the 03 model that like would sort of think and then break everything into tables just showed a leap that you know that that I just hadn't seen for even like the leap between 3.5 to 4 to me you know GPT-3.5 to 4 you know that felt that felt meaningful but nothing is close to as when they introduced reasoning. So, this is sort of like we've gone through like a handful of different phase shifts so to speak you know the initial chat GPT then the reasoning side of thing and now we're in this sort of like computer user harness hive era. So, if that if this can really you know I don't know if you have this this in your life when you use AI but I'm oftentimes like saying I wish you know I could use AI to do X task for me and it succeeds at like 30% of tasks. If it could get to like 80 or 90% that would be a real change in my life.
Summary
- GPT-6 Astra is positioned as a generational leap in AI capabilities, potentially marking the beginning of the AGI era.
- The model excels in benchmarks, scoring 99% on the Arc AGI test, significantly outperforming human scores.
- OpenAI emphasizes practical applications, such as booking tables and creating presentations, highlighting its personal assistant functionalities.
- Initial access to GPT-6 Astra is limited, leading to some controversy regarding its rollout.
- The discussion reflects skepticism about whether the model truly represents AGI, with some believing OpenAI is hedging its claims.
- The competitive landscape with Anthropic is shifting, as GPT-6 Astra may close the gap in AI capabilities.
- The benchmarks indicate a significant improvement in performance and cost-efficiency over previous models.
- The conversation hints at the importance of user experience, with a desire for AI to handle a higher percentage of tasks effectively.
Questions Answered
What is the significance of the release of GPT-6 Astra?
OpenAI's GPT-6 Astra is being touted as a generational leap towards AGI, excelling in various professional and personal assistant tasks. The model is designed to perform tasks quickly and efficiently, positioning itself as a powerful tool for users.
How do experts feel about the implications of GPT-6?
There is a mix of excitement and skepticism regarding GPT-6's capabilities. While benchmarks suggest significant improvements, the true excitement about AGI's arrival is tempered by the practical applications of the model, such as creating presentations or booking tables.
What metaphor is used to describe the announcement of AGI?
The discussion compares the announcement of AGI to expressing love, suggesting that while there is a strong belief in its arrival, there is also a hesitance to fully declare it without certainty. The metaphor emphasizes the emotional weight of such a declaration.
Why is it important to hold AGI as a goal?
Holding AGI as a distant goal prevents disillusionment with current AI capabilities. If AGI is perceived as merely achieving basic tasks, it could lead to disappointment. The discussion suggests that the pursuit of AGI should remain aspirational.
What do the benchmark scores indicate about GPT-6?
GPT-6 Astra has achieved impressive benchmark scores, outperforming competitors while also reducing API costs. This performance is crucial in the competitive landscape against companies like Anthropic, indicating OpenAI's strong position in the market.