transcribe

OpenAI Just Shipped The Best AI In The World (Better Than Mythos)

Limitless Podcast · 33m · transcribed Jun 2026
More from Limitless Podcast Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:00 The most powerful model in the world is here right now. In fact, it's so good that it beats Claude Mythos. [music] OpenAI just released ChatGPT 5.5 and it crushes Claude on every single benchmark. It's the [music] new number one coding model. It can do 20-hour tasks that expert software engineers sometimes can't do. It's already discovered groundbreaking solutions in maths and frontier sciences just such as genetics. And it's cheaper than GPT 5.4. This is the result of 2 years worth of frontier research released in this one single model. In fact, it's so good that an Nvidia engineer said, and I quote, "Losing access to GPT 5.5 feels like I've had a limb amputated."

0:41 >> I think a lot of people are going to compare this to Opus 4.7 and that's fair, but I really think the the true comparison is to Mythos because Sam Altman recently, he just posted something as the model was coming out that felt very much like a jab at Mythos. And we're going to get into the benchmarks comparing them, many of which will actually beat the Claude model. But what I find most interesting about this post is the second paragraph where he says, "We believe in democratization."

1:03 And he mentions specifically, "We have been tracking cybersecurity as a preparedness category for a long time and have built mitigations we believe in that enable us to make capable models broadly available." So this is very much a dig at Mythos, which is, as we all know, privately available, only gated to the companies that are given allowance to it. ChatGPT and OpenAI are like, "Hey, we're going to give you the powerful cybersecurity. We're just going to bake in the precautions into the model so that everyone can have it." And it it ends by saying it's this really sweet thing. It's like, "We we love you and we want you to win. We believe in everyone having access to this intelligence." And I really respect that. And I think it's an awesome way to set the precedence for what the next generation of these models is going to look like. But before we go any further, let's talk about the model itself. It's out right now. If you have a ChatGPT membership, you can go and use it. Go and play with it. You guys, what's the TLDR? What are the high-level things that everyone should know? What's most new and noteworthy about GPT 5.5? Okay, so inspired by your Mythos comparison, the first question that popped into my head is, I use Claude Opus 4.7 every day. So I'm like, is it better than this? Like should should I be switching back to to ChatGPT right now? The answer might be yes. So if we look at the benchmark score right here, GPT 5.5 on the left over here absolutely crushes all the standard benchmarks that these frontier models are weighted against.

2:17 And if you look on the right over here, Claude Opus 4.7, it either doesn't even measure in a particular category or it's completely beaten by GPT 5.5. In fact, the only stat that GPT 5.5 doesn't beat Opus 4.7 in is something called software engineering uh benchmark verified pro or something like that. It's like the pro software coding situation. Um but there's a footnote at the bottom of this blog where OpenAI states, "Anthropic has publicly said that they might have gamed that particular benchmark and they need to be reevaluated." So we might have a complete clean sweep for 5.5 as we see it today. So it's an incredibly powerful model, but a question that popped to my head is, does it actually beat Mythos?

2:58 And we have a direct comparison right here. Yeah, so it shows that it does across some benchmarks. Now again, these benchmarks are pretty fuzzy. We don't know which ones are gamed to do what, but there is a world in which GPT 5.5 will outperform Mythos on some things. Which ones we're not entirely sure. I think as we kind of figure out ways to describe GPT 5.5, it seems as if it's their like first attempt at making a model built for autonomy instead of answers. I think a lot of the benchmarks that they're working on is in agentic coding, things like it handles tasks that are 20 hours long. We'll get into that. It's doing 85% of OpenAI's internal work already. And it also helped rewrite the infrastructure that built it. There is this this amazing quote in the blog post that said, "OpenAI says 5.5 itself helped optimize the stack that serves it. Codex analyzed weeks of production traffic and wrote custom heuristics for load balancing that boosted token generation speed by over 20%." So they're using the model to actually build the model and make it maximally efficient based on the data that it's collected from users like us who are interacting with the model on a daily basis. So it's very smart. It's very clever. It's not just there to give you answers. It's there to think deeply and actually solve problems for you in a way that I think Mythos and a lot of these other frontier models are kind of pivoting towards now. The great thing about this model release is it reveals a few things that OpenAI has as an advantage against, say, a frontier lab like Anthropic. Like it's clear looking at these benchmarks compared to Mythos, which by the way, the entire world is spiraling because of this model because it's going to like have the cybersecurity ability to take over any kind of government system. Uh it's this model is pretty close and Sam is going to be releasing this publicly or OpenAI is going to be releasing it publicly for everyone to use. So a question that popped to my head is, does this mean that it's a matter of compute and OpenAI just simply has more of them? Certainly if you compare um Sam Altman's ability to acquire compute and spend all these trillions of dollars to acquire it um versus Anthropic, Anthropic has been extremely conservative and now they're struggling. Like you know, they recently signed a $5 billion deal with Amazon, which we'll get to later on. But the point is um this is a tale of two stories. Either OpenAI has enough compute and they're about to leapfrog Claude because of that and they're proving that through this model that is a a very good answer to Mythos. Or, and this is the alternative side, Anthropic's Mythos model is just plainly better than 5.5 and these benchmarks aren't actually verified, which is technically kind of true because I don't know how official these things are. These are just through tests that a small set of users have done. So it it's a game of both. I'm sure Anthropic is watching this and thinking, "Hmm, maybe we should roll out Mythos."

5:29 But they don't have the compute. Yeah, they don't have the inference. In fact, speaking of the inference, Sam actually made a post saying that he's really excellent work by the inference team to serve this model so efficiently. He wants to really highlight the fact that to a significant degree, they have become an AI inference company now. And I think that's a really big difference than what was previously stated. Like Anthropic has a really tough time serving compute and we see that. And even if they had Mythos available in a way that was safe, they can't serve it.

5:54 OpenAI can. And we see it reflected in pricing because I mean, we have some pricing for this model, right? And it seems as if it's roughly at par with 4.7 if not slightly better. Slightly it's slightly more expensive, but not by much. So for every million tokens input, it's about the same for Anthropic Opus 4.7 and GPT 5.5. It's $5 in, but the output is $30 for 5.5 per million tokens and $25 per million tokens for 4.7.

6:22 >> more expensive. >> So it's a little more expensive, but here's where you actually have more of a bargain using the more expensive model 5.5. It is cheaper than GPT 5.4 and it uses tokens way more efficiently to think. So what does that mean? If you are an enterprise that wants to, you know, plug in this AI model and not worry about it and just have it power your entire profit engine. Well, you end up losing less tokens so you hit your rate limits in a much slower rate, which means that you end up getting more bang for your buck as long as you use the model like 24/7 or you use it effectively. If you are just kind of out there using 5.5 to like ask questions that you should maybe be asking Google, this is probably not the model for you, but otherwise it's super powerful one.

7:07 Yeah, and if these prices don't mean anything to you, that's fine. As long as you have a $20 monthly subscription. In fact, this is going to be available to free users fairly soon, I believe. But anyone who is a subscriber has access to this. You don't need to use the API. There's nothing fancy. You open up your app on your phone, you go to the web browser, it's there, it's available, ready to go. Now, there's a few interesting things that you can do with this model that haven't previously been possible. And although we don't quite have access to it just yet, we're recording this right as the model got launched, we do have a blog post from OpenAI themselves who are showcasing a few demos. So again, take these with a grain of salt. These are straight from OpenAI, but they are seemingly pretty impressive and pretty noteworthy as to what they're capable of doing. Starting with this space mission application, which is pretty cool and very reminiscent of the moon mission that we just had. Yeah, so if you guys don't know, Josh has a secret. He has many secrets on this show. One is he's a massive space fan. And when he's not hanging out with me, he's doing a space simulations on whatever he can do, right? Well, okay, maybe maybe part of that is a bit of a lie, but with this new app that we're seeing in front of us right now, this was completely live coded using 5.5 and it's used to simulate a specific space mission. Now if this looks very similar, it's because we just had a space mission for the first time we visited or went back to the moon in 53 years, pretty big deal.

8:22 And we can see a pretty accurate simulation going on right here. So as you can see, there's various different toggles. The physics of the entire thing is very important. And that's another point I want to make about this model. It is being used for frontier research not just in AI, but in mathematics, in genetics. Like it made frontier progression on both of these fronts. And so what we're showing here is this is a model that goes way beyond just text and telling you what could be. It actually implements this into a lot of different things and understands the world around it, which is extremely powerful. But we have another one here. We have a we have an earthquake tracker.

8:56 For anyone who wants to make websites, it's so good at making websites. And this is actually one of the strong suits. In this case, there's a few things to highlight on this earthquake tracker. One of them being that it's one just like a pretty elegantly designed website, but two, all of the graphics are interactive. You'll notice that they update dynamically as you hover over them, as you click. It looks very clean. I assume that it is pulling up-to-date information from an API somewhere that it set up. It is just truly competent and capable of doing these kind of longer tail tasks that are a bit more complicated than a static landing page, but have dynamic data, have the richness that you would expect from a high-end, high-quality, polished website, except just built with an AI model from someone who doesn't need to know anything about coding at all. And then for the gamers also, there's another great example of a dungeon game, which is they're they're describing it as a playable 3D dungeon arena prototype built with Codex and GPT models. Now I think this is something novel to this setup where Codex handles the game architecture, the combat systems, the enemy encounters, and then the character models, the character textures and animations, those were created with third-party asset generation tools using something like ImageGen 2.0. So, this is also one of the earlier signs where you can actually merge a lot of these tools together to build something dynamic in a way that you previously couldn't have done before. Yet, actually the quality of this game looks like something out of uh League of Legends or something like that. I think that's what it reminds me of. Like the these games are getting way more high-def than I expected. I know it's just like it's pretty basic for anyone that's watching this. They can kind of like pick with the finer eye, but it's cool. But, for those of you who prefer like the more traditional side of games, um this might be something that you can kind of vibe code in a couple of minutes. Now, it may look basic, uh but theoretically this is like uh a 3D specially aware game. And and that's not something that you could achieve at least very easily with previous models.

10:40 What I love about this as well is it's also they've also created or in included the prompt for all of these things. So, this is something that you can try right now. Like look at this. And the prompt is no more than like what's this? 1 2 3 4 like 12 lines, dude. 12 lines, dude. And you can have like a fully functioning game. You can probably then uh add an extra step or extra prompt saying, "Hey, can you deploy this to Vercel?" And send that to your friends. Now, you can use you have a you have a game. You're a game creator. You're a game developer.

11:05 So, the applications for this model uh cannot be understated. Um I'm going to be very honest. I thought this model was going to be just an iterative upgrade. I didn't think it would get anywhere near Claude Mythos. Uh two stories have now revealed themselves, which is one, it's the answer to Claude Mythos, and two, it's really damn good. I am now convinced that compute is everything, but not in the way that I thought it would be useful. I thought it would be uh hugely for largely for pre-training.

11:31 Um but, to Sam's tweet earlier on and also in Greg Brockman, the president of uh OpenAI's recent interview, they're going all in on inference, test time compute, which just means that if [clears throat] you have more compute and if you have a good enough model, it can do the thing. This thing, like I said, built itself. It's a self-improving model. Very, very impressive. It's good for solving hard problems. It's good for thinking for a long time. In fact, they marketed it as a model that can now think for 20 hours coherently. Straight. Which is almost a full day it can work on a problem. And what you're noticing from this prompt that's on screen is it doesn't take that much to get it going. You don't need to kind of spoon-feed it all the way through anymore. It can make decisions on its own. It can infer conclusions on what you want just based on the the knowledge architecture that it currently has. It's amazingly impressive. In fact, one of the people who got access to it early just posted on X that he's uh posting live as his um prompt is 7 hours into its task. It has been running for over 7 hours. He said, "This is literally never happened before. The models would maybe run for 30 minutes or so." Wow. Or or if you really shout at them after two to three hours. But, he's on 7 plus hours. I think this is going to be fun for people with complicated things. If you really want to make a triple-A feeling video game or a simulator or a really complex website, this is the model to try out and to use it with Codex and see how all these things kind of piece together. It's really I mean, I wasn't I didn't have my hopes very high based on the Opus 4.7 to 4.6 incremental improvement.

12:54 >> Sam. This seems like a very solid improvement over 5.4. Absolutely. And listen, if you are listening to this and you're like, "Listen, I'm not a gamer. I can't I I waste my time with that. I I focus on more serious things." Well, for you serious people, uh if you're a manager at a top company or whatever that might be, this isn't just a toy or a model used for code. As a lot of the examples that we just gave are around coding. You can use this for just admin stuff or managerial work. Like the capability of this model to think more strategically and long-term and understand the context task that you're working towards. Like like we said earlier, for coding specifically, it can work on 20-hour-long expert tasks. That also applies for administrative stuff or things that are more generalized white-collar worker work. Uh and so, in this example, Noam Brown says, "I'm a manager of OpenAI, but like I'm using this model to basically manage my entire team and make sure we're focused on the right things. And guess what? The output of this team and this product has been pretty amazing." So, all around really excellent work by uh the entire team and the infra team specifically as Sam Altman says here. And yeah, I'm I'm looking forward to using this thing. I don't have access to it right now. I've I've refreshed my uh account probably like five times at this point and it hasn't appeared. So, maybe it's like a slow roll out. But, if you're listening to this and you've tried it out, um let us know what you're using it for. Let us know what amazes you. Like I really want to hear more.

14:12 Yeah, OpenAI's had a pretty incredible week. And this comes on the back of their new ImageGen model that they just released, which was also unbelievable. If you haven't seen that episode, we just recorded it yesterday. So, I would go advise you to see because oh my god, it is amazing. We also recorded an episode on Apple's new CEO this week and what that means to the company as well as the hardware race and how this I mean, this model Opus I'm not Opus. This is GPT. GPT 5.5 is very much part of the AGI class of models that is built on Blackwell chips. And we've recorded an entire episode all about that. Very interesting, very fascinating. Also interesting and fascinating because as always, this is the weekly round-up. We have a few other topics to talk about.

14:49 We have some news out of SpaceX, which is a pseudo acquisition. Now, they haven't quite acquired Cursor being the company in question, but they have at least partnered with them with the option to buy Cursor for either $60 or pay $10 for the right to actually work together. This seems like a big deal. This seems like I mean, XAI we could call it SpaceX, but SpaceX AI is taking AI very seriously. They're currently behind. They clearly don't want to be behind. This is a huge step and a huge kind of trust of support in Cursor with this minimum of $10 into accelerating their progress and trying to get themselves into this game.

15:24 This is actually a genius deal, and there are a few stories why it makes that so. So, so let me explain. If you're SpaceX AI, which by the way is a ridiculous name now. Like what what Just call them XAI. Um you are currently harboring 1 to 1.5 million of the Frontier GPUs, mainly Nvidia, in a warehouse. There's one issue. You're not really utilizing all of it because XAI has had a bit of a slow start to training their models. What's the genius idea? Hmm.

15:55 If I rent those out to another company to train their own model, then we can make money from that. Okay. So, that's win number one for SpaceX. But, then they've thought of another thing, which is huh, Grok isn't really good at coding, and we are losing the race every single day we don't update our model at coding because Anthropic and ChatGPT 5.5 is completely running away with it. So, how do they leapfrog and get ahead? They should acquire the company that is using their own GPUs to train a frontier coding model. So, then the question becomes, "Well, who the hell is Cursor?

16:29 What What's the moat that they have? Like why do they have a good shot of training a better coding model than Anthropic and GPT 5.5? Aren't those two companies way ahead?" Well, the answer is not quite so. Cursor for the longest time was the number one platform and tool for people to use to do their vibe coding. Why? Not only did they have access to frontier coding models from Claude and ChatGPT, they also had something called an agent harness. Now, you'll notice in GPT 5.5, it's really good at coding because of something called agentic coding. That is something that Cursor pretty much pioneered. It It's basically the harness, the prompts, the environment that they mold the model or rather that they mold around the model that makes it so good and intuitive and remembers the context across every single project. Like menial things. Like um understanding your GitHub branches and working on separate flows at the same time. A lot of the top software engineers in the world right now use tools like Cursor and agentic coding to be able to pull this off. So, Elon Musk thought, "Hmm. If I give you the GPUs to train a better coding model, which puts gives you a better product, I should have the option to acquire you.

17:36 In acquiring you, I can integrate you with Grok, and Grok somehow becomes the number one coding model over the next year or so depending on if this deal goes. And if the deal falls through and they create a really bad model, well, you pay me $10 for the service." Or I pay you $10 for the service. Not a bad deal. Yeah, it seems like they're they're going to be continuing to work with other companies to accelerate in places that they're weak at currently.

17:59 Because I mean, they they're so strong at building out the hardware and creating these huge data centers. They need someone who can take advantage of all those GPUs. Hopefully, this will help serve that cause. And that's not the only SpaceX news this week. The other is that they have officially filed an S-1, which for those who are not familiar, it means they're going public. It's officially official 100%. They will be going public this year. If there were any doubts, please let them be relinquished. Here we have it. SpaceX will be going public. The most interesting thing from this was I think the share structure of how they're going to be organizing this for Daddy Elon, who's going to be getting quite a big payday if he does well. So, we have on screen here just a series of some of the financials. I mean, we know Starlink as a business has been doing unbelievable.

18:40 They have about $25 in cash, 92 billion assets, 50 billion in liabilities. Dude, that's quite a lot of liabilities on this. My god. They They got a lot of debt, man. I don't know. We'll see we'll see once they finally publish everything. I'm very excited for the first earnings report where you really get a true peek behind the scenes of what's going on there. But, it looks like it's going to be going public at a $1.75 trillion valuation. Now, in terms of pay structure, Elon is posed to get 60 million shares, which is 11 tranches vesting in $500 market cap increments from $1.1 trillion to $6.6 trillion share price.

19:17 Um so, for those unfamiliar with the current ceiling, I think it's Nvidia. Nvidia is what? 5 trillion? Under 5 trillion? Close to 5 It's like 4.3. Okay, so not even close. They're like 20% away from 5 trillion. SpaceX needs to be what is that? Like 20-something percent more valuable than the most valuable company in the world. But, if they do, Elon gets 60 million shares. Now, I haven't done the math on exactly how much that is. Um but, if we we make some assumptions here, the total value at vest looks like it could be about a quarter of a trillion dollars.

19:49 So, pretty good payday for Elon. I think the most important thing is that he's getting a lot of control over this. It seems as if he's going to have 40-something percent control of the company, which is really ultimately what was most important to him as they went public. So, really exciting news. I am hopeful that it happens this June, which we can expect. And it's without a shadow of a doubt going to be the largest IPO in history. I think everyone's going to be talking about it. There is a new vehicle in which some people are investing. I'm actually going to have the founder on the show soon. So, keep an eye out for that one. Yes. Yeah, the space is This is very exciting. Now, in the world of AI hardware, many people think that Nvidia has has run away with the win. And, you know, you could argue that with a $4.3 trillion market cap, not many people are competing. Except, there is one company, Google. Now, you might be thinking Google does all my search engines stuff.

20:36 Well, Google is the only vertically integrated MAG 7 company that is involved or has a frontier capability at every single layer of the AI stack. Now, right at the bottom are these things called Google TPUs, tensor processing units. And they're their version of the GPU. In fact, fun fact, uh Google's Gemini models has never trained on an Nvidia GPU. It's all been their own internal warehouse infrastructure. And they've been working on this thing for 10 years. Now, just today, or rather this week, they released their latest generation of TPUs, the TPU 8T and the TPU 8I. Now, the TPU 8T, T stands for training or pre-training. It is highly optimized for the pre-training part of an AI model. So, this is like the bulk, arguably the more expensive part of training a model. It's like teaching it like, "Hey, these are words. This These are the general fundamental set of facts that you need to know before you can We can kind of like put you out into the world and present you to our users." TPU 8I is specialized or hyper-specialized in inference specifically. Now, the important part about inference is it's being used for so many different things.

21:46 Number one, it's to answer all your different prompts. Whenever you write a prompt and you submit it to an AI model, it is known as inference. It's getting inference. It needs to query the model and make sure it like does the right types of thinking and gives you the right answer. But, the other part of inference is post-training, where a lot of people train the model and then they do more training after the fact by using it to help the model reason and think of other alternative facts before it presents you the actual answer. And that's what that second TPU is. Now, Google's TPUs have been used extensively. In fact, their largest customer is a little-known AI lab known as Anthropic, which currently runs 1.5 million TPUs. So, the argument can be made that TPUs are largely responsible for Claude's and Opus's success. So, very impressive all around, but there's some other facts about this, right?

22:32 Yeah, well, I love that the dual architecture training set up that they have here, being hyper-specific. I mean, the 8T chip in particular, it's built to reduce frontier model development cycles, they said, from months to weeks. And then we have the 8I, which is the reasoning engine, which is specifically served for agentic use to deliver tokens really quick, as fast as possible. And as we know, Anthropic is working closely with them. And also, I mean, Google is making these for themselves. So, I think whoever is working with Google, whoever's kind of focused on these accelerators is probably in for a nice little windfall as it relates to increased velocity of the training and also increased ability to distribute these models. As we know, Anthropic is having a very difficult time with this.

23:11 Now, Nvidia and Jensen are probably feeling a little shook. They got to be feeling a little bit of pressure here. And it seems as if that's why they're pushing to be open source, because if you are a in a closed-source world where everyone is making closed-source models on their own architecture, then the Nvidia edge very quickly disappears. And I mean, I'm looking at these chips in hand. They look beautiful. They're ready to be They're taped out, ready to be manufactured. And I I think you could start getting kind of excited about this new world of accelerated hardware. And we're seeing this happen again and again, because Amazon just made another big investment in who else other than Anthropic. And the deal, I think, is like This has to be close to a record deal. They're owning a tremendous amount of company now. Yep. Uh so, the news here is Amazon announced they're investing $5 billion into Anthropic. Viewers, Anthropic, they've just raised $5 billion.

23:59 Congrats. Um and so, [laughter] the reason why this is important is Well, there's a few reasons. Number one, Anthropic knows that they don't have enough compute. The argument could be made that's why Claude 3 has hasn't been rolled out. Well, hey, hey presto, now you have $5 billion worth more of compute. Uh now, for those of you who didn't know, Amazon is a primary investor already in Anthropic. Before this announcement, they owned around 17% of Anthropic. After this announcement, it's closer to 20%. So, we're talking about one company that's publicly tradeable right now that owns a fifth.

24:32 Is that Is my math right? Yeah, a fifth of the world's leading AI lab, which is pretty crazy. Now, if we look into the stats of this, this is a 5 gigawatt deal, which is more than any single data center that is currently live. It's It's It's actually a multiple of of five. I think uh SpaceX AI's Colossus 2 is the largest right now with their 1 million TBs. So, it's it's going to be 5x larger than the average uh data center that we're seeing right now for AI specifically. And they're aiming to get 1 gigawatt online by the end of the year. Now, the reason why this is so good for both teams is Anthropic already has a close relationship with AWS and Amazon's cloud computing department. So, spinning up more compute clusters is going to be so easy for them. They have a working relationship. They're used to training Claude models on this. So, it shouldn't be too hard to ramp this up. If you're Amazon, hey, welcome back. That $5 billion is going to come right back to you. So, I don't know what kind of like circular economy this is, but it's back and it's very impressive for them. Is it ironic that today Amazon hit an all-time high? Oh, good. Maybe. Maybe not.

25:33 >> I'm holding the stock. I got the stock. I'm holding it. >> Clearly Clearly, they're doing something right. Amazon is a phenomenal company. They're the largest shareholder in Anthropic. It's hard not to be bullish on them. It's hard not to be bullish on the accelerated computing stack. And I think that's probably what Jensen is getting nervous about. That's why Nvidia is pushing open source. And the good news is is he has some help. He has some assistance from the folks overseas in China who have been pumping out unbelievable models all week long as it relates to Kimi and Qwen, our Chinese favorites. Uh we have Kimi K2.6 and Qwen 3.6. There's a lot of digits and numbers. All you need to know is that the best open-source models in the world didn't exist last week. They now exist this week. And they are better at pretty much everything, but exceptional at coding. In fact, word on the street is that some of these models are as good as GPT-5.4 was. And only a few points off of Claude. I mean, these are pretty amazing open-source models that, again, are free to run locally on your machine if you have the machine capability of doing so. That's a big This is a big game-changer. Okay, so typically the story we tell with these open-source models is "Wow, aren't they so amazing?"

26:36 Yeah, they're the good younger brother. They're not as good as the frontier AI labs. That completely changed this week. So, Kimi K2.6 is the latest model from uh Chinese lab called Moonshot Labs. I believe it's Moonshot or Moonshot AI. Um and they released their model, which ends up being as good as coding or at coding as Opus 4.7. And it's 100% open source, like you mentioned, Josh, which means that maybe you could run this on a local device. Now, the answer that you would typically get back from this is "Hey, like listen, it's it's too large to run on my laptop." And that is true.

27:07 But, with the latest Qwen model, which is a 3.6 version, you can run it as an 18 GB sized model, slightly quantized, on your laptop today. So, the point that I want to make about these models isn't exactly the specifics, but across all benchmarks, it's not as good as the frontier AI labs, but it's a few points. That difference and gap has closed massively over the last couple of months, which tells me two things. Number one, China has figured out some kind of groundbreaking way to train their models that they haven't told the West about, and they're going to keep it close guarded and eventually close source their model releases going forwards. And number two, they've figured out a new way to use inference to their benefit. Like, one thing I'm going to highlight here is this new Kimi K2.6 model can code continuously for 12 hours straight using 300 agents. So, the unlock here isn't one model itself. It's spinning up 300 versions of itself and getting it to attack the problem. That's something Sam realized and what he's implementing in 5.5. That's something Opus 4.7 realized and is doing probably similarly with Mythos. So, I have this question here, which is like, "How the heck did China do this?" Well, I think every 3 months if there's a new open model that gets released, they're making these jumps because they're using these models to train themselves. They proved that with Kimi K2.5. There's too many 2.

28:22 whatevers. Um and the same thing is happening with Qwen. It's just all around pretty amazing stuff. Yeah, China's crushing. Okay, so before we go, we have two quick things to hit. The first being one that we missed last week, which we need to touch on quickly. Anthropic has a design tool now. If you are a designer, if you are interested in building web pages, videos, graphics, slideshows, pitch decks, any type of visual asset, Claude now has an entire design suite built just for this purpose. It's called Claude Design. It exists separately. You can access it through the desktop app or on your browser. And it basically allows you to build visual assets in a way that you couldn't previously. Previously with Claude, you had artifacts. An artifact you could generate something dynamic. It could kind of build you a web page. This takes it to a whole new level. You could generate wireframes if you want to try to use less tokens. You could fill it out and create properly created prototypes that are actually clickable.

29:12 It's amazing. The video we're seeing on screen highlights a few of them. Unfortunately, there was a big loser in this, because this sounds like a lot of what that little design company named Figma does. >> Yeah, the little company. The stock market did not love the reaction to that, did it? Nope. Nope. It is down almost 20% on the week. I actually tracked the stock price after the announcement was made. So, like, it wasn't even readily available. It was literally just the tweet. 20 minutes after it was tweeted, the stock was down 6%. So, the the point being, whether this is market speculation or not, like listen, Claude Design isn't as good as Figma. They're working with a few of these different partners such as Canva, but two weeks ago, one of Anthropic's former most execs left the board of Figma, and the rumors was because they were building a competitor. So, it's pretty clear Anthropic is going off to every single sector, whether you're a designer, a software engineer, a mathematician, a research scientist, it doesn't matter. They're going off to everything cuz the model is applicable to everything, and I don't know what this means for certain modes that companies like Figma holds, but it's certainly going to affect stock price.

30:14 Can you do me a favor and click the max button real quick for me just to show the chart? Oh. >> [laughter] >> Yeah, minus 86% since IPO for those who are not watching on screen. It's pretty been pretty bad rough run for Figma. We have to start naming Anthropic the the stock killer, Josh. This is like every single tweet is killing your stock. Okay. How good How good is your accent or impersonation of the of your president of our president, Josh? Pretty hard. Not good. Okay, well, I then we're not going to accept it. I'd like to hear your British take on it though if you're if you're feeling ambitious. Okay, so my my British take on this is this is albeit hilarious and somewhat terrifying that the president of the United States is saying this.

31:00 He commented, okay, on the government's relationship with Anthropic. Now, if you're wondering why on earth he's commenting on it, they're going to be releasing this Claude Metis model. It might be a security risk. It's probably good for the government to have access to this thing and prepare necessarily. The government has been having very important conversations with bankers and governments all around the world to just try and figure out, you know, how best to prepare for this.

31:20 And after having an in-depth discussion with Dario Amodei, which by the way, he blacklisted that CEO and Anthropic entirely from the government using it, he's now rekindling it and saying, "Maybe there's a deal on the line." He goes, and I quote, I'm not going to do the accent, "We'll get along with Anthropic just fine." Trump said on CNN. Okay, we'll get along with Anthropic just fine. I think they could be of great use to us. They're high IQ people. Very good. Very good. They tend to be on the left, radical left, but we get along with them.

31:49 I don't know. That's all I got. But that's that is what you said. Were you practicing that? That's That was actually pretty good. I closed my eyes whilst you were doing that whilst I was laughing and that That was actually pretty good. It sounded like him. Good. It channeled his spirit. It was It was there. It was a good effort. Um but I believe that's it. That is the end of the round up. Josh and I Josh and I are recording this, FYI, it's 4:00 p.m. over here.

32:11 Typically, we're morning birds. We deliver this in the morning, but we waited for the announcement of Spud GPT 5.5 just for you guys, and we're going to be bringing you the cutting edge news every single week. As Josh mentioned, we had three other amazing episodes that we filmed earlier this week. Definitely go check them out. They're all each 20 minutes long. It's your commute to work. It's your gym session if you're not that active. Definitely go check it out and and let us know what you think. But yeah, Josh, any final thoughts?

32:36 >> Call me crazy, but I like the afternoon recordings. I got good energy. I'm like woken up. I'm 100% right now. I'm rocking and rolling. I'm feeling good, so I don't know. Maybe we'll have to lean into this a little bit more, but that's everything. If you've made it this far, if you're still listening to this and you've heard our other episodes, you're caught up. You're done for the week. You can go touch grass. Enjoy your weekend. There will be a lot more to talk about next weekend, but for now, you have fully synchronized with all of the chaos happening on the frontier of AI and technology. Thank you so much for watching. As always, we very much appreciate it. If you enjoyed this episode or any of our previous episodes from this week, don't forget to share them with a friend who you also might enjoy it possibly. We have a newsletter on Substack that goes live twice a week.

33:12 Just went live yesterday. Going live again tomorrow. The Friday issue is a recap of everything that happens this week, which is always fun and exciting. In fact, I'm going to go write that as soon as we finish this episode. So, thank you all for watching. As always, don't forget to [music] subscribe, like, comment, all of the good things, and we will see you guys next week. >> [music]

© transcribe · For agents Built with care and craft by Gokul Rajaram