transcribe

Grok 4.6 is Claude Fable 5, but dirt cheap

Alex Finn · 13m · transcribed Aug 2026
More from Alex Finn Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Introduction to Gro 4.6

What are the claims about Gro 4.6 and how does it compare to other models?

Gro 4.6 is positioned as a competitive AI model that offers quality comparable to more expensive models like Opus 5 and GBT56 at a lower cost and faster speed. The speaker expresses skepticism about benchmarks but indicates that Gro 4.6 can be used effectively in Cursor and Grock Build.

  • Gro 4.6 claims to match the quality of higher-end models at a fraction of the cost.
  • The speaker prefers using Cursor for a more feature-rich experience.
  • Skepticism exists regarding the validity of benchmarks provided by Gro 4.6.
# 2:39

Benchmarking Gro 4.6 Against Competitors

How does Gro 4.6 perform in comparison to Chat GBT and Fable?

In various tests, Gro 4.6 performed well, recreating the Apple website and completing a scavenger hunt faster and cheaper than Chat GBT. While it took slightly longer in some tests, the cost-effectiveness was highlighted.

  • Gro 4.6 outperformed Chat GBT in cost and efficiency during tests.
  • Visual fidelity of Gro 4.6's outputs was noted as less impressive compared to competitors.
  • Gro 4.6's performance in scavenger hunts was significantly better than Chat GBT.
# 5:19

Comparative Performance with Fable and Opus

How does Gro 4.6 stack up against Fable and Opus in various tests?

Gro 4.6 outperformed Fable in bug finding and was cheaper, although Fable completed tasks faster. In gaming simulations, Gro 4.6 showed strong graphics but had some physics issues compared to Opus.

  • Gro 4.6 is cheaper and competitive with Fable and Opus in performance.
  • Graphics quality of Gro 4.6 was noted as superior, but gameplay mechanics had flaws.
  • Fable's higher cost did not justify its performance in certain tests.
# 7:59

Strengths and Limitations of Gro 4.6

What are the strengths and weaknesses of Gro 4.6 for different types of users?

Gro 4.6 is excellent for coding tasks and offers a fast, cost-effective model, but lacks the intuitive features of general-purpose AI applications like Chat GBT. Users focused on coding may prefer Gro 4.6, while those needing broader AI capabilities might find it lacking.

  • Gro 4.6 excels in coding but is not as versatile for general knowledge work.
  • Cursor is a strong coding tool but lacks the comprehensive features of other AI applications.
  • Grockbot is recommended for knowledge work, though it has limitations.
# 10:39

Future Prospects and Recommendations

What are the future prospects for Gro 4.6 and its applications?

The speaker believes Gro 4.6 is rapidly catching up to competitors like Claude and Chat GBT. While Grockbot is highly recommended, the speaker still prefers Chat GBT for general knowledge tasks. There are concerns about Claude's stagnation and potential improvements in the future.

  • Gro 4.6 is becoming competitive with leading AI models.
  • Grockbot is highlighted as a top choice for users needing a general AI agent.
  • Claude's development appears to have slowed, raising questions about its future.

Transcript

0:00 Gro 4.6 just dropped and it's official. SpaceX AI is now a Frontier AI company. Elon claims you can now get above Chad GBT56 and Opus 5 quality for a fraction of the price and much much faster. Well, I'd put it to the test and today you are going to find out if those claims are true. You're also going to find out if you should be switching to Gro 4.6, how to use Gro 4.6, six in which situations you should be using the model. Now, let's lock in and get into it. So, here are the benchmarks they're giving.

0:34 Unfortunately, I think all benchmarks are fake, so we're not going to spend much time here. Basically, what they're claiming, though, is it's just as good as Opus 5 and 56 Soul for a fraction of the price. They're not claiming it's way better or it's the best model ever made. They're saying it's good as Frontier, but you're getting it for way better costs, which is a good thing because Fable and Opus and 56 Soul are very, very, very expensive models. The two places you can be using it right now are Cursor and Grock Build. I lean towards using Cursor. It is a much more fullfeatured AI vibe coding experience.

1:11 It's basically kind of like a stripped down version of the chat GBTF, but it's very, very good. Grock build is solid, but it's just a CLI. There's no interesting harness features I think you'd come to expect from harnesses in August 2026. I lean much more toward using cursor instead of Grock build if you want to use this model. Now, let's talk about if they came through on their claims. I ran this model through the world famous Alex Finn benchmarking test. For those who don't know, it is five different benchmarks. I put it against chat GBT56 soul. I put it against Fable. Here are the results. It beat 56 Soul in my benchmark. So, it went through these five different tests.

1:53 It did a whole scavenger hunt for code on the internet. It did a whole debugging exercise. It built a bunch of simulators. Basically, what it came down to is it scored better than 56 soul. It did it in significantly less time. Did in 18 minutes rather than 24 minutes with GPT. and it did it for basically a third of the price. Now, let's go through those actual results here. We'll starting off with the first one, which is a roller coaster simulator test. Over on the right, 56. Over on the left is Grock. Let's open this up and see what it looks like. As you can see, it built a very nice roller coaster. That is that is the I think the biggest roller coaster out of all the simulators we've done so far. Let's hit ride and see what this is like. This looks great. The sky looks amazing. I love the kind of twilight colors in the sky. Clouds look nice. Nice. And as you can see, the roller coaster is looking good. If we go over to chat GBT and I expand this, I'd say the fidelity is not quite as good.

2:48 It's not quite as pleasant on the eyes. And if we go right here, it looks it looks fine. It looks all right. It did it for a fraction of the price. It actually took a little bit longer on this test, but I think the results came out well, and you spent way less money on it. Here are a couple of the other tests against Chad GBT. A recreation of the Apple website. So, we hand both models the Apple website. We say, "Hey, recreate it as close to graphical fidelity as you possibly can." This is what the Apple website looked like today. As you can see, is the phone, the Air, all the other devices. Ted Lasso, Sabrina Carpenter. Let's see the recreation from Grock. Not too bad.

3:28 Not too bad. Obviously, it doesn't look nearly as good, but it can't use graphic generation. It has to recreate all the graphics itself. Honestly, not too bad. And then if we go over to GPT, GPT looks all right. I like the header better, but the phones look all kind of messed up from there. It doesn't even look like a laptop. all the things down here don't look too great. So, honestly, it it edged it out. Did it actually a little bit longer, too, but much cheaper. Other than that, we have an agent test where we give it tons and tons of files. Then, we basically have it go on a scavenger hunt. Has to use different tools to find things and CSVs and PDFs and different files. It actually dominated GPT. It did it in a looks like a sixth of the time for about a sixth of the cost and it found all and use the tools better. GBT couldn't even find two of the scavenger hunts. Bug fixing. We hand it a open- source library from GitHub, a recent one. We see how many bugs it could find. Both of them found all 13 bugs in the code.

4:32 Grock did it faster. Did it for about half the cost. GBT took about a minute longer for about a dollar more. And overall, it came out that Grock beat GPT pretty handedly, which is pretty shocking. But how did it handle against Fable? Let's check that out. So, Grock actually ended up beating Fable 5. Now, I will say this, the reason why Grock won is because Fable 5 refused to do the scavenger hunt because of its safety guard rails. If you take out Grock's score from here, it is about even when you take out Grock's score from that one. But at the same time, Fable is not letting you do everything. Grock lets you do everything. Grock beat it at about a tenth of the cost. Did it in a minute quicker. Let's check. Fable beat it on Pixel perfect. So the Apple recreation, let's see that real quick.

5:20 Here is the Fable recreation. Some of these actually look really, really strong. some of the devices. Fable actually did this one faster, but again, four times the price. Bug finding, Fable was four times the cost at double the time. And so Grock ended up beating Fable 5, which is really, really impressive. Now, I did some other tests, too. I put it up against Opus, Fable, and 56 Soul on some of these tests. So, here's a flight simulator. I also put Grock against a test with Opus, as well as Fable and Soul in a few other tests here. So, first we have kind of a flappy birds test here. Here is the Gro one.

5:57 Actually looks really, really nice. Oh, next level. So, you get the stars, you get everything. I like that. It does. It runs the simulation a little bit quick there. Oh, things are just collapsing. All right, so the physics are a little bit off. Let's see here. Graphics looks great, though. Let's see what Opus 5 looks like here. Opus 5 looks pretty good, too. Physics actually work well. I like that a lot. There doesn't appear to be second levels though. You just do level Oh, there we go. Next level. Here we go. All right.

6:25 So, now they got new things. there. Physics are a little off too as well. Just seems to collapse before you even do anything. I'd say graphically Grock probably outd does it a little bit gameplay-wise. Opus wins. And then let's check out Fable 5. Fable 5 appears to be completely broken. And then chat GBT 56 Soul. This looks really good, too. I like the graphics. Let's Let's fling this thing. The physics seem off here and everything looks kind of messed up.

6:52 This is probably the weakest of them all. So, let's see how much it cost, how long it took. 9 minutes for Gro was actually longer than 56. Did it as half the price of 56. By far the cheapest out of all of them, and I think it got the best results. Fable produced the worst results and was the most expensive and took the longest. So, let's revisit those claims. Let's see what was good and bad about Gro 46. I've been using all day. I've tested in cursor Grock build. I showed you all my benchmarks.

7:19 Here's what it comes down to. The positives, all the claims are true. It is basically just as good as Opus and Chad GBT56 while being much cheaper and much faster. Those are all 100% true. Here's the issue though. The issue is there's no kind of great generalpurpose harness for this model. The strength right now of Chad GBT is the Chad GBT desktop harness is incredible. It has everything. It does coding really well, kind of like cursor, but it also does the general purpose stuff where it can do computer use really, really well. It can do browser use really, really well.

8:02 It does everything you need to get all your knowledge work done in one single place. Grock doesn't really have that. Grock has cursor. Cursor is good. Cursor is a very good app, but it is really built for coding. Yes, it can do some more generalpurpose stuff, but it's not as strong or as intuitive as a lot of stuff in the chat GBT app. If you are looking to do pure coding, if you're all about vibe coding and that's it. You don't do kind of AI for all of your knowledge work, then this actually is probably the best choice for you. you're going to get an excellent model at an excellent price that's super fast and you can use cursor for all your coding and cursor is excellent at coding. The challenge is it's just not a very great generalpurpose AI agent harness. If you want to do knowledge work with Grock, the best option is Grockbot. Grockbot is a fantastic app. I did a review on it yesterday. You should check that out if you haven't yet. But it's more for your kind of knowledge work. It's not as good for coding and it's a little bit limited with local computer and browser use, but it is very very good. So that is the challenge for Gro right now. It is a great model. It is cheap and it is fast, but it doesn't have that amazing incredible harness just yet. Like Chad GBT desktop app is clawed to a lesser extent. The Claude desktop app is really really good. I wouldn't be surprised if they completely rebrand cursor to like Grock agent or Gro desktop very very soon. The Cursor team built Grockbot and they didn't call it Cursorbot, they call it a Grockbot. So I wouldn't be surprised if they repurpose Cursor to be Grock desktop very soon to be kind of the chat GBT desktop app for SpaceX AI, you know. And that doesn't even mention a lot of the other nice to have tools that Chad GBT and Claude has like Chad GBT voice, Claude voice. Two revolutionary technologies. It's kind of missing from Cursor. The mobile app, right? Cursor just came out with their mobile app. It's very good, but it's not quite the chat GBT and clawed mobile apps. So, who is Gro for? Who should be using it? Well, I think a lot of people should be using it. I think if your workflow is primarily vibe coding, this is kind of the best pure vibe coding model there is right now. It can build just as good as the other models, but do it for way faster, way less cost. If you are an AI power user and you need voice while you're on the go so you can talk to your agent and build things on the go and you want to be, you know, by the pool and you want to send a command to your agent to do and tinker on your computer and do a bunch of things like you're a power user like that, you need kind of a power user harness. Cursor is not quite there yet. Chad GBT is there.

10:51 Claude is basically there. But I will say this, what SpaceX AI has pulled off over the last several months is miraculous. Grock was by far way behind Claude and Chad GBT. Now they are neck andneck. They are right there. So that wouldn't shock me if again cursor turns into Grock desktop and then it's just as good as the other desktop tools and is just as good of a harness. Grockbot absolutely amazing. Everyone should be using Grockbot. I I told you this is like the best kind of for the normie AI agent app I've ever used. It's excellent. That should be used here as well with Grock. I'd still use Chad GBT if you do a ton of general purpose knowledge work and you're a power user that need AI using every device you have. Claude, I'm still using Claude basically just for Fable 5 business planning. I think from a highlevel business strategy perspective, Fable 5 is still the best model. I don't use Claude for literally anything else at all. Claude has basically fallen behind on all levels. I really feel like this desktop app hasn't changed in a very long time. It feels like it's been absolutely months since anything has changed in this app. I'm not sure what's going on with the Claude team. It seems like releases has slowed down dramatically over the last few months. I wonder if it's because of their compute limitations. They just signed a deal with SpaceX AAI to get a whole bunch of compute. Maybe they're getting that onboarded and will be able to release quicker in the future. I wouldn't be surprised if they just have an explosion of releases over the next few weeks. I wouldn't be surprised we get Fable 51 very soon. There's a lot of rumors around that and I'm sure this whole competition flips on its head when that comes out as well. But Grock 46 pound-for-pound is probably the best vibe coding model out there right now.

12:43 If you're tight on money or you're just purely about vibe coding, there's no better model to be using. I do it in Cursor. Again, get Cursor, use it in there. It's great inside Cursor. For me personally, I'll still be using Chad GBT and Claude as well because I have so many of those general purpose use cases I still do with AI agents. I'm going to be doing a full Grockbot boot camp this week in the Vibe Coding Academy. Make sure to sign up for that. Link for that down below. It's the number one AI community on planet Earth. Best decision you ever make joining that. If you learned anything, make sure to subscribe and turn on notifications. Leave a like.

13:14 So grateful you'd watch this video. Seriously, thank you so much.

Summary

SpaceX AI's Gro 4.6 has been released, positioning itself as a competitive alternative to existing AI models like ChatGPT-56 and Fable 5, boasting similar performance at a lower cost and faster execution. The review highlights Gro 4.6's capabilities in coding and simulation tasks, suggesting it is particularly suited for vibe coding, though it lacks a comprehensive general-purpose harness compared to competitors.

- Gro 4.6 claims to match the performance of ChatGPT-56 and Fable 5 while being cheaper and faster.
- Benchmarks show Gro 4.6 outperformed ChatGPT-56 in various coding tasks, completing them in less time and at a lower cost.
- The model excels in coding environments, particularly when used with Cursor, which offers a robust coding experience.
- Gro 4.6 is less effective for general-purpose AI tasks compared to ChatGPT and Claude, which have more comprehensive harnesses.
- Users focused on vibe coding or those with budget constraints may find Gro 4.6 to be the best option.
- Grockbot is recommended for knowledge work, but it lacks the versatility of other AI agents.
- Future updates may enhance Gro 4.6's capabilities and user experience, especially with potential rebranding of Cursor.
- The review emphasizes the need for users to choose the right AI model based on their specific requirements, with Gro 4.6 being ideal for coding tasks.

Questions Answered

What are the claims about Gro 4.6 and how does it compare to other models?

Gro 4.6 is positioned as a competitive AI model that offers quality comparable to more expensive models like Opus 5 and GBT56 at a lower cost and faster speed. The speaker expresses skepticism about benchmarks but indicates that Gro 4.6 can be used effectively in Cursor and Grock Build.

How does Gro 4.6 perform in comparison to Chat GBT and Fable?

In various tests, Gro 4.6 performed well, recreating the Apple website and completing a scavenger hunt faster and cheaper than Chat GBT. While it took slightly longer in some tests, the cost-effectiveness was highlighted.

How does Gro 4.6 stack up against Fable and Opus in various tests?

Gro 4.6 outperformed Fable in bug finding and was cheaper, although Fable completed tasks faster. In gaming simulations, Gro 4.6 showed strong graphics but had some physics issues compared to Opus.

What are the strengths and weaknesses of Gro 4.6 for different types of users?

Gro 4.6 is excellent for coding tasks and offers a fast, cost-effective model, but lacks the intuitive features of general-purpose AI applications like Chat GBT. Users focused on coding may prefer Gro 4.6, while those needing broader AI capabilities might find it lacking.

What are the future prospects for Gro 4.6 and its applications?

The speaker believes Gro 4.6 is rapidly catching up to competitors like Claude and Chat GBT. While Grockbot is highly recommended, the speaker still prefers Chat GBT for general knowledge tasks. There are concerns about Claude's stagnation and potential improvements in the future.

© transcribe · For agents Built with care and craft by Gokul Rajaram