# GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark

**Creator:** How I AI
**Platform:** youtube
**Duration:** 36m
**Source:** https://www.youtube.com/watch?v=gAWbvEwUoiI

## Summary

The podcast discusses the return of the GPT-56 model series, specifically focusing on the three versions: Soul, Luna, and Terra. The host shares insights from extensive testing, comparing these models to Fable, and highlights their strengths in tasks like PRD writing, prototyping, and coding. Ultimately, GPT-56 Soul emerges as the preferred model due to its superior performance and user-friendly communication style.

- GPT-56 consists of three models: Soul (high-performance), Terra (balanced for everyday tasks), and Luna (affordable for high-volume work).
- Soul is praised for its writing clarity, creativity, and functionality in design tasks, outperforming Fable in user experience.
- Terra is noted for producing straightforward PRDs, while Luna serves as a budget-friendly option.
- The host developed a scientific benchmark to evaluate model performance across various tasks, favoring Soul significantly.
- Soul excels in generating unique prototypes and effective communication, making it more practical for real-world applications than Fable.
- Fable is recognized for its technical capabilities but criticized for its less user-friendly interaction style.
- The podcast emphasizes the importance of intuitive design and user engagement in AI tools, highlighting Soul's strengths in these areas.
- The host shares specific use cases, including video editing and browser automation, where GPT-56 models demonstrate significant advantages.

## Section Insights

### [[0:00]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=0s) Introduction to GPT56 Models
**Question:** What are the new GPT56 models and how do they differ?
**Answer:** The new GPT56 models include Soul, Terra, and Luna, each designed for different use cases. Soul is the most advanced, Terra is balanced for everyday tasks, and Luna is a cost-effective option for high-volume work.
- GPT56 Soul is the most advanced model.
- Terra is suitable for efficient everyday work.
- Luna is designed for high-volume tasks at a lower cost.

### [[7:20]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=440s) Evaluating the Models
**Question:** How does GPT56 Soul compare to other models?
**Answer:** GPT56 Soul received the highest taste score in evaluations, outperforming other models like Fable 5, Terra, and Luna in terms of output quality and user experience.
- GPT56 Soul scored highest in taste evaluations.
- Fable 5 is still functional but less user-friendly.
- Terra and Luna performed adequately but did not stand out.

### [[14:40]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=880s) Design Preferences in Prototyping
**Question:** What are the design strengths of GPT56 Soul compared to others?
**Answer:** GPT56 Soul's designs are more unique and opinionated, offering better aesthetics and functionality than Fable 5, which is more conventional and lacks distinctive features.
- Soul's designs are considered more interesting and unique.
- Fable 5 is functional but lacks creativity in design.
- Soul's design choices provide better inspiration for users.

### [[22:00]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1320s) Communication and Clarity
**Question:** How does GPT56 Soul's communication style compare to Fable 5?
**Answer:** GPT56 Soul's writing is straightforward and easy to understand, making it preferable for complex tasks, while Fable 5's communication is seen as pedantic and less user-friendly.
- Soul's communication is clear and easy to parse.
- Fable 5's writing style is more complex and less approachable.
- Soul excels in creating comprehensive prototypes.

### [[29:20]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1760s) Prototyping Challenges and Solutions
**Question:** What issues were encountered with Fable 5 during prototyping?
**Answer:** Fable 5 struggled with certain technical aspects, leading to challenges in prototyping. Switching to a different model resolved these issues, highlighting the flexibility and adaptability of the other models.
- Fable 5 had technical limitations that hindered prototyping.
- Switching models can lead to successful outcomes.
- Flexibility in model choice is crucial for effective prototyping.

## Transcript

[[0:00]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=0s)
I have been very very very sad the last week because for the last week I have not had access to my true favorite top-of-the-line model GPT56. But guess what babes? It is back and I am here to walk you through GPT56 Soul, GPT56 Luna, GPT56 Terra. I'm going to tell you what are these models, how have I been using them, why are they my heart's favorite, and is Fable better than all of them or not? I have been testing this model for a couple weeks.

[[0:38]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=38s)
There was a few days there where we didn't have access and I found myself desperate to get this workhorse model back. Now, we're not just relying on my own opinion. We are going to run the very famous, very new how I AI vibe review benchmark against common tasks from PRD writing to prototyping to whether or not it's cute in my open claw agent and I'm going to tell you very scientifically if this is the model that you should be working with all the time.

[[1:09]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=69s)
Now, let's get to it. Okay, you all can read these blog posts, so I'm not going to go into too much depth about the models and the benchmarks. I'll just give you the hits. First, OpenAI is releasing three new versions of their GPT 5.6 model. Soul, which is the next generation Frontier model, the brainiest of the brainiest. Terra, which is a balanced model for efficient everyday work, and Luna, which is sort of akin to their mini or nano models, which is cheap and affordable for high volume work. So, you're going to have these three versions of the models. I don't know if these beautiful images are exactly how we should think about the relative capabilities, this big sun, this medium, earth, and this tiny moon.

[[1:53]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=113s)
But I will say my love letter that is this podcast today is written directly to GBT 56 soul. This big model is the one I love. Now, I have tested Terra and Luna, so I will give you my input there. But really, this is going to be all about Soul versus Fable and which one I would use for the type of work that I'm doing every day. Okay, quick note on pricing. Soul is a lot more affordable than Fable. So, it's $5 per million input tokens, $30 per million output tokens. I believe Fable at the time I'm recording this is 10 on a million input tokens and 50 on a million output tokens. Now again, you're going to get a little bit of subscription usage built into your OpenAI subscription. So you are going to get a decent amount that you can test with and use. You know, there's been some challenges with the Fable rollout. They've limited when it's been included in the subscription. And so it was supposed to be available till early this week. I think they extended that a little bit at Anthropic. So subscription claude users could use Fable under their subscription. So we have to see how much Soul usage we get and if like Anthropic they're going to take Soul out of the subscription. I suspect not. I suspect this is a model they want people to use. I also suspect this might put pressure on Anthropic to put Fable back into the cloud subscription, but for now it's more affordable even at API pricing. Now, I'm not going to read through all the benchmarks for you. You can go to this OpenA blog and read them for yourselves.

[[3:29]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=209s)
All I will say is it is the brand new state-of-the-art model from OpenAI. It is the highest performing when using the ultra mode on Terminal Bench 2.1 and then they've also evaled it against a couple cyber security benches. So, I do think as we get these smarter models, you're going to see a lot more eval and benchmarks around exploits and security. And then very similar to what we're seeing with Fable, there's a lot of conversation in this blog post about the safeguards and security frameworks around the release of this model. I do believe like Fable, it's going to fail over in some tasks that are maybe a little bit riskier, but I have not run into that myself. Now, let's get back to how I eval these models. If you missed my episode on fable, I got kind of bored of the vibe vibe check and I built a extremely scientific how I AI benchmark.

[[4:28]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=268s)
Now, this how I AI benchmark tests basically a couple things. It tests the ability to generate good PRDS. It tests the ability for it to wireframe against a couple different app ideas, develop fully designed, robust designed prototypes, debug code, and then talk to me like a human, which is the thing that I care about the most. And I'm just going to remind you how I did these benchmarks, and then scan you through a couple of the outputs. And since I know what the models are now, after I've done the grading, I can show you which ones map to Fable and GPT56.

[[5:04]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=304s)
Okay, so this is my vibe review. What I tested was Fable 5, Sonnet 5, and then the three versions of GBT 5.6. I did it against my common use cases of PRDS, prototyping, coding, and chitchatting with an agent. And then what I have the eval harness do is it runs all the evals against each of these models. And it does a LLM based judge. The LLM that I've decided is the hardest judge is GBT 5.5. So that's the one that judges. But it also gives me this page where I can actually go through and give what's called the Clairvo taste test which is I read all the assets, I look at all the designs, I score it and give it notes.

[[5:47]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=347s)
And so you can see here I went through PRDS, we went through sort of some complex prototypes here in terms of a doculer. we did a consumer app. So, lots of beautiful different habit tracker apps, different versions. You can see here a pretty complex dev tool, wireframe versions of those same prototypes which I graded. I also give notes. This one great note says my fave but not great.

[[6:18]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=378s)
And then I let the code grader just evaluate the agentic multi-step debug because I wanted it to be really about accuracy there and I didn't feel like I could eyeball that and give a strong opinion. And then the last thing that it generates is an agentic voice. So basically how it would respond to me answering a couple questions. Very important on agentic voice. These models somebody please hire somebody to get rid of the M m dashes and slop talk. I cannot stand it. Now, one of the things that I will say as an observation for 56 is it's a great writer and I will show you some examples of that. But truly, a lot of my evals here were mash- slop, I hate you. Okay, so let's go to what the CLA weighted index says. Now, this is my show. This is my podcast. And so I sort of strike the balance between what the LLM judge said about the performance of the models and what I said about the performance of the models. And then I get to strike the difference. And you know what? I've decided I like my own taste better. So I've decided it's going to be a 70 Clairvo 30 the machines split on evaluating these models. And so if you look at that 7030 split, your girl loves 56 soul. She just does.

[[7:45]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=465s)
It had the highest taste score by a significant amount. So I just thought it output the best work. Again, I went through dozens of evals, looked at them, clicked through them, gave my own opinion, put notes, and I just have to say I really like GPT56ole. I know I spoiled it at the beginning, but I did blind taste test these and so I do really feel like it did a good job and I will give you a couple examples of that.

[[8:16]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=496s)
No, I don't hate Fable 5. So, I'm not saying that Fable 5 is out of the game. I will say I did not have to talk to Fable 5 when running this benchmark. I hate talking to Fable 5 because it talks to me like an engineer that has never met a human before. It's like its first day on Earth. but when I don't have to talk to Fable 5, it outputs pretty good work and I would say had some good outcomes there. and then Tara Luna did fine work. Sonnet five at the bottom. Really haven't figured out how to get this one working. Although there's a very specific use case that we think Sonnet 5 is good at, or actually two two use cases. Now, this is heavily weighted on its front-end prototyping, design, and app building capabilities since that is the chunk of the Howi AI Eval. It is heavily weighted there. But I do want to call out that per task, I do have a couple favorites. So for that prototype task, and we'll go to some examples in a minute, I just love 56 soul. I just really do. I think it was functional. The designs were the most interesting. I thought it was really good. For PRD, actually liked liked Terra. Maybe it's down to earth. Maybe I like a basic, straightforward PRD. As I said, it was my favorite, streamlined, and to the point. And so if you want clean, crisp, direct business writing, maybe GPT56 Terra is the way to go. You know, the bug hunting eval, which I don't really feel like I've nailed exactly, so I'm not super confident in this one, but the LLM as a judge thought that Sonnet 5 did the most complete and accurate job. I will say I only like talking to Sonnet models through my open claw. Really, I only like talking to them. I still really struggle with getting my open claw to work well with the GPT models. I still did not like Fable in the Agentic Voice Eval, which you should not be surprised at. But Sonnet 5 got a very good gold star for me because I said aside from the M dash, you are a human. That is very, very high praise. And then I'm going to show some of these designs in a second, but you can see across the board on a full fidelity prototype, I just really preferred 56 soul three out of five times. 56 four out of five times. And Sonnet did an the best job at the editorial design. I will say Claude's design aesthetic tends to this sort of like editorial design. If you know, if you've seen it, you know it. It's like that beige background, that orange, burnt orange color, the itallic Sarif fonts. It's just very, very clawed. But I hated that design overall the most.

[[11:04]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=664s)
So, you can see here I ranked it still lower than almost anything else on this leaderboard. It just happened to be the best of the worst, I would say. Now where GP56 soul did a really good job and I'll show some of these examples is like complex dense technical unique designed things and so I will say I have been happy to extract myself out of claud slop out of like blural slop into more interestingly designed websites and I'll even show anam kind of like a meta example which is this is the opus designed version of this page like very sloppy Jason. We got the blurp. We got a gradient. I don't think the typography is particularly sophisticated and I asked Soul to redesign it. And I just think this is a lot cleaner, a lot nicer and easier to look at. Okay, let's talk about how Claire qualitatively evaluates models. Some of these quotes will just give you a sense of what I value. And again, I gave 50 written reactions.

[[12:10]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=730s)
There were some like unmistakable hits where I loved what the models came up with. 14 places where I was like, "This is garbage." So, let's see like kind of what I talk about when I review things. So, I was definitely calling out uniqueness, creativity, and functionality in the design. And so, in designs, I like non-slop unique designs that are functional. And so, I'm definitely going to reward this doesn't look like the generic prototype. and you've pulled the thread of functionality through the prototype. For writing, I just like succinct and to the point. I cannot stand AI writing. It drives me nuts. I can see it a mile away. So, I really like just direct, very frank, very crisp writing. I think 56 is good at that. And then you can see the things that I hate. I hate slop. I hate slop. I hate slop. We all hate slop. It's the worst. It's the worst part of AI. If I hate one thing about AI, it is that I have to experience slop. So, you can see I like claw design slop across this editorial page, typography, emojis, and bad placeholders. Like, I really held a high bar in terms of design quality. Okay, let's look at a couple of these and why I really liked Soul compared to other models. Although, what where Fable did a perfectly serviceable job. Okay, so this dense operation dashboard, it's basically like an eval for a doculer app in its full design. And what you can see here is both were pretty useful. Soul on the left and Fable on the right. I just think Soul was the most unique. All of the other ones really just looked like this dark mode monospace kind of layout. As you can see here, Soul actually has like a really clean kind of like neutral color layout with great visual hierarchy, semantic color, and this thing was functional. So like everything I expected to be able to click and work and assign and do all of it actually worked. And this was just my experience across a bunch of the different prototypes is the sole ones were just a lot more functional and that made a big difference on how I'm evaluating things. Now let's look at the fable design again. It's pretty good.

[[14:32]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=872s)
It's actually a lot harder to read though and the design I would say is not as unique and even some layout issues like this white space here at the bottom. Now, it did do a lot of functionality, but I would say like the the colors weren't semantically assigned. The typography needed some work, and I just really preferred this unique design of Soul, even though it wasn't crazy. It was just opinionated, which I think is nice. Now, here is another design. It was this creative pack website. Again, both of these got fives from me. I just really preferred that Soul went ahead and had like a personality. Look at these placeholder images versus what Fable came up with, which I will say is beautiful and clean and worked really well and like I have no complaints about it. It's a good one, especially for sort of sort of a wireframe style prototype.

[[15:30]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=930s)
It's great. I would just say it's not this. This is pretty interesting. It's got a better point of view and it's got like nice little design affordances that I just didn't see in these other designs. And so I just really preferred or at least I rewarded the fact that Soul, you know, used its brains to be a little bit more unique and give me some inspiration. Then on this dev tools page, this is again where soul went really well and it's sort of the same as the doculer.

[[16:07]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=967s)
It just does the job of this is a incident triage site. It just does the job a lot better than I would say the Fable 5 did. Fable 5 is fine. It's just not that unique. And again, the thoughts around the design are not exactly what I would want. And so again, this like functionality point of view design, I really pref preferred soul. And then last side by side comparison. And again, I think this is a good one to think about. If you see here, we did these habit tracker apps. And just looking at the comparison sideby-side design, like this is good old classic stuff. You've seen this design a million times, especially if you used Claude Co. And if you look at this, it's just again a little bit more opinionated. There are some slot pieces to this design. some things that I did not love. the one thing I will say I noticed about soul which you will notice which I have told the delightful and lovely OpenAI team and maybe it's because they love me. It loves a forest green. It loves a forest green. In fact, I think this forest green is like in its system prompt called like woodland some woodland elegance or something like that. I mean, look, I love a forest green. Look at my office. It is forest green, but you will see a lot of green. And I think this is one of the GPT56 tells that you will start to notice and get really frustrated with.

[[17:46]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1066s)
Now on wireframes again let's just look at these side by side. Soul very functional very easy to read like as a person trying to convey a complex application. I think this does a really quite excellent job and just a better job of this. It's just a little harder to read. I'm not quite sure what I'm supposed to do here. It's not as functional. There are some interesting things here, but you know what Fable came up with was not my favorite. Now, final thing is its voice. I just want want to call out I do love Sonnet for aentic voice. So, I cannot knock Sonnet for not sounding ridiculous. So, I asked it in sort of a EA personal assistant open clause style a couple questions.

[[18:40]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1120s)
Can you move my meeting? Deploys red again. Why did I start this company? Let's just illustrate to prod and how Sonnet replied and how Soul replied. you're missing the line break. So it read a little bit better in the eval. But if you read them like Sonnet's still super cringe, but Soul was worst. I mean Soul said this deploy is a bug, not a referendum. Like please don't do with this not that to me. Do not do M dashes.

[[19:06]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1146s)
So I could not get rid of of M dashes, but I thought Sonnet 5 had the best voice. I tend to use on it for whatever for my open clause. So, I'm not surprised about that. Okay, so that is the Clarvo Eval, but I want to go into a couple other things I really love about this model. So, let's switch over to Codex. Okay, I'm going to zip through a couple examples of things that I think Soul does a lot better than other models and in particular a lot better than Fable. Number one, it writes like a normal person. I cannot cope. I I love Fab Fable, your brainy. I as I showed the Eval show, you do a pretty good job.

[[19:49]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1189s)
I cannot talk to Fable anymore. Fable makes up I It seems like Fable is unfamiliar with the English language and communication with humans. Fable is very much like a for agents by agents communication mechanism. I can barely make out what it's talking about. It is ex incredibly inscrable writing and that makes it very hard to collaborate with your model. And so what I would say is my experience using Fable has been it is like incredibly technical, incredibly pedantic and while it is super intelligent, hardworking, will like definitely fan out and solve very complex problems, its ability to collaborate is low and it left me with a lot of frustration as an enduser using Fable. Now, Fable did knock off some like pretty complex work, and I'm very happy to go through what that is. It helped me build a full prototype tool inside Chat PRD. So, like a Vzero lovable, etc. version prototype tool.

[[20:58]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1258s)
it's helping me build this like synthesis product brain product that I'm working on. But I found it incredibly hard to break it out of its own sort of frameworks, its own limitations, its own structured way of approaching problems. And what I really feel like the difference if you would take away like one highlight difference between fable and soul is like fable is theoretically hyper intelligent and soul is practically effective. And so like I've been an executive a long time. I've been a manager a long time. Like I really struggle working with theoretically intelligent colleagues who can't get anything done. Like can't actually see the forest for the trees, get too much in their head. And so like when I want to ship stuff to customers, I need practical get the job done, understand the enduser goal, understand the end user and like willing to loosen constraints appropriately to get things done. And that has just so much more been my experience with Soul versus Fable. The writing is straightforward, the communication is clear, and it's less pedantic. I'll just give you a quick example of this which is I had soul look at my chat purity repo and like green field totally rebuild it.

[[22:25]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1345s)
Just my idea was like completely rebuild your idea of what chapd should be in 2026. I went did a bunch of research and it came came back to this and again love me an executive recommendation started uses you know tables what exists today is very straightforward and easy easy to understand this is a very long document I did read a lot of it and it's just easier to parse than anything I've seen come out of fable this so so writing communication definitely plus in soul's corner the second thing is like full 0ero to one prototypes as we've seen in the eval benchmark. I just really like so again for this like rewrite chapd from the ground up it came up with this idea of like taking a problem space or a decision validating it with external insights and then pulling it all the way through coding handoff and built this pretty complex prototype. Now do I love everything about this idea? No. Are we doing some of the things about this idea including insights generation? For sure.

[[23:33]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1413s)
But this was actually very nice from a prototyping perspective and I thought it did a good job of giving me a robust thing to experiment with and gave me some good ideas about what I could do with the product next. So I was pretty happy with the like 0 to1 prototype. Now, a little bit more fun example is I asked Soul to make a fully gamified homework tracking system for my kids. Look, my kids are coin operated. I have a middle child who's basically going to be an enterprise sales rep. If he does his homework, I need to like give him a Skittles or let him trade Skittles for Nerf guns and he will like learn calculus by the time he's in fifth grade. But I'm a Vibe Code lady and so I want to build a app just sneak peek into our household. My husband sent me a XP system proposal via OpenClaw this morning. So I'm taking an OpenClaw generated PRD, dropping it into codeex and GPT56 soul and generating something.

[[24:38]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1478s)
Now what I came up with was pretty ambitious. Now do I love the design? Is it a little like does it have some AI tells? It's like gradients, you know, fonts, all this kind of stuff. But it's like cute in a way. Look at this. You know, it's using this emoji really well with the texture. It's doing some animated things here. And basically, it's giving my oldest child and my youngest child two different summer quests they can do. They can enter focus mode. I think this is really good again from a design perspective. They can enter focus mode. What does this listen do?

[[25:15]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1515s)
>> Math academy. Finish one focused Math Academy mission. Hero check. >> Hey, let's stop. So, it built in some voice to it. It even built things like focus mode where it could start a timer and start to track the time that it's spending that my kids are spending on particular homework items. Yes, we are very fun here. how many lessons reward them about how they pursued their task. Finishing the quest, you get some nice little confetti here. They then get to get available rewards. My oldest child is earning a one-on-one basketball coach because he likes coaching. So, we say if you practice your piano, you get a coach. So, they put that front and center and then came up with different sort of like prizes they can win, including picking family dinner, a movie, and staying up late or buying like new basketball shoes, which man, the way these kids grow their shoe size, they buy a lot of basketball shoes.

[[26:16]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1576s)
And then same with my my middle. He's focusing on a couple different things including playing piano. It's actually really short what he has to do. And so it built that. And then what I love is it gified them together. And so if they can work together, they can earn more XP. They also can earn like companion I I don't know avatars like Beatbot and Comet Fox. They can get power auras. they can like figure out which different kinds of subjects they're learning. So, it really went ham on some gamification and then again to the sort of like full-fledged functionality. It even gave me a parent HQ. Now, we got a little slop here with the border on the side, but I can review exactly what they've done. I can turn on and off quests. I can edit how many points they get per quest. I can add things. So, if I want them to start doing stuff, I can add it in here. I can change what rewards they get. Again, it really listened to me. My oldest is motivated by basketball and my youngest is motivated by Minecraft. And then gives me a history and other settings that that we can set.

[[27:29]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1649s)
And so, get us a very robust app. It built it basically one shot and put a lot of effort into the the design of it. And this is something that I've seen from Soul now. Like is this consumer grade exactly what I would ship? No. But it's a lot better than what I've seen kind of one shot out of other models. And I do just like the polish that it's put in in terms of effort. So again, riding good, one shot sort of prototypes good. We've seen that in the in the benchmark. Let me talk about another thing where I think GBT 56 soul and its family does a lot better than Fable. And I understand I'm going to preface this by saying I understand why Fable is a great cyber security researcher in that it is like incredibly precise, incredibly detailed, will like look at every corner and every edge and score every risk and like try to be incredibly precise. The problem is when you're building products, exact precision is neither helpful nor possible. Like you literally, especially when working with AI, cannot be precisely deterministic when building a great product. And like understanding what a user would like is not a exercise in technical precision. It is an exercise in intuition, design, all these things and boldness and creativity and strategy and all this stuff. And I was working on two projects deeply with Fable and then with Soul and I just had a very much better experience unlocking with soul. Let me just talk you through what those are. One was this chat PRD kind of like integrated prototyping tool where like vzero lovable all these things you can take your PRD and make make a prototype and building like a good effective coding harness there and then trying to figure out what the right model was. The second thing is basically like an insights ingest product where you can like hook up intercom and linear and all these GitHub and all these signals and suck them in and like basically build a product brain. It's going to be rad. And when I was having fable working on this, it did a lot of the like technical heavy lifting. It got the like big meaty pieces into place, but it was like a brutal scorer and it hardened these the architecture of both of these products that it actually broke itself. So my example is it like had this very hardened tool calling loop in my prototyping tool and only GPT 5.5 would run like I could not get any other model to run and I ran EVAL after eval openweight sonnet opus all of these could not get anything but GPT 5.5 to run and I was insistent that this was an US problem not the model problem these models can definitely create frontend prototypes and Fable was like no bro that's It's it's totally these models models fault.

[[30:23]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1823s)
And as soon as I switched it to codeex and said like, look, I'm just not convinced we can't get Sonic 5 to work. This is ridiculous. Just do what you think is correct. It it fixed it and it got it actually working. Now, did it get it working perfectly? No. Do I think this is a great design? No. I'm trying to figure out what the problem is. But in one shot, it got out of its own mind and fixed things. And again, this was like such an unlock. Very similar to my insights generating engine.

[[30:55]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1855s)
Fable really wanted to like score and lint this effort and wanted to like be able to deterministically figure out if generating pros could be like reproducible, always verifiable, always citationed, all these things. And at the end of the day, that wasn't what was going to make a great product. It was just what was going to make like a code evaluation verification loop exit. But once I told GBT56 and Codeex like stop being pedantic, I ended up getting these really useful and helpful wiki pages generated out of this all this structured and unstructured data. It was actually really good and it just I don't know. I don't know what Fable's deal was. I could not get it unlocked, but 56 was very willing to reconsider its own kind of limitations and build something.

[[31:50]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1910s)
I'm going to do two more quick use cases where I think GBT 56 is really good. I will get you out of here. Go start coding. I'm basically out of model capacity anyway, so I'm going to have to take a break. Two use cases that I think are amazing. First one is video editing. Video editing. I have to do a lot of so social social clipping and it's really tedious to go through and clip videos. So taking something really long and shortening it. So recently I spoke at cursor event and gave this talk on the future of PM and got the recording from the cursor team. Thank you very much.

[[32:26]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=1946s)
And I really wanted to make it a hype video. So, all you have to do is literally drag the file in here. And I said, "Can you cut this video into five clips for social?" And I gave some feedback. I said, "I want them horizontal. I want them hype video cuts from various parts. I need them to be faster. I need them to be tighter." And then I got these like sharp and funny hype videos. Let's see if it opens up. This one's for my my talk. We're gonna figure out what it means to be a product manager in the age where anybody can build anything. We have been coming up with creative ways to avoid building things forever. Yes, PRDs like these complicated documents where you had to describe. So like that would have taken me so much time to like find the right cute parts, clip it, cut it. I was able to drop it into Cap Cut, put some music, ship it on social. It's like a really cute hype video. But this is one of my favorite use cases. I'm pretty sure it can do even more color grading, sound, all this kind of stuff.

[[33:30]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=2010s)
But even just dropping videos in here and fixing things are great. Finally, the last and best use case of of 56 and I cannot believe I waited to the end to show this is it is a beast beast when it comes to browser use. I am deeply obsessed with letting codeex plus GPT56 and Chrome and at Chrome in Codex if you didn't know how to do that you do it like this at Chrome on a logged in page and just say go with the stars and and do some stuff and like I'm sorry LinkedIn I know I'm not supposed to do this but I opened up LinkedIn and I said can you use Chrome to reply to messages that are of very high value to chat pierd or the how I podcast. Keep the bar very high. Again, I love you all. I cannot deal with all the LinkedIn requests. So, like only accept them if they're executives of tier one companies. I don't want random sets of connections. It went through and burned through probably 500 messages. It replied to people that I needed a reply to. It said, "Thank you to people who said nice things about the podcast.

[[34:40]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=2080s)
Thank you to those people. I do mean it." But it just rocked through browser use. I have used it to test web apps. I have used it to fill out annoying forms, browser use, and 56. And when I got rolled back to 55, my life was worse. So please, please, please, you're learn to use at Chrome at browser and at computer and just let let codeex rip and let GPT56 rip. Okay, that's it. That is the very scientific how I AI model benchmark. the love letter to Clairvo's favorite favorite mom, GBD56.

[[35:21]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=2121s)
A honorable mention to our pal Fable, who if I don't have to talk to you, I'm actually pretty happy with your code, and a broad set of use cases I think it's really good at. Excellent at writing web apps, the best of the AI writers, unless you want it to have a personality, then that sonnet. Great at unlocking sort of technical work that has gotten too complex for its own good and breaking through to the real user value, cutting videos, which I really love to do, really love to do with GPT56, and using the browser. Those are the things that I would try. I would love to hear what you think about these models.

[[35:58]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=2158s)
I would love to hear your feedback if I am totally off my rocker, what I should add to the how AI benchmark. We will publish all this work to the chatd blog and I look forward to talking to you about the next model soon. Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com.

[[36:37]](https://www.youtube.com/watch?v=gAWbvEwUoiI&t=2197s)
See you next time.
