Section Insights
Introduction to Slop Code Bench
What is the significance of the Slop Code Bench benchmark?
The Slop Code Bench is a new benchmark that evaluates code quality and performance of AI models in coding tasks. It highlights the pass rates and costs associated with different models, revealing that even the best models currently only achieve a 33% pass rate.
- Slop Code Bench is an unsaturated benchmark for AI coding models.
- Current top models are only achieving a 33% pass rate.
- The benchmark provides insights into code quality, costs, and defects.
Understanding Pass Rates
What is the difference between strict and loose pass rates?
Strict pass rates measure whether all tests in a suite are passed, while loose pass rates only check if the current challenge is passed. This distinction is crucial for evaluating the reliability of AI models in coding tasks.
- Strict pass rates require passing all historical checkpoints.
- Loose pass rates only require passing the current challenge.
- Understanding these metrics is essential for assessing model performance.
Prompting Strategies in AI Coding
How do different prompting strategies affect AI performance in coding tasks?
The study found that using a prompt to plan before implementing does not significantly improve results compared to directly solving the challenge. This suggests that the effectiveness of AI in coding may not heavily rely on the prompting strategy used.
- Prompting to plan before implementation shows minimal performance difference.
- Directly solving challenges may be equally effective.
- The choice of prompting strategy may not be critical for AI coding success.
AI Decision-Making in Coding
Can AI make effective decisions in coding tasks?
Currently, there is skepticism about AI's ability to make good decisions in coding. While there is potential for improvement, many users have yet to see satisfactory decision-making outcomes from AI models.
- There is doubt about AI's current decision-making capabilities in coding.
- Users have not observed significant improvements in AI decision quality.
- Future advancements may enhance AI's ability to make architectural decisions.
Optimizing Code Quality and Confidence
What strategies can be employed to ensure code quality and confidence in AI-generated code?
To maximize code quality and minimize risks, it is suggested to focus on quick fixes and patchability rather than striving for perfection from the outset. This approach allows for iterative improvements and reduces the likelihood of catastrophic failures.
- Prioritize quick fixes and patchability over perfect initial designs.
- Iterative development can lead to better long-term outcomes.
- Understanding the balance between code readability and maintainability is crucial.
Transcript
0:00 Everyone is trying to figure out what code do you still have to read, what code can you not read, how do you maximize and like how do you avoid disaster. >> What it really boils down to is can you understand the benchmark? So there's this new benchmark going around called slot codebench. >> What percentage of the test did it fail and then how much did it cost? >> This really deviates away from my personal use case where I found that soul is quite good at a lot of tasks.
0:25 Why do you think that it's failing here? All right, y'all. On today's AI that works, we went through a new paper newish. It's originally published in March called Slop Code Bench. And I published some results. We did an initial article last week on testing Opus 5 on Slop Code Bench. We went through the results of that. And then we went through the brand new results I just published this morning looking at how GPT 5.6, six soul fable 5 and Kimmy across two different providers performed on slop codebench compared to the historical results. This is a really interesting benchmark because it's not saturated. the best models at our time are only getting 33% which is like these are Swebench 2024 numbers. so very exciting unsaturated benchmark.
1:10 There's a lot of really cool interesting data about code quality, about cost, about defects. and so this was a really fun episode. I hope you have fun. let's get into it. Dude, >> today's episode I think is going to be incredibly fun. so I'm going to give a little brief intro and then we'll get right into it. For those of you that didn't know, I'm Vibov. This is my co-host Dexter. And this is the AI that works podcast. We where we try and talk about real working AI systems as much as possible. Dexter here runs a company called Human Lair where they help people solve hard problems in difficult code bases using coding agents. I work on building DAML which is a new programming language that's agent first. Today's episode is about benchmarks.
1:52 >> Again, we're going to do another benchmark show. >> There's a reason that they come up so often. And probably for me, the reason that benchmarks come out so often is in a world of total anxiety where everything is unknown, the only thing that we can really know is like some number that someone shares. And we hope that we can trust them. And for me, what I run into is I trust them most of the time.
2:15 But I think what it really boils down to is can you understand the benchmark. So there's this new benchmark going around called Slop Code Bench made by Greg I Gabe Gabe Greg. >> Gabe Oransky, the University of Wisconsin Madison, a wonderful school in a wonderful town for what it's worth. >> There you go. Dexter's been talking about it a little bit on Twitter. So you've probably heard about slop code bench at this point. It was on the front page of hack hacker news when I posted the like initial results last week.
2:41 >> All right, Mr. Humble Brag. but I think the thing about today is we're going to dive deep into it. We're going to really ask some questions. I haven't had the opportunity to dig into it nearly as much while I've read some of the posts. Dexter has been deep in the weeds. So today's conversation is going to be a fun lecture by Dexter Dexter sharing with us what SLOC code is bench is about. You'll hear me interrupt, ask questions along the way. I welcome you all to come ask questions as well and let's just poke at Dexter's brain and see what we learn about Slab Code Bench.
3:10 >> Why don't you give people two fun announcements what's going on this weekend and I will be I'll be ready to share in about a minute. >> For those of you that aren't in San Francisco, we have a special thing going on. if you haven't seen already, let's go to Gouta. and then let's go to Dexter profile and you will see a fun little tweet where we are running our unconference. We've got some amazing speakers coming out. The unconference is unlike other conferences. We don't have we have some pre-planned speakers from people that we know already give great talks, but most of the content is actually driven by the attendees. You sign up, you go to a whiteboard, you submit what things you want to talk about, and every speaker just chooses who they want to hear from next. and we kind of let the conversation flow as it is. This is a thing that we started doing last year. It was one of my favorite events and clearly some other folks enjoy it as well. And if you guys are interested, you can find this either QR code or you can find Dexter's post or my LinkedIn post and sign up and come attend. We do kind of keep it somewhat curated. And for some of the events, we'll try and record the talks and post them if if the speakers are okay with it. But it is Chattam House rules and that's kind of how we get the best talks coming out because when you can say anything and you know it's not going to be cited against your name, you can be honest and you can share real war stories.
4:30 >> People are willing to share a lot more interesting stuff when they can do it off the record. >> Exactly. And this is kind of the goal. Get the best engineers we know, get the best founders we know, and put them into a space. Sadly, no VCs. I apologize in advance, but if you're an engineer and you're coding, this is hopefully a fun weekend, fun way to spend a Saturday. >> Incredible. Great intro. you can find that if you go search our Twitter feeds or whatever. or you can scan the QR code. The recordings are up. You can go see it. let's talk about slop codebench. Amazing. So, I am going to share.
5:07 We'll just do the whole thing for now. so I just posted the article just a second ago. but let's do a little bit of excal real quick. I want to find what was it? Slop >> quick. If you want to give a talk like Dexter, screenshot the Excalad draw and now you know what diagrams he's going to have in his next talk and steal them. So we talked recently about how coding coding agent benchmarks work. and the old world of like write fsbuzz and see and the agent writes it and then we have a unit test that checks it and like the way we build these benchmarks is basically like we use did a human merge this PR off an open- source repo as the signal of like the labeling of the data.
5:54 and then we curate those cases and then we create a verifier and based on the tests that the human wrote basically. and then we talked about basically like the newer benchmarks that do some interesting things like LLM judge models and things like this. but the my biggest problem is like software is really discovering problems as you go. but most benchmarks just give you the whole problem up front. Slop codebench is a benchmark and I'll pull up the website here.
6:22 Slapcobench is a benchmark that is explicitly designed one to be really hard and to be unsaturated by today's models and to show the mo to give the model increasing complexity of problem over time. So that as it works you end up let me pull this up. As it works you end up basically every time the model does a challenge it inherits the previous codebase. And this one actually doesn't have the thing that I like. so basically the way it works is like you give it an initial spec and then it writes some code and then you say, "Oh, actually we also want to add support for something else." And then it writes more code on top of the existing code and you give it more features and more and some of these some of these challenges have like eight plus checkpoints. And so every single time that the problem starts easy and then gets harder and weirder as you go. And you're not just measuring can the model solve a hard and build a complicated thing, but if you just if it's if it doesn't know what's coming, does it create a mountain of slop that makes it harder and harder for it to actually make progress?
7:29 >> Yeah. And the goal here is really it seems like to mimic traditional software patterns where like >> this this is how real teams build things. You don't know the whole thing up front. >> Yeah. You get some intuition, but you don't know. And I think what's interesting here is like when you're trying to train a coding model, it's really hard to put that intuition about where the product is going in the training data. And that's kind of fascinating because the only thing you can measure is did it work, which is different than will it work long term?
7:55 >> Yeah. Will I be able to add new thing? Yeah. Because it's it's it's both position and velocity, right? Is like >> does it do the thing I want today? And will it be easy to add new things, change things, and update it? Do you have a concrete example of a test case here that we can look through just so we can ground ourselves a little bit more concretely? >> with pleasure. So, where's the article? so I actually in the appendix of this article, we have all of the e all of the challenges.
8:27 okay. >> So, XJQ is an interesting one is like you basically for one you build like jQ for XML. >> that reads like XML or HTML from standard in and does like basically selecting things out and then rewrites it. >> Then it adds support for CSS selectors. >> Then it does JSON objects and arrays. Then you would read from a file instead of standard in and accept like UTF8 and then you add JSON export. So it it goes from being a like pull stuff out of XML to like build a generic serializer and deserializer. I mean, yeah, the only mistake here is they use XML, but you know, other than that, it's pretty good.
9:05 >> It's an interesting problem, right? It's a toy problem, right? >> No, I'm joking. I'm joking. I'm being I'm being stupid. >> I love XML, bro. You put XML in your prompts all the time. I've seen it. >> No, I don't. >> You used to >> I don't believe in XML. yeah. >> We're going to go pull up cracking the prompting interview and I'm going to go find footage of you using XML in a prompt. >> the circuit eval is an interesting one, too. I so one of the things that the first time I did this benchmark I did just three of them and I did not use like the Frontier Frontier models.
9:35 We just did like Opus 5 against Opus 4.8 and Sonnet 5. >> and you see Opus 5 got a 25% pass rate 23% pass rate and the other one's got like a 6% pass rate which is basically just Go ahead. >> Yeah. And just to help me understand when you say pass rate, >> so for example, you just showed me, can you Oh, can you open two tabs side by side so you can show me the pass rate and then the checkpoint things too?
10:00 >> Yeah. >> so Oh man, you got to get yourself some some raycast. >> What? >> >> gross. I use an adult. I use a proper window manager for adults, sir. >> Oh, that's actually pretty interesting. I don't know how to do that. That sounds advanced. So when I look at these checkpoints >> and when I look at pass rates, help ground me. What what does that mean? What is a pass rate here? >> So a pass rate is basically I'm actually going to go find this picture from the like run in progress. So basically what we care about is strict pass rate.
10:35 So the model's going to go through >> the right one >> on the right one. You want to zoom it in? >> Oh. Oh, okay. Never mind. >> what were you asking? I was going to say, can you make the screen 50/50 instead of 3060? But yes. >> no, we're just going to look at this one. So, okay, basically what happens is the model does the first challenge. >> and these are all partial passes. Basically, like it did not it did not pass every single like this basically do like blackbox testing. Every every checkpoint emits either like a CLI to run or an HTTP server that like serves results or whatever it is. and so Opus 4.8 8 did the first challenge and did not pass every test. Opus 5 for the first three challenges in circuit eval was able to pass 100% on all of them and then starting in challenge four it started to accumulate defects. and so the strict pass rate is basically like how for how many of them is every single test in the suite running? And so for the second challenge, you can see like Sonnet and Opus actually nailed the first one, but then they failed the next one. and so this is the idea is like, okay, even if you get, you know, let's say you get like number four, you you solve all the new challenges, but you break something from number three, then you fail the strict pass. We don't know the mechanism and the breakdown of which test failed here. I mean, I have the data, but I haven't put it in the report. but does that kind of answer your question of like what pass rate means?
12:08 >> I see. So strictness. Okay. Okay. So there's like loose pass rate like have I passed this checkpoint and there's strict pass rate. Have I passed this checkpoint and every historical checkpoint >> and all the previous ones? Yeah. The the loose pass rate they call it isolated pass rate. >> That makes sense. >> You can look at this in the >> it's kind of like >> paper. >> It's kind of like when coding agents when coding agents work it's like what am I really solving for? Have I solved the current problem in front I have and have I accidentally not broken anything in the past?
12:35 >> Exactly. >> Cool. That makes sense. And so you >> see this in code bases a lot where it runs like the current packages test but then it runs the entire test suite at the very end which >> yeah you want to make sure you didn't regress anything right >> so in this case like when they ran this in March the best model was GPT 5.3 codeex or actually 5.4 but like codeex actually won the isolated pass rate it was able to like win more of the challenges and solve the whole problem but opus 4.6 won the strict pass rate where it's like it it fixed things without regressing.
13:11 >> Does that make sense? >> Yes. >> Cool. so coming back to the the data on pass rates, right? So in the old world, we had a 23% pass rate. in the in the with the Frontier models, we got a little bit better. and so here's the like strict pass rate for all the checkpoints. so they also released new numbers. So GPT 5.5 got 14.8% strict pass rate. we found that soul fable both tied at 33%. And then I ran Kimmy on modal and base 10.
13:49 just to like kind of I just Yeah, >> I had two providers access and they both sponsored inference for this. So I figured I would run them both. It's actually very interesting. Like the modal one technically got got a better strict pass rate, but the base 10 one like got way more. And like my takeaway here is like they're both great providers. We only did one run. Like this is not statistically significant. you should use both of them. Thanks to both of them for sponsoring the inference. I had actually kept the modal one a secret just in case. I didn't want to say which provider it was in case one of them did way worse. I would rather like give them that feedback in private.
14:20 But they were both like in the same arena. >> I think you just give the data, man. You you got to be honest with what it is. It's It's the best thing you can do forever not using them. >> I would feel bad if someone at either company was like, "Hey, we got to give Dex some free inference so we can run this benchmark." And then I published the results and they got roasted, right? Like I would rather give them the opportunity to like fix the provider and and run it again. But anyways, this is not especially if it's like a lot of these things are like basically a core engineering issue sometimes that are really easy to just patch.
14:48 >> Exactly. So anyways, technically the the modal one did a little bit better. but like the interesting thing here is you can see like they all add defects as they go. So a defect is like a failing blackbox test case or like basically you just count as they go over over time. Soul really sucked at this one in particular. and in general like it was kind of the worst although on this medium one it actually did the best out of all of them. Does this chart make sense?
15:19 >> I mean this chart makes sense but I have a very interesting question. Yeah. >> So here's my question here. >> Yeah. >> I think a lot of software engineering is solved by having really good understanding that the problem that you're solving up front. >> So while it's true that you're incrementally adding more software, what I would be really really curious about is what if I gave the whole spec up front of every single problem? >> Oh, you want to compare it compare the defect rate of the whole thing. That's super interesting. I hadn't even thought of that.
15:49 >> Suspicion. My suspicion would be that every single model would nail it significantly better upfront than it will on incremental progress. >> Damn. I have I have three like things I want to try next, but that was not one of them and I think I'm going to incorporate that. I I'll get into like the ways I think this benchmark could be improved or made more interesting or add more dimensions. But that's a really good one. I like that.
16:12 >> That's like my first insight here because I'm like what do the best engineers do? The best engineers provide all the context up front. And if that happens to be true, then really the takeaway is not that these models are bad at long horizon. It's that the humans that are the best at using the models are the ones that can predict the feature the best, which has mostly been true in most of software. That actually is in conflict with my belief which is I believe the best engineers by some definition of engineer like the best product engineers ship small things iterate fast get feedback and like assume that the customer knows more about the problem than they do >> that probably comes from the different domains that we work in and like mostly I've worked in like hardware like algorithm design and a lot of like resource constrained environments like droid system distributed systems are another example of that kind my workload. Like in distured systems, you don't have to ask the customer what to do. You can kind of like if you're able to predict the workload, you can design the right architecture for this. And your intuition for designing the right architecture is how closely can you predict the kinds of workloads you have to deal with. And if you mispredict that, you know, you're not having nines of uptime.
17:22 >> But my take is like, okay, if the customer would be happy with two nines of uptime, then just get it in their hands as quickly as possible. because there's certain things where it's like, "Oh my god, that takes 10 seconds. That's terrible." But then you find out like nobody gives a if it takes 10 seconds because it's happening in the background overnight. >> But I think that's what I mean. That's like part of the resource constraints.
17:42 >> Yeah. So I mean, yes, you can sit and think think for days at a time and try to come up with the perfect problem and perfect solution and perfect architecture. >> but I think there's there's there's there's other ways to approach building software. >> Yeah. anyways that's an interesting thing that we should dig into >> also for a lot of product oriented software sitting and thinking is usually the wrong approach. >> Exactly. >> Like like just to be very clear like what I'm saying only applies to like infrastructure. It applies to like core like performance engineering.
18:14 >> Yeah. >> It really applies to like product design. >> Yep. So, getting back to the data, like you can see here, it's all over the place. Like, modal Kimmy won this challenge. Soul won this challenge. Oh, sorry. Defects left open is bad. So, Fable Kimmy Kimmy won this one. Fable won this one. Fable won this one. Kimmy won this one. >> GPT won this one. Yeah, this is it's very interesting to see like they're kind of all over the And again, like I didn't run the whole benchmark. I only run ran six challenges. I think they have like something like 30 or 40 challenges now. so if I get more money in time, we'll kind of run this for a broader I every time I run this, I kind of double the size.
19:00 >> Okay, I have a question. This really deviates away from my personal use case where I found that soul is quite good at a lot of tasks. Why do you think that this it's failing here? It feels like soul is quite worse in a lot of ways. It is interesting. It only Yeah, it only it only took the prize. It only got first place on Circuit Eval >> and it did quite worse on a couple of them. So yeah, that's I actually am with you, man. I use Soul for everything. I I am so sick of the Fable like the way it talks and Opus 5 the way it talks that I'm just like >> I will deal with shitty UIs and polish them myself by hand and like go look at the CSS in the browser because Fable's better at that. But I don't I'm sick of talking to Fable.
19:49 so it's interesting. Yeah, I don't know the mechanism behind it. what's really interesting is I I had it build up the basically like what is the zoom in on this part of this. Maybe I can >> can't zoom in, man. >> I love Twitter. here we go. So, the first one, the first chart is like the final defect rate. And I had to I had to take it as like the percentage of defects just because the gray dots from the previous run had fewer exercises. So we couldn't use the raw defect count because the new models the the Frontier models did more problems.
20:22 >> but you do see this kind of like inverse curve relationship. But you do see that like both Kimmy models are underneath the fable curve more or less. I mean depending on where you draw it. >> Show me go down. What does the bottom think? >> So the bottom is cost. This is like >> for every one of these dots is a single challenge. And so you have like what percentage of the test did it fail and then how much did it cost?
20:49 >> Okay. >> And so there's kind of this inverse relationship here where you're going like, you know, obviously spending more can tend to correlate to better better like test results, but even Fable on one of the cheap ones has this. >> It seems very correlated. Yeah, it's there's something going on here. So, I don't know. It might be fun to like draw the lines through here, but it's Yeah, it looks like an asmtote going down to zero, which is kind of what we expect.
21:15 >> Spend more money and solve more problems. >> More money problems. >> I mean, so, so this is another thing is like this is all just using a prompt they called just solve. So, they have a couple different prompt options you can pass in. and some of them are just like here's the challenge, go solve it. They have another prompt that is like plan and then implement which is like go read the code, write a plan and then go solve it. And basically what they found is those perform pretty much exactly the same.
21:44 >> but there's other things here where it could be like cool do a Ralph loop and be like cool solve it then review the code then solve the code review. and you could actually see like oh if I spend more money do I get better results? It's interesting to me you're posing you you said something very fascinating which is you asked them all to write a plan and then solve it and you asked them all to just solve it and like it doesn't matter.
22:06 >> I don't I haven't looked at the I haven't looked at the results. I only skimmed them. So like they put a big asterisk on that but my understanding was that the that the just sol that the the plan and then implement didn't make a like big enough difference to to be like a thing you should recommend. So, I have I'm sure you have lots of opinions on that as Mr. as Mr. RPI. but when I think about it, I have what's fascinating is I too have stopped using plan mode in a lot of situations.
22:38 I mean, I I chat with the model and I interact with it and design with it, but it's different than like traditional plan mode. >> You build a >> what is what did we call this? It's the design. Let me find it. Yeah, this thing talks about a design concept which is basically like the thing that is kind of locked up in the context window. It never gets written down into a file, but it's your shared understanding with the model.
23:02 >> Yeah. And what's really fascinating to me is I think most folks that are doing planning, the the problem is that planning is a is a prompt steered by the harness providers and the prompt is only so good for the average workload, >> right? They have to make it as generic as possible versus you being able to just tell it write a file that has these sections and these headers and works this way. >> And I think that's kind of why like I found myself not wanting to plan all the time except for things that like sometimes require planning.
23:34 >> But like you could also just say like ask me questions until you have until we're aligned and then write a doc that looks like this. or I or I talk about like the surface area and not really the full execution plan because if the surface area is well agreed upon then it's easy to go anyway. I I just found that interesting. what what are your thoughts on that? Like upon hearing that at first where it's like the plan didn't make a huge impact. I mean, so I think I actually had this insight a couple months ago, which was basically like last summer, about a year ago, when we first start talking about RPI and for like the three months before that. so in the last like take like 15 months ago through like 12 to 10 months ago people got really into planning and I realized like I think the actual mechanism behind that was having the model write a plan first even if you didn't reset the context window because what you're talking about is like advanced context engineering of like hey let's get all this into a doc and then I'm going to clear the context and start over do do whatever it is >> maybe you're not clearing the context the point is is like planning was a way to get the model to work for longer if you said go build me a SAS, it would like do a little bit and then it would pause and come back to you. but if you wrote a plan, it would kind of keep working until the whole plan was written. And so that let you be a little more hands-off, a little less in the loop. now these models are kind of very very RL to work for longer. So I think like that benefit of planning is now gone. Like it's not useful anymore.
24:55 >> Yeah, that's actually a really good insight on like why planning was useful in the beginning. It was like a checkpointing system in the beginning, but now it's like not that value is gone. It was a It was a way to like in increase the scope of the problem that the model would do unattended. >> Yeah. >> Now, turn out what we ended up wanting to do was only do one part of the plan at a time and do the Ralph Wiggum style thing of like, cool, launch a sub agent to do part one, then I'm going to check it, then launch another sub agent to do part two, or like basically like doing each phase in its own context window.
25:25 >> >> we've been for a while. Does any while we're chatting, do you folks have questions about SLOB codebench and kind of the things that it has? that it kind of asks you to do. But while we wait for that, Dexter, you could keep going. >> Yeah, we'll keep going. But yes, please drop questions in the chat. this is the overall like summary of all this data of like basically like the total cost and the total defect rate. so very interesting to see like for the co for you know six five five to sixx the cost you only save you know a little bit of defect rate.
26:03 >> and defect again is like not the strict pass thing of like did you solve the challenge perfectly but like how many of the tests across all the checkpoints are passing. And so Fable did a little bit better. 2% better, but it costs five times as much. >> Yeah. And for most use cases, it's probably a bad spend, but for some it might be critical. So you can make the decision on your own, which is kind of fascinating.
26:27 >> I'm also looking at this Opus 5 one and like I just want to asterisk again that these three the gray ones did a different set of challenges. so it's really just kind of like directional context. I I wouldn't count this as a comparison here. So, what I'm hearing is Dexter is about to turn into a Kimmy show. >> we are going to we are going to launch our Kimmy provider on on human layer soon. so if you want to play with Kimmy, you can go try it in human layer. I mean, you can get Kimmy from a lot of places. This one actually uses Kimmy via open code. it was not hard to set up, but but yes, technically technically we will we are excited about Kimmy. so here's another one, right? Like if you just look at strict passes, right? Soul and Fable tied. if you wanted a tiebreaker, Fable got a couple more isolated passes where like it didn't solve all the regressions, but it did like solve all the challenges for an individual checkpoint.
27:24 and it's interesting the like the two mo the two Kimies are like one of them got more strict passes, but the other one got more isolated passes. and here's here's an example. Well, here's actual data from like when we ran it last time on the circuit eval checkpoint on the circuit eval test. >> Like soul actually got the farthest with no issues. Fable got almost as far and then like like none of these actually like lined up there. So there's so many dimensions to like how you can understand like model skill and ability.
27:58 so I just thought it was really interesting to just kind of take a bunch of different projections of this data and look at it along different different lines. really cool. >> Yeah. what else we got? Oh yeah. So then we look into the slop codebench people also do a lot of like code quality metric stuff. and they have a lot of custom like obviously cyclomatic complexity is a term that's been around for 50 years in terms of software quality. but they actually I was talking to Gabe and basically like they have 212 or they had two I think they have more now of like slop detectors like different function different like Python a parsers that try to understand like different things that could signify slop and and and then they like I I took the 10 most interesting ones and tried to look at basically like this this percentage is like from the from checkpoint point one to okay checkpoint 3 to checkpoint 8 basically so like how much does the does the slop metric increase from checkpoint 3 to checkpoint 8 because some of them had zero on checkpoint one so I couldn't do percentage increases but like you can see that Kimmy K3 has more than five times the cyclatic complexity like the the complexity grows really fast across the five checkpoints what's also interesting thing is like a lot of these metrics are very close together like they don't change much over the cost of o over the course of the challenge.
29:27 Does this make sense? >> Yeah. >> so the most interesting ones are like cyclatic complexity and cloned percentage like the percentage of lines that are like exact copies of other lines. And then I like this graph dependency entropy one. I don't actually know how they calculate this. I have to dig into the math but you can go read the paper if you want to understand it. so that's like the overall like results based on the circuit eval one.
29:55 soul wrote more code than all the other models. but and then there's an interesting one here that I'm still have to dig into, but basically like according to the data only soul wrote Python tests and the other models all used like scripts and stuff to test their stuff. They did not write like Python unit tests. I think that just you know why I think that's the case though? >> Yeah. >> You started off with a scripting problem and you slowly evolved it when like literally you just have to like what would really happen in a real production scenario at some point you'd be like write some tests in Python.
30:31 >> Yeah. it would just you would come. This is what's interesting about these models is like >> it's not like you wouldn't steer this whole process, but >> yeah, >> if you knew you were going to go start from a scripting problem and go to a production app, >> at some point you'll make that philosophical swap >> of I'm scripting, I'm productizing. >> And >> well, and it's I think it's also they're they're following patterns, right? So if the first model on the first challenge >> does not choose to write unit tests, then the next challenge in the new context window will see the scripts and be like, "Oh, we test this with scripts.
31:06 Okay, we're going to keep doing that." >> Exactly. Because like that's kind of what you told it to do and that's kind of what a real software engineer would do too. And at some point some software engineer is going to stop and reflect and say I guess that is the point of these benchmarks is like can the model stop and reflect and make an orthogonal decision >> code bases. I'd be so pissed. Could you imagine? Could you imagine you're writing your whole codebase, it's running for a whole day, and somehow in the middle of the thing, it decided to just refactor your whole system and say, "No, this is just wrong." And just threw it all away and made it arbitrary decisions on its own.
31:38 >> I mean, if it made a good decision, I'd be happy. I just have never I rarely see it do that. >> I know that's kind of where I'm stuck, too. I haven't seen any of the good decisions yet. >> Yep. If you can feed it good good examples, then that's great. >> you know, I hear myself saying this and what's funny is I used to say this too. I could I used to say the exact same thing. Could you imagine the model writing the whole feature like no way it's going to mess up? It's going to make some bad decisions and clearly I don't feel that way anymore. So I guess it's just a matter of time before like that also becomes true where it starts making architectural decisions in a good way.
32:12 >> we also looked at kind of other other metrics here at the end, right? So the mean function cyclomatic complexity was pretty even. looks like Kimmy on modal had actually like way lower max. So like the most complicated function that modal wrote was more than half as simple as the most complicated function that the other Kimmy wrote. >> silly question. >> Yeah. >> Is all of this me running on the same harness?
32:45 So soul is running in codec cli >> fable is running in cloud code and the two the two Kimmy instances are running in basically like identical open code configs. >> Got it. >> They we also looked at like what percentage of lines triggered at least one like slop rule. They have like an agp like detector for just like is this line slop? so 95% of the code written by Saul triggered at least one SL and they're very aggressive. Like I don't think this means it's bad code. I just think it means like this is like interestingly like if you wanted to compare across different models.
33:21 >> I really would love to see this code. I'd love to know what the examples of slop cuz like I' I'd probably write some slop myself by hand. >> Unfortunately, it's Python only. so you would have to slop fork it into if you wanted to detect Rust. but yes it can be it it it can be done. >> Never mind. I take it all back. It's it I mean it makes sense. It's 95% slob. It's Python.
33:47 >> Yeah. So another interesting one here is you know just looking at across the checkpoints of this challenge like how does the cyclatic complexity change and how does the number of duplicated lines change. It's weird that Opus 5 actually had less duplication, but again, like I don't think this actually says anything. It's just directional. Like it could be that Opus 5 has less duplication because it wrote more chaotic weird code. Like it didn't follow patterns. It just like wrote whatever the heck it wanted in every single case.
34:17 >> Also, the real interesting point here is like the models are roughly the same. That's actually >> all of these are kind of the same on these sorts of metrics. >> Yeah. So like there's no real alpha in a model or a harness actually. >> Yep. But you could say, "Oh, Kimmy is a little more like shoot from the hip than F." >> It's probably within like standard of error, standard deviation of error. >> It doesn't matter.
34:39 >> Exactly. It's also like if you were going to do this properly, you would run it like five times and average everything, which I did not over hundreds of test suites, not just like these few samples. Like you need to look at like different levels of Anyway. >> Yeah. Yeah. So yeah. you know, we'll see. We'll see if if the if the inference providers want to reup and give me some more credits. >> >> how much does this whole thing cost to run on inference?
35:03 >> I use the claude sub for the fable stuff and I use my codec sub for the soul stuff. So, these numbers are not accurate, but it's around $200 if you're paying per token to run six challenges. >> And Kimmy is like around like Yeah. Okay. Not bad. Okay. which is crazy because if you look at the last one that I did, where is it? I think it's here. This one also cost $200 and I used dumber models and and fewer challenges, but it still the cost got up to I don't where's the final cost here?
35:38 Yeah, it still cost $200 because those models wrote like way more code. Like Opus 5 wrote 30,000 lines of code for half as many challenges and 15,000 lines of tests. >> Dude, I'm gonna run Slop Code Bench and just have it run write BML code and see what it does. That would be so fascinating. I think >> yeah, you should. I mean, I think what's nice about Slop Code Bench is like the the code quality metrics are all like implemented specifically for Python, but all the tests are designed to be blackbox where like the model gives a CLI out and you can test it.
36:10 Yeah, that would be super interesting. We should do that on another episode. We should just like that'll be the next pair programming one. We'll just fire it up. There'll be a lot of waiting. there's a lot of waiting in this stuff, but that would be fun. >> I would be I would love to bring that up. >> okay, what else is >> fascinating? It's I think the fascinating thing about all these problems is actually like look software is a incredibly rich dimensional system.
36:34 >> Yes. >> So like there's so many dimension you play. You can change the model, you can change the problem, you can change like the harness, you can change the language, >> and like you can change the prompt, >> and there's so much you can do and like so I don't actually think these problems will get solved because they're so multi-dimensional and they optimize for different things, >> but it's interesting and it, you know, entertains my dumb little what's it called? Goldfish brain for a little bit of time at a time.
36:58 >> It's fun to look at data and I like this. So I decided I really like this poy mandress theme which is a theme created by an autonomous coder collective or something I don't know. Anyways it's a good theme. so this chart is the number of functions that were written across the lifetime of a challenge >> cut against the mean complexity of the functions. And so you have again this like okay if you have lots of more functions those functions are probably simpler and if you have less functions then they may end up being way more complex.
37:28 >> That's cool. Yeah. >> this one I can't there's there's a kind of a line going down and to the right, but it's it's hard to see exactly any any winner or loser here. >> Let's let's recap really quick unless there's more stuff you want to go. >> There's one more. So, I want to talk about like what we do next. So, the thing you said is really interesting of like, hey, I want to what was it?
37:48 What was your proposal? >> You smash all the prompts together in one big goal very upront. >> See see if it does better if you do all of it at once. And if it does, then you guys just need to hire people that can do all of that much faster. >> Yeah. That can do the planning, right? >> That can that can pick the right thing to build. >> I mean, people with good product sense and good engineering sense are generally highly valuable. And I my suspicion is the data would prove that.
38:16 >> Yeah. So my my thought is basically in order to get a ultra strict pass after checkpoint one, you hand a smaller model the same codebase. You give it like sonnet 5 and you let the smaller model go try to do checkpoint two and checkpoint two has to fully pass like basically like the frontier model has to set up the codebase in a way that a small model can solve the next problem. otherwise it doesn't get a full pass.
38:45 Basically this thing has to ace checkpoint two. >> You would have to prompt it differently. Like if you knew you were going to hand off your code to someone else you would code differently, right? >> Yeah. I mean that's the other thing is like so that's part of it giving the whole problem up front having a smaller model do it. again, it's like it's not like how do you how do you raise the score on this, right? It's like what would tell me that the frontier models are good enough to go lights off? That's the whole point of this entire exercise, right? Is like what signal would I have to see in the benchmark data to tell me, oh my god, a model wrote four checkpoints of code and then GPTOSS 12B was able to solve the next checkpoint.
39:26 That is a sign that >> Yeah. Go ahead. That's interesting. That's interesting. I'm That's such an interesting question. What would convince me that I can go completely lights off and have no humans? >> Yeah. So, we're trying to figure that out >> or And I think and that's such an incendiary question, I'm sure, to many people. But even if I were to break it down, like how do I go how do I go lights off and have no humans for certain kinds of problems? Like, for example, managing my website, managing my ad spend, managing >> Exactly. Yeah, any subsystem.
40:00 That's such an interesting problem. I don't know what I would >> need. I'd probably need probably what I need from a candidate when I interview them, which is like >> during the interview process, can you show me one unique insight off of the little data that you have that I have not had myself? And if you can, >> I trust you as a candidate to like work at our company. I mean obviously if you prove me wrong that we that might happen but like that's kind of all you go for.
40:29 >> Yeah. Is where is the alpha? >> Yeah the alpha is showing human reasoning over systems. Well, no. What I'm saying is if a model could give me that insight too. Like yeah, if during the progress of this it implemented it and it was pro even if it was prompted to hey produce some insight and only share it if it's truly unique and the other person isn't going to have and it shared something that was like hey this is a background insight I have I' I'd probably trust that system.
40:55 >> I don't know about you. >> Yeah. I mean we talked about this when we were initially benchmarking Fable, right? It was like hey can this thing tell me something that I missed? like can this that that would impress me and I think you did not find anything interesting but >> no there was nothing interesting and I mean but the coolest the workflow stuff in Babel is great for compressing time. >> Yeah. So yeah, the the three things that I think I would add to this, right, are you know, bench next, right, is like number one, what you said is like give the whole problem up front and compare.
41:35 Number two is have a smaller model try check >> point N + one. If it aces, you get a super strict pass. >> Oh. >> and then the third one I think is apply the because this is what people do in real life is like they apply deterministic litters as a feedback loop during the challenge. >> Yes. I would Yeah. It's probably an unfaithful test without that.
42:08 >> Exactly. Right. Because most people who are saying, "Oh, lights off software factory," they're having other models review the code before it goes in and they're having deterministic llinters set the cyclomatic complexity cap to some number to force the model to write clean code. I don't know if it would actually make things better or not. Like I'm very curious to see what this looks like. It might make them worse if you say, for example, every time Fable writes a thing, we have Soul review it.
42:32 And every time Soul writes a thing, we have Fable review it. >> Yeah. I mean for example for us like we run code rhyme in all our PRs and that gives us a lot of confidence to let stuff run. We run CI/CD. We have a pretty good CI/CD system I would say. >> And that's probably what this that's like number three is like >> that's probably like the true earnest feeling that you really get which is like >> yeah I I think that's good. On the other hand like if you're it depends on who your ICP is. If your ICP is like regular people writing code I actually don't think three is three matters as much. I think one matters. Not one.
43:07 >> Yeah. I care about the frontier of the software factory. I think we're all right now in this in this year 2026. Everyone is trying to figure out what code do you still have to read, what code can you not read, how do you maximize and like how do you avoid disaster? How do you maximize confidence? And how do you like maximize the chance that you won't have some terrible outage that you can't fix because the code got too nasty? You know what? I really think the answer here is you optimize for really quick fixes and patchability as opposed to getting it right the first time around.
43:43 >> Oh, really? You mean the thing I said where you ship the thing and iterate instead of designing the whole thing up front? >> Well, no, you also get it right as much as you can, but you also Sorry, you're right. Never mind. >> Yeah. Thanks, buddy. It's okay. I'm never right on this show, so I I got I got to take them when I can get them. >> >> Tanner's got a really interesting question that I think is good which is like how do you expect configurations like MCPS or skills to influence a benchmark like this? What's your bet? If you had to go off the cuff right now, MCPs make it better where skills make it better or worse.
44:15 >> I mean like it it depends. >> correct. >> I think I think skills are I think skills are hard, right? Like Python is in the weights deeply. if you're going to go write it in BAML, like yes, put the BAML skill in there and you'll get better results. there's some ideas of like, hey, if you can build a really context efficient and high quality code search MCP that is like less tokens and like less searching and finds the right thing faster and basically like uses sub like a a mini like what the morph what Tis and Morph does. They have like an MCP that is a tiny model that's really good at code search. like >> if you index the codebase ahead of time and you never use GP like yes every single thing in the harness can in influence this outcome and like if you're interested in in trying that like yeah go customize the open code harness add an extra skill it was not hard to get Claude to like read the paper pull down the slop codebench repo and make the changes I needed to make so like you can do this >> yeah my personal recommendation is for anyone writing MCPS and skills for entire organizations Try your best to have the minimum set in the context window of every engineer.
45:24 Like it's not to say like don't have any. They're useful, but like >> it's a very high bar what's allowed to be everywhere because it steers everyone and as the models get better, you're really fighting the models more than agreeing with them. >> And I I think this is actually like I'm gonna I'm gonna ask you Vivov another like slightly tangential question, but I think it's very very topical right now with the release of Fable Opus 5 and GPT 5.6. six. what do you think of Boris Churnney's advice and a couple other people's advice of like every time there's a new model, you should throw out your agents MD and your cloud MD and you should throw out all your skills.
46:01 >> I think you should always throw out your agents MD. You should not have an agent MD in your codebase. >> What about the skills? >> I view skills as purpose-built. Like personally, I don't have that many skills in my codebase except for things that I know are not in the model weights for sure. like to like >> that's >> proven to not be in the model weights. not I think they're not in the model weights. Like I am very very >> watched it hill climb and it up over and over again.
46:27 >> Yeah. So in our codebase for example, our type system is very strict when we write BAML code in our Rust codebase. And if you let Opus go wild or Fable go wild or Soul go wild, it will at some point get something wrong because they're bad at they're bad at long mathematical chains where they actually have to like do computation to figure out the right result. And like really what we do is like we reh what Kai on our team has done is he's written a doc that forces the model to be rigorous about the type system and that's been a big godsend to making the model actually work well. But for and if we don't do that the codebase basically veers off turns into like this fragmented system where it has like 50 different type checkers and it breaks the rules in subtle places but fixes other it turns into a slop mess. So I think in that sort of place where it's a hard hard problem and it's a novel problem where we're evolving our types as we go it's never going to be in the skill set so it needs to be sorry it's never going to be in the weights. It needs to be in the train data. Yeah, I think this is the key thing is like >> skills.
47:31 >> Yeah, the skills exist to give the model things that are not in the weights to do default like if you if you say like people used to put in their cloud MD always run the tests after every change. >> You don't need that anymore because the models just do it. one example that we have is like we used to have an agent browser skill with agent browser is a CLI that lets the model like work with a browser. It's very context efficient.
47:54 It's good. It's from versel. we used to have the skill in the repo and we actually threw it out because the model is now knows about agent browser. No, it's same with the GitHub CLI. Like no one would put a GitHub CLI skill in their repo because you've never seen the model screw up a GitHub CLI command. It always gets it right. And so like this is kind of how I think about it is like every time there's a new model you have created skills because you've watched the model try to hill climb against an outcome and you actually end up needing to I guess like you have to put in extra context so that it knows what to do. And so if you have specific things about your repo or specific things about your type system that are in the weights or unlikely to get into the weights, you should definitely do it. But this advice of like don't don't just cargo cult all your skills from model to model the same way you wouldn't cargo cult all your prompts from model to model is is very interesting advice. Like I don't think I go so far as like throw everything out but whenever a new model comes out like user intuition. We went through all of our skills and changed most of the descriptions to because a lot of our description was like you must use this skill if X and the new models Soul and Fable love using skills. They're very good at calling skills. And so we literally had to change like 80% of our skills to be like do not invoke this skill unless the user requests it by name. That is the description of most of our skills because most of the time we're just steering to like use this skill to do this thing.
49:16 >> My yeah my recommendation to people is I I actually 100% agree with Boris u delete all agents MD delete all skills by default and the burden is on you the developer and the team to prove that you need them. And like one reason to need them is, hey, you have an edge team of 10, 20 people that are all doing the same workflow and you want to standardize it. You want to lift the mean up to a certain point. Great reason to add a skill. You want to help build metrics. So you have a core engineering team who's building skills for your company and they're measuring things and making datadriven decisions about what skill is good. Great. Shared skills, shared >> evals on that. Right. Exactly. the same context the the memory thing we were working on two weeks ago of like hey go see what engineers have to keep correcting and build that into your skills or context shards or whatever >> exactly like if you're not everything else like no like just don't hurt yourself you're making your product worse you're making your engineering goals worse >> I actually think when we were going over that project we had kind of I had kind of mocked it up with like information versus instruction of like hey here's how you run the test for this thing versus information of like, oh, our main branch is called Canary instead of instead of main or whatever it is.
50:33 >> Like >> I actually think theformational ones tend to be the ones that get moved across model generations because they're very clearly like here's a thing you don't know and you couldn't know but you need to know in most cases versus the instructions are like here's how to run the rust tests and not like try six different things is like >> that might not need to be carried forward. That might be the kind of thing that the model can just figure out and you're actually d-tuning the model and like lowering your performance by having extra instructions in there.
51:01 >> I'll give you like an example of that. In our codebase, we run a massive monor repo. Monor repos are good and bad. >> The bad part is a model will try and do everything in every subfolder. You don't want that. >> We have a subfolder called BML language >> that belongs in a agent MB that belongs in a skill where it's like, hey, like 99.9999% of the time you're only working in this subfolder. Ignore the rest.
51:24 All right, unless explicitly mentioned and that's a great skill because monory it defaults the model to a good path by default. >> Iman asks, "How do you make claude and codex automatically review each other's work?" You can orchestrate this in a hundred ways. You can use a bash script. You can use a Ralph loop. You can tell I I've told CL you can tell like your claud CLI. You can tell your codeex model to use the Claude CLI. They know how to do this now.
51:49 >> I haven't used folks yet, but like they they apparently have Greile apparently has a way to like check what you're using and shipping and automatically uses the other model to check it, which is kind of interesting. >> Oh, cool. So, you put up a PR and they know, oh, this was made by Claude, so we're going to review it with Opus or with with >> Codeex. Yeah, they kind of have stuff like that. >> Cool. I think that's probably time.
52:12 What skills do you have? That's going to have to be for another episode, unfortunately. >> Yes, indeed. this was tons of fun, folks. Thank you guys for joining. and hopefully you guys found some interesting conversations between the Yaps >> every Tuesday 10:15 a.m. Pacific. >> plus minus or not plusus plus 5 minutes. Yeah, plus zero plus 0 to 5 minutes. >> yeah, we will never be early. Sorry, folks.
Summary
- Slop Code Bench measures AI models' ability to incrementally add complexity to coding tasks, mimicking real-world software development.
- The benchmark is unsaturated, with top models achieving only a 33% pass rate, indicating room for improvement.
- Key metrics include strict pass rates, defect rates, and code quality indicators like cyclomatic complexity and duplication.
- The hosts discuss the importance of understanding benchmarks and the implications of different models' performances on coding quality.
- Incremental coding tasks reveal how models handle evolving requirements, with some models excelling in specific challenges but struggling in others.
- The conversation touches on the potential for future improvements in benchmarking, including testing models with complete specifications upfront.
- The hosts emphasize the need for effective collaboration between models and human engineers to optimize coding outcomes.
Questions Answered
What is the significance of the Slop Code Bench benchmark?
The Slop Code Bench is a new benchmark that evaluates code quality and performance of AI models in coding tasks. It highlights the pass rates and costs associated with different models, revealing that even the best models currently only achieve a 33% pass rate.
What is the difference between strict and loose pass rates?
Strict pass rates measure whether all tests in a suite are passed, while loose pass rates only check if the current challenge is passed. This distinction is crucial for evaluating the reliability of AI models in coding tasks.
How do different prompting strategies affect AI performance in coding tasks?
The study found that using a prompt to plan before implementing does not significantly improve results compared to directly solving the challenge. This suggests that the effectiveness of AI in coding may not heavily rely on the prompting strategy used.
Can AI make effective decisions in coding tasks?
Currently, there is skepticism about AI's ability to make good decisions in coding. While there is potential for improvement, many users have yet to see satisfactory decision-making outcomes from AI models.
What strategies can be employed to ensure code quality and confidence in AI-generated code?
To maximize code quality and minimize risks, it is suggested to focus on quick fixes and patchability rather than striving for perfection from the outset. This approach allows for iterative improvements and reduces the likelihood of catastrophic failures.