Section Insights
Introduction to AI That Works Unconference
What was discussed at the recent unconference?
The unconference brought together diverse experts to explore various topics related to AI, emphasizing the rapid evolution of practices and the need for continuous learning.
- Participants are constantly evolving their methods and practices.
- AI is accelerating the pace of change in the industry.
- The event featured discussions on a wide range of topics, highlighting the diversity of thought.
Challenges with Unattended AI Models
What are the risks of using AI models without oversight?
Unattended AI models tend to produce poor quality code, leading to chaotic outcomes unless actively managed.
- All participants agree that unattended models produce subpar results.
- There is a consensus on the necessity of oversight in AI coding.
- The conversation highlights the importance of maintaining quality control in AI outputs.
Evaluating AI Implementation Loops
How can we effectively evaluate AI implementation?
By breaking down complex problems into smaller, manageable parts, we can validate and optimize the implementation process more effectively.
- Closed-loop evaluations can enhance the efficiency of AI implementations.
- Identifying bottlenecks in planning is crucial for improvement.
- Smaller problem-solving approaches can lead to better overall performance.
Complexity in Code Updates
What are the risks associated with updating code in complex systems?
As systems grow in complexity, the likelihood of missing updates increases, which can lead to significant errors in the code.
- Complex systems are prone to human error during updates.
- The more places that need updating, the higher the risk of mistakes.
- Understanding the structure of code is essential to prevent errors.
Best Practices for Artifact Management
What is the recommended approach for managing code artifacts?
It's advised to store artifacts outside of version control systems like Git, using alternative systems for better organization and accessibility.
- Storing artifacts in Git can lead to inefficiencies.
- Using external systems for artifact management can streamline workflows.
- Integrating cloud solutions can enhance visibility and access to documentation.
Transcript
0:00 We hosted the unconferences last weekend. Every single person is doing something differently than every other person and every single person is changing their mechanism every few months. >> If you can't look back two years and want to kick your own ass, you're probably not learning fast enough. And AI is just compressing that timeline. >> No one is operating the same way that they used to be operating. And no one expects to be operating in the same way they're currently operating.
0:22 >> It's actually really hard to tell who's who's ahead of whom or if the problems are different. It's a very multi-dimensional space. Today's episode we get to talk about the AI that works unconference which is an event that we host about every 3 to four months where we put a bunch of the smartest people that we know in a room together and we just let them build the agenda. We had all sorts of topics and we recaped some of them from cash engineering to slot factories to harness engineering and all the way to building evals to some degree and we talked about some of the learnings that we had what was consistent what was not consistent across engineers in different domains different size of companies and different edgeorgs. If you're interested in hearing what happened, check out this episode and we'll show some examples and eventually we'll at the very end we hit the live Q&A with the audience and see what kinds of questions came up. Let's get started.
1:07 >> I'm so excited to be here with you talking about AI that works. This is the AI that works show, ladies and gentlemen. We are going to talk about things that work in prod. My name is Dex Horthy and I am the CEO and co-founder of Human Layer. We build a multiplayer a collaborative agent coding workspace and I'm joined by Vib. Vib, please introduce yourself.
1:38 >> My name is Vibbof and I'm the co-founder of Boundary and we build a programming language and an observability suite for building code bases that you don't have to read every line of code. >> Wow. Wow. Really? choosing violence today, huh? >> we're just going to start with the violence. >> when dude, when we started this show, we you we I we were not I'm the read the code guy and you're the don't read the code. If anything, it was the opposite.
2:08 If anything, we swapped places. >> Hey, you know, it's a eb and flow of yin and yang. >> six months from now, I'll be saying the opposite and you'll be saying the opposite, and that's okay. I mean, I think the beauty of the world as it goes right now is that it changes so fast. And I personally find that the value of >> like literally just flipping your perspective every now and then to the opposite extreme is probably the healthiest way to really challenge the system as it is today. I don't I just think like if you're trying to stride the middle, you're basically falling behind.
2:38 >> If you were not changing your mind every 3 minutes, you were falling behind. Sorry, every 3 months. Not every 3 minutes. Sorry. It has like I offered a slightly slower time scale than you. So maybe maybe 3 months for me, 3 minutes for you. >> I I had a buddy in college who used to say, you know, if you can't look back two years and want to kick your own ass, you're probably not learning fast enough and AI is just compressing that timeline.
3:02 >> What's your Immod's got a fun question. What do you think about clawed watermark stuff? >> We already have the watermark. It's loadbearing it's a trap. Like if you need a watermark to tell what text was written by Claude, you're not paying attention, buddy. >> Wait, wait, wait. So, catch me up on this. Claude is watermarking text, not just >> they they are putting stuff. So, you know how like when you copy structured text on your laptop and sometimes you paste it into something like if you copy markdown from GitHub and you paste it into Google Docs, it like brings the styles with it. It's like actually it's like, you know, pasting XML and then the the client reads it. They're I guess adding metadata and tags and stuff where basically if you copy text out of cloud and you paste it into something that like supports rich text, there will be little like smidges of >> interesting >> and it will persist through >> it'll persist through some editing. So if you if you edit the right if you delete the right block of text or change it some way, you'll lose that metadata if the client doesn't support it. But I think they're having some something something like that. It'll be like little spans and divs that like change the styles. but yeah, they posted a thing.
4:13 >> Why do you think they're doing that? >> because the EU told them they had to. >> Okay. Well, there you go. That's a good >> Now, why is the EU doing that? you know, just Europe doing Europe things. I love Europe for the record. but this is a very EU coded coded move. >> Interesting. You're trying to do things for the rest of the world, I guess. >> Yep. Signed Provenence metadata. Yeah, it will sign any image it generates gets Yeah. I I mean, I think the images and the video stuff is way more important of like, hey, look, people spreading misinformation with AI using images of things that didn't happen is actually like a useful thing.
4:55 >> Yeah. So like for example, new models will mark AI generated content from day one. Words everywhere you use claude will help detect claude marks. So I guess idea is you can query it very easily to know if it's clawed. This is going to help systems like Reddit and stuff basically like remove like cloud claude speak directly into the stuff which I think is not a bad place if the platform opts into doing that. >> Well, so it's not cloudspeak. It's it's it's other it's like unreadable metadata.
5:22 >> Yeah, that's what I mean. So that's why like Reddit can probably do this in a better way because they're like if they I assume you have a programmatic API that Reddit could check on the back end to know if something is like cloudspeak. >> I see. Yeah. Yeah. I mean it depends. It depends if it's a rich text editor. I'm guessing >> if it's an API I suspect they'll have ways >> like and Reddit does have a rich text editor.
5:45 >> What if they're going to inject like Unicode characters that match the same like they look the same but they're actually Providence metadata for JPEG. I mean, honestly, you know, I I'm kind of torn about all this because like it's telling you it's going to generate it for SVGs, PGs, JPEGs. And like the reason I'm torn is because I'm asking myself a very very important question, which is do I even believe that people care? And the analogy to this, I think, is like cameras. Cameras and locker rooms used to be banned. Now we got Tik Tok reels about locker rooms. And like society just moved on and accepted it as like a norm rather than trying to fight it. As much as I hate >> that's a very interesting example, huh?
6:25 >> I mean, as much as I hate to say it, like I think people are going to be watching content that is fully AI generated and not care. >> Like AI generated, human delivered is probably what I was going to for a while. And I see YouTube already doing this in so many different ways. I watch channels like you can hear the claudisms out of their mouth and they're literally reading out of a teleprompter talking about like talk verbatim saying the script that Claude wrote for them.
6:53 >> No, I was at a wedding this year for the first time and I was like is that a chat GPT tick in your wedding speech? Are you serious? >> I'm not sure. Unconfirmed. But like it was the first time I was like on alert for it. I don't know. It's also like you get to this point of like okay cool. I can I can generate an image with claude and then I can pull it up and I can take out the certific certification that came from claude. Right? You can always remove metadata from the image.
7:16 >> So what that comes down to is now like okay if you take a photo with your iPhone your iPhone's going to sign it using like a hardware like private key basically that proves that this was taken on an iPhone with a real camera. But that stuff can all be cracked. Like there is no it's just like DRM. Like there is no like guaranteed failsafe way to prove that a thing was really made. I think DRM are really good for the majority of use cases like >> No, they're not. People stopped using them because people just kept cracking them.
7:44 >> Yeah. >> Like people don't use DRM anymore >> for a while or do they really not on Netflix and stuff? Like for example, if you're on Apple and you're trying to like FaceTime and share your screen and you're watching TV, it instantly gets blocked like Appleider. >> Yeah. So I mean DRM like I guess Yeah, that that's true. They still like blocks if you're watching something on Hulu and you take a screen cap, it blocks it cuz it's like if you're in the the blast ecosystem, okay, and if you don't have that, then like the Netflix app does the same thing too, right?
8:13 >> I I don't know. I when I hear DRM, I I think of like how they used to like if you if you bought a song from iTunes and then you sent it to somebody, that person couldn't play the file unless they like were logged into your iTunes account. >> Yeah, that stuff I agree. But like the actual like premise that you're going to encode data I think anyway we're way off topic. This is not what we intend to spend 10 minutes talking about. Let's go into what we're going to talk about today which hopefully many of you got the chance to check it out. We or some of you perhaps didn't. We hosted the unconference this last weekend and I think what today's episode of is going to be about is actually going to be recapping what we learned about the unconference topics that came up. Sadly, we do chat and mass rule, so we can't say who did what because we don't yet have approval from every single person out there. I suspect some people spoke about it and what that they were there and that they mentioned it, so we can mention them probably.
9:06 But in general, it was a very fun vibe of learning a lot of different techniques. I think the only the one big takeaway I had and I don't know about you Dexter was that we're all still figuring stuff out. Every single person is doing something differently than every other person and every single person is changing their mechanism every few months. Like no one is operating the same way that they used to be operating and no one expects to be operating in the same way they're currently operating. And the crazy thing is like there is a spectrum of like the the software factory stuff. There's like a spectrum of like read all the code versus like just make your factory work and like use feedback loops and use code to do it. And what's interesting is like sometimes you would think of that as like a linear journey. But like people move back and forth across that spectrum and people are on different sides and it's like it's it's actually really hard to tell who's who's ahead of whom or if the problems are different. It's a very multi-dimensional space. I think the one thing that was very clear was that everyone accepts that if you let models like like run unattended, they will write slop code and you have to be you either have to accept it or you have to do something about it.
10:21 >> Yes. I think there's zero people in the world that don't believe that. It's like the convergence of models running code is slop. Like they converge to I guess they diverge to slop is a better framing. >> Yeah. is what is the entropy thing of like everything everything naturally devolves into chaos. >> Yeah. I yeah I don't think there was a single person that said I mean there was I can't say who but there was someone that was like you can just write code for like forever without reading anything. that was not me by the way.
10:51 I do believe you have to read a minuscule amount of code. >> Yeah. yeah, it was like a just merge merge without review automerge everything kind of thing. It's like yeah, but I mean some people are writing some people I someone is saying they wrote like >> 70 deterministic llinters for Rust. Like they just they just vibe coded a bunch of tools that like look for every single anti-attern. And every time they see an antiattern, they don't update their agents memory. They don't update their prompts or their skills. They actually just go build another llinter so that the next time someone writes code like this, it gets like fed back to the agent.
11:21 >> Yeah. >> Yeah. and tells the agent like, "Hey, here's here's what you did wrong." >> I had a really interesting conversation last night, kind of following up from that where we were talking about like, how do you design llinters in the world of AI? >> And I think the missing thing out of most llinters is some sort of natural language. Like, and the thing is, >> oh yeah, cuz someone was saying, yeah, it was like Deepseek V4 Flash is like small, fast, cheap, and smart. And like if you just write if you can just write a list of a hundred rules of like this should never happen, this should never happen, this should never happen, >> like easier to maintain and like almost as good as a deterministic llinter and like yeah, it's more expensive and it's slower, but like it's close enough that the trade-off is like it's you're you're right on the >> pay let's say you're paying $1,000 per feature to ship slop. Is it worth it to pay like $5 to review that slot?
12:10 Probably. >> Like like $5 to have a lint to run on. It's like 100% a trade-off that I think people would make. >> And I mean, my take with all of this is like it's it's not like your goal should not be to get to like I mean lights it's always good to like set a goal that's farther than you actually are going to end up, right? If you're like, "Hey, I'm going to make a billion dollars this year and you make 200 million, then the goal worked, right?" Like, yeah, you didn't hit the number you said, but like setting ambitious goals is a way to like because you're always, you know, it's like the you halfway and halfway and halfway. like the higher your goal the the the more progress you'll make.
12:45 And so that idea of like I don't if you say your goal lights off fine but like be okay with like hey we went from 50 PR you know 50% of PR is being merged without without any feedback to 90% of PRs being merged without any feedback. But until until you're actually merged a thousand PRs in a month or two and you read all the code and you have no feedback, you should not like stop reading the code until you're very confident of like yes, this has been working for long enough that it's like cuz I don't know Kyle made this tweet this morning here. Let me pull it up. It was just like I think it's already got like 50,000 views. I don't think one should be merging code at a faster rate with AI than they what they were merging before unless they actually have a grasp of the system and they believe that the system is converging.
13:31 >> If you're not reading the code you have no idea how much garbage is in your code. Like you can't say like oh we don't read the code the models are good enough. It's like no if you don't read the code then you don't know you like you cannot say that it's good or not. >> Scroll down. Scroll down. Scroll down right there. >> Right there. >> Yeah. Read the least amount of code you need to ship reliable software and a large enough code base or best engineers don't read all the code. Many more worse engineers ship slob all the time. Yeah, companies work. Anyways, that was like one part of the discourse. What I liked about the unconference though was it was like the audience built the agenda and there was lots of people who like kind of care about code but people shared like their take on Andre Karpathy's LM wiki and how they maintain a like recurring improving memory system over time that like how you can have like all the right information in your context window back to context engineering and stuff. There were people talking about how to build coding harnesses for hardware and build the right feedback loops there.
14:24 >> Do you remember? >> Oh, you got schedule. Nice. real one one person who was there Quinn Quinn law from the pipecat and >> yeah he's always he's he's technically I'm sure he also he gave he gave this talk well he gave the talk >> at AI engineer so but yeah that was good migrating large code bases personal context layer we talked about KV cache engineering like some really like down to the metal like >> I missed that one was it dope Yeah.
14:58 >> I've looked at this stuff before, but I think if you haven't, I think the biggest takeaway is you don't probably realize how much >> things like I learned something new. For example, if you change the reasoning midway, did you know that you invalidate your full cache? >> Like, >> you change the reasoning in the middle >> all the way back. You invalidate the full cache. >> Well, so my my my understanding was like there's certain ways to use the APIs where they will either keep all the reasoning in or they will only keep the most recent reasoning block. No, >> but >> if you take out the reasoning block here, you still keep the previous cache, right?
15:31 >> Nope, you don't. I'll show you why. reason >> that's that's a bad API design. >> No, I'll show you why. GLM grammar. Let's see if I can find it. >> I'm going to make you a whiteboard in case you want to draw it, by the way. >> No, no. I'm just going to pull up chatt for the prompt. Can you get it for me? Yeah, I'll just find out like like the grammar of text for what it looks like for GLM or like for GLM or for OSS 12B.
16:06 >> Yeah. >> or get it for GPT1 12B. I just want to look at the prompt and like the grammar and the prompt. It's I know it's somewhere. Can you just access it online? I just want to show the template of what's going on for it. >> By the way, we only see we only see the unconference thing if you want. >> I'll show this in a second. I don't know. How did you make did you you took a picture of the all the post-its on the whiteboard and this is amazing >> I just take a picture I was like hey Claude do this and it made some mistakes >> but we were able to go fix that along the way.
16:36 >> Yeah. >> and that was pretty like it made very few mistakes. >> Oh yeah. the the the other one I really liked is like kind of the more like human take on the sharing learnings with your engineering team of like hey look everybody's on their own journey learning AI and building like competence with using it for both technical and non-technical work and a lot of people using AI for non-technical work are technical people >> so I'm going to I'm going to share a different screen really fast and like you'll see what I mean so like check this out so like the grammar often has this so reasoning is often one of the earliest messages that goes off into any prompt. So this is like the actual >> the start and the message things are basically those are like special tokens.
17:18 >> Yeah, these are special tokens like well maybe start and system message is a special token. I suspect this is a special token if I were them. But anyway, the point here is like >> see this reasoning stuff right over here. >> Yeah, >> you change the reasoning. You change one of the earliest tokens in your whole prompt. Everything is invalidated. >> No. Oh, I see. You don't mean changing the content of the reasoning. You mean changing the changing the content of when you change the reasoning level?
17:42 That's crazy. >> yeah. >> Why would they make it like that? >> some other interesting facts that I learned is like the reasoning tokens aren't actually hard-coded. So like there's some models that only have medium and high. You can also pass low into them and they mostly work because like they're trained on semantic meaning of what it means. >> At least empirically. Obviously, you're not going to get the same results. Like if you really think about what's happening here is like the model gets trained on like this token yields to some amount of reasoning tokens being favored or not favored >> in the post- training loop. So like when it does that obviously put in low it hasn't it has an approximate meaning of what that means.
18:18 >> So it's allowed to generate so many tokens before it prefers to generate the stop reasoning token. >> Yeah. This is like any other instruction of like hey run the test after you change code. It's like it changes the chance that the next token will be more reasoning or reasoning end and into the actual like messaging and tool calling. >> Exactly. So that was that was really interesting talk. I thought I don't know which of these did you get to attend that you thought was interesting. I love the building one too.
18:48 >> I >> that was >> I was running around doing conference stuff. what how what did you think of the eval talk? What was >> eval talk? I thought the most fascinating part was building evals for your skills. Building evals for skills was a fantastic takeaway. I had not thought of him that way of like if you're going to build a software factory. The idea is if you're going to build a software factory, you can't actually eval most of the stuff that you're trying to do.
19:16 >> it's such like it's like long long context like long like like long horizon tasks and like you can't just look at the trace and say like this is good or bad. >> Yeah. So let me join the room and let me scal really fast. So like the idea is like look there's like two parts of software factory. is planning where you basically build specs and once your specs are complete then you can build like implement one fairly consistent conclusion about the whole thing is that the implement loop is solved like you like coding agents can implement most things and you can eval this pretty well like right here you have some skills you have some other stuff that you end up doing you have like skills you have your environment that you're going to set up >> yeah you have your test and winters >> yeah like all the stuff around that like tests llinters and like if this is bad this is only bad due to a couple of things which is either your skills are wrong your plan was wrong or your environment is set up incorrectly and in that world you can really even you can do a closed loop eval in the implement loop really really fast so if you can do that then the next part just becomes like can you merge and I what we find the biggest bottleneck was really the planning side but The whole part is like because this is so evvalable then what you can also do is you can also go ahead and actually you can also go ahead and actually just validate the implementation loop.
20:45 >> Interesting. Oh yeah because you can take a plan that you already know was implemented well change the skill and then just see if it still implements well. >> Exactly. Exactly. And that's one thing that I found to be very very interesting. >> I mean this is what we talked about on the very first episode of AI that works was like how do you take this large problem and break it down into like things that you can probe at a lower level so you can optimize smaller parts of it and you can figure out like okay the end end is slow and expensive to run and it's like there's lots of edge cases how can I break it down into smaller problems >> another thing verification loops like adversarial review I I'll call that.
21:27 >> Oh my god, I'm so sick of that word. >> Yeah, I know. >> It is not a magic word. Saying the word adversarial does not magically make the code better. I mean, it will make the code slightly better and it will catch all the dumb stuff, but like adversarial review does not solve your problem. >> Yeah. Yeah. But almost everyone had some variant of adversarial review in their process. Almost anyone that was shipping code with AI, they were always like having some sort of loop at the very end. And I think the idea is not >> the idea is not that you should do this as a the idea is like you're going to do oops >> the idea is you're going to do something then you'll have like a loop like an ad loop and then the idea is that this thing will loop before it goes to a human.
22:10 >> You want to delay going to the human as much as possible and this this just increases the odds of like it being correct for the human by the time it gets there. and I thought that was fascinating that almost every independent whe whatever they called it, different people called the different things always had that kind of process coming in. Scorecards for hill climbing. I didn't check that one. Slopfree software factories was >> that was one I went to that was I won't say who it was but it was very good. they talked about like how they have basically like three parts of their loop which is like discuss and decide and brainstorm all of the planning all the parts in planning of like one thing and then they have implement and then they basically have compound. They only have three skills. They have like figure out what we're going to do at a high level that gives us leverage over what's coming next. Do the implementation and then compound is like hey review what what went wrong. Pull in PR comments.
23:02 Pull in agent traces. is pull in user messages and then like figure out and they do like once a week they just get a report of like here's all the things that went wrong this week that like the agent didn't solve and that humans needed to get involved for like prioritize and then they just pick a couple things every week to like build that into the implement system or into the planning system. >> I didn't check out I I actually missed these because I was in the middle of something. but I I mean obviously harness engineering there people we just shared actually I think this they posted about this so I can pull this up.
23:34 >> Oh the the orchestration primitives thing. >> I thought that was kind of nice. It was like some really nice diagrams that they made actually. >> This is Will is one of the OG BAML guys. He built a coding agent in BAML like forever ago. >> Yeah. Yeah, but I think the coolest part is actually he he just like had it go and search through a bunch of stuff >> and like found a bunch of control flows and like I know we've all talked about these processes, but it's just nice to see the entire list of everything kind of >> built out. And what's really fascinating to me personally is I look at this I'm like okay well this literally just looks like code >> and like it's interesting >> this is how you would design a programming language. Well, I don't know about that, but this is definitely how you would like like these strategies are not new. And I think that's probably the most important part here is like none of these strategies of execution patterns are new at all.
24:24 >> So given that they're not new, go ahead. >> Did did you ever did you ever read design patterns that like C++ book? >> I'm actually illiterate and I don't really read, which is a very sad thing about myself to admit. At my first job, I had a 20-minute bus commute. And I would always read on the bus. So, I just read like 40 minutes a day. And I would read. That's how I learned Closure. That's how I read Clean Code. That's how I It's the only way I can read is like it was before phones were super addictive. or I just had managed to not get hit by that. I also had like for two years out of college, I was poor and so I just had a flip phone instead of an iPhone.
25:04 >> dude, that was wild. >> Yeah. Yeah. It was crazy, dude. T9. I'm Yeah. Okay. Pattern. Thank you for sharing that. >> Cool. >> anyways, yeah, design patterns is one of those books. I mean, they had like 40 things of like the observer pattern and the flyweight pattern. It was all these things that like basically like if you like people keep reinventing the same idea over and over again poorly with the wrong algorithm. And it's just like if you keep iterating on this code for a while, you will land at one of these 40 patterns. Factory pattern like Yeah.
25:38 >> Yeah. So, my big gripe with this thing, like it's really nice visually, and honestly, Will did a fantastic job. But my big gripe is just like there's not like 100 patterns. There's literally only like there's a reason like map reduce for example is just map and reduce and that breaks down everything you need to do in distributed systems and like pretty much all other computation workload you have to do. >> And the beauty of most systems is like the minimum that you can have. But it is useful to see concrete examples of certain patterns. And obviously I know map reduce has a couple more constructs but like at at a very high level it is map and reduced. I mean I do think that for many I mean in some ways data structure is very similar to kind of what you're saying texture. It's like while all boils down to the same 40 patterns it is useful to learn like the 100 most common applications of those 40 cap patterns even though that's way more than the 40 patterns themselves >> because it just makes you a better engineer. you if you can do it.
26:34 >> The problem is if you read too many design patterns or you read too much you know you download a Google library and you go spelunking through the Google code you're just like oh I'm going to just make an abstract singleton factory factory for everything and like it's not it's not done in bad faith it's just is like what will naturally h like if you just build the code you will naturally land on these design patterns and if you read about the design patterns you will naturally land on like a lot of probably.
27:01 Yeah, it is a software is a fine balancing act is what I would say. >> okay. So, unfortunately, I think Avery won't be able to join today. his AV is not working. I don't know why. >> I should tell him to just use my mic setup. >> You should just do his slides as if you're Avery. You should just steal his steal his >> valor. Okay. Maybe maybe what we can do is we can have Avery and perhaps one of the other speakers come on next time and actually give their talks because I thought some of the talks were phenomenally well done. Okay. I mean, I have a if you need 20 minutes, I have a talk I'm working on for a launch today that we could talk about that a lot of people seem to so we talked about this a little bit at the unconference. but like is a thing that I think like >> everybody is on this journey towards it's a little bit human layer coded, but it's like basically like where do you put your plan docs and like how do you organize them and where should they live and how do you work with them and like I think the journey that everybody goes on to get there. Was the other guy gonna come? Is he Is he in?
28:01 Did he reply? >> I was out all night and I did not message someone. >> Okay. >> So sadly, but >> I know >> I'm I'm down to charity to talk or honestly if people have questions, we're happy to take questions or like today can just be a light day as well. >> We can answer do we needed design patterns with AI? Isn't AI already knowing it? And I will say that the problem is AI knows it too well and it's in the second it's in the second class of the like, oh, I'm just going to use factory patterns for everything. or it's like I'm going to use an over complicated algorithm for a thing that doesn't need it.
28:33 >> I'll see if Avery has the slides. This I mean this just b Avery's thing just proves this basically. >> Like basically I think what we found is very similar to slop code bench where if you let the AI go wild it is not going to do a good job. >> Yeah. >> As far as I know. >> because it will mess up one out of 10 times. And every one out of 10 times it messes up means that if it's made a 100 decisions it's messed up 10 times.
28:56 >> It compounds, right? It's like, okay, if you put AI in a 100% clean codebase and it messes up one out of 10 times, well, the the the the odds that it's going to mess up is a function of like the cleanliness of the codebase that exists. So, if we come like >> I've got a good example. Yeah, Avery was kind enough to send me the thing. This is kind of interesting. So, like check out this like imagine building a thing that does ticket master and you have to do three things. You have to either open and by the way, all credit to Avery. If you like him, go follow him on Twitter.
29:25 He >> I thought you just said you weren't gonna steal his slides and do them for me. >> Yeah, but this is so good. He just sent it for me. but I'll have him do the full talk later. But you have open, held, and sold. You can represent this as two bulls. And both bulls can never be true. And the consequence of this is if you chose the left representation instead of the right representation, regardless of what it is, and this is Rus code, so excuse us. when you go add a new state, see here, you have one thing to update here. You have to now update 1 2 3 four places. And now you can take this up to the nth degree.
29:59 >> What are the odds that an agent will forget to update one of the places? >> The more complex the system, the more likely it will miss one state and therefore you will be sad. >> And your code will have >> 40 lines of code that can get way more complex with one mistake. Imagine 400,000 lines of code. >> This isn't 40. How many lines of code is this? 1 2 3 >> 4 5 6 7 8 9 10. It's like 20.
30:22 >> Yeah. like it's just guaranteed to mess up. >> Yeah. >> and I'm not I'm not going to steal any more of this thunder because come next week if you want to see the full talk on they've rememeasured slop and had 200 agents implement the same feature and did the stats on like how many what percentage of them got it wrong versus got it right. >> Yeah, that was fantastic. Matas has got a great question. Beyond separate models, what adversarial review approaches do you recommend? how using how many documents? I'll tell you what we do. And Dexter, I'd love your thoughts on this, too, because I'll learn something new. when it comes to separate models, so I don't really care about separate models personally. I just clear the context. That's fine enough. It doesn't really matter. But what I do care about is I'd like to frame adversarial review as a design constraint that probably the model skipped on. So in our case, performance is a very concrete measurable thing. The best way to do adversarial review for us is we just say, is this fast? and you just ask it like what would make this faster and it will literally just come up with better data structures and once you have a better data structure it often will actually clean up the code a lot more like funnily enough faster code is usually better code especially when it comes from like initial model code >> yeah I think I think I put it in this category of like there's this paper we've talked about before called don't waste your back pressure which is like >> if you are just reading all the code and trying to decide if it's correct then then like then yeah you're screwed. So like what you can do is like in the early part of the effort you can use a compiler to check stuff and now you as a human don't have to like check for the symbols in the compiler. you may bring in like a type system as well. I mean type system compiler is kind of the same. and then like basically >> all right fine.
32:08 if you're using a lang but like if you if you use C the compiler is doing the type checking. >> Yeah. Yeah. Yeah. Yeah. >> so if you're using a fake language with fake types, then yes, you need a type system separate from your compiler. >> like Python or TypeScript. Welcome to TypeScript where everything's made up and the types don't matter. >> well, yeah. Anyway, go on. >> This has MCP servers, but this is like using Playright, using external tools, giving the agent back pressure.
32:32 >> What do you guys do though in practice? what the last feature I was building was something where I basically I had this like vision for a thing and I just basically I wrote a PRD Sol Street. I'll pull up the PRD that we made. what is it? V3. So this is a very large task with lots of stuff. I'll talk about what I did. But the first thing we had was oh actually I need the mockup. Let me go find the mockup. I kind of stopped using the doc and I just used this mockup. But somewhere in here I started with just like hey here's the new prompt experience in human layer and then we iterated on this and then I just literally was like okay what does every drop down look like? What is every like ticket? What's the behavior when we're fetching stuff?
33:24 >> What are you sharing? Do you meant to share? >> Oh sorry I'm not sharing. ASR. wow. So, you guys have just been staring at the article. So, I built this PRD basically, which is like the new prompting experience in human layer. and then I went through and was like, okay, what does every drop down look like? What are the options? what happens? What are the loading states? How do we like generate stuff? What do errors look like? What does the directory selector look like?
33:49 >> Etc., etc. Like, we just kind of like went and figured out all of this. And then I just sent Fable on like an infinite loop to go to go implement this. Just like keep going, implement everything. I think it technically made a plan. >> Yeah, exactly. Use the browser to go through and do this. >> Then what I did was I played with it. I vibed it. I had it have codec. So it was like basically like it was like now launch a codec cli codeex cli to review take the feedback.
34:23 and I basically I did this the way I did this was I cued a bunch of messages. So I was like I was like at prompt.md which told it to like track its progress and look at the prd and things like this. And then I did slashcompact and then I said at prompt.md keep going >> and then I did like slashcompact. And so I queued all these up. So like I I don't even use Ralph loops anymore. I just use in human layer you can queue up a bunch of messages and they'll send one at a time and then it was like yeah you know CRLM which is our codeex review skill like review and fix etc. And I just did this like ran this overnight.
34:57 >> and then I came back in the morning and I like you know vibed some changes for for an hour or two. >> Yeah. And then you merged. >> Nope. I looked at the PR and it was 20,000 lines. and it was mostly isolated in its own tree, which is what we like to do when we're riffing out new features. We're like, "Cool. If you need to use shared code, you may not change the code that the main tree is using. If you need to change it like a tiny bit, you can, but don't add extra props. Don't add extra like parameters. Like, you will copy that function and put it in the new tree. And eventually, we'll just throw away the old tree." but we kind of force it to like repeat code if we don't know what the abstractions are. but so for this PR, I then basically like reverse engineered it into seven small PRs.
35:50 >> Seven small PRs of like 3 to 5K each. And so that was actually the interesting planning where we used the code as kind of like the spec of what we're building >> and then we built it into this like PR splitting plan which to build this actually took me like an hour or two because it was asking a bunch of questions and I actually like changed some of the behavior basically. but then this is basically for each one it's basically make a new work tree. Here's your reference code. When you want to copy files like copy the file and delete from it rather than reading the file and like writing the new code. and things like this. So, I'm I'm now on PR number four of this, but each of the PRs was only a couple thousand lines of the last one I merged last night was like 1500 lines. I vibed out the got the behavior right. So, I was like we prototyped to get the shape of it and then I went back and broke it down into things that like Kyle will actually review. And so, here is like one of the PRs that got merged.
36:45 so this is I don't know this was only a couple hundred lines, but >> we do a slightly different approach where it's like we actually do it like so you're doing it more aggressively up front. We do something very we do something slightly different which is like you kind of just merge it ahead of time and then you we kind of just bake in time as a part of engineering process to like refactor recur like retroactively.
37:08 >> Okay. And Matteas asked the question about the thing I was going to talk about right now, which is where do you keep the spec driven dev? Where do you keep your docs? Do you keep it in the repo alongside the code or >> in docs? So this is literally like the journey that everybody freaking goes on, so I'm going to share again. and I didn't finish drawing this, but maybe we can finish on the show. so basically like we hear all this thing about artifacts and we talk about artifacts as like markdown documents of like plans, research, specs, whatever you want to call it, right? so let me give this a border.
37:47 and like they're separate from the code, right? Usually they're like fairly short-lived. and they you use them to produce code. and the first thing that everyone ends up doing is they end up in your code. You have a plan.mmd and a review.mmd and maybe you have a couple for different features and you either get ignore them or you commit them because you want to come back to them later or you want to share them. Eventually you maybe like put them in a special directory. so you have like a folder for all your plans and then you start sharing this across your team and they like build up over time and then the model finds them but they're out of date. and so like it actually like slows you down instead of making you faster. Is it are you following this so far? Like did you go on this journey kind of?
38:29 >> I I had Yeah. I mean I Yes. People initially check in MD files and then you quickly realize you can't do this. >> Yeah. and so like the next thing people usually do is like you scope them by feature and then when the feature is shipped you try to clean up the old ones. You can do this automatically or manually. and then like the real challenge that I hate is like you have work tree one and work tree two and you have feature two over here and feature three over here and your co-orker has feature four and you constantly end up asking yourself like wait what branch was that artifact on? Like what branch was that plan on and like oh I want to reference that plan over here but it's from an older checkout. It's from a feature that didn't get merged yet. So I have to go like find it. it sucks.
39:12 It's hard. I do we this happened with a bunch of customers that we've worked with. the other thing too is like GitHub is really good for reviewing like diffs like when things are changing, >> but you well it's not even bad for markdown. Like it's good for markdown diffs, but like if you just want a document that's not changing that much, I can't comment on this in GitHub. Even if I go look at the raw code, like there's no way to comment on a on a thing. You can only comment on a change.
39:38 >> and so the b the takeaway here is like don't store your artifacts in git. Like put them somewhere else. and so what we like to do is we get ignore this. Some people put them in notion or in Google docs or vib. I know you guys have built a system to do this your like beep thing. and so you get ignore this and then basically you sync it with an external system. So basically what we've built is like there's this flow of like the agent like does like a is talking to the file system.
40:09 actually we'll say the agent calls calls a thing in the harness which touches the file system file system. and then the file system says like okay and then what we have is basically a hook that is >> yeah you guys sync like from the harness itself anytime you read certain files in that plans folder effectively you kind of sync it with like a cloud artifact. Yeah.
40:39 >> Yeah. And so this actually gets pushed to an API where we store the file contents and then you can see them in a web UI. so the hook will like sync this and then this goes back and then the harness says okay the hook is done and then it passes control back to the agent. I'll call this the model technically since the agent is the harness plus the model. and then over here you can serve these and look at them. This is how this is how human layer stuff works when I'm all these docs I'm showing you. That's exactly how these were created. and what's cool there is like the other side of it that we also like to do here is oh, sorry, I stopped sharing my whole screen again, but we'll pop back.
41:20 actually, let me just take a screenshot of this doc. So, like here's our doc here. and so you can view these in the UI now. and then the other thing that happens is like another model running in another harness. Basically what we have built into our sort of meta harness is like we call this like the the demon or the worker. Basically, when when the when the client requests a new session, we start a watcher against that API.
42:08 and so here's your API server, right? the API requests a session. We start a watcher for like basically notify me when task files change when artifact when task like artifact files change. and we group these like per per like thing you're working on. And so if some this agent over here updates it, that change will flow into my local Oops. >> I'm getting you kind of have like a Dropbox like sync kind of mechanism here, right?
42:40 >> Yeah, exactly. So it's just constantly watching this stuff and then this flows into the file system. >> So the outer thing is actually writing to the file system and then when the har when the model wants to read a file the harness goes and touches the file system and this thing has been synced down. so this is how you can kind of like it's a it's built on an electric SQL sync. there is one weird race condition that we don't handle. but what we found is like you probably don't need CRDTS for this for most cases. Like CRDTS would be the way to do this like perfect like like all race conditions handled like no merging. but the scrappy way of doing this is just dropping things on the file system and then the agent can read them. And what's nice here is you don't need any extra MCP tools. so I didn't I didn't mean to tell turn this into a like here's how human layer does this, but basically like the the idea is like you should like make your files accessible to the agent whether it's a Google MCP or an Ocean MCP or just like having some process that just drops them on the file system for you or you or or or yeah and then sorry not or and then once they're available to the agent the agent can read and write them. You want to be able to comment on them. You want to be able to read them. You want to be able to read them from your phone. You want to be able to have them in a a place that is like stored. But the thing is like I don't think you need like you don't need Git for this. Like you need probably like linear history, right? And so you want to be able to say like, okay, what are the past versions of this and what changed? But >> you know what I really want? If I'm completely honest, I actually don't even want versioning or ranking. What I really just want is I just want Dropbox.
44:19 I literally just want to have Dropbox where I point it at a directory and I say this directory is an agent directory. Whatever happens here reflected on the cloud, multiple people might be talking to at the same time. Give me a cdt sync on it and then let me comment and make the comments available via the file system somehow. >> The comments are good. We also we also did give you version history so you can like look at like okay what changed recently and stuff. Yeah, I mean I think Dropbox has like similar kinds of like linear versioning, but like it's like I want this to just be a file system construct which is just like always working. I can just point to it anywhere and it just works. Like if I have to hook into a specific agent that means like when a new one comes out like Pi, I kind of get stuck on it. Matias has a good point about like a one drive folder. So that's kind of the idea. The problem is and this is what I run into.
45:06 It's mostly about edit resolution and how comments work. Edit resolution is really hard to go deal with in a typical CRT system because if you're building an agent-based system, you want to have delayed resolving of edits because if you edit it right now immediately as it does, you kind of want to have the user prompt and say, I actually do want this error to I actually do want this error to propagate up >> and like I'm sorry, I want this edit to immediately interrupt the agent and invalidate its cache or I want to hold off on this edit until after the agent is done working. There's a very important problem to solve there or else you get kind of get screwed.
45:44 >> Yeah, the cloud code harness had this for a little while. I don't know if they still have it, but they basically had like if an agent touches a file >> and then that file changes on disk that the agent hasn't seen, it will inject a system message to remind the agent. You don't see it, but you can see it in the thinking traces. Claude is like, "Oh, the system reminder tells me that the user has edited the file." so that I should and and so I should read it again before I try to edit it. And it's like, see what's different. Yeah.
46:11 >> yeah, we created a context repo that's ties a set of GitLab repos together, gives every engineer one place to understand how the repos fit, write a plan, brand. Yeah. So, this was the second generation of what we built is like we built a separate repo that we would hard link into every repo that had code in it and the agent would read and write there and then we had a hook that was basically like on commit it would also sync that thing. It would just do a straight pushpull. There was no branching in that other repo, but it was a thing that like when certain things happened, it would get synced up. That was even less like real-time sync. That was like on demand sync. But again, like it didn't really matter unless two people were working on the exact same document, which tended to be pretty rare. And the only reason we wanted to go like more real time sync was because we wanted to get to a point where like two people could be like collaborating on one doc with the agent at the same time. but yeah, Narin, that's that's the natural thing. And you will probably like if you keep working on this, you will you you might you may eventually end up with a thing that is more real time. because the craziest thing for me was always like the first time in December or something we were testing this and me and Kyle were sitting next to each other and I was having my agent edit the doc and push it and then he was having his agent edit the doc and push it and we were both staring at the page in GitHub real time when we were storing artifacts in GitHub. That was like the magic moment for me. That was when I was like, "Oh, this this needs to be a product."
47:34 >> Cool. this is cool. I think we should probably close it off here. And I saw there were a couple more questions around sandboxes and more about like process and workflow. >> Is there any Google Doc like system that lets agents and users dynamically update docs and spreadsheets in real time? Cloud keeps making a bunch of new Google Docs for each edit as the API doesn't let it make edits directly. Yeah. This is this is really hard and we should probably get Kyle on to talk about this, but basically like yeah, the the model tools are not good. What you need is something like code mode where the model can what we what we ended up doing is like the model wrote a like markdown edit and then a small model took that edit and turned it into like a snippet of quickjs JavaScript and sorry, give me one sec. Sorry, I'm back. Kyle heard me say his name and he thought I needed something.
48:25 Yeah, I need you to come do a great podcast episode because I'm I'm throwing it over. >> I actually do think the feature is code mode for almost all things like we should actually do an episode maybe we can do a renewed episode and like code mode sandbox and the whole like design patterns that get evoked when you enable the model to write code instead of calling tools cuz I think 100% code there. >> Reese is not allowed back on the podcast. He went and started a competing podcast. He's dead to me.
48:50 >> No, I think we can talk about code mode. I think code mode will be fun. I think there's a >> we've been advising more and more companies to actually do code modas inside of calling for many scenarios. >> It's how all the RLM works. >> Yeah. Anyway, it's 11:09. Andre, sorry we didn't get to your question. I It's for two reasons. One, I have no idea what ED is. I So maybe next time let us know what it is. We'll take a look at it or tweet at us or something.
49:17 >> It's the Employment Development Department of California. >> Eval driven dev. Oh, eval driven dev. We will put a pin in that. We can talk about that forever. >> yeah, go watch the emails. >> Let's do it. >> All right. Thank you everyone. We'll catch you guys everyone.
Summary
- Participants are constantly evolving their methods, with no two individuals operating the same way.
- A key takeaway is that if one doesn't look back and feel they could have done better two years ago, they aren't learning quickly enough.
- The unconference covered various topics, including cash engineering, harness engineering, and building evals.
- There is a consensus that AI-generated code often results in "slop code," necessitating human oversight.
- The conversation included the importance of adversarial review processes to improve code quality.
- Discussions on the use of metadata for AI-generated content and its implications for misinformation were also featured.
- The need for effective documentation and collaboration tools for engineering teams was highlighted, with suggestions for better integration of artifacts and plans within workflows.
- Continuous adaptation and learning are emphasized as essential for staying relevant in the rapidly changing tech environment.
Questions Answered
What was discussed at the recent unconference?
The unconference brought together diverse experts to explore various topics related to AI, emphasizing the rapid evolution of practices and the need for continuous learning.
What are the risks of using AI models without oversight?
Unattended AI models tend to produce poor quality code, leading to chaotic outcomes unless actively managed.
How can we effectively evaluate AI implementation?
By breaking down complex problems into smaller, manageable parts, we can validate and optimize the implementation process more effectively.
What are the risks associated with updating code in complex systems?
As systems grow in complexity, the likelihood of missing updates increases, which can lead to significant errors in the code.
What is the recommended approach for managing code artifacts?
It's advised to store artifacts outside of version control systems like Git, using alternative systems for better organization and accessibility.