Transcript
0:00 What are we going to do today? >> So today we're going to walk you through how to do eval step by step. The real live example on real data. Do we even need eval? I heard Claude code doesn't use eval. >> Oh my gosh, this is a crazy controversy that's been going around. >> There's way too much hype about AI to build a really good AI feature. It's not just a demo. You need to build something that goes to production. I consider AI evals the number one most important new skill for product managers. Where do people at Anthropic and OpenAI go to learn AI evals? It's HL Hussein and Shrea Shunker. This is what your AI agents are actually doing out there in production and that's why looking at the traces is so important.
0:40 >> We show so many demos in class or we just dump this trace into chatbt and we ask was the assistant correct and then chatbt will say yeah absolutely but it will miss all of this nuance. >> Do I need an AI observability tool? I already am paying for whatever data dog or whatever APM tool I already have. What is the difference? >> You don't necessarily need a tool. It's sometimes good to start with one, but if you want to, you can log to CSV file, JSON file, text file, whatever you feel comfortable with.
1:09 >> What are other mistakes like separating the prompt from the product manager that people might be doing in this process that we walk through today that is unintentionally inhibiting them? >> The main thing that's inhibiting people is not doing the error analysis. Before we get into today's episode, if you can do me a quick favor and check if you have a following on Apple and Spotify podcasts and subscribed on YouTube, these are free actions you can take that really help the show grow. And if you become an annual subscriber to my newsletter, did you know that you get access to over $28,000 of premium products? That's right.
1:47 Mobin, Arise, Relay.app, app, Dovetail, Linear, Magic Patterns, Deep Sky, Reforge Build, and Descript. They are all free for an entire year if you become an annual subscriber to my newsletter. So, go take advantage at bundle.ashg.com. And now into today's episode. HL Shrea, welcome to the podcast. >> Thank you for having us again. >> Yeah, super excited. >> What are we going to do today? So today we're going to take you step by step on how to do application specific evals.
2:21 >> Do we even need evals? I heard claude code doesn't use evals. >> Oh my gosh, this is a crazy controversy that's been going around. Absolutely. Everyone needs evals and some people are less rigorous about it because perhaps there's somebody else who's done evals for them upstream. For example, in coding agents, you know that people who are training the models are testing on a bunch of code. So maybe you can easily build a coding application based on you know rel religiously dog fooding your outputs but for most applications that are not just naive applications of foundation models such as what you're building you're going to need some form of evals.
2:58 >> Couldn't agree more. Evals it's about actually improving your product. Maybe you're doing that through dog fooding or maybe you're doing it through the systematic process that we're about to walk you through but you have to do them. So let's get started. So today I'm going to walk you through how to go about evals using a real company that I worked with, Nurture Boss. And Nurture Boss has been very generous and allowing me to use some of their anonymized data as a teaching example. So what is Nurture Boss? Nurture Boss is an is a tool that allows property managers who are managing apartment complexes deal with things like tenant interaction and marketing and sales. And so you can see their website here, nurtureboss.io.
3:44 And you can kind of get a feel for some of their, you know, it's a mobile, you can have a mobile app, you can embed on the website, but you can see here's a example interaction. Do you have any two bedrooms available? And the nurture boss application is interacting with the tenant for you, helping to show listings, schedule appointments, so on and so forth. um you know kind of all the different activities that you might be engaged in as a property manager the their application is helping you manage that with the assistance of AI.
4:18 So it's a really good example because it incorporates all the messiness of a realworld AI application. There is tool calls, there's rag, multi-turn conversations, there's even multiple channels you can interact with the application through voice, text message or chatbot. And so it's kind of a lot of different messiness of the real world. This is not a simplified example. This is something that you will encounter.
4:51 You know, in the real world, your application might have these complexities. So like how do you go about thinking about eval? So when I started working with nurture boss they had something initially that worked but they really wanted to know okay how do we number one figure out what's going wrong and number two like how do we improve the application systematically beyond just doing vibe checks because they already did vibe checks. you know they were using their own application they had some design partners and some initial customers but they wanted to move beyond that and make it really good. Okay so what did so the first thing that you want to start with is some kind of observability.
5:32 So the nurture boss application they instrumented their code and they captured their traces. So let me just show you what a trace looks like. So this is an observability platform called brain trust and it doesn't really matter what you use. There's a lot of popular ones out there. Ones that I see are things like Okay, so Brain Trust is one. There's Langmith is another one that's popular. Arise is another one. It really doesn't matter which one you use. Um, but the reason I'm showing you one is so that you get a feel for what traces might look like and also learn what a trace is. So, this is an example of a trace. Let me just make it big. And before we even get there, some people might be wondering, do I need an AI observability tool? I already am paying for whatever data dog or whatever APM tool I already have. What is the difference between those tools?
6:26 >> Yeah, you don't necessarily need a tool. It's sometimes good to start with one, but if you want to, you can log to CSV file, JSON file, text file, whatever you feel comfortable with. The reason I have it pulled up right here is so that we can read it together and kind of, you know, have something to look at. But if you're using data dog, feel free to log things to data dog to begin with. Um, the most important thing that you're going to want to have is to take notes on your traces, and we'll show you why in a second. Um, one of the things that Shrey and I teach is actually to bibe code your own trace viewer.
7:08 And we'll talk about that in a moment and I can show you note boss did vibe code their own trace viewer eventually. They didn't end up using this. Um but you know sometimes to get started it is handy to use an observability platform but sometimes it's not. Sometimes you know depending on what you want maybe you already have one you feel free to use that. The key is like do the simplest thing you can think of.
7:33 >> Get started. >> Yeah. Um, okay. So, here's a trace and you can see the trace just logs all of the different turns and all of the information that is shown to the LLM. And so, what you have here is a system prompt. You are an AI assistant working as a leasing team member at Harmony, which is the name of a fictitious apartment complex. Your primary role is to respond to text messages from both residents and perspectives uh both residents and prospective residents. So it's very interesting. Uh people will be text messaging this application.
8:12 >> Um and so you know you're engaged with the customer to answer questions, book tours, and drive applications. And there's a whole host of different rules um about how to respond to the customer. And what you'll see here is um you have some rules, okay? Like how to interact with them, determine if the inquiry is from a resident or prospective resident. I won't read all of these for you. There's some property specific information here like URLs.
8:46 This is anonymized. Obviously, the URL is not acme.com. Um but you can get the idea. And uh you know, this is basically the system prompt. So that's the system prompt. You see the first user messages. Um you can see that there's like some kind of logging error. Um it's like don't know who this is or what apartment you're from. So you could see like it's you know this is real world messy logging. But um the first question is I need a onebedroom with the bathroom not connected in floor plans. And then you can see there's a tool call of some kind where we're getting the individual's information. we're getting um something about the availability and the tool is returning a list of apartments here. You can see these are kind of like a list of apartments.
9:36 And then we have the assistant responding with, you know, these apartments here. And um, you know, it's saying, okay, we have like these three apartments with a link to this floor plans page, so on and so forth. Um, and then the, you know, it's a text message. So, this user is saying something about being sick and um you know, not being able to book tours.
10:08 Um and something about one a bathroom connected to the room. Um I'll check for onebedroom apartments. So, the bathroom is not connected. Thank you. And you're welcome. Okay. So, so there's a lot of things that kind of went wrong here. One is what is going on here with um I want a you know I want a bathroom connected to the room but it just said I'll check on that but then it didn't do anything right. Uh so we would hope that it would actually you know do something or if it's not able to do it hand off to a human. So you know that's that's kind of funny.
10:49 >> I do happen to know um that so this response up here is also in markdown. So you can see here um yeah this is a markdown response and this is a text message. So it's gonna be rendered a bit weird in a text message because text messages don't have markdown. And so um this is a bit problematic too with the bold and everything. It's going to it's going to come across uh you know potentially in a weird way with like asterisks and stuff.
11:20 >> Yeah. Yeah. It's going to have asterisks and it's going to have square brackets and stuff like that. If you've been enjoying this demo on how to do AI evals, you are going to love HML and Shrea's course. It is the top grossing course on Maven. It is taken by people from OpenAI to Enthropic to Google to Meta. All of the top AI companies are taking this course to improve their eval. I have secured a massive 35% discount for you guys. So use code AG-VL so that you can get this course. You can learn even more in detail how to write great evals, how to build great AI products that are working in production.
12:00 Check it out at maven.com and look for their course and type in code AG- V L. That means it will only cost you $2,275. The next version of their course starts on October 6th. Today's episode is brought to you by Vant. As a founder, you're moving fast toward product market fit, your next round, or your first big enterprise deal. But with AI accelerating how quickly startups build and ship, security expectations are higher earlier than ever. Getting security and compliance right can unlock growth or stall it if you wait too long.
12:32 With deep integrations and automated workflows built for fastmoving teams, Vanta gets you audit ready fast and keeps you secure with continuous monitoring as your models, infra, and customers evolve. Fast growing startups like Lingchain, Writer, and Curser trust Advant to build a scalable foundation from the start. So go to vanta.com/acosash. That's v a nta.com/ a kh to save $1,000 and join over 10,000 ambitious companies already scaling with.
13:05 One more thing is actually in the first message the user said they wanted the bathroom and bedroom disconnected. Yeah, see I need one bedroom with a bathroom not connected. And then the assistant's first message was, "Here are some bedrooms with bathrooms connected." >> Oh, yes. See, there you go. That's a that's a good observation. So, it actually didn't really help the user. Um, the user just kind of gently reminded them again like, "Hey, I do want a bathroom connected to the room."
13:34 I mean, this is like very messy. You could see like there's misspellings. I could almost not understand what the person was asking. >> Yeah. The the now should be a knot. I do not want a bathroom connected to the room. >> Yeah. So, this is awesome. This is guys, if you are PMs listening, this is what your AI agents are actually doing out there in production a lot of the time. And so, your demo is one thing. It goes well, but then when it goes out in production, there's all this hairiness.
14:00 And that's why looking at the traces is so important. >> Yes. And so, this is exactly why you don't want to have generic metrics. If you try to put helpfulness score, conciseness score, whatever in here, or you try to have AI look through your traces, it's not going to catch stuff like this very well at all because there's a lot of context that you have as a PM and a lot of things that kind of you have taste that you need to reflect on and say, "Hey, this is not a good experience from a product perspective."
14:34 And the language model is not going to do know that because it hasn't been able to read your mind. Mhm. Um, and so >> this happens all the time. We show so many demos in class where we just dump this trace into chatbt. I mean, you can probably even do it now or claude and we ask, "Was the assistant correct?" And then chatbt will say, "Yeah, absolutely." You know, sounds correct. But it will miss all of this nuance that Haml and I and Akos have been mentioning because you know we actually put our product hat on and thought about the user experience a little bit.
15:11 >> Yeah. Let's see what JPD does. >> Oh, look at that. >> So I found that one. Okay, so it figured out that we didn't get the connected bathroom, didn't use. >> But this is hilarious. It says it doesn't filter by bathroom configuration. And the interesting thing is, who knows if that's a filter that the tool provides.
15:43 >> Yeah. >> Assistantly cherrypicked three examples. I mean, maybe that's fine for us, right? Like nobody wants to s put all one bed. I don't want to see a text message of every single apartment. I actually only want to see a couple. Yep. So touchypt might help you a little bit, but you ultimately need to put your human touch on top of this and make sure that correct.
16:13 >> It won't it won't uh know that about the markdown, >> you know. So like I didn't catch that. Yeah, you can um you know I can change the format of this rendering I believe somehow to uh show you the raw. This is actually all rendered as markdown but you know since this is a text message but like you know JPD is not going to know that and there's other examples that we'll see that okay you need to put your you need to kind of have a keen eye about what's going on in the product and bring your whole product knowledge to bear. Um and so yeah, if you try to it it can it can miss a lot of things. Um and so oh yeah, this is like the raw this is another raw way of looking at it in like YAML form. Um but anyways like that's not that's not important. Um so what you could do from here is you need to write a note. So I'm going to go into review mode. Oops. Let me go back to Whoops. Let me go back to that trace.
17:14 Sorry, the wrong wrong hot key. Okay, let me find that trace again. Uh, let me see. Is it this one? Nope. If I try going back, maybe that will save me. No, that didn't save me. Uh, see, is this this one? >> This is also why we encourage people to build their own tools. >> Um, let me see. Is this the right one? Yeah, this is the right one. I think this is it. Yeah. >> Yeah. Okay. So, let me come here >> and Oh, there we go. Okay. Notes. So, I would put a note here. Um, so some issues here would be um, you know, told user that it would check on bathrooms but didn't do it.
18:11 Um say like also did not uh follow user instructions and uh rendered markdown in a text message. And so what you really want to do is this can sound very tedious like what I just did. It sounds like it's very resource and time inensive, but it's really not. Like you just scan the trace and you know if you're familiar with a system prompt, you don't have to read it. Um you're not going to read every system prompt because it's going to be the same really unless you need to. Um, but it's really you can, you know, within about 30 seconds or so or less, you can kind of scan this and say, hm, okay, you can get a sense of like what is happening and you can write some notes. It's the perfection is not key.
19:13 The key is like see what's going wrong in the trace and note what you see and and move on. Um, you don't have to catch everything. Just catch the most important things. And so we can keep going. So let's go on to the next trace and let's see. Let me just You have to edit this out. Uh let me find one with an issue. Okay, here we go. Here's one with the issue. So edit that navigation out.
19:42 Um so we have uh let me just hide some stuff here. So the user in this case is asking this is this is a new trace now. Do you all have one bedroom with study available? I saw it on the virtual tours. And there's a tool call to get information um and availability. And it says yes, hi Priya. We currently have several one-bedroom apartments available but none specifically listed with a study.
20:18 Um so okay, it matches. She did ask for a study. Um, so I gave her one bed one bedrooms instead. And then the user asked, "Can you let me know when one with a study is available?" And the assistant says, "I currently don't have specific information on the available of a onebedroom." So, okay. Is this kind of This is where you get frustrated, right? I asked the question, you just responded with some robotic like I don't have that. Um, and so I would say, yeah, this is a this is a product failure. Um, and what you want to do is say like, okay, um, you want to just note that real quick. So, um, should have handed off to a human or had have better lead nurturing.
21:15 I'm not there's no pun intended with the word nurturing. I'm just that's what came to mind. So, um you know, anything else that you think is wrong with this particular trace? >> Yeah. I don't think you have to get bogged down. It's a good question to ask, but typically would tell people, all right, like think of everything that comes to mind, break them down, move on, right? You want to like get into kind of a flow state here. Like you can debate every trace endlessly and sometimes you see people get stuck in that. So try to just kind of avoid that. We got a problem. Next.
21:50 >> Yeah, I agree with that. Move on. Um, this one I already did, but that's okay. We can we can we can do it again. Um, so let's find the first user message. Let me scroll up. Sorry. Um, okay. So, I'm in California. So, okay, we have to edit the part out where I scroll to this part, but we can start. So, this is a new this is the third trace we're looking at.
22:23 Uh, so the user asks, "I'm in California, looking to relocate to Texas by March 15th. Booven, thanks for sharing. Since you're planning to relocate, blah blah blah. I can help you explore available apartments if you'd like. We can also schedule a virtual tour." Um, he's like, "Yeah, that's great. you know, thank you and okay, I'll arrange a virtual tour for you so you can explore community. What's your preferred date and time? Tomorrow is fine for me. 9:00 a.m. I can schedule your tour for you. That's fine. It schedules a tour.
22:59 And then it says, "Your virtual tour is all set." Looks good, right? Actually, it didn't go so well. And so, the reason it didn't go so well is because there's no such thing as a virtual tour for this apartment. Um, and I don't think there's a such thing as virtual tools for most of their apartments that they have on the platform. And so, you know, there's so the AI called the tour this tour scheduling tool, but it's not doesn't do virtual tours. So, you know, the the platform scheduled in-person tour and maybe the user is confused and it's like, oh great, it's like going to be a virtual tour and this it's kind of a little bit of a disaster. Um and so okay that's fine. So that that problem is there. Um and then you know the date they said January 22nd 2025 the user said wait today is the 22nd. Do you mean tomorrow because like we can't do it today. It's like oh no problem. Um your virtual tour has been you know rescheduled for tomorrow. Okay. But if you look the schedule a tour uh tool was called again.
24:14 So what this means is hey like you just scheduled another tour you didn't reschedu anything. So now now the person has two tours. So and and you know maybe you know that's going to cause a problem for the apartment complex because now they have two tours booked that which are going to be no shows. Um and then you know there's some other questions. how can I go about getting a unit if I'm in California trying to relocate blah blah blah and it's giving some responses here that doesn't that seems okay so you know I have written down the two issues here we don't do rescheduling and we don't do virtual tours only inerson tour so I just wrote some notes so the the idea is like you just keep doing this and so this actually you can do this quite fast you do this for let's say 100 or so traces just write down what you see. Don't try to get into root cause analysis. Don't try to figure out like what went wrong exactly. Just journal, observe freely. Just journal what you see going wrong. If anything, if it's nothing is wrong, you can just skip it. Uh but when you do find something wrong, go ahead and write that down.
25:23 >> Mhm. >> Does that seem Is that clear, Akos? That process. >> Yep. Exactly. >> Okay, great. So, so now you have what we call a bunch of open codes and this is not a so okay this is the start of the most important part of eval which is called error analysis and something that's very approachable to everyone and it's actually it's very important for product managers to be involved in this because a lot of times engineers don't have the context the full context to know if this is good or bad.
25:57 Um, and so what you end up having is you have, let's say, a bunch of these like notes. So I have a spreadsheet open right now with a collection of all the notes that I took. And you know, let's say I did 100 of these. I actually found 40 or so different errors. Um, you might find different number of errors if it's if you're doing it. Um, but here's a collection of these like notes. Okay, great. So, up until this step, you've already learned quite a lot. If you've if you look at 100 traces, you're going to learn and you're going to understand your system better than anyone else. And you're going to have a really like deep understanding of what is wrong. And you might also have a pretty good sense of what you need to work on next already without doing any analysis.
26:51 But it's really good to do analysis of these notes that you took. So how do you do this analysis? So the next step is you categorize these notes. So the term for that is called axial coding. And I'm going to show you how to do this in a spreadsheet. So let me just zoom out a little bit. Um but first one thing you can do is okay you can take these and you can yeah put them into chat GPT or cloud. So what I did is I took um sort of the logs and I said okay please analyze I exported it from here. So uh you know there's let me go back.
27:37 So there's an export button here and I said okay download SCSV downloaded it put it in claude and I said hey there's a metadata field which has an S field called zenote that contains all the open codes and I use the word open codes that's a term of art that LLMs understand because this uh these terms open coding and axial coding which you mentioned open coding is the writing of the notes that's actually a term that's well understood in the field of machine learning but also goes it's been around before machine learning it's been used in the social sciences this kind of process of open coding axial coding is a thing that LMS understand and so um I just say there's a metadata field which has a nested field called znote that contains open codes for analysis of LM logs which are uh that we are conducting please extract all the different open codes and then propose five to six categories that we can create axial codes from. Okay. And then like you know it'll it'll kind of go through and you can like get these categories. So here are some categories like you know capability limitations me representation. Some of these I don't like cuz they're a little bit too broad.
28:56 They're not actionable. I actually not 100% sure what that means. So I might look into it and rename it a bit. Um you know human handoff issues. There's certainly some of that. That's when, hey, you want to escalate to a human being or hand off to human being, but it's not doing it properly. Um, temporal contextual awareness doesn't know what the current date and time is. Um, you know, so there's some categories here. Um, you can refine this. What I like to do is take it to a spreadsheet. So I like have some categories that I kind of have that I maybe have from chat GBT and then I kind of look at them and edit them. So I kind of have these like categories that I sort of edited a bit and I said okay let me just collect these into a list. So that's what this formula does is I just collecting these list of categories into a list. That's all that is.
29:53 >> I have a note here. sometimes you so I think one of the things that's very interesting from looking at your claude is a lot of those axial codes are very vague right like quality or temporal issues and you kind of want to make sure your actual axial codes are not so vague because imagine you're giving them to somebody else to do some labeling with right like something like conversational flow issues might be a little bit better honestly we could even make it a little bit more specific But something like temporal issues, right? Like if Haml told me to go label with temporal issues category, I wouldn't even know what he means. I would want to say like, you know, date formatting error or like something like that, right? So I think that's another place where people get tripped up, which is, you know, just taking things out of the LLM as is and not really thinking about, okay, how do I refine that in a way that's going to give me meaningful error categories.
30:53 Today's episode is brought to you by Jira product discovery. If you're like most product managers, you're probably in Jira tracking tickets and managing the backlog. But what about everything that happens before delivery? Jira product discovery helps you move your discovery, prioritization, and even road mapping work out of spreadsheets and into a purpose-built tool designed for product teams. Capture insights, prioritize what matters, and create road maps you can easily tailor for any audience. And because it's built to work with Jira, everything stays connected from idea to delivery. Used by product teams at Canva, Deliveroo, and even The Economist. Check out why and try it for free today at atlassian.com/roduct-discovery.
31:38 That's a t l a ss i an.com/product-discovery. Jurro product discovery. Build the right thing. Today's episode is also brought to you by my cohort-based coaching program to help you land your dream PM job. I am taking 30 elite PMs to land their jobs at Google, OpenAI, and other $700,000 plus roles. If you want in, check out landpob.com. Once all 30 seats are sold out, that's it. And already seats are going almost every day. So grab yours at landpob.com.
32:14 Today's podcast is brought to you by Pendo, the leading software experience management platform. McKenzie found that 78% of companies are using Genai, but just as many have reported no bottom line improvements. So how do you know if your AI agents are actually working? Are they giving users the wrong answers, creating more work instead of less, improving retention, or hurting it? When your software data and AI data are disconnected, you can answer these questions. But when you bring all your usage data together in one place, you can see what users do before, during, and after they use AI, showing you when agents work, how they help you grow, and when to prioritize on your roadmap.
32:50 Pendo Agent Analytics is the only solution built to do this for product teams. Start measuring your AI's performance with agent analytics at pendo.io/acos. That's pendo.io/ io/ a kh. >> Exactly right. And that's a really important thing to pay attention to. That's why if you reflect on the categories I have in this spreadsheet, they're definitely better than the ones in the claude. >> Yeah, they're different for a reason. >> Um it's because I I iterated on in a bit and that's that's an important principle. You never want to completely hand off the wheel to AI. You want to think about what it's saying. Maybe it helps you to different degrees but you want to see okay like what are the categories here um and actually you might want to go back and forth. So what I did here is I went to those notes which I have here every row is a different note and you know you can use AI. So I used AI categorize the following note into one of the following categories.
33:49 Okay. And what I did is I you know had AI and this is a this is a formula in a spreadsheet. So you can see the prompt. Um and basically like classify each of these notes into one of those categories. And I went back and forth like hm this category actually is not the greatest for this particular note. And I like went and edited the note. I went back to this category field, maybe like added one, deleted one, and like kind of fiddled with it till I was like reasonably happy like, okay, this is like a good set is good enough. Um, one thing I should have added here is like none of the above. Um, which would have been better, but you know, I'm showing you the simple stupid version, which is like get started. I don't want to over complicate it, but that that's what you should do.
34:36 >> And none of the above is is mainly like a means to the end. the end is really having these categories, but sometimes you might have like missed a category. So if you put none of the above here and then your AI does a classification and tells you none of the above, then you can go read those traces again and wonder maybe there's another category that I should add um so I can classify those. >> So when you get so you have these classifications and okay, let me just zoom out so we can see it together. Sorry about that.
35:09 Um, so you have all these classifications and now comes like the powerful part where you will put on well you will have like real superpowers if you do this as a PM that you will go above and beyond and kind of be armed with information that a lot of people usually not armed with and it's counting. So now you can count these issues and you know you can just use a pivot table. That's what I did here is say okay like how many times did I see this?
35:40 So now you have taken a world where it's kind of messy and like you don't really know like what is you know you might not know what is going on. You know that there's some errors and you're and you have this kind of par paralysis of like what do I work on? What do I fix next? What's the most burning problems in my app? And you know you have some data in front of you. you know like hey you're having these conversational flow issues a lot and this conversational flow issue it's actually regarding situations where there's text messages so I happen to know that I can click on this um you know you can say like hey yeah it's like disperate messaging sent inerson tour link um okay there's there's different sometimes it's about text messages sometimes it's just like it's not being here and we can go back to the trace and look at that. It's one of my f favorite things about pivot tables. You can like double click. Um >> you can also make it hierarchical. I like to do this too. So sometimes I like to say I like to break down conversational flow into like three different category subcategories.
36:49 Maybe some will be like repeated messages and some things will be um you know the AI just should have handled this one particular thing better. I don't know. Um so you should I think this is kind of where the subjectivity and your product experience shines right it's like you have to do this process in a way that enables you to make your product better based on the capabilities of your product or what your team can do. So, you know, if you can't have virtual tours, then you can't have virtual tours, that has to somehow be encoded in your system, right? So, >> yeah, definitely. Um, and so, yeah, this could be made better. I think that's, you know, I didn't try to, let's say, make it perfect, but as Freya points out, you can have, you know, subcategories which can help you kind of refine what's going on more. Um, but you can take a look at this. So you can say okay like what do I think is like most important like hey okay maybe the human handoff issue is not happening as much as the conversational flow issue but let's say you feel as a product manager that is a catastrophic error and that the the magnitude of that problem you know the the sort of the impact of that problem is so high that I'm going to prioritize that as number one but you you have some data to back up that this is happening and you have an idea of what's happening and now you have a reason to potentially write emails. Now you're not writing emails in the dark. Now you can write emails in response to actual problems that you are seeing instead of like hallucination score or some AI generated something or the other. Um you know you can motivate this in in thing that you want to fix. Now you don't have to write an eval about every for everything.
38:47 There might be some of these things that might be easy to fix. Like for example, there's this formatting error with output. Um, you know, an example might be using markdown in text messages. You might be able to just fix that. Um, maybe maybe there's no instruction in the prompt at all. Um, and it depends like what kind of eval you need to write. So there's two kinds of evals. One is codebased where you can test something without calling an LLM. So the formatting error without with output you might be able to use a code- based eval for that like hey is the format you do I see markdown elements in this output where there shouldn't be a markdown in which case maybe you should write the eval because it's not going to be expensive. Um, whereas with the LLM as a judge, something more subjective like, hey, you should be handing off to a human, you might need an LLM for that. You might not be able to write a an assertion in code. That's a little bit more expensive of an email. And you have to do it have a judgment call like, okay, is that something that's trivial to fix? It's like I didn't have that I didn't have like uh you know something in my prompt that you know had this instruction. Maybe you you are found like some dumb mistake that you made. Go ahead and fix it. You don't have to get caught up in evals.
40:14 What you want to do is write an eval for something that you think you might want to iterate against. I don't know if there's a better way to say that. I feel like that could have been straight. You think There's a better way to >> No. Okay. >> I think it makes a lot of sense, right? Like so already as a PM like this is a secret sauce for your product. If you don't do this product, you can't p you if you don't do this process, you can't kind of put your own taste or your company's taste into your product. Once you get to this point, that's kind of when paths diverge. I think that's what Hamill is trying to say. Maybe sometimes you figure out, okay, there's some errors that are higher priority than others. I'm going to go and fix those.
40:54 Maybe you want to run these checks at larger scale like maybe you're Meta or you're Google and you're like I can't make a decision based on 42 traces and I need a lot of buy in so I'm just going to you know build a team do a bunch of evals or I'm going to write automated evaluators to check this at scale. Yeah, go for it. We're not going to cover that kind of today. That take our course if you're interested in those um techniques but I think overall I think it's super incredible. Ham started with zero. Right now we're at a place where we know what are the biggest failure modes in a sample of traces are. Right? And most people don't get to this point.
41:31 >> Um and so okay so going further from there um let's say you want to write an eval for the human handoff issue like hey you should be escalating to human being you should be handing off. um it can be really useful to write an eval to help you see all the traces where that flag the traces where that might be happening and help you iterate on that problem. So let's go into how you would go about building that. So here is the prompt for an LLM judge. Now this is just a very basic prompt. It could be made better but I just want to keep it simple. And so um you know you are scoring a leasing assistant um to determine if there's a handoff failure return only true or false. So that's one thing that we teach is you want the LM as a judge to produce a binary score.
42:29 >> Sh you want to talk about why? >> Yeah. So all right I'll try to give us a sick answer for this. The short answer is that people run into a lot of misalignment when they try to use like a liker or a rangeface score. And that's because it's very tedious to check that every single possible LLM output aligns with your preference. Now, when you do a binary score, you only have to check that true align with your TRs and false aligns with your falses, right? It's only two things that you have to check, which makes the process of checking for alignment easier. The other thing is when you're shipping products, right, you make binary decisions. Either this thing was bad or this thing was good. I should fix this or I should not fix this, right? It's not like even if you have a score of like this is 30% failing that gets turned into a binary decision of how you're going are you going to act on it or not, right? So that's kind of why we tell people focus on a binary decision here. it's easier to align and ultimately your business decisions are yes or no decisions.
43:33 >> Yeah, there were some people who are trying to do like one through five scales and stuff, but it seems like LM are not very good at those types of numbered skills. So, it's much better to stay binary. >> Definitely, we need to bring you into our consulting off if you can tell people that. Um, so, okay. Okay, so we have this prompt and you know we have like a list of things that we have you know these seven things that we have where there is a failure you know if there's explicit human handoff um that's request that's that sorry so we have these seven things um an example of one is the user asks to be sent to a human but that's ignored or the there's a policy that you should be transferred but that's not handled properly. There's a sensitive issue um like billing disputes or legal issues that are not adhered to. Um same day walk-in or tour requests, you want to hand that off to a human. So things like that. And then um you know we have some notes about when there's not a handoff failure.
44:50 It's important not to get bogged down in the prompt itself. So, if you're thinking to yourself, can you send me this prompt? Can I copy and paste this prompt? You're asking the wrong question because you need to, you know, so, okay, how do you write this prompt? Um, you want to try to describe the rubric of okay, what is a failure and what is not a failure. You can get LLM to help you bootstrap that, but you should try to edit it. And the key is iterating.
45:26 Okay? It's not necessarily a recipe. And you want to try to have examples. I didn't put examples in here. So, you want to have a section of like maybe some examples. Um, it's not ne necessary to begin with. In a simple case, you may not, you know, you don't need to have examples. I'm just, you know, trying to give you like the most dumbest LLM messages so you can get the concept. The idea is like, you know, you're going to write a prompt. Um, in our class we do have a recipe of like a what you can follow, but you know zooming out from that it's important to just iterate honestly. Um, that's that's what's going to get you the furthest. And so um you know this prompt would be structured differently if it wasn't in a spreadsheet also. Like I'm kind of begging it to return true or false. Um I wouldn't have to beg it if I using an API for example. I could I could do something else. So um this this is the prompt and then what you can do is okay I have my trace here. This is a different tab of the spreadsheet. I'm saying okay um you know this is the trace and I have AI um here and it says assess this LM trace according to these rules and I give it the prompt of my LLM judge. That's all I'm doing.
46:44 >> Nice. Now, here's the neat part. >> AI function built into Google Sheets is that just runs Gemini in the back end. >> Yeah. Yeah. It's okay. I wouldn't say it's amazing. It's some kind of very fast modelish. Um, you know, it's good to get started to get a mental model of what's going on, but I would, you know, I would be a little bit careful using this model for everything. in real life because I'm not too sure about it. But so don't get lost in the sauce of what I'm doing. I'm trying to show the exact I'm trying to give you a mental model, but you might not actually you might want to do something uh you might want to use like a more powerful model uh potentially for LMS to judge.
47:33 >> So okay, we have two columns here. Column G and column H. So the column H is the score outputed by the LM judge. True or false? And most people just stop here. They're like, "Okay, here's my LM judge. I gave it a prompt. Woohoo. We're done." LM judge says like, "It's good. So, we're good, right?" And then what ends up happening is stakeholders they start to feel or observe that there is a dissonance between the out like your evals and actual that the product's performance and they can lose trust in the evals and they start to ask you questions like how do you know this metric like how what is this metric and you tell them okay it's an LM judge like well how do you trust that? A lot of people get stumped there. like, uh, well, that's all we got.
48:26 You don't want to do that. What you want to do is you want to measure the judge against your label. So, remember when we're doing the axial coding, you actually have your own human labels. So, you know, for these various um traces like okay, if they were if this issue existed or not and you can so then you can compute metrics. you can compute how good your LM as a judge is.
48:57 Now in this spreadsheet I have three metrics agreement, TPR and TNR. Now agreement is like the trap metric. The reason it's the trap metric is that's what you might gravitate towards. You know in the naive case you might say okay like we just measure the agreement between the judge and the human. You don't want to do that. The reason is is if you're if this failure is only happening let's say 10% of the time, you can have the dumbest judge in the world have 90% accuracy by just always predicting pass.
49:41 In fact, your LM judge can just be like equals pass. You can introduce a bug that doesn't even call an LLM. It'll be 90% accurate. So, you don't want to do that. That could be very misleading. It can mislead you. So, what you want to do is measure two things. How good is your judge at catching errors? Uh, you know, catching errors that exist and how good is your judge at Sorry, let me rephrase that. I always get Let Ta explain this one so I can give myself a break.
50:14 >> Sure. Sure. Oh, man. I mean, I think you've like said basically most of it. The point is, okay, if you're a product manager and somebody tells you you have high agreement with your judge or they got high agreement with the judge, be a little bit suspicious. Ask them, okay, what's the alignment in the positives or the passes trus? What's the alignment in the falses? And make sure both of them are pretty high individually. If not, then you have to rework your LLM judge prompt.
50:47 >> And if you're confused about this, what why why isn't like if you're not convinced intellectually somehow that like why can I just use agreement HL? Why do I need to like measure positives and negatives separately? You should use the spreadsheet. You should get a spreadsheet like this and you should like do some experiments and say, "Oh, okay." Like what if I just, you know, hardcoded this to false all the time? I think this confusion matrix maybe not necessary. Um, it might confuse people.
51:23 So >> it's going to confuse people. Always does. >> Okay. >> We can't teach the whole course in in a one and a half hour thing. I think we just cut our losses. >> Yeah. Yeah. Okay. So, um, you know, that's kind of there's a lot of things that we didn't cover here. One key thing is how to split your data set so you're not overfitting your judge and you're not inadvertently cheating. That's a that's way too much to get into in a 1hour podcast. There's no way we cover that. But just know that okay, there's a lot of nuance here. How you do this correctly, how you build the judge, how you get confidence. Um there's ways to calculate your metrics. um use this TPR, TNR to like calculate what you know your real accuracy is. That's you know we haven't gone into that. Um there's a lot of nuance on like okay how do you analyze agents like how do you you know if you have lots of steps and lots and lots of handoffs how do you tame that and how do you do analysis of that to catch those errors. Um, another thing we teach is how do you do analysis of retrieval? So, retrieval is like kind of an Achilles heel of a lot of AI systems. And so, a lot of times you have to kind of dive deep and diagnose what's going on with your retrieval step in your rack. And so, there's a host of metrics and analysis you might want to do there. So, there's a lot of things that we didn't cover here, but the reason that we gave you a taste of error analysis because error analysis is the step that most people skip in Evals and it's rarely talked about and it's the thing that's going to give you extreme leverage uh as a PM and you can get there just with counting. I hope that I've convinced you by using this the spreadsheet that is within your reach and you know I don't want to discourage you from using spreadsheets either like feel free to use whatever you're comfortable with sorry go ahead >> so where do you go from here once you have this initial set of metrics how do you go about improving once you have created your initial evals >> right so let's say like this handoff error Eval that we created. Uh what you can do is you can now you have a judge an LLM judge that you like. You feel good enough about this accurate enough. You can use it to score a large sample of all your production traces.
54:01 And now you can find you can learn first of all you can learn more about what is going wrong in those situations. But secondly you can iterate on this problem really fast. You can make some changes to your prompt. And you can calculate your error rate on these test cases that you curated and kind of iterate really fast and say okay like this prompt is working this prompt is not working and you will have a suite of these evals and you can test against all of them to see okay like if you are iterating on this problem are you inadvertently breaking something else and you have some kind of way to a system that you can use to sort of be confident in what you're doing rather than just guessing.
54:48 >> What does a holistic like endstate eval suite look like? >> Sh you want to talk about that? Like how many evals do you usually have in your >> Yeah, it's it's different for every application and it's different for how high stakes the application is. Typically I'll see like you know several codebased evals especially in CI maybe one or two LLM based evals in CI but not really. I do see some people like myself included run LLM powered evaluations like kind of like monitoring or online like every week or so I'll sample some of my traces run my LLM powered evaluators on them and then kind of just look at the score see if anything's off or whatnot. Um, and often I'll see every few weeks that like, oh, there's this new distribution of data or this new cohort of people who are using the tool.
55:43 Like I build AI powered data processing tools. Um, and so I'll see like, oh, there's different document types that have come in or a different set of contracts. I do this for law a lot. So like there's a new type of contract or a new type of document that's come up. Um, and then now I need to like think about it, right? So LLM powered evals, automated evals allow me to really quickly iterate on those. You don't need a hundred of them like just a few is fine.
56:08 >> And what's the role of PM and AI engineer and AI researcher and all of this? How are you working together? Where are the handoffs happening? Like that quick iteration on the system prompt, who's doing that? >> It's a really good question. You know, it depends on the size of the team and the company and the roles. Sometimes these roles are being collapsed into one. um in terms of okay so Jacob Carter the engine the CEO and you know also engineer of this product is the product manager and the AI engineer um allin one so you know he is has a pretty good pulse on like hey is this interaction good or not so that's not feasible all the time um you know in other situations you want the domain expert to be driving especially the error analysis process as much as possible. It might take some training in the beginning to get the right tools surfaced for the PM or get the PM you know able to access the data and might need some engineering like kind of co- what's the right word um pair pair programming in a way or pair pairing so so pairing on error analysis just to feel comfortable but you should try to have one person do the analys so it doesn't become ownorous and usually a product manager is pretty good at that because they have all of the domain expertise to actually judge if something is going wrong. So I would I would bias on the side of having the domain expert or the product manager do this error analysis. And then as far as like writing the prompt is concerned, you do want to try to make it accessible for the project manager to write the prompt.
58:04 So what I've seen in a lot of tools is having a like an admin view where a non-technical person can can edit the prompt. Um I actually have like screenshot of that here. One thing that we talk about a lot in our course is you want to create your own tools to look at your traces. And so NurtureBoss actually created their vibe coded their own tool um to look at traces to remove all the friction of looking at traces because it's so important. And so you know this is pretty simple. You see all the different um channels, voice, email, text, chatbot. You can see like they hide the system prompt by default.
58:47 um you know it's like a very quick and dirty interface on doing this like open coding and axial coding and actually they have a step here that helps them automate the axial coding you see like hey transfer handoff issues tour scheduling blah blah blah so u you know that's worth noting that's something to think about is that's how important error analysis is um so to get back to your question about Okay, how do you how how might you surface the prompt to nontechnical people? So, this is an example where you might have an admin view. So, this is like um you know a real a different real estate agent like hey um you know showing you real estate listings and you might have this admin mode where then you allow someone to fiddle with the prompt and so like this prompt experimentation is really key. Um and so having a way that people can interact with prompts is really helpful. Now, a lot of tools have like prompt playgrounds.
59:59 The only thing that's limiting about most prompt playgrounds is they don't have access to your code, you know, because you might have various tool calls, you might have rag, you might have these things. All your application code is not, you know, in those prompt playgrounds. And so that's why a lot of teams that I see have these like these interfaces where you can like edit the prompt directly in your tool and then like play with it and redo it and whatever. Um so >> this is so nice.
60:30 >> Yeah. So I mean whatever I you know whenever possible you want to expose the prompt to the domain expert because the reason is is because it's English. It's made for the domain expert. It's almost a tragedy to separate the prompt from the product manager because it's it's English. What are other mistakes like separating the prompt from the product manager that people might be doing in this process that we walk through today that is unintentionally inhibiting them?
61:03 The main thing that's inhibiting people is not doing the error analysis. people want to jump straight to, hey, let me let me take a off-the-shelf metric that vendor gives you and and just like create a score that people are very scared of this error analysis. They look at it and they're like, "Oh, I don't have time for this." But it doesn't take that long at all. And it's kind of this thing. It's like a secret club. When you do it just once, you will forever keep doing it. But just like getting over that hump of doing it the first time is just extremely scary for people.
61:42 >> Another common mistake is people people will see this video or I don't know they'll realize okay it's worth doing error analysis but then they think it's worth some human doing error analysis not them. So they'll just outsource it out which again huge pitfall right this is the error analysis is where you build your product right that's where you build your moat. So if you're giving it to someone else then you kind of have no personal touch in your product.
62:11 >> Yeah. Do not outsource this to developers. Um in in you know if you're working on a coding app yes you can there's not really you're the domain expert is the developer but in most cases the domain expert is not the developer. And a lot of people a lot of companies are like oh this AI stuff is like for engineering. Like the whole thing is engineering. let me just shove it over there. They need to figure out whether it's like good or not.
62:38 That's usually the wrong approach. Um >> it's not in engineering skill set. I think that's another interesting thing about today's day and age for PM, especially AI PM, right? You can't expect engineers to be able to do all of these things. uh the people that have been successful at this process either are very very technical PMs or are engineers who are actually PMs and they just think that they're engineers and they they don't realize that they're doing product work, right? So I hope people are kind of convinced that today's day and age you kind of have to have your product knowledge, put your product head on.
63:20 This error analysis is so powerful that um this is a video and we can put it in the show notes of this is Jacob Carter and he just we recorded like a two minute long conversation of how thrilled he was with error analysis. He's actually so thrilled with it that he just he's he thought this is the best thing that's ever happened. Um and he got so much value out of it that yeah like he he kind of stopped there um to begin with and like had so much work to do that he didn't need you know to build evals right away cuz he just found so many things as you can see right here in this picture um that you know he was able to fix. So he did eventually build evals, but you know, starting here gives you a really good grounding and lets you work through issues and like get to evals for things that make sense.
64:17 >> Going back to our beginning, there's so much hype about what you're trying to sell that your AI feature does, but to actually deliver on that hype, you have to go through these errors so that when people are experiencing it in production, they actually get the experience you intended. And this has been our master class in how to do that step by step. If people want to learn more, where can they find you guys? >> So the URL for the course, you can go to eval.info and you can find the course there.
64:48 >> Yeah, or follow us on X. Our websites are on the internet. You know, if you just look us up, we're there. Um, but check out evals.info. I think we've really tried to put together as much information that we can to be freely accessible and available to folks. Um so take a look right you can dive in and I'm sure you will learn things along the way. >> Awesome. Thank you guys so much for the >> I might want to clarify that sorry is um so we mentioned like hey you need to look at traces in production. So you might be wondering like what if your application is not in production what do you do? What if you don't have data?
65:22 What if you don't have traces? Where do you get them? One, try to recruit some friends. Make try to dog food your own app. That's the best thing. For whatever reason, if you're not able to recruit friends, you're not able to dog food your product, which would be kind of sad. But let's say, you know, there could be valid reasons you're not able to do that, you can generate synthetic inputs into your system. And there's a there's a way to do that correctly is kind of what you're doing is you're pretending to be a user.
65:50 You're having an LLM simulate that at scale. Um, that's one of the things that we go into in our course as well. So, so there's ways to bootstrap yourself, but you do need to look at data. >> Amazing. Thank you guys so much for being here. >> Cool. Thanks for having us. >> I hope you enjoyed that episode. If you could take a moment to double check that you have followed on Apple and Spotify podcasts, subscribed on YouTube, left a rating or review on Apple or Spotify, and commented on YouTube. All these things will help the algorithm distribute the show to more and more people. As we distribute the show to more people, we can grow the show, improve the quality of the content and the production to get you better insights to stay ahead in your career.
66:30 Finally, do check out my bundle at bundle.acg.com to get access to nine AI products for an entire year for free. This includes Dovetail, Mobin, Linear, Reforge, Build, Descript, and many other amazing tools that will help you as an AI product manager or builder succeed. I'll see you in the next episode.
Summary
- AI evals are crucial for improving product performance and should not be overlooked.
- Error analysis is a key step that many product managers skip, but it provides valuable insights into system performance.
- Observability tools can help track AI interactions, but simple logging methods (like CSV or JSON) can also be effective.
- The process involves reviewing traces of AI interactions, noting errors, and categorizing them for further analysis.
- Binary scoring for evals is recommended over more complex scoring systems to simplify decision-making.
- Collaboration between product managers and engineers is essential for effective error analysis and prompt iteration.
- Continuous monitoring and adjustment of AI prompts can enhance system performance and user experience.
- The podcast encourages product managers to engage directly in error analysis rather than outsourcing it, as it is vital for maintaining a product's quality and relevance.