transcribe

Redesign Product Metrics for AI Evals

AI Analyst Lab · 20m · transcribed Jun 2026
More from AI Analyst Lab Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:00 Hey everyone, welcome to redesigning product metrics for AIE bells. It's more supplementing product metrics for AIE bells, but I feel like uh chat GPT gave me redesigning as the better hook. So, you just have to do some of this stuff for marketing. But I'm Sean. A little bit about me. I'm a principal data scientist at Entra. We're a legal tech AI company. I've been there for a bit over a year, solely focused on AI evaluations for our data science team.

0:27 been working in data science for a little over a decade across both B2B SAS and consumer product. So you can see some of the companies I worked at up here. Previously, actually before grad school, I was in education a bit. So I was a teacher for a couple years, which is why I like doing lessons like these and also doing them at my day job. One other kind of fun plug, I'm a co-host of a podcast called Data Neighbor. We started about a year ago. We talk all about data as it pertains to AI especially. Um, and we have AI tech company leaders. Some exciting stuff coming in 2026. I think that's relevant for this group is we have a new series we're going to be launching. I'll talk more about that, but it's called measuring the machine and it's solely focused on uh AI evaluation. So, we'll do episodes on that for for a couple months actually. So today what we're going to talk about is going to be a little more of a highle framework. How I think about approaching metrics as I pertain to AI products when I'm working with a prod team. We're not going to be doing like any coding or really like be looking at any data or getting into like the nitty-gritty stuff today. Uh I do have a six week kind of applied hands-on course that I'm kicking off in April where we'll get way more deep. So we'll talk about the foundations which this is a mini preview to. We'll talk a little bit around instrumentation and observability. Uh we'll talk about more importantly, you know, how do you interpret success and failure of AI outputs? Uh and actually more more importantly, how do you connect them back to the to to the actual user value and and the business as well as the whole flow once you've created this metric? How do you operationalize it?

2:14 How do you set up different forms of operationalizing it where you have safe layers within your development environment as well as in production? So who should join the assumption coming in here is you're someone who's working on an AI product. So product managers who are shipping AI powered features, data scientists, analysts who are measuring it and engineers and MLES who are building and maintaining those systems either currently or sometime close to the future. And then last plug on here and we'll get to the good stuff and out of my ad I have squeezed into this presentation is that we have discount code here. AI metrics 30 for those of you joining the session. Okay, the ad's done. Let's get into uh product metrics.

2:53 Why do we have product metrics is a good place to start. I was actually um doing a podcast interview this morning. He he mentioned this this quote I really liked that a metric stops becoming useful once you start trying to optimize it. Product metrics are not about the metric. The metric is not the goal. Product metrics exist to support product decisions. So they're effectively just tools that allow us to answer what are we going to do next? And they should be directly tied to real decisions. The cost of a wrong metric should be the cost of a wrong decision. And the benefit of a really strong metric should be a benefit of making good decisions with your in your product over and over and over again. You know when to ship a product, when to roll back, when to invest further, when to stop. Why do we make product decisions? So product decisions, these all exist just to help our users or our customers accomplish the goal they came to the product for. Humans don't wake up and be like, I can't wait to engage with SAS SAS platforms today.

3:57 That's not their goal. the ways in which a lot of us measure them around Dow and retention and and these kind of metrics don't actually reflect the behavior they're trying to do. I'm going on a trip to Hawaii. I want to book my tickets so I can get to Hawaii or the holidays are coming up or my friends birthday's coming up. I want to get them a gift that they really like. So that's why that's why they're on Amazon. That's why they're on sites booking their their their flights. They're not on there to have I don't know long dwell times. So this creates this gap between what a user is actually trying to do and what we see them doing in product, right? So you can have a user, you can have two users who have very similar observations within our system. They both viewed a page, they're on there for five minutes, they had the same amount of scroll events, but one of those users in the real world, it might come out with, say, we're talking about Amazon or some sort of e-commerce site again that they've identified the right size and great fit of a product that's gotten to them and they're going to keep it and use it forever and someone else could find some basic info or or not even buy it or or they get it and they return it. And so we're a lot of what product metrics are is just trying to bridge that gap between these system outputs, these product outputs, and the actual users's job to be done or value. And so if you worked in in product analytics or product data science or product development, you're probably familiar with the concept of northstar metrics.

5:24 And the reason I go through all of this is because, you know, I want to lay the groundwork of like how strong metrics work today before we get into how AI kind of upsets it. But Northstar metrics are really that bridge between these really frequent and measurable but gameable system outputs and the actual business via the user value that someone's acquiring. And so if you think about it kind of in terms of a hierarchy, you have these leading indicators, these input metrics, these these system outputs at the very bottom.

5:56 These are things that you can get down to the millisecond. They're really easy to to measure and quantify most of the time. At the very top, we have super lagging indicators, things like business outcomes, net revenue, or churn, things that could take us months. So if we make a change right now in the product, we'd see those input metrics change immediately. But if we make a change in the product that's bad, we're not going to see that user turn potentially for months later. And we can't wait months to understand what's going on in the business. The Northstar metric helps us connect these two together so that we can leverage input metrics that we know are grounded in like user value.

6:40 this metric that we basically prioritize like what the team's working on this quarter to like these these bigger business outcomes. So, it's a way to focus the team, prioritize their work, but also to tie their really granular day-to-day movements back to the business and the user. So, what we're going to talk about mostly today is around that lowest level, those driver metrics or input metrics or system outputs or whatever you want to call it. Um, so these are really attractive because they're frequent and measurable as I mentioned before, but they're also pretty high risk because they only work if they directly connect to that northstar. So they're like highly gameable. For instance, you know, if you have an input metric like users clicking on my email and we know if someone clicks on the email, they're probably going to come to my site and they'll get value out of the site. Well, one way to increase those email clicks would just be to spam them with a bunch of more emails, and that's not actually going to end up providing them with whatever user value we're trying to get them out of when they come to our platform later on.

7:44 So, that's kind of what I mean about gameable. There's all sorts of, you know, dark patterns that arise in in products, especially consumer products, keep people on certain sites longer and engage them longer um that got optimized around metrics, but don't actually end in user value. I think it's really common in like engagement metrics. Um, we're not going to go through how to kind of ladder up that input to Northstar to business outcome causality today. We'll actually do this hands-on in that course I mentioned, but I'll also host a lightning lesson in January that goes into some of the methodology because there's different approaches to this. There's stuff that's pretty lightweight correlation. There's the gold standard of experimentation, actual AB test, but that's often not always feasible. And then there's this uh middle area of causal inference modeling. So we'll go through some of those approaches in this upcoming lightning lesson in January if you want to check it out. So that all being said, traditional metrics and traditional measurement assumes this stable product logic, this deterministic code. So we know that a user input will result in a predefined or at least some predefined set of system outputs. There can still be some variance there, but it's a more understandable distribution that we can then more confidently connect to our northstar and business outcomes. Um, AI breaks that assumption around that like really deterministic output because we're dealing with a stochcastic model and so you can have many possible outputs, many outputs like we'll never see and then like an actual output. This is a challenge because, you know, when you're in the early stages of a product, you might see AI products with like amazing demos that are working on like a very small amount of inputs that they've honestly like may have nailed to some degree. Even those demos from some of the AI products like don't even go well on predefined inputs since they're still using a probabistic uh system unless they're faking it. But uh as your user base grows, as the use cases change, even as it doesn't, the same input can have wildly different outputs. And so that's why I think you're seeing a lot of like AI evaluation become really important now cuz we've had a few years of companies since GPT 3.5 come out, start working with this stuff in their product. And some of them have nailed in development, but now as they scale, they're seeing all this output that they weren't able to predict, and they're not hearing about it. and their metrics, they're hearing about it from their customers.

10:14 And by the time you've heard about it from your customer, unless you're working with some customer somewhere who's in a beta and is in experimental mindset, a lot of the time it's too late, especially with all the agentic workflows, which are even harder to diagnose. And so people are already coming into AI products very skeptical and it's really really easy like anything it's really easy to lose trust and then it's really hard to earn it back which is I think is probably one of the reasons where I've seen a lot of teams decide like we need to invest in AI evaluations. This is where at least where I see like the kind of AI metric living in the existing hierarchy as one of these input metrics. These prick metrics are really deterministic previously. Some of them are model scores that are coming out of existing models. But now you're ending up with these like nondeterministic answers that kind of screw up comp or complicate at least the northstar interpretation. High level how I think about this the unit of of observation if previously the unit obser of observation for that input metric was the output of a model where we could have some sort of precision or recall or accuracy score. Now I think it's the output as experienced by the user and this makes it really complicated when we are trying to answer the questions what do we evaluate because it's it really depends on the product and use case and the user for some people evaluation may look like just like correctness or faithfulness for others it's going to be completeness some could be like safety or tone or style but at the end of the day I think a really good place to start for me at least is just to start with user experience. That's what I try to ground myself in every time. And then after I can understand the user experience, I can begin to form hypotheses and translate those to metrics. So process I kind of go through is start by defining the user value and the workflow in the real world you're trying to improve.

12:10 Outside of your products, like name the workflow step, the thing that the person is trying to achieve. Define what bad would be when they're trying to do that work. What good might look like? These two questions are ones I try to answer at this step. What's the job the user is working to accomplish? And then hey, how is AI supposed to help? If you base yourself in that, like you're going to be so far ahead than folks who are trying to take cookie cutter metric suites from other products that don't pertain to their users and plug them in to their own workflow. Start from where the user is. Once you have an understanding of that, you can begin translating it into hypothesis. So if we increase this should say AI quality, we would expect to see X resulting in Y user behavior.

12:56 So you're effectively converting those thoughts around what good and bad look like to claims about how it might show up in the product and then how the user may behave in the product once they realize that. So what would AI output that's helping accomplish that job look like? Can we develop a measurement of the AI output that reflects that and what would we expect the user to do when the AI feature works well? There are many different ways to kind of measure that and judge that. A lot of it comes down to working with some form of ground truth. There's many different approaches to getting ground truth and some are going to be more robust than others. For me, I typically opt for ground truth that can be as close to something that's coming from the user themselves as possible. Whether that's historic user data or some sort of output within the product itself that you can align to. If you don't have users, you're not going to have that. So that's not always the case. This is why we see a lot of subject matter experts and domain experts becoming really important now with their leveraging them for review.

14:00 Other ground truths could be authorative sources. So other documentation you're matching to or simple rule sets. And how we judge against these can vary quite a bit too. Um, it could be purely human review, some sort of like a rubric across many humans. I'm sure you've heard of of model judges for for scaling some of that human-based review and even programmatic checks. One more plug for another lightning lesson coming up in February. A lot of this kind of development of ground truth and validating ground truth and review requires some form of annotation. At least what I found is that the best kind of way to do these annotations and as remove as much cognitive load and effort from that annotation as possible is to reflect it as close to the user experience in the product itself as you can. If you're annotating chat, for instance, having some sort of UI that has a chat interface um is a lot easier than trying to flatten chat into a Google sheet. We'll cover this in depth in the course as well, but I'm also gonna do a little lightning lesson just on like a how to vive code your own custom annotation UIs for your product.

15:06 So check that out. These like uh judgments or ground truth options. The ways you judge these can vary in different ways. You can create metrics that are structural simple checks. You can create ones that are similarity rouge or blue or some bespoke version of that which you're looking at similarity between text or you can do some sort of judgment check review or element the judge. I think the main takeaway though is that you probably want to use more than one of these things. Even if you find out that you can use similarity scores that work really well for your product, you're still going to probably want some sort of semantic check as well in order to to diagnose them. Uh so you'll have you'll end up with many probably different metrics and you'll have to figure out a way to basically prioritize and filter them. A great way I find for doing that is kind of now returning those metrics back to our kind of metric hierarchy and and identifying at least a if not a correlative or causal relationship a logical relationship between them and a northstar valid outcome. So we can identify a smaller suite of kind of key quality indicators due to its unpredictable nature. Even if you've gone through that whole define user value, form a hypothesis, I've linked them back to my north stars and I've filtered and prioritized some suite of metrics I'm feeling really good about.

16:22 It doesn't really end because as we mentioned in in the beginning, as as anything changes with those inputs, the outputs are going to change too and and new cases could potentially arise. You're just not going to have the foresight to basically account for every definition of what the user's needs are upfront in the beginning. And so it is a bit of a continuous process here, not a one-time setup. And so where that lives in the outcome here, you know, I think you kind of insert this new loop. Your system output isn't going directly to inputs as well into your northstar metric. You're you're creating this whole another loop of of defining how do we interpret that system output rather than directly correlating them to north stars. And if you do that, I've found like you're able to iterate a lot faster because you create a metric that people can look at in the leading day-to-day.

17:14 Uh you get better alignment across your team, which just makes conversation so much easier when you're trying to figure out what you're all working on. And you also are able to set up a lot more safety nets for these launches as people are a bit more riskaverse to AI products. Okay, nice. This is good on timing so far. So, few upcoming things. Yeah, check out the course. Feel free to yeah, connect with me on LinkedIn. I think my email is on that course page as well if you have any questions. You've got the promo code there. I also write on Substack with my colleagues uh Hi and Sravia. We both run that podcast data neighbor. So, we talk a lot about data in general. We post our podcast episode kind of readouts on there. And then we've been talking about AI evaluation a bit lately. We'll probably be talking about kind of like powered analytics as well. A lot more on there, too. Um, we got that new series I mentioned coming up in February. So, we're talking to the VPs and leaders from Brain Trust, uh, Arise, Prompt Layer, Trace Loop, a few others, Compos. Check that out. And then we've got those lightning lessons I mentioned are on validating the impact on the business and creating custom UIs.

18:23 I'm sorry I didn't leave enough time for questions. I did want to pass it over to Stella because she has a really great course. >> Oh, thank you Sean. Thank you for leaving time for my shameless plug here. I also teach a AI evals and analytics course on Maven is uh called AI evals and analytics playbook where we share our battle tested playbook that we use in practice industry on evaluating different types of AI products. Yes. So you're all welcome to join us. Our next cohort starts in January. January 16th is going to be a two-eek course. We have one session on Saturday, one session on Sundays for two weeks. So, four sessions in total.

19:03 >> Thanks, Stella. And uh Stella, you have a great uh Substack as well. So, I'd mention follow them on that. >> Thank you for mentioning that. I also dropped it. Also share a lot in Substack of my experience working on AIE valves. Uh you're all welcome to join the conversation. Feel free to uh connect me with connect with me on LinkedIn and also comment in my Substack post. >> Sweet. Thanks, Stella. Um I'm going to close it there. We're at time. That was a bit of a speedun, but uh thanks for joining and if you have any questions, connect me on LinkedIn or or email me or whatever. Thanks everyone.

19:55 Jack back. Dr. Jack.

Summary

Sean, a principal data scientist at Entra, discusses the importance of redesigning product metrics for AI products, emphasizing the need to align metrics with user value and business outcomes. He highlights the challenges posed by AI's stochastic nature and the necessity of continuous evaluation and adaptation of metrics to ensure they effectively guide product decisions.

- Product metrics should support decision-making rather than be the goal themselves.
- Northstar metrics bridge the gap between input metrics and business outcomes, focusing on user value.
- AI products introduce unpredictability, complicating traditional metric evaluation.
- Understanding user workflows is essential for defining relevant metrics and hypotheses.
- Ground truth for evaluating AI outputs should ideally come from user experiences or authoritative sources.
- Continuous iteration and adaptation of metrics are necessary due to changing user needs and product dynamics.
- Collaboration and alignment across teams improve the effectiveness of metric-driven decision-making.
- Upcoming courses and resources on AI evaluation and analytics are available for further learning.
© transcribe · For agents Built with care and craft by Gokul Rajaram