Transcript
0:01 All right, welcome everyone to validating business impact of AI features. We're going to we're going to get started pretty soon here um as people roll in. But uh let's I do have some questions. First question number one um in the chat. Let's wake up the chat a little bit. Can you just drop in the chat if you can see this screen if you can hear me?
0:32 And bonus points if you say something other than yes. What's where are you located at? I'm located in uh South Lake Tahoe. You can say where you're located. What's the weirdest thing on your desk right now? Oh, Vienna, Austria in Dubai, Paris. Okay, this is way cooler group than I thought. I was like, everyone's going to be like, I'm in San Francisco. Oh, this is awesome. Okay, Austin, Texas. Nice. Dang, you guys had a rough lately. You guys have more snow in Austin than I have in Tahoe. Um, Belgium.
1:06 >> This is cool. >> So international. >> Yeah. Oh, you guys are all like up late for us, too. Is it late or early? I think you guys are later. Yeah, like eight hours or something. More than that. Um, nine. Okay. Dang, that's past my bel bedtime, Jacob. So, I appreciate you coming on. Um, okay, cool. 12. What? 12 a.m.? All right. You really want to validate business impact of AI features? I like it. Sweet. Uh, okay.
1:40 Uh, question number two. Okay. So, I'm Sean. Um I'm here with my uh colleagues Hi and Sravia. Um and uh okay question number two. Who has no idea who we are? Who has no idea who me and Sraia are? Or you can say you do know who we are? Any anyone that know or you don't have to say anything that's fine too. Um so I'm Sean. Uh I'm a principal data scientist at Entra. We're a AI legal tech company. I lead uh AI evaluations there for our data science team. Also a AI educator and instructor with Maven which I assume you can all uh kind of figure it out. Uh so I I have two courses here on Maven. AI evals for product development which is kind of what we'll talk about today and then uh another course on AI analytics for builders that I run with high and stra.
2:41 Uh it's about kind of uh teaching folks how to build agentic workflows and use AI tools to effectively like offload a lot of data science analytics tasks. Um maybe hi Sarrai. Do you want to say quick what's up to the room? >> Yeah. Uh hey everybody. My name is Hi. Uh I lead data at uh the same company uh as Sean at Entra and I've had uh probably almost two decades of experience leading data science teams across Silicon Valley big tech companies um names like LinkedIn, Pinterest, Meta, things like that.
3:19 >> Hi, I'm Shravia. I lead uh data science team uh in superhuman previously called Grammarly and pretty excited to be here and also uh you know having another course uh with this group here called AI analytics for builders. Yes. >> Sweet. Thanks guys. Um so hence are going to be monitoring the chat. If you guys have questions as we go through or comments, I mean, we'll try to make it as interactive as possible, as much as you can for something like this. Uh, but uh, you got questions, drop them anytime. Um, you can kind of like plus one or smiley face or whatever on or or down vote, whatever you want to do. I don't know if you want to thumbs down someone else's question, but you can.
4:03 And hi, and Stravia will monitor that and they'll uh, stop me. Otherwise, I'll keep rambling. This is going to run for they've got it scheduled for 60 minutes. My next meeting is not till 1:30 Pacific, so I can stay about 30 minutes over. This might honestly go over anyways, depending on how many questions we have during it. But I want to try and get as much of your questions answered as we possibly can. So, um, let's jump into it though. So, uh, yes, validating business impact of AI features. So what we're going to talk about today um AI evaluations, AI quality. I think a lot of folks I assume in this room have heard about this over the past you know 6 months to a year as a lot of these companies who are building AI products and features who had you know really like pretty cool demos things work pretty well start to scale their products and uh start to see oh shoot like uh I don't really have a measurement around like quality here um so there's been a lot of kind of content literature and courses uh coming up actually I think there's more I us talking to uh my friend Stella the other day who also teaches AI val course on uh on Maven that I recommend checking out.
5:18 Um but we were talking about how there's so much actually content right now around like how to do AI valves, but then there's not that many people actually doing it. So a lot of the content right now is kind of focused really on that AI quality piece. Um the focus of this is then to connect that up the chain to business impact stuff like growth and revenue. And if you think about it, the reason for that is because your come on slide. There we go. your CFO or your exec team, your leadership team, uh your investors, uh your customers and users. honestly like they don't care about like your F1 score or precision or recall or like your success rate of like some obscure uh slice of like an error analysis of like oh my uh my uh we increased the precision for this very uh bespoke part of the product to answer this question by 3%. like no one cares about that besides like the team that's directly working on that and what they really care about like all those people um I mean like leadership execs investors like they care about realistically you know did this increase profit how much did it increase risk because they're trying to make a lot of decisions around investment and prioritization users what do they care about they really just care about you made their life easier you made their life better um you provided them some value. And so the critical answers, the critical questions we really want to answer as we think about AI evaluation and you know evaluation of products in general is is not just like this quality aspect but you know what changed in the product um what changed for users and what changed for business.
7:09 So I want to get um I want to get a little idea of like what's where everyone's at today. So, what I'd like to do is if you could kind of drop drop in the chat a one, two, or three here. Um, and then I'll tell you what these kind of relate to. So, drop in a one if you're and this doesn't have to be like you're with your team at some big company building shipping AI features.
7:32 It could be a personal project, a pass a passion project, whatever. But dropping a one if you're kind of building or playing with or developing some sort of AI feature but um you don't really have like a a metric or metric suites to understand like is the quality good or not. And then you know uh if if if you if you got two two is like okay you got the quality metric but you don't really know like okay I tweaked this change the quality went up but I don't actually know how that translated to dollar actual value. And then number three is like you got features, you're building, you understand quality, you have comments. Okay, ones and twos. I expect a lot of ones and twos.
8:18 If you have if you have some threes, then you can help me lead this session. That would be great. Cool. Yeah, >> between one and two. >> One's two. One, two is great. There's also zero. There's zero's okay, too. Zero is like I just I'm here. I'm not building something yet, and I want to learn. one and twos are really good. Um, so let's talk about kind of those different levels though. Um, our goal today is to provide you with a framework, a simple framework um, to go from level one to level two to level three. So if you're level one, you're going to climb all the way up. If you're level two, you're going to get up to three. We're trying to get you from uh, level one, which is where you're you don't have to take a quiet metric. So you're effectively kind of shipping blind in a way, right? You're it's hard to make kind of confident decisions around investments in a company around AI features when you don't have any sort of measurement around the quality because you're not sure like what to work on. Um the the the the rough thing honestly with shipping blind is like you will get AI quality feedback but it's going to come in the form of qual qualitative feedback from customers and users. Usually the people who give you that qualitative feedback are like the pissed-off ones. So it's like not even it's not going to be fun and then in the worst case scenario it's going to come in form of like a user or or churning user you know. Um so shipping blind uh it's it's like building's like the first stage obviously but we really want to get to level two. level two. I think a lot of the existing content and curriculum around uh AI evaluation out there right now is getting folks from level one to level two. It's really really important step. You need that before you go to level three. Level two, you're like AI quality confident. Um you're you're iterating and developing uh offline and seeing quality go up. Um, but then when you ship it out, like you're not actually sure like in like a month or 3 months or 3/4 like what this actually do to our bottom line? Did it expand users? Did we acquire more users because of it? Did we get better attention? Did we increase our profits?
10:38 Um, so that just makes it hard because with every company there's, you know, tons of things we could be investing in. uh we could be investing in AI features, we could be investing in UX, we could be building new products, we could be iterating on existing products and features. It's really hard to make those decisions and calls and we don't know when what we built is good enough yet. A lot of that can be achieved by validating up to like business impact.
11:06 And so level three, uh, not a lot of folks are are at level three today, not just in this room, but like in the entire industry. Um, and even companies where like they are at level three. They're at level three at like a few features of like many features like like huge companies like metas and stuff like that like they're at level three for some for some stuff, but not everything. And this is where you can like really confidently connect your like AI improvements to business outcomes. So that's where we want to get you today.
11:37 Um, and the framework we're going to use to get there is something that we call the AI impact chain. So this framework, uh, it's there's four links in this chain that bring you from like an AI feature change to eventually revenue, maybe revenue via some other business outcomes. But there's kind of four links in this chain. Uh, there's your feature change. the first link. Then there's AI quality, then there's customer value, and then there's business impact. It's not super mind-blowing crazy chain. Has anyone kind of like seen anything like this before? You can drop in the chat, yes or no, any sort of like this sort of chain from feature change all the way up to business impact with some of these links in between.
12:31 Okay, that's good. then we're here to learn. Um, so you need every single link in this chain to get from feature change to business impact. Uh, anything that's missing here and you're just kind of guessing what your business impact is or what your customer value is, what your quality is. Um yes, formula pyramid maybe like northstar metric pyramid uh where you have kind of like your business impact your northstar metric which is kind of customer value input metrics at the bottom uh which is like what AI quality in this yeah gach good call out. So, if you get all these links to hold together or whatever, if the your whole pyramid is in place, um I like chains that pyramid because you can break a chain and I don't know, pyramids like I don't know how you could have missing things. Um, but uh if if it all holds together, then you're you're proving with evidence and some level of confidence um that whatever feature change you're making to your AI uh product or feature there is actually leading to some business impact in the long term.
13:41 Okay, why is this important? So why can't we just have feature chain AI quality? So AI quality is not necessarily customer value. It's kind of like you think I I I would think it is like quality customer value goes hand in hand. Um but if we kind of really think about what types of metrics these are it becomes very clear how they cannot be connected. So, when I think about quality metrics, and there's probably more than this, but if I were to roll it up into like three categories, and today we're not going to go into how to create quality metrics, but we did a free workshop about a month ago, and you can you can go to data neighbor.com, and you can watch the recording that if you want. It's called redesigning product metrics for AI features. Um, AI quality metrics, they kind of take these three forms. I think there's like a correction rate. So this is like your AI output. How close did you get to what the user actually wanted? And so it's kind of like um could be like an edit distance. It's usually like a range from zero to something. Maybe 0 to 1, 0 to 100, whatever. Um but it's like some kind of like closeness metric of how close you're getting to what they want.
15:00 Uh acceptance rate. This is uh more of like a zero one like is the output you did is the output your product created um accepted by the user or they say like oh this is 100% good move on to the next step in my life. Um that's like a proportion usually. It's like okay 70% of the output was just like one shot good for the user. And then there's error rate. you if you've kind of read some of the existing AI eval uh literature out there or some of the other courses out there, they spend a lot of time focusing on error rates through like error analysis. It usually takes the form of working with a subject matter expert, creating some ground truth, going through and annotating a bunch of AI output, identifying and through those annotations categories of things that went wrong, and then forming uh like precision or precision recall, F1 accuracy metrics around like are we reducing these very uh specific types of errors. It's kind of more like debugging, I would say, similar to like MLOps debugging. um really important.
16:10 It's extremely diagnostic. It can help you get better customer value. It can help you increase your acceptance rate and your correction rate. That's just one very small piece of the puzzle. And these these things do not necessarily mean customer value metrics. So when I talk about customer value metrics, depends on the product. say like we're going to talk mostly about kind of like SAS product today or more like a kind of like a workflow product where it's like hey we're just trying to make someone's life easier trying to offload some of their work for them automate it or make it faster. Uh obviously there's other types of value you can provide like I don't know Netflix or social media or something like that where it's like engagement entertainment based. Uh, we're not really get into that mostly because I don't like social media that much, but um, customer value metrics.
16:59 So, today we'll talk about like time to finish a task. That's like a very clear workflow. Customer value. Are we speeding up something that would take you longer before? That's success. That's customer value. Um, task success rate. It's like just like you go into your job, you're trying to get something done, you succeed or you fail. Are we increasing the number of times you succeed? Is our product increasing the amount of times you're successful? These are more around like kind of customer value metrics. And so you can kind of see where you could end up with AI quality metrics especially around say like error rates but also correction rates like air rates for sure where it's like our output is uh I don't know constantly uh saying that uh the sky is turquoise when it should say the sky is blue.
17:58 That's an error. Let's like fix that. But the customer for whatever product is, this is a terrible example, but maybe they don't really care about that distinction at all. That doesn't actually give them a negative or positive experience, but in terms of your output and error error rate, that does could show up as a metric. And there's many different metrics that can show up like that. You can also have like not just like wrong metrics around AI quality but um like over optimized like trying to make them perfect.
18:34 Everything doesn't have to be perfect like a correction rate for instance. Um you don't have to get to 100% one shot all the time. Sometime like humans like sometimes they want to like hold the reins and like know they have some control over what they're doing too. Maybe you just have to get to 70% and that's what gets you the incremental customer value. Effectively quality metrics they only matter as so long as they are providing incremental customer value. And so you have to have that connection in the chain.
19:11 Customer value then um only matters if it okay we should always bring our customers and users value but there are ways we can create customer value or user value product metrics that don't actually reflect true value in the real world and don't actually uh consequentially lead to business impact. Like an example of this would be something like um a lot of companies over might overoptimize for say like engagement. So uh like Amazon probably definitely doesn't do this because they have their together, but I was on Amazon a couple weeks ago. I bought I recently bought this like Ooni pizza oven. I don't know if you all have heard of these things. are kind of like a wood fire or gas pizza ovens to like try and make like those Italian style pizzas is like one of my things I'm trying to learn this year. And so for these pizza ovens, you have to have this thing called a peel, which is like basically a big paddle, metal or wood paddle with a stick on it and you like shove your pizza in and then you take it out and you're like, "Oh, it's like kind of burnt." And like turn it around, put it back in, out and in. Cuz this thing's super hot. It's like 900° F. And uh so I'm like, "Okay, I got to get this freaking pizza peel thing." And I go on Amazon and like it's just crazy. Like they have there's like ones that are $15 and then there's ones that are like $125 and they have tons of reviews. There's like hundreds of these things. I don't know which one to buy. So to to them on the back end, it's like, "Oh, this dude's like really highly engaged with our site. He's going all over clicking on things, but really I'm just kind of lost and I don't know what to buy."
20:58 Anyone ever kind of experienced that in terms of metrics where they have created some sort of user value metric that they think kind of like is reflective of engagement is a big one of something good behavior the user's having but it's actually like a bad behavior anyone run into I know Hy and Sravia uh we have both because we worked together at a company where that happened many a time.
21:28 Um, so this is kind of like where business impact comes in because uh it's really really easy to make up messy metrics. We're actually going to do a free workshop on this. It's called designing metrics that matter with high. I think it's like February 18th. Again, go to data neighbor.com. You can kind of see this. I I would just say go to data.com and then another window right now. or hi if you want to drop in the chat cuz I'm going to like throw out these names of other free workshops that are upcoming you can check it out see if there's something that resonates with you usually treat moving one step ahead in the journey as positive yeah exactly time spent time spent is like one of those things just like it's so tricky it's like no we want to decrease time but then there's all these platforms that also want to increase it. Yeah, that's a good point, Kesh.
22:27 I think moving through task success and getting through the workflow is definitely like a very positive uh thing. So, I'm I'm very very sure that Amazon's not optimizing on how much time someone's on the page. It's probably like add to cart uh put in your billing checkout cart is probably what they're more optimizing on, but just an example. So business impact metrics are are really important because they can validate if your customer value metrics are good or not. You can correlate those. you can identify causal relationships and you can see oh if this uh customer value metric is going up or down what's it doing to my business impact metric my business outcome in the long run if those are also going up that's probably a good sign like if you're moving customer value in such a way that you're expanding like the seats of a SAS product or retaining through like increased renewal rates or getting more revenue per account or if you're lowering costs or reducing the need for like your your CS time on cases. That's like a really good signal that you are truly providing customer value. It's just very tricky because it's a lagging uh indicator, which is why you need this uh uh kind of four links in your chain to to to get from like the very leading indicators around feature changes all the way to that lagging indicator of business impact. Um common failure modes. So, this is just kind of stuff that comes up that like it seems like they're positive sentiments, but you'll can end up being getting a lot of push back on it if you don't have numbers to back it up. So, things like, hey, the model got better. Uh, AI quality seem to increase because of these metrics. But, um, I've I've definitely been questioned by leadership where it's like, what does that actually mean for the users though? are like, so how how good is it? Like how where do we have to get it to to where the users are going to be happy with it? Um, users seemed happier. Like if you get like bunch of qualitative feedback that could be like highly biased, you know, it depends who you're talking to. For instance, if you have like hand raisers who are like in a beta or something that you're talking to, um, maybe they don't want to be as negative with you. Um, again like makes it still really hard to make like investment resourcing and prioritization decisions just based off like users seem happier. Um and then kind of like other side of the spectrum if you don't connect back the chain if you say like oh like our uh revenue increased or our expansion growth increased basically like you run into these scenarios where a lot of people like to take credit for for those situations, right? It's like okay was it was it your feature change that increased revenue? Was it another team's feature change that increased revenue?
25:25 It's really hard because it's super lagging. This would be like three quarters later. Uh marketing sales like oh we did these marketing campaigns that that increases sales like oh no we like really like grinded like we were the ones who increased revenue. Um sometimes it has nothing to do with a company at all. I worked at a company uh we were Shravi and I worked at this company next door. It's a neighborhood social media app. And I remember there's multiple times where uh suddenly super engaged. Everyone's like daily active users posting everywhere. It's ad supported. We're getting a bunch of revenue from this. Everyone's claiming like, "Oh, this thing I did like a few weeks ago and rolled out. That must have done it. There's like a campaign going on." And then like every time it was like a ice storm in Texas. Um, yeah, we had we had a dude from Austin here. Like it's like I bet you neighbor I bet you Next Door blew up this past week. Um, so just like seasonality, external events, a lot of stuff can move that. So you really have to have a strong understanding of these of this change.
26:30 Feature change, what do we ship AI quality? Did they allocate better customer value? Did users complete work faster or better business impact? Did profit improve without increasing risk? So, pretty straightforward framework. That's like the thing with frameworks, right? It's like um it's like once you see them, it's like oh yeah, it's like of course that's like a no-brainer how we should think about that. If they seem if a framework is like convoluted and hard to understand, it's not a good framework. Um, but also the thing with frameworks is like when we don't have them, it's kind of really hard to just like piece it together from nothing cuz we're all focused on a lot of other stuff. We're all focused like a lot of us are focusing on feature change, AI, quality. It's really hard to find time to think about customer value and then if you get these three parts of change, it's really hard to find the time to carve out for business impact.
27:29 So fairly straightforward concept as most frameworks are. Um all right about 30 minutes in. We'll keep chugging. I want to walk through this in terms of a kind of example here. So in this example, let's imagine that we are all on a team building a product. It is a uh agentic AI data scientist analyst converts like text to SQL. Uh and the reason I picked this is because we have a whole course on this AI analysts for builders. Uh Hy and Stravia and I actually on the side we also do this podcast called data neighbor. Maybe hi you can drop the YouTube link in there or something. But uh the past few months we've interviewed bunch of uh CEOs and and leaders and uh VPs of like BI and Agentic AI agentic analytics uh companies um much like what we're going to talk around here. So if you're kind of interested in this space a little different than AI evaluations but around agentic analytics um check out our podcast over the next couple months every week we'll be releasing an interview with you know people from like the VP of uh VP of Tableau to like CEOs of like startups in the space like live docs and count to a bunch of stuff in between. Um all right but back to our use case. So we have this AI data analyst text to SQL product. If we think about it from the user workflow side, um the user goes in, they ask a question in chat box, AI writes some SQL, presents it to the user, they check if it's good or not. If it is, it runs the query, then it develops a chart, and then exports that chart, and the user has this nice answer to his question or her question in chart form. they can share it with someone else as evidence or whatever it's needed for. Maybe it's like um yeah, there's a very important customer talking to uh an account executive and they need like chart really fast for them. That's definitely happened to me in the past. Um so there's many steps for like failure in this. So I could a user could write the question and then it could return SQL, but it could be totally wrong. Right?
29:57 wrong like group buys, wrong tables, it pulls from, wrong information. It just pulls. It could get it right, but it could have syntax errors. So, it doesn't actually run. It could get it right and run, but the chart could be total trash and just like not be helpful at all. So, now I have to go take this data and like go into like Excel or Tableau or whatever, make my own chart. uh maybe I have to go back and forth and like edit the query myself, ask the question a different way or I can go perfectly I can ask a question and get my chart and be on my way.
30:28 It's like a seatbased we're going to say like it's a seat based model. So, um, you know, for every additional seat we get for people in within an account who are able to use this product at like our for our customers, um, we are $200 a month. And, uh, the current state, what's happening right now is that there's a new version of this. There's a V2 um, of this uh, text to SQL AI analyst. And what we all need to do right now is figure out if we're going to roll out V2. And so I'll tell you a little about V1 versus V2. V1, it's like direct SQL generation. There's no kind of planning step involved. There's no really grounding in like the existing data warehouse schemas or anything or context around that. No validation. Just kind of like does a single loop thing. And you know because there's not like uh this planning step or like validation step it's like a lot lower compute. Uh v2 this new thing we're trying to roll out there's some reasoning steps when it generates the SQL. It actually creates a plan before it generates SQL. So like maybe it figures out like hey here's all the different CTE I'll have to get to.
31:41 Here's things I'll have to filter out. Um there's more context around the existing schemas in the data warehouse. And there's another agent that like does a double check, validates, repairs a SQL. Obviously, a lot higher compute. So, let's imagine we're all remember those level one, two, three. Level one, we're building shipping. Level two we know AI quality. Level three, we know about business impact. Let's all imagine we're level one right now. What do you guys think? Right now, this feature V2, it's out to 5% of people.
32:14 Um, you kind of know some stuff about it. Or do we want to roll it out to more people? Do we want to roll it back? Want to stay at 5%. V2, V1, I guess. Drop in the chat. Which would you which would you prefer right now? V2. V2 sounds pretty nice. Very smart. Okay, let's get a little more information about V2 and V1. So, we won't get AI quality yet, but we do have some latency figures. So, latency is how long it goes from this first step, ask a question to AI writes SQL. How long does that take? V1, it took about 30 seconds.
33:04 And then V2, it now takes two minutes. And latency is really important because you can imagine yourself in the in a position kind of like why I said like maybe like there's some customer that's going to churn. It's your biggest customer and customer support's talking to them or the AE is talking to them and they're chatting with the CEO and the CEO is like what do they want? They're like I don't know they want they have all these data asks. They have all these things they're asking us for right now.
33:30 Like we just got to give them what they want make them happy. and the CEO is now going to the product teams and be like you need to pull this data for me and like you're in the middle of building something else right now. Um, you know, from the customer point of view and so the customer wants to use this this AI analyst to like pull that data for the person who's complaining. Um, yeah, the longer it takes like the more kind of pissed off they're going to be.
33:58 What do you guys think now? So the we know that the response time we added a bunch of kind of planning stages and validation but now it takes four times f longer to answer the question. Yes. And we got V1 speed. It's kind of like what I say is like you can compete on speed, quality, price. Speed's important. Maybe V1 with contextual questions for V2. Yeah. Okay. Let's keep going.
34:28 What's the Oh, man. Rudion, you're you're skipping like five slides, man. You're skipping to the next chain. I got to build I got to build these examples first. But yeah, user experience. I think we know where it's going. Um, all right. What? But we'll we'll play along for a little while. Okay. So, it took longer, but then check this out. So it took longer to get that initial query running, but then the user edits per query. So let's say we go back to our little flow here, ask a question, AI write SQL. The amount of times someone has to like go back and reask the question or kind of edit the SQL and try to run it again. Um that's decreases significantly. So used to you used to have to go back and reask your question like on average like two and a half times. Um and then uh now it's taking like way less times to go back. So like it does take longer to get that initial uh that initial response, but it's a lot closer to a one shot.
35:45 What do we think now? Yeah, good comments in here. I mean, the duration is about the same thing, right? Yeah, it's like Yep. Reasoning adds a check. Yep. Add a context table columns plus V1. Yep. All right. I'm going to skip through a couple of these ones as we build up. So, let's say now we get into like, you know, it takes it takes less edits to go back and forth even though it takes a bit longer. We find that maybe like the syntax errors are a little more. So, you have to go back and tweak those. And then what you kind of have here is you have your your kind of like completion that completion rate when I talk about how close you can get to it. You have kind of your acceptance rate. How often does it just like actually run through?
36:36 And then now we're getting into some of those like error rate metrics. Something like the precision of on customer support question. These are like the kind of things where like you have a golden ground truth data set and you've gone through and annotated and figured out like we categorize these types of question as customer support questions. How does it perform on them? And we actually find that in V1 um it performed it it it got it right 50% of the time like exactly. But on V2, it gets it right 25%. So 50% increase in the amount of times it's getting these customer support questions right. What do we think now? V1 versus V2.
37:19 Yeah. Quality question with weights to just get a number. Yeah, we definitely need to figure out one number. We have too many numbers going. This is like kind of like the challenging thing with uh quality metrics is that like especially the error analysis metrics, they are never ending. You'll constantly be in the iterative loop of expanding your kind of failure modes and you just end up with unlimited list and you really have to assign some sort of weights and prioritize them.
37:54 So, love what I'm seeing in chat. And then, you know, we'll stop playing this game, but like you could imagine like maybe the customer support questions it gets a lot better at, but then the finance questions it gets worse at for some reason. Like it used to like finance questions are hard. like there's all these industry um you know financial calculations but then within a company there's definitely internal interpretations of all those of those metrics. you just filter them and edge case like the hell out of like financial metrics and um and so maybe it actually gets worse at those and and now in this situation actually it's pretty interesting because you'll end up with like a team of people like customer support that could do look at a readout that you have on AI quality metrics and they're like hell yeah ship V2 like this is going to save me a bunch of time and then finance is like no way are you insane like we have to report those numbers out publicly like we're not going to use this at all if you just decreased it to 30%. 40% already kind of sucked. So, um yeah, I think a couple of you got it in like the very beginning, but like you need to get to that customer value. So, hopefully that kind of drove the point home a bit. That's why customer value is important. It's like it breaks that tie and it helps you prioritize which of these metrics is really important.
39:22 Best way to do that's going to be an AB test. Um there's other ways to do it too. So like some I I hear questions before we'll talk about this more, but there's other ways we can we can kind of find these relationships too without without an AB test. But let's see like V1 V2 or V1 the time from to correct chart export. So from the question to exporting used to take 12 minutes now takes 3 minutes. V2 accelerated the chart export workflow by four times.
39:52 Really reducing time. Yeah, exactly, Meredith. Yeah, cost is a is a huge one. We should have been here. Mhm. And I don't I mean I don't know what's going to happen with like the AI industry and all these like proprietary models, but like I know a lot of them are operating at losses right now. So I don't know like how expensive this stuff gets down the line. So keeping an eye on on transaction cost is really important.
40:22 Yeah. So what do you think? V2 V2 ROI for the win. Yes. Okay. So really cool thing about AB tests is that tiebreaker there for sure. What what's actually even cooler about it I think and comes into this whole chain thing is if you do enough AB tests um you can actually start to form the relationship understand the relationship with confidence statistical confidence between those AI quality metrics and your uh and your customer value metrics.
41:00 So we're actually going to I don't have time to go over all that today. As you see we're already 40 minutes in. Uh I will stay over for questions though. Um so we still have some stuff to get through in the next 20 minutes. Um we're going to have a session just focused on this designing experiments for AI features February 19th. Totally free workshop. Pretty similar we'll do today. Um if you're so if you want to join that, if you're liking this so far, take out your phone. Um yeah, scan a QR code or go to bit.ly experiment AI and uh that will get you uh to our signup page.
41:33 You can kind of read about it. It's like the same signup page for this one. So you can tell me if you're interested in it. But what we're going to try and get there, talk about there, teach you how to do is basically say like we can now know with confidence that say like when we decreased edits by 87.5% like this metric back here, user edits per query average. Um maybe that had like we run many AB tests over time. we can draw the conclusion that like that actually is a thing that decreased the workflow time by 75% that 4xed it.
42:06 Maybe these other ones like SQL syntax rate or response latency had no effect on customer value at all. And so with enough AB tests, you're able to kind of like form those relationships, which is pretty cool because now you should still run the AB test, but you can effectively uh run your offline uh AI quality evalu.
42:42 Um, question I get a lot here, uh, is, uh, should we just run AB tests then? Uh, no. Don't skip the AI quality stuff. You still need that because the I'm sure you've heard this many a time. This output from this stuff is like non-deterministic, right? It's like it can feel like you're locked in with the output sometimes and other times it feels like you're just like pulling a slot machine. So you don't just want to release this stuff out in the wild. AI quality metrics are extremely important.
43:20 Um guardrails when you're developing early before you get something in the customer and then once you get those to a place that you're comfortable all depends on the cost and risk as well. You can release that and do an AB test. um you can find those relationships earlier so you know like how high you need to get that up stuff up before you do a DPS. We'll talk more about that at the end as well and obviously we'll have this whole deep dive. Um yeah so bitly experiment AI no probabilistic nondeterministic yeah it's like the same I mean yeah good question um now you have me thinking I mean nondeterministic in a sense like you really I don't know what it's going to say. I guess if you like tweaked the temperature and everything of the models, then you could get to probabilistic, but in some ways I feel like it is like non-determinable, but I use them interchangeably.
44:29 Right? So, we've gone to customer value. We've gone to feature change, AI quality, customer value. This tells us if we want to ship a feature or not. But we still have to find that p last piece in the chain. Um business impact. Can we validate it? How do we understand if that feature change we're we're making is impacting the business? This is really hard to do because it's a lagging indicator. It could be months. It could be quarters. Um that's why like we don't ship like we don't we're not going to wait three quarters to ship something.
45:04 we ship off customer value um because we can measure that in days or weeks. We can measure AI quality and however long it takes you to get the output could be minutes or seconds. Um but we still want to know like what business impact is because it helps us with investments. So basically what I think about AI quality and customer value and business impact I think about AI quality I use in development and helps me know if I want to test this with customers get in front of them. Uh customer value tells me that I'm testing it in customers in production lets me know if I want to roll it out or ship it to everyone else. and then business outcomes, business impact. This tells me what do I want to invest in and prioritize next with respect to the product as a whole. Lots of gray area and overlap between those things, but that's kind of like the three kind of ways I think about it. Um, and the way we do this business impact is something called causal inference.
46:18 AB tests are a form of a causal inference. are actually like the gold standard of causal inference. But there's many other methods of causal inference stuff like I don't know if any of these sound familiar but like uh regression discontinuity design um difference and difference regression panel regression lots of regression regression and econometrics has been around for a long long time. People have been doing this stuff uh double debias machine learning uh propensity score matching course into exact matching.
46:46 There's lots of different models around causal inference. Um, but what the output effectively gets you is learning the relationship between an input and an output. So in this case, we have our customer value. We can ignore AI quality and stuff for now. We have customer value that we have uh yeah R I mean R is like a econom e economists or econometricists dream yeah I started in R2 but I went to the dark side of Python um measuring customer value we know this we know that we reduced 12 to three minutes now we want to learn the relationship between that and some sort of business outcome that we can translate into revenue. Um, regression allows us to make that translation. So, in this case, we talked in the beginning that each seat that we get into an account gives us another $200 a month for the product. So, we're going to try and learn the relationship between uh that time to correct chart export to the number of seats will get gained by an account that month. This is like made up, but this is like how it works.
48:02 Basically, like the coefficient, like in this case, we could say something like every one minute faster, we get someone to export a chart, we will we expect to gain five more C's. This would kind of be the output of those causal inference models. Um, and if you think through this, why does this like make sense? Well, If I can get data faster, I can make decisions faster. If I can be more confident and make decisions faster, then I can acquire more of my own customers or make my customers happy or divert that time that I spent on making decisions or finding data to building other products. And so now I'm willing to buy more seats of this AI analyst tool because uh every additional person I have doing this offloading that analytics and problem solving work that they had to do themselves before uh is now able to effectively give me more time, resources and money to put and invest elsewhere in my business. So you can see just like anything as they make that product better, as we make the product better around faster, we could probably expect increases and like expansion of our existing customer base at this like seat per account level. And then after you have that, it just becomes a math equation. Like once you have that coefficient that's output by the causal inference model um you can do something like okay uh well if we get five more seats per one minute faster export and we increased uh exports by 9 minutes then we get 45 seats more per account. Say we have 100 accounts that's 4500 seats. Um, we know that a seat is $200 per month. So that translates to $900,000 monthly recurring revenue additional. And then if we bring that out to the whole year, now we're at like $10.8 million in incremental annual recurring revenue. And you can put confidence intervals on this. This is obviously like made up. I don't know if people are going to spend I don't know $200 I see for this but uh chat GPT pro and all these companies do charge something like that. Um so you can confidence inter like it's statistics it's like a distribution so maybe it's like in a range of like 8.6 to 13 million AR. But what's really cool here is now because you've connected that chain um we're able to like discern what the customer value impact is on the business and then we can move even higher up to uh the AI feature change. So we'll I I rattle off a bunch of names of causal inference models. Obviously can't teach causal inference right now. Um we are going to have another free workshop on that on February 26. Um and yeah that first number is Raj is like what would come out of our causal inference model 1 minute to five seats 1 minute to one seat one minute to 0.25 25 seats. But that is the crucial number. When you can get that, when you can control for all the other variables that are impacting the acquisition of more seats, when you can understand like the relationship in all of the pieces steps between between like we increased their speed that they got this answer, this person made decisions faster. Making that decisions faster made their manager happy. That manager said, "Hey, we need to get this tool for more people." finance approved it. Like when we can like figure out all that links through a model like this, uh the rest is just simple math. Um so we're going to go into actually what all those types of models are at a high level. I'll probably go like three or four of them. Maybe like panel regression, diff and diff, and like propensity score matching or something.
52:15 If there's certain models you want to learn about, you can drop them in the chat and I can think about uh changing it too. But take out your phone, scan the QR code, check that one out. Bitly causal AI, that'll be in late February. Um and yeah why this matters like when you get this whole chain now when you can identify all those relationships when you can do enough AB tests or cause inference analyses to understand with some level of confidence the difference the relationship between your AI quality metric and your customer value metric and your customer value metric and your business impact. Then you can literally change a feature, get your AI quality evaluation output metrics that takes you like that's just how long you it takes for your eval pipeline to run. That could literally be minutes.
53:12 And then you're just forecasting with those coefficients and you can begin to make predictions of like, hey, this small tweak I made to this feature change may have this huge $10 million impact on the company in the long run. Um, if we want to like get that, let's put more headcount in this area or let's stop focusing on these other things and work on this feature. So, it's extremely helpful um for investment and prioritation uh decisions to get this full length through and kind of went over this, but like I said from the beginning, the CFO doesn't care about your F1 score. Leadership doesn't care about that. This is what they care about. This is like the slide they want to see. They don't want to see um just the edits went from 2.4 to 0.3.
54:07 They want to see how much do we predict within some interval, some confidence interval. Uh, our AR will go up. How do you get calculate that? Oh, this how 4,500 seats went up. Why are the seats going up? Oh, because we made it four times faster. Awesome. Um, this is what this is what you want to be able to report or predict or forecast. And there's definitely uncertainty around it, but it's it's how you get from level one where you're shipping completely blind to level two where you have some idea around call manage all the way to level three where you can like begin to really have some forecasts around the business impact.
54:48 Okay, we have five minutes left. I'm going to stay I'll stay up to 30 minutes over and depending on how many questions we have. But um if you enjoyed this obviously mentioned a few of those things we have up come coming up. Designing experiments for AI features Feb 19. Uh designing metrics that matter. Feb 18. Um was there one causal inference analysis for AI features? Feb 26 or something like that end of February. Check those out. We have like 10 or 11 other workshops. We got some really cool ones actually planned. We have some ones around um doing the experimentation analysis itself uh leveraging uh AI. We have some really cool ones about like running analysis or designing analysis.
55:44 Um with cloud code um I have one around like building custom annotation tools to help um help you create those error analysis kind of metrics. uh a bunch of other stuff. Go to data neighbor.com. Hi, just put it in the chat. Check them out. Sign up for whatever you want. Um also we have in addition to all those free resources, we have a six week cohort. We're going to kick this off in April where we are going to deep dive into a eval from all the way from just like observability instrumentation to um making decisions with like your leadership team and everything in between. So think uh developing ground truth looking at offline versus online evaluations of course getting into the running experiments going really deep into causal inference um all the like common kind of failure modes you run into how do you get this into like a CI/CD pipeline how do you um yeah constantly like iterate and improve on what's working on here um so that'll be a six week cohort starts April 6th uh you can check it out at aievel.ai.
57:00 Aievel.ai or you can just stick out your phone right now. Scan the QR code. Uh it's usually 1,500 bucks. Um discount if you use this promo code biz impact. This is valid for a week. So through Feb 5, uh it's 30% off. So knocks off about $500. Um, I do find that people like like they want to sign up and then it kind of like they forget about it later. So, if you're if you're really interested and you want to sign up immediately, uh, I'll give you an extra 10% off today if you sign up by midnight PT. And just just for reference, like Maven, if Maven has like a satisfaction guaranteed kind of thing. So you can actually drop the course two weeks in and get a full refund or anytime between now and then.
57:54 Um if something comes up and if you have to do something like get a budget approval or something from your work, like I don't know, we're all like humans here. Like I I can extend the promo for you. DM just DM me on LinkedIn. Uh, you can find me on LinkedIn, Sean Butler, or you can email me, Sean Aie. You're probably saying that says Shane Eie. Yeah, it's kind of weird. It's pronounced Sean, spelled Shane.
58:27 That's my little sales pitch. But I do want to go into a couple questions. I can stick around for 30 minutes. Um, we can talk about anything we talked about here today. We talk about the course, we can talk about those upcoming sessions. Um, a couple common questions I do get, uh, I have them here, another tab actually want to pull up because feel like these come up a lot. One I get a lot is around um, like, hey, I can't run an AB test. I don't have um, enough users or my leaders are like really Oh yeah, there will be a recording of course. Sorry, I shipped that. I'll send out a recording. Um I'll send out these this promo code and stuff in the recording too. Um and yeah, just yeah, and if if I forget for some reason, I won't forget. I think it's like automated, but you can just DM me or email me and uh I'll get you what you need. Um so some people say like we can't do AB test because either we don't have enough users. Well, first of all, you have to have some users. So, if you're in the stage where you're building a product, but you don't have customers or users yet, um, on one hand, that's kind of rough because you can't calculate business impact, but on the other hand, it's like that's kind of good. You can just just focus on AI quality at that stage because you can't understand user impact without impacting users. And if you don't have users, you can't impact them. Um, so AI quality is going to be like your focus there. And then you're just gonna want to get you some early users and do more of kind of like qualitative user research discussions with them early on. Don't focus as much on the metrics at that stage.
60:14 Does per user matter? Because if it's high, then maybe you don't need too many. Yeah, exactly. It's all like when you're It depends when you run your experiment what you're going to um what you're going to like have your exposure on. If say like you're want to know like the average time to export per user, then you kind of have to have a lot of users because you have to hit uh a a you have to have a representative sample of like the population of people you want. If you're looking at something like time to expert per page visit, then you could you could use less. The problem with like not having a lot of users is like and you have a power user is you end up biasing just to how they act. So I would say like you actually in the case where if you have like not that many users and you have some that just like use it a lot, you may even want to be like segmenting that out and be like this is how it affects high usage users, this is how it affects low usage users. Um that's a good question though. Something else you can do just like the causal modeling from customer value to business impact. You can do causal modeling from AI quality to customer value. So an AB test an experiment is just a form of causal inference. It's like the gold standard because it really controls for everything. But there is still some uncertainty in there. There is still some randomness associated with it. Um, but if you're if you can't AB test because you don't have enough users and it's going to take you like, you know, three or six months to get enough um events to like reach power and confidence in your results. Then what you can do instead is do a causal inference. um you can basically backfill all of your AI quality metrics like six months and then you can do your causal inference relationship between that historic data to come up with um what the relationship be it's not as accurate as an AB test but it is a totally um acce acceptable kind of um replacement for that and some folks like it's not just like if you don't have enough users it's like if you're a risk giver so I work in legal tech. It's like very uh you know buttoned up industry uh banking, finance, health. A lot of that stuff like really like you kind of don't want to just like throw stuff out there.
62:58 Um there's a lot of risk associated with it. And so doing causal inference before the AB test. You can still do an AB test, but have this kind of like higher confidence call in before it for really big changes that you're not sure about. Um is a nice helpful way to do that. Um, another another question I get is around like leadership or only cares about like latency and cost. Um, which are both really important, right? We talked about uh price, quality, speed.
63:32 Um, so yeah, price and speed, latency cost, those are huge. Uh the way I approach that is basically go into the building of your features with some agreement and contract around like established guard rails like hey we're we're not going to go above this latency. We're not going to go above this cost because we know it's just uh it's basically like the business can't handle that. It's like not an option. And then yeah, no problem Jacob. Yeah, no problem. Um looking forward to to seeing you guys in uh in future sessions. Um so I go in there I have those guardrails and then that way later on even if latency or cost increases you can reframe the discussion around like hey yeah it increased but we are below the guard rails. So rather than focusing on that increase, let's let's focus on the ROI with respect to the incremental gain in customer value and the incremental gain in business impact. Um latency is just a one proxy for time, right? Like resolution time is like the real outcome like we saw here like the time from question to export. Um another question is around instrumentation. This is a lot of metrics.
64:56 Uh AI quality is a lot of metrics. Customer value is a lot of metrics. Business impact is a lot of metrics. Um the other like stuff that's not even there like more like the system metrics around like latency and cost and stuff like that. That's a lot of metrics. All that inventing eventing stuff takes a lot of upfront investment. We are gonna have um in our course an entire week dedicated to instrumentation and oper operational uh an operation observability.
65:29 um that's like the first or second week I think it's second week of the cohort and um so we'll be talking about how we set that up from the system to the AI quality to the customer value to the business impact how we can plan and design plans and work with teams to make sure we get all of that kind of eventing and tracking um ready up front that makes the rest of this analysis possible. uh unrelated our other course actually I just thought of this AI analytics for builders there's actually going to be lessons on how do you use agentic workflows to create tracking specs um so if you're doing AI evals and you want to try and automate some of the AI evals with aentic workflows take both courses actually someone yesterday signed up for both but that was pretty cool we're going to be spending a lot of time together in April with that dude all right what other questions is kind of some of the the first ones I usually that any other questions in the chat or from people in the audience?
66:33 Hi. Do you got any questions? Anything I I didn't mention. I keep hanging around here. If folks want to go to aie eval.ai or scan the QR code and take a look at the syllabus and be like, what's this week talking about? um feel free to do that or if you have any questions about the upcoming workshops or anything today. Do segments and scenarios come up frequently in AI evals? Um segments yes for sure because uh the output of LM or gentic workflow can be drastically different from one user to another. It's not just the output, it's how they respond to it too. So like their inputs are going to change, but also like something that I might be totally cool with as an output high maybe like this sucks. Like so we have um my day job, right? Like we have there's like a spectrum of like uh lawyers that we work with. Um like some are going to be like, "Yeah, this is like uh resonates well with like my style." and other one's like, "Oh, this isn't my style at all." So, segmenting by um like by user persona is definitely a big thing. We kind of talked a little bit about it when we went through the AI quality metrics and we had like the customer support question and the finance question, but segmenting is like a really important diagnostic part of it. Scenarios, could you tell me more what you mean by scenarios?
68:15 Shui, thanks for joining. I think you already bounced, but if you watch the recording, thanks for coming. Hi, did you think of any questions? I think generally I'm curious about kind of like the the audience here. What what's your biggest struggle in the whole AI eval space? I'll answer Rodian's. Uh, I don't know if I'm pronouncing your your name right, so I apologize if it's if I'm not. Um, yeah. What's the user trying to accomplish? Buy a ticket, try to get some info. Is there a definitely definitely the Okay, sweet. Um, yeah, it's kind of like um you know how Thanks, Dragon. Um, happy to help, Rod.
69:05 All right. Um, okay, Rod. Um so yeah scenarios right people have different uh expectations for different products. So for instance some products it's going to be like if you have like yeah let's let's take your example here buying a ticket. If you have an agentic workflow that like you want some want to go in there and say, "Hey, I'm trying to plan a trip to XYZ. Find me the best ticket." Um, you know, my preferences like I like to get there early or I like to roll up late or I want to go when the airport's quiet. Um, or I don't like connections or whatever. uh you kind of want that to really kind of oneshot it because or maybe like ask you a couple follow-ups but like if it takes too long or if it gives you a bunch of wrong tickets um you're just like I could have done this myself by now but then there's other things like I don't know when I interact with like chat GPT or like claude code and it kind of gets me like 60% of the like like 30% of the way there and I talk to it some are. And now I'm like 40 50 70% of the way there. I'm like, "Hey, that's pretty good. That's like 70%'s good." And so like it it definitely matters like what the use case is in terms of like how you're going to be building those eval metrics.
70:37 Some you're going to want to have like a wide spread of different error um error rates, different error rate metrics. some it's going to be really just focused on like what was that like kind of completion uh rate like um yeah that's a great question. We'll definitely talk about that in the cohort um in terms of like how do we uh think of things for like specific uh use cases. Uh there's going to be kind of like a we'll have we're also going to have like a lot of time where it's like just talking about what people are building themselves. So some of the projects will be like reflections and like your own work as well so we can get like into the spec specifics of your use case any any other stuff yeah I know hi said I don't know if there's anything that people like are really struggling with what tools are you using for annotations or analysis okay this is a very good question um so first our podcast thing We do data neighbor that I think high dropped on YouTube earlier. We're doing this other series. It's pretty cool. We're interviewing all these um CEOs of these uh AI evals tools like uh CEO of Brain Trust. Uh we're talking to like senior director of product at Arise. We talked to their head of one of their heads of product um back in August talking to CEO of trace loop compos layer. So there's a lot and I think we got a few more on the docket too. Um but I think a lot of these tools something that I have found like doing annotations like these tools are really good at like getting you the trace and like helping you instrument the data of like all the different steps within like your LM calls and like the agentic workflows. They're really good at kind of uh once you have the annotations in place and you've created like your um your different like error analysis dimensions, they're really good at helping you kind of uh iterate on prompts to like turn it into like LMS judge or um scale it a bit. Annotations thing. I think a lot of things these things are missing. the annotations from the way I see it.
73:04 In order for a subject matter expert to annotate the output of uh AI, especially when the AI output is effectively the product itself, the user interface they have to go through should really reflect the same user interface that the product has for the user. So, I've tried to build annotations through those products. A lot of it's like kind of like balms and rows in a spreadsheet. And I've also tried to do annotations in like a Google sheet before and it just it's really a lot of cognitive load for the subject matter expert. So I can tell you like a couple of examples. one is like if there's two blocks of text like the the output and then of our of our AI output and then like our kind of like golden standard output like what the user wants um in these tools like or in like a Google Sheets or something it's like just huge blocks of text and like 95% of the time the subject matter expert has to just decipher like what the difference is and only 5% of the time probably are they saying why it's good or bad. And so a very simple way to create a custom like annotation around this like is which is what's reflected like more in like products is like if you like have like red lines for like the differences just like when you're like in a Google doc and you have like the suggested edits and you have like the red lines and then the green additions. Um that's not going to mean these products. That's like a a pretty big one for annotation comparison. But what I've done um and colleagues I have done worked with is basically fivecoded custom annotation UIs. It's really really pretty easy.
74:48 We're going to do a workshop on it in March I think and we'll go into it in the course too and in the course I'm going to provide a repo that has like a few key examples of custom annotation tools. But uh you basically can make the back end pretty reliably. Like all it has to do is like read data, annotate, write data. Like those are pretty three simple things to code in the back end. And then like I don't know about front-end development, but like you can vibe code all that kind of thing uh pretty easily. It's pretty it's pretty amazing how you can make your own custom UR. Another example is like um chat or something like you just you really can't annotate chat well in these uh in these annotation tools or in a Google sheet or something because there's so many uh turns in the conversation you just can't really accurately flatten it out into a bunch of like rows and stuff. So that's another really good one where you can like vibe code a custom annotation UI where your subject matter expert can come in and actually be chatting with the product as with just like a user would with the product but then on the side have like an annotation um kind of box where they can like say like oh this response wasn't quite right and if I connect back up to this response earlier you can see it's like diverges from the conversation. So that's my take.
76:19 Yeah. And hi, drop the link in the chat there to our our uh upcoming session on building custom annotation tools. That answer your question, Alexander? Cool. Cool. I can stay for like 14 more minutes and then I got to hop. But any other questions? Yeah, I think Kai had a good one around and questions for the audience like it's like what's the biggest challenge right now?
76:59 Trying to think of what the biggest challenge for me right now is finding the time. I think >> it's like just moving so fast like uh building product is like feels like a lot faster and cheaper now with all like the productivity tools but then like it's like you got to find find that time to set up the entire evaluation workflow but it's really in worth the investment when in my day job like I probably spent four three or four months uh iterating on our AI valves like pipeline testing out different quality metrics relating them back to customer value and business outcomes before I could get all those weights and then really figure out the suite. And it honestly wasn't until that was all set up and did all like that kind of like early investment that I was able to uh be like, okay, now I can actually analyze the data and it's like, oh, it's not working for this segment or or that segment or like we need to get to this goal. Um, but once you have that all set up, it's a it's pretty cool.
78:09 Cool. All right, I'm going to give it two more minutes. Oh, here we go. Here we go. Stay along. Good. Figure out what are the bottlenecks in a uh Yeah. Yeah. Sweet. Dom. Awesome. Hope to see you in some of the other the other ones. Um, Rod, yeah. Uh, common bottlenecks. I mean, instrumentation's a big bottleneck. Um, for sure. Instrumentation and tracking is a big bottleneck. Um, I think agreeing on like what customer value means is a big bottleneck. There's a lot of the custom annotation tools a big bottleneck too actually.
79:04 Yeah, annotation tooling and instrumentation I think are like the two biggest ones. Just getting started also is like a big hard thing to do. It's like the first steps kind of hard. But that's a good question. I'll think about that a little more thoughtfully. if you want to DM me on LinkedIn or something um or email me just to remind me and I can like think about it a bit more.
79:38 Nice. And then there's Oh, nice. Yeah, that that link that uh put in there around redesigning product mix for AI valves. It's kind of like I'd say it's kind of like the higher level version of this. We went into a little more details here. Where do AIPMs hang out? Where do they hang out? Oh, yeah. This is good. Can't find any communities. This workshop's awesome, by the way. Um, okay. Where did they where did any IPMS hang out, man?
80:12 Yeah, somewhere at Lenny's. I'm trying to think of like communities to talk to people. I mean, hi. I feel like you've met a lot of people going to kind of different um conferences and dinners and stuff in San Francisco that we brought onto the podcast. I live in Tahoe. I don't hang out with anyone, man. I hang out with the the Bears and the Coyotes. >> They're everywhere here in San Francisco. If uh if anyone want to drop by, uh I can intro you guys. So but but but just uh uh you know c certainly it's geospaccific I would say uh there's a lot of in-person events if you're in the proximity in the Silicon Valley I would say outside of that it's probably a hit or miss.
80:57 um Slack groups forums, you know, Substack, there's a lot of AIPM's writing on Substack. If you basically go to Lenny's and go to go watch Lenny's podcast, see their guests or go to Maven and look up the people on Maven who are doing like AI product management work, and then you just like look up those people's names on Substack. They all have like pretty awesome newsletters and I think there's like some good community in the comments there and a lot of them are like super awesome respon like like we've had a bunch on our podcast before like super responsive like yeah I'll talk to you.
81:33 Um I guess that'd be my advice. I don't have like a link to a Slack group but we should I mean we should create one for this maybe. Uh, we'll maybe we'll create like a Slack group or like a a Discord or something and I can send that out. All right, I'll do that next week. Yeah, I'll do that. I'll make a Slack group. I'll email everyone who came to this with it. That'll be cool. >> I think it would be also really cool to perhaps in one of the future either lessons or incorporate in the in the course to bring in some AIBMs.
82:10 >> Yeah, that's a good idea. Okay, we did we've had some pretty good ones on our podcast that we could probably bring in. Um, and some other folks we've just talked to that that have been kind of helping us with uh with this. Yeah, trying to think of some people like Alman Con's really good to follow. Colin Matthews, those were two we had in the pod that are like super responsive and they have really good Maven courses.
82:37 Yeah, we'll try to bring someone in for either a lightning lesson or indefinitely in the full cohort. But Aman Khan, it's like head of product at Arise. It's like that AI evals platform. Yeah, that dude's pretty cool. He he did a Lenny's episode, too. It's really good. Definitely worth a listen. Um, all right. Hang out a bit more. Got seven minutes. Nice. Already following Perfectly. Perfect.
83:11 What else? Other questions? What are you guys up to this weekend? I'm going to try that pizza thing I told you about. Been like um going to try that peel thing I eventually got. Try and make some pizzas in the pizza oven for some friends. What else should we talk about? Ski class. Oh, hell yeah. Where'd you say you're from again, Alexander? Austria. Man, I've never skied in Europe. We have Our snow is so rough right now. We It hasn't snowed for three weeks in getting out, but um Oh, you're just watching. Okay. Okay.
83:59 Um Okay. Rod has some AI val thing. Yeah, I said I could talk about ski and snowboarding for for all day, but um for for Rod's questioning around uh structured output in your work or chat or combined. Yeah, so in my work I work more it's like my my day job it's more like um it's more structured output I would say. We do have um yeah, thanks. Hi, I'll see you later. Um we have chat product too. So um I would say the structured output really lends it well stuff well to like uh that completion rate thing where you have like a scale of 0 to one. How close did I get to what the user actually wanted or to acceptance rates? And then I would say the chat stuff is a lot more of that error analysis. you really do have to do a lot of that with the chat kind of stuff. Anything in multi-turn. Um, same thing with like aentic workflows too because you have to really look at like all the different steps that are going on and kind of like check out errors between all those. Uh, the structured output it stuff is easier because you can kind of go like all right, what was the input? What's the output? What's the ground truth? And you'll want to have those other metrics also, but they're more like diagnostic to when you see like, oh my my like cur my completion rate is quite low. Can I figure out why it is? Um yeah, something like chat or like anything multi-turn or multi-step, you're going to want more of those kind of error analysis once. Uh we'll go over like yeah, each of those though in the cohort for sure.
85:47 Great question. Oh, I haven't tried this. boundaryml.com. I'm just pulling it up on my other screen. Ah, first language for building agents. Okay, I want to check this out. Thanks for sharing this. I haven't tried it. Okay, I think I'm going to call it, guys, because my dog is barking like crazy out there.
86:22 I don't know if you can hear her, and I have another meeting in four minutes, so I got to figure out why she's barking, but this was pretty fun. And you guys had some great questions. Um, yeah, I'll send out the recording. Check out aevel.ai for the course. Check out data neighbor.com for all the other free workshops. Um, hit me up with an email if you got any questions or hit me up on LinkedIn. And uh, yeah, this was fun.
86:51 Hope you all got something out of it. Talk to you later. Thanks everyone. Oh, email. Uh, let me drop it in the chat right now. It's Shane Sean spelled Shane Ai.ai. AI. Here you go. Cool. I'll stay on for 10 seconds in case anyone needs to copy the email. 10 9 8 7 6 5 4 3 2 1. Cool. All right.
87:27 Thanks everyone. Appreciate it. Have a great weekend. Hope to talk to you soon. Bye.
Summary
- The AI impact chain consists of four links: feature change, AI quality, customer value, and business impact.
- Many companies struggle to connect AI quality metrics to tangible business outcomes, often shipping products without adequate measurement.
- Validating AI features requires understanding customer value, which can be assessed through metrics like task success rate and time to complete tasks.
- Causal inference methods, including A/B testing, are essential for establishing the relationship between AI quality and business impact.
- Common challenges in AI evaluations include instrumentation, defining customer value, and the need for effective annotation tools.
- The session highlighted the importance of segmenting user data to understand different user experiences and outcomes.
- Upcoming workshops will cover topics such as designing experiments for AI features and causal inference analysis.
- The presenters encouraged participants to engage in discussions and share their experiences to foster a collaborative learning environment.