transcribe

Validate AI Analytics Output in Claude Code with AI

AI Analyst Lab · 1h 5m · transcribed Jun 2026
More from AI Analyst Lab Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:08 Hey, welcome everyone. >> Good morning. >> Welcome, welcome. Good morning, good afternoon, good evening, wherever you are from. >> Um I am just turning the waiting room off right now. There you go. That's off. Um Hey everyone, I'm Sean. Uh joined here uh with Hi and Sravya. Um We'll do intros in a quick second. I think a lot of you have been to our I see some familiar faces in here already.

0:41 So I think a lot of you have already been to many of our sessions. Uh but just to get it started, maybe drop in the chat uh where you're joining from. Want to get an idea of, you know, your city, your country, wherever you are. Are you in the living room? Are you in the office? Are you in your home office? What are your coordinates exactly? Uh I'm located in South Lake Tahoe, California. Nice. >> Ireland.

1:10 Boston, Toronto, a lot of East Coast folks. Ukraine. >> We could be biasing the data by keeping it at 8:00 a.m. PST. You want >> Yeah, actually we we moved our we moved our time zones cuz I feel like uh we were doing a lot of our our uh our things at like noon Pacific time, but that that misses a lot of time zones. >> But look at that, there are Santa Clara, there is also Seattle. Not bad, West Coast. That's awesome for a 8:00 a.m.

1:39 meeting. >> Nice, bright and early. >> Feel us from Nairobi. Very cool. >> Sweet. Um while those roll in, we can do some quick intros. I'm Sean, one of the co-founders of the AI Analyst Lab. Um we do kind of research and education around the field of AI analytic analytics. Uh been in working in data science for a little over a decade.

2:10 Uh past couple years more in the kind of AI evals uh AI analytics space. So, that'll be the topic of today's which we'll get into a little bit. I'll pass off to Hai to do a quick intro. >> Hey everyone. My name is Hai. I've been in the data science and analytics space for around 20 years. Um uh most recently at uh consumer tech companies spent time at Nextdoor. That's where I met uh these folks. Uh co-founder of AI Analyst Lab. We provide a lot of these workshops like these um both free and free and paid. So, very excited to uh to have you all here.

2:50 Over to you, Shravya. >> Hello everyone. I'm Shravya Madipalli. Uh I am one of the co-founders along with Sean and Hai. Uh around 15 years of experience starting with Microsoft. I see some of the Seattle folks. I was there for 6 years. Later moved on to eBay, Nextdoor, and most recently Grammarly which was rebranded as Superhuman. Yeah. So, very excited to share everything that we're doing in the world of AI guys. Agentic analytics um and you know uh it's it's pretty awesome to be part of uh in this domain. Yeah. Over to you, Sean.

3:23 >> Cool. So, we'll get right into it. I know everyone's got busy is busy got a lot of stuff to do if you can't stay for the whole session. We'll send out a recording. Um we'll send out an email afterwards with some of the resources we're going to go through today. We've got some free resources for you. And then we'll send out the recording um I don't know probably probably in a couple days. By Friday at least. Um So, where we headed today?

3:45 So, a lot of the I'm sure a lot of the if you've seen our um kind of work before if you've been messing around with AI and on your own and you're working in the data field, you've probably seen that AI like it can now kind of run an end-to-end analysis in minutes, really. It can write the SQL, it can run the SQL, it can make a chart, it can reason through and make some recommendations, and it can be really polished and confident and also sometimes completely wrong. So, getting the answer kind of has stopped being the hard part, and then now like knowing whether you can even trust the answer is where a lot of the work lies now. Um so, that's what today's about.

4:35 I'll show you uh there's a whole spectrum of ways you can kind of check the work. Um there's ways to do it manually, there's ways to do it completely automated fashion, and then there's there's space in between. We're just going to go through one check live today because we don't have that much time. We have a whole We have multiple courses on this in our first in our 101 boot camp around building AI analyst. Uh it's like building a genetic analytics 101. We go through uh five of these checks three live, two in sync, and then we have an entire uh advanced 201 genetic analytics course where we primarily just focus on on really AI validation and uh context engineering, which happens to be a lot of the ways that you uh increase uh the trustworthiness of your output.

5:18 So, we'll hop into it. I do want to get a couple of reads of the room before we start. So, first one, drop a letter in the chat here uh just to see where folks are. Um A if you've never used Cloud Code before. We're going to be doing a little demo here with Cloud Code. Um B if you have used it but not for analytics. Um and C if you've used it and you've used it for analytics. So, yeah, go ahead in there. Hi Anshravia.

5:45 >> Mostly C's. Wow. >> And then some B's. >> It's good for us to get a pulse, I think, also just on like where things are. If, you know, we asked this we asked this question in almost every single one of our lessons, and we asked it 4 months ago. I'd say everyone was uh A or B. But a lot of people are using this for uh analytics now, which is really awesome, and also just makes it so important to be like, "Hey, now do I trust this thing? How do I set up eval?" I've kind of done the step one. I use it for analytics now. We probably need a D on here of like, "Are you Do you trust when you use it for analytics?"

6:23 Um Okay, good. That's awesome. Let's do this before. If you're an A or B, that's great, too. Um Next question, also just to get a kind of spread of the room. What's your role? And you don't need to write A B C D here, but you know, what's your role? Are you a product manager? Are you Are you a founder? Are you in marketing? Are you uh engineer? Are you designer? Are you in data? Um I just like seeing the chef one. We had the corporate chef the other day, and I'm not That's awesome.

6:55 And then, you know what? I You know what? I started using Claude for cooking a little bit. I've been telling you guys I've been using it for like my fitness and nutrition tracking. And then yesterday it was like, "You had too many fatty foods earlier. You have to have this many grams of carbs." And I was like, "All right, tell me how to do it." And then it like made me this delicious burrito bowl recipe. Tell me how to cook chicken in a little different way, way too.

7:16 Uh which is not the topic of this, but if you guys want to know uh chicken Mexican uh uh high carb uh burrito bowl recipe, where you can cook the chicken from frozen. Doesn't even need to be thawed. Worked pretty well. Took 30 minutes. Thanks, Claude. Lot of founders in here. Chef turned data engineer, yep. Back to chef. Cool. All right. Thanks everyone. Lots of people on the data teams there.

7:47 Um good to get a read of the room who's in here. So, to start off, uh what we're going to even What are we even going to talk about? Uh so, what what is this like AI analytics thing? Cuz this is like a a newer term and it's a bit vague. So, um when I think about it at its like most high level, I would say like agent agentic analytics means something like you ask a question and um says the question in natural language.

8:13 All right? Natural language question, um something like why is checkout dropping? And you have a system that goes and writes the SQL, runs it against your data, builds a chart, hands you back a recommendation. There's no like human in between sitting in mid in the middle of doing it. Uh just kind of does the whole thing. So, input a question, you add with that question, you add uh you know, your data, whatever context you've given it, and then the output is a number and you know, a little story about what happened.

8:47 And usually there's a here's what you should do about it as well. So, a year ago, I think, you know, I've seen a lot a lot of demos on this stuff. We we have run a podcast called Data Neighbor podcast, um where those demos got increasingly better and better and better. You know, about a year ago, I don't know, it kind of seemed like a little bit of like a party trick, like what's really going on in that demo? Is this Is this going to Is this going to work on my data? And now people are really actually shipping real uh decisions off the back of it. So, I think it's it's moved a long way, especially in the past, I'd say, 4 to 6 months. Um so, that's that's the shift and that's why you know, as it scales and as more people are using it for real decision making, it's why that kind of like checking the work part of it suddenly matters a lot more um than it used to.

9:37 Uh so, what makes this different from a person doing the analysis is simple. Uh, when it's wrong, doesn't look wrong. So, you know, you get a really clean-looking number, a clean story, like every single time, whether the answer is right or completely off. There's no like, "Hmm, let me double-check that." Um, and it hands you it basically hands you that wrong uh answer with exactly the same confidence it hands you the right answer, right? So, and that costs you that costs you I think in in a couple of ways, few ways, probably more than this, but the main ways I think it costs you um, is that I've seen happen a few times already. The first is trust. So, if you put a wrong confidently wrong number and you like give it to the board or a leader or even just a teammate, um, you don't really have a lot of that trust built already or you don't have like a kind of like understanding of we're in like experimental mode right now. When somebody catches it, now nobody nobody believes that tool again, right? Nobody believes that approach again. Nobody like maybe people don't like trust you or um, that's maybe that's fair, maybe that's not, but there's a real like uh like trust is like I I highlights to say like trust is like really hard to build and then really easy to lose. So, you don't want to lose trust. That's that's one one thing that happens when it goes wrong.

11:03 And then the second, um, I think is worse, is this like, "Oh, so people all trusted it. It was wrong and then we acted on it." So, we like ship a thing uh that moved the number instead of the thing that actually helped the customer and you you don't find out for like a quarter or something and it just depend it comes back to this uh this system that gave you a wrong answer. So, you can make real wrong product and business decisions off of it, right?

11:27 >> [snorts] >> Um, okay. Let's do another kind let's do a little thinking exercise here. So, drop in the chat. It doesn't have to be one of the things up here, either. Where do you think AI analysis goes wrong? Is it like the sequel? Is it picking the wrong metric? Is it like the recommendation at the end? Or is it something else? Feel free to throw a guess in the chat. Definitions and gotchas.

11:58 Hallucinating data. Wrong context to interpret analysis. Incomplete context. Wrong metric. Everywhere. Context and nuance. Oh, oh, Brian said it almost happened to him today. Yeah, you Yeah, I mean, like knowing about the business is a real That's a real check. Bad storytelling. Company history is incomplete. Cleaning analysis Yeah, these are all really, really good. >> [laughter] >> Probably better to go where it goes right today since the slice is thin.

12:33 Yeah, like that. >> Cool. Yeah, so you can see it goes wrong all over the place. It can go wrong at any of those points, uh which is exactly kind of where I'm headed next here. So, uh This You'll hear this word, um AI eval evals, um come in. So, I'll just I'll give you like a pretty like, I don't know, my like highest-level definition of it. Uh cuz it sounds like a little technical one, but it doesn't really. Uh you know, an eval, it's just a check that tells you how much to trust an answer that you didn't work out yourself. So, the system handed you a number uh or handed you, as some alluded to in the chat here, a story. Um and, you know, you didn't do the math.

13:18 So, an eval's is going to help you figure out how far you can actually lean into that number, how much you can actually trust it. Um Getting the answer, like I said before, it just keeps getting easier and faster. So, knowing how much to trust it, like that's actually going to get easier and faster, too, but that's just like where the work has moved. So, good thing I think about the shape of this, too, it's not just a yes or no stamp that says this answer is true or this answer is false. Like, sometimes we'll have that um and in cases where we actually have that, then it's kind of like you don't even really need the system cuz you're probably pulling from some some other place.

13:57 Um it's more about like how much on a scale that you trust it. Um and you can match that to the stakes of the decision you're going to try to make with that number. So, like you don't have to go and do some crazy rigorous uh eval uh pipeline for every single number you do. If it's a throwaway question that someone's asking just cuz they're curious and they just want some directional understanding, um you can kind of like do some pretty quick, barely barely check that number. A number that you're going to like report to the street, report to the board, or that like a big decision leans on, like a real decision, um you can't really check it like you mean it. So, so trust is like I think it's as like a dial, not a switch.

14:46 Um and then we'll come back to this the to at the end, but like I said before, that this whole discipline of evals for analytics is what our uh Gigantic Analytics 201 um boot camp is all about and that is coming up in a week and a half or I think. It's a week-long boot camp on validation. So, um if we think back to that picture from the start of the input and the output, um question goes in, recommendation comes out, it's easy to think of that picture as as the middle as just like one box that kind of like does the thing. But as we've seen just in the chat with everything that you all just said around like, you know, the different ways this can go wrong, clearly it's not just like one little box in the between that like turns your input into an output. There's a whole pipeline in there. It interprets what you asked, it picks which tables, it picks uh the definition of a metric, it writes the SQL, it runs it, it reads the result, it decides what's even worth saying, and then it writes a recommendation. So that's that's like seven steps right there, and it's probably most definitely not exhaustive.

15:54 Each one of it ha- each one of those steps has multiple places where they can totally go off the tracks. Um and it really matters when you think about this in terms of a pipeline, too, because helps in terms of like prioritization of what you're going to where you're going to want to spend your time in evaluation. An error that's really early in this pipeline, like, you know, like query generation, question intake, analysis, um that's going to carry through to everything after it. So if it grabs the wrong definition of an active user back in step two, every chart, every confident sentence after that is just sitting on a bad number.

16:30 And they're all still going to look clean, right? So um is it right isn't really one question, it's more like many smaller questions um at each of these steps. So uh I just dropped some of the checks, again, not exhaustive, but like some checks we do at each of these steps. So at each one of these steps, there's something you can actually look at. Like uh did it understand the question you asked? Is the SQL correct? Uh is it using the metric the way you mean it? Uh did it pi- did did it pick a like robust, defensible method methodology of analysis?

17:08 Does the chart match the number underneath it? Um every step has its own thing to check. Um and you can keep it pretty manageable cuz each of these checks is really just kind of like one of four kinds and I'll show those in a second. Um but the darker ones here if if I could look at the shading it's like the darker ones are kind of like where I'd I'd spend more of my my time rather than kind of these these ones are like easier fixes.

17:38 Um so there's a few checks also that don't sit in any single step. They run across like the whole um kind of surface area. Um Sorry, I'm just reading the chat here. Seven steps, nine questions. Is slide generator failing? Yeah, probably. What? Oh yeah, the nine questions are going to be like all these checks over here.

18:10 Um there's like so there's multiple questions under under each of these. Um so uh these these checks that don't sit in any steps they they run across the whole thing. So it's like do you get the same answer if you run it through more than one model? So reliability. Can you trace every number back to the exact query that produced it? Um re- re- reliability which is just asking the same question again and again to see if it holds for running underneath everything here.

18:42 Uh you put those per step checks together with these two that are cross cutting across everything and that's kind of our I don't know if you get that you get a pretty robust system not exhaustive a pretty robust robust system of evaluating an analytic analytics pipeline end to end. Um and then so so I guess like from this slide the main thing I want to take take away here is like eval is not a single score.

19:09 It's like many checks at each of these steps plus some that run through the whole kind of gambit. So, it's a lot. So, the obvious question um then is, how do I actually run any of them? Which ones do I run? Um how do I To what level of rigor do I run each of these depending on my specific uh kind of use case? So, you don't you don't really need a different trick for all nine of these um checks down here.

19:42 Um there's really kind of just like a handful of ways to check any output. So, here's four of them. One, check it against a known answer. You've probably heard of this as like ground truth. Not always going to have it. But, if you happen to have a known answer, that's one way to check. Two, grade it against a rubric. Just kind of like, you know, you'd grade an essay, for instance. Um and humans can use those uh rubrics to grade. LLMs can use it. You wouldn't probably start with a human, multiple iterations to make sure that uh rubric is really sound. Uh and then you can all then scale the grading up with like an LLM as a judge.

20:25 Um Three, ask it several different ways and see if the answers line up. And then four, make sure you just show the receipts. The last one's like a very uh low uh level lift and goes a wrong long way around how much you can trust an answer. So, you know, show the actual SQL uh where the actual number came from. This is something we should all probably be doing today in our analysis regardless of AI, right?

20:53 Yeah, in terms of do you models agree, we're actually going to do that uh today. So, got to We're going to go through reliability today. We're going to go just within the same like Opus 4.8 model, but we'll have like a reliability check where you run a few sub-agents on the same question. Um but uh one one thing we've been doing a lot of is like run the same analysis and run it against like um different models like within like the Open family itself, 4.6, 4.8, 4.7, run it against um Codex models, run it against open source models, and see if you can triangulate the same uh answer.

21:36 If you can't, then there's something going on there, right? There could still be something going on there if they all do agree. If they all make the same mistake, um you'd have to have that built baked into your system, but if you're looking at totally different models, um it's a little less uh likely that that would happen. And then there's a few gut checks you can run on any analysis. It's like, did it even run? Did it use a defensible approach? Uh Does it give the same answer if you ask again?

22:05 So maybe I don't know. Like if people are already kind of like So yeah, like actually Brian Brian mentioned it today um like uh he kind of had a uh close call with uh some wrong numbers, but he knew the business pretty well. Might have not been the exact known answer. Like the known answer doesn't have to be like X equals X. It can be like, I know this space and I think about it like this number seems way off.

22:36 Um But which ones of these you already do? And you can add if there's other ones in here, too. That's fine, too. But known answer rubric, asking multiple ways, showing the receipts. How are people kind of um you know, even informally, right? Um Yeah, let's see. Good I mean, good kind of spread against everything. A lot of people are doing four. That's great. A lot of people are doing one. One, sometimes you don't have. Four is really good.

23:08 Low effort ask several ways. Nice. Yeah, one's typical. Yeah, anything novel uh you haven't checked before, you're like you're not going to have one for it. We're going to do a um lightning lesson uh next week called pressure testing AI and uh analysis. Um I'll drop the or if Hi Arshavi, I want to drop the link in there. It's another free workshop. We're going to talk about like what you do when you what you do when you don't have a known answer.

23:44 Um and then I'll have like a QR code for that later on in the presentation, too. >> What's the lesson next week, Sean? >> Yeah, yeah, I think it's like next Wednesday. I think it's the same time slot. >> [snorts] >> It's called like pressure testing AI analysis or something. Um All right, so two caveats. So yeah, go ahead and sign up for for that if you're interested in figuring out like what you do when you don't have ground truths.

24:12 Uh so two caveats here cuz I don't want to oversell all this stuff, right? The first one is about agreement. So asking several ways and getting the same answer back is reassurance. It's not proof the answer's correct. So if the system is wrong for the same reason every time uh it's just going to agree with itself every time and tell you the wrong answer. It's going to you know, tell you nothing. So actually I find the more useful signal when like we're building out agentic analytic systems is is actually disagreement. Because that gives me a direction to like dig into and improve uh the system. So it tells me like what I can fix. Um and we'll get a little bit into that today. Uh in the demo. Uh and then the second caveat is that yeah, you don't run all of this on everything. Like match the effort of your evaluation to the stakes, just like you would do with anything in your job. So, if it's a throwaway question, like whatever, you know, you don't need to go super in-depth on on everything here for a throwaway kind of I'm curious question from someone important where you suddenly have an answer.

25:26 Um and then yeah, but never going to like I said before, like something that's like, you know, publicly reported or you're making a big decision off of or it's going to the board, like yeah, you probably want to give that kind of more robust treatment to that. Um I'm going to give you a sheet later on that'll kind of like help you try and yield a like based on uh where you're at with your company and your decision and what's available with your ground truth.

25:50 Um where what which one of these how how robust robust you want to go. Um like a lot of folks who we've talked to are on leaner teams, right? Like they don't have the resources to go through the full gambit. So, there's a smaller version of this, like the doing the cheaper checks. Um that can make you have a much more better understanding of how much you can trust this, but without having to like spend a ton of cycles on it.

26:22 So, this is our kind of current map of what you might check across the pipeline. It's not exhaustive. It's just kind of where our thinking is right now, but you can see like we said before, there's a lot of checks, whether the SQL's right, whether it picked the defensive method, whether the number of cancels. Today, we're just going to focus on one of them, um and that's reliability. Um and then the you know, this is thing we teach hands-on in the boot camp of the the intro boot camp, you build your own genetic analyst and you build these checks into it. There's also a whole 201 course built purely on this validation and then how to improve it, which turns out to be a lot of context engineering.

27:04 Um we'll get back to this later. I'm also going to send you out a email with some PDFs worksheets around the frameworks we talked about today. Okay, so let me move over to Claude Code. The check we're doing today is that simplest one, reliability, and it's the one you can run no matter how messy your setup is, because it doesn't need an answer key at all. The whole idea is just ask the same question more than once.

27:38 See what holds and see what moves. Let me switch over there. And can everyone see Claude Code? Should be on my desktop here. If it's hard for you to see, you can kind of zoom in via Zoom. >> [snorts] >> Nice. Seems like people can see it.

28:13 Okay, so All right, actually before we start, let's do a kind of prediction. I think I got a slide on this. All right, so same question. I'm going to run the same question. I'm going to run it five times. Do you think we'll get the same answer all five times? Yes or no? I'm going to ask a couple questions. Uh the first one I'm going to ask is around uh this is we're going to look at like a synthetic uh e-commerce website. You can kind of think Amazon.

28:50 Um, I'm going to ask it, how are we doing on checkout conversions? So, given that question, what are your thoughts? Can drop in the chat. Lots of no's. There's a There's pretty low confidence in this system. Pretty low confidence in this system. Yeah, Logan. Logan's like, hell yeah, genetic analytics is the future. Same answer. Okay. Let's try it out. All right. So, there's actually I've created a skill here. Um, and I don't think this skill is pushed yet, but you are all in luck. We haven't actually formally released this yet, but we have two versions of our AI analyst.

29:34 One's AI analyst, one's plus. The plus one's AI analyst open source hasn't really been developed on since like March. And then we have a plus version, which we've been constantly developing on the whole time. That's actually um, available right now to the public. We're just going to move all that stuff into like a V3 or V2 of the I guess V2 of the AI analyst repo, if you already have that on your machine, but maybe um, maybe uh, Hiya Shravya, you can drop the AI analyst plus repo in here.

30:08 As this skill, the reliability one, I don't think is pushed there. Uh, we're working I haven't brought in just for the purpose of this demo, but we have a separate repo that I'll open source that's going to be around like all the different like uh, agentic analytics eval skills. >> [clears throat] >> Okay. So, and then I'm using uh, VS code here, Jennifer.

30:38 Uh, all right. So, basically what this skill does um it fires you ask it a question you tell it like how many uh like uh cycles you want to run on that question. It fires that and independent runs. Um it collects all that uh data and and stores it. And then it um compares it calculates the mean, variance, etc. And then it does a little report that says like, "Hey, is it stable?

31:15 Um is it drift? And how could we improve it?" So, we're going to run What did I say? How are we doing on checkout conversion? I think I have this prompt somewhere. Here it is. Okay, so you're saying I'm going to invoke the reliability skill and say, "How are we doing on checkout conversion?" Run it five times.

31:54 So, it should kick off five separate processes once it gets going. Yeah, so it spins up five independent sub agents. Measures whether answers agree. Launches those concurrently. Eat those tokens. So, you can see it's all running here. And what it's actually doing, we have our data in a Snowflake data warehouse here. And so, each of these agents is going to start hitting Snowflake on their own. We actually have a something we've set up so we can always go back and see the work and like audit the work that's been done. We have a hook here where anytime there's like a a tool call to the Snowflake MCP query, uh we have it record the work for us.

32:55 And I think that's down in the working while this runs, I can show it. Oh man, my computer is slow. Where is working? Just saw it. If we close this down.

33:32 Let me close some other tabs. Okay, so it's returned those. Oh yeah, here it is. Query log. Come on, open up. It'll give me the answer here, but seems like it's going crazy. I wanted to show you guys the query log.

34:17 Well, let's scroll up here. What is the wrong thing? Um all right, so it ran uh five times and I don't know why my computer is going so slow, but you can see it actually came back with uh the same answer for this one every single time. So, it came back with uh 33.2% and uh you can see it actually lists the source here.

34:58 And the reason it does this is because we have a uh metric dictionary that tells Claude what we mean when we say conversion rate. So, if we go to Okay, now we're now we're back at like regular speed here. So, if we go back if we go into our knowledge base and we go to data sets and over market metrics index um there is you know, this is pretty empty, right? Cuz I just added this this morning, but um there is uh I'm going to have to say something.

35:42 Try it again, but uh there is a clear new definition is. And so, if we think about this, it's actually hiding something here. So, this ran um the reason it didn't have variance was cuz we had that in the dictionary, but it's actually hiding um it's hiding something in terms of like, does that is that the definition we actually uh need? So, check in check out conversion actually probably has like several actual readings behind it that could all be defensible like, uh you can count from the moment someone starts checkout to when they pay. Um that's the one that's in our dictionary.

36:21 You can also count every session that ever ever landed. You could count uh cart to check out or check out to payment step or you can count people instead of sessions. Those are all like real definitions and they're all they'd all end up being totally different numbers on this exact same data. So our dictionary where it pins to one of them which is why it came back stable. But if we were in like a an actual meeting, like that definition might not be what our stakeholders were were looking at or wanted to look at. So sometimes I I think it's more helpful to see the divergence first because checkout conversion in my head could have meant something totally different than checkout conversion and say hi or Arabia's head. Reliability just told me the system is consistent with one of those definitions.

37:15 Um but it didn't necessarily say that like that was the right definition. So stable stability is necessary but it's not totally sufficient. Um so we're going to do the same setup with a different question now and this other question is not in our uh metric dictionary here. So we should get uh variance across runs I would think. We're going to ask let's see. We're going to ask about retention rate.

37:51 So for reliability what's our retention rate? And so it's going through the same process again. It's going to kick off five sub agents to run the same question.

38:29 Here's what I was trying to show you before. So, when it runs the SQL, we have it record. That hook has it record the SQL it runs every time. So, you can see um it has like the timestamp. Um this is just an ad hoc analysis. If it's part of like a project was working on, I would tag with that. And then it has the purpose. And then it has the actual uh SQL that was ran. Um so, you can see there's a bunch here on checkout conversion, right?

39:04 Um because you have multiple sub-agents running it at once hitting Snowflake. This is really nice for later on. Um I mean like if I got a bunch of different um answers or if I even got the same answer, um one of the main things that I would have, which is part of like doing uh having this validation is just like showing the receipts, right? And so, I could actually have the exact SQL pointed to it or I could just talk to Claude to and so you can see it's it's just populating more SQL runs as they come in now, but not for retention.

39:40 >> [clears throat] >> But having that kind of like audit um trail is really nice for uh just understanding how much you can trust it and know what is going on. You could hit like control O or something and see what it's running in the background while it's going to check it as well. You could ask Claude, but like I do have concerns about it just like hallucinating. What I've seen is if I don't have this recorded somewhere through a hook where it's actually writing down the exact query it's running. I could also have it go like look at Snowflake query history or something.

40:17 Um then it takes a guess at what it did before looking at its previous context. If you clear out the context, it's not really going to know. Um so like having like a solid audit trail. All right. So, all five runs returned. This time they're staggered. It's going to record um the results for me in a uh JSON with all the runs and it's going to do the with the calculations. And so, you can see this time, like last time, some of the sources were actually looking at a uh metric definition catalog.

40:54 And then this time it basically just says like the source was it made it up on the fly. So, uh these two get pretty close, right? 99.3, 99.2. Um they're like active in month and month plus one. This is active in month after sign up. Sign up core grain. The first one's 9.7% 30-day repeat purchases. Second completed order within 30 days. The first one we're we're asking about retention. Uh this definition 34.9 repeat purchase rate. Purchasers with two plus orders divided by all purchases over the full year.

41:34 78% week one new users that have more than one session in the first 7 days after sign up. Uh and then we talked about these ones. These are like of the month one cohort. How many were active in the month after the sign up month? This one's pretty similar. But these are all pretty solid definitions of retention. Like if you've been in a met in a meeting where you're trying to pick metrics, people you could easily have like four people in the room kind of arguing these different ones cuz they all tell they all help you in different ways. Um conversion. Even if I hadn't had a metric dish definition for like cart conversion, that's a little more cut and dry. I bet we would have got like three out of five matching at least, but it's not that much of a surprise to me where it's like, man, these are just all over the place. So like and like imagine how this shows up in a uh a deck. Imagine if if you have a leader that says, I want to know what retention that just asks that like very like kind of vague, high-level question, but like very like natural, what's our retention rate? And like you could tell them 9.7% or 99.3%.

42:53 Um so the way to to like this tells us like, what am I going to do with this? Well, I'm probably going to go back to my team. I'm going to um like define retention rate with them. This is actually kind of nice. It like gives me some options for retention. Maybe they don't want any of these. Um but it just tells me where to dig, right? And this is clearly a metric definition problem. And then uh to uh fix it, if I wanted everything to stabilize, I would have to get the agreed on definition and then I would just uh add that to like some sort of context somewhere. There's many ways you can manage context. One way could be something like this where you have like a YAML file where you store all your metrics definitions.

43:41 But it's pretty crazy to see the spread. And like that would kill trust. If someone if someone was thinking the uh repeat purchase rate purchasers with two plus orders or all purchases. If your leader was thinking in that in their mind of like what retention is, and then you told them 99.3%, they'd be like, what the hell are you talking about? I'm not I don't believe anything Sean says. >> Okay, so this was just like one of those boxes.

44:15 If we zoom back out, um one check reliability, one box on a big map. Um so notice what it did and didn't do. It caught that retention wasn't stable. That's a big warning. It didn't on its own tell me which reading was the right one. Um I would have to supply that definition. The other boxes on the map cover the rest of it, the receipts, the approach, ground truth, uh where you happen to have it if you have ground truth. And how hard you push on any of this comes from comes back to like your use case and uh what the stakes are. If you're ideally curious, one runs quite preny. If your numbers are driving real decision, then you want to check this stuff like more robustly.

44:58 I just Okay, oh yeah, before we'll go over to questions in a bit here. Um let me grab you real quick. I promise. I did just I posted this right before this session. I think I have it up here actually. There's this. Uh I'll send this in the email, too. But uh this kind of worksheet is like a little bit of a self-assessment. Um helps you figure out which of your analyses actually need this kind of checking, what you've realistically got to check them against.

45:31 You know, do this one first. Tells you where to spend your attention before you spend it. I'll drop a link to this post on it if you want to download the PDF there. But I'll also send this uh in uh an email later today. And then I have another worksheet that I'll share. Let me see if I can pull it up here.

46:02 Um yes. This one gets more into kind of like the different ways. We kind of went over these, right? But the different ways to check anything gut check wise, and then what you're going to do based on your kind of uh low stakes to high stakes effort. So, I'll share both of these in the email, but you can grab that other one from that link right now if you want to check that out in the meantime.

46:33 Um, and then we'll also send a recording out. I think we got a couple of discounts codes. We can we can share those in email, I think. I'll go over it kind of like what's coming up. So, Robbie, feel free to drop the those discounts in the the chat, too. But a few things coming up. And then we'll get into questions. I can stay a little over for questions. I know we got about 12 minutes left. Um, if you actually want to build this, um, here's where here's where you go.

46:59 Every one of these has a QR code around the screen, so you can just pick up your phone right now and kind of click which one you want. Top of the list on the left, this one's free. This is what I mentioned earlier. It's the next lightning lesson, pressure test any AI analysis on Wednesday, June 24th, 8:00 a.m. Pacific. That's going to tell you um how to tell if an analysis actually right when there's not really an answer key to check against. Um, if you scan that, you can register for free right now, just like you did for this one.

47:31 Um, then the courses. So, we have a 101 boot camp. This is where you build your own genetic analytic system. That's July 13th to 17th. We offer it again uh I think end of August. So, if you don't get this one in July, it's going to be another six or seven weeks till we run it again. Uh we just ran it this past weekend. And then 201 is the one built We do go into and we also do go into some eval validation stuff in that 101 since it is kind of like table stakes stuff now.

48:03 And then 201 is built purely on what we did today like validation in context. We also go into multi-models like open source and codex. Um that's going to be in 2 weeks or a week and a half. June 27th to July uh 4th. Um that's And then there's a 5-week one, AI Analytics for Everyone. Uh this is like the pure This one is like the thinking itself. So that runs right now from June 15th to July 19th.

48:34 We're on again August. Um that's about the whole like uh we we we kind of say it's like a mini masters in data science but the execution layer is in uh Claude code rather than in SQL and Python or directly. Uh Whether or not you do any of these, you're going to I'll I'll send you those worksheets and the recording um later today. Um We'll get into questions in a second here. One note on the one promo we are running right now.

49:09 Um the 5-week course, AI Analytics for Everyone. The one about uh the whole analytical thinking and then judgment behind it and how to execute uh in Claude code. We actually just kicked this off this past week so we're on day three. It's a mix of async and live material. So if you want to join this one, um you basically have just missed two office hours at this point. They're both recorded so you're we're leaving that registration up until the end of the week. Um Easily you can catch up. And you self-pace so you take it at your own time and you have you can have the content for forever. It's not like you're uh 5 weeks it goes away. And then the part that uh we think it's really makes it worthwhile is if you register for the five-week course, uh you add a seat for free, or not free, 10 bucks, um in our 101 boot camp, which is normally $900. So, you get that thinking course and the builder course basically together for the price of the five-week course. Um yeah, you scan the QR code on the screen, um or Shravya, or Hi just dropped a Hi just dropped a link in the chat. You can DM us or email us about it, too, if you have got questions. But basically, sign up for your analytics for everyone, and then just DM me and I'll give you a code to get into the $900 boot camp for 10 bucks.

50:33 Okay, let's get into questions. Uh just a few. I'll go I'll run through a few I often get, and then we'll open it up. I see there's already some coming into chat. Um most common one by far is like, you know, what if I don't have a big data team, or my my data's a mess, can I do this? So, yes. This is actually who it helps a lot. Um you can take the leaner path to evals.

50:56 You can lean on the cheaper checks like reliability and making sure it's receipts. Uh you can save that heavy stuff for the handful of numbers that really really have like high stakes behind them. Uh you don't need a super clean like data warehouse to run the same question twice and notice if it disagreed with itself, right? So, there are leaner checks that you can incorp- orate right now, today, that don't require a bunch of technical skill or rigor, or really strong data foundations.

51:27 Um they're going to flag, probably, that you don't have strong data foundations, but that's great. That's like a good step one. Uh next one I get a lot is like, how do you get ground truth when you don't have um the answer key? So, you most certainly almost certainly have ground truth. Um you've just never labeled it as such. So, think about like finance reconciled numbers, uh queries that you've uh written like dozens or hundreds of time, or someone else has. Um it's like a nice head start and it's a real place to anchor to. So, there's probably ground truth um sitting in the context of your business uh that's just not labeled as ground truth, but you can you can kind of like uh reapply it as ground truth uh to to test against.

52:20 But, I've mentioned a little bit about this in terms of like talking about the sub agents, but um we also talked about how we do like multi-model kind of triangulation. So, if multiple models agree, is it right? Um this is the one you got to be careful about. So, agreement, like I said, is reassurance. It's not proof that the number is correct. So, say if models are trained on similar data, they're going to share the same blind spots. So, they can still be confidently wrong together. Um like I said, the more useful signal I find is when they disagree um because that tells me where to dig and gives me um a lot more context about what's going on behind the scenes.

53:02 Um and then if it's stable, um if it gives you the same answer every time, is it right? Kind of similar to the last one. No. A wrong query can be perfectly stable. It'll hand you the same wrong number all day if you execute that query now and over again. This you know, just does that even do with AI analytics? That's just like people running queries. But, I've written plenty of wrong queries. Um so, stable is necessary, but it's not sufficient. Um so, it tells you the system's settled, not that it's settled on the truth.

53:34 Okay, let's get to some questions. I haven't been looking at these, but Hinds and Travia, you, questions that popped up here? Yeah. I missed I can go 10 or 15 minutes after that. Yeah, I think that's good for questions. Balance has a question around two questions. Is there a skill in the repo to help work through metric stuff definition? And the second one is how does this stuff dovetail with choosing North Star metrics in the session on Friday? I think that was last Friday.

54:06 Um I know that session was more on reasoning if the metrics shows past a certain rule break, but would it help when the definition of the metric was not completely agreed on within a team? Yeah, so we have I So, those really good questions. Um the North Star metrics is probably We have a few skills around metric, but I say the North Star metric stuff, which I'm actually pulling out into its own repo right now that I'll hopefully get up in the next week or so. But, it's in that plus one if if you all have known that.

54:37 Um I think that's probably our more robust skill around cuz it does have multi modes around auditing metrics, explaining what good metrics are, um going through the checks of a strong metric. And then you can have it um it will like uh apply those to your metric definition YAML file in there within the workflow. In terms of like Yeah, we haven't created anything, but we maybe we should do this actually. This would be pretty interesting. We haven't created anything in terms [clears throat] of like hey, how do you uh leverage a system here to like get alignment across your team. The North Star metric one I think would be really good because you could have people like come up with ideas, and then you could audit them through that system, and it will give you recommendations and show you the weaknesses, and it has that huge knowledge base that sites like actual case studies, and it sites a lot of that like Amplitude North Star metric framework playbook um that's free online. So, I think it's a nice resource to have uh when you're having that discussion with your team cuz it will tell you you can use like the I think it was like the Northstar metric scale and then like dash explain and have a concept. And um you know, that's that's like basically you're having like an expert on this in the room with you like talking to your team. But I think we could totally create something where it's like hey, record this team meeting with everyone talking about the metrics and kind of live uh give some input or trying to align or even like hit rephrase as like, "Hey, Hi is thinking about this metric in this way. Sean's thinking about it in this way. Um where is the overlap? Why are they diverging?"

56:23 I think we could create something around there. But there's nothing uh right now in there that does that. Really good question and good idea. You can make it two balance. You can make that You can You can extend the Northstar metric thing to do that team thing. >> Thanks for that, Amit. Uh yeah, I might give it a shot at and that framing of like taking in multiple definitions and running them in the Northstar metric, I think is I I think [clears throat] gets that same sort of um level setting for everyone.

56:53 Great answer. Thank you. >> Yeah, yeah, no worries. No worries. >> All right. Let's see. Next question and Al had uh "Some companies don't allow providing full database access to the AI but allow to use it without giving it any company context. What are your thoughts on organizations like banks and fintech that restrict usage of AI due to data security? Cool session, by the way." >> Uh just a matter of time. They'll either rise to the top or burn to the ground. No, I'm just kidding. Uh yeah, uh I think there's uh Listen, like yeah, a big there's a couple of things you could uh do to improve your system. There's one is like building within the system itself like agents, skills, how things are structured. The other is framing your input in a more robust way. Another is adding context.

57:48 I've found by far and especially as the models just progress on their own in terms of capability, context is the thing that really improves them. Something like a metric definition though, I mean I don't think you're like giving any company secrets away. Something about go access all of my data and store those results somewhere is a different thing. So I feel like you know, kind of some of like the just like metric definitions themselves, I don't I don't see that as like kind of going against like those like security restrictions around context.

58:31 Something that we've done at our company or my company I worked at before is anonymized a bunch of data so you could work with it more reliably that way. So we built a whole system that anonymizes data and this is pretty highly sensitive data. It was like legal AI company so you're it's like contracts and stuff. And then there are there was ways to do this more secure like we also use when we when we run things through cloud code and we do not anonymize data, we run it in like a safe like hosted instance like in SageMaker.

59:09 We run it with Amazon Bedrock models. We run it where we have like zero data retention policies. So I think you know, obviously this is a huge question that we get a lot of times like how do I deal with security and a lot of the answers are going to be like just like with any other piece of software that's going to be like legal, security, IT team thing to work with. We can talk through like some of the stuff we've seen successful before, but you know, every every every AI company, every frontier lab is like trying to come up with solutions to make this as secure as possible, too, cuz that's a huge friction point for them.

59:49 >> Yeah. And then possible another freak another option and probably where things are headed is open source models, like you host your own open source locally. Um so you can even run the AI without internet and it certainly doesn't go anywhere. Like your data doesn't go anywhere. It lives in your machine, stuff like that. Um you'll have to have or your company will have to have a really powerful machine if you want it to do the max capability on the analytic side.

60:17 Um but that's I think that's a viable option as well as people are waiting for security clearance and stuff like that. >> Yeah, and we can we actually do that in our our week-long 201 uh boot camp where we talk about validation, context engineering, multi-models, small as models as open source models. Um and I'm pretty bullish on the open source model stuff. I know Hive's been doing a lot of uh research into it. But I mean, if you think about I just look at like the model evolutions that's happened recently.

60:49 Personally, when you have a harness like you when you build a genetic system analytic system to do analytics around a model between Opus 4.6, 4.7, 4.8 and Fable, 4.6 does the just as good, actually better than 4.7 and 4.8 um when it has a system around it. If you have no system around it and you're just like opening up a terminal and there's no repo or anything and you're asking analytics questions, like yeah, Fable blew all those others out of the water, but it was just like totally on par with it when you put a system around it. So, at some point, like those open source models will be at the level of Opus 4.6, and eventually they'll be cheap, or they'll be like smaller, and be able to be at that level. It's just a matter of time.

61:36 So. >> I mean, Deep Seek V4 Pro, um the one that I've been testing quite a bit, seems pretty promising, um but we'll see. >> Yeah, we'll do some free session on the open source stuff at some time soon, and maybe in the next month or two, but if you really want to get into it, join our 201 course. >> All right. Probably the last question here. Uh let me see. Semantic layer ex- explaining definitions, data, metrics takes so, so long. Any good advice on how to cut corners responsibly?

62:10 Uh I can take a crack at this one. >> Yeah. >> Um yeah, uh I mean, Jane, semantic layers is um so, here's how I think about it. It is a very powerful way to govern your enterprise level data across the company. Uh certainly very painful because, and the pain is actually pretty interesting. It's not the execution layer anymore. Like, you can build systems where writing the pipeline, building the data models, and um actually crafting the whatever um warehouse you have, or whatever um uh uh ETL tools you have, like that's not the hard part. That that part is pretty simple. The really long, time-consuming, and people avoid part is like the people and process up front, like um agreeing on the definition, um agreeing on the specifics, agreeing on the filtering, agreeing on what each component actually means, um because they're responsible for explaining the metric themselves, and owning it. Um the the thing that I find at my company is that that is the hard part where people are like, "Oh, yeah, just you know, define something." But then at the end it's like not actually that simple because you know, like we can define something, but if they don't agree later on and you've built everything and reporting and and whatever and like they come back to you and be like, "Hey, I never agreed to this. What's going on?"

63:35 That itself is actually worse than doing the upfront work of like, you know, hey, "This is what we agreed on. This is everything that's aligned." And then everything else is like just a click of a button. I mean, I'm exaggerating here, but that's sort of like the the thing. And so, the top part just cannot be avoided. It's just people things. >> Yeah, if you want to put a line on metrics, just make metrics that go up into the right is what I find. People like those ones. But when they're when they're going down, people are saying like, "This isn't the definition I had in mind."

64:06 Um just joking, but uh okay, I think we're I have a couple other questions on here. You guys can kind of review the slides afterwards if you want to go um go just I guess the last one I would leave with here is will models get good enough and we won't need any of this eval stuff? I think it's actually the opposite. I just based on how I've seen the models progress, I the more of the work we hand off, I think the more we need to check it, not less. So, like reliability, ground truth, whatever what whether was even the right call to make.

64:46 Um none of that goes um away as the models get better. I think they're going to get a lot better at being right, but I think they're also going to get a lot better at being super convincingly wrong, too. So, I would if you're thinking about like where to invest your time right now in agentic analytics, I think getting some form of evals and and understanding the the trustworthiness is is really um high leverage way to use your time.

65:16 Cool, folks. I'll send out that email. Thanks, everyone, for joining. Uh thanks for the engagement and all the questions. We'll see you, hopefully, next Wednesday for that next lightning lesson. All right. Bye. >> Bye, everybody.

Summary

Sean, Hai, and Sravya from AI Analyst Lab discuss the evolution of AI analytics, emphasizing the importance of validating AI-generated outputs to ensure trustworthiness. They highlight the shift from simply obtaining answers to critically evaluating their accuracy, especially as AI tools become more integrated into decision-making processes.

- AI analytics can perform end-to-end analysis quickly, but trust in the results is crucial.
- Validation (evals) of AI outputs is necessary to determine how much to trust the results.
- Errors in AI analytics can occur at various stages, impacting the final output.
- Reliability checks involve asking the same question multiple times to see if the answers are consistent.
- Context and definitions play a significant role in the accuracy of metrics derived from AI analytics.
- The session includes practical demonstrations using Cloud Code for reliability checks.
- Participants are encouraged to assess their own analytics processes and identify areas for improvement.
- Upcoming workshops and boot camps will further explore AI analytics and validation techniques.
© transcribe · For agents Built with care and craft by Gokul Rajaram