Transcript
0:00 My name is Hai. I run the data team at a legal tech company called Ontra and previously spent my entire career at consumer tech companies like Nextdoor, LinkedIn, Pinterest, Meta, and so on and so forth. So, I've been doing this since I guess earlier this year and the three of us have a lot of fun doing it. I'm going to pass it over to Shawn. >> Hello everyone, I'm Shawn. Um yeah, one of the leaders of the AI Analyst Lab and I'm also a principal data scientist at the same company as Hai.
0:36 A legal tech company called Ontra. Uh yeah, I've been in data science for about 10 years and I'd say like the past 2 years I've been focused mostly on like AI and eval and digital tech analytics. So, yeah, really happy to have you all here join us today and talk a little bit about the capability of Cloud Cohesion pertains to analyzing metric tradeoffs. I'll pass it over to Shravya. >> Hello everyone, so excited to see you all here. I am Shravya.
1:05 I've I'm in the data science field for the past probably 14-15 years now. I've been with Microsoft, then eBay, then Nextdoor, that's where Shawn, Hai, and I met and most recently at Superhuman. So, yeah, been part of this journey with Hai and Shawn for quite some time. Very excited to share all our learnings with you. >> Cool. >> Yeah, and the lab we call it that we're all working together on is called AI Analyst Lab and if you're ever interested in some past lessons, past workshops, or upcoming courses, or workshops, take a look at there. Over there.
1:46 We've got just really quick, we also run different workshops and bootcamps and we have one tomorrow on introduction to Cloud Cohesion Analytics uh then we have another one, uh boot camp, just a weekend after, which is 2 days. Um we'll talk a little bit more about what those are uh later in the lesson. Okay, so what are we here to talk about today? Uh we're here to talk about guardrails. So, um like understanding how things move against each other and then how you as either a data professional or a business owner or decision-maker, um come up with plans to sort of mitigate uh the downside that may not be visible to you. So, um so in this slide, assume that uh we have, you know, like a product team optimizing for um something like a mobile checkout uh funnel. So, when you, you know, like um uh when you go buy something, let's let's let's think like Amazon.
2:50 Uh there's the, you know, if you've shopped on your phone, you probably have a uh go through a bunch of payment steps, go through a bunch of add to cards, and uh and finally input your credit cards and stuff like that uh at the end to to to actually check out. So, um you know, like I think it's probably not super surprising that um uh things could go, you know, like one metric could go up and another metric could go down. And uh as a uh rigorous sort of analytics practi- practitioner or data practitioner, um our job is to make sure that uh we cover all the blind spots as early as we can and plan ahead of time as much as we can. So, uh this example here, very simple, it's like, "Hey, assume if or let's say like if something uh caused or some sort of feature that uh that the team uh launched on a checkout uh flow led to conversion rate that goes up 15% perhaps, um but then 2 2 weeks uh, actual net revenue was down. This is probably not that big of a, let's call it like right now still hypothetical, but probably more common than many people sort of like, uh, have the, uh, have the discipline to, uh, to, to uh, to plan ahead of it. Uh, and so this entire lesson, what we're going to do is to really just think through how to use a very simple framework to, uh, to make sure these traps aren't being, uh, aren't being stepped on.
4:25 Who, um, actually would love to get a poll here. Like who, who, who's encountered something similar at your workplace, in your, in your role, or have seen this, uh, play out from time to time, or ever? Is this a common thing? >> I saw this a lot working with, uh, emails and notifications. Lots of, uh, trade-offs in terms of optimizing this. >> Nice, yeah. >> Tristan's got their hand up.
4:57 >> No, I'm saying that I I saw this happen. >> Oh, you saw you saw this happen. >> Oh, cool. I saw this happen in several of my clients. This this this is happening so especially in each e-commerce industry. Um, small shops based out of Texas, we see this not not only for 15% is a very small window for them. Um, sometimes the sales double or triple in in a very in a very short span of time, and then they fall down pretty pretty significantly.
5:26 >> Yeah, wow. Yeah, I mean trade-offs, decisions happens all the time, and so, uh, this is what we're going to talk about is really just, uh, you know, a tool to, uh, help, uh, make it, I guess, uh, more streamlined. Um, I mean tough decisions is still going to need to be made at the end of the day, uh, but, uh, you know, as much of a structure we can put around it, the less taxing it is every time that it feels like an ad hoc thing over and over again.
5:54 Cool. Okay. So, before get into the actual framework and a demo with cloud code, just wanted to sort of like let this group know, if you follow us on LinkedIn or in previous lessons, we've been pretty like we've been pretty open about we've developed something called an AI analyst, which is almost like a senior product data scientist clone in cloud code. So, we've got a agentic system where it it bakes in all the best practices for how to be a really effective product data scientist.
6:34 And so, you know, this is free to to clone. Everybody can go and and and and take a look or use it for your own work or your own interest. It's all cool, open source, completely free, tons of skills, tons of agents that we encoded the, you know, our collective 50 years of knowledge in this field. And you know, from in this lesson, what we're going to be talking about is going to we're going to grab like a very specific piece from the AI analyst um system to just walk through the guardrail component of how you should think about success metrics, but then also like what we call shadow in in this lesson here.
7:23 Cool. And yeah, just quick plug for for folks who are interested in, you know, maybe you've seen the AI analyst repo, maybe you've seen us taught a few lessons in the past, and you've heard it from this from from the audience from this work from this lesson as well, earlier today. We're running a couple things. Tomorrow we have a workshop. This is hands-on. We're going to help you walk through how to install Cloud Code, how to clone the repo, run the real analysis, and all three of us will be there for 3 hours to make sure everybody is totally set up.
7:59 That's sort of like the the shape for the workshop tomorrow. And it's 25 bucks. We just want to make sure we wanted to make it free, but we also want to make sure that people actually do show up. So a small a nominal amount to hopefully get you on to the journey of really up your productivity in the whole data slash analytics space. And then the weekend right after is the boot camp, which is a two-day thing, 4 hours each day, where we're going to walk you through how to build your own AI analyst system. So all the thinking's behind, you know, like what what what goes into spinning up a skill, spinning up an agent, how do you wire them together, all that kind of stuff.
8:45 That's going to be the that's going to be in a week and a little more. We'll come back to the specifics on these at the end. So right now let's go back to the content here. Cool. Okay. So the main idea of this lesson really is is is this. So every success metric has a shadow. So the shadow, it's also called guardrail, is a metric that gets worse when you optimize the success metric, you know, too much. So it could be like I think the you know, we'll get into the examples later in the in a later slide, but a lot of teams don't actually think about these So when the success metric moves, they would immediately say, "Oh, yeah, that's a win. Like, we're we're we're just, you know, like that that looks great. We're we're done." Or or they celebrate too early. Um but if you don't check these shadow metrics or guardrail metrics, then then you really don't have a complete picture of the mechanics of how things are working.
10:02 Um so, you know, in in addition to making sure you make the best product decision or business decision with a holistic picture, you also need to it also helps you teach kind of like, "Hey, when this thing moves, this thing moves like this." And so, that muscle memory is actually, I would argue, more important than the on-the-spot sort of one metric or one instance or one feature that you are or or your company is shipping. So, hopefully that makes sense.
10:35 Cool. Okay, so here are some examples of um shadow pairs. So, things like what left on the left side is success, on the right side is the guardrail. So, first one here is um is you know, like very commonly, if you're in the e-commerce business, conversion rate, for example, so purchase purchase purchase conversion rate or like checkout conversion rate, something like that. And then the guardrail for that could be what is the average order value. So, if you, for example, if you make things so easy to to check out, like to buy, does it how does it affect how much you're buying?
11:18 And these always, you know, like these would move these would move in interesting patterns typically. Um the second example here is sign-ups. So, converting into a for example, like signing up for an account on some website, for example, the the uh the the guardrail for that would be are they retaining like are these sign-ups just, you know, like very low intent people who are just, you know, having a very easy time to get an account or are they actually serious about sticking around, which is the day 30 retention in the in the guardrail component. So, you know, like that's that's kind of like, you know, it's hard to when you start naming these pairs, it's going to be hard to think how having visibility into this uh would, you know, make it easy to game the system, for example. You know, in the case of sign-ups, like the easiest most more a very classic example is like, "Hey, I can always make the sign-up button bigger, right?" And people will convert and so on and so forth, but like, you know, probably most likely they're not going to stick around, the people who sign up because of it.
12:34 And then, oops. And then we have engagements versus NPS. So, net promoter scores or complaints. So, you know, like you can have very what do you call like attention-grabbing sort of stuff that gets people to be very hooked to your product, but then, you know, like your your complaints might go up or your NPS might go down. So, it could be you know, like you could send more emails, but then people might might just get pissed off because of it. And then the final example here is support speed. So, like if you run a customer or I guess a call center of some or like what do you call those like customer experience um teams where, you know, the job is to help people resolve their concerns or whatever, you can imagine. Hey, if their goal on just support speed, like how fast do you close the tickets, then it could be that, uh, the the guardrail for that is, uh, how much of those tickets when when you are so aggressive in closing would actually become open again. Like, you know, the the actual, uh, issue may not have been resolved, but then because the metric itself is to measure against the time to which the ticket closes, uh, the, uh, the guardrail would be, okay, it would suck if, uh, if if, uh, the same issue is open again.
14:02 Does that make sense? Hopefully, this is pretty straightforward. Cool. >> Yeah, this is pretty straightforward. Thanks. >> Great. Okay. So, let's see. So, um, the actual framework here that's, uh, that that is is, uh, helpful. And we walked through two of the sort of like three things here. Um, so, really, there's only three decisions that you really want to think about in terms of, um, pairing guardrails to your success metrics. And, uh, we obviously we talked about two of them. The one The first one is success metric, like, you know, like no surprise. What are you actually trying to improve? And, uh, what are you trying to actually measure to proxy that improvement? So, just some very concrete examples here. A bad success metric would just be, you know, something called engagement, for example. Like, what does that mean? It could be defined billion different ways.
15:00 Um, a good one would be very specific. So, for example, in the whole e-commerce, uh, space, um, you know, could be purchase rate, uh, for your new users or something like that. So, um, very specific, uh, very intentional about the thing that you are actually trying to move, uh, with a feature or with, uh whatever uh product improvements that you may have. And then the second piece is the guardrail, right? Like the uh guardrail itself. Like what is the shadow here?
15:32 Um it is the thing that catches the tradeoff. So, I think a lot of teams that I've seen in the past might be like, "Oh, let's just track everything, or uh let's just uh look at whatever uh we have uh uh that that we don't deem success metric, and see what that looks like." Um the really a really good practice is actually to predefine it up front. So, like be intentional about, "Hey, if we if if we think this is the thing that we're trying to improve, uh what is something that we're not willing to compromise?" Um and so, and also be very specific on that. So, here it's uh you know, if you track everything, that's a that's a bad statement generally. Um a good one is if you list out sort of like the possible things that this could go wrong that you think would be um would be detrimental for the business, that you know would be detrimental for the business, and then just make sure those are explicitly named. Um making it up front is actually one of the most important aspects of this, um because when you already launch the feature, for example, and then you then start to hunt for uh data or metrics to either support or or make the decision on the fly, that's actually really uh that's actually not very very healthy, because then you're you're basically just trying to uh uh trying to use data to kind of like um uh uh you know, like like almost like data theater, uh to to prove a point, if you will.
17:11 Okay. And then the third one here is actually um uh what do you call that? Like, actually put a threshold on the guardrails um, that that you've picked. So, it's one thing to be like, uh, hey, here's my guardrail or a set of guardrail metrics that we care about that they shouldn't, you know, move or whatever. Uh, in reality, if you drive up, for example, a success metric, most likely the guardrail would also move in a different direction. That's that's why it's called a guardrail anyway, uh, because, you know, like, there's a good chance that, uh, that there's some there there's adverse effects on that metric, but you're hoping that the gain of the success metric outpaces that by a lot, for example, or that, you know, in the best case scenario, it doesn't actually move, um, or, you know, even improve. So, um, pre-aligning and pre-defining on what that threshold or the willingness for you to be comfortable with defining let's call it like a feature to be successful if it moves success metrics by this much and it doesn't degrade the guardrail by more than this much, then you have a really powerful combination of a clear, um, criteria that you don't even have to think about when you actually launch the feature or, um, or make the product improvements. And then, on the slide here, there's couple kind of like ways to think about it. One is like alert threshold. So, for example, if it only moves by a little bit, uh, and then like, you know, in that area, um, based on your prior context or your, uh, business domain knowledge about how your products or how your features move, um, you know, maybe 3 to 5% is not a big deal. So, anything below that, you're good to go. Um, uh, and then there's another one which is like, hey, an absolute kill threshold which is if it goes worse than this threshold, then we're absolutely killing the thing." Or, you know, like we'll absolutely pause it and then go figure out what's going on.
19:23 Um so, always having these kind of predefined is going to be helpful so that you're not scrambling after the fact to be like, "Oh, like what does this mean? Like is this good? Is this bad?" And then like make decisions on the spot. Does that make sense for people? Like in terms of, you know, why this is important? Like doing it up front versus doing it afterwards. >> And something else and you might get you might get to this later high is like it's not just about a like you as an individual making these decisions around the alerting kills. It's also about just like being able to align teams. So, you want to take as much of the decision-making process as possible and just do it before any data is even created or looked at. So, everyone's aligned and basically like their opinions aren't biased or polluted after the fact based on on what happens.
20:23 So, like I don't I mean if you like have like a finance team that cares about one about one metric and a sales team that cares about one metric one metric and a product team that cares about another metric. If you're all all aligned in the beginning, then it makes it really easy after everything's said and done to to move forward with whatever your decision is. If you don't have this alignment around thresholds before, then those conversations um can end up like with a lot of conflict. They can get like even like political um and all this just unnecessary stuff that's basically like you can avoid by just setting some some numbers ahead of time in a document that you can point back to later on.
21:07 Cuz everyone's out out there trying to do what's right for the business, but everyone has also like their own priorities, right? >> Yeah, the more complicated, the more tricky, the more sensitive a feature or an experiment is, uh the more important it is to pre-align stuff so that you don't get into Shawn's described scenario where it may even get political. Cuz when people see that uh especially people who are like, "Hey, it affects my area positively." Uh then I'll I'll I'll go I'll probably go crazy to try to, you know, like uh uh not uh not not care about the other areas that that get affected by this.
21:55 All right, cool. So, we'll get into the cloud code component pretty quickly. Uh but uh let's do a very quick exercise here. Um for each of these success metrics, maybe drop it in chat. Uh what guardrail would you put in? So, let's app store rating. If this is the success metric that people want to optimize for, what is uh what is a shadow?
22:32 Shadow for cards is card abandonment. Yeah, that's a good one. Bot response, app store rating, daily active users. Yes, Chris, you got it. That's awesome. Yeah, um some sort of engagement or feature usage. Uh so, um the idea is like, "Oh, you know, app store rating uh may or may not be reflective of the actual um you know, like almost like a qualitative versus quantitative aspects." Um and both can go can go in different directions.
23:05 Uh, maybe we'll do one more and then we'll get into clock code. Push notification, what is the thing? What is the shadow here? What is the guardrail? NPS, push disables, unsubscribes. Yes, great. Yeah. Unsubscribes or um, or anything that indicates uh, just being annoying. Uh, so NPS is is fine. Um, unsubscribe rate is fine. Push disable is all good. These are all in turn is good. Uh, probably more laggy, um, but yeah, it's uh, definitely definitely good idea.
23:46 Cool. Yeah, so seems like everyone is uh, now pretty comfortable looking at what this what the whole guardrail thing looks like. Uh, okay. So, we're going to get into clock code in a bit. Um, just wanted to set up the demo scenario here. So, in the AI analyst agentic system that we open sourced. Um, and again, feel free to download it, take a look, poke around, and uh, build on top of it. Um, the kind of like the the the the fictional data set that we included in the in the repo is uh, a fictional company called Novamart, which is an e-commerce company. Think of it as uh, like Amazon.
24:30 Like uh, yeah, think exactly like Amazon. Uh, it's just a uh, a fake company. Um, and um, uh, let's say if a product team is looking to ship a feature called save for later. So, some sort of button that's like, oh, you know, instead of buying now, like uh, I'll just save for uh, save for a future time. And then you can come back to it pretty pretty pretty easily. Uh, hopefully this is straightforward for folks because of uh, the whole Amazon mechanics. So, think of it as the same thing. Um and, you know, our goal is to understand, you know, for this feature and given this primary metric of 30-day purchase rates, for example, what would be possible guardrails that that we can that we can have. And then we're going to get into cloud code right now to use it pointing at some of the the skills that we have in the in the system to help us answer this question.
25:33 And then we'll see how it works into kind of like future analysis. Uh let me find the share button here. Okay. All right. So, I am switched to the other screen. Can you guys see my cloud code thing? A black screen here? >> Yeah, it has like the overall conversion. Um probably. >> Awesome. That chart has nothing to do with what we're talking about.
26:04 It's a default thing that's that that I was looking at previously. So, Okay. So, for folks who are not familiar, this is VS code. It is an ID environment. Think of it as like a workspace that you set up such that you can have a very intuitive easy to to to to use navigation panel to look at your folder folders and files. So, I'm in AI analyst plus, but think of it as like the the AI analyst repo that we shared. And within it, it's got different agents, different skills, different slash commands and stuff like that. So, if you look at it, it's like it's got a forecast skill, it's got a metric skill. And what we're going to be looking at is the guardrails.
26:54 And if you've never done Cloud Code or anything, it's actually very straightforward. No coding background required and the workshop that we run tomorrow exactly shows how to, you know, get this whole thing set up step by step. Um what I'm doing here is um Cloud Code in terminal. So right here at the bottom is basically my terminal. Um And I just sign into I just went into Cloud Code and and start working there. So what I'm going to do because there's so many different components here in the in the system, I'm just going to ask it to tell us what what what sort of a guardrail or metric pipeline does it have. So I'm going to >> Maybe while you're typing that, hi. Just to get a feel for the room.
27:51 Um curious who here has used Cloud Code for analytics before. Maybe drop a zero in the chat if you've never used Cloud Code before. A one if you've used Cloud Code but not for analytics and a two if you've used it for analytics. It kind of helps us know the audience. Cool. So a lot of people have used Cloud Code. Got a couple who haven't used it. Couple of people who have used it for analytics. Most people have used it but not for analytics so far.
28:17 >> Got it. Cool. Nice. That's a really good poll. All right. So I just I'm firing this prompt. Can you show me an ASCII diagram of the metrics and guardrail pipeline? So it's going to go look and explain to me visually what it is that we have in the Agent X system as per as pertains to metrics and guardrail metrics. So I'm just going to fire that off. Uh Uh, let's see. Code.
28:49 All right, it should be relatively quick, I am hoping. Yeah, so what it's doing is uh burning tokens, um but also looking at uh doing its own thing to show this at the end. Okay, so let me go here. So, this is the illustration of the workflow in the uh in the AI analyst system, so uh metric and guardrail pipeline. Uh it's visually trying to show us what this looks like, what the sequence of steps is, and what it does when it encounters metric and uh uh and you know, like uh how it how it kind of um uh works with success versus guardrail metrics. So, step one in this whole pipeline where, you know, if we use it to do analysis and stuff, um it'll have to help it define certain metrics. So, it's got different uh components to it, and it gets very specific around hey hey, what what is a reusable piece of metric when we say, you know, whatever, like monthly active user, how is it actually defined in the most granular way possible. So, it's got different skills around metric specs, uh and then it forces the system then to pair a guardrail against it, at least one guardrail. So, uh one at least one guardrail for success metrics, and then measures a different dimension. Um it gives some examples around uh if it's quality, then look for something that's quantity, speed versus accuracy, that kind of stuff, and then figure out the acceptable thresholds, and uh what things should move by. Uh it could figure out based on the data that you have. Um uh it but the general recipe is that it will figure out it's it will take it will make sure to bake that in as part of the uh the job, if you will.
30:46 And then there's the kind of like the what do you call that? Like the examples of common pairings that we went through. It will start to it will it will when it actually does the analysis, for example, in an experiment setting, then it will, you know, obviously compute the two and then it will check against the guardrails and then it will have verdicts and the threshold that we just talked about. Right now it's just an example, but you can imagine like it could be specified based on the based on the business and the domain and and and your use case. And then it will give a report, that kind of stuff. So, what I am going to do is um we're going to we're going to do the exact setup that we just looked at from earlier slide. I am planning actually, let me just paste something here. I already have it written down.
31:45 All right, cool. So, this is the prompt the next prompt that I'm going to going to give it. I'm planning to ship a save for later feature on Novamart. So, I have the data set already loaded. It comes pre-installed, if you will, with the AI analyst repo. I'm telling it the primary metric is a 30-day purchase rate. What we're looking for is a plus 15% um increase in that for the save for later feature and just help me design the guardrails for it and then we can actually use it to um to brainstorm slash confirm what would make sense to do.
32:27 And so, yep. And so, it's going to keep thinking a little bit. It will go into the the the guardrail skill. It will go into some of the metric skills to to figure out, you know, hey, in this scenario, where where is my question fitting in the in the sequence here? Cool. Claude Claude Opus 4.7 is not the fastest model, but oh, okay. It's fast enough here.
33:02 All right. Look at the fat of the flattery language. Good feature to guardrail. >> That's one thing it'll never forget to give some flattery. >> Yes, exactly. Okay, cool. So, good feature to guardrail, big risk is cannibalization of immediate purchase. So, notice we do we did the like it understands that the save for later feature is the the thing that, you know, like that you can uh uh that you can come back to hopefully uh in the future when you're not as ready to make the purchase right now.
33:39 Um and then it's flagging that the big risk is that maybe people would be over reliant on the fact that hey, I can always just punt it down to to a later time. So, people would actually not do immediate purchase. So, that's the big idea here, and then it's gotten it's giving some recommended guardrails based on that. Uh the uh so, the guardrail could be same session purchase rates, average order value. It would have the uh the uh what do you call that? Like the acceptable range as well based off of the data that that it has access to on the Novamarts company.
34:19 And it also has its own judgments around why these things should um you know, should exist or what the rationale is for picking some of these picking some of these things. So, it's pretty detailed and uh you know, like right now it's just the interactive kind of like me talking to Claude uh on the screen. Uh but in the actual skill, for example, uh it will actually be able to write out a detailed report once we make some decisions around um you know, these are the guardrails that we care about, uh that kind of stuff. So, um it even says, "Hey, uh three things to nail down before the launch. What are these uh attributions, measurement design, blah blah blah." So, uh when I need you to rewrite the full spec, um so it has access to a metric spec skill so that it can log when we make these decisions that, for example, hey, if I if I think return rate is really important here, then it will log that in the system such that in the future, when I actually use it to analyze an actual experiment or um or analysis, it understands that that's what I care about. So, let's do something here. Let's say uh we don't actually track return rate.
35:38 Uh what else would you suggest? So, I'm trying to show the interactive nature of using it as a brainstorm buddy. So, almost like a you know, like the thing that the phrase that we like to to say a lot in the AI analyst lab is that uh you're getting a an an AI companion effectively. You can delegate your execution to uh to to an AI assistant uh that think that thinks like the best, you know, as as much as possible that we bake in, the best product data scientist out there.
36:16 Cool. Thought partner, yes, Praveen, great point. Thought partner, not a replacement. Uh okay, so two reasonable proxies, it says, "Hey, we don't track return rate." Um then it would be like, "Hey, do the ticket rates, pre-shipment cancellation rates." Uh and then it it will tell me why this would actually proxy the return that we don't have or that we don't track. Uh and then what are the caveats and uh what what are some acceptable ranges that it then um kind of pivots to uh with a new metric like this. So, uh also has recommendations and uh and stuff like that. So, uh so you can go pretty deep with it um in terms of thought partnership, brainstorming, and stuff like that. So, what I'm going to ask next, and then we'll open up go back to the slides and stuff. Hopefully, this gives you an idea. Uh what if return rates increases, but 90-day purchase rate remains flat. So, I'm trying to see So, right here at the top, the some of the recommended guardrails says uh "Hey, return rates and 90-day repeat purchase rates, um I want to see or not Yeah, 90-day purchase rates, um I want to see hey, if these things move in different in a given scenario, what does that mean?"
37:43 Um so, that's um you know, you can you can uh you can figure out pre-aligning up top why what you can expect and what you would do if you see certain things. So, this is really helpful to just, you know, like get your uh get your rationale in the right place. remains flat. Uh What does that tell us? And so, it will continue to think and give us a scenario, and then you can imagine this could be done for a billion different scenarios. So, the the the really helpful thing to do is just to be very comfortable with some of the choices and scenarios for things if they happen this way or that way, if it's unclear.
38:42 Uh 36 seconds. Yep, Opus 4.7 is not the fastest model. Um so, if folks have a way to make it run faster, then let us know. >> Use 4.6. >> Yes. >> We talked about this in the chat, but I I also like to toggle the effort on 4.7 to something that's not as high. >> Yeah.
39:13 >> Okay. So, translation feature is changing what customers buy, not whether they keep coming back. Um so, yeah. So, most likely story generating incremental purchases by meaningful shares or speculative. So, it gives some possible explanations for uh if this were to happen, then that means that. So, assists you to not have to think uh you know, like I mean, you know, it would be a lot easier to just chat like this and get a decision than to than to I guess figure out all the scenarios and then all the different different combinations of what would have happened.
39:56 Okay, cool. I am going to switch back to my slides here. >> Yeah, I question in chat here. Hi from Or David, you want to ask your question? >> Uh let me see. Yeah, which question, sir? >> Uh from David here, are the acceptable ranges derived from your historical data or Claude using best judgment. >> Yeah, so um great question. So, I've done some analysis in the past uh with this exact data sets um quite a bit actually cuz we recorded some lessons in the past um and things like that. Uh so, as you use the repo more, it will know more of uh of the shape of the data sets and understanding of it and what you care about and what you don't. You can also name it explicitly to it like, you know, hey, I care about the threshold being less or more than 5% for example, and then it will it will codify it. So, pretty flexible. Um for this one, I think I just figured it out based on my previous runs.
41:00 >> Got it. Thank you. Appreciate it. >> All right. No problem. Great question. So, okay. So, let me go back to here. Am I showing the right screen? >> I'm seeing the slides. >> Cool. Okay. Same uh black screen here, so I can't really tell. Okay, we did the demo. Um oh, I I wanted to leave the last at least 10 15 minutes for questions, so um couple ways to take the next step if you're ever interested, just like what we talked about earlier in the in the lesson. Tomorrow, again, we have the intro to Cloud Code Analytics. We're going to help you set up the whole thing, install, clone, do actual real analysis, and we'll be there to help you guide every step of the way.
41:49 Um in a week and a half-ish um during the weekends, we have the analytics bootcamp, the Cloud Code Analytics bootcamp. Um for attending this lightning lesson, you'll get 20% off with the code cloud20, and uh this expires tomorrow end of day Pacific time. So, something to Uh we would love to have you join us if you're interested. This is where you're going to build your own AI analyst system such that you know, like you have you can use it on your own data. You can you can have it tailored to your own use case, your own habits and and and your own context. So, you know, learn to you can come and learn to build with us. We've run this once and uh uh pretty good really really good reviews from from our from our students.
42:42 >> Maybe one more plug for the the workshop tomorrow. So, like for the reason we're running this workshop tomorrow is um based on feed feedback from our our our initial boot camp and just kind of what we've seen um kind of bringing uh synthetic analytics and and just Cloud Code more broadly to all the companies that we work at is that the hardest part of and and I think I we saw in the poll earlier that a lot of people were were saying one so they've used Cloud Code before. Um but a lot of the the the the hardest part of this I think is just getting started and I think that like, you know, it can be daunting for a lot of people to get started with Cloud Code or working in terminal especially if like, you're not super technical person. But I mean like once you get started I and it's like a couple hours of friction we find that you're kind of just like really off to the races and dreaming up stuff that like no one else would think of. Like for instance, like my wife's in marketing. She's also like a ceramics artist and she's not technical at all, doesn't know any sort of coding, but um I got her set up with Cloud Code to help her like build out her ceramics business and then she's just asking and creating and developing things that I would have never um of before. So, we want to do this live even though there's like a lot of I'm sure tutorials around there around getting set up. We think that like just spending a few hours with people live is a great way for us to unblock people and have people other people like unblock each other as well and kind of support each other. So, that's why we're doing tomorrow to you know, it's only like 25 bucks.
44:25 Um so, yeah, give it a give it a check out or feel free to DM us if you have more questions about that. It should be pretty fun. >> Yep. Looking forward to seeing quite a few people from from this crowd tomorrow. Okay, cool. So, this is the last slide. We've got roughly 10 minutes left. So, opening up for Q&A if there's been any questions maybe in the chat or if anyone wants to come off mute to ask anything top of mind.
44:59 >> Or feel free to drop questions in the chat if you want. >> Praveen had a question to the follow-up on the thought partner. Praveen, do you want to go ahead and ask question? >> yeah, yeah, sure, sure. Thanks. Thanks for bringing it up. So, I had a question about what do you do you recommend? I think Shawn you mentioned about um I'm just just looking at my question. Sorry. Um You had a you know that you can have So, my question was when you want to have a thought partner, do you go about changing your original cloud code MD file to to give a brutal honest opinion about your metrics or do you go about using another agent that you have you can create I think you can create autonomous agents through cloud right now and use that to stress test you know, what you have what results you you get. So, I'm trying to see what what would be the best approach.
46:05 >> Yeah, I definitely do the sub agent route. Um a couple reasons, like one, you have all this kind of stuff baked in to your system already. That's kind of biasing it, but you also have like if you're working with like your primary kind of like orchestrator agent in Claude code, they're going to be biased by all their context. They're going to be like do they're creating the work. So, like when they judge their own work, they're going to be biased to saying it's correct.
46:33 Even though I do find it it does it will catch itself. So, either if I if I if I'm kind of done with where I'm at with it, I'll I'll sometimes just clear context. And then if I clear context, I'll have it recheck. But, the sub agent it was is usually what I do. And I even have like a skill that's called like it's called like it's like a it's called architect or something, but it's like a plan mode where I'll have multiple personas created specific to whatever I'm analyzing or developing.
47:05 And each of those will basically provide will read whatever's the analysis is, and then they'll separately critique it. And then they'll come together as a group and argue with each other. And then they'll go back and revise their critiques and they'll come back together in the line on like a like a line on their feedback. But, I do find like forcing them into their own context that isn't polluted by everything they built before is a good way to get more kind of reliable answers. Um we're also going to talk about we're not we don't go over it in this coming boot camp. We have it in the advanced boot camp in June where we're going to talk about leveraging both like Codex and Claude code together. So, having totally other Um you know, frontier models critique the work.
48:00 Find is pretty useful as well. >> Yeah. >> What What What happens with me, too, is when you ask it to review the whatever cloud code generated, it tends to like be biased and say yes, oh, this is the right way. But uh if you kick off a sub agent with minimum context, it does a way better job at validating the plan, whatever it is that you're working with. So, I do uh the same as Sean What's Sean shared as well. And uh we plan to share all of this as part of our advanced boot camp. We just realized there's so many questions around this and we don't have time to do that as part of our boot camp that's coming on May 23rd. So, probably 2 weeks after that we'll have an advanced boot camp where we'll share the best practices, like how do you go around like multi-project like, you know, multi-project flows, how what goes into cloud.md and, you know, all of that stuff.
48:53 Yeah. >> Cool. Uh let's see. Uh looks like Tomas, you have a question if you're still on. Or you have two questions. Do you want to ask them live? >> Uh sure. Can you hear me all right? >> Yep. >> Um yeah, so first one is when feeding data to an LLM, uh what format is best? Markdown, CSV, JSON, whatever? Does that matter on the purpose of the analysis or what what should I think about there?
49:39 >> Yeah, I found it doesn't really matter. Um it knows it knows how to how to figure it out, basically. Um and even if, for example, like within uh your, let's call it like CSV for example, if there's like bad formattings or like very irregular stuff, it it figures it out pretty quickly, pretty easily. >> Right. And what about the numbers? So, how how detailed should they should they be or should they be rounded?
50:11 And does that matter if, let's say, I have a data set with thousands of entries? >> Yeah, they can they can just be like raw numbers. So, if you think about this, it's going to be um everything that's ran here um the numbers themselves aren't fed into the LLM. Basically, the LLM reasons through uh Python functions to run that will calculate over the numbers. So, from that point of a Python script, it doesn't matter if you are going to give it a like rounded number or it's more most raw form. So, I'd probably go with like whatever its most raw form is cuz it's the most accurate and then you can work with Claude to decide like from that at the output stage like what's the right kind of uh rounding you want for your audience.
51:16 I don't know how you're Shravi if you have any thoughts on that. >> Yeah, I I think I think that that sounds that sounds right. And then there's we've got questions. Sorry, go ahead, Tom. >> Uh just to follow up there. So, obviously, if if it's going to be handled by by a Python script, then yeah, for sure raw would would be the best. But, what if you're actually going to use the LLM itself to do pattern recognition? Would then would it then be a benefit with uh rounded numbers.
51:53 >> That's a good question. I mean, I think you could. I think I would actually I think I would want the LLM if it was going to do pattern recognition, I would think I would actually want it to create Python helper functions to do the pattern recognition itself. Basically, my my kind of like philosophy with anything with numbers and analytics with LLMs is I want it to do as much of that in a code as possible because whenever you start putting the numbers directly like through the LLM, you start opening yourself up to the probability of hallucination since it's now like in that mode of like predicting the next next token. Whereas, if you have it like create some functions to um to do the pattern recognition itself, then you know it's just reading output.
52:48 Um But that being said, yeah, I think like probably rounding if you're going to do some sort of pattern recognition, I imagine it would be a lot easier for it to like not get tripped up on if you had something with like . and like eight decimals or something like that. I don't know hire Shravya if you have any follow-ups on that. It's a pretty good question. >> Yep. Yeah, I agree. And I I think on the on the point about the put things in code or reason over code, it also expands beyond analytics as well. So, you know, like I think quite a few folks have used quad code but not for analytics. And so, if you've ever built workflows or automations or whatever it is, something that becomes repeatable when you start putting it in code is actually going to guarantee reliable reliability much more than if you're if you have a reason over it's in its raw form all the time. So, I think that carries over to how we think about the analytic side as well.
53:58 >> Thanks a lot. That makes a lot of sense. Thanks. >> Yes, thanks for that thanks for the question. >> Any other questions? >> Got a minute left. >> Where will the recording be available? We will we will send out the recording in an email probably later today or tomorrow. >> Anything else? Yeah, we'll see. >> If no other question, yeah, tomorrow join us for the workshop if you like.
54:30 It should be interesting one getting set up. And then the boot camp's really fun. The May 23 24 boot camp's really fun. We're going to run it monthly if if you can't make it in May. We got another one in June. I think it's like June 13th and 14th or something. I can't remember exactly, but we'll run that monthly. It's got like a I don't know, I think it's got like a 4.95 out of five star on Maven right now.
54:55 Um it's honestly just a good way to like meet other people in the space as well, too. Um yeah. If you have any questions about that, feel free to to ping us. >> I also want to let everyone know we have a Slack workspace where you could join and you know, meet other people with the same in the same area. So, join the Slack workspace. I just gave the link. We have around we have around 600 people. And we'll keep sharing all the workshops there and you know, everything about the courses and other things. So, and you could ask me.
55:27 >> Yeah. >> Yeah. >> We have two workshops next week. We have two other free workshops next week. >> All free lightning lessons. Yes. >> Yeah, I'll probably send out an email to everyone that has those as a reminder, but but on Wednesday we're going to go through root cause analysis in Cloud Code. And on Friday we're going into turning insights into action in Cloud Code. So, those are those are free kind of walk-throughs um just like this where we'll talk about concepts and do a demo and um yeah, and do some Q&A.
56:05 >> Cool. Awesome. Well, thank you all for joining us. Um hope to see you tomorrow. Otherwise, have a good weekend. >> See you, everyone. >> See you, everyone.
Summary
- The speakers, Hai, Shawn, and Shravya, have extensive backgrounds in data science and analytics, primarily in tech companies.
- They introduce the concept of "guardrails," which are metrics that may decline when a success metric improves, highlighting the need for holistic analysis.
- Examples of success metrics and their corresponding guardrails include conversion rates paired with average order value, sign-ups with retention rates, and support speed with ticket re-open rates.
- A structured framework is proposed for defining success metrics, identifying guardrails, and setting acceptable thresholds to guide decision-making.
- The AI Analyst Lab offers workshops and boot camps to help participants learn about Cloud Code and how to implement effective analytics strategies.
- The importance of pre-aligning teams on metrics and thresholds is stressed to avoid conflicts and ensure cohesive decision-making.
- The session includes a live demo of using Cloud Code to analyze metrics and establish guardrails, showcasing the interactive capabilities of the AI system.
- Participants are encouraged to join upcoming workshops for hands-on experience and to engage with a community of data professionals.