transcribe

Run and Analyze A/B Tests End-to-End with Claude Code

AI Analyst Lab · 1h 0m · transcribed Jun 2026
More from AI Analyst Lab Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:00 It's great to see you all. Uh I think most uh most of the people will start joining as well. Uh we could probably start with introductions uh from the start. Um I am Stravia Maripali. Uh I lead growth data science team at superhum. It was previously Grammarly. Uh it's now we we are now superhum. I have experience at uh Microsoft and uh a bunch of places like at eBay and then next door. That's where Sean Hi and I met actually. So Sean and Hi are part of my team. Uh and uh Sean and Hi, do you want to share a bit of intro about yourself too?

0:39 >> Yeah. Yeah. Uh so hey everyone, my name is Hi. Um and uh I'm the head of data at a legal tech AI company called Entra and previously uh worked at big tech in next door, LinkedIn, Pinterest, Meta, so on and so forth leading data science teams. So really happy to be here and uh speaking to you all. >> Hey everyone, I'm Sean. I'm a principal data scientist. Uh a lot of the work I do is around AI evaluation, AI analytics. Um yeah, been in data science for about 10 years. Uh what one thing I like to do in the beginning of these things, I'm I'm located in South Lake Tahoe, California. Where's everyone coming in from? Because usually we have a pretty uh international uh drop it in the chat. Where you from?

1:24 Yeah, northern Utah, India, Toronto. >> It's so cool. Like that's where everyone's from. >> Oh my god, I can imagine from India. It's so late there. Great you guys are interested in joining us from India as well and a bunch of other places. Pretty cool places. >> Ben W, I'm probably 5 minutes away from you. You're in Pacifica Bay Area. >> Nice. Nice. So exciting. So exciting. Okay, we have a pretty packed agenda today, guys. So, as soon as I think we're getting most of the crew in and probably maybe one last question about before we get started, what is something that you're looking forward from this session? Any anything interesting that you'd like to share in chat while the new people who are coming in could share about where they're coming from?

2:22 Okay. Okay. Good. Yeah, cool. Sure. That's fine. Uh I think um I I'll probably get started. Let me start sharing my screen. Yes, we're going to cover all of that, guys. Like I I love all the great responses. So, can you all see my screen? Okay, I see a yes from Sean and Hut. Cool. So what are we going to talk about today? We're going to talk about running and analyzing experiments with cloud code. Yes, I talk about cloud code here, but guess what? We also do cloud web a uh cloud web UI because I know that most of you uh could easily access cloud web and we have some interesting things to share with you at the end as well that you could literally take out of attending this free lesson and use it in your day-to-day work. Okay, so let's get started. Uh, by the way, I'll I'll be presenting today and any questions that you have about what I'm presenting and any questions about other other things that uh we we do at a analyst lab.ai, please u you know push push through your questions in the chat and we'll have Sean and I reply to you as well. Cool.

3:36 So, let's just get started. Okay, let's start with a quick poll. Can you answer to this question? You could type A, B, C, or D. No judgment at all based on what you all feel like. Give me a minute. I'm trying to I can't access the chat. So, uh, Sean and hi, if you see something, you could like, you know, respond to them. But >> a lot of C's. We got a lot of lot of C's. A lot of high performing experimentation folks.

4:07 >> Pretty cool. >> All C's. >> Oh, nice. That's so cool to know actually. Uh we generally try to cater our presentations to everyone because what we try to do here is something that get people who are not not as experienced you know people who are stakeholders not even data scientists how do they get on board into understanding experimentation and how could they uh you know get all the technical details and help from cloud code. So it's absolutely okay even if you are like an A or B uh or so or even D you know so there's no judgment at all but glad to know that most of you are C's uh but let's get started today we'll probably try to do the C that most of you all do but how do you do it with AI cool so let's I'm doing it the other way this time around I want to show you what is the final output that claude code gave me Uh, and let's then dive into the theory behind experimentation because I wanted to share with you all cool things.

5:13 At least it it was pretty cool when I tried to test it out. Let's see if you all can see this. Okay, so this is the prompt that we gave. So this is our repo that we work off pretty like you know detailed. We have a bunch of agents like we have a bunch of things going on. I'm in the repo that I have my data file. This is what I gave as a prompt analyze experiment and this is the CSV file that I used which contains streamflow data u of an AB experiment and using the experiment analyzer agent. This is the agent that we've built. uh and these are the inputs I gave it and I gave it some context around what business context around like you know what is ROI like what is the average lifetime value and you know things like that and it also has access to a bunch of uh like questions and analysis that you have to do when you are given like an experimentation question right and then um I ask it to you know do a final readout in a Google doc I literally paste this thing I don't do anything else you could actually also talk to it with whisper flow if you want to. But let's say I pasted this, right? I'm going to go through the entire demo with you if we have time at the end, but I wanted to show you the end result. So, I'm sharing with you a window. Let me stop sharing and share with you what the output looked like.

6:36 Where am I? Yeah. Okay. So, do you see my Google doc? This is the Google doc that it generated streamflow regional playlist experiment. This is like a demo output. As you see, it has the executive summary. It has question after question like a table with the details and like what happened to these metrics. What's reliability? How does it look like for deep dive segments and you know I don't want to get through the details because there's some interesting stuff that I want to share in the demo directly. But I wanted to share with you that this was the output it generated with one prompt. one single prompt in our repo and that's what we got. Isn't that cool? What do you all think? We'd love to know what you all think. Uh, and let me get started with my presentation back again. Cool. So, now that you saw the end result, um, I wanted to share with you how we got there. what are the multiple pieces that go into this uh like brain of how did it come up with that uh output what went into it and details so before that I want to share with you something so if you want to build this reusable system like one prompt in and an output out obviously you'll have to do some feedback and checking and a bunch of things uh but if you want to understand these how we build these agents and skills we have an AI analyst boot camp coming. It's a two-day beacon intensive program. Um, we have way more than experiments covered there. We also talk about funnels. We talk about root cause analysis, storytelling, like how do you actually get this to present in a not just a do Google doc guys, but also a deck like a presentation. I I have a demo of a presentation that it did as well. I'll share with you at the end.

8:31 But we cover all of that. And for the people who are attending this lightning lesson both live and also who have signed up, we have a code for you claw 20. Please try to use that and leverage this uh to get $120 off. This is our first attempt at doing this boot camp. I know a bunch of my friends and all the uh like you know people I know they're like you're giving this for way uh you know uh shorter am smaller amount than usual but we wanted to share with you all everything that we are learning and pretty excited to share that. The other thing that I want to share is we have an open repo. We have this uh in GitHub uh with like a analyst laban analyst. Uh we Sean Hi and myself share a bunch about this in our LinkedIn post as well.

9:18 Please try to you know get this open source repo clone it and see uh all the things it could do. Share share your feedback. Um what we teach in the analyst boot camp has a lot more things than what we have in the repo. We are actively building what we share for the boot camp. Another interesting thing as well, we have a Slack community. We currently have two free plus members. People growing by the day we people come in, ask questions. You could free join uh free to join anytime and we are there to help like you know navigate through any challenges that you're going through in this space. This is a QR code to get into the boot camp guys. Yeah. Cool. Now let's get started again about the problem.

10:00 So here's a common pattern we see with experiments. We turn it on, we wait, read the dashboard. Green means ship and red means kill. I know a bunch of you are already doing great things uh when you mention that you do the entire analysis. But I'll be honest from my experience, not everyone does the entire workflow or a life cycle of experiments need to be done. Especially when we want to ship things very fast and move fast.

10:27 That's when we keep we we see this happen a bunch as well. So when you do based on just PR green means ship red means kill it's almost like traffic light watching right. So the problem is that treating experiments as a single step instead of a life cycle. So today I'm going to show you the entire life cycle the different stages what are the common ones that kind of get skipped and then we'll demo what happens when the obvious ship it answer turns out to be wrong.

10:58 Cool. What's experiment life cycle? I think a bunch of you who are data scientists that work in the space probably already know it. But for the people who are stakeholders and are newer to tech, um I'd like to like you know share this in deeper detail. Right? So it starts with why uh like why are we experimenting in the first place and the design how do you want to design this experiment, right? And how what is the power like how long do you want to run it and how many users and how do you want to like once you start running it what are some things you need to check for and then once we reach our you know the goal of uh running it for a certain amount of time how do we analyze it and how do we what do we decide uh on do we ramp it kill iterate iterate on top of it and then how do we want to ramp it so we'll cover all of this let me know if I'm going too fast guys.

11:50 Uh, hi Roshan. Please interrupt me if there's something in the messages uh that you want me to stop. I'm currently directly diving in because there's so much content I'd like to share. >> You're good. We're managing the chat, but I'll interrupt you if uh we need to need to answer something directly. >> Okay. For a while already registered for the boot camp. >> Oh, that's awesome. Okay, cool. I'm kind of feeling like I'm giving a big monologue, guys. So, please uh uh you could interrupt me or share anything. So okay let's get uh back to this. So why do we experiment right? So products go through this life cycle. We first build and we want to measure what we build and then we learn right. So experiments basically um so this isn't new. It's the way the lean startup loop works. So experiments are the most critical ones that we discussing today are in the measure step. Without them, your PM probably would say, "Hey, I think user wants dark mode." And your designer would say, "Oh, onboarding is way too long." And maybe your CEO would come and say, "Just change the pricing, guys."

12:54 You know, this is uh uh like there's a lot we could do here. So, all of those are just opinions until we go test them. So, how do we go test them? We first start with design, right? So step one is basically design the experiment before you build anything. Start with a hypothesis and if you notice the structure we have something like we have this like we believe this change will impact this metric because of this reason the because forces you to articulate why you think this works.

13:30 That's the mechanism. Then pick one notar metric and a guardrail. The Nster is what you're optimizing for from this experiment and the guardrail is what you refuse to break. Um, and I want to give a quick plugin about we did a free lightning lesson on how do you define metrics on this. Hi actually ran it a few weeks ago. So if you're interested, please drop in the chat and you know Sean and Hi could respond. Also you could go to AI analystlab.ai.

14:00 there you could see all a bunch of free lessons that we've you know dealt uh we we've shared in this area and u a bunch of things that are coming in future as well. Okay, now coming back here in this uh lightning lesson today we're going to have a running example of a streamflow data set. So what is the experiment that we're going to have uh that we that will run through this entire lightning lesson? We believe showing regional playlists will increase weekly streams because users engage more with music that reflects their local culture. So that's the hypothesis. See if you look at it, we have a because right we think because of that it reflects the local culture and there'll be more people who will engage with it. So this is our statement and this is our experiment that we'll run with. And the northstar that we are aiming to improve is streams per user and the guardrail is churn rate. We don't want people to churn, right? We want them to engage. So we want to ensure that that that guardrail is met.

15:00 Cool. Now now that we have the the statement the hypothesis the because the reason we need to understand how long do we want to run this experiment for right and how many users do we want this experiment to be you know exposed to. So what is the effect size? So is 10% lift easy to detect? Uh so if if you look at it 10 person lift probably is slightly easier to detect when compared to the one person lift one person lift in like control to treatment would need way more data right uh specifically for a significant level it's usually around 5% that's an industry norm and power usually is around 80% that's another norm that we deal with. So for this analysis after the power analysis what we came up with is 6 weeks and 25k users. So what happens? Let's see if we skip this. Uh most of good teams don't skip it. But there are a bunch of teams that I've seen skip it where they think they have this knowledge already from the past experiments. They have the domain hypothesis. They and then they want to move fast. Let's say they blocked on the data scientist. They skip these things and what happens when you skip them is you run too short and you miss the real effects and there's a bunch of peing also that happens. Let's say on day three someone sees the significance and it's such big it's such big of an effect they're like oh let why do we wait on getting the impact on this experiment so let's ship it and then they ship noise because it's just day three right and the and another thing also happens when you don't do the power analysis and all of this the right way you basically can't tell if a null results or no means actually there's no effect or it doesn't mean you have you don't have enough data to to talk about it. Basically, there's not enough power to say you know something, you learned something, right? And like probably the data science folk in this meeting would know that peaking is like the most common way good experiments get ruined.

17:11 Okay, so this is a small sneak peek about what our existing free repo already has. So this is an experiment design pipeline that we have. uh it basically you know comes up with the hypothesis it u you could the power calculation and what exactly happens in the design and what are the guard rails and you know running happens right after this. So every good experiment follows the structure. The system autogenerates all of it once you give it enough context. Hypothesis, power analysis, randomization of design, guard rails and the full pipeline including the decision rules uh also is part of like the the next stages right so today we are picking up after once the design is done how would the experiment run. So that would be part of my demo but I want to share you a sneak peek of what would app what would be if you run it the design skill that we have in the design repo today. So these are literally the charts that it created for the experiment that we are going to talk about today. So on the left you have a feasibility curve.

18:16 What does it show? It basically shows you how many weeks you need to detect different effect sizes. A 5% lift needs around 6 weeks. At 3% lift, we need 16 weeks. Wow, that's in the patient zone, right? And on the right, you have a power sensitivity matrix. Each cell shows you shows basically your statistical power at different combinations of FX size and sample size. The green boundary is where your test becomes reliable. So 80% or power or above basically. So notice that at 12.5k users per arm and a 5% effect we hit exactly 80%. That's why the streamflow experiment if you remember that I gave you this number before in the design basically talks about uh having the experiment run for uh 25k users over 6 weeks and all these to mention it again these charts are autogenerated by the system that we have.

19:18 Cool. Yeah. So now let's talk about running and deciding. So what do we want to do once uh like you know we basically are condensing two steps in the stages here to explain you the theory. We'll go through this in detail if we have time in the demo. So we are condensing two stages here. While the experiment runs check for the SRM right sample ratio mismatch is the randomization clean or not. uh you probably need to watch out for the novelty effects early lift because of you know it's an exciting new feature probably right so there'll a bunch of clicks come from that but not all of that stays the same so early lift u fades if it's not real and when it's time to decide you ship kill or iterate it right and then remember that the statistical significant does not always mean practical significance uh basically a tiny lift that's significant for us it feels like experiment is a win but it not it might not be worth shipping because it adds complexity this is something I deal with day in and day out in my current job in my previous in all my experience I'm sure a bunch of you who are from data science also deal with it a bunch that what is you know good enough to ship even though if it's like a win and when you do ship not every experiment have to go through this 5% 25 50 100% but the ones that affect the your ecosystem of product if it changes a bunch you probably go through that but another important thing that I see a bunch of you know teams miss is having a hold out so ensuring that you ramp gradually especially in case of important like you know UI changes or something that affects users um you basically try to uh ram gradually and you always keep a hold out. So the experiment doesn't end when you just say ship it.

21:22 It doesn't mean that you're done with it. Yeah. So this is what covers what I just mentioned. So you monitor it at each stage especially for the important changes. You basically ramp gradually and you monitor right and you understand uh and you also add a hold out to understand the long-term effects. There are like a bunch of experiments that actually show great results probably at a later stage. For example, I have experience in working with engagement of the product.

21:55 Let's say top of funnel is engagement and bottom of funnel is revenue that you get from these users. Unless you have like metrics like lifetime value which predict the revenue that comes out of these users, it's pretty hard for you to see engagement show an impact in revenue in the experiment time frame. But guess what? If you add all these engagementdriven experiments that are wins to your hold out, what you'd see in over 6 months of time is that you actually see a great improvement in revenue coming through.

22:25 And imagine if you have a hold out uh for like an year or so let's say in like a in like a scenario of you have it for 5 years for some and to understand I'm the long-term effects of improving engagement is going to be so much higher on your on your revenue but you cannot understand that in your two week 3 week or maybe four or five weeks experiment standpoint and that's the reason it's very important to have a hold out >> hey sorry can I interrupt real quick. Is the uh so I know the design uh the design experiment skill is in the repo.

23:02 Is the uh full experiment workflow in the repo as well or is that something you've been iterating on your on your own? >> So I the full experiment thing is in is something I'm iterating on. I'm planning to add it for the boot camp. Uh so currently we have the design one though. So you could actually play with it yourself and see what else you know need to be done to it. So the the repo itself is pretty smart and intelligent. It has a bunch of things in itself. So um yeah, I'll probably talk about it more when I go through the entire like demo.

23:34 Hopefully we have time. I'm trying to go super fast to get to the demo. We'll see. Awesome. Yeah, thanks Sean. Yeah, please interrupt me like this. Cool. Okay, so these are like a bunch of these are like eight questions that I see like you know separate like a good analysis from grade. So not every experiment needs all aid I'll be honest but knowing what to ask means you won't miss those things that matter right so the one thing is basically understanding if the experiment set up correctly did the did the treatment move the metric like what is the reliability what are the so one thing that I see people miss when they're you know in a hurry to ship something before a quarter end because they want like wins out is what are the learnings that you have from the experiment across segments These are something that you wouldn't know unless you dig deeper into understanding oh why what happened for probably uh users that are newer when compared to the matured users you know like asking these questions to look through multiple segments would really help you not just to not just for that experiment related you know decision-m but also all the new iterations that you want from your experiments like that you have in the quarter that are in the backlog that you want to run with your product team, product dev and engineering team. Okay. So, um today we basically are going through all of these through the streamflow data set. I'll actually give you exact prompts to do it yourself in Cloud Web UI and you could we'll have like a one-stop shop prompt and we'll have like multiple pro prompts for every each of these stages as well.

25:23 Okay, time for a live demo. Cool. So, let me just stop sharing here and share my screen and get started. Where am I? Okay, give me one minute if you're trying to Okay, let's go back to the demo. Let's say we have our claude UI right here. Okay. And what you're trying to do here is I have one that I already ran it. If for some reason we're going to have issues with my internet or I don't know cloud usage, we can look back at what I ran it. But let's do this live guys. We have some time. Uh so I have the data set. This is the data set we've generated for this. And I'm going to give this prompt.

26:19 So okay. So what am I doing here? I basically gave it a prompt that I've uploaded a data set. I gave it the context of what the data set is. I gave it what the columns are, what each columns almost like a data dictionary, right? Uh what each columns are and uh I tell it like you know what should be the step one be like what do you look for? uh like sample review mis mismatch the pre-experiment balance and the duration check what's the date range how many weeks you want to run it and stuff let's see what it comes up with one thing I want to let you all know is this is something that I'm facing right now in my company when I'm trying to do things the column names so I didn't give it what are the re what are like each of the columns mean here right for example if you don't give like you don't have a data dictionary of what each column means What claude or like maybe any other uh any other like you know LLM does is it it understands based on what a column name looks like. So we basically u have this issue where let's say there's a column name called anonymous. It takes that into context of what it knows and just takes it as like comes up with its own assumption of what a column name could mean and goes and run with the data. So be very vigilant when you give column names without what the column names mean. In this example, it's pretty obvious. I kind of created uh the data set for the simplicity. I made it so that each column names actually mean what the name uh could you know what you could assume based on the name. But if you're trying to do this at your work, ensure that you also give what each column actually means because not all column names are pretty descriptive. Now let's see what did it give me. So it went through the step one the sample ratio mismatch. It did uh it basically looked at the split and it gave me a pass the verdict. It's pretty cool. No evidence of a randomization or a logging bug. Um and then it basically looked for continuous coariantss. It came up with what are the categorical coariantss and like proportions are virtually identical across arms. made sure that what we are looking at is a pretty balanced data set. And then it came up with a duration check. The experiment spans for six week cohorts and it gave me from when to when and like it also identified a small concern about the week five weeks five and six and basically this is what it says like it it's something that uh this is like worth investigating in step two. So that's what we'll go ahead and do in the step two. How much time? Okay, we have 30 more minutes. I'll try to do the claw web UI uh like demo for the next 5 10 minutes and I'll run in another demo and go through how my you know workflow looks, what is something that I've built for this lightning lesson for experimentation uh in the cloud code after this. Cool. So it's currently thinking through what uh basically my prompt is that I'm asking it to do the treatment effect and also the reliability. I'm telling it what test to run. What is the experimentation t test I want it I want to run uh and you know uh how I want it to be displayed and I also talk about the effect size. Can you give me the effect size and look and also I talk about like what is the post hop power like do you uh give the observed effect size and sample sizes what is the power we achieved I also talk about minimum detectable effect what's the smallest lift we could have detected at 80% power and I also ask it to give me a one sentence summary cool look at what it's giving me uh exactly what uh I mentioned to you guys that I asked Er um and yep so look at the one sentence summary the treatment increased by 14%. Right which is statistically significant. So guess what looks like we have a winner here guys we have our not like we have the metric increased by 14% which is amazing. So there is something that you would see that most of the people would probably go there are a good chunk of people that would probably tell and announce that hey we have a winner because we ran it for the amount of time we wanted to run it. So they did their diligence they didn't peak probably and they got a winner.

31:04 Let's look at the step three and do the segment analysis and think through and look through if what we have is actually a winner or not. So this is a longer one because I'm asking it to look through and break results by all my important dimensions. So the dimensions here are going to be user type if it's an existing user or a new user and I give it five values for the re. So the region has five values and we also have device mobile, mobile or desktop and I ask it for each segment do these same tests. Do can you check for Simpson paradox? Can you check do the guardrail checks and for any degraded segment look for our guardrail which is the churn or not?

31:51 Let's see what it comes up with. It probably might take a bunch of time because this I would say is like the more detailed. Oh, this is pretty fast. Nice. So, look at what we have. So, the lift is brought across most segments, but guess what? It found the guardrail to be degraded for new users. fine the trade but the the trade-off is that existing user churn dropped from them but not as much but for the regions it's pretty fine so the churn treatment users have essentially the same as churn control so we're not loosely losing like heavy or lightweight streamers right so and the bottom line is that the regional playlist feature is a clear win for existing users but is a retention risk for new users the exact population streamflow needs to Hello.

32:48 So if imagine we went to step two and we saw this win and people didn't bother to look through segments of you know like the critical user types we wouldn't have understood that this is what the this is what was the issue right so for the new users it was a retention risk. So now let's dive deep into what now that we know this what do we want to do we basically if you remember going back to the theory that I talked to you about it's not just statistical significance but also practical significance we need to understand the ROI what is the impact of running this experiment and what is the you know return that we're going to get with it so let's do the step four how much time do we have? Okay, I know I'm running fast. Anything uh Sean and Hi you want to share based on what's happening in the chat while this runs?

33:49 >> Uh I think you're good to go. Um there's a lot of questions. They're very desperate. We'll get to the things that we can't answer in chat uh in in a bit in the Q&A. >> Oh, makes sense. Awesome. Sorry guys, I'm trying to run as fast as I can so that I could go through two demos with you. the cloud web UI demo and also the cloud code demo. I want to share all those details and also if possible have at least 15 minutes of Q&A at the end which means we have 10 11 minutes now.

34:18 Cool. If this is going to take time I actually have this run for us. How does this look like? Step four. Let's see if this is still running. Yeah. And another important thing that I wanted to share with you all is that if there are audience like who are listening to us today if you do not understand the details of this because the I I saw a bunch of non- tech folk people who are in like you know our stakeholders from design for product also join we have uh another course called AI analytics for builders that's basically a course where we'll go through it's not a weekend boot camp it's going to be a six week long course where we talk through every detail of how do we set experiments, how what are the metrics and all of that with claw code. So almost like a an enhanced version of the boot camp but that covers 6 weeks long. Yeah, you could look at all of that in a list.ai as well.

35:26 Okay, so it's actually creating the chart. So when I ran it to make sure that I have something for you guys, I didn't do the charts here. So for the demo, I was I was thinking it'll be cool to get a chart and we did get a chart. That's pretty awesome. So this is something literally I'm doing right now with you guys because I was like, let's have a thing with a chart. So I think this is something you can copy paste. Yeah, copy image. How cool is that?

35:54 Okay, so what does this say? So it's basically says that the churn feature reduced churn among existing users saving uh like and then 53,700 users worth 11.8 million in DB but if we shipped to new users the chance spike would burn like so many. So yeah uh this is what we would understand uh what happens when we look at things by segments right and also understand the lifetime value and ROI here. So the count so the conservative scenario talks about deliberately zeros out the turn benefit and still res like still returns 21x the infrastructure cost. So even under the most cautious assumptions this is a strong share for existing users. Yep.

36:45 Okay. Uh so let's go to the step five which is the final prompt that we have. So I'm doing it step after step so that you go through this process with me in for multiple questions but we could have one single one-stop prompt as well that does this entire thing for you and gives you the docs and you know the uh give you the deck and the dock as well. So when we share these results with you over an email, we will basically share with you the every step-by-step prompt and we'll also share with you the detailed uh like you know one-stop shop prompt as well.

37:32 Yeah, look at this. Uh it basically uh talks about per segment decisions. What do we what do we ship? What do we not ship? What do we iterate on? What do we ship? I I'll be honest generally like in real time it wouldn't be as feasible for you to ship like specific areas for example the Pacific region right so you want to do something at the cost of not maintaining it so this is where I would say you would use this agentic system to get to you to a place where it does 80% probably 85 90% of your work but there is definitely 10 15 20% of the work that you still need to too because not unless and until you give it the entire context which probably the more you work with it the more context it gets and the better it gets at giving you these decisions but there could be some reasons that hey shipping it non to non-pacific regions is something not possible for infrastructure we possibly want to you know do the entire ship and then iterate on it right because we don't we might not want like super custom product features as well so that is something that you as a you know data person or a product person or like you know stakeholder would make a call based on all the information that this provides.

38:50 So yeah look at the net annual impact the retention LTV saved and the net ROI. So this is the exhibit executive summary streamflow's regional playlist experiment is a clear win with one critical caveat and it gives like all the details that I shared in the earlier prompts. So in this prompt I actually did not uh share it to give me a Google doc or like a doc but I did in my previous prompt. So look at this. I actually asked it to give me a doc and claude web generated dog guys. It's not a I thought this pretty cool. Look at this.

39:23 It's actually a very and it's something that you could download. It's in a docs that you can edit and share. So um I will share this other prompt as well. So I I wanted to share more detailed prompt and hence I did this for table structure but it could also give you a actual doc that is you know that looks actually pretty good directly shipw worthy. I'm sure you have to do a bunch of changes um based on what you think is the right context but it's it's in a pretty good place already. Okay. So let me stop sharing this and go for clawed code.

40:05 uh prompt called cloud code demo. Okay. Anything quickly, hi or Sean, that you'd want to tell before I jump to the cloud code? >> Uh I'm going to paste in some upcoming uh lessons as well that might be interesting to this group. um as you're pulling up cla code shavia um yeah would be all free and uh they happen over the next uh week or two and uh you know it's also about claude code and just drop it in there definitely register and you'll get the recording afterwards.

40:40 >> Awesome. I I'm generally not seeing any messages guys so sorry if you have questions for me but I quickly saw one message before I'm I'm sharing my screen about did Claude have context about LTV. Yes it did. If people we uh if I didn't give it the context about the uh RO like you know um the ROI or average LTVs it would you know take a bunch of things itself and create scenarios if your average LV was so much this is what you would get if it was so much this is what you would get but I for this demo I gave it the context okay now let's go to cloud code so uh for the people who were at the start of my um of the lightning lesson today. You might have seen me given this context to this um to this cloud code.

41:32 Um I sorry give this prompt to the cloud code and we have a bunch of repos like you know all our boot camp repos here all our a analytics for builders the course that we have we have all of this like in-depth knowledge that it could use to get this information. So let's say this is the analysis. This is the prompt. I'm running I'm running it again. So that oh give me a minute. I actually wanted to run it and ask it to give me something. So it has the prompt to give me the entire and I want to ask it can you share with me a plan before you execute.

42:15 Let's see what this is going to give me. So, it's basically going to look through all of this. Uh, so I just gave it a permission. You could also give it uh like, you know, go to settings and permissions and change that guys to make sure it doesn't ask you for permission every time, but I tried to do it for some things because I want to be a gatekeeper for some permissions. Look at the prompt. Look at its plan that it's generating. So everything that I showed you uh in the cloud web UI it has all of this in its uh like you know in bank on how to generate this analysis. I'll go through this but I want you to check out the prompt that I gave. So analyze experiment in streamlow experiment csv using the experiment analyzer agent. So this is the agent I created to showcase with you all how do we do experiment analysis in cloud code and uh the inputs and all of this and I also talked to it about using another agent at the end then export the final readout as a Google doc using Google doc creator agent and also I talked to it about after the eight question analysis it completes pass the results to experiment readout agent. So I have a bunch of agents and skills that are inbuilt in my system that it uses to come up with an answer to this question I asked and this is the business context that I gave guys like someone asked this question I briefly you saw. Okay. So it actually goes through this entire prompt. This is the plan that it created. This is the data overview. This is the stage one like what it does like all the steps the questions it goes through. uh it actually creates an output for each of these in this. So if you're interested, you could go to this working direct, you could go to the directory like I have my lightning lesson 5 here. So it puts all of my stuff here. Um and stage two, it does the experiment read out. It does this then it does the Google in stage three it does the Google doc export.

44:17 Let's do something to make it easy for you guys to understand. I'll ask it um let me actually talk to it. Can you give me an asky diagram of your entire workflow? So this is what I do most of the times. I literally talk to it, not even type with whisper flow. So this is what it is doing. It takes all this uh knowledge that we built into the experiment analyzer agent and goes through all these questions that I put as part of the agent when to do what for each combo like what do you do when do you pass when do you block or halt and like because you don't want it to run through all of this when you have issues at the SRM itself right so there this is the stage to once we have all of these how do we do the readout so ensure that you generate visualizations if you see we have a bunch of helper files that we have a bunch a good chunk of this I would say 60% of this or 50% of this is already part of our free repo guys so please check out the free repo it's already in u like something that we that Sean and I are sharing and also it's an AI analyst labai as well but I did create a bunch of agents for this lesson specifically. uh I worked over the last weekend for this um to basically understand how do you get these questions but also do the read out in a way that you could literally um share it with your stakeholder right so it reads and synthesizes it's design it's the story board how do you take the context what do you what is the tension and what's the resolution this this narrative is something that we build into our skills and agents and we have different charts and we write the narrative we build the ramp plan and the final one is the Google doc. As simple as we think Google doc is not I mean once we have an agent in skill is probably easy but it's not uh it was not super easy the first time you did it because it does not have as um easy access to where an image sits and where like the text. So you need to work with it to ensure the images and text have you know a gap with each other and the formatting is right. So this is what it is. And let me share with you the final output once again.

46:52 Yep. Sharing. Oh, wait. I'm sharing the wrong one. Let me share the right one. Yeah. Okay, cool. So this was the doc that it shared. If you look at what we've done in the cloud web UI, it's literally the same thing in a Google doc format. I did Google Docs because most of the tech companies I know in Silicon Valley use Google uh uh like you know Google uh workspace. So it basically did this. It did all the formatting, the colors, the you know the original one was not that pretty. It made it pretty because we worked with it to make it pretty. um the charts like what exactly happened with new users and what's in Pacific region and there's a Simpsons paradox that's what it talks about and these are the charts for the guardrail alert and was the experiment run long enough and what is the ROI impact and it talks about the existing users shipping and the new users do not ship and another cool thing is it actually created the deck too I didn't get a uh good enough trying to work on work. So this we also have a Google slide deck uh export uh agent and also reviewer agent but I don't have as as much um as many interesting things on top of it yet.

48:20 That's something I'm going to work over this next week but yeah this is already what it generated not too bad um it basically talks about what like all validity checks being passed what's the primary metric and how does the segment reveal. So look at the segment reveal. So it just did uh like it had some narrative built in. It talks about Simpsons paradox like you know what it gave and lift by region uh churn guard rail like weekly lift trend and finally the business impact and it also gives recommendation.

49:00 Again all of this is nothing I didn't do any code. I didn't do anything. In fact, I spoke to Claude code because we have the system built in the agents and skills. That's the reason why um like we got to this place. Okay, I think that's where my demo ends guys. We have around 11 minutes. We have a hard stop at one unfortunately. So, 11 minutes to take questions.

49:35 Hey Matt. Well, good to see you, Matt. >> Hello. Um, quick question about like accuracy. You seem to be, you know, iterating through this very quickly and just trusting it. Like if let's say you were responsible to, you know, give a report to the CEO and your job was on the line, how would you change the workflow? if you weren't giving a demo. >> I can take it. I can take it, Savia. >> Yeah. >> So, here's here's my philosophy around anyone using cloud code for analysis.

50:08 You're accountable. The human is accountable for the results of AI. No one gets to go to the CEO and say, "Oh, the AI screwed it up." Just like any analyst, like if I gave my results to my manager and he reported that out without taking a look at it, like that's on him, too. So there's many ways to validate this. One thing I like to do is go through and run it on a bunch of back tests of analysis of what I ran in the past. See, you can just even like one off do this yourself. I have it uh basically also take any of the the Python and SQL and I have it put it into a notion document with the results. So I can validate that there too. You can also build like extensive AI evaluation uh kind of metrics around the validity of this too. So if you have like 300 ground truth experiments you've ran in the past whatever year or so or two years, you can run this on there and do some sort of like auto research thing to optimize the the like performance of that in terms of like how accurate the results of this were versus like the human ones back there. So definitely like there's ways to have structured validation like any AI product. Um but at the end of the day I think like you have a very good point like don't just trust it and give it to the CEO like cuz it is your job and you got to like be accountable for that at this point.

51:34 >> Yeah. >> Does that make sense? >> I one thing uh I'd like to add uh to what Sean said, Matt, is that what what's the intent? Like what's the goal of this lightning lesson? Right? The lightning lesson goal is not that hey take this AI agentic system and you know you don't have to work anymore or you know or just you let it work and you give it directly the output not at all what uh the goal is it is pretty smart and intelligent and we could make it do half or way more than half of the work that you basically let it do the work and you decide on what you like or what you have problems with it and work with it, give it the feedback and let it run.

52:17 I agree, Matt. This was a demo. I know it I probably did give it uh like something that I could share with you all, but that does not mean that this is always going to do some great things, right? There is a bunch of things. I'm running it in my company right now, and there are so many things that it also gets wrong. But guess what? the more I work with it, the more I get the skills and agents to a good place, the better it gets in the next iteration. And I'll be honest, I'm banking on the newer model that's going to come out in like maybe 2 3 months. And I don't probably need to work with it as much as I worked with it now to get to to a good place that I could present my findings to CEO directly. Right. Yeah.

53:02 Uh so Joy uh any sorry one one thing ma'am any anything else you'd like to ask or are we good we are we could move >> that's good for now thank you >> okay awesome yeah >> yeah Shabia I had a quick question so um could you maybe share which agent is doing what work in the different steps of the experimentation because I saw you had like multiple agents right so if we could get a quick walk through of which is tasked with doing what in the entire clot whole process.

53:32 >> Yeah, we only have six minutes. So, I'm wondering if I start doing that, it'll probably take the entire time. What I can do is probably in a followup after this lesson, I can try to give some, you know, summary so that you could learn from it. And uh I I'll also be honest, I'm not sure if we could cover all of that in a summary in like even in the even if I take six minutes. So that's the reason we plan to do all of that in a lot of detail in the boot camp and also the a the six week long course where we get into a lot lot of depth. So if you're interested please check them out but we'll also try to uh give you stuff uh for free or in links as much as possible. Okay, awesome.

54:14 >> Davia, there is a really good question from Maxine on what's your perspective on someone junior who hasn't yet built the knowledge or skills to fully validate outputs themselves? Do you think the level of reliance on clot and should be limited or more should more emphasis be placed on learning through feedback from others? >> Oh, I love it. It's I I think this is all the problems that we are dealing with as data leaders in this space like this is something that I'm thinking through literally myself. if I have a team of like you know uh 12 to 15 people like the uh and I I am it's the problem of the hour that we are trying to solve um with evolving AI to me I think you've had it in your question itself I would have accountability on area owners or metric owners or like the the tech leads of the space to ensure that the skills and agents and the context that's built into it is reliable, is up todate, and is something that people could, you know, check and review with the lead before they output something, especially the younger the newer ones until we get it to a place where, you know, the let's say we get it to a place where 95 99% of the reviews always are passing with the tech leads or with the area leads of that area of that metric, then yes, maybe we are getting into the newer system or a new model where it's almost getting it right most of the time with the context we provide. But uh that said it is also very interesting that the recent the newer grads or the recent engineers have way more um flexibility and adaptability to work with AI. So that's I think a great uh thing for students and people who are passing out or who are like newly in this field that they could leverage all these systems to understand and work with how how to work with it to give feedback while relying on the subject matter experts and the area tech leads of data or you know engineers or whatever to ensure that whatever they produce goes through their review.

56:33 Greg. >> Hey guys. Um, could you touch again on how you have Claude uh learn from any mistakes you find when you're double-checking its work? >> Yeah. Uh, you talk to it is I would say the most simplest thing but uh I I mean there multiple things that you u understand when you ask the same question Greg to it. You basically tell it this is a clear discrepancy. I would not expect you to make this mistake.

57:02 what do you need to be corrected in your system to make this happen right so there is a uh clot itself tells you oh I basically don't know this gap and this is why I made this mistake it's it actually tells us where it goes wrong as well along with you yourself identify right because we are the area experts of whenever we ask the question so I would thoroughly encourage people when you work with these systems to work on the areas that you entirely know the nittygritty details about because Only when you know all the nitty-gritty details, you could find these flaws that Claude itself probably can't find or ignores or jumps between loops and tries to assumed things, right? So I I would say half of it that claude itself once you point out that hey this is wrong.

57:50 Why did you get this wrong? gives like a uh like a you know reason of reasoning for why it made that decision and you understand why it made that in that reasoning and a part of it is what you would identify as part of when you look through what did you do to get here and when you look through that workflow you can understand that oh this is probably the reason why it got wrong and you talk to it and ask it as well. Yeah, >> there there's also um I I'll I'll send you this article, Greg. There's this thing that people are doing from Carpathy called uh auto research, which is really cool to like automate some of this where it'll basically it only works if you have a bunch of ground truth to test on. But like I was saying, if you have like a bunch of historic experiments basically like it'll train on the experiments and then it'll try it's an attempt at the experiment and then it'll look at like the historic thing of like oh what's all the stuff I did wrong and actually do its own error analysis that's like by different dimensions of what it got wrong and then it'll feed that back into the agent and say like rebuild your system like whether it's a knowledge base, skills or agents um and then now try it again on another test experiment.

58:58 And you have to create some metric like you could have some like similarity metric or some sort of metric of accuracy and it just like loops like that in a like forever process until it optimizes that metric. Um I've been doing that for a few different systems I've been building out and that's like a really cool way to get it to bake in all of its learning back into the system. Um I'll send that to you. It's like a pretty new thing.

59:22 >> Awesome. Yeah, I think I I was even thinking about just a s maybe simpler tactical technique like how to how to have it correct its own mistakes. So it's equivalent like taking your dog your puppy and rubbing its nose and you know correct it to reward it and punish it when it so it doesn't do the same mistake twice. >> Yeah. A quick answer to that is uh you you you ask it to think about what why it made the mistake and don't do it again and then it's generally pretty good and then you do it incrementally.

59:50 >> Yeah. >> Okay. We are almost at time guys. Uh unfortunately we have a hard but >> yeah I think we're out of time. Um so for people who are looking to ask more follow-up questions I think there's a bunch of engagement here. Thank you all so much. Uh definitely hop into our Slack channel. Uh we'll have you know like if you post those questions we'll be able to answer them uh over there. Uh otherwise we'll send out all the links and stuff.

60:16 >> Awesome. It was amazing guys. Thank you so much for joining joining ours our lightning lesson. Any questions we are always we are all three active on LinkedIn. So you know message us uh comment on our posts what else you need and check out check out more on AI analystlab.ai. Okay. >> Very good. >> Bye. >> Thank you.

Summary

Stravia Maripali and her team from Superhum led a session on running and analyzing experiments using cloud code, emphasizing the importance of a structured experimentation lifecycle. They demonstrated how to design experiments, analyze results, and utilize AI tools to streamline the process, while also addressing common pitfalls and the need for validation in AI outputs.

- Introductions included team members with backgrounds in data science from companies like Microsoft and LinkedIn.
- The session focused on the lifecycle of experiments, including design, execution, and analysis.
- Key concepts discussed included hypothesis formulation, power analysis, and the importance of guardrails in experimentation.
- A live demo showcased how to use cloud code for analyzing experiment data, generating reports, and visualizations.
- The team highlighted the significance of segment analysis to identify potential risks in user retention.
- They introduced an AI analyst boot camp and a Slack community for ongoing learning and support.
- Participants were encouraged to validate AI outputs and maintain accountability for results.
- The session concluded with a Q&A, addressing concerns about the reliability of AI-generated insights and the importance of human oversight.
© transcribe · For agents Built with care and craft by Gokul Rajaram