Transcript
0:00 one. Um yeah, mostly in data science for the past 10 years in the past 2 years been in more of the AI eval's agentic analytics kind of space rather than like kind of classic product data science. Um yeah, Hi, you want to introduce yourself? >> Yeah. Hey everyone, my name is Hi. I work at a legal tech company. Um I lead the data team there. Uh I have roughly 20 years of experience in the data science and analytics space. Uh one of the co-founders of AI Analyst Lab and uh yeah, excited to excited to be here today with you all.
0:37 Shravya, you want to go? >> Yeah, hello everyone. I'm Shravya. Um pretty excited to have the session get started with you all. That Sean's running it. I am co-founder with Hi and Sean for AI Analyst Lab. I have around 15 years of experience in data science starting with Microsoft and most recently at Superhuman. Yep, let's get started. This is uh one of the most awaited uh workshops I could say. >> All right, I got to stop sharing and reshare for a second cuz I got to get my my screens in order.
1:10 But um yeah, today what we're going to talk about is agentic experimentation. Um specifically experiment agentic experimentation um design and analysis. But uh we'll go through a little bit around um we'll kind of talk a bit about like why this is important. I'll give you a quick preview of like kind of like the system that we've been playing around with outputs. Um we're going to share a repo with you uh maybe towards the end or while we're doing it. We'll definitely send it in the emails. Open source repo. It's just something we've been kind of experimenting with around um agentic experimentation. I'm going to say experiment a lot today probably. So this isn't like some of my totally production ready thing, but it is I think kind of uh a way we've been learning and we're just trying to share stuff as as fast as possible as we play with it. And I do think there's a lot of opportunity in space.
2:09 Um we'll talk about like where we think like AI can like help when it comes to experimentation, where probably not going to help yet. Um where the human really still needs evolve. We're going to go through a little tour of like how we're kind of building things out a little bit. Uh we'll do a quick demo. The demo to actually run live, I'm going to run it live, but it's like it kind of takes a while to do all these tests.
2:36 Um so, we'll see how far we get. We might not get to the analysis stage, but we'll probably get through design at least. We'll So, we'll get into cloud code to show you that. And then yeah, obviously we'll share the repo out with anyone so you can go try it at your job. Um and probably make it way better than than we have. So, um kind of start with a some of the challenges of experimentation. Like now I've worked with a lot of teams who ran experiments for it with.
3:03 Um you know, uh teams that have like, you know, like 40 million weekly active users. I've worked with teams that have like a few hundred weekly active users. Um in terms of, you know, different products and uh the teams I've worked with that run tests, well, badly I would say, it's not really for a lack of like tooling or because they're not smart enough or rigorous enough. I think a lot of the time I see tests being ran badly is because the human who's reading the result is the same human who wants a particular result. So, like we ship a feature we believe in. We watch our whatever dashboard or notebook, whatever we have this our our our results um populating. And you know, a lot of times very easy if the number looks good to start building a story around why it's good. It's really easy for us to kind of like move our threshold, our definition of good also, depending on where the result landed. It's really easy to peek early and stop when we're winning.
4:10 It's really easy to explain away like guardrails that kind of broke through the thresholds we had tried to predefine. Um I don't think it's like kind of like a dishonesty thing. I think it's just like how our brains work when we see a number on a screen that we care about on a screen. Um So, I'll show you a little bit before we'll we'll get back into that, but just to show you a little bit about kind of like the output of like the the we're kind of generating with the system just so you can kind of get a preview cuz there's going to be a bit of slides here as we get through go through kind of like what it does. So, the output that kind of generates is like um three parts. It's like your experiment design brief. So, like everything before you do and see any sort of data, uh run any sort of test. Um And so, you can see, you know, uh it has your your whole um your your hypothesis, why you think it's going to work, um the metrics you're going to use, the uh the the effects that you can detect given your sample size, um what the exposure is going to look like, um and then it has some stuff around here where we have like snapshots and dates and ways of tracking this over time so we can go back and audit it. Um The other There's three core documents.
5:32 The other core document is our SRM brief. So, our our our um sample ratio mismatch. Basically, this is something you want to do again before you look at any results. This is like, "Hey, is the treatment and control like evenly kind of distributed. And then the last one is going to be your um actual analysis brief here. So, just to give you an idea of kind of like what the output looks like and um what's in there. So, in this case uh that the analysis brief here uh you know, has like your verdict, um your lift, your confidence interval, um every claim sites like what file and step it came from. Um so, we'll watch we'll go through how this is all built in a genetic system and then we will see how far we get running through the demo and then we'll also have some time for questions.
6:28 Okay, but that's just kind of the output so you don't have to wait till the end to see it. Um and then it's Shravya and hi if there's any questions that come in the chat. Feel free to to let me know. I don't know where I put my chat. Where the heck is it? Oh, there it is. Um okay. So, um one of the things like our kind of like thought process around this too isn't that like like genetic systems aren't solving the math wars. Like the math is not really that hard. People know how to run um a T-test. There's there's Python functions and libraries for all this stuff. Um there's like like dashboards work but a lot of the time like experiments like the decision that is made based off the results of experiments still come can tend to come out from like whoever the loudest person in the room actually is and what they want to see out of it. Um so, that's not like a stats problem.
7:35 That's a uh that's a who is reading the result problem. And that's kind of what we're trying to solve here. So, there's four kind of things that we see that and I'm sure you've all seen it before in your work that can really wreck an experiment. One is when folks kind of move their decision threshold to wherever the result land is. So, you know, we run a test and then after you after you've seen it like maybe maybe you wanted a 2% lift. That's what we set our minimum detectable effect at. That was our goal from the beginning, but then we see at the end that it's like a 1.6% and now suddenly we in our minds have changed our decision threshold to like a 1.5% being the the bar to moving forward.
8:20 It's really easy to do that. Um two you guys you all know about this peaking early um stopping when we're winning, right? So, once we see a lift, it's really hard to unsee it. It's really hard to unsee it. It's really hard to unsee it when your stakeholders have seen it, too. So, um it's also a lot easier when you're under pressure to roll stuff out to like just roll it out at that lift and tell yourself hey like we're being we're being fast, we're being efficient. Um number three explaining away guardrails that broke. Like you can explain away anything. Like everything has a story.
8:56 Like you can spin a story easy. Um so, like explain the guardrails away based on something else um or even even kind of like saying we'll deal with that later, right? Like oh, latency went up. Um but the primary metric moved in the right direction. Oh, we'll fix latency in the next sprint and then yeah, that never really happens. Or maybe it does. Um and then the last one we hear is like changing the metric that's primary.
9:22 Actually, I've seen this a lot. So, like the primary the one we set as primary is like flat. Maybe it's not negative, it's flat. But then we see some other like secondary metrics that go up and suddenly we're like oh, these ones look great. Those are our primary metrics after the fact." Um this is really dangerous, of course, when you have like dashboards that show every single metric that anyone can see. So, every one of these like there's some story to being told after the data arrived. And that story being told rewrites the question we are originally asking. Um and I I don't think it's like lying or dishonesty really. I think it's just like what brains do, especially when you put a bunch of hard work into to doing something.
10:09 So, here's where I think um this stuff actually breaks. Um so, there's kind of like I say three kind of stages at a very, very high level of uh experiments, running experiments, right? There's like the design stage. Um so, you write a hypothesis, you pick a metric, you do you set a threshold, uh you decide the guardrails. Um if you're like loosey-goosey with this stuff, you're basically just like sandbagging or hedging, then it's like you've already made up your decision around what it's going to be if you're not strict at this stage.
10:46 Um instrumentation, that's like more like, you know, like the engineering work, so like the code that assigns variance, um that makes sure the events are firing, the day the test um is actually running. It's pretty straightforward engineering work. Um it's high consequence though, obviously, if you mess it up, but it's pretty procedural. And then analysis, that's reading the data, computing the lift, um deciding to ship or not ship. Um writing the readout.
11:17 And this is where those four habits in the last slide tend to show up, right? After the fact, once you've seen the data. Um where that bias comes in. Uh so, I think where humans like kind of get squeezed and where humans um is is obviously that third one. Um, that's where I think the bias kind of comes in. The middle one is kind of like just engineering. Today, I'm going to talk about the system that we've been building that puts guardrails on more on stages one and three. So, we're deliberately not touching the middle. Like if if we have some agentic system that messes instrumentation up, um, your data is completely wrong and no analytic framework is going to save that.
12:02 Um, that's just like not really like a risk profile we're taking in right now. But, I think in the experiment design stage and experiment analysis stage, um, there's like actually a lot of guardrails we can apply there. Um, so, like platforms that lots of experimentation platforms exist today, they compute uh, statistics correctly. Um, they're going to give you like correct lifts, correct confidence in intervals. Um, but they're going to leave that judgement entirely to you or to your team at exactly the the moment that your judgement is most compromised.
12:41 You are seeing a result after you spend a bunch of time working on something and you really want a certain result and they're just going to like say like here's the numbers, make up make the call. It's like there's just so much opportunity right there to like persuade yourself even subconsciously into making the wrong call. So, those are those all these platforms are more like calculators, right? A calculator can't stop you from asking it the wrong question though or like tweaking your question a bunch of times to to get some result.
13:10 Um, so, the job they take is kind of like the math job. Um, they they don't take like this human in the room job. Um, and then that's kind of where I think, you know, experiments go sideways. Um, and the way is that humans fail experiments is like a way LLMs can fail experiments as well, actually. So, both pattern match towards results that we hope for. So, both given access to our hopes um and say like some like borderline number are going to find a path to our hopes.
13:47 Um that's just like how it's how our minds probably work. Like And so, what this system we're trying to do then is only let each piece do what it's good at and then like add very strict kind of like separation of jobs and then layer of jobs even within each of these layers, too. So, like code Python in this case that their job is to decide is deterministic. So, the same input should get the same answer every time.
14:20 The The verdict comes from a function that can't kind of like flinch when the borderline numbers in front of you. Um that's a thing humans and LLMs are going to be bad at when something's like right on the border. So, we're going to cut them out and we're going to make that a deterministic decision step with code. Um LLMs, they're going to translate. They're going to hold the conversation. Uh they ask the questions that pin the experiment down. They're going to draft the brief. They're going to explain the guardrail.
14:51 Uh write out the kind of like readout. Um things that models are really good at. Um they would flinch if they knew all of the context through everything. So, they're not going to be able to see all that context and they're also not going to make the call itself. They're going to read the call from the code and then do some kind of interpretation around that. And humans judge. Um so, that's like all the work that actually require a judgment. So, like the design review meeting where someone catches that we're testing the wrong thing, Uh the stakeholder conversation about whether this is the right back, a bet, the the team discussions around what what metrics we could use if we're taking like big enough swings, the like kind of like prioritization decisions.
15:39 These are all things where our brains are definitely the the right tool. So, each piece kind of like we want it to live where it can't like flinch or be biased. So, what this actually is, we can kind of explain it in um a few pieces here. Um first, it's a set of agents. So, there's like just like five agents in here. Um one is the understander. It drafts metric definitions um with you as kind of like the co-pilot where you can do that yourself. Uh one is the designer, so that drafts the brief like the experiment brief, design brief, and the data plan. Uh one is the critic, so critiques everything that gets um committed. It doesn't only critique like, "Hey, are we including everything in this like design brief or everything in this analysis brief?" But it critiques the way we got there. Um then there's one that does SQL, another one that does like analyst narrator, so that works around like the pros of the readouts.
16:45 Um seconds are skills. It's like when you hear skills in the cloud code. There's this there's like two primary skills, design um and analyze. So, they're those first and third stage of that experiment process. Um and there's kind of like three size skills in here that you can check out. One's like to look through like the audit trail. Every basically every single decision that gets made is uh cataloged, so later on someone can audit through and see how um the experiment came to that decision.
17:19 Um and there's one for hooking up data and another one for um for uh running the readouts. And then third are just deterministic stats and um kind of decision tree packages. So this just Python terms of like how how to run your power analysis, how to run your T-test, um how do you like walk through and like check all of the kind of like uh decisions of whether or not you're going to your passes test. So agent draft, skills route, and Python decides.
17:53 Um I think we can get more into this when I just open up uh the repo. It'll be easier than explaining it here, but there's kind of two layers uh within this that are like the kind of like uh it's like knowledge base output storage. One's like the project layer, so these are definitions that don't change between experiments. So like what a user is, uh what a session is, what conversion means in your warehouse. Uh those live in semantic models and metrics at the project root. Um they're just YAML files. Uh you can read them, you can edit them, you can you can make them yourself. Uh they're not a database, they're just files. Um but they're things that just exist at the entire kind of like project or level.
18:36 And then there's the experiment layer. So this is the the like kind of like formula of a formalization of a single test. So your hypothesis, your primary metric, uh your decision rule, etc., etc. And so each experiment has its own directory with uh all of its kind of like information that's going through. Um the agent reasons through the second using the first. So when you say something like, "I want to test moving the buy now button."
19:06 Say you're working on like looking at like e-commerce site. The designer agent's going to look at uh your metrics directory. That's uh cataloging your in the root. And it's going to pick conversion rate as a candidate primary metric. And it'll ask you to approve that uh because it matches the word you used. It doesn't invent a metric. It picks from what's already defined there. Um and if you need a metric that doesn't exist, um you can work with the understander agent to kind of create that metric.
19:33 So, we'll watch this in a demo in a moment. Um I do want to kind of burn through these slides here. We're 20 minutes in. I'll try and get to the demo in the next uh maybe 5 minutes. Uh the way I think about it is like like experiment just like a genetic experimentation platform is like a pipeline with a conscience. So, like the pipeline part is very ordinary like ingest data, check it, design a test, collect, analyze, interpret, write it up. And then the conscience part's like at fixed points within that pipeline, stop yourself and refuse to continue unless your certain conditions hold. So, you're kind of trying to remove that human bias that we talked about.
20:16 Um like a critic agent's going to look at the drafted brief and flag problems like loose metrics, ambiguous thresholds, or populations that wouldn't give us enough power. Um So, what does the designer do? When you start a design, this is kind of like the key um command that I think we'll probably go through today. So, design, um you start to design a conversation in our code. The system's going to load your warehouse. Um It doesn't run any queries yet that reveal what's already um in the data.
20:54 It's going to profile uh what tables exist, um what populations are addressable, what events fire. Then it's going to elicit a question from you. Um like the agent will ask the questions that pin down what experiment you're actually trying to run. Um like is this the like what counts as a guardrail? Um, what's the what's the threshold to ship? It's going to draft a brief. That brief has a a hypothesis, the hypothesis, a primary metric, guardrails, explicit splits, a duration, and a threshold.
21:23 Um, a different agent is then going to critique that brief. It doesn't know what you hoped the result would be, like the original designer design uh designer agent does. So, it just checks whether the brief is well-formed and whether the metric definitions hold up. So, you're trying to like remove certain context from certain agents. Um, if it passes, it's going to actually seal that brief. Um, and then that brief won't be looked at again until the analyze step.
21:55 So, the whole design conversation happens before any outcome data uh gets touched. On the analysis side, uh effectively, you're going to run it's going to run through um eight steps and they're very ordered in a very specific way. Um, you have to pass one to get to the next and to eventually get to your ship or learn or or don't ship um decision. So, step one, does the assignment split match what the brief said? So, like if it was a 50/50 actually or was it 57/43 if that that invalid um if it's an invalid SRM, so it doesn't match, then like no lift is even going to get computed. Um, it's not even it's not even going to calculate the lift for you cuz you're going to cheat if it looks like it's up and the experiment is already faulty. So, it won't even continue to the next step. Um, step two, did any guardrail break?
22:55 Uh, again, this is a second step. This is before any lift has been calculated. If your guardrails that were predefined in the design step with very specific thresholds have broke through, the side I'm going to calculate the lift again because on the primary metric because you already said we're you're not going to move forward if those guardrails get passed. So, it's trying to remove the opportunity for you to bias yourself. Like, of course with anything like you can force things through your right. You can go analyze it somewhere else, but it's going to add friction at every point where there's opportunity for you to have biased thinking as a human.
23:35 So, on step three, is the sample big enough? Um again, if it's not, then it's going to halt you here. Um and then only then and only then after it passes those those three, so SRM, guardrail, and uh the sample size, is it going to look at your uh primary lift. If there's no lift at all, it's not going to stop. Um If it if there is lift, it's going to see if it moved in the um the predicted direction.
24:09 Um and it's also going to look at the uh magnitude of the lift. Yeah, if it's even worth um if it's even worth pushing out or if it's um or if it's so small that it doesn't even matter. Um And then step seven, did that hold across the effect window? So, this is like novelty effect or are you seeing that on the last, you know, 1/3 of the exposure period compared to the first 2/3 there's like more than a, I don't know, 25 or 30% drop. Those are thresholds you can set um within here in parameters.
24:48 Um And uh And so, if it's if it has no much effect, again, it's going to it's going to halt. And then based on everything that comes through here, it's going to tell you whether you should ship or not ship or learn or whatever. Okay, I feel like we should get to the demo. This is just kind of saying what I said earlier. Every You know, like we think we should give LLMs as much context as possible, as much data as uh context possible to like make better decisions. We're actually going to take a kind of opposite approach here. Um we're deliberately giving uh the judgment agents less context because the failure mode we're defending against isn't um ignorance, it's motivated reasoning. So, we actually want to remove context um as much as possible. Basically, create like blind judges. So, imagine you're on like a team and you ran your experiment and then you just had someone totally with no context on another team go read all the numbers and tell you whether you're going to do it or not.
25:58 Um so, it's just a way to remove the bias. Okay. Um the output I showed you a little bit in the beginning. We'll show it to you again, but it has those three briefs, the design, SRM analysis. And then it also outputs for different audiences. So, it'll have like a very high-level output for like an exec. Um to something that's extremely detailed with like the sequel SRM for say like an engineer on the team that needs to debug.
26:31 Okay, let's go into Claude code because that's more of the fun part. I don't know if there's any questions you guys want to answer while I'm switching screens here or um >> There's a couple of questions. There's one question, I think two questions around Bayesian AB testing. Uh about people moving towards Bayesian and like, you know, and how do we see Ezoic experimentation there? So, I think there are a couple questions on that.
27:05 >> Yeah, there's some there's some Python uh helpers in here for all different types of tests if you want to run uh Bayesian approach. Um you can. I I haven't done much of that, to be honest. Um so, that's not really what about This isn't really a a uh talk about like, are we going to move to Bayesian experimentation? It's more about like um how are how do Ezoic systems um remove bias and kind of speed up the experiment design analysis process. But there there is like um when you look at the repo, you can you can take a look at some of the uh Bayesian functions in there.
27:47 >> Also, there's a question about do we have an experimentation course specifically? Uh and that's great. Anish, I think shared Ron Kohavi's book, which is great. But we also have A Analytics for Builders, where we do not fully share about experimentation, but we have almost two solid weeks of content on causal analysis. Uh and uh like we'll share probably more about it later in the workshop as well. But yeah. >> Yeah, we got pretty deep. It's not like the it's not like a end-to-end you're going to be the extreme expert everything experimentation. It's more like, here is the kind of, you know, 80% to 90% of how people run experiments in tech today. Um and then, how do you execute it through uh cloud code and Ezoic systems. I don't know if any of the other courses on Maven right now have uh Here's how you do it with Ezoic systems or if they're more like classical. There are some really good courses if you're looking for foundations on that. But we do spend, yeah, a full week or two um just on that.
28:54 I think we'll probably add something or at least some free content. Um yeah, I mean I'm trying to kind of burn through this whole thing, but like when I was trying to to talk think about what I wanted to talk to you, I could easily do like an 8-hour walk through of this agent XP thing we built. Okay. So, we're in cloud code here now. Um So, let's walk through a couple folders and then we'll run some stuff. So, uh two folders.
29:25 Um those ones I told you about before. Uh metrics. Um so, there's like 14 metrics I have I have predefined in here for this project. Just, you know, so like conversion rate, uh add to cart rate, revenue per user, page load time. All things you'd actually track. Say like an e-commerce website's what I'm thinking here. Um each one is just a YAML file. Um And on the other side, there's uh semantic models. Um user, session, order, page event, assignment. Um these are those project level things I told you. They don't change between experiments.
30:05 Uh they're kind of like the vocabulary of your org when it comes to metrics, right? So, agents pick from these. They don't invent them. Um so, we can take a look at one like uh revenue per user maybe. So, you can see here like the type it goes through like what's the type of metric is. It's a typically a mean. Um that drives which stats functions gets called, right? So, if it's like a proportion versus a mean, you're going to run different types of stat tests, so you want to have that um around your metric definitions. um, the numerator is a is the sum of order value um, per uh, user. So, um, that gives it some signal as to where it's going to pull from semantic models. Denominator is distinct to signed users.
30:58 So, that's all in in plain text, right? Uh, if you scroll down here a little bit, you'll see, uh, a common roles. Um, guardrail. So, um, and then it also gives you just like guardrail default number of minus 1%. So, that means any experiment that uses this metric automatically inherits a guardrail that catches a 1% drop in average revenue per user. So, the metric in itself carries its kind of own discipline. And of course, these files, this is where that human judgment comes in, right? This is where you're going to have this YAML file like as a discussion with your team or your stakeholder or whomever and decide what comprises, like how is these metrics actually calculated? What metrics should even even be in this catalog? And then, um, the kind of like default, um, like MDEs or default, if it's a guardrail, um, decreases or increases you would want to guard against. And when you run the experiment itself, you can change that.
32:03 You'll have an opportunity to change that if it's if it's experiment specific, but it is nice to have some defaults there. Um, okay, so let's go ahead and run a design. So, I'm just going to type design and it's going to be running this skill up here. Our design skill. Um, you can read a bit about it. Uh, but it says like, you know, design pre-registers an experiment, walks from intent to semantic models to metrics to hypothesis to brief to data plane um to your seal brief. Um dispatching specialists at each step.
32:45 So, uh the first thing it does is um it's going to load that skill and trying to pull some uh intent for me. So, it says, "What do you want to test? Uh describe the intent in plain English. What change? What surface? Um what where you actually to move? So, um let's say I'll type it out. We're going to use just a simple data in here.
33:19 And what I want to test is I have a prompt over here. So, I want to test whether moving the buy now button above the fold on the product page improves conversion. >> [clears throat] >> And so, what should happen now is the designer agent should start reasoning against um metrics and should start proposing some primary and guardrail metrics and give me um some proposals.
34:02 So, you can see it's listing out these steps that it says it's going to walk through, right? Capture intent. And then there's gates for a lot of these, confirm the intent gate. Um and then it's going to ensure the semantic models and metrics exist. Then it's begin drafting a hypothesis. From that hypothesis, it's going to draft a brief. Uh you can't see these four pending ones, but at that point, it'll have a a critic come in.
34:31 Um so, Yes. It's going to create a new experiment here. So, just created this new experiment ID. Um, and it's just confirming that this is what I'm trying to do. Yes, proceed. Now, it's going to go and check through the semantic models to figure out what data is available.
35:06 And it's not going to make metrics up here, right? It's going to look at the metrics directory, um, that catalog we looked at earlier. So, it says here, yes, semantic models and metrics already cover relevant entities. And it's inferring some stuff here, right? It's going to give me an opportunity later on if I want to use different metrics. But, there's not too many to pick from, so it's pretty easy to to figure out here that I want conversion rate as my primary because I said conversion in my intent sentence here around what I want to test.
35:39 Um If you don't have a metric, um, within your metric catalog here that it thinks you should test, it's going to flag that and, um, it'll give you the opportunity to, uh, create that metric yourself or work with the Understander agent to kind of like co-pilot through creating a metric. Um Once you create that metric with it, it's going to check through your, um, semantic models to see if that even exists in the data available. If If it doesn't have a metric, it's not going to run. If the metric doesn't doesn't, um, exist in the semantic models, it's not going to run.
36:20 So, those are really critical. And it's a it's a way to also just kind of like add some safety nets around the work you're doing. Just having very clear metric definitions and semantic models defined. Otherwise, you know, you can make stuff up. >> [clears throat] >> But there's lots of gates through it throughout here to make sure it's going to check that at each step. Um So, like I said, it picks conversion rate because it matches the word I used.
36:49 Um I think it also proposed some other metrics up there. Come on, Rita. Yeah, proposed a secondary uh metric here, buy now click rate. It's probably going to propose some guardrails when it gets into the brief stage. Sometimes it'll propose them um beforehand. Uh I could intervene at this point and I want if I wanted to it'd swap the metric and say, "Oh, actually I want to use add to cart rate as a primary and keep conversion as secondary."
37:26 Um if I hit like control C and just interrupted it. So, it's hitting some errors while loops through trying to get the hypothesis. Let it work it out. We can take some questions while this uh runs though if there's anything coming up, Shradha or anyone. >> Yeah, there's a quite a bit of questions so we can pause while we run this. So, one question is how do we create the YAML files for the metrics and thresholds?
38:03 >> Yeah, so there's a template here. Um I mean there's two ways to create it, right? Like um this is this is where I think the the human part is important. So, like metric definition uh hopefully somewhere in your company you have some sort of metric definition, some sort of metric catalog, some sort of document. Maybe it's just living in Slack or in someone's head or maybe it's in a SQL query somewhere.
38:35 But like this is not like the automated part of coming up with the definition itself. You can get it into the structure just by working with Claude and explaining like, "Hey, here is um here is how we define this metric, these are the fields um in this semantic model that we pull from. Uh here is the kind of like uh typical MDE um that we want to see. It's It can infer a lot of this stuff too, right? If you're saying like revenue, it's not going to be like, "Oh, the direction of lower revenue is better."
39:12 Um but this is I think a conversation with your team as to like how metrics are defined. We're not going to go into like, "How do you make really strong metrics today?" We're actually going to do that next week for free also. We're going to We have a session next Friday called North Star metrics with AI. We're going to talk about how you can um create really strong metrics that kind of like align with user value um and then you can kind of like uh hand that off to Claude to um align to this uh template.yaml to make the right metrics.
39:46 Um Okay, so this is uh what is triggering right now is I ran this I ran this uh right before we started. I tested out this demo. So, it's actually it caught that and said, "Hey, this experiment um was already ran and this experiment ID um is this a fresh one or application?" Um we're going to say that this is a fresh test so you can kind of see it go through and draft everything from the start.
40:17 What other questions? >> Uh let's see. How is this repo's skills for design and analysis different than the AI analyst repo? >> Yeah, the AI analyst repo um is not purely uh experiment focused. And I think with the AI analyst repo, like a lot of that is I do want a lot of human judgment in there as well.
40:51 Um the use case for Agent XP is like it's really it's only focused on uh A/B tests. Um it's not going to do root cause analysis for you. It's not going to draw you a full deck. Um it's just for experiment design and analysis. And the primary goal of it is to um remove bias wherever possible, whether that bias is from the human or from the the LM. Um ideally, like our goal here would be that anyone on your team can, whether you're a PM, an engineer, a designer, a data scientist, for a simple experiment, should be able to run this and get a really solid experiment design. Now, that doesn't mean that like someone just runs it. Um if you don't have any data resources, like yeah, this is definitely going to be probably better than what you're doing today. But you probably want a a data scientist to review if you're not a data scientist. If you're a data scientist, um I think this speeds up your design um process and allows you still to like focus more critically on like the really complicated uh challenging um experiments that are more nuanced um rather than like the kind of simpler ones. So, uh divert work to more complicated uh work more for our data scientist and then open up experiment design to teams that don't have um a data scientist working with them.
42:26 >> What else? >> More questions. Madhu had a couple questions here. One is where is randomization defined? And then is it possible to randomize? >> Yeah, so you're thinking around like the um there's that those three pieces I talked about, right? There's um the design. There's like the actual like running instrumentation implementation. And then there's the uh analyze. This isn't like a full like it's going to look at your user base and assign everyone to treatment and control. Not yet. It's not something terribly challenging to do. It will check for randomization after the fact in the analyze stage.
43:11 Um but this doesn't do like the assignment and the like um implementation of the test itself. But again, I mean that's that's not something that's terribly hard to do. But it does require like um you know now you're kind of plugging into a system where you are turning features on and off for certain end users. So like, you know, there's whole platforms that are already solved for that. We're we're not really trying to recreate something that like experimentation platforms are already pretty good at. We're trying to take the part where experimentation platforms aren't super good at where they're like giving you the numbers um without like putting guardrails around like did you design that correctly? You know, you're analyzing them correctly and making the right judgment from it.
44:07 Um so we do have some briefs coming out now. So, if we go into experiments and we go to here's our experiment. This is the ID here. Um it's starting to create some readouts. So, it's got our um initial intent readout. It's got a YAML of the hypothesis here. So, you can see it extended that intent we had to a pretty detailed hypothesis. Um, it has rationale around, you know, why we believe that hypothesis. Um it gets into a bit of the kind of leading indicators and guardrail metrics. If we go into the brief YAML it'll actually assign start assigning um numbers to these. So, um there's some default um MDE figures in there for each uh metric. Um, so it'll it'll leverage those uh to start, but once this outputs to us, we can change it at any time. So, here it says like brief drafted with feasible 5% relative MDE. So, what it did is it went and it ran a power analysis and um, you know, below a 5% MDE probably not going to be able to run this test. Um it's not just running the power analysis for um the primary, it's running it for all your guardrails as well, which is something I think a lot of people miss out and forget about.
45:41 They think, oh, we're going to make sure we have enough power to measure our primary metric, but then they don't actually have enough power to know if the guardrail metrics have been triggered. So, it's going to do that for you as well. Um and it'll have those uh those um thresholds that you can't pass. Um it's going to define um the exposure assignment. It's not going to do the the randomization for you like say user, you know, XYZ is it, but it's going to define the the unit of measurement for you.
46:21 And then what's happening right now is after the uh the draft um the brief is drafted by the designer agent, it's being hit by that other agent, the critic agent, to kind of review it and go through a review loop cycle. We can keep taking questions. Yeah, so critic brief consistency on brief. And we'll share this out and you can kind of read through those what each of those agents um do or like, you know, here's our critic agent. The critic agent doesn't just judge brief consistency. It'll say like, "Hey, um it'll it'll look at the analysis versus the brief. Um it'll look at the verdict versus the analysis. It'll even check like, "Hey, are the stats that you pulled during analysis just from the white list and agent XP.stats or did the LLM go and start go rogue and create its own functions?"
47:29 So, trying to add as many guardrails as possible. >> One question from the chat is, "Where do the data get connected?" I think at least that's the flavor of them. Or >> Yeah, so there's like a connect data skill that we've created and I've been working with it. I Right now, I'm just working on a local DuckDB database on my machine with some sample data. You can like any of you just like this stuff's really easy to connect to Snowflake or Databricks or whatever.
48:02 Um you can run this skill and it'll help you wire it up, but you're going to have to create semantic models um that tell uh the agents where to look. So, like we have those over here that tell like, "Hey, what what source table are you going to pull from? What are the dimensions within that table? And then the metrics um over here are going to relate to those semantic models, but you want to have those like clearly defined in your uh repo here so it can have that context about what to pull from where.
48:46 It's not just going rogue and guessing stuff from your warehouse. You probably have something already pretty similar in GitHub in order to create those tables. Um okay, so it has created the brief. Actually, we're doing pretty good on time here. Um we're not going to get through the analyze stuff. You guys can run that on your own. We can do it another session. Um but uh it says, like, here's the brief where it is. Do we want to like seal it and move on to the next step or is there anything we want to change? Um for instance, like the MDE. I'm going to say, "Seal it."
49:22 cuz we have 8 minutes left. And then um it should generate a nice little readout for me. What other questions while we wait for that? Or there's a lot of things going on in chat for folks have questions that haven't been answered, do you want to come off mute and ask?
49:57 >> Or you can type them in. Uh We should put some of these in uh the Slack channel, too, for afterwards. I can also answer some async. Yeah, but Let's see. >> We we can also tell folks, guys, we have a Slack channel. We are at 899 members right now. So, yeah, you could be the luckiest 900 member if you join it. I'll share the Slack link. Uh yeah. For any questions, you could also, you know, post your questions in the Slack channel, too.
50:31 >> Yeah, we should pull these out into the Slack and just answer them all in the Slack. Maybe I'll pull these out and like answer them in a follow-up email or I'll create a a page with the answers to them. Okay, so this the design um portion has wrapped up. So, like yeah, for a lightning lesson, I don't know if I'd demo that again that way cuz it did take 13 minutes end to end. Just kind of a long time for a demo. But, in your real life um like, you know kick this off, go get a cup of coffee.
51:08 Maybe a glass of water. And a small snack. And you come back and you got your design ready for you. So, um it is a pretty quick first pass at the uh experiment design. Yeah, let me share the repo right now before I forget. So, I'll drop the repo in um I just pushed some changes to it. Right before this, I did find some weird stuff. So, I'm going to push some more changes probably this evening. This thing will probably be pretty pretty actively iterated on. It's also pretty early uh days. But, if anyone wants to play around with it feel free.
51:49 I think we could probably open it up for contribution, too. That would be awesome. Uh if anyone wants to contribute to it. Uh let me see where our readout is, though. Design brief. Oh, hey. Didn't give me a nice Maybe it does. You didn't give me the HTML or PDF version of the readout. Just gave me the MD. So, we'll let it uh take another pass at that.
52:26 I'm going to switch back to slides. But, once the readout's finished, I'll um I will show that. Um Okay, we only have 5 minutes left here. I mean, I can stay over a bit if we have questions. Um But, I think I think the main thing here is like you know, a data scientist hasn't doesn't have to spend all their time on the experiments that genuine They can They can spend their time on experiments that genuinely need their judgment.
53:07 Like, the design review conversation where someone says um you're testing the wrong thing, that's still like a human meeting, and we can invest a lot more time in that. Um I think this cycle leveraging like a digital system systems like this really not only it compresses the work and makes it faster, but it does have a real opportunity in this case to add um rigor to it. Um it's like I don't know. Human judgment on on making calls around experimentation, um we talked a lot about how we need to put guard rails on like what agents and LLM's can do and and check their work but like there's a lot of human error um in that occurs when running experiments and I think there's an opportunity here to add some nice um guard rails around it that aren't like crazy rigid through code and have some flexibility around like your specific use case leveraging uh LLM's.
54:07 So once that design brief gets generated I'll I'll show you guys it here. We do have some upcoming uh sessions just or courses in the next couple weeks if anyone's interested and then we can talk about questions too. Next Wednesday uh for anyone here if you haven't I don't know if people have who has and hasn't used Cloud Code but if you've never used Cloud Code before um we do a super intro workshop next Wednesday. Um it's super cheap. We get you a like hold your hand going through and sign Cloud Code and running analysis in there for the first time with our open source AI analyst repo.
54:47 It's like 3 hours. Um it's pretty beginner if you've uh if you use Cloud Code before you probably don't need to do that. Um but in the following weekend we have our full Cloud Code analytics boot camp where we uh teach the framework around building an agentic analytics system yourself. Uh we have you run multiple analyses in there. We have you build on top of our uh repo. We have you port over stuff that you've built from one repo to another. We connect to data warehouses. We connect MCPs. Um this is kind of like our foundational boot camp. Then uh later in June we have our advanced boot camp what gets into AI evaluation validation um and context management and then multi models. So like not just Cloud Code but Codex that opens source models. We'll get a bit into validation, open source, and context management in the weekend boot camp as well, but it's going to be kind of intro async material.
55:51 But both of those are live. They're both like four or five hour sessions. We're also going to run the cloud coding Alex boot camp in July again during weekdays for just two hours a day. If you're more of a weekday person. And then in a couple weeks we're kicking off AI analytics for builders. That's that one Shravan I mentioned where the sec where the fourth week is all on experimentation. Um this is our kind of like think the end-to-end sort of like how do I become a product data scientist in tech course, what are the skills I need from framing questions, working with stakeholders, developing metrics, setting up the tools, doing root cause analysis, doing trend analysis, doing segmentation, doing experimentation, um running causal inference models, uh drafting presentation, doing stakeholder uh readouts, um all of that but executed in cloud code.
56:49 And I don't know, maybe we'll do it we'll do some execution in code x this time too. As that's the getting seems like more and more um robust when it comes to data analysis. There is a promo going on with Maven right now that you all probably have seen where their top courses you can get 25% off with this promo code Maven 100. This is better than any promo that we run. We usually get 10 to 20% off. That expires Sunday uh June 7th at 11:59 p.m. I bet. So, you got a couple more days if you want 25% off.
57:25 That's like site wide off all our top courses. So, I'm sure some other ones people dropped in chat around like experimentation too have that same deal. Uh let's see if the I wonder if it finished the experiment brief. I mean, this is the experiment brief I showed you before. So, this is like what it outputs, but I want to see the one that we actually did. Any Any questions? Well, uh we'll kind of wrap up here. I don't know if Ravi and I have to go, but I can stick around for a bit.
57:58 Yeah, Mohit. >> Hey Sean, hey team. Uh so, quick question. Perhaps I missed uh in this entire course, uh we would we were basically targeting the first and the third pillar of the entire process. First one was about designing and the uh third one was about, you know, the post experiment uh analysis or readouts, which the LLMs would assist us. Um what I probably missed, and that's a question, did you already have like some experiment results in the repo? Uh based upon which the third pillar got executed?
58:34 >> I have some I mean, I'm you I'm working with like a sample data set. So, I have a database locally on my machine my machine that I'm connected to. It's not It's not in the repo, but um where I have uh some experiment results. Well, I have the I have a tables that those tables that the metrics are compa- comprised from. And then, I have a table like you probably would at your company where it's like the exposure and the exposure periods at whatever levels. So, there's a table that's basically like user 1 2 3 4 5 Sean uh is in experiment ABC >> Yeah.
59:15 >> variant treatment this start date this end date. That's how I've usually There's usually some table like that at your company. So, um the analyze step, which we didn't get to today, but you can kind of check out how does it will will aggregate and do all of the the statistics by writing a SQL query that joins to that table to the um tables that comprise the the metrics. >> Yeah. Yeah. Yeah, got it. Got it. Cool.
59:42 >> But yeah, this yeah, this doesn't um this doesn't do that middle part though. Like actually like, you know, building that table and um assigning people to the experiment. >> Yeah, that makes sense. I was I was a little confused where I mean, I was like, "How did the total builder get executed if we do not have experiments?" So, but like you said, it is stored somewhere on on the Yeah. >> Yeah. >> Yeah. >> And I think you could do that. Like I would love I would probably work on building it out to get that middle um stage, but I just think it is like a higher risk stage of like, you know, if I mess something up with the design, it's like I I can go back and and fix that before it starts. Um but if we like screw up the actual like instrumentation and assignment and then run that test for a couple weeks, it's like, "Oh dang, we just We not only ran the data's useless, but we also could have given a bunch of people the wrong experience or something."
60:38 >> Yeah. >> [clears throat] >> Cool. Awesome. Thanks. >> Yeah. What else? I think the design brief is done. I'm just trying to find it on my computer here. Other questions in the chats, Rob, or hi? >> Lots of questions around the Slack channel. Um so, if folks who can't access Slack, um either because you haven't used it before or you don't have an account or if you have an account but you still can't access it, um just reach out to one of us or reach out to me on LinkedIn.
61:23 And uh we'll get you we'll get you in. >> Is it saying that you have to have a certain email domain or something? Is that why that >> Yeah, that's what I'm seeing right now. >> Oh, let me try it. Yeah, yeah. I don't know what is up with that. Let me try it. >> Try the link again. >> Yeah, because I I have a Slack account or with personal email. When I click on it, it says >> It needs Data neighbor AI and all this stuff with the company >> Yeah, they say reach out to the coordinator to create an account something like this.
61:57 >> [clears throat] >> Yeah, we got to Let me try this. >> Can you see the new invite link? >> Yeah, yeah. I think at some point the invite links we have like revert back to only people with the company uh domain [clears throat] can access it. But I think that one I just dropped in or the one Shravya just dropped in, hopefully that works. Let me try that.
62:28 What else? >> link is definitely different with this new one. So, that should work. >> Um any doesn't have to be experiment agentic experiment related questions. There's anything else? How's the tool know the right KPI to IT is? Yeah, so there's um that metric catalog that I showed you. There's like just like 14 metrics in there right now. And so, it inferred for me that it wanted a conversion rate because when I like told it like what I want to run experiment on, I used the word conversion.
63:05 Um if it can't figure it out, it's going to interview you more. If there's no metric in there that matches like what your intent is, it's going to flag that. But um you know, and you could make this more clear, too. I mean, um, we could add a thing in the YAML file, which I don't have, where it's like, "Questions like these." and have example questions, um, should use this type of primary metric. In most cases, like a lot of experiments are going to only have, like, if you're in a specific product area, you're probably focused on very specific metrics. Like, we'll go over this in next week in our North Star metric, um, free session, where, like, you know, each team really has like a primary metric that it's trying to drive, which reflects the user value it's trying to build. So, there's something you could definitely map in there, I think, into the YAML file.
64:02 Good question, though. And also, like, I think it's like, uh, I put this in the the read me on the Agent XP thing, but it's like this is also just like something where, you know, we spent the past, I'd say, since like mid-year last year, we've been pretty focused more on like, we've been focused on the data analysis side because we're more, uh, confident around Agent XP systems being able to do, you know, stuff that's very easily explainable, like, segmentation and root cause, like, correlative kind of, um, behaviors. Um, we haven't really pushed it into a lot of the kind of like, causal relationship stuff, or we do a little bit in builders, because that gets into a lot of kind of judgment around what variables do we create, what transformations do we create on those variables, what methodology do we actually leverage to, um, try and find this causal relationship. Um, so, this is like a little bit of a new frontier for us, um, around experimentation, but it is a lot more cut and dry than say, like, causal inference modeling, where you're like creating a bunch of variables.
65:13 Um which is why we're kind of starting there. But as we learn more, we'll build this out and share it more, too. Nice. Any other questions here in the chat? Any other >> everything. If not, async array. So, if folks Yeah, it's too long of a wall of text here to parse. If anyone has any questions, feel free to raise your hand or just come off mute.
65:49 >> Cool. Well, no worries. Not. We will um Um was there any agent on impact calculation for rollout? Yeah, in the analyze I did I didn't get into the any of the kind of like analyze phase, but there's a there's a bunch of uh checks there. There's like the There's like eight um stage checks. It's not just going to be like, "Hey, is it a above or below your your MDE?" It's going to be like checking the the confidence intervals. It's going to be checking like, "Is the impact actually large enough that you want to roll it out?" Or is it like statistically significant but quite small?
66:33 Um So, maybe we'll do another session on that. Um I think it was a bit too much to try to design and analysis in one session. Cool. Um yeah, so just a reminder intro to cloud coding and Alex workshop's next Wednesday. Um the boot camp is 13th and 14th. Um and then builders is on 15th, so um yeah, that Oh, we do a builders boot camp thing, two for one, if you're interested in that. Basically, if you do the boot camp and then you decide you want to go on to the five-week course, uh we remove whatever you paid for the boot camp from the price of builders.
67:15 So, that's a pretty good deal. Um and then yeah, the Maven 100 deal I think ends on Sunday. So, whether it's our course or another one, I'd recommend checking out their catalog by Sunday. This is probably like one of the best deals you're going to get on Maven anytime soon. We'll send out an email um with the repo in case folks missed it in the chat. Uh though if you look up like AI Analyst Lab um on GitHub you'll find it in in that org, Agent XP. We'll also send out the recording.
67:46 And uh I think that's it. I know. I'll send out the PDF or something it it generated too if you want to check that out. Cool. Thanks everyone. Appreciate the time. Um and hopefully see some of you in our courses or next week in our free sessions. We have two next week. Northstar Metrics and then we have a pretty cool one from Strawbee he's going to do our ads. It's called like What's it called? Share What's the name of it? It's a good name.
68:17 >> Share like basically share analysis any every everywhere using AI. >> Yeah, yeah. So, be around like MCPs and stuff. It's it's it'll be pretty cool. So, um those are both free. Um check those out. All right. Have a great weekend everyone. Bye y'all.
Summary
- Introduction of agentic experimentation, highlighting its significance in data science.
- Discussion on common pitfalls in experimentation, such as biased interpretations of results.
- Presentation of a structured system that includes agents for designing experiments and analyzing results.
- Emphasis on the importance of clear metric definitions and guardrails to prevent bias.
- Overview of the experimentation process, including design, instrumentation, and analysis stages.
- Introduction of a repository for open-source experimentation tools and resources.
- Future workshops and courses on experimentation and data analysis to enhance skills in the field.
- Encouragement for collaboration and contributions to the open-source project.