Transcript
0:00 Uh, we can do some intros and do a few little polls and stuff of the room. Um, so for next about half hour, okay, hold on, confirm. Um, all right, so today we're talking about uh, developing North Star metrics. We're going to talk about like about the next half hour or maybe 40 minutes. I'll try to leave 20 minutes at the end for questions though. It seems we can never do that. It seems like we always just go to the end and I'm like, "We'll stay over." So, we'll see if we can do it today.
0:35 Um, but yeah, we're going to talk through how you actually pick the one metric that kind of tells whether your product is really working. And then I'm going to put that in front of we're going to use Cloud Code and show you a couple of things live in the demo. We're partnering with Amplitude on this one, so we're going to be using Cloud Code for like the system, but actually uh, everything we're teaching here comes from well, it comes from like our experience, comes from a bunch of other research that we're pulling in, but actually like the whole kind of like Amplitude North Star playbook, a lot of the frameworks that the agents and skills and functions that we're going to run in Cloud Code today are based on that.
1:21 It's actually like baked in. And so, I'll show you the the skills and agents today are a little different than what I usually run. Usually we kind of have a skill that has a bunch of information around it, what it should do and it might fire off other skills or fire off agents that then fire off helper functions. In this case, a lot of the the skills and agents go more in depth to read through research and that's curated in knowledge base, read through different rules that apply to different industries when it comes to metrics.
1:54 Um, there's actually an agent that's called like the North Star librarian. So who will go and make sure that any sort of decision that's made around your metric has some sort of source that will pull up to you as well. I don't know. Maybe hi, if you want to look up the North Star playbook from Amplitude. It might be one of those things where you have to like put your email in and then they'll send it to you, but it's pretty good. Um so I wouldn't mind I wouldn't uh I would I would recommend reading through that as well. Although we're going to talk about it here.
2:27 Uh to kick things off though first, um let's do Oh yeah, let's do some intros. I think I recognize a lot of the folks here today. Um you've probably come to some of the other ones, but I'm Shawn. I'm a co-founder of the AI Analyst Lab. Um we do a bunch of free uh workshops around um agentic analytics. We put out a bunch of uh free kind of like email courses, um async stuff around learning how to use the agentic analytics. We've done a lot of primarily with Cloud Code, although although this past month we're getting a lot more into Codex um and open source models as well. Uh we also offer some uh paid courses. We actually have a boot camp coming up this weekend where you build your own agentic system.
3:11 Um it's like a no prereqs required, no code required. Teach you how to build uh your own agentic system. Um and then uh a lot of that's built a lot of that's primed on like our open source agentic analytics system uh at the AI Analyst that we have uh released. I'm sure many of you have access to that repo. And then next week we have a pretty cool course called AI Analyst for Builders, which is like a five-week course around analytical thinking with execution layers in uh in Cloud Code. Um maybe we'll get into Codex on that one a little bit, too. So think of it as like your kind of like you know, I'm a uh product data scientist or data analyst and I want to have analyses that go faster, that go deeper, um that go further.
4:04 Um this is like that or I'm a a PM or engineer or designer or banker. We had like a chef in here the other day actually. Um and I wouldn't like learn how to um to do analysis, but I need to like understand the frameworks first and then now you don't have to like do the coding in R or SQL or Python. You can do it do it first cloud code. So we'll talk more about that later though.
4:25 That's us. Uh Hai, you want to do a quick intro about yourself, Sravya? >> Yeah. Uh hey everyone. My name is Hai. Uh co-founder of AI Analyst Lab. Um as Shawn said, we put out a lot of uh content both on LinkedIn, our emails, and uh and and uh Maven just like this one. Um I've got 20 years or so experience in the data science and analytics space and uh currently utilizing AI for analytics and AI analytic systems for analytics quite a bit.
4:58 Sravya, you want to go next? >> Yeah, yeah. Hello everyone. I'm Sravya. Sorry I had to be off video. Uh but yeah, um I have 15 years experience in data science. I started off career at Microsoft and later worked at Nextdoor and most recently Grammarly, which is uh branded as Superhuman. Uh at Nextdoor is where Shawn, Hai, and I met and that's when we started a bunch of things together and most recently AI Analyst Lab. And uh very excited about our the journey that we're taking from here. Yeah.
5:33 Shawn. >> Sweet. I also just saw Mahawai's that you said you're located I'm in Miami, in Bali. It sounds tough. It sounds like a tough life. Yeah, that rough. I lived in Indonesia for a couple years actually, so I I spent in East Java, but I would go over to Bali every few months. Um All right, cool. Let me get into it.
6:03 So, a couple questions before we kick off just to get a read of the room. We can kind of change the the format of this knowing where people are at. So, first thing here, I just want to see like people's kind of where they're at with Claude Code, where they're at with the Gentic Analytics. So, if you've used Claude Code Um if you've never used Claude Code before, put a zero in the chat. If you've used it, but you haven't used for analytics, put a one.
6:31 Uh if you've used it and used for analytics, put a two. Cool, we got >> ones and >> a couple of people who haven't used it before, lots of ones. Um Nice, Jeremy, welcome back. Sweet. Well, yeah, a lot of our Oh, nice. Hey, I think you get a I think you get to be a one or two if you use Codex. We're just We're saying Claude Code. I I feel like they had like with the coding stuff a little more of like the first mover thing, but um yeah, feel like Codex is doing the same [clears throat] thing these days.
7:10 Um Yeah, and then >> [sighs] >> I don't know. This is like total aside, but I I've been talking to a lot of people as we've like been working with some of the other models too. I just feel like the open-source model stuff's going to catch up so fast to the to being able to do what these other things are doing. It's like the I um I think the stuff I'll do today would do like Opus 4.7, but to be honest, it's like 4.6, 4.7, 4.8 fable in terms of analytics tasks. If you have a Gentic harness around it, there's not much of a a difference. Definitely some step up if there's like vanilla um is asking cloud stuff, but um I feel like we're going to hit some some threshold where the models are so good uh that, you know, that incremental lift uh you don't necessarily need it to to accomplish what you're doing with your day job. But cool, good spread today.
8:03 Um we'll definitely be showing the capability of uh level two today. Um and then second one all right, more and more start North Star metrics cuz we're going to do kind of like some framework discussion here before we even get into analytics. Uh maybe a zero if you haven't really worked with a North Star metric before um and a one if you have. Um let's see.
8:36 Yeah, go go connect with Valen. I like the community building. Cool, nice. This is good. Good even split 50/50. Nice. Um so some of this is going to be new, some of this will be review for people in the beginning. You're one if you're a one, this is some of this will be a bit of a review around what North Star metric is, but we want to set the foundation for it cuz there is, you know, I'd say at least half the people in here haven't worked with uh this North Star metric. Zero and one, Nikita. All right, cool. Uh I mean, I know a lot of people who have heard the word North Star metrics that I've worked with who uh I would say they're basically at the zero level when I see them try to make the metric, but no longer because AI can help them, so they don't have an excuse. Um yeah, so so let's get into it. Um so where I want to start is like think about uh your team your company is kind of like primary dashboards for a second.
9:40 Um And you know, this maybe this is a little dramatic, but if it's anything like most like it's not one metric. It's really easy to add a ton of metrics on there, especially when you have like multiple tabs of dashboards going. Um sometimes I've seen them like especially especially some of like the high-level ones where it's just like or even at the very low team level it's just like 30 or 40 numbers on it. At least most of them have at least like 10. And the question I'd ask is like, you know, for when you look at a dashboard like that that's like just inundated with all these different metrics on it. Like 40 is dramatic, but say even 10. Like how is he how easy is it for your team and for you to all agree and point out one of them and say in a sentence like, "This is the one that tells us that we're winning."
10:33 A lot of people can't do that but they have so many metrics on there. Um we track a ton of activity and we kind of call that measurement. And honestly uh AI makes that worse, not better, because now you can measure basically anything. So it's easier than ever to pour all your effort into the wrong number or collection of wrong numbers. So uh what we'll do, I'm going to give you the the framework here that companies like like Airbnb and and Netflix use to pick uh the one metric that really matters.
11:04 And then I'll show you how we can use AI to really pressure test that metric live in the demo. And we'll turn like a full year of kind of like raw data into into a decision in about I think that one takes like 5 minutes or something to run through the process. Um but the problem basically is that a lot of teams um what a lot of teams measure is activity instead of value. So um it's really easy to make this mistake.
11:33 Things like logins or sign-ups or your monthly revenue, those go up and to the right and it feels really good when that happens. But if you're honest, none of them really tell you whether your customer or your users got more value out of the thing you built, which is like the reason you built it. It's like, you know, like revenue is the reward for us creating the user value, but we build a thing to create the user value. Um revenue is whatever it's just like a kind of steps removed that's just like, "Oh, that that maybe happened." Um you know, I I I pay for plenty of subscriptions that I have just like forgotten to unsubscribe from. Um I'm not getting user value out of them, but some company's getting revenue out of it. Um So it shows up in two ways, either a vanity metric. So one number that looks impressive but means nothing. Like daily active users is a very, very, very common one. Or the opposite, a dashboard with like, you know, 30, 40, 50 numbers and no single one everyone's rallying around. So every team picks their own and you end up arguing in a lot of these meetings instead of being aligned.
12:50 Um And the reason it matters so much is that this metric sits at top the very top of everything. Like if the one you picked is wrong, then every decision underneath it inherits that mistake. And you end up shipping things that moves the number without actually moving the customer. So um So I'll come back to it. Like we're figuring out the right metric. Actually, figuring out the right metric for your own product is week one of what we teach in our AI analytics for builders course um starting on Monday. Um we'll come back and talk to more about that, but like if this isn't an easy thing to do, we're going to talk about it, right? And like we got 44 more minutes here, but we actually spent an entire week um, on this topic because it's so important um, to make that one number correct.
13:41 Um, would love to get a poll from the room like for those of you who said one or even those for you who have said zero, who could think about this, like what's your team's kind of current North Star? Do you have like a clear uh, North Star for your for your team chats? A conversion rate? Conversion. Case time. MAU, ROI.
14:12 You're not going to like what I say about MAU for the MAU people. And you know what? I'm not going to say it. AI's going to say it, so it's I'm not even going to blame me. Loss ratio. Nice transaction rate. Cool. All right, we'll keep going. Um, so what actually makes a metric a real North Star? Um, so there's kind of like seven questions you can go through and I'm actually going to send you guys So, this repo and stuff I'm going to work through is our like plus AI analyst pressure build. We have an open source one called AI analyst.
14:47 And everyone has access to that. This whole North Star thing, which I'll show you later, is in our plus repo. That's currently only open to our course students. I I think we might just open source that, too, later on. Uh, we haven't like decided yet. This North Star stuff, all of us have pushed to main on that one even yet, actually. We'll probably open source that, but in the meantime, um, come on, I cannot click on this. I do have this I'll send this to you guys at the end that kind of has like uh, just like a little worksheet with like some checks around um, so you can check your own North Star so you don't have to remember all this. Um but we'll also send we'll send a recording of this, too.
15:28 Um so actually makes a metric a real North Star. So seven questions. Uh does it capture customer value? Uh reflect your strategy? Is it leading rather than lagging? Can your team move it? Can a normal person understand it? Can you measure it? And is it free of vanity? And then four of those seven are really uh real deal breakers. Like you get all seven of those and it's like, oh yeah, this thing strong. But four of them are are pretty uh are pretty big deal breakers. We'll send the deck out after, James.
16:05 Um this is just I don't have it uh posted anywhere right now, but I'll send the PDF out. Um So customer value, uh leading, and can you move it, and not vanity. Those are Those four are really important. If you miss any one of those four, um basically your North Star metric I I would say is out, no matter how it does on on the other three. Uh the easiest way to see is an example of this, right? So um an example everyone knows, like Airbnb's North Star is nights booked. Um and so if you walk you through that check checklist, a booked night is real value.
16:47 Someone's actually staying somewhere. It leads to money instead of trailing behind it. Uh the team can genuinely move that metric, and it's not something you can fake. Uh so it checks all those boxes, and that's why it holds up. Um I'll tell you there there was a company I once worked for that counted what I mean what a vanity metric that counted their weekly active users. Uh not just the people who are on the platform, um, but of people who would just like open any email from them. Uh, so it's a very gameable, uh, vanity metric, especially when certain, uh, certain like, uh, mail things like like iOS just automatically say your your email's open. So, if you just send more emails, suddenly your weekly active users go up. Don't do stuff like that.
17:40 Um, definitely don't create that metric and then go public and not be able to change it for years. Makes things very very hard. North Star metrics are hard enough when you're like, uh, internal to your team, but some of these inevitably will get out, trickled out to the board. Now you're reporting something to the board. If it's like a if it's not hitting those, uh, seven things on checklists, you're you're giving bad metrics to the board. If you're a public company and then the board wants to have that like reported on in your whatever quarterly earnings call, now you're like you're just stuck with this thing forever. You have to basically go private in order to get rid of that metric. So, it's really important.
18:20 Um, and all these decisions happen off of it. >> [clears throat] >> So, let's get a get check from folks here. I kind of already teased this, but daily active users, one or zero, good North Star or not? Type one for yes, zero for no. Kevin's like, "It's freaking good, man. Shut up, Shawn."
18:54 I guess it depends how you actually define active, which is a very hard thing in itself to define. Cuz Oh, man. Yeah, I mean I've been in I've been in months back and forth with people who try and do like a daily active user and then people like, "Oh, but active is like when you do this thing three times, and this other thing two times, and this thing maybe if you don't do this thing you have to do these other seven things on that day, then we'll count you as active. It's like, oh god. You have just turned uh 20 metrics into one metric somehow, and that doesn't make any sense. Um All right, we'll keep going. We're going to come back to this. We're going to have Claude tell us the answer to this.
19:35 Um So, one metric at the top, uh North Star metric alone isn't enough on its own, cuz you can't reach up and move it directly all the time. What you actually move are things underneath it. Uh these are like input metrics. And there's a simple way to make sure you've kind of covered exhaustively all the uh input metrics that drive our North Star metric. So, there's four kinds, you can think of them as levers. So, there's four kinds of levers. Um there's like how many customers, how much each one of those customers does, how reliably it goes through, and how often they come back.
20:18 And so, those link back to breadth, depth, efficiency, and frequency. BDEF. Those are your levers, and the North Star is the outcome they add up to. So, as those levers go up or down, there should probably some metric at your company or multiple metrics that relates to each of these uh you'll be able to drive your North Star metric. Um we don't just work on these the BDEF ones because uh those are uh a lot of them are gameable, but like the North Star metric itself is not.
20:56 Um let's see. Yeah, one thing that trips a lot of people up, uh your input can't just be like your North Star metric wearing a different hat. This is really easy to do actually. So say you pick like weekly active buyers um from your site as your North Star. That's a head count, right? And then um for how many customers you write like active buyers, that's basically the same number. So you didn't break it down, you just kind of renamed it. Um the way else like to pick a metric you can break into pieces a count of something happening like orders or nights booked. Then the levers underneath are are real and separate. Um and in a second you'll watch a tool catch this mistake on its own.
21:47 Okay, so let me bring in the AI and I can kind of show you a little a bit show you around a little bit this the uh the repo here and then we're going to run some stuff. Yeah, so true Nikita. It all comes back to prioritization of what we're going to work on honestly. Um All right, so this is our AI analyst plus repo. It's like our AI open source AI analyst repo but it has it's like the currently active developed one. So we've been the AI analyst repo we released I think in February. Uh we've been constantly developing on it but in this plus repo um and we'll probably push some of this stuff out to open source um soon but um we do anyone who's uh in any of our courses uh they get access to to this uh private repo as well. Uh what I'm going to show you today I don't even think it's pushed into main on that one but there's a skill here uh North Star and it has uh you know we could take a look at it.
23:01 It has um Yeah, I'll just read the top to you for now. But uh at the high level description of a North Star metric, so North Star metric life cycle coach, it helps PMs design, audit, defend. It could doesn't have to be PMs, it could be anyone. Diagnose and evolve their team's strategic anchor metric across um their North Star metric journey. Um it's cited. Uh so it it it remembers the product across sessions. It grounds every claim in uh the Amplitude playbook, which I were partnering with them today.
23:35 Um it's curated to that casebook. Uh and uh then it just has some things that like when it all when it triggers here. But it has a few different um uh procedures [clears throat] it runs. Uh audit. So this in this case a user has a a candidate metric already. And they want to go through like evaluating to see if it's a good or bad. See if it's weak or strong. Uh we're going to run the audit today as the primary thing we'll we'll run through.
24:05 We'll also run through drivers if I think we have time. Uh triage, should user wants like is this even worth audit? So this is like just like a quicker run of it. Although although honestly the audit takes not that long, you'll see. Um explain, user wants a cited explanation of a North Star metric concept. So this is a little bit different where you went with this one. Rather than having just like a single skill that has kind of all the information in it, we have a whole like Wikipedia of of concepts and beliefs of cases um that to point back to strong North Star metrics of anti-patterns.
24:42 Um >> [clears throat] >> right? Like gap thinking, um insisting on multiple North Star metrics. So it's like it's not just how to like create your metric, it's like all these things that happen in in real life. Um debates. It It different information for different verticals like consumer subscription consumer subscription versus dev tools versus B2B SaaS. Um So, this thing goes into pretty far depth and you can basically have it explain any concept to you as a teacher as well as when it makes any decision it'll cite where it's pulling that from.
25:18 Um Draft, this is something I'm still kind of working on but this is like when you want to design one from scratch and you don't have something to audit. You know, usually where you want to start with is is working with your team and trying to come up with a few yourself and then it'll make recommendations off of you for you but having having a place to start from is really important. Um and then um yeah, inputs and drivers are are pretty similar uh modes of this skill where it kind of decomposes the BDF thing and goes through your history from whatever window of time you select and it tells you how much is each of those um metrics driving your North Star metric.
26:01 Okay, so that's the real high-level. Um yeah, if you join our course you can go through this more. Uh eventually we'll get this open source but it might be a a few weeks cuz we got some other stuff going on right now. But let's go ahead and go into the demo. So, I'm going to earlier, right? I asked you like is DAU a good one or not? So, let's see. We're going to use the North Star skill.
26:28 Come on. Uh if you can't see this that well, you can zoom in on Zoom yourself so I don't have to zoom in here and zoom out all the time. But when you do North Star, I'm going to go into audit mode and I'm going to give it a metric, DAU. These first few ones should run pretty fast. Later on, it's going to I think this next one I'm going to run is going to take a little longer, so we can take questions while it's running.
27:05 So, in this case with DAO, so it's it's opening, it's triggering the skill, right? Um it knows that it needs to run the uh audit mode, so it opens this verb under the skill audit. You can put it in preview, it's still easier to run. So, when it does this, um you know, it it uh pre-filters anything already cleared, so uh if if if if the uh if there's any anti-patterns, it's not even going to run everything downstream.
27:41 It's pointless to go and run this rubric on it if we already know um it's like a anti-pattern uh metric. So, you can see here, pattern refusal confirmed. Anti-pattern page has a real fix recipe with a concrete reframing example. Um so, in this case, DAO is a really common one that people think are North Star metrics, and so uh and we can we can have this explained to us a little bit more, but you can see it refused. Uh DAO matched a canonical bad pattern.
28:14 Um we can ask we'll ask it about this, actually, in a moment. Uh accounts logins not value received. So, it conflates with notifications, required logins, or addictive patterns without improving customer outcomes. Uh so, it conflates value delivery sessions with vanity sessions. Um you can definitely define something where you have a this is a uh a DAO of a value delivery session, then you have to define what a value delivery session is. That's fine, but DOWs itself um it's going to reject that as a uh vanity metric. So, it has a test here. If a metric only goes up, counts raw activity, and you can't explain in one sentence how its movement reflects customer value, it's probably a vanity metric. Um and then it gives you uh some other options. So, uh try replacing daily active users with the value delivering action a user takes when the product works for them. This way, happy deliveries dropped people opening the app in favor of deliveries with no issue. Uh which feature shows correlated with retention and uh CLTV. And so, and then it sources everything, right? So, we could go into uh the wiki cases uh happy deliveries and um you can read all through this kind of uh case study around um why this is more a better uh metric to use. So, we try and really for this one So, you're going to you're going to get people who ask you why. Why this? Why that? We really try and like ground everything in like research that's already done, studies that have already been done.
30:06 Um so, you're not making this this like this is just isn't new stuff. So, we don't need to start from scratch. Um now, if we want we could even have it explain some of this stuff. Oh, it even recommends it here, right? So, it says next steps uh you can either uh we're going to do this actually. Uh we can first have it explain it like if we want to dig into why activity versus value. Maybe we're like pushing back on this.
30:37 >> [clears throat] >> So, we'll say North Star Um and you can look if you wanted it to look at the explain verb over here. Um So basically what this does, it kicks off a agent, the librarian uh agent and um that agent is going to go uh read through our glossary and articles um and resolve the kind of uh the concept that we're asking about. Uh if you want to look at agents Let me close some of this stuff.
31:23 There's an agents folder down here and we have a special group of agents that are all around um Northstar. And you can read more about like what that agent does here. So the dispatcher gives you a resolved article um >> [clears throat] >> adapts the format to the expertise level so we can uh have more in-depth or higher-level explanations, uh applies all the citations to different uh cases or to the Amplitude Playbook.
31:58 Um yeah, it surfaces if there's debate around this topic and there's a whole thing in our knowledge base around debate cuz everything's not cut and dry, everything's not you know, people have different opinions. So you don't have to um you can make your own uh opinions on your own by reading through the debates. Um But what it gave us here is it says uh you know, the Amplitude Northstar A lot of this is going to reiterate what it already just told us but in a little bit more depth.
32:30 Um yeah, so we already read basically the decision rule and the high-level thing before uh DAU, ad impressions, downloads, page views, registered users, uh story points delivered, um yeah, these all have like Oh yeah, story points like feature factory stuff, right? Um time on page. Each of these kind of share that same shape of a raw activity count standing in for value. And I just guarantee I've been on teams where these are the North Star metrics every time. Like but they are actually they're so gameable.
33:09 Um And then so the fix they say is like like I said before like yeah, you could tweak it. You could try to be like, "Hey, a daily active user that actually gets this value." But actually the fix is more like just find what that action that value action was, and that's probably your North Star metric. Okay, so that's a little bit around vanity metrics, a little bit about how the audit works and triggers when you get one. Um So it's it's it is a lot of it's brought in from the Amplitude Playbook, Nikita, but there's also a bunch of other resources around metrics. Um So that's not the only source, but it is a lot of it is derived from that and it's derived from the cases that it's based off cuz that playbook is still sites a lot of um material also goes into that source documentation as well. And it's kind of something it's that's ever-expanding as well. But it's pretty good playbook.
34:04 Um Okay, so let's reframe to another thing. So instead of DAU, we're going to do Where we at on time? Are we pretty good, I think. We're going to do North Star. This is going to take a little while to run, so we can go into some questions while this happens. Audit and we're going to do weekly completed orders.
34:35 Okay, so now I'll try and do a real one. This time it should run the full like seven as full like seven question check. Those four that really matter. But if we talked about and then those deal the deal breakers, customer value, leading indicator, uh you can actually move it. It's not a vanity. And it flags anything, um like if it's off like like completed orders is a little generic on strategy. Um that's the tool, you know, being honest with us, not rubber stamping whatever we type.
35:09 We actually like, you know, we don't want this thing to be sycophantic when we're creating metrics. So we this is kind of a make it hard on us. >> [clears throat] >> Um okay, so it did not pick it up as a vanity metric, so it is not refused. It uh is going to keep going with it. Novamart is this data set we use. It's a synthetic data set we built. It's like an e-commerce site. It's just flagging here that e-commerce isn't one of the industries um we have put into the uh the industry list here.
35:49 So it needs to kind of like reason through that on its own, which is why we are leveraging agents and it's not just like a pure DAG cuz sometimes stuff's going to miss, but it's smart enough to figure out uh what's good for e-commerce. Okay, while this is running, cuz this is going to be a little bit. Do we have any questions? Any questions coming through the chat? Yeah, James, that's the right that's the right one.
36:21 >> I think Val had a question earlier. I don't quite recall what what the what the uh what the context was. Balance, do you want to ask? >> Oh, no. No. That was just a comment to someone else in the in the chat. You can you can skip that one. >> Okay. Cool. Thanks. Um if anyone has other questions, feel free. It doesn't have to be North Star metric or related to could be pure genetic analytics related.
36:56 All right. So, you can see it's going through the questions here, right? So, customer value. >> You have your hand up. >> Yeah, yeah. What's up? >> Um so with a lot of these metrics, how much help or how do you suggest providing context or I guess around the business that you're at? Like do you actually suggest doing that? Like throwing together some sort of a markdown or something saying this is my business. This is kind of where it's at to like also guide that conversation or like right now I'm assuming it's I don't see that context for like what your business is specifically.
37:28 >> Yeah, I I would definitely recommend that. So, we have So, it does have like some some uh let's see. Some kind of businesses in here, but it's nowhere near exhaustive. It's just like some examples and they're pretty high level. Like it has like cuz like each each type of business is going to kind of play a different type of game in a way. Like the e-commerce is kind of like a transaction game where it's like we're just trying to get people to have more transactions, buy more stuff. Um a say like a social media site is more of like an attention or engagement game. You want people on there. They're They're getting entertained, right? They're not We're not They're not there to become more productive. And so, for your business, the more you can add and you can make like a full wiki on your business and clear out all this other stuff, that's going to help it um have the context of like what metric is most relevant to you.
38:29 Like I worked um previously in legal tech and I don't think there's anything on here for legal tech. So we we need to build that out if we were to use this there. >> Good. Thank you. >> Yeah, no problem. Okay, it finished the checklist. So what it did is it went through those seven questions that I showed you in the slides earlier and it actually said this is a it passes as a North Star metric, but it's weak. So it had five of seven checklist criteria pass.
39:05 There's no fatal failures. So like the four fatal questions customer value um not vanity uh actionable I can't remember where where where where the other one is. Uh measurable, like it passed all of those um but it didn't pass uh all all seven. So >> [cough] [clears throat] >> Okay, it actually says like this is one of the reasons it says weak for vanity.
39:38 Weekly completed orders isn't classic bad vanity. It counts a paid transaction not a login, but it's issue blind. It counts orders regardless of returns, defects, or late delivery. And then it gives you a fix. So this is pretty good, right? So the fix would be add a value received qualifier. Weekly completed orders with no issues, meaning they're on time, not returned, and no escalation. That's a pretty strong North Star metric over just like pure orders. So you can see it gets pretty deep. Like if I was on a team, it's like first of all, like would we have gone to orders over just like active users? I don't know, but if we did get there, I don't know that we would have necessarily um filtered out like um returns and such.
40:26 And then vision strategy, it also said this is a bit weak. It says generic, could be any stores North Star metric, non-fatal, optional polish. So, I think um this probably is where it comes in to be like uh kind of like you just said, balance like providing more context. It could probably give us like a a better fix on something that's more specific to this specific e-commerce site. Um I don't know, like if there's some some sort of some sort of order or some sort of product they really like specialize in or something.
40:59 Okay, and then it saves uh a more in-depth um readout of everything here in a markdown file. And so, you know, you can go through and read in more detail if you wanted to around like um did it pass or fail or just like partially pass each of the questions, and then why that occurred. And then for the ones that it failed, what is the way to fix it? So, it gave us like a very like one-liner fix here we can see.
41:39 Goes into more detail around how to fix it. Um And then it also ties you back to always any sort of um kind of like relevant stuff in the industry or cases. So, it gives you similar cases. Um gives you some recommended next steps. Okay. And then this frozen context block thing, once you run this thing for a certain metric, uh I cleared this out before we came here, but it then saves it into a JSON file. So, the time if someone runs it, it can just be like, "Oh, like you ran this last week. We don't need to wait 5 minutes. We already know the answer to this." And it'll just draw it up for us.
42:22 Okay, 15 minutes. I think we're good. I want to I want to show you guys a little more a little deeper demo. So, let's say we have actually uh decided we're going to move forward um with this and now what we want to do is uh we want to understand the drivers of this North Star. So, we're going to say North Star. This one might take a little longer, so we can do questions again.
42:53 Um but we will run through some slides and come back. Uh so, North Star drivers uh what was called? Weekly completed orders. And it might have some follow-up questions based on um the metric. But it's going to kick off this drivers workflow. Which if you wanted to again, you could uh read about this skill up here.
43:35 Yeah. So, the purpose is to decompose what drove the North Star window over that um you know, breadth uh frequency, efficiency, and uh depth. So, it's giving a follow-up questions. It's basically asking if I want like a full like what what date window I want to look at, right? Cuz the uh when you're we're now getting into something like a like hey what's driving in North Star it's going to be what's driving it over a certain time period cuz that can obviously change over time. We're going to say we want the full past year.
44:19 This data is fictional but there's like a few years of data I think 2023 2024 so it should pick like in 2024 Jan to December. So it's going to um basically decompose this into those four. Okay. That was way too fast. What happened is I ran this earlier so I just picked up my last run actually.
44:54 But what it would and that's probably fine cuz we only have 12 minutes left anyways. But um basically what it would do when you do drivers I'm just like I'm telling it tell it weekly completed orders as a North Star breakdown what drove it this year. Um it's pulling the orders breaking the metric into the levers we just talked about and looking what actually moved over the year and then it hands you back an actual report not just like the numbers um in the terminal so you can see it has this report here.
45:27 Um what it found uh or the orders grew six times over the year uh which on the face of it looks like a fantastic year, right? But when you look at the breakdown almost all of that growth like 99% of it is coming from one single lever which is just more buyers showing up. So how often people come back to buy didn't really move. The size of the average order um AOV which is our depth guardrail here actually uh went down um about 9% and the membership program that's supposed to bring people back in is like sitting at 3% or 2% so it's like basically not moving or or adding anything to the uh to the North Star metric itself. So that's that's interesting, right? Like the headline number is telling you it was a great year but when you actually read the breakdown what it's actually saying is you're growing by kind of like renting new customers.
46:31 You're not keeping any of them. Uh and the one lever you'd use to fix that uh of trying to get people to return seems to be broken. So it's the same 6X either way. Um the number can't tell you which one of those was true by itself which is why you have to look at the inputs which kind of goes back to uh what was Nikita was was talking about where it's like, "Okay, if we're going to prioritize work right now, we know breadth is good. We can get a one-time person to come by stuff really easily. We're actually seeing a ton of growth in there but we're not having uh uh returning customers and then whatever these customers are buying over time is actually decreasing in their value. So you know, you probably want to look at uh frequency and um obviously uh depth as uh things you want to increase here and efficiency, too.
47:19 But breadth, you're good. Okay, I think we had some slides on that we could just skip through them. Yeah, I was just going to talk through what it said. Um Yeah, step back. Everything I showed you grading the metric breaking it down throughout the year by the different input metrics, writing the report. The AI did that like really quick, right? Like it does it in minutes.
47:51 Um what it can't do is then decide it can't really decide what you measure in the first place. Um you know, your team your company has to like come up with that context uh to start with. We're working a little bit on having it just like come up with metrics for you from the start based on like kind of like the industry practices. Haven't kind of rolled that out yet, but it's still going to be unique to whatever your team within your your company's doing, especially if you're trying to do something new that people haven't done before. Um, they can't look at that um the North Star metric just like on its own and know like whether to celebrate or worry. It has to drive it down to those input metrics. And then the judgment around like, "Hey, looking at these input metrics, which one of these should we prioritize a team works on based on like how much lift that's going to be?" That's again like that's the human's decision. So, everything else kind of stays the same, but it's pretty good at like helping you create a strong metric and understanding what's driving it.
48:53 Um, so if you want to learn more about the kind of like building these systems and also how to figure out like that kind of human decision part I just talked about. Um, so turning data into turning data into a decision and not just a chart, there's a couple ways to do that. Um, we have a boot camp. We have it coming up tomorrow actually. And then another one we're trying it a little different with a boot camp tomorrow. We have a boot camp for a week weekdays in July for mornings for a couple hours. Um, that's where you build an analytic analytic system yourself.
49:32 Um, and then we have a follow-on. That's our advanced uh AI analytics course in the end of June. And that's like taking it more into uh production. So, it's getting into uh validation, context engineering, multiple models. And then we have our five-week builders course. That kicks off on Monday. The next time it kicks off is uh in August. That's how to frame questions. Basically, it's an entire week on this kind of stuff we talked about today, but then it gets into um you know, obviously, how to frame questions, pick the metric, and dig into the data. But then uh in the subsequent weeks, it goes into like root cause cause analysis, correlative analysis, um uh trend analysis, uh segmentation. And week four, it gets into causal inference and experimentation.
50:25 Uh in the fifth week, it gets into presentation and sharing out with stakeholders. Um the overlap between these courses, James, I'd say is just the main overlap is basically just like uh setting up, installing Cloud Code. >> of us is the overlap. >> Ha ha, yeah, the three of us is the overlap. You guys are the overlap cuz you're going to take them all. I'm just kidding. Although, we do have a uh a two-for-one deal with the five-week and in the first like the build it course, not the scaling one. Um but the overlap uh truly is just that uh in the beginning of uh the boot camp, we just make sure everyone's running got Cloud Code installed, can clone the repo, and run through the analysis the first time.
51:05 And we do that same thing in like week two of the five-week course, but everything else is different. So, the build it and scale it are all about building Agentic systems. This is more around development. And then the AI analytics for builders, this is more about learning the analytical frameworks. So, it's pretty similar to what we did today, actually. We spend, I would say, about 70% of the time, 60% of the time teaching uh the kind of like best-in-class analytics thinking and framework end-to-end. And then we teach you the remaining whatever 30% at like, "Okay, how do you execute all that uh leveraging AI or Agentic systems rather than writing SQL or Python or R directly.
51:49 Um Yeah, so that's I don't know if that that answers your question. And then we do run a two-for-one deal. Um So, uh you save builder bundle boot camp and advanced. Yeah, okay. And then so the uh if you do uh the builder's class the five-week class, we just we we toss in the build that one for free.
52:26 And so yeah, it's basically a two-for-one deal. I think they pair really nicely. They're complimentary. They don't really overlap. You can do one before the other. You could do the boot camp this weekend and then decide you want to take the five-week course uh 6 months from now and I just take whatever you paid for the boot camp this weekend and deduct it from that. You could take the five-week course on Monday and then decide you want to take the boot camp um 6 months from now and I just add you to the boot camp for free. Or you can do add them both at the same time.
52:59 Um Yes. And then we also do a discount with uh the advanced course if you take um uh the boot camp and then want to move on to advanced. There's some codes down here. Any of these if you want to take them solo you get 20% off. But if you take the five-week you you should just take the boot camp as well cuz it's free. So, today Oh yeah, I'll send the recording out in I don't know not too long.
53:36 The price for boot camp plus builder. So, it's 1,800 is the uh builder's course. The five-week course is 1800. And then we put in the boot camp for free. So, it's 1800 for both, so you save yourself 900 bucks. The boot camp on its own is 900. And then you can do 20% off, too. So, if you just want to take the boot camp, you know, it's it's actually like 720 or something. Um and if you and if you guys want to what have more questions about that, uh you can DM me in Slack or on LinkedIn or email me. I'll drop my email in the chat.
54:19 Yeah, boot camp's this weekend and then again in July. We're doing some new stuff with the boot camp, so we really take the feedback from people who've gone to the cohort seriously and we iterate a lot each time. We'll continue to iterate. We do it for a couple of reasons. One is cuz we're learning what people want. As we want to like constantly iterate on the feedback with that. Also, this the industry is moving so fast, so we just are adding new stuff to the boot camp all the time. Like we're going to add um some stuff on open source uh this round, for instance, uh because seems like those models are really catching up. And also, on ICAs I've used like the Fable model, it's really freaking expensive. Like it just like eats your tokens. So, I feel like there's also like we want to be on your side and make sure that uh at your company you're just making like the best decisions and not just like on the newest best model because it's the newest best model, but actually be on the thing that works gets your job done. Which I would say is like a few model releases ago right now, like probably up to 4.6. Um And then This says, "Can this be run in VS Code with Claude as plugin or does it need Claude Code IDE installation?"
55:32 Yeah, you run in VS Code. I run in VS Code. Um but I I do it in terminal on there rather than the plug-in. And then the the terminal, the other reason I do that is cuz like you don't necessarily want to be tied to Anthropic models. You might want to use Google models or OpenAI or open-source models, too. We're not going to get that fully into that in the boot camp. We get to a little bit. The advanced boot camp below we'll spend a bunch of time on that.
56:03 Something else we're doing for we're testing for this boot camp, too, is we're going to have 2 days. We usually just do 2 days, Saturday and Sunday, 4 hours a day. We are going to add a bunch of bonus content this time based on the kind of questions we get people from people. We always get questions around open-source, context management and validation, and data warehouse connections. So, throughout the week after Sunday, the boot camp doesn't really stop. Monday, Tuesday, Wednesday, Thursday, we're going to be releasing async content on those four topics, validation, context management, open-source model, and data data warehouse connections.
56:37 We'll have practice exercises. So, those are optional bonus. We want you to be able to like have it sink in as you use it at work. And then Friday, at the end of the week, we're going to have optional office hours. So, you can do the boot camp on the weekend, apply it at work throughout the week, go through the async content, and then on Friday, if you want, we can answer a bunch of questions live. Of course, the Slack's always open, too.
56:58 And then in July, we're experimenting even more where we are not going to host it on the weekend. We're going to host it on weekdays. So, it's not a like intense 4 hours, but we'll do like 2 hours each morning, and then you can like go apply stuff at work at all day or on your own personal projects, and then come back the next day, ask questions. And it's a little more piecemeal that way. Maybe we'll and we'll experiment with other stuff in the future if there's other ideas.
57:27 Okay, I know we're at 1:00. I can stay over cuz I quit my job last week, so this is all I do. Um but if other people have jobs to get to, that's cool. No worries. Um but I can stay over people have questions. I don't think I had anything else here. Yeah, my last slide's questions. Oh yeah, I'll send out the recording of course and I'll send out the PDF. I'll send out the PDF today with everything. The recording uh either tomorrow I'll probably send it out. And then uh I had that little like worksheet thing that just has the questions and stuff to ask. Um I'll send that out today, too.
58:08 And then sometime in the next few weeks if for those of you who aren't joining our courses and don't have access to AI Analyst Plus, I think in the next few weeks or a month or something we'll probably open source all that. When that happens, I'll kind of blast out our whole email list which all of you are on. Um so you can get access to that, too. But if you want it sooner, um yeah, you you would get it basically uh tomorrow if you took the course.
58:36 Questions. No questions is fine, too. >> I think Nikita had a question about uh so according to this framework, does the North Star metric change over time? >> Oh yeah, like so your North Star metric is definitely going to change over time and that's why we have the why we have this like audit thing. So like um Like it just depends what your team's working on and your team's mission. I guess like if if your team's mission is consistently like the the same thing, then like yeah, it might stay the same for a long time.
59:24 I would say you want to audit your North Star metric quarterly. I'd say the first time you make your NorthStar metric, you want to audit it weekly for like the first month or month and a half to see if you actually picked a good one if you've never created one before cuz it's really easy to be like, this is the best. Everyone agrees. And then you start building stuff and then you see the actual data and how it's moving and it's like do some analysis and it's actually not a great metric and you and you kind of flag all these sort of things. And then I'd say probably revisit it quarterly if you're kind of like what your team's working on and your mission stays like the same, it's probably not going to change much. Could last throughout the entire year. At a very minimum, you definitely want to visit it annually. But I would say like any quarterly planning cycle, think about your NorthStar metric again.
60:13 And then you can like rally around it for that quarter. Depends on like the cycle like the product and company you're working on, too. Okay. Other questions? Cool. All right. This is great. Hey, thanks for the 20 of you stayed over.
60:46 Um and Oh, how do you find NorthStar metrics in the B2B space? Yeah, I think it's just like um it's just going to be like when you when you have B2B like B2B is when it actually becomes really important to uh create a NorthStar metric cuz it's so easy to get caught up in like uh um I'm just developing for like the loudest customer in the room. But like it's the same thing. You want to do it at the user workflow kind of level and hopefully your product teams are kind of organized in such a way that each of them are focused around a user workflow.
61:27 So like um an example of like a good North Star metric for uh one of my one of the team I used to work on where is we were automating out part of someone's workflow trying to make it faster. And so it was like time from the start to the finish of a successfully completed task without any like errors in it. But it's like it's a I see like yeah, think of like what is the user trying to do in the real world.
61:57 And then how can we reflect the success of what they're trying to do in the real world within our product. And then what are metrics that can basically reflect that action in our product. >> [clears throat] >> Yeah, balance. >> Uh so I got another question on this is more on the AI analyst uh lab and it's I'm trying to think about how to apply it like pick a pet project in some my area I'm trying to pick is employee engagement service like around an HCM. I'm going to go through the analysis process. So is the AI analyst the free the free repo you guys have a good starting place to sort of kind of build that out and learn more about like how you set up that repo and adapting it. Or I mean are there any gotchas or places and I can post this in the slack too.
62:49 Uh just thought I wanted I wanted to ask that. >> I think it'd be pretty good for employee engagement survey like cuz that's not going to be I think that actually be a really good place to start cuz that's going to be like either one table or like a CSV or something. You're not going to have to worry about adding a bunch of like semantic models around like hey you have to join to these different tables.
63:14 Although you could if if wanted to like learn specifics about employees later. So, I think that'd be a pretty good start. I would basically download that CSV or can it within into the repo and then start kind of like asking like what question or asking it to like profile the data. There's like a data profiler function and then start asking it like you know, can you help me frame some questions around and then give it some topics and start from there and see where it goes.
63:43 >> Okay. And and my next step was is thinking about it like in a wider platform of like an HCM system where now you have like payroll and like job promotions and stuff like that. That's where you start looking at like, okay, defining the table schemas for what those other tables are and some of those other I think there were schemas and you mentioned something else, too, right? >> Yeah, there's going to be like Yeah, you'll want like the schemas. Um you'll want to have like kind of like these uh semantic models and and kind of like knowledge bases around kind of what we're talking about earlier, the context of your company, like what are the metrics that make sense? Where are the Where are the the definitions to all these things? Like there's just going to be so much stuff that it can kind of take guesses at it, but it will need you to um kind of like formalize in a document somewhere in the repo that it can like read every time so it's not guessing every time.
64:36 >> Okay. And then just the last one, what do you recommend for like a Do you just recommend raw CSVs if you're just trying to like synthesize data to like play with and like check test test how well it works and like kind of verify? Is that >> I think it's a good place to start, but you can connect to like a data warehouse or something pretty easily and it'll run SQL for you. Obviously, there's just like more a little a couple more hoops you have to jump through to like do that connection, but um as long as it's not uh I mean, if it's sensitive like information, you probably don't want to like the CSV onto your computers or something. But like if it's not something that >> It would be synthesized and fake. It would be synthesized and fake.
65:19 >> Yeah, yeah. I would just play with it in CSVs to start. That's what Yeah, I mean in our in our like boot camp stuff, we make like a local database. Um that it queries and stuff since that's like a little more what a lot of people are going to use it for. But a lot of the time when I'm analyzing uh data, I'll just like download some stuff from the internet as a CSV. Um if there's multiple tables, I'll have like I'll make it create a DuckDB database. It's really good at that. But >> Okay.
65:47 >> if it's all in one place, yeah. >> Thank you. >> No problem. Yeah, I think James there is uh there's an interesting thing there around like yeah, customer requirements. I mean, I'm sure you could create s- There's a customer wants and there's customer needs. And one of them provides real value, right? And so it's like this is like just like a hard thing with B2B of like am I building something that provides value to multiple customers or is just this something that this one person really wants and then they're going to leave. That's just like a balancing act for any B2B company that they have to make a decision on. Um but I think like you know, something like uh finance compliance or something where there's like true compliance things, I think you could make out uh like did this You could you go through and and ask the the the North Star metric thing to to run this. Um but like I'm sure it has some stuff in there cuz there's a whole fintech industry thing it has where it's like I bet there's like a thing like uh how fast um and how often do you run through like this without hitting any of like I don't know compliance issues in the workflow. I I don't know because I don't want to work in that space, but I can imagine there's probably something like that.
67:17 Uh we're going to talk about that a bunch in week one of the builders course, though. And then the other thing with the builders course, just to give you guys a little more detail, it's a little different from the boot camp. The boot camp's all live, except for this bonus material we're making. The reason it's all live is because the industry is moving so fast that when we're teaching people how to build a genetic systems, like I don't really want to spend a bunch of time recording myself teaching that for it to change in 2 months and have to record again.
67:46 Uh that's why we do that all live. The uh the five-week course, a lot of that is around kind of like best and practice industry standards for like, you know, analytical thinking and applying these stuff. That hasn't changed for a long time. It's a little more like evergreen. And so we do is all the lesson material is async. They're anywhere from a 5 to a 20-minute videos, about 3 to 4 hours per week. 1 hour is like you should watch all this, and then there's like 3 hours of kind of optional async content. Totally self-paced through the five weeks. Um and then we have live office hours for 3 hours a week. Uh where we do more Q&A about whatever we talked about that week or what we talked about in prior weeks.
68:30 Just depends on who comes. And we stagger those uh office hours based on the time zones of the students who are in the course. So we'll kick it off in the morning Pacific time for the first week, but then we'll do a poll and we'll pick office hours based on who's in there. So like yeah, like we're probably going to be like evening mornings, so we can hit every time zones. Um but we'll talk a lot. We keep talking in extreme depth about kind of uh your specifics of like B2B SaaS metrics there.
69:00 In those live office hours. Or, you know, we can talk in another free thing here at some point. All right. My wife has called me three times in the past 5 minutes, so I think I have to leave. Uh I just keep rejecting her phone call on my on my watch here, so I don't know what's going on, but I think I have to go. But, uh thanks everyone for joining. Uh and thanks for all the engaging questions. And um I think next week next Friday we have one on validating uh validating validating the output of the genetic analytics platforms. This will that'll be pretty good. It'll be like uh kind of like similar thing to like today. We'll talk about framework and then do a little bit of stuff in cloud code. So, uh check that out.
69:43 Hopefully, we'll see you there. Thanks, everyone.
Summary
- North Star metrics should capture customer value, reflect company strategy, and be actionable.
- Common pitfalls include using vanity metrics like daily active users that do not indicate actual user value.
- The framework used by companies like Airbnb and Netflix helps in selecting effective North Star metrics.
- Input metrics (like breadth, depth, efficiency, and frequency) drive the North Star metric and should be monitored for effective decision-making.
- The AI Analyst Lab provides tools and workshops to assist teams in developing and auditing their North Star metrics.
- Regular audits of North Star metrics are recommended to ensure they remain relevant as business goals evolve.
- The discussion includes a live demo of an AI tool that evaluates and suggests improvements for North Star metrics.
- The importance of context in defining metrics is highlighted, especially in B2B environments where user workflows vary significantly.