Section Insights
Introduction to AI Product Management Interviews
What are the key components of AI product management interviews?
The section introduces the two main areas of focus in product management case interviews: product sense and product execution, specifically tailored for AI. It emphasizes the importance of understanding AI product success metrics and outlines the value of the mock interview being presented.
- AI product management interviews focus on product sense and execution.
- Understanding success metrics is crucial for landing a PM job in AI.
- The mock interview aims to provide valuable insights for aspiring AI PMs.
Identifying User Needs for AI Editing Tools
How should we prioritize user groups for AI editing tools?
The discussion revolves around the importance of accommodating various user types, particularly those with high editing fluency. It highlights the need to balance metrics that cater to both novice and experienced users to ensure broad success.
- User needs should be prioritized based on editing fluency.
- Success metrics must accommodate both novice and experienced users.
- Understanding user value is key to defining success metrics.
Defining Positive Metrics for AI Editing Tools
What positive metrics should we consider for measuring success?
The section discusses various positive metrics, including time saved in editing, quality of the final product, and user engagement. It emphasizes the importance of operationalizing these metrics to gauge the effectiveness of the AI editing tools.
- Key metrics include time savings and overall quality of edits.
- Engagement and retention metrics are important for assessing user satisfaction.
- Operationalizing metrics helps in measuring the success of AI tools.
Establishing Trade-offs and Guardrails
What trade-offs and guardrails should we consider for AI editing tools?
This section outlines the need to define trade-offs and guardrails to ensure that the AI tools do not increase editing time or introduce errors. It emphasizes the importance of maintaining a low hallucination rate and ensuring user efficiency.
- Trade-offs must ensure minimal increase in editing time.
- A low hallucination rate is crucial for user trust.
- Guardrails help maintain the effectiveness of AI tools.
Linking Metrics to Business Goals
How can we connect success metrics to business outcomes?
The discussion highlights the importance of linking user engagement metrics to business goals, such as increasing subscriptions and revenue. It suggests that output metrics should be included in the overall success dashboard.
- Success metrics should align with business objectives.
- Output metrics like subscription upgrades are vital for revenue growth.
- A comprehensive dashboard should include both user engagement and revenue metrics.
Transcript
0:00 When it comes to PM case interviews, there are two major buckets. Product sense and product design and product execution and product success metrics. AI has its own flavor of both of these. AI product sense, AI product execution. Today we are giving you the very first full mock interview on YouTube ever published for AI product execution and AI product success metrics. How do you measure the success of launching GPT 5.2? How would you look at a dashboard of metrics for claude code? These are really interesting questions and these are the questions that companies like OpenAI, Anthropic, Meta, Google, all of the best AI companies are asking. So, if you want to land a PM job at one of the top AI companies, if you want to earn that five, six, $700,000 salary, you know, the average stock grant at OpenAI per year is $1.5 million. If you want to land an OpenAIMM job where you're making 1.8 $1.9 million, you have to nail the AI product success metrics interview. So today you are going to get an end toend case example of how to nail those with a 10 out of 10 response. And why are we delivering this value to you for free on YouTube? Well, Dr. Bart and I have been running a cohort program where we take 30 PMs, highly qualified PMs who have PM experience but want to become AIPMs, and we help them do that. We just are completing 9 weeks of the 12-week program of our first cohort and already 30% of the cohort has jobs. People have landed jobs at OpenAI, Anthropic, and other companies like that. So, we are going to share the types of questions that they were asked, how the coaching that we gave them so that they could answer correctly.
1:52 And if you want to go through a similar cohort experience, be sure to check out landpjob.com. That's where you can see all the details for this cohort. We have it is a premium cohort. It is not cheap. It is expensive. But if you are in India, if you are in Europe, we do adjust the pricing to your purchasing power. So be sure to check out cohort 2. We already have sold out onethird of the cohort. There's just a few weeks left for you to secure your spot. So if you want to get a spot to get interview answers like this, be sure to check out landpob.com. And now into today's mock interview.
2:29 Hi Akash, welcome to Descript. I'm Bot. I'll be your interviewer. How are you today? >> Good. Excited about this case interview today. >> So listen, time is limited. I want to get the best out of you. So let's keep the formalities, the weather talk, and jump right into it. We here in the script are launching a new feature called Underlord. So tell me, how would you measure its success? >> Okay, this is a fascinating question. Maybe to start, can we align together on what Underlord is? I'll just pull up the descript and we can just take a look at the feature. Does that sound good?
3:04 Perfect. Awesome. So, I'm going to share my screen. So, let's take a look and see what is the Underlord feature. So, it looks like we have different options. It looks like it's an AI co-editor to help you do anything. It's a chatbased video editor. So, you guys have made it really replace your overall AI editor, which is what I'm used to using. So maybe what we can do is we can jump into a quick project and we can see how Underlord works. So Underlord, you click it on the bottom right. Okay. You can ask it for suggestions.
3:34 It can give you creative ways. You can enhance a composition. Very cool. So it's going to be it's like your editing partner. And it looks like it can even execute on some of these. Let's go ahead and see what it says. Multicam layouts, B-roll, and screen recordings. These are all really good. So, this is giving me a good sense of how this product feature works. It's like your buddy as your editor buddy. And then what I want to confirm is like what type of tool use it has access to.
4:02 So, do you have access to do any of these edits or are you limited? I'm just going to check what kind of tools this agent has access to. Okay. So, you guys have built it so that it can take access of basically all tools. So I just wanted to confirm that in the first I am the PM4 underlord and it has access to all the tools in descript alternative to old style editing which can be frustrating where you want you know what you want to achieve but you don't know how there's lots of menus and well we people always communicated with language so why not do essential thing like video editing in 21st century with your common woods and you are tasked to well assess whether we were successful in this mission or not. Awesome. So, we're the PM of Underlord. We've confirmed that it has access to all of Descript tools. So, you guys have created, you know, MCP or API hooks into all of the editing features. can I take a minute to structure my thoughts and get a sense of where we should go from here?
5:14 >> Take your time. >> All right. All right. So I just created a quick framework. I'm not going to bother to draw all these lines in here because that's just going to distract us. But you can assume that I had like drawn lines from all of these features. So we do the clarifications which we talked about. We'll talk a little bit more about the product and what the product does, the users, the value that the product provides. Talk about positive metrics. So create a bank of positive metrics. Identify a Nordstar metric. Break down that Nordstar. Talk about trade-offs and guardrails. And end with a summary. How does that sound?
5:45 like a very structured answer. I'm happy and looking forward to hearing it. >> Awesome. Cool. So, in terms of clarifications, I don't have too many clarifications. I guess I'm just curious. Did you guys have a certain type of user in mind for this feature when you built it or should I just make some assumptions and think about what that might be? Think of people who are starting their editing journey who want to have an amazing videos like top creators but don't have the tools to do it yet as in skills and knowledge are a little bit not techsavvy enough just to like find the right manual or a video but still are very determined to achieve their goal.
6:32 >> Okay, great. That's really helpful context. So let's see if we can connect those auto. Perfect. So now we talk about the product a little bit. And from the product angle, one of the things that we confirmed already is that it has access to all of Descript's tools. So I'm going to go ahead and start to make some of this smaller here. Let's go ahead and make this smaller so we can start to look at stuff. So has access to all of Descript's tools. Basically the user can chat with it and request anything. it's also kind of like your editing buddy, right? So, we saw that example that we demoed at the beginning where the editing buddy suggested exactly how to improve the recording that I had. So, I think those are the main features of the product. Is there anything else I should be aware of before I move into users? No. Carry on to users.
7:22 >> Awesome. So, now we think about the users. So, you already told us a bit about the users. So actually what I'm going to do is I want to just think about these people a little bit further. People starting their editing journey, top creators, right? So there's going to be people with low editing fluency. They're not going to understand, you know, how to use an editor and so they're going to be very reliant on Underlord. And then there's probably going to be people with like medium editing fluency. Probably nobody with high because you said they're starting their editing journey. But these people, >> it's not limited to particular user, but probably if you have the muscle memory, you don't need the underlow to help you.
8:07 >> Yeah. Well, I imagine there are some use cases for high editing fluency, too, right? Because an AI agent might be able to perform an edit like 10 times or find 10 clips. So, it might be able to do things like add scale for you or save you time by removing all the or double edits or add a b-roll. Definitely there is value for any user using underlord. and well, the question is how how well should we profile that to certain specific user and where we will get the most success and how do we understand this success?
8:44 >> Awesome. Perfect. So in terms of like prioritizing user groups because Underlord is on the homepage, I really feel like the success metrics we need to have need to accommodate any user. And I know that you mentioned that we should think about people starting their editing journey, but I was wondering if I could push back on that a little bit because I imagine that the biggest power user of Descript is like an editor, somebody with high editing fluencies.
9:12 So, I'm thinking like even though our success metric needs to care about people starting their editing journey, we need to make sure it satisfies his high editing fluency. Does that make sense? >> It does. Yes. >> Okay, great. So, in terms of users, we're not going to prioritize any user is just the note I'm going to put here. and the reason being is because this is on our homepage and this is like a persistent agent. So be for those two reasons, we're going to make sure that our success criteria really satisfies everybody. So let's talk about the value that they get. We talked a little bit like they save time, they do it at scale, but I think the value enumerating it in even more detail is going to help us identify the right positive metrics. So one of the things right is it basically reduces the time it takes to edit, right?
10:06 Then it also so basically what we're doing here is like time to export or time to publish. We're kind of reducing that. It also hopefully allows you to do more edits, right? So we basically want to see like more exports, more publishes. as I'm thinking about value, I also want to go ahead and just talk about some of the trade-offs and guardrails, right? So we need to be really careful. And there's this concept of AI evals, right? And so we'll talk about that a little bit more when we get here. But we need to be really careful that it's not you know, it's not hallucinating that it did an edit.
10:43 And we also need to make sure that it has, you know, like a high quality bar for those edits. those are some of the things we'll need to make sure. And so as we think about value, that reminds me of that. These more exports and publish and time to export and publish, I feel like these are almost positive metrics. So I'll move those into the positive metrics. So, as I think about value a little bit more, is there more value we need to think about? Let's see here. Probably another thing it does, like you said, for firsttime editors, I imagine for new customers, one of the key values we're trying to drive from this feature is getting people to complete their first edit, right?
11:20 Because I bet there's a huge amount of churn or drop off from people who they record something, but then they never even complete a first edit, right? So, first edit completion rate for new users on first vid file. That would be another one. So, we're kind of tying these metrics to the value. If you can see here, some of these positive metrics. Let me think a little bit more about this value. And you know what? Maybe like to just get us inspired. Why don't we go back and just take a look at what we were looking at earlier, right?
11:51 With this. So, I'm going to shift back into the product for a moment here. It can do things like add chapter markers. Okay. So that's almost like helping you create like your YouTube description. so if I switch back then into this tab, it's like how would I even describe this? gives you the info to write up the video and publish. You know, things like chapter markers. And if I turn back to this tab and I look a little more, create scene boundaries to split the change the layouts, change the caption. So things like okay so each one of these features there's going to be like a feature specific value. So what I mean by feature specific value is that for the captions for instance we're going to need to make sure that the captions look good that they're accurate that they're in the language they want. Right? Those are just some basic things for example for captions.
12:54 So, because there's so many things going on here, I'm going to go ahead and take one second. I know it's a little bit distracting, but I'm going to take one second to just draw these arrows out so it's a little bit easier in case we revisit this information later. So, this just going to help us trace everything back. So, there's captions as an example. Another example might be like we just saw chapter markers, right? I'm not going to do every future, but I'm just enumerating for one or two so that we get a good sense of what type of success we're looking at. So, for chapter markers, we're going to have to look at there they're smartly divided. I think for chapters, you also want to make sure that they start at a good place. that they're not too short or not too long. Right? So, some of these features, these are the same metrics we would use for that specific tool. And so what I'm thinking after enumerating all these is let's not focus on the tool specific metrics. Let's assume that the PM of that tool is responsible for these type of metrics.
14:00 Is that okay with you? >> Sounds f. >> Okay. So what I'm going to do is I'm going to click on all of these boxes and I'm going to go ahead and say these are like red like or not red. Let's say like these are like yellow like these are not our focus. Yellow not our focus. some other PM. So, I think we've enumerated the value enough. I think the major values, and let's go ahead and draw arrows to those just cuz I want to summarize them. Reduce the time it takes to edit, allows you to do more edits, gets people to complete their first edit, and gives you the info to write up. I think these are the four major values. Did I miss anything? I don't think so. It sounds good. Awesome. So, now we need to get the four We need to across these four major values. Let's go ahead and make sure that we have all the right metrics. So, what I'm going to do is I'm going to give myself some more space here. There we go. And why don't we go ahead and start to connect some of this up just so that we look at it and it looks pretty. I know it takes a second, but for somehow for me visually, it helps me. Okay, so now we're into positive metrics.
15:07 So, we talked about reduce the time it takes to edit. Now I want to make sure that we have sort of a holistic idea of time to export publish. That's right. And then sort of the flip side of this, we need to think about the negative metrics or the trade-offs and guard rails assoc associated with it. So time hallucinating that it didn't edit. That's one. Another one is just like increasing time it takes to edit, right?
15:31 I think there was this hilarious study with GitHub Copilot where engineers were spending more time like editing GitHub copilot and then it made them slower than just writing code by hand. Right? So we don't want that same thing here. So we want to really make sure we reduce the time it takes to export publish. I'm trying to think about what other metrics am I basically what I want to look at is the AI makes a change and then the user changes what the AI did, right? So it's like edit the AI. We want to reduce editing of the AI. So that would actually be like a negative metric we can put down here. And I'm wondering like I'm wondering if I had the greatest framework here because I just separated positive and negative. What I'm going to do just to make it easier for us to tie everything back. So I'm going to change this arrow just so we know what value we were talking about when we come back to this in a second. Editing of the ad.
16:25 Okay. What other things in terms of positive metrics? reduce the time it takes to edit, time to export, publish. Is there anything else around that? I guess it's just like makes you feel like a better editor, right? >> Perhaps something to do with the overall quality of the final edit, like unlocking things that you would not consider yourself. >> I love that. Thank you for bringing that up. unlocking features you wouldn't use. And that was one of the cool values we saw of the product, right? it suggested here's 12 ways to improve that reporting that we had. So unlocking new features the user didn't use. So essentially what we would look at in this type of a metric if we were to operationalize this is like they never used let's say studio sound and now they do. So it's like number of AI tools used going up is essentially the metric we're looking at there. So, engagement, retention, I think both of these are different, right? We could operationalize these differently, right?
17:26 On the engagement side, it would be something like they're opening, they're spending more time per week. Well, is that right? We don't want them to spend more time cuz we just talked about reducing the time >> per edit. It's like they're editing more videos in script. >> We want to increase their efficiency. Really? >> Yep. Editing more videos into script. And then retention, right? like people are coming back, you know, day 7, day 30, these types of retention metrics.
17:53 Okay, so let's do a quick time check here. How much time do we have? We are about 20 minutes in. Okay, I think we're doing fine on time. Let's keep going in terms of this level of depth because I just want to make sure we choose a good north star. So I want to apply the same sort of rigor to all of these values. So if we think about allows you to do more edits, yeah, more exports, publishes, we already talked about that. New number of AI tools, that's kind of similar. So I'll also draw a line there. So that that makes that metric seem pretty attractive. Number of AI tools used.
18:28 Okay, I'll keep that in mind. I won't spend too much time making these arrows look good. Okay. And then is there anything else around allows you to do more edits, more exports, publishes? Well, I guess like when you have a long form, it's like more short forms per long form, right? If I create a 90-minute video and I create 10 clips instead of five, that would probably be a win for Underlord >> and for the users.
18:53 >> Exactly. And then getting people to complete their first edit. This seems like one of the most important. Absolutely. one of the most important, but at the same time, as we just talked about, our power user for Descript since this is going on the homepage is going to be somebody who is a professional editor. Okay? And then gives you the info to write up the video plus publish. So basically like copy paste and ex copy paste this info plus like accept it with minimal edits. Those would be the two metrics that we would look at. Okay. So now I want to talk through what should be the northstar metric. If you feel like we've covered all the positive metrics, do you think so?
19:32 >> I believe so. Yeah. Let's move on. >> Okay. Just for the sake of time, you know, in a real life scenario, I might think about this a little bit more in detail, but given it's an interview, let's continue forward. So, what should be the Nordstar metric? So, let me just talk about the pros and cons of each of these. How about that? So, time to export publish. The pro is this could be an all-encompassing metric. The con is it doesn't really cover quality and you never know like underlord could lead to a higher time to export because people are using more AI tools. So that I think is leaning me then more towards something like number of AI tools used. And we might even operationalize that further into like number of AI tools used with no edits.
20:17 You know like they they used studio sound and they accepted it. They didn't change it any further. They used the automatic layouts and they accepted it. They didn't go in and change more layouts. That's what I mean by this. So there would be I'll give you some examp like when we truly operationalize this. We'll have to do this for each tool which does make it a little bit difficult. but I think we could come up with a synthetic summary metric for it. Okay. So this could be a good one just based on the pros and cons of this.
20:47 Editing more videos into script. Now, that seems like something that really covers everything, right? And even if we did our AB test, let's say we rolled it out 9010, 90 10% of people get underlord, 90 don't. If we start to see the underlord group editing more videos into script, let's say like we operationalize this as like number of exports in 7 or 30 days, then we get a really good signal that they're they're liking the Underlord experience overall.
21:16 And to that point around liking the Underlord experience, I think one thing we should also think about and we haven't put that here. So, let's add that in here is like support requests from users, right? Because theoretically with Underlord, there should be less support requests. If we're creating more support requests, then that's a bad thing, right? So I almost think that's one of the that's one of the guardrail metrics we also need to think about is if we were to choose a north star like time to export, number of AI tools, engagement retention, any of these, we would want to see that they don't have supports increased. So right now where I'm leaning is something like number of exports in 7 to 30 days with the guard rail of support tickets.
22:02 allows you to do more edits. So more exports published. Yeah, that's similar to this number exports more short forms per long form. That's also captured in this number of exports. First edit completion rate. So we will actually see this in number of exports because this would be zero for people who have a first edit completion rate. >> Copy paste this info. This I think is just not as important. So what I'm going to do is I'm going to go ahead and give this again kind of like a yellow coloring like it's not as important.
22:27 It's important but it's not a north star. So based on everything we've talked about and all the pros and cons, I'm leaning towards this as a north star. and support requests is one of the key guard rails. What do you think? >> That's true. The team is not already using it, but go ahead. It's it's your northstar. I think we can proceed with it. >> Sorry, what did you say though? >> No, I mean I I think that's the actual one we're using, but again, it's it's your northstar. I'm here to observe your process and I like what I'm saying. So, go ahead.
22:55 >> Oh, awesome. I love to hear that. That's what you guys are using, too. So, yeah, let's break this metric down. We've kind of discussed like that's our Nordstar. So just to summarize where we've gone, right? We figured out who our users are that we're prioritizing for everybody given the value. We really care about number of exports in 7 to 30 days. So if we were to break this down, there's like a couple different ways we want to break down our northstar, right? Number one is probably by that user type. We already talked about how important it is to look at this for new users. Let's go ahead and make this black again. So we care about how this looks for new users versus power editors. And then another way to break this down. So that's like one vector of breakdown. So let me let me actually give that its own vector, right? So this is like user type.
23:43 Another way to break this down is like type of export. So it's like short form versus long form. We talked about that. Another way to break this down is literally the equation. So let me just think about the equation. And I usually like to think about that for just a second before choosing a Nordstar. So what's the equation going to be? The equation is going to be number of and we also want to look at how much they're editing the AI. So accept it with minimal edits. This needs to be like one of those key guard rails. So let me go ahead and make both these red and thicken them up so that we remember them. Okay. So the equation I'm just trying to think what is the equation?
24:21 Let's go back to the product and just look at the product together. So the equation is basically they hit this export button and then what happens after that right they choose it they publish okay so and then yeah okay so it's really number of publishes in that time frame that's the metric does that sound right it yeah it does perfect so we have our equation for our northstar we have it broken down in three different ways user type of export I feel like we've broken it down sufficiently are we good to move into trade-offs and guardrails go ahead awesome and then for this summary, I want to kind of talk about the whole document and what the executive write up would look like to Andrew Mason and Laura Burkhouser, our CEO and co-founder. So, let's first talk about these trade-offs guardrails. Okay, so the journey we've taken so far is that we looked at high editing, medium editing, low editing. So, we need to think about all three of those. Okay, so for trade-offs guardrails, we need to make sure it's not increasing the time it's edit it taking to edit and it they're they have minimal edits. So, okay, this is another important red one.
25:27 So, for each of these red ones, and then there's hallucination that it didn't edit. We need to make sure it didn't do that. So, I think let's operationalize all of these. So, I feel like it's like a less than 5% hallucination rate. No, probably less than 1% hallucination rate. Then we want to see less than 10% increase in time to edit. Although in reality I would have to look at the data and see like if people are using a bunch more tools maybe this isn't a strong guard rail. assuming no additional tool use per edit or broke down broken down by number of tools used per edit. So I'm thinking of a table essentially you could think of a table where we would say like okay one tool use our A group took 3 minutes our B group took 2 minutes 40 that's fine two tool use our A group took 4 minutes our B group took 5 minutes that would be a problem right because now they're using the same number of tools but it took longer so this would be negative and then this would be positive is kind of what I'm describing does that make sense awesome so increasing time it takes edit. We've operationalized that hallucination. We agree on that. Let's go ahead and give that.
26:43 >> Sorry, I think you're you're narrating mirrorboard while I'm still saying the the script. Yeah. Okay. >> Thank you so much. Increased time it takes to edit. What I had done is just drawn out this A and B case. So 10% less than 10% increase in time to edit. But I kind of mentioned that 10% is arbitrary. We'd have to really look at the data. But this is what I mean by broken down by number of tools. One, two tools. So you can see us kind of imagining three four and then hallucinating that it did an edit. So basically we're going to create an eval around did it actually do what it said it did and we would want less than 1% hallucination rate. Then support requests per user. So basically what we would want is that they don't go up, right?
27:28 >> less than 0% increase or we just that's kind of confusing language. No increase. and then accept it with minimal edits. So, essentially what we're saying here is that it's hard to operationalize this one, but it's like kind of like I was describing earlier with some of these these we're going to have to do it on a feature level. Studio sound with no edits, auto layout with no edits. And as we kind of talked about this, we're going to assume that the AI tool PM itself does it. So, really what we're talking about is coordination level edits. So, how do we operationalize coordination level edits? I think a couple things like this is like probably like people like rage interacting with the chat, right? So, saying I already told you that or you're dumb or you're stupid or something, right? We could kind of operationalize that as like various statements expressing rage. So, what we would want is like these rage interactions with the chat. It's hard to again I'd have to look at the data but I'm going to go ahead and say something like less than 20% of chats are rage interacting with the chat. We'd have to look at the data associate the data with retention. Associate the data with engagement. So we don't know exactly but that's the high level of what I mean by this guardrail. Does that make sense?
28:41 Yes. Yes. Of course. Just wondering if that wouldn't concern user that we are basically listening up on their conversations but at the same time we have to police our own responses so we don't do anything out of the ordinary by accident. So I guess >> we could definitely create a setting for users where they check or uncheck if their conversations are used for training data and if people uncheck then we can give them the option to report a particular conversation. Does that sound good?
29:17 >> It does. Thank you. >> Okay. So we can definitely handle that curveball here. All right. I would potentially like think more about trade-offs, but just looking at time since we're about 35 minutes in, I feel like we got a pretty holistic sense of the trade-offs. Do you agree? >> Yes. Awesome. So, now I think we're ready for the summary. So, if I were the PM for the this feature, I would create a dashboard for myself looking at things like number of exports is my northstar and I would communicate that to everyone. Hey, number of exports is my northstar. But I would also add into my dashboard many of these other many of these other metrics that we looked at.
29:53 So I would want all three of these metrics in my dashboard. >> These are like my positive metrics. I would also want first edit completion rate cuz that's really important. So this might be like my highle positive metric dashboard. And then I'd create this negative metric. So it's kind of like these eight metrics would kind of or these nine metrics I guess would be my major dashboard that I'm looking at in order to assess the success of Underlord. On the positive side, I'm looking at number of exports that they're editing more videos, their day 7, day 30, day 60 retention, their number of AI tools used, and their edit completion rate. And on the negative metrics side, I'm looking at increasing the time it takes to edit, hallucinating that they did an edit, support requests per user, and accepting it with minimal edits. And the final thing I would say is I kind of am in the HL Hussein Shrea Shunker camp of AI evals, which is you have to look at production and synthetic data to do traces to classify more negative scenarios and create evals around them. So I'm just going to go ahead and say like eval driven use neg.
31:00 Awesome. So, this is the dashboard I would look at if I were the PM of the Descript Underlord feature. Looks awesome. Though, I want to throw you a little curveball here because your processor is awesome. You've given me everything a PM would need to see to evaluate the success of the feature. And I'd like him now to like put this process aside. Put aside all like the technical constraints and like feasibility and give me like an alternative metric that you would do like a genie standard that you'd that if you could use which would you use to evaluate success from like a more of a user perspective and from a business perspective. For example, I used to work in Skype and my dream metric was to be able to count how often do people say, "Oh, you're on oh, you're on mute." to to like get this bad experience out of the way. So, so we would do automated tools to like mute you and unmute you automatically or mute unmute you when we detect that you want to speak. Of course, it could never happen. Privacy listening to users, no can do. But that was a dream success metric. What would be your dream success metric?
32:14 >> Yeah, I'm just brainstorming based on what you were saying. So like people basically saying that was awesome or that was great. I think it's kind of the analog of what you just described with Skype inapp survey all being five out of five, no support tickets or questions. really really good social chatter, you know, like people posting on X and YouTube and showing up like on Reddit in an editing forum. just connecting it back to the business goals of Descript, seeing more people sign up for Descript, right? Seeing more people upgrade to higher plans. You know, actually, now that we say that, I kind of missed that that sort of whole bucket of metrics.
32:56 So, let me actually go back there and just summarize those for a second. I probably should have included in these positive metrics some output metrics. And in the AB test, we probably want to look at these too, right? So, some output metrics might be things like upgrading plan, right? Renewing, referring more people. So, all the things that tie back to revenue. I think those actually should go into my dashboard. So, I'm really glad you asked that question. And going back to the genie metric. So I guess the ultimate genie metric is like a combination of that was awesome chats plus more revenue per user plus higher K factor more referrals. it would be like the synthetic metric putting those three together. I feel like that would be my genie metric. Does that sound right? Oh that's like a very personal. So thank you for sharing. I was just wondering how you'll react to that otherwise supportive interview.
33:56 >> Awesome. to to to like to that to that level of like doubt questioning and I love that you immediately reapproach your previous answers to like get the most of them given that there was a hint of maybe I want to hear something more but I want don't want to tell you like directly. >> Yep. Awesome. All right guys, that's the end of the mock interview. So I personally would give myself an 8 out of half out of 10 because Bart needed to remind me about those output metrics.
34:22 So, one of the things you want to do and maybe you should even build it into your framework is always remember I had an input metric here. Always keep this in mind. You can even like create little sticky notes. I like to do this sometimes before an interview is always remember input plus output metrics. And you might even want to think about like some sort of leading versus lagging. I didn't end up talking about that. these are really good. So I did talk about Northstar and guardrail which is like one good way to break down metrics but also consider input output leading lagging. Luckily I had a nice interviewer in BART and as you guys saw I and then went back and fixed my answer. And so that's the thing when somebody gives you an hint don't just assume okay my other answer was good. Go back and fix your other answer. And any interview that anybody gives even if I were the VP of product at the script I was Laura Burkhouser at Descript I could improve. So always take your answers, drop it into claude and chat GPT. Create a Claude skill trained on my articles around this so that you can get feedback and improve further. But this is the high level for a good success metrics interview. Are there any other things you would say I did particularly well that people should make sure to do?
35:34 Bart, as always, the visual structure of the answer is so helpful for the interviewer especially. This is a long answer that take took us what 40 minutes exactly and human being might get confused like miss a detail or like not keep all the narration and here even though I don't see all the screen all the time I do have more stimuli than only speech to get me to focus to remember to understand what is being presented to me. So that's like an if you are to forget everything from this video and remember like a functional advice rather than how to say remember to do all you can to get your answers understood be allow the interviewer to follow them easily and like get the best out of you what you're saying and I loved as I said I loved how Akash did pick up on my hand that Maybe there's something more without telling you because I don't know about you Akash but I had tons of interviewers that were smiling that were friendly that were doing patting me on the back all the time just to fail me a few days later during like the feedback session from the interviewer which always came as a shock. I mean this person who told me I did everything right actually thought that I failed. That's like one thing that you need to remember that they are not in a business of like being your chill reader. They are in the business of finding the best PM and they don't usually want to disturb you but they are also not there as an examiner that will push you to the right answer. They might do that, don't get me wrong, but somebody may be rude and come to you with something like, "Well, why are you asking me? It's your answer." like I've had those two. So, while this is an interview and there's a dialogue, remember that you own your answer and you have the full time to deliver your best to make sure you said all the right things and especially with such visual anchor, be ready to go back to a story point where you missed some important bullet points and add them. Without this visual context, it would be very very hard to reapproach certain aspects of the answer, but that makes it so easy.
38:05 The extra power move on top of this is once you create your visual anchor after the interview, drop this into chat GPT and Claude that's trained either a custom GPT trained on my articles or a Claude skill trained on my articles and say, "What could I have done better?" Add those into your visual anchor. Then send an email to your interviewer half an hour later. Say, "Hey, I really loved our case discussion. It was such a fascinating topic that I had to think about it more and here are a couple other things I thought about." So in this case, we just thought about input output leading lagging, right? That's something we needed to fix. So you could say, "Hey, I thought about the input output metrics, the leading lagging indicators, and here's what I was thinking, and I just created a little mockup of those metrics." Maybe you even go further if you want and you say, "Oh, I also AI prototyped a version of the dashboard just so we could see what the AB test final dashboard would look like." If you go that extra mile at a company like Descript, Descript is a $500 million company. If you go do that, they're you're going to be the only person who does that. Nobody else is going to do that. And so, you're immediately going to stand out. They're going to say, "This is somebody who really loves Descript, who is going above and beyond." So, that can also really help you. that won't help you as much at the Meta and Googles of the world because they're instructed not to look at that type of information, but pretty much everywhere else it will help you. So, if you guys want coaching like this, the type of practice with Bart and I, the type of coaching, feedback on your interview transcripts that's not just from an AI, but from people, then be sure to check out our landp.com cohorts. The next cohort starts the beginning of February, just a few weeks from now. We are going to be helping 30 people land their dream jobs. Are you going to be one of those 30 people? Let us know and we'll see you in the next video. See you there. Thank you.
39:48 Bye-bye. I hope you enjoyed that episode. If you could take a moment to double check that you have followed on Apple and Spotify podcasts, subscribed on YouTube, left a rating or review on Apple or Spotify, and commented on YouTube, all these things will help the algorithm distribute the show to more and more people. As we distribute the show to more people, we can grow the show, improve the quality of the content and the production to get you better insights to stay ahead in your career.
40:14 Finally, do check out my bundle at bundle.ac.com to get access to nine AI products for an entire year for free. This includes Dovetail, Mobin, Linear, Reforge, Build, Descript, and many other amazing tools that will help you as an AI product manager or builder succeed. I'll see you in the next episode.
Summary
- Two main buckets for PM case interviews: product sense/design and execution/success metrics.
- Importance of measuring success metrics for AI products, such as GPT 5.2 and Claude Code.
- Mock interview example for the "Underlord" feature, focusing on defining success metrics.
- Key metrics discussed include time to export, user engagement, and first edit completion rates.
- Emphasis on the need for a North Star metric and guardrails to ensure product quality.
- The interviewer encourages a structured approach, highlighting the importance of visual aids in communication.
- Recommendations for aspiring PMs include continuous improvement and follow-up with interviewers post-interview.
- The video promotes a coaching program aimed at helping PMs transition to AI product management roles.
Questions Answered
What are the key components of AI product management interviews?
The section introduces the two main areas of focus in product management case interviews: product sense and product execution, specifically tailored for AI. It emphasizes the importance of understanding AI product success metrics and outlines the value of the mock interview being presented.
How should we prioritize user groups for AI editing tools?
The discussion revolves around the importance of accommodating various user types, particularly those with high editing fluency. It highlights the need to balance metrics that cater to both novice and experienced users to ensure broad success.
What positive metrics should we consider for measuring success?
The section discusses various positive metrics, including time saved in editing, quality of the final product, and user engagement. It emphasizes the importance of operationalizing these metrics to gauge the effectiveness of the AI editing tools.
What trade-offs and guardrails should we consider for AI editing tools?
This section outlines the need to define trade-offs and guardrails to ensure that the AI tools do not increase editing time or introduce errors. It emphasizes the importance of maintaining a low hallucination rate and ensuring user efficiency.
How can we connect success metrics to business outcomes?
The discussion highlights the importance of linking user engagement metrics to business goals, such as increasing subscriptions and revenue. It suggests that output metrics should be included in the overall success dashboard.