Section Insights
The Role of Agents in Product Development
How are agents influencing product development?
Agents are integral to building features and validating requests, acting as an operating system that connects various tools and stakeholders.
- Agents facilitate communication between stakeholders and product teams.
- They enhance productivity by integrating with tools like Google Workspace and Jira.
- The future of product management involves smaller, more agile teams.
Measuring AI Output Quality
What metrics should be used to evaluate AI outputs?
Quality of AI outputs is measured by their impact on metrics, focusing on outcomes driven by AI rather than just procedural processes.
- Quality is more important than quantity in AI outputs.
- CPOs should focus on creating an operating system for AI collaboration.
- Effective processes enable teams to leverage AI for decision-making.
The Importance of Raw Data in AI
Why is it important to store raw transcripts?
Storing raw transcripts helps maintain fidelity and accuracy in AI outputs, avoiding loss of information through summarization.
- Raw data storage prevents fidelity loss in AI processing.
- Maintaining original transcripts aids in better recall and understanding.
- Experimentation with AI outputs can lead to improved methodologies.
Setting Up Metrics for AI Evaluation
How can metrics like recall be effectively measured?
Metrics can be set up by comparing responses from systems with and without specific skills, allowing for a clear evaluation of AI performance.
- Establishing a control group helps in measuring AI effectiveness.
- Natural language processing can assist in creating and measuring metrics.
- Visual inspection is necessary for initial evaluations before full automation.
AI in Daily Task Management
How can AI assist in managing emails and scheduling?
AI can digest important communications and schedule meetings intelligently, reducing the need for manual checks of emails and calendars.
- AI can prioritize and summarize communications for efficiency.
- It can autonomously schedule meetings based on availability.
- Using AI for responses can streamline communication tasks.
Transcript
0:00 Agents are actually building most of our features. Any stakeholder can talk to an agent and validate any feature request they want. Our agent is an operating system because it's hooked into everything from Google Workspace to Granola, Confluence, Jira, it operates at all. It can fetch all of the transcripts from whatever meeting that happened over whichever period of time. Meet Mikuel Shaglov, the CPO at Ox Classifies, formerly CPO at Azerbaijan's biggest e-commerce marketplace and GPM at Bolt.
0:30 >> We've built everything, our entire stack on top of Open Claw and Hermas agentic scaffolding. I use this knowledge graph to understand how well teams are doing with customer discovery and how well they are connected to the entire organization and the stakeholders. What's going to happen to my job? What actually will happen is that teams will be smaller, leaner, and faster. Not only it's going to change, the lines are going to blur. >> What's going to happen to the future is product management going to exist in the future?
1:00 >> I have good news for you. I think >> before we get into today's show, please take a second to check that you're subscribed on YouTube and following on Apple and Spotify podcasts. If you want access to all of my favorite AI tools, I've gotten them to give you an entire year of their paid plans. Check out bundle.ac.com for an entire year of bolt new, air table, speechify, descript, magic patterns, linear, dovetail, arise, and mobin. And now into today's show.
1:36 I know you guys are overwhelmed with guides on Hermes and OpenClaw and Claude Code and ChatgPT. It seems like there are million AI tools out there and more and more of you are reporting to me that you're not even feeling more productive after using AI. So, I have been searching for those PMs and product leaders who are not just getting like a 10% boost in productivity, but are getting a two 3x boost in productivity. And Mika is one of those people. What he has built at Ox Classifies is an entire operating system. If you've been hearing about buzzwords like OpenClaw and Hermes and saying there's no value for PMs, this episode is the one that will change your mind. He has got his entire team hooked into this operating system that he has built and they are validating features, shipping features, reviewing features. The possibilities for a PM are endless and I can't wait for him to share this all with you. Nikuel, welcome to the podcast.
2:34 >> Glad to be here. >> So, how do people really build a company operating system? How do they measure it? What are the keys to getting the most out of AI as a company? >> All right. So if you think about this like what's the core goal of having AI native teams is to be able to automate as much knowledge as we can because like in previous I call this old paradigm like you used to have a knowledge worker like an employee who has built a context around a certain domain let's say it's it's a deep expertise domain like search recommendations ML and stuff like that and then all of a sudden this person decides to leave the company and takes all the knowledge with him or with her.
3:16 he is a bottleneck effectively a knowledge bottleneck. And the core theme that we're trying to prevent here is the leakage of that knowledge. Because if you could have a single storage of this entire business, customer product technical knowledge in one place then it would increase like the entire value of your organization. That's just one angle to it. And another angle is the better your AI knows your context, the more autonomy you can give it, the higher level tasks you can delegate to it. And let me explain how it actually works in our case. So what you see here is the entire knowledge graph of our company that has been formed over like five months like five months is basically nothing. And you can already see this like web of interconnected particles. And these particles, they cover everything from our contacts, who talks to whom to which projects we're working on to which customers we've talking with, which businesses we've reached out, even our funnel metrics, basically everything. And you can see in the top right corner like there is a specific metric which is a product context coverage. And this is one of my personal KPIs because the higher this percentage of coverage, the more as I mentioned you can delegate to it because right now what what what this means is that it understands our industry like classified our market our business our business model specifically the top line the bottom line the drivers behind it our customers and the value drivers for time at the level of 54%.
5:12 Which already means that it can operate as a capable I would say junior to meet product manager and even make backlo decisions. the more this coverage grows like at numbers of 70 to maybe like 90%. I would not be surprised if this knowledge graph would enable you to do some strategy level work and help business navigate decisions. It's a question of when we're going to get there, but we'll definitely going to get there. Let me maybe zoom in a little bit on on the elements of this knowledge graph. So, you have three layers. So, the first layer is the product. So all the nodes that are connected to the product side of things and you can see like product has been very involved into everything. you have the contacts which is basically all the people talking to each other in different parts of the organization. and there are like personal things that I share like my thoughts my reflections and the personal things that some of the other people share like of course fully anonymized.
6:21 Then another layer is the teams right I I don't really go at a level of actual like say cross functional teams it's more of a clusters or divisions right so if you take buyer and seller experience you can immediately see how interconnected they are within the organization or like the platform teams or like the paying ship teams and you can even drill down at the level of a specific product manager and you can see on a team level what's their knowledge of their existing context which for me is a very strong predictor of how good the team is in terms of customer discovery.
7:01 not only that there is another angle to it. If the team isn't really well connected and you can clearly see that this team has like like very little nodes overlapping nodes with the buyer and seller experience even though a platform is a horizontal layer and they should be ideally they should be intertwined like with each other. you can clearly see that there is some there some instances of silos happening which means that teams should talk more with each other and two I would also raise the question of the stakeholder management like are stakeholders really being involved in the processes within the teams and so on and that's kind of my core tool that I'm using because it it shows me the quality of an entire of the output of the entire organization.
7:58 And >> this is like the coolest visualization I've ever seen. So what tool is this? How are these metrics working? How are they measuring things? >> so you can use like tools like Obsidian for this. I personally created this one from scratch using Fable because it was like a two-lininer prompt basically. the important question that you asked is like super important one which is how do we measure this? and we have a very specific prompt that asks an AI what percentage of knowledge around industry and you have to be very specific. What do you imply by industry?
8:43 In our case we have five verticals. We have real estate, we have auto, we have good sales, we have services and jobs. what percentage of knowledge around the business meaning the business model the actual P&L the drivers and what percentage around customers like different buyer seller customer segments even on the cohort level marketing and so on. which percentage of knowledge do you have at the current point in time taking all of your memories loaded into you directly and what we've seen is that AI has been quite accurate in digitizing this abstract request into a specific number as an output and what I've also what I'm also observing is that the more active product managers are and they typically come with I would say transcript scripts or research studies or like RFD documents and they load it directly into our agent into its memory and I'm seeing that the metric is actually improving over time. It doesn't mean that it's a single source of truth and I think each organization should come up with a specific context metric that would be highly specific and relevant to them.
9:59 but this is something that I found directionally works. >> Here's a quick word from our sponsors. I need to say something that most AI coding companies don't want to hear. Their tools are not built for enterprise. Think about what happens when 20 people in your org start building with AI. Marketing makes a dashboard. Product prototypes a feature. Sales builds a demo tool. Each one looks different. Different components, different styles, no shared design system. And none of the code is pullable. Engineering can't integrate any of it. You've got 20 AI powered prototypes and zero production grade output. Bolt. And you solve this with something called design system agents. You upload your actual design system, your npm packages, your CSS files, your component library. The DS agent builds a custom system from your real code. Now every person building in boldnew across your entire org is using your components, your buttons, your tokens, your brand. And when engineering pulls the code, it matches what they already ship with. That's the difference between an AI tool and an AI platform.
10:57 Check it out at bolt.new/ aos. I used to think I had a retention problem. Turns out I had a messaging problem. I was sending the same onboarding emails to every new user, whether they activated on day one or never logged in again. I had no idea who was slipping or why. Customer.io changed that. Every message I send is now based on what users actually do in the product. Someone hits a key activation moment, they get nudged to the next one. Someone goes quiet, they get a different path entirely. Their AI agent makes it fast. I describe the campaign I want and it builds the full journey form.
11:27 triggers, timing, copy, even branching logic. And when I want to know how something is performing, I just ask the agent directly and it tells me what to do next. They also have an MCP server, which means AI tools like Claude can see directly what's happening in your customer.io workspace. Your segments, your customer data, your attribution, all of it. So instead of explaining your business context every time you need help, Claude already knows it. Notion used customer.io to personalize their onboarding and hit nearly 50% open rate. improved conversion by 6 to 7% with localized campaigns and pushed open rates up another 20% through AB testing.
12:04 The idea is simple. Customer.io helps you deliver more impact from every message you send. If you're a PMR founder and your onboarding is still oneizefits-all, try customer.io at customer.io for us. >> Okay, so this is a totally different way to view a product team. Most product teams, most PMs aren't used to this to if they want to build this. It sounds like there's two things they do. They define the metric and then they create the visualization with fable talk. Let's take a step back here. What is the role of a CPO in an organization like this?
12:39 What is the new agentic CPO doing here? >> That's that's a great question. I think like what has changed significantly is that in the past like the whole team was responsible for the outcomes and outputs right outcomes were measured in how much your metric has moved and outputs how much RFD documents how much code how much prototypes Figma mockups like your team has been able to produce depending on the function. right now what is what has changed and changed pretty significantly is that the cost of this output is like negligible.
13:21 but what becomes more important is the quality and the token consumption. So what do I mean by quality? quality for me is the amount of AI outputs that have significantly moved or moved the metric. So it's it's essentially outcomes but outcomes that were driven by AI outputs and this is how I measure it in terms of the metrics in terms of the actual tooling. I think previously what u CPO was responsible for is building a scaff you know procedural scaffolding on top of the organization.
14:01 What it means is that you have certain like hiring principles, right? your vision. Then if you go a level lower, you have certain processes like planning cadence like a year yearly strategy, then a quarterly strategy, then a sprint planning and then you have like different reviews like design reviews, product reviews and so on. And this allowed the organization to function. But right now what is more important it's not the process itself. It's creating an operating system within which AI could be a collaborator of product managers and where AI could not only help to kind of come up or challenge ideas but it could autonomously make decisions. and the best way to achieve this is that is by owning the entire agentic scaffolding architecture and creating specific rituals creating specific processes and enabling the team to be in touch with those agents all the time training the team to use it actually. let me know if you'd like to to go deeper into any aspects of this. How does this manifest in the PM process? Where are the savings had for teams? Why should a CPO suddenly be spending so much time as an agentic orchestrator? The the biggest reason for this like if you break if you break down the time that product managers on average used to spend probably not not probably actually I estimated this based on my previous experiences around 50% of the time of a typical product manager was sp was spent on processes and rituals.
15:49 you have different business reports, weekly reports, stakeholder management reports, different demos and so on. And this is a repetitive type of task that doesn't really require much cognitive input in many cases. It's just something that you have to do manually. and if we abstract all of this and start delegating this part of work to AI, then essentially what you end up with is you have a product manager that is solely focused on discovery like the where the biggest leverage actually is.
16:35 and a single product manager could actually do the work of two product managers because he he doesn't have the burden of of all the processes that the typical corporate environment imposes on him. >> Wow. Okay. So, we've seen the knowledge graph. We understand the role of the CPO and what it does for a PM. Can you show us what this agent looks like in practice, how a stakeholder might interact with it? There are multiple buckets of tools that this agent can solve and we constantly improve it. So the first bucket is status reporting like you can ask it like what's the status of a specific project.
17:17 Let's say what's the status of auto spare parts catalog integration respond in English because he might get finicky. Okay. And while we're waiting for this there are multiple other things. So status can be you can get a response in whichever shape or form you want directly in Slack in Google doc in Confluence whichever format is more convenient for you. Second thing is it is integrated into all of the workspaces right Google workspace meaning calendar gmail and so on. I don't read my email anymore. An agent does it for me and pings me in case if it's if it's urgent. I don't manage my calendar anymore and it's it's especially difficult when when we live in this interconnected world where everybody is remote in different parts of the globe and just setting a single meeting. Oh, you can clearly see it's all there. Scope listings automated. Yeah, you can clearly see like what's the status. it's quite straightforward. Nice. And did you like have you been tweaking the prompts on how it gives status updates because this is like a pretty good one.
18:35 >> Yeah. Yeah. Yeah. So, it has a massive library of rules and imperatives. I think like 700 lines or so to give you the cleanest and the most factual output without hallucinations. >> Very cool. So, we'll go into how people build that. What other use cases should they know about within Slack that they can accomplish here? like one angle to this let's say you're a stakeholder right and you have a certain feature that you just assume is important to have by what whatever reason like a typical route for a stakeholder in this case is to go directly to a a product manager and harass a product manager with feature requests right but now like we have a gatekeeper which is our agent running under the hood and we train our stakeholders to interact with an agent before reaching out to a product manager.
19:31 >> Oh, >> let's say, yeah, I just woke up and I I have a feature request. I urgently want to have a video submission for outs story like an Instagram because it's fancy cuz I saw somebody else doing it right.
20:04 That's always the reason >> all competitors are doing it. We are behind. yeah. for for sellers. Yes. and this is just like a typical feature request which doesn't have a problem framing. It doesn't have any impact estimates. it doesn't have any rational behind and it doesn't respect the actual priorities that have already been defined for for the organization.
20:35 And but I we'll just give it a little bit of time to think. it will respond with a list of clarifying questions. And in case if after answering all the clarifying questions an agent decides that this feature isn't worth building, then it will politely say so. if it is worth building then this feature will be added to the backlog and escalated to the responsible product manager. an agent has all the organizational structure ma mapped out internally so he knows exactly who the owner of a specific domain is.
21:12 >> Wow. And do you consider yourself the owner of this Saul Goodman agent? A lot of people they want to outsource that to like an AI ops role or something like that. I I think it's a great point and it's a dangerous way to go. I think you should own this for two reasons. Like one reason is the speed of iteration. Like if you own this and you're constantly improving and I'm getting feedback every single day, what works, what doesn't work, and so on. And I immediately like open my IDE or my cloud code and I make the changes right off the bat as we go. It's super fast from feedback to deployment. secondly this actually has an impact on the organization because one it saves time and two it stirs decision- making which means it might have an impact on the entire business and I think CPO should own this. If you delegate this to an engineering team or if you delegate it to it to another person who who doesn't have the skin in the game then you're not going to get the same level of speed and impact.
22:19 >> Okay. So CPOS should be building this now. They want to know how. Can we open up the covers and show people how they build something like this? >> Absolutely. Absolutely. Let me open up my ID. Actually, I will show you two ways. yep. This is just to give an overview. I will I will kind of walk you through the architecture, the brain, the memory, the tools, the skills, and I'll explain it all as we go. so for the architecture here what we're using is a combination of openclaw and hermes.
22:58 why those two? Because open claw is just a great scaffolding. It has everything available out of the box and great engineering support. but Hermes has a unique feature to it which is an automated skill generation and I've actually tested it across five core topics that my team works and I've seen an improvement of in in a recall metric which means that it was way more accurate in responding using those skills by 31%. Which was a dramatic improvement. So, I decided to blend both of those scaffoldings together, in a combination. and it's it's jacked in terms of the capabilities because we've been improving it for like 5 months continuously. when it comes to memory, it has three layers of memory. So, the first layer is what you've seen is a knowledge graph.
23:55 it's like an entire universe of of interconnected knowledge and everything. And you can see the code here and so on. but what is plugged in there? It's a second layer which is a vector database. And the vector database is essentially every single piece of knowledge that an agent receives gets immediately converted into vector. And you need this because most of the requests are fuzzy. and in order for you to get good retrieval, you you do need to have vectors because you might ask some random stuff, right? that that is not related to any specific keyword and an agent needs to be able to match your request to a specific point to a specific number in your knowledge graph. And the third layer and what I've seen makes agents the most robust and they don't lose context is every single conversation is stored in transcripts.
24:55 You can clearly see that every single day the agent stores everything into MD files. And not only this encompasses granola transcripts because all of our meetings are transcribed, but also all of the conversations that I had with an agent and all of the reflection that an agent has. So those three layers of memories, persistent memories, they create this robust knowledge base which improves every single day. >> Wow. How does that exactly how do you architect that properly? So obviously you get a granola enterprise license.
25:29 You have that coming in and >> recording every single meeting but I guess you don't want to just save all the meeting details. You want to strip out some of the personal conversation out of meetings and then you want to write these more condensed synthesized MD files. How do you go from raw transcripts that might have personal information to useful synthesized output? That that's a great point as well and that was my assumption that you actually have to summarize and synthesize the output to be useful but we tested this in in in multiple ways and it turns out that summarization actually hurts a retrieval >> yeah for for two reasons like one is when you summarize something you lose granular details and the devil is always in the nuance and the details.
26:20 And another reason is that when you summarize, you impose a certain template onto whichever conversation that you had. For instance, like the tasks that were solved, the tasks that remain, what is important, what is not important, and then on top of this on top of this template, you you try to kind of shove your your transcript into this template. And we we've noticed like a huge fidelity loss. And I think it was like 20 25% worse recall.
26:52 That's why my personal architectural decision here was let's just store every single transcript, every single conversation that we had in a raw form without any summarization because it cost. >> Quick thought experiment for you. Is there anything in this video you should be trying on your own? If there is, try it. Take a screenshot, post it on LinkedIn X, and tag me. I'd love to see what you're learning. Now, a quick word from our sponsors before we get into the back half of the pod. Do you know how to take an AI product from idea to development to evaluation to deployment and eventually to scale?
27:28 That's exactly what product faculty's AIPM certification helps you do. I even took the course myself. You'll learn directly from Rohan Varma, the product lead working on codecs at OpenAI. You'll go deep into AI prototyping, evaluations, agents, AI native workflows, cloud code, openclaw, latency, cost, guard rails, rag, routing, fine-tuning, and production systems. You'll even build your own AI product as your capstone with unlimited one-on-one support. So, if you want to stop just learning AI and actually build AI products that work, join product faculty's AIPM certification on Maven.
28:05 5,000 plus students have graduated and they have 1,000 plus reviews. Use code Akash 550 to get $550 off your enrollment. I used to live in report purgatory. Every team had a different number. Every weekly review started with someone reconciling spreadsheets. We stopped hiring more analysts and gave the reconciliation to an AI employee instead. Victor is an AI employee that lives on Slack and Microsoft Teams. It connects to 3,000 plus tools your team already uses. ships real deliverables and every action goes through your team for approval first. It's the closest thing I've seen to a small team running like a much larger one. Let me share three things that Victor does that changed how my team operates. First, ask Victor has replaced our Monday metric scramble. Someone types flag any customer whose usage dropped 40% week over week and draft an outreach loom for the account owner to approve. 90 seconds later, the next step is ready. Second scheduled tasks run the work that nobody wants to remember. Every morning, Victor checks overnight support tickets, drafts replies for the on call to approve, and escalates anything mentioning churn.
29:14 Nobody had to ask for it. Finally, spaces ship internal tools in minutes. Ask for renewals dashboard. Victor builds it with a database and post the link to your channel. The team can stop opening four tabs to get the same view. So, stop chatting with AI and start working with it. Get started at victor.com. There's $100 in free credits with no card required at vikor.com/kashgupta3. You can find that link in the description. I want to take a second to talk to you about the fourth cohort of LAN PM job. I trained 30 students in cohort 1, 50 students in cohort 2, and 75 students in cohort 3. And I am bringing back the program for cohort 4.
29:54 It starts in August and it lasts three months where you're going to have intense sessions. a Monday morning session where I go over your resume, behavioral interviews, LinkedIn. On top of that, Bart Choworkski is going to be teaching you the PM fundamentals in 2026, how to write AI PRDs, how to AI prototype with cloud code, all of the key skills you need to freshen up your knowledge for this market. And Ang Vermani is going to be teaching you AI product management. He is an AI product manager at Uber and he is going to teach you how to build AI features that actually work successfully. On top of that, Prasad Ready is going to be doing one-on- ones with you for mock interviews, LinkedIn review, candidate market fit review. So, it is a full package. It is three courses in one for one low fee. So, join at landpob.com.
30:40 >> Nothing. >> Wow. And Hermes and OpenClaw on their own can figure out how to get into which meeting and which meetings are relevant context and they won't just fill up their context window with random meetings. >> Yeah, absolutely. because like what happens is that when you ask a request, what it does it it it it does a search query. and it it it's a blend. It's a hybrid search query. it tries to do exact keyword matching. if it's not successful at this, which I think 75% it's not because it's hard. it's it's a very ambiguous raw context.
31:16 then it does the vector search and vector search allows you to to do this exactly fuzzy retrieval based on probabilities and that pretty much handles like the remainder 75% of use cases. So it only rece retrieves the relevant bit piece of data that is highly specific to your request and it doesn't really overload the prompt with with tokens there. >> Very cool. So that's the memory component. What else do people need to know to build this? I think the the thing that after after deciding on the architectural aspects of this, it is hugely important to create the list of imperatives and as you know LLMs are very biased because of how they are trained and more specifically they are focused on the resemblance of good output rather than the actual results.
32:17 >> >> you know and you can you can do workarounds to solve this and the best way is to have this list of imperatives. For instance, what what what for me is super critical is that there's no fabrications, right? there is if you look here u it's it's just a huge list of different imperatives that we've created over time. It's who you are. It's your voice. it's think before you act because that's that's a terrible thing I've noticed LLM do a lot. they provide you the output that looks plausible but there wasn't that much thinking behind it because it's full of contradictions.
33:02 then there's always facts over guesswork and antiatterns. Another angle which which super pissed me off a lot. I call this fake helpful. For instance, if you ask your agent to book you a meeting and suddenly your tokens in the Google workspace has expired and then the agent says, "Well, I'm sorry I cannot do this, but it it's super easy for you to do. Just open up a calendar, type in Google calendar, name your meeting, choose your time, choose. I mean, it's useless. I mean, it's it's an obvious advice. you don't even have to waste tokens explaining me this and I call this fake helpful and if you are able to put it as an imperative it will save you a lot of time as we go >> but you can clearly see we've we've made a lot of tweaks over time >> so for people who don't know we showed claude MD and soulm what's the difference between those files what should be and what soulm I believe is a part of open claw >> yeah it's a part of open claw it gives effectively like it's a specific context that is being loaded into every single prompt. and claude MD has just the highest priority of them all and soulm is the second in order of priority for openclaw.
34:18 >> So your cloud MD it seemed like you kept that under 100 lines which I think is the advice that Boris Jurnney creator of cloud code gave. >> The stole MD though is like 800 lines. So that can be longer. Yeah, it's it's longer and and you might argue that we are overloading the context and some of those lines might even be contradictory. But what I've noticed was that that's the most robust way that gives me the most accurate output with the best recall. And we we test every single imperative. We actually test on on the basic queries to see across the most important topics that we raise in the organization. like h how how good it is or how bad it is.
34:59 >> All right. What else do people need to know to set this up? >> So another angle is you need to have tools, right? And it has a lot of tools. so since it's interconnected with all of your workspace it is deeply integrated into Google. So it's a separate tool. It's integrated into Atlassian. It's a separate tool. It has automations and automation like the one that we currently have. It just consumes all of the mobile reviews. it writes a digest like a typical supportability team would do and it even tags the people that are responsible and it it can even raise red flags in case if necessary.
35:37 >> And these are all Python files. You just prompt these with natural language in the IDE to Claude and Claude writes these or how does that work? >> I previously I used ID but now I'm I'm doing this for showcase purposes mostly. I'm I'm using cloud app >> because it allows me to do this on the go right here. It it is accessible from my from my mobile phone in case if an urgent thing appears and I want to change anything within the agent.
36:07 >> So do you basically create a GitHub repo that claude code can access? >> Yes, exactly. So every single change is immediately committed into GitHub repo. >> Awesome. So you do everything via GitHub via cloud app. That's so powerful. So you can configure people kind of have this mistaken thing that they need to use the open claw gateway but you can actually do everything through claude and it can manage it on top of Hermes and open claw.
36:33 >> Absolutely. So the only benefit ID gives you is that you can clearly see the the structure of the project. >> Yes. And you went through two idees actually. So you showed us anti-gravity and cursor. Talk to us about when you should be using the claw app versus anti-gravity versus cursor. My use case are the following. So everything that lives in the cloud and doesn't require I would say access to my computer I'm using I'm interacting with it through cloud app. but any specific use cases that I have that require my computer use and mostly these are around browser browser use via different MCPS or Excel use if I want to load a file and get some immediate feedback or data research if I don't really want to load it into agent and I want to do something ad hoc I use an ID for this and why do I have two IDs simply because it allows me to have two autonomous separate instances of agents running at the same time.
37:41 >> Okay, I think I understand the open claw setup part of this. It's mainly through the solemn MD plus open claw is what's giving you that gateway to the Saul Goodman agent in Slack. Explain the Hermes part of this. You had said it's related to the skills and the recall. So the beauty of Hermes as a scaffolding is that it comes with a unique advantage over open claw which is it automatically generates skills based on the tasks that you most frequently request the system to perform and what what we've noticed in testing was that the recall has improved significantly well I mentioned 30% with those automatically generated tasks or skills sorry versus without them.
38:29 So we we have it running under the hood all the time like for instance team hiring evaluation immigration like case building because we have like different people who are relocating locally. So I have a lot of questions around it. Then it it made a decision that we should have a skill for this. >> This the same team hiring evaluation. It also made a decision why do why do I continue doing repetitive tasks? Why not just create a skill and offload all of this onto an agent? So it even kind of takes this meta part of the understanding whether or not you need a skill and makes a decision for you.
39:05 That's that's an amazing part of it. >> Awesome. And how did it help with the recall? You mentioned it also helped there >> plus 31%. So what does it actually mean? so we have five core topics. And those topics are market its business model its product key growth levers selection price buyer activation trust then we have more tactical things like funnel and so on and we have a set of different questions that we ask and we ask like I think like 10 questions for each of those areas. And we compared the non-skll response versus the skill response. And we also looked at how accurate the response in each of those categories for each of those questions was. And what we've seen was that having those skills in place gave us plus 31% more accuracy, which was which was a deal breaker.
40:12 >> How did someone set up like the metrics to measure something like recall? you can do it basically yourself. I'm I'm doing this as a CPO because I think outside of knowledge graph and the percentage of context digitization so so to say you should also be able to track more tactical metrics which is recall and you can ask Claude or whichever system you're working with or your agent directly tell me which areas are the most frequently asked or which I fre most frequently ly interact with and come up with 10 questions per each of those areas which are the most frequently asked and then let's do a comparison basically like one control group versus the treatment group with the skills and see what's the difference there. And that's that's pretty much it >> to use sort of a derogatory word but not in a derogatory way. You're vibing the metric. You're in natural language.
41:20 You're explaining exactly how you want it to work and then it's creating the metric and it can go off and measure it for you. >> That is correct. But I mean for the first pass you have to like at least visually look at the results. >> Mhm. >> but once you have a methodology set up then you can delegate the evals fully to the system. >> All right. Amazing. So now people know how to set this all up. Can you show us some more use cases? What are the best most powerful things people should be using this for?
41:49 >> one cool thing is this is this is also around skills but I think this would be highly practical for people especially in sea level positions. So I have a board of directors like that I'm I'm in constant touch with and they make investment decisions right how much money we as a company should receive at which point in time what would be the payback and so on and so forth and there are multiple people in the board and they have different perspectives but the beauty is the agent was able to abstract their mental models into a set of principles and we called it this the board skill. So in case if you have a pitch deck and you you want to defend the strategy for the next year, then the first thing I do is I run this pitch deck through the board skill to to kind of poke holes and give me brutal feedback on what could go wrong with the defense of this strategy.
42:54 >> Fascinating. So your granola meeting transcript is automatically recording your board meetings. Hermes has automatically created a skill around the profile of those board members. And so when you're creating a board deck and Claude, let's say you use Claude design and you would put in your input >> towards the end of that process, you're going to hit it with a prompt like use the board skill to see how the board would respond so that we can refine this. Is that right? That is exactly right.
43:24 >> Fascinating. That brings up an interesting point for me. We as CPOS, we may not want to let our PMs see what's going on in a board meeting. How do we make sure that information is locked down so that the right groups get access to the right information? >> No, that's a that's a great question. So, and two things. So, the first one, it all boils down to scope of ownership. and the agent has different access rights depending on the person reaching out and that access rights will define the context which this person or employee is able to retrieve and the tools which this a which this person would be able to work with. For instance, board skill is available only to me and to to the executive committee within the company. and the second thing is the privacy aspect right because like not all people really appreciate that their personal meetings are transcribed or recorded. So unless you willingly provide this information to Granola we will not train the model we will not create like the skills on top of this. Although every every employee has a say in this, of course, I'm I'm personally a part of the experiment. That's why I'm completely transparent and I don't care about my privacy at all.
44:49 >> Okay. So, if you slip in that you you know, we're taking care of a sick kid during the weekend, you're okay with that hitting a transcript. But if a particular employee doesn't want it, they can elect to take it out. >> Absolutely. Absolutely. Yes. and from the get-go, we don't really we don't really store the transcripts, the personal transcripts of people's meetings because I think that violates privacy. So, unless you willingly provide us this information, we will not do this.
45:16 >> So, you can select like this is a product trio meeting. This obviously makes it in. This is a product review, but this is just a one-on-one between me and my designer. This is not going to make it in. >> Exactly. Exactly. >> Got it. So Bardex is one really cool use case. What else are you what else should people be using this for? >> Let's see. it responded in a different language. But the thing is so we asked it I want a specific feature and instead he created u a prototype.
45:45 >> Whoa. >> Let's see. Yeah, >> that's quite aic and autonomous. >> It's a gentic and autonomous. and it it sometimes gets finicky depending on which language we interact with and we use different languages like so that's why it might get frustrated at times and it requires an imperative but but here I think it created a prototype which is very close to our actual design system. >> Okay. which is I'm super I mean it's not great in terms of UX, but as a as a oneshot pass I think that's that's an okay okay thing >> as a one shot. I mean it's just so much stuff it's done.
46:28 >> Yeah, pretty much so. And it's very compliant with the actual Nexus design system which is an OX design system at play. >> And do we know what underlying model it hit? Did it hit Fable or does do you have intelligent model routing? How does that work? Well, it it so the thing is there is an agentic orchestrator which makes a decision on which model to use. >> So if it's u if it's a complex complex request or if it's a specific domain area which has high sensitivity or error blast radius most likely it will assign fable onto it. if it's if it's just an execution type of work or status report writing I think it it will assign ous 4.8 date most likely for lowlevel work that doesn't require super accuracy. It will a science on it.
47:19 So token optimization becomes becomes important here. >> So mostly it builds on the enthropic stack. You're not throwing in a codeex or a GLM yet. >> So it's it's also a great point. our engineers actually do u because like if we look at the like the typical triad. So engineers have their own agentic stack, product has their own agentic stack. Of course they they communicate with each other via the same rag and MCPs. but when it comes to product management we use cloud code by subscription pretty much so and we don't really have that much of token spend at the moment. That's why it's not a critical metric for us to track. While for engineers it is critical and they actually have a totally different scaffolding and they have a blend of different models running under the hood for the breakdown of tasks which is typically considered the work of highest complexity. How do you break down an abstract epic into tasks they use the highest cost models like chat GPT 5.66 X six or Fable, but for lowest complexity items, they use like smaller models or sometimes even locally deploy deployed open source models.
48:35 >> Okay. Well, how does this help you with things like your design system? >> Pretty significantly. Let me show you Figma. So, if you look here, right, the entire design system was built from from prompt. >> Wow. and it's very well documented and the components are robust and what is more important was that the components are exhaustive. So for instance you have all of the sizes of the components you have all of the states all of the relevant tokens and it is continuously supported by the agent himself. So in case if during the conversation between engineers and a product manager there is a disconnect right then because of because a certain component is missing or somebody tries to hardcode a component for instance then an agent is well is informed about this and makes a decision to create this component here in a system.
49:36 >> So it's automatically updating the design system based on the work different design teams and product teams are doing. >> Yeah that is correct. And how do you make it auto updating like that? so one is it has a certain tool on on the product side that kind of a collects all of the requests that are coming either directly via Slack or from a coding agent and then it just fires up a development crown job to start building those components on time. Of course there is a review process where a designer takes over and looks like through the components whether they are correct or not like how how well they are documented are there any conflicts are the states exhaustive and so on but that's a review type of a job rather than execution job >> and is the review is the design system skill in Hermes auto improving like is the amount of review that the designer needs to do getting less over time >> it's a great question I I'll have an answer by our next podcast.
50:49 >> Amazing. So that's design system. The next use case you talked to me about which I think it's fascinating and I really want to see is around backlog and roadmap management. What do you use there? >> So you know the biggest problem with backlogs is that every single team has their own textbook. Every single team has their own Excel spreadsheet. there are many apps or tools that try to automate this but I haven't seen a single success successful case.
51:15 that's why my understanding of a backlog is more of an abstraction on top of all of the Excel spreadsheets that team are working with. It's basically every single team has their own spreadsheet. I'm not going to like open them right now because of the NDA things, but it it's the typical typical type of a backlog like you have a you have a project, you have a problem, you have a solution, different links, impact, effort, ROI and so on and so forth. Every single team has their own backlog and they're all interconnected. They're all connected into agent and an agent can fetch all of this ination and can contribute directly to the backlog in case if a feature is definitely worth it and passes the ROI bar and passes the revenue of a responsible product manager.
52:03 >> Fascinating. What about personal assistant? How can you use this as a personal assistant? >> that's my favorite part of a of a flow. Let's maybe I need to plug in the meeting with execs for one and a half hour next week. Not too early, not too late. Suggest something. that's how you would typically talk to your personal assistant. Mhm.
52:40 so right now it does all the retrieval processes that we've talked with that we talked about and the ideal response is he will come back with a list of different time slots that would work for all of the execs who who all have busy calendarers who operate in different time zones all across the globe. and would and I would make a decision on which one to choose. If I'm too lazy, I can even delegate the decision making onto an agent.
53:08 >> So, while that's running, you mentioned you aren't even checking your email anymore. You weren't checking your calendar anymore. So, how is it managing those for you? How is it surfacing the right emails, the right calendar events, getting that prep going? >> So, every single day I get to digest or what's important what what have slipped in terms of Slack messages or emails. I browse through it or you can clearly see Wednesday, July 22nd. Yeah. Cleanest slot, chief. That's Let That's Let That's Let That's Let That's Let That's Let's use it.
53:39 >> Wow. So, it's applying intelligence. It's figuring out when those execs are available. It's looking across time zones. And you can even schedule the meeting directly. >> Yeah. I don't need to open the calendar. So, I'm I'm getting the digest. I review the digest and I make a decision on which emails are actually worth acting on versus the ones that are not or which tasks or or messages I need to respond urgently versus the ones that are not.
54:04 And if I'm not feeling like responding directly, I will ask the agent to come up with a template response. >> Amazing. I know the last area can be a bit of a headache for product leaders. On top of our own job, we need to keep the pipeline of really good PMs. How can this help with recruiting? Oh, actually quite quite a bit. so it has three tools integrated. So the first tool is LinkedIn recruiter. let me try and see whether he will be able to respond.
54:37 The number of product manager candidates we have in the by line just number. Okay. LinkedIn recruiter is responsible for reachouts. So if you have a qu question, let's say I don't know, I want a specific candidate with a very specific domain expertise with this much with this many years of experience located in this specific place just go and find me.
55:11 and if if I'm if if the candidate is found, then I will ask to do reachouts using like predefined templates that we've agreed on with with an agent before. The second thing is it is integrated into our CRM. different companies use different CRM. I think Greenhouse is probably one of the most popular and I don't really interact with the CRM anymore because I do all of the work via the agent. I just type in like like what who is in the pipeline or I provide the feedback directly to an agent. an agent posts this feedback or moves the candidate along the funnel.
55:55 or if needed, I can even build build like the funnel statistics to see you where where the gaps are in terms of the hiring. And the third aspect to it which is quite amazing which is the interview process, right? Because all of the interviews are transcribed. you can get immediate third person opinion on whether or not the candidate really lived up to your expectations. and in case if not of course you are the final decision maker and not the AI it just gives you an alternative point of view which is sobering in many cases.
56:34 then you can ask EI or you can ask an agent to to draft a rejection letter and reach out to the candidate. So this automates probably like 70 75% of the of the recruiting workflow and more importantly it drafts beautiful tailored rejection emails with really deep focus areas for improvement that were derived from the granola transcript. None of the recruiters like that I know actually do this well.
57:07 >> So you get deliver a much better candidate experience as well >> as making it faster and easier for you and potentially even making better decisions because you have the transcript. >> Exactly. Exactly. Yes. >> This has been mindblowing. I want to talk to you about some hot topics now. >> Sure. Let's dive into them. >> What's on everybody's mind right now just in the zeitgeist is what's going to happen to my job. He just showed me how, you know, you might be able to do the work of two PMs with one PM. So, what's going to happen to the future? Is product management going to exist in the future? What's your take?
57:44 >> I have good news for you. I think not only product management is going to exist, I think it's going to thrive. because I think like product managers in product management in large organizations became like burdened with layers of hierarchy with layers of corporate processes and in in many instances product management is about compliance performance status reporting so it's suboptimal product management to be honest and I think what AI really helps you to do it doesn't remove away the decision making from the process it gives you the juiciest bit. It removes a lot of the operational overhead from your work and what product managers are now more focused on is the value discovery because like what AI doesn't really know is what your customers want because AI doesn't really understand the u the customer journey well. Well, it has the funnel metrics obviously but it doesn't it cannot really go offline and talk to a customer. you have to do this. and that's the funnest part in the job because you you don't do all of this you know, theory anymore. You're only focused on delivering the real value and I don't think it's going to get go away anytime sooner. What what actually will happen is that the teams will be smaller, leaner, and faster. So do you anticipate that the ratio of PM to engineer is going to change from what it historically was?
59:22 >> Yes, definitely. I think not only it's going to change it, the lines are going to blur between what a product manager does and an engineering team or an engineering manager does. there's likelihood that they could overstep and be a support system because at the end of the day if a product manager is an orchestrator of of let's say the product quality and token budget right an engineer is an orchestrator of an execution quality and token budget right the same applies for the designer designer is responsible for the consistency of a design system which means that quality and token budget as well. So, the boundaries between those roles would be very blurry.
60:07 >> So, you're a CPO, you're thinking about maybe I need to create a new product team in a particular area. When I was a VP of product at Apollo, we would typically think about, okay, if we're going to create a new front-end product area, we probably want five to seven engineers. We want a PM. We want a designer. When you're thinking about a new product area now, how do you think about staffing that? >> No, it's it's a great question. So I I typically divide them in two different groups. So the first group is high complexity and high error blast radius. I think for those specific domains nothing has pretty much changed and it's it's all around monetization.
60:48 It's all around algorithm algorithm like search ML where every single every where a single tweak might have a dramatic impact on the conversion or on the outcome. here you would need to have a dedicated owner. but for the remainder of the domains like customerf facing domains what I'm currently seeing is that teams can be like easily scaled without adding people. You you might have one product manager owning like three four domains across many platforms at the same time.
61:26 >> Okay. So the PMs you're hiring they feel to me like they're truly AI native PMs. even if they're not building AI features, you're probably hiring people who are really AI forward. What's your process these days to hire PMs? How do you find PMs that are at this level of AI native that they can actually succeed in this environment? >> Great point. So, I typically look at three things. I think the fundamentals haven't changed. but I also added additional qualifiers into the interview. So the fundamentals meaning problem solving and systematic thinking like if you could have both you will adapt in any situation.
62:08 but additional qualifier that I've added is a level of craft when it comes to AI, right? how deep a product manager is in the AI and and there there's a very simple way to test for this is I'm I'm I'm asking like which of the use cases of your daily work you have automated using AI and I get a spectrum of responses if your response is on a on a on a level of well I'm I'm talking with chat GPT using like I don't know terminal and using like web interface and So to me that's kind of a low level of immersion or maturity when it comes to AI tool understanding. But on the other side you might have a very advanced AI agent orchestrator, a product manager who has automated all of his professional work life right and build it entire using AIS and so on. And then I can drill deeper like asking like how do you evaluate this? How do you make decisions on tools, on the brain, on the memory and scaffolding and so on? But I got to be honest with you, it's in the market where I operate, it's a rare skill set. So, product managers who are able to effectively answer those questions, they get way more points. You guys heard it from him, not from me, guys. He is a CPO at one of the greatest companies I would say to work on and he is running it in an AI native way and this is the skill set you need the skill set we taught today. So if you haven't already go play around with Hermes go play around with open claw and if you're a CPO steal Mikail's playbook his team is operating at a level unlike many others. Mikail if they want to learn more about you get in touch with you find your content where can they go?
64:07 well I do have a substack newsletter. It's called corporate waters. you can reach out and get immersed in my thinking and the actual use cases. outside of this you can reach out to me directly over LinkedIn. >> I highly recommend corporate waters. If you guys didn't know, Mikuel and I did a really fun deep dive last year where he looked through all of the interviews he's done in his lengthy career, which by the way, before he was a GPM at Bolt, he was also an ICPM for over a decade at companies like Yandex. So, he collected all that data and we looked into what interviewers are looking for. So, if you want to go see more information from me, Kyle, you can find it in my newsletter or his newsletter. Thank you so much for sharing all this amazing sauce today.
64:54 >> It was a pleasure, Akash. Thank you for having me. >> The craziest thing is that Mikail agreed to open source the information that he used to build all of this. So go check the link in the description for the GitHub repo on his profile. You can fork that and you can begin building this company operating system for yourself. Until the next episode, we'll see you later. I hope you learned as much from today's episode as I did. If you can do one thing that's totally free that would help the show, it would be to check that you're following on Apple and Spotify podcasts. Check that you've left ratings and reviews on those platforms. Check that you're subscribed on YouTube. Leave a like and a comment on this video. And then share it with your friends. We're trying to make better and better podcasts. After 2 years, we think we've gotten something pretty good going. So, let us know what we can do to make it even better, who else we should interview, and we will put on the best shows we possibly can. Finally, don't forget my offer for the bundle. You get an entire year of my paid newsletter, plus my favorite AI tools, bolt, new, air table, speechify, descript, magic patterns, linear, dovetail, arise, and mobin. That's $27,000 worth of value for just $150. So, check that out at bundle.ac.com if it interests you and I can't wait to share our next episode soon.
Summary
- Agents build and validate features, acting as an operating system connected to tools like Google Workspace and Jira.
- A knowledge graph tracks team performance and customer insights, allowing for better decision-making and autonomy for AI.
- Product management roles are evolving to focus on quality and strategic outcomes rather than traditional processes and outputs.
- AI can automate repetitive tasks, freeing product managers to concentrate on discovery and customer engagement.
- The integration of AI tools can lead to smaller, more efficient teams, with blurred lines between roles in product, engineering, and design.
- Metrics for success include knowledge coverage and recall accuracy, which can be continuously improved through AI interactions.
- The system allows for personalized stakeholder interactions, automating feature requests and status updates.
- CPOs should own the AI systems to ensure rapid iteration and alignment with organizational goals.
Questions Answered
How are agents influencing product development?
Agents are integral to building features and validating requests, acting as an operating system that connects various tools and stakeholders.
What metrics should be used to evaluate AI outputs?
Quality of AI outputs is measured by their impact on metrics, focusing on outcomes driven by AI rather than just procedural processes.
Why is it important to store raw transcripts?
Storing raw transcripts helps maintain fidelity and accuracy in AI outputs, avoiding loss of information through summarization.
How can metrics like recall be effectively measured?
Metrics can be set up by comparing responses from systems with and without specific skills, allowing for a clear evaluation of AI performance.
How can AI assist in managing emails and scheduling?
AI can digest important communications and schedule meetings intelligently, reducing the need for manual checks of emails and calendars.