transcribe

How to Build an AI-Native Product Team in 2026 | Charles Zedlewski | Product Growth

Aakash Gupta · 59m · transcribed 7d ago
More from Aakash Gupta Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

The Need for a Centralized Product Repository

Why is there a need for a centralized product repository in the context of AI?

The discussion highlights the overwhelming amount of information and requests that can flood a team's context window, making it difficult to manage. A centralized product repository can streamline the process, allowing team members to quickly access relevant information and understand customer problems more efficiently.

  • Centralized repositories can save significant time in understanding customer issues.
  • AI can help automate the management of product information.
  • The shift towards prototypes over lengthy documentation is becoming more common.
# 11:53

Harnessing AI Models for Product Management

How can product managers effectively utilize AI models?

Product managers are increasingly familiarizing themselves with various AI models and their applications. The discussion emphasizes the importance of selecting the right model for specific tasks and the benefits of experimenting with different open-source harnesses to enhance productivity.

  • Understanding different AI models is crucial for effective product management.
  • Experimentation with open-source harnesses can lead to better task management.
  • Product managers are evolving to become more technical in their approach.
# 23:46

Utilizing Customer Insights for Problem Solving

How does the customer insights tool improve understanding of customer issues?

The customer insights tool aggregates and summarizes customer interactions, allowing teams to quickly grasp customer problems. This tool enhances efficiency by providing immediate access to relevant customer feedback and issues, significantly reducing the time needed to understand complex situations.

  • Customer insights tools can drastically reduce the time needed to identify issues.
  • Automated summarization of customer interactions enhances team responsiveness.
  • Effective use of customer feedback is key to improving product offerings.
# 35:40

Managing Multiple Repositories with Orchestrator

What role does the orchestrator play in managing product repositories?

The orchestrator serves as a centralized tool for monitoring and managing various product repositories. It allows team members to interrogate codebases and track developments across different projects without needing to understand every detail, facilitating better oversight and coordination.

  • Centralized tools like orchestrators improve visibility across multiple projects.
  • They enable teams to manage complex codebases without deep technical knowledge.
  • Effective repository management is essential for product development.
# 47:33

Evaluating AI Performance in Product Features

How is AI performance evaluated in product features?

The evaluation process involves running various tests to determine how well an AI agent performs within the product. This includes analyzing its ability to access information and execute tasks correctly, with insights gained leading to improvements in documentation and feature functionality.

  • Regular evaluation of AI agents is crucial for product reliability.
  • Feedback from AI performance can lead to significant documentation improvements.
  • Understanding AI limitations helps refine product features.

Transcript

0:00 It felt like after a while that the new party foul was flooding your co-workers context window where we all just start launching slop at one another. >> But if we could put all these learnings into one place and so that became this together product repository. >> If I had to do this manually, this would have easily occupied half of my day if not more. There are 19 tickets that were filed in the last 2 months alone. So this is a pretty heavily requested ask.

0:25 So, I kind of get a pretty good sense of what the customer problem is within like 5 minutes instead of maybe a full day. Together, AI just raised an $800 million funding round valuing them at $8.3 billion. I got their product team to show you their exact repo they used to automate their work. We all discovered Chad GPT in 2023 and were horrified at the output and started handwriting. What is a good PRD in the AI era?

0:51 >> I come from Amazon. I used to work at Amazon. So we used to rising like 20 page RPDs. What replaces the bulk of that is actually a prototype. >> We're all going to live in a multimodel, multi-harness world. That's essentially the new bar for UX for the kind of products we build. A lot of these things don't get magically better with AI. >> Where does the line of product manager end and developer begin?

1:22 Before we get into today's show, please take a second to check that you're subscribed on YouTube and following on Apple and Spotify podcasts. If you want access to all of my favorite AI tools, I've gotten them to give you an entire year of their paid plans. Check out bundle.ac.com for an entire year of bolt.new, air table, speechify, descript, magic patterns, linear, dovetail, arise, and mobin. And now into today's show. Charles, so I'm fascinated by how you guys work and how you guys have been using AI. What have been some of the big unlocks for your team in productivity >> together? AI is a platform for powering AI applications and agents. And so we saw going back to the rise back when cursor was starting to hit its growth spurt how what an amazing superpower an AI agent could be for any kind of knowledge worker in in in in addition to in in addition to and including product people. but it it made me think a bit about what would a world be like full of product people and engineers that could all generate as much code and and and and content as they liked. And it wasn't clear to me that the sum of all that was actually necessarily forward progress for us as a startup. And I I imagine it's probably the same for a lot of other organizations. So what we set out to do as a team is ask ourselves what would it mean to use AI but not to use AI to just make ourselves individually productive where we all just start launching slop at one another. but rather AI to make us collectively more productive that we actually that that the code and the content and the research that we did actually advanced the company as a whole.

3:21 And that led us to a bunch of decisions about what we wanted to do centrally and what we wanted to lead people to be free to do individually. And that's largely what brought us to where we are today. >> So how do you coordinate and centralize all this information? >> Hi, I'm Nicolina. I'm a product manager at Together and you know a few months back we all we've been talking about the way that we use AI in the hallway but we never really sat down and discussed how we actually do our day-to-day work and when we did we found that a lot of people had unique workflows that made a lot of sense to the area that they were working in. you know, checking to see the reason behind a node failure, working with customer support, drafting reports, but then the rest of us were doing on the day-to-day basis, we were doing things that were largely the same, researching customer needs, creating PRDs or one pagers. And so we thought, you know, if we're really being thoughtful about how we're constructing these skills and how we're thinking about agents versus skills and now like, you know, loop engineering, what if we could put all these learnings into one place and so that became this together product repository. it's composed mostly of markdown files, some YAML files as well. And what we put in here are both context that is useful for us across the board. So we have this context document or markdown file sorry this context directory and it has a number of groups of the the product the product grouping that we have. So customer intelligence sandboxes but we also have the result of strategy meetings that we have on a regular basis where we divide up product by mission our strategy what we want to accomplish milestone over milestone quarter over quarter so that when we're thinking through our research and we're composing these documents with our agents we can pull this context as it's relevant. And it's also useful to be able to look at the context from other people's sections. So if you're building something that is touching or going to create a a joint workflow. So if I'm creating sandboxes that reinforcement learning users are going to make, I can of course sit down with the product manager from the model shaping team or I can look at the documents that she's already thoughtfully composed about their entry points and how they use SDKs and how an ideal integration might look and kind of work out a really decent proposal before I then put it before her and try to do that leg work. I'm taking advantage of the work she has done on the skills front. It's a mix of things that are useful for anyone outside of a code repository. So everything that's relevant to specific code really shouldn't live here. But if you want to do something like figure out if we are serving all of the generative media models that all of the users are interested on artificial analysis, I can have that skill here. So I could run that as somebody who supports generative media but someone else could as well and they can see the output of that in the terminal or things that we do on a regular basis. So this topline status one is one that I really love. At the end of every sprint we go through all the projects that we work on and we create an update of what's been shipped, what's ongoing, what's coming up next.

6:51 And on this skill you can just decide who what areas you're going to pull from. So for sandboxes, there's a lot of linear projects. but for the SDK and API, that pulls from a lot of projects. so you can pull directly from a repository, see what features got pushed, and then it'll form an update based off of the sources that it was fed, and then you kind of edit it, add additional updates that it missed or context that would be useful for leadership. but that's something that used to take me, I don't know, anywhere from like 10 minutes to 30 minutes, depending on what I'm soothing through with the engineers. That now takes me, you know, a handful of minutes just to type in the the prompts and then put the answers into the the right location there. and as you can see, this is all clawed. Originally, I think we for the most part were we adopted claude code. but a lot of the harnesses that support open-source models are are used to looking at clawed markdown files and being able to run them just as well. And so we've started to move more over to using our own models. one because they're fantastic and it's really nice to use our own models and two it's it's more affordable because clog is I think like many teams are learning across the globe that anthropic and all these other models are starting to become quite pricey and so this is an example of using open code it's the same repository this is one of the skills it's just a a news report that I run in the morning it looks at our like our competitors other model labs just to collect information from what happened over the last 24 hours and I will read it over my morning coffee. It takes a minute so I ran this earlier but you can see it's going to read the skill and kind of look through all of the listed sources that I want it to find news articles for. And then it will create a nice report that I can then click into and build out or read through any information that I want to learn more about that day. So if we return to the context like who owns and maintains this because the worst thing is if you're going to have like outofdate context. So is that one is that like let's say Charles sets the overall product strategy. Is it his deliverable to maintain the strategy document that is in a good format for any AI harness or how do you guys manage these types of behaviors and responsibilities?

9:29 >> I think for the areas that you know the sandboxes or the customer intelligence that I own it would be my responsibility and then on the general it it has been kind of updated based off of larger events. So if we have a a strategy or planning session, then someone would sit down and decide like who's going to sit down and translate that into what material. So there's typically a lot of documents that have been written in the leadup. And so it would be a matter of just translating that into a markdown file. but you know the CEO will write letters or write documents of direction and that's a really it's a very clean moment. you know when direction is being set and I think that speaks a lot to how like well our leadership communicates direction but you know you know that this is a pivot they make it very clear they kind of lay it out for you and so that's kind of your signal and I think as a product team it's just a matter of us like I I tend to update this I think quite a bit but I think if anyone else makes a PR it's like any other PR like if somebody takes the initiative to create a PR or make the the update to the markdown file then you know that's that but as soon as you want to pull from it and you realize it's missing then you go ahead and make that change. and going into the skills.

10:49 So, how do you maintain how do you maintain like what's going to be a team skill and an individual? Maybe somebody has like a different way of writing a PRD. What is the sort of prescription or guidance? You should be using our team PRD. You should be customizing them. What's the line? >> Yeah, I think when it comes to updating skills, there's a rule set on, you know, where each skill should live. You want to keep it closest to the the work you're trying to do. So if it's a skill that is referencing code or is housed within a if it's housed within a a set of work that a team is working on that's always in that one repository then keeping everything colllocated makes a lot of sense and then all the rest of the skills need to go somewhere. So, you're either going to have a personal repository or you're going to have a shared repository and potentially you create a branch where you test out a skill and you say, "Okay, first I I think this is relevant. I think this is repeatable. I'll see if I use this a few times and if you do, then you can push it up to main." And it that's been my approach so far. And if it's something that I think is super niche, then it might just stay in a branch that never gets pushed.

12:03 >> Makes sense. And so everybody can connect into this harness using whatever model they want, but the harness primarily is these contexts and skill files. Is there any other components that people need to know about? >> So there's the the repository and then there's the harness. And the harness there's many open source harnesses. Open code is one of the primary ones that we've we've been watching as it's evolved over time and supported a lot of the models that we support as well.

12:33 but Hermes has also recently published one that I've enjoyed playing around with. And as you figure out which harness you prefer in your day-to-day work or if you're looking for models that suit a specific workflow, you know, I think a lot of product managers are starting to become a bit more familiar with what models are suited to different tasks. But as more open source models are coming out, I think it's really interesting to see, okay, like is GLM how is GLM for coding? How is it for creating these product prototypes or Kimmy for doing analysis? And I really enjoyed that that experimentation because as these models evolve, you feel like you've got a lot more control over pairing the right model for the capability that you're after at the time.

13:22 >> I've been building a lot of AI products lately. My job search OS has 16 different agents. My newsletter has a recommendation engine. And I kept running into the same problem. I'd ship something, it would work in my testing, and then I'd get messages from users saying it's hallucinating or picking the wrong tool. The issue wasn't the prompts or the tools. It was that I wasn't actually evaluating anything. I didn't have a way to see what my agent was actually doing step by step. Every tool call, every decision. That's where Arise comes in. Let me show you. I'm going to open Claude Code and install Arise with just one command. npx skills add Arise AI Arise skills skill. Yes. Now Claude Code already knows how to instrument my agent. I tell it set up tracing to Arise and it automatically analyzes my codebase, figures out where the LLM and tool calls are and adds instrumentation automatically. Now I can see everything, every trace, every span, every decision.

14:18 And more importantly, I can evaluate it. That's the shift. Trace what's happening, evaluate where it fails, then fix it. This trace right here, my resume feedback agent was supposed to pull the company's text stack from the job posting, but instead it hallucinated that they use React when the posting said Python. I never would have caught that without seeing the trace. And here's the part that blew my mind. I asked Claude Code to look at these traces and tell me what I should be evaluating. It came back with four eval criteria I hadn't written. Things like picking the right tool and staying grounded in the input. I wrote the evals, ran them, and found that my agent was making the same kind of mistake about 12% of the time. Claude pushed a fix. I reran the evals and it dropped to under 2%. That whole loop, trace, evaluate, fix, took me about 20 minutes, and now it runs automatically. If you're building AI products and not evaluating them, you're shipping blind. Try Arise free at arise.com and get a year free, a $1,260 value with my bundle, Arise. Check it out. It's one of the top AI evals platforms used by all of the top AI teams for a reason. Amazing. So, Pav, can you show us what this is like in action? How does someone dayto-day use this for customer interviews, writing PRDs, and getting things into a state that they'll hand it over for engineers?

15:38 >> Yeah, Kash. I'm Pavit Alawalia. I'm the product manager for infrastructure at together AAI. So how it really comes into action is basically based on two processes. One is the initial discovery and aligning process and the second part is the build and ship which is more automated and really what this ensures is that as a PM when I'm shipping a you know PR or I'm writing some code it's not noise to engineering it's actually grounded in our best practices and architecture. So let me show you this in action. So I have my cloud code here.

16:08 I'm going to first start with like see a feature that I want to work on, right? So I let's say I want to do some research on a potential feature that I heard some customer complaining about. So one that I'm looking into right now is how many customers are complaining about shared storage resizing. And what this feature is going to do is it's it's basically linked to pylon which is where our support tickets lived to linear which is where all our project and engineering execution is tracked and to notion where some of our internal product documents live and it's basically pulling all the information from there to show me how big is this as a customer problem how many customers are complaining about it give me some verbatims and some tickets that I can go and deep dive into. So you had mentioned that certain parts of the product process are more automated than the others. Which do you feel like are more automated?

17:03 >> Yeah. So the way we see it is the part which is where the decisions are being made right which is defining the feature defining the API surface area the abstraction layer those are very human in the loop right where I'm hands-on with cloud. And the codew writing part, the execution part is actually the part which is automated. And how we would typically do that is like I would basically trigger a goal, right? And say >> give me production ready PRs for a feature to resize. So exactly the feature that we are going to work on, right? Resize the shared volumes for a running cluster without destroying data. Right? Just hypothetically. Now, typically as part of this goal, I would also give it a few other tasks. For example, give me a design doc that I can review with engineering. And first, let's verify on a PC cluster. So we would typically deploy it on a cluster just as a P to validate everything it's working as intended and this is where because goals are amazing this will spin off multiple sub aents and actually start tracking all the pieces and it's going to prompt me for information along the way for example like a PC cluster or more information about the feature and that's where we'll feed it the PRD or maybe a UX prototype that I built along the way. Hm. So very very powerful engineer skills. It sounds like you've have a lot of confidence and the engineering or has built skills that you can just give a goal command like get a production PR ready and it will go pull the relevant skills. It will be able to read your codebase and get something actually production ready.

18:55 >> Absolutely. We actually have a full engineering repo where it's just our architecture details, right? It's a bunch of skills which tells it for things that we've done in the past how to redo it like runbooks engineering architecture design documents which claude has access to through our shared repo and and that's why it's able to build highfidelity PRs in in like a first attempt really. >> Wow. I didn't realize just how good these had gotten at AI native companies.

19:25 So that really then like you were saying in the human in the loop parts gives you a lot more time to focus on those in depth. >> Yeah. So if you see right now it's actually spun off a sub aent to look at how is the data plane working how and then another sub agent to check how is the control plane working and now a new one to look at docs right is there existing internal or external documentation on how we might do this. M and do you have a point of view like you're using looks like cloud code in the terminal. How do what's the easiest way to use a shared team repo harness?

20:00 >> Yeah, so actually if you look at this, it's currently running in my OS directory which is actually pulled a lot of the shared repo artifacts. So for example, Niko shared a few artifacts around research. So it's actually this skill, the research feature research skill which is actually a shared skill that everybody uses in the team. Then there's another one which is my personal favorite which is PR writer. This is actually pretty cool like let me just actually trigger it off right. So resize shared volumes for existing running clusters. So what this actually does is it does a turnbyturn interview of it'll interview me basically on what are the decisions I want to make in this feature and it and if I show you the skill so this is the actual skill it ingests it can take a prototype so if I already have a UX prototype it'll ingest that if I have a running PC it can take that or nothing just like a simple prompt right and then it has access to all the information the shared context that Nico shared about, right? It's pulling all of that in and then it's going to ask me questions turn by turn and it's specifically instructed to not, you know, to push me to challenge me on my assumptions and basically the decisions I'm making and then it goes into drafting a PR and we've given it a question bank basically to kind of ground it in what type of questions typically we need to make for example like what are the trade-offs is there any one-way window decision we are making along the way.

21:37 >> Amazing. So we've spinned off. I think we have currently three agents working for us, right? Research, PR, and PRD. >> Yeah. So this one is actually done. So the research agent came back. It's saying all these customers have asked for this feature in the last 6 months. There are 19 tickets that were filed in the last 2 months alone. So this is a pretty heavily requested ask. I have the source. So I can actually go into these tickets if I want to dig deep exactly what happened and why did it ask why did customers ask this >> and then >> what was pylon again?

22:10 >> Pylon is where our support tickets live. It's the support platform. >> Mhm. >> And then actually it figured out that we had started working on this feature because it has access to linear as well and it was partially implemented but it's broken in some way. So we stopped working on it. So now I'm not duplicating. I'm not like replicating work. I can actually build off somebody else's work or actually talk to that engineer and figure out, hey, what happened? What issues we ran into?

22:36 >> And is this hooked up to the live codebase? So it could go check like in case there's some discrepancy between linear and the codebase. >> Yes, because it's in my personal OS directory, it has a skill which gives it access to all the internal GitHub repos and it knows what's what can map that. By the way, Kosh, later on I can show you like a god's eye view of like what this looks like across all products for repos as well as across all customer feedback.

23:04 >> Okay, I'm excited for that. So bum, here we've got the research report. What's your take on this research report? Is this a like how long would this have taken you? What do you give this out a 1 out of 10? >> Yeah, so if I had to do this manually, this would have easily occupied half of my day, if not more. and this is actually pretty good. It gives me the brief snapshot that I wanted to know which tells me the feature passes the smell test. There are enough customers asking for it. I know most customers are actually want to increase their storage volume, right? So I actually know what's my like MVP use case and then I have a source so I can actually dig dive into some of them to actually get a better sense of the customer pain and it gives me verbatims as well which kind of helps. So I know for example here it call it it calls out a customer called moonlake and we have a separate tool called customer insights where all customer calls live. So let me go into that. So this is the customer insights tool where every customer call that our sales team is getting into all our support tickets are getting summarized and cataloged. So I can actually go into this and actually see what exactly happened with Moon Valley.

24:13 So they wanted a 50 tabyte but there's a lock icon in the UX which is preventing the action and they are basically they reached out to sales team and they basically want a self-s serve workflow right so I kind of get a pretty good sense of what the customer problem is within like 5 minutes instead of maybe a full day and this customer insights tool so you guys have built an MCP server sitting on top of Gong and this has been visualized by Vzero that's what we're looking at right now >> exactly covers both Gong Slack Pylon yeah at least those so it's actually more than just go >> and who maintains this sales ops >> you want to speak to that it's largely automated but >> yes me and my team maintain it so my name is Assan I run the developer experience team here at together and you know me and my team have built a bunch of of this kind of stuff we we maintain it we continue to add new features but it's largely kind of working by itself at this point we have like a daily crown job that runs every day that grabs all of the calls, all of the pylon tickets, all of the like Slack channels, everything that happened in the last 24 hours and it adds it to the database that this is running on. and yeah, what Ponita is showing is the daily screen that like where we show like 5 to 10 kind of insights from customers every single day based on all these calls. We also have an MCP server. We have a chat where you can ask it anything. there there's a lot of kind of facets to this tool as well but yeah my team continues to maintain it and and kind of add new features here and there.

25:44 >> Okay so developer experience maintains like various MCPS what other MCPs are you maintaining? >> Good question. we have this one I mean some of them the the a lot of these tools are exposed differently. this one has MCP server for example. Another tool is called orchestrator. Charles is going to show it off and talk about it momentarily. that one has kind of just a UI and as a wrapper on top of all our our GitHub repos. and yeah, really my team just experiments with a lot of this stuff. We build things that we think may or may be useful for for the the product team and the rest of the company. and some of the internal tools kind of flop and we're like well you know this is this is not used very much and and probably our two some of our most used ones have been this like customer insights MCP and app and and orchestrator and agent evals which I'll also show off. So those are like the the top three that that have been that people have been using.

26:39 >> Sweet. So we get to see all three. Awesome. And Pub, how's our how's our PRD and our PR agents looking? >> Yeah. So, so it's now asking for evidence. So, this is where typically I would run this in a single cloud chat because this is sort of somewhat sequential work. But typically then I would take the output of this research which is pretty detailed for like a starter starting point and I would actually give it to the PR agent to run with. So now it has a sense of the customer problem. So it can ask me more educated questions on hey what are the trade-offs? Should we build as an abstraction? Should it be self-s served?

27:18 When do we pull in support? Those kind of questions. >> Mhm. So almost like managing different agents that feed each other. >> Okay. So now it's asking me, hey, do you want to give me a Figma or a running PC or some screenshots or I can just say no, run with it like let's do it turn by turn by discussing it. >> we all gone through phases of using AI for PRDS. You know, we all discovered Chad GPT in 2023 and were horrified at the output and started handwriting. Then we discovered Claude and said, "Okay, maybe it can write well." And then it almost feels like, you know, people were overloading our colleagues context windows with long documents. What is the what is a good PRD in the AI era?

28:02 >> Yeah. So, that's actually a question we thought a lot about. So where we really use PR like historically like companies started using PRDs as like a gating document where everybody all stakeholders need to come together align on it and only then we'll start building the feature. We see it a little bit differently where it's a tool to trigger ideation and problem solving. That's it. And it's short. It's usually one to two pages. It it defines the customer problem well so that everybody has a shared understanding of the customer problem we are trying to solve. It lays out some solution options, right?

28:33 doesn't have to be fully thought through necessarily but at least like what are the different ways we could go solve this problem right and then a sample user journey like based on the option we want to go with it'll define either a API based journey or a UX based journey to show hey these are the three four steps that a user would go through in in the solved world right once this feature is shipped and that's it like that's enough to actually have an detailed discussion around hey how should we build this feature what should engineering design look like is it even worth building or not? Right? And and that's it. That's what the PR template that this skill uses builds off. And actually what replaces our typtoal six pager like I come from Amazon, right? I used to work at Amazon. So we used to rising like 20 pager PRDS. What replaces the bulk of that is actually a prototype. So we built a separate skill.

29:27 Let me start a new cloud window. And there's a separate skill called UX prototype which actually you give it that one pager that we just discussed and it'll give you a very detailed prompt on how to build it in like a Figma make or any design by code tool like you can use it in VO you could use it in cloud design or even Figma right and that's where a lot more debate will happen because now everybody can visualize it visualize the solution see what's how it's going to interact with the user and That's where we find like a lot of the best critical feedback comes from from engineering, marketing or anybody else.

30:07 >> Do you find you're living outside of cloud design then a lot you're getting this to generate the prompt and then giving it to Figma make? >> Yeah, the only reason I use Figma make is because it's easy to share a Figma make link and others can iterate on top of it. It's easy to click around. Claude builds like HTML files which I which it's hard to keep version tracking and stuff like that. Figma is just like more sharable. That's why we use Figma make.

30:29 Makes sense. So, how do our how do our PRD and PR agents look? >> Yeah. So, okay. So, now it's getting to now it's asking me questions, right? So, it's saying, hey, the evidence splits into two different directions. One is multi-tenant substrate koda bumps and the other is dedicated conversion. which one do you want to go for? I actually want to go in a very specific direction. Just looking at like the feedback we got from customers. I want to say hey we should focus on increasing the attached storage volume sizes or tenants. Tenants here is basically a customer cluster, right? So tenants in a running cluster.

31:16 >> So here's where like your domain expertise, your human judgment is coming in. >> Yeah. like it's not a substitute for me, but it really does two things. A, it speeds up the whole PR writing. Like it solves the blank page problem to a large extent, but it also catches a lot of things that I might not have thought of. Like I often find it's asking me questions that I might have missed if I'm not spending a lot of time thinking about the problem.

31:40 >> So ignore Becca issues. like I'm just giving it like a sample to focus its energy and to give it like reduce the scope so it's not trying to boil the ocean in the problem space. >> Mhm. So a 10 out of 10 that's comes out of it. It's going to be kind of that very short document that accompanies a prototype. >> Yeah. >> If it if this all goes well. >> Yeah. So I actually did this and I have that PR. If you want I can show it to you what that looks like.

32:07 >> Yeah. Let's take a look. >> So this is what the final output looks like, right? So first of all, it's telling like where's the evidence? It it looked we had 20 pilot tickets across 14 customers, right? So straight up, you know, there's evidence and this is what the looks like. It first defines the customer and the business problem like we talked about the 20 confirmed customer issues. It's name dropping some customers just to give us a sense of which customers, right? Is it our top big customers? Is it like a long tale of customers? And then it gives you a sense of like what is the actual pain, right?

32:39 it seems like we cannot update this in the UI and that's basically the ask they just want to selfs serve the volume resizing and that's like pretty obvious from the customer verbatims and then you go into goals and non- goals this is important for scope creep like which is like the number one problem I feel like PMS and and engineering faces so it gives you some goals like it's giving me a goal that hey if done well this should basically handle over 90% of resize requests right it works in both directions upwards and downwards like scale up and scale down billing should adjust immediately. See, this is something I did not think of actually.

33:16 when somebody reduces the size or increases the size, we need to make sure that the billing matches up and they're not over or undercharged for it. And capacity block request. What happens when we don't have enough capacity, right? If it's a fully self-s served experience, when do we actually need to get support involved? And then it definer stories so that you know whoever is reading this and when we are doing a discussion in engineering, everybody understands the key use cases. And then actually the proposal overview, right?

33:45 It's giving me some API specs. it actually gives like a detailed API design as well. So we're going to add to our existing API some changes on how to basically allow for resizing. This is the user story sort of right like what are the different API specs and what is the API calls and then UX as well. End to end user flow happy path what are the steps that a user will follow. So I feel like it always gets the headers right but sometimes the devil is in the details. Is it nailing the details?

34:14 >> So let's look at this one. Right. So admin sets a new size in the console. API validates. This is some internal detail actually. So maybe this part not necessarily needed in the user flow. this is actually internal detail and then it shows the status. It shows that it's available and billing rate is updated. So, it got the user journey right, but it bled some internal detail into it as well. >> So, it it's still going to need a little bit of editing. You can't just immediately take this and start sharing this with your colleagues.

34:47 >> Yeah, absolutely. Actually, at the top it says this is a draft one pager PR. It's not meant to go wide distribution for everyone. And that again kind of goes back to the process that we were talking about where in this initial discovery and design phase, it's human in the loop. So the expectation is that the PM will go and make changes, maybe make some changes to the API, maybe change the scope a little bit before we circulate this widely.

35:14 >> Very cool. So we've got the repo, we've gotten to see it in action. We got a little preview about an orchestrator agent. Charles, can you walk us through this? How do you get to see this bird's eye view at the top of their product ladder? >> So we built this great internal tool which we call orchestrator. The basic idea is each of the individual product people are sort of married up hand and glove with their corresponding engineering team. And Pab kind of gave you a good example of how that works where like he wants to be able to go pretty far into the definition and implementation for some of the things he works on and his the specific engineering team he works with has given him the kind of skills connected to their codebase that he can go do that.

35:58 From my vantage point, what I need to do is be able to like get a check on like where things sit in all kinds of different things that we're building as a company and I don't want to have to like build the equivalent of PubMed and Assan and Nico's environment. So I have this nice tool here called orchestrator and you can see that essentially all the major repositories we have as a company for all of our products are represented here. So if you see I'll give you kind of a quick examp. So you have like let me find pubs tcloud. Yeah. So you can see here we have the repo that pub was just living in t-cloud but you can see that it coexists along with many other repos. So for sake of argument like one of the other products we have we call model shaping which is basically the ability to adapt the behaviors of openweight models and I can decide how I want to interrogate what's going on in that codebase what's going in that product area. I can pick my harness. I can use either open code open code claw or cursor. I can pick my model. So in this case like I tend to be a you know GLM open code kind of guy. and I could ask some question like what's the most recent model we have enabled or supervised supervised fine-tuning right so for we we basically have like more than 30 models that you can adapt and you know fine-tune but these are changing all the time I'm not I don't want to have to go understand the entire world of the product lead for model shaping I just want to understand what's going on in this one specific, you know, to answer this one specific question. and you can see that essentially what we'll do in this case is we'll actually create a sandbox. We'll call in the repo. All this happening in the background and in a minute it's going to basically interrogate that portion of the codebase. and it's going to go research like what are the most recent changes and let me know what happened most recently.

38:01 >> You can see it kind of coitating right now. like the proverbial cooking show. I have like kind of a a a synopsized version of like the conclusion. and in this case it looks like the answer is the last model that we enabled for supervised fine-tuning is the new Nvidia Neatron super 12b model. so I can do this exercise essentially across like any product any any portion of the codebase for that product. and I can and it's not just limited here. I'm just asking questions. But if I also find some part of the product that annoys me and it's something small, I can actually generate a pull request from here as well.

38:47 >> Okay. So you're mainly living in the orchestrator, not inside cloud code and the team repo. >> I use cloud code and the team repo for if there's some requirement. I'm writing myself for some part of the codebase, then I would do that. But if it's me shopping across like all the different products we have and it's like some small UX change or things like this, then I would sooner use the orchestrator because this is basically not just pointed at all the different repos we have as a company, but it's inheriting all of the skills and or MCP servers that are local to each of those repos. So, I don't want to have to build all of that into my local open code just to make one small change. It's a lot simpler to just to just use Orchestrator and it's and I know it's going to have the latest greatest skills in MCP servers for that portion of the codebase.

39:41 >> Got it. So, who are the other users of orchestrator within the company? Like what who is this product exactly built for? >> It's sort of intended for casuals, right? It's intended for so so it's intended for like let's say PubN wants to investigate something in inference or let's say Neilen wants to make a suggestion on on our infrastructure as a service. they don't want to build and replicate all of the specific skills and context native to that particular portion of the codebase. It's a lot simpler to just look at just to use orchestrator where everything is kind of maintained server side.

40:21 >> Okay. So this is like for whenever you're casually working maybe like trying to learn something about another team. This is also creating that connectivity. >> Yeah. One of the things that we talked about as a team was like like like it goes back to this point about how much is shared and how much is individual, right? And there was a point at which I thought, well, wouldn't it be great if everybody knew what everybody else was doing? And there was sort of like we all had the same context and everybody else at all times. And in reality, most people don't have a whole lot of motivation to want to understand all the depth and nuance of the context of somebody else's area. They want to know just the amount they need to get one single question answered or to get one single problem resolved and they're really not interested in the rest. So we sort of we sort of set aside the idea that there was one big broad flat set of context that we're all going to swim in and it's much more like a context hierarchy and some of us belong all the way down to the bottom of the depths of that hierarchy and some just want to traverse the top.

41:26 >> Fascinating. Okay. So if a team company watching this, they wanted to replicate what's being built in orchestrator, how are they going to spin up their own version of this? >> Yeah. So I mean this like from what I from what we see in terms of our own customers, this is becoming increasingly common. Like I was at a AI conference in in Paris the other week and I saw like a like a like a digital services and marketing company called Process and this idea of essentially they they basically already built their equivalent of this same thing where they have like every every portion of the codebase you want to change is all in one place.

42:11 everything can be spun up as a sandbox and generate its own unique pull request or its own questions. So, I I think this set of this set of tools is like not that out of reach for most software development organizations and this idea of like we're all going to live in a multimodel multi-harness world and the main the main endeavor then is how do you organize shared context? That's that's kind of I think what everybody is starting to build. to give a more specific answer, I mean Hassan, what was the total like time invested would you say to build to build orchestrator?

42:51 >> I would say a few weeks of work, maybe a month of work to to to build it of engineering time. yeah, and a lot of it was kind of just like trying to figure stuff out fairly early, trying to figure out the right architecture, the right tools to use, a lot of that stuff. But as Charles said, I think like this has kind of been replicated by by a bunch of other companies. There's like a few open source versions of this as well that that exist out there. So it's it's kind of easier than ever once you have the architecture. But yeah, for our team it took it took a few weeks. So like if we're just drafting the PRD for an internal orchestrator for yourself, what are the key things? It needs to hook into like all of your code bases and your different repos. It needs to inherit all the active MCPs and skills from the different team repos.

43:36 >> It needs a sandboxing mechanism because you're basically going to like essentially like each individual thing you're researching or each pull request you're going to propose that all gets done in a sandbox environment. So you're basically cloning a fraction of the repo, right? The portion of the repo that's necessary in the sandbox to generate that that one pull request. >> Got it. you also going to need some form of a model gateway or router. So you notice I had like a whole choice of models that I could use. so that's that's typically what another another piece of the puzzle. But as the son mentioned, every one of these components I've mentioned so far, there are either like either it's relatively quick to build. There are open-source tools and libraries that let you do these things. And there's commercial shrink wrap software if you don't feel like doing either of the previous two things.

44:26 >> Fascinating. So, we're walking through the whole product development life cycle. The last step, evaluating how these things are going live. Hassan, how are you guys doing agent evaluations? >> Great question. as you said, this is kind of the final step that we do. I'm going to share my screen. I'm going to go over this tool that that our team built called agent evals. And just to give a little bit of context on agent eval you know I I I think the whole world is moving to this agent centric way of doing things right you don't you don't kind of manually write code anymore and and kind of it's moving up the stack with like you just give your agents stuff to do and so you know my team is called developer experience I think Charles is is this close to renaming it agent experience and because it's something we're we're increasingly thinking more about and it's very very top of mind we've rearch rearchitected our docs in kind of this agent first way where it works very very well for for humans obviously but also for agents you actually on some pages on our docs like agent when when an agent is reading them we'll inject something extra if we think it'll help the agent we worked on it together MTP server we worked on together skills and so we're exploring all of these ways that we can make it as easy as possible for agents to use our product and agent eval is kind of the the tool that ties them all together or the tool where where we can actually say like okay like we're we're pretty confident in in in how this works. So, I'm going to share my screen and we're going to go over this.

45:54 Fantastic. Okay. So, agent eval there's a lot here. but agent eval is is the tool we use kind of at the end of this product development life cycle. we have the product. It's shipped right some kind of some major feature or a product or or second version of the product and now we need to validate that it actually works well with agents. So what we do is I tend to work with a lot of the product managers in each of the different areas and we write a series of tests of like the main like here are the main ways that we think users will use a specific product. So let's actually take fine-tuning as an example or we'll take yeah let's take fine tuning as an example. So if I click on fine-tuning this is a prompt that we give for this is one fine-tuning task that we're testing. So here we're saying, hey, we gave it a data set and we said, hey, run a fine-tuning job for this data set and when it's done, spin up this new fine-tuned model as a dedicated endpoint and evaluate it, right? Where it's going to send it some inference requests and then delete the endpoint when you're done, right? So we gave it and this actually touches on a few of our different products. It touches on our finetuning product and our inference product. And so some of these tasks are are a little more simple, some of them are a little more complex. But the point is that the the most important part is to define a series of tasks and give it to this to this tool. And what we do what we do actually has a very similar architecture to orchestrator which we talked about is so what it'll do is it'll spin up a sandbox. It'll spin up right now cloud code and it'll and it'll give us it'll give it this prompt, right? It will give it a together API key and then it gives it whatever it needs for the prompt. In this case, we gave it actually a u a data set. and then it just watches cloud code do its work, right? It's just it lets it do its thing and then it evaluates it at the end. And for us, this is like the best test of like can an agent actually do stuff in your product. We try to make sure we try to make everything in our product we try to make it so that you can do everything in our product for example in our UI as a human but also as an agent through our API or CLI or SDK. So in this case it actually got it correct.

47:58 but we go very very deep here. So we run a lot of different runs. Some of these runs are just using our docs as contacts. Some of them use our RCP server. Some of them use our our our skills. And we can click into every one of these and we can get a lot of information on the run how it worked. We can get some highle info. We can get improvements which this is one of the most important pieces of this. And and this is something we see a lot where we'll ship a feature or or we'll like in the validation stages of a feature and we give it to Asian evals and then it's like oh well I had trouble doing this thing right I I wasn't able to like for example here I said it wasn't able to discover what fine-tunable models we have right obviously this is a this is a big problem if it can't find this and and it said hey like it can't it couldn't find a docs page listing currently fine-tunable models and so we should do this and so we make a ton of improvements to our docs based on this in this case actually what happened is we had a page that listed all the fine-tunable models, but the agent actually couldn't find it. We were able to go into the transcript. You can go into the transcript. You can see the full transcript of this agent and what it did, right? We gave it this initial prompt and you can see all of the different terms. We can see that the the agent started exploring the repo. It started looking at the data set. It like did all of this stuff, right? Like the full the full transcript. and and and we we looked at it and we saw like actually it had trouble finding that one that one page in the docs cuz it's not linked in our main fine-tuning quick start. So then we're like okay well we'll just go submit a PR to our docs and add that page to our fine tuning quick start. So the agent evolve has been responsible for for like dozens of like docs fixes that we've done. and and it's been really really valuable to really understand like and get this bird's eye view from agents on like how good are agents at using this particular tool or this particular API that that we just shipped.

49:47 >> And Kosh I just want to like add on I mean you think about any sizable feature or product that you'd build in the past and the idea about like well did you get the design right? like did it actually, you know, can the can the outside user, whether it's like someone technical like a developer or non-technical be consistently successful using that feature? And think of how long you had to wait to validate that as as a as a product person before. you'd have to wait till you did a bunch of user tests and then you're going to get a bunch of conflicting signal or you're going to have a bunch of developers evaluated and they're all going to say, "Well, I like I don't think this is sufficiently pythonic or I don't like the way you did this syntactically."

50:34 and so it's like it's really hard to get timely signal and it's really hard to get objective signal. and here essentially you know at this point for a lot of our products agents are already the majority user. and so the idea is that the the nice part about that is now did the feature work from a design point of view. This is something that we can validate immediately and continuously. so, it's really powerful to sort of always know where you're at and it's really powerful to have the confidence to know that whatever your documentation says your product is, you can be sure that it actually is that because essentially a few hour every few hours we have agents revalidating that.

51:22 and across time, what we plan to do is keep expanding the range of harnesses and models that we do this for because that's essentially the new bar for UX for the kind of products we build. >> Okay, very cool. So, if your product isn't used by agents yet much, does this still have value? Is this able to kind of simulate what a human is like? And should people still be setting this up? So in in if if for us it is because our users are developers and so they're going to use our SDK and this is basically validating that there's like valid successful paths using our SDK. If we had a more gooey like you know you know web front end intensive product we would have to adapt this tool to use more you know like a visual reasoning model that could actually do the clickthroughs and like interpret the screens and see whether or not that was like naturally intuitive to the model.

52:22 So it could be extended to human ccentric examples. but for our case where agent use is already so popular it works as is. >> M all right. So we've been able to cover front to back from customer research through to actually evaluating how agents would be using your feature. This is kind of the whole product development life cycle. I'm curious where this ends. We kind of drew the line here at create the PRs, create the prototype. Where does the line of product manager end and developer begin?

52:59 >> Trying I want I want to think how to best respond to that without like just like repeating what PN just said at the beginning of our of our whole thing. I think that the the essence of what the product person does versus the essence of what the engineer does, ironically, is probably not all that different than what it was before AI in the sense that the most valuable thing the product person can do is bring a unique insight about the market that's been well validated by lots of internal and external context. And that's essentially what you still saw Pavn do.

53:40 And the essence of what an engineer does to add value to the company is to arrive at a design that is the most efficient way to meet a need. that that is also that is that that that is differentiate that the the essence of the engineer's job is to arrive at a design that most efficiently satisfies a need and is also something that you can maintain and extend and contributes to sort of the long-term architectural strength of the product. So like these two these two the essence of each thing I don't think is actually all that different today than it was before AI. What's different is the convenience with which the product person can reach into the to to to the engineering world and and accomplish small small and mediumsiz tasks and the inverse is to is true as well. The degree to which the engineer can reach into the product management process and answer their own questions. So just like it's possible for me to interrogate the codebase and make small pull requests, it's just as possible for one of my colleagues in engineering to use that customer insights tool and do their own analysis and have their own standing query for for whatever the customers have been asking for in the past month in their area. So, we sort of like it's sort of easier for each side to reach into each other's area to do small things, but in terms of like what each person's supposed to contribute that brings their unique talents and perspectives, I think that's actually the same core as it's ever been.

55:16 >> Amazing. So, I've been preaching to people, get your PM OS, get your team OS, get your company OS. You guys just demonstrated are actually living that reality. So, you're at the very top of the AI adoption curve. And I heard a really interesting observation from Chimat Palhapatia. He said you know our token costs are something like doubling every 70 days but our actual engineering productivity is just up like 5%. When you look at it from a product lens on that like how are your costs on AI growing and are you seeing a tangible like some sort of tangible productivity gain that you can point to from it?

55:56 >> Yeah. So I mean we never were in a place where we were doing like story points or other kind of velocity measures. So I can't prove it on that level but I would definitely say that our our velocity has gained more than 5%. I find claims of 3x to be very suspicious. I don't have any like when you get a team of let's say a dozen engineers and a product manager. there's so much of building software which is discovery debate kind of re-evaluation coordination and a lot of these things don't get magically better with AI. So even if you compress the research and even if you compress the coding and the testing I think that's worth a lot in terms of velocity. I don't I don't know that I I would say our experience has been that like 3x or or some huge multiple like this is is is the case.

56:55 >> M >> and as far as expense goes I think we we went through the same surge that a lot of folks did. It was easier for us to mitigate because we can use our own openweight models and they're a lot less expensive. but the other part is it goes back to this first point which is I think if if you're focused on making teams collectively productive as opposed to yourself feeling individually productive by producing lots of output.

57:23 I think it's I think it's unlikely that you wind up with these crazy like you know token budgets of three times people's salaries or things like this. I don't I don't know that we we ever reach that type of peak. Mhm. I like how measured you guys were in selling the benefits of all of this where none of it was overhyped, but we got to tactically see how it helped both individual ICPMs like Beneath and a product leader like yourself, Charles. All four of you guys, thank you so much for dropping so much insider knowledge about how Together works.

57:57 >> Thank you for having us. >> I hope you guys enjoyed that episode as much as I did. We they have been kind enough to literally open source the tools, the repo, the orchestrator, how they built this. So check the link in description below if you want to get your hands on how Together actually built this and replicate some of this in your own company. I highly recommend you do that. Podcasts are one thing, but actually implementing it, that's when you really get the ROI. So I hope you go do that and we'll see you in the next episode. I hope you learned as much from today's episode as I did. If you can do one thing that's totally free that would help the show, it would be to check that you're following on Apple and Spotify podcasts. Check that you've left ratings and reviews on those platforms.

58:40 Check that you're subscribed on YouTube. Leave a like and a comment on this video. And then share it with your friends. We're trying to make better and better podcasts. After 2 years, we think we've gotten something pretty good going. So, let us know what we can do to make it even better, who else we should interview, and we will put on the best shows we possibly can. Finally, don't forget my offer for the bundle. You get an entire year of my paid newsletter, plus my favorite AI tools. Bolt, new, air table, speechify, descript, magic patterns, linear, dovetail, arise, and mobin. That's $27,000 worth of value for just $150. So, check that out at bundle.ac.com akashi.com if it interests you and I can't wait to share our next episode soon.

Summary

The discussion centers around how Together, an AI-driven company, enhances productivity through a collaborative approach to AI tools and repositories. By centralizing learnings and automating workflows, the team aims to improve collective productivity rather than individual output, leading to more efficient product development processes.

- Together AI has created a product repository to centralize learnings and automate workflows, significantly reducing time spent on tasks.
- The company recently raised $800 million, valuing it at $8.3 billion, indicating strong market interest in their approach.
- The team emphasizes using AI not just for individual productivity but to advance the company's collective goals.
- Product Managers (PMs) are encouraged to create concise PRDs (Product Requirement Documents) that focus on customer problems and potential solutions, often supplemented by prototypes.
- A tool called Orchestrator allows team members to access various product repositories and skills without duplicating efforts, facilitating cross-team collaboration.
- Agent evaluations are conducted to ensure that AI tools function effectively, providing timely feedback and improving documentation based on agent interactions.
- The team believes that while AI can enhance productivity, the core responsibilities of PMs and engineers remain unchanged, focusing on market insights and efficient design.
- Overall, the integration of AI tools has improved workflow efficiency, although the team remains cautious about overstating productivity gains.

Questions Answered

Why is there a need for a centralized product repository in the context of AI?

The discussion highlights the overwhelming amount of information and requests that can flood a team's context window, making it difficult to manage. A centralized product repository can streamline the process, allowing team members to quickly access relevant information and understand customer problems more efficiently.

How can product managers effectively utilize AI models?

Product managers are increasingly familiarizing themselves with various AI models and their applications. The discussion emphasizes the importance of selecting the right model for specific tasks and the benefits of experimenting with different open-source harnesses to enhance productivity.

How does the customer insights tool improve understanding of customer issues?

The customer insights tool aggregates and summarizes customer interactions, allowing teams to quickly grasp customer problems. This tool enhances efficiency by providing immediate access to relevant customer feedback and issues, significantly reducing the time needed to understand complex situations.

What role does the orchestrator play in managing product repositories?

The orchestrator serves as a centralized tool for monitoring and managing various product repositories. It allows team members to interrogate codebases and track developments across different projects without needing to understand every detail, facilitating better oversight and coordination.

How is AI performance evaluated in product features?

The evaluation process involves running various tests to determine how well an AI agent performs within the product. This includes analyzing its ability to access information and execute tasks correctly, with insights gained leading to improvements in documentation and feature functionality.

© transcribe · For agents Built with care and craft by Gokul Rajaram