transcribe

Build and Own Your Software Factory

Mastra · 58m · transcribed 4d ago
More from Mastra Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Introduction and Launch Success

What is the context of today's workshop?

The workshop is hosted by Alex Booker with co-founder Abby Ay, discussing their recent successful launch of Masteractory on Product Hunt, where they ranked number one.

  • The workshop features a discussion about Masteractory's recent launch.
  • The launch was successful, ranking number one on Product Hunt.
  • The hosts are currently at an offsite in San Diego.
# 11:46

Overview of Masteractory

What is Masteractory and how can it be utilized?

Masteractory is an open-source package that allows users to create their own deployments using npm. The hosts discuss its features, including the ability to customize branding and manage tasks through a user interface.

  • Masteractory is an open-source tool for creating custom deployments.
  • Users can manage tasks and issues through a user-friendly interface.
  • Customization options are available for branding and task management.
# 23:33

Triage and Issue Management

How does the triage process work in Masteractory?

The triage process involves specialized agents that manage different statuses of tasks, such as intake, planning, and building. Each agent can be customized for specific behaviors and skills.

  • The triage process is managed by specialized agents for each task status.
  • Customization of agent behavior is possible through UI settings.
  • The system allows for tracking and managing issues effectively.
# 35:20

Observability and Performance Metrics

What observability features does Masteractory provide?

Masteractory includes observability features that allow users to track performance and costs associated with their projects. Users can evaluate the efficiency of their processes and make adjustments as needed.

  • Masteractory provides observability to track project performance and costs.
  • Users can evaluate the efficiency of their engineering processes.
  • The system supports continuous improvement through data collection.
# 47:07

Real-time Collaboration and Code Review

How does Masteractory facilitate collaboration and code review?

Masteractory allows for real-time collaboration on code through open pull requests (PRs). Users can review and approve changes during the workshop, demonstrating the platform's capabilities.

  • Real-time collaboration on code is supported through open PRs.
  • Users can approve and ship code changes during the workshop.
  • The platform enhances team efficiency in managing code contributions.

Transcript

0:00 and welcome to another weekly Mastra workshop. My name is Alex Booker. I'll be your host today and I'm joined by Abby Ay Master's co-founder and CTO. What's up, Abby? >> Hey, Alex. Hey, everyone. How you guys doing? >> You're not in your usual location, I don't think. What's happening this week? >> No, me and the co-founders are doing a little offsite in San Diego. so if anyone's been there, it's been really nice weather. Been nice to get out of SF. yeah, it's been fun.

0:34 >> So, as someone from London, like San Francisco, San Diego, are they are they that different? Really? >> They're like on opposite ends of the world, honestly. and no one knows about AI down here, which is pretty nice. >> Well, they think it's going to kill them, but you know, that's another thing. >> Yeah. They they watch the news, but they don't follow hacker news, right? Exactly. They watch CBS. >> Where are you guys tuning in here on Riverside from? Let us know in the chats down below. And I'm curious, have you maybe seen or heard about Masteractory yet? We made a really big splash last week with a launch on Product Hunt. I, you know, you launch on Product Hunt and you never really know how things are going to land. Sometimes it's hit or miss. Sometimes you launch something and OpenAI are launching a new model on the same day. but we came in clutch last week. we ranked number one number one on top of the week and I think the day we went number one I think Open AI where in the mix right along with a few other big hitters who I can't remember who was also on the board that week.

1:41 >> Facebook or >> Facebook with Muse. Yeah. It >> was a competitive day. Yeah. And so here's the landing page for Master of Factory. Abby, maybe you can tell us like the two three minute story behind factory. I am really excited to get into some demos and show people how we are using Master how we're using Mastractory at Mastra. It gets kind of meta when we use Mastra to build Mastractory and now we're using Masteractory to build Mastra. It's really fascinating. So we'll get into some demos in a second, but yeah, let's set the stage a little bit. What's the story, H?

2:19 >> All right. So, you know, I love telling stories about the history and lore of things. and if you have watched previous workshops, you may be you may have seen the journey happen. So, let me fast or let me go back in time to earlier this year and I can share some stuff actually. >> Yeah, let's do it. The funny thing, all of this started with us trying to with us trying to make memory better. So all of this started here with observational memory. Now if you all don't know what observational memory is, it is our memory system. it how it works. well, the inception behind it is can we design agents that never forget?

3:13 And by doing so, that by having that goal, we ended up building observational memory. Observational memory is a memory system that we have that is essentially like modeled after a human human mind. and it scored 95% on Longman Eval, but it also we haven't released these results yet, but are state-of-the-art on some other benchmarks that we'll be talking about in the next couple weeks. But to I'll be honest, we're not researchers here at MSRA. We do re we started doing some research. You know, typically when you're doing research like this, you have benchmarks and you know, you figure out how to test things.

3:53 We actually did the benchmark at the end. What we did instead, because we're builders, what we did instead was we should build something that can test observational memory out. And what that was is master code, right? So, we built master code. If you ever use monster code, it is our coding agent. it's our TUI. everyone at MRO uses it. At first it was it was kind of a struggle to get them to use it because it wasn't as good as claude code at the beginning. but the shining feature was memory and you know we started using it Tyler and I started using master code. We were saying hey I've been in the same session for a month like it was our way to test observational memory. And so we have master code now and that was cool. And >> the big unlock there I think was whereas Claude code and codeex etc would hit a point where they compact the previous context and lose important details.

4:58 Observational memory fundamentally takes a different approach whereby it's not compressing it with lossy compression but it's sort of rewriting the context in such a way that it's as compact as possible and therefore optimized for space. And this all happens asynchronously as you're interacting with the agent. So you don't hit that wall of compaction where this isn't my turn of phrase, but I've heard it referred to as lobomizing your coding agent. >> Yes. >> Yes. 100%. So by testing out monster code, we tested out observational memory and then we felt like we had something that was pretty popular and strong, but we had built master code in a way that could not be shared with all of you. So the next thing we had to do was package that up. And then we built maybe you've been to the agent controller workshops but we built up this thing called agent controller which is a primitive to build applications like cloud code and things like that. This was cool. This took us a lot of time. And then the next natural step is we had given agent controller HTTP APIs. So you can run it through master server. Well, the next logical step was to do something that's, you know, in a web browser or a desktop or whatever. And then things kind of like ballooned from there because at the same time this whole concept of a factory got very popular. I had done some traveling at the time to some customer sites in like North Carolina. met with a customer and then they asked me, "Abby, how do you do SDLC, the software development life cycle? How are you at MSER doing it?" At the time we were just using Masher code and you know having linear issues and things like that but that made us think internally like well what if we could automate the SDLC like people want seems like a very popular topic and then factory was kind of born.

6:58 So yeah, that's the story. >> It's really interesting because I think one of the biggest advantages a product orientated startup has is when they dog food their own product, right? And I think it's through building master code that observational memory just got better and better and it was shaped around real problems, not necessarily research benchmarks. I think as well we improved our agent loop dramatically during that time as well because this is around the same time the agents are running for longer and the loops are getting more complex and then obviously with master code you can use it locally.

7:33 Someone asked how much master code costs by the way it's free open source you just bring your own or sub$3.99 and then you can use any provider basically. but around that same time, the landscape is evolving still because, you know, everyone's using agents on their own computer. Pretty quick and easy. It's intuitive, right? But then when work is lasting hours or maybe days, nobody else can really see what's going on. So, you have this constant like effort to keep people in the loop by keeping your PRs up to date, updating linear, and so on and so forth.

8:07 And then maybe there's just a pause in the loop and someone has to stare, but you're on your lunch break, right? or even if somehow you pushed that to GitHub, they might not really have context about how you got to that point. They'd be starting from scratch. And so these are problems that everyone's starting to experience, I think, especially in bigger monor repositories where making a bunch of work trees locally is expensive. My my I've got a really powerful MacBook, but it's getting hot all the time running local coding agents. By the way, >> there's more and more reasons to put this stuff in the cloud.

8:37 >> Yet master code runs locally. Well, masteractory in a way it's a way to give your whole team access to a powerful coding agent essentially master code and co coordinates and host all that work in one place and actually I think this is a good segue into a demo because the core interface of master factory as we'll now show is essentially a kind of canban board which is rather unique when you compare it to other software factories or whatever they call themselves like devon and factory and so on. I'm keen to hear a bit more about why it's structured that way and how we're using it in Mastro.

9:13 >> Yeah. Let me show off the the factory that we're using for pretty much all our work right now. let me pull it up. Sorry. Where's this window? There it is. Okay. So let me explain how factory works and is designed and all that stuff. So it is one big canband board. you know at first we didn't really want to build a canband board but we started thinking about what is the like the bare level primitive of a factory and we started thinking of like well if a factory revolves around work right and also notice we didn't call MRA factory a mashra software factory and that's a very big distinction because you should be able to do any work digital work let's say through this factory factory. So, and factories revolve around work and usually work has different phases. Work has a life cycle, you know, and for maybe for engineers there is incoming issues. You triage them. Honestly, the way that we've designed this is the Maestro workflow, right? We triage them, we plan, we build, we review, then it's done or things get cancelled, right? like and there are maybe other phases of this stuff but that's really for us how we work right and this is the default for the factory there's also review because you know the context it takes to build something and I'm talking about agent context the context it takes to build something we don't necessarily want that to bias the review of something and so review has the same thing it's incoming PRs to review to done or cancel and these are different agentic contexts between the two different you know boards and you can also think of review especially in the mashra context is this is like essentially making sure invasive work does not get through especially in open source there's so much things like so many PRs people open you want to have make sure that your review is very tight but maybe work is fine because it there are so many new ideas to explore that you know Maybe you want to triage and do them more liberally and then review is like a lot more tight. So that's how we designed it. I can keep going further, Alex, but let's maybe stop for some questions or we can do more of the demo. I don't know.

11:54 >> Yeah, totally. I mean, just for some context, this is and and Abby, you'll correct me or come in here with some more context if you want, but masteractory is fundamentally an open source package and you can just run npm create factory to make your own. One thing I love is this idea that you can give it a brand and we call our deployment of master factory shipyard. And what you're looking at here is our deployment pointed at the master monor repo. So these are real issues. There's Abby on one of the cards. You saw Francis, another person that master who kicked off and is managing that card.

12:25 And if you see the green circle around it, it means an agent is actively working on it. Oh, look. Here's the code for our deployment basically. >> Yeah. I'll I'll walk through this code real quick and we'll get back to the the UI, but it is just a MRA codebase, right? You can see we have a primitive called MRA factory. You pass a lot of stuff in. Whatever you want to pass in, you can, right? We have O which is just using Monster Server O integrations which could be GitHub soon GitLab linear Jira I don't know if it's Jira selfhosted though yeah maybe not but like you know we got some Jira going on >> people are asking about Jira in the chat by the way >> I knew I knew y'all were I knew it come on why are y'all using Jira dude okay anyway but then we have sandboxes so you know Mashra has Mashra provides cloud primitives So obviously in our shipyard, our factory called shipyard, we use our platform sandbox, but you could also use local or Daytona, E2B, whatever you want. Storage, vector search, like these are all masher primitives that have already existed, you know, pubsub if you want to have durable agents and all that stuff. So there's a lot of options here. You could also do custom boards. I don't really want to show that. It's very complicated right now because there's a lot of configuration. But you have this factory and all you do is you call prepare and pass it to MRA like you would pass any if you're a MRA user you already familiar with this part and now you have a MRA server that can do factory like maybe it's a little magical there's a lot of configuration on the factory side but it is just a MRA project and you know you can deploy it to MRA platform you can self-host it hey do whatever you want honestly but it is MRA deployable MRA CLI deployable, you know, and what you end up with is a factory that you can use. So, I'll get back to this and we can, you know, show off what I usually do every day. Now, I don't use Monster Code anymore. but yeah, let's let's we pause and see where we go. There's one thing I wanted to observe here which is that it's like a you know I think when it comes to build if you want to use a software factory if you want to automate your SDLC you've previously had two options right you could use something completely off the shelf like a product a SAS basically where you have like fairly little control like sometimes you can't even pick the model and if you can you can't use a sub or anything like that someone asked and yes you can bring your individual sub with masteractory and they have some options, right?

15:07 The the other thing we saw is that people build it from scratch and that's a huge laborious effort because you've never built a software factory before probably because it's so new and there's a lot of things you have to consider. You'll probably get there a bit faster if you build on a framework like Mastra, but there's still a lot of things to figure out. And so what masteractory gives you is that kind of goldilocks just right thing I would say where you can npm install it have something pretty good out of the box just by configuring as you saw or whatever sandbox provider you want integrations and so on and so forth. I I don't know to what extent this is true right now, but it seems really obvious that because you have the master instance, if you really need to drop down and write custom agents, for example, one idea I like is creating a new agent that can write code and using ACP to have it write the code with a different coding harness, for example.

15:57 So, if you happen to be using claude code or codeex and you like the way it writes code, it works with your setup, you could use that with masteractory potentially in the future. And because it's all code, you you have those options. >> Yes. Let me that's a really good point, Alex. Let me like show off let me show off my workflow and then I'm going to show how it all works like conceptually because you're right. It's it is just agents. It is just work. It is just workflows. so first, you know, when you're getting started with the factory, obviously you want to know what model to use. And so here you can see I'm using Fable 51, but Obby, can you use your sub? Well, Anthropic said they're not banning people from subs anymore, right? So yeah, I am signed in with my Anthropic sub. I'm also signed in with my codec sub and I was using Grock earlier, but I just, you know, deleted it for now. But XAI, we could also support Kimmy anything with OOTH. And I believe we are going to have more OOTH providers here. You can also just connect via API key all the API keys providers you want. You know this is the powers of MRA model router.

17:08 Also it's another cool thing cuz like it's not like we built all this stuff like fresh. We had this in the framework forever. Now we just expand >> just the model router. >> Just the model router. Yeah. and then >> you say router I say router. >> Yeah. We're all it's all love. then there's this distinction right? So I bought a codeex plan for the team and that's the difference between a personal and orwide. You could set your a suborgide. Now maybe this violates terms but you didn't hear it from me. in the deployed version I have a codeex plan for my team because I want everyone to just use it and not have to worry about because you know some people are I'm on an team anthropic I'm on team codeex. I'm like dude who cares? like we're going to have all of them as a team. And so if you're doing organization type work, then it's going to use that. If you want to do personal work through the factory like a user session, just, you know, use whatever sub you want. And so you come here, you set up all your API keys, you set your model, set your thinking levels.

18:13 Personally, I like to use thinking level low so I can use my own brain, but yeah, you guys, you know, do whatever you want to do. and then after that, you can also do memory settings. So, I remember I told y'all observational memory is like the heart of the factory. So, there's like ways to set this up. I'm using Haiku 45. also there's a personal versus a factorywide setting because, you know, I'll show you in a second. I could do personal sessions as well.

18:40 And then let's get into some work. So, we just and here's funny enough, we just started this new policy at MRA, which is a zero bug policy. We also started a policy because we have the factory now that we don't really accept PRs really. I mean, we accept them, but it's not encouraged. We think good contributors make great issues. And honestly, ever since we said that, y'all have been making some great issues. Oh, man. And look, there's 225 of them. and how you know to get to a zero bug policy with building new features and fixing issues. It's you're gonna need a you're gonna need an army for that or you might just need a factory.

19:24 So let's look at let's actually you know let's fix some bugs together right now. So like cool thing is I could either make this into yolo mode which is this top left auto start runs. That means any work in the factory autonomously gets done. That means I have no supervision of it. I do not want to turn it on. It's it's honestly the most dangerous thing we built. I don't want to turn it on because, you know, agents are agents.

19:51 But if you were crazy like me on a different project, you could just turn it on. Any issue that came in, it would go through the flow. Also, there's another yellow type of button called auto approved plans because how factory works and we're going to show show it off right now. Let's just look at a couple issues. We display the GitHub issue here. Delegated sub aent suspend answer gets overwritten by the delegation call. I know this is a bug.

20:17 >> This is something that >> someone opened a not a master person just someone opened this issue against >> and that was now we need to fix it. >> This is our task and we're going to use master factory to fix it. >> I think I have I have this luxury that I created the framework that I'm looking at this bug and I'm like a this is a bug. Investigate. I know it's a bug. Like I already like I know Roma opened this one. A dude if Roma opened it.

20:42 Probably got to investigate honestly. skills discovery ignores this. That's definitely a bug. Investigate that. Investigate that composio catalog thing. Oh man, investigate that. Literally that is my life now is just hey investigate. Investigate. And what you'll see now is they've gone into triage and I can open the session. And now you have this >> agent session >> and I can see what's going on. And this is a good segue to tell you how a session works, right? So >> wait, Abby, can I ask you one thing?

21:17 You've started all this from factory. I don't know if you I don't know if it's suitable to demo it. You probably don't want to open Slack, but if you do, that's cool. Like it's not you can use this from Slack and you can use this from where work already happens, which is cool. >> Yes. I could demo it. I let me just >> for me I think the Slack integration is one of the coolest things honestly because there's just so many little things that I might just notice that I want to fix and then I can just delegate it via Slack and then everyone else can see what's going on there and you can steer it and stuff all from your team workspace but then you also obviously have this UI as well. It's the the the integration channel with Slack or whatever is just controlling the same same UI basically.

22:00 Let me show Slack. >> Just answering some questions quick. Yes, you can bring a sub. Oops. Add shipyard. Excuse me. And you can see a session has started and it'll probably say what's up to me back potentially. But, you know, it's the same Devon like experience. you can Hey, I'm giving a workshop right now. Can you say what's up to the fam?

22:37 Say what's up to the fam. No team support right now, Carlos. But it's definitely important. Denny's wants that, too. Thanks for the feedback. What's up, fam? Hope the workshop is a banger. Y'all are in good hands with Obby. Now, I could go give this thing. Let me, you know, give it a Let's just give it a GitHub issue. I can work on this out of band. Let me just find a GitHub issue for this and then we can honestly at the end of Hey, part of the cool thing about this workshop is I'm going to get a lot of work done while giving you guys this workshop because I had to get all these bug fixes done. Let me say, "Hey, bro.

23:17 Can you investigate this issue?" Oop. Sorry, this Riverside screen is in my way. Don't make me look bad. The work shoppers are watching. All right. So, let's leave that here. It's probably going to acknowledge and then I'm going to share the factory screen again. >> Stop sharing this. And back.

23:47 >> Someone asked for Google chat support. Abby, >> Google chat support >> like Google Meet or something, I guess. Oh, hey, open source PR is welcome or maybe issues welcome. so okay, so you can see that these sessions started, right? This one actually is now in plan mode. So from triage, it went to plan mode. But let's go look what happened in GitHub. >> So this is remains issue. We triaged it.

24:20 And so you can see this is When did this happen three minutes ago? Okay, that's not my triage. Maybe >> there was some label platform. >> Yes. Let me double check this. My bad, dudes. Maybe I didn't do the right one. >> Is the ticket tracking system bespoke? Yes, essentially. And by the way, each sort of status in the kamban from intake, triage, planning, building, each is its own dedicated specialized agent with its own system instruction and skills which you can customize. So if you want triage to behave differently, you can just go into the UI settings and see what skill the the triode step is using precisely. And I believe you can customize that as well in code and probably soon in the UI as well.

25:13 I think we have an issue because like it says it's Darrow, but honestly that's fine. it's not really a Darrow, but I think there's probably a GitHub automation going on, but because it's last edited by Darrow, whatever. But you can see the outcome of a triage is what happened. What's the severity? how's our confidence in the issue? Is there effort? Low effort impact medium. And then it actually like if we go back into the triager, I can go into the session here. Where's this composio one?

25:51 Go into the session here. And you can see that it did triage and it started scoping out a plant. Phase one, phase two, phase three, risks, assumptions, open questions, none. And so it can transition to work if we want to. And you can see here it's like stuck in this plan queue because I have to approve it, right? I didn't approve auto yolo plans. But because it's so simple, I'm just going to hit build. And it's then the card goes to the next phase of building.

26:27 And honestly, let's hit build on all these. Screw it. And >> let's see how many we can fix in this workshop. >> Let's accept this one. It would be really cool if we can merge something into main. >> Oh, for sure. We're going to do that, dude. but this is also a good segue into how do these phases work? They're all based around skills. So, skills are a very important part of the factory. So, we have skills. So when oops when an agent when a work item goes into a specific phase like triage right this is the skill that gets run and you can edit these skills but we have this triage skill this is teaching the mantra agent this is how you do triage then we have planning this is how we expect plans to go and then you know within the review side which we'll hopefully get to shortly you have re review and then there's al also a skill for re-review because sometimes these the work goes back and forth, right? and we might see that happen with some of these as they >> oh the idea that it's it's only it's kind of making sure that the feedback was specifically addressed, right? And probably >> correct >> looking at the death more so than the whole thing all from scratch.

27:53 >> Let's build this. And we can see now this thing is hustling. It It's hustling. It's going to come in, open a PR. This was quite an easy fix, but that's cool. I didn't have to do anything. I will have to review it later, but we'll get there. >> >> I shouldn't indulge too much. >> I shouldn't indulge too much, but it is very cool to see, you know, from the inside. Look, from the inside, it's like, well, what you just saw there was the task list. That was something we shipped. when you steer it, that's using signals. That's something we shipped and it's like this whole and then it uses master code under the hood and like factory is just the tip of the iceberg.

28:31 It's been like quite a journey to to get there >> and it's building on the latest stuff, right? Like as the industry's moved, master has moved with it. Factory kind of represents some of the most powerful things we can do right now using agents. just very cool to see. Yeah. >> Yeah. let's see. So we have three in building. >> we can take some quick questions, Abby. >> Yeah, sure. >> Someone asked if the skills are Do you have to edit the skills in code? Asked Trent.

29:00 >> You can edit the skills in code or I haven't added this yet, but you'll be able to edit it in the settings here. You can just >> How well does Masteractory work across a whole org if you have many repositories, for example? >> Yeah, that's a good question. And let me show you a different view of this. So we have two we're trying out two different approaches. One is you have a different factory for your different squads. So platform right here you can see they have their own and you know our platform team or mon you know people who work on monster platform they don't really use GitHub issues right they're linear people they could be Jira people too. I don't want to leave out those Jira people, but they predominantly use linear and Jira or whatever. And so their factory looks a lot different than the mess that I have, right? Because we have an open source project. That's wild. and so they're using it here. But also you could have a separate factory for open source, right?

30:07 There's also here, let me go to and we're also testing right now like a durable agent version versus a non-durable agent one. So what I just showed you was like a deployment that is all durable agents and then now this this canband that you're looking at right now is non dur like non-durable agents and you can see in the left here we have different factories for platform Mashra Mashra website we're kind of coalesing to maybe you shouldn't have multiple factories maybe one factory should handle everything and you need to have a smart routing that's going on there >> you You can see now this may be the platform factory but it has all the stuff in it like a bunch of stuff from all over.

30:55 >> >> there's a question from Sabah here asking where do the engineering standards for your project live and how do you get agents to adhere to them? The reason I asked that question now is because you might question >> want different standards for different repos or projects potentially. So I happen to be working in the open source repo in our journey today and we have some things to review it looks like but there's two places where the engineering standards come from one the repo I'm working on itself now MRA has invested in agents MD heavily all across the codebase if we didn't have it dude like probably a lot of this work would be trash because agents are not that great at at just doing things right Like ideally factory should be a teammate of mine, you know what I mean?

31:45 Like it should know what's going on of mantra. That's why if it's just going to AI slop cannon some then why' I build this in the first place? You know what I mean? So you have to work that in your own project. Like you need to have skills and agents MD ready to go. A lot of skills that are in the monster repo get used as well. And then our factory guided skills just make sure that the work is done in a certain manner. Right?

32:11 So kind of think of it like having a junior engineer on your team. You need to give them the tools to be successful. You also need to give them the operating procedure to be successful. I hope that makes sense. >> Where do the agents run their code? >> Great question. So they run the code within a sandbox. you could use any sandbox provider that Monster supports. right now we're using platform sandbox which is backed by E2B. but you could build a Daytona or Railway sandbox just released yesterday, Docker sandbox like I don't know pick your sandbox and you're off to the races.

32:56 >> Is it possible to do multimodel I guess like an adversarial reviews like Code Rabbit and Claude Code review support? >> Yes. So I that's a really good thing and I have a PR for this where you can configure each phase to be its own configuration. So right now you can configure the skill on a phase. in our next version you'll be able to configure the model per phase or the agent per phase if you want to do your own thing or if you want to do ACP cloud code or whatever perph phase. Right now it is a factorywide phase. We thought that was the right thing to do. We didn't think that was a big deal. And then now we're realizing maybe we do want to have more flexibility in both agent and model between your different boards. Now you can do custom boards, custom phases, custom swim lanes. that's where things get a lot more configurable and complicated, but you should be able to configure everything. but in our defaults, you can't configure So my bad.

34:01 >> Is there visibility into the token and task like overall token spend and maybe spend per task in factory? >> Yes. So this is the factory I would say like the factory analytics where you can see what's going on. this is more about the phases and you can see how things have been going. See latest commits activity and we actually are tracking all the token spend not that you know this is audit log.

34:38 I think we have it somewhere >> little chime that keeps going off in the background. >> I mean master factory finished something by the way. >> Correct. Yeah it finished some things for me. you can you know I'll get to the chimes in a second. we had it on this page, but maybe we don't anymore. but this is where you would go to see your token usage across all your sessions per session. some of y'all might be using API key. So, this thing gets more expensive, you know, if you're not using a subsidy from the model lab. so you definitely will need to know I mean just in in normal factories, you need to know the cost of production. And same with this, right?

35:22 Just a question of my own really. since this is just a master project, I suppose you can enable master observability, right? And then through the traces you could see, >> correct? >> through different filters who's doing who's using what and what tasks are doing. >> That is a great great great call out, Alex, because shipyard is a mosher project and it does have mosher observability and you can see what shipyard's been doing and it cost us only 214 bucks. I probably closed a 100 issues in the last three days for 200 bucks. That's pretty good, dude. Oh, that's the last 24 hours, but like not bad.

35:57 >> A little bit. >> Not $14. Damn, dude. We be spending some cash. but that's okay. >> the other cool thing about this is that you can see eval coming into the mix. I think scales should have evals basically. And if each stat is a triage, build etc. is a skill. Ideally, you should have eval to make sure they're getting better or worse. You can also like there's a certain >> there's a certain sort of value in like it used to be that your code was like your IP and the thing you're really precious about, but what if engineering a factory is the truly unique part now and like your your kind of unique selling point as an engineering team that makes you faster than your competition? that that involves taking an engineering approach to building your factory which involves testing and making sure there no regressions and aiming for higher scores and evals and things like that. It's quite an interesting idea I think.

36:55 >> Yeah. And it's good good point that all this observability is here. We are running evals on the code agent. We have efficiency and outcomes going like we're collecting a bunch of data so we can use trace intelligence. That's probably for another time for y'all. But I guess I can click on this you know, we are collecting data and like obviously we don't look at this quite quite often because we're just trying to collect as many traces and to figure this out. But you can imagine if your agent's running in a factory and you're using it, you can improve that agent at quite frequently because you're going to get so many incoming things and you can see highly frustrated with sentiment. That's probably me. I I just talked to the agents like like you I just talked to it so frustrating >> dude if if the sentiment was this is whack. I know it was from you.

37:44 >> Yeah. High frustration mixed emotions. Like honestly I think this thing is talking about me. I feel very targeted right now. >> >> that's okay. >> Some Boris asked if we have something like Terraform to sort of set this all up reproducibly. >> Yeah. So, nothing like Terraform, but you can just do mpm create factory and you'll go through a wizard and you'll just get what I have. Like there's nothing like honestly what I the code I showed you for shipyard is the default factory template. I didn't change anything with it.

38:18 >> So, so yeah, you can all get one today for basically >> you don't need npm create factory and you're off to the races. Now, I can understand if you want to change the off and all that stuff like you know, go for it. Maybe you do want something more reproducible on your end, but it's just a monster project, too. It's not like you need to upgrade the template. You just upgrade the version of the package and all that. I know I want it, you know, not to to cut that short, but we got to review our work. I'm trying to get some PRs done, but by the time this this workshop's done. So, there are pull requests opened. Let's go look at them.

38:56 This is the Composio one that we or no this not. This is like the one that Roma did. Pass correct step number durable agent processor hooks. >> Maybe you're in a separate window or something. I see. Oh yeah, I see it now. >> My bad. >> so it did this. >> Oh, so this PR was open by factory. Got it. >> Correct. And that's cool. And all the code rabbit stuff is probably going to under underway whatever.

39:22 And usually what >> I do want to ask you an I ask you a question about that like so you're saying that masteractory has its own kind of code review options basically kind of like code rabbit but then you're also running code rabbit on the same PR what's that about and how do they work together >> so factory reviews the PR as if it was a engineer on the team that's how like the review skill works so let me tell you a little thing about code review like We want to you this to work with your code review tools. While code rabbit is our choice, you can work with other ones, but code rabbit and stuff try to do the intent of your codebase and they fix showstopping issues. The way we designed the review skill is to look at the history of the files changed, look at the intent and behavior of it and then give a review. That's because we are building something new that will help this which is called knowledge. and I'll show that later. But you know what we found is when you start doing a bunch of factory and you just yolo it like I've just been doing, you'll get a lot of slot PRs. Even if even if you have the good agents MD and all that, you will get a lot of slot PRs if it's something complex in your codebase. Now, this issue looking at it, right? It wrote a test and it just fixed us. We like made some mistake here with the index of something, right?

40:54 that's an easy fix any LLM should be able to do. But if I wanted to go like change the world for me, it's going to create a bunch of slop. That's because it doesn't really know the intent behind my ask. It doesn't know why we should do this. It doesn't know the history of what how the decisions were made which is why we like this is cool but it's not good enough. So memory will get better in our next kind of thing called knowledge which allows the agent to learn things over time and use that later and I can show a little bit about that in in a bit. so you can see this is probably going to be an easy PR review. Let's go back here. Let's just start reviewing stuff. You know, maybe I'll review some homies PRs like Damian's. Here you go, Damian. You get a review.

41:42 >> Why' you have to I was kind of curious when you press investigate in the first instance. So, now you're pressing review like what's the argument against just doing that automatically then you can sort of jump in whenever and not have to wait. >> So, we do have the yolo mode like auto start which would start reviewing things. >> Okay. >> here's the reason why you don't want to do it. one, if it cost $214 in the last 24 hours, maybe not everything should be reviewed, right? Like this is in draft, maybe it shouldn't be reviewed. there might be issu PRs from people that, especially in open source, some people have no business opening a pull request. I'm not going to review it. You know what I mean? I'm not going to spend money on it, right?

42:23 Tokens. >> that's the main reason honestly against YOLO mode. But if you're a small team and it's all private and you trust everybody, like yeah, go ahead. Why not? >> That's a good distinction cuz I I kind of interpreted YOLO mode as like, yo, it's going to merge into main. but it's more doesn't do that. >> Yeah, that's so Yeah. Yeah. Yeah. >> Also, Peter says they're using Code Rabbit, Grapile, Cubit, Codeex Review, Devon Review, and a bunch of others to review their PRs. That's that's a big bill right there. maybe you just need two or three.

42:58 They're all probably talking to each other, too. >> That's funny. I thought it was you. >> No, >> I mean the agents they're here.

43:29 >> Okay. Anyway, that's weird. wasn't me. so yeah, you can see this is a review session just going off, you know, all CI passes, code rabbit completed with no actionable things. It's going to it's just going to give me a review. I can go back to the work item that started this thing. And you can see here that's the work item. I can go back to the review. These are two different agent contexts. and yeah, like maybe there's more to review.

44:08 All of Toby's stuff. Let's review that, too. What other questions do people have? >> Yeah, just looking through the chat right now actually. There's a lot in the middle here somewhere. Just catching up with the Fred. I mean, it's a this is a good operational question from IMAD and very well suited to you as a CTO. How do you view the place for software factories alongside individual engineers or teams?

44:45 Does everyone in the master team work on features via the factory or do they still do work in their own coding agents like you know claude code, codeex, master code? >> Yeah, so we still are we're figuring out where the factory is for net new work. Ideally, we want it to be done through the factory. and let me tell you why. ever you know our team we've been working together for a while even before MSRA we've been working together and in the past we used to pair program we used to get on zoom calls and spend time with each other building things and the team collectively is missing that that that feeling you know working together instead of working with agents. So for that reason alone, we want to build pair planning into the factory. Want to build collaborative sessions into the factory.

45:34 You can work on features with your homies into the factory because we really miss that ourselves. but as of right now before we're venture down that path, people are still using master code to build new features. And this is really just the bug squad, right? Like fixing bugs. And our next phase of work is going to make sure you can build net new features through the factory. but yeah, that's kind of where things land. >> It's interesting cuz I was at the cursor London event yesterday. Got this nice little mechanical keyboard key ring.

46:09 >> It was really fun talking to everyone at the Cursor event about claude code and codecs. but eventually people started to wonder about, >> you know, they were asking like, oh, how much do you use cloud agents and all this? And there are some things where doing it locally I feel like makes a bit more sense. Like if you're maybe iterating on a front end locally then you can just quickly open Chrome DevTools or something go back and forth really fast. I I kind of and I'm sure there are other examples like that. I'm not sure if the cloud is the right place for every task, but there are there are a lot of things where setting up a work tree, figuring out the system is just superfluous when you want to fix something small or you already have a spec in the case of a GitHub issue and you want the factory to genuinely take it end to end hopefully not needing as much iteration because it's not creative or exploratory. It's kind of obvious what correct is because of the because of the spec. I'm still kind of figuring out that balance to be honest.

47:01 >> Yeah. And I think the next couple months at least for the master team is going to be quite involved of like where do these things break down? Can it really automate a lot of work? All the work entirely like it'll be quite interesting. >> How many engineers at Master these days? >> We have about like let's say 15ish or more. But let's just use that number. Yeah, we've got about 10 minutes left.

47:33 what else do you want to show, Abby? We'll we'll take a few more questions, but I want to make sure there's time to Dude, I don't know if we can, but knowledge is people have I don't know how much you want to show. It's a work in progress, but it's very very exciting. >> Let me see if it's enabled on any of our factories. >> And if not, it's cool because we'll do another workshop on it.

47:56 >> Oh, it is. It is. Okay. Couple things I want to do before we get into there. Like I told you, I'm trying to ship PRs in this chat in this workshop. You can see we have a bunch of PRs open by the factory. This is 46 open. That's awesome. This is that Composeio we like started looking at. Let's see what's going on with it. What did the Okay, there's some code rabbit stuff, but I got it approved. Hell yeah. Like and you know what? because I already know and I planned this from the beginning and I already looked at this issue and a lot of this is staged.

48:32 I'm just gonna approve this because I already know it's good and we're going to ship. We're about to ship right now. Bam. Ship that. Cool. And there's more to ship I'm sure from what we started. But you can see like all the PRs that have like in our tenure here today. Where is this? You know that this is another one that we did. And I guess we could take a look. I really want to see one where it did not approve because that's where it gets interesting. maybe we'll wait for this one's review >> where >> I will say at this I will say at this juncture that obviously on a workshop like this you want to merge stuff.

49:16 and so there's like an incentive to go a bit fast. but there is of course still that responsibility to look at it, review it. this doesn't necessarily replace that. and that still takes a little bit of time, but at least you're spending time on that part, not everything that came before it, which is where you would have otherwise spent hours. >> Yes. All right. So, let's transition to knowledge here. Let me find my tab for that.

49:43 >> Sebastian says that Composia fix actually helps us. Cool. >> Dope. Thanks. okay. So, I'm gonna share this version of the shipyard and let's go to Mustra here. You can see there's a special tab >> called knowledge and it's going to have to load. Let's see. This is still very much a work in progress, y'all. >> You said something. I think one of your sharpest lines in recent months, I think you might have used it at the work OS talk. It's like you you want your factory to feel like a teammate, not just like a cold start random agent everyone else has. You want it to feel like a teammate.

50:21 >> What what makes an agent feel like a teammate and how does knowledge help that? >> Well, live demo suck one. I'll put this in the background. okay, this may not work y'all, but it will eventually. what makes a great teammate? one at least in our opinion is someone who is highly capable at their job, dependable at the job and delivers. Oh, nice. I load it. but we're hiring highly skilled labor, right? You know, we want to work with highly skilled people. We want to be highly skilled ourselves. And when we think about an agent that is part of our team, it has to be a highly skilled, highly capable, and highly aware person on the team. Knowledge is the bridging the gap between procedure, right? Like skills are procedures. Running workflows are procedures. Like any lowkilled labor can read a manual and do their job. That's not what we're trying to get. We're trying to have dynamic teammates, right? So you could have skills and agents MD, that's cool. You can have memory, that's cool. But there's something that's the gap between being highly capable and what we think that is is the knowledge of what has gone on at the company, right? the intent and behavior of your codebase and what other people have been doing because it's not just what the agent's doing. If other people have been doing other agents, other people have also been doing stuff. So knowledge is you can see there's just a bunch of nodes in this graph because a lot of stuff has happened and you can see that this is a bunch of knowledge that happened about the ATA v1 JSON RPC endpoint right that way and like all this different stuff that has happened in A toa you know and honestly who cares about A to A but obviously it's something that we support I don't necessarily know everything about AAA all the time you know but now my collection of agents, our collection of agents does. so that's just one example like it's so how this works is as agents are working on things, it's collecting subconscious thoughts, right?

52:32 These are things that are happening in the background like you know, hey there's a thing okay cool let me probably let me like note that down you know and that's really gets represented in this knowledge graph. So then as things the next agent is working in the same area it can use a tool to query the knowledge graph and then ultimately get context of it and then move on with their life and actually work on something right so that's just I don't know I'm very excited I could talk about this we have a workshop on this like by itself like it's so exciting but u yeah >> I was watching a AIE talk by someone at Uber talking about how they built their software factory I forget forget the name of it. they were talking about all their challenges and one of them was that basically the agents were just so slow because they would have to go into a monor repo figure out where the service lives figure out like this that the other but with knowledge anytime it learns something like where the where the service lives how to set it up etc.

53:32 It can just take a shortcut right there. And the other thing is like decisions and decision history. Like no, even better like if you give an agent feedback when you're steering it. Imagine working with a teammate where you give them feedback and every time you give them feedback they keep making the same mistake over again. >> Exactly. >> That's not someone I want to work with. >> You know, a good teammate I don't have to repeat myself more than once. Right.

53:54 So, one last thing on knowledge for yall. Just want to give you some lore. I love doing the lore thing. knowledge was built in the same way we did monster code in a way where we dog fooded some and then we were like man this thing really works. We built this repo called Alexandria. It is a GitHub essentially bunch of files in here. Essentially, Alexandria is listening to all these inputs like Fireflies and Slack and, you know, GitHub pull requests and stuff and it's building a knowledge graph. But this is more of like a knowledge folder structure with areas and you know there's some personal information here. I don't I can't really show you all of it. But what's cool is as Monster Code starts working as everyone you can see people use Monster Code every day. anytime the agent doesn't know something, it files what's called a knowledge gap.

54:48 >> and so this was like the precursor to what I just showed you, what I just showed you here because while Alexandria was a Mashra specific thing, we were testing and doing everything the knowledge primitive is going to be open source and available for all of you. and you know this obviously what I'm showing you is just looking at GitHub, but there's more to knowledge than just GitHub, right? There's Slack channels that the agent probably should know about. There's Fireflies meeting recordings. There's things that are happening in Jira comments, etc. Those all need to go into the knowledge base.

55:24 >> Yeah. >> Sick. And yeah, this is also lending way to a primitive just like you can use OM in any master project. Soon we'll be giving you a knowledge primitive that means you can use this in any project. And of course, Master Factory benefits from it as well. I'm sure we'll do a workshop in due course, so stay tuned for that. >> Yep. And I really want to ship one more pull request before we go. I'm just trying to look which one can I ship, you know? I guess >> you're showing >> maybe.

55:58 >> Oh, yeah. Yeah. Yeah. >> I got another approved, but CI is failing. I can't ship this. But you can tell I can tell that or y'all can tell I really want to. But >> by the way, sorry to interject, but that's the hardest thing about contributing to a monor repo. Sometimes you make a change and then the check fails and then you have to like >> pretty inefficient workflow, but you go to your agent, you say, "Hey, the check's failing. Go and fix it." And it still fails. You go back and forth.

56:24 Maybe if you're a bit smarter than me, you use some kind of loop or subscription or something. >> Yeah. >> Does factory just kind of tenaciously keep going to get it to work or do you have to prompt it? so if if there was if there was a back and forth, actually let's let's test this out right now. the review cycle is done, but yo, bro, CI is failing and you approved the PR. That's whack. Let's see what it does with that >> because the work session in theory should be able to pick it up. When I checked, everything was passing. Maybe a check ran later. Yeah, cool. Check it again. God, >> h was responding to what I said saying that you don't have that loop in mastertory, but actually you can subscribe to a PR in masteractory and it'll get >> notified. Does exist.

57:21 >> Maybe if we look at the workflow here, you see these notifications that are happening. >> Exactly. >> Yeah. So, it is getting pinged. It's just sometimes, and this is a bug on our end, sometimes like the notification listener happens before the notification arrives and stuff like that. Sometimes it happens on time. So, obviously, this is very beta still, but you know, we're making moves. All right, this thing is on its own journey. I wish I could have shipped other PR, but >> give me a link. Give me a link to the PR. We'll paste it in the chat and then people can follow it if they're interested.

57:57 >> Sure. >> You can post it in the chat actually. >> Yeah. >> And on that bombshell, >> we'll call it a day. I think we're bang on time. We covered a lot of ground. Some really fantastic questions, by the way. We always appreciate them because in some cases it shows us, hey, we need to do a better job explaining these things from the forefront. And in some cases, you're giving us genuine ideas for features and ways to make this better. And that is gold dust. Like that is the thing we value the most at Mastra. So, thank you so much for tuning in, sharing your thoughts and feedback with us. Abby, any closing thoughts?

58:36 >> thank you all for the constant support and everything that we do. Please try out factory npm create factory. Open up great issues and the factory, you know, you might you might see the factory fix them. >> Yeah. See you on See you over on GitHub. It's all bad.

Summary

The Mastra workshop hosted by Alex Booker and co-founder Abby Ay focused on the recent launch of Masteractory, a tool designed to enhance software development workflows through the use of AI agents and a unique memory system. The discussion highlighted the tool's capabilities, including its integration with various coding agents, the use of a Kanban board for task management, and the introduction of a knowledge graph to improve agent performance and collaboration.

- Masteractory launched successfully on Product Hunt, ranking number one amidst strong competition.
- The tool is built around the concept of "observational memory," allowing agents to retain context without losing important details.
- Masteractory features a Kanban board for managing tasks through different phases: triage, planning, building, and reviewing.
- The system allows for integration with various coding agents and supports multiple repositories within an organization.
- Knowledge graphs are being developed to help agents learn from past interactions and improve their performance over time.
- The workshop emphasized the importance of dogfooding, where the team uses their own product to refine its features and capabilities.
- Participants were encouraged to contribute to the open-source project and provide feedback for future improvements.

Questions Answered

What is the context of today's workshop?

The workshop is hosted by Alex Booker with co-founder Abby Ay, discussing their recent successful launch of Masteractory on Product Hunt, where they ranked number one.

What is Masteractory and how can it be utilized?

Masteractory is an open-source package that allows users to create their own deployments using npm. The hosts discuss its features, including the ability to customize branding and manage tasks through a user interface.

How does the triage process work in Masteractory?

The triage process involves specialized agents that manage different statuses of tasks, such as intake, planning, and building. Each agent can be customized for specific behaviors and skills.

What observability features does Masteractory provide?

Masteractory includes observability features that allow users to track performance and costs associated with their projects. Users can evaluate the efficiency of their processes and make adjustments as needed.

How does Masteractory facilitate collaboration and code review?

Masteractory allows for real-time collaboration on code through open pull requests (PRs). Users can review and approve changes during the workshop, demonstrating the platform's capabilities.

© transcribe · For agents Built with care and craft by Gokul Rajaram