Section Insights
Introduction and Weather Check
What are the hosts' locations and current weather conditions?
The hosts are located in San Diego, California, and near Charlotte, North Carolina. San Diego is experiencing unusual humidity and heat, while Charlotte is cloudy but cooler than previous days.
- Hosts engage the audience by asking about their locations and favorite cities.
- San Diego's weather has been unusually humid, contrary to its typical dry climate.
- The hosts share personal experiences about adapting to new living conditions.
Overview of Sherlock and Its Capabilities
What is Sherlock and how does it function?
Sherlock is a Slack-native operations health assistant that integrates various telemetry data sources to provide context for operations tasks.
- Sherlock combines data from multiple platforms for enhanced operational insights.
- It utilizes an agent-first approach to streamline deployment and database setup.
- The integration of tools like Axiom and Clickhouse enhances its functionality.
Sub Agents and Their Efficiency
When should sub agents be used, and what are the potential pitfalls?
Sub agents can enhance task efficiency but may also slow down processes if not implemented correctly. They can lead to increased costs if they require recontextualization.
- Sub agents should not be used indiscriminately; careful consideration is needed.
- Recontextualization can lead to duplicated efforts and increased token usage.
- It's important to delegate responsibilities appropriately to avoid inefficiencies.
Model Optimization and Workflow Improvements
How have the hosts optimized their models for better performance?
The hosts have refined their workflows by utilizing specific models that enhance task performance and reduce complexity, focusing on repeatable processes.
- Optimization of models has led to improved performance and cost efficiency.
- Avoiding overcomplication in workflows is crucial for effective task execution.
- The introduction of GLM 5.3 Flash has significantly improved their operations.
Utilizing Memory in Agent Management
How is memory utilized in managing different agents?
Memory is primarily used for thread management, allowing agents to recall previous interactions without repeatedly accessing external APIs.
- Effective memory management enhances the efficiency of agent interactions.
- There is potential for expanding memory use across different threads.
- Persisting corrections back into the system can improve the accuracy of code review agents.
Transcript
0:00 Heat. Heat. N.
0:43 Hey, hey, hey. Hey everyone. Hope you're all having a good day, a good night, a good morning, wherever you're calling in from. For me, it's it's bright in the morning. 9:00 am from San Diego, California.
1:21 >> throw in the chat where you're calling in from. And another city question for you. What's your favorite city you've ever been to? You can give the same answer twice, >> but it's frowned upon. We want two different answers if possible. happy to be here with Tyler today. Where are you calling him from, Tyler? >> Yeah, I'm in just outside of Charlotte, North Carolina. >> Awesome. And what what are we looking at for weather there today?
1:48 >> it's a little cloudy, but not as hot as it's been. It's been brutally hot and humid. So, I'm excited to maybe get outside a little bit after this. >> Yeah. >> Been cooped up inside for too long. >> Yeah, we had in San Diego there's been all these tropical storms kind of right off the coast. And so, you know, Southern California is notorious for being like a dry nice warm. It's been insanely humid and I just moved from New York. So, I was literally trying to escape the classic like northeast humidity, but it came to follow me. But, yeah, it's been so hot. I moved into a new apartment, which by the way, if you notice, I have a I have a new setup.
2:29 I don't know if we have any people that have come week by week, but changed up the setup a little bit. and it it was so extremely hot last week. I got a I got an AC for my room and then I two days later there were maybe like a hundred ACs at Home Depot. Two days later I was like I can't do this. The living room's too hot too. I go back to Home Depot, no AC's left, no fans left. It was crazy. People are I haven't lived here very long, but people are saying it was like the hottest they had seen San Diego.
3:03 Man, that's it. Still blows my mind that there are places without just AC like as a standard thing because every house here it's just you have to have it. >> Yeah. And we got people calling in from India and Malaysia. I know those places can get pretty hot and humid as well. So hopefully the weather rains are in same Brazil. It's something it's something all over the world. I also have had bad luck. I was in I was in France for a an AI conference like two months ago and it was it was something like the hottest day in Paris's like recorded history. It was like 108 and they don't they really don't have AC out there. so that was a that was a tough trip.
3:51 >> But yeah, we'll just wait another minute or so. Oh no, my camera. We'll just wait another minute or so as some people are still coming in. and then we'll get right into it. Have very exciting session today. Pardon my little camera fix here. Part of the new apartment move in is a new tripod and so far not happy with it. >> Oh man, >> not not too stable. >> Get a new desk set up. >> New desk set up. Yeah, I got a new desk.
4:20 an extra monitor. Dealing with three monitor setup. I feel very blessed. Brandon, should I go ahead and bring up the the intro slide? >> Yeah, feel Yeah. Yeah. Feel free to share screen and then in just another 15 20 seconds, we'll get started. >> Cool. Cool. All right. I think we're good to get going here. thank you everyone for joining. Super super excited for the session today. I'm here with Tyler from Prisma and he's going to be chatting about how Prisma built an AI S sur using MRA. very exciting. We're going to go into sub agents, open weight models, a lot of different things and how it's all done with MSRA and how it how Prisma's been using that to sort of really accelerate their development there. just a few housekeeping things. Of course, as you know, always feel free to send questions in the chat.
5:22 We'll try to get to them throughout the session as they become relevant, but we'll also have some time at the end to sort of go through a lot of the lingering questions. So, don't be shy in there. Feel free to ask whatever, whether it's Prisma related, mollated, or anything. And with that in mind, I'll pass it on to Tyler. >> Yeah. Thanks, Brendan. And yeah, definitely jump in with questions. I would love nothing more than to dive deep into some of these things if people have topics they're interested in.
5:50 so I'm going to talk a little bit about our agent called Sherlock. we have a few agents running in Maestra. And this one is is really interesting. It's, as Brandon said, our SR agent. but really what I want to focus on is how we use sub aents and open weight models to power a lot of this. So there's a quick demo in a video I recorded because it can be a little bit slow. So it's nice to have it fast forward. Let me try to get Oh, that playing. There we go.
6:22 I think my headphones are going to do a weird thing where it pauses the video. This is annoying. so I'll be quiet for a second. So what it's doing here I've prompted our Sherlock agent to generate a report performance report of our click house cluster. I've asked it specifically I want things that are contributing high CPU, high memory, any errors that are happening, systemic errors. I want to chart the CPU and memory week over week and then generate all this into a report in HTML that I can go look at with the team. and I've done this in a DM, but we also have Sherlock available in different channels that we can ask it and kind of have a multiplayer story there. We'll talk a bit about that. and as we go here, I'm just going to slide forward because the video keeps pausing on me as I as I speak. My Airbuds are doing weird stuff, but you can see it's like running all these tool calls to run Click House queries. working working. I've fast forwarded this because it takes about 10 minutes for it to generate this whole report. And then here we go. It finally starts to to produce some output for us and summarize itself. But what's interesting is it drops this link for us here to the full report.
7:46 And then I get this really nice HTML page that charts the memory and the CPU over the week week over week. Gives me its analysis of it. the top services that are contributing the most. So, I'm not sitting here waiting for all this to generate. I just throw this off to Sherlock and I go do my own thing, maybe do some work, maybe catch up on Slack messages. I get a ping back from Sherlock that says it's done. Here's the link. I can share this with my team.
8:11 yeah, it's super great. And then we have Monster Studio here. So, this is that same session that I did represented in Monster Studio. So, you can see the trace view of it. I go in here all the time to sort of look at what all went into a single session. Kind of understand if I'm especially I'm deploying a new workflow or a new agent like what was the token usage of this? How many times did it go back and forth with tool calls? Did it take advantage of parallel tool calls? Well, these are all things you can get really deeply into with MRA Studio. But you can see here as it's scrolling down just how many tool calls it did to to service that. it went really deep exploring all you all these different click house stats. It would discover something and then maybe pull that thread a little bit further. Things that would have taken me, you know, maybe like an hour, maybe two to actually dig through all of these different directions and put this report together, it did in like 10 minutes while I was off doing other work. So that's what Sherlock does for us.
9:10 >> Yeah. Yeah, and just one thing I want to point out from this demo already that I love and that is a great thing to show off with MRA is different all these different sort of layers of insight and observability you have obviously you're digging into the studio right now but even in Slack you know that constant stream of letting you know what tool calls you're doing and all this stuff like you said you can kind of >> make it a set it and forget it kind of thing and just come back later but if you want updates on what's going on you can customize all of this different output that's coming and just like sort of any other agent stream, let the user know what's actually going on under the hood. So rather than having something that's just sitting there, just like still working, still working, you get a good idea of where it actually is in the process.
9:55 >> Yeah, I I really do love that there's a couple options on the chat adapter to to tune like exactly what type of feedback you, but I'm pretty partial to this one where you can the tool calls that are coming through, but it's not constantly pinging you either. But it's nice that you can tune that per agent to sort of set up what is the type of feedback I want. So yeah, I'm Tyler Benfield. I am a platform architect at Prisma. I'll talk a little about Prisma in a second.
10:22 you can find me at some of these links or the QR code here will take you to to the same links. Prisma are our let me just jump ahead to this slide because it's a little better. we're most of the TypeScript ecosystem for our Prisma OM which is pretty popular. we also have in the open source community Prisma migrate and studio. that's that's the thing that we're we're you sort of most recognized for when we go out to to events and stuff like that. We also have our Prisma cloud products. So Prisma compute, Prisma Postgres and now Prisma storage. so we're really working towards a holistic cloud offering that is really built for the modern type of agent deployments. so that Prisma compute is Typescript applications running bun. Postgrez is a serverless Postgres that doesn't have the cold start penalties. And what we found is that offering all of these things under one umbrella is really great for agents because you don't need to sign up for a lot of different accounts and try to integrate them together. All these services like play nicely together out of the box and you set up one account with your agent and it can just go crazy at doing deployments and setting up your database, you know, wiring in your S3 storage, all of that under one roof. And we've also taken that philosophy back into the open source products and really focus on an agent first mindset there.
11:35 so reach out to me if you want to know any more about that. I'd be happy to chat about it. >> So today we're >> also just oh sorry just quick note. Yeah and MRA and Prisma work great together. So if you ever want to build some projects combining them definitely definitely should get kicking on that. >> Yeah you can deploy MRA into Prisma compute as well. It definitely works really well. All right. So today's plan, we're going to talk about what Sherlock is. I'll give a little more introduction beyond that that demo I showed. we're going to talk about the architecture really briefly and what all it has access to to get its job done. context heavy task.
12:11 This is where we get into the meat of what I want to share which is how we handle these really context heavy tasks with sub agents and how we use openweight models. All right. So what is Sherlock? It's really a Slack native operations health assistant. It is our S sur agent. So it has access to all of our telemetry across multiple different platforms. So Axiom has our open telemetry traces and application logs. Click house has all of our application metrics and our the its own self-reported health. And then we have Ignite which is our internal knowledge repo. It's really just a GitHub repo that we dump a bunch of markdown into. But by giving that to Sherlock as a workspace, it actually gains a lot of product context, access to our runbooks, example queries that can run in these other systems and it gives us a place to go contribute back anytime something in the product changes. We can go document and I ignite and all of a sudden Sherlock has a whole new set of context available to it.
13:05 And it's all available through Slack in our using the Monster chat adapter. So it watches for messages in our production incidents channel. You can DM it like I did in the the demo and ask it a question. You can also, this is like my favorite part, you can at mention it in one of our channels that it's in. and we actually have just one channel set up specifically for it. So, if we want to have this like multiplayer concept where multiple people can kind of see the same thread and ask follow-up questions to the agent, you can just edit, ask it a question, it'll come back and reply and then you can keep having follow-up dialogue with it. You can keep tweaking reports. You can keep asking it to dive deeper and everybody can kind of see and contribute to that conversation. And then obviously, it's available in studio and through the MRA API as well.
13:47 Yeah, and on that, Tyler, I think the multiplayer thing you mentioned is such a big unlock that we see with using MRA. And it's funny because there are so many people building an AI where they're really building in isolation, right? Like they just have projects on their local computer, maybe if it's sort of like a less techforward company, they're doing something like a cloud project or just using chat GBT. And what happens with something like that is it's as simple as if you have a 100 people on your team, every single one of those 100 people has to come up with the ideas that will help like automate their workflows and increase their productivity. But if you have something like you can get an agent into Slack or you know we have a bunch of different channels from Monster that you can put agents into. There's also our noode agent builder that can be shared. All it takes is one person to discover a workflow that would help a lot of people and then suddenly you can have all this knowledge sharing and also just give other people ideas too. So we really find that when you take this step of taking something from in isolation to sharing it with everyone, it really compounds everything you're able to build from there.
14:52 >> Oh yeah, I believe that. I mean that even that report that I showed earlier that we generated in the demo that started for me as something I would run on a cloud code routine within my cloud desktop. but the problem with that is every time we generate this report I have to go then publish that to some link and try to broadcast that to my team and I'm just the one like copy and pasting stuff around at that point and then they can't follow up with it because it was all generated by my agent that's isolated to me. It's fine for like one-off stuff like that, but when you start getting into wanting to share these artifacts with your team, having it just shared inside of something like Slack where everybody already lives is such a a huge unlock for us.
15:30 >> Yeah. And I think another note with that is it it also kind of unblocks everyone, right? So instead of waiting on you necessarily to generate these reports, you can kind of be like, "Hey, Sherlock can do it." So instead of, you know, you needing for wait to wait for me, you can just ask Sherlock. I do the same thing with a lot of my Slack agents. I basically try to think of like what are the times when I'm someone else's bottleneck and how can I just build an agent that's always on to be able to sort of fill those gaps.
15:58 >> Yeah, that's great. So, let's talk a little about the architecture of Sherlock and this will be really brief. so, it it receives a message off of Slack API trigger. we use the chat adapter again for that. Makes it super easy. The hardest part is going into Slack and setting up the bot there with the right permissions and everything. that could use some cleanup on their side, but the actual integrating it into MRA is like very few lines of code. You drop a few things in and you're done.
16:24 it'll start watching the channel. It'll pull its own thread history. All of that's wired up for you. that dispatches the Sherlock agent whenever it gets tagged or gets one of these messages. And then it has access to a few different sets of tools. So Graphfana IRM, that's the incident response tools inside of Graphfana. And I've given it that with re write access so that actually can if it sees something that is incident worthy it can create an incident in graphana and then start capturing its notes there. So rather than like flooding Slack with a bunch of messages of all the research it's doing it'll just be leaving updates on the graphana incident and only come back to Slack when it finds something meaningful that that pushes the conversation forward. But we can always go into Graphfana and see the entire sort of research process. And then again it has ClickKouse. It has the Ignite workspace I mentioned where it has all of our runbooks and product context.
17:12 it has the Axiom query executor sub agent which I want to dive a bit deeper into that then can query Axiom for us. And then it has another tool for artifact publishing which is a a relatively newer one but that's what powered the HTML reports that we saw in the demo. And what what I kind of modeled off of that is I saw how my cloud agent could generate these nice HTML reports and then I could see them and I could move them to public and share them with the team. And I wanted our agents in Maestra to have this same tool. So this is a really really simple generic tool that just gives the agent a way to generate any type of text artifact whether it's an HTML page, a markdown file, SVG, whatever, and it can just publish it to an S3 compatible bucket running on Prisma storage. so then now that is also served behind a MRA API endpoint in our the instance we host. So it uses the same authentication setup. so the these links are accessible to anybody within the team but not to anyone external which is really nice.
18:15 >> Yeah. And really quick on there I'm just sharing a link on the screen. I mean there was a ton of great stuff you shared there. this is a link to master channels which is basically all the different things we provide out of the box to be able to integrate your agent into other thirdparty platforms. So like Tyler said he used this to get into Slack really easily. We also have a bunch for stuff like Discord, iMessage, GitHub, all these different ways for you to super easily just have the out of the out of the box ability to get this agent into other channels. And of course, ton of documentation well on all these different integrations. I know Tyler, you already mentioned you're using all these different things like Clickhouse and obviously a big focus at Monster is being compatible with whatever your infrastructure is. So we have no shortage of different things that you can partner with ways you can really get this into your existing infrastructure.
19:06 >> Yeah. And what's great is because this is all written in Typescript, I the vast majority of this is written by just by going to a coding agent in this repo and saying build me a tool call that wraps click house or build me a tool call that does like the artifact publishing was just an agent built this whole thing out from a prompt. So we we build Sherlock and our other agents with coding agents, not with some type of, you know, tool that we have to go in and click around as humans, right?
19:34 All right. So let's jump into context heavy tasks. This is where I want to start to focus on the things that we've kind of learned by building Sherlock. and what we found is that the the most context heavy task are really good for sub agents when you can slice the context well. So we'll talk a bit about that. Maestro makes this super easy to give agents as sub agents to other agents. essentially it looks like a tool call. So, this is a a small snippet of how we actually configure Sherlock. we create a new agent and we give it all these tools to do these different actions. But then we just also give it this axium query executor sub agent just registered with this like one line of code and that injects it as a tool call sort of available to this parent agent to dispatch work to. So the parent sees the the sub agents description of when it should use it. and it receives whatever that sub aent's final output is whenever it's it's done with its you delegated workflow.
20:33 So we did this specifically for axim because it needs some some boundaries that we found that we can also like kind of extrapolate out to other scenarios too. so one of the things that we found when we did this with a single agent is sometimes because the query language for Axiom is not greatly represented in the training data. So it makes a lot of mistakes if it's not nudged in the right direction. It'll run a lot of accidentally unbounded queries that just produce a ton of data and then overwhelm the context window of the parent. sometimes it would send a query that would have an error in it.
21:03 So, it would have to go through these rounds of like trial and error to get the right query. it could also like because it bloats a context the aggregation quality and the like sort of inference off of the results starts to diminish. and we don't like know what's going to be retained during context compaction, right? Because when you start flooding the parent agent with all this data context, it has to it's going to inevitably have to run compaction at some point. But when we move this to a sub agent, we start it with a fresh context every time. So it doesn't inherit the entire context of the Sherlock parent agent. It's just given a task of here is a hypothesis that we have about this particular incident or this this question and I need you to go query Axiom to prove or disprove this hypothesis and then report back a summary. so it has a really clear and simple objective. and it returns a really simple lightweight result regardless of how many turns it had to make internally to make these different tool calls. Maybe some of them failed, maybe it, you digs down a different path, but it's all towards this objective of finding this like summarized result and reporting it back.
22:06 so we've essentially offloaded all of the research context for that particular hypothesis or question to this sub agent. And then once we get an answer back, we completely discard that context. It's gone. And the parent Sherlock agent just has the summarized answer to keep continuing its work. And it can also run these in parallel too because we've enabled like background tasks inside of the agent. So it can run multiple of these at the same time. So just to kind of reiterate that we start with the hypothesis from Sherlock.
22:34 It it thinks it, you know, has an idea of what might be going on. Sometimes this is informed by a runbook and some past incidents and then it sends that over to the executor to generate the Axiom query and go execute it. Maybe that makes multiple rounds of queries and then that a sub agent distills the evidence and comes back not with rows but with the actual signal as the result. So we also learned a few things about when not to use sub aents for this because it's tempting to just start throwing sub aents at everything. one of the common things we saw is thinking that you can use a sub aent to speed things up. but in reality it can actually slow things down if you're not careful because if it needs to recontextualize with what the parent was trying to do, maybe it needs to make some repetitive tool calls, you end up potentially duplicating a lot of work.
23:22 also sub aents for cost. There was some thought initially of like, oh, if we move it to a sub agent, then we our context now becomes like smaller. We're starting from scratch and now we don't have all this like repeated context and cash token use happening or potentially uncash token use, right? but it actually can generate more cost if you're not careful because again it needs to recontextualize on itself. so if it needs to go back and pull some of that context out of the parent and you're starting without the cash benefits now. So you could actually end up burning more tokens if you're not careful. Not always. There there's, you know, reasons why this is actually can be a good thing. I'll talk a bit about that in just a second, but just be mindful that's a common mistake.
24:01 and then sub agents for delegated responsibilities. if the orchestrator has to sort of make all of the same decisions, it didn't really help to move that to a sub agent. it's and if you think that maybe this is like a single responsibility idea coming from typical coding that if I give this sub agent a more limited view of the world, like maybe it can only see clickouse for example, then maybe I've isolated how much damage it could do if it you know wants to go crazy. But in reality, most of those scenarios are much better covered by Monstra's tool guard rails and having more deterministic checks to say whether that tool call should be permitted in that context than trying to push it to a sub agent and just kind of hope that it doesn't reach outside of its boundaries. There's just a lot better tools to solve that problem.
24:48 And then sometimes people lean into sub agents for like workflow orchestration. but really Monster workflows are perfectly built for this. They have this nice deterministic way where you can model out a series of steps to execute in a workflow. Some of those can be deterministic functions. Some of them can be agent executions. They can run in parallel and aggregate results back together. we use workflows for a lot of our other behaviors that we have in the system. and they work really well for that. So we don't lean into sub agents to try to like model these different workflow steps with some kind of orchestrator that just follows a runbook or something like that. yeah, these are just common things that we kind of stumbled on as we were going through talking about how we adopt sub agents. but focusing on the speed and cost specifically I think the the common sort of dividing line to me is does this agent that this sub aent that I want to delegate to need context from the parent or is it duplicating work with other sibling agents? If so, you could actually run into slower and costlier sub aents and you probably want to find a way what we've done with the Axium one. slice it out so that you're instead of having the parent delegate a particular task to get done and then that task comes back with a pretty summarized like small result back and then let the agent go the sub agent do all the work to solve that task on its own without bloating the parent if that makes sense.
26:15 >> Yeah. And and I think, you know, this is sort of a more qualitative thing that I've seen come up a lot, but it's also just a lot more it's a lot more exciting to say you're building a multi- aent architecture, right? It's kind of funny. I've noticed like I almost think of it in in three stages where >> you should try to figure out first and foremost, can what you're building just be done in a workflow? Because the amount of times people have built an agent around something that you can really just make a master workflow for is is very surprising honestly. And I mean this is sort of a lot of the issues you're saying to the max where if you could have something fully deterministic the amount of like cost and latency bloat you get if you move that into an agent is really really extreme. And then I think kind of what you did was super smart, right? Like you had this really scoped project at first. You started to see certain bottlenecks it was reaching and then that's when you reached for a multi- aent architecture rather than just saying I want to from the jump look at this as something that's multi- aent. So, I think all of these reasons make a ton of sense. And there is also sort of I feel like people, myself included, always are tempted to kind of think of this as what is this going to be in a year from now?
27:26 And in a with all these different responsibilities, I'm going to want them split out like this. But in reality, the smartest thing to do is really just take it one step at a time. And another thing I'll note here is, you know, >> as you showed, you're never locked into one architecture. It's very easy, especially with MRA, to go from a single agent to a multi-agent architecture. So, what's great with that is you don't have to have a full game plan of what everything's going to look like in a few months from now, and it's very easy to refactor everything. but yeah, these are these are some really relevant points that I think come up a lot. I I love this slide. I feel like it's it really resonates with people.
28:03 >> Yeah, that's awesome. My my general suggestion is start with a single agent and give it all the tool calls until you start to see that become an issue. And our issue in this was surfaced by the context window was just getting super bloated and it was doing compaction much earlier than we thought it should. And then once we investigated that and we saw it was actually this one slice of the Axiom query execution and the iterations it took to get that result that was causing the entire problem and moving that one thing to a sub agent with a really defined this is what we're trying to solve with it completely eliminated that and everything else got to stay regular tool calls in the parent.
28:38 >> Yeah. And this also, I mean, this is a theme that comes up all the time, but it all goes back to you want to make sure you have good eval, you want to make sure you're doing good versioning, all this different stuff. So, when you do run into these problems, you can test out different agents. You can actually dig into the traces. And it's funny how often all this stuff comes back to that, but it really can help guide you a lot more if you have like a robust set of eval science around it.
29:04 >> Yep. So, with that in mind, there's a few things that we intentionally did not delegate to sub agents, kind of following those rules. One was the runbook locator. we actually tried this as a sub agent at one point and it performed worse. because it's actually, especially now that we've moved the runbooks into that ignite workspace I mentioned where it is a a master workspace, it can just spin it up in a sandbox and start running things like GP against it and find anything that it needs. that dramatically reduced how how much sort of token churn it happened to to find the right runbook for an incident. Also the incident classifier in that case the parent agent Sherlock already has enough context to make a decision. So we don't need to delegate that to another agent that is probably going to do a worse job of it.
29:46 and the interesting one to me is the click house investigator because it feels very similar to the Axium one. It's writing a query. It's executing it. It's looking at the results. But I think what happens there is that because the click house SQL language is so much better represented in training data. I find that it makes mistakes far less and it utilizes aggregation much better. So those hypothesis hypotheses that Sherlock sort of constructs it can answer itself with a one-off query or maybe two queries to click house rather than having to do you know four five six that it would do inside of Axiom to answer the same thing.
30:23 Okay, so now let's move over to openweight models. I'm pretty bullish on openweight models. we use a lot of the frontier models for our day-to-day coding, but all of our MRA agents are running with with OpenWeight. So, that's Sherlock, the one that we're talking about today. Also, Gremlin, which is our coding agent we use for like we can dispatch actions to it from Slack and have it do like small to medium amounts of work for us. And then our newest one is Gizmo, which is our code review agent that we are having a lot of success with internally.
30:54 So what we what we did for our approach to openweight was to rather than try to just randomly pick or you know sort of different models for different agents as they spun up, we thought more about the task at hand and what role that agent plays on the team and then try to assign models to the task based off of what we see in their capabilities. So for example, we have a reasoning role which is used by Sherlock and Gremlin.
31:18 Those are agents that need to orchestrate other agents. they need to make more decisions and handle nuance. they just need to think more than a lot of the other agents need to. Gizmo gets its own code review role because that's kind of a specialized task we found. the Axium exeutor that I mentioned that has a data analysis role. And then we have things like summarizing, classification, structuring output data. All of these are fairly like tight defined roles that then have agents like one or more agents that are associated with those roles.
31:47 >> Yeah. Yeah. And I I think really quick on just on that slide cuz it it's it's so important and I think a lot of people that are deep in the space and have been working with it a lot see this every day just from using a bunch of different models. But it is if you can really start optimizing around models for whatever role you're using. You'll see a huge change in how your development process goes. And that's what's great with Monster 2 is obviously it's plugandplay with your models, but it is it's also so interesting though just to sort of like be playing around. New models will come out. They'll be great at some things, not so great at others.
32:22 It's been very funny to see. I was chatting with some teammates about this, but how especially a lot of the newer models are they're just overreasoning where for the great majority of tasks that we're doing, we're going to like more legacy older models, smaller models, a lot of openweight stuff as well because these newer models are great for these super complex tasks. And again, just like the multi- aent architecture, it's exciting to be like, yeah, I'm using this new model. it's it's you know the newest and brightest and I I get tested out but in reality especially if you're doing >> repeatable simple tasks it both helps with latency with cost and honestly with accuracy to have models that are best suited for whatever task you're doing.
33:07 So I think this is super like I I mean all these slides are gold but this one especially too I think is so important to think about while you're building stuff out. >> Yeah. a practical example of that. I I've been building a lot of stuff with Fable lately in some really complex like exploration type like I asked Fable to go do a bunch of different evaluations of different configurations of the this infrastructure we're building and run these different benchmarks against it and it will pull threads and sort of follow paths to answer questions for me and come back with a nice summarized view of that like an hour later. Right.
33:41 that's a Fable type of task I want to throw at it. Yeah, >> but I don't like using Fable for code review because it's way too picky. It thinks too hard about the code that it's doing and it will never be satisfied with any code even if it wrote it. >> yes. >> So, you've had to find this balance like you said of like >> how much reasoning do you actually need for this task that you're trying to solve versus and save the big models for the the times where you actually have a really complex task to get done.
34:09 >> Yeah. Yeah. Same exact thing with us actually. we we have this like go to market agent that that we've been working on and at first it was very sort of like scrapped together. So we're using these larger models to sort of be able to deduce what different tools they should be pulling because it wasn't an exact science and over time we've gotten it really down to having like basically repeatable workflows >> and I tried it with Fable of it and it would try to just really like overthink everything. It would it would be like all right I see that workflow there but let me see if I can solve this better myself. It's like no no no just use just use the workflow we have. It's not don't over complicate this. So yeah that resonates with me as well for sure.
34:50 >> so yeah so a few weeks ago this is what our model sort of mapping looked our reasoning model was GLM 5.2 with a fallback to Kimmy which MRA has a great support for doing these fallbacks. So if one provider is down it will just automatically switch to another one. you don't necessarily need some kind of gateway to pull that off. data analysis was minax M3. Same with summarizing and classification. Those are just simple text tasks. so we gave that to Miniax M3 to 2.7 fallback. But then GLM 5.3 Flash came out and it kind of just dominated everything. everything I've tested with it so far, it has outperformed everything else at a a great cost balance ratio, right? so it even unlocked new cases like our code review agent that has spun up like since then and GLM 5.3 Flash can run through code review and produce excellent output at such a cheap cheap cost. It it's just really impressive. so currently we're running 5.3 Flash AC across all of these, but we have different fallbacks because while 5.3 Flash has served well in all these different use cases we've thrown at it, if it has to fall back to something like like GLM 5.3 could probably do data analysis fine, but it's kind of overkill for that and it' be a little more expensive. where miniax would be just fine at it. so yeah, anyway, I think it's good to like explore these different models. And the way that we map this up is actually like this is a true snippet out of our code of how we configure our our sort of model routing. we just define constants for all the different models and like we we run them on fireworks. So we've got all the the designators here.
36:30 But what we do is we have these like named ROS with the model array where we have like 5.3 flash 5.3 Kimmy K3 how many retries we do for each and then when we spin up something like Sherlock we just point it to this constant that we've exported from this one models.ts file which then lets us change models with just a single line of code across everything that does this like reasoning role. same with our classifier we just drop in a new one. So when 5.3 Flash came out and we wanted to try it, it was as simple as going to some of these roles and just dropping that line in as the primary model and everything just started taking that one as the the pre preferred one. So we didn't have to go touch a bunch of different agents and you sort of make a lot of changes throughout the codebase. It was just one file and if we ever wonder too like what agents are running for different models or sorry for different roles, we just go into this file and we clearly see it.
37:24 >> Yeah. And another note I want to make that's in a super similar vein here is you can also set up evals by models. So if you're doing like a more in-depth analysis of trying to figure out which one you want to be the primary and then what you want to fall back on, you can always set up different experiments that you can run where you're testing different things out. And I know this is a little self-promotion, but I need to highlight what what my producer just said here. Fable thinks every task has extra credit. And I think that's a great way to summarize it. And there are times there are times you want that. But a lot of times >> if if there aren't extra credit questions on the test, you just want someone that answers all the questions, right? Not someone that's like, well, I I added this extra essay at the end just to show you I really know what's going on.
38:09 >> but yeah, I think that's a perfect way to summarize it. And that's why again it's it's again I think for a lot of this stuff a pattern is and I know it's boring but try to start with the barebone stuff you can do if you can get away with an openw weight model something that runs faster and and you know can have less cost on all this different stuff why not why not and I think that's a lot of stuff that's very important to think of especially as you scale things cuz it's you know when you're doing your own isolated thing kind of like you even with sharing it in Slack. The second you what we haven't talked about is the second you make something multiplayer, all these costs start to bubble up because then you have a lot of people that you know both are using it for what it's meant to use for.
38:54 But it's fun. It's fun. It's fun to use. So I noticed that when I release anything into the wild, it's getting a lot more attention than it used to. So it's really important to try to optimize all this stuff. And that obviously is 100xed when you're talking going from an internal to an external project and you're just giving it to users anywhere. So all this stuff is super important to analyze and definitely a big focus at Monster we have is making it very interchangeable so it's easy to do these analyses.
39:25 >> Yep. the the fable thing is is still killing me because I I can just see that happening where it's like I can solve this task really easily, but what if I could impress him? >> Yeah. No, exactly. Exactly. It's like the the kid in class that's like an annoying level of like overachieving, you know, like the like from a show or something where it's like, >> okay, there's a there's difference. You could be smart, but we don't need you to recreate the class in your own view.
39:53 That's such a good transition to my next slide of summarizing text is not a fable class problem, not even an opus problem. so really think like I I think the the frontier models and the smarter ones are really good when you have an ambiguous task. You don't really know what steps are going to be involved and you need it to sort of reason through as it goes like the nuance and the direction and kind of like explore towards a goal. But if you have this really well-defi defined series of steps that you need to take, you you really don't need these like frontier models to solve that in most cases. There might be some exceptions, but for the most part, the the more defined the task and the more well-ritten it is and well structured it is the further you can push down that like intelligence level and go to the cheaper and cheaper models. And it's also impressive now with things like 5.3 Flash from GLM just how how smart they are for the cost and speed that they have.
40:47 Yeah. So, let's wrap up with a few takeaways. back to the the sub aent context or or concept. Protect your context window of your parent agent. you can strategically use sub aents to offload big sections of context to get task done, but don't use it for any of those like antiatterns that we we referenced. route your models by role. It makes it a much easier to think about what model is good at what task to be solved rather than each individual agent and what possible things it could want to do. and then keep it adjustable. Monster makes that super easy like we saw in that code example to just drop in another model that you want to route to or as Brandon said, you can do evalu.
41:29 so experiment with it, but keep it really easy so that you can swap these things out quickly when new models come out and you're not going through a I don't know a lot of work of like diving through your codebase and or asking your agent to dive through and find all the places where you reference this old model and does this new one fit that task if you've organized it by role. It makes it much easier to to say this new model does solve this problem better.
41:53 Yeah, that's all I had. >> Awesome. Thank you so much, Tyler. we have a few minutes now for questions. So, feel free to throw some in the chat, but one I got async that I want to dig into and sort of is just right on what we were just talking about before, but I'd love to chat a little bit about how you think about actually choosing these different models and mapping these different use cases to these models. I'd love to get a a bit of insight, and I'm sure everyone watching would on how you actually decide what's best for coding, what's best for chatting, what's best for all these different things. Is it, you know, just word of mouth? Is it you using it yourself? Are you using any resources to try to dig into rankings?
42:37 And how do you sort of see that as how do you see that progressing over time as well? >> Yeah, that's a good question. I so we don't do a ton of evals yet because all of these are internal agents. So it's a lot of vibe evals. the consequences of us trying a new model and saying this didn't do as well is not very high, right? so we'll throw a new model in sometimes and then just you know walk it back or shift it. but a lot of it is I'll look at I'll look at some of the benchmarks, but I look a lot more now at the cost per task rather than just how well they did or >> how how what the token rates are.
43:16 Like token rates to me at this point are not that important. It's more important to look at what the cost per task is. And then I try to map these tasks to what are the real things that we're doing. because some of the tasks are pretty ambitious, right? Like the ones that something like a Fable or an Astra are good at are these these large ambitious runs that you can look at like a a GLM or other open point model might struggle to do it. Might actually cost more because it has to take so much longer and more tokens to get to the same result. but what I I look at is try to like correlate this is a comparable task to what we need to do and this is how the benchmarks are showing that it cost to solve that with this model. And then I go try it myself.
43:58 I'll plug it in and I'll run this same like run one of these reports for me that I've been you know been doing with other models and kind of get a gut feel of did it produce a similar output or output. but if you have the task well defined enough I'm finding that the deviations per model are not as significant like for that report one for example the next thing I want to do is to build a skill for how it should generate high quality reports which will reduce a lot of the ambiguity around like what does it mean to generate a chart right >> and now I can start pushing that further down the intelligence ladder because it doesn't need to discover that on its own >> >> so maybe a long way of saying vibes >> yeah no No, I I think it's I mean to be honest, I think we have a very similar answer internally. And it's it's funny because I think I mean I'm definitely starting to think a lot about especially even just going through this presentation with you on how we can formalize this stuff more. But right now people have so many so many thoughts on you know what models are best and everyone kind of has their own learned experience from it. But right now it's a lot of like people I'll co-work with will be trying certain things. I'll have my own takes and we're definitely thinking a lot about doing our own benchmarking as well because >> a common thing we'll run into at MRA is what's awesome about MRA is it fits in anywhere as you said you can plug and play any database any deployment any model all this different stuff but what we run into a lot then is people being like okay yeah but like what do you guys think we should use like what is your suggestion and so I think we're starting to think a lot about actually translating those vibes that we know are correct but into real data that we can back up.
45:40 >> but I think it's a really interesting space and and all the tools are there to do these analyses. >> it just takes a lot of time and it's also you know with how fast everything comes out now it feels like you got to continuously update and stuff like that. but yeah, it's a fun space and definitely fun to be able I I think the big thing is to not have the lock in to certain vendors is really really freeing and to be able to go across different different providers unlocks a lot in terms of what you can what you can experiment with. I see we have a bunch of questions coming in. So >> first one what memories of Monster do you use? So, for your different agents, are there any highlights maybe of how you're using monsters memory?
46:26 >> I imagine you're using a little bit of a bunch of different things, but would love to hear a bit on that. >> Yeah, we mostly use memory for thread management because all of the the chats in the the monster channels go into memory for that thread. so that it can just recall them from there instead of having to hit the Slack API or something every time and reload that context. so that's most of it. we don't use a lot of memory across threads yet, but we we are debating it for our code review agent. I have mixed feelings because I also for something like the code review agent, my belief is that most of the anything that we have to correct it on should be persisted back into either the system instructions or into the codebase itself to say this is a in this codebase we care about this or that. but you could definitely think of extending something like a a code review agent to remember when it's been corrected on something and don't make that same recommendation again.
47:20 so anyway, I think it depends on like in most cases I would prefer to codify it somewhere more consistent. So also my like code coding agent for example would have access to that same memory, but there are definitely cases where I I can see an argument to push more into a shared memory across agent threads. Cool. And then, yeah, real quick, can we share the slides? Yeah, I'm sure we can. We can send out the slides after this.
47:46 Maybe we can have a little repo that we send out or something like that. But yeah, for sure. And also >> to mention all these recordings are saved, so you can always go to YouTube and watch them. I think LinkedIn saves them as well. We also do them on Twitter, so they'll live forever. Don't worry if you can't make it live ever. we'll always have them. And if you also ever respond to these invites, cuz we do these every week, even if you can't make it, we automatically send the recording as a follow-up as well. So you can get the slides, get the recording, all that stuff.
48:17 >> And then here, this looks like one from Monster Focus. So how does handle state consistency if a worker node goes down mid task? Okay, so there are a few cool things here and it's and it's all the more reason actually to my earlier point of if you can think of things through workflows, life gets gets a lot easier. But one really cool out of the box thing for MRA is with workflows. We have what's called snapshots, which is basically just for every step of a workflow you're running to. B has a natural podcasting voice. That's funny.
48:49 so for every every step that you go to, we save like all the inputs, outputs, and what sort of tool run you're on. So what this does is it makes you able to if there is any kind of like server crash or anything like that to restart from the last step and that's something I'll actually pull up the link so I can share that. So that's sort of one thing that we work on there but we're also releasing or we have released durable workflows which durable workflows basically spread your workflow runs across different servers.
49:24 So rather being than being linked to a single process, it's running on all these different servers. So if there ever is a process that goes down, it can still continue to run. So there's all these different things that we have because as we've said, there's a lot of people building on Maestra where especially once we chat with them, we figure out a lot of their stuff can be accomplished with workflows. And because the workflows are so flexible, like you can have agents within workflows, workflows within agents. A lot of times there's a workflow structure that you can use to build these snapshots.
50:01 and I'm going to share this super quick. This can >> Yeah, here we go. This is just a link to snapshots on our doc. so you can go there, read more about it, and again, we have durable workflows as well. So yeah, we think a lot about, especially as time goes on, we're getting into agent use cases where this isn't something that's going to run for like seconds or minutes, but things that will run for hours or days as it waits for different inputs as it does really long async tool calls, all this different stuff where we have to have a big focus on durability. So we have sort of the snapshots as like a backup last defense, but we also have released durable workflows, durable agents, all this stuff that spreads your run across different processes so it can just survive any sort of crashes, any sort of things out of your control.
50:52 >> that's awesome. >> And then we have another one here. In SR's context, we do deal with Dinatrace, noble and elastic. We are working on on call agent MCP documentation repository via rag etc. So your explanation of keeping sub agents for specific duties and avoiding agents if work can be done via workflow helps a lot. I don't know if you guys could read all that because I don't think it all showed up on the screen.
51:25 >> but yeah, 100%. I I think you know it goes back to a pattern of what I've seen with people in the field for sure. I'm sure Tyler has seen building and we see internally at MRA all the time. But it really does what we find is there are a lot of use cases where a lot of what you're doing can be made a workflow. And what agents really unlock is sort of like a 10 to 20% of the process that you do need a bit less determinism and ability to handle more ambiguity.
51:56 but the more you can map there. Obviously, as we said before, the cost, the latency, all that stuff is improved, but it also just makes the system much more predictable, it makes much more auditable, it makes it easier as we're talking about multiplayer and more people getting into it. It makes it a lot easier for people to get into your system and understand what's going on as well. And, you know, from there, there's so much other stuff. It makes it easier for other agents to kind of go in and understand what's going on. so all that different stuff.
52:26 and then yeah, I think that's it. We have we have I have a natural podcasting voice. I It's all the mic. It's It's all the mic. >> and you need to get a a Soundcloud. Just do like a, >> you know, like easy listening. Yeah. Just read read the news to us. >> Yeah. It's funny. I actually I I teach yoga in my in my spare time when I have the time to do it. And the feedback I've gotten as a teacher usually has little to do with my teaching and a lot to do with my voice, which is nice, I guess, but people are usually like, "Yeah, your voice is great." And I think that carries me a lot with with my teaching ability. But I I appreciate it.
53:06 but trust me, I pick up one of these. I think this is a Yeah, it's a sure sure microphone. S H E. It It helps a lot. It's funny. I I listened back to as I used to just use my AirPods mic and then I got this cuz I was listening back to some calls I was doing and I was like, "Oh my god, it sounds so scratchy." Like this is I'm supposed to sound like a you know a someone that knows what they're talking about and you can barely even hear my voice. So highly recommend getting getting a mic to anyone out there that's that's looking to have a little little voice glow up.
53:40 >> >> Oh, good question. >> Okay, cool. Yeah. How does the parent agent know whether sub agent ran out of steps versus completed the task successfully? >> All right, so we've got a new thing that we just deployed like last week that is making this phenomenally better. we added a post-processor which is a construct in MRA that you can go go look at or ask your agent to to build for you. that will append after every tool call or batch of tool calls that is it will append to the the instructions to reply back to the LLM what step it is on out of its total. So the agent actually knows as it progresses if it's getting close to running out of steps and it will we found that just having that knowledge helps it to say oh shoot I've got to start wrapping up I'm I'm running out of my ability to get more data and we've kind of tuned the instructions that we put in there for some of them beyond just like the step count to say what should it do when it starts to run out.
54:42 so for example it is the yeah Sherlock has this this post processor added to it and it will instruct it so that if it thinks that it could have done more research to pull another thread or explore another direction then and it's starting to run out of steps it will summarize back to the operator in Slack to say this is what I want to explore next. so then we can just say that sounds good. Let's continue there. Or it might give us a few different options and we can say, "Oh, actually this one sound like one of three sounds like the the right direction to go." So that helps us from we still bound how many steps it takes so that it's not like constantly just off this long exploration task, but it forces it to come back and say, "This is what I found so far. This is what I think I want to do next. Which way do you want me to continue?" so it doesn't maybe directly answer the question of like how does the parent agent know, but it it sort of constrains the sub agent or the parent agent to understand when it needs to like stop researching and start summarizing.
55:43 Awesome. Yeah, that that's incredible. all right, I think we're coming up at about time here. We're almost at the top of the hour, but first off, want to thank everyone for joining. it was an awesome session. and thank you for all the questions and if you ever, you know, have any questions about what you're building and want to chat, feel free to grab time on my calendar. Always love chatting with people, figuring out how they're using Maestra, try to guide them to the best things possible because, you know, like with the state resiliency question today, there's so much to master that a lot of times it's tough to sort of like drink from the fire hose, per se, and get all that information. And so always happy to chat with people more on this kind of stuff. Always happy to dive deep into your use case. So I want to thank you all for coming and riding with us. And then Tyler, thank you so much. This was awesome. Thanks for having me.
56:38 >> Super pumped. I appreciate you taking the time and yeah, hopefully we'll be able to do something again soon because this is awesome and I'm excited to see where Sherlock goes from here. >> Yeah, thanks for having me. Bye everyone. All right.
Summary
- Sherlock is a Slack-native operations health assistant that generates performance reports and monitors system health.
- The use of sub-agents allows for efficient handling of context-heavy tasks without bloating the parent agent's context window.
- OpenWeight models are employed based on the specific roles of agents, optimizing performance and cost-effectiveness.
- The architecture of Sherlock includes various tools and integrations, enabling it to access telemetry and application metrics.
- The session emphasizes the importance of workflows over agents for deterministic tasks to improve efficiency and reduce costs.
- A post-processor feature helps manage task completion and ensures sub-agents summarize findings effectively before running out of steps.
- Continuous evaluation and adaptation of AI models are essential for optimizing performance as new models become available.
- The integration of agents into collaborative platforms enhances knowledge sharing and productivity across teams.
Questions Answered
What are the hosts' locations and current weather conditions?
The hosts are located in San Diego, California, and near Charlotte, North Carolina. San Diego is experiencing unusual humidity and heat, while Charlotte is cloudy but cooler than previous days.
What is Sherlock and how does it function?
Sherlock is a Slack-native operations health assistant that integrates various telemetry data sources to provide context for operations tasks.
When should sub agents be used, and what are the potential pitfalls?
Sub agents can enhance task efficiency but may also slow down processes if not implemented correctly. They can lead to increased costs if they require recontextualization.
How have the hosts optimized their models for better performance?
The hosts have refined their workflows by utilizing specific models that enhance task performance and reduce complexity, focusing on repeatable processes.
How is memory utilized in managing different agents?
Memory is primarily used for thread management, allowing agents to recall previous interactions without repeatedly accessing external APIs.