Section Insights
Introduction to the Workshop
What is the purpose of this workshop?
The workshop is focused on teaching participants how to give their master agents a computer using the Mastra framework.
- The host introduces themselves and the workshop topic.
- Participants are encouraged to engage by sharing their locations.
- The session is being conducted solo, which adds a personal touch.
Overview of Mastra and Workshop Goals
What will participants learn in this workshop?
Participants will learn how to provide their agents with a computer, starting with a local sandbox and progressing to remote sandboxes using Daytona.
- Mastra is an open-source TypeScript framework with additional tools.
- The workshop will cover local and remote sandbox setups.
- Understanding persistence in workspaces is crucial for agent functionality.
Setting Up Remote Sandboxes
How can agents utilize remote sandboxes?
Agents can use remote sandboxes by integrating with platforms like Daytona, which allows for isolated environments on the same server.
- Using Docker is one option for sandboxing, but Daytona offers more convenience.
- Daytona's remote computer feature is highlighted for its ease of use.
- The setup process for Daytona sandboxes is straightforward.
Configuring Sandbox Tools
What configurations are necessary for sandbox tools?
The configuration involves disabling unnecessary file system tools and enabling the execute command tool for specific tasks.
- Selective tool activation enhances security and functionality.
- The demo illustrates the process of executing commands within the sandbox.
- Participants are encouraged to troubleshoot issues that arise during live demonstrations.
Using Browser Agents
How can agents interact with web browsers?
Agents can spin up a local browser to perform tasks like retrieving information from websites, with options for headless operation.
- The browser agent can be configured to run in headless mode for efficiency.
- Participants are shown how to manage browser sessions programmatically.
- The importance of documentation and community resources for troubleshooting is emphasized.
Transcript
1:23 >> Hello and welcome to another master weekly workshop. That's actually my first time rolling the intro, so I'm not even 100% sure if I need to do something to get that to I bet you can hear me, but you can't see me, right? How's it going, Cali? Hey, Bobby have. Hold on, you're Okay, I'm live. You can see me. Perfect. Cool. You know, it's almost like I've done about 50 of these. I'm almost professional about it.
1:57 But we will give ourselves one exception and a little bit of grace today because whereas ordinary I ordinarily I host these workshops with an awesome colleague or expert, today we're flying solo to talk about how to give your master agents a computer. My name's Alex. I'm going to be your host today. I live in London. How about you? If you're joining us here on Riverside, please just take a moment to say hello in the chat. Let us know where you're tuning in from.
2:25 It's always fun to see where people connect from and get to know each other a little bit as well. Daniel's tuning in from Brazil. We've got Donna in Florida. Hey Alf, tuning in from Asaf. Frankfurt. People coming from all over today. I love it. Hey Sandro from Queens. We're going to kick things off properly here in just a second.
3:03 Hey Bonnie from Reading. How's it going? Chris is coming from Dallas. Nice. Okay, folks. Well, with everything all set up and everyone giving given a moment to tune in, let's kick things off good and properly.
3:34 Mhm. Oh, I have to stop the media to share. Okay. I I had a feeling that doing this intro video would cause more problems than than it would help. cuz I can't screen share until I stop the media, which has stopped. Now it's officially stopped. You know, folks, running these things solo is never as easy as it looks. I have to tell you. But that's all good.
4:06 Okay, now we're ready to kick things off properly. Hello, everybody, and welcome to this workshop in which you're going to learn how to give your Maestro agent a computer. Quick little question for you first of all. Who here has heard about Maestro and is using Maestro in their agents already? Give us a thumbs up in the chat or just say so in the chat if you're using Maestro. And if not, let me know as well and I can give you a quick intro.
4:34 Thanks for the feedback there, Sebastian, for the slides are small and I'm really big. I think I can fix that pretty easily. here we go. This is This is not what I'm expecting. Okay, let's stop sharing the screen real quick. Share full screen. We're going to figure this out together.
5:05 For some reason it keeps putting my screen in the small box. >> >> Oh dear. Where's the agent to do this for me? Am I actually going to have to do some independent thoughts to figure out how to fix this? Ugh. Full screen view. No, that's not right. So this will be full Okay, we can do just my screen or we can do picture-in-picture like this for some reason. I guess we'll do just my screen. I mean, you're not going to get too much value, I don't think, out of looking at my face here, so let's just go with the screen.
5:43 Okay, and I'll pop in to show my face occasionally, I suppose. All righty, so Machina, if you don't know, is fundamentally an open-source framework for building agents in TypeScript. We've been around for a couple of years now, almost 30k stars on GitHub. Maybe you can drop us one if you're newer to Machina and you haven't already. Machina is also a platform. We offer you observability studio, so you can collaborate with your team, as well as an agent builder that lets non-technical teammates build agents with primitives you define.
6:14 We offer you a bunch of different primitives to help you build agents, from workflows to memory and sandboxes and browsers. And that's what we're going to talk a little bit about today when we talk about how to give your Machina agent a computer. We're going to look at how to give your agents a local sandbox and a file system. How to bring human into the loop into your agent, so that if you are to trust a computer to control a sandbox or a file system, you still want some control, I'm sure, over what it does.
6:47 We're also going to touch on how to give your Machina agent a remote sandbox using Daytona. And just to confirm, you all can see my screen right now, I think. Yeah. You can't? Oh my goodness, this is a This is a legendary Machina workshop already. What on earth is going on? You can This is Okay, this is really quite strange.
7:19 Okay. We're going to have to take a moment to reset this. So, guys, what kind of agents are you building? Please talk among yourselves or help me out a little bit here by sharing what kind of agent you're building while I figure this out, cuz this does not seem right at all. Media, music, sound effects, chats. Okay, so right now you see me full screen. That's cool. That's what we want.
7:50 And then if we do screen share, let's do entire screen. You should see the whole screen now, I think. Still only me. Okay. And if I do this, you can see me but not the screen. Guys, I promise you this is going to be such an amazing workshop as soon as we actually get started. bear with me here.
8:22 So now you see me huge but the screen is little. All right, folks. What would What would you do in this situation if you >> >> Let's see. Okay, so how about we do this full screen view? No, that's making my thing full view. Damn, this is This is like deleting the DB in prod right now.
8:57 I was thinking I had to stop my webcam, maybe. So I'll do video off and then share screen. And then do full screen. And now you guys are going to give me a huge thumbs up in the chat if you can see the whole screen. You Seriously? Mhm. Okay. So there >> >> I need you to do one thing for me, folks. in the chat, give me a thumbs down if you think I should just stop and try and start again. Or give me a thumbs up if you want to figure this out with me together.
9:37 >> >> Mhm. Okay, let me see. I think the answer has to lie in sharing the screen, right? And then So, right now you're looking at the screen down You can see the screen down here, probably. But, the screen's tiny.
10:08 How This is rough, not going to lie.
10:55 How about this? Perfect, it works. Okay, I'm up. Hey, listen. I just did the same thing I did three times already, but I am not one to look a gift horse in the mouth. All righty, okay. Welcome, everybody. Did you If you've just joined, this workshop has just started. You've missed nothing. This is great timing. We're about to kick things off as we talk about how to give your Master Agent a computer.
11:28 We host these workshops twice a week to help you build better and more capable agents using Mastra. So, check out the link there if you want to see all the upcoming workshops. I'm your host, Alex Booker, developer educator, and focused on developer experience at Mastra. If you don't know about Mastra, we're fundamentally an open-source TypeScript framework, but we also have a observability platform, studio, and agent builder. Today, you're going to learn how to give your agent a computer. That involves giving it first of all, we're going to level this up one step at a time. We're going to start with a local sandbox and file system.
12:05 Then, we're going to see how we can still approve an input on actions that happen in a sandbox. Just because it's isolated doesn't mean we want everything to run loose. Then, we're going to look at how to set up remote sandboxes with Daytona. Daytona is a platform that offers sandboxes, and Mastra has a nice integration. We're also going to look at how to deal with persistence workspaces. So, sandbox by nature is somewhat ephemeral. The idea is that they spin up as a virtual computer, you do some work in an isolated environment, and then eventually it winds down, right? But, what if you want to preserve some files for a long time so that the agent can spin down and then spin back up and work on it? I'll show you how to give your agent a file system.
12:51 And true to the title of the workshop, we're going to look at how to give your agent access to a computer. There are two ways you can do this in Mastra. One is that you can just give your agent access to a browser. It's going to be so cool. I love this demo. Or, you can actually give your agent genuinely a real computer with a file explorer, a browser, the ability to install programs, and fundamentally control a mouse and keyboard to do tasks just like a human would.
13:20 If you have any questions during the stream, please just ask them in the chat. I've got one eye on the chat, and I'll answer questions as they come throughout the stream. And by the way, just if you're wondering about the recording and the code, I'll be sure to send you a link after the stream. As I mentioned, and as you may have just experienced, we usually host these workshops with two people, and I kind of take care of the host role while someone else presents, but today I'm hosting and presenting. And I thought And I love these workshops because they feel a bit more informal somehow. And so, please get involved in the chat. That's the beauty of being here live, and let's go through this together as we look at a few different demos.
14:01 So, I prepared the code for us so that we're not going to waste any time writing new code. Although, if you want to see the agent do something, let me know, and maybe we can make that happen. Let's start with the first demo, which is the local sandbox. So, what is a local sandbox? But, essentially, a way to give your agent tools to execute commands in bash, and tools to give access to the same local file system as the host that is running your master agents.
14:34 This is really common and advantageous when you're building something like a cloud code or a master code, where you want an agent to run on the user's computer, but not be completely unbounded. You might want to set a base path, for example. This will enable the agent to work in files within this path and directory, but not kind of come out of that container to start accessing things like system32, or to go into your root folder and look at environment variables and things like that.
15:05 Let's have a quick look at how it works by spinning up Studio. Master is a TypeScript framework, so you can just call these agents using TypeScript code, maybe in a root handler, maybe as part of your broader system. We also give you this interactive development environment called Studio, which, once you run, is available usually on port 4111. Here's all the agents we have to find in this project. They're obviously kind of demos numbered and things, but these could be sub agents for example or parts of a multi-agent system. Here's our local sandbox and file system. I can say hello and tell it to execute the command 2 + 2 and write the result to result.txt.
15:43 Now, 2 + 2, an agent doesn't really need a command to do that. It does however need a way to run things like bash code to do coding-related tasks or to do things like an open claw agent would or more complex or autonomous tasks. I'm just using this as an example and we write to result.t result.txt. We can then ask the agent to read from result.txt. And if you've ever used Mastra code by the way, which is Mastra's coding agent, this is kind of our version of Claude code. It's all built with Mastra, open source by the way.
16:17 And an an amazing way for us to build Mastra using Mastra. we have some plugins and some skills specifically designed for building Mastra in there, but anybody can use it. You the users this exact same feature. And you can see in settings here that just by virtue of enabling the local file system and the local sandbox, the agent has on it Let's just organize these windows a little bit.
16:47 I I would maybe stop showing my screen at this point, but I I didn't even didn't even touch the screen sharing features right now. just by virtue of having those sandbox and file system values, the agent gets all these tools like read file, write file, edit file, and so on. Let's look at the next demo, configured tools and agents. And so, the idea being that yes, you have those tools now, but you might want to dial in some configuration. For example, you might want to disable some tools. Every tool is enabled by default, but suppose you only want the agent to read files, you could disable the writing tools.
17:25 Here I disable all tools by default, and then I dial in the options. So, I enable the read file tool. For the write file tool, I enable it and then I set required require read before write. This is a really important behavior of tools for touch file systems because suppose you're building a coding agent, you might then go into your text editor and start editing the file, and then you go back to the coding agent, but it has a sort of out of sync view of what the file contents is.
17:54 So, you should always do a read to make sure it has the latest view before attempting a write. You can also enable things like human in the loop, in other words, approval. So, here we set require approval to true. Let me show you the impact of that. And so, we'll go back to agents into our configured workspace tools demo, and this time we're going to ask it to execute the two plus two command. And once again, I know this feels a little tiny bit contrived, but this could be a coding agent, for example, trying to deploy a service. And so, this time it's not allowed to run freely. The human in the loop has to approve it. The cool thing about studio, right, is that you get this UI pre-rendered. It's great for testing, but this is fundamentally a feature of Master I called human in the loop, and you can control all these things with code as well.
18:39 The other options you have with regards to the local file system and sandbox relate to things like containment. For example, making sure that the tools don't try and execute certain commands or they don't try and read files outside of that path. You can read more in the sandboxes documentation. And when we talk about giving an agent a computer, we're generally talking about giving it a sandbox because a sandbox is a type of ephemeral computer, but it's not just a sandbox.
19:08 It's also a browser. It's It's the ability to control a GUI and things like that. Let's look at the next demo and keep things moving. What we're going to do is just make this a bit smaller. So so far what we've done is run the sandbox locally, which again is great for simple server-side applications and for any agent you might put on your user's computer. However, if you're building an agent that has multiple external users, you probably don't want to be using the local file system. It does have some basic options to do with sandboxing and some basic options to do with containment.
19:49 I'll just quickly show you in the docs actually that there's a few options relating to security. So when you enable a local file system, you have the option to turn on native sandbox isolation, and this will basically stop this is actually a really really good idea because sometimes agents do just take a wrong turn. It's not necessarily malicious. And then for example, it might accidentally do a fork bomb, which is where the bash command overloads the computer to the point that it crashes.
20:21 so you it is genuinely genuinely a good idea to enable native sandboxing because it is a guardrail. However, I would not rely on this when you have external users because there's no there's always sort of edge cases and there's always things that you need to consider. it's not a truly isolated environment. That's where this idea of giving your agent a remote sandbox comes into the mix. There's a few different ways you can do this. I mean, first of all, you can run something like Docker on the same machine as your Maestro server, and then you can run them in isolated ways on the same server.
20:56 Generally, it's a bit more convenient to use a service like Daytona or E2B. We also, by the way, at Maestro on the platform enable you to spin up sandboxes very easily. I want to focus on Daytona today. Not only are they a great company, but they have specifically remote computer feature that we're going to need a little bit later in today's demo. And so, let's look at how we can do that next. So, we install the Daytona sandbox package, and instead of giving it a local instead of giving your agent a local sandbox, now you give it a Daytona sandbox. And it really is that simple.
21:34 What we're going to do is go back to studio, and we're going to look at the Daytona sandbox example right here. And we're going to say execute 1 + 1 and write it to results. txt. But before we slap enter, we're going to go to the Daytona website. Bear in mind, I've already set up an application and a key. And here we can see a list of sandboxes. As I mentioned, sandboxes spin up and then they spin down. And these are all right now stopped because I only I must have run them a few hours ago at least at this point and and many days ago in others, at which point they get archived. But check this out when I put them side by side. I love this. You hit send. The agent is going to call the execute command tool. That's basically going to spin up the sandbox. And what I would hope to see over here, just on the right here, is that the sandbox started.
22:22 And it calculated 1 + 1 and wrote it to the home Daytona result.txt directory. Now, with the local sandbox, we gave it a local file system, and that specifically gave it tools relating to writing and reading files. That's the same tool we set the human approval on, and it's the same tool we could disable or enable. but in the case of running a remote sandbox, it's actually pretty beneficial just to let the agent run with bash cuz then it can pipe, it can cat, it can LS, it can work on the file system, and do whatever it needs to do. And you don't have to worry about if it tries to access the roots or if it tries to read something it shouldn't, cuz there's nothing it can't read on there since it's an ephemeral sandbox. I mean, sometimes people misconfigure or maybe misconfigure is a bit of a overstatement. Sometimes people do put sensitive stuff on sandboxes. It's okay to do that as long as you know what you're doing. Generally, you're not going to have generally, in most workflows, that's not going to happen though. So, you don't have to worry about leaking keys and stuff like that.
23:19 And here's the kind of cool thing about Daytona as well is that you can click in to the UI here, go to the terminal, and then you get like an SSH session. So, I can come here and LS, CD into home, CD into the Daytona folder, since that's where it wrote the file. and there's the result. Thought I kind of muddied the output there with one command, but you can see highlighted with my cursor result.txt. And and I can cat it, right? And so, that's really cool. Now, everything is completely isolated. There's a couple of problems though with this setup. I mean, the first is if I create a new chat, what do you think? And let me know in the chat. If I ask it to read result.txt, let me tell you it will find the file.
24:06 What I'm wondering is, do you think it's going to find a file, or do you think it's going to return the result that we wrote previously to the file system? Is it actually going to read the same file? Because in this particular example, the sandbox is basically scoped to the agent. So, any agent run, regardless of the thread or the user, is going to share the same sandbox. Sometimes that's what you want. Other times, it's not. You want to create a sandbox per resource or per user, per organization. I'll show you how to do that in a second.
24:42 First of all, we need to talk a little bit about how to give the agent a file system. Because right now, I mean, it's great that we can write files to Daytona, super convenient, and this is great for like artifacts. I mean, suppose you run a script to generate a PDF or generate a PowerPoint presentation, or something. That script can work in the sandbox, do what it needs to do, and then write it to the file system.
25:05 But, suppose that sandbox goes to sleep and is archived or deleted, it's not a permanent file store. I will say, most sandboxes run and live a lot longer than you might expect. Sometimes they will run for hours, and sometimes they will live for weeks or months, and that can usually be configured depending on your plan with that provider. And so, it's not but they're necessarily ephemeral, but they're certainly not intended as a long-term solution. You also ultimately need a way to get the file out of the sandbox into a different system, right?
25:35 Just back onto the host computer, or maybe into some system that the agent is interested in, like maybe you just want to upload it to Google Drive, for example, in the case of a PowerPoint presentation, or something. Fortunately, Master gives us a few different file system options. The first is the local file system. We spoke a little bit about that already, but you can also enable things like an S3-backed file system, or even a Google Drive file system.
26:02 This works in two ways, and let me show you them. Do let me know if you have any questions. I've got one eye on the chat. And so, here we have an agent set up with the same Daytona sandbox, and a file system powered by S3. I've already preconfigured my access policy and my S3 buckets. And what I'm doing here is defining a mount. What is a mount but a virtual path on the sandbox connected to a virtual file system. So, now if we write any files or read any files from workspace using carts or any other bash command, it's going to actually almost seamlessly via a fuse mounts.
26:44 That's a special sort of software that enables you to create a folder on a computer that isn't actually stored on the computer's disk, it's stored in some cloud, but as far as the environment is concerned, as far as bash is concerned, it can just read and write files as if it was a local file system. When you create some mounts, that also adds tools to the agent. Let me show you. So, we're going to go back to our agents view and then see the Daytona Sandbox with S3 example.
27:13 When we look at the settings here, you can see that we have those sandbox commands like execute, kill process, and so forth, and the file commands. If we use any of these via the agent, they're going to talk to S3, but even better than that, if I just say hello, Oh, Twa mentioned that E2B and Daytona already support persistent sandbox workspaces. I'm sure we could integrate with that. And I can say write, you know, hello. hello to hello.txt in the What do we call it? The workspace in the workspace path.
27:59 Yeah, the fuse mounting is is very cool, because as you can see, so this is actually, okay, so this is funny. it's completely sort of gone around the sandbox and it's just used the right tool to write to hello.txt and then it's read it. And then if I go to S3 in this case, wish me, if I had trouble, if I had trouble navigating the the screen sharing earlier, wish me luck in the S3 interface, but here we go. here are the objects, hello.txt and then we see the content.
28:30 there's a few ways, isn't there, but let's just download it for simplicity. And then it says hello. By the way, I didn't really show you, but if you don't feel like navigating to S3 to figure out what happened, Master also has this workspaces tab, which will let you sort of dig in and see what's going on. We have a bunch of different workspaces on this particular master instance. We want this one, I believe. Here's the workspace folder. As you can see, it's backed by AWS. I've got some old folders and files in here from previous demos.
28:59 And then here's the hello.txt file, which you can copy and do what you want with. So, it's really handy to have this workspaces tab cuz now you can just verify that things are written as you expected. but this doesn't really demonstrate fuse mounting, right? Because this is just using the tools that were added implicitly to the agent. The way the fuse mounting is going to be really useful is if we're trying to take a result from a bash command and put it into a file system. So, if I say execute 2 + 2 and write the results to /workspace/mathresults.txt, there's sort of two options the agent has right now. One is going to be to execute a command and funnel it directly into that path. And because of the fuse mount, that's going to work. but the other option is to use the tools. And what's happening here actually is that I don't think the agent 5.6 saw it. I don't think the model really was sure that it doesn't really know that this has got a fuse mount, basically. So, it's trying to be helpful. It takes the result and writes there.
29:56 this is because it's such a simple example, but I would really wish Sol luck doing this if it was a 5 MB PDF or maybe a PowerPoint presentation, or if you needed to execute shell script and do some complex computation, for example. I think it could really struggle in those cases. But the nice thing about the really nice thing about the tools in master is that we can sort of turn them off, right? So, let's copy some code from here.
30:25 let's just take this one. And then on the tools array, we're going to set enabled false, and then and just manually enable the execute command one because that will enable the sandbox to do what it needs to do. We need to import that. Nice.
30:55 So just to clarify here, what we've done is disabled all the file system tools, then we've enabled the execute command tool, so we've only opted into that one. Now when I head over here and look at the settings, there's only an execute command tool. Look how smooth that was, by the way. Everything just kind of reloaded in the background. So if I do the same demo again, this time let's do execute 10 + 10 and write it, you know, to workspace /foobar.txt.
31:23 And just to prove I'm not doing any trickery, we'll go to the workspace folder. We've only got hello.txt right now. Earlier I asked it to create me a website, which is why there's a few things going on here. But what we're going to see is it'll probably ask for approval due to the way that it was configured, but this time it doesn't have that. Oh, interesting. It thinks that I've exceeded 30 GB of storage space. Really?
31:54 That doesn't That feels like it could be a bug with the tone up, perhaps. let's look into that real quick cuz that's kind of interesting. It might be to do with all my combined Yeah, so basically this is going to be across all my sort of I guess this is across all my different sandboxes. So what do I What do I need to do? Just like delete a bunch? Maybe that'll help. Hm, but to be honest, I think I've used up all my allowance today on this live stream for figuring stuff out in front of people, so we'll come back to that and you can you can take my word for it.
32:28 Cool. Let's look at the next demo. So here's the thing about using the S3 fuse mounts. Actually, Daytona doesn't really have it installed by default. It's kind of a program you have to install in the sandbox. Because it's such a common thing to do, and because Daytona has quite a nice developer experience, if you try and use it and it's not installed, it installs it in the background. But it is a little bit slow to do because every time you spin up a new sandbox, it's going to have to install the program on the sandbox for fuse mounting.
33:01 and so one option you have actually is to create something called a snapshot. A snapshot is basically a configuration of a sandbox. Think about it like a blueprint for the sandbox. A very common use for a snapshot is to basically seed files or to sort of create a sandbox with specific programs pre-installed. But instead of like installing those programs every time a new sandbox is created, you create a snapshot and then it can spin up very fast, very easy.
33:31 How do you do that in Daytona? But by creating a script. In this case, I've got a script called create Daytona snapshots that very simply creates a snapshot with a name and then runs some commands on the base image. In this case, it's using apt to install the S3FS, which is the S3 file system or maybe fuse system, I'm not sure. so that way when I spin up agents in the future, it doesn't have to do the installation.
34:02 Okay. So earlier, we spoke a little bit about how a sandbox there is audio. You can always let Brad who can't hear me know that on YouTube sometimes you get a better experience than Riverside. And there's a couple of questions here as well that we'll get to in a moment. there's a quick question here from Ahmed. would you recommend enabling both file tools and actually command tools if the execute command can do all what you need. I personally disable the file tools if I'm relying on the sandbox, but bear in mind, you might want to write using the execute commands tool, and then read using the get tool, and then you can read the file on the clients.
34:46 And so, there are some situations where you need it, but if you don't know that you need it, you probably don't, is my opinion. I also find that when you use both, you're basically having to maintain two different policies, right? Because you might configure the file system tools to have human in the loop or to have containments enabled, but then a bash command isn't going to respect those same properties. And so, I think that can trip people up sometimes if maybe you are relying on the tool configuration when actually the bash command is not going to respect that. It would be better to define that policy in one place or not even need a policy because the sandbox is isolated.
35:23 and yeah, the big advantage of tools is that you can enable things like you can force truly force things like read before write. One trick is that you can go into the prompt and tell the agent, "Hey, read this thing before you write." You can also go to the agent prompt and say, "Hey, always ask for users to double check." But they are not going to guarantee anything. Whereas if you can enable those on the tools, you can feel very good for like there's no chance actually due to the way it's configured that the tool will run without reading, and the tool won't run without approval in the case of human in the loop.
35:53 So, moving on to this example, you know, we looked earlier a little bit about a basic example. And what's happening here is we're just assigning one instance to every every new instance of the agent. So, they all share the same sandbox. But it's pretty common, I think, to want to create a sandbox per user. And the way you do that basically in Maestra is to rather than assign a Daytona sandbox instance, you create you assign a function, a little factory function, like this, which returns a sandbox. And when you do that, you have the option to configure the ID, which uniquely identifies the sandbox. And so, in this case, we're getting the thread ID. This is a a unique identifier assigned to every conversation, and then we're using that as the unique key, right?
36:39 Whenever you assign a Daytona instance or a sandbox instance to Mastra, Mastra handles the workflow, or the life cycle, rather. So, it will start the sandbox for you. But when you use a function like this, remember to call start because you are now responsible for the life cycle. Also, rather than running every time, you can create a sandbox cache key, which is basically a map of thread IDs and sandbox instances. So, it's highly recommended that you enable this with the same thread ID in order to make things more efficient.
37:10 And so, what we can do now is go back to studio. We'll look at the snap sandbox per thread example. And in thread one, we're going to say write sandbox one to whoami.txt. Yeah, using Docker image in a Docker sandbox is a really good idea. I would I would highly recommend you look into that. All right. Well, we've got an issue with Daytona because the our kind of storage capacity.
37:45 Let's see if we can fix this. So, I think what I want to do basically is delete all my sandboxes. I'm really surprised it's 30 gig, though. Because I I just don't think even I was I've not used this account in a while, so maybe there is some old stuff that's got bigger files. But you can go to Daytona and delete them with the actions thing at the bottom.
38:20 Oh, yeah, actions. Oh, okay, so the contrast is throwing me off a bit here. So, you do actions, delete 20 sandboxes, delete. Good shout. Look at it go. And then hopefully if we go to What was it last time? Limits. Oh, look at that. Okay, we can breathe again. Healthy. Yeah, that was easy. So, let's try that again. Yeah, maybe the maybe it's actually not the files on those sandboxes, but maybe it's the images. I'm quite interested now because we're going to spin up a new sandbox. And now yeah, so basically a sandbox uses 3 GB of storage. And if you look here, that tracks. That 3 gig I thought might have been RAM.
39:05 but it's actually storage. Yeah, so every sandbox uses 3 gigs of storage for the image and to as you say, reserve space on the disk for the instance. yeah, you can definitely use Docker sandbox if Master is running a Docker container. Yeah, that is a cool thing about Daytona. That even though I think with a So, I mentioned a few times today that sandboxes are ephemeral. They are to an extent, right? It's not the same as permanent long long-term storage. but different sandbox providers have different ways of approaching things. Here's Master's integrations page where you can see all the all the different sandboxes. some will basically archive and be wakeable for an indefinite period. But then you you have to factor that into the price, right? In this case, for example, I was reserving storage that I really didn't need to.
39:56 Okay, so we had sandbox one into who am I.txt. Then we're going to go over to this chart and say write sandbox two to who am I.txt. And because the sandbox is now scoped to the Fred, when I let this run, nice. And then go back to the original file and ask it to read who am I.txt. It's going to behave pretty consistently and in line with what you might expect when talking to an agent in a thread.
40:28 this is probably what I would implement if I was building something like a chat GPT, which occasionally does, by the way, run things in a computer. Okay, let's have a look here. We're about 40 minutes in. We've got a couple of demos left to go. These are the best ones. I can't wait to show you these. The first is Mastra using a browser. So, the title of this workshop is about giving your agent a computer. Sometimes you don't need to strictly give your agent a computer. You can also give it a browser. If all you need is the computer to run a browser, you might as well just give the agent a browser directly.
41:01 There's a few ways to make that work in Mastra. The simplest is by creating a local agent browser. This is basically going to spin up Chrome locally and do stuff. So, let's do headless false first cuz I I think it's illustrative. And go to the browser browser agent. I'll tell you all about this in a moment. So, what we'll do is we'll say, "Go to Hacker News and tell me the top post right now."
41:32 And as you can see, it like literally spins up a browser on my computer. I have to close it manually in this case, although I think there's an option to auto close it. If I wasn't sure about that, what I would do is go to the Mastra docs and ask you know, "How do I auto close the local agent browser when headless is false?" That you should highly highly recommend you use this agent, by the way. It's very accurate.
42:03 so, apparently you would have to implement an on close callback or call browser.close or something. This would be quite good in like a coding agent, for example, or maybe you create a hook for when the agent browser is done and then you call browser.close. But but maybe an even better solution actually is just to not run >> >> the is to just not run the browser UI, so headless to true. And what's nice is that if you're in studio And by the way, my memory for this agent is only stored in RAM, so when I reload the dev server, I lose previous conversations. Go to Hacker News and tell me the top post.
42:45 the Mastodon docs actually never used Mentify. We used what did we use before? Docusaurus. It's like a Next.js docs engine. Nextra, I think, just came to mind. Yeah, I think we used Nextra, but it was very very slow and we much prefer Docusaurus actually. And so, check this out. The browser didn't pop up on my machine, but I can see it here. This is actually using a feature of Mastra called browser recorder. We implemented this in studio using the browser viewer.
43:23 And what's really cool about this is if you've ever used a software factory like Devin from Cognition, if it's working on a website, it can like take a it can open the website in a browser, make sure everything is rendering well, take a screenshot, show you. We're using it here in studio to give you an insight into what's going on, but you could use it for everything as well, which is really cool. But what I'm showing you here is the kind of equivalent of running a local sandbox. This is running on the host machine. If you go to the Mastra integrations and go to Let's see, do we have a section for browsers? Yes.
44:00 you can use like the local agent browser, or you can use something like Stagehand from Browserbase. And what StageHand does basically is give you an API key, and then the browser opens in someone's cloud, right? And so, you wouldn't have to worry about spinning up a computer, you can just spin up a browser. But but what if you do want to spin up computer? How can you do that? Well, that's our final demo, and actually it's something we just released recently. So, you're among some of the first to see it.
44:30 So, essentially computer use is a feature of certain sandbox providers. Gitpod supports computer use, which is why I chose it today. And Gitpod supports computer use as well. And Maestro Integrations work with both of those. The best way to demo this, I think, is going to be to put these side by side and go true full screen. Zoom this out a little bit and go to sandboxes. we're just going to say hello to the agent to kick it into gear.
44:56 And I expect this will spin up a new agent. Here it is. And then just like we could look at the the terminal before and SSH into something into the system. You can also browse the file system, by the way, which is nice. There's also a VNC option. I forget what VNC stands for. Virtual something network computing. So, yeah, it's kind of like a remote desktop feature. Got it. And then here's the computer that's now literally live in the cloud for us.
45:28 I can say something like So, can you see it? It'll be a little bit small for you. Oh, we can open it in a new tab. Even better. Where is it? Cuz I want to show it side by side. You can see here the desktop wallpaper and a couple of icons on the desktop. One of them is a home folder. So, I'll say open the home folder. I wish you could see me right now, cuz my I'm smiling from cheek to cheek. This is so cool. Now it opens the the browser.
45:56 open pictures folder. And by the way, you know, doing this one step at a time is obviously nonsense. The cool thing is that the agent can just call these tools like screenshot and click, screenshot and click as many times as it wants to. For example, instead you know, let's see if they can create a text file using the mouse called hello and write hello world in it. Let's see what happens. Why is this useful, right? Because if you want to do something with an agent, but that agent doesn't have for example, or rather if you want your agent to do something with the service, but that service doesn't have an API, maybe that service doesn't literally work in the browser, you might actually Look at it go. It's literally using the computer on my behalf. That's wild. And it can do more, right? It could obviously use a browser, but we can use browser base for that if that's all we need. But it can also like run programs. So it might be using some accounting software. So many systems in the real world are like windows.exe, WinForms, Delphi, old school applications. You could imagine using it for that. And generally giving your agent a computer to do more stuff like download a file and then inspect it or something, or run it through an existing Like why not, right? As long as it's isolated, now your agent can do anything a human can. And that is extremely powerful.
47:16 Okay, folks. I'm going to stop sharing my screen, put the camera back on. we have to be sure that there's nothing else I want to share on screen because if you were here at the beginning, you know that my screen share completely failed and I'm not about to risk it for the biscuit again. So. Let me stop the screen share here and then ask if you answer a few questions and then we'll wrap things up.
47:46 Okay. yeah, so it could try That's a great idea. It could try like LibreOffice or something. That's a great idea. is it Linux Lite? I'm not sure what distribution it is. why I think it's looks like XFCE. It's totally possible. But it's something quite lightweight. Let's see. Let me know if you have any questions. Oh, It's was got some input here. mentioning E2B into E2B is a good one.
48:21 And yes, Mash it does support E2B as well. I might as well working on a project now that has 100 to 200 users. You decided to use one local Docker sandbox for the access with sidecar container architecture. I don't have a and you're looking for a Docker setup example. I don't have one handy to demo. Unfortunately, you can check out the master integrations and you can look into the documentation on Docker. We should definitely create some videos or some demos around that as well.
48:53 Oh, Reyhan's going to a hackathon from Daytona. Is that the one that hosting in Canada maybe? I think I saw Ivan post about that. And yeah, okay, E2B is a self-hosted option. I didn't know that. Mike can't believe how fast we're adding new features. That's right. We're shipping like crazy in Mash right. Trying to give you all the features you need to build a really good agents. And Vibhav is contributing to Copilot. We love Copilot and Mash and Copilot work really great together. So, I'm glad you're here.
49:24 Okay, folks. Well, thank you so much for coming to this workshop in which you learned how to give your Mash agent a remote computer. You can use it for all kinds of things, right? From building coding agents such as software factories to building like your own like why use open claw or Hermes or something when you can build your own with Mash right. And then give it all these things so that the agent can do things on behalf of yourself with full autonomy and hopefully security and isolation as well. We give you all the primitives, but of course you need to be judicious as well whenever you give agent access to these things.
49:57 If you're looking for the recording, you'll get a link in your inbox if you registered on Luma to both the recording and the slides and the code. Not sure if we we don't have any immediate hackathons planned, but we've hosted a few and I'm sure we'll host a few more. Any teasers on next updates for Master features? Well, did you know that the remote sandbox thing was only launched yesterday? So, we've not even really announced that to anybody yet.
50:23 We're also working on a lot of things around observability right now. And next week we're going to be talking about this idea of agent learning, which is where the agent can look at traces, identify failure modes, and then even fix them while you sleep. It's pretty cool, right? Vibhav wants to connect. Yeah, that's cool. Connect with me as well and book a codes at and book a codes on X. Sorry, Yeah, did I say remote sandboxes? That's been around for a while. the computer use integration was shipped just yesterday, so check that out. New books? Hmm, maybe. We've got two books and a few editions of each. Maybe a third book's in the work.
51:02 You'll just have to come and find us at an event and find out. Oh, glad you're having a good experience building with Master Caleb. That's cool. Book giveaway? It's always a book giveaway. Just go to master.ai/book and you yourself can get a copy of the book for free. That's right, Zandra. Okay, folks. Thanks a lot for tuning in. We'll see you next Thursday.
Summary
- Maestro is an open-source TypeScript framework for building agents, offering tools for local and remote sandbox environments.
- Local sandboxes allow agents to execute commands and access a local file system while maintaining security through containment measures.
- Remote sandboxes, such as those provided by Daytona, offer a more secure environment for agents to operate without risking access to sensitive host data.
- Agents can be configured with "human in the loop" features for task approval, enhancing control over automated processes.
- Maestro supports various file systems, including S3 and Google Drive, enabling persistent storage for agent-generated files.
- The workshop demonstrates how to give agents browser capabilities, allowing them to interact with web applications and perform tasks like a human user.
- New features, such as remote computer integration, were recently launched, expanding the functionality of Maestro agents.
- Future workshops will explore agent learning and observability, enhancing the capabilities of agents in identifying and fixing issues autonomously.
Questions Answered
What is the purpose of this workshop?
The workshop is focused on teaching participants how to give their master agents a computer using the Mastra framework.
What will participants learn in this workshop?
Participants will learn how to provide their agents with a computer, starting with a local sandbox and progressing to remote sandboxes using Daytona.
How can agents utilize remote sandboxes?
Agents can use remote sandboxes by integrating with platforms like Daytona, which allows for isolated environments on the same server.
What configurations are necessary for sandbox tools?
The configuration involves disabling unnecessary file system tools and enabling the execute command tool for specific tasks.
How can agents interact with web browsers?
Agents can spin up a local browser to perform tasks like retrieving information from websites, with options for headless operation.