transcribe

Best AI agent harnesses and how I use them (pi, omp, Zcode)

0xSero · 26m · transcribed 18d ago
More from 0xSero Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Introduction to Harnesses

What are the different types of harnesses and their benefits?

The speaker discusses various harnesses, particularly Codeex, which excels in managing tasks outside of coding, such as email and business documents. They emphasize the importance of browser plugins and recommend starting with open-source options.

  • Codeex is effective for non-coding tasks.
  • Browser plugins enhance productivity.
  • Start with open-source tools for flexibility.
# 5:15

Harness Features and Performance

What are the limitations and capabilities of different harnesses?

The speaker highlights that some harnesses lack built-in sub-agents and goals, making them minimalistic. They discuss the performance of different hardware setups running GLM and the importance of prompt processing speed.

  • Minimalist harnesses can be enhanced with additional features.
  • Performance varies significantly based on hardware.
  • Prompt processing speed is crucial for efficiency.
# 10:31

Managing Multiple Tasks with Codeex

How can Codeex be utilized for long-running tasks?

The speaker explains how Codeex can manage long-running tasks like quantizing and pruning models without constant user interaction. It can schedule tasks and automate processes effectively.

  • Codeex automates long-running tasks efficiently.
  • Scheduling tasks reduces the need for user interaction.
  • Automation can enhance productivity in complex workflows.
# 15:47

Automation and Goal Setting in Harnesses

How can automation improve workflow with harnesses?

The speaker discusses the importance of automation in harnesses, allowing for non-interactive work that can run continuously. They provide examples of using automation for tasks like email checking and project management.

  • Automation allows for continuous workflow without manual input.
  • Setting goals can streamline task management.
  • Harnesses can be tailored to individual work needs.
# 21:03

Desktop Apps vs. CLI for Task Management

What are the advantages of using desktop apps over CLI tools?

The speaker compares desktop apps and CLI tools, noting that desktop apps provide better visibility into active sessions and allow for easier task management. They emphasize the integration of automations and plugins in desktop environments.

  • Desktop apps enhance visibility and task management.
  • CLI tools are useful for parallel processing.
  • Integrating automations can improve overall efficiency.

Transcript

0:00 All right. So, hello guys. Today we're going to talk a bit about harnesses and just like my experience with them, the different types that I've seen available and what I think you could benefit from. So, I have been jumping between all of these different types of harnesses. for example, Codeex is really, really good at doing work outside of coding. So, it's really great at like running these quantization processes, managing my email, managing like my business you know, legal documents, working documents.

0:42 the the one thing is I actually really like these browser tools. So, the reason I use Codeex over everything else is because of the browser plugin. I also built my own browser plugin. So, this was like a long time ago now, but I don't even have it installed anymore, but this is a second wave of I forked another browser plugin, but I use it all the time. I'm using Quen 3.827B and this is using PI under the hood. We've all spoken about Pi before. How do we compare browser plugins? I don't think it's necessary to compare browser plugins. I think you should just start with whatever is open source and work your way back from there and whatever is least disruptive to you.

1:24 So if you feel like you're worried about security, you should potentially focus on having a second browser that you do all this from. if you just want to plug it into what you already have going on, I really like Sitegeist. So Sitegeist is the one that that builds on Pi. We've talked about Pi in a earlier session, but here we go. Sitegeist, you can see this one here, I think, is really good. it it takes over the browser and it does all the things that you need it to do. Just you can plug in your subscriptions. it hasn't been updated in a while, but you can take the like codebase and have any agent drive it like pretty effectively.

2:09 It's very simple and clean. So, I use this all the time. I was using it just yesterday. and I would highly recommend. There's also the chat GPT1 if you're already on Codeex. So, what this does is it puts the Codex harness into the browser. It's just the Codex app server, I believe, that you can it's the Codex app server that you can call from the browser and it can control the pages. It can do stuff for you on screen. And a lot of work is done actually through the browser. filing documents, managing your email, legal paperwork, all of that has to go through the browser at some point. but we have like a few different types of harnesses. And let's go with pi dev.

2:51 That's the link to the site. It's just pi.dev. So, any of you that are interested, it's fully open source. very very small codebase. let's close this GBT thing. So pi is it has a very small code base. It's open source. It is re like it has extensions. So it's built on the idea of the agent extensions or plugins. Agent updates itself. So it it embeds the agent with the pi docs.

3:27 that's part of the system prompt is an index of the pi docs. So it could like you can tell it I want to see the price of this session. in euros instead of dollars and it'll be able to go and like write an extension and then hotwire it into the into the footer of the plugin. It is very extendable. So everything can be built in PI, but nothing really is built out of the box that doesn't need to be there. We go back to pi.dev. it's something I recommend everybody read the repo. I know I've talked about this before, so I'm just going to move this here. But the best thing about this is it's just super cheap to run. Like it has a very high cache hit rate. But yeah, there you go. We have a local model.

4:18 Obviously, this is a pretty big local model. I can have codecs like run just a single copy of Quen on a DGX Spark or use like an Intel GPU. I am so annoyed with how this is configured right now. I think we can just use the model to fix it. hey JLM, every time I run pi, it starts with claude sunonnet or like what was it? Claude heel 4.5. I want it to to always start with GLM 5.3 from the home lab provider. and I want all the HomeLab provider models to be in the scoped models. There you go. And you can see it's super fast because it's like it's already like the model's hot basically and it'll figure it out. But this is the the cool thing about this model not the model but the agent is it's just able to like update itself live and I really like that. Yeah, I like I like GLM. I'm a big fan of GLM. I think it's a great model. but it doesn't have sub agents built in. It doesn't have goals or plugin like what is MCP servers. All that is not in this harness. So, it's super minimal. And then you can add that in if you want it by using their on what is GLM running. So, this is GLM 5.3 and it's running on 46000s RTX Pro 6000s, but I also have the DJX Sparks. I have four of them and they also run GLM. It's just much slower on the Sparks, but GLM 5.3 Flash is decent. And then we have these 3090s running stuff in the background as well. But we go back to Pi. We could see Pi is figuring it out and it'll just like eventually get to the point where it works. in the meantime, I'm also going to open something else. I'm going to cd into X and we're going to use OM. But now we just set up OM. So I'm going to check to see if the model is loaded. It should be. And you can see this is running on four sparks. the prefill is quite slow. So prefill is like the time it takes to pro process the prompt. it's like a thousand tokens a second give or take, but with the 6,000s it's like 2,500. So the time between like prompting and getting it running is almost zero. The system prompt for OM is really long. It's like 30,000 tokens which slows things down a little bit.

6:49 But this is OM. So, OM is Oh my PI, that's what it stands for. And OM is built using PI. So, if we go back to the PI GitHub repo, this is the repo and it's relatively small compared to other things that are on the market and it's it's building blocks. So, I'll refer you to the other video shortly. How are we using with reviewer advisor setup? So, the whole advisor thing like we'll we'll set an advisor up. It's just going to be very Let's go and make Kimmy.

7:21 We're gonna make Kimmy the adviser now. And we're gonna start a new thread. I had a chat with either Claude Code or OM yesterday where I asked it to process all the data of like my my Twitter and to see the advice that it would give based on the statistics of my posts. So like what how much better does media posts do versus text posts like the most common words etc. This this was done yesterday. If you can scan through these sessions and find out what I'm talking about you can use sub agents to do so.

7:58 Yeah. So we're going to send that. Now what's going to happen is GLM 5.3 flash is running also on my local infrastructure. it's in OM and it's going to start spreading out sub aents. Usually that's what OM does. I didn't tell it to do that but or I think I did I did actually tell it to do that but pi doesn't even have sub aents. So with pi you would need to tell it to create like a t-mux session and then have it run pi inside of that. So, it's like a little bit messy, but with the OM, it just runs like in the server itself, and it's doing all the searches, finding all the files, you know, doing all the things that it needs to do. All my Pi, it burns way more tokens, but it just has like a lot more features and things going on in it. I haven't seen.

8:48 So, right now, it's doing a sub agent like task. it's writing like what should be done to a sub agent that it's going to then send out and you can basically like decide what model does this task thing. So, if you had I'm just throwing words out there, but if you had a DGX Spark and like an RTX3090 as an example, one of them could manage the other, like you can have one agent on one that's doing like vision and whatever scouting or coding and then another one that is just like driving the smaller or faster model, what whatever it is that you want to do. So, the harness itself has this built into it like really well and I have found it to be very consistent. Also, if the model starts going off track and like doing something it's not supposed to do or if it starts looping and just going in circles in its reasoning, the harness will stop the agent and then it'll use the advisor which is reading if I'm not mistaken it's reading all the tokens in the in the in the chat. but it only comments occasionally. and it usually comments this theory the model that is your whatever you're interacting with.

9:58 So over here I have the advisor as Kimmy K3 and I have the main agent as the GLM flash that I'm running locally. Kimmy K3 I'm using the API. That's another great thing about both PI and OMI PI or OM is that you can have your subscriptions connected and you can also have your own local inference connected and it's relatively easy to use. So this is what I've been doing for coding. Then we have the like problem of meta harnesses. So if you're using quad code and codecs and all this stuff, it it starts getting overwhelming.

10:33 you could have like multiple apps open at the same time and and that eats up at the performance of your machine. So I've been using like things like warp or herder any any terminal multiplexer for all the coding stuff video editing stuff anything like that and then I use the codeex app for these like really longunning tasks. So right now I am quantizing and pruning GLM and DeepSeek and this is being done by codeex. You can see that the goal has been going for 48 hours so far almost 40 hours give or take and I don't have to like interact with it at all. It just keeps going. Each message is like scheduled to hit your session at a certain time. So basically codeex can schedule to wake itself up every 30 minutes to do something. so I just give them these goals. You can see the goal here. What it's doing, it's doing a GLM 5.3 flash EXL campaign on the MI300X on basically it's doing reap observations and then it's doing EXL3 quantization and then it's going to merge the two together and run KLD and B like a little bit of benchmarking before it launches the models. But it it just goes and goes and goes. it's also really good at computer use and browser use. So, I probably shouldn't be testing this right now, but we'll use the Chrome plugin with Brave browser and look at my hugging face and all of my model cards and just like get a feel for how I've designed it and come back to me with suggestions. You could also use the extension the plugin, the Chrome plugin to go onto the hugging face docs and pull any data you need.

12:26 Use the Chrome extension though. don't use the hugging face plugin. if you have a 3090 and one or two DGX sparks, what would you split the OM with? So, I think if you had like that, maybe having Quen 3.827B for the 3090 like in a Q3 and then having something like, I don't know, Deep Seek or GLM or Quen. but that might not be necessary. You could also have on the sparks like the agent. So whatever Quen, GLM, doesn't matter. and then on the 3090 you can run things like text to speech, speech to text, what is the other thing called like embedding models like you you can use it for quantization and there's a whole bunch of things you can do with a 3090. But you can see here by the way that the codecs like it's scanning through stuff right now. you can see that like it's it's managing my browser. So here you could see that codeex is like managing my screen and exploring to see yeah to see like I asked to look at my hugging face and like how everything is organized but what I'm trying to demonstrate is its ability to do computer use while I am also using my computer. So like I can do stuff and codecs can do stuff at the same time without taking over my screen or like getting in the way of what I'm I'm working on. So I'll usually use codecs for all of that. I I think for Yeah, the agents like the agents are they have so much throughput and output like they just put up so much stuff like they create data, they make the traces, they make a mess, they change commands, set like automated like scripts that run on your machine and completely destroy your machine over time if you're not paying attention to it. So that's something that you guys should keep in mind. But codecs like as a harness is phenomenal, especially the plugins. This makes like creative intellectual work so much easier if you have everything connected and Codex can just like go to your GitHub, see what's open, go to your email, see if there's anything relevant.

14:37 open your Slack and and get messages and and process that. it makes life so much easier. I I don't think there's anything like this right now. There's Zcode though. So Zcode, I'm going to bring this here. But where is my screen? So yeah, Zcode is really really great. so here is my here are all the models that I have running in my house and they're running and I can just select them and I can ask any question.

15:11 how many of these recipes are validated? get me like a matrix of the hardware, the recipe, the configuration, the speed sweeps, like all the stuff that's validated. Nothing that isn't validated or pre-ested. So, Zcode is like the closest to Codeex, I I think the Cloud Desktop app is relatively gotten much better recently, but in terms of open source, Zcode is really solid. It does the loops. It does the automations. it has computer use built in. It uses CUA Tricua.

15:47 yeah, it does the Tricua thing. Now if we open this like you can see it's pretty it's relatively close to codeex. It has browser use. It has computer use. and yeah, it does the side conversation stuff and it lets you bring your own models in there. the automations and goals I think are the most important thing because regarding like automations, most of your work doesn't have to be super interactive. I don't think like a lot of work could just be a goal that's running 24 hours a day or like it wakes up every day at 3 p.m. and it checks your emails and then like adds whatever you're doing into a registry. It's like oh or like a a time sheet, hour sheet, like what are you working on? What pull request do you have open? what what does your boss want on Slack? Like all this stuff could be fed through to the agents and these harnesses. eventually you'll build something that is very usable for you.

16:49 even with Quen, I did I did this with Quen yesterday. So this is an example here. I had Quen open. Please continue and run the video when done. So I use them for video editing a lot. like they're not great. Obviously, like having a video editor or being good at it yourself is the way to go, but it saves me a lot of time. I would not make any YouTube videos if I couldn't just give these types of tasks to the agent. Yeah, this is the video I posted. so Quen, I I asked Quen in the browser like use thing, the computer, yeah, the browser thing we were talking at at the start. I asked it to find me flights from San Francisco to a place called Asakusa in in Japan. I said Asakusa just to see if it'll figure out like what I'm talking about. I should have like I could just say Tokyo, which would make it easier, but I told it to find me a hotel and find me a flights. I wouldn't trust this. Like they make horrible recommendations sometimes, but the general theme is that like this is work people pay for. They pay for people to plan their stuff. They pay for people to manage their emails, to do their accounting, their taxes, their video editing. So all this stuff could be relatively automated. You still have to you still have to validate that what's coming out is actually what you need. But a quen is enough to do to do pretty much anything that you would do. if you're focused purely on coding and you like you know you're just best served by the by the models that like the frontier models if you're only in it for coding. But if you're doing other stuff, this is amazing having a personal assistant that is your own.

18:30 so here we go. It's checking what's happening. It's figuring it out. And you can see this is running on the Sparks and it's relatively fast. I think this is something like 50 to 60 tokens a second right now. And GLM is back. So these are like the main things that I'm personally using. And what I would suggest is do most of your coding in a place that has a visible like file tree. You yeah you want to be able to read the code and not many people agree with this but like having an understand but have something where you can read the code because it's going to make life just much easier if you understand like the shape of the repo. You don't have to understand the repo itself. You just have to understand like the shape of how these things connect together.

19:17 there you go. There's the video that a local model just edited over the last day. And you can see that this is Quen that's doing the video editing for a video that got 380 likes and whatever amount of impressions. But this stuff adds up over the day. So you can have like ideas. So when you're going on Twitter, we go on Twitter for a second. When you go on Twitter, like as you're scrolling, you're going to see some ideas. Like here's an idea, here's an idea. All these are ideas of stuff that you can make content about. And now really what you need to do is just bridge the gap between the idea itself and the production of something and then just post that thing. It doesn't have to be perfect. It just needs to kind of work.

20:03 Simple enough, right? practice makes perfect. So there you go. It figured it out. And bam, the video is amazing. So going back like I'm just going to reiterate and we can I'll open it up to questions but to reiterate we have different types of harnesses are good for different things like there's the CLI harnesses that are really in my opinion good for coding because especially if the CLI is in some kind of development environment where the code is readable like what I'm showing you right here. If I can see the code and I can see the shape of the repo, that is going to make me much more effective at what I'm doing, I think. but having having an environment where you can like modify where you can read the code and modify it yourself is going to help.

20:50 getting that habit. so the CLI tools then like you have the open source stuff which I really like which is PI and OM that's all I use in terms of CLIs. maybe cloud code sometimes like if I need Fable, Opus, but I rarely do. I I use it much less than other harnesses in in general. But so yeah, CLI is good for that kind of stuff and like so you can parallelize more. the desktop apps are really good for also like like if you are working on things that are longterm and you need to switch context a lot, having the ability to see like the session name and the status of the session in the sidebar. I'll talk about Omar in a minute, but being able to see this in the sidebar is going to help you like scale your attention as well. you can jump around. Like I know that these are active threads and I can just jump in and see what's going on in there.

21:43 and then the the desktop apps themselves, they usually wire the agent into other stuff. So like automations and goals and loops and plugins. So that is another benefit of using the desktop apps. currently I'm using codecs and zcode for the desktop apps. And then for the CLI I'm just using PI and OM and claude in the warp IDE or ADE or whatever. about Omari I think this is a really good opportunity for the people that are really interested in this technology to build like an operating system. So the philosophy of the the repo has been just contribute. Doesn't matter if you're doing it with AI or not. Like they have good like they have a good system for how they filter through PRs. But the the TLDDR is it's if you want to shape like how the future of operating systems is going to work. I think that's a really good place to start. because they're open to it. Like you're not going to get anything merged into the Linux kernel. most likely you're not going to get anything to change in many of the different operating systems that we have available to us. So I think that's their real offering and then I'm working with them to create like a singleclick experience to get local AI running on a machine based on some like this is it right here. You can see on the bottom right corner I know it's a little hard to see but there's this Quen 3.8. We click load. So basically we scan, we find your hardware and then we get a pre-prepared recipe and configuration and everything and we load that server up and you know in advance like what the tokens a second is going to be what the capabilities are going to be what the best model to run on it.

23:27 There's only one model to run right now for per type of hardware focusing on quens mostly and then from like the when you launch it you can click to open it in open code or claw code or codeex or pi OM doesn't matter like there's a button to open it in any one of these different like harnesses. So you could see it right there like it's loading done and now you can see all the agents like pi, OM, open code, claude, codex, gro, copilot and crush. and you can also share on tail scale and when you share on tail scale it will take that server that you've deployed and it'll put it like it'll get you a tail net IP so that you can share it with your other machines and devices. I think I think this has a lot of potential and promise and doing it at the operating system level is like very valuable. So, I'm just going to open it to questions now. My experience with Open Code has been it's just like it's probably not for people that are going as intense and extreme as I am. it's actually the first thing I usually install on any computer because it gives you free inference. So, I just install Open Code and have it set up tail scale and all that stuff. And it's actually like a really good onboarding experience. Yeah.

24:43 In in regards to the browser plugins, I think it's the model more than it is the plug-in. So Claude models, Kimmy models, GPT models are amazing at browser use. Quen is getting there. It's actually like decent at this. GLM doesn't have vision. The Flash is like a little too code oriented I feel, but I haven't given it enough time. I literally use everything I told you guys I use today every day. It's Pi, OMP, Claude Code, Codeex, what else?

25:17 Zcode. I'm using those every day. Like, I wish there was just one thing that I could use. Like, I would like codeex, but with my local models and all the providers that I would like, if I could get that, like, that would be heaven. Now, I've been able to make it work, but they always change the code and then it's I have to go in and like figure out how they how they're doing their proxying and all that. And I just don't have the energy for it right now. But no one has made a desktop app that is as good as codeex, but completely like open to whatever you want. I've tried T3 code. My app is built on T T3 Code or like the UI of the app is built on T3 Code. It's pretty good. It's it's all right. I mean, it's still not like Codeex. Codeex is really well integrated into all these like plugins and things and that's what makes it value. Thank you guys for taking the time to join and keep learning. Hopefully this was helpful. If it wasn't, leave feedback so that I can improve the next one. Have a good rest of your day.

Summary

The discussion revolves around various harnesses and tools for managing tasks, coding, and automation, with a focus on personal experiences and recommendations. The speaker highlights the effectiveness of tools like Codeex, Pi, and OM for different tasks, emphasizing the importance of browser plugins and local models for productivity.

- Codeex is favored for its browser plugin and ability to manage tasks like email and document filing.
- Pi is an open-source tool with a small codebase, ideal for building extensions and automating tasks.
- OM (Oh my PI) is built on Pi and excels in running sub-agents and automating complex tasks.
- The speaker uses multiple harnesses, including Zcode and Codeex, for different types of work, such as coding and video editing.
- Automation and goal-setting features in these tools help streamline long-term tasks without constant user interaction.
- The importance of having a visible file tree and readable code in CLI environments for effective coding is emphasized.
- Recommendations include using open-source tools for flexibility and the potential for community contributions to shape future technologies.
- The speaker expresses a desire for a unified tool that combines the best features of existing harnesses for a seamless experience.

Questions Answered

What are the different types of harnesses and their benefits?

The speaker discusses various harnesses, particularly Codeex, which excels in managing tasks outside of coding, such as email and business documents. They emphasize the importance of browser plugins and recommend starting with open-source options.

What are the limitations and capabilities of different harnesses?

The speaker highlights that some harnesses lack built-in sub-agents and goals, making them minimalistic. They discuss the performance of different hardware setups running GLM and the importance of prompt processing speed.

How can Codeex be utilized for long-running tasks?

The speaker explains how Codeex can manage long-running tasks like quantizing and pruning models without constant user interaction. It can schedule tasks and automate processes effectively.

How can automation improve workflow with harnesses?

The speaker discusses the importance of automation in harnesses, allowing for non-interactive work that can run continuously. They provide examples of using automation for tasks like email checking and project management.

What are the advantages of using desktop apps over CLI tools?

The speaker compares desktop apps and CLI tools, noting that desktop apps provide better visibility into active sessions and allow for easier task management. They emphasize the integration of automations and plugins in desktop environments.

© transcribe · For agents Built with care and craft by Gokul Rajaram