Section Insights
Optimizing for Human Understanding in Software Development
Why is optimizing for human understanding important in software development?
Maintaining clear communication between agents and humans is crucial for effective software development. It helps developers stay engaged and in control of their projects. The discussion also highlights challenges in creating development environments and enabling non-engineers to contribute to coding.
- Clear communication is essential for effective software development.
- Challenges exist in replicating development environments for all developers.
- Empowering non-engineers to contribute can enhance productivity.
Building a HIPAA Compliant Messaging App
What are the challenges in building a messaging app for doctors?
Creating a messaging app for doctors involves addressing specific compliance requirements, such as HIPAA. The discussion emphasizes the complexities of real-time messaging features and the importance of understanding the product's functionality.
- HIPAA compliance is crucial for messaging apps in healthcare.
- Real-time messaging involves non-trivial engineering tasks.
- Understanding product features is key to successful development.
Automating Work Ingestion in Software Factories
How can software factories automate the ingestion of work?
Automating the process of feeding work into software factories can enhance efficiency. Insights from customer demos and support discussions can serve as a roadmap for feature development. Regularly capturing these insights is essential for prioritizing tasks.
- Automating work ingestion can improve efficiency in software factories.
- Customer demos provide valuable insights for feature development.
- Regularly capturing insights helps prioritize development tasks.
Challenges in Running Development Environments
What are the challenges of managing development environments in large companies?
Managing development environments in large organizations is complex due to the number of internal services and dependencies. The discussion highlights the importance of having control over the environment to ensure effective development processes.
- Large companies face complexities in managing development environments.
- Control over the environment is crucial for effective development.
- The number of services increases the difficulty of managing development.
Adoption of Local vs. Background Agents
What are the trends in the adoption of local and background agents in software development?
There is a mix of preferences for local and background agents among developers. While some prefer local environments, background agents are becoming attractive for long-running tasks to avoid overloading personal devices.
- Developers show varied preferences for local and background agents.
- Background agents are appealing for long-running tasks.
- Managing developer environments is a concern for non-technical users.
Transcript
0:00 I think optimize for human understanding is the most interesting one. The more you're building, the more you have to understand if you can't streamline that almost like agent to human or factory to human communication to help people stay on on board. You basically just end up kind of like noping out and hoping that the agents can solve it. And if you can maintain understanding, then you can maintain kind of control of your destiny. All right, cool. this was a really fun AI that works. we had Cole Murray who is the creator of open inspect which is a software factory system that is open source that you can deploy. We talked about a lot of the challenges there both in terms of like how do you create dev environments? How do you replicate the the dev environment that every developer has running locally into a software factory seems to be the biggest friction point people hit it is a solvable problem but it is a lot of work. and then getting into like how do we help enable non-engineers to ship code whether it's product managers, marketing folks, sales, finance, etc.
0:55 and then I don't have a visual for it, but we also talked to Tyler Brown. He showed us his software factory on their plan or what is it? It's discuss do compound loop and how they let agents do the planning. How they have an orchestrator that manages work across lots of different systems. and how they do a postmortem every week to learn what was missing from our what where did humans have to step in step in and what's one thing we could add to this software factory to help ensure that the next week humans don't have to be involved there and slowly improve.
1:27 >> Yeah. The coolest part being Tyler and his co-founder Parsa have built a HIPPA compliant chat app for doctors which they've grown from revenue all the way from like 200K to 400K in three months using the software factory approach. A system that typically has no room for errors. Seems to work. >> Super exciting. thanks everybody. Let's let's get into the show. What's up y'all? Welcome to AI that works podcast. I'm Dex. I am the CEO and co-founder of Human Layer. We build a multiplayer coding agent workspace with composable building blocks that meet software factory builders where they are. Oh my god, what a terrible sales pitch. No more pitching. I'm gonna hand it over to Vibb who's definitely not going to pitch.
2:04 >> What's up everyone? I'm Vibb. I'm the co-founder over at Boundary and we make a new programming language that's built for agent native codec in the form of BAML. but today's episode is not about code layer. Today's episode is not about BAML. Today's episode is a recap and one two of the most fun talks that we've had at the AI unconference. I believe we're going to have Tyler on the call and we're going to have Cole on the call and both of them gave phenomenal talks about what it actually takes to think about agentic engineering.
2:34 Specifically, I think we're going to talk about how a a real software factory that Tyler operates for his for his I would say like pre-AII business behind the scenes works and how they kind of automate the heck out of it u from every angle from support to engineering. And then Cole works on a really cool spec called open inspect and he does a lot of work around actually putting out software factories and helping engineering teams and folks build around them in their internally so that you can have a thing like ramps agent inside your own codebase.
3:09 >> Yeah, really excited. Really excited to talk to Tyler. Tyler is one of my favorite builders of all time and yeah, he sort of is the is the is the CEO of a company that has been around for a long time and has just kind of been chugging along >> company, right? >> Yeah, he was kind of running it as a side project. and then actually they just like automated the whole thing with AI and now it's like growing and building and they're like revamping the whole It's interesting cuz it's like it's both like okay cool here's one of the best engineers I know building the thing that builds the thing that builds the thing in a really compelling way and also like here's a thing that every business from 5 to 10 years ago wants to do which is like like you know rebuild themselves from the inside to be completely AI native.
3:54 >> Yeah. Exactly. And I think while we load them into the chat, while we go do this, I I want you all to think about something really important when it comes to building these automated systems. We had this for a bunch of other companies kind of in the in the same portfolio as us. We were chatting for a while and the b biggest anxiety that I saw a lot of people make is how do you make the transition?
4:19 What is the transition looking like from I'm a team, we used to ship some code. I see this AI stuff happening, but we just can't get over the hump. And sometimes you can't get over the hump for a variety of reasons. You can't get over the hump because like you just have a 250 person or you have a,000 person or you have a 10,000 person or change management itself is hard. Then there's a whole another element around Oh, Cole's joined. Yeah, we're joined by Cole Murray, who is the maintainer of a project called Open Inspect, which is an open- source software factory system that you can install on your own infrastructure. really excited to have him on and talk about how he's he's building and what he's learning in the trenches with with, you know, real engineers at larger companies trying to figure this whole thing out.
5:03 And I want to pick back up on Vibv was just making a really interesting point that I think we should finish which is basically just like whether when you're two people it's very easy to rebuild your company to be AI native but like it's very hard when you're 250 or a thousand engineers to like you're not just like it's not the end state that matters. It's actually like what are the incremental steps you can take to get there because if you try to do it all in one go it's going to be chaos. I can share like two processes that worked really well that I think are worth thinking about. one is I absolutely think your best engineers should pause for a month and try and build a software factory. Even if they fabulously fail, you will learn so much about your engineering process just by putting your best engineer trying to automate the heck out of it rather than trying to buy it from the outside or rather than trying to just like convert everyone on day one.
5:58 And the second thing that I think that these companies and us were realizing is that pair programming and kind of like kind of AI pilling everyone but not in the way that's like aggressive and like being like you have to do this but genuinely turning into a collaborative exercise where like what is AI but just chatting with the computer and kind of making progress and it's basically pair programming rather than making someone pair program automatically and like like force yourself to learn how to pair program alone. in a box by yourself and also super high pressure to code like 10,000 lines of code every single day. Make them pair program with one of your engineers that knows how to use AI and make them learn the process over time. I mean, that's how I learned it. Dexter right here is the one that taught me how to do this through a pair programming session, no less of like eight hours.
6:45 >> Oh yeah, that was fun. Yeah. that was that was an exhausting day. but yeah, no, I think I think pair programming now is as important as ever. Not pair programming you with the AI, but like I think there's two things. one is like there's a lot of new skills and a lot of new things and everybody on your team is picking up new tips and tricks and different dimensions and so there's a lot of value in cross-pollinating getting people to work together and then there's the other side is agented coding is a little more asynchronous it's a little more back and forth hey I kick off something and then I go check it out and then I you know I go get a cup of coffee I come back and so if you are solo the the you get bored of the you know you you miss the dopamine because the agents are working for longer now you kick something off and then you're like, "What if I did another one and another one and another one?" We've all hit the limits of our personal parallelism. or you go check Twitter if you're me probably and go shout about how bad everybody is at at agents or whatever it is. We're mostly positive.
7:46 but if you're pair programming, you can just pause. If there's someone else there, it kind of forces you to stay on track, right? You pause and go to the whiteboard. You engage with the problem. you look ahead, you talk about what's coming while the agent is working. So, anyways, we already introduced Cole, we got Tyler on here. Tyler, we were explaining to everybody how cool you are, but if you want to do a quick intro, actually both of you, and then I would love to get into Tyler how you're building your software factory and then we'll do some some chat with with Cole and feel free to jump in and interrupt people and all of this.
8:17 >> Awesome. Well, hey, I'm super excited to be here, guys. I'm huge fans of the show. Huge fans. And what what a treat it is to to be on here with the legendary Vibex. my name is Tyler Brown. I run a product called Bloom Text. it's a clinical communication platform for healthcare professionals. we serve about 18,000 patients a month. right now we've had over 30 million messages sent on the platform since launching. Love these guys. We just talked about software factories recently at an unconference. I know you guys recently covered it as well. So, you know, I don't want to cover things that you guys have already covered, but happy to kind of share some of the processes that me and my co-founder are using to build Bloom Text and move super fast. I'll also mention that we have a suite of kind of like products kind of like open source products we use to on the software factory side that are mainly for internal use, but we also kind of share them as well with others.
9:14 >> Awesome. Awesome. Yeah, I would love to I know you have an Excaladra or whatever the evolution of it is. I would love to just have you talk through and we'll jump in and ask questions and interrupt you I'm sure as is tradition on this show. but yeah, I would I would love to learn kind of like what you guys are doing today in your internal Oh, first of all, like how how big how big give a sense of like how big is the team and and and it's >> what's the company size? How long has it been around?
9:40 >> And like if you want to show for like two minutes just the product really fast so we can all anchor ourselves into something a little bit more tangible along the way. And like if you're down to share, how much revenue do you make? >> Okay, cool. Yeah, so it's actually this is going to be a really interesting one. So, it's just me and my co-founder that are working full-time on this. so, two, he's in his 20s. and he's he's sort of what I would call the more AI engineer and I'm more of like the traditional software, you know, traditional been around for 10 years engineer. So we sort of like butt heads a bit on how to do our software factory and I feel like the best of both worlds kind of emerges in our process and design. So I would say he goes a little too fast, I go a little too slow and then we end up somewhere you know the optimal middle ground. Revenue wise we're at we just grew the product from around 270k arr to just under 400 in the last 3 months. separate podcast podcast episode can cover what we did on the growth side to achieve that. but you know software factory is is a part of it but obviously that extends into the business side as well which so interesting interesting side note we can talk about later there but yeah so so that's a quick quick intro. Let me show you the product really quick. I think this is actually a good example cuz you guys will probably recognize this as this is just a messaging app. So I feel like it's very tangible to share. A lot of people, if you've ever tried building a messaging app before, it might sound pretty simple until you get into the weeds and you realize there's a lot of like non-trivial engineering tasks to go along with real-time instant messaging.
11:22 there's a few features I can show that we recently implemented such as message replies, reactions to messages, things like that. and how we approached it from our software factory side. But this is basically what it looks like just to kind of anchor everyone's understanding of our product here. >> So it's kind of like a it and just that from what you were showing me before this it's basically a chat app for doctors >> exactly >> like why do they buy it?
11:52 >> So they buy it because they can't use Microsoft teams they can't use Slack. there are like HIPPA compliant versions but they're >> HIPPA compliant chat that is HIPPA compliant. Yeah. Exactly. Exactly. It's it's HIPPO compliant. It feels like WhatsApp, but it's, you know, you can use it. Okay. So, you know, everybody knows sort of what a software factory is at this point, right? I don't want to like repeat stuff that I think people are already kind of wrapping their heads around, but I feel like it's helpful to kind of start on the, you know, fundamental side. And my goal really is just to like if you guys can all walk away with like one new idea or thing to implement in your own software factory, I'll consider this a success. but this is kind of like how I think about software factories now.
12:36 specifically with the latest iteration of you know agentic coding models because I think a lot has changed and you know initially I thought that you know as models get smarter and more capable we as engineers would get less important right and the stuff we built the harnesses around it would become less important right Boris was kind of alluding to that when you first released cloud code like all of a sudden you don't need cursor anymore you can just use an agent in a command line loop Well, the inverse is true. As models get more capable, you need a better harness and you need more sort of checks and balances around it to make sure that if it goes off for 3 hours, it's not going to come back with something that you throw away. So, this is sort of how I think we should be thinking about software factories in the new world. So before sort of this latest iterations of models you know a lot of our plans were iterative but now that they're so capable we sort of switch to declarative planning which means we we kind of create the end state and we let the agent figure out how to get there.
13:39 That's a huge change in kind of how we approach our planning process. so before you know the bottleneck used to be on the code review side now it's really on the planning side. another big change and I can talk through all of these as well is, you know, a lot of people will work inside the harness before these latest iterations of models and now I think we're at a stage where we can start orchestrating the harness and I know you guys talked about this a little bit on the last episode that you did on software factories. and so, you know, again now like human understanding is sort of the bottleneck in terms of like do I understand what I'm doing with this large plan? and so that's kind of what we optimize for with our plans. and then, you know, we try to create our skills in our factory based off of kind of like tested benchmarks rather than vibes.
14:27 and we're more concerned about cost now just given how expensive everything is getting and will continue to get. and yeah, so so those are the those are kind of like the key fundamentals. Is anything does that kind of track with what you guys how you guys are approaching >> software factories or is it >> optimize? I think I think optimize for human understanding is the most interesting one, right? Is you know I've always said I stole from Jake at at Netflix was the you cannot outsource the thinking and Karpathy basically jumped in and said well you you can outsource the thinking but you can't outsource the understanding. I think they're two sides of the same coin of this idea of like if you don't understand the system you're building and you like the the more you're building the more you have to understand if you can't streamline that almost like agent to human or factory to human communication to help people stay on on board. Like you basically just end up kind of like noping out and hoping that the agents can solve it. And if you can maintain understanding then you can maintain kind of control of your destiny.
15:30 >> Yeah. I'd like to hear too a bit more diving in on the planning review that you guys are doing. I think I see a lot of teams shifting to that as well. And it's kind of a unsolved problem. I see a lot of different teams whether it's using designated tools for this, whether it's just dumping and doing actual like GitHub code reviews on the plan itself. but would love to hear how you guys are approaching that.
15:56 >> There's a really nice slide on here. I'm I'm going to kind of push Tyler towards showing that earlier than perhaps what he planned, which is show us your flow. I think you have a flow on here and Excalibur always like this is what we do during planning. This is what we're doing during during execution loop. This is what we do during XYZ. >> Yeah, absolutely. So, so you know, just before we get there, like I would say this is kind of the thing I think about a lot, which is kind of how do we speed up as much as the work as possible while still avoiding slop, right? That's kind of how we approach this whole thing, right? there's some really I think we already covered this, but this just kind of gives you background on Bloom text here. okay, so this is sort of our plan or this is sort of how we approach it. So this actually looks like pretty simple on the surface, but there's a lot of depth to this. So you know, basically the main takeaway here is there there's three sort of phases to our development work. The first is sort of figuring out the plan.
16:55 And I would say that's kind of the art currently is is creating really really good plans. and that's that's sort of what I would consider to be like pretty ad hoc still. and then there's this phase called the do phase which is almost completely automated at this point where we just sort of hand off our plan to a harness or a factory and we let it do its thing and if everything goes well, you just merge that sucker in right away. however if there's anything wrong with what came out here then so we basically will still do a human review at the end of every you know every PR or every feature but ideally we look at that and we say great but if there's any issues then a postmortem is generated that we then create a GitHub issue for and then every week we look at all of our postmortems and we upgrade the whole harness based off of them and I can show you some examples. Can you show us one of these postmortems?
17:50 >> So wait, before we do that, I want to check one more thing. >> I think and I want opinions here from Cole and you Dex. >> Is this the loop? Just like zoom out for a second. If I am watching this and I want to go build a software factory of some kind, is this the loop that I want to have? This is some sort of planning stage. Then once the planning state is checked off, go to an implementation loop and then there's like some sort of like just like post-processing kind of thing that's just always going to be manual for some steps.
18:24 >> Cole, what do you think? >> What's missing or or how does this map on to >> a macro layer? >> Yeah, I think at a very high level, this is more or less kind of the standard process I see across teams where we do some planning up front. We're creating design docs. how they actually get reviewed I think really depends on the team the level of scope at that point then I think obviously like how the implementation gets done can be a lot more granular right we can like fan this out over multiple sessions we can have multiple agents kind of collaborating on this but at the end of that we're going to have some type of verification process and I think that that varies a lot depending on what you're actually building it could be as simple as integration and end to end tests. It could be, you know, much more formal than that. It really depends on what we're building. But, at the last part, I think, is the review part is probably the more interesting part where I'm seeing teams diverging. And some of that is due to regulatory issues. Some of that is that they're, you know, willing to move fast. they're kind of automating code review now. But I think at a high level, this is kind of the the macro that I'm seeing.
19:36 >> Yeah. I think I guess maybe one missing piece specific to software factories. There is kind of like a a nonhuman in the loop aspect that is not covered here which is like scheduled automations, triggers, things like that. But in terms of kind of human shipping code, that that tracks. >> Yeah. And just just zooming out to the other one, I think I think there's you you don't have on your on your Can you can you go back up to the other picture? I don't I don't see on here a like code code review stage unless the the human reviewer is basically this is like when you have pull request problems those all get logged basically and >> I think verify imple Tyler.
20:17 >> Yeah. Yeah. So this is a good call out. Verifiers in my opinion are the most important thing now. like I sort of consider like code review anything that you can't verify, right? And the goal is to verify as much as possible. and and give the LLM verifiers to use so that it can loop through using them. so there's a lot of different ways you can do verification, right? And I'm sure you guys everyone listening is familiar with a lot of you know like linting verifiers and you know having verifier agents that look at the plan and make sure everything was done correctly. you know, a a lot of the work we do right now is in making sure that when an agent creates a plan or implements something, it has all the data and information it needs, right? In order to >> verify what it's doing.
21:09 >> Dex would call this back pressure >> is the word that used. It's like you're adding >> back on. >> But yeah, we >> done many episodes on back pressure. there's a new layer here that I haven't actually showed here in the diagram. but I would call this the orchestration layer and I think this is the new kind of frontier to add on top of this. So what you're seeing here is a very traditional process that you would go about using for like a big feature. But in production a lot of the stuff we do day-to-day is you know bug fixes, little things. And I would say this can be a bit overkill for that kind of stuff.
21:47 and one of you guys were mentioning earlier like I can only juggle so many things in my head and so I think like how we kind of run our orchestrators is the new frontier and I can dive into how we do that now if you guys want to see that. >> Yeah. One one other thing that would be interesting to hear about is too is like this is this is the like the the production line. But the the other question I think I see in some interesting software factories is like what feeds work into the factory because obviously we have the normal like humans deciding what to build, talking to users, reading bug reports, but the other interesting thing I'd be curious to hear about is like how much how much are you automating the like ingest of I know I know Pearson did a little demo also of like okay he checks post hog screen recordings once a once a week and like pulls in any like UX friction in like how do you detect things that need to be on without having to sit in Slack or sit on support channels. How can you how can you automate feeding that in?
22:42 >> This is such a good question. So, let me show you a really like simple version that I encourage everyone to do. So, each day I have all of these insights >> that I write in notion, right? >> So, each day like all of these ideas are just coming from they're like the exhaust that come from working on your business. So what's cool about BloomTex is we do an average of I would say you know eight customer demos a week and these customer demos are so rich for uncovering new features and for understanding kind of like what to build next. so when you combine that with all the customer support discussions we're having I mean like it's it's a road map for building as much as you want like in indefinitely. Now, I think you can get more sophisticated with how you capture this, but you know, this is something I think anyone can do. You're literally just taking notes on all this stuff. And so, in a given day, I might have like seven things I want to implement, right? How in the world are you supposed to get to all these things?
23:42 one thing I haven't mentioned is that we basically maintain four products simultaneously. So, we have Bloom Text, which is our main production app, but then we have three internal kind of apps that we that we're also using day-to-day for all of our work. And those are more vibe cody apps. but like how do we just kind of like get all the work done across all these things? And so what I do is I actually grab all these and then I throw these into what I call an orchestrator chat. And through the orchestrator chat I'll actually have an agent that manages multiple set multiple sessions of these discuss doops all through one chat.
24:20 >> And that's kind of the secret sauce is you copy your whole day's worth of features. You paste it into a chat with an agent. I can actually show you what these chats look like. for us, we use linear now as kind of our software factory interface. It's not great, but it works for now. and we have an agent cloud code session that just kind of like looks at linear and it assigns agents to linear. It creates sub issues, too. you know, for like overall issues, >> right? So this is the exact pattern that Antonio on our team found where like instead of actually managing linear tickets himself and like the whole discuss you call it the discuss do loop we call the plan implement loop but the same thing he's just like for small tickets my only job is to triage things correctly as small and if I triage things correctly as small I just let the agent rip and I just do a final pass check at the end and it just kind of drives it the whole loop end to end to completion. Is that kind of what you're doing here where like you've triaged it to guaranteed small so you trust an agent enough without him to read all the code? It can only do so much harm. Is that what we call it? You're you're bounding the harm.
25:26 >> Yeah. Yeah. Totally. I mean, you still need to kind of dig into the weeds. And so, it's going back to this question of, you know, how do we speed this up as much as possible while still avoiding slop and stay in control, right? >> And so, if you've got an orchestrator chat, more things can go wrong. I think everyone needs to design an interface for this. And we're still experimenting with the right interface. Can you guys see my screen? I just I just switched applications.
25:49 >> no, we only see the Excalibur. >> You see the browser? >> Okay, that's my bad. >> it's funny. We're going to see you have to share your whole screen. >> So, you know, my co-founder created his own agent IDE. It's it's I don't know why like it's it's very similar to what everyone else has, but what's nice about it is you can always add a feature if you're missing anything. So we have this orchestrator chat called pain chat which actually has the ability to manage and chat with all these individual sessions.
26:17 So this morning I actually had it create like 1 2 3 4 5 six PRs all in a row and it's it's keeping them running under the hood and it's like you know surfacing like whatever is needed. And here's another one. >> Is that running locally or is that running like connected to a cloud environment? >> This one's running locally. but we do have a whole Damon running on an environment that we mainly use through linear. So, this is a great example.
26:46 Customer reported a bug this morning, right? I uploaded all the I uploaded all of the issues u or all of the the chat transcript with the customer. And check this out. This is from linear. I know this is kind of hard to read but look what it did. It created four sub tickets in linear and then assigned agents to each of these. And these are all working on our Damon, which is a Mac Mini. and if you look here, I know this is a lot to read and definitely don't read all this, but basically what I want to what I want to show you guys is this right here. It had access to Sentry. It had access to the transcript of the customer. it was able to look at Post Hog. Like, it was able to kind of get to the root cause because we gave it all of the necessary tools it needed. And so these fixes were kind of verifiable root cause issues.
27:36 that were you know based on those. So boom, there it goes. It creates all these issues right here as tickets and then each one of these is a separate and you can see some of these are already done. Some of these need more more work. This is just kind of a faster way to work especially on the small things. >> Okay. And so one of these gets picked up and then the planner agent picks it up get it gets assigned to the planner agent and then when these plans are approved they'll go to an implement agent.
28:05 That's right. So, it's going back to that same foundation that I showed you guys in this diagram. The only difference is you have an agent on top managing this now, but it's still not automated because that agent on top is is actively triaging the work. It's surfacing. Hey, agent number one has a plan ready for you to review. Agent number two is like you know, just finished and is ready to show you the prototype. So, we're basically going to like a Tinder style swipe left, swipe right, you know, add more comments, style interface. that's, you know, that's kind of where I see us headed with these orchestrator style, management interfaces.
28:46 >> I know you had more to show. I think we're probably almost at time just so we make sure, but, you know, if you if you want to do a deep dive recording, we have done some like longer longer form stuff sometime, ship a feature end, etc. But yeah, Cole, do you want to talk a little bit about Open Inspect and and and the work you're doing? And I'm I'd be cur I'd be most curious to hear kind of some some stories of like, hey, we built it this way cuz we we thought this was what people needed and then we got in the trenches with cuz I know you do a lot of like consulting with people helping them set this up of like what was a surprising thing that you found and and and how did you actually have to like that like was like non-intuitive that you learned after like watching people work.
29:27 >> All righty. So this is just the open inspect interface. I thought we gives us something to look at while we're chatting. but so the short of open inspect is that it is a background agent system. it covers the control plane. It has support for various different sandbox providers. the main ones right now are modal, Daytona, E2B, open computer and there's a couple other ones that are in progress. And then for kind of interfaces, how you can interact with the system. It supports linear GitHub. it has the web client that you see here. and Slack as well. Can I pause really fast because I'm sure there are tons of people that have never ever heard of the word open inspect. So before we go into that exactly like what is the problem that it's trying to solve and why can it not be solved by other things? cuz you said a lot of it seems to talk all the way from compute talent down to like tickets. It's spanning a really big spectrum. Help help me catch up on that really fast. Like what what is the core problem?
30:30 >> Yeah. the core problem we have is that agents are starting to take on very longunning tasks and you want somewhere to push that work ideally off of local host. And a system that you would build for that is typically called a background agent or a cloud agent. Simultaneously within companies we're starting to see this trend where we want to start centralizing systems. We want to start creating kind of core agent infrastructure. whether that's for you know Slack bots, text to SQL bots.
31:02 Go ahead. >> sorry I'm going to take over the screen because I think this would be really useful to draw. so I'm just going to share I'm going to share this. I started I started trying to take notes on what you were what you were saying but I also I sent you a link to this if you want to help draw on it otherwise you keep talking and I will just draw. >> Cool. yeah and so we have this problem basically that we have agents prior to kind of cloud agent systems coming about. These were primarily running on developers laptops and we have a lot of problems that have kind of formed ahead that need a different solution to this and I can dive into kind of the different personas each of their problems but feel free to let me know what you want to hear. I think what I'm hearing, Cole, and correct me if this is not the right words. If I want to build a software factory for my internal team, I have to plug in a lot of systems together all the way from compute resources to systems of record to agent like like surface area of how I want to communicate with them. Whether it be text messages, Slack, whatever the heck I want it to be. Open inspect is some sort of opinionated stack of how to go combine all those together.
32:11 >> Yep. Yeah. More or less. >> Right. And it's like I think Tyler what's in what's gonna be interesting here is like Tyler has built one version of this internally. We've built one version of this internally. >> Dex is also building something like this in a very interesting way. So I'm I think we're going to get into a lot of really interesting details around like the hard parts about this problem. That's probably what I'm most fascinated about especially given your background of like how you work with so many companies of different sizes to go solve this problem.
32:41 like what are the challenges? >> Yeah. I think the orchestrator and the control plane is actually the easiest of the problem. There's a pretty st standard set of primitives here where we take in some user input wherever that comes from. It could be from Slack web wherever. We then need to create agent sessions and basically launch compute to run those. whether you run the agent in the sandbox or not I think separate discussion but basically we want some way to invoke the agent and have a log of what has actually happened with that agent and >> that word comes up again log an agent is nothing more than a log.
33:24 >> Yeah 100%. and so with that then you know you can make a lot of different decisions whether you want to build or buy this but I think the actual hard part here of a system like this is what is actually running in that compute because that is running your dev environment ideally assuming that we want to verify our changes and running that in that compute especially if you're a large company that has a lot of internal services there's kind of a lot of goo to get through there and if you don't have control of that environment it can actually be pretty difficult to do that because you're you know kind of working in a black box and I'm sure people who are familiar with just like bashing the CI right of like submit a fix did it work no >> I'm going to ask you a quick question just again contextualize everyone >> large company means different things can you anchor us on what you mean by large company is it a team with a 100 engineers is it a team with 100 orgs.
34:28 What are we talking about? >> Yeah. with Open Inspect, the biggest team or I guess org you could say that we deployed it to is 500 500 engineers. >> but I think the people count as maybe less of the problem. I think it's more the surface area of your company. As you have more services, as you have more kind of internal dependencies, actually getting a system that is, you know, mirroring that environment becomes increasingly more complex because you have additional services and just complexity that you have to deal with.
35:02 Yeah, we work with a lot of teams where it's like they have, you know, hundreds of engineers, 10-year-old codebase. They have some giant monolithic apps that need very specialized compute just to be, you know, 60 gigs of RAM just to be able to run the app. And then they also have this problem of like there's hundreds of repos and one I need to be able to run an agent session with five to 10 repos checked out. I may not be making edits to all of them, but like the agent needs to know how they work for context or whatever it is and be able to search the actual code and see what the interfaces are. But the bigger problem there is like I may not know which repos are relevant when I start working. If I just drop a an alert in or a page that I got and say like, hey, here's a problem like well I don't want to check out all 200 repos every time I want to do work. I mean this is a thing some people have done and actually seems to be working okay. but it involves like every hour you clone every repo onto a base image and like bake that image and then send it some like store it somewhere so that when people check it out it's like okay we will do a git poll but hopefully not too many things have changed cuz like you also want fast boot times right if I ping a slack bot and doesn't start working for 5 minutes like that sucks.
36:17 Yeah. And that is exactly how open inspect approaches that is that we can pre-clone all these sharing and I'll let you I'll let you go back to demoing or you have the whiteboard up if you want to draw more. >> Yeah. And so you can pre-bake those images. We have the concept of a environment and so if I go here but so I think yeah you can make it seems like you can kind of configure the environment setup. You can set up Docker containers I suppose of like how you deploy everything on stack. I suspect like the whole stack setup, but when you first talk to a company, what percentage of time do you spend on what concept?
36:58 >> So like of I think you said the orchestrator is the easiest part. Sounds like you're like don't spend any time on that. >> most orchestrator is probably the least interesting of the problems. >> Okay. >> So where do you spend the time then? like when you work with them and like let's say you have one week or two weeks to build a software factory and that's all the time that a company's giving one engineer to go build it.
37:21 Where should they spend the time? >> Yeah, I think the dev environment is where you're going to spend the majority of that time getting that set up is kind of the precursor to having valuable work. The phase two then is actually connecting it into the company. And so that could be the production logs, production databases, kind of any context that that agent needs to be able to answer the tasks that you're giving it. And I think then, phase three of that is like, okay, the organization problems that we're about to create.
37:56 a lot of companies now are starting to want non-technicals to be able to contribute. And so now you have marketing, you have sales, you have a lot of people who don't necessarily have the technical expertise at kind of the agent level, but we want a system that they can still utilize to be able to make changes and achieve whatever outcomes they're looking to work on. And so that's where a lot of the time is actually spent of okay, now that we've deployed this system, which typically takes about 2 to 3 weeks to get it kind of set up and integrated. let's start addressing the org problems with this. And depending on the company, there's various different directions that we go, whether it's kind of engineering driving this, whether it's marketing or another team driving it. but eventually they all kind of end up at the same place. So you said dev environment you spend a lot of time on. Can you help us understand what that means? Like when you say you spend a lot of time on this, what are you configuring? Is it like Kubernetes clusters that you're configuring? Are you like, "Oh, you don't have a sand, you don't have a staging environment yet. Let me actually build you a staging environment release pipeline for the first time because you've never done this before even though you're 500% or which I have seen before."
39:13 >> Yeah. Yeah. I mean it very much varies depending on the company. the best case scenario they have a local environment that they can set up whether that's with docker compose whether it's with you know vagrant or whatever other tools they decide to use that's a very easy case the other cases are where they don't have that or they have you know some archaic like handrolled scripts that is how everyone sets everything up and in those cases then we're kind of standardizing how to get that set up the kind of remaining steps.
39:46 Then especially teams that run a lot of web applications, they often don't actually have a like good local reproducible environment that the agent can operate in. And so as an example, in almost every case, the web application has some type of login which is going to be ooth based. Well, if your agent now needs to verify these changes, we have a problem here because the agent is not going to be able to, you know, Google login. I guess it could, but like probably not.
40:18 >> and so we need to then build bypasses into the system that allows the agent to actually be able to test the environment so that it can verify those front end changes. >> When you say dev environment, you're saying a lot more than what I think most people mean by dev environment. I think when I think about dev environment, I'm like, what's the machine I need to set up so I can start writing code on the system? And both you need the you need the language runtime, you need the you need the actual like code checked out, you need network access to any like because most people I we talked about this last time was like most people if you have 200 repos I'm not going to run all 200 just to make the thing work. I'm going to run three or four services and then point them at some shared pool of like developer resources that they're going to use to like the database is going to run over there and I'm going to use the same shared data set as all the other developers.
41:10 >> But the caveat kind of what Cole was talking about is like oh I need login built for my agent. >> That's that's usually not included in what I would I think 90% of people think about when they think about dev environment. And I think what cool you're really saying, >> what you're saying about like the this hardest part of the system is actually not about just dev environment. I might want to use a different word for it and tell me how you guys feel about this. I want to design the fastest testing iteration loop. And part of that is everything I need to do to make sure that my stack runs locally. What's that one thing that they built? SST. Isn't that why SST is so freaking nice?
41:51 thing. >> what was it called? I think it was ST, right? Where like you can >> make open code. >> It was a rapper on Palumi to let you run your serverless functions locally. >> So you Exactly. Cuz like you can't simulate AWS. So your dev cycle is like write the code, push to AWS, go make it run. It's like I'm I'm basically like waiting minutes every time I make a change. I use SST. I'm waiting seconds.
42:14 And like as a human, it's just annoying of why I do this. as an agent, you're to say any other word. If I'm waiting 5 minutes to test the change, it >> it's a clever interface boundary of like, okay, the lambda is a thing where okay, I put that locally versus running the cloud. I think what Cole said about like bashing CI to get it to work is really interesting. Like the iteration is so much slower because like the environment the thing runs in in the CI worker is so different that I can't actually test it locally. but I think what you're talking about like I don't know, we have a we have a contractor who is in a different time zone. he's awesome. and he's just like, "Oh, I'm trying to get the dev stack set up and I can't I can't get it working because I don't have like the keys for work OS or whatever so that I can create like a proper O environment." And he just hacked it. And it's ex exactly what you're talking about. How do I rearchitect my application so that I can run it with as little like upstream exterior dependencies as possible so that it's not complex for the factory to set up and so that the agent can use it and test stuff? He literally mocked out all of work OS in our app so that he could I mean I could have just sent him the keys but it was like it was you know 3:00 a.m. in the US and he was busy so he wanted to get he wanted to get he wanted to get unblocked. Hey guys, sorry sorry to interrupt, but I have to jump.
43:32 we've got an Airbnb we need to leave at now, but oh my gosh, I have so many hot takes about this and I don't want to just drop a hot take and leave. So I will I will withhold. >> No, give us one. Give us one hot take. >> Okay. sandboxes are the wrong primitive for developing applications at large scale companies. >> all right. >> I think that you should give each employee a Mac Mini instead. Pets versus cattle. We talked about this on the last software factory episode. Do you give everybody a bespoke box that they own or do you Anyways, I know you gotta go. we can go deeper. That's >> Yeah, that plus give people dev environ like make sure the dev environment runs exactly the same whether you're local or remote.
44:15 >> % agreed. >> Ian Livingston key card is a key component in that. like routing solutions are a big deal or Yeah. Anyway, sorry guys, I I got to go. But >> >> Super fun. Thanks for having me. And Cole, it's great to meet you. Great to meet you, Cole. >> Meet you as well. >> Okay. See you. See you guys. >> cool. All right. good riddance. that guy, that guy talked way too much. just kidding. We love you, Tyler. cool. Okay. well, we could dig into that or we could keep going on this like dev environment idea.
44:46 I don't know, Cole. What What are you seeing? Are you seeing people wanting to have the option of local and background agents at the same time? Like what does the adoption look like? Are people trying to move everything to background agents or is it just the small stuff that we talked to triage the small stuff and have that go in the background? What are you seeing? >> I see a mix of both. I don't think that you need to you know kind of run everything on the cloud but I do think that it is becoming increasingly attractive as we have longer and longer running agent sessions where you know if I have a 6-hour session running I don't really want that on my laptop because now I my laptop cannot be turned off and you know I have to keep this plugged in overnight and that's a pain. but I I think it really depends. I think for the non-technical users of the company obviously they don't want to maintain a developer environment and a lot of teams who have come across open inspect found it kind of through looking a way to solve that where you know they say hey our PM they want to contribute we don't really want to maintain a dev environment for them you know they're asking like what's docker how do I update this and having a system like this allows that to be completely abstracted way to where they can still operate within this. They can still preview the changes, but they're not actually having to maintain that environment.
46:08 >> And do you find people are actually getting like PRs merged? I mean, obviously like copy changes is easy, but like I what I've seen a lot is like PM will vibe code a feature, give it to the engineers and be like, "See, this is easy." And then the engineers see the code and they're like, "Absolutely. Will we never ever freaking merge this?" like what what is what is the what is I've heard people say things like oh have the PM create the PR as a like prototype and then the intent is like they will close the PR and the engineer will rebuild it properly using that as like a reference of like what is the experience but like I don't know have you have you what have you seen in teams that are like successfully getting nontechnical people to contribute meaningful work is this something that's possible today >> yeah very much so and a lot of this is kind of what my consulting goes into I would say You know, obviously there's the setup process, but things like this, I think, are more organization and, you know, how do we get a company to be quote unquote AI native? And I think a really good way to start with this is for the PMs first, just taking on very small bug fixes.
47:12 you know, you get a customer complaint that comes in and hey, what this modal doesn't close or like, you know, whatever the issue is, let the PM take that on rather than creating an issue and giving that to engineering. Just let them run with it. let them create the pull request and most cases when teams start out the quality of the code is going to be off something is wrong you know there's issues that we have to address and so that part of the iteration I think is very critical to get early on when the scope is still very small for the PMs of just these little bug fixes and that could be adding skills it could be adding documentation llinters I think you know really depends on kind of the issues you're seeing But as the code that the PM is producing becomes more trusted, we can then let the scope of what they can actually ship increase. I think letting them, you know, create proof of concept vibecoded things is great, but the expectation should not be that this is going to go into production. you know, it's a proof of concept. It's great. It allows us to organize around whether we even like this as a concept.
48:19 >> I think that makes sense. I I think I would tie that even back to Tyler's thing of like the compounding stuff like our job is not to you know decide who who does and doesn't get to contribute. Our job is to like let non-technical people try and then figure out how do we need to evolve the guard rails or the skills or the prompts or the llinters so that the next time someone wants to be productive they have they have more feedback and the chance with whatever skill level they already have.
48:51 >> Yeah. >> I mean if it's no different than how you do with engineers too like you onboard a new engineer into into your team even though they're technical they're going to ship a freaking bug. How do you prevent that? Well, you add processes around that system so that like they're good and now you're basically onboarding. You can think of it as you're onboarding an intern onto your team, but this intern is never going to learn how to code. So, you just have to build safety guarders and like these bumper lanes so that they I mean, don't let an intern do a database migration.
49:20 Like, just don't make that possible, >> right? Like, and if you if you do make it possible, like, I'm sorry, you had a process error. Don't blame the intern, blame the process error, patch that process system. Make sure there's no other similar processes at that time that are going to have the same other thing. And like just deal with it. And if you can't support that, then like don't bring interns onto your team, you know? Like that's kind of how I think.
49:47 If you don't want to pay the engineering tax, like bring an intern on was also a process error. Like don't add your marketing team to your database if you don't have an engineer available to support like let's say 100 engineers on your marketing team. Like some percent of your engineering resources have to go towards fixing mistakes that these people are going to make cuz like it's it would be an insane expectation to be like your marketing team will not make any mistakes ever and you're going to build a perfect software factory that has no loss.
50:14 >> Okay. But what is what is different between now and and and before? I mean, I think the analogy of like, hey, you bring on an intern, you better have be ready to assign them a mentor that's going to help them get better and help them like not make mistakes and things like this. You hire an intern with the expectation that that person will become more skilled and will learn as they go. And right now, we're kind of saying like, okay, I'm never going to expect a marketing per a marketing person, even someone who's a little bit technical, to like understand the dangers of database migrations. And so rather than expecting the human to learn and get better, we are forcing ourselves to maybe for better. I'm not saying it's a bad thing, but like we are forced to put all of outsource the learning from the people and put it into the the factory instead.
51:00 >> I don't think it has to be kind of a a hardcut guardrail, right? In the same way that we do allow the intern to eventually grow in their capabilities and we give them a process, right? We start them very small, a very easy scope task. Okay, great. You achieved that. Okay, let's take on a slightly bigger feature. And as they show continued progress, you know, we give them more scope. And I look at the non-technical the same way in that, I think you said it perfectly Vib shouldn't really matter who's prompting the agent because the system should be able to accommodate whoever that is. And you know, you could formalize those guardrails and like you could put in, you know, a a status system, right, where like the pull request or the agent just blocks you if you don't have database certification within the company. obviously I I wouldn't recommend doing that, but you know, you certainly could. And I think >> yeah, >> as people become more familiar with this, I do see a lot of PMs learning more technical concepts to where they're very interested in these things and they're not trying to learn them, you know, at the same level of the engineer, but they understand the stack as their familiaress with the system grows. Yeah, I think that's a great nose to a note to kind of close this out on which is like if you're going to go and build a software factory of some kind.
52:27 One, like there's there's a lot of different components here and like different folks here may have different opinions on what the hardest part is. in Cole's opinion, the hardest the easiest part is orchestrator, the hardest part is setting up the fast testing loop so that your agents can iterate. And like part of that is regardless of what you believe in terms of what products you use, what systems you put together, fundamentally we're all trying to do the same thing. We're trying to help more people contribute to building software faster and help more users at a faster rate than ever before.
53:00 And as we all go down this road, there's a lot of learnings and journey that we have. And like the one shared learning that I suspect every single one of us, including Tyler seem to have, Cole, you have, Dex, you have, we've had on our team, which is it's process engineering all the way. And process engineering means change management. It means mistakes and absorbing mistakes. and >> it means finding and elevating bottlenecks. It means what? Okay, what's the what's the hardest, slowest part of this system? Don't stuff more work in the top of the funnel if if there's a bottleneck here because all work in progress is dead.
53:35 >> Yeah, exactly. And like and don't be afraid to do things more manually like just because you need like Tyler showed this perfectly before where like they have a human review step for parts of their codebase and they recognized some stuff doesn't need human review so they accept that risk and it's always a trade-off of this like risk versus correctness trade-off. So like decide where you are on the spectrum and decide that for every single part of your system. Some parts you don't have to be one or the other for your whole stack.
54:04 It's interesting that open code kind of lets you build kind of stuff. So thank you for sharing all that today. Cole >> open inspect. >> Yes I said open inspect. >> Oh you said open code but >> Oh frick. Sorry. >> Yeah no worries. >> cool. Any any final thoughts from you Cole? Like what's what's exciting? how how can people find you? How can they be helpful? and and yeah, we'll close it out. >> Yeah. you can find me on Twitter, just my full name, Colemurray. For teams looking to deploy software factories, I would love to chat. I think as mentioned the challenges a lot of them are just really org related and how are we going to deploy this system whatever that system means to you within your company and do that in such a way that you achieve whatever outcomes you're looking for and yeah I think more than happy to chat with teams and thanks for having me.
55:00 >> Thanks for thanks for joining Cole.
Summary
- Emphasizing the need for clear communication between agents and humans to maintain control over software development.
- Cole's Open Inspect is an open-source software factory system designed to streamline development environments and processes.
- Tyler's Bloom Text has successfully automated its software factory, enabling rapid growth and efficient feature deployment.
- The importance of creating a robust dev environment that minimizes external dependencies, allowing agents to operate effectively.
- Non-engineers can contribute to coding tasks through small bug fixes, gradually increasing their scope as they gain familiarity with the system.
- The role of postmortems in learning from errors and improving the software factory process.
- The discussion highlights the balance between risk and correctness in software development, advocating for manual reviews where necessary.
- The need for organizations to adapt their processes and manage change effectively to integrate AI and automation into their workflows.
Questions Answered
Why is optimizing for human understanding important in software development?
Maintaining clear communication between agents and humans is crucial for effective software development. It helps developers stay engaged and in control of their projects. The discussion also highlights challenges in creating development environments and enabling non-engineers to contribute to coding.
What are the challenges in building a messaging app for doctors?
Creating a messaging app for doctors involves addressing specific compliance requirements, such as HIPAA. The discussion emphasizes the complexities of real-time messaging features and the importance of understanding the product's functionality.
How can software factories automate the ingestion of work?
Automating the process of feeding work into software factories can enhance efficiency. Insights from customer demos and support discussions can serve as a roadmap for feature development. Regularly capturing these insights is essential for prioritizing tasks.
What are the challenges of managing development environments in large companies?
Managing development environments in large organizations is complex due to the number of internal services and dependencies. The discussion highlights the importance of having control over the environment to ensure effective development processes.
What are the trends in the adoption of local and background agents in software development?
There is a mix of preferences for local and background agents among developers. While some prefer local environments, background agents are becoming attractive for long-running tasks to avoid overloading personal devices.