transcribe

How to Move from Prompt Engineering to Harness Engineering in Testing

Automation Testing with Joe Colantonio · 46m · transcribed Jul 2026
More from Automation Testing with Joe Colantonio Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

The Future of Testing and AI Integration

What is the future of testing in the context of AI?

The future of testing may not involve traditional test writing but rather engineering systems that generate code. Matt Wyn discusses his experiences with AI in software development, emphasizing a shift in mindset towards a hands-off approach where humans may not directly write or read code.

  • The concept of a 'software factory' suggests a new way of building software without direct human coding.
  • AI can change how developers perceive their roles and value in the software development process.
  • Embracing AI tools can lead to innovative practices in testing and development.
# 9:21

Steering AI Agents for Desired Outcomes

How can developers effectively guide AI agents in software development?

Developers can create constraints and contexts for AI agents, allowing them to produce desired software behaviors. This involves setting up validation loops to define what 'good' looks like, which can lead to improved outcomes and a more enjoyable development experience.

  • Establishing guardrails for AI agents can enhance the quality of their outputs.
  • A mindset shift is necessary to leverage AI effectively in software development.
  • The process can rekindle joy in software development similar to Test-Driven Development (TDD).
# 18:42

Leveraging LLMs for Summarization and Insight

How can LLMs be utilized to enhance productivity?

LLMs can be used to summarize long conversations or documents, extracting key points efficiently. By providing relevant context, users can maximize the utility of LLMs, but caution is needed to avoid generating inaccurate information.

  • LLMs excel at summarizing existing information rather than generating new content.
  • Providing context is crucial for effective use of LLMs.
  • Users should be aware of the limitations of LLMs to prevent misinformation.
# 28:04

Continuous Improvement with AI Agents

How can teams improve their interactions with AI agents?

Teams should conduct retrospectives after sessions with AI agents to identify areas for improvement. By analyzing what went wrong and adjusting the context provided to the agents, teams can enhance future interactions and outcomes.

  • Regular retrospectives can lead to better performance from AI agents.
  • Understanding the context and feedback loop is essential for continuous improvement.
  • AI-generated outputs may require human interpretation and refinement.
# 37:25

Evolving Roles in Software Development

What changes are occurring in the roles of testers and developers?

There is a trend towards blending roles in software development, with testers and product owners becoming more involved in coding and implementation. This shift empowers individuals to contribute directly to the codebase, enhancing collaboration and innovation.

  • Roles in software development are becoming more integrated and collaborative.
  • Empowerment of non-developers to code can lead to innovative solutions.
  • The traditional boundaries between roles are blurring in modern software practices.

Transcript

0:00 Hey, what if the future of testing isn't actually writing tests at all, but engineering the system that produces the code? Matt Wyn, who was a longtime Cucumber Core team member, is back on the show after more than a decade, and he spent two years inside Mechanical Orchard, which is a Silicon Valley startup modernizing legacy mainframe code with LLMs. He shares how he brushed up against the idea of the software factory, a hands-off way of building where humans never write or read the code, and it changed how he thinks about his own value. I'm sure a lot of you are going through this as well. So, you definitely want to hear and listen how he went from burying his head in the sand about AI to encoding his entire Birkin coaching practice into 500 lines of markdown and admits that this was horrifying as it was amazing as well.

0:47 Now he's teaching testers and developers how to stop prompting and stop building harnesses and validation loops that let you trust the code you'll actually never read. So if you're drowning in AI generated pull requests or you're not even sure what your job even looks like in aentic world, you want to stay with us all the way to the end because I think if you listen to this episode, it's going to change or reframe how you think about your role in the future. You don't want to miss it, check it out. Hey Matt, welcome back to the show.

1:16 Thanks, Joe. Yeah, it's good to be back. It's been a long time. >> My gosh, I was just looking at it. I think you were on episode 35. I'm almost on episode 600. >> Wow. >> Yeah, it was over a decade ago. >> Blast from the past. >> No doubt. >> I Yeah, I can't I can't say I remember that much about what we talked about last time. I could guess. >> What would you guess? >> Was it to do with BDD? And >> it was indeed. Yes.

1:44 So, I'm just curious to know, it's funny how BDD almost seems like it's the user interface of a lot of testing tools I see now using natural language and it just goes out and does the code behind the scenes. Do you find that odd or do you see that as well? >> I I don't know if I see it. I'm not sure if I'm like spending enough time in that space to to know what all of the testing tools are doing. but I think like the given when then structure is a sort of universal language of how tests are organized right like we used to call it arrange act assert before Chris Matz came up with those three keywords like it's just >> >> it's it's how you know these little nuggets of of behavior like I set something up I do something to the system I see what state it was left in afterwards it's a very like natural unit of describing behavior. So, I can I can see why it's spread.

2:47 >> Absolutely. Also, I'm I'm curious to get your thoughts on this, you know, with the Gentic AI now that a lot of the skills in going to solve a problem in automation or in development or testing goes to an agent instead of a developer. is do you still have the same communication bottlenecks that you found back when BDD was the rage when you were all into it? >> The same communication bottlenecks. I mean I don't know like I think the bottleneck it's in it's really interesting. I'm I'm still puzzling over the the bottleneck thing because, you know, I I I never had one of those t-shirts, but I kind of believed when people had those t-shirts that said typing is not the bottleneck. You know those ones?

3:36 >> Yeah. >> and yet what these machines are doing is saving us a whole load of typing, right? we can sort of wish our wish our wishes at the genie and then the typing comes back at us. We don't have to do the typing anymore. If we can say the right words, we get roughly the typing we wanted. and it is definitely changing things having that. So maybe typing was the bottleneck in some situations. but I still think that a lot of times if you zoom out from the typing, the problem was always shared understanding like shared understanding between the team that are producing the thing to solve the business problem or maybe even shared understanding between those people who are trying to solve a problem and the world out there that has a problem. Like maybe we haven't even understood the problem properly, right? and we're building something that's useless that nobody needs or it's or it's missing the target. So I think it depends how far you zoom out and maybe if yeah we've really understood the business problem we all understand the business problem together and we have a shared understanding of that then typing is the bottleneck right like now it's just a matter of how fast can we crank out value to to meet to match that problem and if we're in that lucky position then then the LLM can really help us I think the LLMs can also help us with some of the other stuff too because We are, I think, currently a bit locked into this this mindset that like the best thing to use LLMs for is to generate code. But they can do lots of stuff to organize and synthesize meaning out of tokens, out of words, like including user interviews to help you to better understand what the problem space is.

5:25 >> Yeah, I love that. And it seems like you've been going more into AI and I I believe what caught my attention is you I think you released a course recently. I saw it on LinkedIn. So maybe talk a little bit about like what you're doing now and maybe well we'll dive into that a little bit more. >> Yeah. So I' I've just had two really a bit over two really fascinating years at this startup in Silicon Valley called Mechanical Orchard. it was founded by Rob Mi who founded Pivotal Labs which is the sort of back in the day was like the most prestigious I would say consultancy that was was wholeheartedly and and proudly using extreme programming practices in their consulting. So they would come in and either like set up an XP team and deliver a thing for you or teach your organization how to use XP and like have it spread around your organization. So where when I landed there what I found was there were all of these people that like me had like years of TDD and pair programming and and a refactoring practice. and we landed there just in this moment when the LLMs were starting to become useful and Rob was just very very very encouraging to all of us to just you know embrace this technology and try and learn as much as we could about it. So I was in this really lucky position to be like amongst all of these really experienced agile people like people who've been doing agile for like 20 years and have a boss that was just encouraging me to to play with tokens and see what we could do with them as much as possible. This is a really cool learning environment.

7:09 We're trying to solve a really hard problem like modernizing legacy code from mainframes like cobalt code trying to understand what these you know those those computers are sometimes like 30 40 years old right these code bases trying to understand what they do modernize them so trying to solve a really hard problem with this new technology with with brilliant people that were really steeped in agile XP practices was just an amazing learning environment so and like one of The opportunities I got with that was I got to go hang out for a week with this guy Justin McCarthy who was the CTO of Strong DM and I don't know if you've seen any of the stuff that they published while he was there but they're basically the people that coined the term the software factory or the dark factory. So, this week with Justin just blew my mind. Basically, on the scale of ambition that we can have about what we can use LLMs to to produce for us, right? Like this this way of building software that's totally handsoff. were were really were not that like their their whole the the rule that they set for themselves to see if it was possible was what if the humans never write the code, never read the code? Like what if that was true?

8:33 What would have to what would have to be true for that to be possible? So if you think about it as a tester, like that's terrifying, right? No human's ever going to read this. >> Why would you want to Yeah. How do how how do we know whether it works? How do we know whether it's going to be performant? How do we know whether it's going to be secure? How do we know whether it's going to be malleable for the for the long term?

8:55 >> but actually so many of these things are things that we can check using LLMs, right? We can we can find out we can inspect the changes and find out whether whether they are going to be malleable, whether they are following security best practices. And if they're not, we can heal the code and make changes to it. And you if you start to move your work up to this level where what you're doing rather than steering the agent itself is you're steering the system that the agents work in. So you're creating these guard rails that the the agents have to work within.

9:37 you're you're using that non-determinism in a way that makes like the you're kind of creating a slope so that the the behavior that the software that you want is going to inevitably fall out of that system. So, it's a really different mindset because you're you're having to think about like what constraints, what context do I give to the LLM so that they're able to work like that. But like the results are phenomenal. Like some of the stuff we built that week with Justin, some of the stuff we built afterwards was just amazing. And yeah, I'm I'm like really excited about this as a as a technique. I think it's it's sort of brought a joy back for me around software development because it feels like TDD again in so many ways like cuz if you want to if you want to be able to go hands off, right, you're going to have to set up this validation loop and and you have to be able to describe to the to the system. You have to be able to set up your harness to tell it what good looks like. So you have to know, you have to be a express what do I mean when I say you know this is this thing has been validated this thing is good. So you have to be a to think about like how do how do I set up that system to validate the the output that the the random probabilistic machines are going to generate and and help it to to hone in on the on the on the solution that you want.

11:15 you don't come off as like a hype guy, you know, like that's that's drinking the the Kool-Aid. So, I'm kind of surprised. I mean, it seems like you're all in on this and this is something I know a lot of developers and testers are terrified by. How did you get there? Like, because I know a lot of people are resisting it or they're like, it's it's BS. It's just LLM's trying to up their their price for when they go, you know?

11:40 >> yeah. Oh, well I I mean I I think I have a lot of reservations about the the capitalism at play here. Like >> I see these three big big players, Anthropic, Open AI, and and Google hoovering up a lot of money that was previously going to humans, right? our bosses are paying for tokens from these people rather than paying us to to to do the work.

12:15 And like when you if you think about what is happening to our industry happening to many other white collar industries, you can see that these machines, right? These machines are amplifiers of whatever is happening in our society. They're amplifying the words the the kinds of words that we use come back at us, right? This is why every like generated PRD that you read that that Claude or Codeex has generated reads like some awful Silicon Valley excerpt, right, with all of these like buzz phrases that it just puts everywhere.

12:56 this is why you know these these AI systems are are biased towards picking minority ethnic people as being more suspect, right? In in police systems. They're just amplifying our the the existing problems we have. And this and the same goes for this like disparity of wealth problem, right? They're just amplifying moving the the money towards an ever smaller number of people. So, like I totally get why people are frustrated about that, but I think that that's like blaming the technology, the large language models, is the is pointing our eye at the wrong place. Like, we should be pointing our eye at the at the capitalist system and and the way it's organized right now because the technology is is really interesting and can do really interesting things. And we've always built software, right? like people have for the for the last several decades people have been writing software that is useless or even detrimental to humanity right and and now we can just generate more of it what a shame but it's it's I yeah I just think we're kind of missing the point by by blaming the technology but the other part of it I think is I I felt this kind of like around September October like I spent a long time actually at Mechanical Orchard around all of this this buzz kind of going I'm just going to bury my head in the sand and hope this thing goes away because I'm I'm good at this and I don't need these tools and sometime around about 18 months ago I started to kind of go hold on a minute they're not going away they're actually some people seem to be getting results out of this I'm curious is I'd better lean into this and try and figure out what it's all about. And kind of with the within 6 months of that or so, I found myself like starting to write these skills. Say example mapping, right? The the thing we probably talked about when I was when I was here 10 years ago.

15:06 >> Yeah. >> I'm like thinking like I could probably get the LLM to to help me with this. I'll write a skill. Oh gosh, this thing is really good at writing good girkin. I had a skill. I used to get paid to go into organizations and coach people on how to write better girkin. I've now encapsulated that in like 500 lines of markdown and you can run it on your feature files and it'll just be like having me in the room kind of like not far off. Maybe not as good jokes, but like it's it's really pretty good.

15:44 And that is sort of horrifying, right? It's it's like amazing and horrifying at the same time because like what is left of me? How am I special anymore if I can encode my knowledge, my skills into into these markdown files and the machines can just do it? Like what am I for? And I think a lot of us who are sort of a bit hypy like or excited let's say like I am have been through that.

16:09 >> I think actually everybody's been through that. Like one one of one of my friends described it as like a grieving. You're >> right. I was just thinking that that the stages of grief, denial, then acceptance, you know. >> Yeah. that yeah, the grieving is real as the as the LLM's like to say. >> So, >> but but yeah, I think that was definitely my experience was was it's like what what this my identity is tied up in >> right >> a bunch of stuff that in being good at a bunch of stuff that these machines can now do. So, what is left of me? Yes.

16:44 >> And I think I've come to a place of understanding that the the the like, you know, like like I said, they're like amplifiers. They're like robot arms. Like I can just reach further. I can do more stuff with these things. And still my my problem solving, my understanding what the problem domain is, my prioritization of what we should solve in the problem domain, my understanding of software architecture and how we want to organize the code is all still necessary in this in this space because these things aren't they're not actually intelligent.

17:25 They're just transformers of text. They take some text and they turn it into some other or tokens and turn it into some other tokens. It's just like which tokens do we point them at so that we get the tokens that we wanted out and we have to still be able to understand the whole process to be able to use them effectively. >> So I agree. I know some people describe it as the ultimate BS machine where it looks right but if you don't know the domain you can get in trouble.

17:52 I don't know if you've experienced that because it does produce what looks like incredible results and then once you start diving in you're like this doesn't make any sense. It's not as good as I thought, you know, and then it starts degrading over time. So, what you described sounds like next level, though. It's almost like a human on the loop where it just does loops and I don't know like >> I don't know what my question is there, but >> Well, I think like the first cuz I I I was actually talking about this yesterday that there's kind of like it's like the Dunning Krueger thing. There's a there's a sort of there's a initially when you're encountering LM, you're like, "Wow."

18:25 >> Yeah. >> This is like magic. Y >> and then you kind of go through this disillusionment where you realize that it it the magic is not as good as you thought it was and it's coming up with a lot of >> And then there's this kind of third I want to say transcendence, but it sounds really pretentious, but but there's this kind of third space you get to where you're like, well, well, of course it's because it's just it's just an inference machine. Like that's all it's doing.

18:51 >> So now I have to think about, well, how do I use Okay, but given that, how do I still use it to my advantage? So, what do I put in? Because the the the thing is is like the the best thing you can get an LLM to do, right, is have a really long conversation like this one. or it's not going to be really long, but you know, have a long conversation or or like sit down, record a transcript with a with a customer, and then give that to the LLM and say, "Right, just give me like the the five most salient points or or write me a proposal based on that conversation I just had with a client."

19:25 And all it's doing is it's it's taking the the input the the information that you already have as the input and and compressing it. It's it's ladder Kesler describes it as like semantic zoom. So I'm going to just zoom out like you would zoom out in a in a map. So it isn't it's not going to lose any information there because it already had all the information. You're just asking it to summarize it using, you know, its giant bucket of all of the words that have ever been written by humankind to organize the structure of it so that it's going to be legible. It's going to be readable for you, but like it's got all of the data there to start off with.

20:05 It's when you ask it to go outside of that to go and rifle in its giant bucket for stuff that it might start making up Like if you if you start with a tiny seed and then you say, "Okay, just go and rifle in a giant bucket and try and make something up for me that's an answer for this." Well, there's a there's a there a much stronger chance now that the the answer that you get back is going to contain some So, it's really about being accepting that that's going to happen and then thinking about either how do I provide it with the right information on the way in.

20:44 So that's like you know my context engineering where what files am I leaving around the place to to guide it and give it signpost which is actually just what we needed to do for ourselves as humans anyway in these code bases right but now we you know we can't rely on folklore it has to be legible in the codebase and then when it has produced something it's probably going to have some in it so what's my harness doing to catch that and find Cuz the thing is humans make mistakes too. This is kind of one of the other like epiphies I had about this was like when I would talk with Rob about this stuff >> Rob like from Rob's level you know humans humans make mistakes too. He couldn't really understand why people were objecting and saying well these things make mistakes because from his point of view humans also make mistakes.

21:36 Mhm. >> And actually it like one of the things we we teach in the course is this idea that I got from Justin of using the three big brains to to review a big a piece of code. So or or a plan to review a thing. If you get like cuz most people are just locked into like whichever you know they're they're an anthropic shop or they're a codeex shop. that if you take your your plan or your your pull request or whatever and you want to review it, if you get an anthropic anus or or a sonnet, a a GPT model and a Gemini model to all review that artifact, just like how if you have diverse humans on a team, those three models are trained with totally different input data. They've got different training weights. They will come back. They will see different things than than the others saw. They'll a lot of them will see the same things, but there'll be there'll be some overlap, but there'll also be unique things that they each see. And that can be a really good way of like playing the averages and giving yourself more information to put into that healing loop >> to to fix the the >> Yeah. that that's something a lot of testers talk about like you shouldn't use the same same LLM to test it that you did to code it and that's a good point if you put it through three or four different LLMs to get a perspective >> I mean I think even just using a new a fresh session from the same model yeah right >> because it's coming at it with a with a fresh pair of eyes >> that's true that's so true so you know so you talked a little bit about the course moving from prompt AI to harness engineering I keep hearing new terms. What is harness harness engineering?

23:28 >> So, this is I'm not sure if I can say her surname correctly. Brigittita Bucklers term. I think she she she's a principal engineer at Thoughtworks and she's written about this a lot. so, the harness is really like the the stuff you put around the LM. it could be just clawed code. but that's kind of not a not a great harness on its own. So, you may well extend Claude code with with skills so that there are going to be prompts that will kick in automatically if you're in a certain situation. you might have hooks that are going to run automatically in in certain situations.

24:11 are deterministic hooks that run, but it's also all of the other infrastructure that you put around the work that the LLM is doing, like when do you run your automated tests? When do you fire up an adversarial sub agent to read the changes that have just been made and tell you all of the things that it thinks are wrong with it. so your harness is like the the the system that you're building around those basic calls to the LLM to try and constrain the work it's doing and narrow it in on where you want it to go.

24:51 >> All right. And so sounds difficult. how how do you get started like like so are we moving like as you said at the beginning we're moving away from like what would make it true if it would code and test itself and just almost loop. So are we moving away from prompting then to where we have the systems all set up ahead of time and then we just become the person on the loop rather than in the in in the loop. Ideally, >> ideally, but you got to start from somewhere, right? So, >> yeah.

25:20 >> >> I think about the the kind of All right. So, like if we zoom back, you know, if we go back a I don't know how long it has to be for you, but like a year, 18 months ago, whatever. Many of us, our experience of using LLMs was copying and pasting this bit of code that that doesn't work or whatever into a chat GPT window in a browser, pressing the help me button, and then copying and pasting the code back out again. It was like Stack Overflow on steroids. That was what we had. Then we had cursor and we get get this like slightly annoying tab complete thing where it was like flashing up with its guess at what you thought the code would need to be.

25:59 but if you are beyond that and you are tending to write your code by sitting in a one of these terminal apps like clawed code and you're sending instructions. You can think about that as as like being in a ripple in a in a programming language like like Ruby where you can program like that, right? But it's kind of laborious and you may well find yourself repeating a lot of the same things over and over, re-explaining things to the machine. And I think it's a fine place to start, but the thing you need to be doing is always asking like why why do I have to do this? Why do I have to explain this to you? What? And the analogy I used the other day was like if you because this is a really common problem you you hear from people is well I prompt the agent and it produces this code but then I have to spend just as much time repairing that code to be like the code I would have written as it would have taken me to write it in the first place.

27:09 >> But like and and I'm not trying to anthropomorphize these things, right? They are machines. >> Yes. But let's imagine that that was a human. You delegated this task to a to a junior and they'd come back with this code. Would you just take it away and fix it or would you maybe you'd sit down with them and coach them, right? And that would be a learning opportunity for that human. There's no point in sitting and coaching in this claude session really because it's going to forget everything when you start a new session unless we can come back to that. But what if you said to yourself, well, I'm disappointed with the work that this human did, but how can I take responsibility for that? What context could I have given them that would have made them more successful? What do what do I wish they'd have known? Should I have written down some architectural decision records so that they knew that, okay, in this application, we use an MVC pattern, so they wouldn't have like written all of their code in in the controller.

28:17 or did they yeah that that's that's a pretty reasonable example. So >> it's like a sprint review >> but with agents >> like a sprint. So a sprint review though would be kind of like yeah would be coming along afterwards. >> Okay. >> And and thinking about yeah how did it go? How could it have it gone better? And I think that is that's really the thing. It's like when you when you finish that ripple session, have a little retro with the agent and start to think about what could we do better next time so that I won't have to keep repeating myself. What was missing? And sure, like you can you can have that conversation with the agent. It won't necessarily give you good guidance.

29:00 You've got to use your skill and knowledge of the of the codebase that you're working in and and what you want. but if you can think about it as okay if this went wrong how can I take responsibility for that what context could I give it next so that next time it will perform better gradually you're going to start getting better results from them and then the other side of it is okay it's produced this wall of text and this is we're doing this it's going to be I guess happening before this podcast comes out. But there's the it'll be on the Maven website, this this free lightning lesson on loading human context because one of the big things we see is like they just spit out a wall of text, right? Either it's like I've got this plan, you know, there's like five pages of plan that you're supposed to read or it generates this giant pull request, you know, 2,000line pull request. What's in it takes a lot of work to figure out.

30:07 but you can rather than reading it line by line like you would have had to in the old days, you can use another LLM to inspect that that thing to give you a projection of that wall of text into something that is meaningful to you. So, if I'm looking at a a large pull request, maybe what I actually want is like a sequence diagram of what are the how are the objects interacting in this change? Well, what's the impact of this change on the behavior of the system?

30:37 How is it how's it architected? I can ask the LLM to generate a diagram for me that that makes that visual, makes it easier for my brain to understand what's in the change. And again, if we're just talking about a transformation where it has all of the information, and what we're doing is is reorganizing that information into a form that's useful for me, there's not going to be in there. It's not going to have to go and reach in its magic bucket, >> right? because it has the context.

31:04 >> Yeah, >> it's not gonna hopefully hallucinate. And I think your course had had a a bullet that interested me. It said something about you're drowning in agentic PRs. And I know I've talked to other people that said, "But all this AI generated code, it's the PRs that are going to be the next bottleneck." >> But this sounds like this is the solution then to that. >> Yeah. And it's scary like >> and I think we still have to go back to this thing of like do we did we indefinitely need all of this code? But if we are going to have the machines generating the code, we can't be expected for to to have to read it all. We have to give ourselves tools to allow us to trust that code. So then then the question becomes okay well this is like the the trick right is is if you say to yourself right I'm not going to read that code but I need to ship it to prod. What do I need to know about that code so that I could trust that I could ship it to prod? What are all of the questions I need to ask? And you'd be surprised how many of those checks you could automate with an LLM. You could get an LLM to do those checks. And sure, like if you don't trust the first one, run it three times, run it five times, five different sessions, right? L, you know, eventually if there was a problem there, it will find it because each of those sessions is a fresh a fresh attempt.

32:30 And gradually you're going to so every time you're you're doing a a task, be thinking about, you know, why do I have to do this? Is this is this not a thing that I could delegate to an LLM to do for me? So the LLM can go and get me the information and then I can make the decisions. >> Absolutely. Another thing that caught my attention was, and maybe this is just me, no one else thinks this, something about working in, brownfield code bases, with landmines all around you. I used to work at GE Healthcare. We worked on a system that actually had a 30-year-old mainframe running it behind the scenes.

33:08 And so when people hear AI, they think, oh, it only needs to be like completely green field, not brownfield. Is that true? like or can you really legitimately use AI for even these older applications? >> Well, it depends what you're trying to do with the older applications. What what we were doing at Mechanical Orchard was taking, you know, the entire legacy mainframe estate and gradually moving it into the cloud into Java.

33:38 And these tools were amazing at helping us to to do that work. Not just I mean I don't want to you know I don't need to sell that business but many many organizations in that space are just taking cobalt and making job. They just they just transpile the the cobalt into into Java. What we were able to do with the LLMs was put them into a loop where okay, you've got this this set of mysterious code and you and here's a a metric of how much coverage you have of of that code. Right. Right.

34:20 Sometimes we would pull it into GNU Cobalt so you could measure measure code coverage with with the GCC tooling. other times we just instrument the the legacy code so you could see the coverage. Either way, here's a feedback loop that says, "Right, this is how much coverage you have. Now, generate me test cases." So, the the LLM could just keep fuzzing until it had a set of characterization tests for that code. And now you can take those characterization tests. It's like a mold. And you can cast a Java version that does exactly the same thing as that that old version did. And all of that is like deterministic checks, right? because you've generated a set of tests that are deterministic, but the LLM was helping you to get there. And it's this thought about like how how can the LLM be useful? How can the transformer be useful and speed up a thing that I could do by hand, but I don't want to have to or it's going to be too laborous to do it by hand. So that was like the rescuing legacy. I think the other thing about Brownfield is so I'm consulting with this company right now. We've got a code base that's like 10 years old and there's a lot of archaeological layers in there, right?

35:30 There's a lot of versions of doing things that are old and decrepit and we don't like and then there are shiny new ways of doing the same thing. >> And one of the issues that people have been having is, you know, when I generate a new thing, it follows the wrong patterns. It follows the old patterns or it gets mixed up and does a kind of mix of the two. not following the the like best practices in the codebase. But this is because those best practices aren't legible. So what we've been working through as as a crew is capturing the implicit folklore architectural decisions that the team have been making over recent years about how they want the code to be so that they're written down in the codebase so that the agents can find them so the agents can generate code that's the shape that they would have generated.

36:19 The other thing I've been doing with them that's really interesting Do you know anything about Rebecca Worf's Brock and her work on design huristics? Have you ever seen any of that stuff? >> No. No. >> So Rebecca like wrote all these wonderful books back in the ' 90s 2000s about object design like responsibility driven design is her thing and >> actually she inspired Eric Evans domain driven design I think in a lot of ways. Yes. Yes. Yes.

36:49 >> but lately Rebecca's been really interested in this this thing of capturing your design huristics. So if you are someone who designs software, what are these rules of thumb that you tend to apply you know your preferences, your tastes about how you like software to be? And an experiment I've been doing with this client is when the humans review poll requests, we have so much data now in these poll requests reviews of what this team's design heristics are because they'll take a a candidate set of code and they'll litter it with comments. They'll cover it with comments saying, "Oh, I wish it was we we we tend to do it like this here or I think I would prefer it if it was like that."

37:36 And I've had an agent go over those poll request reviews to try and distill down some design heruristics that the team has about how they like to organize their code because every code has like every team has like a house style, right? try and distill what the what that house style is and again document that in the codebase so that the agents that are generating new code are more likely to generate it in the style that the team would have written anyway to the team's preferences. It's going to have less you know churn and and comment in in when it goes through a pull request review.

38:15 >> Very cool. I want to do some quick lightning questions. all right. So, what what's the feature of testers? I know you have something in your copy that says code reviewer and chief. Is that where we're all heading? Like where what are we morphing into just one generic role of like thinking engineer? I don't know. >> I don't know. I don't know. I I mean, I definitely like I know some people that just say, you know, we're all we're all just builders now. And I saw stuff happening at mechanical orchard where we had it was definitely like blending together. We had especially what I saw was people on the product side becoming empowered to start making changes to the codebase which is just amazing like people even if it was just a spike like a prototype to be able to make something real and play with it and see it rather than just designing wireframes.

39:10 really really cool. But but I definitely saw people who were coming from a product owner role actually start implementing features in an in in the product. so there was definitely a convergence happening there. And let's see I mean it's been a while since I worked anywhere where there was like a dedicated tester role. I think for a long time I've thought of testing as a specialism within, you know, the the sort of set of T-shaped people that we want to have on a software team anyway. So, some people are their brains are just better wired at thinking about what could go wrong, right? and I don't see why that isn't still useful. In fact, maybe more useful than ever. If we go back to what I was talking about about this sort of what I'm trying to call lean software production way of looking at things, right? This this like hands-off software development way of looking at things where the engineering that we're doing now is we're engineering the system that produces the code rather than engineering the code itself.

40:27 We we really need people that can test that system, right? and can think about where are the leaks in that system, where's it going to go wrong and how do we mitigate that. So I think that skill, that mindset is still just as valuable as ever. but I think it's important to start to really skill up and understand, you know, how these things work. Like read read the attention is all you need paper and like they aren't magic genies.

40:57 They are just transformers. And once we start to kind of recognize like how the technology works, it becomes easier to test it, right? And and and work with it. Same same as when we're testing any any other software, right? Like if we just think of it as a black box, we can't necessarily break it as well as if we as if we look inside. >> Love it. And Matt, really quick, job replacement or more jobs, not less?

41:24 >> Well, you already heard me talking about software. Like I I think as long as there's money around, they'll they'll be paying us to to to make stuff. >> I think yeah, realistically there there may become a crisis where the the money's all going to the Silicon Valley shareholders and there just isn't any money to spend and then they realize there's a problem. But until we get there, I don't know how you say this guy this guy's paradox is I I want to say Javon H Jevans Jevans paradox, right? That you know the you build an extra lane on the freeway and you don't get less congestion, you just get more cars.

42:09 my suspicion is that's what's going to happen with all of this. we're we're just going to get I mean how many if if you're in a if you're in the position of being a product owner, you probably know how often you have to say no in that job. Now you just get to say no less is is kind of what I'm seeing. So I I don't I don't think we've got a problem there. but I think it might reorganize things like I think the way we organize ourselves in teams when we are moving faster when things are changing faster because we are doing more stuff I think it gets harder and harder to share that context across the human brains right so what would have needed a team of eight to get it done like that team of eight if they let's say they can do twice twice as much stuff as they could do before. Maybe actually what you need is four people on that team rather than eight people trying to keep up with twice as much stuff happening, right? So, we have to organize ourselves differently and break that into two teams somehow. and I I don't know yet how that's going to look, but I think that's an issue that we we haven't really thought about yet. I think we're still organizing teams and structuring organizations in quite an old-fashioned way and we haven't really recognized the impact this is going to have on all that.

43:36 >> Absolutely. Okay, Matt, before we go, two-part question. one piece of actual advice to help someone with this AI transformation and the second one is where can people learn more about prompting AI to harness engineering your course? So, piece of advice, probably if you haven't got one already, get a modest subscription for yourself and play with it in your spare time. Make things. Make a little thing for your family that organizes your shopping lists or whatever. Just play with it.

44:10 and get a feel for it in a lowrisk place because you might be surprised how good a good good code it produces and what it feels like to be steering it at the product level rather than thinking about the code. So something that you want >> best way to contact you to find more about find out more about your course. >> Yeah, go to mattwin.net mat wy n.net net. we're also launching, and maybe I'll have it done by the time this podcast comes out now because there's a good deadline, leansoftware.ai.

44:49 That's the the kind of brand we're building around all of this software factory stuff. So, yeah, go to leansoftware.ai and we'll have links to all this awesomeness down below. Thanks again for your automation awesomeness. the links to everything of value we covered in this episode, head on over to testgild.com96. And if the show has helped you in any way, why not rate it and review it in iTunes. Reviews really help in the rankings of the show, and I read each and every one of them. So that's it for this episode of the Test Guild Automation Podcast. I'm Joe, and my mission is to help you succeed with creating end to end full stack automation awesomeness. As always, test everything and keep the good. Cheers.

45:38 Hey, thank you for tuning in. It's incredible to connect with close to 400,000 followers across all our platforms and over 40,000 email subscribers who are at the forefront of automation, testing, and DevOps. If you haven't yet, join our vibrant community at testgild.com, where you become part of our elite circle driving innovation in software testing and automation. And if you're a tool provider or have a service looking to empower our guild with solutions that elevate skills and tackle real world challenges, we're excited to collaborate. Visit testkild.info to explore how we can create transformative experiences together.

46:18 Let's push the boundaries of what we can achieve. >> Testing with loots and liars. The bars began their song. A tune of knowledge, a melody of code. Through the air it spread like wild fire through the land, guiding testers, showing the secrets to behold.

Summary

Matt Wyn discusses the evolving landscape of software testing, emphasizing a shift from traditional coding and testing practices to a more automated, hands-off approach facilitated by large language models (LLMs). He shares insights from his experience at Mechanical Orchard, a startup focused on modernizing legacy code, and explores how testers and developers can adapt to this new reality by engineering systems that produce code rather than writing it themselves.

- The future of testing may involve less direct coding and more focus on engineering systems that generate code.
- Wyn highlights the concept of a "software factory," where humans do not write or read the code, raising questions about trust and validation.
- He encourages testers to embrace AI tools, moving from prompting to harness engineering, which involves creating systems around LLMs to ensure quality and reliability.
- The importance of shared understanding within teams is emphasized, as is the need for testers to adapt their skills to this new environment.
- Wyn discusses the potential for LLMs to assist in modernizing legacy systems, generating tests, and improving code quality through automated checks.
- He acknowledges the fear and uncertainty surrounding AI's impact on jobs but believes it will lead to new roles and opportunities rather than outright replacement.
- Wyn advises individuals to experiment with AI tools in low-risk environments to gain familiarity and confidence.
- For further learning, Wyn recommends visiting his website and the upcoming platform leansoftware.ai for resources on harness engineering and AI in software development.

Questions Answered

What is the future of testing in the context of AI?

The future of testing may not involve traditional test writing but rather engineering systems that generate code. Matt Wyn discusses his experiences with AI in software development, emphasizing a shift in mindset towards a hands-off approach where humans may not directly write or read code.

How can developers effectively guide AI agents in software development?

Developers can create constraints and contexts for AI agents, allowing them to produce desired software behaviors. This involves setting up validation loops to define what 'good' looks like, which can lead to improved outcomes and a more enjoyable development experience.

How can LLMs be utilized to enhance productivity?

LLMs can be used to summarize long conversations or documents, extracting key points efficiently. By providing relevant context, users can maximize the utility of LLMs, but caution is needed to avoid generating inaccurate information.

How can teams improve their interactions with AI agents?

Teams should conduct retrospectives after sessions with AI agents to identify areas for improvement. By analyzing what went wrong and adjusting the context provided to the agents, teams can enhance future interactions and outcomes.

What changes are occurring in the roles of testers and developers?

There is a trend towards blending roles in software development, with testers and product owners becoming more involved in coding and implementation. This shift empowers individuals to contribute directly to the codebase, enhancing collaboration and innovation.

© transcribe · For agents Built with care and craft by Gokul Rajaram