transcribe

Why Developers Hit a Wall at 4 AI Agents

AI Native Dev · 47m · transcribed Jun 2026
More from AI Native Dev Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:00 most people there working interactively with 1 to 2 agents the most experienced engineers, they get stuck for Max, very few people get to five. any of us who work in Claude Code every day human attention is limited. You can only babysit so many agents and you inevitably end up forgetting about one. And this time last year people were debating if AI was even useful we were talking to customers that weren't convinced about how fast they should even move, given that AI quote like just didn't work yet.

0:29 with this massive data set, the scale of it being able to see, across the entire industry what's actually happening out in the wild, helps kind of ground the difference between what you're reading on X and actually is happening at real companies. The AI Native Dev is a podcast for developers and engineering leads at the cutting edge of AI and agentic coding. Join your hosts, Guy Podjarny and me, Simon Maple. Every week as we chat with the most exciting voices in AI and tackle the biggest questions facing developers today.

1:01 This is the AI Native Dev. Back in November, we hosted the first ever in-person AI Native Dev Con in New York. This June 1st and second, we're bringing it to London. It's two days built for AI Native developers and engineering teams, one day full of hands on workshops and one day full of practical talks on agent skills, context engineering, agent orchestration and enablement platforms and how teams are actually shipping AI in production.

1:33 Join us at the Brewery in London near the Barbican for all of that, plus networking parties, giveaways and a room full of people. Building the future of AI native development. You can also join us from anywhere in the world via the live stream. As you're listening to this podcast. You get 30% off your ticket with Code Pod 30. Just head to AI Native Dev Con and we'll see you in London. Hello and welcome to another episode of the AI Native Dev.

2:04 My name is Simon Maple, your host. And joining me today is Nicholas Arcolano, who is the head of AI and research at jellyfish. And Nick and the jellyfish team unveiled a huge amount of data, which we're going to look through, which describes how agentic coding is done at organizations. some really interesting data and findings which show developers creating and merging twice as many pull requests as they were without AI coding tools. As well as that, we're going to be looking at the barrier, how developers hit a barrier when using multiple coding agents in parallel, and what it is we need to do to get beyond that.

2:43 And finally, 2026, it's the year of the CFO. What do our engineering leaders need to bring to the conversation to make sure and show that engineering teams are effective and productive? Nicholas, welcome. How are you? I'm great. I'm happy to be here. This is exciting. Nicholas, tell us a little bit about. Actually tell us a little bit about jellyfish for those who haven't heard of jellyfish before. Yeah. So jellyfish, we're in the AI observability space. And in particular we are focusing on understanding AI transformation at the organizational level.

3:19 So we're tracking what tools our customers are using and pulling in information, you know, not only about, you know, the agents that are using and how they're using them, but also how those things connect to outcomes like the code, the pushing, the quality of that code, ultimately with their business outcomes are. Yeah. Awesome. And actually, with all that you have been reporting fairly regularly on a lot of this data. And I'm calling you Nick, but really, I should call you doctor Nick.

3:47 Right. Because you have a PhD as well. Tell us a little bit about the background behind that. Yeah, I do the my, my Slack avatar is the little, you know, the doctor Nick from some folks around the office call me that. So, yeah, I mean, my my background, you know, I, I've been doing these things. I was in the national research here in the US for years before data science was a thing. And so my, my degree is actually in applied mathematics.

4:17 And I looked at doing kind of inference and signal processing on massive networks for things like cybersecurity defense, and realized that I could start calling myself a data scientist. And and now that title is kind of passé, right? The work to become a data scientist and now it's AI engineer is the exciting. That's it. Yeah. I'm just looking. Right. Absolutely. And I'm just looking through your CV, Nick.

4:47 And, you know, obviously you went to Harvard University. Statistical signal processing, graph theory and network analysis, then to MIT Lincoln Laboratory as a technical member of staff, then to Runkeeper as a senior data scientist, True Motion as director of data science and jellyfish of as head of AI and research. So we have the right person here to talk us through about AI and data. So first of all, tell us a little bit about the and of course, you know, we're going to be talking about about data about trends of adoption and usage through this through this podcast.

5:19 Tell us a little bit about the data that you collect and what you share every month. Yeah, the data is I mean, it's just been phenomenal. The the growth of it, what we've been able to collect. So, you know, we started for a long time. Jellyfish has been around almost a decade. And so, you know, for a long time we've been pulling in things like, you know, get signals commits and pull requests, comments, things from Linear and ADO and JIRA in terms of what tasks people are doing.

5:47 It's kind of the bread and butter of the engineering machine. And about two, a little more than two years ago, you know, when Copilot really started becoming ascendant, we started pulling in data from those APIs to understand just that. People were going to use these tools, sort of basic usage. And that's of all to where now it's just a wealth of these APIs, hooks, open telemetry. So we're able to not only understand, you know, are people using these tools, but, you know, what's what's the token spend look like?

6:20 What models are they using? And just recently we've started getting data at kind of the agent turn level understanding, you know, things like planning and tool calls. And so that's kind of the most exciting frontier of this data, just the scale of it being able to see across the entire industry what's actually happening out in the wild, you know, ground if you're terminally online, which I think a lot of us are these days, given the pace of things that just kind of ground the difference between what you're reading on, on X and what actually is happening in real companies.

6:54 That's it. And that's what I love about this kind of data. It's not it's not a sentiment data. It's not a, you know, a very, very small subset. We're talking about tens of millions. What most $40 million that that were assessed in this. From what I understand from all of that data, what would you say are some of the highlights in terms of in terms of the one liners that really stand out for you? I mean, one of.

7:16 The biggest ones was it's it's amazing how fast it's changed. And you'll you'll remember Simon, you know, this time last year, people were debating if AI was even useful, or at least that was like outside of maybe the the bubble of us that had already bought into the utility of this. But we were talking to customers that weren't convinced about how fast they should even move, given that AI quote like just didn't work yet. And we know what happened at the end of last year.

7:44 I mean, there's been a couple of key changes, obviously, but, you know, the models that came out in the fall and kind of how that advanced engineering. So, you know, we find ourselves now with this massive data set, we've really cemented the understanding that there are real raw coding gains, and we see them out in the wild. know, we find ourselves now with this massive data set, we've really cemented the understanding that there are real raw coding gains, and we see them out in the wild.

8:10 these fancy IDEs. They're using Cursor and Copilot. Those things are increasingly agentic. And without changing other material things about your workflow, two X is about kind of what you can do in terms of just raw code throughput. And that's and so the two x there is the increase in the number of pull requests that are, that are created. Right. And are there other types of applications or types of, I guess, projects whereby you see those increase faster, or types of projects or applications whereby you actually see, you know, AI maybe not as well used.

8:51 Yeah, that's a that's a fascinating question. And we do see that out in the wild. Some of it matches what you would expect, which is, you know, smaller code bases. Newer code bases. We see language differences. Those things, you know, are faster with AI. So things, you know, the languages that we all kind of experience are more AI friendly. And things like Python typescripts, things that are very heavy and, you know, configuration, markdown, YAML, you know, those things move faster.

9:25 Things that move slower are older code bases, big messy distributed code bases. We have results that show you essentially get little to no gains due to AI. You know, as your code base becomes very distributed, where there's just a lot of human work involved, both in terms of mapping together the context of how all this code relates. You know, I think we've, you know, we've all been in the situation and know the people who, you know, you've got the senior engineers in the world, the bodies are buried if you want to get deployed.

9:59 Well, you know, what do we have to do? And how does this relate to that? And it's just knowing, like, oh, you have to change the code in these ten places across these, these repositories. Best of isn't operationalized yet for those big, messy, sprawling code bases. And you just don't see the games due to AI there because the the agents don't know what to do. It's still heavily human in the loop. Yeah, yeah. Super interesting. And I think it'll probably I don't know whether there's enough information there, enough data there to say, well, actually this is the style of application whereby AI will really, really help you or whether that's just, you know, natural because people want to use it in certain situations because they're more familiar themselves.

10:39 They understand, they understand knock on effects and those types of things. But it'd be fascinating to kind of like dig into that, I guess from the pull request point of view, when we talk about the the two x, or is that two x merged pull requests or raised pull requests? Those are. Merged pull requests. So it's two merge pull requests. The the fascinating thing and the thing that's kind of mind boggling is we're starting to see.

11:07 But for a long time throughout the past year, you know, six months even, you weren't seeing downstream increases. So you weren't seeing increases to like the number of features shift. You know, we have a view into things like not just issues, but epics and those things roll up to initiatives. So our customers at jellyfish do a lot of work to use our platform to, you know, map those things into the kind of the business units that they care about, product lines, customers.

11:37 And we weren't seeing movement on those things. And we're starting to so, you know, there's increases, there's more capacity. But where's all that capacity going is kind of this, this mystery that we've been trying to untangle for the past several quarters. Yeah, yeah. Interesting. And I'd love to dig a little bit deeper into the, into into the number of pull requests here. Because the one thing that's very curious and it kind of like, I feel this from my previous role at sneak where, where sneak was very, very good at identifying issues, raising those pull requests.

12:09 And, you know, those pull requests will sometimes get merged. But a lot of the time they won't. And they're there and they're valid pull requests. But there's extra effort that is required to to to merge them. There was some fascinating data here about merge rates of AI generated pull requests versus human pull requests. Tell us a little bit about that. Yeah. So I was just looking at that recently and I published some some results there where, you know, the the average for Q1 for humans was about 80% of pull requests that were open ultimately got merged.

12:47 Any other 20%, the either stay open or they get closed without being merged. There's a variety of reasons that that happens, right? It might I might not have met quality standards. You might have decided it's not a good idea. What we see with PRs that is 6040 instead of 8020. So you're talking about double the amount of press that are kind of dying on the vine, so to speak, there. And, you know, we're digging into what are the reasons for that.

13:17 But anecdotally, when you talk to people, some of that is, you know, workflow differences. So you know, we have we have customers we talk to and I'm sure, you know, people who they're doing these things proactively. A customer request comes in, you just tell an agent to go to fix, you know, so it's just there. And then you decide later if you want that fix. We do have you have the workflow where people might do things, you know, two, three, five different ways.

13:47 You're not sure the architecture you want, you're not sure the way you want the feature to work. So you, you know, do a bunch of different candidates and then you pick one. So that's an intentional kind of throwaway work there. And then there's also the version we hear a lot about where, you know, people see these things on the on the backlog. They see changes and they sort of, you know, to use the term that's falling out of fashion, vibe code a fix.

14:14 And then people realize it's more complicated than they thought. It doesn't it doesn't meet the quality needs. Senior engineer steps in and says, no, no, no, no, no, we didn't we didn't fix this for a reason. There's there's your unearthing demons, right? There's there's there's deep technical issues buried under this seemingly simple thing. So all of these kind of compound to, to give you that the doubling of the rate. And the reason I think it's so interesting is because this gets it.

14:41 The question we all want to know, which is ultimately what does good engine engineering look like and what is it going to cost. And so is it. You know, how much the waste, you know, overhead is there going to be when you're doing highly autonomous things, things with teams of agents? You know, it's not a simple scaling of what people are doing today with 1 or 2 interactive agents. Really interesting. I actually assumed incorrectly that this was as a result of maybe more autonomy on the agent's point of view to I'm going to, you know, do some stuff and create a pull request based on a trigger I found versus a human, actually, you know, initiating these requests and then and then choosing to choosing to pick one.

15:27 So so it sounds like not to paraphrase, but it sounds like is this a flaw in our workflow whereby a GitHub or something like that, you know, a git repository just isn't cut out for agentic development whereby we need a layer there that should expect, you know, multiple versions of a fix and allowing us to pick one. And actually what we're doing is we're we're creating a workaround where rather than send one pull request, which is the fix, we're actually sending multiple and throwing some away.

16:01 It feels like we're we're trying to use the technology here, you know, almost as a, as a means to what we can do with, with a development right now. Yeah. I completely agree with with that. And we I mean, you know, we know that, you know, the whole world of software is littered with tools that maybe aren't being used for the purposes that they were designed for. Very true. You know, some of those persistently just use them in these weird new ways forever and some of them don't.

16:32 So I think history will tell whether, you know, my my kids are doing agentic engineering and asking, why do we do all this weird kid stuff? That's just it's been around for decades and it's not going away. But it's not it's not the right pattern. The question of what is what is the pull request look like is kind of the the core unit of software engineering, I think is a fascinating one because, you know, right now it's so tied to value in the shipping.

17:02 And we see that whole SDLC collapsing in new ways that I think it doesn't necessarily make sense anymore in a truly agentic world. Yeah, yeah. Super interesting. It's amazing how many DIY tasks or home improvement tasks I can do with a hammer. My wife will my wife will swear at me many times. So let's let's move a little bit on to a new topic here talking about AI adoption within organizations. So you've said in the report AI adoption has reached a median of 71% of developer time.

17:33 That seems that seems a lot. Tell us a little bit about how that data is generated, how you're learning from from, you know, AI adoption statistics here. And, you know, is this something that is something across our entire industry or within certain groups? Yeah. The in terms of adoption, it's interesting. We I think for some of us it seems like it's like, why are we even talking about this? All of us use this all the time.

18:00 And you know, jellyfish, we have the benefit of just seeing such a broad swath of companies. And so, you know, all of our customers, they understand that they need to be doing AI coding, they need to be doing agentic coding. And the basic adoption barriers are still very real for a lot of companies. Most of them have gotten over the hurdle of just the basic, you know, security enablement, things that can even get licenses to these things.

18:28 But some of them are still just measuring. Do people even use these tools? And if they do, let's say there's a mandate to use them. Are they just checking in in a performative way or are they actually integrating them into the workflow? You know, I think it's a it's a vanishing minority of people who, you know, are highly resistant to these tools, and maybe they're logging on and using them, but really don't want to what's much more common.

18:54 And so that that 71%, that's weekly active users across the entire developer base that we're tracking about 250,000 developers in total at this point. And so, you know, the median is 71%, P90 is 90%. So 90th percentile is people using it at 90% essentially all the time except on vacation or being stuck in. Even then we're we're using them in meetings all day.

19:25 Almost. The what's the typical background of of these 250,000 like, is there an average developer here Are they working? Are they more open source? So are they more freelancing? Those types of things? Yeah. Our customers. So they're they're paying us to track their AI transformation and how that maps to the business outcomes. So they tend to be, you know, they're they're companies that either build and sell software or software is, you know, part of how they do business.

19:56 But they might be saying apparel company, but they still have 200 engineers because they have operations and they have a digital presence. So, you know, you need jellyfish when you get to be a couple dozen, you know, tens of engineers. And then we have customers up through tens of thousands of engineers. You know, what we don't see are the, you know, five people kind of, you know, in the Valley that are that have started an AI native company yesterday.

20:28 They're they're not using us, you know, because they just got started. They haven't run into the problems of the level of observability. We hope that they do. Yeah. But those are the types of companies we don't see a lot of open source models. So these are a lot of companies. You know they're they're buying kind of big enterprise tools and licenses. So they're they're using Copilot Cursor and Claude Code, you know, kind of Cursor and Claude Code or dominant Claude Code is, as you might expect, has just been ascendant over the past six months.

21:04 And so and they're using those models both through those providers and they're using bedrock models. And so it'll be exciting. Now to see now that you can use the open AI models on Amazon bedrock. Right. See a lot more uptake of those. We have a lot of customers that use bedrock. Yeah. And I'd love to I'd love to go into a per tool actually analysis to kind of understand if there are any trends per per kind of like agentic coding platform.

21:29 But I'd love to understand in terms of, in terms of the weekly active use. And this is this is something that I'm super interested in, the depth of use of these of these tools. How much would you say you have data that show how much people are relying on agentic development in their roles, as well as being a weekly active user? Yeah, that's a that's. A great question. And there's a couple of different lenses on this.

21:58 So one basic lens. So you've got, you know, this this person even touching these tools on a regular basis that's, you know, the very basic table stakes. And then we look at, you know, do people have a reliable habit. Are they using them repeatedly, you know, day after day, week after week. And so one benchmark we said is are you using AI tools regularly three or more times a week. And that persistent, frequent active usage, that's what correlates with and how we are able to see companies that are able to drive up persistent usage across their users.

22:35 They see those productivity gains that I was talking about that to X, and then the level beyond that is autonomy and kind of agentic workflows. And that's where we see the we look at things like what percentage of your peers are being generated in an autonomous fashion. So an agent is able to take a spec and actually do all the work, open the pull request with minimal human interaction. And that's an area where we've seen kind of fascinating growth and really different stories.

23:07 So the, you know, the the elite companies, the 90th percentile that we look at, those folks, they last month, they were they crossed 20% this month were crunching the numbers. It looks like it's closer to 30%. But it's basically been exponential growth. You know, the and they were at, say 2% about a year ago. The median company, that company is just past 2% recently. There are about two and a half, 3% now. So it's really, you know, a story of the leaders here, the people that are figuring out this agentic development.

23:44 They're kind of running away with it in a lot of ways. And everyone else is still kind of figuring out how to get these workflows repeatable, scaled. They might have a small pilot team that could do these things, but they don't. They don't have it, you know, kind of scale to the whole organization. Right? Right. And I think that's very classic of rolling out any kind of technology. So it's it's similar to how we've seen before.

24:10 I'd love to talk about the top x percent. Maybe, you know, those who are the very heavy users of agentic tooling. And it kind of like leans us into something that you refer to as the barrier concept. So these, these more elite, very experienced agent developers who are, you know, using many different agents in, in parallel, tell us what the agentic barrier is and when they reach it. Yeah. So this is something that we observe. So this is work we did in my team.

24:42 And a colleague of mine, it's Tomas Pardiñas. You know, he's gone deep in understanding these workflows. And what we see is, you know, most people there that are kind of working interactively with 1 to 2 agents and we see the most experienced engineers, if they have to do kind of any interactivity with these agents, they get stuck at, you know, for agents, Max, very few people get to five. And it kind of jives with what, you know, any of us who've worked with these, you know, work in Claude Code every day like I do, you just you can.

25:22 Only human attention is limited. You can only babysit so many agents and you inevitably end up forgetting about one. And you're like, oh, I forgot that was even a thing. You know, you missed the push notification to to go click on that terminal window and you forget to go back. And so we find that even at that level of four concurrent agents, you end up spending 80% of your time just focused on one. So you know interactivity just has its limits.

25:49 And so, you know, to break beyond that barrier, you have to get to a fully autonomous mode. You have to be able to to hand off work in its entirety if you have to, to nudge the agent along, if things just kind of grind to a halt in the, you know, the total concurrency is just kind of has a hard limit to it in two minutes. Yeah. And I guess that's almost like allowing instead of the human being, the orchestrator, you're just enabling another agent to do the orchestration, which is essentially handing the whole thing off.

26:18 Do you do you feel like there's a better way that we can work with multiple agents, or do you feel like this is a technology issue? There's what makes this a hard problem to answer is there's clearly technology issues. And we've seen advances in agents interfaces. You just talked about pull requests. You know, lots of us agree that the tools that we're using to to manage code, to manage agents, to inspect what's going on. Are limited.

26:48 And I think all of the the players here in the space are, you know, all of them would love to event invent what the the the new interface, the new IDE for a development is right. That would be that would be $1 trillion invention certainly. But I'm also you know, I have my own skepticism. And part of what I'm trying to understand, you know, dissecting all this data about the SDLC is, you know, what are the limits in terms of actual business and product outcomes.

27:22 So what I mean by that is, you know, there's there's governors, there's a speed of light associated with how fast you can make business decisions, how fast you can get feedback from the market, how fast you can enable your go to market, how fast you can change your product. I hear product leaders talking about starting to deal with user fatigue of products changing too fast. And you don't, you know, you don't even have time to get feedback on whether these features are better because people don't have time to use them, and they becoming increasingly frustrated with how fast the products are changing.

27:53 So understanding when are these throttles on our ability to actually build products, then throttles on how fast it even makes sense to build software, and how interactive you really do want it to be? Yeah, it's a really good point, because I think we've got to consistently remind ourselves that our users are still our users and, you know, we're not well, we're we're sometimes we're sometimes, you know, using technology or agents as our users, as our customers.

28:23 But, you know, the majority of the time it's going to be a user that's going to be using a software in an end user point of view. So, yeah, I guess I guess that's one style where actually slowing down a little bit and having that more interactive flow can be better, because we can introduce more of our thoughts and our, our takes. Are there other ways in which, you know, autonomy or autonomous agents is actually the wrong call?

28:49 And how does that how does that adjust things like team sizes, how we interact with other other folks on our on our teams? Well, certainly. You know, the we see a lot of customers who are, you know, experimented with and thinking about these smaller team sizes and the ratio of product folks to, to builders. You can just do so much more with an engineer now than you used to be able to. It's so, you know, I am certainly of the belief that a good product leader, someone with good, you know, user and business sensibility is just worth their weight in gold or tokens or whatever the most valuable commodity of the moment is.

29:30 So, you know, I think I think that's a thing that is is changing. And there's a real question I find myself asking the question when I sit down to do a thing, you know, how much should I brainstorm with the agent and figure out exactly what I want? And then in that thing off in its entirety, or how much should I kind of work through it and help the, you know, have the agent help me understand what I want?

29:56 You know, I think I'm in a common position that a lot of developers and creators are, that I sit down and even if I think of know what I want. I'm quickly disabused of that notion that I have thought this through in any meaningful way. And engineering, so much of it just depends on these, you know, really kind of sometimes high stakes, you know, high tech decisions about things like architecture, right. Trade offs in the gray area, things.

30:25 Those are really the meat of of true engineering and true product development. And I think we're all trying to figure out how much of that comes through friction and exploration and banging your head against things. And if you outsource too much of that to the agent, do you lose the opportunity to have those insights because you're not in there? Hey everyone! Hope you're enjoying the episode so far. Our team is working really hard behind the scenes to bring you the best guests, so we can have the most informative conversations about agentic development, whether that's talking about the latest tools, the most efficient workflows, or defining best practices.

31:02 But for whatever reason, many of you have yet to subscribe to the channel. If you're enjoying the podcast and want us to continue to bring you the very best content. Please do us a favor and hit that subscribe button. It really does make a difference and lets us continue to improve the quality of our guests and build an even better product for you. All right, back to the episode. I'd love to talk a little bit about, I guess, one of the differences from 2025 to 26, you mentioned 2026 is the year the CFO gets involved, which is which is a very nice way of putting it.

31:35 I guess in 2025, people were much more encouraged to try everything you can. Let's learn much, much more about what is out there and what suits us well. And the only way to really do that because of, you know, everyone having different cultures within their organization is to try it themselves and see what works for them. I guess now people are a little bit more cautious in terms of, you know, spend and trying to think how how do we make the most out of our budget with AI we still want to spend in AI, but we're, we're we're much more thoughtful about our budget.

32:12 I guess from the point of view of data, do you see a difference in 2025 to 2026, in terms of how much variety there is in tools that are being used? Are people focusing more today on a subset of tools, or are they still do you see people still growing in their exploration? We see. So we definitely have seen and we advise people to at this point, so many of the tools have gotten so good that, you know, the risk is kind of, you know, getting analysis paralysis and trying to find the exact right agent, the right workflow.

32:55 Should I do spec-driven development or not? Should I use Codex or Claude Code? You know, for the vast majority of cases, picking a horse and writing it and just getting good at exercising that's building that muscle is the right move. So we do see a lot of that. You know, the the big difference is, you know, the the kind of move to scaling this year. So a lot of folks yeah there was there was budget and urgency to just explore and try things out.

33:27 And now as companies try to scale, we've seen how much more complicated, you know, how much more complex agents have gotten, how much more capable they are, but also how many more tokens they are. You know, the models are beefier. You know, the agents are doing more turns, they're doing more code exploration. So token costs have token consumption is just skyrocketed. And, you know, I think that the tension that engineering leaders are feeling is they're still very much a, you know, it's like the token Max in conversation.

33:59 There's still very much a perception that, you know, adoption and regular use of these things needs to be driven, needs to be supported. And so if you're not using the stuff enough, then you're not you're not going to get there. So you've got to be burning enough fuel to show that you're you could escape, you know, you can achieve some escape velocity, but you have to show your receipts. You know, we're at the point where engineering leaders also have to explain where are all the tokens going is a lot of money.

34:32 And what does that look like? I guess in terms of if the CFO says, look, we've spent all this money, we need to see some level of either efficiency or productivity that our teams can build faster, deliver faster. What does that look like from an engineering manager's point of view? What should they produce back that the CFO would be happy with? Yeah. It's it's a real mess, Simon. The I mean, the challenge is. So that isn't the answer.

34:59 I presume that you give back to the CFO, right? Yeah. No. It's what's fascinating about this right now is that, you know, so there are good metrics. So you can look at raw token spend. Right. That's sort of a basic signal of are we like your like your heart rate. It's a good set. Having a heart rate is a good signal. You're alive, and it stirs more sophisticated questions which is what is the token spend per PR per you know, per feature shift, token spend per outcome.

35:31 And that gives you a sense of what you're actually spending. The the reason I say it's a mess is because, you know, what businesses really care about are things like revenue, right? They care about things that that actually change, you know, the trajectory of the business. And companies generally haven't adapted to think about what does it mean to essentially have infinite capacity to be able to exchange? I mean, I'm kind of the course's level.

36:02 We've entered a world, and in a world where you can just spend infinite money to get things built faster, you can hire infinite robot contractors, so to speak. And companies aren't designed to reason in an infinite capacity world. They're designed to think about. These are the headcount we have. There's so many things we can build, what's going to maximize our business. So that feedback loop of, you know, what would actually matter? You know, that's that's still being built.

36:29 And I think that's hard for people to reason about where should you apply maximum leverage. Yeah. I always it's funny, I always tell my team to focus on outcomes, not output. So when we look at the data here and And we see A2X in pull request merge rate. For me that's output right. It's like you know we're getting PR created. PR merged in terms of the outcome is a good code. Is it quality code. Is that code being patched.

36:58 You know, is the second PR the fixing of the first PR? How do we actually identify if this is truly, you know, being faster, making us putting us in a better place than we would have been without AI coding? Yeah, the. Those are important outcomes for the engineering machine. And we see lots of folks, you know, concerned about quality. And right now there aren't huge smoking guns in terms of quality. You know, I think the main reason for that that we haven't seen, you know, quality go off the rails is that engineering teams, by and large, are responsible stewards of quality.

37:37 So they're not merging bad code. You see the difference in merge rates. So you do see some, you know, upticks in reverse. It's you know, it can be hard to analyze because we also have seen a growth in fixing forward. So as teams get faster, they may not even file bugs or revert code. They may just apply a fixed forward so we can look like just more building and not necessarily fixing a problem. But all of those things, you know, those again, are engineering outputs.

38:11 They're still not necessarily business outputs. So you know. Did you build product or did you just kind of tinker with things that maybe nobody else cares about but engineering? And even if you build new product, you know, I think, you know, we have go to market organizations and things that, you know, if I can wave a magic wand and build five times the products tomorrow for jellyfish, could we sell five times the products? Probably not. We would need, you know, a whole new type of AI enablement and our go to market or to accelerate that team to a place where they can even capture all of that value.

38:47 So that's that's what I mean by, you know, we're in a world where we invented these jet engines and we're still kind of putting them into the cars we used to have. We haven't designed the rest of the the vehicle to accommodate this. One thing that we've massively accelerated, which is code generation. Yeah. And as we talk about, you know, becoming more effective, becoming more productive and scaling that across an organization, I guess, I guess one important role or area that perhaps we've, you know, not looked at as much is more around the, the AI enabler, or maybe it's part of the platform team, the developer experience team, the the group that essentially encourages and helps teams and organizations roll out AI adoption across their organization.

39:36 Do you have experience or any any wisdom you can kind of like share in your in your discussions about how organizations are getting good at these types of rollouts and gaining adoption through sharing within that organization. The it's it's hugely important, Simon, because, you know, we see this kind of uncanny valley where small teams because they're small and also because they tend to be younger and have younger code bases, you know, the company itself, just as have all the baggage of an older company.

40:12 You know, they can move fast because they're small. And we actually see very large companies that have the types of teams you're describing. You know, they they are able to make big investments. And the ones who are the most successful are, you know, they're moving with a purpose and with deliberation to enable folks, they understand that, you know, you need to invest in the tools, you need to invest in the context engineering. You can do the best in training.

40:37 This is this just doesn't happen for free. And where people struggle is in the middle where they don't have those teams, they may not even have ahead of experience. That person certainly doesn't have a whole team of people that can develop tooling. And when you kind of leave every team to figure this out on their own, it can be really challenging. So you know what success looks like. In my experience, what I've seen is putting dedicated resources to it, making making it someone or multiple people's jobs, full time jobs to figure out how to make this transformation in the org.

41:11 You know, making clear investments with money, with training, with time going slow to go fast, you know, and then the thing we talked about before, which is just picking some things and not getting analysis paralysis or stressing about the fact that certainly tomorrow new things are going to come out that makes the thing you just did not the optimal thing to do, but just building those muscles. You know, it comes down to continuous learning and continuous evolution.

41:40 At the end of the day, if you can't build that muscle in your org to just evolve continuously, we're all on this treadmill for a while for the duration, right? It's just going to keep accelerating and you're just going to have to keep learning new tools and evolving your alerts. That's what it looks like. Very interesting. Let's keep going on that. On that thought process of of of giving advice to engineering leads and organizations. What would you say is the biggest thing that folks are getting wrong, or the biggest misconception that engineering leads have today around AI?

42:17 I think one. One big misconception is what we were just talking about would be the flip side of assuming there's some silver bullet. If I just give these these tools to people, they'll just magically work. And I think part of them is conception is we've kind of done the easy part already in the sense that, you know, we gave a lot of people fancier IDEs, and there were real gains associated with that, you know, fancier autocomplete.

42:44 They sort of, you know, they bootstrapped a lot of coding and there was some training involved. But by and large, there have been real gains associated with getting a Copilot or a Cursor in the hands of developers, but then kind of leaving everything else in the process the same. So, you know, the misconception is the understanding that to get to the next level, to get to the promise of, you know what, you and I believe AI native development is really going to look like.

43:11 Those are big cultural changes. The big skill changes, the big architecture changes. They are much bigger rocks to move and require much bigger investments. The payoff is going to be massive. You know, you and I certainly believe that. But I think, you know, some people aren't prepared for how different that is than what we've done so far. The companies. Yeah, it's almost like questioning everything, right? It's like, don't assume that just because you used it the last ten years just fine.

43:42 It's the right way of doing it with AI. And that said, actually, there's probably a lot of things that we should very intentionally keep that AI sometimes makes it easy to drop a lot of the, the, the typical best practices and good hygiene processes of good software development. You know, things like your your usual code reviews and the best practices, your style guides potentially not so much, but the good practices in your workflow. A lot of that we need to maintain because it's sometimes too easy to vibe code and throw something.

44:14 And it's making that assumption of actually, yeah, I'm sure there's no security issues here and and so forth. Yeah. I mean, these. Are, you know, people are having these these kind of soul searching arguments. We've had an argument internally at jellyfish that a lot of folks have had, which is, you know, we've done the two person kind of code review thing of if I open a request, a different person needs to approve it. I can't approve my own pull request for production.

44:41 If an agent wrote the code, is that a different person? Can I review code that a robot wrote that I never read? It's kind of, you know, it logically, it kind of makes sense. And trying to decide, you know, is the, you know, the throttling of, I mean, reviews are becoming a bottleneck. Is that worth the value of having two different eyes on the code? And it you know, I think about I love the way you phrased all that about what things do we want to hold on to?

45:09 Because, you know, one of my kind of, you know, pet peeves or things that I'm kind of trying to understand about the current conversation is this obsession with agent coherence and how long an agent can run on its own. It's such a weird metric, and I understand why, you know, we're grading and benchmarking agents on this, but then when you translate that to the real world, I always use the analogy. You know, if you think of your team is the best engineer, the one who goes off the longest without talking to you into the cave and any questions at all.

45:44 But we, you know, we tend to complain about those folks who are who don't know when to come up for air. So it gets back to our kind of interactivity question of, you know, feedback and communication is a core part of a really well functioning engineering team. And what does that look like, an agentic world. When is that agents communicating to each other, and when is it they need to be communicating to us and having them go off on their own for a very long time is just is a is an anti-pattern.

46:11 It's a dysfunction. Yeah, yeah. So you're interesting. Nick, where can people go to to keep on track? Keep, keep on, on top of all the all the great, the great work you're doing with your reports, keeping up with the data, that kind of if you go to Jellyfish Engineering Trends, we are publishing updates to our benchmarks monthly. So, you know, we have a history of previous months, things around adoption and impact quality, you know, growth and engineering.

46:48 So we have all those things and we're adding new stuff every month. So we just added some token insights. We're going to add some more in the next cut. So the you can't keep up with all the exciting things that we can see in this data. You know, it's, you know, not enough hours in the day or tokens in the world to to. Yeah, I would love to understand about what's actually happening in the real world right now. Amazing.

47:13 Well great shout. Let's keep up to date with that and looking forward to the upcoming reports. Thank you very much, Nick. No. Pleasure. Yeah. Thank you so much. This was fun. Absolutely. And thanks everyone for listening. And be sure to tune in to the next episode. Bye for now. The AI Native Dev is brought to you by the package manager for skills and context. Your hosts are Guy Podjarny and me, Simon Maple. Our producer is Tom Dowler.

47:36 The AI Native Dev is not just a podcast, it's a community. And we host monthly meetups at the Tessl offices in central London. Visit Tessl.io forward slash community to learn more and I hope to see you there.

Summary

The podcast episode features a discussion on the current state of AI in software development, particularly focusing on agentic coding and the impact of AI tools on developer productivity. Nicholas Arcolano from Jellyfish shares insights from their extensive data collection, revealing significant gains in coding efficiency and the challenges faced when using multiple AI agents.

- Developers using AI tools are creating and merging twice as many pull requests compared to those who do not use AI.
- Most engineers effectively interact with 1-2 AI agents, with a barrier to using more due to limited human attention.
- AI adoption among developers has reached a median of 71%, with some companies reporting up to 90% usage.
- Smaller, newer codebases benefit more from AI tools, while older, complex codebases see little to no gains.
- AI-generated pull requests have a lower merge rate (60%) compared to human-generated ones (80%), indicating workflow challenges.
- Companies are advised to invest in dedicated resources for AI integration and avoid analysis paralysis by focusing on a few effective tools.
- The conversation around AI also includes the need for maintaining quality and best practices in software development.
- Engineering leaders must demonstrate the value of AI tools to CFOs, balancing productivity gains with costs associated with token usage.
© transcribe · For agents Built with care and craft by Gokul Rajaram