transcribe

Agentic AI and the Future of Software Development: S3 E4

AMD · 41m · transcribed 13d ago
More from AMD Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Introduction to Claude Code and Boris Churney

What is the background of Boris Churney and the development of Claude Code?

Boris Churney, creator of Claude Code at Anthropic, has a diverse background in engineering and product development, having worked at Meta and Instagram. His journey through various startups has shaped his hands-on approach to code development.

  • Boris has a rich career history that includes roles in engineering, business, and design.
  • His experiences in startups, despite many not succeeding, provided valuable learning opportunities.
  • Claude Code represents a culmination of his diverse skills and experiences.
# 8:16

Evolution of Claude Code's Capabilities

How did Claude Code evolve to become more effective?

The product saw exponential growth starting with Opus four in May 2025, as the model became more sophisticated and capable of functioning like a coworker, leading to a rethinking of workflows.

  • The growth of Claude Code was gradual but accelerated significantly after a key update.
  • The introduction of long-running asynchronous agents has transformed user interaction.
  • Continuous disruption and innovation are integral to the development process.
# 16:32

Skills for Success in Engineering

What skills are essential for engineers in the current landscape?

Successful engineers today are empirical, curious, and autonomous, able to adapt quickly to new data and bring ideas to market without extensive processes.

  • Adaptability and curiosity are crucial traits for engineers in a fast-paced environment.
  • The bottleneck has shifted from engineering to the speed of idea implementation.
  • A CEO-like mindset is beneficial for engineers to drive projects effectively.
# 24:48

Measuring ROI in Technology Development

How should companies approach ROI when implementing new technologies?

Companies often focus too much on cost-cutting rather than improving returns. Encouraging experimentation among all employees can lead to unexpected innovations and better ROI.

  • A balanced approach to ROI should emphasize both investment reduction and return improvement.
  • Creating a culture of experimentation can lead to innovative solutions from unexpected sources.
  • Empowering all team members to contribute can yield surprising results.
# 33:04

Building Trust and Security in AI Models

What measures are taken to ensure trust and security in AI models?

Training models to resist common attacks like prompt injection is crucial for building trust. The architecture includes anti-prompt injection training and runtime classifiers to enhance security.

  • Trust in AI models is built through rigorous training and security measures.
  • The ability to run models safely over extended periods is increasingly important.
  • As AI capabilities grow, so does the need for robust security against potential vulnerabilities.

Transcript

0:03 Welcome to Advanced Insights, where we provide just what the show name suggests, Advanced Insights and some of the most exciting trends and topics in technology. On this episode, I'm joined by Boris Churney. He's the creator and the head of Claude Code at Anthropic. 9 01:00:18,614 --> 01:00:20,783 Boris spent years as a principal engineer at Meta and Instagram before joining Anthropic, where he joined the Labs team 12 01:00:25,705 --> 01:00:27,206 and his spin up project became what is now Claude Code.

0:34 Boris, it is great to have you on this show. I mean, I think about what you've done in terms of driving a Claude Code development, the kind of impact it's had on the industry. I can imagine about a dozen areas we could go in this conversation. But first of all, I just want to welcome you. Thanks for joining us. Yeah, thanks, Mark, for having me here. Okay, so we got to jump in a little bit. And I've just gotten to know you recently, but I read your history.

1:04 I know about you. I follow you. And it's just quite an arc that you've had from a career standpoint. I just would love to hear your thoughts of how you ended up driving the creation of Claude Code. I mean, you were working on Instagram at Meta and and, you know, other focuses even before you even got to Meta. It seems like you've always been really hands on in driving code development. Yeah, so most of my career is tiny startups that mostly didn't work.

1:40 And every time it didn't work, I learned a lot. And every time I did a startup, I wore a different hat. It was sometimes engineering, sometimes business, sometimes products, sometimes design, you know, like at a small company, sometimes you have to do all this stuff. And so from kind of the beginnings of my career, I've kind of felt this need to break down the walls of, you know, engineers do this and designers do this. And at some point, I did a bunch of different startups, I worked in VC for a little bit, but Meta for a little while.

2:16 And something that I've always done as I built products is I built developer tools so that I can build better products. Because there's this sort of intuition that if you have an idea and you want to build it, the easier it is to use your tool, the better you can build it. Imagine if you're, you know, a carpenter and, you know, like, hammers are really badly designed. Let's say it takes like, you know, let's say it takes like three people to hold a hammer to nail, you're going to hammer a lot less often.

2:45 But if you have like a really well designed hammer, and it's very easy to use, and you enjoy using it, then you're going to hammer a lot of nails. And so, you know, like when I build product, I love to build the developer tools that make it even easier to build the product. And the thing that matters is not the tool, the thing that matters is the end result of the product. And so this is the you know, this is the mindset that I took to Meta, this is the mindset that I brought to Anthropic, you 90 01:03:06,449 --> 01:03:07,575 know, at Anthropic, the thing that 91 01:03:07,575 --> 01:03:08,201 matters more than anything is the safety mission.

3:10 This is the thing that brought me there. This is I think that the most important problem to work on. And you know, my role at Anthropic is to 97 01:03:17,752 --> 01:03:19,587 build the tools that let people experience the models, so they can understand where this thing is headed. And that lets us make it a little more safe and a little more capable, and hopefully a little more delightful to use. Well, I tell you, I love hearing that there's, there's a bit of Malcolm Gladwell in your in your store of the art, because you look at what you've done, and you you probably couldn't have done it without that myriad of experiences.

3:44 I'm very fortunate in chip design, I've had a chance to be in super small teams where we had to do really every aspect of chip design. And I feel so fortunate because of my career, it gave me a perspective of how all the myriad of pieces have to come together. And you were so hands on in the myriad of roles. And you think about how in Anthropic, 122 01:04:05,132 --> 01:04:07,468 you're having to put it all together to enable people to have that end result and to have be safe.

4:12 All right. Yeah, absolutely. And you know, like when we hire for the team, this is the kind of people I look forward to. Like I typically don't look for a person that, you know, has a computer science education, and then, you know, like worked and, you know, big tech for for whatever I like, I love people with non traditional backgrounds, I love people with all sorts of different experiences. Because the more that you've seen the more different ways you've approached problems, the more you can bring to this problem.

4:35 Yeah, you have a passion for what you do. And I assume you look for that in the people you're hiring as well. Yeah, yeah. And I love people that are passionate, not just about their work, but about other stuff too. So like, I love people with like, you know, side projects and hobbies and just like if someone has like an amazing personal website, and also does, you know, leather working on the weekend, and they've gotten to a point where they do it at like a very high level.

5:01 That's a really good sign. This is a person that's going to perform all on the job too. I would just love to explore a little bit of how all of that background and you jumped in with Anthropic. 161 01:05:11,574 --> 01:05:14,493 Could you have imagined when you started just thinking like, hey, we could do really a code assist and autocomplete or, you know, just really helping code segments. Tell me when you start to realize that you could progress from there to really these agentic processes that now are significantly advancing code to where basically everyone's saying, look, we're using coding agents.

5:42 We're not coding in the traditional way, which means it's now we've fundamentally transformed software development. And it's all happened in a period of months. How did you imagine that? No. When I joined Anthropic, the 183 01:06:02,458 --> 01:06:04,085 model at the time was on a 3.5. And that's the first model that I think broke through the mainstream. And it made people realize that, oh, man, maybe there's something to this idea of a general model that if you apply it to a particular problem like coding, it can be really good.

6:17 And, you know, like, for Anthropic, this 192 01:06:19,975 --> 01:06:20,643 has just always been the research direction. Like, we have always been focused on safety and security and enterprise and coding. It's just always been this laser focus. And we wanted to build these models for a long time. The reason we want a good coding model is because the model exists as software. And as it gets more intelligent, the way it interacts with the world is through software.

6:44 And so you want the code that it writes to be good so that it can interact well so that we can study it and we can make it safer and better aligned. And so, you know, for a while for Anthropic, we knew we wanted to build 212 01:06:54,927 --> 01:06:55,928 some kind of coding product. And my job was to figure out what's that product. And when I joined, again, it was it was on a 3.5 and all the products at the time, they were fancy autocomplete.

7:05 And, you know, at the time, this was cutting edge. And I remember my first time, you know, using some of these products, it was just amazing. And we felt that very soon the models are going to get to a point where they're going to be much more capable than this. So for coding, it's not a matter of completing a single line. The model is going to be able to write an entire function, an entire file, entire feature, an entire project.

7:29 Now, I think by the end of this year, they're going to be able to build entire businesses. And so, you know, we kind of saw this, you know, that we call this the exponential. This is just the way that, you know, the models are developing. There's these scaling laws that describe the way that model intelligence goes up over time. Somehow, the laws are continuing. And in fact, it's accelerating. And, you know, it's an empirical kind of observation.

7:52 It's not obvious why it keeps happening. But it seems to keep happening. And so all we did is we kind of traced this line, and we're like, all right, right now the model is here, it's going to get here. The product that's here is not going to be enough to let people experience the power of the model. And so we wanted something that was purely agentic. And honestly, Claude code didn't work very well for, you know, the first six months.

8:15 It actually it took time, it took time to catch up. And it wasn't until I would say Opus four, in May of 2025, where it really started to work. And, you know, like, since then, the growth for, you know, the product has been exponential also, it was not at first, but that's when it became exponential. And we've been finding more and more ways to use this, like as the model gets more sophisticated, we figure out more ways to unlock it in the product, we want Claude tag.

8:44 And this is kind of the latest iteration of this. It's long running asynchronous agent that you interact with like a coworker. It has great common sense, it has a great sense of when to jump in and when not to. And, you know, this will just keep happening, we're going to keep disrupting ourselves, we're going to keep figuring out what is, what's the next form factor. You know, you talk about sort of predicting this, this, this line of progress.

9:07 But to me, how did you, how did you think about architecting, you know, really multi agents and how running these multiple agents can actually have us rethink how we implement our workflows. I mean, that, that is what's been shocking to me is, we actually have the opportunity to re-architect, re-engineer how all of us are getting our work done. And I have to wonder, did this, was that architected into your development or is that sort of just a natural outcome of the capabilities that you were developing?

9:46 I think it's a little bit of both. If I think back a couple years, how did, how did we use models? We, you know, like they did this like one line autocomplete and we had agentic workflows, always what we called agentic workflows, but it was essentially these deterministic systems where a step at a time might call out to an LLM to do some bit of computation, but fundamentally, the workflow was deterministic. And I think what's changed as the model got more intelligent is, we've realized actually it's much better to use the model as the coordinator for this workflow.

10:16 And so what you do is the model drives and you give it tools to pull in context. You give it tools to interact with the outside world. You ask it to do something and then you let it figure out how to orchestrate these. You no longer tell it, okay, you do step one, then step two, and then, if not two, then, you know, this, and then else this, and it's just not how it works anymore. You give it a goal.

10:38 You give it the starting point, you give it access to tools and then it goes out. And this has happened, I think, as models got more intelligent. This is just overwhelming the reason. But I think like within Claude code, we've also learned this. At the very beginning, we would kind of spoon feed the model context. And then as the model got more sophisticated, we were like, okay, maybe actually this isn't it. And we moved away from that a little bit.

11:02 And, you know, like, for example, we have this like file called "Claude.md". It's loaded for every session. Anything more and more what we're seeing with customers is they're actually moving more to and they're moving more to tools and more to MCPs because this way the model has a little bit more control over how to load this on. Well, I have to ask, I mean, you guys across Anthropic are 354 01:11:25,364 --> 01:11:26,615 using Claude every day.

11:26 I mean, so you're your customer number one, like before the rest of us see it, you know, you're making sure it accomplishes what you need. Tell us about your your typical day. I mean, is it is it is is Claude like an extension of your your body? Is it creating your work output? Yeah, this is this is sort of the weird thing about how Anthropic works. 365 01:11:50,973 --> 01:11:53,100 At Anthropic Claude is at the center of everything that we do.

11:54 It's every single business process, every single bit of everything everyone does throughout the day. Claude is just at the center of it. A way I sometimes describe it is, let's say you're a new hire and you're onboarding at Anthropic. 374 01:12:07,573 --> 01:12:09,324 Typically, if you have a question like, you know, where is the office? Where's the address? No, I get it. You know, most companies you would go and search a wiki or you would ask a co worker at Anthropic you ask Claude If you have a question like where's the code base and how do I get access to it?

12:22 Again, you don't go to wiki, you go to Claude and you ask Claude. If your questions like, you know, like I took a trip, I have an expense report, how do I file that expense report? And this is how it begins. But then as you kind of get deeper into the work, what remains at the center of everything. And so you know, Claude writes the code, he does the the code review, he does the security review, it brainstorms and comes up, come comes up with new product ideas.

12:48 It aggregates user feedback, it triages incidents. Every single part of the engineering at STLC, Claude is at the center of it too. And you know, like that now now this is happening, not just for engineering, it's happening for product and for design and for marketing and for GTM for every function. This is what we're seeing. And the tools look a little different. It might not be you know, Claude code in a terminal, it might be you know, a tool like cowork instead.

13:13 But it's still Claude at the center. And I think this is actually what I see with the most successful adopters of Claude is you put it just at the center of every process. We were talking a little bit before about my favorite Harvard Business Review article from like from the 90s. And it was talking about like, I think it was like 92 or 96 or something. And it was, it was titled something like computers are here, why are we not taking the productivity benefits?

13:38 And you know, like, this was a big question that people were debating. And essentially, the case the article made was that there were two kinds of companies, there was one kind of company that they took a computer and they put it somewhere in the corner of the office. And then they kept all their, you know, paper and pen processes and all their filing cabinets and everything stayed as is but now there's a computer in the corner.

13:58 These companies are not seeing a productivity improvement. And then there's the companies that took the paper, you know, the filing cabinet and threw it away, they took the paper and pen and burned it. And now there's a computer at the center of every process. And those are the companies that unlock productivity improvements. And, you know, since the beginning of this year at Anthropic, we've seen an 8x 447 01:14:17,202 --> 01:14:19,079 increase in code output per engineer.

14:19 And this is just unheard of in the industry, like, you know, typically, companies see something like a few percent a year. And you know, now our biggest customers that are using Claude code, now they're starting to see something like this, they're seeing like, you know, 50, 100%, 150% improvement. And the thing that we did to get the 8x is we systematically improve a bottleneck at a time. So, you know, first, coding is the bottleneck, we throw Claude at it and Claude does the coding.

14:49 The next bottleneck is the coder view. And, you know, then we have Claude do that. 465 01:14:53,739 --> 01:14:55,324 The next bottleneck is, you know, like generating materials for GTM. You know, maybe we throw Claude at that. 468 01:14:59,578 --> 01:15:01,246 And so just like one step at a time, we un-bottleneck every part of the process. Well, what I love that is because you're driving your development by your own real world needs. It's a self-reinforcing cycle to make sure that the value of where you're spending your resources.

15:16 You're an AI native company, so you're adopting that change at an astonishing rate. But I'd love to hear your experience of how you think about the role of leaders, of engineers. I just think you're ahead of most of the rest of us in starting to see that at play and seeing how you're taking people with you. Stepping back a little bit, engineering is this thing that it's just always been changing.

15:46 And, you know, as engineers, we're no strangers to this. You know, my grandfather actually programmed punch cards in the Soviet Union. You know, like I grew up, I programmed like basic and assembly growing up, and then, you know, I kind of learned higher level languages, like as I went. I worked in JavaScript for a while, and in JavaScript, every engineer knows the frameworks change. Like every month, there's a new set of frameworks, and everyone has to learn new skills.

16:12 And so I think like for engineers, it's always been changing. The languages have always been changing. The frameworks have always been changing. What's happening now is it's accelerating. And so we went up a level of abstraction. We went from, you know, hardware to punch cards, then we went from punch cards to source code. Now we went from manipulating source code to agents. And now we're going up a level again from managing agents to managing like loops and routines.

16:37 And the crazy thing is that these last two steps happen in the span of two years. It's just it's faster than it was before. And I think it will continue to accelerate. And so when I think about the engineers and the people that are really successful, I think it's people that are empirical and curious. So you're able to look at the data and you're able to adjust the approach. You don't always assume the same thing you always do is going to keep working.

17:01 You're able to kind of take feedback from your work. And so if there's new data, you can have a different approach. It's people that are autonomous. Because as engineering is no longer the bottleneck. Now we're actually bottlenecked on the speed with which we can look at ideas and bring them to market. And you know, do so safely. And the people that are really effective at this, they're not people that, you know, like they have to have this big process for how to get an idea to market and they need to collaborate with like 10 other teams to do it.

17:32 They're sort of like one person armies. It's like the people that are most effective are this kind of like CEO archetype. Like you're able to come up with an idea, you're able to talk to users, you're able to look at the data to build, to iterate and to bring it to the market. And we actually see this coming not just from engineers, but from all sorts of functions. I think having this engineering background helps today. But it's actually I don't think it's really essential anymore to doing this kind of thing.

17:58 But it does change the kind of skills that we need to emphasize. You need to be adaptable. You need to be fearless of change. I love what you said. You have to be curious. You have to, you know, really be experimenting. I mean, that's part of what I think these new agentic workflows, new meaning, you know, like we've said, literally in the last, you know, seven months to a year, It changes how you can iterate. So how do you be curious?

18:28 You can experiment now in ways you could never experiment before. I mean, your learning cycles are vastly accelerated, aren't they? Yeah, totally, totally. And I think it's on, you know, it's on leaders to make space. If you don't adjust for this, and you don't create space, everyone's going to keep doing it the old way. And you're going to have to force everyone to do it in a new way. But really, for a lot of people, it's actually just really exciting to get to try all these tools to get to experiment.

18:51 And so I actually see a lot of my role as creating space for people to try stuff out to feel safe doing it. You know, they're not going to get a bad performance review if they, you know, experiment with a new idea, because it might actually work. And my other job is to just give people context. So to give them business context, product context, so that they can make better decisions. But companies that aren't AI native, these are lessons, like what you went through that it's going to take showing best practice, it's going to be, you know, educating the workforce and really driving culture change to be as you described.

19:27 Yeah, yeah, that's right. And our job is to help companies along the way. One thing that I've seen companies go through is this kind of natural transition from not using AI, to every engineer running one agent, to every engineer running 10 agents, then 100 agents, then 1000 agents. And there's sort of these like traits that go with each of these. So you know, with one agent, engineers are using this agent to kind of like write the code, and you're kind of focused on it, you're still kind of single threaded, you're working on one task, then at some point, you kind of get to the point where you kind of trust the agent, and you're like, okay, maybe maybe it's actually doing the right thing, maybe while it works, I can start a second one.

20:06 Yeah. And engineers kind of naturally figure this out, like maybe you'll have like multiple checkouts of the same repository or, you know, multiple views of the same code. So while the agent works, you can, you know, work on something else. And maybe you can scale up to like 10 agents this way, what you're doing is you're essentially around robining through the agents, you start the work in one, you move on to the second, you do the work there, you move on to the third.

20:25 And this is kind of the workflow that we've now support really well in the desktop app. So both for Claude code, and for co work, 655 01:20:33,662 --> 01:20:35,539 you start a session, you just start a second session, you can move on, you can have multiple stuff running in parallel. But then it gets really interesting once you scale this up a little bit more. And I think like when I look at most companies, they're still somewhere between step one and two.

20:47 And when I look at Anthropic, we're 664 01:20:49,302 --> 01:20:49,928 probably somewhere around step three on average. So most engineers are running dozens or hundreds of agents. And the way you do this is you have the agent run another agent. And we have a lot of tools for this. So for example, if you're on the sub agents, you have agents running basically sub agents. Is that how you think about it? Or we've coined that phrase if I don't know if it's industry standard.

21:11 No, no, that's exactly it. So you have the sub agents and the sub agents can start sub agents. So it can it can actually go quite deep. And for Claude code, we support up to five 682 01:21:18,373 --> 01:21:20,375 layers now, this kind of nesting. And there's different ways to support this on the product side. So you know, we've been experimenting with for example, cloud execution. So instead of the agent running locally on your computer, it runs in the cloud.

21:31 And you can start an agent from the desktop app or the mobile app and I actually do a lot of my coding for my phone now. I love that. And it just runs on the cloud. And this is actually what lets you scale up. And then you know, we're continuing to iterate on this to make it even more native and to scale up to 1000s of agents. And the way we get there is with dynamic workflows.

21:51 And this is essentially using Claude to orchestrate just very large teams of agents to do the most complex work, like large code based migrations, like, I think stripe just use this to migrate. You use something like like it would have taken months, and then it took like four days to do like a 10,000 line, Scala to Java migration. There's a lot of companies that are using this for large migrations, like we just migrated Bun from Zig to Rust.

22:15 And this is like a JavaScript runtime. It's a lot of code. And so for the for these kinds of like very, very big workloads, that would have taken weeks or months or years, essentially a way to do it is this divide and conquer where you give the model a lot of agents, and you just tell it like, go, go figure it out, do this big task. So it goes back to what we said earlier, it, it's going to be a journey, it's going to take, it's going to take education.

22:41 And it is going to take culture change. And when you do it right, we actually have similar examples of what you just described, we have, we have created a Rust based application, our lead of AI software development, did it over a weekend for an application that we needed, we thought it would be a competitive gap. And, and over a weekend, now we had to validate it and get it through all the, you know, the test regression and everything where it could be shipped.

23:10 But that just is on was unheard of before. I mean, it, it's just astonishing to me what multi agents can do. Again, you go back to this projection you had, it sounds like that's, this is exactly what you expected, though. Right? I think this is, this was what we expect. Yeah, you know, you never, you never know exactly how it's going to scale. But if, you know, for some reason, you know, this is, this is continuing.

23:37 And, you know, like for, for these traditional scaling laws, and it's funny, the people that wrote the scaling laws paper, like the first few authors like went and started Anthropic, because I 759 01:23:47,063 --> 01:23:48,440 think they knew where it was going, they know that safety becomes very important, and security becomes very important. And so, you know, traditionally, the scaling was say that we scale as a function of the size of the network, the amount of data and the amount of compute.

24:01 And what we're seeing now is it's also a function of the test time compute. And, you know, like, essentially, it's a fancy way of saying how many tokens you throw at it. And so, from a product design point of view, and kind of model training point of view, we want to make it so you can more productively throw more tokens at a problem to get a better result. Because if you do it naively, you spend more tokens, but you don't get a better result.

24:21 And so like one, one example of this is effort levels. So like in the model, you can configure the effort, the higher the effort, the more the maximum amount of tokens the model will be willing to spend on that problem. And so this idea of like multi agent and dynamic workloads, I actually see it as another version of test time compute, you throw more tokens at it, and by orchestrating the agents and having the model orchestrated, you can use these tokens productively to get a better result.

24:44 And Boris, I've heard you state that it's really important to throw those tokens at it, because otherwise, you don't even know the roofline of what you can achieve. Maybe you can share with our audience what you mean by that. Yeah, so I think there's a lot of companies right now that are thinking about ROI, and how you measure this. And I think some companies are actually thinking about it wrong, because they're focused purely on the I part of it.

25:11 And so they think about, okay, here's my investment, how do I reduce it? How do I cost cut? Obviously, this is important. And you should just do this. And you know, there's a bunch of tools for this, you can use opus plan mode, you can use advisor models, you can use a lower effort setting, you can use a cheaper model. So maybe everything doesn't need opus, or you know, maybe you can use haiku, maybe you can use Sonnet. 820 01:25:30,166 --> 01:25:32,085 And of course, you like you write evals to kind of figure this out.

25:33 But I think actually the far more important part is the R. So how do you improve the return that you're saying for models? And I think the biggest lesson here that we've seen from successful customers is, you want to give engineers the freedom to experiment, you want to give everyone the freedom to experiment, so that they come up with new use cases, you want people to experiment, and you want to create this culture where people feel safe experimenting, and they will surprise you.

25:57 It's not going to be your most senior engineer that comes up with like some brilliant idea, it's going to be maybe a new grad, you know, it's not going to be some engineer that figures out, how do you automate this marketing process, it might be a marketing person, someone in the corner of the org that you've never met, but that comes up with a brilliant idea, and they're because they were experimenting. And so the thing you want to encourage upfront is this experimentation, if there's an internal use case, and it takes off, or there's a product, and it takes off, and it uses a lot of tokens, but it's successful.

26:22 That's when you go in, and you want to optimize it. But if you don't do the first part, you'll never find the opportunity. And then once you find the opportunity, then you can optimize it like any other engineering problem. Of course, I love that description, because what you also highlighted is there's a democratization of opportunity. We've actually had junior engineers take a problem, a chip design problem, that was a bug. We had an engineering team had been working for weeks, couldn't solve the bug, and one of our junior engineers said, I think I can solve that, threw a set of a genetic process and came and solved this incredibly remote, 20 things at once had to occur for this chip bug to realize itself.

27:08 This was obviously before we shipped a product. This was in our debug and test phase. It actually shocked to the core a number of our senior engineers, and that actually drove up adoption of these approaches at AMD, because every time you actually have to experience it. You can talk, you can give it lip service, but once you actually have a problem that you couldn't solve or actually experience a huge productivity gain, just leader after leader, I see them immediately converted on the spot to believers.

27:42 That's right. That's right. And you need a certain amount of humility also going in. I think this is a lesson I had to learn over and over and over again. I tried to do things my way, which was before I just gave up and let Claude drive. 899 01:27:55,937 --> 01:27:56,396 I manually did it. For example, maybe this was a year ago, it was debugging a bug also. I remember there was one time I was debugging something and there was a new person that just joined the team.

28:09 I was debugging it by hand. I was running a profile order to figure out what's going on. They just asked Claude to do the same thing. I was like, no, no, no, there's no way Claude won't be able to figure it out. Even you said that. Even I said that because in some ways, I'm biased because I've worked with Claude through all the model generations. Somewhere in my head, I was still stuck on this older model generation, which is just not where the capability was at that point because the model was improving.

28:36 Within 20 minutes, they came up with a solution and I didn't. They had the right one. I've just learned this lesson over and over. I think it's like, as the model gets more advanced, my relationship with the model changes. It's less that I have to handhold it and I have to micromanage it. It becomes more like working with a senior engineer. I trust them. I give them context. I give them goals. I check in. The less I trust them, maybe the newer they are to the problem, the more often I check in.

29:04 Really, it's about setting the right guardrails and then letting the model do the work. Boris, I have a question for you. It is a democratization of what teams can do. What we also see is that people will be running in parallel and we're going to have five ways to drive an improvement out of a workflow. How do you think about that? How do you think about measuring? How do you at Anthropic, in those cases, 953 01:29:34,619 --> 01:29:36,454 pick the solution to go with?

29:37 What's the role of benchmarking and measuring to be able to, in this agentic area, make these decisions and not end up with this massive bifurcation of approaches as we all are trying to solve problems? Yes. I think it's exactly this. There's actually two tools. One is evals, so benchmark. 964 01:29:57,767 --> 01:29:58,684 Then one is vibes. 965 01:29:59,769 --> 01:30:01,312 I think both have their place. For something like a workflow that you're going to repeat maybe thousands, tens of thousands, hundreds of thousands, millions of times, you definitely want to have evals because the evals let you 971 01:30:10,822 --> 01:30:12,448 measure, let's say a new model comes out, you swap it in in the harness and you want to make sure that it improves and this gives you a way to measure it.

30:18 For other things, like maybe a product experience, vibes actually go pretty far 977 01:30:22,625 --> 01:30:24,293 because there's a cost to writing evals. 978 01:30:24,293 --> 01:30:24,752 You don't want to write it for everything. You want to pick and choose. When can intuition when you use something? When does that tell you something? Then when do you actually need to evals to tell you something? When you have this bifurcation and all these different products, I think one of the blessings of having a model is it's actually very easy to migrate and merge code.

30:45 When you pick the approach that you want, usually it's quite easy to just ask Claude, "Hey, migrate all these other call sites to this workflow." That's that. I'm going to shift gears yet again a little bit. I want to talk about something that all of us have to focus on. I think it's just actually centered and Anthropic in the 1002 01:31:05,418 --> 01:31:06,335 approach you all have taken. That's the broad topic of trust and verification. I had on a previous episode, Kathy Pham who lectures at Harvard.

31:16 We talked about governance and how you build in the process. There's one side, there's a governance of how to ensure there's safe AI, and then there's the verification. How do you ensure that you're driving the accuracy that you need? Both come together to provide trust and the deployment of AI. I know it's important to you. You've set up from the outside in our chat here, but maybe we could just spend a minute thinking about how do you build that in?

31:48 Is it a checklist or is it actually built in your own workflows that you run every day? Yeah, it's many, many layers. Exactly like you said, for Anthropic, 1029 01:32:03,684 --> 01:32:05,228 this is just at the core of what we do. As models get increasingly powerful and they get more and more central to the business process, you need it to be safe, you need it to be reliable, you need it to be trustworthy. You need it to be aligned with whatever the intent is of the person that's using the model.

32:21 There's just so many ways that we approach this at different layers. The most basic is model alignment. We have a very large team working on it. We publish a lot of research about it and we have a lot of blog posts about what we're learning. This is just many, many years of research go into this. Essentially, there's a lot of things that go into alignment, but for example, an element of alignment is truthfulness. The model says the right thing.

32:46 Another element is not being sycophantic because a failure mode for models, for example, is agreeing with everything the user is saying, but actually a good model that I like working with will push back. I think with some of our later models, like Opus 4.7, definitely Opus 4.8, we made a lot of progress and so if we train the model well, then if I suggest something and it's a bad idea, the model will push back and it makes me trust it a little bit more.

33:20 There's other layers too. When we talk about trust, another element of this, there's a lot of stuff wrapped into trust, but another element is security. This is another element that goes into training, is for example, anti-prompt injection training. On every model we publish, there's a system card that talks about how resilient is the model to common attacks like prompt injection. Opus 4.7, Opus 4.8, Fable, these are just the least prompt injectable models in the industry by margin of I think like 5 or 10x.

33:52 It's like a really big margin. So you had to architect that in. I mean, how did you do that? There's a lot of training that goes into it. This is just part of the model training process. And we train the models to be resistant to these kinds of attacks and especially when you combine it with a runtime classifier for prompt injection, which we actually also do, the success rate in practice now is near zero. And obviously this is important because as the models do more, as they interact with more systems, being resistant to these very common attacks becomes very important.

34:24 As the model gets more capable, we need a way to run the model for a longer period of time safely. And now there's workloads that run for days, weeks, months at a time. And obviously this means that there can be a person sitting there. And around the same time, what we were realizing is if there's a person that's deciding yes or no every time, at some point the person actually just stops reading and they just say yes, yes, yes, because they're tired of pressing yes.

34:49 And I noticed myself doing this too. I would kind of stop reading the bash commands. I would just say yes. And our security people were also realizing this. And so the thing that they started working on is, is there a way to route this to a classifier so the classifier can decide? And they started this work, it took many months to kind of get this to a mature state. And in the end, through evals, through 1125 01:35:16,544 --> 01:35:18,796 red teaming, through pen testing, we're able to show that this is actually much safer.

35:20 And so at Anthropic, we run on auto mode. 1129 01:35:23,509 --> 01:35:24,760 And we recommend this to all our customers too. Well, Boris, I have to ask you, I mean, you're removing bottlenecks at every stage. But how do you think about scale going forward? I mean, the complexity of what you're taking on at Anthropic is just, as you 1138 01:35:41,068 --> 01:35:42,027 said, growing it dramatically at every release. How do you think about managing scale, actually maintaining that pace?

35:50 Yeah, I think a lot of it comes down to following the scaling laws and continuing to reinvent the product as it scales up. A lot of it is about finding awesome partners to work with to support the scale and support this kind of compute. And I think a lot of it is figuring out the right guard rails so that we can make sure it's safe. Because, again, as the model becomes more and more core to what we do, we have to make sure that it does the right thing.

36:25 But I'm actually, Mark, I'm curious how you think about it. You're no stranger to scaling. And you've seen this through multiple generations for CPUs, for GPUs. How do you think about it? Well, scaling for us at AMD is actually fundamental because the old Moore's law, so Moore's law was that you could double the transistor density, double the performance, but stay at the same power envelope and cost envelope. And that was fantastic for years. But, Boris, about 10 years ago, that started trailing off.

36:56 Not that we don't still, with each new semiconductor node, get more productivity and more density. But the costs are going up. And as you've certainly seen, the power goes up. So we're having to reimagine how we optimize to be able to maintain. We're basically needing to stay on the old Moore's law pace without the physics of the semiconductors supporting that old Moore's law pace. So it's a huge challenge. But we're up for it.

37:27 And frankly, these agentic flows are proving immensely helpful to let us scale because they're helping us manage at a bigger, I'll call it, state space, more variables, more of our chip designs that can be optimized. And I don't think we could continue to scale without these agentic workflows. Literally, it's coming about at the perfect time. It's like a perfect confluence to enable us to continue to scale.

37:58 And you talked about partnership. I mean, that's at the core of who we are at AMD. Our partnership with you, our partnership across the industry, that's also fundamental for us to be able to scale because we like to deeply partner to make sure we are using the best. Like just as the stories you've told us here, we want to be first adopters of those best practices that you're doing at Anthropic. 1208 01:38:22,730 --> 01:38:26,692 So great question for us at AMD for our whole chip industry, paramount of innovating to keep at scale.

38:32 I have to ask the question. I mean, you've talked about really seeing us progress on kind of the line you projected. So I have to ask, what do you see coming in the next two to three years, which in AI time is like an eternity. Two to three years is a really long time. Let's maybe, I'll think about maybe six months. This is just, otherwise my prediction is just going to be way off supper. I think in six months, we are going to see agents running for longer on average.

39:07 We're going to see people running more agents on average. We're going to see them be better aligned with your intent. So it's going to be less course correction, less hand holding. There's going to be a lot more autonomy. So I think it will actually become pretty normal for most people and definitely most engineers and hopefully more than that to be running agents for days or weeks at a time. This would just become a normal thing as opposed to something a few engineers are doing.

39:34 I think by the end of this year, we're going to start to see Claude building bigger and bigger things. So it won't just be features, it's going to be products. We might start to see entire startups being built by Claude. So yeah, progress marches on and we're excited to make it safe and to make it something that's for people to use. Boris, thank you very much. Thanks for sharing your insights. Thanks for your vision and the impact you're having on the industry.

40:05 And my God, it's an exciting future. Thanks so much. Thank you, Mark. I truly enjoyed chatting with Boris. I think Boris really gave us a preview of how work changes going forward in this agentic era of AI. A few takeaways that really stood out. The value of hiring people with non-traditional backgrounds to bring new perspectives and passions. Also, how it Anthropic, Claude is at the 1268 01:40:29,481 --> 01:40:31,567 center of the employee experience. Gentic AI democratizes the opportunity for anyone at any level to make an impactful difference.

40:39 And lastly, when thinking about return on investment, don't just focus on the investment, focus on the return, the freedom of experimentation that provides will be what comes up with the biggest wins. Thank you for joining us today. I look forward to bringing you more in-depth conversations on the cutting edge of technology, industry insights, and visionary perspectives from some of the brightest minds in the industry. You can watch the video show on AMD's YouTube channel, or you can listen to an audio-only platforms, including Spotify, Amazon Music, and YouTube Music.

41:13 Technology is always advancing and I can't wait to share more insights with you next time.

Summary

Boris Churney, creator of Claude Code at Anthropic, discusses the evolution of AI-driven coding tools and their transformative impact on software development. He emphasizes the importance of safety, experimentation, and the integration of AI into everyday workflows, which has led to significant productivity gains at Anthropic.

- Churney's career spans various roles in startups and big tech, emphasizing hands-on experience in product development.
- Claude Code aims to enhance coding efficiency and safety, evolving from simple autocomplete to sophisticated coding agents.
- Anthropic's approach focuses on democratizing AI, allowing employees at all levels to leverage AI tools for impactful contributions.
- The company has seen an 8x increase in code output per engineer through systematic bottleneck removal using Claude.
- Churney advocates for a culture of experimentation, where employees feel safe to innovate and explore new ideas.
- The future of AI in coding includes longer-running agents and the potential for AI to build entire products or startups autonomously.
- Trust and verification in AI are critical, with ongoing efforts to ensure model alignment and security against prompt injection attacks.
- Successful companies integrate AI deeply into their processes, moving away from traditional methods to fully embrace AI's capabilities.

Questions Answered

What is the background of Boris Churney and the development of Claude Code?

Boris Churney, creator of Claude Code at Anthropic, has a diverse background in engineering and product development, having worked at Meta and Instagram. His journey through various startups has shaped his hands-on approach to code development.

How did Claude Code evolve to become more effective?

The product saw exponential growth starting with Opus four in May 2025, as the model became more sophisticated and capable of functioning like a coworker, leading to a rethinking of workflows.

What skills are essential for engineers in the current landscape?

Successful engineers today are empirical, curious, and autonomous, able to adapt quickly to new data and bring ideas to market without extensive processes.

How should companies approach ROI when implementing new technologies?

Companies often focus too much on cost-cutting rather than improving returns. Encouraging experimentation among all employees can lead to unexpected innovations and better ROI.

What measures are taken to ensure trust and security in AI models?

Training models to resist common attacks like prompt injection is crucial for building trust. The architecture includes anti-prompt injection training and runtime classifiers to enhance security.

© transcribe · For agents Built with care and craft by Gokul Rajaram