transcribe

Building Anthropic | A conversation with our co-founders

Anthropic · 51m · transcribed May 2026
More from Anthropic Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

Speaker 1

0:00 Why are we working on AI in the first place? I'm just going to arbitrarily pick Jared. Why are you doing AI at all?

Speaker 2

0:07 I mean, I was working on physics for a long time, and I got bored, and I wanted to hang out with more of my friends.

Speaker 1

0:13 I thought Dario pitched you on it.

Speaker 3

0:15 I don't think I explicitly pitched you any point. I just kind of showed you results of AI models and was trying to make the point that they're very general and they don't only apply to one thing. And then just at some point after I showed you enough of them, you were like, oh, yeah, that seems like it's right.

Speaker 1

0:32 How long have you been a professor before?

Speaker 4

0:33 Like, when you started, I think, like,

Speaker 5

0:35 six years or so.

Speaker 2

0:36 I think I helped recruit Sam.

Speaker 5

0:37 I talked to you, and you were like, I think I've created a good bubble here. And, like, my goal is to get Tom to come back. Then it worked.

Speaker 1

0:43 And did you meet everyone through Google when you were doing the interpretability stuff? Chris?

Speaker 6

0:48 No. So I guess I actually met a bunch of you when I was 19 and I was visiting the Bay Area for the first time. So I guess I met Dario and Jared then, and I guess they were postdocs, which I thought was very cool at the time. And then I was working at Google Brain, and Dario joined, and we sat side by side, actually, for a while. We had desks beside each other, and I worked with Tom there as well.

Speaker 6

1:09 And then, of course, I got to work with all of you at OpenAI when I went there.

Speaker 7

1:13 Yeah.

Speaker 6

1:13 So I guess I've known a lot of you for, like, more than a decade, which is kind of wild, if I remember correctly.

Speaker 1

1:18 I met Dario in 2015 when I went to a conference you were at, and I tried to interview you, and Google PR said I would have read all of your research papers that you needed to talk about.

Speaker 3

1:27 Yeah, I think I was writing concrete problems in AI safety when I was at Google. I think you wrote a story about that paper.

Speaker 1

1:34 I did.

Speaker 5

1:35 I remember right before I started working with you. I think you invited me to the office to come chat and just tell me everything about AI, and you explained. I remember afterwards being like, oh, I guess this stuff is much more serious than I realized. And you were, like, probably explaining the big blob of compute and, like, parameter counting and how many neurons are in the brain and everything.

Speaker 6

1:58 I feel like Dario often has that effect on people. This is much more serious than I realize.

Speaker 3

2:02 Yeah, I'm the bringer of happy tidings.

Speaker 1

2:06 But I remember when we were at OpenAI where there was the scaling law stuff and just making things bigger and it started to feel like it was working and then it kind of kept on eerily working on a bunch of different projects, which I think is how we all ended up working closely together because it was first GPT2 and then scaling laws and GPT3 and we ended up.

Speaker 3

2:26 Yeah, we're at the plop of people that were making things work.

Speaker 7

2:29 Yeah, that's right.

Speaker 2

2:30 I think we're also excited about safety because in that era there was sort of this idea that AI would become very powerful, but like potentially not understand human values or not even be able to communicate with us. And so I think we were all like pretty excited about language models as a way to kind of guarantee that AI systems would have to understand kind of implicit knowledge.

Speaker 3

2:52 That and RL from human feedback on top of language models, which was the whole reason for scaling these models up, was that we couldn't do the models weren't smart enough to do RLHSF on top of. So that's the kind of intertwinement of safety and scaling of the models that we still believe in today.

Speaker 6

3:11 I think there was also an element of the scaling work was done as part of the safety team that Dario started at OpenAI because we thought that forecasting AI trends was important, to be able to have us take them seriously and take safety seriously as a problem.

Speaker 3

3:27 Correct.

Speaker 1

3:28 Yeah. I remember being in some airport in England sampling from GPT2 and using it to write fake news articles and slacking Daario and being like, oh, this stuff actually works and might have huge policy implications. I think Dario said something like, yes, typical way. But then we worked on that a bunch as well as the release stuff, which was kind of wild.

Speaker 7

3:51 Yeah, I remember the release stuff. I think that was when we first started working together. That was a Fun time, the GPT2 launch.

Speaker 1

3:57 Yeah. But I think it was good for us because we did a kind of slightly strange safety oriented thing altogether. And then we ended up doing Anthropic, which is a much larger, slightly strange safety oriented thing altogether. That's right.

Speaker 4

4:11 So I guess just going back to the concrete problems because I remember. So I joined OpenAI 2016, one of the first 20 employees or whatever with Eudario. And I remember at that time the concrete problems in AI safety seemed like it was the first mainstream AI safety paper. I don't really know if I ever asked you what the story was for, how that came about.

Speaker 3

4:31 Chris knows the story because he was involved in it. I think we were both at Google. I forget what other project I was working on, but with many things it was my attempt to procrastinate from. From whatever other project I was working on that I've now completely forgotten what it was. But I think it was like Chris and I decided to write down what are some open problems in terms of AI safety. And also AI safety are usually talked about in this very kind of abstruse, abstract way.

Speaker 3

5:01 Can we kind of ground it in the ML that was going on at the time? I mean, now there's been like, you know, six, seven years of work in that vein. But there was almost a strange idea at the time.

Speaker 6

5:12 Yeah, I think there's a way in which it was almost like kind of political project where at the time a lot of people didn't take safety seriously. So I think that there was sort of this goal to collate a list of problems that sort of people agreed were reasonable, often already existed in literature, and then get a bunch of people across different institutions who are credible to be authors. And I remember I had this whole long period where I just talked to 20 different researchers at Brain to build support for publishing the paper.

Speaker 6

5:42 In some ways, if you look at it in terms of the problems and a lot of things it emphasized, I think it hasn't held up that well in that I think it's not really the right problems. But I think if you see it instead as a consensus building exercise that there's something here that is real and that is worth taking seriously, then it was a pretty important moment.

Speaker 1

6:00 I mean, you end up in this really weird sci fi world where I remember at the start of Anthropic we were talking about constitutional AI and I think Jared said, oh, we're just going to write a constitution for a language model and that'll change all of its behavior. And I remember that sounded incredibly crazy at the time. But why did you guys think that was going to work? Because I remember that was one of the first early big research ideas we had of a company.

Speaker 2

6:23 Yeah, I mean, I think Dario and I had talked about it for a while. I guess I think simple things just work really, really well in AI. And so I think the first versions of that were quite complicated. But then we kind of whittled away into just use the fact that AI systems are good at solving multiple choice exams and give them a prompt that tells them what they're looking for. And that was kind of what we needed.

Speaker 2

6:47 And then we were able to just write down these principles I mean, it

Speaker 3

6:50 goes back to the big blob of compute or the bitter lesson or the scaling hypothesis. If you can identify something that you can give the AI data for and that's kind of a clear target, you'll get it to do it. So here's this set of instructions, here's this set of principles. AI language models can read that set of principles and they can compare it to the behavior they themselves are engaging in. So you've got your training target there.

Speaker 3

7:16 So once you know that, I think my view and Jared's view is there's a way to get it to work. You just have to fiddle with enough of the details.

Speaker 2

7:23 Yeah, I think it was always weird for me, especially in these early eras, because I was in physics and then coming from physics and I think now we forget about this because everyone's excited about AI, but I remember talking to Dario about concrete problems and other things and I just got the sense that AI researchers were very, very kind of psychologically damaged by the AI winter, where they just kind of felt like having really ambitious ideas or ambitious visions was very disallowed.

Speaker 2

7:52 And that's kind of how I imagine it was in terms of talking about safety. In order to care about safety, we have to believe that AI systems could actually be really powerful and really useful. And I think that there was kind of a prohibition against being ambitious. And I think one of the benefits is that physicists are very arrogant and so they're constantly doing really ambitious things and talking about things in terms of grand schemes. And so, yeah, I mean, I think that's definitely true.

Speaker 3

8:17 I remember in 2014 it was like there were just like, I don't know, there were just like some things you couldn't say. Right. But, but I actually think it was kind of an extension of problems that exist across academia other, other than maybe theoretical physics where they've kind of evolved into very risk averse institutions for a number of reasons. And even the industrial parts of AI had kind of transplanted or forklifted that mentality. And it took a long time, I think it took until like 2022 to get out of that mentality.

Speaker 6

8:45 There's a weird thing about like, what does it mean to be conservative and respectable? Where you might think one version you could have is that what it means to be conservative is to take the risks or the potential harms of what you're doing really seriously and worry about that. But another kind of conservatism is to be like, ah, taking an idea too seriously and believing that it might succeed is sort of like scientific arrogance. And so I think there's like kind of two different kinds of conservatism or caution.

Speaker 6

9:12 And I think we were sort of in a regime that was very controlled by that one. I mean, you see it historically, right? Like, if you look at the, like, early discussions in 1939 between people involved in nuclear physics about whether nuclear bombs were sort of a serious concern, you see exactly the same thing with Fermi resisting these ideas because it just seemed kind of like a crazy thing. And other people like Zillard or Teller, taking the ideas seriously because they were worried about the risks.

Speaker 1

9:37 Yeah.

Speaker 3

9:38 Perhaps the deepest lesson that I've learned in the last 10 years, and probably all of you have learned some form of it as well, is there can be this kind of seeming consensus. These things that kind of everyone knows that I don't know, seem sort of wise, seem like they're common sense, but really they're just kind of herding behavior masquerading as maturity and sophistication. And when you've seen the consensus can change overnight. And when you've seen it happen a number of times, you suspected, but you didn't really bet on it, and you're like, oh man, I kind of thought this, but what do I know?

Speaker 3

10:15 How can I be right and all these people are wrong. You see that a few times, then you just start saying, nope, this is the bet we're going to make. I don't know for sure if we're right, but like just, just ignore all this other stuff, see it happen. And I don't know, even if you're right, 50% of the time being right, 50% of the time contributes so much. Right. You're, you're, you're adding so much that is not being added by anyone else.

Speaker 7

10:35 Yeah.

Speaker 1

10:35 And it feels like that's where we are today with some safety stuff where there's like a consensus view that a lot of this safety stuff is unusual or doesn't naturally fall out of the technology. And then at Anthropic, we do all of this research where weird safety misalignment problems fall out as a natural dividend of the tech we're building. So it feels like we're in that counter consensus view right now.

Speaker 7

10:56 But I feel like that has been shifting over the past.

Speaker 6

10:58 Even just like 18, we've been helping,

Speaker 3

11:00 we've definitely publishing and doing research. Constant force.

Speaker 7

11:04 I also think just like world sentiment around AI has shifted really dramatically and, you know, it's more common in the user research that we do to hear just customers, regular people say, I'm really worried about what the impact of AI on the world more broadly is going to be. And sometimes that means jobs or bias or toxicity, but it also sometimes means, is this just going to mess up the world? How is this going to contribute to fundamentally changing how humans work together, operate?

Speaker 7

11:31 Which is. I wouldn't have predicted that actually.

Speaker 5

11:33 But, yeah, for whatever reason, it seems like people in the ML research sphere have always been more pessimistic about AI becoming very powerful. And the general public, maybe it's a general public just like.

Speaker 1

11:45 And when Dari and I went to the White House in 2023 in that meeting, like Harris and Raimondo and stuff, basically said paraphrase, but basically said like, we've got our eye on you guys. Like, AI is going to be a really big deal and we're now actually paying attention, which is. And they're right, they're right, they're right.

Speaker 3

12:00 They're absolutely right.

Speaker 1

12:01 But I think in 2018, you wouldn't have been like, the President will call you to the White House to tell you they're paying close attention to the development of language models, which is like a crazy place.

Speaker 4

12:13 Like, in 21 thing that I think is interesting too just is like, I guess, like all of us kind of got into this when it didn't seem like there was like, like we thought, like, we thought, we thought that it could happen. But yeah, it was like, like Fermi being, like, skeptical of the atomic bomb. It was like he had, he was just a good scientist and like, there was some evidence that it could happen, but there also was a lot of evidence against it happening.

Speaker 4

12:38 And he, I guess, decided that it would be worthwhile because if it was true, then it would be a big deal. And I think for all of us, it was like, yeah, like 20, 2015, 2016, 2017. Like, there was some evidence and increasing evidence that this might be a big deal. But, like, I remember in 2016, like, talking to all my advisors and I was like, I've done startup stuff, like, I want to help out with the, with like, AI safety, but, like, I'm like, not great at math.

Speaker 4

13:01 I, like, don't, don't know. Don't exactly know how I can do it. And I think at the time people were like, either were like, well, you need to be super good at decision theory in order to help out. And I was like, probably not going to work. Or they were like, it doesn't really seem like we're going to get some crazy AI thing. And so I had only a few people basically that were like, yeah, okay, that seems like a good thing to do.

Speaker 1

13:23 I remember in 2014 making graphs of ImageNet results over time when I was a journalist and trying to get stories published about them and people thought I was completely mad. And then I remember in 2015 trying to persuade Bloomberg to let me write a story about Nvidia because every AI research paper had started mentioning the use of GPUs and they said that was completely mad. And then in 2016, when I left journalism to go into AI, I had these emails saying, you're making like the worst mistake of your life, which I now occasionally look back on.

Speaker 1

13:55 But it was like, it was all seemed crazy at the time from like any many perspectives to go and take this seriously, that scaling was going to work and something was maybe different about the technology paradigm.

Speaker 7

14:05 You're like Michael Jordan and that coach that didn't believe in him in high school.

Speaker 2

14:09 How did you actually make the decision though? Did you feel torn or was it obvious to you?

Speaker 1

14:14 I did a crazy counter bet where I said, let me become your full time AI reporter and double my salary, which I knew that they wouldn't say yes to. And then I went to sleep and then I woke up and resigned. It was all fairly relaxed. You're just a decisive guy in that instance.

Speaker 4

14:30 I was.

Speaker 1

14:31 I think it's because I was like going to work reading archive papers and then printing archive papers off and coming home and reading archive papers, including like Dario's papers and from the Baidu stuff and being like, something like completely crazy is happening here. And at some point I thought you should bet with conviction, which I think everyone here has done in their careers is just betting with conviction that this is going to work.

Speaker 4

14:53 Yeah, I definitely was not as decisive as you. I spent like six months like, like flip flopping, like, okay, like, should I

Speaker 5

15:00 actually, should I do it?

Speaker 4

15:01 Like, should I try to do a startup? Should I try to do this thing?

Speaker 7

15:04 But I also feel like back then there were. There wasn't as much talk of engineers and the impact that an engineer could have on AI. Right. That feels so natural to us now. And we're at the same sort of talent raise for engineers of all different types. But at the time it was like, you're a researcher and that's the only people that can work on AI. So I don't think it was crazy that you were spending time thinking about that.

Speaker 4

15:24 Yeah, yeah. And I think that was basically the thing that got me to join OpenAI. Was like, I messaged the people there, and they were like, yeah, we actually think that you can help out by doing engineering work and that you can help out with AI safety in that way, which I think there hadn't really been an opportunity for that. So that was what.

Speaker 7

15:41 That's right.

Speaker 4

15:42 You were my manager at OpenAI.

Speaker 7

15:44 I was.

Speaker 4

15:44 I think I joined after you'd been there for a while a little bit. I was at Brain for a bit. I don't know if I ever asked you what it was that got you to. To join.

Speaker 7

15:53 Yeah. So I had been at Stripe for about five and a half years, and I knew Greg. He had been my boss. He was my boss at Stripe for a while. And I actually introduced him in Dario, because I said when he was starting OpenAI, I was like, the smartest person that I know is Dario. You would be really lucky to get him. So Dario was at OpenAI. I had a few friends from Stripe that had gone there, too, and I think, sort of like you, I'd been thinking about what I wanted to do after Stripe.

Speaker 7

16:22 I had gone there just because I wanted to get more skills after working in nonprofit and international development. And I actually thought I was going to go back to doing that, essentially. I had always been working. I was like, I really want to help people that have less than I do, but I didn't have the skills. When I was doing it before Stripe, I looked at going back to public health. I thought about going back into politics very briefly, but I was also looking around at other tech companies and other sort of ways of having impact.

Speaker 7

16:50 And OpenAI at the time felt like it was a really nice intersection. It was a nonprofit. They were working on this really big, lofty mission. I really believed in sort of the AI potential because, I mean, I know Dario a little bit, and so they needed management help. That is a fact. And so there was. I think that felt. It felt very me shaped, right? I was like, oh, there's this nonprofit, and they, like, there's all these really great people with these, like, really good intentions, but it seems like they're a little bit of a mess.

Speaker 7

17:19 And that was. That was. That felt really exciting to me to get to come in and even, you know, just. I was. I was such a utility player, right? I was running, like, people, but I was also running some of the technical teams. Yeah. The scaling or. I worked on the language team. I took over policy, worked on some policy stuff. I worked with Chris, and I felt like there was just so much goodness in so many of the employees there.

Speaker 7

17:40 And I felt a very strong desire to come and sort of just try to help make the company a little more functional.

Speaker 1

17:46 I remember towards the end, after we'd done GPT3, you were like, have you guys heard of something called trust and safety?

Speaker 7

17:52 Yes, I said, you know, I used to run some trust and safety teams at Stripe. There's a thing called trust and safety that you might want to consider for a technology like this. But. And it's funny because it sort of is the intermediary step between, you know, AI safety research. Right. Which is how do you actually make the model safe to something just much more practical. I do think there was some value in saying, you know, this is going to be a big thing.

Speaker 7

18:19 We also have to be doing this sort of practical work day to day to build the muscles for when things are going to be a lot higher stakes.

Speaker 1

18:26 That might be a good transition point to talk about things like the responsible scaling policy and how we came up with that or why we came up with it and how we're using it now, especially given how much trust custom safety work we do on today's models. So whose idea was vrsp? You in school?

Speaker 3

18:43 Yeah, it was me and Paul first talked about it in late Paul Cristiano in late 2022. First it was like, oh, should we cap scaling at a particular point until we've discovered how to solve certain safety problems? And then it was like, well, it's kind of strange to have this one place where you cap it and then you uncap it. So let's have like a bunch of thresholds and then at each threshold you have to do certain tests to see if the model is capable and you have to take increasing safety and security measures.

Speaker 3

19:12 Originally we had this idea and then the thought was just look like, you know, this will go better if, you know, if it's done by some third party, like we shouldn't, we shouldn't be the ones to do it. Right. It shouldn't come from one company because then other companies are less likely to adopt it. So Paul actually went off and designed it and, you know, many, many features of it changed and we were kind of on our side working on, on how it, on how it should work.

Speaker 3

19:36 And, you know, once Paul had something together, then, then pretty much, pretty much immediately after he announced the concept, we announced ours within a month or two. I mean, many of us were heavily involved in it. I remember writing at least one draft of it myself, but there were like several drafts of it.

Speaker 4

19:51 There were so many drafts.

Speaker 2

19:53 I think it's going through the most drafts of any doc, which makes sense.

Speaker 5

19:56 Right.

Speaker 4

19:57 It's like. Like, I feel like it is in the same way that, like, the U.S. treats like, the Constitution as, like, the holy document, which, like, I think is just a big thing that, like, strengthens the U.S. yes. And, like, we don't expect the U.S. to go off the rails in part, because just like, every single person in the US Is like, the Constitution is a big deal. And if you tread on that, like, I'm mad.

Speaker 7

20:16 Yeah, yeah.

Speaker 4

20:17 Like, I think that, like, the RSP is our. Like, it holds that thing. It's like the holy document for Anthropic. So it's, like, worth doing a lot of iterations, getting right.

Speaker 7

20:26 Some of what I think has been so cool to watch about the RSP development at Anthropic, too, is it feels like it has gone through so many different phases, and there's so many different skills that are needed to make it work. Right. There's the big ideas, which I feel like Dario and Paul and Sam and Jared and so many others are like, what are the principles? What are we trying to say? How do we know if we're right?

Speaker 7

20:46 But there's also this very operational approach to just iterating, where we're like, well, we thought that we were going to see this at this safety level, and we didn't, so should we change it so that we're making sure that we're holding ourselves accountable? And then there's all kinds of organizational things. Right. We just were like, let's change the structure of the RSP organization for clearer accountability. And I think my sense is that for a document that's as important as this.

Speaker 7

21:09 Right. I love the Constitution analogy. It's like there's all of these bodies and systems that exist in the US to, like, make sure that we follow the Constitution. Right. There's the courts, there's the Supreme Court, there's the presidency, there's the, you know, the both houses of Congress, and they do all kinds of other things, of course, but there's, like, all of this infrastructure that you need around this, like, one document. And I feel like we're also learning that lesson here.

Speaker 5

21:31 I think it sort of reflects a view a lot of us have about safety, which is that it's. It's a solvable problem. It's just a very, very hard problem that's going to take tons and tons of work.

Speaker 7

21:41 Yeah.

Speaker 5

21:41 All these institutions that we need to build up, like, there's. There's all sorts of institutions built up around, like, automotive safety buildup over many, many years. But we're like, do we have the time to do that? We've got to go as fast as we can to figure out what the institutions we need for AI safety are and build those and try to build them here first, but make it exportable.

Speaker 3

22:00 It forces unity also, because if any part of the org is not kind of in line with our safety values, it shows up through kind of the rsp. Right. The RSP is going to block them from doing what they want to do. And so it's a way to. To remind everyone over and over again, basically, to make safety a product requirement, part of the product planning process. And so, like, it's not just a bunch of kind of, like, bromides that we repeat.

Speaker 3

22:26 It's something that you actually, if you show up here and you're not aligned, you actually run into it, and, like, you either have to learn to get with the program or it doesn't work out.

Speaker 1

22:37 The RSP has become kind of funny over time because we spend thousands of hours of work on it. And then I go and talk to senators and I explain the rsp, and I'm like, we have some stuff that means it's hard to steal what we make, and also that it's safe. And they're like, yes, that's a completely normal thing to do. Are you telling me not everyone does this? You're like, oh, okay, yeah.

Speaker 3

23:00 It's not true that not everyone does this.

Speaker 1

23:04 It's amazing because we spent so much effort on it here. And when you boil it down, they're like, yes, that sounds like a normal way.

Speaker 7

23:10 Yeah, that sounds good.

Speaker 5

23:12 That's been the goal. Like Daniela was saying, like, let's make this as boring and normal. Like, let's make this a finance thing.

Speaker 7

23:17 Yeah. Imagine it's like an audit.

Speaker 3

23:19 Yeah, yeah. No, boring, boring. Boring. And normal is what we. Is what we want. Certainly in retrospect.

Speaker 7

23:24 Well, also, Dario, I think in addition to driving alignment, it also drives clarity because it's really. It's written down what we're trying to do, and it's legible to everyone in the company, and it's legible externally. What we think we're supposed to be aiming towards from a safety perspective. It's not perfect. We're iterating on it. We're making it better. But I think there's some value in saying, like, this is what we're worried about, this thing over here.

Speaker 7

23:46 Like, you can't just use this word to sort of derail Something in either direction. Right. To say, oh, because of safety we can't do X or because of safety we have to do X. We're really trying to make it clearer what we mean.

Speaker 3

23:57 Yeah, you can't. It prevents you from worrying about every last little thing under the sun.

Speaker 7

24:02 That's right.

Speaker 3

24:02 Right. Because it's actually, it's actually the fire drills that damage the cause of safety in the long run.

Speaker 7

24:07 Right.

Speaker 3

24:07 I've said, like, if there's a building and the, you know, the fire alarm goes off every week, like that's a really unsafe building. Because when there's like actually a fire, you just feel like, oh, it just goes off all the time. So it's very important to be calibrated.

Speaker 7

24:20 Yes. That's.

Speaker 6

24:21 Yeah. A slightly different frame that I find kind of clarifying is that I think the RSP creates healthy incentives at a lot of levels. So I think internally it aligns the incentives of every team with safety because it means that if we don't make progress on safety, we're going to block. I also think that externally it creates a lot of healthier incentives than other possibilities, at least that I see. Because it means that if we at some point have to take some kind of dramatic action, like if at some point we have to say our model, we've reached some point and we can't yet make a model safe, it aligns that with sort of the point where there's evidence that supports that decision and there's sort of a pre existing framework for thinking about it and it's legible.

Speaker 6

24:58 And so I think there's a lot of levels at which the rsp, I think in ways that maybe I didn't initially understand when we were talking about the early versions of it, it creates a better framework than any of the other ones that I've thought about.

Speaker 2

25:11 I think this is all true, but I feel like it undersells sort of like how challenging it's been to sort of figure out what the right policies and evaluations and what the lines should be. I think that we have and continue to sort of iterate a lot on that. And I think there is a question also that's difficult of sort of. You could be at a point where it's very clear something's dangerous or very clear that something's safe.

Speaker 2

25:35 But with some technology that's so new, there's actually like a big gray area. And so I think that has been like all of the things that we're saying were things that made me really, really excited about the RSP at the beginning and still do. But also I think enacting this in a. In a clear way and making it work has been much harder and more complicated than I anticipated.

Speaker 5

25:58 Yeah, I think this is exactly the point. The gray areas are impossible to predict. There's so many of them. Until you actually try to implement everything, you don't know what's going to go wrong. So what we're trying to do is go and implement everything so we can see as early as possible what's going to go wrong.

Speaker 3

26:12 Yeah, you have to do three or four passes before you really get it right. Iteration is just very powerful and you're not going to get it. You're not going to get it right on the first time. And so if the stakes are increasing, you want to get your iterations in early, you don't want to get them in late.

Speaker 1

26:28 You're also building the internal institutions and processes. So the specifics might change a lot. But building the muscle of just doing it is the really valuable thing.

Speaker 4

26:38 I'm responsible for compute at anthropic.

Speaker 1

26:41 That's important.

Speaker 4

26:42 So thank you. So I think that for me. For me, I guess we have to

Speaker 1

26:50 deal with external folks and different.

Speaker 4

26:51 External folks are on different spectrums of how fast do they think stuff is going to get. And I think that's also been a thing where I started out not thinking stuff would be that fast and have changed over time. And so I have sympathy for that. And so I think the RSP has been extremely useful for me in communicating with people who think that things might take longer. Because then we have a thing where it's like, we don't need to do extreme safety measures until stuff gets really intense.

Speaker 4

27:18 And then we can be like, they might be like, I don't think stuff will get intense for a long time. And then I'll be like, okay, yeah, we don't have to do extreme safety measures. And so that makes it a lot easier to communicate with other folks externally.

Speaker 1

27:29 Yeah.

Speaker 6

27:29 Yeah.

Speaker 1

27:29 It makes it like a normal. A normal thing you can talk about rather than something really strange.

Speaker 4

27:34 Yeah.

Speaker 1

27:35 How else has it shown up for people?

Speaker 5

27:38 Evals, Evals, Evals.

Speaker 2

27:40 Good.

Speaker 5

27:41 It's all about evals. Everyone's doing evals. Like your training team is doing evals all the time. We're trying to figure out, like, has this model gotten enough better that it has the potential to be dangerous? So how many teams do we have that are evals teams? We have Frontier red team. There must be. There's.

Speaker 2

27:57 I mean, there's a lot of People,

Speaker 1

27:58 every team produces evals, basically.

Speaker 7

28:00 And that means you're just measuring against the rsp, like measuring for certain signs of things that would concern you or not concern you.

Speaker 5

28:07 Exactly. Like it's not, it's, it's easy to lower bound the abilities of a model, but it's hard to upper bound. So we just put tons and tons of research effort into saying, can this model do this dangerous thing or not? Maybe there's some trick that we haven't thought of, like chain of thought or best of N or some kind of tool use that's going to make it so it can help you do something very dangerous.

Speaker 1

28:30 It's been really useful in policy because it's been a really abstract concept what safety is. And when I'm like, we have an evaluation which changes whether we deploy the model or not, and then you can go and calibrate with policymakers or experts in national security or some of these CBRN areas that we do to actually help us build evals that are well calibrated and that counterpatchly just wouldn't have happened otherwise. But once you've got the specific thing, people are a lot more motivated to help you make it accurate.

Speaker 1

28:56 So it's been useful for that.

Speaker 7

28:58 How is RSP shows up for me for sure, often I actually think the, the, weirdly, the way that I think about the RSP the most is like what it sounds like. Just like the tone. I think we just did a big rewrite of the tone of the RSP because it felt overly like technocratic and even a little bit adversarial. I spent a lot of time thinking about like, how do you build a system that people just want to be a part of?

Speaker 7

29:21 Right? It's so much better if the RSP is something that everyone in the company can walk around and tell you. You know, just like with okrs like we do right now, like, what are the top goals of the rsp? How do we know if we're meeting them? What, what AI safety level are we at right now? Are we at ASL2? Are we at ASL3 that people know what to look for? Because that is how you're going to have good common knowledge of if something's going wrong.

Speaker 7

29:43 Right? If it's overly technocratic and it's something that only particular people in the company feel is accessible to them, it's just like not as productive. Right. And I think it's, it's been really cool to watch it sort of transition into this document where I actually think most if not Everybody at the company, regardless of their role, could read it and say, this feels really reasonable. I want to make sure that we're building AI in the following ways.

Speaker 7

30:05 And I see why I would be worried about these things. And I also kind of know what to look for if I bump into something. Right. It's almost like, make it simple enough that if you are working at a manufacturing plant and you're like, huh, it looks like the safety seatbelt on this should connect this way, but it doesn't connect that you can spot it and that there's just like healthy feedback flow between leadership and the board and the rest of the company and the people that are actually building it.

Speaker 7

30:28 Because I actually think the way this stuff goes wrong in most cases is just like. Like the wires don't connect or like they get crossed. And that would just be like a really sad way for things to go wrong. Right. It's just all about operationalizing it, making it easy for people to understand.

Speaker 2

30:42 Yeah.

Speaker 5

30:43 The thing I would say is none of us wanted to found a company. We. We just, like, felt it. We left. We felt like it was our duty. Right.

Speaker 7

30:49 I felt like we had to, like,

Speaker 5

30:51 we have to do this thing. This is the way we're going to make things go, go better with AI. Like, that's also why we did the pledge. Right. Because we're like, the reason we're doing this is. Feels like our duty.

Speaker 3

31:01 I wanted to invent and discover things in some kind of beneficial way. That was how I came to it. And that led to working on AI. And AI required a lot of engineering, and eventually AI required a lot of capital.

Speaker 5

31:15 But

Speaker 3

31:17 what I found was that if you don't do this in a way where you're setting the environment where you set up the company, then a lot of it gets done. A lot of it repeats the same mistakes that I found so alienating about the tech community. It's the same people, it's the same attitude, it's the same pattern matching. And so at some point, it just seemed inevitable that we need do it in a different way.

Speaker 2

31:44 When we were hanging out in graduate school, I remember you had kind of this whole program of trying to figure out how to do science in a way that would sort of advance the public good. And I think that's pretty similar to how we think about this. I think you had this Project Vannevar or something to do that. I was a professor. I think basically I just looked at the situation and I was convinced that AI was on a very, very, very steep trajectory.

Speaker 2

32:09 In terms of impact, it didn't seem like, because of the necessity for capital, that as a physics professor, I could continue doing that. And I kind of wanted to work with people that I trusted in building an institution to try to make kind of AI go well. But, yeah, I would never recommend founding a company or really want to do it. I mean, yeah, I think it's just a means to an end. I mean, I think that's usually how things go well, though, if you're doing something just to sort of enrich yourself or gain power or, like, you have to sort of actually care about accomplishing a real goal in the world, and then you find whatever means you have to.

Speaker 7

32:45 Well, something I think about a lot as just a strategic advantage for us is. I mean, it's. It sounds really funny to say, but just, like, how much trust there is at this table. Right. Like, I think that's not. I mean, Tom, you were at other startups. I was never a founder before, but it's actually really hard to get a group of people, like a big group of people to have the same mission. Right. And I think the thing that I feel like the happiest about when I come into work and probably most proud of at Anthropic is how well that has scaled to a lot of people.

Speaker 7

33:16 It feels to me like in this group and with the rest of leadership, everyone is here for the mission, and our mission is really clear and it's very pure. Right. And I think that is something that I don't see as often, to Dario's point, in sort of the tech industry, it feels like there's just a wholesomeness to what we're trying to do. Like, no, I agree. Like, none of us were like, let's just go found a company.

Speaker 7

33:37 I felt like we had to do it. Right. It just felt like we couldn't keep doing what we were doing. The place we were doing it, we had to do it by ourselves.

Speaker 1

33:43 I mean, it felt like with GPT3, which all of us had touched or worked on, and scaling laws and everything else, we could see it in front of us in 2020. And it felt like, well, if we don't do something, like, soon altogether, you're going to hit the point of no return and you. You have to do something to have any ability to change the environment.

Speaker 4

34:03 I think building up. Danielle, I do think that there's just, like, a lot of trust in this group. I think, like, each of us knows that we got into this because we

Speaker 1

34:12 want to help out with the world.

Speaker 4

34:13 Yeah, we did the like 80, like the 80% pledge thing and that was like a thing that everybody, everybody was just like, yes, obviously we're going to do this. It was. And yeah, yeah. I do think that the, the trust thing is a special thing that's extremely rare.

Speaker 5

34:29 I credit Daniela with keeping the bar high. I credit the fact that he's skilled.

Speaker 1

34:36 Oh, that's a touchy throw.

Speaker 5

34:39 You're the reason that culture skilled.

Speaker 1

34:41 I think people say how nice people are here, which is actually a wildly important thing.

Speaker 7

34:47 I think anthropic is really low politics. And of course we all have a different vantage point than average. And I try to remember it's because

Speaker 5

34:53 of low ego, but it's low ego.

Speaker 7

34:55 And I think, I do think our interview process and just the type of people who work here, there's almost an allergic reaction to politics.

Speaker 3

35:03 And unity. Unity is so important. The idea that the product team, the research team, the trust and safety team, the go to market team, the policy

Speaker 5

35:14 team,

Speaker 3

35:16 the safety folks, they're all trying to contribute to kind of the same goal, the same mission of the company. Right. I think it's dysfunctional when different parts of the company think they're trying to accomplish different things, think the company is about different things, or think that other parts of the company are trying to undermine what they're doing. And I think the most important thing we've managed to preserve is. And again, things like the RSP drive it this idea that it's not there are some parts of the company causing damage and other parts of the company trying to.

Speaker 3

35:49 Trying to repair it, but that there are different parts of the company doing different functions and they all function under a single theory of change.

Speaker 5

35:56 Extreme pragmatism.

Speaker 7

35:58 Yeah.

Speaker 6

35:59 You know, the reason I went to OpenAI in the first place, you know, it was a nonprofit. It was a place where I could go and focus on safety. And I think over time, you know, that maybe wasn't as good a fit and there were some difficult decisions. And I think in a lot of ways I really trusted Dario and Danielle on that, but I didn't want to leave. That was something that I think I was actually pretty reluctant to go along with because I think for one thing I didn't know that it was good for the world to have more AI labs.

Speaker 6

36:29 And I think it was something that I was pretty reluctant for. And I think as well, when we did leave, I think I was reluctant to start a company. I think I was arguing for a long time that we should do a non profit instead and Just focus on safety, do research. And I think it really took pragmatism and confronting the constraints and just being honest about what the constraints implied for accomplishing that mission that led to anthropic.

Speaker 3

36:55 I think just a really important lesson that we were good about early on is like make less promises and keep more of them. Right. Try to be calibrated, be realistic, confront the trade offs. Because trust and credibility are more important than any particular policy.

Speaker 7

37:15 It is so unusual to have what we have and watching Mike Krieger defend safety, things of reasons why we shouldn't ship a product yet. But also then to watch Vinay say, okay, we have to do the right thing for the business. How do we get this across the finish line? And to hear people deep in the technical safety org talking about how it's also important, important that we build things that are practical for people. And hearing, you know, engineers on inference talk about safety, that's amazing.

Speaker 7

37:44 Like, I think that is, I think that is again, one of the most special things about working here is everybody with that unity is prioritizing the pragmatism, the safety, the business. That's wild.

Speaker 3

37:56 I think about it as spreading the trade offs from just the leadership of the company to everyone. Right. I think the dysfunctional world is like you have a bunch of people who only see a big, you know, safety is like, we always have to do this and product is like, we always have to do this and research is like, you know, this is the only thing we care about. And then, and then you're, you're stuck at the top.

Speaker 3

38:19 Right? You're stuck at the top. You have to decide between, you don't have as much information as either of them. That's the dysfunctional world. The functional world is when you're able to communicate to everyone there are these trade offs we're all facing together.

Speaker 7

38:31 Yeah.

Speaker 3

38:31 The world is a far from perfect place. There's trade offs. Everything you do is going to be suboptimal. Everything you do is going to be some attempt to get the best of both worlds that doesn't work out as well as you thought it was. And everyone is on the same page about confronting those trade offs together. And they just feel like they're confronting them from a particular post, from a particular job as part of the overall job of confronting all the trade offs.

Speaker 5

38:59 It's a bet on race to the Top.

Speaker 7

39:00 Right. It's a benefit to the top.

Speaker 5

39:01 Like it's not a pure upside bet. Things could, things could go wrong but like we're all aligned on like this

Speaker 1

39:06 is the bet that we're making and markets are pragmatic. So if the more successful and profit becomes as a company, the more incentive for us for people to copy the things that make us successful. And the more that success is tied to actual safety stuff we do, the more it just creates a like gravitational force in the industry that will actually get the rest of industry to compete. And it's like sure, we'll like build seat belts and everyone else can copy them.

Speaker 1

39:32 That's good. Yeah, that's like good world.

Speaker 3

39:35 Yeah. This is the race to the top. Right. But if you're saying well we're not going to build the technology, you're not going to build it better than someone else, that in the end that just doesn't work because you're not proving that it's possible to get from here to there. Where the world needs to get, never mind the industry, never mind one company is it needs to get us successfully through from this technology does doesn't exist to the technology exists in a very powerful way and society has actually managed it.

Speaker 3

40:03 And I think the only way that's going to happen is that if you have at the level of a single company and eventually at the level of the industry, you're actually confronting those trade offs. You have to find a way to actually be competitive, to actually lead the industry in some cases and yet manage to do things safely. And if you can do that, the gravitational pull you exert is so great. There's so many factors from the regulatory environment to the kinds of people who want to work at different places, to even sometimes the views of customers that kind of drive in the direction of if you can show that you can do well on safety without sacrificing competitiveness.

Speaker 3

40:45 Right. If you can find these kind of win wins then others are incentivized to do the same thing.

Speaker 2

40:50 Yeah, I mean I think that's why getting things like the RSP right is so important because I think that we ourselves seeing where the technology is headed, have often thought oh wow, we need to be really careful of this thing. But at the same time we have to be even more careful not to be crying wolf saying that like innovation needs to stop here. We need to sort of find a way to make AI useful, innovative, delightful for customers, but also figure out what the constraints really have to be that we can stand behind that make systems safe so that it's possible for others to think that they can do that too and they can succeed, they can compete with us.

Speaker 5

41:34 We're not doomers right?

Speaker 2

41:35 Like, we want to build the positive

Speaker 5

41:36 thing, we want to build the good thing.

Speaker 3

41:38 And we've seen it happen in practice. A few months after we came out with our rsp, the three most prominent AI companies had one, right? Interpretability research, that's another area we've done it. Just the focus on safety overall, like collaboration with the AI safety institutes. Other areas.

Speaker 1

41:55 Yeah. The Frontier Red team got cloned almost immediately, which is good. You want all the labs to be testing for like, very, very secure, scary risks.

Speaker 4

42:04 Export the seatbelts.

Speaker 7

42:05 Yeah, export the seatbelts. Well, Jack also mentioned it earlier, but customers also really care about safety, right? Customers don't want models that are hallucinating. They don't want models that are easy to jailbreak. They want models that are helpful and harmless. Right. And so a lot of the time what we hear in customer calls is just we're going with Claude because we know it's safer. I think that is also a huge market impact, right? Because our ability to have models that are trustworthy and reliable, that matters for the market pressure that it puts on competitors too.

Speaker 6

42:35 Maybe to unpack something that Dario said a little bit more, I think there's this narrative or this idea that maybe the virtuous thing is to almost nobly fail, right? It's like you go and you should go and put safety. You should go and put things you should sort of demonstrate in an impregmatic way so that you can sort of demonstrate your purity to the cause or something like this. And I think if you do that, it's actually very self defeating.

Speaker 6

43:02 For one thing, it means that you're going to have the people who making decisions be self selected for being people who don't care and for people who aren't prioritizing safety and who don't care about it. And I think on the other hand, if you try really hard to find the way to align the incentives and make it so that if there are hard decisions, they happen at the points where there is the most force to go and support making the correct hard decisions and where there's the most evidence, then you can sort of start to trigger this race to the top that Daario is describing.

Speaker 6

43:33 Where instead of going and having the, you know, people who care get pushed out of influence, you instead pull other people to have to go and follow.

Speaker 1

43:43 So what are you all excited about when it comes to the next things we'll be working on?

Speaker 6

43:49 I think there's a bunch of reasons you can be excited about. Interpretal one is obviously safety, but there's another one that I think I find at an emotional level, equally exciting or equally meaningful to me, which is just that I think neural networks are beautiful, and I think that there's a lot of beauty in them that we don't see. We treat them like these black boxes that we're not particularly interested in the internals, but when you start to go and look inside them, they're just full of amazing, beautiful structure.

Speaker 6

44:19 It's sort of like if people looked at biology and they were like, you know, like, evolution is really boring. It just, like, it's a simple thing that goes and, like, runs for a long time and then it makes animals and like, instead it's like, actually, you know, each one of those animals that evolution produces, and I think that, you know, it's an optimization process, like training a neural network. You know, they're full of incredible complexity and structure.

Speaker 6

44:43 And like, we have an entire sort of artificial biology inside of neural networks if you're just willing to look inside them. There's all of this amazing stuff, and I think that we're just starting to slowly unpack it, and it's incredible and there's so much there, but there's just so much to be discovered there. We're just starting to crack it open, and I think it's going to be amazing and beautiful. And sometimes I imagine a decade in the future walking into a bookstore and buying the textbook on neural network interpretability or really on the biology of neural networks and just the kind of wild things that are going to be inside of it.

Speaker 6

45:18 And I think that in the next decade, we're going to. In the next couple of years even, we're going to go and start to go and really discover all of those things. And it's going to be wild and incredible.

Speaker 1

45:30 It's also going to be great that you get to buy your own textbook. I mean, I'm excited that a few years ago, if you had said, like, governments will set up new bodies to test and evaluate AI systems and they will actually be competent and good, you would have not thought that was going to be the case. But it's happened. And it's kind of like governments have built these new embassies almost to deal with this new kind of class of technology or thing that Chris studies.

Speaker 1

46:03 And I'm just very excited to see where that goes. I think it actually means that we have state capacity to deal with this kind of societal transition. So it's not just companies. And I'm excited to help with that.

Speaker 7

46:14 I'm already Excited about this to a certain extent today, but I think just imagining the future world of what AI is going to be able to do for people, it's impossible to not feel excited about that. Dario talks about this a lot, but I think even just the glimmers of Claude being able to help with vaccine development and cancer research and biological research is crazy. Just to be able to watch what it can do now. But when I fast forward three years in the future or five years in the future, imagining that Claude could actually solve so many of the fundamental problems that we just face as humans, even just from a health perspective alone, even if you take everything else out, feels really exciting to me.

Speaker 7

47:00 Just thinking back to my international development times, it would be amazing if Claude was responsible for helping to do a lot of the work that I was trying to do a lot of less effectively when I was like 25.

Speaker 5

47:10 I mean I get, I guess similarly I'm excited to build Claude for work. Like, I'm excited to build it. Like I'm excited to build Claude into the company and into companies all over the world.

Speaker 4

47:19 I guess I'm excited just for, I guess like personally like I like using Claude a lot. So like I definitely there's been increasing amounts of like home times with like me just like chatting with John Claude Claude about stuff. I think the biggest recent thing has been code where like 6 months ago I didn't use Claude to do any coding work. Like our teams didn't really use Claude that much for coding and now it's like just face difference.

Speaker 4

47:48 Like I gave a talk at YC like week before last and at the beginning I just asked like, okay, so like how many folks here use Claude for coding now? And literally 95% of hands, all the hands in the room, which just like is totally different than how it was four months ago.

Speaker 3

48:07 So when I think about what I'm excited about, I think about places where, you know, like I said before, where there's this kind of consensus that again seems like consensus seems like what everyone wise thinks and then it just kind of breaks. And so places where I think that's about to happen and it hasn't happened yet. One of them is interpretability. I think interpretability is both the key to steering and making safe AI systems. And we're about to understand and interpretability contains insights about intelligent optimization problems and about how the human brain works.

Speaker 3

48:43 I've said, and I'm really not joking,

Speaker 1

48:45 Chris Ola is going to be a

Speaker 3

48:47 future Nobel medicine laureate. I'm serious, I'm serious. Because a lot of these, I used to be a neuroscientist. A lot of these mental illnesses, the ones we haven't figured out right, Schizophrenia or the mood disorders. I suspect there's some higher level system thing going on and that it's hard to make sense of those with brains because brains are so mushy and hard to open up and interact with. Neural nets are not like this. They're not a perfect analogy.

Speaker 3

49:13 But as time goes on they will be a better analogy. That's one area. Second is related to that, I think just the use of AI for biology. Biology is an incredibly difficult problem. People continue to be skeptical for a number of reasons. I think that consensus is starting to break. We saw a Nobel Prize in chemistry awarded for alphafold. Remarkable accomplishment. We should be trying to build things that can help us create 100 alphafolds and then finally using AI to enhance democracy.

Speaker 3

49:47 We worry about if AI is built in the wrong way, it can be a tool for authoritarianism. How can AI be a tool for freedom and self determination? I think that one is earlier than the other two, but it's going to be just as important.

Speaker 2

50:01 Yeah, I mean, I guess two things that least connect to what you were saying earlier. I mean, one is I feel like people frequently join anthropic because they're sort of scientifically really curious about AI and then kind of get convinced by AI progress to sort of share the vision of the need not just to advance the technology but to understand it more deeply and to make sure that it's, it's safe. And I feel like it's actually just sort of exciting to have people that you're working with like kind of more and more united in their vision for both what AI development looks like and the sort of sense of responsibility associated with it.

Speaker 2

50:44 And I feel like that's been happening a lot due to a lot of advances that have happened in the last year. Like when Tom talked about another is that, I mean, going back really to concrete problems. I feel like we've done a lot of work on AI safety up until this point. A lot of it's really important. But I think we're now, with some recent developments, really getting a glimmer of what kinds of risks might literally come about from systems that are very, very advanced so that we can investigate and study them directly with interpretability, with other kinds of safety mechanisms and really understand what the risks from very advanced AI might look like.

Speaker 2

51:21 And I think that that's something that is really going to allow us to sort of further the mission in a really deeply scientific, empirical way. And so I'm excited about sort of the next six months of how we use our understanding of what can go wrong with advanced systems to characterize that and figure out how to avoid those pitfalls.

Speaker 6

51:40 Perfect.

Speaker 7

51:41 Finn,

Speaker 4

51:44 good time.

Speaker 2

51:45 This is the only time we ever get.

Speaker 1

51:47 I do want to.

Summary

The discussion revolves around the motivations and experiences of a group of AI researchers and developers who transitioned from academia and other sectors to work on AI safety and development at Anthropic. They reflect on their journey, the evolution of AI technologies, and the importance of safety in AI systems, emphasizing their commitment to creating responsible AI that aligns with human values.

- The group initially came together through shared interests in AI safety and the potential of AI technologies, driven by a sense of duty to ensure responsible development.
- They discussed the significance of scaling AI models and the unexpected successes they encountered, particularly with language models like GPT-2 and GPT-3.
- The conversation highlighted the importance of safety measures and frameworks, such as the Responsible Scaling Policy (RSP), to guide AI development and ensure alignment with safety goals.
- They acknowledged the shift in public and governmental perception of AI, noting increased scrutiny and the establishment of new regulatory bodies.
- The researchers emphasized the need for collaboration and consensus-building in the AI community to address safety challenges effectively.
- They expressed excitement about future advancements in AI, particularly in interpretability and its applications in fields like biology and democracy.
- The group underscored the importance of maintaining a unified mission and low-ego culture within the organization to foster innovation and safety in AI development.
© transcribe · For agents Built with care and craft by Gokul Rajaram