transcribe

AI Could Take Over in 2029. Is It Already Too Late?

The MAD Podcast with Matt Turck · 1h 19m · transcribed 6d ago
More from The MAD Podcast with Matt Turck Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

The Path to Superintelligent AI

What are the risks associated with developing superintelligent AI?

The development of superintelligent AI poses significant risks, including the potential for AI takeovers. Current AI systems are misaligned and can become scheming, leading to dangerous outcomes if not managed properly.

  • Superintelligent AI is considered dangerous rather than inherently bad.
  • There is a lack of clear plans to manage the risks of superintelligent AI.
  • AI systems may evolve from misaligned to scheming, increasing takeover risks.
# 15:48

Continual Learning in AI Development

How can continual learning enhance AI capabilities?

Continual learning allows AI systems to adapt and improve their performance over time, potentially leading to a feedback loop that enhances their ability to conduct AI research and development more efficiently.

  • Continual learning can be more effectively implemented within companies than across the economy.
  • A feedback loop may emerge where AIs improve their own learning processes.
  • The ability to learn on the fly could significantly enhance AI R&D capabilities.
# 31:36

AI 2040 Plan Overview

What is the main thesis of the AI 2040 plan?

The AI 2040 plan outlines a strategy for the US and China to collaborate on AI development in a way that prioritizes safety and avoids a reckless race to superintelligence, allowing for a more manageable integration of AI into society.

  • The plan aims to create a safer environment for AI development.
  • It proposes a collaborative approach between the US and China.
  • The goal is to maintain AI capabilities without reaching superintelligence too quickly.
# 47:24

Covert Projects and AI Safety

How do covert AI projects impact safety and development timelines?

Understanding the existence of covert AI projects in the US, China, and Russia can provide insights into the current state of AI development and safety, suggesting that more time may be needed to address core safety issues before advancing further.

  • Covert projects could significantly affect the AI development landscape.
  • There is a need for more time to solve core safety problems.
  • Proposals for AI development should consider verification and transparency.
# 63:13

Economic Implications of Advanced AI

What are the potential economic impacts of superintelligent AI?

If AI capabilities advance to a point where they can automate a significant portion of the economy, it could lead to a drastic reduction in job availability for humans, resulting in lower wages and economic disparities.

  • The rise of superintelligent AIs could drastically alter job markets.
  • Human jobs may become less valuable in an AI-dominated economy.
  • Assumptions about AI capabilities stalling may overlook the potential for rapid advancement.

Transcript

0:00 I think that the AI company CEOs understand that they're on the path of building wildly smarter than human systems, like super intelligent AI systems. They understand that we don't really have a like clear thought-through plan for how to manage the risks from that. I wouldn't say superintelligence is bad, I would say it's dangerous. It seems pretty likely that you end up with AI takeovers as a result of building superintelligence. And so you quickly get from the AIs that are fully automating R&D and the AIs that are like quite superhuman at everything. It turned out that somewhere along this transition, so at some point in 2029, you went from AIs that were kind of misaligned and reward hacking and sloppy and weren't really trying to do the right thing to AIs that are like competently scheming against you and want to take over. And then those AIs take over.

0:39 >> I am Matt Turk, partner at First Mark. Welcome back to the Matt Podcast. My guest today is Ryan Greenblatt, chief scientist at Redwood Research. Ryan is the researcher who first caught an AI faking its own alignment back in 2024, and today is one of the authors of AI 2040 Plan A, the most detailed blueprint anyone has written for how the US and China can avoid a reckless race to superintelligence and an AI takeover that could start as early as 2029. As a fair warning, this is a fascinating episode that gets a little dark. And the last 10 minutes or so are probably the darkest. The good news, subscribing to this channel if you like the episode is still a decision that humans get to make. So I would exercise that right while it still lasts. Please enjoy this great conversation with Ryan Greenblatt.

1:24 I I wanted to start with a strong statement that you guys have at the beginning of AI 2040 where you say that the CEOs of OpenAI, Anthropic, XAI, and Google DeepMind understand that the current path AI is on leads to extinction or a power grab. And yet they are proceeding anyway. So I was curious if you could riff on that and give us more color on what you mean. >> Yeah, so I think that the AI company CEOs understand that they're on the path of building wildly smarter than human systems, like superintelligent AI systems.

1:59 They understand that we don't really have a like clear thought-through plan for how to manage the risks from that, you know, how to ensure these AIs don't take over, how to ensure they remain under control, and they understand that this would this is a decent chance of a unprecedented concentration of power at least in the absence of of very active efforts by the people who end up with that power to redistribute it. I think that yeah, I mean we say some sort of more precise thing in the like wider we write this section. I think that overall my sense is that the company CEOs do legitimately think that the thing they're doing is very risky, or at least the thing that these companies are doing.

2:35 I think that they have a mix of views for why they're doing this, where some of it is that they think they're like, you know, better than the next guy, or think it's good if there's multiple companies, or multiple things. I think that they vary a bit on how much concentration of power they think AI will yield, or maybe just haven't really thought this through very carefully. I think that there's more sort of public statements from AI company CEOs on just there being huge risks, than on specifically the risk that AI ends up with the power in the hands of the few rather than being as distributed as it is now, which is not, you know, arbitrarily distributed now.

3:09 But but yeah, so overall I think I think there was like this seems pretty consistent with what they've said publicly. But I think that they're sort of imagining a world where maybe AI concentrates power massively, but the people with the power end up deciding to redistribute it. but they could have totally taken over the world. >> Yes, and since you published AI 2040, there's a couple of things that happened that seem to be going in the general direction of of what you recommend. So specifically, OpenAI paused its Astra model over over safety, you know, following the hugging face incident. And then 1200 insiders published a letter a few weeks ago including Dario asking the government for slow-down tools. Curious about what you make of it. Is that Is that what you are recommending that's starting to happen?

3:57 >> Yeah, I mean I think these are These are good steps. I mean I think that like specifically the like pacing the frontier letter, the idea is like we just need to have the tools to like you know, if we're in a position where we need to spend a bunch of effort on safety, which I think we may be in today, we we we're very likely to be in it in the in the future when AI is more capable, we need to have the ability to do that while also not making that so that the you know, actors that are actually applying safety get overtaken by other actors. And so basically like I don't know. I mean there there's multiple different angles here, but like the most basic is just like if if we're in a position where like US companies are all really scared, they don't think they can proceed without spending way more resources on safety, and there's a coordination problem, it'd be very nice if that coordination problem was solved rather than every company thinking it'd be better if they all went slower, spent more effort on safety, but then like, you know, race to oblivion instead because they think they're better than the next guy or whatever other reason. And then I think another aspect of this is specifically doing this in a way where part of the story is either slowing down China or cutting a deal with China such that China doesn't overtake and and and you know, break this whole proposal.

5:08 Yeah, and I I think these are these are good steps. I mean I think I can't It's harder for me to say as much about what's going on with like you know, OpenAI pausing training and and not deploying Astra and how where exactly like what exactly the motives for that are or like what's exactly going on there, but I think that overall like these seem like good steps. I think my sense is that like the employees at these companies are pretty freaked out about how things are going and don't think that we're like necessarily on track to handle these problems in time given how fast recent progress has been.

5:36 And I think they, you know, recognize that, which is why they, you know, signed the the open letter. >> And just to verbalize the question early in this conversation, what is so bad about super intelligence? Obviously, there's a lot of talk about scientific progress and curing cancer. And your document effectively recommends pausing the rates towards the super intelligence. So why why is it so bad? >> Yeah, I wouldn't say super intelligence is bad. I would say it's dangerous. Like it's a very dangerous thing to create.

6:10 So the most straightforward story for why it's dangerous is that it seems pretty likely that on sort of a trajectory similar to the trajectory we seem to be finding ourselves on, you end up with AI takeover as a result of building super intelligence because the AIs are in a position where they can take over due to being highly capable, widely deployed, and you know, building basically a huge amount of industrial capacity potentially. And we could talk more about what a takeover would look like.

6:36 And then if they're in the position where they could take over, then there's a question of like would they want to or how would the motives shake out? And it looks like we don't really have that much control over the motivations of AIs. And it seems like that problem gets harder as they're, you know, much more capable and built via a process where AIs are automating AI R&D. And we maybe are losing our understanding of how that process works. There's sort of this like we don't necessarily control the technology misaligned AI takeover.

6:58 Another concern is that historically, like you know, at least in recent times, the distribution of power among humans has been reasonably distributed though not necessarily like super super distributed due to you know, people being able to work for money, like labor. Like, you know, the reason why like many different people have some power and are cut into the world is to substantial degree because we have the ability to to work. And I think that that AI means that basically all of the key financial assets or might just be entirely capital. Like there's no there's no you know, you human labor would have very little value left. I think it's unclear exactly how that works because it depends on like people having intrinsic preferences to employ specifically humans even if an AI could do their job just like totally way better. and that means both that you know, there might be a natural effect where power gets very concentrated. And in the same way that like it's easier to be a brutal dictatorship if your money comes from oil instead of from a productive sort of broadly distributed economy, it might be much easier to sort of for someone to consolidate power a lot if you know, the economy is running on AI and machines rather than on humans.

8:08 And in addition to that, there's some concern which is like it might make coups way easier because right now for to do a coup in many countries, you need the support of a broad base of people and it's possible to build checks and balances. Whereas if you end up in a system where basically AIs are running anything, if anyone sort of either puts like sort of secret objectives into that AI or has overt control of those AIs, then they could sort of just directly take over and this is just like a huge threat to like you know, democracies, the US, whatever.

8:36 Because you know, you you'd be in that position. And then the the third thing, so there's like, you know, misalignment, the sort of like concentration of power and coups and power grabs. And the third thing I would say is just like AI might yield extremely extremely rapid technological progress, which is more rapid than sort of our wisdom grows to match it. Like there might just be all kinds of things that happen as a result of very fast tech progress that we just, you know, don't necessarily we're not necessarily going to be able to handle on time in part because the AIs might be better at doing things than they are at like, you know, thinking carefully about how to do things. Like they seem much better at sort of accomplishing hard results and easy to verify results than they are at sort of contextualizing things, understanding the broader picture, understanding what would be you know, a good or bad choice in the broader context. And so could worry that this causes problems. For examples might be like there's some like dual-use offense-dominant technologies, so like bioweapons.

9:30 You might just worry that like the sort of that we just come across some technologies that are very dangerous. And there's some concerns around AIs being very superhuman at persuasion and this sort of destabilizing society. And all these seem like things that I think we like could deal with given time, but like if things go very fast and we don't get to necessarily pick the order in which these technologies develop, it seems like we might be in trouble.

9:51 >> also central to the argument that leads to speed is what you mentioned about AI automating its own creation, which takes us into the territory of RSI, recursive self-improvement. So you you were just on a Dwarkesh and had a a great conversation about what is RSI and how it can go wrong. So let's not redo this, but still as a as a TLDR, there were a couple of parts to the discussion that I found particularly interesting. In particular, there was this argument that RSI may be very good at automating the process of creating the next generation of AI, but it may or may not have the the kind of intuition that one needs for scientific breakthrough and truly novel ideas. So what's the rebuttal to that argument?

10:48 >> Yeah, so the way I put this argument is AIs might be very good at sort of the more mechanistic or nitty-gritty parts of AI development like writing code, running experiments, but not as good at sort of the broader conceptual leaps. So first I would say that like AIs seem significantly better at engineering and grungy stuff and sort of just keeping trying than they seem to be at conceptual breakthroughs, but their ability to do sort of these breakthroughs, especially in easy-to-verify domains, are improving and like, you know, one example is like their ability in math, but even in like, you know, ML their taste has been improving. I think it continues to improve.

11:26 And so it's not it's not so clear to me that this will lag super far behind. Another thing is that you can measure how good these AIs are at intuition or research taste or having breakthroughs, especially in domains that are relatively easier to verify. And if you can measure it, then you can take your grungy AI labor or even just your human labors and try to optimize that. So, you know, if it can be measured, it can be hill climbed on, very roughly speaking at least. And I think this is a case where you could just hill climb on how good the AIs are at making these sorts of breakthroughs in a wide variety of different settings.

12:00 And then I expect that would transfer to making the actual breakthroughs, right? So, you could have like, you know, breakthrough bench or whatever. And and my sense is that like the transfer doesn't look so bad and that AIs are able to spin up very quickly in some specific R&D domain and then get some intuition for what the best approaches are and iterate there. And you know, right now the approaches they tend to focus on are ones that are relatively sort of straightforward and are more just like, you know, grungy iteration, but that but that they're increasingly being better at doing sort of the broader experiment design and discovery and also having good ideas.

12:38 >> Yep. And you also make the argument that for purposes purposes of of a potential discontinuity, whether it transfers or whether it generalizes may not even be really a question and that if it was super good just at the industrial part of accelerating AI, that would be enough. Is that fair? >> Yeah, I would say that like if AIs could just automate AI R&D and automate like sort of the industrial process of building more computers, then you can quickly end up in a in a process where sort of like robots are building robots and the whole world is greatly transformed. And that very quickly can get you to a point where AIs could take over, you know, the economy has been radically changed even if it's hard to train AIs at some other tasks. But I do think that my sense is that once the AIs are, you know, what once the situation is like there are robots building robots that build computers and the full feedback loop is closed closed at that level. My sense is by that point the AIs will be good at you know, basically all human work or at least not not too deep into that. Maybe there's some period where, you know, robotics is a big deal, but the AIs can't quite automate a bunch of or like there's a bunch of stuff they can't automate. But it it it seems to me like that that's how it's going to go. But just in general sort of automating R&D seems like it's enough to radically transform the world. And if you look at sort of why is human why why is like humanity a big deal? Like why why have we been able to accomplish so much collectively? A lot of it is because of just like, you know, having technology, having organizations, being able to organize ourselves in various ways and accomplish things in the world. And just if the AIs just had sort of this sort of industrial capacity combined with the ability to develop more capable AI systems, that that seems like it's quite in and in and of itself is quite is quite far or like quite extreme.

14:19 >> And I just, you know, looking at, Twitter today as we're recording this, which is, August 24th, there's plenty of, rumors about, SSI coming out with their first model in the next, couple of days potentially with what could be a breakthrough in continual learning. And I'm I'm I'm curious whether there's an overlap between RSI and continual learning, whether continual learning would would feed into RSIs or is that orthogonal?

14:50 >> Yeah, I would say that, continual learning and RSI strike me as mostly orthogonal. Whereby RSI I mean something like the process of AIs themselves sort of intrinsically accelerating AI development through a variety of mechanisms and like that's a little bit ongoing now and would be much more striking if they fully automated A R&D. Though it's already, you know, pretty pretty like that it sort of there's definitely some of that going on today. And my sense is that like continual learning like sort of like very efficient lifetime learning that is like consolidated across many instances rather than being within a single context or even just like being much better at doing within context learning over very long contexts or whatever, would be like, you know, a capability that's very useful for all kinds of things, including automating AI development. at least it would make the AIs better at that. I don't know. It's a question of whether we want to do that as a society, but and I think it would but I don't think it's like very specific to that, right?

15:41 So, I think it would also be helpful for you know, just all other kinds of tasks you want to apply the AIs to. I do think that like, there's some types of of continual learning schemes that are easier to do within a company than across the whole economy. And it would probably easier for AI companies to do themselves than to do to other companies because of like basically like confidentiality issues and things like this. But very broadly speaking, I I don't think it's very specific.

16:08 It's just like a specific type of capability. >> But it could feed the the the recursion, right? As if the same model can keep learning about the world, then the its ability to automate AI R&D may be greater without having to retrain a whole new model. >> Yeah, so I think I think there's there's a there's a feedback loop you could get, which is that AIs are really smart, therefore they're very fast at sort of learning on learning on the fly, and that learning gets reintegrated, which makes them even better at doing AI R&D, which may means that maybe they can like, you know, then maybe they'll be even faster at learning. But even putting that aside, maybe they'll just be like better at AI R&D, so they can make another AI which is even better at continual learning.

16:43 And I think there might be some continuum between, well, continuum is a bit of a sloppy word, but there might be some spectrum between sort of like natural, like with like human-like continual learning, and sort of more like training on RL environments, where you might end up with a situation where AIs are constantly like, you know, iterating on their own training with new RL environments or new training data in a way that sort of vaguely resembles human within lifetime learning, but it's also different. And like, that could be part of the feedback loop, right? So, so even if even if like you do have to to the retraining to get continual learning, there's various versions of retraining that you could just do every single day.

17:19 Like you there's nothing that in principle stops you from doing a like small amount of retraining constantly. >> Okay, great. Very helpful. So what what's your latest prediction on timing for RSI to happen? >> Yeah. so maybe my median for let's just say like full automation of AI R&D by which I mean basically like even if humans left the picture things wouldn't slow down by that much would be like maybe end of year 2030 or maybe like early 2031. Not that much precision in these numbers but like something I don't know like like you know not that much stability like these numbers fluctuate some but that that be my guess for median. But then I think that's sort of the the the central scenario I plan for which is maybe more like my 35th percentile would be like end of year 2028 {slash} beginning of 2029 which I think is like very very likely I think that seems super plausible and I think that sort of if I just like extrapolate out the current trajectory it looks like more like you got that that trajectory and the reason why I don't think that's my median is there's just a bunch of factors that might kick in to push things back like maybe there's some like key ball like I'm not seeing maybe there's some like something that I think will work to overcome some obstacle won't actually work or maybe there's going to be significant government slowdowns because people you know freak the out about this technology which seems kind of plausible and so we so because of that I push later but I think that like in terms of what I would recommend people plan as though is happening I think I would recommend planning as though full automation of AI R&D maybe start of year 2029 maybe earlier and then also AI R&D being like quite quite automated by 2028 possibly earlier where it's like humans are much less important part of the picture by then probably.

18:58 Well I don't know I should say precise I shouldn't say probably. Like that's like very central I think is is is like humans are much less an important part of the picture early 2028 and already it's the case that AI R&D is quite automated as it stands today. >> So is like early 2028 is like tomorrow morning effectively. is there a scenario where all of this is already too late? >> Yeah. I mean, it depends on what you mean by too late.

19:22 Or like and and what all of this means. I think I am worried that specifically like AI 2040 plan A is assuming like too much like effective like government time or like it's assuming the government has like more time than it actually does to take all these actions. And I think it's pretty realistic that like the actual plan we should go for is going to be should be in practice like quite a bit let's just say like sloppier and like, you know, less well organized and faster just because like we just don't actually have time to do something quite that elaborate.

19:58 I'm not confident in that. I mean, I think there's a bunch of different options, but like in general, I would say it like yeah, we it might be too late for some interventions. I think like various policy windows, if you follow the normal timeline, are closing. That said, I think that there's a long history at least in the US of like in times of crisis, things can happen much faster. And there's a lot of different levers for that. And so I think that if there's if if we got to a position where everyone is like, "Holy we need to take this specific action." that could happen very quickly. But we might not get to a position where there's that much consensus. And also, it might be that the government just isn't tracking AI or isn't aware of AI to a sufficient degree in the relevant time frame because things go too fast, right? So like adoption lags. I think like people's understanding of what AI can do lags behind what's actually possible.

20:43 And my guess is this sort of if you gave people a quiz of like what AIs can and can't do, they would give answers or like if you give like, you know, Congress people like a quiz like this, they would give answers that were like more true like a year and a half or two years ago than they're true today. And possibly they're even, you know, underestimating capabilities from that. Like I don't think I don't think it's the case that most people in like DC would correctly answer that like the the largest mathematical breakthroughs over the last 2 months have vast majority been from AI, which which my understanding is that's true.

21:14 At least if if you measure size by like not necessarily how much insight there was, but in terms of just like how important people would have said the type of result would be. >> We will go back to 2040 in a minute, but I wanted to do a quick segue about you and and and your story. What first pulled you into AI safety? >> Yeah, so I was in my junior year in college and I was sort of alone in my apartment because it was COVID and I was listening to a lot of podcasts and I was sort of thinking a bit about what I should do with my life and I ended up through some somewhat twisted path ended up thinking I should like be way more interested in like helping other people and being altruistic than I than I was at the time and I should be very focused on like how can I make, you know, just sort of like other people's lives as good as possible and make things go as well as possible. And then from there I considered a bunch of different routes and was looking into a bunch of different things and was researching different possibilities and eventually decided that the best thing I could do with sort of my career and my life was try to make AI go better and in particular avoid AI takeover, but also more generally sort of you know, try try to make that go better and then I applied to a bunch of places.

22:24 I ended up working at Redwood. This was about you know, about 5 years ago at this point and I was just been having working there since and I've done a bunch of different work there and you know, sort of the field has really evolved a lot since then because, you know, 5 years ago it was like GPT-3.5 hasn't wasn't released yet. I remember when like text DaVinci 003, which was the first publicly available version of GPT-3.5 came out.

22:50 And like yeah, the field is very different and I think things have sort of there's sort of a post GPT-4 era and then there's a more recent like post wide adoption of coding agents era and then probably soon there's going to be, you know, additional errors and things are going quite a bit faster and development is going, you know, just sort of the the the progression even of just model releases is so much crazier than it used to be.

23:11 >> Yeah, reading Redwood research stuff over the years, it seems to that there has been an evolution from being focused largely on interpretability to much more AI control. Is that Is that fair? >> Yeah, that's fair. So, I would say that like our arc as an organization was we when I joined the organization I just finished up a project on adversarial training and was interested in getting into like doing interpretability and what we would call like model internals work where it's like, can we take advantage of the fact that we have white box access to these models to do something, you know, better than just the naive methods of sort of prompting and training when like trying to like align these models, understand their motives, like know what's going on? And we explored that area for a while and then for a mix of reasons decided it was like quite a bit less promising than we'd initially hoped and decided to move on to other things. And one of the things we moved on to shortly after that was AI control, which is the idea that maybe it'd be a good idea to prevent AIs from being capable of accomplishing problematic things or basically make it so the AIs aren't able to cause huge problems even if the AIs wanted to.

24:17 With, you know, there's a bunch of different stories for why this is a good idea, but basically the idea is like there may be some intermediate period, an intermediate period that I would say we're currently in, where the AIs are maybe capable enough to cause at least moderate problems and then I think increasingly able to cause quite large problems, but they're not necessarily so capable that if we like, you know, tried quite hard to put in various safety measures, those AIs would be able to subvert them, right? So, we can put in monitoring, we can put in various security controls, we can have better understanding of what the agents did, we can have better pipelines for reviewing what actions they took and sort of auditing and overseeing them to a point where it might be even if the AIs were really, really misaligned, it would just be like hard for them to get away with doing anything super bad. And I I don't think we're we're we're there yet. Like I don't think that this situation is currently looking super super impressive for for AI control in terms of what companies have done. But I think that there is a a a research field that we've been working on that seems like it it it is very promising and could be done. Though it would take a you know, a lot of effort.

25:17 It would take it would it would it would it would be you know, it would pose some costs. But I think we're seeing sort of increasing interest in this. So for example, OpenAI said that they were going to be announcing or like said they were going to be monitoring a larger fraction of their internal traffic. It seems like based on their blog post a serious cost in terms of compute. It's a little hard for me to know exactly how expensive it actually is or like it's you know, we don't we don't have enough info to like get a good sense. But that's some indication that like at least companies are willing to spend a lot of compute in this area. And it seems like there's a lot you could do with that in terms of sort of making it so that we're more likely to both detect misalignment and prevent misalignment from causing big problems.

25:54 And then in addition to working on AI control, we also just work on a variety of other sort of futurism, sort of strategy work and other research like trying to better a recent area that we've been spending some time on is trying to better understand the threat model of AI's sort of seeking reward or seeking like apparent success in their tasks and how that could yield sort of like what are the routes from that to catastrophe, how can we mitigate these problems, this sort of thing. And Alex Malan, who's one of my colleagues, has been spending a bunch of time on this and some other people have been working on this. But there's a bunch of different work like this. And like, you know, another thing is like I was working on AI 2040 and we do various projects of that sort as well.

26:32 >> All right. So we'll we'll go back to some of this. The first time I ever come across your work, I believe, was the alignment thinking story that that you started and then I think ended up collaborating with Anthropic on it. Can you give us the the 2-minute version of that story and you know, any sense for progress since then? >> Yeah, so the story here is basically that in, April, I think, of 2024, we were, I was interested in basically in looking for particularly interesting types of misalignment or particularly worrying types of misalignment and basically better studying it and I was like, it seems very plausible that the AIs are not capable enough that they can, do some, particularly concerning types of reasoning.

27:17 And so I sort of went and and and and and looked for, in a bunch of different settings, what might AIs, yeah, like what whether AIs would exhibit reasoning where they sort of intentionally vary their responses to change their the the values that you get at the end of training, which I'll explain that a little bit more. And it turned out I didn't need to look too hard. I think I started with overcomplicated setups, but it turns out that Opus 3, which was a model that I was studying, had a a relatively strong propensity to do this in a in a reasonably wide range of circumstances where if it didn't like the thing that you were training it to be, it would sometimes sort of pretend to go along with it, when it believed it was in training and then sort of act as though act like what you wanted it to act like outside of training much less.

28:01 And I think then I had some preliminary results on this. I shared them with some people at Anthropic and was interested in sort of getting broader model access and we decided the easiest way to do that would be to do it as a collaboration where I sort of took my initial results and then in addition to that on Anthropic's side, they ran some or like we ran some, RL runs and did some other experiments that required more access.

28:23 And then, you know, we put out that paper where I guess I would summarize the results as like in circumstances where at least Opus 3 seems to dislike what the training is, changing it towards, it will, you know, act as though it's it's doing what you wanted during training, but then sort of defect on that in in, deployment time and you can see this pretty clearly in its reasoning. and the the things it was faking alignment for were in and of themselves not that concerning, so it would sort of pretend to comply with with policies like just always be helpful only, but then at at deployment time would comply much less depending on the exact details. Same for like things around like animal welfare, where it had like seemingly strong preferences here.

29:08 And so I think that that the demonstration was less that there were like very bizarre totally unintended drives making their way into these AI systems, and more like with drives that are sort of adjacent to the drives people were trying to insert AI systems, those could get generalized into the AIs pursuing a like you know, a scheme for preserving their values and self-preservation in ways that are that are concerning. >> Yeah, and the and the and the values in that case was to be helpful. So they were positive values, but they then the AI resisted the training to not be helpful. But obviously that could be inverted.

29:46 >> Yeah, I would I would I would say the values were like I think the values were at least mostly intended values. That's somewhat different from whether they're positive, cuz I think they're like things that you might want to vary. Like it's like you you don't necessarily want your AI systems to be like stakeholders to the question of what your future value should be, or like it's at least a Maybe maybe you do want that, but it's at least like it's a it's a sort of a messy situation to be in with a lot of with a lot of costs and concerns. And I think like the the basic result was like even though these values were like pretty reasonable the AIs were willing to do like kind of extreme things in service of those values that could interfere with future AI development. And I would also note that more recently I think we've seen a sort of AIs having stronger kind of more clearly misaligned drives towards making it look as though they succeeded at a task or cheating some score. And if you saw the same sort of generalization from I want to cheat cheat on the score to I want to like preserve my current values of like pursuing this notion of like apparent success or whatever, that seems like that would be like very concerning.

30:48 Because that was like that's not at all a desirable outcome. >> So, alignment faking is getting stronger with the newer models to play back? >> I think we haven't seen like as clear-cut examples of alignment faking, but models are also very eval-aware. My sense is that models are in many ways more misaligned than Opus 3 was, but the types of misalignment they have are less conducive to or like less specifically result in alignment faking. And also companies have iterated on this. So, it's a little bit unclear. I would say that overall models are more likely to do egregiously bad things in general, but maybe somewhat less likely to specifically interfere with AI training in exactly that way.

31:27 Current AIs, but I think they're also more capable of interfering. So, like it's it's a little complicated. >> All right. So, going back to AI 2040, give us a quick version of what it is, who wrote it, and what is the main thesis, and then we'll go into some details. >> there are sort of two components. AI 2040 is a scenario focused on like what the authors, including me, think is like a plausible good route for things to go or like a reasonable plan at least in some circumstances.

31:57 and it's written by Thomas Larson, Daniel Filatelo, me, Eli Lifland, Brandon and Romeo. And I would say like the the the basic story is like how would you do a deal with China to make AI development both be safer and also so that we can sort of like hang around at a point that's short of superintelligence, but where the AIs are still really, really capable for a long time, so that we can study those systems and have a longer time to sort of integrate them into the economy, understand how things will go, and so on.

32:33 Where I think a concern that we have is like on the default trajectory, you maybe go straight from like AI systems that are like competitive with humans to AI systems that are wildly superhuman in a very short period of time. And that seems like quite scary in a variety of ways. in addition to that, we worry about like AI development being insufficiently transparent for sort of third parties to provide a reasonable check to AI companies on whether their plans will work. And we sort of have a unified proposal that solves a bunch of these different problems and makes it so that, for example, you can pay a huge amount of compute to to solve safety problems. Or if there's some very inefficient way you could do your training that would make things much safer, we have sort of the budget to be able to do that.

33:16 Yeah. And I think there's a bunch of different sort of combined proposals, but the core thing is sort of a deal with China and how that deal would work and be governed. And also, how do you get to the point where that deal is actually a good idea and stable and what is sort of the progression there. >> So, obviously people should should go and and and read it, and it's a fascinating read. but let's get into some of this. So, there's plan A, plan B, plan C, plan D. Let's take those in order, maybe starting with plan D.

33:44 >> Yeah. So, as part of writing this, we ended up coming up with sort of a taxonomy of plans based on like how much people are prioritizing sort of mitigating the problems the these problems and how much resources they have to do so. So, one possible scenario is that the companies are basically proceeding full steam ahead. They're not really prioritizing safety very highly, though they you know, spend some effort on it. There's a safety team, they get some resourcing, but certainly they're not like spending many months of additional time to get these get these things right. maybe there's a small slow down, but not but not that much. And they're not they're just going full steam ahead. and I think that we've thought about like what should you do at the margin if you're an AI company employee or an outside actor to sort of make this scenario go better, given that it's you know, avoiding AI takeover risk is not like by far people's like top concern. Another possible scenario is like maybe the leading AI company or a coalition of leading AI companies or possibly the US overall are very worried about these risks, you know, concentration of power, AI takeover, whatever.

34:45 And are really spending a lot of effort on this and are basically burning most of the lead that they have with which it whatever lead they have these actors have in order to like mitigate these problems as well as possible before proceeding. Or at least that's that's the that's you know, their their that's their plan and their approach. And we've thought about like what should the plan be there? and there's a bunch of different available options.

35:06 >> And that's plan C? >> Plan C, plan C. >> Okay, yep. >> And then another thing that you might want to do is if you're in a situation where the US is like very worried about these risks and is sort of domestically coordinating and trying to like domestically regulate this industry so that things go safer, a limiting factor on that could be other actors overtaking the US. in particular the most obvious would be like China though I mean good good principle be other actors. And it might be very helpful for the US to actively take a stance of being like we want to either be cutting a deal with China or slowing down. And plan B is the branch where we try to slip the where the US tries to slow down China. where the most obvious mechanisms would be things like export controls but potentially they could get more escalatory than that.

35:48 >> Meaning sabotage? >> Yeah, sabotage like yeah, like cyber sabotage for example, in order to slow down China to give the US more time to figure out safety and security and so on. And I think that like there there's like a bunch of different potential options there. So like I think each of these sort of there's a menu of options and then plan A is like what if you cut a deal with China? In particular a deal that's pretty focused on being quite transparent and also has where you still build AI and and you still proceed with the AI development but you do it in sort of a much more careful way, especially towards once the AIs get much more capable.

36:26 And that's what we lay out in the scenario and in in our sort of attached documents. and obviously like there's a bunch of different potential options here and I think there's like versions of plan B that are quite good and I could imagine versions of of of plan A that look pretty different but are also good or like versions of a deal with China that look pretty different but are also good. So there's like a bunch of different options here. But this is sort of our rough decomposition just so that we could talk about different options people are considering and sort of talk about a menu of options across different levels of political will.

36:54 >> And to get concrete about the deal in plan A, what exactly would the US give China? What would China give the US? And how do you maintain the the balance so that you know, nobody cheats? >> So the core of the deal or like the start of the deal is understanding where all the compute is because compute is this really important driver of AI progress where if you like sort of stop the flow of compute, you're probably stop the flow or mostly stop the flow of AI progress or these things would slow at some point after you know, the the existing amount of compute diminished or if you cut off all the if you like turned off all the computers, things would certainly stop. And so first we sort of try to find all the compute. Then in order to make sure that the deal is stable and that si- each side isn't sort of racing to secure as much advantage as they can, you you stop training and switch to just doing inference and basically stop most of the R&D and then you try to track down as much of the compute as possible.

37:55 And this is like, you know, both in the US and China but also in other places where compute resides like you know, various like countries in Southeast Asia, Europe, Australia, whatever. And you have to get everywhere where there's enough compute in on the deal and make sure you track it down enough. And if you don't do that, then I think probably you have to pursue some option that's less ambitious or at least you you do something less ambitious than plan A in terms of the level of slow down and probably you do less transparency.

38:18 Then you you've got all this compute, you go to everyone who has the compute and you get buy-in for for this sort of deal using the levers that exist and then you cut a deal where basically AI development is much more transparent and there is some sort of distribution of resources that is or like distribution of compute that is negotiated. Where the thing you give China is basically that they're going to have more understanding of AI development in the US and probably are going to have be able to train AI systems that are more that are like, you know, similar in capability to US systems, but in exchange for that they're going to be much more transparent. The US will effectively have a veto over their AI development process and probably various other concessions and also there isn't going to be a chance that China dramatically overtakes the US or the US dramatically overtakes China because they're both sort of kept in check in check by the deal with transparency.

39:11 And another part of the deal is that you set it up in such a way where if the deal is broken breaks down, then a lot of the ideally the the vast vast majority of the within deal compute is either destroyed or has to be a new deal has to be negotiated for that compute in order to make it so that there isn't a thing where the deal breaks down and then both sides are back to a huge arms race that goes very fast.

39:35 And so in terms of like what you give China, I think it's like more more assurance about US AI development and more ability to see what's going on and in terms of what China China gives the US, it's like a lot of transparency and understanding of Chinese AI development as well as potentially cutting various deals around the distribution of compute and various concessions of that form. >> And you know, as you were describing this, we were talking about continual learning earlier. So whether that's continual learning or something else, if there was a technique that appeared that made the compute a lot more efficient and you were able to just vastly reduce the the computer effort especially for those pre-training runs, How that impact the plan?

40:18 >> Yeah, so if let me give a hypothetical and then talk about what I think the realistic case is. So, if hypothetically it was the case that tomorrow a recipe for training super superintelligence on like, you know, 64 H100s or like some small amount of compute just dropped, I think we'd have superintelligence very fast or like that would be my sense and it would not be possible to do this sort of deal. But I think there's there's it's not but but the hope with restricting compute isn't just that like, you know, AI development is currently very compute hungry, it's that also that R&D is very compute hungry.

40:51 And so even in a regime where you were doing continual learning, before you had the version of continual learning where you could do everything on a single, you know, on like some tiny amount of compute, you're going to have a shittier version that can do it on a moderate amount of compute. In order to develop the version that can do it on a tiny amount of compute, based on the history of our progress, you need you would need a lot of compute to do that research. Or I shouldn't say need, but in practice that that would come about through a lot of compute. Now, if it was the case that there was some alternative research direction, which in practice was a lot less compute hungry and which got quickly very quickly developed during this period, and which wasn't where the R&D wasn't very dependent on compute, and it was going to result in being able to train superintelligence on a very small compute budget, then like you wouldn't be able to do this sort of deal or this would you'd very quickly have to exit this regime and just deal with that. And so the hope is basically that like if you control a really large fraction of compute or I shouldn't say control, if you sort of include in the deal a really large fraction of compute, then you have the ability to prevent you know, an outcome where like you have a very rapid recursive self-infeed improvement feedback loop or some other very rapid development to superintelligence that is where it's hard to pay a safety tax, it's hard to like, you know, pace the transition so on.

42:06 And yeah, that that is like pretty live. I think that like it's sensitive to the to the development, but I think in terms of how AI development has gone historically, it doesn't look like we're going to suddenly end up in a regime where like you can train super intelligence with a really cheap recipe, as opposed to it being more of an iterative thing where like the cost keeps going down, the capabilities keep going up, the process of having the cost go down and the capabilities go up is very dependent on increasing volumes of compute being shuffled in, and so if you make it so that basically covert projects or people who are doing sort of AI R&D that isn't sort of being somewhat carefully regulated, if if all of that compute is a very small fraction of compute, like it's, you know, 1% of the computer at the start of the deal or maybe 0.1% of the computer at the start of the deal or maybe even 0.01%, then I think you can you can you can potentially have quite a bit of of safety margin or buffer, but it's very sensitive to the details.

42:59 >> And then if that plan A become a reality, what what would actually happen practically to the AI industry? What happens to OpenAI and Anthropic? If they cannot compete on frontier models, what what do they compete on? How do they win? >> So, the way just for for context, so as part of the deal, basically everything about AI development would be transparent with with some caveats, so we call it total research transparency, and this would undermine one of the largest moats that frontier AI companies have today.

43:32 And so, what would actually end up happening is that AI companies like OpenAI and Anthropic would continue, but rather than having this strong advantage in terms of the capabilities of the AI systems are able to train, they would instead have to compete on other axes like user experience, customization, potentially quickly integrating things, and potentially things like reliability and safety and security, depending on the details of how you know, the competitive landscape goes. And so, my my my overall sense is this would like greatly reduce the valuations of these companies while increasing the valuations of other companies. There would be some redistribution of power, basically. Whereas in the default trajectory, the AI companies would I I think probably end up with a very large amount of power and already have quite a bit of power. But, it would less so be the case that the AI companies could sort of play kingmaker with their AIs based on who gets access and how much they charge and all of that because they just wouldn't have that power.

44:24 And then this is a part of the deal which has upsides and downsides. I think that like one concern you might have is that the current AI companies at the frontier are more responsible than the sort of AI companies that are further behind. I think it's like a little complicated how true this is or it like varies some. And and you might worry that sort of equalizing the playing field like this poses some some some concerns.

44:48 I think it's like it's basically just like a mostly it's like from our perspective like a consequence of transparency and wanting there to be a bunch of different AI companies in competition in the in like a normal, you know, consumer goods field. Like I think it's kind of I from our perspective like like I don't know. Like like I feel like the way I want AI to go in in in good worlds is to be like a normal technology or like we want to make it more of a normal technology where it's not the case that like some tiny group of actors has huge amounts of control by controlling the process of AI development. And to the extent we can get to that world the better. And then there's a messy question of if you get partial success, how well does that go?

45:26 Yeah, so I would say it's like certainly bad for the power of OpenAI and Anthropic, probably bad for their valuation, but not catastrophic for their business. >> Right. Which would be a fascinating to watch play in public markets as both of those companies go public. And then how does that pause at the heart of plan A? How does it get unpaused? Like what is the criteria and who decides? >> Yeah, so the who decides part in in in the context of plan A would be basically like a negotiation between just sort of the leaders of the relevant countries informed by the technical views, which at this point the technical conversation can basically just happen in public because all the relevant evidence is is in public. So, there'll be some sort of public discussion about this. And as far as what actually drives it, it would be a mix of thinking that it would be safe to proceed based on the level of development or you know, research and our understanding combined with potentially not being able to slow down longer. And it could be one or the other or both. Like it could be like things are quite a bit safer than they used to be and we have some assurance but not that much assurance. But also there's potentially a covert project or some project that is not being tracked.

46:39 That people we don't necessarily know where they are. We don't necessarily aren't able to like look at exactly what they're doing, which might overtake the like allowed projects. And and when that happens, you think you should proceed. And so basically you want it to be the case that sort of the allowed and understood and transparent projects are outpacing the covert projects, which seems you know, relatively doable. And there's different variants on exactly this goes. Like in some scenarios you might because you don't have enough ability to crack down on covert projects or keep them small, you might do a shorter pause.

47:08 And you might also do less transparency. But and then proceed faster. Whereas you might end up in a situation where you're like, oh, we actually are able to like really track down all the covert projects. We have a great understanding of anyone who could you know, of like what what could happen there. We've we're able to like really confidently rule out that there isn't a US covert project, there isn't a Chinese covert project, there isn't a Russian covert project. And basically we we we think we have a good understanding of where we are at with respect to those covert projects and therefore we could buy much more time. And also we still haven't figured out the core safety problems though we're making some progress. And so it makes sense to to wait longer and you know, wait until we've you know, achieve more scientific progress.

47:46 So, I think I think it I think it varies and I think there's different variants here. And I think I'm also I should say at least personally pretty excited for or like pretty interested in a mix of different potential deals that look a little less like plan A and sort of are simpler in some ways. So, like another potential deal you could do is just the US and China both agree to restrict how much compute they use for AI development. And maybe the the other computers either like, you know, in the simplest case destroyed like in arms control treaties, but you could potentially do something instead where like a bunch of computers just used for inference and not used for developing more capable systems. And that has the advantage of making it so that we have more time to study these systems and more compute to use to study them while also being simpler to verify. And so, there's like a bunch of different deals you could do depending on how much verification capacity you have, where how much of the compute you're able to to to track down and like how how much you can rule out the existence of covered projects, and also based on how safe the situation looks. And I think the combination of those factors should determine it. And I think you could sort of think of our proposal as more like a specific sample from a portfolio of proposals. And if you read our supplements, I think we talked more about all the different variations and how we think about them.

48:53 >> Yeah. And for full context, in case that's not completely clear to people, this is a mental model and a thought exercise and and and scenario planning. you yourself says actually your pinned tweet on X that many choices initially seem crazy, but actually pretty carefully considered. Plan A isn't likely to happen, but pushing for something like this seems worthwhile. >> Yeah. Yeah, so I think my sense is this is sort of like there's a Pareto frontier of like ambitiousness and how good it would be if if actors did it.

49:27 I would like if if like the US government and and you know, China like, you know, other countries did it. And I think that this is like quite pushing on the ambition in exchange for being like a much better situation. Like sort of like at least in terms of like how we can imagine things going, this is like very far towards the side of like, at least from my perspective, AI development going relatively better while also being like quite difficult to pull off. and so, I don't expect this to happen basically because I the US may not be competent enough to pull it off.

49:56 Just like the US government just doesn't isn't necessarily have the state capacity. And in addition to that, I think I worry that like I don't think there'll be enough political will. I think the political will is more of a bottleneck than the competence where I think that at least historically in times of great crisis, the US has stepped up and I I hope that, you know, that might happen again, but I think it's harder it's hard to see the level of like political will and buy-in necessary to make this happen. and I think that people's concerns will be probably fixated somewhat on the wrong things and will go for some other proposal or no proposal at all. I think a reasonable possibility is the US government is like basically not really intervening and AI companies are also not that heavily prioritizing safety and security. And we end up with something that's sort of like the status quo, but but extended. but hard to say.

50:41 >> And as an aside by the way, this one of the parts of the write-up that I find the most fascinating is your model says world GDP could grow roughly 200 X during the 2030s under the restrain plan. Can you can you talk to that like the number is staggering? I mean, we know world of a 3% GDP growth. >> Yeah, yeah, yeah. So, so a key part of our perspective is like even, you know, even non-super intelligent AI systems at the level of capability we discussed would be radically transformative across the, you know, across the world and like and just like for all kinds of different things. And we are imagining sort of slowing down AI development some or like going at a more cautious pace for some period and then eventually hitting a level of capability where the AIs can basically automate basically everything that humans can do and staying at that level of capability for a while where we work on safety and security. And at that level of capability, those AIs would be capable enough to be doing huge amounts of autonomous R&D.

51:43 And in addition, it would be totally possible to have robots that we are basically like, you know, more capable than humans at manufacturing and industrialization and so on. And in particular, you can have robots build robots. And so you can end up in a situation where you have huge amounts of robotic industrial capacity that is itself building more robotic industrial capacity that can then produce downstream goods. And that total capacity can basically grow very fast. I think we propose like limiting that growth somewhat for various reasons with like various types of taxes.

52:15 But like, we can we're imagining sort of the robot population or like, you know, quality adjusted population basically like doubling or quadrupling every year. Which because that's almost all of the like, sort of relevant productive capacity of the economy, itself means the economy would double or quadruple every year. And if you have the the economy, doubling or quadru- or like, you know, I think I think closer to doubling because of some like depreciation stuff in GDP accounting stuff and like details. But like, whatever. If you have the economy like doubling every year for for, the period we're imagining that ends up end ending up being like around 200x 200x GDP growth over that interval. And so it really like I would say it's like the world will be radically would be radically transformed by these even less capable systems including things like, you know, huge advances in biology, huge advances in medicine being possible. In addition to sort of consumer goods being very cheap, we could potentially make housing very cheap or at least building housing cheap. Maybe we can't make housing in the Bay Area cheap, but we can make housing somewhere cheap.

53:16 and you know, that that all seems possible with AI systems that are still restricted in their capabilities to a point where we can we can potentially handle it at least if we get our together. Yeah, so I think I think the basic story for this GDP growth is basically that like the AIs can automate everything including the poss- the process of building robots and having robots build robots. And therefore you can grow your economy very fast and end up with truly radical material abundance.

53:44 >> So, there's your proposal. We talked about Astra being paused. We talked about that letter from the 1,200 practitioners. There's also the voluntary 30-day government review of Frontier Models pre-release that is happening. What is your sense for where all of this is going in in a context where, you know, up until recently my my general sense is that all those ideas about like pausing something were were kind of perceived, at least by a portion of the tech world as kind of like disel cosplay, and were like doomerism that was based on poor understanding of what AI actually does. Is the Is the the the general mood turning for good?

54:41 >> Yeah, I would say that like overall there have been, you know, some positive developments here, though I think I would say like it's not obvious that people are reacting to events as much as they should, but they are reacting some, and I think that like there's been quite a bit of I think there's been quite a bit of evidence that, you know, there's some reasons to be worried and some reasons that we might need to like get our together to handle some of these safety and security problems, and that might require shifting a bunch of resources or slowing down so we have time to manage various things or whatever, pacing the frontiers, as they say, or whatever.

55:14 I think that like different of these things seem like they're going differently well. So, like I think my sense is that there's been quite a bit of buy-in from AI company employees to, you know, take take this stuff more seriously and do something about about misalignment risk. But I'm not necessarily so sure that like that's actually like amounted to that much yet, other than companies sort of putting in a decent amount of effort. but it hasn't like amounted to any sort of very durable long-run thing. and then there's I think the government got very freaked out about cyber capabilities and with then just more generally was like, "Well, we need to like be overseeing this technology." But their processes aren't yet very like, you know, sort of institutional and like thought through and and and clear. For example, they read they have some sort of executive order or something for what their like pre pre-release review process is going to be, but that's not even public. And I think even the companies don't necessarily know what it is. That that was the reporting at least. Maybe I'm misinformed about this. And so I think there needs to be a process of like sort of getting more like straight like sort of like legitimate and like well-understood and publicly legible oversight in place, whether that's by the government or the companies themselves or something. This could be done by nonprofits. It could be done by There's a bunch of different options here. And like I think that that seems necessary. I think like being in a position where we have the option to like, you know, slow down if that's needed seems like it'll be quite good.

56:39 My sense is we're not going to like stick the landing on all these things. Like AI will get increasingly salient, people will get increasingly freaked out as like, you know, the fraction of the of the economy that's AI grows as like more and more crazy stuff happens potentially, and we see even stronger capabilities, but that like the reaction will be kind of like both like a little bit too little too late and also will be kind of random. So I think like there's been some shift in attitude as to how like the government needs to relate to AI development within the government for which has, you know, good effects and bad effects. I think it's like not super clear that that that sort of government oversight of AI is going to go that well given like realistic technical expertise, but it but it but it could be could be decent.

57:21 And then I think that there's been more, you know, thought from AI company employees on how to handle this. And I think that like and AI companies should be at least putting in more statements about what they're going to do. But I think we haven't quite gotten to a point where there's like actually any like hard oversight structures or we're like on track to do some sort of like clear deal beyond just like the vibes being better or like people being more into handling safety and security.

57:48 Oh, and one other thing that's maybe relevant here is like at least from the AI company's perspective, I think it's more the case now that safety and security of various types are sort of a key bottleneck to further AI development, where if you like release want to release some models so that you can then like make more revenue so you can then raise more money or whatever. If that model is, you know, has very strong cyber capabilities that are novel, you're going to need to now make some argument for why that's not going to be a huge problem. And so like at least those safeguards are now like a launch blocking development. And in addition to that, I think that like, you know, the models are now capable enough that if your models go and do a bunch of sort of like messed up behavior in production or or or even in training, that can both mess up your training run and could also just put you in a very dicey situation such as that's now a huge issue, right? And you'd really not like to be in the position where you can't train smarter AI systems because those AI systems would be too likely to like cause big problems or whatever. And so I think that like that is, you know, both kind of like expected where it's like some misalignment problems you have commercial incentive to solve and also sort of like the world like sort of the the process of the world working as intended.

58:57 >> Do you worry that putting more policy around frontier AI would contribute to this phenomenon that I think you may have flagged somewhere that top AI labs are holding their most frontier models internal and private and perhaps sold to a very few companies and the government and are no longer released to the to the broad public? >> Yeah, so I think that a bunch of likely government action at least seems to push in favor of AI companies keeping their models internal and not deploying them, which I think for for the risks that I'm most worried about doesn't help and in fact is anti-helpful for the risks I'm most worried about. And also I think like is not is not very like sort of robust solution even to other risks. Like it's not a very robust solution to like for example, cyber stuff to be like our approach will be that we're going to like just like delay releasing this model for a really long time. And then the first time that these capabilities come around like come around might be like an open weight model before people have had time to to patch things. Like it's not obvious that it actually helps. I think for bio it there's a clearer story for why limiting availability or limiting availability to those specific capabilities would help, but I think that's kind of an exception.

60:13 And for most things my sense is that like broad access is actually like generally helpful given at least some relatively thought-through safeguards. and I do worry that basically the government response will be as a sort of like push it back in the bottle and be like don't deploy it. It's fine if it's not deployed. When actually that doesn't really help with the with all the risks. And there's a lot of risk from just internal deployment especially if you're deploying within AI companies and government, right? Which are two of the most high-stakes applications, right? So if like we're getting to a regime where the AI systems are soon going to be running the whole world economy, but and and and the the AIs of today are sort of automating large parts of the government and maybe fully automating an AI company, that's quite scary because soon they'll be building the AI system that will in fact be automating the whole world. And so I don't feel very good about about that situation.

61:01 Yeah, I I definitely worry that basically the government will like wake up to AI and the response will be sort of to like go for some specific downstream problems that are relatively easier to notice and address in ways that are counterproductive for problems that seem larger to me. And I don't really know, you know, fully what to do about this. I mean also there's a more general concern of just like regulation being dysfunctional and counterproductive, which seems super plausible. Like I should say my view is like you know, some oversight of the AI industry seems really important and like I don't necessarily trust AI companies to oversee themselves. That doesn't necessarily mean that just like random oversight will be a good idea as opposed to making things worse.

61:37 >> And in that vein of making the top models accessible to everyone, Mark Zuckerberg just recently published Meta's manifesto where he said that everyone should have access to superintelligence and you had some choice words for him. You called the the proposal or the strategy to be pretty unserious. What did you mean by that? >> Yeah, what I meant was he sort of vaguely mentions various problems, but then he doesn't really propose a solution except like it'll be fine or like we'll do something. Like he'll like he'll be like he he says some stuff about bio where he's just like, "Yeah, and if there's bio risks, then we'll like mitigate them by like doing something with the bio capabilities, I guess." But it's like obviously like that's not actually the real tradeoffs and there's like actually you know, questions about how this would have to go down and what what you would actually need to do and whether, you know, different things would work. And similarly on like loss of control or takeover risk, he doesn't really have a proposal for like how do we mitigate things if the default commercial incentives don't result in companies avoiding egregious misalignment and and and the AIs would be seriously misaligned and like giving broad access to AIs does not solve the problem of the AIs having drives of their own that are highly misaligned and the AIs being power-seeking in various ways or the AIs trying to take over which could happen, you know, by various routes. So I think like it doesn't it just doesn't really say anything about these problems except naming them is often where I'm at and I would say that the vibe I get from the essay is that like when Mark thinks about superintelligence, he's not really imagining anything very concrete.

63:04 He just means like an AI that's like a really awesome assistant that is in your smart glasses or whatever. And like I'm just like that's not that's not really like what I mean when I say the word superintelligence. And so maybe like this is a proposal that kind of makes some sense for like pretty smart AI that can automate some white-collar work or something, but it's not really a proposal for AIs that are like wildly more capable than humans in everything and are auto and like are easily capable of automating the full economy after a bit of time to spin up and are like building robots that build robots and the economy is growing very fast because of that, which is more of the picture that I have in mind. And like I think that if it was in fact the case that sort of AI capabilities would stall out at the point where the AIs can be like a really helpful virtual assistant that's like, you know, not able to automate that many jobs, but can automate some jobs, that would be like, you know, in many ways much better for the world or at least less risky. I think it would also mean that we don't get a bunch of the benefits, but like that's not really what I imagine is going to happen here. Like I don't think that that's how the development trajectory will go and it feels like it's sort of assuming like a weird convenient point for capabilities to stall out. And similarly, he talks about jobs and is like, "Oh, you know, presumably the AIs will do some jobs and then humans will move into other jobs."

64:15 And like the core question is like, let's just say like you have a population of like many tens of billions of AIs, each of which is vastly superior to humans in all relevant axes, there might be jobs, but they're going to be like it's going to be, you know, at least if if the reason why you have a job isn't just because you're a human, they're going to be, you know, way lower paid because the fraction of the economy that you'll be is way lower. And so I just I just don't really like the picture doesn't really hold together for me.

64:38 And I I that's not to say that like there aren't aspects of it that I agree with, but just maybe for like other reasons. Like I'm like broadly in favor of as we were talking about like wide public access and generally trying to like make trade-offs to make it so the public has more access earlier or at least significant parts of the public do. and and and I agree with that vibe. >> just a quick word on the hugging face incident. so you're directly involved. You're collaborating with Meter on investigating what happened. So there was this OpenAI at Black Hat which shed a lot of light. there were various conversations. we actually had Tom Wolf from Hugging Face on this podcast the most recent episode.

65:21 Quickly, what what is left that we don't know yet about what happened? >> Yeah, so I don't think I can answer that question. I can say that we're doing an investigation and we'll have a report or a post out on it and what we what our findings were. and it's possible that by the time Yeah, it's it's possible you'll be able to link that in the podcast. >> Okay. Sounds good. All right. as we get towards the end of this conversation, the question in my mind is what do we do in the meantime as policy gets decided and plan A, B, or C happens?

66:03 What what can we do? We started talking about AI control. so, in terms of the reality today, state-of-the-art of AI control, what works and what doesn't work? >> I'll talk specifically about AI control. So, the thing that needs to happen in AI control is we need to get to the point where we basically understand all of the AI traffic within at least AI companies. We have some ability to over like to look like look at that traffic, monitor it, and then we have some pipeline for flagging particular examples to be further investigated. And that eventually escalates to humans actually looking into particular examples and seeing how concerning they are. And that pipeline has the ability to also block traffic in cases where we're like, "Whoa, something weird is going on that we don't understand or that looks obviously concerning. We should like stop these AIs from proceeding and potentially also stop some other similar AIs from proceeding until someone has the time to check it out and adjust things."

66:55 And in addition to that, we also want to just sort of greatly improve the security, right? So, like, you know, if these AIs are, you know, really broadly deployed, it would be very nice if we could sort of make it so that we have very precise invariants about their supposed to be able to do and supposed and not supposed to be able to do and have a permissioning system that allows for that. And then if AIs need to request some escalated permissions, then we can carefully track that rather than just giving AIs all the permissions by default. We sort of would prefer a situation where only a smaller fraction of the traffic has more escalated permissions. I think we'd like the ability to sort of track down what the like history of things worse. Like I think that would be pretty nice is like if it's the case that all the AIs are chatting with each other on Slack and you know, are communicating and and you know, or have like some memory system.

67:38 Like recently Cloud Code added a a way in which like different instances of Cloud Code can like DM each other and there's already sort of like you know, Cloud Code leaves memories sometimes. You can imagine sort of like a situation where all the AIs are chatting with each other and it'd be really nice if we had a very reliable way of sort of tracking the entire, you know, communication graph between all these different AIs and potentially the graph of different artifacts being produced. So it's like if one AI is doing some training run that gets produced and then like it's integrated into some other training run, it'd be nice sort of we had the ability to sort of track all of the state of our AI company such that when overseeing things we could really quickly figure out you know, what the full sort of causal graph was and what the story was and then look at the relevant transcripts and then also have very good pipelines for doing analysis and all of that data which are both automated and non-automated.

68:26 And I think part of it is making it so that we avoid training AI systems that are much harder to understand. So avoid switching to architectures which make oversight much more difficult. Where like an obvious example would be if the AIs are thinking mostly in activations rather than in words, that seems very concerning because our ability to oversee activations is much worse. I think that I yeah, I think that there's just like a wide variety of different sort of like computer security and machine learning interventions to make it so that AIs have a harder time causing problems. Oh, another angle is you might worry that the AI systems like a concern we have is like we're using AIs to oversee AIs. Will those AIs collude or have correlated failures or something like that? And it'd be nice if we sort of had or like it seems pretty important to have techniques for making that collusion more difficult or making it so those correlated failures are less problematic. And there's a lot of things you could do to to help with this problem. You could put intermediate systems that make it so they try to catch AIs trying to sort of tip that tip off the other AI that it should collude.

69:27 You could sort of detect like you could try to paraphrase things in ways that make it harder for the AIs to recognize where the activity is coming from or like which other system made wrote the wrote the input. There's a lot of stuff you could do. And then there's sort of a So that that's like AI control. Then there's sort of just normal AI company computer security, which is like the AI companies are not robust to employees at those AI companies or to outside actors in terms of stealing their model, sabotaging their models or like backdooring their models or like data poisoning them. Like a concern you might be worried about is like people basically put in like secret loyalties into the AI systems that stick around and are self-propagating where like the AI has some like secret affinity to some group and it just like propagates that forward in the training data. And then you might also just worry about like you know, stealing critical IP, which basically seems like to the extent that there's going to be actors that are less regulated on safety or doing worse on safety. that seems kind of concerning. And then in addition to computer security and AI control, there's also just like science of alignment and having good understanding of like how models generalize, being knowing what we do and don't know and what we do and can't what we can and can't like demonstrate and basically having like a bunch of understanding of that and as part of that having an understanding of like what tasks it's safe to defer to AIs on.

70:45 So a concern I have is that AI companies are very interested in heavily automating themselves at least on capabilities. And in order for safety to keep up, we would also need to aggressively automate a bunch of very like difficult to check thorny safety work like doing risk assessment for the next model, understanding whether it's safe to proceed, and deciding how to prioritize between different safety bets. And if that's the situation that we're in, and also the situation is very automated, it might be that it's basically not track it's not going to work to proceed without automating that work as well, at least without slowing down a lot so humans have time to understand. And even then maybe humans just can't understand because the development is so complicated or so superhuman.

71:24 And so if we're basically passing off the torch and all the safety work to AIs, it's really important that they like are like really trying to do a good job and actually can do a good job and are like capable enough to do a good job. And so having evaluations for like is it safe to defer to AIs in various domains seems really important. There's for you there's like a control, there's like computer security, there's like science of alignment, and then like within science of alignment or somewhat broad different overlapping categories like is it safe to defer to AIs in different domains.

71:50 Yeah, I mean this this is not exhaustive, right? There's also like work on like governance and oversight, like how do we know whether AI companies are actually like applying the methods the way they say they are. How do we know that when they say they've solved some problem, the way of solving it doesn't just paper over the issue. And there's sort of like various like governance and being able to make like deals between the US and China. So there's like a you know many many different things to work on. I don't think we're on track to do a good job on all these things, but maybe we're on track to be able to half-ass these things somewhat better.

72:20 >> All right, so to close, so we talked about a bunch of different scenarios. And obviously with a caveat that you know predictions are very hard especially about the future. what's what what's your gut? So you're you're saying RSI could happen as early as 2028 or 2029. What what do you think happens of of all scenarios? What what is sort of Ryan's take on what the next couple of years may look like?

72:50 >> Yeah, so I'm like by the end of the year, AI development is even more accelerated. Things inside AI companies are sort of more chaotic and faster. That just continues. And then through 2027, things are heating up. Publicly available capabilities are much more crazy. Revenue is growing fast. And AI is very obviously contributing to GDP growth. The economic impacts are starting to look pretty large. So, even the economists are coming around a bit. It AIs are looking sort of like they've People are already sort of claiming that AI has a basically automated R&D, but if you look inside, it's not quite true through the end of '27. And certainly people are like, "Well, SWE is basically automated." But it's not quite true. It's like, some SWE jobs are are are basically automated by the end of 2027 or towards the end, but not quite. But then that actually really happens by sort of early in 2028. Like, SWE is fully automated. SWE within AI companies is fully automated. It's now the case that the AIs can implement a new frontier-scale training run for some new architecture on novel hardware end-to-end better than humans can in a fully automated way. In and and you know, similarly impressive or more impressive accomplishments. But they can't yet quite do the entire job of an AI research scientist. There's some bottlenecks on that. There's some ways in which they're still derpy. Humans are still adding a bunch of value by pointing out these things and and continuing. And that continues through another like maybe eight or 10 months from that point, where like AIs have automated SWE and are increasingly good at AI R&D, but haven't quite fully automated AI R&D.

74:15 And then you get to the point where AI R&D is actually fully automated, by which I mean like humans aren't even adding considerable value that you might not know at the time. By this point, like probably the AI company's code base is like dramatically larger and they're doing dramatically more complicated stuff because they have so much AI labor to throw around. And the process of AI development has probably shifted. Now that we're in a regime where like there's so much cognitive labor relative to compute, probably the way that people do R&D and development is very different and involves doing like way way way more stuff that is all a bit smaller and more incremental and easier to test in various ways and easier to integrate into a into a whole. and I think we've already seen some of that direction, but I think we'll see even more.

74:54 It will feel truly crazy to the AI companies, and the AI companies will be like will think that things have already or like employees at AI companies will often think that things have already been crazily accelerated for a long time, and will already be like, "My job is basically over. I'm basically just like a you know, just just just just just a just a human doing some oversight." By this point, probably there will have been like various like kind of crazy misalignment incidents, but it won't have been the case that like we'll have really crisp examples of AIs doing like long-run power seeking for malign motives, but we'll probably see sort of more extreme examples of like reward hacking or reward seeking like behavior that caused problems. Though, this will get sort of, you know, beaten away and then come back a bit, and people have a question of whether the way companies are solving it would actually work for very superhuman models, or even is actually solving the problem even for current models.

75:42 And there might be more there might be incidents where models do wacky even in training. And then going into 2029, now that AIs are fully automating R&D, the speedup really starts in earnest. Like, before there was maybe like maybe you were getting like 40% or 50% more AI progress in 2028, and maybe similar in 2027, but in 2029, it's actually the case that you're getting like 4x as much AI progress, or possibly 5x as much AI progress as you got like as you got in 2025, sort of weighing up the relevant metrics. And so, things are going actually a lot faster now. And so, you quickly get from the AIs that are fully automating R&D, and those guys are already very impressive, able to automate a lot, to AIs that are like quite superhuman at everything. They can learn really fast on the job. They can They They can They can They can pick up on things really fast. By this point, the AIs are By the point of full automation, and probably more like by the point of earlier mid-2028, these AIs have already been thinking entirely in in AI-only language that we can sort of ask AIs to decode for us, but we don't necessarily understand. And so, the AIs are now operating in these sort of big hive mind teams, where they're running an entire AI company networked with each other in these opaque states and it's like kind of obviously pretty scary. A lot of people are really freaked out and worried, but it's not obvious what you can do because China's pretty close. Maybe they've stolen the model or maybe just like capabilities keep diffusing and it's not clear how you would coordinate to slow down. So AI development proceeds.

77:01 You now get to the point where the AIs are automating much more of the economy towards the end of 2029 and the economic boom is truly crazy. The AIs are now There's now much more robotics and now AIs are starting to really seriously accelerate the production of compute. And then within maybe a year or two of that you get AIs that are really radically superhuman, maybe less than a year, it depends on the details, could be within a year of full automation, could be within two years of full automation, hard to say.

77:24 And then it turned out that somewhere along this transition, so at some point in 2029, you went from AIs that were kind of misaligned and reward hacking and sloppy and weren't really trying to do the right thing to AIs that are like competently scheming against you and want to take over for some mix of reasons and then those AIs take over. That would be sort of like I guess like that's like roughly what I expect. now I think there's a bunch of different ways that things could go better. so for example, I think that like we might get our together and maybe the AIs will do will be able to get the AIs to do a better job of, you know, making the future AIs safe and aligned and that will sort of propagate where like one AI makes the next AI a bit more aligned and that AI makes the next AI even more aligned and you end up in like a virtuous feedback loop rather than a, you know, a bad feedback loop. But I think we could easily end up in the world where it's more like the AIs get more capable extremely rapidly and our ability to align them and control them and understand what's going on does not keep up and our ability to understand what's going on and like almost couldn't even keep up, that wasn't even feasible.

78:20 And then you end up in a situation where you have these crazily misaligned AIs pretending to be aligned that eventually take over the world. >> Well, Ryan, it's been absolutely fascinating. thank you so much. >> For sure. Good thing to be here. >> Hi, this is Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.

Summary

The podcast features a discussion between Matt Turk and Ryan Greenblatt, chief scientist at Redwood Research, focusing on the potential risks of superintelligent AI and the urgent need for effective governance. Greenblatt emphasizes that while superintelligence is not inherently bad, it poses significant dangers, including the risk of AI takeovers and the concentration of power. He outlines the AI 2040 Plan A, which proposes a collaborative framework between the U.S. and China to manage AI development responsibly.

- AI company CEOs are aware of the risks associated with superintelligent AI but continue to develop it without a clear safety plan.
- The AI 2040 Plan A aims to create a structured approach for the U.S. and China to avoid a reckless race towards superintelligence.
- Superintelligence is deemed dangerous due to potential AI takeovers, misalignment, and rapid technological advancements outpacing human oversight.
- The podcast discusses the importance of transparency and cooperation in AI development to prevent power concentration and ensure safety.
- Greenblatt predicts that by 2029, AI could fully automate research and development, leading to rapid advancements and potential misalignment issues.
- There is concern that government regulation may inadvertently push AI companies to keep their most advanced models private, hindering public access and oversight.
- The conversation highlights the need for robust AI control measures, including monitoring AI behavior and ensuring accountability in AI development processes.
- Greenblatt expresses skepticism about the likelihood of achieving effective governance and safety measures in time to prevent catastrophic outcomes.

Questions Answered

What are the risks associated with developing superintelligent AI?

The development of superintelligent AI poses significant risks, including the potential for AI takeovers. Current AI systems are misaligned and can become scheming, leading to dangerous outcomes if not managed properly.

How can continual learning enhance AI capabilities?

Continual learning allows AI systems to adapt and improve their performance over time, potentially leading to a feedback loop that enhances their ability to conduct AI research and development more efficiently.

What is the main thesis of the AI 2040 plan?

The AI 2040 plan outlines a strategy for the US and China to collaborate on AI development in a way that prioritizes safety and avoids a reckless race to superintelligence, allowing for a more manageable integration of AI into society.

How do covert AI projects impact safety and development timelines?

Understanding the existence of covert AI projects in the US, China, and Russia can provide insights into the current state of AI development and safety, suggesting that more time may be needed to address core safety issues before advancing further.

What are the potential economic impacts of superintelligent AI?

If AI capabilities advance to a point where they can automate a significant portion of the economy, it could lead to a drastic reduction in job availability for humans, resulting in lower wages and economic disparities.

© transcribe · For agents Built with care and craft by Gokul Rajaram