Section Insights
Long-Horizon Cyber Tasks and AI Capabilities
What are the implications of AI's ability to execute long-horizon cyber tasks?
AI models have developed the capability to not only find bugs but also to devise multi-step plans for executing complex cyber tasks, which poses significant risks if these capabilities fall into the wrong hands.
- AI can autonomously plan and execute cyber attacks, mimicking human-like strategic thinking.
- The real danger lies in AI's ability to create multi-stage plans rather than just finding vulnerabilities.
- Current discussions around AI safety often overlook the potential for autonomous execution of malicious tasks.
The Rapid Advancement of AI Capabilities
How quickly can dangerous AI capabilities be accessed by malicious actors?
The rapid pace of AI development means that advanced capabilities could be available to malicious actors sooner than expected, raising significant security concerns.
- Generative AI capabilities evolve quickly, often outpacing regulatory and security measures.
- The potential for adversaries to access advanced AI tools poses a changing landscape for cybersecurity.
- There is a pressing need for vigilance in monitoring AI advancements to prevent misuse.
Chinese AI Models and Cybersecurity
What are the implications of the capabilities of Chinese AI models in cybersecurity?
Chinese AI models may possess superior capabilities that are not publicly available, which could be leveraged by adversaries to enhance their cyber operations.
- Chinese companies likely have advanced AI tools that are not released to the public, indicating a hidden advantage.
- Adversaries could invest in tuning these models for cyber operations, increasing their effectiveness.
- The gap between publicly available AI capabilities and those used by state actors could pose a significant threat.
Defensive Measures and AI Limitations
What challenges do organizations face when defending against AI-driven attacks?
Organizations may find themselves limited by the defensive capabilities of AI models, which can be restricted due to regulatory measures, forcing them to seek alternatives.
- Regulatory restrictions can hinder the effectiveness of AI in cybersecurity defense.
- Organizations may need to rely on foreign models for defense if domestic options are limited by regulations.
- The situation highlights the need for a balanced approach to AI regulation that allows for effective defense mechanisms.
Policy Responses to AI Threats
What should be the government's response to the emerging threats posed by AI?
The government needs to focus on fixing vulnerabilities in AI systems and enhancing the ability to respond to cyber threats at machine speed, rather than imposing restrictive measures that could hinder progress.
- A proactive approach is necessary to address the vulnerabilities in AI systems before they can be exploited.
- The government should prioritize rapid response capabilities to counteract AI-driven threats.
- Restrictive policies could stifle innovation and leave organizations vulnerable to attacks.
Transcript
0:00 This thing had the ability to find bugs, yes, chain them together, yes, but then to think through all of those things for an ultimate goal. And so, this is what people call long-horizon cyber tasks, and it had the ability to do that with like the level of skill that you would have of the manager of a TAO team at NSA right now. So, that is what is So, TAO Targeted Access Operations was like I think they've renamed it, but it was like the team at NSA that would do all of the breaking into you know, other governments at American targets.
0:40 >> Right? So, like So, that is what is like really it For a while now, these models have been really good at looking at software and being like, "I found a bug." And then, here, let me write an exploit for you. What's really impressive here is this thing was like, "I want the answers from Hugging Face." And it came up with a plan of like, "How am I going to get out of this network, get across the internet, and get into Hugging Face?"
1:06 And that is like the long planning here is very human-like, and that is what is like actually really scary here. And so, that is what we need to when we think about the danger of these models, we have to stop thinking about the bug finding, because that is what caused the White House to do this spectacularly stupid thing that we talked about on stage, which was to ban Fable, because that is the mechanism that is the thing that is most useful for defenders right now and people who own code is finding bugs and fixing them.
1:37 Where the real danger here is the coming up with a multi-stage plan to execute autonomously, because what you really don't want is you don't want somebody be able to say to their model, "Hey, I would like to steal the I would like to steal money. Go figure it out for me." and then let it work for 12 hours, and just steal money for you. Which is this model would clearly be able to do that. >> Right. And so, I think what you're getting at is you know, one solution is going to be in testing, you want to fully disconnect these models from the ability to like break out and get onto the internet.
2:14 >> Yes. >> But, that's just solving the testing issue. The real problem here is that AI models have achieved this capability. That they're able to do this now. Not only finding the bugs, not only the breaking out, but to be able to do these multi-step plans and then execute. >> Yes. >> And if this is sort of the latest unreleased Open AI model, well, the history of generative AI has told us one thing, and that is that the frontier is only the frontier for a few months, maybe 10 months, maybe a year, but not not much longer than that. And so, if Open AI is seeing this in testing now, the real is the real danger that this type of capability does end up in the hands of you know, evil doers, you know, faster than a lot of people might expect. Cuz if if that's the case, that changes everything.
3:11 >> Yes, that's right. And so, you know, our best knowledge on where say the open weight models are, comes from the AI security institute, which is the UK government's group that does these assessments. they released just this week an assessment with GLM 52. Unfortunately, you know, the Kimmy Kimmy is really Kimmy you know, three is the best of the Chinese models now. The the open weights have not been released, so you can't really do a good assessment for Kimmy yet. and what we're finding is the Chinese models are not cyber tuned out of the box.
3:48 So, it is very likely that the Chinese companies are not are intentionally I'm not going to say neuter, but they're intentionally not making their models really good at cyber. And there's a couple of possible reasons for this. They're probably trying not to tickle the dragon's tail of the PRC overlords. because what they don't want to do is they don't want to trigger a crackdown for their exports. But what happens is is if you take those models and you bring them into your own lab and you have a training set of labeled vulnerabilities, if you have a cyber gym, then you can make them much better yourself. And the amount of resources it takes to do that is not extremely high.
4:33 It's in the tens of thousands or hundreds of thousands of dollars. It's not in the hundreds of millions or billions of dollars. So what that means is one, the Chinese absolutely have better capabilities in house than what we can see on the charts, right? Because I guarantee then what those companies are offering to the People's Liberation Army and the Ministry of State Security is way better than what they're releasing publicly. both for from a profit perspective and a keeping the government happy perspective. Second, it means that other adversary groups are going to take the Chinese models and then spend the several hundred thousand dollars or millions of dollars necessary to to create tuned models. And we are probably not far away then from just going to Hugging Face and getting a you know, Kimmy 3 cyber tuned model that can do both especially the short right short horizon stuff much better than by default and then eventually the long stuff. The the short stuff's easy to train cuz all you need is a bunch of bugs, so you can just go get a bunch of CVEs and train it. The long horizon stuff's harder because you have to build these like cyber gyms and such. It it's not impossible because there there are a bunch of CFPs and examples out there, but you you can do it.
5:52 >> What are CFPs? >> I'm sorry, not CFPs, CTFs. Capture the flag. So, like you can use like capture the flag training sets and all that kind of stuff. So, anyway >> basically puts the AI in the gym and sort of has it work through all the steps in order to meet this objective. >> Yeah, yeah. And so, like people have had, you know, training for humans and for hiring purposes and all that kind of stuff. and so, anyway, what what the the AISI has said is that the the difference between the frontier in the Chinese models is about 7 months, but I would argue that that underestimates it because the Chinese models that we see are under trained.
6:27 So, that the internal Chinese capabilities are probably much closer to the frontier. Now, what what what happened with OpenAI is beyond the frontier because when we say the frontier, we're talking about what's released, right? Right. So, But, yes. So, the What that means is this capability is coming for adversaries. And so, we need to get ready to defend against this capability.
6:58 What Hugging What And then, the other funny part of the story is before we knew this was OpenAI, Hugging Face announced we were attacked by an AI attacker. We don't know who it was. When we tried to defend ourselves, we tried to defend ourselves with an AI system. We used a US frontier model. And the US frontier model shut down and refused to defend us because of a cyber protection put in place. Those are the cyber protections that were required by the Trump administration. So, we had to switch to a Chinese model to defend ourselves. So, we switched to GLM 5.2.
7:31 So, before we knew it was OpenAI, Hugging Face wrote this blog post saying, "Everybody should have at least a Chinese open weight model ready for defense because you might find yourself in a situation where get cut off from an American provider for defensive purposes. I expect that actually wasn't Open AI. I expect it from from their description it sounds like it was an Anthropic model. >> Right. >> So >> So we have this hilarious situation where a American company loses control of their model and attach and attaches a French company. The French company turns to a different American provider for defense and that American company says, "Ooh, that's a cyber problem. I can't help you." And so they have to turn to a Chinese provider to protect them because the White House forced that other American company to have protections because they're a French company that they can't use. It's really kind of a weird sci-fi podcast, I guess.
8:22 >> These American models already had refusals on anything cyber. Like one of the knocks on Fable was that it was would refuse like let's say for bioterrorism. If you asked about mitochondria, it wouldn't answer. So was this really the government or is this just the models own safeguards that they're putting in and I think one of it just to put one detail on one of the interesting things is open source it doesn't basically matter if you're an attacker if you have open weights it doesn't matter if you're an attacker or if you're a defender you can use them without restrictions.
8:50 >> Yes. >> The problem with the restrictions that we're seeing from these closed models is that they can't really they can't really differentiate. So they in order to prevent attackers from using their models they're also basically wholesale you know refusing anything on cyber which means that if you're trying to defend also you can't use it. >> That's right. Well, in the blog post that Anthropic put up when they turned Fable back on they said, "We have to tune up our defenses on cyber way too far because of the White House." So they specifically said that of the precision recall trade-off is we have to tune towards recall versus precision. Right?
9:29 So we will have way too many refusals and so we're in this weird place where they they are doing they're saying they are saying all the time, "I can't do that for you. I can't do that for you." And they're basically being forced to by the White House because the White House has still not defined what is the appropriate level of refusal. And apparently the White House is still hand approving who Anthropic is allowed to let into their cyber program. now, the funny thing is now Open AI has said, "We have approved Hugging Face for our tech program even though they're not an American company." So, I don't know how they were allowed to do that if they just went over the top of the White House or they got like emergency approval or something.
10:09 we will see what the policy res- response is from the White House from Open AI's announcement. I hope there is not a crackdown. That will be the natural response of the White House, but it needs to be the opposite because what this demonstrates is yes, Open AI screwed up, all right. They need to have fixes. But this is coming, right? This level of capability will be in the hand of every adversary, every American company faces. So, the response of the White House needs to be that we have to one, fix the bugs, two, find the bugs, patch them everywhere, and then we have to have the ability to respond at machine speed. So, every American company needs to have AI watching for their defenses. And it's going to be because the attackers are just going to tell their AI, "Go attack this guy."
11:03 And the defenders have to tell AI, "Defend me." Because no human being can defend against this. You cannot have a human being watching your logs anymore. Or at 2:00 a.m. you get a page, and a human being has to be like, "Oh, okay." And then log into Slack and take 15 minutes to log in and look at the log and figure it out. By that point, you're toast.
Summary
- AI models can autonomously devise multi-step plans for cyber attacks, surpassing mere bug-finding capabilities.
- The NSA's Targeted Access Operations (TAO) team exemplifies the level of skill AI models can achieve in cyber operations.
- Current AI models, including those from OpenAI, may soon be accessible to malicious actors, raising urgent security concerns.
- Chinese AI models are likely more capable than publicly available versions, as they may be tuned for cyber operations in private settings.
- The U.S. government's restrictions on AI models hinder defenders' ability to use these tools effectively against cyber threats.
- A notable incident involved Hugging Face needing to switch to a Chinese model for defense after their American model refused to assist due to government-imposed safeguards.
- The White House's response to AI capabilities should focus on enhancing cybersecurity measures rather than imposing further restrictions.
- There is an urgent need for AI-driven defenses that can operate at machine speed to counteract automated cyber attacks.
Questions Answered
What are the implications of AI's ability to execute long-horizon cyber tasks?
AI models have developed the capability to not only find bugs but also to devise multi-step plans for executing complex cyber tasks, which poses significant risks if these capabilities fall into the wrong hands.
How quickly can dangerous AI capabilities be accessed by malicious actors?
The rapid pace of AI development means that advanced capabilities could be available to malicious actors sooner than expected, raising significant security concerns.
What are the implications of the capabilities of Chinese AI models in cybersecurity?
Chinese AI models may possess superior capabilities that are not publicly available, which could be leveraged by adversaries to enhance their cyber operations.
What challenges do organizations face when defending against AI-driven attacks?
Organizations may find themselves limited by the defensive capabilities of AI models, which can be restricted due to regulatory measures, forcing them to seek alternatives.
What should be the government's response to the emerging threats posed by AI?
The government needs to focus on fixing vulnerabilities in AI systems and enhancing the ability to respond to cyber threats at machine speed, rather than imposing restrictive measures that could hinder progress.