transcribe

Inside the Rise of Autonomous AI Hackers: XBOW's Oege de Moor

Sequoia Capital · 8m · transcribed Sep 2026
More from Sequoia Capital Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

The Rise of Autonomous Hacking

What is the significance of autonomous hacking in cybersecurity?

The speaker discusses the evolution of hacking towards fully autonomous systems, emphasizing that those who leverage AI will dominate in cybersecurity, similar to historical military strategies.

  • Autonomous hacking represents a significant shift in cybersecurity.
  • AI-driven attacks will likely outperform traditional human-led efforts.
  • Historical parallels highlight the importance of adopting advanced technologies.
# 1:47

Demonstrating AI's Hacking Capabilities

How effective is the AI tool Xbo in identifying vulnerabilities?

Xbo demonstrated its effectiveness by discovering a critical vulnerability in Bing image search with minimal input, showcasing the potential of AI in cybersecurity.

  • Xbo can autonomously identify severe vulnerabilities with just a URL.
  • The tool's performance challenges the belief that machines can't fully automate hacking.
  • AI tools can be both cost-effective and efficient in cybersecurity.
# 3:35

Advancements in AI Hacking Tools

What advancements have been made in AI hacking tools like Xbo?

Xbo has evolved significantly, outperforming human hackers on platforms like HackerOne, and its capabilities are expected to improve further with advancements in AI models.

  • Xbo's performance has surpassed that of top human hackers.
  • The integration of multiple AI models enhances effectiveness.
  • Continuous advancements in AI will further elevate hacking capabilities.
# 5:23

Limitations of Traditional Code Analysis

What are the limitations of traditional code analysis tools compared to autonomous hacking?

Traditional tools like MyChelle's focus on white box testing and may not effectively identify exploitable vulnerabilities in real-world scenarios, unlike Xbo's black box approach.

  • White box testing may not reveal all exploitable vulnerabilities.
  • Real-world applicability of vulnerabilities is crucial for effective security.
  • Xbo addresses the gaps left by traditional code analysis tools.
# 7:11

The Urgency of Enhancing Cybersecurity Measures

Why is there an urgent need to improve cybersecurity defenses against AI threats?

The speaker emphasizes the need for enhanced cybersecurity measures in light of the rapid evolution of AI-driven attacks, urging stakeholders to prioritize vulnerability detection and response.

  • Cybersecurity stocks dropping in response to AI news is counterintuitive.
  • An arms race in cybersecurity necessitates maximizing AI capabilities.
  • Immediate action is required to stay ahead of potential threats.

Transcript

0:04 >> Thank you very much. You've all heard the story about the breach of the Mexican government. Human hackers used Open AI and Anthropic as assistants in order to achieve a massive data breach. What I want to talk to you about today is autonomous hacking where the AI does all the work without any human assistance.

0:37 The situation in cybersecurity today is akin to the Battle of Nagashino in 1575 in Japan. In this picture on the left-hand side is the army of Oda Nobunaga. Nobunaga was an upstart. He was a minor warlord, but he treated warfare as a system to be optimized. And in particular, he used the very latest weapons, the very latest guns. On the right-hand side is the Takeda clan. The Takeda clan was extremely famous and their cavalry was thought to be invincible.

1:14 They had many well-known warriors who had earned their their their skills in in battles before. But guess who won. The situation in cybersecurity is going to be exactly the same. Those with AI will win. Just to set the scene, let me tell you about one particular vulnerability. A couple of weeks ago, Microsoft announced a remote code execution vulnerability in Bing image search. Bing image search, one of the best secured systems in the world, very well secured by the engineers at Microsoft, but also hammered by thousands of hackers from all over the world trying to get in.

1:58 A remote code execution vulnerability, the very worst kind of vulnerability where you can run arbitrary code on the target system, complete takeover. This vulnerability was found by the product of my company, Xbo. And the only input it needed was the URL. Nothing else. And the cost? $3,000 at list price. That's not what it cost us. So, it's fast, it's cheap, and extremely effective.

2:32 The way Xbo works is very much like a human hacker. It starts by reconnaissance. It sends out a bunch of scouts, agents that discover the attack surface. It prioritizes what endpoints look most juicy, most promising for an attack, and then it goes in and tries every relevant attack type. Despite evidence like this, many human security researchers believe that it's impossible to completely autonomously carry out this task with a machine.

3:05 So, in order to counter the skepticism, already last year my company entered our bot, Xbo, onto the HackerOne platform. HackerOne is this platform that connects companies that want their systems to be tested with ethical hackers who will then go and attack those systems and report what vulnerabilities they find. If they report good vulnerabilities, they get paid a bounty and they get points. Within a few weeks, Xbo first became the number one hacker in the United States, and then in August, it became the number one hacker in the world.

3:38 And I have to stress this is completely black box testing. It's just like the Bing example I mentioned before. You only give it the URL, nothing else. The AI does the work completely autonomously. And that was back in August. The the foundation models that we are building on have enormously progressed since then. This is on a set of open source real web applications. These are not some Mickey Mouse cyber benchmarks. And we started back in March last year.

4:10 We started 37 and then it reached the top of the HackerOne leaderboard with an alloy of Sonnet 40 and Gemini 25. I I can't I can't resist briefly telling you about alloys. So, think of these attacks as a sequence of actions and at every step, you flip a coin to decide what model to ask. Either ask Gemini or ask Sonnet. This is much better than either model separately. It's a bit like like pair programming.

4:41 The two models compensate for each other's mistakes. So, then shortly after Xbo topped the the HackerOne leaderboard, GPT-5 came out. Just extrapolating from its performance, it would have done at least three times better. So, in August, Xbo was a little bit better than the best the best human on HackerOne. With GPT-5, it would have been three times better. And since then, the the models have only gotten better and as you can see, we better collect a new set of benchmarks because it's pretty much saturated.

5:19 So, how should you think about this in in relation to MyChelle's? MyChelle's has been as has mostly been been reported as a tool that reads the source code extremely well and points out potential flaws in the code. This is white box testing. It's not all like what I was talking about about before, purely black box testing. You actually have access to the source code, which of course is an advantage you do. But as an attacker, you don't necessarily.

5:51 The question with these with this code analysis stuff is are the weaknesses actually exploitable in the wild? And if they are exploitable, does it matter? What's the impact? Where can I go if I get into if I can execute remote I I can do remote code execution on a Bing server, where else can I get to? I I can't tell you. And then of course there's many other vulnerabilities that are configuration or deployment problems. You can't actually use them from the source code itself. So, these are the questions that Xbo answers for you.

6:28 If you get to know about a exploits, it's probably already too late. So, you typically like before the Bing example, people publish a CVE to let the world know that there was a vulnerability. Back in 2018, it the delay between publication of a CVE and bad actors exploiting it in the wild was almost two two and a half years. Today, the number has gone negative. For most CVEs, it is already being exploited before the the CVE is even published.

7:08 So, in view of all this evidence, it's incomprehensible to me that whenever there's news about AI and security, cyber traditional cybersecurity stocks drop. This makes no sense at all. We need every possible defense that we can get against these autonomous AI-powered attacks. So, so far I've been preaching like Nostradamus, telling you about all the bad things that might happen. So, let's try and rally the spirit of Nobunaga, that Japanese warrior I talked about at the beginning, and see what can be done.

7:42 So, first of all, all of you, everyone who's working on frontier models, you must maximize the cyber capabilities. No more talk about whether it's safe to do that or not. We're in an arms race, so we have to make sure that we have the very best models to power this type of work. Secondly, we need to enable human security researchers to use this as an extension of their own work in order to to maximize the chances that we find all the vulnerabilities before the bad guys do.

8:13 And finally, you need to prioritize what matters. You need to know whether the bugs are truly exploitable and what their impact is going to be, and Xbo can help with that. We've got about 6 to 9 months to do this. Just extrapolating from the from the from the progress we software engineering agents, in 6 to 9 months we have we will have open weight models that are just as good as MyChelle's and similar models.

8:45 And so, if you want to have a nice Thanksgiving dinner with your family, you better start fixing now. Thank you. >>

Summary

The speaker discusses the emerging threat of autonomous AI-driven hacking, highlighting its potential to surpass human capabilities in cybersecurity. They draw parallels to historical battles, emphasizing that those who leverage AI will dominate in cybersecurity, as evidenced by the success of their AI tool, Xbo, in identifying vulnerabilities more effectively than human hackers.

- Autonomous hacking using AI is becoming a significant threat in cybersecurity.
- The speaker compares the current cybersecurity landscape to the Battle of Nagashino, where technology (AI) will determine success.
- Xbo, an AI tool developed by the speaker's company, autonomously identifies vulnerabilities with minimal input, achieving top rankings on HackerOne.
- Recent vulnerabilities, like the one in Bing image search, illustrate the speed and effectiveness of AI in exploiting security flaws.
- The gap between the discovery of vulnerabilities and their exploitation is shrinking, with many being exploited before they are publicly disclosed.
- The speaker argues for the necessity of enhancing AI capabilities in cybersecurity to counteract autonomous attacks.
- Collaboration between AI tools and human researchers is essential for improving vulnerability detection.
- Urgent action is needed within the next 6 to 9 months to bolster defenses against AI-driven cyber threats.

Questions Answered

What is the significance of autonomous hacking in cybersecurity?

The speaker discusses the evolution of hacking towards fully autonomous systems, emphasizing that those who leverage AI will dominate in cybersecurity, similar to historical military strategies.

How effective is the AI tool Xbo in identifying vulnerabilities?

Xbo demonstrated its effectiveness by discovering a critical vulnerability in Bing image search with minimal input, showcasing the potential of AI in cybersecurity.

What advancements have been made in AI hacking tools like Xbo?

Xbo has evolved significantly, outperforming human hackers on platforms like HackerOne, and its capabilities are expected to improve further with advancements in AI models.

What are the limitations of traditional code analysis tools compared to autonomous hacking?

Traditional tools like MyChelle's focus on white box testing and may not effectively identify exploitable vulnerabilities in real-world scenarios, unlike Xbo's black box approach.

Why is there an urgent need to improve cybersecurity defenses against AI threats?

The speaker emphasizes the need for enhanced cybersecurity measures in light of the rapid evolution of AI-driven attacks, urging stakeholders to prioritize vulnerability detection and response.

© transcribe · For agents Built with care and craft by Gokul Rajaram