Section Insights
Welcome to Black Hat Asia
What is the purpose of Black Hat Asia?
Black Hat Asia serves as a platform for sharing knowledge in the security community, showcasing the journey from learning to leading in cybersecurity. It emphasizes the importance of community engagement and the evolution of security practices.
- Black Hat spans the full career journey in cybersecurity.
- The event aims to share real security happenings and foster knowledge transfer.
- Black Hat India will be launched later this year, expanding the community globally.
- Attendees are encouraged to engage, ask questions, and explore various learning opportunities.
AI in Security: Separating Fact from Fiction
What are the key points regarding AI's role in finding vulnerabilities?
Ari Herbert Voss discusses the dual nature of AI's capabilities in security, highlighting both the real advancements and the aspirational claims. He aims to clarify the fundamentals of AI's impact on security and what it means for the industry.
- AI is increasingly being used to find vulnerabilities and write exploits.
- The talk aims to distinguish between real advancements and aspirational claims in AI.
- Understanding AI fundamentals is crucial for grasping its implications in security.
- The speaker has extensive experience in AI and offensive security.
Future of Offensive Capabilities
How will offensive capabilities evolve in the coming months?
Offensive capabilities are expected to scale, but effective deployment will depend on ergonomics and building practices. The speaker emphasizes that the field is still developing and that initial concerns should not distract from ongoing progress.
- Offensive capabilities will continue to evolve and scale.
- Effective deployment relies on improved building practices.
- The field of offensive security is still in its early stages.
- Initial panic over new models should not overshadow ongoing advancements.
Challenges with Synthetic Data
What are the challenges associated with using synthetic data in AI models?
While synthetic data can provide learning signals, it often lacks richness compared to real data. The speaker discusses the need for intelligent methods to maximize learning from both real and synthetic data while mitigating the downsides of synthetic data.
- Synthetic data can be useful but often lacks the richness of real data.
- Intelligent methods are needed to extract valuable learning signals.
- Understanding the limitations of synthetic data is crucial for effective AI training.
- Future developments will focus on improving data generation techniques.
Advice for Aspiring Cybersecurity Professionals
What advice does the speaker have for students interested in cybersecurity?
The speaker reflects on their own journey and emphasizes the importance of adaptability and continuous learning in the rapidly evolving field of AI and cybersecurity. They encourage students to stay informed and seek out opportunities for growth.
- Students should remain adaptable and open to learning in a fast-paced environment.
- Continuous education is vital in the fields of AI and cybersecurity.
- Networking and engaging with the community can provide valuable insights.
- Practical experience and experimentation are key to developing skills.
Transcript
0:10 Good morning everyone and welcome to day two of Black Hat Asia. It's fantastic to see so many of you here again and a very warm welcome to those joining us today. One of the things I love most about Black Hat is that it spans the full career journey. In fact, this morning's keynote speaker, Ari Herbert Vos, first came to Black Hat as a student attendee, and today he's back on stage to share his work with all of you. That progression from learning to contributing to leading is exactly what this community is about.
0:45 Black Hat exists to share what's really happening in security and to make sure that knowledge is challenged, built on, and passed forward. As part of that, we're continuing to expand the community globally and later this year, we'll be launching Black Hat India in Bengaluru from 27th to the 30th of October. This week, we set the context. We looked at a landscape being reshaped by AI, expanding attack surfaces and increasingly by geopolitics.
1:16 Today you'll see that play out in more depth, particularly as we explore how offensive capabilities are evolving across the program that comes through in the research being shared and the techniques being tested in our briefings. From emerging vulnerabilities and AI security to real world attack methods coming directly from the community and beyond the briefings, there are plenty of ways to engage the stages to continue learning and connecting across the program. the arsenal for hands-on access to tools and techniques, bricks and picks and the lounges, spaces to play, experiment, but also critically to connect.
1:55 So, I'd encourage you to explore, ask questions, and make the most of being here. With that, it's my pleasure to welcome the founder of Blackh and president of Defcon, Jeff Moss. Hey, good morning everyone. Thank you for coming to our second day of the 22nd Black Hat Asia. Okay. So, just a couple of comments before we introduce our day two keynote focused on the subject of AI, which I'm surprised like the room isn't overflowing with people so much misconception, so many hopes and dreams wrapped up in AI and thinking about how quickly it's moved the comments that I would have made a year ago or two years ago on AI and how they might have changed to today. one of the observations is great power competition hasn't changed. It always takes different forms. And so there is a bunch of great power competition going on around AI. And it's basically following into sort of China, US and Europe for a little bit, but primarily I would say US and China with the energy, the compute, the mathematicians, the scientists, the national giant hyperscalers all racing.
3:26 and one thing I tell people is like different nations want different things. They they have different ambitions. They're on different trajectories, different budgets, and so their national interests are to you sometimes might look irrational, but to them it's rational. So, I'll just give you maybe a little glimpse of this. at one point years ago when I was at ICAN and China really wanted a rootname server, their own root name server like I know ABCD FJ L like M root M or something 14th root name server and they're like well that it doesn't really work that way. They really want it and they've wanted it for a long time. So they figured what they would do to try to show us how ready they are to have a root server. They were going to fly in from around the world a hundred of the leading DNS security experts. They had a building a hundred million dollar building built to hold the root server. They had a thousand engineers in the building ready to operate the root server. Right? That is a huge commitment. and this and they're going to fly everybody in to show off like look this is how seriously we take running a root server and all I'm thinking is our root server is like a one use server in Iraq you know it's like right here here's the root server and they are allocating hundreds of millions of dollars to this for national pride and the roots working fine in China but it was a national pride issue think about how much more scaled up that is now in the era of AI, especially now that they're starting to look like there's some actual economic benefits. So, so let me give you an example.
5:36 oh, one final thing with a with a great power competition is we see a lot of theft of model data, training data, poisoning of training data, weights. it could be from competitors, it could be from nation states. so all the normal things that you might see in a military or telecom, it's all happening in AI. So you might be heads down and you might not think that what you're doing is of any interest. Well, maybe not to you, maybe not to your government, but maybe to another government or another competitor. so you can't know everything all the time. I'm just saying that compared to two years ago, the intensity has ratcheted up and so has the spend. also if great power competition sometimes leads to irrational looking behavior, green field opportunities also sometime sometimes lead to irrational looking behavior, right? To spend on AI data centers doesn't make much sense. But the green field opportunity in the eyes of some of the business leaders is unbelievable. If AI is the new electricity, of course you would spend unlimited other people's money to try to capture electricity in a bottle.
7:01 So the opportunity is perceived as unbounded. And if you look at how much money search has made now imagine if if search disintermediated us from a site from an original source imagine how much AI sort of synthetic summarization will disintermediate us even further. On the way over I was reading an article that said search organic search from an enduser to a site is now down about 90%.
7:31 Like end users never get to the source site anymore. they get the summarization and leave and it's destroying ad revenue. Rational behavior on the other hand looks something like what happened with Taiwan semiconductor and the extreme ultraviolet chip manufacturing process. It took decades to manufacture and and master extreme ultraviolet lithography and there were three companies doing it and they're all spending like unlimited money. I think TDK in Japan, there's an American company, there's a Dutch company, and in the end, the three companies came together and said, "We're all spending money trying to do the same thing." And instead, they became a consortium. They all three joined up, and that's how the machine was finalized. See, that seems rational, but we're not in a rational mode right now.
8:28 I guess if I looked at the trajectory of what's worked and what's not, I would say it was really clear, show of hands, how quickly was it obvious to you that the worldwide web was going to be a winner? Like you saw a web page in a browser and an image and you were like, "This is it." Like that was obviously quickly within I don't know months there were guides online of every website. It was how quickly do you think faxes took off?
8:58 Faxes pretty quickly were a winner. Email winner. Cloud winner. Really quickly cloud caught on. We were making fun of it and then it ruled the world. web 3, blockchain, NFTTS, metaverse, VR, like no sounded interesting. didn't catch on. Aenic picture generation of cats. Okay, kind of interesting. But now agents, agents seem to be the thing that's taking off. And so my prediction is agents might be the thing that actually launches AI I guess efficiencies and unlocks new behaviors that are beyond just a rendering of pictures or a predictive autocomplete agents look like that's where the money will be. And so that's why I think it's accelerating so quickly. at Defcon 3, Curtis Carnau, who became a superior court judge in California, gave a talk on, what is it called? agents in the telco context.
10:16 Defcon 3. So that's what 31 years ago, he gave a talk about what's the liability if agents on a telecommunications network start trading or shopping for their cons for their owners, operators, and they make a mistake. Where does the liability lie? So none of this is new. Lawyers and judges have been thinking about this for 30 years. The difference is in the last six months. It's possible now. So you're going to see a huge amount of data from the last three decades unearthed and now they can actually try to apply it and see what happens. So that's why I'm super fascinated and interested in what Ari is going to talk about in his keynote today. there's so Susie mentioned Ari was a student awardee and then became speaker multiple times at Black Hat now is going to be a black hat keynote. he came to Beijing when we did Defcon China in 2018 and 2019 with AI Village and they kind of came to AI of course before it was called AI. I think maybe it was machine learning or expert systems, but it was more from a neuroscience and a mathematics perspective almost 10 years ago, 2016, 10 years ago, and has been in the field ever since. And now runs their own company focused on offense.
11:47 So super fascinating. Glad you're here for the keynote. And it's my pleasure to welcome Ari Herbert Voss. >> >> Good luck. Hello. Oh, I'm very excited to be here today. It's my extreme pleasure to be here to talk about this very important topic because, as I'm sure you're aware over the past couple weeks, you've probably seen a lot of headlines about AI finding vulnerabilities and writing exploits. A lot of this is real. and some of it is also aspirational. So, this talk is about separating those two and getting down to what it means for our industry.
12:27 By way of introduction, I happen to have been the first security hire and a core contributor to GBD3 and to Codex back at OpenAI. And currently I run a company where we automate offensive security at scale for everyone from hot Silicon Valley companies to Fortune 1000 enterprises and financial institutions. In this talk, I want to get four things across. What are the AI fundamentals behind the capability jumps that we're seeing in the news? What are what are the things that are getting better with scale? What hasn't gotten better with scale? And what does this mean for security? I'm also going to reserve some time at the end for some live Q&A.
13:06 So, let's get into talking about some of the fundamentals. LLMs are basically deep learning models that were originally designed to generate text. Many of the AI models we have seen applied to security in the past are trained for what we call discriminative problems. And so this is things like malware classification as a classic example. LLMs are different because they're generating rather than classifying data. A good model for thinking about LM is as an adaptive lookup table or a compressed representation of training data. The lookup occurs when you ask the model to predict the next word in a sequence.
13:38 Much like your smartphone keyboard, when we say words, we really just mean tokens. Tokens are integers that represent parts of words. And LLMs are just fancy lookup tables for tokens. And during training is when we build that table. We create dynamic key value pairs for each input token by generating key query and value vectors for each position in the input token sequence. And these attention weights represent how much focus or attention the model should place on each element in the sequence when processing the current element.
14:07 When you hear about weights in conjunction with models, this is what people are referring to. Weights or parameters can be thought of as a compressed representation of the input data that we sample from in order to complete the input sequence. So, how do you get from building a lookup table to solving math equations and other advanced reasoning capabilities as we're seeing in the news? You might be thinking, Ari, sure, there's probably more going on than simply predicting the next token in a sequence. But no, really, that's really all it is. Broadly speaking, we call this fshot learning. It's called fshot because you don't need a lot of examples. You start with a base model.
14:43 So here's an example of fshot learning in practice in terms of how humans think about things. So throughout our lifetime we see enough data to learn how to read and make connections between concepts. So somebody can ask you a question like perform a mathematical calculation and using what you know about the meaning of translate hex and decimal as we see in this example. you can infer that you need to transform one sequence of characters into another. I didn't need to tell you that mathematically which operations to do because you've already had enough training information already to do that.
15:15 As hackers, we've already seen enough examples of hexadimal and you know maybe read enough hitchhiker's guide to the galaxy to know that in this example 2A is also 42 in decimal. One of the most important things that's happening right now is what we call the scaling hypothesis. And so what that means is for more data plus more compute plus more parameters then you get better performance across a variety of tasks. And this has held surprisingly well over the last about seven plus years.
15:47 What has happened recently is it turns out that raised capabilities are scaling what we call superlinearly rather than linearly. So in in this graph here you see linear goes like this. Super linear goes like this. This means that when you train a model that is twice as big for twice as long on twice as much data, you can get a model that is four times as capable. This is the difference between the last generation of models or the latest generation of models from Anthropic and OpenAI as in the difference between Mythos and Opus and GPD 5.5 which just came out literally today and the previous generation of GPT models.
16:22 Additionally, as models get better at writing software, they get better at breaking it too. Reasoning capabilities are a proxy for vulnerability finding because this is a task that requires pulling in different sources of information and synthesizing it in a new way in order to accomplish a goal. So it moves beyond basic pattern matching. Let's talk about what is getting better with scale. The capability ceiling is rising which means that there's less of a need for complex scaffolding to get the models to do what we want. Opus 4.6 six had near zero exploit success with the same setup that Anthropic used for experimenting with mythos in the press release and research that they released a couple weeks ago. Simpler agentic loops can go further than they ever have before which means that the model itself can now handle a lot more of the process than previously.
17:14 We also have the speed at which vulnerabilities are being discovered is increasing. We have this handy graph courtesy of the folks who put together the zero day clock. And most notably at the very right hand side here, we see that between 2023 and 2026, the average time to exploit has dropped from 5 months to 10 hours. Early reports from Project Glass Wing partners indicate meaningful impact as well. Palaja Network said that they accomplished a year's worth of pentesting in less than three weeks. If you run a bug bounty program, chances are you've also felt this change in the volume of reports you're seeing as bug bounty hunters incorporate models into the bug finding practices.
17:54 Let's also talk about what hasn't gotten better with scale. The capability floor isn't rising as quickly. What this means is that improvements are nonlinear across different classes of vulnerabilities. So from the example that in the entropic report they have this what we call the OF OSS fuzz valuation. So for prior models they have a tiering system for how they think about how to weight different vulnerabilities that they find. So between 200 and 300 they're for tier one and tier 2 type vulnerabilities they have about 200 to 300 type crashes that they found and about one tier three. They don't have an eight and four or five, but for Mythos, they have almost 400 tier one and two crashes, which is a huge jump. They also have a handful of the tier threes and fours, and they also have 10 tier fives from a control flow hijack. This indicates massive gains at low severity around shallow bugs, modest gains at mid-tier bugs, and relatively sparse gains at deep exploitation. That's not to say this is anything to write to not write home about because it's very important. but it does provide some evidence that these gains are nonlinear and that the floor isn't rising as quickly as the ceiling is in terms of capability.
19:14 Models also specifically continue to struggle with certain bug classes that require holding state. This has improved dramatically as context window sizes have have increased. but again it points to this nonlinear improvement. So for these types of bugs, it means things that usually involve you know some amount of time or or other types of of state. concurrency is also a pretty big one too. I also mentioned bug bounty earlier as a bit of subtle foreshadowing because for the volume of slap reports this is also meaningfully increased as well. this is improving but not as fast as the ceiling of what is possible to find.
19:55 Improvements are also nonlinear across exploit steps. So for validation, exploitability, and reliability, these all improve, but much slower. So for now, meaningfully assessing impact is something that still requires more than simply asking a model to just tell you what it thinks is the highest impact. The net result is that better models find more bugs. but that also means it's more outputs for you to sift through. Models do not guarantee that those findings are going to be worth your find.
20:24 What this means for attackers relying on LLMs is that reliability and targeting is still, you know, a bit of a game of chance. Individual attackers need to get lucky when they rely on models to find exploits, but many iterations are required if you want specific impacts on specific targets. Anthropics experiments with mythos still boil down to 198 human reviewed findings that sit behind a much larger pool of automated outputs. The model generates volume, but for end-to-end bug finding, humans are still doing a large amount of filtering, validation, and impact termination.
20:56 Defenders are unfortunately going to get hit by millions of, you know, monkeys with typewriters. And some of those monkeys will write very good exploits. but defenders are going to have to win every time, whereas attackers will only have to get lucky once. So, what does this mean for security? It is possible. It is possible to get similar performance with enough scaffolding on open source models which is very interesting. it also indicates that expert guidance is something that's still quite important here and this has been true since 2023. The amount of expert required scaffolding is decreasing over time but it's still possible to get mythos level results with enough effort and enough expertise.
21:36 Most bugs that cause serious damage are relatively simple and don't require mythos level reasoning, which means that we're in a very interesting period where the bugs that are being found are a bit more shallow, but we're finding them at much higher volumes. And open source is also enabling just about anybody to start using models to find vulnerabilities. We can kind of think about fuzzing as being a mirror for the future. When fuzzing first entered the scene, early consensus was that, you know, we'd maybe solved a bug finding. But the hard part wasn't in generating findings. It ended up becoming understanding which bugs were exploitable and if they could have impact in real world environments. LMS are following that same trajectory.
22:20 There are also some implications here for defense and depth. So running multiple models on the same target leads to a surprisingly low overlap in findings. In the near future, it's unlikely that one model will find everything because we're finding that when you have different experience across different types of models, you can usually find a lot more and have higher coverage. There is currently no single best model that gives full coverage and the best results still come from layering models and approaches.
22:45 In closing, the pace of vulnerability discovery is going to get faster. Open source models and the accessibility of frontier models is going to continue to increase. There are extreme economic pressures in the AI industry to broaden access to these capabilities and that is true for both good and bad use cases. This is probably the best thing that could have happened to our industry because this is the forcing function that we need to do the things we probably should have already been doing.
23:07 It's now extremely easy to find high volumes of basic bugs which happen to be the types that also cause the biggest headache. I'm personally very excited because the ceiling is now higher on interesting and complex vulnerabilities that the LLMs can find. multi-chain, but vulnerabilities that previously took months to build now appear literally overnight with an on an industrial scale. The floor is not rising as quickly, which means there's still a lot of work that will continue to go into prompting, validation, and proctor context to get maximally informative results. The hard work of prioritizing and patching will be even more important in the coming months. Offensive capabilities will continue to scale, but the ergonomics around how you build with them are going to are going to be core to effectively deploying them. There's still a lot to build and our work is cut out for us. So don't let the initial panic of Mythos or any of these other models that are coming out distract you because we're just getting started.
24:00 Now that we've set the stage, I'd like to welcome Jeff back up and we're going to open things up now for Q&A. Okay. So, I've got one or two questions to get things going and then we're going to move to the microphones. There's one on each side. remember to ask a question in the form of a question and we'll take it from there. Okay. So, first off, since we are in Singapore, we're in Southeast Asia, how do you maybe tell me how you think about either the researchers or the work you're seeing that's US, Europe, China, like how do you contextualize what they're focusing on or what you're reading? What do you what's your thought process around those? Are we all pursuing the same thing or are we pursuing different things for different reasons? That's a fantastic question.
24:49 So, when I first got off the plane here, I was affronted with this giant ad for Quen. We don't have any ads for Quen in the US. open source is not as big of a topic in the US. There's HuggingFace and a couple other companies that are really carrying the torch there, but for the most part in the US, a lot of the focus is on the reasoning capabilities that are coming out of the large language models from Anthropic, OpenAI, Google, and Meta. So in this market, it seems like there's a lot more effort in figuring out how to do a lot more of the scaffolding that I was talking about earlier where you can get these smaller models, run them more locally, and be able to get the same performance out of like get mythos level performance out of a collection of smaller models. And so then Europe >> in Europe as well there's sort of similar dynamics because we have Mistral the French model out there that is open source and we're seeing very similar dynamics. The US is very unique in some ways because it has the majority of the large models that are the most capable for reasoning. So everyone else is figuring out different strategies to how to how do you squeeze out that same kind of capability when you don't have the same resources as as the US companies do.
25:58 >> And then China it seems u maybe they have a higher emphasis on inference because they're doing maybe more manufacturing and that might be more important to them than the training or do you see different companies different countries balance between training inference? >> Yeah that's a good question. So, one thing to note is that the GPUs that you use for training are not going to be the same GPUs that you use for inference. The cluster design has to be very different. So, if you're going to be focusing a bit more on using language models, you need to be focusing on building out clusters for inference. And that mirrors what we're starting to see in China, too.
26:33 >> Yeah. So, it's interesting. They'll all start differentiating. So, if there's less manufacturing in America, they might use less inference, but if there's a lot of training, then they'll buy more training GPUs. okay. Fantastic. I think the other one was u maybe a little bit of a hint toward the mythos and open AI and it's sort of like marketing has now entered the chat and so does anything change when marketing is involved or you're just still heads down and it's just sort of an overhead buzz. but it seems like when marketing a certain level of maturity has been achieved when marketing decides it's time to get involved.
27:15 >> That's right. We're in this very interesting period because there's been so much capital that's gone into developing these models that we need to figure out how we're going to get returns. And so marketing can't help but enter the chat. since we're in an industry that is traditionally a cost center, it also means that we have slightly different dynamics than what we'd see out of like the regular SAS companies. because the the way in which we generate value is by reducing risk, showing that you can get more from less and a lot of that also means that we tend to rely a lot on unfortunately fear-based marketing. I do think that there's a lot of concern over how fast things are moving and it's something that we definitely need to be paying attention to. But I think that there's also just a lot more opportunity to be focusing on building and patching and using this energy and this momentum to do a lot of the things that we probably should have just been doing the f, you know, the first time around. But because the economic incentives haven't really been there, this is our opportunity to take advantage of this and fix a lot of this stuff that we should have done the first time.
28:11 >> Okay. So, so I I lied. I have one last question because you just mentioned economics. so one of the tenants of economics is that if something gets cheaper, you do more of it. And so I know there's some been some 10x improvements in efficiencies. And after each 10x improvement, it's not that we are clapping and saying, "Oh, 10 times less energy consumption." It's like, "No, we're going to build more. We're going to build more because now you can do it everywhere. Instead of it being expensive and only IBM can do it, now it's cheaper and anybody can do it." So, what's your observations over the last decade as certain things have gotten cheaper?
28:48 >> Yeah. So, there's two concepts that I think are worth exploring here. there's hedonistic adaptation and then there's the red queen race dynamics that we are experiencing. So by hedonistic adaptation I mean every time something interesting comes out in terms of a capability people get very excited about it for a number of months and then over time you know the the excitement tends to fade. People get kind of you know tired of it. They're like GPD3 is not nearly as exciting as GPD 5.5 is.
29:16 Opus is not nearly as exciting as Mythos is. And it's not to say that the capabilities have dropped off. but it's more that we've gotten used to a certain expectation of what we should be seeing and it's going to be very interesting to see how long we can keep that up. I'm pretty optimistic, but I'm very interested to see how these dynamics play out. I also think that the red queen race dynamics are pretty important here too. because as we find more vulnerabilities and we patch them faster, then the expectation might be that suddenly this is the new normal.
29:44 Like we're just finding tons and tons of stuff much faster. So what I'm alluding to when I was mentioning fuzzing is I mean I'm sure a lot of us have friends in the room that u really like running fuzzers they love finding tons of bugs but just because you're running a fuzzer all the time doesn't mean that we have more secure software naturally. There's still a lot of the human process of figuring out like what are the bugs that matter and honestly a lot of the bugs that tend to cause the huge issues tend to be the ones that are really not that sexy. They're ones that people forgot. And one nice thing about LLMs is they're very good at finding just a lot of the very basic things, the basic checks we probably should have been running the entire time. so I think that's going to change the game a little bit, but I think we're going to find a lot of the same dynamics. Like the more things change, the more they kind of stay the same.
30:26 >> So, so I think earlier when we were talking you were saying there's a 10x improvement in some efficiency and then there was another 10x improvement. So there's 100 times more efficiency. Is there another is this just something that's going to occur sort of like when CPUs got smaller smaller the features got smaller and then at some point okay they're about as small as they can get so now let's compete on power per watt or cycles per watt or efficiency is is there a trajectory and then all of a sudden the game changes and like right now it's raw speed but in the future it's going to be energy >> how do you think the >> future of Yeah >> it's a very it's a very interesting and good question that I don't have a very good answer to Yeah. I think I tend to think about this in terms of Moore's law. Mo's law continues to hold.
31:12 scaling laws continue to hold, but at some point we're kind of all kind of holding on wondering why it's still working. I don't have a very good answer for this one. >> Yeah. I didn't know if there was an equivalent of Moore's law for like training or >> That's the scaling laws. >> Yeah. >> Yeah. Like right now it it it what I mentioned and go back to the graph. Like the most interesting thing for why we're seeing the sort of behaviors that we have now is that super linearity that I mentioned where the more data the more compute that you have. let me go back and find it first before I keep talking.
31:47 Yeah, >> here we are. Okay. So when you train >> Oh yes. >> Okay, there it is. >> Okay, cool. So when you train a model that is twice as big for twice as long and twice as much data, the thinking used to be that you'd get a model that was twice as capable, but today it seems like it's probably closer to four times as capable, which is crazy. So we're starting to see this insane takeoff.
32:12 granted, we only have like two data points. We have Mythos and the GPD 5.5 that came out and we have the previous generation of Opus and I'm forgetting the number for the GPT model, but it's current models that are like as of this month with the previous models that came out not too long ago. And so right now it's still singing like we're starting to see a bit of a takeoff, which is exciting. It's terrifying. but that that's >> Any idea how much does it cost to do twice as much or twice as long? Is it four times as expensive? Is it 16 times as much compute? Is it >> these are all questions that I don't think any of us can answer because of NDA.
32:50 >> Oh, so people know the answer. They just can't say it out loud. Okay. >> Yeah. But I mean it's also what makes like these economic questions so interesting because there is so much capital going into this kind of stuff. and as long as the music keeps playing, we're going to continue finding like some crazy stuff like this. But the music needs to keep playing in order for us to find this kind of stuff. So there's a lot of incentives from the scientists, from the venture capitalists, from like the equity holders in order to continue doing this in part because people just want to see how far this goes. But the the cost of doing that is also starting to be kind of expensive too, >> right? Okay. So let's open up to audience questions. I'll start on the left because yesterday I started on the right.
33:31 >> Thank you very much. Good morning, Eric. this is Andy Chao from FSIC. thank you for the insightful keynote. So, here's my question. In your assessment, who is currently ahead of the autonomous offensive curve? Is it the commercial defenders or is it the financially motivated and state nexus track actors? Thank you. >> Hold on. So, let me make sure that I understand your question. So, you're asking who is ahead, attackers or defenders right now? the commercial side or the financially motivated state nexus side or it can be >> not state not state actors commercial offense commercial defense versus commercially motivated criminal groups.
34:19 >> that's a really good question. I I think the dynamics of open source make this very interesting as far as questions go because I think I mean all all of the things that I've seen in the market, you know, talking to bug bounty hunters, talking to various different types of people that use language models to improve their workflows. I think we're kind of neck and neck at the moment. I I still think that the commercial folks are focused a bit more on like the practical amount of what you can squeeze out of these models and there's a lot of like the economic incentives in order to find as much as possible for your customers. but I think that these criminal groups are starting to catch up pretty quickly, too, because they also have a lot of these incentives. it's it's kind of a non-answer, but I think we're kind of neck and neck, which is very interesting.
35:01 >> Thank you, >> gentleman over here. >> good morning, Ari. thanks for your keynote. I'm Hammond from the University of New South Wales. training data is something that's obviously going to be a big important thing as we need twice as much data for twice larger models, etc. At some point, will that data not run out or are we just getting a lot smarter at using it? >> Yeah, that's a great question. So, this last generation of models has used so much data, but there's also a lot of synthetic data that's starting to end up. There's a lot of problems with synthetic data because sometimes you get learning signals that are not nearly as rich. but there's still a lot that you can get out of it. There's also a lot of focus on alternative methods to see if you can really squeeze the models intelligently. So the if you've heard of thinking machines, it's another like ex OpenAI company that came out.
35:53 a lot of their methods focus on this intelligent way of squeezing as much possible through post-training techniques in order to get more of this learning signal into like the larger pre-trained models. And I think we're going to start seeing a lot more of that where there's if you are looking for particular types of capabilities, you need to figure out like first off, what is the learning signal that you're looking for and where do you get data that represents that. and if you can't get that easily or at a volume that makes sense, then you need to figure out synthetically how do you generate that while also mitigating the downsides of kind of synthetic data that you would the sort of things that come out of using synthetic data.
36:28 >> Thank you. >> Hello. Thank you for sharing that information today and kind of wanted to go back to your your roots of being a student and I work with a lot of students on a regular basis. So, I'm curious, having heard all this about AI and with everything moving as quickly as it is over these last few months, do you have any words of wisdom or suggestions for students, people who are aspiring to go into computer science or cyber security, but they don't know what to do with everything moving so quickly. So, if you have any advice for those students, something I could pass along to them, I would really appreciate it.
37:08 >> Yeah. Oh, that's a question of my own heart. I love that. I think we're in very interesting times right now. I I can't imagine what it's like to be a student right now, but it's both got to be very scary from a job market perspective, but in terms of like being able to learn whatever you want and accelerate the amount of things that you can learn in a very short time, it's kind of a golden renaissance for a lot of people. So, I think in order to be successful, I think a lot of it comes down to figuring out what you're very passionate about, find a mentor in that space, and then leverage all the AI tools that you have available to you in order to learn that and really grow in a direction that you're really excited about because passion will take you very far but preparation will take you even further and having mentorship is really what has served me well and I think will continue to serve people well in this new era that we're in.
37:58 >> Thank you so much. >> hello data engineering student from Australia based in Singapore. I think a lot of the discourse around whenever new frontier models come out is whether the previous models get worse and I think we've heard a bit of that with Claude. I was wondering with your experience having seen under the hood, what does that resource allocation look like? >> Oh, that's that's a good question. So part of why that happens comes down to like just the difficulties of shipping a large language model and how you deploy it across a bunch of different GPUs. So it's not like an intentional now that we have this big thing we can just forget about this one that we have. It's more like how do we serve customers the next best thing with also handling the fact that there is not an infinite number of GPUs out there. There's not an infinite amount of compute or other types of resources. So it is, you know, a known effect in the industry working on the inside. It's not one that people are particularly happy that exists and people are always trying to figure out how to mitigate it. But it is one of the growing pains of building fantastic new technology.
39:09 >> Thank you. >> You have to tell them your story about when you're at OpenAI and you're running the query and then you looked at the billing. >> Oh yeah. Yeah. So working at a large company like this at least when I was there so I was there between 2019 and 2022 2023ish so the economics of doing this was a little bit different but at the time like we were working with Microsoft and I had run I taken off this giant training run it was very expensive it turns out I messed it up and so I look in my console to see how much it costs because I was like oh man I got to go to my manager and figure out what I'm going to do about this. And I kid you not, like it was about a $5 million job that I had just absolutely botched. And I was like, "Oh my god, I'm going to absolutely get fired." but because research was at such a premium and I was on a research team, they kind of looked at it and like, "Well, I guess we're just going to write that one off." like the the capital cost that goes into making these large language models is is quite high. but it's it's an R&D cost for a lot of these places because the expectation is that you're going to mess up a few times, but when you really hit it big, you're going to really hit it big in terms of capabilities that you can squeeze out.
40:21 >> So, like you you welcome to the $5 million screw-up club and you know, and all the other researchers, hey, I got six million, you know. >> Yeah. >> Sorry. >> Okay. >> Oh, do you have a follow-up question? >> Yeah, I have a question for you. I am Tao. I'm from Vietnam. I am a researcher for automotive securityities and I see your talk mention about attacks running continuously for connectetic automotive for how do you we mitigate the risk of resource excession from the s gasway when running this AI driving continuous simulation cycle >> something I'm not sure I fully got the question, but in autonomous vehicle resource exhaustion when an AI simulation is running continuously.
41:19 >> Yes, >> I don't think I quite understand the question. >> Okay, I can respect. Yeah. >> Okay. If you want to find me after, I'm happy to talk. >> Yes. My question about how we mitigate the risk for the resource assessment for the SS ways because in the automotives we like in information IoT for we running this AI as a guest way for how we running this AI revenuous simulation cycle That means when we run manage AI simulation cycle in the S guest way from like IoT. So how how you mitigate the maybe to the maybe so like >> maybe afterward come on up.
42:18 >> Yeah. Yeah. Yeah. >> Yeah. It's a little hard to understand the the nature of question. Okay. >> Yeah. I think it might just be like the audio. >> Yeah. Yeah. can't tell. And no more questions. We have time for questions. Okay. while we're waiting for the next question, I have heard that that mythos and so on destroys capture the flag contests. >> Yeah. >> because they're so well studied, they're sort of bounded problems. And then other people have said that it doesn't perform so well in debugging situations, debug problems. And so people are speculating, well CTFs might move toward like debugging challenges.
42:58 >> And so I'm curious, maybe go back. You talked about the certain types of problems it's not so good at. So if it's good at attack and it's not good at debugging, why? And are there other classes of problems that you're like, "No, that's a tomorrow problem. Today we're doing this thing." >> Yeah. Okay. Me find it on. I'm just going to use this as notes. Let's see. the other section of the talk.
43:32 Okay, so going back to the OSS fuzz evaluation that was included in Anthropics Mythos report. I think this this tells the story I think best because where we see the most massive gains for a model like Mythos, which is arguably at the edge, is among these low severity shallow bugs, which you could also kind of describe a lot of CTF bugs that way because they are, as you say, like pretty well like articulated. they're they're not very natural.
43:59 we're seeing modest gains at these mid-tier bugs, and still sparse gains at like these deeper exploitations. It's not to say that we're not seeing, you know, interesting stuff happening at the deeper exploitation level. but a lot of it is kind of like shooting a shotgun in the dark and then figuring out how to kind of like roughly stitch things together. And we're getting pretty lucky on some of that. over time like models are going to be getting better at this, but right now where we're seeing the most success is in a lot of these like shallow basic bugs that there exists a lot of documentation for and are pretty easy to validate. I think the other thing is models tend to struggle with anything that requires holding state. So things like concern concurrency. So understanding kind of what is happening across different machines and in different levels of memory. They really struggle with that because fundamentally language models don't have a model of of state being held. And the issues that we have with language models around hallucinations are things that will still plague regardless of whether or not you have a model of like whether or not you have like a larger context window or not.
45:03 Like we're seeing improvements here because context windows have really exploded in the last three years. but these sort of issues with the complexity around temporal state analysis and other bugs like that still persist and I anticipate that that will probably continue to be true for as long as we still use transformer-based LLM architectures. >> So I'm curious then maybe to improve debugging does that mean if you're developer you want to make your crash dumps or your your error logs as detailed as possible? Yes. is the more detailed error generation the more of these models and so there'll be like a competitive advantage if your tooling generates better errors than the next guy right or am I thinking about it >> yeah absolutely and even like with mythos and these other large models we find we get better performance if you know what you're doing I think this is going to be continuing to be true is as like a theme across this particular generation of AI that we have in front of us is that if you know what you're doing you know what kind of information you need in order to make a decision and that's the kind information you can provide to a model and instrument it. So if so maybe there's hope for us instead of getting error 17254 popup we'll actually get more meaningful contextual errors because these companies will want to train but as long as it's for end users maybe they're not so interested. So >> yeah the e economic incentives have to be at play here for this to happen.
46:31 >> Yeah. Okay. >> Hello Rajesir from India. I work at day force. so the next big thing that I see is the quantum computing right. So here is my question. So do you foresee quantum computing significantly disrupting the AI landscape especially in terms of in structure and security? >> no and I'll tell you why. So I think of quantum as being orthogonal to AI.
47:04 like the the things that quantum are very good at are things that I think mostly help cryptography or break cryptography. But the type of actions that you do in a computer in order to do AI based tasks are not the same. Like what underlies a lot of the AI capabilities is just being able to multiply matrices together. And with quantum it's actually pretty expensive to multiply matrices. And so you're not really going to see the same gains if quantum continues to >> Is that vector mass? It's just >> Exactly. It's a vector thing.
47:34 >> So, and and quantum is not so good at vector. >> Not not in the ways that it would need to be. Yeah. I mean, think about it in like there's two if you ever had like a classical mathematics education, there's like two branches that they put you into. Either you go down like the algebra path or you go down like the real analysis calculus type of path. and if you're going down algebra, that's a mindset that's much more helpful for solving quantum related problems in cryptography. And if you go down the calculus or the analysis based path, that tends to be a bit more helpful for solving AI problems.
48:07 >> Thank you. You discussed how these models struggle with concurrency and temporal analysis bug hunting. Is there any evidence that the code that they generate is more prone to those types of errors? >> Well, that's a good question. I don't have an immediate answer for you, so let me let me reason through it. I haven't I haven't seen that. I have seen a lot more hallucinations, but as for like true errors, a lot of this is like shooting a gun in the dark and hoping that you hit something.
48:51 yeah, I'll admit that's a great question. I don't have a good answer for you. Thank you. >> Thank you. >> As the capabilities for models in offensive spaces continue to improve, we've seen with project Glass Swing that they've restricted access to models to a select set of companies and they may not release it publicly and they're working on improving their alignment and guard rails. Do you think the alignment and guardrails will continue to improve in a way that allow them to continue public or other companies to get access or do you think they'll continue restricting access as models improve and will possibly governments get involved?
49:29 >> That's a good question. So what we're seeing right now with the terms restriction rhymes with what happened with GPD2. There's a paper that I wrote with my colleagues that are now at Anthropic and some that are still at OpenAI about what we called stage release. And so the idea is that if you have a capability that you're not quite sure how it's going to hit when it goes public, you want to be very thoughtful about how you do the deployment. And so we built up, you know, like a trust and safety team, monitoring, all those sort of things.
49:56 And I built like the first versions of and it was very interesting to see like a lot of the same arguments that we made back then sort of come back with mythos. I think we're in like a very different regime than we were in 2019. And candidly I feel like Anthropic needs to maybe think a little bit differently about how they handle like this stage release because I I don't know if it's necessarily as helpful as it would be if if we were able to give more access to more people because the other thing with a large large language model is that it just can't you can't serve as many customers that way because it is so expensive to serve a very large like a much larger model that we've we've seen before. And so naturally, you kind of end up restricting use anyway, which gives you a little bit more time to, you know, go through and and make sure that your monitoring systems and some of your mitigations are up to snuff. So I think that there's definitely a lot more incentive for governments to step in and figure out how to be helpful. but I think that naturally just because of how it is to deploy the technology, I'm not confident that like artificially gating it will be like any meaningfully different versus just naturally having like a waiting list because it's very hard to serve these things.
51:06 >> Thank you. >> So you maybe hitting on that a little bit, you mentioned a little bit earlier about how mythos you can achieve sort of mythos level if you have a higher skill level and you use open source and you're maybe just a better team. you can get comparable results. So maybe just talk a little bit about how much is open source lagging or are there things open source is doing differently? They exploring different spaces that don't have the immediate economic, you know, returns or how do you think of the two frontier models versus the open source community?
51:40 >> Yeah, that's a great question. So open source is going to lag in terms of like reasoning capabilities because it's just it's very expensive and doesn't make a lot of sense to release something open source that was very expensive for you to train acquired a lot of data that was very expensive for you to gather acquire a lot of GPUs that you could otherwise be using to train other things that give you a competitive advantage in in AI race. So I we're going to continue seeing these with meaningful lagging time. That said, if you are an expert, that doesn't really matter. you can get a lot out of the benefit of like a mythos level system if you understand what you're doing around instrumenting a whole bunch of these smaller systems. and I think that meaningfully this means that like there will be a bit more of a focus in the open source community around like expert-led discussions over expert type instrumentation versus like a more consumer-based focus on a lot more of the like classical foundation model lab capabilities that also come out >> because I'm curious at some point to to the previous gen one's question if somebody says well it's too powerful we can't release it and some developer in Crackkow just says, "Here's my open source version of it." Right? At some point, if things can continue to improve on both sides, the open source, yes, will be lagging, but it'll still probably be pretty powerful in a couple of years, right? And I mean, I don't know if there's going to be the open the internet archive for open source software data training, so not everybody has to train or there'll be sort of like I don't know community efforts to pull national data. so that's I was curious like where is this?
53:20 >> Yeah. Yeah, that's I think it's worth kind of digging into it a little bit more. So, interesting dynamics come out from things like OpenClaw. Like I don't know if people remember like maybe AGI or some of the other harness sort of technologies that people just got similarly excited about. these are all kind of solving like the same problem of figuring out more intelligent ways of instrumenting open source models. So you can try to get the same performance out of those as you would out of like a larger language model that you would get from like a foundation model lab. and I think economically it's it's kind of impossible for you to get that same level of performance exactly, but you can get an approximation. And I think what we're going to see a bit more is this focus on like outputs. Like what are you trying to get out of a language model? And if you can get what you need out of it, then people are going to naturally figure out ways to instrument them much more intelligently. And then you're going to end up with like these more interesting open source communities around things like open claw and whatnot.
54:18 >> Hello, I am Ranzan from Nepal. I have a question. As offensive security evolves into autonomous agents, the speed and scale of initial access attacks like AIdriven spear fishing with inevitably increase. How will the sec ops team adapt the defense posture to catch this agentdriven campaign early in the kill chain? >> How will the sec ops team adapt to >> catch these all early accessdriven campaigns in the kill chain?
54:54 >> Initial access app in the killchain. You mean being aenically or or AI generated and then the devs sec team is going to have to constantly be always looking for >> so let me repeat the question since the use of AIdriven spear fishing attacks are increasing right now right so how will sakeups team the various secups teams throughout the world adapt their oh >> okay how yeah so how are you suggest Secop team seems evolve or adapt in this environment with them.
55:32 >> I think it's important that people get a bit more proactive about, you know, running these tools. You know, there's a whole bunch of different tools that are now starting to pop up, including my own, around, I'm not here to sell. but being able to find problems before, you know, the bad guys do. And I think being more thoughtful about how to shift left within your own organization is really how you win here because it's going to move no matter if you move. so you should probably start looking for things and figure out how to build this into your cycle much faster.
56:00 I I would think about this as like an opportunity for you to or for anybody to really start thinking about like how do you leverage the fact that this is a moment in time where people are finally looking at security and thinking well maybe we should give it a little bit of bigger of a budget. Maybe we should give it a bit more attention and care. so let's take advantage of this particular moment and figure out how we can then encourage people to to find vulnerabilities faster, patch them faster, and start really building up these processes that we really should have been doing in the first place, but the economic incentives were not necessarily aligned with us being able to do that.
56:33 >> Thank you. >> All right, we have time for one more question. Final question. oh. Is that you, Maria? >> >> Okay, go for it. >> Maybe it's not your direct specialization. It's so loud. so for example, I do cyber physical security. It means that I destroy or degrade physical objects by the mean of cyber attacks. And many not many people do it because it allows a huge amount of interdisiplinary knowledge. like you really need to be able to read a lot of engineering documentation books and you have to consume a large amount of information and correlate it correctly.
57:18 So how realistic and because this type of attacks is now on the rise amongst state sponsors threat actors at least from an desire perspective intention or motivation. So how realistically to teach AI bots to become a really very intelligent creative engineer in multidisiplinary domain. >> Okay. Did you get that? >> Cyber threat I mean cyber physical interaction OT and OT systems have a lot of engineering and interconnections and reliance on many other parts.
57:56 Is that further down the chain for AI attacks because it's it's not the same as reading source code and iterating? how realistic is it then for AI to get good at cyber physical? >> That's a great question. Yes. it is further down the road because it's just straight up not in the training data set and nobody at these foundation model labs is particularly aware of how to put it in because it's complex data to collect. Yes. the payoff is not necessarily high for a company that needs to get the returns that you would need in order to you know continue >> playing the music as as you're moving more towards like AGI type systems. So this is an area where I think that like open source and that expertbased community could really like shine in ways that you're not going to see out of like the traditional foundation model lab type approaches. I actually see this can become a large market because this is the type of let's say AI automation or trained bots which state sponsors red actors would love to buy.
58:55 >> Oh, absolutely. Yes, I'm I'm right there with you. We've been thinking about that too. >> Excellent. Thank you. >> All right. So, that wraps up our keynote session. I just have a couple of announcements and I shall release you into the vortex of black hat. Jeff Ari, thank you very much.
Summary
- Black Hat fosters a community that spans all career stages, from learning to leadership.
- The keynote highlighted the impact of AI on cybersecurity, particularly in offensive tactics and vulnerability discovery.
- AI models are improving at generating exploits, with the speed of vulnerability discovery increasing significantly.
- The scaling hypothesis suggests that larger models yield disproportionately better performance, leading to faster exploit development.
- While AI excels at finding shallow vulnerabilities, it struggles with complex issues requiring state retention, such as concurrency.
- The security landscape is changing rapidly, necessitating proactive measures from security operations teams to adapt to AI-driven threats.
- Open-source models are lagging behind commercial models but can still achieve competitive results with expert guidance.
- The economic incentives in AI development are driving both innovation and the need for effective security practices.
Questions Answered
What is the purpose of Black Hat Asia?
Black Hat Asia serves as a platform for sharing knowledge in the security community, showcasing the journey from learning to leading in cybersecurity. It emphasizes the importance of community engagement and the evolution of security practices.
What are the key points regarding AI's role in finding vulnerabilities?
Ari Herbert Voss discusses the dual nature of AI's capabilities in security, highlighting both the real advancements and the aspirational claims. He aims to clarify the fundamentals of AI's impact on security and what it means for the industry.
How will offensive capabilities evolve in the coming months?
Offensive capabilities are expected to scale, but effective deployment will depend on ergonomics and building practices. The speaker emphasizes that the field is still developing and that initial concerns should not distract from ongoing progress.
What are the challenges associated with using synthetic data in AI models?
While synthetic data can provide learning signals, it often lacks richness compared to real data. The speaker discusses the need for intelligent methods to maximize learning from both real and synthetic data while mitigating the downsides of synthetic data.
What advice does the speaker have for students interested in cybersecurity?
The speaker reflects on their own journey and emphasizes the importance of adaptability and continuous learning in the rapidly evolving field of AI and cybersecurity. They encourage students to stay informed and seek out opportunities for growth.