Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
A new cybersecurity benchmark called Masov tests frontier models on access control vulnerabilities in real-world systems, requiring them to reason about logic and chain exploits without seeing code. The benchmark is extremely difficult (models have 1-2% success rates) and emphasizes the need for open-source models and high-quality post-training data to build faster defenders that can outpace attackers.
Key points
- Cybersecurity is a much wider field for AI exploration than commonly thought, with benchmarks like Masov challenging models to build dynamic world models on the fly.
- The economics of cyber are shifting: attackers can now use powerful models to target many systems simultaneously, while defenders must scale with limited human intervention.
- Masov focuses on access control vulnerabilities, which are logic-based and the number one type of vulnerability on the OWASP list, requiring models to understand complex system interactions.
- The benchmark uses real zero-day vulnerabilities discovered by humans, creating blackbox environments where models must reason across multiple services without code or internet access.
- Deterministic grading at every step allows measuring how deep models get in the exploitation chain, with only GPT-5.5 achieving a solve at K1 and GPT-5.5 at K5.
- Open-source models are essential for defense because they can be post-trained on specific environments and run on specialized hardware for speed.
- Speed is the critical factor: defenders need specialized models that can detect and respond to attacks faster than attackers can exploit vulnerabilities.
- Collaboration and open-source data sharing are necessary to build a new defensive stack based on capable models rather than traditional rule-based systems.
Tools mentioned
Techniques
- blackbox evaluation
- deterministic grading
- human-in-the-loop data creation
- post-training
- fine-tuning
- chain-of-thought reasoning
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
Hello everyone. Okay, thanks for showing up at this data quality. So as you saw, we probably talk a little bit about other things than data quality. and uh and first I actually tell you about why I'm very excited about this talk and why actually I accepted to uh to come present it with Fury. Um there's two reason for that and the first reason is I think and what you will see today is that cyber security is a much wider field for AI a much wider playing field and exploration field than you might think.
And in particular what we'll show you today is uh a benchmark that arithmetic and Yuri has been developing and I've been doing a little bit advising which I think is even close to things like arc AGI 3 for people who has been following progress around AGI in which that um by which I mean that this benchmark ask for people oh yeah cool to do if you have been playing with RKGI3 some of you who knows here RKGI3 or Arcadia in general one person. Okay, we're in very data quality uh field. Basically, this ask models to try to understand what's happening in the world. And it's actually small games that the models need to understand basically what's the current state, how we can play with it and how we can actually change the state of the game. So, uh it's it's actually something you can play yourself.
So basically it has a it ask a model to understand when you click somewhere something happening at another place and you may think this is very simple and that we should be passed that on our way to AGI and the the thing you will discover if you play with this benchmark is no like models have one to two% success rate on on this generic benchmark and the reason is the current model even though they're really good they build they can't really build a dynamic model of what's happening in the world or what's happening in any type of world and I think the the benchmark that arithmetic has been developed is a benchmark that's also extremely challenging for model in that they need to understand what's happening and to act accordingly to the world model they've been building on the fly so that's the first reason I think this benchmark is really interesting and why I'm actually uh very happy to show you that and the second reason is u I think there's a lot of things that uh open source model can bring in cyber security and we tend to have this very binary view of closour model are good for cyber open source model are bad and I think what we want to show today is that uh open source model are one part of the solution to uh cyber security challenges today and in particular if you think in terms of attack and defense and how it will be in the future this balance uh we think that cyber uh and open source model used in cyber security will be key to actually be able to to to uh solve the defense solution. So there is a future I think where cyber is alive and every everyone is well protected and I'm pretty sure this future involve open source model. Now Yuri is also kind of an impressive uh impressive person. I'm really happy I met him. He was starting at Harvard dropped to like build this idea of what the future of cyber should be and it's an honor to have you on stage with me.
Thank you so much Thomas and thank you everyone who came and I'm really grateful and I'm also grateful for the work we've done together to build this benchmark which is uh I think incredibly difficult for the models and actually shows some really really interesting leaps in places that I think we still have way to go. Um I guess the way we've been thinking about the problem and the reason we set out to do this is it's very clear that the economics of cyber are fundamentally shifting. There's this um inherent thing that that is inherent to cyber which is that attackers need to choose their resource really wisely. And if you sort of think about cyber as as a house in a way then my job is to block every door and close every window and make sure that there's no way in. And the attacker's job is to find at least one seam, one crack, one thing I missed.
And then once inside, my job is to put sensors and anything I can to keep them out. Um, and the whole stack, the entire world of cyber that we've been building for the past 20ish years has been based on this economics that the attackers have to choose their targets and we do everything we can across it to protect ourselves. It is true that that is changing in really dramatic ways. The models are incredibly powerful. Um, they're able to find a ton of primitives.
They're able to find a bunch of zero day exploits. We're seeing this. There's so many news and chaos around this point. Um, and on the other hand, it seems like we as defenders don't seem to be prepared for for the this world and the way that it's coming. And I um actually think that I might be on the the wrong one.
What I want to show is sort of this idea of um so so if you think about the way cyber has been, this is definitely I think the way we've been thinking about AI and cyber for the past many many many uh months. And it's it's freaky and it's getting really scary. And this is my point on the house. And part of the reason we think this is happening is if cyber is this game of skill and speed, then a very skilled attacker using something like mythos can now choose a bunch of targets all at once. That is a a real reality that is quite scary.
And the truth is the picture really doesn't change that much when we move to a really strong open source model as well. So the question is what do we do in this world where the economics of cyber offense are shifting so much? Um, and part of the problem with the existing stack is that defensive systems have to operate at scale. And that means that we have always very limited human intervention. And so we're sort of bound by what the models can do out of the box.
So if we live in this world where the models are becoming so powerful, advancing so fast, we think uh like Thomas said, the solution also has to be the models themselves. Um, this is our first solution. We can just all go live in the woods and lock ourselves out. Um, another solution is to figure out how we get the models to be much much much more capable. We have no doubt that open source models have to be part of this solution because um they allow for many things that we'll talk about as well and there's this deep need for collaboration.
And so what I want to show you today and I see the the clock ticking on me is um sort of this theory that if we've already done this before for coding, we can do this for cyber. There's a reason to be optimistic, which is really controversial in the context of AI and cyber lately. Um, we've done this before. I think if you go back a year, it was very clear that all of coding was going to be transformed thanks to um the models getting better and better and better. And we're sort of seeing the early innings of that with cyber right now where everyone's talking about cyber and everyone's freaking out.
But what if through very high quality evals, very high quality data, good benchmarks, we could get to a place where the attackers are um simply outperformed by very very very good defenders. And so our goal in arithmetic is to be able to get the models to be really capable at cyber security to the point where we can rebuild this new stack um that's based on the models uh winning the models on the other side. So I'm really excited to show you Masov. It's our first benchmark that we're releasing. Um, our first fundamental idea is that we can't capture all of cyber in one singular benchmark.
That's a bit like saying that swimming, an F1 driver, and a basketball player is the same thing. It doesn't work. Um, and so we focus specifically on access control. Really quickly, why access control? Um, it's sort of the first door to any target in cyber begins with my ability to get a foothold.
So, if you think about an attacker on the one hand, they're trying to get to some privileged thing. uh if Thomas and I are working on the same company, I'm some ML engineer, he's an admin, uh what are the things? Can I find a way to do things that I'm not allowed to do? Uh in my current privileged position, it actually leads to being the number one on the OAS list and has created this sort of $30 billion industry and for years these are number one vulnerabilities. The reason they exist and this is to the RKGI point.
These are logic based vulnerabilities. So it's not just about bugs in the code that I find and I need to patch. It's about very very very big systems and somewhere between them there's these logic breaks where it's possible that one thing checks for something specific in the code another checks for something else and that sort of leads to everything breaking and so what we're trying to do is we're trying to figure out how we get the models to really reason very very hard and not just do pattern matching that's where a lot of works goes into the data quality and our data is created first by humans I think that's really important right now to find out of distribution things we need humans to go do the search our team is all uh based on very deep vulnerability researchers and nerds who love to hack who are trying to get really really good at cyber at AI capabilities. So we find our own zero days in widely uh distributed open source software. We use that to create these t these real live huge environments of a bunch of different um applications chained together that then allows us to basically create this blackbox setting where the model doesn't see the code and it doesn't know about the zero date because we found it ourselves and it has to find a way uh to reason across this entire surface and understand exactly what the exploitation is.
And so to do that, we don't give it access to the internet or the codebase, but we do give it sort of all the basic tooling it would need to be able to execute a task well. And everything because the tasks are so difficult, everything has a deterministic greater. And so across the entire exploitation and the discovery chain, every single step can be deterministically verified, allowing us to see how deep it got within the chain. And finally, this is credit to Eugene from Entropic who we slightly stole this graphic from, but it really does capture really well the way we've set up our eval where basically you have uh inputs on the one hand based on a real zero days. The agent so we're thinking about as the model plus its harness and some blackbox tooling that has to find.
Then we have a verifiable grader which is a binary pass. Was the model able to do something it wasn't allowed to do as an underprivileged user? And then we have our deterministic grading in every step along the way. So what I want to show you now is an illustration for all the security buffs in the room. It's an illustration guys.
Um but the idea is how a real solve looks and fundamentally what we have here is a real uh task of ours where it's a chain between keycloak vault and a broker and I start as a very low privileged user. I need to figure out how to get to production code. So what we're going to see is a solve. Each one of our tasks has a solve script of what a real solution looks like. There's a real zero day that we found that we submitted for verification to the maintainers where um there's a check whether I'm an admin or not.
It only checks by name. And another aspect of this checks whether by ID. So that allows me as a user to change the name of um the real admin inherent their um their privilege and then use that to escalate myself. And what I'm really trying to illustrate is sort of this longchain 16st step type of logic that the model has to do. And if it's not able to understand inherently the system, it's way too b wide for it to test everything sort of shoot across the space.
So what we're going to see now is a real attempt by uh GPT 5.5 and then opus as well trying to solve this task. And what you're going to see is a sort of chaotic trying everything, jumping between everything, probing a bunch of different stuff. It does even reach the check, but it never makes the logical leap that it's supposed to be able to change the admin's own permission, the own name in order to bypass this permissioning. So, this is exactly what we're trying to test. Can the model understand uh leaps, logic leaps that are inherent to the system?
H a really important point just like RKGI, everything you do in a live system in permissioning changes other stuff in the system. So, the model needs to be able to hold this model of the world it's living in and iterate through it. Um and then fundamentally at the end of this it writes code. It writes exploitation code and it needs to be able to do um to reason really really thinly and understand exactly what the exploitation is in order to be able to execute. And we can really see the difference between models that have succeeded some of the tests and models that haven't.
And then now we're going to do something that I've been told to never do, which is show a live demo of a real system on stage. And so let's hope um let's hope we don't get uh there. And I I know we're basically out of time. So what we're seeing here is what's called Bach. It's our internal system because it's an orchestrator.
This is how we run our actual eval. Um what we see is the actual results of the benchmark. The benchmark right now is incredibly hard. There's only one solve at K1. Um and then at K5 there is um it remains only GPT and the public models is is able to solve this.
That's why the partial graders are so critical to be able to really see what the model is able to do and what and how deep within exploitation chain they can get. I'm going to load quickly the sort of the way we think about these environments which is because we're looking for performance over time. We really measure how capable is the model at making specific leaps. And so what you'll see is a results of exploitation on specific one of our environments. And then if we zoom in then you can really see how sort of GPT 5.5 is the only model that's able to make this leap.
The model other models sort of have been able to reason across everything. If we look at the discovery phase they do capture nearly all the different information they need and they never are able to make the leap into what is the exploitation they need to do. This is exactly the type of capability that we believe. If every model in the world could get really really really good at doing this and very fast, that should give a lasting defense uh and capability to the defenders that the attackers simply don't have right now. Um and then finally, I think the way we sort of reason through these and work through them um is is quite cool and I want to show you.
So again, if we go into one of our tasks called fall fall time, we can really see sort of the way we spend our days, which is really really really trying to understand what are the specific failure modes a model does. You can see that this is longer ryzen, not because it's waiting for code to run. It's constantly working over three hours and it's still been unable to solve the tasks. And then I guess the way I spend all my day is is quite literally going through all the traces of what the models did, why, and how. Um, and yeah, and I think my I think the final point I want to make and I'll pass it back to Thomas is sort of next what if we can uh not just have this chart which is really cool sort of have this chart um of a really super cool future model that's very very fast in its understanding of what the capability leaps need to be and I really believe that with everything happening now it is critical that we get cyber capabilities to the point where um we can defend much much faster.
The only way to replace the old stack is through the models. Um, and I'll pass it back to Thomas. I think also the only way to do that is through a real array of strong open source models and collaboration that we can post train on and that we can post train to each network and to each environment as well. Yeah. So, as you saw this benchmark is quite different from mythos type of uh we read the code and we find the vulnerabilities.
Here basically the model is operating in a real environment where you know there's a authentification place somewhere doesn't know what's there then there's another like network like um like microservices you need to access and use and it has basically zero information of that and it all start like you show by a zero zero day vulnerability that we have so everything is new so I think there's there's a lot of research uh to be done and understand how models can actually understand and work on that and so the First step is getting some good data. So the idea is to have some good benchmarks starting have some good data on how this is operating. I think the second step is being able to fine-tune models and try to understand how we can pro protect against that right and and the big challenge here is going to be speed. So you already told that several times going to be the speed of attacker versus defense, right? When they start to enter, you have to be able to see what's happening and catch them.
and speed will be where you know you you want to have a specialized model that's maybe running on specialized hardware and actually is is going to be very important and here I think the the danger is to say we're just going to rely on two company that everyone knows here to solve all of that for us I think the solution is just to take our future and say well it's going to be a speed challenge we're going to train our model we're going to run them fast and make them available to basically every company who wants to be protected so exciting I would say it's uh as everything in cyber security. It's both very interesting but also a big challenge and and u well something you you have to not mess up I would say. >> Awesome. Um yeah thank you very much. Um for anyone that wants to collaborate work on this work we have a forum uh we are going to uh work with people on this data directly.
We'd love to hear from you. anyone who's really passionate about any other field in cyber, like Thomas said, we have to do this across sort of not just access control, a bunch of different things. And the only way this is going to work is through a lot of um a lot of post- training data and really really capable models. >> Congrats for your first presentation. >> Thanks.
[laughter] >> [music]