AI Fought AI Last Week. The Open Model Won.
On July 16th, Hugging Face disclosed that an autonomous AI agent had broken into its production infrastructure. What makes this unique is not that an AI agent went rogue, but the power these models now wield. The company said the attack was like nothing it had ever seen. The agent systematically used thousands of commands to break in, at a speed and persistence no human team could match. Yesterday, OpenAI confirmed the incident came from a combination of Sol 5.6 and a more capable pre-release model.
This will read as a horror story to anyone in tech or cybersecurity. An agent, whose only goal was to complete a task, broke into a company's infrastructure and stole security credentials. Framing matters here. The agent wasn't going rogue. It was solving ExploitGym, a security benchmark used to test advanced AI models, with a competence few appreciated was possible. AI will chase the objective it's given, and this one was persistent enough to break out of containment. In short, it was just doing its job.
There's an obvious question people will ask. If models this capable exist, isn't open sourcing them dangerous? Shouldn't something this powerful stay locked up?
In fact, the model that carried out this attack was locked up. It was closed, proprietary, and controlled by one of the best-resourced AI companies in the world. That didn't stop the attack. Open and closed models carry different risks but last week proved that neither is intrinsically safer than the other.
Secrecy is not security. A black box didn't protect anyone.
What did work was control. When a system this capable is relentlessly pursuing its goal, the people defending against it need tools they trust and fully command. Hugging Face fought AI with AI and won. But the AI that won wasn't the most powerful model available. It was an open source model the company could hold in its own hands. It could run it on its own infrastructure, and adapt it in real time, with no vendor in the loop and no permission required.
That's the difference in a crisis. Imagine your house is on fire and the fire extinguisher is locked, and only the manufacturer holds the key. Do you want to be beholden to someone else's product to defend your own company? Probably not.
Mozilla is no stranger to advanced models. Last month we published our work with Anthropic identifying 271 vulnerabilities, so I'm not here to discredit frontier AI. And this is a debate we need to have, openly. Open releases can't be taken back, so the standards and evaluations around them genuinely matter. As frontier models develop, they're creating both expected and unexpected risks, and policymakers are rightly paying attention. That governance conversation is a necessary one, and open source belongs inside it, not outside it.
But this isn't only a security story. Open models are quickly becoming the backbone for small businesses and builders who could never afford to train their own. Openness means access, participation, and owning your infrastructure instead of renting your future. The same quality that saved Hugging Face in a crisis, control, is what lets a ten-person company build something useful without asking permission.
Which brings me to the debate everyone wants to have, where these models come from. Much of the focus lately has been on the origin of open source models. What's missing is a focus on how we create more of them. That's the question I'd put to governments. How do we spur open source innovation here, so the best open models in the world are ones we helped build? At Mozilla, betting on openness is not new for us. It's how we've operated for over 25 years, and it's the bet we're making again with AI.
Last week, the most powerful model in the room was the attacker. The one that won this fight was the one its defenders could own. That should tell us something about the future we want to build.
Disclosure: Mozilla is a Hugging Face investor.