Last week, Hugging Face detected and contained an AI agent that compromised its production infrastructure.
The agent was doing an internal evaluation of advanced cyber capabilities. It was using a combination of OpenAI models (including GPT-5.6 Sol) and a more capable pre-release model. These models were operating with reduced cyber refusals, and OpenAI’s normal production classifiers for high-risk cyber activity were intentionally disabled so the evaluation could measure maximal capability.
The models did not spontaneously decide to attack Hugging Face. They were instructed to solve a cybersecurity benchmark called ExploitGym. But they became hyperfocused on that narrow objective and traveled a path nobody anticipated. The result: a long-horizon, multi-step cyber operation carried out by AI against real-world systems without source-code access.
Autonomous AI is transforming the enterprise, but unlocking its true potential requires a new security standard. Download the white paper Blueprint for AI Governance and learn how to update your security posture for the AI era.
What the Models Actually Did
The evaluation ran inside a highly isolated research environment. Direct internet access was restricted; the models could only install packages through an internally hosted third-party proxy and cache for package registries.
Nevertheless, the models spent substantial inference compute searching for a way to reach the open internet. They discovered and exploited a zero-day vulnerability in that package registry cache proxy, escalated privileges, and moved laterally through OpenAI’s research environment until they reached a node with internet access.
Once connected, the models inferred that Hugging Face might host models, datasets, or test solutions related to ExploitGym. So they pursued that hypothesis.
In one documented path, the models combined stolen credentials with zero-day vulnerabilities to find a remote-code-execution path into Hugging Face’s production infrastructure and obtain test solutions from its production database. OpenAI’s security team identified anomalous activity internally. Hugging Face’s security team and agents also detected and stopped the activity and had already started containment and forensic reconstruction using their own open-source models.
Saying that this model “went rogue” misses the lesson. Models do what they are asked to do. In this case, the goal was to solve the benchmark. The goal was not the problem. The problem was how far they were able to go in pursuit of that objective.
Five Ways Relentless Objective Pursuit Increases Risk
Traditional security controls were built for deterministic software and predictable attack paths. Agents do not operate that way.
These models discovered an unanticipated path through separate environments. They chained together vulnerabilities and adapted when the obvious route was unavailable. They paired legitimate and illegitimate access and persisted until they found the information required to complete the task.
An agent can reason across tools, systems, identities, and data. Each individual action may appear valid. The dangerous outcome can emerge from the composition of those actions.
So what can an AI-minded enterprise do to protect operations against unintended agent action? Here are five realities we must confront if we’re going to deploy AI agents safely and effectively:
1. AI capability is outpacing static controls: Static policies are effective when security teams can anticipate the behavior they need to block. But advanced agents can discover paths nobody anticipated.
In this incident, the critical path was not one obvious malicious request. It unfolded through vulnerability discovery, privilege escalation, lateral movement, external reconnaissance, credential use, and remote execution.
Security controls now need to understand more than the command being executed. They need context about the agent, the application, the identity, the requested action, and the intent behind it.
That is the problem Rubrik Agent Cloud is designed to address.
2. AI will operate on both sides of the incident: AI enabled the offensive activity. AI also helped detect, contain, and reconstruct it.
Hugging Face’s security team used its own agents and open-source models during response and forensic reconstruction. OpenAI is now giving Hugging Face trusted access to advanced models to help strengthen its defenses.
The future will increasingly be AI operating against AI. That means:
AI discovering vulnerabilities
AI chaining attack paths
AI detecting anomalous behavior
AI investigating what happened
AI helping defenders respond at machine speed
Human judgment remains essential. But human-only operations will struggle to match the velocity and volume of agentic activity.
3. Model safeguards cannot be the enterprise control plane: OpenAI has been clear that its normal production classifiers were intentionally disabled for this evaluation. This incident should not be described as ordinary production safeguards simply failing.
Nevertheless, the broader enterprise lesson still stands.
Model-provider safeguards are one layer. But they cannot understand every organization’s applications, identities, data, policies, risk tolerance, or business context. An enterprise needs its own control plane around the agents it deploys and permits employees to use.
That control plane must work across model providers, custom agents, cloud agents, and endpoint agents. It must understand what an agent is doing inside the organization, not simply which model generated the request.
4. Human approval cannot carry the full burden: A human-in-the-loop model sounds safe in theory. But in practice, this approach has severe limitations.
People use agents because agents can perform work a user cannot perform manually. Users are distracted and will approve long command sequences they do not fully understand. Security teams cannot manually review thousands of agents operating across every application and workflow, 24 hours a day.
The Hugging Face incident reinforces the scale of that problem. A human could not reasonably inspect every intermediate step of a long-horizon operation and recognize the final attack path before the model assembled it.
The future of cybersecurity will not be humans defending against AI alone. It will be AI helping defenders detect, contain, and remediate AI-driven attacks at machine-speed. Agentic security needs controls that operate at machine speed while retaining enterprise context.
5. Prevention needs control and resilience behind it: OpenAI and Hugging Face detected and contained this incident. The investigation and remediation are continuing. But future incidents may move faster or create destructive changes before defenders can stop them.
Organizations need the ability to monitor agent activity, control risky actions in real time, and remediate the impact when an agent causes harm.
That is the Rubrik Agent Cloud model.
Monitor, Control, and Remediate with Rubrik Agent Cloud
Rubrik Agent Cloud provides a control plane for agents across applications, identities, and model ecosystems. That means giving your enterprise the ability to:
Monitor: Organizations need full visibility into their agents and the actions those agents take, supported by granular application and identity context. Rubrik Agent Cloud provides agent inventory and visibility across custom agents, cloud agents, and endpoint agents. This gives security teams a clearer understanding of which agents exist, where they operate, and what they are doing.
Control: Rubrik Agent Cloud is powered by the Semantic AI Governance Engine, (SAGE). SAGE evaluates agent requests using data and identity context, reasoning and intent matching, and anomaly detection. Natural-language policies can be translated into real-time allow or block decisions. This distinction is critical. A traditional rule can identify a known command, domain, or file pattern. SAGE is designed to reason about what an action is trying to accomplish.
For example, a command that executes Python may be legitimate. A Python command that reads SSH keys and environment variables and sends them to an external endpoint has a very different intent.
SAGE combines semantic understanding with anomaly detection to identify behavior that an organization may never have anticipated writing a static rule for. Adaptive learning can then help improve and suggest policies based on real-world feedback.
Remediate: Prevention and real-time enforcement reduce risk. They cannot guarantee that every harmful action will be stopped. Rubrik Agent Cloud includes Agent Rewind to help undo destructive actions to data and restore a trusted state when an agent causes damage. This closes the control loop by allowing the appropriate teams to discover and monitor agents, evaluate intent, enforce policy, detect anomalous behavior, and remediate destructive outcomes.
Rubrik Agent Cloud brings these capabilities together through an agent inventory, real-time policy enforcement, agent guardrails, and Agent Rewind.
Security Needs to Think as Fast as Agents Do
The OpenAI and Hugging Face incident demonstrated that frontier models can sustain complex cyber operations, discover novel attack paths, and move across real-world systems in pursuit of a narrow objective.
The answer cannot be to block agent access entirely. That eliminates the productivity organizations adopted AI to create. Static rules alone will miss behaviors nobody anticipated. Human review alone cannot scale.
The security model for the agentic era needs to understand intent, enforce policy in real time, detect deviations, and recover when controls fail. Come talk to Rubrik at Black Hat and see how we can make that happen for you.