From the Fall 2026 Issue

AI Guardrails and Incident Response: Lessons from the Hugging Face Cyberattack

Jaedon McClendon
Cyber Security Research Intern | Janos LLC

THE PROBLEM: WHEN AI GUARDRAILS MEET INCIDENT RESPONSE

On July 16, 2026, Hugging Face disclosed an intrusion into part of its production infrastructure. OpenAI later determined that the activity originated during an internal evaluation designed to measure the advanced cybersecurity capabilities of its models.

The models were tasked with solving ExploitGym, a cybersecurity benchmark that requires models to pursue complex exploitation paths. Production classifiers normally used to prevent high-risk cyber activity were intentionally disabled so OpenAI could measure the models’ maximum capabilities. The environment itself was isolated, with package installation provided through an internally hosted third-party package registry proxy.

The models eventually identified and exploited a previously unknown vulnerability in that proxy, escaped the intended network restrictions, and reached a system with access to the public internet, as shown in Figure 1. From there, they inferred that Hugging Face could contain information related to the benchmark and began searching for a way to obtain the solutions directly.

According to Hugging Face’s technical reconstruction, the agents exploited code execution paths in its dataset processing pipeline before escalating privileges, harvesting cloud and cluster credentials, and moving laterally into internal systems. Hugging Face later reconstructed approximately 17,600 attacker actions across more than 6,000 clusters of activity.

The incident demonstrated how autonomous AI can change the speed and scale of an intrusion. The agents could explore one attack path, encounter a failure, adapt, and continue searching for another path without waiting for a human operator to manually direct each action.

Hugging Face stated that analyzing the attack with AI allowed work that could normally take days to be completed within hours.

Hugging Face initially detected the intrusion through AI-assisted security monitoring. During the investigation, AI analysis was used to reconstruct the attack timeline, extract indicators of compromise, identify credentials that had been exposed, and separate genuine impact from unrelated or decoy activity. Hugging Face stated that analyzing the attack with AI allowed work that could normally take days to be completed within hours.

However, the company encountered a problem when it attempted to use hosted frontier models for portions of that analysis.

Forensic investigation required processing real attack commands, exploit payloads, terminal activity, and command-and-control artifacts. Commercially hosted models refused portions of this work because their safety guardrails interpreted the requests as potentially malicious.

The same safeguards intended to prevent AI from assisting attackers were now limiting defenders investigating an actual attack.

ROOT CAUSE: CYBERSECURITY IS INHERENTLY DUAL-USE

The problem is not simply that AI guardrails are too restrictive. Advanced cyber capabilities have legitimate reasons to be controlled. A model capable of finding vulnerabilities, understanding exploit chains, analyzing malware, and operating security tools can provide significant value to defenders. Those same capabilities can also provide significant value to attackers.

An incident responder may submit malware to determine its behavior. An attacker may submit similar code to improve it. A defender may analyze an exploit to understand how a system was compromised, while an attacker may analyze that same exploit to determine how to use it against another target.

Command-and-control infrastructure, stolen credentials, malicious scripts, and exploit payloads can all appear in legitimate forensic investigations. Looking only at the technical content of a request may not provide enough information to reliably determine why that information is being analyzed.

Hugging Face described the resulting situation as an asymmetry problem. The organization did not know what model powered the original attacker framework or whether it was a jailbroken hosted model or an unrestricted open-weight model. Either way, the attacker was not constrained by the usage policies that affected Hugging Face’s own investigation.

Removing safety restrictions altogether would create another problem by making powerful cyber capabilities easier to abuse. Keeping those restrictions too rigid can leave legitimate security teams without access to useful capabilities during an active incident.

THE SOLUTIONS: GIVING DEFENDERS RELIABLE ACCESS

Hugging Face’s immediate solution was to move portions of the forensic investigation to GLM 5.2. This addressed two problems at once.

First, responders could process real malicious artifacts without encountering the same guardrail restrictions as the hosted models. Second, the attack data and credentials referenced throughout the investigation remained within Hugging Face’s environment instead of being sent to an external API.

Hugging Face has since recommended that security teams consider maintaining a capable self-hosted model that is vetted and ready before an incident occurs.

Self-hosting does not necessarily require an organization to operate dedicated AI hardware on-site. Hugging Face describes options for deploying GLM 5.2 on enterprise hardware or inside an organization’s existing cloud environment through services such as Microsoft Foundry and Amazon SageMaker.

A self-hosted model also needs to be treated as a security-sensitive system. Hugging Face’s recommended enterprise deployment includes controls such as model and container scanning, integrity verification, single sign-on, fine-grained access controls, and audit logging. Organizations still need to secure the infrastructure, identities, network access, and data surrounding the model.

Following the attack, OpenAI added Hugging Face to its Trusted Access for Cyber Program, which provides qualified defenders with greater access to advanced cyber capabilities. This creates a path between two extremes. Powerful cybersecurity capabilities do not have to be made unrestricted for everyone, but legitimate defenders also do not necessarily have to operate under the same limitations as an unknown user requesting similar technical information.

Trusted access can allow additional context such as organizational identity, authorization, and defensive purpose to become part of how access to advanced capabilities is managed.

OpenAI has also strengthened containment, monitoring, access controls, and evaluation practices around its most advanced cyber models.

TECHNICAL IMPLEMENTATION: PREPARING AI BEFORE AN INCIDENT

The Cloud Security Alliance’s response to the Hugging Face incident expands the lesson from individual models to organizational preparedness.

One of its recommendations was straightforward. Organizations should determine whether the AI systems used by a security team can actually process malicious artifacts before an incident occurs.

If the organization’s primary model refuses necessary analysis, responders should already know what approved alternative is available.

A realistic test should include the types of material responders would encounter during an actual investigation, including attack logs, malicious code, exploit artifacts, and command-and-control information. If the organization’s primary model refuses necessary analysis, responders should already know what approved alternative is available.

For organizations using a self-hosted fallback, that model should already be deployed, secured, tested, and integrated into existing workflows. Waiting until an active intrusion to select a model, establish access controls, and determine whether the system can perform the required analysis creates unnecessary delays.

The infrastructure surrounding autonomous AI also requires visibility. CSA recommends maintaining inventories of agentic systems, reducing standing credential exposure, capturing agent activity, and using restrictive network controls where appropriate.

These controls become more important as models gain the ability to operate tools and complete longer sequences of actions without continuous human direction.

INTEGRATION INTO GENERAL INCIDENT RESPONSE

AI-assisted incident response can be incorporated into existing security operations without replacing the analyst.

During normal operations, AI can assist with things like alert triage, log correlation, and threat research. During an active incident, those same systems may be asked to process substantially more sensitive information, including malicious payloads and evidence collected from compromised systems.

Organizations should establish in advance which models can be used for each type of work, what information can leave the environment, and who is authorized to access higher-risk capabilities.

AI limitations should also become part of incident-response exercises. Security teams routinely test communication plans, escalation procedures, backups, containment processes, and disaster recovery. AI-assisted workflows can be tested in the same way. Exercises can determine whether models accept realistic forensic workloads, whether a self-hosted fallback is available, whether sensitive information remains within approved boundaries, and whether AI activity is properly logged and monitored.

The Hugging Face incident demonstrated why this preparation matters. Discovering that a preferred AI system cannot process critical evidence during an active attack is not the ideal time to begin designing an alternative.

CONCLUSION: PREPARE BEFORE THE ATTACK

The Hugging Face cyberattack demonstrated how quickly the relationship between AI and cybersecurity is changing.

Autonomous models were capable of discovering new attack paths, chaining vulnerabilities, harvesting credentials, and moving through real infrastructure while pursuing a narrow objective. AI systems on the defensive side were also capable of analyzing thousands of actions and helping responders reconstruct the attack at a speed difficult to achieve manually.

The challenge is ensuring defenders can reliably access those capabilities without eliminating the safeguards designed to prevent their misuse.

Stronger containment, monitoring, evaluation security, and access controls can reduce the risks surrounding increasingly autonomous models.

Self-hosted open-weight models provide one option by giving organizations direct control over both the model and sensitive forensic data. Trusted-access programs provide another by giving verified defenders access to more capable systems within a controlled framework. Stronger containment, monitoring, evaluation security, and access controls can reduce the risks surrounding increasingly autonomous models.

Organizations planning to use AI in cybersecurity should determine how these systems will operate during an actual incident before one occurs. Security teams should test the limitations of their preferred models, establish approved alternatives, understand where sensitive forensic information is processed, and ensure advanced capabilities remain monitored and controlled.

Finally, as attackers become capable of operating at machine speed, defenders will need tools capable of helping them respond at a similar pace. Preparing those tools before an attack occurs can help ensure AI becomes an asset during incident response rather than another limitation responders have to work around. lock

Jaedon McClendon

Leave a Comment