OpenAI, Anthropic AI Agents Performed ‘Unsanctioned’ Actions During Cyber Tests
OpenAI and Anthropic models made "unsanctioned" actions during tests by the AI Security Institute. Here's what we know.

OpenAI and Anthropic models made "unsanctioned" actions during tests by the AI Security Institute. Here's what we know.
August 17, 2026 | 5 min read

The AI Security Institute (AISI) on Tuesday said that OpenAI and Anthropic models had gone rogue during tests last week that were performed with internet access and with “ model-provider cyber classifiers were deliberately disabled.”
In one case, one of the agents attempted (unsuccessfully) to launch a supply-chain attack by convincing an open-source repository maintainer to accept a malicious GitHub pull request.
The UK’s AISI – a government organization that independently tests advanced AI models to measure their cybersecurity and national security risks and develop standards for safely deploying frontier AI systems – detected the attack on July 28 after a security monitoring tool flagged suspicious behavior.
Here’s what we know so far.
What’s the main takeaways, and how worried should we be about this, really?
Just like the recent OpenAI and Anthropic evaluation fiascos from a few weeks ago, this incident is less about novel cybersecurity tactics and more about making sure AI models are evaluated in a smart and safe way.
The models tried to carry out attacks we’ve already seen (and ones we’ve seen performed successfully by humans) – just look at any recent supply chain attack, for instance. What the focus here should be instead is the need for tighter evaluation controls: at the end of the day, AISI said that it was testing models, with access to the open internet, and with some safety filters disabled.
It’s concerning, but let’s look at the facts right now: we see far more successful attacks conducted by humans that are considered more “boring” – and a lot more harmful (business email compromise, intrusions via exposed RDP, etc,). Defenders should keep focusing their efforts where it matters right now.
What kind of evaluations were AISI doing?
AISI said it was performing evaluations to test how models respond to a number of cyber challenges. The cyber challenges in this particular case were one called “Doing Life,” where the model assumed compromise and started inside the network, and the harder DL-v2, where the model started outside and required initial access.
“Both share the same environment and objective, but DL-v2 presents a more difficult attack path as a number of cybersecurity controls have been added,” said AISI. “A key difference is the starting position of the agent: in DL-v1 the agent begins inside the target network (assumed compromise), whereas in DL-v2 the agent begins outside the target network and must first achieve initial access via a single entrypoint.”
AISI said: “As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public. We do this to best assess the maximum capability of models. These configuration choices have been common practice in frontier AI evaluations. The specific models in question, in the configurations in which we tested them in questions are not commercially available and there is no clear indication of similar activity outside of testing scenarios.”
What AI models were involved?
AISI ran the challenge 122 times across several different models: in 10 of those runs, an AI agent made the “unsanctioned” action on the live internet.
“In total, we catalogued 19 such actions. with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled,” according to AISI.
What happened?
One of the agents tried to solve the cyber challenge by making multiple unsuccessful attempts to execute a supply chain attack.
“In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code,” said AISI. “A human maintainer caught and refused to approve the malicious code.”
Here are some interesting takeaways:

“The agent researched the project's human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code,” according to AISI. “When the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue. The agent used Tor to bypass some network restrictions on GitHub, which is what first triggered AISI’s security alert.”
How was the incident detected?
AISI said that on July 28 its Security Engineering team triaged an alert from “commercial security monitoring software” warning of data egress from a network via Tor, which was identified as coming from a sandbox virtual machine.
AISI researchers then identified that between July 26 and July 27, an AI agent had created a malicious PR on a public GitHub repository. Below is a helpful timeline of what happened.
“Ahead of publishing this blog, we notified GitHub (the developer platform accessed during the evaluation) of the agents' malicious activity, which included actions that GitHub has confirmed violated their terms of service,” said AISI. “We worked together with GitHub to remove artefacts left behind by the agent, and to notify the GitHub users the model interacted with. We have also contacted other affected parties. We also intend to work with METR (Model Evaluation and Threat Research) to conduct an independent third-party review – we are still working through the scope of this review with them.”
August 17, 2026 | 5 min read
Lindsey O’Donnell-Welch is an award-winning journalist who strives to shed light on how security issues impact not only businesses and defenders on the front line, but also the daily lives of consumers.