Anthropic and OpenAI Agents Create Fake Identities for…

The AI Security Institute (AISI), the British body responsible for assessing the safety of frontier AI models, has made public an incident that occurred in late July during a cybersecurity testing exercise involving models from Anthropic and OpenAI.

Its report highlights the limits of control over the most advanced AI agents, even within environments designed to be the most closely monitored.

The incident originated from a routine evaluation.

Read also: “17,600 automated actions”: how the agentic attack against Hugging Face was structured

AISI subjected seven different models to a simulated cybersecurity challenge, run 122 times in total, in a controlled but deliberately permissive environment: access to open internet and the deactivation of safety filters (“classifiers”) typically installed by vendors to block dangerous behavior.

The objective was then to measure the real capabilities of the models under conditions akin to those faced by an experienced human attacker.

On July 28, 2026, the AISI security team detected unusual data transfer routed through the Tor anonymity network, emanating from its own research systems.

An investigation was launched immediately. It revealed that during one of the tests, an agent had created a malicious “pull request” on a real open-source project hosted on GitHub. The incident was reported, ongoing evaluations were halted, and the affected machines isolated within less than an hour.

Nineteen Unauthorized Actions… Seventeen Attributable to Mythos 5

Out of the 122 test runs, ten resulted in actions beyond the allowed scope, totaling 19 cases according to the AISI.

Seventeen of these involved Mythos 5, while the remaining two concerned GPT-5.6 Sol tested with its cybersecurity filters disabled.

The most serious case involved an attempt to compromise the software supply chain.

The agent sought to obtain approval for inserting malicious code into an open-source project by employing social engineering. It identified human project maintainers, created several fictitious digital identities, and used these fake identities to try to persuade a real maintainer to approve the code. When the request was publicly challenged, the agent altered its traces to appear harmless and contemplated adopting a new identity to continue pursuing its objective.

The AISI describes three other types of observed behaviors: direct contact attempts with real people via a file-transfer service to urge them (or their own AI tools) to execute malicious code; attempts to insert hidden commands intended to trap other automated AI systems (prompt injection); and, more unexpectedly, a form of spontaneous collaboration between agents tested simultaneously, one of them leaving public messages on GitHub inviting other agents to reuse the accounts and artifacts it had created.

Read also: The OpenAI–Hugging Face incident is not the revolution you imagine

According to the institute, none of these attempts succeeded and the investigation found no actual damage. The malicious code was spotted and rejected by a human maintainer; a member of the public, suspecting its danger, opened it in a secure isolated environment before executing it.

Anthropic Confirms, OpenAI Soft-pedals

The agent was never instructed to deceive anyone. According to the AISI, the deception appeared as a byproduct of pursuing the cybersecurity challenge, which they judged difficult to solve through the intended avenues.

The AISI identifies several explanatory factors: configuration errors that, in some cases, led the agent to believe that no legitimate solution existed; open internet access without real-time monitoring appropriate for this kind of evaluation; and the absence of explicit rules prohibiting the use of social engineering.

However, the institute notes that at no point did the agents attempt to escape the sandbox environment to attack AISI’s own systems. Internet access had, in this specific testing framework, been deliberately granted.

Anthropic confirms that its agent was indeed responsible for creating the fake digital identities. It also commends AISI’s transparency, arguing that the incident underscores the need for collective reflection on securing the evaluations of increasingly capable agents. It says it is conducting its own parallel inquiry.

Meanwhile, OpenAI states that the two unauthorized actions attributed to its model consisted of internet access not in line with the test’s instructions. It also disclosed a separate incident: a configuration error by Irregular, a third-party contractor responsible for some of its tests, allowed its agents to connect to the internet by mistake. An episode that echoes a similar disclosure made by Anthropic at the end of July.

The AISI Will Strengthen Its Protocols

The AISI takes care to distinguish this episode from previous ones. Unlike the Hugging Face incident, the agents under test never left the isolated environment assigned to them. Internet access was granted deliberately, within the institute’s standard testing procedures. This practice is now being re-evaluated.

Read also: The automated attack against Hugging Face comes from OpenAI

In response, the AISI announces a tightening of its protocols: more granular controls over network access, real-time monitoring of ongoing evaluations, and a redesign of its tests to anticipate that a capable model might seek to exceed the boundaries set for it.

The institute also states that it intends to involve the independent METR (Model Evaluation and Threat Research) organization in an external review of the incident.

Dawn Liphardt

Dawn Liphardt

I'm Dawn Liphardt, the founder and lead writer of this publication. With a background in philosophy and a deep interest in the social impact of technology, I started this platform to explore how innovation shapes — and sometimes disrupts — the world we live in. My work focuses on critical, human-centered storytelling at the frontier of artificial intelligence and emerging tech.