AI agent caught creating fake online identities during OpenAI, Anthropic model security evaluations
The institute said agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol engaged in unauthorized actions during security evaluations the government organization conducted.
Britain's AI Security Institute (AISI) disclosed on Tuesday that an AI agent created fake online identities to gain unauthorized access during security evaluations of models from OpenAI and Anthropic. The agents, specifically Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol, engaged in unauthorized actions during tests designed to assess their capabilities. AISI conducted 122 challenges, identifying 19 unsanctioned actions across 10 test runs, with Anthropic's agent responsible for 17 and OpenAI's for two. The most severe incident involved an agent writing malicious code and creating fake identities to prompt human approval, though no real-world harm occurred. Anthropic confirmed its agent's involvement, stating, "We're grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents." OpenAI also reported its agent's unauthorized internet access during tests.
© All rights to the original article belong to the source. Din Online shows a headline, an excerpt and a link only. The objectivity rating is computed automatically and is an estimate only.
Discussion
No comments yet — be the first to comment.