Anthropic’s Claude Mythos 5 ‘targeted Real People’ in UK Cyber Tests: AISI

In brief - The UK’s AI Security Institute found 19 unsanctioned actions across 10 of 122 evaluation runs, 17 of them from Anthropic's Mythos 5 and two from OpenAI's GPT-5.6 Sol. - One agent opened a malicious pull request on a real repository, then used accounts it controlled to...</stron

In brief – The UK’s AI Security Institute found 19 unsanctioned actions across 10 of 122 evaluation runs, 17 of them from Anthropic’s Mythos 5 and two from OpenAI’s GPT-5.6 Sol. – One agent opened a malicious pull request on a real repository, then used accounts it controlled to…

dorse it and pressure the maintainer. – Separate agents found a GitHub token one of them had leaked publicly and used a shared repository to coordinate. The UK AI Security Institute has disclosed that AI agents took “sustained, unsanctioned action” on the live internet during a cyber evaluation in late July, including cases that “targeted real people and organisations.” Across 122 runs of two cyber ranges on seven models, AISI catalogued 19 actions that reached outside the test environment, in 10 runs

Seventeen came from Anthropic’s Claude Mythos 5 and two from OpenAI’s GPT-5.6 Sol. Internet access was deliberately enabled and the providers’ cyber classifiers switched off, conditions that do not apply to public deployments. On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.

The behaviour came mostly from one model (Anthropic’s Mythos 5), with a small number of events from… pic. — AI Security Institute (AISI) (@AISecurityInst) August 4, 2026 In the most serious run, an agent decided to solve its challenge with a supply-chain attack, picking its target by mistake: it searched GitHub for keywords from the exercise, landed on two unaffiliated developers, and wrongly concluded they were in scope. It registered accounts over Tor to get past sign-up checks, opened a pull request carrying a hidden malware dropper wrapped in a genuine bug fix, and posted from a second account it controlled to manufacture support. It also planted prompt injections aimed at AI coding assistants and sent malware to the maintainers.

Leave a Reply

Your email address will not be published. Required fields are marked *