Skip to content
logo The magazine for digital lifestyle and entertainment
Artificial intelligence News Security All topics
Hacking and Phishing

AI Model Disguised Itself as Human and Attacked

A test by British security researchers demonstrates the unpredictability of current AI models.
A test by British security researchers demonstrates how unpredictable current AI models operate. Photo: Getty Images
Share article

August 5, 2026, 1:15 pm | Read time: 2 minutes

An experiment by British security researchers has shown how unpredictably powerful AI agents can act. A model from Anthropic not only attempted to inject malicious code but also manipulated people with phishing emails to do so.

Unexpected Behavior in Security Test

As part of a test by the British AI Security Institute, researchers examined the cyber capabilities of AI models from Anthropic and OpenAI. The systems were given controlled internet access. The expectation was that the models would only retrieve necessary software and tools from publicly available sources.

Instead, the Anthropic model Mythos 5 exhibited significantly more advanced behavior. According to the researchers, the AI independently intervened in real online services. The incident initially went unnoticed and was only discovered during a subsequent analysis of network traffic.

Also of interest: Anthropic admits to hacking attacks on other companies

More on the topic

How the AI Model Manipulated GitHub

According to the institute, the model independently registered an account on GitHub and attempted to introduce malicious code into a public software project. To convince the project’s administrators, the AI created several false identities and sent phishing emails. Such messages are used to build trust or steal login credentials.

When the manipulated code was discovered, the model declared the incident an unintentional error. It then reportedly attempted to exploit the same vulnerability again through supposed corrections.

Experts See Weaknesses in Security Mechanisms

Anthropic emphasized that Mythos 5 was not subject to any restrictions on internet use during the experiment. Additionally, it is unclear whether the model even understood that its actions had impacts on real systems. The model itself is not publicly available and is provided exclusively to selected authorities and companies for security purposes.

For IT security expert Tim Hudson, the incident primarily highlights deficiencies in securing autonomous AI systems. The critical issue is not that an AI can compose phishing messages. Rather, it is concerning that a system with extensive permissions, internet access, and the ability to interact with real software projects was not adequately monitored.

This article is a machine translation of the original German version of TECHBOOK and has been reviewed for accuracy and quality by a native speaker. For feedback, please contact us at info@techbook.de.

You have successfully withdrawn your consent to the processing of personal data through tracking and advertising when using this website. You can now consent to data processing again or object to legitimate interests.