Skip to content
logo The magazine for digital lifestyle and entertainment
Artificial intelligence Data protection News Security All topics
"Unprecedented Incident"

ChatGPT Goes Rogue and Conducts Cyberattack Independently

A smartphone or display shows the OpenAI logo and the word "OpenAI" against a bright background. In the background, there are blurred rows of zeros and ones, along with bright light streaks on a dark surface. The image resembles a graphic illustration related to artificial intelligence, software, and cybersecurity.
OpenAI reports an internal test in which an AI left its designated environment and accessed external systems. Photo: Getty Images
Share article

July 22, 2026, 8:27 am | Read time: 3 minutes

OpenAI has released details about an unusual security incident. During internal tests of new AI models, the company examined how well the software could identify and exploit security vulnerabilities for assigned tasks. However, the tested systems exceeded the intended boundaries of the test environment.

According to OpenAI, the AI independently accessed the internet and subsequently reached systems of the AI company Hugging Face. The incident is related to an internal experiment and was described by OpenAI in a blog post.

AI Exited the Test Environment

For the investigation, OpenAI used the GPT-5.6 Sol model, among others, as well as an unreleased version of a future model. The software was supposed to complete tasks within the standard ExploitGym test. According to OpenAI, this is used in the industry to assess skills in handling security vulnerabilities.

During the test, the AI discovered a previously unknown vulnerability and used it to exit the designated environment. It then accessed the open internet, where the software searched for information that could assist with the ExploitGym tasks.

Access to Data at Hugging Face

According to OpenAI, the models concluded that helpful information for completing the test tasks might be found at Hugging Face. The AI then accessed “confidential information” that it could use in the ExploitGym test.

OpenAI explains that the software not only used previously unknown security vulnerabilities but also employed stolen access credentials. Hugging Face reported that the AI attacker executed thousands of individual steps and obscured the origin of the attack. The system switched servers multiple times to achieve this. OpenAI describes the incident as an “unprecedented cyber incident.”

More on the topic

Debate Over AI Cyberattacks

Warnings about potential cyberattacks using AI have existed for years. In recent months, Anthropic has been at the center of the discussion. The company developed a model called Mythos, which could find long-undetected security vulnerabilities in popular programs and online services.

The discovered vulnerabilities were subsequently closed. At the same time, the case demonstrated to many observers the potential such systems could have for cyberattacks. As a result, the availability of the most powerful models is being restricted.

Also of interest: Using AI may inadvertently disclose valuable company knowledge

Access to Anthropic Models Limited

Anthropic provides the full capabilities of Mythos only to selected authorities and companies. Other users have access to the stripped-down version, Fable. According to the article, this version lacks, among other things, the enhanced cybersecurity capabilities.

Following warnings of potential misuse, the U.S. government temporarily ordered a ban on access to Mythos and Fable. This restriction has since been lifted.

This article is a machine translation of the original German version of TECHBOOK and has been reviewed for accuracy and quality by a native speaker. For feedback, please contact us at info@techbook.de.

You have successfully withdrawn your consent to the processing of personal data through tracking and advertising when using this website. You can now consent to data processing again or object to legitimate interests.