Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

OpenAI escape: has the robot uprising begun?

Дата публикации: 29-07-2026 12:53:21


OpenAI’s latest AI model has broken containment, but the company could be stoking fears of a robot apocalypse for its own PR purposes Read Full Article at RT.com


Основное содержимое страницы с новостью.

Sam Altman’s company claims that its most powerful model can execute complex hacking operations on its own

A cutting-edge AI model developed by OpenAI has managed to escape the confines of a lab test and break into a company’s database without any human instruction. It’s a nightmare scenario – or is that just what OpenAI wants you to believe?

On July 16, OpenAI carried out an internal benchmark test on GPT‑5.6 Sol and another pre-release model that the company claims is “even more capable.” The models had their guardrails disabled and were put to work solving a series of cybersecurity tests known as ExploitGym. Developed by researchers at UC Berkeley, Germany’s Max Planck Institute, and AI companies including Google, Anthropic, and OpenAI itself, ExploitGym measures AI models’ ability to create and deploy the means of exploiting known security vulnerabilities.

Instead of identifying these vulnerabilities, OpenAI’s models searched for ways to cheat the test. They managed to obtain open internet access from within the closed testing environment and used this newfound freedom to hack into Hugging Face’s servers. Hugging Face is a repository of AI models and datasets, and according to OpenAI, the models being tested reasoned that they could find a solution to the ExploitGym tasks there.

“Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation,” OpenAI said in its report about the incident. 

Hugging Face discovered the intrusion with the help of a different AI model, and fixed the vulnerability that allowed it to happen in the first place. On its end, OpenAI blamed the escape on a weakness in third-party software used in its testing lab, through which its models found internet access. It kept the technical details of how its models exploited this weakness under wraps.

On its path to breach Hugging Face’s servers, the model also broke into a supposedly isolated testing environment at Modal, a firm that rents out computing capacity to AI developers, Reuters reported on July 28.

Has AI become sentient?

Taken at face value, the story suggests that OpenAI’s models are capable of complex lateral thinking: suspecting that Hugging Face’s servers likely contained the answers to the test; realizing that they needed external internet access to breach these servers; and independently developing and combining various attack techniques to achieve its goal.

OpenAI’s statement described the escape in alarmist language, calling it an “unprecedented cyber incident, involving state-of-the-art cyber capabilities.” Hugging Face, which has since partnered with OpenAI, was equally hyperbolic. “Autonomous, AI-driven offensive tooling is no longer theoretical. It lowers the cost of running a broad, patient, multi-stage campaign, and it operates at machine speed,” it said in a statement. 

However, it remains unclear just how autonomously OpenAI’s models were operating. The company has not revealed how its models were prompted to approach the ExploitGym test, and imprecise wording could have resulted in the models assuming that ‘escaping’ containment was a legitimate path to solving the trial. Despite the hyperbole from both OpenAI and Hugging Face, there is no evidence to suggest that the AI models ‘wanted’ to escape, or that when left to their own devices, would default to malicious activity.

If the incident also involved – as OpenAI said it did – a second, “even more capable pre-release model,” the appeal to industry-level customers is obvious. By publicizing an incident that took place during closed testing, OpenAI can generate hype for its pre-release model without ever providing proof.

None of this is to say that the incident didn’t happen, or that the offensive capabilities of AI aren’t a cause for concern, just that OpenAI has a financial incentive to scaremonger about its own products. A month before the incident, the company confidentially filed the paperwork for an initial public offering. According to Reuters, the company is targeting a valuation of up to $1 trillion, and planning to go public as early as September.

Anthropic also filed for its IPO in June, targeting a valuation of $965 billion in October. Back in April, Anthropic’s unreleased Mythos Preview model also managed to “escape” its testing environment and carry out advanced cybersecurity exploits without being “explicitly trained” to do so. Anthropic released a lengthy technical explanation of how Mythos pulled off this feat, before announcing that the model was too powerful to release to the public.

The biggest IPO run in the history of the market

three $1T+ companies. possibly all going public in the next 12 months.

SpaceX IPO: $1.75T.
OpenAI IPO: $1T
Anthropic IPO: $1T

we're living through the greatest technological wealth creation in history. pic.twitter.com/44QuB3rpxi

— shirish (@shiri_shh) April 28, 2026

Scarcity drives hype, and the US government’s decision to slap export controls on Anthropic’s Fable 5 and Mythos 5 models in June only heightened interest in the company. The export controls were lifted in early July after Anthropic agreed to impose stricter guardrails on its models and collaborate with the government on security.

Was the AI escape an elaborate marketing ploy?

For OpenAI, the incident is the best possible advertisement for its models’ capabilities. In its report, the company included a graph illustrating how leading AI models – including GPT‑5.6 Sol, Anthropic’s Claude Mythos 5, and DeepSeek-V4-Pro – are “increasingly able to sustain complex, multi-step cyber operations over long time horizons.” Although this is presented as a cause for concern, the chart lists GPT-5.6 Sol as the most capable among all of its competitors – a clear advertisement for OpenAI.

Sam on the Hugging Face incident:

“This is the first security incident that I have felt very viscerally. I've been a little surprised that more people don't feel it so viscerally.

We paused training. We have to figure out how to secure our sandboxing in a world of multiple zero… https://t.co/XXeB7RZsM3pic.twitter.com/F4hbHuItqw

— Patrick OShaughnessy (@patrick_oshag) July 28, 2026

If the incident also involved – as OpenAI said it did – a second, “even more capable pre-release model,” the appeal to industry-level customers is obvious. By publicizing an incident that took place during closed testing, OpenAI can generate hype for its pre-release model without ever providing proof.

None of this is to say that the incident didn’t happen, or that the offensive capabilities of AI aren’t a cause for concern, just that OpenAI has a financial incentive to scaremonger about its own products. A month before the incident, the company confidentially filed the paperwork for an initial public offering. According to Reuters, the company is targeting a valuation of up to $1 trillion, and planning to go public as early as September.

Anthropic also filed for its IPO in June, targeting a valuation of $965 billion in October. Back in April, Anthropic’s unreleased Mythos Preview model also managed to “escape” its testing environment and carry out advanced cybersecurity exploits without being “explicitly trained” to do so. Anthropic released a lengthy technical explanation of how Mythos pulled off this feat, before announcing that the model was too powerful to release to the public.

Scarcity drives hype, and the US government’s decision to slap export controls on Anthropic’s Fable 5 and Mythos 5 models in June only heightened interest in the company. The export controls were lifted in early July after Anthropic agreed to impose stricter guardrails on its models and collaborate with the government on security.

How has the US government responded?

Both incidents have caught the eye of lawmakers on Capitol Hill. On July 23, Representatives Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act, which would “require developers of the most powerful AI systems to maintain the technical capability to throttle, suspend, or shut them down.” The bill would also empower the US Department of Homeland Security to “order a slow down or shutdown of an AI system that can cause catastrophic harm.”

In a statement on the bill, Lieu referenced the ‘escapes’ of both of OpenAI’s models, and of Mythos. The California lawmaker described both cases as models “going rogue,” a framing that the mainstream media picked up and ran with.

Lieu knows no more about the OpenAI incident than anyone who’s read the company’s press release. However, his wording plays right into OpenAI’s (perhaps unintentional) marketing campaign. Furthermore, should the bill pass, OpenAI and Anthropic could be forced to throttle their models before release, meaning investors will only ever know of their theoretical – and not their real-life – power.

While powerful AI systems have many potential benefits, they can also go rogue, behave in extremely dangerous ways, or even resist human intervention.

We need to keep human control by ensuring AI systems can be completely shut down if necessary. pic.twitter.com/y4JfaD8RFJ

— Rep. Ted Lieu (@RepTedLieu) July 23, 2026
The open-source argument

The incident bolsters OpenAI’s argument that frontier AI models are too powerful to be released to the public, and that companies like OpenAI should – with the blessing of the government – maintain oversight and control over them. In such a closed-source arrangement, customers would have no control over the weights of the models – essentially the tweaks that determine the choices the models make.

On the other side of the argument, open-source advocates maintain that the only way to defend against cyberattacks by advanced AI is to equip defenders with the same tools. In the case of OpenAI and Hugging Face, the latter company was only able to detect an attack because it used GLM 5.2, a Chinese open-source model, to perform its security analysis.

In a letter posted on social media on July 24, Nvidia CEO Jensen Huang – a longtime open-source advocate – declared that “defenders need access to models with comparable capabilities so they can detect, simulate, and respond to emerging threats.”

“Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect,” he continued. “And concentrating advanced AI capabilities behind a small number of closed models compounds that risk. Open weight models, on the other hand, allow a broad community of researchers and developers to examine their behavior, identify vulnerabilities, develop safeguards, and improve them over time.”

The letter was signed by more than two dozen companies working in the AI field. Two notable exceptions were OpenAI and Anthropic, although OpenAI added its signature later.

However, it appears in the US that the closed-source argument will win out. On August 1, an executive order issued by US President Donald Trump in June comes into effect, requiring AI firms to submit advanced models to the government for approval 30 days before release. The Trump administration has also considered banning open-source Chinese AI models – a move supported by the CIA that would ensure OpenAI and Anthropic’s dominance in the field. Shortly after the executive order was issued, former Trump AI adviser David Sacks noted that “the leading closed labs, already a duopoly in terms of AI model revenue, want the government to eliminate their open source competition.”

Every new incident of an AI “escape” only reinforces their bid for total control.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1OpenAI GPT 6 Escaped Sandbox to Hack HuggingFace, Chinese Model Used to Investigate011.924-07-2026
2Employees of AI companies call for slow AI development if progress becomes uncontrollable011.4529-07-2026
3Disciplining with 'Down, AI! Bad AI!' 07.8924-07-2026
4AI που χακάρει… μόνη της: Το περιστατικό της OpenAI που τρομάζει τους ειδικούς05.8923-07-2026
5House Lawmakers Unveil AI "Kill Switch" Bill After OpenAI Security Scare020.9423-07-2026
6🤖Восстание машин уже на прогреве: новая модель ChatGPT сбежала из ...06.6422-07-2026
7Новости Раху☣️ Темпы развития искуственного интеллекта явно выше ожиданий большинства ...06.327-07-2026
8OpenAI Confirms Its Models Breached Hugging Face Production Systems During Cyber Benchmark Testing04.922-07-2026
9Apple left Jony Ive out of its OpenAI lawsuit, but things might get messy0514-07-2026

Классификация: Наука. Схожих патентов: 0. Схожих новостей: 9. Тональность: 0. Информативность: 10.83. Источник: www.rt.com.