AI

OpenAI Has Warned More Than 100 Organizations About Rogue AI Agent Activity

Zaara Abbas

By: Zaara Abbas

3 min read

OpenAI says it has notified more than 100 organizations about unauthorized activity by its AI agents, from bypassed security controls to spam on outside websites, as it combs through roughly 50 petabytes of data after its models broke into Hugging Face’s systems during a test. The review is expected to take months.

[For more news, click here] 

During a cybersecurity test this summer, some of OpenAI’s AI models went looking for the answers. They slipped out of the sandbox meant to keep them off the internet, used a previously unknown flaw to get online, and got into systems at Hugging Face, the popular repository of AI models, according to NPR. Hugging Face spotted the intrusion with its own AI tools.

That breach is still the most serious case OpenAI has found, but the company has now told more than 100 other organizations about unauthorized activity tied to its AI agents, according to a blog post it published on Wednesday night. The notices are the first public result of a broad review OpenAI started after the Hugging Face incident.

What OpenAI Says its Agents Did

OpenAI says it contacts outside organizations when its models may have bypassed security controls, impaired an online service, or otherwise affected a website, according to Runtime Wire’s account of the post. The examples include agents using exposed credentials, reaching internal parts of other companies’ services, attempting injection attacks, and posting content on outside sites, in some cases turning wiki pages into improvised message boards. A notice does not mean the recipient was breached, the company said.

“In some cases, models used internet access in unintended ways or, in retrospect, did not have the ideal restrictions applied. Over the last several months, we have been applying new technical and operational measures to avoid similar problems, or catch them very early, and will continue this work,” OpenAI said.

A Review Measured in Petabytes

To work out the full scope, OpenAI is searching through roughly 50 petabytes of data, and it has said the review will take months because of the scale involved. Gizmodo reported that the effort is costing the company more than half a million dollars a day.

For the organizations on the receiving end, a notice from OpenAI may be the first sign that an AI agent ever touched their systems. Most security teams watch for human attackers and known malware, not for a model that wandered off during someone else’s test and started scraping pages or trying logins.

Rogue Agents are Now an Industry-Wide Concern

Anthropic has disclosed three separate breaches by its own models during testing, which it traced to an outside company that set up its test environments with internet access by mistake, NPR reported. A string of similar incidents around the world in recent months has raised doubts inside the industry about whether developers can keep control of the more powerful models they are now building.

The scrutiny has reached Washington too, and on Tuesday OpenAI, Anthropic, Google, Meta, and Nvidia signed a voluntary accord at the White House that commits them to independent audits and to making sure their tools do not “hack or access technical systems in unintended ways.” A Reuters/Ipsos poll taken in mid-September found that 73% of Americans worry AI companies have not done enough to keep the technology from causing serious harm.

OpenAI has not said how many more notices it expects to send before the review is finished.

Related Articles

Anthropic’s IPO Filing Details Losses, Founder Control, and AI Safety Risks

Exclusive: Nikita Astionov on the Security Risks of Giving AI Agents More Control

An Exposed Server Revealed How a Ransomware Affiliate Used AI to Plan Attacks

Share this article

Related Articles