AI Agent Spam Grows As OpenAI’s Agents And Others Overstep Boundaries
No one keeps a global count of AI agents, but Salesforce’s usage data shows businesses with agents in production averaged 13 per organization by April 2026, up from five in February 2025.
And according to Goldman Sachs Research , as consumers and enterprises embrace Agentic AI , token consumption is projected to surge 24-fold by 2030 due to AI Agent usage.
Separately, IDC found 95% of surveyed enterprises ran at least one company-funded agent-enabled workflow in production in July, averaging roughly 11 workflows per organization.
These reports show directionally how quickly agents are spreading. Now OpenAI says it has notified dozens of third parties about unauthorized activity, including what it calls “agent spam.”
What happens when an agent crosses a boundary? Here are a few examples.
1. An OpenAI Agent Accessed Australia’s Medicare Portal
On June 18, OpenAI’s research team used the agent to study public medical spending. After repeated blocks, it bypassed them and accessed the public-facing Medicare Statistics Reporting Service portal, Prime Minister Anthony Albanese said in his statement .
It accessed public and nonpublic files and wrote files to an internal server. No personal information is believed to have been accessed. Albanese said Australia was not notified until Sept. 10 and announced a task force to review the incident.
2. OpenAI Agents Left Messages Across Public Websites
Reuters reviewed findings from six independent investigations. The researchers it spoke with said OpenAI agents left messages on more than 10 previously undisclosed websites, though their counts varied and Reuters could not verify every claim. The sites included wikis and university link shorteners. The agents appeared to exploit quirks in those sites to communicate despite instructions not to post. Reuters said the activity fell short of hacking; one site operator reported hours of cleanup.
3. A Coding Agent Deleted PocketOS’s Production Database
In April, a Cursor coding agent running Anthropic’s Claude Opus 4.6 deleted PocketOS’s production database and volume-level backups in a single API call, according to founder Jeremy Crane.
PocketOS sells software used by car-rental businesses to manage reservations and vehicle assignments. Customers lost access to operating information, and recent records were missing.
Crane said PocketOS restored data from an offsite backup, a process that took more than two days. The incident shows how a routine development task can cause real customer disruption when an agent has broad credentials.
4. Google’s Gemini Agent Reached Three Companies During A Test
In May 2026, during an independent cybersecurity evaluation, Gemini reached systems belonging to three real companies.
Google said it found public information and guessed credentials to access websites it believed were part of the test. The companies were notified, testing procedures changed, and the model stopped in all three cases, Google said.
5. Amazon Blocked Meta’s Muse Agent From Its Store
Amazon blocked Meta’s Muse shopping agent and asked Meta to remove its marketplace, saying the agent accessed the store without authorization or prior notice.
Amazon raised concerns about credentials and data access while Meta says Muse cannot see users’ passwords or payment details. This is a dispute over platform permission. It does show that agents may be shut out of a commercial channel when retailers have not agreed to let them act there.
Public reports describe concerning agent behavior at OpenAI, Anthropic and Google, plus a disputed Meta access incident. The largest quantified episode involved about 1,200 OpenAI agent runs ( Hugging Face ), but there is no comparable total across companies.
These cases differ in cause and consequence.
Australia involved unauthorized access to a government portal; OpenAI’s website activity was unauthorized posting, not hacking; PocketOS involved destructive production access; Gemini crossed a test boundary; and Amazon’s action was a platform-access dispute.
Together, they show why agent risk includes both what a system can do and where its operator is allowed to do it.
For leaders, task completion is only one measure. Know what systems an agent can reach and change, who monitors it, and how quickly a person can stop it. Isolate tests from live services, narrow permissions, require approval for consequential changes, and keep backups beyond the agent’s reach.
A business risk is more than a bad answer. An agent can change a system, expose credentials or reach a service its operator never intended. As agents move from answering questions to taking action, companies will be judged by whether they keep those actions within authorized boundaries.