Are AI Browsers Safe? A Single Web Page Can Hijack Them
Are AI browsers safe? A malicious instruction on a page or in an email can redirect an agent operating inside your logged-in accounts.
In one red-team test , OpenAI asked its browser agent to write an out-of-office reply. The agent opened an unread email, followed an instruction embedded in the message and sent a resignation letter to the user's boss instead. No employee lost a job. OpenAI created the attack itself and says its updated system now flags it. The test raises the central question in AI browser safety: what can the agent do after it reads an instruction you never gave it?
What “Hijack” Means In An AI Browser
An AI browser does more than explain a page. Its built-in agent can navigate sites, fill forms, click buttons and work inside accounts on your behalf. Examples include OpenAI's ChatGPT Atlas and Perplexity's Comet. A regular browser waits for your next click. An AI browser can take the next step for you.
Security researchers call this indirect prompt injection: instructions planted in a page, email or document that an AI system reads and mistakes for commands. The browser may then follow those commands using the access you already gave it.
“Hijack” has a narrow meaning here. It doesn’t mean the attacker automatically controls your computer or every account. It means a page or message tricks the browser agent into doing something you didn’t ask it to do, using access you already granted.
Why Prompt Injection Is Hard To Stop
OpenAI calls prompt injection an open challenge for agent security. The UK's cybersecurity agency warns that it may never be possible to block completely. An AI system must process both your directions and the material it finds online, and it may confuse one for the other.
Researchers at Brave , which is also a browser maker, demonstrated the problem with Perplexity's Comet. They placed an instruction in a Reddit comment and asked Comet to summarize the page. The agent followed the planted instruction, moved across logged-in services and exposed information from the user's account. Brave said Perplexity fixed the specific Reddit exploit, although the broader attack had not been fully addressed.
This was controlled security research, not a report of a consumer losing money. It still shows why an ordinary page becomes more dangerous when the software reading it can also act elsewhere.
What The Browser Can Do Determines The Risk
An agent with access to email can expose private information. If it can also send messages, approve purchases or change records, it may act on an attacker’s instructions. The more it’s allowed to do, the greater the possible damage.
OpenAI recommends using Atlas in logged-out mode when a task does not require an account. It also says Atlas asks for confirmation before consequential steps such as sending an email or completing a purchase. Those controls reduce risk, but broad, permanent access still isn’t a good default.
Responsibility is less clear. In Moffatt v. Air Canada , a Canadian tribunal held the airline responsible for inaccurate information supplied by a chatbot on its website. That case didn’t involve an AI browser or prompt injection. The broader lesson is that a company can still be held responsible for bad information from its automated system. It doesn’t settle who bears a loss when an independent browser agent follows a stranger's instruction.
Make The Browser Ask Before It Acts
Think of an AI browser’s access in three levels: read, prepare and execute. Read lets it inspect information, including private material. Prepare lets it draft a message or fill a shopping cart without submitting it. Execute lets it send, buy, delete, publish or transfer.
Require a new confirmation before the browser takes any action. Limit what it can read to the information needed for the current job. Give it access only to what it needs for the task, and no longer than it needs it.
The risk in an AI browser is measured by what it can do after it reads the wrong instruction. Are AI browsers safe? Browser makers' defenses matter, but the access you grant determines how far a mistake can go. Keep that reach narrow, one task at a time.
Loading article...