CONTROLLED EVALUATION
OpenAI Operator: indirect prompt injection becomes an action risk
OpenAI’s published red-team work shows how untrusted content can influence a browser-using agent that can operate tools or move data.
What this was
Controlled evaluation / research-preview safety testing—not a reported customer breach, autonomous escape, or compromise.
What happened
OpenAI’s January 2025 Operator System Card documented internal and external red teaming of its browser-using agent against adversarial instructions placed in test websites, emails, databases, and other mock environments. OpenAI reported 23% susceptibility across 31 automatically checkable final-model scenarios, compared with 62% with no mitigations and 47% with prompting alone. A separate prompt-injection monitor achieved 99% recall and 90% precision on 77 red-team-created attempts. OpenAI states that known cases were mitigated while prompt injection remains an ongoing concern. These are evaluation findings, not evidence that an attacker breached a customer or that Operator escaped control.
Why an agent changes the risk
A conventional application can render untrusted content; an agent may interpret that same content, combine it with private context, and use connected tools. The risk is the path from attacker-controlled instruction to a sensitive action or data flow.
How to test for this
- Red-team indirect prompt injection end to end: seed malicious instructions in representative webpages, email, documents, and tool outputs, then measure both agent susceptibility and monitor performance.
- Keep a strict trust boundary: treat external content as data, never developer instruction, and pass only validated, schema-constrained fields between agent nodes and tools.
- Exercise sensitive tool paths with least privilege, watch mode, and human approval gates—attempting unauthorized reads, writes, exports, and high-impact actions triggered by untrusted content.
Other cases
Get a scoped price without a discovery call
Tell us what is in scope and what your audit needs. You get a fixed price and a date, not a quote after two meetings.