Your AI assistant does what it reads
Set up connected AI tools so a hostile email, web page or shared document cannot turn them against you, and know which steps still need your eyes.
on this page · 0 / 0 checked
You gave an assistant access to your inbox so it could clear the morning for you. To do that job it has to read mail from people you have never met. Every one of those messages is text, and text is the only thing an assistant knows how to act on. That is the feature. It is also the hole.
There is no reliable line inside a language model between the instructions you typed and the words it found in a message, a shared document, a calendar invite or a web page. OWASP puts this first on its list of language model risks, as LLM01 for 2025, and notes that hostile inputs “can affect the model even if they are imperceptible to humans, therefore prompt injections do not need to be human-visible/readable” [1]. This guide is for a solo operator or a small team who has already connected an assistant to real accounts and wants to keep the useful part. It is not for people building an agent platform, and it will not get you through an enterprise security review.
An assistant cannot tell an order from a document
Software security has leaned for decades on one assumption: the program knows which part of its input is a command and which part is data. A database keeps the query separate from the values. A shell knows the program name from its arguments. A language model has no such boundary. Everything arrives as one stream of text and the model decides what to do with all of it.
OWASP defines prompt injection as occurring “when user prompts alter the LLM’s behavior or output in unintended ways” and splits it in two [1]. Direct injection is what somebody types into the chat box. Indirect injection is the one that should worry you, and OWASP describes it as what happens when “an LLM accepts input from external sources, such as websites or files” [1]. The hostile text can sit in white type on a white background, inside an HTML comment, in a document’s tracked changes, in an image’s alt text, or in the invisible text layer of a PDF. You will not see it. The assistant will.
That is why filtering does not close this. There is no character sequence to block. The payload is written in ordinary English, in a language with unlimited ways to say the same thing, and it is indistinguishable in form from the legitimate request sitting next to it in the same window.
Three ingredients have to be present at once
Reading a hostile instruction is not by itself a loss. Damage needs the assistant that read it to also hold something worth taking and to have some way of sending it out. The developer Simon Willison named that combination the lethal trifecta in June 2025: access to private data, exposure to untrusted content, and the ability to communicate externally [8].
Microsoft 365 Copilot showed the complete set in a single bug. CVE-2025-32711, published on 11 June 2025, is recorded as “Ai command injection in M365 Copilot allows an unauthorized attacker to disclose information over a network” and carries a CVSS base score of 9.3, rated critical [6]. Read that description as a parts list. An unauthorized attacker supplied the content. Copilot held information worth disclosing. The network was the way out. Microsoft, which assigned the record itself, classified the flaw under CWE-74, improper neutralization of special elements in output used by a downstream component [6]. That is the same family as SQL injection and command injection, and it is that family for the same reason: nothing enforced the line between instruction and data.
That chain has a CVE number and a vendor who owns it. The shape underneath has neither, and the shape is now the default setup for most people using this software. Turn on a mail connector and a browsing tool in the same assistant and you have assembled all three ingredients yourself, in about ninety seconds, using buttons the vendor put there on purpose.
The vendors have measured this and the number is not zero
Anthropic published figures when it launched Claude for Chrome on 25 August 2025. Across “123 test cases representing 29 different attack scenarios”, deliberately targeted attacks against browser use succeeded 23.6% of the time with no safety mitigations in place. With mitigations, that fell to 11.2% [2]. On “a ‘challenge’ set of four browser-specific attack types”, the defenses took the success rate from 35.7% to 0% [2]. The extension reached Max subscribers on 24 November 2025 and Pro, Team and Enterprise plans on 18 December 2025 [2].
Google described a five-layer defense for Gemini on 13 June 2025, made up of prompt injection content classifiers, security thought reinforcement, markdown sanitization and suspicious URL redaction, a user confirmation framework, and end-user security mitigation notifications [4]. OpenAI was blunter still at the launch of ChatGPT Atlas on 21 October 2025, writing that agents are “susceptible to hidden malicious instructions, which may be hidden in places such as a webpage or email” and that “our safeguards will not stop every attack that emerges as AI agents grow in popularity” [3].
Read those two things together. An 11.2% success rate under deliberate attack is a genuine improvement over 23.6%, and it is also about one in nine. Anthropic’s own documentation for Claude Code states the position plainly: “no system is completely immune to all attacks” [5].
The 11.2% default is Anthropic's measured success rate for deliberately targeted attacks on Claude for Chrome with mitigations enabled [2]. It is a rate for attacks that are actually attempted, not a rate for ordinary browsing. Computed in the page; nothing is sent anywhere.
Take one of the three ingredients away
You are not going to beat this at the model layer, and you do not have to. Every one of these attacks needs all three ingredients. Remove any one and the chain breaks, which is a job you can do with settings you already have access to.
Split the accounts. The assistant that reads inbound mail should not be the same assistant that has your bank connector, your CRM write access and a browsing tool. ChatGPT’s logged-out mode exists for this, and OpenAI puts it in those terms: “You can use agent in logged out mode to limit its access to sensitive data and the risk of it taking actions as you on websites” [3]. Claude for Chrome offers site-level permissions so you can grant or revoke access per website, and it blocks whole high-risk categories including financial services, adult content and pirated content [2].
Cut the scopes to what the job needs. The Model Context Protocol security guidance is explicit that broad tokens create an “expanded blast radius” in which a “stolen broad token enables unrelated tool/resource access”, and it lists “wildcard or omnibus scopes (*, all, full-access)” as a common mistake [7]. When a connector’s consent screen asks for read and write, and the assistant only summarizes, grant read.
Keep the reading and the sending apart in your automations too. If you have a flow in Zapier, n8n or Make that pipes an inbound form submission or email into a model step and then into a send step, you have rebuilt the trifecta in three boxes. Put a fixed template on the outbound step so the model fills fields rather than choosing recipients and content.
The same logic already shows up in tooling. In its Manual mode, Claude Code starts with read-only permissions, can write only to the folder it was started in and its subfolders, and does not auto-approve web-fetching commands such as curl and wget; web fetch also “uses a separate context window to avoid injecting potentially malicious prompts” [5]. Its own best-practice list tells you to “avoid piping untrusted content directly to Claude” [5]. If you point Cursor or any coding agent at a repository you did not write, the README, the issue text and the dependency docs are all untrusted content by this definition.
Put the approval where the damage would be
Confirmation only helps at the step that cannot be undone. Reviewing a draft summary is theatre. Reviewing the moment money moves, a file is shared outside the company, a record is deleted or a message is sent is the whole control.
The vendors have converged on this. Claude for Chrome “asks users before taking high-risk actions like publishing, purchasing, or sharing personal data” [2]. Google’s user confirmation framework requires explicit approval for risky operations, giving the deletion of a calendar event as its example [4]. In Manual mode, Claude Code asks before it edits files, runs tests or executes commands that modify your system [5]. OWASP’s mitigation list says the same thing in the abstract: restrict access to the minimum necessary privileges, require human review before sensitive operations, and clearly distinguish and isolate untrusted external content [1].
The failure mode here is fatigue. Anthropic names it directly and ships allowlisting as “prompt fatigue mitigation” [5]. If you approve twelve prompts an hour you will approve the thirteenth without reading it, so allowlist the reversible things aggressively and leave the irreversible ones prompting. A rule you actually read is worth more than five you have trained yourself to click through. Know which mode you are in before you trust the prompts, too: Claude Code also offers an auto mode in which “a separate classifier model reviews actions instead of you and blocks the ones it judges unsafe” [5]. That is a different safety story from your own eyes on the step, and it is worth choosing on purpose rather than inheriting.
Notice when the assistant has done something you did not ask
Injection shows up as activity you did not initiate. Check the sent folder and the drafts folder of any mailbox an assistant can write to, and glance at the activity log of connected accounts. Review the OAuth grants on your Google and Microsoft accounts on a schedule and revoke connectors you stopped using, because a forgotten grant keeps its scopes indefinitely.
Some of this is being handed to you. Google’s fifth defense layer is end-user security mitigation notifications, built on the principle of “sharing details on attacks that we’ve stopped so users can watch out for similar attacks in the future” [4]. Claude Code requires trust verification the first time it runs in a codebase and when a new MCP server is added, though the documentation notes that this verification is disabled when running non-interactively with the -p flag [5]. The MCP specification requires that a client offering one-click local server setup show “the exact command that will be executed, without truncation” before running it [7]. Read that command. It is the one moment where an attacker’s instructions are still visible to you as text.
What still goes wrong
The honest limit is that none of this detects an attack. It reduces what an attack can do. A well-configured setup still lets a hostile page waste your time, produce a wrong summary, or quietly change a number in a document you then act on, and you will have no notification and no log entry, because from the tool’s point of view nothing unusual happened. Integrity damage is much harder to catch than data theft, and nobody is shipping a good answer to it.
Read-only is also weaker protection than it sounds. An assistant that can only read your data but can still fetch a URL can leak that data through the URL it fetches, which is why the Microsoft 365 Copilot flaw was recorded as disclosure of information over a network rather than as a write of any kind [6]. Google’s markdown sanitization and suspicious URL redaction exist precisely because of that path [4]. If a tool can reach the open internet at all, treat it as having an outbound channel regardless of what its permission screen says about writing.
And the numbers move. The 11.2% figure was measured on one product, in one test set, in 2025 [2]. Individual exploit chains get names, CVE records and owners, as that Copilot flaw did [6]. The gap they exploit gets none of those, because it is a property of how these models read text rather than a defect somebody can close. Plan on the assumption that this stays partly unsolved. Give the assistant what the job needs, keep it away from the combination that hurts, and keep your hand on the steps you would not want reversed.
- 01OWASP — LLM01:2025 Prompt Injectiongenai.owasp.org
- 02Anthropic — Claude for Chromeclaude.com
- 03OpenAI — Introducing ChatGPT Atlasopenai.com
- 04Google — Mitigating prompt injection attacks with a layered defense strategyblog.google
- 05Anthropic — Claude Code securitycode.claude.com
- 06CVE Record — CVE-2025-32711 (Microsoft 365 Copilot)cveawg.mitre.org
- 07Model Context Protocol — Security Best Practicesmodelcontextprotocol.io
- 08Simon Willison — The lethal trifecta for AI agentssimonwillison.net