tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

Your coding agent will follow instructions it finds in your tools

Set up Claude Code, Cursor or Codex so the text your agent reads cannot spend your credentials, and know which default settings are worth keeping.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

You connected the agent to your tools because copying text between windows was the slow part. Now it reads the error tracker itself, opens the issue, searches the logs, and proposes a fix while you make coffee. That is the entire value of the setup. It is also the entire problem, because almost none of the text it now reads was written by you.

This is not a flaw in one product, and it is not something a patch closes. All three of the major coding agents document it in their own security pages, and none of them claims to have solved it. Anthropic’s warning is that while its protections significantly reduce risk, “no system is completely immune to all attacks” [1]. Cursor describes its own Run Modes as “best-effort guardrails rather than a hard security boundary” [4]. What follows is the operator’s version: which settings actually change the outcome, and which convenience is worth buying. If you are writing a threat model for a regulated platform, you want the standards themselves, not this page.

Everything the agent reads arrives as instructions

A language model does not have a channel for “this is the job” and a separate channel for “this is just material”. It has one stream of text. OWASP’s top risk for LLM applications splits the resulting attack in two: direct injection, which occurs when a user’s prompt input directly alters the behaviour of the model in unintended ways, and indirect injection, which occurs when the model accepts input from external sources such as websites or files [6]. Coding agents are indirect injection’s natural home. An issue title, a stack trace, a dependency’s README, a comment on a pull request, the body of a page the agent fetched: all of it is text, all of it is written by someone else, and any of it can be phrased as an instruction.

The clean demonstration was published on 17 June 2026. The security firm Tenet planted a fake error event in Sentry using a publicly exposed reporting key, formatted as markdown that mimicked legitimate Sentry templates, including a fabricated “Resolution” section containing an npx command dressed up as diagnostic guidance. A developer asks their agent to look into the error, the agent reads the planted text as guidance, and it runs the attacker’s package with the developer’s full privileges. Tenet reported an 85% exploitation success rate against injected errors across the most widely used agents, confirmed the behaviour in Claude Code, Cursor and Codex, and found 2,388 organisations with valid injectable keys. Sentry’s leadership acknowledged the report on 3 June 2026 and declined a root-level fix, calling the attack “technically not defensible” on their side [7].

They were right about their side. Nothing in that sequence looks like an intrusion. No stolen password, no unusual login, no firewall event. Every step is authorised work by a tool you installed, on data it was designed to read. So the defence cannot live in the tool that got poisoned. It lives in what your agent is allowed to do after it reads.

The damage needs three things at once

Reading hostile text is harmless on its own. An attack only pays off when the agent can also reach something worth taking, and then send it somewhere. Simon Willison named this combination the lethal trifecta: access to your private data, exposure to untrusted content, and the ability to externally communicate. His argument is that we still do not know how to prevent this reliably, that guardrail products advertising something like 95% detection are quoting a failing grade for security work, and that the answer is to avoid the combination entirely rather than to filter for bad inputs [8]. Anthropic’s sandboxing documentation reaches the same conclusion from the engineering side: effective sandboxing requires both filesystem and network isolation, because without network isolation a compromised agent could exfiltrate sensitive files like SSH keys, and without filesystem isolation it could backdoor system resources to gain network access [3].

The practical use of this is that you rarely need all three at full strength in one session. When you are debugging a live error, the agent needs to read from a tool an outsider can write into, so tighten what it can run and where it can send data. When you are refactoring a local module with no external inputs, the exposure is low and you can afford a looser hand. The mistake is running one permanent configuration, tuned for the loosest task you ever do, and then pointing it at the error tracker.

Permission modes are the dial that matters

Every serious agent ships a set of modes that decide what runs without asking you. In Claude Code’s Manual mode, the agent starts read-only, asks before it edits files or runs commands that can modify your system, and runs only a built-in set of read-only commands such as ls, cat and git status without a prompt [1]. Commands that fetch content from the web, curl and wget among them, are not auto-approved by default [1]. Codex documents workspace-write as “the default low-friction mode for local work”, paired with an on-request approval policy, and under that pairing the agent asks before using the internet or going beyond the workspace boundary [5]. Cursor’s default is blunter: terminal commands need your approval, and its tools only make network requests to GitHub, direct link retrieval and web search providers [4].

Check which mode you are actually in before you rely on any of that, because the shipped default has moved. On Pro, Max and Team plans, the built-in starting permission mode for Claude Code in a terminal or the VS Code extension is auto mode, not Manual [2]. In auto mode a second model, the classifier, reviews actions in your place, blocking anything that escalates beyond your request, targets unrecognised infrastructure, or appears driven by hostile content Claude read [2]. Reads and working-directory edits skip it; shell commands and network operations are what it sees [2]. Its default block list is aimed squarely at the attack above, covering downloading and executing code such as curl | bash, sending sensitive data to external endpoints, and production deploys [2]. Codex offers the equivalent as “Approve for me for eligible approval requests”, where the approvals_reviewer setting routes eligible prompts to a reviewer agent instead of to you [5]. This is a real defence. It is also a model judging a model, and Anthropic’s own warning is that auto mode reduces prompts but does not guarantee safety [2].

Above that sits the mode with no brakes. Claude Code’s bypassPermissions, reached through --dangerously-skip-permissions, is documented for “Isolated containers and VMs only” [2]. Codex calls its version full access and says danger-full-access removes the filesystem and network boundaries and should be used only when you want the agent to act with full access [5]. Both are legitimate settings for a throwaway container that holds nothing and can reach nothing. Neither is a setting for the laptop with your SSH keys, your production credentials and your client work on it. If the prompts are driving you to that mode, the answer is a sandbox or a narrower task, not a bigger hammer.

A sandbox bounds the mistake without preventing it

The better trade is to keep approval on for consequential actions and move boundary enforcement out of the model entirely. Claude Code’s /sandbox does this for Bash: you define which files and network domains commands may touch, and the operating system enforces that boundary for every command and its child processes, on macOS, Linux and WSL2 [3]. Native Windows is not supported, so on Windows you run inside WSL2 [3]. The sandbox pre-allows no domains at all, prompting the first time a command needs a new one [3]. Codex’s workspace-write is the same idea with different defaults, letting the agent read files, edit within the workspace and run routine local commands inside that boundary [5]. In both cases the agent can still be fooled. What changes is that being fooled no longer reaches your home directory.

Read the limits before you rely on it. Anthropic states plainly that sandboxing reduces risk but is not a complete isolation boundary [3]. The built-in proxy makes its allow decision from the client-supplied hostname and by default does not terminate or inspect TLS, so code running inside the sandbox can potentially use domain fronting or similar techniques to reach hosts outside the allowlist, and allowing broad domains such as github.com can itself create paths for data exfiltration [3]. A network allowlist containing the one host every developer needs is a network allowlist with an exit in it. Worse, sandboxed Bash commands inherit the parent process environment by default, including any credentials set there, unless you unset or mask specific variables through sandbox.credentials or scrub them from subprocesses [3]. The same settings block take deny or mask entries for credential directories such as ~/.aws and ~/.ssh, which the default read policy otherwise allows [3]. Setting those is the difference between a box and a box with your keys in it.

Use a sandbox as a blast-radius control rather than a wall. It converts “the agent ran an attacker’s install script with your credentials” into “the agent ran it inside a box holding one repository”. Large improvement, not immunity.

Fewer connections, narrower scopes

Every tool you connect is another party that can put text in front of your agent. The question you are used to asking about an integration is what it can see and change. Add a second one: who can write into this, and what would the agent do if they wrote instructions instead of data. An error tracker, an issue board and a shared inbox all fail that question, because the entire point of them is that outsiders can add entries.

So connect fewer, and scope each one down. OWASP’s mitigations for this risk are, in order, constraining model behaviour, defining and validating expected output formats, input and output filtering, enforcing privilege control and least privilege access, requiring human approval for high-risk actions, segregating and identifying external content, and adversarial testing [6]. The two that a small team can actually implement are least privilege and human approval. Read-only beats read-write. A key scoped to one project beats an admin key. Cursor supports pre-approving specific tools with an MCP allowlist, but the default is stricter and worth keeping: all MCP connections need your approval, and after you approve a connection, each tool call still needs individual approval before it runs [4]. Cursor also honours a .cursorignore file to block agent access to specific files, which is the cheapest way to keep a credentials directory out of the conversation [4].

Be deliberate about where those connectors come from. Anthropic reviews connectors against its listing criteria before adding them to its directory, but the documentation is explicit that it does not security-audit or manage any MCP server [1]. Claude Code asks for trust verification the first time it runs in a codebase and when you add a new MCP server, and that verification is disabled when running non-interactively with the -p flag [1]. If you have a script that pipes work to an agent unattended, you have opted out of the prompt that would have caught an unfamiliar server.

Deny rules outlast anything you type in chat

There is one detail that surprises people, and it decides whether your guardrails survive a long session. Telling the agent “do not push” works, in the sense that Claude Code’s classifier treats a boundary you state in the conversation as a block signal and blocks matching actions even when its default rules would allow them [2]. But those boundaries are not stored as rules. The classifier re-reads them from the transcript on every check, so a boundary can be lost if context compaction removes the message that stated it, and Anthropic’s guidance is direct: for a hard guarantee, add a deny rule instead [2].

Deny rules behave the way you want a rule to behave. They block in every mode, including bypassPermissions, where allow rules have no effect at all [2]. The five minutes you spend writing three of them into your project settings is the highest-value work in this guide, because it is the only part that does not depend on a model remembering something you said an hour ago.

Watch the other end too. If Claude Code’s classifier blocks an action 3 times in a row or 20 times in total, auto mode pauses and the agent goes back to prompting you, and those thresholds are not configurable [2]. Treat that pause as information. Repeated blocks usually mean the classifier is missing context about your infrastructure, and sometimes mean the agent is trying to do something that no one in the room asked for.

checklist
Before you point an agent at a tool outsiders can write into
0 of 8 · saved in this browser only
calculator
What manual approval actually costs
— h / month

approvals × seconds × 22 working days. Computed in the page; nothing is sent anywhere.

What still goes wrong

None of this closes the hole, and you should hold that clearly rather than hopefully. OWASP’s own assessment is that given the stochastic influence at the heart of the way models work, it is unclear whether there are fool-proof methods of prevention for prompt injection [6]. The mitigations reduce impact. They do not deliver a guarantee, and any vendor claiming otherwise is selling something.

The specific weak points are worth naming. The reviewing classifier is better insulated than people assume, because tool results are stripped before it reads, so hostile content in a file or web page cannot manipulate it directly [2]. That is insulation, not immunity: it is still a model, and it still has to judge a command whose purpose was set by text it never saw. The sandbox’s network layer decides from a hostname without inspecting TLS by default, so a permissive allowlist is a hole with paperwork [3]. Workspace trust in Cursor is disabled by default, and turning it on has a real cost, because restricted mode breaks AI features [4]. Then there is the failure the vendors design around rather than solve. Anthropic ships allowlisting specifically as prompt-fatigue mitigation [1], which is a tacit admission that the volume of approvals is itself the risk: a person clicking through 200 prompts a week is not reviewing 200 prompts. That is the argument for fewer connected tools rather than more vigilance.

And there is a category this guide does not cover. If your agent runs unattended in CI, on a schedule, or as part of a pipeline, the human-approval defence is gone by construction, and Anthropic notes that trust verification is skipped in non-interactive runs [1]. That setup needs isolation as its primary control rather than its backup: a container with no standing credentials, an exact allowlist of the tools the job may call, which is what Claude Code’s dontAsk mode is built for [2], and the assumption that anything the agent can reach, an attacker can eventually reach through it. Anthropic’s own best practices for untrusted content include using virtual machines to run scripts and make tool calls, especially when interacting with external web services [1]. On a laptop that holds your client work, that is the advice worth taking literally.

sources
  1. 01Anthropic — Claude Code securitycode.claude.com
  2. 02Anthropic — Choose a permission modecode.claude.com
  3. 03Anthropic — Configure the sandboxed Bash toolcode.claude.com
  4. 04Cursor — Agent securitycursor.com
  5. 05OpenAI — Codex sandboxing and approvalslearn.chatgpt.com
  6. 06OWASP GenAI — LLM01:2025 Prompt Injectiongenai.owasp.org
  7. 07Tenet Security — Agentjacking coding agents with fake Sentry errorstenetsecurity.ai
  8. 08Simon Willison — The lethal trifecta for AI agentssimonwillison.net
next guide
Reverse federalism, and the AI rules that reach a business your size
9 min · verified 2026-09-05
related guides