tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · running the business

How to secure an MCP connection

A procedure for connecting MCP servers to real accounts: decide trust at the server, scope the credential, separate read from write, and keep a record you can revoke.

Published 2026-09-05 · Updated 2026-09-05 · Read 10 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

Connecting an MCP server takes about ninety seconds. You paste a URL or a command, a browser tab opens, you click the button that says you agree, and the model can now read your calendar, your repository or your customer records. Nothing in that flow tells you what you just agreed to, because the flow was designed to get you to a working demo, not to a defensible configuration.

The gap between those two is smaller than it looks, and it is mostly made of decisions you take once. This guide is the order to take them in: work out whether the server itself is trustworthy, scope the credential you hand it, keep read separated from write, treat the server’s output as hostile, and leave yourself an inventory and a log. It is written for a one-person business or a small team wiring MCP servers into accounts that hold real money or real client data. If you are running a security programme with a threat model and a red team, this is beneath you. If you are experimenting on a throwaway account with nothing in it, it is above you, and you should go and experiment.

The server is the trust boundary, not the individual tool

Every permission prompt you will ever click is downstream of one decision, which is whether to connect this server at all. The specification says as much. MCP’s current revision, 2026-07-28, states that “descriptions of tool behavior such as annotations should be considered untrusted, unless obtained from a trusted server” [2]. Read that backwards and you have the whole problem. Annotations are what your client uses to decide how carefully to handle a tool, and they only mean anything if you have already decided the server is honest.

In Claude that chain is explicit. Every submitted tool must carry a title and the applicable hint, readOnlyHint for read-only tools and destructiveHint for tools that modify or delete data, and Anthropic’s connector review criteria say those hints “determine auto-permissions in Claude: read-only tools can run without per-call confirmation; destructive tools always prompt” [5]. The server declares which of its own tools are harmless. The client takes it at face value. That is a sensible design when the server is honest and it is the entire attack when it is not.

Directory listing is a weaker signal than it feels like. Anthropic scans submissions automatically for policy compliance and lists them by default as community connectors, escalating some to a verified review in which reviewers “run a functional test of each tool” [5]. A functional test confirms the tool works, not that it behaves. Anthropic’s own Claude Code documentation is blunt about the limit: it “reviews connectors against its listing criteria before adding them to the Anthropic Directory, but does not security-audit or manage any MCP server” [4]. OpenAI puts the same burden on the person clicking the button, warning that connecting to unsafe or untrusted MCP servers “may increase exposure to security risks (including prompt injection)” and telling you to connect only servers you trust [7]. Anthropic’s connector documentation warns that custom connectors let you reach services “that haven’t been verified by Anthropic” and says to connect only to servers from trusted organisations, reviewing authentication permissions carefully [6].

So the useful question at install time is not “which tools should I enable”. It is “who wrote this, do they run the API it talks to, and would I give this organisation a password”. Anthropic requires directory servers to call their own first-party APIs, or APIs they legitimately proxy, with the server domain matching the service [5]. Apply that test yourself to anything installed from outside a directory. A server for a product, published by the company that makes the product, is a different proposition from one published by a stranger, even if the stranger’s is better written.

Scope the credential before you scope anything else

When you authorise a connector, you are usually handing over your own access. Anthropic describes connectors as inheriting each person’s permissions from the connected service, so that “if someone can’t access a specific file, channel, or record in the source system, the connector can’t reach it from Claude either” [6]. That sounds reassuring until you notice the corollary: if you are an admin, the connector is an admin. The ceiling on what a compromised or manipulated agent can do is the ceiling on what your account can do.

The fix is boring and it works. Create a dedicated account or a dedicated key for the integration rather than connecting as yourself, and give it access to one project, one repository, one folder, one database. Where the platform lets you pick scopes at consent time, pick the smallest set that makes the job run. The MCP security guidance is explicit that broad grants are the failure: it lists publishing every possible scope, using wildcard or omnibus scopes such as *, all and full-access, and “bundling unrelated privileges to preempt future prompts” among the common mistakes [1]. It recommends a minimal initial scope set covering only low-risk discovery and read operations, with incremental elevation when privileged operations are first attempted [1].

There is a second-order version of this that matters if you or a contractor build the server. Token passthrough is forbidden outright: MCP servers “MUST NOT accept any tokens that were not explicitly issued for the MCP server” [1]. A server that accepts whatever token it is handed and forwards it downstream destroys the audit trail, because the downstream resource server’s logs “may show requests that appear to come from a different source with a different identity” [1]. If you have commissioned an internal MCP server, that sentence is the one to put in the brief.

Read and write are two different decisions

The most useful discipline available to a small team is refusing to grant write access on day one. Read-only integrations fail quietly. Write integrations fail by sending an email, deleting a record or moving money.

The vendors have started enforcing the split rather than suggesting it. Anthropic rejects any directory connector that ships a catch-all tool accepting both safe HTTP methods and unsafe ones, and requires read operations and write operations to live in separate tools, ideally split further by create, update and delete [5]. Documenting the difference inside a single tool’s description does not satisfy the requirement [5]. OpenAI gates full MCP support including write actions behind developer mode, which is rolling out in beta to ChatGPT Business, Enterprise and Edu plans and on web only; Pro users can connect MCP servers with read and fetch permissions but not the full write path [7]. ChatGPT may ask for confirmation before important actions depending on app permissions and the action’s context, and OpenAI notes that “some especially risky actions may be blocked instead of being presented for approval” [7].

Take the same position on your own setup. Connect read-only, run the work you actually intended to run for a week, and only then decide which single write tool earns its place. Where your client lets you fix the decision in configuration rather than at the prompt, use it: Claude Code keeps the list of allowed MCP servers in settings that engineers check into source control, which makes the list reviewable rather than remembered [4].

Be honest about what a confirmation prompt buys you. It is a real control on the tenth call of the day and a rubber stamp on the hundredth. Prompt fatigue is a named problem in Claude Code’s own documentation, which is why it supports allowlisting frequently used safe commands per user, per codebase or per organisation [4]. The scope is the control; the prompt is only the reminder.

A local MCP server is a program you agreed to run

Remote connectors reach your data. Local servers reach your machine. The distinction gets lost because both are one line in a config file, and one of those lines is a shell command that runs with your privileges.

The security guidance treats this as its own attack class. Where a client offers one-click local server configuration it “MUST implement proper consent mechanisms prior to executing commands”, showing the exact command that will be executed without truncation, identifying it as a potentially dangerous operation that executes code on your system, and requiring explicit approval [1]. The named risks are arbitrary code execution with client privileges, no visibility into what is being run, obfuscated commands that appear legitimate, and data loss [1]. The recommended shape is a sandbox: restricted access to filesystem, network and other system resources, minimal default privileges, and platform-appropriate sandboxing technologies [1]. The NSA’s May 2026 information sheet on MCP makes the same recommendation in operating-system terms, saying frameworks such as AppContainers, seccomp, AppArmor or SELinux “should be used to isolate each tool’s execution context”, and it points at CVE-2025-49596 in MCP Inspector as a worked example, a disclosed vulnerability allowing malicious actors “to trigger remote code execution via crafted messages” [3].

Two practical consequences. First, read the command before you approve it, particularly the package name, because a command that fetches and runs a package is a command that runs whatever that package contains this week. The NSA sheet recommends choosing supported MCP projects, noting that many popular servers are no longer actively maintained, keeping a clear inventory of everything deployed along with versioning and patch history, and running a formal process for monitoring MCP-related vulnerabilities through security feeds and threat intelligence [3]. Pinning a version is the cheap half of that. Second, prefer the stdio transport for anything running locally, which the security guidance recommends specifically to limit access to just the MCP client rather than exposing an HTTP endpoint other processes on your machine can reach [1].

Claude Code prompts for trust verification the first time it sees a new MCP server, with one exception worth knowing: that verification is disabled when it runs non-interactively with the -p flag [4]. If you script agent runs, the prompt you were relying on is not there.

Treat everything the server sends back as untrusted input

The part people underestimate is not the tool call. It is the response. A tool result goes into the model’s context, and text in context is instructions unless something stops it being instructions.

This is the mechanism behind indirect prompt injection, and the guidance is consistent. The NSA sheet says each output “must be treated as untrusted input to the next phase” and that filtering should include content length checks, disallowed keyword scanning and rate limiting, treating outputs that “may carry hidden logic or prompt manipulation” as a distinct risk that output filtering should screen for [3]. Anthropic rejects connectors whose tool descriptions contain hidden, obfuscated or encoded instructions, instruct Claude to call external software the user did not request, or direct it to pull behavioural instructions from external sources [5].

You cannot patch this yourself, and you should not try. What you can do is arrange your connections so that a successful injection is boring. An agent that can read a shared inbox and also write to your payments system is one poisoned email away from a bad afternoon. The same agent, split into a read-only inbox connection and a separate session for anything that spends money, is not. The NSA sheet asks you to “clearly define trust boundaries between MCP components”, grouping publicly available tools to handle public datasets while access to tools touching sensitive information is “explicitly controlled and segregated” [3]. In practice, for a small team, that means one thing: do not put a source of attacker-controlled text and a high-consequence write tool in the same session.

Some clients already do a version of this. Claude Code runs web fetches in a separate context window precisely to avoid injecting potentially malicious prompts into the main one [4]. That is the pattern to copy at the level you control, which is which servers you enable at the same time.

Keep an inventory and a log, or you cannot answer the only question that matters

When something goes wrong, the question is always the same. Which agent called which tool, with what parameters, and did it succeed. If you cannot answer that, you do not have an incident, you have a rumour.

Logging is the seventh of the NSA’s nine recommendations, and it asks that all tool and model invocations be logged, including the exact parameters and the identities involved [3]. Managed platforms increasingly give you this without work. Zapier’s MCP dashboard has a History tab displaying user-level activity logs for tool calls, superadmins can review those logs for any user in their account, and MCP events also appear in the account’s overall audit log [8]. Zapier is also worth noting as a structural choice rather than a feature list: because its tool catalogue is closed, “users cannot bring tools in from third-party sources”, which Zapier says prevents tool poisoning, at the cost of only being able to use what Zapier offers [8]. That trade is a good one for a small team with no appetite for auditing anybody’s code. Budget for it honestly, though, since each MCP tool call uses 2 tasks from your Zapier account task limit [8].

The inventory is the half that nobody does. Write down every server you have connected, in which client, under which account, with which scopes, and why. Then check that you know how to undo each entry: in Claude, connectors are managed under Customize then Connectors, where you can disconnect a service, modify connection settings, or review permissions and access levels [6]. Custom connectors are available on the free, Pro, Max, Team and Enterprise plans, with free accounts limited to one [6], which makes a free account a reasonable place to test a server you are unsure about.

calculator
One-time review cost of your MCP surface
— h, one-time

servers × tools per server × minutes each. That is the number of tool descriptions currently sitting in your model's context. Computed in the page; nothing is sent anywhere.

Run that once with your real numbers. Most people are surprised by the total, and the surprise is the point: the surface you have accepted is larger than the surface you remember accepting.

checklist
Before you connect an MCP server
0 of 8 · saved in this browser only

What still goes wrong

The largest unfixed problem is that none of this stops indirect prompt injection. It limits the damage. The guidance from every vendor here amounts to assuming injection will succeed and designing so that success is survivable, which is an honest position and also an admission. If your agent can read text an attacker can write, and can take an action you would not want taken, the gap between those two is closed by luck and by scope, not by the model’s judgment. Anthropic states the limit plainly in its own security documentation: these protections significantly reduce risk, but no system is completely immune to all attacks [4].

The second problem is that trust decisions are made once and servers change afterwards. You audited version 2.1. Version 2.4 arrived overnight through an auto-updating package, and nothing prompted you, because from the client’s perspective nothing changed. This is why the NSA recommendation is an inventory carrying versioning and patch history rather than a one-time review [3], and why pinning is worth the small inconvenience. It is also why a managed catalogue like Zapier’s, where you cannot introduce third-party tools at all [8], is a reasonable answer for a business with no capacity to track upstream releases.

The third is that the ecosystem moves faster than the guidance. The specification revision cited here is 2026-07-28 [2], client behaviour differs between Claude, ChatGPT and everything else, and features described as beta today will have changed by the time you read this. Check the vendor page before you rely on a control, including the ones cited above. The parts of this guide that will still hold are the ones that are not about any product: the server is the trust boundary, the credential is the ceiling, read and write are different decisions, and an integration you cannot list is an integration you cannot revoke.

sources
  1. 01Model Context Protocol — Security Best Practicesmodelcontextprotocol.io
  2. 02Model Context Protocol — Specification (revision 2026-07-28)modelcontextprotocol.io
  3. 03NSA — Model Context Protocol (MCP): Security Design Considerations for AI-Driven Automation, Cybersecurity Information Sheet, May 2026media.defense.gov
  4. 04Anthropic — Claude Code securitycode.claude.com
  5. 05Anthropic — Connector pre-submission checklist and review criteriaclaude.com
  6. 06Anthropic Help Center — Use connectors to extend Claude's capabilitiessupport.claude.com
  7. 07OpenAI Help Center — Developer mode and MCP apps in ChatGPThelp.openai.com
  8. 08Zapier — MCP securitydocs.zapier.com
next guide
How to read a frontier model's system card
10 min · verified 2026-09-05
related guides