Running a coding agent while you are not watching
Set up an unattended cloud coding run so the branch waiting in the morning is one you can evaluate in twenty minutes instead of trusting on faith.
on this page · 0 / 0 checked
You close the laptop at six. Somewhere in a data centre, an agent is still working on your repository. In the morning there is a branch, a diff indicator with lines added and removed, and a green check. You have no memory of what happened in hours 2 and 3, because you were asleep, and neither the branch nor the check is going to tell you.
Every serious vendor now works this way. Claude Code on the web runs tasks on Anthropic-managed cloud infrastructure, and sessions persist even if you close your browser [1]. Codex runs tasks in isolated cloud environments and lets you “start work from the web, GitHub, GitLab, Linear, or Slack” [5]. Cursor’s cloud agents run in “isolated VMs in the cloud with full development environments instead of on your local machine” [7]. GitHub’s Copilot coding agent works in “a GitHub Actions-powered environment” and opens “exactly one pull request to address each task it is assigned” [8]. The mechanics are settled. What is not settled is the part that used to happen continuously and now happens once, at the end, by you. This guide is for people who already read diffs. If you cannot tell whether a change is correct, an agent that produces more changes while you sleep is not a shortcut, it is a way to accumulate code nobody has read.
Unattended runs move the cost from writing to reviewing
When you watch an agent work, you correct it constantly and for free. You see it open the wrong file, and you stop it. Those corrections are invisible, and they are doing most of the work of keeping a run on track.
Take the watching away and every correction is deferred to one moment. The agent has not changed. The distribution of your attention has. Four hours of unsupervised work can contain 20 minutes of wrong turn and then three hours built on top of it, and nothing in the session list distinguishes that from four hours of steady progress.
The vendors’ own limits show where the reviewable units are. A Copilot coding agent session has “a maximum execution time of 59 minutes”, the agent “can only work on one branch at a time”, and it “can only make changes in the repository specified when you start a task” [8]. Claude Code cloud sessions stop after a period of inactivity and the VM is reclaimed; reopening one restores the conversation history, but background work that was still running, such as subagents and shell commands, is not restored [1].
Parallelism makes the arithmetic worse rather than better. Anthropic’s docs suggest starting several cloud tasks at once, each claude --cloud command creating its own independent session [1], and note that Claude Code on the web shares rate limits with all other Claude usage on your account, that running tasks in parallel consumes them proportionately, and that there is no separate compute charge for the cloud VM [1]. Cursor’s cloud agents require a paid plan, bill at API pricing for the model you select, and ask you to set a spend limit the first time you use them [7]. Money and rate limits are the costs the vendors meter. Your reading time is the one nobody bills you for, and it is the one that runs out.
The environment is the real permission system
In a supervised session you are the permission system: the tool asks, you approve. Unattended, there is nobody to ask, so the boundary has to be set in advance, in the environment rather than in the conversation.
Claude Code’s cloud environments take one of four network access levels. None allows no outbound network access through the session’s network, Trusted allows an allowlist of package registries, GitHub and cloud SDKs, Full allows any domain, and Custom takes your own list, optionally including the defaults [2]. New environments start at Trusted [2]. Codex is stricter by default: it blocks internet access during the agent phase, while “setup scripts still run with internet access so you can install dependencies”, and where you do open it up, administrators can restrict network requests to GET, HEAD, and OPTIONS so the agent can read the web but not post to it [6].
Credentials are the other half. In Anthropic-hosted environments, git credentials and signing keys stay outside the sandbox and a proxy authenticates on the session’s behalf with scoped credentials [1], and git push works only against the session’s current working branch [2]. Environment variables are not secrets: the docs say plainly that anyone who uses the environment can read the values [2]. On Pro and Max plans there is an alternative, an API credential the agent proxy attaches to requests after they leave the VM, so “the key never reaches Claude, the commands it runs, or the session’s environment variables” [2].
Set all of this before you write the first task, not after the first surprise. The narrowest environment the job can survive in is the correct one, and you will find out in the first run whether you cut too deep.
A task written for an empty room needs a stopping condition
“Improve the error handling” is a fine instruction when you are sitting there. Unattended, it has no end. Anthropic’s guidance for routines is the general rule: the run is autonomous, “so the prompt must be self-contained and explicit about what to do and what success looks like” [3].
The most useful pattern in the docs is to split thinking from doing. Start locally in plan mode, where Claude reads files and proposes an approach without editing source, agree the plan, commit and push it, then send the execution to the cloud with something like claude --cloud "Execute the migration plan in docs/migration-plan.md" [1]. This does two things at once. It gets your judgment into the run before the run starts, and it leaves an artefact in the repository that you can diff the result against afterwards. Note the mechanical detail that catches people out: the cloud VM clones your GitHub remote at your current branch, not your local checkout, so anything uncommitted is invisible to it unless you push first [1].
Give the run a failure route as well as a success condition. An agent told to stop and hand back a draft when it hits something it did not expect fails in a way you can recover in five minutes. An agent told nothing improvises, and improvisation compounds over four hours.
Isolation does not protect you from the content the agent reads
The VM solves the problem of an agent damaging your laptop. It does not solve the problem of an agent believing something it read. OpenAI’s warning is the clearest statement of the risk: “prompt injection can happen when the agent retrieves and follows instructions from untrusted content (for example, a web page or dependency README)”, and the consequences it lists include code or secret exfiltration and leaking material such as commit messages to attacker-controlled servers [6]. Anthropic’s security page lists its own layers of defence and then says the quiet part: “no system is completely immune to all attacks” [4].
Two details make this concrete for unattended runs. First, network isolation is not total. Anthropic notes that when running with network access disabled, Claude Code can still communicate with the Anthropic API, “which may allow data to exit the VM” [1]. Second, the tools you attach travel outside the network allowlist entirely. In routines, every connected MCP connector is included by default, and Claude “can use every tool from an included connector, including writes, without asking for permission during a run” [3]. Removing the ones a task does not need is a 10-second edit and the single highest-value one on the page.
The sharpest example in the docs is the auto-fix feature, where Claude watches a pull request and responds to CI failures and review comments. Anthropic warns that if your repository uses comment-triggered automation such as Atlantis, Terraform Cloud, or custom GitHub Actions that run on issue_comment events, Claude replying on your behalf can trigger those workflows, and advises disabling auto-fix for repositories where a PR comment can deploy infrastructure [1]. The agent did not do anything unusual. It left a comment. The blast radius came from your own plumbing.
Read the transcript before you read the diff
The diff tells you what the agent left behind. The transcript tells you what it did, which is the thing you gave up when you stopped watching. Read them in that order, because the second one explains the first.
Anthropic is unusually direct about what the status indicator means, and it is worth quoting exactly: “A green status in the run list means the session started and exited without an infrastructure error. It does not mean the task in your prompt succeeded” [3]. The same note adds that blocked network requests, missing connector tools and task-level failures all surface in the transcript rather than in the status [3]. An agent that spent an hour retrying a request that a 403 from your own allowlist was always going to reject looks, from the outside, exactly like an agent that worked.
Then read the diff properly. The review surfaces are decent now: Claude Code’s diff view takes inline comments on specific lines and sends them back with your next message [1], and Codex asks you to “review the summary and diff, request a follow-up, or open a pull request when the result is ready” [5]. Check scope first, because files touched outside what you asked for are the cheapest signal that the agent built on a premise you did not give it. Then check the logic, then check whether tests were added or updated for changed behaviour. A green CI run tells you the existing suite still passes, which is a statement about the suite, not about the change.
runs × minutes × 4.33 weeks per month. Computed in the page; nothing is sent anywhere.
A schedule turns one review into a standing commitment
The next step after a manual cloud run is a scheduled one. Claude Code calls these routines: a saved prompt, repositories and connectors that run on a recurring cadence, on an HTTP call to a per-routine endpoint, or in response to GitHub events such as pull requests and releases [3]. The minimum schedule interval is one hour, and there is a daily cap on how many runs can start per account [3].
Two properties of routines deserve more thought than they usually get. They “run autonomously as full Claude Code cloud sessions: there is no permission-mode picker and no approval prompts during a run” [3]. And everything they do wears your face: commits and pull requests carry your GitHub user, and Slack messages, Linear tickets or other connector actions use your linked accounts [3]. A schedule you set up in a confident moment keeps acting as you every night until you remember to pause it.
There are guardrails, and they are worth knowing because they define what a scheduled run cannot quietly do. Claude pushes to branches prefixed with claude/, which are always accepted; a push to any other branch is rejected if the branch is protected, if someone else has an open pull request from it, or if it carries commits authored by someone other than you [3]. Text you pass to a routine through its API trigger arrives wrapped in a block that labels it untrusted and tells Claude not to follow instructions inside it unless the routine’s own prompt says to [3]. That last one is a deliberate design choice against exactly the attack where somebody with your endpoint token tries to give your nightly job new orders.
Before you schedule anything, run it manually three or four times and read every transcript. A routine is not a way to stop reviewing. It is a way to commit to reviewing on a cadence, which is a heavier promise than a one-off run and should be made on evidence.
What still goes wrong
Prompt injection remains unsolved, and the honest position is the one both vendors take in their own documentation rather than any claim you will read elsewhere [4][6]. Isolation reduces the blast radius; it does not make an agent immune to instructions it finds in a README, a web page or an issue comment. Treat any run that reads untrusted content as a run whose output needs more review, not less.
The plumbing has rough edges that surface at the worst moment. Auto-fix cannot react to a merge conflict on its own, because GitHub does not emit a webhook when the base branch advances, so you open the session and ask for a rebase [1]. Repository cloning and pull request creation require GitHub; GitLab, Bitbucket and other remotes can be sent to a Claude cloud session as a local bundle, but the session cannot push results back [1]. Organisations with IP allowlisting enabled find that every Anthropic-hosted cloud session fails with an authentication error, because the session calls the API from Anthropic’s infrastructure and not your network [1]. Both Claude Code on the web and routines are labelled research preview, and the routines page states that behaviour, limits and the API surface may change [1][3].
And the underlying trade does not go away. You have bought parallel execution with your own attention, and attention is the input you cannot scale. Three agents running overnight produce three diffs that all want the same Tuesday morning. If the calculator above returns a number you would not have agreed to in advance, the answer is fewer runs, not faster reading.
- 01Anthropic — Use Claude Code on the webcode.claude.com
- 02Anthropic — Configure cloud environmentscode.claude.com
- 03Anthropic — Automate work with routinescode.claude.com
- 04Anthropic — Claude Code securitycode.claude.com
- 05OpenAI — Codex cloud overviewlearn.chatgpt.com
- 06OpenAI — Codex agent internet accesslearn.chatgpt.com
- 07Cursor — Cloud agentscursor.com
- 08GitHub — About Copilot coding agentdocs.github.com