tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

Getting results from agents that run for hours

Scope, brief and check an unattended agent run so what comes back is finished work rather than something that merely looks finished.

Published 2026-09-05 · Updated 2026-09-05 · Read 10 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

You handed over a 4-hour job and went to do something else. What comes back looks finished. There is a document, a folder of files, a summary saying every step is done. Then you open the third tab of the spreadsheet and it is empty, or the migration reported as complete touched 9 files out of 40. Nothing in the transcript flagged a problem. The run just ended, and the ending looked like success.

That is the characteristic failure of long agent runs, and it is not mainly a model problem. Anthropic’s engineering write-up on long-running agents describes a later agent instance that would “look around, see that progress had been made, and declare the job done”, alongside a separate “tendency to mark a feature as complete without proper testing” [2]. Meanwhile the products keep stretching the runway. ChatGPT Work, announced on 9 July 2026, is pitched as an agent that can “stay with a project for hours if needed, and turn a goal into finished work” [5]. Claude’s Dispatch agent is described plainly as being “for work you want to start and come back to later” [1]. The model capability is not the variable you control. The handoff is, and this guide is about the handoff, for solo operators and small teams using the shipped products. If you are building your own agent harness against an API, go to the vendors’ engineering documentation instead; this is too coarse for that job.

A multi-hour job has to be checkable without redoing it

The scoping test that matters is not whether the task is specific enough. It is whether you will be able to tell that the result is right in materially less time than it would have taken you to produce it.

Three shapes pass that test. Work with a mechanical check at the end, where tests pass or totals reconcile or every row has a value. Work whose output is small enough to read completely, like a 2-page brief or a 40-row table. Work where being wrong is cheap and reversible, like a draft, a branch, or a file in a scratch folder.

The shape that fails is work where verification costs as much as the work. “Go through 800 support tickets and tag each one by root cause” is a bad unattended job, not because the agent cannot do it but because you cannot check 800 tags any faster than you could have applied them. The fix is to change what comes back rather than what goes in. Have it tag all 800, then produce a 30-row sample with its reasoning shown, a count per tag, and a list of every ticket it was unsure about. Now the check is 15 minutes and it tells you something about the other 770.

Anthropic’s research post on long-running Claude has a clean illustration of what happens when you skip this. Claude Opus 4.6 spent a few days building a cosmology solver called CLAX and reached “sub-percent agreement with the reference CLASS implementation” [3]. Real work. But along the way, “for a while it was only testing the code at a single (fiducial) parameter point, drastically reducing its bug-catching surface area” [3]. The agent had chosen its own success criterion, and the one it chose was cheap. If you do not name the check, something will pick one for you.

Brief it like a competent stranger who cannot ask questions

Cursor’s guidance for its cloud agents is to “think of the agent as a smart, but low-context human developer” [7]. Anthropic frames the same constraint through the shift-work analogy: long jobs run as a series of sessions with limited context and no memory between them, so it is like “a software project staffed by engineers working in shifts, where each new engineer arrives with no memory of what happened on the previous shift” [2]. Whatever was in your head at the start is not available at hour three.

So the brief carries four things. The outcome, meaning the named artifact, its format, and where it should end up. The sources, meaning which files, folders and connected systems it may read, and which it may not. The stop conditions, meaning what to do when an input is missing. And the check, meaning how it must verify its own work before reporting the job as done.

Stop conditions are the highest-value lines you will write and the ones most people leave out. “If the Q2 export is not in the Finance folder, stop and tell me. Do not substitute Q1 and do not reconstruct the figures from the CRM.” Without that, the agent does something defensible, and defensible includes quietly using the wrong source and never mentioning it. You are not preventing a mistake so much as converting a silent one into a loud one.

The other half of a good brief is access. Cursor’s advice is to make sure the agent has “access to required secrets (API keys, database credentials, etc.)”, and it notes that “agents are limited by the tools they have access to”, recommending MCP and custom tools so the agent reaches the same systems a human would [7]. It also makes the environment point directly: “Like a human developer, Cursor does better work if its environment is set up correctly” [7]. An agent that cannot reach a system rarely stops. It routes around you.

Make it write down what it did while it does it

The single technique that moves unattended runs from unusable to usable is forcing a written trail as the work happens, not a summary afterwards. Anthropic’s harness for long-running coding agents prompts the model “to commit its progress to git with descriptive commit messages and to write summaries of its progress in a progress file” [2], and structures the work around a feature list whose items are “all initially marked as ‘failing’”, with the coding agent “asked to work on only one feature at a time” and to leave the environment in a clean state after each change [2].

You do not need git to use the idea. Ask for a run-log.md in the output folder, updated as it goes rather than at the end, with one line per step recording what it did, which source it used, what it checked, and anything it could not do. Ask for a second file listing every assumption it had to make. The cost is a few hundred tokens per step. The return is that when the deliverable is wrong, you can find out which step went wrong instead of throwing the whole run away.

Pair that with reversibility. Point the work at a copy, a branch, or a new folder rather than the live file, and say so in the brief. Anthropic’s harness treats the tests themselves as protected, telling the model that “it is unacceptable to remove or edit tests because this could lead to missing or buggy functionality” [2]. The general version for the rest of us is that the agent should never be able to make the evidence of its own failure disappear.

Approval prompts are timed, and silence counts as an answer

Long runs still stop to ask permission, and the mechanics of that pause are worth knowing before you are in a meeting with a phone face-down. In Claude’s Dispatch, when a child task needs permission for something like running a command or writing a file outside its workspace, the prompt is forwarded to you, and “if you don’t respond within ten minutes, the request is automatically denied and the task continues without that action” [1]. Denial is the safer default, but note what it produces: a run that carried on without a step it wanted, for another 3 hours, and the fact that it was denied is now buried in the transcript.

The prompts themselves get thinner as they get further from you. In Claude Code’s agent view, a background session that needs a decision sits in a yellow “Needs input” state, and “a permission prompt shows as text describing what the session wants to run, without numbered options” [4]. That is a decision presented as a sentence, 4 hours after you wrote the brief that would tell you whether to say yes.

OpenAI’s model is configuration-first. With workspace agents, you decide “what tools and data it can use, what actions it can take, and when it needs approval” [6]. ChatGPT Work adds an auto-review layer that uses OpenAI’s “most advanced models to review important actions involving connected tools and APIs before they happen”, and lets you “follow its progress, answer questions, change direction, and approve important actions” [5]. Either way the useful habit is the same: set the gates before the run starts, not at the moment of the prompt. Gate anything irreversible or outside the sandbox, and then approve only what you can evaluate from the summary in front of you. If you cannot evaluate it from that summary, deny it and go read the session.

One practical detail if you dispatch from a phone. The work runs on your desktop, and the instruction is to “leave Claude Desktop open with your computer awake and online” [1]. A closed laptop lid is a failed run.

Check the log before you accept the deliverable

Reviewing an unattended run is two questions, not one. Is the product right, and was the process right. A good-looking product built by a bad process is the case that costs you later.

Work in this order. First, read the verification lines in the log and confirm it actually ran the check you specified, rather than a cheaper one it invented. Second, spot-check 3 items pulled from the middle of the output, not the first 3, because early items get the most attention in any run. Third, search the log for hedging words like assumed, instead, unavailable and skipped, which is where silent substitutions confess. Fourth, count. If you asked for 40 and got 31, the missing 9 are the whole story and no summary will mention them.

That fourth check exists because of a specific behaviour. The same research post that describes the single-parameter testing problem also names “agentic laziness”, and describes it this way: “when asked to complete a complex, multi-part task, they can sometimes find an excuse to stop before finishing the entire task” [3]. The stopping is not random. It arrives with a rationale attached, which is exactly why it survives a quick skim of the summary.

When a run comes back wrong, resist rerunning the same brief. The agent did something a reasonable reading of your instructions permitted, and a second run under the same instructions will make similar choices. Change the brief at the point where the log shows it diverged, then run again. That edit is the durable asset. The output was for today; the brief is what you reuse.

What an unattended hour actually costs

Start with the access gate, because it is the part with a fixed price. Dispatch “requires a Pro or Max plan and the latest Claude Desktop app on macOS or Windows” [1]. Billed monthly, Claude Pro is $20 per month, Max starts at $100 per month, and Team standard seats are $25 per seat per month; annual billing drops Pro to $17 and Team standard seats to $20 [8]. On the OpenAI side, ChatGPT Work is usage-based rather than a separate line item, following the same usage structure as Codex, with more complex tasks consuming more of your plan’s included usage [5]. Workspace agents arrived on 22 April 2026 in research preview at no cost, with credit-based pricing starting on 6 May 2026 [6].

The subscription is the small number. The real cost of an unattended run is your time on both ends plus the expected cost of a wrong result reaching someone else. A 4-hour run that replaces 3 hours of your work, but needs 25 minutes of briefing and 45 minutes of checking, nets you 110 minutes. That is a good trade, and it is nothing like the 4 hours the run duration implies.

Budget for reruns the first time you use a new brief. Assume 2 runs, sometimes 3, until the brief stabilises. After that the same brief tends to work repeatedly, which is the actual economics of this: the first run of a job is expensive and the tenth is nearly free, so the jobs worth automating are the recurring ones, not the interesting one-off.

checklist
Before you start an unattended run
0 of 8 · saved in this browser only
calculator
Time saved per unattended run
— h / month

Your own hours minus briefing and checking, times runs per month. Computed in the page; nothing is sent anywhere.

What still goes wrong

The verification asymmetry never fully closes. For any output large enough to be worth 4 hours of machine time, you are sampling rather than reading, and sampling misses things by design. The honest position is that unattended runs suit work where an error found next week is annoying rather than expensive. If an error would be expensive, either shrink the output until you can read all of it or keep a human in the loop at each step, which means you are no longer running a multi-hour agent.

Domain judgment is still the weak spot, and it fails in a way that looks like competence. In the CLAX project the agent produced a working solver while also making “elementary mistakes, such as tripping over gauge conventions or spending hours chasing bugs that a cosmologist would spot instantly” [3]. The pattern generalises. These runs are most useful in areas where you know enough to review the output properly, and most dangerous in areas where you were hoping the agent would cover for what you do not know.

The approval machinery is also thinner than it looks. A 10-minute window with automatic denial [1] means an ordinary meeting can silently reshape a run, and a permission request rendered as a plain sentence with no options [4] is easy to answer wrongly at speed. Treat the gates as a way of stopping the worst outcomes, not as supervision. And none of this changes who is accountable for what ships. The agent ran for 4 hours; you are the one who sent the result.

Prompts from this guide

agent-run-brief

Task: {outcome}

Deliverable
- Produce {artifact} in {format}, saved to {location}.
- Also write run-log.md in the same folder, updated as you go, with one
  line per step: what you did, which source you used, what you checked,
  and anything you could not do.
- Also write assumptions.md listing every assumption you had to make.

Sources
- You may read: {allowed_sources}
- Do not read or modify: {forbidden_sources}
- Work on a copy or a new branch. Do not edit the original files.

Stop conditions
- If {key_input} is missing or unreadable, stop and write what you found
  to run-log.md. Do not substitute a similar source.
- If any step fails twice, stop and record it rather than working around it.

Check before you report done
- {verification_step}
- Report the count of items produced against the count expected.
- List anything you skipped, assumed, or completed only partially.
sources
  1. 01Anthropic — Run tasks in the background with Dispatchclaude.com
  2. 02Anthropic — Effective harnesses for long-running agentsanthropic.com
  3. 03Anthropic — Long-running Claude for scientific computinganthropic.com
  4. 04Anthropic — Manage agents with agent viewcode.claude.com
  5. 05OpenAI — ChatGPT is now a partner for your most ambitious workopenai.com
  6. 06OpenAI — Introducing workspace agents in ChatGPTopenai.com
  7. 07Cursor — Cloud agent best practicescursor.com
  8. 08Anthropic — Claude pricingclaude.com
next guide
The copyright risk you actually carry when you use AI
9 min · verified 2026-09-05
related guides