tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · running the business

Budgeting for agents that bill by the run

Price a single agent run, set the caps your vendor already offers, and pull the levers that lower cost without lowering what the agent does.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

For a long time the AI line on your card statement was a subscription. A flat price per person, a limit you bumped into on a busy Thursday, and a forecast you could do by counting heads. Agents ended that arrangement. An agent run opens files, calls tools, thinks for a stretch and writes something back, and the cost of one run is nothing like the cost of the next: OpenAI puts a typical Codex task on GPT-5.6 Sol anywhere between 5 and 30 credits [1]. So the vendors put it on a meter. Your seat did not disappear. It stopped being the whole bill and became an allowance with a meter running underneath it.

This is a budgeting problem before it is a technology problem, and it is not hard, but it does need one number you almost certainly do not have: what a single run of your most common agent task costs. This guide is for a solo operator or a small team already running agents on a paid workspace plan, watching the invoice move and unsure why. It is not for anyone building on raw model APIs, where the metering is explicit and the tooling is different, and it is not for an organisation with a procurement function that already does this properly. Every price and rate below was read from the vendor’s own page on 5 September 2026, and vendors change them without much notice.

Three meters, and only one of them is the seat price

The bill has three parts, and confusing them is where most of the surprise comes from.

The first is the seat. ChatGPT Business standard seats are $25 per user per month, or $20 billed annually; premium seats are $125 per user per month, or $100 billed annually [3]. Claude Team is the same shape: $25 per member per month standard, $20 annually, and $125 premium, $100 annually, with a minimum of two members [5]. Claude Pro sits under both at $20 a month, or $17 a month billed annually at $200 up front, and Max starts at $100 a month [6].

The second is the usage each seat already includes, and it is expressed as a multiple rather than a quantity. A ChatGPT Business premium seat gives 5x more usage than a standard seat and drops the 5-hour usage limit [3]. Claude does the same arithmetic against Pro: a Team standard seat gives 1.25x more usage per session, a premium seat 6.25x, with usage limits applied per member rather than to the team as a whole [5]. That last detail matters more than it looks. Per-member limits mean your heaviest user cannot quietly drain everyone else’s headroom, but they also mean buying a second standard seat does not help the person who keeps running out.

The third meter is the one that generates the invoice you did not expect. On ChatGPT Business, that meter runs in credits, and credits can be used across all seat types once included usage is exhausted [2]. On Claude Team and seat-based Enterprise plans, usage credits let members on standard and premium seats keep using Claude, Cowork and Claude Code after they reach their included limits [4]. Anthropic’s usage-based Enterprise plans work differently again, priced as a seat plus usage at API rates, and usage credits do not apply to them [4][6].

Find out which of the three you are actually on before you touch anything else. Two of them are fixed and one is not, and only the third needs managing.

The number to find first is the price of one run

Credit metering comes in two styles, and you need to know which one your agent uses.

Some things are a flat rate per action. In ChatGPT, agent mode costs 30 credits per message, deep research is 50 credits per task, an image generation is 5 credits, and voice is 5 credits per minute; an ordinary GPT-5.6 Sol message is 10 credits, and a Sol Pro message is 50 [1]. Those are easy to budget because they do not vary.

Everything else is token-based, and the formula is published: total credits equals input tokens divided by a million times the input rate, plus the same for cached input, plus the same for output [1]. The rates differ sharply by model. GPT-5.6 Luna is 5 credits per million input tokens and 30 per million output; GPT-5.6 Terra is 50 and 300; GPT-5.6 Sol is 100 and 500; GPT-6 Astra is 250 and 1,250 [1]. OpenAI’s own worked example runs GPT-5.5 over 20,000 input tokens, 80,000 cached input tokens and 5,000 output tokens, and lands at about 7.25 credits [1].

Two things fall out of those rate tables immediately. Output is the expensive stream, costing five to six times input on every model in the table [1]. And cached input is priced at exactly one tenth of fresh input, on every model in the table [1]. Both are levers, and both are covered below.

The number you want is not credits per month. It is credits per unit of work you can name. If your agent produces a weekly client summary, run it three times, read the usage figures, and write down what one summary costs. From there the question of whether the agent is worth running becomes arithmetic rather than a feeling, and you can answer it without a meeting.

calculator
Monthly agent credit burn
— credits / month

people × runs per week × credits per run × 4.33 weeks. Defaults use the published 30 credits for an agent-mode message [1]; replace them with your own measured figures. Computed in the page; nothing is sent anywhere.

Caps go in before you know what the number will be

Both vendors ship spend controls, and almost nobody turns them on until after the first uncomfortable invoice. Do it in the opposite order. A limit set from a guess is more useful than a limit set from evidence, because the guess exists during the month when the damage would happen.

In ChatGPT Business, the controls live in workspace settings under Billing. Owners and admins can manage monthly credit usage limits by seat type and add per-user overrides, and an override takes precedence over the seat-type limit [2]. Purchased credits stay valid for 12 months, and if the workspace does not have enough credits, features that require credits may be unavailable until credits are added [2].

Claude Team and seat-based Enterprise plans give you an organisation-wide monthly limit and an individual per-user monthly limit; the third control, a separate limit for the standard and premium seat tiers, is seat-based Enterprise only [4]. Team owners can enable auto-reload, choosing the balance at which the top-up fires and the amount to add, and admins can read per-member spend in the MTD Spend column under Organization Settings and then Usage [4]. Anthropic is candid about the order of operations: the limit is checked before a request is processed and token consumption is calculated after, so one request can carry you past the cap before subsequent requests are blocked [4]. Set the cap below the number that would actually hurt.

Auto-reload deserves a moment of thought. It is the right setting when an interrupted agent run costs you more than the credits do, and the wrong one when your main risk is a workflow that loops. If you enable it, pair it with a per-user cap, so a single runaway account cannot pull the whole balance down repeatedly.

Four levers that lower the price of a run

Reuse the same context, identically. Cache reads are priced at 0.1x the base input rate, against 1.25x to write a five-minute cache entry and 2x for a one-hour entry; on Claude Fable 5.1 and Mythos 5.1, cache hits and refreshes are 0.025x [7]. The default cache lifetime is five minutes, and the minimum cacheable prompt runs from 512 tokens on the newest models up to 4,096 on some others [7]. The order is what makes it work in practice: caching runs tools, then system, then messages, and a change at one level invalidates that level and everything after it, so editing a tool definition invalidates the system prompt and the messages below it too [7]. Put the stable material first and leave it untouched between runs. Runs that rebuild their context differently every time never cache, and pay full rate forever.

Ask for less output. Output tokens cost five to six times input on the published rate card [1], so an instruction that produces a tight summary rather than an essay moves the expensive number, not the cheap one. Naming a length in the prompt is the highest-yield edit most people never make.

Send the work to the cheaper surface. An agent-mode message is 30 credits and an ordinary Sol message is 10 [1], which means running an agent for a single-step question against material you already have costs three times what asking would. Agents earn their rate when a task needs several tool calls and real multi-step work. For everything else, the chat window is the correct tool and the cheap one.

Batch the reads. If five runs each load the same 40-page policy document, you are paying to load it five times. One run that produces five outputs pays for the context once, at the cost of a longer output stream and a harder-to-debug result. The trade is worth making when the shared context is large and the outputs are short, and not worth making when it is the other way round.

Automation platforms meter the same work on a different clock

Agents that live inside an automation tool are billed by that tool’s meter, not the model vendor’s, and the two rarely resemble each other. Zapier’s Free plan includes 100 tasks a month, Professional starts at $19.99 a month with 750 tasks, and Team starts at $69 a month [8]. A task is used when a Zap successfully moves data or completes an action, and Zapier never charges a task to check for new data, so polling is free [8]. Agents there are metered separately in activities rather than tasks, with 400 activities a month on the free tier and 1,500 on Agents Pro at $400 billed annually, about $33.33 a month; one Zapier MCP tool call uses two tasks from your plan’s quota [8].

The consequence is that the same job can be cheap in one place and expensive in another, and the deciding factor is usually how many times something fires rather than how clever it is. An agent that runs once each working day fires about 22 times a month. The same logic attached to a trigger that fires on every new row in a busy spreadsheet is a different order of magnitude, on a plan that will not tell you until the quota is gone. Before you move an agent into an automation platform, count the triggers, not the steps.

checklist
Before the next billing cycle closes
0 of 8 · saved in this browser only

What still goes wrong

Credits are not dollars, and the rate card publishes no conversion between them [1]. That makes credits a good relative unit for comparing one run against another and a poor one for forecasting a bill, so you still have to open the billing page to turn a credit count into money. Budget in credits, verify in dollars, and do not let anyone present a credit figure as a cost.

The rates themselves move. OpenAI states that GPT-5.6 Sol promotional pricing is available at least through 21 November 2026, which is a plain statement that the price after that date is open [1], and the same rate card records o3 retiring on 26 August 2026 and GPT-5.4 and GPT-5.4 mini on 31 August 2026 [1]. A model retirement is a re-pricing event. Whatever you measured gets measured again when your default model changes underneath you.

The harder limit is that none of this tells you whether the agent is worth having. A cap stops the bill growing; it does not tell you that a run costing 30 credits produced something you would have paid a person to do. That judgment needs a different number, which is what the task was worth to you, and no dashboard will supply it. Cost control without that second number is how teams end up with an agent that is cheap, contained, well governed and pointless.

sources
  1. 01OpenAI Help Center — ChatGPT Rate Cardhelp.openai.com
  2. 02OpenAI Help Center — Managing credits and spend controls in ChatGPT Businesshelp.openai.com
  3. 03OpenAI Help Center — ChatGPT Business overviewhelp.openai.com
  4. 04Anthropic Help Center — Manage usage credits for Team and seat-based Enterprise planssupport.claude.com
  5. 05Anthropic Help Center — What is the Team plan?support.claude.com
  6. 06Claude — Pricingclaude.com
  7. 07Claude Docs — Prompt cachingplatform.claude.com
  8. 08Zapier — Pricingzapier.com
next guide
How to read an AI vendor's safety numbers
9 min · verified 2026-09-05
related guides