Governing AI spend on a small team
Split the bill into fixed seats and metered usage, put a hard ceiling on the metered half, and review it monthly against something other than the invoice.
on this page · 0 / 0 checked
In the first year, AI on your team was a few subscriptions and you knew what each one cost. Then someone wired an agent to the API. Someone else switched the automation tool to pay-per-task so a workflow would stop pausing at month end. A third person moved onto a premium seat because they kept hitting a limit at three in the afternoon. None of those decisions was wrong, and none of them came to you as a purchase. The bill is now partly a subscription and partly a utility meter, and the meter is the half nobody is reading.
This guide is for whoever signs off on software spend at a company of roughly 2 to 50 people, where the AI budget has grown past the point where one person remembers all of it. It is not for enterprises with a committed-spend contract and a vendor rep, and it is not about negotiating discounts. Every price below is a US list price read from the vendor’s own page on 5 September 2026, and vendors move these without much notice.
Your AI bill has two halves and only one of them is a subscription
The fixed half is seats, and it is priced almost identically across the two assistants most small teams use. Claude Team is $20 per seat per month billed annually or $25 billed monthly, for teams of 2 to 150, with a premium seat at $100 annually or $125 monthly carrying 5 times more usage than a standard seat [1]. ChatGPT Business has the same shape: $20 per seat billed annually, $25 monthly, a premium seat at $100 or $125 offering 5 times more usage than standard, and a stated range of 2 to 200 employees [4]. Coding tools sit higher. Cursor Teams is $40 per user per month for a standard seat and $120 for premium [7].
The metered half is tokens and actions, and it has no ceiling built into it. Claude Opus 5 is $5 per million input tokens and $25 per million output; Sonnet 5 is $2 and $10; Haiku 4.5 is $1 and $5; Fable 5.1 is $10 and $50 [2]. On the OpenAI side, GPT-5.6 Sol is $5 input and $30 output, Terra is $2 and $12, and Luna is $0.20 and $1.20 [5]. Some platform features meter separately again: Anthropic lists managed agents at $0.08 per session-hour and web search at $10 per 1,000 searches [1].
Tools that wrap those models add a third layer on top. Cursor’s documentation says every plan includes a pool of model usage, that past it you either add on-demand usage which continues at the same API rates with pay-as-you-go billing or upgrade to a tier with more included usage, and it estimates daily agent users typically at $60 to $100 a month of usage with power users often above $200 [7]. Zapier’s Free plan is 100 tasks a month and Professional starts at $19.99; past the task limit you are either switched to pay-per-task billing at a higher per-task rate or your workflows pause [8].
The fixed half governs itself, because a person approved a specific number and can be asked about it. The metered half was approved as a capability, not as an amount, and that is the whole problem.
Establish what each tool does when you hit the limit, before you set a budget
For every metered tool, there is one mechanical fact worth knowing before anything else: whether it stops at the limit or keeps charging past it. The answer is a property of the product, not of your intentions, and the two behaviours need completely different governance.
OpenAI documents both instruments and keeps them separate. You can set spend alerts on the limits page to send notifications when usage exceeds a certain dollar amount, and you can set a hard spend limit, which stops affected API traffic when tracked spend reaches the limit [6]. An alert tells you. A hard limit acts. Teams routinely set the first and assume they have set the second.
Zapier stops by default. If pay-per-task billing is disabled, your workflows pause once you reach the plan’s task limit [8]. That is a real ceiling, obtained for free, and it is also an outage waiting for a bad week. Cursor’s pricing documentation goes the other way: on-demand usage at API rates or a plan upgrade are the two routes it gives past your included allowance, and it documents no administrative spend cap [7]. If a tool has no cap, the cap is your card.
The useful pattern for a small team is the one OpenAI documents for larger ones. Create separate projects to isolate development and testing from production, and set custom rate and spend limits per project [6]. Two projects is enough. Experiments and anything an agent runs unattended get a low hard cap, on the theory that a runaway loop should die cheaply. Production gets a high cap and an alert well below it, on the theory that you want a phone call rather than a stoppage.
The model you default to is the largest line item nobody chose
The spread between the cheapest and most expensive model on the same menu is wide enough that routing is a bigger lever than any negotiation available to a small team. Haiku 4.5 at $1 and $5 against Fable 5.1 at $10 and $50 is a factor of 10 on both input and output [2]. On OpenAI’s list, Luna at $0.20 and $1.20 against Sol at $5 and $30 is a factor of 25 [5]. Seat tiers repeat the pattern in miniature: a premium seat is 5 times the usage for 5 times the price [1][4].
Left alone, defaults drift up. The frontier model is the one people reach for when a task feels important, and importance is not the same variable as difficulty. The fix is a task inventory rather than a model policy. List the jobs each team actually runs through the tool, then pick the cheapest model that produces an output you would ship, judged on the output and not on how the model feels. Drafting and editing prose, tagging support tickets, extracting fields from documents and summarising a meeting are usually near the bottom of the menu. Generating code that runs against production and reasoning over long, messy inputs usually are not.
Write the mapping down as jobs to price tiers, not jobs to model names. Model names churn and the prices attached to them move in both directions. Sonnet 5 is $2 and $10, while Sonnet 4.6, still sitting on the same price list, is $3 and $15 [2], so a workload pinned to the older model identifier in code is quietly paying 50 percent more than the same call would cost on the current one. Pinning is the right call for reproducibility. It is the wrong call to forget about for a year.
Defaults are Claude Opus 5 list prices. Enter 2 and 10 for Sonnet 5, or 1 and 5 for Haiku 4.5, to price the same volume on a cheaper model. Computed in the page; nothing is sent anywhere.
Two discounts are already on the price list
If your team calls models directly, two reductions are available without talking to anyone, and both are documented on the public pricing pages.
The first is batching. Anthropic gives a 50 percent discount on both input and output tokens for batch processing [2], and OpenAI advertises 50 percent off inputs and outputs with the Batch API [5]. The qualifying condition is patience, not volume. Any job with no human waiting on it belongs there: overnight classification, backfilling summaries across an archive, scoring a week of tickets, regenerating descriptions after a schema change. Teams skip this because the synchronous endpoint already works, which is the same reason nobody moves cold storage.
The second is caching, and it rewards a specific shape of workload. Anthropic prices a 5-minute cache write at 1.25 times base input and a 1-hour write at 2 times, with a cache hit at 10 percent of the standard input price, and states that caching pays off after one cache read for the 5-minute duration or two reads for the hour [2]. OpenAI publishes cached input as its own rate, $0.50 per million against $5.00 uncached on Sol [5]. If every call in a workflow starts with the same long system prompt, policy document or code context, you are currently paying full price to resend it each time.
Neither discount reaches a chat seat. A seat is a flat monthly fee with usage bundled inside it [1][4], and no amount of prompt discipline changes what it costs. That is the argument for keeping repetitive high-volume work on the API side of the house even when a person could do it in the chat window.
Read the numbers monthly, from something that is not the invoice
An invoice is a total arriving four weeks late. The usage data underneath it is available now, at a granularity that answers the only question worth asking, which is what changed.
Anthropic exposes this through the Usage and Cost API, which requires an Admin API key; workspace keys do not work [3]. The usage endpoint is /v1/organizations/usage_report/messages and it filters and groups by API key, workspace, model and service tier; the cost endpoint is /v1/organizations/cost_report and it groups by workspace and by a description field that carries the model [3]. Usage comes in 1-minute, 1-hour or 1-day buckets; cost is daily only. The documentation states that usage and cost data typically appears within 5 minutes of API request completion, and that the API supports polling once per minute for sustained use [3]. OpenAI’s equivalents are the usage tracking dashboard covering the current and past billing cycles, an email notification threshold, and per-API-key usage monitoring on the Usage page once tracking is enabled [6].
You do not need a pipeline to benefit from this. You need one named person, 30 minutes a month, and three questions. Which model line grew, and did the work that model does grow with it. Which API key or workspace grew, and is it attached to something that shipped. Whether any single day looks unlike its neighbours, because a spike that lasts one day is almost always a loop, a retry storm or a test left running, and almost never a person.
Treat the usage export itself with the care you would give any internal report. Per-user and per-workspace consumption data describes what your team is working on and how hard, which is not information you want circulating more widely than the spend decision requires.
What still goes wrong
Every figure here was read from a vendor page on 5 September 2026 and any of them can change next quarter, along with what a plan includes at that price. The structure survives the churn better than the numbers do. Seats and meters, ceilings before budgets, cheapest model that passes, one owner reading the data monthly. Re-check the prices when you re-check the policy.
Hard caps have a failure mode that is worse than the spend they prevent. A limit that stops affected API traffic [6] will eventually fire on the afternoon a client deliverable is due, and the person who hits it will not know why the tool broke. If you set caps, set them per project rather than per organisation, tell people the number exists, and make sure raising it takes minutes rather than a purchase order. A ceiling nobody can lift quickly is not governance, it is a scheduled outage.
The deeper limit is that none of this measures value. Analytics tell you which team spent what on which model. They cannot tell you whether the summaries anyone generated were read, whether the drafted emails were sent, or whether an agent that cost $400 last month replaced work that cost more. A team can run a tidy spend review for a year and never once ask that question, because the numbers in front of them are all about consumption. Somebody has to ask it separately, out loud, and be willing to switch something off.
- 01Claude — Pricingclaude.com
- 02Claude Docs — Pricingplatform.claude.com
- 03Claude Docs — Usage and Cost APIplatform.claude.com
- 04OpenAI — Business pricingopenai.com
- 05OpenAI — API pricingopenai.com
- 06OpenAI — Production best practicesdevelopers.openai.com
- 07Cursor Docs — Pricingcursor.com
- 08Zapier — Pricingzapier.com