tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · running the business

When you let an AI spend money

Set a spending limit an AI agent cannot argue with, keep the approval step switched on, and know which protections you lose when the card is a business card.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

An AI that writes a bad email costs you a redraft. An AI that pays the wrong invoice costs you the invoice, plus the fortnight you spend asking a stranger’s accounts department to send it back. Every other kind of AI mistake is caught before it leaves the building. A payment is caught after. That asymmetry is the entire subject here, and it is why controls that look like overkill on a drafting tool are the floor for a spending one.

You are closer to this than you probably think. ChatGPT’s Instant Checkout completes purchases inside the chat window using encrypted payment tokens that OpenAI says are “only authorized for specific amounts and specific merchants with the user’s permission” [1]. Google’s Agent Payments Protocol, announced on 16 September 2025 with more than 60 organisations behind it, among them American Express, Mastercard and PayPal, defines a format for standing authority that lets an agent buy without a human present at the moment of purchase [3]. This guide is for a solo operator or small-team owner about to let a tool touch a real card. It is not for a finance function with a controller and a procurement policy, and it is not legal advice. The two US rules cited below are here to show you where your protections stop, not to tell you what to do about that.

Money mistakes are the only ones you cannot re-run

Reversibility is the quiet property that makes everything else about AI tolerable. A wrong summary gets corrected. A wrong draft gets rewritten. You pay in attention, and you pay before anything is final.

A payment inverts both halves. You pay in cash, and you find out afterwards, when the money is already sitting in someone else’s account and your leverage is a polite email. The failure modes you have learned to shrug at also change weight. A model that invents a plausible detail is an annoyance in a draft and a wire to the wrong sort code in an invoice. A model that can be talked into an unintended action by text it reads on a web page is a nuisance in your inbox and a loss on your statement.

None of this argues that agents should never touch money. It argues that the controls have to do the work that hesitation does for a human. People pause before a large or unfamiliar payment. An agent does not pause unless the pause is built into something outside the agent.

The limit belongs on the card, not in the prompt

The most common mistake is putting the spending limit in the instructions. “Never spend more than $200” is a sentence in a context window, competing for the model’s attention with everything else in that window, including text the model picked up from a supplier’s web page. It is a request. It can be misread, outweighed, or argued past.

A limit set at the issuer is a decline. The transaction does not reach a judgement call. If you use Stripe Issuing to give an agent its own virtual card, the controls are structured, not conversational: spending limits attached to an interval such as per authorization, weekly or monthly, allowed_categories or blocked_categories for merchant types, allowed_merchant_countries, and allowed or blocked card presence [2]. Where limits overlap, the most restrictive applies [2]. If you set no limit at all, Stripe applies a default of 500 USD per day to the new card, and an unconfigurable default limit of 10,000 USD applies on top of that to each individual authorization; both are automatic, and switching them off means contacting Stripe support rather than changing a parameter [2].

Two details in that documentation are worth copying even if you never touch Stripe. First, spending limits alone do not restrict where money goes; they have to be paired with an allowed or blocked category list to keep an agent inside a type of business [2]. A cap says how much, not to whom. Second, the limit belongs on the card the agent uses and nowhere else. Give the agent its own card with a small float, the way you would give a new contractor an expense card rather than the account your payroll runs from, so that freezing the agent does not freeze you.

The standing mandate is the permission that actually matters

The transaction people worry about is the one an agent shows them for approval. The transaction that costs money is the authority granted weeks earlier and then forgotten.

Both of the emerging standards are explicit about this, which is useful, because it tells you what to read. In the Agentic Commerce Protocol behind ChatGPT’s Instant Checkout, the credential handed to the agent is a token bound to a specific amount and a specific merchant, and OpenAI’s description of the flow is that users “explicitly confirm each step before any action is taken” [1]. Google’s AP2 splits the same idea into two signed objects: an Intent Mandate capturing what you asked for, and a Cart Mandate that “creates a secure, unchangeable record of the exact items and price” [3]. When you are present, you approve and sign the cart. When you are not, you sign a detailed Intent Mandate up front specifying “price limits, timing, and other conditions”, and the agent generates the Cart Mandate on your behalf once those conditions are met, with no approval at the moment of purchase [3].

That second mode is the whole risk, described plainly by the people building it. A delegated mandate is a small contract you sign in advance for purchases you have not seen. Write it the way you would write any authority you cannot supervise: a ceiling, a named merchant or a narrow category, and an expiry date. An open-ended mandate with no end date is the agentic equivalent of a subscription you forgot to cancel, except it can buy things. The upside of the design is real, though. AP2’s sequence from intent to cart to payment is meant to create a “non-repudiable audit trail” of what was authorised, which is exactly the evidence you will want if a payment is later disputed [3].

Approval steps drift off by default

Approval settings are not permanent. They ship with defaults, the defaults change with releases, and nobody sends you a note when the shape of your safety net changes.

Claude in Chrome became generally available on every paid Claude plan on 26 August 2026, and that release changed a default: Claude “will now automatically approve actions it determines to be safe, using the same mechanism as auto mode in Claude Code”, and you “can switch this off in your settings if you’d prefer to continue to approve Claude’s actions manually” [5]. Anthropic pairs that with real restrictions. Claude “asks for permission before accessing financial sites”, and the company says it “strongly advise[s] against using Claude in Chrome to manage or take actions on sensitive information”, with managing financial accounts or investments named in the list [4]. Read that as the vendor telling you where its own confidence ends. OpenAI’s guidance runs in the same direction: limit an agent’s access “to only the sensitive data or credentials it needs to complete the task”, because “it’s safer to ask your agent to do specific things, and not to give it wide latitude to potentially follow harmful instructions from elsewhere like emails” [8].

The practical habit is short. After any update to a tool that can spend, open the settings page and look at the approval toggle before you look at the new features. A control you set once in March is not a control you have in September.

A business card carries fewer rights than your personal one

Most operators assume the consumer protections they know from their own bank follow them into the business. In the United States they do not, and this matters more once an agent is the one initiating payments.

Regulation E is the rule that defines an “unauthorized electronic fund transfer” as one from a consumer’s account “initiated by a person other than the consumer without actual authority to initiate the transfer and from which the consumer receives no benefit”. It hangs that on two other definitions: an “account” is a consumer asset account “established primarily for personal, family, or household purposes”, and a consumer is “a natural person” [6]. Regulation Z, which covers credit cards, does not apply to “an extension of credit primarily for a business, commercial or agricultural purpose”, nor to “an extension of credit to other than a natural person” [7]. A card issued to your company for company spending sits outside both definitions. What you have instead is your issuer’s cardholder agreement and the card network’s rules, which are commercial terms rather than statutory rights.

The consequence is not that you should run agent payments on a personal card. It is that you should find out, before the first agent transaction and not after a bad one, what your issuer’s dispute window and liability terms actually say, and keep the agent’s spending on a card whose terms you have read. If you cannot get a clear answer in ten minutes, that itself is information about how the dispute will go.

Exposure is the cap multiplied by the time before you look

A spending cap does not tell you your risk. A cap plus a review interval does. If an agent can move $200 a day and you read the ledger monthly, you are carrying about $6,000 of exposure to a fault that started on day one. Shorten the interval and the same cap costs a fraction of that.

calculator
Worst case before you notice
— $ exposed

Daily cap × cards × days since your last reconciliation. Computed in the page; nothing is sent anywhere.

The review itself is ten minutes and three questions: whether there is a payment here you do not recognise, whether a recurring amount has crept up, whether anything has gone to a payee you never approved. An agent that paid one wrong invoice this week will pay it again next week, on schedule, until a person notices, so the value of the habit is almost entirely in its frequency rather than its depth.

checklist
Before an agent touches a real card
0 of 8 · saved in this browser only

What still goes wrong

Prompt injection is not solved and the vendors say so. Anthropic says its current configuration “reduces attack success rates to less than 0.08%” against internal testing that combines known effective attack techniques, while stating that the chances of an attack are “still non-zero” [4]. OpenAI, publishing on 7 November 2025, describes prompt injection as a problem that “remains a frontier, challenging research problem” and expects its work on it to be ongoing [8]. A sub-0.1% rate reads as safe until you multiply it by every page an agent reads on your behalf over a year, and the tail of that distribution is where the payments are.

Caps leak at the edges. Stripe notes that spending aggregation is done on a best-effort basis, with a delay of up to 30 seconds between a spend and its aggregation, and that additional tips and fees can be posted later, causing a spending limit to be exceeded [2]. Treat an issuer cap as a firm ceiling with a soft edge rather than a mathematical guarantee, and size it so that a small overshoot is survivable. Undoing a bad purchase is also slower than making one: under the Instant Checkout model, orders, payments and fulfilment are handled by the merchant using its existing systems [1], so a refund is a conversation with a seller, not a button in your assistant.

Finally, the scope. This guide covers one operator letting one agent spend small, repetitive money under caps they control. Once the volume is high enough that nobody could reconcile it by eye, you have crossed into a different problem that wants proper accounting software, segregated approval duties and a named finance owner, and no checklist substitutes for those. The legal points here are US-specific and current as of the retrieval dates in the sources; if a real payment has already gone wrong, the person to call is your issuer, then an accountant, in that order.

sources
  1. 01OpenAI — Buy it in ChatGPT: Instant Checkout and the Agentic Commerce Protocolopenai.com
  2. 02Stripe Docs — Issuing spending controlsdocs.stripe.com
  3. 03Google Cloud — Announcing Agent Payments Protocol (AP2)cloud.google.com
  4. 04Anthropic Help Center — Use Claude in Chrome safelysupport.claude.com
  5. 05Anthropic — Claude in Chrome is generally availableclaude.com
  6. 06eCFR — 12 CFR 1005.2, Regulation E definitionsecfr.gov
  7. 07eCFR — 12 CFR 1026.3, Regulation Z exempt transactionsecfr.gov
  8. 08OpenAI — Understanding prompt injectionsopenai.com
next guide
Which AI rules actually reach a business your size
9 min · verified 2026-09-05
related guides