Budgeting for AI when the price stops falling
Learn to read a vendor price sheet for its expiry dates, work out your real cost per job, and build a budget that survives an increase.
on this page · 0 / 0 checked
For two years the safe assumption about AI costs was that they only went one way. Every few months a cheaper tier arrived, the same job cost less than it had the previous quarter, and anyone quoting a client on AI-assisted work could underprice slightly and expect the gap to close on its own. Plenty of small businesses built their margins on that expectation without ever writing it down as an assumption.
It is written down now, on the vendors’ own pages, and it points in both directions. Google lists Gemini 3.8 Flash at $0.75 per million input tokens “through December 31, 2026” and $1.50 “starting January 1, 2027”, with output going from $3.75 to $7.50 on the same day [2]. OpenAI lists GPT-5.6 Sol at $4 per million input and $20 per million output on standard processing, with a footnote calling that promotional pricing “available at least through November 21, 2026” and no successor rate published anywhere on the page [1]. This guide is about budgeting against numbers that behave like that. It is not for anyone with a negotiated committed-spend agreement, where the price is a contract term rather than a web page, and it will not tell you where prices go next, because nobody knows.
Falling prices were a business decision, not a trend line
A published price is a decision rather than a law of physics, and vendors are still making that decision in your favour in places. Google’s own list has Gemini 3.8 Flash at $0.75 per million input tokens against $1.50 for the earlier 3.5 Flash, so the newer model currently costs half as much to feed [2]. Cuts like that have not stopped. What has changed is the size of the fixed obligations now sitting behind them.
The clearest public example is OpenAI’s. Reporting in July 2026 put its planned infrastructure spending through 2030 at $750 billion, some 25% more than the company had estimated earlier that year, and described a $20 billion data centre campus spanning 1,400 acres northwest of Savannah, drawing at least 3.2 gigawatts from Georgia Power, with OpenAI paying the full cost of the infrastructure and the electric service and Effingham County granting a 50% property tax abatement for 15 years [7]. Take the exact figures with the caution any second-hand number deserves. The shape is what matters, and it is not unique to one company.
A commitment of that kind converts a variable cost into something closer to a fixed one. A vendor renting capacity month to month can shrink when demand disappoints. A vendor that has signed multi-year contracts and broken ground on sites it owns has to service those obligations whether or not revenue arrives on schedule. That does not mean prices must rise. It means the floor under them is higher and more visible than it was, and that price cuts now have to come out of efficiency gains rather than out of a willingness to lose money for market share indefinitely.
For you, the practical consequence is narrow. Stop treating “it will be cheaper by then” as a line in a plan. Treat today’s published price as the price, and any future reduction as a windfall rather than a budget item.
The price sheet already tells you which way each number is going
Vendors are being unusually explicit, and the information is in the fine print rather than the headline rate.
Three patterns are worth learning to spot. The first is an end date with a successor price attached, like Gemini 3.8 Flash doubling on 1 January 2027 [2]. That is the easy case. You can put it in a calendar and decide in advance whether you move, absorb it, or reprice. The second is an end date with nothing after it, like OpenAI’s promotional rate for GPT-5.6 Sol running “at least through November 21, 2026” [1]. That is harder, not easier. An open date means you cannot plan the change, only the possibility of one. The third is no date at all, which is the majority of the sheet and means nothing except that the vendor has not committed to anything.
Generation numbers are not a reliable guide either. On Google’s own page, Gemini 3.5 Flash costs $1.50 per million input and $9.00 per million output, while the newer 3.6, 3.7 and 3.8 Flash models all sit at $0.75 and $3.75 until the new year [2]. Newer is cheaper there. But older Gemini 2.5 Flash is listed at $0.30 for text, image and video input and $2.50 output, cheaper still than any of the 3.x Flash models [2]. And on 1 January 2027, Gemini 3.8 Flash input arrives at $1.50, exactly what the earlier 3.5 Flash costs today [2]. Three generations of releases, and the input rate on that ladder ends up back where it started.
The modifiers matter as much as the base rate, because they are where a vendor allocates scarce capacity by price rather than by announcement. OpenAI’s sheet is four tabs, not one: against standard processing, cached input is charged at 10% of the input rate, writing a cache entry carries a 25% premium, batch and flex run at half price, and fast mode runs at double, with a further 10% uplift for regional processing on models released on or after 5 March 2026 [1]. That spread is wide enough to read the wrong number by accident. GPT-5.6 Sol is $4 and $20 on standard and $2 and $10 on batch, and if you quote yourself the batch rate for interactive work you have understated your costs by half before you start [1]. Google discounts its Batch API by 50%, and charges separately to keep a context cache warm: $0.50 per million tokens per hour on Gemini 3.8 Flash through 31 December 2026, then $1.00 from 1 January 2027 [2]. Anthropic’s published rates run from Haiku 4.5 at $1 and $5 per million to Fable 5.1 at $10 and $50, a factor of ten across one price list [3]. Before you conclude that your costs are rising, check whether you are simply buying the expensive corner of a sheet that has a cheap corner.
Your real unit is cost per job, not cost per million tokens
Per-million-token rates are almost useless for planning because nobody buys a million tokens. You buy a job: one invoice read, one draft written, one support reply, one batch of listings tagged. Until you know what a job costs, a price change is an unquantified worry rather than a number.
Getting it is a twenty-minute exercise. Take the single workflow that accounts for most of your usage, run it once, and read the token counts off the vendor’s usage page for that call. Multiply by how many times a month you run it. Multiply again by the published input and output rates for the model you actually used, on the processing tier you actually used, not the one you meant to use. What comes out is a number in dollars per month. Write it down before you spend a weekend optimising it, because that figure is what decides whether the weekend is worth spending.
Then do the part that matters here. Run the same arithmetic with the price doubled. If your workflow is on Gemini 3.8 Flash, that is not a hypothetical, it is the published rate for 1 January 2027 [2]. If the answer is an extra $30 a month, you have learned that this whole subject is not your problem and you can stop reading about it. If the answer is an extra $600 a month against a fixed-price retainer you signed for a year, you have found a real exposure, and you found it while there was still time to do something.
Priced at Claude Sonnet 5's published $2 per million input and $10 per million output [3]. A 100% rise is what Gemini 3.8 Flash is scheduled to do on 1 January 2027 [2]. Computed in the page; nothing is sent anywhere.
When the per-token price cannot fall, the product changes instead
Watching the price per token is watching one lever out of several, and often not the one that moves.
Consider what has already happened at the free end. OpenAI now runs an advertising-supported free tier, describing advertising as “one pillar of OpenAI’s diversified business model, alongside consumer subscriptions, enterprise offerings, and usage-based APIs” and saying it helps keep ChatGPT available to more than 1 billion weekly active users. The company said that in less than 200 days after launch, ChatGPT Ads had reached $1 billion in annualised revenue run rate, and that advertisers could buy directly through Ads Manager across India, Europe, the Middle East and North Africa from 31 August 2026 [4]. Read that as pricing information, because it is. When the marginal cost of serving a free user cannot go to zero, the user gets monetised another way.
The paid tiers move too, usually through the allowance rather than the price. OpenAI’s own help page says Plus is $20 a month and that subscriptions “may include usage limits such as message caps, especially during high demand”, adding that these limits “may vary based on system conditions” [6]. Anthropic prices Claude Pro at $17 a month on an annual subscription billed at $200 up front, or $20 monthly, with Max starting from $100 and letting you “choose 5x or 20x more usage than Pro” [3]. Notice that the expensive plans are sold in multiples of usage rather than in features. That is what a capacity-constrained business looks like when it prices a subscription.
Retirement is the third lever, and the one that quietly reprices your workflow without any price changing. OpenAI commits to at least 6 months’ notice for generally available models, at least 3 months for specialised variants such as chat, Codex and deep research builds, and says preview models can go with as little as 2 weeks [5]. On 26 August 2026 it notified developers using whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize that those models would be removed from the API on 26 February 2027 [5]. When the model you standardised on goes away, you move to whatever is current, at whatever it costs then. Your budget changed; the price sheet did not.
Price your own work so a vendor’s increase is not your emergency
The exposure that actually hurts a small business is a mismatch of terms. You sold a fixed monthly fee for twelve months. Your supplier reserves the other side of the deal in writing: Anthropic states that price and plans are “subject to change at Anthropic’s discretion” [3], Google prints the date its rates change [2], and OpenAI’s promotional rate carries a floor date and nothing after it [1]. Everything between those two facts is your margin.
Three things fix most of it, and none of them require you to predict anything. Quote AI-heavy work in shorter commitments, or write a plain sentence into the agreement saying that fees may be revisited if third-party model pricing changes by more than some percentage. Ask for the clause rather than assuming a client will refuse it. Second, keep a headroom figure. If AI costs are 8% of what you charge for a piece of work, a doubling takes them to 16% and you survive it; if they are 40%, a doubling ends the product, and you needed to know that before you launched it, not after.
Third, keep the work portable enough that moving is a decision rather than a project. The mechanics are dull and they hold up. Store your prompts and instructions in files you own, not only inside one vendor’s account. Set the model name in one place per workflow instead of scattering it through steps. Once a quarter, run your two most important prompts against a comparable model from a second vendor and note honestly whether the output is acceptable, because the answer changes with every release and a stale answer is worth nothing. The point is not to run two vendors. It is to know the size of the switch before you are forced to make it, and the price ladders make the stakes plain: Haiku 4.5 at $1 and $5 against Fable 5.1 at $10 and $50 on one vendor’s list [3], and GPT-5.6 Luna at $0.20 and $1.20 against GPT-6 Astra at $10 and $50 on the other, rising to $20 and $100 for the same Astra calls in fast mode [1].
What still goes wrong
Every number in this guide was read from the vendor’s own page on 4 September 2026, and the whole argument is that such numbers move. Google has already published its 1 January 2027 increase [2] and OpenAI’s promotional rate has a stated floor date and nothing beyond it [1]. Check the pages rather than trusting this one in six months.
The cost-per-job exercise is also less stable than it looks. Token counts drift when you change a prompt, when a model becomes more verbose, and especially when reasoning or tool use is involved, because the tokens you are billed for include work you never see in the answer. Measure a real run rather than estimating from the length of your prompt, and re-measure after any significant change. If your usage is mostly you typing in a chat window on a subscription, none of the per-token arithmetic applies to you at all. Your exposure is the plan price and the message cap, and the cap is explicitly elastic: OpenAI says Plus limits “may vary based on system conditions” [6], and Anthropic says price and plans are “subject to change at Anthropic’s discretion” [3].
The hedging advice has honest limits too. Keeping a second vendor evaluated costs you a real afternoon each quarter and buys you an option you may never exercise. Prompts are not fully portable, output differs in ways that break downstream steps, and both vendors may be constrained by the same suppliers at the same time anyway. Nobody can tell you whether the price of a given model rises, falls or stays flat next year. The recommendation is only this: stop building plans that require it to fall.
- 01OpenAI — API pricingdevelopers.openai.com
- 02Google — Gemini API pricingai.google.dev
- 03Claude — Pricingclaude.com
- 04OpenAI — A milestone in expanding access to AIopenai.com
- 05OpenAI — Deprecationsdevelopers.openai.com
- 06OpenAI Help Center — What is ChatGPT Plus?help.openai.com
- 07TechCrunch — OpenAI's AI spending spree has ballooned to $750Btechcrunch.com