What the AI chip race actually does to your bill
Trace the custom-silicon race to the two places it reaches you, price your own usage against today's published rates, and stop assuming prices only fall.
on this page · 0 / 0 checked
You will never buy an AI chip. You will never see one, never price one, and never hold a useful opinion about its memory bandwidth. What you will see is a charge on your card every month, or an API bill that arrived bigger than you budgeted, and a vague sense that both numbers are connected to the custom-silicon news you keep half-reading.
They are connected, but not in the direction most of the coverage implies. Every large lab is now spending heavily to make running a model cheaper, and very little of that saving arrives as a smaller number on your statement. It arrives as a cheaper floor, more capacity at the same price, and a top tier that holds its premium. This guide is for someone who pays for AI out of a small business account and wants to know which published prices to plan around. It is not for anyone buying hardware, negotiating a cloud commitment, or trying to pick a stock. Every price below was read from the vendor’s own page on 4 September 2026.
The race is aimed at inference, which is the part you pay for
There are two phases to using a model. Training is the one-time job of building it. Inference is what happens every time you use it, which is every reply, every summarized document, every automation run at 3am. Training gets the announcements. Inference is where your recurring cost lives, and it is what the new chips are built for.
Google lists Ironwood as its seventh-generation TPU, generally available, with 9,216 liquid-cooled chips per pod and a claimed 4x better performance per chip over the previous Trillium generation [4]. The next generation splits the job in two. TPU 8t is described as built for large-scale pre-training, with 9,600 chips in a single superpod and nearly 3x the compute performance per pod over the previous generation. TPU 8i is described as “optimized for post-training and inference” and is claimed to give an “80% performance-per-dollar improvement over previous generations for low-latency inference for large MoE models” [4]. Both are listed as coming soon.
Amazon lists Trainium2 and Trainium3, says Trainium3 has up to 2x the MXFP8 compute throughput of Trainium2, and frames the whole line around “better cost-per-token at production scale” [6]. OpenAI and Broadcom announced a collaboration on 10 gigawatts of custom AI accelerators, with deployments targeted to start in the second half of 2026 and to complete by the end of 2029, on the reasoning that designing its own chips lets OpenAI “embed what it’s learned from developing frontier models and products directly into the hardware” [5].
Read those together and the unit everyone is optimising is the same: cost per token, or performance per dollar. That is the vendor’s cost, not your price. A chip that halves what an answer costs to produce does not halve what you are charged for the answer. What it does is give the vendor room to move, and where that room goes is a separate decision made by a separate team.
Cheaper compute reaches you as more capacity, not a cheaper subscription
If you buy AI as a monthly seat, this is the part that matters, and the evidence is on the pricing page. Claude Pro is $20 per month, or $17 per month billed annually at $200 up front. Max starts at $100 per month, and what it sells is described in multiples of usage, either 5x or 20x more than Pro, with higher output limits. Team seats are $25 per month, or $20 per seat billed annually [7]. Nothing there is a rate you can watch fall. The ladder is built out of usage headroom.
The clearest version of the pattern sits at the bottom of the ladder. ChatGPT’s Free plan lists unlimited text chats with GPT-5.6 Luna, subject to abuse guardrails, along with limited access to uploads, image generation, voice and deep research [8]. On the API price list, that same model is the cheapest OpenAI publishes, at $0.20 per million input tokens and $1.20 per million output [2]. The model a vendor gives away is the one sitting at the bottom of its own price list, and the paid tiers above it are sold on headroom rather than on a lower rate.
So a subscription buyer should stop watching the price and start watching two other things. First, which model your plan actually routes you to by default, since that is where a vendor’s cost savings and cost pressures both show up first. Second, what the included limits are this quarter compared to last. A plan that stayed $20 while quietly doubling what you can do with it has passed you a real saving. A plan that stayed $20 while tightening its weekly limit has quietly raised your price.
On the meter, the cheap tier falls and the frontier keeps its premium
If you buy tokens instead of seats, the same price list contains both directions at once. Anthropic currently lists Claude Sonnet 4.6 at $3 per million input and $15 per million output, and the newer Claude Sonnet 5 at $2 and $10. Claude Haiku 4.5 sits at $1 and $5. Claude Opus 5 is $5 and $25. At the top, Claude Fable 5.1 and Claude Mythos 5.1 are $10 and $50 [1]. Newer was cheaper in the Sonnet line, and the frontier tier is 10x the input price of the cheapest current model on the same page [1].
OpenAI’s list has a wider spread and the same shape. On standard rates, gpt-6-astra is $10 and $50, gpt-5.6-sol is $4 and $20, gpt-5.6-terra is $2 and $12, and gpt-5.6-luna is $0.20 and $1.20 [2]. That is a 50x range in input price inside one vendor, one page, one day. Read that page carefully, because the same four models are printed four times, under Standard, Batch, Flex and Fast mode tabs. Batch and Flex are half the standard rate and Fast mode is double it [2]. Copy a number off the wrong tab and every estimate you build from it is wrong by 2x in one direction or the other.
The practical consequence is not subtle. The largest saving available to you is almost never waiting for the frontier model to get cheaper. It is moving the jobs that do not need the frontier model onto a tier that already costs a fraction. Classification, extraction, tagging, routing, first-draft summarisation and most automation steps do not need the most expensive model on the list. Reserve the top tier for the work where being wrong is expensive, and price the rest at the bottom of the ladder.
One published price is already scheduled to double
The comfortable assumption is that AI prices only fall. The counterexample is already published, with a date on it. Google lists Gemini 3.8 Flash at $0.75 per million input tokens and $3.75 per million output through 31 December 2026, then $1.50 and $7.50 starting 1 January 2027 [3]. That is a doubling of both sides, announced in advance on the vendor’s own pricing page. OpenAI dates a rate too, more quietly: a footnote says GPT-5.6 Sol’s promotional pricing is available at least through 21 November 2026, which tells you the $4 and $20 next to it is not a standing number [2].
Across generations the line is not monotonic either. Gemini 2.5 Flash is listed at $0.30 input and $2.50 output for text, image and video. Gemini 3.5 Flash, a later model, is listed at $1.50 and $9.00. Gemini 3.8 Flash, later again, comes back down to $0.75 and $3.75 for the rest of this year [3]. Prices move in both directions, and a version number is not a discount.
Two habits follow. Read the effective date, not just the number, whenever you take a price off a page. And if you are quoting a client a fixed monthly fee, or setting a product margin, do not build it on an introductory rate without putting the end date in your calendar. A rate that doubles on a published schedule will not send you a warning email in December.
Batching and caching cut more than the next chip generation will
The largest cost reductions available to you today are not silicon. They are two switches in how you call the API, and all three major vendors offer both.
Batch processing takes 50% off input and output on Claude, on OpenAI models, and on Gemini [1][2][3]. Anthropic publishes the arithmetic directly: Claude Opus 5 batched is $2.50 and $12.50 instead of $5 and $25, and Claude Sonnet 5 batched is $1 and $5 [1]. The trade is latency. Anything a human is actively waiting on stays on the standard endpoint. Overnight enrichment, backfills, weekly reports, bulk classification and any queue that tolerates a delay should be batched, and most small operators have more of that work than they think.
Caching is the larger cut. OpenAI prices cached input at a 90% reduction from standard input rates across its models [2]. Anthropic charges a cache read at 0.1x the base input rate, dropping to 0.025x on Claude Fable 5.1 and Claude Mythos 5.1, and charges a premium to write the cache: 1.25x base for a 5-minute cache and 2x for a 1-hour one [1]. Google lists context caching for Gemini 3.8 Flash at $0.075 per million tokens through the end of 2026, rising to $0.15 on 1 January 2027 [3]. If your prompts carry a long fixed preamble, a style guide, a product catalogue, a policy document, the same schema on every call, caching is where your bill actually shrinks.
The write premium is the catch, and it is why caching is not free money. Anthropic puts the break-even in plain terms: at the 5-minute duration a cache pays off after one read, and at the 1-hour duration after two [1]. Cache the stable part of your prompt, not the part that changes per request.
Set against these, TPU 8i’s claimed 80% performance-per-dollar improvement is a vendor figure about a chip that is not shipping, and any part of it that reaches you will arrive slowly and indirectly [4]. Batching and caching are a configuration change you can make this afternoon.
Defaults are Claude Sonnet 5 list rates [1]. Halve the answer to see the batch price, and re-run it with a cheaper tier's rates before you assume you need the expensive one. Computed in the page; nothing is sent anywhere.
Habits for a market where the price moves both ways
Prepay only what you would have spent anyway. Claude Pro at $17 per month billed annually against $20 monthly is about 15% off a roughly $240 decision, and a year of lock-in at that size is a small bet [7]. A large prepaid credit balance on an API meter is a different bet, because the rate underneath it can move in either direction and you have already committed to the old one.
Re-price the ideas you rejected. Keep a short list of automations you turned down because the token cost did not work, and put a recurring quarterly reminder against it. A workflow that failed on price at Claude Sonnet 4.6’s $3 per million input may pass at Claude Haiku 4.5’s $1 [1], and the only thing standing between you and that saving is remembering that you once said no.
Watch volume rather than rate. A halved price on tripled usage is a larger bill, and cheaper tokens are an active invitation to send more of them. The number to review monthly is total spend and total tokens, not the per-million figure you feel good about.
Keep the model name in exactly one place. If the model you call is set in one config value rather than scattered through prompts, scripts and automation steps, moving work down the ladder is one edit rather than an afternoon. That single habit is what turns every price change on every vendor’s page into something you can act on rather than read about.
What still goes wrong
Almost every performance figure in the chip story is the vendor’s own. The 80% performance-per-dollar claim for TPU 8i, the 4x per-chip gain for Ironwood, the up-to-2x throughput for Trainium3 are all marketing pages describing products the vendor sells [4][6]. None of them has been independently verified here, and none of them is a commitment about what you will be charged. A cheaper chip is an input to a pricing decision, not the decision.
The timing is genuinely long. The Broadcom deployments are targeted to start in the second half of 2026 and to complete by the end of 2029 [5]. TPU 8t and TPU 8i are listed as coming soon [4]. Anything you plan around that horizon is a guess, and a business decision that only works if a chip ships on schedule is not a business decision.
Cheaper is also not the same as freer. A lab that designs its own silicon, runs its own models and sets its own prices has more leverage over you, not less, even while the number on the screen goes down. Falling per-token prices can arrive alongside tighter rate limits, changed default models, deprecated versions you had built on, and terms you did not negotiate. The prices in this guide were read on 4 September 2026 and any of them can move; treat every figure here as a snapshot to re-check rather than a standing fact.
- 01Claude Platform Docs — Pricingplatform.claude.com
- 02OpenAI Developer Docs — Pricingdevelopers.openai.com
- 03Google — Gemini API pricingai.google.dev
- 04Google Cloud — Tensor Processing Units (TPU)cloud.google.com
- 05OpenAI — OpenAI and Broadcom announce strategic collaborationopenai.com
- 06AWS — Trainiumaws.amazon.com
- 07Claude — Pricingclaude.com
- 08ChatGPT — Pricingchatgpt.com