When running your own model actually saves money
Work out whether an open-weight stack beats a paid API at your volume, and find the cost cuts that pay off long before you rent a GPU.
on this page · 0 / 0 checked
The gap between what a flagship model API charges per token and what a hosted open-weight model charges is large enough to make owning the stack look obvious. It is also a comparison of two prices, not a comparison of two bills, and the difference between those two things is where small operators lose money. A per-token figure for a model you run yourself excludes the hardware, the serving software, the person who keeps it running and the hours the machine sits idle. A per-token price from a managed API includes all of that, because the provider paid for it and spread it across its customers.
This guide is for solo operators, freelancers and small teams who are looking at an AI bill and wondering whether owning the stack would fix it. It is not for anyone whose decision is already made by policy. If a client contract or a regulator requires the weights to sit on hardware you control, cost is not the argument you are having, and you should price the setup rather than compare it. For everyone else, the honest answer is that the cheap model is available to you without the GPU, and the arithmetic below shows by how much.
A price gap is not a bill
The prices genuinely do span an order of magnitude, and more. Claude Fable 5.1, Anthropic’s flagship, costs $10 per million input tokens and $50 per million output tokens [1]. The small model in the same family, Claude Haiku 4.5, costs $1 and $5 [1], and Gemini 3.5 Flash-Lite costs $0.30 and $2.50 [4]. Hosted open-weight models are cheaper again: Together serves gpt-oss-120B at $0.15 input and $0.60 output, and DeepSeek V4 Flash at $0.14 and $0.28 [5]. Between the top and the bottom of that list there is a factor of about 70 on input and about 180 on output.
So the gap is real. What a headline comparison leaves out is that both sides of it are prices per token from a vendor. A self-hosted number is different in kind. It is a cost per token that only exists once you have paid for a machine, installed a serving stack, tuned it, and pushed enough traffic through it to divide the fixed cost down to something small. The published figure assumes the denominator. Your job is to check whether you have one.
One rented H100 costs about $2,870 a month
Start with the fixed cost, because it is the number that decides most of these cases on its own. An NVIDIA H100 SXM on Lambda’s GPU cloud is $3.99 per GPU-hour on demand [6]. Together lists H100 capacity at the same $3.99 on demand, or $1.99 preemptible [5]. Rent one continuously for 30 days and you have spent about $2,870 before anyone touches it. An A100 SXM with 80GB is $2.79 an hour on Lambda, or roughly $2,000 a month, and a B200 SXM6 is $6.69, or about $4,800 [6].
Now price a plausible small-operator workload against that. Say 100 substantial model calls a day across 22 working days, with 10,000 input tokens and 2,000 output tokens each. That is 2,200 calls, 22 million input tokens and 4.4 million output tokens a month. On Claude Sonnet 5 at $2 and $10 per million [1], that is $44 of input and $44 of output, so $88 a month. On Claude Haiku 4.5 it is $44 [1]. On Gemini 3.5 Flash-Lite it is about $17.60 [4]. On gpt-oss-120B at Together it is about $5.94 [5].
Hold those two figures next to each other. The Sonnet 5 bill for that workload is $88 a month. The single rented GPU is about $2,870 a month. To reach the GPU’s cost in Sonnet 5 tokens you would have to run roughly 72,000 calls a month, about 3,300 every working day. That is the breakeven against the most expensive of the four options priced above, with your own labour valued at zero and the machine assumed to be busy every hour you pay for.
Open weights do not require your own hardware
The step most people skip is that “open model” and “my server” are two separate decisions. Open weights are a licensing fact about the model. Where it runs is a hosting choice, and a serverless provider will run it for you at a per-token price with no fixed cost at all [5]. That is why the comparison above is not really frontier API against self-hosting. It is frontier API against cheap hosted open model against self-hosting, and the middle option captures almost all of the price gap while keeping the fixed cost at zero.
Redo the breakeven with the middle option as the baseline. That same 2,200-call workload costs about $5.94 a month on gpt-oss-120B [5]. One H100 at $3.99 an hour costs about $2,870 [6]. You would need on the order of a million calls a month, roughly 48,000 a working day, before owning the machine beat renting the same class of model by the token. And that is the optimistic version, because it assumes a single GPU could serve that load at all. If it cannot, you are buying two, and the breakeven moves further away.
The discounts you can take without changing anything
Before moving a workload anywhere, take the discounts that are already sitting on the API you use. There are three, and none of them requires changing vendor.
The first is batching. Anthropic’s Message Batches API charges 50% of standard prices, completes most batches in under an hour with a 24-hour ceiling, and accepts up to 100,000 requests or 256MB per batch [2]. OpenAI’s Batch API is also a 50% cost discount against the synchronous API, with a 24-hour completion window, up to 50,000 requests per batch and a 200MB input file [3]. Google lists a Batch API with a 50% cost reduction as well, and its published batch rates are half the standard rates model by model [4]. Anything that does not need an answer while a human waits belongs here: overnight enrichment, bulk classification, transcript summaries, content rewrites.
The second is caching. If your calls repeat a long fixed preamble, a system prompt, a style guide, a product catalogue, you are paying full input price for the same tokens on every call. On Claude, a cache read costs 0.1x the base input price, and 0.025x on Claude Fable 5.1, against 1.25x for a 5-minute cache write and 2x for a 1-hour write [1]. On Gemini 3.8 Flash, cached input is $0.075 per million tokens through 31 December 2026, with storage charged at $0.50 per million tokens per hour [4]. Check that your model offers it at all before you plan around it: Gemini 3.5 Flash-Lite lists context caching as not available [4].
The third is model tiering, which is the least technical and usually the largest. The question to ask of each call is whether it needs the top model. Moving routine extraction and formatting from Claude Fable 5.1 at $10 and $50 down to Claude Haiku 4.5 at $1 and $5 is a 10x cut on both sides of the meter [1], available this afternoon, with no infrastructure to own. Batch those same calls and the cut is 20x [1][2], which is a bigger reduction than moving from Sonnet 5 to gpt-oss-120B would give you [1][5].
Where your breakeven actually sits
Work out your own number rather than borrowing anyone’s. You need four inputs: how many calls you make, how many tokens go in and out of each, and the two prices. The calculator below does the arithmetic for a month of working days.
calls × 22 working days × token cost per call. Prices from the vendor pricing pages cited above. Computed in the page; nothing is sent anywhere.
Compare the result against $2,870, one H100 rented around the clock for a month [6], or against $1,433 at Together’s preemptible H100 rate [5]. If your monthly token bill is smaller than the machine, the question is settled and you should stop pricing GPUs. If it is larger, the second question is utilisation. A GPU rented at an hourly rate bills for every hour you hold it, so a workload that runs in a two-hour burst each night pays for 22 idle hours a day. Batch traffic is exactly the shape that wastes a dedicated machine and exactly the shape that the batch APIs discount by half [2][3].
There is also a cheaper habit worth naming, because it is where small teams accidentally overspend in the other direction. Interactive work usually belongs on a flat plan rather than the meter. Claude Pro is $20 a month billed monthly or $17 a month with an annual subscription, Max starts at $100 a month, and a standard Team seat is $25 a month or $20 billed annually [8]. Keep the API bill for automation, and keep the humans on a subscription.
The costs that never appear in a per-token comparison
Serving a model yourself is a real engineering surface. vLLM is one open-source way to do it, and its own documentation is a fair map of the job: installation paths for GPU, CPU and TPU, hardware pages for Intel Xeon CPUs and Intel XPUs, deployment guides for Docker, Kubernetes and Nginx, and features including paged attention, automatic prefix caching, speculative decoding and several quantisation formats [7]. That list is the good news and the bad news at once. The software is capable and free. It is also a system you now operate: version upgrades, memory tuning, quantisation choices, request queueing, a health check, an on-call plan and a rollback path for the night it stops answering.
None of that is priced per token, and for a solo operator it competes directly with billable work. A saving of $80 a month that costs 12 hours a month to maintain is a pay cut. The threshold worth applying is not “does this reduce inference cost”, it is “does this reduce total cost including the hours it takes from me”. Add to that the migration itself. A prompt tuned against one model is not guaranteed to behave the same against another, so you owe yourself an evaluation on your own tasks before and after, not a vendor’s benchmark on theirs. A vendor’s evaluation of its own product is a starting point, not evidence about your workload.
What still goes wrong
Prices move faster than guides do. Several of the figures above are explicitly temporary: Gemini 3.8 Flash is listed at $0.75 input and $3.75 output through 31 December 2026, rising to $1.50 and $7.50 on 1 January 2027, with its cached input price doubling on the same date [4]. Promotional rates expire, model names change, and a stack you chose on a price gap can lose that gap in a quarter. Re-check the pricing pages before you commit, and prefer decisions that are cheap to reverse.
The arithmetic here also assumes your quality bar survives the cheaper model, and sometimes it does not. A 10x price cut that lifts your error rate enough to need human review has not saved you anything, it has moved the cost from the meter to your afternoon. Run the cheap model against a fixed set of your own real inputs and read the outputs before you switch, not after.
Finally, this guide answers the cost question only. There are legitimate reasons to run your own stack that have nothing to do with money: data that cannot leave your infrastructure, a need to fine-tune on material you will not upload, a client who requires it in writing, or a hedge against a single vendor changing terms. Those are decisions about risk and control. If one of them applies, make it on those grounds and treat the higher cost as the price of the requirement rather than pretending the numbers favour it.
- 01Anthropic — Claude model pricingplatform.claude.com
- 02Anthropic — Message Batches APIplatform.claude.com
- 03OpenAI — Batch API guidedevelopers.openai.com
- 04Google — Gemini API pricingai.google.dev
- 05Together AI — Pricingtogether.ai
- 06Lambda — GPU Cloud pricinglambda.ai
- 07vLLM — Documentationdocs.vllm.ai
- 08Anthropic — Claude plan pricingclaude.com