When owning an AI model beats renting one
Price the real cost of an open-weight model against a closed API at your own volume, and know the three cases where owning actually wins.
on this page · 0 / 0 checked
Every few months a lab publishes a model you can download. The weights sit on Hugging Face, the licence permits commercial use, and the pitch writes itself: stop renting intelligence from a company that can reprice it, retire it, or read what you send. Own it instead. Thinking Machines Lab made that pitch about as clearly as anyone has when it released Inkling in July 2026, a 975-billion-parameter model that activates 41 billion parameters per token, under an Apache 2.0 licence, with its own fine-tuning platform attached [1][2].
The pitch is real. The arithmetic is where it usually comes apart. For most solo operators and small teams, owning a model turns out to mean renting different infrastructure from a different company, with more of the failure modes landing on your desk. What follows is how to tell which case you are in before you spend a quarter finding out. It is not for people whose entire use of AI is a chat subscription, and it is not for teams that already have someone running GPUs in production, because both have made the decision already.
Open weights and open source are different things, and the licence is the difference
The Open Source Initiative published version 1.0 of its Open Source AI Definition on 28 October 2024 [7]. To qualify, a system has to grant four freedoms, use, study, modify and share, and the developer has to release the preferred form for making modifications: information about the training data, the code for training and running the system, and the parameters [7]. Almost nothing marketed as open meets that bar, because data information is the part labs will not publish.
What you actually get varies more than the word “open” suggests. Inkling ships under Apache 2.0, which lets you use, modify and redistribute it commercially as long as you keep the notice [2]. Llama 4 ships under Meta’s community licence, which requires that any product with more than 700 million monthly active users on the release date request a separate licence from Meta, that you “prominently display ‘Built with Llama’” on your site or product documentation, and that any model you train from it “include ‘Llama’ at the beginning of any such AI model name” [6].
The user threshold will never bind you. The other two will. Attribution attaches to your product page, and the naming rule attaches to the fine-tuned model you were planning to treat as your own asset. Read the licence before you build, and read the redistribution clause first, because a fine-tuned derivative is what you are going to end up holding.
The bill moves, it does not disappear
There are three ways to pay for a model, and they are priced on completely different axes.
Renting a closed model is per token. Claude Opus 5 lists at $5 per million input tokens and $25 per million output; Claude Sonnet 5 at $2 and $10; Claude Haiku 4.5 at $1 and $5. Batch processing takes 50% off, and cache hits bill at 0.1 times the base input rate [5]. Nothing is charged when nothing is running.
Renting an open-weight model is also per token, from whoever hosts it. On Together, Inkling lists at $1.00 input and $4.05 output per million tokens. gpt-oss-120B lists at $0.15 and $0.60. DeepSeek V4 Flash lists at $0.14 and $0.28 [4]. Same weights, same licence, same right to walk away with the model, and no infrastructure of your own.
Renting the hardware is per hour. Together’s on-demand rate for a dedicated NVIDIA HGX H100 is $5.49 per hour, currently shown with a promotional rate of $3.99; an HGX B200 is $8.99 per hour [4]. At the list rate, a single H100 left running costs about $4,000 a month, and it bills identically whether you send it two requests or two million.
That third option is the one people picture when they say ownership, and it is almost always the wrong one to start with. A workload of 100 million input tokens and 20 million output tokens a month costs $181 through Together’s Inkling endpoint [4]. You would need roughly twenty times that volume before one dedicated H100 broke even, and a 975-billion-parameter model will not fit on one H100 anyway. For comparison, gpt-oss-120b is engineered specifically so that it “fit[s] on a single 80GB GPU” [8], and it is a seventh the size.
Three reasons justify owning, and they are the only three
The first is insurance. Weights on a disk cannot be deprecated, repriced or rate-limited by anyone. Anthropic’s current price list carries thirteen Claude model versions across four families [5], and the composition of a list like that is not something you control. If a model being withdrawn would break a product you sell, holding the weights is a real hedge, and it is worth paying something for.
The second is customization that prompting cannot reach. If you need consistent output format across thousands of runs, or domain vocabulary a general model keeps getting subtly wrong, or behaviour that survives without a 2,000-word system prompt in front of it, fine-tuning is the tool. Prompting has a ceiling and you will know when you have hit it.
The third is where the data is allowed to go. Residency requirements, air-gapped environments, and client contracts that forbid third-party processing are not solved by a vendor’s privacy policy. They are solved by the inference happening somewhere you control.
If none of those three applies, rent. The right move for most small teams is to keep buying the best closed model for the work and to route cheaper tasks elsewhere, which is a separate discipline covered in routing tasks between AI models and open models good enough for operators.
Fine-tuning is a data problem, and the price list proves it
Tinker is Thinking Machines’ training API. It uses LoRA and runs the infrastructure for you, so you control the training loop without owning the cluster [1][3]. Its price list is instructive precisely because the compute is so cheap. Training on Inkling at a 64K context costs $5.61 per million tokens, with prefill at $1.87 and sampling at $4.68. Inkling-Small trains at $1.73. gpt-oss-120B trains at $0.737 per million tokens and gpt-oss-20B at $0.396. Checkpoint storage is $0.10 per gigabyte-month [3].
Read those numbers again. A five-million-token training set costs about $28 to train on Inkling, and under $4 on gpt-oss-120B [3]. Storing the result costs pennies. Compute is not the constraint and has not been for a while.
The constraint is examples. Fine-tuning needs a few hundred to a few thousand instances of the correct answer, in your format, in your voice, produced or checked by someone who knows what correct looks like in your business. That is unglamorous labour, it cannot be outsourced to the model you are trying to improve, and it is the reason most fine-tuning projects at small companies stall. Before you price a fine-tune, count how many labelled examples you can actually produce this month. If the answer is under a hundred, keep prompting.
Hardware decides what “local” is allowed to mean
Two very different things get called owning a model, and conflating them is the most expensive mistake in this whole area.
The first is a small model on hardware you already have. gpt-oss-20b is a 14GB download that runs “on systems with as little as 16GB memory”, with a 128K context window [8]. That is ownership in the plain sense: no per-token bill, no network call, no third party holding your data. If privacy or offline operation is your reason for looking at open weights, a 20B-class model on a machine you already own may solve the problem outright, and it costs nothing to test.
The second is a frontier-scale open model. gpt-oss-120b is a 65GB download that targets a single 80GB GPU [8]. Inkling is 975 billion total parameters, distributed in BF16 and NVFP4 checkpoints [1][2]. Nobody is running that on a laptop. Owning it means renting machines, which means you have swapped a per-token bill for an hourly one plus the job of keeping the thing up. That trade can be correct. It is never automatic.
Run the decision in two weeks, in this order
Start with your own bill, not with a model. Pull the last three months of API spend and split it by workload into input and output tokens. Most teams find that one workload accounts for the large majority of the cost. That is the only one worth moving, and every hour spent on the others is wasted.
Then evaluate through a per-token host before touching any hardware. Send the same prompts and the same real inputs to the open-weight model and score the results against your own past outputs, not against a public benchmark. At Together’s Inkling rate, a serious evaluation over a few hundred real cases costs single-digit dollars [4], which means there is no excuse for deciding this on reputation.
Only if quality holds and one of the three reasons applies should you price a fine-tune, and only after a fine-tune proves out should you price dedicated capacity. Most teams stop at step two and stay there happily: they get the licence, the portability and the lower per-token rate without operating anything. That is the outcome to aim for, not the consolation prize.
Defaults are Together's list price for Inkling [4]. Compare the result with a dedicated H100 at $5.49/hour, about $4,000 a month left running [4]. Computed in the page; nothing is sent anywhere.
What still goes wrong
Every price in this guide is a list price on one day. Launch discounts expire: Inkling shipped with a “50% discount for a limited time” [1], and Together’s H100 rate is currently shown as a promotional $3.99 against a $5.49 on-demand list [4]. If your case for switching depends on a promotional rate, you do not have a case, you have a window. Re-run the numbers at list price before you commit anything you cannot reverse in a month.
The portability you bought is partial. The weights are yours, but the switching cost was never the weights. It is the prompts, the evaluation set, the retrieval plumbing, the tool definitions and the accumulated small fixes that make the thing work on your data. Moving between models is mostly the work of rebuilding that, which is why holding a copy of Inkling does not by itself make you free of anyone. The exit ramps that matter are covered in AI vendor lock-in exit ramps.
Fine-tuned models also age badly. Your adapter is bound to a base model that will be superseded, and the improvements in the next base model arrive without your customization unless you redo the work. Budget for retraining rather than treating a fine-tune as a finished asset. And none of this helps a team whose real requirement is a signed data processing agreement and a named counterparty to sue; open weights give you control, not accountability, and for some clients accountability is the thing being bought.
- 01Thinking Machines Lab — Introducing Inklingthinkingmachines.ai
- 02Hugging Face — thinkingmachines/Inkling model cardhuggingface.co
- 03Tinker docs — Supported models and pricingtinker-docs.thinkingmachines.ai
- 04Together AI — Pricingtogether.ai
- 05Anthropic — Claude model pricingplatform.claude.com
- 06Meta — Llama 4 Community License Agreementraw.githubusercontent.com
- 07Open Source Initiative — The Open Source AI Definition v1.0opensource.org
- 08Ollama — gpt-oss model libraryollama.com