Deciding when an open model is good enough
Judge an open model on one job you actually do, know which of open's three promises you are buying, and settle the question in an afternoon.
on this page · 0 / 0 checked
Every few months the same claim comes round with a new model name attached. An open model has caught up, it is free to download, and you can stop paying the subscription. Then you look at the leaderboard, find the open entry sitting a few points behind the leader, and have no way to tell whether those few points are the difference between a usable first draft and a wasted afternoon. The claim is usually true. It is just not a decision.
This guide turns it into one. It is for a solo operator or a small team already paying for an assistant or an API bill who wants to know whether an open-weight model can take over part of that work. It is not for anyone planning to fine-tune a model, and it is not for a team with a compliance function that will make this call regardless of what you find. Every score, price and limit below was read from the source’s own page on 5 September 2026, and all of them move.
The gap is small, and it is the wrong thing to measure
Artificial Analysis ranks 196 models on a single intelligence index, 99 of them open weights [1]. At the top sits Claude Fable 5.1 in its highest reasoning setting, scoring 66 [1]. The best open-weights entry is Kimi K3, scoring 60 [1]. Six points, on a scale where the whole top twenty is squeezed between 59 and 66 [1].
Six points also buys less separation than the ranking implies. Kimi K3 sits seventeenth on that list, one point below two configurations of GPT-6 Astra and level with a third [1]. A single number averaged over a fixed set of evaluations is doing a lot of sorting across very little distance, and none of that distance was measured on your work.
The work you were going to hand over is rewriting, summarising, pulling structure out of notes, drafting the email you have written forty times, first-pass code you will read anyway. On those jobs both models are operating well below their ceiling, and the six points get spent on capability neither of them needs to reach. That is the reframe worth keeping: an index tells you which model to reach for when a task is genuinely hard, and almost nothing about a task that is easy for both.
So the question is not whether the open model is as good. It is whether the specific job you have in mind is one where the gap shows up at all. That is answerable, and you are the only person who can answer it, because your work is the only benchmark that is about your work.
Open weights buy three separate things, and you want one of them
The word “open” gets used to mean three different benefits, and people who are sold on one of them often end up testing for another.
The first is that nobody can switch it off. Weights you have downloaded keep working whatever the company that made them does next, and whatever it decides to charge. Nothing else in your stack has that property.
The second is price competition, and it is the one with numbers attached. Kimi K3 is served through OpenRouter by eighteen different companies. The lowest input price on that list is $2.50 per million tokens, on a route charging $14 per million output; Moonshot’s own service charges $3.00 and $15.00; the dearest route runs to $6.00 and $22.50 [4]. Claude Fable 5.1, the model above it on the index, is $10 per million input and $50 per million output from one seller, because there is only one seller [8]. The saving is real, but the structural point matters more: with an open model, no single company sets the price, and if one of them raises it you change a base URL.
At the cheap end the numbers stop being comparable at all. OpenAI’s gpt-oss-120b is served by eighteen providers starting at $0.03 per million input tokens and $0.17 per million output [5]. That is not a discount on $10 and $50. It is a different category of expense, and it is what makes automated, high-volume, low-stakes work worth doing at all.
The third benefit is that your material never leaves your machine, which is the one people say first and test least. It is also the only one of the three that requires real work, and the next section is about why.
Renting an open model is not the same as running one
Almost everyone who says they tested an open model called someone’s API. That is the sensible way to start, and it does deliver two of the three benefits. It does not deliver the third, and the difference is easy to miss because the model file is genuinely the same one.
Read the data policy on the route you are actually using, not the model’s home page. OpenRouter states that it does not store your prompts or responses unless you have explicitly opted into prompt logging in your account settings [6]. Separately, it lets you set “whether you would like to allow routing to providers that may train on your data (according to their own policies)” [6]. That second setting is the one worth reading twice. OpenRouter says it works with providers “wherever possible” to ensure prompts will not be trained on, “but there are exceptions” [6]. Open weights say nothing about the conduct of the person hosting them.
What renting does give you is portability of a kind you cannot get from a closed vendor. When you move between two hosts of the same open model, the model itself is unchanged, so your prompts, your evaluations and your expectations all survive the move. When a closed vendor retires a model, the thing you tuned your prompts against is gone and there is no second supplier to go to. That is the exit ramp, and it is worth something even if you never use it.
What you can actually run on your own machine is a smaller model
Here is where the headline and the practical option come apart. Kimi K3, the open model at the top of the open-weights list, has 2.8 trillion total parameters with roughly 104 billion active per token and a context window of 1,048,576 tokens [2]. Nothing about the licence stops you downloading it [3]. Physics and your budget do. It is a data-centre model that happens to be public.
The models you can genuinely run are further down. Ollama distributes OpenAI’s gpt-oss in two sizes: the 20B version is a 14GB download and runs in “as little as 16GB memory”, and the 120B version is a 65GB download that is built to “fit on a single 80GB GPU” [7]. Both carry a 128K context window and a permissive Apache 2.0 licence [7]. A 16GB laptop is an ordinary machine. An 80GB GPU is not.
Treat these as two decisions wearing one word. “Should I use an open model” is usually a question about price and portability, and it is answered by renting the best open model available. “Should I run a model myself” is a question about data never leaving a machine you control, and it is answered by a considerably smaller model that you should test on its own terms. Confusing them is how people conclude that open models are disappointing, having compared a laptop-sized model against a hosted frontier one.
The licence is short, and its thresholds are far above you
“Open weights” covers a range, from Apache 2.0 with no conditions worth reading [7] to bespoke agreements with revenue triggers in them [3]. The good news for a small operator is that reading the licence takes five minutes and almost always ends in relief.
Take the stricter of the two here. The Kimi K3 License grants the right to use, copy, modify, distribute and sell copies of the model, on two conditions that only bite at scale. If you run a model-as-a-service business and your revenue passes $20 million over any consecutive twelve months, you must enter a separate agreement with Moonshot AI. If you deploy it in a commercial product with more than 100 million monthly active users or more than $20 million in monthly revenue, you must display “Kimi K3” prominently on the user interface. Internal use, and use through Moonshot’s own products or its certified inference partners, is exempt from both [3]. gpt-oss carries none of that, being plain Apache 2.0 [7].
Read the licence once anyway, because the thresholds are the part that changes between models. Then file it and stop thinking about it. The licence is not the reason this decision is hard.
The test that settles it takes an afternoon
Pick one recurring job. Not your workflow, one job, the kind you do weekly and could describe to a temp. Collect ten real items you have already completed, so you know what a good answer looks like without having to judge in the abstract.
Run the same prompt, word for word, through your paid tool and through the open model. Then grade by counting, not by impression: for each of the twenty outputs, how many edits stand between it and something you would send. Impressions favour whichever model you were rooting for. Edit counts do not.
Keep the prompt in a plain text file outside both tools while you do this. It costs nothing, and it means the switch back is as cheap as the switch over. Then decide on a narrow role rather than a replacement. The realistic outcomes are that the open model takes one routine job, becomes the thing you fall back on when your main vendor has an outage, or handles the volume work where the price difference actually compounds.
Defaults are Moonshot's own list price for Kimi K3 [4]. Run it again with 10 and 50 for Claude Fable 5.1 [8], and with 0.03 and 0.17 for gpt-oss-120b [5]. Computed in the page; nothing is sent anywhere.
What still goes wrong
The numbers age fast. Six points between the best proprietary and the best open entry is a 5 September 2026 reading, taken on a day when twenty entries were spread across seven points [1]. One model shipping a new reasoning setting reorders that. The prices move as often. Re-run the comparison when you notice yourself quoting a figure you have not checked, and do not build a business case on a leaderboard position.
Serving is not standardised. Eighteen companies sell access to the same Kimi K3 weights at input prices from $2.50 to $6.00 per million, and the same table reports round-trip latency from 0.35 seconds to 8.53 seconds, throughput from 6 to 115 tokens per second, and three-day uptime from 86.92% to 99.98% [4]. Identical weights do not buy identical service. If a route stops behaving the way it did, the model file is not the only variable to check.
The largest gap is not the model. It is everything wrapped around it. Claude Pro at $17 a month buys chat on web, iOS, Android and desktop, memory across conversations, projects, web search and file creation with code execution [8]. An open model gives you an engine, and if you self-host, the support desk is you at 11pm. That trade is fine when it is deliberate and miserable when it arrives as a surprise, which is the real argument for testing one narrow job rather than migrating. And if you cannot name the job the open model would take, the honest answer is that you are interested in the topic rather than the tool, which is a fine thing to be, but it is not a reason to spend a Saturday.
- 01Artificial Analysis — LLM Leaderboardartificialanalysis.ai
- 02Hugging Face — moonshotai/Kimi-K3 model cardhuggingface.co
- 03Moonshot AI — Kimi K3 Licensehuggingface.co
- 04OpenRouter — Kimi K3 providers and pricingopenrouter.ai
- 05OpenRouter — gpt-oss-120b providers and pricingopenrouter.ai
- 06OpenRouter Docs — Privacy and Loggingopenrouter.ai
- 07Ollama — gpt-ossollama.com
- 08Claude — Pricingclaude.com