tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

Use Chinese models without sending your data to China

Route high-volume, low-risk work to the cheapest tier on the market, price the saving honestly against your current model, and choose where the tokens are actually processed.

Published 2026-09-05 · Updated 2026-09-05 · Read 10 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

Your model bill has one line larger than all the others, and it is almost certainly a background job. Something classifies inbound email, or turns call transcripts into summaries, or drafts the first version of a listing that a person then rewrites. It runs thousands of times a month on the model you originally picked for hard work, because picking a second model for easy work was a decision you never got round to making. That job is where the money is, and it is the only part of your stack where the choice of model is worth an afternoon.

The cheapest tokens among the most-used models come from Chinese labs, and they are cheap by a factor that changes how you plan rather than by a few percent [3][6]. That fact arrives bolted to a second question about where your data ends up, and the two get argued as one thing. They are separable, and separating them is most of the job. This guide is for someone paying an API bill by the token with at least one high-volume automated workflow. It is not for anyone whose entire AI spend is a $20 chat subscription, because there is nothing there to route. It is not procurement advice either, and if a client contract names the vendors you are allowed to process their data with, that contract decides this, not arithmetic.

The cheap tier is where the volume went

OpenRouter ranks models by the number of tokens each one processed through its API, counting both prompt and completion tokens [1]. In its rankings through 3 September 2026, the top five were GLM 5.3 Flash from z-ai at 11.9 trillion tokens, OpenAI’s GPT-5.6 Luna at 11.6 trillion, DeepSeek V4 Flash 0731 at 11.3 trillion, and two Tencent models, Hy4 preview at 11 trillion and Hy3 at 5.25 trillion [1]. Four of the five come from Chinese labs.

Read that for what it is. OpenRouter says outright that the rankings measure adoption, not quality, and that they cover traffic routed through OpenRouter rather than the whole market or usage on a provider’s own API [1]. Routers attract price-sensitive traffic by design. What the table tells you is that routing high-volume work to a Chinese model stopped being an unusual thing to do somewhere in the last year, and that if you are still sending every token to a frontier model out of habit, you are now in a smaller group than you think.

The prices behind that shift are public. GLM 5.3 Flash lists at $0.075 per million input tokens and $0.25 per million output, currently with a 50% discount applied, and 23 providers serve it [3]. DeepSeek V4 Flash 0423 is served by 16 providers, whose prices for it run from $0.0679 input and $0.168 output at DigitalOcean up to $0.21 and $0.56 at the endpoint listed as Azure (US) [2]. For comparison, Claude Sonnet 5 is $2 and $10 per million, Claude Haiku 4.5 is $1 and $5, and Claude Opus 5 is $5 and $25 [5]. OpenAI’s short-context standard rates are $5 and $25 for gpt-6-astra, $2 and $10 for gpt-5.6-sol, and $0.10 and $0.60 for gpt-5.6-luna [6].

That last number matters more than any of the others, and it is the one people skip. The US labs now ship a cheap tier too. The distance from Claude Sonnet 5 to GLM 5.3 Flash is enormous. The distance from GPT-5.6 Luna to GLM 5.3 Flash is not.

The saving lives in the output tokens

Take a real shape. A triage workflow that consumes 40 million input tokens and produces 8 million output tokens in a month, which is a busy but unremarkable automation. On Claude Sonnet 5 at $2 and $10 that is $80 of input and $80 of output, so $160 [5]. On Claude Haiku 4.5 at $1 and $5 it is $80 [5]. On GPT-5.6 Luna at $0.10 and $0.60 it is $4 of input and $4.80 of output, so $8.80 [6]. On GLM 5.3 Flash at $0.075 and $0.25 it is $3 and $2, so $5 [3]. On DeepSeek V4 Flash 0423 at its cheapest listed host it is about $4.06 [2].

The shape of that ladder is the whole lesson. Moving from Sonnet 5 to the cheapest tier saves about $155 a month. Moving from the cheapest US tier to the cheapest Chinese tier saves about $4. The first move is a business decision. The second is a rounding error that people spend weeks arguing about because it has a flag on it.

Notice also that input and output are priced differently, by a factor of 5 or 6 across both published rate cards [5][6]. Two levers attack the two halves. Prompt caching attacks input: Claude charges cache reads at 0.1 times the base input price [5], and OpenAI lists cached input for gpt-5.6-sol at $0.20 against $2 standard [6]. Batch processing attacks both, at 50% off on Claude’s Batch API [5] and the same on OpenAI’s batch rates [6]. If your workflow has a long fixed system prompt and short variable inputs, caching may get you most of the way without changing model at all. If your workflow is short in and long out, caching does almost nothing and only a cheaper tier moves the number.

Do this arithmetic on your own token counts before you do anything else. Not an estimate of your token counts. The actual figures from a month of usage, which every provider dashboard will export.

calculator
Monthly saving from a model switch
— $ / month

Your monthly tokens multiplied by the price difference on each side. Defaults compare Claude Sonnet 5 with GLM 5.3 Flash at list prices read on 5 September 2026. Computed in the page; nothing is sent anywhere.

Route by what reads the output, not by what you have heard about the model

Every model recommendation you read is about a benchmark someone else ran on a task that is not yours. The property that actually decides whether a cheaper model is safe in a given workflow is duller and you already know it: what happens to the output next.

If a person reads the output before it does anything, a cheaper model costs you occasional rework and nothing else. First drafts, internal summaries, suggested tags a human confirms, translations that go to a reviewer. If a deterministic check reads the output, the same applies with less effort. Extraction into a schema that either validates or does not, classification into a fixed set of labels where an unknown label is rejected, code that is run by tests. In both cases a wrong answer is caught by machinery you already have, so the cost of being wrong is measured in retries rather than in reputation.

If the output goes straight to a customer, a regulator, a payment system or a public page with nobody in between, the cheaper model is not a cost decision at all. It is a change to your quality floor, and you should treat it as one.

That framing also tells you how to test. Take 50 real inputs from production, including several that went badly. Run them through the current model and the candidate with identical prompts. Have someone who does not know which output came from which pick the better one, or mark them equivalent. Count the equivalents. Nobody’s published benchmark substitutes for this, because your prompt, your data and your definition of good are all specific to you.

Where the tokens are processed is a separate choice from which model runs

Here is the part that dissolves most of the argument. The cheap models above ship their weights publicly. DeepSeek V4 Flash is published on Hugging Face under the MIT License, a 284 billion parameter mixture-of-experts model with 13 billion parameters active and a 1 million token context [4]. OpenRouter’s GLM 5.3 Flash page links to that model’s weights on Hugging Face and lists 23 providers serving it [3]. DeepSeek V4 Flash 0423 is served by 16, including DigitalOcean, DeepInfra, CoreWeave, Parasail, Alibaba Cloud Int. and Baidu Qianfan [2].

So “which model” and “whose servers” are two dials, not one. Calling DeepSeek’s own API is a decision to send your prompts to DeepSeek, and DeepSeek’s privacy policy states that it stores the information it collects in secure servers located in the People’s Republic of China [8]. Running the same weights through DeepInfra, CoreWeave or the endpoint OpenRouter labels Azure (US) sends the identical prompts to a different company under a different contract [2]. The model is the same. The counterparty, the jurisdiction and the contract you can enforce are not.

This is why the country-of-origin framing produces bad decisions in both directions. It talks people out of a 95% cost reduction on a workflow that processes nothing sensitive, and it talks other people into pasting customer records into a first-party endpoint because the model scored well, without ever reading whose servers that endpoint runs on. The question was never whether the model is Chinese. It is which company receives the bytes, under what terms, and whether you could tell a client the answer without flinching.

The same logic applies in reverse to US vendors. Where processing happens is priced as a feature: OpenAI charges a 10% uplift on regional processing endpoints for data residency, on models released on or after 5 March 2026 [6]. Data location is a product decision with a price tag attached, not a property of a flag.

Keep the switch to one line of configuration

A routing decision you cannot reverse in five minutes is not a strategy, it is a migration. Two habits keep it cheap.

Call models through a layer that speaks to several providers rather than importing one vendor’s SDK at every call site. Then changing model is a configuration value and changing host is another one. If you have baked a provider’s client library into eleven files, every future price comparison starts with a refactor, and so you will not run one.

Then make the data policy explicit rather than inherited. On OpenRouter, the only parameter is a list of provider slugs to allow and ignore is a list to skip, order is a list of slugs to try in order, data_collection set to deny controls whether providers that may store data get used at all, and zdr set to true restricts routing to zero data retention endpoints [7]. There is a quantizations filter too, taking values such as int4, int8 and fp8 [7], which matters because two hosts serving the same open weights at different quantization are not serving quite the same model.

Whatever router or wrapper you use, the equivalent settings exist, and the default is almost always the permissive one. Set them once, in code, next to the model name, so the next person to change the model inherits the policy instead of rediscovering it.

checklist
Before you move a workflow to a cheaper model
0 of 7 · saved in this browser only

What still goes wrong

Prices move, and a good share of the numbers above are promotional. GLM 5.3 Flash’s listed price currently has a 50% discount applied [3], and OpenAI states that GPT-5.6 Sol’s promotional pricing is available at least through 21 November 2026 [6], which is a polite way of saying it will be repriced. A cost model built on introductory rates is a cost model with an expiry date. Re-run the comparison every six months and after any release you were tempted by, and keep the arithmetic somewhere you can find it rather than in the tab you closed.

Open weights solve the jurisdiction question and do not solve the others. You still have no meaningful visibility into training data, and a model can carry preferences or refusals that show up in exactly the topic you care about and nowhere in any benchmark. Test on your own material, including the awkward parts of it. Serving is uneven as well. The same weights across the 16 providers listed for DeepSeek V4 Flash 0423 run from $0.0679 to $0.21 per million input tokens, a spread of more than three times [2], and hosts differ in quantization, which is why OpenRouter exposes quantization as a routing filter at all [7]. An evaluation you ran against one host does not automatically transfer to the cheapest one.

The most common failure is smaller than any of this. You spend an afternoon on a comparison, find a 95% saving, and the saving is $150 a month on a business where an afternoon of your time is worth more than that for a year. Check the size of the prize before you chase it. If the workflow you are optimising is not costing you at least a few hundred dollars a month, the correct move is to leave it on whatever it is running on and go and do something that makes money instead.

sources
  1. 01OpenRouter — Model rankingsopenrouter.ai
  2. 02OpenRouter — DeepSeek V4 Flash 0423 providers and pricingopenrouter.ai
  3. 03OpenRouter — GLM 5.3 Flash providers and pricingopenrouter.ai
  4. 04Hugging Face — deepseek-ai/DeepSeek-V4-Flash model cardhuggingface.co
  5. 05Claude Docs — Pricingplatform.claude.com
  6. 06OpenAI — API pricingdevelopers.openai.com
  7. 07OpenRouter Docs — Provider routingopenrouter.ai
  8. 08DeepSeek — Privacy Policychat.deepseek.com
next guide
The benchmark that separates models is the long one
9 min · verified 2026-09-05
related guides