Chips that can only run one model, and what that changes for you
How etching a model into silicon works, why every AI vendor already charges you less for standing still, and how to spot the work that qualifies.
on this page · 0 / 0 checked
A chip company can now take a model, etch its weights directly into the silicon, and ship a part built around that one model [2]. AMD announced on 6 August 2026 that it had agreed to buy Taalas, the Toronto company doing this, subject to customary closing conditions and regulatory approvals [1]. You will not buy one of these chips. You will probably never knowingly touch one. The reason it is worth ten minutes is the argument underneath it, which applies to AI work you are already paying for.
The argument is that flexibility costs money, and once a workload stops changing, you are paying for flexibility you are not using. That trade is already written into every vendor’s price list, in three places, waiting for anyone who wants it. This guide is about finding the work in your business that has stopped changing, moving it onto the cheaper terms, and knowing what you have signed up for when you do. It is written for people running a metered AI workload, an automation or a script or an API key that bills by the token. If your entire AI use is a chat subscription and a few hundred messages a month, nothing here will move your bill, and you can skip it.
Etching the weights into the chip removes the part that was slow
A normal AI accelerator is a general-purpose machine. The weights sit in memory, the compute sits somewhere else on the card, and the weights are hauled across to the compute over and over while the model generates. AMD’s own description of what it is buying names that cost directly: technology that “optimizes inference dataflows, significantly reducing compute and memory bottlenecks associated with general-purpose architectures” [1].
Taalas shortens the trip. The Register reports that the company’s first test chip, which it called the HC1, was fabbed on TSMC’s 6nm process and has two main regions, a mask-ROM recall fabric where the model weights are etched and an SRAM recall fabric where KV caches and fine-tuning adapters are stored [2]. The same report says the chips “don’t rely on HBM to store the model weights but rather etch them directly into the silicon” [2]. The weights are not held near the compute, they are part of it. Serving Meta’s Llama 3.1 8B, initial benchmarks put HC1 at 16,960 tokens a second, which the report says was 48 times faster than Nvidia’s GPUs and 8.5 times faster than Cerebras’ accelerators when the chip was announced in February [2].
The price of that is severe. Once the chip is made, you are close to stuck with that model. The Register notes that only two layers of metal need to be changed to deploy a new model, which makes updates cheaper and less time-consuming than a full re-spin, and that changes bigger than something like a LoRA adapter require a re-spin anyway [2]. A second-generation part, the HC2, was reported as due out in summer 2026 and aimed at 20 billion parameters [2]. That is the whole shape of the bet in one sentence: give up most of the ability to change the model, get an order of magnitude on the thing you do a billion times.
The direction is inference-specific silicon, not one startup’s idea
It would be easy to file this as a curiosity from one small company that got bought. The broader pattern says otherwise. Google describes its TPUs as “custom accelerators co-designed with open software to power the entire AI lifecycle” and “purpose-built for advanced AI workloads” [3]. Its seventh generation, Ironwood, is generally available, and the next part in the line, TPU 8i, is listed as coming soon, “optimized for post-training and inference”, with an “80% performance-per-dollar improvement over previous generations for low-latency inference for large MoE models” [3].
None of those are model-specific in the Taalas sense. They still run whatever you load. But they are moving along the same axis, which is the one worth watching: hardware is being narrowed from “runs any AI workload” toward “runs inference on the models people actually run, cheaply”. Taalas is the far end of that axis rather than a different axis. The far end tells you which way the near end is travelling.
You collect this trade through the price list, not through hardware
Here is the translation for someone who buys tokens rather than wafers. You will never make a purchasing decision about model-specific silicon. What reaches you is the price of a token, and the price lists already contain the same bargain in three forms.
The first is the model tier. On Claude, Fable 5.1 and Mythos 5.1 are $10 per million input tokens and $50 per million output, Opus 5 is $5 and $25, Sonnet 5 is $2 and $10, and Haiku 4.5 is $1 and $5 [4]. On OpenAI, GPT-6 Astra is $10 input and $50 output, while GPT-5.6 Luna is $0.20 and $1.20 [5]. That is a 50-fold spread on input between the top and bottom of one vendor’s list. Google prices Gemini 3.5 Flash-Lite at $0.30 input and $2.50 output against $0.75 and $3.75 for Gemini 3.8 Flash through 31 December 2026 [6]. The small model is the commoditised, heavily optimised one, and it is priced like it.
The second is batch. Anthropic, OpenAI and Google all take 50 percent off if you send work to a batch endpoint instead of demanding an answer now [4][5][6]. Anthropic applies the discount to both input and output tokens, and its worked example is Sonnet 5 falling from $2 and $10 to $1 and $5 [4]. Nothing about the model changes. You are being paid half your bill for surrendering immediacy on work that did not need it.
The third is caching, which is the most literal version of the argument. Anthropic charges cache hits at 0.1 times the base input price, so $0.20 per million on Sonnet 5, and at 0.025 times on Fable 5.1 and Mythos 5.1, which is $0.25 against a $10 list price [4]. OpenAI lists cached input on GPT-6 Astra at $1.00 against $10.00 standard [5]. Caching pays you specifically for sending the same prefix again. It is a discount for not changing your prompt.
The work that qualifies has stopped changing, and you can list it in an afternoon
The test is not whether the task is important. It is whether the task’s shape has been stable for a month, and whether you would notice if the model behind it were swapped. Extracting six fields from an inbound invoice. Tagging support email by topic. Turning a call transcript into a fixed template. Producing the first draft of a listing description from a spec sheet. These run on the same prompt with different nouns in it, hundreds or thousands of times, and the output is checked by a human or by a rule before it matters. That is the profile.
The opposite profile is the work where you are hiring judgement rather than throughput. The proposal that wins or loses a client, the contract read, the hard debugging session, anything where you would rather have one good answer than four hundred adequate ones. Leave that on the expensive model and stop thinking about it. The saving is not there anyway, because the volume is not there.
Once you have the list, do the arithmetic before the migration, because it decides whether the migration is worth an afternoon.
runs × 30 days × tokens at the published per-million rates [4][5][6]. Run it twice, once at the prices of the model you use now and once at the cheap tier's, and the gap is what standing still is worth. Halve the result again if the work can go through a batch endpoint [4]. Computed in the page; nothing is sent anywhere.
The migration itself is boring and should stay boring. Take thirty real inputs from last month, run them through both models, and read the two sets of outputs side by side rather than reading a benchmark. Keep those thirty inputs and the outputs you accepted, because that set is now your regression check, and you will want it the next time a vendor ships something and you wonder whether to move.
Committing to a model means owning its retirement date
Standing still has one real cost, and it is not capability. It is that the model you standardised on is a product with an end date, and the end is enforced.
Anthropic states that it notifies customers with active deployments for models with upcoming retirements, “providing at least 60 days’ notice before model retirement for publicly released models”, and that “requests to models past the retirement date will fail” [7]. Its published table gives a tentative retirement date per dated model ID, with claude-sonnet-4-5-20250929 listed as not sooner than 29 September 2026 and claude-haiku-4-5-20251001 as not sooner than 15 October 2026, while several older IDs are already retired, including claude-opus-4-1-20250805 on 5 August 2026 [7]. OpenAI gives a longer runway, at least 6 months for generally available models and at least 3 months for specialised variants, but warns that preview models “may be retired with much shorter notice, such as 2 weeks” [8]. Its schedule has several GPT-5 and o3 snapshots, including gpt-5-2025-08-07 and o3-2025-04-16, shutting down on 11 December 2026 [8].
Prices move on a schedule too, and not always downward. Google lists Gemini 3.8 Flash at $0.75 per million input tokens “through December 31, 2026” and $1.50 starting 1 January 2027, with output going from $3.75 to $7.50 on the same date [6]. If you sized a workload on the introductory price, your bill for that model doubles overnight on 1 January 2027.
So the discipline that makes standing still safe is clerical. Write down, in one document, every place a model name is hardcoded: the API keys, the scripts, the Zapier steps, the Make scenarios, the n8n nodes. Record the exact model ID, not the family name, because the retirement schedule is published per ID [7]. Put the earliest retirement date on that list in your calendar with two months to spare, and keep the thirty-item regression set from the previous section next to it. That is the whole cost of commitment, and it is an hour a quarter.
What still goes wrong
The hardware story is not something you can act on, and honest reporting of it stops well short of a recommendation. AMD’s announcement carried no product timeline, no pricing and no list of models, and the deal was still subject to customary closing conditions and regulatory approvals when it was announced [1]. The performance figures come from reporting on initial benchmarks of one first-generation test chip, against a comparison drawn in February, not from an independent benchmark you can reproduce [2]. Treat the whole thing as a signal about where inference costs are heading, not as a purchase you are preparing for.
The cheap tier is also not automatically cheaper. A small model that needs two attempts and a repair prompt to produce an acceptable answer can cost more than one call to the expensive one. Output is the expensive half everywhere: five times input on every Claude model listed above [4], six times on GPT-5.6 Luna [5], and more than eight times on Gemini 3.5 Flash-Lite [6]. A chattier small model erodes its advantage fast. The cost that never appears in the calculator is your review time, and a model that quietly fails a fraction of the time on work nobody checks is not a saving, it is a liability with a lower invoice.
Finally, standing still has a capability cost that accrues quietly. The model you pinned last year was competitive when you pinned it. Nobody sends you a notice when it stops being competitive, only when it is being switched off. The regression set is the defence: rerun it against the current default model twice a year, and let the outputs, rather than the release notes, tell you whether the thing you froze is still good enough to leave frozen.
- 01AMD — AMD Acquires Taalas to Advance Compute Solutions for Rapidly Growing AI Inference Marketnewsroom.amd.com
- 02The Register — AMD acquires AI chip startup Taalas to boost inference performance by etching models into silicontheregister.com
- 03Google Cloud — Cloud TPUcloud.google.com
- 04Anthropic — Claude API pricingplatform.claude.com
- 05OpenAI — API pricingdevelopers.openai.com
- 06Google — Gemini API pricingai.google.dev
- 07Anthropic — Model deprecationsplatform.claude.com
- 08OpenAI — Deprecationsdevelopers.openai.com