The economics of running AI models on your own machine
Work out whether a local-AI machine beats paying for tokens, using memory bandwidth, real API prices and the month it breaks even.
on this page · 0 / 0 checked
Every year or so a machine arrives that can hold a very large model in memory, and the same argument starts again. Buy the hardware once, stop paying for tokens, keep your data in the building. Apple’s current version of that machine is the Mac Studio with M5 Ultra, which starts at $5,499 and can be configured with up to 512GB of unified memory at 1.2TB/s, which Apple says is 50 percent higher bandwidth than before [1]. Apple’s own framing is that you can “run enormous LLMs entirely on device” and “load the largest and most demanding frontier-class open-weight models available today” [1]. Both statements are true. Neither of them is a business case.
This guide is the arithmetic you do before you spend the money: what memory capacity actually buys, what it does not, what your alternative really costs, and the point at which owning beats renting. It is written for solo operators, freelancers and small teams weighing a single machine against a monthly bill. It is not for anyone serving a model to thousands of concurrent users, and it is not a benchmarking guide. If you need throughput numbers for a specific model, you need a benchmark on that model, not an article.
Capacity decides what runs, bandwidth decides how fast it runs
These are two different specifications and the marketing collapses them. Capacity is whether the model loads at all. A model that does not fit in memory does not run, so capacity is a hard gate: below the line nothing works, above it everything does. That is why the 512GB figure gets the headline. It moves the gate.
Bandwidth is how fast weights move while the model is generating, and it sets the speed you actually experience. Here the comparison is less flattering. The M5 Ultra moves data at 1.2TB/s [1]. NVIDIA’s H200 has 141GB of HBM3e at 4.8TB/s, which NVIDIA describes as nearly double the capacity of an H100 with 1.4 times more memory bandwidth [3]. That is a four-fold bandwidth gap against a single datacentre card, and it does not close because you bought more memory. Long, sequential generation feels that gap most: the tokens arrive at the pace the weights can be read.
So the honest summary of what the money buys is a gate, not a speed. You get to run a model that no other machine at this price can hold. You do not get to run it at datacentre pace. If your work is a person waiting on a response, the bandwidth number is the one you will live with every day. If your work is a batch job that runs overnight, it barely matters.
The headline price is not the price of the machine you want
The $5,499 Mac Studio does not have 512GB. The M5 Ultra’s base configuration is 96GB of unified memory [2], and the 512GB option is only offered on the top bin, with the 36-core CPU and 80-core GPU [2]. Apple also stated that the 512GB configuration would not ship with the rest of the line, arriving in late October rather than on the 22 September availability date [1]. Education pricing starts at $5,099 for M5 Ultra and $2,299 for M5 Max, against standard prices of $5,499 and $2,499 [1].
This pattern is not specific to Apple and it is the most reliable way people misjudge local-AI hardware. The capacity in the headline and the price in the headline belong to two different machines. Before comparing anything, price the exact configuration that holds the exact model you intend to run, with the storage you need for several model files, and use that number in every calculation that follows.
It is also worth checking whether you need the top of the range at all. The M5 Max reaches 128GB at up to 614GB/s from $2,499 [1][2]. A good number of open-weight models people actually use in production sit inside that envelope once quantised, and the cheaper machine has roughly half the bandwidth rather than a quarter of it. Buying the largest configuration to run a model you have not yet chosen is the expensive version of this mistake.
Rented hardware is the cheapest way to find out
You can settle most of this empirically for the price of a day. Google Cloud lists on-demand accelerator-optimised machines in Iowa at $88.49 an hour for an A3 High with eight H100s, $84.81 an hour for an A3 Ultra with eight H200s, and $64.44 an hour for an A4 High with eight B200s [4]. Eight H200s is 1,128GB of memory [3], more than twice what the largest Mac Studio holds.
Two things follow. The first is that renting is a diagnostic, not a lifestyle: at $84.81 an hour, running that machine continuously is just over $2,000 a day, which is why nobody sensible leaves one on. The second is that a day of it costs about $2,000 at most and usually far less, and it will tell you whether your model, your prompts and your actual documents produce output you would accept. Spending one day’s rental to de-risk a five-figure purchase is the correct order of operations.
What renting will not tell you is speed on Apple Silicon specifically, because the hardware is different. For that, the substitute is finding published measurements for your specific model on the specific configuration you are considering, and treating anything else as a guess. The number that matters is tokens per second on your workload, and it is not on any spec sheet.
Most API bills are too small to buy a machine with
This is where the case usually collapses, and it collapses quietly because people compare the hardware against a feeling rather than against an invoice.
Current published prices per million tokens: Claude Sonnet 5 is $2 input and $10 output, Haiku 4.5 is $1 and $5, and Opus 5 is $5 and $25 [5]. OpenAI lists GPT-5.6 Luna at $0.20 input and $1.20 output on short context, and GPT-5.6 Terra at $2 and $12 [6]. Gemini 3.8 Flash is $0.75 input and $3.75 output through 31 December 2026, rising to $1.50 and $7.50 on 1 January 2027 [7].
Put a real workload through that. Ten million input tokens and two million output tokens in a month, which is a substantial amount of drafting, summarising and code assistance for one person, costs $40 on Sonnet 5 [5]. Even a team spending $120 a month needs about 46 months of perfect savings to cover a $5,499 machine, and that ignores the electricity, the storage, the time you spend maintaining it and the fact that the machine depreciates while the API price falls. The Gemini line in the pricing table is the pointed part: an introductory price that doubles on a stated date [7]. Vendor prices move in both directions, and they have been moving down faster than hardware depreciates.
price ÷ monthly spend. Ignores electricity, storage and depreciation, so treat the result as the best case. Computed in the page; nothing is sent anywhere.
Privacy is a contract question before it is a hardware question
The strongest argument for local inference is that the data never leaves. That argument is real, and it is also frequently used to justify hardware when a policy setting would have done the job.
Check the terms you already have first. OpenAI states that by default it does not use business data to train its models, that API inputs and outputs may be retained for up to 30 days to provide the service and identify abuse and are then removed unless retention is legally required, and that zero data retention is available for eligible endpoints where the use case qualifies [8]. If your requirement is “our client material must not train a model”, that is met by a contract term. If your requirement is “no third party may hold this data for any period”, the 30-day window is a genuine problem and zero data retention is the thing to ask about by name.
The requirement that hardware genuinely solves is the categorical one: material that cannot cross an organisational or national boundary at all, under a rule you do not get to negotiate. Regulated health records, privileged legal material, classified or export-controlled work. If that is you, the calculator above is the wrong tool, because the alternative is not an API bill, it is not doing the work. If that is not you, write down the exact clause you are trying to satisfy before you spend anything, because a surprising number of privacy requirements dissolve on contact with the vendor’s existing data-processing terms.
Utilisation is the number that decides it
Owned hardware is only cheap when it is busy. A machine running a queue eight hours a day earns its price back. A machine that runs one experiment a fortnight is a very expensive way to avoid a subscription, and it is the most common outcome for individual buyers.
So the last question before purchase is not about the model. It is about the schedule. How many hours a week will something be running on this box, and is that work you are already doing and being paid for, or work you imagine doing once you have the machine? The first justifies the spend. The second reliably does not, and the machine ends up as a fast desktop with an unusually large memory configuration.
There is one genuinely good reason to buy ahead of demonstrated need, which is iteration cost. If you are fine-tuning or repeatedly evaluating an open-weight model, the marginal cost of the next run on owned hardware is close to nothing, and that changes how often you experiment. Rented capacity at $84.81 an hour does not [4]. When the constraint on your work is how many times you can afford to try something, owning the machine removes it. That is a real reason, and it is different from saving money.
What still goes wrong
The model you buy the machine for is not the model you will want in a year. Open weights get superseded, and the successor is often larger. A configuration sized precisely to today’s best open model has no headroom, and unified memory is not upgradable after purchase, so a machine bought at the edge of its capacity ages badly. Buying one tier above what you need is the only defence, and it is expensive.
Capability lags too. The published prices above buy access to frontier models from three vendors [5][6][7], and open weights trail those on hard reasoning by enough that people who switch to local inference for cost reasons often quietly keep a paid subscription for the difficult work. That is a sensible outcome, but it means the hardware did not remove the bill, it reduced it, and the break-even calculation you did assumed removal.
Finally, running your own inference is a job. Model files, quantisation choices, runner updates, the occasional afternoon spent working out why generation slowed down. If nobody on your team wants that job, it becomes nobody’s job, and the machine gradually stops being used for the thing it was bought for. That failure has nothing to do with the specifications and it is the most common one.
- 01Apple — Apple introduces new Mac Studio with M5 Max and M5 Ultraapple.com
- 02Apple — Mac Studio technical specificationsapple.com
- 03NVIDIA — H200 Tensor Core GPUnvidia.com
- 04Google Cloud — Accelerator-optimized machine type pricingcloud.google.com
- 05Anthropic — Claude API pricingplatform.claude.com
- 06OpenAI — API pricingdevelopers.openai.com
- 07Google — Gemini API pricingai.google.dev
- 08OpenAI — Enterprise privacyopenai.com