How to judge an AI performance claim
A procedure for reading benchmark tables, speed numbers and price claims, so you can tell which number is worth a switch and which is noise.
on this page · 0 / 0 checked
A model launches, and with it a table. Bars going up, your current tool one row below, a percentage in bold. Somewhere underneath there is a paragraph of grey text that nobody reads, and that paragraph is where the actual claim lives. The headline says the new model is better. The grey text says on which tasks, in which configuration, against which version, measured how, and by whom.
This guide is a reading procedure for those numbers. You will not need statistics, and you will not need to resolve whether any given model is genuinely ahead. You need to answer one narrow question: does this claim justify changing something I already have working. It is written for a solo operator or a small team paying its own bills. If you run formal evaluations, buy compute at scale, or have a procurement function that demands vendor benchmarks in writing, this is too small for you; you want the methodology papers, not a buyer’s guide.
The setup is part of the claim
A benchmark score is not a property of a model. It is the result of one model, run one way, on one task set, by one party. Change any of those and the number changes, often by more than the gap the chart is drawing your eye to.
Vendors mostly document this, in footnotes. Anthropic’s September 2026 release notes for Claude Fable 5.1 report Humanity’s Last Exam at 60.9% with no tools and 65.0% with tools for the same model [1]. That is a 4-point swing produced entirely by what the model was allowed to use. The same page notes that Fable 5.1 “defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai” [1], which means a benchmark run at high effort was not run in the configuration most people will actually get when they open the app. It also records that the public Terminal-Bench-Science leaderboard uses “3 trials/task, Claude Code harness” [1]. The harness, the scaffolding around the model that gives it tools and retries, is doing part of the work, and a different harness produces a different score for identical weights.
One more setup variable rarely appears elsewhere and matters to you specifically. Anthropic states that Fable 5.1 “was evaluated with its production safeguards enabled”, and that on tasks where those safeguards intervened, Fable 5.1 and Fable 5 “scored a zero on OSWorld 2.0” [1]. That is the honest direction to err in, because safeguards are switched on for you too. A score measured with safety systems disabled is a score of a product you cannot buy.
The error bar is usually bigger than the gap
Most benchmark tables are read as rankings when they should be read as measurements with uncertainty. The numbers move between runs, and the amount they move is frequently larger than the difference between two adjacent rows.
Anthropic puts a figure on it. For Terminal-Bench-Science 0.1, the footnote states that “the standard error is ±3.5–4.5 pts per model” [1]. In the same footnote, the public leaderboard reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%, while Anthropic’s own setup reproduces them at 29.0% and 24.7%, described as “both within noise” [1]. The Fable 5 figure moved 3.3 points between the public leaderboard and the vendor’s own rerun, on the same benchmark and the same weights, with only the setup changed. So any chart where the new model beats the old one by 2 points is telling you nothing you can act on.
Independent evaluators hit the same wall from the cost side. ARC Prize caps its evaluations at “$10,000 USD per run” and states plainly that “a single run is used, we do not average scores across runs” [5]. It handles the resulting variance by requiring agreement between the public and semi-private sets within a stated tolerance, ±10 percentage points for ARC-AGI-1 and ±3 for ARC-AGI-2 [5]. Read that as the professionals’ own admission: repeating these tests properly is expensive, so a lot of published numbers are single samples. Treat a single-digit lead as a tie until someone shows you repeats.
Benchmarks have version numbers, and last month’s table is void
The tests themselves change. Tasks get added, ambiguous ones get removed, graders get rewritten, and the score moves without any model changing at all.
The names give it away when you look. Terminal-Bench 4.0. Terminal-Bench-Science 0.1. OSWorld 2.0. CursorBench 3.2.0 [1]. Those are software versions, and they behave like software versions. Anthropic’s own note on its OSWorld results is the clearest statement of the problem you will find in a vendor document: the scores are “on the benchmark authors’ August 2026 task release”, and “because the task files differ from earlier releases, these numbers aren’t directly comparable to previously published OSWorld 2.0 results, which is why no competitor score is shown” [1]. A vendor declining to show a competitor score because the comparison would be invalid is the behaviour you want. It is also a warning about every third-party chart that does show one.
The practical rule is short. If a comparison table does not state the benchmark version and the date each model was run, you cannot tell whether you are looking at a capability difference or a calendar difference. Older competitor entries were often run on an older task set, at an older effort default, through an older harness.
Who ran the test decides what the number is worth
There is a hierarchy, and it is worth holding in your head because it settles most arguments quickly. At the top, an independent evaluator running a published method on a held-out task set, under rules the vendor does not control. In the middle, a vendor reporting its own numbers with its own setup documented. At the bottom, a number with no method attached.
Most launch tables sit in the middle, and the middle is not uniform within a single page. The same Anthropic document that spells out trial count, harness and standard error for Terminal-Bench-Science 0.1 also carries a GPT-5.6 Sol column on that benchmark and on Terminal-Bench 4.0, with no footnote saying how those competitor numbers were produced [1]. On OSWorld 2.0 the same page shows no competitor at all [1]. Read the footnotes row by row rather than page by page. A vendor can document its own runs carefully and still leave the comparison line unexplained, and the comparison line is the one the chart is built around.
Independent evaluation has teeth when the rules are strict. ARC Prize runs its verified leaderboard on “submissions from trusted partners such as public, high-usage, commercially available model APIs (e.g., OpenAI, xAI, Google, etc.)” and requires “zero data retention agreements with all model providers we test” [5]. Its community leaderboard is a different object: “We do not verify submissions on the community leaderboard by default” [5], and anyone posting a result there is asked to “state clearly the data you tested on, how you tested, and that your results are not verified by ARC Prize” [5]. One organisation, two leaderboards, one verified and one not. Check which one the screenshot came from.
Crowd rankings deserve their own caution, because they look independent and are shaped by access. The Leaderboard Illusion, a 2025 study of Chatbot Arena, found that undisclosed private testing practices “benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired”, and identified 27 private LLM variants tested by Meta in the lead-up to the Llama-4 release [6]. The same study estimated that Google and OpenAI received 19.2% and 20.4% of all arena data respectively, while a combined 83 open-weight models received an estimated 29.7%, and that even limited additional data “can result in relative performance gains of up to 112% on the arena distribution” [6]. None of that makes the leaderboard useless. It makes it a measurement of a system that includes the vendors’ testing strategy, not just their models.
Speed and price claims have separate fine print
Cost claims fail differently from benchmark claims. They are usually true and still wrong about your bill.
Start with what a comparison site’s single price number contains. Artificial Analysis calculates a blended price “assuming a 7:2:1 ratio of cache hit, input, and output tokens” [7] and standardises on OpenAI’s o200k_base tokeniser so models can be compared in the same units [7]. For reasoning models it assumes “2k reasoning tokens” where actual measurements are unavailable [7]. Every one of those is a defensible modelling choice and none of them is your workload. If your work is long documents in and short answers out, the blend is wrong for you in one direction; if you generate long drafts, it is wrong in the other.
Then check the published rate card, because almost every frontier price now has a context step in it. OpenAI lists GPT-6 Astra at $10.00 per million input tokens and $50.00 output at short context, rising to $20.00 and $75.00 at long context [3]. Google’s Gemini 3.1 Pro Preview is $2.00 input and $12.00 output for prompts up to 200k tokens, and $4.00 and $18.00 above that [4]. In both cases the input rate doubles when you cross the step. A price you gathered while testing on 50k-token prompts is not the price you pay once a long thread or a large attachment pushes you over, so find where each candidate’s boundary sits before you compare anything.
Promotional pricing is the other trap. Gemini 3.8 Flash input is listed at “$0.75 through December 31, 2026. $1.50 starting January 1, 2027” [4], and OpenAI notes that GPT-5.6 Sol’s promotional pricing “is available at least through November 21, 2026” [3]. A migration justified by a promotional rate has a deadline attached to its business case, and it is worth writing that date into the decision rather than into a footnote.
The last gap is per-token against per-task. Anthropic prices Claude Opus 5 at $5 per million input tokens and $25 output, and Claude Sonnet 5 at $2 and $10 [2]. Sonnet is nominally 60% cheaper on output. If the cheaper model needs two attempts, or a longer reasoning budget, or your review time to fix what it got wrong, the per-task cost can land higher. Measure the task, not the token.
The number that moves your stack is the one you measured
Everything above is triage. It tells you which claims are worth the cost of testing. The test is the same regardless of how the claim looked.
Pull 10 to 20 real tasks out of your last month, the awkward ones included. Run them through what you use now and the candidate, in the configuration you would actually deploy, which means the effort level your plan defaults to and the tools your account has switched on [1]. Do them in the same sitting, because your standards drift across weeks. Score them the way the recipient will, which is usually a binary of shipped or reworked rather than a percentage. Count the failures that would have cost you something, and note how long each output took you to check. Then set your bar before you look at the totals, and set it well clear of a tie: a candidate that wins 11 of 20 has told you nothing, because a sample that small cannot separate a real edge from luck. Switch on a margin you would still believe after re-running the same set next month.
A claim you repeat becomes your claim
There is a second reason to be strict about this, and it applies the moment you market anything AI-assisted to your own customers.
Performance claims about AI are ordinary advertising claims, and the FTC has enforced them as such. Its Operation AI Comply sweep, announced 25 September 2024, brought five law enforcement actions [8]. One was against DoNotPay, which marketed itself as “the world’s first robot lawyer” and claimed it could “replace the $200-billion-dollar legal industry with artificial intelligence” without substantiating that claim; the settlement required the company to pay $193,000 and prohibited it from making unsubstantiated claims about its service [8]. Then-chair Lina Khan’s line is the durable one: “there is no AI exemption from the laws on the books” [8].
The relevance to a small business is direct. If you repeat a vendor’s unverified number in your own pitch deck, on your pricing page, or in a proposal, you have adopted it. The vendor’s footnote about standard error and effort levels does not travel with the screenshot. Quote your own measured results instead, describe how you got them, and when you must cite a vendor figure, cite it as the vendor’s figure with the date attached.
Price difference × your monthly volume. Defaults are Claude Opus 5 and Claude Sonnet 5 output pricing [2]. Computed in the page; nothing is sent anywhere.
What still goes wrong
The bake-off you run is itself a small sample with a wide error bar, and 20 tasks cannot tell a 3-point difference from luck any better than a vendor’s single run can. What it can tell you is whether the new tool fails on the things you care about, which is a different and more useful question than which model is ahead. Accept that you are testing for disqualifying failures, not for ranking.
Your task set also goes stale. The 20 tasks that represented your work in March may not represent it in September, and a model that lost on the old set can win on the new one. There is no clean fix beyond re-running occasionally and keeping the tasks in a file rather than in your memory. Costs drift under you as well, since published prices change and promotional rates expire on stated dates [3][4], so a decision made on price is worth re-checking when the date arrives.
The largest limit is that none of this covers the claims that matter most and are hardest to test. Nothing in a benchmark table tells you what happens to your data, whether the vendor will still exist in two years, or how the model behaves on the rare, high-consequence task you cannot simulate. Benchmarks measure capability under test conditions. They do not measure reliability under yours, and for anything with legal, financial or medical consequence, a passing score on a public benchmark is not evidence that the work can go out without a human reading it.
- 01Anthropic — Introducing Claude Fable 5.1 and Claude Mythos 5.1anthropic.com
- 02Anthropic — Models overviewplatform.claude.com
- 03OpenAI — API pricingdevelopers.openai.com
- 04Google — Gemini API pricingai.google.dev
- 05ARC Prize — Testing policyarcprize.org
- 06Singh et al. — The Leaderboard Illusion (arXiv 2504.20879)arxiv.org
- 07Artificial Analysis — Methodologyartificialanalysis.ai
- 08FTC — Operation AI Comply press releaseftc.gov