Choosing how hard your AI should think
Set the model tier and the thinking effort per task, so you stop paying flagship rates and flagship waiting time for work that never needed either.
on this page · 0 / 0 checked
At the top of every AI window there is a dropdown with three or four model names in it. Next to it, on most tools now, sits a second control for how hard the model thinks before it answers. Anthropic calls that one effort and gives it five settings [3]. OpenAI calls it reasoning effort and lists seven values [5]. Google calls it thinking level [7]. Most people set both once, near the top, and never touch them again, on the reasonable theory that more thinking cannot hurt.
It can. Not usually in the quality of the answer, though sometimes there too, but in the two things you actually feel: what a month of this costs, and how long you sit watching a cursor blink before you can do the next thing. The two controls are separate, they fail in different ways, and the useful skill is knowing which one to move when an answer comes back wrong. This guide is for solo operators and small teams setting these by hand, in a chat window or in a tool like Zapier or n8n. If you are running an evaluation suite over a production workload, the vendor docs cited here are the primary material and you should read them directly rather than read me summarising them.
Two dials, not one
The first dial is the model tier. Anthropic’s current ladder runs from Claude Haiku 4.5, described as the lowest latency and price, through Claude Sonnet 5 for everyday coding, agent and enterprise workloads, up to Claude Opus 5 for complex agentic coding and enterprise work, with Claude Fable 5.1 as the most capable widely released model [1]. Every vendor ships the same ladder under different names. A higher tier follows complicated instructions more reliably and knows more. It also costs more per word and generally answers slower.
The second dial is how much the model thinks before it commits to an answer. Anthropic’s effort levels are low, medium, high, xhigh and max, with high as the API default [2][3]. The documented behaviour is plain: at low, Claude “minimizes thinking” and “skips thinking for simple tasks where speed matters most”, while at max Claude “always thinks with no constraints on thinking depth” [3]. OpenAI’s reasoning.effort runs none, minimal, low, medium, high, xhigh and max, though not every model supports every value [5]. Google exposes thinking_level, with gemini-3.8-flash defaulting to medium and supporting low, medium and high [7].
These interact, but they are not substitutes for each other. Thinking harder does not give a small model knowledge it does not have, and picking a bigger model does not make it work through a problem step by step if it has been told to answer fast. When an output is wrong, the first useful question is which of the two failed. If the answer was fluent but skipped a step in the logic, that is effort. If it was confidently wrong about the domain, that is tier.
Moving a tier changes a price multiple, moving effort changes a token count
The tier gap is large and it is public. Anthropic’s list prices per million tokens run $1 in and $5 out for Claude Haiku 4.5, $2 and $10 for Sonnet 5, $5 and $25 for Opus 5, and $10 and $50 for Fable 5.1 [4]. OpenAI’s standard short-context rates spread further: GPT-5.6 Luna at $0.10 in and $0.60 out, Terra at $1 and $6, Sol at $2 and $10, GPT-6 Astra at $5 and $25 [6]. That is a factor of about 40 between the cheapest and the most expensive output token inside a single vendor’s own lineup, for work that in many cases is identical.
Effort does not change the rate. It changes how many tokens you buy at that rate. OpenAI states plainly that reasoning tokens are not visible via the API but “still occupy space in the model’s context window and are billed as output tokens” [5]. Google’s pricing note says the same thing in one line: “Response pricing is the sum of output tokens and thinking tokens” [7]. Anthropic’s own guidance for Claude Opus 4.7 warns that stepping up to xhigh means you should “expect meaningfully higher token usage than high” [2]. So the invisible reasoning you never read is billed to you at the visible rate, and turning the dial up on an expensive tier compounds both effects at once.
Two levers cut the bill without touching either dial, and they only exist on the API and automation path rather than in a chat window. Anthropic’s batch processing is 50% off both input and output tokens [4]. Cached input is cheaper again: an Anthropic cache hit costs 10% of the standard input price [4], and OpenAI’s cached input runs about a tenth of standard input, with GPT-6 Astra dropping from $5 to $0.50 per million tokens on short context [6]. If you are sending the same long brief or the same reference document on every run, that is where the money is, not in the effort setting.
The default effort is a guess about somebody else’s workload
Vendors set a default because they have to set something. They are unusually honest that it is a starting point, not a recommendation. Anthropic ships high as the API default and then says in the same document that “the API defaults to high, but the right starting point depends on your model and workload” [2].
You can see the awkwardness inside one vendor’s own advice. For Claude Sonnet 4.6, the recommended default is medium, described as the “best balance of speed, cost, and performance for most applications” and suitable for agentic coding, tool-heavy workflows and code generation [2]. For Claude Opus 4.7, the instruction is to “start with xhigh for coding and agentic use cases”, two notches higher [2]. Same vendor, same page, opposite advice, because the models behave differently. Whatever you inherited by not choosing is unlikely to be the setting either of those recommendations would have given you.
Tier defaults split the same way. Anthropic’s model-choice page offers two routes rather than one answer. The efficiency-first route says that for many applications, starting with a faster, more cost-effective model like Claude Haiku 4.5 can be the optimal approach, and to “upgrade only if necessary for specific capability gaps” [1]. The capability-first route says to implement with the strongest starting point for your task, then “consider increasing efficiency by lowering effort or downgrading models over time with greater workflow optimization” [1]. The page’s own summary line is that most workloads start with Claude Opus 5 [1]. Both routes are offered because the correct one depends on whether your real risk is a wrong answer or a large bill, and only you know which.
Maximum effort is not maximum quality
The most useful sentence in any of these docs is Anthropic advising against its own top setting. On Opus 4.7, the note on max is to “reserve for frontier problems”, because “on most workloads max adds significant cost for relatively small quality gains, and on some structured-output or less intelligence-sensitive tasks it can lead to overthinking” [2]. Overthinking is a real failure mode, not a figure of speech. A model given unlimited room to reason about a task that needed one step will find a second and a third.
The low settings are not degraded modes either. Both OpenAI and Google describe a genuine category of work they belong to. OpenAI’s none is for latency-critical tasks that do not benefit from any reasoning, naming voice, fast information retrieval and classification [5]. Google’s guidance is minimal or low thinking for fact retrieval or classification, the default for comparing concepts or creative reasoning, and maximum thinking for advanced coding, mathematics or multi-step planning [7]. Notice that pulling fields out of a form, sorting messages into buckets and looking something up in a document you pasted all sit in the first group. That is a large share of what small teams actually run.
There is a related trap worth knowing. If an answer gets cut off, the instinct is to raise everything. Anthropic’s guidance gives two remedies for hitting the output limit: raise max_tokens, or “lower the effort level so Claude thinks less and leaves more of the budget for response text” [3]. A truncated answer can be a symptom of too much thinking, not too little. And if you need a hard ceiling on spend, the docs are blunt that effort will not give you one. max_tokens is “a hard cap on total output for the request, thinking and response text combined”, while effort is “soft guidance on how much of that output Claude allocates to thinking” that “doesn’t guarantee a token count” [3].
Three questions settle almost every case
The first question is whether a competent colleague would have to stop and concentrate. Retrieval, classification, reformatting, extracting fields from a document you have pasted in: the vendors route these to their lowest settings themselves [5][7]. Advanced coding, mathematics and multi-step planning are where Google points its maximum thinking level [7], and anything where one wrong step ruins the rest belongs with them. Your intuition about human difficulty is a decent proxy, and it is free.
The second question is what a wrong answer costs. If the downside is that you retype something, stay cheap and read the output. If the downside is a wrong number in front of a client, or a commitment you cannot walk back, buy the care and then check it anyway. The dial is not a substitute for reading the thing.
The third question is how many times this will run. A one-off at maximum effort costs pennies and nobody notices. The same setting on a job that fires on every inbound email is where a 40-times price gap [6] stops being trivia. The reverse also holds, and it is the move most people miss: the design pass deserves the strong setting even when the production runs do not. Work out the prompt and the shape of the workflow on a top tier, then find the cheapest configuration that still passes, which is the capability-first route Anthropic describes, lowering effort or downgrading models as the workflow settles [1]. The fundamentals in how to prompt well will move your results further than a tier upgrade will.
Then make the comparison a habit. Run the same task twice, once at your current setting and once one notch down, and read both outputs yourself before you look at which was which. If you cannot tell them apart, the cheaper one is your new default. Anthropic’s instruction is the same and it is not optional: “The impact of effort levels varies by task type. Evaluate performance on your specific use cases before deploying” [2].
On a subscription, the bill arrives as waiting
Most solo operators never see a per-token invoice, so the token prices above read as somebody else’s problem. They are not, but the currency changes. Claude’s Free plan gives access to Sonnet and Haiku with a 200k context window. Pro is $17 a month with the annual discount, or $20 billed monthly, and adds Opus along with at least 5 times more usage per 5-hour session than Free. Max starts at $100 a month, where you choose 5 times or 20 times more usage than Pro plus priority access at high traffic times, and a Team standard seat is $20 a month billed annually [8].
On a flat plan, over-thinking does not show up on a card statement. It shows up as a limit you hit at 3pm on a Thursday with a deadline in front of you, and as time spent waiting for careful reasoning about a task that needed none. If you have burned through your session limit twice this month, the first thing to try is not a bigger plan. It is dropping the model and the effort on your routine work and keeping the top of both dials for the jobs that earn it. The same logic applies when you are choosing tools in the first place: buy for the work you actually do.
runs × thousand output tokens × price per million ÷ 1,000. Run it at 50 for Claude Fable 5.1 and at 5 for Claude Haiku 4.5 to see the gap [4]. Computed in the page; nothing is sent anywhere.
What still goes wrong
There is no reliable way to know in advance which setting a task needs. Every vendor’s advice reduces to testing on your own cases [2], which means the answer arrives after you have already spent the money at least once. For a task you will run twice, the test costs more than just running it hot and moving on. The three questions above are a shortcut, and shortcuts are wrong sometimes.
Comparisons on a single sample mislead. You run one output at high effort and one at low, cannot tell them apart, drop to low, and then meet the difference three weeks later on the one input that was unusual. The cheap setting fails on the hard cases, which are by definition rare, which is why a two-sample test does not find them. If a task has a tail of awkward inputs, either keep the effort up or build a check that catches the failures rather than trusting a comparison you ran once in March.
Names, prices and levels move under you. Claude Opus 4.6 and Sonnet 4.6 do not support xhigh at all, while Opus 5, Sonnet 5 and Fable 5.1 support all five levels [2], so a configuration copied from a colleague may not even be valid on your model. GPT-6 Astra does not accept none [5]. Prices are provisional too. OpenAI notes that promotional pricing for GPT-5.6 Sol’s standard rates is available at least through November 21, 2026 [6], which is a polite way of saying the number you budgeted against has an expiry date. Anything you build that only works on one specific model at one specific effort level is fragile in a way you will discover at the worst moment.
None of this touches whether the output is true. A cheap model and an expensive one can both be confidently wrong, and the expensive one will be wrong more persuasively. The dial buys you a better shot at hard reasoning. It does not buy you a reason to skip reading the answer.
- 01Anthropic — Choosing a modelplatform.claude.com
- 02Anthropic — Effortplatform.claude.com
- 03Anthropic — Steering thinking and costplatform.claude.com
- 04Anthropic — Claude API pricingplatform.claude.com
- 05OpenAI — Reasoningdevelopers.openai.com
- 06OpenAI — API pricingdevelopers.openai.com
- 07Google — Gemini API thinkingai.google.dev
- 08Anthropic — Claude plans and pricingclaude.com