● Analysis · September 25, 2026 · 6 min

The Same AI Model Can Cost Several Times More Depending on Its Reasoning Level

On September 22, Anthropic and OpenAI launched new models on the same day, both at lower prices: Claude Opus 5.5, and GPT-6 Sol with its smaller sibling GPT-6 Luna. Independent data published by Artificial Analysis brings out a variable that list prices don't show. For Claude Opus 5.5, moving from the default effort (medium) to maximum raises the cost per task from $1.34 to $5.98, 4.5 times more, for 7 extra points on the platform's index. For GPT-6 Sol, the same move costs 4.2 times more. For companies, AI optimization now involves two choices: the model, and the effort level you run it at.

Two launches on the same day, two price cuts

Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, versus $5 and $25 for Opus 5: a 20% cut per token. Tokens are the units of text that model providers bill for. Separately, Anthropic says that at default settings the new model "will cost 40% less than Opus 5 on typical workloads," because on top of the lower price it also uses fewer tokens for the same kind of task. The two figures measure different things. The 20% cut is visible in the price list. The 40% is the vendor's estimate for typical usage, and your company's usage may look different.

On the same day, OpenAI launched GPT-6 Sol, at $2 for input and $10 for output, and GPT-6 Luna, at $0.10 and $0.50, both, according to OpenAI, 50% cheaper than the GPT-5.6 generation's promotional pricing. For an AI agent running hundreds of tasks a day, though, the bill depends on one more setting.

Same model, five effort levels, costs up to 11x apart

The new models can run at several reasoning effort levels, from low to max. The higher the effort, the more "thinking" tokens the model produces before it answers, and those tokens are billed. Artificial Analysis, an independent benchmarking platform, measured both models at every level (index v4.3.2) and calculated the average cost of running one task from the index.

Cost per task and score, at each effort level
30 40 50 60 $0.1 $0.2 $0.5 $1 $2 $5 Cost per task (USD, log scale) Intelligence Index v4.3.2 low · 42 · $0.55 medium · 51 · $1.34 high · 54 · $1.82 xhigh · 56 · $3.46 max · 58 · $5.98 low · 34 · $0.13 medium · 40 · $0.25 high · 43 · $0.37 xhigh · 44 · $0.53 max · 48 · $1.06 Sol non-reasoning · 28 · $0.33
Claude Opus 5.5 GPT-6 Sol

Each point is an effort configuration: level · Intelligence Index score · cost per task (USD). Hollow point: GPT-6 Sol non-reasoning. "Cost per task" = Cost per Intelligence Index Task, the cost of running one index task, not of one solved correctly. Source: Artificial Analysis, Intelligence Index v4.3.2, Claude Opus 5.5 and GPT-6 Sol model pages, values shown on September 25, 2026.

For Opus 5.5, cost per task ranges from $0.55 at low effort to $5.98 at max, nearly 11 times more, while output tokens rise from about 10,000 to 119,000 per task. The score rises from 42 to 58. Going from medium to high adds 3 points for $0.48 more per task. Going from xhigh to max adds 2 points for $2.52 more. The last points of score are the most expensive.

For GPT-6 Sol, costs are lower across the whole scale, from $0.13 at low to $1.06 at max. One detail shows that the link between effort and cost isn't always intuitive: Sol with no reasoning costs $0.33 per task and scores 28, while Sol at low effort costs $0.13 and scores 34. Switching reasoning off entirely does not guarantee the lowest bill.

Cost per task shows what it costs to run a task, not to solve one correctly, and one index point does not translate directly into quality on a company's own tasks. In our September 5 article we showed why cost per task says more than price per token when you compare models. Today's data is on a new version of the index and does not compare with the figures from then, but it adds one observation: cost per task changes even when you don't change the model.

Caching: same model, different bills

The second variable is the context an agent resends on every call: instructions, tool definitions, reference documents. On September 22, OpenAI announced improved caching for GPT-6, with discounts of up to 90% on input tokens already processed, when the same opening part of the prompt is reused within 30 minutes. For Opus 5.5, Anthropic cut the price of cache reads from $0.50 to $0.20 per million tokens, by 60%.

Among the customers OpenAI quotes, Strawberry Browser says it raised its cache hit rate from 83% to 91% in a week and cut its inference costs by 36%. These are customer-reported figures on the vendor's page, without independent verification. The mechanism still holds: two companies using the same model, at the same effort level, can pay different amounts depending on how consistently they structure their prompt and context.

The model you pick isn't always the model that answers

The Opus 5.5 launch page contains one short sentence: "Most cybersecurity tasks will be re-routed to Opus 4.8." Most cybersecurity tasks are automatically redirected to an older model as a safety measure, and biology requests go through the same safeguards as Claude Fable 5.1. For a company running agents on IT or security workflows, this raises a procurement question: which model actually executes the task, and can you see in the logs when a redirect happens? The model name in your configuration stays the same.

What a company should measure before choosing an effort level

The real cost of an AI agent depends on four things at once: the model, the effort level, the context it resends, and how much of that context comes from cache. The recommendations below are MassAI's, based on the data above.

  • Cost per task at each effort level, measured on your own tasks. A test on a real sample says more than any benchmark.
  • Accepted outcome rate. If medium effort produces answers nobody needs to correct, max effort adds no value in that workflow.
  • Output tokens per task. The first sign that an effort level consumes more than it is worth.
  • Cache hit rate and the prompt structure that drives it: stable instructions first, changing data last.
  • Latency and retries. High effort means slower answers, and too little effort can mean more retries.
  • Automatic routing. Whether the provider redirects certain tasks to another model, and whether the redirect shows up in the logs.
  • A short list of tasks that justify max effort. Everything else runs at a lower level by default.

In practice, this means two routing decisions instead of one: which model gets the task, and how much you let it think. A simple, repetitive task can go to a cheap model at low effort, while a high-stakes analysis can justify a strong model at high effort. In the data above, the gap between those two ends exceeds 40 times per task. See also what an operational AI agent built for a company looks like.

Sources: ↗ Anthropic — Introducing Claude Opus 5.5 · ↗ OpenAI — Introducing GPT-6 Sol and Luna · ↗ OpenAI — Better prompt caching for GPT-6 · ↗ Artificial Analysis — Claude Opus 5.5 · ↗ Artificial Analysis — GPT-6 Sol · ↗ Decrypt — the September 22 launches

Want to find out what effort level each step of your workflows deserves?
See how MassAI builds agents →