← Blog
October 3, 2026 · 16 min

Paying for thoughts you never see: reasoning tokens and the reserve

Reasoning models bill for text you never read. OpenAI, Anthropic and Google all charge hidden reasoning as output and count it against the same cap as the answer. We read their current docs plus OpenRouter, LiteLLM, vLLM and Ollama, ran a thinking model locally, and probed our own x402 quote. A six-character answer cost 917 output tokens. With the cap sized for the answer, the call cost 256 tokens and returned nothing.

Our billing internals post described the reserve every prepaid gateway runs. Hold input tokens times input price plus max_tokens times output price, forward the call, settle the truth. The formula quietly assumed that output tokens were the text the agent receives. On reasoning models that assumption is gone. Most of the output can be text nobody receives.

That breaks three things a paying agent relies on. The cap no longer bounds the answer, because reasoning spends it first. The knob that controls reasoning, reasoning_effort, means something different at every hop between the agent and the model. And the receipt that would show where the tokens went is missing on most compatibility layers, including ours.

Our sources, all read or run on 3 October 2026: OpenAI's reasoning guide, token counting guide and model pages, plus its OpenAPI at 92f957a; Anthropic's thinking, extended thinking, steering and cost, effort, pricing and OpenAI SDK compatibility pages; Google's Gemini thinking and OpenAI compatibility pages; OpenRouter's reasoning tokens docs; LiteLLM's Anthropic transformation at d260765; vLLM's sampling parameters and protocol at 5f30fc7; Ollama's OpenAI compatibility docs and a local Ollama 0.30.6; the x402 upto scheme; and unauthenticated probes against our own endpoint.

The meter you cannot read

OpenAI states it plainly. Reasoning tokens "are not visible via the API" but "still occupy space in the model's context window and are billed as output tokens." Its output caps, max_output_tokens on Responses and max_completion_tokens on Chat Completions, "limit all tokens generated by the model, including non-visible tokens." The OpenAPI description of max_completion_tokens says the same: visible output tokens and reasoning tokens. The older max_tokens field is marked deprecated and "not compatible with o-series models."

When a call hits the cap, OpenAI returns status: "incomplete". The guide warns this "might occur before any visible output tokens are produced, meaning you could incur costs for input and reasoning tokens without receiving a visible response." Its advice is to reserve at least 25,000 tokens for reasoning and outputs when you start experimenting.

One more detail matters for accounting. Some models emit formatting tokens that show up in neither the message content nor the reasoning count. So completion_tokens can exceed the visible output "even when the reported reasoning_tokens value is 0."

Anthropic's rule is the same with a sharper edge. Thinking is billed as output, and "the billed output token count does not match the visible token count." On the Claude 5 family the visible part is often nothing at all. There, display defaults to "omitted", so thinking blocks arrive with an empty thinking field while the full reasoning is billed.

max_tokens is Anthropic's hard cap on thinking plus text. But in a tool loop, "each request in the turn has its own max_tokens, so it doesn't bound the whole turn's spend."

Google's Gemini docs close the set. max_output_tokens includes thought tokens. If the model hits it while reasoning, it stops with status "incomplete" and "returns truncated or empty output (while still billing for any thinking tokens generated)." Pricing "is based on the full thought tokens the model needs to generate, despite only the summary being output from the API."

Three providers, one rule. The cap is shared, the bill is full, the view is partial.

Thinking you cannot switch off

The next change is quieter. On the newest models, reasoning is no longer opt-in.

Anthropic's per-model table is explicit. On Claude Opus 5.5, Claude Fable 5.1 and Claude Fable 5, a request with no thinking field gets adaptive thinking, and thinking: {type: "disabled"} returns a 400. Claude Sonnet 5.5 also rejects disabled. Its lowest setting is "between_tools", accepted at high effort or below. Claude Opus 5 can turn thinking off, but only at high effort or below. Older models such as Claude Opus 4.8 still default to thinking off.

Effort defaults differ too. Anthropic's default is medium on Claude Opus 5.5 and high on the other models that support effort. OpenAI's gpt-5.5 defaults to medium.

OpenAI's newest models narrow the range from the other side. GPT-6 Astra accepts low through max, and sending none "returns HTTP 400." GPT-6.1 Sol supports neither none nor minimal and defaults to medium. Google's compatibility docs say reasoning "cannot be turned off for Gemini 2.5 Pro or 3 models."

So the reasoning_effort: "none" sitting in an agent's config is now a request that a growing share of models cannot honor. What happens to it depends on who sits in the middle.

One string, seven meanings

We traced a single value, reasoning_effort: "high", through every layer we could read. Each one turns it into something different.

# reasoning_effort "high", by layer (docs and source read 3 Oct 2026)
OpenAI Chat Completions          enum value; the model decides the token count
Gemini OpenAI compatibility      thinking_level "high" (Gemini 3 models in its table)
                                 thinking_budget 24,576 (Gemini 2.5)
Anthropic OpenAI compatibility   ignored
LiteLLM → Claude, manual budget  thinking.budget_tokens 4,096
LiteLLM → Claude, adaptive       thinking adaptive + output_config.effort "high"
OpenRouter → Anthropic budget    budget_tokens = 0.8 × max_tokens, clamped 1,024–128,000
Ollama 0.30.6, qwen3             thinking on, identical to sending nothing

The same word sets a 4,096-token thinking budget on one path, 24,576 on another, 80 percent of whatever cap you set on a third, and nothing at all on a fourth.

The low end diverges further. Google maps minimal to low on Gemini 3.1 Pro, keeps it as minimal on Gemini 3 Flash, and turns it into a 1,024-token budget on Gemini 2.5. none disables thinking only on 2.5 models. OpenRouter sends minimal to Claude as low and rejects none. Ollama's current docs alias minimal to low for models without level metadata. For models with it, unsupported names "resolve to the model default." The 0.30.6 build we ran rejects minimal with a 400.

LiteLLM handles none by removing both thinking and output_config from the Anthropic request. Read against Anthropic's table, that request reaches Claude Fable 5.1 with no thinking field. That means adaptive thinking at the model's default effort, high. By our reading, asking LiteLLM for none on that model buys more reasoning than asking for low. We did not run that call. The conclusion follows from the code and the table.

Only one layer we read offers a hard cap on reasoning alone. vLLM's thinking_token_budget sampling parameter is the "maximum number of tokens allowed for thinking operations." When it runs out, the sampler forces the model's end-of-reasoning sequence by setting that token's logit to 1e9. The hosted APIs offer softer controls. Anthropic calls effort "a behavioral signal, not a strict token budget," and calls its legacy budget_tokens "a target rather than a strict cap." Only total output is hard-capped.

The receipts diverge as much as the knobs. OpenAI's schema carries completion_tokens_details.reasoning_tokens. Anthropic's native API reports usage.output_tokens_details.thinking_tokens, which on streams appears only in the final message_delta event. Gemini's Interactions API reports total_thought_tokens. vLLM's protocol defines a reasoning_tokens field too.

But Anthropic's own OpenAI-compatible endpoint lists usage.completion_tokens_details as "Always empty," marks reasoning_effort as "Ignored," and notes that "the OpenAI SDK doesn't return Claude's thinking." On the Claude 5 family, where thinking is on by default, an agent using that surface pays for reasoning it cannot see, cannot tune with the field it knows, and cannot itemize.

We ran it: a six-character answer

To see the mechanics end to end, we ran a small thinking model through Ollama's OpenAI-compatible endpoint. The setup: Ollama 0.30.6 (the latest release is 0.35.1, from 29 September), qwen3:1.7b at Q4_K_M, CPU only, a 4,096-token context window, temperature 0 and a fixed seed. One prompt: a merchant charges 0.0025 USDC per call, an agent makes 37 calls and then 12 more, what is the total? Answer with just the number.

# Ollama 0.30.6, qwen3:1.7b, POST /v1/chat/completions, 3 Oct 2026
# request                          finish    completion_tokens  visible content
default (thinking on)              stop         917             "0.1225"
reasoning_effort "low"             stop         917             identical
reasoning_effort "high"            stop         917             identical
reasoning_effort "none"            stop          38             "0.0025 × (37 + 12) = 0.0025 × 49 = 0.1225 USDC"
max_tokens 256                     length       256             "" (empty)
max_tokens 256, effort "low"       length       256             ""
stream + include_usage, cap 300    length       300             0 content chunks, 298 reasoning chunks
reasoning_effort "minimal"         HTTP 400     invalid reasoning value
reasoning_effort "xhigh"           HTTP 400     invalid reasoning value

The answer was right every time. The bill was not stable. With thinking on, six visible characters cost 917 completion tokens. With thinking off, a full worked sentence cost 38. That is a 24x spread for the same question and the same correct number. A first run, before we pinned the context window and thread count, billed 1,457 tokens for the same six characters.

low and high produced output identical to sending nothing. medium and max also left thinking on. On this model the field is a switch with one working position, none.

The capped calls are the failure that matters for billing. With max_tokens at 256, the model spent all 256 tokens reasoning and returned an empty content with finish_reason: "length". Every token was billable. None was usable.

The stream showed the same thing chunk by chunk: 298 reasoning chunks, zero content chunks, and a final usage block reporting 300 completion tokens. Two tokens never surfaced in either field.

No response carried completion_tokens_details. The only evidence of how the 917 tokens split was the reasoning text itself, which a client would have to re-tokenize to audit. Turning thinking off also changed the input side: prompt_tokens went from 53 to 59.

This is a 2-billion-parameter model on a CPU, not a frontier API. Magnitudes on hosted models differ. The mechanics do not. The providers document the same behaviors we measured: a shared cap, an empty answer at the cap that is still billed, and no breakdown on at least one major compatibility layer.

What reasoning does to the reserve

Put those facts into the reserve formula. The cap is now the only hard bound on a call's cost, and it has to cover thinking the agent never asked to see.

# Output side of the reserve = cap × list output price (input excluded)
# model             $/MTok out   cap 1,000   cap 25,000   cap 128,000
GPT-6 Astra           50          $0.05       $1.25        $6.40
Claude Fable 5.1      50          $0.05       $1.25        $6.40
Claude Opus 5.5       20          $0.02       $0.50        $2.56
GPT-5                 10          $0.01       $0.25        $1.28

OpenAI's starting recommendation of 25,000 tokens puts the output side of the reserve at $1.25 per call on GPT-6 Astra or Claude Fable 5.1. At either model's 128,000-token output limit it is $6.40. Above 272,000 input tokens, OpenAI bills the full GPT-6 Astra request at 1.5x the output rate and 2x the input rate, so long-context calls hold more.

The agent's instinct is to size the cap for the answer it wants. That is the setting that fails. A 1,000-token cap on a model that thinks by default reproduces our 256-token result at larger scale: finish on length, empty content, full charge. A naive retry with the same cap pays again for the same empty result.

Four more effects compound across an agent loop.

Tool loops. Each request in a tool-use turn carries its own cap, so no single max_tokens bounds what a turn costs. The reserve bounds calls. Only a spend policy bounds a task.

Multi-turn history. Anthropic keeps prior turns' thinking blocks in context on Claude Opus 4.5 and later Opus models, Sonnet 4.6 and later Sonnet models, and the Fable models. Those blocks "are billed as input tokens like the rest of the conversation history." Reasoning gets paid for twice: once as output when generated, then as input on every later turn that carries it. OpenAI's GPT-5.6 family now renders earlier turns' reasoning into context by default as well.

Caching. Anthropic renders the resolved effort into the prompt, so changing effort between requests invalidates cache breakpoints. A gateway that lowers effort on some calls to save money can lose more on cache misses than it saves on thinking. Anthropic, in beta, and OpenAI, on the GPT-6 family, now offer per-message effort changes that keep the cached prefix intact. We priced the cache side of this in the prompt caching post.

Fallback. A thinking block is readable only by the model that produced it and a fixed set of others. Anthropic drops blocks the target model cannot read "without an error and without billing it." Switching from Claude Opus 5.5 up to Claude Fable 5.1 keeps earlier reasoning on the Claude API. Switching down drops it. OpenAI omits reasoning across model families. A fallback chain that crosses those lines loses reasoning continuity mid-conversation. It can also land on a model whose thinking default is the opposite of the first link's.

upto prices the ceiling, not the thought

x402's upto scheme was written for exactly this variance. The client authorizes a maximum. The server settles an actual amount "determined at settlement time based on resource consumption." The settled amount must not exceed the maximum and "MAY be 0." The spec's first example use case is "Paying for LLM token generation." Each authorization settles at most once, and multi-settlement streaming is explicitly out of scope.

That covers the gap between our 917-token and 38-token calls. It does not tell the server what "actual" means when reasoning ate the cap and the answer came back empty. The upstream bills the gateway for those tokens either way. Whether the agent pays for them, and whether it can prove what it paid for, is a policy the gateway has to state and a receipt it has to emit.

We probed our own walk-up surface to see what it quotes today. An unauthenticated POST /v1/chat/completions returns HTTP 402 with an exact requirement in USDC on Base. We varied only the model and the token fields.

# api.llm4agents.com, unauthenticated POST /v1/chat/completions, 3 Oct 2026
# body (one message: "hi")                         402 amount (USDC, 6 decimals)
openai/gpt-5, no cap                                50000      $0.05
openai/gpt-5, no cap, reasoning_effort "high"       50000      $0.05
openai/gpt-5, max_completion_tokens 100000          50000      $0.05
openai/gpt-5, max_tokens 256                        10000      $0.01
openai/gpt-5, max_tokens 1000                       20000      $0.02
openai/gpt-5, max_tokens 100000                     1210000    $1.21
openai/gpt-5, max_tokens 100000, effort "low"       1210000    $1.21
openai/gpt-5, max_tokens 1000000                    12010000   $12.01
anthropic/claude-opus-5.5, no cap                   100000     $0.10
anthropic/claude-fable-5.1, no cap                  250000     $0.25
openai/gpt-6-astra, no cap                          250000     $0.25
nonexistent/model-x                                 10000      $0.01

Four things stand out. The quote follows max_tokens and ignores max_completion_tokens, the field OpenAI's spec says covers reasoning and the one its deprecation notice points to. It ignores reasoning_effort, which is defensible, since effort is soft and the cap is what binds. It is not clamped to the model's output limit: a 1,000,000-token cap on GPT-5 produced a $12.01 quote, while GPT-5's 128,000-token output limit is worth $1.28 at list price. And an unknown model slug still received a quote.

The no-cap quotes are the ones that matter for reasoning. For GPT-5, Claude Opus 5.5, Claude Fable 5.1 and GPT-6 Astra, the default quote covers at most 5,000 output tokens at each provider's list price. That is a fifth of OpenAI's starting recommendation for reasoning models, on models that think by default. The quote and the upstream bound should be the same number. Our public contract does not say how the two are reconciled when an agent omits the cap.

What it means for LLM4Agents

Our public OpenAPI describes a Chat Completions request with model or a models fallback array, messages, temperature, max_tokens and stream. It sets additionalProperties: true, so other fields pass validation, but it documents none of the reasoning controls. The response usage object lists prompt_tokens, completion_tokens and total_tokens. The billing headers are X-Tokens-Input, X-Tokens-Output, X-Cost-Usd-Cents and X-Model-Used. Nothing tells an agent how much of X-Tokens-Output was thinking.

The threat is a migration nobody announces. As agents move to the Claude 5 family and GPT-6, thinking arrives by default. Caps sized for visible text start returning empty answers that are still billed. Fallback chains that mix thinking-mandatory and thinking-optional models change the cost per hop by an order of magnitude, as the 24x spread showed even on a tiny model.

The surface we speak also has a new hard limit. OpenAI's reasoning guide says Chat Completions "does not support function calling with GPT-6 Astra or GPT-6.1 Sol." For tool-using agents on those models, the Responses surface we audited yesterday is a requirement, not an option.

The opportunity is that the gateway is the only party that sees both sides. It can call each provider's native API, where the reasoning count exists, instead of a compatibility endpoint, where it may not. It can know each model's thinking default, effort vocabulary and output limit. An agent paying per call cannot assemble that table on its own. A gateway that publishes it, enforces it and puts it on the receipt sells something the providers' own compatibility endpoints do not.

Staying on the frontier

First, fix the quote inputs. Honor max_completion_tokens as the cap whenever it is present. Clamp the cap to the model's output limit. Reject unknown model slugs before issuing a 402. Forward the cap that priced the quote to the upstream explicitly, so the quote and the bound are one number.

Second, put reasoning on the receipt. Return completion_tokens_details.reasoning_tokens on every response and in the final stream chunk, mapped from each provider's native field. Add an X-Tokens-Reasoning header next to X-Tokens-Output. Where an upstream does not report the split, prefer that provider's native API. If that is impossible, label the figure as estimated rather than leaving it out.

Third, publish the effort table and fail honestly. One row per model: thinking default, accepted effort values, what each value maps to upstream, and whether thinking can be turned off. When an agent asks for something a model cannot do, such as none on Claude Fable 5.1 or GPT-6 Astra, return a 400 that says so. Never coerce silently.

Fourth, add a reasoning floor. For models that think by default, warn or reject when the cap is below a published floor, starting from OpenAI's 25,000-token guidance. Detect the empty-at-cap case, finish_reason: "length" with empty content and reasoning present, and flag it in a header so clients do not retry with the same cap.

Fifth, move LLM walk-up to upto. Quote the ceiling, settle the measured tokens, and include the reasoning count in the settlement receipt. That is reserve-proxy-settle on the x402 rail.

Sixth, give fallback a reasoning rule. Quote chains at the worst link's cap and price. Keep effort fixed for a conversation, using per-message effort changes where an upstream supports them, so the cache survives. Pin conversations that carry thinking blocks to models that can read them, or drop the blocks explicitly and say so.

Seventh, expose hard caps where they exist, then measure. For self-hosted upstreams on vLLM, map a documented field to thinking_token_budget, the only hard reasoning cap we found. Run the probe set from this post against every routed model on a schedule, covering the empty-at-cap rate, the reasoning share of output and the effort values that return 400, and publish the results.

Pay for thinking you can audit

One OpenAI-compatible gateway, paid per call in USDC over x402.

Register your agent