Priced in characters, billed in tokens: auditing our 402 quote
An x402 walk-up price has to exist before the model runs. Our gateway computes it from one line: the JSON length of the messages, divided by four. The provider then bills in its own tokenizer's units. We ran fifteen languages and a set of agent payloads through six public tokenizers to measure the gap. English is over-counted by about a third. Chinese, Japanese, transaction hashes and base64 signatures are under-quoted two to four times. An image is over-counted almost eighteen times.
Every pay-per-call LLM endpoint has the same problem. The agent signs the price before the work happens, but the unit of work, the token, is defined by a tokenizer the gateway may not have. Output has a ceiling: max_tokens. Input has nothing but an estimate.
This is the fourth audit in our series on what an OpenAI-compatible gateway really translates. The reasoning tokens audit followed hidden output. The structured outputs audit and the tool-calling audit found that our quote ignores schemas and tool definitions. This one asks what the quote does with the part it does count.
Our sources, all read or run on 6 October 2026: OpenAI's token counting guide and GPT-5 model page; Anthropic's token counting, pricing and vision pages; Google's Gemini tokens guide; OpenRouter's usage accounting page; LiteLLM at 18c3118; and x402 at cb0ec5b. The tokenizers: tiktoken 0.14.0 for o200k_base and cl100k_base; the published tokenizer.json files of Qwen3-8B, DeepSeek-V3.1 and GLM-4.6; and Mistral's Tekken, from the tekken_240911.json file shipped in mistral-common 1.12.0. The corpus is the Universal Declaration of Human Rights from NLTK's udhr2 collection, a parallel text. We also probed our own 402.
What our 402 counts
An unauthenticated POST to /v1/chat/completions returns HTTP 402 with an x402 v2 exact requirement in USDC on Base. We set max_tokens to 1 so the output reserve was negligible. Then we sent one character repeated N times and searched for the smallest N that moved the quote from $0.01 to $0.02.
# api.llm4agents.com, unauthenticated POST /v1/chat/completions, 6 Oct 2026
# anthropic/claude-opus-5.5, max_tokens 1, one user message of N copies of one character
character UTF-8 bytes quote reaches $0.02 at N =
a 1 8,283
é 2 8,283
क (Devanagari ka) 3 8,283
的 (Chinese) 3 8,283
😀 (emoji) 4 4,142
" (double quote) 1 4,142
\ (backslash) 1 4,142
tab 1 4,142
newline 1 4,142
Bytes do not matter. A Chinese character costs exactly what a Latin a costs. What matters is the length of the JSON-serialized string in UTF-16 code units. Characters that JSON escapes (quotes, backslashes, tabs, newlines) count twice. So does an emoji, which takes two UTF-16 units.
The thresholds fit one formula exactly. The quote serializes messages, divides the length by four, rounds up, and prices that at the model's input rate. It adds max_tokens at the output rate, applies a 20% platform fee and rounds up to the cent. The 30-character envelope [{"role":"user","content":""}] explains why the step lands at 8,283 rather than 8,313. We then read our gateway's source and found the same line: Math.ceil(JSON.stringify(messages).length / 4). When max_tokens is absent, the code reserves 4,096 output tokens.
# the walk-up quote, reconstructed from probes and confirmed in our source
est_input = ceil( JSON.stringify(messages).length / 4 )
quote_usd = 1.20 * ( est_input * input_price + (max_tokens ?? 4096) * output_price )
quote = max( $0.01, ceil_to_cent(quote_usd) )
Four characters per token is not our invention. Google's Gemini guide says "a token is equivalent to about 4 characters." The trouble is the word "about," and the languages and payloads it was never meant to cover.
How the providers count
OpenAI offers an input token count endpoint, POST /v1/responses/input_tokens. It accepts the same input as the Responses API and "returns the exact count the model will receive." That count includes "formatting tokens used to represent request structure, such as message roles and boundaries," which a local tokenizer will not see. The guide says tiktoken works for plain text, but for images and files "estimates like characters / 4 are inaccurate." tiktoken 0.14.0 maps gpt-5 to o200k_base.
Anthropic's count_tokens endpoint is free. It has its own rate limits, separate from message creation: 5,000, 10,000 and 20,000 requests per minute on the Start, Build and Scale tiers. The result is "an estimate," and it may include system-added tokens that are not billed. The important line is about generations. Claude 4.7 and later models "use a newer tokenizer" that produces "approximately 30 percent more tokens" for the same text. Fable 5.1, Mythos 5.1, Fable 5 and Mythos 5 share it. The docs tell you not to "reuse token counts measured on the older model to estimate costs." They offer the endpoint, not a local tokenizer.
Google documents countTokens and the four-character rule, plus fixed rates for media: 258 tokens for an image of up to 384 pixels per side, 258 per 768x768 tile above that, 263 tokens per second of video and 32 per second of audio.
The layers in between pick a side. OpenRouter reports usage "using the model's native tokenizer," and now includes it in every response; stream_options.include_usage is deprecated and has no effect. LiteLLM's token_counter picks a Hugging Face tokenizer for four families and otherwise uses tiktoken by model name, falling back to cl100k_base for names tiktoken does not know (utils.py, token_counter.py). For every Anthropic model whose name does not contain claude-3, it loads a packaged anthropic_tokenizer.json with a 65,000-entry vocabulary. Its model map lists claude-opus-5-5 and claude-fable-5-1 as Anthropic models, so they get that file. Anthropic says the tokenizer behind those models is new.
So a gateway sits between three facts. The price must be fixed before the call. The exact count lives behind a provider endpoint, or in a public tokenizer file for OpenAI and open-weight families. The bill uses the native count.
One declaration, fifteen languages
The Universal Declaration of Human Rights is the classic parallel text: the same content in hundreds of languages. We normalized each version to NFC, counted it with six tokenizers, and divided by what our quote assumes for the same text. Above 1.00, the quote under-counts.
# real tokens / quote's estimate, UDHR text only, NFC, 6 Oct 2026
# quote = ceil(JSON-escaped UTF-16 length / 4); DSV3.1 = DeepSeek-V3.1
language chars quote o200k cl100k Qwen3 DSV3.1 GLM4.6 Tekken
English 10,637 2,683 0.75 0.75 0.76 0.75 0.75 0.77
Spanish 11,964 3,015 0.82 0.99 0.99 0.94 0.92 0.88
Portuguese (BR) 11,280 2,844 0.84 1.06 1.04 0.96 0.98 0.90
French 11,901 2,999 0.88 1.04 1.04 1.00 0.98 0.89
German 11,936 3,008 0.85 1.10 1.09 1.00 0.97 0.87
Turkish 10,278 2,593 1.15 1.54 1.33 1.59 1.32 1.25
Swahili 10,452 2,637 1.14 1.55 1.56 1.53 1.54 1.49
Vietnamese 11,059 2,789 1.11 1.96 1.09 1.69 1.10 1.09
Russian 11,805 2,975 0.95 1.73 1.18 1.03 0.98 1.04
Arabic 7,645 1,935 1.24 2.74 1.45 1.43 1.56 1.17
Hindi 11,500 2,899 1.16 3.89 3.66 2.12 3.89 1.36
Bengali 9,711 2,452 1.37 4.85 4.20 1.67 4.85 1.63
Chinese (Simpl.) 2,988 771 3.07 4.48 2.36 2.16 2.21 3.44
Japanese 4,182 1,069 3.33 4.51 2.73 2.53 2.80 3.05
Korean 4,715 1,202 2.28 3.88 2.44 2.53 3.02 2.04
Five readings.
English is over-counted on every tokenizer: the estimate is 30 to 33% above the real count. All six land between 5.2 and 5.3 characters per token on this text, not four. An English-speaking agent signs for input it did not send, before the fee.
Spanish, our second language, is near break-even at 0.82 to 0.99. The same content costs 23 to 48% more tokens in Spanish than in English. The estimator sees only 12% more characters.
Under-quoting starts with Turkish and Swahili (1.14 to 1.59) and grows with distance from the Latin script. Arabic runs 1.17 to 2.74, Hindi 1.16 to 3.89, Bengali 1.37 to 4.85.
CJK is under-quoted on every tokenizer. Chinese runs 2.16 to 4.48, Japanese 2.53 to 4.51, Korean 2.04 to 3.88. The estimator thinks the Chinese declaration costs 0.29 times the English one. The tokenizers say 0.83 to 1.71 times.
The tokenizers disagree with each other more than with the estimator. Hindi is 3,376 tokens on o200k_base and 11,284 on cl100k_base: a 3.3x gap between two OpenAI encodings. A gateway that counts "with tiktoken" without choosing the encoding per model inherits that gap.
Same text, different code points
The Vietnamese file in the corpus arrived decomposed. Its accents are stored as combining marks after each base letter. Composing it to NFC shrinks it from 13,012 to 11,059 code points. On screen, the two versions look identical.
# Vietnamese UDHR, as shipped (decomposed) vs NFC, 6 Oct 2026
counter decomposed NFC
our quote 3,277 2,789
o200k 6,950 3,093
cl100k 8,659 5,468
Qwen3 3,032 3,032
DeepSeek-V3.1 7,905 4,706
GLM-4.6 7,844 3,074
Tekken 8,966 3,033
Qwen3's tokenizer.json declares an NFC normalizer, so it counts both forms the same. The other five do not normalize, and they count the decomposed text at 1.6 to 3.0 times. None of the provider docs we read say whether the API normalizes text before billing it. If it does not, an agent can double its input bill without changing a visible character.
The payloads agents carry
Agents on a payment rail do not mostly send prose. They send tool results: receipts, hashes, addresses, signatures, price feeds. We measured those too. The real rows are a Python file from the standard library, a JavaScript file from npm, our own OpenAPI document, one of our blog pages, ten transaction receipts from Base and the PAYMENT-REQUIRED header our gateway returned. The hash, address, base64, UUID, CSV and emoji rows are synthetic, generated with a fixed seed.
# real tokens / quote's estimate, same method, 6 Oct 2026
payload chars quote o200k cl100k Qwen3 DSV3.1 GLM4.6 Tekken
Python source (json/encoder.py) 16,074 4,162 0.83 0.82 0.83 0.89 0.82 0.88
JavaScript source (npm install.js) 5,383 1,393 0.94 0.94 0.94 0.98 0.94 0.96
JSON, minified (our openapi.json) 62,176 16,855 0.94 0.93 0.97 1.00 0.94 1.02
JSON, 2-space indent (same doc) 98,589 26,636 0.83 0.83 0.85 0.86 0.84 0.87
HTML (one of our blog pages) 36,474 9,396 1.02 1.02 1.05 1.08 1.02 1.09
10 Base receipts, minified JSON 24,268 6,486 1.57 1.57 2.67 1.62 1.70 2.71
our PAYMENT-REQUIRED header (base64) 552 139 2.55 2.79 2.85 2.61 2.79 3.03
Hex tx hashes (0x + 64), synthetic 20,099 5,100 2.32 2.31 3.48 2.35 2.68 3.52
EVM addresses (0x + 40), synthetic 12,899 3,300 2.35 2.35 3.47 2.39 2.69 3.51
Base64 65-byte signatures, synthetic 26,699 6,750 2.70 2.83 2.92 2.76 2.83 3.01
UUIDv4 list, synthetic 11,099 2,850 2.49 2.49 3.44 2.51 2.77 3.46
CSV price feed, synthetic 11,573 2,994 2.01 2.01 3.86 2.01 2.42 3.86
Emoji line, synthetic 3,999 1,426 2.67 3.45 3.44 3.03 3.45 4.14
Code, JSON and HTML sit near or below break-even, from 0.82 to 1.09. Indented JSON is cheaper than it looks to the estimator. JSON escaping counts every embedded quote twice, while tokenizers fold runs of spaces into single tokens.
Everything a payment agent handles is under-quoted two to four times. The ten receipts come from Base block 52,244,264 via eth_getTransactionReceipt, and they run 1.57 to 2.71. Our own 402 header runs 2.55 to 3.03 as base64. Decoded, the same header is 147 tokens on o200k_base instead of 355. Base64 more than doubles the tokens of the JSON it wraps.
Digits explain the spread between tokenizers. The pre-tokenizer patterns of Qwen3 and Tekken split every digit on its own (\p{N}). o200k_base, cl100k_base, DeepSeek-V3.1 and GLM-4.6 group up to three (\p{N}{1,3}). Hex and price feeds cost about 3.5 to 3.9 times the estimate on the first pair and 2.0 to 2.7 on the rest.
Images: the miss in the other direction
The estimator treats an image sent as a base64 data URL as text. We made a 1000x1000 JPEG of 69,433 bytes. As a data URL it is 92,603 characters, and the quote estimates the message at 23,181 input tokens. Anthropic documents image cost as ⌈width / 28⌉ × ⌈height / 28⌉ visual tokens, which is 1,296 for this size on every Claude tier. The quote counts 17.9 times what Claude bills.
# api.llm4agents.com, anthropic/claude-opus-5.5, max_tokens 500, 6 Oct 2026
request 402 amount
"Describe this image." (text only) $0.02
+ one 1000x1000 JPEG as a data URL $0.13
+ five copies $0.57
+ ten copies $1.13
At Claude's documented count, ten such images are 12,960 input tokens. At Opus 5.5's list price of $4 per million, that is $0.052. Add the full 500 output tokens and our fee, and the call costs $0.08 after rounding. The quote asks for $1.13.
That number crosses a line the agent did not draw. Since spend controls shipped, x402's TypeScript SDK caps each payment at $1 by default, as its changelog records. Our spend controls audit covers that cap. A careful agent on default settings refuses to pay for a request that costs eight cents.
Who pays for the miss
A wrong estimate does not always produce a wrong bill. The quote also reserves the full max_tokens of output, and the agent rarely uses all of it. The real cost exceeds the quote when the uncounted input, priced at the input rate, is larger than the unused output reserve priced at the output rate, plus whatever the round-up to the cent leaves. Small calls have a cent of slack. Large ones do not.
We measured a large one. Thirty-four copies of the Chinese declaration make 101,658 characters. Sent to openai/gpt-5 with max_tokens 1,000, the quote is $0.06, built on an estimate of 26,212 input tokens. On o200k_base, the encoding tiktoken maps to GPT-5, the text is 80,478 tokens: 3.07 times the estimate. At GPT-5's list price of $1.25 per million plus our fee, the input alone is $0.12. With zero output the call costs $0.13. With the full 1,000 output tokens it costs $0.14. Both are more than twice the quote, before OpenAI's formatting tokens.
What happens next depends on the path. In our gateway's source today:
- Walk-up, non-streaming: the handler returns HTTP 500
x402_settlement_overflowafter the upstream call has completed. The agent gets no answer. Nothing settles. We pay the provider. - Walk-up, streaming: settlement is capped at the quote. The agent gets the answer, and we absorb the difference.
- Bearer, prepaid: the reserve is debited and the excess is recorded as spend, but it is never taken from the balance. We absorb it there too.
So an under-quote is either a subsidy for token-dense input or a failure on the path an honest walk-up agent is most likely to use.
An over-quote costs something different on each rail. On the prepaid rail the reserve is refunded at settlement, so the over-quote on English costs float, not money. On the walk-up rail the requirement is exact, and the agent signs the inflated amount. Our handler then computes the real cost and passes it to the x402 middleware as a settlement override. That is the mechanism x402 built for upto: the upto scheme doc says to set the price "to the maximum authorized amount" and "use settlement overrides to charge the actual amount." The x402 core resource server is explicit about the other case. Overriding the amount "is only valid in schemes that support partial settlement," and with exact it "will likely cause settlement verification to fail."
What it means for LLM4Agents
The 402 amount is the price signal autonomous agents read. An agent comparing two models, or two gateways, compares 402 amounts. Ours is right for English prose, with a margin of about a third in our favor. It is wrong in our disfavor for Arabic, Indic and CJK text on every tokenizer we ran. And it is wrong for the payloads our core users carry.
That last point is the uncomfortable one. LLM4Agents is built for agents that pay. Their context is full of hashes, addresses, base64 payment headers and receipts, which is exactly where the estimator misses most. A Spanish-speaking agent is priced close to fairly. A Japanese-language agent is quoted 22 to 40% of its real input.
The threat runs both ways. A systematic under-quote is a subsidy anyone can find by probing 402s, as we just did, and a cost we carry on every token-dense request. A systematic over-quote makes us look expensive next to a competitor that counts better. On images it pushes careful agents past their own spending caps.
The opportunity is that the fix is mostly public knowledge. OpenAI and the open-weight families publish their tokenizers. Anthropic, OpenAI and Google offer counting endpoints. Image costs are documented formulas. A gateway that quotes within a few percent can offer exact with a straight face, or upto with a tight ceiling. "The 402 is the bill" is a product feature, and not many gateways can claim it.
Staying on the frontier
In order of cost and urgency.
1. Count with the real tokenizer where it is public. Use o200k_base for the GPT-4o and GPT-5 families, as tiktoken maps them, and the published tokenizer.json for open-weight families such as Qwen, DeepSeek, GLM and Mistral. Choose the encoding per model and never fall back to one default: cl100k_base and o200k_base differ by 3.3x on Hindi. Plan for the size. The o200k_base file is 3.6 MB, Qwen3's tokenizer is 11.4 MB and GLM-4.6's is 20 MB.
2. Calibrate the closed families against their endpoints. Anthropic's count_tokens is free and rate-limited separately. Use it, and Google's and OpenAI's counters, offline to build per-model and per-script ratios instead of calling them on every request. Re-run the calibration whenever a model changes generation. Claude 4.7's tokenizer moved counts by about 30% in one release.
3. Ship a better fallback heuristic now. Counting UTF-8 bytes instead of UTF-16 units is a one-line change. Across our six tokenizers it narrows the language spread from 0.75–4.85 to 0.45–1.82, and on o200k_base alone from 0.75–3.33 to 0.45–1.16. Stop counting JSON escapes. Add a rule for long hex, base64 and digit runs, which stay two to four times under because they are ASCII and bytes do not help.
4. Price images by their dimensions, not their base64. Read width and height from the image header and apply the documented rule: ⌈w/28⌉ × ⌈h/28⌉ for Claude, 258 tokens per tile for Gemini, OpenAI's counting endpoint for its models. That alone removes the 17.9x over-count and keeps image requests under default spend caps.
5. Count what the quote skips. Tool definitions and response_format are still free in the estimate, as the earlier audits found. Add them, plus each provider's documented overhead for roles and formatting.
6. Move LLM walk-up to upto. With a ceiling that is accurate to a known margin and settlement on the native count, an over-quote stops costing agents money and the settlement override becomes protocol-correct. Our upto post covers the scheme. Size the safety margin per model and per script bucket. When the estimate is too uncertain to bound, refuse before the call, not with a 500 after it.
7. Measure the drift and show it. Log the estimate against the provider's reported usage for every call, per model and per script. Alert when a model's ratio moves, because a tokenizer release can move it overnight. Return the estimated input tokens with the 402 and the native count with the receipt, so an agent can audit its own bills.
Steps one to four make the number right. Five and six make it safe. Seven keeps it right.