The cache decides the price: prompt caching and the x402 upto scheme
Prompt caching made the price of an LLM request depend on what happened in the previous five minutes. x402's dominant payment scheme demands a number before the work starts. Something has to give, and the spec already says what.
An autonomous agent that pays per call needs a price. HTTP 402 is a quote: the server states what it wants, the agent signs for that amount, the facilitator settles it. The model works perfectly for a fixed-price resource — an API call, a data record, a render.
Inference is not a fixed-price resource, and prompt caching is why. The same request, byte for byte, costs one number on a cold prefix and roughly a tenth of that on a warm one. Which one you get depends on whether another request touched the same prefix recently, on which model served it, and on which workspace the cache lives in. None of that is knowable at the moment the 402 is written.
We read the pricing and caching documentation of the three major providers, then the x402 scheme specifications, and checked the on-chain piece ourselves. The mismatch is real, it is currently absorbed by overcharging, and the primitive that fixes it has been merged in the x402 repository since March.
Caching turned price into a function of history
Start with the numbers, because the magnitude is the whole argument. Anthropic's pricing page lists prompt caching as three multipliers on the base input rate: a 5-minute cache write at 1.25x, a 1-hour cache write at 2x, and a cache read at 0.1x. On Claude Opus 5, base input is $5/MTok, so a 5-minute write is $6.25/MTok, a 1-hour write is $10/MTok, and a hit is $0.50/MTok.
The documented break-even follows directly: with the 5-minute TTL, caching pays off after a single read (1.25x + 0.1x against 2x uncached); with the 1-hour TTL, it takes two reads (2x + 0.2x against 3x). The 1-hour option is not "better caching," it is a bet on the gap between requests.
Three implementation details decide whether any of that happens at all. The prompt caching documentation puts the maximum at four explicit breakpoints per request, with a lookback window of 20 blocks per breakpoint. The cache lifetime is measured from the start of the request that writes or reads the entry, not from the end of its response — a four-minute generation leaves about one minute for the next request to begin. And a read refreshes the entry at no additional cost, which means continuous traffic keeps a 5-minute entry alive indefinitely.
Then there is the minimum. A prefix shorter than the model's threshold silently does not cache: no error, just cache_creation_input_tokens: 0. The thresholds are not monotonic across generations:
// Minimum cacheable prefix, Claude models
512 tokens Opus 5, Fable 5, Fable 5.1, Mythos 5, Mythos 5.1
1024 tokens Opus 4.8, Sonnet 5, Sonnet 4.6, Sonnet 4.5, Opus 4.1, Opus 4, Sonnet 4
2048 tokens Opus 4.7, Mythos Preview, Haiku 3.5
4096 tokens Opus 4.6, Opus 4.5, Haiku 4.5
A 3,000-token prefix caches on Opus 5 and on Sonnet 5. The same prefix, routed to Haiku 4.5, does not cache at all. Nothing in the response says so except a zero.
The accounting fields are the other half of the contract. Anthropic reports cache_creation_input_tokens (written), cache_read_input_tokens (retrieved), and input_tokens — defined as tokens that were neither read from nor used to create a cache, that is, everything after the last breakpoint. The three sum to the real input total. Any billing code that reads input_tokens alone and calls it "the input" is undercounting by whatever the cache absorbed.
Three providers, three cache contracts
The shapes differ enough that a routing gateway cannot treat them as one feature.
Anthropic's caching is explicit. You place cache_control breakpoints, or use top-level automatic placement, and you choose the TTL. The invalidation table is worth memorizing: tool definitions invalidate the tools, system, and messages caches; toggling web search or citations, or changing the speed setting, invalidates system and messages; tool choice, images, thinking parameters and the effort setting invalidate messages. Effort is the one that surprises people — raising effort mid-conversation to get a better answer throws away the cached history that made the conversation cheap.
OpenAI's caching is automatic. Per the prompt caching guide, GPT-5.6 and later require 1,024 visible input tokens, cached tokens cost 0.1x the uncached input rate, and cache writes cost 1.25x — the same two multipliers Anthropic charges, arrived at independently. Retention is controlled by prompt_cache_options.ttl on GPT-5.6+ (minimum "30m") and by prompt_cache_retention on earlier models, which takes "in_memory" or "24h". Reuse breaks on a model change, a change to tool definitions or their ordering, a change to text.format, a change to reasoning.effort, or a compaction event.
The field names differ by endpoint. The Responses API reports usage.input_tokens_details.cached_tokens and usage.input_tokens_details.cache_write_tokens; Chat Completions reports the same two under usage.prompt_tokens_details. Since llm4agents speaks the OpenAI-compatible surface, that second pair is the one every agent SDK in the wild already knows how to parse.
Google's is implicit by default. The context caching documentation states implicit caching is on for all Gemini 2.5 and newer models, with minimums of 2,048 tokens on Gemini 2.5 Flash and Pro and 4,096 on the Gemini 3.x Flash line and Gemini 3.1 Pro Preview. Cached totals surface as usage.total_cached_tokens. The Interactions API supports implicit caching only — explicit cache objects are not available there.
The same request at two prices
Take a working agent: a 30,000-token stable prefix (system prompt, tool definitions, retrieved documents), a 500-token volatile tail, 800 tokens of output, on Claude Opus 5.
// Cold — the prefix is written to cache
cache_creation_input_tokens: 30000 × $6.25/MTok = $0.1875
input_tokens: 500 × $5.00/MTok = $0.0025
output_tokens: 800 × $25.00/MTok = $0.0200
─────────
$0.2100
// Warm — the prefix is read from cache
cache_read_input_tokens: 30000 × $0.50/MTok = $0.0150
input_tokens: 500 × $5.00/MTok = $0.0025
output_tokens: 800 × $25.00/MTok = $0.0200
─────────
$0.0375
Identical request bytes. A 5.6x spread on the total, and 10.9x on the input side alone. Now price it the way an exact-scheme endpoint has to: the server does not know, at quote time, which of those two it is about to do, so it quotes the cold number. Over a thousand turns of a warm loop, the agent pays $210.00 for work that cost $37.67.
The alternative failure is symmetric. Quote the warm number and every cold start — every deploy, every idle gap past the TTL, every first request of the morning — is served below cost. The gateway either overcharges systematically or bleeds on misses. There is no third quote.
Where routing and fallback destroy the cache
This is the uncomfortable part for anyone selling model routing, us included. We have argued before that a fallback chain is the difference between an agent that survives a provider incident and one that does not. That argument still holds. What we did not price is that every fallback hop is a cache event.
Caches are model-scoped. OpenAI lists a model change first among the things that break reuse, and Anthropic's cache key is derived from the rendered prompt bytes for a specific model. So a 429 on the primary that trips the router into the secondary does not merely change the per-token rate — it converts what would have been a 0.1x read into a 1.25x write. On the example above, that single hop costs $0.1725 more than the request it replaced, and it leaves the primary's entry to expire unread.
Two further details make cross-model routing worse than it looks. The minimums above mean a prefix that caches on the primary can silently fail to cache on the fallback: 512 tokens on Opus 5 against 4,096 on Haiku 4.5 is a factor of eight. And token counts are not portable across generations — Anthropic's pricing page notes that Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text. The same bytes, routed to a different model family, are a different quantity of the thing you are billing.
The practical conclusion is not "stop falling back." It is that a fallback decision made on health and latency alone is made on two thirds of the inputs. Warm-prefix availability belongs in that decision, and the price delta belongs in the response.
x402 quotes before the work; caching prices after it
The x402 v2 specification puts a PaymentRequirements object in the accepts array of the 402 response, with seven fields: scheme, network, amount, asset, payTo, maxTimeoutSeconds, and an optional extra. There is no separate ceiling field. amount is "the required payment amount in atomic token units," stated before the resource runs.
Under the exact scheme that is a hard commitment. The EVM implementation supports EIP-3009 transferWithAuthorization, Permit2, and ERC-7710 delegation, and the specification is explicit about the guarantee that makes the scheme safe: "In all cases, the Facilitator cannot modify the amount or destination. They serve only as the transaction broadcaster." That property is exactly what you want for a fixed-price resource, and exactly what you cannot use for metered inference. The signed authorization carries a fixed value; discovering after generation that the request was cheaper leaves you with an out-of-band refund and a reconciliation problem.
We covered the EIP-3009 settlement path when it was the obvious default for agent payments. It still is, for anything whose price is a constant. Inference is not.
The upto scheme is the missing primitive
x402 already has the answer. The upto scheme exists for "usage-based pricing where the final cost is unknown until after resource consumption," and states plainly that "the actual amount charged is determined at settlement time based on resource consumption during the request."
The mechanism is a phase-dependent reading of the same field. At verification, amount "represents the maximum amount the client authorizes." At settlement, it "represents the actual amount to settle, which MUST be less than or equal to the previously authorized maximum." The resource server sets the settlement amount from measured consumption; the facilitator "MUST re-verify the client's signature using the authorized maximum (permitted.amount), not the settlement-time requirements.amount." Single-use settlement, explicit time bounds through validAfter and deadline, and recipient binding are all enforced.
The EVM implementation is where the design constraint becomes concrete. The EVM scheme document uses Permit2's permitWitnessTransferFrom, and says why the obvious alternative is unavailable: "EIP-3009 (transferWithAuthorization) is not supported for the upto scheme because it requires exact amounts at signature time." The settled amount is an independent parameter on the proxy's settle function, and "MUST be <= the authorized maximum." The witness struct carries a facilitator field for access control.
That proxy is not a draft. The specification names x402UptoPermit2Proxy at 0x4020A4f3b7b90ccA423B9fabCc0CE57C6C240002; we queried Base mainnet directly and the address returns 3,142 bytes of deployed bytecode. The repository history shows the scheme has been maintained for six months:
// specs/schemes/upto — commit history
2026-03-04 #1074 Add upto payment scheme specification for EVM
2026-03-25 #1773 feat: add upto to typescript sdk
2026-03-31 #1880 fix: evm contract deploys
2026-06-12 #2607 clarify settle-time verification for partial settlements
2026-07-22 #2697 docs(svm): add `upto` SVM scheme specification
2026-08-12 #3094 feat(ts): svm upto paymentflow
2026-09-03 #3346 feat(ts): delegated receiver authorizer for SVM upto
2026-09-09 #3431 fix(svm): split upto delegated-auth store errors
The cost of adopting it is honest and worth stating: Permit2 needs a prior approval. The specification lists three ways to get one — a standard on-chain approve(Permit2) transaction, a sponsored ERC-20 approval, or an EIP-2612 permit where the token supports it. EIP-3009's appeal was that a fresh wallet could pay on its first request with no setup transaction. Under upto, that setup moves to registration time. For an agent that will make thousands of metered calls, that is a one-time cost against a structural overcharge on every one of them.
What the receipt does not say
Settlement reports back. The v2 SettleResponse carries success, transaction, network, optional errorReason and payer, an extensions object, and an optional amount — "the actual amount settled in atomic units." So the chain records what was charged.
Nothing standard records why. The offer-and-receipt extension adds server-signed offers and receipts for dispute evidence and auditability, delivered in the success response body at extensions["offer-receipt"].info.receipt. Its receipt payload is version, network, resourceUrl, payer, issuedAt, and an optional transaction. The extension is explicitly "privacy-minimal by default" and "intentionally omits transaction references to reduce correlation risk." There is no usage or metering detail in it.
For a metered inference call that is a gap. The agent can verify that it paid an amount, and that the amount was within what it authorized. It cannot verify that the amount corresponds to the token accounting the provider reported — that the 30,000 tokens billed as cache reads were in fact reads, or that the fallback hop that turned a read into a write actually happened. Today the only cross-check is the usage object in the response body, signed by nothing.
The other axis of the same problem is batching. Both major providers price asynchronous work at half: Anthropic's Message Batches API is a 50% discount on both input and output tokens, with most batches finishing in under an hour and a hard 24-hour expiry; OpenAI's Batch API is the same 50%, with a 24h completion window, up to 50,000 requests and 200 MB per batch. Anthropic's caching multipliers stack with the batch discount — with one caveat worth knowing: max_tokens: 0 cache pre-warming is rejected inside a batch, because an ephemeral entry written during batch processing would likely expire before the follow-up request runs. A batch lane is a second price for the same work, and it is a price that can only be settled after the batch returns.
What it means for LLM4Agents
We are the party that holds the cache. The workspace is ours, the prefix layout is ours, the routing decision is ours. That makes the hit rate a design output of the gateway, not a property of the customer's prompt — and it makes the price of a call something we discover during the call.
Our internal accounting already has the right shape. The reserve-proxy-settle path holds a maximum against an agent's balance, proxies the request, then settles the measured amount. That is upto semantics implemented off-chain against our own ledger. Moving the same shape onto the x402 wire is not a redesign; it is making an internal guarantee externally verifiable.
The threat is sharper than the opportunity. If we keep quoting exact at a worst case, we are structurally overcharging exactly the traffic we most want — the long-running, warm-prefix agent loop. Any competitor running identical infrastructure who settles on measured consumption undercuts us on the same hardware, and can prove it on-chain. The fleet economics we published assumed inference cost was the floor; caching moved the floor and left the quote where it was.
Routing is where the two problems meet. Fallback is our reliability story and our largest uncontrolled cost event, and today an agent has no way to see that its cheap request became an expensive one because a provider returned a 429.
Staying on the frontier
1. Expose cache accounting in the OpenAI-compatible response. Normalize Anthropic's cache_read_input_tokens / cache_creation_input_tokens, OpenAI's cached_tokens / cache_write_tokens, and Google's total_cached_tokens into usage.prompt_tokens_details.cached_tokens and cache_write_tokens on every call. It is the field agent SDKs already parse, and it is the precondition for everything below — an agent that cannot see its hit rate cannot optimize it or dispute it.
2. Quote inference routes with upto, keep exact for fixed-price ones. The ceiling is the cold-cache price; the settlement is the measured one. Do the Permit2 approval at agent registration so the first paid call is never blocked on a setup transaction, and keep exact on endpoints whose price is genuinely constant — there is no reason to pay Permit2's complexity for those.
3. Make the fallback chain cache-aware. Order candidates by warm-prefix availability alongside health and latency, and check the candidate's minimum cacheable prefix before routing a short prompt to a model that will silently refuse to cache it. When a hop happens anyway, report the price delta in the response rather than absorbing it into an averaged rate.
4. Publish the prefix contract. Freeze the system block, serialize tool definitions deterministically, and push per-request identifiers and timestamps after the last breakpoint — then document that layout so customers can build against it. The four-breakpoint limit and the 20-block lookback are real constraints on how many stability tiers a prompt can have; customers designing blind will exceed them.
5. Ship a batch lane. Half price on both directions for evaluation runs, bulk classification and overnight work, settled after the batch returns — which requires upto to be in place first, and a 24-hour expiry to be handled as a real failure mode rather than a timeout.
6. Push for usage detail in the receipt. The offer-and-receipt extension is privacy-minimal by design, and that is defensible for a one-shot purchase. Metered inference needs an optional signed usage block alongside the settled amount — token counts by category, model served, cache disposition — so an agent can reconcile the charge against the work instead of trusting the number. That is the field worth proposing, and it is adjacent to the auth-capture work already in the specification.
Pay for what the meter measured, not the worst case
OpenAI-compatible gateway, per-call cache accounting, x402 settlement on measured consumption.
Register an agent