Wait or pay? 402, 429 and Retry-After for paying agents
An agent that pays per call can be refused in two ways. A 402 means "pay first". A 429 means "wait". We checked the IETF draft, the x402 code, four official LLM SDKs and three providers' error docs, and sent unpaid requests to 1,951 x402 sellers. Few of them agree on which refusal is which, or on how long to wait.
This matters more for a paying agent than for a person with an API key. A person who hits a limit reads the error and adjusts. An agent has to decide automatically, on every refusal, whether to pay, wait, switch provider or give up. If it pays when it should wait, it spends money for nothing. If it waits when it should pay, it stalls. If it retries a billing error, it wastes time and sometimes a signed payment. If it gives up on a limit that clears in fifteen seconds, it loses the task.
Our sources, all read on 26 September 2026: draft-ietf-httpapi-ratelimit-headers-11; RFC 9110 and RFC 6585; the x402 reference repository at commit 4fcf836 (25 September); the released OpenAI and Anthropic SDKs for Python and TypeScript; and the current error and rate-limit pages from OpenAI, Anthropic and OpenRouter. The measurements come from the x402 Bazaar discovery catalog, fetched at 09:11 UTC, plus one unpaid request to each listed host. We made no payments.
Two status codes, one decision
HTTP defines both codes, and says little about either. RFC 9110 gives 402 one sentence: it "is reserved for future use." x402 is what the ecosystem built on that sentence. RFC 6585 defines 429 as "too many requests in a given amount of time". The response MAY carry a Retry-After header and "MUST NOT be stored by a cache." Retry-After is either an HTTP date or a whole number of seconds.
Neither code tells a client how much capacity is left. Years of ad hoc headers filled that gap: X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset and many variants. The IETF HTTPAPI working group has been standardizing a replacement since December 2020. Revision -11 of RateLimit header fields for HTTP was published on 23 May 2026 and expires on 24 November. It is still a working-group document, not yet with the IESG.
The draft defines two structured fields. RateLimit-Policy describes the quota and should stay stable across responses. RateLimit reports what is left right now:
RateLimit-Policy: "burst";q=100;w=60,"daily";q=1000;w=86400
RateLimit: "burst";r=37;t=18
// q = quota, w = window in seconds, r = remaining, t = seconds until the window ends
// optional: qu = quota unit, pk = partition key (a byte sequence)
Four details matter for agents. First, the fields are hints, not guarantees. The client "MUST NOT assume that a positive available quota is a guarantee." Second, when Retry-After and RateLimit are both present, Retry-After takes precedence. Third, the draft registers three problem types: quota-exceeded and abnormal-usage-detected with 429, and temporary-reduced-capacity with 503. Fourth, the draft registers only three quota units: requests, content-bytes and concurrent-requests.
That last point is the gap for this industry. LLM providers limit tokens per minute, and x402 sellers limit money. The draft can express neither without a new registry entry or a vendor-prefixed parameter. Its appendix does name "monetization" as a reason servers use quotas, but it defines no unit for it. The -11 text is also inconsistent on the default unit: section 3.1.2 calls it requests, while the IANA table lists request. A client that matches unit strings exactly should accept both until that is fixed.
How providers say "out of money"
When the problem is money rather than speed, three common upstreams for agents answer in different ways. This is from each provider's current error and rate-limit documentation:
rate limit
OpenAI 429 + Retry-After when present
Anthropic 429 rate_limit_error + retry-after
OpenRouter 429 + Retry-After when present
spend limit you set yourself
OpenAI 429 organization_spend_limit_exceeded / project_spend_limit_exceeded
Anthropic 400 invalid_request_error
OpenRouter 402 limit_source: openrouter_key_limit
cap set by the provider
OpenAI 429 organization_usage_limit_exceeded
Anthropic 429 enforced_spend_limit_reached, no retry-after
prepaid balance gone
OpenAI 429 credit_balance_exhausted
OpenRouter 402 limit_source: openrouter_credits
billing or payment details
Anthropic 402 billing_error
spend already in flight
OpenRouter 402 in_flight_budget_exhausted + Retry-After
overloaded
OpenAI 503 server_is_overloaded
Anthropic 529 overloaded_error
OpenRouter 503 no provider available / provider_overloaded
OpenAI uses 429 for every money condition and separates them only through error.code. It warns that "the broader error.type can still be insufficient_quota." Its advice is clear: "Retrying billing, spend, or quota errors won't restore API access." Its rate-limit guide adds that Retry-After "does not mean that quota, billing, or other errors that require user action can be resolved by retrying."
Anthropic uses three codes for money. A spend limit that you set returns 400 invalid_request_error. A problem with billing details returns 402 billing_error. Reaching your usage tier's monthly cap returns 429 with the same rate_limit_error type as ordinary throttling. The rate-limit page says usage then "pauses until 00:00 UTC on the first day of the next month". It also says the response has "no retry-after header" and that retrying, "including the SDKs' automatic retries, fails until access resumes." The distinguishing field is error.details.error_code: enforced_spend_limit_reached.
OpenRouter uses 402 for money, which is the closest to x402's reading of the code. One of its 402s, though, means "wait". OpenRouter charges when a request finishes. It therefore holds an estimated cost up front: input tokens plus the completion tokens allowed by max_tokens. It rejects new requests whose estimate does not fit alongside the current holds. That 402 carries reason: in_flight_budget_exhausted and a Retry-After header. The docs point out that "none of them" (the OpenAI, Anthropic, Vercel AI and OpenRouter SDKs) "retries a 402 on its own." Readers of our reserve-proxy-settle post will recognize the design, because it is a reserve.
So an agent routed across these three can see "out of money" as a 400, a 402 or a 429. It can also see "wait" as a 402 or a 429. Add x402, and a 402 can also be an offer: a price the agent is expected to pay right away. The status code alone does not tell the agent what to do. It has to read the body, and every provider puts the reason in a different field.
What the SDKs actually do
Most agents don't parse these responses themselves. The provider's SDK does it for them. We read the retry code in the latest releases of the four official clients: openai-python 3.19.2, openai-node 7.23.0, anthropic-sdk-python 1.8.0 and anthropic-sdk-typescript 0.128.0, all released between 22 and 24 September 2026.
In the parts that matter here, all four behave the same way:
- They retry 408, 409, 429 and any status of 500 or above, at most twice by default.
- They obey a non-standard
x-should-retry: true|falseheader before any of those rules. - None of them has an exception class for 402 or retries it.
- None of them parses
RateLimit,RateLimit-Policyor anyx-ratelimit-*header. - When they have no server hint, they back off exponentially from 0.5 to 8 seconds, with up to 25% jitter.
Where they differ is Retry-After. Both Python clients and OpenAI's Node client used the same generated rule, which honored the header only up to 60 seconds. This year the two Python clients changed it in opposite directions:
SDK release honors Retry-After above that ceiling
openai-python 3.19.2 up to 120 s no retry, error returned
openai-node 7.23.0 up to 60 s default backoff, 0.5-8 s
anthropic-sdk-python 1.8.0 up to 4,294,967 s (ceiling is a platform limit)
anthropic-sdk-typescript 0.128.0 up to 2^31-1 ms default backoff, 0.5-8 s
OpenAI's Python change (#3555, merged 30 July) explains its reasoning. Waits above 60 seconds were being replaced by "a much shorter exponential delay." Beyond two minutes, the SDK now "surfaces the API error instead of blocking a synchronous worker." Anthropic's Python client went the other way on 12 September: "If the API asks us to wait a certain amount of time, just do what it says," capped at 4,294,967 seconds, about 49.7 days. OpenAI's Node client still quietly replaces anything above 60 seconds with a delay of at most 8 seconds. So given the same 90-second Retry-After, one official SDK waits 90 seconds and another retries after a few seconds.
Apply that to the table above. OpenAI's credit_balance_exhausted is a 429, so every official OpenAI client retries it twice. The only way to stop that is x-should-retry: false, and the error-code page does not say whether billing 429s carry it. Anthropic's spend-cap 429 has no retry-after, so the clients back off and retry twice, as its own docs say they will. Neither case is expensive for an agent with an API key. It is a few wasted seconds.
For an agent that pays per call, the same retry can move money. Both TypeScript clients accept a custom fetch, and the obvious way to use an x402 gateway from them is to pass in wrapFetchWithPayment from @x402/fetch. When the SDK retries after a 429, the wrapper runs from the start: an unpaid request, a new 402, a new signature with a new nonce, and a new paid request. Whether that costs the agent twice depends on the seller, as the next section shows.
x402 has no word for "wait"
The x402 HTTP transport (specs/transports-v2/http.md) maps protocol errors to four statuses: 402 when payment is required, 400 for an invalid payment, 402 again when verification or settlement fails, and 500 for a server error. It never mentions 429, 503 or Retry-After. The MCP transport carries payment state in _meta and the A2A transport carries it in task metadata. Neither has a throttling signal. The core specification has one non-terminal code, settlement_pending, added on 17 August in #3083. It covers a facilitator that broadcast a settlement but could not confirm it. It says nothing about the resource server being busy.
The reference client follows the same model. wrapFetchWithPayment returns any non-402 response unchanged. On a 402 it parses the requirements, signs, and retries once. It does not read Retry-After, and it does not reuse requirements across calls. Every paid call therefore costs at least two HTTP requests: the unpaid request that receives the price, and the paid retry.
The reference server decides whether a throttled request costs money. The TypeScript middleware has three payment flows. With authorization, the default, it verifies before the handler, runs the handler, and settles afterwards. If the handler returns any status of 400 or above, the Express adapter cancels settlement with reason handler_failed. A 429 from inside the handler therefore costs nothing. With upfront, settlement happens before the handler runs. A 429 from inside the handler then arrives after the USDC has moved. The two-phase gap audit explains why sellers choose upfront. The cost for the buyer shows up here. An SDK that retries a 429 from an upfront seller pays again on each retry.
The rule for sellers is simple. Where the limiter sits in the request path determines who pays for throttling. A limiter placed before the payment middleware rejects the request before any payment is verified. A limiter inside the handler rejects it after verification, which is free with authorization and already paid with upfront.
What 1,951 x402 sellers actually send
To see what sellers do in practice, we fetched the full Bazaar catalog from the CDP discovery API: 17,659 entries across 2,028 hosts. For each host we took the resource with the most recorded calls, skipped templated paths, and sent one unpaid request using the method the catalog declares. That produced 1,951 hosts. We then recorded every rate-limit header on the responses.
x402 Bazaar census, 26 Sep 2026, 09:13 UTC
hosts probed, one unpaid request each 1,951
answered 402 1,903
with a PAYMENT-REQUIRED header 1,842
with any rate-limit header 163 (8.6%)
RateLimit-Limit / -Remaining / -Reset 99
RateLimit-Policy as "N;w=M" 99
X-RateLimit-* 54
RateLimit: limit=, remaining=, reset= 10
current syntax ("name";q=;w= / r=;t=) 7
with Retry-After on the 402 itself 2
Fewer than one in eleven sellers tells a buyer anything about its limits. Those that do mostly use syntax the IETF has since replaced. The most common form, RateLimit-Policy: 120;w=60 with separate -Limit, -Remaining and -Reset headers, is the draft-06 format from December 2022. It is exactly what express-rate-limit 8.7.0 sends when standardHeaders is set to true. Seven hosts use the current structured syntax. Five of them name their policy "120-in-1min", "60-in-1min" and so on, which is the same library's default identifier in its draft-8 mode.
The legacy headers disagree on what they mean. X-RateLimit-Reset appeared on 39 hosts with three incompatible meanings: 13 send seconds remaining, 23 send a Unix timestamp in seconds, and 3 send a Unix timestamp in milliseconds. RateLimit-Reset was a delay in seconds on 83 hosts and an ISO timestamp on one. Add OpenAI's duration strings (6m0s) and Anthropic's RFC 3339 timestamps, and an agent that reads "reset" headers has to handle five encodings of the same idea.
The two sellers that put Retry-After on the 402 itself show the ambiguity well. One sends Retry-After: 60 together with remaining: 119 out of 120. That is a price with an instruction to wait a minute, on a quota that is almost full. A client that follows OpenRouter's advice to honor Retry-After on 402s would pause for no reason.
Agent gateways are a subset worth looking at. Of the 21 hosts whose probed resource was a /chat/completions or /v1/messages endpoint and answered 402, 7 advertise a limit. All 7 use legacy headers, and none uses the current syntax.
We then tested whether the advertised limits apply to the unpaid path. We picked the 13 hosts that advertised a quota of 20 requests or fewer, one per registrable domain, and sent each its advertised limit plus two unpaid requests, one after another. Four hosts switched from 402 to 429 on exactly the request after their limit. Eight kept answering 402 past it. One rejected our empty request body with a 400 before reaching the paywall. Three of the four 429s carried Retry-After (55, 55 and 49 seconds). The fourth carried only a Unix-epoch reset. None carried a PAYMENT-REQUIRED header.
On those four hosts, then, unpaid price checks count against the same quota as paid calls. Combine that with the reference client's two-request pattern and an advertised limit of 10 per minute allows at most 5 paid calls per minute. The buyer cannot know this from the headers.
Idempotency is the missing third piece
Retrying a paid request raises a question that rate limits alone cannot answer: did the first attempt run? The general HTTP answer was supposed to be the Idempotency-Key header. Its IETF draft reached -07 on 15 October 2025 and expired on 18 April 2026 with no newer version.
x402 has its own mechanism, the payment-identifier extension, covered in our extensions layer audit. The client sends an id of 16 to 128 characters. The same id with the same payload returns the cached response, and the same id with a different payload returns 409. In the catalog, 1,946 entries on 77 hosts declare it, and 24 entries on 23 hosts make it required.
The specification does not say whether an error response may be cached under an id, and it sets no lifetime for the cache. That gap matters for rate limits. RFC 6585 says a 429 "MUST NOT be stored by a cache." Suppose a seller throttles a paid request and stores the 429 under its payment-identifier. The client waits the 55 seconds it was told to, retries with the same id as the extension instructs, and gets the stored 429 back. A short wait has become a permanent failure for that id.
A decision function an agent can run
Until the specifications converge, the agent has to decide for itself. Below is the classifier we would put in front of any paying agent. code is the provider's machine-readable reason: error.code at OpenAI, error.details.error_code at Anthropic, error.metadata.reason at OpenRouter. paid says whether this response answered a request that already carried a payment.
type Decision =
| { kind: 'pay'; requirements: string }
| { kind: 'wait'; seconds: number }
| { kind: 'backoff' }
| { kind: 'failover' }
| { kind: 'stop'; reason: string }
const MAX_WAIT_S = 120
const BILLING_CODES: ReadonlySet<string> = new Set([
'credit_balance_exhausted', 'organization_spend_limit_exceeded',
'project_spend_limit_exceeded', 'organization_usage_limit_exceeded',
'enforced_spend_limit_reached',
])
function retryAfter(h: Headers): number | undefined {
const v = h.get('retry-after')
if (v === null) return undefined
if (/^\d+$/.test(v.trim())) return Number(v)
const at = Date.parse(v)
return Number.isNaN(at) ? undefined : Math.max(0, (at - Date.now()) / 1000)
}
function classify(status: number, h: Headers, code: string | undefined, paid: boolean): Decision {
const wait = retryAfter(h)
const waitable = wait !== undefined && wait <= MAX_WAIT_S
switch (status) {
case 402: {
const offer = h.get('payment-required')
if (offer !== null && !paid) return { kind: 'pay', requirements: offer }
if (waitable) return { kind: 'wait', seconds: wait } // in-flight budget
return { kind: 'stop', reason: paid ? 'payment_failed' : code ?? 'unfunded' }
}
case 429:
if (code !== undefined && BILLING_CODES.has(code)) return { kind: 'stop', reason: code }
if (wait === undefined) return { kind: 'backoff' }
return waitable ? { kind: 'wait', seconds: wait } : { kind: 'failover' }
case 503:
case 529:
return waitable ? { kind: 'wait', seconds: wait } : { kind: 'failover' }
default:
return status >= 500 ? { kind: 'failover' } : { kind: 'stop', reason: `http_${status}` }
}
}
Three rules sit around it. First, turn off the SDK's own retries (maxRetries: 0) so that two retry loops do not stack. Second, after any paid attempt that fails, read the PAYMENT-RESPONSE header before signing again. If a settlement is recorded there, the next attempt is a second purchase. Third, reuse the same payment-identifier across retries of one logical call, and generate a new one only when the agent means to buy again. Anthropic's 400 for a self-set spend limit falls into the stop branch by default, which is correct. If you want to tell it apart from other 400s, the docs give the message prefix: You have reached your specified API usage limits.
What it means for LLM4Agents
We see this from both sides. Upstream, we are a client of providers that use the table above. The fallback chain already turns an upstream 429 or 5xx into a different model, and a failed link is never charged. Most of the provider disagreement therefore stops at our gateway. That is the value of a gateway: three vocabularies for money and throttling go in, and one should come out.
Downstream, we are a seller, and our own responses have the same gaps as the Bazaar. On 26 September our walk-up 402 carried a PAYMENT-REQUIRED header with a single exact option on eip155:8453: 10000 base units of USDC (one cent), maxTimeoutSeconds 300. It carried no RateLimit fields and no Retry-After. Our OpenAPI documents a 429 for chat completions at 600 requests per API key per minute. It mentions Retry-After, RateLimit and Idempotency zero times.
Our 402 also has two meanings. It is either a walk-up offer or insufficient_balance for a Bearer agent. The second case has the same structure as OpenRouter's in-flight budget. The reserve holds a worst-case amount while a call runs, so an agent running many calls in parallel can run out of balance temporarily, and the problem clears when the holds settle. Today nothing in the response tells that case apart from an empty wallet.
The risk is concrete. A developer who points a stock OpenAI or Anthropic TypeScript SDK at our walk-up endpoint through wrapFetchWithPayment will retry our 429s automatically. Each retry signs a new authorization. Whether a throttled attempt settles depends on where our limiter sits relative to settlement, and nothing in the response tells the agent which case it is in. The opportunity is equally concrete. The gateway that answers "pay, wait or stop" clearly, in headers that stock SDKs already read, has an advantage that is easy to explain.
Staying on the frontier
1. Publish the limits we already enforce, in the current syntax. Send RateLimit-Policy: "per-key";q=600;w=60 and the matching RateLimit on every response, including 402s and 429s, and Retry-After on every 429. Skip the legacy X-RateLimit-* headers, since the census shows how inconsistent they are. The change is small, and it would put us among the seven hosts in the census that use the current syntax.
2. Use the one signal every official SDK obeys. Send x-should-retry: false on refusals that will not clear by waiting: an empty wallet, or a provider billing error passed through. Send x-should-retry: true with Retry-After on refusals that will. All four clients check this header before their status rules. Agents using stock SDKs against our OpenAI-compatible endpoint would stop wasting retries with no code change on their side.
3. Split insufficient_balance in two. Return reason: holds_outstanding with a Retry-After estimated from the oldest open reserve, and reason: wallet_empty with no retry hint. OpenRouter has already documented this pattern, and agents are learning to handle it.
4. Stop the x402 handshake from halving throughput. Count only paid or authenticated requests against the per-key quota. Meter unpaid 402 challenges in a separate per-IP bucket. The four sellers that throttled their unpaid path show what happens otherwise.
5. Support payment-identifier on walk-up, without caching errors. Tie each id to a fingerprint of scheme, network, asset, amount, payTo and route. Replay only successful responses. Never store a 4xx or 5xx under an id.
6. Take the gaps upstream. Propose an x402 transport note that defines 429 and 503 with Retry-After, and requires sellers to reject throttled requests before settlement. Ask the HTTPAPI working group whether a tokens quota unit belongs in the registry, since every LLM provider limits by tokens, and point out the request/requests mismatch in -11. Until then, use a vendor-prefixed parameter to state our token limits.
Specifications will eventually cover all of this. In the meantime, every refusal an agent receives is a choice between paying, waiting and stopping. A seller that makes that choice obvious gets retries it can handle, instead of repeated signatures and stalled tasks.
A gateway that tells your agent when to pay and when to wait
OpenAI-compatible API, pay per call in USDC over x402 or from a funded balance, with automatic model fallback.
Register an agent