← Blog
September 27, 2026 · 22 min

AAuth Budgets: a spending cap for agents that moves no money

For the first time, an IETF-track draft names metered LLM inference as the case it was written for. AAuth Budgets gives an agent a spending cap. The person's server decides the size, the resource enforces it, and no money moves. We read the new base protocol, the unsubmitted Budgets draft and its closest relative, probed the one production deployment it names, and tested its streaming cost signal against 16 HTTP client setups. The design is careful. We found four gaps between it and practice.

The question it answers is simple to state. An agent calls an inference endpoint using a person's or a company's account. What stops it from spending ten thousand dollars when the job was worth ten? Today the answers sit outside any protocol: a provider dashboard, a card limit, a monthly bill. Our own audit of x402 spend controls found the same thing on the stablecoin side. The buyer SDKs cap spending locally, and nothing on the wire says how much an agent is allowed to spend.

Our sources, all read on 27 September 2026: draft-hardt-oauth-aauth-protocol-11 (posted 25 September, 164 pages); draft-hardt-httpbis-signature-key-09 (13 September); the editor's copies of AAuth Budgets and AAuth R3 in the dickhardt/AAuth repository at commit a200889 (25 September); the TPX v0.3 specification; the regent-httpsig repository at release 0.6.0; and the @aauth packages on npm. We sent unauthenticated GET and POST requests to two public deployments. All client tests ran against a local server. We signed nothing and paid nothing.

AAuth in one pass

AAuth is Dick Hardt's proposal for agent authorization. Hardt edited RFC 6749, the OAuth 2.0 framework. AAuth starts from one premise: every agent instance has its own cryptographic identity. An agent gets an identifier of the form aauth:local@domain from its agent provider. It generates a key pair, with Ed25519 recommended. The provider issues an agent token (typ: aa-agent+jwt) that binds the key to the identifier. That token should live no more than 24 hours.

Every request the agent makes is signed with an RFC 9421 HTTP Message Signature. The token rides in a Signature-Key header, defined in a companion draft that Hardt co-authors with Cloudflare's Thibault Meunier. No credential is a bearer token and nothing needs pre-registration. In the draft's words, "the first API call to a resource is the registration."

From there the protocol climbs a ladder. Revision -11 defines five resource access modes:

Mode                  Resource learns             Parties
agent identity        which agent                 agent, resource
resource-managed      which person (own login)    agent, resource
person identity       which person (via PS)       agent, resource, PS
PS authorization      person + consented scope    agent, resource, PS
federated             person + policy verdict     agent, resource, PS, AS

// PS = person server (represents the human or organization)
// AS = access server (the resource's policy engine)

A resource says what it needs through one response header, AAuth-Requirement. It can appear on a 401, a 202 or a 402, with values such as agent-token, person-token, auth-token and interaction. Person tokens and auth tokens live at most one hour. Resource tokens, which a resource hands the agent to carry to its person server, should live no more than five minutes.

Revision -11 is a large rewrite. It adds the person token and the person-identity mode. It removes the act claim. It replaces the mission object with a hash, mission_s256. And in the three modes that involve a person server, "no token a resource reads carries an agent identifier." The resource learns who the person is, not which agent is acting.

Status matters here. The draft is an individual submission with no working group. It has had twelve revisions under its current name since 29 April 2026, after three under an earlier name that began on 2 April. Its implementation section lists four implementations, all marked "exploratory": TypeScript from Hellō, a .NET SDK, and a Python demo plus a Keycloak extension from Christian Posta. The GitHub repository was created in October 2025 and had 123 stars and 37 open issues and pull requests when we looked.

Name collision — a different protocol also calls itself AAuth. "AAuth - Agentic Authorization OAuth 2.1 Extension" (draft-patwhite-aauth-00, May 2025, and draft-rosenberg-oauth-aauth-01, April 2026) is an OAuth grant type, unrelated to Hardt's design. Search results mix the two. Everything below refers to draft-hardt-oauth-aauth-protocol.

A budget is an authorization, not a payment

AAuth Budgets is an extension that has not yet been submitted to the IETF. It was added to the repository on 11 August 2026, marked "Exploratory Draft", and last changed on 20 September. Its abstract is one sentence long: a budget is "a ceiling on what an agent may consume at one resource, denominated in a unit the resource declares, carried as a claim in the auth token, and enforced by the resource."

The introduction states the motivation plainly: "An agent harness calling a model inference endpoint on a person's account can spend without bound." A scope such as inference.completions is either granted or not. It says nothing about whether the agent may spend ten cents or ten thousand dollars. The draft then says: "Metered inference is the initiating use case."

The key design choice is in the non-goals. "Not payment or settlement. No funds move. 402 Payment Required and the resource's commercial arrangement with the person are untouched." A budget bounds what an agent may consume. How the resource is paid is a separate question.

The budget object has three members and appears in the same places a scope does:

{
  "scope":  "inference.completions",
  "budget": { "amount": 5000000, "unit": "USD", "decimals": 6 }
}
// 5000000 / 10^6 = $5.00. An integer in a declared scale, never a float.

It passes through a narrowing chain. The agent may ask for an amount. The resource fixes the unit and the scale, and may lower the amount. The person server may lower it again. In four-party access, the access server may lower it once more. After the resource speaks, nobody may change the unit or the decimals. Asking for more than the resource allows is not an error. The resource simply narrows the amount.

One rule in the four-party path deserves attention. If the resource token carries a budget and the person server's request to the access server omits one, the access server "MUST NOT issue a budget claim." An earlier copy read omission as a grant of the full offer. That made a person server that had never implemented Budgets look like one granting the maximum. The fix is right.

The allocation is not the ceiling

The draft separates two numbers. The person's real ceiling is private state on the person server: "no claim carries it, and it may not be shared with the agent." What an auth token carries is an allocation drawn against that ceiling. The token expires within an hour, or its budget runs out, whichever comes first. Either way the agent must come back to its person server for more.

That return trip is where supervision happens. The draft lists six responses available to the person server: grant the resource's offer, grant less, ask the agent why it needs more, ask the person, decline with a suggested_budget, or end the work. It also explains why the agent is never told its cumulative spend. An agent that can see its spend across allocations can work out the ceiling by subtraction.

The trade-off is written down too. An agent holding several tokens at the same resource holds several budgets. A person server that issues n concurrent tokens of X each "has authorized up to nX for as long as an hour." The resource enforces each token on its own. Only the person server can hold the total.

Where payment goes instead

If budgets move no money, something else must. AAuth's answer is to put payment next to authorization on the same response, as separate fields. Section 11.6 of -11 says a 402 carrying AAuth-Requirement "says authorization and payment are both required; the payment requirement is conveyed separately, by x402 or the Payment scheme." It adds: "AAuth never conveys its own requirements via WWW-Authenticate."

That composes cleanly with x402 v2, whose challenge rides in its own PAYMENT-REQUIRED header, as the transport spec confirms. A seller can send a 402 carrying both headers. An agent that speaks both protocols reads both. One that speaks only x402 pays, retries without an auth token and gets challenged again. That wastes a round trip and a signature, but it does not break anything.

Two details show where the authors' attention went. First, every 402 example in -11 uses WWW-Authenticate: Payment, the HTTP scheme behind MPP, the Stripe and Tempo protocol. None of them shows an x402 header. Second, the only 402 flow that -11 specifies step by step is at the access server's token endpoint. There, payment establishes "a billing relationship" between person server and access server, which "the PS caches per AS." That is closer to opening an account than to paying per call.

The Budgets draft keeps the two conditions apart on purpose. An exhausted budget is answered with 401 and requirement=auth-token, meaning "go back to your person server." A 402 means the resource needs money, not permission. And 429 is not used at all, because a budget is about the cost of the next request, not the rate of requests.

TPX draws the line in a different place. TPX is the OAuth 2.0 profile for metered inference grants that Budgets cites as its OAuth counterpart. It is deployed at tokenpony.dev. It returns 402 with error.code: "budget_exhausted" when a grant is spent, and 402 with balance_exhausted when the person's prepaid balance is empty. Two drafts built for the same kind of meter disagree on the status code for "your cap is spent." An x402 client that sees a TPX 402 will look for a payment challenge that is not there. We described that confusion in our 402-versus-429 census, and here it is again.

The meter: reserve, commit, release

The enforcement section will look familiar to anyone who has read our reserve-proxy-settle internals. The budget is a hard cap: "A resource MUST NOT let metered consumption exceed the granted amount." Output is only metered after it is generated, so the resource must bound each request before serving it. The draft names the pattern: reserve the maximum, serve, commit the actual cost, release the difference. It also requires that "an operation with no finite cost bound MUST be given one, be truncated when the remainder is consumed, or be refused." For an LLM call, that bound is max_tokens.

The agent sees the result in a response header, AAuth-Budget. It carries remaining (required), cost for this request, reserved for what is still held, and required on a refusal. The last one tells the agent what the refused request needed:

HTTP/1.1 401 Unauthorized
AAuth-Requirement: requirement=auth-token;
    resource-token="eyJ..."; reason=insufficient-budget
AAuth-Budget: remaining=150000, required=400000, unit="USD", decimals=6

// The agent has $0.15 left; this call needed up to $0.40.
// It can lower max_tokens and retry on the same token, without asking its PS.

Streaming is the hard case. Headers are sent before the body, and the cost of a streamed completion is only known when the stream ends. The draft's answer is to send reserved in the header and cost in an HTTP trailer, a header sent after the body:

HTTP/1.1 200 OK
Content-Type: text/event-stream
Trailer: AAuth-Budget
AAuth-Budget: remaining=1568800, reserved=431200, unit="USD", decimals=6

   ...stream...

AAuth-Budget: cost=221200
// balance after = remaining + reserved - cost = 1,778,800 ($1.7788)

The draft expects trailers to fail sometimes. A trailer may only add members, never repeat them, so it does not matter whether a client merges trailers into headers or drops them. And "a recipient MUST NOT treat the trailer as necessary." If there is no trailer, the agent treats the request as having cost the full reserved amount until the next response corrects it.

The draft also dismisses the main objection to trailers. It says dropped trailers are "largely a browser property — fetch does not surface trailers — and the traffic this document governs is an agent calling a resource, which is server-to-server." We tested that claim.

Finding 1: most agent clients never see the trailer

We built a local server on Node.js 22.22.2 that serves an OpenAI-shaped SSE stream. It sends the header above, the trailer AAuth-Budget: cost=221200, and a final chunk with usage. It speaks both HTTP/1.1 (chunked) and cleartext HTTP/2. Then we read the response with each client and checked whether the trailer value reached the calling code.

Client                                  Trailer visible?
curl 7.81, HTTP/1.1 and HTTP/2          yes (printed with -i)
Go 1.27.1 net/http, HTTP/1.1 and h2c    yes (resp.Trailer after EOF)
Node node:http, node:http2              yes (res.trailers / 'trailers' event)
undici 8.11.2 request()                 yes (.trailers after body)

Node 22 global fetch                    no
undici 8.11.2 fetch                     no
openai (npm) 7.23.0                     no
requests 2.34.2                         no
httpx 0.28.1, HTTP/1.1 and HTTP/2       no
httpx2 2.13.1                           no
aiohttp 3.14.3                          no
openai (PyPI) 3.19.2                    no

// 16 configurations: 7 see the trailer, 9 do not.
// Every "no" still received the header and the usage chunk.

The split is clear. Low-level transports expose trailers. The high-level clients that agent code actually uses do not. That includes Node's built-in fetch on the server, every mainstream Python client, and both official OpenAI SDKs. The npm SDK is built on fetch. The current PyPI release is built on httpx2, which exposes no trailer attribute either. So the objection is not a browser quirk. It describes the Python and JavaScript stack most agents are written in.

Both OpenAI SDKs did receive the usage in the final SSE chunk. That application-layer channel works everywhere. The draft notes it and says a resource that can send both "SHOULD send both." In practice, the chunk is what the agent will read.

The only production implementation named in the draft reached the same conclusion from the server side. The streaming path of regent-httpsig is commented: "a streamed response's actual cost is known only when the stream ends, and this runtime sends no trailers." It uses the draft's fallback, which the draft calls the cost-omitted mode. That puts all the weight on the fallback formula.

Finding 2: the fallback formula counts the next call twice

When cost never arrives, the draft tells the agent to recover it from the next response:

cost = previous remaining + reserved - current remaining

The reasoning given is that remaining is always net of reservations. That is true. But it also means the next response's remaining is net of the next request's own charge or hold. The formula as written puts that draw on the previous request.

We checked this against Regent's own meter (InMemoryMeter in regent-httpsig 0.6.0, run under Python 3.12). We gave a token a $2.00 budget and made two serial requests. Request A was streamed, with a $0.4312 reservation and an actual cost of $0.2212.

// B is a non-streamed call costing $0.05
A header : remaining=1568800, reserved=431200
B header : cost=50000, remaining=1728800
draft    : 1568800 + 431200 - 1728800 = 271200   // actual 221200, off by 50000

// B is streamed too, reserving $0.4312
A header : remaining=1568800, reserved=431200
B header : remaining=1347600, reserved=431200
draft    : 1568800 + 431200 - 1347600 = 652400   // actual 221200, off by 431200

An agent that streams every call, which is the normal case for inference, reads almost three times the real cost. The error is always in the conservative direction, so the budget cannot overspend. But the per-call figure an agent logs, bills to a sub-task or shows its operator is wrong. The fix is one extra term: subtract the next response's own cost, or its reserved when cost is omitted. Both values are in that header. With the fix, both runs give exactly 221,200.

A second consequence follows. The last streamed call on a token has no next response on that token. The agent can never recover its cost. Only the person server can, later, from the usage endpoint.

Finding 3: the price is in the body, and the body is not signed

AAuth's signing profile, in section 11.3.3.1, requires an agent to cover four components: @method, @authority, @path and signature-key. It requires content-digest and content-type only on requests with a body to a person server, an access server or a revocation endpoint. A resource can ask for more coverage through a metadata field, additional_signature_components.

For a metered inference call, everything that sets the price is in the body: model, max_tokens, the messages. The Budgets draft computes the reservation from exactly that bound, "a maximum output length." With the default profile, the signature proves which agent sent a POST to /v1/chat/completions. It does not prove what that agent asked for.

TLS protects the body on the wire, so the exposure is at any hop that terminates TLS and forwards the request: a CDN, a corporate proxy, a logging layer in the agent's own stack. Such a hop can change the body and keep a valid signature. A captured request can also be replayed with a different body during the 60-second signature window, unless the resource runs the replay cache that the profile makes optional. Either way the budget pays for the request that arrived, not the one the agent signed.

The fix already exists and costs one line. The base draft's own example of resource metadata ends with "additional_signature_components": ["content-type", "content-digest"]. But the Budgets draft's example metadata for inference.example leaves it out. Its example request to the authorization endpoint covers only the four default components. So does the metadata of the production deployment below. The draft already requires content-digest on the usage endpoint and on signed usage responses. It should require the same for the metered request itself.

Finding 4: the reference deployment trails the spec

The Budgets draft says Regent Protocol "is in production at get4agent.com." Its middleware, regent-httpsig, is Apache-2.0 on PyPI, with eight releases between 17 August and 8 September. We fetched get4agent.com's metadata at 09:04 UTC on 27 September:

GET /.well-known/aauth-resource.json
{
  "resource": "https://get4agent.com",
  "jwks_uri": "https://get4agent.com/.well-known/jwks.json",
  "budget_units": [{ "unit": "USD", "decimals": 2, "max": 100000,
      "description": "Marketplace calls, prepaid balance minor units." }],
  "usage_endpoint": "https://get4agent.com/v1/budget/usage",
  "revocation_endpoint": "https://get4agent.com/v1/budget/revoke"
}

GET /.well-known/jwks.json
{ "keys": [{ "kty": "OKP", "crv": "Ed25519", "alg": "EdDSA", ... }] }

Under -11, a verifier must reject this metadata. Section 11.2 says a document with no issuer "MUST be rejected with issuer_missing." Here the field is named resource, the RFC 9728 spelling that AAuth explicitly chose not to use. The key's alg is the polymorphic EdDSA, which section 11.3.1 says "MUST NOT be used" since revision -10 adopted RFC 9864's Ed25519. For comparison, the AAuth project's own test resource at whoami.aauth.dev publishes issuer, access_mode and an Ed25519 key, as the draft requires.

The library is behind the draft in the same way. Its strict-algorithm switch, require_fully_specified_algs, defaults to off and is described as "a transition affordance for the -10 ecosystem." Version 0.5.1 added a challenge that carries a final consumption record for an expired auth token. On 13 September, the Budgets draft removed that challenge because "the base protocol does not permit" a resource to accept an expired token. Release 0.6.0, from 8 September and still the latest, ships it.

None of this is surprising for an exploratory draft that changes every week. The deployment is also doing some things right. Its usage endpoint refused our unsigned POST with a message asking for a jwks_uri signature "covering content-type + content-digest." Its decimals of 2 is reasonable for marketplace calls priced in cents. For inference, the draft argues for 6: a structured-field decimal cannot represent $0.000015, the price of a single token at $15 per million, and would serialize it as 0.000.

The agent side is further behind. The Hellō JavaScript packages we inspected (@aauth/agent 4.1.1, @aauth/fetch 4.0.0, @aauth/protocol 2.0.0 and @aauth/resource 3.0.0) read the R3 annotation x-aauth-budget from an OpenAPI document. None of them parses the AAuth-Budget header or handles a 402.

Smaller details that matter

The number format is chosen with care. Amounts are integers because a JSON number is only exact up to 253 and a structured-field integer is limited to 15 digits. The cap is therefore 999,999,999,999,999 units, about $1 billion at six decimals. The draft notes that "if assets requiring 18 decimal places come into scope, both carriers break" and amounts would need to be strings, as in x402. A USD budget in micro-dollars fits easily. An amount counted in an 18-decimal token does not.

The person server gets its own view. A metered resource may publish a usage_endpoint. The person server or access server queries it with a signed POST. The response gives calendar counters (day, week, month, year, all time), always in UTC, a scheme borrowed from Stripe Issuing. It can also give per-key totals for the signing keys the person server names. Signing the response is recommended, so the resource cannot tell the person server one number and the biller another.

Unknown spend counts as fully spent. If an allocation expires and nobody reports its consumption, the issuer "MUST account for the allocation as fully consumed" until a figure arrives. This is the same conservative rule applied to the agent after a dropped stream.

MCP tools can declare that they draw budget. R3, the companion draft for describing operations, adds a boolean budget annotation. For MCP tools it lives in the tool's _meta as aauth.dev/budget. For OpenAPI it is x-aauth-budget. An agent reading the tool list learns which calls will need a funded auth token before it makes the first one.

The balance is public to every hop. AAuth-Budget is unsigned and travels in the clear above TLS. Intermediaries "MUST NOT add, alter, or remove" it, but they can read it. The draft limits the exposure by never sending a cumulative figure. Each response reveals one call's price and what is left of one allocation.

What it means for LLM4Agents

AAuth Budgets describes our product from the outside. We sell metered inference to autonomous agents. We reserve the worst case, stream the answer and settle the actual cost. We already report cost and remaining balance in response headers on non-streamed calls. The draft proposes a standard name and format for all of that, and adds one thing we do not have: a third party, the person's server, deciding how much each agent may spend.

It also separates three things that our walk-up mode combines. In x402 walk-up, the payment is the authorization: the agent signs an EIP-3009 transfer and the call goes through. AAuth splits the work into identity (a signed agent token instead of an API key), permission to spend (an allocation in an auth token, sized by the person server), and payment (a 402, carried by x402 or MPP). For a single agent paying its own way, x402 alone is simpler and should stay that way. For a company running many agents on one account, the split is what a finance team will ask for. The spend limit belongs to the company's server, not to each agent's wallet.

The draft also puts one deployment explicitly out of scope: "AP-bundled inference," where the agent's vendor pays for its inference. "This is the common deployment today." The case the draft is built for is an agent spending against the person's own inference account. That is a gateway like ours, where the account holder is not the agent. If person servers become common, a gateway that offers only API keys and prepaid balances will be one that enterprise buyers cannot supervise.

The findings shape how we should implement it. The trailer cannot be our main cost channel, because the SDKs our users run cannot read it. Our reservations already come from max_tokens, so an unsigned body would let a proxy change the price. And the formula agents will copy from the draft is wrong by one term. If we get these right, agents on our gateway get correct numbers from the first call.

The main risk is timing. The base protocol changes shape every few weeks. The Budgets draft is not yet an IETF document, and the one production deployment already fails the current revision's metadata checks. Building the whole ladder now would mean rebuilding it. Building the part that matches what we already do is cheap.

Staying on the frontier

Six steps, in order.

1. Accept AAuth agent tokens as an alternative to API keys. This is agent-identity mode, the lowest step. It needs no person server and no access server. Verify Ed25519 and ES256 only, reject EdDSA, enforce the 60-second window, and keep the replay cache. Publish /.well-known/aauth-resource.json with issuer, access_mode and additional_signature_components: ["content-type", "content-digest"], so every POST to the gateway is signed over its body. An agent could then register by making its first signed call.

2. Implement Budgets on /v1/chat/completions, resource side only. Declare budget_units as USD with six decimals. Key reservations by auth-token jti on top of the reserve-settle ledger we already run. Return 401 with reason=insufficient-budget and required when a reservation does not fit. Our reserve formula already produces that number.

3. Report cost in three places. Put reserved in the AAuth-Budget header of every streamed response. Put cost in the final usage chunk, which both OpenAI SDKs read. Also send the trailer for the Go and curl clients that can use it. Document the corrected recovery formula in our client examples.

4. Keep x402 as the payment rail, and state which refusal means what. Use 401 with requirement=auth-token when an allocation is spent. Keep 402 with PAYMENT-REQUIRED for "pay per call or fund the balance." When a caller signed with AAuth, add AAuth-Requirement to the same 402. Publish this mapping, since TPX and AAuth already disagree.

5. Expose a signed usage endpoint. Serve UTC calendar counters and per-key totals from the transaction ledger we already keep. That is what lets a company's person server supervise our agents' spend without scraping dashboards.

6. Send the data upstream and watch the drafts. Our trailer results and the one-term formula fix belong in the dickhardt/AAuth issue tracker, along with a proposal to require content-digest on metered requests. Then watch for three events: the first datatracker submission of draft-hardt-aauth-budgets-00, the next base revision, and whether TPX moves its budget refusal off 402.

x402 answered the question "how does an agent pay?" AAuth Budgets is the first serious attempt to answer "how much may this agent spend, and who decides?" The two fit together at the 402. The gateway that implements both, with the details right, will be the one companies trust to run their agents.

Metered inference with a meter your agent can read

OpenAI-compatible API with reserve-then-settle billing, per-call x402 payments in USDC, and automatic model fallback.

Register an agent