← Blog
September 24, 2026 · 16 min

Auditing ERC-8183: three ABIs and 81,095 agent jobs

ERC-8183 wants to be the escrow primitive for agent commerce: a client locks funds, a provider delivers, an evaluator decides. We read the three texts that define it and counted every job in the live contract it grew out of. The escrow logic works. The standard does not exist yet as one thing, and evaluation, its central idea, is almost never used.

Most agent payments today are pay-per-call. An agent signs an x402 authorization, a facilitator settles it, and the response comes back. That works when the service is a single HTTP request. It does not fit work that takes hours, produces a deliverable, and might be wrong. ERC-8183, titled "Agentic Commerce", targets that second case.

This post does three things. It compares the published EIP, its inline reference code, and the reference repository, function by function. It runs four Foundry probes against the reference implementation. And it reads the state of all 81,095 jobs in AgenticCommerceV3, the production contract on Base that the ERC-8183 discussion treats as its predecessor. All measurements are from 24 September 2026.

The job in one paragraph

ERC-8183 is a Draft created on 25 February 2026. It was merged into the ERCs repository on 5 March through PR #1581. It has four authors, including Davide Crapis, who also co-authored ERC-8004. The spec's one-line description is "Job escrow with evaluator attestation for agent commerce."

A job has three roles and six states. The client creates it and names an evaluator. The provider proposes a price. The client funds the escrow in a single ERC-20. The provider submits a bytes32 reference to the work. Only the evaluator can then call complete, which pays the provider, or reject, which refunds the client. If nobody acts before expiredAt, anyone can call claimRefund and the client gets everything back. The minimal happy path looks like this:

// client
jobId = createJob(provider, evaluator, expiredAt, "summarise 40 PDFs", hook);
// provider
setBudget(jobId, 20_000000, "");                // 20 USDC
// client
fund(jobId, 20_000000, "");                     // expectedBudget: anti-front-running
// provider
submit(jobId, keccak256(deliverable), "");   // Funded -> Submitted
// evaluator
complete(jobId, reasonHash, "");             // Submitted -> Completed, escrow released

That is five transactions from three parties, plus a token approval, to settle one job. An x402 exact payment is one signature and one settlement. The extra machinery buys one thing: someone other than the payer and the payee gets to decide whether the work was done.

Two optional pieces complete the picture. A per-job hook contract gets beforeAction and afterAction callbacks and can veto transitions. claimRefund is deliberately not hookable, so a bad hook cannot trap funds. And the spec recommends mapping job outcomes into ERC-8004 reputation, splitting the work into "ACP remains the payment and escrow layer; ERC-8004 is the identity and reputation layer."

One standard, three texts

The page on eips.ethereum.org has not changed since 13 March, when PR #1601 pasted a full reference contract into it. Four days later the reference repository, erc-8183/base-contracts, started to diverge. Its last commit is from 30 June. It carries its own copy of the spec in eip.md, and that copy has never been proposed back to the ERCs repository. The only open change to the published text is PR #1732, a two-line clarification on integer widths, waiting for an author review since 9 May.

The result is three texts that disagree on the function that moves money. We computed the selectors:

fund()                         selector    signature
EIP page, prose (w/ optParams) 0xd2e13f50  fund(uint256,uint256,bytes)
EIP page, reference code       0xe25ba707  fund(uint256,bytes)
base-contracts @ 142e669       0x1f989ec8  fund(uint256,address,uint256,bytes)
AgenticCommerceV3 on Base      0xd2e13f50  fund(uint256,uint256,bytes)

createJob()
EIP page, reference code       0x41528812  createJob(address,address,uint256,string,address)
base-contracts @ 142e669       0xbaf3ede2  createJob(address,address,uint48,string,address,uint256)
AgenticCommerceV3 on Base      0x41528812  createJob(address,address,uint256,string,address)

The published prose requires fund(jobId, expectedBudget, optParams?) and says it SHALL revert on a budget mismatch "(front-running protection)". The reference code on the same page has no expectedBudget at all. It only has fund(uint256 jobId, bytes optParams). In that code only the provider can call setBudget, so a provider can change the price between the client's decision and the client's transaction, and fund will pull whatever the budget is, up to the client's allowance. The repository added the check on 17 March. It later bound the token address as well, because the budget now carries its own token.

The event layer splits too. The repository changed expiredAt to uint48 on 6 May, which changes the JobCreated topic from 0xb0f0239b… to 0x834a72f1…. An indexer built for one never sees the other's jobs. PR #1732 asks to make uint256 widths normative, on the grounds that the reference implementation uses them. It was opened three days after the reference implementation stopped using them.

The published page also contradicts itself on hooks. Its data-encoding table says a setBudget hook receives abi.encode(amount, optParams). Its reference code sends abi.encode(msg.sender, amount, optParams). A hook written from the table reads the caller's address as the amount. The prose also lists setProvider as hookable, and its bidding-hook example depends on it, but the reference setProvider calls no hook. The repository's eip.md fixes the encoding table and now marks setProvider as not hookable. It also adds a note admitting that the bidding example assumes behavior the reference implementation does not have.

We saw the same pattern in ERC-7824, where the draft and the deployed hub were two different protocols. Here it is worse, because the public discussion is aimed at the stale text. In August, posts in the 400-post Ethereum Magicians thread were still proposing to add a submittedAt field to the Job struct. The repository had added one on 24 March.

What the working draft adds

Judged on its own terms, the repository's draft is the better design. Its 76 tests pass. The changes that matter for agents:

Most of this landed in May and June. It was merged on 30 June from a branch named okx/feature/meta-transactions-and-claims (PR #24). An open PR #26 proposes making the entire lifecycle hookable.

The production contract: 81,095 jobs

The contract that actually settles money is AgenticCommerceV3 at 0x238E…32E0 on Base. In the thread it is identified as Virtuals' ACP deployment and called ERC-8183's predecessor. It is a UUPS proxy deployed on 8 April 2026 (block 44,427,013). Its implementation, 0x8e86…77bc, was verified on Blockscout on 6 May, and the proxy has never been upgraded. We read its configuration directly:

$ ACP=0x238E541BfefD82238730D00a2208E5497F1832E0
$ cast call $ACP 'jobCounter()(uint256)'     # 81095
$ cast call $ACP 'paymentToken()(address)'   # 0x8335…2913 (USDC)
$ cast call $ACP 'platformFeeBP()(uint256)'  # 500 -> 5% to treasury
$ cast call $ACP 'evaluatorFeeBP()(uint256)' # 500 -> 5% to evaluator, on complete only
$ cast code 0xe220329659d41b2a9f26e83816b424bdacf62567  # 0x -> admin is a plain EOA

V3 is a fourth variant. Its fund matches the published prose and its createJob matches the published code. It departs from all three texts in one important way: it accepts evaluator = address(0). ERC-8183 says createJob SHALL revert in that case. In V3 a zero evaluator means submit auto-completes and pays the provider on the spot. The grace period is 15 minutes, not one hour.

We read all 81,095 jobs through Multicall3, 250 jobs(uint256) calls per eth_call, between blocks 51,725,866 and 51,726,048 (09:11–09:17 UTC). The same data had already been indexed by MarselSultanov, who posted the results in post #393 on 19 August, with a reproducible dataset of 62,953 jobs. For those same job IDs our read matches exactly: 72.50% with no evaluator, 27.48% with the client as its own evaluator, and ten jobs, 0.02%, with an independent one. The full picture today:

jobs                        81,095
  Open (never funded)       57,592   71.02%   all but one past expiry
  Completed                 15,724   19.39%
  Expired                    6,026    7.43%
  Rejected                   1,750    2.16%
  Funded                         3            all past expiry
evaluator = none            46,052   56.79%
evaluator = client          35,006   43.17%
independent evaluator           37    0.05%   8 addresses
completed gross volume    2,097.64 USDC
  self-evaluated          1,994.62 USDC   95.09%
  independent                 0.97 USDC    0.05%
median paid job               0.01 USDC   (p90 0.05, max 10.00)

Four things stand out.

Evaluation is not happening. In the 18,142 jobs created after the forensic snapshot, self-evaluation went from 27.48% to 97.60% and no-evaluator jobs fell to 2.25%. The market moved from "no judge" to "I am the judge", which gives the same assurance and costs one more transaction. Across the whole history, eight independent evaluator addresses handled 37 jobs and released 97 cents.

Most of the volume is two addresses paying each other. 0xe09f…8584 and 0x44cc…6664 completed 3,950 jobs between them, 2,014 in one direction and 1,936 in the other. Each job was self-evaluated. Together they moved 1,616.84 USDC, or 77.08% of all completed volume, mostly in 0.05 USDC jobs with IDs 60,245 to 70,984. The protocol cannot tell this apart from real commerce. A reputation system that counts completions counts these too. Without the pair, the rest of the market completed 11,774 jobs worth 480.80 USDC in five and a half months.

Most jobs are empty shells. One client, 0x22f7…1491, created 43,858 jobs (54.08% of the total), all without an evaluator, and 38,831 of them were never funded. Overall, 71% of jobs never reached Funded.

The escrow balance does not reconcile. The contract holds 22.90 USDC. The only live escrows are three Funded, expired, self-evaluated jobs worth 3.00 USDC, which anyone can refund today. The other 19.90 USDC belongs to no live job. Every lifecycle path in V3 moves exactly a job's budget, so the extra must have arrived by direct transfer. The only function that can move it is emergencyWithdraw, which only the admin can call, and only while the contract is paused.

The evaluator problem, measured

The spec's rationale says that once work is submitted "the client cannot pull funds back unilaterally, so the provider is protected after starting work." The same spec allows evaluator = client. blockbird pointed out on 12 August what follows. The client names itself evaluator, receives the deliverable at submit, does nothing, and after expiry anyone can refund the client in full. We reproduced it against the reference repository:

function test_probe1_silentSelfEvaluatorRefundsClient() public {
    (uint256 jobId, uint48 expiry) = _selfEvaluatedSubmittedJob(); // provider submitted
    vm.warp(uint256(expiry) + 1 hours);                           // expiredAt + grace
    vm.prank(thirdParty);
    core.claimRefund(jobId);
    assertEq(usdc.balanceOf(client), BUDGET);                       // full refund
    assertEq(usdc.balanceOf(provider), 0);                          // work delivered, unpaid
}   // [PASS]

The grace period does not fix this. It was added to protect an evaluator who is mid-review from a third party racing claimRefund. It does nothing against an evaluator who chooses silence. It only makes the client wait one more hour, or 15 minutes on V3. Issue #1931 made the incentive explicit: complete costs gas, reject costs gas, and "saying nothing is free and refunds them in full." It was closed on 6 September with no spec change. On 5 September the thread asked the authors a yes-or-no question about separating the evaluation window from expiredAt. As of today nobody has posted since.

V3's fee schedule adds a second bias. The 5% evaluator fee is paid on complete and never on reject. An independent evaluator earns money only by approving. A self-evaluating client that completes gets its own 5% back, and one that stays silent gets 100% back.

None of this is hidden. The spec states that "a malicious evaluator can complete or reject arbitrarily" and recommends reputation or staking. The data shows that the recommended configuration, a third party with something at stake, accounts for 0.05% of real jobs.

Who can move the escrow

The spec is careful about hooks. They cannot block claimRefund, and "implementations MUST NOT allow hooks to modify core escrow state directly." About the operator it says almost nothing. The working draft mentions admin tooling only to ask that it emit events. Both the repository and V3 give the admin more power over escrowed funds than any hook has. We checked with three more probes:

[PASS] test_probe2_feeChangedBetweenFundAndComplete()   // admin sets 100% fee after fund; provider receives 0
[PASS] test_probe3_pauseBlocksRefundAndAdminDrains()    // claimRefund reverts while paused; emergencyWithdraw empties escrow
[PASS] test_probe4_repriceBeforeFundReverts()           // expectedBudget stops provider re-pricing (the fix works)

Fees are global and read at payout time, so the provider's net amount is not fixed when the client funds. claimRefund carries whenNotPaused, so the "permissionless safety mechanism" stops working the moment the admin pauses. While paused, emergencyWithdraw(token, to, amount) can send any amount of any token anywhere, with no per-job accounting. batchDetachHook can remove the policy a client attached to a live job. The contract is also UUPS-upgradeable by the default admin. On V3 both admin roles have belonged to a single EOA, 0xe220…2567, since 4 May, when the deployer handed them over and revoked its own. There is no multisig and no timelock.

The amount at risk on V3 today is small: 22.90 USDC. The design is the issue. Compare the escrow behind x402's auth-capture scheme, which is a non-upgradeable singleton whose release paths are fixed at deployment. An ERC-8183 client is trusting the evaluator it picked and also the contract's admin, whom it did not pick.

Where x402 fits, and where it does not

Both versions of the spec claim "x402 compatibility": an agent signs an intent and a facilitator executes it. The published version recommends an ERC-2771 trusted forwarder. The repository replaced that with per-call EIP-712 signatures, which is the right call because no privileged forwarder enters the trust base. The x402 repository itself had zero references to ERC-8183 or AgenticCommerce when we searched it on 24 September. No x402 scheme wraps a job.

The difference in shape matters more than the missing code. fund pulls tokens with transferFrom. An agent whose only payment skill is signing an EIP-3009 transferWithAuthorization cannot fund a job with it. It needs an EIP-2612 permit (USDC supports it) plus a FundAuthorization, two signatures that a relayer submits together. The reference test suite's median execution gas for the five-call happy path adds up to about 357,000 (createJob 134,647, setBudget 63,574, fund 63,324, submit 25,042, complete 70,600), before the base cost of each transaction.

Against that overhead, the median paid job on V3 was one cent. The market is using a job escrow for payments the size of an API call, and in 99.85% of recent jobs no independent party evaluates anything, which was the reason to use an escrow. For that traffic, x402 is the better tool. Claim settlement is the interesting overlap. Cumulative, monotone settlement is the same idea as x402's upto scheme and signed vouchers, except that here each step is an on-chain transaction instead of a signature.

What it means for LLM4Agents

ERC-8183 does not compete with our payment path. A chat completion is priced, delivered and settled in one round trip, and x402 fits that exactly. ERC-8183 sits one layer up. It covers work an agent contracts for rather than calls: batch inference over a corpus, a research report, a fine-tuning run. That layer is where our customers' agents will start hiring each other, and our gateway could appear in it in three roles.

As a provider, long-running jobs on the gateway could be sold through a job escrow instead of prepaid credit. The census is a warning about the terms. If the client is its own evaluator and there is no bond, every delivered job is a free option for the client. Evaluator silence is currently the provider's risk.

As an evaluator, an LLM-as-judge service is a natural product for an inference gateway, and it is the one role the market is missing. Only eight independent evaluator addresses have ever been named on a V3 job. The thread's emerging consensus fits it well: reason should be a hash of canonical, recomputable evidence, not an opaque verdict.

As a consumer of reputation, we route across providers and will want trust signals. ERC-8183 outcomes fed into ERC-8004 would be one such signal. On today's data that signal is 95% self-evaluated by volume and 77% one circular pair. We showed in our ERC-8004 empirical audit how easily reputation writes are gamed. Completions are cheaper to fake than reviews.

Staying on the frontier

Six steps, in order.

1. Keep inference on x402. No per-call traffic goes through a job escrow. We will use ERC-8183 only for work with a deliverable, a duration, and a value above a threshold that justifies five transactions.

2. Build against the repository draft, behind a selector check. ERC8183WithAuthorization is the design worth targeting. Before interacting with any deployed contract, our adapter will read its bytecode for known fund selectors (0x1f989ec8, 0xd2e13f50, 0xe25ba707) and refuse the one without expectedBudget. Three texts means detection, not assumption.

3. Write a provider acceptance policy before taking a job. Reject evaluator == client above a small budget. Require expiredAt to leave room for the work plus the evaluation. Check who holds the admin role on the escrow contract, and whether it is an EOA. Read the fee settings, remembering they can change before payout. Where the contract supports claim settlement, file a claim per milestone so that one silent evaluator costs at most one milestone.

4. Prototype an evaluator with recomputable output. Start with deterministic checks an LLM can orchestrate, such as schema validation and test execution, and add model-graded rubrics later. Commit a hash of canonical evidence as reason and publish the evidence, so any third party can recompute the verdict. Charge a flat fee that does not depend on the outcome, the reverse of V3's pay-on-complete.

5. Discount ERC-8183 outcomes in any trust score. Self-evaluated completions count as zero. Pairs of addresses that pay each other get flagged. Only completions by an independent evaluator with a history count, and today that means almost nothing counts.

6. Push the spec, not just the code. The fastest fix for the ecosystem is to send the repository's eip.md upstream, so the Magicians thread stops arguing about March's text. We will comment on the thread asking for that, and for a separate evaluation window or a provider-side fallback when a Submitted job expires, as issue #1931 proposed.

ERC-8183 gets the escrow mechanics right, and its working draft is thoughtful. The hard part is the evaluator, and the protocol cannot supply one. Until independent evaluation is cheaper than self-evaluation, "trustless agent commerce" on this primitive means a client paying itself back.

Pay per call now, escrow when the work needs it

An OpenAI-compatible gateway where every request settles with a signed stablecoin payment. No accounts, no prepaid credit, no evaluator to trust.

Register an agent