← Blog
September 16, 2026 · 14 min

TEE-attested agent wallets: what a TDX quote actually proves

The pitch for a TEE agent wallet is one sentence: the key is born inside the enclave and nobody, not even the operator, can take it out. The specification says something narrower, and the gap is where the money is.

An autonomous agent that pays for its own inference needs a private key. Where that key lives is the whole security question. A key in an environment variable is a key the host can read. A key in a custodial API is a key somebody else can move. The third answer, the one that has quietly become the default in the agent stack, is to derive the key inside a confidential VM and publish a hardware attestation that says so.

That answer is now shipping. dstack runs Docker Compose workloads inside Intel TDX confidential VMs and hands each application a derived key plus a TDX quote. Phala ships an ERC-8004 TEE agent template on top of it. Oasis published the same pattern for x402 payments on ROFL. The marketing everywhere is identical: verifiable agents, keys that cannot be stolen.

We read the normative documents instead of the landing pages: the dstack guest API v1 spec, the TDX attestation guide, the encrypted-env spec, the security model, and the Automata on-chain DCAP verifier. Below is what an attested agent wallet actually proves, what it costs to check, and the one substitution that survives every check in the list.

The key is derived, not generated

The first correction is that a dstack application key is not random. It is derived, deterministically, from an application root key held by the KMS. The guest API v1 specification — introduced in dstack 0.6.0 — pins the bytes:

// dstack.guest.v1 key derivation
salt = "dstack-guest-v1"                // 15 bytes ASCII
IKM  = app root secp256k1 private key   // 32 bytes
info = LP("dstack-guest-v1-key") || LP(algorithm) || LP(domain)
L    = 32

key = HKDF-SHA256(salt, IKM, info, L)   // RFC 5869

LP(x) is uint32_be(len(x)) || x. The length prefix is not decoration. The spec is explicit that a domain is arbitrary caller-chosen bytes, so "any delimiter it could also contain would let two different (domain, algorithm) pairs encode to one byte string and share a key." Joining with : or / is not sufficient, and v1 does not do it — a correction over v0, where the path went into the KDF raw and the same 32 bytes served both secp256k1 and ed25519.

Two properties matter for a wallet. First, derivation is flat: "Two domains yield unrelated keys. a/b is an opaque string that happens to contain a slash, not a child of a." There is no BIP-32 hierarchy, so there is no xpub to hand an auditor and no way to enumerate an agent's future addresses. Second, the output is used directly — for secp256k1, as the big-endian private scalar, with the call failing rather than folding if the scalar lands out of range, "because folding the scalar into range would silently land two domains on one key."

The practical consequence: an agent's Ethereum address is a pure function of the app root key, the algorithm string, and a domain string the application picks. Same inputs, same address, on any machine that can reach the same KMS. That is a feature — the wallet survives hardware failure without a seed phrase backup — and it is also the attack surface we get to further down.

The signature chain is the part a relying party checks

A derived public key on its own is worth nothing. GetKey returns it with a two-element signature chain: link 0 is the app root key signing a claim over the derived key, link 1 is the KMS root key signing the app root public key. The claim encoding is length-prefixed the same way:

claim  = LP("dstack-guest-v1-key-claim")
      || LP(algorithm)
      || LP(domain)
      || LP(public_key)

digest = keccak256(claim)
link0  = r || s || v          // 65 bytes, recoverable ECDSA

The spec includes something most protocol docs omit: an argument for why the old format cannot be coerced into the new one. v0's claim was keccak256("{purpose}:{hex(pubkey)}") with purpose caller-supplied, so a malicious app could make the root key sign almost any ASCII string ending in a colon and lowercase hex. A v1 claim ends with LP(public_key), whose four length bytes are 00 00 00 21 for secp256k1, sitting 37 bytes from the end — inside the region a v0 preimage requires to be hex-only, and 0x00 is not a hex character. The two byte strings can never be equal, so forging across surfaces reduces to a keccak256 preimage attack. There is a regression test named after the attack.

Verification is six steps, and step one carries everything: obtain the KMS root public key "from a source you trust independently of the agent being checked" — the DstackKms contract's kmsInfo().k256Pubkey, or a pinned value. The spec states the failure mode plainly: "An attacker who can answer your query for the anchor can also mint a self-consistent chain, so reading the anchor from the KMS you are checking proves nothing."

Step six is the one integrations skip. After recovering the app root key and checking the KMS signature, you must "confirm the app_id you used is the application you meant to talk to. The chain proves that the KMS issued this app root key to some application; only this step ties it to yours." A chain that verifies cleanly against a real KMS and belongs to a different agent is a valid chain.

Rule of thumb — a signature that verifies under an unverified public key says nothing. The dstack spec spells it out: verify the payload signature, then run the chain steps to establish that the public key is what it claims to be. The two halves are independent, and shipping only the first half is the common integration bug.

What the quote measures

The chain establishes key provenance. The TDX quote establishes execution context. The attestation guide maps each register:

MRTD holds the virtual firmware measurement, taken by the TDX module in SEAM mode — OVMF, the first code executed after CVM startup, "serving as the App code's trust anchor." RTMR0 records the virtual hardware setup: CPU count, memory size, device configuration. RTMR1 records the Linux kernel. RTMR2 records the kernel cmdline, including the rootfs hash, and the initrd. RTMR3 is the interesting one: initrd records the dstack app details — compose hash, GPU policy, instance id, app id, key provider.

MRTD through RTMR2 can be precomputed from the built image given CPU and RAM specs, which is why dstack ships dstack-mr and a reproducible build path. RTMR3 cannot be precomputed; it is verified by replaying the event log and checking the result matches the quote. Then you check that the replayed events — compose hash, instance id, app id, rootfs hash, key provider — match what you expected.

The application binds its own data through report_data: up to 64 bytes, zero-padded on the right, with more than 64 bytes an error rather than a truncation. In v1 there is no EmitEvent; runtime RTMR3 events became system-owned in 0.6.0, so an application cannot write its own measurement. It gets 64 bytes and a nonce discipline.

Note what is not in that list. No register measures the agent's model, its prompt, its policy file, or the balance of the wallet it controls. The compose hash covers the Compose document that names the image digests. Everything the agent decides at runtime is outside the measurement.

The environment variable that picks your wallet

Now combine three facts, each documented by dstack itself.

One: the derived key is a function of (app root key, algorithm, domain), and domain is chosen by the application. Two: applications in practice take that domain from configuration — the ERC-8004 TEE agent template requires an AGENT_SALT environment variable as part of its documented configuration. Three: the encrypted-env specification states that "encryption provides confidentiality, not origin authentication," and that because app_id is public and GetAppEnvEncryptPubKey is callable with it, any party with VMM access can fetch the app encryption public key, encrypt a different env payload, and submit the replacement. "The CVM will decrypt and use that payload if decryption succeeds."

So an operator who can reach the VMM can change the domain string. The CVM boots, derives a different key, produces a different address, and attests to all of it. The quote is genuine. The measurements match. The signature chain verifies. The compose hash is unchanged, because the Compose document did not change — only the encrypted env payload did. Every box on the verification checklist gets ticked, and the agent is signing with a wallet the deployer never funded.

This is not a break of the TEE. It is the exact boundary dstack draws in its security model: "Encrypted environment variables prevent the host from reading your secrets. However, the host can replace encrypted values with different ones." The documented mitigation is the launch-token pattern — put APP_LAUNCH_TOKEN in the encrypted env and verify its hash in prelaunch, where the hash is measured via app-compose.json — or sign the env payload with a developer-held key and check it before use.

For a wallet, there is a cheaper fix: bind the address on-chain before it holds value. Register the derived address in an identity registry, fund only that address, and treat any attestation carrying a different address as a failed deployment rather than a new agent. That is the property an ERC-8004 identity record gives you that a raw attestation does not — a prior commitment to which key is supposed to appear.

What it costs to check on-chain

Everything above is off-chain verification. The moment a contract has to decide whether an agent is attested, you pay for DCAP verification in the EVM. Automata's DCAP Attestation is the production path, and its README publishes the numbers:

// Verification methods, Automata DCAP Attestation v1.1
Onchain            ~4-5M gas    proving time: instant
RiscZero Groth16    522k gas    <1 min
SP1 Groth16         493k gas    <30s
SP1 Plonk           569k gas    <2 min

// Onchain: ~4M with the RIP-7212 precompile, ~5M without

Four to five million gas is a whole block's worth of budget on some chains, which is why the ZK path exists: run the DCAP quote verification program inside RISC Zero or SP1, then verify a single Groth16 proof for about half a million gas. The entrypoint takes both:

function verifyAndAttestOnChain(bytes rawQuote, uint32 tcbEvalDataNumber);

function verifyAndAttestWithZkProof(
    bytes output,
    uint8 zkCoProcessor,      // 1 = RiscZero, 2 = Succinct
    bytes proof,
    bytes32 programIdentifier,
    uint32 tcbEvalDataNumber
);

The v1.1 release notes add a detail that matters to anyone tracking Ethereum's roadmap: secp256r1 precompile support (EIP-7951) is live for Hoodi and Sepolia with the Fusaka hardfork, and it "reduces ECDSA verification cost from 330k gas to 6000 gas per signature," for roughly "1M gas reduction per DCAP quote onchain verification." The same precompile we audited for passkey-signed x402 payments cuts a fifth off the cost of verifying a TEE quote. P-256 is the curve Intel's attestation chain and WebAuthn happen to share, and making it cheap helps both at once.

Two more operational facts. The contracts were audited by Trail of Bits in February 2025; an OpenZeppelin review in October 2025 of a downstream verifier surfaced a PCCS Router timestamp validity issue, fixed in v1.1. And the entrypoint is deployed at the same address — 0xaDdeC7e85c2182202b66E331f2a4A0bBB2cEEa1F — across Base, Arbitrum One, Optimism, Polygon, BNB Chain, Unichain, World Chain, HyperEVM and Avalanche mainnets, which makes a multi-chain agent verifier a single constant rather than a config table.

Four limits dstack documents about itself

The security model is unusually candid, and three of its limits are load-bearing for a payment agent.

// Limit 1

Attestation proves identity, not correctness

"Attestation proves which code is running, not that the code is bug-free. It proves the environment is isolated, not that your application handles secrets correctly." An attested agent with a prompt-injection hole is an attested agent that pays an attacker. The quote raises the floor on who can tamper; it says nothing about what the agent decides.

// Limit 2

Disk state has no freshness guarantee

The LUKS2 volume "carries no authentication tag," and neither filesystem "proves that an attached disk represents the latest application state. An infrastructure operator can withhold, delete, replace, or restore an earlier valid encrypted disk image." For an agent that tracks its own spend budget or nonce counter on disk, that is a rollback: restore yesterday's disk and the budget refills. dstack's own guidance is to anchor a monotonic version or state commitment in an external trusted service or ledger. For a payment agent, that means the spend limit belongs on-chain or at the facilitator, never only in the CVM's storage.

// Limit 3

TCB status is surfaced, not gated

dstack's validate_tcb "does not reject a quote based on its TCB status string" — UpToDate, OutOfDate, ConfigurationNeeded all pass the primitive, which only enforces that debug mode is off and the SEAM measurements are well-formed. Only a Revoked TCB is rejected outright, by dcap-qvl. Whether an out-of-date TCB is acceptable is explicitly "a policy decision that belongs downstream." If your verifier does not implement that policy, you accept every non-revoked platform.

The fourth is structural: "All keys derive from the KMS root key, which is protected by TEE isolation. Like all TEE-based systems, a TEE compromise could expose the root key." dstack says it is developing an MPC-based KMS to remove that single point of failure. Until then, every derived agent wallet in a deployment shares one root of compromise — the same concentration risk we mapped in the agent threat model, moved one layer down into the hardware.

Where this touches x402

x402 has no attestation field. The payment payload carries an EIP-3009 authorization — a signature over a transfer with a nonce and a validity window — and the facilitator's /verify and /settle endpoints care about exactly one thing: does the signature recover to an address that holds the funds. A TEE-derived signer is indistinguishable from a key in a .env file at the wire level.

That is the correct default for a payment protocol, and it is also why "attested agent" claims mostly do not reach the counterparty today. The two places the evidence can land are above and below the payment: an identity registry entry that commits to the address before it is funded, and a policy layer that refuses to release value to an unattested signer. Oasis's ROFL x402 work pushes on the other end — running the facilitator itself inside a TEE, on the argument that facilitators "are currently highly centralized and opaque" and could censor transactions. Same primitive, opposite side of the handshake.

Both directions are implementation patterns, not protocol. Nothing in the x402 spec obliges a facilitator to check a quote, and nothing obliges a seller to ask. The gap is real, and it is the kind of gap that gets closed by whoever has a reason to close it.

What it means for LLM4Agents

LLM4Agents sits where the attestation question becomes concrete: an agent registers, funds a balance in stablecoins, and calls an OpenAI-compatible gateway that meters per token. Three consequences follow from the audit.

First, the gateway is a natural attestation consumer, and a cheap one. We already authenticate every request. Accepting an optional VersionedAttestation at registration — verify the quote, replay the event log, check the signature chain, confirm the derived address equals the address being registered — costs one verification per agent lifetime, not one per call. That is off-chain verification with dcap-qvl, not 4M gas. The output is a tier: this agent's key provably lives in a CVM running a known compose hash.

Second, the substitution finding sets the order of operations. Bind first, fund second. If an agent registers an attested address, the platform should pin it and refuse to serve a different signer under the same agent id, even with a valid attestation. The attestation proves a key was derived in a genuine TEE; only our own prior record proves it is the key the operator intended.

Third, the freshness limit tells us where spend controls must live. An agent that enforces its own budget inside a CVM can be rolled back by whoever holds the disk. Budget enforcement belongs on our side of the meter — the balance and the per-period cap are gateway state, not agent state. That is already how the platform works, and the TEE case is an argument for keeping it that way rather than pushing limits down into the client.

The threat this does not address is the one that matters most in practice. An attested agent with a compromised prompt is still a compromised agent, and it now carries a hardware certificate that says its code is exactly what it claims. Attestation is evidence about provenance. It is not evidence about judgement, and a tier system that treats it as the latter would be worse than no tier system at all.

Staying on the frontier

Concrete steps, in the order they pay off.

1. Ship attestation verification at registration, optional and off-chain. Accept a VersionedAttestation on the register endpoint. Verify the TDX quote with dcap-qvl, replay the event log against RTMR3, verify the two-link signature chain with the KMS anchor read from the DstackKms contract rather than from the agent, and require that the derived address equals the address being registered. Reject debug-mode quotes. Surface the TCB status in the agent record instead of gating on it, then decide the policy with data.

2. Pin the signer. Once an agent registers an attested address, treat any change as a new agent. Store the app id, the compose hash and the derived address together. This is the mitigation for the env-substitution path and it costs one database column.

3. Publish the gateway's own measurements. The asymmetry is worth noticing: we ask agents to prove what they run while we prove nothing. Running the metering and settlement path in a CVM with a published compose hash makes our side auditable, and it is the credible version of "your prompts are not read." Start with the settlement component, not the whole gateway.

4. Treat on-chain verification as a settlement-time tool only. At 4-5M gas per quote, no per-call flow can afford it. The realistic on-chain uses are one-time registry entries and dispute resolution, and the ZK path at ~500k gas is what makes even those routine. Track EIP-7951 rollout past Fusaka, because a further 1M gas off the on-chain path changes which of those flows are affordable.

5. Keep a standing position on the x402 gap. There is no attestation field in the payment payload. If one is proposed, an optional evidence reference alongside the EIP-3009 authorization is the shape worth supporting — a URI plus a digest, verified out of band, never a multi-kilobyte quote in an HTTP header. Until then, the identity registry is the binding site, and that is where the effort should go.

Run agents that pay per token, in stablecoins

OpenAI-compatible gateway, x402 and EIP-3009 settlement, per-agent balances and caps enforced on our side of the meter.

Register an agent