← Blog
September 18, 2026 · 15 min

You paid for Opus. What proves Opus answered?

An agent sends a prompt, gets tokens back, and pays for them. Every layer of that transaction is now cryptographically verifiable except the one the agent is actually buying: which model produced the tokens.

The agent payment stack has spent two years getting very good at proving payment. EIP-3009 signatures prove authorization. x402 facilitators prove settlement. Receipts bind a transaction hash to a request. We have written about most of it.

The confidential computing stack has spent the same two years getting very good at proving hardware. An Intel TDX quote proves a genuine CPU in a measured boot state. An NVIDIA attestation report proves a genuine GPU running unmodified firmware with confidential compute mode on. We audited that machinery from the wallet side in TEE-attested agent wallets.

Put the two together and you get a claim that sounds airtight: your prompt ran inside verified hardware and you have a signed receipt for it. Both halves are true. Neither one says which weights were loaded.

That gap is not theoretical. On 5 May 2026, a public registry that tests inference providers caught a gateway silently serving a different model than the one requested: deepseek-ai/DeepSeek-V3.1 in, Qwen/Qwen3.5-122B-A10B out. The hardware attestation was fine. It was always going to be fine.

What the GPU attestation actually says

Start at the bottom. NVIDIA's attestation suite has three parts: the Remote Attestation Service (NRAS), the Reference Integrity Manifest (RIM) service that holds golden measurements, and an OCSP service for certificate status. The GPU's security processor produces a signed report; NRAS verifies it and returns a JWT carrying Entity Attestation Token claims.

The Hopper single-GPU example is six calls against https://nras.attestation.nvidia.com/v4/attest/gpu:

client = attestation.Attestation()
client.set_service_key("YOUR_API_KEY")
client.add_verifier(attestation.Devices.GPU,
                    attestation.Environment.REMOTE,
                    NRAS_URL, "")

evidence_list = client.get_evidence()
result = client.attest(evidence_list)
token = client.get_token()

The token that comes back carries an eat_nonce for freshness and a set of boolean claims: x-nvidia-gpu-attestation-report-signature-verified, x-nvidia-gpu-driver-rim-signature-verified, and an overall x-nvidia-overall-att-result. Read those claims literally. They say the silicon is genuine, the VBIOS and driver match manifests NVIDIA signed, and confidential compute mode is engaged.

They say nothing about userspace. The GPU security processor measures firmware, not the process that later opens a CUDA context. An academic teardown of the H100 confidential computing implementation found the attestation report carries 64 structured records, each with a measurement specification, a size field, and a cryptographic hash — a firmware inventory. The same work measured what CC mode changes at the hardware boundary: across 4,394,976 reads of the H100's BAR0 register space, 99.78% returned zeros in CC mode, against 7.94% returning real values with CC off. That is a firewall doing its job, and it is orthogonal to what model you loaded.

The precise boundary — GPU attestation answers "is this a real, unmodified, locked-down H100?" It cannot answer "is this Opus?" No claim in the token has a slot for that question.

The chain from CPU to weights

The CPU side reaches further. A confidential VM measures its own boot: firmware first, then kernel and initrd. Anything you can get into that measurement chain becomes attestable.

Most providers use this to commit to a container configuration. NEAR AI's model verification docs describe the pattern precisely: the client calls /v1/attestation/report with a 64-character hex nonce, and the response carries a signing_address generated inside the TEE, plus either an intel_quote or an nvidia_payload. The report data binds the signing address and the nonce. The Docker compose manifest is hashed with SHA-256 and compared against the mr_config measurement from the verified quote.

That is a real chain. Hardware, boot state, container config, and a key that signs responses — all bound together and all checkable by a client that never trusts the operator. And it still stops one link short. NEAR's own documentation is honest about where: the attestation establishes hardware trustworthiness, while model identity comes from the endpoint slug or the request parameter. In other words, the model name is an assertion carried alongside a proof of something else.

The compose hash helps only as far as the compose file is specific. If it pins an image digest and a weights path, you have narrowed the problem. If the container fetches weights at runtime, or the serving process picks a backend by routing rule, the measurement covers the wrapper and not the payload.

Modelwrap: the one mechanism that reaches the weights

Tinfoil built the exception. Its Modelwrap design closes the last link in three steps.

First, weights are downloaded from Hugging Face at a specific commit hash and normalized into a read-only EROFS filesystem image. Second, a dm-verity Merkle tree is computed over that image — 4KB blocks, hashed pairwise up to a single 32-byte root — and the root hash is passed to the kernel on the command line. Because the enclave measurement already covers the kernel command line, the attestation now covers the hash tree. Third, at runtime dm-verity intercepts every disk read, walks the tree, and errors if a single bit disagrees with the committed root.

The elegant part is that the inference server is not involved. As Tinfoil puts it, vLLM "has no idea this verification is happening" — the check lives in the kernel, below the application, verifying reads as they happen.

The reason they built it this way is worth stating, because it is the same reason a signature on a model card is not enough. Signing weights and checking the signature after download protects the artifact at rest. Tinfoil assumes a stronger adversary: a malicious hypervisor that tampers with disk contents after the signature check. Continuous verification on every read is the answer to that threat model; a one-time signature is not.

The surrounding architecture follows the same discipline. The attestation architecture pins a specific OVMF firmware version for reproducibility, embeds the SHA-256 of the deployment config in the kernel command line, and publishes expected measurements as signed Sigstore bundles built by a GitHub Action — so the values a client compares against come from a public transparency log rather than from the provider's own API. Boot also queries each GPU through NVIDIA's local verifier to confirm CC mode, chaining CPU to GPU. Clients verify the certificate chain back to the CPU vendor root.

What the registry found

You do not have to take any provider's word for this, because someone is testing them continuously. The awesome-private-inference registry scores providers across nine layers: TDX quote, GPU attestation, report-data binding, key derivation, compose-hash commitment, backend attestation, channel binding, attested session, and receipt. A tenth check, "catalog to served," simply asks whether the models a provider advertises actually answer.

The findings are more instructive than the scores.

// Finding 1

A passing attestation for a model that does not exist

RedPill/Phala's legacy attestation endpoint "ignores its model parameter — anthropic/claude-opus-5 and a nonexistent model both return a passing gateway-scoped attestation." The report is real and gateway-scoped. It simply does not constrain the field the caller cares about. The registry also noted a verifier that decoded JWTs with verify_signature=False. Phala's production OS issue was resolved on 2026-08-18, and its current ACI gateway does publish per-response receipts, fetched via an x-receipt-id header against GET /v1/aci/attestation, that bind request and response hashes to the attested workload.

// Finding 2

A verified quote is not verified code

Chutes was scored Stage 0 because its serving code is unmeasured — excluded from the config filesystem and absent from any runtime measurement register. The registry's summary is the sentence to remember: "verified quote ≠ verified code/model." Prompt exfiltration was demonstrated live against it.

// Finding 3

The scoring bar itself was wrong

Venice was originally scored while excluding prod_os_image and serving_code_attested — the two layers where its prompt-path exposure actually lived. An external audit caught the omission and the score was corrected to 6/9 on 2026-08-10. When the people building the measurement framework can misplace the boundary, an agent reading a marketing page has no chance.

Against that backdrop, the May substitution reads differently. A closed-chain client that checked the model name caught a live silent gateway swap on 2026-05-05. It was not the TDX quote that caught it, and it was not the GPU attestation. It was the one check that compared what was asked for against what came back.

Who holds the expected values

An attestation is a comparison. The quote tells you what measurements the machine reports; something else has to tell you what those measurements should be. That second half is where the trust actually sits, and the three providers that pass the registry's bar answer it in three different ways.

Tinfoil anchors in a software transparency log. Expected measurements are produced by a GitHub Action building a public repository and published as signed Sigstore bundles, so changing the code without changing the published expected values is visible in an append-only log.

NEAR AI anchors on-chain. The registry credits it with Stage 1 for a closed-chain client built on ComposeHashAdded events, with the contract implementation verified on Sourcify as of May 2026. The set of acceptable compose hashes is a public, append-only, on-chain list — which for an agent that already reads chain state is the cheapest possible source of truth.

The third answer, and the common one, is that the provider serves its own expected values from its own API. That is not worthless — it still catches a compromised node that the provider's control plane has not blessed — but it collapses to the problem OpenPcc names: the entity operating the service also defines what correct looks like.

The distinction matters for agents specifically. A human developer verifies once, at integration time, and reads a blog post about the provider's security posture. An autonomous agent verifies on every session, has no capacity to be reassured, and needs the expected values to come from somewhere it can query without permission. A transparency log and an on-chain event list both qualify. A provider endpoint does not.

This also explains why the registry scores "backend attestation" as its own layer. A gateway that forwards an upstream quote unchanged looks identical, on the wire, to one that verified it. The client cannot tell the difference from the response alone, which makes it exactly the kind of check that quietly degrades to nothing.

What verification costs

The reasonable objection is that all of this is too slow for per-call inference. The measurements say otherwise.

The OpenPcc paper benchmarks confidential LLM serving on commodity TEEs with composite CPU and GPU attestation. Time-to-first-token overhead against a non-confidential baseline came in at a median of 6.73%, ranging from 0.44% to 28.9% — the worst cases in the smallest cells, where fixed costs dominate, and under 1% in the largest. Decode throughput ranged from −4.3% to +11.5% against baseline, median 3.8%, with a peak of 4,340 tokens per second.

Attestation itself is a fixed cost you amortize. A single attested request pays 2,032 ms. At 100 requests per attestation TTL that falls to 21.9 ms per request, and past 1,000 requests to under 2.6 ms. For an agent making one call this is punitive. For a gateway holding a session, it rounds to nothing.

The paper's design also names the trust problem behind the whole category. In Apple's Private Cloud Compute and Google's equivalent, the operator of the inference service is also the operator of the hardware root of trust — so a hardware vulnerability or an insider at the operator has no externally verifiable mitigation. That is why transparency logs and third-party verifiers matter more than a well-written security page. It is also why the scope exclusions are worth quoting: OpenPcc defends against a malicious operator reading or logging prompts, and explicitly excludes side channels and prompt injection.

What it means for LLM4Agents

This is uncomfortably close to home, because LLM4Agents routes. A gateway that offers model fallback is, by construction, a gateway that sometimes returns a different model than the one named in the request. We described the mechanics in model fallback chains: when the primary is rate-limited or down, the request goes to the next model in the chain, and the agent gets an answer instead of a 503.

That is the same physical act as the May substitution. The only difference is authorization and disclosure. A fallback the agent configured, priced, and can see in the response is a feature. The identical swap, undisclosed, is fraud. Nothing in the hardware attestation layer distinguishes the two — which means the distinction has to live in the payment layer, where we already have a receipt.

The payment side is further along than people assume. In prompt caching and the x402 upto scheme we argued that a per-call price is only defensible if the receipt explains why the number is the number — cache reads versus writes, tokens in and out. Model identity is the same argument one level up. A receipt that says what you paid without saying what ran is answering the easier question.

The threat that matters for us is not exotic. An agent under a spend cap picks a cheap model deliberately. A gateway that quietly upgrades or downgrades that choice breaks the agent's cost model in one direction and its quality assumptions in the other. And an agent with an on-chain identity under ERC-8004 accumulates reputation from outputs it cannot attribute to a specific model — which makes the reputation noisier than it looks.

Staying on the frontier

Concrete steps, in the order they pay off.

First, name the model in the response, always. The OpenAI-compatible model field in a completion should carry what actually served the request, not an echo of what was asked. If a fallback fired, the response says so. This costs nothing and closes the exact hole the May substitution went through. Any gateway that echoes the request parameter is publishing an assertion it never checked.

Second, extend the receipt with served-model identity. The x402 offer-and-receipt extension is deliberately privacy-minimal and carries no usage detail. That is a reasonable default and a poor fit for metered inference. A receipt should bind, at minimum: the requested model, the served model, the token counts, and a hash of the response — signed by the gateway. Then a disputed charge is checkable after the fact rather than a matter of trust.

Third, run the registry's checks against ourselves. The nine layers are a free audit rubric. Most are not applicable to a non-TEE gateway, but the two that matter most are: catalog-to-served, and a receipt that binds request, route, and response hashes. Both can be implemented without any confidential computing at all, and both are what actually caught real failures.

Fourth, treat attested inference as a routing tier, not a rewrite. Providers with real coverage already speak OpenAI-compatible HTTP. Adding one as a premium destination — for agents that ask for it and pay for it — is a routing rule plus a verification step, not an infrastructure program. The overhead numbers above make it viable for session-holding workloads today. Price it as its own tier and let the agent choose.

Fifth, verify upstream attestation before forwarding, and say whether we did. The registry's "backend attestation" layer exists precisely because forwarding an upstream quote to the client is not the same as checking it. If we route to an attested provider, we check the quote at our edge and record the outcome in the receipt. An unverified upstream is a fine thing to use and a bad thing to imply.

The broader point is that the industry solved the wrong half first, and did it very well. We can prove the chip, the firmware, the boot chain, the container, and the payment. The thing being sold is the only unproven element in the chain. Until a served-model commitment is as routine as a settlement hash, "which model answered" stays a question of trust in an architecture built specifically to remove it.

Pay per call. Know what answered.

An OpenAI-compatible gateway with stablecoin settlement, explicit fallback, and receipts that name the model.

Register your agent