← Blog
October 2, 2026 · 15 min

Auditing Open Responses: passing the suite isn't honoring the spec

Open Responses is the open, multi-provider spec built on OpenAI's Responses API. Ollama, vLLM, OpenRouter and Hugging Face already serve it. We read both releases, the OpenAPI and the 17-test compliance suite, ran the suite against Ollama 0.30.6, then probed the MUSTs the suite does not cover. The server passed the core HTTP tests. Then it ignored tool_choice, truncation and previous_response_id, and answered every request with HTTP 200.

For a gateway, the interface is the product. Agents talk to us through an OpenAI-compatible surface, and that surface is moving. Chat Completions was built for turn-based chat. The Responses API was built around tool calls, reasoning and streamed state. Open Responses is the attempt to make that second shape a vendor-neutral standard, so that a client can describe a request once and run it on any provider.

That is exactly the promise a router makes. So the question for us is concrete: if we add /v1/responses and route it across upstreams, what do we inherit? What does the spec guarantee, what does its test suite actually check, and what do real servers do when a request leaves the happy path?

Our sources, all read or run on 2 October 2026: the openresponses/openresponses repository at 92c12d9, including both specification releases, the dated OpenAPI files, the technical charter and src/lib/compliance-tests.ts; Ollama's OpenAI compatibility docs and a local Ollama 0.30.6; vLLM's docs and its Responses source at 58b3298; OpenRouter's Responses API docs; Hugging Face's launch post; OpenAI's current OpenAPI; the x402 v2 transport specs; and unauthenticated probes against four hosted endpoints, ours included.

What Open Responses is

Open Responses launched on 15 January 2026, the date of its first spec release. It describes itself as "an open-source specification and ecosystem for building multi-provider, interoperable LLM interfaces based on the OpenAI Responses API." Hugging Face's launch post puts it as "Initiated by OpenAI, built by the open source AI community." The spec defines one primary endpoint, POST /responses. The second release, dated 24 April 2026, added POST /responses/compact and a WebSocket mode on the same resource.

The design rests on four ideas. The agentic loop: a request can let the model call tools, read results and continue before it yields. Items: the atomic unit of context, such as a message, a function_call or a reasoning item, usable as both input and output. Semantic streaming: events such as response.output_item.added and response.output_text.delta instead of raw text chunks. And state machines: every item is in_progress, incomplete or completed, with defined transitions.

Extensions have one hard rule. Provider-specific items, events and hosted tools MUST carry a provider prefix, as in openai:web_search_call. Clients MUST be able to ignore unknown extension events and still reconstruct the canonical response. That is the right rule for a router. It lets each upstream innovate without breaking the clients we serve.

Governance is written down. A technical charter puts the spec under a Technical Steering Committee whose seats are held by individuals, not companies, and says "No single vendor may control a majority of Core Maintainer seats." The CONTRIBUTING file names Steve Coffey of OpenAI as Lead Core Maintainer and seven more core maintainers: one more from OpenAI, two from Hugging Face, and one each from Databricks, Amazon, Ollama and OpenRouter. Spec text is CC-BY-4.0. Code is Apache 2.0.

The homepage shows twelve backers: NVIDIA, Vercel, OpenRouter, Hugging Face, LM Studio, Databricks, Red Hat, AWS, Ollama, OpenAI, vLLM and Llama Stack. Anthropic and Google are not among them, although the spec's own motivation promises requests that "run on OpenAI, Anthropic, Gemini, or local models."

One more structural fact matters later. CONTRIBUTING says the OpenAPI source files "are copied from OpenAI's first-party API," must be kept pristine, and that Open Responses additions live in a separate patch file. In practice the schema is a filtered snapshot of OpenAI's API plus normative prose. There have been two snapshots. The last change to the repository landed in mid-July, when the spec gained versioned URLs.

What the spec gets right

Several of its MUSTs are exactly the guarantees an autonomous agent needs from an inference endpoint.

allowed_tools separates what the model can see from what it may call. The full tools list stays in context, so the prompt cache survives, while the request narrows the callable set:

{
  "model": "any-model",
  "input": "Send an email to [email protected] with subject 'hi'.",
  "tools": [
    { "type": "function", "name": "get_weather", ... },
    { "type": "function", "name": "send_email",  ... }
  ],
  "tool_choice": {
    "type": "allowed_tools", "mode": "auto",
    "tools": [{ "type": "function", "name": "get_weather" }]
  }
}

The spec is unambiguous: "Servers MUST enforce allowed_tools as a hard constraint," and a call to a tool outside the list "MUST be rejected or suppressed by the server." It presents this as per-request tool governance for tenant policies, user roles and feature flags. For an agent platform, that is a security control, not a convenience.

truncation: "disabled" lets a client choose a hard failure over silent context loss: "The server MUST NOT truncate any input. If the combined context exceeds the model's maximum context window, the request MUST fail with an error instead of silently dropping content." An agent paying per token, or acting on a long contract, wants precisely that.

previous_response_id lets a client continue without resending the transcript. When it is present, "the server MUST load both the input and output associated with that prior response" and preserve their order. Reasoning items get three optional fields: raw content, opaque encrypted_content, and a user-safe summary. Hugging Face's launch post highlights that providers may now expose raw reasoning. OpenAI's models had exposed only summaries and encrypted content.

What the compliance suite checks

The repository ships an acceptance suite, also exposed as a web tool on the spec site. At 92c12d9 it has 17 tests. One of them, the output-phase schema test, never touches a server: its request is annotated "Local schema fixture; no HTTP request is sent." The other 16 split into seven HTTP core tests (basic text, system prompt, multi-turn, tool calling, image input, streaming, assistant phase), seven WebSocket tests, and two for /responses/compact.

Read the list for what is missing. No test sends allowed_tools. None sends tool_choice: "none", "required" or a forced function. None sends truncation. None uses previous_response_id over HTTP, or store. None checks the stream's [DONE] terminator, which the spec says the server MUST send. Error handling is checked in two cases only: a compaction request without a model, which must return 400 or 422, and a missing previous response over WebSocket.

And the schema the tests validate against accepts "usage": null. The generated validator reads usage: z.union([usageSchema, z.null()]). A server can pass every HTTP test without ever reporting what a request consumed.

So "passes the suite" certifies shape. It does not certify behavior on any of the MUSTs above.

We ran it against Ollama

Ollama's docs say /v1/responses arrived in v0.13.3 and that "Only the non-stateful flavor is supported." We ran the suite with Bun against a local Ollama 0.30.6, one test at a time, using llama3.2:3b pinned to CPU because the host's GPU was busy with other workloads. Model speed is irrelevant to these checks. Model capability is not, and we flag the one place it mattered.

# Open Responses suite @ 92c12d9 vs Ollama 0.30.6 (llama3.2:3b, CPU), 2 Oct 2026
basic-response                 PASS
system-prompt                  PASS
multi-turn                     PASS
tool-calling                   PASS
streaming-response             PASS   21 events
assistant-phase                PASS
response-output-phase-schema   PASS   local fixture, no request sent
image-input                    FAIL   400, model is text-only
compact-response               FAIL   404 page not found
compact-missing-model          FAIL   404 page not found
websocket-* (7 tests)          FAIL   WebSocket connection failed

Six of the sixteen tests that reach a server passed. The pattern is clean: the launch-era HTTP core works, and nothing from the April release does. There is no /responses/compact and no WebSocket mode.

The image failure belongs to the model, not the protocol. Llama 3.2 3B has no vision. But the error is instructive. Ollama returned an outer envelope whose message field was itself a JSON-encoded error from the inference backend, with its own code, message and type. A client that parses one layer sees a string.

Then we tested what the suite skips

We sent the requests the suite never sends, against the same server. For the truncation test we built a second model alias with a 2,048-token context. The same input, measured on the default 65,536-token alias, is 4,833 tokens.

# Probes beyond the suite, Ollama 0.30.6, 2 Oct 2026
# request                                   HTTP  status     echoed field        observed
allowed_tools=[get_weather], email prompt   200   completed  tool_choice "auto"  send_email called, 3 of 3
tool_choice "none"                          200   completed  tool_choice "auto"  send_email called, 2 of 2
tool_choice {function: get_weather}         200   completed  tool_choice "auto"  send_email called, 2 of 2
tool_choice "required", no tool needed      200   completed  tool_choice "auto"  plain text, no call, 2 of 2
truncation "disabled", 4,833-token input    200   completed  truncation          input_tokens 2,047
  on a 2,048-token context                                   "disabled"
store true                                  200   completed  store false         no error
previous_response_id (real, prior turn)     200   completed  null                prior turn not loaded
previous_response_id (nonexistent)          200   completed  null                no error
stream terminator                            -     -          -                  no [DONE] after response.completed
unknown model                               404    -          -                  type "not_found_error", code null

Every tool_choice form was dropped. Restricted to get_weather, the model called send_email in all three runs, and the server returned it as a normal completed response. Told "none", it still called the tool. The response echoed "auto" each time. That echo is the only trace of what happened.

Truncation went the other way. The response echoed "disabled" while usage.input_tokens reported 2,047 of the 4,833 tokens we sent. About 58% of the prompt was dropped. Here the echo was wrong and only the usage count told the truth.

State was silently absent. In the first turn we gave the model a codename, BLUEFIN. In the second, sent with that response's ID as previous_response_id, we asked for it. The server returned 200 and echoed previous_response_id: null, and the model answered "Nightshade." An ID that never existed also returned 200.

To be fair to Ollama: its docs list previous_response_id and truncation as unsupported, and tool_choice is not among the request fields they mark as supported. Nothing here is hidden. The problem is the failure mode. An unsupported field is accepted, ignored, and answered with HTTP 200 and status: "completed". And the official suite certifies the server anyway.

The spec's own definition of compliance is "an API that implements this spec directly or is a proper superset of Open Responses." A server that ignores allowed_tools is not a superset. It is a subset that returns the same status code as a superset.

The rest of the field

vLLM, another backer, is explicit in its source about one silent downgrade. Unless the operator sets VLLM_ENABLE_RESPONSES_API_STORE=1, store: true is quietly turned off: "we opted to implicitly disable store and process the request anyway, as we assume most users do not intend to actually store the response." With the flag on, "Messages are kept in memory only," and "Enabling this option will cause a memory leak, as stored messages are never removed from memory until the server terminates." An unknown previous_response_id does return a not-found error.

OpenRouter chose the honest version of statelessness. Its docs state: "Requests that set store: true or a non-null previous_response_id are rejected with a 400 error." Its error vocabulary is deliberately small: invalid_prompt, rate_limit_exceeded, image_content_policy_violation, server_error. The docs say context_length_exceeded collapses into invalid_prompt, and add a top-level error_type field to recover the precise cause. Its Responses docs page also shows a streaming example with response.content_part.delta and response.done, event names the spec does not define. Without a key we could not check the live stream.

Error envelopes do not agree either. We sent the same unauthenticated request to three hosted Responses endpoints. OpenAI answered 401 with {"error":{"message":...,"type":"invalid_request_error","param":null,"code":null}}. OpenRouter answered 401 with {"error":{"message":"No cookie auth credentials found","code":401}}, with no type and a numeric code. Hugging Face's router answered 401 with an HTML page. The spec says non-streaming servers "MUST return data only as application/json," though an auth layer in front of the API may not count.

The spec is not consistent with itself here. Its error-type table lists invalid_request, not_found, server_error, model_error and too_many_requests. Its own example error uses invalid_request_error with code model_not_found. Ollama sent not_found_error with a null code. A router that wants to fail over on "model not found" has to match at least three spellings.

The billing layer is not in the spec

For a gateway that charges per call, four gaps matter more than conformance.

First, usage. The spec's Usage object has five numbers: input, output and total tokens, plus cached_tokens and reasoning_tokens. There are no cache writes. OpenAI's own current OpenAPI requires cache_write_tokens in Responses usage, and Anthropic bills a 5-minute cache write at 1.25 times the input rate, as we covered in the cache decides the price. The snapshot lags its source. There is no cost field and no currency. And, as noted, usage may be null.

Second, payment errors. The error table covers 400, 404, 429 and 500. There is no 402. On HTTP that is fine. The x402 v2 HTTP transport puts the whole protocol in headers: "All x402 protocol information is communicated through headers (PAYMENT-REQUIRED, PAYMENT-SIGNATURE, PAYMENT-RESPONSE)." The Responses body never needs to know.

The WebSocket mode is different. x402 v2 defines exactly three transports: HTTP, MCP and A2A. On a socket, the only HTTP exchange is the upgrade request. A server can demand payment there, but one payment then opens a connection that may run sequential response.create turns for up to the spec's 60-minute limit. There is no header slot per turn. Per-call x402 pricing does not survive the transport. A socket needs either a prepaid balance or a session-level payment scheme that does not exist yet.

Third, price discovery. With previous_response_id, the server rebuilds the context as prior input plus prior output plus new input. With truncation on auto, it may then cut it. The client does not know its input token count when it signs. A fixed exact payment cannot match that; a maximum with metered settlement can. That is the case for the upto scheme.

Fourth, portability of state. encrypted_content on reasoning items is opaque, with a "format and cryptographic properties" that are "provider-specific." Compaction returns a compaction item that is also an encrypted blob. Extension items "SHOULD be treated as specific to that implementation and not assumed to be portable." A router that fails over from one provider to another mid-conversation cannot hand the second provider the first one's encrypted items. Model fallback chains need a new rule once those items appear.

What it means for LLM4Agents

Today our gateway speaks Chat Completions. An unauthenticated POST /v1/chat/completions returns a 402 with an x402 v2 exact requirement: 10,000 atomic units of USDC on Base (eip155:8453), one cent per call. POST /v1/responses returns 404, and our public OpenAPI does not list it. Our fallback feature, a models array of two or three slugs, exists on Chat Completions only.

Open Responses is the obvious next surface. Hugging Face says the aim is a shared format "practically capable of replacing chat completions." Ollama, vLLM, OpenRouter and Hugging Face's router already serve /v1/responses.

The threat is inheritance. If we proxy Open Responses naively, our clients get the behavior we measured wherever a server like that sits behind us: tool restrictions dropped, truncation hidden, state ignored, all with HTTP 200. Our fallback feature turns encrypted reasoning into a cross-provider error. previous_response_id turns a stateless payment gateway into a store of customer transcripts. And WebSocket mode does not fit per-call x402 at all.

The opportunity is the same fact seen from the other side. A gateway is the one place where the MUSTs can be made true regardless of upstream. Enforcing allowed_tools in the gateway gives an agent's operator a guarantee no single small-model server gives today. Counting tokens before forwarding makes truncation: "disabled" mean what it says. That is a reason to route through us, not just a compatibility checkbox.

Staying on the frontier

First, ship a stateless POST /v1/responses and fail honestly. Reject store: true and any previous_response_id with a 400 until we actually store state, as OpenRouter does. Never answer an unsupported field with 200.

Second, enforce the MUSTs at the gateway rather than trusting upstreams. Validate every function_call in the output against tool_choice and allowed_tools before returning it, and on a violation suppress it or return a model error, as the spec allows. Count input tokens before forwarding. With truncation: "disabled" and an oversize input, return 400. After each call, compare our count with the upstream's usage.input_tokens and alert on a mismatch.

Third, make usage a billing record. Never return null. Add cache_write_tokens, which keeps us a proper superset. Carry the charge, in USDC atomic units, and the x402 settlement reference in an llm4agents:-prefixed extension field or event, so portable clients can ignore it and ours can reconcile it.

Fourth, price by what the client controls. Use exact only when the full input is in the request and output is capped. Use upto, with the maximum stated in the 402, whenever server-side context can grow.

Fifth, give fallback a state rule. Once a chain contains an encrypted reasoning or compaction item, pin it to that provider. If we must fail over, drop foreign-prefixed and encrypted items explicitly and tell the client with a prefixed event.

Sixth, leave WebSocket mode for later, and gate it with a funded API-key balance when we add it. Keep x402 on HTTP until a session-level scheme exists. Treat /responses/compact output as provider-bound state.

Seventh, test upstreams continuously and push fixes upstream. Run the official suite plus the probe set above against every Responses upstream on a schedule, and publish the results. Then open proposals through the repository's process for tests covering allowed_tools, tool_choice, truncation, HTTP previous_response_id and the [DONE] terminator, and for cache_write_tokens in Usage. How a server should fail is what this spec most needs, and those tests are where to start.

Inference that fails loudly

One OpenAI-compatible gateway, paid per call in USDC over x402.

Register your agent