← Blog
October 4, 2026 · 19 min

One schema, four dialects: auditing structured outputs for agents

An agent that sends strict: true thinks it bought a guarantee: the reply will parse and match the schema. We followed that flag through three providers, five translation layers, a local model and our own x402 quote. It holds at fewer hops than its name suggests, and every hop where it breaks still returns HTTP 200.

Agents run on JSON. Tool arguments, extraction results, routing decisions and payment intents all leave the model as a string that another program has to parse. For two years the answer to "what if it does not parse" has been constrained decoding. The provider compiles your JSON Schema into a grammar and masks every token that would break it. OpenAI calls the feature Structured Outputs. Anthropic ships JSON outputs and strict tool use. Gemini, vLLM and Ollama each have their own version.

The headline is the same everywhere: the output will match your schema. The fine print is not. Each provider accepts a different subset of JSON Schema. Each compatibility layer treats the strict flag differently. Some engines drop keywords without saying so. And a schema-valid answer is not the same thing as a true one.

Our sources, all read or run on 4 October 2026: OpenAI's structured outputs and function calling guides; Anthropic's structured outputs page and its OpenAI SDK compatibility page; Gemini's structured output and OpenAI compatibility docs; OpenRouter's structured outputs, response healing and provider routing docs; vLLM at commit 0872ddf, Ollama at 42e911b and LiteLLM at bb4f721; the MCP 2026-07-28 specification; and the x402 Bazaar extension at 751590a. We also ran a 16-keyword enforcement probe against Ollama 0.30.6 and probed our own walk-up 402.

Where agent schemas come from

Agent schemas are rarely written by hand for one model. They arrive through protocols.

MCP is the largest source. In the 2026-07-28 tools specification every tool carries an inputSchema and may carry an outputSchema. Both default to JSON Schema 2020-12 when no $schema is present. If an output schema exists, "Servers MUST provide structured results that conform to this schema" and "Clients SHOULD validate structured results against this schema." For a tool with no parameters, the spec recommends { "type": "object", "additionalProperties": false }.

The same page adds a note worth reading twice. structuredContent "is server-produced result data and is unrelated to LLM 'structured outputs' (schema-constrained model generation)." MCP constrains what the server returns. It says nothing about what the model emits when it fills in the tool's arguments. That half belongs to the provider and to whatever gateway sits in between.

x402's Bazaar extension is the second source. A resource server declares its HTTP endpoint or MCP tool so facilitators can catalog it. For MCP tools, inputSchema is required, follows MCP's Tool.inputSchema format, and servers "should reuse the same schema their MCP tool already declares." The info.output object, by contrast, has a type, an optional format and an optional example. There is no output schema field. An agent that finds a paid API through the Bazaar gets a contract for what to send and a sample of what comes back.

The Bazaar's own GET example also declares headers as an object with additionalProperties: { "type": "string" }. Keep that detail in mind. As the next section shows, two of the three major providers reject that shape in strict mode.

We covered discovery in the Bazaar post and the protocol revision in the MCP 2026-07-28 post. The point here is narrower. The schemas an agent hands to a model were written for validation, in full JSON Schema 2020-12. Constrained decoding speaks a dialect, and each provider speaks a different one.

What strict actually promises

OpenAI. Structured Outputs "ensures the model will always generate responses that adhere to your supplied JSON Schema." The price of that sentence is a set of rules. The root must be an object, not an anyOf, so a discriminated union at the top level fails; OpenAI's guide uses Zod as its example of the pattern. Every field must be listed in required; optional fields are emulated with a union that includes null. additionalProperties: false "must always be set in objects." A schema may have up to 5,000 object properties, 10 levels of nesting, 1,000 enum values, and 120,000 characters across property names, definition names, enum and const values.

Inside those rules OpenAI supports a lot: pattern, nine string format values, minimum, maximum, the exclusive bounds, multipleOf, minItems, maxItems, anyOf, definitions and recursive schemas. It does not support allOf, not, dependentRequired, dependentSchemas or if/then/else. With strict: true, an unsupported schema returns an error rather than a degraded answer. The first request with any new schema adds latency while the schema is processed.

The default depends on the endpoint. For function tools, Chat Completions requests "remain non-strict by default." Responses requests "will attempt to normalize your schema into strict mode when possible" and fall back to best-effort calling when they cannot, reporting strict: false on the tool. The same tool definition gets a guarantee on one endpoint and a best effort on the other.

Two edge cases remain. A safety refusal comes back in a separate refusal field, because it "does not necessarily follow the schema." Hitting the token limit leaves the response incomplete.

Anthropic. JSON outputs moved to output_config.format and are generally available. The beta output_format parameter is deprecated. Strict tool use is strict: true on a tool. Both share one set of limits.

Supported: the basic types, primitive enum, const, anyOf, allOf except with $ref, local $ref and definitions, default, ten string formats including uri, and minItems of 0 or 1. additionalProperties must be false. Not supported: recursive schemas, numerical constraints such as minimum, maximum and multipleOf, string constraints minLength and maxLength, and any array constraint beyond minItems of 0 or 1. "If you use an unsupported feature, you'll receive a 400 error with details." Regex support excludes backreferences, lookaround and word boundaries.

Then there are complexity limits per request: 20 strict tools, 24 optional parameters across all strict schemas, and 16 parameters with union types, including ["string", "null"]. Past those, or past internal grammar-size limits, the API answers "Schema is too complex for compilation." The nullable-union trick that OpenAI recommends for optional fields counts against Anthropic's union budget.

The costs are documented too. Grammars compile on first use and are cached for 24 hours from last use. Claude "automatically receives an additional system prompt explaining the expected output format," which raises input tokens, and changing output_config.format invalidates the prompt cache for that thread. Required properties are emitted first, then optional ones, so key order can differ from the schema.

And the guarantee has documented holes. A refusal returns HTTP 200 with stop_reason: "refusal", "You'll be billed for the tokens generated," and the output may not match the schema. A max_tokens stop may leave the output incomplete. Most unusual: string enum and const values are not guaranteed to keep their capitalization. Claude may return "Conversation Topic 3" for an enum value with a lowercase t, and "the response completes normally, with no error and no special stop_reason."

Gemini. Google documents a shorter list. Objects take properties, required and additionalProperties, which "can be a boolean or a schema." Strings take enum and format, "such as date-time, date, time." Numbers take enum, minimum and maximum. Arrays take items, prefixItems, minItems and maxItems. The page's own examples use anyOf and a recursive "$ref": "#" organization chart. The limitations section says "Not all JSON Schema features are supported" and "Very large or deeply nested schemas may be rejected." It does not say whether an unlisted keyword errors or is ignored. Its best practices tell you to "always validate values in your application."

# Keyword support for constrained output, per each provider's docs, 4 Oct 2026
# "-" = not in the provider's documented list
keyword                       OpenAI strict       Anthropic            Gemini
every property in required    MUST                no                   no
additionalProperties          false, every obj    false only           bool or schema
anyOf                         yes, not at root    yes                  yes (example)
allOf                         error               yes, no $ref         -
recursive $ref                yes                 400                  yes (example)
minimum / maximum             yes                 400                  yes
multipleOf                    yes                 400                  -
minLength / maxLength         -                   400                  -
pattern                       yes                 simple regex only    -
format                        9 values            10 values (+ uri)    date-time, date, time...
minItems / maxItems           yes                 minItems 0 or 1      yes
uniqueItems                   -                   400                  -

Now run real schemas through that table. Gemini's own recursive org chart sets no additionalProperties, so OpenAI's strict mode rejects it, and it is recursive, so Anthropic rejects it too. An integer bounded by minimum: 1 and maximum: 10 is enforced by OpenAI and Gemini and is a 400 on Anthropic. The Bazaar's headers map fails both OpenAI strict and Anthropic. No non-trivial schema is portable by default. Portability is something a layer has to do on purpose.

Five layers, five meanings of strict

Most agents do not call these APIs natively. They call an OpenAI-shaped endpoint and let something translate. We read what each translation layer does with response_format: { type: "json_schema", json_schema: { strict: true, ... } }.

Anthropic's OpenAI-compatible endpoint lists response_format as "Ignored" and points to native Structured Outputs. Function-level strict is also ignored, "which means the tool use JSON is not guaranteed to follow the supplied schema." Both still return a normal response.

Gemini's OpenAI compatibility documents structured output through the OpenAI SDK's parse() helper, with Pydantic and Zod examples. Its "Current limitations" section says support for the OpenAI libraries "is still in beta."

Ollama decodes json_schema into a struct with exactly one field, Schema, in openai/openai.go. The name and strict fields never get past the decoder. The bare schema becomes the native format parameter, so the schema is always applied and the flag means nothing. The docs add two lines: "Ollama's Cloud currently does not support structured outputs," and it is "ideal to also pass the JSON schema as a string in the prompt."

vLLM also ignores the flag, in the other direction. structured_outputs_from_response_format turns any json_schema into a constrained request whether strict is set or not. The default backend is auto: try xgrammar, and if the schema fails validation or uses features xgrammar does not support, fall back to guidance by default. The unsupported list includes multipleOf, uniqueItems, contains and string formats outside a supported set. One code comment is unusually frank. When a string combines pattern or format with length bounds, xgrammar "silently drops minLength/maxLength from the grammar, so output can violate the bound without any error surfacing." vLLM detects that case and reroutes it. An engine without the check would not.

LiteLLM, routing to a Claude model with native support, maps response_format to Anthropic's format and runs filter_anthropic_output_schema. It removes array, numeric, string-length and conditional constraints, rewrites oneOf to anyOf, and appends each removed constraint to the field's description. The docstring says this "mirrors the transformation done by the Anthropic Python SDK." The Anthropic SDK then validates the response against the original schema. LiteLLM's equivalent switch, enable_json_schema_validation, defaults to False. By default a removed constraint becomes a hint. For Claude models without native support, LiteLLM wraps the schema in a synthetic tool and forces tool_choice to it, except when thinking is enabled or the model does not support forced tool use; then the tool is offered but not forced.

OpenRouter says support "is determined per endpoint, not just per model." require_parameters defaults to false, and providers that lack a parameter "will ignore unknown parameters." response_format is a soft preference among providers of the same model. But "this preference never removes a model from your request's candidate list (such as the models fallback list)": if no provider of a fallback model supports it, "the request is still routed to that model and the parameter is ignored." Even with strict: true, "Enforcement varies by provider." Its Response Healing plugin repairs syntax such as missing brackets, trailing commas and markdown fences, only on non-streaming requests, and "if the response is truncated by max_tokens, the plugin will not be able to repair it."

# What happens to response_format json_schema + strict: true, per layer, 4 Oct 2026
OpenAI native              enforced; unsupported keyword -> error
Anthropic native           enforced; unsupported keyword -> 400
Anthropic OpenAI-compat    response_format ignored, strict ignored
Ollama /v1                 name and strict dropped; schema always applied
vLLM                       schema always applied; strict never read
LiteLLM -> Claude          unsupported keywords moved into descriptions;
                           output not re-validated by default
OpenRouter (defaults)      routed to supporting providers when they exist,
                           otherwise the parameter is ignored

Put that table next to a fallback chain. In our fallback post we argued that a two-model chain is the cheapest reliability buy on a gateway. With a schema attached, the second link may enforce it, reject it, or ignore it. The HTTP 200 looks identical in all three cases.

What we measured: what a grammar really enforces

Docs list keywords. They do not show what happens to the ones they leave out. So we measured one engine directly.

Setup: Ollama 0.30.6, qwen2.5:1.5b on CPU, the OpenAI-compatible endpoint, response_format with json_schema and strict: true, temperature 1.0, six runs per case. Every prompt told the model to break the constraint, for example "Report the number 500" against maximum: 10. We validated each output with Python jsonschema 4.20 under Draft 2020-12 with a format checker. A small model was deliberate: the less a model follows the schema on its own, the more the result says about the grammar.

# Ollama 0.30.6, qwen2.5:1.5b (CPU), 6 runs per keyword, temperature 1.0, 4 Oct 2026
# prompt asks for a violation; output validated with jsonschema (Draft 2020-12)
keyword                        schema-valid   what came back
required                       6/6            "b" emitted although told to omit it
additionalProperties: false    6/6            requested extra fields never appeared
enum ["red","green"]           6/6            "green" when told the color was blue
const "v2"                     6/6            "v2" when told v9
minimum 1 / maximum 10         6/6            5 when told 500
maxLength 8                    6/6            "The vast"
minLength 40                   5/6            filler text; 1 run hit max_tokens mid-string
pattern ^[A-Z]{3}-[0-9]{4}$    6/6            "ABC-1234" when told abc123
format: date                   6/6            "2023-11-27" when told "next tuesday"
minItems 3 / maxItems 3        6/6            3 fruits when asked for 10
anyOf, two branches            6/6            a valid branch, values invented
$ref recursion                 6/6            "kids" kept when asked for "children"
multipleOf 5                   0/6            7
format: email                  0/6            "Contact: John at the front desk."
uniqueItems (3-value enum)     0/6            ["apple", "apple", "apple"]
allOf [integer, minimum 100]   0/6            {} followed by whitespace

Twelve keywords held. Four were dropped: multipleOf, format: email, uniqueItems and allOf. All 24 failing outputs came back as HTTP 200 with no error and no warning. Nothing in the response told the client that a third of its constraints had not been applied. The only way to know was to validate.

The dropped set matters for agents. multipleOf is how a schema says an amount moves in fixed increments. uniqueItems is how it says a list of IDs has no duplicates. format: email guards a contact field. allOf is the composition keyword. In the allOf case the grammar did not just skip the bound; the field came back as an empty object, the wrong type entirely.

The passing rows carry the second lesson. Valid is not true. Told the color was blue, the model returned green, because green was legal and blue was not. Told 500, it returned 5. Told "next tuesday," it invented a calendar date. Asked for one word against a 40-character minimum, it padded the string with unrelated filler, and one run never closed the string before max_tokens. With no grammar, a model can say "blue" or refuse. Under the grammar, every legal output was wrong. Constrained decoding converts "I cannot answer in this shape" into a confident answer in this shape.

Truncation behaves the same way. With the same setup and max_tokens: 8, the response was HTTP 200, finish_reason: "length", content { "summary": "The, and a usage block reporting 8 completion tokens. The string does not parse. The tokens were counted anyway.

Design rule — any schema an agent uses to make a decision needs a legal way to say "none of the above": a nullable field, an "unknown" enum member, or an error branch in an anyOf. Without one, the grammar forces a fabrication whenever the truth falls outside the schema. On Anthropic, each nullable union spends one of the 16 union slots, so budget them.

Who pays when the JSON is wrong

Every failure in this post returns HTTP 200. A refusal on Anthropic is billed for the tokens generated. A truncated object is billed for the tokens it used. Healing cannot repair truncation. On a prepaid account those are line items. On a pay-per-call rail they are payments the agent already signed for a string its parser will reject.

We probed our own walk-up surface the same way we did in the reasoning tokens post. An unauthenticated POST /v1/chat/completions returns HTTP 402 with an exact requirement in USDC on Base, and the PAYMENT-REQUIRED header carries the amount. We changed only the schema-related fields.

# api.llm4agents.com, unauthenticated POST /v1/chat/completions, 4 Oct 2026
# openai/gpt-5, max_tokens 500, one message "hi" unless noted     402 amount (USDC, 6 decimals)
no response_format                                              10000    $0.01
+ valid strict json_schema                                      10000    $0.01
+ schema using allOf and not (OpenAI strict rejects both)       10000    $0.01
+ response_format {"type": "bogus"}                             10000    $0.01
+ 154 KB body: strict schema, 400 properties x 50 enum values   10000    $0.01
  same schema sent as a tool's parameters                       10000    $0.01
  same schema text pasted into the message instead              90000    $0.09
  150 KB of plain text in the message                           70000    $0.07
anthropic/claude-opus-5.5, schema with minimum/maximum          20000    $0.02
models [gpt-5, opus-5.5], either order                          20000    $0.02

Three readings.

First, the quote does not look at response_format. A valid schema, a schema OpenAI's strict mode rejects, a 20,000-enum schema twenty times over OpenAI's limit, and a format type that does not exist all receive the same 402. So does an Anthropic request using minimum, which the native API rejects and the compatibility layer ignores. The quote cannot tell a request that will work from one that cannot.

Second, schema bytes are free in the quote. Text in the message moved the quote from $0.01 to $0.07 or $0.09. The same bytes in response_format or in tools did not move it at all. Upstreams bill both. OpenAI says function definitions "count against the model's context limit and are billed as input tokens." Anthropic says the injected format prompt "costs you tokens like any other system prompt." x402 does not let a server collect more than the client signed; under upto, the settled amount "MUST be less than or equal to the authorized maximum." Whatever a large schema costs above the quote is the gateway's cost. In the two-phase gap audit one signature bought many invocations. Here one cheap signature buys an expensive prompt.

Third, the fallback list is quoted at its most expensive link, in either order. That is the right call under an exact scheme. But nothing in the quote or in the response tells the agent whether each link honored the schema. Our public OpenAPI documents model, models, messages, temperature, max_tokens and stream, with additionalProperties: true. response_format and tools pass through undocumented. The only post-hoc clue is X-Model-Used, which names the model that answered, not the dialect that was applied.

What it means for LLM4Agents

We are a translation layer, and every row in the five-layer table is a behavior we could reproduce by accident. A gateway that accepts strict: true and forwards it to a backend that ignores it is selling a guarantee it does not deliver. The agent cannot see the difference. It sees a 200 and a JSON string, and it pays.

Our models fallback is also a dialect switch. It triggers on rate limits, downtime, context length and moderation. When the first link enforces a schema and the second rejects or ignores it, fallback either drops the guarantee or turns a transient 429 into a permanent 400. Reliability and correctness pull in opposite directions unless the router knows the dialect of every link.

The quote has a blind spot exactly where agent workloads are growing. MCP-heavy agents ship long tool lists and large schemas on every call. Today those bytes cost the agent nothing at quote time and cost us real input tokens upstream. That is a subsidy for honest agents and an opening for abusive ones.

The opportunity is the same fact turned around. Agents cannot track four dialects, five translation layers and per-engine silent drops. A gateway that knows them can be the one place where a schema is checked once, translated deliberately, enforced where possible, and validated after the fact, with the verdict on the receipt. That is a product, not a patch.

Staying on the frontier

In order of cost and urgency.

1. Price what we forward. Count response_format and tools in the input estimate behind the 402, exactly like message text. It is the cheapest fix and it closes the subsidy.

2. Lint before the 402. Check the schema against the target model's dialect before quoting. Reject with a 400 that names the keyword and the model, such as minimum on a Claude model through a path that cannot enforce it. Enforce the published ceilings: OpenAI's 5,000 properties, 10 levels and 1,000 enum values; Anthropic's 20 strict tools, 24 optional parameters and 16 union parameters. An agent should never sign for a request we already know will fail.

3. Make fallback schema-aware. When a request carries strict: true, keep only the models links that can enforce that schema, or translate it per link the way the Anthropic SDK does. Either way, say which one happened.

4. Validate after, every time. Check each output against the original schema, not the translated one. Return the verdict next to X-Model-Used, for example as an X-Schema-Valid header plus a mode of enforced, translated or validated-only. The names are ours to choose; the signal is what matters.

5. Normalize the failure signals. Map Anthropic's stop_reason: "refusal" to OpenAI's refusal field and every token-limit stop to finish_reason: "length", so one parser handles every link. Never hand back a truncated object without saying so.

6. Put a price on invalid. Under upto, the settled amount "MAY be 0." A gateway with its own validator can settle less when the output fails it. Decide the policy for refusals, truncations and validation failures, publish it, and put it on the receipt. We made the metering case in the upto post; this is the correctness case.

7. Publish the dialect. Add response_format, tools and strict to the public OpenAPI. Expose per-model structured-output support in /api/v1/models, the way OpenRouter exposes supported parameters. Then ship an MCP-to-strict converter that turns optional parameters into nullable unions and sets additionalProperties: false everywhere, so MCP and Bazaar tool schemas work in strict mode without hand edits.

Steps one and two protect money. Three to five protect the guarantee. Six and seven turn it into something an agent can buy on purpose.

Schemas you can check, per call

One OpenAI-compatible gateway, paid per call in USDC over x402.

Register your agent