Tool calls across a fallback chain: what breaks and what lies
A tool-calling agent loop is a conversation the model writes half of. When a gateway moves that conversation to another model mid-loop, the new model inherits the old one's tool calls, call IDs and signed reasoning. We read four providers, four translation layers and the MCP spec, ran a local model, tested call IDs against Mistral's validator and probed our own x402 quote. Some hops reject the inherited history with a 400. Others accept every tool control and ignore it with a 200.
Tool calling is the primitive under every agent. The model asks for a function, the client runs it, the result goes back, and the loop repeats until the model answers. Four parts of the request shape that loop: tools, tool_choice, parallel_tool_calls and the message history with its call IDs. On one provider they mean one thing. Behind a gateway with a fallback chain, each link reads them in its own dialect.
This is the third audit in a series on what an OpenAI-compatible gateway really translates. The structured outputs audit followed strict: true. The reasoning tokens audit followed hidden output. This one follows the tool envelope: who must call a tool, how many at once, what an ID may look like, and what opaque state rides along with each call.
Our sources, all read or run on 5 October 2026: OpenAI's function calling guide and its OpenAPI spec at 31af4fc; Anthropic's tool use overview, implementation guide, parallel tool use, handle tool calls, thinking, preserved thinking, refusals and fallback, errors, Messages API reference and OpenAI SDK compatibility pages; Google's function calling, thought signatures and Gemini 3 guides plus the API reference; Mistral's function calling guide and mistral-common at f3bb6e8; OpenRouter's tool calling, parameters, provider selection and model fallbacks docs; LiteLLM at 7a7d27c, vLLM at 4a30c4c and Ollama at 42e911b; and the MCP 2026-07-28 tools specification. We also ran a tool-control probe against Ollama 0.30.6, validated call IDs with mistral-common 1.12.0, and probed our own walk-up 402.
Four knobs, four dialects
Start with what each provider documents for forcing and limiting calls.
OpenAI's guide lists four tool_choice modes: auto (the default), required, a forced function, and allowed_tools. The last one restricts calls to a subset without changing the tools list, "so you can maximize savings from prompt caching." "none" imitates passing no functions. parallel_tool_calls defaults to true in the OpenAPI spec, and setting it to false "ensures exactly zero or one tool is called." One caveat matters for strict schemas: on fine-tuned models, when the model calls several functions in one turn, strict mode is disabled for those calls.
Anthropic names four options with different words: auto, any, tool and none. There is no top-level parallel flag. Parallel use is on by default, and you turn it off with disable_parallel_tool_use: true inside the tool_choice object: "It is not a top-level request parameter." With auto that means at most one call; with any or tool, exactly one. When tool_choice is any or tool, the API prefills the assistant turn, so the model writes no text before the call.
Gemini's generateContent API has four modes as well: AUTO, ANY, NONE and VALIDATED. The last one lets the model choose between a call and text, but validates calls with constrained decoding. A subset goes in allowedFunctionNames, which the API reference says "should only be set when the Mode is ANY or VALIDATED." The FunctionCallingConfig object has exactly those two fields. There is nowhere to put parallel_tool_calls: false.
Mistral's guide documents auto, any and none, plus parallel_tool_calls. Its own tokenizer library marks any as "deprecated in favor of required" and accepts both, plus a named tool.
# Forcing and limiting tool calls, per each provider's docs, 5 Oct 2026
control OpenAI Anthropic Gemini Mistral
model decides auto auto AUTO auto
no calls none none NONE none
at least one required any ANY required (any: deprecated)
one named function {type: function} {type: tool} ANY + one name named tool
subset of tools allowed_tools - ANY/VALIDATED + names -
no parallel calls parallel_tool_calls disable_parallel_tool_use no field parallel_tool_calls
= false inside tool_choice = false
A translator can map most rows. The mapping is not the hard part. The hard part is that some targets refuse a row outright, some accept it and do nothing, and some attach state to the response that the next request must return untouched.
Rejected: forced tool use on the newest Claude models
The most consequential row is "at least one." Anthropic's errors page is direct: Claude Opus 5.5, Claude Sonnet 5.5, Claude Fable 5.1 and Claude Mythos 5.1 "don't support forced tool use." Sending {"type": "any"} or {"type": "tool", ...} to any of them, "including on the token counting endpoint," returns a 400 invalid_request_error:
tool_choice: type "tool" and "any" are not supported for this model.
Anthropic's replacement is auto with strict tool use to keep inputs schema-valid, or structured outputs when the response itself needs a fixed shape. Manual extended thinking (thinking: {type: "enabled"}) carries the same restriction on every model. Adaptive thinking does not, except on those four models.
In OpenAI's dialect, tool_choice: "required" is the standard way to say "always act, never chat," and a forced function is the standard way to say "extract into this shape." A translation layer has to map them to any and tool, and on the newest Claude tier both mappings are a 400.
LiteLLM shows the two ways a translator can respond. Its model map at 7a7d27c flags 46 entries with supports_forced_tool_use: false, covering Opus 5.5, Sonnet 5.5 and Fable 5.1 across Anthropic, Bedrock, Vertex, Azure and Databricks, plus Mythos 5.1. For those models, common_utils.py raises a client-side 400 that tells the caller to use auto and prompt for the tool. If drop_params is set, it logs a warning and rewrites the choice to auto, keeping disable_parallel_tool_use.
Both behaviors are defensible. Only one is honest to the caller. The downgrade turns "the model must call a tool" into "the model may call a tool," and the response is a 200 either way. An agent that relies on required to never get free text will get free text, from a model it may not know it was routed to.
Now put it in a chain. A request with models: ["openai/gpt-5", "anthropic/claude-opus-5.5"] and tool_choice: "required" is valid on the first link and invalid on the second. It works until the day the first link returns a 429. Then the reliability feature turns a transient rate limit into a permanent 400, or into a silent downgrade.
Anthropic already wrote the rule for its own fallback. Its server-side fallback, in beta, retries classifier refusals on other Claude models, and one rule governs the list: "The request must be valid as a direct request to every model named. If a fallback model does not support a feature the request uses, the API rejects the request up front." That is the rule a multi-provider chain needs, and it costs nothing to apply before the first token.
Rejected: history another model wrote
Tool controls are per request. History is cumulative. By the third step of a loop, the messages array holds tool calls generated by some model, with IDs minted by some server, and sometimes opaque blobs that server expects back. Four constraints decide whether a different model accepts that history.
Gemini 3 wants its signatures back
Google's thought signatures page states the rule in bold: "When using Gemini 3 models, you must pass back thought signatures during function calling, otherwise you will get a validation error." Validation covers the current turn: every model step after the most recent user message with ordinary content. The first functionCall part in each step must carry its thought_signature. With parallel calls, only the first call has one. Omit it and the request fails with a 400 of the form "Function call FC1 in the 1. content block is missing a thought_signature."
Order is validated too. If the model returned two parallel calls and the client sends them back interleaved with their results (call, result, call, result), the FAQ says the API returns a 400. The expected shape is both calls, then both results.
On the OpenAI-compatible endpoint, the signature rides in a nonstandard field, tool_calls[].extra_content.google.thought_signature, and Google's example marks it "Required and Validated." A client, SDK or gateway that rebuilds tool calls from OpenAI's typed schema, without passing unknown fields through, drops that field. The next step of the loop then fails.
Fallback makes it worse. A history whose current-turn calls came from GPT-5 or Claude has no Gemini signatures at all. Google's answer is a dummy value that skips validation. The FAQ offers "context_engineering_is_the_way_to_go" or "skip_thought_signature_validator", and the Gemini 3 guide repeats the first as the bypass for "transferring a conversation trace from another model." The FAQ also calls injecting custom function call blocks "strongly discouraged."
LiteLLM implements both halves of the workaround. To survive OpenAI clients that drop unknown fields, factory.py embeds the signature in the tool call ID itself, as call_<uuid>__thought__<base64_signature>, and parses it back out on the next request. When a call has no signature, it falls back to a base64-encoded skip_thought_signature_validator. Its own comment calls that a last resort, for use only when no real signature exists.
An ID is not just an ID
The ID trick solves one problem and creates another. Anthropic's Messages API reference constrains tool_use.id and tool_result.tool_use_id to ^[a-zA-Z0-9_-]+$. Base64 contains +, / and =. So when a LiteLLM history that touched Gemini 3 is replayed to Claude, the IDs are invalid. LiteLLM ships a normalize_anthropic_tool_use_id function that strips the __thought__ suffix and replaces any remaining invalid characters with underscores.
Mistral-format models are stricter. In mistral-common at f3bb6e8, the request validators for tokenizer versions v3 through v11 require every tool call ID to match ^[a-zA-Z0-9]{9}$: nine characters, letters and digits only. The v13 validator relaxed this to any non-empty ID other than the literal null. vLLM, when serving with a Mistral tokenizer, does not reject long IDs. Its tokenizers/mistral.py cuts any ID longer than nine characters down to its last nine and logs a warning, before the request reaches mistral-common's validator.
Truncation works when the tail is alphanumeric. It fails when it is not. We installed mistral-common 1.12.0 from that commit and validated a three-message history (user, assistant tool call, tool result) with different IDs:
# mistral-common 1.12.0 @ f3bb6e8, validate_messages, 5 Oct 2026
# "last 9" = what vLLM's truncate_tool_call_ids passes on
# Ollama ID from our probe; OpenAI-style and base64 IDs are synthetic, same shape
ID source ID tested v3-v11 v13
Mistral docs example D681PevKs ok ok
OpenAI-style call_ + 24 chars, last 9 nWzQa5sJd ok ok
Anthropic docs example toolu_..., last 9 917835lq9 ok ok
Ollama call_ + 8 chars, as generated call_2inqttcy reject ok
Ollama call_ + 8 chars, last 9 _2inqttcy reject ok
base64 signature tail, last 9 1ZQ+/kqA= reject ok
The Ollama row is the surprise. In our runs Ollama's IDs were call_ plus eight characters, thirteen in total, so the last nine begin with the underscore. A history produced on a local Ollama model and replayed to a vLLM-served Mistral model on an older tokenizer fails with "Tool call id was _2inqttcy but must be a-z, A-Z, 0-9, with a length of 9." Nothing about the conversation is wrong. Only the shape of the ID is.
vLLM adds one more format of its own. For the Kimi K2 model family, chat_utils.py mints IDs as functions.{name}:{index} instead of its default chatcmpl-tool-<uuid>. A history that crosses into or out of such a model carries IDs no other family would produce.
Tool names MCP allows and providers do not
The tool list has its own dialect problem. The MCP 2026-07-28 tools spec says tool names should be 1 to 128 characters drawn from letters, digits, underscore, hyphen and dot, and gives admin.tools.list as a valid example. It tells clients and proxies that aggregate several servers to resolve collisions, for example by "prefixing tool names with a server identifier."
OpenAI's spec says a function name "must be a-z, A-Z, 0-9, or contain underscores and dashes, with a maximum length of 64." Anthropic's reference allows ^[a-zA-Z0-9_-]{1,128}$. Google's best-practice list says to use function names "without spaces, periods, or dashes." A spec-valid MCP name with a dot falls outside both OpenAI's and Anthropic's documented patterns. A prefixed name that fits Anthropic's 128 characters can overflow OpenAI's 64. A gateway that relays MCP tools has to rename them, keep the mapping, and reverse it on every call the model makes.
Where each result goes
The last history rule is structural. Anthropic requires the tool_result blocks to sit in the user message immediately after the assistant's tool_use message, with all results for a parallel batch in that single message and before any text. OpenAI's chat format sends one tool message per call, so a translator has to merge them. Anthropic's parallel tool use page calls incorrect result formatting the most common reason Claude stops making parallel calls, because it "teaches" Claude to avoid them, and names "a separate user message for each tool result" as the wrong pattern. A careless translator does not get an error. It gets a model that slowly stops batching.
Ignored: controls that come back 200
A rejection is at least visible. The quieter failure is a layer that accepts a control and does nothing with it.
Ollama is the clearest case. Its OpenAI compatibility page at 42e911b marks tools as supported and tool_choice as unsupported. parallel_tool_calls is not listed at all. The request struct in openai/openai.go has no field for either, so both are discarded when the JSON is decoded. In our Open Responses audit we saw the same on the Responses endpoint. This time we measured what it does on Chat Completions.
# Ollama 0.30.6, qwen2.5:1.5b on CPU, /v1/chat/completions, 5 runs each,
# temperature 0.7, 5 Oct 2026. Every response was HTTP 200.
control sent prompt result
tool_choice "none" weather in Paris, use the tool 5/5 called get_weather
tool_choice "required" just say hello, use no tool 0/5 called a tool
tool_choice forcing get_time weather in Paris 5/5 called get_weather
parallel_tool_calls false weather in Paris, London, Tokyo 3/5 returned 3 calls
no parallel control same prompt 2/5 returned 3 calls
Every control was dropped. none did not stop a call. required did not force one. The forced function was replaced by the one the prompt suggested. And parallel_tool_calls: false still produced three calls in three of five runs. The model followed the prompt, not the request, and every response was a well-formed 200.
vLLM honors more, with a twist. Its tool calling docs support auto, required, none and named functions. required and named calls use structured outputs, and auto needs the server flags --enable-auto-tool-choice and --tool-call-parser. parallel_tool_calls: false is applied after generation: tool_calls_utils.py keeps the first tool call and drops the rest, which the model had already generated. With tool_choice: "none", vLLM still puts the tool definitions in the prompt unless the operator starts it with --exclude-tools-when-tool-choice-none. One vLLM path does the honest thing: its generic response-template parser rejects strict tools, required or named choices, and parallel_tool_calls: false, because it can parse those outputs but not constrain them.
Anthropic's OpenAI SDK compatibility page lists tool_choice and parallel_tool_calls as fully supported, and the tool-level strict flag as ignored, so "the tool use JSON is not guaranteed to follow the supplied schema." The same page notes that "most unsupported fields are silently ignored rather than producing errors."
OpenRouter's defaults state the same trade openly. Its provider selection docs say that with require_parameters: false, the default, providers that do not support a parameter "can still receive the request, but will ignore unknown parameters." A short list of parameters acts as a soft preference when choosing among providers of one model: tools, response_format and verbosity. tool_choice and parallel_tool_calls are not on it. The parameters reference gives parallel_tool_calls a default of true; the tool calling guide hedges that it is true "for most models."
Gemini's native API, as noted, has no field for disabling parallel calls. A translator targeting it can drop the flag, reject it, or enforce it the way vLLM does, by cutting the response. Google's OpenAI compatibility page does not mention parallel_tool_calls at all.
Dropped: reasoning that does not travel
The third failure mode is not a 400 and not an ignored flag. It is state that disappears. On Claude, that state is thinking blocks.
Anthropic's thinking docs say that within a tool-use turn, "when you return tool results, you must pass the thinking blocks from the assistant message back to the API, complete and unmodified." Modified blocks get a 400. Across models the rule is softer and silent. The preserved thinking page lists which models can read which models' blocks. Claude Opus 5.5 reads blocks from Opus 5 and earlier Opus, Sonnet and Haiku models, and from Sonnet 5.5 on the Claude API and Google Cloud, but not from Fable or Mythos models. When a conversation moves to a model that cannot read a block, "the blocks are dropped, not rejected," and the model runs the next turns without that reasoning.
Two more rules land directly on gateways. First, Claude Sonnet 5.5's thinking blocks "work only in the account that produced them, or in an account linked to it." A gateway that spreads load across several upstream accounts and sends turn two from a different account than turn one loses the reasoning without an error. Second, on Fable 5.1, Opus 5.5 and Sonnet 5.5, a thinking block stays valid only while the system prompt, the set of tools and every earlier message stay unchanged. "Add, remove, rename, or edit a tool in tools" makes later blocks invalid. The default response to a mismatch is a 400, and the check is enforced by default for accounts created on or after 31 August 2026. Changes to tool_choice, by contrast, are outside the check.
That second rule collides with MCP. The MCP spec lets servers send notifications/tools/list_changed when their tool list changes. A client that appends the new tool to tools mid-session breaks the prefix. Anthropic's documented path is to declare tools up front with defer_loading: true and switch them on with tool_addition blocks in messages, so tools never changes. An OpenAI-shaped request has no way to express that. The translation layer has to.
Anthropic's own fallback shows how much bookkeeping this takes. After a mid-output fallback, its docs tell the client to keep the fallback block exactly where it appeared and to drop the client-side tool_use and thinking blocks that came before it on the next turn. And on the OpenAI-compatible path the question never reaches the client: the compatibility page says "the OpenAI SDK doesn't return Claude's thinking," so an agent on that path has no blocks to send back.
The hidden cost of the tool envelope
Every one of these fields costs tokens the agent does not see. OpenAI says function definitions are "injected into the system message," so they "count against the model's context limit and are billed as input tokens." Anthropic adds a tool use system prompt on top of the definitions, and its size depends on tool_choice:
# Anthropic tool use system prompt, added to every request that has tools
# (tool use overview, 5 Oct 2026), in input tokens
model auto / none any / tool
Claude Opus 5.5 286 not supported
Claude Opus 5 286 406
Claude Sonnet 5 354 474
Claude Opus 4.7 675 804
Claude Haiku 4.5 496 588
Forcing a tool on Opus 5 costs 120 more input tokens per request than letting the model choose. Changing tool_choice between requests also invalidates Anthropic's cached message blocks, per its notes on forcing tool use. On vLLM, tool_choice: "none" still spends prompt tokens on the tool definitions unless the server excludes them.
What our own 402 quotes
We probed our walk-up surface the way we did in the last two audits. An unauthenticated POST to /v1/chat/completions returns HTTP 402 with an x402 v2 exact requirement in USDC on Base (eip155:8453), and the PAYMENT-REQUIRED header carries the amount. We kept the model and max_tokens fixed and changed only the tool envelope.
# api.llm4agents.com, unauthenticated POST /v1/chat/completions, 5 Oct 2026
# anthropic/claude-opus-5.5, max_tokens 500, one get_weather tool unless noted
request body 402 amount
one user message, no tools 121 B $0.02
+ get_weather, tool_choice "auto" 355 B $0.02
+ tool_choice "required" (documented 400) 359 B $0.02
+ tool_choice forcing get_weather (documented 400) 406 B $0.02
+ tool_choice forcing a tool not in tools 409 B $0.02
+ parallel_tool_calls false 362 B $0.02
+ tool result with no tool call before it 396 B $0.02
+ call ID with a base64 tail (outside Anthropic's 1,154 B $0.02
documented ID pattern)
tool description padded to ~100 KB 100 KB $0.02
300 tools 58 KB $0.02
~100 KB tool result in history 101 KB $0.14
~100 KB of tool call arguments in history 101 KB $0.14
~100 KB user message 100 KB $0.14
openai/gpt-5, tool_choice "required" 346 B $0.01
models [gpt-5, opus-5.5], tool_choice "required" 403 B $0.02
Three readings.
First, the quote counts history but not the envelope. A 100 KB tool result or 100 KB of call arguments moves the quote from $0.02 to $0.14, exactly like 100 KB of user text. That is right, because tool results are most of an agent's context. But 100 KB of tool description, or 300 tool definitions, is still free, as the structured outputs audit found for schemas. On Claude the unpriced part also includes the tool use system prompt.
Second, the quote issues a payment requirement for requests the upstream documents as invalid. Forced tool use on Opus 5.5, a tool result with nothing before it, and an ID outside Anthropic's pattern all receive the same $0.02 402. Our walk-up post describes the walk-up quote as final, with no refund path. A request we already know will be refused should never get as far as a signature.
Third, the chain is quoted at its most expensive link, which is right for price, but it is not checked against each link's tool dialect. [gpt-5, opus-5.5] with required is quoted $0.02 and is valid on only one link. Our fallback chains post lists four triggers for moving down the chain: context overflow, rate limits, provider errors and moderation. A 400 for an unsupported tool_choice is none of them, and the public contract does not say what the chain does with it.
What it means for LLM4Agents
The fallback chain is the platform's reliability story, and tool calling is what our customers' agents do all day. This audit says the two do not compose by default. A chain that is safe for a one-shot completion can be unsafe in the middle of a tool loop, because the second link inherits a history it did not write and controls it may not honor.
The failures sort into three classes, and each needs a different answer. Rejections (forced tool use on the Claude 5.5 tier, missing Gemini 3 signatures, nine-character Mistral IDs) are knowable before the request leaves. Ignored controls (Ollama's whole tool envelope, Anthropic-compatible strict, OpenRouter's defaults) are knowable per backend but invisible per response. Dropped reasoning (Claude thinking across models and accounts) is knowable only if the gateway tracks which model and account wrote each turn.
The threat is plain. If a fallback can break a loop, careful agent builders will turn fallback off for tool calls and pin one provider, and the chain stops being a reason to use a gateway. The opportunity is the same fact turned around. An agent cannot keep a table of four providers' ID patterns, signature rules and parallel semantics. A gateway that already terminates every call can: mint IDs, carry signatures, pin turns, refuse impossible chains and say what it applied. That is the difference between a proxy and a runtime for agent loops.
Staying on the frontier
In order of cost and urgency. Steps one and two extend the lint and pricing steps from the structured outputs audit to the tool envelope.
1. Validate every link before the 402. Apply Anthropic's own fallback rule to the whole chain: the request must be valid as a direct request to every model named. Reject forced tool_choice for Opus 5.5, Sonnet 5.5, Fable 5.1 and Mythos 5.1 with a 400 that names the model and suggests auto plus strict tools. Never downgrade silently. If an agent opts into a downgrade, report it.
2. Price the envelope. Count tool definitions in the input estimate behind the quote, plus each provider's documented tool use system prompt for the tool_choice actually sent.
3. Own the tool call IDs. Mint gateway IDs that satisfy every dialect at once. Nine letters and digits pass mistral-common's older validators and Anthropic's pattern, and OpenAI's spec puts no format on the field. Map them to upstream IDs per conversation. Never put provider state inside an ID.
4. Carry opaque state on the side. Store Gemini thought signatures and Claude thinking blocks keyed by our IDs, and reattach them only when the next request goes to a model that can read them. Use Google's dummy signature only as a declared degradation, never as a default.
5. Fail over at turn boundaries. Inside a tool loop, keep the turn on the same model and the same upstream account. Gemini validates only the current turn, and Claude drops unreadable blocks instead of failing, so a new user message is the cheap, safe point to switch models. Mid-turn failover should be opt-in.
6. Normalize MCP tools. Rename dotted or long MCP tool names to fit ^[a-zA-Z0-9_-]{1,64}$, check for collisions, and reverse-map every call. Keep tools stable for a session and bring in tools that appear later through deferred loading on Claude, instead of rewriting the list. Our MCP 2026-07-28 post covers the rest of that spec.
7. Say what was applied. Next to X-Model-Used, return the tool_choice and parallel semantics the answering link actually applied: enforced, translated or ignored. Publish per-model tool support in /api/v1/models, as OpenRouter does with supported parameters. Then add tool-loop conformance tests per link to the per-link eval discipline from the fallback post.
Steps one and two protect money. Three to five protect the loop. Six and seven make the gateway something an agent can reason about.
One endpoint for every link in the chain
OpenAI-compatible, paid per call in USDC over x402.
Register your agent