Skip to main content

Overview

The Responses Format Converter lets OpenAI Responses clients — the OpenAI SDK client.responses API, Codex CLI, and other Responses-based tools — call Anthropic, Gemini, and Chat Completions models. Models that natively support the Responses API are passed through unchanged on the same route.
This route is a stateless, partial implementation of the Responses API. It converts text, images, function calling, reasoning, and streaming. It does not provide server-side response storage or OpenAI hosted tools. See Not Supported before migrating a workflow.

Configuration

SDK Configuration

Set your base URL to:

cURL / Direct API

Usage Examples

Stateless Multi-Turn

Converted models keep no server-side conversation state. Send store: false (or omit store), and carry the history yourself: on every turn, append the complete response.output array of the previous response to input, followed by your new items.
Replay rules:
  • Do not filter, reorder, or rewrite output items. Reasoning items may have an empty summary and exist only to carry encrypted_content; they must stay directly before the function call they were returned with.
  • Every function_call_output must reference a call_id that appears in the same request’s input. Duplicate call_id values in one history are rejected with 400.
  • encrypted_content is opaque provider state — Anthropic thinking signatures or Gemini thought signatures. It is only valid for the provider family that issued it. Content issued by another provider is ignored rather than replayed.

Thinking and Signatures

Gemini 3 models validate thought signatures on function calls in the current turn. This is a native model requirement, not a converter rule; see Google’s Thought signatures guide. The converter’s part is to return each signature inside a reasoning item and put it back on the exact function call part when you replay the item.
Known limitation: Gemini 3 signatures across upstream channels. PipeLLM may serve consecutive requests for the same model through different upstream channels, and this route has no channel affinity. In our validation with the official OpenAI SDK, the client replayed encrypted_content correctly and replays that stayed on the same channel succeeded, but a replay that landed on a different channel was rejected upstream with 400 Thought signature is not valid. This limitation is not resolved yet: a retry may succeed but is not guaranteed, so handle this error in Gemini 3 tool loops.

Gemini thinking levels

For Gemini 3 targets, reasoning.effort accepts low, medium, or high and maps to thinkingLevel; Gemini 2.5 uses thinkingBudget. Native model levels such as minimal are not exposed by this conversion path. The upstream model must support the chosen level.

Supported Features

A stream always ends with exactly one terminal event: response.completed, response.incomplete (output limit or content filter), or response.failed (upstream error or truncated stream). prompt_cache_key, client_metadata, and stream_options.include_obfuscation are accepted but have no effect.

Not Supported

For converted models, anything that cannot be expressed in the target protocol returns 400 instead of being silently dropped: If you need these features with an OpenAI model, use the native /v1/responses route.

Codex CLI

Codex CLI works with this route using a custom provider with wire_api = "responses" and the base URL above. Disable hosted web search for converted models (web_search = "disabled"), because hosted tools are rejected.

Converters Overview

All converter routes

Native Responses

/v1/responses for OpenAI-compatible models

Routing & Protocols

Native routes versus converter routes