> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pipellm.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Responses Format Converter

> Use Claude, Gemini, and Chat Completions models with the OpenAI Responses API format

## Overview

The Responses Format Converter lets OpenAI Responses clients — the OpenAI SDK
`client.responses` API, Codex CLI, and other Responses-based tools — call
**Anthropic**, **Gemini**, and **Chat Completions** models. Models that natively
support the Responses API are passed through unchanged on the same route.

<Warning>
  This route is a **stateless, partial** implementation of the Responses API.
  It converts text, images, function calling, reasoning, and streaming. It does
  not provide server-side response storage or OpenAI hosted tools. See
  [Not Supported](#not-supported) before migrating a workflow.
</Warning>

## Configuration

### SDK Configuration

Set your base URL to:

```
https://api.pipellm.ai/responses/v1
```

### cURL / Direct API

```
https://api.pipellm.ai/responses/v1/responses
```

## Usage Examples

<Tabs>
  <Tab title="Python SDK">
    ```python theme={"dark"}
    from openai import OpenAI

    client = OpenAI(
        api_key="your-pipellm-api-key",
        base_url="https://api.pipellm.ai/responses/v1"
    )

    response = client.responses.create(
        model="claude-sonnet-4-6",  # or "gemini-3-flash-preview"
        input="Hello, how are you?",
        store=False,
    )

    print(response.output_text)
    ```
  </Tab>

  <Tab title="TypeScript SDK">
    ```typescript theme={"dark"}
    import OpenAI from "openai";

    const client = new OpenAI({
      apiKey: "your-pipellm-api-key",
      baseURL: "https://api.pipellm.ai/responses/v1",
    });

    const response = await client.responses.create({
      model: "claude-sonnet-4-6", // or "gemini-3-flash-preview"
      input: "Hello, how are you?",
      store: false,
    });

    console.log(response.output_text);
    ```
  </Tab>

  <Tab title="cURL">
    ```bash theme={"dark"}
    curl https://api.pipellm.ai/responses/v1/responses \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer your-pipellm-api-key" \
      -d '{
        "model": "claude-sonnet-4-6",
        "input": "Hello, how are you?",
        "store": false
      }'
    ```
  </Tab>
</Tabs>

## Stateless Multi-Turn

Converted models keep no server-side conversation state. Send `store: false`
(or omit `store`), and carry the history yourself: on every turn, append the
**complete** `response.output` array of the previous response to `input`,
followed by your new items.

```python theme={"dark"}
tools = [{
    "type": "function",
    "name": "get_weather",
    "description": "Get the weather for a city",
    "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
    },
}]

history = [{"role": "user", "content": "What is the weather in Paris?"}]

first = client.responses.create(
    model="gemini-3-flash-preview",
    input=history,
    tools=tools,
    store=False,
    include=["reasoning.encrypted_content"],
)

# Replay every output item as returned: reasoning items (with
# encrypted_content), function calls, and messages.
history += [item.model_dump(exclude_none=True) for item in first.output]

for item in first.output:
    if item.type == "function_call":
        history.append({
            "type": "function_call_output",
            "call_id": item.call_id,
            "output": '{"temperature_c": 18}',
        })

final = client.responses.create(
    model="gemini-3-flash-preview",
    input=history,
    tools=tools,
    store=False,
    include=["reasoning.encrypted_content"],
)
print(final.output_text)
```

Replay rules:

* Do not filter, reorder, or rewrite output items. Reasoning items may have an
  empty `summary` and exist only to carry `encrypted_content`; they must stay
  directly before the function call they were returned with.
* Every `function_call_output` must reference a `call_id` that appears in the
  same request's `input`. Duplicate `call_id` values in one history are rejected
  with `400`.
* `encrypted_content` is opaque provider state — Anthropic thinking signatures
  or Gemini thought signatures. It is only valid for the provider family that
  issued it. Content issued by another provider is ignored rather than replayed.

## Thinking and Signatures

| Target | What `encrypted_content` carries | If it is not replayed |
| - | - | - |
| Anthropic | Each thinking block’s `signature`, or opaque `data` for redacted thinking | The earlier thinking is left out of the upstream request; the conversation continues without it |
| Gemini | The `thoughtSignature` of function call parts and single thought parts | Gemini 3 models reject the function call turn with `400` |

Gemini 3 models validate thought signatures on function calls in the current
turn. This is a native model requirement, not a converter rule; see Google's
[Thought signatures](https://ai.google.dev/gemini-api/docs/generate-content/thought-signatures)
guide. The converter's part is to return each signature inside a reasoning item
and put it back on the exact function call part when you replay the item.

<Warning>
  **Known limitation: Gemini 3 signatures across upstream channels.** PipeLLM
  may serve consecutive requests for the same model through different upstream
  channels, and this route has no channel affinity. In our validation with the
  official OpenAI SDK, the client replayed `encrypted_content` correctly and
  replays that stayed on the same channel succeeded, but a replay that landed on
  a different channel was rejected upstream with
  `400 Thought signature is not valid`. This limitation is not resolved yet:
  a retry may succeed but is not guaranteed, so handle this error in Gemini 3
  tool loops.
</Warning>

### Gemini thinking levels

For Gemini 3 targets, `reasoning.effort` accepts `low`, `medium`, or `high` and maps to `thinkingLevel`; Gemini 2.5 uses `thinkingBudget`. Native model levels such as `minimal` are not exposed by this conversion path. The upstream model must support the chosen level.

## Supported Features

| Feature | Status |
| - | - |
| Text and image input (`input_text`, `input_image` with URL or data URL) | ✅ Supported |
| `instructions`, `system` / `developer` messages | ✅ Supported |
| Streaming (Responses SSE events) | ✅ Supported |
| Function calling, including streamed arguments | ✅ Supported |
| `reasoning.effort`, `reasoning.summary`, `include: ["reasoning.encrypted_content"]` | ✅ Supported |
| `custom`, `shell` / `local_shell`, `apply_patch`, `namespace` tools | ⚠️ Mapped to function tools and restored in the output; custom tool grammars are passed as a description only, not enforced |
| Function tool `strict` | ⚠️ Only effective for OpenAI targets |
| `usage` | ✅ `input_tokens` includes cached tokens, `output_tokens` includes reasoning tokens |

A stream always ends with exactly one terminal event: `response.completed`,
`response.incomplete` (output limit or content filter), or `response.failed`
(upstream error or truncated stream).

`prompt_cache_key`, `client_metadata`, and
`stream_options.include_obfuscation` are accepted but have no effect.

## Not Supported

For converted models, anything that cannot be expressed in the target protocol
returns `400` instead of being silently dropped:

| Feature | Behavior |
| - | - |
| `store: true`, `previous_response_id`, `conversation`, `background: true` | `400` |
| `GET` / `DELETE /responses/{id}` and other stored-response endpoints | `404` |
| Hosted tools: `web_search`, `file_search`, `code_interpreter`, `mcp`, image generation, computer use | `400` |
| Structured outputs (`text.format` other than plain text), `text.verbosity` | `400` |
| `input_file`, image `file_id`, `item_reference` | `400` |
| `truncation: "auto"`, non-empty `metadata`, `tool_choice` with `allowed_tools` | `400` |
| Unknown top-level fields and unknown fields on tools or input items | `400` |
| Gemini targets: `parallel_tool_calls: false`, `reasoning.effort` other than `low` / `medium` / `high` | `400` |

If you need these features with an OpenAI model, use the native
[`/v1/responses`](/api-reference/openai/responses) route.

## Codex CLI

Codex CLI works with this route using a custom provider with
`wire_api = "responses"` and the base URL above. Disable hosted web search for
converted models (`web_search = "disabled"`), because hosted tools are rejected.

## Related Docs

<Columns cols={3}>
  <Card title="Converters Overview" icon="shuffle" href="/converter/overview">
    All converter routes
  </Card>

  <Card title="Native Responses" icon="sparkles" href="/api-reference/openai/responses">
    `/v1/responses` for OpenAI-compatible models
  </Card>

  <Card title="Routing & Protocols" icon="route" href="/guides/routing-protocols">
    Native routes versus converter routes
  </Card>
</Columns>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.