> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pipellm.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Prompt Caching

> Reuse a long prefix across requests and pay less for it.

When many requests share a long prefix — a system prompt, a tool schema, a document you
keep asking about — caching lets the provider reuse that work instead of reprocessing it.
Cached input is billed at a much lower rate than fresh input.

PipeLLM passes caching through to the upstream provider and meters it separately, so the
savings land on your bill. How you turn it on depends on the protocol.

## Anthropic: explicit

Anthropic caching is opt-in. Mark the end of the prefix you want cached with
`cache_control`, and everything before that point becomes the cache key.

```json theme={"dark"}
{
  "model": "claude-sonnet-4-6",
  "system": [
    {
      "type": "text",
      "text": "<a long style guide>",
      "cache_control": { "type": "ephemeral" }
    }
  ],
  "messages": [{ "role": "user", "content": "Rewrite this paragraph: ..." }]
}
```

The first request pays a **write** price; later requests that match the prefix pay the much
cheaper **read** price. Caches are short-lived — a five-minute window by default, with a
one-hour option priced separately.

Responses report what happened:

```json theme={"dark"}
"usage": {
  "cache_creation_input_tokens": 12043,
  "cache_read_input_tokens": 0
}
```

`cache_creation_input_tokens` means you paid to write. `cache_read_input_tokens` means you
got a hit.

<Warning>
  **Bedrock and Vertex AI accept at most 4 `cache_control` blocks.** When a request carries
  more, PipeLLM keeps the last four — the ones closest to the end of the prompt, which are
  usually the most valuable — and drops the earlier ones before forwarding. The request
  still succeeds, so a request that works on Anthropic's own API can quietly cache less
  here. Keep to four breakpoints if you route across providers.
</Warning>

## OpenAI and Gemini: automatic

No request field to set. The provider detects a repeated prefix on its own and applies a
discount to the cached portion. OpenAI reports it under
`usage.prompt_tokens_details.cached_tokens`.

Because it is automatic, the only thing you control is prompt shape — see below.

## Writing cacheable prompts

Caching keys on an exact prefix match, so the rule is simple: **put what is stable first
and what varies last.**

* Static system prompt, tool definitions, and reference documents at the top.
* The user's current turn at the bottom.
* Nothing variable early on. A timestamp, a request ID, or a shuffled list near the start
  invalidates the whole prefix and you pay full price every time.

A cache only pays off when it is reused inside its lifetime. One-off requests cost slightly
*more* with an explicit cache write, so reserve `cache_control` for prefixes you will
actually hit again within a few minutes.

## Checking the price

Cache rates are per model. Look them up with
[`GET /v1/models/pricing`](/api-reference/model-pricing) — the response carries them
alongside the regular token prices:

```json theme={"dark"}
"pricing": {
  "kind": "token",
  "text": { "prompt": "0.000003", "completion": "0.000015" },
  "cache": {
    "read": "0.0000003",
    "write": "0.00000375",
    "write_1h": "0.0000075"
  }
}
```

A model with no `cache` object does not support caching on the mapping you are routed to.

## Related

* [Get model pricing](/api-reference/model-pricing) — authoritative cache rates
* [Anthropic Messages](/api-reference/anthropic/messages) — where `cache_control` goes


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.