Skip to main content
When many requests share a long prefix — a system prompt, a tool schema, a document you keep asking about — caching lets the provider reuse that work instead of reprocessing it. Cached input is billed at a much lower rate than fresh input. PipeLLM passes caching through to the upstream provider and meters it separately, so the savings land on your bill. How you turn it on depends on the protocol.

Anthropic: explicit

Anthropic caching is opt-in. Mark the end of the prefix you want cached with cache_control, and everything before that point becomes the cache key.
The first request pays a write price; later requests that match the prefix pay the much cheaper read price. Caches are short-lived — a five-minute window by default, with a one-hour option priced separately. Responses report what happened:
cache_creation_input_tokens means you paid to write. cache_read_input_tokens means you got a hit.
Bedrock and Vertex AI accept at most 4 cache_control blocks. When a request carries more, PipeLLM keeps the last four — the ones closest to the end of the prompt, which are usually the most valuable — and drops the earlier ones before forwarding. The request still succeeds, so a request that works on Anthropic’s own API can quietly cache less here. Keep to four breakpoints if you route across providers.

OpenAI and Gemini: automatic

No request field to set. The provider detects a repeated prefix on its own and applies a discount to the cached portion. OpenAI reports it under usage.prompt_tokens_details.cached_tokens. Because it is automatic, the only thing you control is prompt shape — see below.

Writing cacheable prompts

Caching keys on an exact prefix match, so the rule is simple: put what is stable first and what varies last.
  • Static system prompt, tool definitions, and reference documents at the top.
  • The user’s current turn at the bottom.
  • Nothing variable early on. A timestamp, a request ID, or a shuffled list near the start invalidates the whole prefix and you pay full price every time.
A cache only pays off when it is reused inside its lifetime. One-off requests cost slightly more with an explicit cache write, so reserve cache_control for prefixes you will actually hit again within a few minutes.

Checking the price

Cache rates are per model. Look them up with GET /v1/models/pricing — the response carries them alongside the regular token prices:
A model with no cache object does not support caching on the mapping you are routed to.