When many requests share a long prefix — a system prompt, a tool schema, a document you
keep asking about — caching lets the provider reuse that work instead of reprocessing it.
Cached input is billed at a much lower rate than fresh input.
PipeLLM passes caching through to the upstream provider and meters it separately, so the
savings land on your bill. How you turn it on depends on the protocol.
Anthropic: explicit
Anthropic caching is opt-in. Mark the end of the prefix you want cached with
cache_control, and everything before that point becomes the cache key.
The first request pays a write price; later requests that match the prefix pay the much
cheaper read price. Caches are short-lived — a five-minute window by default, with a
one-hour option priced separately.
Responses report what happened:
cache_creation_input_tokens means you paid to write. cache_read_input_tokens means you
got a hit.
Bedrock and Vertex AI accept at most 4 cache_control blocks. When a request carries
more, PipeLLM keeps the last four — the ones closest to the end of the prompt, which are
usually the most valuable — and drops the earlier ones before forwarding. The request
still succeeds, so a request that works on Anthropic’s own API can quietly cache less
here. Keep to four breakpoints if you route across providers.
OpenAI and Gemini: automatic
No request field to set. The provider detects a repeated prefix on its own and applies a
discount to the cached portion. OpenAI reports it under
usage.prompt_tokens_details.cached_tokens.
Because it is automatic, the only thing you control is prompt shape — see below.
Writing cacheable prompts
Caching keys on an exact prefix match, so the rule is simple: put what is stable first
and what varies last.
- Static system prompt, tool definitions, and reference documents at the top.
- The user’s current turn at the bottom.
- Nothing variable early on. A timestamp, a request ID, or a shuffled list near the start
invalidates the whole prefix and you pay full price every time.
A cache only pays off when it is reused inside its lifetime. One-off requests cost slightly
more with an explicit cache write, so reserve cache_control for prefixes you will
actually hit again within a few minutes.
Checking the price
Cache rates are per model. Look them up with
GET /v1/models/pricing — the response carries them
alongside the regular token prices:
A model with no cache object does not support caching on the mapping you are routed to.