Tokenify

Prompt caching

Nothing to enable and nothing to call. Repeated prefixes are recognised upstream, and the discount reaches your invoice.

What happens

When two requests begin with the same tokens — a system prompt, a tool schema, a document you keep re-sending — the supplier serves the shared prefix from its cache instead of recomputing it, and charges a fraction of the input rate for those tokens. We pass that through. On DeepSeek V4.1 Flash a cached token costs about one twenty-fifth of a fresh one. The ratio is the model vendor’s, not ours, so it differs per model — across the catalogue a cache read runs between a third and a twenty-fifth of the input rate.

Most resellers flatten cache pricing into one headline rate and keep the difference. Our invoice lines split fresh and cached tokens and price them separately, so you can check the arithmetic yourself.

Seeing it

Every response reports the split:

"usage": {
  "prompt_tokens": 4890,
  "prompt_tokens_details": { "cached_tokens": 4608 },
  "completion_tokens": 240
}

That is a real measurement: the second of two identical 4,890-token prompts came back with 94% of the prompt served from cache. On the Anthropic endpoint the same figure arrives as cache_read_input_tokens on message_delta.

Earning more hits

Caching matches on a prefix, so the rule is simple: put what does not change first, and what changes last.

  • Keep the system prompt byte-identical. A timestamp, a request id or a shuffled tool list at the top of the prompt invalidates everything after it.
  • Put documents before the question. A long retrieved context followed by a short query caches the expensive half.
  • Order tool schemas deterministically. Serialising a map in Go or Python can reorder keys between runs, which changes the bytes and misses the cache.
  • Agent loops benefit most. Every turn re-sends the whole conversation, so from the second turn onward the bulk of each prompt is a cache hit.

What it costs to store

Nothing. We use implicit caching — the supplier recognises the repeated prefix on its own — so there is no cache to create, no cache to keep alive and no storage charge. You are billed for hits when they happen and nothing when they do not.

Cache misses are not failures

A cache has a lifetime, and a prefix that is not seen again for a while is evicted. A miss simply bills at the ordinary input rate; there is nothing to handle and no error to catch. Rates for each model are on the models page.

Last updated 2026-09-28.