
Cache Read
Cache Read refers to the reuse of already computed intermediate results when operating an AI language model. This saves computing time and costs because unchanging parts of a request don't have to be processed anew every time.
When an AI model processes a text, it computes a series of intermediate values for each word — in technical jargon, key-value pairs, or KV for short. These values cost computing time. If a text changes only slightly, for example because a long system prompt always stays the same, it would be wasteful to recompute the same values every time. This is exactly where cache read comes in: the system stores the already computed intermediate values and, the next time around, simply reads them out of this temporary storage — the so-called cache — instead of generating them anew. The result is the same, but the work only has to be done once.
Cost and speed in model operation
AI models are usually not used just once, but many thousands of times per minute. Each individual request costs computing time on expensive hardware. API providers — that is, services that provide access to a model for a fee — therefore often bill separately: new computations cost more, while a cache read costs less or nothing at all.
Anthropic, the maker of the Claude model, makes this particularly transparent: a so-called cache write — the initial storing of the intermediate values — costs around 25% more there than a normal computation. A subsequent cache read, on the other hand, costs only about 10% of the original price. So anyone sending many requests with identical text segments can significantly reduce their operating costs.
When cache read applies
For a cache read to be possible at all, the beginning of a new input must match the beginning of a previous one. The cached portion must therefore be a shared prefix — a text that precedes the actual new content and has not changed. The model then only computes the new, changed remainder anew; the rest it reads from the cache.
A typical example: a developer builds a chatbot that first receives a long set of rules with every request — such as the bot's personality description or an extensive knowledge base. This part is identical for every user. Without caching, the model would read and compute it a thousand times. With cache read, this happens only once. The individual user question, on the other hand, is new each time and is actually computed anew.
The cache has a limited lifespan. With many providers, stored values expire after a few minutes of inactivity. This means: caching is less worthwhile for sporadic requests, but all the more so for consistently high volume.
Cache read in products and news
In practice, the term appears mainly in the technical documentation of AI API providers — that is, in the guides that explain how developers integrate the model into their own programs. Anthropic and OpenAI have introduced prompt caching as an explicit feature and show in their billing how many tokens — that is, units of text — were processed via cache read.
In financial news surrounding AI companies, the term is relevant when it comes to operating costs and margins. The more efficiently a provider can serve requests through caching, the cheaper it can offer its product — or the higher its profit at the same price. Cache read is therefore not just a technical detail, but a lever in the cost calculations of the entire AI industry.
For ordinary users of a chatbot, the cache is invisible. At most, it manifests itself in shorter response times when a long context has already been cached. However, anyone who develops or evaluates AI services themselves will almost inevitably encounter this term.