prompt caching
Technique that stores the repeated opening part of a prompt so later requests to a language model can reuse it.
Prompt caching is a technique in large language model APIs that stores the part of a prompt repeated across requests, such as system instructions, long documents or earlier conversation history, so later requests can reuse it. Because the model does not reprocess that content from scratch, it lowers the cost of input tokens and reduces response latency. Google's Gemini API calls the same feature context caching.
AI companies including Anthropic, OpenAI and Google offer it as an API feature. Cached content expires after a period without use, after which it must be processed again. It is widely used by chatbots and coding agents that carry on long conversations.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
A forked sub-agent can reuse the parent session’s prompt cache—stored processing of input the model has already read—reducing the cost of an auxiliary check whi…
The DeepSeek V4 Flash run used prompt context caching, which reuses previously processed input; the evaluation recorded a 98% cache hit rate across the open-wei…
A prompt cache stores reusable context so an application does not have to process it in full on every request.