prompt caching

Technique that stores the repeated opening part of a prompt so later requests to a language model can reuse it.

4 articles
Last mentioned

Prompt caching is a technique in large language model APIs that stores the part of a prompt repeated across requests, such as system instructions, long documents or earlier conversation history, so later requests can reuse it. Because the model does not reprocess that content from scratch, it lowers the cost of input tokens and reduces response latency. Google's Gemini API calls the same feature context caching.

AI companies including Anthropic, OpenAI and Google offer it as an API feature. Cached content expires after a period without use, after which it must be processed again. It is widely used by chatbots and coding agents that carry on long conversations.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

…cky sessions matter because moving a conversation between accounts causes prompt cache misses, meaning the provider cannot reuse the stored earlier parts of the…

A forked sub-agent can reuse the parent session’s prompt cache—stored processing of input the model has already read—reducing the cost of an auxiliary check whi…

The DeepSeek V4 Flash run used prompt context caching, which reuses previously processed input; the evaluation recorded a 98% cache hit rate across the open-wei…

A prompt cache stores reusable context so an application does not have to process it in full on every request.


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.