DeepSeek Prompt Caching and API Cost Optimization Guide
DeepSeek automatically caches identical prompt prefixes on the server, providing up to a 90% discount on cache hits. Learn architectural prefix rules and token audit scripts for production efficiency.
On this page
1. Cache mechanics & billing rules
DeepSeek implements automatic server-side Context KV Cache (Prompt Caching). No custom headers or endpoints are required to benefit from caching:
Cache Miss
Initial prompts that have not been cached are billed at the standard input token rate while populating the server prefix cache.
Cache Hit (90% Discount)
Subsequent requests sharing an exact sequential prefix pay only 10% of the standard input token price, while significantly reducing first-token latency.
2. Key prefix engineering rules
Rule 1: Strict sequential prefix match
Caching matches sequentially from character 0. Inserting dynamic data (like timestamps such as Current Time: 2026-10-10) at the start invalidates all subsequent cached context.
Rule 2: Static content first, dynamic last
Anchor large system prompts, API tool schemas, and repository documentation at the beginning of messages; append changing user questions at the end.
3. Off-peak discount scheduling
DeepSeek provides an additional 50% discount during off-peak hours (00:30 ~ 08:30 UTC+8).
Batch Processing Pattern: For offline code audits, data transformation pipelines, and batch document translations, schedule tasks during off-peak windows. Combining off-peak rates with Prompt Caching can reduce overall costs to around 5% of peak baseline prices.
4. Python cache hit ratio & token audit script
Inspect cache statistics from the API response usage object:
from openai import OpenAI
client = OpenAI(
api_key="sk-your_deepseek_api_key",
base_url="https://api.deepseek.com"
)
# Simulate long static context
system_prompt = "You are a code reviewer. " + ("Standard rules... " * 100)
response = client.chat.completions.create(
model="deepseek-chat",
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": "Audit code style."}
]
)
# Extract token usage breakdown
usage = response.usage
prompt_tokens = usage.prompt_tokens
cache_hit_tokens = getattr(usage, "prompt_cache_hit_tokens", 0)
cache_miss_tokens = getattr(usage, "prompt_cache_miss_tokens", prompt_tokens - cache_hit_tokens)
print(f"Total prompt tokens: {prompt_tokens}")
print(f"Cache hit tokens: {cache_hit_tokens} (90% discount applied)")
print(f"Cache miss tokens: {cache_miss_tokens}")
if prompt_tokens > 0:
hit_ratio = (cache_hit_tokens / prompt_tokens) * 100
print(f"Cache hit ratio: {hit_ratio:.2f}%")Troubleshooting
Why is prompt_cache_hit_tokens always 0?
Server-side caching is optimized for prompts exceeding a threshold (typically 64 to 1024 tokens). Short messages or dynamically modified system prompts will not trigger cache hits.
Cache miss during multi-turn conversation
Check if the client application inserts unique session IDs, timestamps, or alters previous turn message ordering.