DeepSeekDSH
Independent community guideNot affiliated with DeepSeek.Download versions and guide references

DeepSeek Prompt Caching and API Cost Optimization Guide

DeepSeek automatically caches identical prompt prefixes on the server, providing up to a 90% discount on cache hits. Learn architectural prefix rules and token audit scripts for production efficiency.

On this page

1. Cache mechanics & billing rules

DeepSeek implements automatic server-side Context KV Cache (Prompt Caching). No custom headers or endpoints are required to benefit from caching:

Cache Miss

Initial prompts that have not been cached are billed at the standard input token rate while populating the server prefix cache.

Cache Hit (90% Discount)

Subsequent requests sharing an exact sequential prefix pay only 10% of the standard input token price, while significantly reducing first-token latency.

2. Key prefix engineering rules

Rule 1: Strict sequential prefix match

Caching matches sequentially from character 0. Inserting dynamic data (like timestamps such as Current Time: 2026-10-10) at the start invalidates all subsequent cached context.

Rule 2: Static content first, dynamic last

Anchor large system prompts, API tool schemas, and repository documentation at the beginning of messages; append changing user questions at the end.

3. Off-peak discount scheduling

DeepSeek provides an additional 50% discount during off-peak hours (00:30 ~ 08:30 UTC+8).

Batch Processing Pattern: For offline code audits, data transformation pipelines, and batch document translations, schedule tasks during off-peak windows. Combining off-peak rates with Prompt Caching can reduce overall costs to around 5% of peak baseline prices.

4. Python cache hit ratio & token audit script

Inspect cache statistics from the API response usage object:

from openai import OpenAI client = OpenAI( api_key="sk-your_deepseek_api_key", base_url="https://api.deepseek.com" ) # Simulate long static context system_prompt = "You are a code reviewer. " + ("Standard rules... " * 100) response = client.chat.completions.create( model="deepseek-chat", messages=[ {"role": "system", "content": system_prompt}, {"role": "user", "content": "Audit code style."} ] ) # Extract token usage breakdown usage = response.usage prompt_tokens = usage.prompt_tokens cache_hit_tokens = getattr(usage, "prompt_cache_hit_tokens", 0) cache_miss_tokens = getattr(usage, "prompt_cache_miss_tokens", prompt_tokens - cache_hit_tokens) print(f"Total prompt tokens: {prompt_tokens}") print(f"Cache hit tokens: {cache_hit_tokens} (90% discount applied)") print(f"Cache miss tokens: {cache_miss_tokens}") if prompt_tokens > 0: hit_ratio = (cache_hit_tokens / prompt_tokens) * 100 print(f"Cache hit ratio: {hit_ratio:.2f}%")

Troubleshooting

Why is prompt_cache_hit_tokens always 0?

Server-side caching is optimized for prompts exceeding a threshold (typically 64 to 1024 tokens). Short messages or dynamically modified system prompts will not trigger cache hits.

Cache miss during multi-turn conversation

Check if the client application inserts unique session IDs, timestamps, or alters previous turn message ordering.

References