DeepSeekDSH
Independent community guideNot affiliated with DeepSeek.Download versions and guide references

DeepSeek Prompt Caching: Slash Coding Agent Token Costs by 80%

How DeepSeek's automatic prefix caching works, exact price differences between cache hits and misses, and how to structure DSH sessions to sustain 80%+ cache hit rates.

On this page
Editorial snapshot · 2026-10-10

In multi-turn coding sessions and agentic loops, input tokens rapidly outpace output tokens as entire repository maps and tool call logs are resent on every turn. DeepSeek provides server-side prompt caching (KV cache reuse) automatically by matching token prefixes, dropping input costs by 75% to 90% without requiring explicit client annotations. Structuring prompts to preserve prefix stability turns long-context agent workflows into an economically sustainable practice.

Server-side prefix matching across 64-token chunks without manual tags

Unlike providers that require explicit cache_control markers in client payloads, DeepSeek implements automatic prompt caching at the gateway layer. When an API request arrives, the cluster hashes the prompt token sequence from beginning to end against existing KV states stored in server memory and SSDs.

Caching operates at a 64-token block granularity. Whenever an incoming request shares an identical prefix with cached states on the serving nodes, the inference engine reuses the precomputed key-value representations. This cuts prefill latency and triggers discounted input billing.

DeepSeek's Multi-head Latent Attention (MLA) architecture underpins this capability by shrinking KV cache memory footprints to one fourth of standard attention models, with secondary offloading to SSD at one eighth the storage requirement. That hardware efficiency allows DeepSeek to keep prefix caching enabled across all public API accounts by default.

Section sources: DeepSeek API Docs — Context Caching (KV Cache) · DeepSeek — Introducing DeepSeek-V4.1-Flash

Cache hit rates cost a fraction of cold prompt misses

DeepSeek separates prompt billing into cache miss and cache hit tiers. For standard DeepSeek-V3 / Chat models, a cache miss costs 2.00 RMB per million tokens, whereas a cache hit drops to 0.50 RMB per million tokens (and as low as 0.10 RMB during off-peak hours). On the high-throughput V4.1 Flash route, cache hit input drops to 0.028 RMB per million tokens.

For one-off prompts, the financial impact remains modest. For an autonomous DSH coding loop that runs for 20 or 30 turns over a substantial codebase, cumulative input tokens regularly exceed 500,000. Under sustained cache hits, over 80% of those tokens qualify for the reduced rate. The table below outlines the official pricing tiers:

Model routeCache hit input (per 1M)Cache miss input (per 1M)Output tokens (per 1M)Input discount
DeepSeek-V3 / Chat (Standard)0.50 RMB ($0.07)2.00 RMB ($0.28)8.00 RMB ($1.10)~75%
DeepSeek-V3 / Chat (Off-peak)0.10 RMB ($0.014)1.00 RMB ($0.14)2.00 RMB ($0.28)Up to 90%
DeepSeek V4.1 Flash0.028 RMB ($0.004)0.14 RMB ($0.020)0.28 RMB ($0.040)~80%
DeepSeek Reasoner (R1)1.00 RMB ($0.14)4.00 RMB ($0.55)16.00 RMB ($2.19)~75%

Figures compiled from DeepSeek's official pricing documentation. Off-peak pricing applies between 00:30 and 08:30 UTC+8. Final billing reflects rates in your DeepSeek console.

Section sources: DeepSeek API Docs — Models and pricing · DeepSeek API Docs — Context Caching (KV Cache)

Structuring prompts for high cache hit rates in DSH

Because DeepSeek enforces strict left-to-right prefix matching, prompt architecture dictates whether caching succeeds. A single dynamic token introduced at the beginning of the prompt invalidates the hash for every subsequent token in the entire payload.

When configuring DeepSeek Harness (DSH) or custom agent pipelines, follow a static-first, append-only convention. Place immutable system guidelines, MCP tool signatures, and repository architecture specifications at the very front of the payload, leaving dynamic tool outputs and user prompts at the tail.

  • Keep system instructions immutable: do not inject dynamic timestamps, random task IDs, or machine status lines at the start of the system message.
  • Freeze MCP tool declarations: maintain consistent tool definitions and parameter order across all turns in a session.
  • Anchor project context early: load repository rules, AGENTS.md, or architectural boundaries immediately after the system prompt so all turns share the same baseline prefix.
  • Use append-only conversation history: append new user messages and tool results to the end of the history array without mutating or reordering prior turns.

Section sources: DeepSeek API Docs — Context Caching (KV Cache) · DeepSeek Harness — Official GitHub Repository

Diagnosing cache misses: usage metrics and architectural pitfalls

Verifying cache efficiency requires inspecting the usage object in each API response. DeepSeek exposes prompt_cache_hit_tokens and prompt_cache_miss_tokens alongside total prompt_tokens. In a healthy multi-turn DSH session, prompt_cache_hit_tokens should account for 80% or more of input tokens from turn two onwards.

Unexpected cache misses usually stem from two causes: idle eviction after prolonged inactivity when the serving cluster reclaims node memory, or naive sliding-window context truncation that slices older lines off the top of the prompt. Effective context pruning compresses intermediate dialogue turns while preserving the top-level prefix intact.

API usage response verification

A response payload reporting prompt_tokens: 32768, prompt_cache_hit_tokens: 28416, and prompt_cache_miss_tokens: 4352 confirms that 86% of input tokens qualified for discounted cache pricing, cutting that turn's input expense by more than two thirds.

Section sources: DeepSeek API Docs — Context Caching (KV Cache) · DeepSeek API Docs — Models and pricing

Continue with model and DSH workflow guides