DeepSeek V4 Pro 0813: What Its Agent Benchmarks Still Measure
Read DeepSeek V4 Pro 0813 correctly: official Harness-based agent scores, current independent evidence, benchmark-version differences, API economics and the planned V4.1 Flash routing change.
On this page
V4 Pro remains useful as a benchmark case because it exposes how much the evaluation system matters. DeepSeek reported 87.9 on its Terminal-Bench 2.1 setup, while newer independent Terminal-Bench 4.0 evidence is far lower under a different benchmark and harness contract. Neither number can be used as the other's replacement. DeepSeek currently plans to route the V4 Pro API name to V4.1 Flash on September 14, 2026, so production users should treat V4 Pro as a versioned baseline and re-test after that provider-side change.
DeepSeek's 0813 table measured a model-and-harness system
The August 13 GA release focused heavily on agent behavior. DeepSeek reported substantial gains in production-oriented coding, repository and automation tasks.
| Benchmark | V4 Pro 0813 |
|---|---|
| HLE, no tools / with tools | 42.7 / 60.0 |
| Terminal Bench 2.1 | 87.9 |
| NL2Repo | 61.5 |
| DeepSWE | 62.7 |
| Toolathlon-Verified | 74.1 |
| AutomationBench (Public) | 31.8 |
| DSBench-Hard | 67.2 |
Section sources: DeepSeek — V4 Pro GA release · DeepSeek API — Change Log
Current independent evidence puts V4 Pro below the newest frontier tier
Artificial Analysis v4.3 scores DeepSeek V4 Pro 0813 at 36 on its Intelligence Index. That is well above the overall model median reported on the model page, but below the current frontier leaders such as GPT-6 Astra and Claude Fable 5.1, both of which score 53 in the same index version.
The value proposition is therefore not simply 'best model'. V4 Pro combines capable reasoning and agent performance with a dramatically lower token price than the most expensive closed models, plus open weights and a 1M context window.
Section sources: Artificial Analysis — DeepSeek V4 Pro 0813 · Artificial Analysis — Intelligence Index v4.3
The Harness and reasoning settings are part of the result
DeepSeek states that public Code Agent tasks for V4 Pro 0813 were evaluated using DeepSeek Harness Minimal Mode with max reasoning effort, top_p 0.95 and temperature 1.0. That means the published result is a model-plus-harness result, not a context-free property of the raw model endpoint.
If another vendor reports a score through a different coding harness, tool policy or benchmark revision, putting the two numbers side by side can create a false ranking. The correct comparison keeps benchmark version, scaffold and effort as consistent as possible.
Because DeepSeek itself used DeepSeek Harness for several public Code Agent evaluations, DSH is not just an interface around the model. The preset and tool surface are part of the measured system.
Section sources: DeepSeek — V4 Pro GA release · DeepSeek API — Change Log
V4 Pro's cost and 1M context explain why it was attractive in production
Before the scheduled September 14 routing change, the current DeepSeek price table shows a 1M context window, up to 384K output tokens, peak cache-miss input at $1.32 per million tokens and output at $3.96. Off-peak rates are half of peak rates.
| Item | V4 Pro 0813 |
|---|---|
| Context | 1M tokens |
| Max output | 384K tokens |
| Peak cache-miss input | $1.32 / 1M |
| Peak output | $3.96 / 1M |
| Off-peak cache-miss input | $0.66 / 1M |
| Off-peak output | $1.98 / 1M |
Section sources: DeepSeek API — Models & Pricing
Its lasting value is long-context, tool-heavy work at a much lower price point
V4 Pro's strongest case was long-context, tool-heavy engineering where per-token economics matter. A 1M window reduces the need to aggressively trim repository context, while the agent benchmarks show that the model can operate effectively when the harness exposes shell, editing and other tools.
It was also a useful open-weight reference point: teams could compare a strong DeepSeek model against proprietary models without accepting frontier-model token pricing.
Section sources: DeepSeek — V4 preview release · Artificial Analysis — DeepSeek V4 Pro 0813
Treat V4 Pro as a versioned baseline while the V4.1 routing change remains scheduled
For a brand-new DeepSeek integration on September 10, 2026, the answer is usually no. DeepSeek now says V4.1 Flash beats V4 Pro on performance, cost, speed and total time, and the V4 Pro model name is scheduled to route to V4.1 Flash on September 14.
V4 Pro is still worth documenting because it is the last stable Pro baseline with a detailed official benchmark set. Use it to understand the progression of DeepSeek's agent performance, but treat V4.1 Flash as the current deployment direction.
Section sources: DeepSeek API — Models & Pricing
DeepSeek — V4 Pro GA release · DeepSeek API — Change Log · Artificial Analysis — DeepSeek V4 Pro 0813 · Artificial Analysis — Intelligence Index v4.3 · DeepSeek API — Models & Pricing · DeepSeek — V4 preview release