DeepSeekDSH
Independent community guideNot affiliated with DeepSeek.Official source snapshot

DeepSeek V4 Pro 0813: What Its Agent Benchmarks Still Measure

Read DeepSeek V4 Pro 0813 correctly: official Harness-based agent scores, current independent evidence, benchmark-version differences, API economics and the planned V4.1 Flash routing change.

Independent editorial review: DeepSeekDSHSource checked: 0.1.5-rc.2 · 2026-09-11
On this page
Editorial snapshot · Sep 12, 2026

V4 Pro remains useful as a benchmark case because it exposes how much the evaluation system matters. DeepSeek reported 87.9 on its Terminal-Bench 2.1 setup, while newer independent Terminal-Bench 4.0 evidence is far lower under a different benchmark and harness contract. Neither number can be used as the other's replacement. DeepSeek currently plans to route the V4 Pro API name to V4.1 Flash on September 14, 2026, so production users should treat V4 Pro as a versioned baseline and re-test after that provider-side change.

Evidence dashboard

V4 Pro: official agent results vs current independent evidence

The same model can look radically different when benchmark version and harness change. Keep those labels attached to every score.

Sep 11, 2026

DeepSeek's August 13 agent results

Public code-agent tasks used DeepSeek Harness Minimal Mode, max effort, top_p 0.95 and temperature 1.0.

Terminal-Bench 2.187.9
DeepSWE62.7
NL2Repo61.5
Toolathlon-Verified74.1
DSBench-Hard67.2

DeepSeek's 0813 table measured a model-and-harness system

The August 13 GA release focused heavily on agent behavior. DeepSeek reported substantial gains in production-oriented coding, repository and automation tasks.

BenchmarkV4 Pro 0813
HLE, no tools / with tools42.7 / 60.0
Terminal Bench 2.187.9
NL2Repo61.5
DeepSWE62.7
Toolathlon-Verified74.1
AutomationBench (Public)31.8
DSBench-Hard67.2

Section sources: DeepSeek — V4 Pro GA release · DeepSeek API — Change Log

Current independent evidence puts V4 Pro below the newest frontier tier

Artificial Analysis v4.3 scores DeepSeek V4 Pro 0813 at 36 on its Intelligence Index. That is well above the overall model median reported on the model page, but below the current frontier leaders such as GPT-6 Astra and Claude Fable 5.1, both of which score 53 in the same index version.

The value proposition is therefore not simply 'best model'. V4 Pro combines capable reasoning and agent performance with a dramatically lower token price than the most expensive closed models, plus open weights and a 1M context window.

Section sources: Artificial Analysis — DeepSeek V4 Pro 0813 · Artificial Analysis — Intelligence Index v4.3

The Harness and reasoning settings are part of the result

DeepSeek states that public Code Agent tasks for V4 Pro 0813 were evaluated using DeepSeek Harness Minimal Mode with max reasoning effort, top_p 0.95 and temperature 1.0. That means the published result is a model-plus-harness result, not a context-free property of the raw model endpoint.

If another vendor reports a score through a different coding harness, tool policy or benchmark revision, putting the two numbers side by side can create a false ranking. The correct comparison keeps benchmark version, scaffold and effort as consistent as possible.

Useful for this site

Because DeepSeek itself used DeepSeek Harness for several public Code Agent evaluations, DSH is not just an interface around the model. The preset and tool surface are part of the measured system.

Section sources: DeepSeek — V4 Pro GA release · DeepSeek API — Change Log

V4 Pro's cost and 1M context explain why it was attractive in production

Before the scheduled September 14 routing change, the current DeepSeek price table shows a 1M context window, up to 384K output tokens, peak cache-miss input at $1.32 per million tokens and output at $3.96. Off-peak rates are half of peak rates.

ItemV4 Pro 0813
Context1M tokens
Max output384K tokens
Peak cache-miss input$1.32 / 1M
Peak output$3.96 / 1M
Off-peak cache-miss input$0.66 / 1M
Off-peak output$1.98 / 1M

Section sources: DeepSeek API — Models & Pricing

Its lasting value is long-context, tool-heavy work at a much lower price point

V4 Pro's strongest case was long-context, tool-heavy engineering where per-token economics matter. A 1M window reduces the need to aggressively trim repository context, while the agent benchmarks show that the model can operate effectively when the harness exposes shell, editing and other tools.

It was also a useful open-weight reference point: teams could compare a strong DeepSeek model against proprietary models without accepting frontier-model token pricing.

Section sources: DeepSeek — V4 preview release · Artificial Analysis — DeepSeek V4 Pro 0813

Treat V4 Pro as a versioned baseline while the V4.1 routing change remains scheduled

For a brand-new DeepSeek integration on September 10, 2026, the answer is usually no. DeepSeek now says V4.1 Flash beats V4 Pro on performance, cost, speed and total time, and the V4 Pro model name is scheduled to route to V4.1 Flash on September 14.

V4 Pro is still worth documenting because it is the last stable Pro baseline with a detailed official benchmark set. Use it to understand the progression of DeepSeek's agent performance, but treat V4.1 Flash as the current deployment direction.

Section sources: DeepSeek API — Models & Pricing

Related model analysis