models/deepseek/deepseek-v4-pro-0813
D
DeepSeek·active

DeepSeek: DeepSeek V4 Pro 0813

DeepSeek's mid-tier model. Long-context specialist with 1.0M window.

Overall score
4.54
/5.00 · ranked #53
Input
$0.660
per 1M tokens
Output
$1.98
per 1M tokens
Context
1.0M
tokens
Blended
$1.65
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on DeepSeek: DeepSeek V4 Pro 0813.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
5.0
Strategic Analysis
5.0
Constrained Rewriting
4.0
Creative Problem Solving
5.0
Tool Calling
4.0
Faithfulness
5.0
Classification
4.0
Long Context
5.0
Safety Calibration
2.0
Persona Consistency
5.0
Agentic Planning
5.0
Multilingual
5.0
Tabular Data
5.0
AIME 2025
98.6
GPQA Diamond
91.7
ARC-AGI-2
61.3
SciCode
51.0
APEX Agents
47.3
Epoch Capabilities Index (ECI)
155.4

What you need to know

DeepSeek V4 Pro 0813 is a high-reasoning model characterized by exceptional performance in structured data and strategic analysis. With a 98.6% score on AIME 2025 and a GPQA Diamond score of 91.7%, it competes with frontier models in complex problem solving and technical accuracy. Its 1.0M token context window is supported by a perfect internal score for long-context retrieval, making it suitable for processing massive datasets without losing coherence.

At a blended cost of $1.65 per million tokens, the model offers a high performance-to-price ratio for its capabilities. It excels in agentic planning, multilingual tasks, and tabular data, consistently hitting the top of internal benchmarks. However, developers should note its low safety calibration score of 2/5, which indicates a tendency to over-refuse benign requests rather than a lack of guardrails.

The model is a strong choice for complex engineering tasks, strategic planning, and high-precision structured output. It is less ideal for applications requiring nuanced safety boundaries or highly rigid constrained rewriting.

Use this if you need frontier-level reasoning and massive context windows at a mid-tier price point. Skip this if your application requires a highly calibrated safety filter to avoid false-positive refusals.

Strengths — Top 3

Structured Output5.0/5.0
Strategic Analysis5.0/5.0
Creative Problem Solving5.0/5.0

Relative weaknesses — Bottom 3

Safety Calibration2.0/5.0
Constrained Rewriting4.0/5.0
Tool Calling4.0/5.0

Similar models

MMeta: Muse Glimmer 30B$0.9754.46NNVIDIA: Nemotron 3 Ultra$1.784.46MMeta: Muse Spark 1.3$3.504.46QQwen 3.7 Max$3.694.62