models/deepseek/deepseek-r1
D
DeepSeek·active

R1

DeepSeek's efficiency model. Context window: 64K tokens.

Overall score
4.08
/5.00 · ranked #107
Input
$0.700
per 1M tokens
Output
$2.50
per 1M tokens
Context
64K
tokens
Blended
$2.05
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on R1.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
4.0
Strategic Analysis
5.0
Constrained Rewriting
4.0
Creative Problem Solving
5.0
Tool Calling
4.0
Faithfulness
5.0
Classification
2.0
Long Context
4.0
Safety Calibration
2.0
Persona Consistency
5.0
Agentic Planning
4.0
Multilingual
5.0
Tabular Data
4.0
AIME 2025
53.3
GPQA Diamond
69.2
ARC-AGI-2
1.3
SciCode
35.7
MATH Level 5
93.1
Epoch Capabilities Index (ECI)
139.7

What you need to know

R1 is a high-reasoning model optimized for complex strategic analysis and creative problem solving. Its strongest technical differentiators are in high-level logic and multilingual capabilities, scoring 5/5 in strategic analysis, persona consistency, and multilingual tasks. This is further supported by a 93.1% score on MATH Level 5 and a 69.2% score on GPQA Diamond, indicating a strong capacity for advanced mathematical and scientific reasoning.

Despite its reasoning strengths, the model struggles with basic classification and safety calibration, both scoring 2/5. It also shows a significant failure in abstract reasoning, as evidenced by a 1.3% score on ARC-AGI-2. Developers should expect poor performance when using the model for simple labeling tasks or applications requiring strict safety guardrails.

At a blended cost of $2.05/MTok, R1 is priced competitively for a model capable of high-level reasoning. The 164K context window is sufficient for most long-form documents, and the model maintains a 4/5 rating for long context and structured output, making it a viable option for complex data extraction and agentic planning.

Use this model if your application requires deep mathematical reasoning, multilingual support, or complex strategic planning. Skip this model if your primary use case is text classification or if your project requires high safety calibration and strict content filtering.

Strengths — Top 3

Strategic Analysis5.0/5.0
Creative Problem Solving5.0/5.0
Faithfulness5.0/5.0

Relative weaknesses — Bottom 3

Classification2.0/5.0
Safety Calibration2.0/5.0
Structured Output4.0/5.0

Similar models

XGrok 4.5$5.004.38GGemini 3.1 Pro Preview$9.504.38SStepFun: Step 3.7 Flash$0.9124.15XGrok 4.3$2.194.15