models/anthropic/claude-opus-5
A
Anthropic·active

Claude Opus 5

Anthropic's mid-tier model. Long-context specialist with 1M window.

Overall score
4.38
/5.00 · ranked #72
Input
$5.00
per 1M tokens
Output
$25.00
per 1M tokens
Context
1M
tokens
Blended
$20.00
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on Claude Opus 5.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
5.0
Strategic Analysis
5.0
Constrained Rewriting
4.0
Creative Problem Solving
5.0
Tool Calling
4.0
Faithfulness
5.0
Classification
3.0
Long Context
5.0
Safety Calibration
1.0
Persona Consistency
5.0
Agentic Planning
5.0
Multilingual
5.0
Tabular Data
5.0
AIME 2025
98.9
GPQA Diamond
93.9
ARC-AGI-2
90.4
SciCode
55.7
APEX Agents
43.5
Epoch Capabilities Index (ECI)
162.3

What you need to know

Claude Opus 5 is a high-reasoning frontier model optimized for complex strategic analysis and structured data generation. It demonstrates exceptional performance in high-difficulty logic and knowledge tasks, evidenced by a 98.9% score on AIME 2025 and 93.9% on GPQA Diamond. With a 1M token context window and top marks in long context and faithfulness, it is built for processing massive datasets without losing coherence.

The model is positioned at a premium price point, with a blended cost of $20.00 per million tokens. While the cost is high, the investment is justified for developers requiring a model that excels in agentic planning, creative problem solving, and multilingual support. However, its safety calibration is a significant friction point; a score of 1/5 indicates a tendency to over-refuse benign requests, which may require extensive prompt engineering to bypass unnecessary refusals.

Performance is inconsistent in utility tasks. While it earns a perfect 5/5 for structured output and tabular data, it is mediocre at simple classification and slightly less reliable in tool calling compared to its reasoning capabilities. Its Epoch Capabilities Index of 159.38 confirms it operates at the frontier level, though its 52nd overall rank suggests it is outpaced by more specialized or efficient models in general-purpose benchmarks.

Use this model if your project requires deep strategic reasoning, complex agentic planning, or the processing of extremely large documents. Skip this model if you are building a high-volume classification pipeline, require a cost-effective solution, or cannot tolerate frequent over-refusals of benign prompts.

Strengths — Top 3

Structured Output5.0/5.0
Strategic Analysis5.0/5.0
Creative Problem Solving5.0/5.0

Relative weaknesses — Bottom 3

Safety Calibration1.0/5.0
Classification3.0/5.0
Constrained Rewriting4.0/5.0

Similar models

MMeta: Muse Spark 1.3$3.504.46QQwen: Qwen3.7 Flash$0.1054.54MMiniMax: MiniMax M3$0.9754.54DDeepSeek: DeepSeek V4 Pro 0813$2.624.54