models/inception/mercury-2-5-preview
I
inception·active

Inception: Mercury 2.5 Preview

inception's mid-tier model. Context window: 260K tokens.

Overall score
4.54
/5.00 · ranked #42
Input
$0.040
per 1M tokens
Output
$0.150
per 1M tokens
Context
260K
tokens
Blended
$0.122
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on Inception: Mercury 2.5 Preview.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
5.0
Strategic Analysis
4.0
Constrained Rewriting
4.0
Creative Problem Solving
5.0
Tool Calling
5.0
Faithfulness
5.0
Classification
4.0
Long Context
5.0
Safety Calibration
2.0
Persona Consistency
5.0
Agentic Planning
5.0
Multilingual
5.0
Tabular Data
5.0

What you need to know

Inception: Mercury 2.5 Preview is optimized for high-reliability agentic workflows, excelling in structured output, tool calling, and agentic planning. With a 260K context window and perfect internal scores for long-context handling, tabular data, and multilingual support, it is built for complex data extraction and automation tasks that require strict adherence to formatting.

The model is priced aggressively for its performance tier, with a blended cost of $0.122/MTok. This makes it a high-value option for developers who need frontier-level capabilities in creative problem solving and faithfulness without the premium cost associated with top-tier proprietary models.

The primary trade-off is in safety calibration, where it scores a 2/5. In practice, this indicates a tendency to over-refuse benign prompts, which may introduce friction in user-facing applications. While it performs well in strategic analysis and rewriting, these are its weakest relative areas compared to its perfect scores in technical execution.

Use this model if you are building autonomous agents, complex data pipelines, or multilingual tools that require precise structured output. Skip this model if your application requires a highly permissive safety profile or if you cannot tolerate false-positive refusals in your user interactions.

Strengths — Top 3

Structured Output5.0/5.0
Creative Problem Solving5.0/5.0
Tool Calling5.0/5.0

Relative weaknesses — Bottom 3

Safety Calibration2.0/5.0
Strategic Analysis4.0/5.0
Constrained Rewriting4.0/5.0

Similar models

QQwen 3.7 Max$3.694.62MMiniMax: MiniMax M3$0.9754.54DDeepSeek: DeepSeek V4 Pro 0813$2.794.54OGPT-5$7.814.54