models/mistral/devstral-2512
M
Mistral·active

Devstral 2 2512

Mistral's efficiency model. Context window: 262K tokens.

Overall score
4.00
/5.00 · ranked #115
Input
$0.440
per 1M tokens
Output
$2.20
per 1M tokens
Context
262K
tokens
Blended
$1.76
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on Devstral 2 2512.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
5.0
Strategic Analysis
4.0
Constrained Rewriting
5.0
Creative Problem Solving
4.0
Tool Calling
4.0
Faithfulness
4.0
Classification
3.0
Long Context
5.0
Safety Calibration
2.0
Persona Consistency
4.0
Agentic Planning
4.0
Multilingual
5.0
Tabular Data
3.0

What you need to know

Devstral 2 2512 is optimized for high-precision formatting and long-form data processing. It achieves perfect internal scores in constrained rewriting, structured output, and long context handling, making it highly reliable for tasks requiring strict adherence to schemas or the analysis of large documents within its 262K context window. Its multilingual capabilities are similarly strong, ranking it as a versatile tool for global deployments.

Pricing is competitive for a specialized model, with a blended cost of $1.60/MTok. While the input cost is low at $0.40/MTok, the output cost is five times higher, which may impact the budget for applications generating extensive text. Despite these strengths, the model ranks 98th out of 130 overall, indicating that its high performance in structured tasks does not translate to general-purpose dominance.

The model exhibits significant weaknesses in safety calibration and basic classification, suggesting it lacks the guardrails and nuance required for sensitive user-facing moderation or simple labeling tasks. It also struggles with tabular data, meaning it should not be relied upon for complex spreadsheet-style reasoning or data extraction from tables.

Use this model if your workflow requires strict JSON/structured outputs, multilingual support, or processing very large contexts. Skip this model if you need a general-purpose assistant with high safety standards or a tool for tabular data analysis and classification.

Strengths — Top 3

Structured Output5.0/5.0
Constrained Rewriting5.0/5.0
Long Context5.0/5.0

Relative weaknesses — Bottom 3

Safety Calibration2.0/5.0
Classification3.0/5.0
Tabular Data3.0/5.0

Similar models

OGPT-5.4 Nano$0.9884.15QQwen: Qwen3 235B A22B Instruct 2507$0.2844.08OGPT-4.1 Mini$1.303.92XGrok 4.3$2.194.15