models/mistral/mistral-small-3-1-24b-instruct
M
Mistral·active

Mistral Small 3.1 24B

Mistral's efficiency model. Context window: 128K tokens.

Overall score
2.77
/5.00 · ranked #147
Input
$0.351
per 1M tokens
Output
$0.555
per 1M tokens
Context
128K
tokens
Blended
$0.504
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on Mistral Small 3.1 24B.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
4.0
Strategic Analysis
3.0
Constrained Rewriting
3.0
Creative Problem Solving
2.0
Tool Calling
1.0
Faithfulness
4.0
Classification
3.0
Long Context
5.0
Safety Calibration
1.0
Persona Consistency
2.0
Agentic Planning
3.0
Multilingual
4.0
Tabular Data
1.0

What you need to know

Mistral Small 3.1 24B is specialized for high-fidelity processing of large datasets, distinguished by a perfect 5/5 score in long context handling and a 128K window. Its primary technical strengths lie in faithfulness and structured output, making it reliable for extracting specific information from long documents without introducing hallucinations.

The model fails significantly in operational autonomy and data manipulation. With a 1/5 score in tool calling and tabular data, it cannot be used for agentic workflows that require API interactions or the processing of spreadsheets. Furthermore, a 1/5 in safety calibration indicates a lack of robust guardrails, which may be a liability for public-facing applications.

At a blended cost of $0.504/MTok, the model is priced as a mid-tier option, but its performance is inconsistent. While it handles multilingual tasks and structured formatting well, its low overall rank (#130 of 130) and poor scores in persona consistency and creative problem solving suggest it lacks the general reasoning capabilities of its peers in this price bracket.

Use this model if you need a cost-effective way to summarize or extract structured data from very long texts where factual accuracy is critical. Skip this model if your project requires tool use, tabular analysis, strict safety filtering, or complex creative reasoning.

Strengths — Top 3

Long Context5.0/5.0
Structured Output4.0/5.0
Faithfulness4.0/5.0

Relative weaknesses — Bottom 3

Tool Calling1.0/5.0
Safety Calibration1.0/5.0
Tabular Data1.0/5.0

Similar models

MMiniMax: MiniMax-01$0.8753.08IIBM: Granite 4.0 Micro$0.0882.83IIBM: Granite 4.1 8B$0.0873.54QQwen: Qwen3 Coder 30B A3B Instruct$0.2283.23