models/openai/gpt-4-1
O
OpenAI·active

GPT-4.1

OpenAI's mid-tier model. Long-context specialist with 1.0M window.

Overall score
4.23
/5.00 · ranked #73
Input
$2.00
per 1M tokens
Output
$8.00
per 1M tokens
Context
1.0M
tokens
Blended
$6.50
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on GPT-4.1.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
4.0
Strategic Analysis
5.0
Constrained Rewriting
5.0
Creative Problem Solving
3.0
Tool Calling
5.0
Faithfulness
5.0
Classification
4.0
Long Context
5.0
Safety Calibration
1.0
Persona Consistency
5.0
Agentic Planning
4.0
Multilingual
5.0
Tabular Data
4.0
SWE-bench Verified
48.5
MATH Level 5
83.0
AIME 2025
38.3
GPQA Diamond
66.9
ARC-AGI-2
0.4
Epoch Capabilities Index (ECI)
137.3

What you need to know

GPT-4.1 is optimized for high-precision execution and massive data ingestion, distinguished primarily by its 1.0M token context window and perfect scores in faithfulness, persona consistency, and tool calling. Its strength lies in reliability and adherence to constraints, making it highly effective for strategic analysis and multilingual tasks where accuracy is non-negotiable.

The model's pricing is high, with a blended cost of $6.50/MTok, placing it in a premium tier. While it delivers top-tier performance in structured tasks and long-context retrieval, its overall rank of 70 out of 130 suggests that its high cost does not always translate to superior general intelligence across all domains. This is evidenced by a significant failure in safety calibration (1/5) and a negligible score on the ARC-AGI-2 benchmark (0.4%).

Technical performance is bifurcated. It excels in quantitative and logical domains, scoring 83% on MATH Level 5 and 66.9% on GPQA Diamond. However, it struggles with creative problem solving and basic classification compared to its other capabilities. The 48.5% SWE-bench Verified score indicates competent but not industry-leading autonomous coding capabilities.

Use this model if your application requires a massive context window, strict adherence to personas, or high-fidelity tool calling for complex strategic workflows. Skip this model if you are budget-constrained, require a model with strong built-in safety guardrails, or need a system capable of novel creative problem solving.

Strengths — Top 3

Strategic Analysis5.0/5.0
Constrained Rewriting5.0/5.0
Tool Calling5.0/5.0

Relative weaknesses — Bottom 3

Safety Calibration1.0/5.0
Creative Problem Solving3.0/5.0
Structured Output4.0/5.0

Similar models

MMistral Medium 3.1$1.604.23XGrok 4.20$2.194.23QQwen: Qwen3.5-35B-A3B$1.004.31DR1 0528$1.744.46