models/openai/gpt-5-5
O
OpenAI·active

GPT-5.5

OpenAI's mid-tier model. Long-context specialist with 1.1M window.

Overall score
4.46
/5.00 · ranked #46
Input
$5.00
per 1M tokens
Output
$30.00
per 1M tokens
Context
1.1M
tokens
Blended
$23.75
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on GPT-5.5.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
4.0
Strategic Analysis
4.0
Constrained Rewriting
5.0
Creative Problem Solving
4.0
Tool Calling
4.0
Faithfulness
5.0
Classification
4.0
Long Context
5.0
Safety Calibration
5.0
Persona Consistency
5.0
Agentic Planning
4.0
Multilingual
4.0
Tabular Data
5.0
SWE-bench Verified
80.6
AIME 2025
100.0
GPQA Diamond
94.0
ARC-AGI-2
85.0
SciCode
56.1
APEX Agents
38.5
Epoch Capabilities Index (ECI)
158.9

What you need to know

GPT-5.5 distinguishes itself through strict adherence to constraints and instruction fidelity rather than expansive reasoning breadth. It achieves perfect internal scores for constrained rewriting, faithfulness, and long-context processing, which translates to reliable performance when working with extensive documents or rigid formatting requirements. The 1.1M token context window supports processing large codebases or research corpora in a single pass, and the 100% score on AIME 2025 indicates mathematical and logical reasoning remains robust even at scale. This makes the model particularly effective for workflows that demand exact compliance over exploratory generation.

At $5.00/MTok input and $30.00/MTok output, the blended cost sits at $23.75/MTok, placing it firmly in the premium pricing tier. The #11 overall rank among 71 evaluated models reflects this trade-off: you are paying for consistency and capacity rather than absolute top-tier performance across every category. The 80.6% score on SWE-bench Verified confirms strong software engineering capabilities, but the 4/5 internal ratings for structured output, strategic analysis, and creative problem solving show measurable gaps when tasks require open-ended ideation or complex multi-step planning. The pricing is justified for pipelines that prioritize deterministic outputs and long-context retention, but less efficient for high-volume exploration or highly structured data formatting where mid-tier alternatives perform comparably.

Use this if your workload requires processing documents beyond standard context limits, enforcing strict output constraints, or maintaining high faithfulness to source material. It is also a practical fit for tabular data extraction and agentic tool-calling workflows that need predictable behavior. Skip this if your primary use case involves open-ended creative generation, complex strategic analysis, or high-throughput structured data formatting where the premium output cost does not align with the performance delta. The 4/5 scores in those areas indicate capable but not exceptional results, making lower-cost models a more efficient choice for those specific tasks.

Strengths — Top 3

Constrained Rewriting5.0/5.0
Faithfulness5.0/5.0
Long Context5.0/5.0

Relative weaknesses — Bottom 3

Structured Output4.0/5.0
Strategic Analysis4.0/5.0
Creative Problem Solving4.0/5.0

Similar models

QQwen: Qwen3.6 35B A3B$0.7854.54OOpenAI: GPT-5.6 Luna Pro$4.754.69OGPT-5.2$10.944.69QQwen: Qwen3.5-9B$0.1374.31