models/openai/gpt-5-6-luna
O
OpenAI·active

OpenAI: GPT-5.6 Luna

OpenAI's flagship model. Long-context specialist with 1.1M window.

Overall score
4.85
/5.00 · ranked #1
Input
$0.200
per 1M tokens
Output
$1.20
per 1M tokens
Context
1.1M
tokens
Blended
$0.950
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on OpenAI: GPT-5.6 Luna.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
5.0
Strategic Analysis
5.0
Constrained Rewriting
5.0
Creative Problem Solving
5.0
Tool Calling
4.0
Faithfulness
5.0
Classification
4.0
Long Context
5.0
Safety Calibration
5.0
Persona Consistency
5.0
Agentic Planning
5.0
Multilingual
5.0
Tabular Data
5.0
AIME 2025
98.3
GPQA Diamond
91.6
ARC-AGI-2
59.5
SciCode
52.5
Epoch Capabilities Index (ECI)
156.2

What you need to know

GPT-5.6 Luna is a high-tier model optimized for complex reasoning and strict adherence to formatting. Its primary differentiator is near-perfect performance in structured output, constrained rewriting, and strategic analysis, making it highly reliable for production pipelines that require precise data schemas. This reliability extends to its reasoning capabilities, evidenced by a 98.3% score on AIME 2025, placing it among the top three models globally for mathematical and logical tasks.

The model handles massive datasets via a 1.1M token context window, maintaining maximum scores for faithfulness and long-context retrieval. While most capabilities are peaked at 5/5, it shows slight relative weakness in classification and tool calling, where it scores 4/5. This suggests that while it can execute these tasks, it is less specialized in them than it is in generative reasoning or structural transformation.

At a blended cost of $4.75/MTok, Luna is priced as a premium model. The cost is justified for high-stakes enterprise applications where accuracy and structural integrity are more valuable than token economy. It provides a high performance-to-cost ratio for complex agentic planning and multilingual tasks, but may be overpriced for simple chat or basic classification.

Use this model if your project requires extreme precision in structured data, high-level strategic reasoning, or the processing of very large documents. Skip this model if your primary use case is simple text classification or if you are operating on a tight budget for high-volume, low-complexity tasks.

Strengths — Top 3

Structured Output5.0/5.0
Strategic Analysis5.0/5.0
Constrained Rewriting5.0/5.0

Relative weaknesses — Bottom 3

Tool Calling4.0/5.0
Classification4.0/5.0
Structured Output5.0/5.0

Similar models

OOpenAI: GPT-5.6 Terra$9.504.77XMiMo-V2.5$0.2454.69ZZ.ai: GLM 5.2$2.524.85QQwen: Qwen3.6 Max Preview$4.884.85