OpenAI: GPT-5.6 Luna
OpenAI's flagship model. Long-context specialist with 1.1M window.
Scores by test
Methodology →What you need to know
GPT-5.6 Luna is a high-tier model optimized for complex reasoning and strict adherence to formatting. Its primary differentiator is near-perfect performance in structured output, constrained rewriting, and strategic analysis, making it highly reliable for production pipelines that require precise data schemas. This reliability extends to its reasoning capabilities, evidenced by a 98.3% score on AIME 2025, placing it among the top three models globally for mathematical and logical tasks.
The model handles massive datasets via a 1.1M token context window, maintaining maximum scores for faithfulness and long-context retrieval. While most capabilities are peaked at 5/5, it shows slight relative weakness in classification and tool calling, where it scores 4/5. This suggests that while it can execute these tasks, it is less specialized in them than it is in generative reasoning or structural transformation.
At a blended cost of $4.75/MTok, Luna is priced as a premium model. The cost is justified for high-stakes enterprise applications where accuracy and structural integrity are more valuable than token economy. It provides a high performance-to-cost ratio for complex agentic planning and multilingual tasks, but may be overpriced for simple chat or basic classification.
Use this model if your project requires extreme precision in structured data, high-level strategic reasoning, or the processing of very large documents. Skip this model if your primary use case is simple text classification or if you are operating on a tight budget for high-volume, low-complexity tasks.
Strengths — Top 3
Relative weaknesses — Bottom 3
Similar models