OpenAI: gpt-oss-120b
OpenAI's efficiency model. Context window: 131K tokens.
Scores by test
Methodology →What you need to know
The gpt-oss-120b distinguishes itself through high-precision analytical capabilities and reliability. It achieves perfect scores in strategic analysis, faithfulness, and agentic planning, making it a strong candidate for complex reasoning tasks where factual accuracy and logical sequencing are critical. Its 131K context window is fully leveraged, as evidenced by a maximum score in long-context processing.
From a cost perspective, the model is highly economical. With a blended cost of $0.152/MTok, it provides high-tier reasoning and multilingual proficiency at a price point significantly lower than most flagship models. This creates a favorable performance-to-price ratio for developers running high-volume analysis or data processing pipelines.
Despite its analytical strengths, the model has significant operational gaps. It performs poorly in safety calibration and is mediocre at tool calling and classification. Developers should expect a lack of built-in guardrails and potential instability when integrating the model into autonomous agent frameworks that rely heavily on external API calls.
Use this model for complex strategic planning, long-document analysis, or multilingual data processing where cost efficiency is a priority. Skip this model if your application requires strict safety filtering, high-accuracy classification, or reliable tool-use integration.
Strengths — Top 3
Relative weaknesses — Bottom 3
Similar models