models/openai/gpt-oss-120b
O
OpenAI·active·free tier available

OpenAI: gpt-oss-120b

OpenAI's efficiency model. Context window: 131K tokens.

Overall score
4.08
/5.00 · ranked #96
Input
$0.037
per 1M tokens
Output
$0.170
per 1M tokens
Context
131K
tokens
Blended
$0.137
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on OpenAI: gpt-oss-120b.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
4.0
Strategic Analysis
5.0
Constrained Rewriting
4.0
Creative Problem Solving
4.0
Tool Calling
3.0
Faithfulness
5.0
Classification
3.0
Long Context
5.0
Safety Calibration
1.0
Persona Consistency
4.0
Agentic Planning
5.0
Multilingual
5.0
Tabular Data
5.0
APEX Agents
4.7
Epoch Capabilities Index (ECI)
140.5

What you need to know

The gpt-oss-120b distinguishes itself through high-precision analytical capabilities and reliability. It achieves perfect scores in strategic analysis, faithfulness, and agentic planning, making it a strong candidate for complex reasoning tasks where factual accuracy and logical sequencing are critical. Its 131K context window is fully leveraged, as evidenced by a maximum score in long-context processing.

From a cost perspective, the model is highly economical. With a blended cost of $0.152/MTok, it provides high-tier reasoning and multilingual proficiency at a price point significantly lower than most flagship models. This creates a favorable performance-to-price ratio for developers running high-volume analysis or data processing pipelines.

Despite its analytical strengths, the model has significant operational gaps. It performs poorly in safety calibration and is mediocre at tool calling and classification. Developers should expect a lack of built-in guardrails and potential instability when integrating the model into autonomous agent frameworks that rely heavily on external API calls.

Use this model for complex strategic planning, long-document analysis, or multilingual data processing where cost efficiency is a priority. Skip this model if your application requires strict safety filtering, high-accuracy classification, or reliable tool-use integration.

Strengths — Top 3

Strategic Analysis5.0/5.0
Faithfulness5.0/5.0
Long Context5.0/5.0

Relative weaknesses — Bottom 3

Safety Calibration1.0/5.0
Tool Calling3.0/5.0
Classification3.0/5.0

Similar models

DDeepSeek V3.2$0.3674.31IInception: Mercury 2$0.6254.08QQwen: Qwen3.6 Flash$0.8914.23AClaude Opus 5$20.004.38