models/xai/grok-4
X
xAI·deprecated 2026-05-15

Grok 4

xAI's efficiency model. Context window: 256K tokens.

Overall score
4.15
/5.00 · ranked #96
Input
$3.00
per 1M tokens
Output
$15.00
per 1M tokens
Context
256K
tokens
Blended
$12.00
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on Grok 4.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
4.0
Strategic Analysis
5.0
Constrained Rewriting
4.0
Creative Problem Solving
3.0
Tool Calling
4.0
Faithfulness
5.0
Classification
4.0
Long Context
5.0
Safety Calibration
2.0
Persona Consistency
5.0
Agentic Planning
3.0
Multilingual
5.0
Tabular Data
5.0

What you need to know

Grok 4 distinguishes itself through high-fidelity handling of complex data structures and long-range context. It achieves a perfect 5/5 score in tabular data processing, strategic analysis, and long context utilization, making it a strong candidate for analyzing massive datasets or extracting insights from extensive documentation. Its reliability is further supported by top marks in faithfulness, multilingual capabilities, and persona consistency.

From a cost perspective, the model is positioned at a premium price point with a blended cost of $12.00/MTok. Given its overall rank of 83 out of 130 models and an average internal score of 4.15, the pricing is high relative to its general performance. Developers are paying a premium specifically for its specialized strengths in data organization and strategic reasoning rather than general-purpose intelligence.

The model struggles with agentic planning and creative problem solving, both scoring 3/5. Additionally, its safety calibration is low at 2/5, which indicates a tendency to over-refuse benign requests. This suggests the model may be overly restrictive, potentially hindering workflows that require nuanced or unconventional prompts.

Use this model if your primary requirement is high-accuracy processing of tabular data or strategic analysis across a 256K context window. Skip this model if you need an autonomous agent for complex planning, a creative brainstorming partner, or a model with flexible, low-friction safety boundaries.

Strengths — Top 3

Strategic Analysis5.0/5.0
Faithfulness5.0/5.0
Long Context5.0/5.0

Relative weaknesses — Bottom 3

Safety Calibration2.0/5.0
Creative Problem Solving3.0/5.0
Agentic Planning3.0/5.0

Similar models

OGPT-5.1$7.814.23QQwen: Qwen3.5-35B-A3B$1.004.31MMistral Medium 3.5$6.004.15SStepFun: Step 3.5 Flash$0.2504.23