models/xai/grok-4-6
X
xAI·active

Grok 4.6

xAI's mid-tier model. Context window: 500K tokens.

Overall score
4.54
/5.00 · ranked #51
Input
$2.00
per 1M tokens
Output
$6.00
per 1M tokens
Context
500K
tokens
Blended
$5.00
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on Grok 4.6.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
5.0
Strategic Analysis
5.0
Constrained Rewriting
4.0
Creative Problem Solving
4.0
Tool Calling
4.0
Faithfulness
5.0
Classification
2.0
Long Context
5.0
Safety Calibration
5.0
Persona Consistency
5.0
Agentic Planning
5.0
Multilingual
5.0
Tabular Data
5.0
AIME 2025
99.2
GPQA Diamond
94.0
ARC-AGI-2
67.1
SciCode
56.5
APEX Agents
65.3
Epoch Capabilities Index (ECI)
156.3

What you need to know

Grok 4.6 is a high-reasoning model optimized for complex logic and structured data. Its primary differentiator is its performance in advanced mathematical and scientific reasoning, evidenced by a 97.8% score on AIME 2025 and 94% on GPQA Diamond. It excels in agentic planning, strategic analysis, and faithfulness, making it reliable for autonomous workflows and high-precision technical tasks.

The model handles massive datasets efficiently with a 500K context window and a perfect 5/5 internal score for long context processing. It is particularly strong at generating tabular data and maintaining persona consistency. However, it struggles significantly with basic classification tasks, where it scores only 2/5, indicating a potential failure in simple labeling or categorization duties.

At a blended cost of $5.00 per million tokens, this is a premium-priced model. The cost is justified for developers requiring high-fidelity structured outputs and complex problem solving, but it is inefficient for routine text processing or simple classification.

Use this model if your application requires deep strategic reasoning, large-scale document analysis, or strict adherence to structured output formats. Skip this model if your primary use case is data classification or if you are operating on a tight budget for high-volume, low-complexity tasks.

Strengths — Top 3

Structured Output5.0/5.0
Strategic Analysis5.0/5.0
Faithfulness5.0/5.0

Relative weaknesses — Bottom 3

Classification2.0/5.0
Constrained Rewriting4.0/5.0
Creative Problem Solving4.0/5.0

Similar models

MMoonshotAI: Kimi K2.6$3.244.62XMiMo-V2.5$0.2454.69QQwen: Qwen3.6 35B A3B$0.7874.54TThinking Machines: Inkling Small$1.014.69