models/moonshot-ai/kimi-k2-7-code
M
Moonshot AI·active

Kimi K2.7 Code

Moonshot AI's flagship model. Context window: 262K tokens.

Overall score
4.62
/5.00 · ranked #18
Input
$0.730
per 1M tokens
Output
$3.50
per 1M tokens
Context
262K
tokens
Blended
$2.81
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on Kimi K2.7 Code.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
5.0
Strategic Analysis
5.0
Constrained Rewriting
4.0
Creative Problem Solving
5.0
Tool Calling
5.0
Faithfulness
5.0
Classification
3.0
Long Context
4.0
Safety Calibration
5.0
Persona Consistency
5.0
Agentic Planning
5.0
Multilingual
5.0
Tabular Data
4.0
AIME 2025
96.4
GPQA Diamond
89.5
SciCode
47.5
APEX Agents
27.6
Epoch Capabilities Index (ECI)
149.7

What you need to know

Kimi K2.7 Code is a high-performance open-weight model optimized for complex logical reasoning and structured data generation. Its most significant technical differentiator is its extreme proficiency in mathematical reasoning, evidenced by a 96.4% score on the AIME 2025 benchmark. This makes it one of the most capable models currently available for high-level quantitative tasks.

The model excels in agentic workflows, earning perfect 5/5 internal scores for tool calling, strategic analysis, and agentic planning. It is particularly reliable for developers requiring strict adherence to formats, as it maintains a perfect score for structured output and faithfulness. While it handles long context up to 262K tokens, its performance in long-context retrieval and classification is slightly lower than its reasoning capabilities, marking its primary relative weakness.

At a blended cost of $2.81/MTok, the model is priced competitively for its rank as the 10th best model out of 110. It offers a high ratio of reasoning capability to cost, especially for developers who prefer the flexibility of an open-weight architecture without sacrificing the performance typically reserved for closed-source frontier models.

Use this model if you are building autonomous agents, solving complex mathematical problems, or require precise structured outputs. Skip this model if your primary use case is simple text classification or if you require a model specifically optimized for maximum retrieval accuracy across its entire 262K context window.

Strengths — Top 3

Structured Output5.0/5.0
Strategic Analysis5.0/5.0
Creative Problem Solving5.0/5.0

Relative weaknesses — Bottom 3

Classification3.0/5.0
Constrained Rewriting4.0/5.0
Long Context4.0/5.0

Similar models

NNex-N2-Pro$0.8134.62GGemini 3.5 Flash Lite$1.954.62MMeta: Muse Spark 1.1$3.504.77AAnthropic: Claude Sonnet 5$8.004.77