Kimi K2.7 Code
Moonshot AI's flagship model. Context window: 262K tokens.
Scores by test
Methodology →What you need to know
Kimi K2.7 Code is a high-performance open-weight model optimized for complex logical reasoning and structured data generation. Its most significant technical differentiator is its extreme proficiency in mathematical reasoning, evidenced by a 96.4% score on the AIME 2025 benchmark. This makes it one of the most capable models currently available for high-level quantitative tasks.
The model excels in agentic workflows, earning perfect 5/5 internal scores for tool calling, strategic analysis, and agentic planning. It is particularly reliable for developers requiring strict adherence to formats, as it maintains a perfect score for structured output and faithfulness. While it handles long context up to 262K tokens, its performance in long-context retrieval and classification is slightly lower than its reasoning capabilities, marking its primary relative weakness.
At a blended cost of $2.81/MTok, the model is priced competitively for its rank as the 10th best model out of 110. It offers a high ratio of reasoning capability to cost, especially for developers who prefer the flexibility of an open-weight architecture without sacrificing the performance typically reserved for closed-source frontier models.
Use this model if you are building autonomous agents, solving complex mathematical problems, or require precise structured outputs. Skip this model if your primary use case is simple text classification or if you require a model specifically optimized for maximum retrieval accuracy across its entire 262K context window.
Strengths — Top 3
Relative weaknesses — Bottom 3
Similar models