Thinking Machines: Inkling Small
thinkingmachines's flagship model. Context window: 524K tokens.
Scores by test
Methodology →What you need to know
Inkling Small distinguishes itself through high-tier reasoning and reliability across almost every internal metric, ranking 13th out of 133 models. It achieves perfect 5/5 scores in critical developer workflows, including tool calling, agentic planning, and structured output. Its 524K context window is paired with a 5/5 long context score, making it a viable option for processing massive datasets without losing faithfulness.
The model is priced competitively for its performance level, with a blended cost of $1.02/MTok. While it excels at complex strategic analysis and creative problem solving, it has a significant deficiency in classification tasks, where it scored only 2/5. This indicates a failure to handle simple labeling or categorization tasks effectively compared to its high performance in more complex reasoning.
Use this model for agentic workflows, complex tool integration, and long-document analysis where high faithfulness is required. Skip this model if your primary use case is high-volume text classification or simple data labeling.
Strengths — Top 3
Relative weaknesses — Bottom 3
Similar models