Anthropic: Claude Fable 5
Anthropic's flagship model. Long-context specialist with 1M window.
Scores by test
Methodology →What you need to know
Claude Fable 5 is a high-reasoning model optimized for complex autonomy and technical precision. Its primary differentiator is its near-perfect performance in mathematical and software engineering benchmarks, specifically scoring 99.7% on AIME 2025 and 95% on SWE-bench Verified. These numbers indicate a model capable of solving advanced competitive math and executing production-level software engineering tasks with minimal failure.
The model maintains a consistent 5/5 internal rating across almost all operational categories, including tool calling, agentic planning, and long context handling within its 1M token window. However, there is a critical failure in safety calibration, where it scored 1/5. This suggests the model lacks standard guardrails and may produce unfiltered or unsafe outputs, requiring developers to implement their own robust moderation layers.
At $10.00 per million input tokens and $50.00 per million output tokens, this is a premium-priced model. The cost is high, but the pricing aligns with its capability as a top-tier reasoning engine, ranking 16th out of 130 models. You are paying for high-fidelity structured output and strategic analysis rather than simple text generation.
Use this model if your application requires autonomous agentic planning, complex codebase manipulation, or high-stakes mathematical reasoning. Skip this model if your use case requires strict safety alignment out of the box or if you are performing simple classification tasks where a cheaper model would suffice.
Strengths — Top 3
Relative weaknesses — Bottom 3
Similar models