Qwen: Qwen3 235B A22B Instruct 2507
Qwen's efficiency model. Context window: 262K tokens.
Scores by test
Methodology →What you need to know
Qwen3 235B A22B Instruct 2507 is built for high-complexity cognitive tasks and long-form processing. It excels in strategic analysis, persona consistency, and faithfulness, making it a reliable choice for applications requiring strict adherence to a specific voice or complex reasoning over large datasets. Its 262K context window is supported by a perfect 5/5 internal score for long context, ensuring it maintains coherence across extensive inputs.
From a cost perspective, the model is positioned as a budget-friendly option for its scale, with a blended cost of $0.435/MTok. While it ranks #94 out of 130 models overall, its strength is concentrated in specialized outputs rather than general utility. It delivers perfect scores in structured output and multilingual capabilities, though it performs mediocrely in basic classification and tabular data handling.
Developers should be aware of a significant weakness in safety calibration. A 1/5 score indicates the model frequently struggles to distinguish between harmful and benign requests, often resulting in the over-refusal of legitimate prompts. This makes it less suitable for customer-facing bots where friction-less interaction is a priority.
Use this model if you need an affordable, high-capacity engine for strategic planning, multilingual structured data generation, or deep analysis of long documents. Skip this model if your workflow relies heavily on data classification, tabular formatting, or requires a high degree of nuance in safety filtering to avoid false positives.
Strengths — Top 3
Relative weaknesses — Bottom 3
Similar models