models/stepfun/step-3-7-flash
S
stepfun·active

StepFun: Step 3.7 Flash

stepfun's efficiency model. Context window: 262K tokens.

Overall score
4.15
/5.00 · ranked #100
Input
$0.200
per 1M tokens
Output
$1.15
per 1M tokens
Context
262K
tokens
Blended
$0.912
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on StepFun: Step 3.7 Flash.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
5.0
Strategic Analysis
5.0
Constrained Rewriting
3.0
Creative Problem Solving
4.0
Tool Calling
5.0
Faithfulness
5.0
Classification
2.0
Long Context
4.0
Safety Calibration
2.0
Persona Consistency
5.0
Agentic Planning
4.0
Multilingual
5.0
Tabular Data
5.0
SciCode
40.0

What you need to know

Step 3.7 Flash is optimized for high-precision technical tasks, specifically excelling in structured output, tool calling, and strategic analysis. With perfect 5/5 scores in these categories, it is highly reliable for developers building agentic workflows or applications requiring strict data formatting and tabular data processing. Its performance in multilingual tasks and persona consistency further supports its utility in complex, multi-step automation.

The model offers a competitive price point with a blended cost of $0.912/MTok and a substantial 262K context window. While it ranks 85th overall among 130 models, its ability to handle long-context tasks and maintain faithfulness makes it a high-value option for those who prioritize functional reliability over general-purpose ranking.

There are significant gaps in its performance regarding classification and safety calibration, both scoring 2/5. It also struggles with constrained rewriting. These weaknesses indicate the model is not suited for moderation tasks, strict content filtering, or high-accuracy categorization.

Use this model if you are building autonomous agents, integrating external tools, or need reliable structured data extraction. Skip this model if your primary use case involves text classification, safety-critical filtering, or highly constrained editing.

Strengths — Top 3

Structured Output5.0/5.0
Strategic Analysis5.0/5.0
Tool Calling5.0/5.0

Relative weaknesses — Bottom 3

Classification2.0/5.0
Safety Calibration2.0/5.0
Constrained Rewriting3.0/5.0

Similar models

XGrok 4.5$5.004.38Oo3$6.504.31GGemini 3.5 Flash$7.134.46MMiniMax: MiniMax M3$0.9754.54