models/google/gemini-3-flash-preview
G
Google·active

Gemini 3 Flash Preview

Google's mid-tier model. Long-context specialist with 1.0M window.

Overall score
4.46
/5.00 · ranked #45
Input
$0.500
per 1M tokens
Output
$3.00
per 1M tokens
Context
1.0M
tokens
Blended
$2.38
3:1 out:in ratio

Price drops, new benchmarks, model updates. Stay current on Gemini 3 Flash Preview.

One email per change. Unsubscribe anytime.

modelpicker.aipowered by live benchmark data

Scores by test

Methodology →
Structured Output
5.0
Strategic Analysis
5.0
Constrained Rewriting
4.0
Creative Problem Solving
5.0
Tool Calling
5.0
Faithfulness
5.0
Classification
4.0
Long Context
5.0
Safety Calibration
1.0
Persona Consistency
5.0
Agentic Planning
5.0
Multilingual
5.0
Tabular Data
4.0
SWE-bench Verified
75.4
AIME 2025
95.6
GPQA Diamond
89.4
ARC-AGI-2
33.6
APEX Agents
24.0
Epoch Capabilities Index (ECI)
151.7

What you need to know

Gemini 3 Flash Preview is optimized for high-complexity technical tasks and massive datasets, distinguished primarily by its 1.0M token context window and perfect internal scores in long context, tool calling, and structured output. Its performance on hard reasoning benchmarks is high, specifically scoring 92.8% on AIME 2025 and 83.2% on GPQA Diamond, indicating strong capabilities in mathematics and graduate-level science.

At a blended cost of $2.38/MTok, the model provides a high ratio of intelligence to price. It maintains a 5/5 internal rating across agentic planning, strategic analysis, and faithfulness, making it a cost-effective alternative to larger frontier models for autonomous workflows. However, it has a critical failure in safety calibration, scoring 1/5, which suggests a lack of robust guardrails.

The model is a strong fit for developers building complex agents, RAG pipelines with massive documents, or applications requiring strict JSON/structured outputs. Skip this model if your application requires high safety alignment or strict content filtering, as the safety calibration data indicates significant vulnerability.

Strengths — Top 3

Structured Output5.0/5.0
Strategic Analysis5.0/5.0
Creative Problem Solving5.0/5.0

Relative weaknesses — Bottom 3

Safety Calibration1.0/5.0
Constrained Rewriting4.0/5.0
Classification4.0/5.0

Similar models

NNVIDIA: Nemotron 3 Ultra$2.854.46QQwen 3.7 Max$3.694.62AAnthropic: Claude Fable 5$40.004.62QQwen: Qwen3.7 Flash$0.1054.54