Gemini 3.7 Flash launched on August 13, 2026, as Google’s workhorse model for coding and agents. The API costs $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026; those rates rise to $1.50 and $7.50 on January 1, 2027.
Research answer

Create a landscape editorial hero image for this Studio Global article: What is Gemini 3.7 Flash, when did it launch, where is it available, and why does Google call it its fastest-growing model ever despite disc. Article summary: Gemini 3.7 Flash is Google’s production “workhorse” Gemini model for complex coding, web development, and tool-using agents. It launched on August 13, 2026, and is available through the Gemini API/AI Studio, Google Antig. Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
Gemini 3.7 Flash is Google’s latest production “workhorse” model, aimed at complex coding, web development, tool use, and multi-step agent workflows. It launched on August 13, 2026, and is available through the Gemini API, Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform.
The model’s appeal is not that it leads every intelligence benchmark. Its stronger proposition is the combination of near-frontier capability, very high output speed, and low initial API pricing—an attractive mix for developers running repeated agent loops at scale.
Google has described Gemini 3.7 Flash as its fastest-growing model ever, but the supplied reporting does not include the figures needed to test that statement. There is no disclosed usage base, measurement period, active-developer count, token volume, or named comparator in the available evidence.
The defensible conclusion is therefore narrower: Google is positioning Flash for rapid developer adoption, and its low introductory price and presence across Google’s developer products support that strategy. The growth superlative should be treated as an unquantified company claim rather than an independently established market statistic.
Gemini 3.7 Flash’s introductory API rates through December 31, 2026 are:
From January 1, 2027, the standard rates become $1.50 per million input tokens and $7.50 per million output tokens—exactly double the introductory prices. The listed discounted or batch-style rates also rise from $0.375/$1.875 to $0.75/$3.75 per million input/output tokens.
That expiry date matters for production planning. A workload that looks unusually inexpensive during the launch period should be costed again using the 2027 rates before developers commit to a long-lived architecture.
Token pricing is also only a partial measure of agent cost. Real-world spending depends on context length, tool calls, retries, cached input, reasoning effort, and output volume. Benchmark-specific costs should not be mistaken for a guaranteed price for a complete production workflow.
Independent Artificial Analysis reporting puts Gemini 3.7 Flash at an Intelligence Index score of 56. That places it above Claude Sonnet 5 at 55 and just below GPT-5.6 Terra at 57 on the cited comparison.
Speed is the model’s clearest differentiator. Artificial Analysis measured approximately 340 output tokens per second and an average time-per-task score of 1.7, reporting that Flash is about 40% faster than GPT-5.6 Terra at maximum reasoning.
Those figures describe measured model or provider performance, not a universal experience for every application. Network conditions, provider routing, prompt size, tool latency, and reasoning settings can all change end-to-end response time.
Google reports substantial improvements over Gemini 3.6 Flash in software engineering, debugging, issue resolution, web development, and agentic workflows. The cited coding results include:
The results suggest a capable coding and terminal-agent model, but not an across-the-board winner. GPT-5.6 Terra remains ahead on the cited Terminal-Bench result, while Flash’s advantage is speed and lower introductory cost.
At high reasoning effort, ARC Prize reports that Gemini 3.7 Flash scored:
These are strong results, particularly given the reported per-task costs. However, ARC-AGI performance should be interpreted as a benchmark signal, not a direct forecast of software-agent reliability. A production agent may make multiple model calls, invoke external tools, retry failed actions, and process much larger contexts than a benchmark task.
The table is not a universal leaderboard. The benchmarks come from different evaluations, and the sources do not establish a single controlled test across every model. Prices may also reflect different hosting, discount, or availability conditions.
The supplied comparisons point to a deliberate positioning choice. Gemini 3.7 Flash is close to the frontier on the cited Intelligence Index, but it does not claim the highest score: GPT-5.6 Terra is slightly ahead at 57, and Terra also leads the cited Terminal-Bench comparison.
Google is instead emphasizing the intelligence–latency–cost trade-off. A model that is fast enough for interactive development and inexpensive enough for repeated calls can be more useful to an agent platform than a slower, costlier model that wins a narrow benchmark. That matters for coding assistants, automated issue resolution, web-development agents, and workflows that repeatedly read context, call tools, inspect results, and try again.
The August 2026 release environment makes that trade-off more important. Qwen3.8-Max adds strong agent and software-engineering competition, GLM-5.3 and Nemotron 3.5 Lightning push on price and efficiency, while Claude Opus 5 and GPT-5.6 target higher-end capability.
Gemini 3.7 Flash is best understood as a high-throughput workhorse rather than the undisputed smartest model. Its strongest case is the combination of a 56 Intelligence Index score, approximately 340-token-per-second generation, solid coding and terminal-agent results, and unusually low introductory pricing.
For developers, the main caveat is financial rather than technical: the $0.75/$3.75 launch rates expire on December 31, 2026, and the standard $1.50/$7.50 rates begin the next day. Google’s growth claim may prove correct, but without supporting usage figures, the evidence currently supports a strategy of making fast, affordable inference the default—not a verified record for model adoption.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Gemini 3.7 Flash launched on August 13, 2026, as Google’s workhorse model for coding and agents.
Gemini 3.7 Flash launched on August 13, 2026, as Google’s workhorse model for coding and agents. The API costs $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026; those rates rise to $1.50 and $7.50 on January 1, 2027.
Google’s “fastest growing model ever” description is not independently verifiable from the supplied evidence because no usage base, growth period, token volume, or comparator is disclosed.