Gemini 3.5 Flash Low is Google's response to a developer revolt over token consumption: the default thinking behavior was burning quotas in under an hour and running tasks at 5.5x the cost of the previous Flash model,... The Low variant cuts token output by roughly 45% compared to the now renamed Medium variant, giv...

Create a landscape editorial hero image for this Studio Global article: What prompted Google to introduce the "Low" thinking level in Gemini 3.5 Flash on Antigravity, and how does this change address developer fr. Article summary: Google introduced the **Gemini 3.5 Flash (Low)** thinking level in Antigravity in direct response to a firestorm of developer backlash triggered by the model's launch at I/O 2026 on May 19. The core problem: Gemini 3.5 F. Topic tags: general, general web, user generated. Reference image context from search candidates: Reference image 1: visual subject "Product naming/UX confusion around Gemini CLI vs Antigravity CLI and broader interface design criticism (zachtratar, kchonyc, teortaxesTex). ## Gemini 3.5 Flash: the main technica" source context "[AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for ..." Reference image 2: visual subject "4M views • 6
When Google launched Gemini 3.5 Flash at I/O 2026 on May 19, it was pitched as the company's strongest agentic and coding model yet . What followed was a week of developer fury, emergency quota changes, and a quiet product pivot that rewrote how the model thinks.
The core issue: Gemini 3.5 Flash's default thinking behavior burned tokens so aggressively that paid Antigravity users were exhausting their quotas in under an hour . Despite per-token pricing that looked competitive on paper—$1.50 per million input tokens and $9.00 per million output tokens—the total cost to complete real tasks told a different story. Artificial Analysis found that running a standard benchmarking suite cost $1,552 with Gemini 3.5 Flash, compared to $282 for the previous Gemini 3 Flash—a 5.5x increase
.
Developer frustration erupted almost immediately. The Antigravity forum, Reddit, and X filled with complaints about extreme quota consumption . Developers who paid for Antigravity's Pro plan reported that their quotas, which previously lasted a full day of work, disappeared within 30 to 60 minutes after switching to Gemini 3.5 Flash
.
One quantified analysis posted to Reddit by u/tadanada explicitly called out the cost inflation, comparing a $1,552 benchmark run for Gemini 3.5 Flash against $278 for Gemini 3 Flash—a 5.6x difference that explained why paid plans were collapsing so quickly .
Google's response came in two waves:
high to medium Even the 9x quota increase didn't fully solve the problem. Some developers reported hitting their weekly Flash lockout within 30 minutes of resuming work after the quota reset .
Gemini 3.5 Flash Low represents a more surgical fix: instead of just giving developers more raw quota (a supply-side bandage), it gave them a way to use fewer tokens per task (a demand-side control).
Google's official documentation describes the Low variant as having been "significantly improved for code and agentic tasks that require fewer steps, offering strong quality at lower latency and cost" . The company states the Low variant generates roughly 45% fewer output tokens than the now-renamed Medium variant
.
For developers, this means they can now set thinking_level: "low".
This effectively gives developers a four-tier dial for reasoning effort—minimal, low, medium, high—instead of a binary choice between "thinking on" and "thinking off" .
One of the biggest API traps in the Gemini 3.5 Flash launch was the unannounced change of the default thinking_level from high to medium. Developers who ported directly from gemini-3-flash-preview without explicitly setting a thinking level were silently getting different reasoning behavior . This meant that even after the Low variant shipped, many developers were still using more tokens than necessary for simple tasks because they hadn't noticed the default had shifted.
The Low variant essentially completes the fix: it gives developers an explicit, documented, and purpose-built level for the kind of cost-sensitive work that the Flash family was originally designed for.
The rollout of Gemini 3.5 Flash Low, combined with the 9x quota increases and default thinking level adjustment, has stabilized the Antigravity developer experience. Developers can now:
thinking_level: "low"The Low variant isn't a replacement for Google's quota increases—it's a complement. Developers who use both the new thinking level and the 9x expanded quotas can now work through meaningful coding sessions without hitting limits or burning through their monthly Antigravity budgets in an afternoon.
Studio Global AI
Use this topic as a starting point for a fresh source-backed answer, then compare citations before you share it.
Gemini 3.5 Flash Low is Google's response to a developer revolt over token consumption: the default thinking behavior was burning quotas in under an hour and running tasks at 5.5x the cost of the previous Flash model,...
Gemini 3.5 Flash Low is Google's response to a developer revolt over token consumption: the default thinking behavior was burning quotas in under an hour and running tasks at 5.5x the cost of the previous Flash model,... The Low variant cuts token output by roughly 45% compared to the now renamed Medium variant, giving developers an explicit control to throttle costs for simple tasks while reserving deep reasoning for complex problems.
This feature landed alongside two emergency 9x quota increases and a quiet change of the default thinking level from 'high' to 'medium'—all within one week of the I/O 2026 launch.