Google Is Making Powerful AI Agents Much Cheaper to Run
Google has released Gemini 3.7 Flash only three weeks after Gemini 3.6 Flash, with a clear focus on coding, autonomous agents and lower operating costs. The new model reaches 65.3% on the DeepSWE software-engineering benchmark and 30.4% on an enterprise automation test, while launching at half the original Gemini 3.6 Flash token price. For developers building agents that may call a model dozens of times to finish one task, the pricing change could matter as much as the benchmark gains.Gemini 3.7 Flash arrived only three weeks after its predecessor
Google introduced Gemini 3.7 Flash on August 13, 2026, describing it as its most intelligent "workhorse" model yet for coding and agents. According to the official Gemini 3.7 Flash announcement, the unusually short release cycle was driven by developer feedback and algorithmic improvements that Google plans to carry into future models.The model is based on Gemini 3.6 Flash rather than a completely new architecture. It retains a context window of up to 1 million tokens, supports 64,000 output tokens and allows developers to choose low, medium or high thinking levels depending on the desired balance between latency, reasoning depth and cost.
The rapid update suggests Google is treating the Flash family less like an occasional model release and more like continuously upgraded infrastructure for production AI agents.
The biggest jump is in long-running coding work
Gemini 3.7 Flash improved sharply on software-engineering evaluations designed to measure more than isolated code generation. Google reports a 65.3% score on DeepSWE v1.1, compared with about 49% for Gemini 3.6 Flash, while FrontierCode 1.1 Main rose from 34.4% to 43.6%.These benchmarks matter for agentic development because an autonomous coding system has to inspect repositories, reason across multiple files, make changes, test them and recover when something fails. Google says the new model is better at adapting to roadblocks, following instructions and completing multi-step plans with fewer failed loops and less manual supervision.
The Gemini 3.7 Flash model card also reports 85.8% on Terminal-bench 2.1, up from 78.0% for Gemini 3.6 Flash. The improvement reinforces Google's positioning of Flash as an execution model rather than simply a faster chatbot.
Enterprise automation jumped from 17% to 30.4%
The improvement is not limited to programming. On AutomationBench, which measures the ability to complete real business workflows, Gemini 3.7 Flash scored 30.4% compared with 17.0% for Gemini 3.6 Flash. That represents roughly a 79% relative increase.Google is already putting those capabilities into Spark, its always-on personal agent for Google AI Pro and Ultra subscribers. Gemini 3.7 Flash now powers Spark with improved tool use for Google Workspace tasks such as consolidating files, drafting emails and updating status documents.
This is the larger strategy behind the release. Google is optimizing Flash for software that receives a goal, decides which tools to use and performs a sequence of actions without requiring the user to approve every individual step.
Google cut the introductory API price in half
Gemini 3.7 Flash launches at $0.75 per 1 million input tokens and $3.75 per 1 million output tokens. Google describes that as half the original Gemini 3.6 Flash price and has temporarily applied the same lower rate to 3.6 Flash as well.The pricing is especially important for agents because a single user request may generate many separate model calls. An agent can plan a task, inspect files, call external tools, evaluate the result, correct mistakes and repeat the process. Every additional loop consumes tokens, so reducing the unit price compounds across an entire workflow rather than saving money only once.
Google's developer documentation confirms that the promotional rate lasts through December 31, 2026. Starting January 1, 2027, standard pricing will double to $1.50 per 1 million input tokens and $7.50 per 1 million output tokens.
The low price is designed to push agents into production
The economics reveal Google's target more clearly than the benchmark charts. A company testing an AI assistant may tolerate an expensive model for occasional prompts. A company deploying thousands of autonomous workflows needs predictable inference costs because each workflow can involve repeated reasoning and tool calls.Gemini 3.7 Flash is therefore positioned between lightweight low-cost models and more expensive frontier systems. Its goal is not necessarily to win every intelligence benchmark. The model card shows stronger competitors on some tests, including GPT-5.6 Terra on DeepSWE and Terminal-bench 2.1. Google's proposition is instead a combination of capability, speed and price that can be deployed at large scale.
That explains why Google calls Flash a workhorse. The competitive question is shifting from which model gives the smartest single answer to which model can economically complete thousands of useful actions every day.
Google is making 3.7 Flash the engine behind its agent stack
The model is already becoming infrastructure inside Google's broader agent platform. Gemini 3.7 Flash is the new default model for the Antigravity agent in Gemini Managed Agents, while developers can also use it directly through Gemini API, Google AI Studio and Android Studio.The model supports the same built-in tool suite as Gemini 3.6 Flash and is explicitly optimized for multi-step agentic workflows, code generation, multimodal reasoning and design adherence. Google also reports an OSWorld-2.0 score of 47.9% for agentic computer use, compared with 33.8% for Gemini 3.6 Flash.
The direction is clear: Google increasingly wants Gemini to operate software, navigate workflows and coordinate tools rather than wait inside a chat window for the next prompt.
The cheap price has an expiration date
Developers evaluating Gemini 3.7 Flash need to separate its current promotional economics from its long-term price. The $0.75 input and $3.75 output rates are introductory, not permanent. Both will double at the start of 2027 unless Google changes its published schedule.That creates an unusual four-month window in which companies can test agent-heavy applications at a substantially reduced inference cost. For prototypes this is attractive, but production teams should calculate unit economics against the announced 2027 rates rather than assuming today's price will remain indefinitely.
Google has effectively combined a technical release with an adoption incentive. Gemini 3.7 Flash is smarter at multi-step work, but the temporary discount is what makes experimenting with large numbers of autonomous agent calls much easier right now.
The AI race is becoming a competition over cost per completed task
Gemini 3.7 Flash shows where the next phase of model competition is heading. Raw intelligence still matters, but agentic systems introduce another metric: how much it costs for a model to actually finish a job.A model that requires fewer retries, handles tools more reliably and costs less per token can outperform a theoretically stronger model economically once hundreds of actions are involved. Google's rapid three-week upgrade cycle and temporary 50% price reduction suggest it sees that shift clearly.
The important number may therefore not be 65.3% on DeepSWE or $0.75 per million input tokens in isolation. It is the combination of better execution and cheaper repeated inference that could make autonomous AI agents practical for far more everyday software and business workflows.
Editorial Team - CoinBotLab