
Google just made its cheapest coding model a lot better, and for now, a lot cheaper too. On August 13, 2026, it released Gemini 3.7 Flash, a workhorse model built for coding and AI agents, and set the price at $0.75 per million input tokens and $3.75 per million output tokens.
Here is the part most headlines skip. That price is introductory. It runs through December 31, 2026, and on January 1 it doubles to $1.50 and $7.50 per million tokens. So the discount has an expiry date. Build on it, and your bill jumps in the new year.
The model itself is a genuine step up for code and agent work. The pricing is the thing to read twice.
What is Gemini 3.7 Flash?
Gemini 3.7 Flash is Google’s latest workhorse model, the cheaper and faster tier of its Gemini family built for high-volume tasks rather than the hardest single problems. In its launch announcement, Google calls it “our most intelligent workhorse model yet for coding and agents.” It handles a 1 million token context window, writes up to 64,000 tokens of output, and takes text, images, audio and video as input while replying in text only. Its knowledge runs to March 2026. You also get tunable thinking levels (low, medium, high), so you can trade speed for reasoning depth on each request. Think of Flash as the engine you run thousands of times a day, not the one you save for a single hard question. That is why a change here matters more to a lot of builders than a flagship launch does.
How much does Gemini 3.7 Flash cost?
Right now it costs $0.75 per million input tokens and $3.75 per million output tokens, and those are introductory rates that expire on December 31, 2026. From January 1, 2027, the standard price is $1.50 per million input and $7.50 per million output, double the launch rate, and that standard rate is roughly where Flash pricing sat before this release. In plain terms, you are getting a better model at a temporary discount, not a permanent price drop. Google spells the numbers out in its developer documentation, so there is no ambiguity about the reset date.
| Pricing tier | Input (per 1M tokens) | Output (per 1M tokens) | When it applies |
|---|---|---|---|
| Introductory | $0.75 | $3.75 | Now through Dec 31, 2026 |
| Standard | $1.50 | $7.50 | From Jan 1, 2027 |
What is the catch with the new pricing?
The catch is that the discount is temporary and the reset is automatic. Wire Gemini 3.7 Flash into an agent that runs at scale, and every token gets twice as expensive on January 1, 2027, with no action from Google needed. High-volume agent loops feel this first, because output tokens are the pricey side and agents generate a lot of them. Two practical moves. Budget your 2027 spend at the standard $1.50 and $7.50 rate, not the promo. And watch your output-token volume now, so the January bill is not a shock. The discount is a real saving for the rest of 2026. Just do not build a business case on a number with a few months left on it.
Is Gemini 3.7 Flash actually better at coding?
Yes, and the jump over the previous Flash is large on Google’s own numbers. On the FrontierCode 1.1 coding test it scores 43.6%, up from 34.4% for Gemini 3.6 Flash. On DeepSWE, a software-engineering benchmark, it hits 65.3% versus 49.0%. Its WebDev Arena rating climbs to 1588 Elo from 1538, and on an agent-automation test it more than doubles, to 30.4% from 17.0%. Google’s model card also lists 85.8% on Terminal-bench 2.1 and 97.0% on a long-context recall test. These are vendor-reported figures, so read them as a direction rather than a promise. Still, the pattern is consistent: fewer failed agent loops and cleaner first-pass code. One trick Google highlights is generating working web and desktop code straight from a design mock.
Where can you use Gemini 3.7 Flash?
You can reach it through Google’s developer tools, its enterprise platform, and the consumer Gemini app. Developers get it in the Gemini API, Google AI Studio, Android Studio and Google Antigravity, with coding help through Gemini Code Assist. Businesses get it inside the Gemini Enterprise Agent Platform. For individuals, it powers Gemini Spark for Google AI Pro and Ultra subscribers across more than 160 countries. The 1 million token context window means it can hold a large codebase, a long PDF, or hours of transcript in a single request, which is part of why Google keeps pointing it at agents that need to keep a lot of state in view. If you already use the wider Gemini family, this is closer to a drop-in upgrade than a new product to learn.
How does this compare to what OpenAI and DeepSeek did this week?
It lands in the middle of a price war over the cheap model tier, and the three big players moved in different directions in the same seven days. Google cut Flash pricing. OpenAI cut the price of its GPT-5.6 Luna model and previewed an ultrafast mode it says runs up to 14 times faster. DeepSeek went the other way and raised prices on its new V4 Pro model, sharply on some workloads, as one industry roundup noted. The real 2026 fight is not over flagship intelligence. It is over the cost of the everyday model that agents call constantly, and when your assistant makes thousands of calls, a few cents per million tokens sets your monthly bill. If you want the wider picture, we broke down what the major AI assistants cost earlier this month. xAI played the same game days before with Grok 4.6, and Meta went cheaper still by putting an agent model on your own laptop.
Should you switch to Gemini 3.7 Flash?
Switch if you run high-volume coding or agent workloads and you already live in Google’s tools, because you get better code output at a discount for the rest of 2026. It fits agents that loop a lot, code generated from design mocks, and long-context jobs like reading a whole repo or a long document in one pass. Hold off if you need top-tier reasoning for hard, single-shot problems, since Flash is the workhorse tier and a heavier model may still win there. It is also worth comparing against OpenAI’s ChatGPT models if your stack is not already tied to Google. Whatever you decide, price your 2027 plan at the standard rate. The model is a clear upgrade. The discount is the part with a clock on it.
So here is the one thing to do today. If Gemini 3.7 Flash fits your stack, start testing it now while it is half price, and drop a calendar note on December 31 to re-check your numbers before the rate doubles. The upgrade is worth having. The promo is worth using. Just do not let the January reset catch you mid-quarter.
Quick questions
When was Gemini 3.7 Flash released? August 13, 2026.
How much does it cost? $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, then $1.50 and $7.50 from January 1, 2027.
What is its context window? 1 million tokens, with up to 64,000 tokens of output.
Is it better than Gemini 3.6 Flash? On Google’s own coding and agent benchmarks, clearly yes, though those numbers are vendor-reported.
Can it handle images or video? It accepts text, images, audio and video as input, and replies in text only.
About this article. Written and fact-checked by the Brandligo editorial desk. AI tooling was used to gather and cross-check sources; every fact and figure here was verified against the primary sources linked above before publication. Published August 15, 2026. If you spot something out of date, tell us at [email protected].