Gemini 3.8 Flash: What's New, How Much It Costs, and When You Should Upgrade
Google launched Gemini 3.8 Flash with improvements in agents, programming, and reasoning. It retains the same price as version 3.7, but may consume more tokens per task.
Gemini 3.8 Flash is now generally available, coming just three weeks after Gemini 3.7 Flash. Google presents it as its most capable Flash model for long-running software engineering, autonomous agents, and complex business workflows. The update specifically improves the execution of agent-based and programming tasks, but introduces an important caveat for those evaluating costs: it can reason more deeply and, as a result, consume more tokens per task.
That changes the practical question. It's not enough to look at the price per million tokens, because Gemini 3.8 Flash has the same introductory price as 3.7. To decide whether it's worth migrating, you have to measure final quality, total number of tokens, latency, tool calls, errors, and human rework. Google even recommends keeping Flash at version 3.7 when computing efficiency is the top priority.
In a nutshell
- Google released Gemini 3.8 Flash on September 2, 2026, as a stable, generally available model.
- It supports up to 1,048,576 input tokens and 65,536 output tokens.
- It accepts text, images, video, audio, and PDFs; its documented output is text.
- The standard introductory API price is US$0.75 per million incoming tokens and US$3.75 per million outgoing tokens through December 31, 2026.
- Effective January 1, 2027, the published standard rates will be US$1.50 for incoming transactions and US$7.50 for outgoing transactions per million tokens.
- The unit price is the same as for Gemini 3.7 Flash, but Google notes that version 3.8 may use more tokens to maximize performance.
- The available reasoning levels are LOW, MEDIUM, and HIGH; MEDIUM is the default setting in Google's business guide. MINIMAL is not supported.
- Gemini 3.7 Flash is still supported, and no end-of-life date has been announced.
What Really Changes with Gemini 3.8 Flash
Gemini 3.8 Flash does not increase the context window or the maximum output compared to 3.7 Flash. Nor does it reduce the unit price. The change is primarily in the model's performance: Google reports greater accuracy and reliability in software engineering, agent-based tasks, and multi-step specialized reasoning.
Google Cloud’s own guide describes the Flash line as the “workhorse” agent of the Gemini 3 family. In version 3.8, this approach shifts toward longer-running and more complex tasks: agents that use tools, multi-step programming flows, multimodal analysis, and business processes where a single response is not sufficient.
To understand why this matters, it's a good idea to first review What Are Artificial Intelligence Agents and How Do They Work?. A more capable model can improve the agent, but its ultimate reliability also depends on permissions, tools, data, validations, and execution rules.
Gemini 3.8 Flash vs. Gemini 3.7 Flash
| Appearance | Gemini 3.8 Flash | Gemini 3.7 Flash |
|---|---|---|
| Status | General Availability | General Availability |
| Maximum context | 1,048,576 tokens | 1,048,576 tokens |
| Maximum output | 65,536 tokens | 65,536 tokens |
| Levels of Thinking | LOW, MEDIUM, HIGH | LOW, MEDIUM, HIGH |
| Standard Introductory Price | US$0.75 input / US$3.75 output via MTok | US$0.75 input / US$3.75 output via MTok |
| Main Focus | Agents, programming, and complex multi-step reasoning | General agents, multistage orchestration, and scheduling |
| Token Usage | It can be larger to maximize performance | Google keeps it as an option when computing efficiency is a priority |
| Announced Retirement | No | No |
This comparison prevents a hasty conclusion: Upgrading to 3.8 is not required. If version 3.7 already meets the quality criteria and the process is optimized for cost or latency, continuing to use it may be a valid decision. If the bottleneck lies in reasoning, tool execution, or long-running programming, version 3.8 warrants evaluation with real workloads.
What Do the Benchmarks Published by Google Show?
Google reports significant improvements in several tests compared to Flash 3.7, but not in all of them. This is a helpful indication because it prevents Flash 3.8 from being presented as a universal improvement.
| Evaluation | Gemini 3.8 Flash | Gemini 3.7 Flash |
|---|---|---|
| Terminal-bench 2.1 | 90,8% | 81,6% |
| SWE-Bench Pro | 61,6% | 60,4% |
| SWE-Atlas | 51,9% | 48,0% |
| τ³-bench Banking | 38,1% | 30,9% |
| Multimodal CharXiv | 86,2% | 84,5% |
| GDP.pdf | 35,0% | 34,0% |
| Humanity's Last Exam | 45,4% | 45,7% |
The results suggest clearer gains in terminal operations, agents, and certain multimodal jobs. Humanity’s Last Exam, on the other hand, appears slightly below 3.7 in Google’s table. This reinforces a necessary practice: benchmarks are meant to guide what to test, not to automatically predict the performance of a specific business workflow.
Price of Gemini 3.8 Flash and the actual cost per task
| Concept | Through December 31, 2026 | Effective January 1, 2027 |
|---|---|---|
| Home | US$0.75 / 1M tokens | US$1.50 / 1M tokens |
| Output, including reasoning tokens | US$3.75 / 1M tokens | US$1.75 / 1 million tokens |
| Cached context reading | US$1 × TP × 0.075 / 1 million tokens | US$0.15 / 1M tokens |
The point that's easiest to miss is that The price per token and the cost per task are not the same. Google is keeping the introductory unit price at 3.8—the same as the initial 3.7—but also acknowledges that the model may use more tokens to achieve better performance, especially at high effort levels.
That is why a cost comparison should include, at a minimum:
- input tokens;
- output and reasoning tokens;
- context served from cache;
- number of calls to tools;
- total execution time;
- retakes or failures;
- minutes of human review and correction.
A model that consumes more tokens may be more cost-effective if it reduces the number of iterations or prevents rework. The opposite can also be true: a simple task may become more expensive without yielding any useful improvement. The only way to know is to measure the entire process.
LOW, MEDIUM, and HIGH: How to Manage Your Thinking and Consumption
Gemini 3.8 Flash allows you to adjust the level of reasoning by thinking_level. Google Cloud documents three valid values: LOW, MEDIUM, and HIGH, with MEDIUM as the default for 3.8 Flash on its enterprise platform. The MINIMAL level is not available and results in a validation error.
| Level | When to evaluate it | What to Watch For |
|---|---|---|
| LOW | Retrieval, fast searches, sorting, or latency-sensitive tasks | Ensure that reduced reasoning ability does not affect accuracy or the ability to follow instructions |
| MEDIUM | Factors and general trends that require a balance between quality and consumption | Total tokens, tools, and consistency across runs |
| HIGH | Complex problems, multi-step reasoning, or demanding multimodal analysis | Higher power consumption, latency, and cost-effectiveness |
It is not advisable to set HIGH as a universal configuration. In a production system, the level of reasoning can be tailored to specific task types: less effort for predictable processes and more effort only where complexity warrants it.
Features Available in the API
The official Google product page documents a comprehensive set of tools for Gemini 3.8 Flash:
- context caching;
- code execution;
- computer use in Preview mode;
- file search;
- function calling;
- grounding with Google Search;
- grounding with Google Maps;
- structured outings;
- configurable reasoning;
- URL context.
It's also important to record what no Here's this variant: The documentation indicates that it does not support image generation, audio generation, or the Live API. It can accept images, audio, video, and PDFs as input, but the model's output is text.
This helps in designing the right architecture. If a process needs to generate an image or maintain a real-time audio experience, that functionality must be handled by another model or component; it should not be assumed simply because Gemini 3.8 Flash supports multimodal input.
Where is Gemini 3.8 Flash available?
Google lists Gemini 3.8 Flash as a stable, generally available model. For developers, it appears in the Gemini API and Google AI Studio, and the announcement also lists Android Studio, Google Antigravity, and Stitch among the access points. In the enterprise environment, Google offers it through the Gemini Enterprise Agent Platform.
For consumers, Google announced access for Google AI Pro and Ultra subscribers across platforms such as the Gemini app, AI Mode in Search, and Gemini in Google Sheets. The exact availability of a feature may vary by product, account, or region, so a company should verify the platform where it plans to implement the feature before designing a permanent dependency.
What to Check When Migrating from Gemini 3.7 or 3.6 Flash
Google has published a specific migration guide. Although changing the model ID is straightforward, there are API rules that could break legacy clients.
- Change the model ID to
gemini-3.8-flash. - Replace
thinking_budgetbythinking_level. Valid values are LOW, MEDIUM, and HIGH. - Remove obsolete sampling parameters. Google states that
temperature,top_kytop_pare ignored by the backend in these Gemini 3 conventions. - Remove unsupported parameters.
frequency_penalty,presence_penaltyycandidate_countcan cause active errors. - Review the conversation history. The model's predefined shifts are not supported, and a history cannot end with a role shift.
model. - Valid function call. Function responses must strictly match the identifier, name, and count of the previous call.
- Run regression tests. Test quality, latency, power consumption, structured format, and tool performance before routing significant traffic.
When Should You Try Version 3.8, and When Might It Be Better to Stick with Version 3.7?
| Status | Reasonable decision | Reason |
|---|---|---|
| Complex agents with multiple tools | Try 3.8 Flash | Google Reports Improvements in Agentive Tasks and Reliability |
| Long-running programs or terminal work | Try 3.8 Flash | It is one of the areas that has shown the most significant improvement in official assessments |
| Stable flow that is highly sensitive to consumption | Compare before migrating; consider 3.7 | Google warns of higher token consumption in version 3.8 and recommends version 3.7 when computational efficiency is a priority |
| Simple and Large-Scale Tasks | Don't assume that 3.8 is necessary | The additional capacity may not offset the actual cost |
| Process with common reasoning errors in 3.7 | Evaluate 3.8 Using Real-Life Examples | The improvement can reduce the number of iterations and rework |
Business use cases where it can add value
Software Development and Maintenance Specialists
The official stance on 3.8 Flash prioritizes software engineering and long-running agents. It can be useful for analyzing a codebase, investigating an incident, running authorized tools, and proposing changes. In production, changes must still go through testing, code review, and deployment checks.
Documentary Research and Knowledge Flows
The combination of broad context, PDFs, search, grounding, and tools makes it possible to build workflows that review documents and sources before producing a summary. This can reduce manual work, but it does not eliminate the need to validate key facts or replace the company’s primary sources.
Multimodal processes
Gemini 3.8 Flash can process text, images, audio, and video. This makes it possible to evaluate processes where information is not provided in a single format—for example, reviewing audiovisual material alongside written documentation. The benefit must be weighed against the cost of processing large files and the level of reasoning required.
Operational agents with tools
Function calls, searches, and computer use expand the range of actions an agent can perform. As an agent’s ability to act increases, so does the impact of an error. For actions that affect customers, money, permissions, or sensitive information, it is advisable to implement the principle of least privilege, maintain activity logs, and require human approval.
The guide to Security and Privacy When Using AI in a Business It includes basic controls that apply regardless of the provider.
Limitations documented by Google
Google DeepMind's model card lists limitations that should be part of any serious evaluation:
- Hallucinations: It may produce incorrect information, just like other foundation models.
- Increased consumption: In certain cases, it uses more tokens to maximize performance, especially when the workload is high.
- Latency and timeouts: Google acknowledges that there may be instances of slow performance or delays.
- Information that is not entirely up to date: The specified cutoff date is March 2026, although Google notes that some domains may behave as if they were limited to January 2025.
- "Computer Use" is still in Preview: It should not be treated as a stable capacity equivalent to the GA core of the model.
- It does not generate images or audio: Its multimodal nature is most evident at the entrance.
In addition, the model card notes that, in automated evaluations, safety performance in languages other than English showed a slight decline compared to 3.7, although Google’s manual review did not find any serious concerns. For a business that operates primarily in Spanish, this is an additional reason to test instructions, rejections, and ambiguous cases in the actual language of use.
Greater agency requires better controls
Google claims to have strengthened security and robustness measures against instruction injection. That does not make an agent a foolproof system. If the model can navigate, read documents, or use tools, a malicious instruction included in an external source could attempt to divert the flow.
An enterprise implementation should retain at least the following controls:
- minimum permissions per tool;
- distinction between reading and writing;
- human approval for actions that are difficult to reverse;
- call logs, tools, and results;
- cost limits and number of steps;
- validation of key data against the original source;
- challenging tests before expanding autonomy.
How to Evaluate Gemini 3.8 Flash Using Real Data
- Choose between 20 and 50 representative tasks. It includes easy, difficult, and ambiguous cases, as well as known bugs in version 3.7.
- Define a correct answer or an acceptance criterion. Without a baseline, you won't be able to measure improvement.
- Run 3.7 and 3.8 under comparable conditions. Record the level of reasoning used.
- Measures the actual cost per task. Don't limit yourself to the MTok rate.
- It logs tool and formatting errors. An agent may make a mistake even if the literal response seems correct.
- Count the human corrections. Less rework can offset higher token consumption.
- Evaluate latency. A better model isn't necessarily better for an interaction that requires an immediate response.
- Keep a control group. Don't migrate all the traffic until you've verified that the improvement is sustained.
Metrics That Actually Help You Make Decisions
| Metrics | What is the answer? |
|---|---|
| Completion Rate | Does it achieve the objective without further intervention? |
| Tokens per task | How much computing power does it actually consume? |
| Cost per Successful Task | How much does it cost to get a usable result—not just a call? |
| Total latency | Is the response time appropriate for the process? |
| Tool-Use Errors | Do you use tools consistently? |
| Human Rework | How long does it take for a person to correct their mistakes? |
| Critical Incidents | Are you making mistakes that prevent you from safely automating the workflow? |
Frequently Asked Questions
Is Gemini 3.8 Flash in preview?
No. Google lists it as a stable, generally available model as of September 2, 2026. Some associated features, such as computer use, do remain in Preview.
Is Gemini 3.8 Flash more expensive than 3.7 Flash?
The standard price per million tokens remains the same during the introductory period, and Google publishes the same future rate for both. However, version 3.8 may use more tokens per task, so the effective cost may indeed be higher. It may also be lower if the higher quality reduces the number of retries and corrections.
Should we upgrade from Flash 3.7 to 3.8?
No. Google has not announced an end-of-life date for Flash 3.7 and explicitly lists it as an option when seeking greater computing efficiency. The migration should be justified based on the organization's own results.
Does Gemini 3.8 Flash generate images?
No. It can accept images as input, but the official documentation states that image generation is not supported. The same applies to audio generation and the Live API.
How much context does Gemini 3.8 Flash support?
The official report documents 1,048,576 incoming tokens and up to 65,536 outgoing tokens.
The update warrants testing, not an automatic migration
Gemini 3.8 Flash is a significant update for agents, scheduling, and complex multi-step tasks. It retains the core components of version 3.7—context, output, and unit price—while aiming to improve the quality of agent-driven work. The cost of this improvement may manifest as higher token usage and, in some cases, increased latency.
For a company, the best approach is to compare versions 3.8 and 3.7 using real-world tasks. If version 3.8 completes more processes correctly and reduces rework, the higher resource consumption may be justified. If the task is already working well and the goal is to minimize computing resources, version 3.7 remains a valid and officially supported option.
You might also be interested in Artificial Intelligence
Resources and Support
We offer a variety of resources to help you become familiar with and understand your resources and opportunities, and to enjoy the benefits of having a digital infrastructure. In addition, we have social programs to support startups and projects with exclusive benefits.
