Google publishes Gemini 3.8 Flash model card

Google publishes Gemini 3.8 Flash model card
News

Google DeepMind has published a model card for Gemini 3.8 Flash, the latest iteration in its Gemini 3 family. The page is dated September 2, 2026. A model card is not a conventional product launch announcement: it documents the model’s intended use, limitations, distribution channels, evaluation approach and safety results. The publication nevertheless gives users and developers a concrete picture of what Google says the model is and how it should be assessed.

Google describes Gemini 3.8 Flash as building on Gemini 3.7 Flash, with performance advances in software engineering and agentic knowledge workflows. The model keeps configurable effort levels, allowing users to adjust the balance between quality, cost and latency. It accepts text, images, audio and video, supports a context window of up to one million tokens and can produce up to 64,000 output tokens. These are specifications reported by Google DeepMind in the model card.

The model is listed across a broad set of Google distribution channels: the Gemini app, Gemini Enterprise Agent Platform, Google AI Studio, Gemini API, Google AI Mode and Google Antigravity. That list matters because access, terms and operational controls can differ between consumer products, enterprise services and developer APIs. Google also says downstream providers can access its models through an API subject to the relevant terms.

The model card reports evaluations covering coding, knowledge work, multimodal capabilities, long context, computer use and scientific reasoning. Google says Gemini 3.8 Flash performs similarly to Gemini 3.7 Flash on its internal safety and tone evaluations, with a slight regression in non-English safety performance. The company also notes possible hallucinations, occasional slowness or timeouts and higher token use at elevated effort levels. Its stated knowledge cutoff is March 2026, although some domains may have information only through January 2025.

For AI users and businesses, the practical news is not simply a new model number. The card makes capability, availability and limitations visible in one primary document, which helps teams compare the model with existing systems and design tests before adoption. Google’s benchmarks and safety findings remain provider-reported evidence, not a guarantee of performance in a particular workflow. Developers should therefore measure accuracy, latency, cost and failure modes on their own tasks, especially when agents can use tools or act on documents.