What Google released on 2 September
Gemini 3.8 Flash is Google's third Flash model in six weeks, and per Google beats its predecessor, 3.7 Flash, on every published benchmark, and even beats Anthropic's Claude Opus 5 on three of them. The model handles text, image, audio, video, and PDF input with a 1 million token context window and up to 64,000 tokens of output, tuned for long-running coding tasks and autonomous agents. In parallel, Google released Gemini 3.8 Flash Cyber, a cybersecurity-specialized variant that, per the company, delivers 'frontier-level performance in vulnerability detection and automated patching'.
Access to this Cyber variant is heavily restricted: it's available exclusively through the new 'Fairwind Program', reserved for government agencies, critical infrastructure operators, and software maintainers who have to apply and be approved. That structurally resembles the access model OpenAI introduced with Daybreak Red for its cyber model, GPT-5.6-Cyber, which we covered recently: at Google too, the most advanced security capability stays reserved for a tightly limited circle, not the regular customer base.
The detail sitting in Google's own documentation
The actually relevant information for budget planning doesn't sit in the product announcement itself, but in Google's technical pricing documentation. It states, verbatim: 'Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.' The currently advertised price of $0.75 per million input tokens and $3.75 per million output tokens is thus explicitly labeled a time-limited introductory offer, not a permanent standard price. On 1 January 2027, the price exactly doubles on both components.
What's notable is what's missing from the documentation: any justification for the price doubling. Unlike Nvidia's price increase, which we covered recently - where the vendor at least pointed to rising memory chip costs as the cause - Google names no external cost driver here. The higher price simply takes effect automatically on the stated date, with no further announcement, explanation, or separate notice to existing customers planned.
The additional, indirect cost factor
Beyond the direct price doubling, Google's documentation names a second, less obvious cost factor: for pure efficiency workloads, Google explicitly recommends staying on Gemini 3.7 Flash, 'which remains fully supported for efficiency-first workloads'. The reason: Gemini 3.8 Flash 'might use more tokens to maximize performance, especially at higher effort levels' on complex tasks. In practice, that means: even before the price doubling at the turn of the year, the actual cost increase from switching from 3.7 to 3.8 Flash can turn out larger than the plain price difference suggests, because the newer model can consume more tokens for the same task.
Why this matters now, not just in 2027
A price step that only takes effect in roughly four months looks low-urgency at first glance. But the decisive point is that this information is already fully known and dated today - unlike, say, Nvidia's memory chip price increase, where the underlying market development remains uncertain, the Gemini 3.8 Flash price step can be factored exactly into your own cost planning for 2027. Anyone building applications or workflows on Gemini 3.8 Flash today is, in effect, already planning around a model whose cost doubles in under four months.
For companies integrating Gemini models into their own applications via API, this is a concrete prompt to review your own cost calculations for 2027 before fixing contracts, internal budgets, or customer quotes based on the current introductory prices. A pricing model explicitly labeled an 'introductory price' shouldn't be treated as a permanent basis in any mid-term calculation - regardless of the specific vendor.
What this means in practice
- For any cost calculation based on Gemini 3.8 Flash, plan from now on with the doubled price that applies from 1 January 2027 - not with the current introductory price, which is explicitly time-limited.
- Check whether your own use case actually needs Gemini 3.8 Flash, or whether Gemini 3.7 Flash - which Google says remains fully supported for efficiency workloads - is sufficient, particularly for tasks that don't need the higher model performance.
- As a general practice for API pricing, check whether an advertised price is labeled an 'introductory price' or time-limited offer before it flows into a mid-term budget or quote calculation - this pattern is showing up increasingly among AI vendors.
- For your own customer quotes or internal budgets that price in AI costs, build in a buffer for announced but not-yet-effective price changes, rather than relying solely on the price current at the time of calculation.
The real value of this analysis isn't a criticism of Google's pricing - a time-limited introductory offer is a legitimate, common business model. The point is that this information sits in Google's own technical documentation, but barely appears in public coverage of the product launch - anyone reading only the benchmark and feature headlines misses a figure that's already concrete enough today to factor into your own planning for 2027.