DeepSeekDSH
Independent community guideNot affiliated with DeepSeek.Official source snapshot

Gemini 3.8 Flash in Production: Multimodal Input, Coding Evidence and Flash Economics

Evaluate Gemini 3.8 Flash using official coding results, current independent evidence, 1M multimodal context and the difference between 2026 promotional and 2027 regular API pricing.

Independent editorial review: DeepSeekDSHSource checked: 0.1.5-rc.2 · 2026-09-11
On this page
Editorial snapshot · Sep 12, 2026

Gemini 3.8 Flash combines low token prices in 2026 with native input for text, images, audio, video and PDFs, which makes it a practical first-pass model for multimodal and high-volume workflows. Benchmark headlines need version labels: launch-era and current composite-index values use different index revisions, so a lower current number is not evidence that the deployed model suddenly regressed.

Flash economics dashboard

Gemini 3.8 Flash: near-frontier workflows at Flash pricing

Use official benchmark results and the current independent index revision without mixing launch-era score scales.

Sep 11, 2026

Google's current model-card evidence

Google positions 3.8 Flash as its most intelligent workhorse model and keeps the 2026 introductory price at the 3.7 Flash level.

DeepSWE v1.173.7%
Vals Finance Agent v261.4%
Legal Agent all-pass10.0%
Context1Mmultimodal input

Google's model card shows serious software-engineering capability at Flash-level pricing

Google DeepMind's September 2026 model card shows meaningful gains over Gemini 3.7 Flash across software engineering and agentic knowledge workflows.

BenchmarkGemini 3.8 FlashGemini 3.7 Flash
DeepSWE v1.173.7%65.3%
Terminal-Bench 2.189.4%85.8%
Terminal-Bench 4.019.1%11.2%
HLE-Verified54.9%53.6%
OSWorld 2.059.0%50.6%
GDPVal-AA v21545 Elo1482 Elo

Section sources: Google DeepMind — Gemini 3.8 Flash model card · Google — Introducing Gemini 3.8 Flash

Coding evidence is strong enough to make Gemini 3.8 Flash a practical workhorse

DeepSWE v1.1 at 73.7% puts Gemini 3.8 Flash close to the largest models in Google's comparison table. Terminal-Bench 2.1 is also very strong at 89.4%. These are agentic software-engineering tasks, not autocomplete benchmarks.

The model is therefore a plausible workhorse for code agents that need to read repositories, use tools and carry tasks over multiple steps — especially when throughput and price matter.

Section sources: Google DeepMind — Gemini 3.8 Flash model card

Terminal-Bench 2.1 and 4.0 demonstrate why version labels are mandatory

Those two numbers are not a contradiction. They demonstrate why benchmark version labels are essential. Terminal-Bench 4.0 is a different and harder evaluation than 2.1, so the raw percentages cannot be compared as if they were the same exam.

This is especially important when comparing DeepSeek launch tables that still include Terminal Bench 2.1 with newer OpenAI and Artificial Analysis tables that emphasize Terminal-Bench 4.0.

Benchmark hygiene

Never write 'Model A scored 89 while Model B scored 58, therefore A is better' until you have confirmed the benchmark version, harness and tool configuration are the same.

Section sources: Google DeepMind — Gemini 3.8 Flash model card · DeepSeek — V4 Pro GA release · OpenAI — GPT-6 Astra

Native audio, video, image and PDF input offer the clearest production advantage

Gemini 3.8 Flash accepts text, images, audio and video with up to a 1M-token context window and 64K text output. That makes it particularly useful for workflows where the agent must reason over screenshots, charts, long documents or video rather than only repository text.

Google also reports strong chart reasoning and long-video results, though those are vendor-reported and should be tested against the exact media types in your own workload.

Section sources: Google DeepMind — Gemini 3.8 Flash model card

Budget with both the 2026 promotional rate and the 2027 regular rate

Google lists an introductory price of $0.75 per million input tokens and $3.75 per million output tokens for Gemini 3.8 Flash, with regular prices shown as $1.50 and $7.50. That is materially below premium frontier-model pricing.

API itemGemini 3.8 Flash
Intro input / 1M$0.75
Intro output / 1M$3.75
Regular input / 1M$1.50
Regular output / 1M$7.50
Context1M tokens
Max output64K tokens

Section sources: Google DeepMind — Gemini 3.8 Flash model card · Google — Introducing Gemini 3.8 Flash

Use Gemini where multimodal throughput matters, then keep a stronger tier for hard terminal failures

The model is most attractive when you need a low-cost multimodal agent, large context and credible coding performance at the same time. Examples include UI QA with screenshots, long-document analysis, code agents that also inspect visual outputs, and high-volume knowledge workflows.

For the hardest autonomous terminal tasks, the current Terminal-Bench 4.0 score shows there is still substantial room above Flash-class performance. Use your own regression set before replacing a stronger frontier model solely because 3.8 Flash is cheaper.

Section sources: Google DeepMind — Gemini 3.8 Flash model card · Artificial Analysis — Intelligence Index v4.3

Related model analysis