Gemini 3.6 Flash

gemini-3.6-flash
OfficialAudio-to-Text

Gemini 3.6 Flash is Google's next-generation lightweight workhorse model released in July 2026. Built for agentic workflows, complex coding, and multimodal tasks, it supports a 1M token input context window and a 64K token output limit. Compared to 3.5 Flash, it reduces output token consumption by ~17% with streamlined reasoning steps and tool calls, significantly lowering overall cost and latency for agent execution. Natively handling text, image, video, audio, and PDF inputs, it excels at computer use and multi-tool orchestration.

Token Type Price (USD) Unit
Input 1.5 Per million tokens
Output 7.5 Per million tokens
Cache Read 0.15 Per million tokens

Read Me

Gemini 3.6 Flash API

Gemini 3.6 Flash is a multimodal large language model released by Google on July 21, 2026, as the latest generation in the Gemini Flash series. It supports five input modalities: text, images, video, audio, and PDF, with text output and a context window of 1,048,576 tokens (1M). Google positions it as a model combining the speed of the Flash series with frontier intelligence, suitable for high-throughput multimodal workloads.

On iCreat, you can call gemini-3.6-flash through an API compatible with OpenAI Chat Completions.

This endpoint supports standard system and user messages, multimodal content arrays, streaming output, and prompt caching, allowing applications already using the OpenAI SDK to integrate directly.

Model Positioning

Gemini 3.6 Flash extends the Gemini Flash series from pure-text reasoning to full multimodal understanding, focusing on video analysis, audio processing, image understanding, PDF document parsing, long-context reasoning, and high-throughput dialogue scenarios.

For teams looking to use a multimodal model at a lower cost while maintaining a million-level context window, it is a practical choice. Gemini 3.6 Flash supports PDF input but does not possess image, video, or audio generation capabilities, supporting only text output.

Core Capabilities

Multimodal Understanding

Gemini 3.6 Flash can simultaneously process text, image, video, audio, and PDF inputs, correlating cross-modal information within a single request, suitable for comprehensive content analysis, video summarization, audio transcription and understanding, PDF document parsing, and other scenarios.

Long-Context Reasoning

The 1,048,576 tokens (1M) context window can accommodate large documents, codebases, and long videos, suitable for single-pass reasoning in ultra-long input scenarios.

Prompt Caching

Repeatedly sent prompt prefixes can hit the cache, with cache read prices at only 10% of the standard input, suitable for high-frequency call scenarios carrying fixed system prompts.

Streaming Output

Supports the stream parameter for real-time streaming returns, allowing applications to present content progressively during generation without waiting for the complete response.

Dynamic Thinking

Dynamic thinking is enabled by default, where the model automatically performs reasoning before answering, without the need to manually configure the thinking parameter.

OpenAI Compatibility

Fully compatible with the OpenAI Chat Completions API, standard parameters such as temperature, max_tokens, top_p, frequency_penalty, presence_penalty, stop, n, response_format can all be used without additional adaptation.

Pricing

Token Type Price
Input $1.50 / million tokens
Output $9.00 / million tokens
Cache Read $0.15 / million tokens

Application Scenarios

  • Multimodal Content Analysis: Simultaneously input video, audio, images, PDF, and text for cross-modal comprehensive analysis and understanding.
  • Long Document Processing: Utilize the 1M context window to process ultra-long documents, codebases, PDFs, and technical materials.
  • Code Generation and Understanding: Knowledge cutoff is March 2026, covering mainstream programming languages and frameworks, supporting function calling and structured output.
  • Intelligent Dialogue Systems: Build multi-turn dialogue systems, supporting streaming output and prompt caching to optimize costs.
  • Content Summarization Generation: Generate text summaries and structured descriptions for video, audio, image, and PDF content.
  • Agent Workflows: Supports function calling and structured output, suitable for Agent orchestration and multi-step task coordination.

Model Comparison

Gemini 3.6 Flash vs. Gemini 3.5 Flash

Dimension Gemini 3.6 Flash Gemini 3.5 Flash
Context Window 1,048,576 tokens 1,048,576 tokens
Input Modalities Text/Image/Video/Audio/PDF Text/Image/Video/Audio
Cache Read $0.15 / million tokens $0.15 / million tokens
Knowledge Cutoff March 2026 January 2026
Release Date July 21, 2026 May 19, 2026
Positioning Latest generation of Flash series, adds PDF input Flash speed + frontier intelligence

Gemini 3.6 Flash vs. Gemini 3.1 Pro, GPT 5.6 Sol

Dimension Gemini 3.6 Flash Gemini 3.1 Pro GPT 5.6 Sol
Positioning Latest generation of Flash series, multimodal understanding Flagship multimodal, professional-level tasks Flagship reasoning, complex professional tasks
Context Window 1,048,576 tokens 1,048,576 tokens ~1,050,000 tokens
Max Output 65,536 tokens 65,536 tokens 128,000 tokens
Official Input Modalities Text, Image, Video, Audio, PDF Text, Image, Video, Audio, PDF Text, Image
Reasoning Control Dynamic thinking (enabled by default) Dynamic thinking (enabled by default) Adjustable depth (none→max)
Input Price $1.50 / million tokens $2.50 / million tokens $5.00 / million tokens
Output Price $9.00 / million tokens $15.00 / million tokens $30.00 / million tokens
Best Suited For High-throughput multimodal workloads Professional-level multimodal tasks Ultra-long output and high-end reasoning

Why Choose Gemini 3.6 Flash?

When workloads require multimodal understanding (video/audio/image/PDF), a million-level context window, and attention to call costs simultaneously, Gemini 3.6 Flash can be chosen. Five-modal input, cache pricing at only 10% of input, 1M context window, and dynamic thinking reasoning make it suitable as the reasoning core for high-throughput multimodal applications.

Through iCreat, teams can integrate directly using the OpenAI-compatible endpoint without self-deployment. The Playground is suitable for prompt-level testing; when integration with tools, state, and multi-turn dialogue is required, the API should be used to build production workflows.

Specifications

Category Description
Model Name Gemini 3.6 Flash
Developer Google
Model ID gemini-3.6-flash
Release Date July 21, 2026
Model Type Multimodal Large Language Model
Context Window 1,048,576 tokens (1M)
Max Output 65,536 tokens
Official Input Modalities Text, Image, Video, Audio, PDF
Output Modality Text
Knowledge Cutoff March 2026
Reasoning Capability Dynamic thinking (enabled by default)
iCreat API Capabilities Compatible with OpenAI Chat Completions API, streaming output, prompt caching
Main Applicable Tasks Multimodal analysis, long-context reasoning, code generation, intelligent dialogue, content summarization, PDF parsing

Architecture

Gemini 3.6 Flash belongs to the Google Gemini Flash series, adopting the multimodal design of the Gemini architecture, supporting cross-modal input encoding and unified text output. The model processes text tokens and multimodal input tokens within the same context window, automatically performing internal reasoning before answering through the dynamic thinking mechanism.

Compared to Gemini 3.1 Pro, 3.6 Flash optimizes reasoning efficiency and call costs while maintaining the same context window and input modalities; compared to Gemini 3.5 Flash, 3.6 Flash adds PDF input support, with the knowledge cutoff updated to March 2026.

Notes

For long outputs and multi-turn dialogues, streaming output is recommended, allowing applications to present progress rather than waiting for the complete response. When setting request budgets, input, output, and cache read usage should be tracked simultaneously.

The 1,048,576 tokens context is a shared request space. When loading multimodal materials, space needs to be reserved for system instructions, dialogue state, and the final answer. Video and audio inputs consume a large number of tokens, so the entire window cannot be filled with source materials.

Multimodal inputs require accessible URLs; direct local file uploads are not supported. The input_audio.format field must match the actual audio format, commonly supporting wav, mp3, flac, ogg, etc.

During evaluation, the completion of the entire task should be observed. It is necessary to measure multimodal understanding accuracy, long-context recall rate, reasoning quality, latency, token usage, and cache hit rate.

Frequently Asked Questions

What is the difference between Gemini 3.6 Flash and Gemini 3.5 Flash?

3.6 Flash adds PDF input support, and the knowledge cutoff is updated from January 2026 to March. Both have the same pricing (input $1.50, output $9.00, cache $0.15). 3.6 Flash is the latest generation of the Flash series, while 3.5 Flash does not support PDF input.

Can I call it directly using the OpenAI SDK?

Yes.

Set base_url to https://api.icreat.ai/v1 and api_key to your iCreat API Key, and it can be used without additional adaptation.

Standard parameters such as temperature, max_tokens, top_p, etc., can all be used.

How to enable streaming output?

Add "stream": true to the request body. The response will be returned chunk by chunk in SSE (Server-Sent Events) format, with each chunk containing delta.content, and the finish_reason of the last chunk being stop.

How is cache read billed?

Repeatedly sent prompt prefixes can hit the cache, priced at $0.15 / million tokens, which is only 10% of the standard input price of $1.50. Suitable for high-frequency call scenarios carrying fixed system prompts.

What audio formats are supported?

The input_audio.format field specifies the audio format, commonly supporting wav, mp3, flac, ogg, etc. This field must match the actual audio format.

How does it compare to GPT 5.6 Sol?

Gemini 3.6 Flash is significantly cheaper (input 3.3 times cheaper, output 3.3 times cheaper) and supports more input modalities (video/audio/PDF), suitable for high-throughput multimodal workloads. GPT 5.6 Sol is stronger in complex reasoning and ultra-long output (128K), suitable for high-end reasoning and professional tasks.

Can the modalities parameter be set to other values?

Currently, only ["text"] is supported. Gemini 3.6 Flash is a text output model and does not have the ability to generate images, video, or audio.