Gemini 3 Flash

MultimodalDeprecatedAvailable
by Google DeepMindModel ID: gemini-3-flash

Google's April 2026 fast multimodal model. Combines Gemini 3 Pro's reasoning with Flash-tier latency and price. Default model in the Gemini app.

Price Β· 1M in / out
$0.60 / $3.60
Context
1,048,576 tokens
Max. output
65,536 tokens
Input β†’ output
Text + Image + Audio + Video β†’ Text
Developer
Google DeepMind
Updated
September 23, 2026

The provider is phasing this model out.

01

Playground

Try Gemini 3 Flash

Chat

$0.60/1M in
Try Gemini 3 Flash

Send a message. The answer arrives in full once the model is done (no streaming).

Max. answer length (tokens)

This run

at most $0.0037 Β· 0.37 credits reserved

Billed by the tokens actually used; the unused part of the reservation is refunded.

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits. Enough for 27 runs of this model.

02

About Gemini 3 Flash

TL;DRAs of September 23, 2026

Gemini 3 Flash is a model by Google DeepMind in the Multimodal category. On Railwail, Gemini 3 Flash costs $0.60 per 1M input tokens and $3.60 per 1M output tokens. The context window holds 1,048,576 tokens, and one response can be up to 65,536 tokens long.

Announced April 22, 2026, Gemini 3 Flash brings Pro-grade reasoning to the Flash latency tier. 1M-token context, fully multimodal (text, image, audio, video), 65K max output. The default model in the Gemini app and AI Mode in Search. PhD-level reasoning on common benchmarks at a fraction of the cost of 3.1 Pro. Recommended for high-throughput agentic workflows, real-time multimodal chat, RAG and consumer applications.

Background

About Google DeepMind

Founded 2010 Β· Mountain View, USA / London, UK

Google DeepMind is the merged AI research organisation formed in April 2023 by combining Google Brain with DeepMind. Demis Hassabis leads the unit as CEO. Flash variants have been Google's high-throughput tier since Gemini 1.5 Flash (May 2024), with Gemini 2.0 Flash (December 2024), 2.5 Flash (mid-2025) and Gemini 3 Flash (April 2026) representing the progression. DeepMind's seminal papers include 'Attention Is All You Need' (2017), AlphaGo (2016), AlphaFold (2018-2021, Nobel Prize 2024) and the Gemini Technical Report.

Visit Google DeepMind

Architecture

Sparse Mixture-of-Experts Transformer (multimodal, latency-optimized)

Gemini 3 Flash was announced April 22, 2026 as the default Flash-tier model and the new default model in the Gemini app and AI Mode in Search. It is a natively multimodal Sparse MoE Transformer engineered to combine Gemini 3 Pro's reasoning quality with Flash-grade latency, efficiency and cost. Pretraining used Google's TPU v6e infrastructure on a multi-trillion-token mixture of web text, code, books, image-text pairs, audio and video frames. Post-training combined supervised fine-tuning, RLHF, RL against verifiable rewards and distillation from larger Gemini 3.1 Pro teacher models. The architecture preserves Gemini's native multimodality across text, image, audio and video, the full tool-use API and Search grounding, while running at a fraction of Pro pricing. Gemini 3 Flash is the recommended default for high-throughput agentic workflows and consumer-facing multimodal chat.

Parameters
Undisclosed (sparse MoE, smaller and sparser than Gemini 3.1 Pro)
Context
1,048,576 tokens

Capabilities

  • Pro-grade reasoning at Flash latency
  • 1,048,576 token context window
  • Natively multimodal: text, image, audio and video
  • Search grounding and Code Execution built into the API
  • Function calling, JSON schema and parallel tool calls
  • Default model in the Gemini app and AI Mode in Search
  • PhD-level reasoning on common benchmarks
  • Available via Vertex AI, AI Studio, Gemini Enterprise, Antigravity and the Gemini app
  • Strong long-video understanding (hour-long clips)
  • Cross-lingual fluency across 100+ languages
  • Best for: high-throughput agentic workflows, real-time multimodal chat, RAG, consumer applications.

Training & license

Pretrained on a multi-trillion-token mixture of web text, code, books, scientific papers, image-text pairs, audio and video frames. Heavily distilled from larger Gemini 3.1 Pro teacher models. Post-training uses supervised fine-tuning, RLHF and RL against verifiable rewards. Knowledge cutoff in late 2025.

License: Proprietary commercial license via Google AI Studio, Vertex AI and the Gemini app. Free tier available in the Gemini app and AI Mode in Search.

Safety testing: Evaluated under Google DeepMind's Frontier Safety Framework v2 with internal red teams and external evaluators.

Known limitations

  • Below Gemini 3.1 Pro on the hardest reasoning and long-context benchmarks
  • Smaller context window than 3.1 Pro (1M vs 2M)
  • Vision can misread dense tables and handwriting
  • Region availability is rolling out in 2026
  • Audio output not yet supported
03

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Input$0.60 / 1M tokens
Output$3.60 / 1M tokens
  • Billed by the tokens each request actually uses.
  • 1 credit = $0.01

Cost calculator

Price calculator

/ req.
/ req.

Total

$0.24

24 credits

Per request

$0.0024 Β· 0.24 credits

Each request is rounded up to 0.01 credits.

04

API

Call Gemini 3 Flash with your Railwail API key. Use this model ID in the request:
curl https://railwail.com/api/v1/chat/completions \
  -H "Authorization: Bearer $RAILWAIL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3-flash",
    "messages": [
      {
        "role": "user",
        "content": "Explain what a vector database is in two sentences."
      }
    ],
    "max_tokens": 1024
  }'
Set your key as RAILWAIL_API_KEYCreate API key
05

Specifications

Model ID
gemini-3-flash
Category
Multimodal
Input
Text, Image, Audio, Video
Output
Text
Context window
1,048,576 tokens
Max. output
65,536 tokens
Billing
By usage (tokens or GPU time)
Lifecycle
Deprecated
Model size
Undisclosed (sparse MoE, smaller and sparser than Gemini 3.1 Pro)
License
Proprietary commercial license via Google AI Studio, Vertex AI and the Gemini app. Free tier available in the Gemini app and AI Mode in Search.
Catalog entry updated
September 23, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • promptrequired

    User message

    Type: Text
    Default: –
    Allowed values: up to 32,000 characters
  • top_p
    Type: Number
    Default: 0.95
    Allowed values: 0 to 1
  • stream
    Type: Yes/no
    Default: false
    Allowed values: –
  • image_url

    Optional image URL to analyze

    Type: Text
    Default: –
    Allowed values: –
  • max_tokens
    Type: Integer
    Default: 4096
    Allowed values: 1 to 32,000
  • temperature
    Type: Number
    Default: 1
    Allowed values: 0 to 2
  • system_prompt

    Optional system instruction

    Type: Text
    Default: –
    Allowed values: up to 8,000 characters

Tags

  • google
  • deepmind
  • balanced
  • multimodal
  • low-latency
  • long-context
  • 1m-context
06

Use cases

What it is used for

  • Default consumer multimodal chat
  • High-throughput agentic workflows
  • Real-time RAG pipelines
  • Long-video summarisation and search
  • Production coding assistants
  • Voice and audio reasoning backends
  • AI Mode in Search and Antigravity workflows
07

Frequently asked questions

What is Gemini 3 Flash?

Gemini 3 Flash is a model by Google DeepMind in the Multimodal category. On Railwail you can call it with an API key through the Railwail API.

How much does Gemini 3 Flash cost on Railwail?

On Railwail, Gemini 3 Flash costs $0.60 per 1M input tokens and $3.60 per 1M output tokens. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.

What is the context window of Gemini 3 Flash?

The context window of Gemini 3 Flash holds 1,048,576 tokens. One response can be up to 65,536 tokens long.

How fast is Gemini 3 Flash?

There are not enough measured runs of Gemini 3 Flash on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Gemini 3 Flash better than BLIP?

That depends on the task. Gemini 3 Flash (Google DeepMind) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare Gemini 3 Flash and BLIP

Can Gemini 3 Flash process images?

Yes. Gemini 3 Flash accepts images as input in addition to text.

How do I use Gemini 3 Flash through the API?

Create a Railwail API key and send your request with the model ID gemini-3-flash. Code examples for curl, Python and JavaScript are in the API section of this page.

08

Comparable models

All in this category

Use Gemini 3 Flash via the API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.