Gemini 3 Flash

MultimodalAuslaufendVerfĂŒgbar
von Google DeepMindModell-ID: gemini-3-flash

Google's April 2026 fast multimodal model. Combines Gemini 3 Pro's reasoning with Flash-tier latency and price. Default model in the Gemini app.

Preis · 1 Mio. In/Out
$ 0.60 / $ 3.60
Kontext
1'048'576 Token
Max. Ausgabe
65'536 Token
Eingabe → Ausgabe
Text + Bild + Audio + Video → Text
Entwickler
Google DeepMind
Aktualisiert
23. September 2026

Der Anbieter lÀsst dieses Modell auslaufen.

01

Playground

Gemini 3 Flash ausprobieren

Chat

$ 0.60/1 Mio. In
Gemini 3 Flash ausprobieren

Schick eine Nachricht. Die Antwort kommt vollstÀndig, sobald das Modell fertig ist (ohne Streaming).

Max. AntwortlÀnge (Token)

Dieser Lauf

höchstens $ 0.0037 · 0.37 Credits vorgemerkt

Abgerechnet werden die tatsÀchlich verbrauchten Token, der Rest der Vormerkung wird erstattet.

Neu hier?

10 Gratis-Credits ($ 0.10) bei Anmeldung mit Google

Nutzbar 24 Stunden nach der Anmeldung, bis zu 5 LĂ€ufe pro Tag und höchstens 2 Credits je Lauf. Andere Anmeldearten starten ohne Guthaben. Reicht fĂŒr 27 LĂ€ufe dieses Modells.

02

Über Gemini 3 Flash

Kurz gesagtStand: 23. September 2026

Gemini 3 Flash ist ein Modell von Google DeepMind aus der Kategorie Multimodal. Über Railwail kostet Gemini 3 Flash $ 0.60 je 1 Mio. Input-Token und $ 3.60 je 1 Mio. Output-Token. Das Kontextfenster umfasst 1'048'576 Token, eine Antwort bis zu 65'536 Token.

Announced April 22, 2026, Gemini 3 Flash brings Pro-grade reasoning to the Flash latency tier. 1M-token context, fully multimodal (text, image, audio, video), 65K max output. The default model in the Gemini app and AI Mode in Search. PhD-level reasoning on common benchmarks at a fraction of the cost of 3.1 Pro. Recommended for high-throughput agentic workflows, real-time multimodal chat, RAG and consumer applications.

Hintergrund

Über Google DeepMind

GegrĂŒndet 2010 · Mountain View, USA / London, UK

Google DeepMind is the merged AI research organisation formed in April 2023 by combining Google Brain with DeepMind. Demis Hassabis leads the unit as CEO. Flash variants have been Google's high-throughput tier since Gemini 1.5 Flash (May 2024), with Gemini 2.0 Flash (December 2024), 2.5 Flash (mid-2025) and Gemini 3 Flash (April 2026) representing the progression. DeepMind's seminal papers include 'Attention Is All You Need' (2017), AlphaGo (2016), AlphaFold (2018-2021, Nobel Prize 2024) and the Gemini Technical Report.

Google DeepMind besuchen

Architektur

Sparse Mixture-of-Experts Transformer (multimodal, latency-optimized)

Gemini 3 Flash was announced April 22, 2026 as the default Flash-tier model and the new default model in the Gemini app and AI Mode in Search. It is a natively multimodal Sparse MoE Transformer engineered to combine Gemini 3 Pro's reasoning quality with Flash-grade latency, efficiency and cost. Pretraining used Google's TPU v6e infrastructure on a multi-trillion-token mixture of web text, code, books, image-text pairs, audio and video frames. Post-training combined supervised fine-tuning, RLHF, RL against verifiable rewards and distillation from larger Gemini 3.1 Pro teacher models. The architecture preserves Gemini's native multimodality across text, image, audio and video, the full tool-use API and Search grounding, while running at a fraction of Pro pricing. Gemini 3 Flash is the recommended default for high-throughput agentic workflows and consumer-facing multimodal chat.

Parameter
Undisclosed (sparse MoE, smaller and sparser than Gemini 3.1 Pro)
Kontext
1'048'576 Token

FĂ€higkeiten

  • Pro-grade reasoning at Flash latency
  • 1,048,576 token context window
  • Natively multimodal: text, image, audio and video
  • Search grounding and Code Execution built into the API
  • Function calling, JSON schema and parallel tool calls
  • Default model in the Gemini app and AI Mode in Search
  • PhD-level reasoning on common benchmarks
  • Available via Vertex AI, AI Studio, Gemini Enterprise, Antigravity and the Gemini app
  • Strong long-video understanding (hour-long clips)
  • Cross-lingual fluency across 100+ languages
  • Best for: high-throughput agentic workflows, real-time multimodal chat, RAG, consumer applications.

Training & Lizenz

Pretrained on a multi-trillion-token mixture of web text, code, books, scientific papers, image-text pairs, audio and video frames. Heavily distilled from larger Gemini 3.1 Pro teacher models. Post-training uses supervised fine-tuning, RLHF and RL against verifiable rewards. Knowledge cutoff in late 2025.

Lizenz: Proprietary commercial license via Google AI Studio, Vertex AI and the Gemini app. Free tier available in the Gemini app and AI Mode in Search.

Sicherheitstests: Evaluated under Google DeepMind's Frontier Safety Framework v2 with internal red teams and external evaluators.

Bekannte Grenzen

  • Below Gemini 3.1 Pro on the hardest reasoning and long-context benchmarks
  • Smaller context window than 3.1 Pro (1M vs 2M)
  • Vision can misread dense tables and handwriting
  • Region availability is rolling out in 2026
  • Audio output not yet supported
03

Preise

Preise in US-Dollar. Abgerechnet wird ĂŒber vorab gekaufte Credits.
Eingabe$ 0.60 / 1 Mio. Token
Ausgabe$ 3.60 / 1 Mio. Token
  • Abgerechnet werden die Token, die jede Anfrage tatsĂ€chlich verbraucht.
  • 1 Credit = $ 0.01

Kostenrechner

Preisrechner

/ Anfr.
/ Anfr.

Gesamt

$ 0.24

24 Credits

Je Anfrage

$ 0.0024 · 0.24 Credits

Jede Anfrage wird auf 0,01 Credits aufgerundet.

04

API

Rufe Gemini 3 Flash mit deinem Railwail-API-SchlĂŒssel auf. Diese Modell-ID gehört in die Anfrage:
curl https://railwail.com/api/v1/chat/completions \
  -H "Authorization: Bearer $RAILWAIL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3-flash",
    "messages": [
      {
        "role": "user",
        "content": "Explain what a vector database is in two sentences."
      }
    ],
    "max_tokens": 1024
  }'
SchlĂŒssel als RAILWAIL_API_KEY setzenAPI-SchlĂŒssel erstellen
05

Spezifikationen

Modell-ID
gemini-3-flash
Entwickler
Google DeepMind
Kategorie
Multimodal
Eingabe
Text, Bild, Audio, Video
Ausgabe
Text
Kontextfenster
1'048'576 Token
Max. Ausgabe
65'536 Token
Abrechnung
Nach Verbrauch (Token bzw. GPU-Zeit)
Lebenszyklus
Auslaufend
Modellgrösse
Undisclosed (sparse MoE, smaller and sparser than Gemini 3.1 Pro)
Lizenz
Proprietary commercial license via Google AI Studio, Vertex AI and the Gemini app. Free tier available in the Gemini app and AI Mode in Search.
Katalogeintrag aktualisiert
23. September 2026

Eingabeparameter

Eingaben und Einstellungen laut Eingabeschema des Modells. Welche davon die API annimmt, zeigt das Beispiel im Abschnitt API.

  • promptPflicht

    User message

    Typ: Text
    Standard: –
    Erlaubte Werte: bis 32'000 Zeichen
  • top_p
    Typ: Zahl
    Standard: 0.95
    Erlaubte Werte: 0 bis 1
  • stream
    Typ: Ja/Nein
    Standard: false
    Erlaubte Werte: –
  • image_url

    Optional image URL to analyze

    Typ: Text
    Standard: –
    Erlaubte Werte: –
  • max_tokens
    Typ: Ganzzahl
    Standard: 4096
    Erlaubte Werte: 1 bis 32'000
  • temperature
    Typ: Zahl
    Standard: 1
    Erlaubte Werte: 0 bis 2
  • system_prompt

    Optional system instruction

    Typ: Text
    Standard: –
    Erlaubte Werte: bis 8'000 Zeichen

Schlagwörter

  • google
  • deepmind
  • balanced
  • multimodal
  • low-latency
  • long-context
  • 1m-context
06

Einsatzgebiete

WofĂŒr es genutzt wird

  • Default consumer multimodal chat
  • High-throughput agentic workflows
  • Real-time RAG pipelines
  • Long-video summarisation and search
  • Production coding assistants
  • Voice and audio reasoning backends
  • AI Mode in Search and Antigravity workflows
07

HĂ€ufige Fragen

Was ist Gemini 3 Flash?

Gemini 3 Flash ist ein Modell von Google DeepMind aus der Kategorie Multimodal. Über Railwail lĂ€sst es sich mit einem API-SchlĂŒssel ĂŒber die Railwail-API aufrufen.

Was kostet Gemini 3 Flash bei Railwail?

Über Railwail kostet Gemini 3 Flash $ 0.60 je 1 Mio. Input-Token und $ 3.60 je 1 Mio. Output-Token. Abgerechnet wird, was jede Anfrage tatsĂ€chlich verbraucht. Bezahlt wird mit vorab gekauften Credits; 1 Credit entspricht $ 0.01.

Wie gross ist das Kontextfenster von Gemini 3 Flash?

Das Kontextfenster von Gemini 3 Flash umfasst 1'048'576 Token. Eine Antwort kann bis zu 65'536 Token lang sein.

Wie schnell ist Gemini 3 Flash?

FĂŒr Gemini 3 Flash gibt es bei Railwail noch zu wenige gemessene LĂ€ufe, um eine Laufzeit anzugeben. Sie hĂ€ngt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Ist Gemini 3 Flash besser als BLIP?

Das hÀngt von der Aufgabe ab. Gemini 3 Flash (Google DeepMind) und BLIP (Salesforce) sind beide Modelle aus der Kategorie Multimodal. Die Vergleichsseite zeigt Preise und Spezifikationen nebeneinander.

Gemini 3 Flash und BLIP vergleichen

Kann Gemini 3 Flash Bilder verarbeiten?

Ja. Gemini 3 Flash nimmt neben Text auch Bilder als Eingabe an.

Wie nutze ich Gemini 3 Flash ĂŒber die API?

Erstelle einen Railwail-API-SchlĂŒssel und sende deine Anfrage mit der Modell-ID gemini-3-flash. Codebeispiele fĂŒr curl, Python und JavaScript stehen im Abschnitt API auf dieser Seite.

08

Vergleichbare Modelle

Alle dieser Kategorie

Gemini 3 Flash ĂŒber die API nutzen

Ein API-SchlĂŒssel fĂŒr alle Modelle auf Railwail. Abgerechnet wird ĂŒber vorab gekaufte Credits, 1 Credit = $ 0.01.