Grok 2 Vision

MultimodalEingestelltNicht verfügbar
von xAIModell-ID: grok-2-vision

xAI's vision-capable Grok 2 snapshot. Image-in, text-out with strong multilingual instruction following.

Status
Nicht verfügbar
Kontext
32.768 Token
Max. Ausgabe
4.096 Token
Eingabe → Ausgabe
Text + Bild → Text
Entwickler
xAI
Aktualisiert
24. September 2026

Grok 2 Vision ist derzeit nicht verfügbar

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert.

Die Angaben auf dieser Seite kannst du weiter nachlesen. Mit einer der verfügbaren Alternativen unten kannst du sofort ein vergleichbares Modell nutzen.

Zu den Alternativen

Der Anbieter hat dieses Modell eingestellt.

01

Vergleichbare Modelle

Alle dieser Kategorie
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ca. 0,00030 $/Lauf

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 $/1 Mio. In

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 $/1 Mio. In

02

Playground

Grok 2 Vision ausprobieren

Keine Eingabemaske

Derzeit nicht verfügbar

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert.

Der Playground ist deaktiviert. Vergleichbare Modelle findest du in derselben Kategorie: Alternativen ansehen

03

Über Grok 2 Vision

Kurz gesagtStand: 24. September 2026

Grok 2 Vision ist ein Modell von xAI aus der Kategorie Multimodal. Über Railwail ist Grok 2 Vision derzeit nicht verfügbar. Das Kontextfenster umfasst 32.768 Token, eine Antwort bis zu 4.096 Token.

Hintergrund

Über xAI

Gegründet 2023 · Palo Alto, California, USA

xAI was founded in March 2023 by Elon Musk together with co-founders from DeepMind, OpenAI, Google Research and Microsoft Research, including Igor Babuschkin, Manuel Kroiss, Yuhuai Wu (now back at Google), Christian Szegedy, Jimmy Ba, Toby Pohlen, Ross Nordeen, Kyle Kosic and Greg Yang. The company is closely affiliated with X (formerly Twitter), Tesla and SpaceX. xAI raised $6B Series B in May 2024 followed by $6B Series C in December 2024 at a reported $50B valuation, with backers including Andreessen Horowitz, Sequoia, Fidelity, Kingdom Holding, Lightspeed and Saudi Prince Alwaleed. The flagship Grok model family launched in late 2023 (Grok-1, briefly open-sourced under Apache 2.0), Grok-2 in August 2024 and Grok-3 in February 2025. Grok 2 Vision arrived in October 2024 as xAI's first multimodal model with image input, made available via the X premium feature and the xAI API.

xAI besuchen

Architektur

Decoder-only Transformer with vision encoder (multimodal LLM)

Grok 2 Vision (model id grok-2-vision-1212 and successors) is a multimodal large language model that adds an image encoder to xAI's Grok 2 text backbone. The architecture follows the now-standard cross-attention multimodal LLM pattern: a Vision Transformer encodes the input image into visual tokens, which are projected into the LLM token space and concatenated with text tokens before the decoder. xAI has not published a technical paper, but the model card mentions a 'mixture of public web data, X data and licensed sources' with a knowledge cutoff in mid-2024. The model accepts up to 10 images per request, with a maximum image side of around 8,000 pixels, and supports the standard chat/completion API with a 131,072-token context window. Grok 2 Vision is positioned as a competitor to GPT-4o and Claude 3.5 Sonnet for chart understanding, OCR-heavy documents and screenshot reasoning. xAI ships safety filters consistent with their stated 'maximum truth-seeking' posture, which is more permissive on controversial content than OpenAI.

Parameter
Undisclosed
Kontext
131.072 Token

Fähigkeiten

  • Image and text input (up to 10 images per request)
  • 131,072-token context window
  • Chart, diagram and screenshot reasoning
  • OCR-heavy document understanding (PDFs as images)
  • Real-time search-grounded responses via X / Grok web tool
  • JSON / structured output and function calling
  • More permissive content policy than OpenAI / Anthropic on controversial topics
  • Best for: chart and screenshot QA, X-integrated agents, code-with-image bug reports

Training & Lizenz

Not disclosed. xAI references 'public web data, licensed third-party data and X user posts that have opted in', with a knowledge cutoff in mid-2024.

Lizenz: Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.

Sicherheitstests: xAI publishes a 'maximum truth-seeking' policy with intentionally lighter content filtering than peers; bias and jailbreak testing is referenced but no formal red-team report.

Bekannte Grenzen

  • Closed weights, hosted only
  • No video or audio input (image-only multimodal)
  • Quality on math / vision benchmarks below GPT-4o and Claude 3.5 Sonnet
  • Lighter safety filtering may produce unsafe content
  • Knowledge cutoff mid-2024 without web tool
04

Preise

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert. Für dieses Modell gibt es derzeit keinen Preis, deshalb lässt es sich nicht ausführen.

05

API

Rufe Grok 2 Vision mit deinem Railwail-API-Schlüssel auf. Diese Modell-ID gehört in die Anfrage:

Kein geprüftes API-Beispiel

Die öffentliche API übergibt ein anderes Eingabeformat, als dieses Modell braucht. Nutze den Playground oben.

06

Spezifikationen

Modell-ID
grok-2-vision
Entwickler
xAI
Kategorie
Multimodal
Eingabe
Text, Bild
Ausgabe
Text
Kontextfenster
32.768 Token
Max. Ausgabe
4.096 Token
Lebenszyklus
Eingestellt
Modellgröße
Undisclosed
Lizenz
Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.
Katalogeintrag aktualisiert
24. September 2026

Schlagwörter

  • xai
  • vision
  • legacy
07

Einsatzgebiete

Wofür es genutzt wird

  • Chart and screenshot question answering
  • OCR-heavy document understanding
  • X-integrated AI assistants and search agents
  • Code-with-image bug analysis
  • Image-grounded customer support
08

Häufige Fragen

Was ist Grok 2 Vision?

Grok 2 Vision ist ein Modell von xAI aus der Kategorie Multimodal. Es steht im Railwail-Katalog, lässt sich derzeit aber nicht ausführen.

Was kostet Grok 2 Vision bei Railwail?

Grok 2 Vision lässt sich über Railwail derzeit nicht ausführen, deshalb gibt es keinen aktuellen Preis. Verfügbare Alternativen mit Preisen stehen weiter unten auf dieser Seite.

Wie groß ist das Kontextfenster von Grok 2 Vision?

Das Kontextfenster von Grok 2 Vision umfasst 32.768 Token. Eine Antwort kann bis zu 4.096 Token lang sein.

Wie schnell ist Grok 2 Vision?

Für Grok 2 Vision gibt es bei Railwail noch zu wenige gemessene Läufe, um eine Laufzeit anzugeben. Sie hängt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Ist Grok 2 Vision besser als BLIP?

Das hängt von der Aufgabe ab. Grok 2 Vision (xAI) und BLIP (Salesforce) sind beide Modelle aus der Kategorie Multimodal. Die Vergleichsseite zeigt Preise und Spezifikationen nebeneinander.

Grok 2 Vision und BLIP vergleichen

Kann Grok 2 Vision Bilder verarbeiten?

Ja. Grok 2 Vision nimmt neben Text auch Bilder als Eingabe an.

Kann ich Grok 2 Vision gerade nutzen?

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert. Die Seite bleibt online; verfügbare Alternativen aus derselben Kategorie stehen weiter unten.

Alle Modelle über eine API

Ein API-Schlüssel für alle Modelle auf Railwail. Abgerechnet wird über vorab gekaufte Credits, 1 Credit = 0,01 $.