Grok 2 Vision

MultimodaalStopgezetNiet beschikbaar
van xAIModel-ID: grok-2-vision

xAI's vision-capable Grok 2 snapshot. Image-in, text-out with strong multilingual instruction following.

Status
Niet beschikbaar
Context
32.768 tokens
Max. uitvoer
4.096 tokens
Invoer → Uitvoer
Tekst + Afbeelding → Tekst
Ontwikkelaar
xAI
Bijgewerkt
24 september 2026

Grok 2 Vision is momenteel niet beschikbaar

Momenteel niet beschikbaar: dit model is gedeactiveerd.

Je kunt de details op deze pagina nog steeds lezen. Kies een van de beschikbare alternatieven hieronder om direct een vergelijkbaar model uit te voeren.

Naar alternatieven

De provider heeft dit model stopgezet.

01

Vergelijkbare modellen

Alle in deze categorie
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ US$ 0,00030/uitvoering

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    US$ 6,00/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    US$ 3,60/1M in

02

Playground

Grok 2 Vision proberen

Geen invoerformulier

Momenteel niet beschikbaar

Momenteel niet beschikbaar: dit model is gedeactiveerd.

De playground is uitgeschakeld. Vergelijkbare modellen vind je in dezelfde categorie: Alternatieven bekijken

03

Over Grok 2 Vision

SamengevatPer 24 september 2026

Grok 2 Vision is een model van xAI in de categorie Multimodaal. Grok 2 Vision is momenteel niet beschikbaar op Railwail. Het contextvenster bevat 32.768 tokens, en een antwoord kan tot 4.096 tokens lang zijn.

Achtergrond

Over xAI

Opgericht 2023 · Palo Alto, California, USA

xAI was founded in March 2023 by Elon Musk together with co-founders from DeepMind, OpenAI, Google Research and Microsoft Research, including Igor Babuschkin, Manuel Kroiss, Yuhuai Wu (now back at Google), Christian Szegedy, Jimmy Ba, Toby Pohlen, Ross Nordeen, Kyle Kosic and Greg Yang. The company is closely affiliated with X (formerly Twitter), Tesla and SpaceX. xAI raised $6B Series B in May 2024 followed by $6B Series C in December 2024 at a reported $50B valuation, with backers including Andreessen Horowitz, Sequoia, Fidelity, Kingdom Holding, Lightspeed and Saudi Prince Alwaleed. The flagship Grok model family launched in late 2023 (Grok-1, briefly open-sourced under Apache 2.0), Grok-2 in August 2024 and Grok-3 in February 2025. Grok 2 Vision arrived in October 2024 as xAI's first multimodal model with image input, made available via the X premium feature and the xAI API.

xAI bezoeken

Architectuur

Decoder-only Transformer with vision encoder (multimodal LLM)

Grok 2 Vision (model id grok-2-vision-1212 and successors) is a multimodal large language model that adds an image encoder to xAI's Grok 2 text backbone. The architecture follows the now-standard cross-attention multimodal LLM pattern: a Vision Transformer encodes the input image into visual tokens, which are projected into the LLM token space and concatenated with text tokens before the decoder. xAI has not published a technical paper, but the model card mentions a 'mixture of public web data, X data and licensed sources' with a knowledge cutoff in mid-2024. The model accepts up to 10 images per request, with a maximum image side of around 8,000 pixels, and supports the standard chat/completion API with a 131,072-token context window. Grok 2 Vision is positioned as a competitor to GPT-4o and Claude 3.5 Sonnet for chart understanding, OCR-heavy documents and screenshot reasoning. xAI ships safety filters consistent with their stated 'maximum truth-seeking' posture, which is more permissive on controversial content than OpenAI.

Parameters
Undisclosed
Context
131.072 tokens

Mogelijkheden

  • Image and text input (up to 10 images per request)
  • 131,072-token context window
  • Chart, diagram and screenshot reasoning
  • OCR-heavy document understanding (PDFs as images)
  • Real-time search-grounded responses via X / Grok web tool
  • JSON / structured output and function calling
  • More permissive content policy than OpenAI / Anthropic on controversial topics
  • Best for: chart and screenshot QA, X-integrated agents, code-with-image bug reports

Training & licentie

Not disclosed. xAI references 'public web data, licensed third-party data and X user posts that have opted in', with a knowledge cutoff in mid-2024.

Licentie: Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.

Veiligheidstests: xAI publishes a 'maximum truth-seeking' policy with intentionally lighter content filtering than peers; bias and jailbreak testing is referenced but no formal red-team report.

Bekende beperkingen

  • Closed weights, hosted only
  • No video or audio input (image-only multimodal)
  • Quality on math / vision benchmarks below GPT-4o and Claude 3.5 Sonnet
  • Lighter safety filtering may produce unsafe content
  • Knowledge cutoff mid-2024 without web tool
04

Prijzen

Momenteel niet beschikbaar: dit model is gedeactiveerd. Er is momenteel geen prijs voor dit model, dus het kan niet worden uitgevoerd.

05

API

Roep Grok 2 Vision aan met je Railwail API-sleutel. Gebruik deze model-ID in het verzoek:

Geen geverifieerd API-voorbeeld

De openbare API geeft een ander invoerformaat door dan dit model nodig heeft. Gebruik de playground hierboven.

06

Specificaties

Model-ID
grok-2-vision
Ontwikkelaar
xAI
Categorie
Multimodaal
Invoer
Tekst, Afbeelding
Uitvoer
Tekst
Contextvenster
32.768 tokens
Max. uitvoer
4.096 tokens
Levenscyclus
Stopgezet
Modelgrootte
Undisclosed
Licentie
Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.
Catalogusitem bijgewerkt
24 september 2026

Tags

  • xai
  • vision
  • legacy
07

Gebruiksscenario's

Waarvoor het wordt gebruikt

  • Chart and screenshot question answering
  • OCR-heavy document understanding
  • X-integrated AI assistants and search agents
  • Code-with-image bug analysis
  • Image-grounded customer support
08

Veelgestelde vragen

Wat is Grok 2 Vision?

Grok 2 Vision is een model van xAI in de categorie Multimodaal. Het staat in de Railwail-catalogus, maar kan momenteel niet worden uitgevoerd.

Hoeveel kost Grok 2 Vision op Railwail?

Grok 2 Vision kan momenteel niet op Railwail worden uitgevoerd, dus er is geen huidige prijs. Beschikbare alternatieven met prijzen staan verderop op deze pagina.

Wat is het contextvenster van Grok 2 Vision?

Het contextvenster van Grok 2 Vision bevat 32.768 tokens. Een antwoord kan tot 4.096 tokens lang zijn.

Hoe snel is Grok 2 Vision?

Er zijn nog niet genoeg gemeten runs van Grok 2 Vision op Railwail om een uitvoeringstijd op te geven. Dit hangt af van de invoer, de instellingen en de belasting bij de provider.

Is Grok 2 Vision beter dan BLIP?

Dat hangt van de taak af. Grok 2 Vision (xAI) en BLIP (Salesforce) zijn beide modellen in de categorie Multimodaal. De vergelijkingspagina toont hun prijzen en specificaties naast elkaar.

Grok 2 Vision en BLIP vergelijken

Kan Grok 2 Vision afbeeldingen verwerken?

Ja. Grok 2 Vision accepteert afbeeldingen als invoer naast tekst.

Kan ik Grok 2 Vision nu gebruiken?

Momenteel niet beschikbaar: dit model is gedeactiveerd. De pagina blijft online; beschikbare alternatieven uit dezelfde categorie staan verderop.

Alle modellen via één API

Één API-sleutel voor elk model op Railwail. Gebruik wordt afgerekend via vooraf gekochte credits, 1 credit = US$ 0,01.