Grok 2 Vision

MultimodalUtgÄttInte tillgÀnglig
av xAIModell-ID: grok-2-vision

xAI's vision-capable Grok 2 snapshot. Image-in, text-out with strong multilingual instruction following.

Status
Inte tillgÀnglig
Kontext
32 768 tokens
Max. utdata
4 096 tokens
Inmatning → utmatning
Text + Bild → Text
Utvecklare
xAI
Uppdaterad
24 september 2026

Grok 2 Vision Àr för nÀrvarande otillgÀnglig

Inte tillgÀnglig för nÀrvarande: denna modell Àr inaktiverad.

Du kan fortfarande lÀsa detaljerna pÄ denna sida. VÀlj ett av de tillgÀngliga alternativen nedan för att köra en jÀmförbar modell direkt.

GĂ„ till alternativ

Leverantören har slutat erbjuda denna modell.

01

JÀmförbara modeller

Alla i denna kategori
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 US$/körning

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 US$/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 US$/1M in

02

Playground

Prova Grok 2 Vision

Ingen inmatningsformulÀr

Inte tillgÀnglig för nÀrvarande

Inte tillgÀnglig för nÀrvarande: denna modell Àr inaktiverad.

Lekplatsen Àr inaktiverad. Du hittar jÀmförbara modeller i samma kategori: Visa alternativ

03

Om Grok 2 Vision

Kort sagtFrÄn och med 24 september 2026

Grok 2 Vision Àr en modell av xAI i kategorin Multimodal. Grok 2 Vision Àr för nÀrvarande inte tillgÀnglig pÄ Railwail. Kontextfönstret innehÄller 32 768 tokens, och ett svar kan vara upp till 4 096 tokens lÄngt.

Bakgrund

Om xAI

Grundat 2023 · Palo Alto, California, USA

xAI was founded in March 2023 by Elon Musk together with co-founders from DeepMind, OpenAI, Google Research and Microsoft Research, including Igor Babuschkin, Manuel Kroiss, Yuhuai Wu (now back at Google), Christian Szegedy, Jimmy Ba, Toby Pohlen, Ross Nordeen, Kyle Kosic and Greg Yang. The company is closely affiliated with X (formerly Twitter), Tesla and SpaceX. xAI raised $6B Series B in May 2024 followed by $6B Series C in December 2024 at a reported $50B valuation, with backers including Andreessen Horowitz, Sequoia, Fidelity, Kingdom Holding, Lightspeed and Saudi Prince Alwaleed. The flagship Grok model family launched in late 2023 (Grok-1, briefly open-sourced under Apache 2.0), Grok-2 in August 2024 and Grok-3 in February 2025. Grok 2 Vision arrived in October 2024 as xAI's first multimodal model with image input, made available via the X premium feature and the xAI API.

Besök xAI

Arkitektur

Decoder-only Transformer with vision encoder (multimodal LLM)

Grok 2 Vision (model id grok-2-vision-1212 and successors) is a multimodal large language model that adds an image encoder to xAI's Grok 2 text backbone. The architecture follows the now-standard cross-attention multimodal LLM pattern: a Vision Transformer encodes the input image into visual tokens, which are projected into the LLM token space and concatenated with text tokens before the decoder. xAI has not published a technical paper, but the model card mentions a 'mixture of public web data, X data and licensed sources' with a knowledge cutoff in mid-2024. The model accepts up to 10 images per request, with a maximum image side of around 8,000 pixels, and supports the standard chat/completion API with a 131,072-token context window. Grok 2 Vision is positioned as a competitor to GPT-4o and Claude 3.5 Sonnet for chart understanding, OCR-heavy documents and screenshot reasoning. xAI ships safety filters consistent with their stated 'maximum truth-seeking' posture, which is more permissive on controversial content than OpenAI.

Parameter
Undisclosed
Kontext
131 072 tokens

Funktioner

  • Image and text input (up to 10 images per request)
  • 131,072-token context window
  • Chart, diagram and screenshot reasoning
  • OCR-heavy document understanding (PDFs as images)
  • Real-time search-grounded responses via X / Grok web tool
  • JSON / structured output and function calling
  • More permissive content policy than OpenAI / Anthropic on controversial topics
  • Best for: chart and screenshot QA, X-integrated agents, code-with-image bug reports

TrÀning & licens

Not disclosed. xAI references 'public web data, licensed third-party data and X user posts that have opted in', with a knowledge cutoff in mid-2024.

Licens: Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.

SĂ€kerhetstestning: xAI publishes a 'maximum truth-seeking' policy with intentionally lighter content filtering than peers; bias and jailbreak testing is referenced but no formal red-team report.

KÀnda begrÀnsningar

  • Closed weights, hosted only
  • No video or audio input (image-only multimodal)
  • Quality on math / vision benchmarks below GPT-4o and Claude 3.5 Sonnet
  • Lighter safety filtering may produce unsafe content
  • Knowledge cutoff mid-2024 without web tool
04

Priser

Inte tillgÀnglig för nÀrvarande: denna modell Àr inaktiverad. Det finns ingen pris för denna modell för nÀrvarande, sÄ den kan inte köras.

05

API

Anropa Grok 2 Vision med din Railwail API-nyckel. AnvÀnd detta modell-ID i begÀran:

Inget verifierat API-exempel

Det offentliga API:et skickar ett annat inmatningsformat Àn vad denna modell behöver. AnvÀnd lekplatsen ovan.

06

Specifikationer

Modell-ID
grok-2-vision
Utvecklare
xAI
Kategori
Multimodal
Inmatning
Text, Bild
Utmatning
Text
Kontextfönster
32 768 tokens
Max. utmatning
4 096 tokens
Livscykel
UtgÄtt
Modellstorlek
Undisclosed
Licens
Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.
KataloginlÀgg uppdaterat
24 september 2026

Taggar

  • xai
  • vision
  • legacy
07

AnvÀndningsfall

Vad det anvÀnds till

  • Chart and screenshot question answering
  • OCR-heavy document understanding
  • X-integrated AI assistants and search agents
  • Code-with-image bug analysis
  • Image-grounded customer support
08

Vanliga frÄgor

Vad Àr Grok 2 Vision?

Grok 2 Vision Àr en modell av xAI i kategorin Multimodal. Den finns i Railwail-katalogen men kan inte köras för nÀrvarande.

Vad kostar Grok 2 Vision pÄ Railwail?

Grok 2 Vision kan inte köras pÄ Railwail för nÀrvarande, sÄ det finns inget aktuellt pris. TillgÀngliga alternativ med priser visas lÀngre ned pÄ denna sida.

Hur stort Àr kontextfönstret för Grok 2 Vision?

Kontextfönstret för Grok 2 Vision innehÄller 32 768 tokens. Ett svar kan vara upp till 4 096 tokens lÄngt.

Hur snabb Àr Grok 2 Vision?

Det finns Ànnu inte tillrÀckligt mÄnga uppmÀtta körningar av Grok 2 Vision pÄ Railwail för att ange en körningstid. Det beror pÄ inmatningen, instÀllningarna och belastningen hos leverantören.

Är Grok 2 Vision bĂ€ttre Ă€n BLIP?

Det beror pÄ uppgiften. Grok 2 Vision (xAI) och BLIP (Salesforce) Àr bÄda modeller i kategorin Multimodal. JÀmförelsesidan visar deras priser och specifikationer sida vid sida.

JÀmför Grok 2 Vision och BLIP

Kan Grok 2 Vision bearbeta bilder?

Ja. Grok 2 Vision accepterar bilder som inmatning utöver text.

Kan jag anvÀnda Grok 2 Vision just nu?

Inte tillgÀnglig för nÀrvarande: denna modell Àr inaktiverad. Sidan förblir online; tillgÀngliga alternativ frÄn samma kategori visas lÀngre ned.

Alla modeller via ett API

En API-nyckel för alla modeller pÄ Railwail. AnvÀndningen debiteras frÄn förbetald kredit, 1 kredit = 0,01 US$.