Grok 2 Vision

MultimodalRetrasIndisponibil
de xAIID model: grok-2-vision

xAI's vision-capable Grok 2 snapshot. Image-in, text-out with strong multilingual instruction following.

Status
Indisponibil
Context
32.768 tokeni
Ieșire max.
4.096 tokeni
Intrare → ieșire
Text + Imagine → Text
Dezvoltator
xAI
Actualizat
24 septembrie 2026

Grok 2 Vision nu este disponibil în acest moment

Indisponibil în prezent: acest model a fost dezactivat.

Poți citi în continuare detaliile pe această pagină. Alege una dintre alternativele disponibile de mai jos pentru a rula imediat un model comparabil.

Mergi la alternative

Furnizorul a retras acest model.

01
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 USD/rulare

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 USD/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 USD/1M in

02

Playground

Încearcă Grok 2 Vision

Fără formular de intrare

Indisponibil în prezent

Indisponibil în prezent: acest model a fost dezactivat.

Playground-ul este dezactivat. Modele comparabile găsești în aceeași categorie: Vezi alternativele

03

Despre Grok 2 Vision

Pe scurtDin 24 septembrie 2026

Grok 2 Vision este un model de xAI din categoria Multimodal. Grok 2 Vision nu este disponibil în prezent pe Railwail. Fereastra de context conține 32.768 token-uri, iar un răspuns poate fi lung de până la 4.096 token-uri.

Fundal

Despre xAI

Fondat 2023 · Palo Alto, California, USA

xAI was founded in March 2023 by Elon Musk together with co-founders from DeepMind, OpenAI, Google Research and Microsoft Research, including Igor Babuschkin, Manuel Kroiss, Yuhuai Wu (now back at Google), Christian Szegedy, Jimmy Ba, Toby Pohlen, Ross Nordeen, Kyle Kosic and Greg Yang. The company is closely affiliated with X (formerly Twitter), Tesla and SpaceX. xAI raised $6B Series B in May 2024 followed by $6B Series C in December 2024 at a reported $50B valuation, with backers including Andreessen Horowitz, Sequoia, Fidelity, Kingdom Holding, Lightspeed and Saudi Prince Alwaleed. The flagship Grok model family launched in late 2023 (Grok-1, briefly open-sourced under Apache 2.0), Grok-2 in August 2024 and Grok-3 in February 2025. Grok 2 Vision arrived in October 2024 as xAI's first multimodal model with image input, made available via the X premium feature and the xAI API.

Vizitează xAI

Arhitectură

Decoder-only Transformer with vision encoder (multimodal LLM)

Grok 2 Vision (model id grok-2-vision-1212 and successors) is a multimodal large language model that adds an image encoder to xAI's Grok 2 text backbone. The architecture follows the now-standard cross-attention multimodal LLM pattern: a Vision Transformer encodes the input image into visual tokens, which are projected into the LLM token space and concatenated with text tokens before the decoder. xAI has not published a technical paper, but the model card mentions a 'mixture of public web data, X data and licensed sources' with a knowledge cutoff in mid-2024. The model accepts up to 10 images per request, with a maximum image side of around 8,000 pixels, and supports the standard chat/completion API with a 131,072-token context window. Grok 2 Vision is positioned as a competitor to GPT-4o and Claude 3.5 Sonnet for chart understanding, OCR-heavy documents and screenshot reasoning. xAI ships safety filters consistent with their stated 'maximum truth-seeking' posture, which is more permissive on controversial content than OpenAI.

Parametri
Undisclosed
Context
131.072 tokeni

Capabilități

  • Image and text input (up to 10 images per request)
  • 131,072-token context window
  • Chart, diagram and screenshot reasoning
  • OCR-heavy document understanding (PDFs as images)
  • Real-time search-grounded responses via X / Grok web tool
  • JSON / structured output and function calling
  • More permissive content policy than OpenAI / Anthropic on controversial topics
  • Best for: chart and screenshot QA, X-integrated agents, code-with-image bug reports

Antrenament & licență

Not disclosed. xAI references 'public web data, licensed third-party data and X user posts that have opted in', with a knowledge cutoff in mid-2024.

Licență: Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.

Teste de siguranță: xAI publishes a 'maximum truth-seeking' policy with intentionally lighter content filtering than peers; bias and jailbreak testing is referenced but no formal red-team report.

Limitări cunoscute

  • Closed weights, hosted only
  • No video or audio input (image-only multimodal)
  • Quality on math / vision benchmarks below GPT-4o and Claude 3.5 Sonnet
  • Lighter safety filtering may produce unsafe content
  • Knowledge cutoff mid-2024 without web tool
04

Prețuri

Indisponibil în prezent: acest model a fost dezactivat. Nu există preț pentru acest model în acest moment, deci nu poate fi executat.

05

API

Apelează Grok 2 Vision cu cheia ta API Railwail. Folosește acest ID de model în cerere:

Niciun exemplu API verificat

API-ul public transmite un format de intrare diferit de ceea ce are nevoie acest model. Folosește playground-ul de mai sus.

06

Specificații

ID model
grok-2-vision
Dezvoltator
xAI
Categorie
Multimodal
Intrare
Text, Imagine
Ieșire
Text
Fereastră de context
32.768 tokeni
Ieșire max.
4.096 tokeni
Ciclu de viață
Retras
Dimensiune model
Undisclosed
Licență
Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.
Intrare catalog actualizată
24 septembrie 2026

Etichete

  • xai
  • vision
  • legacy
07

Cazuri de utilizare

Pentru ce se folosește

  • Chart and screenshot question answering
  • OCR-heavy document understanding
  • X-integrated AI assistants and search agents
  • Code-with-image bug analysis
  • Image-grounded customer support
08

Întrebări frecvente

Ce este Grok 2 Vision?

Grok 2 Vision este un model de xAI din categoria Multimodal. Este listat pe Railwail, dar nu poate fi rulat în acest moment.

Cât costă Grok 2 Vision pe Railwail?

Grok 2 Vision nu poate fi rulat pe Railwail în acest moment, deci nu există preț curent. Alternativele disponibile cu prețuri sunt listate mai jos pe această pagină.

Care este fereastra de context a Grok 2 Vision?

Fereastra de context a Grok 2 Vision conține 32.768 token-uri. Un răspuns poate fi lung de până la 4.096 token-uri.

Cât de rapid este Grok 2 Vision?

Nu sunt suficiente rulări măsurate ale Grok 2 Vision pe Railwail încă pentru a indica un timp de rulare. Depinde de intrare, de setări și de sarcina la furnizor.

Este Grok 2 Vision mai bun decât BLIP?

Depinde de sarcină. Grok 2 Vision (xAI) și BLIP (Salesforce) sunt ambele modele din categoria Multimodal. Pagina de comparație arată prețurile și specificațiile lor una lângă alta.

Compară Grok 2 Vision și BLIP

Poate Grok 2 Vision procesa imagini?

Da. Grok 2 Vision acceptă imagini ca intrare, pe lângă text.

Pot folosi Grok 2 Vision chiar acum?

Indisponibil în prezent: acest model a fost dezactivat. Pagina rămâne online; alternativele disponibile din aceeași categorie sunt listate mai jos.

Toate modelele printr-o singură API

O cheie API pentru fiecare model pe Railwail. Utilizarea se percepe din credite prepay, 1 credit = 0,01 USD.