Grok 2 Vision

MultimodalneWycofanyNiedostępne
od xAIID modelu: grok-2-vision

xAI's vision-capable Grok 2 snapshot. Image-in, text-out with strong multilingual instruction following.

Status
Niedostępne
Kontekst
32 768 tokenów
Maks. wyjście
4096 tokenów
Wejście → Wyjście
Tekst + Obraz → Tekst
Deweloper
xAI
Zaktualizowano
24 września 2026

Grok 2 Vision jest obecnie niedostępny

Obecnie niedostępne: ten model został dezaktywowany.

Możesz nadal przeczytać szczegóły na tej stronie. Wybierz jedną z dostępnych alternatyw poniżej, aby od razu uruchomić porównywalny model.

Przejdź do alternatyw

Dostawca wycofał ten model.

01

Porównywalne modele

Wszystkie w tej kategorii
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 USD/uruchomienie

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 USD/1M wej.

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 USD/1M wej.

02

Playground

Spróbuj Grok 2 Vision

Brak formularza wejściowego

Niedostępny

Obecnie niedostępne: ten model został dezaktywowany.

Plac zabaw jest wyłączony. Porównywalne modele znajdziesz w tej samej kategorii: Przeglądaj alternatywy

03

O Grok 2 Vision

Krótko mówiącStan na 24 września 2026

Grok 2 Vision to model opracowany przez xAI w kategorii Multimodalne. Grok 2 Vision nie jest obecnie dostępny w serwisie Railwail. Okno kontekstu zawiera 32 768 tokenów, a jedna odpowiedź może mieć do 4096 tokenów.

Tło

O xAI

Założona 2023 · Palo Alto, California, USA

xAI was founded in March 2023 by Elon Musk together with co-founders from DeepMind, OpenAI, Google Research and Microsoft Research, including Igor Babuschkin, Manuel Kroiss, Yuhuai Wu (now back at Google), Christian Szegedy, Jimmy Ba, Toby Pohlen, Ross Nordeen, Kyle Kosic and Greg Yang. The company is closely affiliated with X (formerly Twitter), Tesla and SpaceX. xAI raised $6B Series B in May 2024 followed by $6B Series C in December 2024 at a reported $50B valuation, with backers including Andreessen Horowitz, Sequoia, Fidelity, Kingdom Holding, Lightspeed and Saudi Prince Alwaleed. The flagship Grok model family launched in late 2023 (Grok-1, briefly open-sourced under Apache 2.0), Grok-2 in August 2024 and Grok-3 in February 2025. Grok 2 Vision arrived in October 2024 as xAI's first multimodal model with image input, made available via the X premium feature and the xAI API.

Odwiedź xAI

Architektura

Decoder-only Transformer with vision encoder (multimodal LLM)

Grok 2 Vision (model id grok-2-vision-1212 and successors) is a multimodal large language model that adds an image encoder to xAI's Grok 2 text backbone. The architecture follows the now-standard cross-attention multimodal LLM pattern: a Vision Transformer encodes the input image into visual tokens, which are projected into the LLM token space and concatenated with text tokens before the decoder. xAI has not published a technical paper, but the model card mentions a 'mixture of public web data, X data and licensed sources' with a knowledge cutoff in mid-2024. The model accepts up to 10 images per request, with a maximum image side of around 8,000 pixels, and supports the standard chat/completion API with a 131,072-token context window. Grok 2 Vision is positioned as a competitor to GPT-4o and Claude 3.5 Sonnet for chart understanding, OCR-heavy documents and screenshot reasoning. xAI ships safety filters consistent with their stated 'maximum truth-seeking' posture, which is more permissive on controversial content than OpenAI.

Parametry
Undisclosed
Kontekst
131 072 tokenów

Możliwości

  • Image and text input (up to 10 images per request)
  • 131,072-token context window
  • Chart, diagram and screenshot reasoning
  • OCR-heavy document understanding (PDFs as images)
  • Real-time search-grounded responses via X / Grok web tool
  • JSON / structured output and function calling
  • More permissive content policy than OpenAI / Anthropic on controversial topics
  • Best for: chart and screenshot QA, X-integrated agents, code-with-image bug reports

Trening i licencja

Not disclosed. xAI references 'public web data, licensed third-party data and X user posts that have opted in', with a knowledge cutoff in mid-2024.

Licencja: Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.

Testy bezpieczeństwa: xAI publishes a 'maximum truth-seeking' policy with intentionally lighter content filtering than peers; bias and jailbreak testing is referenced but no formal red-team report.

Znane ograniczenia

  • Closed weights, hosted only
  • No video or audio input (image-only multimodal)
  • Quality on math / vision benchmarks below GPT-4o and Claude 3.5 Sonnet
  • Lighter safety filtering may produce unsafe content
  • Knowledge cutoff mid-2024 without web tool
04

Ceny

Obecnie niedostępne: ten model został dezaktywowany. Dla tego modelu nie ma ceny w tej chwili, dlatego nie można go uruchomić.

05

API

Wywołaj Grok 2 Vision za pomocą klucza API Railwail. Użyj tego ID modelu w żądaniu:

Brak zweryfikowanego przykładu API

Publiczny API przekazuje inny format wejściowy niż wymaga ten model. Użyj placu zabaw powyżej.

06

Specyfikacje

ID modelu
grok-2-vision
Deweloper
xAI
Kategoria
Multimodalne
Wejście
Tekst, Obraz
Wyjście
Tekst
Okno kontekstu
32 768 tokenów
Maks. wyjście
4096 tokenów
Cykl życia
Wycofany
Rozmiar modelu
Undisclosed
Licencja
Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.
Wpis w katalogu zaktualizowany
24 września 2026

Tagi

  • xai
  • vision
  • legacy
07

Przypadki użycia

Do czego się go używa

  • Chart and screenshot question answering
  • OCR-heavy document understanding
  • X-integrated AI assistants and search agents
  • Code-with-image bug analysis
  • Image-grounded customer support
08

Często zadawane pytania

Co to jest Grok 2 Vision?

Grok 2 Vision to model opracowany przez xAI w kategorii Multimodalne. Jest wymieniony w katalogu Railwail, ale nie może być uruchomiony w tej chwili.

Ile kosztuje Grok 2 Vision w serwisie Railwail?

Grok 2 Vision nie może być uruchomiony w serwisie Railwail w tej chwili, dlatego nie ma aktualnej ceny. Dostępne alternatywy z cenami są wymienione poniżej na tej stronie.

Jakie jest okno kontekstu Grok 2 Vision?

Okno kontekstu Grok 2 Vision zawiera 32 768 tokenów. Jedna odpowiedź może mieć do 4096 tokenów.

Jak szybki jest Grok 2 Vision?

Dla Grok 2 Vision jest jeszcze zbyt mało zmierzonych przebiegów w serwisie Railwail, aby podać czas przebiegu. Zależy to od wejścia, ustawień i obciążenia u dostawcy.

Czy Grok 2 Vision jest lepszy niż BLIP?

To zależy od zadania. Grok 2 Vision (xAI) i BLIP (Salesforce) to oba modele z kategorii Multimodalne. Strona porównania pokazuje ich ceny i specyfikacje obok siebie.

Porównaj Grok 2 Vision i BLIP

Czy Grok 2 Vision może przetwarzać obrazy?

Tak. Grok 2 Vision akceptuje obrazy jako wejście oprócz tekstu.

Czy mogę używać Grok 2 Vision teraz?

Obecnie niedostępne: ten model został dezaktywowany. Strona pozostaje online; dostępne alternatywy z tej samej kategorii są wymienione poniżej.

Wszystkie modele przez jedno API

Jeden klucz API dla każdego modelu na Railwail. Opłaty pobierane są z przedpłaconych kredytów, 1 kredyt = 0,01 USD.