Grok 2 Vision

MultimodalRetiredUnavailable
by xAIModel ID: grok-2-vision

xAI's vision-capable Grok 2 snapshot. Image-in, text-out with strong multilingual instruction following.

Status
Unavailable
Context
32,768 tokens
Max. output
4,096 tokens
Input โ†’ output
Text + Image โ†’ Text
Developer
xAI
Updated
September 24, 2026

Grok 2 Vision is currently unavailable

Currently unavailable: this model has been deactivated.

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives

The provider has retired this model.

01

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    โ‰ˆ US$0.00030/run

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    US$6.00/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    US$3.60/1M in

02

Playground

Try Grok 2 Vision

No input form

Currently unavailable

Currently unavailable: this model has been deactivated.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

03

About Grok 2 Vision

TL;DRAs of September 24, 2026

Grok 2 Vision is a model by xAI in the Multimodal category. Grok 2 Vision is currently not available on Railwail. The context window holds 32,768 tokens, and one response can be up to 4,096 tokens long.

Background

About xAI

Founded 2023 ยท Palo Alto, California, USA

xAI was founded in March 2023 by Elon Musk together with co-founders from DeepMind, OpenAI, Google Research and Microsoft Research, including Igor Babuschkin, Manuel Kroiss, Yuhuai Wu (now back at Google), Christian Szegedy, Jimmy Ba, Toby Pohlen, Ross Nordeen, Kyle Kosic and Greg Yang. The company is closely affiliated with X (formerly Twitter), Tesla and SpaceX. xAI raised $6B Series B in May 2024 followed by $6B Series C in December 2024 at a reported $50B valuation, with backers including Andreessen Horowitz, Sequoia, Fidelity, Kingdom Holding, Lightspeed and Saudi Prince Alwaleed. The flagship Grok model family launched in late 2023 (Grok-1, briefly open-sourced under Apache 2.0), Grok-2 in August 2024 and Grok-3 in February 2025. Grok 2 Vision arrived in October 2024 as xAI's first multimodal model with image input, made available via the X premium feature and the xAI API.

Visit xAI

Architecture

Decoder-only Transformer with vision encoder (multimodal LLM)

Grok 2 Vision (model id grok-2-vision-1212 and successors) is a multimodal large language model that adds an image encoder to xAI's Grok 2 text backbone. The architecture follows the now-standard cross-attention multimodal LLM pattern: a Vision Transformer encodes the input image into visual tokens, which are projected into the LLM token space and concatenated with text tokens before the decoder. xAI has not published a technical paper, but the model card mentions a 'mixture of public web data, X data and licensed sources' with a knowledge cutoff in mid-2024. The model accepts up to 10 images per request, with a maximum image side of around 8,000 pixels, and supports the standard chat/completion API with a 131,072-token context window. Grok 2 Vision is positioned as a competitor to GPT-4o and Claude 3.5 Sonnet for chart understanding, OCR-heavy documents and screenshot reasoning. xAI ships safety filters consistent with their stated 'maximum truth-seeking' posture, which is more permissive on controversial content than OpenAI.

Parameters
Undisclosed
Context
131,072 tokens

Capabilities

  • Image and text input (up to 10 images per request)
  • 131,072-token context window
  • Chart, diagram and screenshot reasoning
  • OCR-heavy document understanding (PDFs as images)
  • Real-time search-grounded responses via X / Grok web tool
  • JSON / structured output and function calling
  • More permissive content policy than OpenAI / Anthropic on controversial topics
  • Best for: chart and screenshot QA, X-integrated agents, code-with-image bug reports

Training & license

Not disclosed. xAI references 'public web data, licensed third-party data and X user posts that have opted in', with a knowledge cutoff in mid-2024.

License: Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.

Safety testing: xAI publishes a 'maximum truth-seeking' policy with intentionally lighter content filtering than peers; bias and jailbreak testing is referenced but no formal red-team report.

Known limitations

  • Closed weights, hosted only
  • No video or audio input (image-only multimodal)
  • Quality on math / vision benchmarks below GPT-4o and Claude 3.5 Sonnet
  • Lighter safety filtering may produce unsafe content
  • Knowledge cutoff mid-2024 without web tool
04

Pricing

Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.

05

API

Call Grok 2 Vision with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
grok-2-vision
Developer
xAI
Category
Multimodal
Input
Text, Image
Output
Text
Context window
32,768 tokens
Max. output
4,096 tokens
Lifecycle
Retired
Model size
Undisclosed
License
Proprietary commercial API and X Premium product. Generated outputs may be used commercially under the xAI terms.
Catalog entry updated
September 24, 2026

Tags

  • xai
  • vision
  • legacy
07

Use cases

What it is used for

  • Chart and screenshot question answering
  • OCR-heavy document understanding
  • X-integrated AI assistants and search agents
  • Code-with-image bug analysis
  • Image-grounded customer support
08

Frequently asked questions

What is Grok 2 Vision?

Grok 2 Vision is a model by xAI in the Multimodal category. It is listed on Railwail but cannot be run at the moment.

How much does Grok 2 Vision cost on Railwail?

Grok 2 Vision cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of Grok 2 Vision?

The context window of Grok 2 Vision holds 32,768 tokens. One response can be up to 4,096 tokens long.

How fast is Grok 2 Vision?

There are not enough measured runs of Grok 2 Vision on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Grok 2 Vision better than BLIP?

That depends on the task. Grok 2 Vision (xAI) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare Grok 2 Vision and BLIP

Can Grok 2 Vision process images?

Yes. Grok 2 Vision accepts images as input in addition to text.

Can I use Grok 2 Vision right now?

Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.