Llama 3.2 90B Vision (multimodal)

MultimodalUnavailable
by MetaModel ID: llama-3-2-90b-vision-mm

Meta's flagship vision-language model. 90B parameters, image understanding + chat, strong VQA performance.

Status
Unavailable
Context
131,072 tokens
Max. output
8,192 tokens
Input โ†’ output
Text + Image โ†’ Text
Developer
Meta
Updated
June 25, 2026

Llama 3.2 90B Vision (multimodal) is currently unavailable

Currently unavailable: this model has been deactivated.

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    โ‰ˆ $0.00030/run

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    $6.00/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    $3.60/1M in

02

Playground

Try Llama 3.2 90B Vision (multimodal)

Chat

Currently unavailable

Currently unavailable: this model has been deactivated.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try Llama 3.2 90B Vision (multimodal)

Send a message. The answer arrives in full once the model is done (no streaming).

System prompt
Max. answer length (tokens)

This run

No price โ€“ currently unavailable.

New here?

5 free credits ($0.05) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About Llama 3.2 90B Vision (multimodal)

TL;DRAs of June 25, 2026

Llama 3.2 90B Vision (multimodal) is a model by Meta in the Multimodal category. Llama 3.2 90B Vision (multimodal) is currently not available on Railwail. The context window holds 131,072 tokens, and one response can be up to 8,192 tokens long.

Background

About Meta AI (FAIR)

Founded 2013 ยท Menlo Park, California, USA

Meta AI is the research arm of Meta Platforms, established in 2013 as Facebook AI Research (FAIR) by Yann LeCun. FAIR has open-sourced many foundational models including PyTorch, RoBERTa, DETR, SAM and the LLaMA family. LLaMA 1 was released in February 2023, LLaMA 2 in July 2023, LLaMA 3 in April 2024 and LLaMA 3.1 (405B) in July 2024. Llama 3.2 launched in September 2024 at Meta Connect, introducing the first multimodal models in the LLaMA family (vision-enabled 11B and 90B) together with tiny on-device text-only siblings (1B, 3B). All Llama 3.2 vision weights are released under the Llama 3 Community Licence and are widely used by enterprise customers via Meta's partner ecosystem (Hugging Face, AWS Bedrock, Azure AI Studio, Google Vertex, Together AI, Groq, Fireworks).

Visit Meta AI (FAIR)

Architecture

Decoder-only Transformer with cross-attended vision encoder

Llama 3.2 90B Vision combines the 70B-parameter Llama 3.1 text backbone (extended to 90B with vision components) and a Vision Transformer image encoder integrated via cross-attention adapter layers, similar in spirit to Flamingo but reusing the LLaMA architecture. The vision tower processes each image to a sequence of visual tokens which are injected into specific cross-attention layers of the LLM decoder while the original text-only weights remain frozen during the multimodal training stage, preserving text-only performance. Pretraining used 6B image-text pairs followed by multi-stage supervised fine-tuning and Direct Preference Optimisation (DPO) on a curated set of image instructions, math and chart data. The model supports a 128K context window and accepts up to 1120x1120 image inputs natively (with tiling for larger images). It does not support video or audio. Llama 3.2 90B Vision is released under the Llama 3 Community Licence (free for commercial use under 700M MAU).

Parameters
90B
Context
128,000 tokens

Capabilities

  • Open-weights 90B vision-language model under Llama 3 Community Licence
  • 128K token context window
  • Image input up to 1120x1120 with tiling for larger images
  • Chart, diagram, OCR and document understanding
  • Strong on MMMU, MathVista, ChartQA and DocVQA among open-weights models
  • Multilingual: English, German, French, Italian, Portuguese, Spanish, Hindi, Thai
  • Tool use and JSON output via Llama 3.1 alignment recipe
  • Best for: open-weights multimodal apps, on-premise document AI, indie research

Training & license

Pretrained on 6B image-text pairs from public web and licensed sources; supervised fine-tuning and DPO on curated multimodal instruction data. Text knowledge inherited from Llama 3.1 (15T tokens).

License: Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.

Safety testing: Meta publishes a comprehensive model card with red-team findings on CBRN, child-safety and hate-speech vectors, plus Llama Guard 3 and Prompt Guard 2 companion models for production safety.

Known limitations

  • No video or audio input
  • Latency and cost dominated by 90B params; requires multi-GPU serving
  • Licence restricts the largest hyperscaler use cases
  • Vision quality below GPT-4o and Claude 3.5 Sonnet on hardest charts
  • English-centric in vision domain
04

Pricing

Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.

05

API

Call Llama 3.2 90B Vision (multimodal) with your Railwail API key. Use this model ID in the request:
llama-3-2-90b-vision-mmAPI documentationGet an API key

Currently unavailable

The model has no verified price or is deactivated; API calls are refused.

06

Specifications

Model ID
llama-3-2-90b-vision-mm
Developer
Meta
Category
Multimodal
Input
Text, Image
Output
Text
Context window
131,072 tokens
Max. output
8,192 tokens
Model size
90B
License
Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.
Catalog entry updated
June 25, 2026

Tags

  • meta
  • llama
  • multimodal
  • vision
  • open-weights
07

Use cases

What it is used for

  • Open-weights document AI and OCR pipelines
  • On-premise vision-language assistants
  • Chart and diagram understanding for analytics
  • Compliance and regulated-industry multimodal apps
  • Research baselines for vision-language models
08

Frequently asked questions

What is Llama 3.2 90B Vision (multimodal)?

Llama 3.2 90B Vision (multimodal) is a model by Meta in the Multimodal category. It is listed on Railwail but cannot be run at the moment.

How much does Llama 3.2 90B Vision (multimodal) cost on Railwail?

Llama 3.2 90B Vision (multimodal) cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of Llama 3.2 90B Vision (multimodal)?

The context window of Llama 3.2 90B Vision (multimodal) holds 131,072 tokens. One response can be up to 8,192 tokens long.

How fast is Llama 3.2 90B Vision (multimodal)?

There are not enough measured runs of Llama 3.2 90B Vision (multimodal) on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Llama 3.2 90B Vision (multimodal) better than BLIP?

That depends on the task. Llama 3.2 90B Vision (multimodal) (Meta) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare Llama 3.2 90B Vision (multimodal) and BLIP

Can Llama 3.2 90B Vision (multimodal) process images?

Yes. Llama 3.2 90B Vision (multimodal) accepts images as input in addition to text.

Can I use Llama 3.2 90B Vision (multimodal) right now?

Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.