Llama 3.2 90B Vision (multimodal)

MultimodalNicht verfügbar
von MetaModell-ID: llama-3-2-90b-vision-mm

Meta's flagship vision-language model. 90B parameters, image understanding + chat, strong VQA performance.

Status
Nicht verfügbar
Kontext
131.072 Token
Max. Ausgabe
8.192 Token
Eingabe → Ausgabe
Text + Bild → Text
Entwickler
Meta
Aktualisiert
25. Juni 2026

Llama 3.2 90B Vision (multimodal) ist derzeit nicht verfügbar

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert.

Die Angaben auf dieser Seite kannst du weiter nachlesen. Mit einer der verfügbaren Alternativen unten kannst du sofort ein vergleichbares Modell nutzen.

Zu den Alternativen
01

Vergleichbare Modelle

Alle dieser Kategorie
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ca. 0,00030 $/Lauf

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 $/1 Mio. In

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 $/1 Mio. In

02

Playground

Llama 3.2 90B Vision (multimodal) ausprobieren

Chat

Derzeit nicht verfügbar

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert.

Der Playground ist deaktiviert. Vergleichbare Modelle findest du in derselben Kategorie: Alternativen ansehen

Llama 3.2 90B Vision (multimodal) ausprobieren

Schick eine Nachricht. Die Antwort kommt vollständig, sobald das Modell fertig ist (ohne Streaming).

System-Prompt
Max. Antwortlänge (Token)

Dieser Lauf

Kein Preis – derzeit nicht verfügbar.

Neu hier?

5 Gratis-Credits (0,05 $) bei Anmeldung mit Google

Nutzbar 24 Stunden nach der Anmeldung, bis zu 5 Läufe pro Tag und höchstens 2 Credits je Lauf. Andere Anmeldearten starten ohne Guthaben.

03

Über Llama 3.2 90B Vision (multimodal)

Kurz gesagtStand: 25. Juni 2026

Llama 3.2 90B Vision (multimodal) ist ein Modell von Meta aus der Kategorie Multimodal. Über Railwail ist Llama 3.2 90B Vision (multimodal) derzeit nicht verfügbar. Das Kontextfenster umfasst 131.072 Token, eine Antwort bis zu 8.192 Token.

Hintergrund

Über Meta AI (FAIR)

Gegründet 2013 · Menlo Park, California, USA

Meta AI is the research arm of Meta Platforms, established in 2013 as Facebook AI Research (FAIR) by Yann LeCun. FAIR has open-sourced many foundational models including PyTorch, RoBERTa, DETR, SAM and the LLaMA family. LLaMA 1 was released in February 2023, LLaMA 2 in July 2023, LLaMA 3 in April 2024 and LLaMA 3.1 (405B) in July 2024. Llama 3.2 launched in September 2024 at Meta Connect, introducing the first multimodal models in the LLaMA family (vision-enabled 11B and 90B) together with tiny on-device text-only siblings (1B, 3B). All Llama 3.2 vision weights are released under the Llama 3 Community Licence and are widely used by enterprise customers via Meta's partner ecosystem (Hugging Face, AWS Bedrock, Azure AI Studio, Google Vertex, Together AI, Groq, Fireworks).

Meta AI (FAIR) besuchen

Architektur

Decoder-only Transformer with cross-attended vision encoder

Llama 3.2 90B Vision combines the 70B-parameter Llama 3.1 text backbone (extended to 90B with vision components) and a Vision Transformer image encoder integrated via cross-attention adapter layers, similar in spirit to Flamingo but reusing the LLaMA architecture. The vision tower processes each image to a sequence of visual tokens which are injected into specific cross-attention layers of the LLM decoder while the original text-only weights remain frozen during the multimodal training stage, preserving text-only performance. Pretraining used 6B image-text pairs followed by multi-stage supervised fine-tuning and Direct Preference Optimisation (DPO) on a curated set of image instructions, math and chart data. The model supports a 128K context window and accepts up to 1120x1120 image inputs natively (with tiling for larger images). It does not support video or audio. Llama 3.2 90B Vision is released under the Llama 3 Community Licence (free for commercial use under 700M MAU).

Parameter
90B
Kontext
128.000 Token

Fähigkeiten

  • Open-weights 90B vision-language model under Llama 3 Community Licence
  • 128K token context window
  • Image input up to 1120x1120 with tiling for larger images
  • Chart, diagram, OCR and document understanding
  • Strong on MMMU, MathVista, ChartQA and DocVQA among open-weights models
  • Multilingual: English, German, French, Italian, Portuguese, Spanish, Hindi, Thai
  • Tool use and JSON output via Llama 3.1 alignment recipe
  • Best for: open-weights multimodal apps, on-premise document AI, indie research

Training & Lizenz

Pretrained on 6B image-text pairs from public web and licensed sources; supervised fine-tuning and DPO on curated multimodal instruction data. Text knowledge inherited from Llama 3.1 (15T tokens).

Lizenz: Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.

Sicherheitstests: Meta publishes a comprehensive model card with red-team findings on CBRN, child-safety and hate-speech vectors, plus Llama Guard 3 and Prompt Guard 2 companion models for production safety.

Bekannte Grenzen

  • No video or audio input
  • Latency and cost dominated by 90B params; requires multi-GPU serving
  • Licence restricts the largest hyperscaler use cases
  • Vision quality below GPT-4o and Claude 3.5 Sonnet on hardest charts
  • English-centric in vision domain
04

Preise

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert. Für dieses Modell gibt es derzeit keinen Preis, deshalb lässt es sich nicht ausführen.

05

API

Rufe Llama 3.2 90B Vision (multimodal) mit deinem Railwail-API-Schlüssel auf. Diese Modell-ID gehört in die Anfrage:

Derzeit nicht verfügbar

Das Modell hat keinen geprüften Preis oder ist deaktiviert; API-Aufrufe werden abgelehnt.

06

Spezifikationen

Modell-ID
llama-3-2-90b-vision-mm
Entwickler
Meta
Kategorie
Multimodal
Eingabe
Text, Bild
Ausgabe
Text
Kontextfenster
131.072 Token
Max. Ausgabe
8.192 Token
Modellgröße
90B
Lizenz
Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.
Katalogeintrag aktualisiert
25. Juni 2026

Schlagwörter

  • meta
  • llama
  • multimodal
  • vision
  • open-weights
07

Einsatzgebiete

Wofür es genutzt wird

  • Open-weights document AI and OCR pipelines
  • On-premise vision-language assistants
  • Chart and diagram understanding for analytics
  • Compliance and regulated-industry multimodal apps
  • Research baselines for vision-language models
08

Häufige Fragen

Was ist Llama 3.2 90B Vision (multimodal)?

Llama 3.2 90B Vision (multimodal) ist ein Modell von Meta aus der Kategorie Multimodal. Es steht im Railwail-Katalog, lässt sich derzeit aber nicht ausführen.

Was kostet Llama 3.2 90B Vision (multimodal) bei Railwail?

Llama 3.2 90B Vision (multimodal) lässt sich über Railwail derzeit nicht ausführen, deshalb gibt es keinen aktuellen Preis. Verfügbare Alternativen mit Preisen stehen weiter unten auf dieser Seite.

Wie groß ist das Kontextfenster von Llama 3.2 90B Vision (multimodal)?

Das Kontextfenster von Llama 3.2 90B Vision (multimodal) umfasst 131.072 Token. Eine Antwort kann bis zu 8.192 Token lang sein.

Wie schnell ist Llama 3.2 90B Vision (multimodal)?

Für Llama 3.2 90B Vision (multimodal) gibt es bei Railwail noch zu wenige gemessene Läufe, um eine Laufzeit anzugeben. Sie hängt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Ist Llama 3.2 90B Vision (multimodal) besser als BLIP?

Das hängt von der Aufgabe ab. Llama 3.2 90B Vision (multimodal) (Meta) und BLIP (Salesforce) sind beide Modelle aus der Kategorie Multimodal. Die Vergleichsseite zeigt Preise und Spezifikationen nebeneinander.

Llama 3.2 90B Vision (multimodal) und BLIP vergleichen

Kann Llama 3.2 90B Vision (multimodal) Bilder verarbeiten?

Ja. Llama 3.2 90B Vision (multimodal) nimmt neben Text auch Bilder als Eingabe an.

Kann ich Llama 3.2 90B Vision (multimodal) gerade nutzen?

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert. Die Seite bleibt online; verfügbare Alternativen aus derselben Kategorie stehen weiter unten.

Alle Modelle über eine API

Ein API-Schlüssel für alle Modelle auf Railwail. Abgerechnet wird über vorab gekaufte Credits, 1 Credit = 0,01 $.