Concept
aka VLM

Vision-Language Model

Multimodal model that jointly understands images and text in a shared representation.

Definition

Vision-language models accept image and text inputs and reason over both. Architectures range from CLIP-style dual encoders to Llava/Flamingo-style fused decoders. They power visual question answering, OCR-aware reasoning, GUI agents, and robotics policies.

Common use cases

  • Visual Q&A
  • OCR
  • Robotic perception

Related models

gpt
claude
gemini
pixtral

Related terms

    Vision-Language Model — AI Glossary | Railwail