Concept
aka VLM
Vision-Language Model
Multimodal model that jointly understands images and text in a shared representation.
Definition
Vision-language models accept image and text inputs and reason over both. Architectures range from CLIP-style dual encoders to Llava/Flamingo-style fused decoders. They power visual question answering, OCR-aware reasoning, GUI agents, and robotics policies.
Common use cases
- Visual Q&A
- OCR
- Robotic perception
Related models
gpt
claude
gemini
pixtral