Concept
Multimodal
Model that accepts or produces more than one input/output type — text, image, audio, video.
Definition
Multimodal models fuse different modalities — usually by embedding them into a shared latent space — so the system can reason across them. GPT-4o, Gemini, Claude 3.5+ accept text and images; some also handle audio or video. Output multimodality (text-to-image, text-to-video) is the inverse direction.
Common use cases
- Image understanding
- Video summarisation
- Voice assistants
Related models
gpt
claude
gemini