Concept

Multimodal

Model that accepts or produces more than one input/output type — text, image, audio, video.

Definition

Multimodal models fuse different modalities — usually by embedding them into a shared latent space — so the system can reason across them. GPT-4o, Gemini, Claude 3.5+ accept text and images; some also handle audio or video. Output multimodality (text-to-image, text-to-video) is the inverse direction.

Common use cases

  • Image understanding
  • Video summarisation
  • Voice assistants

Related models

gpt
claude
gemini

Related terms

    Multimodal — AI Glossary | Railwail