A screenshot accompanied by a written question gives a model two kinds of information. A multimodal model can process and connect such formats: it might explain a chart while answering a text question, or combine visual and audio information from a video.
Inputs become numerical representations whose relationships are learned during training. Some systems connect separate vision, language and audio components; others use more tightly integrated training. The label does not guarantee support for every input and output format. Those capabilities need to be checked for the particular product.
Uses include scanned documents, diagram explanations and accessibility. For reporting on unusual footage, content analysis and source verification remain complementary tasks. A model may describe an object without knowing where the image originated or whether it was edited. Checking provenance, dates and distribution history answers different questions. Connected to tools and action permissions, a multimodal model can also form part of an AI agent.
