Multimodal

A model that accepts or produces more than one kind of input - text plus images, audio or video.

A multimodal model encodes non-text input into the same representation space as text, so a picture and a sentence about it can be reasoned about together. In practice this means you can paste a screenshot, a chart, a PDF page or a video and ask questions about it in plain language.

Capabilities differ sharply by provider and change often. Image input is now common; audio and video input are not universal; image output is a separate model in most stacks. Check what the specific model version supports rather than assuming parity.

Where this comes up

Prompts, configs and tutorials in the library that touch this term.