Multimodal
A model that accepts or produces more than one kind of input - text plus images, audio or video.
A multimodal model encodes non-text input into the same representation space as text, so a picture and a sentence about it can be reasoned about together. In practice this means you can paste a screenshot, a chart, a PDF page or a video and ask questions about it in plain language.
Capabilities differ sharply by provider and change often. Image input is now common; audio and video input are not universal; image output is a separate model in most stacks. Check what the specific model version supports rather than assuming parity.
Tagged
Where this comes up
Prompts, configs and tutorials in the library that touch this term.
- PromptUniversal
Strict code review of a diff
A review prompt that reports only what would break in production, gives the smallest fix for each finding, and is allowed to say the diff is fine.
- PromptUniversal
Turn a vague request into a written brief
Takes a one-line request from a colleague or client and returns the questions that have to be answered before work starts.
- GuideUniversal
What a prompt actually is
The three parts every working prompt has, why their order changes the answer, and how to tell a vague prompt from a specific one.
- GuideUniversal
Working with long documents and code
Why long sessions drift, what a large context window actually buys you, and how to keep a model on task across a big input.
- GuideUniversal
Giving a model a role and constraints
How naming the audience, the format and the exclusions turns unpredictable answers into repeatable ones.
No-padding everyday default
The one to set if you set only one. Answer first, nothing before it, nothing after it, and no offer to help further.
Tested on Claude Sonnet · May 2026