Glossary

Multimodal Model

A multimodal model can process or generate more than one kind of information, such as text, images, audio or video. In AI search, multimodal input can let users ask questions about screenshots, products or documents, while retrieval may include visual as well as textual evidence.

In plain terms

It is an AI model that works with formats beyond text.

Why it matters

Brands need accurate captions, alt text and surrounding copy because visual assets still require context to be found and interpreted reliably.

How to apply it

  • Give meaningful images descriptive alt text and captions.
  • Place visual evidence near its textual explanation.
  • Test image-led questions separately from text-only prompts.

Example

A buyer uploads a product screenshot and asks an assistant to identify compatible software, combining image understanding with web retrieval.

Sources

Related reading

Related terms

Back to the full glossary (75 terms).

Ranking is no longer enough

You need to be cited, mentioned, and recommended.

Being cited, mentioned, and recommended are three different outcomes, and most brands only ever achieve the first one. Ranking is no longer enough because AI engines answer buyers directly and name only a short list of vendors as the recommendation — everyone else is cited in passing, if at all. Get a free AI Visibility Report to see exactly where your brand appears today across ChatGPT, Google AI Overviews, Gemini, Perplexity, and Copilot, where competitors are winning the recommendation instead, and what's keeping you from moving up the shortlist.