multimodal
Multimodal describes AI that processes more than one type of data, such as text, images, and audio.
Multimodal describes an AI model that accepts or generates two or more types of data, such as text, images, and audio. It contrasts with a unimodal model that handles only one type, such as text.
The term became widespread around 2023, when major large language models such as GPT-4 and Gemini began accepting image input. Computer-use models, which read the screen and operate a computer through mouse and keyboard actions, likewise process screenshots together with text instructions.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
Handle multimodal responses by checking rather than hard-coding nested paths.