Multimodal AI refers to AI systems that can process, combine, or generate more than one type of data, such as text, images, audio, video, or sensor information. By integrating multiple modalities, these systems can support richer perception, reasoning, and interaction than single-modality systems.