Multimodal instruction-tuned describes a model trained to follow user instructions across multiple data modalities, such as text and images. This tuning helps the model align multimodal inputs with task-oriented outputs in a conversational or assistant-like setting.