Multimodal Instruction Foundation Models with Mask Images
2024年1月1日
·
1 min read

We extend vision-language-action (VLA) foundation models with attention-guided mask images, making the correspondence between natural-language and visual instructions explicit so that robots can accurately capture instruction targets and act on them. Through Alpha-channel attention control, we aim at action selection consistent with both language and vision even in cluttered environments.
Related papers:
- Shogo Yanagida, Tatsuya Aoki, Tadahiro Taniguchi, Takato Horii. “Mask-guided VLA: Introducing Vision-Language Instructions with Attention-Guided Mask Images” (in Japanese). Annual Conference of the Japanese Society for Artificial Intelligence (JSAI), June 2026.
- Shogo Yanagida, Tatsuya Aoki, Tadahiro Taniguchi, Takato Horii. “A Robot Foundation Model Understanding Language and Visual Instructions via Alpha-Channel Attention Control” (in Japanese). 43rd Annual Conference of the Robotics Society of Japan, September 2025.

Authors
Takato Horii
(he/him)
Associate Professor
Associate Professor at Graduate School of Engineering Science, Osaka University.
Research interests include cognitive developmental robotics, computational modeling
of emotional development, and human-robot interaction.