Multimodal Instruction Foundation Models with Mask Images

2024年1月1日 · 1 min read
research

We extend vision-language-action (VLA) foundation models with attention-guided mask images, making the correspondence between natural-language and visual instructions explicit so that robots can accurately capture instruction targets and act on them. Through Alpha-channel attention control, we aim at action selection consistent with both language and vision even in cluttered environments.

Related papers:

  • Shogo Yanagida, Tatsuya Aoki, Tadahiro Taniguchi, Takato Horii. “Mask-guided VLA: Introducing Vision-Language Instructions with Attention-Guided Mask Images” (in Japanese). Annual Conference of the Japanese Society for Artificial Intelligence (JSAI), June 2026.
  • Shogo Yanagida, Tatsuya Aoki, Tadahiro Taniguchi, Takato Horii. “A Robot Foundation Model Understanding Language and Visual Instructions via Alpha-Channel Attention Control” (in Japanese). 43rd Annual Conference of the Robotics Society of Japan, September 2025.
Takato Horii
Authors
Associate Professor
Associate Professor at Graduate School of Engineering Science, Osaka University. Research interests include cognitive developmental robotics, computational modeling of emotional development, and human-robot interaction.