Acquisition of Language and Action Concepts from Multimodal Information
Through an integrated cognitive architecture that fuses multimodal information (visual, auditory, and bodily), we study how robots simultaneously acquire and ground action and …
We work on how to design and train robots so that they can act autonomously in the real world. Centering on imitation learning, reinforcement learning, and foundation models such as large language models (LLMs) and vision-language models (VLMs), we build systems that let robots understand their environment, plan actions, and cooperate with humans.
Underlying this work is the question: how does an embodied intelligence connect with the world, and with people? We organize the research into four sub-areas: Knowledge and Concept Learning, Motor Skill Learning, Action Planning and Multi-Robot Coordination, and Human-Robot Interaction.
Through an integrated cognitive architecture that fuses multimodal information (visual, auditory, and bodily), we study how robots simultaneously acquire and ground action and …
Inspired by curiosity-driven exploration in infants, we build learning models in which robot action acquisition and concept formation progress in a mutually reinforcing way. …
Using large language models (LLMs) and vision-language models (VLMs), the robot autonomously generates demonstrations from task instructions, enabling motion-skill learning without …
We study world-model-based reinforcement learning for mobile manipulation that jointly addresses embodiment selection (when to move the base vs. use the arm) and motion planning. …
We apply foundation models such as GPT to humanoid motor control, building unified policies that handle diverse motion tasks including locomotion and whole-body actions. By …
We extend vision-language-action (VLA) foundation models with attention-guided mask images, making the correspondence between natural-language and visual instructions explicit so …
To support foundation-model learning on humanoids and quadruped robots, we develop whole-body teleoperation interfaces and data-collection environments. From affordable …
We systematically survey how data-driven approaches — deep neural networks, reinforcement learning, and large language models — can be applied to robot motion planning, and explore …
We integrate the commonsense reasoning of large language models (LLMs) with structured optimization such as linear programming and dependency graphs, enabling cooperative task …
We develop soft tactile sensors that enable rich bodily contact between robots and humans/environments. From multi-axis force sensors that combine magnetorheological elastomers …
We build frameworks that let autonomous robots explain their decision-making processes and the underlying reasoning to people in an understandable way. Through distributed …
Intent signals from brain-machine interfaces (BMIs) are inherently low-bandwidth and noisy, making fine-grained robot control difficult. We build shared autonomy frameworks in …
We study how the embodiment and dialogue design of robots can support communication between geographically separated people. From remote childcare robots that scaffold infant …
Through explainable machine learning (XAI), we analyze behavioral data from child-robot interactions to estimate toddlers’ temperament and personality. By clarifying in which …