Fan Zhang, Junwei Cao, et al.
IEEE TETC
We describe a multimodal attentive environment system that performs joint audio-visual information processing to enable it to interact intelligently with people. It integrates real-time video and audio processing techniques to detect and track multiple persons in the scene. Speech recognition and eye contact are used to develop a natural human-like communication interface with participants. We have implemented the system as a visually interactive toy robot (VTOYS) and demonstrated it successfully to many people belonging to different age classes. This allows us to explore novel ways of human-machine interactions and novel interfaces-specifically, the new possibilities of the human-machine interaction for the case of the machine having a limited environment perception ability.
Fan Zhang, Junwei Cao, et al.
IEEE TETC
Belle L. Tseng, Zon-Yin Shae, et al.
ICME 2001
Rajeev Gupta, Shourya Roy, et al.
ICAC 2006
David S. Kung
DAC 1998