Sciweavers

ICIP
2003
IEEE

Audio-visual speaker identification using coupled hidden Markov models

15 years 2 months ago
Audio-visual speaker identification using coupled hidden Markov models
In this paper, we investigate the use of the coupled hidden Markov models (CHMM) for the task of audio-visual text dependent speaker identification. Our system determines the identity of the user from a temporal sequence of audio and visual observations obtained from the acoustic speech and the shape of the mouth, respectively. The multi modal observation sequences are then modeled using a set of CHMMs, one for each phoneme-viseme pair and for each person in the database. The use of CHMMs in our system is justified by the capacity of this model to describe the natural audio and visual state asynchrony as well as their conditional dependency over time. To train a CHMM we first train a speaker independent model using expectationmaximization (EM), and then we build a speaker dependent model using maximum a posteriori (MAP) training. Experimental results on XM2VTS database show that our system improves the accuracy of audio-only or video-only speaker identification at all levels of acoust...
Tieyan Fu, Xiao Xing Liu, Lu Hong Liang, Xiaobo Pi
Added 24 Oct 2009
Updated 24 Oct 2009
Type Conference
Year 2003
Where ICIP
Authors Tieyan Fu, Xiao Xing Liu, Lu Hong Liang, Xiaobo Pi, Ara V. Nefian
Comments (0)