Automatic topic detection strategy for information retrieval in spoken document

15 years 9 months ago

Download www.dcs.gla.ac.uk

This paper suggests an alternative solution for the task of spoken document retrieval (SDR). The proposed system runs retrieval on multi-level transcriptions (word and phone) produced by word and phone recognizers respectively, and their outputs are combined. We propose to use latent Dirichlet allocation (LDA) model for capturing the semantic information on word transcription. The LDA model is employed for estimating topic distribution in queries and word transcribed spoken documents, and the matching is performed at the topic level. Acoustic matching between query words and phonetically transcribed spoken documents is performed using phone-based matching algorithm. The results of acoustic and topic level matching methods are compared and shown to be complementary.

Shan Jin, Hemant Misra, Thomas Sikora, Joemon M. J

Real-time Traffic

Image Analysis | Spoken Document Retrieval | Spoken Documents | Topic Level | WIAMIS 2009 |

claim paper

Post Info
More Details (n/a)

Added	21 May 2010
Updated	21 May 2010
Type	Conference
Year	2009
Where	WIAMIS
Authors	Shan Jin, Hemant Misra, Thomas Sikora, Joemon M. Jose

Comments (0)

Sciweavers

Automatic topic detection strategy for information retrieval in spoken document

Image Analysis | Spoken Document Retrieval | Spoken Documents | Topic Level | WIAMIS 2009 |

Explore & Download

Productivity Tools

Document Tools

Image Tools

Sciweavers