Model-Based Estimation of Word Saliency in Text

14 years 4 months ago

Download www.cs.bham.ac.uk

Abstract. We investigate a generative latent variable model for modelbased word saliency estimation for text modelling and classification. The estimation algorithm derived is able to infer the saliency of words with respect to the mixture modelling objective. We demonstrate experimental results showing that common stop-words as well as other corpus-specific common words are automatically down-weighted and this enhances our ability to capture the essential structure in the data, ignoring irrelevant details. As a classifier, our approach improves over the class prediction accuracy of the Naive Bayes classifier in all our experiments. Compared with a recent state of the art text classification method (Dirichlet Compound Multinomial model) we obtained improved results in two out of three benchmark text collections tested, and comparable results on one other data set.

Xin Wang, Ata Kabán

Real-time Traffic

DIS 2006 | Mixture Modelling Objective | Naive Bayes Classifier | Theoretical Computer Science | Word Saliency Estimation |

claim paper

Post Info
More Details (n/a)

Added	22 Aug 2010
Updated	22 Aug 2010
Type	Conference
Year	2006
Where	DIS
Authors	Xin Wang, Ata Kabán

Comments (0)

Sciweavers

Model-Based Estimation of Word Saliency in Text

DIS 2006 | Mixture Modelling Objective | Naive Bayes Classifier | Theoretical Computer Science | Word Saliency Estimation |

Explore & Download

Productivity Tools

Document Tools

Image Tools

Sciweavers