Word-Sense Disambiguation Using Statistical Models of Roget's Categories Trained on Large Corpora

15 years 8 months ago

Download www.aclweb.org

This paper describes a program that disambignates English word senses in unrestricted text using statistical models of the major Roget's Thesaurus categories. Roget's categories serve as approximations of conceptual classes. The categories listed for a word in Roger's index tend to correspond to sense distinctions; thus selecting the most likely category provides a useful level of sense disambiguatiou. The selection of categories is accomplished by identifying and weighting words that are indicative of each category when seen in context, using a Bayesian theoretical framework. Other statistical approaches have required special corpora or hand-labeled training examples for much of the lexicon. Our use of class models overcomes this knowledge acquisition bottleneck, enabling training on unresUicted monolingual text without human intervention. Applied to the 10 million word Grolier's Encyclopedia, the system correctly disambiguated 92% of the instances of 12 polysemou...

David Yarowsky

Real-time Traffic

COLING 1992 | COLING 2008 | English Word Senses | Roget's Categories | Roget's Thesaurus Categories |

claim paper

Added	07 Nov 2010
Updated	07 Nov 2010
Type	Conference
Year	1992
Where	COLING
Authors	David Yarowsky

Sciweavers

Word-Sense Disambiguation Using Statistical Models of Roget's Categories Trained on Large Corpora

COLING 1992 | COLING 2008 | English Word Senses | Roget's Categories | Roget's Thesaurus Categories |

Explore & Download

Productivity Tools

Sciweavers