Parametric Models of Linguistic Count Data

14 years 27 days ago

Download acl.ldc.upenn.edu

It is well known that occurrence counts of words in documents are often modeled poorly by standard distributions like the binomial or Poisson. Observed counts vary more than simple models predict, prompting the use of overdispersed models like Gamma-Poisson or Beta-binomial mixtures as robust alternatives. Another deﬁciency of standard models is due to the fact that most words never occur in a given document, resulting in large amounts of zero counts. We propose using zeroinﬂated models for dealing with this, and evaluate competing models on a Naive Bayes text classiﬁcation task. Simple zero-inﬂated models can account for practically relevant variation, and can be easier to work with than overdispersed models.

Martin Jansche

Real-time Traffic

ACL 2003 | ACL 2007 | Overdispersed Models | Simple Models | Simple Zero-inﬂated Models |

claim paper

Post Info
More Details (n/a)

Added	31 Oct 2010
Updated	31 Oct 2010
Type	Conference
Year	2003
Where	ACL
Authors	Martin Jansche

Comments (0)

Sciweavers

Parametric Models of Linguistic Count Data

ACL 2003 | ACL 2007 | Overdispersed Models | Simple Models | Simple Zero-inﬂated Models |

Explore & Download

Productivity Tools

Document Tools

Image Tools

Sciweavers