On the Relative Hardness of Clustering Corpora

16 years 24 days ago

Download users.dsic.upv.es

Abstract. Clustering is often considered the most important unsupervised learning problem and several clustering algorithms have been proposed over the years. Many of these algorithms have been tested on classical clustering corpora such as Reuters and 20 Newsgroups in order to determine their quality. However, up to now the relative hardness of those corpora has not been determined. The relative clustering hardness of a given corpus may be of high interest, since it would help to determine whether the usual corpora used to benchmark the clustering algorithms are hard enough. Moreover, if it is possible to ﬁnd a set of features involved in the hardness of the clustering task itself, speciﬁc clustering techniques may be used instead of general ones in order to improve the quality of the obtained clusters. In this paper, we are presenting a study of the speciﬁc feature of the vocabulary overlapping among documents of a given corpus. Our preliminary experiments were carried out on t...

David Pinto, Paolo Rosso

Real-time Traffic