Text Clustering with Feature Selection by Using Statistical Data

14 years 19 days ago

Download dblab.mgt.ncu.edu.tw

Abstract-- Feature selection is an important method for improving the efficiency and accuracy of text categorization algorithms by removing redundant and irrelevant terms from the corpus. In this paper, we propose a new supervised feature selection method, named CHIR, which is based on the 2 statistic and new statistical data that can measure the positive termcategory dependency. We also propose a new text clustering algorithm TCFS, which stands for Text Clustering with Feature Selection. TCFS can incorporate CHIR to identify relevant features (i.e., terms) iteratively, and the clustering becomes a learning process. We compared TCFS and the k-means clustering algorithm in combination with different feature selection methods for various real data sets. Our experimental results show that TCFS with CHIR has better clustering accuracy in terms of the F-measure and the purity.

Yanjun Li, Congnan Luo, Soon M. Chung

Real-time Traffic

Feature Selection | Feature Selection Methods | Supervised Feature Selection | TKDE 2008 |

claim paper

Post Info
More Details (n/a)

Added	15 Dec 2010
Updated	15 Dec 2010
Type	Journal
Year	2008
Where	TKDE
Authors	Yanjun Li, Congnan Luo, Soon M. Chung

Comments (0)

Sciweavers

Text Clustering with Feature Selection by Using Statistical Data

Feature Selection | Feature Selection Methods | Supervised Feature Selection | TKDE 2008 |

Explore & Download

Productivity Tools

Document Tools

Image Tools

Sciweavers