CBC: Clustering Based Text Classification Requiring Minimal Labeled Data

16 years 3 hour ago

Download research.microsoft.com

Semi-supervised learning methods construct classifiers using both labeled and unlabeled training data samples. While unlabeled data samples can help to improve the accuracy of trained models to certain extent, existing methods still face difficulties when labeled data is not sufficient and biased against the underlying data distribution. In this paper, we present a clustering based classification (CBC) approach. Using this approach, training data, including both the labeled and unlabeled data, is first clustered with the guidance of the labeled data. Some of unlabeled data samples are then labeled based on the clusters obtained. Discriminative classifiers can subsequently be trained with the expanded labeled dataset. The effectiveness of the proposed method is justified analytically. Related issues such as expanding labeled dataset and interacting clustering with classification are discussed. Our experimental results demonstrated that CBC outperforms existing algorithms when the size o...

Hua-Jun Zeng, Xuanhui Wang, Zheng Chen, Hongjun Lu

Real-time Traffic

Data Mining | ICDM 2003 | Unlabeled Data | Unlabeled Data Samples | Unlabeled Training Data |

claim paper

» Effective multilabel active learning for text classification

» Multilabel ASRS Dataset Classification Using Semi Supervised Subspace Clustering

» DocumentBase Extraction for SingleLabel Text Classification

» Semisupervised Text Classification Using Partitioned EM

» Training Data Cleaning for Text Classification

» Clustering and Classification of Maintenance Logs using Text Data Mining

» A parallel learning algorithm for text classification

» SED supervised experimental design and its application to text classification

Post Info
More Details (n/a)

Added	04 Jul 2010
Updated	04 Jul 2010
Type	Conference
Year	2003
Where	ICDM
Authors	Hua-Jun Zeng, Xuanhui Wang, Zheng Chen, Hongjun Lu, Wei-Ying Ma

Comments (0)

Sciweavers

CBC: Clustering Based Text Classification Requiring Minimal Labeled Data

Data Mining | ICDM 2003 | Unlabeled Data | Unlabeled Data Samples | Unlabeled Training Data |

Explore & Download

Productivity Tools

Sciweavers