

Diversity analysis on imbalanced data sets by using ensemble models

14 years 9 months ago
Diversity analysis on imbalanced data sets by using ensemble models
— Many real-world applications have problems when learning from imbalanced data sets, such as medical diagnosis, fraud detection, and text classification. Very few minority class instances cannot provide sufficient information and result in performance degrading greatly. As a good way to improve the classification performance of weak learner, some ensemblebased algorithms have been proposed to solve class imbalance problem. However, it is still not clear that how diversity affects classification performance especially on minority classes, since diversity is one influential factor of ensemble. This paper explores the impact of diversity on each class and overall performance. As the other influential factor, accuracy is also discussed because of the trade-off between diversity and accuracy. Firstly, three popular re-sampling methods are combined into our ensemble model and evaluated for diversity analysis, which includes under-sampling, over-sampling, and SMOTE [1] – a data gen...
Shuo Wang, Xin Yao
Added 20 May 2010
Updated 20 May 2010
Type Conference
Year 2009
Where CIDM
Authors Shuo Wang, Xin Yao
Comments (0)