Sciweavers

ADMA
2008
Springer

Using Data Mining Methods to Predict Personally Identifiable Information in Emails

14 years 5 months ago
Using Data Mining Methods to Predict Personally Identifiable Information in Emails
Private information management and compliance are important issues nowadays for most of organizations. As a major communication tool for organizations, email is one of the many potential sources for privacy leaks. Information extraction methods have been applied to detect private information in text files. However, since email messages usually consist of low quality text, information extraction methods for private information detection may not achieve good performance. In this paper, we address the problem of predicting the presence of private information in email using data mining and text mining methods. Two prediction models are proposed. The first model is based on association rules that predict one type of private information based on other types of private information identified in emails. The second model is based on classification models that predict private information according to the content of the emails. Experiments on the Enron email dataset show promising results.
Liqiang Geng, Larry Korba, Xin Wang, Yunli Wang, H
Added 01 Jun 2010
Updated 01 Jun 2010
Type Conference
Year 2008
Where ADMA
Authors Liqiang Geng, Larry Korba, Xin Wang, Yunli Wang, Hongyu Liu, Yonghua You
Comments (0)