Sciweavers

BIBE
2005
IEEE

Using Data Mining Techniques to Learn Layouts of Flat-File Biological Datasets

14 years 6 months ago
Using Data Mining Techniques to Learn Layouts of Flat-File Biological Datasets
One of the major problems in biological data integration is that many data sources are stored as flat-files, with a variety of different layouts. Integrating data from such sources can be an extremely time-consuming task. We have been developing data mining techniques to help learn the layout of a dataset in a semi-automatic way. In this paper, we focus on the problem of identifying delimiters for optional fields. Since these fields do not occur in every record, frequency based methods are not able to identify the corresponding delimiters. We present a method which uses contrast analysis on the frequency of sequences to identify such delimiters and help complete the layout descriptions. We demonstrate the effectiveness of this technique using three flat-file biological datasets.
Kaushik Sinha, Xuan Zhang, Ruoming Jin, Gagan Agra
Added 24 Jun 2010
Updated 24 Jun 2010
Type Conference
Year 2005
Where BIBE
Authors Kaushik Sinha, Xuan Zhang, Ruoming Jin, Gagan Agrawal
Comments (0)