Clustered Sequence Representation for Fast Homology Search

15 years 6 months ago

Download web.udl.es

We present a novel approach to managing redundancy in sequence databanks such as GenBank. We store clusters of near-identical sequences as a representative union-sequence and a set of corresponding edits to that sequence. During search, the query is compared to only the union-sequences representing each cluster; cluster members are then only reconstructed and aligned if the union-sequence achieves a sufﬁciently high score. Using this approach with BLAST results in a 27% reduction in collection size and a corresponding 22% decrease in search time with no signiﬁcant change in accuracy. We also describe our method for clustering that uses ﬁngerprinting, an approach that has been successfully applied to collections of text and web documents in Information Retrieval. Our clustering approach is ten times faster on the GenBank nonredundant protein database than the fastest existing approach, CD-HIT. We have integrated our approach into FSA-BLAST, our new Open Source version of BLAST (a...

Michael Cameron, Yaniv Bernstein, Hugh E. Williams

Real-time Traffic