Sciweavers

SIGIR
2006
ACM

Finding near-duplicate web pages: a large-scale evaluation of algorithms

14 years 5 months ago
Finding near-duplicate web pages: a large-scale evaluation of algorithms
Broder et al.’s [3] shingling algorithm and Charikar’s [4] random projection based approach are considered “state-of-theart” algorithms for finding near-duplicate web pages. Both algorithms were either developed at or used by popular web search engines. We compare the two algorithms on a very
Monika Rauch Henzinger
Added 14 Jun 2010
Updated 14 Jun 2010
Type Conference
Year 2006
Where SIGIR
Authors Monika Rauch Henzinger
Comments (0)