Cooperative Crawling

16 years 2 days ago

Download www.cwr.cl

Web crawler design presents many different challenges: architecture, strategies, performance and more. One of the most important research topics concerns improving the selection of “interesting” web pages (for the user), according to importance metrics. Another relevant point is content freshness, i.e. maintaining freshness and consistency of temporary stored copies. For this, the crawler periodically repeats its activity going over stored contents (re-crawling process). In this paper, we propose a scheme to permit a crawler to acquire information about the global state of a website before the crawling process takes place. This scheme requires web server cooperation in order to collect and publish information on its content, useful for enabling a crawler to tune its visit strategy. If this information is unavailable or not updated the crawler still acts in the usual manner. In this sense the proposed scheme is not invasive and is independent from any crawling strategy and architec...

Marina Buzzi

Real-time Traffic

Human Computer Interaction | Important Research Topics | Internet Technology | LAWEB 2003 | Temporary Stored Copies | Web Crawler Design |

claim paper

Added	05 Jul 2010
Updated	05 Jul 2010
Type	Conference
Year	2003
Where	LAWEB
Authors	Marina Buzzi

Sciweavers

Cooperative Crawling

Human Computer Interaction | Important Research Topics | Internet Technology | LAWEB 2003 | Temporary Stored Copies | Web Crawler Design |

Explore & Download

Productivity Tools

Sciweavers