Sciweavers

APWEB
2006
Springer

Sample Sizes for Query Probing in Uncooperative Distributed Information Retrieval

14 years 3 months ago
Sample Sizes for Query Probing in Uncooperative Distributed Information Retrieval
The goal of distributed information retrieval is to support effective searching over multiple document collections. For efficiency, queries should be routed to only those collections that are likely to contain relevant documents, so it is necessary to first obtain information about the content of the target collections. In an uncooperative environment, query probing -- where randomly-chosen queries are used to retrieve a sample of the documents and thus of the lexicon -- has been proposed as a technique for estimating statistical term distributions. In this paper we rebut the claim that a sample of 300 documents is sufficient to provide good coverage of collection terms. We propose a novel sampling strategy and experimentally demonstrate that sample size needs to vary from collection to collection, that our methods achieve good coverage based on variable-sized samples, and that we can use the results of a probe to determine when to stop sampling.
Milad Shokouhi, Falk Scholer, Justin Zobel
Added 20 Aug 2010
Updated 20 Aug 2010
Type Conference
Year 2006
Where APWEB
Authors Milad Shokouhi, Falk Scholer, Justin Zobel
Comments (0)