Sciweavers

HPDC
2010
IEEE

Efficient querying of distributed provenance stores

14 years 1 months ago
Efficient querying of distributed provenance stores
Current projects that automate the collection of provenance information use a centralized architecture for managing the resulting metadata - that is, provenance is gathered at remote hosts and submitted to a central provenance management service. In contrast, we are developing a completely decentralized system with each computer maintaining the authoritative repository of the provenance gathered on it. Our model has several advantages, such as scaling to large amounts of metadata generation, providing low-latency access to provenance metadata about local data, avoiding the need for synchronization with a central service after operating while disconnected from the network, and letting users retain control over their data provenance records. We describe the SPADE project's support for tracking data provenance in distributed environments, including how queries can be optimized with provenance sketches, pre-caching, and caching. Categories and Subject Descriptors H.3.3 [Information S...
Ashish Gehani, Minyoung Kim, Tanu Malik
Added 09 Nov 2010
Updated 09 Nov 2010
Type Conference
Year 2010
Where HPDC
Authors Ashish Gehani, Minyoung Kim, Tanu Malik
Comments (0)