Sciweavers

DCC
2000
IEEE

Summary Structures for Frequency Queries on Large Transaction Sets

14 years 3 months ago
Summary Structures for Frequency Queries on Large Transaction Sets
As large-scale databases become commonplace, there has been signi cant interest in mining them for commercial purposes. One of the basic tasks that underlies many of these mining operations is querying of transaction sets for frequencies of speci ed attribute values. The size of these databases makes it important to develop summary structures capable of high compression ratios as well as supporting fast frequency queries. The nature of the problem and its di erences with respect to traditional text compression allows very high compression ratios. In this paper, we propose a binary trie-based summary structure for representing transaction sets. We demonstrate that this trie structure, when augmented with an appropriate set of horizontal pointers, can support frequency queries several orders of magnitude faster than raw transaction data. We improve the memory characteristics of our scheme by compressing the trie into a Patricia trie and demonstrate that this does not have a signi cant a...
Dow-Yung Yang, Akshay Johar, Ananth Grama, Wojciec
Added 30 Jul 2010
Updated 30 Jul 2010
Type Conference
Year 2000
Where DCC
Authors Dow-Yung Yang, Akshay Johar, Ananth Grama, Wojciech Szpankowski
Comments (0)