A framework for clustering massive-domain data streams

Charu C. Aggarwal

doi:10.1109/ICDE.2009.13

Publication

ICDE 2009

Conference paper

A framework for clustering massive-domain data streams

ICDE 2009

View publication

Abstract

In this paper, we will examine the problem of clustering massive domain data streams. Massive-domain data streams are those in which the number of possible domain values for each attribute are very large and cannot be easily tracked for clustering purposes. Some examples of such streams include IP-address streams, credit-card transaction streams, or streams of sales data over large numbers of items. In such cases, it is well known that even simple stream operations such as counting can be extremely difficult because of the difficulty in maintaining summary information over the different discrete values. The task of clustering is significantly more challenging in such cases, since the intermediate statistics for the different clusters cannot be maintained efficiently. In this paper, we propose a method for clustering massive-domain data streams with the use of sketches. We prove probabilistic results which show that a sketch-based clustering method can provide similar results to an infinitespace clustering algorithm with high probability. We present experimental results which validate these theoretical results, and show that it is possible to approximate the behavior of an infinitespace algorithm accurately. © 2009 IEEE.

Date

08 Jul 2009

Publication

ICDE 2009

Authors

Charu C. Aggarwal

IBM-affiliated at time of publication

Abstract

Date

Publication

Authors

Share