Back to results

University of Washington

How to Use K-means for Big Data Clustering?

Abstract

dc:description.abstract

K-means plays a vital role in data mining, being the simplest and most widely used algorithm under the Euclidean Minimum Sum-of-Squares Clustering (MSSC) model. However, its performance drastically drops when applied to vast amounts of data. Therefore, it is crucial to improve K-means by scaling it to big data using as few of the following computational resources as possible: data, time, and algorithmic ingredients. We introduce a novel parallel scheme that leverages K-means and K-means++ algorithms for big data clustering, offering a "true big data" algorithm that excels in both solution quality and runtime, surpassing classical and recent state-of-the-art MSSC approaches. The new approach naturally implements global search by decomposing the MSSC problem without using additional metaheuristics. On the other hand, this approach can be generalized to a novel metaheuristic, providing fresh perspectives for creating new powerful optimization heuristics. This work shows that data decomposition is the basic approach to solve the big data clustering problem. The empirical success of the new algorithm and its derivatives allowed us to challenge the common belief that more data is required to obtain a good clustering solution. Moreover, the present work questions the established trend that more sophisticated hybrid approaches and algorithms are required to obtain a better clustering solution.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Mussabayev, Ravil
Advisor dc:contributor.advisor
  • Uhlmann, Gunther

Subjects

dc:subject × 9

Rights

dc:rights
Statement dc:rights
  • CC BY-NC-SA
Language dc:language.iso
en_US

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1773/52103
OAI identifier oai:identifier
oai:digital.lib.washington.edu:1773/52103

Chain of custody

source
Harvested from
University of Washington
Base URL
digital.lib.washington.edu/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Mussabayev, Ravil. How to Use K-means for Big Data Clustering?. 2024. https://hdl.handle.net/1773/52103