{"id":{"repo_id":"columbus-state","oai_identifier":"oai:csuepress.columbusstate.edu:theses_dissertations-1231"},"canonical_url":"https://search.dev.ndltd.org/etd/columbus-state/oai:csuepress.columbusstate.edu:theses_dissertations-1231","repository":{"repo_id":"columbus-state","name":"Columbus State University","base_url":"https://csuepress.columbusstate.edu/do/oai/"},"display":{"title":"An Arithmetic-Based Deterministic Centroid Initialization Method for the k-Means Clustering Algorithm","abstract":"<p>One of the greatest challenges in k-means clustering is positioning the initial cluster centers, or centroids, as close to optimal as possible, and doing so in an amount of time deemed reasonable. Traditional fc-means utilizes a randomization process for initializing these centroids, and poor initialization can lead to increased numbers of required clustering iterations to reach convergence, and a greater overall runtime. This research proposes a simple, arithmetic-based deterministic centroid initialization method which is much faster than randomized initialization. Preliminary experiments suggest that this collection of methods, referred to herein as the sharding centroid initialization algorithm family, often outperforms random initialization in terms of the required number of iterations for convergence and overall time-related metrics and is competitive or better in terms of the reported mean sum of squared errors (SSE) metric. Surprisingly, the sharding algorithms often manage to report more advantageous mean SSE values in the instances where their performance is slower than random initialization.</p>","abstract_html":"&lt;p&gt;One of the greatest challenges in k-means clustering is positioning the initial cluster centers, or centroids, as close to optimal as possible, and doing so in an amount of time deemed reasonable. Traditional fc-means utilizes a randomization process for initializing these centroids, and poor initialization can lead to increased numbers of required clustering iterations to reach convergence, and a greater overall runtime. This research proposes a simple, arithmetic-based deterministic centroid initialization method which is much faster than randomized initialization. Preliminary experiments suggest that this collection of methods, referred to herein as the sharding centroid initialization algorithm family, often outperforms random initialization in terms of the required number of iterations for convergence and overall time-related metrics and is competitive or better in terms of the reported mean sum of squared errors (SSE) metric. Surprisingly, the sharding algorithms often manage to report more advantageous mean SSE values in the instances where their performance is slower than random initialization.&lt;/p&gt;","abstract_has_math":false,"creators":["Mayo, Matthew Michael"],"institution":null,"degree_name":"Computer Science - Applied Computing Track","degree_level":"Thesis","degree_discipline":"TSYS School of Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-05-01T07:00:00Z","date_published":"2016-05-01T07:00:00Z","updated_at":"2026-07-24T01:44:55Z","subjects":["K-means","Randomized Initialization","Sum of Squared Errors (SSE)","Centroid","Computer Engineering","Theory and Algorithms"],"languages":["English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://csuepress.columbusstate.edu/theses_dissertations/241","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Mayo, Matthew Michael"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2017-07-25T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["TSYS School of Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Computer Science - Applied Computing Track"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["K-means","Randomized Initialization","Sum of Squared Errors (SSE)","Centroid","Computer Engineering","Theory and Algorithms"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://csuepress.columbusstate.edu/theses_dissertations/241"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>One of the greatest challenges in k-means clustering is positioning the initial cluster centers, or centroids, as close to optimal as possible, and doing so in an amount of time deemed reasonable. Traditional fc-means utilizes a randomization process for initializing these centroids, and poor initialization can lead to increased numbers of required clustering iterations to reach convergence, and a greater overall runtime. This research proposes a simple, arithmetic-based deterministic centroid initialization method which is much faster than randomized initialization. Preliminary experiments suggest that this collection of methods, referred to herein as the sharding centroid initialization algorithm family, often outperforms random initialization in terms of the required number of iterations for convergence and overall time-related metrics and is competitive or better in terms of the reported mean sum of squared errors (SSE) metric. Surprisingly, the sharding algorithms often manage to report more advantageous mean SSE values in the instances where their performance is slower than random initialization.</p>"]},{"key":"dc:title","label":"Title","values":["An Arithmetic-Based Deterministic Centroid Initialization Method for the k-Means Clustering Algorithm"]}]}],"canonical_facts":{"dc:creator":["Mayo, Matthew Michael"],"dc:date.available":["2017-07-25T07:00:00Z"],"dc:description.abstract":["<p>One of the greatest challenges in k-means clustering is positioning the initial cluster centers, or centroids, as close to optimal as possible, and doing so in an amount of time deemed reasonable. Traditional fc-means utilizes a randomization process for initializing these centroids, and poor initialization can lead to increased numbers of required clustering iterations to reach convergence, and a greater overall runtime. This research proposes a simple, arithmetic-based deterministic centroid initialization method which is much faster than randomized initialization. Preliminary experiments suggest that this collection of methods, referred to herein as the sharding centroid initialization algorithm family, often outperforms random initialization in terms of the required number of iterations for convergence and overall time-related metrics and is competitive or better in terms of the reported mean sum of squared errors (SSE) metric. Surprisingly, the sharding algorithms often manage to report more advantageous mean SSE values in the instances where their performance is slower than random initialization.</p>"],"dc:identifier":["https://csuepress.columbusstate.edu/theses_dissertations/241"],"dc:language":["English"],"dc:subject":["K-means","Randomized Initialization","Sum of Squared Errors (SSE)","Centroid","Computer Engineering","Theory and Algorithms"],"dc:title":["An Arithmetic-Based Deterministic Centroid Initialization Method for the k-Means Clustering Algorithm"],"thesis:degree_discipline":["TSYS School of Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Computer Science - Applied Computing Track"]},"updated_at":"2026-07-24T01:44:55Z"}