{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/87410"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/87410","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Bicriterion Clustering and Selecting the Optimal Number of Clusters via Agreement Measure","abstract":"Clustering and classification have been important tools to address a broad range of problems in fields such as image analysis, genomics, and many other areas. Basically, these clustering problems can be simplified as two aspects. The first is to estimate the number of clusters. The second one is to allocate each observation to the clusters. Many different heuristic criteria are available. The representative models are k-means, hierarchical clustering and partitioning around medoids. Among these methods, there exists the problem to select the number of clusters. In addition, some algorithms make use of a starting allocation of the observations, such as k-means, which may contain the inherent bias. Often the data partitioning will suffer lack of consistency across different criteria and algorithms. In this thesis, we propose an approach to select the number of clusters through comparing and optimizing the agreement between two clustering criteria. The intuition is that the clustering randomness from different criteria should be minimized when the true clustering structure is recovered. By maximizing the agreement on allocation of the observations between different methods, it selects the optimal number of clusters and also results in a robust consensus set of clusters. Furthermore we use a number of classification rules to combine the resultant clusters from two algorithms. The favorable performance of the method is demonstrated in simulation studies and fMRI time series application. Finally the asymptotic properties of the agreement statistics are discussed.","abstract_html":"Clustering and classification have been important tools to address a broad range of problems in fields such as image analysis, genomics, and many other areas. Basically, these clustering problems can be simplified as two aspects. The first is to estimate the number of clusters. The second one is to allocate each observation to the clusters. Many different heuristic criteria are available. The representative models are k-means, hierarchical clustering and partitioning around medoids. Among these methods, there exists the problem to select the number of clusters. In addition, some algorithms make use of a starting allocation of the observations, such as k-means, which may contain the inherent bias. Often the data partitioning will suffer lack of consistency across different criteria and algorithms. In this thesis, we propose an approach to select the number of clusters through comparing and optimizing the agreement between two clustering criteria. The intuition is that the clustering randomness from different criteria should be minimized when the true clustering structure is recovered. By maximizing the agreement on allocation of the observations between different methods, it selects the optimal number of clusters and also results in a robust consensus set of clusters. Furthermore we use a number of classification rules to combine the resultant clusters from two algorithms. The favorable performance of the method is demonstrated in simulation studies and fMRI time series application. Finally the asymptotic properties of the agreement statistics are discussed.","abstract_has_math":false,"creators":["Liu, Heng"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Douglas Simpson"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-09-28T16:02:45Z","date_published":"2015-09-28T16:02:45Z","updated_at":"2026-07-22T22:26:30Z","subjects":["Statistics"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["(MiAaPQ)AAI3269964"],"render_values":[{"text":"(MiAaPQ)AAI3269964","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/87410","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Douglas Simpson"]},{"key":"dc:creator","label":"Author","values":["Liu, Heng"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-09-28T16:02:45Z","10000-01-01","2007"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Statistics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/87410","(MiAaPQ)AAI3269964"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Clustering and classification have been important tools to address a broad range of problems in fields such as image analysis, genomics, and many other areas. Basically, these clustering problems can be simplified as two aspects. The first is to estimate the number of clusters. The second one is to allocate each observation to the clusters. Many different heuristic criteria are available. The representative models are k-means, hierarchical clustering and partitioning around medoids. Among these methods, there exists the problem to select the number of clusters. In addition, some algorithms make use of a starting allocation of the observations, such as k-means, which may contain the inherent bias. Often the data partitioning will suffer lack of consistency across different criteria and algorithms. In this thesis, we propose an approach to select the number of clusters through comparing and optimizing the agreement between two clustering criteria. The intuition is that the clustering randomness from different criteria should be minimized when the true clustering structure is recovered. By maximizing the agreement on allocation of the observations between different methods, it selects the optimal number of clusters and also results in a robust consensus set of clusters. Furthermore we use a number of classification rules to combine the resultant clusters from two algorithms. The favorable performance of the method is demonstrated in simulation studies and fMRI time series application. Finally the asymptotic properties of the agreement statistics are discussed.","Made available in DSpace on 2015-09-28T16:02:45Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3269964.pdf: 2472031 bytes, checksum: 9b898974bc4b8462f183d919d52b9968 (MD5) Previous issue date: 2007","Embargo set by: Seth Robbins for item 88691 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","101 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2007."]},{"key":"dc:title","label":"Title","values":["Bicriterion Clustering and Selecting the Optimal Number of Clusters via Agreement Measure"]}]}],"canonical_facts":{"dc:contributor":["Douglas Simpson"],"dc:creator":["Liu, Heng"],"dc:date":["2015-09-28T16:02:45Z","10000-01-01","2007"],"dc:description":["Clustering and classification have been important tools to address a broad range of problems in fields such as image analysis, genomics, and many other areas. Basically, these clustering problems can be simplified as two aspects. The first is to estimate the number of clusters. The second one is to allocate each observation to the clusters. Many different heuristic criteria are available. The representative models are k-means, hierarchical clustering and partitioning around medoids. Among these methods, there exists the problem to select the number of clusters. In addition, some algorithms make use of a starting allocation of the observations, such as k-means, which may contain the inherent bias. Often the data partitioning will suffer lack of consistency across different criteria and algorithms. In this thesis, we propose an approach to select the number of clusters through comparing and optimizing the agreement between two clustering criteria. The intuition is that the clustering randomness from different criteria should be minimized when the true clustering structure is recovered. By maximizing the agreement on allocation of the observations between different methods, it selects the optimal number of clusters and also results in a robust consensus set of clusters. Furthermore we use a number of classification rules to combine the resultant clusters from two algorithms. The favorable performance of the method is demonstrated in simulation studies and fMRI time series application. Finally the asymptotic properties of the agreement statistics are discussed.","Made available in DSpace on 2015-09-28T16:02:45Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3269964.pdf: 2472031 bytes, checksum: 9b898974bc4b8462f183d919d52b9968 (MD5) Previous issue date: 2007","Embargo set by: Seth Robbins for item 88691 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","101 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2007."],"dc:identifier":["http://hdl.handle.net/2142/87410","(MiAaPQ)AAI3269964"],"dc:language":["eng"],"dc:subject":["Statistics"],"dc:title":["Bicriterion Clustering and Selecting the Optimal Number of Clusters via Agreement Measure"],"dc:type":["text"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:30Z"}