{"id":{"repo_id":"ohiolink","oai_identifier":"oai:etd.ohiolink.edu:ucin1353342433"},"canonical_url":"https://search.dev.ndltd.org/etd/ohiolink/oai:etd.ohiolink.edu:ucin1353342433","repository":{"repo_id":"ohiolink","name":"OhioLINK","base_url":"https://etd.ohiolink.edu/acprod/odb_etd/ws/oai/oai"},"display":{"title":"Exploratory Study of Fuzzy Clustering and Set-Distance Based Validation Indexes","abstract":"This thesis is concerned with issues related to clustering. In particular, it addresses the con-vergence speed of fuzzy c-means family of algorithms and cluster validation. The fuzzy c-meansclustering algorithm and its objective function is studied along with a literature review of thespeed of clustering algorithms. After careful examination, several objective functions are derivedby modifying the fuzzy c-means’ objective function.In addition, cluster validation is examined and new set distance based cluster validation indexes(CVI) are proposed which are the ratio of separation between clusters to compactness within acluster. To this end, a new measure of compactness, compactness of a fuzzy partition is presentedand fuzzy derivative of Pompeiu-Hausdorff distance is used as separation.The convergence of fuzzy c-means clustering algorithm is tested on real classification and clus-tering datasets. Under classification datasets, Iris, Breast Cancer Wisconsin and Wine Recognitiondatasets are used. Water Treatment Plant and Libras Movement datasets are used as clusteringdatasets. In classification datasets, the class labels in the data set are used to measure the per-formance. For clustering datasets, Rand index and Jaccard index are used to evaluate clusteringresults.The new set distance based validation indexes are tested on both synthetic and real datasets.Datasets with three, four, five and six clusters are generated by using Gaussian distributions. Theabove mentioned real datasets, Iris, Breast Cancer Wisconsin and Wine Recognition are also used toevaluate the performance of set distance based validation indexes. The result (number of clusters)obtained from the set distance based validation indexes are compared with those obtained from [50]to demonstrate efficiency of set distance based validation indexes and how it considers the structureof underlying data unlike others, [50] in particular.","abstract_html":"This thesis is concerned with issues related to clustering. In particular, it addresses the con-vergence speed of fuzzy c-means family of algorithms and cluster validation. The fuzzy c-meansclustering algorithm and its objective function is studied along with a literature review of thespeed of clustering algorithms. After careful examination, several objective functions are derivedby modifying the fuzzy c-means’ objective function.In addition, cluster validation is examined and new set distance based cluster validation indexes(CVI) are proposed which are the ratio of separation between clusters to compactness within acluster. To this end, a new measure of compactness, compactness of a fuzzy partition is presentedand fuzzy derivative of Pompeiu-Hausdorff distance is used as separation.The convergence of fuzzy c-means clustering algorithm is tested on real classification and clus-tering datasets. Under classification datasets, Iris, Breast Cancer Wisconsin and Wine Recognitiondatasets are used. Water Treatment Plant and Libras Movement datasets are used as clusteringdatasets. In classification datasets, the class labels in the data set are used to measure the per-formance. For clustering datasets, Rand index and Jaccard index are used to evaluate clusteringresults.The new set distance based validation indexes are tested on both synthetic and real datasets.Datasets with three, four, five and six clusters are generated by using Gaussian distributions. Theabove mentioned real datasets, Iris, Breast Cancer Wisconsin and Wine Recognition are also used toevaluate the performance of set distance based validation indexes. The result (number of clusters)obtained from the set distance based validation indexes are compared with those obtained from [50]to demonstrate efficiency of set distance based validation indexes and how it considers the structureof underlying data unlike others, [50] in particular.","abstract_has_math":false,"creators":["Pangaonkar, Manali"],"institution":"University of Cincinnati","degree_name":"MS","degree_level":"masters","degree_discipline":"Engineering and Applied Science: Computer Science","degree_department":null,"school":null,"contributors":["Ralescu, Anca"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012","date_published":"2012","updated_at":"2026-07-24T03:36:23Z","subjects":["Computer Science","Fuzzy Clustering","Cluster Validation","Compactness","Separation","Set Distance","Cluster Comparison"],"languages":["English"],"rights":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://rave.ohiolink.edu/etdc/view?acc_num=ucin1353342433","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ralescu, Anca"]},{"key":"dc:creator","label":"Author","values":["Pangaonkar, Manali"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012"]},{"key":"dc:publisher","label":"Institution","values":["University of Cincinnati / OhioLINK"]},{"key":"dc:type","label":"Dc Type","values":["Electronic Thesis or Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Engineering and Applied Science: Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["MS"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Cincinnati"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science","Fuzzy Clustering","Cluster Validation","Compactness","Separation","Set Distance","Cluster Comparison"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:rights","label":"Dc Rights","values":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://rave.ohiolink.edu/etdc/view?acc_num=ucin1353342433"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This thesis is concerned with issues related to clustering. In particular, it addresses the con-vergence speed of fuzzy c-means family of algorithms and cluster validation. The fuzzy c-meansclustering algorithm and its objective function is studied along with a literature review of thespeed of clustering algorithms. After careful examination, several objective functions are derivedby modifying the fuzzy c-means’ objective function.In addition, cluster validation is examined and new set distance based cluster validation indexes(CVI) are proposed which are the ratio of separation between clusters to compactness within acluster. To this end, a new measure of compactness, compactness of a fuzzy partition is presentedand fuzzy derivative of Pompeiu-Hausdorff distance is used as separation.The convergence of fuzzy c-means clustering algorithm is tested on real classification and clus-tering datasets. Under classification datasets, Iris, Breast Cancer Wisconsin and Wine Recognitiondatasets are used. Water Treatment Plant and Libras Movement datasets are used as clusteringdatasets. In classification datasets, the class labels in the data set are used to measure the per-formance. For clustering datasets, Rand index and Jaccard index are used to evaluate clusteringresults.The new set distance based validation indexes are tested on both synthetic and real datasets.Datasets with three, four, five and six clusters are generated by using Gaussian distributions. Theabove mentioned real datasets, Iris, Breast Cancer Wisconsin and Wine Recognition are also used toevaluate the performance of set distance based validation indexes. The result (number of clusters)obtained from the set distance based validation indexes are compared with those obtained from [50]to demonstrate efficiency of set distance based validation indexes and how it considers the structureof underlying data unlike others, [50] in particular."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf","p.73","621.3 KB"]},{"key":"dc:title","label":"Title","values":["Exploratory Study of Fuzzy Clustering and Set-Distance Based Validation Indexes"]}]}],"canonical_facts":{"dc:contributor":["Ralescu, Anca"],"dc:creator":["Pangaonkar, Manali"],"dc:date":["2012"],"dc:description":["This thesis is concerned with issues related to clustering. In particular, it addresses the con-vergence speed of fuzzy c-means family of algorithms and cluster validation. The fuzzy c-meansclustering algorithm and its objective function is studied along with a literature review of thespeed of clustering algorithms. After careful examination, several objective functions are derivedby modifying the fuzzy c-means’ objective function.In addition, cluster validation is examined and new set distance based cluster validation indexes(CVI) are proposed which are the ratio of separation between clusters to compactness within acluster. To this end, a new measure of compactness, compactness of a fuzzy partition is presentedand fuzzy derivative of Pompeiu-Hausdorff distance is used as separation.The convergence of fuzzy c-means clustering algorithm is tested on real classification and clus-tering datasets. Under classification datasets, Iris, Breast Cancer Wisconsin and Wine Recognitiondatasets are used. Water Treatment Plant and Libras Movement datasets are used as clusteringdatasets. In classification datasets, the class labels in the data set are used to measure the per-formance. For clustering datasets, Rand index and Jaccard index are used to evaluate clusteringresults.The new set distance based validation indexes are tested on both synthetic and real datasets.Datasets with three, four, five and six clusters are generated by using Gaussian distributions. Theabove mentioned real datasets, Iris, Breast Cancer Wisconsin and Wine Recognition are also used toevaluate the performance of set distance based validation indexes. The result (number of clusters)obtained from the set distance based validation indexes are compared with those obtained from [50]to demonstrate efficiency of set distance based validation indexes and how it considers the structureof underlying data unlike others, [50] in particular."],"dc:format":["application/pdf","p.73","621.3 KB"],"dc:identifier":["http://rave.ohiolink.edu/etdc/view?acc_num=ucin1353342433"],"dc:language":["English"],"dc:publisher":["University of Cincinnati / OhioLINK"],"dc:rights":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."],"dc:subject":["Computer Science","Fuzzy Clustering","Cluster Validation","Compactness","Separation","Set Distance","Cluster Comparison"],"dc:title":["Exploratory Study of Fuzzy Clustering and Set-Distance Based Validation Indexes"],"dc:type":["Electronic Thesis or Dissertation"],"thesis:degree_discipline":["Engineering and Applied Science: Computer Science"],"thesis:degree_level":["masters"],"thesis:degree_name":["MS"],"thesis:institution_name":["University of Cincinnati"]},"updated_at":"2026-07-24T03:36:23Z"}