{"id":{"repo_id":"ubc","oai_identifier":"oai:circle.library.ubc.ca:2429/1496"},"canonical_url":"https://search.dev.ndltd.org/etd/ubc/oai:circle.library.ubc.ca:2429/1496","repository":{"repo_id":"ubc","name":"University of British Columbia","base_url":"http://circle.library.ubc.ca/oai/request"},"display":{"title":"Bayesian cluster validation","abstract":"We propose a novel framework based on Bayesian principles for validating clusterings and present efficient algorithms for use with centroid or exemplar based clustering solutions. Our framework treats the data as fixed and introduces perturbations into the clustering procedure. In our algorithms, we scale the distances between points by a random variable whose distribution is tuned against a baseline null dataset. The random variable is integrated out, yielding a soft assignment matrix that gives the behavior under perturbation of the points relative to each of the clusters. From this soft assignment matrix, we are able to visualize inter-cluster behavior, rank clusters, and give a scalar index of the the clustering stability. In a large test on synthetic data, our method matches or outperforms other leading methods at predicting the correct number of clusters. We also present a theoretical analysis of our approach, which suggests that it is useful for high dimensional data.","abstract_html":"We propose a novel framework based on Bayesian principles for validating clusterings and present efficient algorithms for use with centroid or exemplar based clustering solutions. Our framework treats the data as fixed and introduces perturbations into the clustering procedure. In our algorithms, we scale the distances between points by a random variable whose distribution is tuned against a baseline null dataset. The random variable is integrated out, yielding a soft assignment matrix that gives the behavior under perturbation of the points relative to each of the clusters. From this soft assignment matrix, we are able to visualize inter-cluster behavior, rank clusters, and give a scalar index of the the clustering stability. In a large test on synthetic data, our method matches or outperforms other leading methods at predicting the correct number of clusters. We also present a theoretical analysis of our approach, which suggests that it is useful for high dimensional data.","abstract_has_math":false,"creators":["Koepke, Hoyt Adam"],"institution":"University of British Columbia","degree_name":"Master of Science - MSc","degree_level":"master's","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2008,"date_issued":"2008","date_published":"2008","updated_at":"2026-07-24T05:07:37Z","subjects":[],"languages":["eng"],"rights":["Attribution-NonCommercial-NoDerivatives 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by-nc-nd/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2429/1496","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Koepke, Hoyt Adam"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2008"]},{"key":"dc:publisher","label":"Institution","values":["University of British Columbia"]},{"key":"dc:type","label":"Dc Type","values":["Text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["master's"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science - MSc"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of British Columbia"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["http://creativecommons.org/licenses/by-nc-nd/4.0/","Attribution-NonCommercial-NoDerivatives 4.0 International"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2429/1496","http://circle.library.ubc.ca/bitstream/2429/1496/1/ubc_2008_fall_koepke_hoyt.pdf"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We propose a novel framework based on Bayesian principles for validating clusterings and present efficient algorithms for use with centroid or exemplar based clustering solutions. Our framework treats the data as fixed and introduces perturbations into the clustering procedure. In our algorithms, we scale the distances between points by a random variable whose distribution is tuned against a baseline null dataset. The random variable is integrated out, yielding a soft assignment matrix that gives the behavior under perturbation of the points relative to each of the clusters. From this soft assignment matrix, we are able to visualize inter-cluster behavior, rank clusters, and give a scalar index of the the clustering stability. In a large test on synthetic data, our method matches or outperforms other leading methods at predicting the correct number of clusters. We also present a theoretical analysis of our approach, which suggests that it is useful for high dimensional data."]},{"key":"dc:format","label":"Dc Format","values":["4024259","application/pdf"]},{"key":"dc:title","label":"Title","values":["Bayesian cluster validation"]}]}],"canonical_facts":{"dc:creator":["Koepke, Hoyt Adam"],"dc:date":["2008"],"dc:description":["We propose a novel framework based on Bayesian principles for validating clusterings and present efficient algorithms for use with centroid or exemplar based clustering solutions. Our framework treats the data as fixed and introduces perturbations into the clustering procedure. In our algorithms, we scale the distances between points by a random variable whose distribution is tuned against a baseline null dataset. The random variable is integrated out, yielding a soft assignment matrix that gives the behavior under perturbation of the points relative to each of the clusters. From this soft assignment matrix, we are able to visualize inter-cluster behavior, rank clusters, and give a scalar index of the the clustering stability. In a large test on synthetic data, our method matches or outperforms other leading methods at predicting the correct number of clusters. We also present a theoretical analysis of our approach, which suggests that it is useful for high dimensional data."],"dc:format":["4024259","application/pdf"],"dc:identifier":["http://hdl.handle.net/2429/1496","http://circle.library.ubc.ca/bitstream/2429/1496/1/ubc_2008_fall_koepke_hoyt.pdf"],"dc:language":["eng"],"dc:publisher":["University of British Columbia"],"dc:rights":["http://creativecommons.org/licenses/by-nc-nd/4.0/","Attribution-NonCommercial-NoDerivatives 4.0 International"],"dc:title":["Bayesian cluster validation"],"dc:type":["Text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["master's"],"thesis:degree_name":["Master of Science - MSc"],"thesis:institution_name":["University of British Columbia"]},"updated_at":"2026-07-24T05:07:37Z"}