{"id":{"repo_id":"vcu","oai_identifier":"oai:scholarscompass.vcu.edu:etd-2308"},"canonical_url":"https://search.dev.ndltd.org/etd/vcu/oai:scholarscompass.vcu.edu:etd-2308","repository":{"repo_id":"vcu","name":"Virginia Commonwealth University","base_url":"https://scholarscompass.vcu.edu/do/oai/"},"display":{"title":"Normal Mixture Models for Gene Cluster Identification in Two Dimensional Microarray Data","abstract":"This dissertation focuses on methodology specific to microarray data analyses that organize the data in preliminary steps and proposes a cluster analysis method which improves the interpretability of the cluster results. Cluster analysis of microarray data allows samples with similar gene expression values to be discovered and may serve as a useful diagnostic tool. Since microarray data is inherently noisy, data preprocessing steps including smoothing and filtering are discussed. Comparing the results of different clustering methods is complicated by the arbitrariness of the cluster labels. Methods for re-labeling clusters to assess the agreement between the results of different clustering techniques are proposed. Microarray data involve large numbers of observations and generally present as arrays of light intensity values reflecting the degree of activity of the genes. These measurements are often two dimensional in nature since each is associated with an individual sample (cell line) and gene. The usual hierarchical clustering techniques do not easily adapt to this type of problem. These techniques allow only one dimension of the data to be clustered at a time and lose information due to the collapsing of the data in the opposite dimension. A novel clustering technique based on normal mixture distribution models is developed. This method clusters observations that arise from the same normal distribution and allows the data to be simultaneously clustered in two dimensions. The model is fitted using the Expectation/Maximization (EM) algorithm. For every cluster, the posterior probability that an observation belongs to that cluster is calculated. These probabilities allow the analyst to control the cluster assignments, including the use of overlapping clusters. A user friendly program, 2-DCluster, was written to support these methods. This program was written for Microsoft Windows 2000 and XP systems and supports one and two dimensional clustering. The program and sample applications are available at http://etd.vcu.edu. An electronic copy of this dissertation is available at the same address.","abstract_html":"This dissertation focuses on methodology specific to microarray data analyses that organize the data in preliminary steps and proposes a cluster analysis method which improves the interpretability of the cluster results. Cluster analysis of microarray data allows samples with similar gene expression values to be discovered and may serve as a useful diagnostic tool. Since microarray data is inherently noisy, data preprocessing steps including smoothing and filtering are discussed. Comparing the results of different clustering methods is complicated by the arbitrariness of the cluster labels. Methods for re-labeling clusters to assess the agreement between the results of different clustering techniques are proposed. Microarray data involve large numbers of observations and generally present as arrays of light intensity values reflecting the degree of activity of the genes. These measurements are often two dimensional in nature since each is associated with an individual sample (cell line) and gene. The usual hierarchical clustering techniques do not easily adapt to this type of problem. These techniques allow only one dimension of the data to be clustered at a time and lose information due to the collapsing of the data in the opposite dimension. A novel clustering technique based on normal mixture distribution models is developed. This method clusters observations that arise from the same normal distribution and allows the data to be simultaneously clustered in two dimensions. The model is fitted using the Expectation/Maximization (EM) algorithm. For every cluster, the posterior probability that an observation belongs to that cluster is calculated. These probabilities allow the analyst to control the cluster assignments, including the use of overlapping clusters. A user friendly program, 2-DCluster, was written to support these methods. This program was written for Microsoft Windows 2000 and XP systems and supports one and two dimensional clustering. The program and sample applications are available at http://etd.vcu.edu. An electronic copy of this dissertation is available at the same address.","abstract_has_math":false,"creators":["Harvey, Eric Scott"],"institution":null,"degree_name":"Doctor of Philosophy","degree_level":"Dissertation","degree_discipline":"Biostatistics","degree_department":null,"school":null,"contributors":["Dr. Viswanathan Ramakrishnan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2003,"date_issued":"2003-01-01T08:00:00Z","date_published":"2003-01-01T08:00:00Z","updated_at":"2026-07-24T05:55:14Z","subjects":["cluster analysis","microarrays","biostatistics","two dimensional cluster analysis","em algorithm","mixture models","Physical Sciences and Mathematics","Statistics and Probability"],"languages":[],"rights":["© The Author"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://scholarscompass.vcu.edu/etd/1309"],"render_values":[{"text":"https://scholarscompass.vcu.edu/etd/1309","href":"https://scholarscompass.vcu.edu/etd/1309","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.25772/P723-C417","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dr. Viswanathan Ramakrishnan"]},{"key":"dc:creator","label":"Author","values":["Harvey, Eric Scott"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2014-07-09T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Biostatistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["cluster analysis","microarrays","biostatistics","two dimensional cluster analysis","em algorithm","mixture models","Physical Sciences and Mathematics","Statistics and Probability"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["© The Author"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://doi.org/10.25772/P723-C417","https://scholarscompass.vcu.edu/etd/1309"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This dissertation focuses on methodology specific to microarray data analyses that organize the data in preliminary steps and proposes a cluster analysis method which improves the interpretability of the cluster results. Cluster analysis of microarray data allows samples with similar gene expression values to be discovered and may serve as a useful diagnostic tool. Since microarray data is inherently noisy, data preprocessing steps including smoothing and filtering are discussed. Comparing the results of different clustering methods is complicated by the arbitrariness of the cluster labels. Methods for re-labeling clusters to assess the agreement between the results of different clustering techniques are proposed. Microarray data involve large numbers of observations and generally present as arrays of light intensity values reflecting the degree of activity of the genes. These measurements are often two dimensional in nature since each is associated with an individual sample (cell line) and gene. The usual hierarchical clustering techniques do not easily adapt to this type of problem. These techniques allow only one dimension of the data to be clustered at a time and lose information due to the collapsing of the data in the opposite dimension. A novel clustering technique based on normal mixture distribution models is developed. This method clusters observations that arise from the same normal distribution and allows the data to be simultaneously clustered in two dimensions. The model is fitted using the Expectation/Maximization (EM) algorithm. For every cluster, the posterior probability that an observation belongs to that cluster is calculated. These probabilities allow the analyst to control the cluster assignments, including the use of overlapping clusters. A user friendly program, 2-DCluster, was written to support these methods. This program was written for Microsoft Windows 2000 and XP systems and supports one and two dimensional clustering. The program and sample applications are available at http://etd.vcu.edu. An electronic copy of this dissertation is available at the same address."]},{"key":"dc:title","label":"Title","values":["Normal Mixture Models for Gene Cluster Identification in Two Dimensional Microarray Data"]}]}],"canonical_facts":{"dc:contributor":["Dr. Viswanathan Ramakrishnan"],"dc:creator":["Harvey, Eric Scott"],"dc:date.available":["2014-07-09T07:00:00Z"],"dc:description.abstract":["This dissertation focuses on methodology specific to microarray data analyses that organize the data in preliminary steps and proposes a cluster analysis method which improves the interpretability of the cluster results. Cluster analysis of microarray data allows samples with similar gene expression values to be discovered and may serve as a useful diagnostic tool. Since microarray data is inherently noisy, data preprocessing steps including smoothing and filtering are discussed. Comparing the results of different clustering methods is complicated by the arbitrariness of the cluster labels. Methods for re-labeling clusters to assess the agreement between the results of different clustering techniques are proposed. Microarray data involve large numbers of observations and generally present as arrays of light intensity values reflecting the degree of activity of the genes. These measurements are often two dimensional in nature since each is associated with an individual sample (cell line) and gene. The usual hierarchical clustering techniques do not easily adapt to this type of problem. These techniques allow only one dimension of the data to be clustered at a time and lose information due to the collapsing of the data in the opposite dimension. A novel clustering technique based on normal mixture distribution models is developed. This method clusters observations that arise from the same normal distribution and allows the data to be simultaneously clustered in two dimensions. The model is fitted using the Expectation/Maximization (EM) algorithm. For every cluster, the posterior probability that an observation belongs to that cluster is calculated. These probabilities allow the analyst to control the cluster assignments, including the use of overlapping clusters. A user friendly program, 2-DCluster, was written to support these methods. This program was written for Microsoft Windows 2000 and XP systems and supports one and two dimensional clustering. The program and sample applications are available at http://etd.vcu.edu. An electronic copy of this dissertation is available at the same address."],"dc:identifier":["https://doi.org/10.25772/P723-C417","https://scholarscompass.vcu.edu/etd/1309"],"dc:rights":["© The Author"],"dc:subject":["cluster analysis","microarrays","biostatistics","two dimensional cluster analysis","em algorithm","mixture models","Physical Sciences and Mathematics","Statistics and Probability"],"dc:title":["Normal Mixture Models for Gene Cluster Identification in Two Dimensional Microarray Data"],"thesis:degree_discipline":["Biostatistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Doctor of Philosophy"]},"updated_at":"2026-07-24T05:55:14Z"}