{"id":{"repo_id":"umn","oai_identifier":"oai:conservancy.umn.edu:11299/113175"},"canonical_url":"https://search.dev.ndltd.org/etd/umn/oai:conservancy.umn.edu:11299/113175","repository":{"repo_id":"umn","name":"University of Minnesota","base_url":"https://conservancy.umn.edu/server/oai/request"},"display":{"title":"Statistical methods for gene set based significance analysis.","abstract":"Gene set enrichment analysis (GSEA) is a method to identify groups of genes, which are statistically more differentially expressed than all other genes across different treatments within a microarray study. Most of the existing approaches have largely relied on nonparametric methods and require repeated computation of permutation and resampling data to assess the significance of a gene set. In this dissertation, we study parametric approaches for GSEA by formulating the enrichment analysis into a simple model comparison problem. The methods not only gain the flexibility in statistical modeling corresponding to biological problems but also achieve computational efficiency. First, we propose a likelihood based approach assuming a finite mixture model for a two-class comparison problem and the implementation of the analysis is achieved by a likelihood ratio based testing approach. In addition we extend the parametric methods to flexible two-component mixture models for one-sided enrichment analysis which aims to test for enrichment of up (or down) regulation only. Also, we develop chi-square mixture models which incorporate the idea of two-class comparison studies into multiple category microarray experiments. Applications to gene expression data, along with simulations, demonstrate the computational efficiency and the competitive performance of the proposed methods.","abstract_html":"Gene set enrichment analysis (GSEA) is a method to identify groups of genes, which are statistically more differentially expressed than all other genes across different treatments within a microarray study. Most of the existing approaches have largely relied on nonparametric methods and require repeated computation of permutation and resampling data to assess the significance of a gene set. In this dissertation, we study parametric approaches for GSEA by formulating the enrichment analysis into a simple model comparison problem. The methods not only gain the flexibility in statistical modeling corresponding to biological problems but also achieve computational efficiency. First, we propose a likelihood based approach assuming a finite mixture model for a two-class comparison problem and the implementation of the analysis is achieved by a likelihood ratio based testing approach. In addition we extend the parametric methods to flexible two-component mixture models for one-sided enrichment analysis which aims to test for enrichment of up (or down) regulation only. Also, we develop chi-square mixture models which incorporate the idea of two-class comparison studies into multiple category microarray experiments. Applications to gene expression data, along with simulations, demonstrate the computational efficiency and the competitive performance of the proposed methods.","abstract_has_math":false,"creators":["Lee, Sang Mee"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-07","date_published":"2011-07","updated_at":"2026-07-24T05:20:01Z","subjects":["EM","Enrichment analysis","LRT","Mixture model","Biostatistics"],"languages":["en_US"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://purl.umn.edu/113175","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Lee, Sang Mee"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2011-08-17T17:47:00Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2011-08-17T17:47:00Z"]},{"key":"dc:date.issued","label":"Date","values":["2011-07"]},{"key":"dc:type","label":"Dc Type","values":["Thesis or Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["EM","Enrichment analysis","LRT","Mixture model","Biostatistics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://purl.umn.edu/113175"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["University of Minnesota Ph.D. dissertation. July 2011. Major: Biostatistics. Advisors:Dr. Wei Pan and Dr. Baolin Wu. 1 computer file (PDF); viii, 102 pages."]},{"key":"dc:description.abstract","label":"Abstract","values":["Gene set enrichment analysis (GSEA) is a method to identify groups of genes, which are statistically more differentially expressed than all other genes across different treatments within a microarray study. Most of the existing approaches have largely relied on nonparametric methods and require repeated computation of permutation and resampling data to assess the significance of a gene set. In this dissertation, we study parametric approaches for GSEA by formulating the enrichment analysis into a simple model comparison problem. The methods not only gain the flexibility in statistical modeling corresponding to biological problems but also achieve computational efficiency. First, we propose a likelihood based approach assuming a finite mixture model for a two-class comparison problem and the implementation of the analysis is achieved by a likelihood ratio based testing approach. In addition we extend the parametric methods to flexible two-component mixture models for one-sided enrichment analysis which aims to test for enrichment of up (or down) regulation only. Also, we develop chi-square mixture models which incorporate the idea of two-class comparison studies into multiple category microarray experiments. Applications to gene expression data, along with simulations, demonstrate the computational efficiency and the competitive performance of the proposed methods."]},{"key":"dc:title","label":"Title","values":["Statistical methods for gene set based significance analysis."]}]}],"canonical_facts":{"dc:creator":["Lee, Sang Mee"],"dc:date.accessioned":["2011-08-17T17:47:00Z"],"dc:date.available":["2011-08-17T17:47:00Z"],"dc:date.issued":["2011-07"],"dc:description":["University of Minnesota Ph.D. dissertation. July 2011. Major: Biostatistics. Advisors:Dr. Wei Pan and Dr. Baolin Wu. 1 computer file (PDF); viii, 102 pages."],"dc:description.abstract":["Gene set enrichment analysis (GSEA) is a method to identify groups of genes, which are statistically more differentially expressed than all other genes across different treatments within a microarray study. Most of the existing approaches have largely relied on nonparametric methods and require repeated computation of permutation and resampling data to assess the significance of a gene set. In this dissertation, we study parametric approaches for GSEA by formulating the enrichment analysis into a simple model comparison problem. The methods not only gain the flexibility in statistical modeling corresponding to biological problems but also achieve computational efficiency. First, we propose a likelihood based approach assuming a finite mixture model for a two-class comparison problem and the implementation of the analysis is achieved by a likelihood ratio based testing approach. In addition we extend the parametric methods to flexible two-component mixture models for one-sided enrichment analysis which aims to test for enrichment of up (or down) regulation only. Also, we develop chi-square mixture models which incorporate the idea of two-class comparison studies into multiple category microarray experiments. Applications to gene expression data, along with simulations, demonstrate the computational efficiency and the competitive performance of the proposed methods."],"dc:identifier.uri":["http://purl.umn.edu/113175"],"dc:language.iso":["en_US"],"dc:subject":["EM","Enrichment analysis","LRT","Mixture model","Biostatistics"],"dc:title":["Statistical methods for gene set based significance analysis."],"dc:type":["Thesis or Dissertation"]},"updated_at":"2026-07-24T05:20:01Z"}