{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/45589"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/45589","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Statistical modeling of heterogeneous data","abstract":"This dissertation is centered on the modeling of heterogeneous data which is ubiquitous in this digital information age. From the statistical point of view heterogeneous data is composed of dissimilar components, where objects in each component are homogeneous themselves. One such example from the real world is the stock return data, where stocks in the same industry segments tend to move closely together, while different segments tend to have distinct movement patterns. Clustering is one of the most popular ways to characterize data heterogeneity. It is a classical problem of unsupervised learning. We will review major clustering approaches in Chapter 1. In recent years non-parametric Bayesian mixture models have attracted increasing attention in the clustering literature, which is closely related with our work. So we review the Mixture of Dirichlet Process Model in Chapter 2. The main dissertation body consists of three generic statistical methods to model heterogeneity in different scenarios. As data are becoming more and more prevailing today, traditional clustering tasks are often accompanied by additional information about the objects to cluster, known as the side information. The opportunity is that the side information has the potential to complement clustering algorithms to achieve more accurate and meaningful results. In Chapter 3 we describe Two-view Clustering method, a novel non-parametric clustering model that is capable of robustly incorporating noisy side information. We demonstrate the effectiveness of this new model with three real world applications in Chapter 4. Our second work is driven by market segmentation which is a key factor to a modern business's success by accurately recognizing customer groups with varying needs. Market segmentation involves dividing a larger market into sub-markets based upon a variety of factors such customers’ demographic information and product preferences. In Chapter 5 we will propose a multi-task learning framework to solve this problem. Our third work in Chapter 6 tries to solve a problem arising from citation analysis for research evaluation. In bibliometrics one central task is to characterize the statistical distribution of citations. This problem has been regarded as a challenging one for two reasons: (i) the citation distributions of almost all the subject areas are highly right-skewed; (ii) the citation behaviors across various subject areas can be drastically different. We propose a mixture model to formally characterize the statistical distribution of citation data. Based on this model we develop new criteria to evaluate impact of journals and performance of research institutes.","abstract_html":"This dissertation is centered on the modeling of heterogeneous data which is ubiquitous in this digital information age. From the statistical point of view heterogeneous data is composed of dissimilar components, where objects in each component are homogeneous themselves. One such example from the real world is the stock return data, where stocks in the same industry segments tend to move closely together, while different segments tend to have distinct movement patterns. Clustering is one of the most popular ways to characterize data heterogeneity. It is a classical problem of unsupervised learning. We will review major clustering approaches in Chapter 1. In recent years non-parametric Bayesian mixture models have attracted increasing attention in the clustering literature, which is closely related with our work. So we review the Mixture of Dirichlet Process Model in Chapter 2. The main dissertation body consists of three generic statistical methods to model heterogeneity in different scenarios. As data are becoming more and more prevailing today, traditional clustering tasks are often accompanied by additional information about the objects to cluster, known as the side information. The opportunity is that the side information has the potential to complement clustering algorithms to achieve more accurate and meaningful results. In Chapter 3 we describe Two-view Clustering method, a novel non-parametric clustering model that is capable of robustly incorporating noisy side information. We demonstrate the effectiveness of this new model with three real world applications in Chapter 4. Our second work is driven by market segmentation which is a key factor to a modern business&#x27;s success by accurately recognizing customer groups with varying needs. Market segmentation involves dividing a larger market into sub-markets based upon a variety of factors such customers’ demographic information and product preferences. In Chapter 5 we will propose a multi-task learning framework to solve this problem. Our third work in Chapter 6 tries to solve a problem arising from citation analysis for research evaluation. In bibliometrics one central task is to characterize the statistical distribution of citations. This problem has been regarded as a challenging one for two reasons: (i) the citation distributions of almost all the subject areas are highly right-skewed; (ii) the citation behaviors across various subject areas can be drastically different. We propose a mixture model to formally characterize the statistical distribution of citation data. Based on this model we develop new criteria to evaluate impact of journals and performance of research institutes.","abstract_has_math":false,"creators":["Liu, Yufei"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Liang, Feng","Simpson, Douglas G.","Marden, John I.","Chen, Yuguo"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013-08-22T16:48:49Z","date_published":"2013-08-22T16:48:49Z","updated_at":"2026-07-22T22:25:36Z","subjects":["Statistical Learning","Clustering","Non-parametric Bayes","Dirichlet Process","Mixture Model","Heterogeneous Data"],"languages":["en"],"rights":["Copyright 2013 Jeffrey Yufei Liu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/45589","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Liang, Feng","Simpson, Douglas G.","Marden, John I.","Chen, Yuguo"]},{"key":"dc:creator","label":"Author","values":["Liu, Yufei"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2013-08-22T16:48:49Z","2015-08-22T10:00:46Z","2013-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Statistical Learning","Clustering","Non-parametric Bayes","Dirichlet Process","Mixture Model","Heterogeneous Data"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2013 Jeffrey Yufei Liu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/45589"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This dissertation is centered on the modeling of heterogeneous data which is ubiquitous in this digital information age. From the statistical point of view heterogeneous data is composed of dissimilar components, where objects in each component are homogeneous themselves. One such example from the real world is the stock return data, where stocks in the same industry segments tend to move closely together, while different segments tend to have distinct movement patterns. Clustering is one of the most popular ways to characterize data heterogeneity. It is a classical problem of unsupervised learning. We will review major clustering approaches in Chapter 1. In recent years non-parametric Bayesian mixture models have attracted increasing attention in the clustering literature, which is closely related with our work. So we review the Mixture of Dirichlet Process Model in Chapter 2. The main dissertation body consists of three generic statistical methods to model heterogeneity in different scenarios. As data are becoming more and more prevailing today, traditional clustering tasks are often accompanied by additional information about the objects to cluster, known as the side information. The opportunity is that the side information has the potential to complement clustering algorithms to achieve more accurate and meaningful results. In Chapter 3 we describe Two-view Clustering method, a novel non-parametric clustering model that is capable of robustly incorporating noisy side information. We demonstrate the effectiveness of this new model with three real world applications in Chapter 4. Our second work is driven by market segmentation which is a key factor to a modern business's success by accurately recognizing customer groups with varying needs. Market segmentation involves dividing a larger market into sub-markets based upon a variety of factors such customers’ demographic information and product preferences. In Chapter 5 we will propose a multi-task learning framework to solve this problem. Our third work in Chapter 6 tries to solve a problem arising from citation analysis for research evaluation. In bibliometrics one central task is to characterize the statistical distribution of citations. This problem has been regarded as a challenging one for two reasons: (i) the citation distributions of almost all the subject areas are highly right-skewed; (ii) the citation behaviors across various subject areas can be drastically different. We propose a mixture model to formally characterize the statistical distribution of citation data. Based on this model we develop new criteria to evaluate impact of journals and performance of research institutes.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-07-03T15:26:52Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Liu_Yufei.pdf: 4311162 bytes, checksum: 7234ee23944529e19a793d1c9682b1d2 (MD5)","Made available in DSpace on 2013-08-22T16:48:49Z (GMT). No. of bitstreams: 2 Yufei_Liu.pdf: 1898708 bytes, checksum: df3e49bc0a1d91f0e14ea59932349edc (MD5) license.txt: 4057 bytes, checksum: 9f09427bcf7be2b1ff3f43393b952705 (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by Seth Robbins (srobbins@illinois.edu) on 2013-08-22T16:49:43Z Item is restricted until 2015-08-22T16:49:27Z","Restriction data tranferred 2014-07-01T11:32:25-05:00 Original Data Group with Access UIUC Users [automated] Release Date: 2015-08-22 11:49:27 UTC Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 45571 on 2015-08-22T10:00:46Z."]},{"key":"dc:title","label":"Title","values":["Statistical modeling of heterogeneous data"]}]}],"canonical_facts":{"dc:contributor":["Liang, Feng","Simpson, Douglas G.","Marden, John I.","Chen, Yuguo"],"dc:creator":["Liu, Yufei"],"dc:date":["2013-08-22T16:48:49Z","2015-08-22T10:00:46Z","2013-08"],"dc:description":["This dissertation is centered on the modeling of heterogeneous data which is ubiquitous in this digital information age. From the statistical point of view heterogeneous data is composed of dissimilar components, where objects in each component are homogeneous themselves. One such example from the real world is the stock return data, where stocks in the same industry segments tend to move closely together, while different segments tend to have distinct movement patterns. Clustering is one of the most popular ways to characterize data heterogeneity. It is a classical problem of unsupervised learning. We will review major clustering approaches in Chapter 1. In recent years non-parametric Bayesian mixture models have attracted increasing attention in the clustering literature, which is closely related with our work. So we review the Mixture of Dirichlet Process Model in Chapter 2. The main dissertation body consists of three generic statistical methods to model heterogeneity in different scenarios. As data are becoming more and more prevailing today, traditional clustering tasks are often accompanied by additional information about the objects to cluster, known as the side information. The opportunity is that the side information has the potential to complement clustering algorithms to achieve more accurate and meaningful results. In Chapter 3 we describe Two-view Clustering method, a novel non-parametric clustering model that is capable of robustly incorporating noisy side information. We demonstrate the effectiveness of this new model with three real world applications in Chapter 4. Our second work is driven by market segmentation which is a key factor to a modern business's success by accurately recognizing customer groups with varying needs. Market segmentation involves dividing a larger market into sub-markets based upon a variety of factors such customers’ demographic information and product preferences. In Chapter 5 we will propose a multi-task learning framework to solve this problem. Our third work in Chapter 6 tries to solve a problem arising from citation analysis for research evaluation. In bibliometrics one central task is to characterize the statistical distribution of citations. This problem has been regarded as a challenging one for two reasons: (i) the citation distributions of almost all the subject areas are highly right-skewed; (ii) the citation behaviors across various subject areas can be drastically different. We propose a mixture model to formally characterize the statistical distribution of citation data. Based on this model we develop new criteria to evaluate impact of journals and performance of research institutes.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-07-03T15:26:52Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Liu_Yufei.pdf: 4311162 bytes, checksum: 7234ee23944529e19a793d1c9682b1d2 (MD5)","Made available in DSpace on 2013-08-22T16:48:49Z (GMT). No. of bitstreams: 2 Yufei_Liu.pdf: 1898708 bytes, checksum: df3e49bc0a1d91f0e14ea59932349edc (MD5) license.txt: 4057 bytes, checksum: 9f09427bcf7be2b1ff3f43393b952705 (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by Seth Robbins (srobbins@illinois.edu) on 2013-08-22T16:49:43Z Item is restricted until 2015-08-22T16:49:27Z","Restriction data tranferred 2014-07-01T11:32:25-05:00 Original Data Group with Access UIUC Users [automated] Release Date: 2015-08-22 11:49:27 UTC Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 45571 on 2015-08-22T10:00:46Z."],"dc:identifier":["http://hdl.handle.net/2142/45589"],"dc:language":["en"],"dc:rights":["Copyright 2013 Jeffrey Yufei Liu"],"dc:subject":["Statistical Learning","Clustering","Non-parametric Bayes","Dirichlet Process","Mixture Model","Heterogeneous Data"],"dc:title":["Statistical modeling of heterogeneous data"],"dc:type":["text"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:36Z"}