{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/45291"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/45291","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Model selection for correlated data and moment selection from high-dimensional moment conditions","abstract":"High-dimensional correlated data arise frequently in many studies. My primary research interests lie broadly in statistical methodology for correlated data such as longitudinal data and panel data. In this thesis, we address two important but challenging issues: model selection for correlated data with diverging number of parameters and consistent moment selection from high-dimensional moment conditions. Longitudinal data arise frequently in biomedical and genomic research where repeated measurements within subjects are correlated. It is important to select relevant covariates when the dimension of the parameters diverges as the sample size increases. We propose the penalized quadratic inference function to perform model selection and estimation simultaneously in the framework of a diverging number of regression parameters. The penalized quadratic inference function can easily take correlation information from clustered data into account, yet it does not require specifying the likelihood function. This is advantageous compared to existing model selection methods for discrete data with large cluster size. In addition, the proposed approach enjoys the oracle property; it is able to identify non-zero components consistently with probability tending to 1, and any finite linear combination of the estimated non-zero components has an asymptotic normal distribution. We propose an efficient algorithm by selecting an effective tuning parameter to solve the penalized quadratic inference function. Monte Carlo simulation studies have the proposed method selecting the correct model with a high frequency and estimating covariate effects accurately even when the dimension of parameters is high. We illustrate the proposed approach by analyzing periodontal disease data. The generalized method of moments (GMM) approach combines moment conditions optimally to obtain efficient estimation without specifying the full likelihood function. However, the GMM estimator could be infeasible when the number of moment conditions exceeds the sample size. This research intends to address issues arising from the motivating problem where the dimension of estimating equations or moment conditions far exceeds the sample size, such as in selecting informative correlation structure or modeling for dynamic panel data. We propose a Bayesian information type of criterion to select the optimal number of linear combinations of moment conditions. In theory, we show that the proposed criterion leads to consistent selection of the number of principal components for the weighting matrix in the GMM. Monte Carlo studies indicate that the proposed method outperforms existing methods in the sense of reducing bias and improving the efficiency of estimation. We also illustrate a real data example for moment selection using dynamic panel data models.","abstract_html":"High-dimensional correlated data arise frequently in many studies. My primary research interests lie broadly in statistical methodology for correlated data such as longitudinal data and panel data. In this thesis, we address two important but challenging issues: model selection for correlated data with diverging number of parameters and consistent moment selection from high-dimensional moment conditions. Longitudinal data arise frequently in biomedical and genomic research where repeated measurements within subjects are correlated. It is important to select relevant covariates when the dimension of the parameters diverges as the sample size increases. We propose the penalized quadratic inference function to perform model selection and estimation simultaneously in the framework of a diverging number of regression parameters. The penalized quadratic inference function can easily take correlation information from clustered data into account, yet it does not require specifying the likelihood function. This is advantageous compared to existing model selection methods for discrete data with large cluster size. In addition, the proposed approach enjoys the oracle property; it is able to identify non-zero components consistently with probability tending to 1, and any finite linear combination of the estimated non-zero components has an asymptotic normal distribution. We propose an efficient algorithm by selecting an effective tuning parameter to solve the penalized quadratic inference function. Monte Carlo simulation studies have the proposed method selecting the correct model with a high frequency and estimating covariate effects accurately even when the dimension of parameters is high. We illustrate the proposed approach by analyzing periodontal disease data. The generalized method of moments (GMM) approach combines moment conditions optimally to obtain efficient estimation without specifying the full likelihood function. However, the GMM estimator could be infeasible when the number of moment conditions exceeds the sample size. This research intends to address issues arising from the motivating problem where the dimension of estimating equations or moment conditions far exceeds the sample size, such as in selecting informative correlation structure or modeling for dynamic panel data. We propose a Bayesian information type of criterion to select the optimal number of linear combinations of moment conditions. In theory, we show that the proposed criterion leads to consistent selection of the number of principal components for the weighting matrix in the GMM. Monte Carlo studies indicate that the proposed method outperforms existing methods in the sense of reducing bias and improving the efficiency of estimation. We also illustrate a real data example for moment selection using dynamic panel data models.","abstract_has_math":false,"creators":["Cho, Hyun Keun"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Qu, Annie","Douglas, Jeffrey A.","Shao, Xiaofeng","Simpson, Douglas G."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013-08-22T16:35:00Z","date_published":"2013-08-22T16:35:00Z","updated_at":"2026-07-22T22:25:34Z","subjects":["diverging number of parameters","dynamic panel data models","generalized method of moments","high-dimensional moment conditions","moment selection","Longitudinal data","model selection","oracle property","quadratic inference function","smoothly clipped absolute deviation (SCAD)","singularity matrix"],"languages":["en"],"rights":["Copyright 2013 Hyun Keun Cho"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/45291","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Qu, Annie","Douglas, Jeffrey A.","Shao, Xiaofeng","Simpson, Douglas G."]},{"key":"dc:creator","label":"Author","values":["Cho, Hyun Keun"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2013-08-22T16:35:00Z","2013-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["diverging number of parameters","dynamic panel data models","generalized method of moments","high-dimensional moment conditions","moment selection","Longitudinal data","model selection","oracle property","quadratic inference function","smoothly clipped absolute deviation (SCAD)","singularity matrix"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2013 Hyun Keun Cho"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/45291"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["High-dimensional correlated data arise frequently in many studies. My primary research interests lie broadly in statistical methodology for correlated data such as longitudinal data and panel data. In this thesis, we address two important but challenging issues: model selection for correlated data with diverging number of parameters and consistent moment selection from high-dimensional moment conditions. Longitudinal data arise frequently in biomedical and genomic research where repeated measurements within subjects are correlated. It is important to select relevant covariates when the dimension of the parameters diverges as the sample size increases. We propose the penalized quadratic inference function to perform model selection and estimation simultaneously in the framework of a diverging number of regression parameters. The penalized quadratic inference function can easily take correlation information from clustered data into account, yet it does not require specifying the likelihood function. This is advantageous compared to existing model selection methods for discrete data with large cluster size. In addition, the proposed approach enjoys the oracle property; it is able to identify non-zero components consistently with probability tending to 1, and any finite linear combination of the estimated non-zero components has an asymptotic normal distribution. We propose an efficient algorithm by selecting an effective tuning parameter to solve the penalized quadratic inference function. Monte Carlo simulation studies have the proposed method selecting the correct model with a high frequency and estimating covariate effects accurately even when the dimension of parameters is high. We illustrate the proposed approach by analyzing periodontal disease data. The generalized method of moments (GMM) approach combines moment conditions optimally to obtain efficient estimation without specifying the full likelihood function. However, the GMM estimator could be infeasible when the number of moment conditions exceeds the sample size. This research intends to address issues arising from the motivating problem where the dimension of estimating equations or moment conditions far exceeds the sample size, such as in selecting informative correlation structure or modeling for dynamic panel data. We propose a Bayesian information type of criterion to select the optimal number of linear combinations of moment conditions. In theory, we show that the proposed criterion leads to consistent selection of the number of principal components for the weighting matrix in the GMM. Monte Carlo studies indicate that the proposed method outperforms existing methods in the sense of reducing bias and improving the efficiency of estimation. We also illustrate a real data example for moment selection using dynamic panel data models.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-06-18T20:37:03Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Hyun Keun_Cho.pdf: 476665 bytes, checksum: dc2216889b5d5f77163f653fd4d6ba57 (MD5) Cho_Hyun Keun.pdf: 476665 bytes, checksum: dc2216889b5d5f77163f653fd4d6ba57 (MD5)","Made available in DSpace on 2013-08-22T16:35:00Z (GMT). No. of bitstreams: 2 Hyun Keun_Cho.pdf: 476665 bytes, checksum: dc2216889b5d5f77163f653fd4d6ba57 (MD5) license.txt: 4060 bytes, checksum: e3ea00e84c4478adc519a51ee189900c (MD5)"]},{"key":"dc:title","label":"Title","values":["Model selection for correlated data and moment selection from high-dimensional moment conditions"]}]}],"canonical_facts":{"dc:contributor":["Qu, Annie","Douglas, Jeffrey A.","Shao, Xiaofeng","Simpson, Douglas G."],"dc:creator":["Cho, Hyun Keun"],"dc:date":["2013-08-22T16:35:00Z","2013-08"],"dc:description":["High-dimensional correlated data arise frequently in many studies. My primary research interests lie broadly in statistical methodology for correlated data such as longitudinal data and panel data. In this thesis, we address two important but challenging issues: model selection for correlated data with diverging number of parameters and consistent moment selection from high-dimensional moment conditions. Longitudinal data arise frequently in biomedical and genomic research where repeated measurements within subjects are correlated. It is important to select relevant covariates when the dimension of the parameters diverges as the sample size increases. We propose the penalized quadratic inference function to perform model selection and estimation simultaneously in the framework of a diverging number of regression parameters. The penalized quadratic inference function can easily take correlation information from clustered data into account, yet it does not require specifying the likelihood function. This is advantageous compared to existing model selection methods for discrete data with large cluster size. In addition, the proposed approach enjoys the oracle property; it is able to identify non-zero components consistently with probability tending to 1, and any finite linear combination of the estimated non-zero components has an asymptotic normal distribution. We propose an efficient algorithm by selecting an effective tuning parameter to solve the penalized quadratic inference function. Monte Carlo simulation studies have the proposed method selecting the correct model with a high frequency and estimating covariate effects accurately even when the dimension of parameters is high. We illustrate the proposed approach by analyzing periodontal disease data. The generalized method of moments (GMM) approach combines moment conditions optimally to obtain efficient estimation without specifying the full likelihood function. However, the GMM estimator could be infeasible when the number of moment conditions exceeds the sample size. This research intends to address issues arising from the motivating problem where the dimension of estimating equations or moment conditions far exceeds the sample size, such as in selecting informative correlation structure or modeling for dynamic panel data. We propose a Bayesian information type of criterion to select the optimal number of linear combinations of moment conditions. In theory, we show that the proposed criterion leads to consistent selection of the number of principal components for the weighting matrix in the GMM. Monte Carlo studies indicate that the proposed method outperforms existing methods in the sense of reducing bias and improving the efficiency of estimation. We also illustrate a real data example for moment selection using dynamic panel data models.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-06-18T20:37:03Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Hyun Keun_Cho.pdf: 476665 bytes, checksum: dc2216889b5d5f77163f653fd4d6ba57 (MD5) Cho_Hyun Keun.pdf: 476665 bytes, checksum: dc2216889b5d5f77163f653fd4d6ba57 (MD5)","Made available in DSpace on 2013-08-22T16:35:00Z (GMT). No. of bitstreams: 2 Hyun Keun_Cho.pdf: 476665 bytes, checksum: dc2216889b5d5f77163f653fd4d6ba57 (MD5) license.txt: 4060 bytes, checksum: e3ea00e84c4478adc519a51ee189900c (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/45291"],"dc:language":["en"],"dc:rights":["Copyright 2013 Hyun Keun Cho"],"dc:subject":["diverging number of parameters","dynamic panel data models","generalized method of moments","high-dimensional moment conditions","moment selection","Longitudinal data","model selection","oracle property","quadratic inference function","smoothly clipped absolute deviation (SCAD)","singularity matrix"],"dc:title":["Model selection for correlated data and moment selection from high-dimensional moment conditions"],"dc:type":["text"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:34Z"}