{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/108306"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/108306","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Identifiability for latent class models","abstract":"Despite the fact that latent class models have been widely applied and appeared to perform well in various applications, it is well known that inference for latent class models is challenging due to the potential nonidentifiability of the underlying model parameters. In other words, multiple or infinite many sets of parameters of a latent class model could correspond to the same data generating process, i.e., are not statistically differentiable. This thesis studies the identifiability issue of some latent class models arising from real applications in cognitive diagnosis and topic modelings. The main contribution of this thesis include identifiability conditions that are weaker than the existing ones in the current literature, computational tools such as MCMC algorithms and EM algorithms that return estimates and uncertainty quantification of the underlying model parameters, and asymptotic analysis such as consistency study on the proposed estimation procedures for identifiable models. In the first part of the thesis, we consider Cognitive Diagnostic Models (CDMs). CDMs are latent variable models developed to infer latent skills, knowledge, or personalities that underlie responses to educational, psychological, and social science tests and measures. We derive a new set of sufficient conditions for generic identifiability of CDM parameters. An important contribution for practice is that our new generic identifiability conditions are more likely to be satisfied in empirical applications than existing conditions that ensure strict identifiability. For computation, we formulate learning the underlying latent structure as a variable selection problem, and develop a new Bayesian variable selection algorithm that explicitly enforces generic identifiability conditions and monotonicity of item response functions to ensure valid posterior inference. We demonstrate the empirical performance of our algorithm on simulation studies and several real educational testing datasets. In the second part of the thesis, we consider admixture models. The most widely known admixture models are topic models, which model each document by a convex combination of a set of word-frequency vectors, known as topics. Although identifying such latent topics is of primary interest in many applications, it is well-known that the topic parameters are not identifiable. Prior work addressed this issue by imposing stringent conditions on the existence of certain anchor words or on higher order statistics of the data. We derive a new set of identifiability conditions from convex geometry, which are weaker than the conditions from prior work. In addition, we study estimation consistency and establish the convergence rate for topic model with identifiable parameters. We also provide an EM algorithm for computation and demonstrate its competitive performance through simulation studies and real applications.","abstract_html":"Despite the fact that latent class models have been widely applied and appeared to perform well in various applications, it is well known that inference for latent class models is challenging due to the potential nonidentifiability of the underlying model parameters. In other words, multiple or infinite many sets of parameters of a latent class model could correspond to the same data generating process, i.e., are not statistically differentiable. This thesis studies the identifiability issue of some latent class models arising from real applications in cognitive diagnosis and topic modelings. The main contribution of this thesis include identifiability conditions that are weaker than the existing ones in the current literature, computational tools such as MCMC algorithms and EM algorithms that return estimates and uncertainty quantification of the underlying model parameters, and asymptotic analysis such as consistency study on the proposed estimation procedures for identifiable models. In the first part of the thesis, we consider Cognitive Diagnostic Models (CDMs). CDMs are latent variable models developed to infer latent skills, knowledge, or personalities that underlie responses to educational, psychological, and social science tests and measures. We derive a new set of sufficient conditions for generic identifiability of CDM parameters. An important contribution for practice is that our new generic identifiability conditions are more likely to be satisfied in empirical applications than existing conditions that ensure strict identifiability. For computation, we formulate learning the underlying latent structure as a variable selection problem, and develop a new Bayesian variable selection algorithm that explicitly enforces generic identifiability conditions and monotonicity of item response functions to ensure valid posterior inference. We demonstrate the empirical performance of our algorithm on simulation studies and several real educational testing datasets. In the second part of the thesis, we consider admixture models. The most widely known admixture models are topic models, which model each document by a convex combination of a set of word-frequency vectors, known as topics. Although identifying such latent topics is of primary interest in many applications, it is well-known that the topic parameters are not identifiable. Prior work addressed this issue by imposing stringent conditions on the existence of certain anchor words or on higher order statistics of the data. We derive a new set of identifiability conditions from convex geometry, which are weaker than the conditions from prior work. In addition, we study estimation consistency and establish the convergence rate for topic model with identifiable parameters. We also provide an EM algorithm for computation and demonstrate its competitive performance through simulation studies and real applications.","abstract_has_math":false,"creators":["Chen, Yinyin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Liang, Feng","Yang, Yun","Culpepper, Steven Andrew","Narisetty, Naveen Naidu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-08-27T00:50:17Z","date_published":"2020-08-27T00:50:17Z","updated_at":"2026-07-22T22:24:48Z","subjects":["Identifiability","Generic identifiability","Latent class models","Cognitive diagnostic models","Topic models"],"languages":["en"],"rights":["Copyright 2020 Yinyin Chen"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/108306","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Liang, Feng","Yang, Yun","Culpepper, Steven Andrew","Narisetty, Naveen Naidu"]},{"key":"dc:creator","label":"Author","values":["Chen, Yinyin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-08-27T00:50:17Z","2022-08-27T00:51:40Z","2020-05-07","2020-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Identifiability","Generic identifiability","Latent class models","Cognitive diagnostic models","Topic models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Yinyin Chen"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/108306"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Despite the fact that latent class models have been widely applied and appeared to perform well in various applications, it is well known that inference for latent class models is challenging due to the potential nonidentifiability of the underlying model parameters. In other words, multiple or infinite many sets of parameters of a latent class model could correspond to the same data generating process, i.e., are not statistically differentiable. This thesis studies the identifiability issue of some latent class models arising from real applications in cognitive diagnosis and topic modelings. The main contribution of this thesis include identifiability conditions that are weaker than the existing ones in the current literature, computational tools such as MCMC algorithms and EM algorithms that return estimates and uncertainty quantification of the underlying model parameters, and asymptotic analysis such as consistency study on the proposed estimation procedures for identifiable models. In the first part of the thesis, we consider Cognitive Diagnostic Models (CDMs). CDMs are latent variable models developed to infer latent skills, knowledge, or personalities that underlie responses to educational, psychological, and social science tests and measures. We derive a new set of sufficient conditions for generic identifiability of CDM parameters. An important contribution for practice is that our new generic identifiability conditions are more likely to be satisfied in empirical applications than existing conditions that ensure strict identifiability. For computation, we formulate learning the underlying latent structure as a variable selection problem, and develop a new Bayesian variable selection algorithm that explicitly enforces generic identifiability conditions and monotonicity of item response functions to ensure valid posterior inference. We demonstrate the empirical performance of our algorithm on simulation studies and several real educational testing datasets. In the second part of the thesis, we consider admixture models. The most widely known admixture models are topic models, which model each document by a convex combination of a set of word-frequency vectors, known as topics. Although identifying such latent topics is of primary interest in many applications, it is well-known that the topic parameters are not identifiable. Prior work addressed this issue by imposing stringent conditions on the existence of certain anchor words or on higher order statistics of the data. We derive a new set of identifiability conditions from convex geometry, which are weaker than the conditions from prior work. In addition, we study estimation consistency and establish the convergence rate for topic model with identifiable parameters. We also provide an EM algorithm for computation and demonstrate its competitive performance through simulation studies and real applications.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2022-05-01","The student, Yinyin Chen, accepted the attached license on 2020-05-04 at 12:18.","The student, Yinyin Chen, submitted this Dissertation for approval on 2020-05-05 at 00:58.","This Dissertation was approved for publication on 2020-05-07 at 16:59.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15190 on 2020-08-25 at 17:42:45","Made available in DSpace on 2020-08-27T00:50:17Z (GMT). No. of bitstreams: 2 CHEN-DISSERTATION-2020.pdf: 5133742 bytes, checksum: 9375e806034f93a1e2611906b780aa35 (MD5) LICENSE.txt: 4208 bytes, checksum: c05372c3655f1f81aba77be4d60471c2 (MD5) Previous issue date: 2020-05-07","Embargo set by: Seth Robbins for item 115920 Lift date: 2022-08-27T00:50:22Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 115920 Lift date: 2022-08-27T00:51:40Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Identifiability for latent class models"]}]}],"canonical_facts":{"dc:contributor":["Liang, Feng","Yang, Yun","Culpepper, Steven Andrew","Narisetty, Naveen Naidu"],"dc:creator":["Chen, Yinyin"],"dc:date":["2020-08-27T00:50:17Z","2022-08-27T00:51:40Z","2020-05-07","2020-05"],"dc:description":["Despite the fact that latent class models have been widely applied and appeared to perform well in various applications, it is well known that inference for latent class models is challenging due to the potential nonidentifiability of the underlying model parameters. In other words, multiple or infinite many sets of parameters of a latent class model could correspond to the same data generating process, i.e., are not statistically differentiable. This thesis studies the identifiability issue of some latent class models arising from real applications in cognitive diagnosis and topic modelings. The main contribution of this thesis include identifiability conditions that are weaker than the existing ones in the current literature, computational tools such as MCMC algorithms and EM algorithms that return estimates and uncertainty quantification of the underlying model parameters, and asymptotic analysis such as consistency study on the proposed estimation procedures for identifiable models. In the first part of the thesis, we consider Cognitive Diagnostic Models (CDMs). CDMs are latent variable models developed to infer latent skills, knowledge, or personalities that underlie responses to educational, psychological, and social science tests and measures. We derive a new set of sufficient conditions for generic identifiability of CDM parameters. An important contribution for practice is that our new generic identifiability conditions are more likely to be satisfied in empirical applications than existing conditions that ensure strict identifiability. For computation, we formulate learning the underlying latent structure as a variable selection problem, and develop a new Bayesian variable selection algorithm that explicitly enforces generic identifiability conditions and monotonicity of item response functions to ensure valid posterior inference. We demonstrate the empirical performance of our algorithm on simulation studies and several real educational testing datasets. In the second part of the thesis, we consider admixture models. The most widely known admixture models are topic models, which model each document by a convex combination of a set of word-frequency vectors, known as topics. Although identifying such latent topics is of primary interest in many applications, it is well-known that the topic parameters are not identifiable. Prior work addressed this issue by imposing stringent conditions on the existence of certain anchor words or on higher order statistics of the data. We derive a new set of identifiability conditions from convex geometry, which are weaker than the conditions from prior work. In addition, we study estimation consistency and establish the convergence rate for topic model with identifiable parameters. We also provide an EM algorithm for computation and demonstrate its competitive performance through simulation studies and real applications.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2022-05-01","The student, Yinyin Chen, accepted the attached license on 2020-05-04 at 12:18.","The student, Yinyin Chen, submitted this Dissertation for approval on 2020-05-05 at 00:58.","This Dissertation was approved for publication on 2020-05-07 at 16:59.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15190 on 2020-08-25 at 17:42:45","Made available in DSpace on 2020-08-27T00:50:17Z (GMT). No. of bitstreams: 2 CHEN-DISSERTATION-2020.pdf: 5133742 bytes, checksum: 9375e806034f93a1e2611906b780aa35 (MD5) LICENSE.txt: 4208 bytes, checksum: c05372c3655f1f81aba77be4d60471c2 (MD5) Previous issue date: 2020-05-07","Embargo set by: Seth Robbins for item 115920 Lift date: 2022-08-27T00:50:22Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 115920 Lift date: 2022-08-27T00:51:40Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/108306"],"dc:language":["en"],"dc:rights":["Copyright 2020 Yinyin Chen"],"dc:subject":["Identifiability","Generic identifiability","Latent class models","Cognitive diagnostic models","Topic models"],"dc:title":["Identifiability for latent class models"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:48Z"}