{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105645"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105645","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Predicting controlled vocabulary based on text and citations: Case studies in medical subject headings in MEDLINE and patents","abstract":"This dissertation makes three contributions in the area of controlled vocabulary prediction of Medical Subject Headings. The first contribution is a new partial matching measure based on distributional semantics. The second contribution is a probabilistic model based on text similarity and citations. The third contribution is a case study of cross-domain vocabulary prediction in US Patents. Medical subject headings (MeSH) are an important life sciences controlled vocabulary. They are an ideal ground to study controlled vocabulary prediction due to their complexity, hierarchical nature, and practical significance. The dissertation begins with an updated analysis of human indexing consistency in MEDLINE. This study demonstrates the need for partial matching measures to account for indexing variability. Here, I develop four measures combining the MeSH hierarchy and contextual similarity. These measures provide several new tools for evaluating and diagnosing controlled vocabulary models. Next, a generalized predictive model is introduced. This model uses citations and abstract similarity as inputs to a hybrid KNN classifier. Citations and abstracts are found to be complimentary in that they reliably produce unique and relevant candidate terms. Finally, the predictive model is applied to a corpus of approximately 65,000 biomedical US patents. This case study explores differences in the vocabulary of MEDLINE and patents, as well as the prospect for MeSH prediction to open new scholarly opportunities in economics and health policy research.","abstract_html":"This dissertation makes three contributions in the area of controlled vocabulary prediction of Medical Subject Headings. The first contribution is a new partial matching measure based on distributional semantics. The second contribution is a probabilistic model based on text similarity and citations. The third contribution is a case study of cross-domain vocabulary prediction in US Patents. Medical subject headings (MeSH) are an important life sciences controlled vocabulary. They are an ideal ground to study controlled vocabulary prediction due to their complexity, hierarchical nature, and practical significance. The dissertation begins with an updated analysis of human indexing consistency in MEDLINE. This study demonstrates the need for partial matching measures to account for indexing variability. Here, I develop four measures combining the MeSH hierarchy and contextual similarity. These measures provide several new tools for evaluating and diagnosing controlled vocabulary models. Next, a generalized predictive model is introduced. This model uses citations and abstract similarity as inputs to a hybrid KNN classifier. Citations and abstracts are found to be complimentary in that they reliably produce unique and relevant candidate terms. Finally, the predictive model is applied to a corpus of approximately 65,000 biomedical US patents. This case study explores differences in the vocabulary of MEDLINE and patents, as well as the prospect for MeSH prediction to open new scholarly opportunities in economics and health policy research.","abstract_has_math":false,"creators":["Kehoe, Adam K."],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Library & Information Science","degree_department":null,"school":null,"contributors":["Torvik, Vetle I","Smalheiser, Neil R","Dubin, David S","Ludäscher, Bertram","Downie, John S"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-11-26T20:33:52Z","date_published":"2019-11-26T20:33:52Z","updated_at":"2026-07-22T22:24:44Z","subjects":["Controlled vocabulary","Medical Subject Headings","Controlled Vocabulary Prediction"],"languages":["en"],"rights":["Copyright 2019 Adam Kehoe"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105645","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Torvik, Vetle I","Smalheiser, Neil R","Dubin, David S","Ludäscher, Bertram","Downie, John S"]},{"key":"dc:creator","label":"Author","values":["Kehoe, Adam K."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-11-26T20:33:52Z","2019-07-09","2019-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Library & Information Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Controlled vocabulary","Medical Subject Headings","Controlled Vocabulary Prediction"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Adam Kehoe"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105645"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This dissertation makes three contributions in the area of controlled vocabulary prediction of Medical Subject Headings. The first contribution is a new partial matching measure based on distributional semantics. The second contribution is a probabilistic model based on text similarity and citations. The third contribution is a case study of cross-domain vocabulary prediction in US Patents. Medical subject headings (MeSH) are an important life sciences controlled vocabulary. They are an ideal ground to study controlled vocabulary prediction due to their complexity, hierarchical nature, and practical significance. The dissertation begins with an updated analysis of human indexing consistency in MEDLINE. This study demonstrates the need for partial matching measures to account for indexing variability. Here, I develop four measures combining the MeSH hierarchy and contextual similarity. These measures provide several new tools for evaluating and diagnosing controlled vocabulary models. Next, a generalized predictive model is introduced. This model uses citations and abstract similarity as inputs to a hybrid KNN classifier. Citations and abstracts are found to be complimentary in that they reliably produce unique and relevant candidate terms. Finally, the predictive model is applied to a corpus of approximately 65,000 biomedical US patents. This case study explores differences in the vocabulary of MEDLINE and patents, as well as the prospect for MeSH prediction to open new scholarly opportunities in economics and health policy research.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-11-26 without embargo terms","The student, Adam Kehoe, accepted the attached license on 2019-07-03 at 15:13.","The student, Adam Kehoe, submitted this Dissertation for approval on 2019-07-03 at 15:24.","This Dissertation was approved for publication on 2019-07-09 at 10:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14156 on 2019-11-26 at 12:51:06","Made available in DSpace on 2019-11-26T20:33:52Z (GMT). No. of bitstreams: 3 KEHOE-DISSERTATION-2019.pdf: 4065167 bytes, checksum: 5e299d42de2340f7434a260264f52a94 (MD5) LICENSE.txt: 4207 bytes, checksum: c162c22cd6cdbe44e8dc219677c3071e (MD5) PROQUEST_LICENSE.txt: 4553 bytes, checksum: 5f862f3fbd5b2c2fdf22b525acee33ff (MD5) Previous issue date: 2019-07-09"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Predicting controlled vocabulary based on text and citations: Case studies in medical subject headings in MEDLINE and patents"]}]}],"canonical_facts":{"dc:contributor":["Torvik, Vetle I","Smalheiser, Neil R","Dubin, David S","Ludäscher, Bertram","Downie, John S"],"dc:creator":["Kehoe, Adam K."],"dc:date":["2019-11-26T20:33:52Z","2019-07-09","2019-08"],"dc:description":["This dissertation makes three contributions in the area of controlled vocabulary prediction of Medical Subject Headings. The first contribution is a new partial matching measure based on distributional semantics. The second contribution is a probabilistic model based on text similarity and citations. The third contribution is a case study of cross-domain vocabulary prediction in US Patents. Medical subject headings (MeSH) are an important life sciences controlled vocabulary. They are an ideal ground to study controlled vocabulary prediction due to their complexity, hierarchical nature, and practical significance. The dissertation begins with an updated analysis of human indexing consistency in MEDLINE. This study demonstrates the need for partial matching measures to account for indexing variability. Here, I develop four measures combining the MeSH hierarchy and contextual similarity. These measures provide several new tools for evaluating and diagnosing controlled vocabulary models. Next, a generalized predictive model is introduced. This model uses citations and abstract similarity as inputs to a hybrid KNN classifier. Citations and abstracts are found to be complimentary in that they reliably produce unique and relevant candidate terms. Finally, the predictive model is applied to a corpus of approximately 65,000 biomedical US patents. This case study explores differences in the vocabulary of MEDLINE and patents, as well as the prospect for MeSH prediction to open new scholarly opportunities in economics and health policy research.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-11-26 without embargo terms","The student, Adam Kehoe, accepted the attached license on 2019-07-03 at 15:13.","The student, Adam Kehoe, submitted this Dissertation for approval on 2019-07-03 at 15:24.","This Dissertation was approved for publication on 2019-07-09 at 10:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14156 on 2019-11-26 at 12:51:06","Made available in DSpace on 2019-11-26T20:33:52Z (GMT). No. of bitstreams: 3 KEHOE-DISSERTATION-2019.pdf: 4065167 bytes, checksum: 5e299d42de2340f7434a260264f52a94 (MD5) LICENSE.txt: 4207 bytes, checksum: c162c22cd6cdbe44e8dc219677c3071e (MD5) PROQUEST_LICENSE.txt: 4553 bytes, checksum: 5f862f3fbd5b2c2fdf22b525acee33ff (MD5) Previous issue date: 2019-07-09"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105645"],"dc:language":["en"],"dc:rights":["Copyright 2019 Adam Kehoe"],"dc:subject":["Controlled vocabulary","Medical Subject Headings","Controlled Vocabulary Prediction"],"dc:title":["Predicting controlled vocabulary based on text and citations: Case studies in medical subject headings in MEDLINE and patents"],"dc:type":["text"],"thesis:degree_discipline":["Library & Information Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:44Z"}