{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/29711"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/29711","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Toward automatic model adaptation for structured domains","abstract":"In order for a machine learning effort to succeed, an appropriate model must be chosen. This is a difficult task in which one must balance flexibility, so that the model can capture the complexities of the domain, and simplicity, so that the model does not overfit to irrelevant characteristics of the training data. The optimal model is not only a function of the task to which it is applied, but also the amount of training data available. Copious training data can justify a complex model that includes many of the ``true'' domain interaction. But when training data is limited, additional simplifications are necessary. Traditional model selection techniques, that require fitting each of a number of hypothesized models to the training data before selecting one, apply in theory, but are not feasible when the number of possible models is large. In this thesis, we describe steps in a new direction for automatically adapting model flexibility. Our approach leverages prior knowledge of two forms: 1) Qualitative knowledge statements, which describe positive and negative relationships between domain variables, and 2) Structural metadata, which provide categorical assignments for each training instance. In our approach, this prior knowledge is used to implicitly construct a large space of alternative well-formed models. A model adaptation procedure then utilizes the training data to conduct a directed search through the space of possible models. The search requires that relatively few models be fit to the data. Thus, the search is efficient and the risk of overfitting in the model selection process is minimized. We demonstrate our approaches on a variety of machine learning tasks, including military airspace safety prediction, planning operator construction, sports prediction, and document sentiment analysis.","abstract_html":"In order for a machine learning effort to succeed, an appropriate model must be chosen. This is a difficult task in which one must balance flexibility, so that the model can capture the complexities of the domain, and simplicity, so that the model does not overfit to irrelevant characteristics of the training data. The optimal model is not only a function of the task to which it is applied, but also the amount of training data available. Copious training data can justify a complex model that includes many of the ``true&#x27;&#x27; domain interaction. But when training data is limited, additional simplifications are necessary. Traditional model selection techniques, that require fitting each of a number of hypothesized models to the training data before selecting one, apply in theory, but are not feasible when the number of possible models is large. In this thesis, we describe steps in a new direction for automatically adapting model flexibility. Our approach leverages prior knowledge of two forms: 1) Qualitative knowledge statements, which describe positive and negative relationships between domain variables, and 2) Structural metadata, which provide categorical assignments for each training instance. In our approach, this prior knowledge is used to implicitly construct a large space of alternative well-formed models. A model adaptation procedure then utilizes the training data to conduct a directed search through the space of possible models. The search requires that relatively few models be fit to the data. Thus, the search is efficient and the risk of overfitting in the model selection process is minimized. We demonstrate our approaches on a variety of machine learning tasks, including military airspace safety prediction, planning operator construction, sports prediction, and document sentiment analysis.","abstract_has_math":false,"creators":["Levine, Geoffrey C."],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["DeJong, Gerald F.","Roth, Dan","Forsyth, David A.","Kuter, Ugur"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-02-06T20:12:20Z","date_published":"2012-02-06T20:12:20Z","updated_at":"2026-07-22T22:25:27Z","subjects":["Artificial Intelligence","Machine Learning","Statistics","Explanation-Based Learning","Natural Language Processing"],"languages":["en"],"rights":["Copyright 2011 Geoffrey Levine"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/29711","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["DeJong, Gerald F.","Roth, Dan","Forsyth, David A.","Kuter, Ugur"]},{"key":"dc:creator","label":"Author","values":["Levine, Geoffrey C."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-02-06T20:12:20Z","2011-12"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Artificial Intelligence","Machine Learning","Statistics","Explanation-Based Learning","Natural Language Processing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2011 Geoffrey Levine"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/29711"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["In order for a machine learning effort to succeed, an appropriate model must be chosen. This is a difficult task in which one must balance flexibility, so that the model can capture the complexities of the domain, and simplicity, so that the model does not overfit to irrelevant characteristics of the training data. The optimal model is not only a function of the task to which it is applied, but also the amount of training data available. Copious training data can justify a complex model that includes many of the ``true'' domain interaction. But when training data is limited, additional simplifications are necessary. Traditional model selection techniques, that require fitting each of a number of hypothesized models to the training data before selecting one, apply in theory, but are not feasible when the number of possible models is large. In this thesis, we describe steps in a new direction for automatically adapting model flexibility. Our approach leverages prior knowledge of two forms: 1) Qualitative knowledge statements, which describe positive and negative relationships between domain variables, and 2) Structural metadata, which provide categorical assignments for each training instance. In our approach, this prior knowledge is used to implicitly construct a large space of alternative well-formed models. A model adaptation procedure then utilizes the training data to conduct a directed search through the space of possible models. The search requires that relatively few models be fit to the data. Thus, the search is efficient and the risk of overfitting in the model selection process is minimized. We demonstrate our approaches on a variety of machine learning tasks, including military airspace safety prediction, planning operator construction, sports prediction, and document sentiment analysis.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-08-31T20:54:15Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Levine_Geoffrey.zip: 1960362 bytes, checksum: ee145bfc918ec035edd7367ffbf0d8a8 (MD5) Levine_Geoffrey.pdf: 2509271 bytes, checksum: ad0143e7a1f3e56753e6a5a8c8edfcfa (MD5)","Made available in DSpace on 2012-02-06T20:12:20Z (GMT). No. of bitstreams: 3 Levine_Geoffrey.pdf: 2509271 bytes, checksum: ad0143e7a1f3e56753e6a5a8c8edfcfa (MD5) license.txt: 4063 bytes, checksum: c43e02c48d77fc4abb143a67122aedb3 (MD5) Levine_Geoffrey.zip: 1960362 bytes, checksum: ee145bfc918ec035edd7367ffbf0d8a8 (MD5)"]},{"key":"dc:title","label":"Title","values":["Toward automatic model adaptation for structured domains"]}]}],"canonical_facts":{"dc:contributor":["DeJong, Gerald F.","Roth, Dan","Forsyth, David A.","Kuter, Ugur"],"dc:creator":["Levine, Geoffrey C."],"dc:date":["2012-02-06T20:12:20Z","2011-12"],"dc:description":["In order for a machine learning effort to succeed, an appropriate model must be chosen. This is a difficult task in which one must balance flexibility, so that the model can capture the complexities of the domain, and simplicity, so that the model does not overfit to irrelevant characteristics of the training data. The optimal model is not only a function of the task to which it is applied, but also the amount of training data available. Copious training data can justify a complex model that includes many of the ``true'' domain interaction. But when training data is limited, additional simplifications are necessary. Traditional model selection techniques, that require fitting each of a number of hypothesized models to the training data before selecting one, apply in theory, but are not feasible when the number of possible models is large. In this thesis, we describe steps in a new direction for automatically adapting model flexibility. Our approach leverages prior knowledge of two forms: 1) Qualitative knowledge statements, which describe positive and negative relationships between domain variables, and 2) Structural metadata, which provide categorical assignments for each training instance. In our approach, this prior knowledge is used to implicitly construct a large space of alternative well-formed models. A model adaptation procedure then utilizes the training data to conduct a directed search through the space of possible models. The search requires that relatively few models be fit to the data. Thus, the search is efficient and the risk of overfitting in the model selection process is minimized. We demonstrate our approaches on a variety of machine learning tasks, including military airspace safety prediction, planning operator construction, sports prediction, and document sentiment analysis.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-08-31T20:54:15Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Levine_Geoffrey.zip: 1960362 bytes, checksum: ee145bfc918ec035edd7367ffbf0d8a8 (MD5) Levine_Geoffrey.pdf: 2509271 bytes, checksum: ad0143e7a1f3e56753e6a5a8c8edfcfa (MD5)","Made available in DSpace on 2012-02-06T20:12:20Z (GMT). No. of bitstreams: 3 Levine_Geoffrey.pdf: 2509271 bytes, checksum: ad0143e7a1f3e56753e6a5a8c8edfcfa (MD5) license.txt: 4063 bytes, checksum: c43e02c48d77fc4abb143a67122aedb3 (MD5) Levine_Geoffrey.zip: 1960362 bytes, checksum: ee145bfc918ec035edd7367ffbf0d8a8 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/29711"],"dc:language":["en"],"dc:rights":["Copyright 2011 Geoffrey Levine"],"dc:subject":["Artificial Intelligence","Machine Learning","Statistics","Explanation-Based Learning","Natural Language Processing"],"dc:title":["Toward automatic model adaptation for structured domains"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:27Z"}