{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/120545"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/120545","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Learning from the experts: Measuring the policy content of legislation using machine learning","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2025-05-01","abstract_has_math":false,"creators":["Dee, Ethan Adam"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Economics","degree_department":null,"school":null,"contributors":["Krasa, Stefan","Bernhardt, Mark D","Winters, Matthew S","Garlick, Alex"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-05","date_published":"2023-05","updated_at":"2026-07-22T22:24:57Z","subjects":["Machine Learning","Natural Language Processing","Text Classification","Supervised Learning","Text As Data","Legislation","Issue Attention","Policy Agenda","Transformer","Word Embedding","Congress","Legislature"],"languages":["en","eng"],"rights":["Copyright 2023 Ethan Dee"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/120545","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Krasa, Stefan","Bernhardt, Mark D","Winters, Matthew S","Garlick, Alex"]},{"key":"dc:creator","label":"Author","values":["Dee, Ethan Adam"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-05","2023-04-25"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Economics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","Natural Language Processing","Text Classification","Supervised Learning","Text As Data","Legislation","Issue Attention","Policy Agenda","Transformer","Word Embedding","Congress","Legislature"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Ethan Dee"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/120545"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-05-01","The student, Ethan Dee, accepted the attached license on 2023-04-23 at 22:48.","The student, Ethan Dee, submitted this Dissertation for approval on 2023-04-24 at 09:09.","This Dissertation was approved for publication on 2023-04-25 at 11:48.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19114 on 2023-09-01 at 17:21:27","Legislation entails a significant portion of legislative activity, both in volume and in impact, but the high-dimensional information contained within textual data does not lend itself to quantitative analysis. Knowing that a legislator has, for example, been the sponsor of 100 \"Education\" bills requires knowing how to find \"Education\" bills in the first place. Deciding what does and does not count as an \"Education\" bill is the peril of the chorus of researchers who have long pursued ways to measure the policy content of legislation in a tractable way. I present two approaches to classifying legislation into policy areas, offering both generalizable methodological insights for social scientists, as well as two new datasets which classify the universe of state and U.S. Congressional legislation from 2009-2023 into broad, comprehensive policy areas. The first approach presumes a starting point where the researcher has not coded any bills by hand. I begin by refining the \"dictionary method,\" presented in Garlick (2022), using a set of keywords chosen to represent topics, and train a supervised machine learning model to understand the context that often surrounds these keywords, to form predictions for the topics it should assign to each bill. This allows the model to generate predictions for bills which do not contain keywords and form a richer depiction of topics than simplistic keyword-to-topic assignment rules allow. The second approach involves training a supervised machine learning model to emulate the decision function that generated extant hand-coded data. I use the Congressional Bills Project (Adler and Wilkerson 2015) and Pennsylvania Policy Database Project (McLaughlin et al. 2010) hand-coded data to train the model, offering methodological insights regarding how to handle \"noise\" in hand-coder data (such as two hand-coders coding the same bill differently), and present several ways to validate the model's output and generate out-of-sample predictions. Further, I demonstrate how the downstream researcher can use a supervised machine learning model as a companion to their hand-coders, as it offers a unique and compelling perspective on their corpus which can improve the quality of the hand-coded data."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Learning from the experts: Measuring the policy content of legislation using machine learning"]}]}],"canonical_facts":{"dc:contributor":["Krasa, Stefan","Bernhardt, Mark D","Winters, Matthew S","Garlick, Alex"],"dc:creator":["Dee, Ethan Adam"],"dc:date":["2023-05","2023-04-25"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-05-01","The student, Ethan Dee, accepted the attached license on 2023-04-23 at 22:48.","The student, Ethan Dee, submitted this Dissertation for approval on 2023-04-24 at 09:09.","This Dissertation was approved for publication on 2023-04-25 at 11:48.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19114 on 2023-09-01 at 17:21:27","Legislation entails a significant portion of legislative activity, both in volume and in impact, but the high-dimensional information contained within textual data does not lend itself to quantitative analysis. Knowing that a legislator has, for example, been the sponsor of 100 \"Education\" bills requires knowing how to find \"Education\" bills in the first place. Deciding what does and does not count as an \"Education\" bill is the peril of the chorus of researchers who have long pursued ways to measure the policy content of legislation in a tractable way. I present two approaches to classifying legislation into policy areas, offering both generalizable methodological insights for social scientists, as well as two new datasets which classify the universe of state and U.S. Congressional legislation from 2009-2023 into broad, comprehensive policy areas. The first approach presumes a starting point where the researcher has not coded any bills by hand. I begin by refining the \"dictionary method,\" presented in Garlick (2022), using a set of keywords chosen to represent topics, and train a supervised machine learning model to understand the context that often surrounds these keywords, to form predictions for the topics it should assign to each bill. This allows the model to generate predictions for bills which do not contain keywords and form a richer depiction of topics than simplistic keyword-to-topic assignment rules allow. The second approach involves training a supervised machine learning model to emulate the decision function that generated extant hand-coded data. I use the Congressional Bills Project (Adler and Wilkerson 2015) and Pennsylvania Policy Database Project (McLaughlin et al. 2010) hand-coded data to train the model, offering methodological insights regarding how to handle \"noise\" in hand-coder data (such as two hand-coders coding the same bill differently), and present several ways to validate the model's output and generate out-of-sample predictions. Further, I demonstrate how the downstream researcher can use a supervised machine learning model as a companion to their hand-coders, as it offers a unique and compelling perspective on their corpus which can improve the quality of the hand-coded data."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/120545"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Ethan Dee"],"dc:subject":["Machine Learning","Natural Language Processing","Text Classification","Supervised Learning","Text As Data","Legislation","Issue Attention","Policy Agenda","Transformer","Word Embedding","Congress","Legislature"],"dc:title":["Learning from the experts: Measuring the policy content of legislation using machine learning"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Economics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}