University of Illinois at Urbana-Champaign
Learning from the experts: Measuring the policy content of legislation using machine learning
Abstract
dc:descriptionLegislation entails a significant portion of legislative activity, both in volume and in impact, but the high-dimensional information contained within textual data does not lend itself to quantitative analysis. Knowing that a legislator has, for example, been the sponsor of 100 "Education" bills requires knowing how to find "Education" bills in the first place. Deciding what does and does not count as an "Education" bill is the peril of the chorus of researchers who have long pursued ways to measure the policy content of legislation in a tractable way. I present two approaches to classifying legislation into policy areas, offering both generalizable methodological insights for social scientists, as well as two new datasets which classify the universe of state and U.S. Congressional legislation from 2009-2023 into broad, comprehensive policy areas. The first approach presumes a starting point where the researcher has not coded any bills by hand. I begin by refining the "dictionary method," presented in Garlick (2022), using a set of keywords chosen to represent topics, and train a supervised machine learning model to understand the context that often surrounds these keywords, to form predictions for the topics it should assign to each bill. This allows the model to generate predictions for bills which do not contain keywords and form a richer depiction of topics than simplistic keyword-to-topic assignment rules allow. The second approach involves training a supervised machine learning model to emulate the decision function that generated extant hand-coded data. I use the Congressional Bills Project (Adler and Wilkerson 2015) and Pennsylvania Policy Database Project (McLaughlin et al. 2010) hand-coded data to train the model, offering methodological insights regarding how to handle "noise" in hand-coder data (such as two hand-coders coding the same bill differently), and present several ways to validate the model's output and generate out-of-sample predictions. Further, I demonstrate how the downstream researcher can use a supervised machine learning model as a companion to their hand-coders, as it offers a unique and compelling perspective on their corpus which can improve the quality of the hand-coded data.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Economics
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Dee, Ethan Adam
- Contributors dc:contributor
-
- Krasa, Stefan
- Bernhardt, Mark D
- Winters, Matthew S
- Garlick, Alex
Subjects
dc:subject × 12Rights
dc:rights- Statement dc:rights
-
- Copyright 2023 Ethan Dee
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/120545