Cal Poly
Analysis and Usage of Natural Language Features in Success Prediction of Legislative Testimonies
Abstract
dc:description.abstract<p>Committee meetings are a fundamental part of the legislative process in which<br />constituents, lobbyists, and legislators alike can speak on proposed bills at the<br />local and state level. Oftentimes, unspoken “rules” or standards are at play in<br />political processes that can influence the trajectory of a bill, leaving constituents<br />without a political background at an inherent disadvantage when engaging with<br />the legislative process. The work done in this thesis aims to explore the extent to<br />which the language and phraseology of a general public testimony can influence a<br />vote, and examine how this information can be used to promote civic engagement.</p> <p>The Digital Democracy database contains digital records for over 40,000 real<br />testimonies by non-legislator public persons presented at California Legislature<br />committee meetings 2015-2018, along with the speakers’ desired vote outcome<br />and individual legislator votes in that discussion. With this data, we conduct a<br />linguistic analysis that is then leveraged by the Constituent phraseology Analysis<br />Tool (CPAT) to generate a user-based intelligent statistical comparison between<br />a proposed testimony and language patterns that have previously been successful.</p> <p>The following questions are at the core of this research: Which (if any) lan-<br />guage features are correlated with persuasive success in a legislative context?<br />Does the committee’s topic of discussion impact the language features that can<br />lend to a testimony’s success? Can mirroring a legislator’s speech patterns change<br />the probability of the vote going your way? How can this information be used to<br />level the playing field for constituents who want their voices heard?</p> <p>Given the 33 linguistic features developed in this research, supervised classifi-<br />cation models were able to predict testimonial success with up to 85.1% accuracy,<br />indicating that the new features had a significant impact on the prediction of<br />success. Adding these features to the 16 baseline linguistic features developed<br />in Gundala’s [18] research improved the prediction accuracy by up to 2.6%. We<br />also found that balancing the dataset of testimonies drastically impacted the<br />prediction performance metrics, with 93% accuracy achieved for the imbalanced<br />dataset and 60% accuracy after balancing. The Constituent Phraseology Analysis<br />Tool showed promise in the generation of linguistic analysis based on previously<br />successful language patterns, but requires further development before achieving<br />true usability. Additionally, predicting success based on linguistic similarity to a<br />legislator on the committee produced contradictory results. Experiments yielded<br />a 4% increase in predictive accuracy when adding comparative language features<br />to the feature set, but further experimentation with weight distributions revealed<br />only marginal impacts from comparative features.</p>
Degree
thesis:*- Name thesis:degree_name
- MS in Computer Science
- Discipline thesis:degree_discipline
- Computer Science
- Year dc:date.available
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Cossoul, Marine
- Contributors dc:contributor
-
- Foaad Khosmood
- Computer Science
- College of Engineering
Subjects
dc:subject × 7Identifiers
dc:identifier.*- Identifier
- 10.15368/theses.2023.4
- OAI identifier oai:identifier
- oai:digitalcommons.calpoly.edu:theses-4224