{"id":{"repo_id":"calpoly","oai_identifier":"oai:digitalcommons.calpoly.edu:theses-4224"},"canonical_url":"https://search.dev.ndltd.org/etd/calpoly/oai:digitalcommons.calpoly.edu:theses-4224","repository":{"repo_id":"calpoly","name":"Cal Poly","base_url":"https://digitalcommons.calpoly.edu/do/oai/"},"display":{"title":"Analysis and Usage of Natural Language Features in Success Prediction of Legislative Testimonies","abstract":"<p>Committee meetings are a fundamental part of the legislative process in which<br />constituents, lobbyists, and legislators alike can speak on proposed bills at the<br />local and state level. Oftentimes, unspoken “rules” or standards are at play in<br />political processes that can influence the trajectory of a bill, leaving constituents<br />without a political background at an inherent disadvantage when engaging with<br />the legislative process. The work done in this thesis aims to explore the extent to<br />which the language and phraseology of a general public testimony can influence a<br />vote, and examine how this information can be used to promote civic engagement.</p> <p>The Digital Democracy database contains digital records for over 40,000 real<br />testimonies by non-legislator public persons presented at California Legislature<br />committee meetings 2015-2018, along with the speakers’ desired vote outcome<br />and individual legislator votes in that discussion. With this data, we conduct a<br />linguistic analysis that is then leveraged by the Constituent phraseology Analysis<br />Tool (CPAT) to generate a user-based intelligent statistical comparison between<br />a proposed testimony and language patterns that have previously been successful.</p> <p>The following questions are at the core of this research: Which (if any) lan-<br />guage features are correlated with persuasive success in a legislative context?<br />Does the committee’s topic of discussion impact the language features that can<br />lend to a testimony’s success? Can mirroring a legislator’s speech patterns change<br />the probability of the vote going your way? How can this information be used to<br />level the playing field for constituents who want their voices heard?</p> <p>Given the 33 linguistic features developed in this research, supervised classifi-<br />cation models were able to predict testimonial success with up to 85.1% accuracy,<br />indicating that the new features had a significant impact on the prediction of<br />success. Adding these features to the 16 baseline linguistic features developed<br />in Gundala’s [18] research improved the prediction accuracy by up to 2.6%. We<br />also found that balancing the dataset of testimonies drastically impacted the<br />prediction performance metrics, with 93% accuracy achieved for the imbalanced<br />dataset and 60% accuracy after balancing. The Constituent Phraseology Analysis<br />Tool showed promise in the generation of linguistic analysis based on previously<br />successful language patterns, but requires further development before achieving<br />true usability. Additionally, predicting success based on linguistic similarity to a<br />legislator on the committee produced contradictory results. Experiments yielded<br />a 4% increase in predictive accuracy when adding comparative language features<br />to the feature set, but further experimentation with weight distributions revealed<br />only marginal impacts from comparative features.</p>","abstract_html":"&lt;p&gt;Committee meetings are a fundamental part of the legislative process in which&lt;br /&gt;constituents, lobbyists, and legislators alike can speak on proposed bills at the&lt;br /&gt;local and state level. Oftentimes, unspoken “rules” or standards are at play in&lt;br /&gt;political processes that can influence the trajectory of a bill, leaving constituents&lt;br /&gt;without a political background at an inherent disadvantage when engaging with&lt;br /&gt;the legislative process. The work done in this thesis aims to explore the extent to&lt;br /&gt;which the language and phraseology of a general public testimony can influence a&lt;br /&gt;vote, and examine how this information can be used to promote civic engagement.&lt;/p&gt; &lt;p&gt;The Digital Democracy database contains digital records for over 40,000 real&lt;br /&gt;testimonies by non-legislator public persons presented at California Legislature&lt;br /&gt;committee meetings 2015-2018, along with the speakers’ desired vote outcome&lt;br /&gt;and individual legislator votes in that discussion. With this data, we conduct a&lt;br /&gt;linguistic analysis that is then leveraged by the Constituent phraseology Analysis&lt;br /&gt;Tool (CPAT) to generate a user-based intelligent statistical comparison between&lt;br /&gt;a proposed testimony and language patterns that have previously been successful.&lt;/p&gt; &lt;p&gt;The following questions are at the core of this research: Which (if any) lan-&lt;br /&gt;guage features are correlated with persuasive success in a legislative context?&lt;br /&gt;Does the committee’s topic of discussion impact the language features that can&lt;br /&gt;lend to a testimony’s success? Can mirroring a legislator’s speech patterns change&lt;br /&gt;the probability of the vote going your way? How can this information be used to&lt;br /&gt;level the playing field for constituents who want their voices heard?&lt;/p&gt; &lt;p&gt;Given the 33 linguistic features developed in this research, supervised classifi-&lt;br /&gt;cation models were able to predict testimonial success with up to 85.1% accuracy,&lt;br /&gt;indicating that the new features had a significant impact on the prediction of&lt;br /&gt;success. Adding these features to the 16 baseline linguistic features developed&lt;br /&gt;in Gundala’s [18] research improved the prediction accuracy by up to 2.6%. We&lt;br /&gt;also found that balancing the dataset of testimonies drastically impacted the&lt;br /&gt;prediction performance metrics, with 93% accuracy achieved for the imbalanced&lt;br /&gt;dataset and 60% accuracy after balancing. The Constituent Phraseology Analysis&lt;br /&gt;Tool showed promise in the generation of linguistic analysis based on previously&lt;br /&gt;successful language patterns, but requires further development before achieving&lt;br /&gt;true usability. Additionally, predicting success based on linguistic similarity to a&lt;br /&gt;legislator on the committee produced contradictory results. Experiments yielded&lt;br /&gt;a 4% increase in predictive accuracy when adding comparative language features&lt;br /&gt;to the feature set, but further experimentation with weight distributions revealed&lt;br /&gt;only marginal impacts from comparative features.&lt;/p&gt;","abstract_has_math":false,"creators":["Cossoul, Marine"],"institution":null,"degree_name":"MS in Computer Science","degree_level":null,"degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Foaad Khosmood","Computer Science","College of Engineering"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-03-01T08:00:00Z","date_published":"2023-03-01T08:00:00Z","updated_at":"2026-07-24T01:32:13Z","subjects":["Digital Democracy","Natural Language Processing","Machine Learning","Statistical Analysis","Legislative Language","Computer Engineering","Engineering"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["10.15368/theses.2023.4"],"render_values":[{"text":"10.15368/theses.2023.4","href":"https://doi.org/10.15368/theses.2023.4","code":true}]}]},"links":{"outbound_url":"https://digitalcommons.calpoly.edu/theses/2651","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Foaad Khosmood","Computer Science","College of Engineering"]},{"key":"dc:creator","label":"Author","values":["Cossoul, Marine"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2023-03-24T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_name","label":"Degree Name","values":["MS in Computer Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Digital Democracy","Natural Language Processing","Machine Learning","Statistical Analysis","Legislative Language","Computer Engineering","Engineering"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.calpoly.edu/theses/2651","10.15368/theses.2023.4"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Committee meetings are a fundamental part of the legislative process in which<br />constituents, lobbyists, and legislators alike can speak on proposed bills at the<br />local and state level. Oftentimes, unspoken “rules” or standards are at play in<br />political processes that can influence the trajectory of a bill, leaving constituents<br />without a political background at an inherent disadvantage when engaging with<br />the legislative process. The work done in this thesis aims to explore the extent to<br />which the language and phraseology of a general public testimony can influence a<br />vote, and examine how this information can be used to promote civic engagement.</p> <p>The Digital Democracy database contains digital records for over 40,000 real<br />testimonies by non-legislator public persons presented at California Legislature<br />committee meetings 2015-2018, along with the speakers’ desired vote outcome<br />and individual legislator votes in that discussion. With this data, we conduct a<br />linguistic analysis that is then leveraged by the Constituent phraseology Analysis<br />Tool (CPAT) to generate a user-based intelligent statistical comparison between<br />a proposed testimony and language patterns that have previously been successful.</p> <p>The following questions are at the core of this research: Which (if any) lan-<br />guage features are correlated with persuasive success in a legislative context?<br />Does the committee’s topic of discussion impact the language features that can<br />lend to a testimony’s success? Can mirroring a legislator’s speech patterns change<br />the probability of the vote going your way? How can this information be used to<br />level the playing field for constituents who want their voices heard?</p> <p>Given the 33 linguistic features developed in this research, supervised classifi-<br />cation models were able to predict testimonial success with up to 85.1% accuracy,<br />indicating that the new features had a significant impact on the prediction of<br />success. Adding these features to the 16 baseline linguistic features developed<br />in Gundala’s [18] research improved the prediction accuracy by up to 2.6%. We<br />also found that balancing the dataset of testimonies drastically impacted the<br />prediction performance metrics, with 93% accuracy achieved for the imbalanced<br />dataset and 60% accuracy after balancing. The Constituent Phraseology Analysis<br />Tool showed promise in the generation of linguistic analysis based on previously<br />successful language patterns, but requires further development before achieving<br />true usability. Additionally, predicting success based on linguistic similarity to a<br />legislator on the committee produced contradictory results. Experiments yielded<br />a 4% increase in predictive accuracy when adding comparative language features<br />to the feature set, but further experimentation with weight distributions revealed<br />only marginal impacts from comparative features.</p>"]},{"key":"dc:title","label":"Title","values":["Analysis and Usage of Natural Language Features in Success Prediction of Legislative Testimonies"]}]}],"canonical_facts":{"dc:contributor":["Foaad Khosmood","Computer Science","College of Engineering"],"dc:creator":["Cossoul, Marine"],"dc:date.available":["2023-03-24T07:00:00Z"],"dc:description.abstract":["<p>Committee meetings are a fundamental part of the legislative process in which<br />constituents, lobbyists, and legislators alike can speak on proposed bills at the<br />local and state level. Oftentimes, unspoken “rules” or standards are at play in<br />political processes that can influence the trajectory of a bill, leaving constituents<br />without a political background at an inherent disadvantage when engaging with<br />the legislative process. The work done in this thesis aims to explore the extent to<br />which the language and phraseology of a general public testimony can influence a<br />vote, and examine how this information can be used to promote civic engagement.</p> <p>The Digital Democracy database contains digital records for over 40,000 real<br />testimonies by non-legislator public persons presented at California Legislature<br />committee meetings 2015-2018, along with the speakers’ desired vote outcome<br />and individual legislator votes in that discussion. With this data, we conduct a<br />linguistic analysis that is then leveraged by the Constituent phraseology Analysis<br />Tool (CPAT) to generate a user-based intelligent statistical comparison between<br />a proposed testimony and language patterns that have previously been successful.</p> <p>The following questions are at the core of this research: Which (if any) lan-<br />guage features are correlated with persuasive success in a legislative context?<br />Does the committee’s topic of discussion impact the language features that can<br />lend to a testimony’s success? Can mirroring a legislator’s speech patterns change<br />the probability of the vote going your way? How can this information be used to<br />level the playing field for constituents who want their voices heard?</p> <p>Given the 33 linguistic features developed in this research, supervised classifi-<br />cation models were able to predict testimonial success with up to 85.1% accuracy,<br />indicating that the new features had a significant impact on the prediction of<br />success. Adding these features to the 16 baseline linguistic features developed<br />in Gundala’s [18] research improved the prediction accuracy by up to 2.6%. We<br />also found that balancing the dataset of testimonies drastically impacted the<br />prediction performance metrics, with 93% accuracy achieved for the imbalanced<br />dataset and 60% accuracy after balancing. The Constituent Phraseology Analysis<br />Tool showed promise in the generation of linguistic analysis based on previously<br />successful language patterns, but requires further development before achieving<br />true usability. Additionally, predicting success based on linguistic similarity to a<br />legislator on the committee produced contradictory results. Experiments yielded<br />a 4% increase in predictive accuracy when adding comparative language features<br />to the feature set, but further experimentation with weight distributions revealed<br />only marginal impacts from comparative features.</p>"],"dc:identifier":["https://digitalcommons.calpoly.edu/theses/2651","10.15368/theses.2023.4"],"dc:subject":["Digital Democracy","Natural Language Processing","Machine Learning","Statistical Analysis","Legislative Language","Computer Engineering","Engineering"],"dc:title":["Analysis and Usage of Natural Language Features in Success Prediction of Legislative Testimonies"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_name":["MS in Computer Science"]},"updated_at":"2026-07-24T01:32:13Z"}