{"id":{"repo_id":"columbus-state","oai_identifier":"oai:csuepress.columbusstate.edu:theses_dissertations-1330"},"canonical_url":"https://search.dev.ndltd.org/etd/columbus-state/oai:csuepress.columbusstate.edu:theses_dissertations-1330","repository":{"repo_id":"columbus-state","name":"Columbus State University","base_url":"https://csuepress.columbusstate.edu/do/oai/"},"display":{"title":"Automated Essay Evaluation Using Natural Language Processing and Machine Learning","abstract":"<p>The goal of automated essay evaluation is to assign grades to essays and provide feedback using computers. Automated evaluation is increasingly being used in classrooms and online exams. The aim of this project is to develop machine learning models for performing automated essay scoring and evaluate their performance. In this research, a publicly available essay data set was used to train and test the efficacy of the adopted techniques. Natural language processing techniques were used to extract features from essays in the dataset. Three different existing machine learning algorithms were used on the chosen dataset. The data was divided into two parts: training data and testing data. The inter-rater reliability and performance of these models were compared with each other and with human graders. Among the three machine learning models, the random forest performed the best in terms of agreement with human scorers as it achieved the lowest mean absolute error for the test dataset.</p>","abstract_html":"&lt;p&gt;The goal of automated essay evaluation is to assign grades to essays and provide feedback using computers. Automated evaluation is increasingly being used in classrooms and online exams. The aim of this project is to develop machine learning models for performing automated essay scoring and evaluate their performance. In this research, a publicly available essay data set was used to train and test the efficacy of the adopted techniques. Natural language processing techniques were used to extract features from essays in the dataset. Three different existing machine learning algorithms were used on the chosen dataset. The data was divided into two parts: training data and testing data. The inter-rater reliability and performance of these models were compared with each other and with human graders. Among the three machine learning models, the random forest performed the best in terms of agreement with human scorers as it achieved the lowest mean absolute error for the test dataset.&lt;/p&gt;","abstract_has_math":false,"creators":["Ghanta, Harshanthi"],"institution":null,"degree_name":"Computer Science - Applied Computing Track","degree_level":"Thesis","degree_discipline":"TSYS School of Computer Science","degree_department":null,"school":null,"contributors":["Dr. Shamim Khan","Dr. Rania Hodhod","Dr. Hyrum D. Carroll"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-01-01T08:00:00Z","date_published":"2019-01-01T08:00:00Z","updated_at":"2026-07-24T01:45:01Z","subjects":["Automated Essay Evaluation","Machine Learning","Natural Language Processing","Feature Extraction","Computer Sciences","Databases and Information Systems","Numerical Analysis and Scientific Computing"],"languages":["English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://csuepress.columbusstate.edu/theses_dissertations/327","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dr. Shamim Khan","Dr. Rania Hodhod","Dr. Hyrum D. Carroll"]},{"key":"dc:creator","label":"Author","values":["Ghanta, Harshanthi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2020-01-07T08:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["TSYS School of Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Computer Science - Applied Computing Track"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Automated Essay Evaluation","Machine Learning","Natural Language Processing","Feature Extraction","Computer Sciences","Databases and Information Systems","Numerical Analysis and Scientific Computing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://csuepress.columbusstate.edu/theses_dissertations/327"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>The goal of automated essay evaluation is to assign grades to essays and provide feedback using computers. Automated evaluation is increasingly being used in classrooms and online exams. The aim of this project is to develop machine learning models for performing automated essay scoring and evaluate their performance. In this research, a publicly available essay data set was used to train and test the efficacy of the adopted techniques. Natural language processing techniques were used to extract features from essays in the dataset. Three different existing machine learning algorithms were used on the chosen dataset. The data was divided into two parts: training data and testing data. The inter-rater reliability and performance of these models were compared with each other and with human graders. Among the three machine learning models, the random forest performed the best in terms of agreement with human scorers as it achieved the lowest mean absolute error for the test dataset.</p>"]},{"key":"dc:title","label":"Title","values":["Automated Essay Evaluation Using Natural Language Processing and Machine Learning"]}]}],"canonical_facts":{"dc:contributor":["Dr. Shamim Khan","Dr. Rania Hodhod","Dr. Hyrum D. Carroll"],"dc:creator":["Ghanta, Harshanthi"],"dc:date.available":["2020-01-07T08:00:00Z"],"dc:description.abstract":["<p>The goal of automated essay evaluation is to assign grades to essays and provide feedback using computers. Automated evaluation is increasingly being used in classrooms and online exams. The aim of this project is to develop machine learning models for performing automated essay scoring and evaluate their performance. In this research, a publicly available essay data set was used to train and test the efficacy of the adopted techniques. Natural language processing techniques were used to extract features from essays in the dataset. Three different existing machine learning algorithms were used on the chosen dataset. The data was divided into two parts: training data and testing data. The inter-rater reliability and performance of these models were compared with each other and with human graders. Among the three machine learning models, the random forest performed the best in terms of agreement with human scorers as it achieved the lowest mean absolute error for the test dataset.</p>"],"dc:identifier":["https://csuepress.columbusstate.edu/theses_dissertations/327"],"dc:language":["English"],"dc:subject":["Automated Essay Evaluation","Machine Learning","Natural Language Processing","Feature Extraction","Computer Sciences","Databases and Information Systems","Numerical Analysis and Scientific Computing"],"dc:title":["Automated Essay Evaluation Using Natural Language Processing and Machine Learning"],"thesis:degree_discipline":["TSYS School of Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Computer Science - Applied Computing Track"]},"updated_at":"2026-07-24T01:45:01Z"}