{"id":{"repo_id":"columbus-state","oai_identifier":"oai:csuepress.columbusstate.edu:theses_dissertations-1448"},"canonical_url":"https://search.dev.ndltd.org/etd/columbus-state/oai:csuepress.columbusstate.edu:theses_dissertations-1448","repository":{"repo_id":"columbus-state","name":"Columbus State University","base_url":"https://csuepress.columbusstate.edu/do/oai/"},"display":{"title":"A Comparative Study on Feature Extraction and Classification/Clustering of Fake News and Conspiracy Theories from Twitter Data","abstract":"<p>Fake news and conspiracy theories have become largely abundant in the expanding world of social media. They predominantly affect the beliefs and thoughts of the public, resulting in chaos. They have always existed throughout the last few decades. They have been linked to prejudice, revolutions and genocide across history. They have also been known to have propelled people to reject mainstream medicines to an extent where some diseases are recurring in some parts of the world. They impose a serious impact since they are capable of spreading very fast Thus, it is very important to find suitable ways to detect fake news and conspiracy theories in social media, which requires a thorough analysis of their features. This study presents a survey on the various techniques of feature extraction and classification that can be implemented to classify and detect fake news and conspiracy theories from twitter datasets. The results indicate that the tf-idf method of feature extraction, when implemented with the svm classification algorithm, yields the highest accuracy of 99.6% in comparison to the other algorithms i.e. multinomial naive bayes, logistic regression and decision tree. The Bag of Words model yields an accuracy of 52.3% for both multinomial naive bayes and logistic regression algorithms and a lower range of accuracies for the other two algorithms i.e. svm and decision tree . TF-IDF has thus performed better than Bag of Words.</p>","abstract_html":"&lt;p&gt;Fake news and conspiracy theories have become largely abundant in the expanding world of social media. They predominantly affect the beliefs and thoughts of the public, resulting in chaos. They have always existed throughout the last few decades. They have been linked to prejudice, revolutions and genocide across history. They have also been known to have propelled people to reject mainstream medicines to an extent where some diseases are recurring in some parts of the world. They impose a serious impact since they are capable of spreading very fast Thus, it is very important to find suitable ways to detect fake news and conspiracy theories in social media, which requires a thorough analysis of their features. This study presents a survey on the various techniques of feature extraction and classification that can be implemented to classify and detect fake news and conspiracy theories from twitter datasets. The results indicate that the tf-idf method of feature extraction, when implemented with the svm classification algorithm, yields the highest accuracy of 99.6% in comparison to the other algorithms i.e. multinomial naive bayes, logistic regression and decision tree. The Bag of Words model yields an accuracy of 52.3% for both multinomial naive bayes and logistic regression algorithms and a lower range of accuracies for the other two algorithms i.e. svm and decision tree . TF-IDF has thus performed better than Bag of Words.&lt;/p&gt;","abstract_has_math":false,"creators":["Shana, Deb"],"institution":null,"degree_name":"Computer Science - Applied Computing Track","degree_level":"Thesis","degree_discipline":"TSYS School of Computer Science","degree_department":null,"school":null,"contributors":["Dr. Lydia Ray","Dr. Rania Hodhod","Dr. Lixin Wang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-10-01T07:00:00Z","date_published":"2021-10-01T07:00:00Z","updated_at":"2026-07-24T01:45:08Z","subjects":["Conspiracy Theories","Fake News","Feature Extraction","Machine Learning","Algorithims","Computer Sciences","Information Security","Numerical Analysis and Scientific Computing"],"languages":["English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://csuepress.columbusstate.edu/theses_dissertations/446","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dr. Lydia Ray","Dr. Rania Hodhod","Dr. Lixin Wang"]},{"key":"dc:creator","label":"Author","values":["Shana, Deb"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2021-10-29T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["TSYS School of Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Computer Science - Applied Computing Track"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Conspiracy Theories","Fake News","Feature Extraction","Machine Learning","Algorithims","Computer Sciences","Information Security","Numerical Analysis and Scientific Computing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://csuepress.columbusstate.edu/theses_dissertations/446"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Fake news and conspiracy theories have become largely abundant in the expanding world of social media. They predominantly affect the beliefs and thoughts of the public, resulting in chaos. They have always existed throughout the last few decades. They have been linked to prejudice, revolutions and genocide across history. They have also been known to have propelled people to reject mainstream medicines to an extent where some diseases are recurring in some parts of the world. They impose a serious impact since they are capable of spreading very fast Thus, it is very important to find suitable ways to detect fake news and conspiracy theories in social media, which requires a thorough analysis of their features. This study presents a survey on the various techniques of feature extraction and classification that can be implemented to classify and detect fake news and conspiracy theories from twitter datasets. The results indicate that the tf-idf method of feature extraction, when implemented with the svm classification algorithm, yields the highest accuracy of 99.6% in comparison to the other algorithms i.e. multinomial naive bayes, logistic regression and decision tree. The Bag of Words model yields an accuracy of 52.3% for both multinomial naive bayes and logistic regression algorithms and a lower range of accuracies for the other two algorithms i.e. svm and decision tree . TF-IDF has thus performed better than Bag of Words.</p>"]},{"key":"dc:title","label":"Title","values":["A Comparative Study on Feature Extraction and Classification/Clustering of Fake News and Conspiracy Theories from Twitter Data"]}]}],"canonical_facts":{"dc:contributor":["Dr. Lydia Ray","Dr. Rania Hodhod","Dr. Lixin Wang"],"dc:creator":["Shana, Deb"],"dc:date.available":["2021-10-29T07:00:00Z"],"dc:description.abstract":["<p>Fake news and conspiracy theories have become largely abundant in the expanding world of social media. They predominantly affect the beliefs and thoughts of the public, resulting in chaos. They have always existed throughout the last few decades. They have been linked to prejudice, revolutions and genocide across history. They have also been known to have propelled people to reject mainstream medicines to an extent where some diseases are recurring in some parts of the world. They impose a serious impact since they are capable of spreading very fast Thus, it is very important to find suitable ways to detect fake news and conspiracy theories in social media, which requires a thorough analysis of their features. This study presents a survey on the various techniques of feature extraction and classification that can be implemented to classify and detect fake news and conspiracy theories from twitter datasets. The results indicate that the tf-idf method of feature extraction, when implemented with the svm classification algorithm, yields the highest accuracy of 99.6% in comparison to the other algorithms i.e. multinomial naive bayes, logistic regression and decision tree. The Bag of Words model yields an accuracy of 52.3% for both multinomial naive bayes and logistic regression algorithms and a lower range of accuracies for the other two algorithms i.e. svm and decision tree . TF-IDF has thus performed better than Bag of Words.</p>"],"dc:identifier":["https://csuepress.columbusstate.edu/theses_dissertations/446"],"dc:language":["English"],"dc:subject":["Conspiracy Theories","Fake News","Feature Extraction","Machine Learning","Algorithims","Computer Sciences","Information Security","Numerical Analysis and Scientific Computing"],"dc:title":["A Comparative Study on Feature Extraction and Classification/Clustering of Fake News and Conspiracy Theories from Twitter Data"],"thesis:degree_discipline":["TSYS School of Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Computer Science - Applied Computing Track"]},"updated_at":"2026-07-24T01:45:08Z"}