{"id":{"repo_id":"duquesne","oai_identifier":"oai:dsc.duq.edu:etd-1839"},"canonical_url":"https://search.dev.ndltd.org/etd/duquesne/oai:dsc.duq.edu:etd-1839","repository":{"repo_id":"duquesne","name":"Duquesne","base_url":"https://dsc.duq.edu/do/oai/"},"display":{"title":"Authorship Attribution on the Enron Email Corpus","abstract":"In this paper I present authorship attribution on an email corpus. The source I used was the Enron Email Corpus (Cohen, 2009). By reformatting these emails, four test sets were categorized based on the length of each email: Tiny (&#8804; 99 characters), Small (100 to 500 characters), Medium (501 to 999 characters), and Large (&#8805; 1000 characters). The Java Graphical Authorship Attribution Program (JGAAP software) from our Evaluating Variations in Language Laboratory (EVL Lab) was used to perform these tests. Three analysis methods: WEKA RandomForest, WEKA SMO, and Centroid with Cosine Distance were used. Results showed that the Large test set gave the best authorship classification, followed by the Medium, then the Small and the Tiny test sets. WEKA SMO gave better authorship classification than WEKA RandomForest.","abstract_html":"In this paper I present authorship attribution on an email corpus. The source I used was the Enron Email Corpus (Cohen, 2009). By reformatting these emails, four test sets were categorized based on the length of each email: Tiny (&amp;#8804; 99 characters), Small (100 to 500 characters), Medium (501 to 999 characters), and Large (&amp;#8805; 1000 characters). The Java Graphical Authorship Attribution Program (JGAAP software) from our Evaluating Variations in Language Laboratory (EVL Lab) was used to perform these tests. Three analysis methods: WEKA RandomForest, WEKA SMO, and Centroid with Cosine Distance were used. Results showed that the Large test set gave the best authorship classification, followed by the Medium, then the Small and the Tiny test sets. WEKA SMO gave better authorship classification than WEKA RandomForest.","abstract_has_math":false,"creators":["Li, Xuan"],"institution":null,"degree_name":"MS","degree_level":"Immediate Access","degree_discipline":"Computational Mathematics","degree_department":null,"school":null,"contributors":["Patrick Juola","Abhay Gaur"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013-01-01T08:00:00Z","date_published":"2013-01-01T08:00:00Z","updated_at":"2026-07-24T02:10:02Z","subjects":["Authorship attribution","Email classification","Enron","EVL Lab","JGAAP","RANDOMFOREST"],"languages":["English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://dsc.duq.edu/etd/823","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Patrick Juola","Abhay Gaur"]},{"key":"dc:creator","label":"Author","values":["Li, Xuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2018-08-03T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computational Mathematics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Immediate Access"]},{"key":"thesis:degree_name","label":"Degree Name","values":["MS"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Authorship attribution","Email classification","Enron","EVL Lab","JGAAP","RANDOMFOREST"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://dsc.duq.edu/etd/823"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In this paper I present authorship attribution on an email corpus. The source I used was the Enron Email Corpus (Cohen, 2009). By reformatting these emails, four test sets were categorized based on the length of each email: Tiny (&#8804; 99 characters), Small (100 to 500 characters), Medium (501 to 999 characters), and Large (&#8805; 1000 characters). The Java Graphical Authorship Attribution Program (JGAAP software) from our Evaluating Variations in Language Laboratory (EVL Lab) was used to perform these tests. Three analysis methods: WEKA RandomForest, WEKA SMO, and Centroid with Cosine Distance were used. Results showed that the Large test set gave the best authorship classification, followed by the Medium, then the Small and the Tiny test sets. WEKA SMO gave better authorship classification than WEKA RandomForest."]},{"key":"dc:title","label":"Title","values":["Authorship Attribution on the Enron Email Corpus"]}]}],"canonical_facts":{"dc:contributor":["Patrick Juola","Abhay Gaur"],"dc:creator":["Li, Xuan"],"dc:date.available":["2018-08-03T07:00:00Z"],"dc:description.abstract":["In this paper I present authorship attribution on an email corpus. The source I used was the Enron Email Corpus (Cohen, 2009). By reformatting these emails, four test sets were categorized based on the length of each email: Tiny (&#8804; 99 characters), Small (100 to 500 characters), Medium (501 to 999 characters), and Large (&#8805; 1000 characters). The Java Graphical Authorship Attribution Program (JGAAP software) from our Evaluating Variations in Language Laboratory (EVL Lab) was used to perform these tests. Three analysis methods: WEKA RandomForest, WEKA SMO, and Centroid with Cosine Distance were used. Results showed that the Large test set gave the best authorship classification, followed by the Medium, then the Small and the Tiny test sets. WEKA SMO gave better authorship classification than WEKA RandomForest."],"dc:identifier":["https://dsc.duq.edu/etd/823"],"dc:language":["English"],"dc:subject":["Authorship attribution","Email classification","Enron","EVL Lab","JGAAP","RANDOMFOREST"],"dc:title":["Authorship Attribution on the Enron Email Corpus"],"thesis:degree_discipline":["Computational Mathematics"],"thesis:degree_level":["Immediate Access"],"thesis:degree_name":["MS"]},"updated_at":"2026-07-24T02:10:02Z"}