Back to results

Duquesne

Authorship Attribution on the Enron Email Corpus

Abstract

dc:description.abstract

In this paper I present authorship attribution on an email corpus. The source I used was the Enron Email Corpus (Cohen, 2009). By reformatting these emails, four test sets were categorized based on the length of each email: Tiny (≤ 99 characters), Small (100 to 500 characters), Medium (501 to 999 characters), and Large (≥ 1000 characters). The Java Graphical Authorship Attribution Program (JGAAP software) from our Evaluating Variations in Language Laboratory (EVL Lab) was used to perform these tests. Three analysis methods: WEKA RandomForest, WEKA SMO, and Centroid with Cosine Distance were used. Results showed that the Large test set gave the best authorship classification, followed by the Medium, then the Small and the Tiny test sets. WEKA SMO gave better authorship classification than WEKA RandomForest.

Degree

thesis:*
Name thesis:degree_name
MS
Level thesis:degree_level
Immediate Access
Discipline thesis:degree_discipline
Computational Mathematics
Year dc:date.available
2013

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Li, Xuan
Contributors dc:contributor
  • Patrick Juola
  • Abhay Gaur

Subjects

dc:subject × 6

Rights

Language dc:language
English

Identifiers

dc:identifier.*
Repository record dc:identifier
https://dsc.duq.edu/etd/823
OAI identifier oai:identifier
oai:dsc.duq.edu:etd-1839

Chain of custody

source
Harvested from
Duquesne
Base URL
dsc.duq.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Li, Xuan. Authorship Attribution on the Enron Email Corpus. Immediate Access thesis, 2013. https://dsc.duq.edu/etd/823