Eastern Washington University
Support vector machines, N-gram kernels, and text classification
Abstract
dc:description.abstract<p>The expanding popularity of the Internet in recent years has lead to a corresponding increase in the amount of textual data available. This increase is found in the number of web pages, the size and complexity of search engines, and massive volumes of email. For any one attempting to sort through or make sense of this data, one of the fundamental tasks is text classification. Text classification is the task of identifying the category that a given piece of text or document belongs to. In the case of e-mail directed at an on line retailer the categories might be the various product departments. In the case of a search engine the category could be the set of documents relevant to a search topic. In recent years, a new inference method known as Support Vector Machines (SVMs) has been increasingly applied to the task of text classification. The results have been promising and research shows that they outperform several conventional methods. One the key components of SVMs are kernel functions. The choice of kernel function can have substantial effects on the performance of SVMs. In this paper we explore kernels based off of N-grams or consecutive sequences of words.</p>
Degree
thesis:*- Name thesis:degree_name
- Master of Science (MS) in Computer Science
- Level thesis:degree_level
- Thesis: EWU Only
- Discipline thesis:degree_discipline
- Computer Science
- Year
- 2002
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Mill, John
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Access perpetually restricted to EWU users with an active EWU NetID
Identifiers
dc:identifier.*- Repository record dc:identifier
- https://dc.ewu.edu/theses/823
- OAI identifier oai:identifier
- oai:dc.ewu.edu:theses-1827