{"id":{"repo_id":"eastern-wash","oai_identifier":"oai:dc.ewu.edu:theses-1827"},"canonical_url":"https://search.dev.ndltd.org/etd/eastern-wash/oai:dc.ewu.edu:theses-1827","repository":{"repo_id":"eastern-wash","name":"Eastern Washington University","base_url":"https://dc.ewu.edu/do/oai/"},"display":{"title":"Support vector machines, N-gram kernels, and text classification","abstract":"<p>The expanding popularity of the Internet in recent years has lead to a corresponding increase in the amount of textual data available. This increase is found in the number of web pages, the size and complexity of search engines, and massive volumes of email. For any one attempting to sort through or make sense of this data, one of the fundamental tasks is text classification. Text classification is the task of identifying the category that a given piece of text or document belongs to. In the case of e-mail directed at an on line retailer the categories might be the various product departments. In the case of a search engine the category could be the set of documents relevant to a search topic. In recent years, a new inference method known as Support Vector Machines (SVMs) has been increasingly applied to the task of text classification. The results have been promising and research shows that they outperform several conventional methods. One the key components of SVMs are kernel functions. The choice of kernel function can have substantial effects on the performance of SVMs. In this paper we explore kernels based off of N-grams or consecutive sequences of words.</p>","abstract_html":"&lt;p&gt;The expanding popularity of the Internet in recent years has lead to a corresponding increase in the amount of textual data available. This increase is found in the number of web pages, the size and complexity of search engines, and massive volumes of email. For any one attempting to sort through or make sense of this data, one of the fundamental tasks is text classification. Text classification is the task of identifying the category that a given piece of text or document belongs to. In the case of e-mail directed at an on line retailer the categories might be the various product departments. In the case of a search engine the category could be the set of documents relevant to a search topic. In recent years, a new inference method known as Support Vector Machines (SVMs) has been increasingly applied to the task of text classification. The results have been promising and research shows that they outperform several conventional methods. One the key components of SVMs are kernel functions. The choice of kernel function can have substantial effects on the performance of SVMs. In this paper we explore kernels based off of N-grams or consecutive sequences of words.&lt;/p&gt;","abstract_has_math":false,"creators":["Mill, John"],"institution":null,"degree_name":"Master of Science (MS) in Computer Science","degree_level":"Thesis: EWU Only","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2002,"date_issued":"2002-01-01T08:00:00Z","date_published":"2002-01-01T08:00:00Z","updated_at":"2026-07-24T02:12:46Z","subjects":["Graphics and Human Computer Interfaces","Other Computer Sciences"],"languages":[],"rights":["Access perpetually restricted to EWU users with an active EWU NetID"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://dc.ewu.edu/theses/823","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Mill, John"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis: EWU Only"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS) in Computer Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Graphics and Human Computer Interfaces","Other Computer Sciences"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["Access perpetually restricted to EWU users with an active EWU NetID"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://dc.ewu.edu/theses/823"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>The expanding popularity of the Internet in recent years has lead to a corresponding increase in the amount of textual data available. This increase is found in the number of web pages, the size and complexity of search engines, and massive volumes of email. For any one attempting to sort through or make sense of this data, one of the fundamental tasks is text classification. Text classification is the task of identifying the category that a given piece of text or document belongs to. In the case of e-mail directed at an on line retailer the categories might be the various product departments. In the case of a search engine the category could be the set of documents relevant to a search topic. In recent years, a new inference method known as Support Vector Machines (SVMs) has been increasingly applied to the task of text classification. The results have been promising and research shows that they outperform several conventional methods. One the key components of SVMs are kernel functions. The choice of kernel function can have substantial effects on the performance of SVMs. In this paper we explore kernels based off of N-grams or consecutive sequences of words.</p>"]},{"key":"dc:title","label":"Title","values":["Support vector machines, N-gram kernels, and text classification"]}]}],"canonical_facts":{"dc:creator":["Mill, John"],"dc:description.abstract":["<p>The expanding popularity of the Internet in recent years has lead to a corresponding increase in the amount of textual data available. This increase is found in the number of web pages, the size and complexity of search engines, and massive volumes of email. For any one attempting to sort through or make sense of this data, one of the fundamental tasks is text classification. Text classification is the task of identifying the category that a given piece of text or document belongs to. In the case of e-mail directed at an on line retailer the categories might be the various product departments. In the case of a search engine the category could be the set of documents relevant to a search topic. In recent years, a new inference method known as Support Vector Machines (SVMs) has been increasingly applied to the task of text classification. The results have been promising and research shows that they outperform several conventional methods. One the key components of SVMs are kernel functions. The choice of kernel function can have substantial effects on the performance of SVMs. In this paper we explore kernels based off of N-grams or consecutive sequences of words.</p>"],"dc:identifier":["https://dc.ewu.edu/theses/823"],"dc:rights":["Access perpetually restricted to EWU users with an active EWU NetID"],"dc:subject":["Graphics and Human Computer Interfaces","Other Computer Sciences"],"dc:title":["Support vector machines, N-gram kernels, and text classification"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis: EWU Only"],"thesis:degree_name":["Master of Science (MS) in Computer Science"]},"updated_at":"2026-07-24T02:12:46Z"}