Massachusetts Institute of Technology
Automatic classification of documents with an in-depth analysis of information extraction and automatic summarization
Abstract
dc:description.abstractToday, annual information fabrication per capita exceeds two hundred and fifty megabytes. As the amount of data increases, classification and retrieval methods become more necessary to find relevant information. This thesis describes a .Net application (named I-Document) that establishes an automatic classification scheme in a peer-to-peer environment that allows free sharing of academic, business, and personal documents. A Web service architecture for metadata extraction, Information Extraction, Information Retrieval, and text summarization is depicted. Specific details regarding the coding process, competition, business model, and technology employed in the project are also discussed.
Degree
thesis:*- Department dc:contributor.department
- Massachusetts Institute of Technology. Dept. of Civil and Environmental Engineering.
- Grantor dc:publisher
- Massachusetts Institute of Technology
- Year dc:date.issued
- 2004
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Hohm, Joseph Brandon, 1982-
- Advisor dc:contributor.advisor
-
- John R. Williams.
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission.
- Licence dc:rights.uri
- Language dc:language.iso
- eng
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/1721.1/29415
- OAI identifier oai:identifier
- oai:dspace.mit.edu:1721.1/29415