Abstract
dc:description.abstractAcquiring and representing the large body of "common sense" knowledge underlying ordinary human reasoning and communication is a long standing problem in the field of artificial intelligence. This thesis will address the question whether a significant quantity of this knowledge may be acquired by mining natural language content on the Web. Specifically, this thesis emphasizes the representation of knowledge in the form of binary semantic relationships, such as cause, effect, intent, and time, among natural language phrases. The central hypothesis is that seed knowledge collected from volunteers enables automated acquisition of this knowledge from a large, unannotated, general corpus like the Web. A text mining system, ConceptMiner, was developed to evaluate this hypothesis. ConceptMiner leverages web search engines, Information Extraction techniques and the ConceptNet toolkit to analyze Web content for textual evidence indicating common sense relationships.
Degree
thesis:*- Department dc:contributor.department
- Massachusetts Institute of Technology. Dept. of Architecture. Program In Media Arts and Sciences
- Grantor dc:publisher
- Massachusetts Institute of Technology
- Year dc:date.issued
- 2006
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Eslick, Ian S. (Ian Scott)
- Advisor dc:contributor.advisor
-
- Walter Bender, Hugh Herr and Rada Mihalcea.
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission.
- Licence dc:rights.uri
- Language dc:language.iso
- eng
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/1721.1/37385
- OAI identifier oai:identifier
- oai:dspace.mit.edu:1721.1/37385