{"id":{"repo_id":"uno","oai_identifier":"oai:scholarworks.uno.edu:td-1857"},"canonical_url":"https://search.dev.ndltd.org/etd/uno/oai:scholarworks.uno.edu:td-1857","repository":{"repo_id":"uno","name":"University of New Orleans","base_url":"https://scholarworks.uno.edu/do/oai/"},"display":{"title":"A Semi-Supervised Information Extraction Framework for Large Redundant Corpora","abstract":"The vast majority of text freely available on the Internet is not available in a form that computers can understand. There have been numerous approaches to automatically extract information from human- readable sources. The most successful attempts rely on vast training sets of data. Others have succeeded in extracting restricted subsets of the available information. These approaches have limited use and require domain knowledge to be coded into the application. The current thesis proposes a novel framework for Information Extraction. From large sets of documents, the system develops statistical models of the data the user wishes to query which generally avoid the lim- itations and complexity of most Information Extractions systems. The framework uses a semi-supervised approach to minimize human input. It also eliminates the need for external Named Entity Recognition systems by relying on freely available databases. The final result is a query-answering system which extracts information from large corpora with a high degree of accuracy.","abstract_html":"The vast majority of text freely available on the Internet is not available in a form that computers can understand. There have been numerous approaches to automatically extract information from human- readable sources. The most successful attempts rely on vast training sets of data. Others have succeeded in extracting restricted subsets of the available information. These approaches have limited use and require domain knowledge to be coded into the application. The current thesis proposes a novel framework for Information Extraction. From large sets of documents, the system develops statistical models of the data the user wishes to query which generally avoid the lim- itations and complexity of most Information Extractions systems. The framework uses a semi-supervised approach to minimize human input. It also eliminates the need for external Named Entity Recognition systems by relying on freely available databases. The final result is a query-answering system which extracts information from large corpora with a high degree of accuracy.","abstract_has_math":false,"creators":["Normand, Eric"],"institution":null,"degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Abdelguerfi, Mahdi","Richard III, Golden","Tu, Shengru"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2008,"date_issued":"2008-12-19T08:00:00Z","date_published":"2008-12-19T08:00:00Z","updated_at":"2026-07-24T05:28:50Z","subjects":["Information Extraction","Natural Language Processing","Support Vector Machine","Machine Learn- ing","Information Retrieval","unstructured text"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://scholarworks.uno.edu/td/877","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Abdelguerfi, Mahdi","Richard III, Golden","Tu, Shengru"]},{"key":"dc:creator","label":"Author","values":["Normand, Eric"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Information Extraction","Natural Language Processing","Support Vector Machine","Machine Learn- ing","Information Retrieval","unstructured text"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://scholarworks.uno.edu/td/877"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The vast majority of text freely available on the Internet is not available in a form that computers can understand. There have been numerous approaches to automatically extract information from human- readable sources. The most successful attempts rely on vast training sets of data. Others have succeeded in extracting restricted subsets of the available information. These approaches have limited use and require domain knowledge to be coded into the application. The current thesis proposes a novel framework for Information Extraction. From large sets of documents, the system develops statistical models of the data the user wishes to query which generally avoid the lim- itations and complexity of most Information Extractions systems. The framework uses a semi-supervised approach to minimize human input. It also eliminates the need for external Named Entity Recognition systems by relying on freely available databases. The final result is a query-answering system which extracts information from large corpora with a high degree of accuracy."]},{"key":"dc:title","label":"Title","values":["A Semi-Supervised Information Extraction Framework for Large Redundant Corpora"]}]}],"canonical_facts":{"dc:contributor":["Abdelguerfi, Mahdi","Richard III, Golden","Tu, Shengru"],"dc:creator":["Normand, Eric"],"dc:description.abstract":["The vast majority of text freely available on the Internet is not available in a form that computers can understand. There have been numerous approaches to automatically extract information from human- readable sources. The most successful attempts rely on vast training sets of data. Others have succeeded in extracting restricted subsets of the available information. These approaches have limited use and require domain knowledge to be coded into the application. The current thesis proposes a novel framework for Information Extraction. From large sets of documents, the system develops statistical models of the data the user wishes to query which generally avoid the lim- itations and complexity of most Information Extractions systems. The framework uses a semi-supervised approach to minimize human input. It also eliminates the need for external Named Entity Recognition systems by relying on freely available databases. The final result is a query-answering system which extracts information from large corpora with a high degree of accuracy."],"dc:identifier":["https://scholarworks.uno.edu/td/877"],"dc:subject":["Information Extraction","Natural Language Processing","Support Vector Machine","Machine Learn- ing","Information Retrieval","unstructured text"],"dc:title":["A Semi-Supervised Information Extraction Framework for Large Redundant Corpora"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."]},"updated_at":"2026-07-24T05:28:50Z"}