{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/21065"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/21065","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"An experiment in automatic indexing with Korean texts: A comparison of syntactico-statistical and manual methods","abstract":"This study was undertaken in order to develop practical automatic indexing techniques suitable for Korean natural language texts. The study had four purposes: to develop an automatic indexing system for Korean texts, to evaluate the efficiency of the automatic indexing system as compared with a manual indexing system, to compare the effectiveness of weighting algorithms, and to investigate the effect of abstract length.","abstract_html":"This study was undertaken in order to develop practical automatic indexing techniques suitable for Korean natural language texts. The study had four purposes: to develop an automatic indexing system for Korean texts, to evaluate the efficiency of the automatic indexing system as compared with a manual indexing system, to compare the effectiveness of weighting algorithms, and to investigate the effect of abstract length.","abstract_has_math":false,"creators":["Seo, Eun-Gyoung"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Library and Information Science","degree_department":null,"school":null,"contributors":["Smith, Linda C.","Allen, Bryce L.","Davis, Charles H."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-05-07T12:57:19Z","date_published":"2011-05-07T12:57:19Z","updated_at":"2026-07-22T22:25:17Z","subjects":["Library Science","Computer Science"],"languages":["eng"],"rights":["Copyright 1993 Seo, Eun-Gyoung"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["AAI9329159","(UMI)AAI9329159"],"render_values":[{"text":"AAI9329159","href":null,"code":true},{"text":"(UMI)AAI9329159","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/21065","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Smith, Linda C.","Allen, Bryce L.","Davis, Charles H."]},{"key":"dc:creator","label":"Author","values":["Seo, Eun-Gyoung"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2011-05-07T12:57:19Z","10000-01-01","1993"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Library and Information Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Library Science","Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 1993 Seo, Eun-Gyoung"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["AAI9329159","(UMI)AAI9329159","http://hdl.handle.net/2142/21065"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This study was undertaken in order to develop practical automatic indexing techniques suitable for Korean natural language texts. The study had four purposes: to develop an automatic indexing system for Korean texts, to evaluate the efficiency of the automatic indexing system as compared with a manual indexing system, to compare the effectiveness of weighting algorithms, and to investigate the effect of abstract length.","The basic method of this automatic indexing system was to determine the syntactic category of each text word by dictionary look-up, and then to match sequences of category symbols against a dictionary of acceptable patterns. Sequences of text words that matched one of the patterns in the dictionary were extracted as content identifiers. Finally, the system selected highly ranked content identifiers from each document based on statistical (frequency of occurrence) information.","For this experimental study, the Korean text database was constructed manually based on 100 long abstracts and 200 short abstracts covering business subjects. The study assessed how well the set of index terms produced by an automatic indexing technique reflects the major topics described in an indexed document. For the evaluation, a manual index term list was constructed by consultation between two indexers as an external standard to obtain normalized values.","The experimental results showed that the performance of the automatic syntactico-statistical indexing system was comparable to that of other studies which have compared automatic indexing with manual indexing. The WDF system performed better than the IDF system in terms of the ability to present all the correct content identifiers, and the system produced more correct content identifiers in the short abstract group. As a whole, many significant concepts represented in the abstract and recognized by human indexers have been effectively extracted automatically. The extracted concept forms are for the most part comparable to those of manual indexing. Possible enhancements of the automatic syntactico-statistical indexing system are identified which could lead to improved indexing performance.","Made available in DSpace on 2011-05-07T12:57:19Z (GMT). No. of bitstreams: 2 license.txt: 4922 bytes, checksum: 910b249b4beec47e7ab768910c8f966f (MD5) 9329159.pdf: 10765318 bytes, checksum: d9a5486da3d340b31489d7894ffd5568 (MD5) Previous issue date: 1993","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by Howard Ding (hding2@illinois.edu) on 2011-05-07T14:48:14Z Item is restricted indefinitely.","Restriction data tranferred 2014-07-01T11:21:50-05:00 Original Data Group with Access UIUC Users [automated] Release Date: none Reason: ETDs are only available to UIUC Users without author permission","ETDs are only available to UIUC Users without author permission","U of I Only"]},{"key":"dc:title","label":"Title","values":["An experiment in automatic indexing with Korean texts: A comparison of syntactico-statistical and manual methods"]}]}],"canonical_facts":{"dc:contributor":["Smith, Linda C.","Allen, Bryce L.","Davis, Charles H."],"dc:creator":["Seo, Eun-Gyoung"],"dc:date":["2011-05-07T12:57:19Z","10000-01-01","1993"],"dc:description":["This study was undertaken in order to develop practical automatic indexing techniques suitable for Korean natural language texts. The study had four purposes: to develop an automatic indexing system for Korean texts, to evaluate the efficiency of the automatic indexing system as compared with a manual indexing system, to compare the effectiveness of weighting algorithms, and to investigate the effect of abstract length.","The basic method of this automatic indexing system was to determine the syntactic category of each text word by dictionary look-up, and then to match sequences of category symbols against a dictionary of acceptable patterns. Sequences of text words that matched one of the patterns in the dictionary were extracted as content identifiers. Finally, the system selected highly ranked content identifiers from each document based on statistical (frequency of occurrence) information.","For this experimental study, the Korean text database was constructed manually based on 100 long abstracts and 200 short abstracts covering business subjects. The study assessed how well the set of index terms produced by an automatic indexing technique reflects the major topics described in an indexed document. For the evaluation, a manual index term list was constructed by consultation between two indexers as an external standard to obtain normalized values.","The experimental results showed that the performance of the automatic syntactico-statistical indexing system was comparable to that of other studies which have compared automatic indexing with manual indexing. The WDF system performed better than the IDF system in terms of the ability to present all the correct content identifiers, and the system produced more correct content identifiers in the short abstract group. As a whole, many significant concepts represented in the abstract and recognized by human indexers have been effectively extracted automatically. The extracted concept forms are for the most part comparable to those of manual indexing. Possible enhancements of the automatic syntactico-statistical indexing system are identified which could lead to improved indexing performance.","Made available in DSpace on 2011-05-07T12:57:19Z (GMT). No. of bitstreams: 2 license.txt: 4922 bytes, checksum: 910b249b4beec47e7ab768910c8f966f (MD5) 9329159.pdf: 10765318 bytes, checksum: d9a5486da3d340b31489d7894ffd5568 (MD5) Previous issue date: 1993","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by Howard Ding (hding2@illinois.edu) on 2011-05-07T14:48:14Z Item is restricted indefinitely.","Restriction data tranferred 2014-07-01T11:21:50-05:00 Original Data Group with Access UIUC Users [automated] Release Date: none Reason: ETDs are only available to UIUC Users without author permission","ETDs are only available to UIUC Users without author permission","U of I Only"],"dc:identifier":["AAI9329159","(UMI)AAI9329159","http://hdl.handle.net/2142/21065"],"dc:language":["eng"],"dc:rights":["Copyright 1993 Seo, Eun-Gyoung"],"dc:subject":["Library Science","Computer Science"],"dc:title":["An experiment in automatic indexing with Korean texts: A comparison of syntactico-statistical and manual methods"],"dc:type":["text"],"thesis:degree_discipline":["Library and Information Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:17Z"}