{"id":{"repo_id":"nps","oai_identifier":"oai:calhoun.nps.edu:10945/17103"},"canonical_url":"https://search.dev.ndltd.org/etd/nps/oai:calhoun.nps.edu:10945/17103","repository":{"repo_id":"nps","name":"Naval Postgraduate School","base_url":"https://calhoun.nps.edu/server/oai/request"},"display":{"title":"Language identification by statistical analysis.","abstract":"An analysis was conducted of English and Spanish text. The statistical analysis determined the independent probability of letters and the joint probability of various letter combinations for large samples of each language. Various methods were tested in an attempt to utilize these characteristics to identify the language of a short sample text. By use of the joint probability of various vowel-consonant relationships and the Kolmogorov-Smirnov Goodness of Fit Test an identification system was defined that provided a significance level of .0077 for a sample of 107 letters (approximately 21 words). Investigation also showed that the space rate or the interword structure in each language contains a measure of intelligence and was useful in identification","abstract_html":"An analysis was conducted of English and Spanish text. The statistical analysis determined the independent probability of letters and the joint probability of various letter combinations for large samples of each language. Various methods were tested in an attempt to utilize these characteristics to identify the language of a short sample text. By use of the joint probability of various vowel-consonant relationships and the Kolmogorov-Smirnov Goodness of Fit Test an identification system was defined that provided a significance level of .0077 for a sample of 107 letters (approximately 21 words). Investigation also showed that the space rate or the interword structure in each language contains a measure of intelligence and was useful in identification","abstract_has_math":false,"creators":["Rau, Morton David"],"institution":"Monterey, California. Naval Postgraduate School","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Weitzman, R.A.","Power, V.M."],"committee_chairs":[],"committee_members":[],"year":1974,"date_issued":"1974-09","date_published":"1974-09","updated_at":"2026-07-27T20:26:05Z","subjects":[],"languages":["en_US"],"rights":["This publication is a work of the U.S. Government as defined in Title 17, United States Code, Section 101. Copyright protection is not available for this work in the United States."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10945/17103","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Weitzman, R.A.","Power, V.M."]},{"key":"dc:creator","label":"Author","values":["Rau, Morton David"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2012-11-13T23:59:43Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2012-11-13T23:59:43Z"]},{"key":"dc:date.issued","label":"Date","values":["1974-09"]},{"key":"dc:publisher","label":"Institution","values":["Monterey, California. Naval Postgraduate School"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]},{"key":"dc:rights","label":"Dc Rights","values":["This publication is a work of the U.S. Government as defined in Title 17, United States Code, Section 101. Copyright protection is not available for this work in the United States."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10945/17103"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["An analysis was conducted of English and Spanish text. The statistical analysis determined the independent probability of letters and the joint probability of various letter combinations for large samples of each language. Various methods were tested in an attempt to utilize these characteristics to identify the language of a short sample text. By use of the joint probability of various vowel-consonant relationships and the Kolmogorov-Smirnov Goodness of Fit Test an identification system was defined that provided a significance level of .0077 for a sample of 107 letters (approximately 21 words). Investigation also showed that the space rate or the interword structure in each language contains a measure of intelligence and was useful in identification"]},{"key":"dc:title","label":"Title","values":["Language identification by statistical analysis."]}]}],"canonical_facts":{"dc:contributor.advisor":["Weitzman, R.A.","Power, V.M."],"dc:creator":["Rau, Morton David"],"dc:date.accessioned":["2012-11-13T23:59:43Z"],"dc:date.available":["2012-11-13T23:59:43Z"],"dc:date.issued":["1974-09"],"dc:description.abstract":["An analysis was conducted of English and Spanish text. The statistical analysis determined the independent probability of letters and the joint probability of various letter combinations for large samples of each language. Various methods were tested in an attempt to utilize these characteristics to identify the language of a short sample text. By use of the joint probability of various vowel-consonant relationships and the Kolmogorov-Smirnov Goodness of Fit Test an identification system was defined that provided a significance level of .0077 for a sample of 107 letters (approximately 21 words). Investigation also showed that the space rate or the interword structure in each language contains a measure of intelligence and was useful in identification"],"dc:identifier.uri":["https://hdl.handle.net/10945/17103"],"dc:language.iso":["en_US"],"dc:publisher":["Monterey, California. Naval Postgraduate School"],"dc:rights":["This publication is a work of the U.S. Government as defined in Title 17, United States Code, Section 101. Copyright protection is not available for this work in the United States."],"dc:title":["Language identification by statistical analysis."],"dc:type":["Thesis"]},"updated_at":"2026-07-27T20:26:05Z"}