Back to results

Monterey, California. Naval Postgraduate School

Language identification by statistical analysis.

Abstract

dc:description.abstract

An analysis was conducted of English and Spanish text. The statistical analysis determined the independent probability of letters and the joint probability of various letter combinations for large samples of each language. Various methods were tested in an attempt to utilize these characteristics to identify the language of a short sample text. By use of the joint probability of various vowel-consonant relationships and the Kolmogorov-Smirnov Goodness of Fit Test an identification system was defined that provided a significance level of .0077 for a sample of 107 letters (approximately 21 words). Investigation also showed that the space rate or the interword structure in each language contains a measure of intelligence and was useful in identification

Degree

thesis:*
Grantor dc:publisher
Monterey, California. Naval Postgraduate School
Year dc:date.issued
1974

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Rau, Morton David
Advisors dc:contributor.advisor
  • Weitzman, R.A.
  • Power, V.M.

Rights

dc:rights
Statement dc:rights
  • This publication is a work of the U.S. Government as defined in Title 17, United States Code, Section 101. Copyright protection is not available for this work in the United States.
Language dc:language.iso
en_US

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/10945/17103
OAI identifier oai:identifier
oai:calhoun.nps.edu:10945/17103

Chain of custody

source
Harvested from
Naval Postgraduate School
Base URL
calhoun.nps.edu/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
related terms
citation

Rau, Morton David. Language identification by statistical analysis.. Monterey, California. Naval Postgraduate School, 1974. https://hdl.handle.net/10945/17103