{"id":{"repo_id":"unt","oai_identifier":"info:ark/67531/metadc5140"},"canonical_url":"https://search.dev.ndltd.org/etd/unt/info:ark/67531/metadc5140","repository":{"repo_id":"unt","name":"University of North Texas","base_url":"https://digital.library.unt.edu/oai/"},"display":{"title":"The enhancement of machine translation for low-density languages using Web-gathered parallel texts.","abstract":"The majority of the world's languages are poorly represented in informational media like radio, television, newspapers, and the Internet. Translation into and out of these languages may offer a way for speakers of these languages to interact with the wider world, but current statistical machine translation models are only effective with a large corpus of parallel texts - texts in two languages that are translations of one another - which most languages lack. This thesis describes the Babylon project which attempts to alleviate this shortage by supplementing existing parallel texts with texts gathered automatically from the Web -- specifically targeting pages that contain text in a pair of languages. Results indicate that parallel texts gathered from the Web can be effectively used as a source of training data for machine translation and can significantly improve the translation quality for text in a similar domain. However, the small quantity of high-quality low-density language parallel texts on the Web remains a significant obstacle.","abstract_html":"The majority of the world&#x27;s languages are poorly represented in informational media like radio, television, newspapers, and the Internet. Translation into and out of these languages may offer a way for speakers of these languages to interact with the wider world, but current statistical machine translation models are only effective with a large corpus of parallel texts - texts in two languages that are translations of one another - which most languages lack. This thesis describes the Babylon project which attempts to alleviate this shortage by supplementing existing parallel texts with texts gathered automatically from the Web -- specifically targeting pages that contain text in a pair of languages. Results indicate that parallel texts gathered from the Web can be effectively used as a source of training data for machine translation and can significantly improve the translation quality for text in a similar domain. However, the small quantity of high-quality low-density language parallel texts on the Web remains a significant obstacle.","abstract_has_math":false,"creators":["Mohler, Michael Augustine Gaylord"],"institution":"University of North Texas","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Mihalcea, Rada, 1974-","Tarau, Paul","Chen, Jiangping"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2007,"date_issued":"2007-12","date_published":"2007-12","updated_at":"2026-07-24T05:35:09Z","subjects":["machine translation","text alignment","biblical texts","parallel texts","low-density languages","Machine translating."],"languages":["English"],"rights":["Public","Copyright","Mohler, Michael Augustine Gaylord","Copyright is held by the author, unless otherwise noted. All rights reserved."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["oclc: 228503664","https://digital.library.unt.edu/ark:/67531/metadc5140/","ark: ark:/67531/metadc5140"],"render_values":[{"text":"oclc: 228503664","href":null,"code":true},{"text":"https://digital.library.unt.edu/ark:/67531/metadc5140/","href":"https://digital.library.unt.edu/ark:/67531/metadc5140/","code":true},{"text":"ark: ark:/67531/metadc5140","href":null,"code":true}]}]},"links":{"outbound_url":"https://doi.org/10.12794/metadc5140","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Mihalcea, Rada, 1974-","Tarau, Paul","Chen, Jiangping"]},{"key":"dc:creator","label":"Author","values":["Mohler, Michael Augustine Gaylord"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2007-12"]},{"key":"dc:publisher","label":"Institution","values":["University of North Texas"]},{"key":"dc:type","label":"Dc Type","values":["Thesis or Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["machine translation","text alignment","biblical texts","parallel texts","low-density languages","Machine translating."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:rights","label":"Dc Rights","values":["Public","Copyright","Mohler, Michael Augustine Gaylord","Copyright is held by the author, unless otherwise noted. All rights reserved."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["oclc: 228503664","doi: 10.12794/metadc5140","https://digital.library.unt.edu/ark:/67531/metadc5140/","ark: ark:/67531/metadc5140"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The majority of the world's languages are poorly represented in informational media like radio, television, newspapers, and the Internet. Translation into and out of these languages may offer a way for speakers of these languages to interact with the wider world, but current statistical machine translation models are only effective with a large corpus of parallel texts - texts in two languages that are translations of one another - which most languages lack. This thesis describes the Babylon project which attempts to alleviate this shortage by supplementing existing parallel texts with texts gathered automatically from the Web -- specifically targeting pages that contain text in a pair of languages. Results indicate that parallel texts gathered from the Web can be effectively used as a source of training data for machine translation and can significantly improve the translation quality for text in a similar domain. However, the small quantity of high-quality low-density language parallel texts on the Web remains a significant obstacle."]},{"key":"dc:format","label":"Dc Format","values":["Text"]},{"key":"dc:title","label":"Title","values":["The enhancement of machine translation for low-density languages using Web-gathered parallel texts."]}]}],"canonical_facts":{"dc:contributor":["Mihalcea, Rada, 1974-","Tarau, Paul","Chen, Jiangping"],"dc:creator":["Mohler, Michael Augustine Gaylord"],"dc:date":["2007-12"],"dc:description":["The majority of the world's languages are poorly represented in informational media like radio, television, newspapers, and the Internet. Translation into and out of these languages may offer a way for speakers of these languages to interact with the wider world, but current statistical machine translation models are only effective with a large corpus of parallel texts - texts in two languages that are translations of one another - which most languages lack. This thesis describes the Babylon project which attempts to alleviate this shortage by supplementing existing parallel texts with texts gathered automatically from the Web -- specifically targeting pages that contain text in a pair of languages. Results indicate that parallel texts gathered from the Web can be effectively used as a source of training data for machine translation and can significantly improve the translation quality for text in a similar domain. However, the small quantity of high-quality low-density language parallel texts on the Web remains a significant obstacle."],"dc:format":["Text"],"dc:identifier":["oclc: 228503664","doi: 10.12794/metadc5140","https://digital.library.unt.edu/ark:/67531/metadc5140/","ark: ark:/67531/metadc5140"],"dc:language":["English"],"dc:publisher":["University of North Texas"],"dc:rights":["Public","Copyright","Mohler, Michael Augustine Gaylord","Copyright is held by the author, unless otherwise noted. All rights reserved."],"dc:subject":["machine translation","text alignment","biblical texts","parallel texts","low-density languages","Machine translating."],"dc:title":["The enhancement of machine translation for low-density languages using Web-gathered parallel texts."],"dc:type":["Thesis or Dissertation"]},"updated_at":"2026-07-24T05:35:09Z"}