{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/81987"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/81987","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Sequentialized Language Models","abstract":"We then turn to construction of a sequentialized grammatical model of linguistic objects in text compression. We develop the Prediction by Grammatical Match technique, a new compression framework employing a static context-free grammar and an adaptive finite-context statistical model. These compressors are adaptive, general compressors that operate in linear time and bounded space. We show these compressors can deliver substantial reductions in both bits-per-character rates and space usage, and suffer almost no penalty when the grammar does not apply. The new technique rests on three primary technical innovations: an algorithm for designing an optimal, strictly bottom-up parseable metalanguage for a compression scheme comprising multiple grammars; a principled approach to ambiguity and agrammatical text; and an incremental analysis selection algorithm. The metalanguage construction emphasizes lexical left-corner analysis descriptions, with each symbol in a description representing a maximal bundle of bottom-up and top-down information by naming the production introducing the next lexical left-corner item. These three innovations combine into a very powerful compression system that solves an important, long standing problem: efficient and effective use of context-free grammars in general data compression.","abstract_html":"We then turn to construction of a sequentialized grammatical model of linguistic objects in text compression. We develop the Prediction by Grammatical Match technique, a new compression framework employing a static context-free grammar and an adaptive finite-context statistical model. These compressors are adaptive, general compressors that operate in linear time and bounded space. We show these compressors can deliver substantial reductions in both bits-per-character rates and space usage, and suffer almost no penalty when the grammar does not apply. The new technique rests on three primary technical innovations: an algorithm for designing an optimal, strictly bottom-up parseable metalanguage for a compression scheme comprising multiple grammars; a principled approach to ambiguity and agrammatical text; and an incremental analysis selection algorithm. The metalanguage construction emphasizes lexical left-corner analysis descriptions, with each symbol in a description representing a maximal bundle of bottom-up and top-down information by naming the production introducing the next lexical left-corner item. These three innovations combine into a very powerful compression system that solves an important, long standing problem: efficient and effective use of context-free grammars in general data compression.","abstract_has_math":false,"creators":["Lake, John Michael"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["DeJong, Gerald F."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-09-25T20:21:18Z","date_published":"2015-09-25T20:21:18Z","updated_at":"2026-07-22T22:26:17Z","subjects":["Computer Science"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["(MiAaPQ)AAI9990052"],"render_values":[{"text":"(MiAaPQ)AAI9990052","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/81987","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["DeJong, Gerald F."]},{"key":"dc:creator","label":"Author","values":["Lake, John Michael"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-09-25T20:21:18Z","10000-01-01","2000"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/81987","(MiAaPQ)AAI9990052"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We then turn to construction of a sequentialized grammatical model of linguistic objects in text compression. We develop the Prediction by Grammatical Match technique, a new compression framework employing a static context-free grammar and an adaptive finite-context statistical model. These compressors are adaptive, general compressors that operate in linear time and bounded space. We show these compressors can deliver substantial reductions in both bits-per-character rates and space usage, and suffer almost no penalty when the grammar does not apply. The new technique rests on three primary technical innovations: an algorithm for designing an optimal, strictly bottom-up parseable metalanguage for a compression scheme comprising multiple grammars; a principled approach to ambiguity and agrammatical text; and an incremental analysis selection algorithm. The metalanguage construction emphasizes lexical left-corner analysis descriptions, with each symbol in a description representing a maximal bundle of bottom-up and top-down information by naming the production introducing the next lexical left-corner item. These three innovations combine into a very powerful compression system that solves an important, long standing problem: efficient and effective use of context-free grammars in general data compression.","Made available in DSpace on 2015-09-25T20:21:18Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 9990052.pdf: 13518216 bytes, checksum: 5cdd5b829e74803f652e67dd74bf539c (MD5) Previous issue date: 2000","Embargo set by: Seth Robbins for item 83268 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","196 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2000."]},{"key":"dc:title","label":"Title","values":["Sequentialized Language Models"]}]}],"canonical_facts":{"dc:contributor":["DeJong, Gerald F."],"dc:creator":["Lake, John Michael"],"dc:date":["2015-09-25T20:21:18Z","10000-01-01","2000"],"dc:description":["We then turn to construction of a sequentialized grammatical model of linguistic objects in text compression. We develop the Prediction by Grammatical Match technique, a new compression framework employing a static context-free grammar and an adaptive finite-context statistical model. These compressors are adaptive, general compressors that operate in linear time and bounded space. We show these compressors can deliver substantial reductions in both bits-per-character rates and space usage, and suffer almost no penalty when the grammar does not apply. The new technique rests on three primary technical innovations: an algorithm for designing an optimal, strictly bottom-up parseable metalanguage for a compression scheme comprising multiple grammars; a principled approach to ambiguity and agrammatical text; and an incremental analysis selection algorithm. The metalanguage construction emphasizes lexical left-corner analysis descriptions, with each symbol in a description representing a maximal bundle of bottom-up and top-down information by naming the production introducing the next lexical left-corner item. These three innovations combine into a very powerful compression system that solves an important, long standing problem: efficient and effective use of context-free grammars in general data compression.","Made available in DSpace on 2015-09-25T20:21:18Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 9990052.pdf: 13518216 bytes, checksum: 5cdd5b829e74803f652e67dd74bf539c (MD5) Previous issue date: 2000","Embargo set by: Seth Robbins for item 83268 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","196 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2000."],"dc:identifier":["http://hdl.handle.net/2142/81987","(MiAaPQ)AAI9990052"],"dc:language":["eng"],"dc:subject":["Computer Science"],"dc:title":["Sequentialized Language Models"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:17Z"}