{"id":{"repo_id":"emich","oai_identifier":"oai:commons.emich.edu:theses-1000"},"canonical_url":"https://search.dev.ndltd.org/etd/emich/oai:commons.emich.edu:theses-1000","repository":{"repo_id":"emich","name":"Eastern Michigan University","base_url":"https://commons.emich.edu/do/oai/"},"display":{"title":"Generating and presenting string frequency measurements of Project Gutenberg texts","abstract":"<p>The electronic age has increased the range of human capabilities to such an extent that the expectations about appropriate empirical linguistic analysis are changing. A hundred years ago, linguistics was largely an empirical manual process that produced information intended for humans. Today, the world is different as inexpensive computing power and the prevalence of information in electronic format encourages that, whenever possible, information be processed by automated and scalable means and the results be usable and understandable by computers. Creating sustainable and usable observations is best achieved through a standards-based approach that meets long term persistence and usability goals. This thesis presents a scalable architecture for creating linguistic observations in the form of string frequencies measurements and instantiates those measurements in a machine-readable standards-based format called Resource Descriptive Framework (RDF).</p>","abstract_html":"&lt;p&gt;The electronic age has increased the range of human capabilities to such an extent that the expectations about appropriate empirical linguistic analysis are changing. A hundred years ago, linguistics was largely an empirical manual process that produced information intended for humans. Today, the world is different as inexpensive computing power and the prevalence of information in electronic format encourages that, whenever possible, information be processed by automated and scalable means and the results be usable and understandable by computers. Creating sustainable and usable observations is best achieved through a standards-based approach that meets long term persistence and usability goals. This thesis presents a scalable architecture for creating linguistic observations in the form of string frequencies measurements and instantiates those measurements in a machine-readable standards-based format called Resource Descriptive Framework (RDF).&lt;/p&gt;","abstract_has_math":false,"creators":["Reck, Ronald P"],"institution":null,"degree_name":"Master of Arts (MA)","degree_level":"Open Access Thesis","degree_discipline":"English Language and Literature","degree_department":null,"school":null,"contributors":["Anthony Aristar","Helen Aristar-Dry"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2007,"date_issued":"2007-01-01T08:00:00Z","date_published":"2007-01-01T08:00:00Z","updated_at":"2026-07-24T02:16:24Z","subjects":["Metadata","RDF (Document markup language)","Computational linguistics","Linguistics"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://commons.emich.edu/theses/1","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Anthony Aristar","Helen Aristar-Dry"]},{"key":"dc:creator","label":"Author","values":["Reck, Ronald P"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["English Language and Literature"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Open Access Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Arts (MA)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Metadata","RDF (Document markup language)","Computational linguistics","Linguistics"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://commons.emich.edu/theses/1"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>The electronic age has increased the range of human capabilities to such an extent that the expectations about appropriate empirical linguistic analysis are changing. A hundred years ago, linguistics was largely an empirical manual process that produced information intended for humans. Today, the world is different as inexpensive computing power and the prevalence of information in electronic format encourages that, whenever possible, information be processed by automated and scalable means and the results be usable and understandable by computers. Creating sustainable and usable observations is best achieved through a standards-based approach that meets long term persistence and usability goals. This thesis presents a scalable architecture for creating linguistic observations in the form of string frequencies measurements and instantiates those measurements in a machine-readable standards-based format called Resource Descriptive Framework (RDF).</p>"]},{"key":"dc:title","label":"Title","values":["Generating and presenting string frequency measurements of Project Gutenberg texts"]}]}],"canonical_facts":{"dc:contributor":["Anthony Aristar","Helen Aristar-Dry"],"dc:creator":["Reck, Ronald P"],"dc:description.abstract":["<p>The electronic age has increased the range of human capabilities to such an extent that the expectations about appropriate empirical linguistic analysis are changing. A hundred years ago, linguistics was largely an empirical manual process that produced information intended for humans. Today, the world is different as inexpensive computing power and the prevalence of information in electronic format encourages that, whenever possible, information be processed by automated and scalable means and the results be usable and understandable by computers. Creating sustainable and usable observations is best achieved through a standards-based approach that meets long term persistence and usability goals. This thesis presents a scalable architecture for creating linguistic observations in the form of string frequencies measurements and instantiates those measurements in a machine-readable standards-based format called Resource Descriptive Framework (RDF).</p>"],"dc:identifier":["https://commons.emich.edu/theses/1"],"dc:subject":["Metadata","RDF (Document markup language)","Computational linguistics","Linguistics"],"dc:title":["Generating and presenting string frequency measurements of Project Gutenberg texts"],"thesis:degree_discipline":["English Language and Literature"],"thesis:degree_level":["Open Access Thesis"],"thesis:degree_name":["Master of Arts (MA)"]},"updated_at":"2026-07-24T02:16:24Z"}