{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/77501"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/77501","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Syntactically annotated Ngrams for Google Books","abstract":"In this thesis, we present a new edition of the Google Books Ngram Corpus, describing how often words and phrases were used over a period of five centuries, in eight languages; it aggregates data from 6% of all books ever published. This new edition introduces syntactic annotations: words are tagged with their part-of-speech, and head-modifier dependency relationships are recorded. We generate these annotations automatically from the Google Books text, using statistical models that are specifically adapted to the historical text found in these books. The new edition will facilitate the study of linguistic trends, especially those related to the evolution of syntax. We present our initial findings from the annotated Ngrams in the new edition, including studies of the change in various words' primary parts of speech over time, and to find the words most closely related to a given set of topics.","abstract_html":"In this thesis, we present a new edition of the Google Books Ngram Corpus, describing how often words and phrases were used over a period of five centuries, in eight languages; it aggregates data from 6% of all books ever published. This new edition introduces syntactic annotations: words are tagged with their part-of-speech, and head-modifier dependency relationships are recorded. We generate these annotations automatically from the Google Books text, using statistical models that are specifically adapted to the historical text found in these books. The new edition will facilitate the study of linguistic trends, especially those related to the evolution of syntax. We present our initial findings from the annotated Ngrams in the new edition, including studies of the change in various words&#x27; primary parts of speech over time, and to find the words most closely related to a given set of topics.","abstract_has_math":false,"creators":["Lin, Yuri, M. Eng. Massachusetts Institute of Technology"],"institution":"Massachusetts Institute of Technology","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Dept. of Electrical Engineering and Computer Science.","school":null,"contributors":[],"advisors":["Dorothy Curtis and Slav Petrov."],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012","date_published":"2012","updated_at":"2026-07-22T22:21:11Z","subjects":["Electrical Engineering and Computer Science."],"languages":["eng"],"rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"rights_urls":["http://dspace.mit.edu/handle/1721.1/7582"],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1721.1/77501","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Dorothy Curtis and Slav Petrov."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Dept. of Electrical Engineering and Computer Science."]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Massachusetts Institute of Technology. Dept. of Electrical Engineering and Computer Science."]},{"key":"dc:creator","label":"Author","values":["Lin, Yuri, M. Eng. Massachusetts Institute of Technology"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2013-03-01T15:12:48Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2013-03-01T15:12:48Z"]},{"key":"dc:date.issued","label":"Date","values":["2012"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Electrical Engineering and Computer Science."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://dspace.mit.edu/handle/1721.1/7582"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1721.1/77501"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (M. Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2012.","Cataloged from PDF version of thesis.","Includes bibliographical references (p. 101-102)."]},{"key":"dc:description.abstract","label":"Abstract","values":["In this thesis, we present a new edition of the Google Books Ngram Corpus, describing how often words and phrases were used over a period of five centuries, in eight languages; it aggregates data from 6% of all books ever published. This new edition introduces syntactic annotations: words are tagged with their part-of-speech, and head-modifier dependency relationships are recorded. We generate these annotations automatically from the Google Books text, using statistical models that are specifically adapted to the historical text found in these books. The new edition will facilitate the study of linguistic trends, especially those related to the evolution of syntax. We present our initial findings from the annotated Ngrams in the new edition, including studies of the change in various words' primary parts of speech over time, and to find the words most closely related to a given set of topics."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["M.Eng."]},{"key":"dc:title","label":"Title","values":["Syntactically annotated Ngrams for Google Books"]}]}],"canonical_facts":{"dc:contributor.advisor":["Dorothy Curtis and Slav Petrov."],"dc:contributor.department":["Massachusetts Institute of Technology. Dept. of Electrical Engineering and Computer Science."],"dc:contributor.other":["Massachusetts Institute of Technology. Dept. of Electrical Engineering and Computer Science."],"dc:creator":["Lin, Yuri, M. Eng. Massachusetts Institute of Technology"],"dc:date.accessioned":["2013-03-01T15:12:48Z"],"dc:date.available":["2013-03-01T15:12:48Z"],"dc:date.issued":["2012"],"dc:description":["Thesis (M. Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2012.","Cataloged from PDF version of thesis.","Includes bibliographical references (p. 101-102)."],"dc:description.abstract":["In this thesis, we present a new edition of the Google Books Ngram Corpus, describing how often words and phrases were used over a period of five centuries, in eight languages; it aggregates data from 6% of all books ever published. This new edition introduces syntactic annotations: words are tagged with their part-of-speech, and head-modifier dependency relationships are recorded. We generate these annotations automatically from the Google Books text, using statistical models that are specifically adapted to the historical text found in these books. The new edition will facilitate the study of linguistic trends, especially those related to the evolution of syntax. We present our initial findings from the annotated Ngrams in the new edition, including studies of the change in various words' primary parts of speech over time, and to find the words most closely related to a given set of topics."],"dc:description.degree":["M.Eng."],"dc:identifier.uri":["http://hdl.handle.net/1721.1/77501"],"dc:language.iso":["eng"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"dc:rights.uri":["http://dspace.mit.edu/handle/1721.1/7582"],"dc:subject":["Electrical Engineering and Computer Science."],"dc:title":["Syntactically annotated Ngrams for Google Books"],"dc:type":["Thesis"]},"updated_at":"2026-07-22T22:21:11Z"}