{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/69808"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/69808","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"A taxonomy of situated language in natural contexts","abstract":"This thesis develops a multi-modal dataset consisting of transcribed speech along with the locations in which that speech took place. Speech with location attached is called situated language, and is represented here as spatial distributions, or two-dimensional histograms over locations in a home. These histograms are organized in the form of a taxonomy, where one can explore, compare, and contrast various slices along several axes of interest. This dataset is derived from raw data collected as part of the Human Speechome Project, and consists of semi-automatically transcribed spoken language and time-aligned overhead video collected over 15 months in a typical home environment. As part of this thesis, the vocabulary of the child before the age of two is derived from transcription, as well as the age at which the child first produced each of the 658 words in his vocabulary. Locations are derived using an efficient tracking algorithm, developed as part of this thesis, called 2C. This system maintains high accuracy when compared to similar systems, while dramatically reducing processing time, an essential feature when processing a corpus of this size. Spatial distributions are produced for many different cuts through the data, including temporal segments (i.e. morning, day, and night), speaker identities (i.e. mother, father, child), and linguistic content (i.e. per-word, aggregate by word type). Several visualization types and statistics are developed, which prove useful for organizing and exploring the dataset. It will then be shown that spatial distributions contain a wealth of information, and that this information can be exploited in various ways to derive meaningful insights and numerical results from the data.","abstract_html":"This thesis develops a multi-modal dataset consisting of transcribed speech along with the locations in which that speech took place. Speech with location attached is called situated language, and is represented here as spatial distributions, or two-dimensional histograms over locations in a home. These histograms are organized in the form of a taxonomy, where one can explore, compare, and contrast various slices along several axes of interest. This dataset is derived from raw data collected as part of the Human Speechome Project, and consists of semi-automatically transcribed spoken language and time-aligned overhead video collected over 15 months in a typical home environment. As part of this thesis, the vocabulary of the child before the age of two is derived from transcription, as well as the age at which the child first produced each of the 658 words in his vocabulary. Locations are derived using an efficient tracking algorithm, developed as part of this thesis, called 2C. This system maintains high accuracy when compared to similar systems, while dramatically reducing processing time, an essential feature when processing a corpus of this size. Spatial distributions are produced for many different cuts through the data, including temporal segments (i.e. morning, day, and night), speaker identities (i.e. mother, father, child), and linguistic content (i.e. per-word, aggregate by word type). Several visualization types and statistics are developed, which prove useful for organizing and exploring the dataset. It will then be shown that spatial distributions contain a wealth of information, and that this information can be exploited in various ways to derive meaningful insights and numerical results from the data.","abstract_has_math":false,"creators":["Shaw, George Macaulay"],"institution":"Massachusetts Institute of Technology","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Dept. of Architecture. Program in Media Arts and Sciences.","school":null,"contributors":[],"advisors":["Deb Roy."],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011","date_published":"2011","updated_at":"2026-07-22T22:21:44Z","subjects":["Architecture. Program in Media Arts and Sciences."],"languages":["eng"],"rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"rights_urls":["http://dspace.mit.edu/handle/1721.1/7582"],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1721.1/69808","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Deb Roy."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Dept. of Architecture. Program in Media Arts and Sciences."]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Massachusetts Institute of Technology. Dept. of Architecture. Program in Media Arts and Sciences."]},{"key":"dc:creator","label":"Author","values":["Shaw, George Macaulay"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2012-03-16T16:04:54Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2012-03-16T16:04:54Z"]},{"key":"dc:date.issued","label":"Date","values":["2011"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Architecture. Program in Media Arts and Sciences."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://dspace.mit.edu/handle/1721.1/7582"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1721.1/69808"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (S.M.)--Massachusetts Institute of Technology, School of Architecture and Planning, Program in Media Arts and Sciences, 2011.","Cataloged from PDF version of thesis.","Includes bibliographical references (p. 161-163)."]},{"key":"dc:description.abstract","label":"Abstract","values":["This thesis develops a multi-modal dataset consisting of transcribed speech along with the locations in which that speech took place. Speech with location attached is called situated language, and is represented here as spatial distributions, or two-dimensional histograms over locations in a home. These histograms are organized in the form of a taxonomy, where one can explore, compare, and contrast various slices along several axes of interest. This dataset is derived from raw data collected as part of the Human Speechome Project, and consists of semi-automatically transcribed spoken language and time-aligned overhead video collected over 15 months in a typical home environment. As part of this thesis, the vocabulary of the child before the age of two is derived from transcription, as well as the age at which the child first produced each of the 658 words in his vocabulary. Locations are derived using an efficient tracking algorithm, developed as part of this thesis, called 2C. This system maintains high accuracy when compared to similar systems, while dramatically reducing processing time, an essential feature when processing a corpus of this size. Spatial distributions are produced for many different cuts through the data, including temporal segments (i.e. morning, day, and night), speaker identities (i.e. mother, father, child), and linguistic content (i.e. per-word, aggregate by word type). Several visualization types and statistics are developed, which prove useful for organizing and exploring the dataset. It will then be shown that spatial distributions contain a wealth of information, and that this information can be exploited in various ways to derive meaningful insights and numerical results from the data."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["S.M."]},{"key":"dc:title","label":"Title","values":["A taxonomy of situated language in natural contexts"]}]}],"canonical_facts":{"dc:contributor.advisor":["Deb Roy."],"dc:contributor.department":["Massachusetts Institute of Technology. Dept. of Architecture. Program in Media Arts and Sciences."],"dc:contributor.other":["Massachusetts Institute of Technology. Dept. of Architecture. Program in Media Arts and Sciences."],"dc:creator":["Shaw, George Macaulay"],"dc:date.accessioned":["2012-03-16T16:04:54Z"],"dc:date.available":["2012-03-16T16:04:54Z"],"dc:date.issued":["2011"],"dc:description":["Thesis (S.M.)--Massachusetts Institute of Technology, School of Architecture and Planning, Program in Media Arts and Sciences, 2011.","Cataloged from PDF version of thesis.","Includes bibliographical references (p. 161-163)."],"dc:description.abstract":["This thesis develops a multi-modal dataset consisting of transcribed speech along with the locations in which that speech took place. Speech with location attached is called situated language, and is represented here as spatial distributions, or two-dimensional histograms over locations in a home. These histograms are organized in the form of a taxonomy, where one can explore, compare, and contrast various slices along several axes of interest. This dataset is derived from raw data collected as part of the Human Speechome Project, and consists of semi-automatically transcribed spoken language and time-aligned overhead video collected over 15 months in a typical home environment. As part of this thesis, the vocabulary of the child before the age of two is derived from transcription, as well as the age at which the child first produced each of the 658 words in his vocabulary. Locations are derived using an efficient tracking algorithm, developed as part of this thesis, called 2C. This system maintains high accuracy when compared to similar systems, while dramatically reducing processing time, an essential feature when processing a corpus of this size. Spatial distributions are produced for many different cuts through the data, including temporal segments (i.e. morning, day, and night), speaker identities (i.e. mother, father, child), and linguistic content (i.e. per-word, aggregate by word type). Several visualization types and statistics are developed, which prove useful for organizing and exploring the dataset. It will then be shown that spatial distributions contain a wealth of information, and that this information can be exploited in various ways to derive meaningful insights and numerical results from the data."],"dc:description.degree":["S.M."],"dc:identifier.uri":["http://hdl.handle.net/1721.1/69808"],"dc:language.iso":["eng"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"dc:rights.uri":["http://dspace.mit.edu/handle/1721.1/7582"],"dc:subject":["Architecture. Program in Media Arts and Sciences."],"dc:title":["A taxonomy of situated language in natural contexts"],"dc:type":["Thesis"]},"updated_at":"2026-07-22T22:21:44Z"}