{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/92965"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/92965","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"\"Knowing a thing is \"\"a thing\"\": The use of acoustic features in multiword expression extraction\"","abstract":"\"Speakers of a language need to have complex linguistic representations for speaking, often on the level of non-literal, idiomatic expressions like black sheep. Typically, datasets of these so-called multiword expressions come from hand-crafted ontologies or lexicons, because identifying expressions like these in an unsupervised manner is still an unsolved problem in natural language processing. In this thesis I demonstrate that prosodic features, which are helpful in parsing syntax and interpreting meaning, can also be used to identify multiword expressions. To do this, I extracted noun phrases from the Buckeye corpus, which contains spontaneous spoken language, and matched these noun phrases to page titles in Wikipedia, a massive, freely available encyclopedic ontology of entities and phenomena. By incorporating prosodic features into a model that distinguishes between multiword expressions that are found in Wikipedia titles and those that are not, we see increases in classifier performance that suggests that prosodic cues can help with the automatic extraction of multiword expressions from spontaneous speech, helping models and potentially listeners decide whether something is \"\"a thing\"\" or not.\"","abstract_html":"&quot;Speakers of a language need to have complex linguistic representations for speaking, often on the level of non-literal, idiomatic expressions like black sheep. Typically, datasets of these so-called multiword expressions come from hand-crafted ontologies or lexicons, because identifying expressions like these in an unsupervised manner is still an unsolved problem in natural language processing. In this thesis I demonstrate that prosodic features, which are helpful in parsing syntax and interpreting meaning, can also be used to identify multiword expressions. To do this, I extracted noun phrases from the Buckeye corpus, which contains spontaneous spoken language, and matched these noun phrases to page titles in Wikipedia, a massive, freely available encyclopedic ontology of entities and phenomena. By incorporating prosodic features into a model that distinguishes between multiword expressions that are found in Wikipedia titles and those that are not, we see increases in classifier performance that suggests that prosodic cues can help with the automatic extraction of multiword expressions from spontaneous speech, helping models and potentially listeners decide whether something is &quot;&quot;a thing&quot;&quot; or not.&quot;","abstract_has_math":false,"creators":["Jacobs, Cassandra L."],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Fleck, Margaret"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-11-10T18:27:56Z","date_published":"2016-11-10T18:27:56Z","updated_at":"2026-07-22T22:26:35Z","subjects":["Collocations","Phrases","Speech processing","Language models"],"languages":["en"],"rights":["Copyright 2016 Cassandra Jacobs"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/92965","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Fleck, Margaret"]},{"key":"dc:creator","label":"Author","values":["Jacobs, Cassandra L."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-11-10T18:27:56Z","2018-11-11T10:15:28Z","2016-07-19","2016-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Collocations","Phrases","Speech processing","Language models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2016 Cassandra Jacobs"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/92965"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"Speakers of a language need to have complex linguistic representations for speaking, often on the level of non-literal, idiomatic expressions like black sheep. Typically, datasets of these so-called multiword expressions come from hand-crafted ontologies or lexicons, because identifying expressions like these in an unsupervised manner is still an unsolved problem in natural language processing. In this thesis I demonstrate that prosodic features, which are helpful in parsing syntax and interpreting meaning, can also be used to identify multiword expressions. To do this, I extracted noun phrases from the Buckeye corpus, which contains spontaneous spoken language, and matched these noun phrases to page titles in Wikipedia, a massive, freely available encyclopedic ontology of entities and phenomena. By incorporating prosodic features into a model that distinguishes between multiword expressions that are found in Wikipedia titles and those that are not, we see increases in classifier performance that suggests that prosodic cues can help with the automatic extraction of multiword expressions from spontaneous speech, helping models and potentially listeners decide whether something is \"\"a thing\"\" or not.\"","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2018-08-01","The student, Cassandra Jacobs, accepted the attached license on 2016-07-18 at 16:10.","The student, Cassandra Jacobs, submitted this Thesis for approval on 2016-07-18 at 16:13.","This Thesis was approved for publication on 2016-07-19 at 16:37.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9997 on 2016-11-10 at 12:20:58","Made available in DSpace on 2016-11-10T18:27:56Z (GMT). No. of bitstreams: 2 JACOBS-THESIS-2016.pdf: 762704 bytes, checksum: c34b700b7fd5b9c0d8dd7948993f80cd (MD5) LICENSE.txt: 4213 bytes, checksum: d53c4637ed780ec7f29f9bd624ff8f2e (MD5) Previous issue date: 2016-07-19","Embargo set by: Seth Robbins for item 95385 Lift date: 2018-11-10T18:28:02Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 95385 on 2018-11-11T10:15:28Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["\"Knowing a thing is \"\"a thing\"\": The use of acoustic features in multiword expression extraction\""]}]}],"canonical_facts":{"dc:contributor":["Fleck, Margaret"],"dc:creator":["Jacobs, Cassandra L."],"dc:date":["2016-11-10T18:27:56Z","2018-11-11T10:15:28Z","2016-07-19","2016-08"],"dc:description":["\"Speakers of a language need to have complex linguistic representations for speaking, often on the level of non-literal, idiomatic expressions like black sheep. Typically, datasets of these so-called multiword expressions come from hand-crafted ontologies or lexicons, because identifying expressions like these in an unsupervised manner is still an unsolved problem in natural language processing. In this thesis I demonstrate that prosodic features, which are helpful in parsing syntax and interpreting meaning, can also be used to identify multiword expressions. To do this, I extracted noun phrases from the Buckeye corpus, which contains spontaneous spoken language, and matched these noun phrases to page titles in Wikipedia, a massive, freely available encyclopedic ontology of entities and phenomena. By incorporating prosodic features into a model that distinguishes between multiword expressions that are found in Wikipedia titles and those that are not, we see increases in classifier performance that suggests that prosodic cues can help with the automatic extraction of multiword expressions from spontaneous speech, helping models and potentially listeners decide whether something is \"\"a thing\"\" or not.\"","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2018-08-01","The student, Cassandra Jacobs, accepted the attached license on 2016-07-18 at 16:10.","The student, Cassandra Jacobs, submitted this Thesis for approval on 2016-07-18 at 16:13.","This Thesis was approved for publication on 2016-07-19 at 16:37.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9997 on 2016-11-10 at 12:20:58","Made available in DSpace on 2016-11-10T18:27:56Z (GMT). No. of bitstreams: 2 JACOBS-THESIS-2016.pdf: 762704 bytes, checksum: c34b700b7fd5b9c0d8dd7948993f80cd (MD5) LICENSE.txt: 4213 bytes, checksum: d53c4637ed780ec7f29f9bd624ff8f2e (MD5) Previous issue date: 2016-07-19","Embargo set by: Seth Robbins for item 95385 Lift date: 2018-11-10T18:28:02Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 95385 on 2018-11-11T10:15:28Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/92965"],"dc:language":["en"],"dc:rights":["Copyright 2016 Cassandra Jacobs"],"dc:subject":["Collocations","Phrases","Speech processing","Language models"],"dc:title":["\"Knowing a thing is \"\"a thing\"\": The use of acoustic features in multiword expression extraction\""],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:35Z"}