{"id":{"repo_id":"buffalo","oai_identifier":"oai:ubir.buffalo.edu:10477/84076"},"canonical_url":"https://search.dev.ndltd.org/etd/buffalo/oai:ubir.buffalo.edu:10477/84076","repository":{"repo_id":"buffalo","name":"Buffalo","base_url":"https://ubir.buffalo.edu/oai/request"},"display":{"title":"Extraction and Analysis of Self Identity in Twitter Biographies","abstract":"M.S.","abstract_html":"M.S.","abstract_has_math":false,"creators":["Pathak, Arjunil"],"institution":"State University of New York at Buffalo","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Joseph, Kenneth","Computer Science and Engineering"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-06-21T15:47:41Z","date_published":"2022-06-21T15:47:41Z","updated_at":"2026-07-27T19:05:30Z","subjects":["computer science","computer engineering"],"languages":["eng"],"rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/10477/84076","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Joseph, Kenneth","Computer Science and Engineering"]},{"key":"dc:creator","label":"Author","values":["Pathak, Arjunil"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-06-21T15:47:41Z","2020"]},{"key":"dc:publisher","label":"Institution","values":["State University of New York at Buffalo"]},{"key":"dc:type","label":"Dc Type","values":["Text","Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["computer science","computer engineering"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/10477/84076"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["M.S.","An identity is a word or phrase that is used to refer to a particular social group, category, or role which holds a particular set of culturally-constructed meanings, or stereotypes. Most scholarship assumes that a well defined set of identities exist, which can then be used to study the implications of holding a particular identity. However, the identities we express for ourselves and others are often ill-defined. New methods and data are needed to acknowledge complex identities that manifest in the ever evolving domain of social media and to study the downstream impacts of using them to define oneself or another person.This thesis aims to extract and analyze the identities commonly used in social media. For this purpose, ten million biographies of Twitter users are used as the dataset and unigrams are extracted that pertain to identities and preferences. A similarity based network of identities is then generated to identify, study and explain clusters of related identities.The results indicate that a wide range of identities are used on social media, and that these identities cluster together in previously unexplored ways. Further, overlaps between these clusters signify the intricate nature and use of different expressions of identity in ambiguous contexts present in social media. Furthermore, the identity clusters indicate an expected trend of people choosing to identify with activities, attributes, occupations and preferences that they appear to define themselves by. This fragmentation of self identity is in accordance with prevailing literature in the field, but extends our understanding of how the self is presented in the context of social media.","**To request an accessible version of the file(s) associated with this item, contact library@buffalo.edu. Please include the item's persistent URL [http://hdl.handle.net/. . .] in your request.**"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Extraction and Analysis of Self Identity in Twitter Biographies"]}]}],"canonical_facts":{"dc:contributor":["Joseph, Kenneth","Computer Science and Engineering"],"dc:creator":["Pathak, Arjunil"],"dc:date":["2022-06-21T15:47:41Z","2020"],"dc:description":["M.S.","An identity is a word or phrase that is used to refer to a particular social group, category, or role which holds a particular set of culturally-constructed meanings, or stereotypes. Most scholarship assumes that a well defined set of identities exist, which can then be used to study the implications of holding a particular identity. However, the identities we express for ourselves and others are often ill-defined. New methods and data are needed to acknowledge complex identities that manifest in the ever evolving domain of social media and to study the downstream impacts of using them to define oneself or another person.This thesis aims to extract and analyze the identities commonly used in social media. For this purpose, ten million biographies of Twitter users are used as the dataset and unigrams are extracted that pertain to identities and preferences. A similarity based network of identities is then generated to identify, study and explain clusters of related identities.The results indicate that a wide range of identities are used on social media, and that these identities cluster together in previously unexplored ways. Further, overlaps between these clusters signify the intricate nature and use of different expressions of identity in ambiguous contexts present in social media. Furthermore, the identity clusters indicate an expected trend of people choosing to identify with activities, attributes, occupations and preferences that they appear to define themselves by. This fragmentation of self identity is in accordance with prevailing literature in the field, but extends our understanding of how the self is presented in the context of social media.","**To request an accessible version of the file(s) associated with this item, contact library@buffalo.edu. Please include the item's persistent URL [http://hdl.handle.net/. . .] in your request.**"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/10477/84076"],"dc:language":["eng"],"dc:publisher":["State University of New York at Buffalo"],"dc:rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"dc:subject":["computer science","computer engineering"],"dc:title":["Extraction and Analysis of Self Identity in Twitter Biographies"],"dc:type":["Text","Thesis"]},"updated_at":"2026-07-27T19:05:30Z"}