{"id":{"repo_id":"regina","oai_identifier":"oai:uregina.scholaris.ca:10294/6549"},"canonical_url":"https://search.dev.ndltd.org/etd/regina/oai:uregina.scholaris.ca:10294/6549","repository":{"repo_id":"regina","name":"University of Regina","base_url":"https://uregina.scholaris.ca/server/oai/request"},"display":{"title":"A Combinatorial Tweet Clustering Methodology Utilizing Inter and Intra Cosine Similarity","abstract":"Data mining techniques are well known and are often used to analyze and explore datasets for meaningful information. Social media, such as Twitter, has emerged as a source of data where millions of tweets are generated everyday. They include tweets from individuals who share thoughts, commentary and their feelings about a wide variety of subjects. Social media also attracts marketers and businesses for the purpose of advertising, brand imaging and getting feedback from users. Twitter’s significant popularity and mass usage has resulted in a very large dataset where virtually any subject that is queried from the Twitter API may return a vast number of tweets. As a result, these tweets can be related to several distinctly different categories. Data mining a large amount of tweets to classify them into meaningful categories is a challenging task because of the often informal language used, the inclusion of URL links, spam and other irrelevant information. This thesis presents a combinatorial hierarchical clustering methodology that categorizes tweets into meaningful clusters by utilizing inter and intra cluster cosine similarity. Cosine similarity is the degree of relativity between two vectors. This thesis proposes a “Combinatorial Hierarchical Clustering Methodology” as a combination of both agglomerative (Bottom-Up) and divisive (Top-Down) hierarchical clustering approaches that attempts to maximize the clustering effectiveness and quality. The proposed methodology sub-categorizes, divides and combines clusters through an iterative process to help make sorted categories more meaningful. In addition, this approach does not require a-priori information about the numbers of clusters to be formed but rather forms clusters dynamically based on their determined similarity.","abstract_html":"Data mining techniques are well known and are often used to analyze and explore datasets for meaningful information. Social media, such as Twitter, has emerged as a source of data where millions of tweets are generated everyday. They include tweets from individuals who share thoughts, commentary and their feelings about a wide variety of subjects. Social media also attracts marketers and businesses for the purpose of advertising, brand imaging and getting feedback from users. Twitter’s significant popularity and mass usage has resulted in a very large dataset where virtually any subject that is queried from the Twitter API may return a vast number of tweets. As a result, these tweets can be related to several distinctly different categories. Data mining a large amount of tweets to classify them into meaningful categories is a challenging task because of the often informal language used, the inclusion of URL links, spam and other irrelevant information. This thesis presents a combinatorial hierarchical clustering methodology that categorizes tweets into meaningful clusters by utilizing inter and intra cluster cosine similarity. Cosine similarity is the degree of relativity between two vectors. This thesis proposes a “Combinatorial Hierarchical Clustering Methodology” as a combination of both agglomerative (Bottom-Up) and divisive (Top-Down) hierarchical clustering approaches that attempts to maximize the clustering effectiveness and quality. The proposed methodology sub-categorizes, divides and combines clusters through an iterative process to help make sorted categories more meaningful. In addition, this approach does not require a-priori information about the numbers of clusters to be formed but rather forms clusters dynamically based on their determined similarity.","abstract_has_math":false,"creators":["Kaur, Navneet"],"institution":"Faculty of Graduate Studies and Research, University of Regina","degree_name":"Master of Applied Science (MASc)","degree_level":"Master&apos;s","degree_discipline":"Engineering - Software Systems","degree_department":null,"school":null,"contributors":[],"advisors":["Gelowitz, Craig"],"committee_chairs":[],"committee_members":["Benedicenti, Luigi","El-Darieby, Mohamed"],"year":2015,"date_issued":"2015-07","date_published":"2015-07","updated_at":"2026-07-24T04:03:39Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.82465/4493"],"render_values":[{"text":"https://doi.org/10.82465/4493","href":"https://doi.org/10.82465/4493","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10294/6549","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Gelowitz, Craig"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Benedicenti, Luigi","El-Darieby, Mohamed"]},{"key":"dc:creator","label":"Author","values":["Kaur, Navneet"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2015-12-22T19:18:41Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2015-12-22T19:18:41Z"]},{"key":"dc:date.issued","label":"Date","values":["2015-07"]},{"key":"dc:publisher","label":"Institution","values":["Faculty of Graduate Studies and Research, University of Regina"]},{"key":"dc:type","label":"Dc Type","values":["master thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Engineering - Software Systems"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Master&apos;s"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Applied Science (MASc)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Faculty of Graduate Studies and Research, University of Regina"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.82465/4493"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10294/6549"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["A Thesis Submitted to the Faculty of Graduate Studies and Research In Partial Fulfillment of the Requirements for the Degree of Master of Applied Science in Software Systems Engineering, University of Regina. xiii, 88 p."]},{"key":"dc:description.abstract","label":"Abstract","values":["Data mining techniques are well known and are often used to analyze and explore datasets for meaningful information. Social media, such as Twitter, has emerged as a source of data where millions of tweets are generated everyday. They include tweets from individuals who share thoughts, commentary and their feelings about a wide variety of subjects. Social media also attracts marketers and businesses for the purpose of advertising, brand imaging and getting feedback from users. Twitter’s significant popularity and mass usage has resulted in a very large dataset where virtually any subject that is queried from the Twitter API may return a vast number of tweets. As a result, these tweets can be related to several distinctly different categories. Data mining a large amount of tweets to classify them into meaningful categories is a challenging task because of the often informal language used, the inclusion of URL links, spam and other irrelevant information. This thesis presents a combinatorial hierarchical clustering methodology that categorizes tweets into meaningful clusters by utilizing inter and intra cluster cosine similarity. Cosine similarity is the degree of relativity between two vectors. This thesis proposes a “Combinatorial Hierarchical Clustering Methodology” as a combination of both agglomerative (Bottom-Up) and divisive (Top-Down) hierarchical clustering approaches that attempts to maximize the clustering effectiveness and quality. The proposed methodology sub-categorizes, divides and combines clusters through an iterative process to help make sorted categories more meaningful. In addition, this approach does not require a-priori information about the numbers of clusters to be formed but rather forms clusters dynamically based on their determined similarity."]},{"key":"dc:title","label":"Title","values":["A Combinatorial Tweet Clustering Methodology Utilizing Inter and Intra Cosine Similarity"]}]}],"canonical_facts":{"dc:contributor.advisor":["Gelowitz, Craig"],"dc:contributor.committeemember":["Benedicenti, Luigi","El-Darieby, Mohamed"],"dc:creator":["Kaur, Navneet"],"dc:date.accessioned":["2015-12-22T19:18:41Z"],"dc:date.available":["2015-12-22T19:18:41Z"],"dc:date.issued":["2015-07"],"dc:description":["A Thesis Submitted to the Faculty of Graduate Studies and Research In Partial Fulfillment of the Requirements for the Degree of Master of Applied Science in Software Systems Engineering, University of Regina. xiii, 88 p."],"dc:description.abstract":["Data mining techniques are well known and are often used to analyze and explore datasets for meaningful information. Social media, such as Twitter, has emerged as a source of data where millions of tweets are generated everyday. They include tweets from individuals who share thoughts, commentary and their feelings about a wide variety of subjects. Social media also attracts marketers and businesses for the purpose of advertising, brand imaging and getting feedback from users. Twitter’s significant popularity and mass usage has resulted in a very large dataset where virtually any subject that is queried from the Twitter API may return a vast number of tweets. As a result, these tweets can be related to several distinctly different categories. Data mining a large amount of tweets to classify them into meaningful categories is a challenging task because of the often informal language used, the inclusion of URL links, spam and other irrelevant information. This thesis presents a combinatorial hierarchical clustering methodology that categorizes tweets into meaningful clusters by utilizing inter and intra cluster cosine similarity. Cosine similarity is the degree of relativity between two vectors. This thesis proposes a “Combinatorial Hierarchical Clustering Methodology” as a combination of both agglomerative (Bottom-Up) and divisive (Top-Down) hierarchical clustering approaches that attempts to maximize the clustering effectiveness and quality. The proposed methodology sub-categorizes, divides and combines clusters through an iterative process to help make sorted categories more meaningful. In addition, this approach does not require a-priori information about the numbers of clusters to be formed but rather forms clusters dynamically based on their determined similarity."],"dc:identifier.doi":["https://doi.org/10.82465/4493"],"dc:identifier.uri":["https://hdl.handle.net/10294/6549"],"dc:language.iso":["en"],"dc:publisher":["Faculty of Graduate Studies and Research, University of Regina"],"dc:title":["A Combinatorial Tweet Clustering Methodology Utilizing Inter and Intra Cosine Similarity"],"dc:type":["master thesis"],"thesis:degree_discipline":["Engineering - Software Systems"],"thesis:degree_level":["Master&apos;s"],"thesis:degree_name":["Master of Applied Science (MASc)"],"thesis:institution_name":["Faculty of Graduate Studies and Research, University of Regina"]},"updated_at":"2026-07-24T04:03:39Z"}