{"id":{"repo_id":"claremont","oai_identifier":"oai:scholarship.claremont.edu:cgu_etd-1768"},"canonical_url":"https://search.dev.ndltd.org/etd/claremont/oai:scholarship.claremont.edu:cgu_etd-1768","repository":{"repo_id":"claremont","name":"Claremont Graduate University","base_url":"https://scholarship.claremont.edu/do/oai/"},"display":{"title":"Advancing Internet Viewpoint Diversity: A Novel Algorithm and a Corpus Creation Tool","abstract":"<p>A fundamental requirement for Western democracy is an informed and engaged electorate with access to a wide range of viewpoints. However, concerns have arisen regarding how information technology affects the diversity of viewpoints available. In response to an increasingly polarized society and worries surrounding filter bubbles and algorithmic bias, this research presents a novel tool for constructing internet-based topical corpora and an algorithm tailored for viewpoint detection and the curation of diverse search results.Following a comprehensive exploration of viewpoint diversity through the lenses of mass media, social psychology, and information retrieval, this dissertation presents an approach to operationalize viewpoint diversity rooted in a cross-linguistic discourse analysis model. The proposed viewpoint detection algorithm harnesses context-specific sentence embedding features, sentiment features, and topic features. The algorithm's development and refinement are based on the Internet Argument Corpus, compiled by Abbott and his team in 2016. To assess the efficacy of the viewpoint detection algorithm, we need to construct topical corpora from online sources and identify the viewpoints expressed in these documents. To achieve this, we first develop a big data processing architecture for creating indexed corpora from the Common Crawl web archives. The architecture is instantiated into an automated tool that generates an intelligible topical corpus through a series of steps involving processing, filtering, cleaning, and removing duplicate content. Utilizing this tool, we processed approximately 1.2 billion web pages from the Common Crawl dataset, resulting in four distinct corpora. Each corpus includes around 1,000 relevant documents for a specific topic. The viewpoint diversity algorithm was then used to identify the ten most relevant documents that likely represented opposing stances, resulting in a collection of 20 documents per topic. A group of volunteers independently assessed and assigned viewpoints to each document, and then resolved any disagreements collaboratively. The viewpoint detection algorithm was evaluated against the resulting gold standard. It successfully curated a set of documents with balanced viewpoints for the evolution topic, demonstrating its ability to generalize from the Internet Argument Corpus to open internet documents. However, the algorithm did not generalize well for the abortion and gun control topics. We discuss the reasons behind these discrepancies and suggest potential solutions. In summary, this dissertation contributes to enhancing viewpoint diversity and transparency in online content. It offers valuable insights into the challenges and potential solutions for achieving this goal. Our research not only represents a step toward promoting transparency in managing viewpoints on the internet, but also provides tools for researchers and practitioners to utilize extensive textual data more effectively from the web.</p>","abstract_html":"&lt;p&gt;A fundamental requirement for Western democracy is an informed and engaged electorate with access to a wide range of viewpoints. However, concerns have arisen regarding how information technology affects the diversity of viewpoints available. In response to an increasingly polarized society and worries surrounding filter bubbles and algorithmic bias, this research presents a novel tool for constructing internet-based topical corpora and an algorithm tailored for viewpoint detection and the curation of diverse search results.Following a comprehensive exploration of viewpoint diversity through the lenses of mass media, social psychology, and information retrieval, this dissertation presents an approach to operationalize viewpoint diversity rooted in a cross-linguistic discourse analysis model. The proposed viewpoint detection algorithm harnesses context-specific sentence embedding features, sentiment features, and topic features. The algorithm&#x27;s development and refinement are based on the Internet Argument Corpus, compiled by Abbott and his team in 2016. To assess the efficacy of the viewpoint detection algorithm, we need to construct topical corpora from online sources and identify the viewpoints expressed in these documents. To achieve this, we first develop a big data processing architecture for creating indexed corpora from the Common Crawl web archives. The architecture is instantiated into an automated tool that generates an intelligible topical corpus through a series of steps involving processing, filtering, cleaning, and removing duplicate content. Utilizing this tool, we processed approximately 1.2 billion web pages from the Common Crawl dataset, resulting in four distinct corpora. Each corpus includes around 1,000 relevant documents for a specific topic. The viewpoint diversity algorithm was then used to identify the ten most relevant documents that likely represented opposing stances, resulting in a collection of 20 documents per topic. A group of volunteers independently assessed and assigned viewpoints to each document, and then resolved any disagreements collaboratively. The viewpoint detection algorithm was evaluated against the resulting gold standard. It successfully curated a set of documents with balanced viewpoints for the evolution topic, demonstrating its ability to generalize from the Internet Argument Corpus to open internet documents. However, the algorithm did not generalize well for the abortion and gun control topics. We discuss the reasons behind these discrepancies and suggest potential solutions. In summary, this dissertation contributes to enhancing viewpoint diversity and transparency in online content. It offers valuable insights into the challenges and potential solutions for achieving this goal. Our research not only represents a step toward promoting transparency in managing viewpoints on the internet, but also provides tools for researchers and practitioners to utilize extensive textual data more effectively from the web.&lt;/p&gt;","abstract_has_math":false,"creators":["Harwell, Jeffrey"],"institution":null,"degree_name":"Information Systems and Technology, PhD","degree_level":"Open Access Dissertation","degree_discipline":"Center for Information Systems and Technology","degree_department":null,"school":null,"contributors":["Samir Chatterjee","Gondy Leroy"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-01-01T08:00:00Z","date_published":"2023-01-01T08:00:00Z","updated_at":"2026-07-24T01:41:01Z","subjects":["big data","common crawl","machine learning","natural language processing","viewpoint diversity","Systems Science"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://scholarship.claremont.edu/cgu_etd/746","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Samir Chatterjee","Gondy Leroy"]},{"key":"dc:creator","label":"Author","values":["Harwell, Jeffrey"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2024-03-08T08:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Center for Information Systems and Technology"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Open Access Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Information Systems and Technology, PhD"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["big data","common crawl","machine learning","natural language processing","viewpoint diversity","Systems Science"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://scholarship.claremont.edu/cgu_etd/746"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>A fundamental requirement for Western democracy is an informed and engaged electorate with access to a wide range of viewpoints. However, concerns have arisen regarding how information technology affects the diversity of viewpoints available. In response to an increasingly polarized society and worries surrounding filter bubbles and algorithmic bias, this research presents a novel tool for constructing internet-based topical corpora and an algorithm tailored for viewpoint detection and the curation of diverse search results.Following a comprehensive exploration of viewpoint diversity through the lenses of mass media, social psychology, and information retrieval, this dissertation presents an approach to operationalize viewpoint diversity rooted in a cross-linguistic discourse analysis model. The proposed viewpoint detection algorithm harnesses context-specific sentence embedding features, sentiment features, and topic features. The algorithm's development and refinement are based on the Internet Argument Corpus, compiled by Abbott and his team in 2016. To assess the efficacy of the viewpoint detection algorithm, we need to construct topical corpora from online sources and identify the viewpoints expressed in these documents. To achieve this, we first develop a big data processing architecture for creating indexed corpora from the Common Crawl web archives. The architecture is instantiated into an automated tool that generates an intelligible topical corpus through a series of steps involving processing, filtering, cleaning, and removing duplicate content. Utilizing this tool, we processed approximately 1.2 billion web pages from the Common Crawl dataset, resulting in four distinct corpora. Each corpus includes around 1,000 relevant documents for a specific topic. The viewpoint diversity algorithm was then used to identify the ten most relevant documents that likely represented opposing stances, resulting in a collection of 20 documents per topic. A group of volunteers independently assessed and assigned viewpoints to each document, and then resolved any disagreements collaboratively. The viewpoint detection algorithm was evaluated against the resulting gold standard. It successfully curated a set of documents with balanced viewpoints for the evolution topic, demonstrating its ability to generalize from the Internet Argument Corpus to open internet documents. However, the algorithm did not generalize well for the abortion and gun control topics. We discuss the reasons behind these discrepancies and suggest potential solutions. In summary, this dissertation contributes to enhancing viewpoint diversity and transparency in online content. It offers valuable insights into the challenges and potential solutions for achieving this goal. Our research not only represents a step toward promoting transparency in managing viewpoints on the internet, but also provides tools for researchers and practitioners to utilize extensive textual data more effectively from the web.</p>"]},{"key":"dc:title","label":"Title","values":["Advancing Internet Viewpoint Diversity: A Novel Algorithm and a Corpus Creation Tool"]}]}],"canonical_facts":{"dc:contributor":["Samir Chatterjee","Gondy Leroy"],"dc:creator":["Harwell, Jeffrey"],"dc:date.available":["2024-03-08T08:00:00Z"],"dc:description.abstract":["<p>A fundamental requirement for Western democracy is an informed and engaged electorate with access to a wide range of viewpoints. However, concerns have arisen regarding how information technology affects the diversity of viewpoints available. In response to an increasingly polarized society and worries surrounding filter bubbles and algorithmic bias, this research presents a novel tool for constructing internet-based topical corpora and an algorithm tailored for viewpoint detection and the curation of diverse search results.Following a comprehensive exploration of viewpoint diversity through the lenses of mass media, social psychology, and information retrieval, this dissertation presents an approach to operationalize viewpoint diversity rooted in a cross-linguistic discourse analysis model. The proposed viewpoint detection algorithm harnesses context-specific sentence embedding features, sentiment features, and topic features. The algorithm's development and refinement are based on the Internet Argument Corpus, compiled by Abbott and his team in 2016. To assess the efficacy of the viewpoint detection algorithm, we need to construct topical corpora from online sources and identify the viewpoints expressed in these documents. To achieve this, we first develop a big data processing architecture for creating indexed corpora from the Common Crawl web archives. The architecture is instantiated into an automated tool that generates an intelligible topical corpus through a series of steps involving processing, filtering, cleaning, and removing duplicate content. Utilizing this tool, we processed approximately 1.2 billion web pages from the Common Crawl dataset, resulting in four distinct corpora. Each corpus includes around 1,000 relevant documents for a specific topic. The viewpoint diversity algorithm was then used to identify the ten most relevant documents that likely represented opposing stances, resulting in a collection of 20 documents per topic. A group of volunteers independently assessed and assigned viewpoints to each document, and then resolved any disagreements collaboratively. The viewpoint detection algorithm was evaluated against the resulting gold standard. It successfully curated a set of documents with balanced viewpoints for the evolution topic, demonstrating its ability to generalize from the Internet Argument Corpus to open internet documents. However, the algorithm did not generalize well for the abortion and gun control topics. We discuss the reasons behind these discrepancies and suggest potential solutions. In summary, this dissertation contributes to enhancing viewpoint diversity and transparency in online content. It offers valuable insights into the challenges and potential solutions for achieving this goal. Our research not only represents a step toward promoting transparency in managing viewpoints on the internet, but also provides tools for researchers and practitioners to utilize extensive textual data more effectively from the web.</p>"],"dc:identifier":["https://scholarship.claremont.edu/cgu_etd/746"],"dc:subject":["big data","common crawl","machine learning","natural language processing","viewpoint diversity","Systems Science"],"dc:title":["Advancing Internet Viewpoint Diversity: A Novel Algorithm and a Corpus Creation Tool"],"thesis:degree_discipline":["Center for Information Systems and Technology"],"thesis:degree_level":["Open Access Dissertation"],"thesis:degree_name":["Information Systems and Technology, PhD"]},"updated_at":"2026-07-24T01:41:01Z"}