{"id":{"repo_id":"gmu","oai_identifier":"oai:MARS:1920/14391"},"canonical_url":"https://search.dev.ndltd.org/etd/gmu/oai:MARS:1920/14391","repository":{"repo_id":"gmu","name":"George Mason University","base_url":"https://mars.gmu.edu/server/oai/request"},"display":{"title":"Analysis Of Multiple Datasets for COVID-19 Prediction: A Machine Learning Approach","abstract":"SARS-CoV-2, the infectious disease known as COVID-19 resulted in one of the largest pandemics in modern history. Around the world, numerous datasets have been collected to aid researchers in understanding the disease and its epidemiology to allow for quick and efficient diagnosis of COVID-19. This dissertation attempts to investigate some of the datasets and their utility in construction of models for distinguishing COVID-19 from other conditions, such as influenza, based on symptoms. For this purpose, six datasets were investigated. These data sets include the COVID-19 Journal and Screening collected at George Mason University; the daily HealthCheck data that includes self-reported symptoms within the university population; the COVID-CARE Phase I and Phase II datasets which were collected and analyzed through collaboration between George Mason University and Virginia Commonwealth University; and the Case Surveillance Restricted Use Detailed Data from the U.S. Centers for Disease Control and Prevention (CDC) which represents a national database of records of individuals who tested positive for SARS-CoV-2. It is important to note that all of the investigated datasets are very small with the exception of the CDC dataset, Several machine learning methods were applied to the datasets to construct classification models for diagnosing COVID-19 based on symptoms and distinguishing it from other conditions. Temporal representations of data were investigated to see whether inclusion of information when symptoms occur helps in the classification task. The three basic representations include: Flat, in which no time component is added; First Day versus Delayed, in which symptoms present on the first day of illness are examined separately from those that happened later; and Sequential that encodes sequences in which symptoms present in an individual. Several variants of these representations were also studied. This dissertation is concluded by a simulation study that answers a hypothetical question of what could have been done with the data at different times during the pandemic. A week-by-week simulation was performed, and the results were compared.","abstract_html":"SARS-CoV-2, the infectious disease known as COVID-19 resulted in one of the largest pandemics in modern history. Around the world, numerous datasets have been collected to aid researchers in understanding the disease and its epidemiology to allow for quick and efficient diagnosis of COVID-19. This dissertation attempts to investigate some of the datasets and their utility in construction of models for distinguishing COVID-19 from other conditions, such as influenza, based on symptoms. For this purpose, six datasets were investigated. These data sets include the COVID-19 Journal and Screening collected at George Mason University; the daily HealthCheck data that includes self-reported symptoms within the university population; the COVID-CARE Phase I and Phase II datasets which were collected and analyzed through collaboration between George Mason University and Virginia Commonwealth University; and the Case Surveillance Restricted Use Detailed Data from the U.S. Centers for Disease Control and Prevention (CDC) which represents a national database of records of individuals who tested positive for SARS-CoV-2. It is important to note that all of the investigated datasets are very small with the exception of the CDC dataset, Several machine learning methods were applied to the datasets to construct classification models for diagnosing COVID-19 based on symptoms and distinguishing it from other conditions. Temporal representations of data were investigated to see whether inclusion of information when symptoms occur helps in the classification task. The three basic representations include: Flat, in which no time component is added; First Day versus Delayed, in which symptoms present on the first day of illness are examined separately from those that happened later; and Sequential that encodes sequences in which symptoms present in an individual. Several variants of these representations were also studied. This dissertation is concluded by a simulation study that answers a hypothetical question of what could have been done with the data at different times during the pandemic. A week-by-week simulation was performed, and the results were compared.","abstract_has_math":false,"creators":["Mobahi, Hedyeh"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024","date_published":"2024","updated_at":"2026-07-27T19:52:12Z","subjects":["COVID-19","Machine Learning","Sequences","Symptoms","Temporal Data"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["hdl:1920/14391"],"render_values":[{"text":"hdl:1920/14391","href":null,"code":true}]}]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["COVID-19","Machine Learning","Sequences","Symptoms","Temporal Data"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["hdl:1920/14391"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.other","label":"Dc Description Other","values":["SARS-CoV-2, the infectious disease known as COVID-19 resulted in one of the largest pandemics in modern history. Around the world, numerous datasets have been collected to aid researchers in understanding the disease and its epidemiology to allow for quick and efficient diagnosis of COVID-19. This dissertation attempts to investigate some of the datasets and their utility in construction of models for distinguishing COVID-19 from other conditions, such as influenza, based on symptoms. For this purpose, six datasets were investigated. These data sets include the COVID-19 Journal and Screening collected at George Mason University; the daily HealthCheck data that includes self-reported symptoms within the university population; the COVID-CARE Phase I and Phase II datasets which were collected and analyzed through collaboration between George Mason University and Virginia Commonwealth University; and the Case Surveillance Restricted Use Detailed Data from the U.S. Centers for Disease Control and Prevention (CDC) which represents a national database of records of individuals who tested positive for SARS-CoV-2. It is important to note that all of the investigated datasets are very small with the exception of the CDC dataset, Several machine learning methods were applied to the datasets to construct classification models for diagnosing COVID-19 based on symptoms and distinguishing it from other conditions. Temporal representations of data were investigated to see whether inclusion of information when symptoms occur helps in the classification task. The three basic representations include: Flat, in which no time component is added; First Day versus Delayed, in which symptoms present on the first day of illness are examined separately from those that happened later; and Sequential that encodes sequences in which symptoms present in an individual. Several variants of these representations were also studied. This dissertation is concluded by a simulation study that answers a hypothetical question of what could have been done with the data at different times during the pandemic. A week-by-week simulation was performed, and the results were compared."]},{"key":"dc:title","label":"Title","values":["Analysis Of Multiple Datasets for COVID-19 Prediction: A Machine Learning Approach"]}]}],"canonical_facts":{"dc:date.issued":["2024"],"dc:description.other":["SARS-CoV-2, the infectious disease known as COVID-19 resulted in one of the largest pandemics in modern history. Around the world, numerous datasets have been collected to aid researchers in understanding the disease and its epidemiology to allow for quick and efficient diagnosis of COVID-19. This dissertation attempts to investigate some of the datasets and their utility in construction of models for distinguishing COVID-19 from other conditions, such as influenza, based on symptoms. For this purpose, six datasets were investigated. These data sets include the COVID-19 Journal and Screening collected at George Mason University; the daily HealthCheck data that includes self-reported symptoms within the university population; the COVID-CARE Phase I and Phase II datasets which were collected and analyzed through collaboration between George Mason University and Virginia Commonwealth University; and the Case Surveillance Restricted Use Detailed Data from the U.S. Centers for Disease Control and Prevention (CDC) which represents a national database of records of individuals who tested positive for SARS-CoV-2. It is important to note that all of the investigated datasets are very small with the exception of the CDC dataset, Several machine learning methods were applied to the datasets to construct classification models for diagnosing COVID-19 based on symptoms and distinguishing it from other conditions. Temporal representations of data were investigated to see whether inclusion of information when symptoms occur helps in the classification task. The three basic representations include: Flat, in which no time component is added; First Day versus Delayed, in which symptoms present on the first day of illness are examined separately from those that happened later; and Sequential that encodes sequences in which symptoms present in an individual. Several variants of these representations were also studied. This dissertation is concluded by a simulation study that answers a hypothetical question of what could have been done with the data at different times during the pandemic. A week-by-week simulation was performed, and the results were compared."],"dc:identifier":["hdl:1920/14391"],"dc:subject":["COVID-19","Machine Learning","Sequences","Symptoms","Temporal Data"],"dc:title":["Analysis Of Multiple Datasets for COVID-19 Prediction: A Machine Learning Approach"],"dc:type":["Dissertation"]},"updated_at":"2026-07-27T19:52:12Z"}