{"id":{"repo_id":"uthsc","oai_identifier":"oai:digitalcommons.library.tmc.edu:utgsbs_dissertations-2378"},"canonical_url":"https://search.dev.ndltd.org/etd/uthsc/oai:digitalcommons.library.tmc.edu:utgsbs_dissertations-2378","repository":{"repo_id":"uthsc","name":"University of Texas Health Science Center at Houston","base_url":"https://digitalcommons.library.tmc.edu/do/oai/"},"display":{"title":"Addressing the Analytical and Computational Challenges Using Machine Learning in Biomedical Research","abstract":"<p>In the contemporary healthcare field, professionals are confronted with an ever-growing volume of clinical data stored in electronic health records, alongside the genomic data stemming from laboratory experiments. As a response to this deluge of data, the application of machine learning (ML) techniques is gaining popularity since ML techniques have demonstrated an exceptional proficiency in processing big data and deciphering complex nonlinear patterns that are intrinsic to biomedical research.</p> <p>My research leverages ML's capabilities to address the computational challenges spanning diverse areas, including adaptive clinical trial designs, survival analysis, and high-dimensional genetic data analysis. Specifically, Chapter 2 focused on the application of ML in response-adaptive randomization designs. Compared to a traditional equal-randomization design, adaptive randomized trials allocate more patients to the superior treatment arm and increase the overall response rate. In Chapter 3, we proposed a statistical model to impute survival times for censored observations, allowing for a direct application of any ML method in the downstream analysis. We further improve the accuracy of our method using ML regression built on prognostic covariates. In Chapter 4, we developed an artificial neural network-based framework for analyzing longitudinal single-cell RNA sequencing data. Our pipeline achieves: (1) cross-time points cell annotation, (2) detection of novel cell type emerged over time, (3) visualization of cell population evolution, (4) identification of temporal differentially expressed genes. In each chapter, we provide both simulation studies and real-world application results for our methods. In the long run, we aim to lessen the gap between ML and biomedical research, and facilitate healthcare studies.</p>","abstract_html":"&lt;p&gt;In the contemporary healthcare field, professionals are confronted with an ever-growing volume of clinical data stored in electronic health records, alongside the genomic data stemming from laboratory experiments. As a response to this deluge of data, the application of machine learning (ML) techniques is gaining popularity since ML techniques have demonstrated an exceptional proficiency in processing big data and deciphering complex nonlinear patterns that are intrinsic to biomedical research.&lt;/p&gt; &lt;p&gt;My research leverages ML&#x27;s capabilities to address the computational challenges spanning diverse areas, including adaptive clinical trial designs, survival analysis, and high-dimensional genetic data analysis. Specifically, Chapter 2 focused on the application of ML in response-adaptive randomization designs. Compared to a traditional equal-randomization design, adaptive randomized trials allocate more patients to the superior treatment arm and increase the overall response rate. In Chapter 3, we proposed a statistical model to impute survival times for censored observations, allowing for a direct application of any ML method in the downstream analysis. We further improve the accuracy of our method using ML regression built on prognostic covariates. In Chapter 4, we developed an artificial neural network-based framework for analyzing longitudinal single-cell RNA sequencing data. Our pipeline achieves: (1) cross-time points cell annotation, (2) detection of novel cell type emerged over time, (3) visualization of cell population evolution, (4) identification of temporal differentially expressed genes. In each chapter, we provide both simulation studies and real-world application results for our methods. In the long run, we aim to lessen the gap between ML and biomedical research, and facilitate healthcare studies.&lt;/p&gt;","abstract_has_math":false,"creators":["Wang, Yizhuo","<p>0000-0002-1870-0019</p>"],"institution":null,"degree_name":"Doctor of Philosophy (PhD)","degree_level":"Dissertation (PhD)","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Xuelin Huang","Ziyi Li","Bing Z Carter"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-12-01T08:00:00Z","date_published":"2023-12-01T08:00:00Z","updated_at":"2026-07-24T05:48:59Z","subjects":["Biostatistics","machine learning","clinical trial design","adaptive randomization","survival analysis","missing data","bioinformatics","high-dimensional data","scRNA data","Clinical Trials","Data Science","Public Health"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.library.tmc.edu/utgsbs_dissertations/1321","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Xuelin Huang","Ziyi Li","Bing Z Carter"]},{"key":"dc:creator","label":"Author","values":["Wang, Yizhuo","<p>0000-0002-1870-0019</p>"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2024-12-12T08:00:00Z"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation (PhD)"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Biostatistics","machine learning","clinical trial design","adaptive randomization","survival analysis","missing data","bioinformatics","high-dimensional data","scRNA data","Clinical Trials","Data Science","Public Health"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.library.tmc.edu/utgsbs_dissertations/1321"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>In the contemporary healthcare field, professionals are confronted with an ever-growing volume of clinical data stored in electronic health records, alongside the genomic data stemming from laboratory experiments. As a response to this deluge of data, the application of machine learning (ML) techniques is gaining popularity since ML techniques have demonstrated an exceptional proficiency in processing big data and deciphering complex nonlinear patterns that are intrinsic to biomedical research.</p> <p>My research leverages ML's capabilities to address the computational challenges spanning diverse areas, including adaptive clinical trial designs, survival analysis, and high-dimensional genetic data analysis. Specifically, Chapter 2 focused on the application of ML in response-adaptive randomization designs. Compared to a traditional equal-randomization design, adaptive randomized trials allocate more patients to the superior treatment arm and increase the overall response rate. In Chapter 3, we proposed a statistical model to impute survival times for censored observations, allowing for a direct application of any ML method in the downstream analysis. We further improve the accuracy of our method using ML regression built on prognostic covariates. In Chapter 4, we developed an artificial neural network-based framework for analyzing longitudinal single-cell RNA sequencing data. Our pipeline achieves: (1) cross-time points cell annotation, (2) detection of novel cell type emerged over time, (3) visualization of cell population evolution, (4) identification of temporal differentially expressed genes. In each chapter, we provide both simulation studies and real-world application results for our methods. In the long run, we aim to lessen the gap between ML and biomedical research, and facilitate healthcare studies.</p>"]},{"key":"dc:title","label":"Title","values":["Addressing the Analytical and Computational Challenges Using Machine Learning in Biomedical Research"]}]}],"canonical_facts":{"dc:contributor":["Xuelin Huang","Ziyi Li","Bing Z Carter"],"dc:creator":["Wang, Yizhuo","<p>0000-0002-1870-0019</p>"],"dc:date.available":["2024-12-12T08:00:00Z"],"dc:description.abstract":["<p>In the contemporary healthcare field, professionals are confronted with an ever-growing volume of clinical data stored in electronic health records, alongside the genomic data stemming from laboratory experiments. As a response to this deluge of data, the application of machine learning (ML) techniques is gaining popularity since ML techniques have demonstrated an exceptional proficiency in processing big data and deciphering complex nonlinear patterns that are intrinsic to biomedical research.</p> <p>My research leverages ML's capabilities to address the computational challenges spanning diverse areas, including adaptive clinical trial designs, survival analysis, and high-dimensional genetic data analysis. Specifically, Chapter 2 focused on the application of ML in response-adaptive randomization designs. Compared to a traditional equal-randomization design, adaptive randomized trials allocate more patients to the superior treatment arm and increase the overall response rate. In Chapter 3, we proposed a statistical model to impute survival times for censored observations, allowing for a direct application of any ML method in the downstream analysis. We further improve the accuracy of our method using ML regression built on prognostic covariates. In Chapter 4, we developed an artificial neural network-based framework for analyzing longitudinal single-cell RNA sequencing data. Our pipeline achieves: (1) cross-time points cell annotation, (2) detection of novel cell type emerged over time, (3) visualization of cell population evolution, (4) identification of temporal differentially expressed genes. In each chapter, we provide both simulation studies and real-world application results for our methods. In the long run, we aim to lessen the gap between ML and biomedical research, and facilitate healthcare studies.</p>"],"dc:identifier":["https://digitalcommons.library.tmc.edu/utgsbs_dissertations/1321"],"dc:subject":["Biostatistics","machine learning","clinical trial design","adaptive randomization","survival analysis","missing data","bioinformatics","high-dimensional data","scRNA data","Clinical Trials","Data Science","Public Health"],"dc:title":["Addressing the Analytical and Computational Challenges Using Machine Learning in Biomedical Research"],"thesis:degree_level":["Dissertation (PhD)"],"thesis:degree_name":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-24T05:48:59Z"}