University of Toronto
Modeling sources of variation to interpret single cell transcriptomic maps
Abstract
dc:description.abstractUnderstanding how cells are organized and function within an organism is essential for studying biological systems. Single-cell RNA sequencing profiles thousands of individual cells and captures coexisting gene expression programs within a cell. Analyzing data from multiple samples and conditions introduces additional biological variability, including genetic background, age, and disease state, along with technical variability from sample preparation and processing. If not properly addressed, these technical variations can obscure true biological signals. Disentangling these sources of variation and their interactions is crucial for gaining a holistic view of the biological system.The focus of my thesis was threefold: first, to develop methods that disentangle the biological and technical factors, revealing the underlying structure of the data and identifying key contributors to its variability; second, to apply this exploratory approach to healthy rat liver samples, resulting in the creation of the first multi-strain map of the healthy rat liver; and third, to develop a fast and scalable differential expression method for exploring hypothesis-driven queries while accounting for the statistical properties of multi-sample scRNA-seq datasets. To facilitate the exploratory phase of single cell transcriptomics data analysis, I developed sciRED, a novel tool for improving the interpretability of scRNA-seq factor decomposition. sciRED enables the identification and interpretation of critical factors in scRNA-seq data by combining machine learning and statistical approaches. It automatically extracts factors from gene expression data, identifying those explained by known covariates while uncovering novel factors that may represent previously unrecognized biological insights. Using sciRED, we generated the first multi-strain map of the healthy rat liver. This factor decomposition approach reduced ambient RNA contamination, facilitated the annotation of cell populations, and uncovered strain-specific variations in myeloid population of the healthy rat liver. The final component of my thesis focuses on developing FLASH-MM, a fast and scalable differential expression method for scRNA-seq data. FLASH-MM employs an efficient mixed-model framework to address complex hierarchical structures and correlations in multi-sample studies, making it particularly suitable for large datasets. Together, these methods provide an integrative framework for scRNA-seq analysis. Exploratory methods, such as sciRED, allow researchers to identify major sources of variation and explore the data structure without predefined assumptions. These insights can then inform the design of queries using FLASH-MM, guiding the selection of covariates and facilitating hypothesis-driven analyses.
Degree
thesis:*- Department dc:contributor.department
- Molecular Genetics
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Pouyabahar, Delaram
- Advisor dc:contributor.advisor
-
- Bader, Gary D.
Subjects
dc:subject × 6Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/1807/150293
- OAI identifier oai:identifier
- oai:utoronto.scholaris.ca:1807/150293