University of Missouri--Columbia
Applicability and limitations of transcriptomic pipelines in livestock species
Abstract
dc:description.abstractThe genome-to-phenome axis represents a complex relationship influenced by all levels of biology, from specific molecules in cells to broad global diversity. Addressing the fundamental question of genetics- how genetic loci contribute to phenotype- can be challenging. Although there are clear deviations, the central dogma offers an initial step toward that answer, as the primary and secondary products generated from the encoded information serve a wide array of functions and locations. They interact within networks and cascades, facilitating transport, movement, messaging, regulation, and more. The regulation of what, when, and how much of any product influences potential downstream functions. The original question can then be expanded to: In what biological context do variations in genetic loci play a role in regulating products and the terminal observable or measurable phenotype? Answers to this question can provide important insights applicable to food security, conservation efforts, and understanding organisms' responses to disease. Transcriptomic datasets offer a contextualized snapshot of the steady-state levels of RNA from the biological sources from which they are derived, allowing for the identification of variants affecting gene regulation. This information can be combined with phenotypes to explore the loci influencing that trait and how this effect is mediated. To narrow the information available to what is relevant for a selected trait requires high-throughput bioinformatics pipelines to effectively filter and prioritize pertinent genes and transcripts. Livestock species represent a vital branch of the Tree of Life, serving as food sources, model organisms, and essential species within ecological niches. Extending bioinformatics tools from humans to livestock species necessitates evaluating the underlying models, workflows, and assumptions. To establish a means of exploring this dynamic across species and biological contexts, we start with the economically relevant trait of cattle feed efficiency. By investigating a complex trait that exhibits similarities and variations in its regulation across biological contexts, we can elucidate the characteristics of relevant loci. This will provide specific information on metabolic efficiency and broader insights into how loci contribute to complex traits. We present Chapter 1: Integration of Feed Efficiency Associated Variants with Molecular Phenotype QTL as a description of pipeline development and its ability to contribute to finding variants affecting an economically important trait. We first present the development and evaluation of a bulk RNA-sequencing molecular QTL pipeline (xQTL) in cattle, illustrating its applicability in identifying variants that impact feed efficiency through molecular mechanisms. In exploring the question of genetic loci contributing to gene regulation and complex traits, this first project accomplished two main goals. The first was identifying and characterizing variants believed to impact cattle feed efficiency. By building on prior results, we were able to replicate and identify variants and provide discussion into what and how these effects are mediated. We characterize variants that potentially affect 5 feed efficiency traits with mechanisms in 3 distinct tissues and 4 molecular phenotypes. The second was building a bioinformatic pipeline to explore this question within livestock species. The pipeline was designed to be generalized for other species and biological contexts; however, its general applicability was unknown without testing in another species. To address this, I had the opportunity to participate in a study focused on honeybee resistance to a parasitic mite. This experience provided a non-mammalian species with distinct genome characteristics and population structures compared to cattle and other mammals. This enabled an evaluation of the workflow's limitations when applied beyond standard parameters. The trait of interest, Varroa Sensitive Hygiene, further expands this evaluation as it is presented here as a binary behavioral trait rather than a continuous trait like metabolic efficiency. In Chapter 2: Identification of Genomic Regions Affecting Varroa Sensitive Hygiene, I assess the adaptability of the xQTL pipeline to identify variants impacting a behavioral trait of great concern to the commercial bee industry. By analyzing the effects of recombination rate and genome size on xQTL analyses, variants influencing gene regulation were identified in honeybees. This underscores the species-agnostic flexibility of the workflow. However, we could not identify variants with statistical significance for the target trait of Varroa Sensitive Hygiene, possibly due to the highly polygenic nature of behavioral traits and the sample size of 295 individuals. Nevertheless, we identified regions of interest where allele frequencies differ between performing and nonperforming individuals. Based on the evaluation here, the xQTL pipeline, along with the integration workflow described in Chapter 1, will facilitate the identification of these variants at tissue resolution with appropriate sample sizes. The first two chapters highlight that variants affecting complex traits and gene regulation can be identified, but only in tissues, which are heterogeneous aggregates of cell types. The cell, the fundamental unit of life, provides the highest resolution for understanding traits. The emergence of single-cell and nuclei RNA-sequencing (SCSN RNA-seq) has addressed this challenge, with numerous studies demonstrating its potential for discovery. When initially exploring datasets, we observed that, due to data generation and manual exploration, scaling to many samples would be bottlenecked by the time spent on manual quality control, curation, and cell type annotation. We therefore sought to build a pipeline that allows for efficient and reproducible data exploration, facilitating an extension into this biological context. In Chapter 3: Increasing Transcriptomics to Cell Type Resolution Considerations and Applications, we describe the generated pipeline and its utility in creating cell-type transcriptional profiles for identifying the cellular context in which target genes affecting traits are active. We emphasize the need for effective permutations of filtering and clustering thresholds to ensure confident cell-type expression profiles. By applying the pipeline to multiple species, we illustrate that limitations still exist in annotating cell types in a high-throughput manner. However, by prioritizing individual nuclei believed to be of high quality, the cell type profiles of target genes can be further investigated. The evaluation framework presented here provides a means to effectively develop and apply high-throughput bioinformatic pipelines to new species and biological contexts, illuminating the molecular underpinnings of the genome-to-phenome axis of complex traits. Together, the two presented pipelines enable the identification of variants affecting complex traits, offering both a cell type location and a potential mechanism of action for the identified variants. By evaluating the species and biological context, the applicability of these pipelines can be assessed as they are applied to novel traits and challenges.
Degree
thesis:*- Name thesis:degree_name
- Ph. D.
- Level thesis:degree_level
- Doctoral
- Discipline thesis:degree_discipline
- Genetics
- Grantor dc:publisher
- University of Missouri--Columbia
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Stull, Caleb Michael
- Advisor dc:contributor.advisor
-
- Schnabel, Robert D.
Rights
- Language dc:language.iso
- eng, English
Identifiers
dc:identifier.*- OAI identifier oai:identifier
- oai:mospace.umsystem.edu:10355/109526