The University of Texas at Austin
Generative modeling of the tumor microenvironment : deconvolution, completion, and integration
Abstract
dc:description.abstractModern biological measurement technologies provide increasingly diverse views of tissue state, including spatial transcriptomics, spatial proteomics, and multi-sequence medical imaging. However, the data collected by these platforms is often incomplete, aggregated, or distributed across unpaired modalities. These acquisition constraints create three recurring computational problems: deconvolution, where each observation mixes signals from multiple cells; completion, where important channels or sequences are missing; and integration, where complementary modalities are measured on adjacent sections without direct correspondence. Existing methods typically address these settings separately and often rely on strong assumptions such as paired observations, fixed missingness patterns, or predefined feature correspondences. This dissertation develops generative modeling methods for each of these three problems. We introduce BayesTME, a Bayesian framework for deconvolving aggregated spatial transcriptomics measurements without requiring paired single-cell references. BayesTME models spot-level observations using spatially structured priors over latent cell composition and within-cell-type transcriptional variation, enabling joint inference of cell-type proportions, tissue organization, and spatial transcriptional programs. To address incompleteness in biological data, we develop Multi-Channel Diffusion (MCD) and its extension DiffuseMRI, which completes data by learning conditional diffusion models that recover missing channels or imaging sequences from arbitrary observed subsets. By training under randomized masking, these models amortize over the conditioning space, allowing a single model to generalize across many missingness configurations while preserving spatial structure. Finally, we propose MultiTME, a framework for integrating unpaired spatial modalities by learning shared latent representations under cycle-consistency, spatial, and cell-type constraints. These methods demonstrate that structured generative modeling can infer biologically meaningful signals under acquisition constraints, without requiring paired supervision or complete observations. We validate each method on real tissue datasets spanning spatial transcriptomics, spatial proteomics, and MRI, demonstrating improvements in deconvolution accuracy, missing-signal recovery, and cross-modal alignment compared with existing approaches.
Degree
thesis:*- Name thesis:degree_name
- Doctor of Philosophy
- Level thesis:degree_level
- DOCTORAL
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- The University of Texas at Austin
- Year dc:date.issued
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Zhang, Haoran, Ph. D.
- Advisors dc:contributor.advisor
-
- Zhou, Mingyuan (Assistant professor)
- Tansey, Wesley Scott
- Committee members dc:contributor.committeemember
-
- Liu, Qiang
- Huang, Qixing
- Scott, James G
Subjects
dc:subject × 4Rights
- Language dc:language.iso
- English
Identifiers
dc:identifier.*- Identifier URI
- https://doi.org/10.26153/tsw/64258
- OAI identifier oai:identifier
- oai:repositories.lib.utexas.edu:2152/136956