Back to results

University of Cambridge

Toward effective and generalisable machine learning for biosignal time series

Abstract

dc:description.abstract

Biosignal time series collected from wearable devices, such as electrocardiograms (ECG), electroencephalograms (EEG), and Inertial Measurement Units (IMUs), enable continuous monitoring of human physiology and behaviour. Using these signals provides unique opportunities to advance personalised health monitoring, improve early disease detection, and support the development of scaleable and cost-effective healthcare solutions. However, analysing such data presents significant challenges due to its complex nature, as signals are often sparse, irregular, contain missing values, and lack sufficient labels. In contrast, existing state-of-the-art methods are typically developed and validated on clean, regularly sampled biosignals with clinically verified ground-truth annotations, which limits their effectiveness and leads to performance degradation when applied to real-world data. Moreover, most existing models for biosignal time series are developed and evaluated in-domain, where both training and testing are conducted on the same dataset or under the same data collection protocol. Therefore, such models often fail to generalise to new scenarios in real-world healthcare deployments due to distributional shifts across tasks or deployment settings. This thesis addresses these challenges by developing machine learning methods tailored for biosignal time series that focus on self-supervised representation learning, domain adaptation, and foundation modelling for generalisation across tasks and datasets. We demonstrate the efficacy of these methods on a wide range of real-world healthcare tasks, from daily cardio-fitness monitoring to clinical applications. First, to address the scarcity of labelled data and the inherent complexity of biosignal time series, we propose a contrastive self-supervised learning framework specifically designed for biosignals\correction{,} collected across diverse sources, ranging from wearable devices to clinical monitoring systems. Inspired by the success of contrastive learning in computer vision, our method captures the temporal and structural characteristics of biosignals without relying on labels. Then, our new training pipeline learns more effective representations from unlabelled data and improves the downstream task performance. This performance remains comparable to state-of-the-art methods, even when only a small fraction of labelled data is available, thus enhancing label efficiency. Second, to improve model robustness under distribution shifts, we develop a domain adaptation framework with multiple discriminators during fine-tuning. This method learns domain-invariant representations that generalise across source and target domains with differing label distributions, which is a common issue in healthcare due to the high cost and scarcity of gold-standard annotations. We validate this approach on a VO2max prediction task and demonstrate improved adaptation performance under real-world domain shifts. Finally, motivated by the growing success of foundation models in language and vision, we introduce a general-purpose foundation model tailored for biosignal time series to alleviate data complexity and improve model generalisability. Our approach pretrains on heterogeneous datasets that vary in modality, sampling rates, and missingness patterns, and is capable of generalising to a wide range of unseen downstream tasks, from classification to regression, across different biosignal sources. This foundation model provides a scaleable and flexible framework for developing future biosignal-based healthcare applications. Together, this dissertation tackles three central challenges in biosignal time series modelling: (i) data and label scarcity, addressed through contrastive self-supervised learning for label-efficient representation learning; (ii) domain shifts, mitigated by a multi-discriminator domain adaptation framework; and (iii) data and domain heterogeneity, alleviated by a general-purpose biosignal foundation model. These contributions were validated on real-world physiological datasets as well as commonly used machine learning benchmarks, with extensive experiments demonstrating their effectiveness in handling missingness, irregularity, and limited labels while enabling generalisable model deployment across diverse healthcare scenarios. Collectively, this progression from data-efficient modelling to robust domain adaptation and ultimately to scaleable generalisation underscores the potential of deep learning to advance the practical adoption of biosignal-based healthcare solutions.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Wu, Yu
Advisor dc:contributor.advisor
  • Mascolo, Cecilia

Subjects

dc:subject × 4

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
Author Identifier
0009-0004-3709-7472
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/394061

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Wu, Yu. Toward effective and generalisable machine learning for biosignal time series. Doctoral thesis, University of Cambridge, 2025. https://doi.org/10.17863/CAM.124154