Queens University
Data-Driven Pipeline for Learning Discrete behavioural Models of Cyber-Physical Systems
Abstract
dc:description.abstractCyber-Physical Systems (CPS), such as cruise control systems, smart thermostats, and manufacturing controllers, generate continuous streams of sensor data and system parameters that show behavioural patterns. Extracting discrete abstractions from these logs is useful for fault detection, system monitoring, or even reverse engineering when no model of the system is available. The core contribution of this research is the development of an automated pipeline for transforming raw, continuous-valued CPS logs into structured discrete-event behavioural models called automata, while minimizing the need for ground truth about the data or the system that generated the logs. The pipeline consists of two recurring steps. In the discretization stage, high-dimensional time series data are converted into event traces without requiring labelled ground truth. A non-parametric change point detection algorithm identifies statistical changes in sensor readings or system parameters, which are treated as events. These events are then grouped into event types across system variables in an unsupervised manner, requiring no prior knowledge of event categories. The resulting event traces are subsequently passed to the automata learning stage, where a probabilistic timed automata learning algorithm is applied to derive a behavioural model of the system. Both the discretization step and the model learning step are algorithms with hyperparameters. In the absence of prior knowledge about the data, a poor manual choice of hyperparameters can produce uninformative models. To address this and to keep the need for prior knowledge to a minimum, we design the pipeline as a nested particle swarm optimization framework guided by a quality measure from the automata learning algorithm of choice. This enables automatic tuning of discretization and learning hyperparameters through non-convex optimization, eliminating the need for manual intervention. The method is evaluated on two simulated datasets and a real-world one. The strengths and limitations of the pipeline are discussed alongside its performance on the evaluation datasets. The pipeline is able to capture meaningful behaviour on simulated datasets, and the quality metrics obtained on all datasets are encouraging.
Degree
thesis:*- Department dc:contributor.department
- Electrical and Computer Engineering
- Year dc:date.issued
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Kianersi, Nastaran
- Advisor dc:contributor.supervisor
-
- Kauffman, Sean
Subjects
dc:subject × 7Rights
- Language dc:language.iso
- eng
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/1974/36534
- OAI identifier oai:identifier
- oai:queensu.scholaris.ca:1974/36534