Back to results

University of Illinois Urbana-Champaign

Supervisory control with online learning for stabilization and near-optimal performance of time-varying linear systems

Abstract

dc:description

Model-based control methods are widely used in robotics because they use system equations to compute efficient control actions. However, these methods often struggle in real-world situations where the system model is not perfect or where there are unexpected disturbances. In addition, solving nonlinear optimization problems in real time can be too slow or too demanding for systems with limited onboard computing power. To address these challenges, this study proposes a hybrid control approach that combines classical control, optimal planning and online learning. The system we focus on is a 2D quadrotor, modeled as a six-dimensional system controlled using force and torque inputs. At the lower level, we use three different types of controllers: a basic Proportional-Derivative (PD) controller, a trajectory planner using nonlinear programming (NLP), and a control law based on Pontryagin’s Maximum Principle (PMP), which we implement using PyTorch. At the higher level, we add a Multi-Armed Bandit (MAB) layer using the EXP3 algorithm. This layer learns over time which controller performs best based on feedback like tracking error and energy usage. It allows the system to switch between controllers depending on how well they are working at each moment. Our results show that this combination of planning and learning can make the system more reliable and adaptive, even in uncertain environments. While we apply this to a quadrotor, the same idea can be used for many other types of robotic systems.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Industrial Engineering
Grantor
University of Illinois Urbana-Champaign
Year dc:date
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Roy, Dhritiman
Contributors dc:contributor
  • Li, Yingying

Subjects

dc:subject × 4

Rights

dc:rights
Statement dc:rights
  • Copyright 2025 Dhritiman Roy
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/129620

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Roy, Dhritiman. Supervisory control with online learning for stabilization and near-optimal performance of time-varying linear systems. Thesis thesis, University of Illinois Urbana-Champaign, 2025. https://hdl.handle.net/2142/129620