Back to results

ResearchSpace@Auckland

Assembly Video Understanding for Human-Robot Collaboration: Methods and Applications

Abstract

dc:description.abstract

The human-robot relationship is evolving from coexistence and cooperation to collaboration, compassion, and coevolution. In this process, robots are transitioning from machines focused solely on production tasks to reliable partners that prioritise human needs and well-being. This shift requires robots to learn from human expertise, understand human needs and intentions, and collaborate seamlessly with humans on complex tasks. This thesis addresses this need within the context of product assembly. To enable robots to effectively learn from, understand, and collaborate with humans in a shared assembly task, it is necessary to represent the assembly process in a way that both humans and robots can understand. Therefore, a dual-hand human-robot shared assembly process representation method was designed. It decomposes an assembly process into a series of primitive tasks and atomic actions. By rigorously and objectively defining these tasks and actions, the assembly process is uniformly and unambiguously represented, meeting both human intuitive understanding and robotic programming requirements. This work focuses on vision-based methods for understanding assembly processes, as they provide a natural way for robots to interact with humans. However, the lack of datasets capturing complex human assembly actions in industrial scenarios has limited the progress in this field. To address this gap, a human assembly video dataset, HA-ViD, was created. HA-ViD includes common assembly actions, parts, and tools used in industry and comprehensive dual-hand action annotations. Accurately segmenting the sequence of assembly actions is critical for understanding the assembly process, assessing progress, and inferring human intention; therefore, action segmentation is a pivotal technique in this domain. However, existing action segmentation methods are mostly tailored to sports or daily activities and lack the ability to handle the complexity and subtleties of assembly actions. This thesis makes several technical contributions by introducing multiple enhancements to action segmentation algorithms and applying them to various human-robot collaboration (HRC) applications. First, a video-based human fatigue estimation method was developed in the context of human-centric manufacturing. This method introduces a boundary-aware dual-stream action segmentation approach to detect operation types and repetitions, which are then used to estimate human fatigue levels. The estimated fatigue is subsequently applied to optimise human-robot task allocation. Second, a dual-hand action segmentation algorithm, DuHa, was developed to simultaneously segment actions of both hands. To enable real-time performance, DuHa-v2 was proposed, in which the unified action features and implicit object interaction features were designed to reduce computational costs. DuHa-v2 was embedded into a human-robot collaborative assembly framework, which also supports in-process quality checking. The superiority of DuHa and DuHa-v2 over existing action segmentation methods was validated on HA-ViD, and the effectiveness of the proposed framework was confirmed through a real-world case study. Finally, traditional action segmentation methods lack scene adaptability, partly because they conceptualise actions as unified verb-object entities with complete semantics. To overcome this, the dual-hand compositional action segmentation method, DuCAS, was proposed. Instead of segmenting the semantic-complete actions, DuCAS segments action elements—action verb, manipulated object, target object, and tool—and then combines them to form the semantic-complete actions. Furthermore, DuCAS was applied to a framework of multi-modal assembly instruction generation from demonstration videos. In conclusion, this thesis makes four principal contributions advancing assembly video understanding for HRC. First, DH-APR provides the first dual-hand assembly process representation enabling consistent human-robot communication. Second, HA-ViD represents the first industrial assembly dataset with comprehensive dual-hand action annotations. Third, the proposed action segmentation algorithms address existing methods' limitations in hand-object interaction modelling, dual-hand action understanding, real-time capability, and scene adaptability. Fourth, integrated HRC applications demonstrate action understanding's practical effectiveness in fatigue-aware human-robot task allocation, collaborative assembly, and multi-modal instruction generation. Together, these contributions advance the research of both video understanding and HRC.

Degree

thesis:*
Name thesis:degree_name
PhD
Level thesis:degree_level
Doctoral
Discipline thesis:degree_discipline
Mechanical Engineering
Grantor dc:publisher
ResearchSpace@Auckland
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zheng, Hao
Advisors dc:contributor.advisor
  • Xu, Xun
  • Polzer, Jan

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated.

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/2292/74291
OAI identifier oai:identifier
oai:researchspace.auckland.ac.nz:2292/74291

Chain of custody

source
Harvested from
University of Auckland
Base URL
researchspace.auckland.ac.nz/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Zheng, Hao. Assembly Video Understanding for Human-Robot Collaboration: Methods and Applications. Doctoral thesis, ResearchSpace@Auckland, 2024. https://hdl.handle.net/2292/74291