{"id":{"repo_id":"auckland-ms","oai_identifier":"oai:researchspace.auckland.ac.nz:2292/74291"},"canonical_url":"https://search.dev.ndltd.org/etd/auckland-ms/oai:researchspace.auckland.ac.nz:2292/74291","repository":{"repo_id":"auckland-ms","name":"University of Auckland","base_url":"https://researchspace.auckland.ac.nz/server/oai/request"},"display":{"title":"Assembly Video Understanding for Human-Robot Collaboration: Methods and Applications","abstract":"The human-robot relationship is evolving from coexistence and cooperation to collaboration, compassion, and coevolution. In this process, robots are transitioning from machines focused solely on production tasks to reliable partners that prioritise human needs and well-being. This shift requires robots to learn from human expertise, understand human needs and intentions, and collaborate seamlessly with humans on complex tasks. This thesis addresses this need within the context of product assembly. To enable robots to effectively learn from, understand, and collaborate with humans in a shared assembly task, it is necessary to represent the assembly process in a way that both humans and robots can understand. Therefore, a dual-hand human-robot shared assembly process representation method was designed. It decomposes an assembly process into a series of primitive tasks and atomic actions. By rigorously and objectively defining these tasks and actions, the assembly process is uniformly and unambiguously represented, meeting both human intuitive understanding and robotic programming requirements. This work focuses on vision-based methods for understanding assembly processes, as they provide a natural way for robots to interact with humans. However, the lack of datasets capturing complex human assembly actions in industrial scenarios has limited the progress in this field. To address this gap, a human assembly video dataset, HA-ViD, was created. HA-ViD includes common assembly actions, parts, and tools used in industry and comprehensive dual-hand action annotations. Accurately segmenting the sequence of assembly actions is critical for understanding the assembly process, assessing progress, and inferring human intention; therefore, action segmentation is a pivotal technique in this domain. However, existing action segmentation methods are mostly tailored to sports or daily activities and lack the ability to handle the complexity and subtleties of assembly actions. This thesis makes several technical contributions by introducing multiple enhancements to action segmentation algorithms and applying them to various human-robot collaboration (HRC) applications. First, a video-based human fatigue estimation method was developed in the context of human-centric manufacturing. This method introduces a boundary-aware dual-stream action segmentation approach to detect operation types and repetitions, which are then used to estimate human fatigue levels. The estimated fatigue is subsequently applied to optimise human-robot task allocation. Second, a dual-hand action segmentation algorithm, DuHa, was developed to simultaneously segment actions of both hands. To enable real-time performance, DuHa-v2 was proposed, in which the unified action features and implicit object interaction features were designed to reduce computational costs. DuHa-v2 was embedded into a human-robot collaborative assembly framework, which also supports in-process quality checking. The superiority of DuHa and DuHa-v2 over existing action segmentation methods was validated on HA-ViD, and the effectiveness of the proposed framework was confirmed through a real-world case study. Finally, traditional action segmentation methods lack scene adaptability, partly because they conceptualise actions as unified verb-object entities with complete semantics. To overcome this, the dual-hand compositional action segmentation method, DuCAS, was proposed. Instead of segmenting the semantic-complete actions, DuCAS segments action elements—action verb, manipulated object, target object, and tool—and then combines them to form the semantic-complete actions. Furthermore, DuCAS was applied to a framework of multi-modal assembly instruction generation from demonstration videos. In conclusion, this thesis makes four principal contributions advancing assembly video understanding for HRC. First, DH-APR provides the first dual-hand assembly process representation enabling consistent human-robot communication. Second, HA-ViD represents the first industrial assembly dataset with comprehensive dual-hand action annotations. Third, the proposed action segmentation algorithms address existing methods' limitations in hand-object interaction modelling, dual-hand action understanding, real-time capability, and scene adaptability. Fourth, integrated HRC applications demonstrate action understanding's practical effectiveness in fatigue-aware human-robot task allocation, collaborative assembly, and multi-modal instruction generation. Together, these contributions advance the research of both video understanding and HRC.","abstract_html":"The human-robot relationship is evolving from coexistence and cooperation to collaboration, compassion, and coevolution. In this process, robots are transitioning from machines focused solely on production tasks to reliable partners that prioritise human needs and well-being. This shift requires robots to learn from human expertise, understand human needs and intentions, and collaborate seamlessly with humans on complex tasks. This thesis addresses this need within the context of product assembly. To enable robots to effectively learn from, understand, and collaborate with humans in a shared assembly task, it is necessary to represent the assembly process in a way that both humans and robots can understand. Therefore, a dual-hand human-robot shared assembly process representation method was designed. It decomposes an assembly process into a series of primitive tasks and atomic actions. By rigorously and objectively defining these tasks and actions, the assembly process is uniformly and unambiguously represented, meeting both human intuitive understanding and robotic programming requirements. This work focuses on vision-based methods for understanding assembly processes, as they provide a natural way for robots to interact with humans. However, the lack of datasets capturing complex human assembly actions in industrial scenarios has limited the progress in this field. To address this gap, a human assembly video dataset, HA-ViD, was created. HA-ViD includes common assembly actions, parts, and tools used in industry and comprehensive dual-hand action annotations. Accurately segmenting the sequence of assembly actions is critical for understanding the assembly process, assessing progress, and inferring human intention; therefore, action segmentation is a pivotal technique in this domain. However, existing action segmentation methods are mostly tailored to sports or daily activities and lack the ability to handle the complexity and subtleties of assembly actions. This thesis makes several technical contributions by introducing multiple enhancements to action segmentation algorithms and applying them to various human-robot collaboration (HRC) applications. First, a video-based human fatigue estimation method was developed in the context of human-centric manufacturing. This method introduces a boundary-aware dual-stream action segmentation approach to detect operation types and repetitions, which are then used to estimate human fatigue levels. The estimated fatigue is subsequently applied to optimise human-robot task allocation. Second, a dual-hand action segmentation algorithm, DuHa, was developed to simultaneously segment actions of both hands. To enable real-time performance, DuHa-v2 was proposed, in which the unified action features and implicit object interaction features were designed to reduce computational costs. DuHa-v2 was embedded into a human-robot collaborative assembly framework, which also supports in-process quality checking. The superiority of DuHa and DuHa-v2 over existing action segmentation methods was validated on HA-ViD, and the effectiveness of the proposed framework was confirmed through a real-world case study. Finally, traditional action segmentation methods lack scene adaptability, partly because they conceptualise actions as unified verb-object entities with complete semantics. To overcome this, the dual-hand compositional action segmentation method, DuCAS, was proposed. Instead of segmenting the semantic-complete actions, DuCAS segments action elements—action verb, manipulated object, target object, and tool—and then combines them to form the semantic-complete actions. Furthermore, DuCAS was applied to a framework of multi-modal assembly instruction generation from demonstration videos. In conclusion, this thesis makes four principal contributions advancing assembly video understanding for HRC. First, DH-APR provides the first dual-hand assembly process representation enabling consistent human-robot communication. Second, HA-ViD represents the first industrial assembly dataset with comprehensive dual-hand action annotations. Third, the proposed action segmentation algorithms address existing methods&#x27; limitations in hand-object interaction modelling, dual-hand action understanding, real-time capability, and scene adaptability. Fourth, integrated HRC applications demonstrate action understanding&#x27;s practical effectiveness in fatigue-aware human-robot task allocation, collaborative assembly, and multi-modal instruction generation. Together, these contributions advance the research of both video understanding and HRC.","abstract_has_math":false,"creators":["Zheng, Hao"],"institution":"ResearchSpace@Auckland","degree_name":"PhD","degree_level":"Doctoral","degree_discipline":"Mechanical Engineering","degree_department":null,"school":null,"contributors":[],"advisors":["Xu, Xun","Polzer, Jan"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024","date_published":"2024","updated_at":"2026-07-24T01:04:52Z","subjects":["Smart Manufacturing","Human-Robot Collaboration","Video Understanding","Intelligent Assembly","Industry 5.0"],"languages":[],"rights":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."],"rights_urls":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2292/74291","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Xu, Xun","Polzer, Jan"]},{"key":"dc:creator","label":"Author","values":["Zheng, Hao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-12-14T18:54:56Z"]},{"key":"dc:date.issued","label":"Date","values":["2024"]},{"key":"dc:publisher","label":"Institution","values":["ResearchSpace@Auckland"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Mechanical Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["PhD"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["The University of Auckland"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Smart Manufacturing","Human-Robot Collaboration","Video Understanding","Intelligent Assembly","Industry 5.0"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/2292/74291"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The human-robot relationship is evolving from coexistence and cooperation to collaboration, compassion, and coevolution. In this process, robots are transitioning from machines focused solely on production tasks to reliable partners that prioritise human needs and well-being. This shift requires robots to learn from human expertise, understand human needs and intentions, and collaborate seamlessly with humans on complex tasks. This thesis addresses this need within the context of product assembly. To enable robots to effectively learn from, understand, and collaborate with humans in a shared assembly task, it is necessary to represent the assembly process in a way that both humans and robots can understand. Therefore, a dual-hand human-robot shared assembly process representation method was designed. It decomposes an assembly process into a series of primitive tasks and atomic actions. By rigorously and objectively defining these tasks and actions, the assembly process is uniformly and unambiguously represented, meeting both human intuitive understanding and robotic programming requirements. This work focuses on vision-based methods for understanding assembly processes, as they provide a natural way for robots to interact with humans. However, the lack of datasets capturing complex human assembly actions in industrial scenarios has limited the progress in this field. To address this gap, a human assembly video dataset, HA-ViD, was created. HA-ViD includes common assembly actions, parts, and tools used in industry and comprehensive dual-hand action annotations. Accurately segmenting the sequence of assembly actions is critical for understanding the assembly process, assessing progress, and inferring human intention; therefore, action segmentation is a pivotal technique in this domain. However, existing action segmentation methods are mostly tailored to sports or daily activities and lack the ability to handle the complexity and subtleties of assembly actions. This thesis makes several technical contributions by introducing multiple enhancements to action segmentation algorithms and applying them to various human-robot collaboration (HRC) applications. First, a video-based human fatigue estimation method was developed in the context of human-centric manufacturing. This method introduces a boundary-aware dual-stream action segmentation approach to detect operation types and repetitions, which are then used to estimate human fatigue levels. The estimated fatigue is subsequently applied to optimise human-robot task allocation. Second, a dual-hand action segmentation algorithm, DuHa, was developed to simultaneously segment actions of both hands. To enable real-time performance, DuHa-v2 was proposed, in which the unified action features and implicit object interaction features were designed to reduce computational costs. DuHa-v2 was embedded into a human-robot collaborative assembly framework, which also supports in-process quality checking. The superiority of DuHa and DuHa-v2 over existing action segmentation methods was validated on HA-ViD, and the effectiveness of the proposed framework was confirmed through a real-world case study. Finally, traditional action segmentation methods lack scene adaptability, partly because they conceptualise actions as unified verb-object entities with complete semantics. To overcome this, the dual-hand compositional action segmentation method, DuCAS, was proposed. Instead of segmenting the semantic-complete actions, DuCAS segments action elements—action verb, manipulated object, target object, and tool—and then combines them to form the semantic-complete actions. Furthermore, DuCAS was applied to a framework of multi-modal assembly instruction generation from demonstration videos. In conclusion, this thesis makes four principal contributions advancing assembly video understanding for HRC. First, DH-APR provides the first dual-hand assembly process representation enabling consistent human-robot communication. Second, HA-ViD represents the first industrial assembly dataset with comprehensive dual-hand action annotations. Third, the proposed action segmentation algorithms address existing methods' limitations in hand-object interaction modelling, dual-hand action understanding, real-time capability, and scene adaptability. Fourth, integrated HRC applications demonstrate action understanding's practical effectiveness in fatigue-aware human-robot task allocation, collaborative assembly, and multi-modal instruction generation. Together, these contributions advance the research of both video understanding and HRC."]},{"key":"dc:title","label":"Title","values":["Assembly Video Understanding for Human-Robot Collaboration: Methods and Applications"]}]}],"canonical_facts":{"dc:contributor.advisor":["Xu, Xun","Polzer, Jan"],"dc:creator":["Zheng, Hao"],"dc:date.accessioned":["2025-12-14T18:54:56Z"],"dc:date.issued":["2024"],"dc:description.abstract":["The human-robot relationship is evolving from coexistence and cooperation to collaboration, compassion, and coevolution. In this process, robots are transitioning from machines focused solely on production tasks to reliable partners that prioritise human needs and well-being. This shift requires robots to learn from human expertise, understand human needs and intentions, and collaborate seamlessly with humans on complex tasks. This thesis addresses this need within the context of product assembly. To enable robots to effectively learn from, understand, and collaborate with humans in a shared assembly task, it is necessary to represent the assembly process in a way that both humans and robots can understand. Therefore, a dual-hand human-robot shared assembly process representation method was designed. It decomposes an assembly process into a series of primitive tasks and atomic actions. By rigorously and objectively defining these tasks and actions, the assembly process is uniformly and unambiguously represented, meeting both human intuitive understanding and robotic programming requirements. This work focuses on vision-based methods for understanding assembly processes, as they provide a natural way for robots to interact with humans. However, the lack of datasets capturing complex human assembly actions in industrial scenarios has limited the progress in this field. To address this gap, a human assembly video dataset, HA-ViD, was created. HA-ViD includes common assembly actions, parts, and tools used in industry and comprehensive dual-hand action annotations. Accurately segmenting the sequence of assembly actions is critical for understanding the assembly process, assessing progress, and inferring human intention; therefore, action segmentation is a pivotal technique in this domain. However, existing action segmentation methods are mostly tailored to sports or daily activities and lack the ability to handle the complexity and subtleties of assembly actions. This thesis makes several technical contributions by introducing multiple enhancements to action segmentation algorithms and applying them to various human-robot collaboration (HRC) applications. First, a video-based human fatigue estimation method was developed in the context of human-centric manufacturing. This method introduces a boundary-aware dual-stream action segmentation approach to detect operation types and repetitions, which are then used to estimate human fatigue levels. The estimated fatigue is subsequently applied to optimise human-robot task allocation. Second, a dual-hand action segmentation algorithm, DuHa, was developed to simultaneously segment actions of both hands. To enable real-time performance, DuHa-v2 was proposed, in which the unified action features and implicit object interaction features were designed to reduce computational costs. DuHa-v2 was embedded into a human-robot collaborative assembly framework, which also supports in-process quality checking. The superiority of DuHa and DuHa-v2 over existing action segmentation methods was validated on HA-ViD, and the effectiveness of the proposed framework was confirmed through a real-world case study. Finally, traditional action segmentation methods lack scene adaptability, partly because they conceptualise actions as unified verb-object entities with complete semantics. To overcome this, the dual-hand compositional action segmentation method, DuCAS, was proposed. Instead of segmenting the semantic-complete actions, DuCAS segments action elements—action verb, manipulated object, target object, and tool—and then combines them to form the semantic-complete actions. Furthermore, DuCAS was applied to a framework of multi-modal assembly instruction generation from demonstration videos. In conclusion, this thesis makes four principal contributions advancing assembly video understanding for HRC. First, DH-APR provides the first dual-hand assembly process representation enabling consistent human-robot communication. Second, HA-ViD represents the first industrial assembly dataset with comprehensive dual-hand action annotations. Third, the proposed action segmentation algorithms address existing methods' limitations in hand-object interaction modelling, dual-hand action understanding, real-time capability, and scene adaptability. Fourth, integrated HRC applications demonstrate action understanding's practical effectiveness in fatigue-aware human-robot task allocation, collaborative assembly, and multi-modal instruction generation. Together, these contributions advance the research of both video understanding and HRC."],"dc:identifier.uri":["https://hdl.handle.net/2292/74291"],"dc:publisher":["ResearchSpace@Auckland"],"dc:rights":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."],"dc:rights.uri":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"],"dc:subject":["Smart Manufacturing","Human-Robot Collaboration","Video Understanding","Intelligent Assembly","Industry 5.0"],"dc:title":["Assembly Video Understanding for Human-Robot Collaboration: Methods and Applications"],"dc:type":["Thesis"],"thesis:degree_discipline":["Mechanical Engineering"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["PhD"],"thesis:institution_name":["The University of Auckland"]},"updated_at":"2026-07-24T01:04:52Z"}