{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124589"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124589","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Human motion synthesis and compression","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2026-05-01","abstract_has_math":false,"creators":["Li, Zhengyuan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Gui, Liangyan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:02Z","subjects":["Human Motion"],"languages":["en","eng"],"rights":["Copyright 2024 Zhengyuan Li"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124589","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Gui, Liangyan"]},{"key":"dc:creator","label":"Author","values":["Li, Zhengyuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-04-30"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Human Motion"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Zhengyuan Li"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124589"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","The student, Zhengyuan Li, accepted the attached license on 2024-04-26 at 15:45.","The student, Zhengyuan Li, submitted this Thesis for approval on 2024-04-26 at 15:49.","This Thesis was approved for publication on 2024-04-30 at 10:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20670 on 2024-09-16 at 00:44:49","The synthesis of human motion plays a pivotal role in applications ranging from character animation to autonomous driving. Recent advances in human motion synthesis are driven by powerful denoising diffusion models and transformer architectures. This thesis explores two fundamental challenges in human motion synthesis: designing effective architectural frameworks and developing motion compression components with strong reconstruction capabilities and well-conditioned latent spaces. While the current transformer architectures are predominantly temporal-focused, the spatial structure is inherent in human body. we introduce Positional Mask-Guided Spatial-Temporal Fusion (\\ours) -- a novel approach to modeling human motion in a bi-dimensional manner, thus enabling a more nuanced generation of human behavior. Specifically, we design a spatial-temporal transformer architecture with homogeneous and symmetric dual branches for learning representations from human motion sequences. To facilitate the refined interplay between spatial and temporal features, we propose positional masks to guide the fusion process. Extensive experiments demonstrate the state-of-the-art performance of \\ours across tasks and datasets. Efficiently compressing human motion sequences allows for a significant reduction in computational overhead and facilitates more complex analyses and synthesis in constrained environments. In order to build an effective two-person motion compression model, researchers should identify the crucial loss terms, adapt adequate network architecture, and control the variance in the latent space. Through exhaustive experiments, the thesis offers deep insights into the optimal design of motion compression systems for future applications. From those two aspects, the work paves the way for the future research in human motion synthesis."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Human motion synthesis and compression"]}]}],"canonical_facts":{"dc:contributor":["Gui, Liangyan"],"dc:creator":["Li, Zhengyuan"],"dc:date":["2024-05","2024-04-30"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","The student, Zhengyuan Li, accepted the attached license on 2024-04-26 at 15:45.","The student, Zhengyuan Li, submitted this Thesis for approval on 2024-04-26 at 15:49.","This Thesis was approved for publication on 2024-04-30 at 10:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20670 on 2024-09-16 at 00:44:49","The synthesis of human motion plays a pivotal role in applications ranging from character animation to autonomous driving. Recent advances in human motion synthesis are driven by powerful denoising diffusion models and transformer architectures. This thesis explores two fundamental challenges in human motion synthesis: designing effective architectural frameworks and developing motion compression components with strong reconstruction capabilities and well-conditioned latent spaces. While the current transformer architectures are predominantly temporal-focused, the spatial structure is inherent in human body. we introduce Positional Mask-Guided Spatial-Temporal Fusion (\\ours) -- a novel approach to modeling human motion in a bi-dimensional manner, thus enabling a more nuanced generation of human behavior. Specifically, we design a spatial-temporal transformer architecture with homogeneous and symmetric dual branches for learning representations from human motion sequences. To facilitate the refined interplay between spatial and temporal features, we propose positional masks to guide the fusion process. Extensive experiments demonstrate the state-of-the-art performance of \\ours across tasks and datasets. Efficiently compressing human motion sequences allows for a significant reduction in computational overhead and facilitates more complex analyses and synthesis in constrained environments. In order to build an effective two-person motion compression model, researchers should identify the crucial loss terms, adapt adequate network architecture, and control the variance in the latent space. Through exhaustive experiments, the thesis offers deep insights into the optimal design of motion compression systems for future applications. From those two aspects, the work paves the way for the future research in human motion synthesis."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124589"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Zhengyuan Li"],"dc:subject":["Human Motion"],"dc:title":["Human motion synthesis and compression"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}