{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/125606"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/125606","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Dynamic 3D Gaussian tracking for graph-based neural dynamics modeling","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-02-04 without embargo terms","abstract_has_math":false,"creators":["Zhang, Mingtong"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Li, Yunzhu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-07-15","date_published":"2024-07-15","updated_at":"2026-07-22T22:25:02Z","subjects":["Dynamics Model","3d Gaussian Splatting","Action-conditioned Video Prediction","Model-based Planning"],"languages":["en","eng"],"rights":["Copyright 2024 Mingtong Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/125606","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Li, Yunzhu"]},{"key":"dc:creator","label":"Author","values":["Zhang, Mingtong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-07-15","2024-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Dynamics Model","3d Gaussian Splatting","Action-conditioned Video Prediction","Model-based Planning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Mingtong Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/125606"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","The student, Mingtong Zhang, accepted the attached license on 2024-07-10 at 13:14.","The student, Mingtong Zhang, submitted this Thesis for approval on 2024-07-10 at 13:27.","This Thesis was approved for publication on 2024-07-15 at 14:45.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21045 on 2025-02-04 at 21:04:53","Videos of robots interacting with objects encode rich information about the objects' dynamics. However, existing video prediction approaches typically do not explicitly account for the 3D information from videos, such as robot actions and objects' 3D states, limiting their use in real-world robotic applications. In this work, we introduce a comprehensive framework to learn object dynamics directly from multi-view RGB videos by explicitly considering the robot's action trajectories and their effects on scene dynamics. Our approach utilizes the 3D Gaussian representation of 3D Gaussian Splatting (3DGS) to train a particle-based dynamics model using Graph Neural Networks (GNNs). This model operates on sparse control particles downsampled from the densely tracked 3D Gaussian reconstructions, ensuring that the critical dynamics of the scene are captured efficiently and accurately. By learning the neural dynamics model on offline robot interaction data, our method can predict object motions under varying initial configurations and unseen robot actions, providing a robust tool for real-world applications. The 3D transformations of Gaussians can be interpolated from the motions of control particles, enabling the rendering of predicted future object states and achieving action-conditioned video prediction. This capability allows for the generation of realistic and accurate future scenarios based on the robot's actions, facilitating advanced planning and decision-making processes. Furthermore, the dynamics model can be integrated into model-based planning frameworks for object manipulation tasks. By leveraging the learned dynamics, robots can plan and execute complex manipulation tasks with higher precision and reliability. This integration is particularly beneficial for tasks involving deformable materials, where accurate prediction of object behavior is crucial. We conduct extensive experiments on various kinds of deformable materials, including ropes, clothes, and stuffed animals, demonstrating our framework's ability to model complex shapes and dynamics. Our results show significant improvements in prediction accuracy and visual fidelity compared to existing methods, highlighting the effectiveness of incorporating 3D information and explicit action conditioning in dynamics modeling."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Dynamic 3D Gaussian tracking for graph-based neural dynamics modeling"]}]}],"canonical_facts":{"dc:contributor":["Li, Yunzhu"],"dc:creator":["Zhang, Mingtong"],"dc:date":["2024-07-15","2024-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","The student, Mingtong Zhang, accepted the attached license on 2024-07-10 at 13:14.","The student, Mingtong Zhang, submitted this Thesis for approval on 2024-07-10 at 13:27.","This Thesis was approved for publication on 2024-07-15 at 14:45.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21045 on 2025-02-04 at 21:04:53","Videos of robots interacting with objects encode rich information about the objects' dynamics. However, existing video prediction approaches typically do not explicitly account for the 3D information from videos, such as robot actions and objects' 3D states, limiting their use in real-world robotic applications. In this work, we introduce a comprehensive framework to learn object dynamics directly from multi-view RGB videos by explicitly considering the robot's action trajectories and their effects on scene dynamics. Our approach utilizes the 3D Gaussian representation of 3D Gaussian Splatting (3DGS) to train a particle-based dynamics model using Graph Neural Networks (GNNs). This model operates on sparse control particles downsampled from the densely tracked 3D Gaussian reconstructions, ensuring that the critical dynamics of the scene are captured efficiently and accurately. By learning the neural dynamics model on offline robot interaction data, our method can predict object motions under varying initial configurations and unseen robot actions, providing a robust tool for real-world applications. The 3D transformations of Gaussians can be interpolated from the motions of control particles, enabling the rendering of predicted future object states and achieving action-conditioned video prediction. This capability allows for the generation of realistic and accurate future scenarios based on the robot's actions, facilitating advanced planning and decision-making processes. Furthermore, the dynamics model can be integrated into model-based planning frameworks for object manipulation tasks. By leveraging the learned dynamics, robots can plan and execute complex manipulation tasks with higher precision and reliability. This integration is particularly beneficial for tasks involving deformable materials, where accurate prediction of object behavior is crucial. We conduct extensive experiments on various kinds of deformable materials, including ropes, clothes, and stuffed animals, demonstrating our framework's ability to model complex shapes and dynamics. Our results show significant improvements in prediction accuracy and visual fidelity compared to existing methods, highlighting the effectiveness of incorporating 3D information and explicit action conditioning in dynamics modeling."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/125606"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Mingtong Zhang"],"dc:subject":["Dynamics Model","3d Gaussian Splatting","Action-conditioned Video Prediction","Model-based Planning"],"dc:title":["Dynamic 3D Gaussian tracking for graph-based neural dynamics modeling"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}