{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/72891"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/72891","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Unsupervised video segmentation and its application to activity recognition","abstract":"We addressed the fundamental problem of computer vision: segmentation and recognition, in the space-time domain. With the knowledge that generic image segmentation introduces unstable regions due to illumination, com- pression, etc., we utilized temporal information to achieve consistent 3D video segmentation. By exploiting non-local structure in both spatial and temporal space, the instabilities of the segmented regions were alleviated. A segmentation tree was built within every frame, and the label consistency was enforced within each subtree (i.e. spatial clique). By roughly tracking 2D regions across each frame, temporal clique was built in which label consis- tency was enforced as well. The high-order (more than binary) Conditional Random Field (CRF) is designed and solved efficiently. Experimental results demonstrate high-quality segmentation quantitatively and qualitatively. Taking segmented 3D regions, called tubes, as input, we developed an activity recognition framework not only to determine which activity existed in a video but also to locate where it happens. A robust tube feature was extracted with photometric and shape dynamics information. Activity was described as a Parts Activity Model (PAM) with a root template and four- part template under the root. Given the nature of the activity recognition problem that only some parts on the video were used to determine the activity label, we used Multiple Instance Learning (MIL) to formulate the problem. Latent variables included a tube index and the parts location under the root template. Experiments were conducted on three well-known datasets and a state-of-the-art result was achieved.","abstract_html":"We addressed the fundamental problem of computer vision: segmentation and recognition, in the space-time domain. With the knowledge that generic image segmentation introduces unstable regions due to illumination, com- pression, etc., we utilized temporal information to achieve consistent 3D video segmentation. By exploiting non-local structure in both spatial and temporal space, the instabilities of the segmented regions were alleviated. A segmentation tree was built within every frame, and the label consistency was enforced within each subtree (i.e. spatial clique). By roughly tracking 2D regions across each frame, temporal clique was built in which label consis- tency was enforced as well. The high-order (more than binary) Conditional Random Field (CRF) is designed and solved efficiently. Experimental results demonstrate high-quality segmentation quantitatively and qualitatively. Taking segmented 3D regions, called tubes, as input, we developed an activity recognition framework not only to determine which activity existed in a video but also to locate where it happens. A robust tube feature was extracted with photometric and shape dynamics information. Activity was described as a Parts Activity Model (PAM) with a root template and four- part template under the root. Given the nature of the activity recognition problem that only some parts on the video were used to determine the activity label, we used Multiple Instance Learning (MIL) to formulate the problem. Latent variables included a tube index and the parts location under the root template. Experiments were conducted on three well-known datasets and a state-of-the-art result was achieved.","abstract_has_math":false,"creators":["Cheng, Hsien Ting"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Ahuja, Narendra","Forsyth, David A.","Hasegawa-Johnson, Mark A.","Huang, Thomas"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-01-21T19:49:14Z","date_published":"2015-01-21T19:49:14Z","updated_at":"2026-07-22T22:26:07Z","subjects":["segmentation","Video segmentation","Unsupervised clustering","Activity recognition","Multiple instance learning"],"languages":["en"],"rights":["Copyright 2014 Hsien Ting Cheng"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/72891","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ahuja, Narendra","Forsyth, David A.","Hasegawa-Johnson, Mark A.","Huang, Thomas"]},{"key":"dc:creator","label":"Author","values":["Cheng, Hsien Ting"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-01-21T19:49:14Z","2014-12","2015-01-21"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["segmentation","Video segmentation","Unsupervised clustering","Activity recognition","Multiple instance learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2014 Hsien Ting Cheng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/72891"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We addressed the fundamental problem of computer vision: segmentation and recognition, in the space-time domain. With the knowledge that generic image segmentation introduces unstable regions due to illumination, com- pression, etc., we utilized temporal information to achieve consistent 3D video segmentation. By exploiting non-local structure in both spatial and temporal space, the instabilities of the segmented regions were alleviated. A segmentation tree was built within every frame, and the label consistency was enforced within each subtree (i.e. spatial clique). By roughly tracking 2D regions across each frame, temporal clique was built in which label consis- tency was enforced as well. The high-order (more than binary) Conditional Random Field (CRF) is designed and solved efficiently. Experimental results demonstrate high-quality segmentation quantitatively and qualitatively. Taking segmented 3D regions, called tubes, as input, we developed an activity recognition framework not only to determine which activity existed in a video but also to locate where it happens. A robust tube feature was extracted with photometric and shape dynamics information. Activity was described as a Parts Activity Model (PAM) with a root template and four- part template under the root. Given the nature of the activity recognition problem that only some parts on the video were used to determine the activity label, we used Multiple Instance Learning (MIL) to formulate the problem. Latent variables included a tube index and the parts location under the root template. Experiments were conducted on three well-known datasets and a state-of-the-art result was achieved.","Item withdrawn by Laura Spradlin (lspradl2@illinois.edu) on 2014-10-14T13:35:21Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 3 thesisrefs.bib: 18064 bytes, checksum: e45b3d3dda91f27bb14d3fc3852b63dc (MD5) thesis_final.tex: 134134 bytes, checksum: 763ccf9b80b7d5d1cee9ac7564e09b38 (MD5) Cheng_HsienTing.pdf: 10997419 bytes, checksum: 6f02abfa4d10ce7b301a29fd3c7eb77b (MD5)","Made available in DSpace on 2015-01-21T19:49:14Z (GMT). No. of bitstreams: 3 Hsien Ting_Cheng.pdf: 10997419 bytes, checksum: 6f02abfa4d10ce7b301a29fd3c7eb77b (MD5) thesisrefs.bib: 18064 bytes, checksum: e45b3d3dda91f27bb14d3fc3852b63dc (MD5) thesis_final.tex: 134134 bytes, checksum: 763ccf9b80b7d5d1cee9ac7564e09b38 (MD5)"]},{"key":"dc:title","label":"Title","values":["Unsupervised video segmentation and its application to activity recognition"]}]}],"canonical_facts":{"dc:contributor":["Ahuja, Narendra","Forsyth, David A.","Hasegawa-Johnson, Mark A.","Huang, Thomas"],"dc:creator":["Cheng, Hsien Ting"],"dc:date":["2015-01-21T19:49:14Z","2014-12","2015-01-21"],"dc:description":["We addressed the fundamental problem of computer vision: segmentation and recognition, in the space-time domain. With the knowledge that generic image segmentation introduces unstable regions due to illumination, com- pression, etc., we utilized temporal information to achieve consistent 3D video segmentation. By exploiting non-local structure in both spatial and temporal space, the instabilities of the segmented regions were alleviated. A segmentation tree was built within every frame, and the label consistency was enforced within each subtree (i.e. spatial clique). By roughly tracking 2D regions across each frame, temporal clique was built in which label consis- tency was enforced as well. The high-order (more than binary) Conditional Random Field (CRF) is designed and solved efficiently. Experimental results demonstrate high-quality segmentation quantitatively and qualitatively. Taking segmented 3D regions, called tubes, as input, we developed an activity recognition framework not only to determine which activity existed in a video but also to locate where it happens. A robust tube feature was extracted with photometric and shape dynamics information. Activity was described as a Parts Activity Model (PAM) with a root template and four- part template under the root. Given the nature of the activity recognition problem that only some parts on the video were used to determine the activity label, we used Multiple Instance Learning (MIL) to formulate the problem. Latent variables included a tube index and the parts location under the root template. Experiments were conducted on three well-known datasets and a state-of-the-art result was achieved.","Item withdrawn by Laura Spradlin (lspradl2@illinois.edu) on 2014-10-14T13:35:21Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 3 thesisrefs.bib: 18064 bytes, checksum: e45b3d3dda91f27bb14d3fc3852b63dc (MD5) thesis_final.tex: 134134 bytes, checksum: 763ccf9b80b7d5d1cee9ac7564e09b38 (MD5) Cheng_HsienTing.pdf: 10997419 bytes, checksum: 6f02abfa4d10ce7b301a29fd3c7eb77b (MD5)","Made available in DSpace on 2015-01-21T19:49:14Z (GMT). No. of bitstreams: 3 Hsien Ting_Cheng.pdf: 10997419 bytes, checksum: 6f02abfa4d10ce7b301a29fd3c7eb77b (MD5) thesisrefs.bib: 18064 bytes, checksum: e45b3d3dda91f27bb14d3fc3852b63dc (MD5) thesis_final.tex: 134134 bytes, checksum: 763ccf9b80b7d5d1cee9ac7564e09b38 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/72891"],"dc:language":["en"],"dc:rights":["Copyright 2014 Hsien Ting Cheng"],"dc:subject":["segmentation","Video segmentation","Unsupervised clustering","Activity recognition","Multiple instance learning"],"dc:title":["Unsupervised video segmentation and its application to activity recognition"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:07Z"}