{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/121984"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/121984","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Some modules of hierarchical video parsing with transformers for activity localization and recognition","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2024-03-01 without embargo terms","abstract_has_math":false,"creators":["Yu, Mengxuan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Ahuja, Narendra"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-12","date_published":"2023-12","updated_at":"2026-07-22T22:25:00Z","subjects":["Computer Vision","Ai","Video Parsing"],"languages":["en","eng"],"rights":["Copyright 2023 Mengxuan Yu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/121984","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ahuja, Narendra"]},{"key":"dc:creator","label":"Author","values":["Yu, Mengxuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-12","2023-12-01"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Vision","Ai","Video Parsing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Mengxuan Yu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/121984"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","The student, Mengxuan Yu, accepted the attached license on 2023-11-16 at 15:46.","The student, Mengxuan Yu, submitted this Thesis for approval on 2023-11-16 at 15:50.","This Thesis was approved for publication on 2023-12-01 at 10:51.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19916 on 2024-03-01 at 13:14:30","This thesis presents a set of modules of a method for human activity video parsing, with temporal action recognition and localization. The previous works have already achieved very high performances. However, many of them are focusing on short video clips with a single label. The new method described includes a way to parse human activity videos with a sequence of action labels, complex environment, and arbitrary long background clips (the part of the video in which nothing happens). The method applies an encoder combined with LSTM and a self-attentive Transformer to the video frame feature sequence extracted by the I3D model. It uses multiple parsing methods such as CYK parsing and probabilistic inference to decode the result and build the parsing tree efficiently and accurately. The method gives a performance that is a significant improvement in accuracy compared to SoTA methods. The modules presented in this thesis are: (1)Video Tree structure and Vocabulary (2)Video CYK Parsing algorithm (3)Video Grammar Probability Tree, and (4)Mean Average Precision testing"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Some modules of hierarchical video parsing with transformers for activity localization and recognition"]}]}],"canonical_facts":{"dc:contributor":["Ahuja, Narendra"],"dc:creator":["Yu, Mengxuan"],"dc:date":["2023-12","2023-12-01"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","The student, Mengxuan Yu, accepted the attached license on 2023-11-16 at 15:46.","The student, Mengxuan Yu, submitted this Thesis for approval on 2023-11-16 at 15:50.","This Thesis was approved for publication on 2023-12-01 at 10:51.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19916 on 2024-03-01 at 13:14:30","This thesis presents a set of modules of a method for human activity video parsing, with temporal action recognition and localization. The previous works have already achieved very high performances. However, many of them are focusing on short video clips with a single label. The new method described includes a way to parse human activity videos with a sequence of action labels, complex environment, and arbitrary long background clips (the part of the video in which nothing happens). The method applies an encoder combined with LSTM and a self-attentive Transformer to the video frame feature sequence extracted by the I3D model. It uses multiple parsing methods such as CYK parsing and probabilistic inference to decode the result and build the parsing tree efficiently and accurately. The method gives a performance that is a significant improvement in accuracy compared to SoTA methods. The modules presented in this thesis are: (1)Video Tree structure and Vocabulary (2)Video CYK Parsing algorithm (3)Video Grammar Probability Tree, and (4)Mean Average Precision testing"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/121984"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Mengxuan Yu"],"dc:subject":["Computer Vision","Ai","Video Parsing"],"dc:title":["Some modules of hierarchical video parsing with transformers for activity localization and recognition"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:00Z"}