{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132611"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132611","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Analysis of errors in generative image and video models","abstract":"Image and video generative models have become increasingly powerful and produce visuals that are more difficult to distinguish from real ones in recent years. Users can create images and videos to make their imaginations come true easier than ever before. However, on closer examination, these models make interesting mistakes. In this thesis, we first discuss a systematic way of analyzing errors in generated images relating to projective geometry and shadows at a population level. We find that generated images can be reliably distinguished from real images by derived geometric features alone without looking at pixels. Then, we introduce methods that can effectively judge whether a generated video is physically plausible for robotic demonstrations.","abstract_html":"Image and video generative models have become increasingly powerful and produce visuals that are more difficult to distinguish from real ones in recent years. Users can create images and videos to make their imaginations come true easier than ever before. However, on closer examination, these models make interesting mistakes. In this thesis, we first discuss a systematic way of analyzing errors in generated images relating to projective geometry and shadows at a population level. We find that generated images can be reliably distinguished from real images by derived geometric features alone without looking at pixels. Then, we introduce methods that can effectively judge whether a generated video is physically plausible for robotic demonstrations.","abstract_has_math":false,"creators":["Mai, Hanlin"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Lazebnik, Svetlana"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Generative AI","Diffusion Models"],"languages":["en"],"rights":["Copyright 2025 Hanlin Mai"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132611","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lazebnik, Svetlana"]},{"key":"dc:creator","label":"Author","values":["Mai, Hanlin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-12-11"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Generative AI","Diffusion Models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Hanlin Mai"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132611"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Image and video generative models have become increasingly powerful and produce visuals that are more difficult to distinguish from real ones in recent years. Users can create images and videos to make their imaginations come true easier than ever before. However, on closer examination, these models make interesting mistakes. In this thesis, we first discuss a systematic way of analyzing errors in generated images relating to projective geometry and shadows at a population level. We find that generated images can be reliably distinguished from real images by derived geometric features alone without looking at pixels. Then, we introduce methods that can effectively judge whether a generated video is physically plausible for robotic demonstrations.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-12-01","The student, Hanlin Mai, accepted the attached license on 2025-12-11 at 10:55.","The student, Hanlin Mai, submitted this Thesis for approval on 2025-12-11 at 11:05.","This Thesis was approved for publication on 2025-12-11 at 11:45.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22744 on 2026-02-19 at 18:45:09"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Analysis of errors in generative image and video models"]}]}],"canonical_facts":{"dc:contributor":["Lazebnik, Svetlana"],"dc:creator":["Mai, Hanlin"],"dc:date":["2025-12","2025-12-11"],"dc:description":["Image and video generative models have become increasingly powerful and produce visuals that are more difficult to distinguish from real ones in recent years. Users can create images and videos to make their imaginations come true easier than ever before. However, on closer examination, these models make interesting mistakes. In this thesis, we first discuss a systematic way of analyzing errors in generated images relating to projective geometry and shadows at a population level. We find that generated images can be reliably distinguished from real images by derived geometric features alone without looking at pixels. Then, we introduce methods that can effectively judge whether a generated video is physically plausible for robotic demonstrations.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-12-01","The student, Hanlin Mai, accepted the attached license on 2025-12-11 at 10:55.","The student, Hanlin Mai, submitted this Thesis for approval on 2025-12-11 at 11:05.","This Thesis was approved for publication on 2025-12-11 at 11:45.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22744 on 2026-02-19 at 18:45:09"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132611"],"dc:language":["en"],"dc:rights":["Copyright 2025 Hanlin Mai"],"dc:subject":["Generative AI","Diffusion Models"],"dc:title":["Analysis of errors in generative image and video models"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}