{"id":{"repo_id":"brock","oai_identifier":"oai:brocku.scholaris.ca:10464/19106"},"canonical_url":"https://search.dev.ndltd.org/etd/brock/oai:brocku.scholaris.ca:10464/19106","repository":{"repo_id":"brock","name":"Brock University","base_url":"https://brocku.scholaris.ca/server/oai/request"},"display":{"title":"Leveraging the Latent Space for Model Understanding and Optimization","abstract":"In recent years, the field of machine learning has seen massive growth in both the size and quality of models and performance on tasks such as classification or image generation. However, these models are typically limited by two key factors. First, models such as those used in tasks of text-to-image generation lack interpretation. Second, models that leverage the latent space to represent data struggle to capture high-level details. This often results in reconstructions which do not accurately represent the original data. First, to address the issue of interpretability in text-to-image models, we introduce WINOVIS, a novel dataset designed to probe models in their ability to interpret textual prompts. This approach reframes the task of pronoun disambiguation from a single mode of natural language to a multi-model problem involving both visual and textual understanding. Second, we turn our focus to models in image generation such as the VQ-VAE which often struggle to reconstruct images capturing the finer details of the original input image. By introducing lightweight and straightforward modifications to the VQ-VAE’s loss function and dictionary selection process, we enable the reconstruction of images that retain high-level details often absent from the reconstructions produced by the traditional VQ-VAE.","abstract_html":"In recent years, the field of machine learning has seen massive growth in both the size and quality of models and performance on tasks such as classification or image generation. However, these models are typically limited by two key factors. First, models such as those used in tasks of text-to-image generation lack interpretation. Second, models that leverage the latent space to represent data struggle to capture high-level details. This often results in reconstructions which do not accurately represent the original data. First, to address the issue of interpretability in text-to-image models, we introduce WINOVIS, a novel dataset designed to probe models in their ability to interpret textual prompts. This approach reframes the task of pronoun disambiguation from a single mode of natural language to a multi-model problem involving both visual and textual understanding. Second, we turn our focus to models in image generation such as the VQ-VAE which often struggle to reconstruct images capturing the finer details of the original input image. By introducing lightweight and straightforward modifications to the VQ-VAE’s loss function and dictionary selection process, we enable the reconstruction of images that retain high-level details often absent from the reconstructions produced by the traditional VQ-VAE.","abstract_has_math":false,"creators":["Park, Brendan"],"institution":"Brock University","degree_name":"M.Sc. Computer Science","degree_level":"Masters","degree_discipline":"Faculty of Mathematics and Science","degree_department":"Department of Computer Science","school":null,"contributors":[],"advisors":["Li, Yifeng"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-03-26T12:52:00Z","date_published":"2025-03-26T12:52:00Z","updated_at":"2026-07-24T01:23:04Z","subjects":["TECHNOLOGY::Information technology::Computer science::Computer science"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10464/19106","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Li, Yifeng"]},{"key":"dc:contributor.department","label":"Department","values":["Department of Computer Science"]},{"key":"dc:creator","label":"Author","values":["Park, Brendan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-03-26T12:52:00Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-03-26T12:52:00Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-03-26T12:52:00Z"]},{"key":"dc:type","label":"Dc Type","values":["Electronic Thesis or Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Faculty of Mathematics and Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.Sc. Computer Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Brock University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["TECHNOLOGY::Information technology::Computer science::Computer science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10464/19106"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In recent years, the field of machine learning has seen massive growth in both the size and quality of models and performance on tasks such as classification or image generation. However, these models are typically limited by two key factors. First, models such as those used in tasks of text-to-image generation lack interpretation. Second, models that leverage the latent space to represent data struggle to capture high-level details. This often results in reconstructions which do not accurately represent the original data. First, to address the issue of interpretability in text-to-image models, we introduce WINOVIS, a novel dataset designed to probe models in their ability to interpret textual prompts. This approach reframes the task of pronoun disambiguation from a single mode of natural language to a multi-model problem involving both visual and textual understanding. Second, we turn our focus to models in image generation such as the VQ-VAE which often struggle to reconstruct images capturing the finer details of the original input image. By introducing lightweight and straightforward modifications to the VQ-VAE’s loss function and dictionary selection process, we enable the reconstruction of images that retain high-level details often absent from the reconstructions produced by the traditional VQ-VAE."]},{"key":"dc:title","label":"Title","values":["Leveraging the Latent Space for Model Understanding and Optimization"]}]}],"canonical_facts":{"dc:contributor.advisor":["Li, Yifeng"],"dc:contributor.department":["Department of Computer Science"],"dc:creator":["Park, Brendan"],"dc:date.accessioned":["2025-03-26T12:52:00Z"],"dc:date.available":["2025-03-26T12:52:00Z"],"dc:date.issued":["2025-03-26T12:52:00Z"],"dc:description.abstract":["In recent years, the field of machine learning has seen massive growth in both the size and quality of models and performance on tasks such as classification or image generation. However, these models are typically limited by two key factors. First, models such as those used in tasks of text-to-image generation lack interpretation. Second, models that leverage the latent space to represent data struggle to capture high-level details. This often results in reconstructions which do not accurately represent the original data. First, to address the issue of interpretability in text-to-image models, we introduce WINOVIS, a novel dataset designed to probe models in their ability to interpret textual prompts. This approach reframes the task of pronoun disambiguation from a single mode of natural language to a multi-model problem involving both visual and textual understanding. Second, we turn our focus to models in image generation such as the VQ-VAE which often struggle to reconstruct images capturing the finer details of the original input image. By introducing lightweight and straightforward modifications to the VQ-VAE’s loss function and dictionary selection process, we enable the reconstruction of images that retain high-level details often absent from the reconstructions produced by the traditional VQ-VAE."],"dc:identifier.uri":["https://hdl.handle.net/10464/19106"],"dc:language.iso":["eng"],"dc:subject":["TECHNOLOGY::Information technology::Computer science::Computer science"],"dc:title":["Leveraging the Latent Space for Model Understanding and Optimization"],"dc:type":["Electronic Thesis or Dissertation"],"thesis:degree_discipline":["Faculty of Mathematics and Science"],"thesis:degree_level":["Masters"],"thesis:degree_name":["M.Sc. Computer Science"],"thesis:institution_name":["Brock University"]},"updated_at":"2026-07-24T01:23:04Z"}