{"id":{"repo_id":"temple","oai_identifier":"oai:scholarshare.temple.edu:20.500.12613/12145"},"canonical_url":"https://search.dev.ndltd.org/etd/temple/oai:scholarshare.temple.edu:20.500.12613/12145","repository":{"repo_id":"temple","name":"Temple University","base_url":"https://scholarshare.temple.edu/server/oai/request"},"display":{"title":"Evaluating interactions and decision-making by annotators and users in computer vision systems","abstract":"Understanding how humans interact with and are influenced by intelligent systems is essential for improving their design and effectiveness. Although computer vision research has largely centered on algorithmic advances, these increasingly complex systems are ultimately used by people whose decisions and interpretations shape how they function in practice. As models grow more sophisticated and are rapidly deployed in real-world settings, their internal processes often appear as ``black boxes,'' especially to non-experts, creating uncertainty about how to interpret outputs, communicate intent, and how much to rely on system feedback. This disconnect underscores the importance of integrating humans into the design and evaluation of computer vision systems to ensure alignment with user needs and capabilities. This dissertation addresses this gap through a comprehensive evaluation of human involvement throughout the modern computer vision pipeline. We focus on two primary roles: annotators, who create and refine training data, and users, who engage with deployed systems and rely on explanations to inform their decisions. These roles are central to key stages of the modern computer vision pipeline where human decision-making, input, and interpretation directly impact system performance and outcomes: data labeling, training, deployment, and explainability. We conduct human-centered evaluations across these four stages. For the data labeling stage, we examined how varying the amount of context available to annotators influenced efficiency and accuracy in an object matching task. Reduced context improved efficiency without compromising accuracy, while additional context was helpful when objects were less distinctive or image quality was poor. At the training stage, where annotators’ norms and biases shape the training data, we substituted gendered terms in existing image captioning datasets with gender-neutral equivalents to reduce gender bias and examine the impact on model outputs and perceived caption quality. Models trained on neutral data produced fewer gendered descriptions while maintaining, and in some cases, improving descriptive quality. During deployment, we compared two interaction modes for intent communication in multimodal instruction-based image editing systems: post-edit correction and proactive clarification. Although both modes led to similar task performance, post-edit correction helped users better understand how the system functioned and experience a greater sense of communicative ease. Finally, for explainability, we evaluated the impact of AI-generated visual explanations on users’ decision-making in AI-assisted systems. Although the explanations had no impact on performance or confidence, they supported the development of better mental models of system behavior. Users’ AI literacy shaped how the task was approached and how the explanations were utilized. By centering the roles of annotators and users, these contributions collectively identify opportunities to improve efficiency, interpretability, and the overall effectiveness of human interactions across these stages. This dissertation aims to guide the development of human-centered computer vision systems that effectively meet the needs of diverse users.","abstract_html":"Understanding how humans interact with and are influenced by intelligent systems is essential for improving their design and effectiveness. Although computer vision research has largely centered on algorithmic advances, these increasingly complex systems are ultimately used by people whose decisions and interpretations shape how they function in practice. As models grow more sophisticated and are rapidly deployed in real-world settings, their internal processes often appear as ``black boxes,&#x27;&#x27; especially to non-experts, creating uncertainty about how to interpret outputs, communicate intent, and how much to rely on system feedback. This disconnect underscores the importance of integrating humans into the design and evaluation of computer vision systems to ensure alignment with user needs and capabilities. This dissertation addresses this gap through a comprehensive evaluation of human involvement throughout the modern computer vision pipeline. We focus on two primary roles: annotators, who create and refine training data, and users, who engage with deployed systems and rely on explanations to inform their decisions. These roles are central to key stages of the modern computer vision pipeline where human decision-making, input, and interpretation directly impact system performance and outcomes: data labeling, training, deployment, and explainability. We conduct human-centered evaluations across these four stages. For the data labeling stage, we examined how varying the amount of context available to annotators influenced efficiency and accuracy in an object matching task. Reduced context improved efficiency without compromising accuracy, while additional context was helpful when objects were less distinctive or image quality was poor. At the training stage, where annotators’ norms and biases shape the training data, we substituted gendered terms in existing image captioning datasets with gender-neutral equivalents to reduce gender bias and examine the impact on model outputs and perceived caption quality. Models trained on neutral data produced fewer gendered descriptions while maintaining, and in some cases, improving descriptive quality. During deployment, we compared two interaction modes for intent communication in multimodal instruction-based image editing systems: post-edit correction and proactive clarification. Although both modes led to similar task performance, post-edit correction helped users better understand how the system functioned and experience a greater sense of communicative ease. Finally, for explainability, we evaluated the impact of AI-generated visual explanations on users’ decision-making in AI-assisted systems. Although the explanations had no impact on performance or confidence, they supported the development of better mental models of system behavior. Users’ AI literacy shaped how the task was approached and how the explanations were utilized. By centering the roles of annotators and users, these contributions collectively identify opportunities to improve efficiency, interpretability, and the overall effectiveness of human interactions across these stages. This dissertation aims to guide the development of human-centered computer vision systems that effectively meet the needs of diverse users.","abstract_has_math":false,"creators":["Wazzan, Albatool"],"institution":"Temple University. Libraries","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Souvenir, Richard M."],"committee_chairs":[],"committee_members":["MacNeil, Stephen, 1987-","Obradovic, Zoran","Tallapragada, Meghnaa"],"year":2026,"date_issued":"2026-05","date_published":"2026-05","updated_at":"2026-07-27T21:22:14Z","subjects":["Computer science","Human-centered computer vision","Computer and information science"],"languages":["eng"],"rights":["IN COPYRIGHT- This Rights Statement can be used for an Item that is in copyright. Using this statement implies that the organization making this Item available has determined that the Item is in copyright and either is the rights-holder, has obtained permission from the rights-holder(s) to make their Work(s) available, or makes the Item available under an exception or limitation to copyright (including Fair Use) that entitles it to make the Item available."],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://scholarshare.temple.edu/handle/20.500.12613/12145","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Souvenir, Richard M."]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["MacNeil, Stephen, 1987-","Obradovic, Zoran","Tallapragada, Meghnaa"]},{"key":"dc:creator","label":"Author","values":["Wazzan, Albatool"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-06-10T13:35:38Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-06-10T13:35:38Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-05"]},{"key":"dc:publisher","label":"Institution","values":["Temple University. Libraries"]},{"key":"dc:type","label":"Dc Type","values":["Text"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer science","Human-centered computer vision","Computer and information science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["IN COPYRIGHT- This Rights Statement can be used for an Item that is in copyright. Using this statement implies that the organization making this Item available has determined that the Item is in copyright and either is the rights-holder, has obtained permission from the rights-holder(s) to make their Work(s) available, or makes the Item available under an exception or limitation to copyright (including Fair Use) that entitles it to make the Item available."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarshare.temple.edu/handle/20.500.12613/12145"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Understanding how humans interact with and are influenced by intelligent systems is essential for improving their design and effectiveness. Although computer vision research has largely centered on algorithmic advances, these increasingly complex systems are ultimately used by people whose decisions and interpretations shape how they function in practice. As models grow more sophisticated and are rapidly deployed in real-world settings, their internal processes often appear as ``black boxes,'' especially to non-experts, creating uncertainty about how to interpret outputs, communicate intent, and how much to rely on system feedback. This disconnect underscores the importance of integrating humans into the design and evaluation of computer vision systems to ensure alignment with user needs and capabilities. This dissertation addresses this gap through a comprehensive evaluation of human involvement throughout the modern computer vision pipeline. We focus on two primary roles: annotators, who create and refine training data, and users, who engage with deployed systems and rely on explanations to inform their decisions. These roles are central to key stages of the modern computer vision pipeline where human decision-making, input, and interpretation directly impact system performance and outcomes: data labeling, training, deployment, and explainability. We conduct human-centered evaluations across these four stages. For the data labeling stage, we examined how varying the amount of context available to annotators influenced efficiency and accuracy in an object matching task. Reduced context improved efficiency without compromising accuracy, while additional context was helpful when objects were less distinctive or image quality was poor. At the training stage, where annotators’ norms and biases shape the training data, we substituted gendered terms in existing image captioning datasets with gender-neutral equivalents to reduce gender bias and examine the impact on model outputs and perceived caption quality. Models trained on neutral data produced fewer gendered descriptions while maintaining, and in some cases, improving descriptive quality. During deployment, we compared two interaction modes for intent communication in multimodal instruction-based image editing systems: post-edit correction and proactive clarification. Although both modes led to similar task performance, post-edit correction helped users better understand how the system functioned and experience a greater sense of communicative ease. Finally, for explainability, we evaluated the impact of AI-generated visual explanations on users’ decision-making in AI-assisted systems. Although the explanations had no impact on performance or confidence, they supported the development of better mental models of system behavior. Users’ AI literacy shaped how the task was approached and how the explanations were utilized. By centering the roles of annotators and users, these contributions collectively identify opportunities to improve efficiency, interpretability, and the overall effectiveness of human interactions across these stages. This dissertation aims to guide the development of human-centered computer vision systems that effectively meet the needs of diverse users."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Ph.D."]},{"key":"dc:title","label":"Title","values":["Evaluating interactions and decision-making by annotators and users in computer vision systems"]}]}],"canonical_facts":{"dc:contributor.advisor":["Souvenir, Richard M."],"dc:contributor.committeemember":["MacNeil, Stephen, 1987-","Obradovic, Zoran","Tallapragada, Meghnaa"],"dc:creator":["Wazzan, Albatool"],"dc:date.accessioned":["2026-06-10T13:35:38Z"],"dc:date.available":["2026-06-10T13:35:38Z"],"dc:date.issued":["2026-05"],"dc:description.abstract":["Understanding how humans interact with and are influenced by intelligent systems is essential for improving their design and effectiveness. Although computer vision research has largely centered on algorithmic advances, these increasingly complex systems are ultimately used by people whose decisions and interpretations shape how they function in practice. As models grow more sophisticated and are rapidly deployed in real-world settings, their internal processes often appear as ``black boxes,'' especially to non-experts, creating uncertainty about how to interpret outputs, communicate intent, and how much to rely on system feedback. This disconnect underscores the importance of integrating humans into the design and evaluation of computer vision systems to ensure alignment with user needs and capabilities. This dissertation addresses this gap through a comprehensive evaluation of human involvement throughout the modern computer vision pipeline. We focus on two primary roles: annotators, who create and refine training data, and users, who engage with deployed systems and rely on explanations to inform their decisions. These roles are central to key stages of the modern computer vision pipeline where human decision-making, input, and interpretation directly impact system performance and outcomes: data labeling, training, deployment, and explainability. We conduct human-centered evaluations across these four stages. For the data labeling stage, we examined how varying the amount of context available to annotators influenced efficiency and accuracy in an object matching task. Reduced context improved efficiency without compromising accuracy, while additional context was helpful when objects were less distinctive or image quality was poor. At the training stage, where annotators’ norms and biases shape the training data, we substituted gendered terms in existing image captioning datasets with gender-neutral equivalents to reduce gender bias and examine the impact on model outputs and perceived caption quality. Models trained on neutral data produced fewer gendered descriptions while maintaining, and in some cases, improving descriptive quality. During deployment, we compared two interaction modes for intent communication in multimodal instruction-based image editing systems: post-edit correction and proactive clarification. Although both modes led to similar task performance, post-edit correction helped users better understand how the system functioned and experience a greater sense of communicative ease. Finally, for explainability, we evaluated the impact of AI-generated visual explanations on users’ decision-making in AI-assisted systems. Although the explanations had no impact on performance or confidence, they supported the development of better mental models of system behavior. Users’ AI literacy shaped how the task was approached and how the explanations were utilized. By centering the roles of annotators and users, these contributions collectively identify opportunities to improve efficiency, interpretability, and the overall effectiveness of human interactions across these stages. This dissertation aims to guide the development of human-centered computer vision systems that effectively meet the needs of diverse users."],"dc:description.degree":["Ph.D."],"dc:identifier.uri":["https://scholarshare.temple.edu/handle/20.500.12613/12145"],"dc:language.iso":["eng"],"dc:publisher":["Temple University. Libraries"],"dc:rights":["IN COPYRIGHT- This Rights Statement can be used for an Item that is in copyright. Using this statement implies that the organization making this Item available has determined that the Item is in copyright and either is the rights-holder, has obtained permission from the rights-holder(s) to make their Work(s) available, or makes the Item available under an exception or limitation to copyright (including Fair Use) that entitles it to make the Item available."],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Computer science","Human-centered computer vision","Computer and information science"],"dc:title":["Evaluating interactions and decision-making by annotators and users in computer vision systems"],"dc:type":["Text"]},"updated_at":"2026-07-27T21:22:14Z"}