{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/113216"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/113216","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Sonification of the Scene in the Image Environment and Metaverse Using Natural Language","abstract":"This metaverse and computer vision-powered application is designed to serve people with low vision or a visual impairment, ranging from adults to old age. Specifically, we hope to improve the situational awareness of users in a scene by narrating the visual content from their point of view. The user would be able to understand the information through auditory channels as the system would narrate the scene's description using speech technology. This could increase the accessibility of visual-spatial information for the users in a metaverse and later in the physical world. This solution is designed and developed considering the hypothesis that if we enable the narration of a scene's visual content, we can increase the understanding and access to that scene. This study paves the way for VR technology to be used as a training and exploration tool not limited to blind people in generic environments, but applicable to specific domains such as military, healthcare, or architecture and planning. We have run a user study and evaluated our hypothesis about which set of algorithms will perform better for a specific category of tasks - like search or survey - and evaluated the narration algorithms by the user's ratings of naturalness, correctness and satisfaction. The tasks and algorithms have been discussed in detail in the chapters of this thesis.","abstract_html":"This metaverse and computer vision-powered application is designed to serve people with low vision or a visual impairment, ranging from adults to old age. Specifically, we hope to improve the situational awareness of users in a scene by narrating the visual content from their point of view. The user would be able to understand the information through auditory channels as the system would narrate the scene&#x27;s description using speech technology. This could increase the accessibility of visual-spatial information for the users in a metaverse and later in the physical world. This solution is designed and developed considering the hypothesis that if we enable the narration of a scene&#x27;s visual content, we can increase the understanding and access to that scene. This study paves the way for VR technology to be used as a training and exploration tool not limited to blind people in generic environments, but applicable to specific domains such as military, healthcare, or architecture and planning. We have run a user study and evaluated our hypothesis about which set of algorithms will perform better for a specific category of tasks - like search or survey - and evaluated the narration algorithms by the user&#x27;s ratings of naturalness, correctness and satisfaction. The tasks and algorithms have been discussed in detail in the chapters of this thesis.","abstract_has_math":false,"creators":["Wasi, Mohd Sheeban"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Computer Science and Applications","degree_department":"Computer Science and Applications","school":null,"contributors":[],"advisors":[],"committee_chairs":["Polys, Nicholas F."],"committee_members":["McCrickard, D. Scott","Bukvic, Ivica Ico"],"year":2023,"date_issued":"2023-01-17","date_published":"2023-01-17","updated_at":"2026-07-22T22:20:25Z","subjects":["Metaverse","Machine Learning","Computer Vision","X3DOM"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:36223"],"render_values":[{"text":"vt_gsexam:36223","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/10919/113216","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Polys, Nicholas F."]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["McCrickard, D. Scott","Bukvic, Ivica Ico"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science and Applications"]},{"key":"dc:creator","label":"Author","values":["Wasi, Mohd Sheeban"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2023-01-18T09:00:47Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2023-01-18T09:00:47Z"]},{"key":"dc:date.issued","label":"Date","values":["2023-01-17"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science and Applications"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Metaverse","Machine Learning","Computer Vision","X3DOM"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:36223"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/10919/113216"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This metaverse and computer vision-powered application is designed to serve people with low vision or a visual impairment, ranging from adults to old age. Specifically, we hope to improve the situational awareness of users in a scene by narrating the visual content from their point of view. The user would be able to understand the information through auditory channels as the system would narrate the scene's description using speech technology. This could increase the accessibility of visual-spatial information for the users in a metaverse and later in the physical world. This solution is designed and developed considering the hypothesis that if we enable the narration of a scene's visual content, we can increase the understanding and access to that scene. This study paves the way for VR technology to be used as a training and exploration tool not limited to blind people in generic environments, but applicable to specific domains such as military, healthcare, or architecture and planning. We have run a user study and evaluated our hypothesis about which set of algorithms will perform better for a specific category of tasks - like search or survey - and evaluated the narration algorithms by the user's ratings of naturalness, correctness and satisfaction. The tasks and algorithms have been discussed in detail in the chapters of this thesis."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["The solution is built using an object detection algorithm and virtual environments which run on the web browser using X3DOM. The solution would help improve situational awareness for normal people as well as for low vision individuals through speech. On a broader scale, we seek to contribute to accessibility solutions. We have designed four algorithms which will help user to understand the scene information through auditory channels as the system would narrate the scene's description using speech technology. The idea would increase the accessibility of visual-spatial information for the users in a metaverse and later in the physical world."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Sonification of the Scene in the Image Environment and Metaverse Using Natural Language"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Polys, Nicholas F."],"dc:contributor.committeemember":["McCrickard, D. Scott","Bukvic, Ivica Ico"],"dc:contributor.department":["Computer Science and Applications"],"dc:creator":["Wasi, Mohd Sheeban"],"dc:date.accessioned":["2023-01-18T09:00:47Z"],"dc:date.available":["2023-01-18T09:00:47Z"],"dc:date.issued":["2023-01-17"],"dc:description.abstract":["This metaverse and computer vision-powered application is designed to serve people with low vision or a visual impairment, ranging from adults to old age. Specifically, we hope to improve the situational awareness of users in a scene by narrating the visual content from their point of view. The user would be able to understand the information through auditory channels as the system would narrate the scene's description using speech technology. This could increase the accessibility of visual-spatial information for the users in a metaverse and later in the physical world. This solution is designed and developed considering the hypothesis that if we enable the narration of a scene's visual content, we can increase the understanding and access to that scene. This study paves the way for VR technology to be used as a training and exploration tool not limited to blind people in generic environments, but applicable to specific domains such as military, healthcare, or architecture and planning. We have run a user study and evaluated our hypothesis about which set of algorithms will perform better for a specific category of tasks - like search or survey - and evaluated the narration algorithms by the user's ratings of naturalness, correctness and satisfaction. The tasks and algorithms have been discussed in detail in the chapters of this thesis."],"dc:description.abstractgeneral":["The solution is built using an object detection algorithm and virtual environments which run on the web browser using X3DOM. The solution would help improve situational awareness for normal people as well as for low vision individuals through speech. On a broader scale, we seek to contribute to accessibility solutions. We have designed four algorithms which will help user to understand the scene information through auditory channels as the system would narrate the scene's description using speech technology. The idea would increase the accessibility of visual-spatial information for the users in a metaverse and later in the physical world."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:36223"],"dc:identifier.uri":["http://hdl.handle.net/10919/113216"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Metaverse","Machine Learning","Computer Vision","X3DOM"],"dc:title":["Sonification of the Scene in the Image Environment and Metaverse Using Natural Language"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science and Applications"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:20:25Z"}