University of Illinois at Urbana-Champaign
Does the agent say what it sees: an analysis in 3D question answering
Abstract
dc:descriptionThe emerging research topic of 3D Question and Answering involves complex challenges such as semantic language understanding, 3D scene comprehension, and establishing correspondence between language and target objects. In this thesis, I conduct a comprehensive analysis on the state-of-the-art 3D Question Answering benchmark model, ScanQA, by comparing the effectiveness of various auxiliary tasks, identifying limitations in current evaluation methods, and conducting additional ablation studies on the visual component. The findings and potential future directions for the 3D Question Answering task are discussed to assist researchers in their ongoing investigations.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Zhang, Haomeng
- Contributors dc:contributor
-
- Gui, Liangyan
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2023 Haomeng Zhang
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/120420