Back to results

University of Minnesota

Multiple Choice Question Answering using a Large Corpus of Information

Abstract

dc:description.abstract

The amount of natural language data is massive and the potential to harness the information contained within has led to many recent discoveries. In this dissertation I explore only one aspect of learning with the goal of answering multiple choice questions with information from a large corpus of information. I chose this topic because of an internship at NASA’s Jet Propulsion Laboratory, where there is a growing interest in making rovers more autonomous in their field research. Being able to process information and act correctly is a key stepping stone to accomplish this, which is an aspect my dissertation covers. The chapters involve a review on the early embedding methods, and two novel approaches to create multiple choice question answering mechanisms. In Chapter 2 I review popular algorithms to create word and sentence embeddings given the surrounding context. These embeddings are a numerical representation of the language data that can be used in downhill models such as logistic regression. In Chapter 3 I present a novel method to create a domain specific knowledge base that can be querired to answer multiple choice questions from a database of Elementary School science questions. The knowledge base is made up of a graph structure and trained using deep learning techniques. The classifier creates an embedding to represent the question and answers. This embedding is then passed through a feed forward network to determine the probability of a correct answer. We train on questions and general information from a large corpus in a semi-supervised setting. In Chapter 4 I propose a strategy to train a network to simultaneously classify multiple choice questions and learn to generate words relevant to the surrounding context of the question. Using the Transformer architecture in a Generative Adversarial Network as well as an additional classifier is a novel approach to train a network that is robust against data not seen in the training set. This semi-supervised training regiment also uses sentences from a large corpus of information and Reinforcement Learning to better inform the generator of relevant words

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Kinney, Mitchell

Subjects

dc:subject × 4

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/11299/216372
OAI identifier oai:identifier
oai:conservancy.umn.edu:11299/216372

Chain of custody

source
Harvested from
University of Minnesota
Base URL
conservancy.umn.edu/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Kinney, Mitchell. Multiple Choice Question Answering using a Large Corpus of Information. 2020. http://hdl.handle.net/11299/216372