Cornell University
Adaptive Preference Learning With Bandit Feedback: Information Filtering, Dueling Bandits and Incentivizing Exploration
Abstract
dc:description.abstractIn this thesis, we study adaptive preference learning, in which a machine learning system learns users' preferences from feedback while simultaneously using these learned preferences to help them find preferred items. We study three different types of user feedback in three application setting: cardinal feedback with application in information filtering systems, ordinal feedback with application in personalized content recommender systems, and attribute feedback with application in review aggregators. We connect these settings respectively to existing work on classical multi-armed bandits, dueling bandits, and incentivizing exploration. For each type of feedback and application setting, we provide an algorithm and a theoretical analysis bounding its regret. We demonstrate through numerical experiments that our algorithms outperform existing benchmarks.
Degree
thesis:*- Name thesis:degree_name
- Ph. D., Operations Research
- Level thesis:degree_level
- Doctor of Philosophy
- Discipline thesis:degree_discipline
- Operations Research
- Grantor
- Cornell University
- Year dc:date.issued
- 2017
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Chen, Bangrui
- Committee members dc:contributor.committeemember
-
- Topaloglu, Huseyin
- Joachims, Thorsten
Subjects
dc:subject × 9Rights
- Language dc:language.iso
- en_US
Identifiers
dc:identifier.*- Dc Identifier Other
-
ProQuest Submission ID: 10605
ProQuest Publication ID: 10680541 - OAI identifier oai:identifier
- oai:ecommons.cornell.edu:1813/59050