{"id":{"repo_id":"cornell","oai_identifier":"oai:ecommons.cornell.edu:1813/59050"},"canonical_url":"https://search.dev.ndltd.org/etd/cornell/oai:ecommons.cornell.edu:1813/59050","repository":{"repo_id":"cornell","name":"Cornell University","base_url":"https://ecommons.cornell.edu/server/oai/request"},"display":{"title":"Adaptive Preference Learning With Bandit Feedback: Information Filtering, Dueling Bandits and Incentivizing Exploration","abstract":"In this thesis, we study adaptive preference learning, in which a machine learning system learns users' preferences from feedback while simultaneously using these learned preferences to help them find preferred items. We study three different types of user feedback in three application setting: cardinal feedback with application in information filtering systems, ordinal feedback with application in personalized content recommender systems, and attribute feedback with application in review aggregators. We connect these settings respectively to existing work on classical multi-armed bandits, dueling bandits, and incentivizing exploration. For each type of feedback and application setting, we provide an algorithm and a theoretical analysis bounding its regret. We demonstrate through numerical experiments that our algorithms outperform existing benchmarks.","abstract_html":"In this thesis, we study adaptive preference learning, in which a machine learning system learns users&#x27; preferences from feedback while simultaneously using these learned preferences to help them find preferred items. We study three different types of user feedback in three application setting: cardinal feedback with application in information filtering systems, ordinal feedback with application in personalized content recommender systems, and attribute feedback with application in review aggregators. We connect these settings respectively to existing work on classical multi-armed bandits, dueling bandits, and incentivizing exploration. For each type of feedback and application setting, we provide an algorithm and a theoretical analysis bounding its regret. We demonstrate through numerical experiments that our algorithms outperform existing benchmarks.","abstract_has_math":false,"creators":["Chen, Bangrui"],"institution":"Cornell University","degree_name":"Ph. D., Operations Research","degree_level":"Doctor of Philosophy","degree_discipline":"Operations Research","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":["Topaloglu, Huseyin","Joachims, Thorsten"],"year":2017,"date_issued":"2017-12-30","date_published":"2017-12-30","updated_at":"2026-07-24T01:49:04Z","subjects":["Statistics","Operations research","Computer science","adaptive preference learning","bandit feedback","dueling bandits","incentivizing exploration","information filtering","multi-armed bandits"],"languages":["en_US"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.7298/X4251GCQ"],"render_values":[{"text":"https://doi.org/10.7298/X4251GCQ","href":"https://doi.org/10.7298/X4251GCQ","code":true}]},{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["ProQuest Submission ID: 10605","ProQuest Publication ID: 10680541"],"render_values":[{"text":"ProQuest Submission ID: 10605","href":null,"code":true},{"text":"ProQuest Publication ID: 10680541","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/1813/59050","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Topaloglu, Huseyin","Joachims, Thorsten"]},{"key":"dc:creator","label":"Author","values":["Chen, Bangrui"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2018-10-03T19:27:23Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2018-10-03T19:27:23Z"]},{"key":"dc:date.issued","label":"Date","values":["2017-12-30"]},{"key":"dc:type","label":"Dc Type","values":["dissertation or thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Operations Research"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctor of Philosophy"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph. D., Operations Research"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Cornell University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Statistics","Operations research","Computer science","adaptive preference learning","bandit feedback","dueling bandits","incentivizing exploration","information filtering","multi-armed bandits"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.7298/X4251GCQ"]},{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["ProQuest Submission ID: 10605","ProQuest Publication ID: 10680541"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1813/59050"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In this thesis, we study adaptive preference learning, in which a machine learning system learns users' preferences from feedback while simultaneously using these learned preferences to help them find preferred items. We study three different types of user feedback in three application setting: cardinal feedback with application in information filtering systems, ordinal feedback with application in personalized content recommender systems, and attribute feedback with application in review aggregators. We connect these settings respectively to existing work on classical multi-armed bandits, dueling bandits, and incentivizing exploration. For each type of feedback and application setting, we provide an algorithm and a theoretical analysis bounding its regret. We demonstrate through numerical experiments that our algorithms outperform existing benchmarks."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Adaptive Preference Learning With Bandit Feedback: Information Filtering, Dueling Bandits and Incentivizing Exploration"]}]}],"canonical_facts":{"dc:contributor.committeemember":["Topaloglu, Huseyin","Joachims, Thorsten"],"dc:creator":["Chen, Bangrui"],"dc:date.accessioned":["2018-10-03T19:27:23Z"],"dc:date.available":["2018-10-03T19:27:23Z"],"dc:date.issued":["2017-12-30"],"dc:description.abstract":["In this thesis, we study adaptive preference learning, in which a machine learning system learns users' preferences from feedback while simultaneously using these learned preferences to help them find preferred items. We study three different types of user feedback in three application setting: cardinal feedback with application in information filtering systems, ordinal feedback with application in personalized content recommender systems, and attribute feedback with application in review aggregators. We connect these settings respectively to existing work on classical multi-armed bandits, dueling bandits, and incentivizing exploration. For each type of feedback and application setting, we provide an algorithm and a theoretical analysis bounding its regret. We demonstrate through numerical experiments that our algorithms outperform existing benchmarks."],"dc:format.mimetype":["application/pdf"],"dc:identifier.doi":["https://doi.org/10.7298/X4251GCQ"],"dc:identifier.other":["ProQuest Submission ID: 10605","ProQuest Publication ID: 10680541"],"dc:identifier.uri":["https://hdl.handle.net/1813/59050"],"dc:language.iso":["en_US"],"dc:subject":["Statistics","Operations research","Computer science","adaptive preference learning","bandit feedback","dueling bandits","incentivizing exploration","information filtering","multi-armed bandits"],"dc:title":["Adaptive Preference Learning With Bandit Feedback: Information Filtering, Dueling Bandits and Incentivizing Exploration"],"dc:type":["dissertation or thesis"],"thesis:degree_discipline":["Operations Research"],"thesis:degree_level":["Doctor of Philosophy"],"thesis:degree_name":["Ph. D., Operations Research"],"thesis:institution_name":["Cornell University"]},"updated_at":"2026-07-24T01:49:04Z"}