{"id":{"repo_id":"milano","oai_identifier":"oai:air.unimi.it:2434/1038352"},"canonical_url":"https://search.dev.ndltd.org/etd/milano/oai:air.unimi.it:2434/1038352","repository":{"repo_id":"milano","name":"Università degli Studi di Milano","base_url":"https://air.unimi.it/oai/request"},"display":{"title":"ONLINE LEARNING METHODS FOR DIGITAL MARKETS","abstract":"We cast several problems arising from digital markets and economics into an online learning framework, where a learner sequentially interacts with an unknown environment, trying to discover its relevant features to maximize her cumulative reward. After an introduction to online learning in Chapter 1, we start with a study of the bilateral trade problem in Chapter 2. Here, the learner plays the role of a broker whose goal is to increase the value of the market by sequentially interacting with pairs of sellers and buyers, facilitating trades between them. We show how the interplay between the feedback received by the learner and the set of available trading mechanisms affects the attainable regret regimes, devising ad hoc solutions to address the exploration/exploitation dilemma in various types of environments. In Chapter 3, we present an analysis of transparency in repeated first-price auctions. Here, the learner participates in a sequence of first-price auctions to win objects whose exact value is revealed only when she wins the corresponding auction. We show how the level of transparency of the auctioneer (i.e., the amount of information disclosed at the end of each auction) influences the regret rates in different types of environments. In Chapter 4, we study the problem of adaptive optimal taxation. Here, the learner plays the role of a policymaker whose goal is to increase social welfare (seen as a weighted sum of private utility and public revenue) by sequentially setting the tax rate in the labor market. Interestingly, once framed in a formal online learning setting, this problem can be seen as a non-trivial generalization of the classical dynamic pricing problem. Finally, in Chapter 5, we propose an abstract framework generalizing bandits with delayed feedback. This framework allows us to capture scenarios arising frequently in advertising campaigns, where the feedback received comes in the form of delayed and composite income, the sources of which depend on the actions the learner took in the past, and whose exact contributions are not easily identifiable.","abstract_html":"We cast several problems arising from digital markets and economics into an online learning framework, where a learner sequentially interacts with an unknown environment, trying to discover its relevant features to maximize her cumulative reward. After an introduction to online learning in Chapter 1, we start with a study of the bilateral trade problem in Chapter 2. Here, the learner plays the role of a broker whose goal is to increase the value of the market by sequentially interacting with pairs of sellers and buyers, facilitating trades between them. We show how the interplay between the feedback received by the learner and the set of available trading mechanisms affects the attainable regret regimes, devising ad hoc solutions to address the exploration/exploitation dilemma in various types of environments. In Chapter 3, we present an analysis of transparency in repeated first-price auctions. Here, the learner participates in a sequence of first-price auctions to win objects whose exact value is revealed only when she wins the corresponding auction. We show how the level of transparency of the auctioneer (i.e., the amount of information disclosed at the end of each auction) influences the regret rates in different types of environments. In Chapter 4, we study the problem of adaptive optimal taxation. Here, the learner plays the role of a policymaker whose goal is to increase social welfare (seen as a weighted sum of private utility and public revenue) by sequentially setting the tax rate in the labor market. Interestingly, once framed in a formal online learning setting, this problem can be seen as a non-trivial generalization of the classical dynamic pricing problem. Finally, in Chapter 5, we propose an abstract framework generalizing bandits with delayed feedback. This framework allows us to capture scenarios arising frequently in advertising campaigns, where the feedback received comes in the form of delayed and composite income, the sources of which depend on the actions the learner took in the past, and whose exact contributions are not easily identifiable.","abstract_has_math":false,"creators":["COLOMBONI, ROBERTO"],"institution":"Università degli Studi di Milano","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["supervisor: N. Cesa-Bianchi ; co-supervisor: M. Pontil","T. Cesari ; coordinatore: R. Sassi","PONTIL","MASSIMILIANO ; CESARI","TOMMASO","R. Colomboni","CESA BIANCHI, NICOLO' ANTONIO","SASSI, ROBERTO"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-04-17","date_published":"2024-04-17","updated_at":"2026-07-27T20:19:00Z","subjects":["online learning","partial monitoring","multi-armed bandit","two-sided market","first-price auctions","Settore INF/01 - Informatica"],"languages":["eng"],"rights":["info:eu-repo/semantics/openAccess"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["http://dx.doi.org/10.13130/colomboni-roberto_phd2024-04-17","10.13130/colomboni-roberto_phd2024-04-17"],"render_values":[{"text":"http://dx.doi.org/10.13130/colomboni-roberto_phd2024-04-17","href":"http://dx.doi.org/10.13130/colomboni-roberto_phd2024-04-17","code":true},{"text":"10.13130/colomboni-roberto_phd2024-04-17","href":"https://doi.org/10.13130/colomboni-roberto_phd2024-04-17","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/2434/1038352","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["supervisor: N. Cesa-Bianchi ; co-supervisor: M. Pontil","T. Cesari ; coordinatore: R. Sassi","PONTIL","MASSIMILIANO ; CESARI","TOMMASO","R. Colomboni","CESA BIANCHI, NICOLO' ANTONIO","SASSI, ROBERTO"]},{"key":"dc:creator","label":"Author","values":["COLOMBONI, ROBERTO"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-04-17"]},{"key":"dc:publisher","label":"Institution","values":["Università degli Studi di Milano","place:Milano"]},{"key":"dc:relation","label":"Dc Relation","values":["alleditors:PONTIL, MASSIMILIANO ; CESARI, TOMMASO"]},{"key":"dc:type","label":"Dc Type","values":["info:eu-repo/semantics/doctoralThesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["online learning","partial monitoring","multi-armed bandit","two-sided market","first-price auctions","Settore INF/01 - Informatica"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["info:eu-repo/semantics/openAccess"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2434/1038352","http://dx.doi.org/10.13130/colomboni-roberto_phd2024-04-17","10.13130/colomboni-roberto_phd2024-04-17"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We cast several problems arising from digital markets and economics into an online learning framework, where a learner sequentially interacts with an unknown environment, trying to discover its relevant features to maximize her cumulative reward. After an introduction to online learning in Chapter 1, we start with a study of the bilateral trade problem in Chapter 2. Here, the learner plays the role of a broker whose goal is to increase the value of the market by sequentially interacting with pairs of sellers and buyers, facilitating trades between them. We show how the interplay between the feedback received by the learner and the set of available trading mechanisms affects the attainable regret regimes, devising ad hoc solutions to address the exploration/exploitation dilemma in various types of environments. In Chapter 3, we present an analysis of transparency in repeated first-price auctions. Here, the learner participates in a sequence of first-price auctions to win objects whose exact value is revealed only when she wins the corresponding auction. We show how the level of transparency of the auctioneer (i.e., the amount of information disclosed at the end of each auction) influences the regret rates in different types of environments. In Chapter 4, we study the problem of adaptive optimal taxation. Here, the learner plays the role of a policymaker whose goal is to increase social welfare (seen as a weighted sum of private utility and public revenue) by sequentially setting the tax rate in the labor market. Interestingly, once framed in a formal online learning setting, this problem can be seen as a non-trivial generalization of the classical dynamic pricing problem. Finally, in Chapter 5, we propose an abstract framework generalizing bandits with delayed feedback. This framework allows us to capture scenarios arising frequently in advertising campaigns, where the feedback received comes in the form of delayed and composite income, the sources of which depend on the actions the learner took in the past, and whose exact contributions are not easily identifiable."]},{"key":"dc:title","label":"Title","values":["ONLINE LEARNING METHODS FOR DIGITAL MARKETS"]}]}],"canonical_facts":{"dc:contributor":["supervisor: N. Cesa-Bianchi ; co-supervisor: M. Pontil","T. Cesari ; coordinatore: R. Sassi","PONTIL","MASSIMILIANO ; CESARI","TOMMASO","R. Colomboni","CESA BIANCHI, NICOLO' ANTONIO","SASSI, ROBERTO"],"dc:creator":["COLOMBONI, ROBERTO"],"dc:date":["2024-04-17"],"dc:description":["We cast several problems arising from digital markets and economics into an online learning framework, where a learner sequentially interacts with an unknown environment, trying to discover its relevant features to maximize her cumulative reward. After an introduction to online learning in Chapter 1, we start with a study of the bilateral trade problem in Chapter 2. Here, the learner plays the role of a broker whose goal is to increase the value of the market by sequentially interacting with pairs of sellers and buyers, facilitating trades between them. We show how the interplay between the feedback received by the learner and the set of available trading mechanisms affects the attainable regret regimes, devising ad hoc solutions to address the exploration/exploitation dilemma in various types of environments. In Chapter 3, we present an analysis of transparency in repeated first-price auctions. Here, the learner participates in a sequence of first-price auctions to win objects whose exact value is revealed only when she wins the corresponding auction. We show how the level of transparency of the auctioneer (i.e., the amount of information disclosed at the end of each auction) influences the regret rates in different types of environments. In Chapter 4, we study the problem of adaptive optimal taxation. Here, the learner plays the role of a policymaker whose goal is to increase social welfare (seen as a weighted sum of private utility and public revenue) by sequentially setting the tax rate in the labor market. Interestingly, once framed in a formal online learning setting, this problem can be seen as a non-trivial generalization of the classical dynamic pricing problem. Finally, in Chapter 5, we propose an abstract framework generalizing bandits with delayed feedback. This framework allows us to capture scenarios arising frequently in advertising campaigns, where the feedback received comes in the form of delayed and composite income, the sources of which depend on the actions the learner took in the past, and whose exact contributions are not easily identifiable."],"dc:identifier":["https://hdl.handle.net/2434/1038352","http://dx.doi.org/10.13130/colomboni-roberto_phd2024-04-17","10.13130/colomboni-roberto_phd2024-04-17"],"dc:language":["eng"],"dc:publisher":["Università degli Studi di Milano","place:Milano"],"dc:relation":["alleditors:PONTIL, MASSIMILIANO ; CESARI, TOMMASO"],"dc:rights":["info:eu-repo/semantics/openAccess"],"dc:subject":["online learning","partial monitoring","multi-armed bandit","two-sided market","first-price auctions","Settore INF/01 - Informatica"],"dc:title":["ONLINE LEARNING METHODS FOR DIGITAL MARKETS"],"dc:type":["info:eu-repo/semantics/doctoralThesis"]},"updated_at":"2026-07-27T20:19:00Z"}