Back to results

Università degli Studi di Milano

ONLINE LEARNING METHODS FOR DIGITAL MARKETS

Abstract

dc:description

We cast several problems arising from digital markets and economics into an online learning framework, where a learner sequentially interacts with an unknown environment, trying to discover its relevant features to maximize her cumulative reward. After an introduction to online learning in Chapter 1, we start with a study of the bilateral trade problem in Chapter 2. Here, the learner plays the role of a broker whose goal is to increase the value of the market by sequentially interacting with pairs of sellers and buyers, facilitating trades between them. We show how the interplay between the feedback received by the learner and the set of available trading mechanisms affects the attainable regret regimes, devising ad hoc solutions to address the exploration/exploitation dilemma in various types of environments. In Chapter 3, we present an analysis of transparency in repeated first-price auctions. Here, the learner participates in a sequence of first-price auctions to win objects whose exact value is revealed only when she wins the corresponding auction. We show how the level of transparency of the auctioneer (i.e., the amount of information disclosed at the end of each auction) influences the regret rates in different types of environments. In Chapter 4, we study the problem of adaptive optimal taxation. Here, the learner plays the role of a policymaker whose goal is to increase social welfare (seen as a weighted sum of private utility and public revenue) by sequentially setting the tax rate in the labor market. Interestingly, once framed in a formal online learning setting, this problem can be seen as a non-trivial generalization of the classical dynamic pricing problem. Finally, in Chapter 5, we propose an abstract framework generalizing bandits with delayed feedback. This framework allows us to capture scenarios arising frequently in advertising campaigns, where the feedback received comes in the form of delayed and composite income, the sources of which depend on the actions the learner took in the past, and whose exact contributions are not easily identifiable.

Degree

thesis:*
Grantor dc:publisher
Università degli Studi di Milano
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • COLOMBONI, ROBERTO
Contributors dc:contributor
  • supervisor: N. Cesa-Bianchi ; co-supervisor: M. Pontil
  • T. Cesari ; coordinatore: R. Sassi
  • PONTIL
  • MASSIMILIANO ; CESARI
  • TOMMASO
  • R. Colomboni
  • CESA BIANCHI, NICOLO' ANTONIO
  • SASSI, ROBERTO

Subjects

dc:subject × 6

Rights

dc:rights
Statement dc:rights
  • info:eu-repo/semantics/openAccess
Language dc:language
eng

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:air.unimi.it:2434/1038352

Chain of custody

source
Harvested from
Università degli Studi di Milano
Base URL
air.unimi.it/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

COLOMBONI, ROBERTO. ONLINE LEARNING METHODS FOR DIGITAL MARKETS. Università degli Studi di Milano, 2024. https://hdl.handle.net/2434/1038352