Helsingin yliopisto
Methodological and extrinsic challenges in offline evaluation of recommender systems
Abstract
dc:description.abstractRecent studies in recommender systems have shown that performance improvements achieved by state-of-the-art methods could be attributed to the misapplication of offline evaluation protocols. These results highlight the potential for both methodological and extrinsic aspects of evaluation to impact results. Furthermore, specialised recommendation tasks, such as sequential and cross-domain recommendation, additionally pose unique and underexplored challenges for evaluation. In this thesis, we present a series of studies examining different aspects of evaluation in top-n recommendation, sequential recommendation, and cross-domain recommendation. The first study examined sampled evaluation metrics in top-n recommendation. Our results demonstrate greater consistency between sampled and traditional (non-sampled) metrics compared to prior studies, as well as additional advantages, such as the potential for higher discriminative power and robustness against popularity bias in sampled metrics. The second study looked at how the most widely-used data splitting strategy in offline evaluation of sequential recommenders is deeply flawed due to data leakage, which results in performance being significantly overestimated. Our third study investigates user perceptions of cross-domain recommendations. The results show that simply telling users recommendations were based on information from a different domain significantly altered their perceptions of recommendations, lowering both trust and interest. In the final study, we propose a novel meta-evaluation framework based on techniques from psychometric assessment that can be used to investigate various aspects of offline evaluation, including evaluation metrics, the information content of data sets and the role of item popularity and user engagement in recommendation. This thesis contributes to our understanding of a diverse range of evaluation challenges in recommender systems. The findings raise critical concerns regarding the evaluation of recommender systems, showing that standard evaluation methods can result in questionable research findings. Additionally, this thesis contains insights into how evaluation can be improved for several recommendation tasks.
Degree
thesis:*- Grantor dc:publisher
- Helsingin yliopisto
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Liu, Yang
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited.
- Julkaisu on tekijänoikeussäännösten alainen. Teosta voi lukea ja tulostaa henkilökohtaista käyttöä varten. Käyttö kaupallisiin tarkoituksiin on kielletty.
- Publikationen är skyddad av upphovsrätten. Den får läsas och skrivas ut för personligt bruk. Användning i kommersiellt syfte är förbjuden.
- Language dc:language.iso
- eng
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/10138/598761