Abstract
dc:description.abstractThis study investigates the effect of LLM-based text normalization on price prediction from user-generated product descriptions. Using the Mercari Price Suggestion Challenge dataset, we normalize 120,000 item descriptions with GPT-4o-mini and evaluate the impact across three modeling pipelines: a fine-tuned DistilBERT encoder with a linear regression head, a frozen DistilBERT encoder with a linear regression head, and a frozen DistilBERT encoder with an XGBoost regressor. All three pipelines show modest improvement after normalization, with the largest gain in the XGBoost pipeline (1.06% RMSLE) and the smallest in the fine-tuned pipeline (0.21%). Further analysis by description length, price range, and item category reveals that normalization benefits vary substantially across data characteristics and model architectures. The results suggest that normalization benefits are inversely related to model capacity: models with less ability to compensate for noisy input during training benefit most from externally cleaned text. To the best of our knowledge, this is the first study to examine the effect of LLM-based text normalization on a regression task in the e-commerce domain.
Author and committee
dc:creator, dc:contributor.*- Authors dc:creator
-
- Lárus Þóroddsson 2000-
- Orri Kristjánsson 1999-
- Steinar Örn Sólmundsson 1998-
- Contributors dc:contributor
-
- Háskólinn í Reykjavík
Subjects
dc:subject × 8Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/1946/53528
- OAI identifier oai:identifier
- oai:skemman.is:1946/53528