{"id":{"repo_id":"reykjavik","oai_identifier":"oai:skemman.is:1946/53528"},"canonical_url":"https://search.dev.ndltd.org/etd/reykjavik/oai:skemman.is:1946/53528","repository":{"repo_id":"reykjavik","name":"Reykjavík University","base_url":"https://skemman.is/oai/request"},"display":{"title":"Impacts of LLM-Based Text Normalization on Price Prediction","abstract":"This study investigates the effect of LLM-based text normalization on price prediction from user-generated product descriptions. Using the Mercari Price Suggestion Challenge dataset, we normalize 120,000 item descriptions with GPT-4o-mini and evaluate the impact across three modeling pipelines: a fine-tuned DistilBERT encoder with a linear regression head, a frozen DistilBERT encoder with a linear regression head, and a frozen DistilBERT encoder with an XGBoost regressor. All three pipelines show modest improvement after normalization, with the largest gain in the XGBoost pipeline (1.06% RMSLE) and the smallest in the fine-tuned pipeline (0.21%). Further analysis by description length, price range, and item category reveals that normalization benefits vary substantially across data characteristics and model architectures. The results suggest that normalization benefits are inversely related to model capacity: models with less ability to compensate for noisy input during training benefit most from externally cleaned text. To the best of our knowledge, this is the first study to examine the effect of LLM-based text normalization on a regression task in the e-commerce domain.","abstract_html":"This study investigates the effect of LLM-based text normalization on price prediction from user-generated product descriptions. Using the Mercari Price Suggestion Challenge dataset, we normalize 120,000 item descriptions with GPT-4o-mini and evaluate the impact across three modeling pipelines: a fine-tuned DistilBERT encoder with a linear regression head, a frozen DistilBERT encoder with a linear regression head, and a frozen DistilBERT encoder with an XGBoost regressor. All three pipelines show modest improvement after normalization, with the largest gain in the XGBoost pipeline (1.06% RMSLE) and the smallest in the fine-tuned pipeline (0.21%). Further analysis by description length, price range, and item category reveals that normalization benefits vary substantially across data characteristics and model architectures. The results suggest that normalization benefits are inversely related to model capacity: models with less ability to compensate for noisy input during training benefit most from externally cleaned text. To the best of our knowledge, this is the first study to examine the effect of LLM-based text normalization on a regression task in the e-commerce domain.","abstract_has_math":false,"creators":["Lárus Þóroddsson 2000-","Orri Kristjánsson 1999-","Steinar Örn Sólmundsson 1998-"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Háskólinn í Reykjavík"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-06-05T10:49:46Z","date_published":"2026-06-05T10:49:46Z","updated_at":"2026-07-27T20:38:22Z","subjects":["Tölvunarfræði","Hugbúnaðargerð","Gervigreind","Verðskrár","Price lists","Computer science","Software development","Artificial intelligence"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1946/53528","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Háskólinn í Reykjavík"]},{"key":"dc:creator","label":"Author","values":["Lárus Þóroddsson 2000-","Orri Kristjánsson 1999-","Steinar Örn Sólmundsson 1998-"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-06-05T10:49:42Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-06-05T10:49:42Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-06-05T10:49:46Z"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Tölvunarfræði","Hugbúnaðargerð","Gervigreind","Verðskrár","Price lists","Computer science","Software development","Artificial intelligence"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1946/53528"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This study investigates the effect of LLM-based text normalization on price prediction from user-generated product descriptions. Using the Mercari Price Suggestion Challenge dataset, we normalize 120,000 item descriptions with GPT-4o-mini and evaluate the impact across three modeling pipelines: a fine-tuned DistilBERT encoder with a linear regression head, a frozen DistilBERT encoder with a linear regression head, and a frozen DistilBERT encoder with an XGBoost regressor. All three pipelines show modest improvement after normalization, with the largest gain in the XGBoost pipeline (1.06% RMSLE) and the smallest in the fine-tuned pipeline (0.21%). Further analysis by description length, price range, and item category reveals that normalization benefits vary substantially across data characteristics and model architectures. The results suggest that normalization benefits are inversely related to model capacity: models with less ability to compensate for noisy input during training benefit most from externally cleaned text. To the best of our knowledge, this is the first study to examine the effect of LLM-based text normalization on a regression task in the e-commerce domain."]},{"key":"dc:title","label":"Title","values":["Impacts of LLM-Based Text Normalization on Price Prediction"]}]}],"canonical_facts":{"dc:contributor":["Háskólinn í Reykjavík"],"dc:creator":["Lárus Þóroddsson 2000-","Orri Kristjánsson 1999-","Steinar Örn Sólmundsson 1998-"],"dc:date.accessioned":["2026-06-05T10:49:42Z"],"dc:date.available":["2026-06-05T10:49:42Z"],"dc:date.issued":["2026-06-05T10:49:46Z"],"dc:description.abstract":["This study investigates the effect of LLM-based text normalization on price prediction from user-generated product descriptions. Using the Mercari Price Suggestion Challenge dataset, we normalize 120,000 item descriptions with GPT-4o-mini and evaluate the impact across three modeling pipelines: a fine-tuned DistilBERT encoder with a linear regression head, a frozen DistilBERT encoder with a linear regression head, and a frozen DistilBERT encoder with an XGBoost regressor. All three pipelines show modest improvement after normalization, with the largest gain in the XGBoost pipeline (1.06% RMSLE) and the smallest in the fine-tuned pipeline (0.21%). Further analysis by description length, price range, and item category reveals that normalization benefits vary substantially across data characteristics and model architectures. The results suggest that normalization benefits are inversely related to model capacity: models with less ability to compensate for noisy input during training benefit most from externally cleaned text. To the best of our knowledge, this is the first study to examine the effect of LLM-based text normalization on a regression task in the e-commerce domain."],"dc:identifier.uri":["https://hdl.handle.net/1946/53528"],"dc:language.iso":["en"],"dc:subject":["Tölvunarfræði","Hugbúnaðargerð","Gervigreind","Verðskrár","Price lists","Computer science","Software development","Artificial intelligence"],"dc:title":["Impacts of LLM-Based Text Normalization on Price Prediction"],"dc:type":["Thesis"]},"updated_at":"2026-07-27T20:38:22Z"}