Back to results

University of Illinois at Urbana-Champaign

Are language models leaking personal information? Memorization vs. Association

Abstract

dc:description

Large pre-trained language models (PLMs) have transformed the field of natural language processing (NLP) in recent years. PLMs have become the basis for various state-of-the-art NLP systems. Despite the great success of PLMs solving a wide range of NLP tasks, there is rising concern about privacy risks brought with PLMs. For example, recent studies show that PLMs memorize a great portion of training data, including sensitive information, while the information may be leaked unintentionally and utilized by malicious adversaries. In this thesis, we evaluate whether PLMs are prone to leaking personal information and discuss possible reasons behind the privacy leakage. Specifically, we attempt to query PLMs for a target email address with contexts of the email address or prompts containing the owner’s name. We find that PLMs do leak personal information mainly due to memorization. However, the risk of specific personal information being extracted by attackers is low because the models are weak at associating personal identifying information with its owner. We also try to quantify PLMs’ capability of association to help validate the safety of PLM in terms of privacy preserving.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Electrical & Computer Engr
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Shao, Hanyin
Contributors dc:contributor
  • Chang, Kevin Chen-Chuan

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • Copyright 2023 Hanyin Shao
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/120435

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Shao, Hanyin. Are language models leaking personal information? Memorization vs. Association. Thesis thesis, University of Illinois at Urbana-Champaign, 2023. https://hdl.handle.net/2142/120435