Back to results

University of Illinois Urbana-Champaign

Towards similarity learning in security applications

Abstract

dc:description

In today’s world, large amounts of data remain unlabeled, posing a major challenge, especially in security applications, where acquiring high-quality labels is costly and difficult. Without accurate labels, it is hard to train reliable machine learning (ML) models, which limits their effectiveness in real-world scenarios. Similarity learning provides a promising direction by capturing relationships within the data without requiring explicit labels. Instead, it learns from reference pairs by measuring similarity through distance. While simple distance metrics can be used , similarity learning is often combined with deep learning to learn robust feature representations for comparison using predefined or learned similarity measures. How reliable is similarity learning in real-world security applications, particularly when exposed to adversarial threats? Under what conditions can it enhance model generalization and detection performance? This dissertation evaluates the robustness of similarity-learning applications under a realistic threat model by applying adversarial attacks end-to-end, and shows how similarity learning can improve out-of-distribution (OOD) generalization in graph-structured data. Specifically, Chapter 3 presents adversarial attacks targeting perceptual hashing-based reverse image search engines, which use Hamming distance as the similarity metric. By developing advanced attacks and evaluating them end-to-end on real-world systems, our framework successfully subverts several major reverse image search engines. In Chapter 4, we present attacks on vision-based phishing detectors trained using similarity learning. Our framework generates adversarial logos that preserve original brand semantics while bypassing state-of-the-art visual phishing website detectors. Chapter 5 explores how similarity learning, specifically graph contrastive learning (GCL), can complement supervised learning to improve out-of-distribution generalization in graph neural networks (GNNs) under natural distribution shifts. In summary, these studies show that similarity learning-based applications are vulnerable to adversarial attacks, highlighting the need for stronger defenses under realistic end-to-end threat models. At the same time, similarity learning can complement supervised methods by providing diverse feature representations and decision signals, making it valuable for improving out-of-distribution detection under natural distribution shifts.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois Urbana-Champaign
Year dc:date
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Hao, Qingying
Contributors dc:contributor
  • Wang, Gang
  • Gunter, Carl
  • Li, Bo
  • Chandrasekaran, Varun
  • Conti, Mauro

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2025 Qingying Hao
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/130183

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Hao, Qingying. Towards similarity learning in security applications. Dissertation thesis, University of Illinois Urbana-Champaign, 2025. https://hdl.handle.net/2142/130183