University of Illinois at Urbana-Champaign
Node classification with extremely few labels and applications to social networks
Abstract
dc:descriptionThe node classification tasks aim to classify nodes within graph datasets into several classes. The current state of the art can be summarized as the ‘pre-training, fine-tuning’ framework and the ‘pre-training, prompt-tuning’ framework, where the input graph(s) is first pre-trained without knowing the downstream tasks via general-purpose graph learning objectives; and then fine-tuned or prompt-tuned for the downstream tasks using task-specific objectives. Despite multiple previous works of graph pre-training and tuning methods for node classification tasks, current methods still require sufficient labeled nodes for good performance. This thesis proposes a series of novel pre-training and tuning approaches for extremely few-shot node classification and zero-shot clustering tasks. Our framework can be applied to most graph datasets but is also flexible to extend to specific social network applications, including polarization detection and truth-finding. We highlight the following contributions: • We propose a novel interaction-level contrastive objective for graph pre-training. The proposed pre-training paradigm enables better task transferability to node-level and edge-level downstream tasks and finer-grained contrastive pair generation for effective contrastive learning. • We extend the above objective in two social network applications as case studies: polarization detection and truth-finding, by integrating tasks-specific objectives. • We propose a virtual node generation method for graph fine-tuning that optimally expands the propagation of few-shot labels. We first formulate the expected graph propagation from virtual node generation and then propose an efficient solution for optimally deriving the virtual node sets by maximizing node classification confidence of fine-tuning propagation. • We extend the virtual node generation framework into an active learning solution with truth-finding tasks as a case study. • We propose a unified graph prompt-tuning paradigm, which improves the alignment of task formulation between graph pre-training and downstream prompt-tuning. Such alignment further improves transferability from pre-training to downstream objectives by mirroring their formulations.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Cui, Hang
- Contributors dc:contributor
-
- Abdelzaher, Tarek
- Han, Jiawei
- Kaplan, Lance
- Banerjee, Arindam
Subjects
dc:subject × 6Rights
dc:rights- Statement dc:rights
-
- Copyright 2024 Hang Cui
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/127263