{"id":{"repo_id":"uic","oai_identifier":"oai:figshare.com:article/30425152"},"canonical_url":"https://search.dev.ndltd.org/etd/uic/oai:figshare.com:article/30425152","repository":{"repo_id":"uic","name":"University of Illinois - Chicago","base_url":"https://api.figshare.com/v2/oai"},"display":{"title":"Towards Efficient and Scalable Deep Learning on Graph-Structured Data","abstract":"The practical deployment of Graph Neural Networks (GNNs), a primary form of deep learning on graphs, is hindered by intertwined challenges of effectiveness and scalability. This thesis, \"Towards Effective and Scalable Deep Learning on Graph-Structured Data,\" proposes novel methodologies to address these limitations across four main research thrusts. To address scalability in learning node embeddings, one paper introduces CCA-SSG, a self-supervised framework that learns robust node embeddings. It efficiently avoids the computational burden of negative sampling by using a feature decorrelation objective to prevent representational collapse. To enhance MLP-based models, which are faster but less accurate than GNNs, two papers are presented. OrthoReg tackles an \"over-correlation\" issue with a soft orthogonality constraint, making MLPs competitive with leading GNNs. The second framework achieves true end-to-end MLP efficiency by offloading graph computations to a one-time pre-processing step, eliminating iterative complexity. Finally, for training on massive graphs, this thesis proposes Data-Centric Graph Condensation (DCGC). This framework recasts condensation as a distribution matching problem, creating a small, task-agnostic synthetic graph. This approach significantly improves cross-architecture generalization and reduces condensation time compared to traditional gradient-matching techniques. The proposed models are validated on public benchmarks, demonstrating significant improvements in performance and computational efficiency.","abstract_html":"The practical deployment of Graph Neural Networks (GNNs), a primary form of deep learning on graphs, is hindered by intertwined challenges of effectiveness and scalability. This thesis, &quot;Towards Effective and Scalable Deep Learning on Graph-Structured Data,&quot; proposes novel methodologies to address these limitations across four main research thrusts. To address scalability in learning node embeddings, one paper introduces CCA-SSG, a self-supervised framework that learns robust node embeddings. It efficiently avoids the computational burden of negative sampling by using a feature decorrelation objective to prevent representational collapse. To enhance MLP-based models, which are faster but less accurate than GNNs, two papers are presented. OrthoReg tackles an &quot;over-correlation&quot; issue with a soft orthogonality constraint, making MLPs competitive with leading GNNs. The second framework achieves true end-to-end MLP efficiency by offloading graph computations to a one-time pre-processing step, eliminating iterative complexity. Finally, for training on massive graphs, this thesis proposes Data-Centric Graph Condensation (DCGC). This framework recasts condensation as a distribution matching problem, creating a small, task-agnostic synthetic graph. This approach significantly improves cross-architecture generalization and reduces condensation time compared to traditional gradient-matching techniques. The proposed models are validated on public benchmarks, demonstrating significant improvements in performance and computational efficiency.","abstract_has_math":false,"creators":["Hengrui Zhang (6167507)"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-08-01T00:00:00Z","date_published":"2025-08-01T00:00:00Z","updated_at":"2026-07-27T21:34:47Z","subjects":["Computer Science"],"languages":[],"rights":["In Copyright"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.25417/uic.30425152.v1","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Hengrui Zhang (6167507)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-08-01T00:00:00Z"]},{"key":"dc:relation","label":"Dc Relation","values":["https://figshare.com/articles/thesis/Towards_Efficient_and_Scalable_Deep_Learning_on_Graph-Structured_Data/30425152"]},{"key":"dc:type","label":"Dc Type","values":["Text","Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["10.25417/uic.30425152.v1"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The practical deployment of Graph Neural Networks (GNNs), a primary form of deep learning on graphs, is hindered by intertwined challenges of effectiveness and scalability. This thesis, \"Towards Effective and Scalable Deep Learning on Graph-Structured Data,\" proposes novel methodologies to address these limitations across four main research thrusts. To address scalability in learning node embeddings, one paper introduces CCA-SSG, a self-supervised framework that learns robust node embeddings. It efficiently avoids the computational burden of negative sampling by using a feature decorrelation objective to prevent representational collapse. To enhance MLP-based models, which are faster but less accurate than GNNs, two papers are presented. OrthoReg tackles an \"over-correlation\" issue with a soft orthogonality constraint, making MLPs competitive with leading GNNs. The second framework achieves true end-to-end MLP efficiency by offloading graph computations to a one-time pre-processing step, eliminating iterative complexity. Finally, for training on massive graphs, this thesis proposes Data-Centric Graph Condensation (DCGC). This framework recasts condensation as a distribution matching problem, creating a small, task-agnostic synthetic graph. This approach significantly improves cross-architecture generalization and reduces condensation time compared to traditional gradient-matching techniques. The proposed models are validated on public benchmarks, demonstrating significant improvements in performance and computational efficiency."]},{"key":"dc:title","label":"Title","values":["Towards Efficient and Scalable Deep Learning on Graph-Structured Data"]}]}],"canonical_facts":{"dc:creator":["Hengrui Zhang (6167507)"],"dc:date":["2025-08-01T00:00:00Z"],"dc:description":["The practical deployment of Graph Neural Networks (GNNs), a primary form of deep learning on graphs, is hindered by intertwined challenges of effectiveness and scalability. This thesis, \"Towards Effective and Scalable Deep Learning on Graph-Structured Data,\" proposes novel methodologies to address these limitations across four main research thrusts. To address scalability in learning node embeddings, one paper introduces CCA-SSG, a self-supervised framework that learns robust node embeddings. It efficiently avoids the computational burden of negative sampling by using a feature decorrelation objective to prevent representational collapse. To enhance MLP-based models, which are faster but less accurate than GNNs, two papers are presented. OrthoReg tackles an \"over-correlation\" issue with a soft orthogonality constraint, making MLPs competitive with leading GNNs. The second framework achieves true end-to-end MLP efficiency by offloading graph computations to a one-time pre-processing step, eliminating iterative complexity. Finally, for training on massive graphs, this thesis proposes Data-Centric Graph Condensation (DCGC). This framework recasts condensation as a distribution matching problem, creating a small, task-agnostic synthetic graph. This approach significantly improves cross-architecture generalization and reduces condensation time compared to traditional gradient-matching techniques. The proposed models are validated on public benchmarks, demonstrating significant improvements in performance and computational efficiency."],"dc:identifier":["10.25417/uic.30425152.v1"],"dc:relation":["https://figshare.com/articles/thesis/Towards_Efficient_and_Scalable_Deep_Learning_on_Graph-Structured_Data/30425152"],"dc:rights":["In Copyright"],"dc:subject":["Computer Science"],"dc:title":["Towards Efficient and Scalable Deep Learning on Graph-Structured Data"],"dc:type":["Text","Thesis"]},"updated_at":"2026-07-27T21:34:47Z"}