{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105057"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105057","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Learning compact neural network representations with structural priors","abstract":"DSpace SAF Submission Ingestion Package generated from Vireo submission #13747 on 2019-08-22 at 15:07:23","abstract_html":"DSpace SAF Submission Ingestion Package generated from Vireo submission #13747 on 2019-08-22 at 15:07:23","abstract_has_math":false,"creators":["Han, Wei"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Huang, Thomas S.","Hasegawa-Johnson, Mark","Liang, Zhi-Pei","Hwu, Wen-Mei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T20:36:02Z","date_published":"2019-08-23T20:36:02Z","updated_at":"2026-07-22T22:24:44Z","subjects":["neural networks","recurrent neural networks","convolutional neural networks"],"languages":["en"],"rights":["Copyright 2019 Wei Han"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105057","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Huang, Thomas S.","Hasegawa-Johnson, Mark","Liang, Zhi-Pei","Hwu, Wen-Mei"]},{"key":"dc:creator","label":"Author","values":["Han, Wei"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T20:36:02Z","2021-08-24T09:15:16Z","2019-04-18","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["neural networks","recurrent neural networks","convolutional neural networks"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Wei Han"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105057"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["DSpace SAF Submission Ingestion Package generated from Vireo submission #13747 on 2019-08-22 at 15:07:23","Made available in DSpace on 2019-08-23T20:36:02Z (GMT). No. of bitstreams: 2 HAN-DISSERTATION-2019.pdf: 1941888 bytes, checksum: 3e73bb3f2035c5446203dcfa3ad7bb3b (MD5) LICENSE.txt: 4204 bytes, checksum: 9bff5a3e92ba1fbbab9a73eae99a7816 (MD5) Previous issue date: 2019-04-18","Embargo set by: Seth Robbins for item 112176 Lift date: 2021-08-23T20:36:18Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 112176 on 2021-08-24T09:15:16Z.","The development of deep neural networks has taken two directions. On one hand, the networks become deeper and wider, employing drastically more model parameters and consuming more training data. On the other hand, the simplification of the internal structure of the neural networks has also contributed to the success of neural networks. Notably, two important families of neural networks, convolutional neural networks (CNN) and recurrent neural networks (RNN), both introduce certain structural priors and share model parameters internally to simplify the network. In this dissertation, we investigate a few alternative neural network structural priors to learn more compact CNN and RNN models and achieve better parameter-to-performance ratios. We have developed these neural network structures at three different abstraction levels with different target applications for each. First of all, motivated by the filter redundancy in convolutional neural networks, we have studied parameter sharing across filters in a convolutional layer. Instead of the conventional approach treating CNN filters as a set of independent model parameters, we explore the 2D spatial correlation among filters and propose sharing the filter parameters as if they are overlapping slices from a shared 3D tensor. Experiments show the proposed approach can effectively reduce the number of parameters in several state-of-the-art CNN architectures while still maintaining competitive performance. The second problem studied in this dissertation concerns the inter-layer connectivity pattern in RNNs, the family of neural network models specifically designed for time-series data. One longstanding challenge in RNNs is the vanishing gradient problem that hinders the model's capability to learn long-term dependency, the temporal data dependency across many time steps. Motivated by the recent development from two largely unrelated fields---general application of skip-connection to neural networks, and dilated convolution in image and audio related problems---we propose \\textsc{DilatedRNN}, a simple but principled way to construct multi-layer RNNs using multi-resolution recurrent skip connections. The proposed method is conceptually similar to dilated convolution but takes full advantage of the modeling power of RNNs. We show that \\textsc{DilatedRNN} is effective particularly in the problems where long-term dependencies are crucial. Finally, inspired by the structural similarity between CNNs and unrolled single layer RNNs, we also study the parameter sharing across different network layers that may be sequentially connected. Specifically, we focus on the family of feedforward CNNs that have an equivalent RNN form and tie the layer parameters with respect to the standard RNN unrolling rule. Empirically we have found that this family of models not only provides desirable balance between model complexity and performance, but also leads to some novel architecture that can be easily combined with domain knowledge.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-05-01","The student, Wei Han, accepted the attached license on 2019-04-18 at 13:42.","The student, Wei Han, submitted this Dissertation for approval on 2019-04-18 at 13:48.","This Dissertation was approved for publication on 2019-04-18 at 15:15."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Learning compact neural network representations with structural priors"]}]}],"canonical_facts":{"dc:contributor":["Huang, Thomas S.","Hasegawa-Johnson, Mark","Liang, Zhi-Pei","Hwu, Wen-Mei"],"dc:creator":["Han, Wei"],"dc:date":["2019-08-23T20:36:02Z","2021-08-24T09:15:16Z","2019-04-18","2019-05"],"dc:description":["DSpace SAF Submission Ingestion Package generated from Vireo submission #13747 on 2019-08-22 at 15:07:23","Made available in DSpace on 2019-08-23T20:36:02Z (GMT). No. of bitstreams: 2 HAN-DISSERTATION-2019.pdf: 1941888 bytes, checksum: 3e73bb3f2035c5446203dcfa3ad7bb3b (MD5) LICENSE.txt: 4204 bytes, checksum: 9bff5a3e92ba1fbbab9a73eae99a7816 (MD5) Previous issue date: 2019-04-18","Embargo set by: Seth Robbins for item 112176 Lift date: 2021-08-23T20:36:18Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 112176 on 2021-08-24T09:15:16Z.","The development of deep neural networks has taken two directions. On one hand, the networks become deeper and wider, employing drastically more model parameters and consuming more training data. On the other hand, the simplification of the internal structure of the neural networks has also contributed to the success of neural networks. Notably, two important families of neural networks, convolutional neural networks (CNN) and recurrent neural networks (RNN), both introduce certain structural priors and share model parameters internally to simplify the network. In this dissertation, we investigate a few alternative neural network structural priors to learn more compact CNN and RNN models and achieve better parameter-to-performance ratios. We have developed these neural network structures at three different abstraction levels with different target applications for each. First of all, motivated by the filter redundancy in convolutional neural networks, we have studied parameter sharing across filters in a convolutional layer. Instead of the conventional approach treating CNN filters as a set of independent model parameters, we explore the 2D spatial correlation among filters and propose sharing the filter parameters as if they are overlapping slices from a shared 3D tensor. Experiments show the proposed approach can effectively reduce the number of parameters in several state-of-the-art CNN architectures while still maintaining competitive performance. The second problem studied in this dissertation concerns the inter-layer connectivity pattern in RNNs, the family of neural network models specifically designed for time-series data. One longstanding challenge in RNNs is the vanishing gradient problem that hinders the model's capability to learn long-term dependency, the temporal data dependency across many time steps. Motivated by the recent development from two largely unrelated fields---general application of skip-connection to neural networks, and dilated convolution in image and audio related problems---we propose \\textsc{DilatedRNN}, a simple but principled way to construct multi-layer RNNs using multi-resolution recurrent skip connections. The proposed method is conceptually similar to dilated convolution but takes full advantage of the modeling power of RNNs. We show that \\textsc{DilatedRNN} is effective particularly in the problems where long-term dependencies are crucial. Finally, inspired by the structural similarity between CNNs and unrolled single layer RNNs, we also study the parameter sharing across different network layers that may be sequentially connected. Specifically, we focus on the family of feedforward CNNs that have an equivalent RNN form and tie the layer parameters with respect to the standard RNN unrolling rule. Empirically we have found that this family of models not only provides desirable balance between model complexity and performance, but also leads to some novel architecture that can be easily combined with domain knowledge.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2021-05-01","The student, Wei Han, accepted the attached license on 2019-04-18 at 13:42.","The student, Wei Han, submitted this Dissertation for approval on 2019-04-18 at 13:48.","This Dissertation was approved for publication on 2019-04-18 at 15:15."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105057"],"dc:language":["en"],"dc:rights":["Copyright 2019 Wei Han"],"dc:subject":["neural networks","recurrent neural networks","convolutional neural networks"],"dc:title":["Learning compact neural network representations with structural priors"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:44Z"}