{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/104821"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/104821","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Trainability and generalization of small-scale neural networks","abstract":"As deep learning has become solution for various machine learning, artificial intelligence applications, their architectures have been developed accordingly. Modern deep learning applications often use overparameterized setting, which is opposite to what conventional learning theory suggests. While deep neural networks are considered to be less vulnerable to overfitting even with their overparameterized architecture, this project observed that properly trained small-scale networks indeed outperform its larger counterparts. The generalization ability of small-scale networks has been overlooked in many researches and practice, due to their extremely slow convergence speed. This project observed that imbalanced layer-wise gradient norm can hider overall convergence speed of neural networks, and narrow networks are vulnerable to this. This projects investigates possible reasons of convergence failure of small-scale neural networks, and suggests a strategy to alleviate the problem.","abstract_html":"As deep learning has become solution for various machine learning, artificial intelligence applications, their architectures have been developed accordingly. Modern deep learning applications often use overparameterized setting, which is opposite to what conventional learning theory suggests. While deep neural networks are considered to be less vulnerable to overfitting even with their overparameterized architecture, this project observed that properly trained small-scale networks indeed outperform its larger counterparts. The generalization ability of small-scale networks has been overlooked in many researches and practice, due to their extremely slow convergence speed. This project observed that imbalanced layer-wise gradient norm can hider overall convergence speed of neural networks, and narrow networks are vulnerable to this. This projects investigates possible reasons of convergence failure of small-scale neural networks, and suggests a strategy to alleviate the problem.","abstract_has_math":false,"creators":["Song, Myung Hwan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Industrial Engineering","degree_department":null,"school":null,"contributors":["Sun, Ruoyu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T19:51:51Z","date_published":"2019-08-23T19:51:51Z","updated_at":"2026-07-22T22:24:42Z","subjects":["Deep Learning","Neural Networks","Learning Theory"],"languages":["en"],"rights":["Copyright 2019 Myung Hwan Song"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/104821","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Sun, Ruoyu"]},{"key":"dc:creator","label":"Author","values":["Song, Myung Hwan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T19:51:51Z","2019-04-22","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Industrial Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Deep Learning","Neural Networks","Learning Theory"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Myung Hwan Song"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/104821"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["As deep learning has become solution for various machine learning, artificial intelligence applications, their architectures have been developed accordingly. Modern deep learning applications often use overparameterized setting, which is opposite to what conventional learning theory suggests. While deep neural networks are considered to be less vulnerable to overfitting even with their overparameterized architecture, this project observed that properly trained small-scale networks indeed outperform its larger counterparts. The generalization ability of small-scale networks has been overlooked in many researches and practice, due to their extremely slow convergence speed. This project observed that imbalanced layer-wise gradient norm can hider overall convergence speed of neural networks, and narrow networks are vulnerable to this. This projects investigates possible reasons of convergence failure of small-scale neural networks, and suggests a strategy to alleviate the problem.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Myung Hwan Song, accepted the attached license on 2019-04-18 at 14:20.","The student, Myung Hwan Song, submitted this Thesis for approval on 2019-04-18 at 14:21.","This Thesis was approved for publication on 2019-04-22 at 10:37.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13601 on 2019-08-22 at 14:43:15","Made available in DSpace on 2019-08-23T19:51:51Z (GMT). No. of bitstreams: 2 SONG-THESIS-2019.pdf: 953882 bytes, checksum: f2ee5838b1f004ad3fde112879301910 (MD5) LICENSE.txt: 4212 bytes, checksum: 466e262f59dab51ded5d7752d78d39e5 (MD5) Previous issue date: 2019-04-22"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Trainability and generalization of small-scale neural networks"]}]}],"canonical_facts":{"dc:contributor":["Sun, Ruoyu"],"dc:creator":["Song, Myung Hwan"],"dc:date":["2019-08-23T19:51:51Z","2019-04-22","2019-05"],"dc:description":["As deep learning has become solution for various machine learning, artificial intelligence applications, their architectures have been developed accordingly. Modern deep learning applications often use overparameterized setting, which is opposite to what conventional learning theory suggests. While deep neural networks are considered to be less vulnerable to overfitting even with their overparameterized architecture, this project observed that properly trained small-scale networks indeed outperform its larger counterparts. The generalization ability of small-scale networks has been overlooked in many researches and practice, due to their extremely slow convergence speed. This project observed that imbalanced layer-wise gradient norm can hider overall convergence speed of neural networks, and narrow networks are vulnerable to this. This projects investigates possible reasons of convergence failure of small-scale neural networks, and suggests a strategy to alleviate the problem.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Myung Hwan Song, accepted the attached license on 2019-04-18 at 14:20.","The student, Myung Hwan Song, submitted this Thesis for approval on 2019-04-18 at 14:21.","This Thesis was approved for publication on 2019-04-22 at 10:37.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13601 on 2019-08-22 at 14:43:15","Made available in DSpace on 2019-08-23T19:51:51Z (GMT). No. of bitstreams: 2 SONG-THESIS-2019.pdf: 953882 bytes, checksum: f2ee5838b1f004ad3fde112879301910 (MD5) LICENSE.txt: 4212 bytes, checksum: 466e262f59dab51ded5d7752d78d39e5 (MD5) Previous issue date: 2019-04-22"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/104821"],"dc:language":["en"],"dc:rights":["Copyright 2019 Myung Hwan Song"],"dc:subject":["Deep Learning","Neural Networks","Learning Theory"],"dc:title":["Trainability and generalization of small-scale neural networks"],"dc:type":["text"],"thesis:degree_discipline":["Industrial Engineering"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:42Z"}