{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124230"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124230","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Understanding the mechanism of pretraining stabilization heuristics: A variance-oriented perspective","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_has_math":false,"creators":["Liu, Liyuan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei","Ji, Heng","Zhai, ChengXiang","Gao, Jianfeng","Peters, Matthew E"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:00Z","subjects":["Training Stability","Variance","Pretraining"],"languages":["en","eng"],"rights":["Copyright 2024 Liyuan Liu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124230","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei","Ji, Heng","Zhai, ChengXiang","Gao, Jianfeng","Peters, Matthew E"]},{"key":"dc:creator","label":"Author","values":["Liu, Liyuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-04-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Training Stability","Variance","Pretraining"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Liyuan Liu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124230"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, Liyuan Liu, accepted the attached license on 2024-04-04 at 21:10.","The student, Liyuan Liu, submitted this Dissertation for approval on 2024-04-04 at 21:37.","This Dissertation was approved for publication on 2024-04-08 at 16:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20302 on 2024-09-16 at 00:33:41","Language model pretraining has been breaking the glass ceiling for various natural language processing tasks and has been viewed as one of the most significant successes of deep learning, continuously challenging our understanding of learning and cognition. Recently models, including GPT-4 and BART, fueled by an unprecedented scale of computing and data, exhibit unprecedented intelligence, that some even refer to as \"sparks of artificial general intelligence\". The success of large-scale pretraining hinges on intricate engineering heuristics. While the empirical benefits of these heuristics are evident, their underlying mechanisms remain elusive. This dissertation endeavors to demystify the mathematical principles underlying these pretraining heuristics, aiming to illuminate their mechanisms and potentially guide future algorithm developments. Adopting a variance-oriented perspective, my research rigorously inspects the heuristics that are pivotal to the stability of current pretraining practices, emphasizing learning rate warmup, model initialization, and gradient approximation. In this dissertation, I show that these pretraining stabilization heuristics can be coherently elucidated with a unified framework anchored in variance, a classical metric for stability. First, I analyze the variance of adaptive learning rate and model outputs, revealing that both learning rate warmup and model initialization function as variance modulators. Then, I move to explore the variance-bias tradeoff in the discrete variable gradient approximation, i.e., employing a numerical ODE framework, I unveil the underlying dynamics of the approximation bias, achieving second order precision with minimal computational overhead. Besides theoretical results, empirical verifications are conducted to verify the assumptions and applicability of the recognized principles. Building upon these insights, this dissertation introduces novel techniques designed to advance the pretraining practices, including RAdam for learning rate warmup, Admin for Transformer model initialization, ReinMax and SparseMixer for gradient approximation. Under the guidance of the recognized principles, all proposed methods require minimal trial-and-error configurations, thereby emerging as robust and high-perform tools for pretraining practices for adaptations."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Understanding the mechanism of pretraining stabilization heuristics: A variance-oriented perspective"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei","Ji, Heng","Zhai, ChengXiang","Gao, Jianfeng","Peters, Matthew E"],"dc:creator":["Liu, Liyuan"],"dc:date":["2024-05","2024-04-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, Liyuan Liu, accepted the attached license on 2024-04-04 at 21:10.","The student, Liyuan Liu, submitted this Dissertation for approval on 2024-04-04 at 21:37.","This Dissertation was approved for publication on 2024-04-08 at 16:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20302 on 2024-09-16 at 00:33:41","Language model pretraining has been breaking the glass ceiling for various natural language processing tasks and has been viewed as one of the most significant successes of deep learning, continuously challenging our understanding of learning and cognition. Recently models, including GPT-4 and BART, fueled by an unprecedented scale of computing and data, exhibit unprecedented intelligence, that some even refer to as \"sparks of artificial general intelligence\". The success of large-scale pretraining hinges on intricate engineering heuristics. While the empirical benefits of these heuristics are evident, their underlying mechanisms remain elusive. This dissertation endeavors to demystify the mathematical principles underlying these pretraining heuristics, aiming to illuminate their mechanisms and potentially guide future algorithm developments. Adopting a variance-oriented perspective, my research rigorously inspects the heuristics that are pivotal to the stability of current pretraining practices, emphasizing learning rate warmup, model initialization, and gradient approximation. In this dissertation, I show that these pretraining stabilization heuristics can be coherently elucidated with a unified framework anchored in variance, a classical metric for stability. First, I analyze the variance of adaptive learning rate and model outputs, revealing that both learning rate warmup and model initialization function as variance modulators. Then, I move to explore the variance-bias tradeoff in the discrete variable gradient approximation, i.e., employing a numerical ODE framework, I unveil the underlying dynamics of the approximation bias, achieving second order precision with minimal computational overhead. Besides theoretical results, empirical verifications are conducted to verify the assumptions and applicability of the recognized principles. Building upon these insights, this dissertation introduces novel techniques designed to advance the pretraining practices, including RAdam for learning rate warmup, Admin for Transformer model initialization, ReinMax and SparseMixer for gradient approximation. Under the guidance of the recognized principles, all proposed methods require minimal trial-and-error configurations, thereby emerging as robust and high-perform tools for pretraining practices for adaptations."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124230"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Liyuan Liu"],"dc:subject":["Training Stability","Variance","Pretraining"],"dc:title":["Understanding the mechanism of pretraining stabilization heuristics: A variance-oriented perspective"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:00Z"}