{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/115475"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/115475","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"The implicit bias of gradient descent: from linear classifiers to deep networks","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_has_math":false,"creators":["Ji, Ziwei"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Telgarsky, Matus","Chekuri, Chandra","Raginsky, Maxim","Hsu, Daniel"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-05","date_published":"2022-05","updated_at":"2026-07-22T22:24:54Z","subjects":["Implicit bias","test error bounds","gradient descent","stochastic gradient descent","primal-dual analysis"],"languages":["en","eng"],"rights":["Copyright 2022 Ziwei Ji"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/115475","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Telgarsky, Matus","Chekuri, Chandra","Raginsky, Maxim","Hsu, Daniel"]},{"key":"dc:creator","label":"Author","values":["Ji, Ziwei"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-05","2022-04-18"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Implicit bias","test error bounds","gradient descent","stochastic gradient descent","primal-dual analysis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Ziwei Ji"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/115475"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Ziwei Ji, accepted the attached license on 2022-04-16 at 19:26.","The student, Ziwei Ji, submitted this Dissertation for approval on 2022-04-16 at 19:41.","This Dissertation was approved for publication on 2022-04-18 at 13:31.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17727 on 2022-11-11 at 13:15:30","Gradient descent and other first-order methods have been extensively used in deep learning. They can find solutions with not only low training error, but also low test error. This motivates the study of the implicit bias of training algorithms such as gradient descent, meaning we want to understand what special properties are satisfied by gradient descent solutions that lead to good generalization. In this thesis, we focus on gradient descent and its derivatives, show test error bounds and characterize the implicit biases. In detail, for linear classifiers, we consider both the separable setting and nonseparable setting. In the separable case, we first give an 1/t test error bound for stochastic gradient descent. Then for gradient descent, we present a primal-dual framework to analyze its implicit bias, which leads to fast margin maximization algorithms. While the previous results mostly require exponentially-tailed losses, we also show that for general decreasing losses, the implicit bias can still be characterized in terms of the regularization path. In the nonseparable case, we design a nearly-optimal algorithm by combining logistic regression and perceptron. We also characterize the implicit bias of gradient descent via a unique decomposition of the training set. For neural networks, we first provide test error bounds for shallow ReLU networks, using the recent idea of neural tangent kernel, but only requiring a mild overparameterization. Then we show implicit bias results for deep linear networks and deep homogeneous networks, in the form of alignment properties."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["The implicit bias of gradient descent: from linear classifiers to deep networks"]}]}],"canonical_facts":{"dc:contributor":["Telgarsky, Matus","Chekuri, Chandra","Raginsky, Maxim","Hsu, Daniel"],"dc:creator":["Ji, Ziwei"],"dc:date":["2022-05","2022-04-18"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Ziwei Ji, accepted the attached license on 2022-04-16 at 19:26.","The student, Ziwei Ji, submitted this Dissertation for approval on 2022-04-16 at 19:41.","This Dissertation was approved for publication on 2022-04-18 at 13:31.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17727 on 2022-11-11 at 13:15:30","Gradient descent and other first-order methods have been extensively used in deep learning. They can find solutions with not only low training error, but also low test error. This motivates the study of the implicit bias of training algorithms such as gradient descent, meaning we want to understand what special properties are satisfied by gradient descent solutions that lead to good generalization. In this thesis, we focus on gradient descent and its derivatives, show test error bounds and characterize the implicit biases. In detail, for linear classifiers, we consider both the separable setting and nonseparable setting. In the separable case, we first give an 1/t test error bound for stochastic gradient descent. Then for gradient descent, we present a primal-dual framework to analyze its implicit bias, which leads to fast margin maximization algorithms. While the previous results mostly require exponentially-tailed losses, we also show that for general decreasing losses, the implicit bias can still be characterized in terms of the regularization path. In the nonseparable case, we design a nearly-optimal algorithm by combining logistic regression and perceptron. We also characterize the implicit bias of gradient descent via a unique decomposition of the training set. For neural networks, we first provide test error bounds for shallow ReLU networks, using the recent idea of neural tangent kernel, but only requiring a mild overparameterization. Then we show implicit bias results for deep linear networks and deep homogeneous networks, in the form of alignment properties."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/115475"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Ziwei Ji"],"dc:subject":["Implicit bias","test error bounds","gradient descent","stochastic gradient descent","primal-dual analysis"],"dc:title":["The implicit bias of gradient descent: from linear classifiers to deep networks"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:54Z"}