{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113008"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113008","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Dynamical systems perspectives in machine learning","abstract":"We look at two facets of machine learning from a perspective of dynamical systems, that is, the data generated from a dynamical system and the iterative inference algorithm posed as a dynamical system. In the former, we look at time series data which is generated from a mixture of processes. Each process exists for a fixed duration and generates i.i.d categorical data points during that duration. More than one process can coexist at a particular time. The goal is to find the number of such hidden processes and the characteristic categorical distribution of each. This model is motivated by the problem of finding error events in error-logs from a mobile communication network. In the second direction, we consider the problem of regression using a shallow overparameterized neural network. Broadly, we look at training the neural network with the gradient descent algorithm on the squared loss function and discuss the generalization properties of the output of the gradient descent algorithm on an unseen data point. We look at two problems in this setting. First, we discuss the effect of l2 regularization on the squared loss and discuss how different strength of regularization provides a trade-off on the generalization of the neural network. Second, we look at squared loss without regularization and discuss the generalization properties when the true function we are trying to learn belongs to the class of polynomials in the presence of noisy samples. In both the problems, we consider the gradient descent algorithm as a dynamical system and use tools from control theory to analyze this dynamical system.","abstract_html":"We look at two facets of machine learning from a perspective of dynamical systems, that is, the data generated from a dynamical system and the iterative inference algorithm posed as a dynamical system. In the former, we look at time series data which is generated from a mixture of processes. Each process exists for a fixed duration and generates i.i.d categorical data points during that duration. More than one process can coexist at a particular time. The goal is to find the number of such hidden processes and the characteristic categorical distribution of each. This model is motivated by the problem of finding error events in error-logs from a mobile communication network. In the second direction, we consider the problem of regression using a shallow overparameterized neural network. Broadly, we look at training the neural network with the gradient descent algorithm on the squared loss function and discuss the generalization properties of the output of the gradient descent algorithm on an unseen data point. We look at two problems in this setting. First, we discuss the effect of l2 regularization on the squared loss and discuss how different strength of regularization provides a trade-off on the generalization of the neural network. Second, we look at squared loss without regularization and discuss the generalization properties when the true function we are trying to learn belongs to the class of polynomials in the presence of noisy samples. In both the problems, we consider the gradient descent algorithm as a dynamical system and use tools from control theory to analyze this dynamical system.","abstract_has_math":false,"creators":["Satpathi, Siddhartha"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Srikant, Rayadurgam","Beck, Carolyn L","Chatterjee, Sabyasachi","Hu, Bin"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T21:45:32Z","date_published":"2022-01-12T21:45:32Z","updated_at":"2026-07-22T22:24:52Z","subjects":["Neural Network","error-log","LDA","gradient descent"],"languages":["en"],"rights":["Copyright 2021 Siddhartha Satpathi"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113008","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Srikant, Rayadurgam","Beck, Carolyn L","Chatterjee, Sabyasachi","Hu, Bin"]},{"key":"dc:creator","label":"Author","values":["Satpathi, Siddhartha"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T21:45:32Z","2021-07-12","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Neural Network","error-log","LDA","gradient descent"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Siddhartha Satpathi"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113008"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We look at two facets of machine learning from a perspective of dynamical systems, that is, the data generated from a dynamical system and the iterative inference algorithm posed as a dynamical system. In the former, we look at time series data which is generated from a mixture of processes. Each process exists for a fixed duration and generates i.i.d categorical data points during that duration. More than one process can coexist at a particular time. The goal is to find the number of such hidden processes and the characteristic categorical distribution of each. This model is motivated by the problem of finding error events in error-logs from a mobile communication network. In the second direction, we consider the problem of regression using a shallow overparameterized neural network. Broadly, we look at training the neural network with the gradient descent algorithm on the squared loss function and discuss the generalization properties of the output of the gradient descent algorithm on an unseen data point. We look at two problems in this setting. First, we discuss the effect of l2 regularization on the squared loss and discuss how different strength of regularization provides a trade-off on the generalization of the neural network. Second, we look at squared loss without regularization and discuss the generalization properties when the true function we are trying to learn belongs to the class of polynomials in the presence of noisy samples. In both the problems, we consider the gradient descent algorithm as a dynamical system and use tools from control theory to analyze this dynamical system.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Siddhartha Satpathi, accepted the attached license on 2021-07-09 at 16:01.","The student, Siddhartha Satpathi, submitted this Dissertation for approval on 2021-07-09 at 16:13.","This Dissertation was approved for publication on 2021-07-12 at 09:49.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16837 on 2022-01-12 at 12:44:44","Made available in DSpace on 2022-01-12T21:45:32Z (GMT). No. of bitstreams: 3 SATPATHI-DISSERTATION-2021.pdf: 1339292 bytes, checksum: 9cac9a8439794b90ed598f3b5671e627 (MD5) dkwwqfxxvjypxhrzcxwbnbtpkdjjqmsd.zip: 1832706 bytes, checksum: 295ea900c32a4ca127dd9d3508638938 (MD5) LICENSE.txt: 4216 bytes, checksum: c88bdfcbb00b9e772c95a85144043d50 (MD5) Previous issue date: 2021-07-12"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Dynamical systems perspectives in machine learning"]}]}],"canonical_facts":{"dc:contributor":["Srikant, Rayadurgam","Beck, Carolyn L","Chatterjee, Sabyasachi","Hu, Bin"],"dc:creator":["Satpathi, Siddhartha"],"dc:date":["2022-01-12T21:45:32Z","2021-07-12","2021-08"],"dc:description":["We look at two facets of machine learning from a perspective of dynamical systems, that is, the data generated from a dynamical system and the iterative inference algorithm posed as a dynamical system. In the former, we look at time series data which is generated from a mixture of processes. Each process exists for a fixed duration and generates i.i.d categorical data points during that duration. More than one process can coexist at a particular time. The goal is to find the number of such hidden processes and the characteristic categorical distribution of each. This model is motivated by the problem of finding error events in error-logs from a mobile communication network. In the second direction, we consider the problem of regression using a shallow overparameterized neural network. Broadly, we look at training the neural network with the gradient descent algorithm on the squared loss function and discuss the generalization properties of the output of the gradient descent algorithm on an unseen data point. We look at two problems in this setting. First, we discuss the effect of l2 regularization on the squared loss and discuss how different strength of regularization provides a trade-off on the generalization of the neural network. Second, we look at squared loss without regularization and discuss the generalization properties when the true function we are trying to learn belongs to the class of polynomials in the presence of noisy samples. In both the problems, we consider the gradient descent algorithm as a dynamical system and use tools from control theory to analyze this dynamical system.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Siddhartha Satpathi, accepted the attached license on 2021-07-09 at 16:01.","The student, Siddhartha Satpathi, submitted this Dissertation for approval on 2021-07-09 at 16:13.","This Dissertation was approved for publication on 2021-07-12 at 09:49.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16837 on 2022-01-12 at 12:44:44","Made available in DSpace on 2022-01-12T21:45:32Z (GMT). No. of bitstreams: 3 SATPATHI-DISSERTATION-2021.pdf: 1339292 bytes, checksum: 9cac9a8439794b90ed598f3b5671e627 (MD5) dkwwqfxxvjypxhrzcxwbnbtpkdjjqmsd.zip: 1832706 bytes, checksum: 295ea900c32a4ca127dd9d3508638938 (MD5) LICENSE.txt: 4216 bytes, checksum: c88bdfcbb00b9e772c95a85144043d50 (MD5) Previous issue date: 2021-07-12"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113008"],"dc:language":["en"],"dc:rights":["Copyright 2021 Siddhartha Satpathi"],"dc:subject":["Neural Network","error-log","LDA","gradient descent"],"dc:title":["Dynamical systems perspectives in machine learning"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}