{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/46909"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/46909","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Value function approximation architectures for neuro-dynamic programming","abstract":"Neuro-dynamic programming is a class of powerful techniques for approximating the solution to dynamic programming equations. In their most computationally attractive formulations, these techniques provide the approximate solution only within a prescribed finite-dimensional function class. Thus, the question that always arises is how should the function class be chosen? In this dissertation, we first propose an approach using the solutions to associated fluid and diffusion approximations. In order to evaluate this approach, we establish bounds on the approximation errors. Next, we propose a novel parameterized Q-learning algorithm. Q-learning is a model-free method to compute the Q-function associated with an optimal policy, based on observations of states and actions. If the size of a state or a policy space is too large, Q-learning is often not very practical because there are too many Q-function values to update. One way to address this problem is to approximate the Q-function within a function class. However, such methods often require an explicit model of the system, such as the split sampling method introduced by Borkar. The proposed algorithm is a reinforcement learning (RL) method, in which case the system dynamics are not known. This method is designed based on using approximations of the transition kernel of the Markov decision process (MDP). Lastly, we apply the proposed results of value function approximation techniques to several applications. In the power management model, we focus on the processor speed control problem to balance the performance and energy usage. Then we extend the results to the load balancing and the power management problem of geographically distributed data centers with grid regulation. In the cross-layer wireless control problem, the network utility maximization (NUM) and adaptive modulation (AM) are combined to balance the network performance and transmission power. In these applications, we show how to model the real problems by using the MDP model with reasonable assumptions and necessary approximations. Approximations of the value function are obtained for specific models, and evaluated by getting bounds for the errors. These approximate solutions are then used to construct basis functions for learning algorithms in the simulations.","abstract_html":"Neuro-dynamic programming is a class of powerful techniques for approximating the solution to dynamic programming equations. In their most computationally attractive formulations, these techniques provide the approximate solution only within a prescribed finite-dimensional function class. Thus, the question that always arises is how should the function class be chosen? In this dissertation, we first propose an approach using the solutions to associated fluid and diffusion approximations. In order to evaluate this approach, we establish bounds on the approximation errors. Next, we propose a novel parameterized Q-learning algorithm. Q-learning is a model-free method to compute the Q-function associated with an optimal policy, based on observations of states and actions. If the size of a state or a policy space is too large, Q-learning is often not very practical because there are too many Q-function values to update. One way to address this problem is to approximate the Q-function within a function class. However, such methods often require an explicit model of the system, such as the split sampling method introduced by Borkar. The proposed algorithm is a reinforcement learning (RL) method, in which case the system dynamics are not known. This method is designed based on using approximations of the transition kernel of the Markov decision process (MDP). Lastly, we apply the proposed results of value function approximation techniques to several applications. In the power management model, we focus on the processor speed control problem to balance the performance and energy usage. Then we extend the results to the load balancing and the power management problem of geographically distributed data centers with grid regulation. In the cross-layer wireless control problem, the network utility maximization (NUM) and adaptive modulation (AM) are combined to balance the network performance and transmission power. In these applications, we show how to model the real problems by using the MDP model with reasonable assumptions and necessary approximations. Approximations of the value function are obtained for specific models, and evaluated by getting bounds for the errors. These approximate solutions are then used to construct basis functions for learning algorithms in the simulations.","abstract_has_math":false,"creators":["Chen, Wei"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Meyn, Sean P.","Hajek, Bruce","Hutchinson, Seth A.","Nedich, Angelia"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-01-16T18:26:02Z","date_published":"2014-01-16T18:26:02Z","updated_at":"2026-07-22T22:25:38Z","subjects":["Neuro-Dynamic Programming","Parametric Q-learning","Value Function Approximation","Processor Power Management","Data Center Power Management","Cross-Layer Wireless Control"],"languages":["en"],"rights":["Copyright 2013 Wei Chen"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/46909","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Meyn, Sean P.","Hajek, Bruce","Hutchinson, Seth A.","Nedich, Angelia"]},{"key":"dc:creator","label":"Author","values":["Chen, Wei"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014-01-16T18:26:02Z","2016-01-16T11:02:04Z","2013-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Neuro-Dynamic Programming","Parametric Q-learning","Value Function Approximation","Processor Power Management","Data Center Power Management","Cross-Layer Wireless Control"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2013 Wei Chen"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/46909"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Neuro-dynamic programming is a class of powerful techniques for approximating the solution to dynamic programming equations. In their most computationally attractive formulations, these techniques provide the approximate solution only within a prescribed finite-dimensional function class. Thus, the question that always arises is how should the function class be chosen? In this dissertation, we first propose an approach using the solutions to associated fluid and diffusion approximations. In order to evaluate this approach, we establish bounds on the approximation errors. Next, we propose a novel parameterized Q-learning algorithm. Q-learning is a model-free method to compute the Q-function associated with an optimal policy, based on observations of states and actions. If the size of a state or a policy space is too large, Q-learning is often not very practical because there are too many Q-function values to update. One way to address this problem is to approximate the Q-function within a function class. However, such methods often require an explicit model of the system, such as the split sampling method introduced by Borkar. The proposed algorithm is a reinforcement learning (RL) method, in which case the system dynamics are not known. This method is designed based on using approximations of the transition kernel of the Markov decision process (MDP). Lastly, we apply the proposed results of value function approximation techniques to several applications. In the power management model, we focus on the processor speed control problem to balance the performance and energy usage. Then we extend the results to the load balancing and the power management problem of geographically distributed data centers with grid regulation. In the cross-layer wireless control problem, the network utility maximization (NUM) and adaptive modulation (AM) are combined to balance the network performance and transmission power. In these applications, we show how to model the real problems by using the MDP model with reasonable assumptions and necessary approximations. Approximations of the value function are obtained for specific models, and evaluated by getting bounds for the errors. These approximate solutions are then used to construct basis functions for learning algorithms in the simulations.","Item withdrawn by Laura Spradlin (lspradl2@illinois.edu) on 2013-12-02T14:33:51Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Chen_Wei.pdf: 7488417 bytes, checksum: 8821f75e096c84203444be0e87ffaed1 (MD5)","Made available in DSpace on 2014-01-16T18:26:02Z (GMT). No. of bitstreams: 2 Wei_Chen.pdf: 7488417 bytes, checksum: 8821f75e096c84203444be0e87ffaed1 (MD5) license.txt: 4058 bytes, checksum: a1abb7c4bbdb0836afe5bb4e3873ae50 (MD5)","Item marked as restricted to the 'Administrator' Group (id=1) by Seth Robbins (robbins.sd@gmail.com) on 2014-01-16T18:27:35Z Item is restricted until 2016-01-16T18:27:27Z","Restriction data tranferred 2014-07-01T11:36:47-05:00 Original Data Group with Access Administrator Release Date: 2016-01-16 12:27:27 UTC Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 46928 on 2016-01-16T11:02:04Z."]},{"key":"dc:title","label":"Title","values":["Value function approximation architectures for neuro-dynamic programming"]}]}],"canonical_facts":{"dc:contributor":["Meyn, Sean P.","Hajek, Bruce","Hutchinson, Seth A.","Nedich, Angelia"],"dc:creator":["Chen, Wei"],"dc:date":["2014-01-16T18:26:02Z","2016-01-16T11:02:04Z","2013-12"],"dc:description":["Neuro-dynamic programming is a class of powerful techniques for approximating the solution to dynamic programming equations. In their most computationally attractive formulations, these techniques provide the approximate solution only within a prescribed finite-dimensional function class. Thus, the question that always arises is how should the function class be chosen? In this dissertation, we first propose an approach using the solutions to associated fluid and diffusion approximations. In order to evaluate this approach, we establish bounds on the approximation errors. Next, we propose a novel parameterized Q-learning algorithm. Q-learning is a model-free method to compute the Q-function associated with an optimal policy, based on observations of states and actions. If the size of a state or a policy space is too large, Q-learning is often not very practical because there are too many Q-function values to update. One way to address this problem is to approximate the Q-function within a function class. However, such methods often require an explicit model of the system, such as the split sampling method introduced by Borkar. The proposed algorithm is a reinforcement learning (RL) method, in which case the system dynamics are not known. This method is designed based on using approximations of the transition kernel of the Markov decision process (MDP). Lastly, we apply the proposed results of value function approximation techniques to several applications. In the power management model, we focus on the processor speed control problem to balance the performance and energy usage. Then we extend the results to the load balancing and the power management problem of geographically distributed data centers with grid regulation. In the cross-layer wireless control problem, the network utility maximization (NUM) and adaptive modulation (AM) are combined to balance the network performance and transmission power. In these applications, we show how to model the real problems by using the MDP model with reasonable assumptions and necessary approximations. Approximations of the value function are obtained for specific models, and evaluated by getting bounds for the errors. These approximate solutions are then used to construct basis functions for learning algorithms in the simulations.","Item withdrawn by Laura Spradlin (lspradl2@illinois.edu) on 2013-12-02T14:33:51Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Chen_Wei.pdf: 7488417 bytes, checksum: 8821f75e096c84203444be0e87ffaed1 (MD5)","Made available in DSpace on 2014-01-16T18:26:02Z (GMT). No. of bitstreams: 2 Wei_Chen.pdf: 7488417 bytes, checksum: 8821f75e096c84203444be0e87ffaed1 (MD5) license.txt: 4058 bytes, checksum: a1abb7c4bbdb0836afe5bb4e3873ae50 (MD5)","Item marked as restricted to the 'Administrator' Group (id=1) by Seth Robbins (robbins.sd@gmail.com) on 2014-01-16T18:27:35Z Item is restricted until 2016-01-16T18:27:27Z","Restriction data tranferred 2014-07-01T11:36:47-05:00 Original Data Group with Access Administrator Release Date: 2016-01-16 12:27:27 UTC Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 46928 on 2016-01-16T11:02:04Z."],"dc:identifier":["http://hdl.handle.net/2142/46909"],"dc:language":["en"],"dc:rights":["Copyright 2013 Wei Chen"],"dc:subject":["Neuro-Dynamic Programming","Parametric Q-learning","Value Function Approximation","Processor Power Management","Data Center Power Management","Cross-Layer Wireless Control"],"dc:title":["Value function approximation architectures for neuro-dynamic programming"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:38Z"}