{"id":{"repo_id":"ttu","oai_identifier":"oai:ttu-ir.tdl.org:2346/20357"},"canonical_url":"https://search.dev.ndltd.org/etd/ttu/oai:ttu-ir.tdl.org:2346/20357","repository":{"repo_id":"ttu","name":"Texas Technology University","base_url":"https://ttu-ir.tdl.org/server/oai/request"},"display":{"title":"Continuous state Q-learning","abstract":"Q-learning is a solution technique developed to solve classical Markov Decision Processes, MDPs. Markov Decision Processes are models for sequential decision making problems and address many classical control problems. In Chapter I, this paper discusses the model and some standard solution techniques used in Markov Decision Processes and its limitations [6]. Q-learning was developed by Watkins to broaden the scope of problems that dynamic programming, MDP techniques, can solve. Classical Q-learning is a model free solution technique and is therefore able to address a variety of poorly modeled decision problems which were unsolvable using standard MDP techniques. Watkins development of Q-learning is based on Markov Decision Processes with discrete action and state spaces. The model and algorithm associated with classical Q-learning are described in Chapter II. To extend the set of problems which can be addressed using Q-learning, Chapter III addresses solution techniques for poorly modeled problems with continuous state and/or action spaces. The model is slightly altered and the algorithm is adjusted to account for the continuous state and action spaces. Numerical example show that continuous Q-learning does determine the optimal policy over time. Ongoing research is being carried on to improve both the current classical Q-learning method and to prove the convergence in the continuous case.","abstract_html":"Q-learning is a solution technique developed to solve classical Markov Decision Processes, MDPs. Markov Decision Processes are models for sequential decision making problems and address many classical control problems. In Chapter I, this paper discusses the model and some standard solution techniques used in Markov Decision Processes and its limitations [6]. Q-learning was developed by Watkins to broaden the scope of problems that dynamic programming, MDP techniques, can solve. Classical Q-learning is a model free solution technique and is therefore able to address a variety of poorly modeled decision problems which were unsolvable using standard MDP techniques. Watkins development of Q-learning is based on Markov Decision Processes with discrete action and state spaces. The model and algorithm associated with classical Q-learning are described in Chapter II. To extend the set of problems which can be addressed using Q-learning, Chapter III addresses solution techniques for poorly modeled problems with continuous state and/or action spaces. The model is slightly altered and the algorithm is adjusted to account for the continuous state and action spaces. Numerical example show that continuous Q-learning does determine the optimal policy over time. Ongoing research is being carried on to improve both the current classical Q-learning method and to prove the convergence in the continuous case.","abstract_has_math":false,"creators":["Alcorn, Cristy Michele"],"institution":"Texas Tech University","degree_name":"M.S.","degree_level":"Masters","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":1999,"date_issued":"1999-05","date_published":"1999-05","updated_at":"2026-07-24T05:04:47Z","subjects":["Learning models","Markov processes","Dynamic programming"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2346/20357","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Alcorn, Cristy Michele"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2011-02-18T23:39:48Z"]},{"key":"dc:date.issued","label":"Date","values":["1999-05"]},{"key":"dc:publisher","label":"Institution","values":["Texas Tech University"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Texas Tech University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Learning models","Markov processes","Dynamic programming"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/2346/20357"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Q-learning is a solution technique developed to solve classical Markov Decision Processes, MDPs. Markov Decision Processes are models for sequential decision making problems and address many classical control problems. In Chapter I, this paper discusses the model and some standard solution techniques used in Markov Decision Processes and its limitations [6]. Q-learning was developed by Watkins to broaden the scope of problems that dynamic programming, MDP techniques, can solve. Classical Q-learning is a model free solution technique and is therefore able to address a variety of poorly modeled decision problems which were unsolvable using standard MDP techniques. Watkins development of Q-learning is based on Markov Decision Processes with discrete action and state spaces. The model and algorithm associated with classical Q-learning are described in Chapter II. To extend the set of problems which can be addressed using Q-learning, Chapter III addresses solution techniques for poorly modeled problems with continuous state and/or action spaces. The model is slightly altered and the algorithm is adjusted to account for the continuous state and action spaces. Numerical example show that continuous Q-learning does determine the optimal policy over time. Ongoing research is being carried on to improve both the current classical Q-learning method and to prove the convergence in the continuous case."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Continuous state Q-learning"]}]}],"canonical_facts":{"dc:creator":["Alcorn, Cristy Michele"],"dc:date.available":["2011-02-18T23:39:48Z"],"dc:date.issued":["1999-05"],"dc:description.abstract":["Q-learning is a solution technique developed to solve classical Markov Decision Processes, MDPs. Markov Decision Processes are models for sequential decision making problems and address many classical control problems. In Chapter I, this paper discusses the model and some standard solution techniques used in Markov Decision Processes and its limitations [6]. Q-learning was developed by Watkins to broaden the scope of problems that dynamic programming, MDP techniques, can solve. Classical Q-learning is a model free solution technique and is therefore able to address a variety of poorly modeled decision problems which were unsolvable using standard MDP techniques. Watkins development of Q-learning is based on Markov Decision Processes with discrete action and state spaces. The model and algorithm associated with classical Q-learning are described in Chapter II. To extend the set of problems which can be addressed using Q-learning, Chapter III addresses solution techniques for poorly modeled problems with continuous state and/or action spaces. The model is slightly altered and the algorithm is adjusted to account for the continuous state and action spaces. Numerical example show that continuous Q-learning does determine the optimal policy over time. Ongoing research is being carried on to improve both the current classical Q-learning method and to prove the convergence in the continuous case."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["http://hdl.handle.net/2346/20357"],"dc:language.iso":["eng"],"dc:publisher":["Texas Tech University"],"dc:subject":["Learning models","Markov processes","Dynamic programming"],"dc:title":["Continuous state Q-learning"],"dc:type":["Thesis"],"thesis:degree_level":["Masters"],"thesis:degree_name":["M.S."],"thesis:institution_name":["Texas Tech University"]},"updated_at":"2026-07-24T05:04:47Z"}