{"id":{"repo_id":"uoit","oai_identifier":"oai:ontariotechu.scholaris.ca:10155/1108"},"canonical_url":"https://search.dev.ndltd.org/etd/uoit/oai:ontariotechu.scholaris.ca:10155/1108","repository":{"repo_id":"uoit","name":"Ontario Institute of Technology","base_url":"https://ontariotechu.scholaris.ca/server/oai/request"},"display":{"title":"Perpetually playing physics","abstract":"Here we discuss ideas of reinforcement learning and the importance of various aspects of it. We show how reinforcement learning methods based on genetic algorithms can be used to reproduce thermodynamic cycles without prior knowledge of physics. To show this, we introduce an environment that models a simple heat engine. With this, we are able to optimize a neural network based policy to maximize the thermal efficiency for different cases. Using a series of restricted action sets in this environment, our policy was able to reproduce three known thermodynamic cycles. We also introduce an irreversible action, creating an unknown thermodynamic cycle that the agent helps discover, showing how reinforcement learning can find solutions to new problems. We also discuss shortcomings of the method used, the importance of understanding the class of problem being handled, and why some methods can only be used for certain classes of problems.","abstract_html":"Here we discuss ideas of reinforcement learning and the importance of various aspects of it. We show how reinforcement learning methods based on genetic algorithms can be used to reproduce thermodynamic cycles without prior knowledge of physics. To show this, we introduce an environment that models a simple heat engine. With this, we are able to optimize a neural network based policy to maximize the thermal efficiency for different cases. Using a series of restricted action sets in this environment, our policy was able to reproduce three known thermodynamic cycles. We also introduce an irreversible action, creating an unknown thermodynamic cycle that the agent helps discover, showing how reinforcement learning can find solutions to new problems. We also discuss shortcomings of the method used, the importance of understanding the class of problem being handled, and why some methods can only be used for certain classes of problems.","abstract_has_math":false,"creators":["Beeler, Chris"],"institution":"University of Ontario Institute of Technology","degree_name":"Master of Science (MSc)","degree_level":null,"degree_discipline":"Modelling and Computational Science","degree_department":null,"school":null,"contributors":[],"advisors":["van Veen, Lennaert","Tamblyn, Isaac"],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-01","date_published":"2019-08-01","updated_at":"2026-07-24T05:35:20Z","subjects":["Reinforcement learning","Machine learning","Mathematics","Physics","Chemistry"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10155/1108","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["van Veen, Lennaert","Tamblyn, Isaac"]},{"key":"dc:creator","label":"Author","values":["Beeler, Chris"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2019-10-28T19:33:30Z","2022-03-29T17:27:07Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2019-10-28T19:33:30Z","2022-03-29T17:27:07Z"]},{"key":"dc:date.issued","label":"Date","values":["2019-08-01"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Modelling and Computational Science"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MSc)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Ontario Institute of Technology"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Reinforcement learning","Machine learning","Mathematics","Physics","Chemistry"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10155/1108"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Here we discuss ideas of reinforcement learning and the importance of various aspects of it. We show how reinforcement learning methods based on genetic algorithms can be used to reproduce thermodynamic cycles without prior knowledge of physics. To show this, we introduce an environment that models a simple heat engine. With this, we are able to optimize a neural network based policy to maximize the thermal efficiency for different cases. Using a series of restricted action sets in this environment, our policy was able to reproduce three known thermodynamic cycles. We also introduce an irreversible action, creating an unknown thermodynamic cycle that the agent helps discover, showing how reinforcement learning can find solutions to new problems. We also discuss shortcomings of the method used, the importance of understanding the class of problem being handled, and why some methods can only be used for certain classes of problems."]},{"key":"dc:title","label":"Title","values":["Perpetually playing physics"]}]}],"canonical_facts":{"dc:contributor.advisor":["van Veen, Lennaert","Tamblyn, Isaac"],"dc:creator":["Beeler, Chris"],"dc:date.accessioned":["2019-10-28T19:33:30Z","2022-03-29T17:27:07Z"],"dc:date.available":["2019-10-28T19:33:30Z","2022-03-29T17:27:07Z"],"dc:date.issued":["2019-08-01"],"dc:description.abstract":["Here we discuss ideas of reinforcement learning and the importance of various aspects of it. We show how reinforcement learning methods based on genetic algorithms can be used to reproduce thermodynamic cycles without prior knowledge of physics. To show this, we introduce an environment that models a simple heat engine. With this, we are able to optimize a neural network based policy to maximize the thermal efficiency for different cases. Using a series of restricted action sets in this environment, our policy was able to reproduce three known thermodynamic cycles. We also introduce an irreversible action, creating an unknown thermodynamic cycle that the agent helps discover, showing how reinforcement learning can find solutions to new problems. We also discuss shortcomings of the method used, the importance of understanding the class of problem being handled, and why some methods can only be used for certain classes of problems."],"dc:identifier.uri":["https://hdl.handle.net/10155/1108"],"dc:language.iso":["en"],"dc:subject":["Reinforcement learning","Machine learning","Mathematics","Physics","Chemistry"],"dc:title":["Perpetually playing physics"],"dc:type":["Thesis"],"thesis:degree_discipline":["Modelling and Computational Science"],"thesis:degree_name":["Master of Science (MSc)"],"thesis:institution_name":["University of Ontario Institute of Technology"]},"updated_at":"2026-07-24T05:35:20Z"}