{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/123358"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/123358","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"On-node performance optimization of a Monte-Carlo transport code for leadership architectures","abstract":"The tally system in Monte Carlo neutron transport codes accounts for a significant fraction of the total execution time. This project studied the tally performance of a Monte Carlo neutron transport code (i.e., OpenMC) and implemented several optimizations to address the major bottlenecks. First, a comprehensive profiling analysis was carried out on modern Intel micro-architectures (i.e., Intel Xeon Phi and Intel Xeon Platinum 8180) to understand what hardware and settings configurations were optimal. The specific modules and subroutines that were responsible for the performance drop were also highlighted. The first round of optimizations were specific to the information that the profiling analysis provided. Both the nuclide and the reaction index searches were found to be inefficient. As a result, the two searches were improved with the implementation of direct address tables, which have a single search efficiency of O(1) and a small memory footprint.","abstract_html":"The tally system in Monte Carlo neutron transport codes accounts for a significant fraction of the total execution time. This project studied the tally performance of a Monte Carlo neutron transport code (i.e., OpenMC) and implemented several optimizations to address the major bottlenecks. First, a comprehensive profiling analysis was carried out on modern Intel micro-architectures (i.e., Intel Xeon Phi and Intel Xeon Platinum 8180) to understand what hardware and settings configurations were optimal. The specific modules and subroutines that were responsible for the performance drop were also highlighted. The first round of optimizations were specific to the information that the profiling analysis provided. Both the nuclide and the reaction index searches were found to be inefficient. As a result, the two searches were improved with the implementation of direct address tables, which have a single search efficiency of O(1) and a small memory footprint.","abstract_has_math":false,"creators":["Salcedo-Pérez, José Luis."],"institution":"Massachusetts Institute of Technology","degree_name":"Master","degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Department of Nuclear Science and Engineering","school":null,"contributors":[],"advisors":["Benoit Forget and Kord Smith."],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019","date_published":"2019","updated_at":"2026-07-22T22:22:31Z","subjects":["Nuclear Science and Engineering."],"languages":["eng"],"rights":["MIT theses are protected by copyright. They may be viewed, downloaded, or printed from this source but further reproduction or distribution in any format is prohibited without written permission."],"rights_urls":["http://dspace.mit.edu/handle/1721.1/7582"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1721.1/123358","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Benoit Forget and Kord Smith."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Department of Nuclear Science and Engineering","NucEng"]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Massachusetts Institute of Technology. Department of Nuclear Science and Engineering."]},{"key":"dc:creator","label":"Author","values":["Salcedo-Pérez, José Luis."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2020-01-08T19:33:01Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2020-01-08T19:33:01Z"]},{"key":"dc:date.issued","label":"Date","values":["2019"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Nuclear Science and Engineering."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["MIT theses are protected by copyright. They may be viewed, downloaded, or printed from this source but further reproduction or distribution in any format is prohibited without written permission."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://dspace.mit.edu/handle/1721.1/7582"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1721.1/123358"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This electronic version was submitted by the student author. The certified thesis is available in the Institute Archives and Special Collections.","Thesis: S.M., Massachusetts Institute of Technology, Department of Nuclear Science and Engineering, 2019","Cataloged from student-submitted PDF version of thesis.","Includes bibliographical references (pages 83-85)."]},{"key":"dc:description.abstract","label":"Abstract","values":["The tally system in Monte Carlo neutron transport codes accounts for a significant fraction of the total execution time. This project studied the tally performance of a Monte Carlo neutron transport code (i.e., OpenMC) and implemented several optimizations to address the major bottlenecks. First, a comprehensive profiling analysis was carried out on modern Intel micro-architectures (i.e., Intel Xeon Phi and Intel Xeon Platinum 8180) to understand what hardware and settings configurations were optimal. The specific modules and subroutines that were responsible for the performance drop were also highlighted. The first round of optimizations were specific to the information that the profiling analysis provided. Both the nuclide and the reaction index searches were found to be inefficient. As a result, the two searches were improved with the implementation of direct address tables, which have a single search efficiency of O(1) and a small memory footprint.","Moreover, a linear array cache was also introduced to store the following cross-sections: (n, 2n), (n, 3n), (n, 4n), (n, p), (n, [alpha]), and (n, [gamma]). These cross-sections, together with (n; fission), are indispensable to solve the transition matrix of the Bateman equations during transmutation analysis. As a result, pre-computing and storing them before tally-time eliminated redundant computations in the case a high energy particle travels through multiple fuel regions without colliding. Overall, these optimizations resulted in speedups of 2.31x and 2.15x for the Xeon Platinum and Xeon Phi, respectively. Further, this project also presents an alternative method to compute reaction rate tallies. In general, tallying all of the aforementioned seven rates through a Monte Carlo simulation can be quite expensive for realistic light water reactors.","Another approach would be to collapse a very fine-group flux together with a pregenerated multigroup cross section (constructed with the same energy grid). While this approach does provide a 3x speedup in the OpenMC active cycles performance, it also introduces a considerable memory penalty. The issue is that thousands of groups are needed to accurately resolve the (n, [gamma]) rates, most notably that of 238U. This study explores a hybrid approach in which (n, [gamma]) and (n, fission) are handled with a standard reaction rate tally while the remaining reaction rates are computed through the flux tally route. This option provides more flexibility in reducing the total number of groups because the remaining reactions outside of (n, fission) and (n, [gamma]) usually have smoother shapes. Performance was tested on five benchmarks with depleted fuel and increasing geometrical complexity.","Results showed that the hybrid tally method provided decent speedups ranging from 1.30x to 1.75x in the active cycles across all benchmarks. Multiple error analyses were also carried out on the proposed hybrid method; the results show that even when going as low as 300 groups, the eigenvalue is still within 100 pcm of a traditional simulation."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["S.M."]},{"key":"dc:title","label":"Title","values":["On-node performance optimization of a Monte-Carlo transport code for leadership architectures"]}]}],"canonical_facts":{"dc:contributor.advisor":["Benoit Forget and Kord Smith."],"dc:contributor.department":["Massachusetts Institute of Technology. Department of Nuclear Science and Engineering","NucEng"],"dc:contributor.other":["Massachusetts Institute of Technology. Department of Nuclear Science and Engineering."],"dc:creator":["Salcedo-Pérez, José Luis."],"dc:date.accessioned":["2020-01-08T19:33:01Z"],"dc:date.available":["2020-01-08T19:33:01Z"],"dc:date.issued":["2019"],"dc:description":["This electronic version was submitted by the student author. The certified thesis is available in the Institute Archives and Special Collections.","Thesis: S.M., Massachusetts Institute of Technology, Department of Nuclear Science and Engineering, 2019","Cataloged from student-submitted PDF version of thesis.","Includes bibliographical references (pages 83-85)."],"dc:description.abstract":["The tally system in Monte Carlo neutron transport codes accounts for a significant fraction of the total execution time. This project studied the tally performance of a Monte Carlo neutron transport code (i.e., OpenMC) and implemented several optimizations to address the major bottlenecks. First, a comprehensive profiling analysis was carried out on modern Intel micro-architectures (i.e., Intel Xeon Phi and Intel Xeon Platinum 8180) to understand what hardware and settings configurations were optimal. The specific modules and subroutines that were responsible for the performance drop were also highlighted. The first round of optimizations were specific to the information that the profiling analysis provided. Both the nuclide and the reaction index searches were found to be inefficient. As a result, the two searches were improved with the implementation of direct address tables, which have a single search efficiency of O(1) and a small memory footprint.","Moreover, a linear array cache was also introduced to store the following cross-sections: (n, 2n), (n, 3n), (n, 4n), (n, p), (n, [alpha]), and (n, [gamma]). These cross-sections, together with (n; fission), are indispensable to solve the transition matrix of the Bateman equations during transmutation analysis. As a result, pre-computing and storing them before tally-time eliminated redundant computations in the case a high energy particle travels through multiple fuel regions without colliding. Overall, these optimizations resulted in speedups of 2.31x and 2.15x for the Xeon Platinum and Xeon Phi, respectively. Further, this project also presents an alternative method to compute reaction rate tallies. In general, tallying all of the aforementioned seven rates through a Monte Carlo simulation can be quite expensive for realistic light water reactors.","Another approach would be to collapse a very fine-group flux together with a pregenerated multigroup cross section (constructed with the same energy grid). While this approach does provide a 3x speedup in the OpenMC active cycles performance, it also introduces a considerable memory penalty. The issue is that thousands of groups are needed to accurately resolve the (n, [gamma]) rates, most notably that of 238U. This study explores a hybrid approach in which (n, [gamma]) and (n, fission) are handled with a standard reaction rate tally while the remaining reaction rates are computed through the flux tally route. This option provides more flexibility in reducing the total number of groups because the remaining reactions outside of (n, fission) and (n, [gamma]) usually have smoother shapes. Performance was tested on five benchmarks with depleted fuel and increasing geometrical complexity.","Results showed that the hybrid tally method provided decent speedups ranging from 1.30x to 1.75x in the active cycles across all benchmarks. Multiple error analyses were also carried out on the proposed hybrid method; the results show that even when going as low as 300 groups, the eigenvalue is still within 100 pcm of a traditional simulation."],"dc:description.degree":["S.M."],"dc:identifier.uri":["https://hdl.handle.net/1721.1/123358"],"dc:language.iso":["eng"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["MIT theses are protected by copyright. They may be viewed, downloaded, or printed from this source but further reproduction or distribution in any format is prohibited without written permission."],"dc:rights.uri":["http://dspace.mit.edu/handle/1721.1/7582"],"dc:subject":["Nuclear Science and Engineering."],"dc:title":["On-node performance optimization of a Monte-Carlo transport code for leadership architectures"],"dc:type":["Thesis"],"thesis:degree_name":["Master"]},"updated_at":"2026-07-22T22:22:31Z"}