{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/18333"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/18333","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Performance evaluation and enhancement of Dendro","abstract":"DENDRO is a collection of tools for solving Finite Element problems in parallel. This package is written in C++ using the standard template library (STL) and uses the Message Passing (MPI). Dendro uses an octree data-structure to solve image-registration problems using finite element techniques. For analyzing the behavior of the package in terms of speed-up and scalability, it is important to know which part of the package is consuming most of the execution-time. The single node performance and the overall performance of the package is dependent on the code-organization and class-hierarchy. We used the PETSC profiler to collect the performance statistics and instrument the code to know which part of the code takes most of the time. Along with the function-specific execution timings, PETSC profiler also provides the information regarding how many floating point operations is being performed in total and on average (FLOP/second). PETSC also provides information related to memory usage and number of MPI messages and reductions being performed to execute that particular function. We have analyzed these performance-statistics to provide some guidelines to how we can make Dendro more efficient by optimizing certain functions. We obtained around 12X speedup over the performance of (default) Dendro by using compiler-provided optimizations and achieved more than 65% speedup over compiler optimized performance (20X over the naive Dendro performance) by manually tuning some-block of code along with the compiler-optimizations.","abstract_html":"DENDRO is a collection of tools for solving Finite Element problems in parallel. This package is written in C++ using the standard template library (STL) and uses the Message Passing (MPI). Dendro uses an octree data-structure to solve image-registration problems using finite element techniques. For analyzing the behavior of the package in terms of speed-up and scalability, it is important to know which part of the package is consuming most of the execution-time. The single node performance and the overall performance of the package is dependent on the code-organization and class-hierarchy. We used the PETSC profiler to collect the performance statistics and instrument the code to know which part of the code takes most of the time. Along with the function-specific execution timings, PETSC profiler also provides the information regarding how many floating point operations is being performed in total and on average (FLOP/second). PETSC also provides information related to memory usage and number of MPI messages and reductions being performed to execute that particular function. We have analyzed these performance-statistics to provide some guidelines to how we can make Dendro more efficient by optimizing certain functions. We obtained around 12X speedup over the performance of (default) Dendro by using compiler-provided optimizations and achieved more than 65% speedup over compiler optimized performance (20X over the naive Dendro performance) by manually tuning some-block of code along with the compiler-optimizations.","abstract_has_math":false,"creators":["Mukherjee, Jayanta"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Gropp, William D."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-01-14T22:46:33Z","date_published":"2011-01-14T22:46:33Z","updated_at":"2026-07-22T22:25:11Z","subjects":["Performance modeling","performance tuning","optimization","speedup","scalability"],"languages":["en"],"rights":["2010 by Jayanta Mukherjee. All rights reserved."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/18333","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Gropp, William D."]},{"key":"dc:creator","label":"Author","values":["Mukherjee, Jayanta"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2011-01-14T22:46:33Z","2010-12"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Performance modeling","performance tuning","optimization","speedup","scalability"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["2010 by Jayanta Mukherjee. All rights reserved."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/18333"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["DENDRO is a collection of tools for solving Finite Element problems in parallel. This package is written in C++ using the standard template library (STL) and uses the Message Passing (MPI). Dendro uses an octree data-structure to solve image-registration problems using finite element techniques. For analyzing the behavior of the package in terms of speed-up and scalability, it is important to know which part of the package is consuming most of the execution-time. The single node performance and the overall performance of the package is dependent on the code-organization and class-hierarchy. We used the PETSC profiler to collect the performance statistics and instrument the code to know which part of the code takes most of the time. Along with the function-specific execution timings, PETSC profiler also provides the information regarding how many floating point operations is being performed in total and on average (FLOP/second). PETSC also provides information related to memory usage and number of MPI messages and reductions being performed to execute that particular function. We have analyzed these performance-statistics to provide some guidelines to how we can make Dendro more efficient by optimizing certain functions. We obtained around 12X speedup over the performance of (default) Dendro by using compiler-provided optimizations and achieved more than 65% speedup over compiler optimized performance (20X over the naive Dendro performance) by manually tuning some-block of code along with the compiler-optimizations.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2010-12-03T21:40:17Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 8 JayantaMukherjee-ms-thesis.tex: 6431 bytes, checksum: 5f676a6657cb9c1e89eea980c2a28d22 (MD5) DendroRef.bib: 10033 bytes, checksum: 50e6362cf0d6d2e78731a0197126ab04 (MD5) 4-opts.tex: 10329 bytes, checksum: da333a97af6bfe333e0f6d4325850bb3 (MD5) 3-exp.tex: 9195 bytes, checksum: 699cfc830c581a52373d808af04999ed (MD5) 2-related.tex: 7365 bytes, checksum: 38c93219995ef86476d294b79203bf96 (MD5) 1-introduction.tex: 3963 bytes, checksum: f3ab36f7d4c96ca2343299ce4215368d (MD5) Appendix.tex: 9294 bytes, checksum: 6692bbaa74f4837bed849a9e5021c4eb (MD5) Mukherjee_Jayanta.pdf: 881016 bytes, checksum: 666cc63d9905601fdc7e6cb83d0cfb40 (MD5)","Made available in DSpace on 2011-01-14T22:46:33Z (GMT). No. of bitstreams: 9 Appendix.tex: 9294 bytes, checksum: 6692bbaa74f4837bed849a9e5021c4eb (MD5) 1-introduction.tex: 3963 bytes, checksum: f3ab36f7d4c96ca2343299ce4215368d (MD5) 2-related.tex: 7365 bytes, checksum: 38c93219995ef86476d294b79203bf96 (MD5) 3-exp.tex: 9195 bytes, checksum: 699cfc830c581a52373d808af04999ed (MD5) 4-opts.tex: 10329 bytes, checksum: da333a97af6bfe333e0f6d4325850bb3 (MD5) DendroRef.bib: 10033 bytes, checksum: 50e6362cf0d6d2e78731a0197126ab04 (MD5) Mukherjee_Jayanta.pdf: 880978 bytes, checksum: 3a9c15ea6dad9390a028069fd3bbf315 (MD5) JayantaMukherjee-ms-thesis.tex: 6431 bytes, checksum: 5f676a6657cb9c1e89eea980c2a28d22 (MD5) license.txt: 4067 bytes, checksum: 13140cd82d23d51540ef9769512e9207 (MD5)"]},{"key":"dc:title","label":"Title","values":["Performance evaluation and enhancement of Dendro"]}]}],"canonical_facts":{"dc:contributor":["Gropp, William D."],"dc:creator":["Mukherjee, Jayanta"],"dc:date":["2011-01-14T22:46:33Z","2010-12"],"dc:description":["DENDRO is a collection of tools for solving Finite Element problems in parallel. This package is written in C++ using the standard template library (STL) and uses the Message Passing (MPI). Dendro uses an octree data-structure to solve image-registration problems using finite element techniques. For analyzing the behavior of the package in terms of speed-up and scalability, it is important to know which part of the package is consuming most of the execution-time. The single node performance and the overall performance of the package is dependent on the code-organization and class-hierarchy. We used the PETSC profiler to collect the performance statistics and instrument the code to know which part of the code takes most of the time. Along with the function-specific execution timings, PETSC profiler also provides the information regarding how many floating point operations is being performed in total and on average (FLOP/second). PETSC also provides information related to memory usage and number of MPI messages and reductions being performed to execute that particular function. We have analyzed these performance-statistics to provide some guidelines to how we can make Dendro more efficient by optimizing certain functions. We obtained around 12X speedup over the performance of (default) Dendro by using compiler-provided optimizations and achieved more than 65% speedup over compiler optimized performance (20X over the naive Dendro performance) by manually tuning some-block of code along with the compiler-optimizations.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2010-12-03T21:40:17Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 8 JayantaMukherjee-ms-thesis.tex: 6431 bytes, checksum: 5f676a6657cb9c1e89eea980c2a28d22 (MD5) DendroRef.bib: 10033 bytes, checksum: 50e6362cf0d6d2e78731a0197126ab04 (MD5) 4-opts.tex: 10329 bytes, checksum: da333a97af6bfe333e0f6d4325850bb3 (MD5) 3-exp.tex: 9195 bytes, checksum: 699cfc830c581a52373d808af04999ed (MD5) 2-related.tex: 7365 bytes, checksum: 38c93219995ef86476d294b79203bf96 (MD5) 1-introduction.tex: 3963 bytes, checksum: f3ab36f7d4c96ca2343299ce4215368d (MD5) Appendix.tex: 9294 bytes, checksum: 6692bbaa74f4837bed849a9e5021c4eb (MD5) Mukherjee_Jayanta.pdf: 881016 bytes, checksum: 666cc63d9905601fdc7e6cb83d0cfb40 (MD5)","Made available in DSpace on 2011-01-14T22:46:33Z (GMT). No. of bitstreams: 9 Appendix.tex: 9294 bytes, checksum: 6692bbaa74f4837bed849a9e5021c4eb (MD5) 1-introduction.tex: 3963 bytes, checksum: f3ab36f7d4c96ca2343299ce4215368d (MD5) 2-related.tex: 7365 bytes, checksum: 38c93219995ef86476d294b79203bf96 (MD5) 3-exp.tex: 9195 bytes, checksum: 699cfc830c581a52373d808af04999ed (MD5) 4-opts.tex: 10329 bytes, checksum: da333a97af6bfe333e0f6d4325850bb3 (MD5) DendroRef.bib: 10033 bytes, checksum: 50e6362cf0d6d2e78731a0197126ab04 (MD5) Mukherjee_Jayanta.pdf: 880978 bytes, checksum: 3a9c15ea6dad9390a028069fd3bbf315 (MD5) JayantaMukherjee-ms-thesis.tex: 6431 bytes, checksum: 5f676a6657cb9c1e89eea980c2a28d22 (MD5) license.txt: 4067 bytes, checksum: 13140cd82d23d51540ef9769512e9207 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/18333"],"dc:language":["en"],"dc:rights":["2010 by Jayanta Mukherjee. All rights reserved."],"dc:subject":["Performance modeling","performance tuning","optimization","speedup","scalability"],"dc:title":["Performance evaluation and enhancement of Dendro"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:11Z"}