{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/135946"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/135946","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Analysis of Memory Access Patterns for Large Language Model Inference","abstract":"The use of tiered heterogeneous memory systems in HPC workloads is growing in popularity as the increasing memory requirements for these workloads outpace the decline in the cost- per-gigabyte of fast DRAM; however, the Linux kernel has no intelligent strategy to manage these tiered memory systems. Because of this limitation, a great deal of research has been conducted to identify policies that make efficient use of these systems. Much of this prior research focuses on deep learning tasks, while only a few focus on inference for large models. The training and inference workloads for the same type of model are quite different: in training, the task is to continuously update the weights matrices with knowledge gained from each training datum, while in inference, the workload only reads from the weights. Training for neural networks also involves accesses in reverse order to what is used in inference, in a training technique called backpropagation. This thesis presents a memory access pattern heatmap tool that can track evolving access patterns through the lifetime of a workload. This tool is applied to llama.cpp, an LLM inference tool, to identify memory access patterns between remote and local NUMA nodes. The thesis then explores two basic NUMA page placement strategies, where all memory is bound to either the local or remote NUMA nodes to identify the impact of poor NUMA policies on performance and compares them to the default Linux strategy.","abstract_html":"The use of tiered heterogeneous memory systems in HPC workloads is growing in popularity as the increasing memory requirements for these workloads outpace the decline in the cost- per-gigabyte of fast DRAM; however, the Linux kernel has no intelligent strategy to manage these tiered memory systems. Because of this limitation, a great deal of research has been conducted to identify policies that make efficient use of these systems. Much of this prior research focuses on deep learning tasks, while only a few focus on inference for large models. The training and inference workloads for the same type of model are quite different: in training, the task is to continuously update the weights matrices with knowledge gained from each training datum, while in inference, the workload only reads from the weights. Training for neural networks also involves accesses in reverse order to what is used in inference, in a training technique called backpropagation. This thesis presents a memory access pattern heatmap tool that can track evolving access patterns through the lifetime of a workload. This tool is applied to llama.cpp, an LLM inference tool, to identify memory access patterns between remote and local NUMA nodes. The thesis then explores two basic NUMA page placement strategies, where all memory is bound to either the local or remote NUMA nodes to identify the impact of poor NUMA policies on performance and compares them to the default Linux strategy.","abstract_has_math":false,"creators":["Fisher, Max Henry"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Computer Science & Applications","degree_department":"Computer Science and#38; Applications","school":null,"contributors":[],"advisors":[],"committee_chairs":["Nikolopoulos, Dimitrios S."],"committee_members":["Back, Godmar Volker","Li, Huaicheng"],"year":2025,"date_issued":"2025-07-09","date_published":"2025-07-09","updated_at":"2026-07-22T22:19:06Z","subjects":["NUMA","Page Placement","Memory Access Patterns","High Performance Computing","LLM Inference"],"languages":["en"],"rights":["Creative Commons Attribution-ShareAlike 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by-sa/4.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:43726"],"render_values":[{"text":"vt_gsexam:43726","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/135946","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Nikolopoulos, Dimitrios S."]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Back, Godmar Volker","Li, Huaicheng"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science and#38; Applications"]},{"key":"dc:creator","label":"Author","values":["Fisher, Max Henry"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-07-10T08:00:19Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-07-10T08:00:19Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-07-09"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science & Applications"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["NUMA","Page Placement","Memory Access Patterns","High Performance Computing","LLM Inference"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Creative Commons Attribution-ShareAlike 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by-sa/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:43726"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/135946"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The use of tiered heterogeneous memory systems in HPC workloads is growing in popularity as the increasing memory requirements for these workloads outpace the decline in the cost- per-gigabyte of fast DRAM; however, the Linux kernel has no intelligent strategy to manage these tiered memory systems. Because of this limitation, a great deal of research has been conducted to identify policies that make efficient use of these systems. Much of this prior research focuses on deep learning tasks, while only a few focus on inference for large models. The training and inference workloads for the same type of model are quite different: in training, the task is to continuously update the weights matrices with knowledge gained from each training datum, while in inference, the workload only reads from the weights. Training for neural networks also involves accesses in reverse order to what is used in inference, in a training technique called backpropagation. This thesis presents a memory access pattern heatmap tool that can track evolving access patterns through the lifetime of a workload. This tool is applied to llama.cpp, an LLM inference tool, to identify memory access patterns between remote and local NUMA nodes. The thesis then explores two basic NUMA page placement strategies, where all memory is bound to either the local or remote NUMA nodes to identify the impact of poor NUMA policies on performance and compares them to the default Linux strategy."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Scientific computing often involves running programs with very large memory footprints that might not fit in the available memory for the system on which they run. Because expanding the available memory on a system can be expensive, tiered memory systems, which provide the illusion of a fast and large memory by storing some data in regular small, expensive, and fast memory (DRAM) and the rest of the data in large, cheap, and slow memory (NVM), have been growing in popularity. However, effective use of such systems requires an intelligent strategy for moving data between tiers of memory. Past research has explored strategies that leverage data access patterns in programs like those that train machine learning models, but few explore leveraging the access patterns of running machine learning models. This thesis explores how llama.cpp, a program that runs large language models, makes accesses to data and how those patterns change as the program runs. It also explores how differing data placement policies impact performance. It compares how quickly llama.cpp runs when all data is allocated to a slow memory node to when all data is allocated to a fast memory node, reflecting how important these strategies are to fast performance."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Analysis of Memory Access Patterns for Large Language Model Inference"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Nikolopoulos, Dimitrios S."],"dc:contributor.committeemember":["Back, Godmar Volker","Li, Huaicheng"],"dc:contributor.department":["Computer Science and#38; Applications"],"dc:creator":["Fisher, Max Henry"],"dc:date.accessioned":["2025-07-10T08:00:19Z"],"dc:date.available":["2025-07-10T08:00:19Z"],"dc:date.issued":["2025-07-09"],"dc:description.abstract":["The use of tiered heterogeneous memory systems in HPC workloads is growing in popularity as the increasing memory requirements for these workloads outpace the decline in the cost- per-gigabyte of fast DRAM; however, the Linux kernel has no intelligent strategy to manage these tiered memory systems. Because of this limitation, a great deal of research has been conducted to identify policies that make efficient use of these systems. Much of this prior research focuses on deep learning tasks, while only a few focus on inference for large models. The training and inference workloads for the same type of model are quite different: in training, the task is to continuously update the weights matrices with knowledge gained from each training datum, while in inference, the workload only reads from the weights. Training for neural networks also involves accesses in reverse order to what is used in inference, in a training technique called backpropagation. This thesis presents a memory access pattern heatmap tool that can track evolving access patterns through the lifetime of a workload. This tool is applied to llama.cpp, an LLM inference tool, to identify memory access patterns between remote and local NUMA nodes. The thesis then explores two basic NUMA page placement strategies, where all memory is bound to either the local or remote NUMA nodes to identify the impact of poor NUMA policies on performance and compares them to the default Linux strategy."],"dc:description.abstractgeneral":["Scientific computing often involves running programs with very large memory footprints that might not fit in the available memory for the system on which they run. Because expanding the available memory on a system can be expensive, tiered memory systems, which provide the illusion of a fast and large memory by storing some data in regular small, expensive, and fast memory (DRAM) and the rest of the data in large, cheap, and slow memory (NVM), have been growing in popularity. However, effective use of such systems requires an intelligent strategy for moving data between tiers of memory. Past research has explored strategies that leverage data access patterns in programs like those that train machine learning models, but few explore leveraging the access patterns of running machine learning models. This thesis explores how llama.cpp, a program that runs large language models, makes accesses to data and how those patterns change as the program runs. It also explores how differing data placement policies impact performance. It compares how quickly llama.cpp runs when all data is allocated to a slow memory node to when all data is allocated to a fast memory node, reflecting how important these strategies are to fast performance."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:43726"],"dc:identifier.uri":["https://hdl.handle.net/10919/135946"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["Creative Commons Attribution-ShareAlike 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by-sa/4.0/"],"dc:subject":["NUMA","Page Placement","Memory Access Patterns","High Performance Computing","LLM Inference"],"dc:title":["Analysis of Memory Access Patterns for Large Language Model Inference"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science & Applications"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:06Z"}