{"id":{"repo_id":"uic","oai_identifier":"oai:figshare.com:article/32991893"},"canonical_url":"https://search.dev.ndltd.org/etd/uic/oai:figshare.com:article/32991893","repository":{"repo_id":"uic","name":"University of Illinois - Chicago","base_url":"https://api.figshare.com/v2/oai"},"display":{"title":"Performance and Energy Efficiency Insights in LLM Inference Across Hardware Accelerators","abstract":"The rapid growth of LLM-based applications—such as chatbots, coding assistants, and search engines—has driven the need to deliver LLM inference at scale. This workload is highly resource-demanding, requiring substantial compute throughput along with large, high-bandwidth memory. As a result, efficient LLM serving often depends on specialised hardware acceleration. While GPUs continue to dominate, domain-specific accelerators like TPUs and dataflow architectures are becoming increasingly compelling alternatives. In this thesis, we provide a comprehensive empirical performance study of six datacenter-grade GPUs from Nvidia, AMD, and Intel, along with two dataflow AI accelerators from Cerebras and SambaNova, evaluated across fourteen open-source LLMs. Our analysis examines the key factors that influence LLM inference performance, including model size, batch size, quantisation, and multi-GPU scaling across different parallelism strategies. Crucially, we evaluate both performance and energy efficiency to offer an energy-aware comparison across accelerator classes. Our findings show that dataflow AI accelerators deliver an order-of-magnitude speedup in throughput and latency for small batch sizes relative to GPUs. Conversely, GPUs provide larger HBM memory capacities and benefit from a more straightforward programming model, enabling greater flexibility in batch sizing. Together, these results offer actionable insights for optimising LLM inference across a range of deployment settings.","abstract_html":"The rapid growth of LLM-based applications—such as chatbots, coding assistants, and search engines—has driven the need to deliver LLM inference at scale. This workload is highly resource-demanding, requiring substantial compute throughput along with large, high-bandwidth memory. As a result, efficient LLM serving often depends on specialised hardware acceleration. While GPUs continue to dominate, domain-specific accelerators like TPUs and dataflow architectures are becoming increasingly compelling alternatives. In this thesis, we provide a comprehensive empirical performance study of six datacenter-grade GPUs from Nvidia, AMD, and Intel, along with two dataflow AI accelerators from Cerebras and SambaNova, evaluated across fourteen open-source LLMs. Our analysis examines the key factors that influence LLM inference performance, including model size, batch size, quantisation, and multi-GPU scaling across different parallelism strategies. Crucially, we evaluate both performance and energy efficiency to offer an energy-aware comparison across accelerator classes. Our findings show that dataflow AI accelerators deliver an order-of-magnitude speedup in throughput and latency for small batch sizes relative to GPUs. Conversely, GPUs provide larger HBM memory capacities and benefit from a more straightforward programming model, enabling greater flexibility in batch sizing. Together, these results offer actionable insights for optimising LLM inference across a range of deployment settings.","abstract_has_math":false,"creators":["Giacomo Brunetta (23892750)"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-07-15T12:02:04Z","date_published":"2026-07-15T12:02:04Z","updated_at":"2026-07-27T21:33:07Z","subjects":["Computer Science"],"languages":[],"rights":["In Copyright"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.25417/uic.32991893.v1","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Giacomo Brunetta (23892750)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2026-07-15T12:02:04Z"]},{"key":"dc:relation","label":"Dc Relation","values":["https://figshare.com/articles/thesis/Performance_and_Energy_Efficiency_Insights_in_LLM_Inference_Across_Hardware_Accelerators/32991893"]},{"key":"dc:type","label":"Dc Type","values":["Text","Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["10.25417/uic.32991893.v1"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The rapid growth of LLM-based applications—such as chatbots, coding assistants, and search engines—has driven the need to deliver LLM inference at scale. This workload is highly resource-demanding, requiring substantial compute throughput along with large, high-bandwidth memory. As a result, efficient LLM serving often depends on specialised hardware acceleration. While GPUs continue to dominate, domain-specific accelerators like TPUs and dataflow architectures are becoming increasingly compelling alternatives. In this thesis, we provide a comprehensive empirical performance study of six datacenter-grade GPUs from Nvidia, AMD, and Intel, along with two dataflow AI accelerators from Cerebras and SambaNova, evaluated across fourteen open-source LLMs. Our analysis examines the key factors that influence LLM inference performance, including model size, batch size, quantisation, and multi-GPU scaling across different parallelism strategies. Crucially, we evaluate both performance and energy efficiency to offer an energy-aware comparison across accelerator classes. Our findings show that dataflow AI accelerators deliver an order-of-magnitude speedup in throughput and latency for small batch sizes relative to GPUs. Conversely, GPUs provide larger HBM memory capacities and benefit from a more straightforward programming model, enabling greater flexibility in batch sizing. Together, these results offer actionable insights for optimising LLM inference across a range of deployment settings."]},{"key":"dc:title","label":"Title","values":["Performance and Energy Efficiency Insights in LLM Inference Across Hardware Accelerators"]}]}],"canonical_facts":{"dc:creator":["Giacomo Brunetta (23892750)"],"dc:date":["2026-07-15T12:02:04Z"],"dc:description":["The rapid growth of LLM-based applications—such as chatbots, coding assistants, and search engines—has driven the need to deliver LLM inference at scale. This workload is highly resource-demanding, requiring substantial compute throughput along with large, high-bandwidth memory. As a result, efficient LLM serving often depends on specialised hardware acceleration. While GPUs continue to dominate, domain-specific accelerators like TPUs and dataflow architectures are becoming increasingly compelling alternatives. In this thesis, we provide a comprehensive empirical performance study of six datacenter-grade GPUs from Nvidia, AMD, and Intel, along with two dataflow AI accelerators from Cerebras and SambaNova, evaluated across fourteen open-source LLMs. Our analysis examines the key factors that influence LLM inference performance, including model size, batch size, quantisation, and multi-GPU scaling across different parallelism strategies. Crucially, we evaluate both performance and energy efficiency to offer an energy-aware comparison across accelerator classes. Our findings show that dataflow AI accelerators deliver an order-of-magnitude speedup in throughput and latency for small batch sizes relative to GPUs. Conversely, GPUs provide larger HBM memory capacities and benefit from a more straightforward programming model, enabling greater flexibility in batch sizing. Together, these results offer actionable insights for optimising LLM inference across a range of deployment settings."],"dc:identifier":["10.25417/uic.32991893.v1"],"dc:relation":["https://figshare.com/articles/thesis/Performance_and_Energy_Efficiency_Insights_in_LLM_Inference_Across_Hardware_Accelerators/32991893"],"dc:rights":["In Copyright"],"dc:subject":["Computer Science"],"dc:title":["Performance and Energy Efficiency Insights in LLM Inference Across Hardware Accelerators"],"dc:type":["Text","Thesis"]},"updated_at":"2026-07-27T21:33:07Z"}