{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129176"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129176","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Pushing the limits of long context LLM inference via KV cache compression","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Sharma, Akshat"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhang, Minjia"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-05","date_published":"2025-05","updated_at":"2026-07-22T22:25:04Z","subjects":["Systems for LLM","LLM Inference"],"languages":["en","eng"],"rights":["Copyright 2025 Akshat Sharma"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129176","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhang, Minjia"]},{"key":"dc:creator","label":"Author","values":["Sharma, Akshat"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-05","2025-04-20"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Systems for LLM","LLM Inference"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Akshat Sharma"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129176"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Akshat Sharma, accepted the attached license on 2025-04-18 at 14:17.","The student, Akshat Sharma, submitted this Thesis for approval on 2025-04-18 at 14:24.","This Thesis was approved for publication on 2025-04-20 at 16:29.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21668 on 2025-10-19 at 18:09:09","Efficiently deploying/serving LLMs has become remarkably challenging due to their excessive memory and computational requirements. A critical bottleneck in LLM inference is the memory footprint of the Key-Value (KV) cache, particularly in tasks involving long-context understanding and generation. To address these challenges we introduce MiniKV, a hybrid KV cache optimization technique which compresses the KV cache by combining token eviction and 2-bit quantization. Our approach aims to significantly reduce memory usage while maintaining high accuracy on downstream tasks such as question answering, summarization, code generation, and retrieval. Our evaluations demonstrate that MiniKV achieves an 86% reduction in KV cache size while recovering over 98.5% accuracy across downstream tasks. This sets a new state-of-the-art in balancing accuracy and compression, with notable improvements in inference latency and throughput."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Pushing the limits of long context LLM inference via KV cache compression"]}]}],"canonical_facts":{"dc:contributor":["Zhang, Minjia"],"dc:creator":["Sharma, Akshat"],"dc:date":["2025-05","2025-04-20"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Akshat Sharma, accepted the attached license on 2025-04-18 at 14:17.","The student, Akshat Sharma, submitted this Thesis for approval on 2025-04-18 at 14:24.","This Thesis was approved for publication on 2025-04-20 at 16:29.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21668 on 2025-10-19 at 18:09:09","Efficiently deploying/serving LLMs has become remarkably challenging due to their excessive memory and computational requirements. A critical bottleneck in LLM inference is the memory footprint of the Key-Value (KV) cache, particularly in tasks involving long-context understanding and generation. To address these challenges we introduce MiniKV, a hybrid KV cache optimization technique which compresses the KV cache by combining token eviction and 2-bit quantization. Our approach aims to significantly reduce memory usage while maintaining high accuracy on downstream tasks such as question answering, summarization, code generation, and retrieval. Our evaluations demonstrate that MiniKV achieves an 86% reduction in KV cache size while recovering over 98.5% accuracy across downstream tasks. This sets a new state-of-the-art in balancing accuracy and compression, with notable improvements in inference latency and throughput."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129176"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Akshat Sharma"],"dc:subject":["Systems for LLM","LLM Inference"],"dc:title":["Pushing the limits of long context LLM inference via KV cache compression"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}