{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129293"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129293","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"QStore: quantization-aware compressed model storage","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Shah, Raunak"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Park, Yongjoo"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-05-06","date_published":"2025-05-06","updated_at":"2026-07-22T22:25:04Z","subjects":["Compression","Quantization","LLMs","File Formats","Storage","Machine Learning","AI"],"languages":["en","eng"],"rights":["Copyright 2025 Raunak Shah"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129293","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Park, Yongjoo"]},{"key":"dc:creator","label":"Author","values":["Shah, Raunak"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-05-06","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Compression","Quantization","LLMs","File Formats","Storage","Machine Learning","AI"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Raunak Shah"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129293"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Raunak Shah, accepted the attached license on 2025-05-01 at 01:17.","The student, Raunak Shah, submitted this Thesis for approval on 2025-05-01 at 01:30.","This Thesis was approved for publication on 2025-05-06 at 11:52.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22138 on 2025-10-19 at 18:11:27","Modern applications commonly leverage large, multi-modal foundation models. Such use cases often feature complex workflows that demand the storage and usage of models in multiple precisions. However, to make model development accessible to the average user, model providers opt to maintain separate files for each precision (e.g., INT8, BF16); Hence, naively handling model usage in multi-precision workflows (e.g., downloading and storing each precision separately) can incur prohibitive storage costs. We present QStore, a unified, lossless compression format for simultaneously storing a model in two (high and low) precisions. QStore’s encoding scheme stores a pair of different-precision models with even less storage cost versus storing the high-precision model alone: it compresses the low-precision model, then stores a novel representation of the ’conditional information’ present in the high-precision model but not in the low precision model. Then, for model usage, QStore allows direct access to the low precision model via decoding; if the high precision model is required, QStore recovers it losslessly by applying the stored conditional information onto the decoded low-precision model. We evaluate QStore for compressing multiple precisions of popular foundation models, and show that QStore reduces the overall storage footprint by up to 2.2× (45% of the original size) while enabling up to 1.7× and 1.8× faster model saving and loading versus existing approaches."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["QStore: quantization-aware compressed model storage"]}]}],"canonical_facts":{"dc:contributor":["Park, Yongjoo"],"dc:creator":["Shah, Raunak"],"dc:date":["2025-05-06","2025-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Raunak Shah, accepted the attached license on 2025-05-01 at 01:17.","The student, Raunak Shah, submitted this Thesis for approval on 2025-05-01 at 01:30.","This Thesis was approved for publication on 2025-05-06 at 11:52.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22138 on 2025-10-19 at 18:11:27","Modern applications commonly leverage large, multi-modal foundation models. Such use cases often feature complex workflows that demand the storage and usage of models in multiple precisions. However, to make model development accessible to the average user, model providers opt to maintain separate files for each precision (e.g., INT8, BF16); Hence, naively handling model usage in multi-precision workflows (e.g., downloading and storing each precision separately) can incur prohibitive storage costs. We present QStore, a unified, lossless compression format for simultaneously storing a model in two (high and low) precisions. QStore’s encoding scheme stores a pair of different-precision models with even less storage cost versus storing the high-precision model alone: it compresses the low-precision model, then stores a novel representation of the ’conditional information’ present in the high-precision model but not in the low precision model. Then, for model usage, QStore allows direct access to the low precision model via decoding; if the high precision model is required, QStore recovers it losslessly by applying the stored conditional information onto the decoded low-precision model. We evaluate QStore for compressing multiple precisions of popular foundation models, and show that QStore reduces the overall storage footprint by up to 2.2× (45% of the original size) while enabling up to 1.7× and 1.8× faster model saving and loading versus existing approaches."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129293"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Raunak Shah"],"dc:subject":["Compression","Quantization","LLMs","File Formats","Storage","Machine Learning","AI"],"dc:title":["QStore: quantization-aware compressed model storage"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}