{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124607"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124607","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Beyond rules: leveraging Large Language Models for code-data separation in binary disassembly","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2026-05-01","abstract_has_math":false,"creators":["Diwan, Nirav"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Wang, Gang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:02Z","subjects":["Large Language Models","Binary Analysis","Code-data Separation","Unsupervised Domain Adaptation","Security"],"languages":["en","eng"],"rights":["Copyright 2024 Nirav Diwan"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124607","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Wang, Gang"]},{"key":"dc:creator","label":"Author","values":["Diwan, Nirav"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-05-02"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Large Language Models","Binary Analysis","Code-data Separation","Unsupervised Domain Adaptation","Security"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Nirav Diwan"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124607"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","The student, Nirav Diwan, accepted the attached license on 2024-05-01 at 18:24.","The student, Nirav Diwan, submitted this Thesis for approval on 2024-05-01 at 18:54.","This Thesis was approved for publication on 2024-05-02 at 11:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20747 on 2024-09-16 at 00:45:00","Static binary analysis serves as a critical technique for identifying security vulnerabilities in binaries without source code access. The first step of static binary analysis is disassembly, which involves deconstructing the binary file to identify code and data instructions (also known as the code-data separation problem). Current methods for code-data separation assume a fixed or standard file format of the binary file. However, with the proliferation of Internet-of-Things (IoT) devices, new non-standard file formats, which do not conform to a fixed format, are becoming prevalent. This presents a hurdle in performing binary analysis tasks (e.g., detecting security flaws, malware classification, license obligations) for such non-standard file formats. In this work, we examine the code and data distributions for standard and non-standard file formats. Our analysis indicates a distribution shift between standard and non-standard file formats, motivating the need to tackle code-data separation for non-standard file formats. We approach this problem as an unsupervised domain adaptation problem by proposing a pseudo-labeling approach based on Large Language Models. Our best model achieves high performance on standard binary files (F1-Score = 0.99) and non-standard binary files (F1-Score = 0.95). Finally, we discuss our findings and the limitations of our approach."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Beyond rules: leveraging Large Language Models for code-data separation in binary disassembly"]}]}],"canonical_facts":{"dc:contributor":["Wang, Gang"],"dc:creator":["Diwan, Nirav"],"dc:date":["2024-05","2024-05-02"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","The student, Nirav Diwan, accepted the attached license on 2024-05-01 at 18:24.","The student, Nirav Diwan, submitted this Thesis for approval on 2024-05-01 at 18:54.","This Thesis was approved for publication on 2024-05-02 at 11:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20747 on 2024-09-16 at 00:45:00","Static binary analysis serves as a critical technique for identifying security vulnerabilities in binaries without source code access. The first step of static binary analysis is disassembly, which involves deconstructing the binary file to identify code and data instructions (also known as the code-data separation problem). Current methods for code-data separation assume a fixed or standard file format of the binary file. However, with the proliferation of Internet-of-Things (IoT) devices, new non-standard file formats, which do not conform to a fixed format, are becoming prevalent. This presents a hurdle in performing binary analysis tasks (e.g., detecting security flaws, malware classification, license obligations) for such non-standard file formats. In this work, we examine the code and data distributions for standard and non-standard file formats. Our analysis indicates a distribution shift between standard and non-standard file formats, motivating the need to tackle code-data separation for non-standard file formats. We approach this problem as an unsupervised domain adaptation problem by proposing a pseudo-labeling approach based on Large Language Models. Our best model achieves high performance on standard binary files (F1-Score = 0.99) and non-standard binary files (F1-Score = 0.95). Finally, we discuss our findings and the limitations of our approach."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124607"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Nirav Diwan"],"dc:subject":["Large Language Models","Binary Analysis","Code-data Separation","Unsupervised Domain Adaptation","Security"],"dc:title":["Beyond rules: leveraging Large Language Models for code-data separation in binary disassembly"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}