{"id":{"repo_id":"nus","oai_identifier":"oai:scholarbank.nus.edu.sg:10635/309814"},"canonical_url":"https://search.dev.ndltd.org/etd/nus/oai:scholarbank.nus.edu.sg:10635/309814","repository":{"repo_id":"nus","name":"National University of Singapore","base_url":"https://scholarbank.nus.edu.sg/oai/request"},"display":{"title":"MODEL ADAPTATION FOR EDGE AI","abstract":"Deployment of Deep Neural Networks (DNNs) on edge devices presents significant challenges due to their computational demands. Existing model compression techniques often fall short by being oblivious to downstream user-specific tasks. This thesis addresses the challenge of adapting DNN models effectively on resource-limited hardware, advocating for flexible and efficient models. Firstly, we tackle the challenge of on-device adaptation for user-specific models. Given the limited on-chip memory in edge devices, data movement becomes a significant bottleneck, leading to increased energy consumption. We introduce Chameleon, an energy-efficient continual learning framework. Secondly, generic DNN models can lead to wasteful and inefficient inference. We propose CRISP, a class-aware pruning framework that employs a hybrid structured sparsity pattern. Thirdly, while low-precision integer formats are common for reduced precision arithmetic, these formats are inadequate for emerging workloads like transformers. We explore minifloats, reduced-precision floating-point formats that significantly reduce the memory footprint while offering greater flexibility due to their dynamic range. Finally, models such as Vision Transformers (ViTs) struggle with real-time inference on edge devices. We introduce SMS, a token pruning framework that dynamically determines which tokens to merge and to what extent, thereby optimizing model accuracy while adhering to the target average inference budget.","abstract_html":"Deployment of Deep Neural Networks (DNNs) on edge devices presents significant challenges due to their computational demands. Existing model compression techniques often fall short by being oblivious to downstream user-specific tasks. This thesis addresses the challenge of adapting DNN models effectively on resource-limited hardware, advocating for flexible and efficient models. Firstly, we tackle the challenge of on-device adaptation for user-specific models. Given the limited on-chip memory in edge devices, data movement becomes a significant bottleneck, leading to increased energy consumption. We introduce Chameleon, an energy-efficient continual learning framework. Secondly, generic DNN models can lead to wasteful and inefficient inference. We propose CRISP, a class-aware pruning framework that employs a hybrid structured sparsity pattern. Thirdly, while low-precision integer formats are common for reduced precision arithmetic, these formats are inadequate for emerging workloads like transformers. We explore minifloats, reduced-precision floating-point formats that significantly reduce the memory footprint while offering greater flexibility due to their dynamic range. Finally, models such as Vision Transformers (ViTs) struggle with real-time inference on edge devices. We introduce SMS, a token pruning framework that dynamically determines which tokens to merge and to what extent, thereby optimizing model accuracy while adhering to the target average inference budget.","abstract_has_math":false,"creators":["SHIVAM AGGARWAL"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-11-15","date_published":"2024-11-15","updated_at":"2026-07-24T03:32:18Z","subjects":["Quantization","Sparsity","Model Compression","Efficient AI","EdgeAI"],"languages":[],"rights":[],"rights_urls":["https://scholarbank.nus.edu.sg/bitstreams/aedc0f0c-f38b-4720-8a7f-e77860536d79/download"],"identifier_entries":[]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["SHIVAM AGGARWAL"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-11-15"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://scholarbank.nus.edu.sg/handle/10635/309814"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Quantization","Sparsity","Model Compression","Efficient AI","EdgeAI"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://scholarbank.nus.edu.sg/bitstreams/aedc0f0c-f38b-4720-8a7f-e77860536d79/download"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarbank.nus.edu.sg/bitstreams/36a63bb5-6e17-4069-a30f-cd5e4a78f277/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Deployment of Deep Neural Networks (DNNs) on edge devices presents significant challenges due to their computational demands. Existing model compression techniques often fall short by being oblivious to downstream user-specific tasks. This thesis addresses the challenge of adapting DNN models effectively on resource-limited hardware, advocating for flexible and efficient models. Firstly, we tackle the challenge of on-device adaptation for user-specific models. Given the limited on-chip memory in edge devices, data movement becomes a significant bottleneck, leading to increased energy consumption. We introduce Chameleon, an energy-efficient continual learning framework. Secondly, generic DNN models can lead to wasteful and inefficient inference. We propose CRISP, a class-aware pruning framework that employs a hybrid structured sparsity pattern. Thirdly, while low-precision integer formats are common for reduced precision arithmetic, these formats are inadequate for emerging workloads like transformers. We explore minifloats, reduced-precision floating-point formats that significantly reduce the memory footprint while offering greater flexibility due to their dynamic range. Finally, models such as Vision Transformers (ViTs) struggle with real-time inference on edge devices. We introduce SMS, a token pruning framework that dynamically determines which tokens to merge and to what extent, thereby optimizing model accuracy while adhering to the target average inference budget."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["9f3c6f2aa8f96291ef77b0a18484521d","e7bf155a4059222ae86403a58c77814a","135a28864ae70a295050ce2af932fb50"]},{"key":"dc:title","label":"Title","values":["MODEL ADAPTATION FOR EDGE AI"]}]}],"canonical_facts":{"dc:creator":["SHIVAM AGGARWAL"],"dc:date.issued":["2024-11-15"],"dc:description.abstract":["Deployment of Deep Neural Networks (DNNs) on edge devices presents significant challenges due to their computational demands. Existing model compression techniques often fall short by being oblivious to downstream user-specific tasks. This thesis addresses the challenge of adapting DNN models effectively on resource-limited hardware, advocating for flexible and efficient models. Firstly, we tackle the challenge of on-device adaptation for user-specific models. Given the limited on-chip memory in edge devices, data movement becomes a significant bottleneck, leading to increased energy consumption. We introduce Chameleon, an energy-efficient continual learning framework. Secondly, generic DNN models can lead to wasteful and inefficient inference. We propose CRISP, a class-aware pruning framework that employs a hybrid structured sparsity pattern. Thirdly, while low-precision integer formats are common for reduced precision arithmetic, these formats are inadequate for emerging workloads like transformers. We explore minifloats, reduced-precision floating-point formats that significantly reduce the memory footprint while offering greater flexibility due to their dynamic range. Finally, models such as Vision Transformers (ViTs) struggle with real-time inference on edge devices. We introduce SMS, a token pruning framework that dynamically determines which tokens to merge and to what extent, thereby optimizing model accuracy while adhering to the target average inference budget."],"dc:format.checksum.md5":["9f3c6f2aa8f96291ef77b0a18484521d","e7bf155a4059222ae86403a58c77814a","135a28864ae70a295050ce2af932fb50"],"dc:identifier.uri":["https://scholarbank.nus.edu.sg/bitstreams/36a63bb5-6e17-4069-a30f-cd5e4a78f277/download"],"dc:relation.isreferencedby":["https://scholarbank.nus.edu.sg/handle/10635/309814"],"dc:rights":["https://scholarbank.nus.edu.sg/bitstreams/aedc0f0c-f38b-4720-8a7f-e77860536d79/download"],"dc:subject":["Quantization","Sparsity","Model Compression","Efficient AI","EdgeAI"],"dc:title":["MODEL ADAPTATION FOR EDGE AI"],"dc:type":["Thesis"]},"updated_at":"2026-07-24T03:32:18Z"}