{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/116263"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/116263","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"ConvMLP: Hierarchical convolutional MLPS for vision","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_has_math":false,"creators":["Li, Jiachen"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Shi, Humphrey"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08","date_published":"2022-08","updated_at":"2026-07-22T22:24:56Z","subjects":["MLP","Vision"],"languages":["en","eng"],"rights":["Copyright 2022 Jiachen Li"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/116263","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Shi, Humphrey"]},{"key":"dc:creator","label":"Author","values":["Li, Jiachen"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-08","2022-07-18"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["MLP","Vision"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Jiachen Li"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/116263"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Jiachen Li, accepted the attached license on 2022-07-18 at 10:49.","The student, Jiachen Li, submitted this Thesis for approval on 2022-07-18 at 10:56.","This Thesis was approved for publication on 2022-07-18 at 15:34.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18362 on 2022-11-15 at 18:21:24","In the past decade, deep learning has made breakthroughs in most computer vision tasks, and most solutions are based on deep convolution neural networks. Recently, MLP-based architectures, which consist of a sequence of consecutive multi-layer perceptron blocks, have been found to achieve results comparable to those of convolutional and transformer-based methods. However, most MLP-based architectures adopt spatial MLPs which take fixed dimension inputs, therefore making it difficult to apply them to downstream computer vision tasks, such as object detection and semantic segmentation. Moreover, single-stage designs further limit performance in other computer vision tasks and fully connected layers bear heavy computation. To tackle these problems, we propose ConvMLP: a hierarchical Convolutional MLP for visual recognition, which is a lightweight, stage-wise, co-design of convolution layers and MLPs. In particular, ConvMLP-S achieves 76.8\\% top-1 accuracy on ImageNet-1k with 9M parameters and 2.4 GMACs (15\\% and 19\\% of MLP-Mixer-B/16, respectively). Experiments on object detection and semantic segmentation further show that visual representation learned by ConvMLP can be seamlessly transferred and achieve competitive results with fewer parameters."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["ConvMLP: Hierarchical convolutional MLPS for vision"]}]}],"canonical_facts":{"dc:contributor":["Shi, Humphrey"],"dc:creator":["Li, Jiachen"],"dc:date":["2022-08","2022-07-18"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Jiachen Li, accepted the attached license on 2022-07-18 at 10:49.","The student, Jiachen Li, submitted this Thesis for approval on 2022-07-18 at 10:56.","This Thesis was approved for publication on 2022-07-18 at 15:34.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18362 on 2022-11-15 at 18:21:24","In the past decade, deep learning has made breakthroughs in most computer vision tasks, and most solutions are based on deep convolution neural networks. Recently, MLP-based architectures, which consist of a sequence of consecutive multi-layer perceptron blocks, have been found to achieve results comparable to those of convolutional and transformer-based methods. However, most MLP-based architectures adopt spatial MLPs which take fixed dimension inputs, therefore making it difficult to apply them to downstream computer vision tasks, such as object detection and semantic segmentation. Moreover, single-stage designs further limit performance in other computer vision tasks and fully connected layers bear heavy computation. To tackle these problems, we propose ConvMLP: a hierarchical Convolutional MLP for visual recognition, which is a lightweight, stage-wise, co-design of convolution layers and MLPs. In particular, ConvMLP-S achieves 76.8\\% top-1 accuracy on ImageNet-1k with 9M parameters and 2.4 GMACs (15\\% and 19\\% of MLP-Mixer-B/16, respectively). Experiments on object detection and semantic segmentation further show that visual representation learned by ConvMLP can be seamlessly transferred and achieve competitive results with fewer parameters."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/116263"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Jiachen Li"],"dc:subject":["MLP","Vision"],"dc:title":["ConvMLP: Hierarchical convolutional MLPS for vision"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:56Z"}