University of Illinois at Urbana-Champaign
ConvMLP: Hierarchical convolutional MLPS for vision
Abstract
dc:descriptionIn the past decade, deep learning has made breakthroughs in most computer vision tasks, and most solutions are based on deep convolution neural networks. Recently, MLP-based architectures, which consist of a sequence of consecutive multi-layer perceptron blocks, have been found to achieve results comparable to those of convolutional and transformer-based methods. However, most MLP-based architectures adopt spatial MLPs which take fixed dimension inputs, therefore making it difficult to apply them to downstream computer vision tasks, such as object detection and semantic segmentation. Moreover, single-stage designs further limit performance in other computer vision tasks and fully connected layers bear heavy computation. To tackle these problems, we propose ConvMLP: a hierarchical Convolutional MLP for visual recognition, which is a lightweight, stage-wise, co-design of convolution layers and MLPs. In particular, ConvMLP-S achieves 76.8\% top-1 accuracy on ImageNet-1k with 9M parameters and 2.4 GMACs (15\% and 19\% of MLP-Mixer-B/16, respectively). Experiments on object detection and semantic segmentation further show that visual representation learned by ConvMLP can be seamlessly transferred and achieve competitive results with fewer parameters.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Li, Jiachen
- Contributors dc:contributor
-
- Shi, Humphrey
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2022 Jiachen Li
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/116263