Abstract
dc:description.abstractNeural architecture design is crucial in AI development. This thesis first examines the macroscopic architecture of Transformers, challenging the belief that their attention-based token mixer is key. By replacing the attention module with a simple spatial pooling operator, we create PoolFormer, which achieves competitive performance across vision tasks. This supports our hypothesis that the general Transformer architecture, dubbed MetaFormer, is essential for model performance. MetaFormer ensures consistent performance and works well with various token mixers, often yielding state-of-the-art results. Additionally, we address efficiency in neural architecture by proposing InceptionNeXt, which decomposes large-kernel depthwise convolutions into smaller components, inspired by Inceptions. InceptionNeXt improves throughput and maintains performance, offering a more efficient baseline for future architecture design.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- YU WEIHAO