Abstract
dc:description.abstractTransformer-based models have achieved exceptional performance across various tasks but face resource limitations when scaling. This thesis explores strategies to enhance Transformer efficiency. First, we propose WideNet, which optimizes parameter efficiency using parameter-sharing and Mixture-of-Experts, achieving strong results in both vision and language tasks. Second, we investigate transformer configurations, finding that token-level training benefits from deeper, narrower models, while sequence-level tasks face scaling challenges. For tasks requiring longer input sequences, we introduce sequence parallelism, increasing maximum sequence length by 27 times. To address the need for flexible models with fixed computation budgets, we present AdaTape, enabling adaptive computation with elastic sequences. Lastly, we highlight the importance of dataset scaling, revealing the need for proportional scaling of both model parameters and training tokens to achieve compute-optimal results.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- XUE FUZHAO