Back to results

National University of Singapore

TOWARDS EFFICIENT TRANSFORMER SCALING

Abstract

dc:description.abstract

Transformer-based models have achieved exceptional performance across various tasks but face resource limitations when scaling. This thesis explores strategies to enhance Transformer efficiency. First, we propose WideNet, which optimizes parameter efficiency using parameter-sharing and Mixture-of-Experts, achieving strong results in both vision and language tasks. Second, we investigate transformer configurations, finding that token-level training benefits from deeper, narrower models, while sequence-level tasks face scaling challenges. For tasks requiring longer input sequences, we introduce sequence parallelism, increasing maximum sequence length by 27 times. To address the need for flexible models with fixed computation budgets, we present AdaTape, enabling adaptive computation with elastic sequences. Lastly, we highlight the importance of dataset scaling, revealing the need for proportional scaling of both model parameters and training tokens to achieve compute-optimal results.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • XUE FUZHAO

Subjects

dc:subject × 2

Rights

dc:rights

Chain of custody

source
Harvested from
National University of Singapore
Base URL
scholarbank.nus.edu.sg/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

XUE FUZHAO. TOWARDS EFFICIENT TRANSFORMER SCALING. 2024.