{"id":{"repo_id":"nus","oai_identifier":"oai:scholarbank.nus.edu.sg:10635/309586"},"canonical_url":"https://search.dev.ndltd.org/etd/nus/oai:scholarbank.nus.edu.sg:10635/309586","repository":{"repo_id":"nus","name":"National University of Singapore","base_url":"https://scholarbank.nus.edu.sg/oai/request"},"display":{"title":"TOWARDS EFFICIENT TRANSFORMER SCALING","abstract":"Transformer-based models have achieved exceptional performance across various tasks but face resource limitations when scaling. This thesis explores strategies to enhance Transformer efficiency. First, we propose WideNet, which optimizes parameter efficiency using parameter-sharing and Mixture-of-Experts, achieving strong results in both vision and language tasks. Second, we investigate transformer configurations, finding that token-level training benefits from deeper, narrower models, while sequence-level tasks face scaling challenges. For tasks requiring longer input sequences, we introduce sequence parallelism, increasing maximum sequence length by 27 times. To address the need for flexible models with fixed computation budgets, we present AdaTape, enabling adaptive computation with elastic sequences. Lastly, we highlight the importance of dataset scaling, revealing the need for proportional scaling of both model parameters and training tokens to achieve compute-optimal results.","abstract_html":"Transformer-based models have achieved exceptional performance across various tasks but face resource limitations when scaling. This thesis explores strategies to enhance Transformer efficiency. First, we propose WideNet, which optimizes parameter efficiency using parameter-sharing and Mixture-of-Experts, achieving strong results in both vision and language tasks. Second, we investigate transformer configurations, finding that token-level training benefits from deeper, narrower models, while sequence-level tasks face scaling challenges. For tasks requiring longer input sequences, we introduce sequence parallelism, increasing maximum sequence length by 27 times. To address the need for flexible models with fixed computation budgets, we present AdaTape, enabling adaptive computation with elastic sequences. Lastly, we highlight the importance of dataset scaling, revealing the need for proportional scaling of both model parameters and training tokens to achieve compute-optimal results.","abstract_has_math":false,"creators":["XUE FUZHAO"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-08-02","date_published":"2024-08-02","updated_at":"2026-07-24T03:33:22Z","subjects":["Machine Learning","Deep Learning"],"languages":[],"rights":[],"rights_urls":["https://scholarbank.nus.edu.sg/bitstreams/40899bfb-172e-4382-ae5a-7f809f5a0186/download"],"identifier_entries":[]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["XUE FUZHAO"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-08-02"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://scholarbank.nus.edu.sg/handle/10635/309586"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","Deep Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://scholarbank.nus.edu.sg/bitstreams/40899bfb-172e-4382-ae5a-7f809f5a0186/download"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarbank.nus.edu.sg/bitstreams/4efd335a-f8c3-4e7a-ad89-ffcbbfb21d22/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Transformer-based models have achieved exceptional performance across various tasks but face resource limitations when scaling. This thesis explores strategies to enhance Transformer efficiency. First, we propose WideNet, which optimizes parameter efficiency using parameter-sharing and Mixture-of-Experts, achieving strong results in both vision and language tasks. Second, we investigate transformer configurations, finding that token-level training benefits from deeper, narrower models, while sequence-level tasks face scaling challenges. For tasks requiring longer input sequences, we introduce sequence parallelism, increasing maximum sequence length by 27 times. To address the need for flexible models with fixed computation budgets, we present AdaTape, enabling adaptive computation with elastic sequences. Lastly, we highlight the importance of dataset scaling, revealing the need for proportional scaling of both model parameters and training tokens to achieve compute-optimal results."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["9f3c6f2aa8f96291ef77b0a18484521d","e33bbf8658a79e525137aede99369ef7","771af676393ef72aabb149f55ed64406"]},{"key":"dc:title","label":"Title","values":["TOWARDS EFFICIENT TRANSFORMER SCALING"]}]}],"canonical_facts":{"dc:creator":["XUE FUZHAO"],"dc:date.issued":["2024-08-02"],"dc:description.abstract":["Transformer-based models have achieved exceptional performance across various tasks but face resource limitations when scaling. This thesis explores strategies to enhance Transformer efficiency. First, we propose WideNet, which optimizes parameter efficiency using parameter-sharing and Mixture-of-Experts, achieving strong results in both vision and language tasks. Second, we investigate transformer configurations, finding that token-level training benefits from deeper, narrower models, while sequence-level tasks face scaling challenges. For tasks requiring longer input sequences, we introduce sequence parallelism, increasing maximum sequence length by 27 times. To address the need for flexible models with fixed computation budgets, we present AdaTape, enabling adaptive computation with elastic sequences. Lastly, we highlight the importance of dataset scaling, revealing the need for proportional scaling of both model parameters and training tokens to achieve compute-optimal results."],"dc:format.checksum.md5":["9f3c6f2aa8f96291ef77b0a18484521d","e33bbf8658a79e525137aede99369ef7","771af676393ef72aabb149f55ed64406"],"dc:identifier.uri":["https://scholarbank.nus.edu.sg/bitstreams/4efd335a-f8c3-4e7a-ad89-ffcbbfb21d22/download"],"dc:relation.isreferencedby":["https://scholarbank.nus.edu.sg/handle/10635/309586"],"dc:rights":["https://scholarbank.nus.edu.sg/bitstreams/40899bfb-172e-4382-ae5a-7f809f5a0186/download"],"dc:subject":["Machine Learning","Deep Learning"],"dc:title":["TOWARDS EFFICIENT TRANSFORMER SCALING"],"dc:type":["Thesis"]},"updated_at":"2026-07-24T03:33:22Z"}