{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132558"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132558","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Scalable foundation models","abstract":"The continual growth in computational resources and human annotators, driven by advances in hardware architecture and the increasing accessibility of crowdsourcing platforms, has created unprecedented opportunities for artificial intelligence (AI) models. However, without scalable AI solutions (e.g., training algorithms, model architectures), much of this additional compute and annotation may be underutilized or yield diminishing returns. Furthermore, as real-world applications demand increasingly sophisticated AI capabilities, scalable models offer a clear path to achieving higher levels of intelligence by taking full advantage of available computational and annotation resources. This makes scalability not just a technical consideration, but a fundamental requirement for advancing the field of AI in parallel with hardware developments. This dissertation investigates the fundamental trajectory toward scalable foundation models through three subsequent research milestones. (1) Predictable Scaling: We examine scaling laws that govern the development of foundation models, analyzing how models' capabilities correlate with computational resources. Our research establishes principle solutions to forecasting model behaviors and resource requirements across different scales, enabling scientific and reliable scaling of AI models. (2) Scalable Modeling: We explore model architectures and training recipes optimized for multimodal learning, demonstrating how these approaches can effectively utilize increasing computational resources and data to achieve continuous performance improvements. Our findings reveal architectural principles and training strategies that maintain efficiency at scale while avoiding common bottlenecks in previous modeling strategies. (3) Scalable Oversight: We study the scalable post-training approaches that enable continuous model improvement and alignment with human values even as model capabilities expand beyond human expertise. This research introduces novel techniques for scalable supervision that scale in parallel with model complexity and capability, ensuring the responsible advancement of AI models. In Chapter 1, we describe the key research problems and dive deeply into several key featured research in the following chapters. Predictable Scaling In Chapter 2, we study how to estimate the actual capabilities (i.e., downstream performance) in large language models (LLMs) via addressing the challenges of LLMs' emergent abilities. We focus on the pre-training loss as a more computation-efficient metric for performance estimation. We present FLP, a two-stage approach for performance prediction that consists of first estimating a function that maps computational resources (e.g., FLOPs) to the pre-training Loss using a series of sampling models, followed by mapping the pre-training loss to downstream task Performance after the critical \"emergent phase\". Scalable Modeling In Chapter 3, we present a scalable code-guided visual representation learning method and a single transformer architecture for scalable vision-language modeling. A single unified Transformer architecture can effectively addresses the scalability concerns in previous large vision-language models (LVLMs); however, its limited adoption in modern context likely stems from the absence of reliable training recipes that balance both modalities and ensure stable training for billion-scale models. We introduce the first open-source training recipe for developing unified LVLMs, using moderate academic resources (8 x A100 80GB GPUs). In addition, we revisit the next token prediction loss on vision-language pre-training, and argue that this can be a false proxy of the actual capabilities in LVLMs. We propose a new algorithm, ViStruct, to scale up vision-langugage pre-training. The results show that ViStruct scales better with more data and compute. Scalable Oversight In Chapter 4, we investigate a novel approach to AI supervision through learning from AI feedback. We introduce a scalable alignment framework that harnesses the strong capabilities of large language models (LLMs) to guide the development of LVLMs. Our framework advances beyond conventional numerical reward signals by leveraging natural language feedback as a primary mechanism for model optimization and refinement. This methodology enables the systematic refinement of model responses, promoting attributes of helpfulness, truthfulness, and safety and also enhance their capacity for sustained multi-turn interactions. Our approach demonstrates how advanced LLMs can serve as effective supervisors in the training pipeline, offering a scalable solution to the challenge of model alignment. In addition, we utilize the AI feedback to supervise the reasoning consistency of LVLMs. In our curated benchmark that targets the chain-of-thought (CoT) reasoning performance and consistency of LVLMs, the results show that supervising the reasoning process brings better reasoning capabilities in LVLMs.","abstract_html":"The continual growth in computational resources and human annotators, driven by advances in hardware architecture and the increasing accessibility of crowdsourcing platforms, has created unprecedented opportunities for artificial intelligence (AI) models. However, without scalable AI solutions (e.g., training algorithms, model architectures), much of this additional compute and annotation may be underutilized or yield diminishing returns. Furthermore, as real-world applications demand increasingly sophisticated AI capabilities, scalable models offer a clear path to achieving higher levels of intelligence by taking full advantage of available computational and annotation resources. This makes scalability not just a technical consideration, but a fundamental requirement for advancing the field of AI in parallel with hardware developments. This dissertation investigates the fundamental trajectory toward scalable foundation models through three subsequent research milestones. (1) Predictable Scaling: We examine scaling laws that govern the development of foundation models, analyzing how models&#x27; capabilities correlate with computational resources. Our research establishes principle solutions to forecasting model behaviors and resource requirements across different scales, enabling scientific and reliable scaling of AI models. (2) Scalable Modeling: We explore model architectures and training recipes optimized for multimodal learning, demonstrating how these approaches can effectively utilize increasing computational resources and data to achieve continuous performance improvements. Our findings reveal architectural principles and training strategies that maintain efficiency at scale while avoiding common bottlenecks in previous modeling strategies. (3) Scalable Oversight: We study the scalable post-training approaches that enable continuous model improvement and alignment with human values even as model capabilities expand beyond human expertise. This research introduces novel techniques for scalable supervision that scale in parallel with model complexity and capability, ensuring the responsible advancement of AI models. In Chapter 1, we describe the key research problems and dive deeply into several key featured research in the following chapters. Predictable Scaling In Chapter 2, we study how to estimate the actual capabilities (i.e., downstream performance) in large language models (LLMs) via addressing the challenges of LLMs&#x27; emergent abilities. We focus on the pre-training loss as a more computation-efficient metric for performance estimation. We present FLP, a two-stage approach for performance prediction that consists of first estimating a function that maps computational resources (e.g., FLOPs) to the pre-training Loss using a series of sampling models, followed by mapping the pre-training loss to downstream task Performance after the critical &quot;emergent phase&quot;. Scalable Modeling In Chapter 3, we present a scalable code-guided visual representation learning method and a single transformer architecture for scalable vision-language modeling. A single unified Transformer architecture can effectively addresses the scalability concerns in previous large vision-language models (LVLMs); however, its limited adoption in modern context likely stems from the absence of reliable training recipes that balance both modalities and ensure stable training for billion-scale models. We introduce the first open-source training recipe for developing unified LVLMs, using moderate academic resources (8 x A100 80GB GPUs). In addition, we revisit the next token prediction loss on vision-language pre-training, and argue that this can be a false proxy of the actual capabilities in LVLMs. We propose a new algorithm, ViStruct, to scale up vision-langugage pre-training. The results show that ViStruct scales better with more data and compute. Scalable Oversight In Chapter 4, we investigate a novel approach to AI supervision through learning from AI feedback. We introduce a scalable alignment framework that harnesses the strong capabilities of large language models (LLMs) to guide the development of LVLMs. Our framework advances beyond conventional numerical reward signals by leveraging natural language feedback as a primary mechanism for model optimization and refinement. This methodology enables the systematic refinement of model responses, promoting attributes of helpfulness, truthfulness, and safety and also enhance their capacity for sustained multi-turn interactions. Our approach demonstrates how advanced LLMs can serve as effective supervisors in the training pipeline, offering a scalable solution to the challenge of model alignment. In addition, we utilize the AI feedback to supervise the reasoning consistency of LVLMs. In our curated benchmark that targets the chain-of-thought (CoT) reasoning performance and consistency of LVLMs, the results show that supervising the reasoning process brings better reasoning capabilities in LVLMs.","abstract_has_math":false,"creators":["Chen, Yangyi"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng","Zhang, Tong","Peng, Hao","Yang, Zhengyuan","Ping, Wei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["foundation models, scalable AI, multimodal"],"languages":["en"],"rights":["Copyright 2025 Yangyi Chen"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132558","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng","Zhang, Tong","Peng, Hao","Yang, Zhengyuan","Ping, Wei"]},{"key":"dc:creator","label":"Author","values":["Chen, Yangyi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-12-03"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["foundation models, scalable AI, multimodal"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Yangyi Chen"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132558"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The continual growth in computational resources and human annotators, driven by advances in hardware architecture and the increasing accessibility of crowdsourcing platforms, has created unprecedented opportunities for artificial intelligence (AI) models. However, without scalable AI solutions (e.g., training algorithms, model architectures), much of this additional compute and annotation may be underutilized or yield diminishing returns. Furthermore, as real-world applications demand increasingly sophisticated AI capabilities, scalable models offer a clear path to achieving higher levels of intelligence by taking full advantage of available computational and annotation resources. This makes scalability not just a technical consideration, but a fundamental requirement for advancing the field of AI in parallel with hardware developments. This dissertation investigates the fundamental trajectory toward scalable foundation models through three subsequent research milestones. (1) Predictable Scaling: We examine scaling laws that govern the development of foundation models, analyzing how models' capabilities correlate with computational resources. Our research establishes principle solutions to forecasting model behaviors and resource requirements across different scales, enabling scientific and reliable scaling of AI models. (2) Scalable Modeling: We explore model architectures and training recipes optimized for multimodal learning, demonstrating how these approaches can effectively utilize increasing computational resources and data to achieve continuous performance improvements. Our findings reveal architectural principles and training strategies that maintain efficiency at scale while avoiding common bottlenecks in previous modeling strategies. (3) Scalable Oversight: We study the scalable post-training approaches that enable continuous model improvement and alignment with human values even as model capabilities expand beyond human expertise. This research introduces novel techniques for scalable supervision that scale in parallel with model complexity and capability, ensuring the responsible advancement of AI models. In Chapter 1, we describe the key research problems and dive deeply into several key featured research in the following chapters. Predictable Scaling In Chapter 2, we study how to estimate the actual capabilities (i.e., downstream performance) in large language models (LLMs) via addressing the challenges of LLMs' emergent abilities. We focus on the pre-training loss as a more computation-efficient metric for performance estimation. We present FLP, a two-stage approach for performance prediction that consists of first estimating a function that maps computational resources (e.g., FLOPs) to the pre-training Loss using a series of sampling models, followed by mapping the pre-training loss to downstream task Performance after the critical \"emergent phase\". Scalable Modeling In Chapter 3, we present a scalable code-guided visual representation learning method and a single transformer architecture for scalable vision-language modeling. A single unified Transformer architecture can effectively addresses the scalability concerns in previous large vision-language models (LVLMs); however, its limited adoption in modern context likely stems from the absence of reliable training recipes that balance both modalities and ensure stable training for billion-scale models. We introduce the first open-source training recipe for developing unified LVLMs, using moderate academic resources (8 x A100 80GB GPUs). In addition, we revisit the next token prediction loss on vision-language pre-training, and argue that this can be a false proxy of the actual capabilities in LVLMs. We propose a new algorithm, ViStruct, to scale up vision-langugage pre-training. The results show that ViStruct scales better with more data and compute. Scalable Oversight In Chapter 4, we investigate a novel approach to AI supervision through learning from AI feedback. We introduce a scalable alignment framework that harnesses the strong capabilities of large language models (LLMs) to guide the development of LVLMs. Our framework advances beyond conventional numerical reward signals by leveraging natural language feedback as a primary mechanism for model optimization and refinement. This methodology enables the systematic refinement of model responses, promoting attributes of helpfulness, truthfulness, and safety and also enhance their capacity for sustained multi-turn interactions. Our approach demonstrates how advanced LLMs can serve as effective supervisors in the training pipeline, offering a scalable solution to the challenge of model alignment. In addition, we utilize the AI feedback to supervise the reasoning consistency of LVLMs. In our curated benchmark that targets the chain-of-thought (CoT) reasoning performance and consistency of LVLMs, the results show that supervising the reasoning process brings better reasoning capabilities in LVLMs.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Yangyi Chen, accepted the attached license on 2025-12-02 at 14:33.","The student, Yangyi Chen, submitted this Dissertation for approval on 2025-12-02 at 14:46.","This Dissertation was approved for publication on 2025-12-03 at 10:14.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23022 on 2026-02-19 at 18:26:14"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Scalable foundation models"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng","Zhang, Tong","Peng, Hao","Yang, Zhengyuan","Ping, Wei"],"dc:creator":["Chen, Yangyi"],"dc:date":["2025-12","2025-12-03"],"dc:description":["The continual growth in computational resources and human annotators, driven by advances in hardware architecture and the increasing accessibility of crowdsourcing platforms, has created unprecedented opportunities for artificial intelligence (AI) models. However, without scalable AI solutions (e.g., training algorithms, model architectures), much of this additional compute and annotation may be underutilized or yield diminishing returns. Furthermore, as real-world applications demand increasingly sophisticated AI capabilities, scalable models offer a clear path to achieving higher levels of intelligence by taking full advantage of available computational and annotation resources. This makes scalability not just a technical consideration, but a fundamental requirement for advancing the field of AI in parallel with hardware developments. This dissertation investigates the fundamental trajectory toward scalable foundation models through three subsequent research milestones. (1) Predictable Scaling: We examine scaling laws that govern the development of foundation models, analyzing how models' capabilities correlate with computational resources. Our research establishes principle solutions to forecasting model behaviors and resource requirements across different scales, enabling scientific and reliable scaling of AI models. (2) Scalable Modeling: We explore model architectures and training recipes optimized for multimodal learning, demonstrating how these approaches can effectively utilize increasing computational resources and data to achieve continuous performance improvements. Our findings reveal architectural principles and training strategies that maintain efficiency at scale while avoiding common bottlenecks in previous modeling strategies. (3) Scalable Oversight: We study the scalable post-training approaches that enable continuous model improvement and alignment with human values even as model capabilities expand beyond human expertise. This research introduces novel techniques for scalable supervision that scale in parallel with model complexity and capability, ensuring the responsible advancement of AI models. In Chapter 1, we describe the key research problems and dive deeply into several key featured research in the following chapters. Predictable Scaling In Chapter 2, we study how to estimate the actual capabilities (i.e., downstream performance) in large language models (LLMs) via addressing the challenges of LLMs' emergent abilities. We focus on the pre-training loss as a more computation-efficient metric for performance estimation. We present FLP, a two-stage approach for performance prediction that consists of first estimating a function that maps computational resources (e.g., FLOPs) to the pre-training Loss using a series of sampling models, followed by mapping the pre-training loss to downstream task Performance after the critical \"emergent phase\". Scalable Modeling In Chapter 3, we present a scalable code-guided visual representation learning method and a single transformer architecture for scalable vision-language modeling. A single unified Transformer architecture can effectively addresses the scalability concerns in previous large vision-language models (LVLMs); however, its limited adoption in modern context likely stems from the absence of reliable training recipes that balance both modalities and ensure stable training for billion-scale models. We introduce the first open-source training recipe for developing unified LVLMs, using moderate academic resources (8 x A100 80GB GPUs). In addition, we revisit the next token prediction loss on vision-language pre-training, and argue that this can be a false proxy of the actual capabilities in LVLMs. We propose a new algorithm, ViStruct, to scale up vision-langugage pre-training. The results show that ViStruct scales better with more data and compute. Scalable Oversight In Chapter 4, we investigate a novel approach to AI supervision through learning from AI feedback. We introduce a scalable alignment framework that harnesses the strong capabilities of large language models (LLMs) to guide the development of LVLMs. Our framework advances beyond conventional numerical reward signals by leveraging natural language feedback as a primary mechanism for model optimization and refinement. This methodology enables the systematic refinement of model responses, promoting attributes of helpfulness, truthfulness, and safety and also enhance their capacity for sustained multi-turn interactions. Our approach demonstrates how advanced LLMs can serve as effective supervisors in the training pipeline, offering a scalable solution to the challenge of model alignment. In addition, we utilize the AI feedback to supervise the reasoning consistency of LVLMs. In our curated benchmark that targets the chain-of-thought (CoT) reasoning performance and consistency of LVLMs, the results show that supervising the reasoning process brings better reasoning capabilities in LVLMs.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Yangyi Chen, accepted the attached license on 2025-12-02 at 14:33.","The student, Yangyi Chen, submitted this Dissertation for approval on 2025-12-02 at 14:46.","This Dissertation was approved for publication on 2025-12-03 at 10:14.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23022 on 2026-02-19 at 18:26:14"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132558"],"dc:language":["en"],"dc:rights":["Copyright 2025 Yangyi Chen"],"dc:subject":["foundation models, scalable AI, multimodal"],"dc:title":["Scalable foundation models"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}