National University of Singapore
CLOSED-LOOP SCALING: AUTONOMOUS IMPROVEMENT OF LLM AND LVLM REASONING
Abstract
dc:description.abstractAs human-curated data approaches exhaustion, sustaining the improvement of large language models (LLMs) and large vision--language models (LVLMs) demands a paradigm shift. This thesis proposes automatic scaling: a closed-loop framework in which models autonomously improve through their own computation via three layers. Inference-time scaling treats reasoning as search guided by self-evaluation. Training-time scaling internalizes search-discovered knowledge into parameters through iterative preference alignment. Architectural grounding provides structural foundations for sustainable scaling. Through critical analysis, we identify the coherence--correctness gap as a systemic limitation of self-referential scaling and present MVP-Bench, a diagnostic benchmark revealing significant deficits in multi-level visual perception. We propose future directions in dynamic evaluation, agentic scaling, pre-linguistic reasoning foundations, and native multimodal scaling, delineating both the promise and boundaries of autonomous model improvement.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- XIE YUXI