{"id":{"repo_id":"auckland-tech","oai_identifier":"oai:openrepository.aut.ac.nz:10292/19242"},"canonical_url":"https://search.dev.ndltd.org/etd/auckland-tech/oai:openrepository.aut.ac.nz:10292/19242","repository":{"repo_id":"auckland-tech","name":"AUT University","base_url":"https://openrepository.aut.ac.nz/server/oai/request"},"display":{"title":"ChatPPG: Multi-Modal Alignment of Large Language Models for Time-Series Forecasting in Table Tennis","abstract":"In this thesis, we explore the adaptation of large language models (LLMs) for structured time-series forecasting, focusing on predicting table tennis serve landing points. Traditional time-series models rely on specialized architectures, while LLMs are inherently designed for textual data processing, posing challenges in numerical sequence modeling. To address this, we introduce ChatPPG, a multi-modal framework that integrates time-series data into LLMs through structured embeddings, cross-modal attention, and parameter-efficient fine-tuning (i.e., LoRA). Our findings demonstrate that alignment-based approaches significantly enhance forecasting accuracy compared to prompting-based methods, with DeepSeek-R1-Distill-Qwen-1.5B achieving the lowest MSE (0.432) and MAE (0.441). However, our study also highlights a trade-off between accuracy and inference efficiency, as prompting-based methods introduce excessive latency, making them impractical for real-time applications. Ablation experiments further validate the importance of multi-modal feature alignment, interleaved embedding fusion (IEF), and domain-informed prompting, showing that their removal leads to substantial performance degradation. In this thesis, we extend the application of foundation models beyond natural language processing, establishing a scalable and computationally efficient framework for integrating LLMs into structured forecasting tasks. Our future research directions include the development of a fully end-to-end multi-modal sports analytics system, leveraging real-time vision models for spatiotemporal reasoning, as well as the exploration of generative models like stable diffusion for stochastic time-series forecasting. These advancements aim to enhance automated match analysis and intelligent coaching applications, further bridging AI, computer vision, and predictive modeling in sports analytics.","abstract_html":"In this thesis, we explore the adaptation of large language models (LLMs) for structured time-series forecasting, focusing on predicting table tennis serve landing points. Traditional time-series models rely on specialized architectures, while LLMs are inherently designed for textual data processing, posing challenges in numerical sequence modeling. To address this, we introduce ChatPPG, a multi-modal framework that integrates time-series data into LLMs through structured embeddings, cross-modal attention, and parameter-efficient fine-tuning (i.e., LoRA). Our findings demonstrate that alignment-based approaches significantly enhance forecasting accuracy compared to prompting-based methods, with DeepSeek-R1-Distill-Qwen-1.5B achieving the lowest MSE (0.432) and MAE (0.441). However, our study also highlights a trade-off between accuracy and inference efficiency, as prompting-based methods introduce excessive latency, making them impractical for real-time applications. Ablation experiments further validate the importance of multi-modal feature alignment, interleaved embedding fusion (IEF), and domain-informed prompting, showing that their removal leads to substantial performance degradation. In this thesis, we extend the application of foundation models beyond natural language processing, establishing a scalable and computationally efficient framework for integrating LLMs into structured forecasting tasks. Our future research directions include the development of a fully end-to-end multi-modal sports analytics system, leveraging real-time vision models for spatiotemporal reasoning, as well as the exploration of generative models like stable diffusion for stochastic time-series forecasting. These advancements aim to enhance automated match analysis and intelligent coaching applications, further bridging AI, computer vision, and predictive modeling in sports analytics.","abstract_has_math":false,"creators":["Yang, GuangLiang"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-27T18:45:59Z","subjects":[],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["hdl:10292/19242"],"render_values":[{"text":"hdl:10292/19242","href":null,"code":true}]}]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["hdl:10292/19242"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.other","label":"Dc Description Other","values":["In this thesis, we explore the adaptation of large language models (LLMs) for structured time-series forecasting, focusing on predicting table tennis serve landing points. Traditional time-series models rely on specialized architectures, while LLMs are inherently designed for textual data processing, posing challenges in numerical sequence modeling. To address this, we introduce ChatPPG, a multi-modal framework that integrates time-series data into LLMs through structured embeddings, cross-modal attention, and parameter-efficient fine-tuning (i.e., LoRA). Our findings demonstrate that alignment-based approaches significantly enhance forecasting accuracy compared to prompting-based methods, with DeepSeek-R1-Distill-Qwen-1.5B achieving the lowest MSE (0.432) and MAE (0.441). However, our study also highlights a trade-off between accuracy and inference efficiency, as prompting-based methods introduce excessive latency, making them impractical for real-time applications. Ablation experiments further validate the importance of multi-modal feature alignment, interleaved embedding fusion (IEF), and domain-informed prompting, showing that their removal leads to substantial performance degradation. In this thesis, we extend the application of foundation models beyond natural language processing, establishing a scalable and computationally efficient framework for integrating LLMs into structured forecasting tasks. Our future research directions include the development of a fully end-to-end multi-modal sports analytics system, leveraging real-time vision models for spatiotemporal reasoning, as well as the exploration of generative models like stable diffusion for stochastic time-series forecasting. These advancements aim to enhance automated match analysis and intelligent coaching applications, further bridging AI, computer vision, and predictive modeling in sports analytics."]},{"key":"dc:title","label":"Title","values":["ChatPPG: Multi-Modal Alignment of Large Language Models for Time-Series Forecasting in Table Tennis"]}]}],"canonical_facts":{"dc:date.issued":["2025"],"dc:description.other":["In this thesis, we explore the adaptation of large language models (LLMs) for structured time-series forecasting, focusing on predicting table tennis serve landing points. Traditional time-series models rely on specialized architectures, while LLMs are inherently designed for textual data processing, posing challenges in numerical sequence modeling. To address this, we introduce ChatPPG, a multi-modal framework that integrates time-series data into LLMs through structured embeddings, cross-modal attention, and parameter-efficient fine-tuning (i.e., LoRA). Our findings demonstrate that alignment-based approaches significantly enhance forecasting accuracy compared to prompting-based methods, with DeepSeek-R1-Distill-Qwen-1.5B achieving the lowest MSE (0.432) and MAE (0.441). However, our study also highlights a trade-off between accuracy and inference efficiency, as prompting-based methods introduce excessive latency, making them impractical for real-time applications. Ablation experiments further validate the importance of multi-modal feature alignment, interleaved embedding fusion (IEF), and domain-informed prompting, showing that their removal leads to substantial performance degradation. In this thesis, we extend the application of foundation models beyond natural language processing, establishing a scalable and computationally efficient framework for integrating LLMs into structured forecasting tasks. Our future research directions include the development of a fully end-to-end multi-modal sports analytics system, leveraging real-time vision models for spatiotemporal reasoning, as well as the exploration of generative models like stable diffusion for stochastic time-series forecasting. These advancements aim to enhance automated match analysis and intelligent coaching applications, further bridging AI, computer vision, and predictive modeling in sports analytics."],"dc:identifier":["hdl:10292/19242"],"dc:title":["ChatPPG: Multi-Modal Alignment of Large Language Models for Time-Series Forecasting in Table Tennis"],"dc:type":["Thesis"]},"updated_at":"2026-07-27T18:45:59Z"}