{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/137644"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/137644","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Conversational Multimodal LLMs for Food Nutritional Information Retrieval: A Systematic Evaluation","abstract":"Accurate dietary monitoring underpins public health initiatives, chronic disease management, and personalized nutrition, yet manual food logging remains laborious and prone to error. At the same time, vision and language models have reached new levels of capability and offer the potential to infer nutritional content directly from meal photographs with minimal user effort. This thesis systematically evaluates eight state-of-the-art multimodal large language models (including GPT-4o, Qwen 2, and Qwen 2.5, DeepSeek, and LLaVA), in a zero-shot setting on two real-world datasets, Nutrition5k and MetaFood3D. To identify which types of information most improve model predictions, a structured cue ladder protocol was developed. Models first receive only the image, then incrementally gain access to a verified ingredient list, total dish mass, and individual ingredient mass or volume. Single-view and two-view inputs are compared, and two conversational prompting workflows are tested: a predefined sequence of reasoning questions and an agentic self-questioning pipeline that generates clarifying queries on the fly. Results show that GPT-4o achieves the lowest overall mean absolute percentage error (MAPE), with calorie prediction errors improving from approximately 51% in the image-only setting to as low as 29% when provided with total per-ingredient-wise mass. Providing just the total mass reduces calorie prediction error across all models, making it the single most impactful cue. Adding a second image view yields smaller but consistent improvements, in the range of 3-7 percentage points. Finally, the agentic self-questioning workflow consistently outperforms the fixed prompt sequence, particularly in low-context scenarios, with some models showing improvements of over 8 percentage points. Even when granted nearly all available cues, no model attains perfect accuracy, as errors persist in estimating invisible elements such as oils and dressings. These findings clarify the trade-off between user effort and predictive accuracy, and they establish a general evaluation framework and conversational pipeline design that can extend to other applications requiring integrated visual and contextual reasoning. The ultimate aim is to enable an end-user application that delivers reliable nutritional estimates from a simple photograph with minimal additional input.","abstract_html":"Accurate dietary monitoring underpins public health initiatives, chronic disease management, and personalized nutrition, yet manual food logging remains laborious and prone to error. At the same time, vision and language models have reached new levels of capability and offer the potential to infer nutritional content directly from meal photographs with minimal user effort. This thesis systematically evaluates eight state-of-the-art multimodal large language models (including GPT-4o, Qwen 2, and Qwen 2.5, DeepSeek, and LLaVA), in a zero-shot setting on two real-world datasets, Nutrition5k and MetaFood3D. To identify which types of information most improve model predictions, a structured cue ladder protocol was developed. Models first receive only the image, then incrementally gain access to a verified ingredient list, total dish mass, and individual ingredient mass or volume. Single-view and two-view inputs are compared, and two conversational prompting workflows are tested: a predefined sequence of reasoning questions and an agentic self-questioning pipeline that generates clarifying queries on the fly. Results show that GPT-4o achieves the lowest overall mean absolute percentage error (MAPE), with calorie prediction errors improving from approximately 51% in the image-only setting to as low as 29% when provided with total per-ingredient-wise mass. Providing just the total mass reduces calorie prediction error across all models, making it the single most impactful cue. Adding a second image view yields smaller but consistent improvements, in the range of 3-7 percentage points. Finally, the agentic self-questioning workflow consistently outperforms the fixed prompt sequence, particularly in low-context scenarios, with some models showing improvements of over 8 percentage points. Even when granted nearly all available cues, no model attains perfect accuracy, as errors persist in estimating invisible elements such as oils and dressings. These findings clarify the trade-off between user effort and predictive accuracy, and they establish a general evaluation framework and conversational pipeline design that can extend to other applications requiring integrated visual and contextual reasoning. The ultimate aim is to enable an end-user application that delivers reliable nutritional estimates from a simple photograph with minimal additional input.","abstract_has_math":false,"creators":["Bhatambarekar, Gayatri Milind"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Computer Science & Applications","degree_department":"Computer Science and Applications","school":null,"contributors":[],"advisors":[],"committee_chairs":["Meng, Na","Sarkar, Abhijit"],"committee_members":["Brown, Dwayne Christian"],"year":2025,"date_issued":"2025-09-08","date_published":"2025-09-08","updated_at":"2026-07-22T22:20:05Z","subjects":["Large Vision-Language Models","Multimodal Retrieval","Calorie Estimation","Nutrition Analysis","Prompt Engineering"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44582"],"render_values":[{"text":"vt_gsexam:44582","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/137644","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Meng, Na","Sarkar, Abhijit"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Brown, Dwayne Christian"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science and Applications"]},{"key":"dc:creator","label":"Author","values":["Bhatambarekar, Gayatri Milind"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-09-09T08:01:05Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-09-09T08:01:05Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-09-08"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science & Applications"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Large Vision-Language Models","Multimodal Retrieval","Calorie Estimation","Nutrition Analysis","Prompt Engineering"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44582"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/137644"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Accurate dietary monitoring underpins public health initiatives, chronic disease management, and personalized nutrition, yet manual food logging remains laborious and prone to error. At the same time, vision and language models have reached new levels of capability and offer the potential to infer nutritional content directly from meal photographs with minimal user effort. This thesis systematically evaluates eight state-of-the-art multimodal large language models (including GPT-4o, Qwen 2, and Qwen 2.5, DeepSeek, and LLaVA), in a zero-shot setting on two real-world datasets, Nutrition5k and MetaFood3D. To identify which types of information most improve model predictions, a structured cue ladder protocol was developed. Models first receive only the image, then incrementally gain access to a verified ingredient list, total dish mass, and individual ingredient mass or volume. Single-view and two-view inputs are compared, and two conversational prompting workflows are tested: a predefined sequence of reasoning questions and an agentic self-questioning pipeline that generates clarifying queries on the fly. Results show that GPT-4o achieves the lowest overall mean absolute percentage error (MAPE), with calorie prediction errors improving from approximately 51% in the image-only setting to as low as 29% when provided with total per-ingredient-wise mass. Providing just the total mass reduces calorie prediction error across all models, making it the single most impactful cue. Adding a second image view yields smaller but consistent improvements, in the range of 3-7 percentage points. Finally, the agentic self-questioning workflow consistently outperforms the fixed prompt sequence, particularly in low-context scenarios, with some models showing improvements of over 8 percentage points. Even when granted nearly all available cues, no model attains perfect accuracy, as errors persist in estimating invisible elements such as oils and dressings. These findings clarify the trade-off between user effort and predictive accuracy, and they establish a general evaluation framework and conversational pipeline design that can extend to other applications requiring integrated visual and contextual reasoning. The ultimate aim is to enable an end-user application that delivers reliable nutritional estimates from a simple photograph with minimal additional input."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Keeping track of calories and nutrients is essential for health, but entering every detail manually is time-consuming and often inaccurate. Advances in artificial intelligence now make it possible to estimate nutritional content directly from images. This study tests eight leading AI systems, including OpenAI's GPT-4o, that understand both pictures and text. The evaluation uses a step-by-step approach. First, the AI sees only a photo of a meal. Next, it receives a list of ingredients. Then it is told the total weight of the dish. Finally, it gains the weight or volume of each ingredient. We also compare one view versus two views of the meal and two styles of conversation, one following fixed nutrition questions and one where the AI asks its own clarifying questions. Findings show that adding the total dish weight cuts the estimation error. Two views of the meal provide further modest gains. Allowing the AI to generate its own clarifying questions generally improves accuracy over a fixed question script. Even with almost all information provided, the models still struggle with hidden calories in oils or dressings. These insights reveal which pieces of information matter most and how little effort users truly need to log meals. The ultimate goal is a nutrition logging app that asks for a single photograph and perhaps one simple detail, then returns a reliable breakdown of calories and nutrients, making healthy choices easier without the burden of manual entry."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Conversational Multimodal LLMs for Food Nutritional Information Retrieval: A Systematic Evaluation"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Meng, Na","Sarkar, Abhijit"],"dc:contributor.committeemember":["Brown, Dwayne Christian"],"dc:contributor.department":["Computer Science and Applications"],"dc:creator":["Bhatambarekar, Gayatri Milind"],"dc:date.accessioned":["2025-09-09T08:01:05Z"],"dc:date.available":["2025-09-09T08:01:05Z"],"dc:date.issued":["2025-09-08"],"dc:description.abstract":["Accurate dietary monitoring underpins public health initiatives, chronic disease management, and personalized nutrition, yet manual food logging remains laborious and prone to error. At the same time, vision and language models have reached new levels of capability and offer the potential to infer nutritional content directly from meal photographs with minimal user effort. This thesis systematically evaluates eight state-of-the-art multimodal large language models (including GPT-4o, Qwen 2, and Qwen 2.5, DeepSeek, and LLaVA), in a zero-shot setting on two real-world datasets, Nutrition5k and MetaFood3D. To identify which types of information most improve model predictions, a structured cue ladder protocol was developed. Models first receive only the image, then incrementally gain access to a verified ingredient list, total dish mass, and individual ingredient mass or volume. Single-view and two-view inputs are compared, and two conversational prompting workflows are tested: a predefined sequence of reasoning questions and an agentic self-questioning pipeline that generates clarifying queries on the fly. Results show that GPT-4o achieves the lowest overall mean absolute percentage error (MAPE), with calorie prediction errors improving from approximately 51% in the image-only setting to as low as 29% when provided with total per-ingredient-wise mass. Providing just the total mass reduces calorie prediction error across all models, making it the single most impactful cue. Adding a second image view yields smaller but consistent improvements, in the range of 3-7 percentage points. Finally, the agentic self-questioning workflow consistently outperforms the fixed prompt sequence, particularly in low-context scenarios, with some models showing improvements of over 8 percentage points. Even when granted nearly all available cues, no model attains perfect accuracy, as errors persist in estimating invisible elements such as oils and dressings. These findings clarify the trade-off between user effort and predictive accuracy, and they establish a general evaluation framework and conversational pipeline design that can extend to other applications requiring integrated visual and contextual reasoning. The ultimate aim is to enable an end-user application that delivers reliable nutritional estimates from a simple photograph with minimal additional input."],"dc:description.abstractgeneral":["Keeping track of calories and nutrients is essential for health, but entering every detail manually is time-consuming and often inaccurate. Advances in artificial intelligence now make it possible to estimate nutritional content directly from images. This study tests eight leading AI systems, including OpenAI's GPT-4o, that understand both pictures and text. The evaluation uses a step-by-step approach. First, the AI sees only a photo of a meal. Next, it receives a list of ingredients. Then it is told the total weight of the dish. Finally, it gains the weight or volume of each ingredient. We also compare one view versus two views of the meal and two styles of conversation, one following fixed nutrition questions and one where the AI asks its own clarifying questions. Findings show that adding the total dish weight cuts the estimation error. Two views of the meal provide further modest gains. Allowing the AI to generate its own clarifying questions generally improves accuracy over a fixed question script. Even with almost all information provided, the models still struggle with hidden calories in oils or dressings. These insights reveal which pieces of information matter most and how little effort users truly need to log meals. The ultimate goal is a nutrition logging app that asks for a single photograph and perhaps one simple detail, then returns a reliable breakdown of calories and nutrients, making healthy choices easier without the burden of manual entry."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:44582"],"dc:identifier.uri":["https://hdl.handle.net/10919/137644"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Large Vision-Language Models","Multimodal Retrieval","Calorie Estimation","Nutrition Analysis","Prompt Engineering"],"dc:title":["Conversational Multimodal LLMs for Food Nutritional Information Retrieval: A Systematic Evaluation"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science & Applications"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:20:05Z"}