{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/106129"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/106129","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Image captioning using compositional sentiments","abstract":"This thesis presents a method to generate emotional captions of images. An adequate caption should precisely describe the contents in an image. While humans can readily identify the most emotionally salient aspects of an image, many captioning models have difficulties in detecting and generating these non-factual aspects. This is caused by lack of sentiment information in the caption dataset. We solve this issue by preprocessing the text captions in an image captioning dataset with a sentiment analyzer to determine sentiment scores of all images in the training dataset. The model trained from this dataset is able to generate captions that communicate sentiment effectively, without requiring human judges to label sentiment of the training images. The model learns contents of training images, along with embedded word and sentence sentiments. Compared with the model without sentiment, it has better text captioning performance on BLEU-2, which improved from 17.15 to 18.25, and on CIDEr, which improved from 45.21 to 45.68. Automatic sentiment classification of generated captions matches the target sentiment as specified to the captioning system, with accuracy reaching 77.30%, 66.25%, 27.05 on negative, neutral and positive sentiments respectively.","abstract_html":"This thesis presents a method to generate emotional captions of images. An adequate caption should precisely describe the contents in an image. While humans can readily identify the most emotionally salient aspects of an image, many captioning models have difficulties in detecting and generating these non-factual aspects. This is caused by lack of sentiment information in the caption dataset. We solve this issue by preprocessing the text captions in an image captioning dataset with a sentiment analyzer to determine sentiment scores of all images in the training dataset. The model trained from this dataset is able to generate captions that communicate sentiment effectively, without requiring human judges to label sentiment of the training images. The model learns contents of training images, along with embedded word and sentence sentiments. Compared with the model without sentiment, it has better text captioning performance on BLEU-2, which improved from 17.15 to 18.25, and on CIDEr, which improved from 45.21 to 45.68. Automatic sentiment classification of generated captions matches the target sentiment as specified to the captioning system, with accuracy reaching 77.30%, 66.25%, 27.05 on negative, neutral and positive sentiments respectively.","abstract_has_math":false,"creators":["Yang, Yi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-03-02T21:57:50Z","date_published":"2020-03-02T21:57:50Z","updated_at":"2026-07-22T22:24:45Z","subjects":["Image captioning: sentiment"],"languages":["en"],"rights":["Copyright 2019 Yi Yang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/106129","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark"]},{"key":"dc:creator","label":"Author","values":["Yang, Yi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-03-02T21:57:50Z","2019-12-12","2019-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Image captioning: sentiment"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Yi Yang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/106129"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This thesis presents a method to generate emotional captions of images. An adequate caption should precisely describe the contents in an image. While humans can readily identify the most emotionally salient aspects of an image, many captioning models have difficulties in detecting and generating these non-factual aspects. This is caused by lack of sentiment information in the caption dataset. We solve this issue by preprocessing the text captions in an image captioning dataset with a sentiment analyzer to determine sentiment scores of all images in the training dataset. The model trained from this dataset is able to generate captions that communicate sentiment effectively, without requiring human judges to label sentiment of the training images. The model learns contents of training images, along with embedded word and sentence sentiments. Compared with the model without sentiment, it has better text captioning performance on BLEU-2, which improved from 17.15 to 18.25, and on CIDEr, which improved from 45.21 to 45.68. Automatic sentiment classification of generated captions matches the target sentiment as specified to the captioning system, with accuracy reaching 77.30%, 66.25%, 27.05 on negative, neutral and positive sentiments respectively.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-02-28 without embargo terms","The student, Yi Yang, accepted the attached license on 2019-12-11 at 19:26.","The student, Yi Yang, submitted this Thesis for approval on 2019-12-11 at 19:35.","This Thesis was approved for publication on 2019-12-12 at 07:55.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13968 on 2020-02-28 at 17:10:40","Made available in DSpace on 2020-03-02T21:57:50Z (GMT). No. of bitstreams: 2 YANG-THESIS-2019.pdf: 553915 bytes, checksum: 0fbac703cb5ac8de16733c35679e7661 (MD5) LICENSE.txt: 4204 bytes, checksum: a49b3a1ca5e5a400a65aa6c954e8ef6d (MD5) Previous issue date: 2019-12-12"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Image captioning using compositional sentiments"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark"],"dc:creator":["Yang, Yi"],"dc:date":["2020-03-02T21:57:50Z","2019-12-12","2019-12"],"dc:description":["This thesis presents a method to generate emotional captions of images. An adequate caption should precisely describe the contents in an image. While humans can readily identify the most emotionally salient aspects of an image, many captioning models have difficulties in detecting and generating these non-factual aspects. This is caused by lack of sentiment information in the caption dataset. We solve this issue by preprocessing the text captions in an image captioning dataset with a sentiment analyzer to determine sentiment scores of all images in the training dataset. The model trained from this dataset is able to generate captions that communicate sentiment effectively, without requiring human judges to label sentiment of the training images. The model learns contents of training images, along with embedded word and sentence sentiments. Compared with the model without sentiment, it has better text captioning performance on BLEU-2, which improved from 17.15 to 18.25, and on CIDEr, which improved from 45.21 to 45.68. Automatic sentiment classification of generated captions matches the target sentiment as specified to the captioning system, with accuracy reaching 77.30%, 66.25%, 27.05 on negative, neutral and positive sentiments respectively.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-02-28 without embargo terms","The student, Yi Yang, accepted the attached license on 2019-12-11 at 19:26.","The student, Yi Yang, submitted this Thesis for approval on 2019-12-11 at 19:35.","This Thesis was approved for publication on 2019-12-12 at 07:55.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13968 on 2020-02-28 at 17:10:40","Made available in DSpace on 2020-03-02T21:57:50Z (GMT). No. of bitstreams: 2 YANG-THESIS-2019.pdf: 553915 bytes, checksum: 0fbac703cb5ac8de16733c35679e7661 (MD5) LICENSE.txt: 4204 bytes, checksum: a49b3a1ca5e5a400a65aa6c954e8ef6d (MD5) Previous issue date: 2019-12-12"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/106129"],"dc:language":["en"],"dc:rights":["Copyright 2019 Yi Yang"],"dc:subject":["Image captioning: sentiment"],"dc:title":["Image captioning using compositional sentiments"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:45Z"}