{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/110569"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/110569","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"A deeper look into multi-task learning ability of unified text-to-text transformer","abstract":"Structure prediction (SP) tasks are important in natural language understanding in the sense that they provide complex and structured knowledge of the text. Recently, some unified text-to-text transformer models like T5 and TANL have produced competitive results on SP tasks. These models convert SP tasks into a seq2seq problem, where a transformer is used to generate sequences with special tokens representing the extracted spans, labels, and relationships. Compared to many popular Natural Language Understanding models that are designed specifically for the task, the output of the text-to-text transformer is more flexible. With proper format, it could be trained on multiple tasks together and take advantage of the shared knowledge between tasks. To better understand how these models achieve better performance by multi-task learning, we designed several experiments to measure the knowledge transfer ability of a recently proposed model, TANL. In our experiments, we found that the multi-head attention in the decoder can capture the relationship between tasks which leads to performance improvement. Another finding is that TANL may produce many outputs with invalid format when trained from scratch, and starting from a T5 pre-trained model helps to mitigate this problem. Based on these observations and some new intuitions, we proposed an improved version of TANL called SDCT5 (step decomposed and constrained text-to-text Transformer). Preliminary experiment results show that our model can achieve better performance on SP tasks compared to TANL and benefit more from multi-task learning.","abstract_html":"Structure prediction (SP) tasks are important in natural language understanding in the sense that they provide complex and structured knowledge of the text. Recently, some unified text-to-text transformer models like T5 and TANL have produced competitive results on SP tasks. These models convert SP tasks into a seq2seq problem, where a transformer is used to generate sequences with special tokens representing the extracted spans, labels, and relationships. Compared to many popular Natural Language Understanding models that are designed specifically for the task, the output of the text-to-text transformer is more flexible. With proper format, it could be trained on multiple tasks together and take advantage of the shared knowledge between tasks. To better understand how these models achieve better performance by multi-task learning, we designed several experiments to measure the knowledge transfer ability of a recently proposed model, TANL. In our experiments, we found that the multi-head attention in the decoder can capture the relationship between tasks which leads to performance improvement. Another finding is that TANL may produce many outputs with invalid format when trained from scratch, and starting from a T5 pre-trained model helps to mitigate this problem. Based on these observations and some new intuitions, we proposed an improved version of TANL called SDCT5 (step decomposed and constrained text-to-text Transformer). Preliminary experiment results show that our model can achieve better performance on SP tasks compared to TANL and benefit more from multi-task learning.","abstract_has_math":false,"creators":["Cheng, Xiang"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, Chengxiang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-09-17T01:11:17Z","date_published":"2021-09-17T01:11:17Z","updated_at":"2026-07-22T22:24:52Z","subjects":["natural language processing","structure prediction","multi-task learning"],"languages":["en"],"rights":["Copyright 2021 Xiang Cheng"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/110569","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, Chengxiang"]},{"key":"dc:creator","label":"Author","values":["Cheng, Xiang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-09-17T01:11:17Z","2021-04-27","2021-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["natural language processing","structure prediction","multi-task learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Xiang Cheng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/110569"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Structure prediction (SP) tasks are important in natural language understanding in the sense that they provide complex and structured knowledge of the text. Recently, some unified text-to-text transformer models like T5 and TANL have produced competitive results on SP tasks. These models convert SP tasks into a seq2seq problem, where a transformer is used to generate sequences with special tokens representing the extracted spans, labels, and relationships. Compared to many popular Natural Language Understanding models that are designed specifically for the task, the output of the text-to-text transformer is more flexible. With proper format, it could be trained on multiple tasks together and take advantage of the shared knowledge between tasks. To better understand how these models achieve better performance by multi-task learning, we designed several experiments to measure the knowledge transfer ability of a recently proposed model, TANL. In our experiments, we found that the multi-head attention in the decoder can capture the relationship between tasks which leads to performance improvement. Another finding is that TANL may produce many outputs with invalid format when trained from scratch, and starting from a T5 pre-trained model helps to mitigate this problem. Based on these observations and some new intuitions, we proposed an improved version of TANL called SDCT5 (step decomposed and constrained text-to-text Transformer). Preliminary experiment results show that our model can achieve better performance on SP tasks compared to TANL and benefit more from multi-task learning.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-09-16 without embargo terms","The student, Xiang Cheng, accepted the attached license on 2021-04-24 at 10:19.","The student, Xiang Cheng, submitted this Thesis for approval on 2021-04-24 at 10:28.","This Thesis was approved for publication on 2021-04-27 at 09:34.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16544 on 2021-09-16 at 16:47:31","Made available in DSpace on 2021-09-17T01:11:17Z (GMT). No. of bitstreams: 2 CHENG-THESIS-2021.pdf: 1242117 bytes, checksum: c62d59918b73c76eb5fc07f0b10311e2 (MD5) LICENSE.txt: 4208 bytes, checksum: d01d4fe5f48c513e9a750014751422ec (MD5) Previous issue date: 2021-04-27"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["A deeper look into multi-task learning ability of unified text-to-text transformer"]}]}],"canonical_facts":{"dc:contributor":["Zhai, Chengxiang"],"dc:creator":["Cheng, Xiang"],"dc:date":["2021-09-17T01:11:17Z","2021-04-27","2021-05"],"dc:description":["Structure prediction (SP) tasks are important in natural language understanding in the sense that they provide complex and structured knowledge of the text. Recently, some unified text-to-text transformer models like T5 and TANL have produced competitive results on SP tasks. These models convert SP tasks into a seq2seq problem, where a transformer is used to generate sequences with special tokens representing the extracted spans, labels, and relationships. Compared to many popular Natural Language Understanding models that are designed specifically for the task, the output of the text-to-text transformer is more flexible. With proper format, it could be trained on multiple tasks together and take advantage of the shared knowledge between tasks. To better understand how these models achieve better performance by multi-task learning, we designed several experiments to measure the knowledge transfer ability of a recently proposed model, TANL. In our experiments, we found that the multi-head attention in the decoder can capture the relationship between tasks which leads to performance improvement. Another finding is that TANL may produce many outputs with invalid format when trained from scratch, and starting from a T5 pre-trained model helps to mitigate this problem. Based on these observations and some new intuitions, we proposed an improved version of TANL called SDCT5 (step decomposed and constrained text-to-text Transformer). Preliminary experiment results show that our model can achieve better performance on SP tasks compared to TANL and benefit more from multi-task learning.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-09-16 without embargo terms","The student, Xiang Cheng, accepted the attached license on 2021-04-24 at 10:19.","The student, Xiang Cheng, submitted this Thesis for approval on 2021-04-24 at 10:28.","This Thesis was approved for publication on 2021-04-27 at 09:34.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16544 on 2021-09-16 at 16:47:31","Made available in DSpace on 2021-09-17T01:11:17Z (GMT). No. of bitstreams: 2 CHENG-THESIS-2021.pdf: 1242117 bytes, checksum: c62d59918b73c76eb5fc07f0b10311e2 (MD5) LICENSE.txt: 4208 bytes, checksum: d01d4fe5f48c513e9a750014751422ec (MD5) Previous issue date: 2021-04-27"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/110569"],"dc:language":["en"],"dc:rights":["Copyright 2021 Xiang Cheng"],"dc:subject":["natural language processing","structure prediction","multi-task learning"],"dc:title":["A deeper look into multi-task learning ability of unified text-to-text transformer"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}